{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.14"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"},{"sourceId":199703319,"sourceType":"kernelVersion"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Processing - Kaggle - RSNA 2024 Lumbar Spine Degenerative Classification\n\nThis notebook explores how to process the competition data for [Kaggle - RSNA 2024 Lumbar Spine Degenerative Classification](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification) competition\n\n> _The goal of this competition is to create models that can be used to aid in the detection and classification of degenerative spine conditions using lumbar spine MR images. Competitors will develop models that simulate a radiologist's performance in diagnosing spine conditions._\n\nAuthor: `@cyshin971`  \nDate: `10/6/2024`","metadata":{}},{"cell_type":"code","source":"version = '1.0'\nisLocal = False # Set to True when running in local environment\n\n# Due to a lack of disk space when running on Kaggle, only a subset of the data is processed\nkaggle_subset = 10","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:14.921736Z","iopub.execute_input":"2024-10-06T14:52:14.922141Z","iopub.status.idle":"2024-10-06T14:52:14.954952Z","shell.execute_reply.started":"2024-10-06T14:52:14.922097Z","shell.execute_reply":"2024-10-06T14:52:14.953861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🤔 | What this notebook contains","metadata":{}},{"cell_type":"markdown","source":"This notebook explores how to process the DICOM (MRI) competition data using the metadata provided in my previous notebook [Data Analysis RSNA 2024 LSDC (cyshin971)](https://www.kaggle.com/code/chiyoungshin/data-analysis-rsna-2024-lsdc-cyshin971) [6] and provides a template to pad, normalize, and scale MRI data \n\n#### Competition Data Processing\n- Padding images to make them square\n  - Most common ratio (height/width) of images for all series types is `1.0` [[6]](https://www.kaggle.com/code/chiyoungshin/data-analysis-rsna-2024-lsdc-cyshin971)\n- Resizing images using various interpolation methods (`nearest`, `linear`, `cubic`, `area`, `lanczos4`)\n- Normalizing images using various methods:\n  - `min_max` : min-max scaling\n  - `z_score` : z-score normalization (mean and standard deviation)\n  - Normalizing individual instances (`is_series = False`) OR across all instances in a series (`is_series = True`)\n  - Updating normalized location of labels (`norm_x`, `norm_y`)\n    - Dataframes (`test_process_params`, `train_process_params`) for the updated label locations after padding (`delta_x`, `delta_y`) and resizing (`alpha_x`, `alpha_y`) is provided to track the changes in label locations  \n- Scaling images:\n  - `uint8`: scaling image values from `0` ~ `255`\n  - `int8`: scaling image values from `-127` ~ `+127`\n#### Visualizing Processed Images\nCode to visualize processed images with correct annotation of labels\n\n#### Notebook Outputs\nDue to a lack of disk space when running on Kaggle, only a subset of the data is processed. Run the notebook in a local environment to process the entire dataset.\n> NOTE: Please ensure sufficient memory (`288 GB`) on disk prior to running the notebook in a local environment. Compressed zip folder requires `57 GB`\n\n`Version 1` of this notebook provides a dataset (`simple_dataset`) of processed images.  \nNotebook also outputs a json file `proc_info.json` that contains information on the specific methods/parameters used to process the data.\n\n> moniker: `simple_dataset`  \n> version: `processed_v1`  \n> size: `[512, 512]`  \n> interpolation: `area`  \n> is_series: `True`  \n> norm: `min_max`  \n> image_scaling: `None`  ","metadata":{}},{"cell_type":"markdown","source":"### 📜 | Version History","metadata":{}},{"cell_type":"markdown","source":"- `version 1.0` : Initial release version","metadata":{}},{"cell_type":"markdown","source":"### 📌 | References","metadata":{}},{"cell_type":"markdown","source":"1. [RSNA 2024 Lumbar Spine Degenrative Classification - Kaggle Competition][1]\n2. [Consolidated Train Data - Muhammad Tariq Pervez][2]\n3. [Anatomy & Image Visualization Overview-RSNA RAIDS - Abhinav Suri][3]\n4. [Lumbar RSNA 2024: Visualizing + EDA + Sub - Allegich][4]\n5. [RSNA Lumbar Spine Analysis - Satya][5]\n6. [Data Analysis RSNA 2024 LSDC (cyshin971)][6]\n\n[1]: https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification\n[2]: https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/505929\n[3]: https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids\n[4]: https://www.kaggle.com/code/allegich/lumbar-rsna-2024-visualizing-eda-sub\n[5]: https://www.kaggle.com/code/satyaprakashshukl/rsna-lumbar-spine-analysis\n[6]: https://www.kaggle.com/code/chiyoungshin/data-analysis-rsna-2024-lsdc-cyshin971","metadata":{}},{"cell_type":"markdown","source":"# 📚 | Import Libraries ","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\npd.set_option('display.precision', 3) # Set display precisions\nimport cv2\n\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nfrom matplotlib import animation, rc\n\nimport pydicom\nfrom tqdm.notebook import tqdm\nimport joblib","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:14.957018Z","iopub.execute_input":"2024-10-06T14:52:14.957459Z","iopub.status.idle":"2024-10-06T14:52:16.880426Z","shell.execute_reply.started":"2024-10-06T14:52:14.957418Z","shell.execute_reply":"2024-10-06T14:52:16.879402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sys\nprint(\"Python: \" + sys.version)\nprint(\"---------------------------\")\nprint(\"Pandas:\", pd.__version__)\nprint(\"Pydicom:\", pydicom.__version__)\nprint(\"Open-CV:\", cv2.__version__)\nprint(\"---------------------------\")","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:16.881906Z","iopub.execute_input":"2024-10-06T14:52:16.882407Z","iopub.status.idle":"2024-10-06T14:52:16.890430Z","shell.execute_reply.started":"2024-10-06T14:52:16.882358Z","shell.execute_reply":"2024-10-06T14:52:16.889321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🌐 | Global Variables","metadata":{}},{"cell_type":"markdown","source":"### ⚙️ | Configuration","metadata":{}},{"cell_type":"code","source":"class CFG:\n    color_map = plt.cm.bone # 'viridis', 'binary', 'CMRmap', plt.cm.bone\n    analysis_version = 'analysis_v2'\n    load_test_anlysis = True\n    \n    version_str = (version.split(\".\")[0] if version.split(\".\")[1] == '0' else version.replace(\".\", \"_\"))\n    process_version = f'processed_v{version_str}'\n    process_moniker = {'processed_v0_6': 'min_max_series',\n                       'processed_v1': 'simple_dataset'}\n    \n    def check_process_moniker():\n        if CFG.process_version not in CFG.process_moniker.keys():\n            return CFG.process_version\n        else:\n            return CFG.process_moniker[CFG.process_version]\n \nprint(f'Analysis Version: {CFG.analysis_version}')\nprint(f'Process Version: {CFG.check_process_moniker()}')\nprint(f'Load Test Analysis: {CFG.load_test_anlysis}')","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:16.892805Z","iopub.execute_input":"2024-10-06T14:52:16.893167Z","iopub.status.idle":"2024-10-06T14:52:16.908299Z","shell.execute_reply.started":"2024-10-06T14:52:16.893112Z","shell.execute_reply":"2024-10-06T14:52:16.907123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🔎 | Data Features","metadata":{}},{"cell_type":"code","source":"class CONDITION:\n    num_classes = 5\n    class_names = ['Spinal Canal Stenosis',\n                   'Left Neural Foraminal Narrowing','Right Neural Foraminal Narrowing',\n                   'Left Subarticular Stenosis','Right Subarticular Stenosis']\n    label2name = dict(enumerate(class_names))\n    name2label = {v: k for k, v in label2name.items()}\n    \n    short_names = ['SCS',\n                   'l-FN', 'r-FN',\n                   'l-SS', 'r-SS']\n    label2short = dict(enumerate(short_names))\n    short2label = {v: k for k, v in label2short.items()}\n    def short2name(short_name):\n        return CONDITION.label2name[CONDITION.short2label[short_name]]\n    def name2short(class_name):\n        return CONDITION.label2short[CONDITION.name2label[class_name]]\n    \n    sub_names = ['spinal_canal_stenosis',\n                 'left_neural_foraminal_narrowing','right_neural_foraminal_narrowing',\n                 'left_subarticular_stenosis','right_subarticular_stenosis']\n    label2sub = dict(enumerate(sub_names))\n    sub2label = {v: k for k, v in label2sub.items()}\n    def sub2name(sub_name):\n        return CONDITION.label2name[CONDITION.sub2label[sub_name]]\n    def name2sub(class_name):\n        return CONDITION.label2sub[CONDITION.name2label[class_name]]\n    \n    num_sup_classes = 3\n    sup_class_names = ['Mixed', 'Spinal Canal Stenosis', 'Neural Foraminal Narrowing', 'Subarticular Stenosis']\n    sup_label2name = dict(enumerate(sup_class_names))\n    sup_name2label = {v: k for k, v in sup_label2name.items()}\n    name2sup_label = {0: 1, 1: 2, 2: 2, 3: 3, 4: 3}\n    sup2name_label = {0: [], 1: [0], 2: [1, 2], 3: [3, 4]}\n    def name2sup(sub_label):\n        return CONDITION.name2sup_label[sub_label]\n    def sup2name(sup_label):\n        return CONDITION.sup2name_label[sup_label]","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:16.909495Z","iopub.execute_input":"2024-10-06T14:52:16.909928Z","iopub.status.idle":"2024-10-06T14:52:16.922415Z","shell.execute_reply.started":"2024-10-06T14:52:16.909868Z","shell.execute_reply":"2024-10-06T14:52:16.921180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class SERIES:\n    num_classes = 3\n    class_names = ['Sagittal T2/STIR', 'Sagittal T1', 'Axial T2']\n    label2name = dict(enumerate(class_names))\n    name2label = {v: k for k, v in label2name.items()}\n    \n    short_names = ['saggital_t2', 'saggital_t1', 'axial_t2']\n    label2short = dict(enumerate(short_names))\n    short2label = {v: k for k, v in enumerate(short_names)}\n    def short2name(short_name):\n        return SERIES.label2name[SERIES.short2label[short_name]]\n    def name2short(class_name):\n        return SERIES.label2short[SERIES.name2label[class_name]]","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:16.924154Z","iopub.execute_input":"2024-10-06T14:52:16.924595Z","iopub.status.idle":"2024-10-06T14:52:16.944759Z","shell.execute_reply.started":"2024-10-06T14:52:16.924549Z","shell.execute_reply":"2024-10-06T14:52:16.943671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class LEVEL:\n    num_classes = 5\n    class_names = ['L1/L2', 'L2/L3', 'L3/L4', 'L4/L5', 'L5/S1']\n    label2name = dict(enumerate(class_names))\n    name2label = {v: k for k, v in label2name.items()}\n    \n    sub_names = ['l1_l2', 'l2_l3', 'l3_l4', 'l4_l5', 'l5_s1']\n    label2sub = dict(enumerate(sub_names))\n    sub2label = {v: k for k, v in enumerate(sub_names)}\n    def sub2name(sub_name):\n        return LEVEL.label2name[LEVEL.sub2label[sub_name]]\n    def name2sub(class_name):\n        return LEVEL.label2sub[LEVEL.name2label[class_name]]","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:16.946042Z","iopub.execute_input":"2024-10-06T14:52:16.946422Z","iopub.status.idle":"2024-10-06T14:52:16.967089Z","shell.execute_reply.started":"2024-10-06T14:52:16.946382Z","shell.execute_reply":"2024-10-06T14:52:16.966017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class OUTPUT:\n    num_classes = 3\n    class_names = ['Normal/Mild', 'Moderate', 'Severe']\n    label2name = dict(enumerate(class_names))\n    name2label = {v: k for k, v in label2name.items()}\n    \n    colors = {'Normal/Mild': 'green',\n              'Moderate': 'yellow',\n              'Severe': 'red'}","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:16.968618Z","iopub.execute_input":"2024-10-06T14:52:16.969084Z","iopub.status.idle":"2024-10-06T14:52:16.983267Z","shell.execute_reply.started":"2024-10-06T14:52:16.969035Z","shell.execute_reply":"2024-10-06T14:52:16.982150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📁 | Dataset Path ","metadata":{}},{"cell_type":"code","source":"import os\n\nif isLocal: # if running locally or JupyterHub\n    BASE_PATH = \"../data\" # set data path to local (repository) (change for release rev)\n    ANALYSIS_PATH = f\"../data/analysis/{CFG.analysis_version}\"\n    OUTPUT_DIR = \"../data/processed_images\"\n    PROC_DIR = f\"{OUTPUT_DIR}/{CFG.process_version}\"\nelse: # set data path to Kaggle input\n    BASE_PATH = \"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification\"\n    # when running in Kaggle, make sure to check the version of the analysis input data\n    ANALYSIS_PATH = \"/kaggle/input/data-analysis-rsna-2024-lsdc-cyshin971\"\n    OUTPUT_DIR = \"/kaggle/working\"\n    PROC_DIR = f\"{OUTPUT_DIR}/{CFG.process_moniker[CFG.process_version]}\"","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:16.984797Z","iopub.execute_input":"2024-10-06T14:52:16.985244Z","iopub.status.idle":"2024-10-06T14:52:17.000728Z","shell.execute_reply.started":"2024-10-06T14:52:16.985195Z","shell.execute_reply":"2024-10-06T14:52:16.999568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📖 | Meta Data ","metadata":{}},{"cell_type":"markdown","source":"### 🤝 | Helper Functions","metadata":{}},{"cell_type":"code","source":"# Print the Edge Cases\ndef print_edge_cases(e_cases):\n    print(f'Edge Cases: {len(e_cases)}')\n    print(\"---------------------------\")\n    for key in e_cases.keys():\n        print(f'{key}: {len(e_cases[key])}')\n    print(\"---------------------------\")\n\n# Filter a DataFrame for a specific feature\ndef filter_df(df, filter, feature, verbose=False):\n    if verbose:\n        print(f'Original: {df.shape}')\n        display(df.head(3))\n    filt_df = df[df[feature].isin(filter)]\n    if verbose:\n        print(f'Filtered: {df.shape}')\n        display(df.head(3))\n    return filt_df\n\n# Find the study_id, series_id and instance number given the image file path\ndef decode_img_file_path(img_file_path):\n    study_id, series_id, instance = img_file_path.split('_')\n    return int(study_id), int(series_id), int(instance.split('.')[0])","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:17.006704Z","iopub.execute_input":"2024-10-06T14:52:17.007884Z","iopub.status.idle":"2024-10-06T14:52:17.018326Z","shell.execute_reply.started":"2024-10-06T14:52:17.007824Z","shell.execute_reply":"2024-10-06T14:52:17.016896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load Metadata","metadata":{}},{"cell_type":"markdown","source":"This notebook uses the metadata from [Data Analysis RSNA 2024 LSDC (cyshin971)](https://www.kaggle.com/code/chiyoungshin/data-analysis-rsna-2024-lsdc-cyshin971) [6]","metadata":{}},{"cell_type":"code","source":"# Create a DataFrame for the test data if not loading from the sample competition test data\nif not CFG.load_test_anlysis:\n    \n    import glob\n    import re\n\n    # Extract the last three numerical values from each file path of the train images\n    def extract_last_three_values(file_path):\n        # Split the path into parts\n        img_file_numbers = re.findall(r'\\d+', file_path)\n        # Extract the last three values before the file extension\n        last_3_values = img_file_numbers[-3:]\n        return last_3_values\n\n    def create_df(split='test'):\n        images = glob.glob(BASE_PATH + f'/{split}_images/*/*/*.dcm')\n        data = [{'study_id': int(values[0]), 'series_id': int(values[1]), 'instance_number': int(values[2]), 'image_file_path':str(values[0])+'_'+str(values[1])+'_'+str(values[2])+'.dcm'} for file_path in images for values in [extract_last_three_values(file_path)]]\n        df = pd.DataFrame(data)\n        df = df.sort_values(['study_id', 'series_id', 'instance_number']).reset_index(drop=True)\n        \n        return df","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:17.020221Z","iopub.execute_input":"2024-10-06T14:52:17.020623Z","iopub.status.idle":"2024-10-06T14:52:17.035415Z","shell.execute_reply.started":"2024-10-06T14:52:17.020584Z","shell.execute_reply":"2024-10-06T14:52:17.033999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Kaggle Data\nsample_submission_df = pd.read_csv(os.path.join(BASE_PATH, 'sample_submission.csv'))\n\n# Analyzed Train Data for Lumber Spine Degenerative\ntrain_df = pd.read_csv(os.path.join(ANALYSIS_PATH, 'analyzed_train.csv'))\ntrain_series_desc_df = pd.read_csv(os.path.join(ANALYSIS_PATH, 'analyzed_train_series_descriptions.csv'))\nlabeled_train_df = pd.read_csv(os.path.join(ANALYSIS_PATH, 'analyzed_labeled_train.csv'))\n\n# Load Test Data\nif CFG.load_test_anlysis: # Analyzed Test Data for Lumber Spine Degenerative\n    print(f'Loading Analyzed Test Data from {ANALYSIS_PATH}')\n    test_series_desc_df = pd.read_csv(os.path.join(ANALYSIS_PATH, 'analyzed_test_series_descriptions.csv'))\n    test_df = pd.read_csv(os.path.join(ANALYSIS_PATH, 'analyzed_test.csv'))\nelse: # Construct test_df from Kaggle Data\n    print(f'Creating Test Data from {BASE_PATH}')\n    test_series_desc_df = pd.read_csv(os.path.join(BASE_PATH, 'test_series_descriptions.csv'))\n    test_df = create_df('test')    \n    test_df = pd.merge(test_df, test_series_desc_df[['series_id', 'series_description']], on='series_id', how='left')\n\nprint(f'train_df shape: {train_df.shape}')\ndisplay(train_df.head(3))\nprint(f'labeled_train_df shape: {labeled_train_df.shape}')\ndisplay(labeled_train_df.head(3))\nprint(f'train_series_desc_df shape: {train_series_desc_df.shape}')\ndisplay(train_series_desc_df.head(3))\nprint(f'test_df shape: {test_df.shape}')\ndisplay(test_df.head(3))\nprint(f'test_series_desc_df shape: {test_series_desc_df.shape}')\ndisplay(test_series_desc_df.head(3))","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:17.037381Z","iopub.execute_input":"2024-10-06T14:52:17.037857Z","iopub.status.idle":"2024-10-06T14:52:18.058940Z","shell.execute_reply.started":"2024-10-06T14:52:17.037803Z","shell.execute_reply":"2024-10-06T14:52:18.057736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 📐 | Edge Cases","metadata":{}},{"cell_type":"code","source":"import pickle\n\ndef load_edge_cases(file_path):\n    with open(file_path, 'rb') as f:\n        edge_cases = pickle.load(f)\n    return edge_cases\n\ndef find_edge_series(edge_case, verbose=True):\n    series_ids = []\n    for img_file_path in edge_case:\n        _, series_id, _ = decode_img_file_path(img_file_path)\n        if series_id not in series_ids: series_ids.append(series_id)\n    if verbose: print(f'Number of Series: {len(series_ids)}')\n    return series_ids\n\n# Print the Edge Cases\ndef print_edge_cases(e_cases):\n    print(f'Edge Cases: {len(e_cases)}')\n    print(\"---------------------------\")\n    for key in e_cases.keys():\n        print(f'{key}: {len(e_cases[key])}')","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.060741Z","iopub.execute_input":"2024-10-06T14:52:18.061194Z","iopub.status.idle":"2024-10-06T14:52:18.071031Z","shell.execute_reply.started":"2024-10-06T14:52:18.061144Z","shell.execute_reply":"2024-10-06T14:52:18.069089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"edge_cases = load_edge_cases(os.path.join(ANALYSIS_PATH, 'edge_cases.pkl'))\nprint_edge_cases(edge_cases)","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.072728Z","iopub.execute_input":"2024-10-06T14:52:18.073209Z","iopub.status.idle":"2024-10-06T14:52:18.096844Z","shell.execute_reply.started":"2024-10-06T14:52:18.073161Z","shell.execute_reply":"2024-10-06T14:52:18.095809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🩻 | DICOM Files","metadata":{}},{"cell_type":"markdown","source":"### 🤝 | Helper Functions","metadata":{}},{"cell_type":"markdown","source":"Format derived from [Refined Train Data for Lumber Spine Degenerative](https://www.kaggle.com/datasets/tariqcp/train-data-for-rsna-2024-lsdc) [2]\n\n- Image file path (`image_file_path`): the MRI image file path given in the format found in complete_train.csv [[2]](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/505929) `{study_id}_{series_id}_{instance_number}.dcm`\n  > `4003253_702807833_8.dcm`\n\n- DICOM file path (`dcm_file_path`): the file path of the DICOM file in the kaggle dataset `{BASE_PATH}/train_images/{study_id}/{series_id}/{instance_number}.dcm`\n  > `/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/4003253/702807833/8.dcm`","metadata":{}},{"cell_type":"code","source":"# Find the study_id given the series_id\ndef find_study_id(series_id, split='train'):\n    if split == 'test':\n        if series_id not in test_series_desc_df['series_id'].values:\n            print(f\"series_id {series_id} not found in test_series_descriptions.csv\")\n            return -1\n        return test_series_desc_df[test_series_desc_df['series_id'] == series_id]['study_id'].values[0]\n    if series_id not in train_series_desc_df['series_id'].values:\n        print(f\"series_id {series_id} not found in train_series_descriptions.csv\")\n        return -1\n    return train_series_desc_df[train_series_desc_df['series_id'] == series_id]['study_id'].values[0]\n\n# Find the series description given the series_id\ndef find_series_desc(series_id, split='train'):\n    if split == 'test':\n        if series_id not in test_series_desc_df['series_id'].values:\n            print(f\"series_id {series_id} not found in test_series_descriptions.csv\")\n            return -1\n        return test_series_desc_df[test_series_desc_df['series_id'] == series_id]['series_description'].values[0]\n    if series_id not in train_series_desc_df['series_id'].values:\n        print(f\"series_id {series_id} not found in train_series_descriptions.csv\")\n        return -1\n    return train_series_desc_df[train_series_desc_df['series_id'] == series_id]['series_description'].values[0]\n\n# Find the image file path given the series_id and instance number\ndef find_img_file_path(series_id, instance, split='train'):\n    return str(find_study_id(series_id, split))+'_'+str(series_id)+'_'+str(instance)+'.dcm'\n\n# Convert the image file path to the DICOM file path\ndef img_to_dcm_file_path(img_file_path, split='train', processed=False):\n    if processed:\n        img_file_path = img_file_path.replace('.dcm', '.npy')\n        if split == 'test':\n            return os.path.join(PROC_DIR, 'test_images', img_file_path.replace('_', '/'))\n        return os.path.join(PROC_DIR, 'train_images', img_file_path.replace('_', '/'))\n    \n    if split == 'test':\n        return os.path.join(BASE_PATH, 'test_images', img_file_path.replace('_', '/'))\n    return os.path.join(BASE_PATH, 'train_images', img_file_path.replace('_', '/'))","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.099073Z","iopub.execute_input":"2024-10-06T14:52:18.099531Z","iopub.status.idle":"2024-10-06T14:52:18.114125Z","shell.execute_reply.started":"2024-10-06T14:52:18.099486Z","shell.execute_reply":"2024-10-06T14:52:18.112918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Loading Functions","metadata":{}},{"cell_type":"code","source":"# Load MRI image only\ndef load_mri_image(img_file_path, split='train', processed=False):\n    dcm_file_path = img_to_dcm_file_path(img_file_path, split=split, processed=processed)\n    if processed:\n        return np.load(dcm_file_path)\n    return pydicom.dcmread(dcm_file_path).pixel_array\n\n# Load DICOM file (image + metadata)\ndef load_dicom_file(img_file_path, split='train'):\n    dcm_file_path = img_to_dcm_file_path(img_file_path, split=split)\n    dcm_file = pydicom.dcmread(dcm_file_path)\n    \n    # Convert the DICOM dataset to a dictionary to easily access metadata\n    metadata = {tag: dcm_file[tag].value for tag in dcm_file.dir()}\n    tags = metadata.keys()\n    \n    return dcm_file.pixel_array, metadata, tags","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.115812Z","iopub.execute_input":"2024-10-06T14:52:18.116282Z","iopub.status.idle":"2024-10-06T14:52:18.132288Z","shell.execute_reply.started":"2024-10-06T14:52:18.116231Z","shell.execute_reply":"2024-10-06T14:52:18.130890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🖼️ | Data Visualization","metadata":{}},{"cell_type":"markdown","source":"### 🤝 | Helper Functions","metadata":{}},{"cell_type":"code","source":"print(labeled_train_df.columns.values)","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.133709Z","iopub.execute_input":"2024-10-06T14:52:18.134175Z","iopub.status.idle":"2024-10-06T14:52:18.144267Z","shell.execute_reply.started":"2024-10-06T14:52:18.134127Z","shell.execute_reply":"2024-10-06T14:52:18.143273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the label information for a given series_id and instance_number\ndef get_coord_info(series_id, instance_number, processed=False):\n    features = ['condition',\n                'x', 'y',\n                'output',\n                'level']\n    if instance_number in labeled_train_df.loc[labeled_train_df['series_id'] == series_id, 'instance_number'].values:\n        values = labeled_train_df.loc[(labeled_train_df['series_id'] == series_id) &\n                                      (labeled_train_df['instance_number'] == instance_number),\n                                      features].values.T\n        coord_info = {feat: value for feat, value in zip(features, values)}\n        if processed:\n            coord_info['x'] = labeled_train_df.loc[(labeled_train_df['series_id'] == series_id) &\n                                                    (labeled_train_df['instance_number'] == instance_number),\n                                                    'proc_x'].values\n            coord_info['y'] = labeled_train_df.loc[(labeled_train_df['series_id'] == series_id) &\n                                                    (labeled_train_df['instance_number'] == instance_number),\n                                                    'proc_y'].values\n    else: coord_info = {}\n    \n    return coord_info\n\n# Get the window parameters for a given img\ndef find_label_param(img, markersize, series_desc):\n    ms = img.shape[0] * markersize / 100\n    # Offset the annotation for the different series\n    offset = (-ms,0)\n    if SERIES.name2label[series_desc] == SERIES.name2label['Axial T2']: offset = (0,-ms)\n    return ms, offset","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.146147Z","iopub.execute_input":"2024-10-06T14:52:18.146547Z","iopub.status.idle":"2024-10-06T14:52:18.157565Z","shell.execute_reply.started":"2024-10-06T14:52:18.146510Z","shell.execute_reply":"2024-10-06T14:52:18.156236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Annotate the image with the label information\ndef plot_coord(ax, coord_info, ms, figsize, show_info=True, label_offset=(0,0)):\n    rectangles = []\n    labels = []\n    # if coord_info is not empty\n    if coord_info:\n        output_colors = [OUTPUT.colors[out] for out in coord_info['output']]\n        # plot the rectangles\n        for cond, x, y, output_color, level in zip(coord_info['condition'], coord_info['x'], coord_info['y'], output_colors, coord_info['level']):            \n            rectangle = patches.Rectangle((x - ms / 2, y - ms / 2), ms, ms,\n                                               linewidth=1, edgecolor=output_color, facecolor='none')\n            ax.add_patch(rectangle)\n            rectangles.append(rectangle)\n            if show_info:\n                scn = CONDITION.label2short[CONDITION.name2label[cond]]\n                label = ax.text(x+label_offset[0], y+label_offset[1], scn, fontsize=figsize*1.7, color='pink', ha='center', va='bottom', fontweight='bold')\n                labels.append(label)\n                label = ax.text(x+label_offset[0], y+label_offset[1], level, fontsize=figsize*1.6, color='white', ha='center', va='top')\n                labels.append(label)\n    return rectangles, labels","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.159026Z","iopub.execute_input":"2024-10-06T14:52:18.159432Z","iopub.status.idle":"2024-10-06T14:52:18.174048Z","shell.execute_reply.started":"2024-10-06T14:52:18.159392Z","shell.execute_reply":"2024-10-06T14:52:18.172959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 🩻 | DICOM plot","metadata":{}},{"cell_type":"markdown","source":"Function derived from [Anatomy & Image Visualization Overview-RSNA RAIDS - Abhinav Suri](https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids) [3]","metadata":{}},{"cell_type":"code","source":"# Create subplots for displaying images\ndef create_subplots(images, max_images_per_row=1, figsize=4, show_info=False):\n    # Calculate the number of rows needed\n    num_images = len(images)\n    num_rows = (num_images + max_images_per_row - 1) // max_images_per_row  # Ceiling division\n    # Adjust the number of images per row if there are fewer images than the default\n    if num_images < max_images_per_row: max_images_per_row = num_images\n    # Calculate the figure size\n    fs = (max_images_per_row * figsize, num_rows * figsize)\n    if show_info: fs = (max_images_per_row*figsize, num_rows*figsize*1.1)\n    # Create a subplot grid\n    fig, axes = plt.subplots(num_rows, max_images_per_row, figsize=fs)\n    # Flatten axes array for easier looping if there are multiple rows\n    if num_rows > 1 or max_images_per_row > 1: axes = axes.flatten()\n    else: axes = [axes]  # Make it iterable for consistency\n    # Turn off unused subplots\n    for idx in range(num_images, len(axes)):\n        axes[idx].axis('off')\n\n    return fig, axes\n\n# display any images\ndef display_images(images, title='', max_images_per_row=1, figsize=4, show_axis=False):\n    fig, axes = create_subplots(images, max_images_per_row=max_images_per_row, figsize=figsize)\n    \n    for i, img in enumerate(images):\n        ax = axes[i]\n        ax.imshow(img, cmap=CFG.color_map)\n        if not show_axis: ax.axis('off')\n    \n    if title!='': fig.suptitle(title, fontsize=16)\n    plt.tight_layout()\n    plt.show()\n\n# Plot the raw DICOM images\ndef plot_dicom_images(img_file_paths, title='', max_images_per_row=1, figsize=4, markersize=10, show_label=False, show_info=False, split='train', processed=False):   \n    images = []\n    series_ids = []\n    instance_numbers = []\n    coord_infos = []\n    series_descs = []\n    for img_file_path in img_file_paths:\n        images.append(load_mri_image(img_file_path, split=split, processed=processed))\n        _, series_id, instance_number = decode_img_file_path(img_file_path)\n        series_ids.append(series_id)\n        instance_numbers.append(instance_number)\n        coord_infos.append(get_coord_info(series_id, instance_number, processed=processed))\n        series_descs.append(find_series_desc(series_id, split=split))\n    \n    fig, axes = create_subplots(images, max_images_per_row=max_images_per_row, figsize=figsize, show_info=show_info)\n\n    for i, img in enumerate(images):\n        ax = axes[i]\n        \n        if show_label:\n            ms, offset = find_label_param(img, markersize, series_descs[i])\n            plot_coord(ax, coord_infos[i], ms, figsize, show_info, offset)\n\n        ax.title.set_fontsize(figsize*3)\n        ax.title.set_text(f'{series_descs[i]} {series_ids[i]} ({instance_numbers[i]})')\n        ax.imshow(img, cmap=CFG.color_map)\n        if not show_info: ax.axis('off')\n    \n    if title!='': fig.suptitle(title, fontsize=16)\n    fig.tight_layout(rect=[0, 0, 0.98, 0.98])\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.175660Z","iopub.execute_input":"2024-10-06T14:52:18.175999Z","iopub.status.idle":"2024-10-06T14:52:18.196092Z","shell.execute_reply.started":"2024-10-06T14:52:18.175965Z","shell.execute_reply":"2024-10-06T14:52:18.194963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 🧊 | 3D MRI Slides","metadata":{}},{"cell_type":"markdown","source":"Function derived from:\n- [Lumbar RSNA 2024: Visualizing + EDA + Sub - Allegich](https://www.kaggle.com/code/allegich/lumbar-rsna-2024-visualizing-eda-sub) [4]\n- [RSNA Lumbar Spine Analysis - Satya](https://www.kaggle.com/code/satyaprakashshukl/rsna-lumbar-spine-analysis) [5]","metadata":{}},{"cell_type":"code","source":"def plot_3D_slides(series_id, markersize=10, figsize=6, show_label=True, show_info=True, split='train', processed = False):\n    if split == 'train':\n        img_file_paths = train_df[train_df['series_id'] == series_id]['image_file_path'].values\n    else:  # split == 'test'\n        img_file_paths = test_df[test_df['series_id'] == series_id]['image_file_path'].values\n    series_desc = find_series_desc(series_id, split=split)\n\n    images = []\n    for img_file_path in img_file_paths:\n        img = load_mri_image(img_file_path, split=split, processed=processed)\n        if img.max() == 0: continue  # Skip empty images\n        \n        # Get the instance number from the image file path and find the coordinate info\n        if show_label:\n            instance_number = int(img_file_path.split('_')[2].split('.')[0])\n            coord_info = get_coord_info(series_id, instance_number, processed=processed)\n        else: coord_info = {}\n\n        if coord_info:\n            images.append((img, coord_info))  # Store the image with its coordinate info\n        else:\n            images.append((img, None))  # No valid coordinates, append None\n\n    rc('animation', html='jshtml')\n\n    def create_animation(ims):\n        fig, ax = plt.subplots(figsize=(figsize, figsize))\n        plt.axis('off')\n        \n        im = ax.imshow(ims[0][0], cmap=CFG.color_map)  # Display the first image\n        ms, offset = find_label_param(ims[0][0], markersize, series_desc)\n        rects, labels = plot_coord(ax, ims[0][1], ms, figsize, show_info=show_info, label_offset=offset) # Plot the first rectangles\n        text = plt.text(0.05, 0.05, f'Slide {1}', transform=fig.transFigure, fontsize=16, color='darkblue')\n\n        def animate_func(i):\n            im.set_array(ims[i][0])  # Update the image in the animation\n            text.set_text(f'Slide {i + 1}')\n\n            # Remove the previous rectangles and labels\n            for rect in rects:\n                rect.remove()\n            rects.clear()\n            for label in labels:\n                label.remove()\n            labels.clear()\n\n            # Plot the new rectangles and labels\n            ms, offset = find_label_param(ims[i][0], markersize, series_desc)\n            rectangles, labs = plot_coord(ax, ims[i][1], ms, figsize, show_info=show_info, label_offset=offset)\n            for rect in rectangles:\n                rects.append(rect)\n            for label in labs:\n                labels.append(label)\n\n            return [im, text] + rects  # Ensure all rectangles are returned to be updated\n\n        plt.title(f'{str(find_study_id(series_id))}, {str(series_desc)} ({str(series_id)})')\n        plt.close()\n\n        return animation.FuncAnimation(fig, animate_func, frames=len(ims), interval=1000 // 10)\n\n    return create_animation(images)","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.197377Z","iopub.execute_input":"2024-10-06T14:52:18.197721Z","iopub.status.idle":"2024-10-06T14:52:18.213313Z","shell.execute_reply.started":"2024-10-06T14:52:18.197685Z","shell.execute_reply":"2024-10-06T14:52:18.212175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🔁 | Data Processing","metadata":{}},{"cell_type":"code","source":"class PROC:\n    size = (512, 512)\n    is_series = True\n    norm = 'min_max' # 'min_max' 'z_score'\n    image_scaling = None\n    \n    inter_dict = {'nearest': cv2.INTER_NEAREST,\n                  'linear': cv2.INTER_LINEAR,\n                  'cubic': cv2.INTER_CUBIC,\n                  'area': cv2.INTER_AREA,\n                  'lanczos4': cv2.INTER_LANCZOS4}\n    \n    interpolation = 'area'","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.214640Z","iopub.execute_input":"2024-10-06T14:52:18.215172Z","iopub.status.idle":"2024-10-06T14:52:18.228146Z","shell.execute_reply.started":"2024-10-06T14:52:18.215132Z","shell.execute_reply":"2024-10-06T14:52:18.226906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Processed Information","metadata":{}},{"cell_type":"code","source":"import json\n\ndef create_proc_info():\n    proc_info = {}\n    if CFG.process_version in CFG.process_moniker.keys():\n        proc_info['moniker'] = CFG.process_moniker[CFG.process_version]\n    proc_info['version'] = CFG.process_version\n    \n    proc_info['size'] = PROC.size\n    proc_info['interpolation'] = PROC.interpolation\n    proc_info['is_series'] = PROC.is_series\n    proc_info['norm'] = PROC.norm\n    proc_info['image_scaling'] = PROC.image_scaling\n    \n    return proc_info\n\ndef save_proc_info(proc_info):\n    file_path = os.path.join(PROC_DIR, 'proc_info.json')\n    with open(file_path, 'w') as f:\n        json.dump(proc_info, f, indent=4)\n\ndef update_proc(proc_info):\n    PROC.size = proc_info['size']\n    PROC.interpolation = proc_info['interpolation']\n    PROC.is_series = proc_info['is_series']\n    PROC.norm = proc_info['norm']\n    PROC.image_scaling = proc_info['image_scaling']\n\ndef load_proc_info(file_path=None):\n    file_path = os.path.join(file_path, 'proc_info.json')\n    with open(file_path, 'r') as f:\n        proc_info = json.load(f)\n    return proc_info\n\ndef print_proc_info(proc_info):\n    for key, value in proc_info.items():\n        print(f'{key}: {value}')","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.229682Z","iopub.execute_input":"2024-10-06T14:52:18.230058Z","iopub.status.idle":"2024-10-06T14:52:18.241723Z","shell.execute_reply.started":"2024-10-06T14:52:18.230018Z","shell.execute_reply":"2024-10-06T14:52:18.240551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Padding","metadata":{}},{"cell_type":"code","source":"# Pad the image to make it square\ndef pad_image(img):\n    orig_shape = img.shape\n    if orig_shape[0] != orig_shape[1]:\n        # calculate the padding size\n        pad_size = np.abs(orig_shape[0] - orig_shape[1])\n        padding = (pad_size//2, pad_size//2)\n        if sum(orig_shape) % 2 == 1: padding = (pad_size//2, pad_size//2+1)\n        \n        # pad the image to make it square\n        if orig_shape[0] > orig_shape[1]: pad_width = ((0, 0), padding)\n        if orig_shape[1] > orig_shape[0]: pad_width = (padding, (0, 0))\n        pad_img = np.pad(img, pad_width, mode='constant', constant_values=0)\n        \n        # check if the image is square\n        if pad_img.shape[0] != pad_img.shape[1]: print('Error: Image is not square')\n        \n        return pad_img, pad_width\n    return img, ((0, 0), (0, 0))","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.242891Z","iopub.execute_input":"2024-10-06T14:52:18.243504Z","iopub.status.idle":"2024-10-06T14:52:18.257747Z","shell.execute_reply.started":"2024-10-06T14:52:18.243464Z","shell.execute_reply":"2024-10-06T14:52:18.256524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_samples = 1\n\nfiltered_train_df = train_df[train_df['ratio'] == min(train_df['ratio'])].copy()\nlow_ratio_fp=filtered_train_df['image_file_path'].sample(num_samples).values\nfiltered_train_df = train_df[train_df['ratio'] == max(train_df['ratio'])].copy()\nhigh_ratio_fp=filtered_train_df['image_file_path'].sample(num_samples).values\n\nimg_file_paths = np.array([low_ratio_fp[0], high_ratio_fp[0]])\nplot_dicom_images(img_file_paths, title='Original Image', max_images_per_row=2, show_info=True)\n\npadded_imgs = []\nfor img_file_path in img_file_paths:\n    img = load_mri_image(img_file_path)\n    padded_img, _ = pad_image(img)\n    padded_imgs.append(padded_img)\n\ndisplay_images(padded_imgs, title='Padded Image', max_images_per_row=2, show_axis=True)\n\ndel num_samples, filtered_train_df, low_ratio_fp, high_ratio_fp, img_file_paths, padded_imgs","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:18.259389Z","iopub.execute_input":"2024-10-06T14:52:18.260390Z","iopub.status.idle":"2024-10-06T14:52:19.689683Z","shell.execute_reply.started":"2024-10-06T14:52:18.260320Z","shell.execute_reply":"2024-10-06T14:52:19.688395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Normalization","metadata":{}},{"cell_type":"markdown","source":"### Scaling","metadata":{}},{"cell_type":"code","source":"def scale_image(img, dtype=None):\n    if dtype == 'uint8':\n        return (img * 255).astype(np.uint8)\n    if dtype == 'int8':\n        return (img * 127).astype(np.int8)\n    return img","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:19.691410Z","iopub.execute_input":"2024-10-06T14:52:19.691773Z","iopub.status.idle":"2024-10-06T14:52:19.697853Z","shell.execute_reply.started":"2024-10-06T14:52:19.691735Z","shell.execute_reply":"2024-10-06T14:52:19.696387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Normalization Methods","metadata":{}},{"cell_type":"markdown","source":"Normalization derived from:\n- [Lumbar RSNA 2024: Visualizing + EDA + Sub - Allegich](https://www.kaggle.com/code/allegich/lumbar-rsna-2024-visualizing-eda-sub) [4]\n- [RSNA Lumbar Spine Analysis - Satya](https://www.kaggle.com/code/satyaprakashshukl/rsna-lumbar-spine-analysis) [5]","metadata":{}},{"cell_type":"code","source":"# Normalize image using min-max scaling\ndef normalize_min_max(img, min_value, max_value):\n    n_img = (img - min_value)\n    if max_value - min_value != 0: n_img = n_img / (max_value - min_value)\n    return n_img\n\n# Find the mean and standard deviation of the non-zero pixels in the image\ndef stats_image(img):\n    nonzero_pixels = img[np.nonzero(img)] # TODO: Check if this is correct\n    if nonzero_pixels.size == 0:\n        mean = 0\n        std = 0\n    else:\n        mean = np.mean(nonzero_pixels)\n        std = np.std(nonzero_pixels)\n    return mean, std\n\n# Normalize image using z-score\ndef normalize_z_score(img, mean, std):\n    n_img = (img - mean)\n    if std != 0: n_img = n_img / std\n    return n_img","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:19.699624Z","iopub.execute_input":"2024-10-06T14:52:19.699975Z","iopub.status.idle":"2024-10-06T14:52:19.712179Z","shell.execute_reply.started":"2024-10-06T14:52:19.699937Z","shell.execute_reply":"2024-10-06T14:52:19.710938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Normalizing Series vs Instances","metadata":{}},{"cell_type":"markdown","source":"Normalization per Instance:\n- Advantages: Normalizing each instance separately ensures that the full dynamic range of each image is used, which can be beneficial if the intensity range of each instance varies significantly. This approach is often used in image processing tasks where each image needs to be treated independently.\n- Disadvantages: If there is a significant variation in intensity ranges between instances, this could lead to inconsistencies in the relative intensity values across the series.\n\nNormalization over the Whole Series:\n- Advantages: Normalizing over the entire series ensures consistency across all instances, preserving the relative intensity differences between instances. This is particularly important in quantitative analyses or when comparing instances within the same series.\n- Disadvantages: If some instances have a very narrow intensity range compared to others, they might not utilize the full dynamic range after normalization, potentially reducing the contrast in those images.\n\n> In many medical imaging pipelines, normalizing over the whole series is more common because it maintains the integrity of the intensity information across all images, which is crucial for most clinical and research purposes.","metadata":{}},{"cell_type":"code","source":"# Normalize the image using the specified method\ndef normalize_img(img, method='min_max', dtype=None):\n    if method == 'min_max':\n        img = normalize_min_max(img, np.min(img), np.max(img))\n    if method == 'z_score':\n        mean, std = stats_image(img)\n        img = normalize_z_score(img, mean, std)\n    if dtype == 'uint8':\n        img = scale_image(img, dtype)\n    return img\n\n# Normalize the instances individually using the specified method\ndef normalize_instances(instances, method='min_max', dtype=None):\n    normalized_instances = []\n    for img in instances:\n        normalized_instances.append(normalize_img(img, method, dtype))\n    return normalized_instances\n\n# Normalize the series using the specified method\ndef normalize_series(series, method='min_max', dtype=None, verbose=False):\n    normalized_series = []\n    concat_series = series[0]\n\n    for img in series[1:]:\n        concat_series = np.concatenate((concat_series, img), axis=1)\n\n    \n    if method == 'min_max':\n        min_value = np.min(concat_series)\n        max_value = np.max(concat_series)\n        if verbose: print(f'Min Value: {min_value}, Max Value: {max_value}')\n        for img in series:\n            n_img = normalize_min_max(img, min_value, max_value)\n            normalized_series.append(scale_image(n_img, dtype))\n    \n    if method == 'z_score':\n        mean, std = stats_image(concat_series)\n        for img in series:\n            n_img = normalize_z_score(img, mean, std)\n            normalized_series.append(scale_image(n_img, dtype))\n    \n    # TODO: Try histogram equalization\n    # if method == 'histogram':\n    #     return cv2.equalizeHist(series)\n    \n    return normalized_series\n\n# Normalize the images using the specified method\ndef normalize(imgs, is_series=False, method='min_max', dtype=None, verbose=False):\n    if is_series:\n        return normalize_series(imgs, method, dtype, verbose=verbose)\n    return normalize_instances(imgs, method, dtype)","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:19.718242Z","iopub.execute_input":"2024-10-06T14:52:19.718967Z","iopub.status.idle":"2024-10-06T14:52:19.730854Z","shell.execute_reply.started":"2024-10-06T14:52:19.718925Z","shell.execute_reply":"2024-10-06T14:52:19.729594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Process DICOM MRI Images","metadata":{}},{"cell_type":"code","source":"# Check if the images have been processed\ndef check_processed(img_file_paths_list, split=\"train\"):    \n    for img_file_paths in img_file_paths_list:\n        study_id, series_id, _ = decode_img_file_path(img_file_paths[0])\n        for img_file_path in img_file_paths:\n            _ , _ , instance_number = decode_img_file_path(img_file_path)\n            f_dir = (f'{PROC_DIR}/{split}_images/{str(study_id)}/{str(series_id)}')\n            f_name = (f'{instance_number}.npy')\n            if not os.path.exists(os.path.join(f_dir, f_name)):\n                print(f\"{img_file_path} does not exist\")\n                return False\n    return True\n\n# Process each file and create a dictionary of processed parameters\ndef process_img(img_file_paths, is_series=False, method='min_max', size=(320,320), interpolation=cv2.INTER_AREA, dtype=None, split='train', verbose=False):\n    series_params = []\n    study_id, series_id, _ = decode_img_file_path(img_file_paths[0])\n    series_path = f'{PROC_DIR}/{split}_images/{str(study_id)}/{str(series_id)}'\n    os.makedirs(series_path, exist_ok=True)\n    \n    proc_imgs = []\n    \n    for img_file_path in img_file_paths:\n        _ , _ , instance_number = decode_img_file_path(img_file_path)\n        process_params = {}\n        process_params['image_file_path'] = img_file_path\n        \n        # Pad the images\n        padded_img, pad_width = pad_image(load_mri_image(img_file_path, split=split))\n        process_params['delta_x'] = pad_width[1][0]\n        process_params['delta_y'] = pad_width[0][0]\n        \n        # Resize the images\n        resized_img = cv2.resize(padded_img, size, interpolation=interpolation)\n        process_params['alpha_x'] = size[1] / padded_img.shape[1]\n        process_params['alpha_y'] = size[0] / padded_img.shape[0]\n        \n        series_params.append(process_params)\n        proc_imgs.append(resized_img)\n        \n        # Normalize the images\n        if not is_series:\n            norm_img = normalize_instances(resized_img, method, dtype, verbose=verbose)\n        \n            # Save the processed images\n            np.save(f'{series_path}/{instance_number}', norm_img)\n            \n            return series_params\n\n    norm_series = normalize_series(proc_imgs, method, dtype, verbose=verbose)\n    \n    # Save the processed images\n    for i, img_file_path in enumerate(img_file_paths):\n        _ , _ , instance_number = decode_img_file_path(img_file_path)\n        \n        # Save the processed images\n        np.save(f'{series_path}/{instance_number}', norm_series[i])\n    \n    return series_params\n\n# Update the coordinate information with the processed parameters\ndef update_coord_info(labeled_df, process_params):\n    labeled_df = pd.merge(labeled_df, process_params, on='image_file_path', how='left')\n    \n    labeled_df['proc_x'] = (labeled_df['x'] + labeled_df['delta_x']) * labeled_df['alpha_x']\n    labeled_df['proc_y'] = (labeled_df['y'] + labeled_df['delta_y']) * labeled_df['alpha_y']\n    labeled_df.drop(columns=['alpha_x', 'alpha_y', 'delta_x', 'delta_y'], inplace=True)\n    \n    # Move output column to the end of the dataframe\n    columns = labeled_df.columns.tolist()\n    columns.remove('output')\n    columns.append('output')\n    labeled_df = labeled_df[columns]\n    \n    return labeled_df","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:19.732366Z","iopub.execute_input":"2024-10-06T14:52:19.732746Z","iopub.status.idle":"2024-10-06T14:52:19.749151Z","shell.execute_reply.started":"2024-10-06T14:52:19.732706Z","shell.execute_reply":"2024-10-06T14:52:19.747863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the processed directory\nos.makedirs(PROC_DIR, exist_ok=True)\n\n# Get the file paths for each series\ntrain_file_paths = [train_df[train_df['series_id'] == series_id]['image_file_path'].values for series_id in train_df['series_id'].unique()]\n\n# Due to a lack of disk space when running on Kaggle, only a subset of the data is processed\n# Run the notebook in a local environment to process the entire dataset\nif not isLocal:\n    train_file_paths = train_file_paths[:kaggle_subset]\n\n# Check if the data have been processed\nprint(f'Checking {PROC_DIR}/train_images')\nif not check_processed(train_file_paths,split=\"train\"):\n    proc_info = create_proc_info()\n    \n    # Parallelize the processing using joblib\n    train_series_params = joblib.Parallel(n_jobs=-1, backend=\"loky\")(\n        joblib.delayed(process_img)(img_file_paths,\n                                    is_series=PROC.is_series,\n                                    method=PROC.norm,\n                                    size=PROC.size,\n                                    interpolation=PROC.inter_dict[PROC.interpolation],\n                                    dtype=PROC.image_scaling,\n                                    split=\"train\")\n        for img_file_paths in tqdm(train_file_paths, total=len(train_file_paths))\n    )\n    print(f'{PROC_DIR} training images converted.')\n\n    # Save the processed parameters as csv and update the labeled train DataFrame\n    train_process_params = []\n    for series_params in train_series_params:\n        for img_params in series_params:\n            train_process_params.append(img_params)\n    del train_series_params\n    train_process_params = pd.DataFrame(train_process_params)\n    train_process_params.to_csv(os.path.join(PROC_DIR, 'train_process_params.csv'), index=False)\n\n    # Save the processed parameters as json\n    save_proc_info(proc_info)\n\nelse: # If the data have been processed\n    print(f'training images already converted. Check {PROC_DIR}/train_images')\n    \n    # Load the processing information\n    proc_info = load_proc_info(PROC_DIR)\n\n    # Load the processed parameters\n    train_process_params = pd.read_csv(os.path.join(PROC_DIR, 'train_process_params.csv'))\n\nif proc_info['version'] != CFG.process_version:\n    print('Processing information is outdated. Updating...')\n    update_proc(proc_info)\n\n# Update the coordinate information with the processed parameters\nlabeled_train_df = update_coord_info(labeled_train_df, train_process_params)\ndisplay(labeled_train_df.head(3))","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:19.751024Z","iopub.execute_input":"2024-10-06T14:52:19.751419Z","iopub.status.idle":"2024-10-06T14:52:26.214794Z","shell.execute_reply.started":"2024-10-06T14:52:19.751376Z","shell.execute_reply":"2024-10-06T14:52:26.213507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the file paths for each series\ntest_file_paths = [test_df[test_df['series_id'] == series_id]['image_file_path'].values for series_id in test_df['series_id'].unique()]\n\n# Check if the data have been processed\nprint(f'Checking {PROC_DIR}/test_images')\nif not check_processed(test_file_paths,split=\"test\"):\n    \n    # Parallelize the processing using joblib\n    test_series_params = joblib.Parallel(n_jobs=-1, backend=\"loky\")(\n        joblib.delayed(process_img)(img_file_paths,\n                                    is_series=PROC.is_series,\n                                    method=PROC.norm,\n                                    size=PROC.size,\n                                    interpolation=PROC.inter_dict[PROC.interpolation],\n                                    dtype=PROC.image_scaling,\n                                    split=\"test\")\n        for img_file_paths in tqdm(test_file_paths, total=len(test_file_paths))\n    )\n    print(f'{PROC_DIR} testing images converted.')\n\n    # Save the processed parameters as csv and update the labeled train DataFrame\n    test_process_params = []\n    for series_params in test_series_params:\n        for img_params in series_params:\n            test_process_params.append(img_params)\n    del test_series_params\n    test_process_params = pd.DataFrame(test_process_params)\n    test_process_params.to_csv(os.path.join(PROC_DIR, 'test_process_params.csv'), index=False)\n\nelse: # If the data have been processed\n    print(f'testing images already converted. Check {PROC_DIR}/test_images')\n\n    # Load the processed parameters\n    test_process_params = pd.read_csv(os.path.join(PROC_DIR, 'test_process_params.csv'))","metadata":{"execution":{"iopub.status.busy":"2024-10-06T14:52:26.216298Z","iopub.execute_input":"2024-10-06T14:52:26.216664Z","iopub.status.idle":"2024-10-06T14:52:27.608535Z","shell.execute_reply.started":"2024-10-06T14:52:26.216628Z","shell.execute_reply":"2024-10-06T14:52:27.607355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print_proc_info(proc_info)\n\nimg_file_paths=train_file_paths[1][3:6]\nplot_dicom_images(img_file_paths, title='DICOM Image', max_images_per_row=3, show_label=True, show_info=True)\nplot_dicom_images(img_file_paths, title='Processed Image', max_images_per_row=3, show_label=True, show_info=True, processed=True)\n\n_, s_id, _ = decode_img_file_path(img_file_paths[0])\n\ndel img_file_paths","metadata":{"execution":{"iopub.status.busy":"2024-10-06T15:00:20.341144Z","iopub.execute_input":"2024-10-06T15:00:20.341574Z","iopub.status.idle":"2024-10-06T15:00:22.376780Z","shell.execute_reply.started":"2024-10-06T15:00:20.341532Z","shell.execute_reply":"2024-10-06T15:00:22.375376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_3D_slides(s_id, processed=True)","metadata":{"execution":{"iopub.status.busy":"2024-10-06T15:00:32.449878Z","iopub.execute_input":"2024-10-06T15:00:32.450289Z","iopub.status.idle":"2024-10-06T15:00:34.929815Z","shell.execute_reply.started":"2024-10-06T15:00:32.450252Z","shell.execute_reply":"2024-10-06T15:00:34.928253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 💾 | Zip Processed Images","metadata":{}},{"cell_type":"markdown","source":"Set `isSave` to true to zip and save the dataset you created","metadata":{}},{"cell_type":"code","source":"isZip = True\n\nif isZip:\n    import shutil\n\n    output_dir = OUTPUT_DIR\n    output_folder_zip = f'rsna2024_lsdc_{CFG.check_process_moniker()}'\n\n    processed_images_path = os.path.join(output_dir, output_folder_zip)\n    print(f'Creating zip archive at {processed_images_path}.zip')\n\n    # Create a zip archive of PROC_DIR\n    shutil.make_archive(processed_images_path, 'zip', PROC_DIR)\n    print(f'Zip archive created at {processed_images_path}.zip')\n    \n    # List the files in the output folder\n    print(os.listdir(output_dir))","metadata":{"execution":{"iopub.status.busy":"2024-10-06T15:00:46.440327Z","iopub.execute_input":"2024-10-06T15:00:46.440760Z","iopub.status.idle":"2024-10-06T15:01:15.222555Z","shell.execute_reply.started":"2024-10-06T15:00:46.440718Z","shell.execute_reply":"2024-10-06T15:01:15.218869Z"},"trusted":true},"execution_count":null,"outputs":[]}]}