{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"dockerImageVersionId":30732,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## RSNA 2024 Lumbar Spine Degenerative Classification","metadata":{}},{"cell_type":"markdown","source":"## 1. Setup","metadata":{}},{"cell_type":"code","source":"import os\nfrom glob import glob\nfrom tqdm import tqdm\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport missingno as msno\nimport pydicom","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-07-24T07:39:55.703274Z","iopub.execute_input":"2024-07-24T07:39:55.703685Z","iopub.status.idle":"2024-07-24T07:39:57.237827Z","shell.execute_reply.started":"2024-07-24T07:39:55.703650Z","shell.execute_reply":"2024-07-24T07:39:57.235888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_rows', 100)\npd.set_option('display.max_columns', 100)\npd.set_option('display.width', 1000)\n\ncompetition_dataset_directory = Path('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification')","metadata":{"execution":{"iopub.status.busy":"2024-07-24T07:39:57.240369Z","iopub.execute_input":"2024-07-24T07:39:57.241128Z","iopub.status.idle":"2024-07-24T07:39:57.248791Z","shell.execute_reply.started":"2024-07-24T07:39:57.241075Z","shell.execute_reply":"2024-07-24T07:39:57.247413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_images = glob(str(competition_dataset_directory / 'train_images' / '*' / '*' / '*.dcm'))\n\ndf_train = pd.read_csv(competition_dataset_directory / 'train.csv')\ndf_train_series_descriptions = pd.read_csv(competition_dataset_directory / 'train_series_descriptions.csv')\ndf_train_label_coordinates = pd.read_csv(competition_dataset_directory / 'train_label_coordinates.csv')\n\nprint(f'Training Images Count {len(train_images)}')\nprint(f'Training Set Shape: {df_train.shape}')\nprint(f'Train Series Descriptions Shape: {df_train_series_descriptions.shape}')\nprint(f'Train Label Coordinates Shape: {df_train_label_coordinates.shape}')","metadata":{"execution":{"iopub.status.busy":"2024-07-24T07:39:57.250442Z","iopub.execute_input":"2024-07-24T07:39:57.250946Z","iopub.status.idle":"2024-07-24T07:40:44.216179Z","shell.execute_reply.started":"2024-07-24T07:39:57.250900Z","shell.execute_reply":"2024-07-24T07:40:44.214343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Overview\n\nThe problem is to classify five degenerative conditions of the lower spine. There are three core conditions: **spinal canal stenosis**, **neural foraminal narrowing**, and **subarticular stenosis**. The last two of them are considered for each side of the spine, so the targets are left neural foraminal narrowing, right neural foraminal narrowing, left subarticular stenosis, right subarticular stenosis, and spinal canal stenosis. For each imaging study in the dataset, there are three severity scores (Normal/Mild, Moderate, or Severe) given for each of these five conditions at the disc levels L1/L2, L2/L3, L3/L4, L4/L5, and L5/S1.\n\nThe disc levels refer to the specific regions of the lumbar spine where the intervertebral discs are located.\n\n* **L1/L2**: The disc between the first (L1) and second (L2) lumbar vertebrae\n* **L2/L3**: The disc between the second (L2) and third (L3) lumbar vertebrae\n* **L3/L4**: The disc between the third (L3) and fourth (L4) lumbar vertebrae\n* **L4/L5**: The disc between the fourth (L4) and fifth (L5) lumbar vertebrae\n* **L5/S1**: The disc between the fifth lumbar vertebra (L5) and the first sacral vertebra (S1)\n\nThese levels indicate where the spine issues are evaluated and classified. The anatomy of the spine can be seen below.\n\n![anatomy](https://i.ibb.co/pWr25Wh/Screenshot-2024-06-28-at-18-36-06.png)","metadata":{}},{"cell_type":"code","source":"display(df_train)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-24T07:40:44.218497Z","iopub.execute_input":"2024-07-24T07:40:44.218974Z","iopub.status.idle":"2024-07-24T07:40:44.271130Z","shell.execute_reply.started":"2024-07-24T07:40:44.218933Z","shell.execute_reply":"2024-07-24T07:40:44.269920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Conditions\n\nAs already mentioned, there are three conditions exist in the dataset. Their distributions among levels could be different, so they have to be analyzed together.","metadata":{}},{"cell_type":"code","source":"def visualize_condition_counts(df, title, path=None):\n    \n    \"\"\"\n    Visualize condition counts on training set\n    \n    Parameters\n    ----------\n    df: pandas.DataFrame\n        Counts and percentages of conditions\n        \n    title: str\n        Title of the plot\n        \n    path: str, pathlib.Path or None\n        Path of the output file (if path is None, plot is displayed with selected backend)\n    \"\"\"\n    \n    fig, ax = plt.subplots(figsize=(24, 16))\n\n    ax.barh(\n        y=np.arange(df.shape[0] // 3) - 0.2,\n        width=df['count'].values[0::3],\n        height=0.2,\n        align='center',\n        label='Normal/Mild'\n    )\n    ax.barh(\n        y=np.arange(df.shape[0] // 3),\n        width=df['count'].values[1::3],\n        height=0.2,\n        align='center',\n        label='Moderate'\n    )\n    ax.barh(\n        y=np.arange(df.shape[0] // 3) + 0.2,\n        width=df['count'].values[2::3],\n        height=0.2,\n        align='center',\n        label='Severe'\n    )\n\n    ax.set_yticks(np.arange(df.shape[0] // 3))\n    ax.set_yticklabels([\n        f'{level}\\nNormal Count: {normal_count} ({normal_percentage:.2f}%)\\nModerate Count: {moderate_count} ({moderate_percentage:.2f}%)\\nSevere Count: {severe_count} ({severe_percentage:.2f}%)' for level, normal_count, normal_percentage, moderate_count, moderate_percentage, severe_count, severe_percentage, in zip(\n            df['level'].values[0::3],\n            df['count'].values[0::3],\n            df['percentage'].values[0::3],\n            df['count'].values[1::3],\n            df['percentage'].values[1::3],\n            df['count'].values[2::3],\n            df['percentage'].values[2::3],\n        )\n    ])\n    ax.set_xlabel('')\n    ax.tick_params(axis='x', labelsize=17.5, pad=10)\n    ax.tick_params(axis='y', labelsize=17.5, pad=10)\n    ax.set_title(title, size=20, pad=15)\n    ax.legend(loc='best', prop={'size': 18})\n    plt.gca().invert_yaxis()\n\n    plt.show()\n\n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path)\n        plt.close(fig)\n","metadata":{"execution":{"iopub.status.busy":"2024-07-05T06:01:12.791804Z","iopub.execute_input":"2024-07-05T06:01:12.792348Z","iopub.status.idle":"2024-07-05T06:01:12.811334Z","shell.execute_reply.started":"2024-07-05T06:01:12.792298Z","shell.execute_reply":"2024-07-05T06:01:12.810022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.1. Spinal Canal Stenosis\n\nSpinal canal is the passageway that runs through each of the vertebrae and houses the spinal cord. When the space within the spinal canal becomes narrower, it can squeeze the spinal cord and the nerve roots that branch off from it. This narrowing can irritate, compress, or pinch the spinal cord or nerves, leading to back pain and other nerve-related issues such as sciatica. Various conditions and injuries can cause the spinal canal to become narrowed. Spinal stenosis can affect anyone, but it is most common in people over 50.\n\n![spinal_stenosis](https://i.ibb.co/hf82PDJ/17499-spinal-stenosis-02.jpg)\n\nThe condition typically affects two areas of the spine:\n* Lower back (lumbar spinal stenosis): The lumbar spine has five vertebrae in the lower back, labeled L1 to L5, which are the largest in the spine\n* Neck (cervical spinal stenosis): The cervical spine has seven vertebrae in the neck, labeled C1 to C7\n\nAlthough less common, the middle back (thoracic spine) can also be affected by spinal stenosis.","metadata":{}},{"cell_type":"code","source":"spinal_canal_stenosis_columns = [column for column in df_train.columns if column.startswith('spinal_canal_stenosis')]\ndf_train_spinal_canal_stenosis = []\n\nfor column in spinal_canal_stenosis_columns:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'spinal_canal_stenosis'})\n    df_train_spinal_canal_stenosis.append(df)\n    \ndf_train_spinal_canal_stenosis = pd.concat(df_train_spinal_canal_stenosis, axis=0).reset_index(drop=True)\n\ndf_train_spinal_canal_stenosis_counts = df_train_spinal_canal_stenosis.groupby('level').value_counts().reset_index()\ndf_train_spinal_canal_stenosis_counts['severity'] = df_train_spinal_canal_stenosis_counts['spinal_canal_stenosis'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_spinal_canal_stenosis_counts = df_train_spinal_canal_stenosis_counts.sort_values(by=['level', 'severity'], ascending=True)\ndf_train_spinal_canal_stenosis_counts['percentage'] = df_train_spinal_canal_stenosis_counts['count'] / df_train_spinal_canal_stenosis_counts.groupby('level')['count'].transform('sum') * 100","metadata":{"execution":{"iopub.status.busy":"2024-07-04T11:39:49.545487Z","iopub.execute_input":"2024-07-04T11:39:49.545929Z","iopub.status.idle":"2024-07-04T11:39:49.612926Z","shell.execute_reply.started":"2024-07-04T11:39:49.545896Z","shell.execute_reply":"2024-07-04T11:39:49.611528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Severity of spinal canal stenosis within different levels are visualized below. This condition is more common in lower levels since moderate and severe counts consistenly increase until reaching L5/S1 level. L5/S1 level has the lowest moderate and severe counts while L4/L5 is the only level that has more severe counts than moderate counts.","metadata":{}},{"cell_type":"code","source":"visualize_condition_counts(\n    df=df_train_spinal_canal_stenosis_counts,\n    title='Spinal Canal Stenosis Counts On Levels'\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-03T17:41:28.089426Z","iopub.execute_input":"2024-07-03T17:41:28.089776Z","iopub.status.idle":"2024-07-03T17:41:28.682552Z","shell.execute_reply.started":"2024-07-03T17:41:28.089747Z","shell.execute_reply":"2024-07-03T17:41:28.681315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.2. Neural Foraminal Narrowing\n\nForaminal stenosis is the narrowing that occurs in specific areas around the nerves that exit the spinal cord. It is a type of spinal stenosis affecting the neural foramen, which are openings on both sides of the spine. Foraminal stenosis puts pressure on the affected nerves, which can eventually impact the signals traveling through them, causing nerve pain and potentially permanent nerve damage.\n\n![neural_foraminal_stenosis](https://i.ibb.co/7SQMjTk/24856-foraminal-stenosis.jpg)\n\nA neural foramen is an opening where a spinal nerve exits the spine and branches out to other parts of the body. The size of the opening depends on its location in the spine. The location of the foraminal stenosis determines its type:\n\n* Cervical spine (neck): This is the second most common area for foraminal stenosis\n* Thoracic spine (upper and middle back)\n* Lumbar spine (lower back): This is the most common area for foraminal stenosis\n* Sacral spine (far lower back and pelvis)\n* Coccygeal spine (tailbone)\n\nForaminal stenosis appears to be common, especially in people over age 55, and becomes more likely as people age. Some studies suggest that up to 40% of people have at least moderate foraminal stenosis in their lumbar spine by age 60, increasing to about 75% in those aged 80 and older. However, most people with foraminal stenosis are unaware they have it, even if it is severe. Only 17.5% of people with severe foraminal stenosis experience symptoms.","metadata":{}},{"cell_type":"code","source":"neural_foraminal_narrowing_columns = [column for column in df_train.columns if 'neural_foraminal_narrowing' in column]\ndf_train_neural_foraminal_narrowing = []\n\nfor column in neural_foraminal_narrowing_columns:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'neural_foraminal_narrowing'})\n    df_train_neural_foraminal_narrowing.append(df)\n    \ndf_train_neural_foraminal_narrowing = pd.concat(df_train_neural_foraminal_narrowing, axis=0).reset_index(drop=True)\n\ndf_train_neural_foraminal_narrowing_counts = df_train_neural_foraminal_narrowing.groupby('level').value_counts().reset_index()\ndf_train_neural_foraminal_narrowing_counts['severity'] = df_train_neural_foraminal_narrowing_counts['neural_foraminal_narrowing'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_neural_foraminal_narrowing_counts = df_train_neural_foraminal_narrowing_counts.sort_values(by=['level', 'severity'], ascending=True)\ndf_train_neural_foraminal_narrowing_counts['percentage'] = df_train_neural_foraminal_narrowing_counts['count'] / df_train_neural_foraminal_narrowing_counts.groupby('level')['count'].transform('sum') * 100\n\nleft_neural_foraminal_narrowing_columns = [column for column in df_train.columns if column.startswith('left_neural_foraminal_narrowing')]\ndf_train_left_neural_foraminal_narrowing = []\n\nfor column in left_neural_foraminal_narrowing_columns:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'left_neural_foraminal_narrowing'})\n    df_train_left_neural_foraminal_narrowing.append(df)\n    \ndf_train_left_neural_foraminal_narrowing = pd.concat(df_train_left_neural_foraminal_narrowing, axis=0).reset_index(drop=True)\n\ndf_train_left_neural_foraminal_narrowing_counts = df_train_left_neural_foraminal_narrowing.groupby('level').value_counts().reset_index()\ndf_train_left_neural_foraminal_narrowing_counts['severity'] = df_train_left_neural_foraminal_narrowing_counts['left_neural_foraminal_narrowing'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_left_neural_foraminal_narrowing_counts = df_train_left_neural_foraminal_narrowing_counts.sort_values(by=['level', 'severity'], ascending=True)\ndf_train_left_neural_foraminal_narrowing_counts['percentage'] = df_train_left_neural_foraminal_narrowing_counts['count'] / df_train_left_neural_foraminal_narrowing_counts.groupby('level')['count'].transform('sum') * 100\n\nright_neural_foraminal_narrowing_columns = [column for column in df_train.columns if column.startswith('right_neural_foraminal_narrowing')]\ndf_train_right_neural_foraminal_narrowing = []\n\nfor column in right_neural_foraminal_narrowing_columns:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'right_neural_foraminal_narrowing'})\n    df_train_right_neural_foraminal_narrowing.append(df)\n    \ndf_train_right_neural_foraminal_narrowing = pd.concat(df_train_right_neural_foraminal_narrowing, axis=0).reset_index(drop=True)\n\ndf_train_right_neural_foraminal_narrowing_counts = df_train_right_neural_foraminal_narrowing.groupby('level').value_counts().reset_index()\ndf_train_right_neural_foraminal_narrowing_counts['severity'] = df_train_right_neural_foraminal_narrowing_counts['right_neural_foraminal_narrowing'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_right_neural_foraminal_narrowing_counts = df_train_right_neural_foraminal_narrowing_counts.sort_values(by=['level', 'severity'], ascending=True)\ndf_train_right_neural_foraminal_narrowing_counts['percentage'] = df_train_right_neural_foraminal_narrowing_counts['count'] / df_train_right_neural_foraminal_narrowing_counts.groupby('level')['count'].transform('sum') * 100","metadata":{"execution":{"iopub.status.busy":"2024-07-04T11:47:19.457781Z","iopub.execute_input":"2024-07-04T11:47:19.458211Z","iopub.status.idle":"2024-07-04T11:47:19.584410Z","shell.execute_reply.started":"2024-07-04T11:47:19.458181Z","shell.execute_reply":"2024-07-04T11:47:19.582935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Severity of neural foraminal narrowing within different levels are visualized below for left, right and both. This condition is also more common in lower levels. Moderate counts consistenly increase until reaching L5/S1 level and severe counts increase on L5/S1 as well. Severity counts of this condition are consistent between left and right.","metadata":{}},{"cell_type":"code","source":"visualize_condition_counts(\n    df=df_train_neural_foraminal_narrowing_counts,\n    title='Neural Foraminal Narrowing Counts On Levels'\n)\n\nvisualize_condition_counts(\n    df=df_train_left_neural_foraminal_narrowing_counts,\n    title='Left Neural Foraminal Narrowing Counts On Levels'\n)\n\nvisualize_condition_counts(\n    df=df_train_right_neural_foraminal_narrowing_counts,\n    title='Right Neural Foraminal Narrowing Counts On Levels'\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-03T17:41:46.047174Z","iopub.execute_input":"2024-07-03T17:41:46.047554Z","iopub.status.idle":"2024-07-03T17:41:47.848281Z","shell.execute_reply.started":"2024-07-03T17:41:46.047526Z","shell.execute_reply":"2024-07-03T17:41:47.847077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"neural_foraminal_narrowing_columns = [column for column in df_train.columns if 'neural_foraminal_narrowing' in column]\ndf_train_left_right_neural_foraminal_narrowing = []\n\nfor column in neural_foraminal_narrowing_columns[:5]:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'left_neural_foraminal_narrowing'})\n    df_train_left_right_neural_foraminal_narrowing.append(df)\n\nfor column in neural_foraminal_narrowing_columns[5:]:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'right_neural_foraminal_narrowing'})\n    df_train_left_right_neural_foraminal_narrowing.append(df)\n    \ndf_train_left_right_neural_foraminal_narrowing = pd.concat((\n    pd.concat(df_train_left_right_neural_foraminal_narrowing[:5], axis=0).reset_index(drop=True).iloc[:, :1],\n    pd.concat(df_train_left_right_neural_foraminal_narrowing[5:], axis=0).reset_index(drop=True)\n), axis=1, ignore_index=False)\n\ndf_train_left_right_neural_foraminal_narrowing['left_severity'] = df_train_left_right_neural_foraminal_narrowing['left_neural_foraminal_narrowing'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_left_right_neural_foraminal_narrowing['right_severity'] = df_train_left_right_neural_foraminal_narrowing['right_neural_foraminal_narrowing'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_left_right_neural_foraminal_narrowing['level_'] = df_train_left_right_neural_foraminal_narrowing['level'].map({\n    'l1_l2': 0,\n    'l2_l3': 1,\n    'l3_l4': 2,\n    'l4_l5': 3,\n    'l5_s1': 4\n})","metadata":{"execution":{"iopub.status.busy":"2024-07-04T11:55:20.777694Z","iopub.execute_input":"2024-07-04T11:55:20.778112Z","iopub.status.idle":"2024-07-04T11:55:20.825152Z","shell.execute_reply.started":"2024-07-04T11:55:20.778062Z","shell.execute_reply":"2024-07-04T11:55:20.823944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Severity of left and right neural foraminal narrowing have correlation coefficient of 0.437 which shows that they are somewhat correlated with each other.","metadata":{}},{"cell_type":"code","source":"df_neural_foraming_narrowing_correlations = df_train_left_right_neural_foraminal_narrowing[['level_', 'left_severity', 'right_severity']].corr()\ndisplay(df_neural_foraming_narrowing_correlations)","metadata":{"execution":{"iopub.status.busy":"2024-07-04T12:00:00.947226Z","iopub.execute_input":"2024-07-04T12:00:00.949044Z","iopub.status.idle":"2024-07-04T12:00:00.966459Z","shell.execute_reply.started":"2024-07-04T12:00:00.948984Z","shell.execute_reply":"2024-07-04T12:00:00.964712Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.3. Subarticular Stenosis\n\nSubarticular stenosis, also known as lateral recess stenosis, is a condition where the space beneath the facet joints in the spine, known as the lateral recess, becomes narrowed. This space is part of the spinal canal, and its narrowing can compress the spinal nerves as they travel through this area.\n\n![subarticular_stenosis](https://i.ibb.co/F4Nztxm/Lateral-Recess-Stenosis-scaled.jpg)","metadata":{}},{"cell_type":"code","source":"subarticular_stenosis_columns = [column for column in df_train.columns if 'subarticular_stenosis' in column]\ndf_train_subarticular_stenosis = []\n\nfor column in subarticular_stenosis_columns:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'subarticular_stenosis'})\n    df_train_subarticular_stenosis.append(df)\n    \ndf_train_subarticular_stenosis = pd.concat(df_train_subarticular_stenosis, axis=0).reset_index(drop=True)\n\ndf_train_subarticular_stenosis_counts = df_train_subarticular_stenosis.groupby('level').value_counts().reset_index()\ndf_train_subarticular_stenosis_counts['severity'] = df_train_subarticular_stenosis_counts['subarticular_stenosis'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_subarticular_stenosis_counts = df_train_subarticular_stenosis_counts.sort_values(by=['level', 'severity'], ascending=True)\ndf_train_subarticular_stenosis_counts['percentage'] = df_train_subarticular_stenosis_counts['count'] / df_train_subarticular_stenosis_counts.groupby('level')['count'].transform('sum') * 100\n\nleft_subarticular_stenosis_columns = [column for column in df_train.columns if column.startswith('left_subarticular_stenosis')]\ndf_train_left_subarticular_stenosis = []\n\nfor column in left_subarticular_stenosis_columns:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'left_subarticular_stenosis'})\n    df_train_left_subarticular_stenosis.append(df)\n    \ndf_train_left_subarticular_stenosis = pd.concat(df_train_left_subarticular_stenosis, axis=0).reset_index(drop=True)\n\ndf_train_left_subarticular_stenosis_counts = df_train_left_subarticular_stenosis.groupby('level').value_counts().reset_index()\ndf_train_left_subarticular_stenosis_counts['severity'] = df_train_left_subarticular_stenosis_counts['left_subarticular_stenosis'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_left_subarticular_stenosis_counts = df_train_left_subarticular_stenosis_counts.sort_values(by=['level', 'severity'], ascending=True)\ndf_train_left_subarticular_stenosis_counts['percentage'] = df_train_left_subarticular_stenosis_counts['count'] / df_train_left_subarticular_stenosis_counts.groupby('level')['count'].transform('sum') * 100\n\nright_subarticular_stenosis_columns = [column for column in df_train.columns if column.startswith('right_subarticular_stenosis')]\ndf_train_right_subarticular_stenosis = []\n\nfor column in right_subarticular_stenosis_columns:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'right_subarticular_stenosis'})\n    df_train_right_subarticular_stenosis.append(df)\n    \ndf_train_right_subarticular_stenosis = pd.concat(df_train_right_subarticular_stenosis, axis=0).reset_index(drop=True)\n\ndf_train_right_subarticular_stenosis_counts = df_train_right_subarticular_stenosis.groupby('level').value_counts().reset_index()\ndf_train_right_subarticular_stenosis_counts['severity'] = df_train_right_subarticular_stenosis_counts['right_subarticular_stenosis'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_right_subarticular_stenosis_counts = df_train_right_subarticular_stenosis_counts.sort_values(by=['level', 'severity'], ascending=True)\ndf_train_right_subarticular_stenosis_counts['percentage'] = df_train_right_subarticular_stenosis_counts['count'] / df_train_right_subarticular_stenosis_counts.groupby('level')['count'].transform('sum') * 100","metadata":{"execution":{"iopub.status.busy":"2024-07-04T12:05:59.895243Z","iopub.execute_input":"2024-07-04T12:05:59.895719Z","iopub.status.idle":"2024-07-04T12:06:00.016956Z","shell.execute_reply.started":"2024-07-04T12:05:59.895685Z","shell.execute_reply":"2024-07-04T12:06:00.015397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Severity of subarticular stenosis within different levels are visualized below for left, right and both. This condition is also more common in lower levels. However, moderate and severe counts are way higher compared to other two conditions. Severity counts of this condition are consistent between left and right.","metadata":{}},{"cell_type":"code","source":"visualize_condition_counts(\n    df=df_train_subarticular_stenosis_counts,\n    title='Subarticular Stenosis Counts On Levels'\n)\n\nvisualize_condition_counts(\n    df=df_train_left_subarticular_stenosis_counts,\n    title='Left Subarticular Stenosis Counts On Levels'\n)\n\nvisualize_condition_counts(\n    df=df_train_right_subarticular_stenosis_counts,\n    title='Right Subarticular Stenosis Counts On Levels'\n)","metadata":{"execution":{"iopub.status.busy":"2024-06-30T18:15:52.537745Z","iopub.execute_input":"2024-06-30T18:15:52.538167Z","iopub.status.idle":"2024-06-30T18:15:54.137966Z","shell.execute_reply.started":"2024-06-30T18:15:52.538134Z","shell.execute_reply":"2024-06-30T18:15:54.136595Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subarticular_stenosis_columns = [column for column in df_train.columns if 'subarticular_stenosis' in column]\ndf_train_left_right_subarticular_stenosis = []\n\nfor column in subarticular_stenosis_columns[:5]:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'left_subarticular_stenosis'})\n    df_train_left_right_subarticular_stenosis.append(df)\n\nfor column in subarticular_stenosis_columns[5:]:\n    df = df_train[[column]].copy(deep=True)\n    df['level'] = '_'.join(column.split('_')[-2:])\n    df = df.rename(columns={column: 'right_subarticular_stenosis'})\n    df_train_left_right_subarticular_stenosis.append(df)\n    \ndf_train_left_right_subarticular_stenosis = pd.concat((\n    pd.concat(df_train_left_right_subarticular_stenosis[:5], axis=0).reset_index(drop=True).iloc[:, :1],\n    pd.concat(df_train_left_right_subarticular_stenosis[5:], axis=0).reset_index(drop=True)\n), axis=1, ignore_index=False)\n\ndf_train_left_right_subarticular_stenosis['left_severity'] = df_train_left_right_subarticular_stenosis['left_subarticular_stenosis'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_left_right_subarticular_stenosis['right_severity'] = df_train_left_right_subarticular_stenosis['right_subarticular_stenosis'].map({\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n})\ndf_train_left_right_subarticular_stenosis['level_'] = df_train_left_right_subarticular_stenosis['level'].map({\n    'l1_l2': 0,\n    'l2_l3': 1,\n    'l3_l4': 2,\n    'l4_l5': 3,\n    'l5_s1': 4\n})","metadata":{"execution":{"iopub.status.busy":"2024-07-04T12:24:58.900451Z","iopub.execute_input":"2024-07-04T12:24:58.900900Z","iopub.status.idle":"2024-07-04T12:24:58.946896Z","shell.execute_reply.started":"2024-07-04T12:24:58.900868Z","shell.execute_reply":"2024-07-04T12:24:58.945712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Severity of left and right subarticular stenosis have correlation coefficient of 0.581 which shows that they are significantly correlated with each other.","metadata":{}},{"cell_type":"code","source":"df_subarticular_stenosis_correlations = df_train_left_right_subarticular_stenosis[['level_', 'left_severity', 'right_severity']].corr()\ndisplay(df_subarticular_stenosis_correlations)","metadata":{"execution":{"iopub.status.busy":"2024-07-04T12:30:17.984598Z","iopub.execute_input":"2024-07-04T12:30:17.985111Z","iopub.status.idle":"2024-07-04T12:30:18.006457Z","shell.execute_reply.started":"2024-07-04T12:30:17.985060Z","shell.execute_reply":"2024-07-04T12:30:18.003372Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.4. Key Differences\n\n![differences](https://i.ibb.co/GQBKgRG/stenosis.jpg)\n\n* **Lateral (Subarticular) Stenosis**:\n    * Narrowing happens in the lateral recess or subarticular space, which is part of the spinal canal but closer to the sides\n    * Specifically affects the area beneath the facet joints where nerve roots travel before exiting the spine\n    \n\n* **Spinal Canal (Central) Stenosis**:\n    * Narrowing occurs in the central spinal canal, where the spinal cord or cauda equina (in the lumbar spine) is located\n    * Affects the main passageway through which the spinal cord and nerves travel\n\n\n* **Foraminal Stenosis**:\n    * Narrowing occurs in the neural foramina, the openings on either side of the vertebrae through which nerve roots exit the spinal column\n    * Directly affects the spaces through which the nerves exit the spinal canal and extend to other parts of the body","metadata":{}},{"cell_type":"markdown","source":"## 4. Missing Values\n\nMissing values exist on all conditions and counts of them are listed below.\n\n* **spinal_canal_stenosis_l1_l2**: 1\n* **spinal_canal_stenosis_l2_l3**: 1\n* **spinal_canal_stenosis_l3_l4**: 1\n* **spinal_canal_stenosis_l4_l5**: 1\n* **spinal_canal_stenosis_l5_s1**: 1\n* **left_neural_foraminal_narrowing_l1_l2**: 2\n* **left_neural_foraminal_narrowing_l2_l3**: 2\n* **left_neural_foraminal_narrowing_l3_l4**: 2\n* **left_neural_foraminal_narrowing_l4_l5**: 2\n* **left_neural_foraminal_narrowing_l5_s1**: 2\n* **right_neural_foraminal_narrowing_l1_l2**: 8\n* **right_neural_foraminal_narrowing_l2_l3**: 8\n* **right_neural_foraminal_narrowing_l3_l4**: 8\n* **right_neural_foraminal_narrowing_l4_l5**: 8\n* **right_neural_foraminal_narrowing_l5_s1**: 8\n* **left_subarticular_stenosis_l1_l2**: 164\n* **left_subarticular_stenosis_l2_l3**: 82\n* **left_subarticular_stenosis_l3_l4**: 3\n* **left_subarticular_stenosis_l4_l5**: 3\n* **left_subarticular_stenosis_l5_s1**: 11\n* **right_subarticular_stenosis_l1_l2**: 161\n* **right_subarticular_stenosis_l2_l3**: 82\n* **right_subarticular_stenosis_l3_l4**: 2\n* **right_subarticular_stenosis_l4_l5**: 2\n* **right_subarticular_stenosis_l5_s1**: 7\n\nThere are very few missing values on spinal canal stenosis (1) and neural foraminal narrowing (left 2 and right 8). They are consistently missing on all levels so incomplete labels may not be related to lowest vertebrae being not visible on those studies.\n\nThere are more missing values on subarticular stenosis compared to other two conditions and, missing value counts vary between levels. They decrease along with levels so it's not related to lowest vertebrae being not visible either. Missing value counts are consistent between left and right for this condition.","metadata":{}},{"cell_type":"code","source":"msno.matrix(df_train, figsize=(32, 24))\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-04T12:42:02.874459Z","iopub.execute_input":"2024-07-04T12:42:02.875127Z","iopub.status.idle":"2024-07-04T12:42:04.553781Z","shell.execute_reply.started":"2024-07-04T12:42:02.875064Z","shell.execute_reply":"2024-07-04T12:42:04.552491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Dependencies of missing values can be seen more clearly on the dendrogram below. There are 8 groups that are formed:\n* Spinal canal stenosis missing values on all levels\n* Left neural foraminal narrowing missing values on all levels\n* Right neural foraminal narrowing missing values on all levels\n* Left and right subarticular stenosis missing values on l1_l2\n* Left and right subarticular stenosis missing values on l2_l3\n* Left and right subarticular stenosis missing values on l3_l4\n* Left and right subarticular stenosis missing values on l4_l5\n* Left and right subarticular stenosis missing values on l5_s1","metadata":{}},{"cell_type":"code","source":"msno.dendrogram(df_train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-07-03T18:49:37.886715Z","iopub.execute_input":"2024-07-03T18:49:37.887834Z","iopub.status.idle":"2024-07-03T18:49:38.770019Z","shell.execute_reply.started":"2024-07-03T18:49:37.887778Z","shell.execute_reply":"2024-07-03T18:49:38.768831Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Modalities and Planes\n\nMRI (Magnetic Resonance Imaging) modalities refer to the different techniques and sequences used during a scan to capture images of the body's internal structures. An MRI plane refers to the specific orientation or direction in which the images are taken during a scan.\n\nMRI imaging modalities and planes of series are provided in `train_series_descriptions.csv` file. Number of series in studies can be obtained from that data too.","metadata":{}},{"cell_type":"code","source":"train_series = glob(str(competition_dataset_directory / 'train_images' / '*' / '*'))\n\nprint(f'Train Series Descriptions Shape: {df_train_series_descriptions.shape}')\nprint(f'Train Series Count {len(train_series)}')\n\ndisplay(df_train_series_descriptions)","metadata":{"execution":{"iopub.status.busy":"2024-07-10T12:33:45.070232Z","iopub.execute_input":"2024-07-10T12:33:45.070593Z","iopub.status.idle":"2024-07-10T12:33:46.397808Z","shell.execute_reply.started":"2024-07-10T12:33:45.070561Z","shell.execute_reply":"2024-07-10T12:33:46.396412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_counts(df, value, title, path=None):\n    \n    \"\"\"\n    Visualize counts on training set\n    \n    Parameters\n    ----------\n    df: pandas.DataFrame\n        Counts and percentages of conditions\n        \n    title: str\n        Title of the plot\n        \n    path: str, pathlib.Path or None\n        Path of the output file (if path is None, plot is displayed with selected backend)\n    \"\"\"\n    \n    fig, ax = plt.subplots(figsize=(24, 16))\n\n    ax.barh(\n        y=np.arange(df.shape[0]),\n        width=df['count'].values,\n        align='center'\n    )\n\n    ax.set_yticks(np.arange(df.shape[0]))\n    ax.set_yticklabels([\n        f'{val} - Count: {count} ({percentage:.2f}%)' for val, count, percentage in zip(\n            df[value].values,\n            df['count'].values,\n            df['percentage'].values,\n        )\n    ])\n    ax.set_xlabel('')\n    ax.tick_params(axis='x', labelsize=17.5, pad=10)\n    ax.tick_params(axis='y', labelsize=17.5, pad=10)\n    ax.set_title(title, size=20, pad=15)\n    ax.legend(loc='best', prop={'size': 18})\n    plt.gca().invert_yaxis()\n\n    plt.show()\n\n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path)\n        plt.close(fig)\n","metadata":{"execution":{"iopub.status.busy":"2024-07-10T12:33:46.400815Z","iopub.execute_input":"2024-07-10T12:33:46.401222Z","iopub.status.idle":"2024-07-10T12:33:46.414283Z","shell.execute_reply.started":"2024-07-10T12:33:46.401190Z","shell.execute_reply":"2024-07-10T12:33:46.413072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_series_counts = df_train_series_descriptions.groupby('study_id')['series_id'].count().value_counts().reset_index().rename(columns={'series_id': 'series_count'})\ndf_train_series_counts['percentage'] = df_train_series_counts['count'] / df_train_series_counts['count'].sum(axis=0) * 100","metadata":{"_kg_hide-output":true,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-07-07T15:33:17.860401Z","iopub.execute_input":"2024-07-07T15:33:17.861073Z","iopub.status.idle":"2024-07-07T15:33:17.892736Z","shell.execute_reply.started":"2024-07-07T15:33:17.861027Z","shell.execute_reply":"2024-07-07T15:33:17.891283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Number of series in studies can be between 2 and 6, but more than 98% of the studies have 3 or 4 series. Other than that, there are 30 studies that have 5 series, 3 studies that have 2 series and 1 study that has 6 series. ","metadata":{}},{"cell_type":"code","source":"visualize_counts(df=df_train_series_counts, value='series_count', title='Series Counts in Studies')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-07T15:33:18.470507Z","iopub.execute_input":"2024-07-07T15:33:18.471494Z","iopub.status.idle":"2024-07-07T15:33:18.956077Z","shell.execute_reply.started":"2024-07-07T15:33:18.471448Z","shell.execute_reply":"2024-07-07T15:33:18.954905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_series_descriptions['plane'] = df_train_series_descriptions['series_description'].apply(lambda x: str(x).split()[0].strip())\ndf_train_series_descriptions['modality'] = df_train_series_descriptions['series_description'].apply(lambda x: str(x).split()[-1].strip())\n\ndf_train_series_description_counts = df_train_series_descriptions['series_description'].value_counts().reset_index()\ndf_train_series_description_counts['percentage'] = df_train_series_description_counts['count'] / df_train_series_description_counts['count'].sum(axis=0) * 100\n\ndf_train_series_plane_counts = df_train_series_descriptions['plane'].value_counts().reset_index()\ndf_train_series_plane_counts['percentage'] = df_train_series_plane_counts['count'] / df_train_series_plane_counts['count'].sum(axis=0) * 100\n\ndf_train_series_modality_counts = df_train_series_descriptions['modality'].value_counts().reset_index()\ndf_train_series_modality_counts['percentage'] = df_train_series_modality_counts['count'] / df_train_series_modality_counts['count'].sum(axis=0) * 100","metadata":{"execution":{"iopub.status.busy":"2024-07-06T18:40:03.077187Z","iopub.execute_input":"2024-07-06T18:40:03.077568Z","iopub.status.idle":"2024-07-06T18:40:03.103496Z","shell.execute_reply.started":"2024-07-06T18:40:03.077541Z","shell.execute_reply":"2024-07-06T18:40:03.102189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are three unique MRI modalities in the dataset and they are **T1-weighted imaging (T1)**, **T2-weighted imaging (T2)**, and **T2-weighted imaging with Short Tau Inversion Recovery (T2/STIR)**.\n\n&nbsp;\n\n#### 1. T1-Weighted Imaging (T1):\n\n**Characteristics**:\n* High-resolution anatomical detail: Provides clear images of the body's structure\n* Contrast: Good for visualizing fat and differentiating between different types of tissues\n\n**Uses**:\n* Anatomical details: Excellent for assessing normal anatomy and identifying structural abnormalities\n* Pathologies: Useful for detecting fat, subacute hemorrhage, and certain lesions\n* Post-contrast imaging: Often used with gadolinium contrast agents to highlight blood vessels and tumors\n\n**Appearance**:\n* Fat: Appears bright\n* Water: Appears dark\n* Muscle and other soft tissues: Intermediate signal intensity\n\n**Applications**:\n* Brain: Differentiating gray matter from white matter, identifying hemorrhages\n* Spine: Assessing vertebral bodies and disc spaces\n* Abdomen: Visualizing organs and detecting fatty liver disease\n\n&nbsp;\n\n#### 2. T2-Weighted Imaging (T2)\n\n**Characteristics**:\n* Fluid-sensitive: Highlights fluid-containing structures\n* Pathological contrast: Effective for detecting a wide range of pathologies involving water content\n\n**Uses**:\n* Pathologies: Excellent for identifying edema, inflammation, and tumors\n* Conditions: Commonly used to detect cysts, abscesses, and degenerative changes\n\n**Appearance**:\n* Water and fluid: Appear bright\n* Fat: Less bright compared to T1-weighted images\n* Muscle and other soft tissues: Intermediate to dark signal intensity\n\n**Applications**:\n* Brain: Detecting lesions, edema, and demyelinating diseases\n* Spine: Identifying disc herniations, spinal stenosis, and other degenerative conditions\n* Joints: Assessing joint effusions and cartilage integrity\n\n&nbsp;\n\n#### 3. T2-Weighted Imaging with Short Tau Inversion Recovery (T2/STIR)\n\n**Characteristics**:\n* Fat suppression: Specifically suppresses the signal from fat, making it easier to see fluid-related pathologies\n* Enhanced contrast: Improves the visibility of edema and inflammation\n\n**Uses**:\n* Pathologies: Especially useful for detecting edema, inflammation, and bone marrow abnormalities\n* Musculoskeletal imaging: Commonly used to evaluate muscles, tendons, ligaments, and bone marrow\n\n**Appearance**:\n* Water and fluid: Appear bright\n* Fat: Suppressed and appears dark\n* Muscle and other soft tissues: Intermediate signal intensity\n\n**Applications**:\n* Spine: Assessing bone marrow edema and soft tissue abnormalities\n* Musculoskeletal system: Evaluating injuries such as sprains, strains, and muscle tears\n* Oncology: Detecting tumors and assessing their extent\n\n&nbsp;\n\nThese MRI modalities provide complementary information, allowing for a comprehensive assessment of various medical conditions. They can be seen on the image below; T1 (A), T2 (B) and T2/STIR (C).\n\n![modalities](https://i.ibb.co/7GWDc0f/T1-weighted-A-T2-weighted-B-and-STIR-weighted-C-mid-sagittal-images-of-a-58-years.png)\n\n&nbsp;\n\nTheir counts and percentages within the dataset are visualized below. Since more than 98% of the studies have 3 or 4 series, modalities are somewhat equally distributed.","metadata":{}},{"cell_type":"code","source":"visualize_counts(df=df_train_series_modality_counts, value='modality', title='Modality Counts in Series')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-05T07:46:54.486952Z","iopub.execute_input":"2024-07-05T07:46:54.487466Z","iopub.status.idle":"2024-07-05T07:46:54.923270Z","shell.execute_reply.started":"2024-07-05T07:46:54.487414Z","shell.execute_reply":"2024-07-05T07:46:54.921732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are two unique planes in the dataset and they are **Sagittal**, and **Axial**. There is also another plane named **Coronal** but none of the MRIs are given in that plane.\n\n&nbsp;\n\n#### 1. Sagittal Plane:\n\n* **Description**: This plane divides the body into left and right sections\n* **Orientation**: Images are taken from the side, as if you are looking at the body from the side\n* **Common Uses**:\n    * Useful for viewing structures along the midline of the body, such as the spine, brain, and joints\n    * Helps in assessing the alignment and relationships of structures from front to back\n\n&nbsp;\n\n#### 2. Axial Plane:\n\n* **Description**: This plane divides the body into upper (superior) and lower (inferior) sections\n* **Orientation**: Images are taken as if you are looking up from the feet or down from the head, providing horizontal slices of the body\n* **Common Uses**:\n    * Commonly used to view the brain, spine, abdomen, and pelvis\n    * Allows for the examination of cross-sectional anatomy, making it easier to detect abnormalities such as tumors, lesions, or injuries\n\n&nbsp;\n\n#### 3. Coronal Plane:\n\n* **Description**: This plane divides the body into front (anterior) and back (posterior) sections\n* **Orientation**: Images are taken from the front or back, as if you are looking at the body straight on\n* **Common Uses**:\n    * Useful for visualizing the frontal view of organs and structures, such as the brain, heart, lungs, and abdominal organs \n    * Helps in assessing the symmetry and positioning of structures\n\n&nbsp;\n\nThose planes can be seen on the image below; Sagittal (A), Coronal (B) and Axial (C).\n\n![planes](https://i.ibb.co/P15mYbT/planes.png)","metadata":{}},{"cell_type":"code","source":"visualize_counts(df=df_train_series_plane_counts, value='plane', title='Plane Counts in Series')","metadata":{"execution":{"iopub.status.busy":"2024-07-05T07:07:37.817135Z","iopub.execute_input":"2024-07-05T07:07:37.817542Z","iopub.status.idle":"2024-07-05T07:07:38.258987Z","shell.execute_reply.started":"2024-07-05T07:07:37.817511Z","shell.execute_reply":"2024-07-05T07:07:38.257758Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, the counts and percentages of plane and modality combinations are visualized below. It is identical to counts and percentages of modalities.","metadata":{}},{"cell_type":"code","source":"visualize_counts(df=df_train_series_description_counts, value='series_description', title='Series Description Counts in Series')","metadata":{"execution":{"iopub.status.busy":"2024-07-05T08:08:58.576607Z","iopub.execute_input":"2024-07-05T08:08:58.577042Z","iopub.status.idle":"2024-07-05T08:08:59.021100Z","shell.execute_reply.started":"2024-07-05T08:08:58.577007Z","shell.execute_reply":"2024-07-05T08:08:59.019712Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Label Coordinates\n\nLabel coordinates are provided in `train_label_coordinates.csv` file. They are the x/y coordinates for the center of the area where the condition is observed.\n\nStudy level labels and series descriptions can be merged to label coordinates dataframe on id fields.","metadata":{}},{"cell_type":"code","source":"print(f'Train Label Coordinates Shape: {df_train_label_coordinates.shape}')\nprint(f'Annotated Studies {df_train_label_coordinates[\"study_id\"].nunique()}/{df_train.shape[0]}')\nprint(f'Annotated Series {df_train_label_coordinates[\"series_id\"].nunique()}/{df_train_series_descriptions.shape[0]}')\n\ndf_train_label_coordinates = df_train_label_coordinates.merge(\n    df_train,\n    on='study_id',\n    how='left'\n).merge(\n    df_train_series_descriptions,\n    on=['study_id', 'series_id'],\n    how='left'\n)\n\ndisplay(df_train_label_coordinates)","metadata":{"execution":{"iopub.status.busy":"2024-07-14T08:17:22.358400Z","iopub.execute_input":"2024-07-14T08:17:22.358837Z","iopub.status.idle":"2024-07-14T08:17:22.579492Z","shell.execute_reply.started":"2024-07-14T08:17:22.358803Z","shell.execute_reply":"2024-07-14T08:17:22.578072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All of the studies in the dataset have label coordinates except **3008676218**, which is the only study with missing spinal canal stenosis label. 3 series without label coordinates belong to that study.","metadata":{}},{"cell_type":"code","source":"display(df_train.loc[~df_train['study_id'].isin(df_train_label_coordinates['study_id'])])\ndisplay(df_train_series_descriptions.loc[~df_train_series_descriptions['series_id'].isin(df_train_label_coordinates['series_id'])])","metadata":{"execution":{"iopub.status.busy":"2024-07-09T18:33:11.902194Z","iopub.execute_input":"2024-07-09T18:33:11.902609Z","iopub.status.idle":"2024-07-09T18:33:11.936184Z","shell.execute_reply.started":"2024-07-09T18:33:11.902575Z","shell.execute_reply":"2024-07-09T18:33:11.934772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_label_coordinates_counts = df_train_label_coordinates.groupby('study_id')['study_id'].count().value_counts().reset_index().rename(columns={'study_id': 'label_coordinate_count'})\ndf_train_label_coordinates_counts['percentage'] = df_train_label_coordinates_counts['count'] / df_train_label_coordinates_counts['count'].sum(axis=0) * 100","metadata":{"execution":{"iopub.status.busy":"2024-07-09T18:35:51.811777Z","iopub.execute_input":"2024-07-09T18:35:51.812152Z","iopub.status.idle":"2024-07-09T18:35:51.824121Z","shell.execute_reply.started":"2024-07-09T18:35:51.812124Z","shell.execute_reply":"2024-07-09T18:35:51.822755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Number of label coordinate counts in studies can be between 15 and 25, but more than 90% of the studies have 25 label coordinates which is the exact same number as levels (5) multiplied with conditions (5). It shows that most of the studies have label coordinates for all conditions and levels even though the condition is Normal/Mild. In addition to that, number of label coordinate counts being less than 25 could be related to missing conditions.","metadata":{}},{"cell_type":"code","source":"visualize_counts(df=df_train_label_coordinates_counts, value='label_coordinate_count', title='Label Coordinate Counts In Studies')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-09T18:36:56.453856Z","iopub.execute_input":"2024-07-09T18:36:56.454254Z","iopub.status.idle":"2024-07-09T18:36:56.981789Z","shell.execute_reply.started":"2024-07-09T18:36:56.454225Z","shell.execute_reply":"2024-07-09T18:36:56.980634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_label_coordinates_counts = df_train_label_coordinates.groupby(['study_id', 'series_id'])['series_id'].count().value_counts().reset_index().rename(columns={'series_id': 'label_coordinate_count'})\ndf_train_label_coordinates_counts['percentage'] = df_train_label_coordinates_counts['count'] / df_train_label_coordinates_counts['count'].sum(axis=0) * 100","metadata":{"execution":{"iopub.status.busy":"2024-07-09T18:40:09.327828Z","iopub.execute_input":"2024-07-09T18:40:09.328265Z","iopub.status.idle":"2024-07-09T18:40:09.343403Z","shell.execute_reply.started":"2024-07-09T18:40:09.328232Z","shell.execute_reply":"2024-07-09T18:40:09.342122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Number of label coordinate counts in series is between 1 and 10 most of the time, but there is one study that has 15 label coordinates in one series. It shows that label coordinates of each condition are given only for specific modalities. It is very likely that some modalities are taken for certain conditions.","metadata":{}},{"cell_type":"code","source":"visualize_counts(df=df_train_label_coordinates_counts, value='label_coordinate_count', title='Label Coordinate Counts In Studies/Series')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-09T18:40:10.526059Z","iopub.execute_input":"2024-07-09T18:40:10.526517Z","iopub.status.idle":"2024-07-09T18:40:11.147889Z","shell.execute_reply.started":"2024-07-09T18:40:10.526481Z","shell.execute_reply":"2024-07-09T18:40:11.146561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Aggregations of label coordinate counts within condition groups show that some modalities are only taken for certain conditions.\n\n* Spinal canal stenosis conditions are almost always Sagittal T2/STIR, but there are 5 label coordinates (1 series) that are from Sagittal T1\n* Neural foraminal narrowing conditions are always Sagittal T1\n* Subarticular stenosis conditions are always Axial T2\n\nThose modalities and planes could be more effective for detecting the corresponding conditions however other series can be used for predicting the same condition as well.","metadata":{}},{"cell_type":"code","source":"display(df_train_label_coordinates.groupby('condition')['series_description'].value_counts().reset_index())","metadata":{"execution":{"iopub.status.busy":"2024-07-10T12:34:52.546897Z","iopub.execute_input":"2024-07-10T12:34:52.547318Z","iopub.status.idle":"2024-07-10T12:34:52.579362Z","shell.execute_reply.started":"2024-07-10T12:34:52.547280Z","shell.execute_reply":"2024-07-10T12:34:52.578059Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Locations of coordinate labels can be visualized as a heatmap with the function below.","metadata":{}},{"cell_type":"code","source":"def visualize_heatmap(heatmap, title, path=None):\n\n    \"\"\"\n    Visualize given heatmap\n\n    Parameters\n    ----------\n    heatmap: numpy.ndarray of shape (height, width)\n        Heatmap array\n\n    title: str\n        Title of the plot\n\n    path: str or None\n        Path of the output file or None (if path is None, plot is displayed with selected backend)\n    \"\"\"\n\n    fig, ax = plt.subplots(figsize=(8, 8))\n    ax.imshow(heatmap, cmap=plt.cm.hot)\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.tick_params(axis='x', labelsize=15, pad=10)\n    ax.tick_params(axis='y', labelsize=15, pad=10)\n    ax.set_title(title, size=15, pad=12.5, loc='center', wrap=True)\n\n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path)\n        plt.close(fig)\n","metadata":{"execution":{"iopub.status.busy":"2024-07-10T12:44:44.233897Z","iopub.execute_input":"2024-07-10T12:44:44.234382Z","iopub.status.idle":"2024-07-10T12:44:44.244350Z","shell.execute_reply.started":"2024-07-10T12:44:44.234346Z","shell.execute_reply":"2024-07-10T12:44:44.242870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_label_coordinates_sagittal_t1 = df_train_label_coordinates.loc[df_train_label_coordinates['series_description'] == 'Sagittal T1']\nsagittal_t1_heatmap = np.zeros((\n    int(df_train_label_coordinates_sagittal_t1['y'].max()) + 1,\n    int(df_train_label_coordinates_sagittal_t1['x'].max()) + 1,\n))\nfor idx, row in tqdm(df_train_label_coordinates_sagittal_t1.iterrows(), total=df_train_label_coordinates_sagittal_t1.shape[0]):\n    sagittal_t1_heatmap[\n        int(row['y']) - 10:int(row['y']) + 10,\n        int(row['x']) - 10:int(row['x']) + 10,\n        \n    ] += 1\n\ndf_train_label_coordinates_sagittal_t2 = df_train_label_coordinates.loc[df_train_label_coordinates['series_description'] == 'Sagittal T2/STIR']\nsagittal_t2_heatmap = np.zeros((\n    int(df_train_label_coordinates_sagittal_t2['y'].max()) + 1,\n    int(df_train_label_coordinates_sagittal_t2['x'].max()) + 1\n))\nfor idx, row in tqdm(df_train_label_coordinates_sagittal_t2.iterrows(), total=df_train_label_coordinates_sagittal_t2.shape[0]):\n    sagittal_t2_heatmap[\n        int(row['y']) - 10:int(row['y']) + 10,\n        int(row['x']) - 10:int(row['x']) + 10\n    ] += 1\n\ndf_train_label_coordinates_axial_t2 = df_train_label_coordinates.loc[df_train_label_coordinates['series_description'] == 'Axial T2']\naxial_t2_heatmap = np.zeros((\n    int(df_train_label_coordinates_axial_t2['y'].max()) + 1,\n    int(df_train_label_coordinates_axial_t2['x'].max()) + 1\n))\nfor idx, row in tqdm(df_train_label_coordinates_axial_t2.iterrows(), total=df_train_label_coordinates_axial_t2.shape[0]):\n    axial_t2_heatmap[\n        int(row['y']) - 10:int(row['y']) + 10,\n        int(row['x']) - 10:int(row['x']) + 10\n    ] += 1","metadata":{"execution":{"iopub.status.busy":"2024-07-10T12:53:16.149686Z","iopub.execute_input":"2024-07-10T12:53:16.150155Z","iopub.status.idle":"2024-07-10T12:53:21.269341Z","shell.execute_reply.started":"2024-07-10T12:53:16.150121Z","shell.execute_reply":"2024-07-10T12:53:21.268023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Heatmaps show that label coordinates are curving from top left to bottom right which makes sense for sagittal plane but not so much for axial plane. Label coordinates curving on axial plane may be related to them being raw, not normalized. There would be higher overlap if coordinates were normalized with height and width of the slices.","metadata":{}},{"cell_type":"code","source":"visualize_heatmap(sagittal_t1_heatmap, title='Sagittal T1 Label Coordinate Heatmap on All Series')\nvisualize_heatmap(sagittal_t2_heatmap, title='Sagittal T2/STIR Label Coordinate Heatmap on All Series')\nvisualize_heatmap(axial_t2_heatmap, title='Axial T2 Label Coordinate Heatmap on All Series')","metadata":{"execution":{"iopub.status.busy":"2024-07-10T12:55:11.928237Z","iopub.execute_input":"2024-07-10T12:55:11.929723Z","iopub.status.idle":"2024-07-10T12:55:13.394013Z","shell.execute_reply.started":"2024-07-10T12:55:11.929668Z","shell.execute_reply":"2024-07-10T12:55:13.392372Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7. DICOM Metadata\n\nDICOM (Digital Imaging and Communications in Medicine) is a standard protocol for managing, storing, and transmitting medical imaging information. DICOM files contain not only the image data but also a metadata that describes the image and its context. This metadata is organized into DICOM tags and attributes.\n\nDICOM attributes provide specific details about the image or related context. Attributes can be patient demographics, image acquisition parameters, study and series information, equipment details, etc. Attributes are extracted from DICOM files below.","metadata":{}},{"cell_type":"code","source":"metadata = []\ndicom_attributes = [\n    'StudyInstanceUID', 'SeriesInstanceUID', 'PatientID',\n    'ContentDate', 'ContentTime', 'SeriesDescription',\n    'SliceThickness', 'SpacingBetweenSlices', 'PatientPosition',\n    'InstanceNumber', 'ImagePositionPatient', 'ImageOrientationPatient',\n    'SliceLocation', 'SamplesPerPixel', 'PhotometricInterpretation',\n    'Rows', 'Columns', 'PixelSpacing', 'BitsAllocated', 'BitsStored',\n    'HighBit', 'PixelRepresentation', 'WindowCenter', 'WindowWidth',\n    'RescaleIntercept', 'RescaleSlope'\n]\n\ntrain_images_root_directory = competition_dataset_directory / 'train_images'\nstudy_ids = os.listdir(str(train_images_root_directory))\n\nfor study_id in tqdm(study_ids):\n\n    study_directory = train_images_root_directory / study_id\n    series_ids = os.listdir(str(study_directory))\n\n    for series_id in series_ids:\n\n        series_directory = study_directory / series_id\n        dicom_file_names = os.listdir(series_directory)\n\n        for dicom_file_name in dicom_file_names:\n\n            dicom_path = series_directory / dicom_file_name\n            dicom = pydicom.dcmread(dicom_path)\n\n            dicom_dict = {}\n            dicom_dict['slice_id'] = int(dicom_file_name.split('.')[0])\n            for tag in dicom_attributes:\n                try:\n                    dicom_tag_value = dicom[tag].value\n                except KeyError:\n                    dicom_tag_value = np.nan\n                dicom_dict[tag] = dicom_tag_value\n\n            metadata.append(dicom_dict)\n\n\ndf_metadata = pd.DataFrame(metadata)\ndf_metadata = df_metadata.sort_values(by=['StudyInstanceUID', 'SeriesInstanceUID', 'slice_id'], ascending=True).reset_index(drop=True)\ndf_metadata['StudyInstanceUID'] = df_metadata['StudyInstanceUID'].astype(int)\ndf_metadata['SeriesInstanceUID'] = df_metadata['SeriesInstanceUID'].apply(lambda x: int(str(x).split('.')[-1])).astype(int)\ndf_metadata['PatientID'] = df_metadata['PatientID'].astype(int)\ndf_metadata['ContentDate'] = pd.to_datetime(df_metadata['ContentDate'] + ' ' + df_metadata['ContentTime'].apply(lambda x: f'{x[0:2]}:{x[2:4]}'))\ndf_metadata = df_metadata.drop(columns=['PatientID', 'ContentTime'])\n\ndf_metadata = df_metadata.merge(\n    df_train_series_descriptions.rename(columns={\n        'study_id': 'StudyInstanceUID',\n        'series_id': 'SeriesInstanceUID'\n    }),\n    on=['StudyInstanceUID', 'SeriesInstanceUID'],\n    how='left'\n)","metadata":{"execution":{"iopub.status.busy":"2024-07-24T07:44:05.347369Z","iopub.execute_input":"2024-07-24T07:44:05.347878Z","iopub.status.idle":"2024-07-24T08:07:46.631130Z","shell.execute_reply.started":"2024-07-24T07:44:05.347815Z","shell.execute_reply":"2024-07-24T08:07:46.629645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_continuous_metadata_column_histogram(df, title, path=None):\n\n    \"\"\"\n    Visualize histogram of the given continuous column within modalities\n\n    Parameters\n    ----------\n    df: pandas.DataFrame\n        Dataframe with given continuous column and series description as index\n\n    title: str\n        Title of the plot\n\n    path: str, pathlib.Path or None\n        Path of the output file or None (if path is None, plot is displayed with selected backend)\n    \"\"\"\n\n    fig, ax = plt.subplots(figsize=(32, 8), dpi=100)\n    _, bins, _ = ax.hist(df.loc['Axial T2'], bins=24, label='Axial T2', alpha=0.5)\n    ax.hist(df.loc['Sagittal T1'], bins=bins, label='Sagittal T1', alpha=0.5)\n    ax.hist(df.loc['Sagittal T2/STIR'], bins=bins, label='Sagittal T2/STIR', alpha=0.5)\n    ax.tick_params(axis='x', labelsize=15)\n    ax.tick_params(axis='y', labelsize=15)\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.legend(prop={'size': 18})\n    ax.set_title(\n        title + f'''\n        Axial T2 Mean: {np.mean(df.loc['Axial T2']):.2f} Std: {np.std(df.loc['Axial T2']):.2f} Min: {np.min(df.loc['Axial T2']):.2f} Max: {np.max(df.loc['Axial T2']):.2f}\n        Sagittal T1 Mean: {np.mean(df.loc['Sagittal T1']):.2f} Std: {np.std(df.loc['Sagittal T1']):.2f} Min: {np.min(df.loc['Sagittal T1']):.2f} Max: {np.max(df.loc['Sagittal T1']):.2f}\n        Sagittal T2/STIR Mean: {np.mean(df.loc['Sagittal T2/STIR']):.2f} Std: {np.std(df.loc['Sagittal T2/STIR']):.2f} Min: {np.min(df.loc['Sagittal T2/STIR']):.2f} Max: {np.max(df.loc['Sagittal T2/STIR']):.2f}\n        ''',\n        size=15,\n        pad=12.5,\n        loc='center',\n        wrap=True\n    )\n\n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path, bbox_inches='tight')\n        plt.close(fig)","metadata":{"execution":{"iopub.status.busy":"2024-07-24T08:18:28.740151Z","iopub.execute_input":"2024-07-24T08:18:28.740613Z","iopub.status.idle":"2024-07-24T08:18:28.790791Z","shell.execute_reply.started":"2024-07-24T08:18:28.740580Z","shell.execute_reply":"2024-07-24T08:18:28.789440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Histograms of DICOM attributes are visualized for different modalities. Visualizations show that distributions of Sagittal T1 and Sagittal T2/STIR appear to be more similar to each other while distributions of Axial T2 are somewhat different.","metadata":{}},{"cell_type":"code","source":"df_metadata['PixelSpacingX'] = df_metadata['PixelSpacing'].apply(lambda x: float(x[0]))\ndf_metadata['PixelSpacingY'] = df_metadata['PixelSpacing'].apply(lambda x: float(x[1]))\n\ndf_slice_counts = df_metadata.groupby(['StudyInstanceUID', 'series_description'])['series_description'].count().reset_index(level=[0], drop=True).sort_index()\n\nvisualize_continuous_metadata_column_histogram(\n    df=df_slice_counts,\n    title='Series Slice Counts Distribution'\n)\n\ncontinuous_columns = [\n    'SliceThickness', 'SpacingBetweenSlices',\n    'SamplesPerPixel', 'Rows', 'Columns', 'PixelSpacingX', 'PixelSpacingY',\n    'BitsAllocated', 'BitsStored', 'HighBit',\n    'PixelRepresentation',\n    'WindowCenter', 'WindowWidth', 'RescaleIntercept', 'RescaleSlope'\n]\ndf_metadata_continuous_columns = df_metadata.groupby(['StudyInstanceUID', 'series_description'])[continuous_columns].first().reset_index(level=[0], drop=True).sort_index()\nfor column in continuous_columns:\n    visualize_continuous_metadata_column_histogram(\n        df=df_metadata_continuous_columns[column],\n        title=f'Series {column} Distribution'\n    )","metadata":{"execution":{"iopub.status.busy":"2024-07-24T10:37:13.645622Z","iopub.execute_input":"2024-07-24T10:37:13.646929Z","iopub.status.idle":"2024-07-24T10:37:34.240656Z","shell.execute_reply.started":"2024-07-24T10:37:13.646881Z","shell.execute_reply":"2024-07-24T10:37:34.239314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8. MRIs","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}