{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <div  style=\"color:#ff0a6c; border:#0014ff solid;  font-weight:bold; font-size:120%; text-align:center;padding:12.0px; background:#000000\">1. OVERVIEW</div>","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"## Goal","metadata":{}},{"cell_type":"markdown","source":"\"The goal of this competition is to create models that can be used to aid in the detection and classification of degenerative spine conditions using lumbar spine MR images. Competitors will develop models that simulate a radiologist's performance in diagnosing spine conditions.\"","metadata":{}},{"cell_type":"markdown","source":"\n#   Metric overview\n","metadata":{}},{"cell_type":"markdown","source":"We need to predict the **probability of stenosis** for each of the five vertebrae denoted by L1/L2,...  as well as an overall probability of any stenosis in the lumbra spine.  \n\n\"Submissions are evaluated using the average of sample **weighted log losses** and an any_severe_spinal prediction generated by the metric.\"\"\n\nThe sample **weights** are as follows:\n\n- \"1\" - for normal/mild\n- \"2\" - for moderate\n- \"4\" - for severe\n","metadata":{}},{"cell_type":"markdown","source":"#  Anatomical overview","metadata":{}},{"cell_type":"markdown","source":"<center>\n<img src=\"https://i.postimg.cc/0jvKk9qL/maxresdefault.jpg\" width=500>\n</center>","metadata":{}},{"cell_type":"markdown","source":"\n<center>\n<img src=\"https://i.postimg.cc/1tG4YbNM/2187d1a2c4675520a5969d1bf0478d8c.jpg\" width=300>\n</center>","metadata":{}},{"cell_type":"markdown","source":"<center>\n<img src=\"https://i.postimg.cc/hGKwN5vM/63d518b29d0dc6c44ee20d955f90b59c.jpg\" width=600>\n</center>","metadata":{}},{"cell_type":"markdown","source":"# MRI overview","metadata":{}},{"cell_type":"markdown","source":"### MR-slices","metadata":{}},{"cell_type":"markdown","source":"What does **Axial** and **Sagittal** mean? Let's try to understand that:\n\n- **Axial T2** - refers to an axial (cross-sectional) T2-weighted MRI sequence of the spine. T2-weighted images are useful for detecting pathologies like edema, inflammation, or lesions which appear hyperintense (brighter) compared to normal tissues\n\n- **Sagittal T1** - refers to a sagittal (lengthwise) T1-weighted MRI sequence of the spine. T1-weighted images provide good anatomical detail and contrast between different **soft tissues**.\n\n- **Sagittal T2/STIR** refers to either: A sagittal T2-weighted sequence, which is sensitive to detecting lesions, edema, and other pathologies that appear hyperintense. A sagittal STIR (**Short Tau Inversion Recovery**) sequence, which is a **fat-suppressed** T2-weighted sequence that helps highlight lesions and edema by suppressing the bright fat signal\n\nI'll try to make it more clear the difference between the this slices:","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"\n<center>\n<img src=\"https://i.postimg.cc/jqBkh94B/CT-Image-Planes.jpg\" width=500>\n</center>\n","metadata":{}},{"cell_type":"markdown","source":"### Difference between T1 and T2","metadata":{}},{"cell_type":"markdown","source":"<center>\n<img src=\"https://i.postimg.cc/GmFxw3sn/Diagram-shows-the-signal-intensity-of-various-tissues-at-T1-and-T2-weighted-imaging.png\" width=700>\n</center>","metadata":{}},{"cell_type":"markdown","source":"\n### Real case of stenosis and MRI","metadata":{}},{"cell_type":"markdown","source":"<center>\n<img src=\"https://i.postimg.cc/9fp3MjKv/7905-241447-2x.jpg\" width=900>\n</center>","metadata":{}},{"cell_type":"markdown","source":"## Dataset overview","metadata":{}},{"cell_type":"markdown","source":"The dataset we are using is made up of roughly 2000 MR studies. Spine radiology specialists have provided annotations to indicate the presence, vertebral level and location of any lumbar spine stenosis.","metadata":{}},{"cell_type":"markdown","source":"## Code requirements\nThis is a **code competition**, which means that submissions are made through notebooks. Furthermore, the submission notebook is subject to these conditions:\n- run-time (CPU/GPU) <= 9 hours\n- the submission file must be named submission.csv\n- internet access disabled\n- The test set is hidden, and will populated when you submit your notebook\n","metadata":{}},{"cell_type":"markdown","source":"# <div  style=\"color:#ff0a6c; border:#0014ff solid;  font-weight:bold; font-size:120%; text-align:center;padding:12.0px; background:#000000\">2. PREPARATION</div>\n\n","metadata":{}},{"cell_type":"code","source":"# packages\n\n# standard\nimport numpy as np\nimport pandas as pd\n\nimport os\nimport time\n\n# plots\nimport matplotlib.pyplot as plt\nfrom matplotlib import animation, rc\nimport plotly.express as px\nimport seaborn as sns\nimport plotly.express as px\nimport cv2\n\nimport pydicom as dicom # dicom\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\nimport warnings # warning handling\nwarnings.filterwarnings('ignore')\nimport glob\nimport json\nimport collections\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-09-18T06:00:21.618929Z","iopub.execute_input":"2024-09-18T06:00:21.619466Z","iopub.status.idle":"2024-09-18T06:00:24.611265Z","shell.execute_reply.started":"2024-09-18T06:00:21.619440Z","shell.execute_reply":"2024-09-18T06:00:24.610327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# show files\n!ls -l '../input/rsna-2024-lumbar-spine-degenerative-classification'","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:00:36.888507Z","iopub.execute_input":"2024-09-18T06:00:36.889341Z","iopub.status.idle":"2024-09-18T06:00:37.905003Z","shell.execute_reply.started":"2024-09-18T06:00:36.889298Z","shell.execute_reply":"2024-09-18T06:00:37.904086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# configs\ndefault_color_1 = 'darkblue'","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:00:40.665480Z","iopub.execute_input":"2024-09-18T06:00:40.666315Z","iopub.status.idle":"2024-09-18T06:00:40.670493Z","shell.execute_reply.started":"2024-09-18T06:00:40.666274Z","shell.execute_reply":"2024-09-18T06:00:40.669614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read data\npath = '../input/rsna-2024-lumbar-spine-degenerative-classification/'\n\ndf_train_main  = pd.read_csv(path + 'train.csv')\ndf_train_label = pd.read_csv(path + 'train_label_coordinates.csv')\ndf_train_desc  = pd.read_csv(path + 'train_series_descriptions.csv')\ndf_test_desc   = pd.read_csv(path + 'test_series_descriptions.csv')\ndf_sub         = pd.read_csv(path + 'sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:00:43.555789Z","iopub.execute_input":"2024-09-18T06:00:43.556425Z","iopub.status.idle":"2024-09-18T06:00:43.729980Z","shell.execute_reply.started":"2024-09-18T06:00:43.556394Z","shell.execute_reply":"2024-09-18T06:00:43.729200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div  style=\"color:#ff0a6c; border:#0014ff solid;  font-weight:bold; font-size:120%; text-align:center;padding:12.0px; background:#000000\">3. DATA ANALYSIS</div>","metadata":{}},{"cell_type":"code","source":"# structure of train\ndf_train_main.info()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:00:53.278040Z","iopub.execute_input":"2024-09-18T06:00:53.278536Z","iopub.status.idle":"2024-09-18T06:00:53.313563Z","shell.execute_reply.started":"2024-09-18T06:00:53.278503Z","shell.execute_reply":"2024-09-18T06:00:53.312580Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Total Cases: \", len(df_train_main))","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:00.415373Z","iopub.execute_input":"2024-09-18T06:01:00.416045Z","iopub.status.idle":"2024-09-18T06:01:00.422005Z","shell.execute_reply.started":"2024-09-18T06:01:00.416012Z","shell.execute_reply":"2024-09-18T06:01:00.421001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_main.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:02.337038Z","iopub.execute_input":"2024-09-18T06:01:02.337670Z","iopub.status.idle":"2024-09-18T06:01:02.367737Z","shell.execute_reply.started":"2024-09-18T06:01:02.337637Z","shell.execute_reply":"2024-09-18T06:01:02.366841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Structure of train_label:","metadata":{}},{"cell_type":"code","source":"df_train_label.info()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:05.651006Z","iopub.execute_input":"2024-09-18T06:01:05.651335Z","iopub.status.idle":"2024-09-18T06:01:05.671699Z","shell.execute_reply.started":"2024-09-18T06:01:05.651311Z","shell.execute_reply":"2024-09-18T06:01:05.670614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_label.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:08.513374Z","iopub.execute_input":"2024-09-18T06:01:08.513767Z","iopub.status.idle":"2024-09-18T06:01:08.528267Z","shell.execute_reply.started":"2024-09-18T06:01:08.513736Z","shell.execute_reply":"2024-09-18T06:01:08.527191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# look at categories\nfor f in ['instance_number','condition','level']:\n    print(df_train_label[f].value_counts())\n    print('-'*50);print();","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:12.241774Z","iopub.execute_input":"2024-09-18T06:01:12.242632Z","iopub.status.idle":"2024-09-18T06:01:12.265771Z","shell.execute_reply.started":"2024-09-18T06:01:12.242600Z","shell.execute_reply":"2024-09-18T06:01:12.264873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Crosstab condition vs level:","metadata":{}},{"cell_type":"code","source":"pd.crosstab(df_train_label.condition, df_train_label.level)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:15.797762Z","iopub.execute_input":"2024-09-18T06:01:15.798106Z","iopub.status.idle":"2024-09-18T06:01:15.835173Z","shell.execute_reply.started":"2024-09-18T06:01:15.798079Z","shell.execute_reply":"2024-09-18T06:01:15.834238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Combine Tables:\n","metadata":{}},{"cell_type":"code","source":"# join first two tables\ndf_train_step_1 = pd.merge(left=df_train_label, right=df_train_main, how='left', on='study_id').reset_index(drop=True)\ndf_train_step_1.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:19.204860Z","iopub.execute_input":"2024-09-18T06:01:19.205644Z","iopub.status.idle":"2024-09-18T06:01:19.328941Z","shell.execute_reply.started":"2024-09-18T06:01:19.205489Z","shell.execute_reply":"2024-09-18T06:01:19.328018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# join with third table\ndf_train = pd.merge(left=df_train_step_1, right=df_train_desc, how='left', on=['study_id', 'series_id']).reset_index(drop=True)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:21.594660Z","iopub.execute_input":"2024-09-18T06:01:21.595037Z","iopub.status.idle":"2024-09-18T06:01:21.688741Z","shell.execute_reply.started":"2024-09-18T06:01:21.595006Z","shell.execute_reply":"2024-09-18T06:01:21.687871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convert study_id's to categorical\ndf_train.study_id = df_train.study_id.astype('category')\ndf_train.series_id = df_train.series_id.astype('category')\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:23.848482Z","iopub.execute_input":"2024-09-18T06:01:23.849285Z","iopub.status.idle":"2024-09-18T06:01:23.876426Z","shell.execute_reply.started":"2024-09-18T06:01:23.849252Z","shell.execute_reply":"2024-09-18T06:01:23.875449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Coordinates distributions","metadata":{}},{"cell_type":"code","source":"# plot coordinates\ndefault_color_1 = 'black'\nsns.scatterplot(data=df_train, x='x', y='y', color=default_color_1, s = 20,  alpha=0.2)\n\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:27.535233Z","iopub.execute_input":"2024-09-18T06:01:27.535916Z","iopub.status.idle":"2024-09-18T06:01:27.938344Z","shell.execute_reply.started":"2024-09-18T06:01:27.535878Z","shell.execute_reply":"2024-09-18T06:01:27.937507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot coordinate distributions with colored by **condition**","metadata":{}},{"cell_type":"code","source":"p=['#4600c0',  '#00a419', '#111111', '#ffff00', '#ff0000']\nfigs_x=5  \nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nsns.scatterplot(data=df_train, x='x', y='y', hue='condition',  palette=p, s = 20,  alpha=0.2)\nplt.legend(bbox_to_anchor=(1.2,1), loc=2, prop={'size': 10}, markerscale = 2, framealpha=0.8, facecolor='white')\nplt.grid()\nplt.xlim([0, df_train.x.max()])\nplt.ylim([0, df_train.y.max()])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:40.283901Z","iopub.execute_input":"2024-09-18T06:01:40.284594Z","iopub.status.idle":"2024-09-18T06:01:41.940804Z","shell.execute_reply.started":"2024-09-18T06:01:40.284563Z","shell.execute_reply":"2024-09-18T06:01:41.939880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### Plot xy-Density of cases in stenosis split","metadata":{}},{"cell_type":"code","source":"numerical_features =   ['Left Neural Foraminal Narrowing', 'Right Neural Foraminal Narrowing',\n                        'Left Subarticular Stenosis',      'Right Subarticular Stenosis',\n                        'Spinal Canal Stenosis']\ngs=150\nn=2 # num of columns\na=0 \nk=1;\ncolorlabels = 'darkblue'\n\nLabel_size = 10 # Size font of xy labels\nTitle_size = 20 # Size font of Title\nfigs_x=13   \nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nfor i in numerical_features:\n        \n        d = df_train[df_train.condition == i]\n        plt.subplot(1,n, k)\n        #sns.jointplot(data=d, x='x', y='y', color='white', alpha=0.25)\n        plt.hexbin(data=d, x='x', y='y',  gridsize=gs, cmap='CMRmap', bins='log', alpha = 1)\n        \n        #plt.colorbar(label='count in bin')\n        plt.colorbar().set_label(label='count in bin',size=10, color  = 'grey')\n        #plt.colorbar(size=8)\n        \n        #plt.xlabel(f'{i}')\n        #plt.ylabel(f'{j}')\n        plt.tick_params(axis='x', labelsize=Label_size)\n        plt.tick_params(axis='y', labelsize=Label_size)\n        \n        plt.xlim([0, df_train.x.max()])\n        plt.ylim([0, df_train.y.max()])\n        plt.xlabel(f'x', fontsize=Label_size, color = colorlabels)\n        plt.ylabel(f'y', fontsize=Label_size, color = colorlabels) \n        \n        plt.title(f'{i}', color='black', fontsize=Title_size)\n        k=k+1\n        if k == (n+1):    \n            k=1\n            plt.show()\n            plt.figure(figsize=(figs_x, figs_y)) \n","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:47.065818Z","iopub.execute_input":"2024-09-18T06:01:47.066195Z","iopub.status.idle":"2024-09-18T06:01:50.295818Z","shell.execute_reply.started":"2024-09-18T06:01:47.066162Z","shell.execute_reply":"2024-09-18T06:01:50.294889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot coordinates distributions with colored by **level**","metadata":{}},{"cell_type":"code","source":"p=['#4600c0',  '#00a419', '#111111', '#ffff00', '#ff0000']\nfigs_x=5\nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nsns.scatterplot(data=df_train, x='x', y='y', hue='level',  palette=p, s = 20,  alpha=0.3)\nplt.legend(bbox_to_anchor=(1.2,1), loc=2, prop={'size': 10}, markerscale = 2, framealpha=0.8, facecolor='white', reverse=True)\nplt.grid()\nplt.xlim([0, df_train.x.max()])\nplt.ylim([0, df_train.y.max()])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:01:55.391787Z","iopub.execute_input":"2024-09-18T06:01:55.392511Z","iopub.status.idle":"2024-09-18T06:01:57.042759Z","shell.execute_reply.started":"2024-09-18T06:01:55.392474Z","shell.execute_reply":"2024-09-18T06:01:57.041763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.level.unique()\ndf_train.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:02:00.688792Z","iopub.execute_input":"2024-09-18T06:02:00.689529Z","iopub.status.idle":"2024-09-18T06:02:00.717564Z","shell.execute_reply.started":"2024-09-18T06:02:00.689498Z","shell.execute_reply":"2024-09-18T06:02:00.716665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_features =   df_train.level.unique() #L1/L2', 'L2/L3', 'L3/L4', 'L4/L5', 'L5/S1'\ngs=150\nn=2 # num of columns\na=0 \nk=1;\ncolorlabels = 'darkblue'\n\nLabel_size = 10 # Size font of xy labels\nTitle_size = 20 # Size font of Title\nfigs_x=13   \nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nplt.suptitle(\"xy-Density of cases in L-split\", fontsize=Title_size+2, fontweight='bold', y=1.015)\nfor i in numerical_features:\n        \n        #d = df_train[ (df_train.level == i) & (df_train.condition == 'Left Neural Foraminal Narrowing') ]\n        d = df_train[ (df_train.level == i) ]\n        plt.subplot(1,n, k)\n        #sns.jointplot(data=d, x='x', y='y', color='white', alpha=0.25)\n        plt.hexbin(data=d, x='x', y='y',  gridsize=gs, cmap='CMRmap', bins='log', alpha = 1)\n        \n        #plt.colorbar(label='count in bin')\n        plt.colorbar().set_label(label='count in bin',size=10, color  = 'grey')\n        #plt.colorbar(size=8)\n        \n        #plt.xlabel(f'{i}')\n        #plt.ylabel(f'{j}')\n        plt.tick_params(axis='x', labelsize=Label_size)\n        plt.tick_params(axis='y', labelsize=Label_size)\n        \n        plt.xlim([0, df_train.x.max()])\n        plt.ylim([0, df_train.y.max()])\n        plt.xlabel(f'x', fontsize=Label_size, color = colorlabels)\n        plt.ylabel(f'y', fontsize=Label_size, color = colorlabels) \n        \n        plt.title(f'{i}', color='black', fontsize=Title_size)\n        k=k+1\n        if k == (n+1):    \n            k=1\n            plt.show()\n            plt.figure(figsize=(figs_x, figs_y)) \n            ","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:02:02.846780Z","iopub.execute_input":"2024-09-18T06:02:02.847660Z","iopub.status.idle":"2024-09-18T06:02:06.074875Z","shell.execute_reply.started":"2024-09-18T06:02:02.847624Z","shell.execute_reply":"2024-09-18T06:02:06.073943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot stenosis distributions with colored by **MR-split**","metadata":{}},{"cell_type":"code","source":"p=['#000000', '#4600c0', '#ff0000']\nfigs_x=5  \nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nsns.scatterplot(data=df_train, x='x', y='y', hue='series_description',  palette=p, s = 20,  alpha=0.4)\nplt.legend(bbox_to_anchor=(1.2,1), loc=2, prop={'size': 10}, markerscale = 2, framealpha=0.8, facecolor='white')\nplt.grid()\nplt.xlim([0, df_train.x.max()])\nplt.ylim([0, df_train.y.max()])\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-07-13T11:45:46.240958Z","iopub.execute_input":"2024-07-13T11:45:46.241374Z","iopub.status.idle":"2024-07-13T11:45:47.79717Z","shell.execute_reply.started":"2024-07-13T11:45:46.241339Z","shell.execute_reply":"2024-07-13T11:45:47.79611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_features =   df_train.series_description.unique()\ngs=200\nn=2 # num of columns\na=0 \nk=1;\ncolorlabels = 'darkblue'\n\nLabel_size = 10 # Size font of xy labels\nTitle_size = 20 # Size font of Title\nfigs_x=13   \nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nplt.suptitle(\"xy-Density of cases in plane-split\", fontsize=Title_size+2, fontweight='bold', y=1.015)\nfor i in numerical_features:\n        \n        d = df_train[df_train.series_description == i]\n        plt.subplot(1,n, k)\n        #sns.jointplot(data=d, x='x', y='y', color='white', alpha=0.25)\n        plt.hexbin(data=d, x='x', y='y',  gridsize=gs, cmap='CMRmap', bins='log', alpha = 1)\n        \n        #plt.colorbar(label='count in bin')\n        plt.colorbar().set_label(label='count in bin',size=10, color  = 'grey')\n        #plt.colorbar(size=8)\n        \n        plt.tick_params(axis='x', labelsize=Label_size)\n        plt.tick_params(axis='y', labelsize=Label_size)\n        \n        plt.xlim([0, df_train.x.max()])\n        plt.ylim([0, df_train.y.max()])\n        plt.xlabel(f'x', fontsize=Label_size, color = colorlabels)\n        plt.ylabel(f'y', fontsize=Label_size, color = colorlabels) \n        \n        plt.title(f'{i}', color='black', fontsize=Title_size)\n        k=k+1\n        if k == (n+1):    \n            k=1\n            plt.show()\n            plt.figure(figsize=(figs_x, figs_y)) ","metadata":{"execution":{"iopub.status.busy":"2024-07-13T11:45:47.798818Z","iopub.execute_input":"2024-07-13T11:45:47.799737Z","iopub.status.idle":"2024-07-13T11:45:50.339866Z","shell.execute_reply.started":"2024-07-13T11:45:47.799698Z","shell.execute_reply":"2024-07-13T11:45:50.33882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target distribution","metadata":{}},{"cell_type":"markdown","source":"### Let's exploring target distribution in different splits","metadata":{}},{"cell_type":"markdown","source":"At the first, creating the new column by concatenating values from columns entiteld **conditions** and **level**","metadata":{}},{"cell_type":"code","source":"train_label_df = df_train_label.copy()\ntrain_data_df = df_train_main.copy()\ntrain_label_df['new_col'] = df_train_label['condition'].str.lower().str.replace(' ', '_') +  '_' + df_train_label['level'].str.lower().str.replace('/', '_')\n\n# Step 2: Merge the values from train_data_df based on study_id and the newly created column names\ndef get_target_value(row):\n    study_id = row['study_id']\n    new_col = row['new_col']\n    return train_data_df[train_data_df['study_id'] == study_id][new_col].values[0]\n\ntrain_label_df['target'] = train_label_df.apply(get_target_value, axis=1)\n\n# Drop the 'new_col' column \ntrain_label_df.drop(columns=['new_col'], inplace=True)\n\n# Copy the train_label_df to new DataFrame\nfinal_train_df = train_label_df.copy()\n\n# Drop the 'new_col' column \ntrain_label_df.drop(columns=['target'], inplace=True)\n\n#final_train_df.to_csv('final_train.csv')\nfinal_train_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:02:14.968518Z","iopub.execute_input":"2024-09-18T06:02:14.969383Z","iopub.status.idle":"2024-09-18T06:02:33.382855Z","shell.execute_reply.started":"2024-09-18T06:02:14.969349Z","shell.execute_reply":"2024-09-18T06:02:33.381907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Plotting the distribution of the target variable:","metadata":{}},{"cell_type":"code","source":"#plt.suptitle(\"Target class distribution \\n (1 - default, 0 - not default)\", fontsize=15, fontweight='bold', y=1.0)\n\nname = ['Normal/Mild', 'Moderate', 'Severe']\np=['#d0d0d0', '#ffba07', '#ff0000']\nplt.pie(final_train_df.target.value_counts(normalize=True), autopct = '%1.f%%', \ncolors = sns.color_palette(p), startangle=90, wedgeprops=dict(width=0.25), labeldistance=1.2, \ncounterclock=False, radius=1)\nplt.title(f'Total', color='blue', fontsize=Title_size)\nplt.legend(name, loc='lower left',  prop={'size': 12}, markerscale = 3, framealpha=0.8, facecolor='white')\n","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:02:36.010727Z","iopub.execute_input":"2024-09-18T06:02:36.011072Z","iopub.status.idle":"2024-09-18T06:02:36.201670Z","shell.execute_reply.started":"2024-09-18T06:02:36.011044Z","shell.execute_reply":"2024-09-18T06:02:36.200764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Plotting the distribution of the target variable in L/L splits:","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(20, 5))\ncolours = {\"male\": \"#273c75\", \"female\": \"#44bd32\"}\nfor idx, d in enumerate(['foraminal', 'subarticular', 'canal']):\n    diagnosis = list(filter(lambda x: x.find(d) > -1, df_train_main.columns))\n    dff = df_train_main[diagnosis]\n    with warnings.catch_warnings():\n        warnings.simplefilter(action=\"ignore\", category=FutureWarning)\n        value_counts = dff.apply(pd.value_counts).fillna(0).T\n      \n    value_counts.plot(kind='bar', stacked=True, ax=axes[idx], cmap='Set1_r')\n    \n    axes[idx].tick_params(axis='x', labelsize=16) \n    \n    axes[idx].set_title(f\"{d} distribution\",  fontsize=28)\n\n    \n#color=plotdata['gender'].replace(colours)    ","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:02:40.295992Z","iopub.execute_input":"2024-09-18T06:02:40.296369Z","iopub.status.idle":"2024-09-18T06:02:41.549429Z","shell.execute_reply.started":"2024-09-18T06:02:40.296338Z","shell.execute_reply":"2024-09-18T06:02:41.548545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Other view on target variable in L/L splits:","metadata":{}},{"cell_type":"code","source":"n=3 # num of columns\na=0 \nk=1;\ncolorlabels = 'darkblue'\n\np=['#d0d0d0', '#ffba07', '#ff0000']\nlabel = df_train_main.columns.drop('study_id').tolist()        \n        \nLabel_size = 12 # Size font of xy labels\nTitle_size = 15 # Size font of Title\nfigs_x=16\nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nfor i in label:\n    plt.subplot(1, n, k)\n    plt.pie(df_train_main[i].value_counts(normalize=True), autopct = '%1.f%%', \n        colors = sns.color_palette(p), startangle=90, wedgeprops=dict(width=0.25), \n        counterclock=False, radius=1)\n    plt.title(f'{i}', color='blue', fontsize=Title_size)\n    plt.legend(name, loc='lower left',  prop={'size': 12}, markerscale = 3, framealpha=0.8, facecolor='white')\n    k=k+1\n    if k == (n+1):    \n        k=1\n        plt.show()\n        plt.figure(figsize=(figs_x, figs_y))","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:02:44.774362Z","iopub.execute_input":"2024-09-18T06:02:44.775337Z","iopub.status.idle":"2024-09-18T06:02:48.997430Z","shell.execute_reply.started":"2024-09-18T06:02:44.775301Z","shell.execute_reply":"2024-09-18T06:02:48.995899Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's pick a study to see images","metadata":{}},{"cell_type":"code","source":"df_train.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:07.695358Z","iopub.execute_input":"2024-09-18T06:03:07.695716Z","iopub.status.idle":"2024-09-18T06:03:07.717560Z","shell.execute_reply.started":"2024-09-18T06:03:07.695689Z","shell.execute_reply":"2024-09-18T06:03:07.716512Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What do we need to predict?¶\nFor a given study in the test data, we need to predict the severity condition of all types of stenosis at all levels. So from a first impression it seems a very heavy problem, because we need to predict for each of the 25 different possible combinations.","metadata":{}},{"cell_type":"markdown","source":"# <div  style=\"color:#ff0a6c;  border:#0014ff solid; font-weight:bold; font-size:120%; text-align:center;padding:12.0px; background:#000000\">4. MR-IMAGES</div>","metadata":{}},{"cell_type":"markdown","source":"# What is DICOM?","metadata":{}},{"cell_type":"markdown","source":"A .dcm file follows the **Digital Imaging and Communications in Medicine** (DICOM) format. It is the standard format used for storing medical images and related metadata. It dates back to 1983, although it has been revised many times.\n\nWe can use the pydicom library to open and explore these files.","metadata":{}},{"cell_type":"code","source":"import pydicom as dicom\nimport matplotlib.patches as patches","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:16.131970Z","iopub.execute_input":"2024-09-18T06:03:16.132862Z","iopub.status.idle":"2024-09-18T06:03:16.136813Z","shell.execute_reply.started":"2024-09-18T06:03:16.132815Z","shell.execute_reply":"2024-09-18T06:03:16.135761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_label_coordinates=df_train_label","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:18.787685Z","iopub.execute_input":"2024-09-18T06:03:18.788178Z","iopub.status.idle":"2024-09-18T06:03:18.792360Z","shell.execute_reply.started":"2024-09-18T06:03:18.788145Z","shell.execute_reply":"2024-09-18T06:03:18.791411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_label_coordinates[train_label_coordinates.study_id==4003253]","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:21.849773Z","iopub.execute_input":"2024-09-18T06:03:21.850672Z","iopub.status.idle":"2024-09-18T06:03:21.867362Z","shell.execute_reply.started":"2024-09-18T06:03:21.850639Z","shell.execute_reply":"2024-09-18T06:03:21.866490Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Introduction to image visualisation with **imshow()**\n","metadata":{"execution":{"iopub.status.busy":"2024-05-29T20:44:23.044182Z","iopub.execute_input":"2024-05-29T20:44:23.044618Z","iopub.status.idle":"2024-05-29T20:44:23.054793Z","shell.execute_reply.started":"2024-05-29T20:44:23.044586Z","shell.execute_reply":"2024-05-29T20:44:23.053611Z"}}},{"cell_type":"markdown","source":"Let's create random array 100x100:","metadata":{}},{"cell_type":"code","source":"rand_data = np.random.randint(0, 255, (100, 100))\nrand_data","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:25.591178Z","iopub.execute_input":"2024-09-18T06:03:25.592037Z","iopub.status.idle":"2024-09-18T06:03:25.599297Z","shell.execute_reply.started":"2024-09-18T06:03:25.592004Z","shell.execute_reply":"2024-09-18T06:03:25.598097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Visualizing random array with three color map:","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(15, 3))\n\naxes[0].imshow(rand_data, cmap ='viridis')\naxes[1].imshow(rand_data, cmap ='CMRmap')\naxes[2].imshow(rand_data, cmap ='binary')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:30.046823Z","iopub.execute_input":"2024-09-18T06:03:30.047369Z","iopub.status.idle":"2024-09-18T06:03:30.572673Z","shell.execute_reply.started":"2024-09-18T06:03:30.047336Z","shell.execute_reply":"2024-09-18T06:03:30.571715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing MR-images","metadata":{}},{"cell_type":"code","source":"train_label_coordinates['series_description'] = df_train.series_description","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:34.876297Z","iopub.execute_input":"2024-09-18T06:03:34.876936Z","iopub.status.idle":"2024-09-18T06:03:34.882088Z","shell.execute_reply.started":"2024-09-18T06:03:34.876905Z","shell.execute_reply":"2024-09-18T06:03:34.881123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def mrt(id, ser, inst):\n    lag=20\n    path2 = path+'train_images/' + str(id) +'/' + str(ser)+'/' + str(inst) + '.dcm'\n\n    ds = dicom.dcmread(path2)\n    fig, ax = plt.subplots(figsize=(16, 8))\n    from matplotlib.colors import LogNorm \n\n    ax.imshow(ds.pixel_array, cmap ='CMRmap')     # Display the image\n\n    # Create a legend\n    legend_elements = []\n\n    # Plot the coordinates for the current condition\n    ab = train_label_coordinates[(train_label_coordinates.study_id==id) & \n                                          (train_label_coordinates.instance_number==inst)&\n                                         (train_label_coordinates.series_id==ser)]\n\n    a = 25 * max(ds.pixel_array.shape)/640\n    for _, row in ab.iterrows():\n        x, y = row['x'], row['y']\n\n        rect2 = patches.Rectangle((x - a, y - a), 2*a, 2*a, linewidth=2, edgecolor='white', facecolor='none')\n        rect1 = patches.Rectangle((x - a, y - a), 2*a, 2*a, linewidth=2, facecolor='white', alpha = 0.25)\n\n        ax.add_patch(rect2)\n        ax.add_patch(rect1)\n\n        # Add the condition to the legend\n        legend_elements.append(patches.Patch(facecolor='none', edgecolor='r', ))\n\n    # Add title\n    title = f\"{ab.series_description.unique()}, Study: {id}, Series: {ser}, Instance: {inst}\"\n    ax.set_title(title, fontsize=20)\n\n    # Display additional columns:\n    for _, row in ab.iterrows():\n        text = f\"level {row['level']}, {row['condition']}\"\n        ax.text(row['x'] + lag, row['y']+np.random.randint(-15, 15), text, fontsize=10, color='white', verticalalignment='center_baseline')\n    \n    plt.show() ","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:37.571840Z","iopub.execute_input":"2024-09-18T06:03:37.572186Z","iopub.status.idle":"2024-09-18T06:03:37.583665Z","shell.execute_reply.started":"2024-09-18T06:03:37.572156Z","shell.execute_reply":"2024-09-18T06:03:37.582706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #1\nid = 4003253\nser= 702807833\ninst=8\n\nmrt(id, ser, inst)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:39.890110Z","iopub.execute_input":"2024-09-18T06:03:39.890493Z","iopub.status.idle":"2024-09-18T06:03:40.470766Z","shell.execute_reply.started":"2024-09-18T06:03:39.890463Z","shell.execute_reply":"2024-09-18T06:03:40.469786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #2\nid = 4290709089\nser= 4237840455\ninst=11\n\nmrt(id, ser, inst)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:45.875994Z","iopub.execute_input":"2024-09-18T06:03:45.876383Z","iopub.status.idle":"2024-09-18T06:03:46.368897Z","shell.execute_reply.started":"2024-09-18T06:03:45.876340Z","shell.execute_reply":"2024-09-18T06:03:46.368046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #3\nid = 4003253\nser= 1054713880\ninst=4\n\nmrt(id, ser, inst)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:49.430994Z","iopub.execute_input":"2024-09-18T06:03:49.431357Z","iopub.status.idle":"2024-09-18T06:03:49.896139Z","shell.execute_reply.started":"2024-09-18T06:03:49.431328Z","shell.execute_reply":"2024-09-18T06:03:49.895272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #4\nid = 4003253\nser= 2448190387\ninst=11\n\nmrt(id, ser, inst)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:54.103839Z","iopub.execute_input":"2024-09-18T06:03:54.104675Z","iopub.status.idle":"2024-09-18T06:03:54.588685Z","shell.execute_reply.started":"2024-09-18T06:03:54.104644Z","shell.execute_reply":"2024-09-18T06:03:54.587761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3D MR-slides visualisation","metadata":{}},{"cell_type":"code","source":"def MR3d(id, ser):\n    \n    path_to_folder = path + \"train_images/\"+str(id)+'/'+str(ser)\n    def load_dicom(path):\n        dicom = pydicom.read_file(path)\n        data = dicom.pixel_array\n        data = data - np.min(data)\n        if np.max(data) != 0:\n            data = data / np.max(data)\n        data = (data * 255).astype(np.uint8)\n        return data\n\n    rc('animation', html='jshtml')\n\n    def load_dicom(filename):\n        ds = pydicom.dcmread(filename)\n        return ds.pixel_array\n\n    def load_dicom_line(path):\n        t_paths = sorted(\n            glob.glob(os.path.join(path, \"*\")), \n            key=lambda x: int(os.path.splitext(os.path.basename(x))[0].split(\"-\")[-1]),\n        )\n        images = []\n        for filename in t_paths:\n            data = load_dicom(filename)\n            if data.max() == 0:\n                continue\n            images.append(data)\n        return images\n\n    def create_animation(ims):\n        fig = plt.figure(figsize=(6, 6))\n        plt.axis('off')\n        im = plt.imshow(ims[0], cmap=\"CMRmap\")\n        text = plt.text(0.05, 0.05, f'Slide {1}', transform=fig.transFigure, fontsize=16, color='darkblue')\n\n        def animate_func(i):\n            im.set_array(ims[i])\n            text.set_text(f'Slide {i+1}')  \n            return [im]\n        plt.title(f'id = {id}, series = {ser}')\n        \n        plt.close()  \n\n        return animation.FuncAnimation(fig, animate_func, frames=len(ims), interval=1000//10) #24\n\n    images = load_dicom_line(path_to_folder)\n    \n    return create_animation(images)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:03:59.019056Z","iopub.execute_input":"2024-09-18T06:03:59.019409Z","iopub.status.idle":"2024-09-18T06:03:59.032207Z","shell.execute_reply.started":"2024-09-18T06:03:59.019381Z","shell.execute_reply":"2024-09-18T06:03:59.031307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #1\nid = 4003253\nser= 702807833\n\nMR3d(id, ser)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:04:05.514529Z","iopub.execute_input":"2024-09-18T06:04:05.514888Z","iopub.status.idle":"2024-09-18T06:04:08.385471Z","shell.execute_reply.started":"2024-09-18T06:04:05.514859Z","shell.execute_reply":"2024-09-18T06:04:08.384175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #2\nid = 4290709089\nser= 4237840455\n\nMR3d(id, ser)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:04:19.992596Z","iopub.execute_input":"2024-09-18T06:04:19.993376Z","iopub.status.idle":"2024-09-18T06:04:22.523411Z","shell.execute_reply.started":"2024-09-18T06:04:19.993340Z","shell.execute_reply":"2024-09-18T06:04:22.521051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #3\nid = 4003253\nser= 1054713880\n\nMR3d(id, ser)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:04:31.905485Z","iopub.execute_input":"2024-09-18T06:04:31.906091Z","iopub.status.idle":"2024-09-18T06:04:34.204110Z","shell.execute_reply.started":"2024-09-18T06:04:31.906059Z","shell.execute_reply":"2024-09-18T06:04:34.202868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #4\nid = 4003253\nser= 2448190387\n\nMR3d(id, ser)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T06:04:44.634482Z","iopub.execute_input":"2024-09-18T06:04:44.635142Z","iopub.status.idle":"2024-09-18T06:04:51.500777Z","shell.execute_reply.started":"2024-09-18T06:04:44.635109Z","shell.execute_reply":"2024-09-18T06:04:51.499168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DICOM Metadata","metadata":{}},{"cell_type":"code","source":"import os\nfrom glob import glob\nfrom tqdm import tqdm\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport missingno as msno\nimport pydicom\n\nmetadata = []\ndicom_attributes = [\n    'StudyInstanceUID', 'SeriesInstanceUID', 'PatientID',\n    'ContentDate', 'ContentTime', 'SeriesDescription',\n    'SliceThickness', 'SpacingBetweenSlices', 'PatientPosition',\n    'InstanceNumber', 'ImagePositionPatient', 'ImageOrientationPatient',\n    'SliceLocation', 'SamplesPerPixel', 'PhotometricInterpretation',\n    'Rows', 'Columns', 'PixelSpacing', 'BitsAllocated', 'BitsStored',\n    'HighBit', 'PixelRepresentation', 'WindowCenter', 'WindowWidth',\n    'RescaleIntercept', 'RescaleSlope'\n]\ncompetition_dataset_directory = Path('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification')\ndf_train_series_descriptions = pd.read_csv(competition_dataset_directory / 'train_series_descriptions.csv')\ntrain_images_root_directory = competition_dataset_directory / 'train_images'\nstudy_ids = os.listdir(str(train_images_root_directory))\n\nfor study_id in tqdm(study_ids):\n\n    study_directory = train_images_root_directory / study_id\n    series_ids = os.listdir(str(study_directory))\n\n    for series_id in series_ids:\n\n        series_directory = study_directory / series_id\n        dicom_file_names = os.listdir(series_directory)\n\n        for dicom_file_name in dicom_file_names:\n\n            dicom_path = series_directory / dicom_file_name\n            dicom = pydicom.dcmread(dicom_path)\n\n            dicom_dict = {}\n            dicom_dict['slice_id'] = int(dicom_file_name.split('.')[0])\n            for tag in dicom_attributes:\n                try:\n                    dicom_tag_value = dicom[tag].value\n                except KeyError:\n                    dicom_tag_value = np.nan\n                dicom_dict[tag] = dicom_tag_value\n\n            metadata.append(dicom_dict)\n\n\ndf_metadata = pd.DataFrame(metadata)\ndf_metadata = df_metadata.sort_values(by=['StudyInstanceUID', 'SeriesInstanceUID', 'slice_id'], ascending=True).reset_index(drop=True)\ndf_metadata['StudyInstanceUID'] = df_metadata['StudyInstanceUID'].astype(int)\ndf_metadata['SeriesInstanceUID'] = df_metadata['SeriesInstanceUID'].apply(lambda x: int(str(x).split('.')[-1])).astype(int)\ndf_metadata['PatientID'] = df_metadata['PatientID'].astype(int)\ndf_metadata['ContentDate'] = pd.to_datetime(df_metadata['ContentDate'] + ' ' + df_metadata['ContentTime'].apply(lambda x: f'{x[0:2]}:{x[2:4]}'))\ndf_metadata = df_metadata.drop(columns=['PatientID', 'ContentTime'])\n\ndf_metadata = df_metadata.merge(\n    df_train_series_descriptions.rename(columns={\n        'study_id': 'StudyInstanceUID',\n        'series_id': 'SeriesInstanceUID'\n    }),\n    on=['StudyInstanceUID', 'SeriesInstanceUID'],\n    how='left'\n)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T09:19:43.141344Z","iopub.execute_input":"2024-09-18T09:19:43.142272Z","iopub.status.idle":"2024-09-18T09:42:23.933237Z","shell.execute_reply.started":"2024-09-18T09:19:43.142229Z","shell.execute_reply":"2024-09-18T09:42:23.932131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_continuous_metadata_column_histogram(df, title, path=None):\n\n    \"\"\"\n    Visualize histogram of the given continuous column within modalities\n\n    Parameters\n    ----------\n    df: pandas.DataFrame\n        Dataframe with given continuous column and series description as index\n\n    title: str\n        Title of the plot\n\n    path: str, pathlib.Path or None\n        Path of the output file or None (if path is None, plot is displayed with selected backend)\n    \"\"\"\n\n    fig, ax = plt.subplots(figsize=(32, 8), dpi=100)\n    _, bins, _ = ax.hist(df.loc['Axial T2'], bins=24, label='Axial T2', alpha=0.5)\n    ax.hist(df.loc['Sagittal T1'], bins=bins, label='Sagittal T1', alpha=0.5)\n    ax.hist(df.loc['Sagittal T2/STIR'], bins=bins, label='Sagittal T2/STIR', alpha=0.5)\n    ax.tick_params(axis='x', labelsize=15)\n    ax.tick_params(axis='y', labelsize=15)\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.legend(prop={'size': 18})\n    ax.set_title(\n        title + f'''\n        Axial T2 Mean: {np.mean(df.loc['Axial T2']):.2f} Std: {np.std(df.loc['Axial T2']):.2f} Min: {np.min(df.loc['Axial T2']):.2f} Max: {np.max(df.loc['Axial T2']):.2f}\n        Sagittal T1 Mean: {np.mean(df.loc['Sagittal T1']):.2f} Std: {np.std(df.loc['Sagittal T1']):.2f} Min: {np.min(df.loc['Sagittal T1']):.2f} Max: {np.max(df.loc['Sagittal T1']):.2f}\n        Sagittal T2/STIR Mean: {np.mean(df.loc['Sagittal T2/STIR']):.2f} Std: {np.std(df.loc['Sagittal T2/STIR']):.2f} Min: {np.min(df.loc['Sagittal T2/STIR']):.2f} Max: {np.max(df.loc['Sagittal T2/STIR']):.2f}\n        ''',\n        size=15,\n        pad=12.5,\n        loc='center',\n        wrap=True\n    )\n\n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path, bbox_inches='tight')\n        plt.close(fig)","metadata":{"execution":{"iopub.status.busy":"2024-09-18T10:14:13.860333Z","iopub.execute_input":"2024-09-18T10:14:13.861227Z","iopub.status.idle":"2024-09-18T10:14:13.871402Z","shell.execute_reply.started":"2024-09-18T10:14:13.861193Z","shell.execute_reply":"2024-09-18T10:14:13.870612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_metadata['PixelSpacingX'] = df_metadata['PixelSpacing'].apply(lambda x: float(x[0]))\ndf_metadata['PixelSpacingY'] = df_metadata['PixelSpacing'].apply(lambda x: float(x[1]))\n\ndf_slice_counts = df_metadata.groupby(['StudyInstanceUID', 'series_description'])['series_description'].count().reset_index(level=[0], drop=True).sort_index()\n\nvisualize_continuous_metadata_column_histogram(\n    df=df_slice_counts,\n    title='Series Slice Counts Distribution'\n)\n\ncontinuous_columns = [\n    'SliceThickness', 'SpacingBetweenSlices',\n    'SamplesPerPixel', 'Rows', 'Columns', 'PixelSpacingX', 'PixelSpacingY',\n    'BitsAllocated', 'BitsStored', 'HighBit',\n    'PixelRepresentation',\n    'WindowCenter', 'WindowWidth', 'RescaleIntercept', 'RescaleSlope'\n]\ndf_metadata_continuous_columns = df_metadata.groupby(['StudyInstanceUID', 'series_description'])[continuous_columns].first().reset_index(level=[0], drop=True).sort_index()\nfor column in continuous_columns:\n    visualize_continuous_metadata_column_histogram(\n        df=df_metadata_continuous_columns[column],\n        title=f'Series {column} Distribution'\n    )","metadata":{"execution":{"iopub.status.busy":"2024-09-18T10:14:16.199406Z","iopub.execute_input":"2024-09-18T10:14:16.199799Z","iopub.status.idle":"2024-09-18T10:14:30.564431Z","shell.execute_reply.started":"2024-09-18T10:14:16.199749Z","shell.execute_reply":"2024-09-18T10:14:30.563405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div  style=\"color:#ff0a6c;  border:#0014ff solid; font-weight:bold; font-size:120%; text-align:center;padding:12.0px; background:#000000\">5. MODELLING</div>","metadata":{}},{"cell_type":"markdown","source":"Define Competition metrics (for future):","metadata":{}},{"cell_type":"code","source":"import pandas.api.types\nimport sklearn.metrics\n\n\nclass ParticipantVisibleError(Exception):\n    pass\n\n\ndef get_condition(full_location: str) -> str:\n    # Given an input like spinal_canal_stenosis_l1_l2 extracts 'spinal'\n    for injury_condition in ['spinal', 'foraminal', 'subarticular']:\n        if injury_condition in full_location:\n            return injury_condition\n    raise ValueError(f'condition not found in {full_location}')\n    \n    \ndef score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        any_severe_scalar: float\n    ) -> float:\n    '''\n    Pseudocode:\n    1. Calculate the sample weighted log loss for each medical condition:\n    2. Derive a new any_severe label.\n    3. Calculate the sample weighted log loss for the new any_severe label.\n    4. Return the average of all of the label group log losses as the final score, normalized for the number of columns in each group.\n       This mitigates the impact of spinal stenosis having only half as many columns as the other two conditions.\n    '''\n\n    target_levels = ['normal_mild', 'moderate', 'severe']\n\n    # Run basic QC checks on the inputs\n    if not pandas.api.types.is_numeric_dtype(submission[target_levels].values):\n        raise ParticipantVisibleError('All submission values must be numeric')\n\n    if not np.isfinite(submission[target_levels].values).all():\n        raise ParticipantVisibleError('All submission values must be finite')\n\n    if solution[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All labels must be at least zero')\n    if submission[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All predictions must be at least zero')\n\n    solution['study_id']  = solution['row_id'].apply(lambda x: x.split('_')[0])\n    solution['location']  = solution['row_id'].apply(lambda x: '_'.join(x.split('_')[1:]))\n    solution['condition'] = solution['row_id'].apply(get_condition)\n\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n    assert sorted(submission.columns) == sorted(target_levels)\n\n    submission['study_id']  = solution['study_id']\n    submission['location']  = solution['location']\n    submission['condition'] = solution['condition']\n\n    condition_losses  = []\n    condition_weights = []\n    for condition in ['spinal', 'foraminal', 'subarticular']:\n        condition_indices = solution.loc[solution['condition'] == condition].index.values\n        condition_loss = sklearn.metrics.log_loss(\n            y_true=solution.loc[condition_indices, target_levels].values,\n            y_pred=submission.loc[condition_indices, target_levels].values,\n            sample_weight=solution.loc[condition_indices, 'sample_weight'].values\n        )\n        condition_losses.append(condition_loss)\n        condition_weights.append(1 / solution.loc[condition_indices, 'location'].nunique())\n\n        \n    any_severe_spinal_labels      = pd.Series(solution.loc[solution    ['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_weights     = pd.Series(solution.loc[solution    ['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\n    any_severe_spinal_predictions = pd.Series(submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_loss = sklearn.metrics.log_loss(\n        y_true=any_severe_spinal_labels,\n        y_pred=any_severe_spinal_predictions,\n        sample_weight=any_severe_spinal_weights\n    )\n    condition_losses.append(any_severe_spinal_loss)\n    condition_weights.append(any_severe_scalar)\n    return np.average(condition_losses, weights=condition_weights)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model #1 Simple frequencies","metadata":{}},{"cell_type":"code","source":"df_test_desc","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_main.head(3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_unpivoted = df_train_main.melt(id_vars='study_id', var_name='condition', value_name='status')\ndf_unpivoted.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_unpivoted = df_train_main.melt(id_vars='study_id', var_name='condition', value_name='status')\ndf_unpivoted[df_unpivoted.study_id ==4003253]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frequency_table = df_unpivoted.groupby('condition')['status'].value_counts(normalize=True).unstack(fill_value=0)\nfrequency_table = frequency_table.reset_index()\nfrequency_table.rename(columns={ 'Normal/Mild': 'normal_mild', 'Moderate': 'moderate', 'Severe': 'severe'}, inplace=True)\nfrequency_table.round(2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub['condition'] = df_sub['row_id'].str.extract(r'_(.*)')\ndf_sub.round(2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# freq's of all stenosis\ndf_sub = pd.merge(df_sub[['row_id', 'condition']],frequency_table, on='condition', how='inner')[['row_id', 'normal_mild', 'moderate', 'severe']]\ndf_sub.round(2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# setup baseline model just using observed frequencies\n\n# Define the path to  test images directory\ntest_images_dir = \"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_images\"\n\n\n# Get all unique IDs from the filenames in the test images directory\ntest_ids = [filename.split('.')[0] for filename in os.listdir(test_images_dir)]\ntest_ids =[4003252, 4003253] #temporary, for debugging\nunique_ids = list(set(test_ids))\n\n# Generate the row_ids needed for submission by repeating each unique_id for each condition from df_train_main\nconditions = ['left_neural_foraminal_narrowing_l1_l2',\n              'left_neural_foraminal_narrowing_l2_l3',\n              'left_neural_foraminal_narrowing_l3_l4',\n              'left_neural_foraminal_narrowing_l4_l5',\n              'left_neural_foraminal_narrowing_l5_s1',\n              'left_subarticular_stenosis_l1_l2',\n              'left_subarticular_stenosis_l2_l3',\n              'left_subarticular_stenosis_l3_l4',\n              'left_subarticular_stenosis_l4_l5',\n              'left_subarticular_stenosis_l5_s1',\n              'right_neural_foraminal_narrowing_l1_l2',\n              'right_neural_foraminal_narrowing_l2_l3',\n              'right_neural_foraminal_narrowing_l3_l4',\n              'right_neural_foraminal_narrowing_l4_l5',\n              'right_neural_foraminal_narrowing_l5_s1',\n              'right_subarticular_stenosis_l1_l2',\n              'right_subarticular_stenosis_l2_l3',\n              'right_subarticular_stenosis_l3_l4',\n              'right_subarticular_stenosis_l4_l5',\n              'right_subarticular_stenosis_l5_s1',\n              'spinal_canal_stenosis_l1_l2',\n              'spinal_canal_stenosis_l2_l3',\n              'spinal_canal_stenosis_l3_l4',\n              'spinal_canal_stenosis_l4_l5',\n              'spinal_canal_stenosis_l5_s1']\n\nrow_ids = [f\"{id}_{condition}\" for id in unique_ids for condition in conditions]\n\n# Create DataFrame\ndf_submission = pd.DataFrame(row_ids, columns=['row_id'])\n# df_submission['normal_mild'] = 0.333333\n# df_submission['moderate']    = 0.333333\n# df_submission['severe']      = 0.333333\n\ndf_submission['normal_mild'] = 0.4242\ndf_submission['moderate']    = 0.3031\ndf_submission['severe']      = 0.2727\n\n\n# df_submission['normal_mild'] = 0.81\n# df_submission['moderate'] = 0.14\n# df_submission['severe'] = 0.05\n#df_submission","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model #2 XGBoost without tuning","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\ntest_series = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_series_descriptions.csv')\ndf_train = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv')\n\ndf_train_melted = df_train.melt(id_vars=['study_id'], var_name='condition_level', value_name='severity')\ndf_train_melted[['condition', 'level']] = df_train_melted['condition_level'].str.rsplit('_', n=2, expand=True).iloc[:, 1:]\n\nle_severity = LabelEncoder()\ndf_train_melted['severity_encoded'] = le_severity.fit_transform(df_train_melted['severity'])\n\nX_train = df_train_melted[['study_id', 'condition', 'level']]\ny_train = df_train_melted['severity_encoded']\n\nX_train = pd.get_dummies(X_train, columns=['condition', 'level'])\ntest_rows = []\nfor _, row in test_series.iterrows():\n    for condition in ['left_neural_foraminal_narrowing', 'right_neural_foraminal_narrowing', 'left_subarticular_stenosis', 'right_subarticular_stenosis', 'spinal_canal_stenosis']:\n        for level in ['l1_l2', 'l2_l3', 'l3_l4', 'l4_l5', 'l5_s1']:\n            test_rows.append({\n                'study_id': row['study_id'],\n                'condition': condition,\n                'level': level\n            })\n\nX_test = pd.DataFrame(test_rows)\nX_test = pd.get_dummies(X_test, columns=['condition', 'level'])\nX_test = X_test.reindex(columns=X_train.columns, fill_value=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from xgboost import XGBClassifier\n\nmodel = XGBClassifier(random_state=1)\nmodel.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions_proba = model.predict_proba(X_test)\npredictions_df = pd.DataFrame(predictions_proba, columns=le_severity.classes_)\npredictions_df['study_id'] = X_test['study_id'].values\npredictions_df['condition_level'] = X_test.index.map(lambda idx: f\"{test_rows[idx]['condition']}_{test_rows[idx]['level']}\")\npredictions_df['row_id'] = predictions_df['study_id'].astype(str) + '_' + predictions_df['condition_level']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div  style=\"color:#ff0a6c;  border:#0014ff solid; font-weight:bold; font-size:120%; text-align:center;padding:12.0px; background:#000000\">6. SUBMISSION</div>","metadata":{}},{"cell_type":"code","source":"df_sub1 = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/sample_submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the DataFrame to a CSV file for submission\ndf_submission.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normal_mild_value = predictions_df['Normal/Mild'].iloc[0]\nmoderate_value = predictions_df['Moderate'].iloc[0]\nsevere_value = predictions_df['Severe'].iloc[0]\n\n\ndf_sub['normal_mild'] = normal_mild_value/2\ndf_sub['moderate'] =moderate_value*2\ndf_sub['severe'] = 1-(normal_mild_value/2 +moderate_value*2)- severe_value\n\n\nnormal_mild_value = 0.36\nmoderate_value = 0.426666667\nsevere_value = 0.213333333\n\ndf_sub['normal_mild'] = normal_mild_value\ndf_sub['moderate']    = moderate_value\ndf_sub['severe']      = severe_value\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.sample(4)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div  style=\"color:blue;   font-size:120%; text-align:center;padding:12.0px; background:#ffffff\">Thanks for viewing my work.   If you like it give feedback to improve the notebook.     Have a beautiful day my friend! </div>","metadata":{}}]}