{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"dockerImageVersionId":30732,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Lumbar Spine Degenerative Condition Classification Explantory Data Analysis and Data Creation I\n\n<div align=\"center\">\n    <img src=\"https://i.ibb.co/WKDHCCC/RSNA.png\">\n</div>\n\nIn this competition, we aim to develop AI models that can accurately classify degenerative conditions of the lumbar spine using MRI images. Specifically, the objective is to create models that can simulate a radiologist's performance in diagnosing five key lumbar spine degenerative conditions: Left Neural Foraminal Narrowing, Right Neural Foraminal Narrowing, Left Subarticular Stenosis, Right Subarticular Stenosis, and Spinal Canal Stenosis.The goal of this project is to develop AI models to identify and classify degenerative conditions affecting the lumbar spine using MRI scans annotated by spine radiology specialists. Here’s a structured approach to tackle this project:\n\nThis guide will walk you through the data and EDA neccessary to know how to handle the data.\n\n**Did you know:**: This notebook is backend-agnostic? Which means it supports TensorFlow, PyTorch, and JAX backends. However, the best performance can be achieved with `JAX`. Explore further details on [Keras](https://keras.io/keras_3/).\n\n\n\nBy participating in this challenge, we will contribute to the advancement of medical imaging and diagnostic radiology, potentially impacting patient care and treatment outcomes positively. Let's get started on building powerful AI models to enhance the detection and classification of lumbar spine degenerative conditions.\n\n\n\n\n\nIn the this notebook we export datasets used.\n\n\n### My other Notebooks\n- [RNSA | EDA & Dataset Creation I ](https://www.kaggle.com/code/archie40004/rnsa-eda-dataset-creation-i) <- you're reading now\n- [RNSA | EDA & Dataset Creation II ](https://www.kaggle.com/code/archie40004/rnsa-eda-dataset-creation-ii) \n- [RSNA 2024 | RSNA | PreP & Modelling, Training ](https://www.kaggle.com/code/archie40004/rsna-prep-modelling/) ","metadata":{"_uuid":"e2bf8a58-9cb3-4415-9207-847197b6c51a","_cell_guid":"0e092c75-0655-494c-9fcf-722803ac2009","trusted":true}},{"cell_type":"markdown","source":"# 📚 | Import Libraries","metadata":{"_uuid":"d1f11dba-b6b2-40c2-b40b-92c0b83d2454","_cell_guid":"5b3fe4fd-e175-4039-934c-caaf016e18cc","trusted":true}},{"cell_type":"code","source":"!pip install dask\n","metadata":{"execution":{"iopub.status.busy":"2024-07-29T06:53:31.808886Z","iopub.execute_input":"2024-07-29T06:53:31.810410Z","iopub.status.idle":"2024-07-29T06:54:08.151141Z","shell.execute_reply.started":"2024-07-29T06:53:31.810367Z","shell.execute_reply":"2024-07-29T06:54:08.149688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Standard\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport warnings # warning handling\nwarnings.filterwarnings('ignore')\nimport glob\nimport time\nimport collections\nimport os\nimport random\n\n# Plots\nimport matplotlib.pyplot as plt\nfrom matplotlib import animation, rc\nimport plotly.express as px\nimport seaborn as sns\nimport plotly.express as px\nimport cv2\n\n# Dicom\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\nfrom PIL import Image\nimport glob\nfrom concurrent.futures import ProcessPoolExecutor\n\nimport re\nimport pydicom\nimport os\nimport pandas as pd\nimport numpy as np\nimport re\nfrom tqdm import tqdm\nfrom PIL import Image\nimport glob\nfrom concurrent.futures import ProcessPoolExecutor\nimport pickle\nimport pydicom\nimport numpy as np\nimport cv2\nimport zlib\nimport pickle\nimport pandas as pd\nimport os\nimport re\nfrom tqdm import tqdm\nfrom concurrent.futures import ProcessPoolExecutor\nimport bz2\nimport pydicom as dicom\nimport matplotlib.patches as patches\n\nimport numpy as np\nimport pandas as pd\nimport pydicom\nfrom PIL import Image\nfrom concurrent.futures import ProcessPoolExecutor\nimport bz2\nimport pickle\nimport gc\nimport os\nimport re\nimport glob\nfrom tqdm import tqdm","metadata":{"_uuid":"251e71df-a8a1-45c7-ad31-3daaca834f6c","_cell_guid":"12c05666-08e1-41f2-87c9-07adba1b310a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:08.153462Z","iopub.execute_input":"2024-07-29T06:54:08.153822Z","iopub.status.idle":"2024-07-29T06:54:10.531024Z","shell.execute_reply.started":"2024-07-29T06:54:08.153788Z","shell.execute_reply":"2024-07-29T06:54:10.529943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ♻️ | Reproducibility \nSets value for random seed to produce similar result in each run.","metadata":{"_uuid":"ae5e8948-d58f-4387-86f3-4e7094237acb","_cell_guid":"d2655b33-af9c-4cbb-ba74-9fd248a920c7","trusted":true}},{"cell_type":"code","source":"def set_random_seed(seed: int = 42, deterministic: bool = False):\n    random.seed(seed)\n    np.random.seed(seed)\n    os.environ[\"PYTHONHASHSEED\"] = str(seed) \n\nset_random_seed(42)\n","metadata":{"_uuid":"b166ec8b-5ab9-4d01-9bf2-d67185402136","_cell_guid":"75c185fe-3ca3-4d16-898f-2be5d13be3fa","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:10.532460Z","iopub.execute_input":"2024-07-29T06:54:10.533011Z","iopub.status.idle":"2024-07-29T06:54:10.540482Z","shell.execute_reply.started":"2024-07-29T06:54:10.532980Z","shell.execute_reply":"2024-07-29T06:54:10.537975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📁 | Dataset Path","metadata":{"_uuid":"60c9a4f9-9029-4adc-bc35-d716e8c52dea","_cell_guid":"26875908-5ebc-49a8-9d8a-fe04e69d7122","trusted":true}},{"cell_type":"code","source":"Path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification'\ntrain_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/'\ntest_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_images/'\n\ndf_train_main = pd.read_csv(Path+'/train.csv')\ndf_train_label = pd.read_csv(Path+'/train_label_coordinates.csv')\ndf_train_desc = pd.read_csv(Path+'/train_series_descriptions.csv')\ndf_test_desc = pd.read_csv(Path+'/test_series_descriptions.csv')\n","metadata":{"_uuid":"308854c6-a3a6-4db9-a3a8-8a65b53ec580","_cell_guid":"36ba1364-73a6-46d2-953f-38a17ab8685d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:10.544018Z","iopub.execute_input":"2024-07-29T06:54:10.544505Z","iopub.status.idle":"2024-07-29T06:54:10.728378Z","shell.execute_reply.started":"2024-07-29T06:54:10.544457Z","shell.execute_reply":"2024-07-29T06:54:10.727162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📖 | Meta Data\n\nThe dataset comprises MRI scans of the lumbar spine annotated by spine radiology specialists. The goal is to develop AI models to identify and classify degenerative conditions affecting the lumbar spine. The training dataset includes `2,000` MR studies with annotations, and the test set contains an undisclosed number of studies to be evaluated.\n\n## Files\n\n### `train.csv`\n- `study_id`: Unique identifier for each MR study. Each study may include multiple series of images.\n- `spinal_canal_stenosis_[l1_l2, l2_l3, l3_l4, l4_l5, l5_s1]`: Severity labels for spinal canal stenosis at each vertebral level.\n- `left_neural_foraminal_narrowing_[l1_l2, l2_l3, l3_l4, l4_l5, l5_s1]`: Severity labels for left neural foraminal narrowing at each vertebral level.\n- `right_neural_foraminal_narrowing_[l1_l2, l2_l3, l3_l4, l4_l5, l5_s1]`: Severity labels for right neural foraminal narrowing at each vertebral level.\n- `left_subarticular_stenosis_[l1_l2, l2_l3, l3_l4, l4_l5, l5_s1]`: Severity labels for left subarticular stenosis at each vertebral level.\n- `right_subarticular_stenosis_[l1_l2, l2_l3, l3_l4, l4_l5, l5_s1]`: Severity labels for right subarticular stenosis at each vertebral level.\n\n### `train_label_coordinates.csv`\n- `study_id`: Unique identifier for each MR study.\n- `series_id`: Identifier for the imagery series within each study.\n- `instance_number`: The image's order number within the 3D stack.\n- `condition`: The specific condition annotated (e.g., spinal canal stenosis, neural foraminal narrowing, subarticular stenosis).\n- `level`: The vertebral level related to the annotation (e.g., l3_l4).\n- `x`: The x-coordinate for the center of the labeled area.\n- `y`: The y-coordinate for the center of the labeled area.\n\n### `sample_submission.csv`\n- `row_id`: Unique identifier for each prediction, formatted as `study_id_condition_level` (e.g., `12345_spinal_canal_stenosis_l3_l4`).\n- `normal_mild`: Probability prediction for the condition being Normal/Mild.\n- `moderate`: Probability prediction for the condition being Moderate.\n- `severe`: Probability prediction for the condition being Severe.\n\n### `train/test_images/[study_id]/[series_id]/[instance_number].dcm`\n- Contains the MRI imagery data in DICOM format.\n\n### `train/test_series_descriptions.csv`\n- `study_id`: Unique identifier for each MR study.\n- `series_id`: Identifier for the imagery series within each study.\n- `series_description`: Description of the scan's orientation (e.g., Axial T2, Sagittal T1, Sagittal T2/STIR).\n\n> Note that each study may contain multiple series and images, providing a comprehensive view of the lumbar spine. The annotations guide the identification of specific degenerative conditions and their severity at different vertebral levels.","metadata":{"_uuid":"e4e3a188-9258-4d47-94d5-58c2341c6f9a","_cell_guid":"941f408f-2b29-4011-b836-b8eece09c1c6","trusted":true}},{"cell_type":"markdown","source":"The dataset includes MRI scans of the lumbar spine with annotations for various degenerative conditions. The key components of the dataset are:\n- **MRI Images* : Stored in DICOM format, organized by study ID, series ID, and instance number.\n- Annotations: Severity labels for conditions such as spinal canal stenosis, neural foraminal narrowing, and subarticular stenosis at different vertebral levels (L1/L2 to L5/S1).\n- Coordinates: X and Y coordinates for the center of the labeled areas in the images.\n- Severity Scores: Labels indicating the severity of conditions (Normal/Mild, Moderate, Severe).","metadata":{"_uuid":"da7de6c6-3eae-43d8-95e5-2a6507de1e28","_cell_guid":"4f10fc94-068d-4b03-ae77-01b6e85429ef","trusted":true}},{"cell_type":"markdown","source":"# 🎨 | Exploratory Data Analysis (EDA)","metadata":{"_uuid":"5263b283-df48-4604-a472-7d8b30d52a53","_cell_guid":"8363b7a4-7238-4369-ba7a-0a149deae319","trusted":true}},{"cell_type":"code","source":"# structure of train\ndf_train_main.info()","metadata":{"_uuid":"5fea1ff1-d0cc-4d4e-9cc1-71b831da3506","_cell_guid":"a216321e-7428-44e2-a939-0aea8172d408","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:10.729701Z","iopub.execute_input":"2024-07-29T06:54:10.730053Z","iopub.status.idle":"2024-07-29T06:54:10.768876Z","shell.execute_reply.started":"2024-07-29T06:54:10.730021Z","shell.execute_reply":"2024-07-29T06:54:10.767695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_label.info()","metadata":{"_uuid":"ea4348b5-8fbe-4ff6-a5e6-dfb906f7e22c","_cell_guid":"bf168ea9-562f-492d-be08-ef22e73f6975","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:10.770528Z","iopub.execute_input":"2024-07-29T06:54:10.770896Z","iopub.status.idle":"2024-07-29T06:54:10.795189Z","shell.execute_reply.started":"2024-07-29T06:54:10.770866Z","shell.execute_reply":"2024-07-29T06:54:10.793961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_desc.info()","metadata":{"_uuid":"97500de1-1821-4f77-ad8b-4c686e7451d2","_cell_guid":"78bbd38f-97eb-4daa-9448-ddf2ba129493","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:10.796747Z","iopub.execute_input":"2024-07-29T06:54:10.797134Z","iopub.status.idle":"2024-07-29T06:54:10.810178Z","shell.execute_reply.started":"2024-07-29T06:54:10.797092Z","shell.execute_reply":"2024-07-29T06:54:10.808812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# look at categories\nfor f in ['instance_number','condition','level']:\n    print(df_train_label[f].value_counts())\n    print('-'*50);print();","metadata":{"_uuid":"17552555-fcf5-4a4f-a09a-fb1b5b9e1011","_cell_guid":"7390289e-80d9-4d19-b8b3-49ee9c54f960","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:10.811569Z","iopub.execute_input":"2024-07-29T06:54:10.811914Z","iopub.status.idle":"2024-07-29T06:54:10.848777Z","shell.execute_reply.started":"2024-07-29T06:54:10.811884Z","shell.execute_reply":"2024-07-29T06:54:10.847555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# join first two tables\ndf_train_step_1 = pd.merge(left=df_train_label, right=df_train_main, how='left', on='study_id').reset_index(drop=True)\n# join with third table\ndf_train = pd.merge(left=df_train_step_1, right=df_train_desc, how='left', on=['study_id', 'series_id']).reset_index(drop=True)\ndf_train.head()","metadata":{"_uuid":"8d64dcb8-d6da-4bcb-87c6-b4f006c13263","_cell_guid":"2773eaee-4cd6-4479-9b72-738c218c2611","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:10.850305Z","iopub.execute_input":"2024-07-29T06:54:10.851290Z","iopub.status.idle":"2024-07-29T06:54:11.094082Z","shell.execute_reply.started":"2024-07-29T06:54:10.851243Z","shell.execute_reply":"2024-07-29T06:54:11.093000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convert identifiers to categorical\ndf_train.study_id = df_train.study_id.astype('category')\ndf_train.series_id = df_train.series_id.astype('category')\n\n\ndf_test_desc.study_id = df_test_desc.study_id.astype('category')\ndf_test_desc.series_id = df_test_desc.series_id.astype('category')\n","metadata":{"_uuid":"8a7067aa-3ae2-4084-99fd-c44465a9fa2d","_cell_guid":"ff06b9f1-d1bf-4abd-8d9a-6a4bef30d377","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:11.098444Z","iopub.execute_input":"2024-07-29T06:54:11.098831Z","iopub.status.idle":"2024-07-29T06:54:11.112116Z","shell.execute_reply.started":"2024-07-29T06:54:11.098797Z","shell.execute_reply":"2024-07-29T06:54:11.110695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Coordinates distributions","metadata":{"_uuid":"4159280c-fc9d-443e-aaae-427fbc1a7eb6","_cell_guid":"bb968507-5db1-4013-951b-de3ca639703c","trusted":true}},{"cell_type":"code","source":"# Create a scatter plot with Plotly\nfig = px.scatter(df_train, x='x', y='y', opacity=0.6)\n\n# Customize the layout\nfig.update_layout(\n    xaxis_title='X Coordinate',\n    yaxis_title='Y Coordinate',\n    xaxis=dict(showgrid=True),\n    yaxis=dict(showgrid=True)\n)\n\n# Show the plot\nfig.show()","metadata":{"_uuid":"9d97e348-3536-476a-8ad7-f69f98fdab3b","_cell_guid":"8c8bc406-d0e5-4340-b665-0bdbe8537ce2","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:11.113996Z","iopub.execute_input":"2024-07-29T06:54:11.114424Z","iopub.status.idle":"2024-07-29T06:54:13.195385Z","shell.execute_reply.started":"2024-07-29T06:54:11.114391Z","shell.execute_reply":"2024-07-29T06:54:13.193732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot coordinate distributions with colored by condition","metadata":{"_uuid":"c2710ded-7eba-4683-8aa0-02cca7b944bc","_cell_guid":"7271bd0c-30f8-4820-be6c-56ba780bdd6d","trusted":true}},{"cell_type":"code","source":"\n# Define custom color palette\ncustom_palette = ['#4600c0',  '#00a419', '#111111', '#ffff00', '#ff0000']\n\n# Create the plot with Plotly\nfig = px.scatter(df_train, x='x', y='y', color='condition', \n                 color_discrete_sequence=custom_palette, \n                 opacity=0.6, \n                 labels={'x': 'X Coordinate', 'y': 'Y Coordinate'})\n\n# Update the layout for better readability\nfig.update_layout(\n    xaxis=dict(range=[0, df_train['x'].max()]),\n    yaxis=dict(range=[0, df_train['y'].max()]),\n    legend=dict(title='Condition', x=1.05, y=1),\n    width=7*100, height= 7*100\n)\n\n# Show the plot\nfig.show()","metadata":{"_uuid":"5c183a9c-389e-4a3c-a7a1-8088459378cd","_cell_guid":"102c8375-7e13-4454-a92d-9b0769e211a1","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:13.196935Z","iopub.execute_input":"2024-07-29T06:54:13.197330Z","iopub.status.idle":"2024-07-29T06:54:13.366435Z","shell.execute_reply.started":"2024-07-29T06:54:13.197288Z","shell.execute_reply":"2024-07-29T06:54:13.364968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot coordinates distributions with colored by level","metadata":{"_uuid":"9c8c2edf-6688-484a-a3cb-8a6eec2263c8","_cell_guid":"77d379b3-a2ff-44e7-b280-70ee7802f634","trusted":true}},{"cell_type":"code","source":"\n# Define custom color palette\ncustom_palette = ['#4600c0',  '#00a419', '#111111', '#ffff00', '#ff0000']\n\n# Define plot size\nfigs_x = 5  \nfigs_y = 5\n\n# Create the plot with Plotly\nfig = px.scatter(df_train, x='x', y='y', color='level', \n                 color_discrete_sequence=custom_palette, \n                 opacity=0.4,\n                 labels={'x': 'X Coordinate', 'y': 'Y Coordinate'})\n\n# Update the layout for better readability\nfig.update_layout(\n    xaxis=dict(range=[0, df_train['x'].max()]),\n    yaxis=dict(range=[0, df_train['y'].max()]),\n    legend=dict(title='Level', x=1.05, y=1, traceorder='reversed'),\n    width=figs_x*100, height=figs_y*100\n)\n\n# Show the plot\nfig.show()","metadata":{"_uuid":"ee6024e2-bbba-4394-b9f3-375702a46740","_cell_guid":"b980e48a-b521-4e24-9507-9c6bffc845a3","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:13.367889Z","iopub.execute_input":"2024-07-29T06:54:13.368280Z","iopub.status.idle":"2024-07-29T06:54:13.522520Z","shell.execute_reply.started":"2024-07-29T06:54:13.368241Z","shell.execute_reply":"2024-07-29T06:54:13.520961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot stenosis distributions with colored by MR-split","metadata":{"_uuid":"7a8f1301-4500-4d83-be26-430b722035c2","_cell_guid":"4c69c6e1-6623-4efb-93da-583a3fa1e00d","trusted":true}},{"cell_type":"code","source":"\n# Define custom color palette\ncustom_palette = ['#000000', '#4600c0', '#ff0000']\n\n# Define plot size\nfigs_x = 5  \nfigs_y = 5\n\n# Create the plot with Plotly\nfig = px.scatter(df_train, x='x', y='y', color='series_description', \n                 color_discrete_sequence=custom_palette, \n                 opacity=0.4,\n                 labels={'x': 'X Coordinate', 'y': 'Y Coordinate'})\n\n# Update the layout for better readability\nfig.update_layout(\n    xaxis=dict(range=[0, df_train['x'].max()]),\n    yaxis=dict(range=[0, df_train['y'].max()]),\n    legend=dict(title='Series Description', x=1.05, y=1),\n    width=figs_x*100, height=figs_y*100\n)\n\n# Show the plot\nfig.show()","metadata":{"_uuid":"19af07f5-3eb0-4a04-9b4a-370a70e2be37","_cell_guid":"726f389a-1f47-46bc-bfff-9becb47343cf","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:13.524016Z","iopub.execute_input":"2024-07-29T06:54:13.524433Z","iopub.status.idle":"2024-07-29T06:54:13.654194Z","shell.execute_reply.started":"2024-07-29T06:54:13.524401Z","shell.execute_reply":"2024-07-29T06:54:13.652881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_features =   df_train.series_description.unique()\ngs=200\nn=2 # num of columns\na=0 \nk=1;\ncolorlabels = 'darkblue'\n\nLabel_size = 10 # Size font of xy labels\nTitle_size = 20 # Size font of Title\nfigs_x=13   \nfigs_y=5\nplt.figure(figsize=(figs_x, figs_y))    \nplt.suptitle(\"xy-Density of cases in plane-split\", fontsize=Title_size+2, fontweight='bold', y=1.015)\nfor i in numerical_features:\n        \n        d = df_train[df_train.series_description == i]\n        plt.subplot(1,n, k)\n        #sns.jointplot(data=d, x='x', y='y', color='white', alpha=0.25)\n        plt.hexbin(data=d, x='x', y='y',  gridsize=gs, cmap='CMRmap', bins='log', alpha = 1)\n        \n        #plt.colorbar(label='count in bin')\n        plt.colorbar().set_label(label='count in bin',size=10, color  = 'grey')\n        #plt.colorbar(size=8)\n        \n        plt.tick_params(axis='x', labelsize=Label_size)\n        plt.tick_params(axis='y', labelsize=Label_size)\n        \n        plt.xlim([0, df_train.x.max()])\n        plt.ylim([0, df_train.y.max()])\n        plt.xlabel(f'x', fontsize=Label_size, color = colorlabels)\n        plt.ylabel(f'y', fontsize=Label_size, color = colorlabels) \n        \n        plt.title(f'{i}', color='black', fontsize=Title_size)\n        k=k+1\n        if k == (n+1):    \n            k=1\n            plt.show()\n            plt.figure(figsize=(figs_x, figs_y))","metadata":{"_uuid":"a1b20c9e-fedb-4e4a-8ec9-506be18f69e2","_cell_guid":"7ada9407-fc46-4339-a4ea-cfc403857d0a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:13.655793Z","iopub.execute_input":"2024-07-29T06:54:13.656149Z","iopub.status.idle":"2024-07-29T06:54:16.597013Z","shell.execute_reply.started":"2024-07-29T06:54:13.656119Z","shell.execute_reply":"2024-07-29T06:54:16.595733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Target distribution","metadata":{"_uuid":"d8d5ea4a-b0b7-4b78-afdc-89d4a21dadf7","_cell_guid":"75ae1e06-e87a-4e88-ab41-186c12d95a6b","trusted":true}},{"cell_type":"code","source":"train_label_df = df_train_label.copy()\ntrain_data_df = df_train_main.copy()\ntrain_label_df['new_col'] = df_train_label['condition'].str.lower().str.replace(' ', '_') +  '_' + df_train_label['level'].str.lower().str.replace('/', '_')\n\n# Step 2: Merge the values from train_data_df based on study_id and the newly created column names\ndef get_target_value(row):\n    study_id = row['study_id']\n    new_col = row['new_col']\n    return train_data_df[train_data_df['study_id'] == study_id][new_col].values[0]\n\ntrain_label_df['target'] = train_label_df.apply(get_target_value, axis=1)\n\n# Drop the 'new_col' column \ntrain_label_df.drop(columns=['new_col'], inplace=True)\n\n# Copy the train_label_df to new DataFrame\nfinal_train_df = train_label_df.copy()\n\n# Drop the 'new_col' column \ntrain_label_df.drop(columns=['target'], inplace=True)\n\n#final_train_df.to_csv('final_train.csv')\nfinal_train_df.head()","metadata":{"_uuid":"0344a726-b46b-4a6c-bb12-f4cf3cceb518","_cell_guid":"2241e5fe-09ab-4301-9e48-7bb47327ab65","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:16.598640Z","iopub.execute_input":"2024-07-29T06:54:16.598984Z","iopub.status.idle":"2024-07-29T06:54:32.429431Z","shell.execute_reply.started":"2024-07-29T06:54:16.598951Z","shell.execute_reply":"2024-07-29T06:54:32.427041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plotting the distribution of the target variable:","metadata":{"_uuid":"22b056f7-84e6-41cb-a9a9-21f6bb6e578b","_cell_guid":"650200db-5c86-4699-a617-611be2874d6c","trusted":true}},{"cell_type":"code","source":"# Define custom color palette\np = ['#d0d0d0', '#ffba07', '#ff0000']\n\n# Create the plot with Plotly\nfig = px.histogram(final_train_df, x='target', color='target', \n                   color_discrete_sequence=p,\n                   labels={'target': 'Severity'})\n\n# Update the layout for better readability\nfig.update_layout(\n    xaxis_title='Severity Level',\n    yaxis_title='Count',\n    legend_title='Severity',\n    title_font_size=15,\n    title_font_family='Arial',\n    title_font_color='blue',\n    title_x=0.5,\n    legend=dict(x=0.8, y=0.9, traceorder='normal', bgcolor='rgba(255, 255, 255, 0.8)', bordercolor='black')\n)\n\n# Show the plot\nfig.show()","metadata":{"_uuid":"fa504add-78de-4503-a874-05d11365ea7b","_cell_guid":"be016568-791d-46c1-9e10-7a8be7d5f92f","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.430696Z","iopub.status.idle":"2024-07-29T06:54:32.431143Z","shell.execute_reply.started":"2024-07-29T06:54:32.430935Z","shell.execute_reply":"2024-07-29T06:54:32.430953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plotting the distribution of the target variable in L/L splits:","metadata":{"_uuid":"6fb55532-d4ae-4d38-963c-2c16653f7597","_cell_guid":"09e46a6f-9ee6-4f0c-a980-4089fb8b2b7b","trusted":true}},{"cell_type":"code","source":"import plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n# Define custom color palette\ncolours = {\"Normal/Mild\": \"#273c75\", \"Moderate\": \"#44bd32\", \"Severe\": \"#ff0000\"}\n\n# Initialize subplots\nfig = make_subplots(rows=1, cols=3, subplot_titles=(\"Foraminal Distribution\", \"Subarticular Distribution\", \"Canal Distribution\"))\n\n# Define categories to plot\ncategories = ['foraminal', 'subarticular', 'canal']\n\n# Iterate over categories and create bar plots\nfor idx, d in enumerate(categories):\n    diagnosis = list(filter(lambda x: x.find(d) > -1, df_train_main.columns))\n    dff = df_train_main[diagnosis]\n    value_counts = dff.apply(pd.value_counts).fillna(0).T\n    \n    for severity in value_counts.columns:\n        fig.add_trace(go.Bar(\n            x=value_counts.index,\n            y=value_counts[severity],\n            name=severity,\n            marker_color=colours[severity],\n            showlegend=(idx == 0)  # Only show legend for the first subplot\n        ), row=1, col=idx + 1)\n\n# Update layout\nfig.update_layout(\n    title_text=\"Severity Distribution by Condition\",\n    barmode='stack',\n    legend_title_text='Severity',\n    height=500,\n    width=1200\n)\n\n# Show the plot\nfig.show()","metadata":{"_uuid":"7e32ec9b-256b-4671-a29e-b6cad35792a5","_cell_guid":"36ce5a50-89aa-4b6d-9a6a-c8fdf1e5b220","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.432903Z","iopub.status.idle":"2024-07-29T06:54:32.433330Z","shell.execute_reply.started":"2024-07-29T06:54:32.433106Z","shell.execute_reply":"2024-07-29T06:54:32.433121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📷 | Images","metadata":{"_uuid":"3ba0773d-01ef-4ec6-a922-c04463aa2b3a","_cell_guid":"d3849cd3-96b1-41ed-a42b-517a2e26d3dd","trusted":true}},{"cell_type":"markdown","source":"A .dcm file follows the **Digital Imaging and Communications in Medicine** (DICOM) format. It is the standard format used for storing medical images and related metadata. It dates back to 1983, although it has been revised many times.\n\nWe can use the pydicom library to open and explore these files.","metadata":{"_uuid":"f8a94f3c-6314-4f50-9440-9f82fc964000","_cell_guid":"88cae5ef-85d8-4870-844f-0182897ebc03","trusted":true}},{"cell_type":"code","source":"train_label_coordinates=df_train_label","metadata":{"_uuid":"94c38e0c-de82-4fa6-8862-759d825142e1","_cell_guid":"a242e5e8-6ffb-4938-9bd7-dd700cd3cc2c","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.434462Z","iopub.status.idle":"2024-07-29T06:54:32.434845Z","shell.execute_reply.started":"2024-07-29T06:54:32.434660Z","shell.execute_reply":"2024-07-29T06:54:32.434675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_label_coordinates[train_label_coordinates.study_id==4003253]","metadata":{"_uuid":"02f9c37d-916d-4941-a2a8-b4236f2aec77","_cell_guid":"8b6c6d77-f0a5-4942-a704-b6f8d91b35b5","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.436313Z","iopub.status.idle":"2024-07-29T06:54:32.436855Z","shell.execute_reply.started":"2024-07-29T06:54:32.436581Z","shell.execute_reply":"2024-07-29T06:54:32.436604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_label_coordinates[train_label_coordinates.study_id==100206310]","metadata":{"_uuid":"0ae7e21e-928d-4985-9d13-b2dad0e0592d","_cell_guid":"eb710b58-3513-44b2-91dd-01fcdd1e58e4","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.438453Z","iopub.status.idle":"2024-07-29T06:54:32.438980Z","shell.execute_reply.started":"2024-07-29T06:54:32.438705Z","shell.execute_reply":"2024-07-29T06:54:32.438727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualizing MR-images","metadata":{"_uuid":"7161a3bf-6425-41a7-8cb7-da3f9ba1eb93","_cell_guid":"53a6aa86-7ca8-44ad-8cb4-4f0d4069b94a","trusted":true}},{"cell_type":"code","source":"train_label_coordinates['series_description'] = df_train.series_description","metadata":{"_uuid":"d16f19dc-20d3-432e-92bd-46ff69f314a9","_cell_guid":"c11d8215-0625-4136-a108-bea553d0f376","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.440638Z","iopub.status.idle":"2024-07-29T06:54:32.441167Z","shell.execute_reply.started":"2024-07-29T06:54:32.440895Z","shell.execute_reply":"2024-07-29T06:54:32.440917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def mrt(id, ser, inst):\n    lag=20\n    path2 = Path +'/train_images/' + str(id) +'/' + str(ser)+'/' + str(inst) + '.dcm'\n\n    ds = dicom.dcmread(path2)\n    fig, ax = plt.subplots(figsize=(16, 8))\n    from matplotlib.colors import LogNorm \n\n    ax.imshow(ds.pixel_array, cmap ='CMRmap')     # Display the image\n\n    # Create a legend\n    legend_elements = []\n\n    # Plot the coordinates for the current condition\n    ab = train_label_coordinates[(train_label_coordinates.study_id==id) & \n                                          (train_label_coordinates.instance_number==inst)&\n                                         (train_label_coordinates.series_id==ser)]\n\n    a = 25 * max(ds.pixel_array.shape)/640\n    for _, row in ab.iterrows():\n        x, y = row['x'], row['y']\n\n        rect2 = patches.Rectangle((x - a, y - a), 2*a, 2*a, linewidth=2, edgecolor='white', facecolor='none')\n        rect1 = patches.Rectangle((x - a, y - a), 2*a, 2*a, linewidth=2, facecolor='white', alpha = 0.25)\n\n        ax.add_patch(rect2)\n        ax.add_patch(rect1)\n\n        # Add the condition to the legend\n        legend_elements.append(patches.Patch(facecolor='none', edgecolor='r', ))\n\n    # Add title\n    title = f\"{ab.series_description.unique()}, Study: {id}, Series: {ser}, Instance: {inst}\"\n    ax.set_title(title, fontsize=20)\n\n    # Display additional columns:\n    for _, row in ab.iterrows():\n        text = f\"level {row['level']}, {row['condition']}\"\n        ax.text(row['x'] + lag, row['y']+np.random.randint(-15, 15), text, fontsize=10, color='white', verticalalignment='center_baseline')\n    \n    plt.show()","metadata":{"_uuid":"4c30f5aa-c5e0-42cd-82b6-b36ed8c3fa9a","_cell_guid":"7277f4e6-f025-4b60-9c13-46b4e428f609","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.442772Z","iopub.status.idle":"2024-07-29T06:54:32.443327Z","shell.execute_reply.started":"2024-07-29T06:54:32.443026Z","shell.execute_reply":"2024-07-29T06:54:32.443049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #1\nid = 100206310\nser= 1012284084\ninst=8\n\nmrt(id, ser, inst)\n\n#Case #1\nid = 100206310\nser= 1792451510\ninst=8\n\nmrt(id, ser, inst)\n\n#Case #1\nid = 100206310\nser= 2092806862\ninst=8\n\nmrt(id, ser, inst)","metadata":{"_uuid":"26a934a9-03a9-47c7-b9a5-879ea97ba7d5","_cell_guid":"f9a9f5ce-b6c3-4660-ac81-c703ea249f78","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.445563Z","iopub.status.idle":"2024-07-29T06:54:32.446131Z","shell.execute_reply.started":"2024-07-29T06:54:32.445815Z","shell.execute_reply":"2024-07-29T06:54:32.445860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3D MR-slides visualisation","metadata":{"_uuid":"b3573e1a-2b8c-4a1f-ab23-6ee0ea614604","_cell_guid":"6128ab6a-16e9-4313-96ed-7371f4ba87e4","trusted":true}},{"cell_type":"code","source":"def MR3d(id, ser):\n    \n    path_to_folder = Path + \"/train_images/\"+str(id)+'/'+str(ser)\n    def load_dicom(path):\n        dicom = pydicom.read_file(path)\n        data = dicom.pixel_array\n        data = data - np.min(data)\n        if np.max(data) != 0:\n            data = data / np.max(data)\n        data = (data * 255).astype(np.uint8)\n        return data\n\n    rc('animation', html='jshtml')\n\n    def load_dicom(filename):\n        ds = pydicom.dcmread(filename)\n        return ds.pixel_array\n\n    def load_dicom_line(path):\n        t_paths = sorted(\n            glob.glob(os.path.join(path, \"*\")), \n            key=lambda x: int(os.path.splitext(os.path.basename(x))[0].split(\"-\")[-1]),\n        )\n        images = []\n        for filename in t_paths:\n            data = load_dicom(filename)\n            if data.max() == 0:\n                continue\n            images.append(data)\n        return images\n\n    def create_animation(ims):\n        fig = plt.figure(figsize=(6, 6))\n        plt.axis('off')\n        im = plt.imshow(ims[0], cmap=\"CMRmap\")\n        text = plt.text(0.05, 0.05, f'Slide {1}', transform=fig.transFigure, fontsize=16, color='darkblue')\n\n        def animate_func(i):\n            im.set_array(ims[i])\n            return [im]\n        plt.title(f'id = {id}, series = {ser}')\n        \n        plt.close()  \n\n        return animation.FuncAnimation(fig, animate_func, frames=len(ims), interval=1000//10) #24\n\n    images = load_dicom_line(path_to_folder)\n    \n    return create_animation(images)","metadata":{"_uuid":"60ec25b7-5294-4c68-b2f9-82d6c7e629f3","_cell_guid":"0ce2aafa-ee5a-444e-900c-10e959630f9e","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.447679Z","iopub.status.idle":"2024-07-29T06:54:32.448236Z","shell.execute_reply.started":"2024-07-29T06:54:32.447929Z","shell.execute_reply":"2024-07-29T06:54:32.447952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Case #1\n#id = 100206310\nser= 1012284084\n\n#MR3d(id, ser)","metadata":{"_uuid":"7af944e5-ff7a-4081-9ce1-b9d42f41d93a","_cell_guid":"afe18baa-9aab-45a4-b115-7603cc379df5","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-07-29T06:54:32.449551Z","iopub.status.idle":"2024-07-29T06:54:32.450060Z","shell.execute_reply.started":"2024-07-29T06:54:32.449796Z","shell.execute_reply":"2024-07-29T06:54:32.449818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📊 | Dataset Creation","metadata":{"_uuid":"aa3fe299-9298-4094-b34e-b52484f4c034","_cell_guid":"1a75fac3-0110-4952-ab66-51316d6f0c50","trusted":true}},{"cell_type":"code","source":"def reduce_mem_usage(df):\n    \"\"\" Function to reduce the memory usage of a DataFrame by downcasting data types \"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print(f\"Memory usage of dataframe is {start_mem:.2f} MB\")\n\n    for col in df.columns:\n        if col == 'image':  # Skip the 'image' column\n            continue\n\n        col_type = df[col].dtype\n\n        if col_type != object and not pd.api.types.is_categorical_dtype(df[col]):\n            c_min = df[col].min()\n            c_max = df[col].max()\n\n            if str(col_type).startswith('int'):\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n        elif pd.api.types.is_categorical_dtype(df[col]):\n            if df[col].cat.ordered:\n                df[col] = df[col].cat.as_ordered()\n            else:\n                df[col] = df[col].cat.as_unordered()\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print(f\"Memory usage after optimization is: {end_mem:.2f} MB\")\n    print(f\"Decreased by {100 * (start_mem - end_mem) / start_mem:.1f}%\")\n\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-07-29T06:55:21.606514Z","iopub.execute_input":"2024-07-29T06:55:21.606949Z","iopub.status.idle":"2024-07-29T06:55:21.625483Z","shell.execute_reply.started":"2024-07-29T06:55:21.606916Z","shell.execute_reply":"2024-07-29T06:55:21.623872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport gc\nimport numpy as np\nimport glob\nimport pydicom\nfrom PIL import Image\nfrom concurrent.futures import ProcessPoolExecutor, as_completed\nimport re\nfrom tqdm import tqdm\nimport pandas as pd\n\ndef atoi(text):\n    return int(text) if text.isdigit() else text\n\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\n\ndef process_dicom_image(src_path, study_id, series_id, instance_number, output_dir):\n    try:\n        dicom_data = pydicom.dcmread(src_path)\n        image = dicom_data.pixel_array\n        image = (image - image.min()) / (image.max() - image.min() + 1e-6) * 255\n        image = Image.fromarray(image.astype(np.uint8)).convert('L')\n        \n        # Create output directory if it doesn't exist\n        output_path = os.path.join(output_dir, str(study_id), str(series_id))\n        os.makedirs(output_path, exist_ok=True)\n        \n        # Save the image as PNG\n        image_file_path = os.path.join(output_path, f'{instance_number}.png')\n        image.save(image_file_path)\n    except Exception as e:\n        print(f\"Error processing file {src_path}: {e}\")\n\ndef process_image_wrapper(args):\n    process_dicom_image(*args)\n\ndef collect_tasks(df, image_dir, output_dir):\n    tasks = []\n    st_ids = df['study_id'].unique()\n    for si in st_ids:\n        pdf = df[df['study_id'] == si]\n        for _, row in pdf.iterrows():\n            series_id = row['series_id']\n            img_paths = glob.glob(f'{image_dir}/{si}/{series_id}/*.dcm')\n            img_paths = sorted(img_paths, key=natural_keys)\n            for j, impath in enumerate(img_paths):\n                instance_number = j + 1  # Assuming the instance number starts from 1\n                tasks.append((impath, si, series_id, instance_number, output_dir))\n    return tasks\n\ndef process_all_in_parallel(df, image_dir, output_dir):\n    # Collect all tasks\n    tasks = collect_tasks(df, image_dir, output_dir)\n    \n    # Process all tasks in parallel with progress bar\n    with ProcessPoolExecutor() as executor:\n        futures = {executor.submit(process_image_wrapper, task): task for task in tasks}\n        with tqdm(total=len(futures)) as pbar:\n            for future in as_completed(futures):\n                pbar.update(1)\n    \n    # Clear memory\n    gc.collect()\n\n# Clear memory before starting the process\ngc.collect()\n\n# Paths to data directories\ntrain_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/'\ntest_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_images/'\n\n# Output directories for PNG images\ntrain_output_dir = 'train_png_images'\ntest_output_dir = 'test_png_images'\n\n# Process all train tasks in parallel\nprocess_all_in_parallel(df_train_desc, train_image_dir, train_output_dir)\n\n# Process all test tasks in parallel\nprocess_all_in_parallel(df_test_desc, test_image_dir, test_output_dir)\n","metadata":{"execution":{"iopub.status.busy":"2024-07-29T07:32:26.444870Z","iopub.execute_input":"2024-07-29T07:32:26.445385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nimport os\nimport gc\nimport numpy as np\nimport glob\nimport pydicom\nfrom PIL import Image\nfrom concurrent.futures import ProcessPoolExecutor\nimport re\nfrom tqdm import tqdm\nimport pandas as pd\n\ndef atoi(text):\n    return int(text) if text.isdigit() else text\n\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\n\ndef process_dicom_image(src_path, study_id, series_id, instance_number, output_dir):\n    try:\n        dicom_data = pydicom.dcmread(src_path)\n        image = dicom_data.pixel_array\n        image = (image - image.min()) / (image.max() - image.min() + 1e-6) * 255\n        image = Image.fromarray(image.astype(np.uint8)).convert('L')\n        \n        # Create output directory if it doesn't exist\n        output_path = os.path.join(output_dir, str(study_id), str(series_id))\n        os.makedirs(output_path, exist_ok=True)\n        \n        # Save the image as PNG\n        image_file_path = os.path.join(output_path, f'{instance_number}.png')\n        image.save(image_file_path)\n\n        return {'study_id': study_id, 'series_id': series_id, 'instance_number': instance_number, 'image_path': image_file_path}\n    except Exception as e:\n        print(f\"Error processing file {src_path}: {e}\")\n        return None\n\ndef process_image_wrapper(args):\n    return process_dicom_image(*args)\n\ndef collect_tasks(df, data_type='train'):\n    tasks = []\n    st_ids = df['study_id'].unique()\n    for idx, si in enumerate(tqdm(st_ids, total=len(st_ids))):\n        pdf = df[df['study_id'] == si]\n        for _, row in pdf.iterrows():\n            ds = row['series_description'].replace('/', '_')\n            series_id = row['series_id']\n            img_paths = glob.glob(f'/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/{data_type}_images/{si}/{series_id}/*.dcm')\n            img_paths = sorted(img_paths, key=natural_keys)\n            for j, impath in enumerate(img_paths):\n                instance_number = j + 1  # Assuming the instance number starts from 1\n                tasks.append((impath, si, series_id, instance_number, f'{data_type}_png_images'))\n    return tasks\n\ndef process_and_save(tasks, chunk_size=50000):\n    total_chunks = len(tasks) // chunk_size + 1\n    for i in range(total_chunks):\n        chunk_tasks = tasks[i*chunk_size:(i+1)*chunk_size]\n        with ProcessPoolExecutor() as executor:\n            list(tqdm(executor.map(process_image_wrapper, chunk_tasks), total=len(chunk_tasks)))\n        gc.collect()\n\n# Clear memory before starting the process\ngc.collect()\n\n# Paths to data directories\ntrain_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/'\ntest_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_images/'\n\n# Collect tasks for train and test data\ntrain_tasks = collect_tasks(df_train_desc, data_type='train')\ntest_tasks = collect_tasks(df_test_desc, data_type='test')\n\n# Process and save the train tasks\nprocess_and_save(train_tasks)\n\n# Process and save the test tasks\nprocess_and_save(test_tasks)\n'''","metadata":{"execution":{"iopub.status.busy":"2024-07-29T06:55:24.609394Z","iopub.execute_input":"2024-07-29T06:55:24.609814Z","iopub.status.idle":"2024-07-29T07:04:35.268096Z","shell.execute_reply.started":"2024-07-29T06:55:24.609778Z","shell.execute_reply":"2024-07-29T07:04:35.266121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nimport pandas as pd\nimport os\nimport gc\nimport numpy as np\nimport glob\nimport pydicom\nfrom PIL import Image\nfrom concurrent.futures import ProcessPoolExecutor\nimport re\nimport gzip\nfrom tqdm import tqdm\nfrom io import BytesIO\nimport pickle\n\ndef atoi(text):\n    return int(text) if text.isdigit() else text\n\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\n\ndef process_dicom_image(src_path, study_id, series_id, instance_number):\n    try:\n        dicom_data = pydicom.dcmread(src_path)\n        image = dicom_data.pixel_array\n        image = (image - image.min()) / (image.max() - image.min() + 1e-6) * 255\n        image = Image.fromarray(image.astype(np.uint8)).convert('L')\n        image_array = np.array(image, dtype=np.uint8)\n        return {'study_id': study_id, 'series_id': series_id, 'instance_number': instance_number, 'image': image_array}\n    except Exception as e:\n        print(f\"Error processing file {src_path}: {e}\")\n        return None\n\ndef process_image_wrapper(args):\n    return process_dicom_image(*args)\n\ndef collect_tasks(df, data_type='train'):\n    tasks = []\n    st_ids = df['study_id'].unique()\n    for idx, si in enumerate(tqdm(st_ids, total=len(st_ids))):\n        pdf = df[df['study_id'] == si]\n        for _, row in pdf.iterrows():\n            ds = row['series_description'].replace('/', '_')\n            series_id = row['series_id']\n            img_paths = glob.glob(f'/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/{data_type}_images/{si}/{series_id}/*.dcm')\n            img_paths = sorted(img_paths, key=natural_keys)\n            for j, impath in enumerate(img_paths):\n                instance_number = j + 1  # Assuming the instance number starts from 1\n                tasks.append((impath, si, series_id, instance_number))\n    return tasks\n\n\ndef process_and_save(tasks, output_file, chunk_size=50000):\n    total_chunks = len(tasks) // chunk_size + 1\n    for i in range(total_chunks):\n        chunk_tasks = tasks[i*chunk_size:(i+1)*chunk_size]\n        results = []\n        with ProcessPoolExecutor() as executor:\n            results = list(tqdm(executor.map(process_image_wrapper, chunk_tasks), total=len(chunk_tasks)))\n        # Filter out None results\n        results = [res for res in results if res is not None]\n        # Convert results to DataFrame\n        chunk_df = pd.DataFrame(results)\n        # Reduce memory usage\n        chunk_df = reduce_mem_usage(chunk_df)\n        \n        # Save the chunk results as compressed pickle in memory\n        compressed_file_path = f'{output_file}_part_{i}.pkl.gz'\n        with gzip.open(compressed_file_path, 'wb') as f:\n            pickle.dump(chunk_df, f)\n        \n        print(f'Saved {compressed_file_path}')\n        del chunk_tasks\n        del results\n        del chunk_df\n        gc.collect()\n\n# Clear memory before starting the process\ngc.collect()\n\n# Paths to data directories\ntrain_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/'\ntest_image_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_images/'\n\n# Collect tasks for train and test data\ntrain_tasks = collect_tasks(df_train_desc, data_type='train')\ntest_tasks = collect_tasks(df_test_desc, data_type='test')\n\n# Process and save the train tasks in chunks\nprocess_and_save(train_tasks, 'train_images')\n\n# Process and save the test tasks in chunks\nprocess_and_save(test_tasks, 'test_images')\n'''\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-07-29T06:54:32.457331Z","iopub.status.idle":"2024-07-29T06:54:32.457732Z","shell.execute_reply.started":"2024-07-29T06:54:32.457535Z","shell.execute_reply":"2024-07-29T06:54:32.457551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Define conditions and levels\nconditions = [\n    'spinal_canal_stenosis', 'left_neural_foraminal_narrowing', \n    'right_neural_foraminal_narrowing', 'left_subarticular_stenosis', \n    'right_subarticular_stenosis'\n]\nlevels = ['l1_l2', 'l2_l3', 'l3_l4', 'l4_l5', 'l5_s1']\n\nfor condition in conditions:\n    for level in levels:\n        col_name = f'{condition}_{level}'\n        if col_name not in df_train.columns:\n            df_train[col_name] = np.nan  # or another appropriate default value\n\n# Create long format DataFrame\ndf_long_train = pd.melt(df_train, \n                  id_vars=['study_id', 'series_id','series_description', 'instance_number'], \n                  value_vars=[f'{cond}_{level}' for cond in conditions for level in levels], \n                  var_name='condition_level', \n                  value_name='target')\n\n# Split the 'condition_level' column into 'condition' and 'level'\ndf_long_train['condition'] = df_long_train['condition_level'].apply(lambda x: '_'.join(x.split('_')[:-2]))\ndf_long_train['level'] = df_long_train['condition_level'].apply(lambda x: '_'.join(x.split('_')[-2:]).replace('_', '/'))\ndf_long_train = df_long_train.drop(columns=['condition_level'])\n\n# Define the label map\nlabel_map = {'Normal/Mild': 0, 'Moderate': 1, 'Severe': 2}\ndf_long_train['target'] = df_long_train['target'].map(label_map)\n\n# Expand the test DataFrame to have one row per image instance\ntest_data = []\nfor _, row in df_test_desc.iterrows():\n    study_id = row['study_id']\n    series_id = row['series_id']\n    series_description = row['series_description']\n    series_path = f'/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/test_images/{study_id}/{series_id}/'\n    instance_files = sorted(glob.glob(f'{series_path}/*.dcm'), key=natural_keys)\n    for instance_number, _ in enumerate(instance_files, start=1):\n        for condition in conditions:\n            for level in levels:\n                test_data.append([study_id, series_id, series_description, instance_number, condition, level, 0])\n\ndf_long_test = pd.DataFrame(test_data, columns=['study_id', 'series_id', 'series_description', 'instance_number', 'condition', 'level', 'target'])\n\n","metadata":{"execution":{"iopub.status.busy":"2024-07-29T06:54:32.459038Z","iopub.status.idle":"2024-07-29T06:54:32.459479Z","shell.execute_reply.started":"2024-07-29T06:54:32.459278Z","shell.execute_reply":"2024-07-29T06:54:32.459296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Optimize memory usage of the long format DataFrames\ndf_long_train = reduce_mem_usage(df_long_train)\ndf_long_test = reduce_mem_usage(df_long_test)\n\n# Save final DataFrames\ndf_long_train.to_pickle('train_long.pkl')\ndf_long_test.to_pickle('test_long.pkl')","metadata":{"execution":{"iopub.status.busy":"2024-07-29T06:54:32.460835Z","iopub.status.idle":"2024-07-29T06:54:32.461226Z","shell.execute_reply.started":"2024-07-29T06:54:32.461022Z","shell.execute_reply":"2024-07-29T06:54:32.461038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_long_test","metadata":{"execution":{"iopub.status.busy":"2024-07-29T06:54:32.462268Z","iopub.status.idle":"2024-07-29T06:54:32.462662Z","shell.execute_reply.started":"2024-07-29T06:54:32.462472Z","shell.execute_reply":"2024-07-29T06:54:32.462488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Refrecnes\n\nhttps://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage","metadata":{}}]}