{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **RSNA 2024 Lumbar Spine Degenerative Classification Challenge**","metadata":{}},{"cell_type":"markdown","source":"# **1. Introduction**\nThis notebook is dedicated to the RSNA 2024 Lumbar Spine Degenerative Classification Challenge, which aims to develop a machine learning model that identifies and classifies medical conditions affecting the lumbar spine using MRI scans. Given the high prevalence of lumbar spine issues, which are a leading cause of disability worldwide, this challenge holds significant clinical importance.\n\n**Key points about the challenge:**\n\nThe primary objective is to classify five specific conditions:\n* Left Neural Foraminal Narrowing\n* Right Neural Foraminal Narrowing\n* Left Subarticular Stenosis\n* Right Subarticular Stenosis\n* Spinal Canal Stenosis\n\nThese conditions are evaluated at five intervertebral disc levels:\n* L1/L2, \n* L2/L3, \n* L3/L4, \n* L4/L5, \n* L5/S1.\n\nFor each condition at each level, the model must predict the severity, categorized as:\n* Normal/Mild, \n* Moderate,  \n* Severe.\n\nIn this notebook, we will conduct exploratory data analysis (EDA) and visualizations to gain insights into the dataset. The EDA will include examining the distribution of conditions, severity levels, and any missing values in the dataset. Visualizations will be employed to illustrate these distributions, providing a clearer understanding of the data's characteristics.\n\nBy analyzing the data visually, we can identify patterns, trends, and potential challenges that may arise during model training and evaluation. This foundational analysis will inform the subsequent steps in developing a robust machine learning model for the classification of lumbar spine degenerative conditions.\n","metadata":{}},{"cell_type":"markdown","source":"# **2. Import Libraries**","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pydicom\nimport cv2\nfrom glob import glob\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models, applications\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import log_loss\nfrom skimage import exposure\nfrom tqdm import tqdm\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\nimport torchvision.transforms as transforms\nimport torchvision.models as models\nimport warnings\nwarnings.filterwarnings('ignore')\n\n\n# Set random seeds for reproducibility\nnp.random.seed(42)\ntorch.manual_seed(42)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T19:49:51.964066Z","iopub.execute_input":"2024-09-14T19:49:51.965150Z","iopub.status.idle":"2024-09-14T19:50:12.138853Z","shell.execute_reply.started":"2024-09-14T19:49:51.965081Z","shell.execute_reply":"2024-09-14T19:50:12.137621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **3. Load and Explore Data**","metadata":{}},{"cell_type":"code","source":"# Load CSV files\ntrain_df = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv')\ncoordinates_df = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')\nseries_descriptions_df = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv')\n\nprint(\"Train data shape:\", train_df.shape)\nprint(\"Train label coordinates shape:\", coordinates_df.shape)\nprint(\"Train series descriptions shape:\", series_descriptions_df.shape)\n\n# Display first few rows of each dataframe\nprint(\"\\nTrain data head:\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T21:54:06.458835Z","iopub.execute_input":"2024-09-14T21:54:06.459303Z","iopub.status.idle":"2024-09-14T21:54:06.590409Z","shell.execute_reply.started":"2024-09-14T21:54:06.459265Z","shell.execute_reply":"2024-09-14T21:54:06.589097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"\\nTrain label coordinates head:\")\ncoordinates_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:43.913936Z","iopub.execute_input":"2024-09-14T18:08:43.914307Z","iopub.status.idle":"2024-09-14T18:08:43.932645Z","shell.execute_reply.started":"2024-09-14T18:08:43.914267Z","shell.execute_reply":"2024-09-14T18:08:43.931156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"\\nTrain series descriptions head:\")\nseries_descriptions_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:43.936402Z","iopub.execute_input":"2024-09-14T18:08:43.937002Z","iopub.status.idle":"2024-09-14T18:08:43.953319Z","shell.execute_reply.started":"2024-09-14T18:08:43.936945Z","shell.execute_reply":"2024-09-14T18:08:43.951683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display basic information\nprint(\"\\nTrain data info:\")\nprint(train_df.info())\nprint(\"\\nTrain label coordinates info:\")\nprint(coordinates_df.info())\nprint(\"\\nTrain series descriptions info:\")\nprint(series_descriptions_df.info())","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:43.955028Z","iopub.execute_input":"2024-09-14T18:08:43.955581Z","iopub.status.idle":"2024-09-14T18:08:44.366573Z","shell.execute_reply.started":"2024-09-14T18:08:43.955521Z","shell.execute_reply":"2024-09-14T18:08:44.365407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **4. Exploratory Data Analysis (EDA) and Visualization**","metadata":{}},{"cell_type":"code","source":"# Set up plotting style\nplt.style.use('seaborn')\nsns.set_palette(\"deep\")","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:44.367963Z","iopub.execute_input":"2024-09-14T18:08:44.368462Z","iopub.status.idle":"2024-09-14T18:08:44.378588Z","shell.execute_reply.started":"2024-09-14T18:08:44.368421Z","shell.execute_reply":"2024-09-14T18:08:44.377165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the number of studies and series\nnum_studies = train_df['study_id'].nunique()\nnum_series = series_descriptions_df['series_id'].nunique()\nprint(f\"Number of unique studies: {num_studies}\")\nprint(f\"Number of unique series: {num_series}\")\n\n# Count the number of conditions and levels\nconditions = [col for col in train_df.columns if col != 'study_id']\n\nnum_conditions = len(conditions)\nprint(f\"Number of condition-level combinations: {num_conditions}\")","metadata":{"execution":{"iopub.status.busy":"2024-09-14T21:54:19.993632Z","iopub.execute_input":"2024-09-14T21:54:19.994099Z","iopub.status.idle":"2024-09-14T21:54:20.005180Z","shell.execute_reply.started":"2024-09-14T21:54:19.994059Z","shell.execute_reply":"2024-09-14T21:54:20.003647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.1. MRI Image Slice Distribution & Visualization\nvisualize the MRI image slices using the .dcm files.","metadata":{}},{"cell_type":"code","source":"def load_dicom(path):\n    dicom = pydicom.dcmread(path)\n    img = dicom.pixel_array\n    return img\n\ndef preprocess_image(img):\n    # Normalize pixel values\n    img = (img - img.min()) / (img.max() - img.min())\n    # Resize to a consistent size (e.g., 224x224 for many pre-trained models)\n    img = torch.from_numpy(img).unsqueeze(0)  # Add channel dimension\n    img = transforms.Resize((224, 224))(img)\n    return img.squeeze(0).numpy()\n\ndef load_and_preprocess_study(study_id, series_id):\n    study_path = f'/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/{study_id}/{series_id}'\n    images = []\n    for file in sorted(os.listdir(study_path)):\n        if file.endswith('.dcm'):\n            img_path = os.path.join(study_path, file)\n            img = load_dicom(img_path)\n            img = preprocess_image(img)\n            images.append(img)\n    return np.array(images)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:44.393292Z","iopub.execute_input":"2024-09-14T18:08:44.393630Z","iopub.status.idle":"2024-09-14T18:08:44.404718Z","shell.execute_reply.started":"2024-09-14T18:08:44.393593Z","shell.execute_reply":"2024-09-14T18:08:44.403155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Sample MRI Images Visualization**\nThe visualize_sample_images function displays a selection of MRI images from the RSNA 2024 Lumbar Spine Degenerative Classification Dataset, providing an overview of the data quality and the conditions present. Each image is accompanied by a title that includes relevant information such as the study ID, series ID, instance number, condition, and vertebral level, facilitating a quick assessment of the dataset's content. This visualization allows researchers to qualitatively evaluate the clarity of anatomical structures and the visibility of degenerative conditions, which is essential for understanding the dataset and guiding subsequent model development.","metadata":{}},{"cell_type":"code","source":"def visualize_sample_images(train_df, series_descriptions_df, coordinates_df):\n    sample_study = train_df['study_id'].iloc[0]\n    sample_series = series_descriptions_df[series_descriptions_df['study_id'] == sample_study]['series_id'].iloc[0]\n    \n    # Load the sample images\n    sample_images = load_and_preprocess_study(sample_study, sample_series)\n    \n    # Get relevant information for the titles\n    condition_info = coordinates_df[coordinates_df['study_id'] == sample_study]\n    \n    fig, axes = plt.subplots(2, 3, figsize=(15, 10))\n    \n    for i, ax in enumerate(axes.flat):\n        if i < len(sample_images):\n            ax.imshow(sample_images[i], cmap='gray')\n            ax.axis('off')\n            \n            # Get instance number and condition for the current slice\n            instance_number = i + 1  # Assuming instance numbers start at 1\n            instance_data = condition_info[condition_info['instance_number'] == instance_number]\n            \n            if not instance_data.empty:\n                condition = instance_data['condition'].values[0]\n                level = instance_data['level'].values[0]\n                ax.set_title(f'Study: {sample_study}\\nSeries: {sample_series}\\nInstance: {instance_number}\\nCondition: {condition}\\nLevel: {level}', fontsize=10, pad=5)\n            else:\n                ax.set_title(f'Study: {sample_study}\\nSeries: {sample_series}\\nInstance: {instance_number}\\nCondition: N/A\\nLevel: N/A', fontsize=10, pad=5)\n\n    plt.suptitle('Sample Images from the RSNA 2024 Lumbar Spine Degenerative Classification Dataset', fontsize=14)\n    plt.tight_layout()\n    plt.subplots_adjust(top=0.85, hspace=0.4, wspace=0.4)  # Adjust spacing\n    plt.show()\n    \n# Visualize sample images\nvisualize_sample_images(train_df, series_descriptions_df, coordinates_df)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:46.227994Z","iopub.execute_input":"2024-09-14T18:08:46.228519Z","iopub.status.idle":"2024-09-14T18:08:48.072262Z","shell.execute_reply.started":"2024-09-14T18:08:46.228472Z","shell.execute_reply":"2024-09-14T18:08:48.070604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Visualization of Regions of Interest (ROIs)**\nThe **plot_rois_for_series** function visualizes the **Regions of Interest (ROIs)** for all instances in a specified MRI series from the RSNA 2024 Lumbar Spine Degenerative Classification Dataset. It filters the label coordinates based on the provided study and series IDs, then creates subplots for each unique instance number. For each instance, the corresponding DICOM image is loaded and displayed, with ROIs marked in red based on the coordinates from the dataset. The titles of the subplots include the instance number and the associated conditions and vertebral levels, providing context for the visualized images. This function is essential for assessing the locations of significant anatomical features and degenerative conditions within the MRI images, aiding in the evaluation and understanding of the dataset.","metadata":{}},{"cell_type":"code","source":"# Visualize ROIs based on label coordinates for all instances of a series\ndef plot_rois_for_series(study_id, series_id, label_coords, train_df):\n    # Filter label coordinates for the specific study and series\n    instances = label_coords[(label_coords['study_id'] == int(study_id)) & (label_coords['series_id'] == int(series_id))]\n    \n    # Get unique instance numbers for the series\n    unique_instances = instances['instance_number'].unique()\n    \n    # Check if there are any instances to plot\n    if len(unique_instances) == 0:\n        print(\"No instances found for the specified study and series.\")\n        return\n    \n    # Create subplots\n    num_instances = len(unique_instances)\n    cols = 2  # Number of columns for subplots\n    rows = (num_instances + cols - 1) // cols  # Calculate number of rows needed\n    plt.figure(figsize=(15, 5 * rows))  # Adjust figure size based on number of rows\n    \n    for i, instance_number in enumerate(unique_instances):\n        # Load the DICOM image for the current instance\n        dicom_path = f'/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/{study_id}/{series_id}/{instance_number}.dcm'\n        \n        # Check if the DICOM file exists\n        try:\n            dicom = pydicom.dcmread(dicom_path)\n            img = dicom.pixel_array\n        except FileNotFoundError:\n            print(f\"File not found: {dicom_path}\")\n            continue\n        \n        # Create subplot\n        ax = plt.subplot(rows, cols, i + 1)\n        ax.imshow(img, cmap='gray')\n        ax.axis('off')\n        \n        # Plot ROIs for the current instance\n        instance_coords = instances[instances['instance_number'] == instance_number]\n        for _, row in instance_coords.iterrows():\n            plt.plot(row['x'], row['y'], 'ro', markersize=5)  # Plot ROIs\n        \n        # Get relevant conditions for the current instance\n        relevant_conditions = []\n        for _, row in instance_coords.iterrows():\n            condition = row['condition']\n            level = row['level']\n            relevant_conditions.append(f\"{condition}: {level}\")\n        \n        # Create title text for the subplot\n        conditions_text = ', '.join(relevant_conditions)\n        ax.set_title(f'Instance: {instance_number}\\nConditions: {conditions_text}', fontsize=10, wrap=True)\n    \n    plt.tight_layout()\n    plt.show()\n\n# Axial Example \nstudy_id = '100206310'\nseries_id = '1012284084'\nlabel_coords = coordinates_df[(coordinates_df['study_id'] == int(study_id)) & (coordinates_df['series_id'] == int(series_id))]\n\n# Visualize the ROIs for all instances in the series\nplot_rois_for_series(study_id, series_id, label_coords, train_df)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:54.036486Z","iopub.execute_input":"2024-09-14T18:08:54.036989Z","iopub.status.idle":"2024-09-14T18:08:55.990044Z","shell.execute_reply.started":"2024-09-14T18:08:54.036939Z","shell.execute_reply":"2024-09-14T18:08:55.988241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this example, the study_id is set to '100206310' and the series_id is '1012284084', which contains 60 DICOM images. However, the code only displays 5 images because it filters the label_coords DataFrame to find unique instance numbers associated with that series. If the label_coords DataFrame does not have entries for all 60 instances—meaning it lacks the corresponding x and y coordinates for some of those images—only the images for which coordinates are available will be plotted. Therefore, the function creates subplots for each available instance, loading the corresponding DICOM image, plotting the ROIs in red, and adding a title that summarizes the relevant conditions for each instance, ensuring a clear and organized visualization of the data. ","metadata":{}},{"cell_type":"code","source":"# Sagittal Example\nstudy_id = '4003253'\nseries_id = '702807833'\n\nlabel_coords = coordinates_df[(coordinates_df['study_id'] == int(study_id)) & (coordinates_df['series_id'] == int(series_id))]\n\n# Visualize the ROIs for all instances in the series\nplot_rois_for_series(study_id, series_id, label_coords, train_df)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:08:58.792834Z","iopub.execute_input":"2024-09-14T18:08:58.793348Z","iopub.status.idle":"2024-09-14T18:08:59.559376Z","shell.execute_reply.started":"2024-09-14T18:08:58.793301Z","shell.execute_reply":"2024-09-14T18:08:59.557775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.2. Visualize Class Distribution & Missing Data\n\nThis will help us see how the labels are distributed and if there are any missing values in the data.","metadata":{}},{"cell_type":"code","source":"# Calculate the count of missing values for each condition-level combination\nmissing_counts = train_df[conditions].isnull().sum()\nprint(\"\\nCount of missing values for each condition-level combination:\")\nprint(missing_counts)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T21:54:27.279972Z","iopub.execute_input":"2024-09-14T21:54:27.280494Z","iopub.status.idle":"2024-09-14T21:54:27.293508Z","shell.execute_reply.started":"2024-09-14T21:54:27.280450Z","shell.execute_reply":"2024-09-14T21:54:27.292056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check missing data in train.csv\ndef visualize_missing_data(train_df, coordinates_df):\n    # Missing data summary for train_df\n    missing_data_summary = train_df.isnull().sum()\n    plt.figure(figsize=(10, 6))\n    sns.heatmap(train_df.isnull(), cbar=False, cmap=\"viridis\")\n    plt.title('Missing Data in train.csv')\n    plt.show()\n\n#     # Missing data summary for coordinates_df\n#     missing_coords_summary = coordinates_df.isnull().sum()\n#     plt.figure(figsize=(10, 6))\n#     sns.heatmap(coordinates_df.isnull(), cbar=False, cmap=\"magma\")\n#     plt.title('Missing Data in train_label_coordinates.csv')\n#     plt.show()\n    \n# Visualize missing data\nvisualize_missing_data(train_df, coordinates_df)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:09:27.533398Z","iopub.execute_input":"2024-09-14T18:09:27.533855Z","iopub.status.idle":"2024-09-14T18:09:28.390480Z","shell.execute_reply.started":"2024-09-14T18:09:27.533812Z","shell.execute_reply":"2024-09-14T18:09:28.389142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Distribution of Severity Levels for Degenerative Conditions**\nThis below code snippet generates a grid of bar plots to visualize the distribution of severity levels for each degenerative condition in the RSNA 2024 Lumbar Spine Degenerative Classification Dataset. The conditions are plotted in a 6x5 grid, with each subplot displaying the count of normal/mild, moderate, and severe cases for a specific condition. This visualization provides a concise overview of the prevalence and severity of each degenerative condition in the dataset, which is crucial for understanding the class imbalance and guiding the development of robust classification models. By analyzing these distributions, researchers can identify potential challenges, such as rare or underrepresented severity levels, and devise appropriate strategies to address them during model training and evaluation.","metadata":{}},{"cell_type":"code","source":"# Plot the distribution of severity levels for each condition\nplt.figure(figsize=(25, 20))\nfor i, condition in enumerate(conditions):\n    plt.subplot(6, 5, i+1)\n    train_df[condition].value_counts().plot(kind='bar')\n    plt.title(condition, fontsize=8)\n    plt.ylabel('Count', fontsize=8)\n    plt.xticks(rotation=45, fontsize=6)\n    plt.yticks(fontsize=6)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:09:32.252667Z","iopub.execute_input":"2024-09-14T18:09:32.253094Z","iopub.status.idle":"2024-09-14T18:09:37.580143Z","shell.execute_reply.started":"2024-09-14T18:09:32.253040Z","shell.execute_reply":"2024-09-14T18:09:37.578756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Analysis of Severity Distribution Across Conditions**\nThis code snippet analyzes and visualizes the distribution of severity levels across all degenerative conditions in the RSNA 2024 Lumbar Spine Degenerative Classification Dataset. It first melts the condition columns into a single column to aggregate severity counts, then creates a bar plot to display the total counts of each severity level. This visualization helps identify trends and imbalances in severity distribution, which is important for understanding the dataset and informing model training strategies.","metadata":{}},{"cell_type":"code","source":"# Analyze the distribution of conditions and severity\ncondition_columns = [col for col in train_df.columns if col != 'study_id']\nseverity_counts = train_df[condition_columns].melt().value_counts()\nplt.figure(figsize=(12, 6))\nseverity_counts.plot(kind='bar')\nplt.title('Distribution of Severity Across All Conditions')\nplt.xlabel('Severity')\nplt.ylabel('Count')\nplt.xticks(rotation=90, ha='right', fontsize=6)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:09:38.431263Z","iopub.execute_input":"2024-09-14T18:09:38.431666Z","iopub.status.idle":"2024-09-14T18:09:39.654279Z","shell.execute_reply.started":"2024-09-14T18:09:38.431627Z","shell.execute_reply":"2024-09-14T18:09:39.652910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Visualization of Severity Level Distribution**\nThis code generates a bar plot to visualize the distribution of severity levels across the entire RSNA 2024 Lumbar Spine Degenerative Classification Dataset. This plot provides a clear overview of the overall distribution of severity levels in the dataset, which is valuable for understanding class imbalances and guiding the development of effective classification models.","metadata":{}},{"cell_type":"code","source":"# Plot class distribution for severity labels\nseverity_cols = train_df.columns[1:]  # Exclude 'study_id'\nseverity_data = train_df.melt(id_vars='study_id', var_name='condition_level', value_name='severity')\n\nplt.figure(figsize=(10, 5))\nsns.countplot(data=severity_data, x='severity', order=['Normal/Mild', 'Moderate', 'Severe'])\nplt.title(\"Severity Level Distribution\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:09:39.656552Z","iopub.execute_input":"2024-09-14T18:09:39.657049Z","iopub.status.idle":"2024-09-14T18:09:39.924877Z","shell.execute_reply.started":"2024-09-14T18:09:39.657001Z","shell.execute_reply":"2024-09-14T18:09:39.923413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Countplot of Severity Levels for Each Condition**\nThis code snippet creates countplots to visualize the severity level distributions for five specific degenerative conditions in the RSNA 2024 Lumbar Spine Degenerative Classification Dataset. For each condition, a countplot is generated, mapping the severity levels ('Normal/Mild', 'Moderate', 'Severe') to numerical values for plotting. This visualization provides a clear comparison of how severity levels are distributed for each condition, aiding in the assessment of class imbalances and informing model training strategies.","metadata":{}},{"cell_type":"code","source":"# Countplot for each condition's severity levels\nconditions = ['spinal_canal_stenosis', 'left_neural_foraminal_narrowing', \n              'right_neural_foraminal_narrowing', 'left_subarticular_stenosis', \n              'right_subarticular_stenosis']\n\n# Create a figure with subplots\nfig, axes = plt.subplots(3, 2, figsize=(16, 12))\n\n# Initialize a counter for the axes\nax_idx = 0\n\nfor condition in conditions:\n    # Get the data for the current condition\n    data = train_df.filter(like=condition).stack()\n    \n    # Check if there is data to plot\n    if not data.empty:\n        # Create the countplot\n        sns.countplot(ax=axes[ax_idx // 2, ax_idx % 2], \n                      x=data.map({'Normal/Mild': 0, 'Moderate': 1, 'Severe': 2}))\n        axes[ax_idx // 2, ax_idx % 2].set_title(f\"{condition} Severity Distribution\")\n        ax_idx += 1\n\n# Remove empty subplots\nfor i in range(ax_idx, 6):  # 6 is the total number of subplots (3 rows * 2 columns)\n    fig.delaxes(axes[i // 2, i % 2])\n\n# Adjust the layout to accommodate the used subplots\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:09:48.789814Z","iopub.execute_input":"2024-09-14T18:09:48.790846Z","iopub.status.idle":"2024-09-14T18:09:50.023672Z","shell.execute_reply.started":"2024-09-14T18:09:48.790785Z","shell.execute_reply":"2024-09-14T18:09:50.022367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Distribution of Series Descriptions**\nAnalyzes and visualizes the distribution of MRI series descriptions, specifically focusing on the orientations of images such as 'Sagittal T1', 'Axial T2', and 'Sagittal T2/STIR'. By plotting the counts of each series description in a bar chart, it provides insights into the variety of imaging modalities used in the dataset. This analysis is crucial for understanding how different MRI sequences contribute to the assessment of degenerative spine conditions, as each orientation and sequence type can reveal unique anatomical details and pathologies. The visualization allows researchers to identify the prevalence of specific imaging techniques, which can inform the selection of appropriate sequences for diagnosis and treatment planning.","metadata":{}},{"cell_type":"code","source":"# Analyze the distribution of series descriptions\nplt.figure(figsize=(12, 6))\nseries_descriptions_df['series_description'].value_counts().plot(kind='bar')\nplt.title('Distribution of Series Descriptions')\nplt.xlabel('Series Description')\nplt.ylabel('Count')\nplt.xticks(rotation=45, ha='right')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:25:56.454876Z","iopub.execute_input":"2024-09-14T18:25:56.455419Z","iopub.status.idle":"2024-09-14T18:25:56.756903Z","shell.execute_reply.started":"2024-09-14T18:25:56.455374Z","shell.execute_reply":"2024-09-14T18:25:56.755339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.3. Plot Coordinates of Labeled Conditions\nPlot the coordinates of the areas labeled in the train_label_coordinates.csv file.","metadata":{}},{"cell_type":"code","source":"# Plot coordinates for specific conditions\nplt.figure(figsize=(8, 6))\nplt.scatter(coordinates_df['x'], coordinates_df['y'], alpha=0.6)\nplt.title('Coordinate scatter plot for labeled conditions')\nplt.xlabel('X Coordinate')\nplt.ylabel('Y Coordinate')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:09:50.025940Z","iopub.execute_input":"2024-09-14T18:09:50.026353Z","iopub.status.idle":"2024-09-14T18:09:50.517265Z","shell.execute_reply.started":"2024-09-14T18:09:50.026311Z","shell.execute_reply":"2024-09-14T18:09:50.516036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Condition-wise Coordinate Scatter Plot**\nScatter plot to visualize the spatial distribution of different degenerative spine conditions based on their x and y coordinates from the coordinates_df DataFrame. Each unique condition is represented by a distinct color in the scatter plot, allowing for easy differentiation between conditions. The plot provides insights into how various conditions are distributed in the coordinate space, which can be useful for understanding patterns and relationships between different degenerative conditions. By analyzing this scatter plot, researchers can identify clusters or trends that may inform further investigation into the anatomical locations affected by each condition, aiding in the assessment and classification of degenerative spine issues.","metadata":{}},{"cell_type":"code","source":"# Add condition to the scatter plot\nplt.figure(figsize=(8, 6))\nfor condition in coordinates_df['condition'].unique():\n    condition_data = coordinates_df[coordinates_df['condition'] == condition]\n    plt.scatter(condition_data['x'], condition_data['y'], label=condition, alpha=0.6)\n\nplt.title('Condition-wise Coordinate Scatter Plot')\nplt.xlabel('X Coordinate')\nplt.ylabel('Y Coordinate')\nplt.legend()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-09-14T18:09:50.518927Z","iopub.execute_input":"2024-09-14T18:09:50.519387Z","iopub.status.idle":"2024-09-14T18:09:51.817546Z","shell.execute_reply.started":"2024-09-14T18:09:50.519337Z","shell.execute_reply":"2024-09-14T18:09:51.816024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Visualization of MRI Coordinates by Imaging Plane**\nThe **plot_coordinates** function creates a 2x2 subplot to visualize the spatial distribution of MRI coordinates, differentiating between various imaging planes. \nThe function then generates four subplots:\n* **Condition-wise Coordinate Scatter Plot by Imaging Plane:** This subplot displays a scatter plot of coordinates for each condition, with different colors representing specific imaging planes (Sagittal T1, Axial T2, and Sagittal T2/STIR). This allows for a comprehensive comparison of how different conditions are distributed across various imaging modalities.\n* **Sagittal T1:** A dedicated scatter plot for Sagittal T1 images, highlighting the spatial distribution of coordinates in this specific imaging plane.\n* **Axial T2:** A scatter plot focused on Axial T2 images, providing insights into the coordinate distribution in this orientation.\n* **Sagittal T2/STIR:** A scatter plot for Sagittal T2/STIR images, enabling the analysis of coordinate patterns in this imaging sequence.\n\nBy visualizing the MRI coordinates in this manner, researchers can gain valuable insights into how different degenerative spine conditions are represented across various imaging planes. This information can help in understanding the spatial relationships between conditions and guide the selection of appropriate imaging modalities for diagnosis and treatment planning.\n","metadata":{}},{"cell_type":"code","source":"# Function to visualize MRI coordinates with differentiation by imaging plane in subplots\ndef plot_coordinates(coordinates_df, series_descriptions_df):\n    # Merge coordinates_df with series_descriptions_df\n    merged_df = pd.merge(coordinates_df, series_descriptions_df, on=['study_id', 'series_id'])\n\n    # Create a 2x2 subplot with increased spacing\n    fig, axs = plt.subplots(2, 2, figsize=(12, 12), gridspec_kw={'wspace': 0.4, 'hspace': 0.5})\n    \n    # Define colors for specific imaging planes based on dataset\n    plane_colors = {\n        'Sagittal T1': 'blue',\n        'Axial T2': 'green',\n        'Sagittal T2/STIR': 'red'\n    }\n\n    # 1. Condition-wise Coordinate Scatter Plot by Imaging Plane\n    for condition in merged_df['condition'].unique():\n        condition_data = merged_df[merged_df['condition'] == condition]\n        for plane in condition_data['series_description'].unique():\n            plane_data = condition_data[condition_data['series_description'] == plane]\n            color = plane_colors.get(plane, 'gray')  # Default to gray if plane not recognized\n            axs[0, 0].scatter(plane_data['x'], plane_data['y'], label=f'{condition} - {plane}', \n                              color=color, alpha=0.6)\n\n    axs[0, 0].set_title('Condition-wise Coordinate Scatter Plot by Imaging Plane')\n    axs[0, 0].set_xlabel('X Coordinate')\n    axs[0, 0].set_ylabel('Y Coordinate')\n    axs[0, 0].grid(True)\n    axs[0, 0].legend(loc='upper right', fontsize='small')\n\n    # 2. Scatter plot for Sagittal T1\n    sagittal_t1_data = merged_df[merged_df['series_description'] == 'Sagittal T1']\n    axs[0, 1].scatter(sagittal_t1_data['x'], sagittal_t1_data['y'], color='blue', alpha=0.6)\n    axs[0, 1].set_title('Sagittal T1')\n    axs[0, 1].set_xlabel('X Coordinate')\n    axs[0, 1].set_ylabel('Y Coordinate')\n    axs[0, 1].grid(True)\n\n    # 3. Scatter plot for Axial T2\n    axial_t2_data = merged_df[merged_df['series_description'] == 'Axial T2']\n    axs[1, 0].scatter(axial_t2_data['x'], axial_t2_data['y'], color='green', alpha=0.6)\n    axs[1, 0].set_title('Axial T2')\n    axs[1, 0].set_xlabel('X Coordinate')\n    axs[1, 0].set_ylabel('Y Coordinate')\n    axs[1, 0].grid(True)\n\n    # 4. Scatter plot for Sagittal T2/STIR\n    sagittal_t2_stir_data = merged_df[merged_df['series_description'] == 'Sagittal T2/STIR']\n    axs[1, 1].scatter(sagittal_t2_stir_data['x'], sagittal_t2_stir_data['y'], color='red', alpha=0.6)\n    axs[1, 1].set_title('Sagittal T2/STIR')\n    axs[1, 1].set_xlabel('X Coordinate')\n    axs[1, 1].set_ylabel('Y Coordinate')\n    axs[1, 1].grid(True)\n\n    # Adjust layout\n    plt.subplots_adjust(wspace=0.4, hspace=0.5)\n    plt.show()\n\nplot_coordinates(coordinates_df, series_descriptions_df)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T20:27:46.377628Z","iopub.execute_input":"2024-09-14T20:27:46.378876Z","iopub.status.idle":"2024-09-14T20:27:47.791920Z","shell.execute_reply.started":"2024-09-14T20:27:46.378822Z","shell.execute_reply":"2024-09-14T20:27:47.790718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Pixel Intensity Distribution Plot** \nThis plot visualizes the Pixel Intensity Distribution of an MRI image, which shows how frequently different pixel intensity values (grayscale levels) appear in the image. The x-axis represents pixel intensity, ranging from dark (low values) to bright (high values), while the y-axis shows the frequency of each intensity. In MRI scans, most useful information lies within a certain intensity range, and the goal of this plot is to examine the distribution of these pixel values. By analyzing this distribution, we can evaluate whether the image has sufficient contrast or whether preprocessing steps are necessary to enhance its quality.\n\n**Insights from the Output**\n\nThe output shows a high concentration of low-intensity pixel values, which indicates that much of the image consists of darker areas. This could make it difficult for a machine learning model to distinguish important features, such as bone structures and tissues, in its current form. Based on this observation, it is clear that contrast enhancement (such as histogram equalization) might be needed to spread out the pixel values and highlight important anatomical features. Additionally, intensity normalization may be required to ensure consistent pixel values across the dataset, and noise reduction could help to minimize random variations in pixel intensities, improving model performance.","metadata":{}},{"cell_type":"code","source":"def plot_pixel_distribution(study_id, series_id, img_num):\n    # Construct the path to the DICOM image\n    img_path = f\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/{study_id}/{series_id}/{img_num}.dcm\"  # Adjust the path as needed\n    \n    # Check if the file exists\n    if not os.path.exists(img_path):\n        print(f\"File not found: {img_path}\")\n        return\n\n    # Read the DICOM image\n    dicom_image = pydicom.dcmread(img_path).pixel_array\n    \n    # Plot histogram of pixel intensities\n    plt.figure(figsize=(10, 6))\n    plt.hist(dicom_image.ravel(), bins=50, color='c', alpha=0.7)\n    plt.title('Pixel Intensity Distribution')\n    plt.xlabel('Pixel Intensity')\n    plt.ylabel('Frequency')\n    plt.grid(True)\n    plt.show()\n\n\nplot_pixel_distribution(study_id=\"100206310\", series_id=\"1012284084\", img_num=1)","metadata":{"execution":{"iopub.status.busy":"2024-09-14T20:55:04.940430Z","iopub.execute_input":"2024-09-14T20:55:04.941348Z","iopub.status.idle":"2024-09-14T20:55:05.324473Z","shell.execute_reply.started":"2024-09-14T20:55:04.941299Z","shell.execute_reply":"2024-09-14T20:55:05.323157Z"},"trusted":true},"execution_count":null,"outputs":[]}]}