{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"},{"sourceId":9426995,"sourceType":"datasetVersion","datasetId":5726454}],"dockerImageVersionId":30762,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Understanding the Relationship Between `train.csv`, `train_label_coordinates.csv`, and `train_series_descriptions.csv`\n\nIn this competition, we are tasked with predicting the severity of certain spinal conditions at different vertebral levels using MRI scans. The dataset is split into multiple files that provide different layers of information, which, when combined, allow us to create a training pipeline. Here, we explore how the three primary files are related:\n\n---\n\n### 1. **`train.csv`**\nThis file provides the **severity labels** for the different spinal conditions across multiple vertebral levels. For each `study_id`, we have information about conditions like **Spinal Canal Stenosis** or **Left Neural Foraminal Narrowing** at the levels **L1/L2**, **L2/L3**, etc. The severity levels are categorized into:\n- **Normal/Mild**\n- **Moderate**\n- **Severe**\n\n#### Example Structure:\n| study_id  | spinal_canal_stenosis_l1_l2 | spinal_canal_stenosis_l2_l3 | ... |\n| --------- | ----------------------------| ----------------------------| --- |\n| 4003253   | Normal/Mild                  | Normal/Mild                  | ... |\n| 4646740   | Normal/Mild                  | Moderate                     | ... |\n\nThis file provides the ground truth labels, which we will use to train our model. The vertebral levels are specified for each condition, allowing us to track and classify the condition severity for each level.\n\n---\n\n### 2. **`train_label_coordinates.csv`**\nThis file gives us the **precise coordinates** in the MRI images where the conditions are located. For each `study_id` and condition, we have `x` and `y` coordinates that tell us where in the image the region of interest is located. This is crucial for focusing our model's attention on the relevant parts of the scan.\n\n#### Key Columns:\n- **study_id**: Corresponds to the patient or study.\n- **series_id**: Refers to the imaging series where the condition is located.\n- **instance_number**: Refers to the specific slice in the MRI series where the condition is observed.\n- **x, y**: Pixel coordinates of the region of interest (ROI).\n\n#### Example Structure:\n| study_id | series_id   | instance_number | condition             | level   | x         | y        |\n| -------- | ----------- | --------------- | --------------------- | ------- | --------- | -------- |\n| 4003253  | 702807833   | 8               | Spinal Canal Stenosis | L1/L2   | 322.83    | 227.96   |\n| 4003253  | 702807833   | 8               | Spinal Canal Stenosis | L2/L3   | 320.57    | 295.71   |\n\nThis file tells us where the conditions are located within each scan, and we use this information to crop or focus on specific areas of the image during training.\n\n---\n\n### 3. **`train_series_descriptions.csv`**\nThis file helps us identify the **type of MRI series** (scan orientation or modality) used for each `study_id`. It provides metadata for the different MRI series, such as whether the scan is **Sagittal T1**, **Sagittal T2/STIR**, or **Axial T2**. Each of these gives us different views of the spine, and combining them allows for a richer representation of the data.\n\n#### Key Columns:\n- **study_id**: Identifies the patient or study.\n- **series_id**: Uniquely identifies a particular series of MRI images within a study.\n- **series_description**: Describes the orientation or modality of the scan (e.g., Sagittal T1, Axial T2).\n\n#### Example Structure:\n| study_id | series_id   | series_description |\n| -------- | ----------- | ------------------ |\n| 4003253  | 702807833   | Sagittal T2/STIR    |\n| 4003253  | 1054713880  | Sagittal T1         |\n| 4003253  | 2448190387  | Axial T2            |\n\nThis file helps us load the correct MRI images corresponding to each condition and level. For example, we know from `train_label_coordinates.csv` where the condition is located, and we use the `series_description` from this file to load the relevant series for that study.\n\n---\n\n### **Bringing It All Together**\n\nThese three files are interconnected and together form the foundation for building our training dataset:\n\n1. **`train.csv`** provides the labels (severity of conditions) for different vertebral levels and conditions.\n2. **`train_label_coordinates.csv`** provides the precise locations (coordinates) of these conditions within the MRI images.\n3. **`train_series_descriptions.csv`** helps us locate the correct series and modality (Sagittal or Axial views) for each study and condition.\n\n#### Workflow:\n\n1. For each `study_id` in `train.csv`, we use the vertebral level and condition to get the severity labels.\n2. We use the `train_label_coordinates.csv` to retrieve the coordinates of the condition within the MRI images, focusing on specific areas of the scan.\n3. Finally, using `train_series_descriptions.csv`, we load the correct MRI series (Sagittal T1, Axial T2, etc.) for the study and process it using the coordinates and condition information.\n\nBy combining the severity labels, coordinates, and scan types, we can train a deep learning model that learns to predict the severity of conditions from MRI images in a structured and systematic way.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nimport sys\nfrom PIL import Image\nimport cv2\nimport math, random\nimport numpy as np\nimport pandas as pd\nfrom glob import glob\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import KFold\n\nfrom collections import OrderedDict\n\nimport torch\nimport torch.nn.functional as F\nfrom torch import nn\nfrom torch.utils.data import DataLoader, Dataset\nfrom torch.optim import AdamW\n\nimport timm\nfrom timm.utils import ModelEmaV2\nfrom transformers import get_cosine_schedule_with_warmup\n\nfrom torchvision import transforms\n\nimport albumentations as A\n\nfrom sklearn.model_selection import KFold\n\nimport re\nimport pydicom\nfrom typing import Optional\nimport glob\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:38:41.194801Z","iopub.execute_input":"2024-09-23T10:38:41.195723Z","iopub.status.idle":"2024-09-23T10:39:09.734231Z","shell.execute_reply.started":"2024-09-23T10:38:41.195648Z","shell.execute_reply":"2024-09-23T10:39:09.732803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:09.736415Z","iopub.execute_input":"2024-09-23T10:39:09.737246Z","iopub.status.idle":"2024-09-23T10:39:09.778948Z","shell.execute_reply.started":"2024-09-23T10:39:09.737197Z","shell.execute_reply":"2024-09-23T10:39:09.777552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Total Cases: \", len(train))","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:09.780684Z","iopub.execute_input":"2024-09-23T10:39:09.781279Z","iopub.status.idle":"2024-09-23T10:39:09.788410Z","shell.execute_reply.started":"2024-09-23T10:39:09.781216Z","shell.execute_reply":"2024-09-23T10:39:09.787056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.columns","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:09.791548Z","iopub.execute_input":"2024-09-23T10:39:09.792056Z","iopub.status.idle":"2024-09-23T10:39:09.811416Z","shell.execute_reply.started":"2024-09-23T10:39:09.792004Z","shell.execute_reply":"2024-09-23T10:39:09.809718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:09.813351Z","iopub.execute_input":"2024-09-23T10:39:09.813814Z","iopub.status.idle":"2024-09-23T10:39:09.857226Z","shell.execute_reply.started":"2024-09-23T10:39:09.813770Z","shell.execute_reply":"2024-09-23T10:39:09.855753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nimport warnings\nfigure, axis = plt.subplots(1,3, figsize=(20,5)) \nfor idx, d in enumerate(['foraminal', 'subarticular', 'canal']):\n    diagnosis = list(filter(lambda x: x.find(d) > -1, train.columns))\n    dff = train[diagnosis]\n    with warnings.catch_warnings():\n        warnings.simplefilter(action='ignore', category=FutureWarning)\n        value_counts = dff.apply(pd.value_counts).fillna(0).T\n    value_counts.plot(kind='bar', stacked=True, ax=axis[idx])\n    axis[idx].set_title(f'{d} distribution')\n\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:09.859154Z","iopub.execute_input":"2024-09-23T10:39:09.859711Z","iopub.status.idle":"2024-09-23T10:39:11.368039Z","shell.execute_reply.started":"2024-09-23T10:39:09.859648Z","shell.execute_reply":"2024-09-23T10:39:11.366463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# List out all of the Studies we have on patients.\npart_1 = os.listdir('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images')\npart_1 = list(filter(lambda x: x.find('.DS') == -1, part_1))\n\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:11.369591Z","iopub.execute_input":"2024-09-23T10:39:11.370002Z","iopub.status.idle":"2024-09-23T10:39:11.518550Z","shell.execute_reply.started":"2024-09-23T10:39:11.369931Z","shell.execute_reply":"2024-09-23T10:39:11.517429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta_f = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:11.520143Z","iopub.execute_input":"2024-09-23T10:39:11.520556Z","iopub.status.idle":"2024-09-23T10:39:11.541771Z","shell.execute_reply.started":"2024-09-23T10:39:11.520514Z","shell.execute_reply":"2024-09-23T10:39:11.540761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p1 = [(x, f\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/{x}\") for x in part_1]\nmeta_obj = { p[0]: { 'folder_path': p[1], \n                    'SeriesInstanceUIDs': [] \n                   } \n            for p in p1 }\n\n\nfor m in meta_obj:\n    meta_obj[m]['SeriesInstanceUIDs'] = list(\n        filter(lambda x: x.find('.DS') == -1, \n               os.listdir(meta_obj[m]['folder_path'])\n              )\n    )\n\n# grabs the correspoding series descriptions\nfor k in tqdm(meta_obj):\n    for s in meta_obj[k]['SeriesInstanceUIDs']:\n        if 'SeriesDescriptions' not in meta_obj[k]:\n            meta_obj[k]['SeriesDescriptions'] = []\n        try:\n            meta_obj[k]['SeriesDescriptions'].append(\n                df_meta_f[(df_meta_f['study_id'] == int(k)) & \n                (df_meta_f['series_id'] == int(s))]['series_description'].iloc[0])\n        except:\n            print(\"Failed on\", s, k)","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:11.543630Z","iopub.execute_input":"2024-09-23T10:39:11.544190Z","iopub.status.idle":"2024-09-23T10:39:22.295713Z","shell.execute_reply.started":"2024-09-23T10:39:11.544083Z","shell.execute_reply":"2024-09-23T10:39:22.294511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_obj[list(meta_obj.keys())[1]]","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:22.300976Z","iopub.execute_input":"2024-09-23T10:39:22.301376Z","iopub.status.idle":"2024-09-23T10:39:22.308831Z","shell.execute_reply.started":"2024-09-23T10:39:22.301337Z","shell.execute_reply":"2024-09-23T10:39:22.307817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient = train.iloc[1]\npatient","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:22.309994Z","iopub.execute_input":"2024-09-23T10:39:22.310426Z","iopub.status.idle":"2024-09-23T10:39:22.325767Z","shell.execute_reply.started":"2024-09-23T10:39:22.310375Z","shell.execute_reply":"2024-09-23T10:39:22.324376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ptobj = meta_obj[str(patient['study_id'])]\nprint(ptobj)","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:22.327576Z","iopub.execute_input":"2024-09-23T10:39:22.328128Z","iopub.status.idle":"2024-09-23T10:39:22.335654Z","shell.execute_reply.started":"2024-09-23T10:39:22.328050Z","shell.execute_reply":"2024-09-23T10:39:22.334507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get data into the format\n\"\"\"\nim_list_dcm = {\n    '{SeriesInstanceUID}': {\n        'images': [\n            {'SOPInstanceUID': ...,\n             'dicom': PyDicom object\n            },\n            ...,\n        ],\n        'description': # SeriesDescription\n    },\n    ...\n}\n\"\"\"\nim_list_dcm = {}\nfor idx, i in enumerate(ptobj['SeriesInstanceUIDs']):\n    im_list_dcm[i] = {'images': [], 'description': ptobj['SeriesDescriptions'][idx]}\n    images = glob.glob(f\"{ptobj['folder_path']}/{ptobj['SeriesInstanceUIDs'][idx]}/*.dcm\")\n    for j in sorted(images, key=lambda x: int(x.split('/')[-1].replace('.dcm', ''))):\n        im_list_dcm[i]['images'].append({\n            'SOPInstanceUID': j.split('/')[-1].replace('.dcm', ''), \n            'dicom': pydicom.dcmread(j) })","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:22.336951Z","iopub.execute_input":"2024-09-23T10:39:22.337334Z","iopub.status.idle":"2024-09-23T10:39:22.995458Z","shell.execute_reply.started":"2024-09-23T10:39:22.337290Z","shell.execute_reply":"2024-09-23T10:39:22.994227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to display images\ndef display_images(images, title, max_images_per_row=4):\n    # Calculate the number of rows needed\n    num_images = len(images)\n    num_rows = (num_images + max_images_per_row - 1) // max_images_per_row  # Ceiling division\n\n    # Create a subplot grid\n    fig, axes = plt.subplots(num_rows, max_images_per_row, figsize=(5, 1.5 * num_rows))\n    \n    # Flatten axes array for easier looping if there are multiple rows\n    if num_rows > 1:\n        axes = axes.flatten()\n    else:\n        axes = [axes]  # Make it iterable for consistency\n\n    # Plot each image\n    for idx, image in enumerate(images):\n        ax = axes[idx]\n        ax.imshow(image, cmap='gray')  # Assuming grayscale for simplicity, change cmap as needed\n        ax.axis('off')  # Hide axes\n\n    # Turn off unused subplots\n    for idx in range(num_images, len(axes)):\n        axes[idx].axis('off')\n    fig.suptitle(title, fontsize=16)\n\n    plt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:22.996931Z","iopub.execute_input":"2024-09-23T10:39:22.997420Z","iopub.status.idle":"2024-09-23T10:39:23.007334Z","shell.execute_reply.started":"2024-09-23T10:39:22.997370Z","shell.execute_reply":"2024-09-23T10:39:23.005952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in im_list_dcm:\n    display_images([x['dicom'].pixel_array for x in im_list_dcm[i]['images']], \n                   im_list_dcm[i]['description'])","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:23.009205Z","iopub.execute_input":"2024-09-23T10:39:23.010130Z","iopub.status.idle":"2024-09-23T10:39:29.991010Z","shell.execute_reply.started":"2024-09-23T10:39:23.010077Z","shell.execute_reply":"2024-09-23T10:39:29.989882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n\ndf_coor = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')\n\ndf_coor.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:29.992458Z","iopub.execute_input":"2024-09-23T10:39:29.992821Z","iopub.status.idle":"2024-09-23T10:39:30.145351Z","shell.execute_reply.started":"2024-09-23T10:39:29.992783Z","shell.execute_reply":"2024-09-23T10:39:30.143886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_coor_on_img(c, i, title):\n    center_coordinates = (int(c['x']), int(c['y']))\n    radius = 10\n    color = (255, 0, 0)  # Red color in BGR\n    thickness = 2\n    IMG = i['dicom'].pixel_array\n    IMG_normalized = cv2.normalize(IMG, None, alpha=0, beta=255, norm_type=cv2.NORM_MINMAX, dtype=cv2.CV_8U)\n    \n    IMG_with_circle = cv2.circle(IMG_normalized.copy(), center_coordinates, radius, color, thickness)\n    \n    # Convert the image from BGR to RGB for correct color display in matplotlib\n    IMG_with_circle = cv2.cvtColor(IMG_with_circle, cv2.COLOR_BGR2RGB)\n    \n    # Display the image\n    plt.imshow(IMG_with_circle)\n    plt.axis('off')  # Turn off axis numbers and ticks\n    plt.title(title)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:30.146772Z","iopub.execute_input":"2024-09-23T10:39:30.147170Z","iopub.status.idle":"2024-09-23T10:39:30.156267Z","shell.execute_reply.started":"2024-09-23T10:39:30.147131Z","shell.execute_reply":"2024-09-23T10:39:30.155039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"coor_entries = df_coor[df_coor['study_id'] == int(patient['study_id'])]","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:30.157698Z","iopub.execute_input":"2024-09-23T10:39:30.158123Z","iopub.status.idle":"2024-09-23T10:39:30.168849Z","shell.execute_reply.started":"2024-09-23T10:39:30.158083Z","shell.execute_reply":"2024-09-23T10:39:30.167808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Only showing severe cases for this patient\")\nfor idc, c in coor_entries.iterrows():\n    for i in im_list_dcm[str(c['series_id'])]['images']:\n        if int(i['SOPInstanceUID']) == int(c['instance_number']):\n            try:\n                patient_severity = patient[\n                    f\"{c['condition'].lower().replace(' ', '_')}_{c['level'].lower().replace('/', '_')}\"\n                ]\n            except Exception as e:\n                patient_severity = \"unknown severity\"\n            title = f\"{i['SOPInstanceUID']} \\n{c['level']}, {c['condition']}: {patient_severity} \\n{c['x']}, {c['y']}\"\n            if patient_severity == 'Severe':\n                display_coor_on_img(c, i, title)","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:30.170289Z","iopub.execute_input":"2024-09-23T10:39:30.170665Z","iopub.status.idle":"2024-09-23T10:39:30.694938Z","shell.execute_reply.started":"2024-09-23T10:39:30.170627Z","shell.execute_reply":"2024-09-23T10:39:30.693719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_all_coor_on_img(coordinates, image, title):\n    IMG = image['dicom'].pixel_array\n    IMG_normalized = cv2.normalize(IMG, None, alpha=0, beta=255, norm_type=cv2.NORM_MINMAX, dtype=cv2.CV_8U)\n\n    # Loop through each coordinate and draw a circle on the image\n    for _, c in coordinates.iterrows():\n        center_coordinates = (int(c['x']), int(c['y']))\n        radius = 10\n        color = (255, 0, 0)  # Red color in BGR\n        thickness = 2\n        IMG_normalized = cv2.circle(IMG_normalized.copy(), center_coordinates, radius, color, thickness)\n\n    # Convert the image from BGR to RGB for correct color display in matplotlib\n    IMG_with_circles = cv2.cvtColor(IMG_normalized, cv2.COLOR_BGR2RGB)\n\n    # Display the image\n    plt.imshow(IMG_with_circles)\n    plt.axis('off')  # Turn off axis numbers and ticks\n    plt.title(title)\n    plt.show()\n\n\n# Filter coordinates for a specific patient and display all severe cases\ncoor_entries = df_coor[df_coor['study_id'] == int(patient['study_id'])]\nprint(\"Only showing severe cases for this patient\")\n\nfor idc, c in coor_entries.iterrows():\n    for i in im_list_dcm[str(c['series_id'])]['images']:\n        if int(i['SOPInstanceUID']) == int(c['instance_number']):\n            try:\n                patient_severity = patient[\n                    f\"{c['condition'].lower().replace(' ', '_')}_{c['level'].lower().replace('/', '_')}\"\n                ]\n            except Exception as e:\n                patient_severity = \"unknown severity\"\n\n            # Title for the image\n            title = f\"{i['SOPInstanceUID']} \\n{c['level']}, {c['condition']}: {patient_severity}\"\n\n            # Filter to get all coordinates for this image (series_id and instance_number)\n            image_coors = coor_entries[\n                (coor_entries['series_id'] == c['series_id']) &\n                (coor_entries['instance_number'] == c['instance_number'])\n            ]\n\n#             # Only show if the severity is 'Severe'\n            if patient_severity == 'Severe':\n                display_all_coor_on_img(image_coors, i, title)\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:30.696609Z","iopub.execute_input":"2024-09-23T10:39:30.696997Z","iopub.status.idle":"2024-09-23T10:39:31.209251Z","shell.execute_reply.started":"2024-09-23T10:39:30.696940Z","shell.execute_reply":"2024-09-23T10:39:31.208079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# File paths (you can modify them according to your folder structure)\ntrain_csv_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv'\ntrain_label_coordinates_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv'\ntrain_series_descriptions_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv'\noutput_csv_path = '/kaggle/working/RSNA_consolidated_train_data.csv'\n\n# Load the CSV files\ntrain_df = pd.read_csv(train_csv_path)\ntrain_label_coordinates_df = pd.read_csv(train_label_coordinates_path)\ntrain_series_descriptions_df = pd.read_csv(train_series_descriptions_path)\n\n# Ensure that condition names are consistent between files\n# Standardize the condition naming in the label coordinates file (e.g., lowercase and underscores)\ntrain_label_coordinates_df['condition'] = train_label_coordinates_df['condition'].str.lower().str.replace(' ', '_')\n\n# Standardize level format if needed (e.g., making sure L1/L2 is consistent across files)\ntrain_label_coordinates_df['level'] = train_label_coordinates_df['level'].str.lower()\n\n# Standardize the column names for melted train data (condition names)\ncondition_columns = [\n    # Spinal canal stenosis conditions\n    'spinal_canal_stenosis_l1_l2', 'spinal_canal_stenosis_l2_l3', \n    'spinal_canal_stenosis_l3_l4', 'spinal_canal_stenosis_l4_l5', \n    'spinal_canal_stenosis_l5_s1',\n    \n    # Left neural foraminal narrowing conditions\n    'left_neural_foraminal_narrowing_l1_l2', 'left_neural_foraminal_narrowing_l2_l3', \n    'left_neural_foraminal_narrowing_l3_l4', 'left_neural_foraminal_narrowing_l4_l5', \n    'left_neural_foraminal_narrowing_l5_s1',\n    \n    # Right neural foraminal narrowing conditions\n    'right_neural_foraminal_narrowing_l1_l2', 'right_neural_foraminal_narrowing_l2_l3',\n    'right_neural_foraminal_narrowing_l3_l4', 'right_neural_foraminal_narrowing_l4_l5',\n    'right_neural_foraminal_narrowing_l5_s1',\n    \n    # Left subarticular stenosis conditions\n    'left_subarticular_stenosis_l1_l2', 'left_subarticular_stenosis_l2_l3', \n    'left_subarticular_stenosis_l3_l4', 'left_subarticular_stenosis_l4_l5',\n    'left_subarticular_stenosis_l5_s1',\n    \n    # Right subarticular stenosis conditions\n    'right_subarticular_stenosis_l1_l2', 'right_subarticular_stenosis_l2_l3', \n    'right_subarticular_stenosis_l3_l4', 'right_subarticular_stenosis_l4_l5',\n    'right_subarticular_stenosis_l5_s1'\n]\n\n# Melt the condition columns to make condition and severity values accessible\ntrain_df_melted = pd.melt(\n    train_df, \n    id_vars='study_id', \n    value_vars=condition_columns,\n    var_name='condition_level',\n    value_name='severity'\n)\n\n\ntrain_df_melted","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.211153Z","iopub.execute_input":"2024-09-23T10:39:31.211523Z","iopub.status.idle":"2024-09-23T10:39:31.411675Z","shell.execute_reply.started":"2024-09-23T10:39:31.211484Z","shell.execute_reply":"2024-09-23T10:39:31.410439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_series_descriptions_df","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.413209Z","iopub.execute_input":"2024-09-23T10:39:31.413598Z","iopub.status.idle":"2024-09-23T10:39:31.428803Z","shell.execute_reply.started":"2024-09-23T10:39:31.413557Z","shell.execute_reply":"2024-09-23T10:39:31.427370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_label_coordinates_df","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.430193Z","iopub.execute_input":"2024-09-23T10:39:31.430544Z","iopub.status.idle":"2024-09-23T10:39:31.451438Z","shell.execute_reply.started":"2024-09-23T10:39:31.430507Z","shell.execute_reply":"2024-09-23T10:39:31.450165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Standardize the naming of the conditions and levels in `train_df_melted` to match `train_label_coordinates_df`\ntrain_df_melted['condition'] = train_df_melted['condition_level'].apply(lambda x: '_'.join(x.split('_')[:-2]))\ntrain_df_melted['level'] = train_df_melted['condition_level'].apply(lambda x: '/'.join(x.split('_')[-2:]))  # Fix the level formatting\ntrain_df_melted['condition'] = train_df_melted['condition'].str.lower()  # Ensure consistent case and format\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.453401Z","iopub.execute_input":"2024-09-23T10:39:31.453893Z","iopub.status.idle":"2024-09-23T10:39:31.594923Z","shell.execute_reply.started":"2024-09-23T10:39:31.453838Z","shell.execute_reply":"2024-09-23T10:39:31.593874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge train_label_coordinates_df with train_series_descriptions_df\nmerged_df = pd.merge(train_label_coordinates_df, train_series_descriptions_df, on=['study_id', 'series_id'], how='left')\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.596226Z","iopub.execute_input":"2024-09-23T10:39:31.596551Z","iopub.status.idle":"2024-09-23T10:39:31.621839Z","shell.execute_reply.started":"2024-09-23T10:39:31.596515Z","shell.execute_reply":"2024-09-23T10:39:31.620496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"merged_df","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.623369Z","iopub.execute_input":"2024-09-23T10:39:31.623764Z","iopub.status.idle":"2024-09-23T10:39:31.642068Z","shell.execute_reply.started":"2024-09-23T10:39:31.623723Z","shell.execute_reply":"2024-09-23T10:39:31.640766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df_melted","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.643606Z","iopub.execute_input":"2024-09-23T10:39:31.644086Z","iopub.status.idle":"2024-09-23T10:39:31.660941Z","shell.execute_reply.started":"2024-09-23T10:39:31.644031Z","shell.execute_reply":"2024-09-23T10:39:31.659787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Perform the merge\nfinal_df = pd.merge(merged_df, train_df_melted, on=['study_id', 'condition', 'level'], how='left')\n\n# Check where the merge resulted in NaN values in the 'severity' or 'condition_level' column\n# This indicates that these rows from `merged_df` didn't find a match in `train_df_melted`\nmismatched_rows = final_df[final_df['severity'].isna() | final_df['condition_level'].isna()]\n\n# Display the mismatched rows for inspection\nprint(f\"Number of mismatched rows: {len(mismatched_rows)}\")\nprint(mismatched_rows[['study_id', 'condition', 'level']].drop_duplicates().head())  # Display study_id, condition, and level\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.668207Z","iopub.execute_input":"2024-09-23T10:39:31.668581Z","iopub.status.idle":"2024-09-23T10:39:31.767231Z","shell.execute_reply.started":"2024-09-23T10:39:31.668544Z","shell.execute_reply":"2024-09-23T10:39:31.765860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Save the final merged dataframe as a CSV\nfinal_df.to_csv(output_csv_path, index=False)\n\nprint(f\"Consolidated data saved to {output_csv_path}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:31.768681Z","iopub.execute_input":"2024-09-23T10:39:31.769167Z","iopub.status.idle":"2024-09-23T10:39:32.393391Z","shell.execute_reply.started":"2024-09-23T10:39:31.769112Z","shell.execute_reply":"2024-09-23T10:39:32.392258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data = final_df.isnull().sum()\nprint(missing_data)\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:32.394881Z","iopub.execute_input":"2024-09-23T10:39:32.395322Z","iopub.status.idle":"2024-09-23T10:39:32.432637Z","shell.execute_reply.started":"2024-09-23T10:39:32.395281Z","shell.execute_reply":"2024-09-23T10:39:32.431487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nimport glob\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport math\n\n# Load the consolidated CSV file\nconsolidated_csv_path = '/kaggle/input/rsna-consolidated-training-data/RSNA_consolidated_train_data.csv'\ndf = pd.read_csv(consolidated_csv_path)\n\ndef get_images_for_row(row):\n    # Extract study and series information\n    study_id = row['study_id']\n    series_id = row['series_id']\n    description = row['series_description']\n\n    # Path to the folder where the DICOM images are stored (adjust this if necessary)\n    folder_path = f\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/{study_id}/{series_id}/\"\n\n    # Create a dictionary to hold DICOM files\n    im_list_dcm = {'images': [], 'description': description}\n\n    # Use glob to find all DICOM files for this series\n    images = glob.glob(f\"{folder_path}/*.dcm\")\n    \n    # Sort the DICOM files by instance number (or by filename)\n    for dicom_file in sorted(images, key=lambda x: int(x.split('/')[-1].replace('.dcm', ''))):\n        # Read the DICOM file\n        dicom_data = pydicom.dcmread(dicom_file)\n\n        # Store the DICOM data\n        im_list_dcm['images'].append({\n            'SOPInstanceUID': dicom_file.split('/')[-1].replace('.dcm', ''),\n            'dicom': dicom_data\n        })\n\n    return im_list_dcm\n\ndef plot_images_in_grid(rows, images_data):\n    # Get the number of images and calculate the grid size (rows x columns)\n    num_images = len(images_data['images'])\n    grid_size = math.ceil(math.sqrt(num_images))  # Get grid size (square root)\n\n    # Create a grid of subplots\n    fig, axes = plt.subplots(grid_size, grid_size, figsize=(15, 15))\n\n    # Flatten axes for easy iteration\n    axes = axes.flatten()\n\n    # Loop through all the images and their annotations\n    for idx, (image_data, ax) in enumerate(zip(images_data['images'], axes)):\n        if idx >= num_images:\n            break  # If more axes than images, break the loop\n\n        dicom_image = image_data['dicom'].pixel_array\n        ax.imshow(dicom_image, cmap='gray')\n        ax.set_title(f\"Image {idx + 1}\")\n        ax.axis('off')\n\n        # Add annotations for each image (condition, level, severity, coordinates)\n        for _, row in rows.iterrows():\n            if int(row['instance_number']) == int(image_data['SOPInstanceUID']):\n                x, y = row['x'], row['y']\n                condition = row['condition']\n                level = row['level']\n                severity = row['severity']\n\n                # Draw the annotation (ROI)\n                rect = patches.Rectangle((x-25, y-25), 50, 50, linewidth=2, edgecolor='r', facecolor='none')\n                ax.add_patch(rect)\n\n                # Add text annotation\n                ax.text(x, y, f\"{condition} ({level}): {severity}\", color='yellow', fontsize=8,\n                        bbox=dict(facecolor='black', alpha=0.5))\n\n    # Remove any extra subplots in the grid\n    for ax in axes[num_images:]:\n        ax.remove()\n\n    plt.tight_layout()\n    plt.show()\n\n# Example: Fetch images and details for a specific study_id and series_id\nstudy_id = 4003253 #df.iloc[8000]['study_id']\nseries_id = 702807833 #df.iloc[8000]['series_id']\n        \nprint('study_id :',study_id)\nprint('series_id :',series_id)\n# Filter the dataframe to get all rows corresponding to the same study and series\nimage_entries = df[(df['study_id'] == study_id) & (df['series_id'] == series_id)]\n\n# Get the DICOM images for this study and series\nimages_data = get_images_for_row(image_entries.iloc[0])  # Use the first row to fetch DICOM images\n# print(images_data)\n# Show the description and number of images found\nprint(f\"Series Description: {images_data['description']}\")\nprint(f\"Number of DICOM images found: {len(images_data['images'])}\")\n\n# Plot all the images in a grid with annotations for condition, level, severity, and ROI\nplot_images_in_grid(image_entries, images_data)\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:32.434580Z","iopub.execute_input":"2024-09-23T10:39:32.434990Z","iopub.status.idle":"2024-09-23T10:39:36.291283Z","shell.execute_reply.started":"2024-09-23T10:39:32.434931Z","shell.execute_reply":"2024-09-23T10:39:36.289831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nimport glob\nimport pandas as pd\nimport cv2\nimport os\n\n# Load the consolidated CSV file\nconsolidated_csv_path = '/kaggle/input/rsna-consolidated-training-data/RSNA_consolidated_train_data.csv'\ndf = pd.read_csv(consolidated_csv_path)\n\ndef get_images_for_row(row):\n    # Extract study and series information\n    study_id = row['study_id']\n    series_id = row['series_id']\n    description = row['series_description']\n\n    # Path to the folder where the DICOM images are stored (adjust this if necessary)\n    folder_path = f\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/{study_id}/{series_id}/\"\n\n    # Create a dictionary to hold DICOM files\n    im_list_dcm = {'images': [], 'description': description}\n\n    # Use glob to find all DICOM files for this series\n    images = glob.glob(f\"{folder_path}/*.dcm\")\n    \n    # Sort the DICOM files by instance number (or by filename)\n    for dicom_file in sorted(images, key=lambda x: int(x.split('/')[-1].replace('.dcm', ''))):\n        # Read the DICOM file\n        dicom_data = pydicom.dcmread(dicom_file)\n\n        # Store the DICOM data\n        im_list_dcm['images'].append({\n            'SOPInstanceUID': dicom_file.split('/')[-1].replace('.dcm', ''),\n            'dicom': dicom_data\n        })\n\n    return im_list_dcm\n\ndef create_video_from_images(rows, images_data, output_video_path):\n    # Get the frame dimensions from the first image\n    first_dicom_image = images_data['images'][0]['dicom'].pixel_array\n    height, width = first_dicom_image.shape\n\n    # Define the video codec and create VideoWriter object for .webm\n    fourcc = cv2.VideoWriter_fourcc(*'VP80')  # Codec for .webm files\n    fps = 1  # Frames per second\n    out = cv2.VideoWriter(output_video_path, fourcc, fps, (width, height))\n\n    # Loop through all the images and add them to the video\n    for image_data in images_data['images']:\n        dicom_image = image_data['dicom'].pixel_array\n\n        # Normalize the image to 8-bit grayscale\n        dicom_image_normalized = cv2.normalize(dicom_image, None, 0, 255, cv2.NORM_MINMAX).astype('uint8')\n\n        # Convert grayscale image to BGR format (for video writing)\n        dicom_image_bgr = cv2.cvtColor(dicom_image_normalized, cv2.COLOR_GRAY2BGR)\n\n        # Add annotations for each image (condition, level, severity, coordinates)\n        for _, row in rows.iterrows():\n            if int(row['instance_number']) == int(image_data['SOPInstanceUID']):\n                x, y = row['x'], row['y']\n                condition = row['condition']\n                level = row['level']\n                severity = row['severity']\n\n                # Draw the annotation (ROI)\n                rect = cv2.rectangle(dicom_image_bgr, (int(x-10), int(y-10)), (int(x+10), int(y+10)), (0, 0, 255), 2)\n\n                # Add text annotation\n                cv2.putText(dicom_image_bgr, f\"{condition} ({level}): {severity}\", (int(x), int(y)), \n                            cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 255, 255), 1, cv2.LINE_AA)\n\n        # Write the image to the video\n        out.write(dicom_image_bgr)\n\n    # Release the VideoWriter object\n    out.release()\n    print(f\"Video saved to {output_video_path}\")\n\n# Example: Fetch images and details for a specific study_id and series_id\nstudy_id = df.iloc[8800]['study_id']\nseries_id = df.iloc[8800]['series_id']\n\n# Filter the dataframe to get all rows corresponding to the same study and series\nimage_entries = df[(df['study_id'] == study_id) & (df['series_id'] == series_id)]\n\n# Get the DICOM images for this study and series\nimages_data = get_images_for_row(image_entries.iloc[0])  # Use the first row to fetch DICOM images\n\n# Define the output video path as a .webm file\noutput_video_path = '/kaggle/working/dicom_video.webm'\n\n# Create the video\ncreate_video_from_images(image_entries, images_data, output_video_path)\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:36.292696Z","iopub.execute_input":"2024-09-23T10:39:36.293088Z","iopub.status.idle":"2024-09-23T10:39:36.944614Z","shell.execute_reply.started":"2024-09-23T10:39:36.293048Z","shell.execute_reply":"2024-09-23T10:39:36.943257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.display import Video, display\n\n# Path to the saved video\noutput_video_path = '/kaggle/working/dicom_video.webm'\n\n# Display the video in the notebook with controls to pause/play\ndisplay(Video(output_video_path, embed=True, width=600, height=400))\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:36.946131Z","iopub.execute_input":"2024-09-23T10:39:36.946523Z","iopub.status.idle":"2024-09-23T10:39:36.966352Z","shell.execute_reply.started":"2024-09-23T10:39:36.946484Z","shell.execute_reply":"2024-09-23T10:39:36.965018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport pandas as pd\nimport pydicom\nimport numpy as np\n\n# Define the paths for video folders\noutput_dir = '/kaggle/working/output_videos'\nsagittal_t2_stir_dir = os.path.join(output_dir, 'Sagittal_T2_STIR')\naxial_t2_dir = os.path.join(output_dir, 'Axial_T2')\nsagittal_t1_dir = os.path.join(output_dir, 'Sagittal_T1')\n\n# Create directories if they don't exist\nos.makedirs(sagittal_t2_stir_dir, exist_ok=True)\nos.makedirs(axial_t2_dir, exist_ok=True)\nos.makedirs(sagittal_t1_dir, exist_ok=True)\n\n# Load the CSV file\nconsolidated_csv_path = '/kaggle/input/rsna-consolidated-training-data/RSNA_consolidated_train_data.csv'\ndf = pd.read_csv(consolidated_csv_path)\n\n# Function to create and save video\ndef create_and_save_video(image_files, video_path, fps=1, frame_size=(224, 224)):\n    # Define the codec and create a VideoWriter object\n    fourcc = cv2.VideoWriter_fourcc(*'mp4v')\n    video_writer = cv2.VideoWriter(video_path, fourcc, fps, frame_size)\n\n    for image_file in image_files:\n        dicom_data = pydicom.dcmread(image_file)\n        image = dicom_data.pixel_array\n        if image.dtype != np.uint8:\n            image = cv2.normalize(image, None, 0, 255, cv2.NORM_MINMAX).astype(np.uint8)\n        if len(image.shape) == 2:\n            image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n\n        # Resize the image to the target frame size\n        image_resized = cv2.resize(image, frame_size)\n        video_writer.write(image_resized)\n\n    # Release the video writer\n    video_writer.release()\n\n# Loop through each unique combination of study_id and series_id\nunique_videos = df[['study_id', 'series_id', 'series_description']].drop_duplicates()\n\nfor _, row in unique_videos.iterrows():\n    study_id = row['study_id']\n    series_id = row['series_id']\n    series_description = row['series_description']\n\n    # Get all instance numbers for this particular study and series\n    series_rows = df[(df['study_id'] == study_id) & (df['series_id'] == series_id)]\n    instance_numbers = sorted(series_rows['instance_number'].unique())  # Ensure sorting by instance_number\n\n    # Prepare to store the images\n    image_files = []\n\n    # Loop through each instance number and get the corresponding DICOM file\n    for instance_number in instance_numbers:\n        image_path = f\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/{study_id}/{series_id}/{instance_number}.dcm\"\n        if os.path.exists(image_path):\n            image_files.append(image_path)\n\n    # If there are images for this series, create the video\n    if image_files:\n        # Choose the appropriate directory based on the series description\n        if series_description == 'Sagittal T2/STIR':\n            save_dir = sagittal_t2_stir_dir\n        elif series_description == 'Axial T2':\n            save_dir = axial_t2_dir\n        elif series_description == 'Sagittal T1':\n            save_dir = sagittal_t1_dir\n        else:\n            print(f\"Unknown series description: {series_description} for study {study_id}, series {series_id}\")\n            continue\n\n        # Define the path to save the video\n        video_filename = f\"{study_id}_{series_id}.mp4\"\n        video_path = os.path.join(save_dir, video_filename)\n\n        # Create and save the video\n        create_and_save_video(image_files, video_path)\n        print(f\"Saved video for study {study_id}, series {series_id} to {video_path}\")\n    else:\n        print(f\"No images found for study {study_id}, series {series_id}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:39:36.967954Z","iopub.execute_input":"2024-09-23T10:39:36.968471Z","iopub.status.idle":"2024-09-23T10:50:59.974682Z","shell.execute_reply.started":"2024-09-23T10:39:36.968403Z","shell.execute_reply":"2024-09-23T10:50:59.973524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# Load the CSV file\nconsolidated_csv_path = '/kaggle/input/rsna-consolidated-training-data/RSNA_consolidated_train_data.csv'\ndf = pd.read_csv(consolidated_csv_path)\n\n# Group by study_id and series_id and count the number of images (i.e., instance_number)\nimage_counts = df.groupby(['study_id', 'series_id']).size().reset_index(name='image_count')\n\n# Sort the image counts in descending order\nsorted_image_counts = image_counts.sort_values(by='image_count', ascending=False)\n\n# Get the top 5 entries with the maximum image counts\ntop_5_image_counts = sorted_image_counts.head(5)\n\n# Display the result\nprint(\"Top 5 study_id and series_id with the most images:\")\nprint(top_5_image_counts)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:50:59.976047Z","iopub.execute_input":"2024-09-23T10:50:59.976418Z","iopub.status.idle":"2024-09-23T10:51:00.177011Z","shell.execute_reply.started":"2024-09-23T10:50:59.976377Z","shell.execute_reply":"2024-09-23T10:51:00.175571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport glob\n\n# Define the path to the folder where the training images are stored\nimage_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/'\n\n# Dictionary to store the number of images for each study_id and series_id\nimage_counts = {}\n\n# Traverse through the folders and count the images\nfor study_id in os.listdir(image_dir):\n    study_path = os.path.join(image_dir, study_id)\n    if os.path.isdir(study_path):\n        for series_id in os.listdir(study_path):\n            series_path = os.path.join(study_path, series_id)\n            if os.path.isdir(series_path):\n                # Count the number of .dcm files in the series directory\n                num_images = len(glob.glob(os.path.join(series_path, '*.dcm')))\n                \n                # Store the count in the dictionary with (study_id, series_id) as key\n                image_counts[(study_id, series_id)] = num_images\n\n# Sort the dictionary by the image count in descending order\nsorted_image_counts = sorted(image_counts.items(), key=lambda x: x[1], reverse=True)\n\n\n\n# Optional: Save the result to a CSV file\n# import pandas as pd\n# top_5_df = pd.DataFrame(top_5_image_counts, columns=['study_id_series_id', 'image_count'])\n# top_5_df[['study_id', 'series_id']] = top_5_df['study_id_series_id'].apply(pd.Series)\n# top_5_df.drop(columns=['study_id_series_id'], inplace=True)\n# top_5_df.to_csv('/kaggle/working/top_5_image_counts_from_folder.csv', index=False)\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:51:00.178857Z","iopub.execute_input":"2024-09-23T10:51:00.179357Z","iopub.status.idle":"2024-09-23T10:51:32.062657Z","shell.execute_reply.started":"2024-09-23T10:51:00.179294Z","shell.execute_reply":"2024-09-23T10:51:32.060661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the top 5 entries with the highest image counts\ntop_5_image_counts = sorted_image_counts[:2500]\n\n# Print the top 5 study_id and series_id with the most images\nprint(\"Top 5 study_id and series_id with the most images:\")\nfor (study_id, series_id), count in top_5_image_counts:\n    print(f\"Study ID: {study_id}, Series ID: {series_id}, Image Count: {count}\")","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:51:32.065242Z","iopub.execute_input":"2024-09-23T10:51:32.065835Z","iopub.status.idle":"2024-09-23T10:51:32.109145Z","shell.execute_reply.started":"2024-09-23T10:51:32.065765Z","shell.execute_reply":"2024-09-23T10:51:32.107954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Assuming your consolidated CSV contains paths and labels\ncsv_file = '/kaggle/input/rsna-consolidated-training-data/RSNA_consolidated_train_data.csv'  # Update this to the correct CSV path\nimage_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'  # Update this to the correct directory where images are stored\n\n# Load the consolidated CSV\ndf = pd.read_csv(csv_file)\n\n# Define the 5 types (conditions) we want to sample from\ntypes = [\n    'spinal_canal_stenosis',\n    'left_neural_foraminal_narrowing',\n    'right_neural_foraminal_narrowing',\n    'left_subarticular_stenosis',\n    'right_subarticular_stenosis'\n]\n\n# Function to get 200 images from each type\ndef get_images(df, image_dir, types, sample_size=200):\n    images = []\n    for condition in types:\n        # Filter dataframe to only get rows for the current condition\n        df_condition = df[df['condition'] == condition]\n\n        # Randomly sample 200 images for this condition\n        sampled_df = df_condition.sample(sample_size, random_state=42)\n\n        # Load and append images to the list\n        for _, row in sampled_df.iterrows():\n            # Construct the image path\n            study_id = row['study_id']\n            series_id = row['series_id']\n            instance_number = row['instance_number']\n             # Build the image path\n            image_path = os.path.join(image_dir, f\"{study_id}/{series_id}/{instance_number}.dcm\")\n\n            # Load the DICOM image (if needed, convert it to a format compatible with cv2)\n            dicom_data = pydicom.dcmread(image_path)\n            image = dicom_data.pixel_array\n\n            # Convert the image to 8-bit if necessary\n            if image.dtype != np.uint8:\n                image = cv2.normalize(image, None, 0, 255, cv2.NORM_MINMAX).astype(np.uint8)\n\n            # Convert grayscale to RGB (if necessary)\n            if len(image.shape) == 2:  # Check if it's grayscale\n                image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n\n            # Resize the image to a fixed size (e.g., 224x224)\n            image = cv2.resize(image, (224, 224))\n\n            # Append to list\n            images.append(image)\n    \n    return np.array(images)\n\n# Get 1000 images (200 from each of the 5 types)\nimages = get_images(df, image_dir, types, sample_size=200)\n\n# Convert images to float and normalize\nimages = images.astype(np.float32) / 255.0\n\n# Calculate mean and std across the entire dataset (per channel)\nmean = np.mean(images, axis=(0, 1, 2))\nstd = np.std(images, axis=(0, 1, 2))\n\n# Print the calculated mean and std\nprint(f\"Mean: {mean}\")\nprint(f\"Standard Deviation: {std}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:51:32.111100Z","iopub.execute_input":"2024-09-23T10:51:32.111496Z","iopub.status.idle":"2024-09-23T10:51:58.355446Z","shell.execute_reply.started":"2024-09-23T10:51:32.111456Z","shell.execute_reply":"2024-09-23T10:51:58.354266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport pandas as pd\nimport numpy as np\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nfrom sklearn.model_selection import train_test_split\nimport torchvision.transforms as transforms\ntorch.backends.cudnn.benchmark = True\n\nseverity_mapping = {\n    'Normal/Mild': 0,\n    'Moderate': 1,\n    'Severe': 2\n}\n\n\nimport os\nimport pydicom\nimport torch\nfrom torch.utils.data import Dataset\nimport cv2\nimport numpy as np\nimport pandas as pd\n\nclass LumbarSpineVideoDataset(Dataset):\n    def __init__(self, dataframe, image_dir, transform=None, xy_mean=None, xy_std=None, max_frames=200, image_size=(224, 224)):\n        self.dataframe = dataframe\n        self.image_dir = image_dir\n        self.transform = transform\n        self.xy_mean = xy_mean\n        self.xy_std = xy_std\n        self.max_frames = max_frames\n        self.image_size = image_size\n        \n        # Mappings for condition, level, and severity encoding\n        self.condition_mapping = {\n            'spinal_canal_stenosis': 0, 'left_neural_foraminal_narrowing': 1,\n            'right_neural_foraminal_narrowing': 2, 'left_subarticular_stenosis': 3,\n            'right_subarticular_stenosis': 4\n        }\n\n        self.level_mapping = {\n            'l1/l2': 0, 'l2/l3': 1, 'l3/l4': 2, 'l4/l5': 3, 'l5/s1': 4\n        }\n\n        self.severity_mapping = {'Normal/Mild': 0, 'Moderate': 1, 'Severe': 2}\n\n    def __len__(self):\n        # Return the number of unique study and series combinations (each combination represents a video)\n        return len(self.dataframe[['study_id', 'series_id']].drop_duplicates())\n    \n    def __getitem__(self, idx):\n        unique_videos = self.dataframe[['study_id', 'series_id']].drop_duplicates()\n        selected_row = unique_videos.iloc[idx]\n        \n        study_id = selected_row['study_id']\n        series_id = selected_row['series_id']\n        \n        # Get all rows corresponding to the selected study and series (i.e., video)\n        video_rows = self.dataframe[\n            (self.dataframe['study_id'] == study_id) & \n            (self.dataframe['series_id'] == series_id)\n        ]\n        \n        # Load the frames and their annotations\n        frames = []\n        xy_list = []\n        condition_list = []\n        level_list = []\n        severity_list = []\n        frame_numbers = []\n        \n        for _, row in video_rows.iterrows():\n            instance_number = row['instance_number']\n            image_path = os.path.join(self.image_dir, f\"{study_id}/{series_id}/{instance_number}.dcm\")\n            \n            dicom_data = pydicom.dcmread(image_path)\n            image = dicom_data.pixel_array\n            \n            if image.dtype != np.uint8:\n                image = cv2.normalize(image, None, 0, 255, cv2.NORM_MINMAX).astype(np.uint8)\n            if len(image.shape) == 2:  # Convert grayscale to RGB\n                image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n\n            if self.transform:\n                image = self.transform(image)\n                \n            # Append image as a frame\n            frames.append(image)\n            \n            # Normalize (x, y) coordinates\n            x, y = row['x'], row['y']\n            if self.xy_mean is not None and self.xy_std is not None:\n                x = (x - self.xy_mean[0]) / self.xy_std[0]\n                y = (y - self.xy_mean[1]) / self.xy_std[1]\n            xy_list.append([x, y])\n            \n            # Encode condition, level, and severity\n            condition_list.append(self.condition_mapping[row['condition']])\n            level_list.append(self.level_mapping[row['level']])\n            severity_list.append(self.severity_mapping[row['severity']])\n            \n            # Store frame number\n            frame_numbers.append(instance_number)\n\n        # Pad the frames if they are fewer than max_frames\n        num_frames = len(frames)\n        if num_frames < self.max_frames:\n            padding_frames = [torch.ones_like(frames[0]) * 255 for _ in range(self.max_frames - num_frames)]  # White frames\n            frames += padding_frames\n        \n        # Stack the frames into a tensor (batch_size, 3, max_frames, height, width)\n        video_tensor = torch.stack(frames[:self.max_frames], dim=0).permute(3, 0, 1, 2)  # (3, max_frames, height, width)\n        \n        # Stack the xy, condition, level, and severity lists into tensors\n        xy_tensor = torch.tensor(xy_list, dtype=torch.float32).flatten()  # 5 pairs -> shape (10,)\n        condition_tensor = torch.tensor(condition_list, dtype=torch.long)  # shape (5,)\n        level_tensor = torch.tensor(level_list, dtype=torch.long)  # shape (5,)\n        severity_tensor = torch.tensor(severity_list, dtype=torch.long)  # shape (5,)\n        frame_number_tensor = torch.tensor(frame_numbers, dtype=torch.long)  # shape (5,)\n\n        return video_tensor, xy_tensor, condition_tensor, level_tensor, severity_tensor, frame_number_tensor\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:51:58.357839Z","iopub.execute_input":"2024-09-23T10:51:58.358365Z","iopub.status.idle":"2024-09-23T10:51:58.387667Z","shell.execute_reply.started":"2024-09-23T10:51:58.358307Z","shell.execute_reply":"2024-09-23T10:51:58.386261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom torchvision import transforms\nfrom torch.utils.data import DataLoader\n\n# Load the consolidated CSV file\nconsolidated_csv_path = '/kaggle/input/rsna-consolidated-training-data/RSNA_consolidated_train_data.csv'\ndf = pd.read_csv(consolidated_csv_path)\ndf = df.dropna()\n\n# Split the dataset into train and validation (70/30 split)\ntrain_df, val_df = train_test_split(df, test_size=0.3, random_state=42)\n\n# Calculate mean and std of x, y coordinates from the training data\nxy_mean = train_df[['x', 'y']].mean().values\nxy_std = train_df[['x', 'y']].std().values\n\n# Define transformations (must be defined before using in datasets)\ntransform = transforms.Compose([\n    transforms.ToPILImage(),\n    transforms.Resize((224, 224)),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.14617169]*3, std=[0.17529692]*3)\n])\n\n# Define the dataset and DataLoader for training and validation\nimage_dir = \"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images\"\n\ntrain_dataset = LumbarSpineVideoDataset(\n    dataframe=train_df,\n    image_dir=image_dir,\n    transform=transform,\n    xy_mean=xy_mean,\n    xy_std=xy_std\n)\n\nval_dataset = LumbarSpineVideoDataset(\n    dataframe=val_df,\n    image_dir=image_dir,\n    transform=transform,\n    xy_mean=xy_mean,\n    xy_std=xy_std\n)\n\n# Adjusted batch size for video data\nbatch_size = 2\n\nimport torch\n\n\n# Update DataLoader to use the custom collate function\ntrain_loader = DataLoader(train_dataset, batch_size=batch_size, shuffle=True, num_workers=8, pin_memory=True)\nval_loader = DataLoader(val_dataset, batch_size=batch_size, shuffle=False, num_workers=8, pin_memory=True)\n\n# Example: Iterate through a batch from the train_loader\nfor video, xy, condition, level, severity, frame_number in train_loader:\n    print(\"Video shape:\", video.shape)  # Expected: [batch_size, 3, max_frames, height, width]\n    print(\"XY tensor shape:\", xy.shape)  # Expected: [batch_size, num_coordinates]\n    print(\"Condition tensor shape:\", condition.shape)  # Expected: [batch_size, num_conditions]\n    print(\"Level tensor shape:\", level.shape)  # Expected: [batch_size, num_levels]\n    print(\"Severity tensor shape:\", severity.shape)  # Expected: [batch_size, num_severities]\n    print(\"Frame number tensor shape:\", frame_number.shape)  # Expected: [batch_size, num_frames]\n    break\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:51:58.389391Z","iopub.execute_input":"2024-09-23T10:51:58.389843Z","iopub.status.idle":"2024-09-23T10:52:10.627096Z","shell.execute_reply.started":"2024-09-23T10:51:58.389800Z","shell.execute_reply":"2024-09-23T10:52:10.623268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nfrom collections import Counter\nimport numpy as np\n\ndef calculate_class_weights(train_loader):\n    # Initialize a counter for all labels in the training dataset\n    label_counts = Counter()\n\n    # Loop through the train_loader and accumulate the count of each class\n    for _, labels in train_loader:\n        label_counts.update(labels.cpu().numpy())  # Convert tensors to numpy arrays for counting\n\n    # Calculate total number of samples\n    total_samples = sum(label_counts.values())\n\n    # Compute the class weights as the inverse of the frequency\n    class_weights = {cls: total_samples / count for cls, count in label_counts.items()}\n\n    # Convert the class weights to a tensor for use in PyTorch\n    weights_tensor = torch.tensor([class_weights[i] for i in range(len(class_weights))], dtype=torch.float32).to('cuda')\n\n    return weights_tensor\n\n# Example: Calculate class weights from the train_loader\nclass_weights = calculate_class_weights(train_loader)\n","metadata":{"execution":{"iopub.status.busy":"2024-09-23T10:52:10.630925Z","iopub.status.idle":"2024-09-23T10:52:10.632427Z","shell.execute_reply.started":"2024-09-23T10:52:10.631959Z","shell.execute_reply":"2024-09-23T10:52:10.632028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_weights","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Saving the model\ntorch.save(model.state_dict(), 'Best_model_weights.pth')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}