{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"dockerImageVersionId":30762,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**checking for missing values**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport os\nimport pydicom\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nfrom PIL import Image\nimport torchvision.transforms as transforms\n\n\n# File paths\ntrain_csv_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv'\ntrain_label_coords_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv'\n\n# Step 1: Load the CSV files\ntrain_df = pd.read_csv(train_csv_path)\nlabel_coords_df = pd.read_csv(train_label_coords_path)\n\n# Display the first few rows of each dataframe\nprint(\"Train DataFrame:\\n\", train_df.head())\nprint(\"\\nTrain Label Coordinates DataFrame:\\n\", label_coords_df.head())\n\n# Step 2: Data exploration\n# Check for missing values\nprint(\"\\nMissing values in Train DataFrame:\\n\", train_df.isnull().sum())\nprint(\"\\nMissing values in Label Coordinates DataFrame:\\n\", label_coords_df.isnull().sum())\n\n# Step 3: Summarizing the condition severity in the train_df\n\n# Iterate through relevant condition columns to get counts for each severity level\nseverity_columns = [col for col in train_df.columns if 'stenosis' in col or 'narrowing' in col]\nseverity_counts = {}\n\nfor col in severity_columns:\n    col_counts = train_df[col].value_counts()\n    severity_counts[col] = col_counts\n\n# Create a new DataFrame with the counts\nseverity_df = pd.DataFrame(severity_counts).T  # Transpose for better readability\n\n# Display severity counts\nprint(\"\\nCondition Severity Counts:\")\nprint(severity_df)\n\n# Step 4: Visualizing the distribution of conditions\n\n# Fill NaN values with zeros for plotting\nseverity_df = severity_df.fillna(0)\n\n# Plot the severity distribution\nseverity_df.plot(kind='bar', stacked=True, figsize=(15,7))\nplt.title('Distribution of Condition Severity across Different Conditions')\nplt.xlabel('Conditions')\nplt.ylabel('Count')\nplt.xticks(rotation=90)\nplt.show()\n\n# Step 5: Exploring the label coordinates dataframe\n\n# Count the number of unique levels and conditions\nnum_unique_conditions = label_coords_df['condition'].nunique()\nnum_unique_levels = label_coords_df['level'].nunique()\nprint(f\"\\nNumber of unique conditions: {num_unique_conditions}\")\nprint(f\"Number of unique spinal levels: {num_unique_levels}\")\n\n# Visualizing the distribution of spinal levels\nlabel_coords_df['level'].value_counts().plot(kind='bar', figsize=(10,6))\nplt.title('Distribution of Spinal Levels in Label Coordinates')\nplt.xlabel('Spinal Levels')\nplt.ylabel('Count')\nplt.xticks(rotation=90)\nplt.show()\n\n# Checking the x/y coordinate ranges\nprint(\"\\nX Coordinate Range:\")\nprint(f\"Min: {label_coords_df['x'].min()}, Max: {label_coords_df['x'].max()}\")\nprint(\"\\nY Coordinate Range:\")\nprint(f\"Min: {label_coords_df['y'].min()}, Max: {label_coords_df['y'].max()}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-09-08T23:39:27.858990Z","iopub.execute_input":"2024-09-08T23:39:27.859444Z","iopub.status.idle":"2024-09-08T23:39:29.235265Z","shell.execute_reply.started":"2024-09-08T23:39:27.859402Z","shell.execute_reply":"2024-09-08T23:39:29.234275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Distribution of Spinal Levels:\nThe spinal levels L1/L2 to L5/S1 appear to be evenly distributed, which indicates that each spinal level is equally represented in the dataset. This is beneficial for model training, as it ensures that the model will have balanced data across different spinal regions.\n2. Condition Severity Across Different Conditions:\n\"Normal/Mild\" dominates most columns: In almost all conditions, the majority of cases are labeled as \"Normal/Mild\". This suggests that most patients in the dataset have less severe degenerative spinal conditions.\n\"Moderate\" and \"Severe\" cases are less frequent: Conditions like spinal canal stenosis at L4/L5 and neural foraminal narrowing at L4/L5 and L5/S1 tend to have more \"Moderate\" and \"Severe\" cases compared to other conditions. These spinal levels are likely where patients are more prone to severe issues.\n3. Key Observations from the Bar Chart (Condition Severity):\nFor spinal canal stenosis:\nThe severity increases slightly at lower spinal levels. For example, L4/L5 has a higher number of \"Moderate\" and \"Severe\" cases compared to the upper levels like L1/L2.\nFor neural foraminal narrowing:\nThere is a clear increase in severity as you go down the spinal levels, with L4/L5 and L5/S1 showing a significant number of \"Moderate\" and \"Severe\" cases, indicating that these regions are more affected.\nFor subarticular stenosis:\nThe severity follows a similar trend, with lower spinal levels (L4/L5 and L5/S1) showing more \"Severe\" cases.\n4. Missing Values:\nFew Missing Values: The dataset contains very few missing values overall, with a slightly higher number of missing values for the subarticular stenosis condition. These can be handled easily by imputing the missing values or by excluding them from the model if they are not critical.\nConclusion:\nThe data shows that spinal degeneration is more severe in the lower spinal levels (L4/L5, L5/S1), which is a common pattern in medical diagnoses. These areas bear more mechanical stress, which can explain the higher occurrence of moderate and severe cases.\nNeural foraminal narrowing and subarticular stenosis at L4/L5 and L5/S1 seem to be the most problematic conditions in terms of severity.\nThe dominance of \"Normal/Mild\" cases suggests that most patients in this dataset have mild degenerative conditions, but the notable number of severe cases at the lower spine levels will be important for model training and predicting more severe outcomes.\nNext Steps:\nModel Focus: Since severity is higher at lower spinal levels (L4/L5, L5/S1), it might be useful to give special attention to these levels when training your model.\nBalancing the Dataset: Since the data is imbalanced (with many more \"Normal/Mild\" cases), you may want to use techniques like oversampling the \"Moderate\" and \"Severe\" cases or applying class weighting in the loss function during model training.\nHandling Missing Values: Since there are only a few missing values, you can either impute them or drop them without losing much data.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# File path to train_series_descriptions.csv\nseries_desc_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv'\n\n# Step 1: Load the CSV\nseries_desc_df = pd.read_csv(series_desc_path)\n\n# Step 2: Display the first few rows and summary of the data\nprint(\"First few rows of train_series_descriptions.csv:\")\nprint(series_desc_df.head())\n\n# Display summary information about the dataset (to understand its structure)\nprint(\"\\nSummary information:\")\nprint(series_desc_df.info())\n","metadata":{"execution":{"iopub.status.busy":"2024-09-08T23:53:05.278978Z","iopub.execute_input":"2024-09-08T23:53:05.280447Z","iopub.status.idle":"2024-09-08T23:53:05.329836Z","shell.execute_reply.started":"2024-09-08T23:53:05.280391Z","shell.execute_reply":"2024-09-08T23:53:05.328339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nreading the meta data\n\n\"\"\"\n\n# File path to train_series_descriptions.csv\nseries_desc_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv'\n\n# Step 1: Load the CSV\nseries_desc_df = pd.read_csv(series_desc_path)\n\n# Step 2: Display the first few rows of the data\nprint(\"First few rows of train_series_descriptions.csv:\")\nprint(series_desc_df.head())\n\n# Step 3: Display summary information about the dataset (to understand its structure)\nprint(\"\\nSummary information:\")\nprint(series_desc_df.info())\n\n# Step 4: Display all unique series descriptions\nunique_descriptions = series_desc_df['series_description'].unique()\nprint(\"\\nUnique series descriptions:\")\nfor desc in unique_descriptions:\n    print(desc)\n","metadata":{"execution":{"iopub.status.busy":"2024-09-08T23:58:36.720100Z","iopub.execute_input":"2024-09-08T23:58:36.720974Z","iopub.status.idle":"2024-09-08T23:58:36.758510Z","shell.execute_reply.started":"2024-09-08T23:58:36.720921Z","shell.execute_reply":"2024-09-08T23:58:36.757341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nsome images are in gray scale so that they need to be processed as a single channel \nwhile some images are RGB so they need to processed as 3 channels\n\n\n\"\"\"\n\n# File paths\ntrain_csv_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv'\nseries_desc_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv'\ntrain_images_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# Step 1: Load the CSV files\ntrain_df = pd.read_csv(train_csv_path)\nseries_desc_df = pd.read_csv(series_desc_path)\n\n# Ensure that study_id and series_id are integers\ntrain_df['study_id'] = train_df['study_id'].astype(int)\nseries_desc_df['study_id'] = series_desc_df['study_id'].astype(int)\nseries_desc_df['series_id'] = series_desc_df['series_id'].astype(int)\n\n# Step 2: Map the string labels in train.csv to integers\nseverity_map = {'Normal/Mild': 0, 'Moderate': 1, 'Severe': 2}\nfor col in train_df.columns[1:]:\n    train_df[col] = train_df[col].map(severity_map)\n\n# Step 3: Define the dataset class to handle series descriptions\nclass SpineDataset(Dataset):\n    def __init__(self, csv_data, series_data, image_dir, transforms=None, series_type=None):\n        self.csv_data = csv_data\n        self.series_data = series_data\n        self.image_dir = image_dir\n        self.transforms = transforms\n        self.series_type = series_type  # Filter by specific series_description if provided\n\n    def __len__(self):\n        return len(self.csv_data)\n\n    def __getitem__(self, idx):\n        # Get the study ID and labels from the CSV data\n        study_id = int(self.csv_data.iloc[idx]['study_id'])  # Ensure it's an integer\n        labels = self.csv_data.iloc[idx, 1:].values.astype(float)  # Convert labels to float for tensor conversion\n\n        # Filter series descriptions for the given study ID\n        study_series = self.series_data[self.series_data['study_id'] == study_id]\n        \n        # Filter by series type (Sagittal T2/STIR, Sagittal T1, Axial T2)\n        if self.series_type:\n            study_series = study_series[study_series['series_description'] == self.series_type]\n        \n        # If no series found, return dummy image and label\n        if study_series.empty:\n            return torch.zeros((1, 224, 224), dtype=torch.float32), torch.tensor(labels), study_id\n\n        # Select the first series and load the corresponding DICOM image\n        series_id = int(study_series.iloc[0]['series_id'])  # Ensure it's an integer\n        img_path = os.path.join(self.image_dir, str(study_id), str(series_id))\n\n        # Check if the directory exists and contains DICOM files\n        if not os.path.exists(img_path):\n            print(f\"Error: Path {img_path} does not exist for study {study_id}\")\n            return torch.zeros((1, 224, 224), dtype=torch.float32), torch.tensor(labels), study_id\n\n        dicom_files = [f for f in os.listdir(img_path) if f.endswith('.dcm')]\n\n        if dicom_files:\n            dicom_file = os.path.join(img_path, dicom_files[0])  # Load the first DICOM file\n            dicom = pydicom.dcmread(dicom_file)\n            img = dicom.pixel_array\n\n            # Convert to a PIL image for further processing\n            img = Image.fromarray(img)\n\n            # Convert the image to the correct mode (grayscale or RGB)\n            if img.mode == 'L':  # Grayscale images\n                img = img.convert('L')  # Explicitly set to single channel\n            else:\n                img = img.convert('RGB')  # Set multi-channel images to RGB\n\n            # Apply transformations (resize, normalize, etc.)\n            if self.transforms:\n                img = self.transforms(img)\n\n            return img, torch.tensor(labels), study_id\n        else:\n            print(f\"No DICOM files found in {img_path} for study {study_id}\")\n            # If no image is found, return a dummy tensor\n            img = torch.zeros((1, 224, 224), dtype=torch.float32)\n            return img, torch.tensor(labels), study_id\n\n# Step 4: Define the transformations\nimage_transforms = transforms.Compose([\n    transforms.Resize((224, 224)),  # Resize all images to 224x224\n    transforms.ToTensor(),  # Convert image to a PyTorch tensor\n    transforms.Normalize(mean=[0.5], std=[0.5])  # Normalize the grayscale image\n])\n\n# Step 5: Create a function to visualize samples for all series types\ndef visualize_all_series_types():\n    series_types = ['Axial T2', 'Sagittal T2/STIR', 'Sagittal T1']  # The three series types\n\n    for series_type in series_types:\n        print(f\"\\nVisualizing series type: {series_type}\\n\")\n        \n        # Create the dataset and dataloader for the current series type\n        dataset = SpineDataset(csv_data=train_df, series_data=series_desc_df, image_dir=train_images_path, transforms=image_transforms, series_type=series_type)\n        dataloader = DataLoader(dataset, batch_size=4, shuffle=True, num_workers=4, pin_memory=True)\n        \n        # Visualize a sample of images and their labels\n        visualize_sample_images(dataloader, series_type)\n\n# Step 6: Visualize a sample of images and their labels\ndef visualize_sample_images(dataloader, series_type):\n    try:\n        images, labels, study_ids = next(iter(dataloader))\n    except StopIteration:\n        print(f\"No samples found for series type: {series_type}\")\n        return\n\n    images = images.to(device)\n    labels = labels.to(device)\n\n    # Move images and labels to CPU for visualization\n    images = images.cpu().numpy()\n    labels = labels.cpu().numpy()\n\n    fig, ax = plt.subplots(1, 4, figsize=(20, 5))  # Displaying 4 images in a row\n    for i in range(4):\n        ax[i].imshow(images[i][0], cmap='gray')  # Display the first channel as grayscale\n        ax[i].set_title(f\"Study ID: {study_ids[i]}\\nLabel: {labels[i]}\")\n        ax[i].axis('off')\n    plt.show()\n\n# Step 7: Move data to GPU if available\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n\n# Step 8: Call the function to visualize samples from all series types\nvisualize_all_series_types()\n","metadata":{"execution":{"iopub.status.busy":"2024-09-09T00:07:44.142411Z","iopub.execute_input":"2024-09-09T00:07:44.143530Z","iopub.status.idle":"2024-09-09T00:07:48.225581Z","shell.execute_reply.started":"2024-09-09T00:07:44.143477Z","shell.execute_reply":"2024-09-09T00:07:48.223703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nvisualize x and y\n\n\"\"\"\n\nimport os\nimport pandas as pd\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# File paths\ntrain_label_coords_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv'\ntrain_images_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# Load the CSV file containing coordinates\nlabel_coords_df = pd.read_csv(train_label_coords_path)\n\n# Function to load DICOM image\ndef load_dicom_image(study_id, series_id, instance_number):\n    image_path = os.path.join(train_images_path, str(study_id), str(series_id), f\"{instance_number}.dcm\")\n    if os.path.exists(image_path):\n        dicom = pydicom.dcmread(image_path)\n        img = dicom.pixel_array  # Extract the pixel data\n        return img\n    else:\n        print(f\"Image {image_path} not found.\")\n        return None\n\n# Function to plot the image and overlay the key points\ndef plot_image_with_points(study_id, series_id, instance_number, points):\n    img = load_dicom_image(study_id, series_id, instance_number)\n    \n    if img is not None:\n        plt.imshow(img, cmap='gray')\n        plt.scatter(points['x'], points['y'], c='red', s=40, label=\"Key Points\")\n        plt.legend()\n        plt.title(f\"Study ID: {study_id}, Series ID: {series_id}, Instance: {instance_number}\")\n        plt.axis('off')\n        plt.show()\n    else:\n        print(f\"Could not load the image for Study ID: {study_id}, Series ID: {series_id}, Instance: {instance_number}\")\n\n# Sample visualization for the first few rows of the CSV\nfor index, row in label_coords_df.iterrows():\n    study_id = row['study_id']\n    series_id = row['series_id']\n    instance_number = row['instance_number']\n    points = {'x': row['x'], 'y': row['y']}\n    \n    # Plot the image with points\n    plot_image_with_points(study_id, series_id, instance_number, points)\n    \n    # Limit the number of visualizations for demonstration (optional)\n    if index >= 5:\n        break\n","metadata":{"execution":{"iopub.status.busy":"2024-09-09T00:11:11.122319Z","iopub.execute_input":"2024-09-09T00:11:11.122840Z","iopub.status.idle":"2024-09-09T00:11:12.976246Z","shell.execute_reply.started":"2024-09-09T00:11:11.122763Z","shell.execute_reply":"2024-09-09T00:11:12.975027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nfull data visualization\n\n\"\"\"\n# File paths\ntrain_label_coords_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv'\nseries_desc_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv'\ntrain_images_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# Load the CSV files containing coordinates and series descriptions\nlabel_coords_df = pd.read_csv(train_label_coords_path)\nseries_desc_df = pd.read_csv(series_desc_path)\n\n# Ensure the correct data types\nseries_desc_df['study_id'] = series_desc_df['study_id'].astype(int)\nseries_desc_df['series_id'] = series_desc_df['series_id'].astype(int)\n\n# Function to load DICOM image\ndef load_dicom_image(study_id, series_id, instance_number):\n    image_path = os.path.join(train_images_path, str(study_id), str(series_id), f\"{instance_number}.dcm\")\n    if os.path.exists(image_path):\n        dicom = pydicom.dcmread(image_path)\n        img = dicom.pixel_array  # Extract the pixel data\n        return img\n    else:\n        print(f\"Image {image_path} not found.\")\n        return None\n\n# Function to get the series description\ndef get_series_description(study_id, series_id):\n    desc_row = series_desc_df[(series_desc_df['study_id'] == study_id) & (series_desc_df['series_id'] == series_id)]\n    if not desc_row.empty:\n        return desc_row.iloc[0]['series_description']\n    else:\n        return 'Description Not Found'\n\n# Function to plot the image and overlay the key points with description\ndef plot_image_with_points(study_id, series_id, instance_number, points, description):\n    img = load_dicom_image(study_id, series_id, instance_number)\n    \n    if img is not None:\n        plt.imshow(img, cmap='gray')\n        plt.scatter(points['x'], points['y'], c='red', s=40, label=\"Key Points\")\n        plt.legend()\n        plt.title(f\"Study ID: {study_id}, Series ID: {series_id}, Instance: {instance_number}\\nDescription: {description}\")\n        plt.axis('off')\n        plt.show()\n    else:\n        print(f\"Could not load the image for Study ID: {study_id}, Series ID: {series_id}, Instance: {instance_number}\")\n\n# Sample visualization for the first few rows of the CSV\nfor index, row in label_coords_df.iterrows():\n    study_id = row['study_id']\n    series_id = row['series_id']\n    instance_number = row['instance_number']\n    points = {'x': row['x'], 'y': row['y']}\n    \n    # Fetch the series description\n    description = get_series_description(study_id, series_id)\n    \n    # Plot the image with points and description\n    plot_image_with_points(study_id, series_id, instance_number, points, description)\n    \n    \n    if index >= 5:\n        break\n","metadata":{"execution":{"iopub.status.busy":"2024-09-09T00:12:36.499730Z","iopub.execute_input":"2024-09-09T00:12:36.500195Z","iopub.status.idle":"2024-09-09T00:12:38.557861Z","shell.execute_reply.started":"2024-09-09T00:12:36.500154Z","shell.execute_reply":"2024-09-09T00:12:38.556766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}