{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"dockerImageVersionId":30746,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**1. Understand the Problem**\n\n* Familiarize yourself with degenerative spine conditions and how they are diagnosed using MRI.\n\n*  Review the competition details, including the evaluation metrics and submission format.","metadata":{}},{"cell_type":"markdown","source":"**1.Project Setup and Data Loading**\n\n** Objective**: To classify lumbar spine degenerative conditions using medical images.\n\nData Source: You have a dataset containing medical images stored in a folder structure and corresponding labels in a CSV file.\n\n**Key Files:**\ntrain_images: Folder containing DICOM images of lumbar spine scans.\n\ntrain_label_coordinates.csv: CSV file containing labels for each image, including study IDs, series IDs, and specific conditions.\n\n2. Data Preprocessing\nReading Labels: You loaded the labels from the CSV file into a Pandas DataFrame using pd.read_csv(). This CSV contained columns like study_id, series_id, condition, etc.\nIterating Over Images: You wrote a loop to iterate through the folder structure (organized by study ID and series ID) to load images using the pydicom library.\nLabel Extraction: For each image, you matched the corresponding label from the DataFrame based on study_id and series_id. If no label was found, you assigned a default value.\n\n3. Image Processing\n\nData Augmentation: You used the ImageDataGenerator from Keras to apply data augmentation techniques like rotation, shifting, shearing, and zooming to artificially expand the dataset and improve model generalization.\nImage Preparation: After loading, you converted images into NumPy arrays for easier manipulation and feeding into the neural network.\n\n4. Model Architecture\nConvolutional Neural Network (CNN):\nLayers:\nConv2D Layers: Extract spatial features from the images using convolution operations.\nMaxPooling2D Layers: Downsample the image to reduce computational complexity.\nFlatten Layer: Converts the 2D matrices to a 1D vector to feed into fully connected layers.\nDense Layers: Fully connected layers to perform classification.\nDropout Layer: Regularization technique to prevent overfitting by randomly setting a fraction of input units to 0.\nOutput Layer: A softmax layer to predict one of the three classes (Normal/Mild, Moderate, Severe).\nCompilation: You compiled the model using the Adam optimizer and sparse_categorical_crossentropy loss function, which is suitable for multi-class classification.\n\n5. Model Training\nTraining the Model:\nYou trained the model on the augmented image data using the fit() method.\nValidation: During training, you monitored the model's performance on a validation set to prevent overfitting and adjust hyperparameters if necessary.\n\n6. Model Evaluation\n\nOnce training is complete, you will evaluate the model using the validation set to assess its performance. This typically involves checking metrics like accuracy and loss, as well as visualizing the results with plots.\n\n7. Summary\n\nGoal: The primary goal is to classify images of lumbar spine conditions into different severity levels using a deep learning model.\n\nProcess:\nData Loading: Load and preprocess image data and corresponding labels.\nImage Processing: Apply techniques like data augmentation to enhance model performance.\nModel Design: Design a CNN model tailored for image classification tasks.\nTraining: Train the model on augmented data and monitor its performance on a validation set.\nEvaluation (Next Step): Assess the model's accuracy and loss, and fine-tune if necessary.\nThis workflow provides a clear and concise overview of your work, from the initial setup to the upcoming steps in model evaluation.","metadata":{}},{"cell_type":"markdown","source":"**1. Import Libraries**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport pydicom\nimport cv2\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MultiLabelBinarizer\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T09:44:39.787392Z","iopub.execute_input":"2024-08-13T09:44:39.788353Z","iopub.status.idle":"2024-08-13T09:44:39.938695Z","shell.execute_reply.started":"2024-08-13T09:44:39.788315Z","shell.execute_reply":"2024-08-13T09:44:39.93773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nbase_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification'\ntrain_desc_path = os.path.join(base_path,'train_series_descriptions.csv')\ntrain_label_path = os.path.join(base_path,'train_label_coordinates.csv')\ntrain_csv_path = os.path.join(base_path,'train.csv')\ntrain_folder = os.path.join(base_path,'train_images')\ntest_folder = os.path.join(base_path,'test_images')","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:09:13.540142Z","iopub.execute_input":"2024-08-13T11:09:13.540575Z","iopub.status.idle":"2024-08-13T11:09:13.54733Z","shell.execute_reply.started":"2024-08-13T11:09:13.540542Z","shell.execute_reply":"2024-08-13T11:09:13.546062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**2. Load Data**","metadata":{}},{"cell_type":"code","source":"# Define the paths to the directories and files\ntrain_images_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\ntrain_csv_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv'\ntrain_label_coordinates_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv'\ntrain_series_descriptions_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv'\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T10:45:06.987264Z","iopub.execute_input":"2024-08-13T10:45:06.988358Z","iopub.status.idle":"2024-08-13T10:45:06.993481Z","shell.execute_reply.started":"2024-08-13T10:45:06.988322Z","shell.execute_reply":"2024-08-13T10:45:06.992237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\ntrain_desc = pd.read_csv(train_desc_path)\ntrain_label = pd.read_csv(train_label_path)\ntrain = pd.read_csv(train_csv_path)","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:32:35.688277Z","iopub.execute_input":"2024-08-13T11:32:35.688691Z","iopub.status.idle":"2024-08-13T11:32:35.814966Z","shell.execute_reply.started":"2024-08-13T11:32:35.688659Z","shell.execute_reply":"2024-08-13T11:32:35.813827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_desc.head()","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:33:12.916473Z","iopub.execute_input":"2024-08-13T11:33:12.916916Z","iopub.status.idle":"2024-08-13T11:33:12.934141Z","shell.execute_reply.started":"2024-08-13T11:33:12.916885Z","shell.execute_reply":"2024-08-13T11:33:12.932942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_desc.tail()","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:33:37.133063Z","iopub.execute_input":"2024-08-13T11:33:37.133485Z","iopub.status.idle":"2024-08-13T11:33:37.145577Z","shell.execute_reply.started":"2024-08-13T11:33:37.133452Z","shell.execute_reply":"2024-08-13T11:33:37.144218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_label","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:33:59.638505Z","iopub.execute_input":"2024-08-13T11:33:59.638949Z","iopub.status.idle":"2024-08-13T11:33:59.65925Z","shell.execute_reply.started":"2024-08-13T11:33:59.638917Z","shell.execute_reply":"2024-08-13T11:33:59.65805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:34:26.99797Z","iopub.execute_input":"2024-08-13T11:34:26.998406Z","iopub.status.idle":"2024-08-13T11:34:27.024778Z","shell.execute_reply.started":"2024-08-13T11:34:26.998371Z","shell.execute_reply":"2024-08-13T11:34:27.023538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.merge(train_label, train_desc, on = ['study_id', 'series_id'], how = 'inner')","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:36:14.982645Z","iopub.execute_input":"2024-08-13T11:36:14.983101Z","iopub.status.idle":"2024-08-13T11:36:15.011117Z","shell.execute_reply.started":"2024-08-13T11:36:14.983068Z","shell.execute_reply":"2024-08-13T11:36:15.009641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:36:25.224875Z","iopub.execute_input":"2024-08-13T11:36:25.225288Z","iopub.status.idle":"2024-08-13T11:36:25.243696Z","shell.execute_reply.started":"2024-08-13T11:36:25.225257Z","shell.execute_reply":"2024-08-13T11:36:25.242199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load training labels\ntrain_labels = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv')\ntrain_label_coords = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')\n\n# Display the first few rows\nprint(train_labels.head())\nprint(train_label_coords.head())","metadata":{"execution":{"iopub.status.busy":"2024-08-13T08:27:31.705892Z","iopub.execute_input":"2024-08-13T08:27:31.70634Z","iopub.status.idle":"2024-08-13T08:27:31.896741Z","shell.execute_reply.started":"2024-08-13T08:27:31.70631Z","shell.execute_reply":"2024-08-13T08:27:31.895596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the structure of the training dataset\nprint(train_df.info())\nprint(train_df.describe())\n\n# Check for missing values\nprint(train_df.isnull().sum())\n\n# Do the same for label coordinates\nprint(label_coords_df.info())\nprint(label_coords_df.describe())\nprint(label_coords_df.isnull().sum())\n\n# And for series descriptions\nprint(train_series_desc_df.info())\nprint(train_series_desc_df.describe())\nprint(train_series_desc_df.isnull().sum())\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T09:43:44.572443Z","iopub.execute_input":"2024-08-13T09:43:44.572873Z","iopub.status.idle":"2024-08-13T09:43:44.682498Z","shell.execute_reply.started":"2024-08-13T09:43:44.572839Z","shell.execute_reply":"2024-08-13T09:43:44.681254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"****3. Image Preprocessing****","metadata":{}},{"cell_type":"code","source":"# List the directories inside the train_images folder\ntrain_images_dir = os.path.join(data_dir, 'train_images')\nstudies = os.listdir(train_images_dir)\n\n# Print the first few study IDs\nprint(\"Available study IDs:\", studies[:5])\n\n# Now, list the series within a study\nsample_study_dir = os.path.join(train_images_dir, studies[0])\nseries = os.listdir(sample_study_dir)\n\nprint(\"Available series in the first study:\", series)\n\n# List the DICOM files in the first series\nsample_series_dir = os.path.join(sample_study_dir, series[0])\ndicom_files = os.listdir(sample_series_dir)\n\nprint(\"Available DICOM files in the first series:\", dicom_files[:5])\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T09:44:06.399463Z","iopub.execute_input":"2024-08-13T09:44:06.399931Z","iopub.status.idle":"2024-08-13T09:44:06.530136Z","shell.execute_reply.started":"2024-08-13T09:44:06.399897Z","shell.execute_reply":"2024-08-13T09:44:06.52898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the path to a sample DICOM image\nsample_image_path = os.path.join(sample_series_dir, dicom_files[0])\n\n# Load the DICOM file\ndicom_image = pydicom.dcmread(sample_image_path)\n\n# Display the image using matplotlib\nplt.imshow(dicom_image.pixel_array, cmap=plt.cm.bone)\nplt.axis('off')  # Hide the axis\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T09:44:46.386355Z","iopub.execute_input":"2024-08-13T09:44:46.386787Z","iopub.status.idle":"2024-08-13T09:44:46.626118Z","shell.execute_reply.started":"2024-08-13T09:44:46.386724Z","shell.execute_reply":"2024-08-13T09:44:46.624966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# Load training labels\ntrain_labels = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv')\n\n# Print the first few rows and columns for debugging\nprint(train_labels.columns)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T10:11:17.721548Z","iopub.execute_input":"2024-08-13T10:11:17.722719Z","iopub.status.idle":"2024-08-13T10:11:17.75402Z","shell.execute_reply.started":"2024-08-13T10:11:17.722663Z","shell.execute_reply":"2024-08-13T10:11:17.752702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nimport matplotlib.pyplot as plt\n\n# Path to the specific DICOM file\nfile_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/1012375618/352098527/7.dcm'\n\n# Load the DICOM file\ntry:\n    dicom_img = pydicom.dcmread(file_path)\n    \n    # Extract pixel data\n    img = dicom_img.pixel_array\n    \n    # Display the image\n    plt.imshow(img, cmap='gray')  # Use 'gray' colormap for DICOM images\n    plt.title('DICOM Image')\n    plt.axis('off')  # Hide axis\n    plt.show()\n\nexcept Exception as e:\n    print(f\"Error loading DICOM file: {e}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T10:16:37.474519Z","iopub.execute_input":"2024-08-13T10:16:37.474992Z","iopub.status.idle":"2024-08-13T10:16:37.670347Z","shell.execute_reply.started":"2024-08-13T10:16:37.474958Z","shell.execute_reply":"2024-08-13T10:16:37.66926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# Path to the folder containing DICOM files\nfolder_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images/1020394063/1523561649'\n\n# Get all DICOM files in the folder\ndicom_files = [f for f in os.listdir(folder_path) if f.endswith('.dcm')]\n\n# Load and display each DICOM file\nfor dicom_file in dicom_files:\n    file_path = os.path.join(folder_path, dicom_file)\n    \n    try:\n        # Load the DICOM file\n        dicom_img = pydicom.dcmread(file_path)\n        \n        # Extract pixel data\n        img = dicom_img.pixel_array\n        \n        # Display the image\n        plt.imshow(img, cmap='gray')  # Use 'gray' colormap for DICOM images\n        plt.title(dicom_file)  # Display the filename as the title\n        plt.axis('off')  # Hide axis\n        plt.show()\n        \n    except Exception as e:\n        print(f\"Error loading DICOM file {dicom_file}: {e}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T10:17:51.070392Z","iopub.execute_input":"2024-08-13T10:17:51.070813Z","iopub.status.idle":"2024-08-13T10:17:55.444053Z","shell.execute_reply.started":"2024-08-13T10:17:51.070784Z","shell.execute_reply":"2024-08-13T10:17:55.443004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_folder","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:37:12.292467Z","iopub.execute_input":"2024-08-13T11:37:12.292911Z","iopub.status.idle":"2024-08-13T11:37:12.300035Z","shell.execute_reply.started":"2024-08-13T11:37:12.29288Z","shell.execute_reply":"2024-08-13T11:37:12.298793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['image_path'] = train_folder + '/' + train_data.study_id.astype(str) + '/' + train_data.series_id.astype(str) + '/' + train_data.instance_number.astype(str) + '.dcm'","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:43:56.73527Z","iopub.execute_input":"2024-08-13T11:43:56.735574Z","iopub.status.idle":"2024-08-13T11:43:56.875001Z","shell.execute_reply.started":"2024-08-13T11:43:56.735547Z","shell.execute_reply":"2024-08-13T11:43:56.873877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:52:33.359665Z","iopub.execute_input":"2024-08-13T11:52:33.360187Z","iopub.status.idle":"2024-08-13T11:52:33.377283Z","shell.execute_reply.started":"2024-08-13T11:52:33.360152Z","shell.execute_reply":"2024-08-13T11:52:33.376047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.columns","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:59:27.45359Z","iopub.execute_input":"2024-08-13T11:59:27.454983Z","iopub.status.idle":"2024-08-13T11:59:27.462779Z","shell.execute_reply.started":"2024-08-13T11:59:27.454931Z","shell.execute_reply":"2024-08-13T11:59:27.46149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:59:46.110408Z","iopub.execute_input":"2024-08-13T11:59:46.110856Z","iopub.status.idle":"2024-08-13T11:59:46.14909Z","shell.execute_reply.started":"2024-08-13T11:59:46.110826Z","shell.execute_reply":"2024-08-13T11:59:46.147848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# Assuming you have already set the train_folder variable correctly\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\ndef display_image(study_id, series_id, instance_number):\n    # Construct the file path\n    file_path = os.path.join(train_folder, study_id, series_id, f'{instance_number + 1}.dcm')  # +1 to start from 1\n    # Check if the file exists\n    if os.path.exists(file_path):\n        dicom_image = pydicom.dcmread(file_path).pixel_array\n        plt.imshow(dicom_image, cmap='gray')\n        plt.axis('off')\n        plt.show()\n    else:\n        print(f'File not found: {file_path}')\n\n# Update this call with actual values from your dataset\ndisplay_image('4003253', '702807833', 7)  # Change 7 to the actual instance number you want to display\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:57:31.398028Z","iopub.execute_input":"2024-08-13T11:57:31.398537Z","iopub.status.idle":"2024-08-13T11:57:32.477193Z","shell.execute_reply.started":"2024-08-13T11:57:31.398503Z","shell.execute_reply":"2024-08-13T11:57:32.476091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study_id = '4003253'\nseries_id = '702807833'\nseries_path = os.path.join(train_folder, study_id, series_id)\nprint(os.listdir(series_path))  # This will list all the DICOM files in the specified series\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:57:46.815561Z","iopub.execute_input":"2024-08-13T11:57:46.81693Z","iopub.status.idle":"2024-08-13T11:57:46.830287Z","shell.execute_reply.started":"2024-08-13T11:57:46.816886Z","shell.execute_reply":"2024-08-13T11:57:46.828614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# Assuming you have already set the train_folder variable correctly\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\ndef display_images(study_id, series_id, instance_numbers):\n    plt.figure(figsize=(15, 10))  # Set the figure size for better visualization\n    for i, instance_number in enumerate(instance_numbers):\n        # Construct the file path (add 1 if the instance_number starts from 0)\n        file_path = os.path.join(train_folder, study_id, series_id, f'{instance_number + 1}.dcm')\n        # Check if the file exists\n        if os.path.exists(file_path):\n            dicom_image = pydicom.dcmread(file_path).pixel_array\n            plt.subplot(1, len(instance_numbers), i + 1)  # Create subplots\n            plt.imshow(dicom_image, cmap='gray')\n            plt.axis('off')\n            plt.title(f'Instance: {instance_number + 1}')  # Title for each image\n        else:\n            print(f'File not found: {file_path}')\n    plt.tight_layout()  # Adjust subplots to fit into the figure area.\n    plt.show()\n\n# Update this call with actual values from your dataset\nstudy_id = '4003253'\nseries_id = '702807833'\ninstance_numbers = [7, 8, 9, 10, 11]  # Replace with the actual instance numbers you want to display\ndisplay_images(study_id, series_id, instance_numbers)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T11:58:46.200501Z","iopub.execute_input":"2024-08-13T11:58:46.201017Z","iopub.status.idle":"2024-08-13T11:58:47.270289Z","shell.execute_reply.started":"2024-08-13T11:58:46.20098Z","shell.execute_reply":"2024-08-13T11:58:47.268936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\n\n# Path to the train_images folder\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# List all study IDs (folders) in the train folder\nstudy_ids = os.listdir(train_folder)\nprint(\"Available Study IDs:\", study_ids)\n\n# Pick a study ID and list all series IDs (folders)\nseries_ids = os.listdir(os.path.join(train_folder, study_ids[0]))\nprint(f\"Available Series IDs for {study_ids[0]}:\", series_ids)\n\n# Pick a series ID and list all instance files\ninstance_files = os.listdir(os.path.join(train_folder, study_ids[0], series_ids[0]))\nprint(f\"Available Instance Files for {study_ids[0]}/{series_ids[0]}:\", instance_files)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:19:30.596987Z","iopub.execute_input":"2024-08-13T12:19:30.59741Z","iopub.status.idle":"2024-08-13T12:19:30.609414Z","shell.execute_reply.started":"2024-08-13T12:19:30.597376Z","shell.execute_reply":"2024-08-13T12:19:30.608363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**. Visualizing a DICOM Image**","metadata":{}},{"cell_type":"code","source":"import pydicom\nimport matplotlib.pyplot as plt\n\ndef display_image(study_id, series_id, instance_number):\n    file_path = os.path.join(train_folder, study_id, series_id, f'{instance_number}.dcm')\n    dicom_image = pydicom.dcmread(file_path).pixel_array\n    plt.imshow(dicom_image, cmap='gray')\n    plt.axis('off')\n    plt.show()\n\n# Example: Replace with actual study_id, series_id, and instance_number\ndisplay_image(study_ids[0], series_ids[0], instance_files[0].split('.')[0])\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:20:10.859012Z","iopub.execute_input":"2024-08-13T12:20:10.859503Z","iopub.status.idle":"2024-08-13T12:20:11.048651Z","shell.execute_reply.started":"2024-08-13T12:20:10.859467Z","shell.execute_reply":"2024-08-13T12:20:11.047014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Automate Random Selection for Visualization**","metadata":{}},{"cell_type":"markdown","source":"To visualize a random image from the dataset:","metadata":{}},{"cell_type":"code","source":"import random\n\n# Randomly select a study ID, series ID, and instance file\nrandom_study = random.choice(study_ids)\nrandom_series = random.choice(os.listdir(os.path.join(train_folder, random_study)))\nrandom_instance = random.choice(os.listdir(os.path.join(train_folder, random_study, random_series)))\n\n# Display the randomly selected image\ndisplay_image(random_study, random_series, random_instance.split('.')[0])\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:21:00.768512Z","iopub.execute_input":"2024-08-13T12:21:00.768954Z","iopub.status.idle":"2024-08-13T12:21:00.94127Z","shell.execute_reply.started":"2024-08-13T12:21:00.768923Z","shell.execute_reply":"2024-08-13T12:21:00.939842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# Path to the train_images folder\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# Function to display multiple DICOM images\ndef display_images(image_arrays, titles, ncols=4):\n    nrows = len(image_arrays) // ncols + (1 if len(image_arrays) % ncols else 0)\n    fig, axes = plt.subplots(nrows, ncols, figsize=(15, 5 * nrows))\n    \n    for i, (image, title) in enumerate(zip(image_arrays, titles)):\n        ax = axes.flat[i]\n        ax.imshow(image, cmap='gray')\n        ax.set_title(title)\n        ax.axis('off')\n    \n    # Hide any remaining subplots if there are empty spots\n    for i in range(len(image_arrays), nrows * ncols):\n        axes.flat[i].axis('off')\n\n    plt.tight_layout()\n    plt.show()\n\n# Select a few study IDs to visualize\nselected_studies = study_ids[:3]  # Adjust the number based on your need\n\nimage_arrays = []\ntitles = []\n\n# Iterate through selected studies and series to collect images\nfor study_id in selected_studies:\n    study_path = os.path.join(train_folder, study_id)\n    series_ids = os.listdir(study_path)\n    \n    for series_id in series_ids[:2]:  # Adjust the number based on your need\n        series_path = os.path.join(study_path, series_id)\n        instance_files = os.listdir(series_path)\n        \n        for instance_file in instance_files[:3]:  # Adjust the number based on your need\n            dicom_path = os.path.join(series_path, instance_file)\n            dicom_data = pydicom.dcmread(dicom_path)\n            image_arrays.append(dicom_data.pixel_array)\n            titles.append(f'{study_id}/{series_id}/{instance_file}')\n        \nprint(f\"Selected {len(image_arrays)} images for visualization.\")\n\n# Display the selected images\ndisplay_images(image_arrays, titles)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:31:57.251806Z","iopub.execute_input":"2024-08-13T12:31:57.252459Z","iopub.status.idle":"2024-08-13T12:32:01.223248Z","shell.execute_reply.started":"2024-08-13T12:31:57.25242Z","shell.execute_reply":"2024-08-13T12:32:01.221725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Key Regions Highlighted: Enhancing Spinal Degeneration Visualization**","metadata":{}},{"cell_type":"code","source":"pip install opencv-python\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:38:18.828066Z","iopub.execute_input":"2024-08-13T12:38:18.828602Z","iopub.status.idle":"2024-08-13T12:38:34.933983Z","shell.execute_reply.started":"2024-08-13T12:38:18.828565Z","shell.execute_reply":"2024-08-13T12:38:34.932491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Highlight Areas in the Image**","metadata":{}},{"cell_type":"code","source":"import os\nimport pydicom\nimport cv2\nimport matplotlib.pyplot as plt\n\n# Path to the train_images folder\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# Function to display images with highlighted regions\ndef highlight_region(image, points, color=(255, 0, 0), thickness=2):\n    \"\"\"\n    Draw rectangles around the points of interest in the image.\n    \n    Args:\n        image: The original image array.\n        points: List of top-left and bottom-right points for rectangles.\n        color: Color of the rectangle in BGR format.\n        thickness: Thickness of the rectangle border.\n        \n    Returns:\n        Image with rectangles drawn.\n    \"\"\"\n    for (x, y, w, h) in points:\n        cv2.rectangle(image, (x, y), (x + w, y + h), color, thickness)\n    return image\n\n# Select a study and series\nselected_study = study_ids[0]\nselected_series = os.listdir(os.path.join(train_folder, selected_study))[0]\n\n# Path to DICOM files\ndicom_files = os.listdir(os.path.join(train_folder, selected_study, selected_series))\n\n# Load a DICOM file and convert it to an OpenCV image\ndicom_path = os.path.join(train_folder, selected_study, selected_series, dicom_files[0])\ndicom_data = pydicom.dcmread(dicom_path)\nimage = dicom_data.pixel_array\n\n# Convert the image to 8-bit format (if necessary) for OpenCV\nimage_8bit = cv2.convertScaleAbs(image, alpha=(255.0/65535.0))\n\n# Define the regions of interest (ROI) manually or through a model (x, y, width, height)\n# Example: Highlighting a small portion at the center of the image\nroi = [(150, 200, 50, 50), (300, 400, 100, 100)]  # Example points\n\n# Highlight the region\nhighlighted_image = highlight_region(image_8bit, roi, color=(0, 255, 0), thickness=3)\n\n# Display the image with highlighted regions\nplt.figure(figsize=(10, 10))\nplt.imshow(highlighted_image, cmap='gray')\nplt.axis('off')\nplt.title('Highlighted Spine Regions')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:39:19.119108Z","iopub.execute_input":"2024-08-13T12:39:19.119551Z","iopub.status.idle":"2024-08-13T12:39:19.447902Z","shell.execute_reply.started":"2024-08-13T12:39:19.119517Z","shell.execute_reply":"2024-08-13T12:39:19.446739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# Path to the train_images folder\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# Step 1: List available Study IDs\nstudy_ids = os.listdir(train_folder)\nprint(\"Available Study IDs:\", study_ids)\n\n# Pick the first study ID for demonstration (you can change this as needed)\nstudy_id = study_ids[0]\n\n# Step 2: List available Series IDs for the chosen Study ID\nseries_ids = os.listdir(os.path.join(train_folder, study_id))\nprint(f\"Available Series IDs for {study_id}:\", series_ids)\n\n# Pick the first series ID for demonstration\nseries_id = series_ids[0]\n\n# Step 3: List available Instance Files for the chosen Series ID\ninstance_files = os.listdir(os.path.join(train_folder, study_id, series_id))\nprint(f\"Available Instance Files for {study_id}/{series_id}:\", instance_files)\n\n# Pick the first instance file for demonstration\ninstance_number = instance_files[0]  # e.g., '1.dcm'\n\n# Function to create a heatmap overlay on the DICOM image\ndef create_heatmap(image, critical_points, intensity=255):\n    heatmap = np.zeros_like(image, dtype=np.float32)\n    \n    for point in critical_points:\n        # Increase heatmap intensity at the critical points\n        cv2.circle(heatmap, point, radius=10, color=intensity, thickness=-1)\n    \n    # Normalize the heatmap to be between 0 and 1\n    heatmap = cv2.normalize(heatmap, None, alpha=0, beta=1, norm_type=cv2.NORM_MINMAX)\n    \n    # Convert heatmap to color (e.g., using a colormap)\n    heatmap_color = cv2.applyColorMap(np.uint8(heatmap * 255), cv2.COLORMAP_JET)\n    \n    # Ensure the original image is 3 channels\n    if len(image.shape) == 2:  # If the image is grayscale\n        image = cv2.cvtColor(image, cv2.COLOR_GRAY2BGR)\n\n    # Overlay the heatmap on the original image\n    overlay = cv2.addWeighted(image, 0.5, heatmap_color, 0.5, 0)\n    \n    return overlay\n\n# Load the DICOM image\nfile_path = os.path.join(train_folder, study_id, series_id, instance_number)\ndicom_data = pydicom.dcmread(file_path)\ndicom_image = dicom_data.pixel_array\n\n# If the image is 2D, make sure it's in the correct format (e.g., uint8)\nif dicom_image.dtype != np.uint8:\n    dicom_image = (dicom_image / dicom_image.max() * 255).astype(np.uint8)\n\n# Define critical points (for demonstration, let's assume some random points)\n# In a real scenario, these would come from an analysis or model output\ncritical_points = [(100, 100), (150, 120), (200, 200)]  # Replace with actual points\n\n# Create the heatmap overlay\nheatmap_overlay = create_heatmap(dicom_image, critical_points)\n\n# Display the result\nplt.figure(figsize=(10, 10))\nplt.subplot(1, 2, 1)\nplt.title('Original DICOM Image')\nplt.imshow(dicom_image, cmap='gray')\nplt.axis('off')\n\nplt.subplot(1, 2, 2)\nplt.title('Heatmap Overlay')\nplt.imshow(heatmap_overlay)\nplt.axis('off')\n\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:43:43.582927Z","iopub.execute_input":"2024-08-13T12:43:43.584544Z","iopub.status.idle":"2024-08-13T12:43:44.06704Z","shell.execute_reply.started":"2024-08-13T12:43:43.584491Z","shell.execute_reply":"2024-08-13T12:43:44.065843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Steps for 3D Reconstruction of Spine**\n\n**Load the DICOM Series:**\nLoad the series of DICOM images that represent cross-sections of the spine.\n\n* Stack the Images:\nStack the 2D images into a 3D NumPy array, where each slice corresponds to a DICOM image.\n\n* Create a Volume Rendering:\nUse a volume rendering technique to visualize the 3D structure.\n\n* Surface Reconstruction (Optional):\nIf you need to create a mesh from the volume, you can use techniques like Marching Cubes or similar algorithms available in libraries like skimage.","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pydicom\nimport matplotlib.pyplot as plt\nfrom mpl_toolkits.mplot3d import Axes3D\n\n# Path to the train_images folder\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# Step 1: List available Study IDs\nstudy_ids = os.listdir(train_folder)\nstudy_id = study_ids[0]  # Choose the first study ID for demonstration\n\n# Step 2: List available Series IDs\nseries_ids = os.listdir(os.path.join(train_folder, study_id))\nseries_id = series_ids[0]  # Choose the first series ID for demonstration\n\n# Step 3: List available Instance Files\ninstance_files = sorted(os.listdir(os.path.join(train_folder, study_id, series_id)))\ndicom_images = []\n\n# Step 4: Load DICOM images into a 3D array\nfor instance in instance_files:\n    file_path = os.path.join(train_folder, study_id, series_id, instance)\n    dicom_data = pydicom.dcmread(file_path)\n    dicom_image = dicom_data.pixel_array\n    \n    # Normalize the image to 0-255 for better visualization\n    dicom_image = (dicom_image / np.max(dicom_image) * 255).astype(np.uint8)\n    \n    dicom_images.append(dicom_image)\n\n# Stack images to create a 3D volume\nvolume = np.stack(dicom_images, axis=-1)\n\n# Step 5: Visualize a few slices of the 3D volume\nfig = plt.figure(figsize=(10, 10))\nnum_slices = len(dicom_images)\nfor i in range(0, num_slices, num_slices // 10):  # Show 10 slices\n    ax = fig.add_subplot(5, 2, i // (num_slices // 10) + 1)\n    ax.imshow(volume[:, :, i], cmap='gray')\n    ax.axis('off')\n    ax.set_title(f'Slice {i + 1}')\nplt.tight_layout()\nplt.show()\n\n# Step 6: 3D Visualization (optional)\n# Create a grid of points for the 3D volume\nx, y, z = np.indices(volume.shape)\n\n# Create a mask for non-zero voxels\nmask = volume > 0\n\n# Set up the 3D plot\nfig = plt.figure()\nax = fig.add_subplot(111, projection='3d')\nax.scatter(x[mask], y[mask], z[mask], c='r', s=1)  # Use red points for the volume\n\nax.set_xlabel('X axis')\nax.set_ylabel('Y axis')\nax.set_zlabel('Z axis')\nplt.title('3D Reconstruction of Spine')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:46:50.58086Z","iopub.execute_input":"2024-08-13T12:46:50.581267Z","iopub.status.idle":"2024-08-13T12:48:08.525958Z","shell.execute_reply.started":"2024-08-13T12:46:50.581237Z","shell.execute_reply":"2024-08-13T12:48:08.524952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary Report for Lumbar Spine Degenerative Classification Project**","metadata":{}},{"cell_type":"markdown","source":"The goal of this project is to classify lumbar spine images based on degeneration levels (Normal/Mild, Moderate, Severe) using a Convolutional Neural Network (CNN). The dataset consists of DICOM images of lumbar spines, and we aim to help medical professionals identify and diagnose spinal issues efficiently.","metadata":{}},{"cell_type":"markdown","source":"**Data Exploration**","metadata":{}},{"cell_type":"code","source":"import os\n\n# Path to the train_images folder\ntrain_folder = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\n\n# List all study IDs (folders) in the train folder\nstudy_ids = os.listdir(train_folder)\nprint(\"Available Study IDs:\", study_ids)\n\n# Pick a study ID and list all series IDs (folders)\nseries_ids = os.listdir(os.path.join(train_folder, study_ids[0]))\nprint(f\"Available Series IDs for {study_ids[0]}:\", series_ids)\n\n# Pick a series ID and list all instance files\ninstance_files = os.listdir(os.path.join(train_folder, study_ids[0], series_ids[0]))\nprint(f\"Available Instance Files for {study_ids[0]}/{series_ids[0]}:\", instance_files)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:55:28.959614Z","iopub.execute_input":"2024-08-13T12:55:28.960674Z","iopub.status.idle":"2024-08-13T12:55:28.97373Z","shell.execute_reply.started":"2024-08-13T12:55:28.960639Z","shell.execute_reply":"2024-08-13T12:55:28.97238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Model Development**","metadata":{}},{"cell_type":"markdown","source":"A CNN architecture was defined for image classification, using data augmentation to enhance model performance.","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Define the model architecture\nmodel = Sequential([\n    Conv2D(32, (3, 3), activation='relu', input_shape=(height, width, channels)),\n    MaxPooling2D(pool_size=(2, 2)),\n    Flatten(),\n    Dense(128, activation='relu'),\n    Dropout(0.5),\n    Dense(3, activation='softmax')  # Assuming three output classes: Normal/Mild, Moderate, Severe\n])\n\n# Compile the model\nmodel.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])\n\n# Data augmentation\ndatagen = ImageDataGenerator(\n    rotation_range=20,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    shear_range=0.2,\n    zoom_range=0.2,\n    horizontal_flip=True,\n    fill_mode='nearest'\n)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-13T12:55:51.567222Z","iopub.execute_input":"2024-08-13T12:55:51.567632Z","iopub.status.idle":"2024-08-13T12:55:51.94554Z","shell.execute_reply.started":"2024-08-13T12:55:51.567598Z","shell.execute_reply":"2024-08-13T12:55:51.944231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()\n","metadata":{"execution":{"iopub.status.busy":"2024-08-14T06:49:32.741396Z","iopub.execute_input":"2024-08-14T06:49:32.741861Z","iopub.status.idle":"2024-08-14T06:49:32.778171Z","shell.execute_reply.started":"2024-08-14T06:49:32.741828Z","shell.execute_reply":"2024-08-14T06:49:32.776962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout, Input\nfrom tensorflow.keras.optimizers import Adam\n\n# Define input image dimensions and number of classes\nimg_height = 150  # Set to the height of your images\nimg_width = 150   # Set to the width of your images\nnum_classes = 10  # Set to the number of classes in your dataset\n\n# Initialize the model\nmodel = Sequential()\n\n# Input Layer\nmodel.add(Input(shape=(img_height, img_width, 3)))\n\n# Convolutional Layer 1\nmodel.add(Conv2D(32, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\n\n# Convolutional Layer 2\nmodel.add(Conv2D(64, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\n\n# Convolutional Layer 3\nmodel.add(Conv2D(128, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\n\n# Flatten the features\nmodel.add(Flatten())\n\n# Fully Connected Layer 1\nmodel.add(Dense(512, activation='relu'))\nmodel.add(Dropout(0.5))\n\n# Output Layer\nmodel.add(Dense(num_classes, activation='softmax'))\n\n# Compile the model\nmodel.compile(optimizer=Adam(),\n              loss='sparse_categorical_crossentropy',\n              metrics=['accuracy'])\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T06:49:09.394303Z","iopub.execute_input":"2024-08-22T06:49:09.3951Z","iopub.status.idle":"2024-08-22T06:49:09.493726Z","shell.execute_reply.started":"2024-08-22T06:49:09.395066Z","shell.execute_reply":"2024-08-22T06:49:09.493012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next steps involve training and evaluating your CNN model. Here's how you can proceed:\n\n1. Data Augmentation and Generators:\n","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Initialize ImageDataGenerator for data augmentation\ntrain_datagen = ImageDataGenerator(\n    rescale=1./255,        # Normalize pixel values to [0, 1]\n    rotation_range=40,     # Random rotations\n    width_shift_range=0.2, # Random horizontal shifts\n    height_shift_range=0.2,# Random vertical shifts\n    shear_range=0.2,       # Random shearing\n    zoom_range=0.2,        # Random zoom\n    horizontal_flip=True,  # Random horizontal flips\n    fill_mode='nearest'    # Filling mode for new pixels\n)\n\n# Create an ImageDataGenerator for validation data (without augmentation)\nval_datagen = ImageDataGenerator(rescale=1./255)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T06:52:51.812474Z","iopub.execute_input":"2024-08-22T06:52:51.812871Z","iopub.status.idle":"2024-08-22T06:52:51.818885Z","shell.execute_reply.started":"2024-08-22T06:52:51.812843Z","shell.execute_reply":"2024-08-22T06:52:51.817848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**2.Create Data Generators:**","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Define paths\ntrain_images_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images'\ntrain_csv_path = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train.csv'\n\n# Initialize ImageDataGenerator for data augmentation\ntrain_datagen = ImageDataGenerator(\n    rescale=1./255,        # Normalize pixel values to [0, 1]\n    rotation_range=40,     # Random rotations\n    width_shift_range=0.2, # Random horizontal shifts\n    height_shift_range=0.2,# Random vertical shifts\n    shear_range=0.2,       # Random shearing\n    zoom_range=0.2,        # Random zoom\n    horizontal_flip=True,  # Random horizontal flips\n    fill_mode='nearest'    # Filling mode for new pixels\n)\n\n# Create an ImageDataGenerator for validation data (without augmentation)\nval_datagen = ImageDataGenerator(rescale=1./255)\n\n# Create data generators\ntrain_generator = train_datagen.flow_from_directory(\n    train_images_path,\n    target_size=(img_height, img_width),\n    batch_size=32,\n    class_mode='sparse'  # Use 'sparse' for integer labels\n)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T06:56:22.050096Z","iopub.execute_input":"2024-08-22T06:56:22.050532Z","iopub.status.idle":"2024-08-22T06:56:25.157762Z","shell.execute_reply.started":"2024-08-22T06:56:22.050504Z","shell.execute_reply":"2024-08-22T06:56:25.156865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check a batch from the train generator\nx_batch, y_batch = next(train_generator)\nprint(f\"Batch image shape: {x_batch.shape}\")\nprint(f\"Batch labels shape: {y_batch.shape}\")\n\n# Check a batch from the validation generator\nx_val_batch, y_val_batch = next(validation_generator)\nprint(f\"Validation batch image shape: {x_val_batch.shape}\")\nprint(f\"Validation batch labels shape: {y_val_batch.shape}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:04:22.527678Z","iopub.execute_input":"2024-08-22T07:04:22.528052Z","iopub.status.idle":"2024-08-22T07:04:22.534292Z","shell.execute_reply.started":"2024-08-22T07:04:22.528025Z","shell.execute_reply":"2024-08-22T07:04:22.533188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Verify the data generators\nprint(f\"Number of training samples: {train_generator.samples}\")\nprint(f\"Number of validation samples: {validation_generator.samples}\")\nprint(f\"Batch size: {train_generator.batch_size}\")\nprint(f\"Classes: {train_generator.class_indices}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:09:10.965555Z","iopub.execute_input":"2024-08-22T07:09:10.965926Z","iopub.status.idle":"2024-08-22T07:09:10.971861Z","shell.execute_reply.started":"2024-08-22T07:09:10.965898Z","shell.execute_reply":"2024-08-22T07:09:10.970946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if there are images loaded\nx_batch, y_batch = next(train_generator)\nprint(f\"Batch image shape: {x_batch.shape}\")\nprint(f\"Batch labels shape: {y_batch.shape}\")\n\nx_val_batch, y_val_batch = next(validation_generator)\nprint(f\"Validation batch image shape: {x_val_batch.shape}\")\nprint(f\"Validation batch labels shape: {y_val_batch.shape}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:09:45.788511Z","iopub.execute_input":"2024-08-22T07:09:45.788849Z","iopub.status.idle":"2024-08-22T07:09:45.794797Z","shell.execute_reply.started":"2024-08-22T07:09:45.788826Z","shell.execute_reply":"2024-08-22T07:09:45.793947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense\n\n# Example model definition\nmodel = Sequential([\n    Conv2D(32, (3, 3), activation='relu', input_shape=(img_height, img_width, 3)),\n    MaxPooling2D((2, 2)),\n    Flatten(),\n    Dense(128, activation='relu'),\n    Dense(len(train_generator.class_indices), activation='softmax')\n])\n\nmodel.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:10:18.676928Z","iopub.execute_input":"2024-08-22T07:10:18.677429Z","iopub.status.idle":"2024-08-22T07:10:18.71726Z","shell.execute_reply.started":"2024-08-22T07:10:18.677396Z","shell.execute_reply":"2024-08-22T07:10:18.716533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Define image dimensions\nimg_height = 150\nimg_width = 150\n\n# Initialize ImageDataGenerator for data augmentation with validation split\ntrain_datagen = ImageDataGenerator(\n    rescale=1./255,        # Normalize pixel values to [0, 1]\n    rotation_range=40,     # Random rotations\n    width_shift_range=0.2, # Random horizontal shifts\n    height_shift_range=0.2,# Random vertical shifts\n    shear_range=0.2,       # Random shearing\n    zoom_range=0.2,        # Random zoom\n    horizontal_flip=True,  # Random horizontal flips\n    fill_mode='nearest',   # Filling mode for new pixels\n    validation_split=0.2   # Split data into training and validation sets\n)\n\n# Create data generators with validation split\ntrain_generator = train_datagen.flow_from_directory(\n    '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images',  # Absolute path to the training images\n    target_size=(img_height, img_width),\n    batch_size=32,\n    class_mode='sparse',    # Use 'sparse' for integer labels\n    subset='training'       # Use 'training' subset\n)\n\nvalidation_generator = train_datagen.flow_from_directory(\n    '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images',  # Absolute path to the training images\n    target_size=(img_height, img_width),\n    batch_size=32,\n    class_mode='sparse',    # Use 'sparse' for integer labels\n    subset='validation'     # Use 'validation' subset\n)\n\n# Check if data generators are correctly set up\nprint(f\"Training images: {train_generator.samples}, Validation images: {validation_generator.samples}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:18:41.000468Z","iopub.execute_input":"2024-08-22T07:18:41.001229Z","iopub.status.idle":"2024-08-22T07:18:45.254106Z","shell.execute_reply.started":"2024-08-22T07:18:41.0012Z","shell.execute_reply":"2024-08-22T07:18:45.253183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if data generators are yielding batches correctly\ntry:\n    x_batch, y_batch = next(train_generator)\n    x_val_batch, y_val_batch = next(validation_generator)\n    print(f\"Train batch image shape: {x_batch.shape}\")\n    print(f\"Train batch labels shape: {y_batch.shape}\")\n    print(f\"Validation batch image shape: {x_val_batch.shape}\")\n    print(f\"Validation batch labels shape: {y_val_batch.shape}\")\nexcept Exception as e:\n    print(f\"Error with data generators: {e}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:21:29.452347Z","iopub.execute_input":"2024-08-22T07:21:29.452722Z","iopub.status.idle":"2024-08-22T07:21:29.459702Z","shell.execute_reply.started":"2024-08-22T07:21:29.452695Z","shell.execute_reply":"2024-08-22T07:21:29.458602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example of model compilation\nfrom tensorflow.keras.optimizers import Adam\n\nmodel.compile(\n    optimizer=Adam(),\n    loss='sparse_categorical_crossentropy',\n    metrics=['accuracy']\n)\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:21:54.352087Z","iopub.execute_input":"2024-08-22T07:21:54.352479Z","iopub.status.idle":"2024-08-22T07:21:54.36126Z","shell.execute_reply.started":"2024-08-22T07:21:54.352452Z","shell.execute_reply":"2024-08-22T07:21:54.360334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Define image dimensions\nimg_height = 150\nimg_width = 150\n\n# Initialize ImageDataGenerator for data augmentation with validation split\ntrain_datagen = ImageDataGenerator(\n    rescale=1./255,        # Normalize pixel values to [0, 1]\n    rotation_range=40,     # Random rotations\n    width_shift_range=0.2, # Random horizontal shifts\n    height_shift_range=0.2,# Random vertical shifts\n    shear_range=0.2,       # Random shearing\n    zoom_range=0.2,        # Random zoom\n    horizontal_flip=True,  # Random horizontal flips\n    fill_mode='nearest',   # Filling mode for new pixels\n    validation_split=0.2   # Split data into training and validation sets\n)\n\n# Create data generators with validation split\ntrain_generator = train_datagen.flow_from_directory(\n    '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images',  # Absolute path to the training images\n    target_size=(img_height, img_width),\n    batch_size=32,\n    class_mode='sparse',    # Use 'sparse' for integer labels\n    subset='training'       # Use 'training' subset\n)\n\nvalidation_generator = train_datagen.flow_from_directory(\n    '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_images',  # Absolute path to the training images\n    target_size=(img_height, img_width),\n    batch_size=32,\n    class_mode='sparse',    # Use 'sparse' for integer labels\n    subset='validation'     # Use 'validation' subset\n)\n\n# Check if data generators are correctly set up\nprint(f\"Training images: {train_generator.samples}, Validation images: {validation_generator.samples}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-08-22T07:23:24.880337Z","iopub.execute_input":"2024-08-22T07:23:24.880975Z","iopub.status.idle":"2024-08-22T07:23:29.798704Z","shell.execute_reply.started":"2024-08-22T07:23:24.880946Z","shell.execute_reply":"2024-08-22T07:23:29.797701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. Train the Model:\n\nFit the model using the data generators.","metadata":{}},{"cell_type":"markdown","source":"**Additional Tips:**\n\n* Utilize Cross-Validation: Implement cross-validation techniques to ensure the robustness of your model.\n\n* Feature Engineering: Experiment with different ways to preprocess the images and create additional features from the MRI scans.\n\n* Model Ensembling: Consider ensembling multiple models to improve overall performance.","metadata":{}}]}