{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA Lumbar Spine Challenge\n\n## Acknowledge\n\nThis is a fork from a notebook original written by Abhinav Suri MPH, Andrew Wentland MD PhD, Hari Trivedi MD \n– Radiology Artificial Intelligence Data Standards Committee","metadata":{}},{"cell_type":"markdown","source":"## Competition Info\n\nLink: https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification","metadata":{}},{"cell_type":"markdown","source":"###  Overview\n\nThe goal of this competition is to create models that can be used to aid in the detection and classification of degenerative spine conditions using lumbar spine MR images. Competitors will develop models that simulate a radiologist's performance in diagnosing spine conditions.\n\n### Description\n\nLow back pain is the leading cause of disability worldwide, according to the World Health Organization, affecting 619 million people in 2020. Most people experience low back pain at some point in their lives, with the frequency increasing with age. Pain and restricted mobility are often symptoms of spondylosis, a set of degenerative spine conditions including degeneration of intervertebral discs and subsequent narrowing of the spinal canal (spinal stenosis), subarticular recesses, or neural foramen with associated compression or irritations of the nerves in the low back.\n\nMagnetic resonance imaging (MRI) provides a detailed view of the lumbar spine vertebra, discs and nerves, enabling radiologists to assess the presence and severity of these conditions. Proper diagnosis and grading of these conditions help guide treatment and potential surgery to help alleviate back pain and improve overall health and quality of life for patients.\n\nRSNA has teamed with the American Society of Neuroradiology (ASNR) to conduct this competition exploring whether artificial intelligence can be used to aid in the detection and classification of degenerative spine conditions using lumbar spine MR images.\n\nThe challenge will focus on the classification of five lumbar spine degenerative conditions: Left Neural Foraminal Narrowing, Right Neural Foraminal Narrowing, Left Subarticular Stenosis, Right Subarticular Stenosis, and Spinal Canal Stenosis. For each imaging study in the dataset, we’ve provided severity scores (Normal/Mild, Moderate, or Severe) for each of the five conditions across the intervertebral disc levels L1/L2, L2/L3, L3/L4, L4/L5, and L5/S1.\n\nTo create the ground truth dataset, the RSNA challenge planning task force collected imaging data sourced from eight sites on five continents. This multi-institutional, expertly curated dataset promises to improve standardized classification of degenerative lumbar spine conditions and enable development of tools to automate accurate and rapid disease classification.\n\nChallenge winners will be recognized at an event during the RSNA 2024 annual meeting. For more information on the challenge, contact RSNA Informatics staff at informatics@rsna.org.","metadata":{}},{"cell_type":"markdown","source":"## Expected Directory Structure\nWe expect the following to be in your working directory to run this notebook.\n\n```\n.\n├── ExploreData.ipynb **This notebook\n└── /kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/\n    ├── test_images/\n    │   ├── 1005139/\n    │   │   └── 609308237/\n    │   │       ├── 1.dcm\n    │   │       └── ...\n    │   └── ...\n    ├── test_series_descriptions.csv\n    ├── train_images/\n    │   ├── 4003253/\n    │   │   └── 702807833/\n    │   │       ├── 1.dcm\n    │   │       └── ...\n    │   └── ...\n    ├── train_label_coordinates.csv\n    ├── train_series_descriptions.csv\n    └── train.csv\n```","metadata":{}},{"cell_type":"markdown","source":"## Loading Diagnosis Information","metadata":{}},{"cell_type":"markdown","source":"In this part of the notebook, we'll load in information about the diagnoses present in our dataset to give an overview of the distribution of cases.","metadata":{}},{"cell_type":"code","source":"import os\nimport glob\nimport warnings\n\nfrom pathlib import Path\n\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport cv2\nimport pydicom\nimport numpy as np\nfrom tqdm import tqdm\n","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:18.181331Z","iopub.execute_input":"2024-09-25T11:24:18.181591Z","iopub.status.idle":"2024-09-25T11:24:19.388735Z","shell.execute_reply.started":"2024-09-25T11:24:18.181566Z","shell.execute_reply":"2024-09-25T11:24:19.387772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_ROOT_PATH = Path('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/')","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:19.3903Z","iopub.execute_input":"2024-09-25T11:24:19.390683Z","iopub.status.idle":"2024-09-25T11:24:19.394747Z","shell.execute_reply.started":"2024-09-25T11:24:19.390657Z","shell.execute_reply":"2024-09-25T11:24:19.39388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(DATA_ROOT_PATH / 'train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:19.395757Z","iopub.execute_input":"2024-09-25T11:24:19.39608Z","iopub.status.idle":"2024-09-25T11:24:19.429133Z","shell.execute_reply.started":"2024-09-25T11:24:19.396054Z","shell.execute_reply":"2024-09-25T11:24:19.428454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"total_cases = len(train)\nprint(\"Total Cases: \", total_cases)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:19.431024Z","iopub.execute_input":"2024-09-25T11:24:19.431278Z","iopub.status.idle":"2024-09-25T11:24:19.436019Z","shell.execute_reply.started":"2024-09-25T11:24:19.431256Z","shell.execute_reply":"2024-09-25T11:24:19.435144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Missing data:\")\ntotal_cases - train.count()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:19.437364Z","iopub.execute_input":"2024-09-25T11:24:19.437647Z","iopub.status.idle":"2024-09-25T11:24:19.460477Z","shell.execute_reply.started":"2024-09-25T11:24:19.437625Z","shell.execute_reply":"2024-09-25T11:24:19.459616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:19.461704Z","iopub.execute_input":"2024-09-25T11:24:19.46205Z","iopub.status.idle":"2024-09-25T11:24:19.48753Z","shell.execute_reply.started":"2024-09-25T11:24:19.462017Z","shell.execute_reply":"2024-09-25T11:24:19.486746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categories = set()\n\nfor c in train.drop(columns=[\"study_id\"]).columns:\n    categories |= set(train[c].unique().tolist())\n    \ncategories","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:19.488741Z","iopub.execute_input":"2024-09-25T11:24:19.489018Z","iopub.status.idle":"2024-09-25T11:24:19.501183Z","shell.execute_reply.started":"2024-09-25T11:24:19.488989Z","shell.execute_reply":"2024-09-25T11:24:19.500283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"figure, axis = plt.subplots(1,3, figsize=(20,5)) \nfor idx, d in enumerate(['foraminal', 'subarticular', 'canal']):\n    diagnosis = list(filter(lambda x: x.find(d) > -1, train.columns))\n    dff = train[diagnosis]\n    with warnings.catch_warnings():\n        warnings.simplefilter(action='ignore', category=FutureWarning)\n        value_counts = dff.apply(pd.value_counts).fillna(0).T\n    value_counts.plot(kind='bar', stacked=True, ax=axis[idx])\n    axis[idx].set_title(f'{d} distribution')","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:19.502445Z","iopub.execute_input":"2024-09-25T11:24:19.502715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note from the original author:\n> As expected many of our patients have normal/mild grades for each of the diagnoses categories. You'll also see that for some of the diagnoses (particularly subarticular stenosis category), there are missing data. This is due to the fact that some of the images do not have those regions visualized (more specifically the most superior vertebral bodies were less likely to make it into the imaging field).","metadata":{}},{"cell_type":"markdown","source":"## Loading in images","metadata":{}},{"cell_type":"markdown","source":"### Grab metadata for each scan.\nFor each scan let's create an object with the following structure:\n\n```\nmeta_obj = {\n    StudyInstanceUID: {\n        'folder_path': ... # path to the folder,\n        'SeriesInstanceUIDs': [ Array of the SeriesInstanceUIDs ],\n        'SeriesDescriptions' [ Array of the Series Descriptions ]\n    }, ...\n}\n```","metadata":{}},{"cell_type":"code","source":"# List out all of the Studies we have on patients.\nTRAIN_IMAGES_PATH = DATA_ROOT_PATH / 'train_images'\npart_1 = os.listdir(TRAIN_IMAGES_PATH)\npart_1 = list(filter(lambda x: x.find('.DS') == -1, part_1))","metadata":{"execution":{"iopub.status.idle":"2024-09-25T11:24:20.852988Z","shell.execute_reply.started":"2024-09-25T11:24:20.709081Z","shell.execute_reply":"2024-09-25T11:24:20.852154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta_f = pd.read_csv(DATA_ROOT_PATH / 'train_series_descriptions.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:20.857148Z","iopub.execute_input":"2024-09-25T11:24:20.85743Z","iopub.status.idle":"2024-09-25T11:24:20.873404Z","shell.execute_reply.started":"2024-09-25T11:24:20.857405Z","shell.execute_reply":"2024-09-25T11:24:20.872582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p1 = [(x, TRAIN_IMAGES_PATH / x) for x in part_1]\nmeta_obj = {\n    p[0]: {\n        'folder_path': p[1], \n        'SeriesInstanceUIDs': [],\n        'SeriesDescriptions': []  # NEW\n    } \n    for p in p1\n}","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:20.874611Z","iopub.execute_input":"2024-09-25T11:24:20.875109Z","iopub.status.idle":"2024-09-25T11:24:20.898498Z","shell.execute_reply.started":"2024-09-25T11:24:20.875083Z","shell.execute_reply":"2024-09-25T11:24:20.89766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for m in meta_obj:\n    meta_obj[m]['SeriesInstanceUIDs'] = list(\n        filter(lambda x: x.find('.DS') == -1, \n               os.listdir(meta_obj[m]['folder_path'])\n              )\n    )","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:20.899669Z","iopub.execute_input":"2024-09-25T11:24:20.899978Z","iopub.status.idle":"2024-09-25T11:24:26.105599Z","shell.execute_reply.started":"2024-09-25T11:24:20.899934Z","shell.execute_reply":"2024-09-25T11:24:26.104616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# grabs the correspoding series descriptions\nfor k in tqdm(meta_obj):\n    for s in meta_obj[k]['SeriesInstanceUIDs']:\n        try:\n            meta_obj[k]['SeriesDescriptions'].append(\n                df_meta_f[(df_meta_f['study_id'] == int(k)) & \n                (df_meta_f['series_id'] == int(s))]['series_description'].iloc[0])\n        except:\n            print(\"Failed on\", s, k)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:26.1071Z","iopub.execute_input":"2024-09-25T11:24:26.107386Z","iopub.status.idle":"2024-09-25T11:24:29.350153Z","shell.execute_reply.started":"2024-09-25T11:24:26.107361Z","shell.execute_reply":"2024-09-25T11:24:29.349132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_obj[list(meta_obj.keys())[1]]","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:29.351263Z","iopub.execute_input":"2024-09-25T11:24:29.351578Z","iopub.status.idle":"2024-09-25T11:24:29.358302Z","shell.execute_reply.started":"2024-09-25T11:24:29.351552Z","shell.execute_reply":"2024-09-25T11:24:29.357358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Pull up images for one patient","metadata":{}},{"cell_type":"markdown","source":"Now that we've made our meta obj, let's pull up all the images for one patient.","metadata":{}},{"cell_type":"code","source":"patient = train.iloc[1]","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:29.359536Z","iopub.execute_input":"2024-09-25T11:24:29.359945Z","iopub.status.idle":"2024-09-25T11:24:29.366308Z","shell.execute_reply.started":"2024-09-25T11:24:29.359918Z","shell.execute_reply":"2024-09-25T11:24:29.365336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ptobj = meta_obj[str(patient['study_id'])]","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:29.367446Z","iopub.execute_input":"2024-09-25T11:24:29.36775Z","iopub.status.idle":"2024-09-25T11:24:29.374821Z","shell.execute_reply.started":"2024-09-25T11:24:29.367727Z","shell.execute_reply":"2024-09-25T11:24:29.373986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(ptobj)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:29.376019Z","iopub.execute_input":"2024-09-25T11:24:29.376443Z","iopub.status.idle":"2024-09-25T11:24:29.383914Z","shell.execute_reply.started":"2024-09-25T11:24:29.376413Z","shell.execute_reply":"2024-09-25T11:24:29.383093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get data into the format\n\"\"\"\nim_list_dcm = {\n    '{SeriesInstanceUID}': {\n        'images': [\n            {'SOPInstanceUID': ...,\n             'dicom': PyDicom object\n            },\n            ...,\n        ],\n        'description': # SeriesDescription\n    },\n    ...\n}\n\"\"\"\nim_list_dcm = {}\nfor idx, i in enumerate(ptobj['SeriesInstanceUIDs']):\n    im_list_dcm[i] = {'images': [], 'description': ptobj['SeriesDescriptions'][idx]}\n    images = glob.glob(f\"{ptobj['folder_path']}/{ptobj['SeriesInstanceUIDs'][idx]}/*.dcm\")\n    for j in sorted(images, key=lambda x: int(x.split('/')[-1].replace('.dcm', ''))):\n        im_list_dcm[i]['images'].append({\n            'SOPInstanceUID': j.split('/')[-1].replace('.dcm', ''), \n            'dicom': pydicom.dcmread(j) })","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:29.385016Z","iopub.execute_input":"2024-09-25T11:24:29.385334Z","iopub.status.idle":"2024-09-25T11:24:29.937527Z","shell.execute_reply.started":"2024-09-25T11:24:29.385303Z","shell.execute_reply":"2024-09-25T11:24:29.936637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to display images\ndef display_images(images, title, max_images_per_row=4):\n    # Calculate the number of rows needed\n    num_images = len(images)\n    num_rows = (num_images + max_images_per_row - 1) // max_images_per_row  # Ceiling division\n\n    # Create a subplot grid\n    fig, axes = plt.subplots(num_rows, max_images_per_row, figsize=(5, 1.5 * num_rows))\n    \n    # Flatten axes array for easier looping if there are multiple rows\n    if num_rows > 1:\n        axes = axes.flatten()\n    else:\n        axes = [axes]  # Make it iterable for consistency\n\n    # Plot each image\n    for idx, image in enumerate(images):\n        ax = axes[idx]\n        ax.imshow(image, cmap='gray')  # Assuming grayscale for simplicity, change cmap as needed\n        ax.axis('off')  # Hide axes\n\n    # Turn off unused subplots\n    for idx in range(num_images, len(axes)):\n        axes[idx].axis('off')\n    fig.suptitle(title, fontsize=16)\n\n    plt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:29.938711Z","iopub.execute_input":"2024-09-25T11:24:29.939209Z","iopub.status.idle":"2024-09-25T11:24:29.947209Z","shell.execute_reply.started":"2024-09-25T11:24:29.939176Z","shell.execute_reply":"2024-09-25T11:24:29.946221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in im_list_dcm:\n    display_images([x['dicom'].pixel_array for x in im_list_dcm[i]['images']], \n                   im_list_dcm[i]['description'])","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:29.94822Z","iopub.execute_input":"2024-09-25T11:24:29.948516Z","iopub.status.idle":"2024-09-25T11:24:35.587727Z","shell.execute_reply.started":"2024-09-25T11:24:29.948491Z","shell.execute_reply":"2024-09-25T11:24:35.58675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Look at the coordinates of pathologies\n\nWe can also display the coordinates of the pathologies that are annotated for each of the patients.","metadata":{}},{"cell_type":"code","source":"df_coor = pd.read_csv(DATA_ROOT_PATH / 'train_label_coordinates.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:35.58886Z","iopub.execute_input":"2024-09-25T11:24:35.589161Z","iopub.status.idle":"2024-09-25T11:24:35.694875Z","shell.execute_reply.started":"2024-09-25T11:24:35.589134Z","shell.execute_reply":"2024-09-25T11:24:35.69408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_coor.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:35.69615Z","iopub.execute_input":"2024-09-25T11:24:35.696871Z","iopub.status.idle":"2024-09-25T11:24:35.708421Z","shell.execute_reply.started":"2024-09-25T11:24:35.696837Z","shell.execute_reply":"2024-09-25T11:24:35.707501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_coor_on_img(c, i, title):\n    center_coordinates = (int(c['x']), int(c['y']))\n    radius = 10\n    color = (255, 0, 0)  # Red color in BGR\n    thickness = 2\n    IMG = i['dicom'].pixel_array\n    IMG_normalized = cv2.normalize(IMG, None, alpha=0, beta=255, norm_type=cv2.NORM_MINMAX, dtype=cv2.CV_8U)\n    \n    IMG_with_circle = cv2.circle(IMG_normalized.copy(), center_coordinates, radius, color, thickness)\n    \n    # Convert the image from BGR to RGB for correct color display in matplotlib\n    IMG_with_circle = cv2.cvtColor(IMG_with_circle, cv2.COLOR_BGR2RGB)\n    \n    # Display the image\n    plt.imshow(IMG_with_circle)\n    plt.axis('off')  # Turn off axis numbers and ticks\n    plt.title(title)\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:35.709538Z","iopub.execute_input":"2024-09-25T11:24:35.709858Z","iopub.status.idle":"2024-09-25T11:24:35.717431Z","shell.execute_reply.started":"2024-09-25T11:24:35.709832Z","shell.execute_reply":"2024-09-25T11:24:35.716612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"coor_entries = df_coor[df_coor['study_id'] == int(patient['study_id'])]","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:35.718416Z","iopub.execute_input":"2024-09-25T11:24:35.718742Z","iopub.status.idle":"2024-09-25T11:24:35.72681Z","shell.execute_reply.started":"2024-09-25T11:24:35.718715Z","shell.execute_reply":"2024-09-25T11:24:35.725921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Only showing severe cases for this patient\")\nfor idc, c in coor_entries.iterrows():\n    for i in im_list_dcm[str(c['series_id'])]['images']:\n        if int(i['SOPInstanceUID']) == int(c['instance_number']):\n            try:\n                patient_severity = patient[\n                    f\"{c['condition'].lower().replace(' ', '_')}_{c['level'].lower().replace('/', '_')}\"\n                ]\n            except Exception as e:\n                patient_severity = \"unknown severity\"\n            title = f\"{i['SOPInstanceUID']} \\n{c['level']}, {c['condition']}: {patient_severity} \\n{c['x']}, {c['y']}\"\n            if patient_severity == 'Severe':\n                display_coor_on_img(c, i, title)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:35.727924Z","iopub.execute_input":"2024-09-25T11:24:35.728316Z","iopub.status.idle":"2024-09-25T11:24:36.069689Z","shell.execute_reply.started":"2024-09-25T11:24:35.728292Z","shell.execute_reply":"2024-09-25T11:24:36.068682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pandas.api.types\nimport sklearn.metrics\n\n\nclass ParticipantVisibleError(Exception):\n    pass\n\n\ndef get_condition(full_location: str) -> str:\n    # Given an input like spinal_canal_stenosis_l1_l2 extracts 'spinal'\n    for injury_condition in ['spinal', 'foraminal', 'subarticular']:\n        if injury_condition in full_location:\n            return injury_condition\n    raise ValueError(f'condition not found in {full_location}')\n\n\ndef score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        any_severe_scalar: float\n    ) -> float:\n    '''\n    Pseudocode:\n    1. Calculate the sample weighted log loss for each medical condition:\n    2. Derive a new any_severe label.\n    3. Calculate the sample weighted log loss for the new any_severe label.\n    4. Return the average of all of the label group log losses as the final score, normalized for the number of columns in each group.\n       This mitigates the impact of spinal stenosis having only half as many columns as the other two conditions.\n    '''\n\n    target_levels = ['normal_mild', 'moderate', 'severe']\n\n    # Run basic QC checks on the inputs\n    if not pandas.api.types.is_numeric_dtype(submission[target_levels].values):\n        raise ParticipantVisibleError('All submission values must be numeric')\n\n    if not np.isfinite(submission[target_levels].values).all():\n        raise ParticipantVisibleError('All submission values must be finite')\n\n    if solution[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All labels must be at least zero')\n    if submission[target_levels].min().min() < 0:\n        raise ParticipantVisibleError('All predictions must be at least zero')\n\n    solution['study_id'] = solution['row_id'].apply(lambda x: x.split('_')[0])\n    solution['location'] = solution['row_id'].apply(lambda x: '_'.join(x.split('_')[1:]))\n    solution['condition'] = solution['row_id'].apply(get_condition)\n\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n    assert sorted(submission.columns) == sorted(target_levels)\n\n    submission['study_id'] = solution['study_id']\n    submission['location'] = solution['location']\n    submission['condition'] = solution['condition']\n\n    condition_losses = []\n    condition_weights = []\n    for condition in ['spinal', 'foraminal', 'subarticular']:\n        condition_indices = solution.loc[solution['condition'] == condition].index.values\n        condition_loss = sklearn.metrics.log_loss(\n            y_true=solution.loc[condition_indices, target_levels].values,\n            y_pred=submission.loc[condition_indices, target_levels].values,\n            sample_weight=solution.loc[condition_indices, 'sample_weight'].values\n        )\n        condition_losses.append(condition_loss)\n        condition_weights.append(1 / solution.loc[condition_indices, 'location'].nunique())\n\n        \n    any_severe_spinal_labels = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_weights = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\n    any_severe_spinal_predictions = pd.Series(submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n    any_severe_spinal_loss = sklearn.metrics.log_loss(\n        y_true=any_severe_spinal_labels,\n        y_pred=any_severe_spinal_predictions,\n        sample_weight=any_severe_spinal_weights\n    )\n    condition_losses.append(any_severe_spinal_loss)\n    condition_weights.append(any_severe_scalar)\n    return np.average(condition_losses, weights=condition_weights)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:36.071197Z","iopub.execute_input":"2024-09-25T11:24:36.071757Z","iopub.status.idle":"2024-09-25T11:24:36.586245Z","shell.execute_reply.started":"2024-09-25T11:24:36.071726Z","shell.execute_reply":"2024-09-25T11:24:36.58546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install --quiet git+https://github.com/facebookresearch/segment-anything-2/","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:24:36.587536Z","iopub.execute_input":"2024-09-25T11:24:36.587988Z","iopub.status.idle":"2024-09-25T11:30:08.217399Z","shell.execute_reply.started":"2024-09-25T11:24:36.587935Z","shell.execute_reply":"2024-09-25T11:30:08.215842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\n\nimport numpy as np\nfrom PIL import Image\nimport pandas as pd\nimport pydicom\nimport matplotlib.pyplot as plt\n\nfrom sam2.sam2_image_predictor import SAM2ImagePredictor","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:08.226545Z","iopub.execute_input":"2024-09-25T11:30:08.227347Z","iopub.status.idle":"2024-09-25T11:30:12.468831Z","shell.execute_reply.started":"2024-09-25T11:30:08.227304Z","shell.execute_reply":"2024-09-25T11:30:12.468014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIR = Path(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification\")\nIMG_DIR = DATA_DIR / \"train_images\"\nCOOR_PATH = DATA_DIR / \"train_label_coordinates.csv\"","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:12.469911Z","iopub.execute_input":"2024-09-25T11:30:12.470387Z","iopub.status.idle":"2024-09-25T11:30:14.090338Z","shell.execute_reply.started":"2024-09-25T11:30:12.47036Z","shell.execute_reply":"2024-09-25T11:30:14.089254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv(COOR_PATH)\ndf = df.set_index([\"study_id\", \"series_id\", \"instance_number\"])[[\"x\", \"y\", \"condition\"]]\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:14.091647Z","iopub.execute_input":"2024-09-25T11:30:14.091946Z","iopub.status.idle":"2024-09-25T11:30:14.252016Z","shell.execute_reply.started":"2024-09-25T11:30:14.09192Z","shell.execute_reply":"2024-09-25T11:30:14.251125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df[df[\"condition\"].str.contains(\"Spinal\")]","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:14.253035Z","iopub.execute_input":"2024-09-25T11:30:14.253306Z","iopub.status.idle":"2024-09-25T11:30:14.281353Z","shell.execute_reply.started":"2024-09-25T11:30:14.253282Z","shell.execute_reply":"2024-09-25T11:30:14.280399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_img_path(study_id, series_id, instance_number):\n    img_path = IMG_DIR / str(study_id) / str(series_id) / f\"{instance_number}.dcm\"\n    return img_path\n\n\ndef load_dcm_img(path: Path) -> np.ndarray:\n    dicom = pydicom.read_file(path)\n    img: np.ndarray = dicom.pixel_array\n    img = img.clip(np.percentile(img, 1), np.percentile(img, 99))\n    img = img - np.min(img)\n    img = img / np.max(img)\n    img = (img * 255).astype(np.uint8)\n    return img\n\n\n# https://github.com/facebookresearch/segment-anything-2/blob/main/notebooks/image_predictor_example.ipynb\ndef show_mask(mask, ax):\n    color = np.array([30/255, 144/255, 255/255, 0.6])\n    h, w = mask.shape[-2:]\n    mask = mask.astype(np.uint8)\n    mask_image =  mask.reshape(h, w, 1) * color.reshape(1, 1, -1)\n    ax.imshow(mask_image)\n\n    \ndef show_points(coords, labels, ax, marker_size=375):\n    pos_points = coords[labels==1]\n    neg_points = coords[labels==0]\n    ax.scatter(\n        pos_points[:, 0],\n        pos_points[:, 1],\n        color='green',\n        marker='o',\n        s=marker_size,\n        edgecolor='green',\n        linewidth=1.25,\n        alpha=0.5  # Set transparency to 50%\n    )\n    ax.scatter(\n        neg_points[:, 0],\n        neg_points[:, 1],\n        color='red',\n        marker='o',\n        s=marker_size,\n        edgecolor='red',\n        linewidth=1.25,\n        alpha=0.5  # Set transparency to 50%\n    )\n\n    \n    \ndef show_segmentation(img, masks, input_points, input_labels, title=\"\") -> None:\n    fig, (ax1, ax2) = plt.subplots(ncols=2, figsize=(10, 10))\n\n    ax1.imshow(img)\n    show_points(input_points, input_labels, ax1)\n    ax1.axis(\"off\")\n\n    ax2.imshow(img)\n    show_mask(masks,ax2)\n    ax2.axis(\"off\")\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:14.282617Z","iopub.execute_input":"2024-09-25T11:30:14.283328Z","iopub.status.idle":"2024-09-25T11:30:14.296114Z","shell.execute_reply.started":"2024-09-25T11:30:14.283293Z","shell.execute_reply":"2024-09-25T11:30:14.295202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor = SAM2ImagePredictor.from_pretrained(\"facebook/sam2-hiera-large\")","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:14.297215Z","iopub.execute_input":"2024-09-25T11:30:14.29751Z","iopub.status.idle":"2024-09-25T11:30:27.265688Z","shell.execute_reply.started":"2024-09-25T11:30:14.297487Z","shell.execute_reply":"2024-09-25T11:30:27.264661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for (study_id, series_id, instance_number), chunk in df.groupby(level=[0, 1, 2])[[\"x\", \"y\"]]:\n    \n    img_path = get_img_path(study_id, series_id, instance_number)\n    img = load_dcm_img(img_path)\n    img = img[..., None].repeat(3, -1)\n    \n    input_points = chunk.values[0:1]\n    input_labels = np.ones((input_points.shape[0], ))\n    \n    print(input_points)\n    print(input_labels)\n    \n    predictor.set_image(img)\n    \n    masks, scores, logits = predictor.predict(\n        point_coords=input_points,\n        point_labels=input_labels,\n        multimask_output=False,\n    )\n    \n    print(\"scores:\", scores)\n    print(\"logits\", logits)\n    \n    # you can try more images in this loop\n    show_segmentation(img, masks, input_points, input_labels)\n    \n    break","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:27.267011Z","iopub.execute_input":"2024-09-25T11:30:27.267945Z","iopub.status.idle":"2024-09-25T11:30:28.793843Z","shell.execute_reply.started":"2024-09-25T11:30:27.26791Z","shell.execute_reply":"2024-09-25T11:30:28.792918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Only showing severe cases for this patient\")\nfor idc, c in coor_entries.iterrows():\n    for i in im_list_dcm[str(c['series_id'])]['images']:\n        if int(i['SOPInstanceUID']) == int(c['instance_number']):\n            try:\n                patient_severity = patient[\n                    f\"{c['condition'].lower().replace(' ', '_')}_{c['level'].lower().replace('/', '_')}\"\n                ]\n            except Exception as e:\n                patient_severity = \"unknown severity\"\n            title = f\"{i['SOPInstanceUID']} \\n{c['level']}, {c['condition']}: {patient_severity} \\n{c['x']}, {c['y']}\"\n            if patient_severity == 'Severe':\n                display_coor_on_img(c, i, title)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:30:28.795016Z","iopub.execute_input":"2024-09-25T11:30:28.795314Z","iopub.status.idle":"2024-09-25T11:30:29.164322Z","shell.execute_reply.started":"2024-09-25T11:30:28.795289Z","shell.execute_reply":"2024-09-25T11:30:29.163293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import cv2\nimport numpy as np\n\n# Load the image in grayscale\nimg_file_path = str(IMG_DIR / str(study_id) / str(series_id) / f\"{instance_number}.dcm\")\nimage = load_dcm_img(img_file_path)\n\nprint(image)\n\n# Equalize the histogram\nimage = cv2.equalizeHist(image)\n\n# Apply median smoothing\nimage = cv2.medianBlur(image, 13)\n\n# Create a temporary image for thresholding\ntmp = np.zeros_like(image)\n\n# Apply binary thresholding to remove unwanted information\n_, tmp = cv2.threshold(image, 0, 120, cv2.THRESH_BINARY)\n\n# Subtract the thresholded image from the original\nimage = cv2.subtract(image, tmp)\n\n# Create a cross-shaped structuring element\nkernel_cross = cv2.getStructuringElement(cv2.MORPH_CROSS, (5, 5))\n\n# Apply morphological gradient operation to enhance edges\nimage_cross = cv2.morphologyEx(image, cv2.MORPH_GRADIENT, kernel_cross, iterations=3)\n\n# Show the result\nplt.figure(figsize=(10, 8))\nplt.imshow(image)\nplt.title(\"Original Image\")\nplt.axis('off')  # Hide axis\nplt.show()\nplt.figure(figsize=(10, 8))\nplt.imshow(image_cross)\nplt.title(\"Processed Image\")\nplt.axis('off')  # Hide axis\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T11:50:21.489804Z","iopub.execute_input":"2024-09-25T11:50:21.490731Z","iopub.status.idle":"2024-09-25T11:50:22.141142Z","shell.execute_reply.started":"2024-09-25T11:50:21.490693Z","shell.execute_reply":"2024-09-25T11:50:22.140271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:58:18.5504Z","iopub.execute_input":"2024-09-25T13:58:18.550873Z","iopub.status.idle":"2024-09-25T13:58:18.563186Z","shell.execute_reply.started":"2024-09-25T13:58:18.550848Z","shell.execute_reply":"2024-09-25T13:58:18.562404Z"},"trusted":true},"execution_count":null,"outputs":[]}]}