{
  "id": 539453,
  "title": "3rd Place Solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/sonyspine-s-tkmn-moyashii-3rd-place-solution",
  "author_name": "",
  "post_date": "2024-10-28T10:42:41.113Z",
  "votes": 65,
  "comment_count": 10,
  "views": 0,
  "content": "<h1>3rd Place Solution</h1>\n<p>First and foremost, we would like to express our deepest gratitude to Kaggle and the competition organizers for providing this wonderful opportunity. We also thank all the participants for making this competition engaging and insightful.</p>\n<h2>Summary</h2>\n<p>We constructed a general two-stage pipeline:</p>\n<ul>\n<li><strong>Stage 1</strong>: Crop sagittal images at each disc level and crop axial images using disc level assignments and spinal canal positions.</li>\n<li><strong>Stage 2</strong>: Use a <strong>Center Classifier</strong> to classify the severity of Spinal Canal Stenosis and a <strong>Side Classifier</strong> to classify the severity of Neural Foraminal Narrowing and Subarticular Stenosis.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fed20bc73bd85e207bf670b752e038a32%2Fpipeline_overview.png?generation=1728438360443270&amp;alt=media\" alt=\"overview\"></p>\n<p>We will explain each process in detail below.</p>\n<h2>Stage 1</h2>\n<p>The responsibility of Stage 1 is to extract the information necessary for estimating disease severity from the input data.</p>\n<h3>1. Disc Level Keypoint Detector (CenterNet)</h3>\n<p>We built a CenterNet-based 2D keypoint detector using EfficientNetB6 as the backbone and FPN as the neck. By inputting sagittal images near the center of the body, we estimate the coordinates of each disc level. For training data, we used sagittal images from RSNA2024 that have coordinates of Spinal Canal Stenosis at all levels, as well as the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">Coordinate Pretraining Dataset</a>. By using the trained model to generate pseudo-labels on unused RSNA2024 data, we ultimately utilized all RSNA2024 data. Recognizing from several discussions that label noise existed, we manually reviewed all annotations and corrected erroneous labels by hand.</p>\n<h3>2. Crop Level</h3>\n<p>We cropped the sagittal images at each disc level. To ensure diversity in the input data, we adopted multiple cropping settings. There was almost no difference in accuracy due to cropping settings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc7d43fb7b862bdf4c95f8fcc0f18883e%2Fsagittal.png?generation=1728437481708829&amp;alt=media\" alt=\"crop_sagittal\"></p>\n<h3>3. Assign Level</h3>\n<p>Using the output of the Disc Level Keypoint Detector, we assign arbitrary disc levels to the axial slices. The processing flow is as follows:</p>\n<ol>\n<li>Convert the image coordinates of disc levels to real-world coordinates.</li>\n<li>Estimate the vertebral positions of L1, L2, …, S1 from the midpoints of each disc level (e.g., L1/L2, L2/L3, etc.). Since we cannot obtain coordinates for T12/L1 and S1/S2, we pseudo-calculate the coordinates for L1 and S1.</li>\n<li>Calculate the intersection points between the line segments connecting adjacent vertebrae and the axial planes, and assign the corresponding disc levels.</li>\n</ol>\n<h3>4. Spinal Canal Keypoint Detector (CenterNet)</h3>\n<p>We constructed a CenterNet-based 2D keypoint detector using EfficientNetB4 as the backbone and FPN as the neck. By inputting axial images, we estimate the coordinates of the spinal canal. Since the Y-coordinate can be accurately estimated from the results of the Disc Level Keypoint Detector but estimating the X-coordinate is challenging, we introduced this detector. For training data, we used axial images from RSNA2024 that have coordinates of Spinal Canal Stenosis.</p>\n<h3>5. Crop Spinal</h3>\n<p>Using the output from the Spinal Canal Keypoint Detector, we crop the necessary regions centered on the spinal canal. To ensure diversity in the input data, we cropped at multiple sizes. There was almost no difference in accuracy due to cropping methods.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fa41a778fc592d7f632e5659178fcbcd6%2Faxial.png?generation=1728437501864432&amp;alt=media\" alt=\"crop_axial\"></p>\n<h2>Stage 2</h2>\n<p>The responsibility of Stage 2 is to estimate the severity of each condition using the outputs from Stage 1.</p>\n<h3>6. Center Classifier (2D-Encoder + Attention)</h3>\n<p>We created a classification model to estimate the severity of Spinal Canal Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. For sagittal T1 and sagittal T2/STIR, we input 15 slices at equal intervals into an encoder to generate feature representations for each slice. For axial T2, we input 10 slices at equal intervals. These slice features are then input into an attention mechanism to learn the relationships between slices. To ensure model diversity, we created two models with different head structures. To improve accuracy, we used auxiliary losses such as the severity of other conditions and slice-level predictions. Additionally, increasing the loss weight for the Severe class, due to the metric specifications, was effective. Test-time augmentation (TTA) by flipping axial images also improved the leaderboard score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2F2cd3d34e44fcde5d8f429c4b2eee042d%2Fclassifier.png?generation=1728438379066904&amp;alt=media\" alt=\"classifier\"></p>\n<p>After verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels to estimate the severity of Spinal Canal Stenosis. During training, we treated each group of slices as an independent data point without depending on the disc level. In other words, the model is designed to consistently learn the characteristics of Spinal Canal Stenosis from slice groups at any disc level and predict the severity based on those features, without specializing in any specific disc level. This allowed us to secure five times the amount of data per condition, which we believe contributed to the improvement in accuracy.</p>\n<p>We used the following encoders:</p>\n<ul>\n<li>ResNet18 (160x160, 224x224)</li>\n<li>MNasNet-S (224x224)</li>\n<li>EfficientNet-B4 (224x224)</li>\n<li>EfficientNetV2-RW (224x224)</li>\n<li>EfficientNetV2-S (224x224)</li>\n<li>ConvNeXt-N (224x224, 320x320)</li>\n<li>ConvNeXt-T (224x224, 320x320)</li>\n<li>MaxViT-N (256x256)</li>\n</ul>\n<h4>Training</h4>\n<p>The basic training settings are as follows:</p>\n<ul>\n<li>10–20 epochs</li>\n<li>AdamW with learning rate <code>lr=0.000025</code>, OneCycleLR scheduler (Warmup for 3/10 steps of the total)</li>\n<li>Batch size: 2–8</li>\n<li>Cross-Entropy Weight: <code>[1.0, 2.0, 4.0]</code></li>\n<li><code>drop_path_rate</code> = 0.2 or 0.3</li>\n<li>Augmentations:<ul>\n<li><code>RandomBrightnessContrast</code></li>\n<li><code>Blur</code></li>\n<li><code>Distortion</code></li>\n<li><code>ShiftScaleRotate</code></li>\n<li><code>CoarseDropout</code></li>\n<li><code>Mixup</code> (Optional)</li></ul></li>\n</ul>\n<h3>7. Split LR</h3>\n<p>We designed preprocessing steps for training and inference of the Side Classifier. We split the sagittal and axial images into the left and right sides of the body. For the right-side data, we reversed the order of sagittal slices and horizontally flipped the axial images. This allowed us to handle the left and right sides uniformly and effectively doubled the amount of data available for training.</p>\n<h3>8. Side Classifier (2D-Encoder + Attention)</h3>\n<p>We created a classification model to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. The model structure and the number of input slices are identical to those of the Center Classifier.</p>\n<p>After verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels and arbitrary sides (left or right) to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis. By applying the Split LR preprocessing and flipping the right-side images to increase data, prediction accuracy improved. During training, we treated each group of slices as an independent data point without distinguishing disc levels or sides of the body. In other words, the model is designed to learn the characteristics of these conditions from the input slice groups, regardless of disc level or side, and estimate the severity based on those features. This allowed us to secure ten times the amount of data per condition, which we believe contributed to the improvement in accuracy.</p>\n<h2>Team Validation Strategy</h2>\n<ul>\n<li><strong>StratifiedKFold</strong><ul>\n<li><code>y</code>: Number of moderate or higher severity cases included in one study</li>\n<li><code>groups</code>: <code>study_id</code></li></ul></li>\n<li>Reference code: <a href=\"https://www.kaggle.com/code/artemtprv/lumbar-rsna-2024-eda-3d-visualization\" target=\"_blank\">Lumbar RSNA 2024 EDA + 3D Visualizationn</a></li>\n</ul>\n<h2>Pseudo Labeling</h2>\n<p>For items without ground truth labels, we used the predictions of the trained model as soft labels. The change in accuracy due to the presence or absence of pseudo-labels was not significant, but we introduced the use of pseudo-labels as an option to ensure model diversity.</p>\n<h2>Ensemble</h2>\n<p>The highest CV score for a single model was <strong>0.3858</strong>, achieved by combining the Center Classifier (Type B) ConvNeXt-T and the Side Classifier (Type B) ConvNeXt-N.</p>\n<p>The ensemble CV score was <strong>0.3643</strong>, obtained by simply averaging 30 models (15 Center models and 15 Side models) with diversity in input image types, model architectures, data augmentations, auxiliary losses, pseudo-labels, etc.</p>\n<h2>Post-processing</h2>\n<p>We applied Temperature Scaling with a temperature of <strong>0.91</strong> to the logits of Spinal Canal Stenosis, sharpening the predicted probabilities.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc92c6364e79e86a61586452ca13407a2%2Ftemperature_scaling.png?generation=1728437602634965&amp;alt=media\" alt=\"temperature_scaling\"></p>\n<h2>What Didn't Work</h2>\n<ul>\n<li>One-stage solution</li>\n<li>Multi-level Multi-disease models</li>\n<li>Multi-level Single-disease models</li>\n<li>Models specialized for each disc level</li>\n<li>Models specialized for each side of the body</li>\n<li>3D-CNN</li>\n<li>2.5D-CNN + Attention</li>\n<li>2D-CNN + LSTM</li>\n<li>Focal Loss</li>\n<li>Long epochs</li>\n</ul>\n<h2>Code (Updated on 2024-10-27)</h2>\n<p><a href=\"https://github.com/Moyasii/Kaggle-2024-RSNA-Pub\" target=\"_blank\">https://github.com/Moyasii/Kaggle-2024-RSNA-Pub</a></p>\n<h2>Video (Updated on 2024-10-28)</h2>\n<p><a href=\"https://www.youtube.com/watch?v=e2uRj5f9Lms&amp;ab_channel=sugupoko\" target=\"_blank\">https://www.youtube.com/watch?v=e2uRj5f9Lms&amp;ab_channel=sugupoko</a></p>",
  "messages": [
    {
      "id": "3012379",
      "postDate": "10/09/2024 01:47:44",
      "content": "<h1>3rd Place Solution</h1>\n<p>First and foremost, we would like to express our deepest gratitude to Kaggle and the competition organizers for providing this wonderful opportunity. We also thank all the participants for making this competition engaging and insightful.</p>\n<h2>Summary</h2>\n<p>We constructed a general two-stage pipeline:</p>\n<ul>\n<li><strong>Stage 1</strong>: Crop sagittal images at each disc level and crop axial images using disc level assignments and spinal canal positions.</li>\n<li><strong>Stage 2</strong>: Use a <strong>Center Classifier</strong> to classify the severity of Spinal Canal Stenosis and a <strong>Side Classifier</strong> to classify the severity of Neural Foraminal Narrowing and Subarticular Stenosis.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fed20bc73bd85e207bf670b752e038a32%2Fpipeline_overview.png?generation=1728438360443270&amp;alt=media\" alt=\"overview\"></p>\n<p>We will explain each process in detail below.</p>\n<h2>Stage 1</h2>\n<p>The responsibility of Stage 1 is to extract the information necessary for estimating disease severity from the input data.</p>\n<h3>1. Disc Level Keypoint Detector (CenterNet)</h3>\n<p>We built a CenterNet-based 2D keypoint detector using EfficientNetB6 as the backbone and FPN as the neck. By inputting sagittal images near the center of the body, we estimate the coordinates of each disc level. For training data, we used sagittal images from RSNA2024 that have coordinates of Spinal Canal Stenosis at all levels, as well as the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">Coordinate Pretraining Dataset</a>. By using the trained model to generate pseudo-labels on unused RSNA2024 data, we ultimately utilized all RSNA2024 data. Recognizing from several discussions that label noise existed, we manually reviewed all annotations and corrected erroneous labels by hand.</p>\n<h3>2. Crop Level</h3>\n<p>We cropped the sagittal images at each disc level. To ensure diversity in the input data, we adopted multiple cropping settings. There was almost no difference in accuracy due to cropping settings.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc7d43fb7b862bdf4c95f8fcc0f18883e%2Fsagittal.png?generation=1728437481708829&amp;alt=media\" alt=\"crop_sagittal\"></p>\n<h3>3. Assign Level</h3>\n<p>Using the output of the Disc Level Keypoint Detector, we assign arbitrary disc levels to the axial slices. The processing flow is as follows:</p>\n<ol>\n<li>Convert the image coordinates of disc levels to real-world coordinates.</li>\n<li>Estimate the vertebral positions of L1, L2, …, S1 from the midpoints of each disc level (e.g., L1/L2, L2/L3, etc.). Since we cannot obtain coordinates for T12/L1 and S1/S2, we pseudo-calculate the coordinates for L1 and S1.</li>\n<li>Calculate the intersection points between the line segments connecting adjacent vertebrae and the axial planes, and assign the corresponding disc levels.</li>\n</ol>\n<h3>4. Spinal Canal Keypoint Detector (CenterNet)</h3>\n<p>We constructed a CenterNet-based 2D keypoint detector using EfficientNetB4 as the backbone and FPN as the neck. By inputting axial images, we estimate the coordinates of the spinal canal. Since the Y-coordinate can be accurately estimated from the results of the Disc Level Keypoint Detector but estimating the X-coordinate is challenging, we introduced this detector. For training data, we used axial images from RSNA2024 that have coordinates of Spinal Canal Stenosis.</p>\n<h3>5. Crop Spinal</h3>\n<p>Using the output from the Spinal Canal Keypoint Detector, we crop the necessary regions centered on the spinal canal. To ensure diversity in the input data, we cropped at multiple sizes. There was almost no difference in accuracy due to cropping methods.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fa41a778fc592d7f632e5659178fcbcd6%2Faxial.png?generation=1728437501864432&amp;alt=media\" alt=\"crop_axial\"></p>\n<h2>Stage 2</h2>\n<p>The responsibility of Stage 2 is to estimate the severity of each condition using the outputs from Stage 1.</p>\n<h3>6. Center Classifier (2D-Encoder + Attention)</h3>\n<p>We created a classification model to estimate the severity of Spinal Canal Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. For sagittal T1 and sagittal T2/STIR, we input 15 slices at equal intervals into an encoder to generate feature representations for each slice. For axial T2, we input 10 slices at equal intervals. These slice features are then input into an attention mechanism to learn the relationships between slices. To ensure model diversity, we created two models with different head structures. To improve accuracy, we used auxiliary losses such as the severity of other conditions and slice-level predictions. Additionally, increasing the loss weight for the Severe class, due to the metric specifications, was effective. Test-time augmentation (TTA) by flipping axial images also improved the leaderboard score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2F2cd3d34e44fcde5d8f429c4b2eee042d%2Fclassifier.png?generation=1728438379066904&amp;alt=media\" alt=\"classifier\"></p>\n<p>After verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels to estimate the severity of Spinal Canal Stenosis. During training, we treated each group of slices as an independent data point without depending on the disc level. In other words, the model is designed to consistently learn the characteristics of Spinal Canal Stenosis from slice groups at any disc level and predict the severity based on those features, without specializing in any specific disc level. This allowed us to secure five times the amount of data per condition, which we believe contributed to the improvement in accuracy.</p>\n<p>We used the following encoders:</p>\n<ul>\n<li>ResNet18 (160x160, 224x224)</li>\n<li>MNasNet-S (224x224)</li>\n<li>EfficientNet-B4 (224x224)</li>\n<li>EfficientNetV2-RW (224x224)</li>\n<li>EfficientNetV2-S (224x224)</li>\n<li>ConvNeXt-N (224x224, 320x320)</li>\n<li>ConvNeXt-T (224x224, 320x320)</li>\n<li>MaxViT-N (256x256)</li>\n</ul>\n<h4>Training</h4>\n<p>The basic training settings are as follows:</p>\n<ul>\n<li>10–20 epochs</li>\n<li>AdamW with learning rate <code>lr=0.000025</code>, OneCycleLR scheduler (Warmup for 3/10 steps of the total)</li>\n<li>Batch size: 2–8</li>\n<li>Cross-Entropy Weight: <code>[1.0, 2.0, 4.0]</code></li>\n<li><code>drop_path_rate</code> = 0.2 or 0.3</li>\n<li>Augmentations:<ul>\n<li><code>RandomBrightnessContrast</code></li>\n<li><code>Blur</code></li>\n<li><code>Distortion</code></li>\n<li><code>ShiftScaleRotate</code></li>\n<li><code>CoarseDropout</code></li>\n<li><code>Mixup</code> (Optional)</li></ul></li>\n</ul>\n<h3>7. Split LR</h3>\n<p>We designed preprocessing steps for training and inference of the Side Classifier. We split the sagittal and axial images into the left and right sides of the body. For the right-side data, we reversed the order of sagittal slices and horizontally flipped the axial images. This allowed us to handle the left and right sides uniformly and effectively doubled the amount of data available for training.</p>\n<h3>8. Side Classifier (2D-Encoder + Attention)</h3>\n<p>We created a classification model to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. The model structure and the number of input slices are identical to those of the Center Classifier.</p>\n<p>After verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels and arbitrary sides (left or right) to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis. By applying the Split LR preprocessing and flipping the right-side images to increase data, prediction accuracy improved. During training, we treated each group of slices as an independent data point without distinguishing disc levels or sides of the body. In other words, the model is designed to learn the characteristics of these conditions from the input slice groups, regardless of disc level or side, and estimate the severity based on those features. This allowed us to secure ten times the amount of data per condition, which we believe contributed to the improvement in accuracy.</p>\n<h2>Team Validation Strategy</h2>\n<ul>\n<li><strong>StratifiedKFold</strong><ul>\n<li><code>y</code>: Number of moderate or higher severity cases included in one study</li>\n<li><code>groups</code>: <code>study_id</code></li></ul></li>\n<li>Reference code: <a href=\"https://www.kaggle.com/code/artemtprv/lumbar-rsna-2024-eda-3d-visualization\" target=\"_blank\">Lumbar RSNA 2024 EDA + 3D Visualizationn</a></li>\n</ul>\n<h2>Pseudo Labeling</h2>\n<p>For items without ground truth labels, we used the predictions of the trained model as soft labels. The change in accuracy due to the presence or absence of pseudo-labels was not significant, but we introduced the use of pseudo-labels as an option to ensure model diversity.</p>\n<h2>Ensemble</h2>\n<p>The highest CV score for a single model was <strong>0.3858</strong>, achieved by combining the Center Classifier (Type B) ConvNeXt-T and the Side Classifier (Type B) ConvNeXt-N.</p>\n<p>The ensemble CV score was <strong>0.3643</strong>, obtained by simply averaging 30 models (15 Center models and 15 Side models) with diversity in input image types, model architectures, data augmentations, auxiliary losses, pseudo-labels, etc.</p>\n<h2>Post-processing</h2>\n<p>We applied Temperature Scaling with a temperature of <strong>0.91</strong> to the logits of Spinal Canal Stenosis, sharpening the predicted probabilities.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc92c6364e79e86a61586452ca13407a2%2Ftemperature_scaling.png?generation=1728437602634965&amp;alt=media\" alt=\"temperature_scaling\"></p>\n<h2>What Didn't Work</h2>\n<ul>\n<li>One-stage solution</li>\n<li>Multi-level Multi-disease models</li>\n<li>Multi-level Single-disease models</li>\n<li>Models specialized for each disc level</li>\n<li>Models specialized for each side of the body</li>\n<li>3D-CNN</li>\n<li>2.5D-CNN + Attention</li>\n<li>2D-CNN + LSTM</li>\n<li>Focal Loss</li>\n<li>Long epochs</li>\n</ul>\n<h2>Code (Updated on 2024-10-27)</h2>\n<p><a href=\"https://github.com/Moyasii/Kaggle-2024-RSNA-Pub\" target=\"_blank\">https://github.com/Moyasii/Kaggle-2024-RSNA-Pub</a></p>\n<h2>Video (Updated on 2024-10-28)</h2>\n<p><a href=\"https://www.youtube.com/watch?v=e2uRj5f9Lms&amp;ab_channel=sugupoko\" target=\"_blank\">https://www.youtube.com/watch?v=e2uRj5f9Lms&amp;ab_channel=sugupoko</a></p>",
      "rawMarkdown": "# 3rd Place Solution\n\nFirst and foremost, we would like to express our deepest gratitude to Kaggle and the competition organizers for providing this wonderful opportunity. We also thank all the participants for making this competition engaging and insightful.\n\n## Summary\n\nWe constructed a general two-stage pipeline:\n\n- **Stage 1**: Crop sagittal images at each disc level and crop axial images using disc level assignments and spinal canal positions.\n- **Stage 2**: Use a **Center Classifier** to classify the severity of Spinal Canal Stenosis and a **Side Classifier** to classify the severity of Neural Foraminal Narrowing and Subarticular Stenosis.\n\n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fed20bc73bd85e207bf670b752e038a32%2Fpipeline_overview.png?generation=1728438360443270&alt=media)\n\nWe will explain each process in detail below.\n\n## Stage 1\n\nThe responsibility of Stage 1 is to extract the information necessary for estimating disease severity from the input data.\n\n### 1. Disc Level Keypoint Detector (CenterNet)\n\nWe built a CenterNet-based 2D keypoint detector using EfficientNetB6 as the backbone and FPN as the neck. By inputting sagittal images near the center of the body, we estimate the coordinates of each disc level. For training data, we used sagittal images from RSNA2024 that have coordinates of Spinal Canal Stenosis at all levels, as well as the [Coordinate Pretraining Dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500). By using the trained model to generate pseudo-labels on unused RSNA2024 data, we ultimately utilized all RSNA2024 data. Recognizing from several discussions that label noise existed, we manually reviewed all annotations and corrected erroneous labels by hand.\n\n### 2. Crop Level\n\nWe cropped the sagittal images at each disc level. To ensure diversity in the input data, we adopted multiple cropping settings. There was almost no difference in accuracy due to cropping settings.\n\n![crop_sagittal](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc7d43fb7b862bdf4c95f8fcc0f18883e%2Fsagittal.png?generation=1728437481708829&alt=media)\n\n### 3. Assign Level\n\nUsing the output of the Disc Level Keypoint Detector, we assign arbitrary disc levels to the axial slices. The processing flow is as follows:\n\n1. Convert the image coordinates of disc levels to real-world coordinates.\n2. Estimate the vertebral positions of L1, L2, ..., S1 from the midpoints of each disc level (e.g., L1/L2, L2/L3, etc.). Since we cannot obtain coordinates for T12/L1 and S1/S2, we pseudo-calculate the coordinates for L1 and S1.\n3. Calculate the intersection points between the line segments connecting adjacent vertebrae and the axial planes, and assign the corresponding disc levels.\n\n### 4. Spinal Canal Keypoint Detector (CenterNet)\n\nWe constructed a CenterNet-based 2D keypoint detector using EfficientNetB4 as the backbone and FPN as the neck. By inputting axial images, we estimate the coordinates of the spinal canal. Since the Y-coordinate can be accurately estimated from the results of the Disc Level Keypoint Detector but estimating the X-coordinate is challenging, we introduced this detector. For training data, we used axial images from RSNA2024 that have coordinates of Spinal Canal Stenosis.\n\n### 5. Crop Spinal\n\nUsing the output from the Spinal Canal Keypoint Detector, we crop the necessary regions centered on the spinal canal. To ensure diversity in the input data, we cropped at multiple sizes. There was almost no difference in accuracy due to cropping methods.\n\n![crop_axial](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fa41a778fc592d7f632e5659178fcbcd6%2Faxial.png?generation=1728437501864432&alt=media)\n\n## Stage 2\n\nThe responsibility of Stage 2 is to estimate the severity of each condition using the outputs from Stage 1.\n\n### 6. Center Classifier (2D-Encoder + Attention)\n\nWe created a classification model to estimate the severity of Spinal Canal Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. For sagittal T1 and sagittal T2/STIR, we input 15 slices at equal intervals into an encoder to generate feature representations for each slice. For axial T2, we input 10 slices at equal intervals. These slice features are then input into an attention mechanism to learn the relationships between slices. To ensure model diversity, we created two models with different head structures. To improve accuracy, we used auxiliary losses such as the severity of other conditions and slice-level predictions. Additionally, increasing the loss weight for the Severe class, due to the metric specifications, was effective. Test-time augmentation (TTA) by flipping axial images also improved the leaderboard score.\n\n![classifier](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2F2cd3d34e44fcde5d8f429c4b2eee042d%2Fclassifier.png?generation=1728438379066904&alt=media)\n\nAfter verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels to estimate the severity of Spinal Canal Stenosis. During training, we treated each group of slices as an independent data point without depending on the disc level. In other words, the model is designed to consistently learn the characteristics of Spinal Canal Stenosis from slice groups at any disc level and predict the severity based on those features, without specializing in any specific disc level. This allowed us to secure five times the amount of data per condition, which we believe contributed to the improvement in accuracy.\n\nWe used the following encoders:\n\n- ResNet18 (160x160, 224x224)\n- MNasNet-S (224x224)\n- EfficientNet-B4 (224x224)\n- EfficientNetV2-RW (224x224)\n- EfficientNetV2-S (224x224)\n- ConvNeXt-N (224x224, 320x320)\n- ConvNeXt-T (224x224, 320x320)\n- MaxViT-N (256x256)\n\n#### Training\n\nThe basic training settings are as follows:\n\n- 10–20 epochs\n- AdamW with learning rate `lr=0.000025`, OneCycleLR scheduler (Warmup for 3/10 steps of the total)\n- Batch size: 2–8\n- Cross-Entropy Weight: `[1.0, 2.0, 4.0]`\n- `drop_path_rate` = 0.2 or 0.3\n- Augmentations:\n  - `RandomBrightnessContrast`\n  - `Blur`\n  - `Distortion`\n  - `ShiftScaleRotate`\n  - `CoarseDropout`\n  - `Mixup` (Optional)\n\n### 7. Split LR\n\nWe designed preprocessing steps for training and inference of the Side Classifier. We split the sagittal and axial images into the left and right sides of the body. For the right-side data, we reversed the order of sagittal slices and horizontally flipped the axial images. This allowed us to handle the left and right sides uniformly and effectively doubled the amount of data available for training.\n\n### 8. Side Classifier (2D-Encoder + Attention)\n\nWe created a classification model to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. The model structure and the number of input slices are identical to those of the Center Classifier.\n\nAfter verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels and arbitrary sides (left or right) to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis. By applying the Split LR preprocessing and flipping the right-side images to increase data, prediction accuracy improved. During training, we treated each group of slices as an independent data point without distinguishing disc levels or sides of the body. In other words, the model is designed to learn the characteristics of these conditions from the input slice groups, regardless of disc level or side, and estimate the severity based on those features. This allowed us to secure ten times the amount of data per condition, which we believe contributed to the improvement in accuracy.\n\n## Team Validation Strategy\n\n- **StratifiedKFold**\n  - `y`: Number of moderate or higher severity cases included in one study\n  - `groups`: `study_id`\n- Reference code: [Lumbar RSNA 2024 EDA + 3D Visualizationn](https://www.kaggle.com/code/artemtprv/lumbar-rsna-2024-eda-3d-visualization)\n\n## Pseudo Labeling\n\nFor items without ground truth labels, we used the predictions of the trained model as soft labels. The change in accuracy due to the presence or absence of pseudo-labels was not significant, but we introduced the use of pseudo-labels as an option to ensure model diversity.\n\n## Ensemble\n\nThe highest CV score for a single model was **0.3858**, achieved by combining the Center Classifier (Type B) ConvNeXt-T and the Side Classifier (Type B) ConvNeXt-N.\n\nThe ensemble CV score was **0.3643**, obtained by simply averaging 30 models (15 Center models and 15 Side models) with diversity in input image types, model architectures, data augmentations, auxiliary losses, pseudo-labels, etc.\n\n## Post-processing\n\nWe applied Temperature Scaling with a temperature of **0.91** to the logits of Spinal Canal Stenosis, sharpening the predicted probabilities.\n\n![temperature_scaling](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc92c6364e79e86a61586452ca13407a2%2Ftemperature_scaling.png?generation=1728437602634965&alt=media)\n\n## What Didn't Work\n\n- One-stage solution\n- Multi-level Multi-disease models\n- Multi-level Single-disease models\n- Models specialized for each disc level\n- Models specialized for each side of the body\n- 3D-CNN\n- 2.5D-CNN + Attention\n- 2D-CNN + LSTM\n- Focal Loss\n- Long epochs\n\n## Code (Updated on 2024-10-27)\n\nhttps://github.com/Moyasii/Kaggle-2024-RSNA-Pub\n\n## Video (Updated on 2024-10-28)\n\nhttps://www.youtube.com/watch?v=e2uRj5f9Lms&ab_channel=sugupoko",
      "votes": null
    },
    {
      "id": "3012397",
      "postDate": "10/09/2024 02:21:41",
      "content": "<h1>Not in the Final Submission</h1>\n<ul>\n<li>We decided not to use the following external dataset <a href=\"https://tianchi.aliyun.com/dataset/79463\" target=\"_blank\">Spinal Disease Dataset</a> one week before the deadline. The reason was that we were unsure if the content of the \"Terms of Use\" complied with Kaggle's rules. Using this dataset showed an approximate performance improvement of +0.005 in a single model. </li>\n<li>The DICOM information were corrupted, making it difficult to determine which image it corresponded to.  </li>\n</ul>",
      "rawMarkdown": "# Not in the Final Submission\n- We decided not to use the following external dataset [Spinal Disease Dataset](https://tianchi.aliyun.com/dataset/79463) one week before the deadline. The reason was that we were unsure if the content of the \"Terms of Use\" complied with Kaggle's rules. Using this dataset showed an approximate performance improvement of +0.005 in a single model. \n- The DICOM information were corrupted, making it difficult to determine which image it corresponded to.",
      "votes": null
    },
    {
      "id": "3012532",
      "postDate": "10/09/2024 06:28:49",
      "content": "<p>Great explanations! Do you think that temperature scaling had a bigger impact because of the noise in the dataset or just because 🤔</p>",
      "rawMarkdown": "Great explanations! Do you think that temperature scaling had a bigger impact because of the noise in the dataset or just because 🤔",
      "votes": null
    },
    {
      "id": "3012618",
      "postDate": "10/09/2024 08:09:39",
      "content": "<p>Thank you for sharing your detailed solution. Congratulations, great work!</p>",
      "rawMarkdown": "Thank you for sharing your detailed solution. Congratulations, great work!",
      "votes": null
    },
    {
      "id": "3013054",
      "postDate": "10/09/2024 16:02:58",
      "content": "<p>Congratulations on winning the 3rd prize in this competition. Thanks for sharing the details of your solution with diagrams and images. </p>",
      "rawMarkdown": "Congratulations on winning the 3rd prize in this competition. Thanks for sharing the details of your solution with diagrams and images.",
      "votes": null
    },
    {
      "id": "3013384",
      "postDate": "10/10/2024 02:58:17",
      "content": "<p>Thanks for sharing and explaining your solution in depth, awesome work!</p>",
      "rawMarkdown": "Thanks for sharing and explaining your solution in depth, awesome work!",
      "votes": null
    },
    {
      "id": "3014608",
      "postDate": "10/11/2024 12:15:32",
      "content": "<p>Thanks for sharing and explaining your solution in depth, awesome work!</p>",
      "rawMarkdown": "Thanks for sharing and explaining your solution in depth, awesome work!",
      "votes": null
    },
    {
      "id": "3014935",
      "postDate": "10/11/2024 18:10:05",
      "content": "<p>Great explanations! </p>",
      "rawMarkdown": "Great explanations!",
      "votes": null
    },
    {
      "id": "3020103",
      "postDate": "10/17/2024 06:53:58",
      "content": "<p>Shocking! Thank you for this very detailed solution, which has made me realize the huge gap between me and the masters. I'm also looking forward to your code!!! I believe I can learn a lot from here.</p>",
      "rawMarkdown": "Shocking! Thank you for this very detailed solution, which has made me realize the huge gap between me and the masters. I'm also looking forward to your code!!! I believe I can learn a lot from here.",
      "votes": null
    },
    {
      "id": "3021195",
      "postDate": "10/18/2024 08:52:22",
      "content": "<p>Did you use DICOM format images, or did you switch to another format (like PNG)?\"</p>",
      "rawMarkdown": "Did you use DICOM format images, or did you switch to another format (like PNG)?\"",
      "votes": null
    },
    {
      "id": "3024953",
      "postDate": "10/22/2024 07:33:38",
      "content": "<p>Any updates on the code?</p>",
      "rawMarkdown": "Any updates on the code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3012397,
      "author_name": "sugupoko",
      "author_url": "",
      "post_date": "10/09/2024 02:21:41",
      "content": "<h1>Not in the Final Submission</h1>\n<ul>\n<li>We decided not to use the following external dataset <a href=\"https://tianchi.aliyun.com/dataset/79463\" target=\"_blank\">Spinal Disease Dataset</a> one week before the deadline. The reason was that we were unsure if the content of the \"Terms of Use\" complied with Kaggle's rules. Using this dataset showed an approximate performance improvement of +0.005 in a single model. </li>\n<li>The DICOM information were corrupted, making it difficult to determine which image it corresponded to.  </li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3012532,
      "author_name": "zshashz",
      "author_url": "",
      "post_date": "10/09/2024 06:28:49",
      "content": "<p>Great explanations! Do you think that temperature scaling had a bigger impact because of the noise in the dataset or just because 🤔</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3012618,
      "author_name": "marek3000",
      "author_url": "",
      "post_date": "10/09/2024 08:09:39",
      "content": "<p>Thank you for sharing your detailed solution. Congratulations, great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3013054,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "10/09/2024 16:02:58",
      "content": "<p>Congratulations on winning the 3rd prize in this competition. Thanks for sharing the details of your solution with diagrams and images. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3013384,
      "author_name": "ajpath",
      "author_url": "",
      "post_date": "10/10/2024 02:58:17",
      "content": "<p>Thanks for sharing and explaining your solution in depth, awesome work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3014608,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "10/11/2024 12:15:32",
      "content": "<p>Thanks for sharing and explaining your solution in depth, awesome work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3014935,
      "author_name": "humayrakhanomrime",
      "author_url": "",
      "post_date": "10/11/2024 18:10:05",
      "content": "<p>Great explanations! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3020103,
      "author_name": "switch9527",
      "author_url": "",
      "post_date": "10/17/2024 06:53:58",
      "content": "<p>Shocking! Thank you for this very detailed solution, which has made me realize the huge gap between me and the masters. I'm also looking forward to your code!!! I believe I can learn a lot from here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3021195,
      "author_name": "danielpopov",
      "author_url": "",
      "post_date": "10/18/2024 08:52:22",
      "content": "<p>Did you use DICOM format images, or did you switch to another format (like PNG)?\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3024953,
      "author_name": "lordpatil",
      "author_url": "",
      "post_date": "10/22/2024 07:33:38",
      "content": "<p>Any updates on the code?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3012379": "# 3rd Place Solution\n\nFirst and foremost, we would like to express our deepest gratitude to Kaggle and the competition organizers for providing this wonderful opportunity. We also thank all the participants for making this competition engaging and insightful.\n\n## Summary\n\nWe constructed a general two-stage pipeline:\n\n- **Stage 1**: Crop sagittal images at each disc level and crop axial images using disc level assignments and spinal canal positions.\n- **Stage 2**: Use a **Center Classifier** to classify the severity of Spinal Canal Stenosis and a **Side Classifier** to classify the severity of Neural Foraminal Narrowing and Subarticular Stenosis.\n\n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fed20bc73bd85e207bf670b752e038a32%2Fpipeline_overview.png?generation=1728438360443270&alt=media)\n\nWe will explain each process in detail below.\n\n## Stage 1\n\nThe responsibility of Stage 1 is to extract the information necessary for estimating disease severity from the input data.\n\n### 1. Disc Level Keypoint Detector (CenterNet)\n\nWe built a CenterNet-based 2D keypoint detector using EfficientNetB6 as the backbone and FPN as the neck. By inputting sagittal images near the center of the body, we estimate the coordinates of each disc level. For training data, we used sagittal images from RSNA2024 that have coordinates of Spinal Canal Stenosis at all levels, as well as the [Coordinate Pretraining Dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500). By using the trained model to generate pseudo-labels on unused RSNA2024 data, we ultimately utilized all RSNA2024 data. Recognizing from several discussions that label noise existed, we manually reviewed all annotations and corrected erroneous labels by hand.\n\n### 2. Crop Level\n\nWe cropped the sagittal images at each disc level. To ensure diversity in the input data, we adopted multiple cropping settings. There was almost no difference in accuracy due to cropping settings.\n\n![crop_sagittal](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc7d43fb7b862bdf4c95f8fcc0f18883e%2Fsagittal.png?generation=1728437481708829&alt=media)\n\n### 3. Assign Level\n\nUsing the output of the Disc Level Keypoint Detector, we assign arbitrary disc levels to the axial slices. The processing flow is as follows:\n\n1. Convert the image coordinates of disc levels to real-world coordinates.\n2. Estimate the vertebral positions of L1, L2, ..., S1 from the midpoints of each disc level (e.g., L1/L2, L2/L3, etc.). Since we cannot obtain coordinates for T12/L1 and S1/S2, we pseudo-calculate the coordinates for L1 and S1.\n3. Calculate the intersection points between the line segments connecting adjacent vertebrae and the axial planes, and assign the corresponding disc levels.\n\n### 4. Spinal Canal Keypoint Detector (CenterNet)\n\nWe constructed a CenterNet-based 2D keypoint detector using EfficientNetB4 as the backbone and FPN as the neck. By inputting axial images, we estimate the coordinates of the spinal canal. Since the Y-coordinate can be accurately estimated from the results of the Disc Level Keypoint Detector but estimating the X-coordinate is challenging, we introduced this detector. For training data, we used axial images from RSNA2024 that have coordinates of Spinal Canal Stenosis.\n\n### 5. Crop Spinal\n\nUsing the output from the Spinal Canal Keypoint Detector, we crop the necessary regions centered on the spinal canal. To ensure diversity in the input data, we cropped at multiple sizes. There was almost no difference in accuracy due to cropping methods.\n\n![crop_axial](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fa41a778fc592d7f632e5659178fcbcd6%2Faxial.png?generation=1728437501864432&alt=media)\n\n## Stage 2\n\nThe responsibility of Stage 2 is to estimate the severity of each condition using the outputs from Stage 1.\n\n### 6. Center Classifier (2D-Encoder + Attention)\n\nWe created a classification model to estimate the severity of Spinal Canal Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. For sagittal T1 and sagittal T2/STIR, we input 15 slices at equal intervals into an encoder to generate feature representations for each slice. For axial T2, we input 10 slices at equal intervals. These slice features are then input into an attention mechanism to learn the relationships between slices. To ensure model diversity, we created two models with different head structures. To improve accuracy, we used auxiliary losses such as the severity of other conditions and slice-level predictions. Additionally, increasing the loss weight for the Severe class, due to the metric specifications, was effective. Test-time augmentation (TTA) by flipping axial images also improved the leaderboard score.\n\n![classifier](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2F2cd3d34e44fcde5d8f429c4b2eee042d%2Fclassifier.png?generation=1728438379066904&alt=media)\n\nAfter verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels to estimate the severity of Spinal Canal Stenosis. During training, we treated each group of slices as an independent data point without depending on the disc level. In other words, the model is designed to consistently learn the characteristics of Spinal Canal Stenosis from slice groups at any disc level and predict the severity based on those features, without specializing in any specific disc level. This allowed us to secure five times the amount of data per condition, which we believe contributed to the improvement in accuracy.\n\nWe used the following encoders:\n\n- ResNet18 (160x160, 224x224)\n- MNasNet-S (224x224)\n- EfficientNet-B4 (224x224)\n- EfficientNetV2-RW (224x224)\n- EfficientNetV2-S (224x224)\n- ConvNeXt-N (224x224, 320x320)\n- ConvNeXt-T (224x224, 320x320)\n- MaxViT-N (256x256)\n\n#### Training\n\nThe basic training settings are as follows:\n\n- 10–20 epochs\n- AdamW with learning rate `lr=0.000025`, OneCycleLR scheduler (Warmup for 3/10 steps of the total)\n- Batch size: 2–8\n- Cross-Entropy Weight: `[1.0, 2.0, 4.0]`\n- `drop_path_rate` = 0.2 or 0.3\n- Augmentations:\n  - `RandomBrightnessContrast`\n  - `Blur`\n  - `Distortion`\n  - `ShiftScaleRotate`\n  - `CoarseDropout`\n  - `Mixup` (Optional)\n\n### 7. Split LR\n\nWe designed preprocessing steps for training and inference of the Side Classifier. We split the sagittal and axial images into the left and right sides of the body. For the right-side data, we reversed the order of sagittal slices and horizontally flipped the axial images. This allowed us to handle the left and right sides uniformly and effectively doubled the amount of data available for training.\n\n### 8. Side Classifier (2D-Encoder + Attention)\n\nWe created a classification model to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis from sagittal T1, sagittal T2/STIR, and axial T2 images. The model structure and the number of input slices are identical to those of the Center Classifier.\n\nAfter verifying various patterns of input-output combinations for the model, we found that the most effective approach was to input groups of sagittal T1, sagittal T2, and axial T2 slices at arbitrary disc levels and arbitrary sides (left or right) to estimate the severity of Neural Foraminal Narrowing and Subarticular Stenosis. By applying the Split LR preprocessing and flipping the right-side images to increase data, prediction accuracy improved. During training, we treated each group of slices as an independent data point without distinguishing disc levels or sides of the body. In other words, the model is designed to learn the characteristics of these conditions from the input slice groups, regardless of disc level or side, and estimate the severity based on those features. This allowed us to secure ten times the amount of data per condition, which we believe contributed to the improvement in accuracy.\n\n## Team Validation Strategy\n\n- **StratifiedKFold**\n  - `y`: Number of moderate or higher severity cases included in one study\n  - `groups`: `study_id`\n- Reference code: [Lumbar RSNA 2024 EDA + 3D Visualizationn](https://www.kaggle.com/code/artemtprv/lumbar-rsna-2024-eda-3d-visualization)\n\n## Pseudo Labeling\n\nFor items without ground truth labels, we used the predictions of the trained model as soft labels. The change in accuracy due to the presence or absence of pseudo-labels was not significant, but we introduced the use of pseudo-labels as an option to ensure model diversity.\n\n## Ensemble\n\nThe highest CV score for a single model was **0.3858**, achieved by combining the Center Classifier (Type B) ConvNeXt-T and the Side Classifier (Type B) ConvNeXt-N.\n\nThe ensemble CV score was **0.3643**, obtained by simply averaging 30 models (15 Center models and 15 Side models) with diversity in input image types, model architectures, data augmentations, auxiliary losses, pseudo-labels, etc.\n\n## Post-processing\n\nWe applied Temperature Scaling with a temperature of **0.91** to the logits of Spinal Canal Stenosis, sharpening the predicted probabilities.\n\n![temperature_scaling](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6560839%2Fc92c6364e79e86a61586452ca13407a2%2Ftemperature_scaling.png?generation=1728437602634965&alt=media)\n\n## What Didn't Work\n\n- One-stage solution\n- Multi-level Multi-disease models\n- Multi-level Single-disease models\n- Models specialized for each disc level\n- Models specialized for each side of the body\n- 3D-CNN\n- 2.5D-CNN + Attention\n- 2D-CNN + LSTM\n- Focal Loss\n- Long epochs\n\n## Code (Updated on 2024-10-27)\n\nhttps://github.com/Moyasii/Kaggle-2024-RSNA-Pub\n\n## Video (Updated on 2024-10-28)\n\nhttps://www.youtube.com/watch?v=e2uRj5f9Lms&ab_channel=sugupoko",
    "3012397": "# Not in the Final Submission\n- We decided not to use the following external dataset [Spinal Disease Dataset](https://tianchi.aliyun.com/dataset/79463) one week before the deadline. The reason was that we were unsure if the content of the \"Terms of Use\" complied with Kaggle's rules. Using this dataset showed an approximate performance improvement of +0.005 in a single model. \n- The DICOM information were corrupted, making it difficult to determine which image it corresponded to.",
    "3012532": "Great explanations! Do you think that temperature scaling had a bigger impact because of the noise in the dataset or just because 🤔",
    "3012618": "Thank you for sharing your detailed solution. Congratulations, great work!",
    "3013054": "Congratulations on winning the 3rd prize in this competition. Thanks for sharing the details of your solution with diagrams and images.",
    "3013384": "Thanks for sharing and explaining your solution in depth, awesome work!",
    "3014608": "Thanks for sharing and explaining your solution in depth, awesome work!",
    "3014935": "Great explanations!",
    "3020103": "Shocking! Thank you for this very detailed solution, which has made me realize the huge gap between me and the masters. I'm also looking forward to your code!!! I believe I can learn a lot from here.",
    "3021195": "Did you use DICOM format images, or did you switch to another format (like PNG)?\"",
    "3024953": "Any updates on the code?"
  },
  "source": "meta"
}