{
  "id": 541279,
  "title": "Summary of Top Team Solutions",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/541279",
  "author_name": "",
  "post_date": "2024-10-18T14:09:28.461821300Z",
  "votes": 29,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to all who participated in the competition.<br>\nI learned a lot from the solutions that were made public after the competition. <br>\nI read through the top solutions and compiled my own insights, so I’d like to share them here. If any of these solutions catch your interest, I’ve included links to the original posts, so please check them out and upvote!<br>\nOnce again, I’d like to express my gratitude to all the participants who shared their solutions.</p>\n<p>Note:<br>\nPlease forgive me if my interpretation may be incorrect. If there are any errors, I would appreciate it if you could let me know in the comments or otherwise.</p>\n<h1>Summary of Top Team Solutions</h1>\n<h2>1. Approach</h2>\n<ul>\n<li><p>Almost all teams adopted a two-stage pipeline (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, etc…)</p>\n<ul>\n<li>In the 1st stage, they estimated slice indices, keypoints, etc., and cropped the regions of interest at each level based on these estimations.</li>\n<li>In the 2nd stage, they used the cropped regions as inputs for label classification models to predict severity scores for each level.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fb7834266579f475e9e5db6d4ace15f41%2F2stage_pipeline_image.jpg?generation=1729260379466636&amp;alt=media\" alt=\"\"></li></ul></li>\n<li><p>By cropping the regions at each level and treating them equally, they effectively increased the data size by a factor of 5 (for NFN and SS, treating the left and right sides equally resulted in a 5x2=10x increase). This seems to have been effective in a competition where the amount of data was somewhat limited.</p></li>\n<li><p>On the other hand, a few teams adopted a single-stage pipeline (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439\" target=\"_blank\">7th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548\" target=\"_blank\">8th</a>). <br>\nNote: This means they didn’t crop the regions, so the term \"single-stage\" might be a bit misleading.</p>\n<ul>\n<li>The single-stage pipeline had the advantage of avoiding errors that occur during the region cropping estimation and being able to consider the overall context of the image. Although it is more challenging, by modeling it well, they were able to achieve accuracy on par with or better than the two-stage pipeline.</li>\n<li>For example, the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548\" target=\"_blank\">8th</a> team trained keypoint detection (heatmap prediction) as a subtask, and from the weighted feature map obtained by the heatmap, they classified the classes. Without cropping, they managed to guide the model’s attention to the regions of interest effectively. (Note: There might be a misunderstanding as I haven’t fully decoded the source code.)</li></ul></li>\n</ul>\n<h2>2. Preprocessing (1st Stage Processing)</h2>\n<h4>Slice Index Estimation</h4>\n<ul>\n<li><p>Many top teams implemented the process of estimating slice indices (instance_number) suitable for predicting severity scores from a group of MRI slices using either \"a trained slice estimation model\" or \"rule-based\" methods.</p></li>\n<li><p>The approach of training a slice estimation model involved using 2D/3D models to output the optimal slice index from the MRI slice volume and estimate the best slice index for each level (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, etc…).</p>\n<ul>\n<li>Since some patients have curved spines, causing the optimal slice index to vary by level, this approach was advantageous in addressing such cases.</li>\n<li>Tasks for this approach included predicting the relative position of the slice index (regression task), predicting the distance of each slice from the ground truth slice (regression task), or classifying whether each slice is the ground truth (classification task).</li></ul></li>\n<li><p>The rule-based approach often involved selecting the central slice for SCS and slices at a certain distance from the center for NFN as the optimal slice index (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690\" target=\"_blank\">9th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>, etc…).</p></li>\n<li><p>Meanwhile, there were also teams that did not estimate slice indices at all (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012395\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, etc…).</p>\n<ul>\n<li>These teams seemed to implicitly resolve this by inputting multiple slices into the label classification model in the 2nd stage. By training the model with a sufficient number of slices, it is assumed that the model implicitly learned to select the slices most useful for prediction.</li></ul></li>\n</ul>\n<h4>Keypoint Estimation</h4>\n<ul>\n<li>A relatively large number of teams used 2D/2.5D regression models to directly estimate x and y relative coordinates. However, some teams estimated keypoints by outputting heatmaps through a segmentation model (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690\" target=\"_blank\">9th</a>, etc…) or by using keypoint detectors like CenterNet (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>).</li>\n<li>Additionally, some teams did not prepare an independent keypoint estimation model but instead trained a model that simultaneously outputs both slice indices and keypoints (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539472\" target=\"_blank\">5th</a>).</li>\n</ul>\n<h4>Axial Slice Level Correspondence</h4>\n<ul>\n<li>Axial slices correspond to specific levels, but since this correspondence is not explicitly provided, teams needed to estimate the relationships in some way.</li>\n<li>Due to the availability of <a href=\"(https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approach?scriptVersionId=191553260\" target=\"_blank\">helpful notebooks</a>), most teams aligned Axial slices with Sagittal slices using DICOM metadata and then estimated the corresponding level of each Axial slice based on the keypoints predicted from the Sagittal slices.</li>\n<li>An alternative approach was used by the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a> team, which utilized the alignment information as supplementary data while also training a model to predict which level each Axial slice belonged to.</li>\n</ul>\n<h4>Level Region Cropping</h4>\n<ul>\n<li><p>Most teams determined the cropping range for each level based on the coordinates obtained from keypoint estimation.</p>\n<ul>\n<li>Many teams used fixed values to set the cropping size based on the estimated coordinates, but some teams dynamically determined the size using the distance between keypoints of adjacent levels (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>).</li></ul></li>\n<li><p>There were also teams that did not use keypoints, but instead employed object detection models like YOLO or SpineNetV2 to directly output the regions of interest (ROIs) (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587\" target=\"_blank\">10th</a>).</p></li>\n</ul>\n<h2>3. Classification Process (2nd Stage Processing)</h2>\n<h4>Architecture</h4>\n<ul>\n<li><p>A common point among top teams was the use of Multiple Instance Learning (MIL) models that take multiple slices as input (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, etc…).</p>\n<ul>\n<li>While there were some differences in the details, many teams seemed to use a model architecture along the following lines:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fdba39e3ac9826122837181c473d6b9be%2FMIL.jpg?generation=1729260429260887&amp;alt=media\" alt=\"\"></li></ul>\n<ol>\n<li>Extract feature vectors for each slice using a 2D backbone model.</li>\n<li>Pass the feature vectors of each slice through an LSTM/GRU to capture temporal information.</li>\n<li>Aggregate the feature vectors of the slices using Attention Pooling or Global Average Pooling (GAP).</li>\n<li>Input the aggregated feature vector into the final head to output the classification result.</li></ol>\n<ul>\n<li>The choice of 2D backbone models varied among teams (with ConvNeXt and EfficientNetv2 being somewhat popular), but they all commonly used smaller models.</li>\n<li>For feature vector aggregation, many teams used Attention Pooling.</li></ul></li>\n<li><p>As for the number of input slices, teams that performed slice index estimation typically used around 3 to 5 slices centered on the estimated index, while teams that did not estimate slice indices used around 10 to 30 slices at equal intervals (padding with dummy slices if necessary).</p></li>\n<li><p>Some teams also employed 2.5D models, which treated consecutive slices stacked along the channel dimension like normal images (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>, etc…).</p></li>\n<li><p>Due to space limitations, I won’t go into details about individual architectures, but each team had unique modeling approaches. If you're interested, I recommend checking out the discussions for more in-depth information.</p></li>\n</ul>\n<h4>Multi-view Input</h4>\n<ul>\n<li><p>Generally, teams followed the annotations, using Sagittal T2 for SCS prediction, Sagittal T1 for NFN prediction, and Axial for SS prediction. However, some teams adopted multi-modal architectures that input images from multiple views.</p></li>\n<li><p>There were two main approaches for implementing multi-view input models. One approach was to concatenate the feature vectors just before the final head, while the other was to combine the feature vectors at the 2D backbone stage.</p>\n<ul>\n<li><p>The first approach involved processing each view independently using the same steps as in a single-view model (2D backbone -&gt; LSTM/GPU -&gt; aggregate) and concatenating the feature vectors just before the final head to merge them into a single vector (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690\" target=\"_blank\">9th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587\" target=\"_blank\">10th</a>, etc…).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F7c8f974b7505b96c98c7b36c1683e5f0%2Fmulti_view1.jpg?generation=1729260461167251&amp;alt=media\" alt=\"\"></p></li>\n<li><p>The second approach involved inputting the feature vectors from all slices of all views into a Transformer Encoder right after the 2D backbone stage, allowing it to learn relationships between the feature vectors. These were then aggregated into a single vector using Attention Pooling or GAP (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, etc…).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F1b0036f5634c2fae35d141e7a378d56a%2Fmulti_view2.jpg?generation=1729260482124477&amp;alt=media\" alt=\"\"></p></li>\n<li><p>There didn't seem to be a clearly superior method in this competition, but the insights gained from effectively fusing features from different sources could be useful for handling similar tasks in the future.</p></li></ul></li>\n</ul>\n<h4>Augmentation</h4>\n<ul>\n<li>To make the 2nd stage models more robust against errors introduced by slice index estimation or keypoint estimation in the 1st stage, random shifting of keypoint coordinates or slice indices during training was an effective augmentation strategy (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540001\" target=\"_blank\">12th</a>, etc…).</li>\n<li>The <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a> place team used 27 cropping patterns during training ([original slice index, ±1 shifted slice index] * [original keypoint, keypoints shifted in one of 8 directions] = 3 * 9 = 27). During inference, they applied a similar technique by generating predictions for these 27 patterns and averaging the results, similar to Test-Time Augmentation (TTA), which greatly improved performance.</li>\n</ul>\n<h4>Noisy Data Removal</h4>\n<ul>\n<li><p>The annotations provided in this competition were noisy (e.g., samples clearly labeled as \"normal\" were given \"severe\" labels), and some teams improved their scores by removing such data (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>).</p></li>\n<li><p>To detect noisy data, the 2nd team used ensemble out-of-fold (oof) predictions and excluded samples where the difference between the prediction and ground truth label was greater than 0.8, treating them as noise.</p></li>\n<li><p>However, when I tried a similar approach in a late submission, the score consistently dropped. I couldn’t replicate the improvement in my environment.</p>\n<ul>\n<li>The possible reasons could be issues with my reproduction process or insufficient performance of the model used for noise detection. In any case, I realized that this method might backfire under certain conditions, so careful experimentation and validation are necessary when applying it in future competitions.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "3021441",
      "postDate": "10/18/2024 14:09:28",
      "content": "<p>Thanks to all who participated in the competition.<br>\nI learned a lot from the solutions that were made public after the competition. <br>\nI read through the top solutions and compiled my own insights, so I’d like to share them here. If any of these solutions catch your interest, I’ve included links to the original posts, so please check them out and upvote!<br>\nOnce again, I’d like to express my gratitude to all the participants who shared their solutions.</p>\n<p>Note:<br>\nPlease forgive me if my interpretation may be incorrect. If there are any errors, I would appreciate it if you could let me know in the comments or otherwise.</p>\n<h1>Summary of Top Team Solutions</h1>\n<h2>1. Approach</h2>\n<ul>\n<li><p>Almost all teams adopted a two-stage pipeline (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, etc…)</p>\n<ul>\n<li>In the 1st stage, they estimated slice indices, keypoints, etc., and cropped the regions of interest at each level based on these estimations.</li>\n<li>In the 2nd stage, they used the cropped regions as inputs for label classification models to predict severity scores for each level.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fb7834266579f475e9e5db6d4ace15f41%2F2stage_pipeline_image.jpg?generation=1729260379466636&amp;alt=media\" alt=\"\"></li></ul></li>\n<li><p>By cropping the regions at each level and treating them equally, they effectively increased the data size by a factor of 5 (for NFN and SS, treating the left and right sides equally resulted in a 5x2=10x increase). This seems to have been effective in a competition where the amount of data was somewhat limited.</p></li>\n<li><p>On the other hand, a few teams adopted a single-stage pipeline (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439\" target=\"_blank\">7th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548\" target=\"_blank\">8th</a>). <br>\nNote: This means they didn’t crop the regions, so the term \"single-stage\" might be a bit misleading.</p>\n<ul>\n<li>The single-stage pipeline had the advantage of avoiding errors that occur during the region cropping estimation and being able to consider the overall context of the image. Although it is more challenging, by modeling it well, they were able to achieve accuracy on par with or better than the two-stage pipeline.</li>\n<li>For example, the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548\" target=\"_blank\">8th</a> team trained keypoint detection (heatmap prediction) as a subtask, and from the weighted feature map obtained by the heatmap, they classified the classes. Without cropping, they managed to guide the model’s attention to the regions of interest effectively. (Note: There might be a misunderstanding as I haven’t fully decoded the source code.)</li></ul></li>\n</ul>\n<h2>2. Preprocessing (1st Stage Processing)</h2>\n<h4>Slice Index Estimation</h4>\n<ul>\n<li><p>Many top teams implemented the process of estimating slice indices (instance_number) suitable for predicting severity scores from a group of MRI slices using either \"a trained slice estimation model\" or \"rule-based\" methods.</p></li>\n<li><p>The approach of training a slice estimation model involved using 2D/3D models to output the optimal slice index from the MRI slice volume and estimate the best slice index for each level (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, etc…).</p>\n<ul>\n<li>Since some patients have curved spines, causing the optimal slice index to vary by level, this approach was advantageous in addressing such cases.</li>\n<li>Tasks for this approach included predicting the relative position of the slice index (regression task), predicting the distance of each slice from the ground truth slice (regression task), or classifying whether each slice is the ground truth (classification task).</li></ul></li>\n<li><p>The rule-based approach often involved selecting the central slice for SCS and slices at a certain distance from the center for NFN as the optimal slice index (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690\" target=\"_blank\">9th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>, etc…).</p></li>\n<li><p>Meanwhile, there were also teams that did not estimate slice indices at all (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012395\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, etc…).</p>\n<ul>\n<li>These teams seemed to implicitly resolve this by inputting multiple slices into the label classification model in the 2nd stage. By training the model with a sufficient number of slices, it is assumed that the model implicitly learned to select the slices most useful for prediction.</li></ul></li>\n</ul>\n<h4>Keypoint Estimation</h4>\n<ul>\n<li>A relatively large number of teams used 2D/2.5D regression models to directly estimate x and y relative coordinates. However, some teams estimated keypoints by outputting heatmaps through a segmentation model (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690\" target=\"_blank\">9th</a>, etc…) or by using keypoint detectors like CenterNet (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>).</li>\n<li>Additionally, some teams did not prepare an independent keypoint estimation model but instead trained a model that simultaneously outputs both slice indices and keypoints (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539472\" target=\"_blank\">5th</a>).</li>\n</ul>\n<h4>Axial Slice Level Correspondence</h4>\n<ul>\n<li>Axial slices correspond to specific levels, but since this correspondence is not explicitly provided, teams needed to estimate the relationships in some way.</li>\n<li>Due to the availability of <a href=\"(https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approach?scriptVersionId=191553260\" target=\"_blank\">helpful notebooks</a>), most teams aligned Axial slices with Sagittal slices using DICOM metadata and then estimated the corresponding level of each Axial slice based on the keypoints predicted from the Sagittal slices.</li>\n<li>An alternative approach was used by the <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a> team, which utilized the alignment information as supplementary data while also training a model to predict which level each Axial slice belonged to.</li>\n</ul>\n<h4>Level Region Cropping</h4>\n<ul>\n<li><p>Most teams determined the cropping range for each level based on the coordinates obtained from keypoint estimation.</p>\n<ul>\n<li>Many teams used fixed values to set the cropping size based on the estimated coordinates, but some teams dynamically determined the size using the distance between keypoints of adjacent levels (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>).</li></ul></li>\n<li><p>There were also teams that did not use keypoints, but instead employed object detection models like YOLO or SpineNetV2 to directly output the regions of interest (ROIs) (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587\" target=\"_blank\">10th</a>).</p></li>\n</ul>\n<h2>3. Classification Process (2nd Stage Processing)</h2>\n<h4>Architecture</h4>\n<ul>\n<li><p>A common point among top teams was the use of Multiple Instance Learning (MIL) models that take multiple slices as input (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, etc…).</p>\n<ul>\n<li>While there were some differences in the details, many teams seemed to use a model architecture along the following lines:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fdba39e3ac9826122837181c473d6b9be%2FMIL.jpg?generation=1729260429260887&amp;alt=media\" alt=\"\"></li></ul>\n<ol>\n<li>Extract feature vectors for each slice using a 2D backbone model.</li>\n<li>Pass the feature vectors of each slice through an LSTM/GRU to capture temporal information.</li>\n<li>Aggregate the feature vectors of the slices using Attention Pooling or Global Average Pooling (GAP).</li>\n<li>Input the aggregated feature vector into the final head to output the classification result.</li></ol>\n<ul>\n<li>The choice of 2D backbone models varied among teams (with ConvNeXt and EfficientNetv2 being somewhat popular), but they all commonly used smaller models.</li>\n<li>For feature vector aggregation, many teams used Attention Pooling.</li></ul></li>\n<li><p>As for the number of input slices, teams that performed slice index estimation typically used around 3 to 5 slices centered on the estimated index, while teams that did not estimate slice indices used around 10 to 30 slices at equal intervals (padding with dummy slices if necessary).</p></li>\n<li><p>Some teams also employed 2.5D models, which treated consecutive slices stacked along the channel dimension like normal images (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>, etc…).</p></li>\n<li><p>Due to space limitations, I won’t go into details about individual architectures, but each team had unique modeling approaches. If you're interested, I recommend checking out the discussions for more in-depth information.</p></li>\n</ul>\n<h4>Multi-view Input</h4>\n<ul>\n<li><p>Generally, teams followed the annotations, using Sagittal T2 for SCS prediction, Sagittal T1 for NFN prediction, and Axial for SS prediction. However, some teams adopted multi-modal architectures that input images from multiple views.</p></li>\n<li><p>There were two main approaches for implementing multi-view input models. One approach was to concatenate the feature vectors just before the final head, while the other was to combine the feature vectors at the 2D backbone stage.</p>\n<ul>\n<li><p>The first approach involved processing each view independently using the same steps as in a single-view model (2D backbone -&gt; LSTM/GPU -&gt; aggregate) and concatenating the feature vectors just before the final head to merge them into a single vector (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091\" target=\"_blank\">1st</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690\" target=\"_blank\">9th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587\" target=\"_blank\">10th</a>, etc…).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F7c8f974b7505b96c98c7b36c1683e5f0%2Fmulti_view1.jpg?generation=1729260461167251&amp;alt=media\" alt=\"\"></p></li>\n<li><p>The second approach involved inputting the feature vectors from all slices of all views into a Transformer Encoder right after the 2D backbone stage, allowing it to learn relationships between the feature vectors. These were then aggregated into a single vector using Attention Pooling or GAP (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453\" target=\"_blank\">3rd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443\" target=\"_blank\">4th</a>, etc…).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F1b0036f5634c2fae35d141e7a378d56a%2Fmulti_view2.jpg?generation=1729260482124477&amp;alt=media\" alt=\"\"></p></li>\n<li><p>There didn't seem to be a clearly superior method in this competition, but the insights gained from effectively fusing features from different sources could be useful for handling similar tasks in the future.</p></li></ul></li>\n</ul>\n<h4>Augmentation</h4>\n<ul>\n<li>To make the 2nd stage models more robust against errors introduced by slice index estimation or keypoint estimation in the 1st stage, random shifting of keypoint coordinates or slice indices during training was an effective augmentation strategy (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569\" target=\"_blank\">11th</a>, <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540001\" target=\"_blank\">12th</a>, etc…).</li>\n<li>The <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803\" target=\"_blank\">2nd</a> place team used 27 cropping patterns during training ([original slice index, ±1 shifted slice index] * [original keypoint, keypoints shifted in one of 8 directions] = 3 * 9 = 27). During inference, they applied a similar technique by generating predictions for these 27 patterns and averaging the results, similar to Test-Time Augmentation (TTA), which greatly improved performance.</li>\n</ul>\n<h4>Noisy Data Removal</h4>\n<ul>\n<li><p>The annotations provided in this competition were noisy (e.g., samples clearly labeled as \"normal\" were given \"severe\" labels), and some teams improved their scores by removing such data (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452\" target=\"_blank\">2nd</a>).</p></li>\n<li><p>To detect noisy data, the 2nd team used ensemble out-of-fold (oof) predictions and excluded samples where the difference between the prediction and ground truth label was greater than 0.8, treating them as noise.</p></li>\n<li><p>However, when I tried a similar approach in a late submission, the score consistently dropped. I couldn’t replicate the improvement in my environment.</p>\n<ul>\n<li>The possible reasons could be issues with my reproduction process or insufficient performance of the model used for noise detection. In any case, I realized that this method might backfire under certain conditions, so careful experimentation and validation are necessary when applying it in future competitions.</li></ul></li>\n</ul>",
      "rawMarkdown": "Thanks to all who participated in the competition.\nI learned a lot from the solutions that were made public after the competition. \nI read through the top solutions and compiled my own insights, so I’d like to share them here. If any of these solutions catch your interest, I’ve included links to the original posts, so please check them out and upvote!\nOnce again, I’d like to express my gratitude to all the participants who shared their solutions.\n\nNote:\nPlease forgive me if my interpretation may be incorrect. If there are any errors, I would appreciate it if you could let me know in the comments or otherwise.\n\n# Summary of Top Team Solutions\n## 1. Approach\n- Almost all teams adopted a two-stage pipeline ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), etc...)\n  - In the 1st stage, they estimated slice indices, keypoints, etc., and cropped the regions of interest at each level based on these estimations.\n  - In the 2nd stage, they used the cropped regions as inputs for label classification models to predict severity scores for each level.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fb7834266579f475e9e5db6d4ace15f41%2F2stage_pipeline_image.jpg?generation=1729260379466636&alt=media)\n\n- By cropping the regions at each level and treating them equally, they effectively increased the data size by a factor of 5 (for NFN and SS, treating the left and right sides equally resulted in a 5x2=10x increase). This seems to have been effective in a competition where the amount of data was somewhat limited.\n\n- On the other hand, a few teams adopted a single-stage pipeline ([7th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439), [8th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548)). \nNote: This means they didn’t crop the regions, so the term \"single-stage\" might be a bit misleading.\n  - The single-stage pipeline had the advantage of avoiding errors that occur during the region cropping estimation and being able to consider the overall context of the image. Although it is more challenging, by modeling it well, they were able to achieve accuracy on par with or better than the two-stage pipeline.\n  - For example, the [8th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548) team trained keypoint detection (heatmap prediction) as a subtask, and from the weighted feature map obtained by the heatmap, they classified the classes. Without cropping, they managed to guide the model’s attention to the regions of interest effectively. (Note: There might be a misunderstanding as I haven’t fully decoded the source code.)\n\n## 2. Preprocessing (1st Stage Processing)\n#### Slice Index Estimation\n- Many top teams implemented the process of estimating slice indices (instance_number) suitable for predicting severity scores from a group of MRI slices using either \"a trained slice estimation model\" or \"rule-based\" methods.\n\n- The approach of training a slice estimation model involved using 2D/3D models to output the optimal slice index from the MRI slice volume and estimate the best slice index for each level ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), etc...).\n  - Since some patients have curved spines, causing the optimal slice index to vary by level, this approach was advantageous in addressing such cases.\n  - Tasks for this approach included predicting the relative position of the slice index (regression task), predicting the distance of each slice from the ground truth slice (regression task), or classifying whether each slice is the ground truth (classification task).\n\n- The rule-based approach often involved selecting the central slice for SCS and slices at a certain distance from the center for NFN as the optimal slice index ([9th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569), etc...).\n\n- Meanwhile, there were also teams that did not estimate slice indices at all ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012395), [3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), [4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), etc...).\n  - These teams seemed to implicitly resolve this by inputting multiple slices into the label classification model in the 2nd stage. By training the model with a sufficient number of slices, it is assumed that the model implicitly learned to select the slices most useful for prediction.\n\n\n#### Keypoint Estimation\n- A relatively large number of teams used 2D/2.5D regression models to directly estimate x and y relative coordinates. However, some teams estimated keypoints by outputting heatmaps through a segmentation model ([4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), [9th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690), etc...) or by using keypoint detectors like CenterNet ([3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453)).\n- Additionally, some teams did not prepare an independent keypoint estimation model but instead trained a model that simultaneously outputs both slice indices and keypoints ([5th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539472)).\n\n#### Axial Slice Level Correspondence\n- Axial slices correspond to specific levels, but since this correspondence is not explicitly provided, teams needed to estimate the relationships in some way.\n- Due to the availability of [helpful notebooks]((https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approach?scriptVersionId=191553260)), most teams aligned Axial slices with Sagittal slices using DICOM metadata and then estimated the corresponding level of each Axial slice based on the keypoints predicted from the Sagittal slices.\n- An alternative approach was used by the [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803) team, which utilized the alignment information as supplementary data while also training a model to predict which level each Axial slice belonged to.\n\n#### Level Region Cropping\n- Most teams determined the cropping range for each level based on the coordinates obtained from keypoint estimation.\n  - Many teams used fixed values to set the cropping size based on the estimated coordinates, but some teams dynamically determined the size using the distance between keypoints of adjacent levels ([4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569)).\n \n- There were also teams that did not use keypoints, but instead employed object detection models like YOLO or SpineNetV2 to directly output the regions of interest (ROIs) ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [10th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587)).\n\n## 3. Classification Process (2nd Stage Processing)\n#### Architecture\n- A common point among top teams was the use of Multiple Instance Learning (MIL) models that take multiple slices as input ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), etc...).\n  - While there were some differences in the details, many teams seemed to use a model architecture along the following lines:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fdba39e3ac9826122837181c473d6b9be%2FMIL.jpg?generation=1729260429260887&alt=media)\n\n    1. Extract feature vectors for each slice using a 2D backbone model.\n    2. Pass the feature vectors of each slice through an LSTM/GRU to capture temporal information.\n    3. Aggregate the feature vectors of the slices using Attention Pooling or Global Average Pooling (GAP).\n    4. Input the aggregated feature vector into the final head to output the classification result.\n\n  - The choice of 2D backbone models varied among teams (with ConvNeXt and EfficientNetv2 being somewhat popular), but they all commonly used smaller models.\n  - For feature vector aggregation, many teams used Attention Pooling.\n- As for the number of input slices, teams that performed slice index estimation typically used around 3 to 5 slices centered on the estimated index, while teams that did not estimate slice indices used around 10 to 30 slices at equal intervals (padding with dummy slices if necessary).\n- Some teams also employed 2.5D models, which treated consecutive slices stacked along the channel dimension like normal images ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569), etc...).\n- Due to space limitations, I won’t go into details about individual architectures, but each team had unique modeling approaches. If you're interested, I recommend checking out the discussions for more in-depth information.\n\n#### Multi-view Input\n- Generally, teams followed the annotations, using Sagittal T2 for SCS prediction, Sagittal T1 for NFN prediction, and Axial for SS prediction. However, some teams adopted multi-modal architectures that input images from multiple views.\n- There were two main approaches for implementing multi-view input models. One approach was to concatenate the feature vectors just before the final head, while the other was to combine the feature vectors at the 2D backbone stage.\n  - The first approach involved processing each view independently using the same steps as in a single-view model (2D backbone -> LSTM/GPU -> aggregate) and concatenating the feature vectors just before the final head to merge them into a single vector ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [9th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690), [10th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587), etc...).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F7c8f974b7505b96c98c7b36c1683e5f0%2Fmulti_view1.jpg?generation=1729260461167251&alt=media)\n\n  - The second approach involved inputting the feature vectors from all slices of all views into a Transformer Encoder right after the 2D backbone stage, allowing it to learn relationships between the feature vectors. These were then aggregated into a single vector using Attention Pooling or GAP ([3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), [4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), etc...).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F1b0036f5634c2fae35d141e7a378d56a%2Fmulti_view2.jpg?generation=1729260482124477&alt=media)\n\n  - There didn't seem to be a clearly superior method in this competition, but the insights gained from effectively fusing features from different sources could be useful for handling similar tasks in the future.\n\n#### Augmentation\n- To make the 2nd stage models more robust against errors introduced by slice index estimation or keypoint estimation in the 1st stage, random shifting of keypoint coordinates or slice indices during training was an effective augmentation strategy ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569), [12th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540001), etc...).\n- The [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803) place team used 27 cropping patterns during training ([original slice index, ±1 shifted slice index] * [original keypoint, keypoints shifted in one of 8 directions] = 3 * 9 = 27). During inference, they applied a similar technique by generating predictions for these 27 patterns and averaging the results, similar to Test-Time Augmentation (TTA), which greatly improved performance.\n\n#### Noisy Data Removal\n- The annotations provided in this competition were noisy (e.g., samples clearly labeled as \"normal\" were given \"severe\" labels), and some teams improved their scores by removing such data ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452)).\n- To detect noisy data, the 2nd team used ensemble out-of-fold (oof) predictions and excluded samples where the difference between the prediction and ground truth label was greater than 0.8, treating them as noise.\n\n- However, when I tried a similar approach in a late submission, the score consistently dropped. I couldn’t replicate the improvement in my environment.\n  - The possible reasons could be issues with my reproduction process or insufficient performance of the model used for noise detection. In any case, I realized that this method might backfire under certain conditions, so careful experimentation and validation are necessary when applying it in future competitions.",
      "votes": null
    },
    {
      "id": "3045881",
      "postDate": "11/14/2024 22:02:14",
      "content": "<p>Thank you for your great summary.</p>\n<p>I wonder why the use of LSTM/GRU over the use of more complex attention mechanism. Do you have an idea ?</p>",
      "rawMarkdown": "Thank you for your great summary.\n\nI wonder why the use of LSTM/GRU over the use of more complex attention mechanism. Do you have an idea ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3045881,
      "author_name": "hugovergnesx",
      "author_url": "",
      "post_date": "11/14/2024 22:02:14",
      "content": "<p>Thank you for your great summary.</p>\n<p>I wonder why the use of LSTM/GRU over the use of more complex attention mechanism. Do you have an idea ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3021441": "Thanks to all who participated in the competition.\nI learned a lot from the solutions that were made public after the competition. \nI read through the top solutions and compiled my own insights, so I’d like to share them here. If any of these solutions catch your interest, I’ve included links to the original posts, so please check them out and upvote!\nOnce again, I’d like to express my gratitude to all the participants who shared their solutions.\n\nNote:\nPlease forgive me if my interpretation may be incorrect. If there are any errors, I would appreciate it if you could let me know in the comments or otherwise.\n\n# Summary of Top Team Solutions\n## 1. Approach\n- Almost all teams adopted a two-stage pipeline ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), etc...)\n  - In the 1st stage, they estimated slice indices, keypoints, etc., and cropped the regions of interest at each level based on these estimations.\n  - In the 2nd stage, they used the cropped regions as inputs for label classification models to predict severity scores for each level.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fb7834266579f475e9e5db6d4ace15f41%2F2stage_pipeline_image.jpg?generation=1729260379466636&alt=media)\n\n- By cropping the regions at each level and treating them equally, they effectively increased the data size by a factor of 5 (for NFN and SS, treating the left and right sides equally resulted in a 5x2=10x increase). This seems to have been effective in a competition where the amount of data was somewhat limited.\n\n- On the other hand, a few teams adopted a single-stage pipeline ([7th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439), [8th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548)). \nNote: This means they didn’t crop the regions, so the term \"single-stage\" might be a bit misleading.\n  - The single-stage pipeline had the advantage of avoiding errors that occur during the region cropping estimation and being able to consider the overall context of the image. Although it is more challenging, by modeling it well, they were able to achieve accuracy on par with or better than the two-stage pipeline.\n  - For example, the [8th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539548) team trained keypoint detection (heatmap prediction) as a subtask, and from the weighted feature map obtained by the heatmap, they classified the classes. Without cropping, they managed to guide the model’s attention to the regions of interest effectively. (Note: There might be a misunderstanding as I haven’t fully decoded the source code.)\n\n## 2. Preprocessing (1st Stage Processing)\n#### Slice Index Estimation\n- Many top teams implemented the process of estimating slice indices (instance_number) suitable for predicting severity scores from a group of MRI slices using either \"a trained slice estimation model\" or \"rule-based\" methods.\n\n- The approach of training a slice estimation model involved using 2D/3D models to output the optimal slice index from the MRI slice volume and estimate the best slice index for each level ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), etc...).\n  - Since some patients have curved spines, causing the optimal slice index to vary by level, this approach was advantageous in addressing such cases.\n  - Tasks for this approach included predicting the relative position of the slice index (regression task), predicting the distance of each slice from the ground truth slice (regression task), or classifying whether each slice is the ground truth (classification task).\n\n- The rule-based approach often involved selecting the central slice for SCS and slices at a certain distance from the center for NFN as the optimal slice index ([9th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569), etc...).\n\n- Meanwhile, there were also teams that did not estimate slice indices at all ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012395), [3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), [4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), etc...).\n  - These teams seemed to implicitly resolve this by inputting multiple slices into the label classification model in the 2nd stage. By training the model with a sufficient number of slices, it is assumed that the model implicitly learned to select the slices most useful for prediction.\n\n\n#### Keypoint Estimation\n- A relatively large number of teams used 2D/2.5D regression models to directly estimate x and y relative coordinates. However, some teams estimated keypoints by outputting heatmaps through a segmentation model ([4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), [9th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690), etc...) or by using keypoint detectors like CenterNet ([3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453)).\n- Additionally, some teams did not prepare an independent keypoint estimation model but instead trained a model that simultaneously outputs both slice indices and keypoints ([5th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539472)).\n\n#### Axial Slice Level Correspondence\n- Axial slices correspond to specific levels, but since this correspondence is not explicitly provided, teams needed to estimate the relationships in some way.\n- Due to the availability of [helpful notebooks]((https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approach?scriptVersionId=191553260)), most teams aligned Axial slices with Sagittal slices using DICOM metadata and then estimated the corresponding level of each Axial slice based on the keypoints predicted from the Sagittal slices.\n- An alternative approach was used by the [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803) team, which utilized the alignment information as supplementary data while also training a model to predict which level each Axial slice belonged to.\n\n#### Level Region Cropping\n- Most teams determined the cropping range for each level based on the coordinates obtained from keypoint estimation.\n  - Many teams used fixed values to set the cropping size based on the estimated coordinates, but some teams dynamically determined the size using the distance between keypoints of adjacent levels ([4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569)).\n \n- There were also teams that did not use keypoints, but instead employed object detection models like YOLO or SpineNetV2 to directly output the regions of interest (ROIs) ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [10th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587)).\n\n## 3. Classification Process (2nd Stage Processing)\n#### Architecture\n- A common point among top teams was the use of Multiple Instance Learning (MIL) models that take multiple slices as input ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452), [3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), etc...).\n  - While there were some differences in the details, many teams seemed to use a model architecture along the following lines:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fdba39e3ac9826122837181c473d6b9be%2FMIL.jpg?generation=1729260429260887&alt=media)\n\n    1. Extract feature vectors for each slice using a 2D backbone model.\n    2. Pass the feature vectors of each slice through an LSTM/GRU to capture temporal information.\n    3. Aggregate the feature vectors of the slices using Attention Pooling or Global Average Pooling (GAP).\n    4. Input the aggregated feature vector into the final head to output the classification result.\n\n  - The choice of 2D backbone models varied among teams (with ConvNeXt and EfficientNetv2 being somewhat popular), but they all commonly used smaller models.\n  - For feature vector aggregation, many teams used Attention Pooling.\n- As for the number of input slices, teams that performed slice index estimation typically used around 3 to 5 slices centered on the estimated index, while teams that did not estimate slice indices used around 10 to 30 slices at equal intervals (padding with dummy slices if necessary).\n- Some teams also employed 2.5D models, which treated consecutive slices stacked along the channel dimension like normal images ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569), etc...).\n- Due to space limitations, I won’t go into details about individual architectures, but each team had unique modeling approaches. If you're interested, I recommend checking out the discussions for more in-depth information.\n\n#### Multi-view Input\n- Generally, teams followed the annotations, using Sagittal T2 for SCS prediction, Sagittal T1 for NFN prediction, and Axial for SS prediction. However, some teams adopted multi-modal architectures that input images from multiple views.\n- There were two main approaches for implementing multi-view input models. One approach was to concatenate the feature vectors just before the final head, while the other was to combine the feature vectors at the 2D backbone stage.\n  - The first approach involved processing each view independently using the same steps as in a single-view model (2D backbone -> LSTM/GPU -> aggregate) and concatenating the feature vectors just before the final head to merge them into a single vector ([1st](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091), [9th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690), [10th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587), etc...).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F7c8f974b7505b96c98c7b36c1683e5f0%2Fmulti_view1.jpg?generation=1729260461167251&alt=media)\n\n  - The second approach involved inputting the feature vectors from all slices of all views into a Transformer Encoder right after the 2D backbone stage, allowing it to learn relationships between the feature vectors. These were then aggregated into a single vector using Attention Pooling or GAP ([3rd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539453), [4th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443), etc...).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F1b0036f5634c2fae35d141e7a378d56a%2Fmulti_view2.jpg?generation=1729260482124477&alt=media)\n\n  - There didn't seem to be a clearly superior method in this competition, but the insights gained from effectively fusing features from different sources could be useful for handling similar tasks in the future.\n\n#### Augmentation\n- To make the 2nd stage models more robust against errors introduced by slice index estimation or keypoint estimation in the 1st stage, random shifting of keypoint coordinates or slice indices during training was an effective augmentation strategy ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803), [11th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539569), [12th](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540001), etc...).\n- The [2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452#3012803) place team used 27 cropping patterns during training ([original slice index, ±1 shifted slice index] * [original keypoint, keypoints shifted in one of 8 directions] = 3 * 9 = 27). During inference, they applied a similar technique by generating predictions for these 27 patterns and averaging the results, similar to Test-Time Augmentation (TTA), which greatly improved performance.\n\n#### Noisy Data Removal\n- The annotations provided in this competition were noisy (e.g., samples clearly labeled as \"normal\" were given \"severe\" labels), and some teams improved their scores by removing such data ([2nd](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452)).\n- To detect noisy data, the 2nd team used ensemble out-of-fold (oof) predictions and excluded samples where the difference between the prediction and ground truth label was greater than 0.8, treating them as noise.\n\n- However, when I tried a similar approach in a late submission, the score consistently dropped. I couldn’t replicate the improvement in my environment.\n  - The possible reasons could be issues with my reproduction process or insufficient performance of the model used for noise detection. In any case, I realized that this method might backfire under certain conditions, so careful experimentation and validation are necessary when applying it in future competitions.",
    "3045881": "Thank you for your great summary.\n\nI wonder why the use of LSTM/GRU over the use of more complex attention mechanism. Do you have an idea ?"
  },
  "source": "meta"
}