{
  "id": 539452,
  "title": "2nd place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539452",
  "author_name": "YujiAriyasu",
  "post_date": "2024-10-09T01:46:47.027000",
  "votes": 99,
  "comment_count": 22,
  "views": 0,
  "content": "<p>First, I would like to express my gratitude to Kaggle and RSNA for hosting this excellent competition.<br>\nI'm also grateful to my teammates who worked hard till the end.<br>\nOur solution is a simple blend of our individual predictions and small post-processing. My teammates will likely share their solutions in the replies to this post. I'll describe my solution and post-processing below.</p>\n<p>inference code: <a href=\"https://www.kaggle.com/code/yujiariyasu/rsna-lumbar-spine-2nd-place-solution\" target=\"_blank\">https://www.kaggle.com/code/yujiariyasu/rsna-lumbar-spine-2nd-place-solution</a></p>\n<h1>Summary</h1>\n<p>My solution is an ensemble of small models. I worked separately on axial and sagittal. Additionally, I created separate models for each target.<br>\nBasically, all models predict 3 targets: ['normal_mild', 'moderate', 'severe']. I used models that treat data from different levels and left/right as the same, without considering these distinctions. In the end, I used the team's ensemble oof to remove noisy labal data and retrain the classification model.</p>\n<h1>Axial</h1>\n<p>First, I classify which slices to use for predicting each level. I used <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> code for this - thank you always for your significant contributions!<br>\nNext, I estimate the regions within each image to use for severity prediction. I trained YOLOX using the provided data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2Fcdd5586eea2b4020457c4ad8b69b07a4%2F2024-10-18%2013.31.18.png?generation=1729225900120808&amp;alt=media\"></p>\n<p>Finally, I trained classification models using ConvNeXt Small. For spinal predictions, I directly use the regions estimated by YOLOX. For non-spinal predictions, I use only the left or right half of the image, allowing me to treat left and right labels equally.</p>\n<h1>Sagittal</h1>\n<p>First, I classify slices suitable for predicting spinal and subarticular targets. I used 2.5D images and a simple Timm model.<br>\nNext, I estimate regions for each level within the images. I trained YOLOX using data shared by my teammate <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> - thank you for your excellent contribution!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F4e96087f8f0251caf9d7553da5d55c9f%2F2024-10-18%2013.35.14.png?generation=1729226157298856&amp;alt=media\"></p>\n<p>Using boxes, level each level horizontally and then crop.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F81d7c0015e4a5bf975cd54d1e100c6ef%2F2024-10-18%2013.39.21.png?generation=1729226474453903&amp;alt=media\"></p>\n<p>Finally, I perform classification using a MIL model that accepts 5 images. The backbone is ConvNeXt Small. For spinal and subarticular, I use 5 slices centered on those predicted in the 1st stage. For foraminal, I use 5 slices centered between the spinal and subarticular slices. Some models use T1/T2 in separate channels, while others use only one.</p>\n<h1>noise reduction</h1>\n<p>My teammate discovered label noise in train dataset, so we removed samples with high loss. Using our ensemble oof (CV: 0.3687), we excluded samples where the difference between the label and the predicted value was 0.8 or greater. Due to imbalanced data, we needed to apply coefficients to the moderate and severe categories. This magic improved our score by 1% on both public and private leaderboards. I came up with this idea just two days before the deadline, so I didn't have time to try various methods or coefficients. There are likely better approaches.</p>\n<p>This is a brief overview of my solution. There are many more intricate details that I couldn't include here.</p>\n<p>training code: <a href=\"https://github.com/yujiariyasu/rsna_2024_lumbar_spine_degenerative_classification\" target=\"_blank\">https://github.com/yujiariyasu/rsna_2024_lumbar_spine_degenerative_classification</a></p>\n<h1>Ensemble and Post-processing</h1>\n<p>I simply weighted-averaged the predictions of each member, then applied post-processing only to spinal predictions.<br>\nFor each study, I multiplied the highest predicted spinal-severe value among the 5 levels by 1.25.</p>\n<p>Again, thanks to my teammates!</p>",
  "messages": [
    {
      "id": 3012378,
      "postDate": "2024-10-09T01:46:47.027Z",
      "content": "<p>First, I would like to express my gratitude to Kaggle and RSNA for hosting this excellent competition.<br>\nI'm also grateful to my teammates who worked hard till the end.<br>\nOur solution is a simple blend of our individual predictions and small post-processing. My teammates will likely share their solutions in the replies to this post. I'll describe my solution and post-processing below.</p>\n<p>inference code: <a href=\"https://www.kaggle.com/code/yujiariyasu/rsna-lumbar-spine-2nd-place-solution\" target=\"_blank\">https://www.kaggle.com/code/yujiariyasu/rsna-lumbar-spine-2nd-place-solution</a></p>\n<h1>Summary</h1>\n<p>My solution is an ensemble of small models. I worked separately on axial and sagittal. Additionally, I created separate models for each target.<br>\nBasically, all models predict 3 targets: ['normal_mild', 'moderate', 'severe']. I used models that treat data from different levels and left/right as the same, without considering these distinctions. In the end, I used the team's ensemble oof to remove noisy labal data and retrain the classification model.</p>\n<h1>Axial</h1>\n<p>First, I classify which slices to use for predicting each level. I used <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> code for this - thank you always for your significant contributions!<br>\nNext, I estimate the regions within each image to use for severity prediction. I trained YOLOX using the provided data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2Fcdd5586eea2b4020457c4ad8b69b07a4%2F2024-10-18%2013.31.18.png?generation=1729225900120808&amp;alt=media\"></p>\n<p>Finally, I trained classification models using ConvNeXt Small. For spinal predictions, I directly use the regions estimated by YOLOX. For non-spinal predictions, I use only the left or right half of the image, allowing me to treat left and right labels equally.</p>\n<h1>Sagittal</h1>\n<p>First, I classify slices suitable for predicting spinal and subarticular targets. I used 2.5D images and a simple Timm model.<br>\nNext, I estimate regions for each level within the images. I trained YOLOX using data shared by my teammate <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> - thank you for your excellent contribution!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F4e96087f8f0251caf9d7553da5d55c9f%2F2024-10-18%2013.35.14.png?generation=1729226157298856&amp;alt=media\"></p>\n<p>Using boxes, level each level horizontally and then crop.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F81d7c0015e4a5bf975cd54d1e100c6ef%2F2024-10-18%2013.39.21.png?generation=1729226474453903&amp;alt=media\"></p>\n<p>Finally, I perform classification using a MIL model that accepts 5 images. The backbone is ConvNeXt Small. For spinal and subarticular, I use 5 slices centered on those predicted in the 1st stage. For foraminal, I use 5 slices centered between the spinal and subarticular slices. Some models use T1/T2 in separate channels, while others use only one.</p>\n<h1>noise reduction</h1>\n<p>My teammate discovered label noise in train dataset, so we removed samples with high loss. Using our ensemble oof (CV: 0.3687), we excluded samples where the difference between the label and the predicted value was 0.8 or greater. Due to imbalanced data, we needed to apply coefficients to the moderate and severe categories. This magic improved our score by 1% on both public and private leaderboards. I came up with this idea just two days before the deadline, so I didn't have time to try various methods or coefficients. There are likely better approaches.</p>\n<p>This is a brief overview of my solution. There are many more intricate details that I couldn't include here.</p>\n<p>training code: <a href=\"https://github.com/yujiariyasu/rsna_2024_lumbar_spine_degenerative_classification\" target=\"_blank\">https://github.com/yujiariyasu/rsna_2024_lumbar_spine_degenerative_classification</a></p>\n<h1>Ensemble and Post-processing</h1>\n<p>I simply weighted-averaged the predictions of each member, then applied post-processing only to spinal predictions.<br>\nFor each study, I multiplied the highest predicted spinal-severe value among the 5 levels by 1.25.</p>\n<p>Again, thanks to my teammates!</p>",
      "rawMarkdown": "First, I would like to express my gratitude to Kaggle and RSNA for hosting this excellent competition.\nI'm also grateful to my teammates who worked hard till the end.\nOur solution is a simple blend of our individual predictions and small post-processing. My teammates will likely share their solutions in the replies to this post. I'll describe my solution and post-processing below.\n\ninference code: https://www.kaggle.com/code/yujiariyasu/rsna-lumbar-spine-2nd-place-solution\n\n# Summary\nMy solution is an ensemble of small models. I worked separately on axial and sagittal. Additionally, I created separate models for each target.\nBasically, all models predict 3 targets: ['normal_mild', 'moderate', 'severe']. I used models that treat data from different levels and left/right as the same, without considering these distinctions. In the end, I used the team's ensemble oof to remove noisy labal data and retrain the classification model.\n\n# Axial\nFirst, I classify which slices to use for predicting each level. I used @hengck23 code for this - thank you always for your significant contributions!\nNext, I estimate the regions within each image to use for severity prediction. I trained YOLOX using the provided data.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2Fcdd5586eea2b4020457c4ad8b69b07a4%2F2024-10-18%2013.31.18.png?generation=1729225900120808&alt=media\" width=\"800\">\n\n\nFinally, I trained classification models using ConvNeXt Small. For spinal predictions, I directly use the regions estimated by YOLOX. For non-spinal predictions, I use only the left or right half of the image, allowing me to treat left and right labels equally.\n\n# Sagittal\nFirst, I classify slices suitable for predicting spinal and subarticular targets. I used 2.5D images and a simple Timm model.\nNext, I estimate regions for each level within the images. I trained YOLOX using data shared by my teammate @brendanartley - thank you for your excellent contribution!\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F4e96087f8f0251caf9d7553da5d55c9f%2F2024-10-18%2013.35.14.png?generation=1729226157298856&alt=media\" width=\"400\">\n\nUsing boxes, level each level horizontally and then crop.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F81d7c0015e4a5bf975cd54d1e100c6ef%2F2024-10-18%2013.39.21.png?generation=1729226474453903&alt=media\" width=\"400\">\n\nFinally, I perform classification using a MIL model that accepts 5 images. The backbone is ConvNeXt Small. For spinal and subarticular, I use 5 slices centered on those predicted in the 1st stage. For foraminal, I use 5 slices centered between the spinal and subarticular slices. Some models use T1/T2 in separate channels, while others use only one.\n\n# noise reduction\nMy teammate discovered label noise in train dataset, so we removed samples with high loss. Using our ensemble oof (CV: 0.3687), we excluded samples where the difference between the label and the predicted value was 0.8 or greater. Due to imbalanced data, we needed to apply coefficients to the moderate and severe categories. This magic improved our score by 1% on both public and private leaderboards. I came up with this idea just two days before the deadline, so I didn't have time to try various methods or coefficients. There are likely better approaches.\n\nThis is a brief overview of my solution. There are many more intricate details that I couldn't include here.\n\ntraining code: https://github.com/yujiariyasu/rsna_2024_lumbar_spine_degenerative_classification\n\n# Ensemble and Post-processing\nI simply weighted-averaged the predictions of each member, then applied post-processing only to spinal predictions.\nFor each study, I multiplied the highest predicted spinal-severe value among the 5 levels by 1.25.\n\nAgain, thanks to my teammates!",
      "votes": 99
    },
    {
      "id": 3012803,
      "postDate": "2024-10-09T11:39:35.723Z",
      "content": "<p>Thank you to my teammates <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a>, <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, and <a href=\"https://www.kaggle.com/kevin0912\" target=\"_blank\">@kevin0912</a> for a great competition. Below are more specifics regarding my part of the solution.</p>\n<h2>Stage 1a - Slice Localization</h2>\n<h3>Foraminal</h3>\n<p>I trained a CNN-transformer model taking the 3D sagittal T1 series as input and predicting the 20 targets: the distance (using ImagePositionPatient0) from the target foramen slice as specified in the ground truth coordinates for both left and right foramina at each level (10 targets) as well as a classification label for the target foramen slice (1 for the exact slice, 0.5 for adjacent slices for each side at each level (10 targets). This was optimized using a combined smooth L1 and BCE loss. During inference, I found that taking the slice with the minimum predicted distance was better than using the max classification score. </p>\n<h3>Spinal (Sagittal)</h3>\n<p>Same as above, except using sagittal T2 series as input and predicting 10 targets (no left and right for spinal canal).</p>\n<h3>Subarticular</h3>\n<p>Using the ground truth coordinates I assigned each axial T2 slice to one of 11 labels. If the slice was positive for either left or right subarticular, then it was assigned an intervertebral level (e.g., L1/L2, L2/L3). If the slice was in between two intervertebral levels, it was assigned the level in between (e.g., L2, L3). If it was above L1 or below S1 it was assigned L1 or S1, for a total of 11 classes. I also added 2 classes for right and left subarticular, as well as 10 targets for the distance (using ImagePositionPatient2) from the target slice for each level and side. </p>\n<p>I also used the <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s mapping code to map the predicted sagittal T2 coordinates for each spinal canal level to the axial plane. Then, for each level, I could calculate the distance of each axial slice from the predicted sagittal T2 spinal canal level. This resulted in a vector of length 5 for each slice, representing the distance of each axial slice from the predicted sagittal T2 spinal canal location along the axial z-axis for each level. I also trained a CNN-transformer model as above, except for subarticular, I concatenated this vector to the image embedding of the slice before inputting into the transformer head. I noticed that for axial T2 slices, there were several studies where the levels were offset by 1 (e.g., L1/L2 predicted as L2/L3 or vice versa). Adding in the additional information of mapped sagittal T2 coordinates improved this significantly. </p>\n<p>For subarticular, I found it was more effective to use the slice with maximum classification score instead of minimum distance. </p>\n<h3>Spinal (Axial)</h3>\n<p>Radiologists (myself included) also use axial T2 images to grade spinal stenosis. At my institution we pretty much only look at axial T2s for this. Thus I found it very helpful to include spinal canal stenosis models trained on axial T2 images. To identify the slice, I just took the average of the right and left subarticular slices (oftentimes the same or offset by 1). </p>\n<h2>Stage 1b - Keypoint Localization</h2>\n<p>For foraminal, spinal (sagittal), and subarticular, I trained a regression model on 2D images using the ground truth coordinates to predict the keypoint coordinates. For foraminal and spinal (sagittal), this involved predicting each level. For subarticular, this involved predicting each side. </p>\n<p>For many cases, the levels (for foraminal/spinal) or left/right (for subarticular) were on different slices. I just propagated the coordinates across all slices where there was a keypoint to make training easier. </p>\n<p>For axial spinal keypoint localization, I took the average of the right and left subarticular coordinates and multiplied the y-coordinate by 1.05, since the center of the canal is typically a bit lower than the subarticular zones (assuming the image is conventionally oriented).</p>\n<h2>Stage 2 - Stenosis Grading</h2>\n<p>Using the ground truth coordinates, I generated crops for each target of interest. I found that training on crops generated using ground truth coordinates was slightly better than training on crops generated using predicted coordinates from my localization models. </p>\n<p>For the most part, I generated crops with 3 image channels, where the center channel was the target slice and the 2 other channels were the adjacent slices on either side. For subarticular, I also trained a few models with 5 image channels (using 2 adjacent slices on either side instead of 1). </p>\n<p>I also generated crops using slice offset of 1 in either direction as well as slightly offset in the xy-plane of the image, for a total of 27 crops per target ([original slice + 1 slice offset in either direction] x [original coordinates + 8 in-plane spatial offsets]). This means using 3 image channels I actually had information across 5 slices. This augmented the training data and helped account for minor errors in slice/keypoint localization. </p>\n<p>Each crop model was trained to predict 3 classes: normal_mild, moderate, and severe. At inference, all of the crop predictions for a target were aggregated using either median or mean. This can be thought of as a form of test-time augmentation (TTA) and improved performance significantly. The loss was a sample-weighted log loss similar to the competition metric. </p>\n<p>I trained simple 2D CNNs treating images as 3-channel images, as well as attention pooling models and MIL models (<a href=\"https://github.com/mahmoodlab/CLAM)\" target=\"_blank\">https://github.com/mahmoodlab/CLAM)</a>, using different backbones. The simple CNNs actually did the best, but the other models added diversity to the ensemble. I used a variety of backbones, including maxvit_tiny_tf_512, coatnet_1_rw_224, NFNet-0, and CSN-ResNet101 (which is actually a model for 3D data, so for this model the channel dimension was treated as a spatial dimension).</p>\n<p>This resulted in 4 groups of models: foraminal (sagittal T1), spinal (sagittal T2), subarticular (axial T2), and spinal (axial T2). Spinal was averaged using weight of 1 for sagittal and 0.5 for axial, though there was not much difference from equal weighting.</p>\n<p>My part of the model had CV of 0.3805 and LB of 0.35. I did not test my retrained models on LB after the noise reduction step described by Yuji, but I imagine it would have been 0.34.</p>\n<h2>Note on handling multiple axial T2 series</h2>\n<p>Many studies had multiple axial T2 series. Several axial T2 series did not have all of the levels. The subarticular slice localization model was also used to predict whether a level was present, since it predicted what level each slice was at. If a level was duplicated across two axial T2 series, then crops from both were used. For spinal axial T2 prediction, if a level was not predicted to be present, the spinal sagittal model prediction was used on its own. </p>\n<h2>Highlights</h2>\n<p>I think the top 3 most important elements of my solution were:</p>\n<ul>\n<li>Training and aggregating over multiple crops (especially slices)</li>\n<li>Using both sagittal T2 and axial T2 series to predict spinal canal stenosis</li>\n<li>Noise reduction step (not unique to my solution)</li>\n</ul>\n<p>Noise reduction was a last-minute attempt for us to improve our score. I noticed during the last week as I was reviewing incorrect predictions that there were a fair number of incorrect labels. However, I thought the models could learn through the noise and that removing these samples would not change anything. Yuji ended up testing the idea, which resulted in a notable improvement in our LB score - and thankfully our final models and submissions made it in time.</p>",
      "rawMarkdown": "Thank you to my teammates @yujiariyasu, @brendanartley, and @kevin0912 for a great competition. Below are more specifics regarding my part of the solution.\n\n## Stage 1a - Slice Localization\n\n### Foraminal\n\nI trained a CNN-transformer model taking the 3D sagittal T1 series as input and predicting the 20 targets: the distance (using ImagePositionPatient0) from the target foramen slice as specified in the ground truth coordinates for both left and right foramina at each level (10 targets) as well as a classification label for the target foramen slice (1 for the exact slice, 0.5 for adjacent slices for each side at each level (10 targets). This was optimized using a combined smooth L1 and BCE loss. During inference, I found that taking the slice with the minimum predicted distance was better than using the max classification score. \n\n### Spinal (Sagittal)\n\nSame as above, except using sagittal T2 series as input and predicting 10 targets (no left and right for spinal canal).\n\n### Subarticular\n\nUsing the ground truth coordinates I assigned each axial T2 slice to one of 11 labels. If the slice was positive for either left or right subarticular, then it was assigned an intervertebral level (e.g., L1/L2, L2/L3). If the slice was in between two intervertebral levels, it was assigned the level in between (e.g., L2, L3). If it was above L1 or below S1 it was assigned L1 or S1, for a total of 11 classes. I also added 2 classes for right and left subarticular, as well as 10 targets for the distance (using ImagePositionPatient2) from the target slice for each level and side. \n\nI also used the @hengck23's mapping code to map the predicted sagittal T2 coordinates for each spinal canal level to the axial plane. Then, for each level, I could calculate the distance of each axial slice from the predicted sagittal T2 spinal canal level. This resulted in a vector of length 5 for each slice, representing the distance of each axial slice from the predicted sagittal T2 spinal canal location along the axial z-axis for each level. I also trained a CNN-transformer model as above, except for subarticular, I concatenated this vector to the image embedding of the slice before inputting into the transformer head. I noticed that for axial T2 slices, there were several studies where the levels were offset by 1 (e.g., L1/L2 predicted as L2/L3 or vice versa). Adding in the additional information of mapped sagittal T2 coordinates improved this significantly. \n\nFor subarticular, I found it was more effective to use the slice with maximum classification score instead of minimum distance. \n\n### Spinal (Axial)\n\nRadiologists (myself included) also use axial T2 images to grade spinal stenosis. At my institution we pretty much only look at axial T2s for this. Thus I found it very helpful to include spinal canal stenosis models trained on axial T2 images. To identify the slice, I just took the average of the right and left subarticular slices (oftentimes the same or offset by 1). \n\n## Stage 1b - Keypoint Localization\n\nFor foraminal, spinal (sagittal), and subarticular, I trained a regression model on 2D images using the ground truth coordinates to predict the keypoint coordinates. For foraminal and spinal (sagittal), this involved predicting each level. For subarticular, this involved predicting each side. \n\nFor many cases, the levels (for foraminal/spinal) or left/right (for subarticular) were on different slices. I just propagated the coordinates across all slices where there was a keypoint to make training easier. \n\nFor axial spinal keypoint localization, I took the average of the right and left subarticular coordinates and multiplied the y-coordinate by 1.05, since the center of the canal is typically a bit lower than the subarticular zones (assuming the image is conventionally oriented).\n\n## Stage 2 - Stenosis Grading\n\nUsing the ground truth coordinates, I generated crops for each target of interest. I found that training on crops generated using ground truth coordinates was slightly better than training on crops generated using predicted coordinates from my localization models. \n\nFor the most part, I generated crops with 3 image channels, where the center channel was the target slice and the 2 other channels were the adjacent slices on either side. For subarticular, I also trained a few models with 5 image channels (using 2 adjacent slices on either side instead of 1). \n\nI also generated crops using slice offset of 1 in either direction as well as slightly offset in the xy-plane of the image, for a total of 27 crops per target ([original slice + 1 slice offset in either direction] x [original coordinates + 8 in-plane spatial offsets]). This means using 3 image channels I actually had information across 5 slices. This augmented the training data and helped account for minor errors in slice/keypoint localization. \n\nEach crop model was trained to predict 3 classes: normal_mild, moderate, and severe. At inference, all of the crop predictions for a target were aggregated using either median or mean. This can be thought of as a form of test-time augmentation (TTA) and improved performance significantly. The loss was a sample-weighted log loss similar to the competition metric. \n\nI trained simple 2D CNNs treating images as 3-channel images, as well as attention pooling models and MIL models (https://github.com/mahmoodlab/CLAM), using different backbones. The simple CNNs actually did the best, but the other models added diversity to the ensemble. I used a variety of backbones, including maxvit_tiny_tf_512, coatnet_1_rw_224, NFNet-0, and CSN-ResNet101 (which is actually a model for 3D data, so for this model the channel dimension was treated as a spatial dimension).\n\nThis resulted in 4 groups of models: foraminal (sagittal T1), spinal (sagittal T2), subarticular (axial T2), and spinal (axial T2). Spinal was averaged using weight of 1 for sagittal and 0.5 for axial, though there was not much difference from equal weighting.\n\nMy part of the model had CV of 0.3805 and LB of 0.35. I did not test my retrained models on LB after the noise reduction step described by Yuji, but I imagine it would have been 0.34.\n\n## Note on handling multiple axial T2 series\n\nMany studies had multiple axial T2 series. Several axial T2 series did not have all of the levels. The subarticular slice localization model was also used to predict whether a level was present, since it predicted what level each slice was at. If a level was duplicated across two axial T2 series, then crops from both were used. For spinal axial T2 prediction, if a level was not predicted to be present, the spinal sagittal model prediction was used on its own. \n\n## Highlights\n\nI think the top 3 most important elements of my solution were:\n- Training and aggregating over multiple crops (especially slices)\n- Using both sagittal T2 and axial T2 series to predict spinal canal stenosis\n- Noise reduction step (not unique to my solution)\n\nNoise reduction was a last-minute attempt for us to improve our score. I noticed during the last week as I was reviewing incorrect predictions that there were a fair number of incorrect labels. However, I thought the models could learn through the noise and that removing these samples would not change anything. Yuji ended up testing the idea, which resulted in a notable improvement in our LB score - and thankfully our final models and submissions made it in time.",
      "votes": 24,
      "replies": [
        {
          "id": 3015331,
          "postDate": "2024-10-12T08:45:53.787Z",
          "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> ,</p>\n<p>I have few questions regarding the solution,</p>\n<ol>\n<li>In you CNN-transformer model, did you use image metadata as input to transformer or not ?</li>\n<li>What image sizes did you use ?</li>\n<li>In Foraminal (sagittal T1) image slices you stated that minimum predicted distance was better than using the max classification score. What was the observation for Spinal (Sagittal T2) series?</li>\n</ol>",
          "rawMarkdown": "Congratulations @vaillant ,\n\nI have few questions regarding the solution,\n1. In you CNN-transformer model, did you use image metadata as input to transformer or not ?\n2. What image sizes did you use ?\n3. In Foraminal (sagittal T1) image slices you stated that minimum predicted distance was better than using the max classification score. What was the observation for Spinal (Sagittal T2) series?",
          "replies": [
            {
              "id": 3016960,
              "postDate": "2024-10-14T11:21:27.947Z",
              "content": "<ol>\n<li><p>No image metadata was used.</p></li>\n<li><p>Crops were all 64 x 64.</p></li>\n<li><p>For sagittal T2 spinal, using minimum distance was very slightly better. </p></li>\n</ol>",
              "rawMarkdown": "1. No image metadata was used.\n\n2. Crops were all 64 x 64.\n\n3. For sagittal T2 spinal, using minimum distance was very slightly better. ",
              "votes": 3
            },
            {
              "id": 3191100,
              "postDate": "2025-05-01T11:42:15.100Z",
              "content": "<p>Congratulations, I am very amazed by your work!</p>\n<p>What confuses me is that the original MRI image shapes vary from 272,272 to 1024,1024, <br>\nso 64,64 would cover a different area for different sizes of images.</p>\n<p>I wonder what size of subarticular zone you have in your cropped images.<br>\nSorry if the question is trivial, I have no previous experience</p>",
              "rawMarkdown": "Congratulations, I am very amazed by your work!\n\nWhat confuses me is that the original MRI image shapes vary from 272,272 to 1024,1024, \nso 64,64 would cover a different area for different sizes of images.\n\nI wonder what size of subarticular zone you have in your cropped images.\nSorry if the question is trivial, I have no previous experience"
            }
          ]
        }
      ]
    },
    {
      "id": 3012395,
      "postDate": "2024-10-09T02:16:11.853Z",
      "content": "<h2>Some more details</h2>\n<p>Thanks to my awesome teammates, I learned a lot from them and we worked hard for this result. Also, thanks to RSNA and Kaggle for hosting this competition.</p>\n<p>TLDR; My solution is a 2-stage pipeline that uses sagittal images only. Spinal and subarticular labels are inferred from sagittal T2s and foraminal labels from sagittal T1s. I use many data augmentation tricks, pseudo labeling, TTA, and more. For details, keep reading!</p>\n<h2>Stage 1</h2>\n<p>In this stage, I predict the location of each disc using the competition coordinates and then crop around the area of each disc. I use the coordinate dataset <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">here</a>, which contains corrected coordinates and additional coordinates for Sagittal T2s.</p>\n<p>The T1 and T2 pipelines are very similar in this stage. I pass the eight middle frames of each image sequence into a 2D backbone with an attention head to predict ten coordinates (two per disc). For T1s, each level has one coordinate corresponding to the right and left foraminal zones. For T2s, each level has one coordinate behind each disc and one in front of each disc. The performance of the localization models was quite strong, with predictions falling within 5 pixels of the ground truth coordinates ~99.1% of the time.</p>\n<p>I take the average distance between levels as the height. For T1s, the width is the same as the height, but for T2s, I take the average disc width. You can see an example crop in the figure below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe1da057424be2a25dd7d9b47eedd5df0%2Fstage1.jpg?generation=1728439805566813&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<h2>Stage 2</h2>\n<p>In this stage, we classify the labels for each cropped sequence. Each sequence is treated equally, regardless of the disc level. I used <code>0.5*loss_with_labels + 0.5*loss_with_pseudolabels</code> during training. Using pseudo labels was the only denoising method I tried, and in hindsight, I should have experimented with other denoising strategies sooner.</p>\n<p>For T1s, I pass the middle 24 frames into an encoder, pass the output sequence of embeddings into an LSTM layer, and then use attention to pool the embeddings into a single feature representation. During training, I apply heavy augmentations, flip the sequence order, combine left/right sides from different sequences, and apply manifold mixup to the embeddings. The idea was to try and force the model to extract information from both sides of the sequence. For T2s, the process is similar. It differs in that it uses between 5 - 16 frames and different image sizes. The differing sizes increased the diversity of predictions. Finally, nine rotations provide TTA during inference.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F715b5e3979d5ff4e15a8c7b536f08a82%2Fstage2.jpg?generation=1729177127889019&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p>I also worked on an Axial T2 pipeline that improved my CV by 0.01-0.02, but it was not strong enough to improve our overall team CV. I have cleaned up the codebase, and released it <a href=\"https://github.com/brendanartley/RSNA-2024-Competition\" target=\"_blank\">here</a>. Happy Kaggling 😀</p>\n<p>--</p>\n<p>Final note, although <a href=\"https://www.kaggle.com/kevin0912\" target=\"_blank\">@kevin0912</a> did not have models in the final ensemble, his contributions/discussions helped us achieve better performance overall as a team. Thanks again to all my teammates!</p>",
      "rawMarkdown": "## Some more details\n\nThanks to my awesome teammates, I learned a lot from them and we worked hard for this result. Also, thanks to RSNA and Kaggle for hosting this competition.\n\nTLDR; My solution is a 2-stage pipeline that uses sagittal images only. Spinal and subarticular labels are inferred from sagittal T2s and foraminal labels from sagittal T1s. I use many data augmentation tricks, pseudo labeling, TTA, and more. For details, keep reading!\n\n## Stage 1\n\nIn this stage, I predict the location of each disc using the competition coordinates and then crop around the area of each disc. I use the coordinate dataset [here](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset), which contains corrected coordinates and additional coordinates for Sagittal T2s.\n\nThe T1 and T2 pipelines are very similar in this stage. I pass the eight middle frames of each image sequence into a 2D backbone with an attention head to predict ten coordinates (two per disc). For T1s, each level has one coordinate corresponding to the right and left foraminal zones. For T2s, each level has one coordinate behind each disc and one in front of each disc. The performance of the localization models was quite strong, with predictions falling within 5 pixels of the ground truth coordinates ~99.1% of the time.\n\nI take the average distance between levels as the height. For T1s, the width is the same as the height, but for T2s, I take the average disc width. You can see an example crop in the figure below.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe1da057424be2a25dd7d9b47eedd5df0%2Fstage1.jpg?generation=1728439805566813&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n## Stage 2\n\nIn this stage, we classify the labels for each cropped sequence. Each sequence is treated equally, regardless of the disc level. I used `0.5*loss_with_labels + 0.5*loss_with_pseudolabels` during training. Using pseudo labels was the only denoising method I tried, and in hindsight, I should have experimented with other denoising strategies sooner.\n\nFor T1s, I pass the middle 24 frames into an encoder, pass the output sequence of embeddings into an LSTM layer, and then use attention to pool the embeddings into a single feature representation. During training, I apply heavy augmentations, flip the sequence order, combine left/right sides from different sequences, and apply manifold mixup to the embeddings. The idea was to try and force the model to extract information from both sides of the sequence. For T2s, the process is similar. It differs in that it uses between 5 - 16 frames and different image sizes. The differing sizes increased the diversity of predictions. Finally, nine rotations provide TTA during inference.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F715b5e3979d5ff4e15a8c7b536f08a82%2Fstage2.jpg?generation=1729177127889019&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\nI also worked on an Axial T2 pipeline that improved my CV by 0.01-0.02, but it was not strong enough to improve our overall team CV. I have cleaned up the codebase, and released it [here](https://github.com/brendanartley/RSNA-2024-Competition). Happy Kaggling 😀\n\n--\n\nFinal note, although @kevin0912 did not have models in the final ensemble, his contributions/discussions helped us achieve better performance overall as a team. Thanks again to all my teammates!",
      "votes": 22,
      "replies": [
        {
          "id": 3013052,
          "postDate": "2024-10-09T16:01:37.600Z",
          "content": "<p>Congratulations on winning the second prize in this competition. Also, thanks for sharing the details of your solution,</p>",
          "rawMarkdown": "Congratulations on winning the second prize in this competition. Also, thanks for sharing the details of your solution,",
          "votes": 1
        },
        {
          "id": 3016783,
          "postDate": "2024-10-14T06:47:28.527Z",
          "content": "<p>According to the dataset, <code>Sagittal T2</code> is only associated with <code>Spinal Canal Stenosis</code>, but according to 2nd stage work flow, it's <code>Foraminal</code> which is associated with <code>Sagittal T1</code> and <code>Axial T2</code>; </p>\n<pre><code>merged_df = train_coords.merge(train_desc, how=, on=[, ])\nmerged_df.groupby([])[].value_counts()\n</code></pre>",
          "rawMarkdown": "According to the dataset, `Sagittal T2` is only associated with `Spinal Canal Stenosis`, but according to 2nd stage work flow, it's `Foraminal` which is associated with `Sagittal T1` and `Axial T2`; \n\n```python\nmerged_df = train_coords.merge(train_desc, how='inner', on=['study_id', 'series_id'])\nmerged_df.groupby(['condition'])['series_description'].value_counts()\n```",
          "replies": [
            {
              "id": 3018353,
              "postDate": "2024-10-15T17:33:41.870Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/samu2505\" target=\"_blank\">@samu2505</a>, just updated the figure.</p>",
              "rawMarkdown": "Thanks @samu2505, just updated the figure.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3080822,
      "postDate": "2024-12-25T19:28:44.103Z",
      "content": "<p>Our prediction book saves and runs in 2 minutes, everything completes, but it continues to send, it takes too long and timed out on our last try. The change of libraries was added in the figure, but the solution to the real problem was not fully understood. How can we optimize our prediction book?</p>",
      "rawMarkdown": "\nOur prediction book saves and runs in 2 minutes, everything completes, but it continues to send, it takes too long and timed out on our last try. The change of libraries was added in the figure, but the solution to the real problem was not fully understood. How can we optimize our prediction book?",
      "votes": 3
    },
    {
      "id": 3014853,
      "postDate": "2024-10-11T16:38:03.763Z",
      "content": "<p>congrats brother! that's actually great :-)</p>",
      "rawMarkdown": "congrats brother! that's actually great :-)",
      "votes": 1
    },
    {
      "id": 3012626,
      "postDate": "2024-10-09T08:13:41.080Z",
      "content": "<p>Congratulations, great work!</p>",
      "rawMarkdown": "Congratulations, great work!",
      "votes": 1
    },
    {
      "id": 3018905,
      "postDate": "2024-10-16T05:51:48.050Z",
      "content": "<p>Thanks for sharing your contribution and insights. Congratulatios!</p>",
      "rawMarkdown": "Thanks for sharing your contribution and insights. Congratulatios!"
    },
    {
      "id": 3017248,
      "postDate": "2024-10-14T17:16:48.167Z",
      "content": "<p>congrats! clever solutions</p>",
      "rawMarkdown": "congrats! clever solutions"
    },
    {
      "id": 3016894,
      "postDate": "2024-10-14T09:04:36.593Z",
      "content": "<p>Thanks for sharing your contribution and insights. Nice work.</p>",
      "rawMarkdown": "Thanks for sharing your contribution and insights. Nice work."
    },
    {
      "id": 3014606,
      "postDate": "2024-10-11T12:14:51.283Z",
      "content": "<p>Congratss!</p>",
      "rawMarkdown": "Congratss!"
    },
    {
      "id": 3013824,
      "postDate": "2024-10-10T15:21:07.117Z",
      "content": "<p>Congratulations and thanks for sharing your insights!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your insights!"
    },
    {
      "id": 3013649,
      "postDate": "2024-10-10T11:37:55.073Z",
      "content": "<p>Congratulations Yuji! Great work!</p>",
      "rawMarkdown": "Congratulations Yuji! Great work!\n"
    },
    {
      "id": 3013174,
      "postDate": "2024-10-09T17:48:59.237Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3013473,
      "postDate": "2024-10-10T07:02:42.723Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!",
      "votes": 1
    },
    {
      "id": 3013414,
      "postDate": "2024-10-10T05:03:44.970Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": 1
    },
    {
      "id": 3016878,
      "postDate": "2024-10-14T08:52:00.903Z",
      "content": "<p>Great, thanks. <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> </p>",
      "rawMarkdown": "Great, thanks. @yujiariyasu "
    },
    {
      "id": 3014929,
      "postDate": "2024-10-11T18:05:24.910Z",
      "content": "<p>Congratulations! </p>",
      "rawMarkdown": "Congratulations! "
    }
  ],
  "comments": [
    {
      "id": 3012803,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2024-10-09T11:39:35.723000",
      "content": "<p>Thank you to my teammates <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a>, <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, and <a href=\"https://www.kaggle.com/kevin0912\" target=\"_blank\">@kevin0912</a> for a great competition. Below are more specifics regarding my part of the solution.</p>\n<h2>Stage 1a - Slice Localization</h2>\n<h3>Foraminal</h3>\n<p>I trained a CNN-transformer model taking the 3D sagittal T1 series as input and predicting the 20 targets: the distance (using ImagePositionPatient0) from the target foramen slice as specified in the ground truth coordinates for both left and right foramina at each level (10 targets) as well as a classification label for the target foramen slice (1 for the exact slice, 0.5 for adjacent slices for each side at each level (10 targets). This was optimized using a combined smooth L1 and BCE loss. During inference, I found that taking the slice with the minimum predicted distance was better than using the max classification score. </p>\n<h3>Spinal (Sagittal)</h3>\n<p>Same as above, except using sagittal T2 series as input and predicting 10 targets (no left and right for spinal canal).</p>\n<h3>Subarticular</h3>\n<p>Using the ground truth coordinates I assigned each axial T2 slice to one of 11 labels. If the slice was positive for either left or right subarticular, then it was assigned an intervertebral level (e.g., L1/L2, L2/L3). If the slice was in between two intervertebral levels, it was assigned the level in between (e.g., L2, L3). If it was above L1 or below S1 it was assigned L1 or S1, for a total of 11 classes. I also added 2 classes for right and left subarticular, as well as 10 targets for the distance (using ImagePositionPatient2) from the target slice for each level and side. </p>\n<p>I also used the <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s mapping code to map the predicted sagittal T2 coordinates for each spinal canal level to the axial plane. Then, for each level, I could calculate the distance of each axial slice from the predicted sagittal T2 spinal canal level. This resulted in a vector of length 5 for each slice, representing the distance of each axial slice from the predicted sagittal T2 spinal canal location along the axial z-axis for each level. I also trained a CNN-transformer model as above, except for subarticular, I concatenated this vector to the image embedding of the slice before inputting into the transformer head. I noticed that for axial T2 slices, there were several studies where the levels were offset by 1 (e.g., L1/L2 predicted as L2/L3 or vice versa). Adding in the additional information of mapped sagittal T2 coordinates improved this significantly. </p>\n<p>For subarticular, I found it was more effective to use the slice with maximum classification score instead of minimum distance. </p>\n<h3>Spinal (Axial)</h3>\n<p>Radiologists (myself included) also use axial T2 images to grade spinal stenosis. At my institution we pretty much only look at axial T2s for this. Thus I found it very helpful to include spinal canal stenosis models trained on axial T2 images. To identify the slice, I just took the average of the right and left subarticular slices (oftentimes the same or offset by 1). </p>\n<h2>Stage 1b - Keypoint Localization</h2>\n<p>For foraminal, spinal (sagittal), and subarticular, I trained a regression model on 2D images using the ground truth coordinates to predict the keypoint coordinates. For foraminal and spinal (sagittal), this involved predicting each level. For subarticular, this involved predicting each side. </p>\n<p>For many cases, the levels (for foraminal/spinal) or left/right (for subarticular) were on different slices. I just propagated the coordinates across all slices where there was a keypoint to make training easier. </p>\n<p>For axial spinal keypoint localization, I took the average of the right and left subarticular coordinates and multiplied the y-coordinate by 1.05, since the center of the canal is typically a bit lower than the subarticular zones (assuming the image is conventionally oriented).</p>\n<h2>Stage 2 - Stenosis Grading</h2>\n<p>Using the ground truth coordinates, I generated crops for each target of interest. I found that training on crops generated using ground truth coordinates was slightly better than training on crops generated using predicted coordinates from my localization models. </p>\n<p>For the most part, I generated crops with 3 image channels, where the center channel was the target slice and the 2 other channels were the adjacent slices on either side. For subarticular, I also trained a few models with 5 image channels (using 2 adjacent slices on either side instead of 1). </p>\n<p>I also generated crops using slice offset of 1 in either direction as well as slightly offset in the xy-plane of the image, for a total of 27 crops per target ([original slice + 1 slice offset in either direction] x [original coordinates + 8 in-plane spatial offsets]). This means using 3 image channels I actually had information across 5 slices. This augmented the training data and helped account for minor errors in slice/keypoint localization. </p>\n<p>Each crop model was trained to predict 3 classes: normal_mild, moderate, and severe. At inference, all of the crop predictions for a target were aggregated using either median or mean. This can be thought of as a form of test-time augmentation (TTA) and improved performance significantly. The loss was a sample-weighted log loss similar to the competition metric. </p>\n<p>I trained simple 2D CNNs treating images as 3-channel images, as well as attention pooling models and MIL models (<a href=\"https://github.com/mahmoodlab/CLAM)\" target=\"_blank\">https://github.com/mahmoodlab/CLAM)</a>, using different backbones. The simple CNNs actually did the best, but the other models added diversity to the ensemble. I used a variety of backbones, including maxvit_tiny_tf_512, coatnet_1_rw_224, NFNet-0, and CSN-ResNet101 (which is actually a model for 3D data, so for this model the channel dimension was treated as a spatial dimension).</p>\n<p>This resulted in 4 groups of models: foraminal (sagittal T1), spinal (sagittal T2), subarticular (axial T2), and spinal (axial T2). Spinal was averaged using weight of 1 for sagittal and 0.5 for axial, though there was not much difference from equal weighting.</p>\n<p>My part of the model had CV of 0.3805 and LB of 0.35. I did not test my retrained models on LB after the noise reduction step described by Yuji, but I imagine it would have been 0.34.</p>\n<h2>Note on handling multiple axial T2 series</h2>\n<p>Many studies had multiple axial T2 series. Several axial T2 series did not have all of the levels. The subarticular slice localization model was also used to predict whether a level was present, since it predicted what level each slice was at. If a level was duplicated across two axial T2 series, then crops from both were used. For spinal axial T2 prediction, if a level was not predicted to be present, the spinal sagittal model prediction was used on its own. </p>\n<h2>Highlights</h2>\n<p>I think the top 3 most important elements of my solution were:</p>\n<ul>\n<li>Training and aggregating over multiple crops (especially slices)</li>\n<li>Using both sagittal T2 and axial T2 series to predict spinal canal stenosis</li>\n<li>Noise reduction step (not unique to my solution)</li>\n</ul>\n<p>Noise reduction was a last-minute attempt for us to improve our score. I noticed during the last week as I was reviewing incorrect predictions that there were a fair number of incorrect labels. However, I thought the models could learn through the noise and that removing these samples would not change anything. Yuji ended up testing the idea, which resulted in a notable improvement in our LB score - and thankfully our final models and submissions made it in time.</p>",
      "votes": 24,
      "replies": [
        {
          "id": 3015331,
          "author_name": "Rohit Chaudhari",
          "author_url": "",
          "post_date": "2024-10-12T08:45:53.787000",
          "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> ,</p>\n<p>I have few questions regarding the solution,</p>\n<ol>\n<li>In you CNN-transformer model, did you use image metadata as input to transformer or not ?</li>\n<li>What image sizes did you use ?</li>\n<li>In Foraminal (sagittal T1) image slices you stated that minimum predicted distance was better than using the max classification score. What was the observation for Spinal (Sagittal T2) series?</li>\n</ol>",
          "votes": 0,
          "replies": [
            {
              "id": 3016960,
              "author_name": "Ian Pan",
              "author_url": "",
              "post_date": "2024-10-14T11:21:27.947000",
              "content": "<ol>\n<li><p>No image metadata was used.</p></li>\n<li><p>Crops were all 64 x 64.</p></li>\n<li><p>For sagittal T2 spinal, using minimum distance was very slightly better. </p></li>\n</ol>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3191100,
              "author_name": "Rasim Berke Turan",
              "author_url": "",
              "post_date": "2025-05-01T11:42:15.100000",
              "content": "<p>Congratulations, I am very amazed by your work!</p>\n<p>What confuses me is that the original MRI image shapes vary from 272,272 to 1024,1024, <br>\nso 64,64 would cover a different area for different sizes of images.</p>\n<p>I wonder what size of subarticular zone you have in your cropped images.<br>\nSorry if the question is trivial, I have no previous experience</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3012395,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-10-09T02:16:11.853000",
      "content": "<h2>Some more details</h2>\n<p>Thanks to my awesome teammates, I learned a lot from them and we worked hard for this result. Also, thanks to RSNA and Kaggle for hosting this competition.</p>\n<p>TLDR; My solution is a 2-stage pipeline that uses sagittal images only. Spinal and subarticular labels are inferred from sagittal T2s and foraminal labels from sagittal T1s. I use many data augmentation tricks, pseudo labeling, TTA, and more. For details, keep reading!</p>\n<h2>Stage 1</h2>\n<p>In this stage, I predict the location of each disc using the competition coordinates and then crop around the area of each disc. I use the coordinate dataset <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">here</a>, which contains corrected coordinates and additional coordinates for Sagittal T2s.</p>\n<p>The T1 and T2 pipelines are very similar in this stage. I pass the eight middle frames of each image sequence into a 2D backbone with an attention head to predict ten coordinates (two per disc). For T1s, each level has one coordinate corresponding to the right and left foraminal zones. For T2s, each level has one coordinate behind each disc and one in front of each disc. The performance of the localization models was quite strong, with predictions falling within 5 pixels of the ground truth coordinates ~99.1% of the time.</p>\n<p>I take the average distance between levels as the height. For T1s, the width is the same as the height, but for T2s, I take the average disc width. You can see an example crop in the figure below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe1da057424be2a25dd7d9b47eedd5df0%2Fstage1.jpg?generation=1728439805566813&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<h2>Stage 2</h2>\n<p>In this stage, we classify the labels for each cropped sequence. Each sequence is treated equally, regardless of the disc level. I used <code>0.5*loss_with_labels + 0.5*loss_with_pseudolabels</code> during training. Using pseudo labels was the only denoising method I tried, and in hindsight, I should have experimented with other denoising strategies sooner.</p>\n<p>For T1s, I pass the middle 24 frames into an encoder, pass the output sequence of embeddings into an LSTM layer, and then use attention to pool the embeddings into a single feature representation. During training, I apply heavy augmentations, flip the sequence order, combine left/right sides from different sequences, and apply manifold mixup to the embeddings. The idea was to try and force the model to extract information from both sides of the sequence. For T2s, the process is similar. It differs in that it uses between 5 - 16 frames and different image sizes. The differing sizes increased the diversity of predictions. Finally, nine rotations provide TTA during inference.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F715b5e3979d5ff4e15a8c7b536f08a82%2Fstage2.jpg?generation=1729177127889019&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p>I also worked on an Axial T2 pipeline that improved my CV by 0.01-0.02, but it was not strong enough to improve our overall team CV. I have cleaned up the codebase, and released it <a href=\"https://github.com/brendanartley/RSNA-2024-Competition\" target=\"_blank\">here</a>. Happy Kaggling 😀</p>\n<p>--</p>\n<p>Final note, although <a href=\"https://www.kaggle.com/kevin0912\" target=\"_blank\">@kevin0912</a> did not have models in the final ensemble, his contributions/discussions helped us achieve better performance overall as a team. Thanks again to all my teammates!</p>",
      "votes": 22,
      "replies": [
        {
          "id": 3013052,
          "author_name": "C R Suthikshn Kumar",
          "author_url": "",
          "post_date": "2024-10-09T16:01:37.600000",
          "content": "<p>Congratulations on winning the second prize in this competition. Also, thanks for sharing the details of your solution,</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3016783,
          "author_name": "samu2505",
          "author_url": "",
          "post_date": "2024-10-14T06:47:28.527000",
          "content": "<p>According to the dataset, <code>Sagittal T2</code> is only associated with <code>Spinal Canal Stenosis</code>, but according to 2nd stage work flow, it's <code>Foraminal</code> which is associated with <code>Sagittal T1</code> and <code>Axial T2</code>; </p>\n<pre><code>merged_df = train_coords.merge(train_desc, how=, on=[, ])\nmerged_df.groupby([])[].value_counts()\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 3018353,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-10-15T17:33:41.870000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/samu2505\" target=\"_blank\">@samu2505</a>, just updated the figure.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3080822,
      "author_name": "İbrahim Öztürk",
      "author_url": "",
      "post_date": "2024-12-25T19:28:44.103000",
      "content": "<p>Our prediction book saves and runs in 2 minutes, everything completes, but it continues to send, it takes too long and timed out on our last try. The change of libraries was added in the figure, but the solution to the real problem was not fully understood. How can we optimize our prediction book?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3014853,
      "author_name": "Evanhis",
      "author_url": "",
      "post_date": "2024-10-11T16:38:03.763000",
      "content": "<p>congrats brother! that's actually great :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3012626,
      "author_name": "Marek Nurzynski",
      "author_url": "",
      "post_date": "2024-10-09T08:13:41.080000",
      "content": "<p>Congratulations, great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3018905,
      "author_name": "Tanishk Patil",
      "author_url": "",
      "post_date": "2024-10-16T05:51:48.050000",
      "content": "<p>Thanks for sharing your contribution and insights. Congratulatios!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3017248,
      "author_name": "Mingguangyi Yang",
      "author_url": "",
      "post_date": "2024-10-14T17:16:48.167000",
      "content": "<p>congrats! clever solutions</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3016894,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2024-10-14T09:04:36.593000",
      "content": "<p>Thanks for sharing your contribution and insights. Nice work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3014606,
      "author_name": "MrSimple",
      "author_url": "",
      "post_date": "2024-10-11T12:14:51.283000",
      "content": "<p>Congratss!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3013824,
      "author_name": "Gauthier Avite",
      "author_url": "",
      "post_date": "2024-10-10T15:21:07.117000",
      "content": "<p>Congratulations and thanks for sharing your insights!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3013649,
      "author_name": "Priyash_07",
      "author_url": "",
      "post_date": "2024-10-10T11:37:55.073000",
      "content": "<p>Congratulations Yuji! Great work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3013174,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-09T17:48:59.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3013473,
      "author_name": "AstonJin",
      "author_url": "",
      "post_date": "2024-10-10T07:02:42.723000",
      "content": "<p>congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3013414,
      "author_name": "Abhishek",
      "author_url": "",
      "post_date": "2024-10-10T05:03:44.970000",
      "content": "<p>Congratulations!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3016878,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2024-10-14T08:52:00.903000",
      "content": "<p>Great, thanks. <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3014929,
      "author_name": "Humayra Khanom Rime",
      "author_url": "",
      "post_date": "2024-10-11T18:05:24.910000",
      "content": "<p>Congratulations! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3012378": "First, I would like to express my gratitude to Kaggle and RSNA for hosting this excellent competition.\nI'm also grateful to my teammates who worked hard till the end.\nOur solution is a simple blend of our individual predictions and small post-processing. My teammates will likely share their solutions in the replies to this post. I'll describe my solution and post-processing below.\n\ninference code: https://www.kaggle.com/code/yujiariyasu/rsna-lumbar-spine-2nd-place-solution\n\n# Summary\nMy solution is an ensemble of small models. I worked separately on axial and sagittal. Additionally, I created separate models for each target.\nBasically, all models predict 3 targets: ['normal_mild', 'moderate', 'severe']. I used models that treat data from different levels and left/right as the same, without considering these distinctions. In the end, I used the team's ensemble oof to remove noisy labal data and retrain the classification model.\n\n# Axial\nFirst, I classify which slices to use for predicting each level. I used @hengck23 code for this - thank you always for your significant contributions!\nNext, I estimate the regions within each image to use for severity prediction. I trained YOLOX using the provided data.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2Fcdd5586eea2b4020457c4ad8b69b07a4%2F2024-10-18%2013.31.18.png?generation=1729225900120808&alt=media\" width=\"800\">\n\n\nFinally, I trained classification models using ConvNeXt Small. For spinal predictions, I directly use the regions estimated by YOLOX. For non-spinal predictions, I use only the left or right half of the image, allowing me to treat left and right labels equally.\n\n# Sagittal\nFirst, I classify slices suitable for predicting spinal and subarticular targets. I used 2.5D images and a simple Timm model.\nNext, I estimate regions for each level within the images. I trained YOLOX using data shared by my teammate @brendanartley - thank you for your excellent contribution!\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F4e96087f8f0251caf9d7553da5d55c9f%2F2024-10-18%2013.35.14.png?generation=1729226157298856&alt=media\" width=\"400\">\n\nUsing boxes, level each level horizontally and then crop.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3460291%2F81d7c0015e4a5bf975cd54d1e100c6ef%2F2024-10-18%2013.39.21.png?generation=1729226474453903&alt=media\" width=\"400\">\n\nFinally, I perform classification using a MIL model that accepts 5 images. The backbone is ConvNeXt Small. For spinal and subarticular, I use 5 slices centered on those predicted in the 1st stage. For foraminal, I use 5 slices centered between the spinal and subarticular slices. Some models use T1/T2 in separate channels, while others use only one.\n\n# noise reduction\nMy teammate discovered label noise in train dataset, so we removed samples with high loss. Using our ensemble oof (CV: 0.3687), we excluded samples where the difference between the label and the predicted value was 0.8 or greater. Due to imbalanced data, we needed to apply coefficients to the moderate and severe categories. This magic improved our score by 1% on both public and private leaderboards. I came up with this idea just two days before the deadline, so I didn't have time to try various methods or coefficients. There are likely better approaches.\n\nThis is a brief overview of my solution. There are many more intricate details that I couldn't include here.\n\ntraining code: https://github.com/yujiariyasu/rsna_2024_lumbar_spine_degenerative_classification\n\n# Ensemble and Post-processing\nI simply weighted-averaged the predictions of each member, then applied post-processing only to spinal predictions.\nFor each study, I multiplied the highest predicted spinal-severe value among the 5 levels by 1.25.\n\nAgain, thanks to my teammates!",
    "3012803": "Thank you to my teammates @yujiariyasu, @brendanartley, and @kevin0912 for a great competition. Below are more specifics regarding my part of the solution.\n\n## Stage 1a - Slice Localization\n\n### Foraminal\n\nI trained a CNN-transformer model taking the 3D sagittal T1 series as input and predicting the 20 targets: the distance (using ImagePositionPatient0) from the target foramen slice as specified in the ground truth coordinates for both left and right foramina at each level (10 targets) as well as a classification label for the target foramen slice (1 for the exact slice, 0.5 for adjacent slices for each side at each level (10 targets). This was optimized using a combined smooth L1 and BCE loss. During inference, I found that taking the slice with the minimum predicted distance was better than using the max classification score. \n\n### Spinal (Sagittal)\n\nSame as above, except using sagittal T2 series as input and predicting 10 targets (no left and right for spinal canal).\n\n### Subarticular\n\nUsing the ground truth coordinates I assigned each axial T2 slice to one of 11 labels. If the slice was positive for either left or right subarticular, then it was assigned an intervertebral level (e.g., L1/L2, L2/L3). If the slice was in between two intervertebral levels, it was assigned the level in between (e.g., L2, L3). If it was above L1 or below S1 it was assigned L1 or S1, for a total of 11 classes. I also added 2 classes for right and left subarticular, as well as 10 targets for the distance (using ImagePositionPatient2) from the target slice for each level and side. \n\nI also used the @hengck23's mapping code to map the predicted sagittal T2 coordinates for each spinal canal level to the axial plane. Then, for each level, I could calculate the distance of each axial slice from the predicted sagittal T2 spinal canal level. This resulted in a vector of length 5 for each slice, representing the distance of each axial slice from the predicted sagittal T2 spinal canal location along the axial z-axis for each level. I also trained a CNN-transformer model as above, except for subarticular, I concatenated this vector to the image embedding of the slice before inputting into the transformer head. I noticed that for axial T2 slices, there were several studies where the levels were offset by 1 (e.g., L1/L2 predicted as L2/L3 or vice versa). Adding in the additional information of mapped sagittal T2 coordinates improved this significantly. \n\nFor subarticular, I found it was more effective to use the slice with maximum classification score instead of minimum distance. \n\n### Spinal (Axial)\n\nRadiologists (myself included) also use axial T2 images to grade spinal stenosis. At my institution we pretty much only look at axial T2s for this. Thus I found it very helpful to include spinal canal stenosis models trained on axial T2 images. To identify the slice, I just took the average of the right and left subarticular slices (oftentimes the same or offset by 1). \n\n## Stage 1b - Keypoint Localization\n\nFor foraminal, spinal (sagittal), and subarticular, I trained a regression model on 2D images using the ground truth coordinates to predict the keypoint coordinates. For foraminal and spinal (sagittal), this involved predicting each level. For subarticular, this involved predicting each side. \n\nFor many cases, the levels (for foraminal/spinal) or left/right (for subarticular) were on different slices. I just propagated the coordinates across all slices where there was a keypoint to make training easier. \n\nFor axial spinal keypoint localization, I took the average of the right and left subarticular coordinates and multiplied the y-coordinate by 1.05, since the center of the canal is typically a bit lower than the subarticular zones (assuming the image is conventionally oriented).\n\n## Stage 2 - Stenosis Grading\n\nUsing the ground truth coordinates, I generated crops for each target of interest. I found that training on crops generated using ground truth coordinates was slightly better than training on crops generated using predicted coordinates from my localization models. \n\nFor the most part, I generated crops with 3 image channels, where the center channel was the target slice and the 2 other channels were the adjacent slices on either side. For subarticular, I also trained a few models with 5 image channels (using 2 adjacent slices on either side instead of 1). \n\nI also generated crops using slice offset of 1 in either direction as well as slightly offset in the xy-plane of the image, for a total of 27 crops per target ([original slice + 1 slice offset in either direction] x [original coordinates + 8 in-plane spatial offsets]). This means using 3 image channels I actually had information across 5 slices. This augmented the training data and helped account for minor errors in slice/keypoint localization. \n\nEach crop model was trained to predict 3 classes: normal_mild, moderate, and severe. At inference, all of the crop predictions for a target were aggregated using either median or mean. This can be thought of as a form of test-time augmentation (TTA) and improved performance significantly. The loss was a sample-weighted log loss similar to the competition metric. \n\nI trained simple 2D CNNs treating images as 3-channel images, as well as attention pooling models and MIL models (https://github.com/mahmoodlab/CLAM), using different backbones. The simple CNNs actually did the best, but the other models added diversity to the ensemble. I used a variety of backbones, including maxvit_tiny_tf_512, coatnet_1_rw_224, NFNet-0, and CSN-ResNet101 (which is actually a model for 3D data, so for this model the channel dimension was treated as a spatial dimension).\n\nThis resulted in 4 groups of models: foraminal (sagittal T1), spinal (sagittal T2), subarticular (axial T2), and spinal (axial T2). Spinal was averaged using weight of 1 for sagittal and 0.5 for axial, though there was not much difference from equal weighting.\n\nMy part of the model had CV of 0.3805 and LB of 0.35. I did not test my retrained models on LB after the noise reduction step described by Yuji, but I imagine it would have been 0.34.\n\n## Note on handling multiple axial T2 series\n\nMany studies had multiple axial T2 series. Several axial T2 series did not have all of the levels. The subarticular slice localization model was also used to predict whether a level was present, since it predicted what level each slice was at. If a level was duplicated across two axial T2 series, then crops from both were used. For spinal axial T2 prediction, if a level was not predicted to be present, the spinal sagittal model prediction was used on its own. \n\n## Highlights\n\nI think the top 3 most important elements of my solution were:\n- Training and aggregating over multiple crops (especially slices)\n- Using both sagittal T2 and axial T2 series to predict spinal canal stenosis\n- Noise reduction step (not unique to my solution)\n\nNoise reduction was a last-minute attempt for us to improve our score. I noticed during the last week as I was reviewing incorrect predictions that there were a fair number of incorrect labels. However, I thought the models could learn through the noise and that removing these samples would not change anything. Yuji ended up testing the idea, which resulted in a notable improvement in our LB score - and thankfully our final models and submissions made it in time.",
    "3012395": "## Some more details\n\nThanks to my awesome teammates, I learned a lot from them and we worked hard for this result. Also, thanks to RSNA and Kaggle for hosting this competition.\n\nTLDR; My solution is a 2-stage pipeline that uses sagittal images only. Spinal and subarticular labels are inferred from sagittal T2s and foraminal labels from sagittal T1s. I use many data augmentation tricks, pseudo labeling, TTA, and more. For details, keep reading!\n\n## Stage 1\n\nIn this stage, I predict the location of each disc using the competition coordinates and then crop around the area of each disc. I use the coordinate dataset [here](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset), which contains corrected coordinates and additional coordinates for Sagittal T2s.\n\nThe T1 and T2 pipelines are very similar in this stage. I pass the eight middle frames of each image sequence into a 2D backbone with an attention head to predict ten coordinates (two per disc). For T1s, each level has one coordinate corresponding to the right and left foraminal zones. For T2s, each level has one coordinate behind each disc and one in front of each disc. The performance of the localization models was quite strong, with predictions falling within 5 pixels of the ground truth coordinates ~99.1% of the time.\n\nI take the average distance between levels as the height. For T1s, the width is the same as the height, but for T2s, I take the average disc width. You can see an example crop in the figure below.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe1da057424be2a25dd7d9b47eedd5df0%2Fstage1.jpg?generation=1728439805566813&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n## Stage 2\n\nIn this stage, we classify the labels for each cropped sequence. Each sequence is treated equally, regardless of the disc level. I used `0.5*loss_with_labels + 0.5*loss_with_pseudolabels` during training. Using pseudo labels was the only denoising method I tried, and in hindsight, I should have experimented with other denoising strategies sooner.\n\nFor T1s, I pass the middle 24 frames into an encoder, pass the output sequence of embeddings into an LSTM layer, and then use attention to pool the embeddings into a single feature representation. During training, I apply heavy augmentations, flip the sequence order, combine left/right sides from different sequences, and apply manifold mixup to the embeddings. The idea was to try and force the model to extract information from both sides of the sequence. For T2s, the process is similar. It differs in that it uses between 5 - 16 frames and different image sizes. The differing sizes increased the diversity of predictions. Finally, nine rotations provide TTA during inference.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F715b5e3979d5ff4e15a8c7b536f08a82%2Fstage2.jpg?generation=1729177127889019&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\nI also worked on an Axial T2 pipeline that improved my CV by 0.01-0.02, but it was not strong enough to improve our overall team CV. I have cleaned up the codebase, and released it [here](https://github.com/brendanartley/RSNA-2024-Competition). Happy Kaggling 😀\n\n--\n\nFinal note, although @kevin0912 did not have models in the final ensemble, his contributions/discussions helped us achieve better performance overall as a team. Thanks again to all my teammates!",
    "3080822": "\nOur prediction book saves and runs in 2 minutes, everything completes, but it continues to send, it takes too long and timed out on our last try. The change of libraries was added in the figure, but the solution to the real problem was not fully understood. How can we optimize our prediction book?",
    "3014853": "congrats brother! that's actually great :-)",
    "3012626": "Congratulations, great work!",
    "3018905": "Thanks for sharing your contribution and insights. Congratulatios!",
    "3017248": "congrats! clever solutions",
    "3016894": "Thanks for sharing your contribution and insights. Nice work.",
    "3014606": "Congratss!",
    "3013824": "Congratulations and thanks for sharing your insights!",
    "3013649": "Congratulations Yuji! Great work!\n",
    "3013174": "",
    "3013473": "congratulations!",
    "3013414": "Congratulations!!",
    "3016878": "Great, thanks. @yujiariyasu ",
    "3014929": "Congratulations! "
  }
}