{
  "id": 541813,
  "title": "6th Place Solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/nvspine-6th-place-solution",
  "author_name": "",
  "post_date": "2024-10-24T09:09:14.980Z",
  "votes": 20,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about MRI data and how to create strong models for it. <br>\nThis competition was very challenging (and time-consuming !) because it had 3 underlying modalities (SCS, NFN, SS) and 2 tasks (disk localization, severity classification).</p>\n<h2>TLDR</h2>\n<p>Our solution is an ensemble of multiple models which train on study or series level for each individual intervertebral disc level. Crops of the original MRI are isolated in a first stage and the severity of the condition is predicted using the sequence of crops. Final results are aggregated using an MLP model that directly optimizes the competition metric.</p>\n<h2>Cross validation</h2>\n<p>A common cross validation approach was used within the team with four folds. Below is the tracking of CV against leaderboard scores. Our experience from previous RSNA competitions led us to expect a small yet reasonable shake-up, and we trusted our CV more than public LB.</p>\n<p><a href=\"https://ibb.co/LzxRtby\"><img src=\"https://i.ibb.co/yn6PyGC/cvlb.png\" alt=\"cvlb\"></a></p>\n<h2>Data &amp; Augmentations</h2>\n<p>All data was sourced from competition data, except for the spinenet model weights. Some of our models used <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> publicly shared coordinates (thanks a lot!). </p>\n<p>For preprocessing we mostly used min-max normalisation, and sometimes added windowing based on dicom window width and window length. We relied on albumentations for augmentations, and used a variety of approaches for the different pipelines:</p>\n<ul>\n<li><strong>Model 1 :</strong><ul>\n<li>ShiftScaleRotate, ElasticTransform</li>\n<li>RandomGamma, RandomBrightnessContrast</li>\n<li>MotionBlur, GaussianBlur</li>\n<li>MixUp for the Sagittal only models.</li>\n<li>Flipping left / right and the associated targets for some models.</li></ul></li>\n<li><strong>Model 2 :</strong><ul>\n<li>ShiftScaleRotate</li>\n<li>RandomBrightnessContrast</li>\n<li>CoarseDropout</li>\n<li>RandomCrop</li>\n<li>For localisation of crops, augmenting the target coord position was very effective</li>\n<li>Random shift start and end instance number of crop </li></ul></li>\n<li><strong>Model 3 :</strong><ul>\n<li>ShiftScaleRotate</li>\n<li>RandomBrightnessContrast</li>\n<li>CoarseDropout</li>\n<li>RandomCrop</li>\n<li>ChannelDropout</li></ul></li>\n</ul>\n<h2>Models</h2>\n<h3>Model 1</h3>\n<p>The code to this part of the pipeline is available on github</p>\n<blockquote>\n  <p><a href=\"https://github.com/TheoViel/kaggle_rsna_lumbar_spine\" target=\"_blank\">https://github.com/TheoViel/kaggle_rsna_lumbar_spine</a></p>\n</blockquote>\n<p>It achieves ~0.42 private LB, using only the sagittal data.</p>\n<p><a href=\"https://ibb.co/kcm5KVW\"><img src=\"https://i.ibb.co/HghB7Lk/RSNA-spine-theo-drawio.png\" alt=\"RSNA-spine-theo-drawio\"></a></p>\n<p><strong>Coordinates model :</strong> A simple <code>coatnet_rmlp_2_rw_384</code> is trained on all the center frames of the sagittal images to predict the 5 (x, y) coordinates associated with the disk injuries. It is trained with the MSE on 10 classes, and used to generate crops. Crop generation is done by selecting the square of 20% of the image size around the predicted ROI center.</p>\n<p><strong>Classification models :</strong> On the crop above, we once again trained CoAtNets, this time <code>coatnet_1_rw_224</code> and <code>coatnet_2_rw_224</code> and with a RNN layer to incorporate the 3D information. <br>\nWe initially wanted to train separate models for each injury since each image modality had its target. For instance the SCS only models are trained to predict the 3 severity classes on the Sagittal T2 images, by sampling the 5 (or 3) frames at the center of the stack. But handling the two sagittal modalities together worked better, and even showed decent performance on the SS task. Further improvements came from adding more frames (5 is not enough to capture the SCS and SS signal) and using 3 RNN heads that had access to different frames - for the left, right and center (scs) targets. We also added MixUp and trained models for 10 epochs (vs 5 for the SCS models) and with a higher learning rate (1e-3 vs 5e-4 for the SCS models). </p>\n<p><strong>MLP :</strong> The classification models are trained with the CE, and tweaked to maximize the AUC. The MLP model is here to aggregate predictions, and account for the competition metric. Surprisingly, what worked best here is to consider each target independently, i.e. the scs_l1_l2 does not interact with the scs_l2_l3 features nor the nfn_left_l1_l2 features in the MLP. The logits layer weights for the different levels are shared though, i.e. the model consists of 3 MLP (one for SCS, one for SS and one for SCS).</p>\n<h3>Model 2</h3>\n<p>The code to this part of the pipeline is available on github</p>\n<blockquote>\n  <p><a href=\"https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\" target=\"_blank\">https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/</a></p>\n</blockquote>\n<p><a href=\"https://ibb.co/2hCR7GR\"><img src=\"https://i.ibb.co/nmY4sx4/RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx.png\" alt=\"RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx\"></a></p>\n<p>For Sagittal and Axial images we train severity classification on series and individual intervertebral disc level. Therefore we need to isolate each individual disc within the dicom. </p>\n<p><strong>Axial xy-localisation :</strong> We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of <code>efficientnetv2_rw_t</code> with a linear head and L1-loss is used with learning rate 1e-4 for 16 epochs with a batchsize of 16. The dicom images are resized to 384 and images are fed to the model individually. </p>\n<p><strong>Sagittal xy-localisation :</strong> We use spinenet to predict the left top corner point of the intervertebral disc. Many of the spinenet points are shifted incorrectly by one disc. We use the train coordinates file to identify the shifted ones and shift them by one disc. We then retrain the localisation on the corrected spinenet predictions. We exclude from training series where the label is too far from the spinenet prediction. The same training model and procedure is used as axial xy-localisation, except we have an xy label for each vertebrae and we add a mask label to indicate slices with no spinenet xy prediction (in this case, xy-loss is masked).</p>\n<p><strong>Sagittal z-localisation :</strong> We use the train coordinates file to learn the z-location (the annotated instance number). For all series we predict the annotated instance of spinal_canal_stenosis as well as left and right neural_foraminal_narrowing instance number. For sagittal t2, we mask the foraminal loss and for sagittal t1 we loss the spinal canal loss. L1 loss is used where the target is the number of annotations on the instance, between 0 and 5. Same training procedure as above.</p>\n<p><strong>Axial z-localisation :</strong> We leverage sagittal point to axial level mapping from the repo M-Scan repo - code here - to map the xyz position of the sagittal t2 series to the z-position of the patient’s axial series. This identifies which instance number, or slice, contains which intervertebral disc. Inspiration from Ian Pan, Heng and others who shared this approach earlier on in the competition.</p>\n<p><strong>Axial severity classification :</strong> Crops for stage 2 are made by taking the distance of right to left annotation and extending outward either side by half this distance. On the z-position the sequence is cropped by finding the max distance from one level to the next and extending outward either side by this number of slices from the level’s center slice. Therefore we use a variable number of slices per series/level. A 2.5d model of <code>efficientnetv2_rw_t</code> with a bidirectional single layer RNN head (512 dim) and weighted CE loss is used. We use learning rate 4e-4 for 5 epochs with a batch size of 8 series/levels with all crops resized to 256 dim.<br>\nWe use two separate axial models. One to predict the Axial t2 severity, and one to predict the Sagittal t2 severity. The model was not effective in predicting Sagittal t1 severity. </p>\n<p><strong>Sagittal severity classification :</strong> Crops for stage 2 are made by taking the max distance of one vertebrae to the next and extending outward either direction for each level from the levels center xy position. On the z-position the sequence is cropped by excluding slices which were predicted to have no spinenet annotation (as seen above under Sagittal xy-localisation above). Sagittal t1 and t2 labels and series were trained together.  </p>\n<h3>Model 3</h3>\n<p>The code to this part of the pipeline is available on github</p>\n<blockquote>\n  <p><a href=\"https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\" target=\"_blank\">https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/</a></p>\n</blockquote>\n<p><a href=\"https://ibb.co/WxgRc9H\"><img src=\"https://i.ibb.co/j6gjLK8/Screenshot-2024-10-19-at-09-45-26.png\" alt=\"Screenshot-2024-10-19-at-09-45-26\"></a></p>\n<p><strong>Axial xy-localisation :</strong> We use the train coordinates file to learn the xy-location of the labelled points for each slice. A backbone of <code>tf_efficientnetv2_s</code> with a linear head and L1-loss. The dicom images are resized to 384 and images are fed to the model individually. </p>\n<p><strong>Sagittal xy-localisation :</strong> We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of <code>tf_efficientnetv2_s</code> with a linear head and L1-loss is used. The dicom images are resized to 384 and images are fed to the model individually. It was beneficial to train a single model on a combined dataset of T1 and T2 sagittal images </p>\n<p><strong>Axial z-localisation :</strong> Same approach as Model 2. We derived axial z localisation for each vertebrae from T1/ T2 Sagittal xy-localization from same study using 3d coordinates. 3d coordinates are derived separately for T1 and T2 and then averaged. </p>\n<p><strong>Axial Crops :</strong> Multiple axial series for a study are combined by sorting by ImagePositionPatient. Then median of x1/y1 and x2/y2 predictions for all slices are used to create 112x224 sized bounding boxes for a study. For each slice of the axial series a intervertebral disk (IVD) is assigned using the closest derived z coordinates. A maximum distance of 20 was used. So for each IVD a 3d bounding box with height x width of 112x224 and varying number of slices is cropped and saved to disk as a npy file. </p>\n<p><strong>Sagittal Crops :</strong> 3d crops for stage 2 are made by cropping a bounding box of 28% width / length of original dicom around the x/y coordinates predicted from stage 1 and using all slices. Thereby all the x/y predictions for all slices of a series are aggregated using median to increase robustness.</p>\n<p><strong>Study-Level IVD model :</strong> For a single study-level IVD the 3 view 3d-crops (Axial 3d crop, Sagittal T1 3d crop, Sagittal T2 3d crop) are loaded and reshaped to 144x288x9, 144x144x9, 144x144x9 cubes. If a view is not available it is replaced with zeros. The 2 sagittal views are concatenated on the width axis resulting in a 144x288x9 cube and then further concatenate with the axial view cube on the height axis to a single 288x288x9 3d-image. Dropping one of the views with a chance of 10% is used as additional augmentation. <br>\nThe model is of 2.5D nature and has 3 tf_efficientnetv2_s backbones where conv_pw of each InvertedResidual layer has been patched to apply 3d instead of 2d convolution. The 288x288x9 3d-image is fed to each backbone individually to predict neural_foraminal_narrowing, spinal_canal_stenosis, subarticular_stenosis separately using a CrossEntropyLoss. Additionally a loss is predicted for having a severe spinal_canal_stenosis condition in any of the 5 IVDs belonging to a study. The 4 losses are averaged and optimised as a single loss, which represents the competition metric.</p>\n<h2>Model ensembling</h2>\n<p>Ensembling is done in the MLP of the 1st pipeline. Predicted logits of all our models are concatenated together for the three classes. Here are our final scores:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>SCS</th>\n<th>NFN</th>\n<th>SS</th>\n<th>ANY</th>\n<th>CV</th>\n<th></th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loss</td>\n<td>0.260</td>\n<td>0.475</td>\n<td>0.538</td>\n<td>0.255</td>\n<td>0.382</td>\n<td></td>\n<td>0.355</td>\n<td>0.401</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not help</h2>\n<ul>\n<li>SpineNet and other medical specific models for condition level modelling. </li>\n<li>Pseudo labelling on external data</li>\n<li>Denoising techniques</li>\n<li>Encoder-decoder architectures to add an auxiliary injury localization task</li>\n<li>Bi-encoder architectures to jointly learn on axial and sagittal images</li>\n</ul>\n<p><em>Thanks for reading !</em></p>",
  "messages": [
    {
      "id": "3024323",
      "postDate": "10/21/2024 14:09:29",
      "content": "<p>Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about MRI data and how to create strong models for it. <br>\nThis competition was very challenging (and time-consuming !) because it had 3 underlying modalities (SCS, NFN, SS) and 2 tasks (disk localization, severity classification).</p>\n<h2>TLDR</h2>\n<p>Our solution is an ensemble of multiple models which train on study or series level for each individual intervertebral disc level. Crops of the original MRI are isolated in a first stage and the severity of the condition is predicted using the sequence of crops. Final results are aggregated using an MLP model that directly optimizes the competition metric.</p>\n<h2>Cross validation</h2>\n<p>A common cross validation approach was used within the team with four folds. Below is the tracking of CV against leaderboard scores. Our experience from previous RSNA competitions led us to expect a small yet reasonable shake-up, and we trusted our CV more than public LB.</p>\n<p><a href=\"https://ibb.co/LzxRtby\"><img src=\"https://i.ibb.co/yn6PyGC/cvlb.png\" alt=\"cvlb\"></a></p>\n<h2>Data &amp; Augmentations</h2>\n<p>All data was sourced from competition data, except for the spinenet model weights. Some of our models used <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> publicly shared coordinates (thanks a lot!). </p>\n<p>For preprocessing we mostly used min-max normalisation, and sometimes added windowing based on dicom window width and window length. We relied on albumentations for augmentations, and used a variety of approaches for the different pipelines:</p>\n<ul>\n<li><strong>Model 1 :</strong><ul>\n<li>ShiftScaleRotate, ElasticTransform</li>\n<li>RandomGamma, RandomBrightnessContrast</li>\n<li>MotionBlur, GaussianBlur</li>\n<li>MixUp for the Sagittal only models.</li>\n<li>Flipping left / right and the associated targets for some models.</li></ul></li>\n<li><strong>Model 2 :</strong><ul>\n<li>ShiftScaleRotate</li>\n<li>RandomBrightnessContrast</li>\n<li>CoarseDropout</li>\n<li>RandomCrop</li>\n<li>For localisation of crops, augmenting the target coord position was very effective</li>\n<li>Random shift start and end instance number of crop </li></ul></li>\n<li><strong>Model 3 :</strong><ul>\n<li>ShiftScaleRotate</li>\n<li>RandomBrightnessContrast</li>\n<li>CoarseDropout</li>\n<li>RandomCrop</li>\n<li>ChannelDropout</li></ul></li>\n</ul>\n<h2>Models</h2>\n<h3>Model 1</h3>\n<p>The code to this part of the pipeline is available on github</p>\n<blockquote>\n  <p><a href=\"https://github.com/TheoViel/kaggle_rsna_lumbar_spine\" target=\"_blank\">https://github.com/TheoViel/kaggle_rsna_lumbar_spine</a></p>\n</blockquote>\n<p>It achieves ~0.42 private LB, using only the sagittal data.</p>\n<p><a href=\"https://ibb.co/kcm5KVW\"><img src=\"https://i.ibb.co/HghB7Lk/RSNA-spine-theo-drawio.png\" alt=\"RSNA-spine-theo-drawio\"></a></p>\n<p><strong>Coordinates model :</strong> A simple <code>coatnet_rmlp_2_rw_384</code> is trained on all the center frames of the sagittal images to predict the 5 (x, y) coordinates associated with the disk injuries. It is trained with the MSE on 10 classes, and used to generate crops. Crop generation is done by selecting the square of 20% of the image size around the predicted ROI center.</p>\n<p><strong>Classification models :</strong> On the crop above, we once again trained CoAtNets, this time <code>coatnet_1_rw_224</code> and <code>coatnet_2_rw_224</code> and with a RNN layer to incorporate the 3D information. <br>\nWe initially wanted to train separate models for each injury since each image modality had its target. For instance the SCS only models are trained to predict the 3 severity classes on the Sagittal T2 images, by sampling the 5 (or 3) frames at the center of the stack. But handling the two sagittal modalities together worked better, and even showed decent performance on the SS task. Further improvements came from adding more frames (5 is not enough to capture the SCS and SS signal) and using 3 RNN heads that had access to different frames - for the left, right and center (scs) targets. We also added MixUp and trained models for 10 epochs (vs 5 for the SCS models) and with a higher learning rate (1e-3 vs 5e-4 for the SCS models). </p>\n<p><strong>MLP :</strong> The classification models are trained with the CE, and tweaked to maximize the AUC. The MLP model is here to aggregate predictions, and account for the competition metric. Surprisingly, what worked best here is to consider each target independently, i.e. the scs_l1_l2 does not interact with the scs_l2_l3 features nor the nfn_left_l1_l2 features in the MLP. The logits layer weights for the different levels are shared though, i.e. the model consists of 3 MLP (one for SCS, one for SS and one for SCS).</p>\n<h3>Model 2</h3>\n<p>The code to this part of the pipeline is available on github</p>\n<blockquote>\n  <p><a href=\"https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\" target=\"_blank\">https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/</a></p>\n</blockquote>\n<p><a href=\"https://ibb.co/2hCR7GR\"><img src=\"https://i.ibb.co/nmY4sx4/RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx.png\" alt=\"RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx\"></a></p>\n<p>For Sagittal and Axial images we train severity classification on series and individual intervertebral disc level. Therefore we need to isolate each individual disc within the dicom. </p>\n<p><strong>Axial xy-localisation :</strong> We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of <code>efficientnetv2_rw_t</code> with a linear head and L1-loss is used with learning rate 1e-4 for 16 epochs with a batchsize of 16. The dicom images are resized to 384 and images are fed to the model individually. </p>\n<p><strong>Sagittal xy-localisation :</strong> We use spinenet to predict the left top corner point of the intervertebral disc. Many of the spinenet points are shifted incorrectly by one disc. We use the train coordinates file to identify the shifted ones and shift them by one disc. We then retrain the localisation on the corrected spinenet predictions. We exclude from training series where the label is too far from the spinenet prediction. The same training model and procedure is used as axial xy-localisation, except we have an xy label for each vertebrae and we add a mask label to indicate slices with no spinenet xy prediction (in this case, xy-loss is masked).</p>\n<p><strong>Sagittal z-localisation :</strong> We use the train coordinates file to learn the z-location (the annotated instance number). For all series we predict the annotated instance of spinal_canal_stenosis as well as left and right neural_foraminal_narrowing instance number. For sagittal t2, we mask the foraminal loss and for sagittal t1 we loss the spinal canal loss. L1 loss is used where the target is the number of annotations on the instance, between 0 and 5. Same training procedure as above.</p>\n<p><strong>Axial z-localisation :</strong> We leverage sagittal point to axial level mapping from the repo M-Scan repo - code here - to map the xyz position of the sagittal t2 series to the z-position of the patient’s axial series. This identifies which instance number, or slice, contains which intervertebral disc. Inspiration from Ian Pan, Heng and others who shared this approach earlier on in the competition.</p>\n<p><strong>Axial severity classification :</strong> Crops for stage 2 are made by taking the distance of right to left annotation and extending outward either side by half this distance. On the z-position the sequence is cropped by finding the max distance from one level to the next and extending outward either side by this number of slices from the level’s center slice. Therefore we use a variable number of slices per series/level. A 2.5d model of <code>efficientnetv2_rw_t</code> with a bidirectional single layer RNN head (512 dim) and weighted CE loss is used. We use learning rate 4e-4 for 5 epochs with a batch size of 8 series/levels with all crops resized to 256 dim.<br>\nWe use two separate axial models. One to predict the Axial t2 severity, and one to predict the Sagittal t2 severity. The model was not effective in predicting Sagittal t1 severity. </p>\n<p><strong>Sagittal severity classification :</strong> Crops for stage 2 are made by taking the max distance of one vertebrae to the next and extending outward either direction for each level from the levels center xy position. On the z-position the sequence is cropped by excluding slices which were predicted to have no spinenet annotation (as seen above under Sagittal xy-localisation above). Sagittal t1 and t2 labels and series were trained together.  </p>\n<h3>Model 3</h3>\n<p>The code to this part of the pipeline is available on github</p>\n<blockquote>\n  <p><a href=\"https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\" target=\"_blank\">https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/</a></p>\n</blockquote>\n<p><a href=\"https://ibb.co/WxgRc9H\"><img src=\"https://i.ibb.co/j6gjLK8/Screenshot-2024-10-19-at-09-45-26.png\" alt=\"Screenshot-2024-10-19-at-09-45-26\"></a></p>\n<p><strong>Axial xy-localisation :</strong> We use the train coordinates file to learn the xy-location of the labelled points for each slice. A backbone of <code>tf_efficientnetv2_s</code> with a linear head and L1-loss. The dicom images are resized to 384 and images are fed to the model individually. </p>\n<p><strong>Sagittal xy-localisation :</strong> We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of <code>tf_efficientnetv2_s</code> with a linear head and L1-loss is used. The dicom images are resized to 384 and images are fed to the model individually. It was beneficial to train a single model on a combined dataset of T1 and T2 sagittal images </p>\n<p><strong>Axial z-localisation :</strong> Same approach as Model 2. We derived axial z localisation for each vertebrae from T1/ T2 Sagittal xy-localization from same study using 3d coordinates. 3d coordinates are derived separately for T1 and T2 and then averaged. </p>\n<p><strong>Axial Crops :</strong> Multiple axial series for a study are combined by sorting by ImagePositionPatient. Then median of x1/y1 and x2/y2 predictions for all slices are used to create 112x224 sized bounding boxes for a study. For each slice of the axial series a intervertebral disk (IVD) is assigned using the closest derived z coordinates. A maximum distance of 20 was used. So for each IVD a 3d bounding box with height x width of 112x224 and varying number of slices is cropped and saved to disk as a npy file. </p>\n<p><strong>Sagittal Crops :</strong> 3d crops for stage 2 are made by cropping a bounding box of 28% width / length of original dicom around the x/y coordinates predicted from stage 1 and using all slices. Thereby all the x/y predictions for all slices of a series are aggregated using median to increase robustness.</p>\n<p><strong>Study-Level IVD model :</strong> For a single study-level IVD the 3 view 3d-crops (Axial 3d crop, Sagittal T1 3d crop, Sagittal T2 3d crop) are loaded and reshaped to 144x288x9, 144x144x9, 144x144x9 cubes. If a view is not available it is replaced with zeros. The 2 sagittal views are concatenated on the width axis resulting in a 144x288x9 cube and then further concatenate with the axial view cube on the height axis to a single 288x288x9 3d-image. Dropping one of the views with a chance of 10% is used as additional augmentation. <br>\nThe model is of 2.5D nature and has 3 tf_efficientnetv2_s backbones where conv_pw of each InvertedResidual layer has been patched to apply 3d instead of 2d convolution. The 288x288x9 3d-image is fed to each backbone individually to predict neural_foraminal_narrowing, spinal_canal_stenosis, subarticular_stenosis separately using a CrossEntropyLoss. Additionally a loss is predicted for having a severe spinal_canal_stenosis condition in any of the 5 IVDs belonging to a study. The 4 losses are averaged and optimised as a single loss, which represents the competition metric.</p>\n<h2>Model ensembling</h2>\n<p>Ensembling is done in the MLP of the 1st pipeline. Predicted logits of all our models are concatenated together for the three classes. Here are our final scores:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>SCS</th>\n<th>NFN</th>\n<th>SS</th>\n<th>ANY</th>\n<th>CV</th>\n<th></th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loss</td>\n<td>0.260</td>\n<td>0.475</td>\n<td>0.538</td>\n<td>0.255</td>\n<td>0.382</td>\n<td></td>\n<td>0.355</td>\n<td>0.401</td>\n</tr>\n</tbody>\n</table>\n<h2>What did not help</h2>\n<ul>\n<li>SpineNet and other medical specific models for condition level modelling. </li>\n<li>Pseudo labelling on external data</li>\n<li>Denoising techniques</li>\n<li>Encoder-decoder architectures to add an auxiliary injury localization task</li>\n<li>Bi-encoder architectures to jointly learn on axial and sagittal images</li>\n</ul>\n<p><em>Thanks for reading !</em></p>",
      "rawMarkdown": "Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about MRI data and how to create strong models for it. \nThis competition was very challenging (and time-consuming !) because it had 3 underlying modalities (SCS, NFN, SS) and 2 tasks (disk localization, severity classification).\n\n## TLDR\nOur solution is an ensemble of multiple models which train on study or series level for each individual intervertebral disc level. Crops of the original MRI are isolated in a first stage and the severity of the condition is predicted using the sequence of crops. Final results are aggregated using an MLP model that directly optimizes the competition metric.\n\n## Cross validation \n\nA common cross validation approach was used within the team with four folds. Below is the tracking of CV against leaderboard scores. Our experience from previous RSNA competitions led us to expect a small yet reasonable shake-up, and we trusted our CV more than public LB.\n\n<a href=\"https://ibb.co/LzxRtby\"><img src=\"https://i.ibb.co/yn6PyGC/cvlb.png\" alt=\"cvlb\" border=\"0\"></a>\n\n## Data & Augmentations\nAll data was sourced from competition data, except for the spinenet model weights. Some of our models used @brendanartley publicly shared coordinates (thanks a lot!). \n\nFor preprocessing we mostly used min-max normalisation, and sometimes added windowing based on dicom window width and window length. We relied on albumentations for augmentations, and used a variety of approaches for the different pipelines:\n\n- **Model 1 :**\n  - ShiftScaleRotate, ElasticTransform\n  - RandomGamma, RandomBrightnessContrast\n  - MotionBlur, GaussianBlur\n  - MixUp for the Sagittal only models.\n  - Flipping left / right and the associated targets for some models.\n- **Model 2 :**\n  - ShiftScaleRotate\n  - RandomBrightnessContrast\n  - CoarseDropout\n  - RandomCrop\n  - For localisation of crops, augmenting the target coord position was very effective\n  - Random shift start and end instance number of crop \n- **Model 3 :**\n  - ShiftScaleRotate\n  - RandomBrightnessContrast\n  - CoarseDropout\n  - RandomCrop\n  - ChannelDropout\n\n## Models\n\n### Model 1\n\nThe code to this part of the pipeline is available on github\n> https://github.com/TheoViel/kaggle_rsna_lumbar_spine\n\nIt achieves ~0.42 private LB, using only the sagittal data.\n\n<a href=\"https://ibb.co/kcm5KVW\"><img src=\"https://i.ibb.co/HghB7Lk/RSNA-spine-theo-drawio.png\" alt=\"RSNA-spine-theo-drawio\" border=\"0\"></a>\n\n\n**Coordinates model :** A simple `coatnet_rmlp_2_rw_384` is trained on all the center frames of the sagittal images to predict the 5 (x, y) coordinates associated with the disk injuries. It is trained with the MSE on 10 classes, and used to generate crops. Crop generation is done by selecting the square of 20% of the image size around the predicted ROI center.\n\n**Classification models :** On the crop above, we once again trained CoAtNets, this time `coatnet_1_rw_224` and `coatnet_2_rw_224` and with a RNN layer to incorporate the 3D information. \nWe initially wanted to train separate models for each injury since each image modality had its target. For instance the SCS only models are trained to predict the 3 severity classes on the Sagittal T2 images, by sampling the 5 (or 3) frames at the center of the stack. But handling the two sagittal modalities together worked better, and even showed decent performance on the SS task. Further improvements came from adding more frames (5 is not enough to capture the SCS and SS signal) and using 3 RNN heads that had access to different frames - for the left, right and center (scs) targets. We also added MixUp and trained models for 10 epochs (vs 5 for the SCS models) and with a higher learning rate (1e-3 vs 5e-4 for the SCS models). \n\n**MLP :** The classification models are trained with the CE, and tweaked to maximize the AUC. The MLP model is here to aggregate predictions, and account for the competition metric. Surprisingly, what worked best here is to consider each target independently, i.e. the scs_l1_l2 does not interact with the scs_l2_l3 features nor the nfn_left_l1_l2 features in the MLP. The logits layer weights for the different levels are shared though, i.e. the model consists of 3 MLP (one for SCS, one for SS and one for SCS).\n\n### Model 2\n\nThe code to this part of the pipeline is available on github\n> https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\n\n<a href=\"https://ibb.co/2hCR7GR\"><img src=\"https://i.ibb.co/nmY4sx4/RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx.png\" alt=\"RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx\" border=\"0\"></a>\n\nFor Sagittal and Axial images we train severity classification on series and individual intervertebral disc level. Therefore we need to isolate each individual disc within the dicom. \n\n**Axial xy-localisation :** We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of `efficientnetv2_rw_t` with a linear head and L1-loss is used with learning rate 1e-4 for 16 epochs with a batchsize of 16. The dicom images are resized to 384 and images are fed to the model individually. \n\n**Sagittal xy-localisation :** We use spinenet to predict the left top corner point of the intervertebral disc. Many of the spinenet points are shifted incorrectly by one disc. We use the train coordinates file to identify the shifted ones and shift them by one disc. We then retrain the localisation on the corrected spinenet predictions. We exclude from training series where the label is too far from the spinenet prediction. The same training model and procedure is used as axial xy-localisation, except we have an xy label for each vertebrae and we add a mask label to indicate slices with no spinenet xy prediction (in this case, xy-loss is masked).\n\n**Sagittal z-localisation :** We use the train coordinates file to learn the z-location (the annotated instance number). For all series we predict the annotated instance of spinal_canal_stenosis as well as left and right neural_foraminal_narrowing instance number. For sagittal t2, we mask the foraminal loss and for sagittal t1 we loss the spinal canal loss. L1 loss is used where the target is the number of annotations on the instance, between 0 and 5. Same training procedure as above.\n\n**Axial z-localisation :** We leverage sagittal point to axial level mapping from the repo M-Scan repo - code here - to map the xyz position of the sagittal t2 series to the z-position of the patient’s axial series. This identifies which instance number, or slice, contains which intervertebral disc. Inspiration from Ian Pan, Heng and others who shared this approach earlier on in the competition.\n\n**Axial severity classification :** Crops for stage 2 are made by taking the distance of right to left annotation and extending outward either side by half this distance. On the z-position the sequence is cropped by finding the max distance from one level to the next and extending outward either side by this number of slices from the level’s center slice. Therefore we use a variable number of slices per series/level. A 2.5d model of `efficientnetv2_rw_t` with a bidirectional single layer RNN head (512 dim) and weighted CE loss is used. We use learning rate 4e-4 for 5 epochs with a batch size of 8 series/levels with all crops resized to 256 dim.\nWe use two separate axial models. One to predict the Axial t2 severity, and one to predict the Sagittal t2 severity. The model was not effective in predicting Sagittal t1 severity. \n\n**Sagittal severity classification :** Crops for stage 2 are made by taking the max distance of one vertebrae to the next and extending outward either direction for each level from the levels center xy position. On the z-position the sequence is cropped by excluding slices which were predicted to have no spinenet annotation (as seen above under Sagittal xy-localisation above). Sagittal t1 and t2 labels and series were trained together.  \n\n### Model 3 \n\nThe code to this part of the pipeline is available on github\n> https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\n\n<a href=\"https://ibb.co/WxgRc9H\"><img src=\"https://i.ibb.co/j6gjLK8/Screenshot-2024-10-19-at-09-45-26.png\" alt=\"Screenshot-2024-10-19-at-09-45-26\" border=\"0\"></a>\n\n**Axial xy-localisation :** We use the train coordinates file to learn the xy-location of the labelled points for each slice. A backbone of `tf_efficientnetv2_s` with a linear head and L1-loss. The dicom images are resized to 384 and images are fed to the model individually. \n\n**Sagittal xy-localisation :** We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of `tf_efficientnetv2_s` with a linear head and L1-loss is used. The dicom images are resized to 384 and images are fed to the model individually. It was beneficial to train a single model on a combined dataset of T1 and T2 sagittal images \n\n**Axial z-localisation :** Same approach as Model 2. We derived axial z localisation for each vertebrae from T1/ T2 Sagittal xy-localization from same study using 3d coordinates. 3d coordinates are derived separately for T1 and T2 and then averaged. \n\n**Axial Crops :** Multiple axial series for a study are combined by sorting by ImagePositionPatient. Then median of x1/y1 and x2/y2 predictions for all slices are used to create 112x224 sized bounding boxes for a study. For each slice of the axial series a intervertebral disk (IVD) is assigned using the closest derived z coordinates. A maximum distance of 20 was used. So for each IVD a 3d bounding box with height x width of 112x224 and varying number of slices is cropped and saved to disk as a npy file. \n\n**Sagittal Crops :** 3d crops for stage 2 are made by cropping a bounding box of 28% width / length of original dicom around the x/y coordinates predicted from stage 1 and using all slices. Thereby all the x/y predictions for all slices of a series are aggregated using median to increase robustness.\n\n**Study-Level IVD model :** For a single study-level IVD the 3 view 3d-crops (Axial 3d crop, Sagittal T1 3d crop, Sagittal T2 3d crop) are loaded and reshaped to 144x288x9, 144x144x9, 144x144x9 cubes. If a view is not available it is replaced with zeros. The 2 sagittal views are concatenated on the width axis resulting in a 144x288x9 cube and then further concatenate with the axial view cube on the height axis to a single 288x288x9 3d-image. Dropping one of the views with a chance of 10% is used as additional augmentation. \nThe model is of 2.5D nature and has 3 tf_efficientnetv2_s backbones where conv_pw of each InvertedResidual layer has been patched to apply 3d instead of 2d convolution. The 288x288x9 3d-image is fed to each backbone individually to predict neural_foraminal_narrowing, spinal_canal_stenosis, subarticular_stenosis separately using a CrossEntropyLoss. Additionally a loss is predicted for having a severe spinal_canal_stenosis condition in any of the 5 IVDs belonging to a study. The 4 losses are averaged and optimised as a single loss, which represents the competition metric.\n\n\n## Model ensembling\n\nEnsembling is done in the MLP of the 1st pipeline. Predicted logits of all our models are concatenated together for the three classes. Here are our final scores:\n\n|           | SCS  | NFN  | SS   | ANY  | CV   |      |  Public LB | Private LB |\n|---------|------|------|------|------|------|------|------------|------------|\n| Loss  | 0.260| 0.475| 0.538| 0.255| 0.382|      | 0.355      | 0.401      |\n\n## What did not help\n\n- SpineNet and other medical specific models for condition level modelling. \n- Pseudo labelling on external data\n- Denoising techniques\n- Encoder-decoder architectures to add an auxiliary injury localization task\n- Bi-encoder architectures to jointly learn on axial and sagittal images\n\n\n*Thanks for reading !*",
      "votes": null
    },
    {
      "id": "3024986",
      "postDate": "10/22/2024 08:21:35",
      "content": "<p>thank you <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>  and your temates i learn a lot from this competition , this is very informative competition about computer vision in medical images , it is one one of the hardest competition i ever participating in it, i wrote a lot of torch code , debugging , sometimes i am getting lost …, but this RSNA2024 is good preparation for RSNA2025 that i am waiting for , i want to also to thank all the participants that share their notebooks especially <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> </p>",
      "rawMarkdown": "thank you @darraghdog  and your temates i learn a lot from this competition , this is very informative competition about computer vision in medical images , it is one one of the hardest competition i ever participating in it, i wrote a lot of torch code , debugging , sometimes i am getting lost ..., but this RSNA2024 is good preparation for RSNA2025 that i am waiting for , i want to also to thank all the participants that share their notebooks especially @junkoda",
      "votes": null
    },
    {
      "id": "3025002",
      "postDate": "10/22/2024 08:46:05",
      "content": "<p>great. good work</p>",
      "rawMarkdown": "great. good work",
      "votes": null
    },
    {
      "id": "3027295",
      "postDate": "10/24/2024 16:32:33",
      "content": "<p>It's helpful for us!</p>",
      "rawMarkdown": "It's helpful for us!",
      "votes": null
    },
    {
      "id": "3029580",
      "postDate": "10/27/2024 13:27:16",
      "content": "<p>excellent work</p>",
      "rawMarkdown": "excellent work",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3024986,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "10/22/2024 08:21:35",
      "content": "<p>thank you <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>  and your temates i learn a lot from this competition , this is very informative competition about computer vision in medical images , it is one one of the hardest competition i ever participating in it, i wrote a lot of torch code , debugging , sometimes i am getting lost …, but this RSNA2024 is good preparation for RSNA2025 that i am waiting for , i want to also to thank all the participants that share their notebooks especially <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3025002,
      "author_name": "kasai147",
      "author_url": "",
      "post_date": "10/22/2024 08:46:05",
      "content": "<p>great. good work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3027295,
      "author_name": "prashantkumaryt",
      "author_url": "",
      "post_date": "10/24/2024 16:32:33",
      "content": "<p>It's helpful for us!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3029580,
      "author_name": "rahma11",
      "author_url": "",
      "post_date": "10/27/2024 13:27:16",
      "content": "<p>excellent work</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3024323": "Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about MRI data and how to create strong models for it. \nThis competition was very challenging (and time-consuming !) because it had 3 underlying modalities (SCS, NFN, SS) and 2 tasks (disk localization, severity classification).\n\n## TLDR\nOur solution is an ensemble of multiple models which train on study or series level for each individual intervertebral disc level. Crops of the original MRI are isolated in a first stage and the severity of the condition is predicted using the sequence of crops. Final results are aggregated using an MLP model that directly optimizes the competition metric.\n\n## Cross validation \n\nA common cross validation approach was used within the team with four folds. Below is the tracking of CV against leaderboard scores. Our experience from previous RSNA competitions led us to expect a small yet reasonable shake-up, and we trusted our CV more than public LB.\n\n<a href=\"https://ibb.co/LzxRtby\"><img src=\"https://i.ibb.co/yn6PyGC/cvlb.png\" alt=\"cvlb\" border=\"0\"></a>\n\n## Data & Augmentations\nAll data was sourced from competition data, except for the spinenet model weights. Some of our models used @brendanartley publicly shared coordinates (thanks a lot!). \n\nFor preprocessing we mostly used min-max normalisation, and sometimes added windowing based on dicom window width and window length. We relied on albumentations for augmentations, and used a variety of approaches for the different pipelines:\n\n- **Model 1 :**\n  - ShiftScaleRotate, ElasticTransform\n  - RandomGamma, RandomBrightnessContrast\n  - MotionBlur, GaussianBlur\n  - MixUp for the Sagittal only models.\n  - Flipping left / right and the associated targets for some models.\n- **Model 2 :**\n  - ShiftScaleRotate\n  - RandomBrightnessContrast\n  - CoarseDropout\n  - RandomCrop\n  - For localisation of crops, augmenting the target coord position was very effective\n  - Random shift start and end instance number of crop \n- **Model 3 :**\n  - ShiftScaleRotate\n  - RandomBrightnessContrast\n  - CoarseDropout\n  - RandomCrop\n  - ChannelDropout\n\n## Models\n\n### Model 1\n\nThe code to this part of the pipeline is available on github\n> https://github.com/TheoViel/kaggle_rsna_lumbar_spine\n\nIt achieves ~0.42 private LB, using only the sagittal data.\n\n<a href=\"https://ibb.co/kcm5KVW\"><img src=\"https://i.ibb.co/HghB7Lk/RSNA-spine-theo-drawio.png\" alt=\"RSNA-spine-theo-drawio\" border=\"0\"></a>\n\n\n**Coordinates model :** A simple `coatnet_rmlp_2_rw_384` is trained on all the center frames of the sagittal images to predict the 5 (x, y) coordinates associated with the disk injuries. It is trained with the MSE on 10 classes, and used to generate crops. Crop generation is done by selecting the square of 20% of the image size around the predicted ROI center.\n\n**Classification models :** On the crop above, we once again trained CoAtNets, this time `coatnet_1_rw_224` and `coatnet_2_rw_224` and with a RNN layer to incorporate the 3D information. \nWe initially wanted to train separate models for each injury since each image modality had its target. For instance the SCS only models are trained to predict the 3 severity classes on the Sagittal T2 images, by sampling the 5 (or 3) frames at the center of the stack. But handling the two sagittal modalities together worked better, and even showed decent performance on the SS task. Further improvements came from adding more frames (5 is not enough to capture the SCS and SS signal) and using 3 RNN heads that had access to different frames - for the left, right and center (scs) targets. We also added MixUp and trained models for 10 epochs (vs 5 for the SCS models) and with a higher learning rate (1e-3 vs 5e-4 for the SCS models). \n\n**MLP :** The classification models are trained with the CE, and tweaked to maximize the AUC. The MLP model is here to aggregate predictions, and account for the competition metric. Surprisingly, what worked best here is to consider each target independently, i.e. the scs_l1_l2 does not interact with the scs_l2_l3 features nor the nfn_left_l1_l2 features in the MLP. The logits layer weights for the different levels are shared though, i.e. the model consists of 3 MLP (one for SCS, one for SS and one for SCS).\n\n### Model 2\n\nThe code to this part of the pipeline is available on github\n> https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\n\n<a href=\"https://ibb.co/2hCR7GR\"><img src=\"https://i.ibb.co/nmY4sx4/RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx.png\" alt=\"RSNA-2024-RSNA-2024-Lumbar-Spine-Degenerative-Classification-pptx\" border=\"0\"></a>\n\nFor Sagittal and Axial images we train severity classification on series and individual intervertebral disc level. Therefore we need to isolate each individual disc within the dicom. \n\n**Axial xy-localisation :** We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of `efficientnetv2_rw_t` with a linear head and L1-loss is used with learning rate 1e-4 for 16 epochs with a batchsize of 16. The dicom images are resized to 384 and images are fed to the model individually. \n\n**Sagittal xy-localisation :** We use spinenet to predict the left top corner point of the intervertebral disc. Many of the spinenet points are shifted incorrectly by one disc. We use the train coordinates file to identify the shifted ones and shift them by one disc. We then retrain the localisation on the corrected spinenet predictions. We exclude from training series where the label is too far from the spinenet prediction. The same training model and procedure is used as axial xy-localisation, except we have an xy label for each vertebrae and we add a mask label to indicate slices with no spinenet xy prediction (in this case, xy-loss is masked).\n\n**Sagittal z-localisation :** We use the train coordinates file to learn the z-location (the annotated instance number). For all series we predict the annotated instance of spinal_canal_stenosis as well as left and right neural_foraminal_narrowing instance number. For sagittal t2, we mask the foraminal loss and for sagittal t1 we loss the spinal canal loss. L1 loss is used where the target is the number of annotations on the instance, between 0 and 5. Same training procedure as above.\n\n**Axial z-localisation :** We leverage sagittal point to axial level mapping from the repo M-Scan repo - code here - to map the xyz position of the sagittal t2 series to the z-position of the patient’s axial series. This identifies which instance number, or slice, contains which intervertebral disc. Inspiration from Ian Pan, Heng and others who shared this approach earlier on in the competition.\n\n**Axial severity classification :** Crops for stage 2 are made by taking the distance of right to left annotation and extending outward either side by half this distance. On the z-position the sequence is cropped by finding the max distance from one level to the next and extending outward either side by this number of slices from the level’s center slice. Therefore we use a variable number of slices per series/level. A 2.5d model of `efficientnetv2_rw_t` with a bidirectional single layer RNN head (512 dim) and weighted CE loss is used. We use learning rate 4e-4 for 5 epochs with a batch size of 8 series/levels with all crops resized to 256 dim.\nWe use two separate axial models. One to predict the Axial t2 severity, and one to predict the Sagittal t2 severity. The model was not effective in predicting Sagittal t1 severity. \n\n**Sagittal severity classification :** Crops for stage 2 are made by taking the max distance of one vertebrae to the next and extending outward either direction for each level from the levels center xy position. On the z-position the sequence is cropped by excluding slices which were predicted to have no spinenet annotation (as seen above under Sagittal xy-localisation above). Sagittal t1 and t2 labels and series were trained together.  \n\n### Model 3 \n\nThe code to this part of the pipeline is available on github\n> https://github.com/darraghdog/kaggle-rsna-2024-6th-place-solution-model2/\n\n<a href=\"https://ibb.co/WxgRc9H\"><img src=\"https://i.ibb.co/j6gjLK8/Screenshot-2024-10-19-at-09-45-26.png\" alt=\"Screenshot-2024-10-19-at-09-45-26\" border=\"0\"></a>\n\n**Axial xy-localisation :** We use the train coordinates file to learn the xy-location of the labelled points for each slice. A backbone of `tf_efficientnetv2_s` with a linear head and L1-loss. The dicom images are resized to 384 and images are fed to the model individually. \n\n**Sagittal xy-localisation :** We use the train coordinates file to learn the xy-location of the labelled point for each slice. A backbone of `tf_efficientnetv2_s` with a linear head and L1-loss is used. The dicom images are resized to 384 and images are fed to the model individually. It was beneficial to train a single model on a combined dataset of T1 and T2 sagittal images \n\n**Axial z-localisation :** Same approach as Model 2. We derived axial z localisation for each vertebrae from T1/ T2 Sagittal xy-localization from same study using 3d coordinates. 3d coordinates are derived separately for T1 and T2 and then averaged. \n\n**Axial Crops :** Multiple axial series for a study are combined by sorting by ImagePositionPatient. Then median of x1/y1 and x2/y2 predictions for all slices are used to create 112x224 sized bounding boxes for a study. For each slice of the axial series a intervertebral disk (IVD) is assigned using the closest derived z coordinates. A maximum distance of 20 was used. So for each IVD a 3d bounding box with height x width of 112x224 and varying number of slices is cropped and saved to disk as a npy file. \n\n**Sagittal Crops :** 3d crops for stage 2 are made by cropping a bounding box of 28% width / length of original dicom around the x/y coordinates predicted from stage 1 and using all slices. Thereby all the x/y predictions for all slices of a series are aggregated using median to increase robustness.\n\n**Study-Level IVD model :** For a single study-level IVD the 3 view 3d-crops (Axial 3d crop, Sagittal T1 3d crop, Sagittal T2 3d crop) are loaded and reshaped to 144x288x9, 144x144x9, 144x144x9 cubes. If a view is not available it is replaced with zeros. The 2 sagittal views are concatenated on the width axis resulting in a 144x288x9 cube and then further concatenate with the axial view cube on the height axis to a single 288x288x9 3d-image. Dropping one of the views with a chance of 10% is used as additional augmentation. \nThe model is of 2.5D nature and has 3 tf_efficientnetv2_s backbones where conv_pw of each InvertedResidual layer has been patched to apply 3d instead of 2d convolution. The 288x288x9 3d-image is fed to each backbone individually to predict neural_foraminal_narrowing, spinal_canal_stenosis, subarticular_stenosis separately using a CrossEntropyLoss. Additionally a loss is predicted for having a severe spinal_canal_stenosis condition in any of the 5 IVDs belonging to a study. The 4 losses are averaged and optimised as a single loss, which represents the competition metric.\n\n\n## Model ensembling\n\nEnsembling is done in the MLP of the 1st pipeline. Predicted logits of all our models are concatenated together for the three classes. Here are our final scores:\n\n|           | SCS  | NFN  | SS   | ANY  | CV   |      |  Public LB | Private LB |\n|---------|------|------|------|------|------|------|------------|------------|\n| Loss  | 0.260| 0.475| 0.538| 0.255| 0.382|      | 0.355      | 0.401      |\n\n## What did not help\n\n- SpineNet and other medical specific models for condition level modelling. \n- Pseudo labelling on external data\n- Denoising techniques\n- Encoder-decoder architectures to add an auxiliary injury localization task\n- Bi-encoder architectures to jointly learn on axial and sagittal images\n\n\n*Thanks for reading !*",
    "3024986": "thank you @darraghdog  and your temates i learn a lot from this competition , this is very informative competition about computer vision in medical images , it is one one of the hardest competition i ever participating in it, i wrote a lot of torch code , debugging , sometimes i am getting lost ..., but this RSNA2024 is good preparation for RSNA2025 that i am waiting for , i want to also to thank all the participants that share their notebooks especially @junkoda",
    "3025002": "great. good work",
    "3027295": "It's helpful for us!",
    "3029580": "excellent work"
  },
  "source": "meta"
}