{
  "id": 540091,
  "title": "1st place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/540091",
  "author_name": "NANACHI",
  "post_date": "2024-10-12T15:01:47.310000",
  "votes": 105,
  "comment_count": 21,
  "views": 0,
  "content": "<p>First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!</p>\n<p>Furthermore, I'd like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's discussion and notebook were the starting point for my solution, and <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> 's <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">this dataset</a> helped my coordinate prediction models. I was deeply impressed by their contributions to the Kaggle community.</p>\n<p>This is my first solution write-up, so please feel free to leave any comments or suggestions for improvement!</p>\n<h2>Summary</h2>\n<p>My solution is 2 stage approach, creating <code>test_label_coordinates.csv</code> and predicting severity. Furthermore, I separated 1st stage into instance_number prediction and coordinate prediction. Therefore I prepared 3 type of model, instance_number prediction model, coordinate prediction model and severity prediction model. The pipeline is shown in the following figure. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F83a6d286c875d8fb9ed5ff50513cbf11%2Frsna_pipeline_overview.png?generation=1728723901143537&amp;alt=media\" alt=\"pipeline\"></p>\n<h2>1st stage: test_label_coordinates creation</h2>\n<p>In the 1st stage, I use 2 type of models, 3D convolution model and 2D convolution model. These models are very simple, encoder + level-separated heads. </p>\n<h3>instance_number prediction (sagittal)</h3>\n<p>In this part, I used simple 3D ConvNeXt to predict instance_number for each level. Data that is fed into models is just normalized from 0 to 1, sorted by dicom's metadata and padded 32 to depth direction to align shape. Data preprocessing is shown in the following figure (scs example). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Ffe356b60c418b7bcef760bafa9d36210%2Fscs_volume_example.png?generation=1728729008294363&amp;alt=media\" alt=\"scs_volume_example\"></p>\n<p>In training models, I trained models 2 tasks, regression and classification, and I used L1 Loss and Cross Entropy Loss respectively. In the classification task, these heads output (bs, 32) shape logits for each level. In the regression task, these heads output (bs, 3) shape vectors for each level. (bs, 3) shape vector means (x, y, z) and I used z for depth prediction, (x, y) were used auxiliary loss. In the regression task, I normalized coordinate labels 0 to 1 for stabilizing models during training. Concretely, I used label (x', y', z') = (x/width, y/height, z/32). The model architecture is shown in the following image (scs example). I implemented 3D ConvNeXt for this task (to implement 3D ConvNeXt, I referred to <a href=\"https://github.com/FrancescoSaverioZuppichini/ConvNext\" target=\"_blank\">this repo</a>). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Fbdf35cf44c9f268063c9b76d8698be19%2Frsna_instance_number_prediction_model_scs_example.png?generation=1728730136962978&amp;alt=media\" alt=\"instance_number_prediction_scs_example\"></p>\n<p>The results of instance_number prediction models are shown in the following table (sagt2, scs). </p>\n<table>\n<thead>\n<tr>\n<th>model/error</th>\n<th>+-0</th>\n<th>+-1</th>\n<th>+-2</th>\n<th>error&gt;+-2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>cls</td>\n<td>71.08%</td>\n<td>27.04%</td>\n<td>1.43%</td>\n<td>0.44%</td>\n</tr>\n<tr>\n<td>reg</td>\n<td>67.48%</td>\n<td>30.59%</td>\n<td>1.61%</td>\n<td>0.31%</td>\n</tr>\n</tbody>\n</table>\n<p>I ensembled this 2 type of predictions using median for each level (actually I used 5 fold for each task). </p>\n<h3>coordinate prediction(sagittal)</h3>\n<p>In coordinate prediction task, I used 2d encoder + level-separated heads, almost same as instance_number regression model. Data is 3 channel image. The image is picked up using median of instance_number of L1 ~ S1. Then the data processed normalization and reshaping (512x512). Labels are (x', y') = (x/width, y/height)<br>\nfor each level, same as instance_number regression, and also I used L1 loss. The model architecture is shown in the following figure. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F9585ce9ac5e65d46ba4c0bf759d19ba4%2Frsna_coordinate_prediction_model_scs_example.png?generation=1728733391395184&amp;alt=media\" alt=\"coordinate_prediction_model_scs_example\"></p>\n<p>I used ConvNeXt-base and Efficientnet-v2-l for this task. Before I train these models, I trained these models using <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> 's <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">dataset</a>. These pretrained models were slightly better than pretrained models that were trained using imagenet. I ensembled these predictions using mean. </p>\n<h3>instance_number calculation and coordinate prediction (axial)</h3>\n<p>For instance_number prediction of axial, I borrowed <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's method (notebook is <a href=\"https://www.kaggle.com/code/hengck23/2d-to-3d-projection-for-dicom/notebook\" target=\"_blank\">here</a>). Then I predicted coordinates of axial, same as coordinate prediction for sagittal. </p>\n<h2>2nd stage: severity prediction</h2>\n<p>For the 2nd stage, I attempted simple 2.5D model and MIL. 2.5D model can be implemented easily, however, MIL was better than simple 2.5D at final. </p>\n<h2>preprocessing</h2>\n<h3>Cropping method</h3>\n<p>My preprocessing strategy is cropping. For example, I cropped sagt2 image for scs; </p>\n<ol>\n<li>pick up 5 images (center is an image that was assigned instance_number)</li>\n<li>reshape 512x512</li>\n<li>crop images using the coordinate (96 pix left and 32 pix right from coordinate x, 40 pix upper and 40 pix lower from coordinate y)</li>\n</ol>\n<p>After cropping an image, the image can be like the figure below (sagt2 for scs, L1/L2). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F3091793ada3e32fe2c2f4c022e10bf93%2Frsna_sagt2_cropped_image.png?generation=1728738536688017&amp;alt=media\" alt=\"scs_cropped_image\"></p>\n<p>sagt2, sagt1 and axial were cropped for each classification task. The following tables are representing cropping range from (x, y) coordinate. </p>\n<p><strong>for scs</strong></p>\n<table>\n<thead>\n<tr>\n<th>type</th>\n<th>left</th>\n<th>right</th>\n<th>upper</th>\n<th>lower</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>sagt2</td>\n<td>96</td>\n<td>32</td>\n<td>40</td>\n<td>40</td>\n</tr>\n<tr>\n<td>axial</td>\n<td>96</td>\n<td>96</td>\n<td>96</td>\n<td>96</td>\n</tr>\n</tbody>\n</table>\n<p>Note that when I crop images from axial, I picked up left or right  subarticular stenosis coordinate randomly, and for adjusting cropping point, I added +-20 to ss coordinate x. As a result, cropping range can be like the following figure (the example is right ss coordinate x + 20). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F55cb454b535e50eed67dc7e03d3da6f2%2Frsna_ax_for_scs.png?generation=1728739918254610&amp;alt=media\" alt=\"axial_for_scs_cropping\"></p>\n<p><strong>for nfn</strong></p>\n<table>\n<thead>\n<tr>\n<th>type</th>\n<th>left</th>\n<th>right</th>\n<th>upper</th>\n<th>lower</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>sagt1 (both left and right)</td>\n<td>96</td>\n<td>64</td>\n<td>32</td>\n<td>32</td>\n</tr>\n<tr>\n<td>axial (right)</td>\n<td>144</td>\n<td>48</td>\n<td>96</td>\n<td>96</td>\n</tr>\n<tr>\n<td>axial (left)</td>\n<td>48</td>\n<td>144</td>\n<td>96</td>\n<td>96</td>\n</tr>\n</tbody>\n</table>\n<p><strong>for ss</strong></p>\n<table>\n<thead>\n<tr>\n<th>type</th>\n<th>left</th>\n<th>right</th>\n<th>upper</th>\n<th>lower</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>axial (right)</td>\n<td>144</td>\n<td>48</td>\n<td>96</td>\n<td>96</td>\n</tr>\n<tr>\n<td>axial (left)</td>\n<td>48</td>\n<td>144</td>\n<td>96</td>\n<td>96</td>\n</tr>\n</tbody>\n</table>\n<p>The following image is the range of cropping axial for right subarticular stenosis. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F916b910dd75444be91a822e2e19a2f9b%2Frsna_axial_cropping_for_ss.png?generation=1728740665640827&amp;alt=media\" alt=\"axial_cropping_for_ss_right\"></p>\n<h3>data augmentations</h3>\n<p>I used several augmentations like below; </p>\n<p><em>Before cropping</em></p>\n<ul>\n<li>random shift of coordinate x and y (-10~+10 pix)</li>\n<li>random shift of instance_number (-2~+2. shifting probability was decided error probability of each instance_number prediction models)</li>\n</ul>\n<p><em>After cropping</em></p>\n<ul>\n<li>RandomBrightnessContrast(p=0.25)</li>\n<li>ShiftScaleRotate(shift_limit=0.1, scale_limit=(-0.1, 0.1), rotate_limit=20, p=0.5)</li>\n</ul>\n<p>Especially, random shift of instance_number was crucial for robustness of error of 1st stage. </p>\n<h2>model architecture</h2>\n<p>My model architectures are shown in following figures. <br>\n<strong>[EDITED]</strong>  I have updated the figure illustrating the model architecture to correct an error in the previous version. <code>aux_attn_score</code> in the code below is fed into cross entropy loss directly. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1466c45fe85d9cc404d5047150dba7c0%2Frsna_severity_prediction_model_for_scs_fixed.png?generation=1728897262477106&amp;alt=media\" alt=\"severity_prediction_model_scs_fixed\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1c2a33a4eac3252349847fa6723ebd72%2Frsna_severity_prediction_model_for_ss_fixed.png?generation=1728897365864286&amp;alt=media\" alt=\"severity_prediction_model_ss_fixed\"></p>\n<p>I used ConvNeXt-small and Efficientnet-v2-s as the encoder. After implementing Attention-based MIL, my public LB score was improved from 0.37 -&gt; 0.35. Then, adding bi-LSTM, aux losses and ensembling improve my score from 0.35 to 0.33. bi-LSTM + Attention-based MIL was implemented like below. </p>\n<pre><code> (nn.Module):\n     ():\n        (LSTMMIL, ).__init__()\n        .lstm = nn.LSTM(input_dim, input_dim//, num_layers=, batch_first=, dropout=, bidirectional=)\n        .aux_attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, )\n        )\n        .attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, )\n        )\n     ():\n        batch_size, num_instances, input_dim = bags.size()\n        bags_lstm, _ = .lstm(bags)\n        attn_scores = .attention(bags_lstm).squeeze(-)\n        aux_attn_scores = .aux_attention(bags_lstm).squeeze(-)\n        attn_weights = torch.softmax(attn_scores, dim=-)\n        weighted_instances = torch.bmm(attn_weights.unsqueeze(), bags_lstm).squeeze()\n\n         weighted_instances, aux_attn_scores\n</code></pre>\n<h2>what didn't work</h2>\n<ul>\n<li>MAMBA and Self-Attention instead of bi-LSTM</li>\n<li>sharing weight between aux_attention layer and attention layer</li>\n<li>sagt1 image for scs, sagt1 and sagt2 image for ss, sagt2 image for nfn</li>\n<li>long epochs (I used 7 epochs for convnext-small and 14 epochs for efficientnet-v2-s)</li>\n<li>large models (convnext-large &lt; convnext-base &lt; convnext-small in my experiments)</li>\n<li>vision transformers (I think this was my problem. but convolution models were better than vits in my experiments)</li>\n</ul>\n<h2>code</h2>\n<p>All training code is implemented in google colaboratory. All models are used for <a href=\"https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook\" target=\"_blank\">this inference code</a>. Following links are pairs of model name &amp; training notebook link. You can check these model name in the <a href=\"https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook\" target=\"_blank\">inference code</a>. </p>\n<p>You can train on google colaboratory environment with T4 + high memory. </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/wadakoki/rsna-spine-final-models/data\" target=\"_blank\">models</a></li>\n</ul>\n<h3>instance number prediction models (SCS)</h3>\n<ul>\n<li>scs_depth_1024_ssr: <a href=\"https://colab.research.google.com/drive/1JbSFgIwxlviyXb6uHv4vfbyqyjbdCRkw?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>scs_depth: <a href=\"https://colab.research.google.com/drive/19YylaxYLYk1q6IOHpfi9UMnhbNUQBYML?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>scs_depth_1024_ssr_l1: <a href=\"https://colab.research.google.com/drive/11fV56U5hPL2IiRzxaiLmuwyjCrygWsgO?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>instance number prediction models (NFN)</h3>\n<ul>\n<li>nfn_depth_1024_ssr: <a href=\"https://colab.research.google.com/drive/1sIU9Aun1S1vla_W-4fZ24tGcn-IFD0a_?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_depth: <a href=\"https://colab.research.google.com/drive/1EPcR7F5p2SvcgpaJdNKw-0vYyc7eYjwr?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_depth_1024_ssr_l1: <a href=\"https://colab.research.google.com/drive/1Gk4Db4tjhxUEL3uSLiRVlG1MeeGx6K6l?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate prediction models (SCS)</h3>\n<ul>\n<li>scs_detect_pre: <a href=\"https://colab.research.google.com/drive/1qIXQRLkLFyXzyvP9jA2gaX6YZ4_qta27?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>scs_detect_pre_effv2l: <a href=\"https://colab.research.google.com/drive/18vB2qrrBxC7Q4dwVQDR-ioPOY46oEbfu?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate prediction models (NFN)</h3>\n<ul>\n<li>nfn_detect_pre: <a href=\"https://colab.research.google.com/drive/1IPiJDgPDXxOqNTbppZzM89n3ZuPKeWiV?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_detect_pre_effv2l: <a href=\"https://colab.research.google.com/drive/1eKkZFqKUWZIacJrJYGM1PswYYSi3ztjx?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate prediction models (SS)</h3>\n<ul>\n<li>ss_detect: <a href=\"https://colab.research.google.com/drive/1J3Pj8RMbDm5mG0vvBztyrSRK4G8-NLnU?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>severity prediction models (SCS)</h3>\n<ul>\n<li>_scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_convnext-s_for_exp: <a href=\"https://colab.research.google.com/drive/1dWJUGhubs067mJ0GaIOZn8xUt_-1507_?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>_scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_effv2s_for_exp: <a href=\"https://colab.research.google.com/drive/1SsqZOCv7eSYZfqcu5V6ufsHbPp3yN94X?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>severity prediction models (NFN)</h3>\n<ul>\n<li>nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_convnext-s: <a href=\"https://colab.research.google.com/drive/1GtC4vtGVo2sY1Mr6cZf7J1IltO2poAEW?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_effv2s: <a href=\"https://colab.research.google.com/drive/1PJ5qU6szTPagEWCMXBYZ3WZPzK5L6_SJ?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>severity prediction models (SS)</h3>\n<ul>\n<li>ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_effv2s: <a href=\"https://colab.research.google.com/drive/1CUmMoQeUy2ataJoDffEHTe338d66Sl3A?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_convnext-s: <a href=\"https://colab.research.google.com/drive/1KnagfSmFJ69HfASLCfzCOCrZBZyL3y7-?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate pretrained models</h3>\n<p><a href=\"https://colab.research.google.com/drive/1TdmZCT86dFTiP2vd2uMPtNMaAsJKPkc2?usp=sharing\" target=\"_blank\">this notebook</a> is pre-training code for coordinate prediction models with <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">this dataset</a>. model checkpoints are in <a href=\"https://www.kaggle.com/datasets/wadakoki/rsna-pretrained-models-for-coordinate/data\" target=\"_blank\">this dataset</a></p>",
  "messages": [
    {
      "id": 3015549,
      "postDate": "2024-10-12T15:01:47.310Z",
      "content": "<p>First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!</p>\n<p>Furthermore, I'd like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's discussion and notebook were the starting point for my solution, and <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> 's <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">this dataset</a> helped my coordinate prediction models. I was deeply impressed by their contributions to the Kaggle community.</p>\n<p>This is my first solution write-up, so please feel free to leave any comments or suggestions for improvement!</p>\n<h2>Summary</h2>\n<p>My solution is 2 stage approach, creating <code>test_label_coordinates.csv</code> and predicting severity. Furthermore, I separated 1st stage into instance_number prediction and coordinate prediction. Therefore I prepared 3 type of model, instance_number prediction model, coordinate prediction model and severity prediction model. The pipeline is shown in the following figure. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F83a6d286c875d8fb9ed5ff50513cbf11%2Frsna_pipeline_overview.png?generation=1728723901143537&amp;alt=media\" alt=\"pipeline\"></p>\n<h2>1st stage: test_label_coordinates creation</h2>\n<p>In the 1st stage, I use 2 type of models, 3D convolution model and 2D convolution model. These models are very simple, encoder + level-separated heads. </p>\n<h3>instance_number prediction (sagittal)</h3>\n<p>In this part, I used simple 3D ConvNeXt to predict instance_number for each level. Data that is fed into models is just normalized from 0 to 1, sorted by dicom's metadata and padded 32 to depth direction to align shape. Data preprocessing is shown in the following figure (scs example). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Ffe356b60c418b7bcef760bafa9d36210%2Fscs_volume_example.png?generation=1728729008294363&amp;alt=media\" alt=\"scs_volume_example\"></p>\n<p>In training models, I trained models 2 tasks, regression and classification, and I used L1 Loss and Cross Entropy Loss respectively. In the classification task, these heads output (bs, 32) shape logits for each level. In the regression task, these heads output (bs, 3) shape vectors for each level. (bs, 3) shape vector means (x, y, z) and I used z for depth prediction, (x, y) were used auxiliary loss. In the regression task, I normalized coordinate labels 0 to 1 for stabilizing models during training. Concretely, I used label (x', y', z') = (x/width, y/height, z/32). The model architecture is shown in the following image (scs example). I implemented 3D ConvNeXt for this task (to implement 3D ConvNeXt, I referred to <a href=\"https://github.com/FrancescoSaverioZuppichini/ConvNext\" target=\"_blank\">this repo</a>). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Fbdf35cf44c9f268063c9b76d8698be19%2Frsna_instance_number_prediction_model_scs_example.png?generation=1728730136962978&amp;alt=media\" alt=\"instance_number_prediction_scs_example\"></p>\n<p>The results of instance_number prediction models are shown in the following table (sagt2, scs). </p>\n<table>\n<thead>\n<tr>\n<th>model/error</th>\n<th>+-0</th>\n<th>+-1</th>\n<th>+-2</th>\n<th>error&gt;+-2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>cls</td>\n<td>71.08%</td>\n<td>27.04%</td>\n<td>1.43%</td>\n<td>0.44%</td>\n</tr>\n<tr>\n<td>reg</td>\n<td>67.48%</td>\n<td>30.59%</td>\n<td>1.61%</td>\n<td>0.31%</td>\n</tr>\n</tbody>\n</table>\n<p>I ensembled this 2 type of predictions using median for each level (actually I used 5 fold for each task). </p>\n<h3>coordinate prediction(sagittal)</h3>\n<p>In coordinate prediction task, I used 2d encoder + level-separated heads, almost same as instance_number regression model. Data is 3 channel image. The image is picked up using median of instance_number of L1 ~ S1. Then the data processed normalization and reshaping (512x512). Labels are (x', y') = (x/width, y/height)<br>\nfor each level, same as instance_number regression, and also I used L1 loss. The model architecture is shown in the following figure. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F9585ce9ac5e65d46ba4c0bf759d19ba4%2Frsna_coordinate_prediction_model_scs_example.png?generation=1728733391395184&amp;alt=media\" alt=\"coordinate_prediction_model_scs_example\"></p>\n<p>I used ConvNeXt-base and Efficientnet-v2-l for this task. Before I train these models, I trained these models using <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> 's <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">dataset</a>. These pretrained models were slightly better than pretrained models that were trained using imagenet. I ensembled these predictions using mean. </p>\n<h3>instance_number calculation and coordinate prediction (axial)</h3>\n<p>For instance_number prediction of axial, I borrowed <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's method (notebook is <a href=\"https://www.kaggle.com/code/hengck23/2d-to-3d-projection-for-dicom/notebook\" target=\"_blank\">here</a>). Then I predicted coordinates of axial, same as coordinate prediction for sagittal. </p>\n<h2>2nd stage: severity prediction</h2>\n<p>For the 2nd stage, I attempted simple 2.5D model and MIL. 2.5D model can be implemented easily, however, MIL was better than simple 2.5D at final. </p>\n<h2>preprocessing</h2>\n<h3>Cropping method</h3>\n<p>My preprocessing strategy is cropping. For example, I cropped sagt2 image for scs; </p>\n<ol>\n<li>pick up 5 images (center is an image that was assigned instance_number)</li>\n<li>reshape 512x512</li>\n<li>crop images using the coordinate (96 pix left and 32 pix right from coordinate x, 40 pix upper and 40 pix lower from coordinate y)</li>\n</ol>\n<p>After cropping an image, the image can be like the figure below (sagt2 for scs, L1/L2). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F3091793ada3e32fe2c2f4c022e10bf93%2Frsna_sagt2_cropped_image.png?generation=1728738536688017&amp;alt=media\" alt=\"scs_cropped_image\"></p>\n<p>sagt2, sagt1 and axial were cropped for each classification task. The following tables are representing cropping range from (x, y) coordinate. </p>\n<p><strong>for scs</strong></p>\n<table>\n<thead>\n<tr>\n<th>type</th>\n<th>left</th>\n<th>right</th>\n<th>upper</th>\n<th>lower</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>sagt2</td>\n<td>96</td>\n<td>32</td>\n<td>40</td>\n<td>40</td>\n</tr>\n<tr>\n<td>axial</td>\n<td>96</td>\n<td>96</td>\n<td>96</td>\n<td>96</td>\n</tr>\n</tbody>\n</table>\n<p>Note that when I crop images from axial, I picked up left or right  subarticular stenosis coordinate randomly, and for adjusting cropping point, I added +-20 to ss coordinate x. As a result, cropping range can be like the following figure (the example is right ss coordinate x + 20). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F55cb454b535e50eed67dc7e03d3da6f2%2Frsna_ax_for_scs.png?generation=1728739918254610&amp;alt=media\" alt=\"axial_for_scs_cropping\"></p>\n<p><strong>for nfn</strong></p>\n<table>\n<thead>\n<tr>\n<th>type</th>\n<th>left</th>\n<th>right</th>\n<th>upper</th>\n<th>lower</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>sagt1 (both left and right)</td>\n<td>96</td>\n<td>64</td>\n<td>32</td>\n<td>32</td>\n</tr>\n<tr>\n<td>axial (right)</td>\n<td>144</td>\n<td>48</td>\n<td>96</td>\n<td>96</td>\n</tr>\n<tr>\n<td>axial (left)</td>\n<td>48</td>\n<td>144</td>\n<td>96</td>\n<td>96</td>\n</tr>\n</tbody>\n</table>\n<p><strong>for ss</strong></p>\n<table>\n<thead>\n<tr>\n<th>type</th>\n<th>left</th>\n<th>right</th>\n<th>upper</th>\n<th>lower</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>axial (right)</td>\n<td>144</td>\n<td>48</td>\n<td>96</td>\n<td>96</td>\n</tr>\n<tr>\n<td>axial (left)</td>\n<td>48</td>\n<td>144</td>\n<td>96</td>\n<td>96</td>\n</tr>\n</tbody>\n</table>\n<p>The following image is the range of cropping axial for right subarticular stenosis. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F916b910dd75444be91a822e2e19a2f9b%2Frsna_axial_cropping_for_ss.png?generation=1728740665640827&amp;alt=media\" alt=\"axial_cropping_for_ss_right\"></p>\n<h3>data augmentations</h3>\n<p>I used several augmentations like below; </p>\n<p><em>Before cropping</em></p>\n<ul>\n<li>random shift of coordinate x and y (-10~+10 pix)</li>\n<li>random shift of instance_number (-2~+2. shifting probability was decided error probability of each instance_number prediction models)</li>\n</ul>\n<p><em>After cropping</em></p>\n<ul>\n<li>RandomBrightnessContrast(p=0.25)</li>\n<li>ShiftScaleRotate(shift_limit=0.1, scale_limit=(-0.1, 0.1), rotate_limit=20, p=0.5)</li>\n</ul>\n<p>Especially, random shift of instance_number was crucial for robustness of error of 1st stage. </p>\n<h2>model architecture</h2>\n<p>My model architectures are shown in following figures. <br>\n<strong>[EDITED]</strong>  I have updated the figure illustrating the model architecture to correct an error in the previous version. <code>aux_attn_score</code> in the code below is fed into cross entropy loss directly. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1466c45fe85d9cc404d5047150dba7c0%2Frsna_severity_prediction_model_for_scs_fixed.png?generation=1728897262477106&amp;alt=media\" alt=\"severity_prediction_model_scs_fixed\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1c2a33a4eac3252349847fa6723ebd72%2Frsna_severity_prediction_model_for_ss_fixed.png?generation=1728897365864286&amp;alt=media\" alt=\"severity_prediction_model_ss_fixed\"></p>\n<p>I used ConvNeXt-small and Efficientnet-v2-s as the encoder. After implementing Attention-based MIL, my public LB score was improved from 0.37 -&gt; 0.35. Then, adding bi-LSTM, aux losses and ensembling improve my score from 0.35 to 0.33. bi-LSTM + Attention-based MIL was implemented like below. </p>\n<pre><code> (nn.Module):\n     ():\n        (LSTMMIL, ).__init__()\n        .lstm = nn.LSTM(input_dim, input_dim//, num_layers=, batch_first=, dropout=, bidirectional=)\n        .aux_attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, )\n        )\n        .attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, )\n        )\n     ():\n        batch_size, num_instances, input_dim = bags.size()\n        bags_lstm, _ = .lstm(bags)\n        attn_scores = .attention(bags_lstm).squeeze(-)\n        aux_attn_scores = .aux_attention(bags_lstm).squeeze(-)\n        attn_weights = torch.softmax(attn_scores, dim=-)\n        weighted_instances = torch.bmm(attn_weights.unsqueeze(), bags_lstm).squeeze()\n\n         weighted_instances, aux_attn_scores\n</code></pre>\n<h2>what didn't work</h2>\n<ul>\n<li>MAMBA and Self-Attention instead of bi-LSTM</li>\n<li>sharing weight between aux_attention layer and attention layer</li>\n<li>sagt1 image for scs, sagt1 and sagt2 image for ss, sagt2 image for nfn</li>\n<li>long epochs (I used 7 epochs for convnext-small and 14 epochs for efficientnet-v2-s)</li>\n<li>large models (convnext-large &lt; convnext-base &lt; convnext-small in my experiments)</li>\n<li>vision transformers (I think this was my problem. but convolution models were better than vits in my experiments)</li>\n</ul>\n<h2>code</h2>\n<p>All training code is implemented in google colaboratory. All models are used for <a href=\"https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook\" target=\"_blank\">this inference code</a>. Following links are pairs of model name &amp; training notebook link. You can check these model name in the <a href=\"https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook\" target=\"_blank\">inference code</a>. </p>\n<p>You can train on google colaboratory environment with T4 + high memory. </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/wadakoki/rsna-spine-final-models/data\" target=\"_blank\">models</a></li>\n</ul>\n<h3>instance number prediction models (SCS)</h3>\n<ul>\n<li>scs_depth_1024_ssr: <a href=\"https://colab.research.google.com/drive/1JbSFgIwxlviyXb6uHv4vfbyqyjbdCRkw?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>scs_depth: <a href=\"https://colab.research.google.com/drive/19YylaxYLYk1q6IOHpfi9UMnhbNUQBYML?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>scs_depth_1024_ssr_l1: <a href=\"https://colab.research.google.com/drive/11fV56U5hPL2IiRzxaiLmuwyjCrygWsgO?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>instance number prediction models (NFN)</h3>\n<ul>\n<li>nfn_depth_1024_ssr: <a href=\"https://colab.research.google.com/drive/1sIU9Aun1S1vla_W-4fZ24tGcn-IFD0a_?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_depth: <a href=\"https://colab.research.google.com/drive/1EPcR7F5p2SvcgpaJdNKw-0vYyc7eYjwr?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_depth_1024_ssr_l1: <a href=\"https://colab.research.google.com/drive/1Gk4Db4tjhxUEL3uSLiRVlG1MeeGx6K6l?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate prediction models (SCS)</h3>\n<ul>\n<li>scs_detect_pre: <a href=\"https://colab.research.google.com/drive/1qIXQRLkLFyXzyvP9jA2gaX6YZ4_qta27?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>scs_detect_pre_effv2l: <a href=\"https://colab.research.google.com/drive/18vB2qrrBxC7Q4dwVQDR-ioPOY46oEbfu?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate prediction models (NFN)</h3>\n<ul>\n<li>nfn_detect_pre: <a href=\"https://colab.research.google.com/drive/1IPiJDgPDXxOqNTbppZzM89n3ZuPKeWiV?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_detect_pre_effv2l: <a href=\"https://colab.research.google.com/drive/1eKkZFqKUWZIacJrJYGM1PswYYSi3ztjx?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate prediction models (SS)</h3>\n<ul>\n<li>ss_detect: <a href=\"https://colab.research.google.com/drive/1J3Pj8RMbDm5mG0vvBztyrSRK4G8-NLnU?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>severity prediction models (SCS)</h3>\n<ul>\n<li>_scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_convnext-s_for_exp: <a href=\"https://colab.research.google.com/drive/1dWJUGhubs067mJ0GaIOZn8xUt_-1507_?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>_scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_effv2s_for_exp: <a href=\"https://colab.research.google.com/drive/1SsqZOCv7eSYZfqcu5V6ufsHbPp3yN94X?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>severity prediction models (NFN)</h3>\n<ul>\n<li>nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_convnext-s: <a href=\"https://colab.research.google.com/drive/1GtC4vtGVo2sY1Mr6cZf7J1IltO2poAEW?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_effv2s: <a href=\"https://colab.research.google.com/drive/1PJ5qU6szTPagEWCMXBYZ3WZPzK5L6_SJ?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>severity prediction models (SS)</h3>\n<ul>\n<li>ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_effv2s: <a href=\"https://colab.research.google.com/drive/1CUmMoQeUy2ataJoDffEHTe338d66Sl3A?usp=sharing\" target=\"_blank\">notebook</a></li>\n<li>ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_convnext-s: <a href=\"https://colab.research.google.com/drive/1KnagfSmFJ69HfASLCfzCOCrZBZyL3y7-?usp=sharing\" target=\"_blank\">notebook</a></li>\n</ul>\n<h3>coordinate pretrained models</h3>\n<p><a href=\"https://colab.research.google.com/drive/1TdmZCT86dFTiP2vd2uMPtNMaAsJKPkc2?usp=sharing\" target=\"_blank\">this notebook</a> is pre-training code for coordinate prediction models with <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500\" target=\"_blank\">this dataset</a>. model checkpoints are in <a href=\"https://www.kaggle.com/datasets/wadakoki/rsna-pretrained-models-for-coordinate/data\" target=\"_blank\">this dataset</a></p>",
      "rawMarkdown": "First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!\n\nFurthermore, I'd like to thank @hengck23 and @brendanartley. @hengck23 's discussion and notebook were the starting point for my solution, and @brendanartley 's [this dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500) helped my coordinate prediction models. I was deeply impressed by their contributions to the Kaggle community.\n\nThis is my first solution write-up, so please feel free to leave any comments or suggestions for improvement!\n\n## Summary\n\nMy solution is 2 stage approach, creating `test_label_coordinates.csv` and predicting severity. Furthermore, I separated 1st stage into instance_number prediction and coordinate prediction. Therefore I prepared 3 type of model, instance_number prediction model, coordinate prediction model and severity prediction model. The pipeline is shown in the following figure. \n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F83a6d286c875d8fb9ed5ff50513cbf11%2Frsna_pipeline_overview.png?generation=1728723901143537&alt=media)\n\n## 1st stage: test_label_coordinates creation\n\nIn the 1st stage, I use 2 type of models, 3D convolution model and 2D convolution model. These models are very simple, encoder + level-separated heads. \n\n### instance_number prediction (sagittal)\n\nIn this part, I used simple 3D ConvNeXt to predict instance_number for each level. Data that is fed into models is just normalized from 0 to 1, sorted by dicom's metadata and padded 32 to depth direction to align shape. Data preprocessing is shown in the following figure (scs example). \n\n![scs_volume_example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Ffe356b60c418b7bcef760bafa9d36210%2Fscs_volume_example.png?generation=1728729008294363&alt=media)\n\nIn training models, I trained models 2 tasks, regression and classification, and I used L1 Loss and Cross Entropy Loss respectively. In the classification task, these heads output (bs, 32) shape logits for each level. In the regression task, these heads output (bs, 3) shape vectors for each level. (bs, 3) shape vector means (x, y, z) and I used z for depth prediction, (x, y) were used auxiliary loss. In the regression task, I normalized coordinate labels 0 to 1 for stabilizing models during training. Concretely, I used label (x', y', z') = (x/width, y/height, z/32). The model architecture is shown in the following image (scs example). I implemented 3D ConvNeXt for this task (to implement 3D ConvNeXt, I referred to [this repo](https://github.com/FrancescoSaverioZuppichini/ConvNext)). \n\n![instance_number_prediction_scs_example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Fbdf35cf44c9f268063c9b76d8698be19%2Frsna_instance_number_prediction_model_scs_example.png?generation=1728730136962978&alt=media)\n\nThe results of instance_number prediction models are shown in the following table (sagt2, scs). \n\n| model/error | +-0 | +-1 | +-2 | error>+-2 | \n| --- | --- | --- | --- | --- |\n|cls| 71.08% | 27.04% | 1.43% | 0.44% |\n|reg| 67.48% | 30.59% | 1.61% | 0.31% |\n\n\nI ensembled this 2 type of predictions using median for each level (actually I used 5 fold for each task). \n\n### coordinate prediction(sagittal)\n\nIn coordinate prediction task, I used 2d encoder + level-separated heads, almost same as instance_number regression model. Data is 3 channel image. The image is picked up using median of instance_number of L1 ~ S1. Then the data processed normalization and reshaping (512x512). Labels are (x', y') = (x/width, y/height)\nfor each level, same as instance_number regression, and also I used L1 loss. The model architecture is shown in the following figure. \n\n![coordinate_prediction_model_scs_example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F9585ce9ac5e65d46ba4c0bf759d19ba4%2Frsna_coordinate_prediction_model_scs_example.png?generation=1728733391395184&alt=media)\n\nI used ConvNeXt-base and Efficientnet-v2-l for this task. Before I train these models, I trained these models using @brendanartley 's [dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500). These pretrained models were slightly better than pretrained models that were trained using imagenet. I ensembled these predictions using mean. \n\n### instance_number calculation and coordinate prediction (axial)\n\nFor instance_number prediction of axial, I borrowed @hengck23 's method (notebook is [here](https://www.kaggle.com/code/hengck23/2d-to-3d-projection-for-dicom/notebook)). Then I predicted coordinates of axial, same as coordinate prediction for sagittal. \n\n## 2nd stage: severity prediction\n\nFor the 2nd stage, I attempted simple 2.5D model and MIL. 2.5D model can be implemented easily, however, MIL was better than simple 2.5D at final. \n\n## preprocessing\n\n### Cropping method\n\nMy preprocessing strategy is cropping. For example, I cropped sagt2 image for scs; \n\n1. pick up 5 images (center is an image that was assigned instance_number)\n2. reshape 512x512\n3. crop images using the coordinate (96 pix left and 32 pix right from coordinate x, 40 pix upper and 40 pix lower from coordinate y)\n\nAfter cropping an image, the image can be like the figure below (sagt2 for scs, L1/L2). \n\n![scs_cropped_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F3091793ada3e32fe2c2f4c022e10bf93%2Frsna_sagt2_cropped_image.png?generation=1728738536688017&alt=media)\n\nsagt2, sagt1 and axial were cropped for each classification task. The following tables are representing cropping range from (x, y) coordinate. \n\n**for scs**\n\n| type | left | right | upper | lower |\n| --- | --- | --- | --- | --- |\n| sagt2 | 96 | 32 | 40 | 40 |\n| axial | 96 | 96 | 96 | 96 |\n\nNote that when I crop images from axial, I picked up left or right  subarticular stenosis coordinate randomly, and for adjusting cropping point, I added +-20 to ss coordinate x. As a result, cropping range can be like the following figure (the example is right ss coordinate x + 20). \n\n![axial_for_scs_cropping](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F55cb454b535e50eed67dc7e03d3da6f2%2Frsna_ax_for_scs.png?generation=1728739918254610&alt=media)\n\n**for nfn**\n\n| type | left | right | upper | lower |\n| --- | --- | --- | --- | --- |\n| sagt1 (both left and right)| 96 | 64 | 32 | 32 |\n| axial (right) | 144 | 48 | 96 | 96 |\n| axial (left) | 48 | 144 | 96 | 96 |\n\n**for ss**\n\n| type | left | right | upper | lower |\n| --- | --- | --- | --- | --- |\n| axial (right) | 144 | 48 | 96 | 96 |\n| axial (left) | 48 | 144 | 96 | 96 |\n\nThe following image is the range of cropping axial for right subarticular stenosis. \n\n![axial_cropping_for_ss_right](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F916b910dd75444be91a822e2e19a2f9b%2Frsna_axial_cropping_for_ss.png?generation=1728740665640827&alt=media)\n\n### data augmentations\n\nI used several augmentations like below; \n\n*Before cropping*\n\n* random shift of coordinate x and y (-10~+10 pix)\n* random shift of instance_number (-2~+2. shifting probability was decided error probability of each instance_number prediction models)\n\n*After cropping*\n\n* RandomBrightnessContrast(p=0.25)\n* ShiftScaleRotate(shift_limit=0.1, scale_limit=(-0.1, 0.1), rotate_limit=20, p=0.5)\n\nEspecially, random shift of instance_number was crucial for robustness of error of 1st stage. \n\n\n## model architecture\n\nMy model architectures are shown in following figures. \n**[EDITED]**  I have updated the figure illustrating the model architecture to correct an error in the previous version. `aux_attn_score` in the code below is fed into cross entropy loss directly. \n\n![severity_prediction_model_scs_fixed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1466c45fe85d9cc404d5047150dba7c0%2Frsna_severity_prediction_model_for_scs_fixed.png?generation=1728897262477106&alt=media)\n![severity_prediction_model_ss_fixed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1c2a33a4eac3252349847fa6723ebd72%2Frsna_severity_prediction_model_for_ss_fixed.png?generation=1728897365864286&alt=media)\n\nI used ConvNeXt-small and Efficientnet-v2-s as the encoder. After implementing Attention-based MIL, my public LB score was improved from 0.37 -> 0.35. Then, adding bi-LSTM, aux losses and ensembling improve my score from 0.35 to 0.33. bi-LSTM + Attention-based MIL was implemented like below. \n\n```python\nclass LSTMMIL(nn.Module):\n    def __init__(self, input_dim):\n        super(LSTMMIL, self).__init__()\n        self.lstm = nn.LSTM(input_dim, input_dim//2, num_layers=2, batch_first=True, dropout=0.1, bidirectional=True)\n        self.aux_attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, 1)\n        )\n        self.attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, 1)\n        )\n    def forward(self, bags):\n        batch_size, num_instances, input_dim = bags.size()\n        bags_lstm, _ = self.lstm(bags)\n        attn_scores = self.attention(bags_lstm).squeeze(-1)\n        aux_attn_scores = self.aux_attention(bags_lstm).squeeze(-1)\n        attn_weights = torch.softmax(attn_scores, dim=-1)\n        weighted_instances = torch.bmm(attn_weights.unsqueeze(1), bags_lstm).squeeze(1)\n\n        return weighted_instances, aux_attn_scores\n```\n\n## what didn't work\n\n* MAMBA and Self-Attention instead of bi-LSTM\n* sharing weight between aux_attention layer and attention layer\n* sagt1 image for scs, sagt1 and sagt2 image for ss, sagt2 image for nfn\n* long epochs (I used 7 epochs for convnext-small and 14 epochs for efficientnet-v2-s)\n* large models (convnext-large < convnext-base < convnext-small in my experiments)\n* vision transformers (I think this was my problem. but convolution models were better than vits in my experiments)\n\n## code\n\nAll training code is implemented in google colaboratory. All models are used for [this inference code](https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook). Following links are pairs of model name & training notebook link. You can check these model name in the [inference code](https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook). \n\nYou can train on google colaboratory environment with T4 + high memory. \n\n- [models](https://www.kaggle.com/datasets/wadakoki/rsna-spine-final-models/data)\n\n### instance number prediction models (SCS)\n\n- scs_depth_1024_ssr: [notebook](https://colab.research.google.com/drive/1JbSFgIwxlviyXb6uHv4vfbyqyjbdCRkw?usp=sharing)\n- scs_depth: [notebook](https://colab.research.google.com/drive/19YylaxYLYk1q6IOHpfi9UMnhbNUQBYML?usp=sharing)\n- scs_depth_1024_ssr_l1: [notebook](https://colab.research.google.com/drive/11fV56U5hPL2IiRzxaiLmuwyjCrygWsgO?usp=sharing)\n\n### instance number prediction models (NFN)\n- nfn_depth_1024_ssr: [notebook](https://colab.research.google.com/drive/1sIU9Aun1S1vla_W-4fZ24tGcn-IFD0a_?usp=sharing)\n- nfn_depth: [notebook](https://colab.research.google.com/drive/1EPcR7F5p2SvcgpaJdNKw-0vYyc7eYjwr?usp=sharing)\n- nfn_depth_1024_ssr_l1: [notebook](https://colab.research.google.com/drive/1Gk4Db4tjhxUEL3uSLiRVlG1MeeGx6K6l?usp=sharing)\n\n### coordinate prediction models (SCS)\n\n- scs_detect_pre: [notebook](https://colab.research.google.com/drive/1qIXQRLkLFyXzyvP9jA2gaX6YZ4_qta27?usp=sharing)\n- scs_detect_pre_effv2l: [notebook](https://colab.research.google.com/drive/18vB2qrrBxC7Q4dwVQDR-ioPOY46oEbfu?usp=sharing)\n\n### coordinate prediction models (NFN)\n\n- nfn_detect_pre: [notebook](https://colab.research.google.com/drive/1IPiJDgPDXxOqNTbppZzM89n3ZuPKeWiV?usp=sharing)\n- nfn_detect_pre_effv2l: [notebook](https://colab.research.google.com/drive/1eKkZFqKUWZIacJrJYGM1PswYYSi3ztjx?usp=sharing)\n\n### coordinate prediction models (SS)\n\n- ss_detect: [notebook](https://colab.research.google.com/drive/1J3Pj8RMbDm5mG0vvBztyrSRK4G8-NLnU?usp=sharing)\n\n### severity prediction models (SCS)\n\n- _scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_convnext-s_for_exp: [notebook](https://colab.research.google.com/drive/1dWJUGhubs067mJ0GaIOZn8xUt_-1507_?usp=sharing)\n- _scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_effv2s_for_exp: [notebook](https://colab.research.google.com/drive/1SsqZOCv7eSYZfqcu5V6ufsHbPp3yN94X?usp=sharing)\n\n### severity prediction models (NFN)\n\n- nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_convnext-s: [notebook](https://colab.research.google.com/drive/1GtC4vtGVo2sY1Mr6cZf7J1IltO2poAEW?usp=sharing)\n- nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_effv2s: [notebook](https://colab.research.google.com/drive/1PJ5qU6szTPagEWCMXBYZ3WZPzK5L6_SJ?usp=sharing)\n\n### severity prediction models (SS)\n\n- ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_effv2s: [notebook](https://colab.research.google.com/drive/1CUmMoQeUy2ataJoDffEHTe338d66Sl3A?usp=sharing)\n- ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_convnext-s: [notebook](https://colab.research.google.com/drive/1KnagfSmFJ69HfASLCfzCOCrZBZyL3y7-?usp=sharing)\n\n### coordinate pretrained models\n\n[this notebook](https://colab.research.google.com/drive/1TdmZCT86dFTiP2vd2uMPtNMaAsJKPkc2?usp=sharing) is pre-training code for coordinate prediction models with [this dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500). model checkpoints are in [this dataset](https://www.kaggle.com/datasets/wadakoki/rsna-pretrained-models-for-coordinate/data)",
      "votes": 105
    },
    {
      "id": 3015576,
      "postDate": "2024-10-12T16:01:39.137Z",
      "content": "<p>Great first write up and congrats on the strong finish!</p>\n<p>Did the Attention-based MIL improvements improve your local CV as much as the LB? Also, do you know how much of the improvement was due to the auxillary loss?</p>",
      "rawMarkdown": "Great first write up and congrats on the strong finish!\n\nDid the Attention-based MIL improvements improve your local CV as much as the LB? Also, do you know how much of the improvement was due to the auxillary loss?\n\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 3016126,
          "postDate": "2024-10-13T11:55:09.073Z",
          "content": "<p>Thank you Bartley! Also, congratulations on winning a gold medal!</p>\n<blockquote>\n  <p>Attention-based MIL improvements improve your local CV as much as the LB?</p>\n</blockquote>\n<p>Improvements of my local cv by MIL are smaller than the improvement of LB. I lost log of actual improvements of cv, but I remember the improvements are less than 0.012. However public LB improved 0.3729 -&gt; 0.3588 (diff is 0.0141) and private LB improved 0.4259 -&gt; 0.4062 (diff is 0.0197). </p>\n<blockquote>\n  <p>do you know how much of the improvement was due to the auxillary loss?</p>\n</blockquote>\n<p>Actually, I submitted bi-LSTM+aux loss model at the same time, so I don't have the result of the improvements on LB by aux losses. The table below is the result of validation data of scs (same seed, same preprocess, same hyper params and same architecture with/without aux head)</p>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>without aux</th>\n<th>with aux</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.254</td>\n<td>0.240</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.283</td>\n<td>0.254</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.267</td>\n<td>0.254</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.264</td>\n<td>0.252</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.244</td>\n<td>0.261</td>\n</tr>\n<tr>\n<td>mean</td>\n<td>0.2624</td>\n<td>0.2522</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Thank you Bartley! Also, congratulations on winning a gold medal!\n\n> Attention-based MIL improvements improve your local CV as much as the LB?\n\nImprovements of my local cv by MIL are smaller than the improvement of LB. I lost log of actual improvements of cv, but I remember the improvements are less than 0.012. However public LB improved 0.3729 -> 0.3588 (diff is 0.0141) and private LB improved 0.4259 -> 0.4062 (diff is 0.0197). \n\n> do you know how much of the improvement was due to the auxillary loss?\n\nActually, I submitted bi-LSTM+aux loss model at the same time, so I don't have the result of the improvements on LB by aux losses. The table below is the result of validation data of scs (same seed, same preprocess, same hyper params and same architecture with/without aux head)\n\n|fold | without aux | with aux |\n| --- | --- | --- |\n| 0 | 0.254 | 0.240 |\n| 1 | 0.283 | 0.254 |\n| 2 | 0.267 | 0.254 |\n| 3 | 0.264 | 0.252 |\n| 4 | 0.244 | 0.261 |\n| mean | 0.2624 | 0.2522 |\n\n",
          "votes": 2,
          "replies": [
            {
              "id": 3016176,
              "postDate": "2024-10-13T13:08:46.380Z",
              "content": "<p>Interesting, the improvement for aux loss is a more significant than I thought. Thanks!</p>",
              "rawMarkdown": "Interesting, the improvement for aux loss is a more significant than I thought. Thanks!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3016736,
      "postDate": "2024-10-14T05:32:46.077Z",
      "content": "<p><code>In the classification task, these heads output (bs, 32) shape logits for each level.</code>; what targets were used since there are 5 levels for each image?</p>",
      "rawMarkdown": "`In the classification task, these heads output (bs, 32) shape logits for each level. `; what targets were used since there are 5 levels for each image?",
      "votes": 1
    },
    {
      "id": 3016650,
      "postDate": "2024-10-14T02:15:48.707Z",
      "content": "<p>Congratulations! Could you explain in detail why auxiliary loss is effective? I understand it should be nearly equivalent to the original architecture.</p>",
      "rawMarkdown": "Congratulations! Could you explain in detail why auxiliary loss is effective? I understand it should be nearly equivalent to the original architecture.",
      "votes": 2,
      "replies": [
        {
          "id": 3016791,
          "postDate": "2024-10-14T07:02:11.390Z",
          "content": "<p>Thank you!</p>\n<p>In my opinion, auxiliary loss makes it possible to add context to feature vectors. For example, I added depth aux head after bi-LSTM. This is because I wanted to incorporate the context of which slice the annotators were focusing on into the feature vectors output by the LSTM. I think this allows the main attention layer to function more effectively.</p>",
          "rawMarkdown": "Thank you!\n\nIn my opinion, auxiliary loss makes it possible to add context to feature vectors. For example, I added depth aux head after bi-LSTM. This is because I wanted to incorporate the context of which slice the annotators were focusing on into the feature vectors output by the LSTM. I think this allows the main attention layer to function more effectively.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3510943,
      "postDate": "2026-08-10T01:37:21.663Z",
      "content": "<p>This is very useful for my project, as i am working on similar predictive layer for Spine CSM , would like to connect with you further to seek your advice </p>",
      "rawMarkdown": "This is very useful for my project, as i am working on similar predictive layer for Spine CSM , would like to connect with you further to seek your advice "
    },
    {
      "id": 3105794,
      "postDate": "2025-01-24T01:58:51.063Z",
      "content": "<p>Could you kindly provide the necessary file? Without it, we are unable to run the notebook.</p>",
      "rawMarkdown": "Could you kindly provide the necessary file? Without it, we are unable to run the notebook."
    },
    {
      "id": 3041412,
      "postDate": "2024-11-10T10:06:48.170Z",
      "content": "<p>Hello! I’d like to ask about the three model notebooks in the instance number prediction models (SCS). I assume that these three models aim to achieve the same objective, with the only difference being their model structures—is that correct? Will these three models be integrated in the inference code? If my understanding is incorrect, could you please point it out? Thank you! Apologies for my slow progress with reading the code; so far, I’ve only reviewed these three notebooks.</p>",
      "rawMarkdown": "Hello! I’d like to ask about the three model notebooks in the instance number prediction models (SCS). I assume that these three models aim to achieve the same objective, with the only difference being their model structures—is that correct? Will these three models be integrated in the inference code? If my understanding is incorrect, could you please point it out? Thank you! Apologies for my slow progress with reading the code; so far, I’ve only reviewed these three notebooks."
    },
    {
      "id": 3036211,
      "postDate": "2024-11-04T11:03:47.003Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wadakoki\" target=\"_blank\">@wadakoki</a> what hardware did you use for the competition? Was it just the free quota on Google Colab Notebooks?</p>",
      "rawMarkdown": "Hi @wadakoki what hardware did you use for the competition? Was it just the free quota on Google Colab Notebooks?"
    },
    {
      "id": 3032725,
      "postDate": "2024-10-31T09:40:41.890Z",
      "content": "<p>thanks for sharing. Great solution and explanation.</p>",
      "rawMarkdown": "thanks for sharing. Great solution and explanation."
    },
    {
      "id": 3021408,
      "postDate": "2024-10-18T13:19:38.073Z",
      "content": "<p>Amazing Presentation <a href=\"https://www.kaggle.com/wadakoki\" target=\"_blank\">@wadakoki</a> </p>",
      "rawMarkdown": "Amazing Presentation @wadakoki "
    },
    {
      "id": 3021156,
      "postDate": "2024-10-18T08:28:04.493Z",
      "content": "<p>You're number one!</p>",
      "rawMarkdown": "You're number one!"
    },
    {
      "id": 3018270,
      "postDate": "2024-10-15T16:34:27.020Z",
      "content": "<p>congra man you deserve  it</p>",
      "rawMarkdown": "congra man you deserve  it"
    },
    {
      "id": 3016653,
      "postDate": "2024-10-14T02:26:04.727Z",
      "content": "<p>thanks for sharing. Piece of artist solution cake.</p>",
      "rawMarkdown": "thanks for sharing. Piece of artist solution cake."
    },
    {
      "id": 3018155,
      "postDate": "2024-10-15T14:48:50.443Z",
      "content": "<p>Good work.!!! 👍</p>",
      "rawMarkdown": "Good work.!!! 👍"
    },
    {
      "id": 3017994,
      "postDate": "2024-10-15T12:45:19.727Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 3017771,
      "postDate": "2024-10-15T07:32:04.183Z",
      "content": "<p>Thanks for sharing your approach. </p>",
      "rawMarkdown": "Thanks for sharing your approach. "
    },
    {
      "id": 3016963,
      "postDate": "2024-10-14T11:22:53.183Z",
      "content": "<p>good work!</p>",
      "rawMarkdown": "good work!"
    },
    {
      "id": 3032724,
      "postDate": "2024-10-31T09:39:40.360Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3018310,
      "postDate": "2024-10-15T17:00:45.067Z",
      "content": "<p>Congratulations! 😃</p>",
      "rawMarkdown": "Congratulations! 😃",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3015576,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-10-12T16:01:39.137000",
      "content": "<p>Great first write up and congrats on the strong finish!</p>\n<p>Did the Attention-based MIL improvements improve your local CV as much as the LB? Also, do you know how much of the improvement was due to the auxillary loss?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3016126,
          "author_name": "NANACHI",
          "author_url": "",
          "post_date": "2024-10-13T11:55:09.073000",
          "content": "<p>Thank you Bartley! Also, congratulations on winning a gold medal!</p>\n<blockquote>\n  <p>Attention-based MIL improvements improve your local CV as much as the LB?</p>\n</blockquote>\n<p>Improvements of my local cv by MIL are smaller than the improvement of LB. I lost log of actual improvements of cv, but I remember the improvements are less than 0.012. However public LB improved 0.3729 -&gt; 0.3588 (diff is 0.0141) and private LB improved 0.4259 -&gt; 0.4062 (diff is 0.0197). </p>\n<blockquote>\n  <p>do you know how much of the improvement was due to the auxillary loss?</p>\n</blockquote>\n<p>Actually, I submitted bi-LSTM+aux loss model at the same time, so I don't have the result of the improvements on LB by aux losses. The table below is the result of validation data of scs (same seed, same preprocess, same hyper params and same architecture with/without aux head)</p>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>without aux</th>\n<th>with aux</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.254</td>\n<td>0.240</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.283</td>\n<td>0.254</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.267</td>\n<td>0.254</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.264</td>\n<td>0.252</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.244</td>\n<td>0.261</td>\n</tr>\n<tr>\n<td>mean</td>\n<td>0.2624</td>\n<td>0.2522</td>\n</tr>\n</tbody>\n</table>",
          "votes": 2,
          "replies": [
            {
              "id": 3016176,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-10-13T13:08:46.380000",
              "content": "<p>Interesting, the improvement for aux loss is a more significant than I thought. Thanks!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3016736,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "2024-10-14T05:32:46.077000",
      "content": "<p><code>In the classification task, these heads output (bs, 32) shape logits for each level.</code>; what targets were used since there are 5 levels for each image?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3016650,
      "author_name": "Mingjie Wang",
      "author_url": "",
      "post_date": "2024-10-14T02:15:48.707000",
      "content": "<p>Congratulations! Could you explain in detail why auxiliary loss is effective? I understand it should be nearly equivalent to the original architecture.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3016791,
          "author_name": "NANACHI",
          "author_url": "",
          "post_date": "2024-10-14T07:02:11.390000",
          "content": "<p>Thank you!</p>\n<p>In my opinion, auxiliary loss makes it possible to add context to feature vectors. For example, I added depth aux head after bi-LSTM. This is because I wanted to incorporate the context of which slice the annotators were focusing on into the feature vectors output by the LSTM. I think this allows the main attention layer to function more effectively.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3510943,
      "author_name": "A D Vashishta",
      "author_url": "",
      "post_date": "2026-08-10T01:37:21.663000",
      "content": "<p>This is very useful for my project, as i am working on similar predictive layer for Spine CSM , would like to connect with you further to seek your advice </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3105794,
      "author_name": "tim062912",
      "author_url": "",
      "post_date": "2025-01-24T01:58:51.063000",
      "content": "<p>Could you kindly provide the necessary file? Without it, we are unable to run the notebook.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3041412,
      "author_name": "Switch9527",
      "author_url": "",
      "post_date": "2024-11-10T10:06:48.170000",
      "content": "<p>Hello! I’d like to ask about the three model notebooks in the instance number prediction models (SCS). I assume that these three models aim to achieve the same objective, with the only difference being their model structures—is that correct? Will these three models be integrated in the inference code? If my understanding is incorrect, could you please point it out? Thank you! Apologies for my slow progress with reading the code; so far, I’ve only reviewed these three notebooks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3036211,
      "author_name": "homiecal",
      "author_url": "",
      "post_date": "2024-11-04T11:03:47.003000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wadakoki\" target=\"_blank\">@wadakoki</a> what hardware did you use for the competition? Was it just the free quota on Google Colab Notebooks?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3032725,
      "author_name": "Kenny",
      "author_url": "",
      "post_date": "2024-10-31T09:40:41.890000",
      "content": "<p>thanks for sharing. Great solution and explanation.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3021408,
      "author_name": "Malabh Bakshi",
      "author_url": "",
      "post_date": "2024-10-18T13:19:38.073000",
      "content": "<p>Amazing Presentation <a href=\"https://www.kaggle.com/wadakoki\" target=\"_blank\">@wadakoki</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3021156,
      "author_name": "SUZUKI SEIYA",
      "author_url": "",
      "post_date": "2024-10-18T08:28:04.493000",
      "content": "<p>You're number one!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3018270,
      "author_name": "Yisak Birhanu Bule",
      "author_url": "",
      "post_date": "2024-10-15T16:34:27.020000",
      "content": "<p>congra man you deserve  it</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3016653,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2024-10-14T02:26:04.727000",
      "content": "<p>thanks for sharing. Piece of artist solution cake.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3018155,
      "author_name": "Vinothkumar Sekar",
      "author_url": "",
      "post_date": "2024-10-15T14:48:50.443000",
      "content": "<p>Good work.!!! 👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3017994,
      "author_name": "Rudrersh Asagodu",
      "author_url": "",
      "post_date": "2024-10-15T12:45:19.727000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3017771,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2024-10-15T07:32:04.183000",
      "content": "<p>Thanks for sharing your approach. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3016963,
      "author_name": "Usaid Ahmad",
      "author_url": "",
      "post_date": "2024-10-14T11:22:53.183000",
      "content": "<p>good work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3032724,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-31T09:39:40.360000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3018310,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-15T17:00:45.067000",
      "content": "<p>Congratulations! 😃</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3015549": "First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!\n\nFurthermore, I'd like to thank @hengck23 and @brendanartley. @hengck23 's discussion and notebook were the starting point for my solution, and @brendanartley 's [this dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500) helped my coordinate prediction models. I was deeply impressed by their contributions to the Kaggle community.\n\nThis is my first solution write-up, so please feel free to leave any comments or suggestions for improvement!\n\n## Summary\n\nMy solution is 2 stage approach, creating `test_label_coordinates.csv` and predicting severity. Furthermore, I separated 1st stage into instance_number prediction and coordinate prediction. Therefore I prepared 3 type of model, instance_number prediction model, coordinate prediction model and severity prediction model. The pipeline is shown in the following figure. \n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F83a6d286c875d8fb9ed5ff50513cbf11%2Frsna_pipeline_overview.png?generation=1728723901143537&alt=media)\n\n## 1st stage: test_label_coordinates creation\n\nIn the 1st stage, I use 2 type of models, 3D convolution model and 2D convolution model. These models are very simple, encoder + level-separated heads. \n\n### instance_number prediction (sagittal)\n\nIn this part, I used simple 3D ConvNeXt to predict instance_number for each level. Data that is fed into models is just normalized from 0 to 1, sorted by dicom's metadata and padded 32 to depth direction to align shape. Data preprocessing is shown in the following figure (scs example). \n\n![scs_volume_example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Ffe356b60c418b7bcef760bafa9d36210%2Fscs_volume_example.png?generation=1728729008294363&alt=media)\n\nIn training models, I trained models 2 tasks, regression and classification, and I used L1 Loss and Cross Entropy Loss respectively. In the classification task, these heads output (bs, 32) shape logits for each level. In the regression task, these heads output (bs, 3) shape vectors for each level. (bs, 3) shape vector means (x, y, z) and I used z for depth prediction, (x, y) were used auxiliary loss. In the regression task, I normalized coordinate labels 0 to 1 for stabilizing models during training. Concretely, I used label (x', y', z') = (x/width, y/height, z/32). The model architecture is shown in the following image (scs example). I implemented 3D ConvNeXt for this task (to implement 3D ConvNeXt, I referred to [this repo](https://github.com/FrancescoSaverioZuppichini/ConvNext)). \n\n![instance_number_prediction_scs_example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2Fbdf35cf44c9f268063c9b76d8698be19%2Frsna_instance_number_prediction_model_scs_example.png?generation=1728730136962978&alt=media)\n\nThe results of instance_number prediction models are shown in the following table (sagt2, scs). \n\n| model/error | +-0 | +-1 | +-2 | error>+-2 | \n| --- | --- | --- | --- | --- |\n|cls| 71.08% | 27.04% | 1.43% | 0.44% |\n|reg| 67.48% | 30.59% | 1.61% | 0.31% |\n\n\nI ensembled this 2 type of predictions using median for each level (actually I used 5 fold for each task). \n\n### coordinate prediction(sagittal)\n\nIn coordinate prediction task, I used 2d encoder + level-separated heads, almost same as instance_number regression model. Data is 3 channel image. The image is picked up using median of instance_number of L1 ~ S1. Then the data processed normalization and reshaping (512x512). Labels are (x', y') = (x/width, y/height)\nfor each level, same as instance_number regression, and also I used L1 loss. The model architecture is shown in the following figure. \n\n![coordinate_prediction_model_scs_example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F9585ce9ac5e65d46ba4c0bf759d19ba4%2Frsna_coordinate_prediction_model_scs_example.png?generation=1728733391395184&alt=media)\n\nI used ConvNeXt-base and Efficientnet-v2-l for this task. Before I train these models, I trained these models using @brendanartley 's [dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500). These pretrained models were slightly better than pretrained models that were trained using imagenet. I ensembled these predictions using mean. \n\n### instance_number calculation and coordinate prediction (axial)\n\nFor instance_number prediction of axial, I borrowed @hengck23 's method (notebook is [here](https://www.kaggle.com/code/hengck23/2d-to-3d-projection-for-dicom/notebook)). Then I predicted coordinates of axial, same as coordinate prediction for sagittal. \n\n## 2nd stage: severity prediction\n\nFor the 2nd stage, I attempted simple 2.5D model and MIL. 2.5D model can be implemented easily, however, MIL was better than simple 2.5D at final. \n\n## preprocessing\n\n### Cropping method\n\nMy preprocessing strategy is cropping. For example, I cropped sagt2 image for scs; \n\n1. pick up 5 images (center is an image that was assigned instance_number)\n2. reshape 512x512\n3. crop images using the coordinate (96 pix left and 32 pix right from coordinate x, 40 pix upper and 40 pix lower from coordinate y)\n\nAfter cropping an image, the image can be like the figure below (sagt2 for scs, L1/L2). \n\n![scs_cropped_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F3091793ada3e32fe2c2f4c022e10bf93%2Frsna_sagt2_cropped_image.png?generation=1728738536688017&alt=media)\n\nsagt2, sagt1 and axial were cropped for each classification task. The following tables are representing cropping range from (x, y) coordinate. \n\n**for scs**\n\n| type | left | right | upper | lower |\n| --- | --- | --- | --- | --- |\n| sagt2 | 96 | 32 | 40 | 40 |\n| axial | 96 | 96 | 96 | 96 |\n\nNote that when I crop images from axial, I picked up left or right  subarticular stenosis coordinate randomly, and for adjusting cropping point, I added +-20 to ss coordinate x. As a result, cropping range can be like the following figure (the example is right ss coordinate x + 20). \n\n![axial_for_scs_cropping](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F55cb454b535e50eed67dc7e03d3da6f2%2Frsna_ax_for_scs.png?generation=1728739918254610&alt=media)\n\n**for nfn**\n\n| type | left | right | upper | lower |\n| --- | --- | --- | --- | --- |\n| sagt1 (both left and right)| 96 | 64 | 32 | 32 |\n| axial (right) | 144 | 48 | 96 | 96 |\n| axial (left) | 48 | 144 | 96 | 96 |\n\n**for ss**\n\n| type | left | right | upper | lower |\n| --- | --- | --- | --- | --- |\n| axial (right) | 144 | 48 | 96 | 96 |\n| axial (left) | 48 | 144 | 96 | 96 |\n\nThe following image is the range of cropping axial for right subarticular stenosis. \n\n![axial_cropping_for_ss_right](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F916b910dd75444be91a822e2e19a2f9b%2Frsna_axial_cropping_for_ss.png?generation=1728740665640827&alt=media)\n\n### data augmentations\n\nI used several augmentations like below; \n\n*Before cropping*\n\n* random shift of coordinate x and y (-10~+10 pix)\n* random shift of instance_number (-2~+2. shifting probability was decided error probability of each instance_number prediction models)\n\n*After cropping*\n\n* RandomBrightnessContrast(p=0.25)\n* ShiftScaleRotate(shift_limit=0.1, scale_limit=(-0.1, 0.1), rotate_limit=20, p=0.5)\n\nEspecially, random shift of instance_number was crucial for robustness of error of 1st stage. \n\n\n## model architecture\n\nMy model architectures are shown in following figures. \n**[EDITED]**  I have updated the figure illustrating the model architecture to correct an error in the previous version. `aux_attn_score` in the code below is fed into cross entropy loss directly. \n\n![severity_prediction_model_scs_fixed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1466c45fe85d9cc404d5047150dba7c0%2Frsna_severity_prediction_model_for_scs_fixed.png?generation=1728897262477106&alt=media)\n![severity_prediction_model_ss_fixed](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5867589%2F1c2a33a4eac3252349847fa6723ebd72%2Frsna_severity_prediction_model_for_ss_fixed.png?generation=1728897365864286&alt=media)\n\nI used ConvNeXt-small and Efficientnet-v2-s as the encoder. After implementing Attention-based MIL, my public LB score was improved from 0.37 -> 0.35. Then, adding bi-LSTM, aux losses and ensembling improve my score from 0.35 to 0.33. bi-LSTM + Attention-based MIL was implemented like below. \n\n```python\nclass LSTMMIL(nn.Module):\n    def __init__(self, input_dim):\n        super(LSTMMIL, self).__init__()\n        self.lstm = nn.LSTM(input_dim, input_dim//2, num_layers=2, batch_first=True, dropout=0.1, bidirectional=True)\n        self.aux_attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, 1)\n        )\n        self.attention = nn.Sequential(\n            nn.Tanh(),\n            nn.Linear(input_dim, 1)\n        )\n    def forward(self, bags):\n        batch_size, num_instances, input_dim = bags.size()\n        bags_lstm, _ = self.lstm(bags)\n        attn_scores = self.attention(bags_lstm).squeeze(-1)\n        aux_attn_scores = self.aux_attention(bags_lstm).squeeze(-1)\n        attn_weights = torch.softmax(attn_scores, dim=-1)\n        weighted_instances = torch.bmm(attn_weights.unsqueeze(1), bags_lstm).squeeze(1)\n\n        return weighted_instances, aux_attn_scores\n```\n\n## what didn't work\n\n* MAMBA and Self-Attention instead of bi-LSTM\n* sharing weight between aux_attention layer and attention layer\n* sagt1 image for scs, sagt1 and sagt2 image for ss, sagt2 image for nfn\n* long epochs (I used 7 epochs for convnext-small and 14 epochs for efficientnet-v2-s)\n* large models (convnext-large < convnext-base < convnext-small in my experiments)\n* vision transformers (I think this was my problem. but convolution models were better than vits in my experiments)\n\n## code\n\nAll training code is implemented in google colaboratory. All models are used for [this inference code](https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook). Following links are pairs of model name & training notebook link. You can check these model name in the [inference code](https://www.kaggle.com/code/wadakoki/rsna-infer-pipeline-public/notebook). \n\nYou can train on google colaboratory environment with T4 + high memory. \n\n- [models](https://www.kaggle.com/datasets/wadakoki/rsna-spine-final-models/data)\n\n### instance number prediction models (SCS)\n\n- scs_depth_1024_ssr: [notebook](https://colab.research.google.com/drive/1JbSFgIwxlviyXb6uHv4vfbyqyjbdCRkw?usp=sharing)\n- scs_depth: [notebook](https://colab.research.google.com/drive/19YylaxYLYk1q6IOHpfi9UMnhbNUQBYML?usp=sharing)\n- scs_depth_1024_ssr_l1: [notebook](https://colab.research.google.com/drive/11fV56U5hPL2IiRzxaiLmuwyjCrygWsgO?usp=sharing)\n\n### instance number prediction models (NFN)\n- nfn_depth_1024_ssr: [notebook](https://colab.research.google.com/drive/1sIU9Aun1S1vla_W-4fZ24tGcn-IFD0a_?usp=sharing)\n- nfn_depth: [notebook](https://colab.research.google.com/drive/1EPcR7F5p2SvcgpaJdNKw-0vYyc7eYjwr?usp=sharing)\n- nfn_depth_1024_ssr_l1: [notebook](https://colab.research.google.com/drive/1Gk4Db4tjhxUEL3uSLiRVlG1MeeGx6K6l?usp=sharing)\n\n### coordinate prediction models (SCS)\n\n- scs_detect_pre: [notebook](https://colab.research.google.com/drive/1qIXQRLkLFyXzyvP9jA2gaX6YZ4_qta27?usp=sharing)\n- scs_detect_pre_effv2l: [notebook](https://colab.research.google.com/drive/18vB2qrrBxC7Q4dwVQDR-ioPOY46oEbfu?usp=sharing)\n\n### coordinate prediction models (NFN)\n\n- nfn_detect_pre: [notebook](https://colab.research.google.com/drive/1IPiJDgPDXxOqNTbppZzM89n3ZuPKeWiV?usp=sharing)\n- nfn_detect_pre_effv2l: [notebook](https://colab.research.google.com/drive/1eKkZFqKUWZIacJrJYGM1PswYYSi3ztjx?usp=sharing)\n\n### coordinate prediction models (SS)\n\n- ss_detect: [notebook](https://colab.research.google.com/drive/1J3Pj8RMbDm5mG0vvBztyrSRK4G8-NLnU?usp=sharing)\n\n### severity prediction models (SCS)\n\n- _scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_convnext-s_for_exp: [notebook](https://colab.research.google.com/drive/1dWJUGhubs067mJ0GaIOZn8xUt_-1507_?usp=sharing)\n- _scs_classify_5ch_axsagt2-lstm-mil_auxloss_auxdepth_effv2s_for_exp: [notebook](https://colab.research.google.com/drive/1SsqZOCv7eSYZfqcu5V6ufsHbPp3yN94X?usp=sharing)\n\n### severity prediction models (NFN)\n\n- nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_convnext-s: [notebook](https://colab.research.google.com/drive/1GtC4vtGVo2sY1Mr6cZf7J1IltO2poAEW?usp=sharing)\n- nfn_classify_5ch_axsagt1-lstm-mil_auxloss_auxdepth_2shift_effv2s: [notebook](https://colab.research.google.com/drive/1PJ5qU6szTPagEWCMXBYZ3WZPzK5L6_SJ?usp=sharing)\n\n### severity prediction models (SS)\n\n- ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_effv2s: [notebook](https://colab.research.google.com/drive/1CUmMoQeUy2ataJoDffEHTe338d66Sl3A?usp=sharing)\n- ss_classify_5ch_ax-lstm-mil_auxloss_auxdepth_convnext-s: [notebook](https://colab.research.google.com/drive/1KnagfSmFJ69HfASLCfzCOCrZBZyL3y7-?usp=sharing)\n\n### coordinate pretrained models\n\n[this notebook](https://colab.research.google.com/drive/1TdmZCT86dFTiP2vd2uMPtNMaAsJKPkc2?usp=sharing) is pre-training code for coordinate prediction models with [this dataset](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500). model checkpoints are in [this dataset](https://www.kaggle.com/datasets/wadakoki/rsna-pretrained-models-for-coordinate/data)",
    "3015576": "Great first write up and congrats on the strong finish!\n\nDid the Attention-based MIL improvements improve your local CV as much as the LB? Also, do you know how much of the improvement was due to the auxillary loss?\n\n\n",
    "3016736": "`In the classification task, these heads output (bs, 32) shape logits for each level. `; what targets were used since there are 5 levels for each image?",
    "3016650": "Congratulations! Could you explain in detail why auxiliary loss is effective? I understand it should be nearly equivalent to the original architecture.",
    "3510943": "This is very useful for my project, as i am working on similar predictive layer for Spine CSM , would like to connect with you further to seek your advice ",
    "3105794": "Could you kindly provide the necessary file? Without it, we are unable to run the notebook.",
    "3041412": "Hello! I’d like to ask about the three model notebooks in the instance number prediction models (SCS). I assume that these three models aim to achieve the same objective, with the only difference being their model structures—is that correct? Will these three models be integrated in the inference code? If my understanding is incorrect, could you please point it out? Thank you! Apologies for my slow progress with reading the code; so far, I’ve only reviewed these three notebooks.",
    "3036211": "Hi @wadakoki what hardware did you use for the competition? Was it just the free quota on Google Colab Notebooks?",
    "3032725": "thanks for sharing. Great solution and explanation.",
    "3021408": "Amazing Presentation @wadakoki ",
    "3021156": "You're number one!",
    "3018270": "congra man you deserve  it",
    "3016653": "thanks for sharing. Piece of artist solution cake.",
    "3018155": "Good work.!!! 👍",
    "3017994": "Congratulations!",
    "3017771": "Thanks for sharing your approach. ",
    "3016963": "good work!",
    "3032724": "",
    "3018310": "Congratulations! 😃"
  }
}