{
  "id": 539472,
  "title": "5th place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/two-people-5th-place-solution",
  "author_name": "",
  "post_date": "2024-10-10T00:44:08.130Z",
  "votes": 35,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I would like to express my gratitude to kaggle and rsna for organizing such a wonderful competition. I also want to thank <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>, who teamed up with me once again.</p>\n<h1>Summary</h1>\n<p>Our team's approach consists of the following main components. </p>\n<ul>\n<li>stage1 : <strong>heatmap-based detection + gaussian-expanding-label + external-dataset</strong></li>\n<li>stage2 : <strong>2.5d model(cnn + rnn) + level-wise sequence modeling + two-step training</strong></li>\n<li>augmentation : <strong>cutmix(p=1.0)</strong></li>\n<li>ensemble : <strong>various backbone ensemble + tta-like ensemble</strong></li>\n</ul>\n<h1>Stage1</h1>\n<h3>heatmap-based detection</h3>\n<p>Drawing inspiration from <a href=\"https://paperswithcode.com/task/keypoint-detection\" target=\"_blank\">keypoint detection</a>, we developed a heatmap-based model to identify 25 classes. We needed to develop 3 models, each designed to predict the given labels for their respective inputs.</p>\n<ul>\n<li>sagittal_t2 -&gt; spinal canal stenosis(5 classes)</li>\n<li>sagittal_t1 -&gt; neural foraminal narrowing(10 classes)</li>\n<li>axial_t2 -&gt; subarticular stenosis(10 classes)</li>\n</ul>\n<h3>gaussian-expanding-label</h3>\n<p>In the early stages of the competition, we used the given points as labels, but this resulted in slower training due to class imbalance. To address this, we applied a gaussian filter to the x and y coordinates, and for the z-axis, we multiplied by 0.5 as we moved further from the target frame, effectively increasing the area of the overall labels. This helped improve the convergence speed of the models and the z-axis accuracy.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2F69ed9f40fda16547c482d67810081a25%2Fheatmap-image.png?generation=1728448201888004&amp;alt=media\" alt=\"\"></p>\n<h3>external-dataset</h3>\n<p>While the performance with 3d unet was good, the 2d unet combined with a sequential model demonstrated higher accuracy related to the z-axis. Therefore, we ultimately opted for a 2d unet along with a sequential model (transformer, lstm).</p>\n<p>For the backbone, efficientnet_b5 provided the best performance. For the axial_t2, we found that increasing the maximum length to accommodate longer sequences improved performance. Additionally, leveraging the <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">public dataset</a> allowed us to make further improvements.</p>\n<h1>Stage2</h1>\n<h3>2.5d model(cnn + rnn)</h3>\n<p>We used the detection coordinates obtained from stage 1, cropping along the z-axis by ±2 and the x, y axes by ±32, and then resized the result (5, 64, 64) -&gt; (5, 128, 128) for use in stage 2. The structure of our model is similar to a typical 2.5d model(cnn + rnn), but our team added an additional module to model the relationships between classes. In the early stages of the competition, we modeled the 25 classes using lstm. </p>\n<h3>level-wise sequence modeling</h3>\n<p>However, upon examining the provided data labels, we were able to make the following analysis:</p>\n<blockquote>\n  <p>When symptom 1 is present at the level, there is a high probability that symptoms 2 and 3 will also be present at the same level. </p>\n</blockquote>\n<p>Therefore, we modified our approach to model only the classes at the same level, rather than all 25 classes. This adjustment significantly improved our score. </p>\n<pre><code>x = x.reshape(-, , , self.hidden_size)\nx = x.permute(, , , )\nx = x.reshape(-, , self.hidden_size)\n\nx, _ = self.rnn2(x)\n\nx = x.reshape(-, , , self.hidden_size)\nx = x.permute(, , , )\nx = x.reshape(-, , self.hidden_size)\n</code></pre>\n<p>In the later stages of the competition, we also tried concatenating the results of sequence modeling only at the same level and modeling only the same region. However, this approach did not perform better than the results from modeling only at the same level. Additionally, we implemented changes like skip connections, which we then used for our ensemble.</p>\n<p>In the case of cnn, we experimented with models like regnet and efficientnet, but convnext demonstrated the best performance.</p>\n<h3>two-step training</h3>\n<p>In the early stages of the competition, we trained our model using a loss function that closely followed the competition metric. However, this led to overfitting on the weighted labels, resulting in poor auc score. To improve the auc while still performing well on the competition metric, our team implemented a two-step training approach.</p>\n<p><strong>1st-step(pretraining)</strong><br>\nWe focused on maximizing the auc score by training the model's overall parameters without using weighted loss and any loss.</p>\n<p><strong>2nd-step(finetuning)</strong><br>\nWe employed weighted loss and any loss, freezing the model's backbone and training only the head parameters to optimize for the competition metric.</p>\n<p>Through this method, our team was able to significantly improve our scores compared to simply training with weighted loss and any loss.</p>\n<h1>Augmentation</h1>\n<h3>cutmix(p=1.0)</h3>\n<p>When training stage 2, we observed that the model quickly began to overfit. To prevent overfitting, we tried various methods, including flip, rotate, brightness, contrast, blur, and mixup. Among these, cutmix played the most significant role in increasing the auc score. In fact, using cutmix with p=1.0 resulted in the highest auc score.</p>\n<p>Additionally, we experimented with various methods, such as randomly adding ±1 at the z from stage1 or flipping the left and right labels. However, these approaches did not result in significant score improvements.</p>\n<h1>Ensemble</h1>\n<p>Based on these methods, we developed various stage 1 and stage 2 models and performed an ensemble.</p>\n<h3>various backbone ensemble</h3>\n<ul>\n<li>stage1 : max length</li>\n<li>stage1 : cnn backbone(regnety_002, efficientnet_b5)</li>\n<li>stage1 : whether it has fixed (x, y) coordinates or dynamic (x, y) coordinates according to the z-axis.</li>\n<li>stage2 : rnn modeling(skip connection, sequence modeling axis)</li>\n<li>stage2 : cnn backbone(convnext_small, convnext_tiny, caformer_s18, pvt_v2_b3)</li>\n</ul>\n<h3>tta-like ensemble</h3>\n<p>Additionally, the ensemble method that yielded the highest score on the private leaderboard was similar to test-time augmentation (tta). Instead of combining the stage 1 models developed by team members and passing them to stage 2 models, we inferred stage 2 models for each individual stage 1 model and then performed an ensemble. </p>\n<h1>Not worked</h1>\n<ul>\n<li>adding the mask from stage 1 as the stage 2's cnn channel</li>\n<li>bigger cnn backbone for stage2</li>\n<li>label smoothing </li>\n</ul>\n<h1>Code</h1>\n<p><a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> 's part<br>\n<a href=\"https://github.com/ElFazouani/RSNA-2024-Lumbar-Spine-Degenerative-Classification\" target=\"_blank\">https://github.com/ElFazouani/RSNA-2024-Lumbar-Spine-Degenerative-Classification</a></p>\n<p><a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> 's part<br>\n<a href=\"https://github.com/siwooyong/RSNA-2024-Lumbar-Spine-Degenerative-Classification\" target=\"_blank\">https://github.com/siwooyong/RSNA-2024-Lumbar-Spine-Degenerative-Classification</a></p>\n<p>inference notebook<br>\n<a href=\"https://www.kaggle.com/code/ahmedelfazouan/rsna-inference\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/rsna-inference</a></p>",
  "messages": [
    {
      "id": "3012473",
      "postDate": "10/09/2024 04:51:44",
      "content": "<p>I would like to express my gratitude to kaggle and rsna for organizing such a wonderful competition. I also want to thank <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>, who teamed up with me once again.</p>\n<h1>Summary</h1>\n<p>Our team's approach consists of the following main components. </p>\n<ul>\n<li>stage1 : <strong>heatmap-based detection + gaussian-expanding-label + external-dataset</strong></li>\n<li>stage2 : <strong>2.5d model(cnn + rnn) + level-wise sequence modeling + two-step training</strong></li>\n<li>augmentation : <strong>cutmix(p=1.0)</strong></li>\n<li>ensemble : <strong>various backbone ensemble + tta-like ensemble</strong></li>\n</ul>\n<h1>Stage1</h1>\n<h3>heatmap-based detection</h3>\n<p>Drawing inspiration from <a href=\"https://paperswithcode.com/task/keypoint-detection\" target=\"_blank\">keypoint detection</a>, we developed a heatmap-based model to identify 25 classes. We needed to develop 3 models, each designed to predict the given labels for their respective inputs.</p>\n<ul>\n<li>sagittal_t2 -&gt; spinal canal stenosis(5 classes)</li>\n<li>sagittal_t1 -&gt; neural foraminal narrowing(10 classes)</li>\n<li>axial_t2 -&gt; subarticular stenosis(10 classes)</li>\n</ul>\n<h3>gaussian-expanding-label</h3>\n<p>In the early stages of the competition, we used the given points as labels, but this resulted in slower training due to class imbalance. To address this, we applied a gaussian filter to the x and y coordinates, and for the z-axis, we multiplied by 0.5 as we moved further from the target frame, effectively increasing the area of the overall labels. This helped improve the convergence speed of the models and the z-axis accuracy.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2F69ed9f40fda16547c482d67810081a25%2Fheatmap-image.png?generation=1728448201888004&amp;alt=media\" alt=\"\"></p>\n<h3>external-dataset</h3>\n<p>While the performance with 3d unet was good, the 2d unet combined with a sequential model demonstrated higher accuracy related to the z-axis. Therefore, we ultimately opted for a 2d unet along with a sequential model (transformer, lstm).</p>\n<p>For the backbone, efficientnet_b5 provided the best performance. For the axial_t2, we found that increasing the maximum length to accommodate longer sequences improved performance. Additionally, leveraging the <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">public dataset</a> allowed us to make further improvements.</p>\n<h1>Stage2</h1>\n<h3>2.5d model(cnn + rnn)</h3>\n<p>We used the detection coordinates obtained from stage 1, cropping along the z-axis by ±2 and the x, y axes by ±32, and then resized the result (5, 64, 64) -&gt; (5, 128, 128) for use in stage 2. The structure of our model is similar to a typical 2.5d model(cnn + rnn), but our team added an additional module to model the relationships between classes. In the early stages of the competition, we modeled the 25 classes using lstm. </p>\n<h3>level-wise sequence modeling</h3>\n<p>However, upon examining the provided data labels, we were able to make the following analysis:</p>\n<blockquote>\n  <p>When symptom 1 is present at the level, there is a high probability that symptoms 2 and 3 will also be present at the same level. </p>\n</blockquote>\n<p>Therefore, we modified our approach to model only the classes at the same level, rather than all 25 classes. This adjustment significantly improved our score. </p>\n<pre><code>x = x.reshape(-, , , self.hidden_size)\nx = x.permute(, , , )\nx = x.reshape(-, , self.hidden_size)\n\nx, _ = self.rnn2(x)\n\nx = x.reshape(-, , , self.hidden_size)\nx = x.permute(, , , )\nx = x.reshape(-, , self.hidden_size)\n</code></pre>\n<p>In the later stages of the competition, we also tried concatenating the results of sequence modeling only at the same level and modeling only the same region. However, this approach did not perform better than the results from modeling only at the same level. Additionally, we implemented changes like skip connections, which we then used for our ensemble.</p>\n<p>In the case of cnn, we experimented with models like regnet and efficientnet, but convnext demonstrated the best performance.</p>\n<h3>two-step training</h3>\n<p>In the early stages of the competition, we trained our model using a loss function that closely followed the competition metric. However, this led to overfitting on the weighted labels, resulting in poor auc score. To improve the auc while still performing well on the competition metric, our team implemented a two-step training approach.</p>\n<p><strong>1st-step(pretraining)</strong><br>\nWe focused on maximizing the auc score by training the model's overall parameters without using weighted loss and any loss.</p>\n<p><strong>2nd-step(finetuning)</strong><br>\nWe employed weighted loss and any loss, freezing the model's backbone and training only the head parameters to optimize for the competition metric.</p>\n<p>Through this method, our team was able to significantly improve our scores compared to simply training with weighted loss and any loss.</p>\n<h1>Augmentation</h1>\n<h3>cutmix(p=1.0)</h3>\n<p>When training stage 2, we observed that the model quickly began to overfit. To prevent overfitting, we tried various methods, including flip, rotate, brightness, contrast, blur, and mixup. Among these, cutmix played the most significant role in increasing the auc score. In fact, using cutmix with p=1.0 resulted in the highest auc score.</p>\n<p>Additionally, we experimented with various methods, such as randomly adding ±1 at the z from stage1 or flipping the left and right labels. However, these approaches did not result in significant score improvements.</p>\n<h1>Ensemble</h1>\n<p>Based on these methods, we developed various stage 1 and stage 2 models and performed an ensemble.</p>\n<h3>various backbone ensemble</h3>\n<ul>\n<li>stage1 : max length</li>\n<li>stage1 : cnn backbone(regnety_002, efficientnet_b5)</li>\n<li>stage1 : whether it has fixed (x, y) coordinates or dynamic (x, y) coordinates according to the z-axis.</li>\n<li>stage2 : rnn modeling(skip connection, sequence modeling axis)</li>\n<li>stage2 : cnn backbone(convnext_small, convnext_tiny, caformer_s18, pvt_v2_b3)</li>\n</ul>\n<h3>tta-like ensemble</h3>\n<p>Additionally, the ensemble method that yielded the highest score on the private leaderboard was similar to test-time augmentation (tta). Instead of combining the stage 1 models developed by team members and passing them to stage 2 models, we inferred stage 2 models for each individual stage 1 model and then performed an ensemble. </p>\n<h1>Not worked</h1>\n<ul>\n<li>adding the mask from stage 1 as the stage 2's cnn channel</li>\n<li>bigger cnn backbone for stage2</li>\n<li>label smoothing </li>\n</ul>\n<h1>Code</h1>\n<p><a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> 's part<br>\n<a href=\"https://github.com/ElFazouani/RSNA-2024-Lumbar-Spine-Degenerative-Classification\" target=\"_blank\">https://github.com/ElFazouani/RSNA-2024-Lumbar-Spine-Degenerative-Classification</a></p>\n<p><a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> 's part<br>\n<a href=\"https://github.com/siwooyong/RSNA-2024-Lumbar-Spine-Degenerative-Classification\" target=\"_blank\">https://github.com/siwooyong/RSNA-2024-Lumbar-Spine-Degenerative-Classification</a></p>\n<p>inference notebook<br>\n<a href=\"https://www.kaggle.com/code/ahmedelfazouan/rsna-inference\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/rsna-inference</a></p>",
      "rawMarkdown": "I would like to express my gratitude to kaggle and rsna for organizing such a wonderful competition. I also want to thank @ahmedelfazouan, who teamed up with me once again.\n\n# Summary\nOur team's approach consists of the following main components. \n- stage1 : **heatmap-based detection + gaussian-expanding-label + external-dataset**\n- stage2 : **2.5d model(cnn + rnn) + level-wise sequence modeling + two-step training**\n- augmentation : **cutmix(p=1.0)**\n- ensemble : **various backbone ensemble + tta-like ensemble**\n\n# Stage1\n### heatmap-based detection\nDrawing inspiration from [keypoint detection](https://paperswithcode.com/task/keypoint-detection), we developed a heatmap-based model to identify 25 classes. We needed to develop 3 models, each designed to predict the given labels for their respective inputs.\n\n- sagittal_t2 -> spinal canal stenosis(5 classes)\n- sagittal_t1 -> neural foraminal narrowing(10 classes)\n- axial_t2 -> subarticular stenosis(10 classes)\n\n### gaussian-expanding-label\nIn the early stages of the competition, we used the given points as labels, but this resulted in slower training due to class imbalance. To address this, we applied a gaussian filter to the x and y coordinates, and for the z-axis, we multiplied by 0.5 as we moved further from the target frame, effectively increasing the area of the overall labels. This helped improve the convergence speed of the models and the z-axis accuracy.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2F69ed9f40fda16547c482d67810081a25%2Fheatmap-image.png?generation=1728448201888004&alt=media)\n\n### external-dataset\nWhile the performance with 3d unet was good, the 2d unet combined with a sequential model demonstrated higher accuracy related to the z-axis. Therefore, we ultimately opted for a 2d unet along with a sequential model (transformer, lstm).\n\nFor the backbone, efficientnet_b5 provided the best performance. For the axial_t2, we found that increasing the maximum length to accommodate longer sequences improved performance. Additionally, leveraging the [public dataset](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset) allowed us to make further improvements.\n\n# Stage2\n### 2.5d model(cnn + rnn)\nWe used the detection coordinates obtained from stage 1, cropping along the z-axis by ±2 and the x, y axes by ±32, and then resized the result (5, 64, 64) -> (5, 128, 128) for use in stage 2. The structure of our model is similar to a typical 2.5d model(cnn + rnn), but our team added an additional module to model the relationships between classes. In the early stages of the competition, we modeled the 25 classes using lstm. \n\n### level-wise sequence modeling\nHowever, upon examining the provided data labels, we were able to make the following analysis:\n\n>When symptom 1 is present at the level, there is a high probability that symptoms 2 and 3 will also be present at the same level. \n\nTherefore, we modified our approach to model only the classes at the same level, rather than all 25 classes. This adjustment significantly improved our score. \n\n```python\nx = x.reshape(-1, 5, 5, self.hidden_size)\nx = x.permute(0, 2, 1, 3)\nx = x.reshape(-1, 5, self.hidden_size)\n\nx, _ = self.rnn2(x)\n\nx = x.reshape(-1, 5, 5, self.hidden_size)\nx = x.permute(0, 2, 1, 3)\nx = x.reshape(-1, 25, self.hidden_size)\n```\n\nIn the later stages of the competition, we also tried concatenating the results of sequence modeling only at the same level and modeling only the same region. However, this approach did not perform better than the results from modeling only at the same level. Additionally, we implemented changes like skip connections, which we then used for our ensemble.\n\nIn the case of cnn, we experimented with models like regnet and efficientnet, but convnext demonstrated the best performance.\n\n### two-step training\nIn the early stages of the competition, we trained our model using a loss function that closely followed the competition metric. However, this led to overfitting on the weighted labels, resulting in poor auc score. To improve the auc while still performing well on the competition metric, our team implemented a two-step training approach.\n\n**1st-step(pretraining)**\nWe focused on maximizing the auc score by training the model's overall parameters without using weighted loss and any loss.\n\n**2nd-step(finetuning)**\nWe employed weighted loss and any loss, freezing the model's backbone and training only the head parameters to optimize for the competition metric.\n\nThrough this method, our team was able to significantly improve our scores compared to simply training with weighted loss and any loss.\n\n# Augmentation\n### cutmix(p=1.0)\nWhen training stage 2, we observed that the model quickly began to overfit. To prevent overfitting, we tried various methods, including flip, rotate, brightness, contrast, blur, and mixup. Among these, cutmix played the most significant role in increasing the auc score. In fact, using cutmix with p=1.0 resulted in the highest auc score.\n\nAdditionally, we experimented with various methods, such as randomly adding ±1 at the z from stage1 or flipping the left and right labels. However, these approaches did not result in significant score improvements.\n\n# Ensemble\nBased on these methods, we developed various stage 1 and stage 2 models and performed an ensemble.\n\n### various backbone ensemble\n- stage1 : max length\n- stage1 : cnn backbone(regnety_002, efficientnet_b5)\n- stage1 : whether it has fixed (x, y) coordinates or dynamic (x, y) coordinates according to the z-axis.\n- stage2 : rnn modeling(skip connection, sequence modeling axis)\n- stage2 : cnn backbone(convnext_small, convnext_tiny, caformer_s18, pvt_v2_b3)\n\n### tta-like ensemble\nAdditionally, the ensemble method that yielded the highest score on the private leaderboard was similar to test-time augmentation (tta). Instead of combining the stage 1 models developed by team members and passing them to stage 2 models, we inferred stage 2 models for each individual stage 1 model and then performed an ensemble. \n\n# Not worked\n- adding the mask from stage 1 as the stage 2's cnn channel\n- bigger cnn backbone for stage2\n- label smoothing \n\n# Code\n@ahmedelfazouan 's part\nhttps://github.com/ElFazouani/RSNA-2024-Lumbar-Spine-Degenerative-Classification\n\n@siwooyong 's part\nhttps://github.com/siwooyong/RSNA-2024-Lumbar-Spine-Degenerative-Classification\n\ninference notebook\nhttps://www.kaggle.com/code/ahmedelfazouan/rsna-inference",
      "votes": null
    },
    {
      "id": "3012511",
      "postDate": "10/09/2024 05:55:47",
      "content": "<p>Congratulations, really interesting findings! Cutmix really does work well. Did you try experimenting with more augmentation for bigger cnn backbones? really counterintuitive to me why bigger model didn't perform better.</p>",
      "rawMarkdown": "Congratulations, really interesting findings! Cutmix really does work well. Did you try experimenting with more augmentation for bigger cnn backbones? really counterintuitive to me why bigger model didn't perform better.",
      "votes": null
    },
    {
      "id": "3012531",
      "postDate": "10/09/2024 06:25:23",
      "content": "<p>Thank you for the great teamwork <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> !<br>\nI had a lot of fun working on this competition.</p>",
      "rawMarkdown": "Thank you for the great teamwork @siwooyong !\nI had a lot of fun working on this competition.",
      "votes": null
    },
    {
      "id": "3012541",
      "postDate": "10/09/2024 06:42:36",
      "content": "<p>Working with you is always the best! 😀</p>",
      "rawMarkdown": "Working with you is always the best! 😀",
      "votes": null
    },
    {
      "id": "3012552",
      "postDate": "10/09/2024 06:59:44",
      "content": "<p>I tried using bigger models like convnext_base and maxvit_small, but their performance was worse. I haven't conducted detailed experiments, so I'm not certain, but I have two hypotheses. 😀</p>\n<p>The first is that for bigger models to perform well, they require a sufficient dataset and a certain level of problem complexity. In this competition, those conditions weren't met, which likely led to overfitting and resulted in poorer performance compared to the smaller models.</p>\n<p>The second is that I used an image size of (128, 128) instead of (224, 224), which may have been more detrimental to the bigger models.</p>",
      "rawMarkdown": "I tried using bigger models like convnext_base and maxvit_small, but their performance was worse. I haven't conducted detailed experiments, so I'm not certain, but I have two hypotheses. 😀\n\nThe first is that for bigger models to perform well, they require a sufficient dataset and a certain level of problem complexity. In this competition, those conditions weren't met, which likely led to overfitting and resulted in poorer performance compared to the smaller models.\n\nThe second is that I used an image size of (128, 128) instead of (224, 224), which may have been more detrimental to the bigger models.",
      "votes": null
    },
    {
      "id": "3012581",
      "postDate": "10/09/2024 07:37:17",
      "content": "<p>Ahh I see, that makes much more sense.</p>",
      "rawMarkdown": "Ahh I see, that makes much more sense.",
      "votes": null
    },
    {
      "id": "3013056",
      "postDate": "10/09/2024 16:04:27",
      "content": "<p>Congratulations on winning the 5th prize in this competition. Thanks for sharing the details of your approach with specifics. </p>",
      "rawMarkdown": "Congratulations on winning the 5th prize in this competition. Thanks for sharing the details of your approach with specifics.",
      "votes": null
    },
    {
      "id": "3014683",
      "postDate": "10/11/2024 13:38:36",
      "content": "<p>Congratulations! Thanks a lot for adding all details. I opened our inference notebook but there are 2 privates datasets, any plan to make those public <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15158801%2F345de06e4da305f6f3ae3445647974d9%2FScreenshot%202024-10-11%20at%209.36.51AM.png?generation=1728653909629306&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratulations! Thanks a lot for adding all details. I opened our inference notebook but there are 2 privates datasets, any plan to make those public ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15158801%2F345de06e4da305f6f3ae3445647974d9%2FScreenshot%202024-10-11%20at%209.36.51AM.png?generation=1728653909629306&alt=media)",
      "votes": null
    },
    {
      "id": "3014692",
      "postDate": "10/11/2024 13:46:46",
      "content": "<p>The datasets are public now.</p>",
      "rawMarkdown": "The datasets are public now.",
      "votes": null
    },
    {
      "id": "3015451",
      "postDate": "10/12/2024 12:39:43",
      "content": "<p>Congradulations! <br>\nThanks for your sharing! Could you introduce your loss function of the stage_1 in the code to me breifly?I am a rookie and know little about Unet :)</p>",
      "rawMarkdown": "Congradulations! \nThanks for your sharing! Could you introduce your loss function of the stage_1 in the code to me breifly?I am a rookie and know little about Unet :)",
      "votes": null
    },
    {
      "id": "3015487",
      "postDate": "10/12/2024 13:35:24",
      "content": "<p>thanks, stage 1 model is composed of two main modules, the first one is used to detect areas where x and y are present, and the other to detect z,<br>\nthe loss is also divided into two parts, the first half to measure the performance of x and y and the other for z.</p>",
      "rawMarkdown": "thanks, stage 1 model is composed of two main modules, the first one is used to detect areas where x and y are present, and the other to detect z,\nthe loss is also divided into two parts, the first half to measure the performance of x and y and the other for z.",
      "votes": null
    },
    {
      "id": "3018733",
      "postDate": "10/16/2024 03:03:58",
      "content": "<p>Congratulations! Really an eye-opener to a new thought process of object detection.</p>",
      "rawMarkdown": "Congratulations! Really an eye-opener to a new thought process of object detection.",
      "votes": null
    },
    {
      "id": "3018909",
      "postDate": "10/16/2024 05:55:57",
      "content": "<p>Congratiulation!</p>",
      "rawMarkdown": "Congratiulation!",
      "votes": null
    },
    {
      "id": "3037942",
      "postDate": "11/06/2024 11:50:02",
      "content": "<p>Congratulations. I have a question about the code of <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> 's part. It seems that you use the separately processed 'coords_rsna_improved.csv' file. When it is Sagittal T2/STIR, you create a side column and use it. Could you please explain how to create this file?</p>",
      "rawMarkdown": "Congratulations. I have a question about the code of @ahmedelfazouan 's part. It seems that you use the separately processed 'coords_rsna_improved.csv' file. When it is Sagittal T2/STIR, you create a side column and use it. Could you please explain how to create this file?",
      "votes": null
    },
    {
      "id": "3038344",
      "postDate": "11/06/2024 21:13:00",
      "content": "<p>It was shared in this public dataset <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "It was shared in this public dataset [here](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3012511,
      "author_name": "zshashz",
      "author_url": "",
      "post_date": "10/09/2024 05:55:47",
      "content": "<p>Congratulations, really interesting findings! Cutmix really does work well. Did you try experimenting with more augmentation for bigger cnn backbones? really counterintuitive to me why bigger model didn't perform better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3012552,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "10/09/2024 06:59:44",
          "content": "<p>I tried using bigger models like convnext_base and maxvit_small, but their performance was worse. I haven't conducted detailed experiments, so I'm not certain, but I have two hypotheses. 😀</p>\n<p>The first is that for bigger models to perform well, they require a sufficient dataset and a certain level of problem complexity. In this competition, those conditions weren't met, which likely led to overfitting and resulted in poorer performance compared to the smaller models.</p>\n<p>The second is that I used an image size of (128, 128) instead of (224, 224), which may have been more detrimental to the bigger models.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3012581,
              "author_name": "zshashz",
              "author_url": "",
              "post_date": "10/09/2024 07:37:17",
              "content": "<p>Ahh I see, that makes much more sense.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3012531,
      "author_name": "ahmedelfazouan",
      "author_url": "",
      "post_date": "10/09/2024 06:25:23",
      "content": "<p>Thank you for the great teamwork <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> !<br>\nI had a lot of fun working on this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3012541,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "10/09/2024 06:42:36",
          "content": "<p>Working with you is always the best! 😀</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3013056,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "10/09/2024 16:04:27",
      "content": "<p>Congratulations on winning the 5th prize in this competition. Thanks for sharing the details of your approach with specifics. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3014683,
      "author_name": "karelbecerra",
      "author_url": "",
      "post_date": "10/11/2024 13:38:36",
      "content": "<p>Congratulations! Thanks a lot for adding all details. I opened our inference notebook but there are 2 privates datasets, any plan to make those public <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15158801%2F345de06e4da305f6f3ae3445647974d9%2FScreenshot%202024-10-11%20at%209.36.51AM.png?generation=1728653909629306&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3014692,
          "author_name": "ahmedelfazouan",
          "author_url": "",
          "post_date": "10/11/2024 13:46:46",
          "content": "<p>The datasets are public now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3015451,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "10/12/2024 12:39:43",
      "content": "<p>Congradulations! <br>\nThanks for your sharing! Could you introduce your loss function of the stage_1 in the code to me breifly?I am a rookie and know little about Unet :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3015487,
          "author_name": "ahmedelfazouan",
          "author_url": "",
          "post_date": "10/12/2024 13:35:24",
          "content": "<p>thanks, stage 1 model is composed of two main modules, the first one is used to detect areas where x and y are present, and the other to detect z,<br>\nthe loss is also divided into two parts, the first half to measure the performance of x and y and the other for z.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3018733,
      "author_name": "sidharthserjy",
      "author_url": "",
      "post_date": "10/16/2024 03:03:58",
      "content": "<p>Congratulations! Really an eye-opener to a new thought process of object detection.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3018909,
      "author_name": "tanishkpatil",
      "author_url": "",
      "post_date": "10/16/2024 05:55:57",
      "content": "<p>Congratiulation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3037942,
      "author_name": "jinstat",
      "author_url": "",
      "post_date": "11/06/2024 11:50:02",
      "content": "<p>Congratulations. I have a question about the code of <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> 's part. It seems that you use the separately processed 'coords_rsna_improved.csv' file. When it is Sagittal T2/STIR, you create a side column and use it. Could you please explain how to create this file?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3038344,
          "author_name": "ahmedelfazouan",
          "author_url": "",
          "post_date": "11/06/2024 21:13:00",
          "content": "<p>It was shared in this public dataset <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3012473": "I would like to express my gratitude to kaggle and rsna for organizing such a wonderful competition. I also want to thank @ahmedelfazouan, who teamed up with me once again.\n\n# Summary\nOur team's approach consists of the following main components. \n- stage1 : **heatmap-based detection + gaussian-expanding-label + external-dataset**\n- stage2 : **2.5d model(cnn + rnn) + level-wise sequence modeling + two-step training**\n- augmentation : **cutmix(p=1.0)**\n- ensemble : **various backbone ensemble + tta-like ensemble**\n\n# Stage1\n### heatmap-based detection\nDrawing inspiration from [keypoint detection](https://paperswithcode.com/task/keypoint-detection), we developed a heatmap-based model to identify 25 classes. We needed to develop 3 models, each designed to predict the given labels for their respective inputs.\n\n- sagittal_t2 -> spinal canal stenosis(5 classes)\n- sagittal_t1 -> neural foraminal narrowing(10 classes)\n- axial_t2 -> subarticular stenosis(10 classes)\n\n### gaussian-expanding-label\nIn the early stages of the competition, we used the given points as labels, but this resulted in slower training due to class imbalance. To address this, we applied a gaussian filter to the x and y coordinates, and for the z-axis, we multiplied by 0.5 as we moved further from the target frame, effectively increasing the area of the overall labels. This helped improve the convergence speed of the models and the z-axis accuracy.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2F69ed9f40fda16547c482d67810081a25%2Fheatmap-image.png?generation=1728448201888004&alt=media)\n\n### external-dataset\nWhile the performance with 3d unet was good, the 2d unet combined with a sequential model demonstrated higher accuracy related to the z-axis. Therefore, we ultimately opted for a 2d unet along with a sequential model (transformer, lstm).\n\nFor the backbone, efficientnet_b5 provided the best performance. For the axial_t2, we found that increasing the maximum length to accommodate longer sequences improved performance. Additionally, leveraging the [public dataset](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset) allowed us to make further improvements.\n\n# Stage2\n### 2.5d model(cnn + rnn)\nWe used the detection coordinates obtained from stage 1, cropping along the z-axis by ±2 and the x, y axes by ±32, and then resized the result (5, 64, 64) -> (5, 128, 128) for use in stage 2. The structure of our model is similar to a typical 2.5d model(cnn + rnn), but our team added an additional module to model the relationships between classes. In the early stages of the competition, we modeled the 25 classes using lstm. \n\n### level-wise sequence modeling\nHowever, upon examining the provided data labels, we were able to make the following analysis:\n\n>When symptom 1 is present at the level, there is a high probability that symptoms 2 and 3 will also be present at the same level. \n\nTherefore, we modified our approach to model only the classes at the same level, rather than all 25 classes. This adjustment significantly improved our score. \n\n```python\nx = x.reshape(-1, 5, 5, self.hidden_size)\nx = x.permute(0, 2, 1, 3)\nx = x.reshape(-1, 5, self.hidden_size)\n\nx, _ = self.rnn2(x)\n\nx = x.reshape(-1, 5, 5, self.hidden_size)\nx = x.permute(0, 2, 1, 3)\nx = x.reshape(-1, 25, self.hidden_size)\n```\n\nIn the later stages of the competition, we also tried concatenating the results of sequence modeling only at the same level and modeling only the same region. However, this approach did not perform better than the results from modeling only at the same level. Additionally, we implemented changes like skip connections, which we then used for our ensemble.\n\nIn the case of cnn, we experimented with models like regnet and efficientnet, but convnext demonstrated the best performance.\n\n### two-step training\nIn the early stages of the competition, we trained our model using a loss function that closely followed the competition metric. However, this led to overfitting on the weighted labels, resulting in poor auc score. To improve the auc while still performing well on the competition metric, our team implemented a two-step training approach.\n\n**1st-step(pretraining)**\nWe focused on maximizing the auc score by training the model's overall parameters without using weighted loss and any loss.\n\n**2nd-step(finetuning)**\nWe employed weighted loss and any loss, freezing the model's backbone and training only the head parameters to optimize for the competition metric.\n\nThrough this method, our team was able to significantly improve our scores compared to simply training with weighted loss and any loss.\n\n# Augmentation\n### cutmix(p=1.0)\nWhen training stage 2, we observed that the model quickly began to overfit. To prevent overfitting, we tried various methods, including flip, rotate, brightness, contrast, blur, and mixup. Among these, cutmix played the most significant role in increasing the auc score. In fact, using cutmix with p=1.0 resulted in the highest auc score.\n\nAdditionally, we experimented with various methods, such as randomly adding ±1 at the z from stage1 or flipping the left and right labels. However, these approaches did not result in significant score improvements.\n\n# Ensemble\nBased on these methods, we developed various stage 1 and stage 2 models and performed an ensemble.\n\n### various backbone ensemble\n- stage1 : max length\n- stage1 : cnn backbone(regnety_002, efficientnet_b5)\n- stage1 : whether it has fixed (x, y) coordinates or dynamic (x, y) coordinates according to the z-axis.\n- stage2 : rnn modeling(skip connection, sequence modeling axis)\n- stage2 : cnn backbone(convnext_small, convnext_tiny, caformer_s18, pvt_v2_b3)\n\n### tta-like ensemble\nAdditionally, the ensemble method that yielded the highest score on the private leaderboard was similar to test-time augmentation (tta). Instead of combining the stage 1 models developed by team members and passing them to stage 2 models, we inferred stage 2 models for each individual stage 1 model and then performed an ensemble. \n\n# Not worked\n- adding the mask from stage 1 as the stage 2's cnn channel\n- bigger cnn backbone for stage2\n- label smoothing \n\n# Code\n@ahmedelfazouan 's part\nhttps://github.com/ElFazouani/RSNA-2024-Lumbar-Spine-Degenerative-Classification\n\n@siwooyong 's part\nhttps://github.com/siwooyong/RSNA-2024-Lumbar-Spine-Degenerative-Classification\n\ninference notebook\nhttps://www.kaggle.com/code/ahmedelfazouan/rsna-inference",
    "3012511": "Congratulations, really interesting findings! Cutmix really does work well. Did you try experimenting with more augmentation for bigger cnn backbones? really counterintuitive to me why bigger model didn't perform better.",
    "3012531": "Thank you for the great teamwork @siwooyong !\nI had a lot of fun working on this competition.",
    "3012541": "Working with you is always the best! 😀",
    "3012552": "I tried using bigger models like convnext_base and maxvit_small, but their performance was worse. I haven't conducted detailed experiments, so I'm not certain, but I have two hypotheses. 😀\n\nThe first is that for bigger models to perform well, they require a sufficient dataset and a certain level of problem complexity. In this competition, those conditions weren't met, which likely led to overfitting and resulted in poorer performance compared to the smaller models.\n\nThe second is that I used an image size of (128, 128) instead of (224, 224), which may have been more detrimental to the bigger models.",
    "3012581": "Ahh I see, that makes much more sense.",
    "3013056": "Congratulations on winning the 5th prize in this competition. Thanks for sharing the details of your approach with specifics.",
    "3014683": "Congratulations! Thanks a lot for adding all details. I opened our inference notebook but there are 2 privates datasets, any plan to make those public ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15158801%2F345de06e4da305f6f3ae3445647974d9%2FScreenshot%202024-10-11%20at%209.36.51AM.png?generation=1728653909629306&alt=media)",
    "3014692": "The datasets are public now.",
    "3015451": "Congradulations! \nThanks for your sharing! Could you introduce your loss function of the stage_1 in the code to me breifly?I am a rookie and know little about Unet :)",
    "3015487": "thanks, stage 1 model is composed of two main modules, the first one is used to detect areas where x and y are present, and the other to detect z,\nthe loss is also divided into two parts, the first half to measure the performance of x and y and the other for z.",
    "3018733": "Congratulations! Really an eye-opener to a new thought process of object detection.",
    "3018909": "Congratiulation!",
    "3037942": "Congratulations. I have a question about the code of @ahmedelfazouan 's part. It seems that you use the separately processed 'coords_rsna_improved.csv' file. When it is Sagittal T2/STIR, you create a side column and use it. Could you please explain how to create this file?",
    "3038344": "It was shared in this public dataset [here](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset)"
  },
  "source": "meta"
}