{
  "id": 539510,
  "title": "13th place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/spirit-bomb-13th-place-solution",
  "author_name": "",
  "post_date": "2024-10-10T10:26:18.897Z",
  "votes": 21,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>2 stage system</h1>\n<h2>Summary</h2>\n<p>Similar to other competitors, we implemented a 2-stage system composed of three core models (plus an additional one for pretraining):</p>\n<ul>\n<li>Keypoints: Localizing intervertebral spots.</li>\n<li>Levels: Classifying axial images into spinal levels.</li>\n<li>Pretraining (Patch): Classifying individual patches.</li>\n<li>Sequence: A 2D model combined with a Transformer for final predictions.</li>\n</ul>\n<p>For detailed training specifics, please refer to the accompanying code.</p>\n<h2>Models</h2>\n<h3>1. Keypoints</h3>\n<p>We initially experimented with bounding box detectors and segmentation models, but keypoints provided the best results and flexibility. The model outputs 6 pairs of XY coordinates (12 total), where 5 pairs are used for sagittal images, and 1 pair is designated for axial images. We incorporated the <a href=\"https://www.kaggle.com/code/brendanartley/lumbar-coordinate-dataset-code\" target=\"_blank\">corrected dataset labels</a> shared by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. This allowed us to use tilted crops, although the performance improvement was minimal.</p>\n<p>The loss function was based on Euclidean distance, with masking applied to avoid propagating axial keypoint errors in sagittal images and vice versa. We also introduced position encoding to the backbone features before pooling to enhance the model’s performance.</p>\n<h3>2. Levels</h3>\n<p>For axial images, we developed a straightforward image model that predicts the corresponding level. Some metadata was concatenated with the backbone features to enhance the model's performance.</p>\n<p>We observed occasional prediction inconsistencies (e.g., L5-S1 predicted next to L1-L2 within the same series). To address this, we redefined the task as a regression problem, predicting values between 0.0 and 1.0 at 0.25 intervals (e.g., 0.0 for L1-L2, 0.25 for L2-L3, etc.), using Mean Squared Error (MSE) as the loss function. This approach penalized predictions further from the true value more heavily.</p>\n<h3>3. Sequence</h3>\n<p>Leveraging anatomical symmetries, we trained one model with study-level-side inputs. </p>\n<ul>\n<li>Input: For example, one row would be all patches in the L1-L2 right side, and another row all patches in the L4-L5 left side.</li>\n<li>Output: neural_foraminal_narrowing, subarticular_stenosis, spinal_canal_stenosis.</li>\n</ul>\n<p>Details are explained further below.</p>\n<h4>Input</h4>\n<p>We used the keypoints and levels models to create patches, with each axial patch corresponding to one or more spinal levels. Since the levels model is a regressor, we applied a threshold with a tolerance range to capture additional context. For example, patches corresponding to L2-L3 (0.25 output) could span from 0.10 to 0.40.</p>\n<p>To ensure consistency, we maintained uniform pixel spacing across each plane. We used 96x96 patches. Sagittal patches were offset at the top and bottom to prevent level leakage.</p>\n<p>Some examples:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F7d69bbb4c4a112042c288b1780157350%2Fpatches.png?generation=1728463071124063&amp;alt=media\" alt=\"\"></p>\n<h4>Pretraining: Patch model</h4>\n<p>Before training the sequence model, we pre-trained a patch model using individual patches with the same loss function (defined later). This step significantly improved the convergence of the sequence model.</p>\n<h4>Architecture</h4>\n<p>2D model + encoder-decoder transformer. </p>\n<p>The architecture combined a 2D image model with an encoder-decoder transformer. The best backbones we found were <code>regnetz_b16.ra3_in1k</code> and <code>hgnet_tiny.ssld_in1k</code>.</p>\n<p>Metadata (modified XYZ world coordinates) was fed into the encoder, while image features were passed to the decoder. We added three learnable vectors to the decoder, each representing one disease classification (analogous to CLS tokens). Various pooling strategies were tested, including max/mean pooling and attention pooling.</p>\n<h4>Reshapable model</h4>\n<p>The model was designed to be flexible, allowing it to process either an entire study or individual level-sides. This adaptability enabled us to train the model at the level-side while evaluating it on the whole study.</p>\n<p>What to do with spinal?</p>\n<p>Since spinal stenosis affects the center rather than specific sides, we duplicated the label for both sides during training. At inference, we aggregated the learnable vectors from both sides (those added in the decoder) before passing them through the classification head. We tested multiple aggregation methods and ultimately chose max pooling.</p>\n<h4>Loss</h4>\n<p>We used the competition-specific loss function implemented in PyTorch. Although we explored several variations, such as a focal version of the competition loss, none yielded significantly better results.</p>\n<h2>Further information</h2>\n<table>\n<thead>\n<tr>\n<th>description</th>\n<th>backbone</th>\n<th>local avg cv</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best public: ensemble 4 (dropped one) TTA 20</td>\n<td>regnet</td>\n<td>0.4027</td>\n<td>0.3452</td>\n<td>0.4121</td>\n</tr>\n<tr>\n<td>best local: ensemble 5 TTA 40</td>\n<td>hgnet</td>\n<td>0.3883</td>\n<td>0.3612</td>\n<td>0.4215</td>\n</tr>\n<tr>\n<td>best private: ensemble 5 TTA 16</td>\n<td>regnet</td>\n<td>0.4076</td>\n<td>0.3472</td>\n<td>0.4118</td>\n</tr>\n</tbody>\n</table>\n<p>The two selected submissions were the best local and the best in the public leaderboard. The best in the public leaderboard was the 3rd best in the private leaderboard.</p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/claverru/RSNA-lumbar\" target=\"_blank\">https://github.com/claverru/RSNA-lumbar</a></p>",
  "messages": [
    {
      "id": "3012689",
      "postDate": "10/09/2024 09:05:00",
      "content": "<h1>2 stage system</h1>\n<h2>Summary</h2>\n<p>Similar to other competitors, we implemented a 2-stage system composed of three core models (plus an additional one for pretraining):</p>\n<ul>\n<li>Keypoints: Localizing intervertebral spots.</li>\n<li>Levels: Classifying axial images into spinal levels.</li>\n<li>Pretraining (Patch): Classifying individual patches.</li>\n<li>Sequence: A 2D model combined with a Transformer for final predictions.</li>\n</ul>\n<p>For detailed training specifics, please refer to the accompanying code.</p>\n<h2>Models</h2>\n<h3>1. Keypoints</h3>\n<p>We initially experimented with bounding box detectors and segmentation models, but keypoints provided the best results and flexibility. The model outputs 6 pairs of XY coordinates (12 total), where 5 pairs are used for sagittal images, and 1 pair is designated for axial images. We incorporated the <a href=\"https://www.kaggle.com/code/brendanartley/lumbar-coordinate-dataset-code\" target=\"_blank\">corrected dataset labels</a> shared by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. This allowed us to use tilted crops, although the performance improvement was minimal.</p>\n<p>The loss function was based on Euclidean distance, with masking applied to avoid propagating axial keypoint errors in sagittal images and vice versa. We also introduced position encoding to the backbone features before pooling to enhance the model’s performance.</p>\n<h3>2. Levels</h3>\n<p>For axial images, we developed a straightforward image model that predicts the corresponding level. Some metadata was concatenated with the backbone features to enhance the model's performance.</p>\n<p>We observed occasional prediction inconsistencies (e.g., L5-S1 predicted next to L1-L2 within the same series). To address this, we redefined the task as a regression problem, predicting values between 0.0 and 1.0 at 0.25 intervals (e.g., 0.0 for L1-L2, 0.25 for L2-L3, etc.), using Mean Squared Error (MSE) as the loss function. This approach penalized predictions further from the true value more heavily.</p>\n<h3>3. Sequence</h3>\n<p>Leveraging anatomical symmetries, we trained one model with study-level-side inputs. </p>\n<ul>\n<li>Input: For example, one row would be all patches in the L1-L2 right side, and another row all patches in the L4-L5 left side.</li>\n<li>Output: neural_foraminal_narrowing, subarticular_stenosis, spinal_canal_stenosis.</li>\n</ul>\n<p>Details are explained further below.</p>\n<h4>Input</h4>\n<p>We used the keypoints and levels models to create patches, with each axial patch corresponding to one or more spinal levels. Since the levels model is a regressor, we applied a threshold with a tolerance range to capture additional context. For example, patches corresponding to L2-L3 (0.25 output) could span from 0.10 to 0.40.</p>\n<p>To ensure consistency, we maintained uniform pixel spacing across each plane. We used 96x96 patches. Sagittal patches were offset at the top and bottom to prevent level leakage.</p>\n<p>Some examples:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F7d69bbb4c4a112042c288b1780157350%2Fpatches.png?generation=1728463071124063&amp;alt=media\" alt=\"\"></p>\n<h4>Pretraining: Patch model</h4>\n<p>Before training the sequence model, we pre-trained a patch model using individual patches with the same loss function (defined later). This step significantly improved the convergence of the sequence model.</p>\n<h4>Architecture</h4>\n<p>2D model + encoder-decoder transformer. </p>\n<p>The architecture combined a 2D image model with an encoder-decoder transformer. The best backbones we found were <code>regnetz_b16.ra3_in1k</code> and <code>hgnet_tiny.ssld_in1k</code>.</p>\n<p>Metadata (modified XYZ world coordinates) was fed into the encoder, while image features were passed to the decoder. We added three learnable vectors to the decoder, each representing one disease classification (analogous to CLS tokens). Various pooling strategies were tested, including max/mean pooling and attention pooling.</p>\n<h4>Reshapable model</h4>\n<p>The model was designed to be flexible, allowing it to process either an entire study or individual level-sides. This adaptability enabled us to train the model at the level-side while evaluating it on the whole study.</p>\n<p>What to do with spinal?</p>\n<p>Since spinal stenosis affects the center rather than specific sides, we duplicated the label for both sides during training. At inference, we aggregated the learnable vectors from both sides (those added in the decoder) before passing them through the classification head. We tested multiple aggregation methods and ultimately chose max pooling.</p>\n<h4>Loss</h4>\n<p>We used the competition-specific loss function implemented in PyTorch. Although we explored several variations, such as a focal version of the competition loss, none yielded significantly better results.</p>\n<h2>Further information</h2>\n<table>\n<thead>\n<tr>\n<th>description</th>\n<th>backbone</th>\n<th>local avg cv</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best public: ensemble 4 (dropped one) TTA 20</td>\n<td>regnet</td>\n<td>0.4027</td>\n<td>0.3452</td>\n<td>0.4121</td>\n</tr>\n<tr>\n<td>best local: ensemble 5 TTA 40</td>\n<td>hgnet</td>\n<td>0.3883</td>\n<td>0.3612</td>\n<td>0.4215</td>\n</tr>\n<tr>\n<td>best private: ensemble 5 TTA 16</td>\n<td>regnet</td>\n<td>0.4076</td>\n<td>0.3472</td>\n<td>0.4118</td>\n</tr>\n</tbody>\n</table>\n<p>The two selected submissions were the best local and the best in the public leaderboard. The best in the public leaderboard was the 3rd best in the private leaderboard.</p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/claverru/RSNA-lumbar\" target=\"_blank\">https://github.com/claverru/RSNA-lumbar</a></p>",
      "rawMarkdown": "# 2 stage system\n\n## Summary\n\nSimilar to other competitors, we implemented a 2-stage system composed of three core models (plus an additional one for pretraining):\n\n- Keypoints: Localizing intervertebral spots.\n- Levels: Classifying axial images into spinal levels.\n- Pretraining (Patch): Classifying individual patches.\n- Sequence: A 2D model combined with a Transformer for final predictions.\n\nFor detailed training specifics, please refer to the accompanying code.\n\n## Models\n\n### 1. Keypoints\n\nWe initially experimented with bounding box detectors and segmentation models, but keypoints provided the best results and flexibility. The model outputs 6 pairs of XY coordinates (12 total), where 5 pairs are used for sagittal images, and 1 pair is designated for axial images. We incorporated the [corrected dataset labels](https://www.kaggle.com/code/brendanartley/lumbar-coordinate-dataset-code) shared by @brendanartley. This allowed us to use tilted crops, although the performance improvement was minimal.\n\nThe loss function was based on Euclidean distance, with masking applied to avoid propagating axial keypoint errors in sagittal images and vice versa. We also introduced position encoding to the backbone features before pooling to enhance the model’s performance.\n\n### 2. Levels\n\nFor axial images, we developed a straightforward image model that predicts the corresponding level. Some metadata was concatenated with the backbone features to enhance the model's performance.\n\nWe observed occasional prediction inconsistencies (e.g., L5-S1 predicted next to L1-L2 within the same series). To address this, we redefined the task as a regression problem, predicting values between 0.0 and 1.0 at 0.25 intervals (e.g., 0.0 for L1-L2, 0.25 for L2-L3, etc.), using Mean Squared Error (MSE) as the loss function. This approach penalized predictions further from the true value more heavily.\n\n### 3. Sequence\n\nLeveraging anatomical symmetries, we trained one model with study-level-side inputs. \n- Input: For example, one row would be all patches in the L1-L2 right side, and another row all patches in the L4-L5 left side.\n- Output: neural_foraminal_narrowing, subarticular_stenosis, spinal_canal_stenosis.\n\nDetails are explained further below.\n\n#### Input\n\nWe used the keypoints and levels models to create patches, with each axial patch corresponding to one or more spinal levels. Since the levels model is a regressor, we applied a threshold with a tolerance range to capture additional context. For example, patches corresponding to L2-L3 (0.25 output) could span from 0.10 to 0.40.\n\nTo ensure consistency, we maintained uniform pixel spacing across each plane. We used 96x96 patches. Sagittal patches were offset at the top and bottom to prevent level leakage.\n\nSome examples:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F7d69bbb4c4a112042c288b1780157350%2Fpatches.png?generation=1728463071124063&alt=media)\n\n#### Pretraining: Patch model\n\nBefore training the sequence model, we pre-trained a patch model using individual patches with the same loss function (defined later). This step significantly improved the convergence of the sequence model.\n\n#### Architecture\n\n2D model + encoder-decoder transformer. \n\nThe architecture combined a 2D image model with an encoder-decoder transformer. The best backbones we found were `regnetz_b16.ra3_in1k` and `hgnet_tiny.ssld_in1k`.\n\nMetadata (modified XYZ world coordinates) was fed into the encoder, while image features were passed to the decoder. We added three learnable vectors to the decoder, each representing one disease classification (analogous to CLS tokens). Various pooling strategies were tested, including max/mean pooling and attention pooling.\n\n#### Reshapable model\n\nThe model was designed to be flexible, allowing it to process either an entire study or individual level-sides. This adaptability enabled us to train the model at the level-side while evaluating it on the whole study.\n\nWhat to do with spinal?\n\nSince spinal stenosis affects the center rather than specific sides, we duplicated the label for both sides during training. At inference, we aggregated the learnable vectors from both sides (those added in the decoder) before passing them through the classification head. We tested multiple aggregation methods and ultimately chose max pooling.\n\n#### Loss\n\nWe used the competition-specific loss function implemented in PyTorch. Although we explored several variations, such as a focal version of the competition loss, none yielded significantly better results.\n\n## Further information\n\n| description | backbone | local avg cv | public | private |\n| --- | --- | --- | --- | --- |\n| best public: ensemble 4 (dropped one) TTA 20 | regnet | 0.4027 | 0.3452 | 0.4121  | \n| best local: ensemble 5 TTA 40 | hgnet | 0.3883 | 0.3612 | 0.4215  | \n| best private: ensemble 5 TTA 16 | regnet | 0.4076 | 0.3472 | 0.4118  | \n\n\nThe two selected submissions were the best local and the best in the public leaderboard. The best in the public leaderboard was the 3rd best in the private leaderboard.\n\n\n## Code\n\nhttps://github.com/claverru/RSNA-lumbar",
      "votes": null
    },
    {
      "id": "3019312",
      "postDate": "10/16/2024 12:36:56",
      "content": "<p>Congrats on the gold and becoming Competitions Master!!</p>",
      "rawMarkdown": "Congrats on the gold and becoming Competitions Master!!",
      "votes": null
    },
    {
      "id": "3019916",
      "postDate": "10/17/2024 02:46:06",
      "content": "<p>Congratulations on this impressive development! great job</p>",
      "rawMarkdown": "Congratulations on this impressive development! great job",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3019312,
      "author_name": "bacterio",
      "author_url": "",
      "post_date": "10/16/2024 12:36:56",
      "content": "<p>Congrats on the gold and becoming Competitions Master!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3019916,
      "author_name": "darwinberrio",
      "author_url": "",
      "post_date": "10/17/2024 02:46:06",
      "content": "<p>Congratulations on this impressive development! great job</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3012689": "# 2 stage system\n\n## Summary\n\nSimilar to other competitors, we implemented a 2-stage system composed of three core models (plus an additional one for pretraining):\n\n- Keypoints: Localizing intervertebral spots.\n- Levels: Classifying axial images into spinal levels.\n- Pretraining (Patch): Classifying individual patches.\n- Sequence: A 2D model combined with a Transformer for final predictions.\n\nFor detailed training specifics, please refer to the accompanying code.\n\n## Models\n\n### 1. Keypoints\n\nWe initially experimented with bounding box detectors and segmentation models, but keypoints provided the best results and flexibility. The model outputs 6 pairs of XY coordinates (12 total), where 5 pairs are used for sagittal images, and 1 pair is designated for axial images. We incorporated the [corrected dataset labels](https://www.kaggle.com/code/brendanartley/lumbar-coordinate-dataset-code) shared by @brendanartley. This allowed us to use tilted crops, although the performance improvement was minimal.\n\nThe loss function was based on Euclidean distance, with masking applied to avoid propagating axial keypoint errors in sagittal images and vice versa. We also introduced position encoding to the backbone features before pooling to enhance the model’s performance.\n\n### 2. Levels\n\nFor axial images, we developed a straightforward image model that predicts the corresponding level. Some metadata was concatenated with the backbone features to enhance the model's performance.\n\nWe observed occasional prediction inconsistencies (e.g., L5-S1 predicted next to L1-L2 within the same series). To address this, we redefined the task as a regression problem, predicting values between 0.0 and 1.0 at 0.25 intervals (e.g., 0.0 for L1-L2, 0.25 for L2-L3, etc.), using Mean Squared Error (MSE) as the loss function. This approach penalized predictions further from the true value more heavily.\n\n### 3. Sequence\n\nLeveraging anatomical symmetries, we trained one model with study-level-side inputs. \n- Input: For example, one row would be all patches in the L1-L2 right side, and another row all patches in the L4-L5 left side.\n- Output: neural_foraminal_narrowing, subarticular_stenosis, spinal_canal_stenosis.\n\nDetails are explained further below.\n\n#### Input\n\nWe used the keypoints and levels models to create patches, with each axial patch corresponding to one or more spinal levels. Since the levels model is a regressor, we applied a threshold with a tolerance range to capture additional context. For example, patches corresponding to L2-L3 (0.25 output) could span from 0.10 to 0.40.\n\nTo ensure consistency, we maintained uniform pixel spacing across each plane. We used 96x96 patches. Sagittal patches were offset at the top and bottom to prevent level leakage.\n\nSome examples:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F7d69bbb4c4a112042c288b1780157350%2Fpatches.png?generation=1728463071124063&alt=media)\n\n#### Pretraining: Patch model\n\nBefore training the sequence model, we pre-trained a patch model using individual patches with the same loss function (defined later). This step significantly improved the convergence of the sequence model.\n\n#### Architecture\n\n2D model + encoder-decoder transformer. \n\nThe architecture combined a 2D image model with an encoder-decoder transformer. The best backbones we found were `regnetz_b16.ra3_in1k` and `hgnet_tiny.ssld_in1k`.\n\nMetadata (modified XYZ world coordinates) was fed into the encoder, while image features were passed to the decoder. We added three learnable vectors to the decoder, each representing one disease classification (analogous to CLS tokens). Various pooling strategies were tested, including max/mean pooling and attention pooling.\n\n#### Reshapable model\n\nThe model was designed to be flexible, allowing it to process either an entire study or individual level-sides. This adaptability enabled us to train the model at the level-side while evaluating it on the whole study.\n\nWhat to do with spinal?\n\nSince spinal stenosis affects the center rather than specific sides, we duplicated the label for both sides during training. At inference, we aggregated the learnable vectors from both sides (those added in the decoder) before passing them through the classification head. We tested multiple aggregation methods and ultimately chose max pooling.\n\n#### Loss\n\nWe used the competition-specific loss function implemented in PyTorch. Although we explored several variations, such as a focal version of the competition loss, none yielded significantly better results.\n\n## Further information\n\n| description | backbone | local avg cv | public | private |\n| --- | --- | --- | --- | --- |\n| best public: ensemble 4 (dropped one) TTA 20 | regnet | 0.4027 | 0.3452 | 0.4121  | \n| best local: ensemble 5 TTA 40 | hgnet | 0.3883 | 0.3612 | 0.4215  | \n| best private: ensemble 5 TTA 16 | regnet | 0.4076 | 0.3472 | 0.4118  | \n\n\nThe two selected submissions were the best local and the best in the public leaderboard. The best in the public leaderboard was the 3rd best in the private leaderboard.\n\n\n## Code\n\nhttps://github.com/claverru/RSNA-lumbar",
    "3019312": "Congrats on the gold and becoming Competitions Master!!",
    "3019916": "Congratulations on this impressive development! great job"
  },
  "source": "meta"
}