{
  "id": 539486,
  "title": "7th solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/hlip-7th-solution",
  "author_name": "",
  "post_date": "2024-10-15T02:03:44.320Z",
  "votes": 27,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Summary</h1>\n<p>First, we would like to express our gratitude to Kaggle and RSNA for hosting this great competition.<br><br>\nOur solution is an ensemble of 1 stage and 2 stage.</p>\n<h1>lhwcv part</h1>\n<p>I'm very grateful to my teammates, I learned a lot from them. <br></p>\n<p>2 stage methods, components:</p>\n<ul>\n<li><strong>stage1</strong>: keypoint regression (both 2d and 3d)</li>\n<li><strong>stage2_model1</strong>: crop by keypoint, fuse 3 view on feature space</li>\n<li><strong>stage2_model2</strong>: crop by keypoint, single view for 3 types of condition</li>\n<li><strong>ensemble</strong>: various backbone + flip TTA</li>\n</ul>\n<h2>stage1</h2>\n<p>keypoint regression (both 2d and 3d) <br></p>\n<p>Sagittal:</p>\n<ul>\n<li><strong>keypoint-sag-2d</strong>: use timm model, first sample 10 slice for each series (same methods in ITK's baseline)<br>\nthen input with the center slice.</li>\n<li><strong>keypoint-sag-3d</strong>: use timm-3d model,sagittal T2 -&gt; 5 xyz points, T1-&gt; 10 xyz points</li>\n</ul>\n<p>Axial:</p>\n<ul>\n<li><strong>keypoint-axial-3d</strong>: use timm-3d model, 10 xyz points </li>\n<li><strong>keypoint-axial-level-cls</strong>: 2.5d (CNN + LSTM + attention) for level classification (z), <br>\nand another 2d model for xy regression</li>\n</ul>\n<p>Used backbones, which is searched by ITK:</p>\n<ul>\n<li>densenet161</li>\n<li>convnext small</li>\n<li>gluon_resnet152_v1s</li>\n<li>convformer_s36.sail_in22k_ft_in1k_384</li>\n</ul>\n<h2>stage2_model1</h2>\n<p>fuse 3 view on feature space <br><br>\nThe keypoint models used are <strong>keypoint-sag-2d</strong> and <strong>keypoint-axial-3d</strong>.  <br>\nFor the sagittal plane, we sample 10 slices for each series, cropping the center 5 slices on T2, and cropping 5 slices from the left and 5 slices from the right on T1, for each level.  <br>\nFor the axial plane, we sample 5 slices for each level based on the Z-coordinate of the point.  <br>\nThe final input size is: x -&gt; <strong>(bs, 5_cond, 5_level, 5, crop_h, crop_w)</strong>.  <br>\nThe model architecture is as follows:</p>\n<pre><code>\ny, axial_embs = .axial_model(axial_x) \nys2, sag_embs = .sag_model(x)\n\nsag_embs = sag_embs.permute(, , , ).reshape(b, , -)  \nembs = torch.cat((sag_embs, axial_embs), dim=-)\nys = .out_linear(embs).reshape(b, , , )  \nys = ys.permute(, , , )  \nys = ys + ys2\n</code></pre>\n<h2>stage2_model2</h2>\n<p>The defect of stage2_model1 is that it uses 2D keypoints on the sagittal plane and does not handle multiple series on the axial plane very well.  <br>\nThe keypoint models used are <strong>keypoint-sag-3d</strong> and <strong>keypoint-axial-level-cls</strong>.  <br>\nwe crop the 3D ROI image using (z_imgs, crop_h, crop_w) based on XYZ points, where z_imgs can be either 3 or 5; we found that 3 works best for my purposes.  <br>\nFor the axial plane, there may be multiple detections, but we only used the one with the highest confidence (not necessarily the best).</p>\n<h2>ensemble</h2>\n<p>stage2_model1 + stage2_model2 with many backbones, TTA: hflip for sagittal, vflip for axial <br><br>\n<strong>results</strong>: cv: 0.373  LB: 0.35  PB: 0.40</p>\n<p><strong>backbones</strong>: </p>\n<ul>\n<li>pvt-b1</li>\n<li>pvt-b2</li>\n<li>convnext-tiny</li>\n<li>convnext-small</li>\n<li>densenet161</li>\n</ul>\n<h1>hecgck &amp; lhwcv ensemble</h1>\n<p>Team submission notebook can be found at:  <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6e4d08b10ead7a1c9b3904f1d83bd626%2FWechatIMG1025.jpg?generation=1728957811388315&amp;alt=media\" alt=\"\"><br>\nTeam post-submission notebook can be found at:  <br>\n<a href=\"https://www.kaggle.com/code/hengck23/post-lhw-v24-ensemble-add-heng\" target=\"_blank\">https://www.kaggle.com/code/hengck23/post-lhw-v24-ensemble-add-heng</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F07bfb7f227ec83b40bc3b7a92c5f7092%2FWechatIMG1026.jpg?generation=1728957822586591&amp;alt=media\" alt=\"\"></p>\n<h1>hecgck part</h1>\n<p>see: <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439</a></p>\n<h2>Code</h2>\n<p>Infer Notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng</a><br>\nTrain:</li>\n<li><a href=\"https://github.com/lhwcv/solution-rsna-2024-lumbar-spine\" target=\"_blank\">https://github.com/lhwcv/solution-rsna-2024-lumbar-spine</a></li>\n<li><a href=\"https://github.com/hengck23/solution-rsna-2024-lumbar-spine\" target=\"_blank\">https://github.com/hengck23/solution-rsna-2024-lumbar-spine</a></li>\n</ul>",
  "messages": [
    {
      "id": "3012554",
      "postDate": "10/09/2024 07:04:19",
      "content": "<h1>Summary</h1>\n<p>First, we would like to express our gratitude to Kaggle and RSNA for hosting this great competition.<br><br>\nOur solution is an ensemble of 1 stage and 2 stage.</p>\n<h1>lhwcv part</h1>\n<p>I'm very grateful to my teammates, I learned a lot from them. <br></p>\n<p>2 stage methods, components:</p>\n<ul>\n<li><strong>stage1</strong>: keypoint regression (both 2d and 3d)</li>\n<li><strong>stage2_model1</strong>: crop by keypoint, fuse 3 view on feature space</li>\n<li><strong>stage2_model2</strong>: crop by keypoint, single view for 3 types of condition</li>\n<li><strong>ensemble</strong>: various backbone + flip TTA</li>\n</ul>\n<h2>stage1</h2>\n<p>keypoint regression (both 2d and 3d) <br></p>\n<p>Sagittal:</p>\n<ul>\n<li><strong>keypoint-sag-2d</strong>: use timm model, first sample 10 slice for each series (same methods in ITK's baseline)<br>\nthen input with the center slice.</li>\n<li><strong>keypoint-sag-3d</strong>: use timm-3d model,sagittal T2 -&gt; 5 xyz points, T1-&gt; 10 xyz points</li>\n</ul>\n<p>Axial:</p>\n<ul>\n<li><strong>keypoint-axial-3d</strong>: use timm-3d model, 10 xyz points </li>\n<li><strong>keypoint-axial-level-cls</strong>: 2.5d (CNN + LSTM + attention) for level classification (z), <br>\nand another 2d model for xy regression</li>\n</ul>\n<p>Used backbones, which is searched by ITK:</p>\n<ul>\n<li>densenet161</li>\n<li>convnext small</li>\n<li>gluon_resnet152_v1s</li>\n<li>convformer_s36.sail_in22k_ft_in1k_384</li>\n</ul>\n<h2>stage2_model1</h2>\n<p>fuse 3 view on feature space <br><br>\nThe keypoint models used are <strong>keypoint-sag-2d</strong> and <strong>keypoint-axial-3d</strong>.  <br>\nFor the sagittal plane, we sample 10 slices for each series, cropping the center 5 slices on T2, and cropping 5 slices from the left and 5 slices from the right on T1, for each level.  <br>\nFor the axial plane, we sample 5 slices for each level based on the Z-coordinate of the point.  <br>\nThe final input size is: x -&gt; <strong>(bs, 5_cond, 5_level, 5, crop_h, crop_w)</strong>.  <br>\nThe model architecture is as follows:</p>\n<pre><code>\ny, axial_embs = .axial_model(axial_x) \nys2, sag_embs = .sag_model(x)\n\nsag_embs = sag_embs.permute(, , , ).reshape(b, , -)  \nembs = torch.cat((sag_embs, axial_embs), dim=-)\nys = .out_linear(embs).reshape(b, , , )  \nys = ys.permute(, , , )  \nys = ys + ys2\n</code></pre>\n<h2>stage2_model2</h2>\n<p>The defect of stage2_model1 is that it uses 2D keypoints on the sagittal plane and does not handle multiple series on the axial plane very well.  <br>\nThe keypoint models used are <strong>keypoint-sag-3d</strong> and <strong>keypoint-axial-level-cls</strong>.  <br>\nwe crop the 3D ROI image using (z_imgs, crop_h, crop_w) based on XYZ points, where z_imgs can be either 3 or 5; we found that 3 works best for my purposes.  <br>\nFor the axial plane, there may be multiple detections, but we only used the one with the highest confidence (not necessarily the best).</p>\n<h2>ensemble</h2>\n<p>stage2_model1 + stage2_model2 with many backbones, TTA: hflip for sagittal, vflip for axial <br><br>\n<strong>results</strong>: cv: 0.373  LB: 0.35  PB: 0.40</p>\n<p><strong>backbones</strong>: </p>\n<ul>\n<li>pvt-b1</li>\n<li>pvt-b2</li>\n<li>convnext-tiny</li>\n<li>convnext-small</li>\n<li>densenet161</li>\n</ul>\n<h1>hecgck &amp; lhwcv ensemble</h1>\n<p>Team submission notebook can be found at:  <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6e4d08b10ead7a1c9b3904f1d83bd626%2FWechatIMG1025.jpg?generation=1728957811388315&amp;alt=media\" alt=\"\"><br>\nTeam post-submission notebook can be found at:  <br>\n<a href=\"https://www.kaggle.com/code/hengck23/post-lhw-v24-ensemble-add-heng\" target=\"_blank\">https://www.kaggle.com/code/hengck23/post-lhw-v24-ensemble-add-heng</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F07bfb7f227ec83b40bc3b7a92c5f7092%2FWechatIMG1026.jpg?generation=1728957822586591&amp;alt=media\" alt=\"\"></p>\n<h1>hecgck part</h1>\n<p>see: <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439</a></p>\n<h2>Code</h2>\n<p>Infer Notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng</a><br>\nTrain:</li>\n<li><a href=\"https://github.com/lhwcv/solution-rsna-2024-lumbar-spine\" target=\"_blank\">https://github.com/lhwcv/solution-rsna-2024-lumbar-spine</a></li>\n<li><a href=\"https://github.com/hengck23/solution-rsna-2024-lumbar-spine\" target=\"_blank\">https://github.com/hengck23/solution-rsna-2024-lumbar-spine</a></li>\n</ul>",
      "rawMarkdown": "# Summary\nFirst, we would like to express our gratitude to Kaggle and RSNA for hosting this great competition.<br/>\nOur solution is an ensemble of 1 stage and 2 stage.\n\n# lhwcv part\nI'm very grateful to my teammates, I learned a lot from them. <br/>\n\n2 stage methods, components:\n- **stage1**: keypoint regression (both 2d and 3d)\n- **stage2_model1**: crop by keypoint, fuse 3 view on feature space\n- **stage2_model2**: crop by keypoint, single view for 3 types of condition\n- **ensemble**: various backbone + flip TTA\n\n## stage1\nkeypoint regression (both 2d and 3d) <br/>\n\nSagittal:\n- **keypoint-sag-2d**: use timm model, first sample 10 slice for each series (same methods in ITK's baseline)\n   then input with the center slice.\n- **keypoint-sag-3d**: use timm-3d model,sagittal T2 -> 5 xyz points, T1-> 10 xyz points\n\nAxial:\n- **keypoint-axial-3d**: use timm-3d model, 10 xyz points \n- **keypoint-axial-level-cls**: 2.5d (CNN + LSTM + attention) for level classification (z), \n  and another 2d model for xy regression\n  \nUsed backbones, which is searched by ITK:\n- densenet161\n- convnext small\n- gluon_resnet152_v1s\n- convformer_s36.sail_in22k_ft_in1k_384\n\n\n## stage2_model1\nfuse 3 view on feature space <br/>\nThe keypoint models used are **keypoint-sag-2d** and **keypoint-axial-3d**.  \nFor the sagittal plane, we sample 10 slices for each series, cropping the center 5 slices on T2, and cropping 5 slices from the left and 5 slices from the right on T1, for each level.  \nFor the axial plane, we sample 5 slices for each level based on the Z-coordinate of the point.  \nThe final input size is: x -> **(bs, 5_cond, 5_level, 5, crop_h, crop_w)**.  \nThe model architecture is as follows:\n\n```python\n# train \ny, axial_embs = self.axial_model(axial_x) # bs, \nys2, sag_embs = self.sag_model(x)\n# sag_embs: bs, 5_cond, 1_level, 128\nsag_embs = sag_embs.permute(0, 2, 1, 3).reshape(b, 1, -1)  # bs, 1, 5*128\nembs = torch.cat((sag_embs, axial_embs), dim=-1)\nys = self.out_linear(embs).reshape(b, 1, 5, 3)  # bs, 1_level, 5_cond, 3\nys = ys.permute(0, 2, 1, 3)  # # bs, 5_cond, 1_level, 3\nys = ys + ys2\n```\n\n## stage2_model2\nThe defect of stage2_model1 is that it uses 2D keypoints on the sagittal plane and does not handle multiple series on the axial plane very well.  \nThe keypoint models used are **keypoint-sag-3d** and **keypoint-axial-level-cls**.  \nwe crop the 3D ROI image using (z_imgs, crop_h, crop_w) based on XYZ points, where z_imgs can be either 3 or 5; we found that 3 works best for my purposes.  \nFor the axial plane, there may be multiple detections, but we only used the one with the highest confidence (not necessarily the best).\n\n\n## ensemble\nstage2_model1 + stage2_model2 with many backbones, TTA: hflip for sagittal, vflip for axial <br/>\n**results**: cv: 0.373  LB: 0.35  PB: 0.40\n\n**backbones**: \n- pvt-b1\n- pvt-b2\n- convnext-tiny\n- convnext-small\n- densenet161\n\n\n# hecgck & lhwcv ensemble\nTeam submission notebook can be found at:  \nhttps://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6e4d08b10ead7a1c9b3904f1d83bd626%2FWechatIMG1025.jpg?generation=1728957811388315&alt=media)\nTeam post-submission notebook can be found at:  \nhttps://www.kaggle.com/code/hengck23/post-lhw-v24-ensemble-add-heng\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F07bfb7f227ec83b40bc3b7a92c5f7092%2FWechatIMG1026.jpg?generation=1728957822586591&alt=media)\n\n\n\n# hecgck part\nsee: \nhttps://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439\n\n## Code\nInfer Notebook:\n- https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\nTrain:\n- https://github.com/lhwcv/solution-rsna-2024-lumbar-spine\n- https://github.com/hengck23/solution-rsna-2024-lumbar-spine",
      "votes": null
    },
    {
      "id": "3012928",
      "postDate": "10/09/2024 14:01:38",
      "content": "<p>Congratulations on your seventh place. Do you have a specific notebook</p>",
      "rawMarkdown": "Congratulations on your seventh place. Do you have a specific notebook",
      "votes": null
    },
    {
      "id": "3014773",
      "postDate": "10/11/2024 15:20:01",
      "content": "<p>Looking forward to your code, thanks.</p>",
      "rawMarkdown": "Looking forward to your code, thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3012928,
      "author_name": "icw1lee",
      "author_url": "",
      "post_date": "10/09/2024 14:01:38",
      "content": "<p>Congratulations on your seventh place. Do you have a specific notebook</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3014773,
      "author_name": "nextnlp",
      "author_url": "",
      "post_date": "10/11/2024 15:20:01",
      "content": "<p>Looking forward to your code, thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3012554": "# Summary\nFirst, we would like to express our gratitude to Kaggle and RSNA for hosting this great competition.<br/>\nOur solution is an ensemble of 1 stage and 2 stage.\n\n# lhwcv part\nI'm very grateful to my teammates, I learned a lot from them. <br/>\n\n2 stage methods, components:\n- **stage1**: keypoint regression (both 2d and 3d)\n- **stage2_model1**: crop by keypoint, fuse 3 view on feature space\n- **stage2_model2**: crop by keypoint, single view for 3 types of condition\n- **ensemble**: various backbone + flip TTA\n\n## stage1\nkeypoint regression (both 2d and 3d) <br/>\n\nSagittal:\n- **keypoint-sag-2d**: use timm model, first sample 10 slice for each series (same methods in ITK's baseline)\n   then input with the center slice.\n- **keypoint-sag-3d**: use timm-3d model,sagittal T2 -> 5 xyz points, T1-> 10 xyz points\n\nAxial:\n- **keypoint-axial-3d**: use timm-3d model, 10 xyz points \n- **keypoint-axial-level-cls**: 2.5d (CNN + LSTM + attention) for level classification (z), \n  and another 2d model for xy regression\n  \nUsed backbones, which is searched by ITK:\n- densenet161\n- convnext small\n- gluon_resnet152_v1s\n- convformer_s36.sail_in22k_ft_in1k_384\n\n\n## stage2_model1\nfuse 3 view on feature space <br/>\nThe keypoint models used are **keypoint-sag-2d** and **keypoint-axial-3d**.  \nFor the sagittal plane, we sample 10 slices for each series, cropping the center 5 slices on T2, and cropping 5 slices from the left and 5 slices from the right on T1, for each level.  \nFor the axial plane, we sample 5 slices for each level based on the Z-coordinate of the point.  \nThe final input size is: x -> **(bs, 5_cond, 5_level, 5, crop_h, crop_w)**.  \nThe model architecture is as follows:\n\n```python\n# train \ny, axial_embs = self.axial_model(axial_x) # bs, \nys2, sag_embs = self.sag_model(x)\n# sag_embs: bs, 5_cond, 1_level, 128\nsag_embs = sag_embs.permute(0, 2, 1, 3).reshape(b, 1, -1)  # bs, 1, 5*128\nembs = torch.cat((sag_embs, axial_embs), dim=-1)\nys = self.out_linear(embs).reshape(b, 1, 5, 3)  # bs, 1_level, 5_cond, 3\nys = ys.permute(0, 2, 1, 3)  # # bs, 5_cond, 1_level, 3\nys = ys + ys2\n```\n\n## stage2_model2\nThe defect of stage2_model1 is that it uses 2D keypoints on the sagittal plane and does not handle multiple series on the axial plane very well.  \nThe keypoint models used are **keypoint-sag-3d** and **keypoint-axial-level-cls**.  \nwe crop the 3D ROI image using (z_imgs, crop_h, crop_w) based on XYZ points, where z_imgs can be either 3 or 5; we found that 3 works best for my purposes.  \nFor the axial plane, there may be multiple detections, but we only used the one with the highest confidence (not necessarily the best).\n\n\n## ensemble\nstage2_model1 + stage2_model2 with many backbones, TTA: hflip for sagittal, vflip for axial <br/>\n**results**: cv: 0.373  LB: 0.35  PB: 0.40\n\n**backbones**: \n- pvt-b1\n- pvt-b2\n- convnext-tiny\n- convnext-small\n- densenet161\n\n\n# hecgck & lhwcv ensemble\nTeam submission notebook can be found at:  \nhttps://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6e4d08b10ead7a1c9b3904f1d83bd626%2FWechatIMG1025.jpg?generation=1728957811388315&alt=media)\nTeam post-submission notebook can be found at:  \nhttps://www.kaggle.com/code/hengck23/post-lhw-v24-ensemble-add-heng\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F07bfb7f227ec83b40bc3b7a92c5f7092%2FWechatIMG1026.jpg?generation=1728957822586591&alt=media)\n\n\n\n# hecgck part\nsee: \nhttps://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539439\n\n## Code\nInfer Notebook:\n- https://www.kaggle.com/code/hengck23/lhw-v24-ensemble-add-heng\nTrain:\n- https://github.com/lhwcv/solution-rsna-2024-lumbar-spine\n- https://github.com/hengck23/solution-rsna-2024-lumbar-spine",
    "3012928": "Congratulations on your seventh place. Do you have a specific notebook",
    "3014773": "Looking forward to your code, thanks."
  },
  "source": "meta"
}