{
  "id": 539587,
  "title": "10th place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539587",
  "author_name": "",
  "post_date": "2024-10-09T16:53:56.075870700Z",
  "votes": 19,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, we would like to thank the hosts and the Kaggle team for organizing this wonderful competition. And to all the participants who fought through the fierce competition, great job. Various solutions have been shared, and there is so much we can learn from them.</p>\n<h1>Pipeline overview</h1>\n<p>Our pipeline consists of the following four components.</p>\n<ol>\n<li>Stage1: Keypoint model</li>\n<li>Stage1: Level crop logic for MRIs</li>\n<li>Stage2: Multi-view Classification model</li>\n<li>Ensemble</li>\n<li>Things did not work</li>\n</ol>\n<h1>Stage1: Key point model</h1>\n<p>At first, we developed a model that predicts keypoints with confidence on each slice image.<br>\nWe used efficientnetv2-m as the backbone.</p>\n<ul>\n<li>sagittal T1/T2 (5-level keypoints and confidence)</li>\n<li>axial: (left / right key points and confidence)<br>\nThese outputs are used to fix the level on cropping images.</li>\n</ul>\n<p>example:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Feb6f65c0c4afc20efb54a9d630434d18%2Fkeypoint.png?generation=1728490803598369&amp;alt=media\" alt=\"keypoint\"></p>\n<h1>Stage1: Level crop logic for MRIs</h1>\n<h2>Crop logic for Sagittal T1 and Sagittal T2/STIR</h2>\n<p>We used SpineNetV2 to crop the Sagittal images. This Model is super incredible and we were able to crop the T12 - S1 vertebrae and the corresponding intervertebral area without any further effort!!<br>\nConsidering that SpineNetV2 makes mistakes in predicting the label of the vertebrae(ex: predict L5-S1 as L4-L5), we used the label predicted by Keypoint model to revise the label.</p>\n<h2>Crop logic for Axial T2</h2>\n<h3>Axial Vertebrae/Disc Detector</h3>\n<p>First, we manually annotated about 16000 images to train a vertebrae/disc yolov10 based detector. With this detector, we can crop the Axial T2 on image level as well as getting the center of vertebrae/disc, which will be the key factor for making Cross-Reference Method work. <br>\nexample:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F5dde79111ceeca57e83596e400417eff%2Faxial-anno.png?generation=1728491272188260&amp;alt=media\" alt=\"axial annotate\"></p>\n<h3>Cross-Reference Method</h3>\n<p>Our solution of cropping Axial T2 slices is based on Ian Pan’s Cross-Reference Images in Different MRI Planes.<br>\nIan Pan’s Cross-Reference Images in Different MRI Planes has shown the possibility of projecting the axial image to the certain y coordinates of the Sagittal Slice. However, using ImagePositionPatient is not enough because of the way Axial T2 MRI is scanned which is shared by hengck23. <br>\nexample of Ian Pan's Notebook: (study id: 38281420)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Fd26d75b3a44dea41cc45109d6c998c77%2Fcross-view-ian-pan.png?generation=1728491685893591&amp;alt=media\" alt=\"cross view ian pan\"></p>\n<p>Instead of ImagePositionPatient, we calculated the “CenterPosition”, which means the Position of the center of vertebrae/disc, using ImagePositionPatient and the center coordinates of vertebrae/disc(the center of bbox by yolov10). With the “CenterPosition”, we managed to accurately project the axial image to Sagittal Slice. </p>\n<pre><code> ():\n    col, row = coordinate\n    \n    row_cosines = np.array(image_orientation[:])\n    col_cosines = np.array(image_orientation[:])\n\n    \n    origin = np.array(image_position)\n\n    \n    row_spacing, col_spacing = pixel_spacing\n\n    spatial_coordinate = origin + row * row_spacing * row_cosines + col * col_spacing * col_cosines\n\n     spatial_coordinate\n</code></pre>\n<p>example of our solution: (study id: 38281420)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F4bad2202c9d1988e2f98c25ecaa948f0%2Fcross-view-koo.png?generation=1728491780710721&amp;alt=media\" alt=\"cross view rihan\"></p>\n<p>The rest thing is simple, we used the vertebrae polygon information generated by SpineNetV2 to figure out the y range(y1, y2) of l{i}_l{i+1}, and axial images whose y coordinates are within (y1, y2) are the member of l{i}_l{i+1}.<br>\nFor training data, we manually fixed about 50 studies which has mistake in level prediction.</p>\n<h1>Stage2: Multi-view Classification model</h1>\n<p>We created a 2.5D model that takes input of the sagittal T1/T2 and axial level crops (voxels) for each study.</p>\n<ul>\n<li>input size: 224 or 336</li>\n<li>depth: 20 (pad if needed)</li>\n<li>Smaller backbone gave better results, so we adopted convnext base, convnext tiny, base, swin base, eva02-small and eca_nfnet_l1.</li>\n<li>Augmentations are also very important in this competition to avoid overfitting.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F6f3422a6a94be8cc264def9d582eeaf5%2Frsna2024-pipeline.png?generation=1728491933078740&amp;alt=media\" alt=\"rsna2024-pipeline\"></p>\n<p>Judging from reading other participants' solutions, I regret that we should have tried more to create models that use single view as input.</p>\n<h1>Ensemble</h1>\n<p>We adjusted ensemble weights by each symptom. Ensemble by probability gave a better cv than ensemble by logit. Our best cv is 0.3758, which is also the best score in private LB.<br>\nThe weight for each model is below:</p>\n<p><strong>spinal(cv:0.2419, any sever cv: 0.2465):</strong></p>\n<ul>\n<li>convnext base(yiemon): 0.07</li>\n<li>swin base(yiemon): 0.16</li>\n<li>convnext tiny(yiemon): 0.11</li>\n<li>eva02 small(yiemon): 0.13</li>\n<li>convnext base(rihan):0.13</li>\n<li>eva02 small(no mixup, rihan): 0.07</li>\n<li>eva02 small(mixup, rihan): 0.07</li>\n<li>eca nfnet l1(rihan): 0.11</li>\n<li>swin base(rihan): 0.15</li>\n</ul>\n<p><strong>foraminal(cv: 0.4744):</strong></p>\n<ul>\n<li>convnext base(yiemon): 0.07</li>\n<li>swin base(yiemon): 0.15</li>\n<li>convnext tiny(yiemon): 0.07</li>\n<li>eva02 small(yiemon): 0.06</li>\n<li>convnext base(rihan):0.12</li>\n<li>eva02 small(no mixup, rihan): 0.12</li>\n<li>eva02 small(mixup, rihan): 0.12</li>\n<li>eca nfnet l1(rihan): 0.15</li>\n<li>swin base(rihan): 0.14</li>\n</ul>\n<p><strong>subarticular(cv: 0.5403):</strong></p>\n<ul>\n<li>convnext base(yiemon): 0.12</li>\n<li>swin base(yiemon): 0.11</li>\n<li>convnext tiny(yiemon): 0.07</li>\n<li>eva02 small(yiemon): 0.13</li>\n<li>convnext base(rihan):0.14</li>\n<li>eva02 small(no mixup, rihan): 0.08</li>\n<li>eva02 small(mixup, rihan): 0.12</li>\n<li>eca nfnet l1(rihan): 0.13</li>\n<li>swin base(rihan): 0.1</li>\n</ul>\n<p>Our best result is </p>\n<pre><code>: . \n LB: . \n LB: .\n</code></pre>\n<h1>Things did not work</h1>\n<ul>\n<li>3D model</li>\n<li>pseudo Label and pretraining on external datasets</li>\n<li>Segmentation Model using external dataset</li>\n</ul>",
  "messages": [
    {
      "id": "3013103",
      "postDate": "10/09/2024 16:53:56",
      "content": "<p>First of all, we would like to thank the hosts and the Kaggle team for organizing this wonderful competition. And to all the participants who fought through the fierce competition, great job. Various solutions have been shared, and there is so much we can learn from them.</p>\n<h1>Pipeline overview</h1>\n<p>Our pipeline consists of the following four components.</p>\n<ol>\n<li>Stage1: Keypoint model</li>\n<li>Stage1: Level crop logic for MRIs</li>\n<li>Stage2: Multi-view Classification model</li>\n<li>Ensemble</li>\n<li>Things did not work</li>\n</ol>\n<h1>Stage1: Key point model</h1>\n<p>At first, we developed a model that predicts keypoints with confidence on each slice image.<br>\nWe used efficientnetv2-m as the backbone.</p>\n<ul>\n<li>sagittal T1/T2 (5-level keypoints and confidence)</li>\n<li>axial: (left / right key points and confidence)<br>\nThese outputs are used to fix the level on cropping images.</li>\n</ul>\n<p>example:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Feb6f65c0c4afc20efb54a9d630434d18%2Fkeypoint.png?generation=1728490803598369&amp;alt=media\" alt=\"keypoint\"></p>\n<h1>Stage1: Level crop logic for MRIs</h1>\n<h2>Crop logic for Sagittal T1 and Sagittal T2/STIR</h2>\n<p>We used SpineNetV2 to crop the Sagittal images. This Model is super incredible and we were able to crop the T12 - S1 vertebrae and the corresponding intervertebral area without any further effort!!<br>\nConsidering that SpineNetV2 makes mistakes in predicting the label of the vertebrae(ex: predict L5-S1 as L4-L5), we used the label predicted by Keypoint model to revise the label.</p>\n<h2>Crop logic for Axial T2</h2>\n<h3>Axial Vertebrae/Disc Detector</h3>\n<p>First, we manually annotated about 16000 images to train a vertebrae/disc yolov10 based detector. With this detector, we can crop the Axial T2 on image level as well as getting the center of vertebrae/disc, which will be the key factor for making Cross-Reference Method work. <br>\nexample:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F5dde79111ceeca57e83596e400417eff%2Faxial-anno.png?generation=1728491272188260&amp;alt=media\" alt=\"axial annotate\"></p>\n<h3>Cross-Reference Method</h3>\n<p>Our solution of cropping Axial T2 slices is based on Ian Pan’s Cross-Reference Images in Different MRI Planes.<br>\nIan Pan’s Cross-Reference Images in Different MRI Planes has shown the possibility of projecting the axial image to the certain y coordinates of the Sagittal Slice. However, using ImagePositionPatient is not enough because of the way Axial T2 MRI is scanned which is shared by hengck23. <br>\nexample of Ian Pan's Notebook: (study id: 38281420)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Fd26d75b3a44dea41cc45109d6c998c77%2Fcross-view-ian-pan.png?generation=1728491685893591&amp;alt=media\" alt=\"cross view ian pan\"></p>\n<p>Instead of ImagePositionPatient, we calculated the “CenterPosition”, which means the Position of the center of vertebrae/disc, using ImagePositionPatient and the center coordinates of vertebrae/disc(the center of bbox by yolov10). With the “CenterPosition”, we managed to accurately project the axial image to Sagittal Slice. </p>\n<pre><code> ():\n    col, row = coordinate\n    \n    row_cosines = np.array(image_orientation[:])\n    col_cosines = np.array(image_orientation[:])\n\n    \n    origin = np.array(image_position)\n\n    \n    row_spacing, col_spacing = pixel_spacing\n\n    spatial_coordinate = origin + row * row_spacing * row_cosines + col * col_spacing * col_cosines\n\n     spatial_coordinate\n</code></pre>\n<p>example of our solution: (study id: 38281420)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F4bad2202c9d1988e2f98c25ecaa948f0%2Fcross-view-koo.png?generation=1728491780710721&amp;alt=media\" alt=\"cross view rihan\"></p>\n<p>The rest thing is simple, we used the vertebrae polygon information generated by SpineNetV2 to figure out the y range(y1, y2) of l{i}_l{i+1}, and axial images whose y coordinates are within (y1, y2) are the member of l{i}_l{i+1}.<br>\nFor training data, we manually fixed about 50 studies which has mistake in level prediction.</p>\n<h1>Stage2: Multi-view Classification model</h1>\n<p>We created a 2.5D model that takes input of the sagittal T1/T2 and axial level crops (voxels) for each study.</p>\n<ul>\n<li>input size: 224 or 336</li>\n<li>depth: 20 (pad if needed)</li>\n<li>Smaller backbone gave better results, so we adopted convnext base, convnext tiny, base, swin base, eva02-small and eca_nfnet_l1.</li>\n<li>Augmentations are also very important in this competition to avoid overfitting.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F6f3422a6a94be8cc264def9d582eeaf5%2Frsna2024-pipeline.png?generation=1728491933078740&amp;alt=media\" alt=\"rsna2024-pipeline\"></p>\n<p>Judging from reading other participants' solutions, I regret that we should have tried more to create models that use single view as input.</p>\n<h1>Ensemble</h1>\n<p>We adjusted ensemble weights by each symptom. Ensemble by probability gave a better cv than ensemble by logit. Our best cv is 0.3758, which is also the best score in private LB.<br>\nThe weight for each model is below:</p>\n<p><strong>spinal(cv:0.2419, any sever cv: 0.2465):</strong></p>\n<ul>\n<li>convnext base(yiemon): 0.07</li>\n<li>swin base(yiemon): 0.16</li>\n<li>convnext tiny(yiemon): 0.11</li>\n<li>eva02 small(yiemon): 0.13</li>\n<li>convnext base(rihan):0.13</li>\n<li>eva02 small(no mixup, rihan): 0.07</li>\n<li>eva02 small(mixup, rihan): 0.07</li>\n<li>eca nfnet l1(rihan): 0.11</li>\n<li>swin base(rihan): 0.15</li>\n</ul>\n<p><strong>foraminal(cv: 0.4744):</strong></p>\n<ul>\n<li>convnext base(yiemon): 0.07</li>\n<li>swin base(yiemon): 0.15</li>\n<li>convnext tiny(yiemon): 0.07</li>\n<li>eva02 small(yiemon): 0.06</li>\n<li>convnext base(rihan):0.12</li>\n<li>eva02 small(no mixup, rihan): 0.12</li>\n<li>eva02 small(mixup, rihan): 0.12</li>\n<li>eca nfnet l1(rihan): 0.15</li>\n<li>swin base(rihan): 0.14</li>\n</ul>\n<p><strong>subarticular(cv: 0.5403):</strong></p>\n<ul>\n<li>convnext base(yiemon): 0.12</li>\n<li>swin base(yiemon): 0.11</li>\n<li>convnext tiny(yiemon): 0.07</li>\n<li>eva02 small(yiemon): 0.13</li>\n<li>convnext base(rihan):0.14</li>\n<li>eva02 small(no mixup, rihan): 0.08</li>\n<li>eva02 small(mixup, rihan): 0.12</li>\n<li>eca nfnet l1(rihan): 0.13</li>\n<li>swin base(rihan): 0.1</li>\n</ul>\n<p>Our best result is </p>\n<pre><code>: . \n LB: . \n LB: .\n</code></pre>\n<h1>Things did not work</h1>\n<ul>\n<li>3D model</li>\n<li>pseudo Label and pretraining on external datasets</li>\n<li>Segmentation Model using external dataset</li>\n</ul>",
      "rawMarkdown": "First of all, we would like to thank the hosts and the Kaggle team for organizing this wonderful competition. And to all the participants who fought through the fierce competition, great job. Various solutions have been shared, and there is so much we can learn from them.\n\n# Pipeline overview\nOur pipeline consists of the following four components.\n1. Stage1: Keypoint model\n2. Stage1: Level crop logic for MRIs\n3. Stage2: Multi-view Classification model\n4. Ensemble\n5. Things did not work\n\n# Stage1: Key point model\nAt first, we developed a model that predicts keypoints with confidence on each slice image.\nWe used efficientnetv2-m as the backbone.\n- sagittal T1/T2 (5-level keypoints and confidence)\n- axial: (left / right key points and confidence)\nThese outputs are used to fix the level on cropping images.\n\nexample:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Feb6f65c0c4afc20efb54a9d630434d18%2Fkeypoint.png?generation=1728490803598369&alt=media\" alt=\"keypoint\" width=\"700\" height=\"700\">\n\n# Stage1: Level crop logic for MRIs\n## Crop logic for Sagittal T1 and Sagittal T2/STIR\nWe used SpineNetV2 to crop the Sagittal images. This Model is super incredible and we were able to crop the T12 - S1 vertebrae and the corresponding intervertebral area without any further effort!!\nConsidering that SpineNetV2 makes mistakes in predicting the label of the vertebrae(ex: predict L5-S1 as L4-L5), we used the label predicted by Keypoint model to revise the label.\n\n## Crop logic for Axial T2\n### Axial Vertebrae/Disc Detector\nFirst, we manually annotated about 16000 images to train a vertebrae/disc yolov10 based detector. With this detector, we can crop the Axial T2 on image level as well as getting the center of vertebrae/disc, which will be the key factor for making Cross-Reference Method work. \nexample:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F5dde79111ceeca57e83596e400417eff%2Faxial-anno.png?generation=1728491272188260&alt=media\" alt=\"axial annotate\" width=\"300\" height=\"250\">\n\n### Cross-Reference Method\nOur solution of cropping Axial T2 slices is based on Ian Pan’s Cross-Reference Images in Different MRI Planes.\nIan Pan’s Cross-Reference Images in Different MRI Planes has shown the possibility of projecting the axial image to the certain y coordinates of the Sagittal Slice. However, using ImagePositionPatient is not enough because of the way Axial T2 MRI is scanned which is shared by hengck23. \nexample of Ian Pan's Notebook: (study id: 38281420)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Fd26d75b3a44dea41cc45109d6c998c77%2Fcross-view-ian-pan.png?generation=1728491685893591&alt=media\" alt=\"cross view ian pan\" width=\"300\" height=\"300\">\n\nInstead of ImagePositionPatient, we calculated the “CenterPosition”, which means the Position of the center of vertebrae/disc, using ImagePositionPatient and the center coordinates of vertebrae/disc(the center of bbox by yolov10). With the “CenterPosition”, we managed to accurately project the axial image to Sagittal Slice. \n```python\ndef calculate_spatial_coordinates(coordinate, image_orientation, image_position, pixel_spacing):\n    col, row = coordinate\n    # ImageOrientationPatient\n    row_cosines = np.array(image_orientation[:3])\n    col_cosines = np.array(image_orientation[3:])\n   \n    # ImagePositionPatient\n    origin = np.array(image_position)\n   \n    # PixelSpacing\n    row_spacing, col_spacing = pixel_spacing\n   \n    spatial_coordinate = origin + row * row_spacing * row_cosines + col * col_spacing * col_cosines\n   \n    return spatial_coordinate\n```\nexample of our solution: (study id: 38281420)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F4bad2202c9d1988e2f98c25ecaa948f0%2Fcross-view-koo.png?generation=1728491780710721&alt=media\" alt=\"cross view rihan\" width=\"300\" height=\"300\">\n\nThe rest thing is simple, we used the vertebrae polygon information generated by SpineNetV2 to figure out the y range(y1, y2) of l{i}_l{i+1}, and axial images whose y coordinates are within (y1, y2) are the member of l{i}_l{i+1}.\nFor training data, we manually fixed about 50 studies which has mistake in level prediction.\n\n# Stage2: Multi-view Classification model\nWe created a 2.5D model that takes input of the sagittal T1/T2 and axial level crops (voxels) for each study.\n- input size: 224 or 336\n- depth: 20 (pad if needed)\n- Smaller backbone gave better results, so we adopted convnext base, convnext tiny, base, swin base, eva02-small and eca_nfnet_l1.\n- Augmentations are also very important in this competition to avoid overfitting.\n\n![rsna2024-pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F6f3422a6a94be8cc264def9d582eeaf5%2Frsna2024-pipeline.png?generation=1728491933078740&alt=media)\n\nJudging from reading other participants' solutions, I regret that we should have tried more to create models that use single view as input.\n\n# Ensemble\nWe adjusted ensemble weights by each symptom. Ensemble by probability gave a better cv than ensemble by logit. Our best cv is 0.3758, which is also the best score in private LB.\nThe weight for each model is below:\n\n**spinal(cv:0.2419, any sever cv: 0.2465):**\n- convnext base(yiemon): 0.07\n- swin base(yiemon): 0.16\n- convnext tiny(yiemon): 0.11\n- eva02 small(yiemon): 0.13\n- convnext base(rihan):0.13\n- eva02 small(no mixup, rihan): 0.07\n- eva02 small(mixup, rihan): 0.07\n- eca nfnet l1(rihan): 0.11\n- swin base(rihan): 0.15\n\n\n**foraminal(cv: 0.4744):**\n- convnext base(yiemon): 0.07\n- swin base(yiemon): 0.15\n- convnext tiny(yiemon): 0.07\n- eva02 small(yiemon): 0.06\n- convnext base(rihan):0.12\n- eva02 small(no mixup, rihan): 0.12\n- eva02 small(mixup, rihan): 0.12\n- eca nfnet l1(rihan): 0.15\n- swin base(rihan): 0.14\n\n\n**subarticular(cv: 0.5403):**\n- convnext base(yiemon): 0.12\n- swin base(yiemon): 0.11\n- convnext tiny(yiemon): 0.07\n- eva02 small(yiemon): 0.13\n- convnext base(rihan):0.14\n- eva02 small(no mixup, rihan): 0.08\n- eva02 small(mixup, rihan): 0.12\n- eca nfnet l1(rihan): 0.13\n- swin base(rihan): 0.1\n\nOur best result is \n```\nCV: 0.3758 \nPublic LB: 0.34 \nPrivate LB: 0.40\n```\n# Things did not work\n- 3D model\n- pseudo Label and pretraining on external datasets\n- Segmentation Model using external dataset",
      "votes": null
    },
    {
      "id": "3032530",
      "postDate": "10/31/2024 02:48:07",
      "content": "<p>I have learned a great deal from your solution.<br>\nIf possible, would you be willing to share your training code?<br>\nI am very interested in your pipeline and would like to learn about the detailed processing from the actual implementation.</p>",
      "rawMarkdown": "I have learned a great deal from your solution.\nIf possible, would you be willing to share your training code?\nI am very interested in your pipeline and would like to learn about the detailed processing from the actual implementation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3032530,
      "author_name": "miyamotodaiya",
      "author_url": "",
      "post_date": "10/31/2024 02:48:07",
      "content": "<p>I have learned a great deal from your solution.<br>\nIf possible, would you be willing to share your training code?<br>\nI am very interested in your pipeline and would like to learn about the detailed processing from the actual implementation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3013103": "First of all, we would like to thank the hosts and the Kaggle team for organizing this wonderful competition. And to all the participants who fought through the fierce competition, great job. Various solutions have been shared, and there is so much we can learn from them.\n\n# Pipeline overview\nOur pipeline consists of the following four components.\n1. Stage1: Keypoint model\n2. Stage1: Level crop logic for MRIs\n3. Stage2: Multi-view Classification model\n4. Ensemble\n5. Things did not work\n\n# Stage1: Key point model\nAt first, we developed a model that predicts keypoints with confidence on each slice image.\nWe used efficientnetv2-m as the backbone.\n- sagittal T1/T2 (5-level keypoints and confidence)\n- axial: (left / right key points and confidence)\nThese outputs are used to fix the level on cropping images.\n\nexample:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Feb6f65c0c4afc20efb54a9d630434d18%2Fkeypoint.png?generation=1728490803598369&alt=media\" alt=\"keypoint\" width=\"700\" height=\"700\">\n\n# Stage1: Level crop logic for MRIs\n## Crop logic for Sagittal T1 and Sagittal T2/STIR\nWe used SpineNetV2 to crop the Sagittal images. This Model is super incredible and we were able to crop the T12 - S1 vertebrae and the corresponding intervertebral area without any further effort!!\nConsidering that SpineNetV2 makes mistakes in predicting the label of the vertebrae(ex: predict L5-S1 as L4-L5), we used the label predicted by Keypoint model to revise the label.\n\n## Crop logic for Axial T2\n### Axial Vertebrae/Disc Detector\nFirst, we manually annotated about 16000 images to train a vertebrae/disc yolov10 based detector. With this detector, we can crop the Axial T2 on image level as well as getting the center of vertebrae/disc, which will be the key factor for making Cross-Reference Method work. \nexample:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F5dde79111ceeca57e83596e400417eff%2Faxial-anno.png?generation=1728491272188260&alt=media\" alt=\"axial annotate\" width=\"300\" height=\"250\">\n\n### Cross-Reference Method\nOur solution of cropping Axial T2 slices is based on Ian Pan’s Cross-Reference Images in Different MRI Planes.\nIan Pan’s Cross-Reference Images in Different MRI Planes has shown the possibility of projecting the axial image to the certain y coordinates of the Sagittal Slice. However, using ImagePositionPatient is not enough because of the way Axial T2 MRI is scanned which is shared by hengck23. \nexample of Ian Pan's Notebook: (study id: 38281420)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2Fd26d75b3a44dea41cc45109d6c998c77%2Fcross-view-ian-pan.png?generation=1728491685893591&alt=media\" alt=\"cross view ian pan\" width=\"300\" height=\"300\">\n\nInstead of ImagePositionPatient, we calculated the “CenterPosition”, which means the Position of the center of vertebrae/disc, using ImagePositionPatient and the center coordinates of vertebrae/disc(the center of bbox by yolov10). With the “CenterPosition”, we managed to accurately project the axial image to Sagittal Slice. \n```python\ndef calculate_spatial_coordinates(coordinate, image_orientation, image_position, pixel_spacing):\n    col, row = coordinate\n    # ImageOrientationPatient\n    row_cosines = np.array(image_orientation[:3])\n    col_cosines = np.array(image_orientation[3:])\n   \n    # ImagePositionPatient\n    origin = np.array(image_position)\n   \n    # PixelSpacing\n    row_spacing, col_spacing = pixel_spacing\n   \n    spatial_coordinate = origin + row * row_spacing * row_cosines + col * col_spacing * col_cosines\n   \n    return spatial_coordinate\n```\nexample of our solution: (study id: 38281420)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F4bad2202c9d1988e2f98c25ecaa948f0%2Fcross-view-koo.png?generation=1728491780710721&alt=media\" alt=\"cross view rihan\" width=\"300\" height=\"300\">\n\nThe rest thing is simple, we used the vertebrae polygon information generated by SpineNetV2 to figure out the y range(y1, y2) of l{i}_l{i+1}, and axial images whose y coordinates are within (y1, y2) are the member of l{i}_l{i+1}.\nFor training data, we manually fixed about 50 studies which has mistake in level prediction.\n\n# Stage2: Multi-view Classification model\nWe created a 2.5D model that takes input of the sagittal T1/T2 and axial level crops (voxels) for each study.\n- input size: 224 or 336\n- depth: 20 (pad if needed)\n- Smaller backbone gave better results, so we adopted convnext base, convnext tiny, base, swin base, eva02-small and eca_nfnet_l1.\n- Augmentations are also very important in this competition to avoid overfitting.\n\n![rsna2024-pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5688805%2F6f3422a6a94be8cc264def9d582eeaf5%2Frsna2024-pipeline.png?generation=1728491933078740&alt=media)\n\nJudging from reading other participants' solutions, I regret that we should have tried more to create models that use single view as input.\n\n# Ensemble\nWe adjusted ensemble weights by each symptom. Ensemble by probability gave a better cv than ensemble by logit. Our best cv is 0.3758, which is also the best score in private LB.\nThe weight for each model is below:\n\n**spinal(cv:0.2419, any sever cv: 0.2465):**\n- convnext base(yiemon): 0.07\n- swin base(yiemon): 0.16\n- convnext tiny(yiemon): 0.11\n- eva02 small(yiemon): 0.13\n- convnext base(rihan):0.13\n- eva02 small(no mixup, rihan): 0.07\n- eva02 small(mixup, rihan): 0.07\n- eca nfnet l1(rihan): 0.11\n- swin base(rihan): 0.15\n\n\n**foraminal(cv: 0.4744):**\n- convnext base(yiemon): 0.07\n- swin base(yiemon): 0.15\n- convnext tiny(yiemon): 0.07\n- eva02 small(yiemon): 0.06\n- convnext base(rihan):0.12\n- eva02 small(no mixup, rihan): 0.12\n- eva02 small(mixup, rihan): 0.12\n- eca nfnet l1(rihan): 0.15\n- swin base(rihan): 0.14\n\n\n**subarticular(cv: 0.5403):**\n- convnext base(yiemon): 0.12\n- swin base(yiemon): 0.11\n- convnext tiny(yiemon): 0.07\n- eva02 small(yiemon): 0.13\n- convnext base(rihan):0.14\n- eva02 small(no mixup, rihan): 0.08\n- eva02 small(mixup, rihan): 0.12\n- eca nfnet l1(rihan): 0.13\n- swin base(rihan): 0.1\n\nOur best result is \n```\nCV: 0.3758 \nPublic LB: 0.34 \nPrivate LB: 0.40\n```\n# Things did not work\n- 3D model\n- pseudo Label and pretraining on external datasets\n- Segmentation Model using external dataset",
    "3032530": "I have learned a great deal from your solution.\nIf possible, would you be willing to share your training code?\nI am very interested in your pipeline and would like to learn about the detailed processing from the actual implementation."
  },
  "source": "meta"
}