{
  "id": 539459,
  "title": "14th place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/itysh-14th-place-solution",
  "author_name": "",
  "post_date": "2024-10-09T04:06:49.513Z",
  "votes": 31,
  "comment_count": 2,
  "views": 0,
  "content": "<h3>Models</h3>\n<ul>\n<li>Single-stage multi-condition (s12): 0.412 CV</li>\n<li>Multi-stage multi-condition (s1+s2): 0.404 CV</li>\n<li>Multi-stage single-condition (@tamotamo shared it <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539459#3012444\" target=\"_blank\">here</a>): 0.400 CV</li>\n</ul>\n<h3>Data</h3>\n<p>CLAHE normalization of the input.  I projected input points of interest provided by orgs for specific sources and can locate all 25 points in each of the inputs. I used <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/528653\" target=\"_blank\">publicly shared key point annotation</a> for a two-stage setup. For external data, I used <a href=\"https://huggingface.co/datasets/cdoswald/SPIDER\" target=\"_blank\">Spider </a>(dataset for segmentation of specific vertebra and disks) in my very first experiments and single-stage model pre-training.</p>\n<h3>Single-stage model (s12)</h3>\n<p><strong>Architecture</strong>:  Attention-based aggregation (similar to SED used in Bird-call competition): I predict 3d voxel grid (at reduced res corresponding to 1/token size) of attention weights and predicted targets. Then I apply spatial softmax on attention and weight the targets accordingly summing them along the spatial dimensions. It results in scalar Bx25 predictions. I use <a href=\"https://www.kaggle.com/code/junkoda/optimize-the-evaluation-metric\" target=\"_blank\">competition metric approximation</a> as the loss function. In addition, I consider aux loss on attention using Gaussians to approximate annotated key points.<br>\nAs an alternative, I consider 2 cross-attention layer decoder applied to each slice independently. The decoder takes 25 learnable queries as the input and each predicts the target, xy position, and probability p that this slice is valid for a particular target. The final prediction is weighted (based on p) target and xy. The loss is CE on p, Competition metric approximation on the target, and MSE on xy. The idea here is that the decoder trying to predict the position of the specific key point is also aggregating the information providing the target class. This approach performs similarly to attention-based aggregation.</p>\n<p><strong>Backbone</strong>: Since the model must be able to predict specific levels for visually indistinguishable vertebra it should have access to the global image content (i.e. ViT-based architecture is preferable), and in early experiments with Spider I saw that DINOv2 (with registers) can assign levels to vertebra reliably. Therefore, I used DINOv2 B backbone taking a sequence of input frames (N,3,H,W). For side views the backbone is augmented with <strong>zero initialized LSTM adapters</strong> to perform sequence mixing and is pre-trained on the Spider segmentation dataset. For axial views, I added a <strong>LSTM mixing layer</strong> after the backbone.</p>\n<p>The image setup: 16x448x448 for side input and 48x332x332 for axial input. If the sequence is shorter, images are repeated.</p>\n<h3>2-stage multi-condition (s1+s2)</h3>\n<p>s1 is a simple single-slice segmentation model applied to the center slice and trained to predict 10 key points shared publicly. I used Sagittal T1 + Sagittal T2/STIR sources and then reprojected corresponding coordinates to find appropriate slices in Axial T2 and also cross-check the sources and drop points in the side view. I do not do s1 on axial view because multiple groups make the data messy and the task is quite harder than point detection in a lateral view. The backbone is DINOv2 again because I need to predict 10 distinguishable key points. With 10 instead of 5 key poins I can derive the pox orientation and the size. Input size 448x448.</p>\n<p>Then I apply s2 for each source. In the side view I crop a small stack of boxes around proper areas of interest. As a result, I produce 5 stacks of image crops 5x16x3x192x192 for side views. Axial cropping is done based on the plane corresponding to the projected key point and is not so aggressive laterally. So I produce sequences of 5x8x3x320x320 (making sure the group=view angle is the same for all selected images, and replicate images if the number of selected slices is insufficient). The model is simple ConvNeXtv2 nano + LSTM mixing layer + concat pooling over sequence + the head for side view, and in axial view I use DINOv2 instead. The competition metric approximation is used as the loss function.</p>\n<h3>Aggregation</h3>\n<p>I use a simple weighting of each model independent for each condition (apply to logits). The matrix below shows the contribution of each of 7 sources (s2 Sagittal T2/STIR, s2 Sagittal T1, s2 Axial T2, s12  Sagittal T2/STIR, s2 Sagittal T1, s2 Axial, Multi-stage single-condition) to 5 targets (SCS, L NFN, R NFN, L SS, RSS). Some sources are particularly important for specific targets. <br>\nw =    [[0.1228, 0.0025, 0.0031, 0.0612, 0.1021],<br>\n        [0.0549, 0.2573, 0.1362, 0.0484, 0.1198],<br>\n        [0.1712, 0.0790, 0.0269, 0.2366, 0.1519],<br>\n        [0.1308, 0.0068, 0.0373, 0.1285, 0.0949],<br>\n        [0.0161, 0.1436, 0.2533, 0.0101, 0.0072],<br>\n        [0.0689, 0.0220, 0.0208, 0.0797, 0.1492],<br>\n        [0.4353, 0.4888, 0.5225, 0.4354, 0.3749]]<br>\nCV 0.379, LB 0.35/0.41</p>",
  "messages": [
    {
      "id": "3012411",
      "postDate": "10/09/2024 02:55:36",
      "content": "<h3>Models</h3>\n<ul>\n<li>Single-stage multi-condition (s12): 0.412 CV</li>\n<li>Multi-stage multi-condition (s1+s2): 0.404 CV</li>\n<li>Multi-stage single-condition (@tamotamo shared it <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539459#3012444\" target=\"_blank\">here</a>): 0.400 CV</li>\n</ul>\n<h3>Data</h3>\n<p>CLAHE normalization of the input.  I projected input points of interest provided by orgs for specific sources and can locate all 25 points in each of the inputs. I used <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/528653\" target=\"_blank\">publicly shared key point annotation</a> for a two-stage setup. For external data, I used <a href=\"https://huggingface.co/datasets/cdoswald/SPIDER\" target=\"_blank\">Spider </a>(dataset for segmentation of specific vertebra and disks) in my very first experiments and single-stage model pre-training.</p>\n<h3>Single-stage model (s12)</h3>\n<p><strong>Architecture</strong>:  Attention-based aggregation (similar to SED used in Bird-call competition): I predict 3d voxel grid (at reduced res corresponding to 1/token size) of attention weights and predicted targets. Then I apply spatial softmax on attention and weight the targets accordingly summing them along the spatial dimensions. It results in scalar Bx25 predictions. I use <a href=\"https://www.kaggle.com/code/junkoda/optimize-the-evaluation-metric\" target=\"_blank\">competition metric approximation</a> as the loss function. In addition, I consider aux loss on attention using Gaussians to approximate annotated key points.<br>\nAs an alternative, I consider 2 cross-attention layer decoder applied to each slice independently. The decoder takes 25 learnable queries as the input and each predicts the target, xy position, and probability p that this slice is valid for a particular target. The final prediction is weighted (based on p) target and xy. The loss is CE on p, Competition metric approximation on the target, and MSE on xy. The idea here is that the decoder trying to predict the position of the specific key point is also aggregating the information providing the target class. This approach performs similarly to attention-based aggregation.</p>\n<p><strong>Backbone</strong>: Since the model must be able to predict specific levels for visually indistinguishable vertebra it should have access to the global image content (i.e. ViT-based architecture is preferable), and in early experiments with Spider I saw that DINOv2 (with registers) can assign levels to vertebra reliably. Therefore, I used DINOv2 B backbone taking a sequence of input frames (N,3,H,W). For side views the backbone is augmented with <strong>zero initialized LSTM adapters</strong> to perform sequence mixing and is pre-trained on the Spider segmentation dataset. For axial views, I added a <strong>LSTM mixing layer</strong> after the backbone.</p>\n<p>The image setup: 16x448x448 for side input and 48x332x332 for axial input. If the sequence is shorter, images are repeated.</p>\n<h3>2-stage multi-condition (s1+s2)</h3>\n<p>s1 is a simple single-slice segmentation model applied to the center slice and trained to predict 10 key points shared publicly. I used Sagittal T1 + Sagittal T2/STIR sources and then reprojected corresponding coordinates to find appropriate slices in Axial T2 and also cross-check the sources and drop points in the side view. I do not do s1 on axial view because multiple groups make the data messy and the task is quite harder than point detection in a lateral view. The backbone is DINOv2 again because I need to predict 10 distinguishable key points. With 10 instead of 5 key poins I can derive the pox orientation and the size. Input size 448x448.</p>\n<p>Then I apply s2 for each source. In the side view I crop a small stack of boxes around proper areas of interest. As a result, I produce 5 stacks of image crops 5x16x3x192x192 for side views. Axial cropping is done based on the plane corresponding to the projected key point and is not so aggressive laterally. So I produce sequences of 5x8x3x320x320 (making sure the group=view angle is the same for all selected images, and replicate images if the number of selected slices is insufficient). The model is simple ConvNeXtv2 nano + LSTM mixing layer + concat pooling over sequence + the head for side view, and in axial view I use DINOv2 instead. The competition metric approximation is used as the loss function.</p>\n<h3>Aggregation</h3>\n<p>I use a simple weighting of each model independent for each condition (apply to logits). The matrix below shows the contribution of each of 7 sources (s2 Sagittal T2/STIR, s2 Sagittal T1, s2 Axial T2, s12  Sagittal T2/STIR, s2 Sagittal T1, s2 Axial, Multi-stage single-condition) to 5 targets (SCS, L NFN, R NFN, L SS, RSS). Some sources are particularly important for specific targets. <br>\nw =    [[0.1228, 0.0025, 0.0031, 0.0612, 0.1021],<br>\n        [0.0549, 0.2573, 0.1362, 0.0484, 0.1198],<br>\n        [0.1712, 0.0790, 0.0269, 0.2366, 0.1519],<br>\n        [0.1308, 0.0068, 0.0373, 0.1285, 0.0949],<br>\n        [0.0161, 0.1436, 0.2533, 0.0101, 0.0072],<br>\n        [0.0689, 0.0220, 0.0208, 0.0797, 0.1492],<br>\n        [0.4353, 0.4888, 0.5225, 0.4354, 0.3749]]<br>\nCV 0.379, LB 0.35/0.41</p>",
      "rawMarkdown": "### Models \n- Single-stage multi-condition (s12): 0.412 CV\n- Multi-stage multi-condition (s1+s2): 0.404 CV\n- Multi-stage single-condition (@tamotamo shared it [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539459#3012444)): 0.400 CV\n\n### Data\nCLAHE normalization of the input.  I projected input points of interest provided by orgs for specific sources and can locate all 25 points in each of the inputs. I used [publicly shared key point annotation](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/528653) for a two-stage setup. For external data, I used [Spider ](https://huggingface.co/datasets/cdoswald/SPIDER)(dataset for segmentation of specific vertebra and disks) in my very first experiments and single-stage model pre-training.\n\n### Single-stage model (s12)\n**Architecture**:  Attention-based aggregation (similar to SED used in Bird-call competition): I predict 3d voxel grid (at reduced res corresponding to 1/token size) of attention weights and predicted targets. Then I apply spatial softmax on attention and weight the targets accordingly summing them along the spatial dimensions. It results in scalar Bx25 predictions. I use [competition metric approximation](https://www.kaggle.com/code/junkoda/optimize-the-evaluation-metric) as the loss function. In addition, I consider aux loss on attention using Gaussians to approximate annotated key points.\nAs an alternative, I consider 2 cross-attention layer decoder applied to each slice independently. The decoder takes 25 learnable queries as the input and each predicts the target, xy position, and probability p that this slice is valid for a particular target. The final prediction is weighted (based on p) target and xy. The loss is CE on p, Competition metric approximation on the target, and MSE on xy. The idea here is that the decoder trying to predict the position of the specific key point is also aggregating the information providing the target class. This approach performs similarly to attention-based aggregation.\n\n**Backbone**: Since the model must be able to predict specific levels for visually indistinguishable vertebra it should have access to the global image content (i.e. ViT-based architecture is preferable), and in early experiments with Spider I saw that DINOv2 (with registers) can assign levels to vertebra reliably. Therefore, I used DINOv2 B backbone taking a sequence of input frames (N,3,H,W). For side views the backbone is augmented with **zero initialized LSTM adapters** to perform sequence mixing and is pre-trained on the Spider segmentation dataset. For axial views, I added a **LSTM mixing layer** after the backbone.\n\nThe image setup: 16x448x448 for side input and 48x332x332 for axial input. If the sequence is shorter, images are repeated.\n\n### 2-stage multi-condition (s1+s2)\ns1 is a simple single-slice segmentation model applied to the center slice and trained to predict 10 key points shared publicly. I used Sagittal T1 + Sagittal T2/STIR sources and then reprojected corresponding coordinates to find appropriate slices in Axial T2 and also cross-check the sources and drop points in the side view. I do not do s1 on axial view because multiple groups make the data messy and the task is quite harder than point detection in a lateral view. The backbone is DINOv2 again because I need to predict 10 distinguishable key points. With 10 instead of 5 key poins I can derive the pox orientation and the size. Input size 448x448.\n\nThen I apply s2 for each source. In the side view I crop a small stack of boxes around proper areas of interest. As a result, I produce 5 stacks of image crops 5x16x3x192x192 for side views. Axial cropping is done based on the plane corresponding to the projected key point and is not so aggressive laterally. So I produce sequences of 5x8x3x320x320 (making sure the group=view angle is the same for all selected images, and replicate images if the number of selected slices is insufficient). The model is simple ConvNeXtv2 nano + LSTM mixing layer + concat pooling over sequence + the head for side view, and in axial view I use DINOv2 instead. The competition metric approximation is used as the loss function.\n\n### Aggregation\nI use a simple weighting of each model independent for each condition (apply to logits). The matrix below shows the contribution of each of 7 sources (s2 Sagittal T2/STIR, s2 Sagittal T1, s2 Axial T2, s12  Sagittal T2/STIR, s2 Sagittal T1, s2 Axial, Multi-stage single-condition) to 5 targets (SCS, L NFN, R NFN, L SS, RSS). Some sources are particularly important for specific targets. \nw =    [[0.1228, 0.0025, 0.0031, 0.0612, 0.1021],\n        [0.0549, 0.2573, 0.1362, 0.0484, 0.1198],\n        [0.1712, 0.0790, 0.0269, 0.2366, 0.1519],\n        [0.1308, 0.0068, 0.0373, 0.1285, 0.0949],\n        [0.0161, 0.1436, 0.2533, 0.0101, 0.0072],\n        [0.0689, 0.0220, 0.0208, 0.0797, 0.1492],\n        [0.4353, 0.4888, 0.5225, 0.4354, 0.3749]]\nCV 0.379, LB 0.35/0.41",
      "votes": null
    },
    {
      "id": "3012444",
      "postDate": "10/09/2024 03:39:54",
      "content": "<h1><strong>Summray</strong></h1>\n<p>My pipeline consists of three steps: 1) Slice position (Z) detection, 2) Coordinates (XY) detection, and 3) classification. All steps are performed independently for the three types of conditions: spinal canal stenosis (SS), neural foraminal narrowing (NFN), and subarticular stenosis (SS). Additionally, NFN, SCS, and SS correspond to the image types Sagittal T1, Sagittal T2, and Axial T2, respectively.</p>\n<h2><strong>1. Slice position (Z) detection</strong></h2>\n<p>This model can detect the slices where disease exists. Slice positions to detect is from train_label_coordinates.csv</p>\n<ul>\n<li>input (N, C, H, W)<ul>\n<li>The C are neighboring images</li>\n<li>The N refers to the number of images obtained at regular intervals in the stack.</li>\n<li>Sagittal T1/T2 (10, 5, 224, 224)</li>\n<li>Axial T2 (20, 3, 224, 224)</li></ul></li>\n<li>Backbone<ul>\n<li>efficientnet_b0</li>\n<li>convnextv2_tiny</li></ul></li>\n<li>model structure<ul>\n<li>The images are encoded to obtain BxN×D feature vectors.</li>\n<li>The model is split into two branches: one for level classifier and one for right/left classifier.<ul>\n<li>BxNx1280-&gt;LSTM-&gt;reshape(BNx1280)-&gt;classification head</li></ul></li></ul></li>\n</ul>\n<h2><strong>2. Coordinates (XY) detection</strong></h2>\n<p>In step 1, I detected the slices to use, i.e., determined the Z position for analysis. In step 2, I made segmentation models to detect XY coordinates.</p>\n<ul>\n<li>input (C, H, W)<ul>\n<li>The C are neighboring images</li>\n<li>(3, 224, 224)</li></ul></li>\n<li>output (C, H, W)<ul>\n<li>(2, 224, 224) for SS (Axial T2)<ul>\n<li>Channels correspond to Left/Right</li></ul></li>\n<li>(5, 224, 224) for NFN (Sagittal T1) and SCS (Sagittal T2)<ul>\n<li>Channels correspond to L1/L2…L5/S1</li></ul></li></ul></li>\n<li>Model<ul>\n<li>Unet from segmentation_models_pytorch</li>\n<li>resnet18</li></ul></li>\n</ul>\n<h2><strong>3. Classification</strong></h2>\n<p>I could predict XYZ coordinates for levels and left/right from step 1 and 2. In this step, I made model to classify 3 classes (Normal/Moderate/Severe).</p>\n<ul>\n<li>Detection of XYZ coordinates<ul>\n<li>Determine candidates of Z positions from step 1<ul>\n<li>This Z position was not so accurate, I used these positions as start positions to search accurate positions from step 2 mask</li></ul></li>\n<li>Search best Z positions and determine XY<ul>\n<li>Calculate mask areas from neighboring slices and the slice with the largest area assigned as the best position</li>\n<li>Calculate centroid to determine XY</li></ul></li></ul></li>\n<li>Input<ul>\n<li>CenterCrop (3, 128, 110), Resize(160, 160) for SS</li>\n<li>CenterCrop (3, 90, 128), Resize(160, 160) for NFN and SCS</li></ul></li>\n<li>Backbone<ul>\n<li>efficientnet_b0</li>\n<li>efficentnetv2s</li>\n<li>convnextv2_tiny</li>\n<li>convnextv2_nano</li></ul></li>\n</ul>",
      "rawMarkdown": "#  **Summray**\nMy pipeline consists of three steps: 1) Slice position (Z) detection, 2) Coordinates (XY) detection, and 3) classification. All steps are performed independently for the three types of conditions: spinal canal stenosis (SS), neural foraminal narrowing (NFN), and subarticular stenosis (SS). Additionally, NFN, SCS, and SS correspond to the image types Sagittal T1, Sagittal T2, and Axial T2, respectively.\n\n## **1. Slice position (Z) detection**\nThis model can detect the slices where disease exists. Slice positions to detect is from train_label_coordinates.csv\n- input (N, C, H, W)\n  - The C are neighboring images\n  - The N refers to the number of images obtained at regular intervals in the stack.\n  - Sagittal T1/T2 (10, 5, 224, 224)\n  - Axial T2 (20, 3, 224, 224)\n- Backbone\n  - efficientnet_b0\n  - convnextv2_tiny\n- model structure\n  - The images are encoded to obtain BxN×D feature vectors.\n  - The model is split into two branches: one for level classifier and one for right/left classifier.\n      - BxNx1280->LSTM->reshape(BNx1280)->classification head\n\n## **2. Coordinates (XY) detection**\nIn step 1, I detected the slices to use, i.e., determined the Z position for analysis. In step 2, I made segmentation models to detect XY coordinates.\n- input (C, H, W)\n  - The C are neighboring images\n  - (3, 224, 224)\n- output (C, H, W)\n  - (2, 224, 224) for SS (Axial T2)\n      - Channels correspond to Left/Right\n  - (5, 224, 224) for NFN (Sagittal T1) and SCS (Sagittal T2)\n      - Channels correspond to L1/L2...L5/S1\n- Model\n  - Unet from segmentation_models_pytorch\n  - resnet18\n\n## **3. Classification**\nI could predict XYZ coordinates for levels and left/right from step 1 and 2. In this step, I made model to classify 3 classes (Normal/Moderate/Severe).\n- Detection of XYZ coordinates\n    - Determine candidates of Z positions from step 1\n      - This Z position was not so accurate, I used these positions as start positions to search accurate positions from step 2 mask\n    - Search best Z positions and determine XY\n      - Calculate mask areas from neighboring slices and the slice with the largest area assigned as the best position\n        - Calculate centroid to determine XY\n- Input\n  - CenterCrop (3, 128, 110), Resize(160, 160) for SS\n  - CenterCrop (3, 90, 128), Resize(160, 160) for NFN and SCS\n- Backbone\n  - efficientnet_b0\n  - efficentnetv2s\n  - convnextv2_tiny\n  - convnextv2_nano",
      "votes": null
    },
    {
      "id": "3012914",
      "postDate": "10/09/2024 13:51:27",
      "content": "<p>Congratulations on your 14th place. Do you have a specific notebook</p>",
      "rawMarkdown": "Congratulations on your 14th place. Do you have a specific notebook",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3012444,
      "author_name": "tamotamo",
      "author_url": "",
      "post_date": "10/09/2024 03:39:54",
      "content": "<h1><strong>Summray</strong></h1>\n<p>My pipeline consists of three steps: 1) Slice position (Z) detection, 2) Coordinates (XY) detection, and 3) classification. All steps are performed independently for the three types of conditions: spinal canal stenosis (SS), neural foraminal narrowing (NFN), and subarticular stenosis (SS). Additionally, NFN, SCS, and SS correspond to the image types Sagittal T1, Sagittal T2, and Axial T2, respectively.</p>\n<h2><strong>1. Slice position (Z) detection</strong></h2>\n<p>This model can detect the slices where disease exists. Slice positions to detect is from train_label_coordinates.csv</p>\n<ul>\n<li>input (N, C, H, W)<ul>\n<li>The C are neighboring images</li>\n<li>The N refers to the number of images obtained at regular intervals in the stack.</li>\n<li>Sagittal T1/T2 (10, 5, 224, 224)</li>\n<li>Axial T2 (20, 3, 224, 224)</li></ul></li>\n<li>Backbone<ul>\n<li>efficientnet_b0</li>\n<li>convnextv2_tiny</li></ul></li>\n<li>model structure<ul>\n<li>The images are encoded to obtain BxN×D feature vectors.</li>\n<li>The model is split into two branches: one for level classifier and one for right/left classifier.<ul>\n<li>BxNx1280-&gt;LSTM-&gt;reshape(BNx1280)-&gt;classification head</li></ul></li></ul></li>\n</ul>\n<h2><strong>2. Coordinates (XY) detection</strong></h2>\n<p>In step 1, I detected the slices to use, i.e., determined the Z position for analysis. In step 2, I made segmentation models to detect XY coordinates.</p>\n<ul>\n<li>input (C, H, W)<ul>\n<li>The C are neighboring images</li>\n<li>(3, 224, 224)</li></ul></li>\n<li>output (C, H, W)<ul>\n<li>(2, 224, 224) for SS (Axial T2)<ul>\n<li>Channels correspond to Left/Right</li></ul></li>\n<li>(5, 224, 224) for NFN (Sagittal T1) and SCS (Sagittal T2)<ul>\n<li>Channels correspond to L1/L2…L5/S1</li></ul></li></ul></li>\n<li>Model<ul>\n<li>Unet from segmentation_models_pytorch</li>\n<li>resnet18</li></ul></li>\n</ul>\n<h2><strong>3. Classification</strong></h2>\n<p>I could predict XYZ coordinates for levels and left/right from step 1 and 2. In this step, I made model to classify 3 classes (Normal/Moderate/Severe).</p>\n<ul>\n<li>Detection of XYZ coordinates<ul>\n<li>Determine candidates of Z positions from step 1<ul>\n<li>This Z position was not so accurate, I used these positions as start positions to search accurate positions from step 2 mask</li></ul></li>\n<li>Search best Z positions and determine XY<ul>\n<li>Calculate mask areas from neighboring slices and the slice with the largest area assigned as the best position</li>\n<li>Calculate centroid to determine XY</li></ul></li></ul></li>\n<li>Input<ul>\n<li>CenterCrop (3, 128, 110), Resize(160, 160) for SS</li>\n<li>CenterCrop (3, 90, 128), Resize(160, 160) for NFN and SCS</li></ul></li>\n<li>Backbone<ul>\n<li>efficientnet_b0</li>\n<li>efficentnetv2s</li>\n<li>convnextv2_tiny</li>\n<li>convnextv2_nano</li></ul></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3012914,
          "author_name": "icw1lee",
          "author_url": "",
          "post_date": "10/09/2024 13:51:27",
          "content": "<p>Congratulations on your 14th place. Do you have a specific notebook</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3012411": "### Models \n- Single-stage multi-condition (s12): 0.412 CV\n- Multi-stage multi-condition (s1+s2): 0.404 CV\n- Multi-stage single-condition (@tamotamo shared it [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539459#3012444)): 0.400 CV\n\n### Data\nCLAHE normalization of the input.  I projected input points of interest provided by orgs for specific sources and can locate all 25 points in each of the inputs. I used [publicly shared key point annotation](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/528653) for a two-stage setup. For external data, I used [Spider ](https://huggingface.co/datasets/cdoswald/SPIDER)(dataset for segmentation of specific vertebra and disks) in my very first experiments and single-stage model pre-training.\n\n### Single-stage model (s12)\n**Architecture**:  Attention-based aggregation (similar to SED used in Bird-call competition): I predict 3d voxel grid (at reduced res corresponding to 1/token size) of attention weights and predicted targets. Then I apply spatial softmax on attention and weight the targets accordingly summing them along the spatial dimensions. It results in scalar Bx25 predictions. I use [competition metric approximation](https://www.kaggle.com/code/junkoda/optimize-the-evaluation-metric) as the loss function. In addition, I consider aux loss on attention using Gaussians to approximate annotated key points.\nAs an alternative, I consider 2 cross-attention layer decoder applied to each slice independently. The decoder takes 25 learnable queries as the input and each predicts the target, xy position, and probability p that this slice is valid for a particular target. The final prediction is weighted (based on p) target and xy. The loss is CE on p, Competition metric approximation on the target, and MSE on xy. The idea here is that the decoder trying to predict the position of the specific key point is also aggregating the information providing the target class. This approach performs similarly to attention-based aggregation.\n\n**Backbone**: Since the model must be able to predict specific levels for visually indistinguishable vertebra it should have access to the global image content (i.e. ViT-based architecture is preferable), and in early experiments with Spider I saw that DINOv2 (with registers) can assign levels to vertebra reliably. Therefore, I used DINOv2 B backbone taking a sequence of input frames (N,3,H,W). For side views the backbone is augmented with **zero initialized LSTM adapters** to perform sequence mixing and is pre-trained on the Spider segmentation dataset. For axial views, I added a **LSTM mixing layer** after the backbone.\n\nThe image setup: 16x448x448 for side input and 48x332x332 for axial input. If the sequence is shorter, images are repeated.\n\n### 2-stage multi-condition (s1+s2)\ns1 is a simple single-slice segmentation model applied to the center slice and trained to predict 10 key points shared publicly. I used Sagittal T1 + Sagittal T2/STIR sources and then reprojected corresponding coordinates to find appropriate slices in Axial T2 and also cross-check the sources and drop points in the side view. I do not do s1 on axial view because multiple groups make the data messy and the task is quite harder than point detection in a lateral view. The backbone is DINOv2 again because I need to predict 10 distinguishable key points. With 10 instead of 5 key poins I can derive the pox orientation and the size. Input size 448x448.\n\nThen I apply s2 for each source. In the side view I crop a small stack of boxes around proper areas of interest. As a result, I produce 5 stacks of image crops 5x16x3x192x192 for side views. Axial cropping is done based on the plane corresponding to the projected key point and is not so aggressive laterally. So I produce sequences of 5x8x3x320x320 (making sure the group=view angle is the same for all selected images, and replicate images if the number of selected slices is insufficient). The model is simple ConvNeXtv2 nano + LSTM mixing layer + concat pooling over sequence + the head for side view, and in axial view I use DINOv2 instead. The competition metric approximation is used as the loss function.\n\n### Aggregation\nI use a simple weighting of each model independent for each condition (apply to logits). The matrix below shows the contribution of each of 7 sources (s2 Sagittal T2/STIR, s2 Sagittal T1, s2 Axial T2, s12  Sagittal T2/STIR, s2 Sagittal T1, s2 Axial, Multi-stage single-condition) to 5 targets (SCS, L NFN, R NFN, L SS, RSS). Some sources are particularly important for specific targets. \nw =    [[0.1228, 0.0025, 0.0031, 0.0612, 0.1021],\n        [0.0549, 0.2573, 0.1362, 0.0484, 0.1198],\n        [0.1712, 0.0790, 0.0269, 0.2366, 0.1519],\n        [0.1308, 0.0068, 0.0373, 0.1285, 0.0949],\n        [0.0161, 0.1436, 0.2533, 0.0101, 0.0072],\n        [0.0689, 0.0220, 0.0208, 0.0797, 0.1492],\n        [0.4353, 0.4888, 0.5225, 0.4354, 0.3749]]\nCV 0.379, LB 0.35/0.41",
    "3012444": "#  **Summray**\nMy pipeline consists of three steps: 1) Slice position (Z) detection, 2) Coordinates (XY) detection, and 3) classification. All steps are performed independently for the three types of conditions: spinal canal stenosis (SS), neural foraminal narrowing (NFN), and subarticular stenosis (SS). Additionally, NFN, SCS, and SS correspond to the image types Sagittal T1, Sagittal T2, and Axial T2, respectively.\n\n## **1. Slice position (Z) detection**\nThis model can detect the slices where disease exists. Slice positions to detect is from train_label_coordinates.csv\n- input (N, C, H, W)\n  - The C are neighboring images\n  - The N refers to the number of images obtained at regular intervals in the stack.\n  - Sagittal T1/T2 (10, 5, 224, 224)\n  - Axial T2 (20, 3, 224, 224)\n- Backbone\n  - efficientnet_b0\n  - convnextv2_tiny\n- model structure\n  - The images are encoded to obtain BxN×D feature vectors.\n  - The model is split into two branches: one for level classifier and one for right/left classifier.\n      - BxNx1280->LSTM->reshape(BNx1280)->classification head\n\n## **2. Coordinates (XY) detection**\nIn step 1, I detected the slices to use, i.e., determined the Z position for analysis. In step 2, I made segmentation models to detect XY coordinates.\n- input (C, H, W)\n  - The C are neighboring images\n  - (3, 224, 224)\n- output (C, H, W)\n  - (2, 224, 224) for SS (Axial T2)\n      - Channels correspond to Left/Right\n  - (5, 224, 224) for NFN (Sagittal T1) and SCS (Sagittal T2)\n      - Channels correspond to L1/L2...L5/S1\n- Model\n  - Unet from segmentation_models_pytorch\n  - resnet18\n\n## **3. Classification**\nI could predict XYZ coordinates for levels and left/right from step 1 and 2. In this step, I made model to classify 3 classes (Normal/Moderate/Severe).\n- Detection of XYZ coordinates\n    - Determine candidates of Z positions from step 1\n      - This Z position was not so accurate, I used these positions as start positions to search accurate positions from step 2 mask\n    - Search best Z positions and determine XY\n      - Calculate mask areas from neighboring slices and the slice with the largest area assigned as the best position\n        - Calculate centroid to determine XY\n- Input\n  - CenterCrop (3, 128, 110), Resize(160, 160) for SS\n  - CenterCrop (3, 90, 128), Resize(160, 160) for NFN and SCS\n- Backbone\n  - efficientnet_b0\n  - efficentnetv2s\n  - convnextv2_tiny\n  - convnextv2_nano",
    "3012914": "Congratulations on your 14th place. Do you have a specific notebook"
  },
  "source": "meta"
}