{
  "id": 539548,
  "title": "8th Place Solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/k-mataro-8th-place-solution",
  "author_name": "",
  "post_date": "2024-10-13T16:16:48Z",
  "votes": 32,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to RSNA and Kaggle team for hosting such interesting challenge. Thanks <a href=\"https://www.kaggle.com/stgkrtua\" target=\"_blank\">@stgkrtua</a> for teaming up with me. This was my first time participating in a medical competition, but I enjoyed it very much.</p>\n<h2>Overview</h2>\n<p>Our core architecture of our solution is shown in the following figure. It consists of 2D classifier and 1D classifier.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F6e38e1e8473ed960c3b594e14d382c38%2Frsna2024png.png?generation=1728480886830728&amp;alt=media\"></p>\n<h3>2D classifier:</h3>\n<ul>\n<li>Takes a 2-dimensional image as input.</li>\n<li>Predicts 25x3 severity labels (=5 conditions x 5 levels x 3 class).</li>\n</ul>\n<h3>1D classifier:</h3>\n<ul>\n<li>Stack the output of the 2D classifier.</li>\n<li>The final prediction is predicted by considering the 1st-stage predictions for every instance ID.</li>\n</ul>\n<p>One of the uniqueness of our model is the <strong>feature extraction</strong> of 1st stage.<br>\nInstead of cropping out the important parts of an image, we trained our model to extract the features corresponding to these regions. So we trained the detector as a sub task and used the heatmap as the weight of the feature extraction like an attention.</p>\n<p>This model provides several advantages, </p>\n<ul>\n<li>it enables to consider the overall <strong>context of the image</strong>.</li>\n<li>it gives robust outputs that are <strong>less sensitive to ambiguities in detection</strong>.</li>\n</ul>\n<h2>Other findings</h2>\n<h3>SSL</h3>\n<p>Self-supervised learning slightly improved the public/private score. The model is trained to learn the multi-view similality by using the xyz coordinate. Unfortunately, CV score didn't change. I guess the labels of test dataset are cleaner than train dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F95c172fb3a91bf3cc93871ddbee0ca71%2Fssl.png?generation=1728481052071515&amp;alt=media\"><br>\nRed area in the figure shows the high similatiry area between two images.</p>\n<h3>Relabeling</h3>\n<p>As you know, there were many inconsistencies in the labels. <br>\nI created a super cool annotation tool and tried to correct the labels myself, but I couldn't do it well due to the lack of domain knowledge. If anyone has made corrections, I would appreciate it if they could share the results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fee61db48bd2b8de9c120e9985a435f21%2Fannotation.png?generation=1728481276648808&amp;alt=media\"><br>\n<strong>Creating the annotation tool was one of the most enjoyable parts</strong> of the competition, even though it didn't end up being very useful.</p>\n<h2>Code</h2>\n<p>The inference code is available at<br>\n<a href=\"https://www.kaggle.com/kmat2019/rsna8th-inference\" target=\"_blank\">https://www.kaggle.com/kmat2019/rsna8th-inference</a></p>\n<p>Due to a minor bug, the score is slightly better than the final submission.</p>",
  "messages": [
    {
      "id": "3012909",
      "postDate": "10/09/2024 13:49:02",
      "content": "<p>Thanks to RSNA and Kaggle team for hosting such interesting challenge. Thanks <a href=\"https://www.kaggle.com/stgkrtua\" target=\"_blank\">@stgkrtua</a> for teaming up with me. This was my first time participating in a medical competition, but I enjoyed it very much.</p>\n<h2>Overview</h2>\n<p>Our core architecture of our solution is shown in the following figure. It consists of 2D classifier and 1D classifier.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F6e38e1e8473ed960c3b594e14d382c38%2Frsna2024png.png?generation=1728480886830728&amp;alt=media\"></p>\n<h3>2D classifier:</h3>\n<ul>\n<li>Takes a 2-dimensional image as input.</li>\n<li>Predicts 25x3 severity labels (=5 conditions x 5 levels x 3 class).</li>\n</ul>\n<h3>1D classifier:</h3>\n<ul>\n<li>Stack the output of the 2D classifier.</li>\n<li>The final prediction is predicted by considering the 1st-stage predictions for every instance ID.</li>\n</ul>\n<p>One of the uniqueness of our model is the <strong>feature extraction</strong> of 1st stage.<br>\nInstead of cropping out the important parts of an image, we trained our model to extract the features corresponding to these regions. So we trained the detector as a sub task and used the heatmap as the weight of the feature extraction like an attention.</p>\n<p>This model provides several advantages, </p>\n<ul>\n<li>it enables to consider the overall <strong>context of the image</strong>.</li>\n<li>it gives robust outputs that are <strong>less sensitive to ambiguities in detection</strong>.</li>\n</ul>\n<h2>Other findings</h2>\n<h3>SSL</h3>\n<p>Self-supervised learning slightly improved the public/private score. The model is trained to learn the multi-view similality by using the xyz coordinate. Unfortunately, CV score didn't change. I guess the labels of test dataset are cleaner than train dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F95c172fb3a91bf3cc93871ddbee0ca71%2Fssl.png?generation=1728481052071515&amp;alt=media\"><br>\nRed area in the figure shows the high similatiry area between two images.</p>\n<h3>Relabeling</h3>\n<p>As you know, there were many inconsistencies in the labels. <br>\nI created a super cool annotation tool and tried to correct the labels myself, but I couldn't do it well due to the lack of domain knowledge. If anyone has made corrections, I would appreciate it if they could share the results.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fee61db48bd2b8de9c120e9985a435f21%2Fannotation.png?generation=1728481276648808&amp;alt=media\"><br>\n<strong>Creating the annotation tool was one of the most enjoyable parts</strong> of the competition, even though it didn't end up being very useful.</p>\n<h2>Code</h2>\n<p>The inference code is available at<br>\n<a href=\"https://www.kaggle.com/kmat2019/rsna8th-inference\" target=\"_blank\">https://www.kaggle.com/kmat2019/rsna8th-inference</a></p>\n<p>Due to a minor bug, the score is slightly better than the final submission.</p>",
      "rawMarkdown": "Thanks to RSNA and Kaggle team for hosting such interesting challenge. Thanks @stgkrtua for teaming up with me. This was my first time participating in a medical competition, but I enjoyed it very much.\n\n## Overview\nOur core architecture of our solution is shown in the following figure. It consists of 2D classifier and 1D classifier.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F6e38e1e8473ed960c3b594e14d382c38%2Frsna2024png.png?generation=1728480886830728&alt=media\" width=\"720\">\n\n### 2D classifier:\n- Takes a 2-dimensional image as input.\n- Predicts 25x3 severity labels (=5 conditions x 5 levels x 3 class).\n\n### 1D classifier:\n- Stack the output of the 2D classifier.\n- The final prediction is predicted by considering the 1st-stage predictions for every instance ID.\n\nOne of the uniqueness of our model is the **feature extraction** of 1st stage.\nInstead of cropping out the important parts of an image, we trained our model to extract the features corresponding to these regions. So we trained the detector as a sub task and used the heatmap as the weight of the feature extraction like an attention.\n\nThis model provides several advantages, \n- it enables to consider the overall **context of the image**.\n- it gives robust outputs that are **less sensitive to ambiguities in detection**.\n\n## Other findings\n### SSL\nSelf-supervised learning slightly improved the public/private score. The model is trained to learn the multi-view similality by using the xyz coordinate. Unfortunately, CV score didn't change. I guess the labels of test dataset are cleaner than train dataset.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F95c172fb3a91bf3cc93871ddbee0ca71%2Fssl.png?generation=1728481052071515&alt=media\" width=\"640\">\nRed area in the figure shows the high similatiry area between two images.\n\n### Relabeling\nAs you know, there were many inconsistencies in the labels. \nI created a super cool annotation tool and tried to correct the labels myself, but I couldn't do it well due to the lack of domain knowledge. If anyone has made corrections, I would appreciate it if they could share the results.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fee61db48bd2b8de9c120e9985a435f21%2Fannotation.png?generation=1728481276648808&alt=media\" width=\"640\">\n**Creating the annotation tool was one of the most enjoyable parts** of the competition, even though it didn't end up being very useful.\n\n## Code\nThe inference code is available at\nhttps://www.kaggle.com/kmat2019/rsna8th-inference\n\nDue to a minor bug, the score is slightly better than the final submission.",
      "votes": null
    },
    {
      "id": "3013509",
      "postDate": "10/10/2024 08:18:45",
      "content": "<p>Congratulations. The approach is pretty interesting. May I ask some questions:</p>\n<ol>\n<li>Which images did you choose to train the 1st stage and the 2nd stage? Images with coordinate annotations or all the images?</li>\n<li>In the 1st stage, not all images have all 5 level annotations, how did you determine the labels?<br>\nThank you</li>\n</ol>",
      "rawMarkdown": "Congratulations. The approach is pretty interesting. May I ask some questions:\n1. Which images did you choose to train the 1st stage and the 2nd stage? Images with coordinate annotations or all the images?\n2. In the 1st stage, not all images have all 5 level annotations, how did you determine the labels?\nThank you",
      "votes": null
    },
    {
      "id": "3013821",
      "postDate": "10/10/2024 15:16:22",
      "content": "<p>Congratulations! Any place where I can find the code?</p>",
      "rawMarkdown": "Congratulations! Any place where I can find the code?",
      "votes": null
    },
    {
      "id": "3014571",
      "postDate": "10/11/2024 11:17:53",
      "content": "<p>Hi! I'm very impressed with SSL approach, can you please elaborate how do you define similarity labels between slices from different planes? </p>\n<p>I thought about making a transformer model that aggregates embeddings from different planes, however I wasn't able to come up with distance function</p>",
      "rawMarkdown": "Hi! I'm very impressed with SSL approach, can you please elaborate how do you define similarity labels between slices from different planes? \n\nI thought about making a transformer model that aggregates embeddings from different planes, however I wasn't able to come up with distance function",
      "votes": null
    },
    {
      "id": "3014610",
      "postDate": "10/11/2024 12:15:46",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "3014816",
      "postDate": "10/11/2024 16:03:04",
      "content": "<ol>\n<li>All images were used for training, but the ratio was adjusted to ensure that both labeled and unlabeled images were sampled nearly equally.</li>\n<li>We set the loss weight to less than 10% for non-annotated data.</li>\n</ol>",
      "rawMarkdown": "1. All images were used for training, but the ratio was adjusted to ensure that both labeled and unlabeled images were sampled nearly equally.\n2. We set the loss weight to less than 10% for non-annotated data.",
      "votes": null
    },
    {
      "id": "3014831",
      "postDate": "10/11/2024 16:15:48",
      "content": "<p>I simply computed the xyz map from dicom data for each images, and trained a model to learn a similarity metric between image pairs based on the distance of their corresponding xyz maps.</p>\n<pre><code> = cos_sim(embedding_0.reshape(batch, num_pixel, , num_ch), embedding_1.reshape(batch, , num_pixel, num_ch))\n = dist(xyz_map_0.reshape(batch, num_pixel, , ), xyz_map_1.reshape(batch, , num_pixel, )) &lt; threshold\n</code></pre>",
      "rawMarkdown": "I simply computed the xyz map from dicom data for each images, and trained a model to learn a similarity metric between image pairs based on the distance of their corresponding xyz maps.\n```\nsimilarity_matrix = cos_sim(embedding_0.reshape(batch, num_pixel, 1, num_ch), embedding_1.reshape(batch, 1, num_pixel, num_ch))\nlabel_matrix = dist(xyz_map_0.reshape(batch, num_pixel, 1, 3), xyz_map_1.reshape(batch, 1, num_pixel, 3)) < threshold\n```",
      "votes": null
    },
    {
      "id": "3014836",
      "postDate": "10/11/2024 16:18:32",
      "content": "<p>It will be released within a few days.</p>",
      "rawMarkdown": "It will be released within a few days.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3013509,
      "author_name": "namgalielei",
      "author_url": "",
      "post_date": "10/10/2024 08:18:45",
      "content": "<p>Congratulations. The approach is pretty interesting. May I ask some questions:</p>\n<ol>\n<li>Which images did you choose to train the 1st stage and the 2nd stage? Images with coordinate annotations or all the images?</li>\n<li>In the 1st stage, not all images have all 5 level annotations, how did you determine the labels?<br>\nThank you</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 3014816,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "10/11/2024 16:03:04",
          "content": "<ol>\n<li>All images were used for training, but the ratio was adjusted to ensure that both labeled and unlabeled images were sampled nearly equally.</li>\n<li>We set the loss weight to less than 10% for non-annotated data.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3013821,
      "author_name": "karelbecerra",
      "author_url": "",
      "post_date": "10/10/2024 15:16:22",
      "content": "<p>Congratulations! Any place where I can find the code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3014836,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "10/11/2024 16:18:32",
          "content": "<p>It will be released within a few days.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3014571,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "10/11/2024 11:17:53",
      "content": "<p>Hi! I'm very impressed with SSL approach, can you please elaborate how do you define similarity labels between slices from different planes? </p>\n<p>I thought about making a transformer model that aggregates embeddings from different planes, however I wasn't able to come up with distance function</p>",
      "votes": null,
      "replies": [
        {
          "id": 3014831,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "10/11/2024 16:15:48",
          "content": "<p>I simply computed the xyz map from dicom data for each images, and trained a model to learn a similarity metric between image pairs based on the distance of their corresponding xyz maps.</p>\n<pre><code> = cos_sim(embedding_0.reshape(batch, num_pixel, , num_ch), embedding_1.reshape(batch, , num_pixel, num_ch))\n = dist(xyz_map_0.reshape(batch, num_pixel, , ), xyz_map_1.reshape(batch, , num_pixel, )) &lt; threshold\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3014610,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "10/11/2024 12:15:46",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3012909": "Thanks to RSNA and Kaggle team for hosting such interesting challenge. Thanks @stgkrtua for teaming up with me. This was my first time participating in a medical competition, but I enjoyed it very much.\n\n## Overview\nOur core architecture of our solution is shown in the following figure. It consists of 2D classifier and 1D classifier.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F6e38e1e8473ed960c3b594e14d382c38%2Frsna2024png.png?generation=1728480886830728&alt=media\" width=\"720\">\n\n### 2D classifier:\n- Takes a 2-dimensional image as input.\n- Predicts 25x3 severity labels (=5 conditions x 5 levels x 3 class).\n\n### 1D classifier:\n- Stack the output of the 2D classifier.\n- The final prediction is predicted by considering the 1st-stage predictions for every instance ID.\n\nOne of the uniqueness of our model is the **feature extraction** of 1st stage.\nInstead of cropping out the important parts of an image, we trained our model to extract the features corresponding to these regions. So we trained the detector as a sub task and used the heatmap as the weight of the feature extraction like an attention.\n\nThis model provides several advantages, \n- it enables to consider the overall **context of the image**.\n- it gives robust outputs that are **less sensitive to ambiguities in detection**.\n\n## Other findings\n### SSL\nSelf-supervised learning slightly improved the public/private score. The model is trained to learn the multi-view similality by using the xyz coordinate. Unfortunately, CV score didn't change. I guess the labels of test dataset are cleaner than train dataset.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2F95c172fb3a91bf3cc93871ddbee0ca71%2Fssl.png?generation=1728481052071515&alt=media\" width=\"640\">\nRed area in the figure shows the high similatiry area between two images.\n\n### Relabeling\nAs you know, there were many inconsistencies in the labels. \nI created a super cool annotation tool and tried to correct the labels myself, but I couldn't do it well due to the lack of domain knowledge. If anyone has made corrections, I would appreciate it if they could share the results.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2938236%2Fee61db48bd2b8de9c120e9985a435f21%2Fannotation.png?generation=1728481276648808&alt=media\" width=\"640\">\n**Creating the annotation tool was one of the most enjoyable parts** of the competition, even though it didn't end up being very useful.\n\n## Code\nThe inference code is available at\nhttps://www.kaggle.com/kmat2019/rsna8th-inference\n\nDue to a minor bug, the score is slightly better than the final submission.",
    "3013509": "Congratulations. The approach is pretty interesting. May I ask some questions:\n1. Which images did you choose to train the 1st stage and the 2nd stage? Images with coordinate annotations or all the images?\n2. In the 1st stage, not all images have all 5 level annotations, how did you determine the labels?\nThank you",
    "3013821": "Congratulations! Any place where I can find the code?",
    "3014571": "Hi! I'm very impressed with SSL approach, can you please elaborate how do you define similarity labels between slices from different planes? \n\nI thought about making a transformer model that aggregates embeddings from different planes, however I wasn't able to come up with distance function",
    "3014610": "Congratulations!",
    "3014816": "1. All images were used for training, but the ratio was adjusted to ensure that both labeled and unlabeled images were sampled nearly equally.\n2. We set the loss weight to less than 10% for non-annotated data.",
    "3014831": "I simply computed the xyz map from dicom data for each images, and trained a model to learn a similarity metric between image pairs based on the distance of their corresponding xyz maps.\n```\nsimilarity_matrix = cos_sim(embedding_0.reshape(batch, num_pixel, 1, num_ch), embedding_1.reshape(batch, 1, num_pixel, num_ch))\nlabel_matrix = dist(xyz_map_0.reshape(batch, num_pixel, 1, 3), xyz_map_1.reshape(batch, 1, num_pixel, 3)) < threshold\n```",
    "3014836": "It will be released within a few days."
  },
  "source": "meta"
}