{
  "id": 539981,
  "title": "15th Place Solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/derpmaxxing-15th-place-solution",
  "author_name": "",
  "post_date": "2024-10-11T19:12:45.330898600Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Congrats to all winners and those seeing themselves better data scientists than who they were before completing this competition! Despite being 2 ranks short of gold, we find this competition really interesting as there is no trivial way to approach this competition, which makes it much more interesting.</p>\n<p>Most importantly I'd like to thank <a href=\"https://www.kaggle.com/viktorcikojevic\" target=\"_blank\">@viktorcikojevic</a> for teaming with me on this (any loads of past) competition. Without him, there's no chance I'd have come this far.</p>\n<h2>Overview</h2>\n<p>On a higher level, our pipeline is depicted in the following diagram:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F458043a2fd2439bce786b15e1f27a367%2Fkaggle-overview-image.png?generation=1728672784413838&amp;alt=media\" alt=\"\"></p>\n<p>It consists of three stages:</p>\n<ol>\n<li>Keypoint detection stage</li>\n<li>Crop proposal stage</li>\n<li>Crop classification stage</li>\n</ol>\n<h3>a brief word on what motivated this design choice:</h3>\n<ul>\n<li>This pipeline should allow all models to train at a <em>per-image</em> level, rather than <em>per-patient</em> level. There is only about ~2000 patients, so we thought this would be a better way to utilize all of the data and prevent overfitting. </li>\n<li>A lot of information are inferrable between each model's output, which allowed us to include a helpful bias to the model. For example, we know T2 runs right at the center of the person, so all T1 keypoints on their left are left T1 keypoints, same for the right hand side. This means we can take some shortcuts on what the model must learn.</li>\n</ul>\n<h5>When it comes to implementation, it means we did the following:</h5>\n<ul>\n<li>At the keypoint detection stage, our T1 keypoint model will only predict 5 classes: L1/L2, L2/L3, L3/L4, L4/L5, L5/S1 <strong>and not 10</strong>. No sides are predicted at this stage.</li>\n<li>At the keypoint detection stage, our axial keypoint model will only predict 2 classes: left and right keypoint, <strong>not 10</strong>. No levels are predicted</li>\n<li>the crop classifier's job is to take a crop in, and output 3 logits - mild / moderate / severe - it doesn't predict the condition</li>\n<li>the Crop proposal stage does 3 main jobs<ul>\n<li>fill in the sides of each T1 keypoint (because the keypoint model only knows the levels)</li>\n<li>fill in the level of each Axial keypoint (because the keypoint model only knows the sides)</li>\n<li>aggregate per-image predictions into the final 25 keypoints for each patient.</li></ul></li>\n</ul>\n<p>In the sections below we will describe each stage in more detail.</p>\n<h2>Keypoint Detection Stage</h2>\n<p>Here we developed a segmentation model that outputs a map of keypoints for each image. For each of the conditions, we train a SMP (Segmentation Model Pytorch) model with <code>timm</code> backbone. The model takes the 3 consecutive <code>instance_number</code> channels as inputs and outputs a multi-channel heatmaps.</p>\n<h3>T2 Keypoint models</h3>\n<p>For T2 models, we just use the dataset shared by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> here <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset</a>.</p>\n<p>With us joining when there's on month remaining, we did not find ways to exploit the left keypoints, so we just train with the right keypoints only.</p>\n<h3>T1 &amp; Axial Keypoint models</h3>\n<p>For T1 and Axial keypoint models, we just trust the labelled keypoints as-is and trained our model using those. Mostly because we're lazy (to remove all the labelling noise), and also we're a little short on time</p>\n<p>One reason we think we can afford some labelling noise is the fact that we took shortcuts to minimise what labels the models are trained on, i.e. the T1 model doesn't care about the sides of the keypoint, so we're robust to side flips, and the Axial model doesn't care about the level, so we're robust to any noise wrt. level labels.</p>\n<h2>Crop Proposal Heuristic</h2>\n<p>We use heuristics to go from output keypoints to the final 25 keypoints for the patient. It needs to accomplish 3 things:</p>\n<ul>\n<li>fill in the sides of each T1 keypoint (because the T1 keypoint model only knows the levels)</li>\n<li>fill in the level of each Axial keypoint (because Axial keypoint model only knows the sides)</li>\n<li>aggregate per-image predictions into the final 25 keypoints for each patient.</li>\n</ul>\n<p>The overview of the Crop proposal heuristics is shown here:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2Ff88af07f0fca5f22c6367086e1c8edb1%2Fkaggle-crop-image.png?generation=1728673662221061&amp;alt=media\" alt=\"\"></p>\n<h3>Step 1: Argmax T2 Instance Number</h3>\n<p>This step is rather simple: for each of the 5 T2 level, we find the instance number with the highest confidence from the keypoint model. After this step ends, we have 5 T2 keypoints for the patient.</p>\n<h3>Step 1: Infer T1 Sides</h3>\n<p>Because we know the T2 keypoints from the previous step, we can now compute the XYZ position in the world coordinate using the dicom's metadata.<br>\nDoing so will tell us the XYZ coordinates of the spine, where the X axis points from the right hand of the patient to the left. So now inferring the sides<br>\nis rather trivial:</p>\n<blockquote>\n  <p>for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint.</p>\n</blockquote>\n<p>This operation has an accuracy of 97% on determining the side of each T1 keypoint, the remaining 3% is either the T1 keypoint not being detected at all, or labelling noise.</p>\n<p>There are some edge case that needs handling, which turns this step's logic into</p>\n<blockquote>\n  <p>for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint</p>\n  <p>but if T2 on level L1/L2 is not detected, then use T2 on level L2/L3 instead, if that's missing too then use t2 on L3/L4, and so on</p>\n</blockquote>\n<h3>Step 3: Argmax T1 instance number</h3>\n<p>Pretty much step 1 applied to T1 keypoints - for each side and level, we find which <code>instance_number</code> has the highest confidence. <br>\nThis step gives us 10 final T1 keypoints for the patient.</p>\n<h3>Step 4: Guess Missing T1 Keypoint</h3>\n<p>Under any incident of T1 keypoints missing, if its \"twin\" exists, then just mirror it over and call it a day</p>\n<blockquote>\n  <p>if T1 left L1/L2 went missing, mirror T1 right L1/L2 around T2 L1/L2, and blindly claim that's the T1 left L1/L2 location</p>\n</blockquote>\n<p>This step gives minor improvements to the cv (around 0.001 to 0.002)</p>\n<h3>Step 5: Infer Axial Levels</h3>\n<p>Because we know T2 keypoints, each have their levels. Then we can use that to guess what level each axial keypoint should have.</p>\n<blockquote>\n  <p>for a given Axial keypoint, look for its closest T2 keypoint using its XYZ coordinates, and take that T2 point's level as the Axial point's level</p>\n</blockquote>\n<p>This step is vulnerable to missing T2 keypoints (e.g. if T2's L2/L3 is missing, no Axial keypoint can correctly have L2/L3 as level). To combat this issue, <br>\nwe just linearly interpolate all missing T2 keypoints before inferring axial levels.</p>\n<p>This operation also has an accuracy of 97% on levels of each Axial keypoint, the remaining 3% is mostly driven from the T2 keypoints being wrong, which affects the level lookups, or the axial keypoints being missing altogether. </p>\n<h3>Step 6: Argmax Axial Instance Number</h3>\n<p>Pretty much step 1, but now applied to Axial keypoints - for each side and level, we find which instance number has the highest confidence. <br>\nThis step gives us 10 final Axial keypoints for the patient.</p>\n<h3>Step 7: Guess Missing Axial Keypoints</h3>\n<p>This step sadly doesn't exist for us. I tried many approach in imputing these missing axial keypoints, but none of them really work out that well. <br>\nOne reason is that unlike T1 keypoints where normally only one point goes missing and mirroring fixes the issue, the Axial keypoints of the same level normally go missing together.</p>\n<p>All experiments here leads to worse cv score, so we just accepted the fact that we just can't do this and miss keypoints..</p>\n<h2>Crop Classification Stage</h2>\n<p>At the end of our pipeline is a classification model is a 9-class model that's construced as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F58d44332885eac17491d601cbde2043c%2Fcrop-cls.png?generation=1728673803309682&amp;alt=media\" alt=\"\"></p>\n<p>It's a <code>timm</code> backbone and 3 heads, one for each series type. Each head outputs 3 classes - the severity of the given image <em>if</em> the image is from that series type. For example, if an image is an axial image, the forward pass sends it though red path of timm backbone + axial head, and only the last 3 entries (the subarticular severities) are filled, the T1 and T2 logits are filled with -100 so that the probabilities are 0 after softmax.</p>\n<hr>\n<p>Thank you very much for reading this far. We hope this has been informational, or at least entertaining :) </p>",
  "messages": [
    {
      "id": "3014987",
      "postDate": "10/11/2024 19:12:45",
      "content": "<p>Congrats to all winners and those seeing themselves better data scientists than who they were before completing this competition! Despite being 2 ranks short of gold, we find this competition really interesting as there is no trivial way to approach this competition, which makes it much more interesting.</p>\n<p>Most importantly I'd like to thank <a href=\"https://www.kaggle.com/viktorcikojevic\" target=\"_blank\">@viktorcikojevic</a> for teaming with me on this (any loads of past) competition. Without him, there's no chance I'd have come this far.</p>\n<h2>Overview</h2>\n<p>On a higher level, our pipeline is depicted in the following diagram:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F458043a2fd2439bce786b15e1f27a367%2Fkaggle-overview-image.png?generation=1728672784413838&amp;alt=media\" alt=\"\"></p>\n<p>It consists of three stages:</p>\n<ol>\n<li>Keypoint detection stage</li>\n<li>Crop proposal stage</li>\n<li>Crop classification stage</li>\n</ol>\n<h3>a brief word on what motivated this design choice:</h3>\n<ul>\n<li>This pipeline should allow all models to train at a <em>per-image</em> level, rather than <em>per-patient</em> level. There is only about ~2000 patients, so we thought this would be a better way to utilize all of the data and prevent overfitting. </li>\n<li>A lot of information are inferrable between each model's output, which allowed us to include a helpful bias to the model. For example, we know T2 runs right at the center of the person, so all T1 keypoints on their left are left T1 keypoints, same for the right hand side. This means we can take some shortcuts on what the model must learn.</li>\n</ul>\n<h5>When it comes to implementation, it means we did the following:</h5>\n<ul>\n<li>At the keypoint detection stage, our T1 keypoint model will only predict 5 classes: L1/L2, L2/L3, L3/L4, L4/L5, L5/S1 <strong>and not 10</strong>. No sides are predicted at this stage.</li>\n<li>At the keypoint detection stage, our axial keypoint model will only predict 2 classes: left and right keypoint, <strong>not 10</strong>. No levels are predicted</li>\n<li>the crop classifier's job is to take a crop in, and output 3 logits - mild / moderate / severe - it doesn't predict the condition</li>\n<li>the Crop proposal stage does 3 main jobs<ul>\n<li>fill in the sides of each T1 keypoint (because the keypoint model only knows the levels)</li>\n<li>fill in the level of each Axial keypoint (because the keypoint model only knows the sides)</li>\n<li>aggregate per-image predictions into the final 25 keypoints for each patient.</li></ul></li>\n</ul>\n<p>In the sections below we will describe each stage in more detail.</p>\n<h2>Keypoint Detection Stage</h2>\n<p>Here we developed a segmentation model that outputs a map of keypoints for each image. For each of the conditions, we train a SMP (Segmentation Model Pytorch) model with <code>timm</code> backbone. The model takes the 3 consecutive <code>instance_number</code> channels as inputs and outputs a multi-channel heatmaps.</p>\n<h3>T2 Keypoint models</h3>\n<p>For T2 models, we just use the dataset shared by <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> here <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset</a>.</p>\n<p>With us joining when there's on month remaining, we did not find ways to exploit the left keypoints, so we just train with the right keypoints only.</p>\n<h3>T1 &amp; Axial Keypoint models</h3>\n<p>For T1 and Axial keypoint models, we just trust the labelled keypoints as-is and trained our model using those. Mostly because we're lazy (to remove all the labelling noise), and also we're a little short on time</p>\n<p>One reason we think we can afford some labelling noise is the fact that we took shortcuts to minimise what labels the models are trained on, i.e. the T1 model doesn't care about the sides of the keypoint, so we're robust to side flips, and the Axial model doesn't care about the level, so we're robust to any noise wrt. level labels.</p>\n<h2>Crop Proposal Heuristic</h2>\n<p>We use heuristics to go from output keypoints to the final 25 keypoints for the patient. It needs to accomplish 3 things:</p>\n<ul>\n<li>fill in the sides of each T1 keypoint (because the T1 keypoint model only knows the levels)</li>\n<li>fill in the level of each Axial keypoint (because Axial keypoint model only knows the sides)</li>\n<li>aggregate per-image predictions into the final 25 keypoints for each patient.</li>\n</ul>\n<p>The overview of the Crop proposal heuristics is shown here:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2Ff88af07f0fca5f22c6367086e1c8edb1%2Fkaggle-crop-image.png?generation=1728673662221061&amp;alt=media\" alt=\"\"></p>\n<h3>Step 1: Argmax T2 Instance Number</h3>\n<p>This step is rather simple: for each of the 5 T2 level, we find the instance number with the highest confidence from the keypoint model. After this step ends, we have 5 T2 keypoints for the patient.</p>\n<h3>Step 1: Infer T1 Sides</h3>\n<p>Because we know the T2 keypoints from the previous step, we can now compute the XYZ position in the world coordinate using the dicom's metadata.<br>\nDoing so will tell us the XYZ coordinates of the spine, where the X axis points from the right hand of the patient to the left. So now inferring the sides<br>\nis rather trivial:</p>\n<blockquote>\n  <p>for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint.</p>\n</blockquote>\n<p>This operation has an accuracy of 97% on determining the side of each T1 keypoint, the remaining 3% is either the T1 keypoint not being detected at all, or labelling noise.</p>\n<p>There are some edge case that needs handling, which turns this step's logic into</p>\n<blockquote>\n  <p>for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint</p>\n  <p>but if T2 on level L1/L2 is not detected, then use T2 on level L2/L3 instead, if that's missing too then use t2 on L3/L4, and so on</p>\n</blockquote>\n<h3>Step 3: Argmax T1 instance number</h3>\n<p>Pretty much step 1 applied to T1 keypoints - for each side and level, we find which <code>instance_number</code> has the highest confidence. <br>\nThis step gives us 10 final T1 keypoints for the patient.</p>\n<h3>Step 4: Guess Missing T1 Keypoint</h3>\n<p>Under any incident of T1 keypoints missing, if its \"twin\" exists, then just mirror it over and call it a day</p>\n<blockquote>\n  <p>if T1 left L1/L2 went missing, mirror T1 right L1/L2 around T2 L1/L2, and blindly claim that's the T1 left L1/L2 location</p>\n</blockquote>\n<p>This step gives minor improvements to the cv (around 0.001 to 0.002)</p>\n<h3>Step 5: Infer Axial Levels</h3>\n<p>Because we know T2 keypoints, each have their levels. Then we can use that to guess what level each axial keypoint should have.</p>\n<blockquote>\n  <p>for a given Axial keypoint, look for its closest T2 keypoint using its XYZ coordinates, and take that T2 point's level as the Axial point's level</p>\n</blockquote>\n<p>This step is vulnerable to missing T2 keypoints (e.g. if T2's L2/L3 is missing, no Axial keypoint can correctly have L2/L3 as level). To combat this issue, <br>\nwe just linearly interpolate all missing T2 keypoints before inferring axial levels.</p>\n<p>This operation also has an accuracy of 97% on levels of each Axial keypoint, the remaining 3% is mostly driven from the T2 keypoints being wrong, which affects the level lookups, or the axial keypoints being missing altogether. </p>\n<h3>Step 6: Argmax Axial Instance Number</h3>\n<p>Pretty much step 1, but now applied to Axial keypoints - for each side and level, we find which instance number has the highest confidence. <br>\nThis step gives us 10 final Axial keypoints for the patient.</p>\n<h3>Step 7: Guess Missing Axial Keypoints</h3>\n<p>This step sadly doesn't exist for us. I tried many approach in imputing these missing axial keypoints, but none of them really work out that well. <br>\nOne reason is that unlike T1 keypoints where normally only one point goes missing and mirroring fixes the issue, the Axial keypoints of the same level normally go missing together.</p>\n<p>All experiments here leads to worse cv score, so we just accepted the fact that we just can't do this and miss keypoints..</p>\n<h2>Crop Classification Stage</h2>\n<p>At the end of our pipeline is a classification model is a 9-class model that's construced as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F58d44332885eac17491d601cbde2043c%2Fcrop-cls.png?generation=1728673803309682&amp;alt=media\" alt=\"\"></p>\n<p>It's a <code>timm</code> backbone and 3 heads, one for each series type. Each head outputs 3 classes - the severity of the given image <em>if</em> the image is from that series type. For example, if an image is an axial image, the forward pass sends it though red path of timm backbone + axial head, and only the last 3 entries (the subarticular severities) are filled, the T1 and T2 logits are filled with -100 so that the probabilities are 0 after softmax.</p>\n<hr>\n<p>Thank you very much for reading this far. We hope this has been informational, or at least entertaining :) </p>",
      "rawMarkdown": "Congrats to all winners and those seeing themselves better data scientists than who they were before completing this competition! Despite being 2 ranks short of gold, we find this competition really interesting as there is no trivial way to approach this competition, which makes it much more interesting.\n\nMost importantly I'd like to thank @viktorcikojevic for teaming with me on this (any loads of past) competition. Without him, there's no chance I'd have come this far.\n\n## Overview\n\nOn a higher level, our pipeline is depicted in the following diagram:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F458043a2fd2439bce786b15e1f27a367%2Fkaggle-overview-image.png?generation=1728672784413838&alt=media)\n\nIt consists of three stages:\n1. Keypoint detection stage\n1. Crop proposal stage\n1. Crop classification stage\n\n### a brief word on what motivated this design choice:\n\n- This pipeline should allow all models to train at a *per-image* level, rather than *per-patient* level. There is only about ~2000 patients, so we thought this would be a better way to utilize all of the data and prevent overfitting. \n- A lot of information are inferrable between each model's output, which allowed us to include a helpful bias to the model. For example, we know T2 runs right at the center of the person, so all T1 keypoints on their left are left T1 keypoints, same for the right hand side. This means we can take some shortcuts on what the model must learn.\n\n##### When it comes to implementation, it means we did the following:\n- At the keypoint detection stage, our T1 keypoint model will only predict 5 classes: L1/L2, L2/L3, L3/L4, L4/L5, L5/S1 **and not 10**. No sides are predicted at this stage.\n- At the keypoint detection stage, our axial keypoint model will only predict 2 classes: left and right keypoint, **not 10**. No levels are predicted\n- the crop classifier's job is to take a crop in, and output 3 logits - mild / moderate / severe - it doesn't predict the condition\n- the Crop proposal stage does 3 main jobs\n  - fill in the sides of each T1 keypoint (because the keypoint model only knows the levels)\n  - fill in the level of each Axial keypoint (because the keypoint model only knows the sides)\n  - aggregate per-image predictions into the final 25 keypoints for each patient.\n\n\nIn the sections below we will describe each stage in more detail.\n\n\n## Keypoint Detection Stage\n\nHere we developed a segmentation model that outputs a map of keypoints for each image. For each of the conditions, we train a SMP (Segmentation Model Pytorch) model with `timm` backbone. The model takes the 3 consecutive `instance_number` channels as inputs and outputs a multi-channel heatmaps.\n\n### T2 Keypoint models\nFor T2 models, we just use the dataset shared by @brendanartley here https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset.\n\nWith us joining when there's on month remaining, we did not find ways to exploit the left keypoints, so we just train with the right keypoints only.\n\n### T1 & Axial Keypoint models\nFor T1 and Axial keypoint models, we just trust the labelled keypoints as-is and trained our model using those. Mostly because we're lazy (to remove all the labelling noise), and also we're a little short on time\n\nOne reason we think we can afford some labelling noise is the fact that we took shortcuts to minimise what labels the models are trained on, i.e. the T1 model doesn't care about the sides of the keypoint, so we're robust to side flips, and the Axial model doesn't care about the level, so we're robust to any noise wrt. level labels.\n\n## Crop Proposal Heuristic\n\nWe use heuristics to go from output keypoints to the final 25 keypoints for the patient. It needs to accomplish 3 things:\n\n- fill in the sides of each T1 keypoint (because the T1 keypoint model only knows the levels)\n- fill in the level of each Axial keypoint (because Axial keypoint model only knows the sides)\n- aggregate per-image predictions into the final 25 keypoints for each patient.\n\nThe overview of the Crop proposal heuristics is shown here:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2Ff88af07f0fca5f22c6367086e1c8edb1%2Fkaggle-crop-image.png?generation=1728673662221061&alt=media)\n\n### Step 1: Argmax T2 Instance Number\n\nThis step is rather simple: for each of the 5 T2 level, we find the instance number with the highest confidence from the keypoint model. After this step ends, we have 5 T2 keypoints for the patient.\n\n### Step 1: Infer T1 Sides\n\nBecause we know the T2 keypoints from the previous step, we can now compute the XYZ position in the world coordinate using the dicom's metadata.\nDoing so will tell us the XYZ coordinates of the spine, where the X axis points from the right hand of the patient to the left. So now inferring the sides\nis rather trivial:\n> for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint.\n\nThis operation has an accuracy of 97% on determining the side of each T1 keypoint, the remaining 3% is either the T1 keypoint not being detected at all, or labelling noise.\n\nThere are some edge case that needs handling, which turns this step's logic into\n> for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint\n> \n> but if T2 on level L1/L2 is not detected, then use T2 on level L2/L3 instead, if that's missing too then use t2 on L3/L4, and so on\n\n### Step 3: Argmax T1 instance number\n\nPretty much step 1 applied to T1 keypoints - for each side and level, we find which `instance_number` has the highest confidence. \nThis step gives us 10 final T1 keypoints for the patient.\n\n### Step 4: Guess Missing T1 Keypoint\n\nUnder any incident of T1 keypoints missing, if its \"twin\" exists, then just mirror it over and call it a day\n\n> if T1 left L1/L2 went missing, mirror T1 right L1/L2 around T2 L1/L2, and blindly claim that's the T1 left L1/L2 location\n\nThis step gives minor improvements to the cv (around 0.001 to 0.002)\n\n### Step 5: Infer Axial Levels\n\nBecause we know T2 keypoints, each have their levels. Then we can use that to guess what level each axial keypoint should have.\n\n> for a given Axial keypoint, look for its closest T2 keypoint using its XYZ coordinates, and take that T2 point's level as the Axial point's level\n\nThis step is vulnerable to missing T2 keypoints (e.g. if T2's L2/L3 is missing, no Axial keypoint can correctly have L2/L3 as level). To combat this issue, \nwe just linearly interpolate all missing T2 keypoints before inferring axial levels.\n\nThis operation also has an accuracy of 97% on levels of each Axial keypoint, the remaining 3% is mostly driven from the T2 keypoints being wrong, which affects the level lookups, or the axial keypoints being missing altogether. \n\n### Step 6: Argmax Axial Instance Number\n\nPretty much step 1, but now applied to Axial keypoints - for each side and level, we find which instance number has the highest confidence. \nThis step gives us 10 final Axial keypoints for the patient.\n\n### Step 7: Guess Missing Axial Keypoints\n\nThis step sadly doesn't exist for us. I tried many approach in imputing these missing axial keypoints, but none of them really work out that well. \nOne reason is that unlike T1 keypoints where normally only one point goes missing and mirroring fixes the issue, the Axial keypoints of the same level normally go missing together.\n\nAll experiments here leads to worse cv score, so we just accepted the fact that we just can't do this and miss keypoints..\n\n\n## Crop Classification Stage\n\nAt the end of our pipeline is a classification model is a 9-class model that's construced as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F58d44332885eac17491d601cbde2043c%2Fcrop-cls.png?generation=1728673803309682&alt=media)\n\nIt's a `timm` backbone and 3 heads, one for each series type. Each head outputs 3 classes - the severity of the given image *if* the image is from that series type. For example, if an image is an axial image, the forward pass sends it though red path of timm backbone + axial head, and only the last 3 entries (the subarticular severities) are filled, the T1 and T2 logits are filled with -100 so that the probabilities are 0 after softmax.\n\n---\n\nThank you very much for reading this far. We hope this has been informational, or at least entertaining :)",
      "votes": null
    },
    {
      "id": "3015054",
      "postDate": "10/11/2024 23:30:56",
      "content": "<p>Congratulations, do you have open source code</p>",
      "rawMarkdown": "Congratulations, do you have open source code",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3015054,
      "author_name": "icw1lee",
      "author_url": "",
      "post_date": "10/11/2024 23:30:56",
      "content": "<p>Congratulations, do you have open source code</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3014987": "Congrats to all winners and those seeing themselves better data scientists than who they were before completing this competition! Despite being 2 ranks short of gold, we find this competition really interesting as there is no trivial way to approach this competition, which makes it much more interesting.\n\nMost importantly I'd like to thank @viktorcikojevic for teaming with me on this (any loads of past) competition. Without him, there's no chance I'd have come this far.\n\n## Overview\n\nOn a higher level, our pipeline is depicted in the following diagram:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F458043a2fd2439bce786b15e1f27a367%2Fkaggle-overview-image.png?generation=1728672784413838&alt=media)\n\nIt consists of three stages:\n1. Keypoint detection stage\n1. Crop proposal stage\n1. Crop classification stage\n\n### a brief word on what motivated this design choice:\n\n- This pipeline should allow all models to train at a *per-image* level, rather than *per-patient* level. There is only about ~2000 patients, so we thought this would be a better way to utilize all of the data and prevent overfitting. \n- A lot of information are inferrable between each model's output, which allowed us to include a helpful bias to the model. For example, we know T2 runs right at the center of the person, so all T1 keypoints on their left are left T1 keypoints, same for the right hand side. This means we can take some shortcuts on what the model must learn.\n\n##### When it comes to implementation, it means we did the following:\n- At the keypoint detection stage, our T1 keypoint model will only predict 5 classes: L1/L2, L2/L3, L3/L4, L4/L5, L5/S1 **and not 10**. No sides are predicted at this stage.\n- At the keypoint detection stage, our axial keypoint model will only predict 2 classes: left and right keypoint, **not 10**. No levels are predicted\n- the crop classifier's job is to take a crop in, and output 3 logits - mild / moderate / severe - it doesn't predict the condition\n- the Crop proposal stage does 3 main jobs\n  - fill in the sides of each T1 keypoint (because the keypoint model only knows the levels)\n  - fill in the level of each Axial keypoint (because the keypoint model only knows the sides)\n  - aggregate per-image predictions into the final 25 keypoints for each patient.\n\n\nIn the sections below we will describe each stage in more detail.\n\n\n## Keypoint Detection Stage\n\nHere we developed a segmentation model that outputs a map of keypoints for each image. For each of the conditions, we train a SMP (Segmentation Model Pytorch) model with `timm` backbone. The model takes the 3 consecutive `instance_number` channels as inputs and outputs a multi-channel heatmaps.\n\n### T2 Keypoint models\nFor T2 models, we just use the dataset shared by @brendanartley here https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset.\n\nWith us joining when there's on month remaining, we did not find ways to exploit the left keypoints, so we just train with the right keypoints only.\n\n### T1 & Axial Keypoint models\nFor T1 and Axial keypoint models, we just trust the labelled keypoints as-is and trained our model using those. Mostly because we're lazy (to remove all the labelling noise), and also we're a little short on time\n\nOne reason we think we can afford some labelling noise is the fact that we took shortcuts to minimise what labels the models are trained on, i.e. the T1 model doesn't care about the sides of the keypoint, so we're robust to side flips, and the Axial model doesn't care about the level, so we're robust to any noise wrt. level labels.\n\n## Crop Proposal Heuristic\n\nWe use heuristics to go from output keypoints to the final 25 keypoints for the patient. It needs to accomplish 3 things:\n\n- fill in the sides of each T1 keypoint (because the T1 keypoint model only knows the levels)\n- fill in the level of each Axial keypoint (because Axial keypoint model only knows the sides)\n- aggregate per-image predictions into the final 25 keypoints for each patient.\n\nThe overview of the Crop proposal heuristics is shown here:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2Ff88af07f0fca5f22c6367086e1c8edb1%2Fkaggle-crop-image.png?generation=1728673662221061&alt=media)\n\n### Step 1: Argmax T2 Instance Number\n\nThis step is rather simple: for each of the 5 T2 level, we find the instance number with the highest confidence from the keypoint model. After this step ends, we have 5 T2 keypoints for the patient.\n\n### Step 1: Infer T1 Sides\n\nBecause we know the T2 keypoints from the previous step, we can now compute the XYZ position in the world coordinate using the dicom's metadata.\nDoing so will tell us the XYZ coordinates of the spine, where the X axis points from the right hand of the patient to the left. So now inferring the sides\nis rather trivial:\n> for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint.\n\nThis operation has an accuracy of 97% on determining the side of each T1 keypoint, the remaining 3% is either the T1 keypoint not being detected at all, or labelling noise.\n\nThere are some edge case that needs handling, which turns this step's logic into\n> for a given T1 keypoint on level L1/L2, if it's X value is higher than its T2 on level L1/L2, then it's a left keypoint, otherwise it's the right keypoint\n> \n> but if T2 on level L1/L2 is not detected, then use T2 on level L2/L3 instead, if that's missing too then use t2 on L3/L4, and so on\n\n### Step 3: Argmax T1 instance number\n\nPretty much step 1 applied to T1 keypoints - for each side and level, we find which `instance_number` has the highest confidence. \nThis step gives us 10 final T1 keypoints for the patient.\n\n### Step 4: Guess Missing T1 Keypoint\n\nUnder any incident of T1 keypoints missing, if its \"twin\" exists, then just mirror it over and call it a day\n\n> if T1 left L1/L2 went missing, mirror T1 right L1/L2 around T2 L1/L2, and blindly claim that's the T1 left L1/L2 location\n\nThis step gives minor improvements to the cv (around 0.001 to 0.002)\n\n### Step 5: Infer Axial Levels\n\nBecause we know T2 keypoints, each have their levels. Then we can use that to guess what level each axial keypoint should have.\n\n> for a given Axial keypoint, look for its closest T2 keypoint using its XYZ coordinates, and take that T2 point's level as the Axial point's level\n\nThis step is vulnerable to missing T2 keypoints (e.g. if T2's L2/L3 is missing, no Axial keypoint can correctly have L2/L3 as level). To combat this issue, \nwe just linearly interpolate all missing T2 keypoints before inferring axial levels.\n\nThis operation also has an accuracy of 97% on levels of each Axial keypoint, the remaining 3% is mostly driven from the T2 keypoints being wrong, which affects the level lookups, or the axial keypoints being missing altogether. \n\n### Step 6: Argmax Axial Instance Number\n\nPretty much step 1, but now applied to Axial keypoints - for each side and level, we find which instance number has the highest confidence. \nThis step gives us 10 final Axial keypoints for the patient.\n\n### Step 7: Guess Missing Axial Keypoints\n\nThis step sadly doesn't exist for us. I tried many approach in imputing these missing axial keypoints, but none of them really work out that well. \nOne reason is that unlike T1 keypoints where normally only one point goes missing and mirroring fixes the issue, the Axial keypoints of the same level normally go missing together.\n\nAll experiments here leads to worse cv score, so we just accepted the fact that we just can't do this and miss keypoints..\n\n\n## Crop Classification Stage\n\nAt the end of our pipeline is a classification model is a 9-class model that's construced as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11372549%2F58d44332885eac17491d601cbde2043c%2Fcrop-cls.png?generation=1728673803309682&alt=media)\n\nIt's a `timm` backbone and 3 heads, one for each series type. Each head outputs 3 classes - the severity of the given image *if* the image is from that series type. For example, if an image is an axial image, the forward pass sends it though red path of timm backbone + axial head, and only the last 3 entries (the subarticular severities) are filled, the T1 and T2 logits are filled with -100 so that the probabilities are 0 after softmax.\n\n---\n\nThank you very much for reading this far. We hope this has been informational, or at least entertaining :)",
    "3015054": "Congratulations, do you have open source code"
  },
  "source": "meta"
}