{
  "id": 539690,
  "title": "9th place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539690",
  "author_name": "Adam Narai",
  "post_date": "2024-10-10T10:17:56.133000",
  "votes": 29,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Thanks to the organisers and congrats to the winners and all participants of this competition. I’m especially glad we worked with MRI data this year, as it gave me the opportunity to apply my experience in this field.</p>\n<p>Training code: <a href=\"https://github.com/adamnarai/kaggle-rsna-2024\" target=\"_blank\">https://github.com/adamnarai/kaggle-rsna-2024</a>  <br>\nInference code: <a href=\"https://www.kaggle.com/code/adamnarai/rsna2024-two-stage-split-global-3-base?scriptVersionId=199792909\" target=\"_blank\">https://www.kaggle.com/code/adamnarai/rsna2024-two-stage-split-global-3-base?scriptVersionId=199792909</a></p>\n<h2>Summary</h2>\n<p>Stage 1: Gaussian heatmap-based keypoint detection using DeepLabV3Plus, with separate models for each of the three series types<br>\nStage 2: Level-wise ROI classification using an ensemble of 2.5D models with GRU head and ResNet18/Swin-Tiny/ConvNeXt-Nano bases</p>\n<h2>Keypoint detection</h2>\n<p>I created a gaussian heatmap for each coordinate and trained DeepLabV3Plus models with resnet34 encoder separately for the three series types (Sag T2, Sag T1, Axi). Coordinates were then defined as the argmax of each predicted map. The inputs were 5, 3 and 3 slices (as channels), respectively for the three series types resized to 512x512 pixels and intensity normalised. Slices were selected from the middle for Sag T2 series and from predetermined mm positions relative to the middle for Sag T1 series. For the Axi series the relevant Sag T2 coordinate was projected into the axial series space and slices were selected around the one closest to this coordinate. I used rotation, sheer, channel shuffle (p=0.5), and one of sharpening or motion blur (p=0.5) augmentations. Using AdamW optimizer with 0.001 base learning rate, batch size 16 and CosineAnnealingLR scheduler I trained the three models for 30, 20 and 10 epochs, respectively with MSE loss and 5-fold cross validation.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F933188b82687f65f63b0028409e40850%2Frsna2024_seg_model2x.png?generation=1728553648423715&amp;alt=media\" alt=\"\"></p>\n<h2>Classification</h2>\n<p>I extracted 50 mm x 50 mm ROIs centred on the coordinates with 5 slices (as channel) resized to either 128x128 pixels (ResNet18 base) or 224x224 pixels (Swin-Tiny, and ConvNeXt-Nano bases) and normalised their intensity. Slices were selected the same way as for keypoint detection, only with level-dependent mm positions for Sag T1 in this case. I used rotation, sheer, and one of sharpening or motion blur (p=0.5) augmentations. I also used channel shuffle (p=0.5) only for the Axi series, since despite using a recurrent network head, it significantly improved performance. The classifiers were 2.5D models based on resnet18, swin_tiny and convnext_nano feature extractors with GRU head. Using AdamW optimizer with 0.001 or 0.00003 base learning rate, batch size 16 and StepLR or CosineAnnealingLR scheduler I trained the models for 3-5 epochs with cross entropy loss and 5-fold cross validation (optimal parameters depended on model type). I used two approaches: “split” models with separate one-output models for each core condition (spinal canal stenosis, neural foraminal narrowing and subarticular stenosis) and “global” models with outputs for all five conditions (including the two sides).</p>\n<h3>Split models</h3>\n<p>Spinal: Using Sag T2 and Axi (centred on mean of left and right coordinates) ROIs.  <br>\nForaminal: Using Sag T1 ROIs, same model for both sides.  <br>\nSubarticular: Using Axi ROIs, same model for both sides, but right side images are mirrored.  <br>\nPredictions were obtained for each level and side using these three models and concatenated to make the 25 outputs per study.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2Ff3c84f957b2678f463e92ea32f40d6e9%2Frsna_2024_split_model2x.png?generation=1728553678267138&amp;alt=media\" alt=\"\"></p>\n<h3>Global model</h3>\n<p>Sag T2, left/right Sag T2 and left/right (right mirrored) Axi ROIs were used as inputs to the model, resulting in predictions for all 5 conditions at a given level. The same feature extractor model was used for both left and right side.  <br>\nPredictions were obtained for each level and concatenated to make the 25 outputs per study.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F094459632507cb2c3512e4ef9baf3300%2Frsna2024_global_model2x.png?generation=1728553689737251&amp;alt=media\" alt=\"\"></p>\n<h2>Submission</h2>\n<p>My final submission was an ensemble of 6 models, combining both split and global models with three different feature extractors ResNet18, Swin-Tiny, and ConvNeXt-Nano. Each of these models was further a 5-fold ensemble, and the final predictions were the simple mean of these ensembles.</p>",
  "messages": [
    {
      "id": 3013599,
      "postDate": "2024-10-10T10:17:56.133Z",
      "content": "<p>Thanks to the organisers and congrats to the winners and all participants of this competition. I’m especially glad we worked with MRI data this year, as it gave me the opportunity to apply my experience in this field.</p>\n<p>Training code: <a href=\"https://github.com/adamnarai/kaggle-rsna-2024\" target=\"_blank\">https://github.com/adamnarai/kaggle-rsna-2024</a>  <br>\nInference code: <a href=\"https://www.kaggle.com/code/adamnarai/rsna2024-two-stage-split-global-3-base?scriptVersionId=199792909\" target=\"_blank\">https://www.kaggle.com/code/adamnarai/rsna2024-two-stage-split-global-3-base?scriptVersionId=199792909</a></p>\n<h2>Summary</h2>\n<p>Stage 1: Gaussian heatmap-based keypoint detection using DeepLabV3Plus, with separate models for each of the three series types<br>\nStage 2: Level-wise ROI classification using an ensemble of 2.5D models with GRU head and ResNet18/Swin-Tiny/ConvNeXt-Nano bases</p>\n<h2>Keypoint detection</h2>\n<p>I created a gaussian heatmap for each coordinate and trained DeepLabV3Plus models with resnet34 encoder separately for the three series types (Sag T2, Sag T1, Axi). Coordinates were then defined as the argmax of each predicted map. The inputs were 5, 3 and 3 slices (as channels), respectively for the three series types resized to 512x512 pixels and intensity normalised. Slices were selected from the middle for Sag T2 series and from predetermined mm positions relative to the middle for Sag T1 series. For the Axi series the relevant Sag T2 coordinate was projected into the axial series space and slices were selected around the one closest to this coordinate. I used rotation, sheer, channel shuffle (p=0.5), and one of sharpening or motion blur (p=0.5) augmentations. Using AdamW optimizer with 0.001 base learning rate, batch size 16 and CosineAnnealingLR scheduler I trained the three models for 30, 20 and 10 epochs, respectively with MSE loss and 5-fold cross validation.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F933188b82687f65f63b0028409e40850%2Frsna2024_seg_model2x.png?generation=1728553648423715&amp;alt=media\" alt=\"\"></p>\n<h2>Classification</h2>\n<p>I extracted 50 mm x 50 mm ROIs centred on the coordinates with 5 slices (as channel) resized to either 128x128 pixels (ResNet18 base) or 224x224 pixels (Swin-Tiny, and ConvNeXt-Nano bases) and normalised their intensity. Slices were selected the same way as for keypoint detection, only with level-dependent mm positions for Sag T1 in this case. I used rotation, sheer, and one of sharpening or motion blur (p=0.5) augmentations. I also used channel shuffle (p=0.5) only for the Axi series, since despite using a recurrent network head, it significantly improved performance. The classifiers were 2.5D models based on resnet18, swin_tiny and convnext_nano feature extractors with GRU head. Using AdamW optimizer with 0.001 or 0.00003 base learning rate, batch size 16 and StepLR or CosineAnnealingLR scheduler I trained the models for 3-5 epochs with cross entropy loss and 5-fold cross validation (optimal parameters depended on model type). I used two approaches: “split” models with separate one-output models for each core condition (spinal canal stenosis, neural foraminal narrowing and subarticular stenosis) and “global” models with outputs for all five conditions (including the two sides).</p>\n<h3>Split models</h3>\n<p>Spinal: Using Sag T2 and Axi (centred on mean of left and right coordinates) ROIs.  <br>\nForaminal: Using Sag T1 ROIs, same model for both sides.  <br>\nSubarticular: Using Axi ROIs, same model for both sides, but right side images are mirrored.  <br>\nPredictions were obtained for each level and side using these three models and concatenated to make the 25 outputs per study.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2Ff3c84f957b2678f463e92ea32f40d6e9%2Frsna_2024_split_model2x.png?generation=1728553678267138&amp;alt=media\" alt=\"\"></p>\n<h3>Global model</h3>\n<p>Sag T2, left/right Sag T2 and left/right (right mirrored) Axi ROIs were used as inputs to the model, resulting in predictions for all 5 conditions at a given level. The same feature extractor model was used for both left and right side.  <br>\nPredictions were obtained for each level and concatenated to make the 25 outputs per study.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F094459632507cb2c3512e4ef9baf3300%2Frsna2024_global_model2x.png?generation=1728553689737251&amp;alt=media\" alt=\"\"></p>\n<h2>Submission</h2>\n<p>My final submission was an ensemble of 6 models, combining both split and global models with three different feature extractors ResNet18, Swin-Tiny, and ConvNeXt-Nano. Each of these models was further a 5-fold ensemble, and the final predictions were the simple mean of these ensembles.</p>",
      "rawMarkdown": "Thanks to the organisers and congrats to the winners and all participants of this competition. I’m especially glad we worked with MRI data this year, as it gave me the opportunity to apply my experience in this field.\n\nTraining code: https://github.com/adamnarai/kaggle-rsna-2024  \nInference code: https://www.kaggle.com/code/adamnarai/rsna2024-two-stage-split-global-3-base?scriptVersionId=199792909\n\n## Summary\n\nStage 1: Gaussian heatmap-based keypoint detection using DeepLabV3Plus, with separate models for each of the three series types\nStage 2: Level-wise ROI classification using an ensemble of 2.5D models with GRU head and ResNet18/Swin-Tiny/ConvNeXt-Nano bases\n\n## Keypoint detection\n\nI created a gaussian heatmap for each coordinate and trained DeepLabV3Plus models with resnet34 encoder separately for the three series types (Sag T2, Sag T1, Axi). Coordinates were then defined as the argmax of each predicted map. The inputs were 5, 3 and 3 slices (as channels), respectively for the three series types resized to 512x512 pixels and intensity normalised. Slices were selected from the middle for Sag T2 series and from predetermined mm positions relative to the middle for Sag T1 series. For the Axi series the relevant Sag T2 coordinate was projected into the axial series space and slices were selected around the one closest to this coordinate. I used rotation, sheer, channel shuffle (p=0.5), and one of sharpening or motion blur (p=0.5) augmentations. Using AdamW optimizer with 0.001 base learning rate, batch size 16 and CosineAnnealingLR scheduler I trained the three models for 30, 20 and 10 epochs, respectively with MSE loss and 5-fold cross validation.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F933188b82687f65f63b0028409e40850%2Frsna2024_seg_model2x.png?generation=1728553648423715&alt=media)\n\n## Classification\n\nI extracted 50 mm x 50 mm ROIs centred on the coordinates with 5 slices (as channel) resized to either 128x128 pixels (ResNet18 base) or 224x224 pixels (Swin-Tiny, and ConvNeXt-Nano bases) and normalised their intensity. Slices were selected the same way as for keypoint detection, only with level-dependent mm positions for Sag T1 in this case. I used rotation, sheer, and one of sharpening or motion blur (p=0.5) augmentations. I also used channel shuffle (p=0.5) only for the Axi series, since despite using a recurrent network head, it significantly improved performance. The classifiers were 2.5D models based on resnet18, swin\\_tiny and convnext\\_nano feature extractors with GRU head. Using AdamW optimizer with 0.001 or 0.00003 base learning rate, batch size 16 and StepLR or CosineAnnealingLR scheduler I trained the models for 3-5 epochs with cross entropy loss and 5-fold cross validation (optimal parameters depended on model type). I used two approaches: “split” models with separate one-output models for each core condition (spinal canal stenosis, neural foraminal narrowing and subarticular stenosis) and “global” models with outputs for all five conditions (including the two sides).\n\n### Split models\n\nSpinal: Using Sag T2 and Axi (centred on mean of left and right coordinates) ROIs.  \nForaminal: Using Sag T1 ROIs, same model for both sides.  \nSubarticular: Using Axi ROIs, same model for both sides, but right side images are mirrored.  \nPredictions were obtained for each level and side using these three models and concatenated to make the 25 outputs per study.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2Ff3c84f957b2678f463e92ea32f40d6e9%2Frsna_2024_split_model2x.png?generation=1728553678267138&alt=media)\n\n### Global model\n\nSag T2, left/right Sag T2 and left/right (right mirrored) Axi ROIs were used as inputs to the model, resulting in predictions for all 5 conditions at a given level. The same feature extractor model was used for both left and right side.  \nPredictions were obtained for each level and concatenated to make the 25 outputs per study.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F094459632507cb2c3512e4ef9baf3300%2Frsna2024_global_model2x.png?generation=1728553689737251&alt=media)\n\n## Submission\n\nMy final submission was an ensemble of 6 models, combining both split and global models with three different feature extractors ResNet18, Swin-Tiny, and ConvNeXt-Nano. Each of these models was further a 5-fold ensemble, and the final predictions were the simple mean of these ensembles.",
      "votes": 29
    },
    {
      "id": 3017684,
      "postDate": "2024-10-15T05:43:16.670Z",
      "content": "<p>Congratulations!!! I appreciate having the opportunity to explore your approach, because your solution are elegant and your code is clean.</p>\n<p>Were 5, 3 and 3 slices (as channels) chosen as hyperparameters respectively for the three types of series experimentally? Or were there any additional arguments based on the dataset?</p>\n<p>Did you try other solutions instead of DeepLabV3Plus and how did you settle on it?</p>",
      "rawMarkdown": "Congratulations!!! I appreciate having the opportunity to explore your approach, because your solution are elegant and your code is clean.\n\nWere 5, 3 and 3 slices (as channels) chosen as hyperparameters respectively for the three types of series experimentally? Or were there any additional arguments based on the dataset?\n\nDid you try other solutions instead of DeepLabV3Plus and how did you settle on it?",
      "votes": 1,
      "replies": [
        {
          "id": 3018180,
          "postDate": "2024-10-15T15:11:12.683Z",
          "content": "<p>Thanks! I'm glad you find my solution and codes useful.<br>\nYes, slice numbers were chosen experimentally, if I remember correctly, using 3 slices had basically the same performance or even better than using 5 and I generally prefer to use the simpler/faster model when the performance difference is negligible. My hypothesis is that for foraminal narrowing and subarticular stenosis there are only 1-2 meaningful slices in most of the cases, so that's why using more does not really help.<br>\nI tried a few different segmentation models from the SMP package and experimented with different encoders and DeepLabV3+ with resnet34 seemed to work best. However, keypoint detection models quickly reached very good performance, so I haven't done too much finetuning on them.</p>",
          "rawMarkdown": "Thanks! I'm glad you find my solution and codes useful.\nYes, slice numbers were chosen experimentally, if I remember correctly, using 3 slices had basically the same performance or even better than using 5 and I generally prefer to use the simpler/faster model when the performance difference is negligible. My hypothesis is that for foraminal narrowing and subarticular stenosis there are only 1-2 meaningful slices in most of the cases, so that's why using more does not really help.\nI tried a few different segmentation models from the SMP package and experimented with different encoders and DeepLabV3+ with resnet34 seemed to work best. However, keypoint detection models quickly reached very good performance, so I haven't done too much finetuning on them.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3014914,
      "postDate": "2024-10-11T17:47:21.903Z",
      "content": "<p>Congratulations!<br>\nI would like to run your training script by cloning the github repository. Could you write a walk through explaining the steps to take to perform the training phase? I tried to run train.py, but I get errors like:<br>\n<code>FileNotFoundError: [Errno 2] No such file or directory: ../AdamNarai_Solution/data/raw/../processed/train_label_coordinates.csv</code><br>\nIs there any data preparation step? </p>",
      "rawMarkdown": "Congratulations!\nI would like to run your training script by cloning the github repository. Could you write a walk through explaining the steps to take to perform the training phase? I tried to run train.py, but I get errors like:\n```FileNotFoundError: [Errno 2] No such file or directory: ../AdamNarai_Solution/data/raw/../processed/train_label_coordinates.csv```\nIs there any data preparation step? ",
      "votes": 1,
      "replies": [
        {
          "id": 3014956,
          "postDate": "2024-10-11T18:24:17.840Z",
          "content": "<p>Thanks!<br>\nI'm planning to add proper documentation to the repo and maybe restructure the codes to make it more intuitive to reproduce my final solution, I just haven't had the time yet.<br>\nIn short, you should run preproc/normalize_coordinates.py first (this should fix your error) and then you can train the models using main.py. Model parameters can be tuned using the YAML files in the experiments folder, the current parameters are set for the ResNet models.</p>",
          "rawMarkdown": "Thanks!\nI'm planning to add proper documentation to the repo and maybe restructure the codes to make it more intuitive to reproduce my final solution, I just haven't had the time yet.\nIn short, you should run preproc/normalize_coordinates.py first (this should fix your error) and then you can train the models using main.py. Model parameters can be tuned using the YAML files in the experiments folder, the current parameters are set for the ResNet models.",
          "replies": [
            {
              "id": 3015643,
              "postDate": "2024-10-12T17:33:39.393Z",
              "content": "<p>In <code>dataset.py</code> script, line 237 we have,</p>\n<p><code>series_list = series_coords['series_id'].unique().tolist()</code> <br>\nbut I believe it shall be <br>\n<code>series_list = series_coords['series_id'].tolist()</code><br>\notherwise the code on the next line<br>\n<code>series_id = self.most_frequent(series_list)</code><br>\nwill be ineffective or am I missing something here?</p>\n<p>Also, why the (normalized?) instance number is hard-coded here (<code>dataset.py:444</code>)<br>\n<code>instance_number = 16.8</code></p>",
              "rawMarkdown": "In `dataset.py` script, line 237 we have,\n\n`series_list = series_coords['series_id'].unique().tolist()` \nbut I believe it shall be \n`series_list = series_coords['series_id'].tolist()`\notherwise the code on the next line\n`series_id = self.most_frequent(series_list)`\nwill be ineffective or am I missing something here?\n\nAlso, why the (normalized?) instance number is hard-coded here (`dataset.py:444`)\n`instance_number = 16.8`\n",
              "votes": 1
            },
            {
              "id": 3015687,
              "postDate": "2024-10-12T19:03:45.773Z",
              "content": "<p>Yeah, that's a bug, nice catch!</p>\n<p>The instance_number_type argument determines the interpretation of the instance_number, in this case it is the mm position relative to the center of the image volume. 16.8 was calculated as the average across the whole training data showing relatively small variance. You could interpret it as an anatomical parameter, as long as the image volume is centered on the spine (as it is standard practice) this should reliably estimate the position of the neural foramen (at least with the use of multiple slices several millimeters apart, as it is the case here). </p>",
              "rawMarkdown": "Yeah, that's a bug, nice catch!\n\nThe instance_number_type argument determines the interpretation of the instance_number, in this case it is the mm position relative to the center of the image volume. 16.8 was calculated as the average across the whole training data showing relatively small variance. You could interpret it as an anatomical parameter, as long as the image volume is centered on the spine (as it is standard practice) this should reliably estimate the position of the neural foramen (at least with the use of multiple slices several millimeters apart, as it is the case here). "
            },
            {
              "id": 3024324,
              "postDate": "2024-10-21T14:09:51.803Z",
              "content": "<p>could you tell me how to solve this question?<br>\nFileNotFoundError: [Errno 2] No such file or directory: 'G:\\competition\\kaggle-rsna-2024-master\\models\\rsna-2024-giddy-monkey-1266\\splits.csv'</p>",
              "rawMarkdown": "could you tell me how to solve this question?\nFileNotFoundError: [Errno 2] No such file or directory: 'G:\\\\competition\\\\kaggle-rsna-2024-master\\\\models\\\\rsna-2024-giddy-monkey-1266\\\\splits.csv'",
              "votes": 1
            },
            {
              "id": 3030651,
              "postDate": "2024-10-28T18:16:41.647Z",
              "content": "<p>Sorry, the splits.csv file is not uploaded with the models since it is not used for inference (only for the OOF validation). Now you can download it from here: <a href=\"https://www.kaggle.com/datasets/adamnarai/rsna-2024-training-splits\" target=\"_blank\">https://www.kaggle.com/datasets/adamnarai/rsna-2024-training-splits</a><br>\nThe same file was used for all models here, so you can just copy it into each model folder.</p>",
              "rawMarkdown": "Sorry, the splits.csv file is not uploaded with the models since it is not used for inference (only for the OOF validation). Now you can download it from here: https://www.kaggle.com/datasets/adamnarai/rsna-2024-training-splits\nThe same file was used for all models here, so you can just copy it into each model folder."
            }
          ]
        }
      ]
    },
    {
      "id": 3013846,
      "postDate": "2024-10-10T15:38:55.203Z",
      "content": "<p>Congratulations!!! I would like to explore a little bit more your solution cause is well documented and a lot of details.</p>\n<p>I opened your submission notebook but datasets are private. Any plan to make it public? Thank you!</p>",
      "rawMarkdown": "Congratulations!!! I would like to explore a little bit more your solution cause is well documented and a lot of details.\n\nI opened your submission notebook but datasets are private. Any plan to make it public? Thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 3013872,
          "postDate": "2024-10-10T16:25:18.747Z",
          "content": "<p>Thanks! Sure, I just set them to public.</p>",
          "rawMarkdown": "Thanks! Sure, I just set them to public."
        }
      ]
    },
    {
      "id": 3013812,
      "postDate": "2024-10-10T14:51:48.527Z",
      "content": "<p>Great write-up <a href=\"https://www.kaggle.com/adamnarai\" target=\"_blank\">@adamnarai</a>.</p>\n<p>How much improvement did you get by including Split models and Global models? Was one approach better than the other?</p>",
      "rawMarkdown": "Great write-up @adamnarai.\n\nHow much improvement did you get by including Split models and Global models? Was one approach better than the other?",
      "votes": 1,
      "replies": [
        {
          "id": 3013862,
          "postDate": "2024-10-10T16:12:04.443Z",
          "content": "<p>Thanks! <br>\nThe split and global models were quite similar, but generally global models were slightly better and combining the two gave a 2-3% improvement on CV score. <br>\nWith the resnet+swin+convnext ensemble this effect was much smaller, the split+global combination resulted in about 1% improvement on CV score.</p>",
          "rawMarkdown": "Thanks! \nThe split and global models were quite similar, but generally global models were slightly better and combining the two gave a 2-3% improvement on CV score. \nWith the resnet+swin+convnext ensemble this effect was much smaller, the split+global combination resulted in about 1% improvement on CV score.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 3014561,
      "postDate": "2024-10-11T10:58:14.353Z",
      "content": "<p>Thanks for your sharing! Do you mean that you choose 5 or 3 instances(images) from each seires_id? If not, I wanna know what's the shape of the input data in the classification stage. I am a rookie, hoping for your kind answer. :)</p>",
      "rawMarkdown": "Thanks for your sharing! Do you mean that you choose 5 or 3 instances(images) from each seires_id? If not, I wanna know what's the shape of the input data in the classification stage. I am a rookie, hoping for your kind answer. :)",
      "replies": [
        {
          "id": 3014598,
          "postDate": "2024-10-11T12:06:11.457Z",
          "content": "<p>Yes, exactly, 5 or 3 images as channels for stage 1, so the inputs were sized like 512x512x5 or 512x512x3 for keypoint detection and 128x128x5 or 224x224x5 in the classification stage.</p>",
          "rawMarkdown": "Yes, exactly, 5 or 3 images as channels for stage 1, so the inputs were sized like 512x512x5 or 512x512x3 for keypoint detection and 128x128x5 or 224x224x5 in the classification stage.",
          "replies": [
            {
              "id": 3014614,
              "postDate": "2024-10-11T12:18:37.097Z",
              "content": "<p>Thanks for your reply! I also wanna know how you choose the images in the training and inference? Do you choose them randomly or other ways? Besides, for some study, there are two series for the same description so how do you handle this?</p>",
              "rawMarkdown": "Thanks for your reply! I also wanna know how you choose the images in the training and inference? Do you choose them randomly or other ways? Besides, for some study, there are two series for the same description so how do you handle this?"
            },
            {
              "id": 3014620,
              "postDate": "2024-10-11T12:28:57.327Z",
              "content": "<p>You can find this in the description above: \"Slices were selected from the middle for Sag T2 series and from predetermined mm positions relative to the middle for Sag T1 series. For the Axi series the relevant Sag T2 coordinate was projected into the axial series space and slices were selected around the one closest to this coordinate.\"<br>\nI simply choose the first series for sagittal images, since in the training set it was very rare to have more than one series. For axial images, I checked all series to find the closest slice.</p>",
              "rawMarkdown": "You can find this in the description above: \"Slices were selected from the middle for Sag T2 series and from predetermined mm positions relative to the middle for Sag T1 series. For the Axi series the relevant Sag T2 coordinate was projected into the axial series space and slices were selected around the one closest to this coordinate.\"\nI simply choose the first series for sagittal images, since in the training set it was very rare to have more than one series. For axial images, I checked all series to find the closest slice."
            }
          ]
        }
      ]
    },
    {
      "id": 3025834,
      "postDate": "2024-10-23T07:18:02.597Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3017684,
      "author_name": "prince_lvov",
      "author_url": "",
      "post_date": "2024-10-15T05:43:16.670000",
      "content": "<p>Congratulations!!! I appreciate having the opportunity to explore your approach, because your solution are elegant and your code is clean.</p>\n<p>Were 5, 3 and 3 slices (as channels) chosen as hyperparameters respectively for the three types of series experimentally? Or were there any additional arguments based on the dataset?</p>\n<p>Did you try other solutions instead of DeepLabV3Plus and how did you settle on it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3018180,
          "author_name": "Adam Narai",
          "author_url": "",
          "post_date": "2024-10-15T15:11:12.683000",
          "content": "<p>Thanks! I'm glad you find my solution and codes useful.<br>\nYes, slice numbers were chosen experimentally, if I remember correctly, using 3 slices had basically the same performance or even better than using 5 and I generally prefer to use the simpler/faster model when the performance difference is negligible. My hypothesis is that for foraminal narrowing and subarticular stenosis there are only 1-2 meaningful slices in most of the cases, so that's why using more does not really help.<br>\nI tried a few different segmentation models from the SMP package and experimented with different encoders and DeepLabV3+ with resnet34 seemed to work best. However, keypoint detection models quickly reached very good performance, so I haven't done too much finetuning on them.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3014914,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2024-10-11T17:47:21.903000",
      "content": "<p>Congratulations!<br>\nI would like to run your training script by cloning the github repository. Could you write a walk through explaining the steps to take to perform the training phase? I tried to run train.py, but I get errors like:<br>\n<code>FileNotFoundError: [Errno 2] No such file or directory: ../AdamNarai_Solution/data/raw/../processed/train_label_coordinates.csv</code><br>\nIs there any data preparation step? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3014956,
          "author_name": "Adam Narai",
          "author_url": "",
          "post_date": "2024-10-11T18:24:17.840000",
          "content": "<p>Thanks!<br>\nI'm planning to add proper documentation to the repo and maybe restructure the codes to make it more intuitive to reproduce my final solution, I just haven't had the time yet.<br>\nIn short, you should run preproc/normalize_coordinates.py first (this should fix your error) and then you can train the models using main.py. Model parameters can be tuned using the YAML files in the experiments folder, the current parameters are set for the ResNet models.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3015643,
              "author_name": "Rasoul Mojtahedzadeh",
              "author_url": "",
              "post_date": "2024-10-12T17:33:39.393000",
              "content": "<p>In <code>dataset.py</code> script, line 237 we have,</p>\n<p><code>series_list = series_coords['series_id'].unique().tolist()</code> <br>\nbut I believe it shall be <br>\n<code>series_list = series_coords['series_id'].tolist()</code><br>\notherwise the code on the next line<br>\n<code>series_id = self.most_frequent(series_list)</code><br>\nwill be ineffective or am I missing something here?</p>\n<p>Also, why the (normalized?) instance number is hard-coded here (<code>dataset.py:444</code>)<br>\n<code>instance_number = 16.8</code></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3015687,
              "author_name": "Adam Narai",
              "author_url": "",
              "post_date": "2024-10-12T19:03:45.773000",
              "content": "<p>Yeah, that's a bug, nice catch!</p>\n<p>The instance_number_type argument determines the interpretation of the instance_number, in this case it is the mm position relative to the center of the image volume. 16.8 was calculated as the average across the whole training data showing relatively small variance. You could interpret it as an anatomical parameter, as long as the image volume is centered on the spine (as it is standard practice) this should reliably estimate the position of the neural foramen (at least with the use of multiple slices several millimeters apart, as it is the case here). </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3024324,
              "author_name": "wudushang",
              "author_url": "",
              "post_date": "2024-10-21T14:09:51.803000",
              "content": "<p>could you tell me how to solve this question?<br>\nFileNotFoundError: [Errno 2] No such file or directory: 'G:\\competition\\kaggle-rsna-2024-master\\models\\rsna-2024-giddy-monkey-1266\\splits.csv'</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3030651,
              "author_name": "Adam Narai",
              "author_url": "",
              "post_date": "2024-10-28T18:16:41.647000",
              "content": "<p>Sorry, the splits.csv file is not uploaded with the models since it is not used for inference (only for the OOF validation). Now you can download it from here: <a href=\"https://www.kaggle.com/datasets/adamnarai/rsna-2024-training-splits\" target=\"_blank\">https://www.kaggle.com/datasets/adamnarai/rsna-2024-training-splits</a><br>\nThe same file was used for all models here, so you can just copy it into each model folder.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3013846,
      "author_name": "karelbecerra",
      "author_url": "",
      "post_date": "2024-10-10T15:38:55.203000",
      "content": "<p>Congratulations!!! I would like to explore a little bit more your solution cause is well documented and a lot of details.</p>\n<p>I opened your submission notebook but datasets are private. Any plan to make it public? Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3013872,
          "author_name": "Adam Narai",
          "author_url": "",
          "post_date": "2024-10-10T16:25:18.747000",
          "content": "<p>Thanks! Sure, I just set them to public.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3013812,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-10-10T14:51:48.527000",
      "content": "<p>Great write-up <a href=\"https://www.kaggle.com/adamnarai\" target=\"_blank\">@adamnarai</a>.</p>\n<p>How much improvement did you get by including Split models and Global models? Was one approach better than the other?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3013862,
          "author_name": "Adam Narai",
          "author_url": "",
          "post_date": "2024-10-10T16:12:04.443000",
          "content": "<p>Thanks! <br>\nThe split and global models were quite similar, but generally global models were slightly better and combining the two gave a 2-3% improvement on CV score. <br>\nWith the resnet+swin+convnext ensemble this effect was much smaller, the split+global combination resulted in about 1% improvement on CV score.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3014561,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2024-10-11T10:58:14.353000",
      "content": "<p>Thanks for your sharing! Do you mean that you choose 5 or 3 instances(images) from each seires_id? If not, I wanna know what's the shape of the input data in the classification stage. I am a rookie, hoping for your kind answer. :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3014598,
          "author_name": "Adam Narai",
          "author_url": "",
          "post_date": "2024-10-11T12:06:11.457000",
          "content": "<p>Yes, exactly, 5 or 3 images as channels for stage 1, so the inputs were sized like 512x512x5 or 512x512x3 for keypoint detection and 128x128x5 or 224x224x5 in the classification stage.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3014614,
              "author_name": "I2nfinit3y",
              "author_url": "",
              "post_date": "2024-10-11T12:18:37.097000",
              "content": "<p>Thanks for your reply! I also wanna know how you choose the images in the training and inference? Do you choose them randomly or other ways? Besides, for some study, there are two series for the same description so how do you handle this?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3014620,
              "author_name": "Adam Narai",
              "author_url": "",
              "post_date": "2024-10-11T12:28:57.327000",
              "content": "<p>You can find this in the description above: \"Slices were selected from the middle for Sag T2 series and from predetermined mm positions relative to the middle for Sag T1 series. For the Axi series the relevant Sag T2 coordinate was projected into the axial series space and slices were selected around the one closest to this coordinate.\"<br>\nI simply choose the first series for sagittal images, since in the training set it was very rare to have more than one series. For axial images, I checked all series to find the closest slice.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3025834,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-23T07:18:02.597000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3013599": "Thanks to the organisers and congrats to the winners and all participants of this competition. I’m especially glad we worked with MRI data this year, as it gave me the opportunity to apply my experience in this field.\n\nTraining code: https://github.com/adamnarai/kaggle-rsna-2024  \nInference code: https://www.kaggle.com/code/adamnarai/rsna2024-two-stage-split-global-3-base?scriptVersionId=199792909\n\n## Summary\n\nStage 1: Gaussian heatmap-based keypoint detection using DeepLabV3Plus, with separate models for each of the three series types\nStage 2: Level-wise ROI classification using an ensemble of 2.5D models with GRU head and ResNet18/Swin-Tiny/ConvNeXt-Nano bases\n\n## Keypoint detection\n\nI created a gaussian heatmap for each coordinate and trained DeepLabV3Plus models with resnet34 encoder separately for the three series types (Sag T2, Sag T1, Axi). Coordinates were then defined as the argmax of each predicted map. The inputs were 5, 3 and 3 slices (as channels), respectively for the three series types resized to 512x512 pixels and intensity normalised. Slices were selected from the middle for Sag T2 series and from predetermined mm positions relative to the middle for Sag T1 series. For the Axi series the relevant Sag T2 coordinate was projected into the axial series space and slices were selected around the one closest to this coordinate. I used rotation, sheer, channel shuffle (p=0.5), and one of sharpening or motion blur (p=0.5) augmentations. Using AdamW optimizer with 0.001 base learning rate, batch size 16 and CosineAnnealingLR scheduler I trained the three models for 30, 20 and 10 epochs, respectively with MSE loss and 5-fold cross validation.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F933188b82687f65f63b0028409e40850%2Frsna2024_seg_model2x.png?generation=1728553648423715&alt=media)\n\n## Classification\n\nI extracted 50 mm x 50 mm ROIs centred on the coordinates with 5 slices (as channel) resized to either 128x128 pixels (ResNet18 base) or 224x224 pixels (Swin-Tiny, and ConvNeXt-Nano bases) and normalised their intensity. Slices were selected the same way as for keypoint detection, only with level-dependent mm positions for Sag T1 in this case. I used rotation, sheer, and one of sharpening or motion blur (p=0.5) augmentations. I also used channel shuffle (p=0.5) only for the Axi series, since despite using a recurrent network head, it significantly improved performance. The classifiers were 2.5D models based on resnet18, swin\\_tiny and convnext\\_nano feature extractors with GRU head. Using AdamW optimizer with 0.001 or 0.00003 base learning rate, batch size 16 and StepLR or CosineAnnealingLR scheduler I trained the models for 3-5 epochs with cross entropy loss and 5-fold cross validation (optimal parameters depended on model type). I used two approaches: “split” models with separate one-output models for each core condition (spinal canal stenosis, neural foraminal narrowing and subarticular stenosis) and “global” models with outputs for all five conditions (including the two sides).\n\n### Split models\n\nSpinal: Using Sag T2 and Axi (centred on mean of left and right coordinates) ROIs.  \nForaminal: Using Sag T1 ROIs, same model for both sides.  \nSubarticular: Using Axi ROIs, same model for both sides, but right side images are mirrored.  \nPredictions were obtained for each level and side using these three models and concatenated to make the 25 outputs per study.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2Ff3c84f957b2678f463e92ea32f40d6e9%2Frsna_2024_split_model2x.png?generation=1728553678267138&alt=media)\n\n### Global model\n\nSag T2, left/right Sag T2 and left/right (right mirrored) Axi ROIs were used as inputs to the model, resulting in predictions for all 5 conditions at a given level. The same feature extractor model was used for both left and right side.  \nPredictions were obtained for each level and concatenated to make the 25 outputs per study.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8375965%2F094459632507cb2c3512e4ef9baf3300%2Frsna2024_global_model2x.png?generation=1728553689737251&alt=media)\n\n## Submission\n\nMy final submission was an ensemble of 6 models, combining both split and global models with three different feature extractors ResNet18, Swin-Tiny, and ConvNeXt-Nano. Each of these models was further a 5-fold ensemble, and the final predictions were the simple mean of these ensembles.",
    "3017684": "Congratulations!!! I appreciate having the opportunity to explore your approach, because your solution are elegant and your code is clean.\n\nWere 5, 3 and 3 slices (as channels) chosen as hyperparameters respectively for the three types of series experimentally? Or were there any additional arguments based on the dataset?\n\nDid you try other solutions instead of DeepLabV3Plus and how did you settle on it?",
    "3014914": "Congratulations!\nI would like to run your training script by cloning the github repository. Could you write a walk through explaining the steps to take to perform the training phase? I tried to run train.py, but I get errors like:\n```FileNotFoundError: [Errno 2] No such file or directory: ../AdamNarai_Solution/data/raw/../processed/train_label_coordinates.csv```\nIs there any data preparation step? ",
    "3013846": "Congratulations!!! I would like to explore a little bit more your solution cause is well documented and a lot of details.\n\nI opened your submission notebook but datasets are private. Any plan to make it public? Thank you!",
    "3013812": "Great write-up @adamnarai.\n\nHow much improvement did you get by including Split models and Global models? Was one approach better than the other?",
    "3014561": "Thanks for your sharing! Do you mean that you choose 5 or 3 instances(images) from each seires_id? If not, I wanna know what's the shape of the input data in the classification stage. I am a rookie, hoping for your kind answer. :)",
    "3025834": ""
  }
}