{
  "id": 539443,
  "title": "4th place solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/539443",
  "author_name": "tattaka",
  "post_date": "2024-10-09T00:21:51.816000",
  "votes": 76,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Congrats to all prize and medal winners! This year's RSNA competition required us to carefully handle data and build a pipeline, which was a lot of fun. We share our solution.</p>\n<h1>Summary</h1>\n<p>Our solution detects the keypoint, which is the region of interest in the symptom, and builds a classification model using the surrounding crops as input.<br>\nThe results of each model are refined by the stacking model and submitted as the final result.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F11c196e4f15041d980ae018aa1ba5e5c%2Foverall_pipeline.png?generation=1728433161180972&amp;alt=media\" alt=\"\"></p>\n<h1>Keypoint detection Model</h1>\n<h2>Disk Level detection model and Keypoint detection for Axial, Sagittal T1/T2 (@yu4u)</h2>\n<p>Resize each axial slice to 128×128 and use a 2.5DCNN + LSTM model to estimate which level (L1, L2, …, S1) each slice belongs to. Subsequently, detect the boundary slices between each level. From these slices (up to five), use a UNet model to detect the left and right keypoints.<br>\nSimilarly, resize the sagittal T1 slices to 128×128 and use a 2.5D CNN model to identify the left and right slices belonging to the foraminal zone where keypoints should be detected. Then, individually detect keypoints for five levels from these left and right slices.<br>\nFor sagittal T2/STIR, simply extract the middle slice of the series and detect keypoints for the five levels.</p>\n<h2>Keypoint detection model for Sagittal (@tattaka)</h2>\n<p>We resized each of the Sagittal T1, T2/STIR images to 20x256x256 and predicted the xy coordinates of the keypoints.<br>\nThe xy coordinates of the keypoints were taken from the shared <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">Lumbar Coordinate Dataset</a>.<br>\nFor the backbone, we used caformer_s18, convnext_tiny, resnetrs50, and swinv2_tiny, and applied SCSE attention to the UNet Decoder.<br>\nAs for the loss function, we used BCELoss * 0.2 + DICELoss * 0.8.</p>\n<h1>Classification model</h1>\n<h2>Multi-view input, multi-condition output model (@tattaka)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc5d845a64af1ad8c7e9da045b73d3469%2Ftattaka_model.png?generation=1728433227519640&amp;alt=media\" alt=\"\"></p>\n<p>We crop the areas around the keypoints inferred from the volumes of Sagittal T1, Sagittal T2/STIR, and Axial T2, and use them as inputs to classify the conditions at each level.<br>\nEach image is cropped to a size that is twice the distance between the neighboring keypoints. For sagittal images, padding is applied if all slices are fewer than 30, and linear interpolation is used if there are more. (There is also a model variation that simply uses linear interpolation to resize to 20.)<br>\nFor axial images, slices are taken from the range of ±2 around the predicted gaps between the discs.<br>\nAfter the images are input into a 2D model backbone, features are extracted from each slice, and the final output is obtained using a transformer encoder and attention pooling on the extracted features. </p>\n<p>The model used for the final submission includes variations such as:</p>\n<ul>\n<li>A model that processes cropped slices with one or two backbones,</li>\n<li>Multiple augmentation patterns,</li>\n<li>Various preprocessing patterns for multiple slices.</li>\n</ul>\n<p>The backbones used were caformer_s18, resnetrs50, <a href=\"https://github.com/naver-ai/rdnet\" target=\"_blank\">rdnet_tiny</a>, and maxxvitv2_nano.<br>\nA key technique to successfully train this model is to apply attention pooling to the features before inputting them into the transformer and calculate the auxiliary loss.</p>\n<h2>Single-view input, single-condition output model (@yu4u)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F03b00abca721a6790b300042f5879de5%2Fyu4u_pipeline.png?generation=1728433245673485&amp;alt=media\" alt=\"\"></p>\n<p>In this part, severity score is estimated via cropped images.</p>\n<p>For Sagittal T1 and Sagittal T2/STIR images, the cropping scale is determined based on the average distance between the keypoints of the five levels, and patches are cropped centered on the keypoints. For Axial T2 images, an affine transformation is applied to position the left and right keypoints at specific locations within the patch before cropping. Subsequently, the cropped images are input into a 2.5D CNN model to calculate the severity score. For Axial T2 images, the model is used not only to predict subarticular stenosis but also spinal canal stenosis. Other combinations did not yield significant results.</p>\n<h1>Ensemble and stacking</h1>\n<h2>Nelder-Mead guided stacking MLP</h2>\n<p>We constructed a stacking model using an MLP.<br>\nThe key feature of this model is that, in addition to the standard skip connection, the output optimized by the Nelder-Mead method is added to the model output. (In other words, the model learns the difference between the ground truth and the Nelder-Mead results.)<br>\nThe inputs for ss, scs, and any consist only of the results from their respective classification models, while nfn is fed the concatenated outputs of scs, ss, and nfn.</p>\n<h2>Stacking LightGBM and XGBoost</h2>\n<p>The outputs of the individual models are stacked using LightGBM and XGBoost. In this stacking approach, the same model is used separately for each level, and only inputs of the same type as the output target are utilized. As a result, the input dimensions are equal to the number of models multiplied by three. Using predictions for different targets was not effective.</p>\n<h1>Source code and notebooks</h1>\n<h2>source code</h2>\n<ul>\n<li>tattaka's part: <a href=\"https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public\" target=\"_blank\">https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public</a> </li>\n<li>yu4u's part: <a href=\"https://github.com/yu4u/kaggle-rsna2024-4th\" target=\"_blank\">https://github.com/yu4u/kaggle-rsna2024-4th</a></li>\n</ul>\n<h2>notebook</h2>\n<ul>\n<li>best submission notebook: <a href=\"https://www.kaggle.com/code/ren4yu/rsna2024-ensemble-submission-stacking-tattakav2/notebook?scriptVersionId=200046802\" target=\"_blank\">https://www.kaggle.com/code/ren4yu/rsna2024-ensemble-submission-stacking-tattakav2/notebook?scriptVersionId=200046802</a></li>\n<li>MLP stacking for scs: <a href=\"https://www.kaggle.com/code/tattaka/rsna2024-stacking\" target=\"_blank\">https://www.kaggle.com/code/tattaka/rsna2024-stacking</a></li>\n<li>MLP stacking for nfn: <a href=\"https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn\" target=\"_blank\">https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn</a></li>\n<li>MLP stacking for ss: <a href=\"https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn\" target=\"_blank\">https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn</a></li>\n</ul>",
  "messages": [
    {
      "id": 3012347,
      "postDate": "2024-10-09T00:21:51.817Z",
      "content": "<p>Congrats to all prize and medal winners! This year's RSNA competition required us to carefully handle data and build a pipeline, which was a lot of fun. We share our solution.</p>\n<h1>Summary</h1>\n<p>Our solution detects the keypoint, which is the region of interest in the symptom, and builds a classification model using the surrounding crops as input.<br>\nThe results of each model are refined by the stacking model and submitted as the final result.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F11c196e4f15041d980ae018aa1ba5e5c%2Foverall_pipeline.png?generation=1728433161180972&amp;alt=media\" alt=\"\"></p>\n<h1>Keypoint detection Model</h1>\n<h2>Disk Level detection model and Keypoint detection for Axial, Sagittal T1/T2 (@yu4u)</h2>\n<p>Resize each axial slice to 128×128 and use a 2.5DCNN + LSTM model to estimate which level (L1, L2, …, S1) each slice belongs to. Subsequently, detect the boundary slices between each level. From these slices (up to five), use a UNet model to detect the left and right keypoints.<br>\nSimilarly, resize the sagittal T1 slices to 128×128 and use a 2.5D CNN model to identify the left and right slices belonging to the foraminal zone where keypoints should be detected. Then, individually detect keypoints for five levels from these left and right slices.<br>\nFor sagittal T2/STIR, simply extract the middle slice of the series and detect keypoints for the five levels.</p>\n<h2>Keypoint detection model for Sagittal (@tattaka)</h2>\n<p>We resized each of the Sagittal T1, T2/STIR images to 20x256x256 and predicted the xy coordinates of the keypoints.<br>\nThe xy coordinates of the keypoints were taken from the shared <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">Lumbar Coordinate Dataset</a>.<br>\nFor the backbone, we used caformer_s18, convnext_tiny, resnetrs50, and swinv2_tiny, and applied SCSE attention to the UNet Decoder.<br>\nAs for the loss function, we used BCELoss * 0.2 + DICELoss * 0.8.</p>\n<h1>Classification model</h1>\n<h2>Multi-view input, multi-condition output model (@tattaka)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc5d845a64af1ad8c7e9da045b73d3469%2Ftattaka_model.png?generation=1728433227519640&amp;alt=media\" alt=\"\"></p>\n<p>We crop the areas around the keypoints inferred from the volumes of Sagittal T1, Sagittal T2/STIR, and Axial T2, and use them as inputs to classify the conditions at each level.<br>\nEach image is cropped to a size that is twice the distance between the neighboring keypoints. For sagittal images, padding is applied if all slices are fewer than 30, and linear interpolation is used if there are more. (There is also a model variation that simply uses linear interpolation to resize to 20.)<br>\nFor axial images, slices are taken from the range of ±2 around the predicted gaps between the discs.<br>\nAfter the images are input into a 2D model backbone, features are extracted from each slice, and the final output is obtained using a transformer encoder and attention pooling on the extracted features. </p>\n<p>The model used for the final submission includes variations such as:</p>\n<ul>\n<li>A model that processes cropped slices with one or two backbones,</li>\n<li>Multiple augmentation patterns,</li>\n<li>Various preprocessing patterns for multiple slices.</li>\n</ul>\n<p>The backbones used were caformer_s18, resnetrs50, <a href=\"https://github.com/naver-ai/rdnet\" target=\"_blank\">rdnet_tiny</a>, and maxxvitv2_nano.<br>\nA key technique to successfully train this model is to apply attention pooling to the features before inputting them into the transformer and calculate the auxiliary loss.</p>\n<h2>Single-view input, single-condition output model (@yu4u)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F03b00abca721a6790b300042f5879de5%2Fyu4u_pipeline.png?generation=1728433245673485&amp;alt=media\" alt=\"\"></p>\n<p>In this part, severity score is estimated via cropped images.</p>\n<p>For Sagittal T1 and Sagittal T2/STIR images, the cropping scale is determined based on the average distance between the keypoints of the five levels, and patches are cropped centered on the keypoints. For Axial T2 images, an affine transformation is applied to position the left and right keypoints at specific locations within the patch before cropping. Subsequently, the cropped images are input into a 2.5D CNN model to calculate the severity score. For Axial T2 images, the model is used not only to predict subarticular stenosis but also spinal canal stenosis. Other combinations did not yield significant results.</p>\n<h1>Ensemble and stacking</h1>\n<h2>Nelder-Mead guided stacking MLP</h2>\n<p>We constructed a stacking model using an MLP.<br>\nThe key feature of this model is that, in addition to the standard skip connection, the output optimized by the Nelder-Mead method is added to the model output. (In other words, the model learns the difference between the ground truth and the Nelder-Mead results.)<br>\nThe inputs for ss, scs, and any consist only of the results from their respective classification models, while nfn is fed the concatenated outputs of scs, ss, and nfn.</p>\n<h2>Stacking LightGBM and XGBoost</h2>\n<p>The outputs of the individual models are stacked using LightGBM and XGBoost. In this stacking approach, the same model is used separately for each level, and only inputs of the same type as the output target are utilized. As a result, the input dimensions are equal to the number of models multiplied by three. Using predictions for different targets was not effective.</p>\n<h1>Source code and notebooks</h1>\n<h2>source code</h2>\n<ul>\n<li>tattaka's part: <a href=\"https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public\" target=\"_blank\">https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public</a> </li>\n<li>yu4u's part: <a href=\"https://github.com/yu4u/kaggle-rsna2024-4th\" target=\"_blank\">https://github.com/yu4u/kaggle-rsna2024-4th</a></li>\n</ul>\n<h2>notebook</h2>\n<ul>\n<li>best submission notebook: <a href=\"https://www.kaggle.com/code/ren4yu/rsna2024-ensemble-submission-stacking-tattakav2/notebook?scriptVersionId=200046802\" target=\"_blank\">https://www.kaggle.com/code/ren4yu/rsna2024-ensemble-submission-stacking-tattakav2/notebook?scriptVersionId=200046802</a></li>\n<li>MLP stacking for scs: <a href=\"https://www.kaggle.com/code/tattaka/rsna2024-stacking\" target=\"_blank\">https://www.kaggle.com/code/tattaka/rsna2024-stacking</a></li>\n<li>MLP stacking for nfn: <a href=\"https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn\" target=\"_blank\">https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn</a></li>\n<li>MLP stacking for ss: <a href=\"https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn\" target=\"_blank\">https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn</a></li>\n</ul>",
      "rawMarkdown": "Congrats to all prize and medal winners! This year's RSNA competition required us to carefully handle data and build a pipeline, which was a lot of fun. We share our solution.\n\n# Summary\nOur solution detects the keypoint, which is the region of interest in the symptom, and builds a classification model using the surrounding crops as input.\nThe results of each model are refined by the stacking model and submitted as the final result.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F11c196e4f15041d980ae018aa1ba5e5c%2Foverall_pipeline.png?generation=1728433161180972&alt=media)\n\n# Keypoint detection Model\n## Disk Level detection model and Keypoint detection for Axial, Sagittal T1/T2 (@yu4u)\n\nResize each axial slice to 128×128 and use a 2.5DCNN + LSTM model to estimate which level (L1, L2, …, S1) each slice belongs to. Subsequently, detect the boundary slices between each level. From these slices (up to five), use a UNet model to detect the left and right keypoints.\nSimilarly, resize the sagittal T1 slices to 128×128 and use a 2.5D CNN model to identify the left and right slices belonging to the foraminal zone where keypoints should be detected. Then, individually detect keypoints for five levels from these left and right slices.\nFor sagittal T2/STIR, simply extract the middle slice of the series and detect keypoints for the five levels.\n\n\n## Keypoint detection model for Sagittal (@tattaka)\n\nWe resized each of the Sagittal T1, T2/STIR images to 20x256x256 and predicted the xy coordinates of the keypoints.\nThe xy coordinates of the keypoints were taken from the shared [Lumbar Coordinate Dataset](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset).\nFor the backbone, we used caformer_s18, convnext_tiny, resnetrs50, and swinv2_tiny, and applied SCSE attention to the UNet Decoder.\nAs for the loss function, we used BCELoss * 0.2 + DICELoss * 0.8.\n\n# Classification model\n\n## Multi-view input, multi-condition output model (@tattaka)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc5d845a64af1ad8c7e9da045b73d3469%2Ftattaka_model.png?generation=1728433227519640&alt=media)\n\nWe crop the areas around the keypoints inferred from the volumes of Sagittal T1, Sagittal T2/STIR, and Axial T2, and use them as inputs to classify the conditions at each level.\nEach image is cropped to a size that is twice the distance between the neighboring keypoints. For sagittal images, padding is applied if all slices are fewer than 30, and linear interpolation is used if there are more. (There is also a model variation that simply uses linear interpolation to resize to 20.)\nFor axial images, slices are taken from the range of ±2 around the predicted gaps between the discs.\nAfter the images are input into a 2D model backbone, features are extracted from each slice, and the final output is obtained using a transformer encoder and attention pooling on the extracted features. \n\nThe model used for the final submission includes variations such as:\n* A model that processes cropped slices with one or two backbones,\n* Multiple augmentation patterns,\n* Various preprocessing patterns for multiple slices.\n\nThe backbones used were caformer_s18, resnetrs50, [rdnet_tiny](https://github.com/naver-ai/rdnet), and maxxvitv2_nano.\nA key technique to successfully train this model is to apply attention pooling to the features before inputting them into the transformer and calculate the auxiliary loss.\n\n## Single-view input, single-condition output model (@yu4u)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F03b00abca721a6790b300042f5879de5%2Fyu4u_pipeline.png?generation=1728433245673485&alt=media)\n\nIn this part, severity score is estimated via cropped images.\n\nFor Sagittal T1 and Sagittal T2/STIR images, the cropping scale is determined based on the average distance between the keypoints of the five levels, and patches are cropped centered on the keypoints. For Axial T2 images, an affine transformation is applied to position the left and right keypoints at specific locations within the patch before cropping. Subsequently, the cropped images are input into a 2.5D CNN model to calculate the severity score. For Axial T2 images, the model is used not only to predict subarticular stenosis but also spinal canal stenosis. Other combinations did not yield significant results.\n\n\n# Ensemble and stacking\n\n## Nelder-Mead guided stacking MLP\nWe constructed a stacking model using an MLP.\nThe key feature of this model is that, in addition to the standard skip connection, the output optimized by the Nelder-Mead method is added to the model output. (In other words, the model learns the difference between the ground truth and the Nelder-Mead results.)\nThe inputs for ss, scs, and any consist only of the results from their respective classification models, while nfn is fed the concatenated outputs of scs, ss, and nfn.\n\n## Stacking LightGBM and XGBoost\nThe outputs of the individual models are stacked using LightGBM and XGBoost. In this stacking approach, the same model is used separately for each level, and only inputs of the same type as the output target are utilized. As a result, the input dimensions are equal to the number of models multiplied by three. Using predictions for different targets was not effective.\n\n# Source code and notebooks\n## source code\n* tattaka's part: https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public \n* yu4u's part: https://github.com/yu4u/kaggle-rsna2024-4th\n\n## notebook\n* best submission notebook: https://www.kaggle.com/code/ren4yu/rsna2024-ensemble-submission-stacking-tattakav2/notebook?scriptVersionId=200046802\n* MLP stacking for scs: https://www.kaggle.com/code/tattaka/rsna2024-stacking\n* MLP stacking for nfn: https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn\n* MLP stacking for ss: https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn",
      "votes": 76
    },
    {
      "id": 3017635,
      "postDate": "2024-10-15T04:40:25.970Z",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a><br>\nCongrats for prize and Share your solution ! 🎉</p>\n<p>We can't have access to your MLP stacking codes. Please check access setting 🙏</p>",
      "rawMarkdown": "@tattaka @ren4yu\nCongrats for prize and Share your solution ! 🎉\n\nWe can't have access to your MLP stacking codes. Please check access setting 🙏",
      "votes": 1,
      "replies": [
        {
          "id": 3018031,
          "postDate": "2024-10-15T13:16:43.533Z",
          "content": "<p>Thank you, I have changed the permissions of the notebook for MLP stacking to public.</p>",
          "rawMarkdown": "Thank you, I have changed the permissions of the notebook for MLP stacking to public.\n"
        }
      ]
    },
    {
      "id": 3012740,
      "postDate": "2024-10-09T09:57:27.453Z",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring 🧠🧠🧠</p>\n<p>How hard is it to make your multi-view, multi-condition architecture up and running? How many iterations on the architecture/experiments worth of time does it take? And do you recommend less experienced kagglers to try and construct something similar or do you think it requires much more intuiton about the cv domain and thorough knowledge of valid computer vision achitectures and it would be better to focus on careful data processing and other basic things rather than assembling such \"monsters\"?</p>",
      "rawMarkdown": "@tattaka Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring 🧠🧠🧠\n\nHow hard is it to make your multi-view, multi-condition architecture up and running? How many iterations on the architecture/experiments worth of time does it take? And do you recommend less experienced kagglers to try and construct something similar or do you think it requires much more intuiton about the cv domain and thorough knowledge of valid computer vision achitectures and it would be better to focus on careful data processing and other basic things rather than assembling such \"monsters\"?",
      "votes": 2,
      "replies": [
        {
          "id": 3013378,
          "postDate": "2024-10-10T02:38:57.680Z",
          "content": "<blockquote>\n  <p>How hard is it to make your multi-view, multi-condition architecture up and running?   </p>\n</blockquote>\n<p>In my setup, the model was sensitive to hyperparameters and did not work with all backbones.</p>\n<blockquote>\n  <p>How many iterations on the architecture/experiments worth of time does it take?</p>\n</blockquote>\n<p>Experiments with multi-view architectures took at least 6 hours to complete a 5-fold experiment. Most of the time, I validated the effectiveness using only 1 fold.</p>\n<blockquote>\n  <p>And do you recommend less experienced kagglers to try and construct something similar or do you think it requires much more intuiton about the cv domain and thorough knowledge of valid computer vision achitectures and it would be better to focus on careful data processing and other basic things rather than assembling such \"monsters\"?</p>\n</blockquote>\n<p>This is just my personal opinion, but I recommend starting by building a simple model or pipeline and checking if it works. You can then use this pipeline as a starting point for experiments. By testing hypotheses and exploring things that go against your intuition, you can improve the pipeline.</p>",
          "rawMarkdown": "> How hard is it to make your multi-view, multi-condition architecture up and running?   \n\nIn my setup, the model was sensitive to hyperparameters and did not work with all backbones.\n\n> How many iterations on the architecture/experiments worth of time does it take?\n\nExperiments with multi-view architectures took at least 6 hours to complete a 5-fold experiment. Most of the time, I validated the effectiveness using only 1 fold.\n\n> And do you recommend less experienced kagglers to try and construct something similar or do you think it requires much more intuiton about the cv domain and thorough knowledge of valid computer vision achitectures and it would be better to focus on careful data processing and other basic things rather than assembling such \"monsters\"?\n\nThis is just my personal opinion, but I recommend starting by building a simple model or pipeline and checking if it works. You can then use this pipeline as a starting point for experiments. By testing hypotheses and exploring things that go against your intuition, you can improve the pipeline.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3012358,
      "postDate": "2024-10-09T00:43:42.013Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>, the solution is very elegant, I liked your use of so many heads for study-level predictions to ensemble the approaches, what was your OOF CV and did it follow the LB well?</p>",
      "rawMarkdown": "Congratulations @tattaka @ren4yu, the solution is very elegant, I liked your use of so many heads for study-level predictions to ensemble the approaches, what was your OOF CV and did it follow the LB well?",
      "votes": 2,
      "replies": [
        {
          "id": 3012364,
          "postDate": "2024-10-09T00:59:36.560Z",
          "content": "<p>The final CV scores were:</p>\n<pre><code> score: .\n score: .\n score: .\n score: .\n</code></pre>\n<p>In our solution, the correlation between CV and LB was good.</p>",
          "rawMarkdown": "The final CV scores were:\n```\nss score: 0.53005\nforaminal score: 0.47115\nscs score: 0.24895\nany score: 0.24815\n```\nIn our solution, the correlation between CV and LB was good.",
          "votes": 3,
          "replies": [
            {
              "id": 3012369,
              "postDate": "2024-10-09T01:16:09.010Z",
              "content": "<p>My CV is like this (mine alone, ensembled is 0.380):</p>\n<pre><code>(,\n {: ,\n  : ,\n  : ,\n  : })\n</code></pre>\n<p>Could you double check my OOF submission file by calculating the score on your function, thank you!</p>\n<p>For us 0.446 CV was 0.39 LB and 0.380 CV was 0.4 LB</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna-2024-oof\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna-2024-oof</a></p>",
              "rawMarkdown": "My CV is like this (mine alone, ensembled is 0.380):\n\n```python\n(0.3866988309688447,\n {'spinal': 0.25158152964121944,\n  'foraminal': 0.5008341775129096,\n  'subarticular': 0.5491287326102894,\n  'spinal_severe': 0.24525088411096035})\n```\n\nCould you double check my OOF submission file by calculating the score on your function, thank you!\n\nFor us 0.446 CV was 0.39 LB and 0.380 CV was 0.4 LB\n\nhttps://www.kaggle.com/datasets/harshitsheoran/rsna-2024-oof\n"
            },
            {
              "id": 3012375,
              "postDate": "2024-10-09T01:38:20.573Z",
              "content": "<p>When you calculate your score using <a href=\"https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public/blob/main/src/stage2/exp107/scoring.ipynb\" target=\"_blank\">our notebook</a> , you get</p>\n<pre><code>\"Score (0.38779225210879453, {: 0.251212030446364, : 0.4966134331970362, : 0.5588322928130882, : 0.24451125197868967})\".\n</code></pre>\n<p>There are some slight differences, but it seems that the oof score is not incorrect.</p>",
              "rawMarkdown": "When you calculate your score using [our notebook](https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public/blob/main/src/stage2/exp107/scoring.ipynb) , you get\n```\n\"Score (0.38779225210879453, {'spinal': 0.251212030446364, 'foraminal': 0.4966134331970362, 'subarticular': 0.5588322928130882, 'any': 0.24451125197868967})\".\n```\nThere are some slight differences, but it seems that the oof score is not incorrect.\n\n\n\n\n\n\n"
            },
            {
              "id": 3012377,
              "postDate": "2024-10-09T01:41:55.373Z",
              "content": "<p>Thank you, so, it was just us missing something huge on subarticular, and LB not following the rest, its stupidly frustrating</p>",
              "rawMarkdown": "Thank you, so, it was just us missing something huge on subarticular, and LB not following the rest, its stupidly frustrating"
            },
            {
              "id": 3012403,
              "postDate": "2024-10-09T02:38:17.643Z",
              "content": "<p>My team single model CV: 0.4, subarticular: 0.57, spinal: 0.27, foraminal: 0.51 🥲</p>",
              "rawMarkdown": "My team single model CV: 0.4, subarticular: 0.57, spinal: 0.27, foraminal: 0.51 🥲",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3035188,
      "postDate": "2024-11-03T05:37:50.660Z",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> may I ask with the classification and keypoint detection models, do you use pre-trained weights in the backbones? or do you train from scratch?</p>\n<p>What is the typical approach when training models in a different domain, i.e. medical data MRIs, when the weights were trained on ImageNet?</p>",
      "rawMarkdown": "@tattaka @ren4yu may I ask with the classification and keypoint detection models, do you use pre-trained weights in the backbones? or do you train from scratch?\n\nWhat is the typical approach when training models in a different domain, i.e. medical data MRIs, when the weights were trained on ImageNet?"
    },
    {
      "id": 3019908,
      "postDate": "2024-10-17T02:40:53.603Z",
      "content": "<p>Congratulations on this impressive development! I would like to know if you think that using a classic machine learning model could have achieved similar performance to your solution, and in which cases you consider it better to use traditional machine learning models versus deep learning models.</p>",
      "rawMarkdown": "Congratulations on this impressive development! I would like to know if you think that using a classic machine learning model could have achieved similar performance to your solution, and in which cases you consider it better to use traditional machine learning models versus deep learning models."
    },
    {
      "id": 3015234,
      "postDate": "2024-10-12T05:41:13.467Z",
      "content": "<p>Great! Congratulations. I haven't participated in any competition yet. Can you please share some tips so that I can do well?</p>",
      "rawMarkdown": "Great! Congratulations. I haven't participated in any competition yet. Can you please share some tips so that I can do well?"
    },
    {
      "id": 3015047,
      "postDate": "2024-10-11T23:08:38.780Z",
      "content": "<p>Interesting use of \"Nelder-Mead guided stacking\". Did you compare the performance between this method and a single MLP head?</p>",
      "rawMarkdown": "Interesting use of \"Nelder-Mead guided stacking\". Did you compare the performance between this method and a single MLP head?",
      "replies": [
        {
          "id": 3015107,
          "postDate": "2024-10-12T02:26:00.710Z",
          "content": "<p>It deteriorates by 0.01 compared to a single MLP head (worse than the simple Nelder-Mead).</p>",
          "rawMarkdown": "It deteriorates by 0.01 compared to a single MLP head (worse than the simple Nelder-Mead).",
          "votes": 1
        }
      ]
    },
    {
      "id": 3014607,
      "postDate": "2024-10-11T12:15:08.997Z",
      "content": "<p>Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring !!!!</p>",
      "rawMarkdown": "Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring !!!!"
    },
    {
      "id": 3013050,
      "postDate": "2024-10-09T15:59:56.417Z",
      "content": "<p>Congratulations on winning the 4th prize in this competition. Thanks for sharing the explanation about your approach with good diagrams.</p>",
      "rawMarkdown": "Congratulations on winning the 4th prize in this competition. Thanks for sharing the explanation about your approach with good diagrams."
    },
    {
      "id": 3016527,
      "postDate": "2024-10-13T22:51:23.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3013127,
      "postDate": "2024-10-09T17:06:31.767Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3012381,
      "postDate": "2024-10-09T01:54:00.953Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3012630,
      "postDate": "2024-10-09T08:15:10.367Z",
      "content": "<p>thank you so much  for sharing</p>",
      "rawMarkdown": "thank you so much  for sharing"
    },
    {
      "id": 3012580,
      "postDate": "2024-10-09T07:35:17.883Z",
      "content": "<p>Thank you for sharing.Helpful</p>",
      "rawMarkdown": "Thank you for sharing.Helpful"
    }
  ],
  "comments": [
    {
      "id": 3017635,
      "author_name": "kaerururu",
      "author_url": "",
      "post_date": "2024-10-15T04:40:25.970000",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a><br>\nCongrats for prize and Share your solution ! 🎉</p>\n<p>We can't have access to your MLP stacking codes. Please check access setting 🙏</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3018031,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2024-10-15T13:16:43.533000",
          "content": "<p>Thank you, I have changed the permissions of the notebook for MLP stacking to public.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3012740,
      "author_name": "Ivan Vybornov",
      "author_url": "",
      "post_date": "2024-10-09T09:57:27.453000",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring 🧠🧠🧠</p>\n<p>How hard is it to make your multi-view, multi-condition architecture up and running? How many iterations on the architecture/experiments worth of time does it take? And do you recommend less experienced kagglers to try and construct something similar or do you think it requires much more intuiton about the cv domain and thorough knowledge of valid computer vision achitectures and it would be better to focus on careful data processing and other basic things rather than assembling such \"monsters\"?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3013378,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2024-10-10T02:38:57.680000",
          "content": "<blockquote>\n  <p>How hard is it to make your multi-view, multi-condition architecture up and running?   </p>\n</blockquote>\n<p>In my setup, the model was sensitive to hyperparameters and did not work with all backbones.</p>\n<blockquote>\n  <p>How many iterations on the architecture/experiments worth of time does it take?</p>\n</blockquote>\n<p>Experiments with multi-view architectures took at least 6 hours to complete a 5-fold experiment. Most of the time, I validated the effectiveness using only 1 fold.</p>\n<blockquote>\n  <p>And do you recommend less experienced kagglers to try and construct something similar or do you think it requires much more intuiton about the cv domain and thorough knowledge of valid computer vision achitectures and it would be better to focus on careful data processing and other basic things rather than assembling such \"monsters\"?</p>\n</blockquote>\n<p>This is just my personal opinion, but I recommend starting by building a simple model or pipeline and checking if it works. You can then use this pipeline as a starting point for experiments. By testing hypotheses and exploring things that go against your intuition, you can improve the pipeline.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3012358,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2024-10-09T00:43:42.013000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>, the solution is very elegant, I liked your use of so many heads for study-level predictions to ensemble the approaches, what was your OOF CV and did it follow the LB well?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3012364,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2024-10-09T00:59:36.560000",
          "content": "<p>The final CV scores were:</p>\n<pre><code> score: .\n score: .\n score: .\n score: .\n</code></pre>\n<p>In our solution, the correlation between CV and LB was good.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3012369,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2024-10-09T01:16:09.010000",
              "content": "<p>My CV is like this (mine alone, ensembled is 0.380):</p>\n<pre><code>(,\n {: ,\n  : ,\n  : ,\n  : })\n</code></pre>\n<p>Could you double check my OOF submission file by calculating the score on your function, thank you!</p>\n<p>For us 0.446 CV was 0.39 LB and 0.380 CV was 0.4 LB</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna-2024-oof\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna-2024-oof</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3012375,
              "author_name": "tattaka",
              "author_url": "",
              "post_date": "2024-10-09T01:38:20.573000",
              "content": "<p>When you calculate your score using <a href=\"https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public/blob/main/src/stage2/exp107/scoring.ipynb\" target=\"_blank\">our notebook</a> , you get</p>\n<pre><code>\"Score (0.38779225210879453, {: 0.251212030446364, : 0.4966134331970362, : 0.5588322928130882, : 0.24451125197868967})\".\n</code></pre>\n<p>There are some slight differences, but it seems that the oof score is not incorrect.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3012377,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2024-10-09T01:41:55.373000",
              "content": "<p>Thank you, so, it was just us missing something huge on subarticular, and LB not following the rest, its stupidly frustrating</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3012403,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2024-10-09T02:38:17.643000",
              "content": "<p>My team single model CV: 0.4, subarticular: 0.57, spinal: 0.27, foraminal: 0.51 🥲</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3035188,
      "author_name": "homiecal",
      "author_url": "",
      "post_date": "2024-11-03T05:37:50.660000",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> may I ask with the classification and keypoint detection models, do you use pre-trained weights in the backbones? or do you train from scratch?</p>\n<p>What is the typical approach when training models in a different domain, i.e. medical data MRIs, when the weights were trained on ImageNet?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3019908,
      "author_name": "Darwin Berrio",
      "author_url": "",
      "post_date": "2024-10-17T02:40:53.603000",
      "content": "<p>Congratulations on this impressive development! I would like to know if you think that using a classic machine learning model could have achieved similar performance to your solution, and in which cases you consider it better to use traditional machine learning models versus deep learning models.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3015234,
      "author_name": "Humayra Khanom Rime",
      "author_url": "",
      "post_date": "2024-10-12T05:41:13.467000",
      "content": "<p>Great! Congratulations. I haven't participated in any competition yet. Can you please share some tips so that I can do well?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3015047,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-10-11T23:08:38.780000",
      "content": "<p>Interesting use of \"Nelder-Mead guided stacking\". Did you compare the performance between this method and a single MLP head?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3015107,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2024-10-12T02:26:00.710000",
          "content": "<p>It deteriorates by 0.01 compared to a single MLP head (worse than the simple Nelder-Mead).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3014607,
      "author_name": "MrSimple",
      "author_url": "",
      "post_date": "2024-10-11T12:15:08.997000",
      "content": "<p>Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring !!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3013050,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-10-09T15:59:56.417000",
      "content": "<p>Congratulations on winning the 4th prize in this competition. Thanks for sharing the explanation about your approach with good diagrams.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3016527,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-13T22:51:23.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3013127,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-09T17:06:31.767000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3012381,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-09T01:54:00.953000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3012630,
      "author_name": "Abdullah Al-Shami",
      "author_url": "",
      "post_date": "2024-10-09T08:15:10.367000",
      "content": "<p>thank you so much  for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3012580,
      "author_name": "Deepak",
      "author_url": "",
      "post_date": "2024-10-09T07:35:17.883000",
      "content": "<p>Thank you for sharing.Helpful</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3012347": "Congrats to all prize and medal winners! This year's RSNA competition required us to carefully handle data and build a pipeline, which was a lot of fun. We share our solution.\n\n# Summary\nOur solution detects the keypoint, which is the region of interest in the symptom, and builds a classification model using the surrounding crops as input.\nThe results of each model are refined by the stacking model and submitted as the final result.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F11c196e4f15041d980ae018aa1ba5e5c%2Foverall_pipeline.png?generation=1728433161180972&alt=media)\n\n# Keypoint detection Model\n## Disk Level detection model and Keypoint detection for Axial, Sagittal T1/T2 (@yu4u)\n\nResize each axial slice to 128×128 and use a 2.5DCNN + LSTM model to estimate which level (L1, L2, …, S1) each slice belongs to. Subsequently, detect the boundary slices between each level. From these slices (up to five), use a UNet model to detect the left and right keypoints.\nSimilarly, resize the sagittal T1 slices to 128×128 and use a 2.5D CNN model to identify the left and right slices belonging to the foraminal zone where keypoints should be detected. Then, individually detect keypoints for five levels from these left and right slices.\nFor sagittal T2/STIR, simply extract the middle slice of the series and detect keypoints for the five levels.\n\n\n## Keypoint detection model for Sagittal (@tattaka)\n\nWe resized each of the Sagittal T1, T2/STIR images to 20x256x256 and predicted the xy coordinates of the keypoints.\nThe xy coordinates of the keypoints were taken from the shared [Lumbar Coordinate Dataset](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset).\nFor the backbone, we used caformer_s18, convnext_tiny, resnetrs50, and swinv2_tiny, and applied SCSE attention to the UNet Decoder.\nAs for the loss function, we used BCELoss * 0.2 + DICELoss * 0.8.\n\n# Classification model\n\n## Multi-view input, multi-condition output model (@tattaka)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fc5d845a64af1ad8c7e9da045b73d3469%2Ftattaka_model.png?generation=1728433227519640&alt=media)\n\nWe crop the areas around the keypoints inferred from the volumes of Sagittal T1, Sagittal T2/STIR, and Axial T2, and use them as inputs to classify the conditions at each level.\nEach image is cropped to a size that is twice the distance between the neighboring keypoints. For sagittal images, padding is applied if all slices are fewer than 30, and linear interpolation is used if there are more. (There is also a model variation that simply uses linear interpolation to resize to 20.)\nFor axial images, slices are taken from the range of ±2 around the predicted gaps between the discs.\nAfter the images are input into a 2D model backbone, features are extracted from each slice, and the final output is obtained using a transformer encoder and attention pooling on the extracted features. \n\nThe model used for the final submission includes variations such as:\n* A model that processes cropped slices with one or two backbones,\n* Multiple augmentation patterns,\n* Various preprocessing patterns for multiple slices.\n\nThe backbones used were caformer_s18, resnetrs50, [rdnet_tiny](https://github.com/naver-ai/rdnet), and maxxvitv2_nano.\nA key technique to successfully train this model is to apply attention pooling to the features before inputting them into the transformer and calculate the auxiliary loss.\n\n## Single-view input, single-condition output model (@yu4u)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F03b00abca721a6790b300042f5879de5%2Fyu4u_pipeline.png?generation=1728433245673485&alt=media)\n\nIn this part, severity score is estimated via cropped images.\n\nFor Sagittal T1 and Sagittal T2/STIR images, the cropping scale is determined based on the average distance between the keypoints of the five levels, and patches are cropped centered on the keypoints. For Axial T2 images, an affine transformation is applied to position the left and right keypoints at specific locations within the patch before cropping. Subsequently, the cropped images are input into a 2.5D CNN model to calculate the severity score. For Axial T2 images, the model is used not only to predict subarticular stenosis but also spinal canal stenosis. Other combinations did not yield significant results.\n\n\n# Ensemble and stacking\n\n## Nelder-Mead guided stacking MLP\nWe constructed a stacking model using an MLP.\nThe key feature of this model is that, in addition to the standard skip connection, the output optimized by the Nelder-Mead method is added to the model output. (In other words, the model learns the difference between the ground truth and the Nelder-Mead results.)\nThe inputs for ss, scs, and any consist only of the results from their respective classification models, while nfn is fed the concatenated outputs of scs, ss, and nfn.\n\n## Stacking LightGBM and XGBoost\nThe outputs of the individual models are stacked using LightGBM and XGBoost. In this stacking approach, the same model is used separately for each level, and only inputs of the same type as the output target are utilized. As a result, the input dimensions are equal to the number of models multiplied by three. Using predictions for different targets was not effective.\n\n# Source code and notebooks\n## source code\n* tattaka's part: https://github.com/tattaka/rsna-2024-lumbar-spine-degenerative-classification-public \n* yu4u's part: https://github.com/yu4u/kaggle-rsna2024-4th\n\n## notebook\n* best submission notebook: https://www.kaggle.com/code/ren4yu/rsna2024-ensemble-submission-stacking-tattakav2/notebook?scriptVersionId=200046802\n* MLP stacking for scs: https://www.kaggle.com/code/tattaka/rsna2024-stacking\n* MLP stacking for nfn: https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn\n* MLP stacking for ss: https://www.kaggle.com/code/tattaka/rsna2024-stacking-nfn",
    "3017635": "@tattaka @ren4yu\nCongrats for prize and Share your solution ! 🎉\n\nWe can't have access to your MLP stacking codes. Please check access setting 🙏",
    "3012740": "@tattaka Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring 🧠🧠🧠\n\nHow hard is it to make your multi-view, multi-condition architecture up and running? How many iterations on the architecture/experiments worth of time does it take? And do you recommend less experienced kagglers to try and construct something similar or do you think it requires much more intuiton about the cv domain and thorough knowledge of valid computer vision achitectures and it would be better to focus on careful data processing and other basic things rather than assembling such \"monsters\"?",
    "3012358": "Congratulations @tattaka @ren4yu, the solution is very elegant, I liked your use of so many heads for study-level predictions to ensemble the approaches, what was your OOF CV and did it follow the LB well?",
    "3035188": "@tattaka @ren4yu may I ask with the classification and keypoint detection models, do you use pre-trained weights in the backbones? or do you train from scratch?\n\nWhat is the typical approach when training models in a different domain, i.e. medical data MRIs, when the weights were trained on ImageNet?",
    "3019908": "Congratulations on this impressive development! I would like to know if you think that using a classic machine learning model could have achieved similar performance to your solution, and in which cases you consider it better to use traditional machine learning models versus deep learning models.",
    "3015234": "Great! Congratulations. I haven't participated in any competition yet. Can you please share some tips so that I can do well?",
    "3015047": "Interesting use of \"Nelder-Mead guided stacking\". Did you compare the performance between this method and a single MLP head?",
    "3014607": "Thanks for sharing! Your Multi-view input, multi-condition output is truly inspiring !!!!",
    "3013050": "Congratulations on winning the 4th prize in this competition. Thanks for sharing the explanation about your approach with good diagrams.",
    "3016527": "",
    "3013127": "",
    "3012381": "",
    "3012630": "thank you so much  for sharing",
    "3012580": "Thank you for sharing.Helpful"
  }
}