{
  "id": 533240,
  "title": "Image-level object detection approach with Yolo [LB 0.54]",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/533240",
  "author_name": "",
  "post_date": "2024-09-10T07:23:37.785118300Z",
  "votes": 50,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Most of the public notebooks have approached this competition by building classification models on sequences of images and making study-level predictions directly. </p>\n<p>I'd like to propose a bottom up approach by building object detection models for each modality to predict bounding boxes for level and condition at the same time. For example, a neural_foraminal_narrowing model learns to predict 5 levels * 2 positions (left/right) * 3 conditions = 30 labels. Then, a simple aggregation method (taking the maximum within each \"condition_level\" ) is used to combine the results from images to obtain study-level predictions. </p>\n<p>Evaluation and Submission notebook: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-yolo-approach/notebook\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-yolo-approach/notebook</a></p>\n<p>Training notebooks:<br>\nSCS: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-scs/notebook\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-scs/notebook</a><br>\nNFN: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-nfn\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-nfn</a><br>\nSS: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-ss\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-ss</a> </p>",
  "messages": [
    {
      "id": "2984973",
      "postDate": "09/10/2024 07:23:37",
      "content": "<p>Most of the public notebooks have approached this competition by building classification models on sequences of images and making study-level predictions directly. </p>\n<p>I'd like to propose a bottom up approach by building object detection models for each modality to predict bounding boxes for level and condition at the same time. For example, a neural_foraminal_narrowing model learns to predict 5 levels * 2 positions (left/right) * 3 conditions = 30 labels. Then, a simple aggregation method (taking the maximum within each \"condition_level\" ) is used to combine the results from images to obtain study-level predictions. </p>\n<p>Evaluation and Submission notebook: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-yolo-approach/notebook\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-yolo-approach/notebook</a></p>\n<p>Training notebooks:<br>\nSCS: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-scs/notebook\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-scs/notebook</a><br>\nNFN: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-nfn\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-nfn</a><br>\nSS: <a href=\"https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-ss\" target=\"_blank\">https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-ss</a> </p>",
      "rawMarkdown": "Most of the public notebooks have approached this competition by building classification models on sequences of images and making study-level predictions directly. \n\nI'd like to propose a bottom up approach by building object detection models for each modality to predict bounding boxes for level and condition at the same time. For example, a neural_foraminal_narrowing model learns to predict 5 levels * 2 positions (left/right) * 3 conditions = 30 labels. Then, a simple aggregation method (taking the maximum within each \"condition_level\" ) is used to combine the results from images to obtain study-level predictions. \n\nEvaluation and Submission notebook: https://www.kaggle.com/code/namgalielei/lsdc-yolo-approach/notebook\n\nTraining notebooks:\nSCS: https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-scs/notebook\nNFN: https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-nfn\nSS: https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-ss",
      "votes": null
    },
    {
      "id": "2985038",
      "postDate": "09/10/2024 08:44:33",
      "content": "<p>For the future. Be careful with <code>ultralytics</code> <strong>AGPL-3.0 license</strong>! As they can essentially assert open-source requirements for any project that incorporates their code.</p>\n<p>You can use 'original' YOLO implementation by WongKinYiu without any risk</p>",
      "rawMarkdown": "For the future. Be careful with `ultralytics` **AGPL-3.0 license**! As they can essentially assert open-source requirements for any project that incorporates their code.\n\nYou can use 'original' YOLO implementation by WongKinYiu without any risk",
      "votes": null
    },
    {
      "id": "2985060",
      "postDate": "09/10/2024 09:24:44",
      "content": "<p>Is it usable in this comp<br>\n? </p>",
      "rawMarkdown": "Is it usable in this comp\n?",
      "votes": null
    },
    {
      "id": "2985068",
      "postDate": "09/10/2024 09:34:13",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "2985077",
      "postDate": "09/10/2024 09:45:25",
      "content": "<p>Thank you so very much for sharing this <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a>. </p>\n<p>I wanted to ask you: </p>\n<p>During training, are you storing the bounding box annotated images in the working directory or passing its coordinates to the model<br>\nOr is the model figuring out the bounding boxes by itself? </p>",
      "rawMarkdown": "Thank you so very much for sharing this @namgalielei. \n\nI wanted to ask you: \n\nDuring training, are you storing the bounding box annotated images in the working directory or passing its coordinates to the model\nOr is the model figuring out the bounding boxes by itself?",
      "votes": null
    },
    {
      "id": "2985105",
      "postDate": "09/10/2024 10:31:06",
      "content": "<p>Apologies. I understood how it is being passed. I leave an explanation for anyone confused: </p>\n<ul>\n<li>each data fold file has a label and image file. The label consists a pin point of the anamoly in the MRI. This is being fed to the model. </li>\n</ul>",
      "rawMarkdown": "Apologies. I understood how it is being passed. I leave an explanation for anyone confused: \n\n- each data fold file has a label and image file. The label consists a pin point of the anamoly in the MRI. This is being fed to the model.",
      "votes": null
    },
    {
      "id": "2985282",
      "postDate": "09/10/2024 13:46:43",
      "content": "<p>Thanks for your sharing!</p>",
      "rawMarkdown": "Thanks for your sharing!",
      "votes": null
    },
    {
      "id": "2985389",
      "postDate": "09/10/2024 15:18:16",
      "content": "<p>Thanks for sharing. Just a quick question, did you do some preprocessing of training data， in order to use YOLO? </p>",
      "rawMarkdown": "Thanks for sharing. Just a quick question, did you do some preprocessing of training data， in order to use YOLO?",
      "votes": null
    },
    {
      "id": "2985473",
      "postDate": "09/10/2024 17:00:42",
      "content": "<p>Part of the winning requirements is making your stuff public anyway. Not just the code but your process etc.</p>\n<p>So it should not matter</p>",
      "rawMarkdown": "Part of the winning requirements is making your stuff public anyway. Not just the code but your process etc.\n\nSo it should not matter",
      "votes": null
    },
    {
      "id": "2985474",
      "postDate": "09/10/2024 17:02:47",
      "content": "<p>Good work!</p>",
      "rawMarkdown": "Good work!",
      "votes": null
    },
    {
      "id": "2985731",
      "postDate": "09/11/2024 00:23:50",
      "content": "<p>I only remove study with null labels</p>",
      "rawMarkdown": "I only remove study with null labels",
      "votes": null
    },
    {
      "id": "2986066",
      "postDate": "09/11/2024 10:33:45",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> !  </p>\n<p>I wanted to ask you, why did you choose to train separate models for each condition instead of view (t1 sag,t2 sag, axial)? Is there any particular reason for  this design choice? What do you think of one model per view ? My intuition is YOLO might train better on similar looking images. Do let me know your thoughts if possible.</p>",
      "rawMarkdown": "Hey @namgalielei !  \n\nI wanted to ask you, why did you choose to train separate models for each condition instead of view (t1 sag,t2 sag, axial)? Is there any particular reason for  this design choice? What do you think of one model per view ? My intuition is YOLO might train better on similar looking images. Do let me know your thoughts if possible.",
      "votes": null
    },
    {
      "id": "2986166",
      "postDate": "09/11/2024 12:43:35",
      "content": "<p>Actually the view (modality) corresponds to conditions: t1 sag -&gt; nfn, axial t2 -&gt; ss, t2 sag -&gt; scs</p>",
      "rawMarkdown": "Actually the view (modality) corresponds to conditions: t1 sag -> nfn, axial t2 -> ss, t2 sag -> scs",
      "votes": null
    },
    {
      "id": "2986193",
      "postDate": "09/11/2024 13:23:40",
      "content": "<p>Thanks for clarrification!</p>",
      "rawMarkdown": "Thanks for clarrification!",
      "votes": null
    },
    {
      "id": "2986543",
      "postDate": "09/11/2024 18:44:16",
      "content": "<p>As far as I know, you can use the Yolo V8 if you do not plan to receive a cash prize. If the license limits commercial use, then the cash prize will not be paid, but they will give a medal. Of course, I can understand something wrong. Have a good day!</p>",
      "rawMarkdown": "As far as I know, you can use the Yolo V8 if you do not plan to receive a cash prize. If the license limits commercial use, then the cash prize will not be paid, but they will give a medal. Of course, I can understand something wrong. Have a good day!",
      "votes": null
    },
    {
      "id": "2987497",
      "postDate": "09/12/2024 17:58:13",
      "content": "<p>Thanks for sharing .</p>\n<p>How do you define the size of box around that area? <br>\nHow do you know which size is  fine ?</p>",
      "rawMarkdown": "Thanks for sharing .\n\nHow do you define the size of box around that area? \nHow do you know which size is  fine ?",
      "votes": null
    },
    {
      "id": "2988266",
      "postDate": "09/13/2024 15:25:00",
      "content": "<p>Thank you for sharing your work. I have found that some models have too many categories and an imbalance of categories, resulting in almost no detection results for some classes during inference. Even if I train each model for 5 folds and lower the detection threshold, I cannot solve this problem.</p>",
      "rawMarkdown": "Thank you for sharing your work. I have found that some models have too many categories and an imbalance of categories, resulting in almost no detection results for some classes during inference. Even if I train each model for 5 folds and lower the detection threshold, I cannot solve this problem.",
      "votes": null
    },
    {
      "id": "2990164",
      "postDate": "09/16/2024 04:03:34",
      "content": "<p>Just define a fixed small size relative to the image size. It does not matter much because we only care about the whether the findings appear on these images.</p>",
      "rawMarkdown": "Just define a fixed small size relative to the image size. It does not matter much because we only care about the whether the findings appear on these images.",
      "votes": null
    },
    {
      "id": "2990226",
      "postDate": "09/16/2024 05:45:21",
      "content": "<p>This suggestion reflects a great deal of thought in training separate models for different modalities, and the combination with a straightforward aggregation method makes it an elegant solution. Thank you for sharing your insights.</p>",
      "rawMarkdown": "This suggestion reflects a great deal of thought in training separate models for different modalities, and the combination with a straightforward aggregation method makes it an elegant solution. Thank you for sharing your insights.",
      "votes": null
    },
    {
      "id": "2992089",
      "postDate": "09/18/2024 07:01:31",
      "content": "<p>Thanks for your sharing.</p>",
      "rawMarkdown": "Thanks for your sharing.",
      "votes": null
    },
    {
      "id": "2997846",
      "postDate": "09/25/2024 00:45:38",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2997888",
      "postDate": "09/25/2024 02:34:17",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "3008810",
      "postDate": "10/07/2024 06:07:52",
      "content": "<p>I have a problem, In a 3D volume I detect 3 L5/S1, 2 is normal, 1 is severe. So, the predicted label is expected to be either normal or severe?</p>",
      "rawMarkdown": "I have a problem, In a 3D volume I detect 3 L5/S1, 2 is normal, 1 is severe. So, the predicted label is expected to be either normal or severe?",
      "votes": null
    },
    {
      "id": "3008900",
      "postDate": "10/07/2024 08:56:43",
      "content": "<p>Ideally, if there is \"severe\" finding at any slice, then the whole sequence should be classified as \"severe\". However, these Yolo models I've trained don't have such reliability, and the label distribution is highly unbalanced towards \"normal/mild\", in this case I simply take the max probability within each severity across all the slices, then normalize by dividing by the sum. You might find a better aggregation technique</p>",
      "rawMarkdown": "Ideally, if there is \"severe\" finding at any slice, then the whole sequence should be classified as \"severe\". However, these Yolo models I've trained don't have such reliability, and the label distribution is highly unbalanced towards \"normal/mild\", in this case I simply take the max probability within each severity across all the slices, then normalize by dividing by the sum. You might find a better aggregation technique",
      "votes": null
    },
    {
      "id": "3008903",
      "postDate": "10/07/2024 09:06:29",
      "content": "<p>I have experiment with end-to-end and two stage method. I think an end-to-end approach is necessary for this problem, but the drawback is a lack of data, leading to overfitting very quickly. A two-stage method requires rule configuration and with a 38% LB, it's also prone to overfitting. Additionally, two-stage methods require a deep understanding of the problem to manually configure rules.</p>",
      "rawMarkdown": "I have experiment with end-to-end and two stage method. I think an end-to-end approach is necessary for this problem, but the drawback is a lack of data, leading to overfitting very quickly. A two-stage method requires rule configuration and with a 38% LB, it's also prone to overfitting. Additionally, two-stage methods require a deep understanding of the problem to manually configure rules.",
      "votes": null
    },
    {
      "id": "3012707",
      "postDate": "10/09/2024 09:11:26",
      "content": "<p>Can you share your final contest plan，Thank you very much！😁</p>",
      "rawMarkdown": "Can you share your final contest plan，Thank you very much！😁",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2985038,
      "author_name": "samson8",
      "author_url": "",
      "post_date": "09/10/2024 08:44:33",
      "content": "<p>For the future. Be careful with <code>ultralytics</code> <strong>AGPL-3.0 license</strong>! As they can essentially assert open-source requirements for any project that incorporates their code.</p>\n<p>You can use 'original' YOLO implementation by WongKinYiu without any risk</p>",
      "votes": null,
      "replies": [
        {
          "id": 2985060,
          "author_name": "arindamroy23",
          "author_url": "",
          "post_date": "09/10/2024 09:24:44",
          "content": "<p>Is it usable in this comp<br>\n? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2985473,
              "author_name": "vsahin",
              "author_url": "",
              "post_date": "09/10/2024 17:00:42",
              "content": "<p>Part of the winning requirements is making your stuff public anyway. Not just the code but your process etc.</p>\n<p>So it should not matter</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2986543,
              "author_name": "zaakciiru",
              "author_url": "",
              "post_date": "09/11/2024 18:44:16",
              "content": "<p>As far as I know, you can use the Yolo V8 if you do not plan to receive a cash prize. If the license limits commercial use, then the cash prize will not be paid, but they will give a medal. Of course, I can understand something wrong. Have a good day!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2985068,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "09/10/2024 09:34:13",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2985077,
      "author_name": "arindamroy23",
      "author_url": "",
      "post_date": "09/10/2024 09:45:25",
      "content": "<p>Thank you so very much for sharing this <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a>. </p>\n<p>I wanted to ask you: </p>\n<p>During training, are you storing the bounding box annotated images in the working directory or passing its coordinates to the model<br>\nOr is the model figuring out the bounding boxes by itself? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2985105,
          "author_name": "arindamroy23",
          "author_url": "",
          "post_date": "09/10/2024 10:31:06",
          "content": "<p>Apologies. I understood how it is being passed. I leave an explanation for anyone confused: </p>\n<ul>\n<li>each data fold file has a label and image file. The label consists a pin point of the anamoly in the MRI. This is being fed to the model. </li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2985282,
      "author_name": "rayandsun",
      "author_url": "",
      "post_date": "09/10/2024 13:46:43",
      "content": "<p>Thanks for your sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2985389,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "09/10/2024 15:18:16",
      "content": "<p>Thanks for sharing. Just a quick question, did you do some preprocessing of training data， in order to use YOLO? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2985731,
          "author_name": "namgalielei",
          "author_url": "",
          "post_date": "09/11/2024 00:23:50",
          "content": "<p>I only remove study with null labels</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2985474,
      "author_name": "xianhellg",
      "author_url": "",
      "post_date": "09/10/2024 17:02:47",
      "content": "<p>Good work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2986066,
      "author_name": "arindamroy23",
      "author_url": "",
      "post_date": "09/11/2024 10:33:45",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> !  </p>\n<p>I wanted to ask you, why did you choose to train separate models for each condition instead of view (t1 sag,t2 sag, axial)? Is there any particular reason for  this design choice? What do you think of one model per view ? My intuition is YOLO might train better on similar looking images. Do let me know your thoughts if possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2986166,
          "author_name": "namgalielei",
          "author_url": "",
          "post_date": "09/11/2024 12:43:35",
          "content": "<p>Actually the view (modality) corresponds to conditions: t1 sag -&gt; nfn, axial t2 -&gt; ss, t2 sag -&gt; scs</p>",
          "votes": null,
          "replies": [
            {
              "id": 2986193,
              "author_name": "arindamroy23",
              "author_url": "",
              "post_date": "09/11/2024 13:23:40",
              "content": "<p>Thanks for clarrification!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2987497,
      "author_name": "greenpy",
      "author_url": "",
      "post_date": "09/12/2024 17:58:13",
      "content": "<p>Thanks for sharing .</p>\n<p>How do you define the size of box around that area? <br>\nHow do you know which size is  fine ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2990164,
          "author_name": "namgalielei",
          "author_url": "",
          "post_date": "09/16/2024 04:03:34",
          "content": "<p>Just define a fixed small size relative to the image size. It does not matter much because we only care about the whether the findings appear on these images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2988266,
      "author_name": "zznznb",
      "author_url": "",
      "post_date": "09/13/2024 15:25:00",
      "content": "<p>Thank you for sharing your work. I have found that some models have too many categories and an imbalance of categories, resulting in almost no detection results for some classes during inference. Even if I train each model for 5 folds and lower the detection threshold, I cannot solve this problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2990226,
      "author_name": "rishantenis",
      "author_url": "",
      "post_date": "09/16/2024 05:45:21",
      "content": "<p>This suggestion reflects a great deal of thought in training separate models for different modalities, and the combination with a straightforward aggregation method makes it an elegant solution. Thank you for sharing your insights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2992089,
      "author_name": "arielzhao",
      "author_url": "",
      "post_date": "09/18/2024 07:01:31",
      "content": "<p>Thanks for your sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2997846,
      "author_name": "syamiyer",
      "author_url": "",
      "post_date": "09/25/2024 00:45:38",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2997888,
      "author_name": "lunarecho1",
      "author_url": "",
      "post_date": "09/25/2024 02:34:17",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3008810,
      "author_name": "quan0095",
      "author_url": "",
      "post_date": "10/07/2024 06:07:52",
      "content": "<p>I have a problem, In a 3D volume I detect 3 L5/S1, 2 is normal, 1 is severe. So, the predicted label is expected to be either normal or severe?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3008900,
          "author_name": "namgalielei",
          "author_url": "",
          "post_date": "10/07/2024 08:56:43",
          "content": "<p>Ideally, if there is \"severe\" finding at any slice, then the whole sequence should be classified as \"severe\". However, these Yolo models I've trained don't have such reliability, and the label distribution is highly unbalanced towards \"normal/mild\", in this case I simply take the max probability within each severity across all the slices, then normalize by dividing by the sum. You might find a better aggregation technique</p>",
          "votes": null,
          "replies": [
            {
              "id": 3008903,
              "author_name": "quan0095",
              "author_url": "",
              "post_date": "10/07/2024 09:06:29",
              "content": "<p>I have experiment with end-to-end and two stage method. I think an end-to-end approach is necessary for this problem, but the drawback is a lack of data, leading to overfitting very quickly. A two-stage method requires rule configuration and with a 38% LB, it's also prone to overfitting. Additionally, two-stage methods require a deep understanding of the problem to manually configure rules.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3012707,
      "author_name": "hulkkk",
      "author_url": "",
      "post_date": "10/09/2024 09:11:26",
      "content": "<p>Can you share your final contest plan，Thank you very much！😁</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2984973": "Most of the public notebooks have approached this competition by building classification models on sequences of images and making study-level predictions directly. \n\nI'd like to propose a bottom up approach by building object detection models for each modality to predict bounding boxes for level and condition at the same time. For example, a neural_foraminal_narrowing model learns to predict 5 levels * 2 positions (left/right) * 3 conditions = 30 labels. Then, a simple aggregation method (taking the maximum within each \"condition_level\" ) is used to combine the results from images to obtain study-level predictions. \n\nEvaluation and Submission notebook: https://www.kaggle.com/code/namgalielei/lsdc-yolo-approach/notebook\n\nTraining notebooks:\nSCS: https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-scs/notebook\nNFN: https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-nfn\nSS: https://www.kaggle.com/code/namgalielei/lsdc-train-yolo-ss",
    "2985038": "For the future. Be careful with `ultralytics` **AGPL-3.0 license**! As they can essentially assert open-source requirements for any project that incorporates their code.\n\nYou can use 'original' YOLO implementation by WongKinYiu without any risk",
    "2985060": "Is it usable in this comp\n?",
    "2985068": "Thanks for sharing",
    "2985077": "Thank you so very much for sharing this @namgalielei. \n\nI wanted to ask you: \n\nDuring training, are you storing the bounding box annotated images in the working directory or passing its coordinates to the model\nOr is the model figuring out the bounding boxes by itself?",
    "2985105": "Apologies. I understood how it is being passed. I leave an explanation for anyone confused: \n\n- each data fold file has a label and image file. The label consists a pin point of the anamoly in the MRI. This is being fed to the model.",
    "2985282": "Thanks for your sharing!",
    "2985389": "Thanks for sharing. Just a quick question, did you do some preprocessing of training data， in order to use YOLO?",
    "2985473": "Part of the winning requirements is making your stuff public anyway. Not just the code but your process etc.\n\nSo it should not matter",
    "2985474": "Good work!",
    "2985731": "I only remove study with null labels",
    "2986066": "Hey @namgalielei !  \n\nI wanted to ask you, why did you choose to train separate models for each condition instead of view (t1 sag,t2 sag, axial)? Is there any particular reason for  this design choice? What do you think of one model per view ? My intuition is YOLO might train better on similar looking images. Do let me know your thoughts if possible.",
    "2986166": "Actually the view (modality) corresponds to conditions: t1 sag -> nfn, axial t2 -> ss, t2 sag -> scs",
    "2986193": "Thanks for clarrification!",
    "2986543": "As far as I know, you can use the Yolo V8 if you do not plan to receive a cash prize. If the license limits commercial use, then the cash prize will not be paid, but they will give a medal. Of course, I can understand something wrong. Have a good day!",
    "2987497": "Thanks for sharing .\n\nHow do you define the size of box around that area? \nHow do you know which size is  fine ?",
    "2988266": "Thank you for sharing your work. I have found that some models have too many categories and an imbalance of categories, resulting in almost no detection results for some classes during inference. Even if I train each model for 5 folds and lower the detection threshold, I cannot solve this problem.",
    "2990164": "Just define a fixed small size relative to the image size. It does not matter much because we only care about the whether the findings appear on these images.",
    "2990226": "This suggestion reflects a great deal of thought in training separate models for different modalities, and the combination with a straightforward aggregation method makes it an elegant solution. Thank you for sharing your insights.",
    "2992089": "Thanks for your sharing.",
    "2997846": "Thanks for sharing!",
    "2997888": "Thanks for sharing",
    "3008810": "I have a problem, In a 3D volume I detect 3 L5/S1, 2 is normal, 1 is severe. So, the predicted label is expected to be either normal or severe?",
    "3008900": "Ideally, if there is \"severe\" finding at any slice, then the whole sequence should be classified as \"severe\". However, these Yolo models I've trained don't have such reliability, and the label distribution is highly unbalanced towards \"normal/mild\", in this case I simply take the max probability within each severity across all the slices, then normalize by dividing by the sum. You might find a better aggregation technique",
    "3008903": "I have experiment with end-to-end and two stage method. I think an end-to-end approach is necessary for this problem, but the drawback is a lack of data, leading to overfitting very quickly. A two-stage method requires rule configuration and with a 38% LB, it's also prone to overfitting. Additionally, two-stage methods require a deep understanding of the problem to manually configure rules.",
    "3012707": "Can you share your final contest plan，Thank you very much！😁"
  },
  "source": "meta"
}