{
  "id": 265583,
  "title": "9th place solution",
  "url": "/competitions/siim-covid19-detection/discussion/265583",
  "author_name": "Psi",
  "post_date": "2021-08-16T07:46:49.046000",
  "votes": 50,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the hosts for this interesting competition. In the following, we want to give a summary of the solution of Team Watercooled: <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>. As always, thanks to all team members contributing equally to the solution.</p>\n<h3>Summary</h3>\n<p>Our solution is based on a blend of separate study (classification) and image (detection) models. Most study models are regularized by additional segmentation loss. Detection models include EfficientDet and Yolo models and are blended using WBF. We only rely on provided competition data and do not utilize any external data.</p>\n<h3>Preprocessing and CV</h3>\n<p>For common preprocessing, we transform the original DICOM files to PNG images and rescale the size to either 25% or 50% of the original size to speed up data loading and training runtimes. We handled duplicate images as suggested by the hosts (<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a>) by only taking images with boxes if several images belong to the same study.  </p>\n<p>For validation, we employ 5-fold cross validation where the same study is not overlapping between folds. Overall, we can see quite decent CV and LB correlation, within a certain random range which is to be expected given the small data and nature of the problem and validation metric.</p>\n<h3>Study models</h3>\n<p>Our final blend contains nine models, each one fitted across 5 folds, leading to an overall average of 45 fits. For each model, we use cross entropy loss on all four classes, and also process the outputs with softmax before averaging. For us, softmax was clearly superior as each image can always only have a single target. All study models are regularized by additional segmentation loss using a Unet decoder. In detail, these are the different models blended:</p>\n<ol>\n<li>Backbone: tf_efficientnet_b0, Image size: (512,512)</li>\n<li>Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnetv2_l, Image size: (512,512), Note: Ben preprocessing</li>\n<li>Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)</li>\n<li>Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)</li>\n</ol>\n<p>For augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.</p>\n<h3>Image models</h3>\n<p>For detection, we used both EfficientDet and Yolo models.</p>\n<h4>Efficient Det models</h4>\n<p>For most EfficientDet models, we regularize the detection part with additional study classification, but only use the detection part as output. For a further discussion on hybrid models, see below. In detail, we trained the following EfficientDet models, each for 5-folds:</p>\n<ol>\n<li>Backbone: tf_efficientdet_d0, Image size: (512,512)</li>\n<li>Backbone: tf_efficientdet_d0, Image size: (512,512)</li>\n<li>Backbone: tf_efficientdet_d3, Image size: (512,512)</li>\n<li>Backbone: efficientdet_q2, Image size: (768,768)</li>\n</ol>\n<p>For augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.</p>\n<h4>Yolo models</h4>\n<p>We trained several yolo models, were 3/ 4 uses opacity as the only class and has no bounding box if the image had none as label. For one yolo model, however we created a full-image boundingbox with label “none” as a second label</p>\n<ol>\n<li>yolov5s (512) opacity</li>\n<li>yolov5x (512) opacity</li>\n<li>yolov5m (640) opacity</li>\n<li>yolov5s (512) opacity + none</li>\n</ol>\n<h4>None predictions</h4>\n<p>One peculiarity of this competition was to figure out how to best predict none for the detection part. Several competitors decided to have a separate binary classification model, but this appeared to be quite limiting to us. So in the end, we decided to make none predictions by combining image and study predictions. In detail, we calculate:</p>\n<p>none = 0.5 * (1-max(box_confidence)) + 0.5 * negative + 0.2 * atypical</p>\n<p>The reason for including atypical for the prediction for none, is the fact that \"Bounding boxes were not placed on pleural effusions, or pneumothoraces.” and these fall into the category of atypical.</p>\n<h4>Blending</h4>\n<p>We blended all outputs using WBF with IOU 0.5. To speed up blending, we only take the top75 boxes for each model output for a given image.</p>\n<h3>A note on hybrid models</h3>\n<p>Originally, we started training only hybrid models, i.e. combining study and image predictions in a combined EfficientDet model where we re-use the EfficientNet encoder then for both detections and classifications. This worked really well, and also has a good built-in regularization effect for both detections and classifications. On study level this actually worked similarly as the segmentation regularization. However, as we got tiny improvements after splitting up the models, we stuck to that. But a more streamlined solution can totally use these hybrid models with very similar expected results.</p>\n<h3>A note on external data</h3>\n<p>Reading other solution posts, it appears we have, not for the first time, wasted some further improvements by not utilizing external data. We tried a little bit to incorporate ChestX, without much success, but have not dug too deep and not approached other datasets. It appears to be specifically useful in Chest XRay competitions to use all data available, even if you do not see immediate gains on CV. This might be specifically true as it is hard to employ mixup/cutmix techniques to multiply the actual available data. Sometimes data also stems from similar/same datasets and small effects might also only appear on the leaderboard. </p>\n<h3>Links</h3>\n<p>Training: <a href=\"https://github.com/ChristofHenkel/kaggle-siim-covid-detection-9th-place\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-siim-covid-detection-9th-place</a><br>\nInference: <a href=\"https://www.kaggle.com/ilu000/siim-covid-detection-9th-place\" target=\"_blank\">https://www.kaggle.com/ilu000/siim-covid-detection-9th-place</a><br>\nVideo: <a href=\"https://www.youtube.com/watch?v=lRJyB-gi-nk\" target=\"_blank\">https://www.youtube.com/watch?v=lRJyB-gi-nk</a></p>",
  "messages": [
    {
      "id": 1474633,
      "postDate": "2021-08-16T07:46:49.047Z",
      "content": "<p>Thanks to Kaggle and the hosts for this interesting competition. In the following, we want to give a summary of the solution of Team Watercooled: <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>. As always, thanks to all team members contributing equally to the solution.</p>\n<h3>Summary</h3>\n<p>Our solution is based on a blend of separate study (classification) and image (detection) models. Most study models are regularized by additional segmentation loss. Detection models include EfficientDet and Yolo models and are blended using WBF. We only rely on provided competition data and do not utilize any external data.</p>\n<h3>Preprocessing and CV</h3>\n<p>For common preprocessing, we transform the original DICOM files to PNG images and rescale the size to either 25% or 50% of the original size to speed up data loading and training runtimes. We handled duplicate images as suggested by the hosts (<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a>) by only taking images with boxes if several images belong to the same study.  </p>\n<p>For validation, we employ 5-fold cross validation where the same study is not overlapping between folds. Overall, we can see quite decent CV and LB correlation, within a certain random range which is to be expected given the small data and nature of the problem and validation metric.</p>\n<h3>Study models</h3>\n<p>Our final blend contains nine models, each one fitted across 5 folds, leading to an overall average of 45 fits. For each model, we use cross entropy loss on all four classes, and also process the outputs with softmax before averaging. For us, softmax was clearly superior as each image can always only have a single target. All study models are regularized by additional segmentation loss using a Unet decoder. In detail, these are the different models blended:</p>\n<ol>\n<li>Backbone: tf_efficientnet_b0, Image size: (512,512)</li>\n<li>Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)</li>\n<li>Backbone: tf_efficientnetv2_l, Image size: (512,512), Note: Ben preprocessing</li>\n<li>Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)</li>\n<li>Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)</li>\n</ol>\n<p>For augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.</p>\n<h3>Image models</h3>\n<p>For detection, we used both EfficientDet and Yolo models.</p>\n<h4>Efficient Det models</h4>\n<p>For most EfficientDet models, we regularize the detection part with additional study classification, but only use the detection part as output. For a further discussion on hybrid models, see below. In detail, we trained the following EfficientDet models, each for 5-folds:</p>\n<ol>\n<li>Backbone: tf_efficientdet_d0, Image size: (512,512)</li>\n<li>Backbone: tf_efficientdet_d0, Image size: (512,512)</li>\n<li>Backbone: tf_efficientdet_d3, Image size: (512,512)</li>\n<li>Backbone: efficientdet_q2, Image size: (768,768)</li>\n</ol>\n<p>For augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.</p>\n<h4>Yolo models</h4>\n<p>We trained several yolo models, were 3/ 4 uses opacity as the only class and has no bounding box if the image had none as label. For one yolo model, however we created a full-image boundingbox with label “none” as a second label</p>\n<ol>\n<li>yolov5s (512) opacity</li>\n<li>yolov5x (512) opacity</li>\n<li>yolov5m (640) opacity</li>\n<li>yolov5s (512) opacity + none</li>\n</ol>\n<h4>None predictions</h4>\n<p>One peculiarity of this competition was to figure out how to best predict none for the detection part. Several competitors decided to have a separate binary classification model, but this appeared to be quite limiting to us. So in the end, we decided to make none predictions by combining image and study predictions. In detail, we calculate:</p>\n<p>none = 0.5 * (1-max(box_confidence)) + 0.5 * negative + 0.2 * atypical</p>\n<p>The reason for including atypical for the prediction for none, is the fact that \"Bounding boxes were not placed on pleural effusions, or pneumothoraces.” and these fall into the category of atypical.</p>\n<h4>Blending</h4>\n<p>We blended all outputs using WBF with IOU 0.5. To speed up blending, we only take the top75 boxes for each model output for a given image.</p>\n<h3>A note on hybrid models</h3>\n<p>Originally, we started training only hybrid models, i.e. combining study and image predictions in a combined EfficientDet model where we re-use the EfficientNet encoder then for both detections and classifications. This worked really well, and also has a good built-in regularization effect for both detections and classifications. On study level this actually worked similarly as the segmentation regularization. However, as we got tiny improvements after splitting up the models, we stuck to that. But a more streamlined solution can totally use these hybrid models with very similar expected results.</p>\n<h3>A note on external data</h3>\n<p>Reading other solution posts, it appears we have, not for the first time, wasted some further improvements by not utilizing external data. We tried a little bit to incorporate ChestX, without much success, but have not dug too deep and not approached other datasets. It appears to be specifically useful in Chest XRay competitions to use all data available, even if you do not see immediate gains on CV. This might be specifically true as it is hard to employ mixup/cutmix techniques to multiply the actual available data. Sometimes data also stems from similar/same datasets and small effects might also only appear on the leaderboard. </p>\n<h3>Links</h3>\n<p>Training: <a href=\"https://github.com/ChristofHenkel/kaggle-siim-covid-detection-9th-place\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-siim-covid-detection-9th-place</a><br>\nInference: <a href=\"https://www.kaggle.com/ilu000/siim-covid-detection-9th-place\" target=\"_blank\">https://www.kaggle.com/ilu000/siim-covid-detection-9th-place</a><br>\nVideo: <a href=\"https://www.youtube.com/watch?v=lRJyB-gi-nk\" target=\"_blank\">https://www.youtube.com/watch?v=lRJyB-gi-nk</a></p>",
      "rawMarkdown": "Thanks to Kaggle and the hosts for this interesting competition. In the following, we want to give a summary of the solution of Team Watercooled: @christofhenkel, @ilu000, @philippsinger. As always, thanks to all team members contributing equally to the solution.\n\n###Summary\n\nOur solution is based on a blend of separate study (classification) and image (detection) models. Most study models are regularized by additional segmentation loss. Detection models include EfficientDet and Yolo models and are blended using WBF. We only rely on provided competition data and do not utilize any external data.\n\n###Preprocessing and CV\n\nFor common preprocessing, we transform the original DICOM files to PNG images and rescale the size to either 25% or 50% of the original size to speed up data loading and training runtimes. We handled duplicate images as suggested by the hosts (https://www.kaggle.com/c/siim-covid19-detection/discussion/246597) by only taking images with boxes if several images belong to the same study.  \n\nFor validation, we employ 5-fold cross validation where the same study is not overlapping between folds. Overall, we can see quite decent CV and LB correlation, within a certain random range which is to be expected given the small data and nature of the problem and validation metric.\n\n###Study models\n\nOur final blend contains nine models, each one fitted across 5 folds, leading to an overall average of 45 fits. For each model, we use cross entropy loss on all four classes, and also process the outputs with softmax before averaging. For us, softmax was clearly superior as each image can always only have a single target. All study models are regularized by additional segmentation loss using a Unet decoder. In detail, these are the different models blended:\n\n1. Backbone: tf_efficientnet_b0, Image size: (512,512)\n2. Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)\n3. Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)\n4. Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)\n5. Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)\n6. Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)\n7. Backbone: tf_efficientnetv2_l, Image size: (512,512), Note: Ben preprocessing\n8. Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)\n9. Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)\n\nFor augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.\n\n###Image models\n\nFor detection, we used both EfficientDet and Yolo models.\n\n####Efficient Det models\n\nFor most EfficientDet models, we regularize the detection part with additional study classification, but only use the detection part as output. For a further discussion on hybrid models, see below. In detail, we trained the following EfficientDet models, each for 5-folds:\n\n1. Backbone: tf_efficientdet_d0, Image size: (512,512)\n2. Backbone: tf_efficientdet_d0, Image size: (512,512)\n3. Backbone: tf_efficientdet_d3, Image size: (512,512)\n4. Backbone: efficientdet_q2, Image size: (768,768)\n\nFor augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.\n\n####Yolo models\n\nWe trained several yolo models, were 3/ 4 uses opacity as the only class and has no bounding box if the image had none as label. For one yolo model, however we created a full-image boundingbox with label “none” as a second label\n\n1. yolov5s (512) opacity\n2. yolov5x (512) opacity\n3. yolov5m (640) opacity\n4. yolov5s (512) opacity + none\n\n####None predictions\n\nOne peculiarity of this competition was to figure out how to best predict none for the detection part. Several competitors decided to have a separate binary classification model, but this appeared to be quite limiting to us. So in the end, we decided to make none predictions by combining image and study predictions. In detail, we calculate:\n\nnone = 0.5 * (1-max(box_confidence)) + 0.5 * negative + 0.2 * atypical\n\nThe reason for including atypical for the prediction for none, is the fact that \"Bounding boxes were not placed on pleural effusions, or pneumothoraces.” and these fall into the category of atypical.\n\n####Blending\nWe blended all outputs using WBF with IOU 0.5. To speed up blending, we only take the top75 boxes for each model output for a given image.\n\n###A note on hybrid models\n\nOriginally, we started training only hybrid models, i.e. combining study and image predictions in a combined EfficientDet model where we re-use the EfficientNet encoder then for both detections and classifications. This worked really well, and also has a good built-in regularization effect for both detections and classifications. On study level this actually worked similarly as the segmentation regularization. However, as we got tiny improvements after splitting up the models, we stuck to that. But a more streamlined solution can totally use these hybrid models with very similar expected results.\n\n### A note on external data\nReading other solution posts, it appears we have, not for the first time, wasted some further improvements by not utilizing external data. We tried a little bit to incorporate ChestX, without much success, but have not dug too deep and not approached other datasets. It appears to be specifically useful in Chest XRay competitions to use all data available, even if you do not see immediate gains on CV. This might be specifically true as it is hard to employ mixup/cutmix techniques to multiply the actual available data. Sometimes data also stems from similar/same datasets and small effects might also only appear on the leaderboard. \n\n### Links\n\nTraining: https://github.com/ChristofHenkel/kaggle-siim-covid-detection-9th-place\nInference: https://www.kaggle.com/ilu000/siim-covid-detection-9th-place\nVideo: https://www.youtube.com/watch?v=lRJyB-gi-nk",
      "votes": 50
    },
    {
      "id": 1475305,
      "postDate": "2021-08-16T15:03:53.643Z",
      "content": "<p>Yes, pretraining on external data is the key here. I also missed it. Lesson learned.</p>",
      "rawMarkdown": "Yes, pretraining on external data is the key here. I also missed it. Lesson learned.",
      "votes": 5,
      "replies": [
        {
          "id": 1475543,
          "postDate": "2021-08-16T17:36:50.620Z",
          "content": "<p>Agree, we also skipped this part thinking it won't help.</p>",
          "rawMarkdown": "Agree, we also skipped this part thinking it won't help."
        },
        {
          "id": 1475618,
          "postDate": "2021-08-16T19:04:13.383Z",
          "content": "<p>At least you pseudo-tagged, we didnt even do that 😄</p>",
          "rawMarkdown": "At least you pseudo-tagged, we didnt even do that 😄"
        },
        {
          "id": 1475635,
          "postDate": "2021-08-16T19:19:33.370Z",
          "content": "<p>Unfortunately pseudo-labeling caused slight drop in my private LB, even though it helped both CV and public LB</p>",
          "rawMarkdown": "Unfortunately pseudo-labeling caused slight drop in my private LB, even though it helped both CV and public LB"
        },
        {
          "id": 1475637,
          "postDate": "2021-08-16T19:20:58.223Z",
          "content": "<p>Interesting, did the ones without pseudo score better in private? Was one of your selected ones without?</p>",
          "rawMarkdown": "Interesting, did the ones without pseudo score better in private? Was one of your selected ones without?"
        },
        {
          "id": 1475655,
          "postDate": "2021-08-16T19:33:34.853Z",
          "content": "<p>I noticed the drop in my submission records. Both my selected were with pseudo labels, I didn't expect a drop in private for pseudo labeling to be honest. </p>",
          "rawMarkdown": "I noticed the drop in my submission records. Both my selected were with pseudo labels, I didn't expect a drop in private for pseudo labeling to be honest. "
        },
        {
          "id": 1475661,
          "postDate": "2021-08-16T19:39:02.710Z",
          "content": "<p>study-level only submissions:<br>\nwith pseudo labels: private 0.385 public 0.405<br>\nwithout pseudo labels: private 0.386 public 0.401<br>\nCV improved 0.005-0.006</p>",
          "rawMarkdown": "study-level only submissions:\nwith pseudo labels: private 0.385 public 0.405\nwithout pseudo labels: private 0.386 public 0.401\nCV improved 0.005-0.006",
          "votes": 2
        },
        {
          "id": 1475690,
          "postDate": "2021-08-16T19:47:59.647Z",
          "content": "<p>We missed pretraining on any external data too. Pseudo labels did give us a good amount of boost on cv, although it didn't reflect back that much on the leaderboard. <br>\n<a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Also, It seems you used higher image size in classification models but not in detection models, while it was the opposite for us in terms of cv/lb boost. Did you experiment with larger image sizes detection models?</p>",
          "rawMarkdown": "We missed pretraining on any external data too. Pseudo labels did give us a good amount of boost on cv, although it didn't reflect back that much on the leaderboard. \n@philippsinger Also, It seems you used higher image size in classification models but not in detection models, while it was the opposite for us in terms of cv/lb boost. Did you experiment with larger image sizes detection models?"
        },
        {
          "id": 1477469,
          "postDate": "2021-08-17T13:45:08.007Z",
          "content": "<p>Yes we did but didnt give any noticable improvements. In general, image size didnt seem to be important in this competition as the models (over)fit the data very easily.</p>",
          "rawMarkdown": "Yes we did but didnt give any noticable improvements. In general, image size didnt seem to be important in this competition as the models (over)fit the data very easily.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1476071,
      "postDate": "2021-08-17T01:17:15.547Z",
      "content": "<blockquote>\n  <p>Bounding boxes were not placed on pleural effusions, or pneumothoraces</p>\n</blockquote>\n<p>You are looking at the data very carefully. See you in another competition!</p>",
      "rawMarkdown": "> Bounding boxes were not placed on pleural effusions, or pneumothoraces\n\nYou are looking at the data very carefully. See you in another competition!",
      "votes": 1,
      "replies": [
        {
          "id": 1480126,
          "postDate": "2021-08-18T20:14:35.893Z",
          "content": "<p>Actually the hosts posted this also somewhere :)</p>",
          "rawMarkdown": "Actually the hosts posted this also somewhere :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1632345,
      "postDate": "2021-12-29T14:21:45.120Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> , you use WBF in only study model, or on image model, or both?</p>",
      "rawMarkdown": "Hi @philippsinger , you use WBF in only study model, or on image model, or both?"
    },
    {
      "id": 1480517,
      "postDate": "2021-08-19T04:03:07.763Z",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>,  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> It seems that you guys have used a total of <code>8</code> detection models including <code>4</code> <strong>YOLOs</strong> and <code>4</code> <strong>EffDets</strong>. </p>\n<ul>\n<li>So, if we consider <strong>5 folds</strong> did you guys used <code>8*5=40</code> models?</li>\n<li>How did you merged boxes from all models? Did you first blend all folds of a model then merged all models?</li>\n</ul>",
      "rawMarkdown": "@philippsinger,  @christofhenkel, @ilu000 It seems that you guys have used a total of `8` detection models including `4` **YOLOs** and `4` **EffDets**. \n* So, if we consider **5 folds** did you guys used `8*5=40` models?\n* How did you merged boxes from all models? Did you first blend all folds of a model then merged all models?",
      "replies": [
        {
          "id": 1482686,
          "postDate": "2021-08-20T07:36:53.913Z",
          "content": "<p>Yes 40 models sounds about right.<br>\nWe blended all 40 models in the end with WBF.</p>",
          "rawMarkdown": "Yes 40 models sounds about right.\nWe blended all 40 models in the end with WBF.",
          "votes": 1
        },
        {
          "id": 1483166,
          "postDate": "2021-08-20T13:34:48.210Z",
          "content": "<p>Thanks for the answer :) </p>",
          "rawMarkdown": "Thanks for the answer :) "
        },
        {
          "id": 1495472,
          "postDate": "2021-08-29T15:22:51.430Z",
          "content": "<blockquote>\n  <p>How did you merged boxes from all models? Did you first blend all folds of a model then merged all models?</p>\n</blockquote>\n<p>We blended the models in one step, but we only kept the top 75 boxes of each model for each image. This speeds up the WBF a lot and on CV did not hurt much. </p>",
          "rawMarkdown": ">How did you merged boxes from all models? Did you first blend all folds of a model then merged all models?\n\nWe blended the models in one step, but we only kept the top 75 boxes of each model for each image. This speeds up the WBF a lot and on CV did not hurt much. "
        }
      ]
    },
    {
      "id": 1475426,
      "postDate": "2021-08-16T16:05:14.117Z",
      "content": "<p>Congratulations to everyone on your team ! </p>\n<p>Do you have tried to pre-train on chestX dataset after the competition ? If yes by how much did it improve your CV ? </p>",
      "rawMarkdown": "Congratulations to everyone on your team ! \n\nDo you have tried to pre-train on chestX dataset after the competition ? If yes by how much did it improve your CV ? ",
      "replies": [
        {
          "id": 1475619,
          "postDate": "2021-08-16T19:05:37.520Z",
          "content": "<p>No, not really time to go into detail here. It probably also needs a few iterations to get the right settings.</p>",
          "rawMarkdown": "No, not really time to go into detail here. It probably also needs a few iterations to get the right settings.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1474972,
      "postDate": "2021-08-16T11:23:17.117Z",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Hi, is it <code>efficientdet_q2</code> or <code>efficientdet_d2</code> ?</p>",
      "rawMarkdown": "@philippsinger Hi, is it ` efficientdet_q2` or ` efficientdet_d2` ?",
      "replies": [
        {
          "id": 1474995,
          "postDate": "2021-08-16T11:42:56.513Z",
          "content": "<p><code>efficientdet_q2</code></p>",
          "rawMarkdown": "`efficientdet_q2`"
        },
        {
          "id": 1474998,
          "postDate": "2021-08-16T11:47:04.443Z",
          "content": "<p>would be able to tell the difference? why <code>q</code> ?</p>",
          "rawMarkdown": "would be able to tell the difference? why `q` ?"
        },
        {
          "id": 1475009,
          "postDate": "2021-08-16T11:50:32.907Z",
          "content": "<p>Has a different FPN layout, you can check in timm effdet package.</p>",
          "rawMarkdown": "Has a different FPN layout, you can check in timm effdet package.",
          "votes": 2
        },
        {
          "id": 1475011,
          "postDate": "2021-08-16T11:51:11.433Z",
          "content": "<p>thanks for the info</p>",
          "rawMarkdown": "thanks for the info"
        }
      ]
    },
    {
      "id": 1622197,
      "postDate": "2021-12-18T13:14:48.903Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1475305,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2021-08-16T15:03:53.643000",
      "content": "<p>Yes, pretraining on external data is the key here. I also missed it. Lesson learned.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1475543,
          "author_name": "Zhanseri Ikram",
          "author_url": "",
          "post_date": "2021-08-16T17:36:50.620000",
          "content": "<p>Agree, we also skipped this part thinking it won't help.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1475618,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-16T19:04:13.383000",
          "content": "<p>At least you pseudo-tagged, we didnt even do that 😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1475635,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2021-08-16T19:19:33.370000",
          "content": "<p>Unfortunately pseudo-labeling caused slight drop in my private LB, even though it helped both CV and public LB</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1475637,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-16T19:20:58.223000",
          "content": "<p>Interesting, did the ones without pseudo score better in private? Was one of your selected ones without?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1475655,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2021-08-16T19:33:34.853000",
          "content": "<p>I noticed the drop in my submission records. Both my selected were with pseudo labels, I didn't expect a drop in private for pseudo labeling to be honest. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1475661,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2021-08-16T19:39:02.710000",
          "content": "<p>study-level only submissions:<br>\nwith pseudo labels: private 0.385 public 0.405<br>\nwithout pseudo labels: private 0.386 public 0.401<br>\nCV improved 0.005-0.006</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1475690,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2021-08-16T19:47:59.647000",
          "content": "<p>We missed pretraining on any external data too. Pseudo labels did give us a good amount of boost on cv, although it didn't reflect back that much on the leaderboard. <br>\n<a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Also, It seems you used higher image size in classification models but not in detection models, while it was the opposite for us in terms of cv/lb boost. Did you experiment with larger image sizes detection models?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1477469,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-17T13:45:08.007000",
          "content": "<p>Yes we did but didnt give any noticable improvements. In general, image size didnt seem to be important in this competition as the models (over)fit the data very easily.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1476071,
      "author_name": "YujiAriyasu",
      "author_url": "",
      "post_date": "2021-08-17T01:17:15.547000",
      "content": "<blockquote>\n  <p>Bounding boxes were not placed on pleural effusions, or pneumothoraces</p>\n</blockquote>\n<p>You are looking at the data very carefully. See you in another competition!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1480126,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-18T20:14:35.893000",
          "content": "<p>Actually the hosts posted this also somewhere :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1632345,
      "author_name": "Thái Hoàng Nhân - CS",
      "author_url": "",
      "post_date": "2021-12-29T14:21:45.120000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> , you use WBF in only study model, or on image model, or both?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480517,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-08-19T04:03:07.763000",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>,  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> It seems that you guys have used a total of <code>8</code> detection models including <code>4</code> <strong>YOLOs</strong> and <code>4</code> <strong>EffDets</strong>. </p>\n<ul>\n<li>So, if we consider <strong>5 folds</strong> did you guys used <code>8*5=40</code> models?</li>\n<li>How did you merged boxes from all models? Did you first blend all folds of a model then merged all models?</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 1482686,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-20T07:36:53.913000",
          "content": "<p>Yes 40 models sounds about right.<br>\nWe blended all 40 models in the end with WBF.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483166,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-08-20T13:34:48.210000",
          "content": "<p>Thanks for the answer :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1495472,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-08-29T15:22:51.430000",
          "content": "<blockquote>\n  <p>How did you merged boxes from all models? Did you first blend all folds of a model then merged all models?</p>\n</blockquote>\n<p>We blended the models in one step, but we only kept the top 75 boxes of each model for each image. This speeds up the WBF a lot and on CV did not hurt much. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1475426,
      "author_name": "Tom Darmon",
      "author_url": "",
      "post_date": "2021-08-16T16:05:14.117000",
      "content": "<p>Congratulations to everyone on your team ! </p>\n<p>Do you have tried to pre-train on chestX dataset after the competition ? If yes by how much did it improve your CV ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1475619,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-16T19:05:37.520000",
          "content": "<p>No, not really time to go into detail here. It probably also needs a few iterations to get the right settings.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1474972,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-08-16T11:23:17.117000",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Hi, is it <code>efficientdet_q2</code> or <code>efficientdet_d2</code> ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1474995,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-16T11:42:56.513000",
          "content": "<p><code>efficientdet_q2</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1474998,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-08-16T11:47:04.443000",
          "content": "<p>would be able to tell the difference? why <code>q</code> ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1475009,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-16T11:50:32.907000",
          "content": "<p>Has a different FPN layout, you can check in timm effdet package.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1475011,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-08-16T11:51:11.433000",
          "content": "<p>thanks for the info</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1622197,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-12-18T13:14:48.903000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1474633": "Thanks to Kaggle and the hosts for this interesting competition. In the following, we want to give a summary of the solution of Team Watercooled: @christofhenkel, @ilu000, @philippsinger. As always, thanks to all team members contributing equally to the solution.\n\n###Summary\n\nOur solution is based on a blend of separate study (classification) and image (detection) models. Most study models are regularized by additional segmentation loss. Detection models include EfficientDet and Yolo models and are blended using WBF. We only rely on provided competition data and do not utilize any external data.\n\n###Preprocessing and CV\n\nFor common preprocessing, we transform the original DICOM files to PNG images and rescale the size to either 25% or 50% of the original size to speed up data loading and training runtimes. We handled duplicate images as suggested by the hosts (https://www.kaggle.com/c/siim-covid19-detection/discussion/246597) by only taking images with boxes if several images belong to the same study.  \n\nFor validation, we employ 5-fold cross validation where the same study is not overlapping between folds. Overall, we can see quite decent CV and LB correlation, within a certain random range which is to be expected given the small data and nature of the problem and validation metric.\n\n###Study models\n\nOur final blend contains nine models, each one fitted across 5 folds, leading to an overall average of 45 fits. For each model, we use cross entropy loss on all four classes, and also process the outputs with softmax before averaging. For us, softmax was clearly superior as each image can always only have a single target. All study models are regularized by additional segmentation loss using a Unet decoder. In detail, these are the different models blended:\n\n1. Backbone: tf_efficientnet_b0, Image size: (512,512)\n2. Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)\n3. Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)\n4. Backbone: tf_efficientnet_b7_ns, Image size: (1024,1024)\n5. Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)\n6. Backbone: tf_efficientnet_b5_ns, Image size: (1024,1024)\n7. Backbone: tf_efficientnetv2_l, Image size: (512,512), Note: Ben preprocessing\n8. Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)\n9. Backbone: xcit_small_24_p16_224_dist, Image size: (640,640)\n\nFor augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.\n\n###Image models\n\nFor detection, we used both EfficientDet and Yolo models.\n\n####Efficient Det models\n\nFor most EfficientDet models, we regularize the detection part with additional study classification, but only use the detection part as output. For a further discussion on hybrid models, see below. In detail, we trained the following EfficientDet models, each for 5-folds:\n\n1. Backbone: tf_efficientdet_d0, Image size: (512,512)\n2. Backbone: tf_efficientdet_d0, Image size: (512,512)\n3. Backbone: tf_efficientdet_d3, Image size: (512,512)\n4. Backbone: efficientdet_q2, Image size: (768,768)\n\nFor augmentations we utilized a mix of ShiftScaleRotate, HorizontalFlip, RandomBrightnessContrast and Cutout. The elaborated models can slightly differ in minor hyperparameter settings.\n\n####Yolo models\n\nWe trained several yolo models, were 3/ 4 uses opacity as the only class and has no bounding box if the image had none as label. For one yolo model, however we created a full-image boundingbox with label “none” as a second label\n\n1. yolov5s (512) opacity\n2. yolov5x (512) opacity\n3. yolov5m (640) opacity\n4. yolov5s (512) opacity + none\n\n####None predictions\n\nOne peculiarity of this competition was to figure out how to best predict none for the detection part. Several competitors decided to have a separate binary classification model, but this appeared to be quite limiting to us. So in the end, we decided to make none predictions by combining image and study predictions. In detail, we calculate:\n\nnone = 0.5 * (1-max(box_confidence)) + 0.5 * negative + 0.2 * atypical\n\nThe reason for including atypical for the prediction for none, is the fact that \"Bounding boxes were not placed on pleural effusions, or pneumothoraces.” and these fall into the category of atypical.\n\n####Blending\nWe blended all outputs using WBF with IOU 0.5. To speed up blending, we only take the top75 boxes for each model output for a given image.\n\n###A note on hybrid models\n\nOriginally, we started training only hybrid models, i.e. combining study and image predictions in a combined EfficientDet model where we re-use the EfficientNet encoder then for both detections and classifications. This worked really well, and also has a good built-in regularization effect for both detections and classifications. On study level this actually worked similarly as the segmentation regularization. However, as we got tiny improvements after splitting up the models, we stuck to that. But a more streamlined solution can totally use these hybrid models with very similar expected results.\n\n### A note on external data\nReading other solution posts, it appears we have, not for the first time, wasted some further improvements by not utilizing external data. We tried a little bit to incorporate ChestX, without much success, but have not dug too deep and not approached other datasets. It appears to be specifically useful in Chest XRay competitions to use all data available, even if you do not see immediate gains on CV. This might be specifically true as it is hard to employ mixup/cutmix techniques to multiply the actual available data. Sometimes data also stems from similar/same datasets and small effects might also only appear on the leaderboard. \n\n### Links\n\nTraining: https://github.com/ChristofHenkel/kaggle-siim-covid-detection-9th-place\nInference: https://www.kaggle.com/ilu000/siim-covid-detection-9th-place\nVideo: https://www.youtube.com/watch?v=lRJyB-gi-nk",
    "1475305": "Yes, pretraining on external data is the key here. I also missed it. Lesson learned.",
    "1476071": "> Bounding boxes were not placed on pleural effusions, or pneumothoraces\n\nYou are looking at the data very carefully. See you in another competition!",
    "1632345": "Hi @philippsinger , you use WBF in only study model, or on image model, or both?",
    "1480517": "@philippsinger,  @christofhenkel, @ilu000 It seems that you guys have used a total of `8` detection models including `4` **YOLOs** and `4` **EffDets**. \n* So, if we consider **5 folds** did you guys used `8*5=40` models?\n* How did you merged boxes from all models? Did you first blend all folds of a model then merged all models?",
    "1475426": "Congratulations to everyone on your team ! \n\nDo you have tried to pre-train on chestX dataset after the competition ? If yes by how much did it improve your CV ? ",
    "1474972": "@philippsinger Hi, is it ` efficientdet_q2` or ` efficientdet_d2` ?",
    "1622197": ""
  }
}