{
  "id": 266571,
  "title": "5th Place Solution - Complete Writeup + Code ",
  "url": "/competitions/siim-covid19-detection/discussion/266571",
  "author_name": "",
  "post_date": "2021-08-19T15:25:29.401092Z",
  "votes": 28,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi everyone, sorry for taking a bit longer to publish our complete solution. It took us several days to clean the code files, update the GitHub repo, and finishing with the write-up.<br>\nThanks to the Kaggle team and SIIM, FIBASIO, RSNA, and all the sponsors who hosted this challenging and interesting covid19 classification and detection competition! Also, thanks to my teammates <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> &amp; <a href=\"https://www.kaggle.com/benihime91\" target=\"_blank\">@benihime91</a> without whom it wouldn't have been possible.</p>\n<p>Representing the  <strong><em>Ayushman Nischay Shivam</em></strong> team, in this post I’m going to explain our winning solution in detail.</p>\n<p>Also, we have already made our inference notebook along with model weights public: you may visit that with this <a href=\"https://www.kaggle.com/nischaydnk/604e8587410a-v2m-bin-weighted\" target=\"_blank\">link</a></p>\n<p>Our all codes related to training or preprocessing data codes are also made public: <a href=\"https://github.com/benihime91/SIIM-COVID19-DETECTION-KAGGLE\" target=\"_blank\">https://github.com/benihime91/SIIM-COVID19-DETECTION-KAGGLE</a></p>\n<p>As <a href=\"https://www.kaggle.com/benihime91\" target=\"_blank\">@benihime91</a> has also talked and summarised a bit about our solution in this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263945\" target=\"_blank\">post</a> , I will try to cover up everything in detail.</p>\n<h2>Overview</h2>\n<p><strong>Our winning blend consists of :</strong></p>\n<p>11 multiclass classification models with 4 different architectures.<br>\n2 x 5 fold (10) binary classification models with 2 different architectures.<br>\n5 x 5 fold (25) object detection models with 5 different architectures.</p>\n<p><a href=\"https://ibb.co/0DcqQwR\"><img src=\"https://i.ibb.co/p0x2nmB/Screenshot-2021-08-19-at-8-27-40-PM.png\" alt=\"Screenshot-2021-08-19-at-8-27-40-PM\"></a></p>\n<h2>Image Data Used</h2>\n<p>We only used competition <a href=\"https://www.kaggle.com/c/siim-covid19-detection/data\" target=\"_blank\">data</a> for training models. <em>No external image data was used.</em></p>\n<h1>Models Summary</h1>\n<h2>Study Level Models</h2>\n<p>All of our study models were trained with variants of Efficientnet models. Other architectures like Resnet, Densenet, transformer-based models didn’t perform well for us. Our models were pretrained on imagenet and didn’t use any external x-ray data directly / indirectly during the competition. Some models were trained on multiple stages which includes finetuning with a reduced learning rate or increasing the image size.</p>\n<p>Baseline Architectures used in our final study level solution:</p>\n<ul>\n<li>Efficientnet v2m</li>\n<li>Efficientnet v2l</li>\n<li>Efficientnet B5</li>\n<li>Efficientnet B7</li>\n</ul>\n<p><a href=\"https://ibb.co/7nMKjnN\"><img src=\"https://i.ibb.co/T4jtY4q/Screenshot-2021-08-19-at-5-50-19-PM.png\" alt=\"Screenshot-2021-08-19-at-5-50-19-PM\"></a></p>\n<h2>Efficientnet v2m:</h2>\n<ul>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>512 x 512(5- fold) &amp; 640x640(fold 0) &amp; 1024x1024(fold 0) image size.</p></li>\n<li><p>PCAM pooling + SAM attention map used in stage2</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were generated with PCAM pooling + DANET attention map in stage 1.</p></li>\n<li><p>Ranger optimizer, Cosine Scheduler with warmup were used.</p></li>\n<li><p>Activation layers of the model were replaced with Mish activation</p>\n<h2>Efficientnet v2l:</h2></li>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>512 x 512(5 folds) &amp; 640x640 (fold 1) image size</p></li>\n<li><p>PCAM pooling + SAM attention map used in training and finetune stage.</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were introduced using the out of folds predictions of the Efficientnet v2m model mentioned above.</p></li>\n<li><p>Ranger optimizer, Cosine Scheduler with warmup were used.</p></li>\n<li><p>Activation layers of the model were replaced with Mish activation</p></li>\n</ul>\n<h2>Efficientnet B5:</h2>\n<ul>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>640 x 640 image size</p></li>\n<li><p>average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were introduced using the out of folds predictions of the efficientnet B5 model with similar configs.</p></li>\n<li><p>AdamW optimizer with OneCycleLR scheduling was used.</p>\n<h2>Efficientnet B7:</h2></li>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>640 x 640 image size</p></li>\n<li><p>average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with similar configs as of the current model.</p></li>\n<li><p>AdamW optimizer with OneCycleLR scheduling was used.</p></li>\n</ul>\n<p><strong>As described above for models individually, some of the strategies which were quite common in our models and gave us a good amount of boost were:</strong></p>\n<ol>\n<li>Noisy Student training</li>\n<li>Horizontal Flip Test Time Augmentation</li>\n<li>Attention Head</li>\n<li>Auxilliary Loss using Segmentation Masks</li>\n<li>Fine-tuning</li>\n</ol>\n<h2>Image Level Solution { Binary Classification }:</h2>\n<p>Very Similar to study level models, our binary model was trained with Efficientnet B6. Again our model was trained with imagenet weights without any pretraining on <strong>external data</strong>. <br>\nFor None predictions, we noticed that if duplicates are ignored, all none predictions were the same as the “Negative For Pneumonia” Class which was in study predictions.</p>\n<p>So, our final binary predictions were a weighted average of Efficientnet binary predictions and study-based Efficientnet- v2m(5 fold) <strong>Negative for Pneumonia</strong> predictions whose training was explained above in the study level solution.</p>\n<p><a href=\"https://ibb.co/7g7d83g\"><img src=\"https://i.ibb.co/3ft9Gxf/Screenshot-2021-08-19-at-7-08-09-PM.png\" alt=\"Screenshot-2021-08-19-at-7-08-09-PM\"></a></p>\n<h2>Efficientnet B6:</h2>\n<ul>\n<li>Pretrained imagenet weights were used from timm models. </li>\n<li>512 x 512 image size</li>\n<li>average pooling was used in the training and finetune stage.</li>\n<li>We used two separate segmentation heads for binary model efficientnet B6(1st after 3rd Block, 2nd after 6th Block )</li>\n<li>Loss fnc: BCE{Binary} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss1} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss2} </li>\n<li>Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with the same configs.</li>\n<li>Ranger optimizer, Cosine Scheduler with warmup were used.</li>\n</ul>\n<h2>Classification models based Ensemble{None + multiclass}:</h2>\n<p><a href=\"https://ibb.co/ZxM4G6N\"><img src=\"https://i.ibb.co/pdLVbvn/Screenshot-2021-08-19-at-6-08-45-PM.png\" alt=\"Screenshot-2021-08-19-at-6-08-45-PM\"></a></p>\n<p><strong>Study Level:</strong> We simply took the mean of 11 models predictions for each class based on their fold-wise results and ensemble boost. Further, they were blended with efficientnet v2m(5 fold) with weights 0.85 - 0.15.</p>\n<p><strong>Image Level:</strong> For none predictions, as mentioned before we took the weighted average of Efficientnet B6( trained on Binary Classification) and Efficientnet v2m (same as study model). </p>\n<p>Weights for all the ensembles mentioned above were solely determined by best validation score and diversity.</p>\n<h2>Image Level Solution{ Object Detection }:</h2>\n<p><a href=\"https://ibb.co/WHypL7j\"><img src=\"https://i.ibb.co/x2j8Nwd/Screenshot-2021-08-19-at-6-10-01-PM.png\" alt=\"Screenshot-2021-08-19-at-6-10-01-PM\"></a></p>\n<p>For the object detection part, our final solution used five models(5 fold each), all having different baseline architecture.</p>\n<p><strong><em>Summary of each object detection model:</em></strong></p>\n<p><strong>Efficientnet - D5:</strong> It was trained on just training data. The image size used was 512 x 512. It was trained in two stages, in the second stage it was finetuned with a lower learning rate. The exponential moving average(EMA) was also used in this model’s training.</p>\n<p><strong>Efficientnet - D3:</strong> It was trained on the training data + public test data pseudo labels generated from an ensemble of decent scoring object detection models. Image size used was the default for efficientdet D3 which is 896 x 896. It was also trained in two stages as efficientdet D5. The exponential moving average(EMA) was also used in this model’s training.</p>\n<p><strong>Yolo - v5l6:</strong> It was trained on the training data + public test data pseudo labels. The image size used was 640 x 640 for training. Some images without Bounding Boxes were also included in training data( 20% ). </p>\n<p><strong>Yolo - v5x:</strong> It was trained on the training data + public test data pseudo labels. Image size used was the default for Yolo-v5x which is 640 x 640. Some images without Bounding Boxes were also included in training data( 20% ).</p>\n<p><strong>RetinaNet:</strong> The backbone used for retinanet was resnext101_64x4d. Image used was <br>\n(1333,800). Pseudo labels weren’t used, only training data with bounding boxes were used in the training part.</p>\n<h2>Post Processing</h2>\n<p><a href=\"https://ibb.co/fG06QP7\"><img src=\"https://i.ibb.co/yBsvVbJ/Screenshot-2021-08-19-at-6-12-53-PM.png\" alt=\"Screenshot-2021-08-19-at-6-12-53-PM\"></a></p>\n<p>Post-processing gave us a great amount of boost in both oof score as well as a leaderboard ( + 0.004).<br>\nAs we didn’t use any end-to-end solution and trained models for each level separately,  <br>\nthe idea behind using post-processing was to somehow consume both binary predictions and detection model predictions in a form of ensemble. <br>\nDue to the high amount of diversity in both types of predictions, we were able to achieve that boost. <br>\nWe tried several types of merge both predictions which include mean, weighted average, geometric mean, etc. Out of which power ensemble outperformed everyone in both public leaderboard and validation score. Therefore, we planned to stick with the <strong><em>power ensemble</em></strong> in final submissions. </p>\n<h2>Refrences</h2>\n<p>Some of the key references which helped us in the competition.</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240233\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240233</a> by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a><br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230</a> by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a><br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616#1241561\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616#1241561</a> by <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a><br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/239918\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/239918</a> by <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> <br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a> by hosts <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a></p>\n<p>Thank you all, I've tried my best to cover everything we did in the competition in just one topic. I hope it would be a help to you.</p>",
  "messages": [
    {
      "id": "1481564",
      "postDate": "08/19/2021 15:25:29",
      "content": "<p>Hi everyone, sorry for taking a bit longer to publish our complete solution. It took us several days to clean the code files, update the GitHub repo, and finishing with the write-up.<br>\nThanks to the Kaggle team and SIIM, FIBASIO, RSNA, and all the sponsors who hosted this challenging and interesting covid19 classification and detection competition! Also, thanks to my teammates <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> &amp; <a href=\"https://www.kaggle.com/benihime91\" target=\"_blank\">@benihime91</a> without whom it wouldn't have been possible.</p>\n<p>Representing the  <strong><em>Ayushman Nischay Shivam</em></strong> team, in this post I’m going to explain our winning solution in detail.</p>\n<p>Also, we have already made our inference notebook along with model weights public: you may visit that with this <a href=\"https://www.kaggle.com/nischaydnk/604e8587410a-v2m-bin-weighted\" target=\"_blank\">link</a></p>\n<p>Our all codes related to training or preprocessing data codes are also made public: <a href=\"https://github.com/benihime91/SIIM-COVID19-DETECTION-KAGGLE\" target=\"_blank\">https://github.com/benihime91/SIIM-COVID19-DETECTION-KAGGLE</a></p>\n<p>As <a href=\"https://www.kaggle.com/benihime91\" target=\"_blank\">@benihime91</a> has also talked and summarised a bit about our solution in this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263945\" target=\"_blank\">post</a> , I will try to cover up everything in detail.</p>\n<h2>Overview</h2>\n<p><strong>Our winning blend consists of :</strong></p>\n<p>11 multiclass classification models with 4 different architectures.<br>\n2 x 5 fold (10) binary classification models with 2 different architectures.<br>\n5 x 5 fold (25) object detection models with 5 different architectures.</p>\n<p><a href=\"https://ibb.co/0DcqQwR\"><img src=\"https://i.ibb.co/p0x2nmB/Screenshot-2021-08-19-at-8-27-40-PM.png\" alt=\"Screenshot-2021-08-19-at-8-27-40-PM\"></a></p>\n<h2>Image Data Used</h2>\n<p>We only used competition <a href=\"https://www.kaggle.com/c/siim-covid19-detection/data\" target=\"_blank\">data</a> for training models. <em>No external image data was used.</em></p>\n<h1>Models Summary</h1>\n<h2>Study Level Models</h2>\n<p>All of our study models were trained with variants of Efficientnet models. Other architectures like Resnet, Densenet, transformer-based models didn’t perform well for us. Our models were pretrained on imagenet and didn’t use any external x-ray data directly / indirectly during the competition. Some models were trained on multiple stages which includes finetuning with a reduced learning rate or increasing the image size.</p>\n<p>Baseline Architectures used in our final study level solution:</p>\n<ul>\n<li>Efficientnet v2m</li>\n<li>Efficientnet v2l</li>\n<li>Efficientnet B5</li>\n<li>Efficientnet B7</li>\n</ul>\n<p><a href=\"https://ibb.co/7nMKjnN\"><img src=\"https://i.ibb.co/T4jtY4q/Screenshot-2021-08-19-at-5-50-19-PM.png\" alt=\"Screenshot-2021-08-19-at-5-50-19-PM\"></a></p>\n<h2>Efficientnet v2m:</h2>\n<ul>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>512 x 512(5- fold) &amp; 640x640(fold 0) &amp; 1024x1024(fold 0) image size.</p></li>\n<li><p>PCAM pooling + SAM attention map used in stage2</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were generated with PCAM pooling + DANET attention map in stage 1.</p></li>\n<li><p>Ranger optimizer, Cosine Scheduler with warmup were used.</p></li>\n<li><p>Activation layers of the model were replaced with Mish activation</p>\n<h2>Efficientnet v2l:</h2></li>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>512 x 512(5 folds) &amp; 640x640 (fold 1) image size</p></li>\n<li><p>PCAM pooling + SAM attention map used in training and finetune stage.</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were introduced using the out of folds predictions of the Efficientnet v2m model mentioned above.</p></li>\n<li><p>Ranger optimizer, Cosine Scheduler with warmup were used.</p></li>\n<li><p>Activation layers of the model were replaced with Mish activation</p></li>\n</ul>\n<h2>Efficientnet B5:</h2>\n<ul>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>640 x 640 image size</p></li>\n<li><p>average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were introduced using the out of folds predictions of the efficientnet B5 model with similar configs.</p></li>\n<li><p>AdamW optimizer with OneCycleLR scheduling was used.</p>\n<h2>Efficientnet B7:</h2></li>\n<li><p>Pretrained imagenet weights were used from timm models. </p></li>\n<li><p>640 x 640 image size</p></li>\n<li><p>average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.</p></li>\n<li><p>Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} </p></li>\n<li><p>Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with similar configs as of the current model.</p></li>\n<li><p>AdamW optimizer with OneCycleLR scheduling was used.</p></li>\n</ul>\n<p><strong>As described above for models individually, some of the strategies which were quite common in our models and gave us a good amount of boost were:</strong></p>\n<ol>\n<li>Noisy Student training</li>\n<li>Horizontal Flip Test Time Augmentation</li>\n<li>Attention Head</li>\n<li>Auxilliary Loss using Segmentation Masks</li>\n<li>Fine-tuning</li>\n</ol>\n<h2>Image Level Solution { Binary Classification }:</h2>\n<p>Very Similar to study level models, our binary model was trained with Efficientnet B6. Again our model was trained with imagenet weights without any pretraining on <strong>external data</strong>. <br>\nFor None predictions, we noticed that if duplicates are ignored, all none predictions were the same as the “Negative For Pneumonia” Class which was in study predictions.</p>\n<p>So, our final binary predictions were a weighted average of Efficientnet binary predictions and study-based Efficientnet- v2m(5 fold) <strong>Negative for Pneumonia</strong> predictions whose training was explained above in the study level solution.</p>\n<p><a href=\"https://ibb.co/7g7d83g\"><img src=\"https://i.ibb.co/3ft9Gxf/Screenshot-2021-08-19-at-7-08-09-PM.png\" alt=\"Screenshot-2021-08-19-at-7-08-09-PM\"></a></p>\n<h2>Efficientnet B6:</h2>\n<ul>\n<li>Pretrained imagenet weights were used from timm models. </li>\n<li>512 x 512 image size</li>\n<li>average pooling was used in the training and finetune stage.</li>\n<li>We used two separate segmentation heads for binary model efficientnet B6(1st after 3rd Block, 2nd after 6th Block )</li>\n<li>Loss fnc: BCE{Binary} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss1} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss2} </li>\n<li>Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with the same configs.</li>\n<li>Ranger optimizer, Cosine Scheduler with warmup were used.</li>\n</ul>\n<h2>Classification models based Ensemble{None + multiclass}:</h2>\n<p><a href=\"https://ibb.co/ZxM4G6N\"><img src=\"https://i.ibb.co/pdLVbvn/Screenshot-2021-08-19-at-6-08-45-PM.png\" alt=\"Screenshot-2021-08-19-at-6-08-45-PM\"></a></p>\n<p><strong>Study Level:</strong> We simply took the mean of 11 models predictions for each class based on their fold-wise results and ensemble boost. Further, they were blended with efficientnet v2m(5 fold) with weights 0.85 - 0.15.</p>\n<p><strong>Image Level:</strong> For none predictions, as mentioned before we took the weighted average of Efficientnet B6( trained on Binary Classification) and Efficientnet v2m (same as study model). </p>\n<p>Weights for all the ensembles mentioned above were solely determined by best validation score and diversity.</p>\n<h2>Image Level Solution{ Object Detection }:</h2>\n<p><a href=\"https://ibb.co/WHypL7j\"><img src=\"https://i.ibb.co/x2j8Nwd/Screenshot-2021-08-19-at-6-10-01-PM.png\" alt=\"Screenshot-2021-08-19-at-6-10-01-PM\"></a></p>\n<p>For the object detection part, our final solution used five models(5 fold each), all having different baseline architecture.</p>\n<p><strong><em>Summary of each object detection model:</em></strong></p>\n<p><strong>Efficientnet - D5:</strong> It was trained on just training data. The image size used was 512 x 512. It was trained in two stages, in the second stage it was finetuned with a lower learning rate. The exponential moving average(EMA) was also used in this model’s training.</p>\n<p><strong>Efficientnet - D3:</strong> It was trained on the training data + public test data pseudo labels generated from an ensemble of decent scoring object detection models. Image size used was the default for efficientdet D3 which is 896 x 896. It was also trained in two stages as efficientdet D5. The exponential moving average(EMA) was also used in this model’s training.</p>\n<p><strong>Yolo - v5l6:</strong> It was trained on the training data + public test data pseudo labels. The image size used was 640 x 640 for training. Some images without Bounding Boxes were also included in training data( 20% ). </p>\n<p><strong>Yolo - v5x:</strong> It was trained on the training data + public test data pseudo labels. Image size used was the default for Yolo-v5x which is 640 x 640. Some images without Bounding Boxes were also included in training data( 20% ).</p>\n<p><strong>RetinaNet:</strong> The backbone used for retinanet was resnext101_64x4d. Image used was <br>\n(1333,800). Pseudo labels weren’t used, only training data with bounding boxes were used in the training part.</p>\n<h2>Post Processing</h2>\n<p><a href=\"https://ibb.co/fG06QP7\"><img src=\"https://i.ibb.co/yBsvVbJ/Screenshot-2021-08-19-at-6-12-53-PM.png\" alt=\"Screenshot-2021-08-19-at-6-12-53-PM\"></a></p>\n<p>Post-processing gave us a great amount of boost in both oof score as well as a leaderboard ( + 0.004).<br>\nAs we didn’t use any end-to-end solution and trained models for each level separately,  <br>\nthe idea behind using post-processing was to somehow consume both binary predictions and detection model predictions in a form of ensemble. <br>\nDue to the high amount of diversity in both types of predictions, we were able to achieve that boost. <br>\nWe tried several types of merge both predictions which include mean, weighted average, geometric mean, etc. Out of which power ensemble outperformed everyone in both public leaderboard and validation score. Therefore, we planned to stick with the <strong><em>power ensemble</em></strong> in final submissions. </p>\n<h2>Refrences</h2>\n<p>Some of the key references which helped us in the competition.</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240233\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240233</a> by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a><br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230</a> by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a><br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616#1241561\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616#1241561</a> by <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a><br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/239918\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/239918</a> by <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> <br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a> by hosts <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a></p>\n<p>Thank you all, I've tried my best to cover everything we did in the competition in just one topic. I hope it would be a help to you.</p>",
      "rawMarkdown": "Hi everyone, sorry for taking a bit longer to publish our complete solution. It took us several days to clean the code files, update the GitHub repo, and finishing with the write-up.\nThanks to the Kaggle team and SIIM, FIBASIO, RSNA, and all the sponsors who hosted this challenging and interesting covid19 classification and detection competition! Also, thanks to my teammates @shivamcyborg & @benihime91 without whom it wouldn't have been possible.\n\n\nRepresenting the  ***Ayushman Nischay Shivam*** team, in this post I’m going to explain our winning solution in detail.\n\nAlso, we have already made our inference notebook along with model weights public: you may visit that with this [link](https://www.kaggle.com/nischaydnk/604e8587410a-v2m-bin-weighted)\n\nOur all codes related to training or preprocessing data codes are also made public: https://github.com/benihime91/SIIM-COVID19-DETECTION-KAGGLE\n\nAs @benihime91 has also talked and summarised a bit about our solution in this [post](https://www.kaggle.com/c/siim-covid19-detection/discussion/263945) , I will try to cover up everything in detail.\n\n## Overview\n\n**Our winning blend consists of :**\n\n11 multiclass classification models with 4 different architectures.\n2 x 5 fold (10) binary classification models with 2 different architectures.\n5 x 5 fold (25) object detection models with 5 different architectures.\n\n<a href=\"https://ibb.co/0DcqQwR\"><img src=\"https://i.ibb.co/p0x2nmB/Screenshot-2021-08-19-at-8-27-40-PM.png\" alt=\"Screenshot-2021-08-19-at-8-27-40-PM\" border=\"0\"></a>\n\n\n## Image Data Used\n\nWe only used competition [data](https://www.kaggle.com/c/siim-covid19-detection/data) for training models. *No external image data was used.*\n\n# Models Summary\n\n\n## Study Level Models \n   \n   \nAll of our study models were trained with variants of Efficientnet models. Other architectures like Resnet, Densenet, transformer-based models didn’t perform well for us. Our models were pretrained on imagenet and didn’t use any external x-ray data directly / indirectly during the competition. Some models were trained on multiple stages which includes finetuning with a reduced learning rate or increasing the image size.\n\nBaseline Architectures used in our final study level solution:\n\n   - Efficientnet v2m\n   - Efficientnet v2l\n   - Efficientnet B5\n   - Efficientnet B7\n\n   \n<a href=\"https://ibb.co/7nMKjnN\"><img src=\"https://i.ibb.co/T4jtY4q/Screenshot-2021-08-19-at-5-50-19-PM.png\" alt=\"Screenshot-2021-08-19-at-5-50-19-PM\" border=\"0\"></a>\n\n## Efficientnet v2m: \n\n- Pretrained imagenet weights were used from timm models. \n- 512 x 512(5- fold) & 640x640(fold 0) & 1024x1024(fold 0) image size.\n- PCAM pooling + SAM attention map used in stage2\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were generated with PCAM pooling + DANET attention map in stage 1.\n- Ranger optimizer, Cosine Scheduler with warmup were used.\n- Activation layers of the model were replaced with Mish activation\n\n\n ## Efficientnet v2l:\n- Pretrained imagenet weights were used from timm models. \n- 512 x 512(5 folds) & 640x640 (fold 1) image size\n- PCAM pooling + SAM attention map used in training and finetune stage.\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were introduced using the out of folds predictions of the Efficientnet v2m model mentioned above.\n- Ranger optimizer, Cosine Scheduler with warmup were used.\n- Activation layers of the model were replaced with Mish activation\n\n\n## Efficientnet B5: \n- Pretrained imagenet weights were used from timm models. \n- 640 x 640 image size\n- average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were introduced using the out of folds predictions of the efficientnet B5 model with similar configs.\n- AdamW optimizer with OneCycleLR scheduling was used.\n\n\n ## Efficientnet B7:\n- Pretrained imagenet weights were used from timm models. \n- 640 x 640 image size\n- average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with similar configs as of the current model.\n- AdamW optimizer with OneCycleLR scheduling was used.\n\n\n**As described above for models individually, some of the strategies which were quite common in our models and gave us a good amount of boost were:**\n\n  1. Noisy Student training\n  2. Horizontal Flip Test Time Augmentation\n  3. Attention Head\n  4. Auxilliary Loss using Segmentation Masks\n  5. Fine-tuning\n\n\n## Image Level Solution { Binary Classification }:\n\nVery Similar to study level models, our binary model was trained with Efficientnet B6. Again our model was trained with imagenet weights without any pretraining on **external data**. \nFor None predictions, we noticed that if duplicates are ignored, all none predictions were the same as the “Negative For Pneumonia” Class which was in study predictions.\n\nSo, our final binary predictions were a weighted average of Efficientnet binary predictions and study-based Efficientnet- v2m(5 fold) **Negative for Pneumonia** predictions whose training was explained above in the study level solution.\n\n<a href=\"https://ibb.co/7g7d83g\"><img src=\"https://i.ibb.co/3ft9Gxf/Screenshot-2021-08-19-at-7-08-09-PM.png\" alt=\"Screenshot-2021-08-19-at-7-08-09-PM\" border=\"0\"></a>\n\n\n## Efficientnet B6:\n\n - Pretrained imagenet weights were used from timm models. \n - 512 x 512 image size\n - average pooling was used in the training and finetune stage.\n - We used two separate segmentation heads for binary model efficientnet B6(1st after 3rd Block, 2nd after 6th Block )\n - Loss fnc: BCE{Binary} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss1} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss2} \n - Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with the same configs.\n - Ranger optimizer, Cosine Scheduler with warmup were used.\n\n\n## Classification models based Ensemble{None + multiclass}:\n\n<a href=\"https://ibb.co/ZxM4G6N\"><img src=\"https://i.ibb.co/pdLVbvn/Screenshot-2021-08-19-at-6-08-45-PM.png\" alt=\"Screenshot-2021-08-19-at-6-08-45-PM\" border=\"0\"></a>\n\n\n**Study Level:** We simply took the mean of 11 models predictions for each class based on their fold-wise results and ensemble boost. Further, they were blended with efficientnet v2m(5 fold) with weights 0.85 - 0.15.\n\n**Image Level:** For none predictions, as mentioned before we took the weighted average of Efficientnet B6( trained on Binary Classification) and Efficientnet v2m (same as study model). \n\nWeights for all the ensembles mentioned above were solely determined by best validation score and diversity.\n\n\n## Image Level Solution{ Object Detection }:\n\n\n<a href=\"https://ibb.co/WHypL7j\"><img src=\"https://i.ibb.co/x2j8Nwd/Screenshot-2021-08-19-at-6-10-01-PM.png\" alt=\"Screenshot-2021-08-19-at-6-10-01-PM\" border=\"0\"></a>\n\nFor the object detection part, our final solution used five models(5 fold each), all having different baseline architecture.\n\n\n***Summary of each object detection model:***\n\n**Efficientnet - D5:** It was trained on just training data. The image size used was 512 x 512. It was trained in two stages, in the second stage it was finetuned with a lower learning rate. The exponential moving average(EMA) was also used in this model’s training.\n\n\n**Efficientnet - D3:** It was trained on the training data + public test data pseudo labels generated from an ensemble of decent scoring object detection models. Image size used was the default for efficientdet D3 which is 896 x 896. It was also trained in two stages as efficientdet D5. The exponential moving average(EMA) was also used in this model’s training.\n\n\n**Yolo - v5l6:** It was trained on the training data + public test data pseudo labels. The image size used was 640 x 640 for training. Some images without Bounding Boxes were also included in training data( 20% ). \n\n\n**Yolo - v5x:** It was trained on the training data + public test data pseudo labels. Image size used was the default for Yolo-v5x which is 640 x 640. Some images without Bounding Boxes were also included in training data( 20% ).\n\n\n**RetinaNet:** The backbone used for retinanet was resnext101_64x4d. Image used was \n(1333,800). Pseudo labels weren’t used, only training data with bounding boxes were used in the training part.\n\n\n\n## Post Processing\n\n<a href=\"https://ibb.co/fG06QP7\"><img src=\"https://i.ibb.co/yBsvVbJ/Screenshot-2021-08-19-at-6-12-53-PM.png\" alt=\"Screenshot-2021-08-19-at-6-12-53-PM\" border=\"0\"></a>\n\n\nPost-processing gave us a great amount of boost in both oof score as well as a leaderboard ( + 0.004).\nAs we didn’t use any end-to-end solution and trained models for each level separately,  \nthe idea behind using post-processing was to somehow consume both binary predictions and detection model predictions in a form of ensemble. \nDue to the high amount of diversity in both types of predictions, we were able to achieve that boost. \nWe tried several types of merge both predictions which include mean, weighted average, geometric mean, etc. Out of which power ensemble outperformed everyone in both public leaderboard and validation score. Therefore, we planned to stick with the ***power ensemble*** in final submissions. \n\n \n## Refrences\n\nSome of the key references which helped us in the competition.\n\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/240233 by @hengck23\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230 by @ttahara\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616#1241561 by @moewie94\nhttps://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637 by @cdeotte\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/239918 by @xhlulu \nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/246597 by hosts @paras42\n\nThank you all, I've tried my best to cover everything we did in the competition in just one topic. I hope it would be a help to you.",
      "votes": null
    },
    {
      "id": "1481723",
      "postDate": "08/19/2021 16:37:32",
      "content": "<p>Thanks for the explanation ! </p>",
      "rawMarkdown": "Thanks for the explanation !",
      "votes": null
    },
    {
      "id": "1481788",
      "postDate": "08/19/2021 17:02:45",
      "content": "<p>My pleasure :)</p>",
      "rawMarkdown": "My pleasure :)",
      "votes": null
    },
    {
      "id": "1481879",
      "postDate": "08/19/2021 17:49:50",
      "content": "<p>Quite Informative 👍<br>\nThanks</p>",
      "rawMarkdown": "Quite Informative 👍\nThanks",
      "votes": null
    },
    {
      "id": "1485902",
      "postDate": "08/22/2021 13:53:37",
      "content": "<p>Congratulations ! Your solution is clear and well written ! </p>\n<p>Thank you for taking the time to share it 👏</p>",
      "rawMarkdown": "Congratulations ! Your solution is clear and well written ! \n\nThank you for taking the time to share it 👏",
      "votes": null
    },
    {
      "id": "1488132",
      "postDate": "08/24/2021 05:59:10",
      "content": "<p>I'm glad you liked it, thank you.</p>",
      "rawMarkdown": "I'm glad you liked it, thank you.",
      "votes": null
    },
    {
      "id": "1488137",
      "postDate": "08/24/2021 06:01:35",
      "content": "<p>Congratulations! This probably is a lot of work. I just have one question though. How do you assign the ensemble's statistical weights? It seems like a that you started with 0.2 confidence and then shifted the numbers using regression, right?</p>",
      "rawMarkdown": "Congratulations! This probably is a lot of work. I just have one question though. How do you assign the ensemble's statistical weights? It seems like a that you started with 0.2 confidence and then shifted the numbers using regression, right?",
      "votes": null
    },
    {
      "id": "1488309",
      "postDate": "08/24/2021 08:09:09",
      "content": "<p>Hi, for giving weights in an ensemble. We considered several factors which include single models public leaderboard score, oof score, and correlation with other models. <br>\nThe weights mentioned above were first optimized to maximize the overall oof after ensemble, then I tweaked them a little based on their public leaderboard score. </p>",
      "rawMarkdown": "Hi, for giving weights in an ensemble. We considered several factors which include single models public leaderboard score, oof score, and correlation with other models. \nThe weights mentioned above were first optimized to maximize the overall oof after ensemble, then I tweaked them a little based on their public leaderboard score.",
      "votes": null
    },
    {
      "id": "1492367",
      "postDate": "08/27/2021 05:55:34",
      "content": "<p>Congratulations 🎉!! Well written and honestly simpler than the 40+ ensemble solutions 😊.</p>",
      "rawMarkdown": "Congratulations 🎉!! Well written and honestly simpler than the 40+ ensemble solutions 😊.",
      "votes": null
    },
    {
      "id": "1494290",
      "postDate": "08/28/2021 15:00:48",
      "content": "<p>congrats ! one question tho, how did you generate this csv 'df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv'?<br>\nthanks</p>",
      "rawMarkdown": "congrats ! one question tho, how did you generate this csv 'df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv'?\nthanks",
      "votes": null
    },
    {
      "id": "1494393",
      "postDate": "08/28/2021 16:29:06",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sparkyjunior\" target=\"_blank\">@sparkyjunior</a> so what we basically did was we trained <code>efficientnetb5</code> , generated OOF predictions. We then took these OOF predictions of <code>efficientnetb5</code> and trainined <code>efficientnetb6</code> using noisy student training.  </p>\n<p>Finally, we took the <code>negative</code> predictions from both the trained models <code>(b5+b6)/2</code>.  This becomes the <code>df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv</code></p>\n<p>We then take these predictions and take them as the OOF predictions for the <code>none</code> class and then again train another <code>efficientnetb6</code> using noisy student training method with the above <code>none</code> OOFs and original  <code>none</code> labels, which becomes one of the models in our binary ensemble.</p>\n<p>I know it's a bit confusing 😂</p>",
      "rawMarkdown": "Hey @sparkyjunior so what we basically did was we trained `efficientnetb5` , generated OOF predictions. We then took these OOF predictions of `efficientnetb5` and trainined `efficientnetb6` using noisy student training.  \n\nFinally, we took the `negative` predictions from both the trained models `(b5+b6)/2`.  This becomes the `df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv`\n\nWe then take these predictions and take them as the OOF predictions for the `none` class and then again train another `efficientnetb6` using noisy student training method with the above `none` OOFs and original  `none` labels, which becomes one of the models in our binary ensemble.\n\nI know it's a bit confusing 😂",
      "votes": null
    },
    {
      "id": "1494469",
      "postDate": "08/28/2021 17:36:27",
      "content": "<p>well thats one way to explain it. thanks man 😂❤️</p>",
      "rawMarkdown": "well thats one way to explain it. thanks man 😂❤️",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1481723,
      "author_name": "akpmpr",
      "author_url": "",
      "post_date": "08/19/2021 16:37:32",
      "content": "<p>Thanks for the explanation ! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1481788,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "08/19/2021 17:02:45",
          "content": "<p>My pleasure :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1481879,
      "author_name": "yashguptansut",
      "author_url": "",
      "post_date": "08/19/2021 17:49:50",
      "content": "<p>Quite Informative 👍<br>\nThanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1485902,
      "author_name": "tomdarmon",
      "author_url": "",
      "post_date": "08/22/2021 13:53:37",
      "content": "<p>Congratulations ! Your solution is clear and well written ! </p>\n<p>Thank you for taking the time to share it 👏</p>",
      "votes": null,
      "replies": [
        {
          "id": 1488132,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "08/24/2021 05:59:10",
          "content": "<p>I'm glad you liked it, thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1488137,
      "author_name": "thanhdvan",
      "author_url": "",
      "post_date": "08/24/2021 06:01:35",
      "content": "<p>Congratulations! This probably is a lot of work. I just have one question though. How do you assign the ensemble's statistical weights? It seems like a that you started with 0.2 confidence and then shifted the numbers using regression, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1488309,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "08/24/2021 08:09:09",
          "content": "<p>Hi, for giving weights in an ensemble. We considered several factors which include single models public leaderboard score, oof score, and correlation with other models. <br>\nThe weights mentioned above were first optimized to maximize the overall oof after ensemble, then I tweaked them a little based on their public leaderboard score. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1492367,
      "author_name": "sauravmaheshkar",
      "author_url": "",
      "post_date": "08/27/2021 05:55:34",
      "content": "<p>Congratulations 🎉!! Well written and honestly simpler than the 40+ ensemble solutions 😊.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1494290,
      "author_name": "sparkyjunior",
      "author_url": "",
      "post_date": "08/28/2021 15:00:48",
      "content": "<p>congrats ! one question tho, how did you generate this csv 'df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv'?<br>\nthanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1494393,
          "author_name": "benihime91",
          "author_url": "",
          "post_date": "08/28/2021 16:29:06",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/sparkyjunior\" target=\"_blank\">@sparkyjunior</a> so what we basically did was we trained <code>efficientnetb5</code> , generated OOF predictions. We then took these OOF predictions of <code>efficientnetb5</code> and trainined <code>efficientnetb6</code> using noisy student training.  </p>\n<p>Finally, we took the <code>negative</code> predictions from both the trained models <code>(b5+b6)/2</code>.  This becomes the <code>df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv</code></p>\n<p>We then take these predictions and take them as the OOF predictions for the <code>none</code> class and then again train another <code>efficientnetb6</code> using noisy student training method with the above <code>none</code> OOFs and original  <code>none</code> labels, which becomes one of the models in our binary ensemble.</p>\n<p>I know it's a bit confusing 😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494469,
          "author_name": "sparkyjunior",
          "author_url": "",
          "post_date": "08/28/2021 17:36:27",
          "content": "<p>well thats one way to explain it. thanks man 😂❤️</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1481564": "Hi everyone, sorry for taking a bit longer to publish our complete solution. It took us several days to clean the code files, update the GitHub repo, and finishing with the write-up.\nThanks to the Kaggle team and SIIM, FIBASIO, RSNA, and all the sponsors who hosted this challenging and interesting covid19 classification and detection competition! Also, thanks to my teammates @shivamcyborg & @benihime91 without whom it wouldn't have been possible.\n\n\nRepresenting the  ***Ayushman Nischay Shivam*** team, in this post I’m going to explain our winning solution in detail.\n\nAlso, we have already made our inference notebook along with model weights public: you may visit that with this [link](https://www.kaggle.com/nischaydnk/604e8587410a-v2m-bin-weighted)\n\nOur all codes related to training or preprocessing data codes are also made public: https://github.com/benihime91/SIIM-COVID19-DETECTION-KAGGLE\n\nAs @benihime91 has also talked and summarised a bit about our solution in this [post](https://www.kaggle.com/c/siim-covid19-detection/discussion/263945) , I will try to cover up everything in detail.\n\n## Overview\n\n**Our winning blend consists of :**\n\n11 multiclass classification models with 4 different architectures.\n2 x 5 fold (10) binary classification models with 2 different architectures.\n5 x 5 fold (25) object detection models with 5 different architectures.\n\n<a href=\"https://ibb.co/0DcqQwR\"><img src=\"https://i.ibb.co/p0x2nmB/Screenshot-2021-08-19-at-8-27-40-PM.png\" alt=\"Screenshot-2021-08-19-at-8-27-40-PM\" border=\"0\"></a>\n\n\n## Image Data Used\n\nWe only used competition [data](https://www.kaggle.com/c/siim-covid19-detection/data) for training models. *No external image data was used.*\n\n# Models Summary\n\n\n## Study Level Models \n   \n   \nAll of our study models were trained with variants of Efficientnet models. Other architectures like Resnet, Densenet, transformer-based models didn’t perform well for us. Our models were pretrained on imagenet and didn’t use any external x-ray data directly / indirectly during the competition. Some models were trained on multiple stages which includes finetuning with a reduced learning rate or increasing the image size.\n\nBaseline Architectures used in our final study level solution:\n\n   - Efficientnet v2m\n   - Efficientnet v2l\n   - Efficientnet B5\n   - Efficientnet B7\n\n   \n<a href=\"https://ibb.co/7nMKjnN\"><img src=\"https://i.ibb.co/T4jtY4q/Screenshot-2021-08-19-at-5-50-19-PM.png\" alt=\"Screenshot-2021-08-19-at-5-50-19-PM\" border=\"0\"></a>\n\n## Efficientnet v2m: \n\n- Pretrained imagenet weights were used from timm models. \n- 512 x 512(5- fold) & 640x640(fold 0) & 1024x1024(fold 0) image size.\n- PCAM pooling + SAM attention map used in stage2\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were generated with PCAM pooling + DANET attention map in stage 1.\n- Ranger optimizer, Cosine Scheduler with warmup were used.\n- Activation layers of the model were replaced with Mish activation\n\n\n ## Efficientnet v2l:\n- Pretrained imagenet weights were used from timm models. \n- 512 x 512(5 folds) & 640x640 (fold 1) image size\n- PCAM pooling + SAM attention map used in training and finetune stage.\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were introduced using the out of folds predictions of the Efficientnet v2m model mentioned above.\n- Ranger optimizer, Cosine Scheduler with warmup were used.\n- Activation layers of the model were replaced with Mish activation\n\n\n## Efficientnet B5: \n- Pretrained imagenet weights were used from timm models. \n- 640 x 640 image size\n- average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were introduced using the out of folds predictions of the efficientnet B5 model with similar configs.\n- AdamW optimizer with OneCycleLR scheduling was used.\n\n\n ## Efficientnet B7:\n- Pretrained imagenet weights were used from timm models. \n- 640 x 640 image size\n- average pooling + sCSE attention map along with multi-head attention was used in the training and finetune stage.\n- Loss fnc: BCE{4-class} + [0.75* lovasz_loss + 0.25* BCE ]{Segmentation loss} \n- Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with similar configs as of the current model.\n- AdamW optimizer with OneCycleLR scheduling was used.\n\n\n**As described above for models individually, some of the strategies which were quite common in our models and gave us a good amount of boost were:**\n\n  1. Noisy Student training\n  2. Horizontal Flip Test Time Augmentation\n  3. Attention Head\n  4. Auxilliary Loss using Segmentation Masks\n  5. Fine-tuning\n\n\n## Image Level Solution { Binary Classification }:\n\nVery Similar to study level models, our binary model was trained with Efficientnet B6. Again our model was trained with imagenet weights without any pretraining on **external data**. \nFor None predictions, we noticed that if duplicates are ignored, all none predictions were the same as the “Negative For Pneumonia” Class which was in study predictions.\n\nSo, our final binary predictions were a weighted average of Efficientnet binary predictions and study-based Efficientnet- v2m(5 fold) **Negative for Pneumonia** predictions whose training was explained above in the study level solution.\n\n<a href=\"https://ibb.co/7g7d83g\"><img src=\"https://i.ibb.co/3ft9Gxf/Screenshot-2021-08-19-at-7-08-09-PM.png\" alt=\"Screenshot-2021-08-19-at-7-08-09-PM\" border=\"0\"></a>\n\n\n## Efficientnet B6:\n\n - Pretrained imagenet weights were used from timm models. \n - 512 x 512 image size\n - average pooling was used in the training and finetune stage.\n - We used two separate segmentation heads for binary model efficientnet B6(1st after 3rd Block, 2nd after 6th Block )\n - Loss fnc: BCE{Binary} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss1} + [0.5* lovasz_loss + 0.5* BCE ]{Segmentation loss2} \n - Noisy labels were introduced using the out of folds predictions of the efficientnet B6(640 image size) model with the same configs.\n - Ranger optimizer, Cosine Scheduler with warmup were used.\n\n\n## Classification models based Ensemble{None + multiclass}:\n\n<a href=\"https://ibb.co/ZxM4G6N\"><img src=\"https://i.ibb.co/pdLVbvn/Screenshot-2021-08-19-at-6-08-45-PM.png\" alt=\"Screenshot-2021-08-19-at-6-08-45-PM\" border=\"0\"></a>\n\n\n**Study Level:** We simply took the mean of 11 models predictions for each class based on their fold-wise results and ensemble boost. Further, they were blended with efficientnet v2m(5 fold) with weights 0.85 - 0.15.\n\n**Image Level:** For none predictions, as mentioned before we took the weighted average of Efficientnet B6( trained on Binary Classification) and Efficientnet v2m (same as study model). \n\nWeights for all the ensembles mentioned above were solely determined by best validation score and diversity.\n\n\n## Image Level Solution{ Object Detection }:\n\n\n<a href=\"https://ibb.co/WHypL7j\"><img src=\"https://i.ibb.co/x2j8Nwd/Screenshot-2021-08-19-at-6-10-01-PM.png\" alt=\"Screenshot-2021-08-19-at-6-10-01-PM\" border=\"0\"></a>\n\nFor the object detection part, our final solution used five models(5 fold each), all having different baseline architecture.\n\n\n***Summary of each object detection model:***\n\n**Efficientnet - D5:** It was trained on just training data. The image size used was 512 x 512. It was trained in two stages, in the second stage it was finetuned with a lower learning rate. The exponential moving average(EMA) was also used in this model’s training.\n\n\n**Efficientnet - D3:** It was trained on the training data + public test data pseudo labels generated from an ensemble of decent scoring object detection models. Image size used was the default for efficientdet D3 which is 896 x 896. It was also trained in two stages as efficientdet D5. The exponential moving average(EMA) was also used in this model’s training.\n\n\n**Yolo - v5l6:** It was trained on the training data + public test data pseudo labels. The image size used was 640 x 640 for training. Some images without Bounding Boxes were also included in training data( 20% ). \n\n\n**Yolo - v5x:** It was trained on the training data + public test data pseudo labels. Image size used was the default for Yolo-v5x which is 640 x 640. Some images without Bounding Boxes were also included in training data( 20% ).\n\n\n**RetinaNet:** The backbone used for retinanet was resnext101_64x4d. Image used was \n(1333,800). Pseudo labels weren’t used, only training data with bounding boxes were used in the training part.\n\n\n\n## Post Processing\n\n<a href=\"https://ibb.co/fG06QP7\"><img src=\"https://i.ibb.co/yBsvVbJ/Screenshot-2021-08-19-at-6-12-53-PM.png\" alt=\"Screenshot-2021-08-19-at-6-12-53-PM\" border=\"0\"></a>\n\n\nPost-processing gave us a great amount of boost in both oof score as well as a leaderboard ( + 0.004).\nAs we didn’t use any end-to-end solution and trained models for each level separately,  \nthe idea behind using post-processing was to somehow consume both binary predictions and detection model predictions in a form of ensemble. \nDue to the high amount of diversity in both types of predictions, we were able to achieve that boost. \nWe tried several types of merge both predictions which include mean, weighted average, geometric mean, etc. Out of which power ensemble outperformed everyone in both public leaderboard and validation score. Therefore, we planned to stick with the ***power ensemble*** in final submissions. \n\n \n## Refrences\n\nSome of the key references which helped us in the competition.\n\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/240233 by @hengck23\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230 by @ttahara\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616#1241561 by @moewie94\nhttps://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637 by @cdeotte\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/239918 by @xhlulu \nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/246597 by hosts @paras42\n\nThank you all, I've tried my best to cover everything we did in the competition in just one topic. I hope it would be a help to you.",
    "1481723": "Thanks for the explanation !",
    "1481788": "My pleasure :)",
    "1481879": "Quite Informative 👍\nThanks",
    "1485902": "Congratulations ! Your solution is clear and well written ! \n\nThank you for taking the time to share it 👏",
    "1488132": "I'm glad you liked it, thank you.",
    "1488137": "Congratulations! This probably is a lot of work. I just have one question though. How do you assign the ensemble's statistical weights? It seems like a that you started with 0.2 confidence and then shifted the numbers using regression, right?",
    "1488309": "Hi, for giving weights in an ensemble. We considered several factors which include single models public leaderboard score, oof score, and correlation with other models. \nThe weights mentioned above were first optimized to maximize the overall oof after ensemble, then I tweaked them a little based on their public leaderboard score.",
    "1492367": "Congratulations 🎉!! Well written and honestly simpler than the 40+ ensemble solutions 😊.",
    "1494290": "congrats ! one question tho, how did you generate this csv 'df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv'?\nthanks",
    "1494393": "Hey @sparkyjunior so what we basically did was we trained `efficientnetb5` , generated OOF predictions. We then took these OOF predictions of `efficientnetb5` and trainined `efficientnetb6` using noisy student training.  \n\nFinally, we took the `negative` predictions from both the trained models `(b5+b6)/2`.  This becomes the `df_study_split_binary_negative_eb5ns_eb6eb6_ns_4024.csv`\n\nWe then take these predictions and take them as the OOF predictions for the `none` class and then again train another `efficientnetb6` using noisy student training method with the above `none` OOFs and original  `none` labels, which becomes one of the models in our binary ensemble.\n\nI know it's a bit confusing 😂",
    "1494469": "well thats one way to explain it. thanks man 😂❤️"
  },
  "source": "meta"
}