{
  "id": 226618,
  "title": "In Gauss (God) we Trust! 26th place solution",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/writeups/in-gauss-we-trust-in-gauss-god-we-trust-26th-place",
  "author_name": "",
  "post_date": "2021-05-04T11:15:16.060Z",
  "votes": 37,
  "comment_count": 12,
  "views": 0,
  "content": "<h1>Transfer Learning</h1>\n<p>As we all know, if we train on <code>imagenet</code> weights, we may take quite a while to converge, even if we finetune it. The intuition is simple, <code>imagenet</code> were trained on many common items in life, and none of them resemble closely to the image structures of X-rays, therefore, the model may have a hard time detecting shapes and details from the X-rays. We can of course unfreeze all the layers and retrain them from scratch, using the State of the Art models' backbone, however, due to limited hardware, we decided it is best to use what others have trained. After all, it is much easier to stand on the shoulder of giants like <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">ammarali</a>. Consequently, I conveniently used a set of <code>pretrained</code> weights trained specifically on this dataset as a starting point. The weights and ideas can be found <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910\" target=\"_blank\">here</a></strong>.<br>\nWe used a few models and found out that <code>resnet200d</code> has the best results on this set of training images. There is no deep reason why this model outperforms other State of the Art models, but using <code>gradcam</code> we can see how the model sees the images.</p>\n<h2>Cross Validation Strategy</h2>\n<p>Knowing how to choose a robust and leak free cross validation strategy is extremely important. Data leakage can cause you to have blind confidence on your model. We are also guilty of committing one since we trained our models with the NiH pretrained weights, without taking into consideration if the weights overlap with the training and validation folds information.</p>\n<ul>\n<li>Cross Validation Strategy: Multi-Label Stratified Group KFold(K=<strong>5</strong>) - using sin's method to split folds will ensure that there is no leakage. Quoting almost verbatim from <code>sklearn</code> :</li>\n<li>Assuming that some data is Independent and Identically Distributed (i.i.d.) is making the assumption that all samples stem from the same generative process and that the generative process is assumed to have no memory of past generated samples.**</li>\n</ul>\n<h2>Therefore it is paramount to ensure that amongst this 3255 <strong>unique</strong> patients, we need to ensure that each unique patients' images DO NOT appear in the validation fold. That is to say, if patient John Doe has 100 X-ray images, but during our 5-fold splits, he has 70 images in Fold 1-4, while 30 images are in Fold 5, then if we were to train on Fold 1-4 and validate on Fold 5, there may be potential leakage and the model will predict with confidence for John Doe's images. This is under the assumption that John Doe's data does not fulfill the i.i.d process.</h2>\n<h2>Model Architectures, Training Parameters &amp; Augmentations</h2>\n<p>We built-upon Tawara’s Multi-head model for our best scoring models. In particular, we experimented with the activation functions and dropout rates. We found models with <code>Swish</code> activation in the <code>multi-head</code> component of the network to perform best in our experiments. Our best scoring single model is a multi-head model with a <code>resnet200d</code> backbone. In particular, one single fold of <code>resnet200d</code> gives a private score of 0.970. <br>\nWe started experimenting with <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">sin's</a> <a href=\"https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965\" target=\"_blank\">pipeline</a>  which is similar to qishen ha's pipeline back in Melanoma and used <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">Tawara's</a> multihead approach. We did not have time to experiment with the 3-4 stage training as we joined the competition late.</p>\n<ul>\n<li>model:<ul>\n<li><strong>backbone</strong>: <code>ResNet200D</code> and <code>SeResNet152d</code></li>\n<li><strong>classifier/multi-head:</strong> independent&nbsp;<strong>Spatial-Attention Module</strong>&nbsp;and MLP by Target Group(ETT(3), NGT(4), CVC(3), and Swan(1))</li>\n<li><strong>NOTE: I use&nbsp;<a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">the pre-trained model</a>&nbsp;shared by <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> .</strong>&nbsp;Thanks!</li></ul></li>\n</ul>\n<h2>Preprocessing</h2>\n<p><a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">Rueben Schmidt</a> mentioned in this <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146\" target=\"_blank\">post</a> here that some images have black borders around them. I removed them during both the training and inference process. There was no significant increase on the LB score, even if there was, it is in the 3-4th decimal places, but I noticed my local cv to increase. Thus I decided to remove for all. After all, if I keep this consistent in both training and inference, I reckon that no surprise factor would pop out. </p>\n<h2>Augmentation</h2>\n<p>In particular, we made use of a different <code>Normalization</code> parameter which is more accustomed to the X-ray pretrained images. Thanks Tawara again! Heavy augmentations are used during <strong>Train-Time-Augmentation.</strong> But during <strong>Test-Time-Augmentation,</strong> we merely used a <code>HorizontalFlip</code> with 100% probability, and only used <code>tta_steps=1</code>. </p>\n<pre><code>augmentations_class: AlbumentationsAugmentation\naugmentations_train:\n AlbumentationsAugmentation:\n   - name: RandomResizedCrop\n     params:\n       height: 640\n       width: 640\n       scale: [0.9, 1.0]\n       p: 1.0\n   - name: HorizontalFlip\n     params:\n       p: 0.5\n   - name: ShiftScaleRotate\n     params:\n       shift_limit: 0.2\n       scale_limit: 0.2\n       rotate_limit: 20\n       border_mode: 0\n       value: 0\n       mask_value: 0\n       p: 0.5\n   - name: HueSaturationValue\n     params:\n       hue_shift_limit: 10\n       sat_shift_limit: 10\n       val_shift_limit: 10\n       p: 0.7\n   - name: RandomBrightnessContrast\n     params:\n       brightness_limit: [-0.2, 0.2]\n       contrast_limit: [-0.2, 0.2]\n       p: 0.7\n   - name: CLAHE\n     params:\n       clip_limit: [1,4]\n       p: 0.5\n   - name: JpegCompression\n     params:\n       p: 0.2\n   - name: IAAPiecewiseAffine\n     params:\n       p: 0.2\n   - name: IAASharpen\n     params:\n       p: 0.2\n   - name: Cutout\n     params: \n       # use int(image_size * 0.1)\n       max_h_size: 64\n       max_w_size: 64\n       num_holes: 5\n       p: 0.5\n   - name: Resize\n     params:\n       height: 640\n       width: 640\n       p: 1.0\n   - name: Normalize\n     params:\n       mean: [0.4887381077884414]\n       std: [0.23064819430546407]\n       p: 1.0\n   - name: ToTensorV2\n     params:\n       p: 1.0\naugmentations_val:\n AlbumentationsAugmentation:\n   - name: Resize\n     params:\n       height: 640\n       width: 640\n       p: 1.0\n   - name: Normalize\n     params:\n       mean: [0.4887381077884414]\n       std: [0.23064819430546407]\n       p: 1.0\n   - name: ToTensorV2\n     params:\n       p: 1.0\n</code></pre>\n<hr>\n<h2>Batch Size and Tricks</h2>\n<p>Due to hardware limitation, we can barely fit in anything more than a <code>batch_size</code> of 8. We quote the well known fact <a href=\"https://arxiv.org/abs/1609.04836\" target=\"_blank\">here</a>:</p>\n<blockquote>\n  <p>It has been observed in practice that when using a larger batch there is a degradation in the quality of the model, as measured by its ability to generalize […]<br>\n  large-batch methods tend to converge to sharp minimizers of the training and testing functions—and as is well known, sharp minima lead to poorer generalization. In contrast, small-batch methods consistently converge to flat minimizers, and our experiments support a commonly held view that this is due to the inherent noise in the gradient estimation.<br>\n  The above shows that large batch size may <code>fit</code> the model too well, as the model will learn features of the dataset in less iterations, and may memorize this particular dataset's features, leading to overfitting and poor generalization. However, too small a batch size causes our convergence to go too slow, empirically, we take 32 or 64 as the ideal batch size in this competition. </p>\n  <h2>We used both <code>torch.amp</code> and <code>gradient accumulation</code> to be able to fit more batch sizes. We did not freeze the <code>batch_norm</code> layers, which still yielded great results. What we should have done is to experiment more on how to freeze the batch norm layers properly, as I believe that it may help. In the end, we used a batch size of 8 and fit 4 iterations using <code>gradient accumulation</code>  and trained a total number of 20 epochs to get a local CV score of roughly 0.969.</h2>\n  <h2>Optimizer, Scheduler and Loss</h2>\n  <p>Nothing too fancy here, although we really wanted to try out <code>Focal Loss</code> in this setting. The configuration can be seen here. But note that we incorporated <code>GradualWarmUpScheduler</code> along with <code>CosineAnnealingLR</code>.</p>\n</blockquote>\n<pre><code>scheduler: CosineAnnealingLR\nscheduler_params: # Note that in params we must put 1.e instead of 1e\n CosineAnnealingLR:\n   T_max: 16\n   eta_min: 1.e-7\n   last_epoch: -1\n   verbose: True   \ntrain_step_scheduler: False\nval_step_scheduler: False\noptimizer: Adam\noptimizer_params:\n Adam:\n   lr: 0.00002\n   betas:\n     - 0.9\n     - 0.999\n   eps: 1.e-7\n   weight_decay: 0\n   amsgrad: False\ncriterion_train: BCEWithLogitsLoss\ncriterion_val: BCEWithLogitsLoss\ncriterion_params:\n CrossEntropyLoss:\n   weight: null\n   size_average: null\n   ignore_index: -100\n   reduce: null\n   reduction: mean\n LabelSmoothingLoss:\n   classes: 2\n   smoothing: 0.05\n   dim: -1\n</code></pre>\n<h2>Activation Function</h2>\n<pre><code>import torch\nsigmoid = torch.nn.Sigmoid()\nclass Swish(torch.autograd.Function):\n   @staticmethod\n   def forward(ctx, i):\n       result = i * sigmoid(i)\n       ctx.save_for_backward(i)\n       return result\n   @staticmethod\n   def backward(ctx, grad_output):\n       i = ctx.saved_variables[0]\n       sigmoid_i = sigmoid(i)\n       return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\nclass Swish_Module(torch.nn.Module):\n   def forward(self, x):\n       return Swish.apply(x)\n</code></pre>\n<hr>\n<p>Our second best performing model is also a multi-head model with swish activation in the heads, but with a <code>SeResNet152d</code> backbone (<code>seresnet152d</code>).<br>\nDuring training, we use gradient accumulation so that the bath size can scale up eventually to our desired sizes. For the second model, it scales to 16. The training parameters for our second best performing model above are:</p>\n<pre><code>image_size = 768\nseed = 42\nwarmup_epo = 1\ninit_lr = 1e-4\nbatch_size = 4\nvalid_batch_size = 4\nn_epochs = 30\nwarmup_factor = 10\nnum_workers = 4\niters_to_accumulate = 4\nuse_amp = True\ndebug = False\nearly_stop = 10\n</code></pre>\n<p>We also used the <code>albumentations</code> library to perform augmentations on the datasets. For the second model, the augmentations for the trainingand validation datasets are as follows:</p>\n<pre><code>transforms_train = albumentations.Compose(\n   [\n       albumentations.RandomResizedCrop(image_size, image_size, scale=(0.9, 1), p=1),\n       albumentations.HorizontalFlip(p=0.5),\n       albumentations.ShiftScaleRotate(p=0.5),\n       albumentations.HueSaturationValue(\n           hue_shift_limit=10, sat_shift_limit=10, val_shift_limit=10, p=0.7\n       ),\n       albumentations.RandomBrightnessContrast(\n           brightness_limit=(-0.2, 0.2), contrast_limit=(-0.2, 0.2), p=0.7\n       ),\n       albumentations.CLAHE(clip_limit=(1, 4), p=0.5),\n       albumentations.OneOf(\n           [\n               albumentations.OpticalDistortion(distort_limit=1.0),\n               albumentations.GridDistortion(num_steps=5, distort_limit=1.0),\n               albumentations.ElasticTransform(alpha=3),\n           ],\n           p=0.2,\n       ),\n       albumentations.OneOf(\n           [\n               albumentations.GaussNoise(var_limit=[10, 50]),\n               albumentations.GaussianBlur(),\n               albumentations.MotionBlur(),\n               albumentations.MedianBlur(),\n           ],\n           p=0.2,\n       ),\n       albumentations.Resize(image_size, image_size),\n       albumentations.OneOf(\n           [\n               JpegCompression(),\n               Downscale(scale_min=0.1, scale_max=0.15),\n           ],\n           p=0.2,\n       ),\n       IAAPiecewiseAffine(p=0.2),\n       IAASharpen(p=0.2),\n       albumentations.Cutout(\n           max_h_size=int(image_size * 0.1),\n           max_w_size=int(image_size * 0.1),\n           num_holes=5,\n           p=0.5,\n       ),\n       albumentations.Normalize(),\n   ]\n)\ntransforms_valid = albumentations.Compose(\n   [albumentations.Resize(image_size, image_size), albumentations.Normalize()]\n)\n</code></pre>\n<h2>Selected Submissions</h2>\n<p>Our best submission to the competition comprised of a weighted (convex) ensemble of two models </p>\n<ul>\n<li><code>Multi-Head ResNet200d</code> and</li>\n<li><code>Multi-head SeResNet152d</code><br>\nboth pretrained on NiH data with Swish activation. The weights of the ensemble were determined by forward selection, inspired by Chris Delotte’s original implementation in the Melanoma competition. In summary, the ensemble is of the form: <br>\n$$w_1 \\times \\text{(mean of predictions for model (1))} + w_2 \\times \\text{(mean of predictions for model (2))}$$<br>\nwhere $w_1 = 0.595$ and $w_2 = 0.405$.<br>\nOur best submission obtained a public score of 0.968 and a private score of 0.972. The notebook that illustrates the forward selection approach can be found <a href=\"https://www.kaggle.com/khoongweihao/ranzcr-forward-selection-oof-ensemble?scriptVersionId=56891135\" target=\"_blank\">here</a>, but at version 5 for the weights here. The dataset containing all our OOFs and respective submission files can be found <a href=\"https://www.kaggle.com/khoongweihao/ranzcr-oof-and-subs\" target=\"_blank\">here</a>, where the description in the dataset contains all the models and their scores. The weights for model (1) can be found <a href=\"https://www.kaggle.com/reighns/ranzcrweights\" target=\"_blank\">here</a>, and the 5-folds used in inference are <code>multihead_resnet200d_fold0_best_loss.pth</code>, <code>multihead_resnet200d_fold1_best_AUC (1).pth</code>, <code>multihead_resnet200d_fold2_best_AUC.pth</code>, <code>multihead_resnet200d_fold3_best_AUC.pth</code>, and <code>multihead_resnet200d_fold4_best_AUC.pth</code>. The weights for model (2) can be found <a href=\"https://www.kaggle.com/khoongweihao/ranzcr-multihead-model-weights\" target=\"_blank\">here</a>, and the 5-folds used in inference are <code>grad_accum_multihead_seresnet152d_swish_fold0_best_AUC.pth</code>, <code>grad_accum_multihead_seresnet152d_swish_fold1_best_AUC.pth</code>, <code>grad_accum_multihead_seresnet152d_swish_fold2_best_AUC.pth</code>, <code>grad_accum_multihead_seresnet152d_swish_fold3_best_AUC.pth</code>, and <code>grad_accum_multihead_seresnet152d_swish_fold4_best_AUC.pth</code>.</li>\n</ul>\n<h2>Our second selected submission obtained a public score of 0.967 and a private score of 0.971. The submission contained only model (2) above, where its 5-folds were inferenced. The weights are the same as above. Note that when inferencing the models, we used Sin’s pipeline available <a href=\"https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\" target=\"_blank\">here</a>.</h2>\n<h1>Conclusion</h1>\n<p>What we could have done better:</p>\n<ul>\n<li>Use more variety of <code>classifier head</code> like <code>GeM</code>.</li>\n<li>Use more variety of <code>backbone</code> and WE JUST CANNOT MAKE <code>efficietnet</code> work. 😐</li>\n<li>Use <a href=\"http://neptune.ai\" target=\"_blank\">Neptune.ai</a> to log our experiments as soon things start to get messy.</li>\n<li>Experiment on 3-4 stage training.</li>\n<li>Pseudo Labelling</li>\n<li>Knowledge Distillation</li>\n<li>Experiment more on maximizing AUC during ensembles. <code>rank_pct</code> etc.<br>\nEDIT: I also express my gratitude to <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> as well, he has provided a lot of tips and insights.<br>\nThank you to Kaggle and the community for hosting this competition. I have learned so much just by standing on the shoulder of giants. <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> to name a few. Taking their ideas and incorporating them into our own pipeline made things work. Of course, I would really like to thank my teammate and buddy <a href=\"https://www.kaggle.com/khoongweihao\" target=\"_blank\">@khoongweihao</a> for all the help he provided me in the past year. </li>\n</ul>",
  "messages": [
    {
      "id": "1241576",
      "postDate": "03/17/2021 05:33:24",
      "content": "<h1>Transfer Learning</h1>\n<p>As we all know, if we train on <code>imagenet</code> weights, we may take quite a while to converge, even if we finetune it. The intuition is simple, <code>imagenet</code> were trained on many common items in life, and none of them resemble closely to the image structures of X-rays, therefore, the model may have a hard time detecting shapes and details from the X-rays. We can of course unfreeze all the layers and retrain them from scratch, using the State of the Art models' backbone, however, due to limited hardware, we decided it is best to use what others have trained. After all, it is much easier to stand on the shoulder of giants like <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">ammarali</a>. Consequently, I conveniently used a set of <code>pretrained</code> weights trained specifically on this dataset as a starting point. The weights and ideas can be found <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910\" target=\"_blank\">here</a></strong>.<br>\nWe used a few models and found out that <code>resnet200d</code> has the best results on this set of training images. There is no deep reason why this model outperforms other State of the Art models, but using <code>gradcam</code> we can see how the model sees the images.</p>\n<h2>Cross Validation Strategy</h2>\n<p>Knowing how to choose a robust and leak free cross validation strategy is extremely important. Data leakage can cause you to have blind confidence on your model. We are also guilty of committing one since we trained our models with the NiH pretrained weights, without taking into consideration if the weights overlap with the training and validation folds information.</p>\n<ul>\n<li>Cross Validation Strategy: Multi-Label Stratified Group KFold(K=<strong>5</strong>) - using sin's method to split folds will ensure that there is no leakage. Quoting almost verbatim from <code>sklearn</code> :</li>\n<li>Assuming that some data is Independent and Identically Distributed (i.i.d.) is making the assumption that all samples stem from the same generative process and that the generative process is assumed to have no memory of past generated samples.**</li>\n</ul>\n<h2>Therefore it is paramount to ensure that amongst this 3255 <strong>unique</strong> patients, we need to ensure that each unique patients' images DO NOT appear in the validation fold. That is to say, if patient John Doe has 100 X-ray images, but during our 5-fold splits, he has 70 images in Fold 1-4, while 30 images are in Fold 5, then if we were to train on Fold 1-4 and validate on Fold 5, there may be potential leakage and the model will predict with confidence for John Doe's images. This is under the assumption that John Doe's data does not fulfill the i.i.d process.</h2>\n<h2>Model Architectures, Training Parameters &amp; Augmentations</h2>\n<p>We built-upon Tawara’s Multi-head model for our best scoring models. In particular, we experimented with the activation functions and dropout rates. We found models with <code>Swish</code> activation in the <code>multi-head</code> component of the network to perform best in our experiments. Our best scoring single model is a multi-head model with a <code>resnet200d</code> backbone. In particular, one single fold of <code>resnet200d</code> gives a private score of 0.970. <br>\nWe started experimenting with <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">sin's</a> <a href=\"https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965\" target=\"_blank\">pipeline</a>  which is similar to qishen ha's pipeline back in Melanoma and used <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">Tawara's</a> multihead approach. We did not have time to experiment with the 3-4 stage training as we joined the competition late.</p>\n<ul>\n<li>model:<ul>\n<li><strong>backbone</strong>: <code>ResNet200D</code> and <code>SeResNet152d</code></li>\n<li><strong>classifier/multi-head:</strong> independent&nbsp;<strong>Spatial-Attention Module</strong>&nbsp;and MLP by Target Group(ETT(3), NGT(4), CVC(3), and Swan(1))</li>\n<li><strong>NOTE: I use&nbsp;<a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">the pre-trained model</a>&nbsp;shared by <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> .</strong>&nbsp;Thanks!</li></ul></li>\n</ul>\n<h2>Preprocessing</h2>\n<p><a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">Rueben Schmidt</a> mentioned in this <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146\" target=\"_blank\">post</a> here that some images have black borders around them. I removed them during both the training and inference process. There was no significant increase on the LB score, even if there was, it is in the 3-4th decimal places, but I noticed my local cv to increase. Thus I decided to remove for all. After all, if I keep this consistent in both training and inference, I reckon that no surprise factor would pop out. </p>\n<h2>Augmentation</h2>\n<p>In particular, we made use of a different <code>Normalization</code> parameter which is more accustomed to the X-ray pretrained images. Thanks Tawara again! Heavy augmentations are used during <strong>Train-Time-Augmentation.</strong> But during <strong>Test-Time-Augmentation,</strong> we merely used a <code>HorizontalFlip</code> with 100% probability, and only used <code>tta_steps=1</code>. </p>\n<pre><code>augmentations_class: AlbumentationsAugmentation\naugmentations_train:\n AlbumentationsAugmentation:\n   - name: RandomResizedCrop\n     params:\n       height: 640\n       width: 640\n       scale: [0.9, 1.0]\n       p: 1.0\n   - name: HorizontalFlip\n     params:\n       p: 0.5\n   - name: ShiftScaleRotate\n     params:\n       shift_limit: 0.2\n       scale_limit: 0.2\n       rotate_limit: 20\n       border_mode: 0\n       value: 0\n       mask_value: 0\n       p: 0.5\n   - name: HueSaturationValue\n     params:\n       hue_shift_limit: 10\n       sat_shift_limit: 10\n       val_shift_limit: 10\n       p: 0.7\n   - name: RandomBrightnessContrast\n     params:\n       brightness_limit: [-0.2, 0.2]\n       contrast_limit: [-0.2, 0.2]\n       p: 0.7\n   - name: CLAHE\n     params:\n       clip_limit: [1,4]\n       p: 0.5\n   - name: JpegCompression\n     params:\n       p: 0.2\n   - name: IAAPiecewiseAffine\n     params:\n       p: 0.2\n   - name: IAASharpen\n     params:\n       p: 0.2\n   - name: Cutout\n     params: \n       # use int(image_size * 0.1)\n       max_h_size: 64\n       max_w_size: 64\n       num_holes: 5\n       p: 0.5\n   - name: Resize\n     params:\n       height: 640\n       width: 640\n       p: 1.0\n   - name: Normalize\n     params:\n       mean: [0.4887381077884414]\n       std: [0.23064819430546407]\n       p: 1.0\n   - name: ToTensorV2\n     params:\n       p: 1.0\naugmentations_val:\n AlbumentationsAugmentation:\n   - name: Resize\n     params:\n       height: 640\n       width: 640\n       p: 1.0\n   - name: Normalize\n     params:\n       mean: [0.4887381077884414]\n       std: [0.23064819430546407]\n       p: 1.0\n   - name: ToTensorV2\n     params:\n       p: 1.0\n</code></pre>\n<hr>\n<h2>Batch Size and Tricks</h2>\n<p>Due to hardware limitation, we can barely fit in anything more than a <code>batch_size</code> of 8. We quote the well known fact <a href=\"https://arxiv.org/abs/1609.04836\" target=\"_blank\">here</a>:</p>\n<blockquote>\n  <p>It has been observed in practice that when using a larger batch there is a degradation in the quality of the model, as measured by its ability to generalize […]<br>\n  large-batch methods tend to converge to sharp minimizers of the training and testing functions—and as is well known, sharp minima lead to poorer generalization. In contrast, small-batch methods consistently converge to flat minimizers, and our experiments support a commonly held view that this is due to the inherent noise in the gradient estimation.<br>\n  The above shows that large batch size may <code>fit</code> the model too well, as the model will learn features of the dataset in less iterations, and may memorize this particular dataset's features, leading to overfitting and poor generalization. However, too small a batch size causes our convergence to go too slow, empirically, we take 32 or 64 as the ideal batch size in this competition. </p>\n  <h2>We used both <code>torch.amp</code> and <code>gradient accumulation</code> to be able to fit more batch sizes. We did not freeze the <code>batch_norm</code> layers, which still yielded great results. What we should have done is to experiment more on how to freeze the batch norm layers properly, as I believe that it may help. In the end, we used a batch size of 8 and fit 4 iterations using <code>gradient accumulation</code>  and trained a total number of 20 epochs to get a local CV score of roughly 0.969.</h2>\n  <h2>Optimizer, Scheduler and Loss</h2>\n  <p>Nothing too fancy here, although we really wanted to try out <code>Focal Loss</code> in this setting. The configuration can be seen here. But note that we incorporated <code>GradualWarmUpScheduler</code> along with <code>CosineAnnealingLR</code>.</p>\n</blockquote>\n<pre><code>scheduler: CosineAnnealingLR\nscheduler_params: # Note that in params we must put 1.e instead of 1e\n CosineAnnealingLR:\n   T_max: 16\n   eta_min: 1.e-7\n   last_epoch: -1\n   verbose: True   \ntrain_step_scheduler: False\nval_step_scheduler: False\noptimizer: Adam\noptimizer_params:\n Adam:\n   lr: 0.00002\n   betas:\n     - 0.9\n     - 0.999\n   eps: 1.e-7\n   weight_decay: 0\n   amsgrad: False\ncriterion_train: BCEWithLogitsLoss\ncriterion_val: BCEWithLogitsLoss\ncriterion_params:\n CrossEntropyLoss:\n   weight: null\n   size_average: null\n   ignore_index: -100\n   reduce: null\n   reduction: mean\n LabelSmoothingLoss:\n   classes: 2\n   smoothing: 0.05\n   dim: -1\n</code></pre>\n<h2>Activation Function</h2>\n<pre><code>import torch\nsigmoid = torch.nn.Sigmoid()\nclass Swish(torch.autograd.Function):\n   @staticmethod\n   def forward(ctx, i):\n       result = i * sigmoid(i)\n       ctx.save_for_backward(i)\n       return result\n   @staticmethod\n   def backward(ctx, grad_output):\n       i = ctx.saved_variables[0]\n       sigmoid_i = sigmoid(i)\n       return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\nclass Swish_Module(torch.nn.Module):\n   def forward(self, x):\n       return Swish.apply(x)\n</code></pre>\n<hr>\n<p>Our second best performing model is also a multi-head model with swish activation in the heads, but with a <code>SeResNet152d</code> backbone (<code>seresnet152d</code>).<br>\nDuring training, we use gradient accumulation so that the bath size can scale up eventually to our desired sizes. For the second model, it scales to 16. The training parameters for our second best performing model above are:</p>\n<pre><code>image_size = 768\nseed = 42\nwarmup_epo = 1\ninit_lr = 1e-4\nbatch_size = 4\nvalid_batch_size = 4\nn_epochs = 30\nwarmup_factor = 10\nnum_workers = 4\niters_to_accumulate = 4\nuse_amp = True\ndebug = False\nearly_stop = 10\n</code></pre>\n<p>We also used the <code>albumentations</code> library to perform augmentations on the datasets. For the second model, the augmentations for the trainingand validation datasets are as follows:</p>\n<pre><code>transforms_train = albumentations.Compose(\n   [\n       albumentations.RandomResizedCrop(image_size, image_size, scale=(0.9, 1), p=1),\n       albumentations.HorizontalFlip(p=0.5),\n       albumentations.ShiftScaleRotate(p=0.5),\n       albumentations.HueSaturationValue(\n           hue_shift_limit=10, sat_shift_limit=10, val_shift_limit=10, p=0.7\n       ),\n       albumentations.RandomBrightnessContrast(\n           brightness_limit=(-0.2, 0.2), contrast_limit=(-0.2, 0.2), p=0.7\n       ),\n       albumentations.CLAHE(clip_limit=(1, 4), p=0.5),\n       albumentations.OneOf(\n           [\n               albumentations.OpticalDistortion(distort_limit=1.0),\n               albumentations.GridDistortion(num_steps=5, distort_limit=1.0),\n               albumentations.ElasticTransform(alpha=3),\n           ],\n           p=0.2,\n       ),\n       albumentations.OneOf(\n           [\n               albumentations.GaussNoise(var_limit=[10, 50]),\n               albumentations.GaussianBlur(),\n               albumentations.MotionBlur(),\n               albumentations.MedianBlur(),\n           ],\n           p=0.2,\n       ),\n       albumentations.Resize(image_size, image_size),\n       albumentations.OneOf(\n           [\n               JpegCompression(),\n               Downscale(scale_min=0.1, scale_max=0.15),\n           ],\n           p=0.2,\n       ),\n       IAAPiecewiseAffine(p=0.2),\n       IAASharpen(p=0.2),\n       albumentations.Cutout(\n           max_h_size=int(image_size * 0.1),\n           max_w_size=int(image_size * 0.1),\n           num_holes=5,\n           p=0.5,\n       ),\n       albumentations.Normalize(),\n   ]\n)\ntransforms_valid = albumentations.Compose(\n   [albumentations.Resize(image_size, image_size), albumentations.Normalize()]\n)\n</code></pre>\n<h2>Selected Submissions</h2>\n<p>Our best submission to the competition comprised of a weighted (convex) ensemble of two models </p>\n<ul>\n<li><code>Multi-Head ResNet200d</code> and</li>\n<li><code>Multi-head SeResNet152d</code><br>\nboth pretrained on NiH data with Swish activation. The weights of the ensemble were determined by forward selection, inspired by Chris Delotte’s original implementation in the Melanoma competition. In summary, the ensemble is of the form: <br>\n$$w_1 \\times \\text{(mean of predictions for model (1))} + w_2 \\times \\text{(mean of predictions for model (2))}$$<br>\nwhere $w_1 = 0.595$ and $w_2 = 0.405$.<br>\nOur best submission obtained a public score of 0.968 and a private score of 0.972. The notebook that illustrates the forward selection approach can be found <a href=\"https://www.kaggle.com/khoongweihao/ranzcr-forward-selection-oof-ensemble?scriptVersionId=56891135\" target=\"_blank\">here</a>, but at version 5 for the weights here. The dataset containing all our OOFs and respective submission files can be found <a href=\"https://www.kaggle.com/khoongweihao/ranzcr-oof-and-subs\" target=\"_blank\">here</a>, where the description in the dataset contains all the models and their scores. The weights for model (1) can be found <a href=\"https://www.kaggle.com/reighns/ranzcrweights\" target=\"_blank\">here</a>, and the 5-folds used in inference are <code>multihead_resnet200d_fold0_best_loss.pth</code>, <code>multihead_resnet200d_fold1_best_AUC (1).pth</code>, <code>multihead_resnet200d_fold2_best_AUC.pth</code>, <code>multihead_resnet200d_fold3_best_AUC.pth</code>, and <code>multihead_resnet200d_fold4_best_AUC.pth</code>. The weights for model (2) can be found <a href=\"https://www.kaggle.com/khoongweihao/ranzcr-multihead-model-weights\" target=\"_blank\">here</a>, and the 5-folds used in inference are <code>grad_accum_multihead_seresnet152d_swish_fold0_best_AUC.pth</code>, <code>grad_accum_multihead_seresnet152d_swish_fold1_best_AUC.pth</code>, <code>grad_accum_multihead_seresnet152d_swish_fold2_best_AUC.pth</code>, <code>grad_accum_multihead_seresnet152d_swish_fold3_best_AUC.pth</code>, and <code>grad_accum_multihead_seresnet152d_swish_fold4_best_AUC.pth</code>.</li>\n</ul>\n<h2>Our second selected submission obtained a public score of 0.967 and a private score of 0.971. The submission contained only model (2) above, where its 5-folds were inferenced. The weights are the same as above. Note that when inferencing the models, we used Sin’s pipeline available <a href=\"https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965\" target=\"_blank\">here</a>.</h2>\n<h1>Conclusion</h1>\n<p>What we could have done better:</p>\n<ul>\n<li>Use more variety of <code>classifier head</code> like <code>GeM</code>.</li>\n<li>Use more variety of <code>backbone</code> and WE JUST CANNOT MAKE <code>efficietnet</code> work. 😐</li>\n<li>Use <a href=\"http://neptune.ai\" target=\"_blank\">Neptune.ai</a> to log our experiments as soon things start to get messy.</li>\n<li>Experiment on 3-4 stage training.</li>\n<li>Pseudo Labelling</li>\n<li>Knowledge Distillation</li>\n<li>Experiment more on maximizing AUC during ensembles. <code>rank_pct</code> etc.<br>\nEDIT: I also express my gratitude to <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> as well, he has provided a lot of tips and insights.<br>\nThank you to Kaggle and the community for hosting this competition. I have learned so much just by standing on the shoulder of giants. <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> to name a few. Taking their ideas and incorporating them into our own pipeline made things work. Of course, I would really like to thank my teammate and buddy <a href=\"https://www.kaggle.com/khoongweihao\" target=\"_blank\">@khoongweihao</a> for all the help he provided me in the past year. </li>\n</ul>",
      "rawMarkdown": "# Transfer Learning\n\nAs we all know, if we train on `imagenet` weights, we may take quite a while to converge, even if we finetune it. The intuition is simple, `imagenet` were trained on many common items in life, and none of them resemble closely to the image structures of X-rays, therefore, the model may have a hard time detecting shapes and details from the X-rays. We can of course unfreeze all the layers and retrain them from scratch, using the State of the Art models' backbone, however, due to limited hardware, we decided it is best to use what others have trained. After all, it is much easier to stand on the shoulder of giants like [ammarali](https://www.kaggle.com/ammarali32). Consequently, I conveniently used a set of `pretrained` weights trained specifically on this dataset as a starting point. The weights and ideas can be found **[here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910)**.\n\nWe used a few models and found out that `resnet200d` has the best results on this set of training images. There is no deep reason why this model outperforms other State of the Art models, but using `gradcam` we can see how the model sees the images.\n\n## Cross Validation Strategy\n\nKnowing how to choose a robust and leak free cross validation strategy is extremely important. Data leakage can cause you to have blind confidence on your model. We are also guilty of committing one since we trained our models with the NiH pretrained weights, without taking into consideration if the weights overlap with the training and validation folds information.\n\n- Cross Validation Strategy: Multi-Label Stratified Group KFold(K=**5**) - using sin's method to split folds will ensure that there is no leakage. Quoting almost verbatim from `sklearn` :\n\n- Assuming that some data is Independent and Identically Distributed (i.i.d.) is making the assumption that all samples stem from the same generative process and that the generative process is assumed to have no memory of past generated samples.**\n\nTherefore it is paramount to ensure that amongst this 3255 **unique** patients, we need to ensure that each unique patients' images DO NOT appear in the validation fold. That is to say, if patient John Doe has 100 X-ray images, but during our 5-fold splits, he has 70 images in Fold 1-4, while 30 images are in Fold 5, then if we were to train on Fold 1-4 and validate on Fold 5, there may be potential leakage and the model will predict with confidence for John Doe's images. This is under the assumption that John Doe's data does not fulfill the i.i.d process.\n---\n\n## Model Architectures, Training Parameters & Augmentations\n\nWe built-upon Tawara’s Multi-head model for our best scoring models. In particular, we experimented with the activation functions and dropout rates. We found models with `Swish` activation in the `multi-head` component of the network to perform best in our experiments. Our best scoring single model is a multi-head model with a `resnet200d` backbone. In particular, one single fold of `resnet200d` gives a private score of 0.970. \n\nWe started experimenting with [sin's](https://www.kaggle.com/underwearfitting) [pipeline](https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965)  which is similar to qishen ha's pipeline back in Melanoma and used [Tawara's](https://www.kaggle.com/ttahara) multihead approach. We did not have time to experiment with the 3-4 stage training as we joined the competition late.\n\n- model:\n    - **backbone**: `ResNet200D` and `SeResNet152d`\n    - **classifier/multi-head:** independent **Spatial-Attention Module** and MLP by Target Group(ETT(3), NGT(4), CVC(3), and Swan(1))\n    - **NOTE: I use [the pre-trained model](https://www.kaggle.com/ammarali32/startingpointschestx) shared by @ammarali32 .** Thanks!\n\n## Preprocessing\n\n[Rueben Schmidt](https://www.kaggle.com/reubenschmidt) mentioned in this [post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146) here that some images have black borders around them. I removed them during both the training and inference process. There was no significant increase on the LB score, even if there was, it is in the 3-4th decimal places, but I noticed my local cv to increase. Thus I decided to remove for all. After all, if I keep this consistent in both training and inference, I reckon that no surprise factor would pop out. \n\n## Augmentation\n\nIn particular, we made use of a different `Normalization` parameter which is more accustomed to the X-ray pretrained images. Thanks Tawara again! Heavy augmentations are used during **Train-Time-Augmentation.** But during **Test-Time-Augmentation,** we merely used a `HorizontalFlip` with 100% probability, and only used `tta_steps=1`. \n\n```yaml\naugmentations_class: AlbumentationsAugmentation\naugmentations_train:\n  AlbumentationsAugmentation:\n    - name: RandomResizedCrop\n      params:\n        height: 640\n        width: 640\n        scale: [0.9, 1.0]\n        p: 1.0\n    - name: HorizontalFlip\n      params:\n        p: 0.5\n    - name: ShiftScaleRotate\n      params:\n        shift_limit: 0.2\n        scale_limit: 0.2\n        rotate_limit: 20\n        border_mode: 0\n        value: 0\n        mask_value: 0\n        p: 0.5\n    - name: HueSaturationValue\n      params:\n        hue_shift_limit: 10\n        sat_shift_limit: 10\n        val_shift_limit: 10\n        p: 0.7\n    - name: RandomBrightnessContrast\n      params:\n        brightness_limit: [-0.2, 0.2]\n        contrast_limit: [-0.2, 0.2]\n        p: 0.7\n    - name: CLAHE\n      params:\n        clip_limit: [1,4]\n        p: 0.5\n    - name: JpegCompression\n      params:\n        p: 0.2\n    - name: IAAPiecewiseAffine\n      params:\n        p: 0.2\n    - name: IAASharpen\n      params:\n        p: 0.2\n    - name: Cutout\n      params: \n        # use int(image_size * 0.1)\n        max_h_size: 64\n        max_w_size: 64\n        num_holes: 5\n        p: 0.5\n    - name: Resize\n      params:\n        height: 640\n        width: 640\n        p: 1.0\n    - name: Normalize\n      params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]\n        p: 1.0\n    - name: ToTensorV2\n      params:\n        p: 1.0\naugmentations_val:\n  AlbumentationsAugmentation:\n    - name: Resize\n      params:\n        height: 640\n        width: 640\n        p: 1.0\n    - name: Normalize\n      params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]\n        p: 1.0\n    - name: ToTensorV2\n      params:\n        p: 1.0\n```\n\n---\n\n## Batch Size and Tricks\n\nDue to hardware limitation, we can barely fit in anything more than a `batch_size` of 8. We quote the well known fact [here](https://arxiv.org/abs/1609.04836):\n\n> It has been observed in practice that when using a larger batch there is a degradation in the quality of the model, as measured by its ability to generalize [...]\n\n> large-batch methods tend to converge to sharp minimizers of the training and testing functions—and as is well known, sharp minima lead to poorer generalization. In contrast, small-batch methods consistently converge to flat minimizers, and our experiments support a commonly held view that this is due to the inherent noise in the gradient estimation.\n\nThe above shows that large batch size may `fit` the model too well, as the model will learn features of the dataset in less iterations, and may memorize this particular dataset's features, leading to overfitting and poor generalization. However, too small a batch size causes our convergence to go too slow, empirically, we take 32 or 64 as the ideal batch size in this competition. \n\nWe used both `torch.amp` and `gradient accumulation` to be able to fit more batch sizes. We did not freeze the `batch_norm` layers, which still yielded great results. What we should have done is to experiment more on how to freeze the batch norm layers properly, as I believe that it may help. In the end, we used a batch size of 8 and fit 4 iterations using `gradient accumulation`  and trained a total number of 20 epochs to get a local CV score of roughly 0.969.\n\n---\n\n## Optimizer, Scheduler and Loss\n\nNothing too fancy here, although we really wanted to try out `Focal Loss` in this setting. The configuration can be seen here. But note that we incorporated `GradualWarmUpScheduler` along with `CosineAnnealingLR`.\n\n```yaml\nscheduler: CosineAnnealingLR\nscheduler_params: # Note that in params we must put 1.e instead of 1e\n  CosineAnnealingLR:\n    T_max: 16\n    eta_min: 1.e-7\n    last_epoch: -1\n    verbose: True   \ntrain_step_scheduler: False\nval_step_scheduler: False\noptimizer: Adam\noptimizer_params:\n  Adam:\n    lr: 0.00002\n    betas:\n      - 0.9\n      - 0.999\n    eps: 1.e-7\n    weight_decay: 0\n    amsgrad: False\ncriterion_train: BCEWithLogitsLoss\ncriterion_val: BCEWithLogitsLoss\ncriterion_params:\n  CrossEntropyLoss:\n    weight: null\n    size_average: null\n    ignore_index: -100\n    reduce: null\n    reduction: mean\n  LabelSmoothingLoss:\n    classes: 2\n    smoothing: 0.05\n    dim: -1\n```\n\n## Activation Function\n\n```python\nimport torch\n\nsigmoid = torch.nn.Sigmoid()\n\nclass Swish(torch.autograd.Function):\n    @staticmethod\n    def forward(ctx, i):\n        result = i * sigmoid(i)\n        ctx.save_for_backward(i)\n        return result\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        i = ctx.saved_variables[0]\n        sigmoid_i = sigmoid(i)\n        return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\n\nclass Swish_Module(torch.nn.Module):\n    def forward(self, x):\n        return Swish.apply(x)\n```\n\n---\n\nOur second best performing model is also a multi-head model with swish activation in the heads, but with a `SeResNet152d` backbone (`seresnet152d`).\n\nDuring training, we use gradient accumulation so that the bath size can scale up eventually to our desired sizes. For the second model, it scales to 16. The training parameters for our second best performing model above are:\n\n```python\nimage_size = 768\nseed = 42\nwarmup_epo = 1\ninit_lr = 1e-4\nbatch_size = 4\nvalid_batch_size = 4\nn_epochs = 30\nwarmup_factor = 10\nnum_workers = 4\niters_to_accumulate = 4\nuse_amp = True\ndebug = False\nearly_stop = 10\n```\n\nWe also used the `albumentations` library to perform augmentations on the datasets. For the second model, the augmentations for the trainingand validation datasets are as follows:\n\n```python\ntransforms_train = albumentations.Compose(\n    [\n        albumentations.RandomResizedCrop(image_size, image_size, scale=(0.9, 1), p=1),\n        albumentations.HorizontalFlip(p=0.5),\n        albumentations.ShiftScaleRotate(p=0.5),\n        albumentations.HueSaturationValue(\n            hue_shift_limit=10, sat_shift_limit=10, val_shift_limit=10, p=0.7\n        ),\n        albumentations.RandomBrightnessContrast(\n            brightness_limit=(-0.2, 0.2), contrast_limit=(-0.2, 0.2), p=0.7\n        ),\n        albumentations.CLAHE(clip_limit=(1, 4), p=0.5),\n        albumentations.OneOf(\n            [\n                albumentations.OpticalDistortion(distort_limit=1.0),\n                albumentations.GridDistortion(num_steps=5, distort_limit=1.0),\n                albumentations.ElasticTransform(alpha=3),\n            ],\n            p=0.2,\n        ),\n        albumentations.OneOf(\n            [\n                albumentations.GaussNoise(var_limit=[10, 50]),\n                albumentations.GaussianBlur(),\n                albumentations.MotionBlur(),\n                albumentations.MedianBlur(),\n            ],\n            p=0.2,\n        ),\n        albumentations.Resize(image_size, image_size),\n        albumentations.OneOf(\n            [\n                JpegCompression(),\n                Downscale(scale_min=0.1, scale_max=0.15),\n            ],\n            p=0.2,\n        ),\n        IAAPiecewiseAffine(p=0.2),\n        IAASharpen(p=0.2),\n        albumentations.Cutout(\n            max_h_size=int(image_size * 0.1),\n            max_w_size=int(image_size * 0.1),\n            num_holes=5,\n            p=0.5,\n        ),\n        albumentations.Normalize(),\n    ]\n)\n\ntransforms_valid = albumentations.Compose(\n    [albumentations.Resize(image_size, image_size), albumentations.Normalize()]\n)\n```\n\n## Selected Submissions\n\nOur best submission to the competition comprised of a weighted (convex) ensemble of two models \n\n- `Multi-Head ResNet200d` and\n- `Multi-head SeResNet152d`\n\nboth pretrained on NiH data with Swish activation. The weights of the ensemble were determined by forward selection, inspired by Chris Delotte’s original implementation in the Melanoma competition. In summary, the ensemble is of the form: \n\n$$w_1 \\times \\text{(mean of predictions for model (1))} + w_2 \\times \\text{(mean of predictions for model (2))}$$\n\nwhere $w_1 = 0.595$ and $w_2 = 0.405$.\n\nOur best submission obtained a public score of 0.968 and a private score of 0.972. The notebook that illustrates the forward selection approach can be found [here](https://www.kaggle.com/khoongweihao/ranzcr-forward-selection-oof-ensemble?scriptVersionId=56891135), but at version 5 for the weights here. The dataset containing all our OOFs and respective submission files can be found [here](https://www.kaggle.com/khoongweihao/ranzcr-oof-and-subs), where the description in the dataset contains all the models and their scores. The weights for model (1) can be found [here](https://www.kaggle.com/reighns/ranzcrweights), and the 5-folds used in inference are `multihead_resnet200d_fold0_best_loss.pth`, `multihead_resnet200d_fold1_best_AUC (1).pth`, `multihead_resnet200d_fold2_best_AUC.pth`, `multihead_resnet200d_fold3_best_AUC.pth`, and `multihead_resnet200d_fold4_best_AUC.pth`. The weights for model (2) can be found [here](https://www.kaggle.com/khoongweihao/ranzcr-multihead-model-weights), and the 5-folds used in inference are `grad_accum_multihead_seresnet152d_swish_fold0_best_AUC.pth`, `grad_accum_multihead_seresnet152d_swish_fold1_best_AUC.pth`, `grad_accum_multihead_seresnet152d_swish_fold2_best_AUC.pth`, `grad_accum_multihead_seresnet152d_swish_fold3_best_AUC.pth`, and `grad_accum_multihead_seresnet152d_swish_fold4_best_AUC.pth`.\n\nOur second selected submission obtained a public score of 0.967 and a private score of 0.971. The submission contained only model (2) above, where its 5-folds were inferenced. The weights are the same as above. Note that when inferencing the models, we used Sin’s pipeline available [here](https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965).\n\n---\n\n# Conclusion\n\nWhat we could have done better:\n\n- Use more variety of `classifier head` like `GeM`.\n- Use more variety of `backbone` and WE JUST CANNOT MAKE `efficietnet` work. 😐\n- Use [Neptune.ai](http://neptune.ai) to log our experiments as soon things start to get messy.\n- Experiment on 3-4 stage training.\n- Pseudo Labelling\n- Knowledge Distillation\n- Experiment more on maximizing AUC during ensembles. `rank_pct` etc.\n\n\nEDIT: I also express my gratitude to @bjoernholzhauer as well, he has provided a lot of tips and insights.\n\nThank you to Kaggle and the community for hosting this competition. I have learned so much just by standing on the shoulder of giants. @ttahara @underwearfitting @hengck23 @ammarali32 @cdeotte @haqishen to name a few. Taking their ideas and incorporating them into our own pipeline made things work. Of course, I would really like to thank my teammate and buddy @khoongweihao for all the help he provided me in the past year.",
      "votes": null
    },
    {
      "id": "1241963",
      "postDate": "03/17/2021 09:57:28",
      "content": "<p>Congratulations. Thanks for great detailed description </p>",
      "rawMarkdown": "Congratulations. Thanks for great detailed description",
      "votes": null
    },
    {
      "id": "1241964",
      "postDate": "03/17/2021 09:58:37",
      "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> Thanks, you are one of the giants whose shoulder we stood on :) </p>",
      "rawMarkdown": "ammarali32 Thanks, you are one of the giants whose shoulder we stood on :)",
      "votes": null
    },
    {
      "id": "1241976",
      "postDate": "03/17/2021 10:07:43",
      "content": "<p>Thanks, I am really glad to hear that and you are the giants who scored so well and got on the top 2%</p>",
      "rawMarkdown": "Thanks, I am really glad to hear that and you are the giants who scored so well and got on the top 2%",
      "votes": null
    },
    {
      "id": "1242035",
      "postDate": "03/17/2021 11:00:38",
      "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> Congratulations on Silver Finish Hongnan . Great and Detailed Solution Writeup </p>",
      "rawMarkdown": "reighns Congratulations on Silver Finish Hongnan . Great and Detailed Solution Writeup",
      "votes": null
    },
    {
      "id": "1242058",
      "postDate": "03/17/2021 11:23:11",
      "content": "<p>Congrats on your silver and thank you for the details!</p>",
      "rawMarkdown": "Congrats on your silver and thank you for the details!",
      "votes": null
    },
    {
      "id": "1242115",
      "postDate": "03/17/2021 12:10:47",
      "content": "<p>Thanks a lot! Congrats to you too! Been following you since Melanoma. Haha</p>",
      "rawMarkdown": "Thanks a lot! Congrats to you too! Been following you since Melanoma. Haha",
      "votes": null
    },
    {
      "id": "1242117",
      "postDate": "03/17/2021 12:11:10",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> for all the great resources. Saves us a lot of time. </p>",
      "rawMarkdown": "Thanks @usharengaraju for all the great resources. Saves us a lot of time.",
      "votes": null
    },
    {
      "id": "1242127",
      "postDate": "03/17/2021 12:22:29",
      "content": "<p>Thanks for the detailed write up and congrats on silver!</p>\n<p>We tried different approaches for rank ensembles and had a gain in our final score. It can be a worth to try. <br>\nCongrats again!</p>",
      "rawMarkdown": "Thanks for the detailed write up and congrats on silver!\n\nWe tried different approaches for rank ensembles and had a gain in our final score. It can be a worth to try. \nCongrats again!",
      "votes": null
    },
    {
      "id": "1242269",
      "postDate": "03/17/2021 14:05:29",
      "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> congrats on the silver medal and thanks a lot for sharing such a detailed summary! </p>",
      "rawMarkdown": "reighns congrats on the silver medal and thanks a lot for sharing such a detailed summary!",
      "votes": null
    },
    {
      "id": "1242312",
      "postDate": "03/17/2021 14:39:28",
      "content": "<p>Well done my friend, you deserve it!</p>",
      "rawMarkdown": "Well done my friend, you deserve it!",
      "votes": null
    },
    {
      "id": "1245267",
      "postDate": "03/19/2021 16:23:55",
      "content": "<p>did you find out the mean and std of the images as you have written <br>\n<code>params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]</code><br>\nand by <code>but using gradcam</code> are you meaning gradient accumulation or \"Gradient-weighted Class Activation Mapping,\"?</p>",
      "rawMarkdown": "did you find out the mean and std of the images as you have written \n`params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]`\nand by ` but using gradcam` are you meaning gradient accumulation or \"Gradient-weighted Class Activation Mapping,\"?",
      "votes": null
    },
    {
      "id": "1245276",
      "postDate": "03/19/2021 16:30:27",
      "content": "<p>Great Solution <br>\nand Congratulation on Silver!<br>\nI never used to try any new things into backbone</p>",
      "rawMarkdown": "Great Solution \nand Congratulation on Silver!\nI never used to try any new things into backbone",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1241963,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "03/17/2021 09:57:28",
      "content": "<p>Congratulations. Thanks for great detailed description </p>",
      "votes": null,
      "replies": [
        {
          "id": 1241964,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/17/2021 09:58:37",
          "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> Thanks, you are one of the giants whose shoulder we stood on :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1241976,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "03/17/2021 10:07:43",
          "content": "<p>Thanks, I am really glad to hear that and you are the giants who scored so well and got on the top 2%</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1242035,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/17/2021 11:00:38",
      "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> Congratulations on Silver Finish Hongnan . Great and Detailed Solution Writeup </p>",
      "votes": null,
      "replies": [
        {
          "id": 1242117,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/17/2021 12:11:10",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> for all the great resources. Saves us a lot of time. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1242058,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "03/17/2021 11:23:11",
      "content": "<p>Congrats on your silver and thank you for the details!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1242115,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/17/2021 12:10:47",
          "content": "<p>Thanks a lot! Congrats to you too! Been following you since Melanoma. Haha</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1242127,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "03/17/2021 12:22:29",
      "content": "<p>Thanks for the detailed write up and congrats on silver!</p>\n<p>We tried different approaches for rank ensembles and had a gain in our final score. It can be a worth to try. <br>\nCongrats again!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242269,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "03/17/2021 14:05:29",
      "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> congrats on the silver medal and thanks a lot for sharing such a detailed summary! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242312,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "03/17/2021 14:39:28",
      "content": "<p>Well done my friend, you deserve it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1245267,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "03/19/2021 16:23:55",
      "content": "<p>did you find out the mean and std of the images as you have written <br>\n<code>params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]</code><br>\nand by <code>but using gradcam</code> are you meaning gradient accumulation or \"Gradient-weighted Class Activation Mapping,\"?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1245276,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "03/19/2021 16:30:27",
      "content": "<p>Great Solution <br>\nand Congratulation on Silver!<br>\nI never used to try any new things into backbone</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1241576": "# Transfer Learning\n\nAs we all know, if we train on `imagenet` weights, we may take quite a while to converge, even if we finetune it. The intuition is simple, `imagenet` were trained on many common items in life, and none of them resemble closely to the image structures of X-rays, therefore, the model may have a hard time detecting shapes and details from the X-rays. We can of course unfreeze all the layers and retrain them from scratch, using the State of the Art models' backbone, however, due to limited hardware, we decided it is best to use what others have trained. After all, it is much easier to stand on the shoulder of giants like [ammarali](https://www.kaggle.com/ammarali32). Consequently, I conveniently used a set of `pretrained` weights trained specifically on this dataset as a starting point. The weights and ideas can be found **[here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910)**.\n\nWe used a few models and found out that `resnet200d` has the best results on this set of training images. There is no deep reason why this model outperforms other State of the Art models, but using `gradcam` we can see how the model sees the images.\n\n## Cross Validation Strategy\n\nKnowing how to choose a robust and leak free cross validation strategy is extremely important. Data leakage can cause you to have blind confidence on your model. We are also guilty of committing one since we trained our models with the NiH pretrained weights, without taking into consideration if the weights overlap with the training and validation folds information.\n\n- Cross Validation Strategy: Multi-Label Stratified Group KFold(K=**5**) - using sin's method to split folds will ensure that there is no leakage. Quoting almost verbatim from `sklearn` :\n\n- Assuming that some data is Independent and Identically Distributed (i.i.d.) is making the assumption that all samples stem from the same generative process and that the generative process is assumed to have no memory of past generated samples.**\n\nTherefore it is paramount to ensure that amongst this 3255 **unique** patients, we need to ensure that each unique patients' images DO NOT appear in the validation fold. That is to say, if patient John Doe has 100 X-ray images, but during our 5-fold splits, he has 70 images in Fold 1-4, while 30 images are in Fold 5, then if we were to train on Fold 1-4 and validate on Fold 5, there may be potential leakage and the model will predict with confidence for John Doe's images. This is under the assumption that John Doe's data does not fulfill the i.i.d process.\n---\n\n## Model Architectures, Training Parameters & Augmentations\n\nWe built-upon Tawara’s Multi-head model for our best scoring models. In particular, we experimented with the activation functions and dropout rates. We found models with `Swish` activation in the `multi-head` component of the network to perform best in our experiments. Our best scoring single model is a multi-head model with a `resnet200d` backbone. In particular, one single fold of `resnet200d` gives a private score of 0.970. \n\nWe started experimenting with [sin's](https://www.kaggle.com/underwearfitting) [pipeline](https://www.kaggle.com/underwearfitting/single-fold-training-of-resnet200d-lb0-965)  which is similar to qishen ha's pipeline back in Melanoma and used [Tawara's](https://www.kaggle.com/ttahara) multihead approach. We did not have time to experiment with the 3-4 stage training as we joined the competition late.\n\n- model:\n    - **backbone**: `ResNet200D` and `SeResNet152d`\n    - **classifier/multi-head:** independent **Spatial-Attention Module** and MLP by Target Group(ETT(3), NGT(4), CVC(3), and Swan(1))\n    - **NOTE: I use [the pre-trained model](https://www.kaggle.com/ammarali32/startingpointschestx) shared by @ammarali32 .** Thanks!\n\n## Preprocessing\n\n[Rueben Schmidt](https://www.kaggle.com/reubenschmidt) mentioned in this [post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146) here that some images have black borders around them. I removed them during both the training and inference process. There was no significant increase on the LB score, even if there was, it is in the 3-4th decimal places, but I noticed my local cv to increase. Thus I decided to remove for all. After all, if I keep this consistent in both training and inference, I reckon that no surprise factor would pop out. \n\n## Augmentation\n\nIn particular, we made use of a different `Normalization` parameter which is more accustomed to the X-ray pretrained images. Thanks Tawara again! Heavy augmentations are used during **Train-Time-Augmentation.** But during **Test-Time-Augmentation,** we merely used a `HorizontalFlip` with 100% probability, and only used `tta_steps=1`. \n\n```yaml\naugmentations_class: AlbumentationsAugmentation\naugmentations_train:\n  AlbumentationsAugmentation:\n    - name: RandomResizedCrop\n      params:\n        height: 640\n        width: 640\n        scale: [0.9, 1.0]\n        p: 1.0\n    - name: HorizontalFlip\n      params:\n        p: 0.5\n    - name: ShiftScaleRotate\n      params:\n        shift_limit: 0.2\n        scale_limit: 0.2\n        rotate_limit: 20\n        border_mode: 0\n        value: 0\n        mask_value: 0\n        p: 0.5\n    - name: HueSaturationValue\n      params:\n        hue_shift_limit: 10\n        sat_shift_limit: 10\n        val_shift_limit: 10\n        p: 0.7\n    - name: RandomBrightnessContrast\n      params:\n        brightness_limit: [-0.2, 0.2]\n        contrast_limit: [-0.2, 0.2]\n        p: 0.7\n    - name: CLAHE\n      params:\n        clip_limit: [1,4]\n        p: 0.5\n    - name: JpegCompression\n      params:\n        p: 0.2\n    - name: IAAPiecewiseAffine\n      params:\n        p: 0.2\n    - name: IAASharpen\n      params:\n        p: 0.2\n    - name: Cutout\n      params: \n        # use int(image_size * 0.1)\n        max_h_size: 64\n        max_w_size: 64\n        num_holes: 5\n        p: 0.5\n    - name: Resize\n      params:\n        height: 640\n        width: 640\n        p: 1.0\n    - name: Normalize\n      params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]\n        p: 1.0\n    - name: ToTensorV2\n      params:\n        p: 1.0\naugmentations_val:\n  AlbumentationsAugmentation:\n    - name: Resize\n      params:\n        height: 640\n        width: 640\n        p: 1.0\n    - name: Normalize\n      params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]\n        p: 1.0\n    - name: ToTensorV2\n      params:\n        p: 1.0\n```\n\n---\n\n## Batch Size and Tricks\n\nDue to hardware limitation, we can barely fit in anything more than a `batch_size` of 8. We quote the well known fact [here](https://arxiv.org/abs/1609.04836):\n\n> It has been observed in practice that when using a larger batch there is a degradation in the quality of the model, as measured by its ability to generalize [...]\n\n> large-batch methods tend to converge to sharp minimizers of the training and testing functions—and as is well known, sharp minima lead to poorer generalization. In contrast, small-batch methods consistently converge to flat minimizers, and our experiments support a commonly held view that this is due to the inherent noise in the gradient estimation.\n\nThe above shows that large batch size may `fit` the model too well, as the model will learn features of the dataset in less iterations, and may memorize this particular dataset's features, leading to overfitting and poor generalization. However, too small a batch size causes our convergence to go too slow, empirically, we take 32 or 64 as the ideal batch size in this competition. \n\nWe used both `torch.amp` and `gradient accumulation` to be able to fit more batch sizes. We did not freeze the `batch_norm` layers, which still yielded great results. What we should have done is to experiment more on how to freeze the batch norm layers properly, as I believe that it may help. In the end, we used a batch size of 8 and fit 4 iterations using `gradient accumulation`  and trained a total number of 20 epochs to get a local CV score of roughly 0.969.\n\n---\n\n## Optimizer, Scheduler and Loss\n\nNothing too fancy here, although we really wanted to try out `Focal Loss` in this setting. The configuration can be seen here. But note that we incorporated `GradualWarmUpScheduler` along with `CosineAnnealingLR`.\n\n```yaml\nscheduler: CosineAnnealingLR\nscheduler_params: # Note that in params we must put 1.e instead of 1e\n  CosineAnnealingLR:\n    T_max: 16\n    eta_min: 1.e-7\n    last_epoch: -1\n    verbose: True   \ntrain_step_scheduler: False\nval_step_scheduler: False\noptimizer: Adam\noptimizer_params:\n  Adam:\n    lr: 0.00002\n    betas:\n      - 0.9\n      - 0.999\n    eps: 1.e-7\n    weight_decay: 0\n    amsgrad: False\ncriterion_train: BCEWithLogitsLoss\ncriterion_val: BCEWithLogitsLoss\ncriterion_params:\n  CrossEntropyLoss:\n    weight: null\n    size_average: null\n    ignore_index: -100\n    reduce: null\n    reduction: mean\n  LabelSmoothingLoss:\n    classes: 2\n    smoothing: 0.05\n    dim: -1\n```\n\n## Activation Function\n\n```python\nimport torch\n\nsigmoid = torch.nn.Sigmoid()\n\nclass Swish(torch.autograd.Function):\n    @staticmethod\n    def forward(ctx, i):\n        result = i * sigmoid(i)\n        ctx.save_for_backward(i)\n        return result\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        i = ctx.saved_variables[0]\n        sigmoid_i = sigmoid(i)\n        return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\n\nclass Swish_Module(torch.nn.Module):\n    def forward(self, x):\n        return Swish.apply(x)\n```\n\n---\n\nOur second best performing model is also a multi-head model with swish activation in the heads, but with a `SeResNet152d` backbone (`seresnet152d`).\n\nDuring training, we use gradient accumulation so that the bath size can scale up eventually to our desired sizes. For the second model, it scales to 16. The training parameters for our second best performing model above are:\n\n```python\nimage_size = 768\nseed = 42\nwarmup_epo = 1\ninit_lr = 1e-4\nbatch_size = 4\nvalid_batch_size = 4\nn_epochs = 30\nwarmup_factor = 10\nnum_workers = 4\niters_to_accumulate = 4\nuse_amp = True\ndebug = False\nearly_stop = 10\n```\n\nWe also used the `albumentations` library to perform augmentations on the datasets. For the second model, the augmentations for the trainingand validation datasets are as follows:\n\n```python\ntransforms_train = albumentations.Compose(\n    [\n        albumentations.RandomResizedCrop(image_size, image_size, scale=(0.9, 1), p=1),\n        albumentations.HorizontalFlip(p=0.5),\n        albumentations.ShiftScaleRotate(p=0.5),\n        albumentations.HueSaturationValue(\n            hue_shift_limit=10, sat_shift_limit=10, val_shift_limit=10, p=0.7\n        ),\n        albumentations.RandomBrightnessContrast(\n            brightness_limit=(-0.2, 0.2), contrast_limit=(-0.2, 0.2), p=0.7\n        ),\n        albumentations.CLAHE(clip_limit=(1, 4), p=0.5),\n        albumentations.OneOf(\n            [\n                albumentations.OpticalDistortion(distort_limit=1.0),\n                albumentations.GridDistortion(num_steps=5, distort_limit=1.0),\n                albumentations.ElasticTransform(alpha=3),\n            ],\n            p=0.2,\n        ),\n        albumentations.OneOf(\n            [\n                albumentations.GaussNoise(var_limit=[10, 50]),\n                albumentations.GaussianBlur(),\n                albumentations.MotionBlur(),\n                albumentations.MedianBlur(),\n            ],\n            p=0.2,\n        ),\n        albumentations.Resize(image_size, image_size),\n        albumentations.OneOf(\n            [\n                JpegCompression(),\n                Downscale(scale_min=0.1, scale_max=0.15),\n            ],\n            p=0.2,\n        ),\n        IAAPiecewiseAffine(p=0.2),\n        IAASharpen(p=0.2),\n        albumentations.Cutout(\n            max_h_size=int(image_size * 0.1),\n            max_w_size=int(image_size * 0.1),\n            num_holes=5,\n            p=0.5,\n        ),\n        albumentations.Normalize(),\n    ]\n)\n\ntransforms_valid = albumentations.Compose(\n    [albumentations.Resize(image_size, image_size), albumentations.Normalize()]\n)\n```\n\n## Selected Submissions\n\nOur best submission to the competition comprised of a weighted (convex) ensemble of two models \n\n- `Multi-Head ResNet200d` and\n- `Multi-head SeResNet152d`\n\nboth pretrained on NiH data with Swish activation. The weights of the ensemble were determined by forward selection, inspired by Chris Delotte’s original implementation in the Melanoma competition. In summary, the ensemble is of the form: \n\n$$w_1 \\times \\text{(mean of predictions for model (1))} + w_2 \\times \\text{(mean of predictions for model (2))}$$\n\nwhere $w_1 = 0.595$ and $w_2 = 0.405$.\n\nOur best submission obtained a public score of 0.968 and a private score of 0.972. The notebook that illustrates the forward selection approach can be found [here](https://www.kaggle.com/khoongweihao/ranzcr-forward-selection-oof-ensemble?scriptVersionId=56891135), but at version 5 for the weights here. The dataset containing all our OOFs and respective submission files can be found [here](https://www.kaggle.com/khoongweihao/ranzcr-oof-and-subs), where the description in the dataset contains all the models and their scores. The weights for model (1) can be found [here](https://www.kaggle.com/reighns/ranzcrweights), and the 5-folds used in inference are `multihead_resnet200d_fold0_best_loss.pth`, `multihead_resnet200d_fold1_best_AUC (1).pth`, `multihead_resnet200d_fold2_best_AUC.pth`, `multihead_resnet200d_fold3_best_AUC.pth`, and `multihead_resnet200d_fold4_best_AUC.pth`. The weights for model (2) can be found [here](https://www.kaggle.com/khoongweihao/ranzcr-multihead-model-weights), and the 5-folds used in inference are `grad_accum_multihead_seresnet152d_swish_fold0_best_AUC.pth`, `grad_accum_multihead_seresnet152d_swish_fold1_best_AUC.pth`, `grad_accum_multihead_seresnet152d_swish_fold2_best_AUC.pth`, `grad_accum_multihead_seresnet152d_swish_fold3_best_AUC.pth`, and `grad_accum_multihead_seresnet152d_swish_fold4_best_AUC.pth`.\n\nOur second selected submission obtained a public score of 0.967 and a private score of 0.971. The submission contained only model (2) above, where its 5-folds were inferenced. The weights are the same as above. Note that when inferencing the models, we used Sin’s pipeline available [here](https://www.kaggle.com/underwearfitting/resnet200d-public-benchmark-2xtta-lb0-965).\n\n---\n\n# Conclusion\n\nWhat we could have done better:\n\n- Use more variety of `classifier head` like `GeM`.\n- Use more variety of `backbone` and WE JUST CANNOT MAKE `efficietnet` work. 😐\n- Use [Neptune.ai](http://neptune.ai) to log our experiments as soon things start to get messy.\n- Experiment on 3-4 stage training.\n- Pseudo Labelling\n- Knowledge Distillation\n- Experiment more on maximizing AUC during ensembles. `rank_pct` etc.\n\n\nEDIT: I also express my gratitude to @bjoernholzhauer as well, he has provided a lot of tips and insights.\n\nThank you to Kaggle and the community for hosting this competition. I have learned so much just by standing on the shoulder of giants. @ttahara @underwearfitting @hengck23 @ammarali32 @cdeotte @haqishen to name a few. Taking their ideas and incorporating them into our own pipeline made things work. Of course, I would really like to thank my teammate and buddy @khoongweihao for all the help he provided me in the past year.",
    "1241963": "Congratulations. Thanks for great detailed description",
    "1241964": "ammarali32 Thanks, you are one of the giants whose shoulder we stood on :)",
    "1241976": "Thanks, I am really glad to hear that and you are the giants who scored so well and got on the top 2%",
    "1242035": "reighns Congratulations on Silver Finish Hongnan . Great and Detailed Solution Writeup",
    "1242058": "Congrats on your silver and thank you for the details!",
    "1242115": "Thanks a lot! Congrats to you too! Been following you since Melanoma. Haha",
    "1242117": "Thanks @usharengaraju for all the great resources. Saves us a lot of time.",
    "1242127": "Thanks for the detailed write up and congrats on silver!\n\nWe tried different approaches for rank ensembles and had a gain in our final score. It can be a worth to try. \nCongrats again!",
    "1242269": "reighns congrats on the silver medal and thanks a lot for sharing such a detailed summary!",
    "1242312": "Well done my friend, you deserve it!",
    "1245267": "did you find out the mean and std of the images as you have written \n`params:\n        mean: [0.4887381077884414]\n        std: [0.23064819430546407]`\nand by ` but using gradcam` are you meaning gradient accumulation or \"Gradient-weighted Class Activation Mapping,\"?",
    "1245276": "Great Solution \nand Congratulation on Silver!\nI never used to try any new things into backbone"
  },
  "source": "meta"
}