{
  "id": 226684,
  "title": "606th place: good CV & ensembling don't make up for mediocre individual models",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/writeups/bj-rn-606th-place-good-cv-ensembling-don-t-make-up",
  "author_name": "",
  "post_date": "2021-03-17T10:35:17.177Z",
  "votes": 19,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I considered whether it's worth it to write this up given my finishing position and how amazing the pipelines of some of the winning solutions are, but thought that I learnt some interesting stuff.</p>\n<h1>TL;DR</h1>\n<p><strong>Diverse mediocre individual models</strong> + <strong>stratified group-5-fold CV</strong> + <strong>ensembling based on CV</strong> lead to decent results that are not good enough for a medal. My ensembling was a regularized form of arithmetic, logit-scale, rank and power averages (more details below). I trained only using main labels, no multiple training stages, no pseudo-labelling, no use of extra annotations. I added two public notebooks to my final submissions, because I just could not match their public LB. Bottom line: Rigorous CV and good ensembling only help so much, if your basic models aren't good enough.</p>\n<h1>Interesting learnings/missed opportunities</h1>\n<ul>\n<li><strong>Biggest missed opportunities</strong>: Self-distillation, using soft-labelled external data and using the extra annotations. I realized these were very serious options, but chose to focus on other things.</li>\n<li><strong>Rigorous CV + ensembling helps</strong>: One of the things that worked well for me and that was to be expected.<ul>\n<li>Despite my mediocre individual models this got me to within 0.001 of the bronze medal ranks, so that's not as bad as the rank alone sounds.</li>\n<li>There were no unpleasant surprises here: What worked in my stratified group-5-fold CV also worked on the public LB and the private LB.</li>\n<li>A lot of the basic ensembling techniques I researched and <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">mentioned</a> seemed to work well, as one would expect based on previous competitions: power averaging, rank averaging, simple arithmetic averages (on probability and logit scale) etc. (see below for more).</li></ul></li>\n<li><strong>Image size</strong>: With this task, large(-ish) image sizes seem to help a lot, in the end I used 380 x 380 to 750 x 750 for various models.</li>\n<li><strong>GPU memory/BatchNorm freezing</strong>: I used a GTX 1080 Ti and Kaggle notebooks, thus, large models like ResNet-200D were a problem. I used small batch sizes (4) + gradient accumulation (to go to 32) + mixed precision training + freezing BatchNorm layers as a workaround, which worked okay, but resulted in slightly lower performance than others got <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085\" target=\"_blank\">with better hardware</a>. Still, I could only train the larger models at home (perhaps TPU might have worked?), because it usually took about 1 hour per epoch to train e.g. ResNet-200D.</li>\n<li><strong>Single channel models</strong> (incl. without ImageNet pre-training):<ul>\n<li>I had assumed that one good approach should be to train single channel (e.g. xse_resnext18 with Mish activation, DenseNets etc.) networks from scratch with only the competition data, especially DenseNet had been heavily advocated by <a href=\"https://uwspace.uwaterloo.ca/bitstream/handle/10012/16290/Riasatian_Abtin.pdf?sequence=3&amp;isAllowed=y\" target=\"_blank\">one thesis</a> that I read. I did not get good results with this (probably my fault).</li>\n<li>I also failed to get anywhere with single channel Convolution - BatchNorm - ReLU (CBR) networks (see <a href=\"https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html\" target=\"_blank\">this Google AI blog</a> on this, it is explored in much more detail in <a href=\"https://maithraraghu.com/assets/files/thesis_final.pdf\" target=\"_blank\">this thesis</a>) trained from scratch.</li>\n<li>Perhaps I needed to pre-train such models on large external datasets (either grayscale ImageNet or one of the larger chest X-ray datasets)?</li>\n<li>By the way, why is there not a <strong>repository of grayscale ImageNet models</strong>?? It seems obvious and there's some <a href=\"https://openaccess.thecvf.com/content_eccv_2018_workshops/w33/html/Xie_Pre-training_on_Grayscale_ImageNet_Improves_Medical_Image_Classification_ECCVW_2018_paper.html\" target=\"_blank\">papers to support this</a> that pre-training on grayscale ImageNet should help for medical image classification (besides making models smaller).</li></ul></li>\n<li><strong>You need time</strong>, especially if you don't have an existing good pipeline for X-ray images, yet. <ul>\n<li>Certainly, in terms of competition standings entering so late was not helpful for me. However, since the main purpose of Kaggle for me is to learn, that's alright since I learnt a lot.</li>\n<li>I concentrated on other competitions and only seriously started with 3 weeks to go. I guess that was not a lot of time, especially given my limited hardware.</li>\n<li>I actually \"entered\"  on day 1, when some work colleagues asked me about this new Kaggle competition. So, I tried how one of my sets of <code>fastai</code> <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">training</a> and <a href=\"https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4\" target=\"_blank\">inference</a> notebooks from the Cassava Leaf Disease competition would perform. Initially, that was the second best public notebook score and scored in the bronze area, but it very quickly dropped to the bottom of the LB. </li></ul></li>\n<li><strong>Classification head</strong>: While with smaller models a more complex 2-3 layer classification head worked great (but usually required training the head for 2-3 epochs first with the rest of the model frozen), I ran into trouble with its BatchNorm layers, when using small batch sizes, so for the large models I did not use it (without it, just training the final layer with the rest frozen did not seem to be useful).</li>\n<li><strong>Inference time</strong>: Mixed precision inference helped and seemed to not affect performance in a meaningful way. It was interesting to see what other teams did like only including models from some folds.</li>\n</ul>\n<h1>Best versions of individual models (not all in final ensemble)</h1>\n<p>As mentioned, my CV - public LB correlation seemed decent. This plot gives an idea and  shows some of my main models (more details in the table below):<br>\n<img src=\"https://i.imgur.com/X0i7TUt.jpg\" alt=\"Plot of CV scores vs. public LB score\"></p>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>Model</th>\n<th>Pre-trained</th>\n<th>Size</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0*</td>\n<td>EfficientNet-B4</td>\n<td>ImageNet</td>\n<td>380</td>\n<td>?</td>\n<td>0.945</td>\n<td>0.953</td>\n</tr>\n<tr>\n<td>1</td>\n<td>ResNeXt-50-32x4d (minimal aug.)</td>\n<td>ImageNet</td>\n<td>600</td>\n<td>0.9344</td>\n<td>0.949</td>\n<td>0.954</td>\n</tr>\n<tr>\n<td>2</td>\n<td>ResNeXt-50-32x4d (more aug.)</td>\n<td>ImageNet</td>\n<td>600</td>\n<td>0.9347</td>\n<td>0.950</td>\n<td>0.955</td>\n</tr>\n<tr>\n<td>3</td>\n<td>EfficientNet-B2 noisy student</td>\n<td>ImageNet+</td>\n<td>260</td>\n<td>0.9011</td>\n<td>0.925</td>\n<td>0.929</td>\n</tr>\n<tr>\n<td>4</td>\n<td>EfficientNet-B4 noisy student</td>\n<td>ImageNet+</td>\n<td>380</td>\n<td>0.9207</td>\n<td>0.943</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>DenseNet-blur-121d</td>\n<td>ImageNet</td>\n<td>224</td>\n<td>0.8717</td>\n<td>?</td>\n<td>?</td>\n</tr>\n<tr>\n<td>6</td>\n<td>DenseNet-blur-121d (with SWA)</td>\n<td>ImageNet</td>\n<td>224</td>\n<td>0.8699</td>\n<td>?</td>\n<td>?</td>\n</tr>\n<tr>\n<td>7</td>\n<td>EfficientNet-B5</td>\n<td>ImageNet</td>\n<td>512</td>\n<td>0.922</td>\n<td>0.943**</td>\n<td>0.948</td>\n</tr>\n<tr>\n<td>8</td>\n<td>Xception</td>\n<td>ImageNet</td>\n<td>750</td>\n<td>0.936</td>\n<td>0.957</td>\n<td>0.958</td>\n</tr>\n<tr>\n<td>9</td>\n<td>Inception v3</td>\n<td>ImageNet</td>\n<td>704</td>\n<td>0.936</td>\n<td>0.954</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>10</td>\n<td>ResNet-200D</td>\n<td>ImageNet</td>\n<td>640</td>\n<td>0.947</td>\n<td>0.958</td>\n<td>0.963</td>\n</tr>\n<tr>\n<td>11</td>\n<td>SE-ResNet-152D</td>\n<td>ImageNet</td>\n<td>704</td>\n<td>0.948</td>\n<td>0.960</td>\n<td>0.963</td>\n</tr>\n<tr>\n<td>Y*</td>\n<td>ResNet-200D</td>\n<td>ImageNet</td>\n<td>512</td>\n<td>?</td>\n<td>0.965</td>\n<td>0.967</td>\n</tr>\n<tr>\n<td>Z*</td>\n<td>SE-ResNet-152D-320</td>\n<td>ImageNet</td>\n<td>640</td>\n<td>?</td>\n<td>0.962</td>\n<td>0.964</td>\n</tr>\n<tr>\n<td>Y+Z*</td>\n<td>Blended 0.6:0.4</td>\n<td>-</td>\n<td>-</td>\n<td>?</td>\n<td>0.965</td>\n<td>0.968</td>\n</tr>\n</tbody>\n</table>\n<p>The models marked with a * were not part of my cross-validation scheme. Y and Z are public notebooks that I added to one of my final submissions. For other models I got a 5-fold CV using the same stratified group-5-fold CV scheme that I could then use for ensembling. For several models TTA did not make much of a difference, for the EfficientNet-B5 marked with ** it made a difference from 0.940 to 0.943 on the public LB. </p>\n<h1>Ensembling model(s)</h1>\n<p>I boosted my CV by 0.008 over my best individual model and the LB by the exact same amount. <br>\nThe figure below shows the correlation of the public LB predictions of the different models. Unsurprisingly my CV (and LB) scores did not improve, when I used models with correlations &gt;&gt;0.95 together. That's how models 1 and 9 were elimintated from the final ensemble, while models 3, 5 and 6 were just too weak to add much.</p>\n<p><img src=\"https://i.imgur.com/2MIvXcN.jpg\" alt=\"Correlation of public LB predictions\"></p>\n<p>My main interesting idea was this:</p>\n<ul>\n<li>I saw different models performed differently on different targets in CV. So, simple unweighted averaging are likely non-ideal.</li>\n<li>I did not want to simply optimize blending weights on the full set of OOF predictions in order to avoid overfitting. </li>\n<li>I created an averaging model (very simple, there's just models * targets parameters that are soft-maxed to force them to sum to 1 for each target) that I fitted by fold. To avoid the overfitting, I used weight decay:<ul>\n<li>weight decay of about 1 was good for a batch size of 128 in terms of maximizing mean ROC curve - SD ROC curve across folds</li>\n<li>We penalize the non-softmaxed weighting coefficients (\"betas\") towards all being equally zero. I.e. we are basically penalizing towards a simple arithmetic average and any deviation from that has to be robust across folds to be \"accepted\". </li></ul></li>\n</ul>\n<pre><code>class MyAverager(nn.Module):\n    def __init__(self, n_models, n_targets):         \n        super(MyAverager, self).__init__()\n        self.betas = nn.Parameter(torch.randn(size=(n_models, n_targets)))\n        self.softmax = nn.Softmax(dim=0)        \n\n    def forward(self, inputs):\n        # Assume input tensors indexed as sample, model, target\n        # self.betas for weights indexed as model, target\n        wgts = self.softmax(self.betas)\n        x = torch.mul(inputs, wgts).sum(dim=1)        \n        return x\n</code></pre>\n<p>I did the same thing for logit transformed model predictions and other transformations.</p>\n<p>The public notebooks I added, I just added with weights that I guesstimated based on their performance on the LB vs. the performance of my CVed models. I looked at how much weight my ensembling model gives to models with that kind of gap and picked those weights with some discounting (I should trust my CV more than a public notebook selected on public LB).</p>\n<h1>Selected submissions</h1>\n<p>I managed to select my 2nd (no real difference to the best score) and 7th best scores on the private LB, so the submission selection based on my CV went well.</p>\n<ul>\n<li>Selection 1:<ul>\n<li>First level models: 2,4,7,8,10,11,Y and Z</li>\n<li>Second level (optimized using PyTorch averaging model; public NBs added with 1.5 x weight): weighted average of probabilities, weigthed power (0.76) average, weighted rank averaging</li>\n<li>Third level: Equally weighted rank averaging</li>\n<li>Public LB 0.966, private LB 0.969</li></ul></li>\n<li>Selection 2:<ul>\n<li>First level models: 2,4,7,8,10,11,Y and Z</li>\n<li>Second level stacking (optimized using PyTorch averaging model; public NBs added with 1 x weight): weighted average of probabilities, two types of weigthed power average (power 0.57735 and 0.76), weighted rank averaging</li>\n<li>Third level: Equally weighted rank averaging</li>\n<li>Public LB 0.966, private LB 0.968</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1241981",
      "postDate": "03/17/2021 10:12:10",
      "content": "<p>I considered whether it's worth it to write this up given my finishing position and how amazing the pipelines of some of the winning solutions are, but thought that I learnt some interesting stuff.</p>\n<h1>TL;DR</h1>\n<p><strong>Diverse mediocre individual models</strong> + <strong>stratified group-5-fold CV</strong> + <strong>ensembling based on CV</strong> lead to decent results that are not good enough for a medal. My ensembling was a regularized form of arithmetic, logit-scale, rank and power averages (more details below). I trained only using main labels, no multiple training stages, no pseudo-labelling, no use of extra annotations. I added two public notebooks to my final submissions, because I just could not match their public LB. Bottom line: Rigorous CV and good ensembling only help so much, if your basic models aren't good enough.</p>\n<h1>Interesting learnings/missed opportunities</h1>\n<ul>\n<li><strong>Biggest missed opportunities</strong>: Self-distillation, using soft-labelled external data and using the extra annotations. I realized these were very serious options, but chose to focus on other things.</li>\n<li><strong>Rigorous CV + ensembling helps</strong>: One of the things that worked well for me and that was to be expected.<ul>\n<li>Despite my mediocre individual models this got me to within 0.001 of the bronze medal ranks, so that's not as bad as the rank alone sounds.</li>\n<li>There were no unpleasant surprises here: What worked in my stratified group-5-fold CV also worked on the public LB and the private LB.</li>\n<li>A lot of the basic ensembling techniques I researched and <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221\" target=\"_blank\">mentioned</a> seemed to work well, as one would expect based on previous competitions: power averaging, rank averaging, simple arithmetic averages (on probability and logit scale) etc. (see below for more).</li></ul></li>\n<li><strong>Image size</strong>: With this task, large(-ish) image sizes seem to help a lot, in the end I used 380 x 380 to 750 x 750 for various models.</li>\n<li><strong>GPU memory/BatchNorm freezing</strong>: I used a GTX 1080 Ti and Kaggle notebooks, thus, large models like ResNet-200D were a problem. I used small batch sizes (4) + gradient accumulation (to go to 32) + mixed precision training + freezing BatchNorm layers as a workaround, which worked okay, but resulted in slightly lower performance than others got <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085\" target=\"_blank\">with better hardware</a>. Still, I could only train the larger models at home (perhaps TPU might have worked?), because it usually took about 1 hour per epoch to train e.g. ResNet-200D.</li>\n<li><strong>Single channel models</strong> (incl. without ImageNet pre-training):<ul>\n<li>I had assumed that one good approach should be to train single channel (e.g. xse_resnext18 with Mish activation, DenseNets etc.) networks from scratch with only the competition data, especially DenseNet had been heavily advocated by <a href=\"https://uwspace.uwaterloo.ca/bitstream/handle/10012/16290/Riasatian_Abtin.pdf?sequence=3&amp;isAllowed=y\" target=\"_blank\">one thesis</a> that I read. I did not get good results with this (probably my fault).</li>\n<li>I also failed to get anywhere with single channel Convolution - BatchNorm - ReLU (CBR) networks (see <a href=\"https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html\" target=\"_blank\">this Google AI blog</a> on this, it is explored in much more detail in <a href=\"https://maithraraghu.com/assets/files/thesis_final.pdf\" target=\"_blank\">this thesis</a>) trained from scratch.</li>\n<li>Perhaps I needed to pre-train such models on large external datasets (either grayscale ImageNet or one of the larger chest X-ray datasets)?</li>\n<li>By the way, why is there not a <strong>repository of grayscale ImageNet models</strong>?? It seems obvious and there's some <a href=\"https://openaccess.thecvf.com/content_eccv_2018_workshops/w33/html/Xie_Pre-training_on_Grayscale_ImageNet_Improves_Medical_Image_Classification_ECCVW_2018_paper.html\" target=\"_blank\">papers to support this</a> that pre-training on grayscale ImageNet should help for medical image classification (besides making models smaller).</li></ul></li>\n<li><strong>You need time</strong>, especially if you don't have an existing good pipeline for X-ray images, yet. <ul>\n<li>Certainly, in terms of competition standings entering so late was not helpful for me. However, since the main purpose of Kaggle for me is to learn, that's alright since I learnt a lot.</li>\n<li>I concentrated on other competitions and only seriously started with 3 weeks to go. I guess that was not a lot of time, especially given my limited hardware.</li>\n<li>I actually \"entered\"  on day 1, when some work colleagues asked me about this new Kaggle competition. So, I tried how one of my sets of <code>fastai</code> <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">training</a> and <a href=\"https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4\" target=\"_blank\">inference</a> notebooks from the Cassava Leaf Disease competition would perform. Initially, that was the second best public notebook score and scored in the bronze area, but it very quickly dropped to the bottom of the LB. </li></ul></li>\n<li><strong>Classification head</strong>: While with smaller models a more complex 2-3 layer classification head worked great (but usually required training the head for 2-3 epochs first with the rest of the model frozen), I ran into trouble with its BatchNorm layers, when using small batch sizes, so for the large models I did not use it (without it, just training the final layer with the rest frozen did not seem to be useful).</li>\n<li><strong>Inference time</strong>: Mixed precision inference helped and seemed to not affect performance in a meaningful way. It was interesting to see what other teams did like only including models from some folds.</li>\n</ul>\n<h1>Best versions of individual models (not all in final ensemble)</h1>\n<p>As mentioned, my CV - public LB correlation seemed decent. This plot gives an idea and  shows some of my main models (more details in the table below):<br>\n<img src=\"https://i.imgur.com/X0i7TUt.jpg\" alt=\"Plot of CV scores vs. public LB score\"></p>\n<table>\n<thead>\n<tr>\n<th>#</th>\n<th>Model</th>\n<th>Pre-trained</th>\n<th>Size</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0*</td>\n<td>EfficientNet-B4</td>\n<td>ImageNet</td>\n<td>380</td>\n<td>?</td>\n<td>0.945</td>\n<td>0.953</td>\n</tr>\n<tr>\n<td>1</td>\n<td>ResNeXt-50-32x4d (minimal aug.)</td>\n<td>ImageNet</td>\n<td>600</td>\n<td>0.9344</td>\n<td>0.949</td>\n<td>0.954</td>\n</tr>\n<tr>\n<td>2</td>\n<td>ResNeXt-50-32x4d (more aug.)</td>\n<td>ImageNet</td>\n<td>600</td>\n<td>0.9347</td>\n<td>0.950</td>\n<td>0.955</td>\n</tr>\n<tr>\n<td>3</td>\n<td>EfficientNet-B2 noisy student</td>\n<td>ImageNet+</td>\n<td>260</td>\n<td>0.9011</td>\n<td>0.925</td>\n<td>0.929</td>\n</tr>\n<tr>\n<td>4</td>\n<td>EfficientNet-B4 noisy student</td>\n<td>ImageNet+</td>\n<td>380</td>\n<td>0.9207</td>\n<td>0.943</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>DenseNet-blur-121d</td>\n<td>ImageNet</td>\n<td>224</td>\n<td>0.8717</td>\n<td>?</td>\n<td>?</td>\n</tr>\n<tr>\n<td>6</td>\n<td>DenseNet-blur-121d (with SWA)</td>\n<td>ImageNet</td>\n<td>224</td>\n<td>0.8699</td>\n<td>?</td>\n<td>?</td>\n</tr>\n<tr>\n<td>7</td>\n<td>EfficientNet-B5</td>\n<td>ImageNet</td>\n<td>512</td>\n<td>0.922</td>\n<td>0.943**</td>\n<td>0.948</td>\n</tr>\n<tr>\n<td>8</td>\n<td>Xception</td>\n<td>ImageNet</td>\n<td>750</td>\n<td>0.936</td>\n<td>0.957</td>\n<td>0.958</td>\n</tr>\n<tr>\n<td>9</td>\n<td>Inception v3</td>\n<td>ImageNet</td>\n<td>704</td>\n<td>0.936</td>\n<td>0.954</td>\n<td>0.956</td>\n</tr>\n<tr>\n<td>10</td>\n<td>ResNet-200D</td>\n<td>ImageNet</td>\n<td>640</td>\n<td>0.947</td>\n<td>0.958</td>\n<td>0.963</td>\n</tr>\n<tr>\n<td>11</td>\n<td>SE-ResNet-152D</td>\n<td>ImageNet</td>\n<td>704</td>\n<td>0.948</td>\n<td>0.960</td>\n<td>0.963</td>\n</tr>\n<tr>\n<td>Y*</td>\n<td>ResNet-200D</td>\n<td>ImageNet</td>\n<td>512</td>\n<td>?</td>\n<td>0.965</td>\n<td>0.967</td>\n</tr>\n<tr>\n<td>Z*</td>\n<td>SE-ResNet-152D-320</td>\n<td>ImageNet</td>\n<td>640</td>\n<td>?</td>\n<td>0.962</td>\n<td>0.964</td>\n</tr>\n<tr>\n<td>Y+Z*</td>\n<td>Blended 0.6:0.4</td>\n<td>-</td>\n<td>-</td>\n<td>?</td>\n<td>0.965</td>\n<td>0.968</td>\n</tr>\n</tbody>\n</table>\n<p>The models marked with a * were not part of my cross-validation scheme. Y and Z are public notebooks that I added to one of my final submissions. For other models I got a 5-fold CV using the same stratified group-5-fold CV scheme that I could then use for ensembling. For several models TTA did not make much of a difference, for the EfficientNet-B5 marked with ** it made a difference from 0.940 to 0.943 on the public LB. </p>\n<h1>Ensembling model(s)</h1>\n<p>I boosted my CV by 0.008 over my best individual model and the LB by the exact same amount. <br>\nThe figure below shows the correlation of the public LB predictions of the different models. Unsurprisingly my CV (and LB) scores did not improve, when I used models with correlations &gt;&gt;0.95 together. That's how models 1 and 9 were elimintated from the final ensemble, while models 3, 5 and 6 were just too weak to add much.</p>\n<p><img src=\"https://i.imgur.com/2MIvXcN.jpg\" alt=\"Correlation of public LB predictions\"></p>\n<p>My main interesting idea was this:</p>\n<ul>\n<li>I saw different models performed differently on different targets in CV. So, simple unweighted averaging are likely non-ideal.</li>\n<li>I did not want to simply optimize blending weights on the full set of OOF predictions in order to avoid overfitting. </li>\n<li>I created an averaging model (very simple, there's just models * targets parameters that are soft-maxed to force them to sum to 1 for each target) that I fitted by fold. To avoid the overfitting, I used weight decay:<ul>\n<li>weight decay of about 1 was good for a batch size of 128 in terms of maximizing mean ROC curve - SD ROC curve across folds</li>\n<li>We penalize the non-softmaxed weighting coefficients (\"betas\") towards all being equally zero. I.e. we are basically penalizing towards a simple arithmetic average and any deviation from that has to be robust across folds to be \"accepted\". </li></ul></li>\n</ul>\n<pre><code>class MyAverager(nn.Module):\n    def __init__(self, n_models, n_targets):         \n        super(MyAverager, self).__init__()\n        self.betas = nn.Parameter(torch.randn(size=(n_models, n_targets)))\n        self.softmax = nn.Softmax(dim=0)        \n\n    def forward(self, inputs):\n        # Assume input tensors indexed as sample, model, target\n        # self.betas for weights indexed as model, target\n        wgts = self.softmax(self.betas)\n        x = torch.mul(inputs, wgts).sum(dim=1)        \n        return x\n</code></pre>\n<p>I did the same thing for logit transformed model predictions and other transformations.</p>\n<p>The public notebooks I added, I just added with weights that I guesstimated based on their performance on the LB vs. the performance of my CVed models. I looked at how much weight my ensembling model gives to models with that kind of gap and picked those weights with some discounting (I should trust my CV more than a public notebook selected on public LB).</p>\n<h1>Selected submissions</h1>\n<p>I managed to select my 2nd (no real difference to the best score) and 7th best scores on the private LB, so the submission selection based on my CV went well.</p>\n<ul>\n<li>Selection 1:<ul>\n<li>First level models: 2,4,7,8,10,11,Y and Z</li>\n<li>Second level (optimized using PyTorch averaging model; public NBs added with 1.5 x weight): weighted average of probabilities, weigthed power (0.76) average, weighted rank averaging</li>\n<li>Third level: Equally weighted rank averaging</li>\n<li>Public LB 0.966, private LB 0.969</li></ul></li>\n<li>Selection 2:<ul>\n<li>First level models: 2,4,7,8,10,11,Y and Z</li>\n<li>Second level stacking (optimized using PyTorch averaging model; public NBs added with 1 x weight): weighted average of probabilities, two types of weigthed power average (power 0.57735 and 0.76), weighted rank averaging</li>\n<li>Third level: Equally weighted rank averaging</li>\n<li>Public LB 0.966, private LB 0.968</li></ul></li>\n</ul>",
      "rawMarkdown": "I considered whether it's worth it to write this up given my finishing position and how amazing the pipelines of some of the winning solutions are, but thought that I learnt some interesting stuff.\n\n# TL;DR\n\n**Diverse mediocre individual models** + **stratified group-5-fold CV** + **ensembling based on CV** lead to decent results that are not good enough for a medal. My ensembling was a regularized form of arithmetic, logit-scale, rank and power averages (more details below). I trained only using main labels, no multiple training stages, no pseudo-labelling, no use of extra annotations. I added two public notebooks to my final submissions, because I just could not match their public LB. Bottom line: Rigorous CV and good ensembling only help so much, if your basic models aren't good enough.\n\n# Interesting learnings/missed opportunities\n\n* **Biggest missed opportunities**: Self-distillation, using soft-labelled external data and using the extra annotations. I realized these were very serious options, but chose to focus on other things.\n* **Rigorous CV + ensembling helps**: One of the things that worked well for me and that was to be expected.\n * Despite my mediocre individual models this got me to within 0.001 of the bronze medal ranks, so that's not as bad as the rank alone sounds.\n * There were no unpleasant surprises here: What worked in my stratified group-5-fold CV also worked on the public LB and the private LB.\n * A lot of the basic ensembling techniques I researched and [mentioned](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221) seemed to work well, as one would expect based on previous competitions: power averaging, rank averaging, simple arithmetic averages (on probability and logit scale) etc. (see below for more).\n* **Image size**: With this task, large(-ish) image sizes seem to help a lot, in the end I used 380 x 380 to 750 x 750 for various models.\n* **GPU memory/BatchNorm freezing**: I used a GTX 1080 Ti and Kaggle notebooks, thus, large models like ResNet-200D were a problem. I used small batch sizes (4) + gradient accumulation (to go to 32) + mixed precision training + freezing BatchNorm layers as a workaround, which worked okay, but resulted in slightly lower performance than others got [with better hardware](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085). Still, I could only train the larger models at home (perhaps TPU might have worked?), because it usually took about 1 hour per epoch to train e.g. ResNet-200D.\n* **Single channel models** (incl. without ImageNet pre-training):\n * I had assumed that one good approach should be to train single channel (e.g. xse_resnext18 with Mish activation, DenseNets etc.) networks from scratch with only the competition data, especially DenseNet had been heavily advocated by [one thesis](https://uwspace.uwaterloo.ca/bitstream/handle/10012/16290/Riasatian_Abtin.pdf?sequence=3&isAllowed=y) that I read. I did not get good results with this (probably my fault).\n * I also failed to get anywhere with single channel Convolution - BatchNorm - ReLU (CBR) networks (see [this Google AI blog](https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html) on this, it is explored in much more detail in [this thesis](https://maithraraghu.com/assets/files/thesis_final.pdf)) trained from scratch.\n * Perhaps I needed to pre-train such models on large external datasets (either grayscale ImageNet or one of the larger chest X-ray datasets)?\n * By the way, why is there not a **repository of grayscale ImageNet models**?? It seems obvious and there's some [papers to support this](https://openaccess.thecvf.com/content_eccv_2018_workshops/w33/html/Xie_Pre-training_on_Grayscale_ImageNet_Improves_Medical_Image_Classification_ECCVW_2018_paper.html) that pre-training on grayscale ImageNet should help for medical image classification (besides making models smaller).\n* **You need time**, especially if you don't have an existing good pipeline for X-ray images, yet. \n * Certainly, in terms of competition standings entering so late was not helpful for me. However, since the main purpose of Kaggle for me is to learn, that's alright since I learnt a lot.\n * I concentrated on other competitions and only seriously started with 3 weeks to go. I guess that was not a lot of time, especially given my limited hardware.\n * I actually \"entered\"  on day 1, when some work colleagues asked me about this new Kaggle competition. So, I tried how one of my sets of `fastai` [training](https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb) and [inference](https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4) notebooks from the Cassava Leaf Disease competition would perform. Initially, that was the second best public notebook score and scored in the bronze area, but it very quickly dropped to the bottom of the LB. \n* **Classification head**: While with smaller models a more complex 2-3 layer classification head worked great (but usually required training the head for 2-3 epochs first with the rest of the model frozen), I ran into trouble with its BatchNorm layers, when using small batch sizes, so for the large models I did not use it (without it, just training the final layer with the rest frozen did not seem to be useful).\n* **Inference time**: Mixed precision inference helped and seemed to not affect performance in a meaningful way. It was interesting to see what other teams did like only including models from some folds.\n\n# Best versions of individual models (not all in final ensemble)\n\nAs mentioned, my CV - public LB correlation seemed decent. This plot gives an idea and  shows some of my main models (more details in the table below):\n![Plot of CV scores vs. public LB score](https://i.imgur.com/X0i7TUt.jpg)\n\n| # | Model | Pre-trained | Size | CV | Public | Private |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0* | EfficientNet-B4 | ImageNet | 380 | ? | 0.945 | 0.953 |\n| 1 | ResNeXt-50-32x4d (minimal aug.) | ImageNet | 600 | 0.9344 | 0.949 | 0.954 |\n| 2 | ResNeXt-50-32x4d (more aug.) | ImageNet | 600 | 0.9347 | 0.950 | 0.955 |\n| 3 | EfficientNet-B2 noisy student | ImageNet+ | 260 | 0.9011 | 0.925 | 0.929 |\n|4 | EfficientNet-B4 noisy student | ImageNet+ | 380 | 0.9207 | 0.943 | | 0.948 |\n|5 | DenseNet-blur-121d | ImageNet | 224 | 0.8717 | ? | ? |\n|6 | DenseNet-blur-121d (with SWA) | ImageNet | 224 | 0.8699 | ? | ? |\n|7 | EfficientNet-B5 | ImageNet | 512 | 0.922 | 0.943** | 0.948 |\n|8 | Xception | ImageNet | 750 | 0.936 | 0.957 | 0.958 |\n|9 | Inception v3 | ImageNet | 704 | 0.936 | 0.954 | 0.956 |\n|10 | ResNet-200D | ImageNet | 640 | 0.947 | 0.958 | 0.963 |\n| 11 | SE-ResNet-152D | ImageNet | 704 | 0.948 | 0.960 | 0.963 |\n|Y* | ResNet-200D | ImageNet | 512 | ? | 0.965 | 0.967 |\n|Z* | SE-ResNet-152D-320 | ImageNet | 640 | ? | 0.962 | 0.964 |\n|Y+Z*| Blended 0.6:0.4 | - | - | ? | 0.965 | 0.968 |\n\nThe models marked with a * were not part of my cross-validation scheme. Y and Z are public notebooks that I added to one of my final submissions. For other models I got a 5-fold CV using the same stratified group-5-fold CV scheme that I could then use for ensembling. For several models TTA did not make much of a difference, for the EfficientNet-B5 marked with ** it made a difference from 0.940 to 0.943 on the public LB. \n\n\n\n\n# Ensembling model(s)\n\nI boosted my CV by 0.008 over my best individual model and the LB by the exact same amount. \nThe figure below shows the correlation of the public LB predictions of the different models. Unsurprisingly my CV (and LB) scores did not improve, when I used models with correlations >>0.95 together. That's how models 1 and 9 were elimintated from the final ensemble, while models 3, 5 and 6 were just too weak to add much.\n\n![Correlation of public LB predictions](https://i.imgur.com/2MIvXcN.jpg)\n\nMy main interesting idea was this:\n* I saw different models performed differently on different targets in CV. So, simple unweighted averaging are likely non-ideal.\n* I did not want to simply optimize blending weights on the full set of OOF predictions in order to avoid overfitting. \n* I created an averaging model (very simple, there's just models * targets parameters that are soft-maxed to force them to sum to 1 for each target) that I fitted by fold. To avoid the overfitting, I used weight decay:\n * weight decay of about 1 was good for a batch size of 128 in terms of maximizing mean ROC curve - SD ROC curve across folds\n * We penalize the non-softmaxed weighting coefficients (\"betas\") towards all being equally zero. I.e. we are basically penalizing towards a simple arithmetic average and any deviation from that has to be robust across folds to be \"accepted\". \n```\nclass MyAverager(nn.Module):\n    def __init__(self, n_models, n_targets):         \n        super(MyAverager, self).__init__()\n        self.betas = nn.Parameter(torch.randn(size=(n_models, n_targets)))\n        self.softmax = nn.Softmax(dim=0)        \n\n    def forward(self, inputs):\n        # Assume input tensors indexed as sample, model, target\n        # self.betas for weights indexed as model, target\n        wgts = self.softmax(self.betas)\n        x = torch.mul(inputs, wgts).sum(dim=1)        \n        return x\n```\nI did the same thing for logit transformed model predictions and other transformations.\n\nThe public notebooks I added, I just added with weights that I guesstimated based on their performance on the LB vs. the performance of my CVed models. I looked at how much weight my ensembling model gives to models with that kind of gap and picked those weights with some discounting (I should trust my CV more than a public notebook selected on public LB).\n\n# Selected submissions\n\nI managed to select my 2nd (no real difference to the best score) and 7th best scores on the private LB, so the submission selection based on my CV went well.\n\n* Selection 1:\n * First level models: 2,4,7,8,10,11,Y and Z\n * Second level (optimized using PyTorch averaging model; public NBs added with 1.5 x weight): weighted average of probabilities, weigthed power (0.76) average, weighted rank averaging\n * Third level: Equally weighted rank averaging\n * Public LB 0.966, private LB 0.969\n* Selection 2:\n * First level models: 2,4,7,8,10,11,Y and Z\n * Second level stacking (optimized using PyTorch averaging model; public NBs added with 1 x weight): weighted average of probabilities, two types of weigthed power average (power 0.57735 and 0.76), weighted rank averaging\n * Third level: Equally weighted rank averaging\n * Public LB 0.966, private LB 0.968",
      "votes": null
    },
    {
      "id": "1242084",
      "postDate": "03/17/2021 11:47:54",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Thanks for sharing the detailed solution writeup</p>",
      "rawMarkdown": "bjoernholzhauer Thanks for sharing the detailed solution writeup",
      "votes": null
    },
    {
      "id": "1242195",
      "postDate": "03/17/2021 13:23:34",
      "content": "<p>I really like your lucid posts and answers. Looking forward to team up one day. And yes, I’ve learnt a lot from your replies :) </p>\n<p>Thanks!</p>",
      "rawMarkdown": "I really like your lucid posts and answers. Looking forward to team up one day. And yes, I’ve learnt a lot from your replies :) \n\nThanks!",
      "votes": null
    },
    {
      "id": "1242257",
      "postDate": "03/17/2021 13:53:25",
      "content": "<p>Thanks. Congratulations on placing high up in the silver medals!</p>",
      "rawMarkdown": "Thanks. Congratulations on placing high up in the silver medals!",
      "votes": null
    },
    {
      "id": "1244158",
      "postDate": "03/18/2021 18:16:29",
      "content": "<p>Your solution really looks very impressive. I wish to try it. maybe you could have try using external dataset</p>",
      "rawMarkdown": "Your solution really looks very impressive. I wish to try it. maybe you could have try using external dataset",
      "votes": null
    },
    {
      "id": "1245754",
      "postDate": "03/20/2021 05:41:41",
      "content": "<p>Nice writeup.  Thanks for sharing your clear, well laid out process.  Please do continue to  post.</p>",
      "rawMarkdown": "Nice writeup.  Thanks for sharing your clear, well laid out process.  Please do continue to  post.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1242084,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/17/2021 11:47:54",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Thanks for sharing the detailed solution writeup</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242195,
      "author_name": "reighns",
      "author_url": "",
      "post_date": "03/17/2021 13:23:34",
      "content": "<p>I really like your lucid posts and answers. Looking forward to team up one day. And yes, I’ve learnt a lot from your replies :) </p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1242257,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/17/2021 13:53:25",
          "content": "<p>Thanks. Congratulations on placing high up in the silver medals!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1244158,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "03/18/2021 18:16:29",
      "content": "<p>Your solution really looks very impressive. I wish to try it. maybe you could have try using external dataset</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1245754,
      "author_name": "socated",
      "author_url": "",
      "post_date": "03/20/2021 05:41:41",
      "content": "<p>Nice writeup.  Thanks for sharing your clear, well laid out process.  Please do continue to  post.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1241981": "I considered whether it's worth it to write this up given my finishing position and how amazing the pipelines of some of the winning solutions are, but thought that I learnt some interesting stuff.\n\n# TL;DR\n\n**Diverse mediocre individual models** + **stratified group-5-fold CV** + **ensembling based on CV** lead to decent results that are not good enough for a medal. My ensembling was a regularized form of arithmetic, logit-scale, rank and power averages (more details below). I trained only using main labels, no multiple training stages, no pseudo-labelling, no use of extra annotations. I added two public notebooks to my final submissions, because I just could not match their public LB. Bottom line: Rigorous CV and good ensembling only help so much, if your basic models aren't good enough.\n\n# Interesting learnings/missed opportunities\n\n* **Biggest missed opportunities**: Self-distillation, using soft-labelled external data and using the extra annotations. I realized these were very serious options, but chose to focus on other things.\n* **Rigorous CV + ensembling helps**: One of the things that worked well for me and that was to be expected.\n * Despite my mediocre individual models this got me to within 0.001 of the bronze medal ranks, so that's not as bad as the rank alone sounds.\n * There were no unpleasant surprises here: What worked in my stratified group-5-fold CV also worked on the public LB and the private LB.\n * A lot of the basic ensembling techniques I researched and [mentioned](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211221) seemed to work well, as one would expect based on previous competitions: power averaging, rank averaging, simple arithmetic averages (on probability and logit scale) etc. (see below for more).\n* **Image size**: With this task, large(-ish) image sizes seem to help a lot, in the end I used 380 x 380 to 750 x 750 for various models.\n* **GPU memory/BatchNorm freezing**: I used a GTX 1080 Ti and Kaggle notebooks, thus, large models like ResNet-200D were a problem. I used small batch sizes (4) + gradient accumulation (to go to 32) + mixed precision training + freezing BatchNorm layers as a workaround, which worked okay, but resulted in slightly lower performance than others got [with better hardware](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224085). Still, I could only train the larger models at home (perhaps TPU might have worked?), because it usually took about 1 hour per epoch to train e.g. ResNet-200D.\n* **Single channel models** (incl. without ImageNet pre-training):\n * I had assumed that one good approach should be to train single channel (e.g. xse_resnext18 with Mish activation, DenseNets etc.) networks from scratch with only the competition data, especially DenseNet had been heavily advocated by [one thesis](https://uwspace.uwaterloo.ca/bitstream/handle/10012/16290/Riasatian_Abtin.pdf?sequence=3&isAllowed=y) that I read. I did not get good results with this (probably my fault).\n * I also failed to get anywhere with single channel Convolution - BatchNorm - ReLU (CBR) networks (see [this Google AI blog](https://ai.googleblog.com/2019/12/understanding-transfer-learning-for.html) on this, it is explored in much more detail in [this thesis](https://maithraraghu.com/assets/files/thesis_final.pdf)) trained from scratch.\n * Perhaps I needed to pre-train such models on large external datasets (either grayscale ImageNet or one of the larger chest X-ray datasets)?\n * By the way, why is there not a **repository of grayscale ImageNet models**?? It seems obvious and there's some [papers to support this](https://openaccess.thecvf.com/content_eccv_2018_workshops/w33/html/Xie_Pre-training_on_Grayscale_ImageNet_Improves_Medical_Image_Classification_ECCVW_2018_paper.html) that pre-training on grayscale ImageNet should help for medical image classification (besides making models smaller).\n* **You need time**, especially if you don't have an existing good pipeline for X-ray images, yet. \n * Certainly, in terms of competition standings entering so late was not helpful for me. However, since the main purpose of Kaggle for me is to learn, that's alright since I learnt a lot.\n * I concentrated on other competitions and only seriously started with 3 weeks to go. I guess that was not a lot of time, especially given my limited hardware.\n * I actually \"entered\"  on day 1, when some work colleagues asked me about this new Kaggle competition. So, I tried how one of my sets of `fastai` [training](https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb) and [inference](https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4) notebooks from the Cassava Leaf Disease competition would perform. Initially, that was the second best public notebook score and scored in the bronze area, but it very quickly dropped to the bottom of the LB. \n* **Classification head**: While with smaller models a more complex 2-3 layer classification head worked great (but usually required training the head for 2-3 epochs first with the rest of the model frozen), I ran into trouble with its BatchNorm layers, when using small batch sizes, so for the large models I did not use it (without it, just training the final layer with the rest frozen did not seem to be useful).\n* **Inference time**: Mixed precision inference helped and seemed to not affect performance in a meaningful way. It was interesting to see what other teams did like only including models from some folds.\n\n# Best versions of individual models (not all in final ensemble)\n\nAs mentioned, my CV - public LB correlation seemed decent. This plot gives an idea and  shows some of my main models (more details in the table below):\n![Plot of CV scores vs. public LB score](https://i.imgur.com/X0i7TUt.jpg)\n\n| # | Model | Pre-trained | Size | CV | Public | Private |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0* | EfficientNet-B4 | ImageNet | 380 | ? | 0.945 | 0.953 |\n| 1 | ResNeXt-50-32x4d (minimal aug.) | ImageNet | 600 | 0.9344 | 0.949 | 0.954 |\n| 2 | ResNeXt-50-32x4d (more aug.) | ImageNet | 600 | 0.9347 | 0.950 | 0.955 |\n| 3 | EfficientNet-B2 noisy student | ImageNet+ | 260 | 0.9011 | 0.925 | 0.929 |\n|4 | EfficientNet-B4 noisy student | ImageNet+ | 380 | 0.9207 | 0.943 | | 0.948 |\n|5 | DenseNet-blur-121d | ImageNet | 224 | 0.8717 | ? | ? |\n|6 | DenseNet-blur-121d (with SWA) | ImageNet | 224 | 0.8699 | ? | ? |\n|7 | EfficientNet-B5 | ImageNet | 512 | 0.922 | 0.943** | 0.948 |\n|8 | Xception | ImageNet | 750 | 0.936 | 0.957 | 0.958 |\n|9 | Inception v3 | ImageNet | 704 | 0.936 | 0.954 | 0.956 |\n|10 | ResNet-200D | ImageNet | 640 | 0.947 | 0.958 | 0.963 |\n| 11 | SE-ResNet-152D | ImageNet | 704 | 0.948 | 0.960 | 0.963 |\n|Y* | ResNet-200D | ImageNet | 512 | ? | 0.965 | 0.967 |\n|Z* | SE-ResNet-152D-320 | ImageNet | 640 | ? | 0.962 | 0.964 |\n|Y+Z*| Blended 0.6:0.4 | - | - | ? | 0.965 | 0.968 |\n\nThe models marked with a * were not part of my cross-validation scheme. Y and Z are public notebooks that I added to one of my final submissions. For other models I got a 5-fold CV using the same stratified group-5-fold CV scheme that I could then use for ensembling. For several models TTA did not make much of a difference, for the EfficientNet-B5 marked with ** it made a difference from 0.940 to 0.943 on the public LB. \n\n\n\n\n# Ensembling model(s)\n\nI boosted my CV by 0.008 over my best individual model and the LB by the exact same amount. \nThe figure below shows the correlation of the public LB predictions of the different models. Unsurprisingly my CV (and LB) scores did not improve, when I used models with correlations >>0.95 together. That's how models 1 and 9 were elimintated from the final ensemble, while models 3, 5 and 6 were just too weak to add much.\n\n![Correlation of public LB predictions](https://i.imgur.com/2MIvXcN.jpg)\n\nMy main interesting idea was this:\n* I saw different models performed differently on different targets in CV. So, simple unweighted averaging are likely non-ideal.\n* I did not want to simply optimize blending weights on the full set of OOF predictions in order to avoid overfitting. \n* I created an averaging model (very simple, there's just models * targets parameters that are soft-maxed to force them to sum to 1 for each target) that I fitted by fold. To avoid the overfitting, I used weight decay:\n * weight decay of about 1 was good for a batch size of 128 in terms of maximizing mean ROC curve - SD ROC curve across folds\n * We penalize the non-softmaxed weighting coefficients (\"betas\") towards all being equally zero. I.e. we are basically penalizing towards a simple arithmetic average and any deviation from that has to be robust across folds to be \"accepted\". \n```\nclass MyAverager(nn.Module):\n    def __init__(self, n_models, n_targets):         \n        super(MyAverager, self).__init__()\n        self.betas = nn.Parameter(torch.randn(size=(n_models, n_targets)))\n        self.softmax = nn.Softmax(dim=0)        \n\n    def forward(self, inputs):\n        # Assume input tensors indexed as sample, model, target\n        # self.betas for weights indexed as model, target\n        wgts = self.softmax(self.betas)\n        x = torch.mul(inputs, wgts).sum(dim=1)        \n        return x\n```\nI did the same thing for logit transformed model predictions and other transformations.\n\nThe public notebooks I added, I just added with weights that I guesstimated based on their performance on the LB vs. the performance of my CVed models. I looked at how much weight my ensembling model gives to models with that kind of gap and picked those weights with some discounting (I should trust my CV more than a public notebook selected on public LB).\n\n# Selected submissions\n\nI managed to select my 2nd (no real difference to the best score) and 7th best scores on the private LB, so the submission selection based on my CV went well.\n\n* Selection 1:\n * First level models: 2,4,7,8,10,11,Y and Z\n * Second level (optimized using PyTorch averaging model; public NBs added with 1.5 x weight): weighted average of probabilities, weigthed power (0.76) average, weighted rank averaging\n * Third level: Equally weighted rank averaging\n * Public LB 0.966, private LB 0.969\n* Selection 2:\n * First level models: 2,4,7,8,10,11,Y and Z\n * Second level stacking (optimized using PyTorch averaging model; public NBs added with 1 x weight): weighted average of probabilities, two types of weigthed power average (power 0.57735 and 0.76), weighted rank averaging\n * Third level: Equally weighted rank averaging\n * Public LB 0.966, private LB 0.968",
    "1242084": "bjoernholzhauer Thanks for sharing the detailed solution writeup",
    "1242195": "I really like your lucid posts and answers. Looking forward to team up one day. And yes, I’ve learnt a lot from your replies :) \n\nThanks!",
    "1242257": "Thanks. Congratulations on placing high up in the silver medals!",
    "1244158": "Your solution really looks very impressive. I wish to try it. maybe you could have try using external dataset",
    "1245754": "Nice writeup.  Thanks for sharing your clear, well laid out process.  Please do continue to  post."
  },
  "source": "meta"
}