{
  "id": 220879,
  "title": "375th place private LB (278th place public)",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/bj-rn-375th-place-private-lb-278th-place-public",
  "author_name": "",
  "post_date": "2021-02-20T00:13:12.817Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So, after the latest updates, I moved up a few places on the LB and slipped into the bronze medals. Within two days I've gone from zero medals to my second bronze medal. 😃 On this competition, I worked on my own, while on the  Rainforest Connection Species Audio Detection competition](<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection</a>) I had the pleasure to work with a great team. So, I thought I'd write up some of the more interesting things I tried on this competition.</p>\n<p>My best selected solution was a weighted power averaging ensemble (i.e. squaring the predictions of each model and then averaging, followed by picking the class with the highest score) of</p>\n<ul>\n<li><strong>EfficientNet-B5</strong> (i.e. not noisy student, somehow my noisy student version seemed worse on CV - I failed to figure out why) on image size 456 by 456 with TTA</li>\n<li><strong>EfficientNet-B4-noisy student</strong> on image size 512 by 512 with TTA</li>\n<li><strong>ResNeXt-50 (32x4d)</strong> with image size 512 by 512</li>\n<li><strong>Vision Transformer</strong> (<code>vit_base_patch16_224</code>) image size 224 by 224</li>\n</ul>\n<p>I did two different test-time augmentation approaches for the two EfficientNets and no TTA for the others (yes, yes, I clearly have an EfficientNet bias).</p>\n<p><strong>Things that worked</strong>:</p>\n<ul>\n<li><strong>label smoothing cross-entropy</strong> with epsilon of 0.1 to 0.2 or so as a loss function (not surprising, given that it was originally proposed for noisy labels)</li>\n<li><strong>focal loss</strong> (not in my selected solutions, but in CV it was very similar to label smoothing cross-entropy)</li>\n<li><strong>Ensembling</strong> (not surprise there, either)<ul>\n<li><strong>Power averaging</strong> was behind my best selected submission. This looked somewhat promising on CV and not that bad on the public LB (while I selected a weighted arithmetic mean based on LB, which did worse - in part, I also just wanted to select a second submission that was not too similar to the other one). </li>\n<li>I got some even better private LB numbers with <strong>rank averaging</strong> (in silver territory, but I did not select it), but admittedly my very best private LB score was a simple weighted arithmetic mean of two of the models. See below for what did not work so well (i.e. selection the best submission)…</li>\n<li>Why did I think power and rank averaging would work? I kind of figured that accuracy is in some sense a rank based metric (you only care about the first rank of course) and those methods had just been helpful in the <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">Rainforest Connection Species Audio Detection competition</a> that used mean average precision (a very rank based metric).</li></ul></li>\n<li>Using <strong>diverse models</strong> in my ensemble, even if it does hurt on the public LB (I hedged my bets a bit and tried to bet on a half-way house that probably relied on the public LB to much, which cost me a bit in ranks).<ul>\n<li>The idea was to have diverse image sizes as inputs, somewhat diverse models and diverse augmentations during training (and TTA).</li>\n<li>Nevertheless, it seems the two EfficientNets were different enough that having both helped. I probably would not have added the EfficientNet-B4-NS, if it had not been for the public LB score of <a href=\"https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903\" target=\"_blank\">this notebook</a>. That notebook also made me stop trying other ResNet-types and just go with ResNeXt-50 (32x4d) (who knows whether that was the right decision).</li>\n<li>My choice of architectures was perhaps a bit limited I mostly picked these architectures based on what I could get as pre-trained models via <code>timm</code> and perhaps payed too much attention to what others did.  </li></ul></li>\n<li><strong>Gradient accumulation</strong> to use bigger batch sizes (72) with EfficientNet-B5 (necessary when using a GPU on Kaggle or my own 1080-Ti). It appeared to help a tiny bit in CV to use a batch size &gt;32 despite the whole \"<a href=\"https://twitter.com/ylecun/status/989610208497360896?lang=en\" target=\"_blank\">Friends dont let friends use minibatches larger than 32</a>\" <a href=\"https://arxiv.org/abs/1804.07612\" target=\"_blank\">thing</a>.</li>\n<li><strong>flat learning rate before cosine annealing learning rate schedule</strong> that I learnt about from <a href=\"https://www.kaggle.com/abhishek/tez-faster-and-easier-training-for-leaf-detection\" target=\"_blank\">Abishek Thakur's notebook</a></li>\n<li>Weight decay - whenever I tried it, it helped.</li>\n<li><code>fastai</code> + <code>timm</code> as e.g. described in this <a href=\"https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai\" target=\"_blank\">great notebook</a> by <a href=\"https://www.kaggle.com/muellerzr\" target=\"_blank\">@muellerzr</a> (thanks, again, for that nice starter kit) and PyTorch. You only appreciate how smoothly that works with CPU and GPU when you try to use the TPU with PyTorch, which kind of worked, but was a bit painful.</li>\n</ul>\n<p><strong>Things that did not work for me (mostly my fault, some of these clearly worked out better for others):</strong></p>\n<ul>\n<li>Selecting the best submissions - I did 27 submission before the deadline and I selected in terms of the private LB the 17th and 20th best of these.<ul>\n<li>I suspect I was not sufficiently disciplined with running everything I tried in the last few days through a proper CV (yes, yes, I know, that's a terrible sin…) and was too influenced by the public LB.</li>\n<li>I wonder whether those that did best that were super disciplined and perhaps even used repeated k-fold?! </li>\n<li>To some extent the ordering of my various submissions may have had a strong element of luck (accuracy is such a low information content metric and I did not manage to) and dropping 97 places is not that bad, but I aspire to do better next time.</li></ul></li>\n<li>Getting value out of the 2019 competition data, but I probably just invested insufficient time. It just always made my CV worse even after de-duplication, so I gave up.</li>\n<li>Stochastic weight averaging: it kept reliably improving my loss function in CV, but not accuracy (it seemed I got the best accuracy when very slightly overfitting in terms of (label-smoothing) cross-entropy - which I understand is sort of expected). I wonder whether I should nevertheless put a model with SWA in my ensemble.</li>\n<li>Really large models: I initially thought that the biggest EfficientNet I could train with TPU would be best, then I scaled my \"ambitions\" back to the largest one I could comfortably cross-validate (B5) some experiments for… </li>\n<li>Doing two competitions that ended within 24 hours of each other. Yes, I got two medals, but I am left feeling I would have liked to rather do one of them more deeply and do a better job / contribute more to the one where I teamed up.</li>\n</ul>",
  "messages": [
    {
      "id": "1211059",
      "postDate": "02/19/2021 23:35:15",
      "content": "<p>So, after the latest updates, I moved up a few places on the LB and slipped into the bronze medals. Within two days I've gone from zero medals to my second bronze medal. 😃 On this competition, I worked on my own, while on the  Rainforest Connection Species Audio Detection competition](<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection</a>) I had the pleasure to work with a great team. So, I thought I'd write up some of the more interesting things I tried on this competition.</p>\n<p>My best selected solution was a weighted power averaging ensemble (i.e. squaring the predictions of each model and then averaging, followed by picking the class with the highest score) of</p>\n<ul>\n<li><strong>EfficientNet-B5</strong> (i.e. not noisy student, somehow my noisy student version seemed worse on CV - I failed to figure out why) on image size 456 by 456 with TTA</li>\n<li><strong>EfficientNet-B4-noisy student</strong> on image size 512 by 512 with TTA</li>\n<li><strong>ResNeXt-50 (32x4d)</strong> with image size 512 by 512</li>\n<li><strong>Vision Transformer</strong> (<code>vit_base_patch16_224</code>) image size 224 by 224</li>\n</ul>\n<p>I did two different test-time augmentation approaches for the two EfficientNets and no TTA for the others (yes, yes, I clearly have an EfficientNet bias).</p>\n<p><strong>Things that worked</strong>:</p>\n<ul>\n<li><strong>label smoothing cross-entropy</strong> with epsilon of 0.1 to 0.2 or so as a loss function (not surprising, given that it was originally proposed for noisy labels)</li>\n<li><strong>focal loss</strong> (not in my selected solutions, but in CV it was very similar to label smoothing cross-entropy)</li>\n<li><strong>Ensembling</strong> (not surprise there, either)<ul>\n<li><strong>Power averaging</strong> was behind my best selected submission. This looked somewhat promising on CV and not that bad on the public LB (while I selected a weighted arithmetic mean based on LB, which did worse - in part, I also just wanted to select a second submission that was not too similar to the other one). </li>\n<li>I got some even better private LB numbers with <strong>rank averaging</strong> (in silver territory, but I did not select it), but admittedly my very best private LB score was a simple weighted arithmetic mean of two of the models. See below for what did not work so well (i.e. selection the best submission)…</li>\n<li>Why did I think power and rank averaging would work? I kind of figured that accuracy is in some sense a rank based metric (you only care about the first rank of course) and those methods had just been helpful in the <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">Rainforest Connection Species Audio Detection competition</a> that used mean average precision (a very rank based metric).</li></ul></li>\n<li>Using <strong>diverse models</strong> in my ensemble, even if it does hurt on the public LB (I hedged my bets a bit and tried to bet on a half-way house that probably relied on the public LB to much, which cost me a bit in ranks).<ul>\n<li>The idea was to have diverse image sizes as inputs, somewhat diverse models and diverse augmentations during training (and TTA).</li>\n<li>Nevertheless, it seems the two EfficientNets were different enough that having both helped. I probably would not have added the EfficientNet-B4-NS, if it had not been for the public LB score of <a href=\"https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903\" target=\"_blank\">this notebook</a>. That notebook also made me stop trying other ResNet-types and just go with ResNeXt-50 (32x4d) (who knows whether that was the right decision).</li>\n<li>My choice of architectures was perhaps a bit limited I mostly picked these architectures based on what I could get as pre-trained models via <code>timm</code> and perhaps payed too much attention to what others did.  </li></ul></li>\n<li><strong>Gradient accumulation</strong> to use bigger batch sizes (72) with EfficientNet-B5 (necessary when using a GPU on Kaggle or my own 1080-Ti). It appeared to help a tiny bit in CV to use a batch size &gt;32 despite the whole \"<a href=\"https://twitter.com/ylecun/status/989610208497360896?lang=en\" target=\"_blank\">Friends dont let friends use minibatches larger than 32</a>\" <a href=\"https://arxiv.org/abs/1804.07612\" target=\"_blank\">thing</a>.</li>\n<li><strong>flat learning rate before cosine annealing learning rate schedule</strong> that I learnt about from <a href=\"https://www.kaggle.com/abhishek/tez-faster-and-easier-training-for-leaf-detection\" target=\"_blank\">Abishek Thakur's notebook</a></li>\n<li>Weight decay - whenever I tried it, it helped.</li>\n<li><code>fastai</code> + <code>timm</code> as e.g. described in this <a href=\"https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai\" target=\"_blank\">great notebook</a> by <a href=\"https://www.kaggle.com/muellerzr\" target=\"_blank\">@muellerzr</a> (thanks, again, for that nice starter kit) and PyTorch. You only appreciate how smoothly that works with CPU and GPU when you try to use the TPU with PyTorch, which kind of worked, but was a bit painful.</li>\n</ul>\n<p><strong>Things that did not work for me (mostly my fault, some of these clearly worked out better for others):</strong></p>\n<ul>\n<li>Selecting the best submissions - I did 27 submission before the deadline and I selected in terms of the private LB the 17th and 20th best of these.<ul>\n<li>I suspect I was not sufficiently disciplined with running everything I tried in the last few days through a proper CV (yes, yes, I know, that's a terrible sin…) and was too influenced by the public LB.</li>\n<li>I wonder whether those that did best that were super disciplined and perhaps even used repeated k-fold?! </li>\n<li>To some extent the ordering of my various submissions may have had a strong element of luck (accuracy is such a low information content metric and I did not manage to) and dropping 97 places is not that bad, but I aspire to do better next time.</li></ul></li>\n<li>Getting value out of the 2019 competition data, but I probably just invested insufficient time. It just always made my CV worse even after de-duplication, so I gave up.</li>\n<li>Stochastic weight averaging: it kept reliably improving my loss function in CV, but not accuracy (it seemed I got the best accuracy when very slightly overfitting in terms of (label-smoothing) cross-entropy - which I understand is sort of expected). I wonder whether I should nevertheless put a model with SWA in my ensemble.</li>\n<li>Really large models: I initially thought that the biggest EfficientNet I could train with TPU would be best, then I scaled my \"ambitions\" back to the largest one I could comfortably cross-validate (B5) some experiments for… </li>\n<li>Doing two competitions that ended within 24 hours of each other. Yes, I got two medals, but I am left feeling I would have liked to rather do one of them more deeply and do a better job / contribute more to the one where I teamed up.</li>\n</ul>",
      "rawMarkdown": "So, after the latest updates, I moved up a few places on the LB and slipped into the bronze medals. Within two days I've gone from zero medals to my second bronze medal. 😃 On this competition, I worked on my own, while on the  Rainforest Connection Species Audio Detection competition](https://www.kaggle.com/c/rfcx-species-audio-detection) I had the pleasure to work with a great team. So, I thought I'd write up some of the more interesting things I tried on this competition.\n\nMy best selected solution was a weighted power averaging ensemble (i.e. squaring the predictions of each model and then averaging, followed by picking the class with the highest score) of\n* **EfficientNet-B5** (i.e. not noisy student, somehow my noisy student version seemed worse on CV - I failed to figure out why) on image size 456 by 456 with TTA\n* **EfficientNet-B4-noisy student** on image size 512 by 512 with TTA\n* **ResNeXt-50 (32x4d)** with image size 512 by 512\n* **Vision Transformer** (`vit_base_patch16_224`) image size 224 by 224\n\nI did two different test-time augmentation approaches for the two EfficientNets and no TTA for the others (yes, yes, I clearly have an EfficientNet bias).\n\n**Things that worked**:\n* **label smoothing cross-entropy** with epsilon of 0.1 to 0.2 or so as a loss function (not surprising, given that it was originally proposed for noisy labels)\n* **focal loss** (not in my selected solutions, but in CV it was very similar to label smoothing cross-entropy)\n* **Ensembling** (not surprise there, either)\n * **Power averaging** was behind my best selected submission. This looked somewhat promising on CV and not that bad on the public LB (while I selected a weighted arithmetic mean based on LB, which did worse - in part, I also just wanted to select a second submission that was not too similar to the other one). \n * I got some even better private LB numbers with **rank averaging** (in silver territory, but I did not select it), but admittedly my very best private LB score was a simple weighted arithmetic mean of two of the models. See below for what did not work so well (i.e. selection the best submission)...\n * Why did I think power and rank averaging would work? I kind of figured that accuracy is in some sense a rank based metric (you only care about the first rank of course) and those methods had just been helpful in the [Rainforest Connection Species Audio Detection competition](https://www.kaggle.com/c/rfcx-species-audio-detection) that used mean average precision (a very rank based metric).\n* Using **diverse models** in my ensemble, even if it does hurt on the public LB (I hedged my bets a bit and tried to bet on a half-way house that probably relied on the public LB to much, which cost me a bit in ranks).\n * The idea was to have diverse image sizes as inputs, somewhat diverse models and diverse augmentations during training (and TTA).\n * Nevertheless, it seems the two EfficientNets were different enough that having both helped. I probably would not have added the EfficientNet-B4-NS, if it had not been for the public LB score of [this notebook](https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903). That notebook also made me stop trying other ResNet-types and just go with ResNeXt-50 (32x4d) (who knows whether that was the right decision).\n * My choice of architectures was perhaps a bit limited I mostly picked these architectures based on what I could get as pre-trained models via `timm` and perhaps payed too much attention to what others did.  \n* **Gradient accumulation** to use bigger batch sizes (72) with EfficientNet-B5 (necessary when using a GPU on Kaggle or my own 1080-Ti). It appeared to help a tiny bit in CV to use a batch size >32 despite the whole \"[Friends dont let friends use minibatches larger than 32](https://twitter.com/ylecun/status/989610208497360896?lang=en)\" [thing](https://arxiv.org/abs/1804.07612).\n* **flat learning rate before cosine annealing learning rate schedule** that I learnt about from [Abishek Thakur's notebook](https://www.kaggle.com/abhishek/tez-faster-and-easier-training-for-leaf-detection)\n* Weight decay - whenever I tried it, it helped.\n* `fastai` + `timm` as e.g. described in this [great notebook](https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai) by @muellerzr (thanks, again, for that nice starter kit) and PyTorch. You only appreciate how smoothly that works with CPU and GPU when you try to use the TPU with PyTorch, which kind of worked, but was a bit painful.\n\n**Things that did not work for me (mostly my fault, some of these clearly worked out better for others):**\n* Selecting the best submissions - I did 27 submission before the deadline and I selected in terms of the private LB the 17th and 20th best of these.\n * I suspect I was not sufficiently disciplined with running everything I tried in the last few days through a proper CV (yes, yes, I know, that's a terrible sin...) and was too influenced by the public LB.\n * I wonder whether those that did best that were super disciplined and perhaps even used repeated k-fold?! \n * To some extent the ordering of my various submissions may have had a strong element of luck (accuracy is such a low information content metric and I did not manage to) and dropping 97 places is not that bad, but I aspire to do better next time.\n* Getting value out of the 2019 competition data, but I probably just invested insufficient time. It just always made my CV worse even after de-duplication, so I gave up.\n* Stochastic weight averaging: it kept reliably improving my loss function in CV, but not accuracy (it seemed I got the best accuracy when very slightly overfitting in terms of (label-smoothing) cross-entropy - which I understand is sort of expected). I wonder whether I should nevertheless put a model with SWA in my ensemble.\n* Really large models: I initially thought that the biggest EfficientNet I could train with TPU would be best, then I scaled my \"ambitions\" back to the largest one I could comfortably cross-validate (B5) some experiments for... \n* Doing two competitions that ended within 24 hours of each other. Yes, I got two medals, but I am left feeling I would have liked to rather do one of them more deeply and do a better job / contribute more to the one where I teamed up.",
      "votes": null
    },
    {
      "id": "1211081",
      "postDate": "02/19/2021 23:49:22",
      "content": "<p>Congratulations on being a competitions expert. I guess we followed the same path :)</p>",
      "rawMarkdown": "Congratulations on being a competitions expert. I guess we followed the same path :)",
      "votes": null
    },
    {
      "id": "1211092",
      "postDate": "02/20/2021 00:05:50",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>! Thanks for sharing your experiments. I always enjoy reading your contributions to the discussion forum in every competition.</p>",
      "rawMarkdown": "Congratulations @bjoernholzhauer! Thanks for sharing your experiments. I always enjoy reading your contributions to the discussion forum in every competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1211081,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/19/2021 23:49:22",
      "content": "<p>Congratulations on being a competitions expert. I guess we followed the same path :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211092,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "02/20/2021 00:05:50",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>! Thanks for sharing your experiments. I always enjoy reading your contributions to the discussion forum in every competition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1211059": "So, after the latest updates, I moved up a few places on the LB and slipped into the bronze medals. Within two days I've gone from zero medals to my second bronze medal. 😃 On this competition, I worked on my own, while on the  Rainforest Connection Species Audio Detection competition](https://www.kaggle.com/c/rfcx-species-audio-detection) I had the pleasure to work with a great team. So, I thought I'd write up some of the more interesting things I tried on this competition.\n\nMy best selected solution was a weighted power averaging ensemble (i.e. squaring the predictions of each model and then averaging, followed by picking the class with the highest score) of\n* **EfficientNet-B5** (i.e. not noisy student, somehow my noisy student version seemed worse on CV - I failed to figure out why) on image size 456 by 456 with TTA\n* **EfficientNet-B4-noisy student** on image size 512 by 512 with TTA\n* **ResNeXt-50 (32x4d)** with image size 512 by 512\n* **Vision Transformer** (`vit_base_patch16_224`) image size 224 by 224\n\nI did two different test-time augmentation approaches for the two EfficientNets and no TTA for the others (yes, yes, I clearly have an EfficientNet bias).\n\n**Things that worked**:\n* **label smoothing cross-entropy** with epsilon of 0.1 to 0.2 or so as a loss function (not surprising, given that it was originally proposed for noisy labels)\n* **focal loss** (not in my selected solutions, but in CV it was very similar to label smoothing cross-entropy)\n* **Ensembling** (not surprise there, either)\n * **Power averaging** was behind my best selected submission. This looked somewhat promising on CV and not that bad on the public LB (while I selected a weighted arithmetic mean based on LB, which did worse - in part, I also just wanted to select a second submission that was not too similar to the other one). \n * I got some even better private LB numbers with **rank averaging** (in silver territory, but I did not select it), but admittedly my very best private LB score was a simple weighted arithmetic mean of two of the models. See below for what did not work so well (i.e. selection the best submission)...\n * Why did I think power and rank averaging would work? I kind of figured that accuracy is in some sense a rank based metric (you only care about the first rank of course) and those methods had just been helpful in the [Rainforest Connection Species Audio Detection competition](https://www.kaggle.com/c/rfcx-species-audio-detection) that used mean average precision (a very rank based metric).\n* Using **diverse models** in my ensemble, even if it does hurt on the public LB (I hedged my bets a bit and tried to bet on a half-way house that probably relied on the public LB to much, which cost me a bit in ranks).\n * The idea was to have diverse image sizes as inputs, somewhat diverse models and diverse augmentations during training (and TTA).\n * Nevertheless, it seems the two EfficientNets were different enough that having both helped. I probably would not have added the EfficientNet-B4-NS, if it had not been for the public LB score of [this notebook](https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903). That notebook also made me stop trying other ResNet-types and just go with ResNeXt-50 (32x4d) (who knows whether that was the right decision).\n * My choice of architectures was perhaps a bit limited I mostly picked these architectures based on what I could get as pre-trained models via `timm` and perhaps payed too much attention to what others did.  \n* **Gradient accumulation** to use bigger batch sizes (72) with EfficientNet-B5 (necessary when using a GPU on Kaggle or my own 1080-Ti). It appeared to help a tiny bit in CV to use a batch size >32 despite the whole \"[Friends dont let friends use minibatches larger than 32](https://twitter.com/ylecun/status/989610208497360896?lang=en)\" [thing](https://arxiv.org/abs/1804.07612).\n* **flat learning rate before cosine annealing learning rate schedule** that I learnt about from [Abishek Thakur's notebook](https://www.kaggle.com/abhishek/tez-faster-and-easier-training-for-leaf-detection)\n* Weight decay - whenever I tried it, it helped.\n* `fastai` + `timm` as e.g. described in this [great notebook](https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai) by @muellerzr (thanks, again, for that nice starter kit) and PyTorch. You only appreciate how smoothly that works with CPU and GPU when you try to use the TPU with PyTorch, which kind of worked, but was a bit painful.\n\n**Things that did not work for me (mostly my fault, some of these clearly worked out better for others):**\n* Selecting the best submissions - I did 27 submission before the deadline and I selected in terms of the private LB the 17th and 20th best of these.\n * I suspect I was not sufficiently disciplined with running everything I tried in the last few days through a proper CV (yes, yes, I know, that's a terrible sin...) and was too influenced by the public LB.\n * I wonder whether those that did best that were super disciplined and perhaps even used repeated k-fold?! \n * To some extent the ordering of my various submissions may have had a strong element of luck (accuracy is such a low information content metric and I did not manage to) and dropping 97 places is not that bad, but I aspire to do better next time.\n* Getting value out of the 2019 competition data, but I probably just invested insufficient time. It just always made my CV worse even after de-duplication, so I gave up.\n* Stochastic weight averaging: it kept reliably improving my loss function in CV, but not accuracy (it seemed I got the best accuracy when very slightly overfitting in terms of (label-smoothing) cross-entropy - which I understand is sort of expected). I wonder whether I should nevertheless put a model with SWA in my ensemble.\n* Really large models: I initially thought that the biggest EfficientNet I could train with TPU would be best, then I scaled my \"ambitions\" back to the largest one I could comfortably cross-validate (B5) some experiments for... \n* Doing two competitions that ended within 24 hours of each other. Yes, I got two medals, but I am left feeling I would have liked to rather do one of them more deeply and do a better job / contribute more to the one where I teamed up.",
    "1211081": "Congratulations on being a competitions expert. I guess we followed the same path :)",
    "1211092": "Congratulations @bjoernholzhauer! Thanks for sharing your experiments. I always enjoy reading your contributions to the discussion forum in every competition."
  },
  "source": "meta"
}