{
  "id": 177924,
  "title": "40th Place Summary - In Chris and CV we trust",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/billo-40th-place-summary-in-chris-and-cv-we-trust",
  "author_name": "",
  "post_date": "2020-08-27T22:03:04.913Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi Kaggle fam,</p>\n<p>Apologies this was a little delayed – was quite surprised at the result and took some time to retrace my own steps, but thought I would share a few findings that might add to what’s been shared already. </p>\n<p>Firstly massive shoutout to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for his triple-stratified TFrecords which saved a heap of time, and also the interesting techniques he kindly shared which I’ve still yet to try out and fully understand – I like many other noivces look forward to learning more from you in future competitions!</p>\n<p>Like most teams, I realised early on that training different nets on images of different sizes generated varying results in terms CV across validation folds, so I thought ensembles that maximise local CV might be the way to go. My submissions were simple weighted averages of a bunch of EfficientNets of sizes (B4-B7), mostly initialised with noisy-student weights, trained on image sizes 384, 512 and 768 with TTA (15) across 5 validation folds. For augmentation, I believe the original author of the techniques I used is <a href=\"https://www.kaggle.com/agentauers\" target=\"_blank\">@agentauers</a>, so a huge thank you for sharing them. Previous competition data was also used for increasing the diversity of the ensemble.</p>\n<p>Something I found interesting was that with my setup, larger nets + larger images seemed to have generally performed a little better than smaller nets + smaller images (both CV &amp; private LB). Even when ensembling, smaller nets did not help with CV much at all and hence were largely not part of my final submissions. I’m curious to hear if that was the case for anyone else (sorry if it’s been talked about already).</p>\n<p>As for metadata, I had a hunch it might be useful so had been including it in some of my submissions – following Chris’s advice of a roughly 90/10 (image/meta) split. Shoutout to <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> for his notebook, which worked very well for ensembling. I’d be interested to hear more from others who used metadata and how everyone incorporated it into their submissions.</p>\n<p>Thank you to the organisers and Kaggle fam for this wonderful competition. Here’s hoping for many more to come.</p>\n<p>Subs:</p>\n<table>\n<thead>\n<tr>\n<th>Sub</th>\n<th>No. of Effnets</th>\n<th>External Data</th>\n<th>Metadata</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>18</td>\n<td>Y</td>\n<td>N</td>\n<td>0.9521</td>\n<td>0.9408</td>\n</tr>\n<tr>\n<td>2</td>\n<td>18</td>\n<td>Y</td>\n<td>Y</td>\n<td>0.9535</td>\n<td>0.9434</td>\n</tr>\n<tr>\n<td>3</td>\n<td>45</td>\n<td>Y</td>\n<td>Y</td>\n<td>0.9447</td>\n<td>0.9344</td>\n</tr>\n<tr>\n<td>Unused smaller ensemble</td>\n<td>6</td>\n<td>Y</td>\n<td>Y</td>\n<td>0.9544</td>\n<td>0.9405</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "988212",
      "postDate": "08/27/2020 22:02:07",
      "content": "<p>Hi Kaggle fam,</p>\n<p>Apologies this was a little delayed – was quite surprised at the result and took some time to retrace my own steps, but thought I would share a few findings that might add to what’s been shared already. </p>\n<p>Firstly massive shoutout to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for his triple-stratified TFrecords which saved a heap of time, and also the interesting techniques he kindly shared which I’ve still yet to try out and fully understand – I like many other noivces look forward to learning more from you in future competitions!</p>\n<p>Like most teams, I realised early on that training different nets on images of different sizes generated varying results in terms CV across validation folds, so I thought ensembles that maximise local CV might be the way to go. My submissions were simple weighted averages of a bunch of EfficientNets of sizes (B4-B7), mostly initialised with noisy-student weights, trained on image sizes 384, 512 and 768 with TTA (15) across 5 validation folds. For augmentation, I believe the original author of the techniques I used is <a href=\"https://www.kaggle.com/agentauers\" target=\"_blank\">@agentauers</a>, so a huge thank you for sharing them. Previous competition data was also used for increasing the diversity of the ensemble.</p>\n<p>Something I found interesting was that with my setup, larger nets + larger images seemed to have generally performed a little better than smaller nets + smaller images (both CV &amp; private LB). Even when ensembling, smaller nets did not help with CV much at all and hence were largely not part of my final submissions. I’m curious to hear if that was the case for anyone else (sorry if it’s been talked about already).</p>\n<p>As for metadata, I had a hunch it might be useful so had been including it in some of my submissions – following Chris’s advice of a roughly 90/10 (image/meta) split. Shoutout to <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> for his notebook, which worked very well for ensembling. I’d be interested to hear more from others who used metadata and how everyone incorporated it into their submissions.</p>\n<p>Thank you to the organisers and Kaggle fam for this wonderful competition. Here’s hoping for many more to come.</p>\n<p>Subs:</p>\n<table>\n<thead>\n<tr>\n<th>Sub</th>\n<th>No. of Effnets</th>\n<th>External Data</th>\n<th>Metadata</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>18</td>\n<td>Y</td>\n<td>N</td>\n<td>0.9521</td>\n<td>0.9408</td>\n</tr>\n<tr>\n<td>2</td>\n<td>18</td>\n<td>Y</td>\n<td>Y</td>\n<td>0.9535</td>\n<td>0.9434</td>\n</tr>\n<tr>\n<td>3</td>\n<td>45</td>\n<td>Y</td>\n<td>Y</td>\n<td>0.9447</td>\n<td>0.9344</td>\n</tr>\n<tr>\n<td>Unused smaller ensemble</td>\n<td>6</td>\n<td>Y</td>\n<td>Y</td>\n<td>0.9544</td>\n<td>0.9405</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Hi Kaggle fam,\n\nApologies this was a little delayed – was quite surprised at the result and took some time to retrace my own steps, but thought I would share a few findings that might add to what’s been shared already. \n\nFirstly massive shoutout to @cdeotte for his triple-stratified TFrecords which saved a heap of time, and also the interesting techniques he kindly shared which I’ve still yet to try out and fully understand – I like many other noivces look forward to learning more from you in future competitions!\n\nLike most teams, I realised early on that training different nets on images of different sizes generated varying results in terms CV across validation folds, so I thought ensembles that maximise local CV might be the way to go. My submissions were simple weighted averages of a bunch of EfficientNets of sizes (B4-B7), mostly initialised with noisy-student weights, trained on image sizes 384, 512 and 768 with TTA (15) across 5 validation folds. For augmentation, I believe the original author of the techniques I used is @agentauers, so a huge thank you for sharing them. Previous competition data was also used for increasing the diversity of the ensemble.\n\nSomething I found interesting was that with my setup, larger nets + larger images seemed to have generally performed a little better than smaller nets + smaller images (both CV & private LB). Even when ensembling, smaller nets did not help with CV much at all and hence were largely not part of my final submissions. I’m curious to hear if that was the case for anyone else (sorry if it’s been talked about already).\n\nAs for metadata, I had a hunch it might be useful so had been including it in some of my submissions – following Chris’s advice of a roughly 90/10 (image/meta) split. Shoutout to @titericz for his notebook, which worked very well for ensembling. I’d be interested to hear more from others who used metadata and how everyone incorporated it into their submissions.\n\nThank you to the organisers and Kaggle fam for this wonderful competition. Here’s hoping for many more to come.\n\nSubs:\n| Sub | No. of Effnets | External Data | Metadata | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- |\n| 1 | 18 | Y | N | 0.9521 | 0.9408 |\n| 2 | 18 | Y | Y | 0.9535 | 0.9434 |\n| 3 | 45 | Y | Y | 0.9447 | 0.9344 |\n| Unused smaller ensemble | 6 | Y | Y | 0.9544 | 0.9405 |",
      "votes": null
    },
    {
      "id": "988661",
      "postDate": "08/28/2020 07:27:08",
      "content": "<p>Congratulations and thanks for sharing your approach.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your approach.",
      "votes": null
    },
    {
      "id": "989315",
      "postDate": "08/28/2020 17:52:40",
      "content": "<p>Congratulations Billo on achieving solo silver. </p>\n<p>I also built models of all image sizes (128x128 thru 1024x1024) and all EfficientNet sizes (B0 thru B7). Similar to you only image sizes (384 thru 1024) and EfficientNet (B3 thru B7) increased my local CV. And 256x256 helped slightly.</p>",
      "rawMarkdown": "Congratulations Billo on achieving solo silver. \n\nI also built models of all image sizes (128x128 thru 1024x1024) and all EfficientNet sizes (B0 thru B7). Similar to you only image sizes (384 thru 1024) and EfficientNet (B3 thru B7) increased my local CV. And 256x256 helped slightly.",
      "votes": null
    },
    {
      "id": "990798",
      "postDate": "08/29/2020 21:20:44",
      "content": "<p>Cheers Chris! Great to hear that wasn't a one-off and you also had something similar going on. </p>\n<p>Thanks again for your generosity in sharing - you're an absolute legend. Please keep it up! 🙏</p>",
      "rawMarkdown": "Cheers Chris! Great to hear that wasn't a one-off and you also had something similar going on. \n\nThanks again for your generosity in sharing - you're an absolute legend. Please keep it up! 🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 988661,
      "author_name": "karrak3256",
      "author_url": "",
      "post_date": "08/28/2020 07:27:08",
      "content": "<p>Congratulations and thanks for sharing your approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 989315,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/28/2020 17:52:40",
      "content": "<p>Congratulations Billo on achieving solo silver. </p>\n<p>I also built models of all image sizes (128x128 thru 1024x1024) and all EfficientNet sizes (B0 thru B7). Similar to you only image sizes (384 thru 1024) and EfficientNet (B3 thru B7) increased my local CV. And 256x256 helped slightly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 990798,
          "author_name": "billo97",
          "author_url": "",
          "post_date": "08/29/2020 21:20:44",
          "content": "<p>Cheers Chris! Great to hear that wasn't a one-off and you also had something similar going on. </p>\n<p>Thanks again for your generosity in sharing - you're an absolute legend. Please keep it up! 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "988212": "Hi Kaggle fam,\n\nApologies this was a little delayed – was quite surprised at the result and took some time to retrace my own steps, but thought I would share a few findings that might add to what’s been shared already. \n\nFirstly massive shoutout to @cdeotte for his triple-stratified TFrecords which saved a heap of time, and also the interesting techniques he kindly shared which I’ve still yet to try out and fully understand – I like many other noivces look forward to learning more from you in future competitions!\n\nLike most teams, I realised early on that training different nets on images of different sizes generated varying results in terms CV across validation folds, so I thought ensembles that maximise local CV might be the way to go. My submissions were simple weighted averages of a bunch of EfficientNets of sizes (B4-B7), mostly initialised with noisy-student weights, trained on image sizes 384, 512 and 768 with TTA (15) across 5 validation folds. For augmentation, I believe the original author of the techniques I used is @agentauers, so a huge thank you for sharing them. Previous competition data was also used for increasing the diversity of the ensemble.\n\nSomething I found interesting was that with my setup, larger nets + larger images seemed to have generally performed a little better than smaller nets + smaller images (both CV & private LB). Even when ensembling, smaller nets did not help with CV much at all and hence were largely not part of my final submissions. I’m curious to hear if that was the case for anyone else (sorry if it’s been talked about already).\n\nAs for metadata, I had a hunch it might be useful so had been including it in some of my submissions – following Chris’s advice of a roughly 90/10 (image/meta) split. Shoutout to @titericz for his notebook, which worked very well for ensembling. I’d be interested to hear more from others who used metadata and how everyone incorporated it into their submissions.\n\nThank you to the organisers and Kaggle fam for this wonderful competition. Here’s hoping for many more to come.\n\nSubs:\n| Sub | No. of Effnets | External Data | Metadata | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- |\n| 1 | 18 | Y | N | 0.9521 | 0.9408 |\n| 2 | 18 | Y | Y | 0.9535 | 0.9434 |\n| 3 | 45 | Y | Y | 0.9447 | 0.9344 |\n| Unused smaller ensemble | 6 | Y | Y | 0.9544 | 0.9405 |",
    "988661": "Congratulations and thanks for sharing your approach.",
    "989315": "Congratulations Billo on achieving solo silver. \n\nI also built models of all image sizes (128x128 thru 1024x1024) and all EfficientNet sizes (B0 thru B7). Similar to you only image sizes (384 thru 1024) and EfficientNet (B3 thru B7) increased my local CV. And 256x256 helped slightly.",
    "990798": "Cheers Chris! Great to hear that wasn't a one-off and you also had something similar going on. \n\nThanks again for your generosity in sharing - you're an absolute legend. Please keep it up! 🙏"
  },
  "source": "meta"
}