{
  "id": 118250,
  "title": "34th Place Solution + Code",
  "url": "/competitions/understanding_cloud_organization/writeups/karl-hornlund-34th-place-solution-code",
  "author_name": "",
  "post_date": "2019-11-20T10:43:04.565183400Z",
  "votes": 12,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>Congrats to the winners! </p>\n\n<p>Code for my solution <a href=\"https://github.com/khornlund/understanding-cloud-organization\">here</a>.</p>\n\n<p>Explanation copied below.</p>\n\n<h1>Summary</h1>\n\n<h2>Results</h2>\n\n<p>| Rank | Score | Percentile |\n| --- | --- | --- |\n| 34 | 0.66385 | Top 2.2% |</p>\n\n<h2>Strategy</h2>\n\n<p>Originally I had an idea early on very similar to <a href=\"https://arxiv.org/pdf/1911.04252.pdf\">this</a> recent paper. I was going to train a model on the ground truthed data, and then iteratively create pseudo labels for unlabelled data and train on that. I figured this was a good opportunity for such a strategy because there was very little training data (~5000 images), so there was a lot to be gained by generating more training samples. And, because this was not a synchronous kernel competition, I'd be able to create as large an ensemble as I like.</p>\n\n<p>Then I realised how noisy the image labels were, and wasn't so sure that pseudo labels would work very well. In particular, I noticed that the validation scores of my models was super noisy - using the same configuration with a different random seed resulted in serious metric differences. I figured I would give up on trying to fine tune individual models and instead focus on engineering a system that would allow me to train and ensemble <em>lots</em> of models.</p>\n\n<p>I developed functionality to allow me to automate the configuration, training, and inference of models.</p>\n\n<p>I trained an ensemble of ~120 models, using a variety of encoder/decoder combinations. I first averaged them together by their encoder/decoder combinations (eg. all the efficientnet-b2 FPN get averaged together). Then I averaged these mini-ensembles together using a weighted average.</p>\n\n<p>With about a week of the competition to go, I saw the Noisy Student paper. I was getting decent results on the LB and figured I'd give pseudo labelling a go. I downloaded ~4200 images using the same resolution and locations as the official data, generated pseudo labels for them, and trained a new ensemble of ~50 models.</p>\n\n<p>I only finished training the pseudo labelled models in time to make a few submissions on the final day, and managed to get up to 0.67739 (9th place) on the public LB - but that actually only scored 0.66331 (~45th) on the private LB. My other selected submission was a weighted average of my past 25 submissions, which scored 0.67574 on the public LB and 0.66385 (34th) on the private LB.</p>\n\n<p>I had a few unselected submissions that scored 0.666+ (~18th), the best of which funnily enough came from a mini-ensemble of only efficientnet-b2-Unet models.</p>\n\n<h2>Reflection</h2>\n\n<p>Looking back I realise I made a pretty big mistake not capturing the appropriate metrics for thorough local CV. I was only recording dice coefficient using a threshold of 0.5, and so I wasn't well informed to pick a threshold for my submissions.</p>\n\n<p>Also, while the models were each trained on a random 80% of the data, and evaluated on the remaining 20%, this was only done at a per-model level. I didn't keep a hold-out set to validate the ensembles against. Because we only had ~5000 training samples, I got a bit greedy with training data here.</p>\n\n<p>I was hoping that by keeping logs of all my experiments, after a while I'd be able to identify which randomly generated configurations (eg. learning rate) worked better than others. This didn't turn out to be the case! I should have spent more time fine tuning each model, as the law of diminishing returns was coming into effect as the size of my ensemble grew.</p>\n\n<h1>Details</h1>\n\n<h2>Ensemble Pipeline</h2>\n\n<p>See <code>uco.ensemble.py</code> for implementation.</p>\n\n<p>Each training experiment is configured using a YAML file which gets loaded into a dictionary. I set up a class to randomise these parameters, so I could leave it to run while at work/sleep and it would cycle through different architectures, loss functions, and other parameters.</p>\n\n<p>After each training epoch the model would be evaluated on a 20% validation set. The mean dice score was tracked throughout training, and when the training completed (either after a set number of epochs or early stopping) only the best scoring checkpoint would be saved. I set a cutoff mean dice score, and threw away models that scored under that.</p>\n\n<p>The saved checkpoint would be loaded, and run inference on the test data. I saved out the <em>raw</em> (sigmoid) predictions of each model to HDF5. I scaled by 250 and rounded to integers so I could save as <code>uint8</code> to save disk space.</p>\n\n<p>These raw predictions would be grouped by (encoder, decoder) pair, and averaged together weighted by mean dice scores. Then the groups would be averaged together, with parameterised weights.</p>\n\n<p>By saving out the results at each stage to HDF5 (raw predictions, group averages, and total averages), I could re-run any part of the pipeline with ease.</p>\n\n<p>I did the above for both segmentation and classification models. The details below are just for the segmentation models.</p>\n\n<h2>Models</h2>\n\n<p>I used <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">segmentation_models.pytorch</a>\n(SMP) for segmentation, and used <a href=\"https://github.com/rwightman/pytorch-image-models\">pytorch-image-models</a> (TIIM) for classification.</p>\n\n<p><strong>Encoders</strong></p>\n\n<ul>\n<li>efficientnet B0, B2, B5, B6</li>\n<li>resnext 101_32x8d</li>\n<li>se_resnext 101_32x8d</li>\n<li>inceptionresnet v2, v4</li>\n<li>dpn 131</li>\n<li>densenet 161</li>\n</ul>\n\n<p><strong>Decoders</strong></p>\n\n<ul>\n<li>FPN</li>\n<li>Unet</li>\n</ul>\n\n<p>I had terrible results with LinkNet and PSPNet.</p>\n\n<h2>Training</h2>\n\n<p><strong>GPU</strong>\nRTX 2080Ti.</p>\n\n<p><strong>Loss</strong>\nI used BCE + Dice with BCE weight ~U(0.65, 0.75) and dice weight 1 - BCE.</p>\n\n<p>I used BCE + Lovasz with BCE weight ~U(0.83, 0.92) and lovasz 1 - BCE.</p>\n\n<p><strong>Learning Rate</strong>\nEncoder ~U(5e-5, 9e-5)\nDecoder ~U(3e-3, 5e-3)</p>\n\n<p><strong>Optimizer</strong>\nRAdam / <a href=\"https://github.com/catalyst-team/catalyst/blob/master/catalyst/contrib/optimizers/qhadamw.py\">QHAdamW</a></p>\n\n<p><strong>Augmentation</strong>\nCompositions are in <code>data_loader.augmentation.py</code>.</p>\n\n<p>I made one custom augmentation - I modified Cutout to apply to masks. I wasn't sure if this would actually be better than only applying Cutout to the image - because the ground truth bounding boxes were large and covered areas that actually weren't very cloudy. It wasn't obvious from my experiments which worked better - but they both helped, so I just added them both to the available random configuration options for training.</p>\n\n<p><strong>Image Sizes</strong>\nI wanted to use images sizes divisible by 32 so they would work without rounding effects, so I used the following which maintained the original 1400:2100 aspect ratio:</p>\n\n<ul>\n<li>256x384</li>\n<li>320x480</li>\n<li>384x576</li>\n<li>448x672</li>\n</ul>\n\n<p>Most models were trained using 320x480. I didn't notice any improvement using larger image sizes, but I figured it might help the ensemble to use diverse sizes.</p>\n\n<p><strong>Pseudo Labels</strong>\nI used my ensemble trained on the official training data to predict masks for the ~4000 images I downloaded. I then removed any images without masks, and trained on the rest.</p>\n\n<p>In contrast to some of the other people that used pseudo labels, I did not make my thresholds harsher for selecting pseudo labels. My rationale was that since most images included 2+ classes, increasing the thresholds to be 'safe' would likely mean missing the 2nd class in many images - leading to lots of false negative labels in my pseudo labels.</p>\n\n<p>I used a <a href=\"https://github.com/khornlund/pytorch-balanced-sampler\">balanced sampler</a> to include 4 pseudo labelled samples per batch (typically batch sizes were 10-16).</p>\n\n<h2>Post-Processing</h2>\n\n<p><strong>TTA</strong>\nI used flips from <a href=\"https://github.com/qubvel/ttach\">TTAch</a></p>\n\n<p><strong>Segmentation Thresholds</strong>\nI experimented with a bunch of different ways to threshold positive predictions, as\nthe dice metric penalises false positives so heavily.</p>\n\n<p>I started out by using the following threshold rule:</p>\n\n<ol>\n<li>Outputs must have N pixels above some <em>top threshold</em>. I started out using N ~ 8000 for each class, and a top threshold of ~0.57.</li>\n<li>For predictions that pass (1), produce a binary mask using <em>bot threshold</em> of ~0.4.</li>\n</ol>\n\n<p>I used the continuous output of the classifier to modulate these thresholds. Ie. if the classifier was high, I would reduce the min size requirement, or the top threshold.</p>\n\n<p>In the end I simply used maximum pixel prediction and no min size.</p>\n\n<p>The distribution of predictions for the different classes is actually pretty interesting:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2116899%2Ff665763ea6d3668f7514f747997ad71d%2Faverage-prediction-distribution.png?generation=1574246988235284&amp;alt=media\" alt=\"\"></p>\n\n<p>Class 1 has very nice bimodal distribution. This suggests it was the easiest to learn.</p>",
  "messages": [
    {
      "id": "677578",
      "postDate": "11/20/2019 10:43:04",
      "content": "<p>Hi all,</p>\n\n<p>Congrats to the winners! </p>\n\n<p>Code for my solution <a href=\"https://github.com/khornlund/understanding-cloud-organization\">here</a>.</p>\n\n<p>Explanation copied below.</p>\n\n<h1>Summary</h1>\n\n<h2>Results</h2>\n\n<p>| Rank | Score | Percentile |\n| --- | --- | --- |\n| 34 | 0.66385 | Top 2.2% |</p>\n\n<h2>Strategy</h2>\n\n<p>Originally I had an idea early on very similar to <a href=\"https://arxiv.org/pdf/1911.04252.pdf\">this</a> recent paper. I was going to train a model on the ground truthed data, and then iteratively create pseudo labels for unlabelled data and train on that. I figured this was a good opportunity for such a strategy because there was very little training data (~5000 images), so there was a lot to be gained by generating more training samples. And, because this was not a synchronous kernel competition, I'd be able to create as large an ensemble as I like.</p>\n\n<p>Then I realised how noisy the image labels were, and wasn't so sure that pseudo labels would work very well. In particular, I noticed that the validation scores of my models was super noisy - using the same configuration with a different random seed resulted in serious metric differences. I figured I would give up on trying to fine tune individual models and instead focus on engineering a system that would allow me to train and ensemble <em>lots</em> of models.</p>\n\n<p>I developed functionality to allow me to automate the configuration, training, and inference of models.</p>\n\n<p>I trained an ensemble of ~120 models, using a variety of encoder/decoder combinations. I first averaged them together by their encoder/decoder combinations (eg. all the efficientnet-b2 FPN get averaged together). Then I averaged these mini-ensembles together using a weighted average.</p>\n\n<p>With about a week of the competition to go, I saw the Noisy Student paper. I was getting decent results on the LB and figured I'd give pseudo labelling a go. I downloaded ~4200 images using the same resolution and locations as the official data, generated pseudo labels for them, and trained a new ensemble of ~50 models.</p>\n\n<p>I only finished training the pseudo labelled models in time to make a few submissions on the final day, and managed to get up to 0.67739 (9th place) on the public LB - but that actually only scored 0.66331 (~45th) on the private LB. My other selected submission was a weighted average of my past 25 submissions, which scored 0.67574 on the public LB and 0.66385 (34th) on the private LB.</p>\n\n<p>I had a few unselected submissions that scored 0.666+ (~18th), the best of which funnily enough came from a mini-ensemble of only efficientnet-b2-Unet models.</p>\n\n<h2>Reflection</h2>\n\n<p>Looking back I realise I made a pretty big mistake not capturing the appropriate metrics for thorough local CV. I was only recording dice coefficient using a threshold of 0.5, and so I wasn't well informed to pick a threshold for my submissions.</p>\n\n<p>Also, while the models were each trained on a random 80% of the data, and evaluated on the remaining 20%, this was only done at a per-model level. I didn't keep a hold-out set to validate the ensembles against. Because we only had ~5000 training samples, I got a bit greedy with training data here.</p>\n\n<p>I was hoping that by keeping logs of all my experiments, after a while I'd be able to identify which randomly generated configurations (eg. learning rate) worked better than others. This didn't turn out to be the case! I should have spent more time fine tuning each model, as the law of diminishing returns was coming into effect as the size of my ensemble grew.</p>\n\n<h1>Details</h1>\n\n<h2>Ensemble Pipeline</h2>\n\n<p>See <code>uco.ensemble.py</code> for implementation.</p>\n\n<p>Each training experiment is configured using a YAML file which gets loaded into a dictionary. I set up a class to randomise these parameters, so I could leave it to run while at work/sleep and it would cycle through different architectures, loss functions, and other parameters.</p>\n\n<p>After each training epoch the model would be evaluated on a 20% validation set. The mean dice score was tracked throughout training, and when the training completed (either after a set number of epochs or early stopping) only the best scoring checkpoint would be saved. I set a cutoff mean dice score, and threw away models that scored under that.</p>\n\n<p>The saved checkpoint would be loaded, and run inference on the test data. I saved out the <em>raw</em> (sigmoid) predictions of each model to HDF5. I scaled by 250 and rounded to integers so I could save as <code>uint8</code> to save disk space.</p>\n\n<p>These raw predictions would be grouped by (encoder, decoder) pair, and averaged together weighted by mean dice scores. Then the groups would be averaged together, with parameterised weights.</p>\n\n<p>By saving out the results at each stage to HDF5 (raw predictions, group averages, and total averages), I could re-run any part of the pipeline with ease.</p>\n\n<p>I did the above for both segmentation and classification models. The details below are just for the segmentation models.</p>\n\n<h2>Models</h2>\n\n<p>I used <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">segmentation_models.pytorch</a>\n(SMP) for segmentation, and used <a href=\"https://github.com/rwightman/pytorch-image-models\">pytorch-image-models</a> (TIIM) for classification.</p>\n\n<p><strong>Encoders</strong></p>\n\n<ul>\n<li>efficientnet B0, B2, B5, B6</li>\n<li>resnext 101_32x8d</li>\n<li>se_resnext 101_32x8d</li>\n<li>inceptionresnet v2, v4</li>\n<li>dpn 131</li>\n<li>densenet 161</li>\n</ul>\n\n<p><strong>Decoders</strong></p>\n\n<ul>\n<li>FPN</li>\n<li>Unet</li>\n</ul>\n\n<p>I had terrible results with LinkNet and PSPNet.</p>\n\n<h2>Training</h2>\n\n<p><strong>GPU</strong>\nRTX 2080Ti.</p>\n\n<p><strong>Loss</strong>\nI used BCE + Dice with BCE weight ~U(0.65, 0.75) and dice weight 1 - BCE.</p>\n\n<p>I used BCE + Lovasz with BCE weight ~U(0.83, 0.92) and lovasz 1 - BCE.</p>\n\n<p><strong>Learning Rate</strong>\nEncoder ~U(5e-5, 9e-5)\nDecoder ~U(3e-3, 5e-3)</p>\n\n<p><strong>Optimizer</strong>\nRAdam / <a href=\"https://github.com/catalyst-team/catalyst/blob/master/catalyst/contrib/optimizers/qhadamw.py\">QHAdamW</a></p>\n\n<p><strong>Augmentation</strong>\nCompositions are in <code>data_loader.augmentation.py</code>.</p>\n\n<p>I made one custom augmentation - I modified Cutout to apply to masks. I wasn't sure if this would actually be better than only applying Cutout to the image - because the ground truth bounding boxes were large and covered areas that actually weren't very cloudy. It wasn't obvious from my experiments which worked better - but they both helped, so I just added them both to the available random configuration options for training.</p>\n\n<p><strong>Image Sizes</strong>\nI wanted to use images sizes divisible by 32 so they would work without rounding effects, so I used the following which maintained the original 1400:2100 aspect ratio:</p>\n\n<ul>\n<li>256x384</li>\n<li>320x480</li>\n<li>384x576</li>\n<li>448x672</li>\n</ul>\n\n<p>Most models were trained using 320x480. I didn't notice any improvement using larger image sizes, but I figured it might help the ensemble to use diverse sizes.</p>\n\n<p><strong>Pseudo Labels</strong>\nI used my ensemble trained on the official training data to predict masks for the ~4000 images I downloaded. I then removed any images without masks, and trained on the rest.</p>\n\n<p>In contrast to some of the other people that used pseudo labels, I did not make my thresholds harsher for selecting pseudo labels. My rationale was that since most images included 2+ classes, increasing the thresholds to be 'safe' would likely mean missing the 2nd class in many images - leading to lots of false negative labels in my pseudo labels.</p>\n\n<p>I used a <a href=\"https://github.com/khornlund/pytorch-balanced-sampler\">balanced sampler</a> to include 4 pseudo labelled samples per batch (typically batch sizes were 10-16).</p>\n\n<h2>Post-Processing</h2>\n\n<p><strong>TTA</strong>\nI used flips from <a href=\"https://github.com/qubvel/ttach\">TTAch</a></p>\n\n<p><strong>Segmentation Thresholds</strong>\nI experimented with a bunch of different ways to threshold positive predictions, as\nthe dice metric penalises false positives so heavily.</p>\n\n<p>I started out by using the following threshold rule:</p>\n\n<ol>\n<li>Outputs must have N pixels above some <em>top threshold</em>. I started out using N ~ 8000 for each class, and a top threshold of ~0.57.</li>\n<li>For predictions that pass (1), produce a binary mask using <em>bot threshold</em> of ~0.4.</li>\n</ol>\n\n<p>I used the continuous output of the classifier to modulate these thresholds. Ie. if the classifier was high, I would reduce the min size requirement, or the top threshold.</p>\n\n<p>In the end I simply used maximum pixel prediction and no min size.</p>\n\n<p>The distribution of predictions for the different classes is actually pretty interesting:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2116899%2Ff665763ea6d3668f7514f747997ad71d%2Faverage-prediction-distribution.png?generation=1574246988235284&amp;alt=media\" alt=\"\"></p>\n\n<p>Class 1 has very nice bimodal distribution. This suggests it was the easiest to learn.</p>",
      "rawMarkdown": "Hi all,\n\nCongrats to the winners! \n\nCode for my solution [here](https://github.com/khornlund/understanding-cloud-organization).\n\nExplanation copied below.\n\nSummary\n=======\n\nResults\n---------\n\n| Rank | Score | Percentile |\n| --- | --- | --- |\n| 34 | 0.66385 | Top 2.2% |\n\nStrategy\n-----------\nOriginally I had an idea early on very similar to [this](https://arxiv.org/pdf/1911.04252.pdf) recent paper. I was going to train a model on the ground truthed data, and then iteratively create pseudo labels for unlabelled data and train on that. I figured this was a good opportunity for such a strategy because there was very little training data (~5000 images), so there was a lot to be gained by generating more training samples. And, because this was not a synchronous kernel competition, I'd be able to create as large an ensemble as I like.\n\nThen I realised how noisy the image labels were, and wasn't so sure that pseudo labels would work very well. In particular, I noticed that the validation scores of my models was super noisy - using the same configuration with a different random seed resulted in serious metric differences. I figured I would give up on trying to fine tune individual models and instead focus on engineering a system that would allow me to train and ensemble *lots* of models.\n\nI developed functionality to allow me to automate the configuration, training, and inference of models.\n\nI trained an ensemble of ~120 models, using a variety of encoder/decoder combinations. I first averaged them together by their encoder/decoder combinations (eg. all the efficientnet-b2 FPN get averaged together). Then I averaged these mini-ensembles together using a weighted average.\n\nWith about a week of the competition to go, I saw the Noisy Student paper. I was getting decent results on the LB and figured I'd give pseudo labelling a go. I downloaded ~4200 images using the same resolution and locations as the official data, generated pseudo labels for them, and trained a new ensemble of ~50 models.\n\nI only finished training the pseudo labelled models in time to make a few submissions on the final day, and managed to get up to 0.67739 (9th place) on the public LB - but that actually only scored 0.66331 (~45th) on the private LB. My other selected submission was a weighted average of my past 25 submissions, which scored 0.67574 on the public LB and 0.66385 (34th) on the private LB.\n\nI had a few unselected submissions that scored 0.666+ (~18th), the best of which funnily enough came from a mini-ensemble of only efficientnet-b2-Unet models.\n\nReflection\n------------\nLooking back I realise I made a pretty big mistake not capturing the appropriate metrics for thorough local CV. I was only recording dice coefficient using a threshold of 0.5, and so I wasn't well informed to pick a threshold for my submissions.\n\nAlso, while the models were each trained on a random 80% of the data, and evaluated on the remaining 20%, this was only done at a per-model level. I didn't keep a hold-out set to validate the ensembles against. Because we only had ~5000 training samples, I got a bit greedy with training data here.\n\nI was hoping that by keeping logs of all my experiments, after a while I'd be able to identify which randomly generated configurations (eg. learning rate) worked better than others. This didn't turn out to be the case! I should have spent more time fine tuning each model, as the law of diminishing returns was coming into effect as the size of my ensemble grew.\n\nDetails\n=====\n\nEnsemble Pipeline\n----------------------\nSee ``uco.ensemble.py`` for implementation.\n\nEach training experiment is configured using a YAML file which gets loaded into a dictionary. I set up a class to randomise these parameters, so I could leave it to run while at work/sleep and it would cycle through different architectures, loss functions, and other parameters.\n\nAfter each training epoch the model would be evaluated on a 20% validation set. The mean dice score was tracked throughout training, and when the training completed (either after a set number of epochs or early stopping) only the best scoring checkpoint would be saved. I set a cutoff mean dice score, and threw away models that scored under that.\n\nThe saved checkpoint would be loaded, and run inference on the test data. I saved out the *raw* (sigmoid) predictions of each model to HDF5. I scaled by 250 and rounded to integers so I could save as ``uint8`` to save disk space.\n\nThese raw predictions would be grouped by (encoder, decoder) pair, and averaged together weighted by mean dice scores. Then the groups would be averaged together, with parameterised weights.\n\nBy saving out the results at each stage to HDF5 (raw predictions, group averages, and total averages), I could re-run any part of the pipeline with ease.\n\nI did the above for both segmentation and classification models. The details below are just for the segmentation models.\n\nModels\n---------\nI used [segmentation_models.pytorch](https://github.com/qubvel/segmentation_models.pytorch)\n(SMP) for segmentation, and used [pytorch-image-models](https://github.com/rwightman/pytorch-image-models) (TIIM) for classification.\n\n**Encoders**\n\n- efficientnet B0, B2, B5, B6\n- resnext 101_32x8d\n- se_resnext 101_32x8d\n- inceptionresnet v2, v4\n- dpn 131\n- densenet 161\n\n**Decoders**\n\n- FPN\n- Unet\n\nI had terrible results with LinkNet and PSPNet.\n\nTraining\n----------\n\n**GPU**\nRTX 2080Ti.\n\n**Loss**\nI used BCE + Dice with BCE weight ~U(0.65, 0.75) and dice weight 1 - BCE.\n\nI used BCE + Lovasz with BCE weight ~U(0.83, 0.92) and lovasz 1 - BCE.\n\n**Learning Rate**\nEncoder ~U(5e-5, 9e-5)\nDecoder ~U(3e-3, 5e-3)\n\n**Optimizer**\nRAdam / [QHAdamW](https://github.com/catalyst-team/catalyst/blob/master/catalyst/contrib/optimizers/qhadamw.py)\n\n**Augmentation**\nCompositions are in `data_loader.augmentation.py`.\n\nI made one custom augmentation - I modified Cutout to apply to masks. I wasn't sure if this would actually be better than only applying Cutout to the image - because the ground truth bounding boxes were large and covered areas that actually weren't very cloudy. It wasn't obvious from my experiments which worked better - but they both helped, so I just added them both to the available random configuration options for training.\n\n**Image Sizes**\nI wanted to use images sizes divisible by 32 so they would work without rounding effects, so I used the following which maintained the original 1400:2100 aspect ratio:\n\n- 256x384\n- 320x480\n- 384x576\n- 448x672\n\nMost models were trained using 320x480. I didn't notice any improvement using larger image sizes, but I figured it might help the ensemble to use diverse sizes.\n\n**Pseudo Labels**\nI used my ensemble trained on the official training data to predict masks for the ~4000 images I downloaded. I then removed any images without masks, and trained on the rest.\n\nIn contrast to some of the other people that used pseudo labels, I did not make my thresholds harsher for selecting pseudo labels. My rationale was that since most images included 2+ classes, increasing the thresholds to be 'safe' would likely mean missing the 2nd class in many images - leading to lots of false negative labels in my pseudo labels.\n\nI used a [balanced sampler](https://github.com/khornlund/pytorch-balanced-sampler) to include 4 pseudo labelled samples per batch (typically batch sizes were 10-16).\n\nPost-Processing\n-------------------\n\n**TTA**\nI used flips from [TTAch](https://github.com/qubvel/ttach)\n\n**Segmentation Thresholds**\nI experimented with a bunch of different ways to threshold positive predictions, as\nthe dice metric penalises false positives so heavily.\n\nI started out by using the following threshold rule:\n\n1. Outputs must have N pixels above some *top threshold*. I started out using N ~ 8000 for each class, and a top threshold of ~0.57.\n2. For predictions that pass (1), produce a binary mask using *bot threshold* of ~0.4.\n\nI used the continuous output of the classifier to modulate these thresholds. Ie. if the classifier was high, I would reduce the min size requirement, or the top threshold.\n\nIn the end I simply used maximum pixel prediction and no min size.\n\nThe distribution of predictions for the different classes is actually pretty interesting:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2116899%2Ff665763ea6d3668f7514f747997ad71d%2Faverage-prediction-distribution.png?generation=1574246988235284&amp;alt=media)\n\n\nClass 1 has very nice bimodal distribution. This suggests it was the easiest to learn.",
      "votes": null
    },
    {
      "id": "677610",
      "postDate": "11/20/2019 11:42:01",
      "content": "<p>Congrats! And thanks for sharing your solution 😄 </p>",
      "rawMarkdown": "Congrats! And thanks for sharing your solution 😄",
      "votes": null
    },
    {
      "id": "677714",
      "postDate": "11/20/2019 14:18:06",
      "content": "<p>Congratulations\nThanks for Sharing your Approach &amp; Insights! <a href=\"/khornlund\">@khornlund</a> </p>",
      "rawMarkdown": "Congratulations\nThanks for Sharing your Approach &amp; Insights! @khornlund",
      "votes": null
    },
    {
      "id": "677748",
      "postDate": "11/20/2019 14:50:52",
      "content": "<p>Thank you for sharing your writeup and codes. Learnt a lot from your last post as well in Steel.</p>\n\n<blockquote>\n  <p>Class 1 has very nice bimodal distribution. This suggests it was the easiest to learn.</p>\n</blockquote>\n\n<p>Could you give me some reference/links to learn more about these distributions and how having <code>bimodal</code> distribution helps in learning?</p>",
      "rawMarkdown": "Thank you for sharing your writeup and codes. Learnt a lot from your last post as well in Steel.\n&gt; Class 1 has very nice bimodal distribution. This suggests it was the easiest to learn.\n\nCould you give me some reference/links to learn more about these distributions and how having `bimodal` distribution helps in learning?",
      "votes": null
    },
    {
      "id": "677784",
      "postDate": "11/20/2019 15:32:33",
      "content": "<p>Congratulations! Thanks for sharing. I am a beginner was just curious about this, it's good to read the write up.  </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing. I am a beginner was just curious about this, it's good to read the write up.",
      "votes": null
    },
    {
      "id": "677824",
      "postDate": "11/20/2019 16:18:54",
      "content": "<p>I think what he means is this. The PDF plots are all max pixel values of a specific mask type. These numbers approximate the probability that an image has a specific mask type.</p>\n\n<p>If probability is below 0.3, we are confident that there should not be a mask. If probability is above 0.7, we are confident there should be a mask. When probability is between 0.3 and 0.7 we are unsure. A bimodal distribution has most of its predictions below 0.3 and above 0.7 and therefore doesn't have many confusing middle values.</p>",
      "rawMarkdown": "I think what he means is this. The PDF plots are all max pixel values of a specific mask type. These numbers approximate the probability that an image has a specific mask type.\n\nIf probability is below 0.3, we are confident that there should not be a mask. If probability is above 0.7, we are confident there should be a mask. When probability is between 0.3 and 0.7 we are unsure. A bimodal distribution has most of its predictions below 0.3 and above 0.7 and therefore doesn't have many confusing middle values.",
      "votes": null
    },
    {
      "id": "677828",
      "postDate": "11/20/2019 16:24:22",
      "content": "<p>Class 1 is Flower. My classifier also had the easiest time with Flowers achieving 85% accuracy and 0.92 AUC.</p>\n\n<pre><code>Classifier Accuracy\nFish : AUC = 0.815, ACC = 0.741\nFlower : AUC = 0.92, ACC = 0.85\nGravel : AUC = 0.808, ACC = 0.735\nSugar : AUC = 0.84, ACC = 0.786\nOVERALL: AUC = 0.8577, ACC = 0.7779\n</code></pre>",
      "rawMarkdown": "Class 1 is Flower. My classifier also had the easiest time with Flowers achieving 85% accuracy and 0.92 AUC.\n\n    Classifier Accuracy\n    Fish : AUC = 0.815, ACC = 0.741\n    Flower : AUC = 0.92, ACC = 0.85\n    Gravel : AUC = 0.808, ACC = 0.735\n    Sugar : AUC = 0.84, ACC = 0.786\n    OVERALL: AUC = 0.8577, ACC = 0.7779",
      "votes": null
    },
    {
      "id": "677841",
      "postDate": "11/20/2019 16:38:07",
      "content": "<p>Congrats Karl. Ensembling lots of models is a smart approach for this noisy data. It looks like you have some awesome models. You were able to ensemble your public LB to a high score. I think if you had a proper CV where all models used the same folds, you could ensemble your CV to an equally high score and achieve a higher private LB. After comp analysis shows that CV and private LB matched perfectly (but CV and public LB not perfectly).</p>\n\n<p>I like how you efficiently saved raw predictions to disk. That's a neat trick. I see that you modulated your thresholds with your classifier output again this comp. I meant to try that after reading that from you in Steel but forgot. My experiments confirm that that is a good trick. When building classifiers I found that segmentation models could predict empty masks better than pure classifiers (that don't use training masks). Therefore your trick is a way to ensemble the prediction ability of segmentation models with the prediction ability of pure classifiers.</p>",
      "rawMarkdown": "Congrats Karl. Ensembling lots of models is a smart approach for this noisy data. It looks like you have some awesome models. You were able to ensemble your public LB to a high score. I think if you had a proper CV where all models used the same folds, you could ensemble your CV to an equally high score and achieve a higher private LB. After comp analysis shows that CV and private LB matched perfectly (but CV and public LB not perfectly).\n\nI like how you efficiently saved raw predictions to disk. That's a neat trick. I see that you modulated your thresholds with your classifier output again this comp. I meant to try that after reading that from you in Steel but forgot. My experiments confirm that that is a good trick. When building classifiers I found that segmentation models could predict empty masks better than pure classifiers (that don't use training masks). Therefore your trick is a way to ensemble the prediction ability of segmentation models with the prediction ability of pure classifiers.",
      "votes": null
    },
    {
      "id": "677977",
      "postDate": "11/20/2019 20:51:56",
      "content": "<p>Exactly :)</p>",
      "rawMarkdown": "Exactly :)",
      "votes": null
    },
    {
      "id": "678176",
      "postDate": "11/21/2019 03:57:39",
      "content": "<p>Congrats, thanks for sharing</p>\n\n<p>the method train and label, and train similar to reinforcement learning, cool method</p>",
      "rawMarkdown": "Congrats, thanks for sharing\n\nthe method train and label, and train similar to reinforcement learning, cool method",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 677610,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "11/20/2019 11:42:01",
      "content": "<p>Congrats! And thanks for sharing your solution 😄 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 677714,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "11/20/2019 14:18:06",
      "content": "<p>Congratulations\nThanks for Sharing your Approach &amp; Insights! <a href=\"/khornlund\">@khornlund</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 677748,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "11/20/2019 14:50:52",
      "content": "<p>Thank you for sharing your writeup and codes. Learnt a lot from your last post as well in Steel.</p>\n\n<blockquote>\n  <p>Class 1 has very nice bimodal distribution. This suggests it was the easiest to learn.</p>\n</blockquote>\n\n<p>Could you give me some reference/links to learn more about these distributions and how having <code>bimodal</code> distribution helps in learning?</p>",
      "votes": null,
      "replies": [
        {
          "id": 677824,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/20/2019 16:18:54",
          "content": "<p>I think what he means is this. The PDF plots are all max pixel values of a specific mask type. These numbers approximate the probability that an image has a specific mask type.</p>\n\n<p>If probability is below 0.3, we are confident that there should not be a mask. If probability is above 0.7, we are confident there should be a mask. When probability is between 0.3 and 0.7 we are unsure. A bimodal distribution has most of its predictions below 0.3 and above 0.7 and therefore doesn't have many confusing middle values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677828,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/20/2019 16:24:22",
          "content": "<p>Class 1 is Flower. My classifier also had the easiest time with Flowers achieving 85% accuracy and 0.92 AUC.</p>\n\n<pre><code>Classifier Accuracy\nFish : AUC = 0.815, ACC = 0.741\nFlower : AUC = 0.92, ACC = 0.85\nGravel : AUC = 0.808, ACC = 0.735\nSugar : AUC = 0.84, ACC = 0.786\nOVERALL: AUC = 0.8577, ACC = 0.7779\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677977,
          "author_name": "khornlund",
          "author_url": "",
          "post_date": "11/20/2019 20:51:56",
          "content": "<p>Exactly :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 677784,
      "author_name": "garima85gupta",
      "author_url": "",
      "post_date": "11/20/2019 15:32:33",
      "content": "<p>Congratulations! Thanks for sharing. I am a beginner was just curious about this, it's good to read the write up.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 677841,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/20/2019 16:38:07",
      "content": "<p>Congrats Karl. Ensembling lots of models is a smart approach for this noisy data. It looks like you have some awesome models. You were able to ensemble your public LB to a high score. I think if you had a proper CV where all models used the same folds, you could ensemble your CV to an equally high score and achieve a higher private LB. After comp analysis shows that CV and private LB matched perfectly (but CV and public LB not perfectly).</p>\n\n<p>I like how you efficiently saved raw predictions to disk. That's a neat trick. I see that you modulated your thresholds with your classifier output again this comp. I meant to try that after reading that from you in Steel but forgot. My experiments confirm that that is a good trick. When building classifiers I found that segmentation models could predict empty masks better than pure classifiers (that don't use training masks). Therefore your trick is a way to ensemble the prediction ability of segmentation models with the prediction ability of pure classifiers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 678176,
      "author_name": "jt120lz",
      "author_url": "",
      "post_date": "11/21/2019 03:57:39",
      "content": "<p>Congrats, thanks for sharing</p>\n\n<p>the method train and label, and train similar to reinforcement learning, cool method</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "677578": "Hi all,\n\nCongrats to the winners! \n\nCode for my solution [here](https://github.com/khornlund/understanding-cloud-organization).\n\nExplanation copied below.\n\nSummary\n=======\n\nResults\n---------\n\n| Rank | Score | Percentile |\n| --- | --- | --- |\n| 34 | 0.66385 | Top 2.2% |\n\nStrategy\n-----------\nOriginally I had an idea early on very similar to [this](https://arxiv.org/pdf/1911.04252.pdf) recent paper. I was going to train a model on the ground truthed data, and then iteratively create pseudo labels for unlabelled data and train on that. I figured this was a good opportunity for such a strategy because there was very little training data (~5000 images), so there was a lot to be gained by generating more training samples. And, because this was not a synchronous kernel competition, I'd be able to create as large an ensemble as I like.\n\nThen I realised how noisy the image labels were, and wasn't so sure that pseudo labels would work very well. In particular, I noticed that the validation scores of my models was super noisy - using the same configuration with a different random seed resulted in serious metric differences. I figured I would give up on trying to fine tune individual models and instead focus on engineering a system that would allow me to train and ensemble *lots* of models.\n\nI developed functionality to allow me to automate the configuration, training, and inference of models.\n\nI trained an ensemble of ~120 models, using a variety of encoder/decoder combinations. I first averaged them together by their encoder/decoder combinations (eg. all the efficientnet-b2 FPN get averaged together). Then I averaged these mini-ensembles together using a weighted average.\n\nWith about a week of the competition to go, I saw the Noisy Student paper. I was getting decent results on the LB and figured I'd give pseudo labelling a go. I downloaded ~4200 images using the same resolution and locations as the official data, generated pseudo labels for them, and trained a new ensemble of ~50 models.\n\nI only finished training the pseudo labelled models in time to make a few submissions on the final day, and managed to get up to 0.67739 (9th place) on the public LB - but that actually only scored 0.66331 (~45th) on the private LB. My other selected submission was a weighted average of my past 25 submissions, which scored 0.67574 on the public LB and 0.66385 (34th) on the private LB.\n\nI had a few unselected submissions that scored 0.666+ (~18th), the best of which funnily enough came from a mini-ensemble of only efficientnet-b2-Unet models.\n\nReflection\n------------\nLooking back I realise I made a pretty big mistake not capturing the appropriate metrics for thorough local CV. I was only recording dice coefficient using a threshold of 0.5, and so I wasn't well informed to pick a threshold for my submissions.\n\nAlso, while the models were each trained on a random 80% of the data, and evaluated on the remaining 20%, this was only done at a per-model level. I didn't keep a hold-out set to validate the ensembles against. Because we only had ~5000 training samples, I got a bit greedy with training data here.\n\nI was hoping that by keeping logs of all my experiments, after a while I'd be able to identify which randomly generated configurations (eg. learning rate) worked better than others. This didn't turn out to be the case! I should have spent more time fine tuning each model, as the law of diminishing returns was coming into effect as the size of my ensemble grew.\n\nDetails\n=====\n\nEnsemble Pipeline\n----------------------\nSee ``uco.ensemble.py`` for implementation.\n\nEach training experiment is configured using a YAML file which gets loaded into a dictionary. I set up a class to randomise these parameters, so I could leave it to run while at work/sleep and it would cycle through different architectures, loss functions, and other parameters.\n\nAfter each training epoch the model would be evaluated on a 20% validation set. The mean dice score was tracked throughout training, and when the training completed (either after a set number of epochs or early stopping) only the best scoring checkpoint would be saved. I set a cutoff mean dice score, and threw away models that scored under that.\n\nThe saved checkpoint would be loaded, and run inference on the test data. I saved out the *raw* (sigmoid) predictions of each model to HDF5. I scaled by 250 and rounded to integers so I could save as ``uint8`` to save disk space.\n\nThese raw predictions would be grouped by (encoder, decoder) pair, and averaged together weighted by mean dice scores. Then the groups would be averaged together, with parameterised weights.\n\nBy saving out the results at each stage to HDF5 (raw predictions, group averages, and total averages), I could re-run any part of the pipeline with ease.\n\nI did the above for both segmentation and classification models. The details below are just for the segmentation models.\n\nModels\n---------\nI used [segmentation_models.pytorch](https://github.com/qubvel/segmentation_models.pytorch)\n(SMP) for segmentation, and used [pytorch-image-models](https://github.com/rwightman/pytorch-image-models) (TIIM) for classification.\n\n**Encoders**\n\n- efficientnet B0, B2, B5, B6\n- resnext 101_32x8d\n- se_resnext 101_32x8d\n- inceptionresnet v2, v4\n- dpn 131\n- densenet 161\n\n**Decoders**\n\n- FPN\n- Unet\n\nI had terrible results with LinkNet and PSPNet.\n\nTraining\n----------\n\n**GPU**\nRTX 2080Ti.\n\n**Loss**\nI used BCE + Dice with BCE weight ~U(0.65, 0.75) and dice weight 1 - BCE.\n\nI used BCE + Lovasz with BCE weight ~U(0.83, 0.92) and lovasz 1 - BCE.\n\n**Learning Rate**\nEncoder ~U(5e-5, 9e-5)\nDecoder ~U(3e-3, 5e-3)\n\n**Optimizer**\nRAdam / [QHAdamW](https://github.com/catalyst-team/catalyst/blob/master/catalyst/contrib/optimizers/qhadamw.py)\n\n**Augmentation**\nCompositions are in `data_loader.augmentation.py`.\n\nI made one custom augmentation - I modified Cutout to apply to masks. I wasn't sure if this would actually be better than only applying Cutout to the image - because the ground truth bounding boxes were large and covered areas that actually weren't very cloudy. It wasn't obvious from my experiments which worked better - but they both helped, so I just added them both to the available random configuration options for training.\n\n**Image Sizes**\nI wanted to use images sizes divisible by 32 so they would work without rounding effects, so I used the following which maintained the original 1400:2100 aspect ratio:\n\n- 256x384\n- 320x480\n- 384x576\n- 448x672\n\nMost models were trained using 320x480. I didn't notice any improvement using larger image sizes, but I figured it might help the ensemble to use diverse sizes.\n\n**Pseudo Labels**\nI used my ensemble trained on the official training data to predict masks for the ~4000 images I downloaded. I then removed any images without masks, and trained on the rest.\n\nIn contrast to some of the other people that used pseudo labels, I did not make my thresholds harsher for selecting pseudo labels. My rationale was that since most images included 2+ classes, increasing the thresholds to be 'safe' would likely mean missing the 2nd class in many images - leading to lots of false negative labels in my pseudo labels.\n\nI used a [balanced sampler](https://github.com/khornlund/pytorch-balanced-sampler) to include 4 pseudo labelled samples per batch (typically batch sizes were 10-16).\n\nPost-Processing\n-------------------\n\n**TTA**\nI used flips from [TTAch](https://github.com/qubvel/ttach)\n\n**Segmentation Thresholds**\nI experimented with a bunch of different ways to threshold positive predictions, as\nthe dice metric penalises false positives so heavily.\n\nI started out by using the following threshold rule:\n\n1. Outputs must have N pixels above some *top threshold*. I started out using N ~ 8000 for each class, and a top threshold of ~0.57.\n2. For predictions that pass (1), produce a binary mask using *bot threshold* of ~0.4.\n\nI used the continuous output of the classifier to modulate these thresholds. Ie. if the classifier was high, I would reduce the min size requirement, or the top threshold.\n\nIn the end I simply used maximum pixel prediction and no min size.\n\nThe distribution of predictions for the different classes is actually pretty interesting:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2116899%2Ff665763ea6d3668f7514f747997ad71d%2Faverage-prediction-distribution.png?generation=1574246988235284&amp;alt=media)\n\n\nClass 1 has very nice bimodal distribution. This suggests it was the easiest to learn.",
    "677610": "Congrats! And thanks for sharing your solution 😄",
    "677714": "Congratulations\nThanks for Sharing your Approach &amp; Insights! @khornlund",
    "677748": "Thank you for sharing your writeup and codes. Learnt a lot from your last post as well in Steel.\n&gt; Class 1 has very nice bimodal distribution. This suggests it was the easiest to learn.\n\nCould you give me some reference/links to learn more about these distributions and how having `bimodal` distribution helps in learning?",
    "677784": "Congratulations! Thanks for sharing. I am a beginner was just curious about this, it's good to read the write up.",
    "677824": "I think what he means is this. The PDF plots are all max pixel values of a specific mask type. These numbers approximate the probability that an image has a specific mask type.\n\nIf probability is below 0.3, we are confident that there should not be a mask. If probability is above 0.7, we are confident there should be a mask. When probability is between 0.3 and 0.7 we are unsure. A bimodal distribution has most of its predictions below 0.3 and above 0.7 and therefore doesn't have many confusing middle values.",
    "677828": "Class 1 is Flower. My classifier also had the easiest time with Flowers achieving 85% accuracy and 0.92 AUC.\n\n    Classifier Accuracy\n    Fish : AUC = 0.815, ACC = 0.741\n    Flower : AUC = 0.92, ACC = 0.85\n    Gravel : AUC = 0.808, ACC = 0.735\n    Sugar : AUC = 0.84, ACC = 0.786\n    OVERALL: AUC = 0.8577, ACC = 0.7779",
    "677841": "Congrats Karl. Ensembling lots of models is a smart approach for this noisy data. It looks like you have some awesome models. You were able to ensemble your public LB to a high score. I think if you had a proper CV where all models used the same folds, you could ensemble your CV to an equally high score and achieve a higher private LB. After comp analysis shows that CV and private LB matched perfectly (but CV and public LB not perfectly).\n\nI like how you efficiently saved raw predictions to disk. That's a neat trick. I see that you modulated your thresholds with your classifier output again this comp. I meant to try that after reading that from you in Steel but forgot. My experiments confirm that that is a good trick. When building classifiers I found that segmentation models could predict empty masks better than pure classifiers (that don't use training masks). Therefore your trick is a way to ensemble the prediction ability of segmentation models with the prediction ability of pure classifiers.",
    "677977": "Exactly :)",
    "678176": "Congrats, thanks for sharing\n\nthe method train and label, and train similar to reinforcement learning, cool method"
  },
  "source": "meta"
}