{
  "id": 107947,
  "title": "Kernels for 0.926 on private test set (37-50 place on LB)",
  "url": "/competitions/aptos2019-blindness-detection/writeups/peter-lex-kernels-for-0-926-on-private-test-set-37",
  "author_name": "",
  "post_date": "2019-09-08T09:06:36.670Z",
  "votes": 13,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>Thanks to APTOS and Kaggle for hosting the competition.</p>\n\n<p>I learned a lot in this competition and had a lot of fun searching the solution space with my amazing team mate <a href=\"/nemethpeti\">@nemethpeti</a>! Congratulations to Peter for making it to the top 3% in his first Kaggle competition!</p>\n\n<p>I thought I'd make public the solution that would have landed us somewhere between 37th and 50th had we selected it 😢Since we never felt confident in our validation strategy, we ended up selected the submission that has the best LB score and another with the most diverse ensemble of models. Neither of these turned out to be the best on the public LB.</p>\n\n<h2>Final ensemble</h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-256\">EfficientNet B3 - ord regression -  256px - radius reduction preproc - flip TTA</a></li>\n<li><a href=\"https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-300\">EfficinetNet B3 - ord regression - 300px - radius reduction preproc - flip TTA</a></li>\n<li><a href=\"https://www.kaggle.com/lextoumbourou/bd-densenet201-2015val-psu3-2019-val\">DenseNet101 - ord regression - 320px with Ben’s preprocessing - no TTA</a></li>\n</ul>\n\n<p>We blended the models with a very simple weighted average using weights: 0.4, 0.4 and 0.2, respectively.</p>\n\n<p><a href=\"https://www.kaggle.com/lextoumbourou/2-x-b3-densenet201-blend-0-926-on-private-lb\">Blend Kernel</a></p>\n\n<h2>Training notes</h2>\n\n<ul>\n<li><p><strong>Batch size</strong>: 64\nAny less than this and I found training to be very unstable.</p></li>\n<li><p><strong>Data</strong>: concat 2019 + 2015 training sets. Thanks to <a href=\"/tanlikesmath\">@tanlikesmath</a> for providing a <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">resized dataset</a>.\nI downsampled class 0 to be the same size as class 2. To ensure that the whole dataset was used I resampled class 0 each epoch.</p></li>\n<li><p><strong>Validation</strong>: 10000 examples from the 2015 test set with class 0 downsampled to match class 2. Thanks to <a href=\"/benjaminwarner\">@benjaminwarner</a> for providing a <a href=\"https://www.kaggle.com/benjaminwarner/resized-2015-2019-blindness-detection-images\">resized dataset</a> of the 2015 test data.</p></li>\n<li><p><strong>Preprocessing</strong>: Preprocessing copied from <a href=\"https://www.kaggle.com/joorarkesteijn/fast-cropping-preprocessing-and-augmentation\">joorarkesteijn's kernel</a> which used ideas from <a href=\"https://www.kaggle.com/ratthachat/aptos-updated-preprocessing-ben-s-cropping\">Neuron Engineer's kernel</a>. We used the gaussian blur subtraction method from Ben preprocessing for the DenseNet model only. The images were normalised using appropriate values for the model's pretrained weights.</p></li>\n<li><p><strong>Augmentations</strong>: flip_lr, brightness, contrast, rotate(360)</p></li>\n<li><p><strong>Transfer learning</strong>: all models were initialised with ImageNet weights.</p></li>\n<li><p><strong>Model head</strong>: <a href=\"https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets\">multiclass (ordinal regression) outputs</a>. For DenseNet, the penultimate FC layer was altered to have an output size of 2046.</p></li>\n<li><p><strong>Loss</strong>: BCEWithLogitsLoss with modified label smoothing.\nWe converted the ordinal regression from <code>[1, 1, 0, 0, 0]</code> labels into <code>[0.95, 0.95, 0.05, 0.05, 0.05]</code>.</p></li>\n<li><p><strong>Optimiser</strong>: Adam (fast.ai default)</p></li>\n<li><p><strong>Pseudo-labelling</strong>: add all test labels from our best submissions. This appeared to help a lot with our results on the public set, though it's debatable how useful it was on the private set.</p></li>\n<li><p><strong>Train</strong>: train just head for one epoch, then unfreeze all layers and train 15 epochs using cosign annealing.\nBest LR found using the technique from <a href=\"https://arxiv.org/pdf/1803.09820\">one cycle</a>. I also saved the best validation loss checkpoints.</p></li>\n<li><p><strong>Hardware</strong>: all training was exclusively done on Kaggle kernels.</p></li>\n<li><p><strong>Software</strong>: PyTorch and Fast.ai.</p></li>\n</ul>\n\n<h2>Secret sauce</h2>\n\n<p>Downsampling helped a lot in the early days of the comp. I found using the technique of resampling class 0 each epoch helped to reduce overfitting.</p>\n\n<p>Label smoothing helped a lot with overfitting and having a stable LB score.</p>\n\n<p>Ensembling models with slightly different preprocessing had big impact on our LB score in the early days. However, as our individual models improved, the improvement of this technique decreased.</p>\n\n<p>Pseudo labelling gave us some big improvements to our single model scores  (+1-2%) on the public set which helped with motivation. However, it doesn't appear to have had a significant effect on the private results for those models.</p>\n\n<h2>Other notes</h2>\n\n<p>I used ordinal regression based as per <a href=\"https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets\">my kernel</a>, this made it a bit of a battle to find a good way to ensemble our models as Peter's best model had used regression and classification for his best models. We tried linear stacking and converting my outputs to floats with various techniques and blending our outputs, In the end, our best kernel wasn't our combined ensemble, but our individual models were improved significantly by pooling our knowledge.</p>",
  "messages": [
    {
      "id": "620934",
      "postDate": "09/08/2019 04:38:04",
      "content": "<p>Hi all,</p>\n\n<p>Thanks to APTOS and Kaggle for hosting the competition.</p>\n\n<p>I learned a lot in this competition and had a lot of fun searching the solution space with my amazing team mate <a href=\"/nemethpeti\">@nemethpeti</a>! Congratulations to Peter for making it to the top 3% in his first Kaggle competition!</p>\n\n<p>I thought I'd make public the solution that would have landed us somewhere between 37th and 50th had we selected it 😢Since we never felt confident in our validation strategy, we ended up selected the submission that has the best LB score and another with the most diverse ensemble of models. Neither of these turned out to be the best on the public LB.</p>\n\n<h2>Final ensemble</h2>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-256\">EfficientNet B3 - ord regression -  256px - radius reduction preproc - flip TTA</a></li>\n<li><a href=\"https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-300\">EfficinetNet B3 - ord regression - 300px - radius reduction preproc - flip TTA</a></li>\n<li><a href=\"https://www.kaggle.com/lextoumbourou/bd-densenet201-2015val-psu3-2019-val\">DenseNet101 - ord regression - 320px with Ben’s preprocessing - no TTA</a></li>\n</ul>\n\n<p>We blended the models with a very simple weighted average using weights: 0.4, 0.4 and 0.2, respectively.</p>\n\n<p><a href=\"https://www.kaggle.com/lextoumbourou/2-x-b3-densenet201-blend-0-926-on-private-lb\">Blend Kernel</a></p>\n\n<h2>Training notes</h2>\n\n<ul>\n<li><p><strong>Batch size</strong>: 64\nAny less than this and I found training to be very unstable.</p></li>\n<li><p><strong>Data</strong>: concat 2019 + 2015 training sets. Thanks to <a href=\"/tanlikesmath\">@tanlikesmath</a> for providing a <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">resized dataset</a>.\nI downsampled class 0 to be the same size as class 2. To ensure that the whole dataset was used I resampled class 0 each epoch.</p></li>\n<li><p><strong>Validation</strong>: 10000 examples from the 2015 test set with class 0 downsampled to match class 2. Thanks to <a href=\"/benjaminwarner\">@benjaminwarner</a> for providing a <a href=\"https://www.kaggle.com/benjaminwarner/resized-2015-2019-blindness-detection-images\">resized dataset</a> of the 2015 test data.</p></li>\n<li><p><strong>Preprocessing</strong>: Preprocessing copied from <a href=\"https://www.kaggle.com/joorarkesteijn/fast-cropping-preprocessing-and-augmentation\">joorarkesteijn's kernel</a> which used ideas from <a href=\"https://www.kaggle.com/ratthachat/aptos-updated-preprocessing-ben-s-cropping\">Neuron Engineer's kernel</a>. We used the gaussian blur subtraction method from Ben preprocessing for the DenseNet model only. The images were normalised using appropriate values for the model's pretrained weights.</p></li>\n<li><p><strong>Augmentations</strong>: flip_lr, brightness, contrast, rotate(360)</p></li>\n<li><p><strong>Transfer learning</strong>: all models were initialised with ImageNet weights.</p></li>\n<li><p><strong>Model head</strong>: <a href=\"https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets\">multiclass (ordinal regression) outputs</a>. For DenseNet, the penultimate FC layer was altered to have an output size of 2046.</p></li>\n<li><p><strong>Loss</strong>: BCEWithLogitsLoss with modified label smoothing.\nWe converted the ordinal regression from <code>[1, 1, 0, 0, 0]</code> labels into <code>[0.95, 0.95, 0.05, 0.05, 0.05]</code>.</p></li>\n<li><p><strong>Optimiser</strong>: Adam (fast.ai default)</p></li>\n<li><p><strong>Pseudo-labelling</strong>: add all test labels from our best submissions. This appeared to help a lot with our results on the public set, though it's debatable how useful it was on the private set.</p></li>\n<li><p><strong>Train</strong>: train just head for one epoch, then unfreeze all layers and train 15 epochs using cosign annealing.\nBest LR found using the technique from <a href=\"https://arxiv.org/pdf/1803.09820\">one cycle</a>. I also saved the best validation loss checkpoints.</p></li>\n<li><p><strong>Hardware</strong>: all training was exclusively done on Kaggle kernels.</p></li>\n<li><p><strong>Software</strong>: PyTorch and Fast.ai.</p></li>\n</ul>\n\n<h2>Secret sauce</h2>\n\n<p>Downsampling helped a lot in the early days of the comp. I found using the technique of resampling class 0 each epoch helped to reduce overfitting.</p>\n\n<p>Label smoothing helped a lot with overfitting and having a stable LB score.</p>\n\n<p>Ensembling models with slightly different preprocessing had big impact on our LB score in the early days. However, as our individual models improved, the improvement of this technique decreased.</p>\n\n<p>Pseudo labelling gave us some big improvements to our single model scores  (+1-2%) on the public set which helped with motivation. However, it doesn't appear to have had a significant effect on the private results for those models.</p>\n\n<h2>Other notes</h2>\n\n<p>I used ordinal regression based as per <a href=\"https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets\">my kernel</a>, this made it a bit of a battle to find a good way to ensemble our models as Peter's best model had used regression and classification for his best models. We tried linear stacking and converting my outputs to floats with various techniques and blending our outputs, In the end, our best kernel wasn't our combined ensemble, but our individual models were improved significantly by pooling our knowledge.</p>",
      "rawMarkdown": "Hi all,\n\nThanks to APTOS and Kaggle for hosting the competition.\n\nI learned a lot in this competition and had a lot of fun searching the solution space with my amazing team mate @nemethpeti! Congratulations to Peter for making it to the top 3% in his first Kaggle competition!\n\nI thought I'd make public the solution that would have landed us somewhere between 37th and 50th had we selected it 😢Since we never felt confident in our validation strategy, we ended up selected the submission that has the best LB score and another with the most diverse ensemble of models. Neither of these turned out to be the best on the public LB.\n\n## Final ensemble\n\n* [EfficientNet B3 - ord regression -  256px - radius reduction preproc - flip TTA](https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-256)\n* [EfficinetNet B3 - ord regression - 300px - radius reduction preproc - flip TTA](https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-300)\n* [DenseNet101 - ord regression - 320px with Ben’s preprocessing - no TTA](https://www.kaggle.com/lextoumbourou/bd-densenet201-2015val-psu3-2019-val)\n\nWe blended the models with a very simple weighted average using weights: 0.4, 0.4 and 0.2, respectively.\n\n[Blend Kernel](https://www.kaggle.com/lextoumbourou/2-x-b3-densenet201-blend-0-926-on-private-lb)\n\n## Training notes\n\n* **Batch size**: 64\n  Any less than this and I found training to be very unstable.\n\n* **Data**: concat 2019 + 2015 training sets. Thanks to @tanlikesmath for providing a [resized dataset](https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized).\n  I downsampled class 0 to be the same size as class 2. To ensure that the whole dataset was used I resampled class 0 each epoch.\n\n* **Validation**: 10000 examples from the 2015 test set with class 0 downsampled to match class 2. Thanks to @benjaminwarner for providing a [resized dataset](https://www.kaggle.com/benjaminwarner/resized-2015-2019-blindness-detection-images) of the 2015 test data.\n\n* **Preprocessing**: Preprocessing copied from [joorarkesteijn's kernel](https://www.kaggle.com/joorarkesteijn/fast-cropping-preprocessing-and-augmentation) which used ideas from [Neuron Engineer's kernel](https://www.kaggle.com/ratthachat/aptos-updated-preprocessing-ben-s-cropping). We used the gaussian blur subtraction method from Ben preprocessing for the DenseNet model only. The images were normalised using appropriate values for the model's pretrained weights.\n\n* **Augmentations**: flip_lr, brightness, contrast, rotate(360)\n\n* **Transfer learning**: all models were initialised with ImageNet weights.\n\n* **Model head**: [multiclass (ordinal regression) outputs](https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets). For DenseNet, the penultimate FC layer was altered to have an output size of 2046.\n\n* **Loss**: BCEWithLogitsLoss with modified label smoothing.\n  We converted the ordinal regression from `[1, 1, 0, 0, 0]` labels into `[0.95, 0.95, 0.05, 0.05, 0.05]`.\n\n* **Optimiser**: Adam (fast.ai default)\n\n* **Pseudo-labelling**: add all test labels from our best submissions. This appeared to help a lot with our results on the public set, though it's debatable how useful it was on the private set.\n\n* **Train**: train just head for one epoch, then unfreeze all layers and train 15 epochs using cosign annealing.\nBest LR found using the technique from [one cycle](https://arxiv.org/pdf/1803.09820). I also saved the best validation loss checkpoints.\n\n* **Hardware**: all training was exclusively done on Kaggle kernels.\n\n* **Software**: PyTorch and Fast.ai.\n\n## Secret sauce\n\nDownsampling helped a lot in the early days of the comp. I found using the technique of resampling class 0 each epoch helped to reduce overfitting.\n\nLabel smoothing helped a lot with overfitting and having a stable LB score.\n\nEnsembling models with slightly different preprocessing had big impact on our LB score in the early days. However, as our individual models improved, the improvement of this technique decreased.\n\nPseudo labelling gave us some big improvements to our single model scores  (+1-2%) on the public set which helped with motivation. However, it doesn't appear to have had a significant effect on the private results for those models.\n\n## Other notes\n\n I used ordinal regression based as per [my kernel](https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets), this made it a bit of a battle to find a good way to ensemble our models as Peter's best model had used regression and classification for his best models. We tried linear stacking and converting my outputs to floats with various techniques and blending our outputs, In the end, our best kernel wasn't our combined ensemble, but our individual models were improved significantly by pooling our knowledge.",
      "votes": null
    },
    {
      "id": "620941",
      "postDate": "09/08/2019 04:52:15",
      "content": "<p>Congratulations\nGreat work\nAnd Thanks for sharing your approach and insights.!!</p>",
      "rawMarkdown": "Congratulations\nGreat work\nAnd Thanks for sharing your approach and insights.!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 620941,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "09/08/2019 04:52:15",
      "content": "<p>Congratulations\nGreat work\nAnd Thanks for sharing your approach and insights.!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "620934": "Hi all,\n\nThanks to APTOS and Kaggle for hosting the competition.\n\nI learned a lot in this competition and had a lot of fun searching the solution space with my amazing team mate @nemethpeti! Congratulations to Peter for making it to the top 3% in his first Kaggle competition!\n\nI thought I'd make public the solution that would have landed us somewhere between 37th and 50th had we selected it 😢Since we never felt confident in our validation strategy, we ended up selected the submission that has the best LB score and another with the most diverse ensemble of models. Neither of these turned out to be the best on the public LB.\n\n## Final ensemble\n\n* [EfficientNet B3 - ord regression -  256px - radius reduction preproc - flip TTA](https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-256)\n* [EfficinetNet B3 - ord regression - 300px - radius reduction preproc - flip TTA](https://www.kaggle.com/lextoumbourou/bd-efficientnetb3-2015-val-sz-300)\n* [DenseNet101 - ord regression - 320px with Ben’s preprocessing - no TTA](https://www.kaggle.com/lextoumbourou/bd-densenet201-2015val-psu3-2019-val)\n\nWe blended the models with a very simple weighted average using weights: 0.4, 0.4 and 0.2, respectively.\n\n[Blend Kernel](https://www.kaggle.com/lextoumbourou/2-x-b3-densenet201-blend-0-926-on-private-lb)\n\n## Training notes\n\n* **Batch size**: 64\n  Any less than this and I found training to be very unstable.\n\n* **Data**: concat 2019 + 2015 training sets. Thanks to @tanlikesmath for providing a [resized dataset](https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized).\n  I downsampled class 0 to be the same size as class 2. To ensure that the whole dataset was used I resampled class 0 each epoch.\n\n* **Validation**: 10000 examples from the 2015 test set with class 0 downsampled to match class 2. Thanks to @benjaminwarner for providing a [resized dataset](https://www.kaggle.com/benjaminwarner/resized-2015-2019-blindness-detection-images) of the 2015 test data.\n\n* **Preprocessing**: Preprocessing copied from [joorarkesteijn's kernel](https://www.kaggle.com/joorarkesteijn/fast-cropping-preprocessing-and-augmentation) which used ideas from [Neuron Engineer's kernel](https://www.kaggle.com/ratthachat/aptos-updated-preprocessing-ben-s-cropping). We used the gaussian blur subtraction method from Ben preprocessing for the DenseNet model only. The images were normalised using appropriate values for the model's pretrained weights.\n\n* **Augmentations**: flip_lr, brightness, contrast, rotate(360)\n\n* **Transfer learning**: all models were initialised with ImageNet weights.\n\n* **Model head**: [multiclass (ordinal regression) outputs](https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets). For DenseNet, the penultimate FC layer was altered to have an output size of 2046.\n\n* **Loss**: BCEWithLogitsLoss with modified label smoothing.\n  We converted the ordinal regression from `[1, 1, 0, 0, 0]` labels into `[0.95, 0.95, 0.05, 0.05, 0.05]`.\n\n* **Optimiser**: Adam (fast.ai default)\n\n* **Pseudo-labelling**: add all test labels from our best submissions. This appeared to help a lot with our results on the public set, though it's debatable how useful it was on the private set.\n\n* **Train**: train just head for one epoch, then unfreeze all layers and train 15 epochs using cosign annealing.\nBest LR found using the technique from [one cycle](https://arxiv.org/pdf/1803.09820). I also saved the best validation loss checkpoints.\n\n* **Hardware**: all training was exclusively done on Kaggle kernels.\n\n* **Software**: PyTorch and Fast.ai.\n\n## Secret sauce\n\nDownsampling helped a lot in the early days of the comp. I found using the technique of resampling class 0 each epoch helped to reduce overfitting.\n\nLabel smoothing helped a lot with overfitting and having a stable LB score.\n\nEnsembling models with slightly different preprocessing had big impact on our LB score in the early days. However, as our individual models improved, the improvement of this technique decreased.\n\nPseudo labelling gave us some big improvements to our single model scores  (+1-2%) on the public set which helped with motivation. However, it doesn't appear to have had a significant effect on the private results for those models.\n\n## Other notes\n\n I used ordinal regression based as per [my kernel](https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets), this made it a bit of a battle to find a good way to ensemble our models as Peter's best model had used regression and classification for his best models. We tried linear stacking and converting my outputs to floats with various techniques and blending our outputs, In the end, our best kernel wasn't our combined ensemble, but our individual models were improved significantly by pooling our knowledge.",
    "620941": "Congratulations\nGreat work\nAnd Thanks for sharing your approach and insights.!!"
  },
  "source": "meta"
}