{
  "id": 77355,
  "title": "231st place keras solution",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/wonderbolts-231st-place-keras-solution",
  "author_name": "",
  "post_date": "2019-03-05T20:52:25.050Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Congratulations to all the winners and great thanks to discussion participants!</p>\n\n<p>We started the competition using Keras because we were familiar with it and then preferred to stick to it. </p>\n\n<p>At first, we were just trying different models (Inception, ResNets, Xception) and some small architectures (MobileNet, SqueezeNet) and chose to use pretrained ResNet50 from Keras applications. After long time with it and no result more than 0.47 LB we switched to Inception networks. </p>\n\n<p>All models were trained on single fold: 85% train and 15% test.\nWe splitted data using iterative stratification package from discussion.\nFrom external data we used only rare classes and excluded all duplicates from training.</p>\n\n<h2>Things that worked</h2>\n\n<ul>\n<li>F1 loss and BCE with weighted positive predictions</li>\n<li>SGD with warm restarts (cosine learning rate, initial lr=0.1)</li>\n<li>Log-dampened class weights</li>\n<li>Oversampling of rare classes by 5x</li>\n<li>Undersampling classes 0 and 25 by 0.5x</li>\n<li>\"Standard\" augmentations (flip rl/ud, shift, rotate, shear, brightness)</li>\n<li>Squeeze-Excitation module</li>\n<li>Weight decay</li>\n<li>Randomly dropping green channel and all classes in 1% of images led to +0.04 LB</li>\n<li>BatchNorm running mean/variance update on test set</li>\n</ul>\n\n<h2>Things that did not work</h2>\n\n<ul>\n<li>Imagenet weights (we still wonder why)</li>\n<li>Focal loss</li>\n<li>Threshold fitting</li>\n<li>GapNet-like things</li>\n<li>Convolutional block attention module</li>\n<li>Filling empty predictions with most popular class</li>\n<li>Mixup, probably because of the specialty of the green channel</li>\n<li>Heavy augmentations (blur, various types of noise, dropout, contrast, elastic transformations, sharpen/emboss)</li>\n<li>Small crops (less than 40% of number of pixels)</li>\n<li>AdamW</li>\n<li>Global average + global max pooling</li>\n<li>TTA (but not hurt)</li>\n<li>BatchNorm in dense layers reduced performance dramatically</li>\n</ul>\n\n<h2>Things that we regret:</h2>\n\n<ul>\n<li>Realized the way we overfit (high recall, low precision) too late (2 days before the end of the competition)</li>\n<li>Not using cross-validation</li>\n<li>Not ensembling different architectures</li>\n<li>Closer to final: using Keras</li>\n</ul>\n\n<h2>Top solution</h2>\n\n<p>Top performing solution scored 0.505 on private LB. We didn't choose it because it performed worse both in local validation and on public LB. \nIt was an ensemble of all the best SE-BNInception (6 models) with fixed threshold. 4 of them were RGBY and 2 were RGB.</p>\n\n<h2>Hardware</h2>\n\n<p>We used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).\nOccasionally we used standard workstations with 1060 and 8-16GB RAM.</p>",
  "messages": [
    {
      "id": "454523",
      "postDate": "01/11/2019 19:28:12",
      "content": "<p>Congratulations to all the winners and great thanks to discussion participants!</p>\n\n<p>We started the competition using Keras because we were familiar with it and then preferred to stick to it. </p>\n\n<p>At first, we were just trying different models (Inception, ResNets, Xception) and some small architectures (MobileNet, SqueezeNet) and chose to use pretrained ResNet50 from Keras applications. After long time with it and no result more than 0.47 LB we switched to Inception networks. </p>\n\n<p>All models were trained on single fold: 85% train and 15% test.\nWe splitted data using iterative stratification package from discussion.\nFrom external data we used only rare classes and excluded all duplicates from training.</p>\n\n<h2>Things that worked</h2>\n\n<ul>\n<li>F1 loss and BCE with weighted positive predictions</li>\n<li>SGD with warm restarts (cosine learning rate, initial lr=0.1)</li>\n<li>Log-dampened class weights</li>\n<li>Oversampling of rare classes by 5x</li>\n<li>Undersampling classes 0 and 25 by 0.5x</li>\n<li>\"Standard\" augmentations (flip rl/ud, shift, rotate, shear, brightness)</li>\n<li>Squeeze-Excitation module</li>\n<li>Weight decay</li>\n<li>Randomly dropping green channel and all classes in 1% of images led to +0.04 LB</li>\n<li>BatchNorm running mean/variance update on test set</li>\n</ul>\n\n<h2>Things that did not work</h2>\n\n<ul>\n<li>Imagenet weights (we still wonder why)</li>\n<li>Focal loss</li>\n<li>Threshold fitting</li>\n<li>GapNet-like things</li>\n<li>Convolutional block attention module</li>\n<li>Filling empty predictions with most popular class</li>\n<li>Mixup, probably because of the specialty of the green channel</li>\n<li>Heavy augmentations (blur, various types of noise, dropout, contrast, elastic transformations, sharpen/emboss)</li>\n<li>Small crops (less than 40% of number of pixels)</li>\n<li>AdamW</li>\n<li>Global average + global max pooling</li>\n<li>TTA (but not hurt)</li>\n<li>BatchNorm in dense layers reduced performance dramatically</li>\n</ul>\n\n<h2>Things that we regret:</h2>\n\n<ul>\n<li>Realized the way we overfit (high recall, low precision) too late (2 days before the end of the competition)</li>\n<li>Not using cross-validation</li>\n<li>Not ensembling different architectures</li>\n<li>Closer to final: using Keras</li>\n</ul>\n\n<h2>Top solution</h2>\n\n<p>Top performing solution scored 0.505 on private LB. We didn't choose it because it performed worse both in local validation and on public LB. \nIt was an ensemble of all the best SE-BNInception (6 models) with fixed threshold. 4 of them were RGBY and 2 were RGB.</p>\n\n<h2>Hardware</h2>\n\n<p>We used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).\nOccasionally we used standard workstations with 1060 and 8-16GB RAM.</p>",
      "rawMarkdown": "Congratulations to all the winners and great thanks to discussion participants!\n\nWe started the competition using Keras because we were familiar with it and then preferred to stick to it. \n\nAt first, we were just trying different models (Inception, ResNets, Xception) and some small architectures (MobileNet, SqueezeNet) and chose to use pretrained ResNet50 from Keras applications. After long time with it and no result more than 0.47 LB we switched to Inception networks. \n\nAll models were trained on single fold: 85% train and 15% test.\nWe splitted data using iterative stratification package from discussion.\nFrom external data we used only rare classes and excluded all duplicates from training.\n\n## Things that worked ##\n\n- F1 loss and BCE with weighted positive predictions\n- SGD with warm restarts (cosine learning rate, initial lr=0.1)\n- Log-dampened class weights\n- Oversampling of rare classes by 5x\n- Undersampling classes 0 and 25 by 0.5x\n- \"Standard\" augmentations (flip rl/ud, shift, rotate, shear, brightness)\n- Squeeze-Excitation module\n- Weight decay\n- Randomly dropping green channel and all classes in 1% of images led to +0.04 LB\n- BatchNorm running mean/variance update on test set\n\n## Things that did not work ##\n- Imagenet weights (we still wonder why)\n- Focal loss\n- Threshold fitting\n- GapNet-like things\n- Convolutional block attention module\n- Filling empty predictions with most popular class\n- Mixup, probably because of the specialty of the green channel\n- Heavy augmentations (blur, various types of noise, dropout, contrast, elastic transformations, sharpen/emboss)\n- Small crops (less than 40% of number of pixels)\n- AdamW\n- Global average + global max pooling\n- TTA (but not hurt)\n- BatchNorm in dense layers reduced performance dramatically\n\n## Things that we regret: ##\n- Realized the way we overfit (high recall, low precision) too late (2 days before the end of the competition)\n- Not using cross-validation\n- Not ensembling different architectures\n- Closer to final: using Keras\n\n## Top solution ##\nTop performing solution scored 0.505 on private LB. We didn't choose it because it performed worse both in local validation and on public LB. \nIt was an ensemble of all the best SE-BNInception (6 models) with fixed threshold. 4 of them were RGBY and 2 were RGB.\n\n## Hardware ##\nWe used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).\nOccasionally we used standard workstations with 1060 and 8-16GB RAM.",
      "votes": null
    },
    {
      "id": "454651",
      "postDate": "01/11/2019 23:26:15",
      "content": "<p>Thanks for sharing. We also tried different things for filling the empty predictions, but found nothing that worked</p>",
      "rawMarkdown": "Thanks for sharing. We also tried different things for filling the empty predictions, but found nothing that worked",
      "votes": null
    },
    {
      "id": "1079465",
      "postDate": "11/16/2020 05:12:27",
      "content": "<p>Hi, your solution was really impressive. I am new to Kaggle and just looking for some competitions for exercise. I wonder how you play with the squeeze-and-excitation structure with pretrained model from Keras, since Keras is well encapsulated and I wonder how to regularly insert layers of SEnet into pretrained networks like ResNet50. Plus if I insert SE layers into ResNet50, is it still appropriate to use the pretrained weights of it?<br>\nI'll really appreciate it for your reply. Thanks a lot.</p>",
      "rawMarkdown": "Hi, your solution was really impressive. I am new to Kaggle and just looking for some competitions for exercise. I wonder how you play with the squeeze-and-excitation structure with pretrained model from Keras, since Keras is well encapsulated and I wonder how to regularly insert layers of SEnet into pretrained networks like ResNet50. Plus if I insert SE layers into ResNet50, is it still appropriate to use the pretrained weights of it?\nI'll really appreciate it for your reply. Thanks a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1079465,
      "author_name": "ruijunzhou",
      "author_url": "",
      "post_date": "11/16/2020 05:12:27",
      "content": "<p>Hi, your solution was really impressive. I am new to Kaggle and just looking for some competitions for exercise. I wonder how you play with the squeeze-and-excitation structure with pretrained model from Keras, since Keras is well encapsulated and I wonder how to regularly insert layers of SEnet into pretrained networks like ResNet50. Plus if I insert SE layers into ResNet50, is it still appropriate to use the pretrained weights of it?<br>\nI'll really appreciate it for your reply. Thanks a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454651,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "01/11/2019 23:26:15",
      "content": "<p>Thanks for sharing. We also tried different things for filling the empty predictions, but found nothing that worked</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454523": "Congratulations to all the winners and great thanks to discussion participants!\n\nWe started the competition using Keras because we were familiar with it and then preferred to stick to it. \n\nAt first, we were just trying different models (Inception, ResNets, Xception) and some small architectures (MobileNet, SqueezeNet) and chose to use pretrained ResNet50 from Keras applications. After long time with it and no result more than 0.47 LB we switched to Inception networks. \n\nAll models were trained on single fold: 85% train and 15% test.\nWe splitted data using iterative stratification package from discussion.\nFrom external data we used only rare classes and excluded all duplicates from training.\n\n## Things that worked ##\n\n- F1 loss and BCE with weighted positive predictions\n- SGD with warm restarts (cosine learning rate, initial lr=0.1)\n- Log-dampened class weights\n- Oversampling of rare classes by 5x\n- Undersampling classes 0 and 25 by 0.5x\n- \"Standard\" augmentations (flip rl/ud, shift, rotate, shear, brightness)\n- Squeeze-Excitation module\n- Weight decay\n- Randomly dropping green channel and all classes in 1% of images led to +0.04 LB\n- BatchNorm running mean/variance update on test set\n\n## Things that did not work ##\n- Imagenet weights (we still wonder why)\n- Focal loss\n- Threshold fitting\n- GapNet-like things\n- Convolutional block attention module\n- Filling empty predictions with most popular class\n- Mixup, probably because of the specialty of the green channel\n- Heavy augmentations (blur, various types of noise, dropout, contrast, elastic transformations, sharpen/emboss)\n- Small crops (less than 40% of number of pixels)\n- AdamW\n- Global average + global max pooling\n- TTA (but not hurt)\n- BatchNorm in dense layers reduced performance dramatically\n\n## Things that we regret: ##\n- Realized the way we overfit (high recall, low precision) too late (2 days before the end of the competition)\n- Not using cross-validation\n- Not ensembling different architectures\n- Closer to final: using Keras\n\n## Top solution ##\nTop performing solution scored 0.505 on private LB. We didn't choose it because it performed worse both in local validation and on public LB. \nIt was an ensemble of all the best SE-BNInception (6 models) with fixed threshold. 4 of them were RGBY and 2 were RGB.\n\n## Hardware ##\nWe used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).\nOccasionally we used standard workstations with 1060 and 8-16GB RAM.",
    "454651": "Thanks for sharing. We also tried different things for filling the empty predictions, but found nothing that worked",
    "1079465": "Hi, your solution was really impressive. I am new to Kaggle and just looking for some competitions for exercise. I wonder how you play with the squeeze-and-excitation structure with pretrained model from Keras, since Keras is well encapsulated and I wonder how to regularly insert layers of SEnet into pretrained networks like ResNet50. Plus if I insert SE layers into ResNet50, is it still appropriate to use the pretrained weights of it?\nI'll really appreciate it for your reply. Thanks a lot."
  },
  "source": "meta"
}