{
  "id": 168507,
  "title": "[12th place] Solution Overview",
  "url": "/competitions/alaska2-image-steganalysis/writeups/ian-pan-felipe-kitamura-12th-place-solution-overvi",
  "author_name": "",
  "post_date": "2020-07-21T11:06:31.223Z",
  "votes": 51,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Congrats to all the winners and participants. Thank you to the organizers for this interesting competition! </p>\n<p>Wow, <a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> and I are super surprised that we somehow ended up in the gold medal range! Our solution doesn't really have a lot of novelty, but here it is, in brief:</p>\n<p>4-model ensemble:</p>\n<ul>\n<li>EfficientNet-B8 (0.930 -&gt; 0.920)</li>\n<li>EfficientNet-B5, change initial stride 2 to stride 1 (0.926 -&gt; 0.920)</li>\n<li>EfficientNet-B4, change initial stride 2 to stride 1 (0.928 -&gt; 0.920)</li>\n<li>EfficientNet-B0, change first two stride 2 to stride 1 (0.930 -&gt; 0.921)</li>\n</ul>\n<p>All models were trained on original 512x512 RGB input with flip augmentation.</p>\n<ul>\n<li>Initialization: ImageNet</li>\n<li>4 classes</li>\n<li>Multisample dropout, dropout 0.2</li>\n<li>Concat pooling (concatenate global max and global average pooling)</li>\n<li>Vanilla cross-entropy loss</li>\n<li>AdamW optimizer, cosine decay, initial LR 5.0e-4, 50 epochs. B4 trained slightly differently, using OneCycleLR (30% warmup, peak LR 3.0e-4, 30 epochs) with original striding, then fine-tuned for 30 epochs with initial stride 1</li>\n<li>Single 90/10 split</li>\n<li>Models were trained on 4x Quadro RTX 6000 24GB using a batch size of around 32-40. B8/B5 models were trained on 8x V100 32GB. We didn't start seriously participating in this competition until the last 2 weeks, so without these hardware resources it would have been very difficult to do well. This is probably the competition where I used the most GPU compute, as I usually don't even bother with multi-GPU training.</li>\n</ul>\n<p>It was tough to get local CV because I kept running into a bug with the wAUC computation… In order to fix it, it made all the wAUCs for the models fall into the 0.34-0.35 range. Still unsure what was going on!</p>\n<p>Best single model public LB 0.930 (private 0.921) -&gt; simple ensemble LB 0.936 (private 0.928).</p>\n<p>Other things we tried:</p>\n<ul>\n<li>TTA: many other teams had improvements with TTA, but we did not. For example, the B0 model went from 0.930 to 0.926 on public LB with TTA. <strong>However</strong>, this improved private LB performance from 0.920 to 0.924! </li>\n<li>Mixup: we experimented with a mixup strategy where we mixed together the cover and stego images of the same source. It took much longer to train, though, so we abandoned the idea. </li>\n<li>ResNet/ResNeSt: didn't work as well. Removing the max pooling helped, but EfficientNets were better. MixNet also performed well, but EfficientNets were the best overall. </li>\n</ul>",
  "messages": [
    {
      "id": "937337",
      "postDate": "07/21/2020 00:08:43",
      "content": "<p>Congrats to all the winners and participants. Thank you to the organizers for this interesting competition! </p>\n<p>Wow, <a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> and I are super surprised that we somehow ended up in the gold medal range! Our solution doesn't really have a lot of novelty, but here it is, in brief:</p>\n<p>4-model ensemble:</p>\n<ul>\n<li>EfficientNet-B8 (0.930 -&gt; 0.920)</li>\n<li>EfficientNet-B5, change initial stride 2 to stride 1 (0.926 -&gt; 0.920)</li>\n<li>EfficientNet-B4, change initial stride 2 to stride 1 (0.928 -&gt; 0.920)</li>\n<li>EfficientNet-B0, change first two stride 2 to stride 1 (0.930 -&gt; 0.921)</li>\n</ul>\n<p>All models were trained on original 512x512 RGB input with flip augmentation.</p>\n<ul>\n<li>Initialization: ImageNet</li>\n<li>4 classes</li>\n<li>Multisample dropout, dropout 0.2</li>\n<li>Concat pooling (concatenate global max and global average pooling)</li>\n<li>Vanilla cross-entropy loss</li>\n<li>AdamW optimizer, cosine decay, initial LR 5.0e-4, 50 epochs. B4 trained slightly differently, using OneCycleLR (30% warmup, peak LR 3.0e-4, 30 epochs) with original striding, then fine-tuned for 30 epochs with initial stride 1</li>\n<li>Single 90/10 split</li>\n<li>Models were trained on 4x Quadro RTX 6000 24GB using a batch size of around 32-40. B8/B5 models were trained on 8x V100 32GB. We didn't start seriously participating in this competition until the last 2 weeks, so without these hardware resources it would have been very difficult to do well. This is probably the competition where I used the most GPU compute, as I usually don't even bother with multi-GPU training.</li>\n</ul>\n<p>It was tough to get local CV because I kept running into a bug with the wAUC computation… In order to fix it, it made all the wAUCs for the models fall into the 0.34-0.35 range. Still unsure what was going on!</p>\n<p>Best single model public LB 0.930 (private 0.921) -&gt; simple ensemble LB 0.936 (private 0.928).</p>\n<p>Other things we tried:</p>\n<ul>\n<li>TTA: many other teams had improvements with TTA, but we did not. For example, the B0 model went from 0.930 to 0.926 on public LB with TTA. <strong>However</strong>, this improved private LB performance from 0.920 to 0.924! </li>\n<li>Mixup: we experimented with a mixup strategy where we mixed together the cover and stego images of the same source. It took much longer to train, though, so we abandoned the idea. </li>\n<li>ResNet/ResNeSt: didn't work as well. Removing the max pooling helped, but EfficientNets were better. MixNet also performed well, but EfficientNets were the best overall. </li>\n</ul>",
      "rawMarkdown": "Congrats to all the winners and participants. Thank you to the organizers for this interesting competition! \n\nWow, @felipekitamura and I are super surprised that we somehow ended up in the gold medal range! Our solution doesn't really have a lot of novelty, but here it is, in brief:\n\n4-model ensemble:\n- EfficientNet-B8 (0.930 -&gt; 0.920)\n- EfficientNet-B5, change initial stride 2 to stride 1 (0.926 -&gt; 0.920)\n- EfficientNet-B4, change initial stride 2 to stride 1 (0.928 -&gt; 0.920)\n- EfficientNet-B0, change first two stride 2 to stride 1 (0.930 -&gt; 0.921)\n\nAll models were trained on original 512x512 RGB input with flip augmentation.\n\n- Initialization: ImageNet\n- 4 classes\n- Multisample dropout, dropout 0.2\n- Concat pooling (concatenate global max and global average pooling)\n- Vanilla cross-entropy loss\n- AdamW optimizer, cosine decay, initial LR 5.0e-4, 50 epochs. B4 trained slightly differently, using OneCycleLR (30% warmup, peak LR 3.0e-4, 30 epochs) with original striding, then fine-tuned for 30 epochs with initial stride 1\n- Single 90/10 split\n- Models were trained on 4x Quadro RTX 6000 24GB using a batch size of around 32-40. B8/B5 models were trained on 8x V100 32GB. We didn't start seriously participating in this competition until the last 2 weeks, so without these hardware resources it would have been very difficult to do well. This is probably the competition where I used the most GPU compute, as I usually don't even bother with multi-GPU training.\n\nIt was tough to get local CV because I kept running into a bug with the wAUC computation... In order to fix it, it made all the wAUCs for the models fall into the 0.34-0.35 range. Still unsure what was going on!\n\nBest single model public LB 0.930 (private 0.921) -&gt; simple ensemble LB 0.936 (private 0.928).\n\nOther things we tried:\n- TTA: many other teams had improvements with TTA, but we did not. For example, the B0 model went from 0.930 to 0.926 on public LB with TTA. **However**, this improved private LB performance from 0.920 to 0.924! \n- Mixup: we experimented with a mixup strategy where we mixed together the cover and stego images of the same source. It took much longer to train, though, so we abandoned the idea. \n- ResNet/ResNeSt: didn't work as well. Removing the max pooling helped, but EfficientNets were better. MixNet also performed well, but EfficientNets were the best overall.",
      "votes": null
    },
    {
      "id": "937344",
      "postDate": "07/21/2020 00:21:44",
      "content": "<p>Congrats <a href=\"/vaillant\">@vaillant</a> and <a href=\"/felipekitamura\">@felipekitamura</a>. Definitely changing first layer stride to 1 is one of the keys here.</p>",
      "rawMarkdown": "Congrats @vaillant and @felipekitamura. Definitely changing first layer stride to 1 is one of the keys here.",
      "votes": null
    },
    {
      "id": "937346",
      "postDate": "07/21/2020 00:23:21",
      "content": "<p>wow, LB is my best single model. ensemble gets lower LB score☹️ . I am looking forward your simple ensemble.</p>",
      "rawMarkdown": "wow, LB is my best single model. ensemble gets lower LB score☹️ . I am looking forward your simple ensemble.",
      "votes": null
    },
    {
      "id": "937357",
      "postDate": "07/21/2020 00:50:42",
      "content": "<p>Thanks <a href=\"/titericz\">@titericz</a>. Congrats to you too!</p>\n\n<p>It's an honor to receive congrats from the grand master <a href=\"/titericz\">@titericz</a>!</p>\n\n<p>Ainda mora nos EUA?</p>",
      "rawMarkdown": "Thanks @titericz. Congrats to you too!\n\nIt's an honor to receive congrats from the grand master @titericz!\n\nAinda mora nos EUA?",
      "votes": null
    },
    {
      "id": "937360",
      "postDate": "07/21/2020 00:55:54",
      "content": "<p>Congrats! The method is simple but efficient!\nI failed to come up with changing stride to 1 during the competition, I will use it in my next competition. Thanks for your solution!</p>",
      "rawMarkdown": "Congrats! The method is simple but efficient!\nI failed to come up with changing stride to 1 during the competition, I will use it in my next competition. Thanks for your solution!",
      "votes": null
    },
    {
      "id": "937363",
      "postDate": "07/21/2020 00:59:45",
      "content": "<p>Congrats！ \nDoes anyone can express what is the insight of changing first layer stride to 1. Thanks！</p>",
      "rawMarkdown": "Congrats！ \nDoes anyone can express what is the insight of changing first layer stride to 1. Thanks！",
      "votes": null
    },
    {
      "id": "937380",
      "postDate": "07/21/2020 01:22:32",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> congrats on the result. As someone pretty new to image competitions may I ask if this is the line that you modified?</p>\n<p><a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/35dbb4b72c2f587ae8b499a9d405939b41929cda/efficientnet_pytorch/model.py#L170\" target=\"_blank\">https://github.com/lukemelas/EfficientNet-PyTorch/blob/35dbb4b72c2f587ae8b499a9d405939b41929cda/efficientnet_pytorch/model.py#L170</a></p>\n<p>Here you set <code>stride=1</code>?</p>\n<p>I am also curious to know what your experimentation process looks like?</p>\n<p>A few more questions that come to mind - what optimizer and learning rate schedule did you use for your different models?</p>",
      "rawMarkdown": "vaillant congrats on the result. As someone pretty new to image competitions may I ask if this is the line that you modified?\n\nhttps://github.com/lukemelas/EfficientNet-PyTorch/blob/35dbb4b72c2f587ae8b499a9d405939b41929cda/efficientnet_pytorch/model.py#L170\n\nHere you set `stride=1`?\n\nI am also curious to know what your experimentation process looks like?\n\nA few more questions that come to mind - what optimizer and learning rate schedule did you use for your different models?",
      "votes": null
    },
    {
      "id": "937388",
      "postDate": "07/21/2020 01:39:12",
      "content": "<p>My intuition was that having the network operate on full resolution input for as long as possible would improve performance since we are trying to detect small, subtle changes in the image. As was obvious early on, downsampling the input (e.g., to 256x256) worsened performance significantly. </p>",
      "rawMarkdown": "My intuition was that having the network operate on full resolution input for as long as possible would improve performance since we are trying to detect small, subtle changes in the image. As was obvious early on, downsampling the input (e.g., to 256x256) worsened performance significantly.",
      "votes": null
    },
    {
      "id": "937389",
      "postDate": "07/21/2020 01:40:26",
      "content": "<p>I used the implementation from <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a>. </p>\n<pre><code>model = timm.models.efficientnet_b4()\nmodel.conv_stem.stride = 1\n</code></pre>",
      "rawMarkdown": "I used the implementation from https://github.com/rwightman/pytorch-image-models. \n\n```\nmodel = timm.models.efficientnet_b4()\nmodel.conv_stem.stride = 1\n```",
      "votes": null
    },
    {
      "id": "937391",
      "postDate": "07/21/2020 01:40:37",
      "content": "<p><a href=\"/gzblue\">@gzblue</a></p>\n\n<p>\"Does anyone can express what is the insight of changing first layer stride to 1. Thanks！\"</p>\n\n<p>capture high frequency.\nupsampling the input image (2x enlarge) should gives better results, but is not feasible in this case.</p>\n\n<p>bookmarked this! It is a common trick in image based kaggle competition</p>",
      "rawMarkdown": "gzblue\n\n\"Does anyone can express what is the insight of changing first layer stride to 1. Thanks！\"\n\ncapture high frequency.\nupsampling the input image (2x enlarge) should gives better results, but is not feasible in this case.\n\nbookmarked this! It is a common trick in image based kaggle competition",
      "votes": null
    },
    {
      "id": "937403",
      "postDate": "07/21/2020 01:59:51",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> - another quick question - did you have to use any tricks to be able to train the higher efficientnets? B5 &amp; B8 for example did you use gradient accumulation, apex, etc?</p>",
      "rawMarkdown": "vaillant - another quick question - did you have to use any tricks to be able to train the higher efficientnets? B5 &amp; B8 for example did you use gradient accumulation, apex, etc?",
      "votes": null
    },
    {
      "id": "937416",
      "postDate": "07/21/2020 02:16:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>,</p>\n<p>Simple but effective solution indeed. Can you help me understand the hardware set up (Colab/Kaggle TPU/GPU) you used to train the model. How long did it take for each model to train. And did you use keras/pytorch for it ?</p>",
      "rawMarkdown": "Hi @vaillant,\n\nSimple but effective solution indeed. Can you help me understand the hardware set up (Colab/Kaggle TPU/GPU) you used to train the model. How long did it take for each model to train. And did you use keras/pytorch for it ?",
      "votes": null
    },
    {
      "id": "937452",
      "postDate": "07/21/2020 03:14:47",
      "content": "<p>Congrats on result and thanks for sharing your solution <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>!</p>",
      "rawMarkdown": "Congrats on result and thanks for sharing your solution @vaillant!",
      "votes": null
    },
    {
      "id": "937585",
      "postDate": "07/21/2020 05:03:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> 💯💯💯</p>\n<p>Have you used TPU for training ? For me anything greater that B3 took a very long time to finish one epoch.</p>",
      "rawMarkdown": "Congrats @vaillant 💯💯💯\n\nHave you used TPU for training ? For me anything greater that B3 took a very long time to finish one epoch.",
      "votes": null
    },
    {
      "id": "937648",
      "postDate": "07/21/2020 05:43:00",
      "content": "<p>That's clever. Keeping original Image resolution as long as possible certainly helps.<br>\nIndeed, we are trying to detect extremely low signal and nothing guarantees that a harsh pooling will preserve this subtle piece of information.<br>\nAnother Kaggler (<a href=\"https://www.kaggle.com/davidaustin\" target=\"_blank\">@davidaustin</a>) made this observation that goes along the same direction the <br>\n<em>\"deep learning topologies that give the best results all seem to have a common block, which is the MBConv block associated with the MobileNetV2 block sequence (EfficientNet, MobileNet, MixNet, MNasNet) […]  I believe there's a key learning to derive here […]  it has to do with the widening of the network with 1x1 convolutions to enhance the channelwise separation […] It would also partially explain why some other standard network topologies (ie Resnet, DenseNet) don't perform well.\"</em><br>\nGreat findings</p>",
      "rawMarkdown": "That's clever. Keeping original Image resolution as long as possible certainly helps.\nIndeed, we are trying to detect extremely low signal and nothing guarantees that a harsh pooling will preserve this subtle piece of information.\nAnother Kaggler ([@davidaustin](https://www.kaggle.com/davidaustin)) made this observation that goes along the same direction the \n*\"deep learning topologies that give the best results all seem to have a common block, which is the MBConv block associated with the MobileNetV2 block sequence (EfficientNet, MobileNet, MixNet, MNasNet) [...]  I believe there's a key learning to derive here [...]  it has to do with the widening of the network with 1x1 convolutions to enhance the channelwise separation [...] It would also partially explain why some other standard network topologies (ie Resnet, DenseNet) don't perform well.\"*\nGreat findings",
      "votes": null
    },
    {
      "id": "938186",
      "postDate": "07/21/2020 11:48:14",
      "content": "<p>This is real intelligent engineering. Simple but great work. Congrats <a href=\"/vaillant\">@vaillant</a> </p>",
      "rawMarkdown": "This is real intelligent engineering. Simple but great work. Congrats @vaillant",
      "votes": null
    },
    {
      "id": "938977",
      "postDate": "07/21/2020 23:13:11",
      "content": "<p>Hi, we used 4 24GB Quadro RTX6000 and also rented 8 32GB V100s from AWS for 2 days. The B5 model took 2 days on 8 V100s to train, and other models took 3-5 days on 4 24GB GPUs. </p>",
      "rawMarkdown": "Hi, we used 4 24GB Quadro RTX6000 and also rented 8 32GB V100s from AWS for 2 days. The B5 model took 2 days on 8 V100s to train, and other models took 3-5 days on 4 24GB GPUs.",
      "votes": null
    },
    {
      "id": "939619",
      "postDate": "07/22/2020 10:38:01",
      "content": "<p>Congratulations! I'm amazed that you can reach the gold medal in two weeks.\nI have one question, how did you implement Concat pooling and Multisample dropout? Is there any easy way to do this?</p>",
      "rawMarkdown": "Congratulations! I'm amazed that you can reach the gold medal in two weeks.\nI have one question, how did you implement Concat pooling and Multisample dropout? Is there any easy way to do this?",
      "votes": null
    },
    {
      "id": "939800",
      "postDate": "07/22/2020 13:16:59",
      "content": "<p>Using PyTorch, it's quite easy. Here is the implementation of concat pooling, borrowed from fast.ai:</p>\n<pre><code>class AdaptiveConcatPool2d(nn.Module):\n\n    def forward(self, x):\n        return torch.cat((F.adaptive_avg_pool2d(x, 1), F.adaptive_max_pool2d(x, 1)), dim=1)\n</code></pre>\n<p>This is what I used for multisample dropout in the model's forward pass. I believe it originated from <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>:</p>\n<pre><code>x = torch.mean(\n    torch.stack(\n        [self.fc(self.dropout(features)) for _ in range(5)],\n        dim=0,\n    ),\n    dim=0,\n)\n</code></pre>",
      "rawMarkdown": "Using PyTorch, it's quite easy. Here is the implementation of concat pooling, borrowed from fast.ai:\n\n```\nclass AdaptiveConcatPool2d(nn.Module):\n\n    def forward(self, x):\n        return torch.cat((F.adaptive_avg_pool2d(x, 1), F.adaptive_max_pool2d(x, 1)), dim=1)\n```\n\nThis is what I used for multisample dropout in the model's forward pass. I believe it originated from @haqishen:\n\n```\nx = torch.mean(\n    torch.stack(\n        [self.fc(self.dropout(features)) for _ in range(5)],\n        dim=0,\n    ),\n    dim=0,\n)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937380,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "07/21/2020 01:22:32",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> congrats on the result. As someone pretty new to image competitions may I ask if this is the line that you modified?</p>\n<p><a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/35dbb4b72c2f587ae8b499a9d405939b41929cda/efficientnet_pytorch/model.py#L170\" target=\"_blank\">https://github.com/lukemelas/EfficientNet-PyTorch/blob/35dbb4b72c2f587ae8b499a9d405939b41929cda/efficientnet_pytorch/model.py#L170</a></p>\n<p>Here you set <code>stride=1</code>?</p>\n<p>I am also curious to know what your experimentation process looks like?</p>\n<p>A few more questions that come to mind - what optimizer and learning rate schedule did you use for your different models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 937389,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "07/21/2020 01:40:26",
          "content": "<p>I used the implementation from <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a>. </p>\n<pre><code>model = timm.models.efficientnet_b4()\nmodel.conv_stem.stride = 1\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937403,
          "author_name": "rdizzl3",
          "author_url": "",
          "post_date": "07/21/2020 01:59:51",
          "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> - another quick question - did you have to use any tricks to be able to train the higher efficientnets? B5 &amp; B8 for example did you use gradient accumulation, apex, etc?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937416,
      "author_name": "shwetank3",
      "author_url": "",
      "post_date": "07/21/2020 02:16:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>,</p>\n<p>Simple but effective solution indeed. Can you help me understand the hardware set up (Colab/Kaggle TPU/GPU) you used to train the model. How long did it take for each model to train. And did you use keras/pytorch for it ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 938977,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "07/21/2020 23:13:11",
          "content": "<p>Hi, we used 4 24GB Quadro RTX6000 and also rented 8 32GB V100s from AWS for 2 days. The B5 model took 2 days on 8 V100s to train, and other models took 3-5 days on 4 24GB GPUs. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937452,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "07/21/2020 03:14:47",
      "content": "<p>Congrats on result and thanks for sharing your solution <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937585,
      "author_name": "vishnurapps",
      "author_url": "",
      "post_date": "07/21/2020 05:03:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> 💯💯💯</p>\n<p>Have you used TPU for training ? For me anything greater that B3 took a very long time to finish one epoch.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937344,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/21/2020 00:21:44",
      "content": "<p>Congrats <a href=\"/vaillant\">@vaillant</a> and <a href=\"/felipekitamura\">@felipekitamura</a>. Definitely changing first layer stride to 1 is one of the keys here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 937357,
          "author_name": "felipekitamura",
          "author_url": "",
          "post_date": "07/21/2020 00:50:42",
          "content": "<p>Thanks <a href=\"/titericz\">@titericz</a>. Congrats to you too!</p>\n\n<p>It's an honor to receive congrats from the grand master <a href=\"/titericz\">@titericz</a>!</p>\n\n<p>Ainda mora nos EUA?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937346,
      "author_name": "sharksbeer",
      "author_url": "",
      "post_date": "07/21/2020 00:23:21",
      "content": "<p>wow, LB is my best single model. ensemble gets lower LB score☹️ . I am looking forward your simple ensemble.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937360,
      "author_name": "charliezhao",
      "author_url": "",
      "post_date": "07/21/2020 00:55:54",
      "content": "<p>Congrats! The method is simple but efficient!\nI failed to come up with changing stride to 1 during the competition, I will use it in my next competition. Thanks for your solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937363,
      "author_name": "gzblue0926",
      "author_url": "",
      "post_date": "07/21/2020 00:59:45",
      "content": "<p>Congrats！ \nDoes anyone can express what is the insight of changing first layer stride to 1. Thanks！</p>",
      "votes": null,
      "replies": [
        {
          "id": 937388,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "07/21/2020 01:39:12",
          "content": "<p>My intuition was that having the network operate on full resolution input for as long as possible would improve performance since we are trying to detect small, subtle changes in the image. As was obvious early on, downsampling the input (e.g., to 256x256) worsened performance significantly. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937391,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/21/2020 01:40:37",
          "content": "<p><a href=\"/gzblue\">@gzblue</a></p>\n\n<p>\"Does anyone can express what is the insight of changing first layer stride to 1. Thanks！\"</p>\n\n<p>capture high frequency.\nupsampling the input image (2x enlarge) should gives better results, but is not feasible in this case.</p>\n\n<p>bookmarked this! It is a common trick in image based kaggle competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937648,
          "author_name": "remicogranne",
          "author_url": "",
          "post_date": "07/21/2020 05:43:00",
          "content": "<p>That's clever. Keeping original Image resolution as long as possible certainly helps.<br>\nIndeed, we are trying to detect extremely low signal and nothing guarantees that a harsh pooling will preserve this subtle piece of information.<br>\nAnother Kaggler (<a href=\"https://www.kaggle.com/davidaustin\" target=\"_blank\">@davidaustin</a>) made this observation that goes along the same direction the <br>\n<em>\"deep learning topologies that give the best results all seem to have a common block, which is the MBConv block associated with the MobileNetV2 block sequence (EfficientNet, MobileNet, MixNet, MNasNet) […]  I believe there's a key learning to derive here […]  it has to do with the widening of the network with 1x1 convolutions to enhance the channelwise separation […] It would also partially explain why some other standard network topologies (ie Resnet, DenseNet) don't perform well.\"</em><br>\nGreat findings</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 938186,
      "author_name": "kalyanpichuka",
      "author_url": "",
      "post_date": "07/21/2020 11:48:14",
      "content": "<p>This is real intelligent engineering. Simple but great work. Congrats <a href=\"/vaillant\">@vaillant</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 939619,
      "author_name": "ajtryt2",
      "author_url": "",
      "post_date": "07/22/2020 10:38:01",
      "content": "<p>Congratulations! I'm amazed that you can reach the gold medal in two weeks.\nI have one question, how did you implement Concat pooling and Multisample dropout? Is there any easy way to do this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 939800,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "07/22/2020 13:16:59",
          "content": "<p>Using PyTorch, it's quite easy. Here is the implementation of concat pooling, borrowed from fast.ai:</p>\n<pre><code>class AdaptiveConcatPool2d(nn.Module):\n\n    def forward(self, x):\n        return torch.cat((F.adaptive_avg_pool2d(x, 1), F.adaptive_max_pool2d(x, 1)), dim=1)\n</code></pre>\n<p>This is what I used for multisample dropout in the model's forward pass. I believe it originated from <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>:</p>\n<pre><code>x = torch.mean(\n    torch.stack(\n        [self.fc(self.dropout(features)) for _ in range(5)],\n        dim=0,\n    ),\n    dim=0,\n)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "937337": "Congrats to all the winners and participants. Thank you to the organizers for this interesting competition! \n\nWow, @felipekitamura and I are super surprised that we somehow ended up in the gold medal range! Our solution doesn't really have a lot of novelty, but here it is, in brief:\n\n4-model ensemble:\n- EfficientNet-B8 (0.930 -&gt; 0.920)\n- EfficientNet-B5, change initial stride 2 to stride 1 (0.926 -&gt; 0.920)\n- EfficientNet-B4, change initial stride 2 to stride 1 (0.928 -&gt; 0.920)\n- EfficientNet-B0, change first two stride 2 to stride 1 (0.930 -&gt; 0.921)\n\nAll models were trained on original 512x512 RGB input with flip augmentation.\n\n- Initialization: ImageNet\n- 4 classes\n- Multisample dropout, dropout 0.2\n- Concat pooling (concatenate global max and global average pooling)\n- Vanilla cross-entropy loss\n- AdamW optimizer, cosine decay, initial LR 5.0e-4, 50 epochs. B4 trained slightly differently, using OneCycleLR (30% warmup, peak LR 3.0e-4, 30 epochs) with original striding, then fine-tuned for 30 epochs with initial stride 1\n- Single 90/10 split\n- Models were trained on 4x Quadro RTX 6000 24GB using a batch size of around 32-40. B8/B5 models were trained on 8x V100 32GB. We didn't start seriously participating in this competition until the last 2 weeks, so without these hardware resources it would have been very difficult to do well. This is probably the competition where I used the most GPU compute, as I usually don't even bother with multi-GPU training.\n\nIt was tough to get local CV because I kept running into a bug with the wAUC computation... In order to fix it, it made all the wAUCs for the models fall into the 0.34-0.35 range. Still unsure what was going on!\n\nBest single model public LB 0.930 (private 0.921) -&gt; simple ensemble LB 0.936 (private 0.928).\n\nOther things we tried:\n- TTA: many other teams had improvements with TTA, but we did not. For example, the B0 model went from 0.930 to 0.926 on public LB with TTA. **However**, this improved private LB performance from 0.920 to 0.924! \n- Mixup: we experimented with a mixup strategy where we mixed together the cover and stego images of the same source. It took much longer to train, though, so we abandoned the idea. \n- ResNet/ResNeSt: didn't work as well. Removing the max pooling helped, but EfficientNets were better. MixNet also performed well, but EfficientNets were the best overall.",
    "937344": "Congrats @vaillant and @felipekitamura. Definitely changing first layer stride to 1 is one of the keys here.",
    "937346": "wow, LB is my best single model. ensemble gets lower LB score☹️ . I am looking forward your simple ensemble.",
    "937357": "Thanks @titericz. Congrats to you too!\n\nIt's an honor to receive congrats from the grand master @titericz!\n\nAinda mora nos EUA?",
    "937360": "Congrats! The method is simple but efficient!\nI failed to come up with changing stride to 1 during the competition, I will use it in my next competition. Thanks for your solution!",
    "937363": "Congrats！ \nDoes anyone can express what is the insight of changing first layer stride to 1. Thanks！",
    "937380": "vaillant congrats on the result. As someone pretty new to image competitions may I ask if this is the line that you modified?\n\nhttps://github.com/lukemelas/EfficientNet-PyTorch/blob/35dbb4b72c2f587ae8b499a9d405939b41929cda/efficientnet_pytorch/model.py#L170\n\nHere you set `stride=1`?\n\nI am also curious to know what your experimentation process looks like?\n\nA few more questions that come to mind - what optimizer and learning rate schedule did you use for your different models?",
    "937388": "My intuition was that having the network operate on full resolution input for as long as possible would improve performance since we are trying to detect small, subtle changes in the image. As was obvious early on, downsampling the input (e.g., to 256x256) worsened performance significantly.",
    "937389": "I used the implementation from https://github.com/rwightman/pytorch-image-models. \n\n```\nmodel = timm.models.efficientnet_b4()\nmodel.conv_stem.stride = 1\n```",
    "937391": "gzblue\n\n\"Does anyone can express what is the insight of changing first layer stride to 1. Thanks！\"\n\ncapture high frequency.\nupsampling the input image (2x enlarge) should gives better results, but is not feasible in this case.\n\nbookmarked this! It is a common trick in image based kaggle competition",
    "937403": "vaillant - another quick question - did you have to use any tricks to be able to train the higher efficientnets? B5 &amp; B8 for example did you use gradient accumulation, apex, etc?",
    "937416": "Hi @vaillant,\n\nSimple but effective solution indeed. Can you help me understand the hardware set up (Colab/Kaggle TPU/GPU) you used to train the model. How long did it take for each model to train. And did you use keras/pytorch for it ?",
    "937452": "Congrats on result and thanks for sharing your solution @vaillant!",
    "937585": "Congrats @vaillant 💯💯💯\n\nHave you used TPU for training ? For me anything greater that B3 took a very long time to finish one epoch.",
    "937648": "That's clever. Keeping original Image resolution as long as possible certainly helps.\nIndeed, we are trying to detect extremely low signal and nothing guarantees that a harsh pooling will preserve this subtle piece of information.\nAnother Kaggler ([@davidaustin](https://www.kaggle.com/davidaustin)) made this observation that goes along the same direction the \n*\"deep learning topologies that give the best results all seem to have a common block, which is the MBConv block associated with the MobileNetV2 block sequence (EfficientNet, MobileNet, MixNet, MNasNet) [...]  I believe there's a key learning to derive here [...]  it has to do with the widening of the network with 1x1 convolutions to enhance the channelwise separation [...] It would also partially explain why some other standard network topologies (ie Resnet, DenseNet) don't perform well.\"*\nGreat findings",
    "938186": "This is real intelligent engineering. Simple but great work. Congrats @vaillant",
    "938977": "Hi, we used 4 24GB Quadro RTX6000 and also rented 8 32GB V100s from AWS for 2 days. The B5 model took 2 days on 8 V100s to train, and other models took 3-5 days on 4 24GB GPUs.",
    "939619": "Congratulations! I'm amazed that you can reach the gold medal in two weeks.\nI have one question, how did you implement Concat pooling and Multisample dropout? Is there any easy way to do this?",
    "939800": "Using PyTorch, it's quite easy. Here is the implementation of concat pooling, borrowed from fast.ai:\n\n```\nclass AdaptiveConcatPool2d(nn.Module):\n\n    def forward(self, x):\n        return torch.cat((F.adaptive_avg_pool2d(x, 1), F.adaptive_max_pool2d(x, 1)), dim=1)\n```\n\nThis is what I used for multisample dropout in the model's forward pass. I believe it originated from @haqishen:\n\n```\nx = torch.mean(\n    torch.stack(\n        [self.fc(self.dropout(features)) for _ in range(5)],\n        dim=0,\n    ),\n    dim=0,\n)\n```"
  },
  "source": "meta"
}