{
  "id": 154351,
  "title": "1st place solution",
  "url": "/competitions/herbarium-2020-fgvc7/writeups/vinbdi-medicalimagingteam-1st-place-solution",
  "author_name": "",
  "post_date": "2020-05-28T06:37:20.627Z",
  "votes": 21,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Firstly, I would like to thank the host team and Kaggle for hosting such an interesting competition as well as my teammates (<a href=\"/dattran2346\">@dattran2346</a>, <a href=\"/moewie94\">@moewie94</a>) for their hard work and contributions.</p>\n\n<p><strong>Model Architecture</strong></p>\n\n<p>We used EfficientNet B3/4/5 (noisy student version) with original image sizes (320x320, 380x380, 456x456), Inception V4 and SEResNeXt50 with image size 448x448 as our backbones. \nFor the EfficientNets, we first extracted feature maps from block 4, 5, and 6, forwarded through  one-squeeze multi-excitation modules (<a href=\"https://eccv2018.org/openaccess/content_ECCV_2018/papers/Ming_Sun_Multi-Attention_Multi-Class_Constraint_ECCV_2018_paper.pdf\">OSME</a>), before pooling and concatenating them into a \"local\" features vector. The ArcFace(s=16, m=0.1) loss + BNNeck combination was then used with \"head\" block features as input to help our models learn better \"global\" representation. \nThis pipeline was trained jointly end-to-end with Cross Entropy losses (no label smoothing) and <a href=\"https://papers.nips.cc/paper/7344-maximum-entropy-fine-grained-classification.pdf\">entropy losses</a>; at inference time, we averaged the 3 classifiers' logits (local classifier, ArcFace and BNNeck).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1134433%2Fb488481298423a5b4b3ecdbf549b1ad2%2Fmodel.png?generation=1590638315354457&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Training</strong></p>\n\n<p>We employed a two-stage training which is similar to 4th ranked team. In the first stage, only samples from the top 5000 most frequent categories were used. After obtaining a good model (measured in terms of validation f1-score), we employed the weight imprinting trick to initialize the classifiers' weights for the remaining 27093 categories then re-trained the entire network until convergence. Below are our training configs.\n+ Stage 1: 75 epochs, batch size = 2048 (gradient accumulation), cosine scheduler with linear warmup and AdamW optimizer, balanced sampling (randomly sample 75 images per category per epoch).\n+ Stage 2: Increase number of epochs to 105, balanced sampling (randomly sample 25 images per category per epoch).\n+ Augmentations: <code>ResizeShorterSide(random.randint(img_size + 32, img_size + 128)),  RandomCrop(img_size),  RandAugment(num_ops=2, magnitude=random.randint(1, 11), \n RandomErasing()</code></p>\n\n<p><strong>Scores</strong>\nWe reported some single model scores without any post-processing here\n+ Inception V4 - Public 0.79712 - Private 0.77781\n+ B3 - Public 0.80570 - Private 0.78182\n+ B4 - Public 0.83014 - Private 0.81488\n+ B5 - Public 0.84774 - Private 0.83309</p>\n\n<p><strong>Ensemble</strong> \n We took the code and idea in this<a href=\"https://www.kaggle.com/mathormad/5th-place-solution-stacking-pipeline\"> notebook</a> from last year RSNA challenge, modified them to build a simple 2-layer stacking network to combine predictions from our single models. The stacking ensemble obtained public score 0.85505 and private score 0.84180.\n```\nKERNEL_SIZE = 9\nNUM_CHANNELS = 128</p>\n\n<p>class StackingCNN(nn.Module):\n    def <strong>init</strong>(self, num_models, num_channels=NUM_CHANNELS):\n        super(StackingCNN, self).<strong>init</strong>()\n        self.base = nn.Sequential(OrderedDict([\n            ('conv1', nn.Conv2d(1, num_channels,\n                kernel_size=(num_models, 1),\n                padding=(int((num_models - 1) / 2), 0))),\n            ('relu1', nn.ReLU(inplace=True)),\n            ('conv2', nn.Conv2d(num_channels, num_channels,\n                                kernel_size=(1, KERNEL_SIZE), padding=(0, int((KERNEL_SIZE - 1) / 2)))),\n            ('relu2', nn.ReLU(inplace=True))\n        ]))</p>\n\n<pre><code>def forward(self, images, labels=None):\n    with autocast():\n        x = self.base(images)\n        x = x.permute(0, 3, 1, 2)\n        x = fast_global_avg_pool_2d(x)\n        if self.training:\n            losses = {\n                \"avg_loss\": F.cross_entropy(x, labels)\n                }\n            return losses\n        else:\n            outputs = {\"logits\": x}\n            return outputs\n</code></pre>\n\n<p>```</p>\n\n<p><strong>Postprocessing</strong>\nWe exploited the fact that the test set has the number of examples per species capped at a maximum of 10. For each species, we first sorted the predictions based on their top 1 confidence scores. For predictions with confidence scores outside top 10, we changed them to top 2 category instead. This simple trick gave pretty consistent boost (0.2 - 0.3 %) on public LB for every single submission file. Our final submission scored 0.85797 on public and 0.84484 on private.</p>",
  "messages": [
    {
      "id": "864601",
      "postDate": "05/28/2020 04:58:22",
      "content": "<p>Firstly, I would like to thank the host team and Kaggle for hosting such an interesting competition as well as my teammates (<a href=\"/dattran2346\">@dattran2346</a>, <a href=\"/moewie94\">@moewie94</a>) for their hard work and contributions.</p>\n\n<p><strong>Model Architecture</strong></p>\n\n<p>We used EfficientNet B3/4/5 (noisy student version) with original image sizes (320x320, 380x380, 456x456), Inception V4 and SEResNeXt50 with image size 448x448 as our backbones. \nFor the EfficientNets, we first extracted feature maps from block 4, 5, and 6, forwarded through  one-squeeze multi-excitation modules (<a href=\"https://eccv2018.org/openaccess/content_ECCV_2018/papers/Ming_Sun_Multi-Attention_Multi-Class_Constraint_ECCV_2018_paper.pdf\">OSME</a>), before pooling and concatenating them into a \"local\" features vector. The ArcFace(s=16, m=0.1) loss + BNNeck combination was then used with \"head\" block features as input to help our models learn better \"global\" representation. \nThis pipeline was trained jointly end-to-end with Cross Entropy losses (no label smoothing) and <a href=\"https://papers.nips.cc/paper/7344-maximum-entropy-fine-grained-classification.pdf\">entropy losses</a>; at inference time, we averaged the 3 classifiers' logits (local classifier, ArcFace and BNNeck).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1134433%2Fb488481298423a5b4b3ecdbf549b1ad2%2Fmodel.png?generation=1590638315354457&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Training</strong></p>\n\n<p>We employed a two-stage training which is similar to 4th ranked team. In the first stage, only samples from the top 5000 most frequent categories were used. After obtaining a good model (measured in terms of validation f1-score), we employed the weight imprinting trick to initialize the classifiers' weights for the remaining 27093 categories then re-trained the entire network until convergence. Below are our training configs.\n+ Stage 1: 75 epochs, batch size = 2048 (gradient accumulation), cosine scheduler with linear warmup and AdamW optimizer, balanced sampling (randomly sample 75 images per category per epoch).\n+ Stage 2: Increase number of epochs to 105, balanced sampling (randomly sample 25 images per category per epoch).\n+ Augmentations: <code>ResizeShorterSide(random.randint(img_size + 32, img_size + 128)),  RandomCrop(img_size),  RandAugment(num_ops=2, magnitude=random.randint(1, 11), \n RandomErasing()</code></p>\n\n<p><strong>Scores</strong>\nWe reported some single model scores without any post-processing here\n+ Inception V4 - Public 0.79712 - Private 0.77781\n+ B3 - Public 0.80570 - Private 0.78182\n+ B4 - Public 0.83014 - Private 0.81488\n+ B5 - Public 0.84774 - Private 0.83309</p>\n\n<p><strong>Ensemble</strong> \n We took the code and idea in this<a href=\"https://www.kaggle.com/mathormad/5th-place-solution-stacking-pipeline\"> notebook</a> from last year RSNA challenge, modified them to build a simple 2-layer stacking network to combine predictions from our single models. The stacking ensemble obtained public score 0.85505 and private score 0.84180.\n```\nKERNEL_SIZE = 9\nNUM_CHANNELS = 128</p>\n\n<p>class StackingCNN(nn.Module):\n    def <strong>init</strong>(self, num_models, num_channels=NUM_CHANNELS):\n        super(StackingCNN, self).<strong>init</strong>()\n        self.base = nn.Sequential(OrderedDict([\n            ('conv1', nn.Conv2d(1, num_channels,\n                kernel_size=(num_models, 1),\n                padding=(int((num_models - 1) / 2), 0))),\n            ('relu1', nn.ReLU(inplace=True)),\n            ('conv2', nn.Conv2d(num_channels, num_channels,\n                                kernel_size=(1, KERNEL_SIZE), padding=(0, int((KERNEL_SIZE - 1) / 2)))),\n            ('relu2', nn.ReLU(inplace=True))\n        ]))</p>\n\n<pre><code>def forward(self, images, labels=None):\n    with autocast():\n        x = self.base(images)\n        x = x.permute(0, 3, 1, 2)\n        x = fast_global_avg_pool_2d(x)\n        if self.training:\n            losses = {\n                \"avg_loss\": F.cross_entropy(x, labels)\n                }\n            return losses\n        else:\n            outputs = {\"logits\": x}\n            return outputs\n</code></pre>\n\n<p>```</p>\n\n<p><strong>Postprocessing</strong>\nWe exploited the fact that the test set has the number of examples per species capped at a maximum of 10. For each species, we first sorted the predictions based on their top 1 confidence scores. For predictions with confidence scores outside top 10, we changed them to top 2 category instead. This simple trick gave pretty consistent boost (0.2 - 0.3 %) on public LB for every single submission file. Our final submission scored 0.85797 on public and 0.84484 on private.</p>",
      "rawMarkdown": "Firstly, I would like to thank the host team and Kaggle for hosting such an interesting competition as well as my teammates (@dattran2346, @moewie94) for their hard work and contributions.\n\n**Model Architecture**\n\nWe used EfficientNet B3/4/5 (noisy student version) with original image sizes (320x320, 380x380, 456x456), Inception V4 and SEResNeXt50 with image size 448x448 as our backbones. \nFor the EfficientNets, we first extracted feature maps from block 4, 5, and 6, forwarded through  one-squeeze multi-excitation modules ([OSME](https://eccv2018.org/openaccess/content_ECCV_2018/papers/Ming_Sun_Multi-Attention_Multi-Class_Constraint_ECCV_2018_paper.pdf)), before pooling and concatenating them into a \"local\" features vector. The ArcFace(s=16, m=0.1) loss + BNNeck combination was then used with \"head\" block features as input to help our models learn better \"global\" representation. \nThis pipeline was trained jointly end-to-end with Cross Entropy losses (no label smoothing) and [entropy losses](https://papers.nips.cc/paper/7344-maximum-entropy-fine-grained-classification.pdf); at inference time, we averaged the 3 classifiers' logits (local classifier, ArcFace and BNNeck).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1134433%2Fb488481298423a5b4b3ecdbf549b1ad2%2Fmodel.png?generation=1590638315354457&amp;alt=media)\n\n**Training**\n\nWe employed a two-stage training which is similar to 4th ranked team. In the first stage, only samples from the top 5000 most frequent categories were used. After obtaining a good model (measured in terms of validation f1-score), we employed the weight imprinting trick to initialize the classifiers' weights for the remaining 27093 categories then re-trained the entire network until convergence. Below are our training configs.\n+ Stage 1: 75 epochs, batch size = 2048 (gradient accumulation), cosine scheduler with linear warmup and AdamW optimizer, balanced sampling (randomly sample 75 images per category per epoch).\n+ Stage 2: Increase number of epochs to 105, balanced sampling (randomly sample 25 images per category per epoch).\n+ Augmentations: `ResizeShorterSide(random.randint(img_size + 32, img_size + 128)),  RandomCrop(img_size),  RandAugment(num_ops=2, magnitude=random.randint(1, 11), \n RandomErasing()`\n\n**Scores**\nWe reported some single model scores without any post-processing here\n+ Inception V4 - Public 0.79712 - Private 0.77781\n+ B3 - Public 0.80570 - Private 0.78182\n+ B4 - Public 0.83014 - Private 0.81488\n+ B5 - Public 0.84774 - Private 0.83309\n\n**Ensemble** \n We took the code and idea in this[ notebook](https://www.kaggle.com/mathormad/5th-place-solution-stacking-pipeline) from last year RSNA challenge, modified them to build a simple 2-layer stacking network to combine predictions from our single models. The stacking ensemble obtained public score 0.85505 and private score 0.84180.\n```\nKERNEL_SIZE = 9\nNUM_CHANNELS = 128\n\nclass StackingCNN(nn.Module):\n    def __init__(self, num_models, num_channels=NUM_CHANNELS):\n        super(StackingCNN, self).__init__()\n        self.base = nn.Sequential(OrderedDict([\n            ('conv1', nn.Conv2d(1, num_channels,\n                kernel_size=(num_models, 1),\n                padding=(int((num_models - 1) / 2), 0))),\n            ('relu1', nn.ReLU(inplace=True)),\n            ('conv2', nn.Conv2d(num_channels, num_channels,\n                                kernel_size=(1, KERNEL_SIZE), padding=(0, int((KERNEL_SIZE - 1) / 2)))),\n            ('relu2', nn.ReLU(inplace=True))\n        ]))\n\n    def forward(self, images, labels=None):\n        with autocast():\n            x = self.base(images)\n            x = x.permute(0, 3, 1, 2)\n            x = fast_global_avg_pool_2d(x)\n            if self.training:\n                losses = {\n                    \"avg_loss\": F.cross_entropy(x, labels)\n                    }\n                return losses\n            else:\n                outputs = {\"logits\": x}\n                return outputs\n```\n\n**Postprocessing**\nWe exploited the fact that the test set has the number of examples per species capped at a maximum of 10. For each species, we first sorted the predictions based on their top 1 confidence scores. For predictions with confidence scores outside top 10, we changed them to top 2 category instead. This simple trick gave pretty consistent boost (0.2 - 0.3 %) on public LB for every single submission file. Our final submission scored 0.85797 on public and 0.84484 on private.",
      "votes": null
    },
    {
      "id": "864682",
      "postDate": "05/28/2020 06:09:00",
      "content": "<p>great solution, good job !\nwhen i compare your solution to ours (2nd place), i think what gave you the extra edge to win (0.15%) is the cool ensambling technique, while we did a simple averaging.</p>\n\n<p>how do you train the \"StackingCNN\" ? do you keep aside a dedicated validation set just for that ?</p>",
      "rawMarkdown": "great solution, good job !\nwhen i compare your solution to ours (2nd place), i think what gave you the extra edge to win (0.15%) is the cool ensambling technique, while we did a simple averaging.\n\nhow do you train the \"StackingCNN\" ? do you keep aside a dedicated validation set just for that ?",
      "votes": null
    },
    {
      "id": "864707",
      "postDate": "05/28/2020 06:33:31",
      "content": "<p>thanks for your words 👍. Your team's TResNet was a very exciting architecture !</p>\n\n<p>I definitely agreed that the stacking cnn was our extra edge to win in this comp.\nI constructed a mini-train set (about 80k images) with a similar distribution to the test set (all 32k categories, no more than 10 samples per cat.) to train the stacking cnn. Validation set was the same one as what we used to checkpoint our single models.</p>",
      "rawMarkdown": "thanks for your words 👍. Your team's TResNet was a very exciting architecture !\n\nI definitely agreed that the stacking cnn was our extra edge to win in this comp.\nI constructed a mini-train set (about 80k images) with a similar distribution to the test set (all 32k categories, no more than 10 samples per cat.) to train the stacking cnn. Validation set was the same one as what we used to checkpoint our single models.",
      "votes": null
    },
    {
      "id": "864814",
      "postDate": "05/28/2020 07:48:43",
      "content": "<p>great job.\nI am curios to know if you have done some ablation to see how much each of the tricks adds to the baseline (i assume you had baseline). Namely:\n- local features\n- ArchFace\n- BNNeck\n- weights imprinting</p>\n\n<p>Also, correct me if i'm wrong your epoch in stage 1 is 75x32093=~2.4M. What;s the hardware used and how long it takes to converge?\nI must say i didn't expect 100+epochs but it seems both top teams had very long training.</p>",
      "rawMarkdown": "great job.\nI am curios to know if you have done some ablation to see how much each of the tricks adds to the baseline (i assume you had baseline). Namely:\n- local features\n- ArchFace\n- BNNeck\n- weights imprinting\n\nAlso, correct me if i'm wrong your epoch in stage 1 is 75x32093=~2.4M. What;s the hardware used and how long it takes to converge?\nI must say i didn't expect 100+epochs but it seems both top teams had very long training.",
      "votes": null
    },
    {
      "id": "864868",
      "postDate": "05/28/2020 08:32:31",
      "content": "<p>We tried all tricks with a ResNet50 to get quick feedback. Our simplest baseline was just a ResNet50 with ArcFace head. Adding other things on top of ArcFace don't improve the performance of ArcFace head per se; however, we now have extra classifiers' logits to ensemble 😄. I would say that our single model isn't \"single\" in the purest sense 😄.\nIn stage 1, we have 75x5000 = 375000 samples per epoch. In stage 2, we have 25 x 32093 = 802k samples per epoch. We train all models for a total of 75+105=180 epochs.\nWe have a box of 4 2080ti for prototyping + a dgx1 for longer training.</p>",
      "rawMarkdown": "We tried all tricks with a ResNet50 to get quick feedback. Our simplest baseline was just a ResNet50 with ArcFace head. Adding other things on top of ArcFace don't improve the performance of ArcFace head per se; however, we now have extra classifiers' logits to ensemble 😄. I would say that our single model isn't \"single\" in the purest sense 😄.\nIn stage 1, we have 75x5000 = 375000 samples per epoch. In stage 2, we have 25 x 32093 = 802k samples per epoch. We train all models for a total of 75+105=180 epochs.\nWe have a box of 4 2080ti for prototyping + a dgx1 for longer training.",
      "votes": null
    },
    {
      "id": "864890",
      "postDate": "05/28/2020 08:49:44",
      "content": "<p>thanks for the quick reply. I wander why you choose ArcFace and not standard softmax for your baseline !? Is that it just works better for few shot tasks. A lot to learn for me  obviously :)</p>",
      "rawMarkdown": "thanks for the quick reply. I wander why you choose ArcFace and not standard softmax for your baseline !? Is that it just works better for few shot tasks. A lot to learn for me  obviously :)",
      "votes": null
    },
    {
      "id": "866612",
      "postDate": "05/29/2020 14:42:54",
      "content": "<p>Hi <a href=\"/andy2709\">@andy2709</a> \ni am trying to better understand the stacking CNN.\nwhat you stack is the feature of each CNN (~2048), or the logits (~32,000) ?</p>\n\n<p>why \"NUM_CHANNELS = 128\" ?</p>",
      "rawMarkdown": "Hi @andy2709 \ni am trying to better understand the stacking CNN.\nwhat you stack is the feature of each CNN (~2048), or the logits (~32,000) ?\n\nwhy \"NUM_CHANNELS = 128\" ?",
      "votes": null
    },
    {
      "id": "866699",
      "postDate": "05/29/2020 16:01:25",
      "content": "<p>Hey <a href=\"/mrt23564\">@mrt23564</a>, \nI used the models logits as input to the stacking net. The input shape will be [batch size, 1, num models, 32093]. In my case, num models = 5.\nI tried setting the num_channels to 32 and 64 first and found out that the stacking net couldn't overfit the mini-train set probably due to it having too few parameters. Increasing to 128, 256 or something larger solved this issue. In the end, I went with 128.</p>",
      "rawMarkdown": "Hey @mrt23564, \nI used the models logits as input to the stacking net. The input shape will be [batch size, 1, num models, 32093]. In my case, num models = 5.\nI tried setting the num_channels to 32 and 64 first and found out that the stacking net couldn't overfit the mini-train set probably due to it having too few parameters. Increasing to 128, 256 or something larger solved this issue. In the end, I went with 128.",
      "votes": null
    },
    {
      "id": "866734",
      "postDate": "05/29/2020 16:27:07",
      "content": "<p>got it, thanks. 👍 </p>",
      "rawMarkdown": "got it, thanks. 👍",
      "votes": null
    },
    {
      "id": "867035",
      "postDate": "05/29/2020 23:24:52",
      "content": "<p>If your intent is using a \"single model\" for this task without needing extra heads or classifiers, then you can try adding triplet loss (or preferably soft-triplet loss) that works directly on the embedding vector \"head\"  in addition to CE loss (or maybe ArcFace loss) can improve your result on this task (metric learning has shown to be crucial when you have small intra-class variation in your dataset), without extra parameters that the ArcFace requires.\nAbout BNNeck, I experimented with BNNeck in my Person-ReID research, it provides relatively small improvement yet an important one, and from my experience, I found that you can achieve the same effect simply by disabling the bias in the \"main\" classification layer.\nBNNeck or disabling FC bias parameter works only if you combine it with either:\n1. Triplet loss on the embedding (\"head\") (after L2 normalization of the features) \n2. ArcFace head (as a secondary classifier that branches from the head), similar to how it was done here. </p>",
      "rawMarkdown": "If your intent is using a \"single model\" for this task without needing extra heads or classifiers, then you can try adding triplet loss (or preferably soft-triplet loss) that works directly on the embedding vector \"head\"  in addition to CE loss (or maybe ArcFace loss) can improve your result on this task (metric learning has shown to be crucial when you have small intra-class variation in your dataset), without extra parameters that the ArcFace requires.\nAbout BNNeck, I experimented with BNNeck in my Person-ReID research, it provides relatively small improvement yet an important one, and from my experience, I found that you can achieve the same effect simply by disabling the bias in the \"main\" classification layer.\nBNNeck or disabling FC bias parameter works only if you combine it with either:\n1. Triplet loss on the embedding (\"head\") (after L2 normalization of the features) \n2. ArcFace head (as a secondary classifier that branches from the head), similar to how it was done here.",
      "votes": null
    },
    {
      "id": "867508",
      "postDate": "05/30/2020 11:27:49",
      "content": "<p>Hey <a href=\"/andy2709\">@andy2709</a> congratulations on your 1st place.\nCorrect me if I'm wrong but having 3 \"heavy\" classifiers of 32k in a single model for sure eats a lot of gpu memory, forcing you to reduce your \"actual\" batch for each step (therefore needing gradient accumulation to compensate for that).\n1. Did you use the same \"batch size\" of 2048 in stage 1 and stage 2?\n2. I wonder how much time per epoch it took for B5 during stage 2?</p>\n\n<p>other than that I think StackingCNN and OSME are nice :)</p>",
      "rawMarkdown": "Hey @andy2709 congratulations on your 1st place.\nCorrect me if I'm wrong but having 3 \"heavy\" classifiers of 32k in a single model for sure eats a lot of gpu memory, forcing you to reduce your \"actual\" batch for each step (therefore needing gradient accumulation to compensate for that).\n1. Did you use the same \"batch size\" of 2048 in stage 1 and stage 2?\n2. I wonder how much time per epoch it took for B5 during stage 2?\n\nother than that I think StackingCNN and OSME are nice :)",
      "votes": null
    },
    {
      "id": "867973",
      "postDate": "05/30/2020 19:33:47",
      "content": "<p>HI, <a href=\"/andy2709\">@andy2709</a> </p>\n\n<p>Would you mind to share how do you validate your model? Which validation set did you sample?</p>",
      "rawMarkdown": "HI, @andy2709 \n\nWould you mind to share how do you validate your model? Which validation set did you sample?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 864682,
      "author_name": "mrt23564",
      "author_url": "",
      "post_date": "05/28/2020 06:09:00",
      "content": "<p>great solution, good job !\nwhen i compare your solution to ours (2nd place), i think what gave you the extra edge to win (0.15%) is the cool ensambling technique, while we did a simple averaging.</p>\n\n<p>how do you train the \"StackingCNN\" ? do you keep aside a dedicated validation set just for that ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 864707,
          "author_name": "andy2709",
          "author_url": "",
          "post_date": "05/28/2020 06:33:31",
          "content": "<p>thanks for your words 👍. Your team's TResNet was a very exciting architecture !</p>\n\n<p>I definitely agreed that the stacking cnn was our extra edge to win in this comp.\nI constructed a mini-train set (about 80k images) with a similar distribution to the test set (all 32k categories, no more than 10 samples per cat.) to train the stacking cnn. Validation set was the same one as what we used to checkpoint our single models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 866612,
          "author_name": "mrt23564",
          "author_url": "",
          "post_date": "05/29/2020 14:42:54",
          "content": "<p>Hi <a href=\"/andy2709\">@andy2709</a> \ni am trying to better understand the stacking CNN.\nwhat you stack is the feature of each CNN (~2048), or the logits (~32,000) ?</p>\n\n<p>why \"NUM_CHANNELS = 128\" ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 866699,
          "author_name": "andy2709",
          "author_url": "",
          "post_date": "05/29/2020 16:01:25",
          "content": "<p>Hey <a href=\"/mrt23564\">@mrt23564</a>, \nI used the models logits as input to the stacking net. The input shape will be [batch size, 1, num models, 32093]. In my case, num models = 5.\nI tried setting the num_channels to 32 and 64 first and found out that the stacking net couldn't overfit the mini-train set probably due to it having too few parameters. Increasing to 128, 256 or something larger solved this issue. In the end, I went with 128.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 866734,
          "author_name": "mrt23564",
          "author_url": "",
          "post_date": "05/29/2020 16:27:07",
          "content": "<p>got it, thanks. 👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 864814,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "05/28/2020 07:48:43",
      "content": "<p>great job.\nI am curios to know if you have done some ablation to see how much each of the tricks adds to the baseline (i assume you had baseline). Namely:\n- local features\n- ArchFace\n- BNNeck\n- weights imprinting</p>\n\n<p>Also, correct me if i'm wrong your epoch in stage 1 is 75x32093=~2.4M. What;s the hardware used and how long it takes to converge?\nI must say i didn't expect 100+epochs but it seems both top teams had very long training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 864868,
          "author_name": "andy2709",
          "author_url": "",
          "post_date": "05/28/2020 08:32:31",
          "content": "<p>We tried all tricks with a ResNet50 to get quick feedback. Our simplest baseline was just a ResNet50 with ArcFace head. Adding other things on top of ArcFace don't improve the performance of ArcFace head per se; however, we now have extra classifiers' logits to ensemble 😄. I would say that our single model isn't \"single\" in the purest sense 😄.\nIn stage 1, we have 75x5000 = 375000 samples per epoch. In stage 2, we have 25 x 32093 = 802k samples per epoch. We train all models for a total of 75+105=180 epochs.\nWe have a box of 4 2080ti for prototyping + a dgx1 for longer training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 864890,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "05/28/2020 08:49:44",
          "content": "<p>thanks for the quick reply. I wander why you choose ArcFace and not standard softmax for your baseline !? Is that it just works better for few shot tasks. A lot to learn for me  obviously :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 867035,
          "author_name": "hussam789",
          "author_url": "",
          "post_date": "05/29/2020 23:24:52",
          "content": "<p>If your intent is using a \"single model\" for this task without needing extra heads or classifiers, then you can try adding triplet loss (or preferably soft-triplet loss) that works directly on the embedding vector \"head\"  in addition to CE loss (or maybe ArcFace loss) can improve your result on this task (metric learning has shown to be crucial when you have small intra-class variation in your dataset), without extra parameters that the ArcFace requires.\nAbout BNNeck, I experimented with BNNeck in my Person-ReID research, it provides relatively small improvement yet an important one, and from my experience, I found that you can achieve the same effect simply by disabling the bias in the \"main\" classification layer.\nBNNeck or disabling FC bias parameter works only if you combine it with either:\n1. Triplet loss on the embedding (\"head\") (after L2 normalization of the features) \n2. ArcFace head (as a secondary classifier that branches from the head), similar to how it was done here. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 867508,
      "author_name": "hussam789",
      "author_url": "",
      "post_date": "05/30/2020 11:27:49",
      "content": "<p>Hey <a href=\"/andy2709\">@andy2709</a> congratulations on your 1st place.\nCorrect me if I'm wrong but having 3 \"heavy\" classifiers of 32k in a single model for sure eats a lot of gpu memory, forcing you to reduce your \"actual\" batch for each step (therefore needing gradient accumulation to compensate for that).\n1. Did you use the same \"batch size\" of 2048 in stage 1 and stage 2?\n2. I wonder how much time per epoch it took for B5 during stage 2?</p>\n\n<p>other than that I think StackingCNN and OSME are nice :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 867973,
      "author_name": "discoholic",
      "author_url": "",
      "post_date": "05/30/2020 19:33:47",
      "content": "<p>HI, <a href=\"/andy2709\">@andy2709</a> </p>\n\n<p>Would you mind to share how do you validate your model? Which validation set did you sample?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864601": "Firstly, I would like to thank the host team and Kaggle for hosting such an interesting competition as well as my teammates (@dattran2346, @moewie94) for their hard work and contributions.\n\n**Model Architecture**\n\nWe used EfficientNet B3/4/5 (noisy student version) with original image sizes (320x320, 380x380, 456x456), Inception V4 and SEResNeXt50 with image size 448x448 as our backbones. \nFor the EfficientNets, we first extracted feature maps from block 4, 5, and 6, forwarded through  one-squeeze multi-excitation modules ([OSME](https://eccv2018.org/openaccess/content_ECCV_2018/papers/Ming_Sun_Multi-Attention_Multi-Class_Constraint_ECCV_2018_paper.pdf)), before pooling and concatenating them into a \"local\" features vector. The ArcFace(s=16, m=0.1) loss + BNNeck combination was then used with \"head\" block features as input to help our models learn better \"global\" representation. \nThis pipeline was trained jointly end-to-end with Cross Entropy losses (no label smoothing) and [entropy losses](https://papers.nips.cc/paper/7344-maximum-entropy-fine-grained-classification.pdf); at inference time, we averaged the 3 classifiers' logits (local classifier, ArcFace and BNNeck).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1134433%2Fb488481298423a5b4b3ecdbf549b1ad2%2Fmodel.png?generation=1590638315354457&amp;alt=media)\n\n**Training**\n\nWe employed a two-stage training which is similar to 4th ranked team. In the first stage, only samples from the top 5000 most frequent categories were used. After obtaining a good model (measured in terms of validation f1-score), we employed the weight imprinting trick to initialize the classifiers' weights for the remaining 27093 categories then re-trained the entire network until convergence. Below are our training configs.\n+ Stage 1: 75 epochs, batch size = 2048 (gradient accumulation), cosine scheduler with linear warmup and AdamW optimizer, balanced sampling (randomly sample 75 images per category per epoch).\n+ Stage 2: Increase number of epochs to 105, balanced sampling (randomly sample 25 images per category per epoch).\n+ Augmentations: `ResizeShorterSide(random.randint(img_size + 32, img_size + 128)),  RandomCrop(img_size),  RandAugment(num_ops=2, magnitude=random.randint(1, 11), \n RandomErasing()`\n\n**Scores**\nWe reported some single model scores without any post-processing here\n+ Inception V4 - Public 0.79712 - Private 0.77781\n+ B3 - Public 0.80570 - Private 0.78182\n+ B4 - Public 0.83014 - Private 0.81488\n+ B5 - Public 0.84774 - Private 0.83309\n\n**Ensemble** \n We took the code and idea in this[ notebook](https://www.kaggle.com/mathormad/5th-place-solution-stacking-pipeline) from last year RSNA challenge, modified them to build a simple 2-layer stacking network to combine predictions from our single models. The stacking ensemble obtained public score 0.85505 and private score 0.84180.\n```\nKERNEL_SIZE = 9\nNUM_CHANNELS = 128\n\nclass StackingCNN(nn.Module):\n    def __init__(self, num_models, num_channels=NUM_CHANNELS):\n        super(StackingCNN, self).__init__()\n        self.base = nn.Sequential(OrderedDict([\n            ('conv1', nn.Conv2d(1, num_channels,\n                kernel_size=(num_models, 1),\n                padding=(int((num_models - 1) / 2), 0))),\n            ('relu1', nn.ReLU(inplace=True)),\n            ('conv2', nn.Conv2d(num_channels, num_channels,\n                                kernel_size=(1, KERNEL_SIZE), padding=(0, int((KERNEL_SIZE - 1) / 2)))),\n            ('relu2', nn.ReLU(inplace=True))\n        ]))\n\n    def forward(self, images, labels=None):\n        with autocast():\n            x = self.base(images)\n            x = x.permute(0, 3, 1, 2)\n            x = fast_global_avg_pool_2d(x)\n            if self.training:\n                losses = {\n                    \"avg_loss\": F.cross_entropy(x, labels)\n                    }\n                return losses\n            else:\n                outputs = {\"logits\": x}\n                return outputs\n```\n\n**Postprocessing**\nWe exploited the fact that the test set has the number of examples per species capped at a maximum of 10. For each species, we first sorted the predictions based on their top 1 confidence scores. For predictions with confidence scores outside top 10, we changed them to top 2 category instead. This simple trick gave pretty consistent boost (0.2 - 0.3 %) on public LB for every single submission file. Our final submission scored 0.85797 on public and 0.84484 on private.",
    "864682": "great solution, good job !\nwhen i compare your solution to ours (2nd place), i think what gave you the extra edge to win (0.15%) is the cool ensambling technique, while we did a simple averaging.\n\nhow do you train the \"StackingCNN\" ? do you keep aside a dedicated validation set just for that ?",
    "864707": "thanks for your words 👍. Your team's TResNet was a very exciting architecture !\n\nI definitely agreed that the stacking cnn was our extra edge to win in this comp.\nI constructed a mini-train set (about 80k images) with a similar distribution to the test set (all 32k categories, no more than 10 samples per cat.) to train the stacking cnn. Validation set was the same one as what we used to checkpoint our single models.",
    "864814": "great job.\nI am curios to know if you have done some ablation to see how much each of the tricks adds to the baseline (i assume you had baseline). Namely:\n- local features\n- ArchFace\n- BNNeck\n- weights imprinting\n\nAlso, correct me if i'm wrong your epoch in stage 1 is 75x32093=~2.4M. What;s the hardware used and how long it takes to converge?\nI must say i didn't expect 100+epochs but it seems both top teams had very long training.",
    "864868": "We tried all tricks with a ResNet50 to get quick feedback. Our simplest baseline was just a ResNet50 with ArcFace head. Adding other things on top of ArcFace don't improve the performance of ArcFace head per se; however, we now have extra classifiers' logits to ensemble 😄. I would say that our single model isn't \"single\" in the purest sense 😄.\nIn stage 1, we have 75x5000 = 375000 samples per epoch. In stage 2, we have 25 x 32093 = 802k samples per epoch. We train all models for a total of 75+105=180 epochs.\nWe have a box of 4 2080ti for prototyping + a dgx1 for longer training.",
    "864890": "thanks for the quick reply. I wander why you choose ArcFace and not standard softmax for your baseline !? Is that it just works better for few shot tasks. A lot to learn for me  obviously :)",
    "866612": "Hi @andy2709 \ni am trying to better understand the stacking CNN.\nwhat you stack is the feature of each CNN (~2048), or the logits (~32,000) ?\n\nwhy \"NUM_CHANNELS = 128\" ?",
    "866699": "Hey @mrt23564, \nI used the models logits as input to the stacking net. The input shape will be [batch size, 1, num models, 32093]. In my case, num models = 5.\nI tried setting the num_channels to 32 and 64 first and found out that the stacking net couldn't overfit the mini-train set probably due to it having too few parameters. Increasing to 128, 256 or something larger solved this issue. In the end, I went with 128.",
    "866734": "got it, thanks. 👍",
    "867035": "If your intent is using a \"single model\" for this task without needing extra heads or classifiers, then you can try adding triplet loss (or preferably soft-triplet loss) that works directly on the embedding vector \"head\"  in addition to CE loss (or maybe ArcFace loss) can improve your result on this task (metric learning has shown to be crucial when you have small intra-class variation in your dataset), without extra parameters that the ArcFace requires.\nAbout BNNeck, I experimented with BNNeck in my Person-ReID research, it provides relatively small improvement yet an important one, and from my experience, I found that you can achieve the same effect simply by disabling the bias in the \"main\" classification layer.\nBNNeck or disabling FC bias parameter works only if you combine it with either:\n1. Triplet loss on the embedding (\"head\") (after L2 normalization of the features) \n2. ArcFace head (as a secondary classifier that branches from the head), similar to how it was done here.",
    "867508": "Hey @andy2709 congratulations on your 1st place.\nCorrect me if I'm wrong but having 3 \"heavy\" classifiers of 32k in a single model for sure eats a lot of gpu memory, forcing you to reduce your \"actual\" batch for each step (therefore needing gradient accumulation to compensate for that).\n1. Did you use the same \"batch size\" of 2048 in stage 1 and stage 2?\n2. I wonder how much time per epoch it took for B5 during stage 2?\n\nother than that I think StackingCNN and OSME are nice :)",
    "867973": "HI, @andy2709 \n\nWould you mind to share how do you validate your model? Which validation set did you sample?"
  },
  "source": "meta"
}