{
  "id": 137353,
  "title": "Potential 6th place solution with a single model ! [all about layers]",
  "url": "/competitions/bengaliai-cv19/discussion/137353",
  "author_name": "Satwik",
  "post_date": "2020-03-20T10:16:15.965000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello! After quite a rough LB shakeup , we ended being 317th in the private LB from 79th in public, however , scrolling down through our submissions , we did find one low scoring public kernel , which could have given us ~70th place on private with an LB of 9356. However , on using post-processing as Chris mentioned <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">here</a> , \nwe managed to hit a private LB score of 9557 , with public LB 9703! Which would potentially give us 6th place as the private leaderboard stands right now.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3459992%2F95ceecb84fb83408fa0a27fc8dad1c60%2Foof.PNG?generation=1584697688199528&amp;alt=media\" alt=\"\"></p>\n\n<p>This was really unexpected for us , because it was a really low score for the public LB and we had been stuck there for about a month.  </p>\n\n<h2>The Solution -</h2>\n\n<p>Briefly , the solution is really simple. \n- We use a single model , a SEResNeXt50 from pretrainedmodels, with 3 different output heads , one for each class, as used by iafoss in <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">this</a> notebook. \n- Data used was 128x128 images , cropped using <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">this</a>  notebook.\n- For training , we used a basic Shift , Scale , Rotate , and Mixup only. Training was done for ONLY 30 epochs over a OneCycle LR policy , and Over9000 optimizer.</p>\n\n<p>With this , we get a public LB of ~966 , and private LB ~92. Now, this is where the magic comes.\nWith the following changes , we make a straight jump to a private LB of 9557, and public LB 9703, which is the main part of our solution.</p>\n\n<ul>\n<li>Change every AdaptiveAvgPool2D layer to GeM layer</li>\n<li>Change every AdaptiveConcatPool2D layer to GeM layer</li>\n<li>Removed Dropout and Weight Decay (refer to <a href=\"https://arxiv.org/abs/1912.11370\">this paper</a> )</li>\n<li>Used Group Normalization instead of Batch Normalization</li>\n<li>Added weight standardization to all vanilla Conv2D layers</li>\n</ul>\n\n<p>Codes for the same:-</p>\n\nGeneralized Mean Pooling (GeM) -\n\n<p><code>def gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)</code></p>\n\n<p><code>class GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'</code></p>\n\nCon2D with Weight Standardization -\n\n<p>`class Conv2d(nn.Conv2d):</p>\n\n<pre><code>def __init__(self, in_channels, out_channels, kernel_size, stride=1,\n             padding=0, dilation=1, groups=1, bias=True):\n    super(Conv2d, self).__init__(in_channels, out_channels, kernel_size, stride,\n             padding, dilation, groups, bias)\n\ndef forward(self, x):\n    weight = self.weight\n    weight_mean = weight.mean(dim=1, keepdim=True).mean(dim=2,\n                              keepdim=True).mean(dim=3, keepdim=True)\n    weight = weight - weight_mean\n    std = weight.view(weight.size(0), -1).std(dim=1).view(-1, 1, 1, 1) + 1e-5\n    weight = weight / std.expand_as(weight)\n    return F.conv2d(x, weight, self.bias, self.stride,\n                    self.padding, self.dilation, self.groups)`\n</code></pre>\n\nCodes to convert the model -\n\n<p><code>def convert_to_gem(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.AdaptiveAvgPool2d):\n            setattr(model, child_name, GeM())\n        else:\n            convert_to_gem(child)</code></p>\n\n<p><code>def convert_to_conv2d(model):\n    for child_name, child in model.named_children():\n        if child_name not in ['fc1','fc2']:\n            if isinstance(child, nn.Conv2d):\n                in_feat = child.in_channels\n                out_feat = child.out_channels\n                ker_size = child.kernel_size\n                stride = child.stride\n                padding = child.padding\n                dilation = child.dilation\n                groups = child.groups\n                setattr(model, child_name, Conv2d(in_channels=in_feat, out_channels=out_feat, kernel_size=ker_size, stride=stride,\n                                                 padding = padding, dilation=dilation, groups=groups))\n            else:\n                convert_to_conv2d(child)</code></p>\n\n<p><code>def convert_to_groupnorm(model):\n    for child_name, child in model.named_children():\n            if isinstance(child, nn.BatchNorm2d):\n                num_features = child.num_features\n                setattr(model, child_name, GroupNorm(num_groups=32, num_channels=num_features))\n            else:\n                convert_to_groupnorm(child)</code></p>\n\n<p>If you have any ideas , or suggestions please let us know! We also have a feeling using non-cropped images , and some other augmentations would have given us even better results. \nThanks for reading! </p>\n\n<p>EDIT - \nHere are the notebooks for this -</p>\n\n<p>Inference - <a href=\"https://www.kaggle.com/p4rallax/private-0-9557/\">https://www.kaggle.com/p4rallax/private-0-9557/</a></p>\n\n<p>Training - <a href=\"https://www.kaggle.com/virajbagal/potential-6th-place-lb-magic-layers-pp\">https://www.kaggle.com/virajbagal/potential-6th-place-lb-magic-layers-pp</a></p>",
  "messages": [
    {
      "id": 780491,
      "postDate": "2020-03-20T10:16:15.967Z",
      "content": "<p>Hello! After quite a rough LB shakeup , we ended being 317th in the private LB from 79th in public, however , scrolling down through our submissions , we did find one low scoring public kernel , which could have given us ~70th place on private with an LB of 9356. However , on using post-processing as Chris mentioned <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">here</a> , \nwe managed to hit a private LB score of 9557 , with public LB 9703! Which would potentially give us 6th place as the private leaderboard stands right now.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3459992%2F95ceecb84fb83408fa0a27fc8dad1c60%2Foof.PNG?generation=1584697688199528&amp;alt=media\" alt=\"\"></p>\n\n<p>This was really unexpected for us , because it was a really low score for the public LB and we had been stuck there for about a month.  </p>\n\n<h2>The Solution -</h2>\n\n<p>Briefly , the solution is really simple. \n- We use a single model , a SEResNeXt50 from pretrainedmodels, with 3 different output heads , one for each class, as used by iafoss in <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">this</a> notebook. \n- Data used was 128x128 images , cropped using <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">this</a>  notebook.\n- For training , we used a basic Shift , Scale , Rotate , and Mixup only. Training was done for ONLY 30 epochs over a OneCycle LR policy , and Over9000 optimizer.</p>\n\n<p>With this , we get a public LB of ~966 , and private LB ~92. Now, this is where the magic comes.\nWith the following changes , we make a straight jump to a private LB of 9557, and public LB 9703, which is the main part of our solution.</p>\n\n<ul>\n<li>Change every AdaptiveAvgPool2D layer to GeM layer</li>\n<li>Change every AdaptiveConcatPool2D layer to GeM layer</li>\n<li>Removed Dropout and Weight Decay (refer to <a href=\"https://arxiv.org/abs/1912.11370\">this paper</a> )</li>\n<li>Used Group Normalization instead of Batch Normalization</li>\n<li>Added weight standardization to all vanilla Conv2D layers</li>\n</ul>\n\n<p>Codes for the same:-</p>\n\nGeneralized Mean Pooling (GeM) -\n\n<p><code>def gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)</code></p>\n\n<p><code>class GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'</code></p>\n\nCon2D with Weight Standardization -\n\n<p>`class Conv2d(nn.Conv2d):</p>\n\n<pre><code>def __init__(self, in_channels, out_channels, kernel_size, stride=1,\n             padding=0, dilation=1, groups=1, bias=True):\n    super(Conv2d, self).__init__(in_channels, out_channels, kernel_size, stride,\n             padding, dilation, groups, bias)\n\ndef forward(self, x):\n    weight = self.weight\n    weight_mean = weight.mean(dim=1, keepdim=True).mean(dim=2,\n                              keepdim=True).mean(dim=3, keepdim=True)\n    weight = weight - weight_mean\n    std = weight.view(weight.size(0), -1).std(dim=1).view(-1, 1, 1, 1) + 1e-5\n    weight = weight / std.expand_as(weight)\n    return F.conv2d(x, weight, self.bias, self.stride,\n                    self.padding, self.dilation, self.groups)`\n</code></pre>\n\nCodes to convert the model -\n\n<p><code>def convert_to_gem(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.AdaptiveAvgPool2d):\n            setattr(model, child_name, GeM())\n        else:\n            convert_to_gem(child)</code></p>\n\n<p><code>def convert_to_conv2d(model):\n    for child_name, child in model.named_children():\n        if child_name not in ['fc1','fc2']:\n            if isinstance(child, nn.Conv2d):\n                in_feat = child.in_channels\n                out_feat = child.out_channels\n                ker_size = child.kernel_size\n                stride = child.stride\n                padding = child.padding\n                dilation = child.dilation\n                groups = child.groups\n                setattr(model, child_name, Conv2d(in_channels=in_feat, out_channels=out_feat, kernel_size=ker_size, stride=stride,\n                                                 padding = padding, dilation=dilation, groups=groups))\n            else:\n                convert_to_conv2d(child)</code></p>\n\n<p><code>def convert_to_groupnorm(model):\n    for child_name, child in model.named_children():\n            if isinstance(child, nn.BatchNorm2d):\n                num_features = child.num_features\n                setattr(model, child_name, GroupNorm(num_groups=32, num_channels=num_features))\n            else:\n                convert_to_groupnorm(child)</code></p>\n\n<p>If you have any ideas , or suggestions please let us know! We also have a feeling using non-cropped images , and some other augmentations would have given us even better results. \nThanks for reading! </p>\n\n<p>EDIT - \nHere are the notebooks for this -</p>\n\n<p>Inference - <a href=\"https://www.kaggle.com/p4rallax/private-0-9557/\">https://www.kaggle.com/p4rallax/private-0-9557/</a></p>\n\n<p>Training - <a href=\"https://www.kaggle.com/virajbagal/potential-6th-place-lb-magic-layers-pp\">https://www.kaggle.com/virajbagal/potential-6th-place-lb-magic-layers-pp</a></p>",
      "rawMarkdown": "Hello! After quite a rough LB shakeup , we ended being 317th in the private LB from 79th in public, however , scrolling down through our submissions , we did find one low scoring public kernel , which could have given us ~70th place on private with an LB of 9356. However , on using post-processing as Chris mentioned [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021) , \nwe managed to hit a private LB score of 9557 , with public LB 9703! Which would potentially give us 6th place as the private leaderboard stands right now.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3459992%2F95ceecb84fb83408fa0a27fc8dad1c60%2Foof.PNG?generation=1584697688199528&amp;alt=media)\n\nThis was really unexpected for us , because it was a really low score for the public LB and we had been stuck there for about a month.  \n\n## The Solution -\n\nBriefly , the solution is really simple. \n- We use a single model , a SEResNeXt50 from pretrainedmodels, with 3 different output heads , one for each class, as used by iafoss in [this](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964) notebook. \n- Data used was 128x128 images , cropped using [this](https://www.kaggle.com/iafoss/image-preprocessing-128x128)  notebook.\n- For training , we used a basic Shift , Scale , Rotate , and Mixup only. Training was done for ONLY 30 epochs over a OneCycle LR policy , and Over9000 optimizer.\n\nWith this , we get a public LB of ~966 , and private LB ~92. Now, this is where the magic comes.\nWith the following changes , we make a straight jump to a private LB of 9557, and public LB 9703, which is the main part of our solution.\n\n- Change every AdaptiveAvgPool2D layer to GeM layer\n- Change every AdaptiveConcatPool2D layer to GeM layer\n- Removed Dropout and Weight Decay (refer to [this paper](https://arxiv.org/abs/1912.11370) )\n- Used Group Normalization instead of Batch Normalization\n- Added weight standardization to all vanilla Conv2D layers\n\nCodes for the same:-\n\n###### Generalized Mean Pooling (GeM) -\n\n`def gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)`\n\n`class GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'`\n\n\n###### Con2D with Weight Standardization -\n\n`class Conv2d(nn.Conv2d):\n\n    def __init__(self, in_channels, out_channels, kernel_size, stride=1,\n                 padding=0, dilation=1, groups=1, bias=True):\n        super(Conv2d, self).__init__(in_channels, out_channels, kernel_size, stride,\n                 padding, dilation, groups, bias)\n\n    def forward(self, x):\n        weight = self.weight\n        weight_mean = weight.mean(dim=1, keepdim=True).mean(dim=2,\n                                  keepdim=True).mean(dim=3, keepdim=True)\n        weight = weight - weight_mean\n        std = weight.view(weight.size(0), -1).std(dim=1).view(-1, 1, 1, 1) + 1e-5\n        weight = weight / std.expand_as(weight)\n        return F.conv2d(x, weight, self.bias, self.stride,\n                        self.padding, self.dilation, self.groups)`\n\n\n###### Codes to convert the model -\n\n`def convert_to_gem(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.AdaptiveAvgPool2d):\n            setattr(model, child_name, GeM())\n        else:\n            convert_to_gem(child)`\n\n\n`def convert_to_conv2d(model):\n    for child_name, child in model.named_children():\n        if child_name not in ['fc1','fc2']:\n            if isinstance(child, nn.Conv2d):\n                in_feat = child.in_channels\n                out_feat = child.out_channels\n                ker_size = child.kernel_size\n                stride = child.stride\n                padding = child.padding\n                dilation = child.dilation\n                groups = child.groups\n                setattr(model, child_name, Conv2d(in_channels=in_feat, out_channels=out_feat, kernel_size=ker_size, stride=stride,\n                                                 padding = padding, dilation=dilation, groups=groups))\n            else:\n                convert_to_conv2d(child)`\n\n\n`def convert_to_groupnorm(model):\n    for child_name, child in model.named_children():\n            if isinstance(child, nn.BatchNorm2d):\n                num_features = child.num_features\n                setattr(model, child_name, GroupNorm(num_groups=32, num_channels=num_features))\n            else:\n                convert_to_groupnorm(child)`\n\n\nIf you have any ideas , or suggestions please let us know! We also have a feeling using non-cropped images , and some other augmentations would have given us even better results. \nThanks for reading! \n\n\n\n\nEDIT - \nHere are the notebooks for this -\n\nInference - https://www.kaggle.com/p4rallax/private-0-9557/\n\nTraining - https://www.kaggle.com/virajbagal/potential-6th-place-lb-magic-layers-pp\n\n\n\n\n\n",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "780491": "Hello! After quite a rough LB shakeup , we ended being 317th in the private LB from 79th in public, however , scrolling down through our submissions , we did find one low scoring public kernel , which could have given us ~70th place on private with an LB of 9356. However , on using post-processing as Chris mentioned [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/136021) , \nwe managed to hit a private LB score of 9557 , with public LB 9703! Which would potentially give us 6th place as the private leaderboard stands right now.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3459992%2F95ceecb84fb83408fa0a27fc8dad1c60%2Foof.PNG?generation=1584697688199528&amp;alt=media)\n\nThis was really unexpected for us , because it was a really low score for the public LB and we had been stuck there for about a month.  \n\n## The Solution -\n\nBriefly , the solution is really simple. \n- We use a single model , a SEResNeXt50 from pretrainedmodels, with 3 different output heads , one for each class, as used by iafoss in [this](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964) notebook. \n- Data used was 128x128 images , cropped using [this](https://www.kaggle.com/iafoss/image-preprocessing-128x128)  notebook.\n- For training , we used a basic Shift , Scale , Rotate , and Mixup only. Training was done for ONLY 30 epochs over a OneCycle LR policy , and Over9000 optimizer.\n\nWith this , we get a public LB of ~966 , and private LB ~92. Now, this is where the magic comes.\nWith the following changes , we make a straight jump to a private LB of 9557, and public LB 9703, which is the main part of our solution.\n\n- Change every AdaptiveAvgPool2D layer to GeM layer\n- Change every AdaptiveConcatPool2D layer to GeM layer\n- Removed Dropout and Weight Decay (refer to [this paper](https://arxiv.org/abs/1912.11370) )\n- Used Group Normalization instead of Batch Normalization\n- Added weight standardization to all vanilla Conv2D layers\n\nCodes for the same:-\n\n###### Generalized Mean Pooling (GeM) -\n\n`def gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)`\n\n`class GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'`\n\n\n###### Con2D with Weight Standardization -\n\n`class Conv2d(nn.Conv2d):\n\n    def __init__(self, in_channels, out_channels, kernel_size, stride=1,\n                 padding=0, dilation=1, groups=1, bias=True):\n        super(Conv2d, self).__init__(in_channels, out_channels, kernel_size, stride,\n                 padding, dilation, groups, bias)\n\n    def forward(self, x):\n        weight = self.weight\n        weight_mean = weight.mean(dim=1, keepdim=True).mean(dim=2,\n                                  keepdim=True).mean(dim=3, keepdim=True)\n        weight = weight - weight_mean\n        std = weight.view(weight.size(0), -1).std(dim=1).view(-1, 1, 1, 1) + 1e-5\n        weight = weight / std.expand_as(weight)\n        return F.conv2d(x, weight, self.bias, self.stride,\n                        self.padding, self.dilation, self.groups)`\n\n\n###### Codes to convert the model -\n\n`def convert_to_gem(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.AdaptiveAvgPool2d):\n            setattr(model, child_name, GeM())\n        else:\n            convert_to_gem(child)`\n\n\n`def convert_to_conv2d(model):\n    for child_name, child in model.named_children():\n        if child_name not in ['fc1','fc2']:\n            if isinstance(child, nn.Conv2d):\n                in_feat = child.in_channels\n                out_feat = child.out_channels\n                ker_size = child.kernel_size\n                stride = child.stride\n                padding = child.padding\n                dilation = child.dilation\n                groups = child.groups\n                setattr(model, child_name, Conv2d(in_channels=in_feat, out_channels=out_feat, kernel_size=ker_size, stride=stride,\n                                                 padding = padding, dilation=dilation, groups=groups))\n            else:\n                convert_to_conv2d(child)`\n\n\n`def convert_to_groupnorm(model):\n    for child_name, child in model.named_children():\n            if isinstance(child, nn.BatchNorm2d):\n                num_features = child.num_features\n                setattr(model, child_name, GroupNorm(num_groups=32, num_channels=num_features))\n            else:\n                convert_to_groupnorm(child)`\n\n\nIf you have any ideas , or suggestions please let us know! We also have a feeling using non-cropped images , and some other augmentations would have given us even better results. \nThanks for reading! \n\n\n\n\nEDIT - \nHere are the notebooks for this -\n\nInference - https://www.kaggle.com/p4rallax/private-0-9557/\n\nTraining - https://www.kaggle.com/virajbagal/potential-6th-place-lb-magic-layers-pp\n\n\n\n\n\n"
  }
}