{
  "id": 128911,
  "title": "Efficientnet - Trials and Analysis",
  "url": "/competitions/bengaliai-cv19/discussion/128911",
  "author_name": "",
  "post_date": "2020-02-04T07:22:34.274059800Z",
  "votes": 23,
  "comment_count": 20,
  "views": 0,
  "content": "<p>EfficientNet Models: [Single model , Single Fold]</p>\n\n<p>*<em>Study -1 *</em></p>\n\n<p>Things I tried:\nB3 - 300 x 300 LB: 0.9639 [40 epochs]\nB3 - 300 x 300 + cutmix LB: 0.9695 [40 epochs]</p>\n\n<p>B7 - 128 LB : 0.9683 [40 epochs]\nB7 - 224 LB : 0.9685 [15 epochs]</p>\n\n<p>Things from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to <a href=\"/iafoss\">@iafoss</a>\n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size</p>\n\n<p>*<em>Study -2 *</em>\nMost models will get WxH in the range of 16x16 or something similar. So if you use \"Conv\" in the tails it will work. But for efficientnet. Its 2560x4x4.  So using a conv layer on the tail doesn't going to capture much information.</p>\n\n<p><strong>Gem (Half Precision)</strong>\ndef gem(x, p=3, eps=1e-6):\n    x = x.double()  # x=x.to(torch.float32)  # comment this during inference\n    x = F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n    return x.half() # Comment this line in inference code use ## return x </p>\n\n<p>class GeM(nn.Module):\n    def <strong>init</strong>(self, p=3, eps=1e-6):\n        super(GeM,self).<strong>init</strong>()\n        # super().<strong>init</strong>()\n        self.p = Parameter(torch.ones(1)*p)\n        # print(self.p.dtype)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def <strong>repr</strong>(self):\n        return self.<strong>class</strong>._<em>name</em>_ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'</p>\n\n<p>I kindly request fellows who are trying efficientnet to share and discuss in this thread. So we can work and improve the scores together.</p>\n\n<p>*<em>Study -3 *</em>\nScore:\nB3 - 300 x 300 + cutmix LB: 0.9702 [40 epochs] Random Sampling\nB3 - 300 x 300 + cutmix LB: 0.9680 [40 epochs] Stratified Sampling + Data Cleaning\nB3 - 300 x 300 + cutmix LB: 0.9687 [40 epochs] Data Cleaning\nTail similar to  <a href=\"/iafoss\">@iafoss</a></p>\n\n<p>It seems Stratified Sampling + Data Cleaning affects the results. </p>\n\n<p>Data Cleaning:\nRemoving the graphemes that were having similar. The similar graphemes were reported in\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a> </p>",
  "messages": [
    {
      "id": "736440",
      "postDate": "02/04/2020 07:22:34",
      "content": "<p>EfficientNet Models: [Single model , Single Fold]</p>\n\n<p>*<em>Study -1 *</em></p>\n\n<p>Things I tried:\nB3 - 300 x 300 LB: 0.9639 [40 epochs]\nB3 - 300 x 300 + cutmix LB: 0.9695 [40 epochs]</p>\n\n<p>B7 - 128 LB : 0.9683 [40 epochs]\nB7 - 224 LB : 0.9685 [15 epochs]</p>\n\n<p>Things from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to <a href=\"/iafoss\">@iafoss</a>\n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size</p>\n\n<p>*<em>Study -2 *</em>\nMost models will get WxH in the range of 16x16 or something similar. So if you use \"Conv\" in the tails it will work. But for efficientnet. Its 2560x4x4.  So using a conv layer on the tail doesn't going to capture much information.</p>\n\n<p><strong>Gem (Half Precision)</strong>\ndef gem(x, p=3, eps=1e-6):\n    x = x.double()  # x=x.to(torch.float32)  # comment this during inference\n    x = F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n    return x.half() # Comment this line in inference code use ## return x </p>\n\n<p>class GeM(nn.Module):\n    def <strong>init</strong>(self, p=3, eps=1e-6):\n        super(GeM,self).<strong>init</strong>()\n        # super().<strong>init</strong>()\n        self.p = Parameter(torch.ones(1)*p)\n        # print(self.p.dtype)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def <strong>repr</strong>(self):\n        return self.<strong>class</strong>._<em>name</em>_ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'</p>\n\n<p>I kindly request fellows who are trying efficientnet to share and discuss in this thread. So we can work and improve the scores together.</p>\n\n<p>*<em>Study -3 *</em>\nScore:\nB3 - 300 x 300 + cutmix LB: 0.9702 [40 epochs] Random Sampling\nB3 - 300 x 300 + cutmix LB: 0.9680 [40 epochs] Stratified Sampling + Data Cleaning\nB3 - 300 x 300 + cutmix LB: 0.9687 [40 epochs] Data Cleaning\nTail similar to  <a href=\"/iafoss\">@iafoss</a></p>\n\n<p>It seems Stratified Sampling + Data Cleaning affects the results. </p>\n\n<p>Data Cleaning:\nRemoving the graphemes that were having similar. The similar graphemes were reported in\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123859\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123859</a> </p>",
      "rawMarkdown": "EfficientNet Models: [Single model , Single Fold]\n\n**Study -1 **\n\nThings I tried:\nB3 - 300 x 300 LB: 0.9639 [40 epochs]\nB3 - 300 x 300 + cutmix LB: 0.9695 [40 epochs]\n\nB7 - 128 LB : 0.9683 [40 epochs]\nB7 - 224 LB : 0.9685 [15 epochs]\n\nThings from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to @iafoss\n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size\n\n**Study -2 **\nMost models will get WxH in the range of 16x16 or something similar. So if you use \"Conv\" in the tails it will work. But for efficientnet. Its 2560x4x4.  So using a conv layer on the tail doesn't going to capture much information.\n\n**Gem (Half Precision)**\ndef gem(x, p=3, eps=1e-6):\n    x = x.double()  # x=x.to(torch.float32)  # comment this during inference\n    x = F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n    return x.half() # Comment this line in inference code use ## return x \n\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        # super().__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        # print(self.p.dtype)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n\nI kindly request fellows who are trying efficientnet to share and discuss in this thread. So we can work and improve the scores together.\n\n\n**Study -3 **\nScore:\nB3 - 300 x 300 + cutmix LB: 0.9702 [40 epochs] Random Sampling\nB3 - 300 x 300 + cutmix LB: 0.9680 [40 epochs] Stratified Sampling + Data Cleaning\nB3 - 300 x 300 + cutmix LB: 0.9687 [40 epochs] Data Cleaning\nTail similar to  @iafoss\n\nIt seems Stratified Sampling + Data Cleaning affects the results. \n\nData Cleaning:\nRemoving the graphemes that were having similar. The similar graphemes were reported in\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123859",
      "votes": null
    },
    {
      "id": "736613",
      "postDate": "02/04/2020 11:46:47",
      "content": "<p>thanks for sharing. </p>",
      "rawMarkdown": "thanks for sharing.",
      "votes": null
    },
    {
      "id": "736669",
      "postDate": "02/04/2020 12:49:55",
      "content": "<p>Thanks for sharing such insight for <strong>EfficientNet</strong>. I was also trying with Efficient most of the time but failed to score that high until now. My set-up was something like this:</p>\n\n<p>```\nmodel: eff b5\nimg_size: 128\nbatch_size: 40\nopt: adam\naugmentation: augmix\nepoch: 20</p>\n\n<p>LB: 0.9604\n<code>``\nI think you are using **PyTorch**, I'm in **Keras**. I found</code>cutmix<code>a bit complicated for me to implement. However, as far as the tail part is concern, I used</code>GAP<code>followed by basic dense-drop etc.  Right now, I am thinking I'm getting less leverage for using **Keras** for</code>cutmix-mixup-gem<code>:( - Anyway,</code>B7 - 128 LB : 0.9683 [40 epochs]<code>seems you didn't used **cutmix** and using bigger model for long epoch (long for at least B7) brought that score! Would you please inform some of your basic hyper-parameter setup, like</code>batch_size, optimizer, scheduler`?</p>",
      "rawMarkdown": "Thanks for sharing such insight for **EfficientNet**. I was also trying with Efficient most of the time but failed to score that high until now. My set-up was something like this:\n\n```\nmodel: eff b5\nimg_size: 128\nbatch_size: 40\nopt: adam\naugmentation: augmix\nepoch: 20\n\nLB: 0.9604\n```\nI think you are using **PyTorch**, I'm in **Keras**. I found `cutmix` a bit complicated for me to implement. However, as far as the tail part is concern, I used `GAP` followed by basic dense-drop etc.  Right now, I am thinking I'm getting less leverage for using **Keras** for `cutmix-mixup-gem` :( - Anyway, `B7 - 128 LB : 0.9683 [40 epochs]` seems you didn't used **cutmix** and using bigger model for long epoch (long for at least B7) brought that score! Would you please inform some of your basic hyper-parameter setup, like `batch_size, optimizer, scheduler`?",
      "votes": null
    },
    {
      "id": "736774",
      "postDate": "02/04/2020 14:44:31",
      "content": "<p>I reverted back to B3 bcz after many tests, my results with B3 was way better than B7. Using large model increases chances for overfitting. So better go with B0-B5.</p>\n\n<p>I havent tuned my hyper-params yet\nBatch size =50\nI cant give further info reg optimizer or scheduler as I am about to join a team. May be in later stages.</p>",
      "rawMarkdown": "I reverted back to B3 bcz after many tests, my results with B3 was way better than B7. Using large model increases chances for overfitting. So better go with B0-B5.\n\nI havent tuned my hyper-params yet\nBatch size =50\nI cant give further info reg optimizer or scheduler as I am about to join a team. May be in later stages.",
      "votes": null
    },
    {
      "id": "736807",
      "postDate": "02/04/2020 15:33:48",
      "content": "<p>Hi, <a href=\"/dhakshiin1601\">@dhakshiin1601</a>, thank you for sharing. I  also use effcientnet and my best lb is 0.9682, cv0.9692, eff4 224x224 input with cutmix/mixup and random shift rotate. Have you try swish replace Mish? Does it decrease the performance?</p>",
      "rawMarkdown": "Hi, @dhakshiin1601, thank you for sharing. I  also use effcientnet and my best lb is 0.9682, cv0.9692, eff4 224x224 input with cutmix/mixup and random shift rotate. Have you try swish replace Mish? Does it decrease the performance?",
      "votes": null
    },
    {
      "id": "736847",
      "postDate": "02/04/2020 16:25:07",
      "content": "<p>Question why you need <code>conv layer</code> in the tail ? My current <code>B0</code> can reach me top 50. In case you are wondering what I use   for B0:\ncut the model before <code>Pool</code>. This how it looks after cutting. \n<code>\n    (3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n    (4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n</code></p>\n\n<p>Add  0.5 * (<code>AdaptiveAvgPool2d</code> +<code>AdaptiveMaxPool2d</code>) , <code>Flatten()</code> + <code>nn.Linear</code>for each class</p>\n\n<p>Also check out this excellent post for another idea  using custom tail: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125819\">https://www.kaggle.com/c/bengaliai-cv19/discussion/125819</a></p>\n\n<p>btw I use <code>EffNet</code> from this repo: \n<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a></p>\n\n<p>Also for mixup you have to train 80+ epochs </p>",
      "rawMarkdown": "Question why you need `conv layer` in the tail ? My current `B0` can reach me top 50. In case you are wondering what I use   for B0:\ncut the model before `Pool`. This how it looks after cutting. \n```\n    (3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n    (4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n```\n\nAdd  0.5 * (`AdaptiveAvgPool2d` +`AdaptiveMaxPool2d`) , `Flatten()` + `nn.Linear `for each class\n\nAlso check out this excellent post for another idea  using custom tail: https://www.kaggle.com/c/bengaliai-cv19/discussion/125819\n\nbtw I use `EffNet` from this repo: \nhttps://github.com/rwightman/pytorch-image-models\n\nAlso for mixup you have to train 80+ epochs",
      "votes": null
    },
    {
      "id": "736867",
      "postDate": "02/04/2020 16:50:00",
      "content": "<p><a href=\"/drhabib\">@drhabib</a>  Is your current ranking is efficientnet?  😄 </p>",
      "rawMarkdown": "drhabib  Is your current ranking is efficientnet?  😄",
      "votes": null
    },
    {
      "id": "736877",
      "postDate": "02/04/2020 16:57:30",
      "content": "<p>its single model serenext50 =) still working hard to break 0.99  mark  =) </p>",
      "rawMarkdown": "its single model serenext50 =) still working hard to break 0.99  mark  =)",
      "votes": null
    },
    {
      "id": "736987",
      "postDate": "02/04/2020 19:16:11",
      "content": "<p>How did you implement EfficientNet for Keras <a href=\"/ipythonx\">@ipythonx</a>? Tried it while importing some models but always failed when sumbitting because I pip installed the model. Even when saving the model in .h5, I had some issues with the loss function not being implemented without the pip install. I'd be really curious to know!</p>",
      "rawMarkdown": "How did you implement EfficientNet for Keras @ipythonx? Tried it while importing some models but always failed when sumbitting because I pip installed the model. Even when saving the model in .h5, I had some issues with the loss function not being implemented without the pip install. I'd be really curious to know!",
      "votes": null
    },
    {
      "id": "737004",
      "postDate": "02/04/2020 19:48:44",
      "content": "<p>Refer to this for importing models into kernals: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128508\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128508</a></p>",
      "rawMarkdown": "Refer to this for importing models into kernals: https://www.kaggle.com/c/bengaliai-cv19/discussion/128508",
      "votes": null
    },
    {
      "id": "737118",
      "postDate": "02/04/2020 23:48:07",
      "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a> I also faced similar problem. However, I now train my model in offline.</p>",
      "rawMarkdown": "maxlenormand I also faced similar problem. However, I now train my model in offline.",
      "votes": null
    },
    {
      "id": "737313",
      "postDate": "02/05/2020 07:11:14",
      "content": "<p>Thanks a lot <a href=\"/greatgamedota\">@greatgamedota</a> ! I'll definitely try this! This would probably allow me to gain a few places on the leaderboard, my CV with EfficientNet was better than with DenseNet121.</p>\n\n<p>So <a href=\"/ipythonx\">@ipythonx</a> if I understand well you couldn't submit your models with EfficientNet?</p>",
      "rawMarkdown": "Thanks a lot @greatgamedota ! I'll definitely try this! This would probably allow me to gain a few places on the leaderboard, my CV with EfficientNet was better than with DenseNet121.\n\nSo @ipythonx if I understand well you couldn't submit your models with EfficientNet?",
      "votes": null
    },
    {
      "id": "737332",
      "postDate": "02/05/2020 07:37:13",
      "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a> I trained my model offline, and uploaded the trained weight to kernel. I didn't train my model in kaggle kernel. </p>",
      "rawMarkdown": "maxlenormand I trained my model offline, and uploaded the trained weight to kernel. I didn't train my model in kaggle kernel.",
      "votes": null
    },
    {
      "id": "737461",
      "postDate": "02/05/2020 11:21:44",
      "content": "<p>Thanks for sharing <a href=\"/dhakshiin1601\">@dhakshiin1601</a> will try efficientnet and report back</p>",
      "rawMarkdown": "Thanks for sharing @dhakshiin1601 will try efficientnet and report back",
      "votes": null
    },
    {
      "id": "737671",
      "postDate": "02/05/2020 16:27:25",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> That's what I had understood, I was just asking how you managed to submit an EfficientNet, because I had issues linked to the pip installing. I'll try the tricks proposed earlier though!</p>",
      "rawMarkdown": "ipythonx That's what I had understood, I was just asking how you managed to submit an EfficientNet, because I had issues linked to the pip installing. I'll try the tricks proposed earlier though!",
      "votes": null
    },
    {
      "id": "737678",
      "postDate": "02/05/2020 16:33:57",
      "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a> something similar like this, and their were some tails at the end of the model. As I said, I trained the model and uploaded the optimized weights and load likewise. \n```\n!pip install ../input/efficientnet-keras-source-code/repository/qubvel-efficientnet-c993591</p>\n\n<p>import efficientnet.keras as efn \nefnet = efn.EfficientNetB0(weights=None, include_top = False, input_shape=(128, 128, 3))\n```</p>",
      "rawMarkdown": "maxlenormand something similar like this, and their were some tails at the end of the model. As I said, I trained the model and uploaded the optimized weights and load likewise. \n```\n!pip install ../input/efficientnet-keras-source-code/repository/qubvel-efficientnet-c993591\n\nimport efficientnet.keras as efn \nefnet = efn.EfficientNetB0(weights=None, include_top = False, input_shape=(128, 128, 3))\n```",
      "votes": null
    },
    {
      "id": "737816",
      "postDate": "02/05/2020 20:00:20",
      "content": "<p>Could you tell me how did you manage to use EfficienNets when we are not allowed to use internet on our kernels. I usually use EfficientNetB0 but I'm currently not doing that because I cannot pip install anything. Could you help me out ? </p>",
      "rawMarkdown": "Could you tell me how did you manage to use EfficienNets when we are not allowed to use internet on our kernels. I usually use EfficientNetB0 but I'm currently not doing that because I cannot pip install anything. Could you help me out ?",
      "votes": null
    },
    {
      "id": "738058",
      "postDate": "02/06/2020 05:16:36",
      "content": "<p>Thanks <a href=\"/drhabib\">@drhabib</a> for your insights, I will update the tail as per your suggestion. I am currently using Lukemelas.</p>\n\n<p>One doubt regarding your cut procedure:\n<strong>*<em>Method 1</em>*</strong>\n*<em>Before Head:</em>*\n{Prev layers ..}</p>\n\n<p><strong>Head for each label</strong>\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear</p>\n\n<p>or \n<strong>*<em>Method 2</em>*</strong>\n*<em>Before Head:</em>*\n{Prev layers ..}\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)</p>\n\n<p><strong>Head for each label</strong>\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear</p>",
      "rawMarkdown": "Thanks @drhabib for your insights, I will update the tail as per your suggestion. I am currently using Lukemelas.\n\nOne doubt regarding your cut procedure:\n****Method 1****\n**Before Head:**\n{Prev layers ..}\n\n**Head for each label**\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear\n\nor \n****Method 2****\n**Before Head:**\n{Prev layers ..}\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n\n**Head for each label**\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear",
      "votes": null
    },
    {
      "id": "738518",
      "postDate": "02/06/2020 16:27:42",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> doing 0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) gave you better results than GeM()? btw thanks  for all the information</p>\n\n<p>And how many epochs did you train your B0? im having so bad results with effnet and idk why :\\</p>",
      "rawMarkdown": "drhabib doing 0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) gave you better results than GeM()? btw thanks  for all the information\n\nAnd how many epochs did you train your B0? im having so bad results with effnet and idk why :\\",
      "votes": null
    },
    {
      "id": "740919",
      "postDate": "02/10/2020 00:27:07",
      "content": "<p>Hi there, I am using efficientnet b2 at 128x128, should I or should I not use dense layers before my 3 output layers after I use an AveragePooling layer. I tried using them with b0, I got a pretty bad result even though I got excellent results using dense layers with wide-resnets. </p>",
      "rawMarkdown": "Hi there, I am using efficientnet b2 at 128x128, should I or should I not use dense layers before my 3 output layers after I use an AveragePooling layer. I tried using them with b0, I got a pretty bad result even though I got excellent results using dense layers with wide-resnets.",
      "votes": null
    },
    {
      "id": "742754",
      "postDate": "02/11/2020 14:24:42",
      "content": "<p>Is stratified splitting of data working ?? </p>\n\n<p>I have tried MultilabelStratifiedKFold. Is it offering any boost in performance ??</p>",
      "rawMarkdown": "Is stratified splitting of data working ?? \n\nI have tried MultilabelStratifiedKFold. Is it offering any boost in performance ??",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 736613,
      "author_name": "tiandaye",
      "author_url": "",
      "post_date": "02/04/2020 11:46:47",
      "content": "<p>thanks for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 736669,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "02/04/2020 12:49:55",
      "content": "<p>Thanks for sharing such insight for <strong>EfficientNet</strong>. I was also trying with Efficient most of the time but failed to score that high until now. My set-up was something like this:</p>\n\n<p>```\nmodel: eff b5\nimg_size: 128\nbatch_size: 40\nopt: adam\naugmentation: augmix\nepoch: 20</p>\n\n<p>LB: 0.9604\n<code>``\nI think you are using **PyTorch**, I'm in **Keras**. I found</code>cutmix<code>a bit complicated for me to implement. However, as far as the tail part is concern, I used</code>GAP<code>followed by basic dense-drop etc.  Right now, I am thinking I'm getting less leverage for using **Keras** for</code>cutmix-mixup-gem<code>:( - Anyway,</code>B7 - 128 LB : 0.9683 [40 epochs]<code>seems you didn't used **cutmix** and using bigger model for long epoch (long for at least B7) brought that score! Would you please inform some of your basic hyper-parameter setup, like</code>batch_size, optimizer, scheduler`?</p>",
      "votes": null,
      "replies": [
        {
          "id": 736774,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "02/04/2020 14:44:31",
          "content": "<p>I reverted back to B3 bcz after many tests, my results with B3 was way better than B7. Using large model increases chances for overfitting. So better go with B0-B5.</p>\n\n<p>I havent tuned my hyper-params yet\nBatch size =50\nI cant give further info reg optimizer or scheduler as I am about to join a team. May be in later stages.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 736987,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/04/2020 19:16:11",
          "content": "<p>How did you implement EfficientNet for Keras <a href=\"/ipythonx\">@ipythonx</a>? Tried it while importing some models but always failed when sumbitting because I pip installed the model. Even when saving the model in .h5, I had some issues with the loss function not being implemented without the pip install. I'd be really curious to know!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737004,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/04/2020 19:48:44",
          "content": "<p>Refer to this for importing models into kernals: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128508\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128508</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737118,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "02/04/2020 23:48:07",
          "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a> I also faced similar problem. However, I now train my model in offline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737313,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/05/2020 07:11:14",
          "content": "<p>Thanks a lot <a href=\"/greatgamedota\">@greatgamedota</a> ! I'll definitely try this! This would probably allow me to gain a few places on the leaderboard, my CV with EfficientNet was better than with DenseNet121.</p>\n\n<p>So <a href=\"/ipythonx\">@ipythonx</a> if I understand well you couldn't submit your models with EfficientNet?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737332,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "02/05/2020 07:37:13",
          "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a> I trained my model offline, and uploaded the trained weight to kernel. I didn't train my model in kaggle kernel. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737671,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/05/2020 16:27:25",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> That's what I had understood, I was just asking how you managed to submit an EfficientNet, because I had issues linked to the pip installing. I'll try the tricks proposed earlier though!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737678,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "02/05/2020 16:33:57",
          "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a> something similar like this, and their were some tails at the end of the model. As I said, I trained the model and uploaded the optimized weights and load likewise. \n```\n!pip install ../input/efficientnet-keras-source-code/repository/qubvel-efficientnet-c993591</p>\n\n<p>import efficientnet.keras as efn \nefnet = efn.EfficientNetB0(weights=None, include_top = False, input_shape=(128, 128, 3))\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740919,
          "author_name": "namansingh2803",
          "author_url": "",
          "post_date": "02/10/2020 00:27:07",
          "content": "<p>Hi there, I am using efficientnet b2 at 128x128, should I or should I not use dense layers before my 3 output layers after I use an AveragePooling layer. I tried using them with b0, I got a pretty bad result even though I got excellent results using dense layers with wide-resnets. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 736807,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "02/04/2020 15:33:48",
      "content": "<p>Hi, <a href=\"/dhakshiin1601\">@dhakshiin1601</a>, thank you for sharing. I  also use effcientnet and my best lb is 0.9682, cv0.9692, eff4 224x224 input with cutmix/mixup and random shift rotate. Have you try swish replace Mish? Does it decrease the performance?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 736847,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "02/04/2020 16:25:07",
      "content": "<p>Question why you need <code>conv layer</code> in the tail ? My current <code>B0</code> can reach me top 50. In case you are wondering what I use   for B0:\ncut the model before <code>Pool</code>. This how it looks after cutting. \n<code>\n    (3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n    (4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n</code></p>\n\n<p>Add  0.5 * (<code>AdaptiveAvgPool2d</code> +<code>AdaptiveMaxPool2d</code>) , <code>Flatten()</code> + <code>nn.Linear</code>for each class</p>\n\n<p>Also check out this excellent post for another idea  using custom tail: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125819\">https://www.kaggle.com/c/bengaliai-cv19/discussion/125819</a></p>\n\n<p>btw I use <code>EffNet</code> from this repo: \n<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a></p>\n\n<p>Also for mixup you have to train 80+ epochs </p>",
      "votes": null,
      "replies": [
        {
          "id": 736867,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/04/2020 16:50:00",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  Is your current ranking is efficientnet?  😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 736877,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/04/2020 16:57:30",
          "content": "<p>its single model serenext50 =) still working hard to break 0.99  mark  =) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737816,
          "author_name": "namansingh2803",
          "author_url": "",
          "post_date": "02/05/2020 20:00:20",
          "content": "<p>Could you tell me how did you manage to use EfficienNets when we are not allowed to use internet on our kernels. I usually use EfficientNetB0 but I'm currently not doing that because I cannot pip install anything. Could you help me out ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738058,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "02/06/2020 05:16:36",
          "content": "<p>Thanks <a href=\"/drhabib\">@drhabib</a> for your insights, I will update the tail as per your suggestion. I am currently using Lukemelas.</p>\n\n<p>One doubt regarding your cut procedure:\n<strong>*<em>Method 1</em>*</strong>\n*<em>Before Head:</em>*\n{Prev layers ..}</p>\n\n<p><strong>Head for each label</strong>\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear</p>\n\n<p>or \n<strong>*<em>Method 2</em>*</strong>\n*<em>Before Head:</em>*\n{Prev layers ..}\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)</p>\n\n<p><strong>Head for each label</strong>\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738518,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "02/06/2020 16:27:42",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> doing 0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) gave you better results than GeM()? btw thanks  for all the information</p>\n\n<p>And how many epochs did you train your B0? im having so bad results with effnet and idk why :\\</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 737461,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "02/05/2020 11:21:44",
      "content": "<p>Thanks for sharing <a href=\"/dhakshiin1601\">@dhakshiin1601</a> will try efficientnet and report back</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 742754,
      "author_name": "dhakshiin1601",
      "author_url": "",
      "post_date": "02/11/2020 14:24:42",
      "content": "<p>Is stratified splitting of data working ?? </p>\n\n<p>I have tried MultilabelStratifiedKFold. Is it offering any boost in performance ??</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "736440": "EfficientNet Models: [Single model , Single Fold]\n\n**Study -1 **\n\nThings I tried:\nB3 - 300 x 300 LB: 0.9639 [40 epochs]\nB3 - 300 x 300 + cutmix LB: 0.9695 [40 epochs]\n\nB7 - 128 LB : 0.9683 [40 epochs]\nB7 - 224 LB : 0.9685 [15 epochs]\n\nThings from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to @iafoss\n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size\n\n**Study -2 **\nMost models will get WxH in the range of 16x16 or something similar. So if you use \"Conv\" in the tails it will work. But for efficientnet. Its 2560x4x4.  So using a conv layer on the tail doesn't going to capture much information.\n\n**Gem (Half Precision)**\ndef gem(x, p=3, eps=1e-6):\n    x = x.double()  # x=x.to(torch.float32)  # comment this during inference\n    x = F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n    return x.half() # Comment this line in inference code use ## return x \n\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        # super().__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        # print(self.p.dtype)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n\nI kindly request fellows who are trying efficientnet to share and discuss in this thread. So we can work and improve the scores together.\n\n\n**Study -3 **\nScore:\nB3 - 300 x 300 + cutmix LB: 0.9702 [40 epochs] Random Sampling\nB3 - 300 x 300 + cutmix LB: 0.9680 [40 epochs] Stratified Sampling + Data Cleaning\nB3 - 300 x 300 + cutmix LB: 0.9687 [40 epochs] Data Cleaning\nTail similar to  @iafoss\n\nIt seems Stratified Sampling + Data Cleaning affects the results. \n\nData Cleaning:\nRemoving the graphemes that were having similar. The similar graphemes were reported in\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123859",
    "736613": "thanks for sharing.",
    "736669": "Thanks for sharing such insight for **EfficientNet**. I was also trying with Efficient most of the time but failed to score that high until now. My set-up was something like this:\n\n```\nmodel: eff b5\nimg_size: 128\nbatch_size: 40\nopt: adam\naugmentation: augmix\nepoch: 20\n\nLB: 0.9604\n```\nI think you are using **PyTorch**, I'm in **Keras**. I found `cutmix` a bit complicated for me to implement. However, as far as the tail part is concern, I used `GAP` followed by basic dense-drop etc.  Right now, I am thinking I'm getting less leverage for using **Keras** for `cutmix-mixup-gem` :( - Anyway, `B7 - 128 LB : 0.9683 [40 epochs]` seems you didn't used **cutmix** and using bigger model for long epoch (long for at least B7) brought that score! Would you please inform some of your basic hyper-parameter setup, like `batch_size, optimizer, scheduler`?",
    "736774": "I reverted back to B3 bcz after many tests, my results with B3 was way better than B7. Using large model increases chances for overfitting. So better go with B0-B5.\n\nI havent tuned my hyper-params yet\nBatch size =50\nI cant give further info reg optimizer or scheduler as I am about to join a team. May be in later stages.",
    "736807": "Hi, @dhakshiin1601, thank you for sharing. I  also use effcientnet and my best lb is 0.9682, cv0.9692, eff4 224x224 input with cutmix/mixup and random shift rotate. Have you try swish replace Mish? Does it decrease the performance?",
    "736847": "Question why you need `conv layer` in the tail ? My current `B0` can reach me top 50. In case you are wondering what I use   for B0:\ncut the model before `Pool`. This how it looks after cutting. \n```\n    (3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n    (4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n```\n\nAdd  0.5 * (`AdaptiveAvgPool2d` +`AdaptiveMaxPool2d`) , `Flatten()` + `nn.Linear `for each class\n\nAlso check out this excellent post for another idea  using custom tail: https://www.kaggle.com/c/bengaliai-cv19/discussion/125819\n\nbtw I use `EffNet` from this repo: \nhttps://github.com/rwightman/pytorch-image-models\n\nAlso for mixup you have to train 80+ epochs",
    "736867": "drhabib  Is your current ranking is efficientnet?  😄",
    "736877": "its single model serenext50 =) still working hard to break 0.99  mark  =)",
    "736987": "How did you implement EfficientNet for Keras @ipythonx? Tried it while importing some models but always failed when sumbitting because I pip installed the model. Even when saving the model in .h5, I had some issues with the loss function not being implemented without the pip install. I'd be really curious to know!",
    "737004": "Refer to this for importing models into kernals: https://www.kaggle.com/c/bengaliai-cv19/discussion/128508",
    "737118": "maxlenormand I also faced similar problem. However, I now train my model in offline.",
    "737313": "Thanks a lot @greatgamedota ! I'll definitely try this! This would probably allow me to gain a few places on the leaderboard, my CV with EfficientNet was better than with DenseNet121.\n\nSo @ipythonx if I understand well you couldn't submit your models with EfficientNet?",
    "737332": "maxlenormand I trained my model offline, and uploaded the trained weight to kernel. I didn't train my model in kaggle kernel.",
    "737461": "Thanks for sharing @dhakshiin1601 will try efficientnet and report back",
    "737671": "ipythonx That's what I had understood, I was just asking how you managed to submit an EfficientNet, because I had issues linked to the pip installing. I'll try the tricks proposed earlier though!",
    "737678": "maxlenormand something similar like this, and their were some tails at the end of the model. As I said, I trained the model and uploaded the optimized weights and load likewise. \n```\n!pip install ../input/efficientnet-keras-source-code/repository/qubvel-efficientnet-c993591\n\nimport efficientnet.keras as efn \nefnet = efn.EfficientNetB0(weights=None, include_top = False, input_shape=(128, 128, 3))\n```",
    "737816": "Could you tell me how did you manage to use EfficienNets when we are not allowed to use internet on our kernels. I usually use EfficientNetB0 but I'm currently not doing that because I cannot pip install anything. Could you help me out ?",
    "738058": "Thanks @drhabib for your insights, I will update the tail as per your suggestion. I am currently using Lukemelas.\n\nOne doubt regarding your cut procedure:\n****Method 1****\n**Before Head:**\n{Prev layers ..}\n\n**Head for each label**\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear\n\nor \n****Method 2****\n**Before Head:**\n{Prev layers ..}\n(3): Conv2d(320, 1280, kernel_size=(1, 1), stride=(1, 1), bias=False)\n(4): BatchNorm2d(1280, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n\n**Head for each label**\n0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) , Flatten() + nn.Linear",
    "738518": "drhabib doing 0.5 * (AdaptiveAvgPool2d +AdaptiveMaxPool2d) gave you better results than GeM()? btw thanks  for all the information\n\nAnd how many epochs did you train your B0? im having so bad results with effnet and idk why :\\",
    "740919": "Hi there, I am using efficientnet b2 at 128x128, should I or should I not use dense layers before my 3 output layers after I use an AveragePooling layer. I tried using them with b0, I got a pretty bad result even though I got excellent results using dense layers with wide-resnets.",
    "742754": "Is stratified splitting of data working ?? \n\nI have tried MultilabelStratifiedKFold. Is it offering any boost in performance ??"
  },
  "source": "meta"
}