{
  "id": 103743,
  "title": "EfficientNet B6 & B7 Weights Released",
  "url": "/competitions/aptos2019-blindness-detection/discussion/103743",
  "author_name": "",
  "post_date": "2019-08-11T12:05:46.588252200Z",
  "votes": 33,
  "comment_count": 25,
  "views": 0,
  "content": "<p>I saw that B6 and B7 weights were made public recently, along with the improvements from using AutoAugment for preprocessing. Has anyone tried using for this competition yet? Any improvements over B5? </p>\n\n<p>TF: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet</a>\nPyTorch: <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\">https://github.com/lukemelas/EfficientNet-PyTorch</a></p>",
  "messages": [
    {
      "id": "596861",
      "postDate": "08/11/2019 12:05:46",
      "content": "<p>I saw that B6 and B7 weights were made public recently, along with the improvements from using AutoAugment for preprocessing. Has anyone tried using for this competition yet? Any improvements over B5? </p>\n\n<p>TF: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet</a>\nPyTorch: <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\">https://github.com/lukemelas/EfficientNet-PyTorch</a></p>",
      "rawMarkdown": "I saw that B6 and B7 weights were made public recently, along with the improvements from using AutoAugment for preprocessing. Has anyone tried using for this competition yet? Any improvements over B5? \n\nTF: https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\nPyTorch: https://github.com/lukemelas/EfficientNet-PyTorch",
      "votes": null
    },
    {
      "id": "596990",
      "postDate": "08/11/2019 16:04:59",
      "content": "<p>Thanks Tom! in cases that our kaggle GPU is OK for handling the big two, I will try them this week.</p>",
      "rawMarkdown": "Thanks Tom! in cases that our kaggle GPU is OK for handling the big two, I will try them this week.",
      "votes": null
    },
    {
      "id": "597025",
      "postDate": "08/11/2019 17:06:26",
      "content": "<p>Awesome, if I get some time I will try too. Let me know how you get on :)</p>",
      "rawMarkdown": "Awesome, if I get some time I will try too. Let me know how you get on :)",
      "votes": null
    },
    {
      "id": "597053",
      "postDate": "08/11/2019 18:12:10",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "597187",
      "postDate": "08/12/2019 01:29:00",
      "content": "<p>B7 scored less than B5 when I tried! Also, the new EfficientNet pre-trained models(PyTorch) with AutoAugment are doing worse than the Standard ones. I must be doing something very wrong 🙁 </p>",
      "rawMarkdown": "B7 scored less than B5 when I tried! Also, the new EfficientNet pre-trained models(PyTorch) with AutoAugment are doing worse than the Standard ones. I must be doing something very wrong 🙁",
      "votes": null
    },
    {
      "id": "597505",
      "postDate": "08/12/2019 12:40:33",
      "content": "<p>Given Kaggle GPU constraints, it is bit difficult to train B7 on it. But you can experiment with B6 and Good news is you will definitely get better score with B6 as compare to B5 :)</p>",
      "rawMarkdown": "Given Kaggle GPU constraints, it is bit difficult to train B7 on it. But you can experiment with B6 and Good news is you will definitely get better score with B6 as compare to B5 :)",
      "votes": null
    },
    {
      "id": "597573",
      "postDate": "08/12/2019 14:22:24",
      "content": "<p>Thank you for posting. One important thing about <code>EfficientNet</code> models that they are designed to take in to account input image dimensions.</p>\n\n<p>So if you want to squeeze every last droplet from your model make sure to use same image resolutions =)</p>\n\n<p><code>\nefficientnet-b0-224\nefficientnet-b1-240\nefficientnet-b2-260\nefficientnet-b3-300\nefficientnet-b4-380\nefficientnet-b5-456\nefficientnet-b6-528\nefficientnet-b7-600\n</code>\nYou can get more info about the design and why resolution matters from the official paper:  <a href=\"https://arxiv.org/abs/1905.11946\">https://arxiv.org/abs/1905.11946</a></p>",
      "rawMarkdown": "Thank you for posting. One important thing about `EfficientNet ` models that they are designed to take in to account input image dimensions.\n\nSo if you want to squeeze every last droplet from your model make sure to use same image resolutions =)\n\n```\nefficientnet-b0-224\nefficientnet-b1-240\nefficientnet-b2-260\nefficientnet-b3-300\nefficientnet-b4-380\nefficientnet-b5-456\nefficientnet-b6-528\nefficientnet-b7-600\n```\nYou can get more info about the design and why resolution matters from the official paper:  https://arxiv.org/abs/1905.11946",
      "votes": null
    },
    {
      "id": "597703",
      "postDate": "08/12/2019 17:15:14",
      "content": "<p>Great comment, I wasn't aware of this, thanks!</p>",
      "rawMarkdown": "Great comment, I wasn't aware of this, thanks!",
      "votes": null
    },
    {
      "id": "597983",
      "postDate": "08/13/2019 01:34:40",
      "content": "<p>I wasn't are of this, too. Thank you for sharing!</p>",
      "rawMarkdown": "I wasn't are of this, too. Thank you for sharing!",
      "votes": null
    },
    {
      "id": "598237",
      "postDate": "08/13/2019 10:34:28",
      "content": "<p>Hi Tom, it seems that Keras B6/7 weights hasn't come out yet, so I will to wait a bit :)</p>\n\n<p><strong>EDIT</strong> For keras user, \nIt appears at the moment that if we use <code>!git clone https://github.com/qubvel/efficientnet.git</code>\ninstead of <code>!pip3 install efficientnet</code>, we will get the latest version which has new B6/B7 implementation and pretrained weights ... :)</p>\n\n<p><strong>UPDATED</strong> Hi Tom <a href=\"/taindow\">@taindow</a>, I cannot make B6 converges nicely in 9-hour kernel (I also faced some technical difficulty in order to continue training in another kernel) ... I think it is promising if we can train on a local machine ... However, since I am focusing only on kernel-only training, I will move to other approaches :)</p>",
      "rawMarkdown": "Hi Tom, it seems that Keras B6/7 weights hasn't come out yet, so I will to wait a bit :)\n\n**EDIT** For keras user, \nIt appears at the moment that if we use `!git clone https://github.com/qubvel/efficientnet.git`\ninstead of `!pip3 install efficientnet`, we will get the latest version which has new B6/B7 implementation and pretrained weights ... :)\n\n**UPDATED** Hi Tom @taindow, I cannot make B6 converges nicely in 9-hour kernel (I also faced some technical difficulty in order to continue training in another kernel) ... I think it is promising if we can train on a local machine ... However, since I am focusing only on kernel-only training, I will move to other approaches :)",
      "votes": null
    },
    {
      "id": "598316",
      "postDate": "08/13/2019 13:06:44",
      "content": "<p>Nice! Thanks for sharing! Remember to replace the Batch Normalization layers if you are using very small batch sizes with B6 and B7. <a href=\"https://arxiv.org/pdf/1803.08494.pdf\">Group Normalization</a> can be a great alternative to Batch Normalization in this case.</p>",
      "rawMarkdown": "Nice! Thanks for sharing! Remember to replace the Batch Normalization layers if you are using very small batch sizes with B6 and B7. [Group Normalization](https://arxiv.org/pdf/1803.08494.pdf) can be a great alternative to Batch Normalization in this case.",
      "votes": null
    },
    {
      "id": "598677",
      "postDate": "08/13/2019 21:38:01",
      "content": "<p>I do believe that bigger architecture should give a better accuracy (B6/B7) also taking into account the image size that goes along with it to get the maximum , which also means that we would have to use a very small batch size even with mixed precision (FP16) for inference. Which makes it even more time consuming (with all the preprocessing that we already do) to predict for the private test set. Did anyone run an inference kernel on B6/B7 yet ? If so did you make a note of how long it took ? Please share the details.</p>",
      "rawMarkdown": "I do believe that bigger architecture should give a better accuracy (B6/B7) also taking into account the image size that goes along with it to get the maximum , which also means that we would have to use a very small batch size even with mixed precision (FP16) for inference. Which makes it even more time consuming (with all the preprocessing that we already do) to predict for the private test set. Did anyone run an inference kernel on B6/B7 yet ? If so did you make a note of how long it took ? Please share the details.",
      "votes": null
    },
    {
      "id": "598763",
      "postDate": "08/14/2019 02:24:18",
      "content": "<p>Actually image size may matter MORE for those architectures that weren't designed to take image size into account, as their budget for width, depth and image size are less balanced. I think it's ok as long as your image size is similar to the recommended ones, those sizes are just \"efficient\" with respect to computational power.</p>",
      "rawMarkdown": "Actually image size may matter MORE for those architectures that weren't designed to take image size into account, as their budget for width, depth and image size are less balanced. I think it's ok as long as your image size is similar to the recommended ones, those sizes are just \"efficient\" with respect to computational power.",
      "votes": null
    },
    {
      "id": "598947",
      "postDate": "08/14/2019 09:08:56",
      "content": "<p>Larger image size does not seem to work for me. It takes too long to train. </p>",
      "rawMarkdown": "Larger image size does not seem to work for me. It takes too long to train.",
      "votes": null
    },
    {
      "id": "599175",
      "postDate": "08/14/2019 15:44:56",
      "content": "<p>How can you remove the Batch Normalization layers?</p>",
      "rawMarkdown": "How can you remove the Batch Normalization layers?",
      "votes": null
    },
    {
      "id": "600417",
      "postDate": "08/16/2019 05:02:26",
      "content": "<p>But with b6(and 500+ image size) what's the biggest batch size for a kaggle kernel? 8?</p>",
      "rawMarkdown": "But with b6(and 500+ image size) what's the biggest batch size for a kaggle kernel? 8?",
      "votes": null
    },
    {
      "id": "600468",
      "postDate": "08/16/2019 06:25:19",
      "content": "<p>Thank you very much. \nI read the paper &amp; also knew that, resolution is part of their \"compounding scaling\" method. Then also didn't got idea, resolution is important for this model.</p>",
      "rawMarkdown": "Thank you very much. \nI read the paper &amp; also knew that, resolution is part of their \"compounding scaling\" method. Then also didn't got idea, resolution is important for this model.",
      "votes": null
    },
    {
      "id": "600471",
      "postDate": "08/16/2019 06:27:26",
      "content": "<p>May be you're trying without changing resolution.\nSee DrHB's above <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/103743#597573\">this</a> comment.\nAfter that, problem might get solved. Hope so.</p>",
      "rawMarkdown": "May be you're trying without changing resolution.\nSee DrHB's above [this](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/103743#597573) comment.\nAfter that, problem might get solved. Hope so.",
      "votes": null
    },
    {
      "id": "600524",
      "postDate": "08/16/2019 08:28:06",
      "content": "<p>I'm very much aware of that and have been using matching resolution since my 1st submission with EfficientNets. It is to be noted that with increasing resolution, batch size decreases so much that the training takes ages to reach a reasonable score on the Validation set. I've to train longer maybe!</p>",
      "rawMarkdown": "I'm very much aware of that and have been using matching resolution since my 1st submission with EfficientNets. It is to be noted that with increasing resolution, batch size decreases so much that the training takes ages to reach a reasonable score on the Validation set. I've to train longer maybe!",
      "votes": null
    },
    {
      "id": "600720",
      "postDate": "08/16/2019 13:47:06",
      "content": "<p>True. I also need to halve my batch size &amp; epochs to fit in 9 hour limit.</p>",
      "rawMarkdown": "True. I also need to halve my batch size &amp; epochs to fit in 9 hour limit.",
      "votes": null
    },
    {
      "id": "601192",
      "postDate": "08/17/2019 06:51:00",
      "content": "<p>Surprisingly, after using 300 as the dimension for training B3, my LB score dropped about 2 %. My best score is .794 using B3 and image size 224. Used <a href=\"/drhabib\">@drhabib</a> 's method for training both networks and used the same parameters for training.  What am I doing wrong? Is there any other trick I should use?</p>",
      "rawMarkdown": "Surprisingly, after using 300 as the dimension for training B3, my LB score dropped about 2 %. My best score is .794 using B3 and image size 224. Used @drhabib 's method for training both networks and used the same parameters for training.  What am I doing wrong? Is there any other trick I should use?",
      "votes": null
    },
    {
      "id": "601217",
      "postDate": "08/17/2019 07:56:09",
      "content": "<p>Same here...</p>",
      "rawMarkdown": "Same here...",
      "votes": null
    },
    {
      "id": "601219",
      "postDate": "08/17/2019 08:00:44",
      "content": "<p>Have you decreased batch size for training with higher resolution?</p>",
      "rawMarkdown": "Have you decreased batch size for training with higher resolution?",
      "votes": null
    },
    {
      "id": "601249",
      "postDate": "08/17/2019 09:12:52",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "601298",
      "postDate": "08/17/2019 11:26:26",
      "content": "<p><a href=\"/tahsin\">@tahsin</a>, same here. \nReason might be overfitting. As public is just 15%, can't rely on public LB.</p>",
      "rawMarkdown": "tahsin, same here. \nReason might be overfitting. As public is just 15%, can't rely on public LB.",
      "votes": null
    },
    {
      "id": "602804",
      "postDate": "08/19/2019 14:06:11",
      "content": "<p>Take a look at this kernel: <a href=\"https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019\">EfficientNetB5 with Keras (APTOS 2019)</a> from input line 15 to 17.\nThe code is super clear.</p>",
      "rawMarkdown": "Take a look at this kernel: [EfficientNetB5 with Keras (APTOS 2019)](https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019) from input line 15 to 17.\nThe code is super clear.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 596990,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "08/11/2019 16:04:59",
      "content": "<p>Thanks Tom! in cases that our kaggle GPU is OK for handling the big two, I will try them this week.</p>",
      "votes": null,
      "replies": [
        {
          "id": 597025,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "08/11/2019 17:06:26",
          "content": "<p>Awesome, if I get some time I will try too. Let me know how you get on :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 598237,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "08/13/2019 10:34:28",
          "content": "<p>Hi Tom, it seems that Keras B6/7 weights hasn't come out yet, so I will to wait a bit :)</p>\n\n<p><strong>EDIT</strong> For keras user, \nIt appears at the moment that if we use <code>!git clone https://github.com/qubvel/efficientnet.git</code>\ninstead of <code>!pip3 install efficientnet</code>, we will get the latest version which has new B6/B7 implementation and pretrained weights ... :)</p>\n\n<p><strong>UPDATED</strong> Hi Tom <a href=\"/taindow\">@taindow</a>, I cannot make B6 converges nicely in 9-hour kernel (I also faced some technical difficulty in order to continue training in another kernel) ... I think it is promising if we can train on a local machine ... However, since I am focusing only on kernel-only training, I will move to other approaches :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 597053,
      "author_name": "harshthaker",
      "author_url": "",
      "post_date": "08/11/2019 18:12:10",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 597187,
      "author_name": "kirankunapuli",
      "author_url": "",
      "post_date": "08/12/2019 01:29:00",
      "content": "<p>B7 scored less than B5 when I tried! Also, the new EfficientNet pre-trained models(PyTorch) with AutoAugment are doing worse than the Standard ones. I must be doing something very wrong 🙁 </p>",
      "votes": null,
      "replies": [
        {
          "id": 600471,
          "author_name": "prashantkikani",
          "author_url": "",
          "post_date": "08/16/2019 06:27:26",
          "content": "<p>May be you're trying without changing resolution.\nSee DrHB's above <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/103743#597573\">this</a> comment.\nAfter that, problem might get solved. Hope so.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 600524,
          "author_name": "kirankunapuli",
          "author_url": "",
          "post_date": "08/16/2019 08:28:06",
          "content": "<p>I'm very much aware of that and have been using matching resolution since my 1st submission with EfficientNets. It is to be noted that with increasing resolution, batch size decreases so much that the training takes ages to reach a reasonable score on the Validation set. I've to train longer maybe!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 600720,
          "author_name": "prashantkikani",
          "author_url": "",
          "post_date": "08/16/2019 13:47:06",
          "content": "<p>True. I also need to halve my batch size &amp; epochs to fit in 9 hour limit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 597505,
      "author_name": "monsterspy",
      "author_url": "",
      "post_date": "08/12/2019 12:40:33",
      "content": "<p>Given Kaggle GPU constraints, it is bit difficult to train B7 on it. But you can experiment with B6 and Good news is you will definitely get better score with B6 as compare to B5 :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 600417,
          "author_name": "homoalways",
          "author_url": "",
          "post_date": "08/16/2019 05:02:26",
          "content": "<p>But with b6(and 500+ image size) what's the biggest batch size for a kaggle kernel? 8?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 597573,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "08/12/2019 14:22:24",
      "content": "<p>Thank you for posting. One important thing about <code>EfficientNet</code> models that they are designed to take in to account input image dimensions.</p>\n\n<p>So if you want to squeeze every last droplet from your model make sure to use same image resolutions =)</p>\n\n<p><code>\nefficientnet-b0-224\nefficientnet-b1-240\nefficientnet-b2-260\nefficientnet-b3-300\nefficientnet-b4-380\nefficientnet-b5-456\nefficientnet-b6-528\nefficientnet-b7-600\n</code>\nYou can get more info about the design and why resolution matters from the official paper:  <a href=\"https://arxiv.org/abs/1905.11946\">https://arxiv.org/abs/1905.11946</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 597703,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "08/12/2019 17:15:14",
          "content": "<p>Great comment, I wasn't aware of this, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 597983,
          "author_name": "soongja",
          "author_url": "",
          "post_date": "08/13/2019 01:34:40",
          "content": "<p>I wasn't are of this, too. Thank you for sharing!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 598763,
          "author_name": "homoalways",
          "author_url": "",
          "post_date": "08/14/2019 02:24:18",
          "content": "<p>Actually image size may matter MORE for those architectures that weren't designed to take image size into account, as their budget for width, depth and image size are less balanced. I think it's ok as long as your image size is similar to the recommended ones, those sizes are just \"efficient\" with respect to computational power.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 600468,
          "author_name": "prashantkikani",
          "author_url": "",
          "post_date": "08/16/2019 06:25:19",
          "content": "<p>Thank you very much. \nI read the paper &amp; also knew that, resolution is part of their \"compounding scaling\" method. Then also didn't got idea, resolution is important for this model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601192,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/17/2019 06:51:00",
          "content": "<p>Surprisingly, after using 300 as the dimension for training B3, my LB score dropped about 2 %. My best score is .794 using B3 and image size 224. Used <a href=\"/drhabib\">@drhabib</a> 's method for training both networks and used the same parameters for training.  What am I doing wrong? Is there any other trick I should use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601217,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/17/2019 07:56:09",
          "content": "<p>Same here...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601219,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/17/2019 08:00:44",
          "content": "<p>Have you decreased batch size for training with higher resolution?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601249,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/17/2019 09:12:52",
          "content": "<p>yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601298,
          "author_name": "prashantkikani",
          "author_url": "",
          "post_date": "08/17/2019 11:26:26",
          "content": "<p><a href=\"/tahsin\">@tahsin</a>, same here. \nReason might be overfitting. As public is just 15%, can't rely on public LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 598316,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "08/13/2019 13:06:44",
      "content": "<p>Nice! Thanks for sharing! Remember to replace the Batch Normalization layers if you are using very small batch sizes with B6 and B7. <a href=\"https://arxiv.org/pdf/1803.08494.pdf\">Group Normalization</a> can be a great alternative to Batch Normalization in this case.</p>",
      "votes": null,
      "replies": [
        {
          "id": 599175,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/14/2019 15:44:56",
          "content": "<p>How can you remove the Batch Normalization layers?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602804,
          "author_name": "andreaamico90",
          "author_url": "",
          "post_date": "08/19/2019 14:06:11",
          "content": "<p>Take a look at this kernel: <a href=\"https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019\">EfficientNetB5 with Keras (APTOS 2019)</a> from input line 15 to 17.\nThe code is super clear.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 598677,
      "author_name": "sachinprabhu",
      "author_url": "",
      "post_date": "08/13/2019 21:38:01",
      "content": "<p>I do believe that bigger architecture should give a better accuracy (B6/B7) also taking into account the image size that goes along with it to get the maximum , which also means that we would have to use a very small batch size even with mixed precision (FP16) for inference. Which makes it even more time consuming (with all the preprocessing that we already do) to predict for the private test set. Did anyone run an inference kernel on B6/B7 yet ? If so did you make a note of how long it took ? Please share the details.</p>",
      "votes": null,
      "replies": [
        {
          "id": 598947,
          "author_name": "harshthaker",
          "author_url": "",
          "post_date": "08/14/2019 09:08:56",
          "content": "<p>Larger image size does not seem to work for me. It takes too long to train. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "596861": "I saw that B6 and B7 weights were made public recently, along with the improvements from using AutoAugment for preprocessing. Has anyone tried using for this competition yet? Any improvements over B5? \n\nTF: https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\nPyTorch: https://github.com/lukemelas/EfficientNet-PyTorch",
    "596990": "Thanks Tom! in cases that our kaggle GPU is OK for handling the big two, I will try them this week.",
    "597025": "Awesome, if I get some time I will try too. Let me know how you get on :)",
    "597053": "Thanks for sharing :)",
    "597187": "B7 scored less than B5 when I tried! Also, the new EfficientNet pre-trained models(PyTorch) with AutoAugment are doing worse than the Standard ones. I must be doing something very wrong 🙁",
    "597505": "Given Kaggle GPU constraints, it is bit difficult to train B7 on it. But you can experiment with B6 and Good news is you will definitely get better score with B6 as compare to B5 :)",
    "597573": "Thank you for posting. One important thing about `EfficientNet ` models that they are designed to take in to account input image dimensions.\n\nSo if you want to squeeze every last droplet from your model make sure to use same image resolutions =)\n\n```\nefficientnet-b0-224\nefficientnet-b1-240\nefficientnet-b2-260\nefficientnet-b3-300\nefficientnet-b4-380\nefficientnet-b5-456\nefficientnet-b6-528\nefficientnet-b7-600\n```\nYou can get more info about the design and why resolution matters from the official paper:  https://arxiv.org/abs/1905.11946",
    "597703": "Great comment, I wasn't aware of this, thanks!",
    "597983": "I wasn't are of this, too. Thank you for sharing!",
    "598237": "Hi Tom, it seems that Keras B6/7 weights hasn't come out yet, so I will to wait a bit :)\n\n**EDIT** For keras user, \nIt appears at the moment that if we use `!git clone https://github.com/qubvel/efficientnet.git`\ninstead of `!pip3 install efficientnet`, we will get the latest version which has new B6/B7 implementation and pretrained weights ... :)\n\n**UPDATED** Hi Tom @taindow, I cannot make B6 converges nicely in 9-hour kernel (I also faced some technical difficulty in order to continue training in another kernel) ... I think it is promising if we can train on a local machine ... However, since I am focusing only on kernel-only training, I will move to other approaches :)",
    "598316": "Nice! Thanks for sharing! Remember to replace the Batch Normalization layers if you are using very small batch sizes with B6 and B7. [Group Normalization](https://arxiv.org/pdf/1803.08494.pdf) can be a great alternative to Batch Normalization in this case.",
    "598677": "I do believe that bigger architecture should give a better accuracy (B6/B7) also taking into account the image size that goes along with it to get the maximum , which also means that we would have to use a very small batch size even with mixed precision (FP16) for inference. Which makes it even more time consuming (with all the preprocessing that we already do) to predict for the private test set. Did anyone run an inference kernel on B6/B7 yet ? If so did you make a note of how long it took ? Please share the details.",
    "598763": "Actually image size may matter MORE for those architectures that weren't designed to take image size into account, as their budget for width, depth and image size are less balanced. I think it's ok as long as your image size is similar to the recommended ones, those sizes are just \"efficient\" with respect to computational power.",
    "598947": "Larger image size does not seem to work for me. It takes too long to train.",
    "599175": "How can you remove the Batch Normalization layers?",
    "600417": "But with b6(and 500+ image size) what's the biggest batch size for a kaggle kernel? 8?",
    "600468": "Thank you very much. \nI read the paper &amp; also knew that, resolution is part of their \"compounding scaling\" method. Then also didn't got idea, resolution is important for this model.",
    "600471": "May be you're trying without changing resolution.\nSee DrHB's above [this](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/103743#597573) comment.\nAfter that, problem might get solved. Hope so.",
    "600524": "I'm very much aware of that and have been using matching resolution since my 1st submission with EfficientNets. It is to be noted that with increasing resolution, batch size decreases so much that the training takes ages to reach a reasonable score on the Validation set. I've to train longer maybe!",
    "600720": "True. I also need to halve my batch size &amp; epochs to fit in 9 hour limit.",
    "601192": "Surprisingly, after using 300 as the dimension for training B3, my LB score dropped about 2 %. My best score is .794 using B3 and image size 224. Used @drhabib 's method for training both networks and used the same parameters for training.  What am I doing wrong? Is there any other trick I should use?",
    "601217": "Same here...",
    "601219": "Have you decreased batch size for training with higher resolution?",
    "601249": "yes",
    "601298": "tahsin, same here. \nReason might be overfitting. As public is just 15%, can't rely on public LB.",
    "602804": "Take a look at this kernel: [EfficientNetB5 with Keras (APTOS 2019)](https://www.kaggle.com/carlolepelaars/efficientnetb5-with-keras-aptos-2019) from input line 15 to 17.\nThe code is super clear."
  },
  "source": "meta"
}