{
  "id": 207744,
  "title": "Using different Image Size for Transfer Learning (EfficientNet)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207744",
  "author_name": "",
  "post_date": "2020-12-31T04:58:16.800689300Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Update: <a href=\"https://www.pyimagesearch.com/2019/06/24/change-input-shape-dimensions-for-fine-tuning-with-keras/\" target=\"_blank\">This link here might explain slightly</a></p>\n<p><a href=\"https://stackoverflow.com/questions/46036522/defining-model-in-keras-include-top-true\" target=\"_blank\">This link is important to read!</a></p>\n<p>I have made a switch from TensorFlow to PyTorch recently. I use a famous <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch/tree/master/geffnet\" target=\"_blank\">Github repo</a> for training on <code>EfficientNets</code>. I wrote the model initiation class as follows:</p>\n<pre><code>    class CustomEfficientNet(nn.Module):\n        def __init__(self, config: type, pretrained: bool=True):\n            super().__init__()\n            self.config = config\n            self.model = geffnet.create_model(\n                model_name='EfficientNetB5',\n                pretrained=pretrained)\n            n_features = self.model.classifier.in_features\n            self.model.classifier = nn.Linear(n_features, num_classes=5)\n\n\n        def forward(self, input_neurons):\n            output_predictions = self.model(input_neurons)\n            return output_predictions\n</code></pre>\n<p>In addition, in my <code>transforms</code>, I tend to use <code>Resize(img_size = 512, img_size=512)</code> for my training. So the question here is, the official input size for <code>EfficientNetB5</code> is 456x456, but I used 512x512 or even 256x256 and get very decent results. Is this normal? Or did I miss out on the source code where the author will resize into the native resolution for you?</p>\n<p>PS: This seems to be the norm in all the PyTorch Tutorials I saw on Kaggle. My full code can be seen here in this <a href=\"https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline/\" target=\"_blank\">notebook</a> here in the event if anyone wants to explain to me my question (I really did not resize the images to the native resolution!!!); Also this is not a promotion of my content, I just want to clarify everything possible as leaving logic gaps is a no go for me. </p>",
  "messages": [
    {
      "id": "1133285",
      "postDate": "12/31/2020 04:58:16",
      "content": "<p>Update: <a href=\"https://www.pyimagesearch.com/2019/06/24/change-input-shape-dimensions-for-fine-tuning-with-keras/\" target=\"_blank\">This link here might explain slightly</a></p>\n<p><a href=\"https://stackoverflow.com/questions/46036522/defining-model-in-keras-include-top-true\" target=\"_blank\">This link is important to read!</a></p>\n<p>I have made a switch from TensorFlow to PyTorch recently. I use a famous <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch/tree/master/geffnet\" target=\"_blank\">Github repo</a> for training on <code>EfficientNets</code>. I wrote the model initiation class as follows:</p>\n<pre><code>    class CustomEfficientNet(nn.Module):\n        def __init__(self, config: type, pretrained: bool=True):\n            super().__init__()\n            self.config = config\n            self.model = geffnet.create_model(\n                model_name='EfficientNetB5',\n                pretrained=pretrained)\n            n_features = self.model.classifier.in_features\n            self.model.classifier = nn.Linear(n_features, num_classes=5)\n\n\n        def forward(self, input_neurons):\n            output_predictions = self.model(input_neurons)\n            return output_predictions\n</code></pre>\n<p>In addition, in my <code>transforms</code>, I tend to use <code>Resize(img_size = 512, img_size=512)</code> for my training. So the question here is, the official input size for <code>EfficientNetB5</code> is 456x456, but I used 512x512 or even 256x256 and get very decent results. Is this normal? Or did I miss out on the source code where the author will resize into the native resolution for you?</p>\n<p>PS: This seems to be the norm in all the PyTorch Tutorials I saw on Kaggle. My full code can be seen here in this <a href=\"https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline/\" target=\"_blank\">notebook</a> here in the event if anyone wants to explain to me my question (I really did not resize the images to the native resolution!!!); Also this is not a promotion of my content, I just want to clarify everything possible as leaving logic gaps is a no go for me. </p>",
      "rawMarkdown": "Update: [This link here might explain slightly](https://www.pyimagesearch.com/2019/06/24/change-input-shape-dimensions-for-fine-tuning-with-keras/)\n\n[This link is important to read!](https://stackoverflow.com/questions/46036522/defining-model-in-keras-include-top-true)\n\nI have made a switch from TensorFlow to PyTorch recently. I use a famous [Github repo](https://github.com/rwightman/gen-efficientnet-pytorch/tree/master/geffnet) for training on `EfficientNets`. I wrote the model initiation class as follows:\n\n```\n    class CustomEfficientNet(nn.Module):\n        def __init__(self, config: type, pretrained: bool=True):\n            super().__init__()\n            self.config = config\n            self.model = geffnet.create_model(\n                model_name='EfficientNetB5',\n                pretrained=pretrained)\n            n_features = self.model.classifier.in_features\n            self.model.classifier = nn.Linear(n_features, num_classes=5)\n            \n    \n        def forward(self, input_neurons):\n            output_predictions = self.model(input_neurons)\n            return output_predictions\n```\n\nIn addition, in my `transforms`, I tend to use `Resize(img_size = 512, img_size=512)` for my training. So the question here is, the official input size for `EfficientNetB5` is 456x456, but I used 512x512 or even 256x256 and get very decent results. Is this normal? Or did I miss out on the source code where the author will resize into the native resolution for you?\n\nPS: This seems to be the norm in all the PyTorch Tutorials I saw on Kaggle. My full code can be seen here in this [notebook](https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline/) here in the event if anyone wants to explain to me my question (I really did not resize the images to the native resolution!!!); Also this is not a promotion of my content, I just want to clarify everything possible as leaving logic gaps is a no go for me.",
      "votes": null
    },
    {
      "id": "1133304",
      "postDate": "12/31/2020 05:18:11",
      "content": "<p>I think it's the global pooling layer to make efficientnet can apply to any input size image.<br>\nAlso check this paper, <a href=\"https://arxiv.org/abs/1906.06423\" target=\"_blank\">https://arxiv.org/abs/1906.06423</a>, hope this can help you.</p>",
      "rawMarkdown": "I think it's the global pooling layer to make efficientnet can apply to any input size image.\nAlso check this paper, https://arxiv.org/abs/1906.06423, hope this can help you.",
      "votes": null
    },
    {
      "id": "1133424",
      "postDate": "12/31/2020 07:47:03",
      "content": "<p>Start training with a smaller size (not too small), and then fine-tune it to train only the fully connected layer, which can greatly shorten the training time to obtain the best results. Of course, there are differences in the image enhancement used. For example, as mentioned in the paper, the author uses RandomResizedCrop during the training phase, and uses CenterCrop when fine-tuning. I don't know if I understand it right?</p>",
      "rawMarkdown": "Start training with a smaller size (not too small), and then fine-tune it to train only the fully connected layer, which can greatly shorten the training time to obtain the best results. Of course, there are differences in the image enhancement used. For example, as mentioned in the paper, the author uses RandomResizedCrop during the training phase, and uses CenterCrop when fine-tuning. I don't know if I understand it right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1133304,
      "author_name": "newbiejailer",
      "author_url": "",
      "post_date": "12/31/2020 05:18:11",
      "content": "<p>I think it's the global pooling layer to make efficientnet can apply to any input size image.<br>\nAlso check this paper, <a href=\"https://arxiv.org/abs/1906.06423\" target=\"_blank\">https://arxiv.org/abs/1906.06423</a>, hope this can help you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1133424,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/31/2020 07:47:03",
      "content": "<p>Start training with a smaller size (not too small), and then fine-tune it to train only the fully connected layer, which can greatly shorten the training time to obtain the best results. Of course, there are differences in the image enhancement used. For example, as mentioned in the paper, the author uses RandomResizedCrop during the training phase, and uses CenterCrop when fine-tuning. I don't know if I understand it right?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1133285": "Update: [This link here might explain slightly](https://www.pyimagesearch.com/2019/06/24/change-input-shape-dimensions-for-fine-tuning-with-keras/)\n\n[This link is important to read!](https://stackoverflow.com/questions/46036522/defining-model-in-keras-include-top-true)\n\nI have made a switch from TensorFlow to PyTorch recently. I use a famous [Github repo](https://github.com/rwightman/gen-efficientnet-pytorch/tree/master/geffnet) for training on `EfficientNets`. I wrote the model initiation class as follows:\n\n```\n    class CustomEfficientNet(nn.Module):\n        def __init__(self, config: type, pretrained: bool=True):\n            super().__init__()\n            self.config = config\n            self.model = geffnet.create_model(\n                model_name='EfficientNetB5',\n                pretrained=pretrained)\n            n_features = self.model.classifier.in_features\n            self.model.classifier = nn.Linear(n_features, num_classes=5)\n            \n    \n        def forward(self, input_neurons):\n            output_predictions = self.model(input_neurons)\n            return output_predictions\n```\n\nIn addition, in my `transforms`, I tend to use `Resize(img_size = 512, img_size=512)` for my training. So the question here is, the official input size for `EfficientNetB5` is 456x456, but I used 512x512 or even 256x256 and get very decent results. Is this normal? Or did I miss out on the source code where the author will resize into the native resolution for you?\n\nPS: This seems to be the norm in all the PyTorch Tutorials I saw on Kaggle. My full code can be seen here in this [notebook](https://www.kaggle.com/reighns/complete-and-reusable-pytorch-pipeline/) here in the event if anyone wants to explain to me my question (I really did not resize the images to the native resolution!!!); Also this is not a promotion of my content, I just want to clarify everything possible as leaving logic gaps is a no go for me.",
    "1133304": "I think it's the global pooling layer to make efficientnet can apply to any input size image.\nAlso check this paper, https://arxiv.org/abs/1906.06423, hope this can help you.",
    "1133424": "Start training with a smaller size (not too small), and then fine-tune it to train only the fully connected layer, which can greatly shorten the training time to obtain the best results. Of course, there are differences in the image enhancement used. For example, as mentioned in the paper, the author uses RandomResizedCrop during the training phase, and uses CenterCrop when fine-tuning. I don't know if I understand it right?"
  },
  "source": "meta"
}