{
  "id": 202720,
  "title": "EfficicentNet Input Size / Keras Question",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202720",
  "author_name": "",
  "post_date": "2020-12-11T15:21:37.701112Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi Marching Learning  expert,<br>\nI implemented my model using pretrained EfficientNetB4 and have a  input layer, and  augmentation layer with input shape of (380, 380, 3) as suggested. </p>\n<pre><code>base_mode = EfficientNetB4(...)\nmodel = tf.keras.Sequential()\nmodel.add (tf.keras.Input(shape=IMAGE_SHAPE) )\nmodel.add(img_augmentation)\n\nmodel.add(base_model)\n</code></pre>\n<p>However,  without any change,  I can feed the model with images size of (400, 400, 3), or any other sizes (like (512, 512, 3)).   Why is that?   is there a automatic resize layer?  Just curious.  </p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1109352",
      "postDate": "12/11/2020 15:21:37",
      "content": "<p>Hi Marching Learning  expert,<br>\nI implemented my model using pretrained EfficientNetB4 and have a  input layer, and  augmentation layer with input shape of (380, 380, 3) as suggested. </p>\n<pre><code>base_mode = EfficientNetB4(...)\nmodel = tf.keras.Sequential()\nmodel.add (tf.keras.Input(shape=IMAGE_SHAPE) )\nmodel.add(img_augmentation)\n\nmodel.add(base_model)\n</code></pre>\n<p>However,  without any change,  I can feed the model with images size of (400, 400, 3), or any other sizes (like (512, 512, 3)).   Why is that?   is there a automatic resize layer?  Just curious.  </p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi Marching Learning  expert,\nI implemented my model using pretrained EfficientNetB4 and have a  input layer, and  augmentation layer with input shape of (380, 380, 3) as suggested. \n\n    base_mode = EfficientNetB4(...)\n    model = tf.keras.Sequential()\n    model.add (tf.keras.Input(shape=IMAGE_SHAPE) )\n    model.add(img_augmentation)\n\n    model.add(base_model)\n\nHowever,  without any change,  I can feed the model with images size of (400, 400, 3), or any other sizes (like (512, 512, 3)).   Why is that?   is there a automatic resize layer?  Just curious.  \n\nThanks.",
      "votes": null
    },
    {
      "id": "1109436",
      "postDate": "12/11/2020 17:13:34",
      "content": "<p>The CNN is fully convolutional, whatever input size you feed the CNN, the convolution layers will output a feature map based on your input size. Large input size = bigger feature map size.<br>\nI think you will find your answer in this topic: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">CNN input size explained.</a></p>",
      "rawMarkdown": "The CNN is fully convolutional, whatever input size you feed the CNN, the convolution layers will output a feature map based on your input size. Large input size = bigger feature map size.\nI think you will find your answer in this topic: [CNN input size explained.](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147)",
      "votes": null
    },
    {
      "id": "1109459",
      "postDate": "12/11/2020 17:44:02",
      "content": "<p>No, the efficientnet model does not resize the input images. It's just that different input image sizes yield different output feature sizes. For instance, efficientnetb0 with an input image size (512, 512, 3) yields a total feature set of (16, 16, 1280) whereas the same model with an input image size of (256, 256, 3) yields a total feature set of (8, 8, 1280). The feature set is actually 1280 squares of 8x8 pix (or 16x16 in the former case). Afterwards, these features are usually passed into a globalaveragepooling2d (outputs the average of the value in a square) or a globalmaxpooling2d (outputs the maximum of the values in a square) layer to reduce the squares into a 1d array of 1280 features.</p>",
      "rawMarkdown": "No, the efficientnet model does not resize the input images. It's just that different input image sizes yield different output feature sizes. For instance, efficientnetb0 with an input image size (512, 512, 3) yields a total feature set of (16, 16, 1280) whereas the same model with an input image size of (256, 256, 3) yields a total feature set of (8, 8, 1280). The feature set is actually 1280 squares of 8x8 pix (or 16x16 in the former case). Afterwards, these features are usually passed into a globalaveragepooling2d (outputs the average of the value in a square) or a globalmaxpooling2d (outputs the maximum of the values in a square) layer to reduce the squares into a 1d array of 1280 features.",
      "votes": null
    },
    {
      "id": "1109620",
      "postDate": "12/11/2020 22:15:31",
      "content": "<p>Thanks for the reply. The linked Notebook had great explanation on the exact question I have. It is really helpful to understand how CNN work.</p>",
      "rawMarkdown": "Thanks for the reply. The linked Notebook had great explanation on the exact question I have. It is really helpful to understand how CNN work.",
      "votes": null
    },
    {
      "id": "1109627",
      "postDate": "12/11/2020 22:26:49",
      "content": "<p>Thanks for the reply.   A follow up question would be do think I should train the Efficientnet with different Sizes or just the recommended input size of 380?</p>",
      "rawMarkdown": "Thanks for the reply.   A follow up question would be do think I should train the Efficientnet with different Sizes or just the recommended input size of 380?",
      "votes": null
    },
    {
      "id": "1109655",
      "postDate": "12/11/2020 23:27:03",
      "content": "<p>There is no definite answer for your question. Try and see what works better.</p>",
      "rawMarkdown": "There is no definite answer for your question. Try and see what works better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1109436,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "12/11/2020 17:13:34",
      "content": "<p>The CNN is fully convolutional, whatever input size you feed the CNN, the convolution layers will output a feature map based on your input size. Large input size = bigger feature map size.<br>\nI think you will find your answer in this topic: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">CNN input size explained.</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1109620,
          "author_name": "luqing2",
          "author_url": "",
          "post_date": "12/11/2020 22:15:31",
          "content": "<p>Thanks for the reply. The linked Notebook had great explanation on the exact question I have. It is really helpful to understand how CNN work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109459,
      "author_name": "tolgadincer",
      "author_url": "",
      "post_date": "12/11/2020 17:44:02",
      "content": "<p>No, the efficientnet model does not resize the input images. It's just that different input image sizes yield different output feature sizes. For instance, efficientnetb0 with an input image size (512, 512, 3) yields a total feature set of (16, 16, 1280) whereas the same model with an input image size of (256, 256, 3) yields a total feature set of (8, 8, 1280). The feature set is actually 1280 squares of 8x8 pix (or 16x16 in the former case). Afterwards, these features are usually passed into a globalaveragepooling2d (outputs the average of the value in a square) or a globalmaxpooling2d (outputs the maximum of the values in a square) layer to reduce the squares into a 1d array of 1280 features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109627,
          "author_name": "luqing2",
          "author_url": "",
          "post_date": "12/11/2020 22:26:49",
          "content": "<p>Thanks for the reply.   A follow up question would be do think I should train the Efficientnet with different Sizes or just the recommended input size of 380?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109655,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "12/11/2020 23:27:03",
          "content": "<p>There is no definite answer for your question. Try and see what works better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1109352": "Hi Marching Learning  expert,\nI implemented my model using pretrained EfficientNetB4 and have a  input layer, and  augmentation layer with input shape of (380, 380, 3) as suggested. \n\n    base_mode = EfficientNetB4(...)\n    model = tf.keras.Sequential()\n    model.add (tf.keras.Input(shape=IMAGE_SHAPE) )\n    model.add(img_augmentation)\n\n    model.add(base_model)\n\nHowever,  without any change,  I can feed the model with images size of (400, 400, 3), or any other sizes (like (512, 512, 3)).   Why is that?   is there a automatic resize layer?  Just curious.  \n\nThanks.",
    "1109436": "The CNN is fully convolutional, whatever input size you feed the CNN, the convolution layers will output a feature map based on your input size. Large input size = bigger feature map size.\nI think you will find your answer in this topic: [CNN input size explained.](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147)",
    "1109459": "No, the efficientnet model does not resize the input images. It's just that different input image sizes yield different output feature sizes. For instance, efficientnetb0 with an input image size (512, 512, 3) yields a total feature set of (16, 16, 1280) whereas the same model with an input image size of (256, 256, 3) yields a total feature set of (8, 8, 1280). The feature set is actually 1280 squares of 8x8 pix (or 16x16 in the former case). Afterwards, these features are usually passed into a globalaveragepooling2d (outputs the average of the value in a square) or a globalmaxpooling2d (outputs the maximum of the values in a square) layer to reduce the squares into a 1d array of 1280 features.",
    "1109620": "Thanks for the reply. The linked Notebook had great explanation on the exact question I have. It is really helpful to understand how CNN work.",
    "1109627": "Thanks for the reply.   A follow up question would be do think I should train the Efficientnet with different Sizes or just the recommended input size of 380?",
    "1109655": "There is no definite answer for your question. Try and see what works better."
  },
  "source": "meta"
}