{
  "id": 213545,
  "title": "Help! Efficientnet doesn't train properly if tensor is converted to float.",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/213545",
  "author_name": "",
  "post_date": "2021-01-23T09:21:01.870678Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>If I apply the following augmentations to my images as shown:</p>\n<p><code>train_augmentations = A.Compose([\n            A.RandomCrop(image_size, image_size, p=1),\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),\n            A.Flip(p=0.5),\n            A.ShiftScaleRotate(p=0.5),\n            A.HueSaturationValue(p=0.5, hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2),\n            A.RandomBrightnessContrast(p=0.5, brightness_limit=(-0.2,0.2), contrast_limit=(-0.2, 0.2)),\n            A.ToFloat()\n            ], p=1)</code></p>\n<p>My efficientnet doesn't train properly, train accuracy goes up but val accuracy just gets stuck at 0.1+. But if I remove the last augmentation, ToFloat(), then it trains fine. Could someone please explain to me why? I didn't have this problem with Xception/Inception.. </p>\n<p>Edit: I believe the problem might not lie with converting to float itself, but rather the rescaling by 1/255 when the transformation ToFloat() is applied.. Still don't understand why though.</p>",
  "messages": [
    {
      "id": "1165852",
      "postDate": "01/23/2021 09:21:01",
      "content": "<p>If I apply the following augmentations to my images as shown:</p>\n<p><code>train_augmentations = A.Compose([\n            A.RandomCrop(image_size, image_size, p=1),\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),\n            A.Flip(p=0.5),\n            A.ShiftScaleRotate(p=0.5),\n            A.HueSaturationValue(p=0.5, hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2),\n            A.RandomBrightnessContrast(p=0.5, brightness_limit=(-0.2,0.2), contrast_limit=(-0.2, 0.2)),\n            A.ToFloat()\n            ], p=1)</code></p>\n<p>My efficientnet doesn't train properly, train accuracy goes up but val accuracy just gets stuck at 0.1+. But if I remove the last augmentation, ToFloat(), then it trains fine. Could someone please explain to me why? I didn't have this problem with Xception/Inception.. </p>\n<p>Edit: I believe the problem might not lie with converting to float itself, but rather the rescaling by 1/255 when the transformation ToFloat() is applied.. Still don't understand why though.</p>",
      "rawMarkdown": "If I apply the following augmentations to my images as shown:\n\n`train_augmentations = A.Compose([\n            A.RandomCrop(image_size, image_size, p=1),\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),\n            A.Flip(p=0.5),\n            A.ShiftScaleRotate(p=0.5),\n            A.HueSaturationValue(p=0.5, hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2),\n            A.RandomBrightnessContrast(p=0.5, brightness_limit=(-0.2,0.2), contrast_limit=(-0.2, 0.2)),\n            A.ToFloat()\n            ], p=1)`\n\nMy efficientnet doesn't train properly, train accuracy goes up but val accuracy just gets stuck at 0.1+. But if I remove the last augmentation, ToFloat(), then it trains fine. Could someone please explain to me why? I didn't have this problem with Xception/Inception.. \n\nEdit: I believe the problem might not lie with converting to float itself, but rather the rescaling by 1/255 when the transformation ToFloat() is applied.. Still don't understand why though.",
      "votes": null
    },
    {
      "id": "1165920",
      "postDate": "01/23/2021 10:09:51",
      "content": "<p>Look at the first few layers of Efficientnet model - they contain Rescaling layer and Normalization layer - so you don't need to rescale the input going into the model.</p>\n<pre><code>```\nLayer (type)                    Output Shape         Param #     Connected to                     \n==================================================================================================\ninput_1 (InputLayer)            [(None, 512, 512, 3) 0                                            \n__________________________________________________________________________________________________\nrescaling (Rescaling)           (None, 512, 512, 3)  0           input_1[0][0]                    \n__________________________________________________________________________________________________\nnormalization (Normalization)   (None, 512, 512, 3)  7           rescaling[0][0]      \n</code></pre>\n<p>```</p>",
      "rawMarkdown": "Look at the first few layers of Efficientnet model - they contain Rescaling layer and Normalization layer - so you don't need to rescale the input going into the model.\n\n```\n```\nLayer (type)                    Output Shape         Param #     Connected to                     \n==================================================================================================\ninput_1 (InputLayer)            [(None, 512, 512, 3) 0                                            \n__________________________________________________________________________________________________\nrescaling (Rescaling)           (None, 512, 512, 3)  0           input_1[0][0]                    \n__________________________________________________________________________________________________\nnormalization (Normalization)   (None, 512, 512, 3)  7           rescaling[0][0]      \n```\n```",
      "votes": null
    },
    {
      "id": "1165931",
      "postDate": "01/23/2021 10:21:47",
      "content": "<p>Thanks Jimmy! </p>",
      "rawMarkdown": "Thanks Jimmy!",
      "votes": null
    },
    {
      "id": "1166567",
      "postDate": "01/23/2021 17:32:01",
      "content": "<p>Your welcome - sadly I learned this the hard way - spent lots of hours trying to understand why code similiar to yours did not work :)</p>\n<p>There is a post buried somewhere in this Discussion were someone pointed this out for Efficientnets.  I found the post after lots of wasted hours.</p>\n<p>I know as a good bit of advice that I should read all the recent Discussion posts every day - but I also know that I very often fail to follow my own advice. :)   </p>",
      "rawMarkdown": "Your welcome - sadly I learned this the hard way - spent lots of hours trying to understand why code similiar to yours did not work :)\n\nThere is a post buried somewhere in this Discussion were someone pointed this out for Efficientnets.  I found the post after lots of wasted hours.\n\nI know as a good bit of advice that I should read all the recent Discussion posts every day - but I also know that I very often fail to follow my own advice. :)",
      "votes": null
    },
    {
      "id": "1166926",
      "postDate": "01/23/2021 23:07:42",
      "content": "<p>Maybe a further helping comment according to my understanding… The above screenshot seems to come from the <code>tensorflow.keras.application</code> package. However if you use PyPI <code>efficientnet==1.1.1</code> (because the tensorflow version of the TPU environment does not include efficientnet yet) then this rescaling and normalization layers are not present. As far as I know, each tensorflow.keras.application (and also PyPI efficientnet package) usually have a <code>preprocess_input</code> funtion in the corresponding module that you should apply to the image with floating values in the range <code>[0., 255.]</code>.</p>",
      "rawMarkdown": "Maybe a further helping comment according to my understanding... The above screenshot seems to come from the ``tensorflow.keras.application`` package. However if you use PyPI ``efficientnet==1.1.1`` (because the tensorflow version of the TPU environment does not include efficientnet yet) then this rescaling and normalization layers are not present. As far as I know, each tensorflow.keras.application (and also PyPI efficientnet package) usually have a ``preprocess_input`` funtion in the corresponding module that you should apply to the image with floating values in the range ``[0., 255.]``.",
      "votes": null
    },
    {
      "id": "1167012",
      "postDate": "01/24/2021 01:27:36",
      "content": "<p>Thanks, that's good to know!</p>",
      "rawMarkdown": "Thanks, that's good to know!",
      "votes": null
    },
    {
      "id": "1167116",
      "postDate": "01/24/2021 04:10:46",
      "content": "<p>Thanks. I just noticed this detail.</p>",
      "rawMarkdown": "Thanks. I just noticed this detail.",
      "votes": null
    },
    {
      "id": "1167150",
      "postDate": "01/24/2021 05:05:46",
      "content": "<p><a href=\"https://www.kaggle.com/morodertobias\" target=\"_blank\">tmoroder</a></p>\n<p>Your correct - my screenshot was from the tf.keras.application version of efficientnet.  The scaling and normalization issues had me confused for almost two weeks.</p>\n<p>The lesson I learned - ALWAYS generate a model.summary and take a close look at the first several layers and the last couple.  I generated the summaries but never looked at them.  </p>",
      "rawMarkdown": "[tmoroder](https://www.kaggle.com/morodertobias)\n\n\nYour correct - my screenshot was from the tf.keras.application version of efficientnet.  The scaling and normalization issues had me confused for almost two weeks.\n\nThe lesson I learned - ALWAYS generate a model.summary and take a close look at the first several layers and the last couple.  I generated the summaries but never looked at them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1165920,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/23/2021 10:09:51",
      "content": "<p>Look at the first few layers of Efficientnet model - they contain Rescaling layer and Normalization layer - so you don't need to rescale the input going into the model.</p>\n<pre><code>```\nLayer (type)                    Output Shape         Param #     Connected to                     \n==================================================================================================\ninput_1 (InputLayer)            [(None, 512, 512, 3) 0                                            \n__________________________________________________________________________________________________\nrescaling (Rescaling)           (None, 512, 512, 3)  0           input_1[0][0]                    \n__________________________________________________________________________________________________\nnormalization (Normalization)   (None, 512, 512, 3)  7           rescaling[0][0]      \n</code></pre>\n<p>```</p>",
      "votes": null,
      "replies": [
        {
          "id": 1165931,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/23/2021 10:21:47",
          "content": "<p>Thanks Jimmy! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1166567,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/23/2021 17:32:01",
          "content": "<p>Your welcome - sadly I learned this the hard way - spent lots of hours trying to understand why code similiar to yours did not work :)</p>\n<p>There is a post buried somewhere in this Discussion were someone pointed this out for Efficientnets.  I found the post after lots of wasted hours.</p>\n<p>I know as a good bit of advice that I should read all the recent Discussion posts every day - but I also know that I very often fail to follow my own advice. :)   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1166926,
          "author_name": "morodertobias",
          "author_url": "",
          "post_date": "01/23/2021 23:07:42",
          "content": "<p>Maybe a further helping comment according to my understanding… The above screenshot seems to come from the <code>tensorflow.keras.application</code> package. However if you use PyPI <code>efficientnet==1.1.1</code> (because the tensorflow version of the TPU environment does not include efficientnet yet) then this rescaling and normalization layers are not present. As far as I know, each tensorflow.keras.application (and also PyPI efficientnet package) usually have a <code>preprocess_input</code> funtion in the corresponding module that you should apply to the image with floating values in the range <code>[0., 255.]</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1167012,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/24/2021 01:27:36",
          "content": "<p>Thanks, that's good to know!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1167116,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "01/24/2021 04:10:46",
          "content": "<p>Thanks. I just noticed this detail.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1167150,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/24/2021 05:05:46",
          "content": "<p><a href=\"https://www.kaggle.com/morodertobias\" target=\"_blank\">tmoroder</a></p>\n<p>Your correct - my screenshot was from the tf.keras.application version of efficientnet.  The scaling and normalization issues had me confused for almost two weeks.</p>\n<p>The lesson I learned - ALWAYS generate a model.summary and take a close look at the first several layers and the last couple.  I generated the summaries but never looked at them.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1165852": "If I apply the following augmentations to my images as shown:\n\n`train_augmentations = A.Compose([\n            A.RandomCrop(image_size, image_size, p=1),\n            A.CoarseDropout(p=0.5),\n            A.Cutout(p=0.5),\n            A.Flip(p=0.5),\n            A.ShiftScaleRotate(p=0.5),\n            A.HueSaturationValue(p=0.5, hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2),\n            A.RandomBrightnessContrast(p=0.5, brightness_limit=(-0.2,0.2), contrast_limit=(-0.2, 0.2)),\n            A.ToFloat()\n            ], p=1)`\n\nMy efficientnet doesn't train properly, train accuracy goes up but val accuracy just gets stuck at 0.1+. But if I remove the last augmentation, ToFloat(), then it trains fine. Could someone please explain to me why? I didn't have this problem with Xception/Inception.. \n\nEdit: I believe the problem might not lie with converting to float itself, but rather the rescaling by 1/255 when the transformation ToFloat() is applied.. Still don't understand why though.",
    "1165920": "Look at the first few layers of Efficientnet model - they contain Rescaling layer and Normalization layer - so you don't need to rescale the input going into the model.\n\n```\n```\nLayer (type)                    Output Shape         Param #     Connected to                     \n==================================================================================================\ninput_1 (InputLayer)            [(None, 512, 512, 3) 0                                            \n__________________________________________________________________________________________________\nrescaling (Rescaling)           (None, 512, 512, 3)  0           input_1[0][0]                    \n__________________________________________________________________________________________________\nnormalization (Normalization)   (None, 512, 512, 3)  7           rescaling[0][0]      \n```\n```",
    "1165931": "Thanks Jimmy!",
    "1166567": "Your welcome - sadly I learned this the hard way - spent lots of hours trying to understand why code similiar to yours did not work :)\n\nThere is a post buried somewhere in this Discussion were someone pointed this out for Efficientnets.  I found the post after lots of wasted hours.\n\nI know as a good bit of advice that I should read all the recent Discussion posts every day - but I also know that I very often fail to follow my own advice. :)",
    "1166926": "Maybe a further helping comment according to my understanding... The above screenshot seems to come from the ``tensorflow.keras.application`` package. However if you use PyPI ``efficientnet==1.1.1`` (because the tensorflow version of the TPU environment does not include efficientnet yet) then this rescaling and normalization layers are not present. As far as I know, each tensorflow.keras.application (and also PyPI efficientnet package) usually have a ``preprocess_input`` funtion in the corresponding module that you should apply to the image with floating values in the range ``[0., 255.]``.",
    "1167012": "Thanks, that's good to know!",
    "1167116": "Thanks. I just noticed this detail.",
    "1167150": "[tmoroder](https://www.kaggle.com/morodertobias)\n\n\nYour correct - my screenshot was from the tf.keras.application version of efficientnet.  The scaling and normalization issues had me confused for almost two weeks.\n\nThe lesson I learned - ALWAYS generate a model.summary and take a close look at the first several layers and the last couple.  I generated the summaries but never looked at them."
  },
  "source": "meta"
}