{
  "id": 203095,
  "title": "Correct way to unfreeze EfficientNet weights",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/203095",
  "author_name": "",
  "post_date": "2020-12-13T17:37:01.276878400Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I see a lot of notebooks using EfficientNet, most of them starting the training from scratch rather than freezing the earlier convolutional layers but if at all at a later stage you decide to do fine-tuning on a publicly shared model this is the correct way to do it, please make sure you do not <strong>unfreeze the BatchNormalization layer</strong>.</p>\n<p>The reason for this is that the BN layer consist of two non trainable parameters mean and variance and if you set the layer to trainable as well the updates applied to the non-trainable weights will suddenly destroy what the model has learned.</p>",
  "messages": [
    {
      "id": "1111406",
      "postDate": "12/13/2020 17:37:01",
      "content": "<p>I see a lot of notebooks using EfficientNet, most of them starting the training from scratch rather than freezing the earlier convolutional layers but if at all at a later stage you decide to do fine-tuning on a publicly shared model this is the correct way to do it, please make sure you do not <strong>unfreeze the BatchNormalization layer</strong>.</p>\n<p>The reason for this is that the BN layer consist of two non trainable parameters mean and variance and if you set the layer to trainable as well the updates applied to the non-trainable weights will suddenly destroy what the model has learned.</p>",
      "rawMarkdown": "I see a lot of notebooks using EfficientNet, most of them starting the training from scratch rather than freezing the earlier convolutional layers but if at all at a later stage you decide to do fine-tuning on a publicly shared model this is the correct way to do it, please make sure you do not **unfreeze the BatchNormalization layer**.\n\nThe reason for this is that the BN layer consist of two non trainable parameters mean and variance and if you set the layer to trainable as well the updates applied to the non-trainable weights will suddenly destroy what the model has learned.",
      "votes": null
    },
    {
      "id": "1112818",
      "postDate": "12/14/2020 23:15:20",
      "content": "<p>Thanks. I came across this article that seems to back up what you are saying. Are you suggesting to freeze earlier layers just for a few epochs and then unfreeze but apply differential learning rates?</p>\n<p><a href=\"https://towardsdatascience.com/transfer-learning-using-differential-learning-rates-638455797f00\" target=\"_blank\">https://towardsdatascience.com/transfer-learning-using-differential-learning-rates-638455797f00</a></p>",
      "rawMarkdown": "Thanks. I came across this article that seems to back up what you are saying. Are you suggesting to freeze earlier layers just for a few epochs and then unfreeze but apply differential learning rates?\n\nhttps://towardsdatascience.com/transfer-learning-using-differential-learning-rates-638455797f00",
      "votes": null
    },
    {
      "id": "1113340",
      "postDate": "12/15/2020 11:31:24",
      "content": "<p>You have some code in your shared kernels - commented out - that appeared to be the way you handled the BN layers.  For some reason the \"If NOT\" did not work.  </p>\n<p>My unfreeze for all layers but the BN was to set the base model to true <br>\nand than loop thru the layers and changing the BN back to False.</p>\n<p>On my system I have almost zero success in inserting code into these posts - so the crap below is the best I can get for my code !!</p>\n<p>base_efficientnet.trainable = True<br>\nfor layer in base_efficientnet.layers:</p>\n<pre><code>if isinstance(layer, tf.keras.layers.BatchNormalization): \n    layer.trainable = False\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "You have some code in your shared kernels - commented out - that appeared to be the way you handled the BN layers.  For some reason the \"If NOT\" did not work.  \n\nMy unfreeze for all layers but the BN was to set the base model to true \nand than loop thru the layers and changing the BN back to False.\n\nOn my system I have almost zero success in inserting code into these posts - so the crap below is the best I can get for my code !!\n\nbase_efficientnet.trainable = True\nfor layer in base_efficientnet.layers:\n\n    if isinstance(layer, tf.keras.layers.BatchNormalization): \n        layer.trainable = False\n\n\n`",
      "votes": null
    },
    {
      "id": "1113384",
      "postDate": "12/15/2020 12:08:56",
      "content": "<p>Hi Matthew, I wanted to convey that if you are training the architecture from scratch using image net or noisy student weights that is ok but if you are doing finetuning on a public model, then the correct way to go about is you freeze the earlier layers of CNN and let the later layers learn but even in the later layers you should not keep batch normalization on.</p>",
      "rawMarkdown": "Hi Matthew, I wanted to convey that if you are training the architecture from scratch using image net or noisy student weights that is ok but if you are doing finetuning on a public model, then the correct way to go about is you freeze the earlier layers of CNN and let the later layers learn but even in the later layers you should not keep batch normalization on.",
      "votes": null
    },
    {
      "id": "1113407",
      "postDate": "12/15/2020 12:23:01",
      "content": "<p>Hi,<br>\nWhen you set the entire network to trainable, it doesn't matters. This freezing of BN layers only matters when you are finetuning according to what I have read on Keras official documentation.</p>",
      "rawMarkdown": "Hi,\nWhen you set the entire network to trainable, it doesn't matters. This freezing of BN layers only matters when you are finetuning according to what I have read on Keras official documentation.",
      "votes": null
    },
    {
      "id": "1113661",
      "postDate": "12/15/2020 16:07:13",
      "content": "<p>Thanks for the valuable insight <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a>.</p>",
      "rawMarkdown": "Thanks for the valuable insight @harveenchadha.",
      "votes": null
    },
    {
      "id": "1113897",
      "postDate": "12/15/2020 19:13:05",
      "content": "<p>Guess I confused things with my post.  I was not commenting on what layers to freeze or when to freeze them but only indicating that your code</p>\n<p><code>if not isinstance(layer, tf.keras.layers.BatchNormalization):</code></p>\n<p>did not work on my local system - the \"if not\" did not correctly identify the BN layers.  But I could see the BN layers using \"if\".</p>",
      "rawMarkdown": "Guess I confused things with my post.  I was not commenting on what layers to freeze or when to freeze them but only indicating that your code\n\n`if not isinstance(layer, tf.keras.layers.BatchNormalization): `\n\ndid not work on my local system - the \"if not\" did not correctly identify the BN layers.  But I could see the BN layers using \"if\".",
      "votes": null
    },
    {
      "id": "1115209",
      "postDate": "12/16/2020 04:53:34",
      "content": "<p>We indeed want our model to learn this data distribution, why shouldn't we not unfreeze the BN layer</p>",
      "rawMarkdown": "We indeed want our model to learn this data distribution, why shouldn't we not unfreeze the BN layer",
      "votes": null
    },
    {
      "id": "1115217",
      "postDate": "12/16/2020 05:06:38",
      "content": "<p>Please read this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F3a57c930a3d0027cd1c6b333e0db4823%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095184111587&amp;alt=media\" alt=\"Pic\"></p>",
      "rawMarkdown": "Please read this:\n\n![Pic](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F3a57c930a3d0027cd1c6b333e0db4823%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095184111587&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1112818,
      "author_name": "kagglethomas88",
      "author_url": "",
      "post_date": "12/14/2020 23:15:20",
      "content": "<p>Thanks. I came across this article that seems to back up what you are saying. Are you suggesting to freeze earlier layers just for a few epochs and then unfreeze but apply differential learning rates?</p>\n<p><a href=\"https://towardsdatascience.com/transfer-learning-using-differential-learning-rates-638455797f00\" target=\"_blank\">https://towardsdatascience.com/transfer-learning-using-differential-learning-rates-638455797f00</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1113384,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "12/15/2020 12:08:56",
          "content": "<p>Hi Matthew, I wanted to convey that if you are training the architecture from scratch using image net or noisy student weights that is ok but if you are doing finetuning on a public model, then the correct way to go about is you freeze the earlier layers of CNN and let the later layers learn but even in the later layers you should not keep batch normalization on.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1113340,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "12/15/2020 11:31:24",
      "content": "<p>You have some code in your shared kernels - commented out - that appeared to be the way you handled the BN layers.  For some reason the \"If NOT\" did not work.  </p>\n<p>My unfreeze for all layers but the BN was to set the base model to true <br>\nand than loop thru the layers and changing the BN back to False.</p>\n<p>On my system I have almost zero success in inserting code into these posts - so the crap below is the best I can get for my code !!</p>\n<p>base_efficientnet.trainable = True<br>\nfor layer in base_efficientnet.layers:</p>\n<pre><code>if isinstance(layer, tf.keras.layers.BatchNormalization): \n    layer.trainable = False\n</code></pre>\n<p>`</p>",
      "votes": null,
      "replies": [
        {
          "id": 1113407,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "12/15/2020 12:23:01",
          "content": "<p>Hi,<br>\nWhen you set the entire network to trainable, it doesn't matters. This freezing of BN layers only matters when you are finetuning according to what I have read on Keras official documentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1113897,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/15/2020 19:13:05",
          "content": "<p>Guess I confused things with my post.  I was not commenting on what layers to freeze or when to freeze them but only indicating that your code</p>\n<p><code>if not isinstance(layer, tf.keras.layers.BatchNormalization):</code></p>\n<p>did not work on my local system - the \"if not\" did not correctly identify the BN layers.  But I could see the BN layers using \"if\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1113661,
      "author_name": "puzuwe",
      "author_url": "",
      "post_date": "12/15/2020 16:07:13",
      "content": "<p>Thanks for the valuable insight <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115209,
      "author_name": "gvsaikumar",
      "author_url": "",
      "post_date": "12/16/2020 04:53:34",
      "content": "<p>We indeed want our model to learn this data distribution, why shouldn't we not unfreeze the BN layer</p>",
      "votes": null,
      "replies": [
        {
          "id": 1115217,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "12/16/2020 05:06:38",
          "content": "<p>Please read this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F3a57c930a3d0027cd1c6b333e0db4823%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095184111587&amp;alt=media\" alt=\"Pic\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1111406": "I see a lot of notebooks using EfficientNet, most of them starting the training from scratch rather than freezing the earlier convolutional layers but if at all at a later stage you decide to do fine-tuning on a publicly shared model this is the correct way to do it, please make sure you do not **unfreeze the BatchNormalization layer**.\n\nThe reason for this is that the BN layer consist of two non trainable parameters mean and variance and if you set the layer to trainable as well the updates applied to the non-trainable weights will suddenly destroy what the model has learned.",
    "1112818": "Thanks. I came across this article that seems to back up what you are saying. Are you suggesting to freeze earlier layers just for a few epochs and then unfreeze but apply differential learning rates?\n\nhttps://towardsdatascience.com/transfer-learning-using-differential-learning-rates-638455797f00",
    "1113340": "You have some code in your shared kernels - commented out - that appeared to be the way you handled the BN layers.  For some reason the \"If NOT\" did not work.  \n\nMy unfreeze for all layers but the BN was to set the base model to true \nand than loop thru the layers and changing the BN back to False.\n\nOn my system I have almost zero success in inserting code into these posts - so the crap below is the best I can get for my code !!\n\nbase_efficientnet.trainable = True\nfor layer in base_efficientnet.layers:\n\n    if isinstance(layer, tf.keras.layers.BatchNormalization): \n        layer.trainable = False\n\n\n`",
    "1113384": "Hi Matthew, I wanted to convey that if you are training the architecture from scratch using image net or noisy student weights that is ok but if you are doing finetuning on a public model, then the correct way to go about is you freeze the earlier layers of CNN and let the later layers learn but even in the later layers you should not keep batch normalization on.",
    "1113407": "Hi,\nWhen you set the entire network to trainable, it doesn't matters. This freezing of BN layers only matters when you are finetuning according to what I have read on Keras official documentation.",
    "1113661": "Thanks for the valuable insight @harveenchadha.",
    "1113897": "Guess I confused things with my post.  I was not commenting on what layers to freeze or when to freeze them but only indicating that your code\n\n`if not isinstance(layer, tf.keras.layers.BatchNormalization): `\n\ndid not work on my local system - the \"if not\" did not correctly identify the BN layers.  But I could see the BN layers using \"if\".",
    "1115209": "We indeed want our model to learn this data distribution, why shouldn't we not unfreeze the BN layer",
    "1115217": "Please read this:\n\n![Pic](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1295006%2F3a57c930a3d0027cd1c6b333e0db4823%2FScreenshot%202020-12-16%20at%2010.33.41%20AM.png?generation=1608095184111587&alt=media)"
  },
  "source": "meta"
}