{
  "id": 214094,
  "title": "Is it good to tune BatchNormalization layers also while fine tuning any model?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214094",
  "author_name": "",
  "post_date": "2021-01-25T09:24:40.641515200Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I saw big difference between the score of my models on CV. Can anyone explain something about this.</p>\n<p>Community , Kindly provide your views on this.</p>",
  "messages": [
    {
      "id": "1168993",
      "postDate": "01/25/2021 09:24:40",
      "content": "<p>I saw big difference between the score of my models on CV. Can anyone explain something about this.</p>\n<p>Community , Kindly provide your views on this.</p>",
      "rawMarkdown": "I saw big difference between the score of my models on CV. Can anyone explain something about this.\n\nCommunity , Kindly provide your views on this.",
      "votes": null
    },
    {
      "id": "1169764",
      "postDate": "01/25/2021 18:16:19",
      "content": "<p>So which one is better in your case? Would be interested in this discussion question too, because I am unsure what to do with BatchNormalization layers myself. For instance this keras.io code example <a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/#tips-for-fine-tuning-efficientnet\" target=\"_blank\">Image classification via fine-tuning with EfficientNet</a> explicitly recommends to keep BatchNormalization layers non-trainable even during fine tuning (to not destroy the mean and variance estimates as far as I understand it). However I do not see this, and other techniques from this code example, applied too often in the Kaggle notebooks I look at.</p>",
      "rawMarkdown": "So which one is better in your case? Would be interested in this discussion question too, because I am unsure what to do with BatchNormalization layers myself. For instance this keras.io code example [Image classification via fine-tuning with EfficientNet](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/#tips-for-fine-tuning-efficientnet) explicitly recommends to keep BatchNormalization layers non-trainable even during fine tuning (to not destroy the mean and variance estimates as far as I understand it). However I do not see this, and other techniques from this code example, applied too often in the Kaggle notebooks I look at.",
      "votes": null
    },
    {
      "id": "1169781",
      "postDate": "01/25/2021 18:24:42",
      "content": "<p>I am experimenting on this only currently. But i haven't found major differences in my CV score , might the reason being something else.<br>\nBut , there was a bit difference between training time. Like on freezing the batchnormalization layers , the training time reduced a lot , recieving the same score on 20 epochs as it was doing in around 35-40 epochs. <br>\nAlso , the keras page says that we need to unfreeze the whole block , so i am bit confused in that too , that how to approach it .<br>\nAlthough , if i get something helpful from it , i will surely share it with the community and listen to there views too .</p>",
      "rawMarkdown": "I am experimenting on this only currently. But i haven't found major differences in my CV score , might the reason being something else.\nBut , there was a bit difference between training time. Like on freezing the batchnormalization layers , the training time reduced a lot , recieving the same score on 20 epochs as it was doing in around 35-40 epochs. \nAlso , the keras page says that we need to unfreeze the whole block , so i am bit confused in that too , that how to approach it .\nAlthough , if i get something helpful from it , i will surely share it with the community and listen to there views too .",
      "votes": null
    },
    {
      "id": "1169810",
      "postDate": "01/25/2021 18:40:50",
      "content": "<p>Thanks… The \"unfreeze the whole block\" actually does not refer to the BatchNormalitzion, but rather to EfficientNet as I understand it. If you unfreeze some layers within a block, the recommendation is to unfreeze also all remaining layers of that block (except the BatchNormalitzion).</p>",
      "rawMarkdown": "Thanks... The \"unfreeze the whole block\" actually does not refer to the BatchNormalitzion, but rather to EfficientNet as I understand it. If you unfreeze some layers within a block, the recommendation is to unfreeze also all remaining layers of that block (except the BatchNormalitzion).",
      "votes": null
    },
    {
      "id": "1169816",
      "postDate": "01/25/2021 18:44:21",
      "content": "<p>yeah currently i am doing that only. <br>\nWill see how does it works 😅😅</p>",
      "rawMarkdown": "yeah currently i am doing that only. \nWill see how does it works 😅😅",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1169764,
      "author_name": "morodertobias",
      "author_url": "",
      "post_date": "01/25/2021 18:16:19",
      "content": "<p>So which one is better in your case? Would be interested in this discussion question too, because I am unsure what to do with BatchNormalization layers myself. For instance this keras.io code example <a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/#tips-for-fine-tuning-efficientnet\" target=\"_blank\">Image classification via fine-tuning with EfficientNet</a> explicitly recommends to keep BatchNormalization layers non-trainable even during fine tuning (to not destroy the mean and variance estimates as far as I understand it). However I do not see this, and other techniques from this code example, applied too often in the Kaggle notebooks I look at.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1169781,
          "author_name": "prashantarorat",
          "author_url": "",
          "post_date": "01/25/2021 18:24:42",
          "content": "<p>I am experimenting on this only currently. But i haven't found major differences in my CV score , might the reason being something else.<br>\nBut , there was a bit difference between training time. Like on freezing the batchnormalization layers , the training time reduced a lot , recieving the same score on 20 epochs as it was doing in around 35-40 epochs. <br>\nAlso , the keras page says that we need to unfreeze the whole block , so i am bit confused in that too , that how to approach it .<br>\nAlthough , if i get something helpful from it , i will surely share it with the community and listen to there views too .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1169810,
          "author_name": "morodertobias",
          "author_url": "",
          "post_date": "01/25/2021 18:40:50",
          "content": "<p>Thanks… The \"unfreeze the whole block\" actually does not refer to the BatchNormalitzion, but rather to EfficientNet as I understand it. If you unfreeze some layers within a block, the recommendation is to unfreeze also all remaining layers of that block (except the BatchNormalitzion).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1169816,
          "author_name": "prashantarorat",
          "author_url": "",
          "post_date": "01/25/2021 18:44:21",
          "content": "<p>yeah currently i am doing that only. <br>\nWill see how does it works 😅😅</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1168993": "I saw big difference between the score of my models on CV. Can anyone explain something about this.\n\nCommunity , Kindly provide your views on this.",
    "1169764": "So which one is better in your case? Would be interested in this discussion question too, because I am unsure what to do with BatchNormalization layers myself. For instance this keras.io code example [Image classification via fine-tuning with EfficientNet](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/#tips-for-fine-tuning-efficientnet) explicitly recommends to keep BatchNormalization layers non-trainable even during fine tuning (to not destroy the mean and variance estimates as far as I understand it). However I do not see this, and other techniques from this code example, applied too often in the Kaggle notebooks I look at.",
    "1169781": "I am experimenting on this only currently. But i haven't found major differences in my CV score , might the reason being something else.\nBut , there was a bit difference between training time. Like on freezing the batchnormalization layers , the training time reduced a lot , recieving the same score on 20 epochs as it was doing in around 35-40 epochs. \nAlso , the keras page says that we need to unfreeze the whole block , so i am bit confused in that too , that how to approach it .\nAlthough , if i get something helpful from it , i will surely share it with the community and listen to there views too .",
    "1169810": "Thanks... The \"unfreeze the whole block\" actually does not refer to the BatchNormalitzion, but rather to EfficientNet as I understand it. If you unfreeze some layers within a block, the recommendation is to unfreeze also all remaining layers of that block (except the BatchNormalitzion).",
    "1169816": "yeah currently i am doing that only. \nWill see how does it works 😅😅"
  },
  "source": "meta"
}