{
  "id": 172882,
  "title": "Tips for fine tuning EfficientNet",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172882",
  "author_name": "",
  "post_date": "2020-08-06T20:50:51.966062400Z",
  "votes": 36,
  "comment_count": 15,
  "views": 0,
  "content": "<h1>Tips for fine-tuning EfficientNet</h1>\n<p>On unfreezing layers:</p>\n<ul>\n<li>The <code>BathcNormalization</code> layers need to be kept frozen. If they are also turned to trainable, the first epoch after unfreezing will significantly reduce accuracy.</li>\n<li>In some cases, it may be beneficial to open up only a portion of layers instead of unfreezing all. This will make fine-tuning much faster when going to larger models like <code>B7</code>.</li>\n<li><strong>Each block needs to be all turned on or off</strong>. This is because the architecture includes a shortcut from the first layer to the last layer for each block. Not respecting blocks also significantly harms the final performance.</li>\n</ul>\n<p>Some other tips for utilizing EfficientNet:</p>\n<ul>\n<li>Larger variants of <code>EfficientNet</code> <strong>do not guarantee improved performance</strong>, especially for tasks <strong>with less data</strong> or <strong>fewer classes</strong>. In such a case, the larger variant of EfficientNet chosen, the harder it is to tune hyperparameters.</li>\n<li>EMA (<strong>Exponential Moving Average</strong>) is very helpful in training EfficientNet from scratch, but not so much for transfer learning.</li>\n<li>Do not use the <code>RMSprop</code> setup as in the original paper for transfer learning. The momentum and learning rate are too high for transfer learning. It will easily corrupt the pretrained weight and blow up the loss. A quick check is to see if loss (as <code>categorical cross-entropy</code>) is getting significantly larger than <code>log(NUM_CLASSES</code>) after the same epoch. If so, the initial <code>learning rate/momentum</code> is too high.</li>\n<li>Smaller batch size benefit validation accuracy, possibly due to effectively providing regularization.</li>\n<li>Using the latest EfficientNet weights.</li>\n</ul>\n<p><a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">Source</a>.</p>",
  "messages": [
    {
      "id": "960993",
      "postDate": "08/06/2020 20:50:51",
      "content": "<h1>Tips for fine-tuning EfficientNet</h1>\n<p>On unfreezing layers:</p>\n<ul>\n<li>The <code>BathcNormalization</code> layers need to be kept frozen. If they are also turned to trainable, the first epoch after unfreezing will significantly reduce accuracy.</li>\n<li>In some cases, it may be beneficial to open up only a portion of layers instead of unfreezing all. This will make fine-tuning much faster when going to larger models like <code>B7</code>.</li>\n<li><strong>Each block needs to be all turned on or off</strong>. This is because the architecture includes a shortcut from the first layer to the last layer for each block. Not respecting blocks also significantly harms the final performance.</li>\n</ul>\n<p>Some other tips for utilizing EfficientNet:</p>\n<ul>\n<li>Larger variants of <code>EfficientNet</code> <strong>do not guarantee improved performance</strong>, especially for tasks <strong>with less data</strong> or <strong>fewer classes</strong>. In such a case, the larger variant of EfficientNet chosen, the harder it is to tune hyperparameters.</li>\n<li>EMA (<strong>Exponential Moving Average</strong>) is very helpful in training EfficientNet from scratch, but not so much for transfer learning.</li>\n<li>Do not use the <code>RMSprop</code> setup as in the original paper for transfer learning. The momentum and learning rate are too high for transfer learning. It will easily corrupt the pretrained weight and blow up the loss. A quick check is to see if loss (as <code>categorical cross-entropy</code>) is getting significantly larger than <code>log(NUM_CLASSES</code>) after the same epoch. If so, the initial <code>learning rate/momentum</code> is too high.</li>\n<li>Smaller batch size benefit validation accuracy, possibly due to effectively providing regularization.</li>\n<li>Using the latest EfficientNet weights.</li>\n</ul>\n<p><a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">Source</a>.</p>",
      "rawMarkdown": "# Tips for fine-tuning EfficientNet\nOn unfreezing layers:\n\n- The `BathcNormalization` layers need to be kept frozen. If they are also turned to trainable, the first epoch after unfreezing will significantly reduce accuracy.\n- In some cases, it may be beneficial to open up only a portion of layers instead of unfreezing all. This will make fine-tuning much faster when going to larger models like `B7`.\n- **Each block needs to be all turned on or off**. This is because the architecture includes a shortcut from the first layer to the last layer for each block. Not respecting blocks also significantly harms the final performance.\n\nSome other tips for utilizing EfficientNet:\n- Larger variants of `EfficientNet` **do not guarantee improved performance**, especially for tasks **with less data** or **fewer classes**. In such a case, the larger variant of EfficientNet chosen, the harder it is to tune hyperparameters.\n- EMA (**Exponential Moving Average**) is very helpful in training EfficientNet from scratch, but not so much for transfer learning.\n- Do not use the `RMSprop` setup as in the original paper for transfer learning. The momentum and learning rate are too high for transfer learning. It will easily corrupt the pretrained weight and blow up the loss. A quick check is to see if loss (as `categorical cross-entropy`) is getting significantly larger than `log(NUM_CLASSES`) after the same epoch. If so, the initial `learning rate/momentum` is too high.\n- Smaller batch size benefit validation accuracy, possibly due to effectively providing regularization.\n- Using the latest EfficientNet weights.\n\n[Source](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/).",
      "votes": null
    },
    {
      "id": "961106",
      "postDate": "08/06/2020 23:12:39",
      "content": "<p>Why would you keep BatchNorm frozen? BatchNorm contains values that are trained to approximate the mean and standard deviation of the dataset the model was originally trained on...ImageNet, not the dataset you are trying to fine-tune on.  BatchNorm will try to correct your input of melanoma data as it tries to differ from ImageNet data.  I would think that if you were going to unfreeze anything it would be the BatchNorm layers.  </p>",
      "rawMarkdown": "Why would you keep BatchNorm frozen? BatchNorm contains values that are trained to approximate the mean and standard deviation of the dataset the model was originally trained on...ImageNet, not the dataset you are trying to fine-tune on.  BatchNorm will try to correct your input of melanoma data as it tries to differ from ImageNet data.  I would think that if you were going to unfreeze anything it would be the BatchNorm layers.",
      "votes": null
    },
    {
      "id": "961155",
      "postDate": "08/07/2020 00:38:22",
      "content": "<p>I think it's a common practice to freeze the <code>bn layer</code> while fine-tuning. When we unfreeze a model that contains <code>BatchNormalization</code> layers in order to do fine-tuning, we should keep the <code>BatchNormalization</code> layers in inference mode by passing <code>training=False</code> when calling the base model. Otherwise, the updates applied to the non-trainable weights will suddenly destroy what the model has learned.</p>\n<p>IMO, the <code>bn_layer</code> should be unfrozen if I can use a larger <code>batch size</code>, otherwise, it might not be stable properly.</p>",
      "rawMarkdown": "I think it's a common practice to freeze the `bn layer` while fine-tuning. When we unfreeze a model that contains `BatchNormalization` layers in order to do fine-tuning, we should keep the `BatchNormalization` layers in inference mode by passing `training=False` when calling the base model. Otherwise, the updates applied to the non-trainable weights will suddenly destroy what the model has learned.\n\nIMO, the `bn_layer` should be unfrozen if I can use a larger `batch size`, otherwise, it might not be stable properly.",
      "votes": null
    },
    {
      "id": "961544",
      "postDate": "08/07/2020 09:20:07",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a>, can't understand why freezing batch norm layer would create more stability or better results.<br>\nNot saying that this is not true but it seems very counter-intuitive. Would you have a link to some content talking about this?</p>",
      "rawMarkdown": "I agree with @brianfeeny, can't understand why freezing batch norm layer would create more stability or better results.\nNot saying that this is not true but it seems very counter-intuitive. Would you have a link to some content talking about this?",
      "votes": null
    },
    {
      "id": "961623",
      "postDate": "08/07/2020 10:54:01",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>  hmm, I see. I've already given the source link at the very end of this post, please follow the link. </p>\n<p>However, IMO, it actually depends on the custom data set. If it's way more <strong>different</strong> than the original ImageNet data set and we need to <strong>fine-tune</strong> the pre-trained weights on it and if we have to set <strong>lower batch size</strong> then, in this case, I think it's a bit risky to set <code>bn</code> layer <strong>unfreeze</strong>. But if we can accumulate a larger <code>batch size</code> (<code>Gradient Accumulation</code>), we can then set the <code>bn</code> layer unfreeze. Also, with lower <code>batch size</code>, we can also replace <code>BatchNormalization</code> with <code>Group</code> or <code>Instance</code> normalization (if we want).</p>\n<p>Also, anyone can check directly from <a href=\"https://keras.io/guides/transfer_learning/\" target=\"_blank\">here</a>, written by François Chollet about fine-tuning - section:<strong>Fine-tuning</strong>, title <code>Important notes about BatchNormalization layer</code>. </p>",
      "rawMarkdown": "optimo  hmm, I see. I've already given the source link at the very end of this post, please follow the link. \n\nHowever, IMO, it actually depends on the custom data set. If it's way more **different** than the original ImageNet data set and we need to **fine-tune** the pre-trained weights on it and if we have to set **lower batch size** then, in this case, I think it's a bit risky to set `bn` layer **unfreeze**. But if we can accumulate a larger `batch size` (`Gradient Accumulation`), we can then set the `bn` layer unfreeze. Also, with lower `batch size`, we can also replace `BatchNormalization` with `Group` or `Instance` normalization (if we want).\n\nAlso, anyone can check directly from [here](https://keras.io/guides/transfer_learning/), written by François Chollet about fine-tuning - section:**Fine-tuning**, title `Important notes about BatchNormalization layer`.",
      "votes": null
    },
    {
      "id": "961642",
      "postDate": "08/07/2020 11:07:35",
      "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> rather than downvoting, simple giving constructive feedback gives more sense, I believe. I can easily ignore that but It seems, it's happening <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#961105\" target=\"_blank\">in other places</a>. 😅</p>",
      "rawMarkdown": "brianfeeny rather than downvoting, simple giving constructive feedback gives more sense, I believe. I can easily ignore that but It seems, it's happening [in other places](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#961105). 😅",
      "votes": null
    },
    {
      "id": "961665",
      "postDate": "08/07/2020 11:23:34",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I don't downvote.  If somehow something said I downvoted something then it was by accident.  Edit: So I just read up on downvoting and if I had downvoted it would show in blue (to me), but your downvoted comments are in grey.  </p>",
      "rawMarkdown": "ipythonx I don't downvote.  If somehow something said I downvoted something then it was by accident.  Edit: So I just read up on downvoting and if I had downvoted it would show in blue (to me), but your downvoted comments are in grey.",
      "votes": null
    },
    {
      "id": "961932",
      "postDate": "08/07/2020 16:19:39",
      "content": "<p>I found adding dense head improved the model performance. Adding three Fully Connected Layers of sizes 512, 256 and 128 helped a lot. \nRectified Adam optimizer seems to work slightly better than adam. \nCyclical Learning rates and Focal loss also work better. </p>",
      "rawMarkdown": "I found adding dense head improved the model performance. Adding three Fully Connected Layers of sizes 512, 256 and 128 helped a lot. \nRectified Adam optimizer seems to work slightly better than adam. \nCyclical Learning rates and Focal loss also work better.",
      "votes": null
    },
    {
      "id": "962057",
      "postDate": "08/07/2020 19:00:00",
      "content": "<p><a href=\"/ajaykumar7778\">@ajaykumar7778</a> So you are saying that you found doing 512, 256, 128 FC layers and then finally 1 output, worked better than just a dense head?  This was for just the image data? with or without dropout and batch normalization?  My own experiments showed better scores without batch norm or dropout.</p>",
      "rawMarkdown": "ajaykumar7778 So you are saying that you found doing 512, 256, 128 FC layers and then finally 1 output, worked better than just a dense head?  This was for just the image data? with or without dropout and batch normalization?  My own experiments showed better scores without batch norm or dropout.",
      "votes": null
    },
    {
      "id": "962119",
      "postDate": "08/07/2020 20:31:47",
      "content": "<p>Interesting. For me adding dense + bn + relu + dropout blocks at the head decreased performance (by a small amount) compared to just a single linear classifier. Maybe I can try 2 layer Dense + Relu/Mish instead.</p>",
      "rawMarkdown": "Interesting. For me adding dense + bn + relu + dropout blocks at the head decreased performance (by a small amount) compared to just a single linear classifier. Maybe I can try 2 layer Dense + Relu/Mish instead.",
      "votes": null
    },
    {
      "id": "962154",
      "postDate": "08/07/2020 21:16:40",
      "content": "<p>Interesting, I tried adding dropout layer but haven't tried dense layer yet…</p>",
      "rawMarkdown": "Interesting, I tried adding dropout layer but haven't tried dense layer yet...",
      "votes": null
    },
    {
      "id": "962316",
      "postDate": "08/08/2020 03:31:40",
      "content": "<p>Yes. Adding 3 FC layers instead of a single dense head worked better. This is only for image data. I did not use BatchNorm, but used a dropout of 0.3 for each of the FC layers.</p>",
      "rawMarkdown": "Yes. Adding 3 FC layers instead of a single dense head worked better. This is only for image data. I did not use BatchNorm, but used a dropout of 0.3 for each of the FC layers.",
      "votes": null
    },
    {
      "id": "962504",
      "postDate": "08/08/2020 07:15:59",
      "content": "<p><a href=\"/ajaykumar7778\">@ajaykumar7778</a> thanks for sharing. do your experiments show that BN is no needed in your FC layers  </p>\n\n<p><a href=\"/arroqc\">@arroqc</a>  i'm using similar head - 2-3 blocks of dense + prelu + bn + dropout. the improvement is tiny. i wonder where's the best place for BN? after dense or after activation layer?</p>",
      "rawMarkdown": "ajaykumar7778 thanks for sharing. do your experiments show that BN is no needed in your FC layers  \n\n@arroqc  i'm using similar head - 2-3 blocks of dense + prelu + bn + dropout. the improvement is tiny. i wonder where's the best place for BN? after dense or after activation layer?",
      "votes": null
    },
    {
      "id": "963004",
      "postDate": "08/08/2020 15:35:17",
      "content": "<p>It's an open question. If you look for it on the web you'll find a lot of different answer. I think it has become common practice to do Dense + BN + activation + dropout. But like many things in machine learning you can just try it :)</p>",
      "rawMarkdown": "It's an open question. If you look for it on the web you'll find a lot of different answer. I think it has become common practice to do Dense + BN + activation + dropout. But like many things in machine learning you can just try it :)",
      "votes": null
    },
    {
      "id": "963207",
      "postDate": "08/08/2020 18:50:23",
      "content": "<p><a href=\"/arroqc\">@arroqc</a> indeed, there are different views from big figures in the DL field. so one more thing to try.   </p>",
      "rawMarkdown": "arroqc indeed, there are different views from big figures in the DL field. so one more thing to try.",
      "votes": null
    },
    {
      "id": "1159798",
      "postDate": "01/19/2021 13:27:26",
      "content": "<p>Do you know a Kaggle notebook applying fine-tuning on EfficientNet using PyTorch? I would be very interested to see a working example.</p>",
      "rawMarkdown": "Do you know a Kaggle notebook applying fine-tuning on EfficientNet using PyTorch? I would be very interested to see a working example.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961623,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "08/07/2020 10:54:01",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>  hmm, I see. I've already given the source link at the very end of this post, please follow the link. </p>\n<p>However, IMO, it actually depends on the custom data set. If it's way more <strong>different</strong> than the original ImageNet data set and we need to <strong>fine-tune</strong> the pre-trained weights on it and if we have to set <strong>lower batch size</strong> then, in this case, I think it's a bit risky to set <code>bn</code> layer <strong>unfreeze</strong>. But if we can accumulate a larger <code>batch size</code> (<code>Gradient Accumulation</code>), we can then set the <code>bn</code> layer unfreeze. Also, with lower <code>batch size</code>, we can also replace <code>BatchNormalization</code> with <code>Group</code> or <code>Instance</code> normalization (if we want).</p>\n<p>Also, anyone can check directly from <a href=\"https://keras.io/guides/transfer_learning/\" target=\"_blank\">here</a>, written by François Chollet about fine-tuning - section:<strong>Fine-tuning</strong>, title <code>Important notes about BatchNormalization layer</code>. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1159798,
      "author_name": "pheaboo",
      "author_url": "",
      "post_date": "01/19/2021 13:27:26",
      "content": "<p>Do you know a Kaggle notebook applying fine-tuning on EfficientNet using PyTorch? I would be very interested to see a working example.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 961106,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "08/06/2020 23:12:39",
      "content": "<p>Why would you keep BatchNorm frozen? BatchNorm contains values that are trained to approximate the mean and standard deviation of the dataset the model was originally trained on...ImageNet, not the dataset you are trying to fine-tune on.  BatchNorm will try to correct your input of melanoma data as it tries to differ from ImageNet data.  I would think that if you were going to unfreeze anything it would be the BatchNorm layers.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 961155,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "08/07/2020 00:38:22",
          "content": "<p>I think it's a common practice to freeze the <code>bn layer</code> while fine-tuning. When we unfreeze a model that contains <code>BatchNormalization</code> layers in order to do fine-tuning, we should keep the <code>BatchNormalization</code> layers in inference mode by passing <code>training=False</code> when calling the base model. Otherwise, the updates applied to the non-trainable weights will suddenly destroy what the model has learned.</p>\n<p>IMO, the <code>bn_layer</code> should be unfrozen if I can use a larger <code>batch size</code>, otherwise, it might not be stable properly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961544,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/07/2020 09:20:07",
          "content": "<p>I agree with <a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a>, can't understand why freezing batch norm layer would create more stability or better results.<br>\nNot saying that this is not true but it seems very counter-intuitive. Would you have a link to some content talking about this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961642,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "08/07/2020 11:07:35",
          "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> rather than downvoting, simple giving constructive feedback gives more sense, I believe. I can easily ignore that but It seems, it's happening <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#961105\" target=\"_blank\">in other places</a>. 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961665,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/07/2020 11:23:34",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I don't downvote.  If somehow something said I downvoted something then it was by accident.  Edit: So I just read up on downvoting and if I had downvoted it would show in blue (to me), but your downvoted comments are in grey.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 961932,
      "author_name": "ajaykumar7778",
      "author_url": "",
      "post_date": "08/07/2020 16:19:39",
      "content": "<p>I found adding dense head improved the model performance. Adding three Fully Connected Layers of sizes 512, 256 and 128 helped a lot. \nRectified Adam optimizer seems to work slightly better than adam. \nCyclical Learning rates and Focal loss also work better. </p>",
      "votes": null,
      "replies": [
        {
          "id": 962057,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/07/2020 19:00:00",
          "content": "<p><a href=\"/ajaykumar7778\">@ajaykumar7778</a> So you are saying that you found doing 512, 256, 128 FC layers and then finally 1 output, worked better than just a dense head?  This was for just the image data? with or without dropout and batch normalization?  My own experiments showed better scores without batch norm or dropout.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962119,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "08/07/2020 20:31:47",
          "content": "<p>Interesting. For me adding dense + bn + relu + dropout blocks at the head decreased performance (by a small amount) compared to just a single linear classifier. Maybe I can try 2 layer Dense + Relu/Mish instead.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962154,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "08/07/2020 21:16:40",
          "content": "<p>Interesting, I tried adding dropout layer but haven't tried dense layer yet…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962316,
          "author_name": "ajaykumar7778",
          "author_url": "",
          "post_date": "08/08/2020 03:31:40",
          "content": "<p>Yes. Adding 3 FC layers instead of a single dense head worked better. This is only for image data. I did not use BatchNorm, but used a dropout of 0.3 for each of the FC layers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962504,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "08/08/2020 07:15:59",
          "content": "<p><a href=\"/ajaykumar7778\">@ajaykumar7778</a> thanks for sharing. do your experiments show that BN is no needed in your FC layers  </p>\n\n<p><a href=\"/arroqc\">@arroqc</a>  i'm using similar head - 2-3 blocks of dense + prelu + bn + dropout. the improvement is tiny. i wonder where's the best place for BN? after dense or after activation layer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963004,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "08/08/2020 15:35:17",
          "content": "<p>It's an open question. If you look for it on the web you'll find a lot of different answer. I think it has become common practice to do Dense + BN + activation + dropout. But like many things in machine learning you can just try it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963207,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "08/08/2020 18:50:23",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> indeed, there are different views from big figures in the DL field. so one more thing to try.   </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "960993": "# Tips for fine-tuning EfficientNet\nOn unfreezing layers:\n\n- The `BathcNormalization` layers need to be kept frozen. If they are also turned to trainable, the first epoch after unfreezing will significantly reduce accuracy.\n- In some cases, it may be beneficial to open up only a portion of layers instead of unfreezing all. This will make fine-tuning much faster when going to larger models like `B7`.\n- **Each block needs to be all turned on or off**. This is because the architecture includes a shortcut from the first layer to the last layer for each block. Not respecting blocks also significantly harms the final performance.\n\nSome other tips for utilizing EfficientNet:\n- Larger variants of `EfficientNet` **do not guarantee improved performance**, especially for tasks **with less data** or **fewer classes**. In such a case, the larger variant of EfficientNet chosen, the harder it is to tune hyperparameters.\n- EMA (**Exponential Moving Average**) is very helpful in training EfficientNet from scratch, but not so much for transfer learning.\n- Do not use the `RMSprop` setup as in the original paper for transfer learning. The momentum and learning rate are too high for transfer learning. It will easily corrupt the pretrained weight and blow up the loss. A quick check is to see if loss (as `categorical cross-entropy`) is getting significantly larger than `log(NUM_CLASSES`) after the same epoch. If so, the initial `learning rate/momentum` is too high.\n- Smaller batch size benefit validation accuracy, possibly due to effectively providing regularization.\n- Using the latest EfficientNet weights.\n\n[Source](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/).",
    "961106": "Why would you keep BatchNorm frozen? BatchNorm contains values that are trained to approximate the mean and standard deviation of the dataset the model was originally trained on...ImageNet, not the dataset you are trying to fine-tune on.  BatchNorm will try to correct your input of melanoma data as it tries to differ from ImageNet data.  I would think that if you were going to unfreeze anything it would be the BatchNorm layers.",
    "961155": "I think it's a common practice to freeze the `bn layer` while fine-tuning. When we unfreeze a model that contains `BatchNormalization` layers in order to do fine-tuning, we should keep the `BatchNormalization` layers in inference mode by passing `training=False` when calling the base model. Otherwise, the updates applied to the non-trainable weights will suddenly destroy what the model has learned.\n\nIMO, the `bn_layer` should be unfrozen if I can use a larger `batch size`, otherwise, it might not be stable properly.",
    "961544": "I agree with @brianfeeny, can't understand why freezing batch norm layer would create more stability or better results.\nNot saying that this is not true but it seems very counter-intuitive. Would you have a link to some content talking about this?",
    "961623": "optimo  hmm, I see. I've already given the source link at the very end of this post, please follow the link. \n\nHowever, IMO, it actually depends on the custom data set. If it's way more **different** than the original ImageNet data set and we need to **fine-tune** the pre-trained weights on it and if we have to set **lower batch size** then, in this case, I think it's a bit risky to set `bn` layer **unfreeze**. But if we can accumulate a larger `batch size` (`Gradient Accumulation`), we can then set the `bn` layer unfreeze. Also, with lower `batch size`, we can also replace `BatchNormalization` with `Group` or `Instance` normalization (if we want).\n\nAlso, anyone can check directly from [here](https://keras.io/guides/transfer_learning/), written by François Chollet about fine-tuning - section:**Fine-tuning**, title `Important notes about BatchNormalization layer`.",
    "961642": "brianfeeny rather than downvoting, simple giving constructive feedback gives more sense, I believe. I can easily ignore that but It seems, it's happening [in other places](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#961105). 😅",
    "961665": "ipythonx I don't downvote.  If somehow something said I downvoted something then it was by accident.  Edit: So I just read up on downvoting and if I had downvoted it would show in blue (to me), but your downvoted comments are in grey.",
    "961932": "I found adding dense head improved the model performance. Adding three Fully Connected Layers of sizes 512, 256 and 128 helped a lot. \nRectified Adam optimizer seems to work slightly better than adam. \nCyclical Learning rates and Focal loss also work better.",
    "962057": "ajaykumar7778 So you are saying that you found doing 512, 256, 128 FC layers and then finally 1 output, worked better than just a dense head?  This was for just the image data? with or without dropout and batch normalization?  My own experiments showed better scores without batch norm or dropout.",
    "962119": "Interesting. For me adding dense + bn + relu + dropout blocks at the head decreased performance (by a small amount) compared to just a single linear classifier. Maybe I can try 2 layer Dense + Relu/Mish instead.",
    "962154": "Interesting, I tried adding dropout layer but haven't tried dense layer yet...",
    "962316": "Yes. Adding 3 FC layers instead of a single dense head worked better. This is only for image data. I did not use BatchNorm, but used a dropout of 0.3 for each of the FC layers.",
    "962504": "ajaykumar7778 thanks for sharing. do your experiments show that BN is no needed in your FC layers  \n\n@arroqc  i'm using similar head - 2-3 blocks of dense + prelu + bn + dropout. the improvement is tiny. i wonder where's the best place for BN? after dense or after activation layer?",
    "963004": "It's an open question. If you look for it on the web you'll find a lot of different answer. I think it has become common practice to do Dense + BN + activation + dropout. But like many things in machine learning you can just try it :)",
    "963207": "arroqc indeed, there are different views from big figures in the DL field. so one more thing to try.",
    "1159798": "Do you know a Kaggle notebook applying fine-tuning on EfficientNet using PyTorch? I would be very interested to see a working example."
  },
  "source": "meta"
}