{
  "id": 169980,
  "title": "Tip for those using frozen models",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169980",
  "author_name": "",
  "post_date": "2020-07-26T01:14:55.297516400Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>For those that are freezing a pre-trained model and just training the head/classifier, you may wish to unfreeze the <code>BatchNorm</code> layers.  The <code>BatchNorm</code> layers are trained to the std and mean of the original data, ImageNet for example, and not your data, SIIM-ISIC Melanoma Classification.  The <code>BatchNorm</code> layers will try to <em>correct</em> your data to be more like ImageNet.  You can unfreeze <code>BatchNorm</code> layers (PyTorch) like so:</p>\n\n<p><code>\nfor name, param in model.named_parameters():\n    if(\"bn\" not in name):\n        param.requires_grad = False\n</code></p>\n\n<p>Because <code>BatchNorm</code> is scattered throughout the entire model this means that backpropagation is going to have to calculate gradients over basically the entire model.  It won't use them for layers you have frozen, but remember everything is chained so you have to calculate all layers between the ones you care about.</p>\n\n<p>My own experiments have shown that using a purely frozen ImageNet model is not the way to go for this competition, that you want to perhaps load pre-trained weights, but then train it on our new data.  But I realize some people have strategies that involve freezing only parts of the model.</p>",
  "messages": [
    {
      "id": "945573",
      "postDate": "07/26/2020 01:14:55",
      "content": "<p>For those that are freezing a pre-trained model and just training the head/classifier, you may wish to unfreeze the <code>BatchNorm</code> layers.  The <code>BatchNorm</code> layers are trained to the std and mean of the original data, ImageNet for example, and not your data, SIIM-ISIC Melanoma Classification.  The <code>BatchNorm</code> layers will try to <em>correct</em> your data to be more like ImageNet.  You can unfreeze <code>BatchNorm</code> layers (PyTorch) like so:</p>\n\n<p><code>\nfor name, param in model.named_parameters():\n    if(\"bn\" not in name):\n        param.requires_grad = False\n</code></p>\n\n<p>Because <code>BatchNorm</code> is scattered throughout the entire model this means that backpropagation is going to have to calculate gradients over basically the entire model.  It won't use them for layers you have frozen, but remember everything is chained so you have to calculate all layers between the ones you care about.</p>\n\n<p>My own experiments have shown that using a purely frozen ImageNet model is not the way to go for this competition, that you want to perhaps load pre-trained weights, but then train it on our new data.  But I realize some people have strategies that involve freezing only parts of the model.</p>",
      "rawMarkdown": "For those that are freezing a pre-trained model and just training the head/classifier, you may wish to unfreeze the `BatchNorm` layers.  The `BatchNorm` layers are trained to the std and mean of the original data, ImageNet for example, and not your data, SIIM-ISIC Melanoma Classification.  The `BatchNorm` layers will try to *correct* your data to be more like ImageNet.  You can unfreeze `BatchNorm` layers (PyTorch) like so:\n\n```\nfor name, param in model.named_parameters():\n    if(\"bn\" not in name):\n        param.requires_grad = False\n```\n\nBecause `BatchNorm` is scattered throughout the entire model this means that backpropagation is going to have to calculate gradients over basically the entire model.  It won't use them for layers you have frozen, but remember everything is chained so you have to calculate all layers between the ones you care about.\n\nMy own experiments have shown that using a purely frozen ImageNet model is not the way to go for this competition, that you want to perhaps load pre-trained weights, but then train it on our new data.  But I realize some people have strategies that involve freezing only parts of the model.",
      "votes": null
    },
    {
      "id": "950510",
      "postDate": "07/29/2020 12:46:55",
      "content": "<p>Imagenet data is completely different from the one to be used in this competition ,  so i guess if unfreeze our full model and train it for around 20 epochs , we can get much better results.\nAlthough , i haven't tried this experiment to just unfreeze the batch_norm layers. But will surely try it.\nThank You: 👍 </p>",
      "rawMarkdown": "Imagenet data is completely different from the one to be used in this competition ,  so i guess if unfreeze our full model and train it for around 20 epochs , we can get much better results.\nAlthough , i haven't tried this experiment to just unfreeze the batch_norm layers. But will surely try it.\nThank You: 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 950510,
      "author_name": "prashantarorat",
      "author_url": "",
      "post_date": "07/29/2020 12:46:55",
      "content": "<p>Imagenet data is completely different from the one to be used in this competition ,  so i guess if unfreeze our full model and train it for around 20 epochs , we can get much better results.\nAlthough , i haven't tried this experiment to just unfreeze the batch_norm layers. But will surely try it.\nThank You: 👍 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "945573": "For those that are freezing a pre-trained model and just training the head/classifier, you may wish to unfreeze the `BatchNorm` layers.  The `BatchNorm` layers are trained to the std and mean of the original data, ImageNet for example, and not your data, SIIM-ISIC Melanoma Classification.  The `BatchNorm` layers will try to *correct* your data to be more like ImageNet.  You can unfreeze `BatchNorm` layers (PyTorch) like so:\n\n```\nfor name, param in model.named_parameters():\n    if(\"bn\" not in name):\n        param.requires_grad = False\n```\n\nBecause `BatchNorm` is scattered throughout the entire model this means that backpropagation is going to have to calculate gradients over basically the entire model.  It won't use them for layers you have frozen, but remember everything is chained so you have to calculate all layers between the ones you care about.\n\nMy own experiments have shown that using a purely frozen ImageNet model is not the way to go for this competition, that you want to perhaps load pre-trained weights, but then train it on our new data.  But I realize some people have strategies that involve freezing only parts of the model.",
    "950510": "Imagenet data is completely different from the one to be used in this competition ,  so i guess if unfreeze our full model and train it for around 20 epochs , we can get much better results.\nAlthough , i haven't tried this experiment to just unfreeze the batch_norm layers. But will surely try it.\nThank You: 👍"
  },
  "source": "meta"
}