{
  "id": 206183,
  "title": "Discriminative learning rates",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206183",
  "author_name": "",
  "post_date": "2020-12-23T14:44:26.975059100Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>For me it works fine to first fine-tune the final layers for 2 or 3 epochs (while keeping the rest of the model frozen) and then training the whole model for more epochs. Whether those initial epochs with the backbone frozen help seems to depend on how many layers you are putting on top (with more than 1 layer, it seems to help, while with a single layer it did not help). </p>\n<p>However, what does not seem to work (i.e. I've not found a way of doing this that does better in CV than the alternative) for me is discriminative learning rates - i.e. the above strategy, but where in the second stage the learning rate is lower for early layers in the network and higher learning in later layers. Are others seeing this, too?</p>\n<p>If this is not just me messing up somehow, is there some reason why we'd expect this? E.g. are these photos sufficiently far from ImageNet (possibly due to there being no relevant artificial objects, animals or people?) that refining even early layers a lot makes sense?</p>",
  "messages": [
    {
      "id": "1123853",
      "postDate": "12/23/2020 14:44:26",
      "content": "<p>For me it works fine to first fine-tune the final layers for 2 or 3 epochs (while keeping the rest of the model frozen) and then training the whole model for more epochs. Whether those initial epochs with the backbone frozen help seems to depend on how many layers you are putting on top (with more than 1 layer, it seems to help, while with a single layer it did not help). </p>\n<p>However, what does not seem to work (i.e. I've not found a way of doing this that does better in CV than the alternative) for me is discriminative learning rates - i.e. the above strategy, but where in the second stage the learning rate is lower for early layers in the network and higher learning in later layers. Are others seeing this, too?</p>\n<p>If this is not just me messing up somehow, is there some reason why we'd expect this? E.g. are these photos sufficiently far from ImageNet (possibly due to there being no relevant artificial objects, animals or people?) that refining even early layers a lot makes sense?</p>",
      "rawMarkdown": "For me it works fine to first fine-tune the final layers for 2 or 3 epochs (while keeping the rest of the model frozen) and then training the whole model for more epochs. Whether those initial epochs with the backbone frozen help seems to depend on how many layers you are putting on top (with more than 1 layer, it seems to help, while with a single layer it did not help). \n\nHowever, what does not seem to work (i.e. I've not found a way of doing this that does better in CV than the alternative) for me is discriminative learning rates - i.e. the above strategy, but where in the second stage the learning rate is lower for early layers in the network and higher learning in later layers. Are others seeing this, too?\n\nIf this is not just me messing up somehow, is there some reason why we'd expect this? E.g. are these photos sufficiently far from ImageNet (possibly due to there being no relevant artificial objects, animals or people?) that refining even early layers a lot makes sense?",
      "votes": null
    },
    {
      "id": "1124155",
      "postDate": "12/23/2020 17:48:53",
      "content": "<p>For me it did help a little when using differential learning rates however i dont freeze any weights. I use 3e-4for the base model and 2x that learning rate for the classifier (only one layer)</p>",
      "rawMarkdown": "For me it did help a little when using differential learning rates however i dont freeze any weights. I use 3e-4for the base model and 2x that learning rate for the classifier (only one layer)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1124155,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "12/23/2020 17:48:53",
      "content": "<p>For me it did help a little when using differential learning rates however i dont freeze any weights. I use 3e-4for the base model and 2x that learning rate for the classifier (only one layer)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1123853": "For me it works fine to first fine-tune the final layers for 2 or 3 epochs (while keeping the rest of the model frozen) and then training the whole model for more epochs. Whether those initial epochs with the backbone frozen help seems to depend on how many layers you are putting on top (with more than 1 layer, it seems to help, while with a single layer it did not help). \n\nHowever, what does not seem to work (i.e. I've not found a way of doing this that does better in CV than the alternative) for me is discriminative learning rates - i.e. the above strategy, but where in the second stage the learning rate is lower for early layers in the network and higher learning in later layers. Are others seeing this, too?\n\nIf this is not just me messing up somehow, is there some reason why we'd expect this? E.g. are these photos sufficiently far from ImageNet (possibly due to there being no relevant artificial objects, animals or people?) that refining even early layers a lot makes sense?",
    "1124155": "For me it did help a little when using differential learning rates however i dont freeze any weights. I use 3e-4for the base model and 2x that learning rate for the classifier (only one layer)"
  },
  "source": "meta"
}