{
  "id": 169985,
  "title": "Differential Learning Rates",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169985",
  "author_name": "",
  "post_date": "2020-07-26T02:52:22.937363800Z",
  "votes": 11,
  "comment_count": 14,
  "views": 0,
  "content": "<p>There are probably not many hyperparameters more important than the learning rate in deep learning.  You should definitely spend time finding out what is the best learning rate for your model.  The difference between setting an ideal learning rate and an arbitrary one is huge.  </p>\n\n<p>With many Kaggle competitions, including this one, many are leveraging Transfer Learning.  A typical approach is to not only use an existing model, for example, one of the many used for ImageNet competitions but to also use its pre-trained weights.  A typical approach is to use most of the model but remove its head, which is the classifier, and replace it with one more suited for the task at hand.  So what you end up with, is a model that is very well trained, possibly over days or weeks on a very large dataset, and a novel classifier that starts off with completely random weights.  Intuition says that these two parts to the model should possibly be trained at different rates.  Enter Differential Learning Rates.</p>\n\n<p>You can train the main model at one learning rate, and the new classifier at another learning rate.  This is accomplished in PyTorch using optimizer parameter groups:</p>\n\n<p><code>\nbase_lr = .001\noptimizer = optimizer.AdamW([\n{ 'params': model.layer9.parameters(), 'lr': base_lr/3},\n{ 'params': model.classifier.parameters(), 'lr': base_lr /6},\n], lr=base_lr)\n</code></p>\n\n<p>So what this would do, is use <code>base_lr</code>/3 for <code>model.layer9</code> and <code>base_lr</code> /6 for <code>model.classifier</code> and then for everything else, it would use <code>base_lr</code>. </p>\n\n<p>This can likely produce better results than a \"One Size Fits All\" approach of trying to find one ideal learning rate for a pre-trained model along with new novel additions.  You still need to find the ideal learning rates, however, the point is you have two very different parts to the model and you may wish to use two differentiated rates.</p>",
  "messages": [
    {
      "id": "945613",
      "postDate": "07/26/2020 02:52:22",
      "content": "<p>There are probably not many hyperparameters more important than the learning rate in deep learning.  You should definitely spend time finding out what is the best learning rate for your model.  The difference between setting an ideal learning rate and an arbitrary one is huge.  </p>\n\n<p>With many Kaggle competitions, including this one, many are leveraging Transfer Learning.  A typical approach is to not only use an existing model, for example, one of the many used for ImageNet competitions but to also use its pre-trained weights.  A typical approach is to use most of the model but remove its head, which is the classifier, and replace it with one more suited for the task at hand.  So what you end up with, is a model that is very well trained, possibly over days or weeks on a very large dataset, and a novel classifier that starts off with completely random weights.  Intuition says that these two parts to the model should possibly be trained at different rates.  Enter Differential Learning Rates.</p>\n\n<p>You can train the main model at one learning rate, and the new classifier at another learning rate.  This is accomplished in PyTorch using optimizer parameter groups:</p>\n\n<p><code>\nbase_lr = .001\noptimizer = optimizer.AdamW([\n{ 'params': model.layer9.parameters(), 'lr': base_lr/3},\n{ 'params': model.classifier.parameters(), 'lr': base_lr /6},\n], lr=base_lr)\n</code></p>\n\n<p>So what this would do, is use <code>base_lr</code>/3 for <code>model.layer9</code> and <code>base_lr</code> /6 for <code>model.classifier</code> and then for everything else, it would use <code>base_lr</code>. </p>\n\n<p>This can likely produce better results than a \"One Size Fits All\" approach of trying to find one ideal learning rate for a pre-trained model along with new novel additions.  You still need to find the ideal learning rates, however, the point is you have two very different parts to the model and you may wish to use two differentiated rates.</p>",
      "rawMarkdown": "There are probably not many hyperparameters more important than the learning rate in deep learning.  You should definitely spend time finding out what is the best learning rate for your model.  The difference between setting an ideal learning rate and an arbitrary one is huge.  \n\nWith many Kaggle competitions, including this one, many are leveraging Transfer Learning.  A typical approach is to not only use an existing model, for example, one of the many used for ImageNet competitions but to also use its pre-trained weights.  A typical approach is to use most of the model but remove its head, which is the classifier, and replace it with one more suited for the task at hand.  So what you end up with, is a model that is very well trained, possibly over days or weeks on a very large dataset, and a novel classifier that starts off with completely random weights.  Intuition says that these two parts to the model should possibly be trained at different rates.  Enter Differential Learning Rates.\n\nYou can train the main model at one learning rate, and the new classifier at another learning rate.  This is accomplished in PyTorch using optimizer parameter groups:\n\n```\nbase_lr = .001\noptimizer = optimizer.AdamW([\n{ 'params': model.layer9.parameters(), 'lr': base_lr/3},\n{ 'params': model.classifier.parameters(), 'lr': base_lr /6},\n], lr=base_lr)\n```\n\nSo what this would do, is use `base_lr`/3 for `model.layer9` and `base_lr` /6 for `model.classifier` and then for everything else, it would use `base_lr`. \n\nThis can likely produce better results than a \"One Size Fits All\" approach of trying to find one ideal learning rate for a pre-trained model along with new novel additions.  You still need to find the ideal learning rates, however, the point is you have two very different parts to the model and you may wish to use two differentiated rates.",
      "votes": null
    },
    {
      "id": "945701",
      "postDate": "07/26/2020 04:57:50",
      "content": "<p>Thanks for the tip. In TensorFlow, you can accomplish the same thing using <code>tf.keras.backend.stop_gradient()</code>. For example if your model is</p>\n\n<pre><code>base = EfficientNetB0()\nx = base(x)\nx = head(x)\n</code></pre>\n\n<p>Then you can do the following</p>\n\n<pre><code>base = EfficientNetB0()\nx = base(x)\nx = 0.5*x + 0.5*tf.keras.backend.stop_gradient(x)\nx = head(x) \n</code></pre>\n\n<p>Then the head will train with twice the learning rate of the backbone.</p>",
      "rawMarkdown": "Thanks for the tip. In TensorFlow, you can accomplish the same thing using `tf.keras.backend.stop_gradient()`. For example if your model is\n\n    base = EfficientNetB0()\n    x = base(x)\n    x = head(x)\n    \nThen you can do the following\n\n    base = EfficientNetB0()\n    x = base(x)\n    x = 0.5*x + 0.5*tf.keras.backend.stop_gradient(x)\n    x = head(x) \n\nThen the head will train with twice the learning rate of the backbone.",
      "votes": null
    },
    {
      "id": "945770",
      "postDate": "07/26/2020 05:56:38",
      "content": "<p>Do you use this technique often?</p>",
      "rawMarkdown": "Do you use this technique often?",
      "votes": null
    },
    {
      "id": "945881",
      "postDate": "07/26/2020 07:39:54",
      "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> thanks for opening the discussion.</p>\n\n<p>Could you share some intuition about setting smaller learning rates to the last layers? or any reading?</p>\n\n<p>Intuitively I would do the opposite : \n- first layers are learning basic filters, that detect simple edges, textures and shape, they do not need to adapt much since they are quite universal and don't depend from previous layers.\n- last layers are trying to put everything together for your specific problem so they probably need to be more flexible than the previous layers hence a bigger learning rate\n- in finetuning one way to go is freeze all the layers but the last one. So the basic case is 0 learning rate for first layers and positive for the last. So why would you switch that approach to bigger learning rates at the beginning?</p>\n\n<p>Did you have good results with that approach? Thanks!</p>",
      "rawMarkdown": "brianfeeny thanks for opening the discussion.\n\nCould you share some intuition about setting smaller learning rates to the last layers? or any reading?\n\nIntuitively I would do the opposite : \n- first layers are learning basic filters, that detect simple edges, textures and shape, they do not need to adapt much since they are quite universal and don't depend from previous layers.\n- last layers are trying to put everything together for your specific problem so they probably need to be more flexible than the previous layers hence a bigger learning rate\n- in finetuning one way to go is freeze all the layers but the last one. So the basic case is 0 learning rate for first layers and positive for the last. So why would you switch that approach to bigger learning rates at the beginning?\n\nDid you have good results with that approach? Thanks!",
      "votes": null
    },
    {
      "id": "946421",
      "postDate": "07/26/2020 15:10:34",
      "content": "<p>If you read previous competition winning solutions, using differential learning rates has helped winners. However, I have never been successful with it. I have tried it in NLP and Image comps with transfer learning but it never increases my CV LB. Perhaps it doesn't help me because i always use simple light heads like</p>\n\n<pre><code> # SIMPLE HEAD\n x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = Dense(1,activation='sigmoid')(x)\n</code></pre>\n\n<p>And simple light heads don't require larger learning rate or more time to learn. Perhaps differential learning is more important if you have a medium head</p>\n\n<pre><code># MEDIUM HEAD\n x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = Dense(128)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(64)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(1,activation='sigmoid')(x)\n</code></pre>\n\n<p>Or fully connected heavy head</p>\n\n<pre><code># HEAVY HEAD (because no pooling)\n x = base(input)\n x = Flatten()(x)\n x = Dense(64)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(32)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(1,activation='sigmoid')(x)\n</code></pre>",
      "rawMarkdown": "If you read previous competition winning solutions, using differential learning rates has helped winners. However, I have never been successful with it. I have tried it in NLP and Image comps with transfer learning but it never increases my CV LB. Perhaps it doesn't help me because i always use simple light heads like\n\n     # SIMPLE HEAD\n     x = base(input)\n     x = GlobalAveragePooling2D()(x)\n     x = Dense(1,activation='sigmoid')(x)\n\nAnd simple light heads don't require larger learning rate or more time to learn. Perhaps differential learning is more important if you have a medium head\n\n    # MEDIUM HEAD\n     x = base(input)\n     x = GlobalAveragePooling2D()(x)\n     x = Dense(128)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(64)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(1,activation='sigmoid')(x)\n\nOr fully connected heavy head\n\n    # HEAVY HEAD (because no pooling)\n     x = base(input)\n     x = Flatten()(x)\n     x = Dense(64)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(32)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(1,activation='sigmoid')(x)",
      "votes": null
    },
    {
      "id": "946808",
      "postDate": "07/26/2020 21:22:52",
      "content": "<p>I didn't mean to imply that you would use smaller learning rates on the classifier, I was just showing an example.  What I have done is trained say Meta and CNN separately and found different ideal learning rates for each.  So that is one way I could use differential learning rates.  Also, yes I use a pre-trained model (Transfer Learning), and so I hope to start experimenting with Differential Learning rates myself.</p>",
      "rawMarkdown": "I didn't mean to imply that you would use smaller learning rates on the classifier, I was just showing an example.  What I have done is trained say Meta and CNN separately and found different ideal learning rates for each.  So that is one way I could use differential learning rates.  Also, yes I use a pre-trained model (Transfer Learning), and so I hope to start experimenting with Differential Learning rates myself.",
      "votes": null
    },
    {
      "id": "946821",
      "postDate": "07/26/2020 21:40:03",
      "content": "<p>Ok thank you, sorry I did not get it was just an example!</p>",
      "rawMarkdown": "Ok thank you, sorry I did not get it was just an example!",
      "votes": null
    },
    {
      "id": "946849",
      "postDate": "07/26/2020 22:31:28",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> thanks for sharing.  So I am trying to understand where you splice in your head.  Take for example EfficientNet. In PyTorch the <code>forward</code> function  looks like so:</p>\n\n<p>```\n        # Convolution layers\n        x = self.extract_features(inputs)</p>\n\n<pre><code>    # Pooling and final linear layer\n    x = self._avg_pooling(x)\n    x = x.flatten(start_dim=1)\n    x = self._dropout(x)\n    x = self._fc(x)\n</code></pre>\n\n<p>```</p>\n\n<p>Do you just replace <code>_fc</code> with your head, or do you remove the <code>pooling/flatten/dropout/_fc</code> and replace it with your head?  </p>\n\n<p>What I see a lot of people do is just swap out <code>_fc</code>........they have to, because the number of output features is different.  Typically with a dense head, and then take that down to 1 feature/binary.  But since you are changing the pooling, I figured you probably are splicing in before the pooling and replacing that with your head.</p>",
      "rawMarkdown": "cdeotte thanks for sharing.  So I am trying to understand where you splice in your head.  Take for example EfficientNet. In PyTorch the `forward` function  looks like so:\n\n```\n        # Convolution layers\n        x = self.extract_features(inputs)\n\n        # Pooling and final linear layer\n        x = self._avg_pooling(x)\n        x = x.flatten(start_dim=1)\n        x = self._dropout(x)\n        x = self._fc(x)\n```\n\nDo you just replace `_fc` with your head, or do you remove the `pooling/flatten/dropout/_fc` and replace it with your head?  \n\nWhat I see a lot of people do is just swap out `_fc`........they have to, because the number of output features is different.  Typically with a dense head, and then take that down to 1 feature/binary.  But since you are changing the pooling, I figured you probably are splicing in before the pooling and replacing that with your head.",
      "votes": null
    },
    {
      "id": "946941",
      "postDate": "07/27/2020 00:58:56",
      "content": "<p>I consider the head <code>pooling/flatten/dropout/_fc</code> because you can build a head without <code>pooling</code> if you wish to utilize the spatial information of the base CNN feature maps. For example AlexNet and VGGNet don't use <code>GlobalAveragePooling2D()</code> in their heads. (But most state of the art CNN do use <code>GlobalPooling</code> in their heads and using global pooling usually works better).</p>",
      "rawMarkdown": "I consider the head `pooling/flatten/dropout/_fc` because you can build a head without `pooling` if you wish to utilize the spatial information of the base CNN feature maps. For example AlexNet and VGGNet don't use `GlobalAveragePooling2D()` in their heads. (But most state of the art CNN do use `GlobalPooling` in their heads and using global pooling usually works better).",
      "votes": null
    },
    {
      "id": "947091",
      "postDate": "07/27/2020 04:29:37",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> don't you have to flatten though?  I realize you use TF (I use PyTorch), but basically, I am assuming you have come into your head, a tensor of size B, C, H, W.  I don't know much about GlobalAveragePooling because we don't have exactly that in PyTorch, but basically I assume its g going to take an average across your image dimensions and store those in your Channel dimension, so now you will end up with B, C, H, W.  H, W will both be equal to \"1\" at this point, since you took an average.  Now how do you go into a Dense layer with a 4D Tensor?  I would have thought you would \"flatten\" it and then go into Dense.</p>\n\n<p>So I would have thought your Simple Head looks like:</p>\n\n<p>```</p>\n\n<h1>SIMPLE HEAD</h1>\n\n<p>x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = flatten(x)\n x = Dense(1,activation='sigmoid')(x)\n```</p>",
      "rawMarkdown": "cdeotte don't you have to flatten though?  I realize you use TF (I use PyTorch), but basically, I am assuming you have come into your head, a tensor of size B, C, H, W.  I don't know much about GlobalAveragePooling because we don't have exactly that in PyTorch, but basically I assume its g going to take an average across your image dimensions and store those in your Channel dimension, so now you will end up with B, C, H, W.  H, W will both be equal to \"1\" at this point, since you took an average.  Now how do you go into a Dense layer with a 4D Tensor?  I would have thought you would \"flatten\" it and then go into Dense.\n\nSo I would have thought your Simple Head looks like:\n\n```\n# SIMPLE HEAD\n x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = flatten(x)\n x = Dense(1,activation='sigmoid')(x)\n```",
      "votes": null
    },
    {
      "id": "947146",
      "postDate": "07/27/2020 05:15:29",
      "content": "<p>In TensorFlow Keras, global average pooling returns <code>(batch_size, channels)</code>. Documentation <a href=\"https://keras.io/api/layers/pooling_layers/global_average_pooling2d/\">here</a>.</p>\n\n<p>But conceptually you are correct. Before global average pooling, we have <code>(batch_size, height, width, channels)</code> then global average pooling changes this to <code>(batch_size, 1, 1, channels)</code> and then TensorFlow Keras returns <code>(batch_size, channels)</code>, so it flattens for us.</p>",
      "rawMarkdown": "In TensorFlow Keras, global average pooling returns `(batch_size, channels)`. Documentation [here][1].\n\nBut conceptually you are correct. Before global average pooling, we have `(batch_size, height, width, channels)` then global average pooling changes this to `(batch_size, 1, 1, channels)` and then TensorFlow Keras returns `(batch_size, channels)`, so it flattens for us.\n\n[1]: https://keras.io/api/layers/pooling_layers/global_average_pooling2d/",
      "votes": null
    },
    {
      "id": "947165",
      "postDate": "07/27/2020 05:36:45",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> thanks that makes perfect sense</p>",
      "rawMarkdown": "cdeotte thanks that makes perfect sense",
      "votes": null
    },
    {
      "id": "948305",
      "postDate": "07/27/2020 19:39:11",
      "content": "<p>I am not quite sure how can you set learning rates in efficient net, there are many many layers in the efficientnet, do you mean you want to make different only last one plus additional last one? Or better - can you point to the full example?</p>",
      "rawMarkdown": "I am not quite sure how can you set learning rates in efficient net, there are many many layers in the efficientnet, do you mean you want to make different only last one plus additional last one? Or better - can you point to the full example?",
      "votes": null
    },
    {
      "id": "948358",
      "postDate": "07/27/2020 21:05:09",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> there is nothing special about EfficientNet.  You can set learning rates for any parameters you wish.  typically you would want to differentiate the rates as you get closer to the head.  Or you can simply just do a separate rate for the head.  EfficientNet has many layers yes, so you can treat them all mostly the same, but you may wish to change the rate for the final layers, or at least the classifier.  you simply reference the layer by name, as I do above.</p>",
      "rawMarkdown": "jacekpoplawski there is nothing special about EfficientNet.  You can set learning rates for any parameters you wish.  typically you would want to differentiate the rates as you get closer to the head.  Or you can simply just do a separate rate for the head.  EfficientNet has many layers yes, so you can treat them all mostly the same, but you may wish to change the rate for the final layers, or at least the classifier.  you simply reference the layer by name, as I do above.",
      "votes": null
    },
    {
      "id": "948364",
      "postDate": "07/27/2020 21:11:09",
      "content": "<p>Basically you said I can do it in many ways, what I mean is that you present example of way which is helpful, because I don't know what idea will work.</p>",
      "rawMarkdown": "Basically you said I can do it in many ways, what I mean is that you present example of way which is helpful, because I don't know what idea will work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 945701,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/26/2020 04:57:50",
      "content": "<p>Thanks for the tip. In TensorFlow, you can accomplish the same thing using <code>tf.keras.backend.stop_gradient()</code>. For example if your model is</p>\n\n<pre><code>base = EfficientNetB0()\nx = base(x)\nx = head(x)\n</code></pre>\n\n<p>Then you can do the following</p>\n\n<pre><code>base = EfficientNetB0()\nx = base(x)\nx = 0.5*x + 0.5*tf.keras.backend.stop_gradient(x)\nx = head(x) \n</code></pre>\n\n<p>Then the head will train with twice the learning rate of the backbone.</p>",
      "votes": null,
      "replies": [
        {
          "id": 945770,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/26/2020 05:56:38",
          "content": "<p>Do you use this technique often?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946421,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/26/2020 15:10:34",
          "content": "<p>If you read previous competition winning solutions, using differential learning rates has helped winners. However, I have never been successful with it. I have tried it in NLP and Image comps with transfer learning but it never increases my CV LB. Perhaps it doesn't help me because i always use simple light heads like</p>\n\n<pre><code> # SIMPLE HEAD\n x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = Dense(1,activation='sigmoid')(x)\n</code></pre>\n\n<p>And simple light heads don't require larger learning rate or more time to learn. Perhaps differential learning is more important if you have a medium head</p>\n\n<pre><code># MEDIUM HEAD\n x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = Dense(128)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(64)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(1,activation='sigmoid')(x)\n</code></pre>\n\n<p>Or fully connected heavy head</p>\n\n<pre><code># HEAVY HEAD (because no pooling)\n x = base(input)\n x = Flatten()(x)\n x = Dense(64)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(32)(x)\n x = BatchNormalization()(x)\n x = Activation('relu')(x)\n x = Dense(1,activation='sigmoid')(x)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946849,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/26/2020 22:31:28",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> thanks for sharing.  So I am trying to understand where you splice in your head.  Take for example EfficientNet. In PyTorch the <code>forward</code> function  looks like so:</p>\n\n<p>```\n        # Convolution layers\n        x = self.extract_features(inputs)</p>\n\n<pre><code>    # Pooling and final linear layer\n    x = self._avg_pooling(x)\n    x = x.flatten(start_dim=1)\n    x = self._dropout(x)\n    x = self._fc(x)\n</code></pre>\n\n<p>```</p>\n\n<p>Do you just replace <code>_fc</code> with your head, or do you remove the <code>pooling/flatten/dropout/_fc</code> and replace it with your head?  </p>\n\n<p>What I see a lot of people do is just swap out <code>_fc</code>........they have to, because the number of output features is different.  Typically with a dense head, and then take that down to 1 feature/binary.  But since you are changing the pooling, I figured you probably are splicing in before the pooling and replacing that with your head.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946941,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/27/2020 00:58:56",
          "content": "<p>I consider the head <code>pooling/flatten/dropout/_fc</code> because you can build a head without <code>pooling</code> if you wish to utilize the spatial information of the base CNN feature maps. For example AlexNet and VGGNet don't use <code>GlobalAveragePooling2D()</code> in their heads. (But most state of the art CNN do use <code>GlobalPooling</code> in their heads and using global pooling usually works better).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947091,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/27/2020 04:29:37",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> don't you have to flatten though?  I realize you use TF (I use PyTorch), but basically, I am assuming you have come into your head, a tensor of size B, C, H, W.  I don't know much about GlobalAveragePooling because we don't have exactly that in PyTorch, but basically I assume its g going to take an average across your image dimensions and store those in your Channel dimension, so now you will end up with B, C, H, W.  H, W will both be equal to \"1\" at this point, since you took an average.  Now how do you go into a Dense layer with a 4D Tensor?  I would have thought you would \"flatten\" it and then go into Dense.</p>\n\n<p>So I would have thought your Simple Head looks like:</p>\n\n<p>```</p>\n\n<h1>SIMPLE HEAD</h1>\n\n<p>x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = flatten(x)\n x = Dense(1,activation='sigmoid')(x)\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947146,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/27/2020 05:15:29",
          "content": "<p>In TensorFlow Keras, global average pooling returns <code>(batch_size, channels)</code>. Documentation <a href=\"https://keras.io/api/layers/pooling_layers/global_average_pooling2d/\">here</a>.</p>\n\n<p>But conceptually you are correct. Before global average pooling, we have <code>(batch_size, height, width, channels)</code> then global average pooling changes this to <code>(batch_size, 1, 1, channels)</code> and then TensorFlow Keras returns <code>(batch_size, channels)</code>, so it flattens for us.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947165,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/27/2020 05:36:45",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> thanks that makes perfect sense</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 945881,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "07/26/2020 07:39:54",
      "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> thanks for opening the discussion.</p>\n\n<p>Could you share some intuition about setting smaller learning rates to the last layers? or any reading?</p>\n\n<p>Intuitively I would do the opposite : \n- first layers are learning basic filters, that detect simple edges, textures and shape, they do not need to adapt much since they are quite universal and don't depend from previous layers.\n- last layers are trying to put everything together for your specific problem so they probably need to be more flexible than the previous layers hence a bigger learning rate\n- in finetuning one way to go is freeze all the layers but the last one. So the basic case is 0 learning rate for first layers and positive for the last. So why would you switch that approach to bigger learning rates at the beginning?</p>\n\n<p>Did you have good results with that approach? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 946808,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/26/2020 21:22:52",
          "content": "<p>I didn't mean to imply that you would use smaller learning rates on the classifier, I was just showing an example.  What I have done is trained say Meta and CNN separately and found different ideal learning rates for each.  So that is one way I could use differential learning rates.  Also, yes I use a pre-trained model (Transfer Learning), and so I hope to start experimenting with Differential Learning rates myself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946821,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/26/2020 21:40:03",
          "content": "<p>Ok thank you, sorry I did not get it was just an example!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 948305,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "07/27/2020 19:39:11",
      "content": "<p>I am not quite sure how can you set learning rates in efficient net, there are many many layers in the efficientnet, do you mean you want to make different only last one plus additional last one? Or better - can you point to the full example?</p>",
      "votes": null,
      "replies": [
        {
          "id": 948358,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/27/2020 21:05:09",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> there is nothing special about EfficientNet.  You can set learning rates for any parameters you wish.  typically you would want to differentiate the rates as you get closer to the head.  Or you can simply just do a separate rate for the head.  EfficientNet has many layers yes, so you can treat them all mostly the same, but you may wish to change the rate for the final layers, or at least the classifier.  you simply reference the layer by name, as I do above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948364,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/27/2020 21:11:09",
          "content": "<p>Basically you said I can do it in many ways, what I mean is that you present example of way which is helpful, because I don't know what idea will work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "945613": "There are probably not many hyperparameters more important than the learning rate in deep learning.  You should definitely spend time finding out what is the best learning rate for your model.  The difference between setting an ideal learning rate and an arbitrary one is huge.  \n\nWith many Kaggle competitions, including this one, many are leveraging Transfer Learning.  A typical approach is to not only use an existing model, for example, one of the many used for ImageNet competitions but to also use its pre-trained weights.  A typical approach is to use most of the model but remove its head, which is the classifier, and replace it with one more suited for the task at hand.  So what you end up with, is a model that is very well trained, possibly over days or weeks on a very large dataset, and a novel classifier that starts off with completely random weights.  Intuition says that these two parts to the model should possibly be trained at different rates.  Enter Differential Learning Rates.\n\nYou can train the main model at one learning rate, and the new classifier at another learning rate.  This is accomplished in PyTorch using optimizer parameter groups:\n\n```\nbase_lr = .001\noptimizer = optimizer.AdamW([\n{ 'params': model.layer9.parameters(), 'lr': base_lr/3},\n{ 'params': model.classifier.parameters(), 'lr': base_lr /6},\n], lr=base_lr)\n```\n\nSo what this would do, is use `base_lr`/3 for `model.layer9` and `base_lr` /6 for `model.classifier` and then for everything else, it would use `base_lr`. \n\nThis can likely produce better results than a \"One Size Fits All\" approach of trying to find one ideal learning rate for a pre-trained model along with new novel additions.  You still need to find the ideal learning rates, however, the point is you have two very different parts to the model and you may wish to use two differentiated rates.",
    "945701": "Thanks for the tip. In TensorFlow, you can accomplish the same thing using `tf.keras.backend.stop_gradient()`. For example if your model is\n\n    base = EfficientNetB0()\n    x = base(x)\n    x = head(x)\n    \nThen you can do the following\n\n    base = EfficientNetB0()\n    x = base(x)\n    x = 0.5*x + 0.5*tf.keras.backend.stop_gradient(x)\n    x = head(x) \n\nThen the head will train with twice the learning rate of the backbone.",
    "945770": "Do you use this technique often?",
    "945881": "brianfeeny thanks for opening the discussion.\n\nCould you share some intuition about setting smaller learning rates to the last layers? or any reading?\n\nIntuitively I would do the opposite : \n- first layers are learning basic filters, that detect simple edges, textures and shape, they do not need to adapt much since they are quite universal and don't depend from previous layers.\n- last layers are trying to put everything together for your specific problem so they probably need to be more flexible than the previous layers hence a bigger learning rate\n- in finetuning one way to go is freeze all the layers but the last one. So the basic case is 0 learning rate for first layers and positive for the last. So why would you switch that approach to bigger learning rates at the beginning?\n\nDid you have good results with that approach? Thanks!",
    "946421": "If you read previous competition winning solutions, using differential learning rates has helped winners. However, I have never been successful with it. I have tried it in NLP and Image comps with transfer learning but it never increases my CV LB. Perhaps it doesn't help me because i always use simple light heads like\n\n     # SIMPLE HEAD\n     x = base(input)\n     x = GlobalAveragePooling2D()(x)\n     x = Dense(1,activation='sigmoid')(x)\n\nAnd simple light heads don't require larger learning rate or more time to learn. Perhaps differential learning is more important if you have a medium head\n\n    # MEDIUM HEAD\n     x = base(input)\n     x = GlobalAveragePooling2D()(x)\n     x = Dense(128)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(64)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(1,activation='sigmoid')(x)\n\nOr fully connected heavy head\n\n    # HEAVY HEAD (because no pooling)\n     x = base(input)\n     x = Flatten()(x)\n     x = Dense(64)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(32)(x)\n     x = BatchNormalization()(x)\n     x = Activation('relu')(x)\n     x = Dense(1,activation='sigmoid')(x)",
    "946808": "I didn't mean to imply that you would use smaller learning rates on the classifier, I was just showing an example.  What I have done is trained say Meta and CNN separately and found different ideal learning rates for each.  So that is one way I could use differential learning rates.  Also, yes I use a pre-trained model (Transfer Learning), and so I hope to start experimenting with Differential Learning rates myself.",
    "946821": "Ok thank you, sorry I did not get it was just an example!",
    "946849": "cdeotte thanks for sharing.  So I am trying to understand where you splice in your head.  Take for example EfficientNet. In PyTorch the `forward` function  looks like so:\n\n```\n        # Convolution layers\n        x = self.extract_features(inputs)\n\n        # Pooling and final linear layer\n        x = self._avg_pooling(x)\n        x = x.flatten(start_dim=1)\n        x = self._dropout(x)\n        x = self._fc(x)\n```\n\nDo you just replace `_fc` with your head, or do you remove the `pooling/flatten/dropout/_fc` and replace it with your head?  \n\nWhat I see a lot of people do is just swap out `_fc`........they have to, because the number of output features is different.  Typically with a dense head, and then take that down to 1 feature/binary.  But since you are changing the pooling, I figured you probably are splicing in before the pooling and replacing that with your head.",
    "946941": "I consider the head `pooling/flatten/dropout/_fc` because you can build a head without `pooling` if you wish to utilize the spatial information of the base CNN feature maps. For example AlexNet and VGGNet don't use `GlobalAveragePooling2D()` in their heads. (But most state of the art CNN do use `GlobalPooling` in their heads and using global pooling usually works better).",
    "947091": "cdeotte don't you have to flatten though?  I realize you use TF (I use PyTorch), but basically, I am assuming you have come into your head, a tensor of size B, C, H, W.  I don't know much about GlobalAveragePooling because we don't have exactly that in PyTorch, but basically I assume its g going to take an average across your image dimensions and store those in your Channel dimension, so now you will end up with B, C, H, W.  H, W will both be equal to \"1\" at this point, since you took an average.  Now how do you go into a Dense layer with a 4D Tensor?  I would have thought you would \"flatten\" it and then go into Dense.\n\nSo I would have thought your Simple Head looks like:\n\n```\n# SIMPLE HEAD\n x = base(input)\n x = GlobalAveragePooling2D()(x)\n x = flatten(x)\n x = Dense(1,activation='sigmoid')(x)\n```",
    "947146": "In TensorFlow Keras, global average pooling returns `(batch_size, channels)`. Documentation [here][1].\n\nBut conceptually you are correct. Before global average pooling, we have `(batch_size, height, width, channels)` then global average pooling changes this to `(batch_size, 1, 1, channels)` and then TensorFlow Keras returns `(batch_size, channels)`, so it flattens for us.\n\n[1]: https://keras.io/api/layers/pooling_layers/global_average_pooling2d/",
    "947165": "cdeotte thanks that makes perfect sense",
    "948305": "I am not quite sure how can you set learning rates in efficient net, there are many many layers in the efficientnet, do you mean you want to make different only last one plus additional last one? Or better - can you point to the full example?",
    "948358": "jacekpoplawski there is nothing special about EfficientNet.  You can set learning rates for any parameters you wish.  typically you would want to differentiate the rates as you get closer to the head.  Or you can simply just do a separate rate for the head.  EfficientNet has many layers yes, so you can treat them all mostly the same, but you may wish to change the rate for the final layers, or at least the classifier.  you simply reference the layer by name, as I do above.",
    "948364": "Basically you said I can do it in many ways, what I mean is that you present example of way which is helpful, because I don't know what idea will work."
  },
  "source": "meta"
}