{
  "id": 104012,
  "title": "Should you change your learning rate depending on your model's complexity, e.g. moving from resnet18 to resnet50?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/104012",
  "author_name": "",
  "post_date": "2019-08-13T15:45:31.672014700Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Again, this is my first competition so I'm trying to learn a lot about the theory at the same time. Suppose I found a good learning rate for resnet18, can I assume this value will be good for other models, say resnet50, densenet121, densenet201, etc - or this value is mostly model-dependent?</p>\n\n<p>Furthermore, will this value be equally good for 6-channels vs 3-channels inputs? We could also think about image size and probably a lot more factors.</p>\n\n<p>I guess a broader question would be, for any change you make to your model architecture, to what extent should you consider \"adapting\" the model architecture to this change?</p>\n\n<p>I know this is a hyperparameter we should thoroughly \"get right\", but where should the obsession stop?</p>",
  "messages": [
    {
      "id": "598459",
      "postDate": "08/13/2019 15:45:31",
      "content": "<p>Again, this is my first competition so I'm trying to learn a lot about the theory at the same time. Suppose I found a good learning rate for resnet18, can I assume this value will be good for other models, say resnet50, densenet121, densenet201, etc - or this value is mostly model-dependent?</p>\n\n<p>Furthermore, will this value be equally good for 6-channels vs 3-channels inputs? We could also think about image size and probably a lot more factors.</p>\n\n<p>I guess a broader question would be, for any change you make to your model architecture, to what extent should you consider \"adapting\" the model architecture to this change?</p>\n\n<p>I know this is a hyperparameter we should thoroughly \"get right\", but where should the obsession stop?</p>",
      "rawMarkdown": "Again, this is my first competition so I'm trying to learn a lot about the theory at the same time. Suppose I found a good learning rate for resnet18, can I assume this value will be good for other models, say resnet50, densenet121, densenet201, etc - or this value is mostly model-dependent?\n\nFurthermore, will this value be equally good for 6-channels vs 3-channels inputs? We could also think about image size and probably a lot more factors.\n\nI guess a broader question would be, for any change you make to your model architecture, to what extent should you consider \"adapting\" the model architecture to this change?\n\nI know this is a hyperparameter we should thoroughly \"get right\", but where should the obsession stop?",
      "votes": null
    },
    {
      "id": "598469",
      "postDate": "08/13/2019 15:56:13",
      "content": "<blockquote>\n  <p>or this value is mostly model-dependent?</p>\n</blockquote>\n\n<p>It is model-dependent</p>\n\n<p></p>\n\n<p>The gradients in value and number will differ from each other in resnet18 and resnet50, because the number of layers are different in architecture.</p>",
      "rawMarkdown": "&gt;or this value is mostly model-dependent?\n\nIt is model-dependent\n\n![](https://miro.medium.com/max/1400/0*00BrbBeDrFOjocpK.)\n\n\nThe gradients in value and number will differ from each other in resnet18 and resnet50, because the number of layers are different in architecture.",
      "votes": null
    },
    {
      "id": "598551",
      "postDate": "08/13/2019 17:49:36",
      "content": "<p>Thanks <a href=\"/cyberia\">@cyberia</a> this is helpful.  </p>\n\n<p>Moving beyond the original question, and assuming finding the best lr for a giving model relies in the realm of trial and error between a lr of 0.1 and 1e-6, I still wonder:</p>\n\n<p>Do you guys (experts) use a standardized trial-and-error procedure to find the best lr for a given model? This step is pretty resource and time consuming.</p>\n\n<p>I am aware of LR schedulers and adaptive LRs, but looking for a well formulated reasoning from more experienced people here.</p>",
      "rawMarkdown": "Thanks @cyberia this is helpful.  \n\nMoving beyond the original question, and assuming finding the best lr for a giving model relies in the realm of trial and error between a lr of 0.1 and 1e-6, I still wonder:\n\nDo you guys (experts) use a standardized trial-and-error procedure to find the best lr for a given model? This step is pretty resource and time consuming.\n\nI am aware of LR schedulers and adaptive LRs, but looking for a well formulated reasoning from more experienced people here.",
      "votes": null
    },
    {
      "id": "598573",
      "postDate": "08/13/2019 18:31:16",
      "content": "<p>i am also having a hard time tuning the hyperparameters. That kept me thinking whether would it be better to start with a larger model and smaller image sizes - then you do not need to tune the hyperparameters twice compared to using smaller models and having large imager sizes for experimenting. Definitely would love to hear other experts' input.</p>",
      "rawMarkdown": "i am also having a hard time tuning the hyperparameters. That kept me thinking whether would it be better to start with a larger model and smaller image sizes - then you do not need to tune the hyperparameters twice compared to using smaller models and having large imager sizes for experimenting. Definitely would love to hear other experts' input.",
      "votes": null
    },
    {
      "id": "598655",
      "postDate": "08/13/2019 20:58:37",
      "content": "<p>Learning rate finders are available in the <a href=\"https://docs.fast.ai/basic_train.html#lr_find\">fastai</a> library, and code for these are available for other libraries too, like <a href=\"https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/\">Keras</a>.</p>\n\n<p>I will say that I have found that for me, the learning rate finder most of the time leads me to choose the same learning rate regardless of model complexity, and seems to be more dependent on the dataset.</p>",
      "rawMarkdown": "Learning rate finders are available in the [fastai](https://docs.fast.ai/basic_train.html#lr_find) library, and code for these are available for other libraries too, like [Keras](https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/).\n\nI will say that I have found that for me, the learning rate finder most of the time leads me to choose the same learning rate regardless of model complexity, and seems to be more dependent on the dataset.",
      "votes": null
    },
    {
      "id": "598983",
      "postDate": "08/14/2019 10:22:13",
      "content": "<p>Have a look here <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/61476#359324\">https://www.kaggle.com/c/home-credit-default-risk/discussion/61476#359324</a> \nit is way better answer to your question than I can provide. Written by a Grand Master <a href=\"https://www.kaggle.com/tilii7\">https://www.kaggle.com/tilii7</a>.</p>\n\n<p>Unfortunately there is no silver bullet LR value for various architectures. I use hit and trial to get a feel of how it converges, ideally it should be done in a scientific manner without hit and trial as there is no room for luck in Science.</p>",
      "rawMarkdown": "Have a look here https://www.kaggle.com/c/home-credit-default-risk/discussion/61476#359324 \nit is way better answer to your question than I can provide. Written by a Grand Master https://www.kaggle.com/tilii7.\n\nUnfortunately there is no silver bullet LR value for various architectures. I use hit and trial to get a feel of how it converges, ideally it should be done in a scientific manner without hit and trial as there is no room for luck in Science.",
      "votes": null
    },
    {
      "id": "599197",
      "postDate": "08/14/2019 16:40:38",
      "content": "<p>Thanks a lot <a href=\"/cyberia\">@cyberia</a>, really helpful</p>",
      "rawMarkdown": "Thanks a lot @cyberia, really helpful",
      "votes": null
    },
    {
      "id": "601490",
      "postDate": "08/17/2019 17:09:45",
      "content": "<p>Thanks for this <a href=\"/tanlikesmath\">@tanlikesmath</a> , I personally used this for pytorch <a href=\"https://github.com/davidtvs/pytorch-lr-finder\">https://github.com/davidtvs/pytorch-lr-finder</a> and it was fairly helpful. Cheers</p>",
      "rawMarkdown": "Thanks for this @tanlikesmath , I personally used this for pytorch https://github.com/davidtvs/pytorch-lr-finder and it was fairly helpful. Cheers",
      "votes": null
    },
    {
      "id": "602250",
      "postDate": "08/18/2019 20:22:41",
      "content": "<p>FYI I use a cosine annealing scheduler from 3e-4 to 1e-7 and it seems to work well for many model architectures</p>",
      "rawMarkdown": "FYI I use a cosine annealing scheduler from 3e-4 to 1e-7 and it seems to work well for many model architectures",
      "votes": null
    },
    {
      "id": "602371",
      "postDate": "08/19/2019 02:25:31",
      "content": "<p>6 channels vs 3 channels are difference because channel is function which is extracting image features so more channel the better and high accuracy. PC version model channels has high accuracy than mobile version model because they has more channels than mobile model(e.g ssd-mobile).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F770123%2F3bb746bc8cbb8673175b3c9735acafd3%2F1_aq0q7gCvuNUqnMHh4cpnIw.png?generation=1566181480673540&amp;alt=media\" alt=\"\">\nrestNet-18 is the lowest of FLOPs.</p>",
      "rawMarkdown": "6 channels vs 3 channels are difference because channel is function which is extracting image features so more channel the better and high accuracy. PC version model channels has high accuracy than mobile version model because they has more channels than mobile model(e.g ssd-mobile).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F770123%2F3bb746bc8cbb8673175b3c9735acafd3%2F1_aq0q7gCvuNUqnMHh4cpnIw.png?generation=1566181480673540&amp;alt=media)\nrestNet-18 is the lowest of FLOPs.",
      "votes": null
    },
    {
      "id": "602437",
      "postDate": "08/19/2019 04:21:14",
      "content": "<p>Are you using PyTorch's <code>CosineAnnealingLR</code>?</p>",
      "rawMarkdown": "Are you using PyTorch's `CosineAnnealingLR`?",
      "votes": null
    },
    {
      "id": "602864",
      "postDate": "08/19/2019 15:27:51",
      "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> if you are talking about the one from pytorch-ignite, yes!</p>",
      "rawMarkdown": "lorenzofabbri92 if you are talking about the one from pytorch-ignite, yes!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 598469,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "08/13/2019 15:56:13",
      "content": "<blockquote>\n  <p>or this value is mostly model-dependent?</p>\n</blockquote>\n\n<p>It is model-dependent</p>\n\n<p></p>\n\n<p>The gradients in value and number will differ from each other in resnet18 and resnet50, because the number of layers are different in architecture.</p>",
      "votes": null,
      "replies": [
        {
          "id": 598551,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/13/2019 17:49:36",
          "content": "<p>Thanks <a href=\"/cyberia\">@cyberia</a> this is helpful.  </p>\n\n<p>Moving beyond the original question, and assuming finding the best lr for a giving model relies in the realm of trial and error between a lr of 0.1 and 1e-6, I still wonder:</p>\n\n<p>Do you guys (experts) use a standardized trial-and-error procedure to find the best lr for a given model? This step is pretty resource and time consuming.</p>\n\n<p>I am aware of LR schedulers and adaptive LRs, but looking for a well formulated reasoning from more experienced people here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 598983,
          "author_name": "cyberia",
          "author_url": "",
          "post_date": "08/14/2019 10:22:13",
          "content": "<p>Have a look here <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/61476#359324\">https://www.kaggle.com/c/home-credit-default-risk/discussion/61476#359324</a> \nit is way better answer to your question than I can provide. Written by a Grand Master <a href=\"https://www.kaggle.com/tilii7\">https://www.kaggle.com/tilii7</a>.</p>\n\n<p>Unfortunately there is no silver bullet LR value for various architectures. I use hit and trial to get a feel of how it converges, ideally it should be done in a scientific manner without hit and trial as there is no room for luck in Science.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599197,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/14/2019 16:40:38",
          "content": "<p>Thanks a lot <a href=\"/cyberia\">@cyberia</a>, really helpful</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 598573,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "08/13/2019 18:31:16",
      "content": "<p>i am also having a hard time tuning the hyperparameters. That kept me thinking whether would it be better to start with a larger model and smaller image sizes - then you do not need to tune the hyperparameters twice compared to using smaller models and having large imager sizes for experimenting. Definitely would love to hear other experts' input.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 598655,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "08/13/2019 20:58:37",
      "content": "<p>Learning rate finders are available in the <a href=\"https://docs.fast.ai/basic_train.html#lr_find\">fastai</a> library, and code for these are available for other libraries too, like <a href=\"https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/\">Keras</a>.</p>\n\n<p>I will say that I have found that for me, the learning rate finder most of the time leads me to choose the same learning rate regardless of model complexity, and seems to be more dependent on the dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 601490,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/17/2019 17:09:45",
          "content": "<p>Thanks for this <a href=\"/tanlikesmath\">@tanlikesmath</a> , I personally used this for pytorch <a href=\"https://github.com/davidtvs/pytorch-lr-finder\">https://github.com/davidtvs/pytorch-lr-finder</a> and it was fairly helpful. Cheers</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 602250,
      "author_name": "michelml",
      "author_url": "",
      "post_date": "08/18/2019 20:22:41",
      "content": "<p>FYI I use a cosine annealing scheduler from 3e-4 to 1e-7 and it seems to work well for many model architectures</p>",
      "votes": null,
      "replies": [
        {
          "id": 602437,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/19/2019 04:21:14",
          "content": "<p>Are you using PyTorch's <code>CosineAnnealingLR</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602864,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/19/2019 15:27:51",
          "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> if you are talking about the one from pytorch-ignite, yes!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 602371,
      "author_name": "elsa1717",
      "author_url": "",
      "post_date": "08/19/2019 02:25:31",
      "content": "<p>6 channels vs 3 channels are difference because channel is function which is extracting image features so more channel the better and high accuracy. PC version model channels has high accuracy than mobile version model because they has more channels than mobile model(e.g ssd-mobile).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F770123%2F3bb746bc8cbb8673175b3c9735acafd3%2F1_aq0q7gCvuNUqnMHh4cpnIw.png?generation=1566181480673540&amp;alt=media\" alt=\"\">\nrestNet-18 is the lowest of FLOPs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "598459": "Again, this is my first competition so I'm trying to learn a lot about the theory at the same time. Suppose I found a good learning rate for resnet18, can I assume this value will be good for other models, say resnet50, densenet121, densenet201, etc - or this value is mostly model-dependent?\n\nFurthermore, will this value be equally good for 6-channels vs 3-channels inputs? We could also think about image size and probably a lot more factors.\n\nI guess a broader question would be, for any change you make to your model architecture, to what extent should you consider \"adapting\" the model architecture to this change?\n\nI know this is a hyperparameter we should thoroughly \"get right\", but where should the obsession stop?",
    "598469": "&gt;or this value is mostly model-dependent?\n\nIt is model-dependent\n\n![](https://miro.medium.com/max/1400/0*00BrbBeDrFOjocpK.)\n\n\nThe gradients in value and number will differ from each other in resnet18 and resnet50, because the number of layers are different in architecture.",
    "598551": "Thanks @cyberia this is helpful.  \n\nMoving beyond the original question, and assuming finding the best lr for a giving model relies in the realm of trial and error between a lr of 0.1 and 1e-6, I still wonder:\n\nDo you guys (experts) use a standardized trial-and-error procedure to find the best lr for a given model? This step is pretty resource and time consuming.\n\nI am aware of LR schedulers and adaptive LRs, but looking for a well formulated reasoning from more experienced people here.",
    "598573": "i am also having a hard time tuning the hyperparameters. That kept me thinking whether would it be better to start with a larger model and smaller image sizes - then you do not need to tune the hyperparameters twice compared to using smaller models and having large imager sizes for experimenting. Definitely would love to hear other experts' input.",
    "598655": "Learning rate finders are available in the [fastai](https://docs.fast.ai/basic_train.html#lr_find) library, and code for these are available for other libraries too, like [Keras](https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/).\n\nI will say that I have found that for me, the learning rate finder most of the time leads me to choose the same learning rate regardless of model complexity, and seems to be more dependent on the dataset.",
    "598983": "Have a look here https://www.kaggle.com/c/home-credit-default-risk/discussion/61476#359324 \nit is way better answer to your question than I can provide. Written by a Grand Master https://www.kaggle.com/tilii7.\n\nUnfortunately there is no silver bullet LR value for various architectures. I use hit and trial to get a feel of how it converges, ideally it should be done in a scientific manner without hit and trial as there is no room for luck in Science.",
    "599197": "Thanks a lot @cyberia, really helpful",
    "601490": "Thanks for this @tanlikesmath , I personally used this for pytorch https://github.com/davidtvs/pytorch-lr-finder and it was fairly helpful. Cheers",
    "602250": "FYI I use a cosine annealing scheduler from 3e-4 to 1e-7 and it seems to work well for many model architectures",
    "602371": "6 channels vs 3 channels are difference because channel is function which is extracting image features so more channel the better and high accuracy. PC version model channels has high accuracy than mobile version model because they has more channels than mobile model(e.g ssd-mobile).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F770123%2F3bb746bc8cbb8673175b3c9735acafd3%2F1_aq0q7gCvuNUqnMHh4cpnIw.png?generation=1566181480673540&amp;alt=media)\nrestNet-18 is the lowest of FLOPs.",
    "602437": "Are you using PyTorch's `CosineAnnealingLR`?",
    "602864": "lorenzofabbri92 if you are talking about the one from pytorch-ignite, yes!"
  },
  "source": "meta"
}