{
  "id": 581275,
  "title": "Model is extremly sensitive to learning rate. Why is that the case?",
  "url": "/competitions/birdclef-2025/discussion/581275",
  "author_name": "Tim",
  "post_date": "2025-05-29T15:54:29.506000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>My models (I tried ViT and CNN like backbones) are extremely sensitive to the learning rate. That does not feel right. Currently I'm using AdamW with \\gamma=10^{-4}. However, when increasing it, the model is not learning at all. </p>\n<p>Has anyone else experienced this? What could cause such extreme learning rate sensitivity? Could it be something about the input scale, loss surface, optimizer setup, or architecture?</p>\n<p>Any pointers or debugging ideas are appreciated!</p>",
  "messages": [
    {
      "id": 3212305,
      "postDate": "2025-05-29T15:54:29.507Z",
      "content": "<p>My models (I tried ViT and CNN like backbones) are extremely sensitive to the learning rate. That does not feel right. Currently I'm using AdamW with \\gamma=10^{-4}. However, when increasing it, the model is not learning at all. </p>\n<p>Has anyone else experienced this? What could cause such extreme learning rate sensitivity? Could it be something about the input scale, loss surface, optimizer setup, or architecture?</p>\n<p>Any pointers or debugging ideas are appreciated!</p>",
      "rawMarkdown": "My models (I tried ViT and CNN like backbones) are extremely sensitive to the learning rate. That does not feel right. Currently I'm using AdamW with \\gamma=10^{-4}. However, when increasing it, the model is not learning at all. \n\nHas anyone else experienced this? What could cause such extreme learning rate sensitivity? Could it be something about the input scale, loss surface, optimizer setup, or architecture?\n\nAny pointers or debugging ideas are appreciated!",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3212305": "My models (I tried ViT and CNN like backbones) are extremely sensitive to the learning rate. That does not feel right. Currently I'm using AdamW with \\gamma=10^{-4}. However, when increasing it, the model is not learning at all. \n\nHas anyone else experienced this? What could cause such extreme learning rate sensitivity? Could it be something about the input scale, loss surface, optimizer setup, or architecture?\n\nAny pointers or debugging ideas are appreciated!"
  }
}