{
  "id": 327147,
  "title": "Question about LR scheduler",
  "url": "/competitions/birdclef-2022/discussion/327147",
  "author_name": "",
  "post_date": "2022-05-25T20:27:21.996363900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In this comp we dealt with problems in making the models learn continuously. Every model we trained stopped learning (reducing loss) after epoch ~16 and a few of them resumed learning only in epochs 23 to 27. This results in a lot of epochs and time spent with no improvements.</p>\n<p>Is this a known behaviour of CosineAnnealingLR? Is there some way to avoid that? I don't remember this pattern in other image classification models (with different datasets) I trained with this scheduler, but I may just have forgotten it occured. Could someone shed a light on this?</p>\n<p><a href=\"https://postimg.cc/0McHgB5g\" target=\"_blank\"><img src=\"https://i.postimg.cc/7hkv2FpL/plot.png\" alt=\"plot.png\"></a></p>",
  "messages": [
    {
      "id": "1801523",
      "postDate": "05/25/2022 20:27:21",
      "content": "<p>In this comp we dealt with problems in making the models learn continuously. Every model we trained stopped learning (reducing loss) after epoch ~16 and a few of them resumed learning only in epochs 23 to 27. This results in a lot of epochs and time spent with no improvements.</p>\n<p>Is this a known behaviour of CosineAnnealingLR? Is there some way to avoid that? I don't remember this pattern in other image classification models (with different datasets) I trained with this scheduler, but I may just have forgotten it occured. Could someone shed a light on this?</p>\n<p><a href=\"https://postimg.cc/0McHgB5g\" target=\"_blank\"><img src=\"https://i.postimg.cc/7hkv2FpL/plot.png\" alt=\"plot.png\"></a></p>",
      "rawMarkdown": "In this comp we dealt with problems in making the models learn continuously. Every model we trained stopped learning (reducing loss) after epoch ~16 and a few of them resumed learning only in epochs 23 to 27. This results in a lot of epochs and time spent with no improvements.\n\nIs this a known behaviour of CosineAnnealingLR? Is there some way to avoid that? I don't remember this pattern in other image classification models (with different datasets) I trained with this scheduler, but I may just have forgotten it occured. Could someone shed a light on this?\n\n[![plot.png](https://i.postimg.cc/7hkv2FpL/plot.png)](https://postimg.cc/0McHgB5g)",
      "votes": null
    },
    {
      "id": "1801526",
      "postDate": "05/25/2022 20:32:22",
      "content": "<p>I also noticed this can be avoided by reducing the LR. But this is not really a solution… the models didn't get better and it took a lot more time.</p>",
      "rawMarkdown": "I also noticed this can be avoided by reducing the LR. But this is not really a solution... the models didn't get better and it took a lot more time.",
      "votes": null
    },
    {
      "id": "1801836",
      "postDate": "05/26/2022 07:40:23",
      "content": "<p>validation staying constant while training loss dips….I thought this is a textbook-typical behavior of a model slowly overfitting itself, and I'm quite sure this probably isn't even only about CosineAnnealing either. Our team had encountered this phonomena as well, but even when we submit models after training ~40 epochs there were no improvements in public LB score.</p>",
      "rawMarkdown": "validation staying constant while training loss dips....I thought this is a textbook-typical behavior of a model slowly overfitting itself, and I'm quite sure this probably isn't even only about CosineAnnealing either. Our team had encountered this phonomena as well, but even when we submit models after training ~40 epochs there were no improvements in public LB score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801526,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "05/25/2022 20:32:22",
      "content": "<p>I also noticed this can be avoided by reducing the LR. But this is not really a solution… the models didn't get better and it took a lot more time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801836,
      "author_name": "jayeonyi",
      "author_url": "",
      "post_date": "05/26/2022 07:40:23",
      "content": "<p>validation staying constant while training loss dips….I thought this is a textbook-typical behavior of a model slowly overfitting itself, and I'm quite sure this probably isn't even only about CosineAnnealing either. Our team had encountered this phonomena as well, but even when we submit models after training ~40 epochs there were no improvements in public LB score.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801523": "In this comp we dealt with problems in making the models learn continuously. Every model we trained stopped learning (reducing loss) after epoch ~16 and a few of them resumed learning only in epochs 23 to 27. This results in a lot of epochs and time spent with no improvements.\n\nIs this a known behaviour of CosineAnnealingLR? Is there some way to avoid that? I don't remember this pattern in other image classification models (with different datasets) I trained with this scheduler, but I may just have forgotten it occured. Could someone shed a light on this?\n\n[![plot.png](https://i.postimg.cc/7hkv2FpL/plot.png)](https://postimg.cc/0McHgB5g)",
    "1801526": "I also noticed this can be avoided by reducing the LR. But this is not really a solution... the models didn't get better and it took a lot more time.",
    "1801836": "validation staying constant while training loss dips....I thought this is a textbook-typical behavior of a model slowly overfitting itself, and I'm quite sure this probably isn't even only about CosineAnnealing either. Our team had encountered this phonomena as well, but even when we submit models after training ~40 epochs there were no improvements in public LB score."
  },
  "source": "meta"
}