{
  "id": 264958,
  "title": "Is a custom learning rates needed with the Adam optimiser ",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/264958",
  "author_name": "",
  "post_date": "2021-08-13T23:53:50.114985200Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I notice many PyTourch notebooks use of a custom learning  rate (i.e., CosineAnnealingLR)  with the Adam optimiser. </p>\n<p>I understood that  Adam was an adaptive learning rate algorithm  designed specifically for training deep neural networks so is the customer learning rate necessary?</p>",
  "messages": [
    {
      "id": "1471092",
      "postDate": "08/13/2021 23:53:50",
      "content": "<p>I notice many PyTourch notebooks use of a custom learning  rate (i.e., CosineAnnealingLR)  with the Adam optimiser. </p>\n<p>I understood that  Adam was an adaptive learning rate algorithm  designed specifically for training deep neural networks so is the customer learning rate necessary?</p>",
      "rawMarkdown": "I notice many PyTourch notebooks use of a custom learning  rate (i.e., CosineAnnealingLR)  with the Adam optimiser. \n\nI understood that  Adam was an adaptive learning rate algorithm  designed specifically for training deep neural networks so is the customer learning rate necessary?",
      "votes": null
    },
    {
      "id": "1471256",
      "postDate": "08/14/2021 04:45:51",
      "content": "<p>Yep, decay algorithms (like CosineAnnealingLR) is external (your) signal for optimizer change internal parameters (and rights prior knowledge about changing one is very helpfull for more accurate optimization). </p>",
      "rawMarkdown": "Yep, decay algorithms (like CosineAnnealingLR) is external (your) signal for optimizer change internal parameters (and rights prior knowledge about changing one is very helpfull for more accurate optimization).",
      "votes": null
    },
    {
      "id": "1471352",
      "postDate": "08/14/2021 06:39:42",
      "content": "<p><strong>Adam is very picky optimizer</strong>( Things like warmup and cosine annealing are must to do with it. Also it may be quite sensitive to lr selection. So be mindful (and/or use more modern optimizer).</p>",
      "rawMarkdown": "**Adam is very picky optimizer**( Things like warmup and cosine annealing are must to do with it. Also it may be quite sensitive to lr selection. So be mindful (and/or use more modern optimizer).",
      "votes": null
    },
    {
      "id": "1471799",
      "postDate": "08/14/2021 13:23:49",
      "content": "<p>I'm a happy user of adam with cosine annealing scheduler, but you make me will to try other optimizers.</p>",
      "rawMarkdown": "I'm a happy user of adam with cosine annealing scheduler, but you make me will to try other optimizers.",
      "votes": null
    },
    {
      "id": "1474809",
      "postDate": "08/16/2021 09:25:13",
      "content": "<p><a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a>, thank you for the insights. What do you suggest AdamW?</p>",
      "rawMarkdown": "lafoss, thank you for the insights. What do you suggest AdamW?",
      "votes": null
    },
    {
      "id": "1475770",
      "postDate": "08/16/2021 20:26:56",
      "content": "<p>Most often I use RAdam+LARS+LookAhead, but it may not work well for some networks like EfficientNet. Also out of the quite recent optimizer, you can check Madgrad. <br>\nBut my recommendation would be just take some small training setup (so u can get results for each run within 10min-1h), and run a number of experiments to see the specifics of a particular optimizer you chose (like lr sensitivity, need for warmup, training curves, wd, etc… and another thing is that optimizers may exhibit slightly different behavior with different network families). It is not just blind replacing an optimizer and hoping for a better score. And if the things are right, another optimizer may give you just a slight improvement.</p>",
      "rawMarkdown": "Most often I use RAdam+LARS+LookAhead, but it may not work well for some networks like EfficientNet. Also out of the quite recent optimizer, you can check Madgrad. \nBut my recommendation would be just take some small training setup (so u can get results for each run within 10min-1h), and run a number of experiments to see the specifics of a particular optimizer you chose (like lr sensitivity, need for warmup, training curves, wd, etc... and another thing is that optimizers may exhibit slightly different behavior with different network families). It is not just blind replacing an optimizer and hoping for a better score. And if the things are right, another optimizer may give you just a slight improvement.",
      "votes": null
    },
    {
      "id": "1476305",
      "postDate": "08/17/2021 04:37:27",
      "content": "<p>I concur, it is more important to be able to test rapidly ideas than to train a huge model.  I tuned my pipeline using a setting that was taking about 10 minutes per epoch with a V100 GPU. It would run in less than 20 minutes per epoch on Kaggle kernels.</p>",
      "rawMarkdown": "I concur, it is more important to be able to test rapidly ideas than to train a huge model.  I tuned my pipeline using a setting that was taking about 10 minutes per epoch with a V100 GPU. It would run in less than 20 minutes per epoch on Kaggle kernels.",
      "votes": null
    },
    {
      "id": "1561213",
      "postDate": "10/27/2021 12:20:31",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1471256,
      "author_name": "miklgr500",
      "author_url": "",
      "post_date": "08/14/2021 04:45:51",
      "content": "<p>Yep, decay algorithms (like CosineAnnealingLR) is external (your) signal for optimizer change internal parameters (and rights prior knowledge about changing one is very helpfull for more accurate optimization). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1471352,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "08/14/2021 06:39:42",
      "content": "<p><strong>Adam is very picky optimizer</strong>( Things like warmup and cosine annealing are must to do with it. Also it may be quite sensitive to lr selection. So be mindful (and/or use more modern optimizer).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1471799,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/14/2021 13:23:49",
          "content": "<p>I'm a happy user of adam with cosine annealing scheduler, but you make me will to try other optimizers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1474809,
          "author_name": "kevinmcisaac",
          "author_url": "",
          "post_date": "08/16/2021 09:25:13",
          "content": "<p><a href=\"https://www.kaggle.com/lafoss\" target=\"_blank\">@lafoss</a>, thank you for the insights. What do you suggest AdamW?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1475770,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "08/16/2021 20:26:56",
          "content": "<p>Most often I use RAdam+LARS+LookAhead, but it may not work well for some networks like EfficientNet. Also out of the quite recent optimizer, you can check Madgrad. <br>\nBut my recommendation would be just take some small training setup (so u can get results for each run within 10min-1h), and run a number of experiments to see the specifics of a particular optimizer you chose (like lr sensitivity, need for warmup, training curves, wd, etc… and another thing is that optimizers may exhibit slightly different behavior with different network families). It is not just blind replacing an optimizer and hoping for a better score. And if the things are right, another optimizer may give you just a slight improvement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1476305,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/17/2021 04:37:27",
          "content": "<p>I concur, it is more important to be able to test rapidly ideas than to train a huge model.  I tuned my pipeline using a setting that was taking about 10 minutes per epoch with a V100 GPU. It would run in less than 20 minutes per epoch on Kaggle kernels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1561213,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 12:20:31",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1471092": "I notice many PyTourch notebooks use of a custom learning  rate (i.e., CosineAnnealingLR)  with the Adam optimiser. \n\nI understood that  Adam was an adaptive learning rate algorithm  designed specifically for training deep neural networks so is the customer learning rate necessary?",
    "1471256": "Yep, decay algorithms (like CosineAnnealingLR) is external (your) signal for optimizer change internal parameters (and rights prior knowledge about changing one is very helpfull for more accurate optimization).",
    "1471352": "**Adam is very picky optimizer**( Things like warmup and cosine annealing are must to do with it. Also it may be quite sensitive to lr selection. So be mindful (and/or use more modern optimizer).",
    "1471799": "I'm a happy user of adam with cosine annealing scheduler, but you make me will to try other optimizers.",
    "1474809": "lafoss, thank you for the insights. What do you suggest AdamW?",
    "1475770": "Most often I use RAdam+LARS+LookAhead, but it may not work well for some networks like EfficientNet. Also out of the quite recent optimizer, you can check Madgrad. \nBut my recommendation would be just take some small training setup (so u can get results for each run within 10min-1h), and run a number of experiments to see the specifics of a particular optimizer you chose (like lr sensitivity, need for warmup, training curves, wd, etc... and another thing is that optimizers may exhibit slightly different behavior with different network families). It is not just blind replacing an optimizer and hoping for a better score. And if the things are right, another optimizer may give you just a slight improvement.",
    "1476305": "I concur, it is more important to be able to test rapidly ideas than to train a huge model.  I tuned my pipeline using a setting that was taking about 10 minutes per epoch with a V100 GPU. It would run in less than 20 minutes per epoch on Kaggle kernels.",
    "1561213": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}