{
  "id": 229557,
  "title": "MADGRAD: A New Optimizer to Try",
  "url": "/competitions/bms-molecular-translation/discussion/229557",
  "author_name": "",
  "post_date": "2021-03-30T18:22:12.271003Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Facebook Research has recently introduced a new optimizer called <strong>MADGRAD</strong>: Momentumized, Adaptive, Dual Averaged Gradient Method. I have not tried it yet, but this is what they claim: </p>\n<blockquote>\n  <p>We introduce MADGRAD, a novel optimization method in the family of AdaGrad adaptive gradient methods. MADGRAD shows excellent performance on deep learning optimization problems from multiple fields, including classification and image-to-image tasks in vision, and recurrent and bidirectionally-masked models in natural language processing. For each of these tasks, MADGRAD matches or outperforms both SGD and ADAM in test set performance, even on problems for which adaptive methods normally perform poorly.</p>\n</blockquote>\n<p>If you work in PyTorch, you can do <code>!pip install madgrad</code> and plug it into the pipeline. I plan to give it a try soon and see if it can be useful on this data.</p>\n<p>Links:</p>\n<ul>\n<li>GitHub: <a href=\"https://github.com/facebookresearch/madgrad\" target=\"_blank\">https://github.com/facebookresearch/madgrad</a></li>\n<li>documentation: <a href=\"https://madgrad.readthedocs.io/en/latest/\" target=\"_blank\">https://madgrad.readthedocs.io/en/latest/</a></li>\n<li>paper: <a href=\"https://arxiv.org/abs/2101.11075\" target=\"_blank\">https://arxiv.org/abs/2101.11075</a></li>\n</ul>",
  "messages": [
    {
      "id": "1257332",
      "postDate": "03/30/2021 18:22:12",
      "content": "<p>Facebook Research has recently introduced a new optimizer called <strong>MADGRAD</strong>: Momentumized, Adaptive, Dual Averaged Gradient Method. I have not tried it yet, but this is what they claim: </p>\n<blockquote>\n  <p>We introduce MADGRAD, a novel optimization method in the family of AdaGrad adaptive gradient methods. MADGRAD shows excellent performance on deep learning optimization problems from multiple fields, including classification and image-to-image tasks in vision, and recurrent and bidirectionally-masked models in natural language processing. For each of these tasks, MADGRAD matches or outperforms both SGD and ADAM in test set performance, even on problems for which adaptive methods normally perform poorly.</p>\n</blockquote>\n<p>If you work in PyTorch, you can do <code>!pip install madgrad</code> and plug it into the pipeline. I plan to give it a try soon and see if it can be useful on this data.</p>\n<p>Links:</p>\n<ul>\n<li>GitHub: <a href=\"https://github.com/facebookresearch/madgrad\" target=\"_blank\">https://github.com/facebookresearch/madgrad</a></li>\n<li>documentation: <a href=\"https://madgrad.readthedocs.io/en/latest/\" target=\"_blank\">https://madgrad.readthedocs.io/en/latest/</a></li>\n<li>paper: <a href=\"https://arxiv.org/abs/2101.11075\" target=\"_blank\">https://arxiv.org/abs/2101.11075</a></li>\n</ul>",
      "rawMarkdown": "Facebook Research has recently introduced a new optimizer called **MADGRAD**: Momentumized, Adaptive, Dual Averaged Gradient Method. I have not tried it yet, but this is what they claim: \n> We introduce MADGRAD, a novel optimization method in the family of AdaGrad adaptive gradient methods. MADGRAD shows excellent performance on deep learning optimization problems from multiple fields, including classification and image-to-image tasks in vision, and recurrent and bidirectionally-masked models in natural language processing. For each of these tasks, MADGRAD matches or outperforms both SGD and ADAM in test set performance, even on problems for which adaptive methods normally perform poorly.\n\nIf you work in PyTorch, you can do `!pip install madgrad` and plug it into the pipeline. I plan to give it a try soon and see if it can be useful on this data.\n\nLinks:\n- GitHub: https://github.com/facebookresearch/madgrad\n- documentation: https://madgrad.readthedocs.io/en/latest/\n- paper: https://arxiv.org/abs/2101.11075",
      "votes": null
    },
    {
      "id": "1257599",
      "postDate": "03/31/2021 01:32:40",
      "content": "<p>Tried it in a different competition. It's pretty solid.</p>",
      "rawMarkdown": "Tried it in a different competition. It's pretty solid.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1257599,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "03/31/2021 01:32:40",
      "content": "<p>Tried it in a different competition. It's pretty solid.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1257332": "Facebook Research has recently introduced a new optimizer called **MADGRAD**: Momentumized, Adaptive, Dual Averaged Gradient Method. I have not tried it yet, but this is what they claim: \n> We introduce MADGRAD, a novel optimization method in the family of AdaGrad adaptive gradient methods. MADGRAD shows excellent performance on deep learning optimization problems from multiple fields, including classification and image-to-image tasks in vision, and recurrent and bidirectionally-masked models in natural language processing. For each of these tasks, MADGRAD matches or outperforms both SGD and ADAM in test set performance, even on problems for which adaptive methods normally perform poorly.\n\nIf you work in PyTorch, you can do `!pip install madgrad` and plug it into the pipeline. I plan to give it a try soon and see if it can be useful on this data.\n\nLinks:\n- GitHub: https://github.com/facebookresearch/madgrad\n- documentation: https://madgrad.readthedocs.io/en/latest/\n- paper: https://arxiv.org/abs/2101.11075",
    "1257599": "Tried it in a different competition. It's pretty solid."
  },
  "source": "meta"
}