{
  "id": 240316,
  "title": "LAMB - \"Large Batch Optimization for Deep Learning: Training BERT in 76 minutes\"",
  "url": "/competitions/bms-molecular-translation/discussion/240316",
  "author_name": "",
  "post_date": "2021-05-19T08:38:23.623489200Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Here is an optimizer that maybe can help.</p>\n<p>\"With the advent of large scale datasets, training large deep neural networks, even using computationally efficient optimization methods like Stochastic gradient descent (SGD), has become particularly challenging. For instance, training state-of-the-art deep learning models like BERT and ResNet-50 takes 3 days on 16 TPUv3 chips and 29 hours on 8 Tesla P100 gpus respectively (Devlin et al., 2018;He et al., 2016). Thus, there is a growing interest to develop optimization solutions to tackle this critical issue. The goal of this paper is to investigate and develop optimization techniques to accelerate training large deep neural networks, mostly focusing on approaches based on variants of SGD.\"</p>\n<p><a href=\"https://arxiv.org/abs/1904.00962\" target=\"_blank\">https://arxiv.org/abs/1904.00962</a></p>\n<p><a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/LAMB\" target=\"_blank\">https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/LAMB</a></p>",
  "messages": [
    {
      "id": "1314566",
      "postDate": "05/19/2021 08:38:23",
      "content": "<p>Here is an optimizer that maybe can help.</p>\n<p>\"With the advent of large scale datasets, training large deep neural networks, even using computationally efficient optimization methods like Stochastic gradient descent (SGD), has become particularly challenging. For instance, training state-of-the-art deep learning models like BERT and ResNet-50 takes 3 days on 16 TPUv3 chips and 29 hours on 8 Tesla P100 gpus respectively (Devlin et al., 2018;He et al., 2016). Thus, there is a growing interest to develop optimization solutions to tackle this critical issue. The goal of this paper is to investigate and develop optimization techniques to accelerate training large deep neural networks, mostly focusing on approaches based on variants of SGD.\"</p>\n<p><a href=\"https://arxiv.org/abs/1904.00962\" target=\"_blank\">https://arxiv.org/abs/1904.00962</a></p>\n<p><a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/LAMB\" target=\"_blank\">https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/LAMB</a></p>",
      "rawMarkdown": "Here is an optimizer that maybe can help.\n\n\"With the advent of large scale datasets, training large deep neural networks, even using computationally efficient optimization methods like Stochastic gradient descent (SGD), has become particularly challenging. For instance, training state-of-the-art deep learning models like BERT and ResNet-50 takes 3 days on 16 TPUv3 chips and 29 hours on 8 Tesla P100 gpus respectively (Devlin et al., 2018;He et al., 2016). Thus, there is a growing interest to develop optimization solutions to tackle this critical issue. The goal of this paper is to investigate and develop optimization techniques to accelerate training large deep neural networks, mostly focusing on approaches based on variants of SGD.\"\n\nhttps://arxiv.org/abs/1904.00962\n\nhttps://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/LAMB",
      "votes": null
    },
    {
      "id": "1315216",
      "postDate": "05/19/2021 16:23:51",
      "content": "<p>This is awesome. Will definitely look into this.</p>",
      "rawMarkdown": "This is awesome. Will definitely look into this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1315216,
      "author_name": "ankitp013",
      "author_url": "",
      "post_date": "05/19/2021 16:23:51",
      "content": "<p>This is awesome. Will definitely look into this.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1314566": "Here is an optimizer that maybe can help.\n\n\"With the advent of large scale datasets, training large deep neural networks, even using computationally efficient optimization methods like Stochastic gradient descent (SGD), has become particularly challenging. For instance, training state-of-the-art deep learning models like BERT and ResNet-50 takes 3 days on 16 TPUv3 chips and 29 hours on 8 Tesla P100 gpus respectively (Devlin et al., 2018;He et al., 2016). Thus, there is a growing interest to develop optimization solutions to tackle this critical issue. The goal of this paper is to investigate and develop optimization techniques to accelerate training large deep neural networks, mostly focusing on approaches based on variants of SGD.\"\n\nhttps://arxiv.org/abs/1904.00962\n\nhttps://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/LAMB",
    "1315216": "This is awesome. Will definitely look into this."
  },
  "source": "meta"
}