{
  "id": 148095,
  "title": "ZeRO & DeepSpeed: Optimizing 100 Billion Parameters! ",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/148095",
  "author_name": "Innat",
  "post_date": "2020-05-03T07:36:20.073000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'm not using it yet, just sharing because I think it's really promising.  It's a gift from <strong>Microsoft</strong>, don't worry it won't crash! 😂  </p>\n\n<p>&gt; <strong>DeepSpeed</strong> is compatible with <strong>PyTorch</strong>. One piece of that library, called <strong>ZeRO</strong>, is a new parallelized optimizer that greatly reduces the resources needed for model and data parallelism while massively increasing the number of parameters that can be trained.  </p>\n\n<p>source: \n- <a href=\"https://www.microsoft.com/en-us/research/blog/zero-deepspeed-new-system-optimizations-enable-training-models-with-over-100-billion-parameters/\">ZeRO &amp; DeepSpeed: New system optimizations enable training models with over 100 billion parameters</a></p>\n\n<ul>\n<li><a href=\"https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/\">Turing-NLG: A 17-billion-parameter language model by Microsoft</a></li>\n</ul>\n\n<p>git code: <a href=\"https://github.com/microsoft/DeepSpeed\">https://github.com/microsoft/DeepSpeed</a></p>",
  "messages": [
    {
      "id": 831155,
      "postDate": "2020-05-03T07:36:20.073Z",
      "content": "<p>I'm not using it yet, just sharing because I think it's really promising.  It's a gift from <strong>Microsoft</strong>, don't worry it won't crash! 😂  </p>\n\n<p>&gt; <strong>DeepSpeed</strong> is compatible with <strong>PyTorch</strong>. One piece of that library, called <strong>ZeRO</strong>, is a new parallelized optimizer that greatly reduces the resources needed for model and data parallelism while massively increasing the number of parameters that can be trained.  </p>\n\n<p>source: \n- <a href=\"https://www.microsoft.com/en-us/research/blog/zero-deepspeed-new-system-optimizations-enable-training-models-with-over-100-billion-parameters/\">ZeRO &amp; DeepSpeed: New system optimizations enable training models with over 100 billion parameters</a></p>\n\n<ul>\n<li><a href=\"https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/\">Turing-NLG: A 17-billion-parameter language model by Microsoft</a></li>\n</ul>\n\n<p>git code: <a href=\"https://github.com/microsoft/DeepSpeed\">https://github.com/microsoft/DeepSpeed</a></p>",
      "rawMarkdown": "I'm not using it yet, just sharing because I think it's really promising.  It's a gift from **Microsoft**, don't worry it won't crash! 😂  \n\n&gt; **DeepSpeed** is compatible with **PyTorch**. One piece of that library, called **ZeRO**, is a new parallelized optimizer that greatly reduces the resources needed for model and data parallelism while massively increasing the number of parameters that can be trained.  \n\nsource: \n- [ZeRO &amp; DeepSpeed: New system optimizations enable training models with over 100 billion parameters](https://www.microsoft.com/en-us/research/blog/zero-deepspeed-new-system-optimizations-enable-training-models-with-over-100-billion-parameters/)\n\n- [Turing-NLG: A 17-billion-parameter language model by Microsoft](https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/)\n\ngit code: https://github.com/microsoft/DeepSpeed",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "831155": "I'm not using it yet, just sharing because I think it's really promising.  It's a gift from **Microsoft**, don't worry it won't crash! 😂  \n\n&gt; **DeepSpeed** is compatible with **PyTorch**. One piece of that library, called **ZeRO**, is a new parallelized optimizer that greatly reduces the resources needed for model and data parallelism while massively increasing the number of parameters that can be trained.  \n\nsource: \n- [ZeRO &amp; DeepSpeed: New system optimizations enable training models with over 100 billion parameters](https://www.microsoft.com/en-us/research/blog/zero-deepspeed-new-system-optimizations-enable-training-models-with-over-100-billion-parameters/)\n\n- [Turing-NLG: A 17-billion-parameter language model by Microsoft](https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/)\n\ngit code: https://github.com/microsoft/DeepSpeed"
  }
}