{
  "id": 126702,
  "title": "BERT & Friends Reference",
  "url": "/competitions/tensorflow2-question-answering/discussion/126702",
  "author_name": "",
  "post_date": "2020-01-19T14:11:09.821465200Z",
  "votes": 37,
  "comment_count": 21,
  "views": 0,
  "content": "<p><img src=\"https://s3.amazonaws.com/images.seroundtable.com/google-bert-global-1575952051.jpg\" alt=\"\"></p>\n\n<p>I gathered here few papers, blog posts, repositories about BERT and variants.</p>\n\n<ol>\n<li>Jacob Devlin Ming-Wei Chang Kenton Lee Kristina Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, <a href=\"https://arxiv.org/abs/1810.04805\">https://arxiv.org/abs/1810.04805</a>  </li>\n<li>Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, Veselin Stoyanov, RoBERTa: A Robustly Optimized BERT Pretraining Approach, <a href=\"https://arxiv.org/abs/1907.11692\">https://arxiv.org/abs/1907.11692</a>    </li>\n<li>Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut, ALBERT: A Lite BERT for Self-supervised Learning of Language Representations, <a href=\"https://arxiv.org/abs/1909.11942\">https://arxiv.org/abs/1909.11942</a>   </li>\n<li>Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, Qun Liu, TinyBERT: Distilling BERT for Natural Language Understanding, <a href=\"https://arxiv.org/abs/1909.10351\">https://arxiv.org/abs/1909.10351</a>  </li>\n<li>J.S. McCarley, Pruning a BERT-based Question Answering Model, <a href=\"https://arxiv.org/abs/1910.06360\">https://arxiv.org/abs/1910.06360</a>   </li>\n<li>Rani Horev, BERT Explained: State of the art language model for NLP, <a href=\"https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270\">https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270</a>    </li>\n<li>Less Wright, Meet ALBERT: a new ‘Lite BERT’ from Google &amp; Toyota with State of the Art NLP performance and 18x fewer parameters,  <a href=\"https://medium.com/@lessw/meet-albert-a-new-lite-bert-from-google-toyota-with-state-of-the-art-nlp-performance-and-18x-df8f7b58fa28\">https://medium.com/@lessw/meet-albert-a-new-lite-bert-from-google-toyota-with-state-of-the-art-nlp-performance-and-18x-df8f7b58fa28</a>   </li>\n<li>Suleiman Khan, BERT, RoBERTa, DistilBERT, XLNet — which one to use?, <a href=\"https://towardsdatascience.com/bert-roberta-distilbert-xlnet-which-one-to-use-3d5ab82ba5f8\">https://towardsdatascience.com/bert-roberta-distilbert-xlnet-which-one-to-use-3d5ab82ba5f8</a>  </li>\n<li>Radu Soricut and Zhenzhong Lan, ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations,  <a href=\"https://ai.googleblog.com/2019/12/albert-lite-bert-for-self-supervised.html\">https://ai.googleblog.com/2019/12/albert-lite-bert-for-self-supervised.html</a>    </li>\n<li>Arun Maiya, BERT Text Classification in 3 Lines of Code Using Keras, <a href=\"https://towardsdatascience.com/bert-text-classification-in-3-lines-of-code-using-keras-264db7e7a358\">https://towardsdatascience.com/bert-text-classification-in-3-lines-of-code-using-keras-264db7e7a358</a>   </li>\n<li>Aaron (Ari) Bornstein, Beyond Word Embeddings Part 2: Word Vectors and NLP Modeling from BoW to BERT, <a href=\"https://towardsdatascience.com/beyond-word-embeddings-part-2-word-vectors-nlp-modeling-from-bow-to-bert-4ebd4711d0ec\">https://towardsdatascience.com/beyond-word-embeddings-part-2-word-vectors-nlp-modeling-from-bow-to-bert-4ebd4711d0ec</a>   </li>\n<li>Miguel Romero Calvo, Dissecting BERT Part 1: The Encoder, <a href=\"https://medium.com/dissecting-bert/dissecting-bert-part-1-d3c3d495cdb3\">https://medium.com/dissecting-bert/dissecting-bert-part-1-d3c3d495cdb3</a>    </li>\n<li>Francisco Ingham,  Understanding BERT Part 2: BERT Specifics, <a href=\"https://medium.com/dissecting-bert/dissecting-bert-part2-335ff2ed9c73\">https://medium.com/dissecting-bert/dissecting-bert-part2-335ff2ed9c73</a>   </li>\n<li>Miguel Romero Calvo, Dissecting BERT Appendix: The Decoder, <a href=\"https://medium.com/dissecting-bert/dissecting-bert-appendix-the-decoder-3b86f66b0e5f\">https://medium.com/dissecting-bert/dissecting-bert-appendix-the-decoder-3b86f66b0e5f</a>   </li>\n<li>Jay Alammar, The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning), <a href=\"http://jalammar.github.io/illustrated-bert/\">http://jalammar.github.io/illustrated-bert/</a>   </li>\n<li>Jay Alammar, A Visual Guide to Using BERT for the First Time, <a href=\"http://jalammar.github.io/\">http://jalammar.github.io/</a>   </li>\n<li>Fast-BERT, <a href=\"https://github.com/kaushaltrivedi/fast-bert\">https://github.com/kaushaltrivedi/fast-bert</a>    </li>\n<li>Chi Sun, Xipeng Qiu, Yige Xu, Xuanjing Huang, How to Fine-Tune BERT for Text Classification?, <a href=\"https://arxiv.org/abs/1905.05583\">https://arxiv.org/abs/1905.05583</a></li>\n</ol>",
  "messages": [
    {
      "id": "723103",
      "postDate": "01/19/2020 14:11:09",
      "content": "<p><img src=\"https://s3.amazonaws.com/images.seroundtable.com/google-bert-global-1575952051.jpg\" alt=\"\"></p>\n\n<p>I gathered here few papers, blog posts, repositories about BERT and variants.</p>\n\n<ol>\n<li>Jacob Devlin Ming-Wei Chang Kenton Lee Kristina Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, <a href=\"https://arxiv.org/abs/1810.04805\">https://arxiv.org/abs/1810.04805</a>  </li>\n<li>Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, Veselin Stoyanov, RoBERTa: A Robustly Optimized BERT Pretraining Approach, <a href=\"https://arxiv.org/abs/1907.11692\">https://arxiv.org/abs/1907.11692</a>    </li>\n<li>Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut, ALBERT: A Lite BERT for Self-supervised Learning of Language Representations, <a href=\"https://arxiv.org/abs/1909.11942\">https://arxiv.org/abs/1909.11942</a>   </li>\n<li>Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, Qun Liu, TinyBERT: Distilling BERT for Natural Language Understanding, <a href=\"https://arxiv.org/abs/1909.10351\">https://arxiv.org/abs/1909.10351</a>  </li>\n<li>J.S. McCarley, Pruning a BERT-based Question Answering Model, <a href=\"https://arxiv.org/abs/1910.06360\">https://arxiv.org/abs/1910.06360</a>   </li>\n<li>Rani Horev, BERT Explained: State of the art language model for NLP, <a href=\"https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270\">https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270</a>    </li>\n<li>Less Wright, Meet ALBERT: a new ‘Lite BERT’ from Google &amp; Toyota with State of the Art NLP performance and 18x fewer parameters,  <a href=\"https://medium.com/@lessw/meet-albert-a-new-lite-bert-from-google-toyota-with-state-of-the-art-nlp-performance-and-18x-df8f7b58fa28\">https://medium.com/@lessw/meet-albert-a-new-lite-bert-from-google-toyota-with-state-of-the-art-nlp-performance-and-18x-df8f7b58fa28</a>   </li>\n<li>Suleiman Khan, BERT, RoBERTa, DistilBERT, XLNet — which one to use?, <a href=\"https://towardsdatascience.com/bert-roberta-distilbert-xlnet-which-one-to-use-3d5ab82ba5f8\">https://towardsdatascience.com/bert-roberta-distilbert-xlnet-which-one-to-use-3d5ab82ba5f8</a>  </li>\n<li>Radu Soricut and Zhenzhong Lan, ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations,  <a href=\"https://ai.googleblog.com/2019/12/albert-lite-bert-for-self-supervised.html\">https://ai.googleblog.com/2019/12/albert-lite-bert-for-self-supervised.html</a>    </li>\n<li>Arun Maiya, BERT Text Classification in 3 Lines of Code Using Keras, <a href=\"https://towardsdatascience.com/bert-text-classification-in-3-lines-of-code-using-keras-264db7e7a358\">https://towardsdatascience.com/bert-text-classification-in-3-lines-of-code-using-keras-264db7e7a358</a>   </li>\n<li>Aaron (Ari) Bornstein, Beyond Word Embeddings Part 2: Word Vectors and NLP Modeling from BoW to BERT, <a href=\"https://towardsdatascience.com/beyond-word-embeddings-part-2-word-vectors-nlp-modeling-from-bow-to-bert-4ebd4711d0ec\">https://towardsdatascience.com/beyond-word-embeddings-part-2-word-vectors-nlp-modeling-from-bow-to-bert-4ebd4711d0ec</a>   </li>\n<li>Miguel Romero Calvo, Dissecting BERT Part 1: The Encoder, <a href=\"https://medium.com/dissecting-bert/dissecting-bert-part-1-d3c3d495cdb3\">https://medium.com/dissecting-bert/dissecting-bert-part-1-d3c3d495cdb3</a>    </li>\n<li>Francisco Ingham,  Understanding BERT Part 2: BERT Specifics, <a href=\"https://medium.com/dissecting-bert/dissecting-bert-part2-335ff2ed9c73\">https://medium.com/dissecting-bert/dissecting-bert-part2-335ff2ed9c73</a>   </li>\n<li>Miguel Romero Calvo, Dissecting BERT Appendix: The Decoder, <a href=\"https://medium.com/dissecting-bert/dissecting-bert-appendix-the-decoder-3b86f66b0e5f\">https://medium.com/dissecting-bert/dissecting-bert-appendix-the-decoder-3b86f66b0e5f</a>   </li>\n<li>Jay Alammar, The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning), <a href=\"http://jalammar.github.io/illustrated-bert/\">http://jalammar.github.io/illustrated-bert/</a>   </li>\n<li>Jay Alammar, A Visual Guide to Using BERT for the First Time, <a href=\"http://jalammar.github.io/\">http://jalammar.github.io/</a>   </li>\n<li>Fast-BERT, <a href=\"https://github.com/kaushaltrivedi/fast-bert\">https://github.com/kaushaltrivedi/fast-bert</a>    </li>\n<li>Chi Sun, Xipeng Qiu, Yige Xu, Xuanjing Huang, How to Fine-Tune BERT for Text Classification?, <a href=\"https://arxiv.org/abs/1905.05583\">https://arxiv.org/abs/1905.05583</a></li>\n</ol>",
      "rawMarkdown": "![](https://s3.amazonaws.com/images.seroundtable.com/google-bert-global-1575952051.jpg)\n\nI gathered here few papers, blog posts, repositories about BERT and variants.\n\n1. Jacob Devlin Ming-Wei Chang Kenton Lee Kristina Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, https://arxiv.org/abs/1810.04805  \n2. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, Veselin Stoyanov, RoBERTa: A Robustly Optimized BERT Pretraining Approach, https://arxiv.org/abs/1907.11692    \n3. Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut, ALBERT: A Lite BERT for Self-supervised Learning of Language Representations, https://arxiv.org/abs/1909.11942   \n4. Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, Qun Liu, TinyBERT: Distilling BERT for Natural Language Understanding, https://arxiv.org/abs/1909.10351  \n5. J.S. McCarley, Pruning a BERT-based Question Answering Model, https://arxiv.org/abs/1910.06360   \n6. Rani Horev, BERT Explained: State of the art language model for NLP, https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270    \n7. Less Wright, Meet ALBERT: a new ‘Lite BERT’ from Google &amp; Toyota with State of the Art NLP performance and 18x fewer parameters,  https://medium.com/@lessw/meet-albert-a-new-lite-bert-from-google-toyota-with-state-of-the-art-nlp-performance-and-18x-df8f7b58fa28   \n8. Suleiman Khan, BERT, RoBERTa, DistilBERT, XLNet — which one to use?, https://towardsdatascience.com/bert-roberta-distilbert-xlnet-which-one-to-use-3d5ab82ba5f8  \n9.  Radu Soricut and Zhenzhong Lan, ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations,  https://ai.googleblog.com/2019/12/albert-lite-bert-for-self-supervised.html    \n10. Arun Maiya, BERT Text Classification in 3 Lines of Code Using Keras, https://towardsdatascience.com/bert-text-classification-in-3-lines-of-code-using-keras-264db7e7a358   \n11. Aaron (Ari) Bornstein, Beyond Word Embeddings Part 2: Word Vectors and NLP Modeling from BoW to BERT, https://towardsdatascience.com/beyond-word-embeddings-part-2-word-vectors-nlp-modeling-from-bow-to-bert-4ebd4711d0ec   \n12. Miguel Romero Calvo, Dissecting BERT Part 1: The Encoder, https://medium.com/dissecting-bert/dissecting-bert-part-1-d3c3d495cdb3    \n13. Francisco Ingham,  Understanding BERT Part 2: BERT Specifics, https://medium.com/dissecting-bert/dissecting-bert-part2-335ff2ed9c73   \n14. Miguel Romero Calvo, Dissecting BERT Appendix: The Decoder, https://medium.com/dissecting-bert/dissecting-bert-appendix-the-decoder-3b86f66b0e5f   \n15. Jay Alammar, The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning), http://jalammar.github.io/illustrated-bert/   \n16. Jay Alammar, A Visual Guide to Using BERT for the First Time, http://jalammar.github.io/   \n17. Fast-BERT, https://github.com/kaushaltrivedi/fast-bert    \n18. Chi Sun, Xipeng Qiu, Yige Xu, Xuanjing Huang, How to Fine-Tune BERT for Text Classification?, https://arxiv.org/abs/1905.05583",
      "votes": null
    },
    {
      "id": "723140",
      "postDate": "01/19/2020 14:48:19",
      "content": "<p>Great \nVery Helpful\nThanks for Sharing <a href=\"/gpreda\">@gpreda</a> </p>",
      "rawMarkdown": "Great \nVery Helpful\nThanks for Sharing @gpreda",
      "votes": null
    },
    {
      "id": "723162",
      "postDate": "01/19/2020 15:17:28",
      "content": "<p>Would definitely add here Jay Alammar's posts. </p>",
      "rawMarkdown": "Would definitely add here Jay Alammar's posts.",
      "votes": null
    },
    {
      "id": "723164",
      "postDate": "01/19/2020 15:19:19",
      "content": "<p>Thank you for the suggestion. I will.</p>",
      "rawMarkdown": "Thank you for the suggestion. I will.",
      "votes": null
    },
    {
      "id": "723231",
      "postDate": "01/19/2020 17:39:09",
      "content": "<p>Really helpful\nthank you for sharing. <a href=\"/gpreda\">@gpreda</a> </p>",
      "rawMarkdown": "Really helpful\nthank you for sharing. @gpreda",
      "votes": null
    },
    {
      "id": "723484",
      "postDate": "01/20/2020 05:07:24",
      "content": "<p>Great work.</p>\n\n<p>Thank you for sharing.</p>",
      "rawMarkdown": "Great work.\n\nThank you for sharing.",
      "votes": null
    },
    {
      "id": "723680",
      "postDate": "01/20/2020 10:19:27",
      "content": "<p>Nice gathering! Thank you for sharing!</p>\n\n<p>I found also the paper <a href=\"https://arxiv.org/abs/1905.05583\">\"How to Fine-Tune BERT for Text Classification?\"</a> especially interesting.\nIt presents a lot of different approaches to fine-tune BERT &amp; his Friends.\nReally worth reading!</p>",
      "rawMarkdown": "Nice gathering! Thank you for sharing!\n\nI found also the paper [\"How to Fine-Tune BERT for Text Classification?\"](https://arxiv.org/abs/1905.05583) especially interesting.\nIt presents a lot of different approaches to fine-tune BERT &amp; his Friends.\nReally worth reading!",
      "votes": null
    },
    {
      "id": "723681",
      "postDate": "01/20/2020 10:24:41",
      "content": "<p>Thank you for the suggestion. I included in the list.</p>",
      "rawMarkdown": "Thank you for the suggestion. I included in the list.",
      "votes": null
    },
    {
      "id": "723734",
      "postDate": "01/20/2020 11:46:25",
      "content": "<p>Great work! Thanks for sharing <a href=\"/gpreda\">@gpreda</a> 😄 👍 </p>",
      "rawMarkdown": "Great work! Thanks for sharing @gpreda 😄 👍",
      "votes": null
    },
    {
      "id": "723848",
      "postDate": "01/20/2020 14:25:11",
      "content": "<p>Great resources, thanks!</p>",
      "rawMarkdown": "Great resources, thanks!",
      "votes": null
    },
    {
      "id": "724595",
      "postDate": "01/21/2020 09:26:35",
      "content": "<p>Thanks for share 💯 </p>",
      "rawMarkdown": "Thanks for share 💯",
      "votes": null
    },
    {
      "id": "726043",
      "postDate": "01/22/2020 18:37:32",
      "content": "<p>Can You suggest which the papers explains how to train own BERT model on own language copus? The most papers I found is about using pre-trained model.</p>",
      "rawMarkdown": "Can You suggest which the papers explains how to train own BERT model on own language copus? The most papers I found is about using pre-trained model.",
      "votes": null
    },
    {
      "id": "726055",
      "postDate": "01/22/2020 18:50:18",
      "content": "<p><a href=\"/peterpirog\">@peterpirog</a> The paper \"RoBERTa: A Robustly Optimized BERT Pretraining Approach\" describes well how to train a \"virgin model\". It is available <a href=\"https://arxiv.org/abs/1907.11692\">here</a>.</p>",
      "rawMarkdown": "peterpirog The paper \"RoBERTa: A Robustly Optimized BERT Pretraining Approach\" describes well how to train a \"virgin model\". It is available [here](https://arxiv.org/abs/1907.11692).",
      "votes": null
    },
    {
      "id": "726076",
      "postDate": "01/22/2020 19:27:55",
      "content": "<p>Thank You, I would like to train BERT on polish language corpus I hope this paper will be useful for me :)</p>",
      "rawMarkdown": "Thank You, I would like to train BERT on polish language corpus I hope this paper will be useful for me :)",
      "votes": null
    },
    {
      "id": "726083",
      "postDate": "01/22/2020 19:39:35",
      "content": "<p><a href=\"/peterpirog\">@peterpirog</a> I never read it but I think this paper should help you --&gt; <a href=\"https://arxiv.org/abs/1911.03894\">CamemBERT: a Tasty French Language Model</a>\nThey describe how they trained BERT on french language corpus.</p>\n\n<p>Hope it will help!</p>",
      "rawMarkdown": "peterpirog I never read it but I think this paper should help you --&gt; [CamemBERT: a Tasty French Language Model](https://arxiv.org/abs/1911.03894)\nThey describe how they trained BERT on french language corpus.\n\nHope it will help!",
      "votes": null
    },
    {
      "id": "726741",
      "postDate": "01/23/2020 07:58:16",
      "content": "<p>The reason is that training BERT requires large computational effort. Here is a small post about compared training time using TPU vs. using GPU for training BERT. With a large GPU cluster takes few days: <a href=\"https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert/\">https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert/</a></p>",
      "rawMarkdown": "The reason is that training BERT requires large computational effort. Here is a small post about compared training time using TPU vs. using GPU for training BERT. With a large GPU cluster takes few days: https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert/",
      "votes": null
    },
    {
      "id": "727262",
      "postDate": "01/23/2020 15:27:40",
      "content": "<p>I can agree with authors of the paper suggested above: \n \"Despite their success, most available models have either been trained on English data or on the concatenation of data in multiple languages. This makes practical use of such models—in all languages except English—very limited.\"\nFor example polish langage has 7 declensions and and many forms of verb (tenses, person, number, voice, ascpect and mode) so number of unique words is huge, even creating of representative  lanuage corpus is problematic.\nI hope theese papers will help me to create BERT model for polish. </p>",
      "rawMarkdown": "I can agree with authors of the paper suggested above: \n \"Despite their success, most available models have either been trained on English data or on the concatenation of data in multiple languages. This makes practical use of such models—in all languages except English—very limited.\"\nFor example polish langage has 7 declensions and and many forms of verb (tenses, person, number, voice, ascpect and mode) so number of unique words is huge, even creating of representative  lanuage corpus is problematic.\nI hope theese papers will help me to create BERT model for polish.",
      "votes": null
    },
    {
      "id": "727328",
      "postDate": "01/23/2020 16:25:37",
      "content": "<p>I just found something else!</p>\n\n<p>While searching in the <a href=\"https://huggingface.co/transformers/pretrained_models.html\">pretrained models in the library transformers</a> I found <code>bert-base-multilingual-cased</code>, <code>bert-base-multilingual-uncased</code> that has been both trained with multiple languages with Polish included.\nTake time to look carefuly at all the list, I saw other pretrained models that looks interesting like <code>xlm-roberta-base</code> or <code>xlm-mlm-100-1280</code>.</p>",
      "rawMarkdown": "I just found something else!\n\nWhile searching in the [pretrained models in the library transformers](https://huggingface.co/transformers/pretrained_models.html) I found `bert-base-multilingual-cased`, `bert-base-multilingual-uncased` that has been both trained with multiple languages with Polish included.\nTake time to look carefuly at all the list, I saw other pretrained models that looks interesting like `xlm-roberta-base` or `xlm-mlm-100-1280`.",
      "votes": null
    },
    {
      "id": "727340",
      "postDate": "01/23/2020 16:36:00",
      "content": "<p>It should save you days of training time!</p>",
      "rawMarkdown": "It should save you days of training time!",
      "votes": null
    },
    {
      "id": "727584",
      "postDate": "01/23/2020 21:02:30",
      "content": "<p>Thx, I have to check it :)</p>",
      "rawMarkdown": "Thx, I have to check it :)",
      "votes": null
    },
    {
      "id": "733316",
      "postDate": "01/31/2020 00:24:59",
      "content": "<p>Thanks for great summary</p>",
      "rawMarkdown": "Thanks for great summary",
      "votes": null
    },
    {
      "id": "738289",
      "postDate": "02/06/2020 10:47:50",
      "content": "<p>Great job! ,Thanks for sharing.</p>",
      "rawMarkdown": "Great job! ,Thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 723140,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "01/19/2020 14:48:19",
      "content": "<p>Great \nVery Helpful\nThanks for Sharing <a href=\"/gpreda\">@gpreda</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 723162,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "01/19/2020 15:17:28",
      "content": "<p>Would definitely add here Jay Alammar's posts. </p>",
      "votes": null,
      "replies": [
        {
          "id": 723164,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "01/19/2020 15:19:19",
          "content": "<p>Thank you for the suggestion. I will.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 723231,
      "author_name": "yashagrawal300",
      "author_url": "",
      "post_date": "01/19/2020 17:39:09",
      "content": "<p>Really helpful\nthank you for sharing. <a href=\"/gpreda\">@gpreda</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 723484,
      "author_name": "prashant111",
      "author_url": "",
      "post_date": "01/20/2020 05:07:24",
      "content": "<p>Great work.</p>\n\n<p>Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 723680,
      "author_name": "maroberti",
      "author_url": "",
      "post_date": "01/20/2020 10:19:27",
      "content": "<p>Nice gathering! Thank you for sharing!</p>\n\n<p>I found also the paper <a href=\"https://arxiv.org/abs/1905.05583\">\"How to Fine-Tune BERT for Text Classification?\"</a> especially interesting.\nIt presents a lot of different approaches to fine-tune BERT &amp; his Friends.\nReally worth reading!</p>",
      "votes": null,
      "replies": [
        {
          "id": 723681,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "01/20/2020 10:24:41",
          "content": "<p>Thank you for the suggestion. I included in the list.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 723734,
      "author_name": "mashlyn",
      "author_url": "",
      "post_date": "01/20/2020 11:46:25",
      "content": "<p>Great work! Thanks for sharing <a href=\"/gpreda\">@gpreda</a> 😄 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 723848,
      "author_name": "anlgrbz",
      "author_url": "",
      "post_date": "01/20/2020 14:25:11",
      "content": "<p>Great resources, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 724595,
      "author_name": "dasmehdixtr",
      "author_url": "",
      "post_date": "01/21/2020 09:26:35",
      "content": "<p>Thanks for share 💯 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 726043,
      "author_name": "peterpirog",
      "author_url": "",
      "post_date": "01/22/2020 18:37:32",
      "content": "<p>Can You suggest which the papers explains how to train own BERT model on own language copus? The most papers I found is about using pre-trained model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 726055,
          "author_name": "maroberti",
          "author_url": "",
          "post_date": "01/22/2020 18:50:18",
          "content": "<p><a href=\"/peterpirog\">@peterpirog</a> The paper \"RoBERTa: A Robustly Optimized BERT Pretraining Approach\" describes well how to train a \"virgin model\". It is available <a href=\"https://arxiv.org/abs/1907.11692\">here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 726076,
          "author_name": "peterpirog",
          "author_url": "",
          "post_date": "01/22/2020 19:27:55",
          "content": "<p>Thank You, I would like to train BERT on polish language corpus I hope this paper will be useful for me :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 726083,
          "author_name": "maroberti",
          "author_url": "",
          "post_date": "01/22/2020 19:39:35",
          "content": "<p><a href=\"/peterpirog\">@peterpirog</a> I never read it but I think this paper should help you --&gt; <a href=\"https://arxiv.org/abs/1911.03894\">CamemBERT: a Tasty French Language Model</a>\nThey describe how they trained BERT on french language corpus.</p>\n\n<p>Hope it will help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 726741,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "01/23/2020 07:58:16",
          "content": "<p>The reason is that training BERT requires large computational effort. Here is a small post about compared training time using TPU vs. using GPU for training BERT. With a large GPU cluster takes few days: <a href=\"https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert/\">https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 727262,
          "author_name": "peterpirog",
          "author_url": "",
          "post_date": "01/23/2020 15:27:40",
          "content": "<p>I can agree with authors of the paper suggested above: \n \"Despite their success, most available models have either been trained on English data or on the concatenation of data in multiple languages. This makes practical use of such models—in all languages except English—very limited.\"\nFor example polish langage has 7 declensions and and many forms of verb (tenses, person, number, voice, ascpect and mode) so number of unique words is huge, even creating of representative  lanuage corpus is problematic.\nI hope theese papers will help me to create BERT model for polish. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 727328,
          "author_name": "maroberti",
          "author_url": "",
          "post_date": "01/23/2020 16:25:37",
          "content": "<p>I just found something else!</p>\n\n<p>While searching in the <a href=\"https://huggingface.co/transformers/pretrained_models.html\">pretrained models in the library transformers</a> I found <code>bert-base-multilingual-cased</code>, <code>bert-base-multilingual-uncased</code> that has been both trained with multiple languages with Polish included.\nTake time to look carefuly at all the list, I saw other pretrained models that looks interesting like <code>xlm-roberta-base</code> or <code>xlm-mlm-100-1280</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 727340,
          "author_name": "maroberti",
          "author_url": "",
          "post_date": "01/23/2020 16:36:00",
          "content": "<p>It should save you days of training time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 727584,
      "author_name": "peterpirog",
      "author_url": "",
      "post_date": "01/23/2020 21:02:30",
      "content": "<p>Thx, I have to check it :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 733316,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/31/2020 00:24:59",
      "content": "<p>Thanks for great summary</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 738289,
      "author_name": "",
      "author_url": "",
      "post_date": "02/06/2020 10:47:50",
      "content": "<p>Great job! ,Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "723103": "![](https://s3.amazonaws.com/images.seroundtable.com/google-bert-global-1575952051.jpg)\n\nI gathered here few papers, blog posts, repositories about BERT and variants.\n\n1. Jacob Devlin Ming-Wei Chang Kenton Lee Kristina Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, https://arxiv.org/abs/1810.04805  \n2. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, Veselin Stoyanov, RoBERTa: A Robustly Optimized BERT Pretraining Approach, https://arxiv.org/abs/1907.11692    \n3. Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut, ALBERT: A Lite BERT for Self-supervised Learning of Language Representations, https://arxiv.org/abs/1909.11942   \n4. Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, Qun Liu, TinyBERT: Distilling BERT for Natural Language Understanding, https://arxiv.org/abs/1909.10351  \n5. J.S. McCarley, Pruning a BERT-based Question Answering Model, https://arxiv.org/abs/1910.06360   \n6. Rani Horev, BERT Explained: State of the art language model for NLP, https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270    \n7. Less Wright, Meet ALBERT: a new ‘Lite BERT’ from Google &amp; Toyota with State of the Art NLP performance and 18x fewer parameters,  https://medium.com/@lessw/meet-albert-a-new-lite-bert-from-google-toyota-with-state-of-the-art-nlp-performance-and-18x-df8f7b58fa28   \n8. Suleiman Khan, BERT, RoBERTa, DistilBERT, XLNet — which one to use?, https://towardsdatascience.com/bert-roberta-distilbert-xlnet-which-one-to-use-3d5ab82ba5f8  \n9.  Radu Soricut and Zhenzhong Lan, ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations,  https://ai.googleblog.com/2019/12/albert-lite-bert-for-self-supervised.html    \n10. Arun Maiya, BERT Text Classification in 3 Lines of Code Using Keras, https://towardsdatascience.com/bert-text-classification-in-3-lines-of-code-using-keras-264db7e7a358   \n11. Aaron (Ari) Bornstein, Beyond Word Embeddings Part 2: Word Vectors and NLP Modeling from BoW to BERT, https://towardsdatascience.com/beyond-word-embeddings-part-2-word-vectors-nlp-modeling-from-bow-to-bert-4ebd4711d0ec   \n12. Miguel Romero Calvo, Dissecting BERT Part 1: The Encoder, https://medium.com/dissecting-bert/dissecting-bert-part-1-d3c3d495cdb3    \n13. Francisco Ingham,  Understanding BERT Part 2: BERT Specifics, https://medium.com/dissecting-bert/dissecting-bert-part2-335ff2ed9c73   \n14. Miguel Romero Calvo, Dissecting BERT Appendix: The Decoder, https://medium.com/dissecting-bert/dissecting-bert-appendix-the-decoder-3b86f66b0e5f   \n15. Jay Alammar, The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning), http://jalammar.github.io/illustrated-bert/   \n16. Jay Alammar, A Visual Guide to Using BERT for the First Time, http://jalammar.github.io/   \n17. Fast-BERT, https://github.com/kaushaltrivedi/fast-bert    \n18. Chi Sun, Xipeng Qiu, Yige Xu, Xuanjing Huang, How to Fine-Tune BERT for Text Classification?, https://arxiv.org/abs/1905.05583",
    "723140": "Great \nVery Helpful\nThanks for Sharing @gpreda",
    "723162": "Would definitely add here Jay Alammar's posts.",
    "723164": "Thank you for the suggestion. I will.",
    "723231": "Really helpful\nthank you for sharing. @gpreda",
    "723484": "Great work.\n\nThank you for sharing.",
    "723680": "Nice gathering! Thank you for sharing!\n\nI found also the paper [\"How to Fine-Tune BERT for Text Classification?\"](https://arxiv.org/abs/1905.05583) especially interesting.\nIt presents a lot of different approaches to fine-tune BERT &amp; his Friends.\nReally worth reading!",
    "723681": "Thank you for the suggestion. I included in the list.",
    "723734": "Great work! Thanks for sharing @gpreda 😄 👍",
    "723848": "Great resources, thanks!",
    "724595": "Thanks for share 💯",
    "726043": "Can You suggest which the papers explains how to train own BERT model on own language copus? The most papers I found is about using pre-trained model.",
    "726055": "peterpirog The paper \"RoBERTa: A Robustly Optimized BERT Pretraining Approach\" describes well how to train a \"virgin model\". It is available [here](https://arxiv.org/abs/1907.11692).",
    "726076": "Thank You, I would like to train BERT on polish language corpus I hope this paper will be useful for me :)",
    "726083": "peterpirog I never read it but I think this paper should help you --&gt; [CamemBERT: a Tasty French Language Model](https://arxiv.org/abs/1911.03894)\nThey describe how they trained BERT on french language corpus.\n\nHope it will help!",
    "726741": "The reason is that training BERT requires large computational effort. Here is a small post about compared training time using TPU vs. using GPU for training BERT. With a large GPU cluster takes few days: https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert/",
    "727262": "I can agree with authors of the paper suggested above: \n \"Despite their success, most available models have either been trained on English data or on the concatenation of data in multiple languages. This makes practical use of such models—in all languages except English—very limited.\"\nFor example polish langage has 7 declensions and and many forms of verb (tenses, person, number, voice, ascpect and mode) so number of unique words is huge, even creating of representative  lanuage corpus is problematic.\nI hope theese papers will help me to create BERT model for polish.",
    "727328": "I just found something else!\n\nWhile searching in the [pretrained models in the library transformers](https://huggingface.co/transformers/pretrained_models.html) I found `bert-base-multilingual-cased`, `bert-base-multilingual-uncased` that has been both trained with multiple languages with Polish included.\nTake time to look carefuly at all the list, I saw other pretrained models that looks interesting like `xlm-roberta-base` or `xlm-mlm-100-1280`.",
    "727340": "It should save you days of training time!",
    "727584": "Thx, I have to check it :)",
    "733316": "Thanks for great summary",
    "738289": "Great job! ,Thanks for sharing."
  },
  "source": "meta"
}