{
  "id": 141961,
  "title": "A list of promising Models. ",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141961",
  "author_name": "",
  "post_date": "2020-04-08T09:11:47.576870100Z",
  "votes": 25,
  "comment_count": 14,
  "views": 0,
  "content": "<ul>\n<li><p>https://arxiv.org/pdf/1801.06146.pdf\"&gt;Universal Language Model FIne-Tuning </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1907.11692\">RoBERTa</a> | <a href=\"https://arxiv.org/abs/1909.11942\">ALBERT</a> | <a href=\"https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf\">Generative Pre-Training (GPT)</a></p></li>\n<li><p>https://arxiv.org/pdf/1906.08237v2.pdf\"&gt;XLNet | <a href=\"https://github.com/zihangdai/xlnet\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1905.07129v3.pdf\">ERNIE</a> | <a href=\"https://github.com/thunlp/ERNIE\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1910.10683.pdf\">Text-to-Text Transfer Transformer (T5)</a> | <a href=\"https://github.com/google-research/text-to-text-transfer-transformer\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1911.04070v1.pdf\">Binary-Partitioning Transformer</a> | <a href=\"https://github.com/yzh119/BPT\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1909.01259.pdf\">Neural Attentive Bag-of-Entities Model for Text Classification</a> | <a href=\"https://github.com/wikipedia2vec/wikipedia2vec/tree/master/examples/text_classification\">Code</a></p></li>\n<li><p><a href=\"https://www.aclweb.org/anthology/N19-1408.pdf\">Rethinking Complex Neural Network Architectures for Document Classification</a> | <a href=\"https://github.com/castorini/hedwig\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/2001.04451.pdf\">Reformer: The Efficient Transformer</a> | <a href=\"https://github.com/google/trax/tree/master/trax/models/reformer\">Code</a></p></li>\n<li><p><a href=\"https://openreview.net/pdf?id=r1xMH1BtvB\">ELECTRA</a> | <a href=\"https://github.com/huggingface/transformers\">Code</a></p></li>\n</ul>",
  "messages": [
    {
      "id": "801253",
      "postDate": "04/08/2020 09:11:47",
      "content": "<ul>\n<li><p>https://arxiv.org/pdf/1801.06146.pdf\"&gt;Universal Language Model FIne-Tuning </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1907.11692\">RoBERTa</a> | <a href=\"https://arxiv.org/abs/1909.11942\">ALBERT</a> | <a href=\"https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf\">Generative Pre-Training (GPT)</a></p></li>\n<li><p>https://arxiv.org/pdf/1906.08237v2.pdf\"&gt;XLNet | <a href=\"https://github.com/zihangdai/xlnet\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1905.07129v3.pdf\">ERNIE</a> | <a href=\"https://github.com/thunlp/ERNIE\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1910.10683.pdf\">Text-to-Text Transfer Transformer (T5)</a> | <a href=\"https://github.com/google-research/text-to-text-transfer-transformer\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1911.04070v1.pdf\">Binary-Partitioning Transformer</a> | <a href=\"https://github.com/yzh119/BPT\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1909.01259.pdf\">Neural Attentive Bag-of-Entities Model for Text Classification</a> | <a href=\"https://github.com/wikipedia2vec/wikipedia2vec/tree/master/examples/text_classification\">Code</a></p></li>\n<li><p><a href=\"https://www.aclweb.org/anthology/N19-1408.pdf\">Rethinking Complex Neural Network Architectures for Document Classification</a> | <a href=\"https://github.com/castorini/hedwig\">Code</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/2001.04451.pdf\">Reformer: The Efficient Transformer</a> | <a href=\"https://github.com/google/trax/tree/master/trax/models/reformer\">Code</a></p></li>\n<li><p><a href=\"https://openreview.net/pdf?id=r1xMH1BtvB\">ELECTRA</a> | <a href=\"https://github.com/huggingface/transformers\">Code</a></p></li>\n</ul>",
      "rawMarkdown": "[Universal Language Model FIne-Tuning ](chrome-extension://cbnaodkpfinfiipjblikofhlhlcickei/src/pdfviewer/web/viewer.html?file=https://arxiv.org/pdf/1801.06146.pdf)\n\n- [RoBERTa](https://arxiv.org/abs/1907.11692) | [ALBERT](https://arxiv.org/abs/1909.11942) | [Generative Pre-Training (GPT)](https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf)\n\n- [XLNet](chrome-extension://cbnaodkpfinfiipjblikofhlhlcickei/src/pdfviewer/web/viewer.html?file=https://arxiv.org/pdf/1906.08237v2.pdf) | [Code](https://github.com/zihangdai/xlnet)\n\n- [ERNIE](https://arxiv.org/pdf/1905.07129v3.pdf) | [Code](https://github.com/thunlp/ERNIE)\n\n- [Text-to-Text Transfer Transformer (T5)](https://arxiv.org/pdf/1910.10683.pdf) | [Code](https://github.com/google-research/text-to-text-transfer-transformer)\n\n- [Binary-Partitioning Transformer](https://arxiv.org/pdf/1911.04070v1.pdf) | [Code](https://github.com/yzh119/BPT)\n\n- [Neural Attentive Bag-of-Entities Model for Text Classification](https://arxiv.org/pdf/1909.01259.pdf) | [Code](https://github.com/wikipedia2vec/wikipedia2vec/tree/master/examples/text_classification)\n\n- [Rethinking Complex Neural Network Architectures for Document Classification](https://www.aclweb.org/anthology/N19-1408.pdf) | [Code](https://github.com/castorini/hedwig)\n\n- [Reformer: The Efficient Transformer](https://arxiv.org/pdf/2001.04451.pdf) | [Code](https://github.com/google/trax/tree/master/trax/models/reformer)\n\n- [ELECTRA](https://openreview.net/pdf?id=r1xMH1BtvB) | [Code](https://github.com/huggingface/transformers)",
      "votes": null
    },
    {
      "id": "801330",
      "postDate": "04/08/2020 11:13:38",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍",
      "votes": null
    },
    {
      "id": "801369",
      "postDate": "04/08/2020 12:27:52",
      "content": "<p>Thanks. What about DistillBERT?</p>",
      "rawMarkdown": "Thanks. What about DistillBERT?",
      "votes": null
    },
    {
      "id": "801374",
      "postDate": "04/08/2020 12:29:52",
      "content": "<p>Yes, it's also. Sorry, I forgot to mention. However, people may easily get informed if they go through BERT.</p>",
      "rawMarkdown": "Yes, it's also. Sorry, I forgot to mention. However, people may easily get informed if they go through BERT.",
      "votes": null
    },
    {
      "id": "801569",
      "postDate": "04/08/2020 15:42:04",
      "content": "<p>Yup, HuggingFace's transformers library is just awsome!</p>",
      "rawMarkdown": "Yup, HuggingFace's transformers library is just awsome!",
      "votes": null
    },
    {
      "id": "801593",
      "postDate": "04/08/2020 15:56:35",
      "content": "<p>Home for NLP 😄 </p>",
      "rawMarkdown": "Home for NLP 😄",
      "votes": null
    },
    {
      "id": "801952",
      "postDate": "04/09/2020 01:02:26",
      "content": "<p>This is great. Very helpful. Thanks for putting it here. </p>",
      "rawMarkdown": "This is great. Very helpful. Thanks for putting it here.",
      "votes": null
    },
    {
      "id": "802183",
      "postDate": "04/09/2020 07:58:34",
      "content": "<p>A lot of these are not multi-lingual AFAIK though.</p>",
      "rawMarkdown": "A lot of these are not multi-lingual AFAIK though.",
      "votes": null
    },
    {
      "id": "802234",
      "postDate": "04/09/2020 09:00:43",
      "content": "<p>Yes, you're right. I observed some people were trying with translated datasets too, so that's why.  And also, I was reading some SOTA papers, wanted to make a small summary of some interesting approaches. :)</p>",
      "rawMarkdown": "Yes, you're right. I observed some people were trying with translated datasets too, so that's why.  And also, I was reading some SOTA papers, wanted to make a small summary of some interesting approaches. :)",
      "votes": null
    },
    {
      "id": "802582",
      "postDate": "04/09/2020 16:25:05",
      "content": "<p>MultiFit is another one you can add to the list. Made by fast.ai </p>\n\n<p><a href=\"https://arxiv.org/abs/1909.04761\">https://arxiv.org/abs/1909.04761</a></p>",
      "rawMarkdown": "MultiFit is another one you can add to the list. Made by fast.ai \n\nhttps://arxiv.org/abs/1909.04761",
      "votes": null
    },
    {
      "id": "805681",
      "postDate": "04/12/2020 23:26:01",
      "content": "<p>Great. 1 vote from me</p>",
      "rawMarkdown": "Great. 1 vote from me",
      "votes": null
    },
    {
      "id": "806099",
      "postDate": "04/13/2020 12:45:08",
      "content": "<p>Great summary!  upvoted</p>\n\n<p>I guess RoBERTa is the only one with expertise in multilingual text classification...!?</p>",
      "rawMarkdown": "Great summary!  upvoted\n\nI guess RoBERTa is the only one with expertise in multilingual text classification...!?",
      "votes": null
    },
    {
      "id": "806114",
      "postDate": "04/13/2020 12:55:58",
      "content": "<p>Yes, for multi-lingual cases, AFAIK, <code>xlm</code>, <code>m-bert</code>, <code>xlm-roberta</code>, <code>multifit</code> are such models. However since we are allowed to translate the non-eng text to eng, so we can use mono-lingual models as well.</p>",
      "rawMarkdown": "Yes, for multi-lingual cases, AFAIK, `xlm`, `m-bert`, `xlm-roberta`, `multifit` are such models. However since we are allowed to translate the non-eng text to eng, so we can use mono-lingual models as well.",
      "votes": null
    },
    {
      "id": "806377",
      "postDate": "04/13/2020 17:08:27",
      "content": "<p>Yeah absolutely.\nThanks.</p>",
      "rawMarkdown": "Yeah absolutely.\nThanks.",
      "votes": null
    },
    {
      "id": "1057339",
      "postDate": "10/22/2020 15:10:21",
      "content": "<p>I recently made a kaggle kernel using the Reformer architecture for Named Entity Recognition and achieved an accuracy of 85%. Please check out the kernel and comment if you have any suggestions. You can find the kernel <a href=\"https://www.kaggle.com/sauravmaheshkar/trax-ner-using-reformer\" target=\"_blank\">here</a>. Could be used for Translation as well, by modifying the code a little bit. 😊</p>",
      "rawMarkdown": "I recently made a kaggle kernel using the Reformer architecture for Named Entity Recognition and achieved an accuracy of 85%. Please check out the kernel and comment if you have any suggestions. You can find the kernel [here](https://www.kaggle.com/sauravmaheshkar/trax-ner-using-reformer). Could be used for Translation as well, by modifying the code a little bit. 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1057339,
      "author_name": "sauravmaheshkar",
      "author_url": "",
      "post_date": "10/22/2020 15:10:21",
      "content": "<p>I recently made a kaggle kernel using the Reformer architecture for Named Entity Recognition and achieved an accuracy of 85%. Please check out the kernel and comment if you have any suggestions. You can find the kernel <a href=\"https://www.kaggle.com/sauravmaheshkar/trax-ner-using-reformer\" target=\"_blank\">here</a>. Could be used for Translation as well, by modifying the code a little bit. 😊</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 801330,
      "author_name": "binaicrai",
      "author_url": "",
      "post_date": "04/08/2020 11:13:38",
      "content": "<p>👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 801369,
      "author_name": "parmarsuraj99",
      "author_url": "",
      "post_date": "04/08/2020 12:27:52",
      "content": "<p>Thanks. What about DistillBERT?</p>",
      "votes": null,
      "replies": [
        {
          "id": 801374,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/08/2020 12:29:52",
          "content": "<p>Yes, it's also. Sorry, I forgot to mention. However, people may easily get informed if they go through BERT.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 801569,
          "author_name": "parmarsuraj99",
          "author_url": "",
          "post_date": "04/08/2020 15:42:04",
          "content": "<p>Yup, HuggingFace's transformers library is just awsome!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 801593,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/08/2020 15:56:35",
          "content": "<p>Home for NLP 😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 801952,
      "author_name": "tejashshah",
      "author_url": "",
      "post_date": "04/09/2020 01:02:26",
      "content": "<p>This is great. Very helpful. Thanks for putting it here. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 802183,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "04/09/2020 07:58:34",
      "content": "<p>A lot of these are not multi-lingual AFAIK though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 802234,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/09/2020 09:00:43",
          "content": "<p>Yes, you're right. I observed some people were trying with translated datasets too, so that's why.  And also, I was reading some SOTA papers, wanted to make a small summary of some interesting approaches. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 802582,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "04/09/2020 16:25:05",
      "content": "<p>MultiFit is another one you can add to the list. Made by fast.ai </p>\n\n<p><a href=\"https://arxiv.org/abs/1909.04761\">https://arxiv.org/abs/1909.04761</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 805681,
      "author_name": "podsyp",
      "author_url": "",
      "post_date": "04/12/2020 23:26:01",
      "content": "<p>Great. 1 vote from me</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 806099,
      "author_name": "namanj27",
      "author_url": "",
      "post_date": "04/13/2020 12:45:08",
      "content": "<p>Great summary!  upvoted</p>\n\n<p>I guess RoBERTa is the only one with expertise in multilingual text classification...!?</p>",
      "votes": null,
      "replies": [
        {
          "id": 806114,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "04/13/2020 12:55:58",
          "content": "<p>Yes, for multi-lingual cases, AFAIK, <code>xlm</code>, <code>m-bert</code>, <code>xlm-roberta</code>, <code>multifit</code> are such models. However since we are allowed to translate the non-eng text to eng, so we can use mono-lingual models as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 806377,
          "author_name": "namanj27",
          "author_url": "",
          "post_date": "04/13/2020 17:08:27",
          "content": "<p>Yeah absolutely.\nThanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "801253": "[Universal Language Model FIne-Tuning ](chrome-extension://cbnaodkpfinfiipjblikofhlhlcickei/src/pdfviewer/web/viewer.html?file=https://arxiv.org/pdf/1801.06146.pdf)\n\n- [RoBERTa](https://arxiv.org/abs/1907.11692) | [ALBERT](https://arxiv.org/abs/1909.11942) | [Generative Pre-Training (GPT)](https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf)\n\n- [XLNet](chrome-extension://cbnaodkpfinfiipjblikofhlhlcickei/src/pdfviewer/web/viewer.html?file=https://arxiv.org/pdf/1906.08237v2.pdf) | [Code](https://github.com/zihangdai/xlnet)\n\n- [ERNIE](https://arxiv.org/pdf/1905.07129v3.pdf) | [Code](https://github.com/thunlp/ERNIE)\n\n- [Text-to-Text Transfer Transformer (T5)](https://arxiv.org/pdf/1910.10683.pdf) | [Code](https://github.com/google-research/text-to-text-transfer-transformer)\n\n- [Binary-Partitioning Transformer](https://arxiv.org/pdf/1911.04070v1.pdf) | [Code](https://github.com/yzh119/BPT)\n\n- [Neural Attentive Bag-of-Entities Model for Text Classification](https://arxiv.org/pdf/1909.01259.pdf) | [Code](https://github.com/wikipedia2vec/wikipedia2vec/tree/master/examples/text_classification)\n\n- [Rethinking Complex Neural Network Architectures for Document Classification](https://www.aclweb.org/anthology/N19-1408.pdf) | [Code](https://github.com/castorini/hedwig)\n\n- [Reformer: The Efficient Transformer](https://arxiv.org/pdf/2001.04451.pdf) | [Code](https://github.com/google/trax/tree/master/trax/models/reformer)\n\n- [ELECTRA](https://openreview.net/pdf?id=r1xMH1BtvB) | [Code](https://github.com/huggingface/transformers)",
    "801330": "👍",
    "801369": "Thanks. What about DistillBERT?",
    "801374": "Yes, it's also. Sorry, I forgot to mention. However, people may easily get informed if they go through BERT.",
    "801569": "Yup, HuggingFace's transformers library is just awsome!",
    "801593": "Home for NLP 😄",
    "801952": "This is great. Very helpful. Thanks for putting it here.",
    "802183": "A lot of these are not multi-lingual AFAIK though.",
    "802234": "Yes, you're right. I observed some people were trying with translated datasets too, so that's why.  And also, I was reading some SOTA papers, wanted to make a small summary of some interesting approaches. :)",
    "802582": "MultiFit is another one you can add to the list. Made by fast.ai \n\nhttps://arxiv.org/abs/1909.04761",
    "805681": "Great. 1 vote from me",
    "806099": "Great summary!  upvoted\n\nI guess RoBERTa is the only one with expertise in multilingual text classification...!?",
    "806114": "Yes, for multi-lingual cases, AFAIK, `xlm`, `m-bert`, `xlm-roberta`, `multifit` are such models. However since we are allowed to translate the non-eng text to eng, so we can use mono-lingual models as well.",
    "806377": "Yeah absolutely.\nThanks.",
    "1057339": "I recently made a kaggle kernel using the Reformer architecture for Named Entity Recognition and achieved an accuracy of 85%. Please check out the kernel and comment if you have any suggestions. You can find the kernel [here](https://www.kaggle.com/sauravmaheshkar/trax-ner-using-reformer). Could be used for Translation as well, by modifying the code a little bit. 😊"
  },
  "source": "meta"
}