{
  "id": 70880,
  "title": "ULMFit,Transformer Models for classification (SoTA Approaches)",
  "url": "/competitions/quora-insincere-questions-classification/discussion/70880",
  "author_name": "",
  "post_date": "2018-11-08T04:42:43.430735900Z",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>While Google BERT has shown to achieve SoTA results on various NLP tasks, it is too expensive to train on kernel, and ULMFit achives SoTA results on text classification and takes reasonable compute.  </p>\n\n<ul>\n<li><p>ULMFit: <a href=\"https://arxiv.org/abs/1801.06146\">Paper</a>,<a href=\"http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html\">fastai Blog</a>, <a href=\"https://www.youtube.com/watch?v=zxJJ0T54HX8\">Video (External)</a>, </p></li>\n<li><p>Google BERT: <a href=\"https://arxiv.org/abs/1810.04805\">Paper</a></p></li>\n</ul>\n\n<p>Also you can use AWD LSTM as well which forms the core of ULMFit.</p>\n\n<ul>\n<li>AWD-LSTM: <a href=\"https://arxiv.org/pdf/1708.02182.pdf\">Paper</a>, <a href=\"https://github.com/salesforce/awd-lstm-lm\">Github</a></li>\n</ul>\n\n<p>Transformer Models can also be used for better results\n- Attention is all you need: <a href=\"https://arxiv.org/abs/1706.03762\">Paper</a></p>\n\n<ul>\n<li><p>ELMO: <a href=\"https://arxiv.org/abs/1802.05365\">Paper</a></p></li>\n<li><p>OpenAI Transformer : <a href=\"https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf\">Paper</a>, <a href=\"https://blog.openai.com/language-unsupervised/\">OpenAI Blog</a>, <a href=\"https://github.com/huggingface/pytorch-openai-transformer-lm\">Pytorch implementation</a></p></li>\n</ul>\n\n<p>Not sure how transformer models will scale well on kernels yet.</p>\n\n<p>Finally for more papers and resources, check <a href=\"https://github.com/sebastianruder/NLP-progress\">here</a></p>",
  "messages": [
    {
      "id": "417321",
      "postDate": "11/08/2018 04:42:43",
      "content": "<p>While Google BERT has shown to achieve SoTA results on various NLP tasks, it is too expensive to train on kernel, and ULMFit achives SoTA results on text classification and takes reasonable compute.  </p>\n\n<ul>\n<li><p>ULMFit: <a href=\"https://arxiv.org/abs/1801.06146\">Paper</a>,<a href=\"http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html\">fastai Blog</a>, <a href=\"https://www.youtube.com/watch?v=zxJJ0T54HX8\">Video (External)</a>, </p></li>\n<li><p>Google BERT: <a href=\"https://arxiv.org/abs/1810.04805\">Paper</a></p></li>\n</ul>\n\n<p>Also you can use AWD LSTM as well which forms the core of ULMFit.</p>\n\n<ul>\n<li>AWD-LSTM: <a href=\"https://arxiv.org/pdf/1708.02182.pdf\">Paper</a>, <a href=\"https://github.com/salesforce/awd-lstm-lm\">Github</a></li>\n</ul>\n\n<p>Transformer Models can also be used for better results\n- Attention is all you need: <a href=\"https://arxiv.org/abs/1706.03762\">Paper</a></p>\n\n<ul>\n<li><p>ELMO: <a href=\"https://arxiv.org/abs/1802.05365\">Paper</a></p></li>\n<li><p>OpenAI Transformer : <a href=\"https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf\">Paper</a>, <a href=\"https://blog.openai.com/language-unsupervised/\">OpenAI Blog</a>, <a href=\"https://github.com/huggingface/pytorch-openai-transformer-lm\">Pytorch implementation</a></p></li>\n</ul>\n\n<p>Not sure how transformer models will scale well on kernels yet.</p>\n\n<p>Finally for more papers and resources, check <a href=\"https://github.com/sebastianruder/NLP-progress\">here</a></p>",
      "rawMarkdown": "While Google BERT has shown to achieve SoTA results on various NLP tasks, it is too expensive to train on kernel, and ULMFit achives SoTA results on text classification and takes reasonable compute.  \n\n- ULMFit: [Paper](https://arxiv.org/abs/1801.06146),[fastai Blog](http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html), [Video (External)](https://www.youtube.com/watch?v=zxJJ0T54HX8), \n\n- Google BERT: [Paper](https://arxiv.org/abs/1810.04805)\n\nAlso you can use AWD LSTM as well which forms the core of ULMFit.\n\n- AWD-LSTM: [Paper](https://arxiv.org/pdf/1708.02182.pdf), [Github](https://github.com/salesforce/awd-lstm-lm)\n\nTransformer Models can also be used for better results\n- Attention is all you need: [Paper](https://arxiv.org/abs/1706.03762)\n\n- ELMO: [Paper](https://arxiv.org/abs/1802.05365)\n\n- OpenAI Transformer : [Paper](https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf), [OpenAI Blog](https://blog.openai.com/language-unsupervised/), [Pytorch implementation](https://github.com/huggingface/pytorch-openai-transformer-lm)\n\nNot sure how transformer models will scale well on kernels yet.\n\nFinally for more papers and resources, check [here](https://github.com/sebastianruder/NLP-progress)",
      "votes": null
    },
    {
      "id": "417343",
      "postDate": "11/08/2018 06:02:03",
      "content": "<p>As far as I know this competition disallows external data and internet connection for kernel, which means pretrained models can't be used, so I think the SOTA approaches (unsupervised learning on a large corpus then finetune on a specific task) is not viable.</p>",
      "rawMarkdown": "As far as I know this competition disallows external data and internet connection for kernel, which means pretrained models can't be used, so I think the SOTA approaches (unsupervised learning on a large corpus then finetune on a specific task) is not viable.",
      "votes": null
    },
    {
      "id": "417351",
      "postDate": "11/08/2018 06:20:09",
      "content": "<p>I really hope organizers of this competition will allow using external data or will extend available data to language models and other modern approaches.</p>",
      "rawMarkdown": "I really hope organizers of this competition will allow using external data or will extend available data to language models and other modern approaches.",
      "votes": null
    },
    {
      "id": "417407",
      "postDate": "11/08/2018 08:01:44",
      "content": "<p>I think transformer models can be adapted to kernels ( I didn't try it though ^^) </p>",
      "rawMarkdown": "I think transformer models can be adapted to kernels ( I didn't try it though ^^)",
      "votes": null
    },
    {
      "id": "449070",
      "postDate": "01/02/2019 15:47:51",
      "content": "<p>See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: <a href=\"https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub\">https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub</a></p>",
      "rawMarkdown": "See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 417343,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "11/08/2018 06:02:03",
      "content": "<p>As far as I know this competition disallows external data and internet connection for kernel, which means pretrained models can't be used, so I think the SOTA approaches (unsupervised learning on a large corpus then finetune on a specific task) is not viable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417351,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "11/08/2018 06:20:09",
      "content": "<p>I really hope organizers of this competition will allow using external data or will extend available data to language models and other modern approaches.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417407,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "11/08/2018 08:01:44",
      "content": "<p>I think transformer models can be adapted to kernels ( I didn't try it though ^^) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449070,
      "author_name": "manuelsh",
      "author_url": "",
      "post_date": "01/02/2019 15:47:51",
      "content": "<p>See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: <a href=\"https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub\">https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "417321": "While Google BERT has shown to achieve SoTA results on various NLP tasks, it is too expensive to train on kernel, and ULMFit achives SoTA results on text classification and takes reasonable compute.  \n\n- ULMFit: [Paper](https://arxiv.org/abs/1801.06146),[fastai Blog](http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html), [Video (External)](https://www.youtube.com/watch?v=zxJJ0T54HX8), \n\n- Google BERT: [Paper](https://arxiv.org/abs/1810.04805)\n\nAlso you can use AWD LSTM as well which forms the core of ULMFit.\n\n- AWD-LSTM: [Paper](https://arxiv.org/pdf/1708.02182.pdf), [Github](https://github.com/salesforce/awd-lstm-lm)\n\nTransformer Models can also be used for better results\n- Attention is all you need: [Paper](https://arxiv.org/abs/1706.03762)\n\n- ELMO: [Paper](https://arxiv.org/abs/1802.05365)\n\n- OpenAI Transformer : [Paper](https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf), [OpenAI Blog](https://blog.openai.com/language-unsupervised/), [Pytorch implementation](https://github.com/huggingface/pytorch-openai-transformer-lm)\n\nNot sure how transformer models will scale well on kernels yet.\n\nFinally for more papers and resources, check [here](https://github.com/sebastianruder/NLP-progress)",
    "417343": "As far as I know this competition disallows external data and internet connection for kernel, which means pretrained models can't be used, so I think the SOTA approaches (unsupervised learning on a large corpus then finetune on a specific task) is not viable.",
    "417351": "I really hope organizers of this competition will allow using external data or will extend available data to language models and other modern approaches.",
    "417407": "I think transformer models can be adapted to kernels ( I didn't try it though ^^)",
    "449070": "See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub"
  },
  "source": "meta"
}