{
  "id": 70780,
  "title": "The year of transfer learning in NLP?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/70780",
  "author_name": "Yury Kashnitsky",
  "post_date": "2018-11-07T09:59:42.416000",
  "votes": 45,
  "comment_count": 4,
  "views": 0,
  "content": "<p>According to Jeremy Howard, Sebastian Ruder &amp; co., \"2018 is the year of transfer learning in NLP\" (a nice <a href=\"http://ruder.io/nlp-imagenet/\">post</a> by Sebastian on the topic). Well... this statement is still needed to be verified, more benchmarks needed. But indeed, the idea that you can get value from bulks of unlabeled data to improve classification for your scarce labeled data, is very promising (at least, for 2 of my business tasks, it's extremely relevant). So word vectors (w2V, GloVe, fasttext etc.) have already been used for quite a while. But a more interested approach with language model embeddings is proposed in  ELMo, ULMFiT, and OpenAI's \"sentiment neuron\" etc.</p>\n\n<p>It's really a very interesting idea to test: whether indeed language models provide better embeddings than \"classic\" approaches. It's bit sad to see only data for the competition limited to these 4 types of embeddings. In my opinion, Kaggle can gradually turn into a platform for testing scientific hypotheses as well (like in this case, whether language models outperform simpler approaches), but we don't see it in this competition.</p>\n\n<p>But anyway, , thanks a lot to Quora and organizers, it's very cool to have such a dataset for experiments. </p>",
  "messages": [
    {
      "id": 416807,
      "postDate": "2018-11-07T09:59:42.417Z",
      "content": "<p>According to Jeremy Howard, Sebastian Ruder &amp; co., \"2018 is the year of transfer learning in NLP\" (a nice <a href=\"http://ruder.io/nlp-imagenet/\">post</a> by Sebastian on the topic). Well... this statement is still needed to be verified, more benchmarks needed. But indeed, the idea that you can get value from bulks of unlabeled data to improve classification for your scarce labeled data, is very promising (at least, for 2 of my business tasks, it's extremely relevant). So word vectors (w2V, GloVe, fasttext etc.) have already been used for quite a while. But a more interested approach with language model embeddings is proposed in  ELMo, ULMFiT, and OpenAI's \"sentiment neuron\" etc.</p>\n\n<p>It's really a very interesting idea to test: whether indeed language models provide better embeddings than \"classic\" approaches. It's bit sad to see only data for the competition limited to these 4 types of embeddings. In my opinion, Kaggle can gradually turn into a platform for testing scientific hypotheses as well (like in this case, whether language models outperform simpler approaches), but we don't see it in this competition.</p>\n\n<p>But anyway, , thanks a lot to Quora and organizers, it's very cool to have such a dataset for experiments. </p>",
      "rawMarkdown": "According to Jeremy Howard, Sebastian Ruder &amp; co., \"2018 is the year of transfer learning in NLP\" (a nice [post][1] by Sebastian on the topic). Well... this statement is still needed to be verified, more benchmarks needed. But indeed, the idea that you can get value from bulks of unlabeled data to improve classification for your scarce labeled data, is very promising (at least, for 2 of my business tasks, it's extremely relevant). So word vectors (w2V, GloVe, fasttext etc.) have already been used for quite a while. But a more interested approach with language model embeddings is proposed in  ELMo, ULMFiT, and OpenAI's \"sentiment neuron\" etc.\n\nIt's really a very interesting idea to test: whether indeed language models provide better embeddings than \"classic\" approaches. It's bit sad to see only data for the competition limited to these 4 types of embeddings. In my opinion, Kaggle can gradually turn into a platform for testing scientific hypotheses as well (like in this case, whether language models outperform simpler approaches), but we don't see it in this competition.\n\nBut anyway, , thanks a lot to Quora and organizers, it's very cool to have such a dataset for experiments. \n\n\n\n\n  [1]: http://ruder.io/nlp-imagenet/",
      "votes": 44
    },
    {
      "id": 417637,
      "postDate": "2018-11-08T14:57:53.023Z",
      "content": "<p>Couldn't agree more. Hopefully, the data is available for download, so we can at least compare the two methods on our own machines and see if it's worth pushing the rules to allow transfer learning from language models :)</p>",
      "rawMarkdown": "Couldn't agree more. Hopefully, the data is available for download, so we can at least compare the two methods on our own machines and see if it's worth pushing the rules to allow transfer learning from language models :)",
      "votes": 1
    },
    {
      "id": 449071,
      "postDate": "2019-01-02T15:48:47.263Z",
      "content": "<p>See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: <a href=\"https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub\">https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub</a></p>",
      "rawMarkdown": "See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub"
    },
    {
      "id": 417562,
      "postDate": "2018-11-08T13:10:22.223Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 417638,
          "postDate": "2018-11-08T14:58:46.607Z",
          "content": "<p><a href=\"http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html\">http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html</a></p>",
          "rawMarkdown": "http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 417637,
      "author_name": "Quentin Retourne",
      "author_url": "",
      "post_date": "2018-11-08T14:57:53.023000",
      "content": "<p>Couldn't agree more. Hopefully, the data is available for download, so we can at least compare the two methods on our own machines and see if it's worth pushing the rules to allow transfer learning from language models :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 449071,
      "author_name": "ManuelSH",
      "author_url": "",
      "post_date": "2019-01-02T15:48:47.263000",
      "content": "<p>See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: <a href=\"https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub\">https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 417562,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-08T13:10:22.223000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 417638,
          "author_name": "Quentin Retourne",
          "author_url": "",
          "post_date": "2018-11-08T14:58:46.607000",
          "content": "<p><a href=\"http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html\">http://nlp.fast.ai/classification/2018/05/15/introducting-ulmfit.html</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "416807": "According to Jeremy Howard, Sebastian Ruder &amp; co., \"2018 is the year of transfer learning in NLP\" (a nice [post][1] by Sebastian on the topic). Well... this statement is still needed to be verified, more benchmarks needed. But indeed, the idea that you can get value from bulks of unlabeled data to improve classification for your scarce labeled data, is very promising (at least, for 2 of my business tasks, it's extremely relevant). So word vectors (w2V, GloVe, fasttext etc.) have already been used for quite a while. But a more interested approach with language model embeddings is proposed in  ELMo, ULMFiT, and OpenAI's \"sentiment neuron\" etc.\n\nIt's really a very interesting idea to test: whether indeed language models provide better embeddings than \"classic\" approaches. It's bit sad to see only data for the competition limited to these 4 types of embeddings. In my opinion, Kaggle can gradually turn into a platform for testing scientific hypotheses as well (like in this case, whether language models outperform simpler approaches), but we don't see it in this competition.\n\nBut anyway, , thanks a lot to Quora and organizers, it's very cool to have such a dataset for experiments. \n\n\n\n\n  [1]: http://ruder.io/nlp-imagenet/",
    "417637": "Couldn't agree more. Hopefully, the data is available for download, so we can at least compare the two methods on our own machines and see if it's worth pushing the rules to allow transfer learning from language models :)",
    "449071": "See here an implementation of ULMFIT with fastai library, excluding the first step which trains with an external disallowed corpus: https://www.kaggle.com/manuelsh/ulmfit-from-fast-ai-pub",
    "417562": ""
  }
}