{
  "id": 159806,
  "title": "Some NLP terminology",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/159806",
  "author_name": "Yassine Alouini",
  "post_date": "2020-06-18T18:51:21.127000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Some of the ML terms that I have heard/read employed for describing some NLP tasks (some of these aren't specific to NLP):</p>\n\n<ul>\n<li><strong>Fine-tuning</strong> on dataset X: taking a pre-trained model and training it on a different dataset, usually a smaller and more specialized one.</li>\n<li><strong>Fine-tuning</strong> on task  Z: taking a pre-trained model and training it to predict outputs for task Z.</li>\n<li><strong>Transfer learning</strong>: using a model trained for task X to do task Y. Sometimes task X is generic whereas Y is specific (X: MLM, Y: binary label prediction).</li>\n<li><strong>Pre-trained</strong> model: for example <a href=\"https://huggingface.co/roberta-base\"><code>roberta-base</code></a> is a pre-trained model</li>\n<li><strong>Backbone/base</strong> model: the first layer of the model that is frozen. Nowadays, it is usually a transformer variation.</li>\n<li><strong>Tokenization</strong>: extracting tokens from the input text. There are many different <a href=\"https://github.com/huggingface/tokenizers\">tokenizers</a>. </li>\n<li><strong>MLM</strong> (Masked Language Model): masking some tokens at random and predicting them.</li>\n<li><strong>NER</strong> (Named entity recognition): predicting the type of each entity (name, verb, pronoun, date, place, and so on)</li>\n<li><strong>NSP</strong> (Next Sentence Prediction): predicting whether the second sentence follows the first one (in logical/causal way) or not. </li>\n<li><strong>GLUE</strong>: General Language Understanding Evaluation, a popular NLP benchmark</li>\n<li><strong>AdamW</strong>: Adam with weight decay, a very popular gradient descent optimizer. More about it <a href=\"https://www.fast.ai/2018/07/02/adam-weight-decay/\">here</a>. </li>\n</ul>\n\n<p>Let me know if these  terms are correct/used  (especially the first four terms related to training tasks) and if not please suggest corrections. Also, before leaving, check this great <a href=\"https://www.kaggle.com/rftexas/nlp-cheatsheet-master-nlp\"><strong>notebook</strong></a> packed with many more terms. Thanks in advance for your comments! </p>",
  "messages": [
    {
      "id": 892264,
      "postDate": "2020-06-18T18:51:21.127Z",
      "content": "<p>Some of the ML terms that I have heard/read employed for describing some NLP tasks (some of these aren't specific to NLP):</p>\n\n<ul>\n<li><strong>Fine-tuning</strong> on dataset X: taking a pre-trained model and training it on a different dataset, usually a smaller and more specialized one.</li>\n<li><strong>Fine-tuning</strong> on task  Z: taking a pre-trained model and training it to predict outputs for task Z.</li>\n<li><strong>Transfer learning</strong>: using a model trained for task X to do task Y. Sometimes task X is generic whereas Y is specific (X: MLM, Y: binary label prediction).</li>\n<li><strong>Pre-trained</strong> model: for example <a href=\"https://huggingface.co/roberta-base\"><code>roberta-base</code></a> is a pre-trained model</li>\n<li><strong>Backbone/base</strong> model: the first layer of the model that is frozen. Nowadays, it is usually a transformer variation.</li>\n<li><strong>Tokenization</strong>: extracting tokens from the input text. There are many different <a href=\"https://github.com/huggingface/tokenizers\">tokenizers</a>. </li>\n<li><strong>MLM</strong> (Masked Language Model): masking some tokens at random and predicting them.</li>\n<li><strong>NER</strong> (Named entity recognition): predicting the type of each entity (name, verb, pronoun, date, place, and so on)</li>\n<li><strong>NSP</strong> (Next Sentence Prediction): predicting whether the second sentence follows the first one (in logical/causal way) or not. </li>\n<li><strong>GLUE</strong>: General Language Understanding Evaluation, a popular NLP benchmark</li>\n<li><strong>AdamW</strong>: Adam with weight decay, a very popular gradient descent optimizer. More about it <a href=\"https://www.fast.ai/2018/07/02/adam-weight-decay/\">here</a>. </li>\n</ul>\n\n<p>Let me know if these  terms are correct/used  (especially the first four terms related to training tasks) and if not please suggest corrections. Also, before leaving, check this great <a href=\"https://www.kaggle.com/rftexas/nlp-cheatsheet-master-nlp\"><strong>notebook</strong></a> packed with many more terms. Thanks in advance for your comments! </p>",
      "rawMarkdown": "Some of the ML terms that I have heard/read employed for describing some NLP tasks (some of these aren't specific to NLP):\n\n- **Fine-tuning** on dataset X: taking a pre-trained model and training it on a different dataset, usually a smaller and more specialized one.\n- **Fine-tuning** on task  Z: taking a pre-trained model and training it to predict outputs for task Z.\n- **Transfer learning**: using a model trained for task X to do task Y. Sometimes task X is generic whereas Y is specific (X: MLM, Y: binary label prediction).\n- **Pre-trained** model: for example [`roberta-base`](https://huggingface.co/roberta-base) is a pre-trained model\n- **Backbone/base** model: the first layer of the model that is frozen. Nowadays, it is usually a transformer variation.\n- **Tokenization**: extracting tokens from the input text. There are many different [tokenizers](https://github.com/huggingface/tokenizers). \n- **MLM** (Masked Language Model): masking some tokens at random and predicting them.\n- **NER** (Named entity recognition): predicting the type of each entity (name, verb, pronoun, date, place, and so on)\n- **NSP** (Next Sentence Prediction): predicting whether the second sentence follows the first one (in logical/causal way) or not. \n- **GLUE**: General Language Understanding Evaluation, a popular NLP benchmark\n- **AdamW**: Adam with weight decay, a very popular gradient descent optimizer. More about it [here](https://www.fast.ai/2018/07/02/adam-weight-decay/). \n\n\nLet me know if these  terms are correct/used  (especially the first four terms related to training tasks) and if not please suggest corrections. Also, before leaving, check this great [**notebook**](https://www.kaggle.com/rftexas/nlp-cheatsheet-master-nlp) packed with many more terms. Thanks in advance for your comments! ",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "892264": "Some of the ML terms that I have heard/read employed for describing some NLP tasks (some of these aren't specific to NLP):\n\n- **Fine-tuning** on dataset X: taking a pre-trained model and training it on a different dataset, usually a smaller and more specialized one.\n- **Fine-tuning** on task  Z: taking a pre-trained model and training it to predict outputs for task Z.\n- **Transfer learning**: using a model trained for task X to do task Y. Sometimes task X is generic whereas Y is specific (X: MLM, Y: binary label prediction).\n- **Pre-trained** model: for example [`roberta-base`](https://huggingface.co/roberta-base) is a pre-trained model\n- **Backbone/base** model: the first layer of the model that is frozen. Nowadays, it is usually a transformer variation.\n- **Tokenization**: extracting tokens from the input text. There are many different [tokenizers](https://github.com/huggingface/tokenizers). \n- **MLM** (Masked Language Model): masking some tokens at random and predicting them.\n- **NER** (Named entity recognition): predicting the type of each entity (name, verb, pronoun, date, place, and so on)\n- **NSP** (Next Sentence Prediction): predicting whether the second sentence follows the first one (in logical/causal way) or not. \n- **GLUE**: General Language Understanding Evaluation, a popular NLP benchmark\n- **AdamW**: Adam with weight decay, a very popular gradient descent optimizer. More about it [here](https://www.fast.ai/2018/07/02/adam-weight-decay/). \n\n\nLet me know if these  terms are correct/used  (especially the first four terms related to training tasks) and if not please suggest corrections. Also, before leaving, check this great [**notebook**](https://www.kaggle.com/rftexas/nlp-cheatsheet-master-nlp) packed with many more terms. Thanks in advance for your comments! "
  }
}