{
  "id": 148709,
  "title": "Review Your NLP Knowledge",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/148709",
  "author_name": "",
  "post_date": "2020-05-05T11:07:12.406000",
  "votes": 0,
  "comment_count": 12,
  "views": 0,
  "content": "<p><strong>1. Abbreviated Words in NLP:</strong></p>\n\n<ul>\n<li>CNN: Convolutional neural network</li>\n<li>RNN: Recurrent neural network</li>\n<li>LSTM: Long Short Term Memory</li>\n<li>BERT: Bidirectional Encoder Representations from Transformers.</li>\n<li>POS: parts of speech.</li>\n<li>DTM: Document Term Matrix.</li>\n<li>NER: name entity recognition.</li>\n<li>NLG: Natural Language Generation.</li>\n<li>NLU: Natural Language Understanding.</li>\n<li>TF IDF: Term Frequency–Inverse Document Frequency.</li>\n<li>re: Regular expression.</li>\n<li>LDA: Latent Dirichlet Allocation.</li>\n<li>LSI: Latent Semantic Indexing.</li>\n<li>NMF: Non-Negative Matrix Factorization.</li>\n<li>NLTK: Natural Language Toolkit</li>\n</ul>\n\n<p><strong>2. Some Common Steps for NLP Problems:</strong></p>\n\n<ul>\n<li>Sentence Segmentation: break the text apart into separate sentences</li>\n<li>Tokenization: split Sentence to words</li>\n<li>Stemming: process of reducing words to their word stem for example thinking→ think</li>\n<li>Lemmatizing: for example worse→ bad</li>\n<li>POS tags: Predicting Parts of Speech for Each Token</li>\n<li>Identifying Stop Words: like “and”, “the”</li>\n<li>Name entity recognition: detect nouns with the real world concepts.</li>\n<li>Text classification</li>\n<li>Chunking</li>\n<li>Coreference resolution\n<a href=\"https://medium.com/@ageitgey/natural-language-processing-is-fun-9a0bff37854e\">Read more…</a></li>\n</ul>\n\n<p><strong>3. Applications of NLP in The Real World:</strong></p>\n\n<ul>\n<li>Multilingual Toxic Comment Classification 👍 </li>\n<li>Personal assistant applications</li>\n<li>Fighting spam</li>\n<li>Chatbots</li>\n<li>Managing the Advertisement</li>\n<li>Sentiment analysis</li>\n<li>Text classification</li>\n<li>Text summarization</li>\n<li>Toxicity Classification</li>\n<li>Name entity recognization</li>\n<li>Part of speech tagging</li>\n<li>Language model building</li>\n<li>Machine translation</li>\n<li>Spell checking</li>\n<li>Speech recognition</li>\n<li>Character recognition\n<a href=\"https://www.wonderflow.co/nlp-examples/\">Read more…</a></li>\n</ul>\n\n<p><strong>4. Python Library for NLP:</strong>\n- NLTK\n- spaCy\n- Huggingface\n- Gensim : is a python library specifically for Topic Modelling.\n- Pattern\n- Stanford CoreNLP\n- Polyglot\n- TextBlob\n- re: python library for regular expression\n- WordCloud\n- allennlp: an open-source NLP research library, built on PyTorch\n<a href=\"https://kleiber.me/blog/2018/02/25/top-10-python-nlp-libraries-2018/\">Read more…</a></p>\n\n<p><strong>5. A few terms in NLP:</strong></p>\n\n<ul>\n<li>Stop words</li>\n<li>Punctuation</li>\n<li>Word embedding</li>\n<li>Word segmentation</li>\n<li>Text summarization</li>\n<li>Regular expression</li>\n<li>Morphological segmentation</li>\n<li>Named entity recognition</li>\n<li>Corpus: A collection of texts</li>\n<li>Document-Term Matrix</li>\n<li>n-gram: tokenize sentences by n words combination</li>\n<li>LDA (Latent Dirichlet Allocation): a technique for topic modelling.\n<a href=\"https://www.kdnuggets.com/2017/02/natural-language-processing-key-terms-explained.html\">Read more…</a></li>\n</ul>\n\n<p>This is not my compilation but was very useful to me! Thanks to mjbahmani for the great work +1</p>",
  "messages": [
    {
      "id": 853873,
      "postDate": "2020-05-19T14:58:39.983Z",
      "content": "<p>Fantastic summary! Thank you so much! 👍 </p>",
      "rawMarkdown": "Fantastic summary! Thank you so much! 👍 ",
      "votes": 1
    },
    {
      "id": 850980,
      "postDate": "2020-05-17T08:17:31.873Z",
      "content": "<p>Amazing work <a href=\"/moradnejad\">@moradnejad</a> . This is a great summary for beginners like me in the field of NLP. It will be helpful to all beginners as a references to knowledge.\nThanks a lot.</p>",
      "rawMarkdown": "Amazing work @moradnejad . This is a great summary for beginners like me in the field of NLP. It will be helpful to all beginners as a references to knowledge.\nThanks a lot.",
      "votes": 1
    },
    {
      "id": 850414,
      "postDate": "2020-05-16T15:48:01.323Z",
      "content": "<p><a href=\"/moradnejad\">@moradnejad</a> this is a great summary, thanks for the brief. </p>",
      "rawMarkdown": "@moradnejad this is a great summary, thanks for the brief. ",
      "votes": 1
    },
    {
      "id": 844604,
      "postDate": "2020-05-12T17:42:09.470Z",
      "content": "<p>Thanks man. This is great just like your deep learning cheat sheet. awesome 👍 </p>",
      "rawMarkdown": "Thanks man. This is great just like your deep learning cheat sheet. awesome 👍 ",
      "votes": 1
    },
    {
      "id": 834803,
      "postDate": "2020-05-05T19:33:12.303Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": 1,
      "replies": [
        {
          "id": 842845,
          "postDate": "2020-05-11T16:55:20.690Z",
          "content": "<p>Glad to be helpful :)</p>",
          "rawMarkdown": "Glad to be helpful :)"
        }
      ]
    },
    {
      "id": 834255,
      "postDate": "2020-05-05T12:32:35.090Z",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": 1,
      "replies": [
        {
          "id": 834284,
          "postDate": "2020-05-05T13:05:41.247Z",
          "content": "<p>You're welcome 👍 </p>",
          "rawMarkdown": "You're welcome 👍 "
        }
      ]
    },
    {
      "id": 834167,
      "postDate": "2020-05-05T11:07:12.407Z",
      "content": "<p><strong>1. Abbreviated Words in NLP:</strong></p>\n\n<ul>\n<li>CNN: Convolutional neural network</li>\n<li>RNN: Recurrent neural network</li>\n<li>LSTM: Long Short Term Memory</li>\n<li>BERT: Bidirectional Encoder Representations from Transformers.</li>\n<li>POS: parts of speech.</li>\n<li>DTM: Document Term Matrix.</li>\n<li>NER: name entity recognition.</li>\n<li>NLG: Natural Language Generation.</li>\n<li>NLU: Natural Language Understanding.</li>\n<li>TF IDF: Term Frequency–Inverse Document Frequency.</li>\n<li>re: Regular expression.</li>\n<li>LDA: Latent Dirichlet Allocation.</li>\n<li>LSI: Latent Semantic Indexing.</li>\n<li>NMF: Non-Negative Matrix Factorization.</li>\n<li>NLTK: Natural Language Toolkit</li>\n</ul>\n\n<p><strong>2. Some Common Steps for NLP Problems:</strong></p>\n\n<ul>\n<li>Sentence Segmentation: break the text apart into separate sentences</li>\n<li>Tokenization: split Sentence to words</li>\n<li>Stemming: process of reducing words to their word stem for example thinking→ think</li>\n<li>Lemmatizing: for example worse→ bad</li>\n<li>POS tags: Predicting Parts of Speech for Each Token</li>\n<li>Identifying Stop Words: like “and”, “the”</li>\n<li>Name entity recognition: detect nouns with the real world concepts.</li>\n<li>Text classification</li>\n<li>Chunking</li>\n<li>Coreference resolution\n<a href=\"https://medium.com/@ageitgey/natural-language-processing-is-fun-9a0bff37854e\">Read more…</a></li>\n</ul>\n\n<p><strong>3. Applications of NLP in The Real World:</strong></p>\n\n<ul>\n<li>Multilingual Toxic Comment Classification 👍 </li>\n<li>Personal assistant applications</li>\n<li>Fighting spam</li>\n<li>Chatbots</li>\n<li>Managing the Advertisement</li>\n<li>Sentiment analysis</li>\n<li>Text classification</li>\n<li>Text summarization</li>\n<li>Toxicity Classification</li>\n<li>Name entity recognization</li>\n<li>Part of speech tagging</li>\n<li>Language model building</li>\n<li>Machine translation</li>\n<li>Spell checking</li>\n<li>Speech recognition</li>\n<li>Character recognition\n<a href=\"https://www.wonderflow.co/nlp-examples/\">Read more…</a></li>\n</ul>\n\n<p><strong>4. Python Library for NLP:</strong>\n- NLTK\n- spaCy\n- Huggingface\n- Gensim : is a python library specifically for Topic Modelling.\n- Pattern\n- Stanford CoreNLP\n- Polyglot\n- TextBlob\n- re: python library for regular expression\n- WordCloud\n- allennlp: an open-source NLP research library, built on PyTorch\n<a href=\"https://kleiber.me/blog/2018/02/25/top-10-python-nlp-libraries-2018/\">Read more…</a></p>\n\n<p><strong>5. A few terms in NLP:</strong></p>\n\n<ul>\n<li>Stop words</li>\n<li>Punctuation</li>\n<li>Word embedding</li>\n<li>Word segmentation</li>\n<li>Text summarization</li>\n<li>Regular expression</li>\n<li>Morphological segmentation</li>\n<li>Named entity recognition</li>\n<li>Corpus: A collection of texts</li>\n<li>Document-Term Matrix</li>\n<li>n-gram: tokenize sentences by n words combination</li>\n<li>LDA (Latent Dirichlet Allocation): a technique for topic modelling.\n<a href=\"https://www.kdnuggets.com/2017/02/natural-language-processing-key-terms-explained.html\">Read more…</a></li>\n</ul>\n\n<p>This is not my compilation but was very useful to me! Thanks to mjbahmani for the great work +1</p>",
      "rawMarkdown": "**1. Abbreviated Words in NLP:**\n\n- CNN: Convolutional neural network\n- RNN: Recurrent neural network\n- LSTM: Long Short Term Memory\n- BERT: Bidirectional Encoder Representations from Transformers.\n- POS: parts of speech.\n- DTM: Document Term Matrix.\n- NER: name entity recognition.\n- NLG: Natural Language Generation.\n- NLU: Natural Language Understanding.\n- TF IDF: Term Frequency–Inverse Document Frequency.\n- re: Regular expression.\n- LDA: Latent Dirichlet Allocation.\n- LSI: Latent Semantic Indexing.\n- NMF: Non-Negative Matrix Factorization.\n- NLTK: Natural Language Toolkit\n\n**2. Some Common Steps for NLP Problems:**\n\n- Sentence Segmentation: break the text apart into separate sentences\n- Tokenization: split Sentence to words\n- Stemming: process of reducing words to their word stem for example thinking→ think\n- Lemmatizing: for example worse→ bad\n- POS tags: Predicting Parts of Speech for Each Token\n- Identifying Stop Words: like “and”, “the”\n- Name entity recognition: detect nouns with the real world concepts.\n- Text classification\n- Chunking\n- Coreference resolution\n [Read more…](https://medium.com/@ageitgey/natural-language-processing-is-fun-9a0bff37854e)\n\n**3. Applications of NLP in The Real World:**\n\n- Multilingual Toxic Comment Classification 👍 \n- Personal assistant applications\n- Fighting spam\n- Chatbots\n- Managing the Advertisement\n- Sentiment analysis\n- Text classification\n- Text summarization\n- Toxicity Classification\n- Name entity recognization\n- Part of speech tagging\n- Language model building\n- Machine translation\n- Spell checking\n- Speech recognition\n- Character recognition\n[Read more…](https://www.wonderflow.co/nlp-examples/)\n\n**4. Python Library for NLP:**\n- NLTK\n- spaCy\n- Huggingface\n- Gensim : is a python library specifically for Topic Modelling.\n- Pattern\n- Stanford CoreNLP\n- Polyglot\n- TextBlob\n- re: python library for regular expression\n- WordCloud\n- allennlp: an open-source NLP research library, built on PyTorch\n[Read more…](https://kleiber.me/blog/2018/02/25/top-10-python-nlp-libraries-2018/)\n\n**5. A few terms in NLP:**\n\n- Stop words\n- Punctuation\n- Word embedding\n- Word segmentation\n- Text summarization\n- Regular expression\n- Morphological segmentation\n- Named entity recognition\n- Corpus: A collection of texts\n- Document-Term Matrix\n- n-gram: tokenize sentences by n words combination\n- LDA (Latent Dirichlet Allocation): a technique for topic modelling.\n[Read more…](https://www.kdnuggets.com/2017/02/natural-language-processing-key-terms-explained.html)\n\nThis is not my compilation but was very useful to me! Thanks to mjbahmani for the great work +1"
    },
    {
      "id": 854468,
      "postDate": "2020-05-20T03:37:50.843Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 849705,
      "postDate": "2020-05-16T02:16:24.167Z",
      "content": "<p>Thanks !</p>",
      "rawMarkdown": "Thanks !",
      "votes": 1
    },
    {
      "id": 849330,
      "postDate": "2020-05-15T17:02:22.293Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 848473,
      "postDate": "2020-05-15T02:12:51.493Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 853873,
      "author_name": "Zoe H",
      "author_url": "",
      "post_date": "2020-05-19T14:58:39.983000",
      "content": "<p>Fantastic summary! Thank you so much! 👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 850980,
      "author_name": "roblex nana",
      "author_url": "",
      "post_date": "2020-05-17T08:17:31.873000",
      "content": "<p>Amazing work <a href=\"/moradnejad\">@moradnejad</a> . This is a great summary for beginners like me in the field of NLP. It will be helpful to all beginners as a references to knowledge.\nThanks a lot.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 850414,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-05-16T15:48:01.323000",
      "content": "<p><a href=\"/moradnejad\">@moradnejad</a> this is a great summary, thanks for the brief. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 844604,
      "author_name": "aamirhussain",
      "author_url": "",
      "post_date": "2020-05-12T17:42:09.470000",
      "content": "<p>Thanks man. This is great just like your deep learning cheat sheet. awesome 👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 834803,
      "author_name": "Marc Serra",
      "author_url": "",
      "post_date": "2020-05-05T19:33:12.303000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 842845,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-11T16:55:20.690000",
          "content": "<p>Glad to be helpful :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 834255,
      "author_name": "Alberto Maria Falletta",
      "author_url": "",
      "post_date": "2020-05-05T12:32:35.090000",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 834284,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-05T13:05:41.247000",
          "content": "<p>You're welcome 👍 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 854468,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-20T03:37:50.843000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 849705,
      "author_name": "Mohammad Essam",
      "author_url": "",
      "post_date": "2020-05-16T02:16:24.167000",
      "content": "<p>Thanks !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 849330,
      "author_name": "Aditya Singh",
      "author_url": "",
      "post_date": "2020-05-15T17:02:22.293000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 848473,
      "author_name": "shivan kumar",
      "author_url": "",
      "post_date": "2020-05-15T02:12:51.493000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "853873": "Fantastic summary! Thank you so much! 👍 ",
    "850980": "Amazing work @moradnejad . This is a great summary for beginners like me in the field of NLP. It will be helpful to all beginners as a references to knowledge.\nThanks a lot.",
    "850414": "@moradnejad this is a great summary, thanks for the brief. ",
    "844604": "Thanks man. This is great just like your deep learning cheat sheet. awesome 👍 ",
    "834803": "Thank you for sharing!",
    "834255": "Thanks for sharing :)",
    "834167": "**1. Abbreviated Words in NLP:**\n\n- CNN: Convolutional neural network\n- RNN: Recurrent neural network\n- LSTM: Long Short Term Memory\n- BERT: Bidirectional Encoder Representations from Transformers.\n- POS: parts of speech.\n- DTM: Document Term Matrix.\n- NER: name entity recognition.\n- NLG: Natural Language Generation.\n- NLU: Natural Language Understanding.\n- TF IDF: Term Frequency–Inverse Document Frequency.\n- re: Regular expression.\n- LDA: Latent Dirichlet Allocation.\n- LSI: Latent Semantic Indexing.\n- NMF: Non-Negative Matrix Factorization.\n- NLTK: Natural Language Toolkit\n\n**2. Some Common Steps for NLP Problems:**\n\n- Sentence Segmentation: break the text apart into separate sentences\n- Tokenization: split Sentence to words\n- Stemming: process of reducing words to their word stem for example thinking→ think\n- Lemmatizing: for example worse→ bad\n- POS tags: Predicting Parts of Speech for Each Token\n- Identifying Stop Words: like “and”, “the”\n- Name entity recognition: detect nouns with the real world concepts.\n- Text classification\n- Chunking\n- Coreference resolution\n [Read more…](https://medium.com/@ageitgey/natural-language-processing-is-fun-9a0bff37854e)\n\n**3. Applications of NLP in The Real World:**\n\n- Multilingual Toxic Comment Classification 👍 \n- Personal assistant applications\n- Fighting spam\n- Chatbots\n- Managing the Advertisement\n- Sentiment analysis\n- Text classification\n- Text summarization\n- Toxicity Classification\n- Name entity recognization\n- Part of speech tagging\n- Language model building\n- Machine translation\n- Spell checking\n- Speech recognition\n- Character recognition\n[Read more…](https://www.wonderflow.co/nlp-examples/)\n\n**4. Python Library for NLP:**\n- NLTK\n- spaCy\n- Huggingface\n- Gensim : is a python library specifically for Topic Modelling.\n- Pattern\n- Stanford CoreNLP\n- Polyglot\n- TextBlob\n- re: python library for regular expression\n- WordCloud\n- allennlp: an open-source NLP research library, built on PyTorch\n[Read more…](https://kleiber.me/blog/2018/02/25/top-10-python-nlp-libraries-2018/)\n\n**5. A few terms in NLP:**\n\n- Stop words\n- Punctuation\n- Word embedding\n- Word segmentation\n- Text summarization\n- Regular expression\n- Morphological segmentation\n- Named entity recognition\n- Corpus: A collection of texts\n- Document-Term Matrix\n- n-gram: tokenize sentences by n words combination\n- LDA (Latent Dirichlet Allocation): a technique for topic modelling.\n[Read more…](https://www.kdnuggets.com/2017/02/natural-language-processing-key-terms-explained.html)\n\nThis is not my compilation but was very useful to me! Thanks to mjbahmani for the great work +1",
    "854468": "",
    "849705": "Thanks !",
    "849330": "Thanks for sharing !",
    "848473": "Thank you for sharing!"
  }
}