{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Toxic Comments Classification - NLP & ML Project","metadata":{}},{"cell_type":"markdown","source":"With the advent of social networks, blogs and online chats the human psychology is changing: **likes, comments and tags have a real and concrete role in our lives**, going from a \"virtual\" world to reality. The impact of **toxic comments affects the engagement of users** in conversations. Many people stop expressing themselves and give up in presence of a threat of abuse and harassment online.\n\nThis means that **platforms and users loose the ability to be involved in constructive conversations.**  Moreover, the severity and viciousness of these comments has evolved into something much more sinister such as the 2021 US capitol riots and several cases of suicide, even among very young people. Identify toxic comments can **lead communities to facilitate and improve online conversations and exchange of opinions**.","metadata":{}},{"cell_type":"markdown","source":"## Abstract\nAfter importing the data, that is raw comments, we proceed with an **initial cleaning of the text**, with the aim of removing textual elements that are not to be considered during the subsequent analysis (mentions, hashtags, repeated spaces, ...). We then proceed with some **classic NLP tasks, namely the removal of stopwords, tokenization and stemming.** The raw text is then transformed into a list of words, called tokens. **We turn the resulting words into numbers, more precisely arrays of numbers.** We will use three approaches: the first is very simple and based on word frequencies (TF-IDF), the others are more complex (Word2Vec, Doc2Vec). **The vectors resulting from these operations will be used by machine learning models** to make predictions about new comments. In particular **we will test different models**: logistic regression, naive bayes and linear svm. At the end we will choose **the combination of encoding and model that perform better**, and we will proceed with the **fine tuning, looking for the best parameters and the best threshold.** I also developed **a demo application** for this classifier, hosted on my Hugging Face. \n\nThe following image shows graphically the workflow.","metadata":{}},{"cell_type":"markdown","source":"## Steps of the Project\nThis is an end-to-end project covering all the necessary steps from text preprocessing to final classification and tuning.\n\nThe project consists of the following main sections (wich may consist of different subsections):","metadata":{}},{"cell_type":"markdown","source":"1. **Importing Tools**\n2. **Importing Data and Cleaning**\n3. **Preprocessing Text - NLP pipeline**\n4. **Some Visualization**\n5. **Splitting the data in training and testing sets**\n6. **Different Embeddings: TF-IDF, Word2Vec, Doc2Vec**\n7. **Dealing with Imbalanced Classes**\n8. **Features Normalization**\n9. **Algorithm Selection and Hyperparameter tuning**\n10. **Performance of the model**\n11. **ROC, AUC, Precision-Recall Curve**\n12. **Threshold Tuning**\n13. **Demo App & Further Steps**\n14. **Drawing conclusions — Summary**","metadata":{}},{"cell_type":"markdown","source":"## Dataset","metadata":{}},{"cell_type":"markdown","source":"The dataset for this project is hosted on Kaggle and was part of a competition. According to the description on Kaggle, *The Conversation AI team, a research initiative founded by Jigsaw and Google (both a part of Alphabet) are working on tools to help improve online conversation. One area of focus is the study of negative online behaviors, like toxic comments*.\n\n**The goal of this project is slightly different from the original one: here i will perform a binary classification of the comments** (instead of the multiclass multioutput defined in the original competition).\n\nI will use only the train set provided by Kaggle and i will proceed with the usual splitting into train and test sets. **The dataset consists of 143346 non-toxic comments and 16225 toxic comments. A comment is considered toxic if belongs to at least one of the toxicity categories provided in the original dataset.** The original toxicity labels have been assigned in a manual way by humans, according to the description of the dataset available on Kaggle.\n\nThe dataset contains the raw comments, extracted as they were originally posted by users. **This comments have been extracted from Wikipedia talks pages.** For more details about the dataset, please refer to the competition page on Kaggle, at <https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge>.","metadata":{}},{"cell_type":"markdown","source":"## 1. Importing Tools","metadata":{}},{"cell_type":"markdown","source":"For this project are required the common data science libraries. In addition i use **imblearn** to deal with imbalanced datasets, **nltk** and **spacy** for basic NLP tasks, **gensim** for advanced text embedding and **wordcloud** to provide visualizations base on the words in the corpus.","metadata":{}},{"cell_type":"code","source":"# UTILITY\nfrom PIL import Image\nfrom joblib import dump\nimport pandas as pd\nimport numpy as np\nfrom collections import Counter\nfrom sklearn.preprocessing import MinMaxScaler, Normalizer\nimport warnings\n\n# IMBALANCED LEARN\nfrom imblearn.over_sampling import SMOTE\n\n# NLP\nimport re\nfrom nltk.corpus import wordnet\nfrom nltk.tokenize import word_tokenize\nfrom nltk import SnowballStemmer\nimport spacy\nfrom gensim.models.doc2vec import Doc2Vec, TaggedDocument, Word2Vec\nfrom gensim.test.utils import get_tmpfile\nfrom sklearn.feature_extraction.text import TfidfVectorizer\n\n# MODELS\nfrom sklearn.model_selection import train_test_split, RandomizedSearchCV\nfrom sklearn.svm import LinearSVC\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.naive_bayes import MultinomialNB\n\n# EVALUATION\nfrom sklearn.metrics import confusion_matrix, accuracy_score, precision_score, recall_score,\\\n     f1_score, roc_auc_score, roc_curve, precision_recall_curve\n\n# VISUALIZATION\nfrom wordcloud import WordCloud\nimport matplotlib.pyplot as plt\n\n# INIT\nwarnings.filterwarnings('ignore')\nNUM_OF_THREADS = 10 # set > 1 to use multithreading support","metadata":{"execution":{"iopub.status.busy":"2022-06-30T11:57:20.782876Z","iopub.execute_input":"2022-06-30T11:57:20.784309Z","iopub.status.idle":"2022-06-30T11:57:20.793322Z","shell.execute_reply.started":"2022-06-30T11:57:20.784259Z","shell.execute_reply":"2022-06-30T11:57:20.79219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Importing Data and Cleaning","metadata":{}},{"cell_type":"markdown","source":"The first step of this project is importing and **cleaning textual data**. For this purpose we start with some **basic regex-based text cleaning** operations. \n\nIn particular:\n\n1. remove numbers and words concatenated with numbers (IE h4ck3r)\n2. remove emails and mentions (words with @)\n3. remove hashtags (words with #)\n4. remove URLs\n5. replace the ASCII '�' symbol with '8'\n6. remove multiple spaces","metadata":{}},{"cell_type":"code","source":"# read data from csv file\ndata = pd.read_csv(\"../input/traindata/train.csv\")\n# drop na and reset index\ndata.dropna(inplace=True)\ndata.reset_index(inplace=True, drop=True)\n# column name to work with\nname = \"comment_text\"","metadata":{"execution":{"iopub.status.busy":"2022-06-30T11:58:10.15022Z","iopub.execute_input":"2022-06-30T11:58:10.151217Z","iopub.status.idle":"2022-06-30T11:58:11.269329Z","shell.execute_reply.started":"2022-06-30T11:58:10.151178Z","shell.execute_reply":"2022-06-30T11:58:11.268208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# function to clean the text performin some simple regex pattern matching\ndef apply_regex(corpus):\n    corpus = corpus.apply(lambda x: re.sub(\"\\S*\\d\\S*\",\" \", x))          # removes numbers and words concatenated with numbers (IE h4ck3r)\n    corpus = corpus.apply(lambda x: re.sub(\"\\S*@\\S*\\s?\",\" \", x))        # removes emails and mentions (words with @)\n    corpus = corpus.apply(lambda x: re.sub(\"\\S*#\\S*\\s?\",\" \", x))        # removes hashtags (words with #)\n    corpus = corpus.apply(lambda x: re.sub(r'http\\S+', ' ', x))         # removes URLs\n    corpus = corpus.apply(lambda x: re.sub(r'[^a-zA-Z0-9 ]', ' ',x))    # keeps numbers and letters\n    corpus = corpus.apply(lambda x: x.replace(u'\\ufffd', '8'))          # replaces the ASCII '�' symbol with '8'\n    corpus = corpus.apply(lambda x: re.sub(' +', ' ', x))               # removes multiple spaces\n    return corpus","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# apply the function and clean the data\ndata[name] = apply_regex(data[name])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we said, in this project we focus **on a binary classification task**, so we create a new **binary column called *label*.** The value of this attribute is 1 if the comments belongs to at least oone of the toxicity categories.","metadata":{}},{"cell_type":"code","source":"# ASSIGN A BINARY LABEL -- 1 is toxic, 0 is not\ndata[\"label\"] = np.where(data[\"toxic\"]==1, 1, \n                np.where(data[\"severe_toxic\"]==1, 1, \n                np.where(data[\"obscene\"]==1, 1, \n                np.where(data[\"threat\"]==1, 1,\n                np.where(data[\"insult\"]==1, 1,\n                np.where(data[\"identity_hate\"]==1, 1, 0))))))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After this simple operation, we check if we are in an imbalanced situation. **The imbalanced ratio is almost 9:1**. I will deal with this problem using **SMOTE to oversample the minority class in order to balance the classes** with a 1:1 ratio. The following instruction shows the actual samples for each class. The balancing process **will be performed after the encoding** of the text into vecotrs, due to the fact that SMOTE works on numerical data.","metadata":{}},{"cell_type":"code","source":"# Count of samples for each class\nCounter(data[\"label\"])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Preprocessing Text - NLP pipeline","metadata":{}},{"cell_type":"markdown","source":"After the inital steps, we need to proceed with a **NLP pipeline**. There are some steps which are common for NLP problems like text classification, in particular we need to:\n\n1. Remove stop words and tokenize the text\n2. Stemming the tokens\n\n**Tokens are the building blocks of Natural Language**. With tokenization we break a text into smaller units, named tokens. Tokenization can be performed in different ways, for example by separating at the level of words, characters or subwords. **In this project we tokenize words.**\n\n**Stop word removal** is another common process in NLP. The goal is **removing common words that appears in most of the documents** in the corpus. **This words have no significance** since carry a lower amount of information. Generally, articles and pronouns are identified as stop words.\n\n**Stemming** is a technique used to take words at their root form (remove inflection), helping to normalize text. Same words with different inflection can be redundant in the NLP process.","metadata":{}},{"cell_type":"code","source":"# TOKENIZE TEXT - we use the Spacy library stopwords\nspacy_model = spacy.load(\"en_core_web_sm\")\nstop_words = spacy_model.Defaults.stop_words","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TOKENIZE TEXT and STOP WORDS REMOVAL - execution (removes also the words shorter than 2 and longer than 15 chars)\ndef tokenize(doc):\n    tokens_1 = word_tokenize(str(doc))\n    return [word.lower() for word in tokens_1 if len(word) > 1 and len(word) < 15 and word not in stop_words and not word.isdigit()]\n    \ndata[\"tokenized\"] = data[name].apply(tokenize)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# STEMMING\nstemmer = SnowballStemmer(language=\"english\")\n\ndef applyStemming(listOfTokens):\n    return [stemmer.stem(token) for token in listOfTokens]\n\ndata['stemmed'] = data['tokenized'].apply(applyStemming)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The final result** of the NLP preprocessing is shown in the next cell. **We can see the result of all the steps**, from raw text to stemmed text. For the next steps, **we will use the stemmed text.**","metadata":{}},{"cell_type":"code","source":"data.sample(3)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Some Visualizations","metadata":{}},{"cell_type":"markdown","source":"In this section we go throught some visualizations that will help us to better understand the data we are working with. This visualizations are **word clouds of the toxic and non-toxic raw comments.** This visualization are obtained using the *wordcloud* library. Word Clouds are visual displays of text data for simple text analysis. **Word Clouds display the most prominent or frequent words in a body of text.** ","metadata":{}},{"cell_type":"markdown","source":"First we divide toxic and non-toxic comment based on the binary label.","metadata":{}},{"cell_type":"code","source":"# TAKING TOXIC COMMENTS\nthreat_text = data[data[\"label\"]==1][\"comment_text\"]\nneg_text = pd.Series(threat_text).str.cat(sep=' ')\n\n# TAKING NON-TOXIC COMMENTS\nnon_threat_text = data[data[\"label\"]==0][\"comment_text\"]\npos_text = pd.Series(non_threat_text).str.cat(sep=' ')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At this point, we generate the wordclouds. The choice of the color palette is arbitrary, i tried to select colors that enfatize the toxicity of the words.","metadata":{}},{"cell_type":"code","source":"# Generate toxic word cloud\nwordcloud_neg = WordCloud(collocations=False, \n                            width=1500, \n                            height=600, \n                            max_font_size=150\n                        ).generate(neg_text)\nplt.figure(figsize=(15, 10))\nplt.imshow(wordcloud_neg.recolor(colormap=\"GnBu\"), interpolation='bilinear')\nplt.axis(\"off\")\nplt.title(f\"Most common words assosiated with toxic comment\", size=20)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Generate non-toxic word cloud\nwordcloud_pos = WordCloud(collocations=False, \n                            width=1500, \n                            height=600, \n                            max_font_size=150\n                        ).generate(pos_text)\nplt.figure(figsize=(15, 10))\nplt.imshow(wordcloud_pos.recolor(colormap=\"Accent\"), interpolation='bilinear')\nplt.axis(\"off\")\nplt.title(f\"Most common words assosiated with non-toxic comment\", size=20)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from this visualizations, while some words are present in both the categories, other are **really caracterizing the toxic comments.**","metadata":{}},{"cell_type":"markdown","source":"## 5. Splitting the Data in Training and Testing sets","metadata":{}},{"cell_type":"markdown","source":"The next step consist of **splitting the dataset into train and test sets** using the usual *train_test_split* functionality provided by sklearn.","metadata":{}},{"cell_type":"code","source":"# Even if for the competition on kaggle the gave also the \"test\" dataset, i decided to use only the \"train\" set as a unique dataset.\nX_train, X_test, y_train, y_test = train_test_split(data[\"stemmed\"], data[\"label\"], random_state=42)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Different Embeddings: TF-IDF, Word2Vec, Doc2Vec","metadata":{}},{"cell_type":"markdown","source":"Now that we have the test and train sets, we need to **encode the stemmed text**, that now is a vector of words, **into a Vector Space Model (VSM)** that means **converting the array of words into an array of numbers**, that will be used by a machine learning algorithm for the actual classification task.\n\nConverting text into a VSM is a **non trivial aspect, involving some domain-based reasoning.** In this project i will encode the text using **three different methods**:\n\n1. Basic TF-IDF (frequency based) approach\n2. More advanced Word2Vec model (custom trained on the corpus)\n3. More advanced Doc2Vec model (custom trained on the corpus)\n\nThis approaches will be explained in the following subsections.","metadata":{}},{"cell_type":"markdown","source":"### TF-IDF Vectorizer","metadata":{}},{"cell_type":"markdown","source":"**Term Frequency–Inverse Document Frequency** (TF-IDF) is a popular technique used to **compute words importance across all the documents** in the corpus. In particular, the assumption is that a word that appears more times in a document but doesn't appear in all the others, is important for that document. TF-IDF **assigns a weight to each word based on the occurence frequency**. At the end, words that are frequent in general will have a lower weight (carry less information for a specific document).\n\nWe represents a set of words (for example a comment) with an **array containing the scores of the words**.\n\nThis methods belongs to the \"Bag of Words\" category and share **some disadvantages** in respect to more advance methods. For instance, **word order is lost and the context is not considered**.","metadata":{}},{"cell_type":"code","source":"# commodity methos to avoid running tokenization again\ndef do_nothing(lemmas):\n    return lemmas","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The stemmed text is passed to the sklearn *TfidfVectorizer* specifying a number of features equal to 500. This is an arbitrary choice, and in order to have a fair comparison with other method, this parameter will be setted to 500 for all the three methods.","metadata":{}},{"cell_type":"code","source":"# encodin text into vectors\ntfid = TfidfVectorizer(lowercase=False, tokenizer=do_nothing, max_features=500)\n\ntrain_vectors_tfidf = tfid.fit_transform(X_train).toarray()\ntest_vectors_tfidf = tfid.transform(X_test).toarray()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Word2Vec: Model Creation and Persistence Store","metadata":{}},{"cell_type":"markdown","source":"Word2vec is a NLP technique published in 2013. This algorithm uses a **neural network to learn word representation** from a large corpus of text. The goal for the model is to **learn a semantic representation** of the words. That means similar words have a similar representation and are \"near\" in the Vector Space Model. \n\nWord2Vec represent each word as a vector, but in this project we are working with comments, that are a set of words. In order to obtain a representation of a document (a comment) we **take the average of the representations of the words that make that document (comment).** This is a very simple but still effective approach.\n\nWe use Word2Vec in order to create a **continuous bag-of-words (CBOW) representation**. With CBOW the current words is predicted from a window of contextual words. Note that the order of the context words is not taken into consideration.","metadata":{}},{"cell_type":"markdown","source":"We start by setting the parameter for the train of the model. Once ther model is trained, we save it on the disk.","metadata":{}},{"cell_type":"code","source":"# execute only if in order to retrain the model or if the trained model is not present in the folder \"./w2vec.model\"\n# seting the parameter for the w2v model\nw2v = Word2Vec(\n                sg=0,                   # 1 for skip-gram, 0 for CBOW\n                vector_size=500,        # size of the resulting vector\n                workers=NUM_OF_THREADS, # multithreading\n                epochs=100,             # number of epochs\n                seed=42                 # seed for reproducibility\n            )            \n\n# building the vocabulary and training the model\nw2v.build_vocab(X_train, progress_per=50000)\nw2v.train(X_train, total_examples=w2v.corpus_count, epochs=30, report_delay=1)\n\n# saving the model to the disk in order to avoid training again\nw2v.save(\"w2vec.model\")\nprint(\"Model Saved\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Word2Vec Embedding using the Trained Model","metadata":{}},{"cell_type":"markdown","source":"At this point, we can load the trained model and use it for the actual embedding.","metadata":{}},{"cell_type":"code","source":"wordVec = Word2Vec.load(\"./w2vec.model\").wv","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following function takes the tokens of each document (comment) and look up in the model for the vectorial representation of each token. All the representations of the tokens of each document are averaged together in order to get a vectorial representation of the whole comment. **Basically a comment is represented as the average of the words representations of the words that make that comment.**","metadata":{}},{"cell_type":"code","source":"# Creating a feature vector by averaging all embeddings for all sentences\ndef embedding_feats(corpus, w2v_model):\n    DIMENSION = 500\n    zero_vector = np.zeros(DIMENSION)\n    feats = []\n    for tokens in corpus:\n        feat_for_this =  np.zeros(DIMENSION)\n        count_for_this = 0 + 1e-5 # to avoid divide-by-zero \n        for token in tokens:\n            if token in w2v_model:\n                feat_for_this += w2v_model[token]\n                count_for_this +=1\n        if(count_for_this!=0):\n            feats.append(feat_for_this/count_for_this) \n        else:\n            feats.append(zero_vector)\n    return feats\n\n\ntrain_vectors_w2v = embedding_feats(X_train, wordVec)\ntest_vectors_w2v = embedding_feats(X_test, wordVec)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The resulting vectors will be used in the next steps.","metadata":{}},{"cell_type":"markdown","source":"### Doc2Vec: Model Creation and Persistence Store","metadata":{}},{"cell_type":"markdown","source":"**Doc2Vec model**, as opposite to Word2Vec model, is used to **create a vector representation of a group of words** considered as a single unit. With this approach we directly represent a document (comment) into a vector. Doc2Vec is **based on Word2Vec but in addition has a unit named Paragraph Vector**. \n\nDuring the trainin phase, the paragraph vector is learned together with the words. The words representation is shared across all the documets and permit to learn the paragraph vector.\n\nThe **Distributed Bag-Of-Words (DBOW)** model helps in guessing the context words from a target word. The **Distributed Memory (DM)**  model helps in guessing a single word from a context of words. In this project we use the DBOW approach.","metadata":{}},{"cell_type":"markdown","source":"We prepare the data to be trained and set up the parameters. Once the model is trained, we save it on the disk.","metadata":{}},{"cell_type":"code","source":"# execute only if in order to retrain the model or if the trained model is not present in the folder \"./d2v_comments.model\"\n#prepare training data in doc2vec format:\ntrain_doc2vec = [TaggedDocument((d), tags=[str(i)]) for i, d in enumerate(X_train)]\n\n#Train a doc2vec model to learn representations.\nmodel = Doc2Vec(\n                vector_size=500,                # size of the resulting vector\n                dm=0,                           #dm=0 -> we use a DBOW representation\n                epochs=100,                     # number of epochs\n                seed=42,                        # seed for reproducibility\n                workers=NUM_OF_THREADS,         # multithreading\n                )        \n\n# building the vocabulary and training the model\nmodel.build_vocab(train_doc2vec)\nmodel.train(train_doc2vec, total_examples=model.corpus_count, epochs=model.epochs)\n\n# saving the model to the disk in order to avoid training again\nmodel.save(\"d2v_comments.model\")\nprint(\"Model Saved\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Doc2Vec Embedding using the Trained Model","metadata":{}},{"cell_type":"markdown","source":"We load the trained model and with the *infer_vector* method we generate a vectorized representation of each document. This instruction will take some time.","metadata":{}},{"cell_type":"code","source":"#Infer the feature representation for training and test data using the trained model\nmodel_d2v= Doc2Vec.load(\"d2v_comments.model\")\n\n#infer in multiple steps to get a stable representation. \ntrain_vectors_d2v =  [model_d2v.infer_vector(tokens, epochs=70) for tokens in X_train]\ntest_vectors_d2v = [model_d2v.infer_vector(tokens, epochs=70) for tokens in X_test]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At the end of this section we have **three different encodings obtained with different methods**. The next section is about dealing with imbalanced classes.","metadata":{}},{"cell_type":"markdown","source":"## 7. Dealing with Imbalanced Classes","metadata":{}},{"cell_type":"markdown","source":"The dataset is **imbalanced with a ratio of almost 9:1.** In order to deal with this problem we use the *imbalanced learn* libraries. In particular i'm using the **SMOTE (synthetic minority oversampling technique) algorithm to produce new synthetic data points**. This new samples are generated using vector operations (this is why this operation must be done after the vector encoding phase). With the default parameters, the algorithm will **completely balance the dataset**, that will end up having a ratio of 1:1, that means same number of observation for both the classes.\n\nIn particular, SMOTE selects a minority class instance (a) at random and finds its k nearest minority class neighbors. Then, one of the k nearest neighbors is choosen at random (b). Using any distance metric, the difference in distance between the feature vector (a) and its neighbor (b) is computed. Now, this difference is multiplied by any random value in (0,1] and is added to the previous feature vector (a) in order to obtain the synthetic sample. \n\nThe operation is applied to train and test data resulting from each of the encoding methods used, in order to keep different approaches separated. I decided to use SMOTE because according to *\"V. Rupapara et al.: Impact of SMOTE on Imbalanced Text Features for Toxic Comments Classification\"* this approach is very effective in this kind of applications. After some tests i realized that this operations really seems to **boost the performance of the final classification model** for this project.","metadata":{}},{"cell_type":"code","source":"# SMOTE for the tf-idf encoding\nsm_tfidf = SMOTE(random_state=42, n_jobs=NUM_OF_THREADS)\ntrain_vectors_tfidf, y_train_tfidf = sm_tfidf.fit_resample(train_vectors_tfidf, y_train)\ntest_vectors_tfidf, y_test_tfidf = sm_tfidf.fit_resample(test_vectors_tfidf, y_test)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SMOTE for the W2V encoding\nsm_w2v = SMOTE(random_state=42, n_jobs=NUM_OF_THREADS)\ntrain_vectors_w2v, y_train_w2v = sm_w2v.fit_resample(train_vectors_w2v, y_train)\ntest_vectors_w2v, y_test_w2v = sm_w2v.fit_resample(test_vectors_w2v, y_test)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SMOTE for the D2V encoding\nsm_d2v = SMOTE(random_state=42, n_jobs=NUM_OF_THREADS)\ntrain_vectors_d2v, y_train_d2v = sm_d2v.fit_resample(train_vectors_d2v, y_train)\ntest_vectors_d2v, y_test_d2v = sm_d2v.fit_resample(test_vectors_d2v, y_test)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8. Features Normalization","metadata":{}},{"cell_type":"markdown","source":"In this section **we normalize the data** using *normalizer* from sklearn. *normalizer* **normalizes samples individually to unit norm.** Each sample is rescaled independently so that its norm (l1, l2 or inf) equals one.\nScaling is a **common operation for text classification**. We execute a normalization for each method used to encode the text. **The resulting normalized vectors will be fed as input to the classification models.** After some tests, this approach seems to improve the final performance of the model.\n\nWe normalize the vectors obtained from the different encoding techniques.","metadata":{}},{"cell_type":"markdown","source":"### TF-IDF Normalization","metadata":{}},{"cell_type":"code","source":"# NORMALIZING TF-IDF VECTORS\nnorm_TFIDF = Normalizer(copy=False)\nnorm_train_tfidf = norm_TFIDF.fit_transform(train_vectors_tfidf)\nnorm_test_tfidf = norm_TFIDF.transform(test_vectors_tfidf)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Word2Vec Normalization","metadata":{}},{"cell_type":"code","source":"# NORMALIZING WORD2VEC VECTORS\nnorm_W2V = Normalizer(copy=False)\nnorm_train_w2v = norm_W2V.fit_transform(train_vectors_w2v)\nnorm_test_w2v = norm_W2V.transform(test_vectors_w2v)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Doc2Vec Normalization","metadata":{}},{"cell_type":"code","source":"# NORMALIZING DOC2VEC VECTORS\nnorm_D2V = Normalizer(copy=False)\nnorm_train_d2v = norm_D2V.fit_transform(train_vectors_d2v)\nnorm_test_d2v = norm_D2V.transform(test_vectors_d2v)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At the end of this section, **we have a train and a test set for each encoding method used.** This vectors will be used for the classification task.","metadata":{}},{"cell_type":"markdown","source":"## 9. Algorithm Selection & Hyperparameter Tuning","metadata":{}},{"cell_type":"markdown","source":"In this section **different baseline models will be tested** for each encoding methods used.\nBasically for each encoding techniques (TF-IDF, Word2VEc, Doc2Vec) will be tested 3 models (for a total of 9 models):\n\n1. Linear SVC baseline\n2. Logistic Regression baseline\n3. Naive Bayes baseline\n\nA custom function will be used to **reason on the performance of these models** and, at the end, one model will be selected. The selected combination of model and encoding will be fine tuned.","metadata":{}},{"cell_type":"markdown","source":"We create a custom function to test programmatically different model baselines.","metadata":{}},{"cell_type":"code","source":"# generic function to create different models\ndef create_models(seed=42, classW=\"balanced\"):\n    models = []\n    # we can append more than one model to test\n    models.append(('LinearSVC', LinearSVC(random_state=seed, class_weight=classW)))\n    models.append(('Logit', LogisticRegression(random_state=seed, class_weight=classW, n_jobs=NUM_OF_THREADS, max_iter=300)))\n    models.append(('NaiveBayesMN', MultinomialNB()))\n    return models","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to evaluate performance of different techniques and models, a custom function is used. The following function **return a dataframe with some evaluation metrics**. <br><br>\nThe metrics considered are:\n* Accuracy\n* F1 score\n* Recall\n* Precision\n* Specificity\n* True Positive num\n* True Negative num\n* False Positive num\n* False Negative num\n* Size of the Test Set","metadata":{}},{"cell_type":"code","source":"# Returns a dataframe with classification metrics and confusion matrix values\ndef make_classification_score(y_test, predictions, modelName):\n    tn, fp, fn, tp = confusion_matrix(y_test, predictions).ravel()\n    prec=precision_score(y_test, predictions)\n    rec=recall_score(y_test, predictions)\n    f1=f1_score(y_test, predictions)\n    acc=accuracy_score(y_test, predictions)\n    # specificity\n    spec=tn/(tn+fp)\n\n    score = {'Model': [modelName], 'Accuracy': [acc], 'f1': [f1], 'Recall': [rec], 'Precision': [prec], \\\n        'Specificity': [spec], 'TP': [tp], 'TN': [tn], 'FP': [fp], 'FN': [fn], 'y_test size': [len(y_test)]}\n    df_score = pd.DataFrame(data=score)\n    return df_score","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Another custom function is used in order to **automatically train - test - evaluate** all the models defined.","metadata":{}},{"cell_type":"code","source":"# generic function to test the created models\ndef test_models(models, X_train, y_train, X_test, y_test):\n    score = pd.DataFrame()\n    for name, model in models:\n        # in case of nayve bayes we scale the values\n        if name==\"NaiveBayesMN\":\n            scaler = MinMaxScaler()\n            X_train = scaler.fit_transform(X_train)\n            X_test = scaler.transform(X_test)\n        # fit the model with the training data\n        model.fit(X_train, y_train)\n        # make predictions with the testing data\n        preds = model.predict(X_test)\n        # compute score\n        tmp = make_classification_score(y_test, preds, name)\n        score = pd.concat([score, tmp])\n    return score","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We proceed with the actual **creation of the models**, after this step, different encoding approaches are tested in the following sub sections.","metadata":{}},{"cell_type":"code","source":"models = create_models()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Testing the TF-IDF Approach","metadata":{}},{"cell_type":"code","source":"results = test_models(models, norm_train_tfidf, y_train_tfidf, norm_test_tfidf, y_test_tfidf)\nresults","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Testing the Word2Vec Approach","metadata":{}},{"cell_type":"code","source":"results = test_models(models, norm_train_w2v, y_train_w2v, norm_test_w2v, y_test_w2v)\nresults","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Testing the Doc2Vec Approach","metadata":{}},{"cell_type":"code","source":"results = test_models(models, norm_train_d2v, y_train_d2v, norm_test_d2v, y_test_d2v)\nresults","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### RandomizedSearch to find the Best Parameters","metadata":{}},{"cell_type":"markdown","source":"As result of the previous section, we identified a good encoding and a good model: **Logistic Regression with Doc2Vec seems better than the other candidates**. Now we perform an **hyperparameter search** in order to **find the best parameters**. For this purpose we **set up a parameter grid for the model**. The complete list of parameter for each model can be found on the documentation provided by sklearn.","metadata":{}},{"cell_type":"code","source":"#### LOGIT PARAM GRID ####\nlogit_param = {\"penalty\" : [\"l1\",\"l2\",\"elasticnet\"],\n                \"tol\" : [0.0004, 0.004, 0.04, 0.4],\n                \"C\" : [0.5, 1, 1.5, 2, 5],\n                \"random_state\" : [42],\n                \"solver\" : [\"saga\", \"sag\", \"liblinear\", \"lbfgs\", \"newton-cg\"],\n                \"max_iter\" : [200,400,600,800]\n            }","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We run the **randomized search, with a cross validation folds set to 8 and 15 iterations.** At the end we can get also the best parameters found by the search. Here another option is *GridSearchCV* that tests all the possible combinations of parameters. Even if a grid search represents a more exhaustive search, it turns out to be computationally too complex and long compared to a randomized search. Basically, in contrast to GridSearchCV, not all parameter values are tried out, but rather a fixed number of parameter settings is sampled from the specified distributions. **The number of parameter settings that are tried is given by n_iter.** \n\nThe random search approach could lead to a non-optimal combination of parameter (a sort of local optimum), so **we are not sure that we have found the real best parameters**.","metadata":{}},{"cell_type":"code","source":"# running the search and fitting the model\nfinal_model = RandomizedSearchCV(LogisticRegression(), \n                                logit_param, \n                                random_state=42, \n                                cv=8, \n                                verbose=-1, \n                                n_jobs=NUM_OF_THREADS, \n                                n_iter=15)\n                                \nfinal_model.fit(norm_train_d2v, y_train_d2v)\n# make predictions with the testing data\npreds = final_model.predict(norm_test_d2v)\n# compute score\nmake_classification_score(y_test_d2v, preds, \"logit_tuned_d2v\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# return the possibly best estimator\nfinal_model.best_estimator_","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The previous step returned the best parameters for the algorithms. Now we can use the *predict_proba()* method to **predict the probability of belonging to one class or another.** We will use this probabilities in other to tune a custom threshold.","metadata":{}},{"cell_type":"code","source":"# getting predicted probability values\nprobability = final_model.predict_proba(norm_test_d2v)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We conclude this section by **materializing the fitted model on the disc**, in order to use it in other application or avoid retraining.","metadata":{}},{"cell_type":"code","source":"dump(final_model, 'toxicCommModel.joblib')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10. Performance of the Model","metadata":{}},{"cell_type":"markdown","source":"Evaluating the performance of the model is a very important step. In order to do so, i created two custom functions. The first to **convert the probability of belonging to a class into a final binary class** using a specified threshold. The second one return a dataframe with all the **metrics for the evaluation of the model**. This functions are described in the following sections. We use the probability from the previous step to assign to each sample a final class.","metadata":{}},{"cell_type":"markdown","source":"### Custom Function to convert a probability to a class using a threshold t","metadata":{}},{"cell_type":"markdown","source":"This function takes the probability from *predict_proba()* and a threshold value. **Then convert the probability to a binary class [0,1]**. Note that in the function we **consider the prbability to belong to the \"1\" class (toxic comment).**","metadata":{}},{"cell_type":"code","source":"# from probability to class value based on custom threshold\ndef probs_to_prediction(probs, threshold):\n    pred=[]\n    for x in probs[:,1]:\n        if x>threshold:\n            pred.append(1)\n        else:\n            pred.append(0)\n    return pred","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Evaluation with a Threshold = 0.5\nBefore doing some advanced threshold tuning, it is important to **check that results with a threshold t=0.5 are the same obtained previously** using the *predict()* method. **Results are consistent because t=0.5 is the default threshold implementated by the algorithm.**","metadata":{}},{"cell_type":"code","source":"predictions=probs_to_prediction(probability, 0.5)\nmake_classification_score(y_test_d2v, predictions, \"logit, t=0.5\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The result is not bad and all the considered metrics are over .91. **We still have some False Positives and False Negatives**: this leads us to a **more advanced threshold tuning** in order to decide, based on the application, **which is the important metric to maximize** (precision or recall based on the cost of a FN and the cost of a FP)","metadata":{}},{"cell_type":"markdown","source":"## 11. ROC, AUC, Precision-Recall Curve","metadata":{}},{"cell_type":"markdown","source":"Before tuning the threshold of our classifier, in this section we plot different curves. This plots will help to decide the **\"trade off\" between all the metrics**. The threshold choice starts from this trade off.","metadata":{}},{"cell_type":"markdown","source":"### ROC AUC and Curve Plot","metadata":{}},{"cell_type":"markdown","source":"By definition an ROC curve (receiver operating characteristic curve) is a graph showing the **performance of a classification model at all the classification thresholds**. First we compute the Area Under The roc Curve (AUC) and then we plot the curve. **A value of the AUC=1 refers to a perfect classifier, AUC=0.5 refers to a random classifier.**","metadata":{}},{"cell_type":"code","source":"# calculate roc auc score\nprint(\"SVM: ROC AUC = %.4f\" % roc_auc_score(y_test_d2v, probability[:, 1]))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We got an **AUC value equal to 0.9794**, not bad. Now we **plot the curve**. The dots represent **different threshold values.**","metadata":{}},{"cell_type":"code","source":"# calculate roc curve\nmodel_fpr, model_tpr, _ = roc_curve(y_test_d2v, probability[:, 1])\n# plot roc curve\nplt.plot(model_fpr, model_tpr, marker='.', label='Logit')\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate (recall)')\nplt.legend()\nplt.title(\"ROC Curve\")\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Precision-Recall Cruve Plot","metadata":{}},{"cell_type":"markdown","source":"In a situation of unbalanced classes it is better to use a Precision - Recall curve instead of a ROC curve. Even if we balanced the dataset, caould be usefull to look at this curve anyway.<br>\nThe precision-recall curve shows the **tradeoff between precision and recall** for different threshold. **We can choose a value of the threshold that matches to de desired values of precision and recall.**","metadata":{}},{"cell_type":"code","source":"# calculate precision-recall curve\nmodel_precision, model_recall, thresholds = precision_recall_curve(y_test_d2v, probability[:, 1])\n# plot the curve\nplt.plot(model_recall, model_precision, marker='.', label='Logit')\nplt.xlabel('Recall')\nplt.ylabel('Precision')\nplt.legend()\nplt.title('Precision-Recall Curve')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 12. Threshold Tuning","metadata":{}},{"cell_type":"markdown","source":"**The choice of the important metric** to maximize (precision vs recall) **depends on the domain and the application itself**. For example, in spam detection if an important email were marked as spam when it actually isn’t, this could be disastrous as you would miss it. As opposite, in medicine a false negative would be much more damaging.\n\n**Based on the context** in which we use a toxic comments classification model, **we can decide which metric to maximize based on the cost of a false positive or false negative**, there is no a right or wrong answer here. In an application for kids (YouTube Kids) we would like to maximize the recall as a False Negative detains an higher misclassification cost. In another context we could maximize precision in order to avoid the reporting of non-toxic comments.\n\n**Let's try, for instance, to maximize precision that means cost of FP > cost of FN.** This operation requires some tests at different thresholds until we are satisfied with the number of FN and FN.","metadata":{}},{"cell_type":"code","source":"# plotting precision-recall curve with the choosen threshold\nplt.plot(model_recall, model_precision, marker='.', label='Logit')\nplt.plot(model_recall[43000], model_precision[43000], \"ro\", label=\"threshold\")\nplt.xlabel('Recall')\nplt.ylabel('Precision')\nplt.legend()\nplt.title('Precision and Recall values for choosen Threshold')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The value of the choosen threshold is:","metadata":{}},{"cell_type":"code","source":"# print the threshold value\nprint(\"Threshold value = %.4f\" % thresholds[43000])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we see, the **number of false positive with this threshold is 266 and the precision is 0.9898**. On the other hand, the recall dropped to 0.730.","metadata":{}},{"cell_type":"code","source":"# choosen threshold results\npredictions=probs_to_prediction(probability, thresholds[43000])\nmake_classification_score(y_test_d2v, predictions, \"logit, custom t\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can try also different solutions. **It is up to us to decide what is better given the actual costs** of misclassification.","metadata":{}},{"cell_type":"markdown","source":"## 13. Demo App & Further Steps","metadata":{}},{"cell_type":"markdown","source":"With the pipeline and the trained models, is possible to **develop a demo application with a GUI**. These models can also be **embedded in existing applications**.\n\nFor this purpose, i developed a demo application to play a bit with the tuned models. The application is hosted at <https://huggingface.co/spaces/EdBianchi/Social_Toximeter>.","metadata":{}},{"cell_type":"markdown","source":"The demo consist of **three sections**, the **first one allow to write a comment and display the probabilities** assigned to each class by the classifier.","metadata":{}},{"cell_type":"markdown","source":"<img src=\"../input/images/app1.png\" width=\"550\" height=\"350\">","metadata":{}},{"cell_type":"markdown","source":"**The second one shows the final classification label for the comment at three different thresholds**, 0.5 is the default thresh, 0.937 is the thresh selected after the tuning and 0.999 is a more extreme thresh. **This allow to understand how the output class changes based on the threshold** (and the importance of choiching a correct threshold).","metadata":{}},{"cell_type":"markdown","source":"<img src=\"../input/images/app2.png\" width=\"550\" height=\"200\">","metadata":{}},{"cell_type":"markdown","source":"**The last section shows the output of the NLP preprocerssing pipeline**, this are the tokens that are embedded using the trained D2V model.","metadata":{}},{"cell_type":"markdown","source":"<img src=\"../input/images/app3.png\" width=\"500\" height=\"350\">","metadata":{}},{"cell_type":"markdown","source":"## 14. Drawing Conclusions - Summary","metadata":{}},{"cell_type":"markdown","source":"We started by **cleaning the data** and performing a common **NLP pipeline**. After, we used **different embedding methods**, from a frequency based one to more advanced ones. With the resulting vectors **we performed class balancing and normalization**. At the end we **tested different baseline** models and **tuned the hyperparameters** of the best performing model. At the end we selected a context in which precision is the important metric to maximize, and **we tuned the threshold**.\n\nI am satisfied with this result of this work, and surprised about the effect of SMOTE on the performance of the model. I think that working on natural language, as well as being the basis of important everyday applications, is a stimulating and intriguing challenge: the whole society in which we live is based on language (which allows us to communicate).\n\nI tried to follow and cover **all the main steps involved in a real Data Science pipeline** applied to a NLP project and be **as precise as possible in the description** of all the passages. Although the notebook shows a linear process, **machine learning projects tend to be iterative rather than linear processes**, where previous steps are often revisited as we learn more about the problem to solve. I did lots of experiments with the encoding techniques, with normalization and class balancing in order to find this solution.\n\nI conclude keeping in my mind that \"*All models are wrong, but some are useful*\", i hope this can be a useful one.","metadata":{}},{"cell_type":"markdown","source":"Bianchi Edoardo, 2022","metadata":{}}]}