{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## [Workbook 1](https://www.kaggle.com/sabasiddiqi/workbook-1-text-pre-processing-for-beginners) - Text Preprocessing for Beginners - Data Cleaning\n<br>\n**Level** : Beginner\n\nThis notebook discusses **Text Data Preprocessing** for **NLP Problems** using Toxic Comment Classification Dataset. Data comprises of large number of Wikipedia comments which have been labeled by human raters for toxic behavior\n\n\nNext Workbook : [Workbook 2 - Text Preprocessing for Beginners - Feature Extraction](https://www.kaggle.com/sabasiddiqi/workbook-2-text-preprocessing-feature-extraction) \n\nTo skip the initial steps (reading data, text extraction from data), Jump to [Text Pre-Processing Steps](#jump).","metadata":{"_uuid":"bb56331fa1541341295222fdf292593d6ca57960"}},{"cell_type":"markdown","source":"Starting by importing required libraries.","metadata":{"_uuid":"f187dd217d98c4dcd57c7b806fe0eda25b048721"}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport re\nimport string\nfrom string import digits\nfrom nltk.corpus import stopwords\nfrom nltk.stem import WordNetLemmatizer\nfrom tqdm import tqdm","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-12T18:27:16.014064Z","iopub.execute_input":"2022-08-12T18:27:16.014404Z","iopub.status.idle":"2022-08-12T18:27:16.019529Z","shell.execute_reply.started":"2022-08-12T18:27:16.014363Z","shell.execute_reply":"2022-08-12T18:27:16.018241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reading training and test data from CSV file and saving as Pandas' Dataframe","metadata":{"_uuid":"4b81ba34af6ad302da74b215e71035409883d88d"}},{"cell_type":"code","source":"train = pd.read_csv('../input/tagsdata/separatedTagData.csv')\ntrain_data=train.drop(train.columns[2], axis=1) \ntrain_data.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:28:46.928376Z","iopub.execute_input":"2022-08-12T18:28:46.928677Z","iopub.status.idle":"2022-08-12T18:28:46.950532Z","shell.execute_reply.started":"2022-08-12T18:28:46.928585Z","shell.execute_reply":"2022-08-12T18:28:46.949423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comments=train_data.iloc[:,1]\ncomments.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:22.113579Z","iopub.execute_input":"2022-08-12T18:27:22.114047Z","iopub.status.idle":"2022-08-12T18:27:22.121430Z","shell.execute_reply.started":"2022-08-12T18:27:22.113800Z","shell.execute_reply":"2022-08-12T18:27:22.120398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels=train_data.iloc[:,0]\nlabels.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:23.064751Z","iopub.execute_input":"2022-08-12T18:27:23.065067Z","iopub.status.idle":"2022-08-12T18:27:23.071172Z","shell.execute_reply.started":"2022-08-12T18:27:23.065023Z","shell.execute_reply":"2022-08-12T18:27:23.070277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now extracting comments from train and test data, and storing their index for later use.\nMerging comments for both train and test, so that Preprocessing Steps can be performed on both at same time.","metadata":{"_uuid":"28f108017087cdcd4140318eaf26440e9b9763e2"}},{"cell_type":"markdown","source":"<br>\n###   <a id=\"jump\">Basic Text Preprocessing Steps - Cleaning </a>\n<br>\nNow that we have comments, its time to process them to convert them into a form that can be fed to classifier.\n\nTo do so following basic steps are performed and to get a better idea of what these steps do, an example is added as well. \n\n**“You are annoying!!! goJumpOff4Cliff pleaseeeeeeee”**\n* Step 1 - [Remove punctuation](#1) →** You are annoying goJumpOff4Cliff pleaseeeeeeee**\n* Step 2 - [Remove digits](#2)→ ** You are annoying goJumpOffCliff please**\n* Step 3 - [Split combined words](#3) → **You are annoying go Jump Off Cliff please**\n* Step 4 - [Convert to lowercase](#4) →   ** your are annoying go jump off cliff please**\n* Step 5 - [Split each sentence using delimiter](#5) →   ** your, are, annoying, go, jump, off, cliff, please**\n* Step 6 - [Remove stop words](#6) →       **annoying, jump, cliff **\n* Step 7 - [Convert Word to Base Form](#7) →                      **annoy, jump, cliff** \n\nPlease note that order of steps matter here, if step number 4 is performed before Step 3, we wont be able to split the Combined words like **goJumpOffCliff**.","metadata":{"_uuid":"2a9173fa5cf23077316c23f581010c7b5e43ca1b"}},{"cell_type":"markdown","source":"<a id=\"1\">Step 1 - Remove Punctuation</a>","metadata":{"_uuid":"1bccb9dcd1609c1015e5655a1d3510730d06c028"}},{"cell_type":"code","source":"c=comments.str.translate(str.maketrans(' ', ' ', string.punctuation))\nc.head()","metadata":{"_uuid":"cc43012244bf477c8aeaa66e13481030a2f0f655","execution":{"iopub.status.busy":"2022-08-12T18:27:34.295327Z","iopub.execute_input":"2022-08-12T18:27:34.295593Z","iopub.status.idle":"2022-08-12T18:27:34.303445Z","shell.execute_reply.started":"2022-08-12T18:27:34.295537Z","shell.execute_reply":"2022-08-12T18:27:34.301929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\">Step 2 - Remove Digits </a>\n\nRemoving \\n and digits","metadata":{"_uuid":"d6787905090303a072f5da8bb0d6690f542c4f92"}},{"cell_type":"code","source":"c=c.str.translate(str.maketrans(' ', ' ', '\\n'))\nc=c.str.translate(str.maketrans(' ', ' ', digits))\nc.head()","metadata":{"_uuid":"3e20a6ef51edc389a0d971f844ff6e55e5bc03e1","execution":{"iopub.status.busy":"2022-08-12T18:27:35.261896Z","iopub.execute_input":"2022-08-12T18:27:35.262360Z","iopub.status.idle":"2022-08-12T18:27:35.271689Z","shell.execute_reply.started":"2022-08-12T18:27:35.262316Z","shell.execute_reply":"2022-08-12T18:27:35.270394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"> Step 3 - Split combined words </a>\n\nFor instance, converting **whyAreYou** to **why Are You **","metadata":{"_uuid":"83046568d2d665eb948b2883b7d836ab59b33421"}},{"cell_type":"code","source":"c=c.apply(lambda tweet: re.sub(r'([a-z])([A-Z])',r'\\1 \\2',tweet))\nc.head()","metadata":{"_uuid":"fa3f07d0489c52923c0d26a0dcd245c2673c898e","execution":{"iopub.status.busy":"2022-08-12T18:27:36.082120Z","iopub.execute_input":"2022-08-12T18:27:36.082687Z","iopub.status.idle":"2022-08-12T18:27:36.090787Z","shell.execute_reply.started":"2022-08-12T18:27:36.082640Z","shell.execute_reply":"2022-08-12T18:27:36.089252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"> Step 4 - Convert to lowercase </a>\n","metadata":{"_uuid":"00569c5680aaa623750dd5adf56f74cd9d3a3ea3"}},{"cell_type":"code","source":"c=c.str.lower()\nc.head()","metadata":{"_uuid":"44476b98e5ea4e337fcd58cc013e86bece14b894","execution":{"iopub.status.busy":"2022-08-12T18:27:36.877666Z","iopub.execute_input":"2022-08-12T18:27:36.878036Z","iopub.status.idle":"2022-08-12T18:27:36.884841Z","shell.execute_reply.started":"2022-08-12T18:27:36.877998Z","shell.execute_reply":"2022-08-12T18:27:36.883888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"> Step 5 - Split each sentence using delimiter </a>\n\nConverting each sentence to list of words. We are doing it to keep necessary words in the upcoming steps and descarding the rest.","metadata":{"_uuid":"55798f058bed4976fad95983289fbd5c3190ce16"}},{"cell_type":"code","source":"c=c.str.split()\nc.head()","metadata":{"_uuid":"d33ed62a40c4f5ea689190cef26d819f944639d8","execution":{"iopub.status.busy":"2022-08-12T18:27:37.727678Z","iopub.execute_input":"2022-08-12T18:27:37.728181Z","iopub.status.idle":"2022-08-12T18:27:37.734872Z","shell.execute_reply.started":"2022-08-12T18:27:37.728147Z","shell.execute_reply":"2022-08-12T18:27:37.733891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"> Step 6 - Remove Stop Words </a>\n\nStop words are the most common words in a language and mostly filtered in NLP problems.","metadata":{"_uuid":"31402aa2f7c2c697d302861e035adc7afd29ffe0"}},{"cell_type":"code","source":"stop = set(stopwords.words('english'))\nc=c.apply(lambda x: [item for item in x if item not in stop])\nc.head()    ","metadata":{"_uuid":"d730ddbc6ba04053b45b339a4081c43fa67d3ba3","execution":{"iopub.status.busy":"2022-08-12T18:27:39.006788Z","iopub.execute_input":"2022-08-12T18:27:39.007026Z","iopub.status.idle":"2022-08-12T18:27:39.017021Z","shell.execute_reply.started":"2022-08-12T18:27:39.006987Z","shell.execute_reply":"2022-08-12T18:27:39.015694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"> Step 7 - Convert Word to Base Form or Lematize </a> \n\nConverting each word to its base form e.g. trying to try, or tried to try for simplification; using **WordNetLemmatizer** function from **NLTK** library.","metadata":{"_uuid":"cef1f657ac3b141292aa1c174c64f3c4c6417c68"}},{"cell_type":"code","source":"from tqdm import tqdm\nlemmatizer = WordNetLemmatizer()\ncom=[]\nfor y in tqdm(c):\n    new=[]\n    for x in y:\n        z=lemmatizer.lemmatize(x)\n        z=lemmatizer.lemmatize(z,'v')\n        new.append(z)\n    y=new\n    com.append(y)","metadata":{"_uuid":"702e54ae6c5f03d284cc837f51b24401a4e0180a","execution":{"iopub.status.busy":"2022-08-12T18:27:40.593739Z","iopub.execute_input":"2022-08-12T18:27:40.594242Z","iopub.status.idle":"2022-08-12T18:27:40.606389Z","shell.execute_reply.started":"2022-08-12T18:27:40.594199Z","shell.execute_reply":"2022-08-12T18:27:40.605443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Data obtained after Lemmatization is in array form, and is converted to Dataframe in the next step.","metadata":{"_uuid":"48813f5e0a8bc5df6c43f10f2e49ede0c5036718"}},{"cell_type":"code","source":"clean_data=pd.DataFrame(np.array(com), index=comments.index,columns={'comment_text'})\nclean_data['comment_text']=clean_data['comment_text'].str.join(\" \")\nprint(clean_data.head())","metadata":{"_uuid":"44a940c65a571bcbf4d7e4fc474ca9558190e66f","execution":{"iopub.status.busy":"2022-08-12T18:27:42.166577Z","iopub.execute_input":"2022-08-12T18:27:42.167025Z","iopub.status.idle":"2022-08-12T18:27:42.177423Z","shell.execute_reply.started":"2022-08-12T18:27:42.166994Z","shell.execute_reply":"2022-08-12T18:27:42.176191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Separating Train and Test Comments using the index stored earlier.","metadata":{"_uuid":"f07a2152a0da7e0e23f52993a78b2bd29ce76ff2"}},{"cell_type":"code","source":"print(\"PreProcessed Train Data : \",clean_data.head(5))","metadata":{"_uuid":"5cff6c73cf88b0f335477fbd6f1f421672e6b73d","execution":{"iopub.status.busy":"2022-08-12T18:27:43.707754Z","iopub.execute_input":"2022-08-12T18:27:43.708195Z","iopub.status.idle":"2022-08-12T18:27:43.719003Z","shell.execute_reply.started":"2022-08-12T18:27:43.708139Z","shell.execute_reply":"2022-08-12T18:27:43.717831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Merging comments and labels for training data set and ids for test data set.","metadata":{"_uuid":"4e87b3ce4ff75c0428cf6c02660fb1ab4a150070"}},{"cell_type":"code","source":"frames=[clean_data,labels]\ntrain_result = pd.concat(frames,axis=1)\nprint(train_result.head())","metadata":{"_uuid":"f370ed62a2c4b40c0528572d74fb7824386eb32a","execution":{"iopub.status.busy":"2022-08-12T18:27:44.838547Z","iopub.execute_input":"2022-08-12T18:27:44.839002Z","iopub.status.idle":"2022-08-12T18:27:44.851596Z","shell.execute_reply.started":"2022-08-12T18:27:44.838957Z","shell.execute_reply":"2022-08-12T18:27:44.850591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Saving data in csv format to use it in different notebook, or you can continue working in the same notebook.","metadata":{"_uuid":"77daf9455c751f8ef69e3de525fabce6d77c396c"}},{"cell_type":"code","source":"train_result.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:46.136186Z","iopub.execute_input":"2022-08-12T18:27:46.136640Z","iopub.status.idle":"2022-08-12T18:27:46.148841Z","shell.execute_reply.started":"2022-08-12T18:27:46.136586Z","shell.execute_reply":"2022-08-12T18:27:46.147687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.feature_extraction.text import CountVectorizer\nfrom sklearn.feature_extraction.text import TfidfTransformer","metadata":{"_uuid":"d3849add3d70a89ecb6bc9d2c8b67fe2af9ff913","execution":{"iopub.status.busy":"2022-08-12T18:27:46.757999Z","iopub.execute_input":"2022-08-12T18:27:46.758383Z","iopub.status.idle":"2022-08-12T18:27:46.761781Z","shell.execute_reply.started":"2022-08-12T18:27:46.758351Z","shell.execute_reply":"2022-08-12T18:27:46.760974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Empty Comment Cells In Train: \",clean_data['comment_text'].isna().sum())\ntrain_data = train_result","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:47.522445Z","iopub.execute_input":"2022-08-12T18:27:47.523227Z","iopub.status.idle":"2022-08-12T18:27:47.529318Z","shell.execute_reply.started":"2022-08-12T18:27:47.523153Z","shell.execute_reply":"2022-08-12T18:27:47.528402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train= train_data\ntrain_comments=train.iloc[:,0]\ntrain_labels=train.iloc[:,1]\nprint(\"Train Comments Shape : \",train_comments.shape)\nprint(\"Train Labels Shape :\",train_labels.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:48.277290Z","iopub.execute_input":"2022-08-12T18:27:48.278498Z","iopub.status.idle":"2022-08-12T18:27:48.284487Z","shell.execute_reply.started":"2022-08-12T18:27:48.278450Z","shell.execute_reply":"2022-08-12T18:27:48.283847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vectorizer = CountVectorizer(analyzer = 'word',stop_words='english',max_features=10000)\ntrain_comments_count=vectorizer.fit(train_comments).transform(train_comments)\nprint(\"Term Frequency Matrix(TF): \\n\",train_comments_count.toarray())\nprint(\"Verifying that TF is not empty by checking the sum \",train_comments_count.toarray().sum() )","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:48.847621Z","iopub.execute_input":"2022-08-12T18:27:48.848078Z","iopub.status.idle":"2022-08-12T18:27:48.856429Z","shell.execute_reply.started":"2022-08-12T18:27:48.848032Z","shell.execute_reply":"2022-08-12T18:27:48.855821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf_transformer = TfidfTransformer()\ntf_transformer.fit(train_comments_count)\ntrain_tfidf = tf_transformer.transform(train_comments_count)\nprint(\"Train TF-IDF Matrix Shape: \",train_tfidf.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:50.122871Z","iopub.execute_input":"2022-08-12T18:27:50.123233Z","iopub.status.idle":"2022-08-12T18:27:50.130790Z","shell.execute_reply.started":"2022-08-12T18:27:50.123204Z","shell.execute_reply":"2022-08-12T18:27:50.129842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.naive_bayes import MultinomialNB\nfrom tqdm import tqdm\nclf = MultinomialNB()\nclf.get_params()\nclf.get_params().keys()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:51.616321Z","iopub.execute_input":"2022-08-12T18:27:51.616570Z","iopub.status.idle":"2022-08-12T18:27:51.623682Z","shell.execute_reply.started":"2022-08-12T18:27:51.616526Z","shell.execute_reply":"2022-08-12T18:27:51.622380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport scipy.sparse\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.metrics import make_scorer\nfrom sklearn.model_selection import GridSearchCV\nimport matplotlib.pyplot as plt\n\nX=train_tfidf\ny = train_labels\nclf_class = MultinomialNB(alpha=0.01)\nclf_class.fit(X,y)\nclasses=clf_class.classes_","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:27:53.720400Z","iopub.execute_input":"2022-08-12T18:27:53.720820Z","iopub.status.idle":"2022-08-12T18:27:53.728425Z","shell.execute_reply.started":"2022-08-12T18:27:53.720777Z","shell.execute_reply":"2022-08-12T18:27:53.727402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Add Test Data Here (please add minimum two inputs to keep it as a series","metadata":{}},{"cell_type":"code","source":"data = ['user black listing',\n        'lacks information',\n       'headphones are not used feedback']","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:29:40.844169Z","iopub.execute_input":"2022-08-12T18:29:40.844532Z","iopub.status.idle":"2022-08-12T18:29:40.848803Z","shell.execute_reply.started":"2022-08-12T18:29:40.844442Z","shell.execute_reply":"2022-08-12T18:29:40.847431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comments=pd.DataFrame(data,columns = ['labels'])\ncomments = comments.squeeze()\nprint(comments)\nc=comments.str.translate(str.maketrans(' ', ' ', string.punctuation))\nc=c.str.translate(str.maketrans(' ', ' ', '\\n'))\nc=c.str.translate(str.maketrans(' ', ' ', digits))\nc=c.apply(lambda tweet: re.sub(r'([a-z])([A-Z])',r'\\1 \\2',tweet))\nc=c.str.lower()\nc=c.str.split()\nstop = set(stopwords.words('english'))\nc=c.apply(lambda x: [item for item in x if item not in stop])\nlemmatizer = WordNetLemmatizer()\ncom=[]\nfor y in tqdm(c):\n    new=[]\n    for x in y:\n        z=lemmatizer.lemmatize(x)\n        z=lemmatizer.lemmatize(z,'v')\n        new.append(z)\n    y=new\n    com.append(y)\nclean_data=pd.DataFrame(np.array(com), index=comments.index,columns={'comment_text'})\nclean_data['comment_text']=clean_data['comment_text'].str.join(\" \")","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:29:41.966559Z","iopub.execute_input":"2022-08-12T18:29:41.966826Z","iopub.status.idle":"2022-08-12T18:29:41.984588Z","shell.execute_reply.started":"2022-08-12T18:29:41.966792Z","shell.execute_reply":"2022-08-12T18:29:41.983342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_comments_count = vectorizer.transform(clean_data['comment_text'])\ntest_tfidf = tf_transformer.transform(test_comments_count)\npred_conf_clf = clf_class.predict_proba(test_tfidf)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:29:42.185446Z","iopub.execute_input":"2022-08-12T18:29:42.185670Z","iopub.status.idle":"2022-08-12T18:29:42.191520Z","shell.execute_reply.started":"2022-08-12T18:29:42.185641Z","shell.execute_reply":"2022-08-12T18:29:42.190353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for p_index, p_val in enumerate(pred_conf_clf):\n    print(\"Text - \", clean_data['comment_text'][p_index])\n    for index, val in enumerate(p_val):\n        print(classes[index], round(val*100, 2) ,\"%\")\n    print(\"Max Probability -> \", classes[np.argmax(p_val)],round(p_val.max()*100, 2) ,\"%\\n\" )\n   ","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:29:42.403035Z","iopub.execute_input":"2022-08-12T18:29:42.403302Z","iopub.status.idle":"2022-08-12T18:29:42.417123Z","shell.execute_reply.started":"2022-08-12T18:29:42.403241Z","shell.execute_reply":"2022-08-12T18:29:42.414632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}