{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction**\nIn this notebook,I use some simple techniques to classify text. There are many models/techniques that is used to text classify but here I use simple approach so that one can understand easily how to preprocess data and classify text.\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-20T20:25:12.471410Z","iopub.execute_input":"2023-02-20T20:25:12.472148Z","iopub.status.idle":"2023-02-20T20:25:27.544789Z","shell.execute_reply.started":"2023-02-20T20:25:12.472062Z","shell.execute_reply":"2023-02-20T20:25:27.543506Z"}}},{"cell_type":"markdown","source":"## **Table of Contents**\n* Import tools/libraries\n* Load data set\n* Target visualization\n* Text Cleaning\n* Remove Stopward\n* Stemming\n* TF-IDF model\n* Split data set\n* Logistic regression Model\n* NB model\n* Stochastic Gradient Descent","metadata":{}},{"cell_type":"markdown","source":"### **All the necessary tools/libraries**\nAll the tools that help us to complete this task.","metadata":{}},{"cell_type":"code","source":"import re\nimport string\nimport numpy as np \nimport random\nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\nfrom wordcloud import WordCloud, STOPWORDS, ImageColorGenerator\nimport nltk\nfrom nltk.corpus import stopwords\nfrom nltk.tokenize import word_tokenize\nimport os\nimport spacy\nimport random\nfrom spacy.util import compounding\nfrom spacy.util import minibatch\nfrom collections import defaultdict\nfrom collections import Counter\nimport keras\nfrom keras.models import Sequential\nfrom keras.initializers import Constant\nfrom keras.layers import (LSTM, \n                          Embedding, \n                          BatchNormalization,\n                          Dense, \n                          TimeDistributed, \n                          Dropout, \n                          Bidirectional,\n                          Flatten, \n                          GlobalMaxPool1D)\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers.embeddings import Embedding\nfrom keras.callbacks import ModelCheckpoint, ReduceLROnPlateau\n\nfrom sklearn.metrics import (\n    precision_score, \n    recall_score, \n    f1_score, \n    classification_report,\n    accuracy_score\n)\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-21T08:23:46.346858Z","iopub.execute_input":"2023-02-21T08:23:46.347175Z","iopub.status.idle":"2023-02-21T08:23:58.826000Z","shell.execute_reply.started":"2023-02-21T08:23:46.347105Z","shell.execute_reply":"2023-02-21T08:23:58.823675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Load data set**\nThis data set contain 6 columns but for text classification we need just two columns namely 1st columns and last columns. The columns name for this data set is not clear, so to fixed it we change the columns name.","metadata":{}},{"cell_type":"code","source":"columns  = [\"sentiment\", \"ids\", \"date\", \"flag\", \"user\", \"text\"]\ndata = pd.read_csv(\"/kaggle/input/sentiment140/training.1600000.processed.noemoticon.csv\", encoding = \"ISO-8859-1\", names = columns)\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:39:07.980864Z","iopub.execute_input":"2023-02-21T07:39:07.981478Z","iopub.status.idle":"2023-02-21T07:39:15.004070Z","shell.execute_reply.started":"2023-02-21T07:39:07.981445Z","shell.execute_reply":"2023-02-21T07:39:15.003107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we work on only two columns namely sentiment and text so we can separate thoes two columns.","metadata":{}},{"cell_type":"code","source":"data = data[['sentiment','text']]\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:39:15.008743Z","iopub.execute_input":"2023-02-21T07:39:15.010980Z","iopub.status.idle":"2023-02-21T07:39:15.069650Z","shell.execute_reply.started":"2023-02-21T07:39:15.010942Z","shell.execute_reply":"2023-02-21T07:39:15.068691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Target visualization**\nour target is sentiment, and we have to predict when it is positive or negative. So, as our data set contain 0 and 4 so we need to replace thoes value as negative and possitive, and later we will encode thoes values as 0 and 1.","metadata":{}},{"cell_type":"code","source":"class_dict = {0:'negative', 4:'positive'}\ndata['sentiment'] = data['sentiment'].apply(lambda x:  class_dict[x])\ncount = data['sentiment'].value_counts()\ncount.plot(kind='bar')","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:39:18.818630Z","iopub.execute_input":"2023-02-21T07:39:18.818997Z","iopub.status.idle":"2023-02-21T07:39:19.443702Z","shell.execute_reply.started":"2023-02-21T07:39:18.818964Z","shell.execute_reply":"2023-02-21T07:39:19.442592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Text Cleaning**\nOur data set is not clear, it contains uppercase, brackets, links, punctuation and so many things. We need to remove thoes things from our data. Here, we will use re library to fixed thoes things.","metadata":{}},{"cell_type":"code","source":"\ndef clean_text(text):\n    text = str(text).lower()\n    text = re.sub('\\[.*?\\]', '', text)\n    text = re.sub('https?://\\S+|www\\.\\S+', '', text)\n    text = re.sub('<.*?>+', '', text)\n    text = re.sub('[%s]' % re.escape(string.punctuation), '', text)\n    text = re.sub('\\n', '', text)\n    text = re.sub('\\w*\\d\\w*', '', text)\n    return text","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:39:27.098877Z","iopub.execute_input":"2023-02-21T07:39:27.099242Z","iopub.status.idle":"2023-02-21T07:39:27.108993Z","shell.execute_reply.started":"2023-02-21T07:39:27.099211Z","shell.execute_reply":"2023-02-21T07:39:27.106202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['text'] = data['text'].apply(clean_text)\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:39:33.721089Z","iopub.execute_input":"2023-02-21T07:39:33.721472Z","iopub.status.idle":"2023-02-21T07:40:00.470743Z","shell.execute_reply.started":"2023-02-21T07:39:33.721440Z","shell.execute_reply":"2023-02-21T07:40:00.469822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Remove Stopwords**\nIn Natural Language Processing (NLP), stop words are commonly occurring words that are filtered out before or after processing of text data. Stop words are usually words that do not contribute much to the meaning of a sentence or document, and are therefore not considered useful for text analysis. Examples of stop words include \"the\", \"and\", \"a\", \"an\", \"in\", \"of\", \"is\", \"to\", \"that\", \"it\", and so on.Removing stop words from a text can help reduce the dimensionality of the dataset, which can make analysis more efficient and effective.","metadata":{}},{"cell_type":"code","source":"stop_words = stopwords.words('english')\n\ndef remove_stopwords(text):\n    text = ' '.join(word for word in text.split(' ') if word not in stop_words)\n    return text\n    \ndata['text'] = data['text'].apply(remove_stopwords)\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:40:12.404869Z","iopub.execute_input":"2023-02-21T07:40:12.405244Z","iopub.status.idle":"2023-02-21T07:41:01.245594Z","shell.execute_reply.started":"2023-02-21T07:40:12.405198Z","shell.execute_reply":"2023-02-21T07:41:01.244635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Stemming**\nStemming is a technique used in natural language processing (NLP) to reduce words to their base or root form, which is called the stem. The goal of stemming is to reduce the inflectional and derivational forms of words to a common base form, which can simplify text analysis and improve the accuracy of text-based applications such as search engines, sentiment analysis, and text classification.\n\nFor example, the words \"running\", \"runs\", and \"ran\" can be reduced to their stem \"run\", which can help to identify them as variants of the same word and thus improve the accuracy of analysis.","metadata":{}},{"cell_type":"code","source":"stemmer = nltk.SnowballStemmer(\"english\")\n\ndef stemm_text(text):\n    text = ' '.join(stemmer.stem(word) for word in text.split(' '))\n    return text\n\ndata['text'] = data['text'].apply(stemm_text)\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:44:33.593168Z","iopub.execute_input":"2023-02-21T07:44:33.594309Z","iopub.status.idle":"2023-02-21T07:47:13.480100Z","shell.execute_reply.started":"2023-02-21T07:44:33.594263Z","shell.execute_reply":"2023-02-21T07:47:13.479069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Target Encoding**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\na = LabelEncoder()\na.fit(data['sentiment'])\n\ndata['sentiment'] = a.transform(data['sentiment'])\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:47:24.090806Z","iopub.execute_input":"2023-02-21T07:47:24.091271Z","iopub.status.idle":"2023-02-21T07:47:24.393327Z","shell.execute_reply.started":"2023-02-21T07:47:24.091229Z","shell.execute_reply":"2023-02-21T07:47:24.392270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Split data set**","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nx = data['text']\ny = data['sentiment']\nx_train, x_test, y_train, y_test = train_test_split(x, y, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:47:28.675362Z","iopub.execute_input":"2023-02-21T07:47:28.675728Z","iopub.status.idle":"2023-02-21T07:47:28.963382Z","shell.execute_reply.started":"2023-02-21T07:47:28.675696Z","shell.execute_reply":"2023-02-21T07:47:28.962399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **TF-IDF Vectorizer**\nTF-IDF stands for term frequency-inverse document frequency. It is a numerical statistic that is used to reflect how important a word is to a document in a collection or corpus of documents. It is a widely used model in information retrieval and text mining to measure the relevance of a word to a document.The main idea behind the TF-IDF model is to give more weight to words that appear frequently in a document but rarely in the corpus. This is done by computing two values:\n\nTerm Frequency (TF): The number of times a term appears in a document. A term that appears frequently in a document is likely to be important to the meaning of that document.\n\nInverse Document Frequency (IDF): The logarithmically scaled inverse fraction of the number of documents that contain the word. A term that appears in many documents is likely to be less important than a term that appears in only a few documents.\n","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_extraction.text import TfidfVectorizer\nvectoriser = TfidfVectorizer(ngram_range=(1,2),max_features=3891472)\nvectoriser.fit(x_train)\n\nx_train = vectoriser.transform(x_train)\nx_test  = vectoriser.transform(x_test)\n#print('No. of feature_words: ', len(vectoriser.get_feature_names()))","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:47:33.035989Z","iopub.execute_input":"2023-02-21T07:47:33.036464Z","iopub.status.idle":"2023-02-21T07:48:41.224103Z","shell.execute_reply.started":"2023-02-21T07:47:33.036421Z","shell.execute_reply":"2023-02-21T07:48:41.223053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Logistic Regression Model**\nLogistic regression is a type of classification algorithm used in natural language processing (NLP) for text classification tasks. In text classification, the goal is to predict the category or label of a given text document","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nmodel = LogisticRegression()\nmodel.fit(x_train, y_train)\n\n\ny_pred = model.predict(x_test)\naccuracy = accuracy_score(y_test, y_pred)\nprint(f\"Test accuracy: {accuracy:.4f}\")\nprint(classification_report(y_test, y_pred))\n","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:49:04.250170Z","iopub.execute_input":"2023-02-21T07:49:04.250551Z","iopub.status.idle":"2023-02-21T07:51:00.924928Z","shell.execute_reply.started":"2023-02-21T07:49:04.250519Z","shell.execute_reply":"2023-02-21T07:51:00.923625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Confusion Matrix**\nWe use confusion metrics to see the performance of classification algorithm. Confusion matrics represent the value of true positive(tp), false positive(fp) ,true negative(tn) and false negative(fn).","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score ,confusion_matrix, classification_report\nconfution_lg = confusion_matrix(y_test, y_pred) #confusion metrics\nsns.heatmap(confution_lg, linewidths=0.01, annot=True,fmt= '.1f', color='red') #heat map","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:51:55.692045Z","iopub.execute_input":"2023-02-21T07:51:55.692427Z","iopub.status.idle":"2023-02-21T07:51:55.954361Z","shell.execute_reply.started":"2023-02-21T07:51:55.692393Z","shell.execute_reply":"2023-02-21T07:51:55.953285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Prediction**","metadata":{}},{"cell_type":"code","source":"s = ['it is a bad question']\ns = vectoriser.transform(s)\nsentiment = model.predict(s)\nsentiment","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:52:03.266055Z","iopub.execute_input":"2023-02-21T07:52:03.266435Z","iopub.status.idle":"2023-02-21T07:52:03.348450Z","shell.execute_reply.started":"2023-02-21T07:52:03.266400Z","shell.execute_reply":"2023-02-21T07:52:03.347341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Naive Bayes Model**","metadata":{}},{"cell_type":"code","source":"# Create a Multinomial Naive Bayes model\nfrom sklearn.naive_bayes import MultinomialNB\nnb = MultinomialNB()\n\n# Train the model\nnb.fit(x_train, y_train)\n# Evaluate the model on the test set\ny_p = nb.predict(x_test)\naccuracy = accuracy_score(y_test, y_p)\nprint(f\"Test accuracy: {accuracy:.4f}\")\nprint(classification_report(y_test, y_p))","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:52:09.375449Z","iopub.execute_input":"2023-02-21T07:52:09.375809Z","iopub.status.idle":"2023-02-21T07:52:10.752795Z","shell.execute_reply.started":"2023-02-21T07:52:09.375779Z","shell.execute_reply":"2023-02-21T07:52:10.750881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Stochastic Gradient Descent**","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import SGDClassifier\nsgd = SGDClassifier()\nsgd.fit(x_train, y_train)\npred = sgd.predict(x_test)\nprint(\"test accuracy score of Stochastic Gradient Descent = \", accuracy_score(y_test, pred)*100)","metadata":{"execution":{"iopub.status.busy":"2023-02-21T07:52:17.905128Z","iopub.execute_input":"2023-02-21T07:52:17.905830Z","iopub.status.idle":"2023-02-21T07:52:21.092385Z","shell.execute_reply.started":"2023-02-21T07:52:17.905794Z","shell.execute_reply":"2023-02-21T07:52:21.091399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Special thanks to https://www.kaggle.com/code/andreshg/nlp-glove-bert-tf-idf-lstm-explained/notebook#3.-Data-Pre-processing-%F0%9F%9B%A0 this notebook.","metadata":{}}]}