{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Objective\n\nSentiment analysis is a common use case of NLP where the idea is to classify the tweet as positive, negative or neutral depending upon the text in the tweet. This problem goes a way ahead and expects us to also determine the words in the tweet which decide the polarity of the tweet.\n\n## Understanding the Evaluation Metric\n\nThe metric in this competition is the word-level **Jaccard score**. Jaccard Score is a measure of how similar/dissimilar two sets are.  The higher the score, the more similar the two strings. The idea is to find the number of common tokens and divide it by the total number of unique tokens. Its expressed in the mathematical terms by,\n\n![](https://imgur.com/lMHa8CL.png)\n\n![](https://images.deepai.org/glossary-terms/jaccard-index-391304.jpg)\n\n[Source](https://en.wikipedia.org/wiki/Jaccard_index)\n\nHere is a great example to understand the Jaccard Similarity Metric in an inutitve way.Refer to the main blog for more details:[FIVE MOST POPULAR SIMILARITY MEASURES IMPLEMENTATION IN PYTHON](https://dataaspirant.com/2015/04/11/five-most-popular-similarity-measures-implementation-in-python/)\n\n![](https://i0.wp.com/dataaspirant.com/wp-content/uploads/2015/04/jaccaard2.png?resize=768%2C307&ssl=1)\n\n**Here is how one can implement the jaccard score in Python:**","metadata":{}},{"cell_type":"code","source":"\ndef jaccard(str1, str2): \n    a = set(str1.lower().split()) \n    b = set(str2.lower().split())\n    c = a.intersection(b)\n    return float(len(c)) / (len(a) + len(b) - len(c))\n\n\nSentence_1 = 'Life well spent is life good'\nSentence_2 = 'Life is an art and it is good so far'\nSentence_3 = 'Life is good'\n\n    \nprint(jaccard(Sentence_1,Sentence_2))\nprint(jaccard(Sentence_1,Sentence_3))","metadata":{"execution":{"iopub.status.busy":"2022-07-22T05:58:05.040321Z","iopub.execute_input":"2022-07-22T05:58:05.040665Z","iopub.status.idle":"2022-07-22T05:58:05.049396Z","shell.execute_reply.started":"2022-07-22T05:58:05.040630Z","shell.execute_reply":"2022-07-22T05:58:05.048563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now begin with the Exploratory data analysis. Let's quickly go over the Table of contents to get an idea what we shall be covering in this notebook:\n\n# Table of Contents\n\n- 1. Dataset and Dependencies\n   - 1.1 Importing the necessary libraries\n   - 1.2 Reading in the Dataset\n- 2. General EDA\n   - 2.1 Missing Vales\n- 3. Analysis of the Sentiment Column\n   - 3.1 An example of each sentiment\n   - 3.2 Distribution of the Sentiment column\n- 4. Text Data Preprocessing\n- 5. Text Statistics\n   - 5.1 Sentence Length\n   - 5.2 Word Count\n   - 5.3 Word Frequency\n- 6. Ngram Analysis\n- 7. Exploring the selected_text column\n- 8  Wordclouds\n- 9. Extracting the sentiment terms : Resources\n  ","metadata":{}},{"cell_type":"markdown","source":"# 1. Datasets & Dependencies\n## 1.1 Importing the necessary libraries","metadata":{}},{"cell_type":"code","source":"from IPython.core.interactiveshell import InteractiveShell\nInteractiveShell.ast_node_interactivity = 'all'\n\n!pip install chart_studio\n!pip install textstat\n\nimport numpy as np \nimport pandas as pd \n\n# text processing libraries\nimport re\nimport string\nimport nltk\nfrom nltk.corpus import stopwords\n\n\n# Visualisation libraries\nimport matplotlib.pyplot as plt\nimport plotly.graph_objs as go\nimport chart_studio.plotly as py\nimport plotly.figure_factory as ff\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\n\n# sklearn \nfrom sklearn import model_selection\nfrom sklearn.feature_extraction.text import CountVectorizer,TfidfVectorizer\n\n# File system manangement\nimport os\n\n# Pytorch\nimport torch\n\n#Transformers\nfrom transformers import BertTokenizer\n\n# Suppress warnings \nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-22T05:58:05.051358Z","iopub.execute_input":"2022-07-22T05:58:05.052038Z","iopub.status.idle":"2022-07-22T05:58:16.593621Z","shell.execute_reply.started":"2022-07-22T05:58:05.051948Z","shell.execute_reply":"2022-07-22T05:58:16.593075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir('../input/')","metadata":{"execution":{"iopub.status.busy":"2022-07-22T05:58:16.594489Z","iopub.execute_input":"2022-07-22T05:58:16.594704Z","iopub.status.idle":"2022-07-22T05:58:16.601118Z","shell.execute_reply.started":"2022-07-22T05:58:16.594675Z","shell.execute_reply":"2022-07-22T05:58:16.600114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.2 Reading the datasets","metadata":{}},{"cell_type":"code","source":"#Training data\ntrain = pd.read_csv('../input/tweet-sentiment-extraction/train.csv')\ntest = pd.read_csv('../input/tweet-sentiment-extraction/test.csv')\nprint('Training data shape: ', train.shape)\nprint('Testing data shape: ', test.shape)\n\n# First few rows of the training dataset\ntrain.head()\n\n# First few rows of the testing dataset\ntest.head()","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","execution":{"iopub.status.busy":"2022-07-22T05:58:16.602529Z","iopub.execute_input":"2022-07-22T05:58:16.602886Z","iopub.status.idle":"2022-07-22T05:58:16.691907Z","shell.execute_reply.started":"2022-07-22T05:58:16.602828Z","shell.execute_reply":"2022-07-22T05:58:16.690902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The columns denote the following:\n\n* The `textID` of a tweet\n* The `text` of a tweet\n* The `selected text` which determines the polarity of the tweet\n* `sentiment` of the tweet\n\nThe test dataset doesn't have the selected text column which needs to be identified.","metadata":{}},{"cell_type":"markdown","source":"# 2. General EDA\n## 2.1 Missing Values treatment in the dataset\n","metadata":{}},{"cell_type":"code","source":"train.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:01:59.096085Z","iopub.execute_input":"2022-07-22T06:01:59.096638Z","iopub.status.idle":"2022-07-22T06:01:59.109633Z","shell.execute_reply.started":"2022-07-22T06:01:59.096594Z","shell.execute_reply":"2022-07-22T06:01:59.108385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:01:01.348549Z","iopub.execute_input":"2022-07-22T06:01:01.348868Z","iopub.status.idle":"2022-07-22T06:01:01.357424Z","shell.execute_reply.started":"2022-07-22T06:01:01.348830Z","shell.execute_reply":"2022-07-22T06:01:01.356492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.dropna(axis=0, how='any', inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:01:50.297161Z","iopub.execute_input":"2022-07-22T06:01:50.297525Z","iopub.status.idle":"2022-07-22T06:01:50.319776Z","shell.execute_reply.started":"2022-07-22T06:01:50.297480Z","shell.execute_reply":"2022-07-22T06:01:50.318640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 감성 column 분석\n\ntrain[train['sentiment'] == 'positive']['text'].values[0]","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:04:55.030220Z","iopub.execute_input":"2022-07-22T06:04:55.030804Z","iopub.status.idle":"2022-07-22T06:04:55.042545Z","shell.execute_reply.started":"2022-07-22T06:04:55.030763Z","shell.execute_reply":"2022-07-22T06:04:55.041344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['sentiment'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:05:39.420652Z","iopub.execute_input":"2022-07-22T06:05:39.421033Z","iopub.status.idle":"2022-07-22T06:05:39.435602Z","shell.execute_reply.started":"2022-07-22T06:05:39.420969Z","shell.execute_reply":"2022-07-22T06:05:39.434693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['sentiment'].value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:06:09.197851Z","iopub.execute_input":"2022-07-22T06:06:09.198290Z","iopub.status.idle":"2022-07-22T06:06:09.213488Z","shell.execute_reply.started":"2022-07-22T06:06:09.198230Z","shell.execute_reply":"2022-07-22T06:06:09.212463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['sentiment'].value_counts(normalize=True).iplot()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:06:57.761874Z","iopub.execute_input":"2022-07-22T06:06:57.762351Z","iopub.status.idle":"2022-07-22T06:06:58.479820Z","shell.execute_reply.started":"2022-07-22T06:06:57.762152Z","shell.execute_reply":"2022-07-22T06:06:58.479005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['sentiment'].value_counts(normalize=True).iplot(kind='bar', yTitle='Ratio', linecolor='black', \n                                                      opacity=0.7, color='red', theme='pearl',\n                                                      bargap=0.6, gridcolor='white', title='dist'\n                                        )","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:10:28.588326Z","iopub.execute_input":"2022-07-22T06:10:28.588686Z","iopub.status.idle":"2022-07-22T06:10:28.999789Z","shell.execute_reply.started":"2022-07-22T06:10:28.588601Z","shell.execute_reply":"2022-07-22T06:10:28.998496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# text preprocessing helper functions\n\ndef clean_text(text):\n    '''Make text lowercase, remove text in square brackets,remove links,remove punctuation\n    and remove words containing numbers.'''\n    text = text.lower()\n    text = re.sub('\\[.*?\\]', '', text)\n    text = re.sub('https?://\\S+|www\\.\\S+', '', text)\n    text = re.sub('<.*?>+', '', text)\n    text = re.sub('[%s]' % re.escape(string.punctuation), '', text)\n    text = re.sub('\\n', '', text)\n    text = re.sub('\\w*\\d\\w*', '', text)\n    return text\n\n\ndef text_preprocessing(text):\n    \"\"\"\n    Cleaning and parsing the text.\n\n    \"\"\"\n    tokenizer = nltk.tokenize.RegexpTokenizer(r'\\w+')\n    nopunc = clean_text(text)\n    tokenized_text = tokenizer.tokenize(nopunc)\n    #remove_stopwords = [w for w in tokenized_text if w not in stopwords.words('english')]\n    combined_text = ' '.join(tokenized_text)\n    return combined_text","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:12:15.558578Z","iopub.execute_input":"2022-07-22T06:12:15.558938Z","iopub.status.idle":"2022-07-22T06:12:15.568169Z","shell.execute_reply.started":"2022-07-22T06:12:15.558893Z","shell.execute_reply":"2022-07-22T06:12:15.567294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['text_clean'] = train['text'].apply(str).apply(lambda x: text_preprocessing(x))\ntest['text_clean'] = test['text'].apply(str).apply(lambda x: text_preprocessing(x))\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:14:25.941400Z","iopub.execute_input":"2022-07-22T06:14:25.941921Z","iopub.status.idle":"2022-07-22T06:14:27.936347Z","shell.execute_reply.started":"2022-07-22T06:14:25.941884Z","shell.execute_reply":"2022-07-22T06:14:27.935478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# analyzing text 통계\n\ntrain['text_len'] = train['text_clean'].astype(str).apply(len)\ntrain.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:16:24.968943Z","iopub.execute_input":"2022-07-22T06:16:24.969567Z","iopub.status.idle":"2022-07-22T06:16:25.004468Z","shell.execute_reply.started":"2022-07-22T06:16:24.969525Z","shell.execute_reply":"2022-07-22T06:16:25.003415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['text_word_count'] = train['text_clean'].apply(lambda x: len(str(x).split()))","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:19:58.832157Z","iopub.execute_input":"2022-07-22T06:19:58.832447Z","iopub.status.idle":"2022-07-22T06:19:58.871815Z","shell.execute_reply.started":"2022-07-22T06:19:58.832418Z","shell.execute_reply":"2022-07-22T06:19:58.870623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:20:01.476555Z","iopub.execute_input":"2022-07-22T06:20:01.477093Z","iopub.status.idle":"2022-07-22T06:20:01.490856Z","shell.execute_reply.started":"2022-07-22T06:20:01.477053Z","shell.execute_reply":"2022-07-22T06:20:01.490111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pos = train[train['sentiment']=='positive']\nneg = train[train['sentiment']=='negative']\nneutral = train[train['sentiment']=='neutral']","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:21:29.547577Z","iopub.execute_input":"2022-07-22T06:21:29.548178Z","iopub.status.idle":"2022-07-22T06:21:29.579109Z","shell.execute_reply.started":"2022-07-22T06:21:29.548138Z","shell.execute_reply":"2022-07-22T06:21:29.578358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pos['text_len'].iplot(kind='hist',bins=100)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:22:52.858942Z","iopub.execute_input":"2022-07-22T06:22:52.859251Z","iopub.status.idle":"2022-07-22T06:22:53.525529Z","shell.execute_reply.started":"2022-07-22T06:22:52.859221Z","shell.execute_reply":"2022-07-22T06:22:53.524615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trace0 = go.Box(y=pos['text_len'], name='pos', marker=dict(color='red',))\ntrace1 = go.Box(y=neg['text_len'], name='neg', marker=dict(color='green',))\ntrace2 = go.Box(y=neutral['text_len'], name='neu', marker=dict(color='blue',))\n\ndata = [trace0, trace1, trace2]\nlayout = go.Layout(title='len')\n\nfig = go.Figure(data=data, layout=layout)\niplot(fig)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:28:38.418447Z","iopub.execute_input":"2022-07-22T06:28:38.418770Z","iopub.status.idle":"2022-07-22T06:28:38.773199Z","shell.execute_reply.started":"2022-07-22T06:28:38.418723Z","shell.execute_reply":"2022-07-22T06:28:38.772343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from transformers import BertTokenizer","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:30:51.458777Z","iopub.execute_input":"2022-07-22T06:30:51.459277Z","iopub.status.idle":"2022-07-22T06:30:51.462522Z","shell.execute_reply.started":"2022-07-22T06:30:51.459241Z","shell.execute_reply":"2022-07-22T06:30:51.461847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenizer = BertTokenizer.from_pretrained('bert-base-uncased', do_lower_case=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:31:50.528397Z","iopub.execute_input":"2022-07-22T06:31:50.528882Z","iopub.status.idle":"2022-07-22T06:31:50.690423Z","shell.execute_reply.started":"2022-07-22T06:31:50.528837Z","shell.execute_reply":"2022-07-22T06:31:50.689274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"text_obs = train['text'][10]\nprint(text_obs)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:32:48.188227Z","iopub.execute_input":"2022-07-22T06:32:48.188645Z","iopub.status.idle":"2022-07-22T06:32:48.194618Z","shell.execute_reply.started":"2022-07-22T06:32:48.188614Z","shell.execute_reply":"2022-07-22T06:32:48.193194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(tokenizer.tokenize(text_obs, add_special_tokens=True))","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:33:34.628587Z","iopub.execute_input":"2022-07-22T06:33:34.628887Z","iopub.status.idle":"2022-07-22T06:33:34.634759Z","shell.execute_reply.started":"2022-07-22T06:33:34.628846Z","shell.execute_reply":"2022-07-22T06:33:34.633797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_token = tokenizer.tokenize(text_obs, add_special_tokens=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:34:44.692563Z","iopub.execute_input":"2022-07-22T06:34:44.693228Z","iopub.status.idle":"2022-07-22T06:34:44.697988Z","shell.execute_reply.started":"2022-07-22T06:34:44.693184Z","shell.execute_reply":"2022-07-22T06:34:44.697133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(tokenizer.convert_tokens_to_ids(data_token))","metadata":{"execution":{"iopub.status.busy":"2022-07-22T06:34:56.908288Z","iopub.execute_input":"2022-07-22T06:34:56.908792Z","iopub.status.idle":"2022-07-22T06:34:56.913789Z","shell.execute_reply.started":"2022-07-22T06:34:56.908754Z","shell.execute_reply":"2022-07-22T06:34:56.912614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Bert 를 위한 전처리","metadata":{},"execution_count":null,"outputs":[]}]}