{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Abstract\nAfter having built [the baseline model (version 0.0)](https://www.kaggle.com/lukeshrek/classify-toxic-question) with Logistics Regression and TF-IDF Bi-grams Vectorizer, I personally found that model is still very simple and can be further improved to give better results.\n\nIn this version 1.0, I will proceed to go deeper into data analysis and visualization, improve the preprocessing steps and build the model based on the Bidirectional GRU network.\n\nThis notebook represent Data Exploration and Analysis works. ","metadata":{}},{"cell_type":"markdown","source":"## Initial Configurations","metadata":{}},{"cell_type":"markdown","source":"In this notebook I will use various package to support data visualize and analyze process.\n\n* Textstat: Textstat is an easy to use library to calculate statistics from text. It helps determine readability, complexity, and grade level.\n* Chart Studio: Chart Studio provides a web-service for hosting graphs. Graphs are saved inside online Chart Studio account.","metadata":{}},{"cell_type":"code","source":"!pip install textstat\n!pip install chart_studio","metadata":{"_uuid":"afe6eadd-248c-4c79-a349-39dbc03f0e8e","_cell_guid":"34a6222f-715a-4e1b-bc4e-897a077620ec","collapsed":false,"jupyter":{"outputs_hidden":false},"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-06-10T20:10:24.144035Z","iopub.execute_input":"2021-06-10T20:10:24.144403Z","iopub.status.idle":"2021-06-10T20:10:36.957090Z","shell.execute_reply.started":"2021-06-10T20:10:24.144324Z","shell.execute_reply":"2021-06-10T20:10:36.955920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These packages are available for installing using pip.","metadata":{}},{"cell_type":"markdown","source":"Code cell below are modules and libraries used in this notebook:\n\n* os: provides functions for interacting with the operating system. \n* json: built-in package which can be used to work with JSON data.\n* string: provides additional tools to manipulate strings.\n* math: built-in module for mathematical tasks.\n* collections: implements specialized container datatypes providing alternatives to Python’s general purpose built-in containers, dict, list, set, and tuple.\n* warnings: filterwarnings(action) with action as \"ignore\" to suppress all warnings.\n* statistics: provides functions for calculating mathematical statistics of numeric (Real-valued) data.\n* tqdm: output a progress bar by wrapping around any iterable.\n\n\n* NumPy: NumPy is the fundamental package for array computing with Python.\n* pandas: pandas is a fast, powerful, flexible and easy to use open source data analysis and manipulation tool, built on top of the Python programming language.\n* Matplotlib: Matplotlib is a comprehensive library for creating static, animated, and interactive visualizations in Python.\n* Seaborn: Seaborn is a Python data visualization library based on matplotlib. It provides a high-level interface for drawing attractive and informative statistical graphics.\n* WordCloud: Word cloud is a technique for visualising frequent words in a text where the size of the words represents their frequency.\n* plotly: An open-source, interactive data visualization library for Python.\n* spaCy: spaCy is a free open-source library for Natural Language Processing in Python. It features NER, POS tagging, dependency parsing, word vectors and more.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport json\nimport string\nimport math\n\nfrom statistics import *\nfrom tqdm import tqdm\n\n# Factory function to supply missing values\nfrom collections import defaultdict\n\n# Warnings control\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport numpy as np\nimport pandas as pd\nfrom pandas.io.json import json_normalize\npd.options.mode.chained_assignment = None\npd.options.display.max_columns = 999\n\n# Textstat\nimport textstat\n\n# Imports for visualization\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\ncolor = sns.color_palette()\n\n# plotly based imports\nfrom plotly import tools\nimport plotly.offline as py\nfrom plotly.offline import init_notebook_mode, iplot\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.figure_factory as ff\n\n# Wordcloud library\nfrom wordcloud import WordCloud, ImageColorGenerator, STOPWORDS\n\n# spaCy based imports\nimport spacy\nfrom spacy.lang.en.stop_words import STOP_WORDS\nfrom spacy.lang.en import English\n\n# spaCy Parser for questions\npunctuations = string.punctuation\nstopwords = list(STOP_WORDS)\nparser = English()","metadata":{"_uuid":"f334d49e-f799-4f31-80d5-7cdb232dd40f","_cell_guid":"5ce0f078-ffc0-4f46-8a7e-fbba0596b35c","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:10:36.959770Z","iopub.execute_input":"2021-06-10T20:10:36.960073Z","iopub.status.idle":"2021-06-10T20:10:41.472223Z","shell.execute_reply.started":"2021-06-10T20:10:36.960042Z","shell.execute_reply":"2021-06-10T20:10:41.470809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Introduction","metadata":{}},{"cell_type":"markdown","source":"## Problem Description\n        \nAn existential problem for any major website today is how to handle toxic and divisive content. Quora wants to tackle this problem head-on to keep their platform a place where users can feel safe sharing their knowledge with the world.\n\nQuora is a platform that empowers people to learn from each other. On Quora, people can ask questions and connect with others who contribute unique insights and quality answers. A key challenge is to weed out insincere questions - those founded upon false premises, or that intend to make a statement rather than look for helpful answers.\n\nIn this competition, Kagglers will develop models that identify and flag insincere questions. To date, Quora has employed both machine learning and manual review to address this problem. More scalable methods could be developed to detect toxic and misleading content.","metadata":{}},{"cell_type":"markdown","source":"## Data Description\nIn this competition the model should be able to detect whether a question asked on Quora is sincere or not. An insincere question is defined as a question intended to make a statement rather than look for helpful answers. Some characteristics that can signify that a question is insincere:\n\n* Has a non-neutral tone\n    * Has an exaggerated tone to underscore a point about a group of people\n    * Is rhetorical and meant to imply a statement about a group of people\n* Is disparaging or inflammatory\n    * Suggests a discriminatory idea against a protected class of people, or seeks confirmation of a stereotype\n    * Makes disparaging attacks/insults against a specific person or group of people\n    * Based on an outlandish premise about a group of people\n    * Disparages against a characteristic that is not fixable and not measurable\n* Isn't grounded in reality\n    * Based on false information, or contains absurd assumptions\n* Uses sexual content (incest, bestiality, pedophilia) for shock value, and not to seek genuine answers\n\nThe training data includes the question that was asked, and whether it was identified as insincere (target = 1) or not (target = 0).","metadata":{}},{"cell_type":"markdown","source":"# 2. Data Exploration and Visualization","metadata":{}},{"cell_type":"markdown","source":"## Input","metadata":{}},{"cell_type":"markdown","source":"We will take a look at the input data directory","metadata":{}},{"cell_type":"code","source":"!ls ../input/quora-insincere-questions-classification","metadata":{"execution":{"iopub.status.busy":"2021-06-10T20:10:41.473674Z","iopub.execute_input":"2021-06-10T20:10:41.473924Z","iopub.status.idle":"2021-06-10T20:10:42.214088Z","shell.execute_reply.started":"2021-06-10T20:10:41.473902Z","shell.execute_reply":"2021-06-10T20:10:42.213149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* train.csv - the training set\n* test.csv - the test set\n* sample_submission.csv - A sample submission in the correct format\n* enbeddings/ - Folder containing word embeddings.\n\nExternal data sources are not allowed to use. The following embeddings are given to us which can be used for building our models.","metadata":{}},{"cell_type":"markdown","source":"Let's see how train data and test data are distributed.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/quora-insincere-questions-classification/train.csv\")\ntest_df = pd.read_csv(\"../input/quora-insincere-questions-classification/test.csv\")\nprint(\"Train shape: \", train_df.shape)\nprint(\"Test shape: \", test_df.shape)","metadata":{"_uuid":"1e6cb82d-1804-48fd-9957-a2241fe9b777","_cell_guid":"2b265759-77d2-47d2-b981-72863ef86503","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:10:42.215928Z","iopub.execute_input":"2021-06-10T20:10:42.216345Z","iopub.status.idle":"2021-06-10T20:10:47.708036Z","shell.execute_reply.started":"2021-06-10T20:10:42.216301Z","shell.execute_reply":"2021-06-10T20:10:47.706819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-10T20:10:47.711939Z","iopub.execute_input":"2021-06-10T20:10:47.712219Z","iopub.status.idle":"2021-06-10T20:10:47.747673Z","shell.execute_reply.started":"2021-06-10T20:10:47.712193Z","shell.execute_reply":"2021-06-10T20:10:47.746651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.dtypes","metadata":{"execution":{"iopub.status.busy":"2021-06-10T20:10:47.751837Z","iopub.execute_input":"2021-06-10T20:10:47.752133Z","iopub.status.idle":"2021-06-10T20:10:47.760647Z","shell.execute_reply.started":"2021-06-10T20:10:47.752107Z","shell.execute_reply":"2021-06-10T20:10:47.759697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* qid - unique question identifier\n* question_text - Quora question text\n* target - a question labeled \"insincere\" has a value of 1, otherwise 0","metadata":{}},{"cell_type":"markdown","source":"## Embeddings ","metadata":{}},{"cell_type":"markdown","source":"Basically embeddings folder are zipped so we cannot show its details by using ```!ls```, but information about the allowed embeddings are available in the competition's data description.","metadata":{}},{"cell_type":"code","source":"!ls ../input/quora-insincere-questions-classification/embeddings.zip","metadata":{"execution":{"iopub.status.busy":"2021-06-10T20:10:47.762654Z","iopub.execute_input":"2021-06-10T20:10:47.763151Z","iopub.status.idle":"2021-06-10T20:10:48.502918Z","shell.execute_reply.started":"2021-06-10T20:10:47.763101Z","shell.execute_reply":"2021-06-10T20:10:48.501625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n* glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n* paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n* wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html","metadata":{}},{"cell_type":"markdown","source":"## Target Count and Target Distribution","metadata":{}},{"cell_type":"markdown","source":"It can be easily detected that the dataset is imbalance, in which the vast majority of questions are sincere, and only a small number are insincere. Let us look at the distribution of the target variable by plotting a bar chart and a pie chart to understand more.","metadata":{}},{"cell_type":"code","source":"# Target count\ncnt_srs = train_df['target'].value_counts()\ntrace = go.Bar(\nx=cnt_srs.index,\n    y=cnt_srs.values,\n    marker=dict(\n        color=cnt_srs.values,\n        colorscale = 'Picnic',\n        reversescale = True\n    ),\n)\n\nlayout = go.Layout(\n    title='Target Count',\n    font=dict(size=18)\n)\n\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig, filename=\"TargetCount\")\n\n# Target distribution\nlabels = (np.array(cnt_srs.index))\nsizes = (np.array((cnt_srs / cnt_srs.sum())*100))\n\ntrace = go.Pie(labels=labels, values=sizes)\nlayout = go.Layout(\n    title='Target distribution',\n    font=dict(size=18),\n    width=600,\n    height=600,\n)\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig, filename=\"usertype\")","metadata":{"_uuid":"e5f592fc-9e36-471d-9f56-0b568135fcbe","_cell_guid":"00dfc9c7-e1f7-41b5-836b-b8e5183a0217","execution":{"iopub.status.busy":"2021-06-10T20:10:48.504559Z","iopub.execute_input":"2021-06-10T20:10:48.504862Z","iopub.status.idle":"2021-06-10T20:10:49.439354Z","shell.execute_reply.started":"2021-06-10T20:10:48.504830Z","shell.execute_reply":"2021-06-10T20:10:49.438387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So about 6% of the training data are insincere questions (target=1) and rest of them are sincere.","metadata":{}},{"cell_type":"markdown","source":"## WordCloud of questions","metadata":{}},{"cell_type":"markdown","source":"To find the most frequently occuring words in questions, I built word clouds of a random sample of 1000 insincere and 1000 sincere questions on the ```question_text``` column.\n\nFirst of all, questions are concatenated into a single string. Next, the question string is splitted into a dictionary and each words are counted uniquely to find the time that word appears. Finally, ```generate_from_frequencies``` from WordCloud libraries is used to plot the word cloud from frequency dictionary.","metadata":{}},{"cell_type":"code","source":"# Split sentences into a dictionary of uniquely occuring words and their frequencies\ndef word_freq_dict(text):\n    # Convert text into word list\n    wordList = text.split()\n    # Generate word freq dictionary\n    wordFreqDict = {word: wordList.count(word) for word in wordList}\n    return wordFreqDict\n\n# Plot a wordcloud from a word frequency dictionary\ndef word_cloud_from_frequency(word_freq_dict, title, figure_size=(10,6)):\n    wordcloud.generate_from_frequencies(word_freq_dict)\n    plt.figure(figsize=figure_size)\n    plt.imshow(wordcloud)\n    plt.axis(\"off\")\n    plt.title(title)\n    plt.show()","metadata":{"_uuid":"00d0d2b2-9e04-4499-8c0c-bc256fcb54c4","_cell_guid":"26e7e1cf-464b-45c3-8b85-e7787cedd870","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:10:49.441059Z","iopub.execute_input":"2021-06-10T20:10:49.441443Z","iopub.status.idle":"2021-06-10T20:10:49.447464Z","shell.execute_reply.started":"2021-06-10T20:10:49.441404Z","shell.execute_reply":"2021-06-10T20:10:49.446648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Wordcloud of a random sample of 1000 insincere questions\ninsincere_questions = train_df.question_text[train_df['target'] == 1]\ninsincere_sample = \" \".join(insincere_questions.sample(1000, random_state=1).values)\ninsincere_word_freq = word_freq_dict(insincere_sample)\nwordcloud = WordCloud(width= 5000,\n    height=3000,\n    max_words=200,\n    colormap='Reds',\n    background_color='white')\n\nword_cloud_from_frequency(insincere_word_freq, \"Most Frequent Words in a sample of 1000 raw questions flagged insincere\")","metadata":{"_uuid":"3f8b55c5-337c-4f6f-bdfb-c4c741218236","_cell_guid":"2ac673b5-489e-45fc-a4a2-fce3a4e83ec8","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:10:49.448983Z","iopub.execute_input":"2021-06-10T20:10:49.449255Z","iopub.status.idle":"2021-06-10T20:11:38.789254Z","shell.execute_reply.started":"2021-06-10T20:10:49.449218Z","shell.execute_reply":"2021-06-10T20:11:38.788211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Wordcloud of a random sample of 1000 sincere questions\nsincere_questions = train_df.question_text[train_df['target'] == 0]\nsincere_sample = \" \".join(sincere_questions.sample(1000, random_state=1).values)\nsincere_word_freq = word_freq_dict(sincere_sample)\nwordcloud = WordCloud(width= 5000,\n    height=3000,\n    max_words=200,\n    colormap='Blues',\n    background_color='white')\n\nword_cloud_from_frequency(sincere_word_freq, \"Most Frequent Words in a sample of 1000 raw questions flagged sincere\")","metadata":{"_uuid":"af1572ca-6d93-480d-ba05-14763eff0593","_cell_guid":"68ef8659-97bc-4744-8847-8ddb28569c60","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:11:38.790937Z","iopub.execute_input":"2021-06-10T20:11:38.791475Z","iopub.status.idle":"2021-06-10T20:12:21.062661Z","shell.execute_reply.started":"2021-06-10T20:11:38.791435Z","shell.execute_reply":"2021-06-10T20:12:21.061749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are various of word presents in both types of questions (obviously).\nBut these two wordcloud seems to be not useful at all, which can be explained by two main reason: (i) wordcloud is made by gathering only 2000 of over 1000000 questions, so total figure can be way more different, and (ii) too many noise and uninformative words(such as what, when, etc., which obviously appears in a question no matter it is toxic or not).\n\nMaybe it is a good idea to look at the most frequent words in each of the classes separately. ","metadata":{}},{"cell_type":"markdown","source":"## Word n-grams Count Plot\n\nIn the following cell, I have created 1, 2 and 3-gram count plot with similar process.\nBefore looking up for most useful frequent grams, it has to be ensure that all of the stopwords need to be eliminated from the counter. Words are all lowercased before zipping a number of words (based on parameter ```n_gram``` value). Finally, a frequency dictionary is created and count all words appear.","metadata":{}},{"cell_type":"markdown","source":"### Unigram","metadata":{}},{"cell_type":"code","source":"train_insincere_df = train_df[train_df[\"target\"]==1]\ntrain_sincere_df = train_df[train_df[\"target\"]==0]\n\nstopwords = set(STOPWORDS)\nmore_stopwords = {'one', 'br', 'Po', 'th', 'sayi', 'fo', 'Unknown'}\nstopwords = stopwords.union(more_stopwords)\n    \n# N-gram generation\ndef generate_ngrams(text, n_gram=1):\n    token = [token for token in text.lower().split(\" \") if token != \"\" if token not in STOPWORDS]\n    ngrams = zip(*[token[i:] for i in range(n_gram)])\n    return [\" \".join(ngram) for ngram in ngrams]\n\n# Horizontal bar chart\ndef horizontal_bar_chart(df, color):\n    trace = go.Bar(\n        y=df[\"word\"].values[::-1],\n        x=df[\"wordcount\"].values[::-1],\n        showlegend=False,\n        orientation = 'h',\n        marker=dict(\n            color=color,\n        ),\n    )\n    return trace","metadata":{"_uuid":"fdc9aaca-611b-43f0-9bb3-f9ed2638870f","_cell_guid":"0151240c-ee34-4069-a35e-92e6209a94fc","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:12:21.063889Z","iopub.execute_input":"2021-06-10T20:12:21.064378Z","iopub.status.idle":"2021-06-10T20:12:21.185126Z","shell.execute_reply.started":"2021-06-10T20:12:21.064341Z","shell.execute_reply":"2021-06-10T20:12:21.184319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the bar chart from sincere questions #\nfreq_dict = defaultdict(int)\nfor sent in train_sincere_df[\"question_text\"]:\n    for word in generate_ngrams(sent):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace0 = horizontal_bar_chart(fd_sorted.head(50), 'blue')\n\n# Get the bar chart from insincere questions #\nfreq_dict = defaultdict(int)\nfor sent in train_insincere_df[\"question_text\"]:\n    for word in generate_ngrams(sent):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace1 = horizontal_bar_chart(fd_sorted.head(50), 'red')\n\n# Creating two subplots\nfig = tools.make_subplots(rows=1, cols=2, vertical_spacing=0.04,\n                          subplot_titles=[\"Frequent words of sincere questions\", \n                                          \"Frequent words of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=900, paper_bgcolor='rgb(233,233,233)', title=\"Word Count Plots\")\npy.iplot(fig, filename='word-plots')","metadata":{"execution":{"iopub.status.busy":"2021-06-10T20:12:21.186385Z","iopub.execute_input":"2021-06-10T20:12:21.186669Z","iopub.status.idle":"2021-06-10T20:12:30.501227Z","shell.execute_reply.started":"2021-06-10T20:12:21.186641Z","shell.execute_reply":"2021-06-10T20:12:30.500434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Observations from unigram count plot***\n\n* Some of the top words are common across both the classes like 'people', 'will', 'think' etc.\n* Many of top words in sincere questions (after excluding the both common ones) used for describe and comparison purpose: 'best', 'good', 'much', etc.\n* The other top words in insincere questions (after excluding the both common ones) often involve matters that may be sensitive or controversial: 'trump', 'women', 'white', etc.","metadata":{}},{"cell_type":"markdown","source":"### Bigram","metadata":{}},{"cell_type":"markdown","source":"As mentioned above, bigram and trigram count plot are made by the same process as unigram: Lowercase non-stopword words, zip into grams and count. ","metadata":{}},{"cell_type":"code","source":"freq_dict = defaultdict(int)\nfor sent in train_sincere_df[\"question_text\"]:\n    for word in generate_ngrams(sent,2):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace0 = horizontal_bar_chart(fd_sorted.head(50), 'blue')\n\n\nfreq_dict = defaultdict(int)\nfor sent in train_insincere_df[\"question_text\"]:\n    for word in generate_ngrams(sent,2):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace1 = horizontal_bar_chart(fd_sorted.head(50), 'red')\n\n# Creating two subplots\nfig = tools.make_subplots(rows=1, cols=2, vertical_spacing=0.04,horizontal_spacing=0.15,\n                          subplot_titles=[\"Frequent bigrams of sincere questions\", \n                                          \"Frequent bigrams of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=900, paper_bgcolor='rgb(233,233,233)', title=\"Bigram Count Plots\")\npy.iplot(fig, filename='word-plots')","metadata":{"_uuid":"7bc61d2b-37dd-42ac-b7ab-876a0583fdcf","_cell_guid":"bb651707-6369-4a9d-bd1e-4a68070343a1","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:12:30.502553Z","iopub.execute_input":"2021-06-10T20:12:30.502811Z","iopub.status.idle":"2021-06-10T20:12:45.133238Z","shell.execute_reply.started":"2021-06-10T20:12:30.502784Z","shell.execute_reply":"2021-06-10T20:12:45.132475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Observations from bigram count plot***\n\n* Some of the top bigrams are common across both the classes are the combinations of a top word among both classes and a top word among one class, especially among sincere questions, for example: 'best way'. People tend to find useful answer for the best method to do something, therefore, they may need to provide informative question.\n* It can be inspected that several topics appear: 'computer science', 'machine learning', 'tv shows', etc. \n* The top bigrams in insincere questions still involve matters that may be sensitive or controversial (and significantly related to people, religion, politics) such as 'donald trump', 'white people', 'black people', etc. This observation is quite reasonable and similar to characteristics (provided in the competition description) that can signify that a question is insincere.","metadata":{}},{"cell_type":"markdown","source":"### Trigram","metadata":{}},{"cell_type":"code","source":"freq_dict = defaultdict(int)\nfor sent in train_sincere_df[\"question_text\"]:\n    for word in generate_ngrams(sent,3):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace0 = horizontal_bar_chart(fd_sorted.head(50), 'blue')\n\n\nfreq_dict = defaultdict(int)\nfor sent in train_insincere_df[\"question_text\"]:\n    for word in generate_ngrams(sent,3):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace1 = horizontal_bar_chart(fd_sorted.head(50), 'red')\n\n# Creating two subplots\nfig = tools.make_subplots(rows=1, cols=2, vertical_spacing=0.04, horizontal_spacing=0.2,\n                          subplot_titles=[\"Frequent trigrams of sincere questions\", \n                                          \"Frequent trigrams of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=1200, paper_bgcolor='rgb(233,233,233)', title=\"Trigram Count Plots\")\npy.iplot(fig, filename='word-plots')","metadata":{"_uuid":"88701aa8-f1fa-4fe8-a522-5509143d896d","_cell_guid":"c8bbdd71-1ddd-40ca-a5ad-cdc0c12c1546","execution":{"iopub.status.busy":"2021-06-10T20:12:45.134419Z","iopub.execute_input":"2021-06-10T20:12:45.134670Z","iopub.status.idle":"2021-06-10T20:12:59.187494Z","shell.execute_reply.started":"2021-06-10T20:12:45.134645Z","shell.execute_reply":"2021-06-10T20:12:59.186453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Observations from trigram count plot***\n\n* We can now easily inspect many trigrams are combinations of popular uni- and bigrams. So further plot count (```n>3```) is now unnecessary.\n* More detailed and informational phrases appear: 'black lives matter', 'gun control advocates', etc. \n* Unexpectedly, people's ages were mentioned frequently in the questions, in both categories. Perhaps people really care a lot about characteristics such as physical, psychological, cognitive level of others - traits that are influenced by certain age.","metadata":{}},{"cell_type":"markdown","source":"## Meta Features\n\nFor further exploration, let's create some meta features and then look at how they are distributed between the classes. The ones that we will create are\n\n* Number of words in the text\n* Number of unique words in the text\n* Number of characters in the text\n* Number of stopwords\n* Number of punctuations\n* Number of upper case words\n* Number of title case words\n* Average length of the words","metadata":{}},{"cell_type":"code","source":"# Number of words in the text\ntrain_df[\"num_words\"] = train_df[\"question_text\"].apply(lambda x: len(str(x).split()))\ntest_df[\"num_words\"] = test_df[\"question_text\"].apply(lambda x: len(str(x).split()))\n\n# Number of unique words in the text\ntrain_df[\"num_unique_words\"] = train_df[\"question_text\"].apply(lambda x: len(set(str(x).split())))\ntest_df[\"num_unique_words\"] = test_df[\"question_text\"].apply(lambda x: len(set(str(x).split())))\n\n# Number of characters in the text\ntrain_df[\"num_chars\"] = train_df[\"question_text\"].apply(lambda x: len(str(x)))\ntest_df[\"num_chars\"] = test_df[\"question_text\"].apply(lambda x: len(str(x)))\n\n# Number of stopwords in the text\ntrain_df[\"num_stopwords\"] = train_df[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\ntest_df[\"num_stopwords\"] = test_df[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\n\n# Number of punctuations in the text\ntrain_df[\"num_punctuations\"] =train_df['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\ntest_df[\"num_punctuations\"] =test_df['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\n\n# Number of title case words in the text\ntrain_df[\"num_words_upper\"] = train_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\ntest_df[\"num_words_upper\"] = test_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\n\n# Number of title case words in the text\ntrain_df[\"num_words_title\"] = train_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\ntest_df[\"num_words_title\"] = test_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\n\n# Average length of the words in the text\ntrain_df[\"mean_word_len\"] = train_df[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ntest_df[\"mean_word_len\"] = test_df[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))","metadata":{"_uuid":"039a1621-c5a1-4ded-9e05-63d0271a14b9","_cell_guid":"78187b37-fb57-421c-96eb-e72f11eafbc9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:12:59.188692Z","iopub.execute_input":"2021-06-10T20:12:59.188935Z","iopub.status.idle":"2021-06-10T20:13:47.263252Z","shell.execute_reply.started":"2021-06-10T20:12:59.188912Z","shell.execute_reply":"2021-06-10T20:13:47.262208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plotting meta features","metadata":{}},{"cell_type":"markdown","source":"To have a better view of how these meta features are distributed I use some box plots.\n\nIn descriptive statistics, a box plot or boxplot (also known as box and whisker plot) is a type of chart often used in explanatory data analysis. Box plots visually show the distribution of numerical data and skewness through displaying the data quartiles (or percentiles) and averages.\n\nBox plots show the five-number summary of a set of data: including the minimum score, first (lower) quartile, median, third (upper) quartile, and maximum score.\n\n* Minimum Score: The lowest score, excluding outliers (shown at the end of the left whisker).\n* Lower Quartile: Twenty-five percent of scores fall below the lower quartile value (also known as the first quartile).\n* Median: The median marks the mid-point of the data and is shown by the line that divides the box into two parts (sometimes known as the second quartile). Half the scores are greater than or equal to this value and half are less.\n* Upper Quartile: Seventy-five percent of the scores fall below the upper quartile value (also known as the third quartile). Thus, 25% of data are above this value.\n* Maximum Score: The highest score, excluding outliers (shown at the end of the right whisker).\n\nSome other summary are:\n* Whiskers: The upper and lower whiskers represent scores outside the middle 50% (i.e. the lower 25% of scores and the upper 25% of scores).\n* The Interquartile Range (or IQR): This is the box plot showing the middle 50% of scores (i.e., the range between the 25th and 75th percentile).","metadata":{}},{"cell_type":"code","source":"# Truncate some extreme values for better visuals\ntrain_df['num_words'].loc[train_df['num_words']>60] = 60 \ntrain_df['num_punctuations'].loc[train_df['num_punctuations']>10] = 10 \ntrain_df['num_chars'].loc[train_df['num_chars']>350] = 350\n\n# Box plot making\nf, axes = plt.subplots(3, 1, figsize=(10,20))\nsns.boxplot(x='target', y='num_words', data=train_df, ax=axes[0])\naxes[0].set_xlabel('Target', fontsize=12)\naxes[0].set_title(\"Number of words in each class\", fontsize=15)\n\nsns.boxplot(x='target', y='num_chars', data=train_df, ax=axes[1])\naxes[1].set_xlabel('Target', fontsize=12)\naxes[1].set_title(\"Number of characters in each class\", fontsize=15)\n\nsns.boxplot(x='target', y='num_punctuations', data=train_df, ax=axes[2])\naxes[2].set_xlabel('Target', fontsize=12)\n#plt.ylabel('Number of punctuations in text', fontsize=12)\naxes[2].set_title(\"Number of punctuations in each class\", fontsize=15)\nplt.show()","metadata":{"_uuid":"1b000ed2-9c9f-412e-a019-402becb87bc6","_cell_guid":"ac64fbc4-25f5-445f-b97f-3a3cf9ea8d2d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:14:00.215149Z","iopub.execute_input":"2021-06-10T20:14:00.215703Z","iopub.status.idle":"2021-06-10T20:14:00.679788Z","shell.execute_reply.started":"2021-06-10T20:14:00.215657Z","shell.execute_reply":"2021-06-10T20:14:00.678298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Observations from Meta Features*** \n\nWe can see that the insincere questions have more number of words as well as characters compared to sincere questions. So this might be a useful feature in our model.","metadata":{}},{"cell_type":"markdown","source":"## Detailed Statistics for given data","metadata":{}},{"cell_type":"markdown","source":"```textstat``` package is used for this purpose. \n\nBefore any further analyze, all questions must be tokenized. I used spaCy tokenizer to perform the task, following with ```tqdm``` to visualize progress bars.","metadata":{}},{"cell_type":"code","source":"def spacy_tokenizer(sentence):\n    mytokens = parser(sentence)\n    mytokens = [ word.lemma_.lower().strip() if word.lemma_ != \"-PRON-\" else word.lower_ for word in mytokens ]\n    mytokens = [ word for word in mytokens if word not in stopwords and word not in punctuations ]\n    mytokens = \" \".join([i for i in mytokens])\n    return mytokens\n\ntqdm.pandas()\nsincere_questions = train_sincere_df[\"question_text\"].progress_apply(spacy_tokenizer)\ninsincere_questions = train_insincere_df[\"question_text\"].progress_apply(spacy_tokenizer)","metadata":{"_uuid":"602677b2-c1cb-4beb-b5d9-73aa68f186f7","_cell_guid":"61315703-0e55-4b56-96f2-04cd6e1a6e07","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.724961Z","iopub.status.idle":"2021-06-10T20:13:47.725397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To illustrate some popular summarized figures, I created a table data for all plots (which will be described below). \n\nThe figures used are provided by ```math``` module: Mean, Standard Deviation, Variance, Median, Max, Min.","metadata":{}},{"cell_type":"code","source":"# One function for all plots\ndef plot_readability(a,b,title,bins=0.1,colors=['#0000FF', '#FF0000']):\n    trace1 = ff.create_distplot([a,b], [\"Sincere questions\",\"Insincere questions\"], bin_size=bins, colors=colors, show_rug=False)\n    trace1['layout'].update(title=title)\n    iplot(trace1, filename='Distplot')\n    table_data= [[\"Statistical Measures\",\"Sincere questions\",\"Insincere questions\"],\n                [\"Mean\",mean(a),mean(b)],\n                [\"Standard Deviation\",pstdev(a),pstdev(b)],\n                [\"Variance\",pvariance(a),pvariance(b)],\n                [\"Median\",median(a),median(b)],\n                [\"Maximum value\",max(a),max(b)],\n                [\"Minimum value\",min(a),min(b)]]\n    trace2 = ff.create_table(table_data)\n    iplot(trace2, filename='Table')","metadata":{"_uuid":"4140c484-1af0-4cd6-ba1d-a168a1bad4d9","_cell_guid":"9b682caa-502a-4356-8286-2d1960e601d5","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.726577Z","iopub.status.idle":"2021-06-10T20:13:47.727152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Syllable Analysis","metadata":{}},{"cell_type":"code","source":"syllable_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.syllable_count))\nsyllable_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.syllable_count))\nplot_readability(syllable_sincere,syllable_insincere,\"Syllable Analysis\",5)","metadata":{"_uuid":"a0504b95-cd98-41a9-925a-2ab821b7ffa6","_cell_guid":"5522b476-1951-4748-a48d-6397eb1f69e8","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.728628Z","iopub.status.idle":"2021-06-10T20:13:47.729223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Lexicon Analysis","metadata":{}},{"cell_type":"code","source":"lexicon_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.lexicon_count))\nlexicon_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.lexicon_count))\nplot_readability(lexicon_sincere,lexicon_insincere,\"Lexicon Analysis\",4)","metadata":{"_uuid":"de25541e-de39-44c8-bdf9-85b7e25d3c1a","_cell_guid":"d5d0cea5-eb67-47c7-9a10-69a9ba47b2db","execution":{"iopub.status.busy":"2021-06-10T20:13:47.730371Z","iopub.status.idle":"2021-06-10T20:13:47.730995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Question Length","metadata":{}},{"cell_type":"code","source":"length_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(len))\nlength_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(len))\nplot_readability(length_sincere,length_insincere,\"Question Length\",40)","metadata":{"_uuid":"bbd0f798-cfbe-4a20-b17d-9bd3029e77cf","_cell_guid":"8616dbe6-ab04-41b1-a963-5159f6380988","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.732451Z","iopub.status.idle":"2021-06-10T20:13:47.733037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Average Syllables per Word","metadata":{}},{"cell_type":"code","source":"spw_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.avg_syllables_per_word))\nspw_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.avg_syllables_per_word))\nplot_readability(spw_sincere,spw_insincere,\"Average syllables per word\",0.2)","metadata":{"_uuid":"0f1dda3c-f810-4cc1-ad73-dfc9a0bbc27a","_cell_guid":"287f0775-5eff-4ab0-9b72-39cd1fe62eaa","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.734389Z","iopub.status.idle":"2021-06-10T20:13:47.734981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Average Letter per Word","metadata":{}},{"cell_type":"code","source":"lpw_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.avg_letter_per_word))\nlpw_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.avg_letter_per_word))\nplot_readability(lpw_sincere,lpw_insincere,\"Average letters per word\",2)","metadata":{"_uuid":"f7d0e9e0-b0ef-4bcd-ab97-de7c9e818ddf","_cell_guid":"1d273ea3-8823-4f8e-9b7e-63255c5402f4","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.736061Z","iopub.status.idle":"2021-06-10T20:13:47.736740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Readability Features","metadata":{}},{"cell_type":"markdown","source":"**Flesch reading ease**\n\nIn the Flesch reading-ease test, higher scores indicate material that is easier to read; lower numbers mark passages that are more difficult to read. The formula for the Flesch reading-ease score (FRES) test is:\n\n<img src=\"https://latex.codecogs.com/svg.image?206.835&space;-&space;1.015\\left(\\frac{total\\;words}{total\\;sentences}\\right)&space;-&space;84.6\\left(\\frac{total\\;syllables}{total\\;words}\\right)\" title=\"206.835 - 1.015\\left(\\frac{total\\;words}{total\\;sentences}\\right) - 84.6\\left(\\frac{total\\;syllables}{total\\;words}\\right)\" />\n\nScores can be interpreted as shown in the table below.\n\n|     Score    |  School level (US) |                                  Notes                                  |\n|:------------:|:------------------:|:-----------------------------------------------------------------------:|\n| 100.00–90.00 | 5th grade          | Very easy to read. Easily understood by an average 11-year-old student. |\n| 90.0–80.0    | 6th grade          | Easy to read. Conversational English for consumers.                     |\n| 80.0–70.0    | 7th grade          | Fairly easy to read.                                                    |\n| 70.0–60.0    | 8th & 9th grade    | Plain English. Easily understood by 13- to 15-year-old students.        |\n| 60.0–50.0    | 10th to 12th grade | Fairly difficult to read.                                               |\n| 50.0–30.0    | College            | Difficult to read.                                                      |\n| 30.0–10.0    | College graduate   | Very difficult to read. Best understood by university graduates.        |\n| 10.0–0.0     | Professional       | Extremely difficult to read. Best understood by university graduates.   |","metadata":{}},{"cell_type":"code","source":"fre_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.flesch_reading_ease))\nfre_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.flesch_reading_ease))\nplot_readability(fre_sincere,fre_insincere,\"Flesch Reading Ease\",20)","metadata":{"_uuid":"cb236d17-c278-4683-b419-888cece214c7","_cell_guid":"f18f89ff-cbcb-4c4f-93c9-0dc25e20dee9","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.737823Z","iopub.status.idle":"2021-06-10T20:13:47.738460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generally speaking, questions given in the dataset are fairly easy to read (mean value is ~70-75).","metadata":{}},{"cell_type":"markdown","source":"**The Flesch-Kincaid Grade Level**\n\nThe \"Flesch–Kincaid Grade Level Formula\" presents a score as a U.S. grade level, making it easier to judge the readability level of texts. It can also mean the number of years of education generally required to understand this text, relevant when the formula results in a number greater than 10. The grade level is calculated with the following formula:\n\n<img src=\"https://latex.codecogs.com/svg.image?0.39&space;-&space;11.8\\left(\\frac{total\\;words}{total\\;sentences}\\right)&space;-&space;15.59\\left(\\frac{total\\;syllables}{total\\;words}\\right)\" title=\"0.39 - 11.8\\left(\\frac{total\\;words}{total\\;sentences}\\right) - 15.59\\left(\\frac{total\\;syllables}{total\\;words}\\right)\" />\n\nThe result is a number that corresponds with a U.S. grade level. ","metadata":{}},{"cell_type":"code","source":"fkg_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.flesch_kincaid_grade))\nfkg_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.flesch_kincaid_grade))\nplot_readability(fkg_sincere,fkg_insincere,\"Flesch Kincaid Grade\",4)","metadata":{"_uuid":"900ab858-c448-406f-aa11-1687ea153695","_cell_guid":"3928c844-b810-441e-b153-942686ee8b40","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.740151Z","iopub.status.idle":"2021-06-10T20:13:47.740511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The Fog Scale (Gunning FOG Formula)**\n\nThe Gunning fog index is a readability test for English writing. The index estimates the years of formal education a person needs to understand the text on the first reading. \n\nThe Gunning fog index is calculated with the following formula:\n\n<img src=\"https://latex.codecogs.com/svg.image?0.4\\left&space;[&space;\\left&space;(&space;\\frac{words}{sentences}&space;\\right&space;)&plus;&space;100&space;&space;\\left&space;(&space;\\frac{complex\\;words}{words}&space;\\right)&space;\\right]\" title=\"0.4\\left [ \\left ( \\frac{words}{sentences} \\right )+ 100 \\left ( \\frac{complex\\;words}{words} \\right) \\right]\" />\n\nThe fog index is commonly used to confirm that text can be read easily by the intended audience. Texts for a wide audience generally need a fog index less than 12. Texts requiring near-universal understanding generally need an index less than 8.","metadata":{}},{"cell_type":"code","source":"fog_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.gunning_fog))\nfog_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.gunning_fog))\nplot_readability(fog_sincere,fog_insincere,\"The Fog Scale (Gunning FOG Formula)\",4)","metadata":{"_uuid":"979ecc9d-394b-416c-836c-61d1b0e4974e","_cell_guid":"82a77e34-356f-4959-8bf7-c33b72decffe","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.742221Z","iopub.status.idle":"2021-06-10T20:13:47.742637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Automated Readability Index**\n\nThe automated readability index (ARI) is a readability test for English texts, designed to gauge the understandability of a text. Like the Flesch–Kincaid grade level, Gunning fog index, SMOG index, Fry readability formula, and Coleman–Liau index, it produces an approximate representation of the US grade level needed to comprehend the text.\n\nThe formula for calculating the automated readability index is given below:\n\n<img src=\"https://latex.codecogs.com/svg.image?4.71\\left(\\frac{character}{words}\\right)&space;+&space;0.5\\left(\\frac{words}{sentences}\\right)&space;-&space;21.43\" title=\"4.71\\left(\\frac{character}{words}\\right) + 0.5\\left(\\frac{words}{sentences}\\right) - 21.43\" />\n\n| Score | Age   | Grade Level        |\n|-------|-------|--------------------|\n| 1     | 5-6   | Kindergarten       |\n| 2     | 6-7   | First/Second Grade |\n| 3     | 7-9   | Third Grade        |\n| 4     | 9-10  | Fourth Grade       |\n| 5     | 10-11 | Fifth Grade        |\n| 6     | 11-12 | Sixth Grade        |\n| 7     | 12-13 | Seventh Grade      |\n| 8     | 13-14 | Eighth Grade       |\n| 9     | 14-15 | Ninth Grade        |\n| 10    | 15-16 | Tenth Grade        |\n| 11    | 16-17 | Eleventh Grade     |\n| 12    | 17-18 | Twelfth grade      |\n| 13    | 18-24 | College student    |\n| 14    | 24+   | Professor          |","metadata":{}},{"cell_type":"code","source":"ari_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.automated_readability_index))\nari_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.automated_readability_index))\nplot_readability(ari_sincere,ari_insincere,\"Automated Readability Index\",10)","metadata":{"_uuid":"089fcb49-3579-4006-97d8-f3cd70b32905","_cell_guid":"d22b85d9-14bb-45ed-9dc4-c4ac99d660b1","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.743700Z","iopub.status.idle":"2021-06-10T20:13:47.744049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The Coleman-Liau Index**\n\nThe Coleman–Liau index is a readability test designed by Meri Coleman and T. L. Liau to gauge the understandability of a text. Its output approximates the U.S. grade level thought necessary to comprehend the text.\n\nLike the ARI but unlike most of the other indices, Coleman–Liau relies on characters instead of syllables per word.\n\nThe Coleman–Liau index is calculated with the following formula:\n\n<img src=\"https://latex.codecogs.com/svg.image?CLI&space;=&space;0.0588L&space;-&space;0.296S&space;-&space;15.8\" title=\"CLI = 0.0588L - 0.296S - 15.8\" />","metadata":{}},{"cell_type":"code","source":"cli_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.coleman_liau_index))\ncli_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.coleman_liau_index))\nplot_readability(cli_sincere,cli_insincere,\"The Coleman-Liau Index\",10)","metadata":{"_uuid":"37957f50-c221-4b13-bda6-2115d1fcebad","_cell_guid":"256f5eaf-382c-4358-bc93-0413f2460964","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.745101Z","iopub.status.idle":"2021-06-10T20:13:47.745483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Linsear Write Formula**\n\nThe standard Linsear Write metric Lw runs on a 100-word sample:\n1. For each \"easy word\", defined as words with 2 syllables or less, add 1 point.\n2. For each \"hard word\", defined as words with 3 syllables or more, add 3 points.\n3. Divide the points by the number of sentences in the 100-word sample.\n4. Adjust the provisional result r:\n* If r > 20, Lw = r / 2.\n* If r ≤ 20, Lw = r / 2 - 1.\nThe result is a \"grade level\" measure, reflecting the estimated years of education needed to read the text fluently","metadata":{}},{"cell_type":"code","source":"lwf_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.linsear_write_formula))\nlwf_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.linsear_write_formula))\nplot_readability(lwf_sincere,lwf_insincere,\"Linsear Write Formula\",2)","metadata":{"_uuid":"ac31ac67-842b-4320-8795-bd11e19afb67","_cell_guid":"ba0b6025-f6e9-4e18-b5d8-ad5d930d9e6e","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.746524Z","iopub.status.idle":"2021-06-10T20:13:47.747078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dale–Chall Readability Formula**\n\nThe Dale–Chall readability formula is a readability test that provides a numeric gauge of the comprehension difficulty that readers come upon when reading a text. It uses a list of 3000 words that groups of fourth-grade American students could reliably understand, considering any word not on that list to be difficult.\n\nThe formula for calculating the raw score of the Dale–Chall readability score (1948) is given below:\n\n<img src=\"https://latex.codecogs.com/svg.image?0.1579\\left&space;(&space;&space;\\frac{difficult\\;words}{words}\\ast&space;&space;100\\right&space;)&space;&plus;&space;0.0496\\left&space;(&space;\\frac{words}{sentences}&space;\\right&space;)\" title=\"0.1579\\left ( \\frac{difficult\\;words}{words}\\ast 100\\right ) + 0.0496\\left ( \\frac{words}{sentences} \\right )\" />\n\nIf the percentage of difficult words is above 5%, then add 3.6365 to the raw score to get the adjusted score, otherwise the adjusted score is equal to the raw score.\n\n|     Score    |                                 Notes                                |\n|:------------:|:--------------------------------------------------------------------:|\n| 4.9 or lower | easily understood by an average 4th-grade student or lower           |\n| 5.0–5.9      | easily understood by an average 5th or 6th-grade student             |\n| 6.0–6.9      | easily understood by an average 7th or 8th-grade student             |\n| 7.0–7.9      | easily understood by an average 9th or 10th-grade student            |\n| 8.0–8.9      | easily understood by an average 11th or 12th-grade student           |\n| 9.0–9.9      | easily understood by an average 13th to 15th-grade (college) student |","metadata":{}},{"cell_type":"code","source":"dcr_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(textstat.dale_chall_readability_score))\ndcr_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(textstat.dale_chall_readability_score))\nplot_readability(dcr_sincere,dcr_insincere,\"Dale-Chall Readability Score\",1)","metadata":{"_uuid":"42ce7786-ee1a-4735-b24a-5fb4d9d42d97","_cell_guid":"f705b889-0fe7-4de4-a4a9-d69a67ca4fca","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.747924Z","iopub.status.idle":"2021-06-10T20:13:47.748460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Readability Consensus based upon all the above tests**\n\nThe estimated school grade level required to understand the text based on all above tests.","metadata":{}},{"cell_type":"code","source":"def consensus_all(text):\n    return textstat.text_standard(text,float_output=True)\n\ncon_sincere = np.array(train_sincere_df[\"question_text\"].progress_apply(consensus_all))\ncon_insincere = np.array(train_insincere_df[\"question_text\"].progress_apply(consensus_all))\nplot_readability(con_sincere,con_insincere,\"Readability Consensus based upon all the above tests\",2)","metadata":{"_uuid":"11e916de-80ed-4219-af5f-274aff42d165","_cell_guid":"f2a4a488-2a98-4ded-aa0e-e85261c1e5cf","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-06-10T20:13:47.749305Z","iopub.status.idle":"2021-06-10T20:13:47.749822Z"},"trusted":true},"execution_count":null,"outputs":[]}]}