{"cells":[{"metadata":{"_uuid":"a8f9622945156d6337ba73c481da2de7efef7384"},"cell_type":"markdown","source":"# <div style=\"text-align: center\">A Data Science Framework for Quora </div>\n### <div align=\"center\"><b>Quite Practical and Far from any Theoretical Concepts</b></div>\n<img src='http://s9.picofile.com/file/8342477368/kq.png'>\n<div style=\"text-align:center\">last update: <b>19/01/2019</b></div>\n\n"},{"metadata":{"_uuid":"750903cc2679d39058f56df6c6c040be02b748df"},"cell_type":"markdown","source":" <a id=\"1\"></a> <br>\n## 1- Introduction\n<font color=\"red\">Quora</font> has defined a competition in **Kaggle**. A realistic and attractive data set for data scientists.\non this notebook, I will provide a **comprehensive** approach to solve Quora classification problem for **beginners**.\n\nI am open to getting your feedback for improving this **kernel**."},{"metadata":{"_uuid":"cda11210a88d6484112cbe2c3624225328326c6a"},"cell_type":"markdown","source":"<a id=\"top\"></a> <br>\n## Notebook  Content\n1. [Introduction](#1)\n1. [Data Science Workflow for Quora](#2)\n1. [Problem Definition](#3)\n    1. [Business View](#31)\n        1. [Real world Application Vs Competitions](#311)\n    1. [What is a insincere question?](#32)\n    1. [How can we find insincere question?](#33)\n1. [Problem feature](#4)\n    1. [Aim](#41)\n    1. [Variables](#42)\n    1. [ Inputs & Outputs](#43)\n1. [Select Framework](#5)\n    1. [Import](#51)\n    1. [Version](#52)\n    1. [Setup](#53)\n1. [Exploratory data analysis](#6)\n    1. [Data Collection](#61)\n        1. [Features](#611)\n        1. [Explorer Dataset](#612)\n    1. [Data Cleaning](#62)\n    1. [Data Preprocessing](#63)\n        1. [Is data set imbalance?](#631)\n        1. [Some Feature Engineering](#632)\n    1. [Data Visualization](#64)\n        1. [countplot](#641)\n        1. [pie plot](#642)\n        1. [Histogram](#643)\n        1. [violin plot](#645)\n        1. [kdeplot](#646)\n1. [Apply Learning](#7)\n1. [Conclusion](#8)\n1. [References](#9)"},{"metadata":{"_uuid":"e11b73b618b0f6e4335520ef80267c6d577d1ba5"},"cell_type":"markdown","source":"<a id=\"2\"></a> <br>\n## 2- A Data Science Workflow for Quora\nOf course, the same solution can not be provided for all problems, so the best way is to create a **general framework** and adapt it to new problem.\n\n**You can see my workflow in the below image** :\n\n <img src=\"http://s8.picofile.com/file/8342707700/workflow2.png\"  />\n\n**You should feel free\tto\tadjust \tthis\tchecklist \tto\tyour needs**\n###### [Go to top](#top)"},{"metadata":{"_uuid":"600be852c0d28e7c0c5ebb718904ab15a536342c"},"cell_type":"markdown","source":"<a id=\"3\"></a> <br>\n## 3- Problem Definition\nI think one of the important things when you start a new machine learning project is Defining your problem. that means you should understand business problem.( **Problem Formalization**)\n> **we will be predicting whether a question asked on Quora is sincere or not.**\n<a id=\"31\"></a> <br>\n## 3-1 About Quora\nQuora is a platform that empowers people to learn from each other. On Quora, people can ask questions and connect with others who contribute unique insights and quality answers. A key challenge is to weed out insincere questions -- those founded upon false premises, or that intend to make a statement rather than look for helpful answers.\n<a id=\"32\"></a> <br>\n## 3-2 Business View \nAn existential problem for any major website today is how to handle toxic and divisive content. **Quora** wants to tackle this problem head-on to keep their platform a place where users can feel safe sharing their knowledge with the world.\n\n**Quora** is a platform that empowers people to learn from each other. On Quora, people can ask questions and connect with others who contribute unique insights and quality answers. A key challenge is to weed out insincere questions -- those founded upon false premises, or that intend to make a statement rather than look for helpful answers.\n\nIn this kernel, I will develop models that identify and flag insincere questions.we Help Quora uphold their policy of “Be Nice, Be Respectful” and continue to be a place for sharing and growing the world’s knowledge.\n<a id=\"321\"></a> <br>\n### 3-2-1 Real world Application Vs Competitions\nJust a simple comparison between real-world apps with competitions:\n<img src=\"http://s9.picofile.com/file/8339956300/reallife.png\" height=\"600\" width=\"500\" />\n<a id=\"33\"></a> <br>\n## 3-3 What is a insincere question?\nIs defined as a question intended to make a **statement** rather than look for **helpful answers**.\n<img src='http://s8.picofile.com/file/8342711526/Quora_moderation.png'>\n<a id=\"34\"></a> <br>\n## 3-4 How can we find insincere question?\nSome characteristics that can signify that a question is insincere:\n\n1. **Has a non-neutral tone**\n    1. Has an exaggerated tone to underscore a point about a group of people\n    1. Is rhetorical and meant to imply a statement about a group of people\n1. **Is disparaging or inflammatory**\n    1. Suggests a discriminatory idea against a protected class of people, or seeks confirmation of a stereotype\n    1. Makes disparaging attacks/insults against a specific person or group of people\n    1. Based on an outlandish premise about a group of people\n    1. Disparages against a characteristic that is not fixable and not measurable\n1. **Isn't grounded in reality**\n    1. Based on false information, or contains absurd assumptions\n    1. Uses sexual content (incest, bestiality, pedophilia) for shock value, and not to seek genuine answers\n    ###### [Go to top](#top)"},{"metadata":{"_uuid":"556980c672d2f7b2a4ee943b9d13b88de6e41e04"},"cell_type":"markdown","source":"<a id=\"4\"></a> <br>\n## 4- Problem Feature\nProblem Definition has three steps that have illustrated in the picture below:\n\n1. Aim\n1. Variable\n1. Inputs & Outputs\n\n\n\n\n\n<a id=\"41\"></a> <br>\n### 4-1 Aim\nWe will be predicting whether a question asked on Quora is **sincere** or not.\n\n\n<a id=\"42\"></a> <br>\n### 4-2 Variables\n\n1. qid - unique question identifier\n1. question_text - Quora question text\n1. target - a question labeled \"insincere\" has a value of 1, otherwise 0\n\n<a id=\"43\"></a> <br>\n### 4-3 Inputs & Outputs\nwe use train.csv and test.csv as Input and we should upload a  submission.csv as Output\n\n\n**<< Note >>**\n> You must answer the following question:\nHow does your company expect to use and benefit from **your model**.\n###### [Go to top](#top)"},{"metadata":{"_uuid":"fbedcae8843986c2139f18dad4b5f313e6535ac5"},"cell_type":"markdown","source":"<a id=\"5\"></a> <br>\n## 5- Select Framework\nAfter problem definition and problem feature, we should select our framework to solve the problem.\nWhat we mean by the framework is that  the programming languages you use and by what modules the problem will be solved.\n###### [Go to top](#top)"},{"metadata":{"_uuid":"c90e261f3b150e10aaec1f34ab3be768acf7aa25"},"cell_type":"markdown","source":"<a id=\"52\"></a> <br>\n## 5-2 Import"},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import accuracy_score\nfrom wordcloud import WordCloud as wc\nfrom nltk.corpus import stopwords\nimport matplotlib.pylab as pylab\nimport matplotlib.pyplot as plt\nfrom pandas import get_dummies\nimport matplotlib as mpl\nimport seaborn as sns\nimport pandas as pd\nimport numpy as np\nimport matplotlib\nimport warnings\nimport sklearn\nimport string\nimport scipy\nimport numpy\nimport nltk\nimport json\nimport sys\nimport csv\nimport os","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1c2beac253f7ddddcc2e1aa26dc850d5b87268f3"},"cell_type":"markdown","source":"<a id=\"53\"></a> <br>\n## 5-3 version"},{"metadata":{"_kg_hide-input":true,"_uuid":"9ffe2f1e5995150c8138f9e98509c7525fb230b4","trusted":true},"cell_type":"code","source":"print('matplotlib: {}'.format(matplotlib.__version__))\nprint('sklearn: {}'.format(sklearn.__version__))\nprint('scipy: {}'.format(scipy.__version__))\nprint('seaborn: {}'.format(sns.__version__))\nprint('pandas: {}'.format(pd.__version__))\nprint('numpy: {}'.format(np.__version__))\nprint('Python: {}'.format(sys.version))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"431bf889ae401c1089a13835356c13f2b6a06f6c"},"cell_type":"markdown","source":"<a id=\"54\"></a> <br>\n## 5-4 Setup\n\nA few tiny adjustments for better **code readability**"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"8645feedee1145c2df1268a697b0b8773858ad1a"},"cell_type":"code","source":"sns.set(style='white', context='notebook', palette='deep')\npylab.rcParams['figure.figsize'] = 12,8\nwarnings.filterwarnings('ignore')\nmpl.style.use('ggplot')\nsns.set_style('white')\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c5048b61a4837c8826551c8871609973ebbe3847"},"cell_type":"markdown","source":"<a id=\"55\"></a> <br>\n## 5-5 NLTK\nIn this kernel, we use the NLTK library So, before we begin the next step, we will first introduce this library.\n<img src='https://arts.unimelb.edu.au/__data/assets/image/0005/2735348/nltk.jpg' width=300 height=300>"},{"metadata":{"_kg_hide-input":true,"_uuid":"adadeb7a83d0bc711a779948197c40841b10f1ca","trusted":true},"cell_type":"code","source":"from nltk.tokenize import sent_tokenize, word_tokenize\n \ndata = \"All work and no play makes jack a dull boy, all work and no play\"\nprint(word_tokenize(data))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"04ff1a533119d589baee777c21194a951168b0c7"},"cell_type":"markdown","source":"<a id=\"6\"></a> <br>\n## 6- EDA\nBy the end of the section, you'll be able to answer these questions and more, while generating graphics that are both insightful and beautiful.  then We will review analytical and statistical operations:\n\n1. Data Collection\n1. Visualization\n1. Data Cleaning\n1. Data Preprocessing\n<img src=\"http://s9.picofile.com/file/8338476134/EDA.png\" width=400 height=400>\n\n ###### [Go to top](#top)"},{"metadata":{"_uuid":"cedecea930b278f86292367cc28d2996a235a169"},"cell_type":"markdown","source":"<a id=\"61\"></a> <br>\n## 6-1 Data Collection\nI start Collection Data by the training and testing datasets into **Pandas DataFrames**.\n###### [Go to top](#top)"},{"metadata":{"_kg_hide-input":true,"_uuid":"9269ae851b744856bce56840637030a16a5877e1","trusted":true},"cell_type":"code","source":"train = pd.read_csv('../input/train.csv')\ntest = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"58ed9c838069f54de5cf90b20a774c3e236149b3"},"cell_type":"markdown","source":"**<< Note 1 >>**\n\n* Each **row** is an observation (also known as : sample, example, instance, record).\n* Each **column** is a feature (also known as: Predictor, attribute, Independent Variable, input, regressor, Covariate).\n###### [Go to top](#top)"},{"metadata":{"_uuid":"4708d70e39d1ae861bbf34411cf03d07f261fceb","trusted":true},"cell_type":"code","source":"train.sample(1) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f8e7a84ab982504d7263b1812fa66bba78bddbdc","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"test.sample(1) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3483fbc1e932d9f387703a796248963e77cefa1d"},"cell_type":"markdown","source":"Or you can use others command to explorer dataset, such as "},{"metadata":{"_uuid":"08a94b16129d4c231b64d4691374e18aa80f1d80","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train.tail(1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"581b90e6a869c3793472c7edd59091d6d6342fb2"},"cell_type":"markdown","source":"<a id=\"611\"></a> <br>\n## 6-1-1 Features\nFeatures can be from following types:\n* numeric\n* categorical\n* ordinal\n* datetime\n* coordinates\n\nFind the type of features in **Qoura dataset**?!\n\nFor getting some information about the dataset you can use **info()** command."},{"metadata":{"_kg_hide-input":true,"_uuid":"ca840f02925751186f87e402fcb5f637ab1ab8a0","trusted":true},"cell_type":"code","source":"print(train.info())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4cbcf76344a6e3c8e841ccf1f43bf00d040a06a1","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(test.info())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73ab30f86273b590a51fc363d9bf78c2709558fa"},"cell_type":"markdown","source":"<a id=\"612\"></a> <br>\n## 6-1-2 Explorer Dataset\n\n###### [Go to top](#top)"},{"metadata":{"_kg_hide-input":true,"_uuid":"4b45251be7be77333051fe738639104ae1005fa5","trusted":true},"cell_type":"code","source":"# shape for train and test\nprint('Shape of train:',train.shape)\nprint('Shape of test:',test.shape)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"c64e9d3e0bf394fb833de94a0fc5c34f69fce24c","trusted":true},"cell_type":"code","source":"#columns*rows\ntrain.size","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7b5fd1034cd591ebd29fba1c77d342ec2b408d13"},"cell_type":"markdown","source":"After loading the data via **pandas**, we should checkout what the content is, description and via the following:"},{"metadata":{"_kg_hide-input":true,"_uuid":"edd043f8feb76cfe51b79785302ca4936ceb7b51","trusted":true},"cell_type":"code","source":"type(train)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"edd043f8feb76cfe51b79785302ca4936ceb7b51","trusted":true},"cell_type":"code","source":"type(test)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1b8b6f0c962a59e5258e74ed9e740a4aaf7c8113","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c288c3dc8656a872a8529368812546e434d3a22"},"cell_type":"markdown","source":"To pop up 5 random rows from the data set, we can use **sample(5)**  function and find the type of features."},{"metadata":{"_uuid":"09eb18d1fcf4a2b73ba2f5ddce99dfa521681140","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train.sample(5) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8280749a19af32869978c61941d1dea306632d71"},"cell_type":"markdown","source":"<a id=\"62\"></a> <br>\n## 6-2 Data Cleaning\n\n###### [Go to top](#top)"},{"metadata":{"_uuid":"a6315bf510cecb907b2d23aad25faf6ccad32ac4"},"cell_type":"markdown","source":"How many NA elements in every column!!\n\nGood news, it is Zero!\n\nTo check out how many null info are on the dataset, we can use **isnull().sum()**."},{"metadata":{"_kg_hide-input":true,"_uuid":"675f72fb58d83c527f71819e71ed8e17f81126f5","trusted":true},"cell_type":"code","source":"train.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5faa6528c6667060c05268757ff46e211b4fea3f"},"cell_type":"markdown","source":"But if we had , we can just use **dropna()**(be careful sometimes you should not do this!)"},{"metadata":{"_kg_hide-input":true,"_uuid":"e8e124ca20643ad307d9bfdc34328d548c6ddcbc","trusted":true},"cell_type":"code","source":"# remove rows that have NA's\nprint('Before Droping',train.shape)\ntrain = train.dropna()\nprint('After Droping',train.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"277e1998627d6a3ddeff4e913a6b8c3dc81dec96"},"cell_type":"markdown","source":"\nWe can get a quick idea of how many instances (rows) and how many attributes (columns) the data contains with the shape property."},{"metadata":{"_uuid":"c2f1eaf0b6dfdc7cc4dace04614e99ed56425d00"},"cell_type":"markdown","source":"To print dataset **columns**, we can use columns atribute."},{"metadata":{"_uuid":"909d61b33ec06249d0842e6115597bbacf21163f","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train.columns","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3458838205be4c7fbff88e95ef69934e13e2199b"},"cell_type":"markdown","source":"You see number of unique item for Target  with command below:"},{"metadata":{"_kg_hide-input":true,"_uuid":"c7937700664991b29bdb0b3f04942c59498da760","trusted":true},"cell_type":"code","source":"train_target = train['target'].values\n\nnp.unique(train_target)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d824cb29e135dc5ae98964e71ec0adc0e05ebd43"},"cell_type":"markdown","source":"YES, quora problem is a **binary classification**! :)"},{"metadata":{"_uuid":"ae08b544a8d4202c7d0a47ec83d685e81c91a66d"},"cell_type":"markdown","source":"To check the first 5 rows of the data set, we can use head(5)."},{"metadata":{"_kg_hide-input":true,"_uuid":"5899889553c3416b27e93efceddb106eb71f5156","trusted":true},"cell_type":"code","source":"train.head(5) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1150b6ac3d82562aefd5c64f9f01accee5eace4d"},"cell_type":"markdown","source":"Or to check out last 5 row of the data set, we use tail() function."},{"metadata":{"_uuid":"79339442ff1f53ae1054d794337b9541295d3305","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train.tail() ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c8a1cc36348c68fb98d6cb28aa9919fc5f2892f3"},"cell_type":"markdown","source":"To give a **statistical summary** about the dataset, we can use **describe()**\n"},{"metadata":{"_uuid":"3f7211e96627b9a81c5b620a9ba61446f7719ea3","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train.describe() ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"10bdb8246f66c14043392806cae714f688cc8251"},"cell_type":"markdown","source":"As you can see, the statistical information that this command gives us is not suitable for this type of data\n**describe() is more useful for numerical data sets**"},{"metadata":{"_uuid":"6c8c838f497c66a227975fb9a2f588e431f0c568"},"cell_type":"markdown","source":"**<< Note 2 >>**\nin pandas's data frame you can perform some query such as \"where\""},{"metadata":{"_uuid":"c8c8d9fd63d9bdb601183aeb4f1435affeb8a596","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train.where(train ['target']==1).count()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"33fc33a18489b438a884819d99dc00a02b113be8"},"cell_type":"markdown","source":"As you can see in the below in python, it is so easy perform some query on the dataframe:"},{"metadata":{"_uuid":"8b545ff7e8367c5ab9c1db710f70b6936ac8422c","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train[train['target']>1]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2788d023986eca622f7db9e1d64c2a4e02737ddb"},"cell_type":"markdown","source":"Some examples of questions that they are insincere"},{"metadata":{"_uuid":"d517b2b99a455a6b89c238faf1647515b8a67d87","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train[train['target']==1].head(5)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4b67d109a0cec1a5475b863bbce8aa3ac9d2d4fb"},"cell_type":"markdown","source":"<a id=\"631\"></a> <br>\n## 6-3-1 Is data set imbalance?\n"},{"metadata":{"_uuid":"4218d492753322c50142021833efb24cfdfc6ad3","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_target.mean()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8e058f90ca403f00d91d0405a7d8822dc7d6de55"},"cell_type":"markdown","source":"A large part of the data is unbalanced, but **how can we  solve it?**"},{"metadata":{"_kg_hide-input":true,"_uuid":"dc6340ee1b637d192e29cbc8d3744ae6351b9c8b","trusted":true},"cell_type":"code","source":"train[\"target\"].value_counts()\n# data is imbalance","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"873b2cfdee04b8ba087df1c4bf01ae69ef2f1c52"},"cell_type":"markdown","source":"<a id=\"632\"></a> <br>\n## 6-3-2 Exploreing Question"},{"metadata":{"_kg_hide-input":true,"_uuid":"e445c859d7c43857cfbf370ff20060a5341d3c89","trusted":true},"cell_type":"code","source":"question = train['question_text']\ni=0\nfor q in question[:5]:\n    i=i+1\n    print('sample '+str(i)+':' ,q)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"fa78f61df85a9c76bd092dbf6d6bcec4b6b2631f","trusted":true},"cell_type":"code","source":"text_withnumber = train['question_text']\nresult = ''.join([i for i in text_withnumber if not i.isdigit()])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c50c6c2683c5c08a6c9c34b75be61567d5993fa0"},"cell_type":"markdown","source":"<a id=\"632\"></a> <br>\n## 6-3-2 Some Feature Engineering"},{"metadata":{"_uuid":"c641754f26e07c368596af3054268f1b3b764921"},"cell_type":"markdown","source":"[NLTK](https://www.nltk.org/) is one of the leading platforms for working with human language data and Python, the module NLTK is used for natural language processing. NLTK is literally an acronym for Natural Language Toolkit.\n\nWe get a set of **English stop** words using the line"},{"metadata":{"_kg_hide-input":true,"_uuid":"10ca7d56255b95fc774fff5adf7b4273ec7a1ea2","trusted":true},"cell_type":"code","source":"#from nltk.corpus import stopwords\neng_stopwords = set(stopwords.words(\"english\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3f5af107041b279ce723761f37f4ffebae2b22a3"},"cell_type":"markdown","source":"The returned list stopWords contains **179 stop words**  on my computer.\nYou can view the length or contents of this array with the lines:"},{"metadata":{"_kg_hide-input":true,"_uuid":"eca2d53bfae70c55b3b5b0e2c244826465cb478b","trusted":true},"cell_type":"code","source":"print(len(eng_stopwords))\nprint(eng_stopwords)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f049c6d9633200496ed97f8066257849b4824da"},"cell_type":"markdown","source":"The metafeatures that we'll create based on  SRK's  EDAs, [sudalairajkumar](http://http://www.kaggle.com/sudalairajkumar/simple-feature-engg-notebook-spooky-author) and [tunguz](https://www.kaggle.com/tunguz/just-some-simple-eda) are:\n1. Number of words in the text\n1. Number of unique words in the text\n1. Number of characters in the text\n1. Number of stopwords\n1. Number of punctuations\n1. Number of upper case words\n1. Number of title case words\n1. Average length of the words\n\n###### [Go to top](#top)"},{"metadata":{"_uuid":"f4982fc699bcb147513c247b9f4d86b02902eded"},"cell_type":"markdown","source":"Number of words in the text "},{"metadata":{"_kg_hide-input":true,"_uuid":"5b29fbd86ab48be6bd84fcac6fb6bca84d4b8792","trusted":true},"cell_type":"code","source":"train[\"num_words\"] = train[\"question_text\"].apply(lambda x: len(str(x).split()))\ntest[\"num_words\"] = test[\"question_text\"].apply(lambda x: len(str(x).split()))\nprint('maximum of num_words in train',train[\"num_words\"].max())\nprint('min of num_words in train',train[\"num_words\"].min())\nprint(\"maximum of  num_words in test\",test[\"num_words\"].max())\nprint('min of num_words in train',test[\"num_words\"].min())\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"83becd8affc2abab2252e065f77c80dc0dcf53be"},"cell_type":"markdown","source":"Number of unique words in the text"},{"metadata":{"_kg_hide-input":true,"_uuid":"72aebb943122982b891c959fa9fa36224adcb2fc","trusted":true},"cell_type":"code","source":"train[\"num_unique_words\"] = train[\"question_text\"].apply(lambda x: len(set(str(x).split())))\ntest[\"num_unique_words\"] = test[\"question_text\"].apply(lambda x: len(set(str(x).split())))\nprint('maximum of num_unique_words in train',train[\"num_unique_words\"].max())\nprint('mean of num_unique_words in train',train[\"num_unique_words\"].mean())\nprint(\"maximum of num_unique_words in test\",test[\"num_unique_words\"].max())\nprint('mean of num_unique_words in train',test[\"num_unique_words\"].mean())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cb2719fa417f2c3fabea9b6582081738ecdf678b"},"cell_type":"markdown","source":"Number of characters in the text "},{"metadata":{"_kg_hide-input":true,"_uuid":"a7029af9cfed9eb2e624d7177887e111a71054ff","trusted":true},"cell_type":"code","source":"\ntrain[\"num_chars\"] = train[\"question_text\"].apply(lambda x: len(str(x)))\ntest[\"num_chars\"] = test[\"question_text\"].apply(lambda x: len(str(x)))\nprint('maximum of num_chars in train',train[\"num_chars\"].max())\nprint(\"maximum of num_chars in test\",test[\"num_chars\"].max())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ddd289db2420b4f3fee7268fef94926688afd203"},"cell_type":"markdown","source":"Number of stopwords in the text"},{"metadata":{"_kg_hide-input":true,"_uuid":"086e229b918087420c33b57c7ad51d6723cf70f7","trusted":true},"cell_type":"code","source":"train[\"num_stopwords\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in eng_stopwords]))\ntest[\"num_stopwords\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in eng_stopwords]))\nprint('maximum of num_stopwords in train',train[\"num_stopwords\"].max())\nprint(\"maximum of num_stopwords in test\",test[\"num_stopwords\"].max())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"428a009f1a3b00b73ec7d6e8558aebc995e42594"},"cell_type":"markdown","source":"Number of punctuations in the text"},{"metadata":{"_kg_hide-input":true,"_uuid":"947abd63c51d74dc33c2891fb1e1b9381d9da23c","trusted":true},"cell_type":"code","source":"\ntrain[\"num_punctuations\"] =train['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\ntest[\"num_punctuations\"] =test['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\nprint('maximum of num_punctuations in train',train[\"num_punctuations\"].max())\nprint(\"maximum of num_punctuations in test\",test[\"num_punctuations\"].max())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a93e8fd32c2ff7dffd81f62ff3b6b3a5975d6836"},"cell_type":"markdown","source":"Number of title case words in the text"},{"metadata":{"_kg_hide-input":true,"_uuid":"82c95fcf5848ca383a6a84501fe74fef371392d1","trusted":true},"cell_type":"code","source":"\ntrain[\"num_words_upper\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\ntest[\"num_words_upper\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\nprint('maximum of num_words_upper in train',train[\"num_words_upper\"].max())\nprint(\"maximum of num_words_upper in test\",test[\"num_words_upper\"].max())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5bc4c2b642cb8adf12fd6dbb01616079a454d384"},"cell_type":"markdown","source":"Number of title case words in the text"},{"metadata":{"_kg_hide-input":true,"_uuid":"b938bdcfe7c418f5b4d57c9fd21c77d8bf4d3f06","trusted":true},"cell_type":"code","source":"\ntrain[\"num_words_title\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\ntest[\"num_words_title\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\nprint('maximum of num_words_title in train',train[\"num_words_title\"].max())\nprint(\"maximum of num_words_title in test\",test[\"num_words_title\"].max())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c81bb79d34d242f74b8ed650a8998efdb29e38b"},"cell_type":"markdown","source":" Average length of the words in the text "},{"metadata":{"_kg_hide-input":true,"_uuid":"3058236ff8702754ee4132f7eb705dd54f354af4","trusted":true},"cell_type":"code","source":"\ntrain[\"mean_word_len\"] = train[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ntest[\"mean_word_len\"] = test[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\nprint('mean_word_len in train',train[\"mean_word_len\"].max())\nprint(\"mean_word_len in test\",test[\"mean_word_len\"].max())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c91162602814ba230ab9fe30f9941ac6409133b9"},"cell_type":"markdown","source":"We add some new feature to train and test data set now, print columns agains"},{"metadata":{"_uuid":"05cae032149a7c79a92a3b2bf80185c483d0e976","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(train.columns)\ntrain.head(1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aa882e5bcdc7d5f440489eff75d1d225269655a4"},"cell_type":"markdown","source":"**<< Note >>**\n>**Preprocessing and generation pipelines depend on a model type**"},{"metadata":{"_uuid":"f453f5a76194116a73ce8ae5c98de980dc8b5758"},"cell_type":"markdown","source":"## What is Tokenizer?"},{"metadata":{"trusted":true,"_uuid":"b25495da861c8f917ff91c5c9296b954de03983f"},"cell_type":"code","source":"import nltk\nmystring = \"I love Kaggle\"\nmystring2 = \"I'd love to participate in kaggle competitions.\"\nnltk.word_tokenize(mystring)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cb760f3b6a02e2b4d5f371bb1fb5376f1cf7db88"},"cell_type":"code","source":"nltk.word_tokenize(mystring2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"055772bd170aa8018aabd85106b76675802c33b3"},"cell_type":"markdown","source":"<a id=\"64\"></a> <br>\n## 6-4 Data Visualization\n**Data visualization**  is the presentation of data in a pictorial or graphical format. It enables decision makers to see analytics presented visually, so they can grasp difficult concepts or identify new patterns.\n\n> * Two** important rules** for Data visualization:\n>     1. Do not put too little information\n>     1. Do not put too much information\n\n###### [Go to top](#top)"},{"metadata":{"_uuid":"5d991f5a4a9e4fffcbcee4a51b3cf1cd95007427"},"cell_type":"markdown","source":"<a id=\"641\"></a> <br>\n## 6-4-1 CountPlot"},{"metadata":{"_kg_hide-input":true,"_uuid":"1b54931579ed4e3004369a59fe9c6f23b97719de","trusted":true},"cell_type":"code","source":"ax=sns.countplot(x='target',hue=\"target\", data=train  ,linewidth=5,edgecolor=sns.color_palette(\"dark\", 3))\nplt.title('Is data set imbalance?');","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"eb15e8fed181179a086bf0db7dc21eaabb4eb088","trusted":true},"cell_type":"code","source":"ax = sns.countplot(y=\"target\", hue=\"target\", data=train)\nplt.title('Is data set imbalance?');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"be2b936baaa6dcc2d574a861c75584ed04d3589e"},"cell_type":"markdown","source":"<a id=\"642\"></a> <br>\n## 6-4-2  Pie Plot"},{"metadata":{"_kg_hide-input":true,"_uuid":"4a2332f8c87da0a4f8cc31f587ce547470a0d615","trusted":true},"cell_type":"code","source":"\nax=train['target'].value_counts().plot.pie(explode=[0,0.1],autopct='%1.1f%%' ,shadow=True)\nax.set_title('target')\nax.set_ylabel('')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee5a0017d867f90a5b89a78866582e73ceef5a05","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"#plt.pie(train['target'],autopct='%1.1f%%')\n \n#plt.axis('equal')\n#plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fbe8c50bcc1b632f42dd249e27a9a7c14517fd29"},"cell_type":"markdown","source":"<a id=\"643\"></a> <br>\n## 6-4-3  Histogram"},{"metadata":{"_kg_hide-input":true,"_uuid":"e41be26587175bee1bfe63e43eca1dd3f445c082","trusted":true},"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(20,10))\ntrain[train['target']==0].num_words.plot.hist(ax=ax[0],bins=20,edgecolor='black',color='red')\nax[0].set_title('target= 0')\nx1=list(range(0,85,5))\nax[0].set_xticks(x1)\ntrain[train['target']==1].num_words.plot.hist(ax=ax[1],color='green',bins=20,edgecolor='black')\nax[1].set_title('target= 1')\nx2=list(range(0,85,5))\nax[1].set_xticks(x2)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"a7d0a12e7f719781cd09e0d6df6ae1e780a6b1ab","trusted":true},"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(18,8))\ntrain[['target','num_words']].groupby(['target']).mean().plot.bar(ax=ax[0])\nax[0].set_title('num_words vs target')\nsns.countplot('num_words',hue='target',data=train,ax=ax[1])\nax[1].set_title('num_words:target=0 vs target=1')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"7870ad9dc4a007463301513963d6a2c1fe978aa4","trusted":true},"cell_type":"code","source":"# histograms\ntrain.hist(figsize=(15,20))\nplt.figure()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"e065ebff5374a9ab83df9c099a05962eb3645934","trusted":true},"cell_type":"code","source":"train[\"num_words\"].hist();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a0df20cff46a2bcdeef476e797d1535fd66850c4"},"cell_type":"markdown","source":"<a id=\"644\"></a> <br>\n## 6-4-4 Violin Plot"},{"metadata":{"_kg_hide-input":true,"_uuid":"fe8ca6c82ce44d745ad1ebc942826cf03e2d9895","trusted":true},"cell_type":"code","source":"sns.violinplot(data=train,x=\"target\", y=\"num_words\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8436fb0e0a159e9136a44dcd3eb9aae65a4a3b8b","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"sns.violinplot(data=train,x=\"target\", y=\"num_words_upper\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b08771239bf613afa9f89457fdf25c044279940e"},"cell_type":"markdown","source":"<a id=\"645\"></a> <br>\n## 6-4-5 KdePlot"},{"metadata":{"_kg_hide-input":true,"_uuid":"92f2ce0f1ff05002d196522a6a62579f0dba6ef3","trusted":true},"cell_type":"code","source":"sns.FacetGrid(train, hue=\"target\", size=5).map(sns.kdeplot, \"num_words\").add_legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a38e752250257db85554f00b9b440e5d968d4c7a"},"cell_type":"markdown","source":"<a id=\"646\"></a> <br>\n## 6-4-6 BoxPlot"},{"metadata":{"_uuid":"ede2021504e4c4d9eb53c9428f643f11e827f666","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train['num_words'].loc[train['num_words']>60] = 60 #truncation for better visuals\naxes= sns.boxplot(x='target', y='num_words', data=train)\naxes.set_xlabel('Target', fontsize=12)\naxes.set_title(\"Number of words in each class\", fontsize=15)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"05a3de95dca3933c118bbaca719a7bd6e4c007cd","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train['num_chars'].loc[train['num_chars']>350] = 350 #truncation for better visuals\n\naxes= sns.boxplot(x='target', y='num_chars', data=train)\naxes.set_xlabel('Target', fontsize=12)\naxes.set_title(\"Number of num_chars in each class\", fontsize=15)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"08745636a4aef9797daf0f52610cdd84d6cfd8f7"},"cell_type":"markdown","source":"<a id=\"646\"></a> <br>\n## 6-4-6 WordCloud"},{"metadata":{"_uuid":"582eb9b3a4e7ba61cea78b803c9fee55326f9940","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def generate_wordcloud(text): \n    wordcloud = wc(relative_scaling = 1.0,stopwords = eng_stopwords).generate(text)\n    fig,ax = plt.subplots(1,1,figsize=(10,10))\n    ax.imshow(wordcloud, interpolation='bilinear')\n    ax.axis(\"off\")\n    ax.margins(x=0, y=0)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3d192ec8ece7238fdd16327e7a2e81fcaf0ec18c","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"text =\" \".join(train.question_text)\ngenerate_wordcloud(text)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"97adc471c068fbd8d36ca19a4db0d98b0924c731"},"cell_type":"markdown","source":"-----------------\n<a id=\"8\"></a> <br>\n# 8- Conclusion"},{"metadata":{"_uuid":"1adfb5ba84e0f1d8fba58a2fca30546ead095047","collapsed":true},"cell_type":"markdown","source":"This kernel is not completed yet , I have tried to cover all the parts related to the process of **Quora problem** with a variety of Python packages and I know that there are still some problems then I hope to get your feedback to improve it.\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}