{"cells":[{"metadata":{"_uuid":"0130311d0de047caa7f7d76637bd8eb635a36927"},"cell_type":"markdown","source":"***Exploratory Data Analysis (EDA) for Quora Insincere Questions Competition***"},{"metadata":{"_uuid":"5e438f15523a5a31c10bfe4b9be3156323f085cd"},"cell_type":"markdown","source":"Let's import a set of standard useful packages that will help us do our exploratory analysis.\n\nWe will use these to explore the Quora data in various ways."},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport time\nfrom tqdm import tqdm\n\nfrom sklearn.metrics import f1_score\nfrom sklearn.model_selection import KFold\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.naive_bayes import GaussianNB, MultinomialNB, BernoulliNB\n\nimport nltk\n#nltk.download()\nfrom nltk.corpus import stopwords\nimport string\n\nfrom scipy.sparse import hstack\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport random\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5a8e359f9d3fa43b08d7353ee42ba20ee966ad9f"},"cell_type":"code","source":"# controls whether we work with only a subset of the data\nSAMPLE_DATA = None","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"288b2eae0026b09c3dcee227631556da127ef5cc"},"cell_type":"markdown","source":"Next let's take a look at the input files provided as part of this competition."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"print(os.listdir(\"../input\"))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dab942fb829b6806cb15b882a78ab0c4be2ce6b6"},"cell_type":"markdown","source":"Finally, let's load the training set and make sure the data is loaded correctly."},{"metadata":{"_uuid":"1b5b599a7885b34340140a505da203b49db3e202","trusted":true},"cell_type":"code","source":"train = pd.read_csv('../input/train.csv')\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"81b2f6dee596a50fd621e5f16e6296efd7578204"},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"74a96087757018680a8acf13bd6a1531892edd0e"},"cell_type":"markdown","source":"Let's see how prevalent the two classes are:"},{"metadata":{"_uuid":"1e398f9251179e3407f1a3c0f2236735ef3c5f14"},"cell_type":"markdown","source":"There are far fewer examples of insincere questions (target=1) than sincere questions (target=0)."},{"metadata":{"trusted":true,"_uuid":"828942774930c0a14fd547a1b3ef1c6b2f213ccb"},"cell_type":"code","source":"sns.countplot(x='target',data=train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9a10b4980e23938c34520e4bdf306c0979fbf8fe"},"cell_type":"markdown","source":"Before going any further,  let's take a look at what some random `question_text` looks like so we get a better idea of the format and length of the questions."},{"metadata":{"_uuid":"bbf61bc22f89a18a88bbd3b004c776a1c0285b28","scrolled":false,"trusted":true},"cell_type":"code","source":"train_sample = train.sample(n=50)\ntrain_sample.question_text.head(n=25)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c9e46490f18526c9c17acc37c396b1e24c7970e"},"cell_type":"markdown","source":"And let's look at some which have been marked as Insincere:"},{"metadata":{"_uuid":"eba8979a37661c730c01bd959c8671d5c91ffaa3","scrolled":true,"trusted":true},"cell_type":"code","source":"train[train.target == 1].question_text","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c247d18d951fec80fc4610b90402880a0ee1071"},"cell_type":"markdown","source":"Let's think about some useful attributes we might like to add to our training data.  Some that spring to mind are based solely on the content of the questions:\n1. Length of question\n2. Number of tokens\n3. Number of stop words\n4. Number of puncuation marks\n5. mean word length\n"},{"metadata":{"_uuid":"534b49124f6eac5dee0eee947abb9a373ccee96f"},"cell_type":"markdown","source":"These meta-attributes may turn out to be useful later as we attempt to predict which questions are insincere."},{"metadata":{"_uuid":"2b5d32cc00f4c4f5382692d64ba3e9d3504661b3"},"cell_type":"markdown","source":"To speed up the calculation of metafeatures, we could code against a the sample subset so we can iterate more quickly."},{"metadata":{"_uuid":"b2ec2381454eb0442d277526113a787d56a586c4","trusted":true},"cell_type":"code","source":"# temporarily calculate metafeatures against only train_sample to speed up iterative development\nif SAMPLE_DATA:\n    train = train_sample","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f8c44ff112ce9077a07c36d0e7448a17e1200709"},"cell_type":"markdown","source":"Before we proceed with computing metafeatures we should stop and normalize the text so that all words are tokenized - lower cased, whitespace removed, etc. so that later computations can deal with uniform input."},{"metadata":{"_uuid":"a346ebf43e521023c4f16e107f56f39ac2d849cd","trusted":true},"cell_type":"code","source":"from nltk import word_tokenize\n%timeit train['question_tokens'] = train['question_text'].str.lower().str.split(' ')\n#train['question_tokens'].head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6af56ae697663d6ae857b17e8277c62d755a0ec4"},"cell_type":"markdown","source":"Let's add columns to the training set for length of question and number of tokens."},{"metadata":{"_uuid":"759a03da2c6858d6b9c478da730b271a0a07d481","trusted":true},"cell_type":"code","source":"train['question_length'] = train['question_text'].str.len()\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aba554d370fc7077c9f251f25a2d5554aa4e7c4f","trusted":true},"cell_type":"code","source":"train['question_tokens_length'] = train['question_tokens'].str.len()\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b11cbd768c30c4891aa328d17e055bd9873a8e0a"},"cell_type":"markdown","source":"Using the NLTK package we can import the list of English stop words in order to compute stop word numbers."},{"metadata":{"_uuid":"f771c2af9c10d39855450aa78dca82ae0e4f5c64","trusted":true},"cell_type":"code","source":"eng_stopwords = set(stopwords.words(\"english\"))\nrandom.sample(eng_stopwords,15)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0203d389ac27a909d1e1b3f971359975fa6eafe8"},"cell_type":"markdown","source":"Let's define a function that we will apply to each question in the training set which will compute the number of stop words found in that question."},{"metadata":{"_uuid":"74ffcb9b5987e09c5475fce366d98d463b64d5bd","trusted":true},"cell_type":"code","source":"def num_stopwords(question_tokens):\n    return len([w for w in question_tokens if w in eng_stopwords])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8c8af523d4525fef28063ef33a42e0cfb554f9c3","scrolled":false,"trusted":true},"cell_type":"code","source":"train['question_num_stopwords'] = train['question_tokens'].apply(lambda question_tokens: num_stopwords(question_tokens))\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1a4a38d1107b9f44bcf5e4185a12904ef9beb5b0"},"cell_type":"markdown","source":"Python provides a string.punctuation method which we can use to count punctuation within the questions."},{"metadata":{"_uuid":"35333269f61d26a3e2296474c6822e88013ff84e","trusted":true},"cell_type":"code","source":"def num_punctuation(question_text):\n    punctuation_marks = list(string.punctuation)\n    return len([t for t in list(question_text) if t in punctuation_marks])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6ec6325b4e00ba5c98d021a4fa7efb8cee71fe55","trusted":true},"cell_type":"code","source":"train['num_punctuation'] = train['question_text'].apply(lambda question_text:num_punctuation(question_text))\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5b2771a7d7cddebac3a78b47d9dd223e9810040e"},"cell_type":"markdown","source":"Now let's compute a few metafeatures which use these features such as percent of tokens which are stopwords and punctuation."},{"metadata":{"_uuid":"adbced2c23a0df0a7fa03bcbf645320a828dd3ec","trusted":true},"cell_type":"code","source":"train['percent_punctuation'] = train['num_punctuation'] / train['question_length']\ntrain['percent_stopwords'] =train['question_num_stopwords']/train['question_tokens_length']\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0ec02871f65a0c15feaf1081fc42141790fddeb4"},"cell_type":"markdown","source":"Let's compute the mean word (token) length of each question."},{"metadata":{"_uuid":"e11b54f13d5d164976687dd0de519175c56e7aac","trusted":true},"cell_type":"code","source":"def mean_word_length(row):\n    return np.mean([len(str(w)) for w in row['question_tokens']])\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"30d3405585effc31bb878f277c33a90cb186289f","trusted":true},"cell_type":"code","source":"train['mean_word_length'] = train['question_tokens'].apply(lambda question_tokens:np.mean([len(w) for w in question_tokens]))\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"16d0f2c08cab6f16a9563af947f863a00989ee6e"},"cell_type":"markdown","source":"Let's get parts of speech for the tokens."},{"metadata":{"_uuid":"fd97ee3017b7722b79bf791459e185ea2e80f1c2","trusted":true},"cell_type":"code","source":"def get_pos(question_tokens):\n    pos_list = nltk.pos_tag([w for w in question_tokens if w])\n    return [pos[1] for pos in pos_list]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1611aa9c8fc1ae8a78ae53f8a4b94c3df292d16c","scrolled":false,"trusted":true},"cell_type":"code","source":"train['pos'] = train['question_tokens'].apply(lambda question_tokens:get_pos(question_tokens))\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8e388279f5d5baf12b336202ff9335910ea18ff1"},"cell_type":"markdown","source":"Let's compute the most frequent parts of speech for each question."},{"metadata":{"_uuid":"8f33f9273205c5a8742ad14fedd23373b9a38cf1","trusted":true},"cell_type":"code","source":"from statistics import mode\nimport statistics\n\ndef get_mode_of_pos(pos):\n    if not pos:\n        return ''\n    else:\n        poses = [p for p in pos if isinstance(p, str)]\n        if len(poses):\n            try:\n                m = mode(poses)\n                return m\n            except statistics.StatisticsError as e:\n                return ''","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0c93480048ec9455ca134157e2b48181639a8766","trusted":true},"cell_type":"code","source":"train['pos_mode'] = train['pos'].apply(lambda pos:get_mode_of_pos(pos))\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b79b879d323801513cde46652461accde789672e"},"cell_type":"code","source":"train['pos_first'] = train['pos'].str.slice(0,1)\ntrain['pos_last'] = train['pos'].str.slice(-1,1)\n#train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"82c887d25a9bcccbbade31c1ef1336378735db9a"},"cell_type":"markdown","source":"Let's make some visualizations based on these metafeatures."},{"metadata":{"_uuid":"7e0211886c6d1a4e57850ab1a5583c0b299e617f","trusted":true},"cell_type":"code","source":"sns.boxplot(x='target',y=\"percent_stopwords\",data=train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e3806d893533ababb5d3d87a7be4f8ad1f647ed0","trusted":true},"cell_type":"code","source":"sns.boxplot(x='target',y=\"question_length\",data=train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"104f0e56d412224edf9b96fe927fdde358bb1593","trusted":true},"cell_type":"code","source":"sns.boxplot(x='target',y=\"percent_punctuation\",data=train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5690e03b8c8a9b58818eaa963a8a4766a4eb4111","trusted":true},"cell_type":"code","source":"sns.boxplot(x='target',y=\"question_tokens_length\",data=train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7dd3b0a04403dbab782bac80e9791b4d8138015a","trusted":true},"cell_type":"code","source":"sns.boxplot(x='target',y=\"mean_word_length\",data=train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e91eb1d641826d6966cb3c606c0ef0870daaca44","trusted":true},"cell_type":"code","source":"sns.boxplot(x='target',y=\"mean_word_length\",hue=\"pos_mode\",data=train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8f9145d14cb73a194e3127a7f89afaa7a39cf38c"},"cell_type":"code","source":"sns.catplot(x='target',y=\"pos_mode\",data=train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9f853168396c958291314b08f3dff367aab657db"},"cell_type":"code","source":"train.hist(figsize=(15,20))\nplt.figure()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d8a1fb4732f5bfd2b5bd17c3bc5b4ffaeb620c23"},"cell_type":"markdown","source":"Lastly, let's save the preprocessed training data with metafeatures for future use."},{"metadata":{"trusted":true,"_uuid":"04f601c76f0cf79765fb82aa6a22ea9443939498"},"cell_type":"code","source":"train.to_csv('train_preprocessed.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"19ce20fa2aea99fcd00de35acd42fcce54181270"},"cell_type":"markdown","source":""}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}