{"cells":[{"metadata":{"_uuid":"65c61ad3cf834bb9353d31af919ed381b33ee927"},"cell_type":"markdown","source":"> ### Objective"},{"metadata":{"_uuid":"b8574fb73b68b8e47e39a01c8dc07cd86cccd15a"},"cell_type":"markdown","source":"In this competition you will be predicting whether a question asked on Quora is sincere or not.\n\nAn insincere question is defined as a question intended to make a statement rather than look for helpful answers. Some characteristics that can signify that a question is insincere:\n\n**Has a non-neutral tone**\n* Has an exaggerated tone to underscore a point about a group of people\n* Is rhetorical and meant to imply a statement about a group of people\n\n**Is disparaging or inflammatory**\n* Suggests a discriminatory idea against a protected class of people, or seeks confirmation of a stereotype\n* Makes disparaging attacks/insults against a specific person or group of people \n* Based on an outlandish premise about a group of people \n* Disparages against a characteristic that is not fixable and not measurable \n\n**Isn't grounded in reality**\n* Based on false information, or contains absurd assumptions\n* Uses sexual content (incest, bestiality, pedophilia) for shock value, and not to seek genuine answers"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport nltk\nfrom nltk.corpus import stopwords\nimport string\nfrom sklearn.model_selection import train_test_split                # to split the data\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import accuracy_score   \nfrom sklearn.metrics import classification_report, confusion_matrix\n\neng_stopwords = set(stopwords.words(\"english\"))\npd.options.mode.chained_assignment = None\n\nimport os\nprint(os.listdir(\"../input\"))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"236ea2079b8c4e581c89b4f63b9edb179b073a25"},"cell_type":"markdown","source":"#### The files that are provided by the kaggle team:"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"!ls ../input/","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"10690ab98b7d9ee5817f2432981e1f6409cb566f"},"cell_type":"markdown","source":"**File descriptions**\n\n* train.csv - the training set\n* test.csv - the test set\n* sample_submission.csv - A sample submission in the correct format\n* enbeddings/ - (see below)"},{"metadata":{"trusted":true,"_uuid":"e1b0b3ff51b65532ff333dd436c909d726530bbf"},"cell_type":"code","source":"# List the embeddings provided by kaggle team\n!ls ../input/embeddings/","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"822ed0b2eaf8e3c9d4449954fd004bdc581d291e"},"cell_type":"markdown","source":"**the below embedding are also downloadable from here**:\n\n* GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n* glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n* paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n* wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html"},{"metadata":{"trusted":true,"_uuid":"e48c075e9de5d090163af64f07f036afd3502bee"},"cell_type":"code","source":"## Read the train and test dataset and check the top few lines ##\ntrain_df = pd.read_csv(\"../input/train.csv\")\ntest_df = pd.read_csv(\"../input/test.csv\")\nprint(\"Number of rows in train dataset : \",train_df.shape[0])\nprint(\"Number of rows in test dataset : \",test_df.shape[0]) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8772632fc3ebbbeef737ab17aa254428d06b0c23"},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"23c1d479bda0d79d57a48db264198124d6d709e5"},"cell_type":"code","source":"#Check for the class-categorization count and also the class imbalance\ncnt_srs = train_df['target'].value_counts()\n\nplt.figure(figsize=(8,4))\nsns.barplot(cnt_srs.index, cnt_srs.values, alpha=0.8)\nplt.ylabel('Number of Occurrences', fontsize=12)\nplt.xlabel('target', fontsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"51ef0e831e21a370755aa293e76fd5eee177b401"},"cell_type":"markdown","source":"* As the distribution of target variable is varying a lot, hence accuracy is not the metric we need to look, May be F1-score."},{"metadata":{"trusted":true,"_uuid":"749a91bac61acbbe31504896b7f743c9d43a2f24"},"cell_type":"code","source":"# Let us print some lines of each of the questions cagtegory in quora to try and understand their writing style if possible.\ngrouped_df = train_df.groupby('target')\nfor name, group in grouped_df:\n    print(\"Target Name :\", name)\n    cnt =0\n    for ind, row in group.iterrows():\n        print(row['question_text'])\n        cnt += 1\n        if cnt == 2:\n            break\n    print(\"\\n\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b4cefc5e4eb0e662aa0d49b48dad729b1fb20874"},"cell_type":"markdown","source":"**Feature Engineering:**\n    \n* Now let us come try to do some feature engineering. This consists of two main parts.\n\n**Meta features -** features that are extracted from the text like number of words, number of stop words, number of punctuations etc\n\n**Text based features -** features directly based on the text / words like frequency, svd, word2vec etc."},{"metadata":{"_uuid":"8ea8e09fddd29512535f0116cb4facd88befd938"},"cell_type":"markdown","source":"**Meta Features:**\n    \n We will start with creating meta featues and see how good are they at predicting the  authors. \n\n The feature list is as follows:\n        \n* Number of words in the text\n* Number of unique words in the text\n* Number of characters in the text\n* Number of stopwords \n* Number of punctuations\n* Number of upper case words\n* Number of title case words\n* Average length of the words"},{"metadata":{"trusted":true,"_uuid":"2193a881ebca35e25a48c772d284181f129ca9eb"},"cell_type":"code","source":"# Number of words in the text \ntrain_df[\"num_words\"] = train_df[\"question_text\"].apply(lambda x: len(str(x).split()))\ntest_df[\"num_words\"] = test_df[\"question_text\"].apply(lambda x: len(str(x).split()))\n\n## Number of unique words in the text ##\ntrain_df[\"num_unique_words\"] = train_df[\"question_text\"].apply(lambda x: len(set(str(x).split())))\ntest_df[\"num_unique_words\"] = test_df[\"question_text\"].apply(lambda x: len(set(str(x).split())))\n\n## Number of characters in the text ##\ntrain_df[\"num_chars\"] = train_df[\"question_text\"].apply(lambda x: len(str(x)))\ntest_df[\"num_chars\"] = test_df[\"question_text\"].apply(lambda x: len(str(x)))\n\n## Number of stopwords in the text ##\ntrain_df[\"num_stopwords\"] = train_df[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in eng_stopwords]))\ntest_df[\"num_stopwords\"] = test_df[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in eng_stopwords]))\n\n## Number of punctuations in the text ##\ntrain_df[\"num_punctuations\"] =train_df['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\ntest_df[\"num_punctuations\"] =test_df['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\n\n## Number of title case words in the text ##\ntrain_df[\"num_words_upper\"] = train_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\ntest_df[\"num_words_upper\"] = test_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\n\n## Number of title case words in the text ##\ntrain_df[\"num_words_title\"] = train_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\ntest_df[\"num_words_title\"] = test_df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\n\n## Average length of the words in the text ##\ntrain_df[\"mean_word_len\"] = train_df[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ntest_df[\"mean_word_len\"] = test_df[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"450da0abe7be999ac940da6620eda6b4bbb45f1d"},"cell_type":"markdown","source":"* Let us now plot some of our new variables to see of they will be helpful in predictions."},{"metadata":{"trusted":true,"_uuid":"1a02201cbd009d07cec67d28f1ef3a8e0081fc29"},"cell_type":"code","source":"train_df.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3afbff4f4f2415e4dedeccfe3190e57823f4365f"},"cell_type":"code","source":"train_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"53c3f3bc54628f4fff3a80b8ef9b87b0f18bf7cd"},"cell_type":"code","source":"train_df['num_words'].loc[train_df['num_words']>50] = 50 #truncation for better visuals\nplt.figure(figsize=(10,6))\nsns.boxplot(x='target', y='num_words', data=train_df)\nplt.xlabel('target category', fontsize=12)\nplt.ylabel('Number of words in text', fontsize=12)\nplt.title(\"Number of words by target category\", fontsize=15)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a37dc5a855382bd32669ef3d3ca6630889b9cb6a"},"cell_type":"code","source":"train_df['num_punctuations'].loc[train_df['num_punctuations']>10] = 10 #truncation for better visuals\nplt.figure(figsize=(10,6))\nsns.boxplot(x='target', y='num_punctuations', data=train_df)\nplt.xlabel('target Name', fontsize=12)\nplt.ylabel('Number of puntuations in text', fontsize=12)\nplt.title(\"Number of punctuations by target category\", fontsize=15)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f88db34eb355c151b5c2fd4b76f5e27cd19a0087"},"cell_type":"code","source":"train_df['num_chars'].loc[train_df['num_chars']>300] = 300 #truncation for better visuals\nplt.figure(figsize=(10,6))\nsns.boxplot(x='target', y='num_chars', data=train_df)\nplt.xlabel('target Name', fontsize=12)\nplt.ylabel('Number of characters in text', fontsize=12)\nplt.title(\"Number of characters by target category\", fontsize=15)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fad0f7e29d04a093fb746d0ef5a33602503d5c2f"},"cell_type":"markdown","source":"![](http://)* In all the cases shown above , we observe that Number of words, Number of punctuations, Number of characters are more for insincere Text. "},{"metadata":{"_uuid":"610af22ae6f41a071ccb671b204801f5fd90b8cf"},"cell_type":"markdown","source":"**Modelling will be done in Next stage...stay tuned!!**"},{"metadata":{"_uuid":"fd87a8b5a8c8dfe2f81b438ea44fbed935083bc1"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"74b903d2e76bb29c8d7ad7362f5e17edc93bc86a"},"cell_type":"markdown","source":"**References**\n"},{"metadata":{"_uuid":"c36321a06ef79d483fe61ec58577654e9a00b5b8"},"cell_type":"markdown","source":"https://www.kaggle.com/sudalairajkumar/simple-feature-engg-notebook-spooky-author#"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}