{"cells":[{"metadata":{"_uuid":"6d1ba4f1f1ffa4c5f4b7fcae0e539b5af457aee9"},"cell_type":"markdown","source":"**Thanks for viewing my Kernel! If you like my work and find it useful, please leave an upvote! :)**\n\n**Key Insights:** \n\n* 6.2% of the questions in training data are insincere\n* Insincere questions are dominated by words like trump, women, white, men, indian, muslims, black, americans, girls, indians, sex and india. More reference to specific groups of people of directly a person i.e. Donald Trump. \n* Top 3 bigrams in the insincere questions are 'Donald Trump', 'White People' and 'Black People'. Questions related to race are highly insincere. Presence of Chinese people, Indian muslims, Indian girls, North Indians, Indian women and White Women confirm the same.\n* Insincere questions are related to hypothetical scenarios, age, race, etc\n* Sincere questions are related to tips, advices, suggestions, facts, etc. \n* Insincere questions have more words, characters, stop words and punctuations"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"fb5fde08c825b8f478fb128e524b1a660fb3d173"},"cell_type":"code","source":"from IPython.display import Image\nImage(filename=\"../input/quora-image/quora.jpg\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport random\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\n\nfrom wordcloud import WordCloud, STOPWORDS\nfrom nltk.corpus import stopwords\nfrom collections import defaultdict\nimport string\n\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn import linear_model\nimport eli5\n\nimport os\nprint(os.listdir(\"../input/quora-insincere-questions-classification\"))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train = pd.read_csv('../input/quora-insincere-questions-classification/train.csv')\ntest = pd.read_csv('../input/quora-insincere-questions-classification/test.csv')\nsub = pd.read_csv('../input/quora-insincere-questions-classification/sample_submission.csv')\n\nprint('Train data: \\nRows: {}\\nCols: {}'.format(train.shape[0],train.shape[1]))\nprint(train.columns)\n\nprint('\\nTest data: \\nRows: {}\\nCols: {}'.format(test.shape[0],test.shape[1]))\nprint(test.columns)\n\nprint('\\nSubmission data: \\nRows: {}\\nCols: {}'.format(sub.shape[0],sub.shape[1]))\nprint(sub.columns)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7eebe70dd20b4ca455c1e9ecdd50cb18d07903a"},"cell_type":"markdown","source":"**6.2 % of train questions are insincere**"},{"metadata":{"trusted":true,"_uuid":"f237b111a5a6230d260f346e770ed5850970f5f5","_kg_hide-input":true},"cell_type":"code","source":"temp = train['target'].value_counts(normalize=True).reset_index()\n\ncolors = ['#4f92ff', '#4ffff0']\nexplode = (0.05, 0.05)\n \nplt.pie(temp['target'], explode=explode, labels=temp['index'], colors=colors,\n         autopct='%1.1f%%', shadow=True, startangle=0)\n \nfig = plt.gcf()\nfig.set_size_inches(12, 6)\nfig.suptitle('% Target Distribution', fontsize=16)\nplt.rcParams['font.size'] = 14\nplt.axis('equal')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"55cc3521dd4c9f00eb677a31bbe9e61a32060dc2"},"cell_type":"markdown","source":"Thanks to [SRK's](https://www.kaggle.com/sudalairajkumar) exploratory [kernel](https://www.kaggle.com/sudalairajkumar/simple-exploration-notebook-qiqc) for the custom functions. I modified them a bit so that I can reuse for any dataframe and n-gram combination. "},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"add986af56f31baa64d35ea0d3fe11cca74376de"},"cell_type":"code","source":"def ngram_extractor(text, n_gram):\n    token = [token for token in text.lower().split(\" \") if token != \"\" if token not in STOPWORDS]\n    ngrams = zip(*[token[i:] for i in range(n_gram)])\n    return [\" \".join(ngram) for ngram in ngrams]\n\n# Function to generate a dataframe with n_gram and top max_row frequencies\ndef generate_ngrams(df, col, n_gram, max_row):\n    temp_dict = defaultdict(int)\n    for question in df[col]:\n        for word in ngram_extractor(question, n_gram):\n            temp_dict[word] += 1\n    temp_df = pd.DataFrame(sorted(temp_dict.items(), key=lambda x: x[1])[::-1]).head(max_row)\n    temp_df.columns = [\"word\", \"wordcount\"]\n    return temp_df\n\ndef comparison_plot(df_1,df_2,col_1,col_2, space):\n    fig, ax = plt.subplots(1, 2, figsize=(20,10))\n    \n    sns.barplot(x=col_2, y=col_1, data=df_1, ax=ax[0], color=\"palegreen\")\n    sns.barplot(x=col_2, y=col_1, data=df_2, ax=ax[1], color=\"palegreen\")\n\n    ax[0].set_xlabel('Word count', size=14, color=\"green\")\n    ax[0].set_ylabel('Words', size=14, color=\"green\")\n    ax[0].set_title('Top words in sincere questions', size=18, color=\"green\")\n\n    ax[1].set_xlabel('Word count', size=14, color=\"green\")\n    ax[1].set_ylabel('Words', size=14, color=\"green\")\n    ax[1].set_title('Top words in insincere questions', size=18, color=\"green\")\n\n    fig.subplots_adjust(wspace=space)\n    \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ed642810c199f86d4cb899ef124094ce6e714de7"},"cell_type":"markdown","source":"**Top 20 1-gram words in sincere and insincere questions**\n* Sincere questions are dominated by words like best, will, people, good, one, etc. with no reference to any specific nouns.  Some of these words are high even in insincere words - meaning they are not significant to the classification. \n* Insincere questions are dominated by words like trump, women, white, men, indian, muslims, black, americans, girls, indians, sex and india. More reference to specific groups of people of directly a person i.e. Donald Trump. "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"ac7ae64d7a87d47d2f48c878f469d95319def26c"},"cell_type":"code","source":"sincere_1gram = generate_ngrams(train[train[\"target\"]==0], 'question_text', 1, 20)\ninsincere_1gram = generate_ngrams(train[train[\"target\"]==1], 'question_text', 1, 20)\n\ncomparison_plot(sincere_1gram,insincere_1gram,'word','wordcount', 0.25)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"50e28e54c3db347d117d4d22116eb5ea136ec69a"},"cell_type":"markdown","source":"**Top 20 2-gram words in sincere and insincere questions**\n* Top 3 bigrams in the insincere questions are 'Donald Trump', 'White People' and 'Black People'. Questions related to race are highly insincere. \n* Presence of Chinese people, Indian muslims, Indian girls, North Indians, Indian women and White Women confirm the same.\n* Sincere questions have best way, year old, will happen, etc. as the top ones. No clear trend there but 'best' is the key word to look for. "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"bfc93be17187ec6ed88cab7339866532c6a837b8"},"cell_type":"code","source":"sincere_2gram = generate_ngrams(train[train[\"target\"]==0], 'question_text', 2, 20)\ninsincere_2gram = generate_ngrams(train[train[\"target\"]==1], 'question_text', 2, 20)\n\ncomparison_plot(sincere_2gram,insincere_2gram,'word','wordcount', .35)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"deac2ac9c083dda1dde6b2d0ff19afa081b82ffb"},"cell_type":"markdown","source":"**Top 20 3-gram words in sincere and insincere questions**\n* Insincere questions are related to hypothetical scenarios, age, race, etc\n* Sincere questions are related to tips, advices, suggestions, facts, etc. "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"f3a879c023f602b5d41aed3093aa2dd56ce11dab"},"cell_type":"code","source":"sincere_3gram = generate_ngrams(train[train[\"target\"]==0], 'question_text', 3, 20)\ninsincere_3gram = generate_ngrams(train[train[\"target\"]==1], 'question_text', 3, 20)\n\ncomparison_plot(sincere_3gram,insincere_3gram,'word','wordcount', .45)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d39a3c869e48083d42a19fbffcf5725f68285330"},"cell_type":"markdown","source":"**Insincere questions have more words per question**"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"6b1b5e646f03d5fc339c300169f6bf5479108b08"},"cell_type":"code","source":"# Number of words in the questions\ntrain[\"word_count\"] = train[\"question_text\"].apply(lambda x: len(str(x).split()))\ntest[\"word_count\"] = test[\"question_text\"].apply(lambda x: len(str(x).split()))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"word_count\", y=\"target\", data=train, ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Word Count', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Word Count distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"13545f8390f3874d11d1a26331d1f3b751ba5ac6"},"cell_type":"code","source":"# Number of unique words in the questions\ntrain[\"unique_word_count\"] = train[\"question_text\"].apply(lambda x: len(set(str(x).split())))\ntest[\"unique_word_count\"] = test[\"question_text\"].apply(lambda x: len(set(str(x).split())))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"unique_word_count\", y=\"target\", data=train, ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Unique Word Count', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Unique Word Count distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b76476418ddb1e475046980dbc40e8880e5611f0"},"cell_type":"markdown","source":"**Insincere questions have more characters than sincere questions**"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"179495370298a150e89cf7899d08e1189999981d"},"cell_type":"code","source":"# Number of characters in the questions\ntrain[\"char_length\"] = train[\"question_text\"].apply(lambda x: len(str(x)))\ntest[\"char_length\"] = test[\"question_text\"].apply(lambda x: len(str(x)))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"char_length\", y=\"target\", data=train, ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Character Length', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Character Length distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"573072e0272f3d7db793867e14dc25b611f234b7"},"cell_type":"markdown","source":"**Insincere questions have more stop words than sincere questions**"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"7635644d0885af517da46d945c26bba84d736159"},"cell_type":"code","source":"# Number of stop words in the questions\ntrain[\"stop_words_count\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\ntest[\"stop_words_count\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"stop_words_count\", y=\"target\", data=train, ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of stop words', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Number of Stop Words distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"de1a577605d772301173d22a7388ce711a8d758d"},"cell_type":"markdown","source":"**Insincere questions have more punctuations**"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"4265c6a73866b421881ea194ca934ea49b5ff1d8"},"cell_type":"code","source":"# Number of punctuations in the questions\ntrain[\"punc_count\"] = train[\"question_text\"].apply(lambda x: len([c for c in str(x) if c in string.punctuation]))\ntest[\"punc_count\"] = test[\"question_text\"].apply(lambda x: len([c for c in str(x) if c in string.punctuation]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"punc_count\", y=\"target\", data=train[train['punc_count']<train['punc_count'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of punctuations', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Punctuation distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6e7aa2ba7500bdb1df9957ffafb33a03a5d257ff"},"cell_type":"markdown","source":"**More upper case words in Sincere questions**"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"f7a6fba9fc1a4ad7e3f1e88aced0951c6d754587"},"cell_type":"code","source":"# Number of upper case words in the questions\ntrain[\"upper_words\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\ntest[\"upper_words\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"upper_words\", y=\"target\", data=train[train['upper_words']<train['upper_words'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of Upper case words', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Upper case words distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"2ee2e9ea1494da98bc8a0fc9b0110ddf7d65327d"},"cell_type":"code","source":"# Number of title words in the questions\ntrain[\"title_words\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\ntest[\"title_words\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"title_words\", y=\"target\", data=train[train['title_words']<train['title_words'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of Title words', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Title words distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"dbfbb43edf02e8fa9358249a5474b54db9c49597"},"cell_type":"code","source":"# Mean word length in the questions\ntrain[\"word_length\"] = train[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ntest[\"word_length\"] = test[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"word_length\", y=\"target\", data=train[train['word_length']<train['word_length'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Mean word length', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Distribution of mean word length', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"847cc2fbeac831b3a92c76d8ceeb7c9b365c4937"},"cell_type":"markdown","source":"**A base model with vectorized matrix shows that insincere words are predominantly due to identification of religion, nationality, race, caste, political affiliation, etc. **"},{"metadata":{"trusted":true,"_uuid":"a506049980bafb488294e761da6174a144e99d81"},"cell_type":"code","source":"# Get the tfidf vectors\ntfidf_vec = TfidfVectorizer(stop_words='english', ngram_range=(1,3))\ntfidf_vec.fit_transform(train['question_text'].values.tolist() + test['question_text'].values.tolist())\ntrain_tfidf = tfidf_vec.transform(train['question_text'].values.tolist())\ntest_tfidf = tfidf_vec.transform(test['question_text'].values.tolist())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e76b5292ef0b01d66f7815a0b7760dae3397267"},"cell_type":"code","source":"y_train = train[\"target\"].values\n\nx_train = train_tfidf\nx_test = test_tfidf\n\nmodel = linear_model.LogisticRegression(C=5., solver='sag')\nmodel.fit(x_train, y_train)\ny_test = model.predict_proba(x_test)[:,1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ce3505661488b6e42733a0909c1fe04e49848ed5"},"cell_type":"code","source":"eli5.show_weights(model, vec=tfidf_vec, top=100, feature_filter=lambda x: x != '<BIAS>')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f6695255f48781ead442a076d33974fa88436d1"},"cell_type":"code","source":"sub['prediction'] = y_test\nsub.to_csv('baseline_submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0af8d3de82a2137440430079421b02270fbcd57f"},"cell_type":"markdown","source":"**More to come!!!**\n* Note to self: Questions marks are present in the words. Should remove them while processing."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}