{"cells":[{"metadata":{"_uuid":"dd949d8c86abb1718d8f6f1f658fc0dfb2429809"},"cell_type":"markdown","source":"# Quora Insincere Questions Classification"},{"metadata":{"_uuid":"2275a55daa6a5dff2ab2965796f4fb0766ca8822"},"cell_type":"markdown","source":"## Detect toxic content to improve online conversations"},{"metadata":{"_uuid":"7e4199353ca9c1134c32ba81f30973d6831afae1"},"cell_type":"markdown","source":"## Problem\n\n* **Handle toxic and disivie content / miseleading content**"},{"metadata":{"_uuid":"435c8ef2d315b56480f4fdb7d2049dbd8a74819a"},"cell_type":"markdown","source":"* **To do that we have to develop a model that identify and flag insincere questions**"},{"metadata":{"_uuid":"79eebd8067de178b729156ede5cb04dc96d8d958"},"cell_type":"markdown","source":"## Evaluation"},{"metadata":{"_uuid":"9d4bc29c43715a455883fa4f4a446c9147b99684"},"cell_type":"markdown","source":"* **For each qid in the testset, predict the corresponding questions_text:**\n\n * **Is Insincere => 1**\n\n * **Is Sincere => 0**\n\n* **Submissions are evaluated on F1 Score between the predicted and the observed targets**"},{"metadata":{"_uuid":"f94813892c456114585ceefb5baee2cf1881db8c"},"cell_type":"markdown","source":"## Data Fields"},{"metadata":{"_uuid":"e692da197db736b90ade2f7711e848f273e0ee12"},"cell_type":"markdown","source":"* **qid : unique question identifer**\n\n* **question_text : Quora question text**\n\n* **target : a question labeled << insincere >> has a value of 1, otherwise 0**"},{"metadata":{"trusted":true,"_uuid":"ec9266ebccd23a0c214c6e17bfa02dac3a7270f4"},"cell_type":"code","source":"import os\nimport numpy as np \nimport pandas as pd \n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5306e257e42ad3a1fb157c6f7b82da80d293ed66"},"cell_type":"code","source":"import nltk\nfrom sklearn.pipeline import Pipeline\nfrom nltk.corpus import stopwords\nfrom string import punctuation \nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.feature_extraction.text import CountVectorizer,TfidfVectorizer\nfrom sklearn import model_selection\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import f1_score","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9593cff560e60db77a34e02d8ef61e5dbf96aa7d"},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f6d20d91c4a519e683fa03dad1dc7260a22d3edf"},"cell_type":"markdown","source":"## Load the dataset"},{"metadata":{"trusted":true,"_uuid":"1ee38aeadba07001f92124415b711dbb091cd04c"},"cell_type":"code","source":"path = \"../input\"\ndf_train = pd.read_csv(os.path.join(path, \"train.csv\"))\ndf_test = pd.read_csv(os.path.join(path, \"test.csv\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ea1bab61378e22fbde92102117e76a893ee9573f"},"cell_type":"code","source":"df_train.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4d71752bcf2ceee114aaa6fe31706e1260fcae85"},"cell_type":"code","source":"df_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b2425404625e68013a958f30dce35d1025de7f93"},"cell_type":"code","source":"df_test.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"75fff275d47794e2122344f86fa5864815f0d2e5"},"cell_type":"code","source":"df_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1dded5db9eeaff6829a3adb5a4453ec1508b3953"},"cell_type":"code","source":"df_train[\"target\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c9370a3e0b13ee8311da62bd166a716a54630ee"},"cell_type":"markdown","source":"## Exploratory Data Analysis"},{"metadata":{"trusted":true,"_uuid":"395c0411d05c1b0870f926d96c258cf57965a5e7"},"cell_type":"code","source":"insincere = df_train[df_train[\"target\"] == 1]\nsincere = df_train[df_train[\"target\"] == 0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ccacf32a00cf5708ac7f15f68d25cac28891872c"},"cell_type":"code","source":"sincere.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cc82ab1a6a637068d22a1f63129b341df3d8819a"},"cell_type":"code","source":"insincere.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4ff96fa74112c02fe87c41657b46bf15c09bb106"},"cell_type":"markdown","source":"### Distribution of Sincere / Insincere questions"},{"metadata":{"trusted":true,"_uuid":"e5aa100c1cdbdb31e0772e6868b2b5a1192be37c"},"cell_type":"code","source":"question_class = df_train[\"target\"].value_counts()\nquestion_class.plot(kind= \"bar\", color= [\"blue\", \"orange\"])\nplt.title(\"Bar chart\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"369f10cddf022b9ceac4590de6fe3458ca0d7ae2"},"cell_type":"code","source":"print(df_train[\"target\"].value_counts())\nprint(sum(df_train[\"target\"] == 1) / sum(df_train[\"target\"] == 0) * 100, \"percent of questions are insincere.\")\nprint(100 - sum(df_train[\"target\"] == 1) / sum(df_train[\"target\"] == 0) * 100, \"percent of questions are sincere\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e0e6b3e51eaf6dbde66af031a64461e57fdb3629"},"cell_type":"markdown","source":"**We have a Unbalenced Data**"},{"metadata":{"_uuid":"9a97e65a9b9f24829051c896493da25c30755518"},"cell_type":"markdown","source":"### Insincere Word cloud"},{"metadata":{"trusted":true,"_uuid":"7915bbc99fd8a5f6250cbe985473565094da9e3c"},"cell_type":"code","source":"from wordcloud import WordCloud, STOPWORDS\nstopwords = set(STOPWORDS)\n\n# Generate a word cloud image\nsincere_wordcloud = WordCloud(width=600, height=400).generate(str(sincere[\"question_text\"]))\n#Positive Word cloud\nplt.figure(figsize=(10,8), facecolor=\"black\")\nplt.imshow(sincere_wordcloud)\nplt.axis(\"off\")\nplt.tight_layout(pad=0)\nplt.show();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"010873fc883e30bb5efa0618d56e9abd756eab01"},"cell_type":"markdown","source":"### Insincere Word cloud"},{"metadata":{"trusted":true,"_uuid":"71d839a50ef3143a721a2890f247222324d616f4"},"cell_type":"code","source":"from wordcloud import WordCloud, STOPWORDS\nstopwords = set(STOPWORDS)\n\n# Generate a word cloud image\ninsincere_wordcloud = WordCloud(width=600, height=400).generate(str(insincere[\"question_text\"]))\n#Positive Word cloud\nplt.figure(figsize=(10,8), facecolor=\"black\")\nplt.imshow(insincere_wordcloud)\nplt.axis(\"off\")\nplt.tight_layout(pad=0)\nplt.show();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b23dc419cb99af4e12fc365847f3ac0db9e3f906"},"cell_type":"markdown","source":"## Feature Engineering"},{"metadata":{"_uuid":"31f56db557ab8fc643d5f0b50b86082d56769afc"},"cell_type":"markdown","source":"### Number of words in the text"},{"metadata":{"trusted":true,"_uuid":"321849103d2f1e9d2103804a4d3d72b60afb7876"},"cell_type":"code","source":"df_train[\"number_words\"] = df_train[\"question_text\"].apply(lambda x: len(x.split()))\ndf_test[\"number_words\"] = df_test[\"question_text\"].apply(lambda x: len(x.split()))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d7c11fbc42ef14c4af5cadfb6801838386f3d2a4"},"cell_type":"markdown","source":"### Number of unique words in the text"},{"metadata":{"trusted":true,"_uuid":"ff78030d9a3acadb4fff7c6ac1b486431d01a112"},"cell_type":"code","source":"df_train[\"num_unique_words\"] = df_train[\"question_text\"].apply(lambda x: len(set(str(x).split())))\ndf_test[\"num_unique_words\"] = df_test[\"question_text\"].apply(lambda x: len(set(str(x).split())))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"db326affbd7401a41b1346f294b19ea4bbb429fc"},"cell_type":"markdown","source":"### Number of characters in the text"},{"metadata":{"trusted":true,"_uuid":"2a96cb0ff778ec1fe57c4d6e36949e9c6a5597ba"},"cell_type":"code","source":"df_train[\"num_chars\"] = df_train[\"question_text\"].apply(lambda x: len(str(x)))\ndf_test[\"num_chars\"] = df_test[\"question_text\"].apply(lambda x: len(str(x)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a5528d5def5194e93b450cb6c0b68660f3444bb"},"cell_type":"markdown","source":"### Number of Stopwords in text"},{"metadata":{"trusted":true,"_uuid":"3ac5b6acb1af13c80235650329a1850dc7c111d7"},"cell_type":"code","source":"from nltk.corpus import stopwords \nstop_words = set(stopwords.words(\"english\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f2d4fd155385f44aaba5be22e9537605931310a4"},"cell_type":"code","source":"df_train[\"num_stopwords\"] = df_train[\"question_text\"].apply(lambda x : len([nw for nw in str(x).split() if nw.lower() in stop_words]))\ndf_test[\"num_stopwords\"] = df_test[\"question_text\"].apply(lambda x : len([nw for nw in str(x).split() if nw.lower() in stop_words]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"59371e332a81e2edf585bc3feff40784b2a3b041"},"cell_type":"markdown","source":"### Number of punctuations in text"},{"metadata":{"trusted":true,"_uuid":"afe997b6b05b22884773626dc7dadef57fbbd84d"},"cell_type":"code","source":"df_train[\"num_punctuation\"] = df_train[\"question_text\"].apply(lambda x : len([np for np in str(x) if np in punctuation]))\ndf_test[\"num_punctuation\"] = df_test[\"question_text\"].apply(lambda x : len([np for np in str(x) if np in punctuation]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3fa9d93ea292df5ca017f16e0e5523430e3722bb"},"cell_type":"markdown","source":"### Number of Upper case and Lower case in text"},{"metadata":{"trusted":true,"_uuid":"eb7382e2a7869f13f3636be2d6038d5137af66bf"},"cell_type":"code","source":"df_train[\"num_uppercase\"] = df_train[\"question_text\"].apply(lambda x : len([nu for nu in str(x).split() if nu.isupper()]))\ndf_test[\"num_uppercase\"] = df_test[\"question_text\"].apply(lambda x : len([nu for nu in str(x).split() if nu.isupper()]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a147e86865ee8547f15818eae4a0079ce64077dc"},"cell_type":"code","source":"df_train[\"num_lowercase\"] = df_train[\"question_text\"].apply(lambda x : len([nl for nl in str(x).split() if nl.islower()]))\ndf_test[\"num_lowercase\"] = df_test[\"question_text\"].apply(lambda x : len([nl for nl in str(x).split() if nl.islower()]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3241491aad2789cb0d71e05adb4f00a21b2b8854"},"cell_type":"markdown","source":"### Number of title in text"},{"metadata":{"trusted":true,"_uuid":"524f7e7b75471851a5af86541cfd8b20c86e373d"},"cell_type":"code","source":"df_train[\"num_title\"] = df_train[\"question_text\"].apply(lambda x : len([nl for nl in str(x).split() if nl.istitle()]))\ndf_test[\"num_title\"] = df_test[\"question_text\"].apply(lambda x : len([nl for nl in str(x).split() if nl.istitle()]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"499865af52b6f2e882248282947472dcbf98d496"},"cell_type":"code","source":"df_train[df_train[\"target\"] == 1].describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d9a30445a4f6ae15dcbb789070de3401161ade84"},"cell_type":"code","source":"df_train[df_train[\"target\"] == 0].describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee67ce6c9105421a35deec3725117e2127c2fd7a"},"cell_type":"markdown","source":"### Box plot of number of words according to the target"},{"metadata":{"trusted":true,"_uuid":"407249a41f38daa9acc9cf2e63cac5e7d1d531d6"},"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(12, 10)\nsns.boxplot(data=df_train, y=\"number_words\", x=\"target\",orient=\"v\")\nax.set(xlabel=\"target\", ylabel=\"number of words\", title=\"Box plot of number of words according to the target\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c6473dbecedccfa787a38e512159c8b1e394e9b0"},"cell_type":"markdown","source":"### Box plot of number of unique words according to the target"},{"metadata":{"trusted":true,"_uuid":"c1f422c1c12e321caed91d36bc8d30335e9dde2f"},"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(12, 10)\nsns.boxplot(data=df_train, y=\"num_unique_words\", x=\"target\",orient=\"v\")\nax.set(xlabel=\"target\", ylabel=\"number of unique words\", title=\"Box plot of number of unique words according to the target\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73cb8d96538371e6ce15a8943590599e6a2bd1d4"},"cell_type":"markdown","source":"### Box plot of number of characters according to the target"},{"metadata":{"trusted":true,"_uuid":"5f90e02fda06739bf08202ddfec636d8350dcbf2"},"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(12, 10)\nsns.boxplot(data=df_train, y=\"num_chars\", x=\"target\",orient=\"v\")\nax.set(xlabel=\"target\", ylabel=\"number of characters\", title=\"Box plot of number of characters according to the target\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"acea6ad702831f36396694dbd19801f204ee5f95"},"cell_type":"markdown","source":"### Box plot of number of stopwords according to the target"},{"metadata":{"trusted":true,"_uuid":"de194a2069b657fb35c9c83ab95a48b8eaa3dbd4"},"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(12, 10)\nsns.boxplot(data=df_train, y=\"num_stopwords\", x=\"target\",orient=\"v\")\nax.set(xlabel=\"target\", ylabel=\"number of stopwords\", title=\"Box plot of number of stopwords according to the target\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"521b5d63e9a40c03572f335cdd31e6c7f6a7fb89"},"cell_type":"markdown","source":"### Box plot of number of lowercases according to the target"},{"metadata":{"trusted":true,"_uuid":"a8e96ad9a43a48e1bdba05735810edead8f7644f"},"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(12, 10)\nsns.boxplot(data=df_train, y=\"num_lowercase\", x=\"target\",orient=\"v\")\nax.set(xlabel=\"target\", ylabel=\"number of lowercases\", title=\"Box plot of number of lowercases according to the target\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"550ca63001a8f1a53beb0027063a5279d7a58fe6"},"cell_type":"markdown","source":"### Box plot of number of titles according to the target"},{"metadata":{"trusted":true,"_uuid":"28c3a9872b34ebeb848823908cb987f95a21f54d"},"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(12, 10)\nsns.boxplot(data=df_train, y=\"num_title\", x=\"target\",orient=\"v\")\nax.set(xlabel=\"target\", ylabel=\"number of titles\", title=\"Box plot of number of titles according to the target\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"37dc3a0c43c939fbfa2e6e8c66ac2972dd4dc609"},"cell_type":"markdown","source":"## Remove Stopwords and Punctuation"},{"metadata":{"trusted":true,"_uuid":"c67c887aefe5a1f627aaf00be7f57907699ea788"},"cell_type":"code","source":"def text_process(question):\n    nopunc = [char for char in question if char not in punctuation]\n    nopunc = \"\".join(nopunc)\n    meaning = [word for word in nopunc.split() if word.lower() not in stopwords.words(\"english\")]\n    return( \" \".join( meaning ))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fbdf6a3b68a7d608fbd205e8c65a1945045772d6"},"cell_type":"code","source":"clean_train = df_train[\"question_text\"].apply(text_process)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"983f3a47abbe914cfce42a2244e046cbdeef7ca0"},"cell_type":"code","source":"clean_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"51b349dfe5155b199ce9f8468bee498816f9ab59"},"cell_type":"code","source":"clean_test = df_test[\"question_text\"].apply(text_process)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f1b486575432ebb03776d94740028d82cc1deff6"},"cell_type":"code","source":"clean_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d5c7defa80924ee51908f97ca820d357e72c1d2b"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"392656da4123b76fb77103406ac5bd812fe7013b"},"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(clean_train, df_train.target.values, test_size=0.2, stratify = df_train.target.values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d5f3cc31a283f0801c0e0399ff8dc0d57b5a0041"},"cell_type":"code","source":"X_train.shape, X_val.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7c172aab8baf54e33c5f984cb3adf4d33c9e90f7"},"cell_type":"code","source":"y_train.shape, y_val.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"969db78002c70c841614bc2ffff6407b411e9301"},"cell_type":"markdown","source":"## Model Building"},{"metadata":{"trusted":true,"_uuid":"6a203f6629781642d05386dc83a3ca41f2adbc82"},"cell_type":"code","source":"pipeline = Pipeline([(\"cv\",CountVectorizer(analyzer=\"word\",ngram_range=(1,4),max_df=0.9)),\n                     (\"clf\",LogisticRegression(solver=\"saga\", class_weight=\"balanced\", C=0.45, max_iter=250, verbose=1))])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"37d8a20957145f7e1f7fae9cb8b2c4d549ad8a5f"},"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(clean_train, df_train.target.values, test_size=0.1, stratify = df_train.target.values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"793b34ce95739b1adc7128afae959dfef530c5a4"},"cell_type":"code","source":"lr_model = pipeline.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"58b883eb04c348d33a25efe0c895a0109183787f"},"cell_type":"code","source":"lr_model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"375d263e736b24d2a013c194fc4c9653cc13b784"},"cell_type":"code","source":"y_pred = lr_model.predict(X_val)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"233c0fb4f81ea0549bd470ea1b1d83d8993b27de"},"cell_type":"markdown","source":"## Evaluating Results"},{"metadata":{"trusted":true,"_uuid":"be41a096b44fd4cb92bb746ad2122395681b5781"},"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, classification_report\ncm = confusion_matrix(y_val, y_pred)\nsns.heatmap(cm, cmap=\"Blues\", annot=True, square=True, fmt=\".0f\");","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"75262dd65e5b23f8a167e42a3460ae741d40c505"},"cell_type":"code","source":"print(classification_report(y_val, y_pred))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e415b4d7b6243ab4e8e18a38d592de54b9937e65"},"cell_type":"code","source":"y_pred_final = pipeline.predict(clean_test)\ny_pred_final","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"39ac2e86f7aa2098e14f6c41ff4d3b3bd5995976"},"cell_type":"markdown","source":"## Formatting Final Results"},{"metadata":{"trusted":true,"_uuid":"0f815801bb4a6a1dfc3441c6cae4c09f4e5f1833"},"cell_type":"code","source":"df_sub = pd.DataFrame({\"qid\":df_test[\"qid\"], \"prediction\":y_pred_final})\ndf_sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b221812ecdf3ffa5c4700af5c901ec4c04bd80a0"},"cell_type":"code","source":"df_sub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.6.7"}},"nbformat":4,"nbformat_minor":1}