{"cells":[{"metadata":{"_uuid":"14e59926bdb58188ab43a548fab42f08f450e1be"},"cell_type":"markdown","source":"![](https://kaggle2.blob.core.windows.net/competitions/kaggle/10737/logos/thumb76_76.png?t=2018-10-22-18-50-58t=2018-10-24-17-14-05)\n<h2 style=\"color:  #aa2200;text-align: center;\" id=\"Quora-Insincere-Questions-Classification\">Quora Insincere Questions Classification</h2>\n<ul style=\"align:justify;color: #aa2200;\">\n <li><span style=\"text-color: #aa2200;\">An existential problem for any major website today is how to handle toxic and divisive content. Quora wants to tackle this problem head-on to keep their platform a place where users can feel safe sharing their knowledge with the world.<a style=\"color: #aa2200;\" href=\"https://www.quora.com/\" rel=\"nofollow\">Quora</a>&nbsp;is a platform that empowers people to learn from each other. On Quora, people can ask questions and connect with others who contribute unique insights and quality answers. A key challenge is to weed out insincere questions -- those founded upon false premises, or that intend to make a statement rather than look for helpful answers.</span></li>\n<li><span style=\"text-color: #aa2200;\">In this competition, Kagglers will develop models that identify and flag insincere questions. To date, Quora has employed both machine learning and manual review to address this problem. With your help, they can develop more scalable methods to detect toxic and misleading content.</span></li>\n<li><span style=\"text-color: #aa2200;\">Here's your chance to combat online trolls at scale. Help Quora uphold their policy of &ldquo;Be Nice, Be Respectful&rdquo; and continue to be a place for sharing and growing the world&rsquo;s knowledge.</span></li>\n</ul>\n<hr>\n<h2 style=\"color: #aa2200;text-align: Left;\" id=\"Quora-Insincere-Questions-Classification\">Outline of the Notebook</h2>\n<hr>\n<h4 style=\"color: #aa2200;text-align: Left;\" >Objective : Predict whether a question asked on Quora is sincere or not. This is a kernels only comeptition.</h4>\n<ul>\n<li><a href=\"#1.Data-Reading-and-Investigation\" style=\"color:#aa2200;\">1.Data Reading and Investigation</a></li>\n<li><a href=\"#2.Target-variable-Distribution-Check\" style=\"color:#aa2200;\">2.Target variable Distribution Check</a></li>\n<li><a href=\"#3.Word-Cloud-of-Sincere-and-Insincere-Target-Class\" style=\"color:#aa2200;\">3.Word Cloud of Sincere and Insincere Target Class</a></li>\n<li><a href=\"#4.-Word-Count-by-Sincere-Vs-Insincere\" style=\"color:#aa2200;\">4. Word Count by Sincere Vs Insincere</a></li>\n<li><a href=\"#5.-Bi-gram-by-Sincere-Vs-Insincere\" style=\"color:#aa2200;\">5. Bi-gram by Sincere Vs Insincere</a></li>\n<li><a href=\"#6.-Tri-gram-by-Sincere-Vs-Insincere\" style=\"color:#aa2200;\">6. Tri-gram by Sincere Vs Insincere</a></li>\n<li><a href=\"#7.-Meta-Feature-Engineering\" style=\"color:#aa2200;\">7. Meta Feature Engineering</a></li>\n<li><a href=\"#8.-Model-Training-using-Vowpal-Wabbit-Algorithm\" style=\"color:#aa2200;\">8. Model Training using Vowpal Wabbit Algorithm</a></li>\n<li><a href=\"#9.-Model-Training\" style=\"color:#aa2200;\">9. Model Training</a></li>\n<li><a href=\"#10.-F1-Score-Graph\" style=\"color:#aa2200;\">10. F1-Score Graph</a></li>\n<li><a href=\"#11.-Final-Submission\" style=\"color:#aa2200;\">11. Final Submission</a></li>\n</ul>\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# https://www.kaggle.com/sudalairajkumar/simple-exploration-notebook-qiqc\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n# https://www.kaggle.com/arunkumarramanan/market-data-nn-baseline\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nfrom vowpalwabbit import pyvw\nimport sklearn.metrics as metrics\nfrom gensim.parsing import preprocessing as prep\nimport plotly.tools as tls\nimport warnings\n# from plotly.tools import FigureFactory as FF \nwarnings.filterwarnings('ignore')\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nplt.style.use(\"fivethirtyeight\")\n%matplotlib inline\nimport seaborn as sns\nimport numpy as np\nimport plotly.figure_factory as ff\n\n\n######### Function\ndef mis_value_graph(data):\n#     data.isnull().sum().plot(kind=\"bar\", figsize = (20,10), fontsize = 20)\n#     plt.xlabel(\"Columns\", fontsize = 20)\n#     plt.ylabel(\"Value Count\", fontsize = 20)\n#     plt.title(\"Total Missing Value By Column\", fontsize = 20)\n#     for i in range(len(data)):\n#          colors.append(generate_color())\n            \n    data = [\n    go.Bar(\n        x = data.columns,\n        y = data.isnull().sum(),\n        name = 'Unknown Assets',\n        textfont=dict(size=20),\n        marker=dict(\n#         color= colors,\n        line=dict(\n            color=generate_color(),\n            width=2,\n        ), opacity = 0.45\n    )\n    ),\n    ]\n    layout= go.Layout(\n        title= '\"Total Missing Value By Column\"',\n        xaxis= dict(title='Columns', ticklen=5, zeroline=False, gridwidth=2),\n        yaxis=dict(title='Value Count', ticklen=5, gridwidth=2),\n        showlegend=True\n    )\n    fig= go.Figure(data=data, layout=layout)\n    py.iplot(fig, filename='skin')\n    \n\ndef mis_impute(data):\n    for i in data.columns:\n        if data[i].dtype == \"object\":\n            data[i] = data[i].fillna(\"other\")\n        elif (data[i].dtype == \"int64\" or data[i].dtype == \"float64\"):\n            data[i] = data[i].fillna(data[i].mean())\n        else:\n            pass\n    return data\n\n\nimport random\n\ndef generate_color():\n    color = '#{:02x}{:02x}{:02x}'.format(*map(lambda x: random.randint(0, 255), range(3)))\n    return color","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c0fd001524ee68affb1fbc49d31226333cdb7686"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" > 1.Data Reading and Investigation</h3>"},{"metadata":{"trusted":true,"_uuid":"a1766d07c2c1d5cd9c401b9e3a24d664020b3e90"},"cell_type":"code","source":"train = pd.read_csv(\"../input/quora-insincere-questions-classification/train.csv\")\nmis_value_graph(train)\nprint(\"Train Shape:\",train.shape)\ndisplay(train.isna().sum().to_frame())\nprint(\"=====Train Data Column Types=====\")\ndisplay(train.dtypes)\nprint(\"=====Train Data=====\")\ndisplay(train.head())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"30eed18ebe0e0a572ec6d151b8883a860d96aaf4","_kg_hide-input":false},"cell_type":"code","source":"test = pd.read_csv(\"../input/quora-insincere-questions-classification/test.csv\")\nmis_value_graph(train)\nprint(\"Test Shape:\",test.shape)\ndisplay(test.isna().sum().to_frame())\nprint(\"=====Test Data Column Types=====\")\ndisplay(test.dtypes)\nprint(\"=====Test Data=====\")\ndisplay(train.head())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f84e8824d30b21ea1f7ee6c453a0d3f39667de31"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >2.Target variable Distribution Check</h3>"},{"metadata":{"trusted":true,"_uuid":"a7c68f8a091895bc9e62676a8a3c1d62b9e1f007","_kg_hide-input":true},"cell_type":"code","source":"colors = ['#FEBFB3', '#aa2200', '#aa2222', '#aa22aa']\ntrace1 = go.Pie(\nlabels = ['Sincere','Insincere'],\nvalues = train.target.value_counts(),\ntextfont=dict(size=20),\nmarker=dict(colors=colors,line=dict(color='#aa2200', width=2)), hole = 0.45)\nlayout = dict(title = \"Sincere vs Insincere Comments\")\ndata = [trace1]\npy.iplot(dict(data=data, layout=layout), filename='basic-line')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"808f676a8d5037553ac34611c71c1ba902859cc1","_kg_hide-input":true},"cell_type":"code","source":"## target count ##\ncnt_srs = train['target'].value_counts()\ntrace = go.Bar(\n    x=['Sincere','Insincere'],\n    y=cnt_srs.values,\n    marker=dict(\n        color=cnt_srs.values,\n        colorscale = 'RdBu', # ['Greys', 'YlGnBu', 'Greens', 'YlOrRd', 'Bluered', 'RdBu',\n#             'Reds', 'Blues', 'Picnic', 'Rainbow', 'Portland', 'Jet',\n#             'Hot', 'Blackbody', 'Earth', 'Electric', 'Viridis', 'Cividis']       \n        reversescale = True\n    ),\n)\nlayout = go.Layout(\n    title='Target Count',\n    font=dict(size=18)\n)\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig, filename=\"TargetCount\")\ndisplay('We can see that clearly here class imbalance Problem')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9c79e487adf30c5cfe291d3e0a55f50f7fbb81c7"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >3.Word Cloud of Sincere and Insincere Target Class</h3>"},{"metadata":{"trusted":true,"_uuid":"cbf975c660c764ded5fa2619ba1d777fd79d9e5b","_kg_hide-input":true},"cell_type":"code","source":"from wordcloud import WordCloud, STOPWORDS\nfrom PIL import Image\n\n# Thanks : https://www.kaggle.com/aashita/word-clouds-of-various-shapes ##\ndef plot_wordcloud(text, mask=None, max_words=200, max_font_size=100, figure_size=(24.0,16.0), \n                   title = None, title_size=40, image_color=False):\n    stopwords = set(STOPWORDS)\n    more_stopwords = {'one', 'br', 'Po', 'th', 'sayi', 'fo', 'Unknown'}\n    stopwords = stopwords.union(more_stopwords)\n\n    wordcloud = WordCloud(background_color='white',\n                    stopwords = stopwords,\n                    max_words = max_words,\n                    max_font_size = max_font_size, \n                    random_state = 42,\n                    width=500, \n                    height=300,\n                    mask = mask)\n    wordcloud.generate(str(text))\n    \n    plt.figure(figsize=figure_size)\n    if image_color:\n        image_colors = ImageColorGenerator(mask);\n        plt.imshow(wordcloud.recolor(color_func=image_colors), interpolation=\"bilinear\");\n        plt.title(title, fontdict={'size': title_size,  \n                                  'verticalalignment': 'bottom'})\n    else:\n        plt.imshow(wordcloud);\n        plt.title(title, fontdict={'size': title_size, 'color': '#aa2200', \n                                  'verticalalignment': 'bottom'})\n    plt.axis('off');\n    plt.tight_layout()  \n    \n\ncomments_mask = np.array(Image.open(\"../input/quora24/img_44218.png\"))\nplot_wordcloud(train[\"question_text\"], comments_mask,title = 'Question Words Frequency of Quora')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"af63c1527a64d9f8c04bc6b2bc0f54de7ff1fd1f"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >4. Word Count by Sincere Vs Insincere</h3>"},{"metadata":{"trusted":true,"_uuid":"058613d6d390f8b279e118d55be0e134823c1421","_kg_hide-input":true,"scrolled":false},"cell_type":"code","source":"# https://www.kaggle.com/sudalairajkumar/simple-exploration-notebook-qiqc\nfrom collections import defaultdict\ntrain1_df = train[train[\"target\"]==1]\ntrain0_df = train[train[\"target\"]==0]\n\n## custom function for ngram generation ##\ndef generate_ngrams(text, n_gram=1):\n    token = [token for token in text.lower().split(\" \") if token != \"\" if token not in STOPWORDS]\n    ngrams = zip(*[token[i:] for i in range(n_gram)])\n    return [\" \".join(ngram) for ngram in ngrams]\n\n## custom function for horizontal bar chart ##\ndef horizontal_bar_chart(df, color):\n    trace = go.Bar(\n        y=df[\"word\"].values[::-1],\n        x=df[\"wordcount\"].values[::-1],\n        showlegend=False,\n        orientation = 'h',\n        marker=dict(\n            color=[i for j in range(100) for i in ['#aa2200','#2b6dad',generate_color(),generate_color(),generate_color()]],\n        ),\n    )\n    return trace\n\n## Get the bar chart from sincere questions ##\nfreq_dict = defaultdict(int)\nfor sent in train0_df[\"question_text\"]:\n    for word in generate_ngrams(sent):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\nprint(\"Frequency Trigram for Sincere Question\")\ndisplay(fd_sorted.head(10))\ntrace0 = horizontal_bar_chart(fd_sorted.head(30), 'blue')\n\n## Get the bar chart from insincere questions ##\nfreq_dict = defaultdict(int)\nfor sent in train1_df[\"question_text\"]:\n    for word in generate_ngrams(sent):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\nprint(\"Frequency Trigram for Insincere Question\")\ndisplay(fd_sorted.head(10))\ntrace1 = horizontal_bar_chart(fd_sorted.head(50), color = 'blue')\n\n# Creating two subplots\nfig = tls.make_subplots(rows=1, cols=2, vertical_spacing=0.04,subplot_titles=[\"Frequent words of sincere questions\",\"Frequent words of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=900,title=\"Word Count Plots\")\npy.iplot(fig, filename='word-plots')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b6569e2d860965753dc56967dbdce4ad345e8ef1"},"cell_type":"markdown","source":"<h6 style=\"align:justify;color: #aa2200;\"> Words Counts Insights</h6>\n\n<ul style=\"align:justify;color: #aa2200;\">\n<li><span style=\"text-color: #aa2200;\">We can see that their are so many comman words people use to write their question here.</span></li>\n<li><span style=\"text-color: #aa2200;\"><b>In Insincere Graph</b> you can check insincere bar plot <b>`people,trump, women, will, think, many, white, men, indian. muslims` are most used words more than 4k times</b>.</span></li>\n<li><span style=\"text-color: #aa2200;\"><b>In Sincere Graph</b>you can check Sincere bar plot <b>`Best, will, people, good, one, make, think, many, much, someone etc.`  are most used words more than 30k times.</b></span></li>\n</ul>\n<h3 style=\"color: #aa2200;text-align: Left;\" >5. Bi-gram by Sincere Vs Insincere</h3>"},{"metadata":{"trusted":true,"_uuid":"f7aaed035a056c0d4c0c002380c467b91d90cba6","_kg_hide-input":true,"scrolled":false},"cell_type":"code","source":"freq_dict = defaultdict(int)\nfor sent in train0_df[\"question_text\"]:\n    for word in generate_ngrams(sent,2):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\nprint(\"Frequency Trigram for Sincere Question\")\ndisplay(fd_sorted.head(10))\ntrace0 = horizontal_bar_chart(fd_sorted.head(30), generate_color())\n\n\nfreq_dict = defaultdict(int)\nfor sent in train1_df[\"question_text\"]:\n    for word in generate_ngrams(sent,2):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\nprint(\"Frequency Trigram for Insincere Question\")\ndisplay(fd_sorted.head(10))\ntrace1 = horizontal_bar_chart(fd_sorted.head(30), generate_color())\n\n# Creating two subplots\nfig = tls.make_subplots(rows=1, cols=2, vertical_spacing=0.04,horizontal_spacing=0.15,\n                          subplot_titles=[\"Frequent bigrams of sincere questions\", \n                                          \"Frequent bigrams of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=900,title=\"Bigram Word pair Count Plots\")\npy.iplot(fig, filename='word-plots')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"97162b7a1257eb0808be3e2fecdefb4ca2b8a4c5"},"cell_type":"markdown","source":"<h6 style=\"align:justify;color: #aa2200;\"> Bigram Insights</h6>\n<ul style=\"align:justify;color: #aa2200;\">\n<li><span style=\"text-color: #aa2200;\">A <b>bigram</b> is a sequence of two adjacent elements from a string of tokens, which are typically letters, syllables, or words. A bigram is an n-gram for n=2</span></li>\n<li><span style=\"text-color: #aa2200;\"><b>In Sincere Graph</b> you can check Sincere bar plot <b>`Best way, Year old, will happen` are most used words more than 2000 times.</b></span></li>\n<li><span style=\"text-color: #aa2200;\"><b>In Insincere Graph</b> you can check Insincere bar plot <b>`Donald Trump, White Perople, black people, many people etc.` are most used words more than 500 times.</b></span></li></ul>\n\n<h3 style=\"color: #aa2200;text-align: Left;\" >6. Tri-gram by Sincere Vs Insincere</h3>"},{"metadata":{"_kg_hide-input":true,"trusted":true,"scrolled":false,"_uuid":"ccb54a49fef1d82c24b07cee1bbaf98aefc79bd6"},"cell_type":"code","source":"freq_dict = defaultdict(int)\nfor sent in train0_df[\"question_text\"]:\n    for word in generate_ngrams(sent,3):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\nprint(\"Frequency Trigram for Sincere Question\")\ndisplay(fd_sorted.head(10))\ntrace0 = horizontal_bar_chart(fd_sorted.head(30), 'green')\n\n\nfreq_dict = defaultdict(int)\nfor sent in train1_df[\"question_text\"]:\n    for word in generate_ngrams(sent,3):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\nprint(\"Frequency Trigram for Insincere Question\")\ndisplay(fd_sorted.head(10))\ntrace1 = horizontal_bar_chart(fd_sorted.head(30), 'green')\n\n# Creating two subplots\nfig = tls.make_subplots(rows=1, cols=2, vertical_spacing=0.5, horizontal_spacing=0.2,\n                          subplot_titles=[\"Frequent trigrams of sincere questions\", \n                                          \"Frequent trigrams of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=900, title=\"Trigram Count Plots\")\npy.iplot(fig, filename='word-plots')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"441344e9a090ffa006318a924e3c7cf9506a2fc0"},"cell_type":"markdown","source":"<h6 style=\"align:justify;color: #aa2200;\"> Trigram Insights</h6>\n<ul style=\"align:justify;color: #aa2200;\">\n<li><span style=\"text-color: #aa2200;\">A <b>Trigra</b> is a sequence of <b>three adjacent elements</b>from a <b>string of tokens</b>, which are typically <b>letters, syllables, or words</b>. A <b>Trigram</b>is an n-gram for <b>n=3</b></span></li>\n<li><span style=\"text-color: #aa2200;\"><b>In Sincere Graph</b> you can check Sincere bar plot <b>`tips someone starting, someone starting work, useful tips someone, advice give someone,short-term business travelers, hotels short-term business, good hotels short-term, give someone moving, good bad neighborhoods, best known for? etc.`*are most used words more than 400 times.</b></span></li>\n<li><span style=\"text-color: #aa2200;\"><b>In Insincere Graph</b> you can check Sincere bar plot <b>`will donald trump, black lives matter, long will take, kim jong un,12 year old, people still believe,14 year old,united states america.`are most used words more than 30 times.</b></span></li></ul>"},{"metadata":{"_uuid":"edf997c30ef3d1b14e0d15e8165c541b79d21c7e"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >7. Meta Feature Engineering</h3>\n<p style=\"color: #aa2200;text-align: Left;\">(From SRK's Diary)</p>\n\n<ul style=\"align:justify;color: #aa2200;\">\n<p style=\"color: #aa2200;text-align: Left;\">Now let us create some meta features and then look at how they are distributed between the classes. The ones that we will create are</p>\n<li><span style=\"text-color: #aa2200;\">Number of words in the text       </span></li>\n<li><span style=\"text-color: #aa2200;\">Number of unique words in the text</span></li>\n<li><span style=\"text-color: #aa2200;\">Number of characters in the text  </span></li>\n<li><span style=\"text-color: #aa2200;\">Number of stopwords               </span></li>\n<li><span style=\"text-color: #aa2200;\">Number of punctuations            </span></li>\n<li><span style=\"text-color: #aa2200;\">Number of upper case words        </span></li>\n<li><span style=\"text-color: #aa2200;\">Number of title case words        </span></li>\n<li><span style=\"text-color: #aa2200;\">Average length of the words       </span></li>\n</ul>"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"589fc7d529cb3678ec7b08f48452d887a230fa2e"},"cell_type":"code","source":"import string\n## Number of words in the text ##\ntrain[\"num_words\"] = train[\"question_text\"].apply(lambda x: len(str(x).split()))\ntest[\"num_words\"] = test[\"question_text\"].apply(lambda x: len(str(x).split()))\n\n## Number of unique words in the text ##\ntrain[\"num_unique_words\"] = train[\"question_text\"].apply(lambda x: len(set(str(x).split())))\ntest[\"num_unique_words\"] = test[\"question_text\"].apply(lambda x: len(set(str(x).split())))\n\n## Number of characters in the text ##\ntrain[\"num_chars\"] = train[\"question_text\"].apply(lambda x: len(str(x)))\ntest[\"num_chars\"] = test[\"question_text\"].apply(lambda x: len(str(x)))\n\n## Number of stopwords in the text ##\ntrain[\"num_stopwords\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\ntest[\"num_stopwords\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\n\n## Number of punctuations in the text ##\ntrain[\"num_punctuations\"] =train['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\ntest[\"num_punctuations\"] =test['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\n\n## Number of title case words in the text ##\ntrain[\"num_words_upper\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\ntest[\"num_words_upper\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\n\n## Number of title case words in the text ##\ntrain[\"num_words_title\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\ntest[\"num_words_title\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\n\n## Average length of the words in the text ##\ntrain[\"mean_word_len\"] = train[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ntest[\"mean_word_len\"] = test[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\n\nprint(\"Train shape:\",train.shape)\nprint(\"Test Shape:\",test.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7ea814ca651a746e7e01493f890236d0e6ce1adb"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >8. Model Training using Vowpal Wabbit Algorithm</h3>\n<ul style=\"align:justify;color: #aa2200;\">\n<li><span style=\"text-color: #aa2200;\">Vowpal Wabbit. Vowpal Wabbit (also known as \"VW\") is an open source fast out-of-core machine learning system library and program developed originally at Yahoo! Research, and currently at Microsoft Research. It was started and is led by John Langford.</span></li>\n<li><span style=\"text-color: #aa2200;\">The Vowpal Wabbit (VW) is a project started at Yahoo! Research and now sponsored by Microsoft Research. Started and led by John Langford, VW focuses on fast learning by building an intrinsically fast learning algorithm. John gave two guest lectures to us on AllReduce and Bandits during NYU Big Data class this semester. From what I see, he is a reputed researcher and really passionate about online learning algorithm.</span></li>\n<li><span style=\"text-color: #aa2200;\">The Vowpal Wabbit name is oddly pronounced and strange. Langford explained that Vowpal Wabbit is how Elmer Fudd would pronounce “Vorpal Rabbit”. As for “Vorpal”, if you Google it, you will find “Vorpal Bunny”, which is also known as a “killer rabbit” in a popular computer game. Maybe this is exactly what he wants VW to be – cute but also powerful and fast.</span></li>\n<li><span style=\"text-color: #aa2200;\">VW supports a number of machine learning problems, importance weighting, a selection of loss functions and optimization algorithms, like SGD (Stochastic Gradient Descent), BFGS (a a popular algorithm for parameter estimation), conjugate gradient etc. It has been used to learn a sparse terafeature (i.e. 1012 sparse features) dataset on 1000 nodes in one hour, which beats all current machine linear learning algorithms. According to its tutorial on John Langford’s GitHub, VW is about a factor of 3 faster than svmsgd on the RCV1 example, which is a collection for text categorization. </span></li>\n<li><span style=\"text-color: #aa2200;\">For More Reading : https://www.zinkov.com/posts/2013-08-13-vowpal-tutorial/</span></li>\n</ul>\n\n\n"},{"metadata":{"trusted":true,"_uuid":"ecce2445c54e0826403f0360627c71896ac5f399"},"cell_type":"code","source":"# https://www.kaggle.com/hippskill/vowpal-wabbit-starter-pack\nclass Tokenizer(object):\n    def __call__(self, doc): \n        striped = prep.strip_punctuation(doc)\n        striped = prep.strip_tags(striped)\n        striped = prep.strip_multiple_whitespaces(striped).lower()\n        return striped\n    \nclass FilterRareWords(object):\n    def __init__(self):\n        self.cv = defaultdict(int)\n    def fit(self, texts):\n        for text in texts:\n            for word in text.split():\n                self.cv[word] += 1\n    def __call__(self, text):\n        return ' '.join([self.filter_word(word) for word in text.split()])\n    def filter_word(self, word):\n        return '' if self.cv[word] < 2 else word\n\ntokenizer = Tokenizer()\nfilter_words = FilterRareWords()\n\ndisplay(train[train['target'] == 1].head())\n\ntrain['question_text'] = train['question_text'].apply(tokenizer)\n\nfilter_words.fit(train['question_text'])\ntrain['question_text'] = train['question_text'].apply(filter_words)\npos_weight = train['target'].sum() / train.shape[0]\ndisplay(train.head())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d2aaea8fca390552f1027567185462fe7c24c6e7"},"cell_type":"code","source":"def make_vw_feature_line(label, importance, text):\n    return '{} {} |text {}'.format(label, importance, text)\n\ndef make_vw_corpus(texts, labels):\n    for text, label in zip(texts, labels):\n        if label == 1.0:\n            cur_feautre = make_vw_feature_line('1', 1 - pos_weight, text)\n        else:\n            cur_feautre = make_vw_feature_line('-1', pos_weight, text)\n        yield cur_feautre","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"35a7e4225501fed23cef591dc7e088d1494a9c38"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nX_train, X_test = train_test_split(train, test_size=0.1, shuffle=True, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ecf5ec4910feee5da48a6f3d71b1190c9146aaba"},"cell_type":"code","source":"vw = pyvw.vw(\n    quiet=True,\n    loss_function='logistic',\n    link='logistic',\n    b=29,\n    ngram=2,\n    skips=1,\n    random_seed=42,\n    l1=3.4742122764e-09,\n    l2=1.24232077629e-11,\n    learning_rate=0.751849318433,\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5364cdd463b7a865003e65bd083048a6a4ec5f53"},"cell_type":"code","source":"def get_pred(feature):\n    ex = vw.example(feature)\n    pred = vw.predict(ex)\n    ex.finish()\n    return pred","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"42c053ec1bbc47b64282ae6521e094abf5c41ccd"},"cell_type":"code","source":"feature_map =  ['question_text', 'num_words', 'num_unique_words',\n       'num_chars', 'num_stopwords', 'num_punctuations', 'num_words_upper',\n       'num_words_title', 'mean_word_len']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"efafed8888a0e7d254950a857a25bde3d4ffc9e0"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >9. Model Training</h3>"},{"metadata":{"trusted":true,"_uuid":"6b5643f65a74ea1d4fa72324ad81225845f2a492"},"cell_type":"code","source":"%%time\nfor fit_iter in range(7):\n    for num, feature in enumerate(make_vw_corpus(X_train['question_text'], X_train['target'])):\n        ex = vw.example(feature)\n        vw.learn(ex)\n        ex.finish()\n        \n    print('pass num {} done'.format(fit_iter))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"60373a7176e45de4f53ed838637058c007e09888"},"cell_type":"code","source":"pred = np.array([get_pred(x) for x in make_vw_corpus(X_test['question_text'], X_test['target'])])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ea073fe5bc58085256c423fcaed69660774bd6ff"},"cell_type":"code","source":"thresholds = np.linspace(0, 1, 100)\nf1_scores = [metrics.f1_score(X_test['target'], pred > threshold) for threshold in thresholds]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"504cf62e6c94bbc148d35a06a48533ea74e78fca"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >10. F1-Score Graph</h3>"},{"metadata":{"trusted":true,"_uuid":"30ccf97d9c16259bcafdbe77c858621b03e410a7"},"cell_type":"code","source":"plt.figure(figsize=(20,8))\nplt.plot(thresholds, f1_scores)\nplt.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7d87bd71e38f6535af1bb366f2f714edf6337f91"},"cell_type":"code","source":"print('best f1 score is {} with threshold {}'.format(np.max(f1_scores), thresholds[np.argmax(f1_scores)]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"494437d0f469c992edd01669a43c841918749bd6"},"cell_type":"markdown","source":"<h3 style=\"color: #aa2200;text-align: Left;\" >11. Final Submission</h3>"},{"metadata":{"trusted":true,"_uuid":"f02644ef1d47aea6d38ecf5083e109dab4744077"},"cell_type":"code","source":"test['question_text'] = test['question_text'].apply(tokenizer)\ntest['question_text'] = test['question_text'].apply(filter_words)\n\npred = np.array([get_pred(x) for x in make_vw_corpus(test['question_text'], [1] * len(test))])\n\nexample = pd.read_csv('../input/quora-insincere-questions-classification/sample_submission.csv')\nexample['prediction'] = (pred > thresholds[np.argmax(f1_scores)]).astype(int)\nexample.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"08ad46d430858d4239ace6e4b60bf88e01e85cc0"},"cell_type":"markdown","source":"### Suggestions are welcome...I want to improve this kernel more...Upvote if you like it!!!\n### Thanks for reading🙏🙏🙏"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}