{"cells":[{"metadata":{"_uuid":"521178d5e7c0e427d5aff12df5d6b64aacc495ec"},"cell_type":"markdown","source":"# About the dataset\nAn existential problem for any major website today is how to handle toxic and divisive content. Quora wants to tackle this problem head-on to keep their platform a place where users can feel safe sharing their knowledge with the world.\n\nQuora is a platform that empowers people to learn from each other. On Quora, people can ask questions and connect with others who contribute unique insights and quality answers. A key challenge is to weed out insincere questions -- those founded upon false premises, or that intend to make a statement rather than look for helpful answers.\n\nIn this competition, Kagglers will develop models that identify and flag insincere questions. To date, Quora has employed both machine learning and manual review to address this problem. With your help, they can develop more scalable methods to detect toxic and misleading content.\n\nHere's your chance to combat online trolls at scale. Help Quora uphold their policy of “Be Nice, Be Respectful” and continue to be a place for sharing and growing the world’s knowledge."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Usual Imports\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport string\nimport random\nimport operator\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.decomposition import NMF, LatentDirichletAllocation, TruncatedSVD\nfrom statistics import *\nfrom sklearn.feature_extraction.text import CountVectorizer\nimport concurrent.futures\nimport time\nimport pyLDAvis.sklearn\nfrom pylab import bone, pcolor, colorbar, plot, show, rcParams, savefig\nimport textstat\nimport warnings\nimport nltk\nwarnings.filterwarnings('ignore')\n\n%matplotlib inline\nimport os\nprint(os.listdir(\"../input\"))\n\n# Plotly based imports for visualization\nfrom plotly import tools\nimport plotly.plotly as py\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.figure_factory as ff\n\n# spaCy based imports\nimport spacy\nfrom spacy.lang.en.stop_words import STOP_WORDS\nfrom spacy.lang.en import English\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0e2da97e5cd6e24ce093f3e6ccfa5bb514b5fb49"},"cell_type":"code","source":"quora_train = pd.read_csv(\"../input/train.csv\")\nquora_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"239f186671cd2dcea82cf3e34b994eb29fab5279"},"cell_type":"code","source":"# SpaCy Parser for questions\npunctuations = string.punctuation\nstopwords = list(STOP_WORDS)\n\nparser = English()\ndef spacy_tokenizer(sentence):\n    mytokens = parser(sentence)\n    mytokens = [ word.lemma_.lower().strip() if word.lemma_ != \"-PRON-\" else word.lower_ for word in mytokens ]\n    mytokens = [ word for word in mytokens if word not in stopwords and word not in punctuations ]\n    mytokens = \" \".join([i for i in mytokens])\n    return mytokens","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"57656d248a9e4c3972c812c61d68c8d215704732"},"cell_type":"code","source":"tqdm.pandas()\nsincere_questions = quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(spacy_tokenizer)\ninsincere_questions = quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(spacy_tokenizer)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9dfcbf580ef8ca2e5e4d12e52450d80f3b5f2bdc"},"cell_type":"markdown","source":"# Statistics for the given data\nWe will use the ```textstat``` package by kaggler Shivam Bansal(@shivamb) for this purpose. "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"1dcafbcda483ca683367a2166060b5b023d1177e"},"cell_type":"code","source":"# One function for all plots\ndef plot_readability(a,b,title,bins=0.1,colors=['#3A4750', '#F64E8B']):\n    trace1 = ff.create_distplot([a,b], [\"Sincere questions\",\"Insincere questions\"], bin_size=bins, colors=colors, show_rug=False)\n    trace1['layout'].update(title=title)\n    iplot(trace1, filename='Distplot')\n    table_data= [[\"Statistical Measures\",\"Sincere questions\",\"Insincere questions\"],\n                [\"Mean\",mean(a),mean(b)],\n                [\"Standard Deviation\",pstdev(a),pstdev(b)],\n                [\"Variance\",pvariance(a),pvariance(b)],\n                [\"Median\",median(a),median(b)],\n                [\"Maximum value\",max(a),max(b)],\n                [\"Minimum value\",min(a),min(b)]]\n    trace2 = ff.create_table(table_data)\n    iplot(trace2, filename='Table')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2140035bf3df84d6489036653a6eaccb18911b7b"},"cell_type":"markdown","source":"## 1. Syllable Analysis"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"f7d0eb537dc163cc89ac96c1aaa0d46020a6fcb7"},"cell_type":"code","source":"syllable_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.syllable_count))\nsyllable_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.syllable_count))\nplot_readability(syllable_sincere,syllable_insincere,\"Syllable Analysis\",5)\n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cb6470a612cd1ccbfee2ffd52c48d96f6d6bc719"},"cell_type":"markdown","source":"## 2. Lexicon Analysis"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"23e1f6ef520b477f539ae8c7e4437508040ed957"},"cell_type":"code","source":"lexicon_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.lexicon_count))\nlexicon_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.lexicon_count))\nplot_readability(lexicon_sincere,lexicon_insincere,\"Lexicon Analysis\",4,['#C65D17','#DDB967'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0fd53c188a05c89d3174691b442e4fd82ff49dbb"},"cell_type":"markdown","source":"## 3. Question length"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"ff4d9e201b7923a89b2a1351c1f32461e97ddade"},"cell_type":"code","source":"length_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(len))\nlength_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(len))\nplot_readability(length_sincere,length_insincere,\"Question Length\",40,['#C65D17','#DDB967'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"87996e35f4164a8330d62effb12fce40f275b42f"},"cell_type":"markdown","source":"## 4. Average syllables per word in a question "},{"metadata":{"trusted":true,"_uuid":"99fd809d3687c950a050ef2752a54237679b362a","_kg_hide-input":true},"cell_type":"code","source":"spw_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.avg_syllables_per_word))\nspw_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.avg_syllables_per_word))\nplot_readability(spw_sincere,spw_insincere,\"Average syllables per word\",0.2,['#8D99AE','#EF233C'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"22c95c6df254cf22a9beed663aca06ab26502f6b"},"cell_type":"markdown","source":"## 5. Average letters per word in a question "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"5dca43a3c100c8935ae53c058869fa551d515bf3"},"cell_type":"code","source":"lpw_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.avg_letter_per_word))\nlpw_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.avg_letter_per_word))\nplot_readability(lpw_sincere,lpw_insincere,\"Average letters per word\",2,['#8491A3','#2B2D42'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c6348beb1f227eaa52133d2b0f673aa027f5a988"},"cell_type":"markdown","source":"## 6. Readability features\nThis basically returns the readability statistics for given text. "},{"metadata":{"_uuid":"29a2b90e03437a64ded6bc13988f6c9ed1b4d0b2"},"cell_type":"markdown","source":"### 6.1 The Flesch Reading Ease formula\nThe following table can be helpful to assess the ease of readability in a document.\n###    Score\t- Difficulty\n* 90-100 - Very Easy\n* 80-89 -\tEasy\n* 70-79 -\tFairly Easy\n* 60-69 -\tStandard\n* 50-59 -\tFairly Difficult\n* 30-49 -\tDifficult\n* 0-29 -\tVery Confusing\n\nRead More: [Wikipedia](https://en.wikipedia.org/wiki/Flesch%E2%80%93Kincaid_readability_tests)"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"c8b2e765bd0f00c616010a4f61dafce6042022df"},"cell_type":"code","source":"fre_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.flesch_reading_ease))\nfre_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.flesch_reading_ease))\nplot_readability(fre_sincere,fre_insincere,\"Flesch Reading Ease\",20)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d507510e6dad4992673b4d7986e09f43d304ac6e"},"cell_type":"markdown","source":"## 6.2 The Flesch-Kincaid Grade Level\nReturns the Flesch-Kincaid Grade of the given text. This is a grade formula in that a score of 9.3 means that a ninth grader would be able to read the document.\n\nRead More: [Wikipedia](https://en.wikipedia.org/wiki/Flesch%E2%80%93Kincaid_readability_tests)"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"041cee01087b1850d5fe3e2668b4204d04c02ea8"},"cell_type":"code","source":"fkg_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.flesch_kincaid_grade))\nfkg_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.flesch_kincaid_grade))\nplot_readability(fkg_sincere,fkg_insincere,\"Flesch Kincaid Grade\",4,['#C1D37F','#491F21'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ae6bbc8fb0e503d5a783a5d41b4e34266b902528"},"cell_type":"markdown","source":"## 6.3 The Fog Scale (Gunning FOG Formula)\nReturns the FOG index of the given text. This is a grade formula in that a score of 9.3 means that a ninth grader would be able to read the document.\n\nRead More: [Wikipedia](https://en.wikipedia.org/wiki/Gunning_fog_index)"},{"metadata":{"trusted":true,"_uuid":"afabe851daf96813b0b0d9c02ad7ba67edadaa32","_kg_hide-input":true},"cell_type":"code","source":"fog_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.gunning_fog))\nfog_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.gunning_fog))\nplot_readability(fog_sincere,fog_insincere,\"The Fog Scale (Gunning FOG Formula)\",4,['#E2D58B','#CDE77F'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3bd9752f79307430a91e93dd1b10c03a64d22ff2"},"cell_type":"markdown","source":"## 6.4 Automated Readability Index\nReturns the ARI (Automated Readability Index) which outputs a number that approximates the grade level needed to comprehend the text.For example if the ARI is 6.5, then the grade level to comprehend the text is 6th to 7th grade.\n\nRead More: [Wikipedia](https://en.wikipedia.org/wiki/Automated_readability_index)"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"e2c17b2d23dfbf84ac19da580ca346337c77d098"},"cell_type":"code","source":"ari_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.automated_readability_index))\nari_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.automated_readability_index))\nplot_readability(ari_sincere,ari_insincere,\"Automated Readability Index\",10,['#488286','#FF934F'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2de5019ce4b05dd23732cc03d5e9de81931a3a1f"},"cell_type":"markdown","source":"## 6.5 The Coleman-Liau Index\nReturns the grade level of the text using the Coleman-Liau Formula. This is a grade formula in that a score of 9.3 means that a ninth grader would be able to read the document.\n\nRead More: [Wikipedia](https://en.wikipedia.org/wiki/Coleman%E2%80%93Liau_index)"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"b9dc34b66bd79f547eb3ecb837f0d35aa9ea3af8"},"cell_type":"code","source":"cli_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.coleman_liau_index))\ncli_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.coleman_liau_index))\nplot_readability(cli_sincere,cli_insincere,\"The Coleman-Liau Index\",10,['#8491A3','#2B2D42'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"46f10b49383d2c232f7849c22666a62662f15763"},"cell_type":"markdown","source":"## 6.6 Linsear Write Formula\nReturns the grade level of the text using the Coleman-Liau Formula. This is a grade formula in that a score of 9.3 means that a ninth grader would be able to read the document.\n\nRead More: [Wikipedia](https://en.wikipedia.org/wiki/Linsear_Write)"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"1aebd06cdcc79da444c0d9c4858c0734546e2038"},"cell_type":"code","source":"lwf_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.linsear_write_formula))\nlwf_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.linsear_write_formula))\nplot_readability(lwf_sincere,lwf_insincere,\"Linsear Write Formula\",2,['#8D99AE','#EF233C'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"308bad057fdbab108d177e4cf6362599ed84a12d"},"cell_type":"markdown","source":"## 6.7 Dale-Chall Readability Score\nDifferent from other tests, since it uses a lookup table of the most commonly used 3000 English words. Thus it returns the grade level using the New Dale-Chall Formula.\n\n**Score** - **Understood by**\n* 4.9 or lower - average 4th-grade student or lower\n* 5.0–5.9\t- average 5th or 6th-grade student\n* 6.0–6.9\t- average 7th or 8th-grade student\n* 7.0–7.9\t- average 9th or 10th-grade student\n* 8.0–8.9\t- average 11th or 12th-grade student\n* 9.0–9.9\t- average 13th to 15th-grade (college) student\n\nRead More: [Wikipedia](https://en.wikipedia.org/wiki/Dale%E2%80%93Chall_readability_formula)"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"ce38049fb5b35e7aa2c7eb53ec678d14220a6ad2"},"cell_type":"code","source":"dcr_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(textstat.dale_chall_readability_score))\ndcr_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(textstat.dale_chall_readability_score))\nplot_readability(dcr_sincere,dcr_insincere,\"Dale-Chall Readability Score\",1,['#C65D17','#DDB967'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2e71396cf26be5d953c94386d0bdf67cb793e443"},"cell_type":"markdown","source":"## 6.8 Readability Consensus based upon all the above tests\nBased upon all the above tests, returns the estimated school grade level required to understand the text."},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"c657365e26b56db23819db12907534a363c5d579"},"cell_type":"code","source":"def consensus_all(text):\n    return textstat.text_standard(text,float_output=True)\n\ncon_sincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 0].progress_apply(consensus_all))\ncon_insincere = np.array(quora_train[\"question_text\"][quora_train[\"target\"] == 1].progress_apply(consensus_all))\nplot_readability(con_sincere,con_insincere,\"Readability Consensus based upon all the above tests\",2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5fb406aa01ae4e1c849c5d7a16fb6edbf344f9f3"},"cell_type":"markdown","source":"# N-gram analysis\nIn the fields of computational linguistics and probability, an n-gram is a contiguous sequence of n items from a given sample of text or speech. The items can be phonemes, syllables, letters, words or base pairs according to the application. The n-grams typically are collected from a text or speech corpus. When the items are words, n-grams may also be called shingles."},{"metadata":{"trusted":true,"_uuid":"2e33f94f14f96679a59e0dca18f5a62e88066f8e"},"cell_type":"code","source":"def word_generator(text):\n    word = list(text.split())\n    return word\ndef bigram_generator(text):\n    bgram = list(nltk.bigrams(text.split()))\n    bgram = [' '.join((a, b)) for (a, b) in bgram]\n    return bgram\ndef trigram_generator(text):\n    tgram = list(nltk.trigrams(text.split()))\n    tgram = [' '.join((a, b, c)) for (a, b, c) in tgram]\n    return tgram\nsincere_words = sincere_questions.progress_apply(word_generator)\ninsincere_words = insincere_questions.progress_apply(word_generator)\nsincere_bigrams = sincere_questions.progress_apply(bigram_generator)\ninsincere_bigrams = insincere_questions.progress_apply(bigram_generator)\nsincere_trigrams = sincere_questions.progress_apply(trigram_generator)\ninsincere_trigrams = insincere_questions.progress_apply(trigram_generator)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5d77455bdb80a87c169aca723a856ba4d23d7461","_kg_hide-input":true},"cell_type":"code","source":"color_brewer = ['#57B8FF','#B66D0D','#009FB7','#FBB13C','#FE6847','#4FB5A5','#8C9376','#F29F60','#8E1C4A','#85809B','#515B5D','#9EC2BE','#808080','#9BB58E','#5C0029','#151515','#A63D40','#E9B872','#56AA53','#CE6786','#449339','#2176FF','#348427','#671A31','#106B26','#008DD5','#034213','#BC2F59','#939C44','#ACFCD9','#1D3950','#9C5414','#5DD9C1','#7B6D49','#8120FF','#F224F2','#C16D45','#8A4F3D','#616B82','#443431','#340F09']\n\ndef ngram_visualizer(v,t):\n    X = v.values\n    Y = v.index\n    trace = [go.Bar(\n                y=Y,\n                x=X,\n                orientation = 'h',\n                marker=dict(color=color_brewer, line=dict(color='rgb(8,48,107)',width=1.5,)),\n                opacity = 0.6\n    )]\n    layout = go.Layout(\n        title=t,\n        margin = go.Margin(\n            l = 200,\n            r = 400\n        )\n    )\n\n    fig = go.Figure(data=trace, layout=layout)\n    iplot(fig, filename='horizontal-bar')\n    \ndef ngram_plot(ngrams,title):\n    ngram_list = []\n    for i in tqdm(ngrams.values, total=ngrams.shape[0]):\n        ngram_list.extend(i)\n    random.shuffle(color_brewer)\n    ngram_visualizer(pd.Series(ngram_list).value_counts()[:20],title)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d82e0ca33b3e827d2a0c3768bc7d3ea90f3625b9"},"cell_type":"code","source":"# Top Sincere words\nngram_plot(sincere_words,\"Top Sincere Words\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ccd33b9977ca7cec79716e5e229f64876dcd9ce2"},"cell_type":"code","source":"# Top Insincere words\nngram_plot(insincere_words,\"Top Insincere Words\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d62a75afab2d23315949701adf592a907337c456","_kg_hide-input":true},"cell_type":"code","source":"# Sincere Bigrams\nngram_plot(sincere_bigrams,\"Top 20 Sincere Bigrams\")","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"ceea79a3e10c97be7e55244c8c4d90d69c82a567"},"cell_type":"code","source":"# Insincere Bigrams\nngram_plot(insincere_bigrams,\"Top 20 Insincere Bigrams\")","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"scrolled":true,"_uuid":"d30155acff09e0a5a8e63fe19bbae29c3bc8cccd"},"cell_type":"code","source":"# Sincere Trigrams\nngram_plot(sincere_trigrams,\"Top 20 Sincere Trigrams\")","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"b96fecabbdc4a314dd69024934ebdafe1e9f86ef"},"cell_type":"code","source":"# Insincere Trigrams\nngram_plot(insincere_trigrams,\"Top 20 Insincere Trigrams\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5e0d0f7c353809e4349e7476ff36adf02cb12cd0"},"cell_type":"markdown","source":"# What is topic-modelling?\nIn machine learning and natural language processing, a topic model is a type of statistical model for discovering the abstract \"topics\" that occur in a collection of documents. Topic modeling is a frequently used text-mining tool for discovery of hidden semantic structures in a text body. Intuitively, given that a document is about a particular topic, one would expect particular words to appear in the document more or less frequently: \"dog\" and \"bone\" will appear more often in documents about dogs, \"cat\" and \"meow\" will appear in documents about cats, and \"the\" and \"is\" will appear equally in both. A document typically concerns multiple topics in different proportions; thus, in a document that is 10% about cats and 90% about dogs, there would probably be about 9 times more dog words than cat words.\n\nThe \"topics\" produced by topic modeling techniques are clusters of similar words. A topic model captures this intuition in a mathematical framework, which allows examining a set of documents and discovering, based on the statistics of the words in each, what the topics might be and what each document's balance of topics is. It involves various techniques of dimensionality reduction(mostly non-linear) and unsupervised learning like LDA, SVD, autoencoders etc."},{"metadata":{"_kg_hide-input":false,"_uuid":"9eb19d64839d5e2c2ae2bc9075085656aef4b9fa"},"cell_type":"markdown","source":"## Count Vectorizers for the data"},{"metadata":{"trusted":true,"_uuid":"289b9c8dd4ec1a0b40b384e84e8be8e19bcaf4b4"},"cell_type":"code","source":"vectorizer_sincere = CountVectorizer(min_df=5, max_df=0.9, stop_words='english', lowercase=True, token_pattern='[a-zA-Z\\-][a-zA-Z\\-]{2,}')\nsincere_questions_vectorized = vectorizer_sincere.fit_transform(sincere_questions)\nvectorizer_insincere = CountVectorizer(min_df=5, max_df=0.9, stop_words='english', lowercase=True, token_pattern='[a-zA-Z\\-][a-zA-Z\\-]{2,}')\ninsincere_questions_vectorized = vectorizer_insincere.fit_transform(insincere_questions)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"af05e0dfe36d400da8acde3208ba199695175919"},"cell_type":"markdown","source":"## Applying Latent Dirichlet Allocation(LDA) models"},{"metadata":{"trusted":true,"_uuid":"f61ccee8c43a7126b99a3965b4a5c7d5b81be0bb"},"cell_type":"code","source":"# Latent Dirichlet Allocation Model\nlda_sincere = LatentDirichletAllocation(n_components=10, max_iter=5, learning_method='online',verbose=True)\nsincere_lda = lda_sincere.fit_transform(sincere_questions_vectorized)\nlda_insincere = LatentDirichletAllocation(n_components=10, max_iter=5, learning_method='online',verbose=True)\ninsincere_lda = lda_insincere.fit_transform(insincere_questions_vectorized)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"44fa86115507320444624f22f84323cc66e3a118"},"cell_type":"markdown","source":"## Printing keywords"},{"metadata":{"trusted":true,"_uuid":"ccac6156cdb0ff6d65d454d8b406ac6e291930e9"},"cell_type":"code","source":"# Functions for printing keywords for each topic\ndef selected_topics(model, vectorizer, top_n=10):\n    for idx, topic in enumerate(model.components_):\n        print(\"Topic %d:\" % (idx))\n        print([(vectorizer.get_feature_names()[i], topic[i])\n                        for i in topic.argsort()[:-top_n - 1:-1]]) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6f5e0b2b1905a7f1001085fb2ed10508b1354a41"},"cell_type":"code","source":"# Keywords for topics clustered by Latent Dirichlet Allocation\nprint(\"Sincere questions LDA Model:\")\nselected_topics(lda_sincere, vectorizer_sincere)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"665abde73fd307adcaf029f1371a2703cdc841fb"},"cell_type":"code","source":"print(\"Insincere questions LDA Model:\")\nselected_topics(lda_insincere, vectorizer_insincere)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"83b7e1144594f548ef642629b758170e88e615b9"},"cell_type":"markdown","source":"# Visualizing LDA results of sincere questions with pyLDAvis"},{"metadata":{"trusted":true,"_uuid":"3627bfb0dea9129369b7d0ab158026507ffdf2b2"},"cell_type":"code","source":"pyLDAvis.enable_notebook()\ndash = pyLDAvis.sklearn.prepare(lda_sincere, sincere_questions_vectorized, vectorizer_sincere, mds='tsne')\ndash","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"512c381221df3dae33f52d15fa98aa231efa2c78"},"cell_type":"markdown","source":"So, the sincere questions mostly deal with topics like education, relationships, work life, product reviews, elements of life etc."},{"metadata":{"_uuid":"b09583a5850f538ee9d72bd6023b514c89257b31"},"cell_type":"markdown","source":"# Visualizing LDA results of insincere questions"},{"metadata":{"trusted":true,"_uuid":"cbac8ad03cf84ebeafa9b022cdaad87fffc5d73b"},"cell_type":"code","source":"pyLDAvis.enable_notebook()\ndash = pyLDAvis.sklearn.prepare(lda_insincere, insincere_questions_vectorized, vectorizer_insincere, mds='tsne')\ndash","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0f89b581ab0517c698117cc3ca7cb1a58570b8f7"},"cell_type":"markdown","source":"The insincere questions, however deal with racism(there is a lot of mention of race here), homosexuality, politics, American elections, religion, terrotism and sex etc.\n\nBut words related to India and America appear in both sincere and insincere questions. Kaggle or quora, we are everywhere. "},{"metadata":{"_uuid":"dd81a0b5437f240e37a31a30a09e811175e83f39"},"cell_type":"markdown","source":"### Show your appreciation by UPVOTES. I welcome suggestions to improve this kernel further."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}