{"cells":[{"metadata":{"_uuid":"9f18ee38147fda56c02178a5b524877918fae805"},"cell_type":"markdown","source":"*[CAVEAT: THIS NOTEBOOK TAKES A QUITE SOME TIME TO LOAD DUE TO BOKEH RENDERING THE PLOTS AND IT ALSO TAKES MORE THEN AN HOUR TO RUN - NOTE IF YOU FORK IT]*\n\n----\n# Introduction\n\nThis kernel will be an exploration into the target variable and how it is distributed across the structure of the training data to see if any potential information or patterns can be gleaned going forward. Since classical treatments of text data normally comes with the challenges of high dimensionality (using term frequencies or term frequency inverse document frequencies), the plan therefore in this kernel is to visually explore the target variable in some lower dimensional space. \n\nWe will explore two methods of representing the data in a lower dimensional space, with the first being using the truncated SVD method to linear reduce the dimensions of the term frequency representation of the text data - i.e. a method known as Latent Semantic Analysis (LSA). The second method would be to utilize document embeddings via the Doc2Vec method, which learns a lower dimensional projection for each question. In these lower dimensional spaces, we can finally then utilize the manifold learning method of the t-Distributed Stochastic Neighbor Embedding (t-SNE) technique to further reduce the dimensionality for target variable visualization.\n\nThe kernel is structured as follows:\n\n1. Text preprocessing on question text via standard NLP\n2. T-SNE on LSA feature space: \n3. T-SNE on Doc2Vec space: \n\nLet's go"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Importing the relevant libraries\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nfrom sklearn.feature_extraction.text import TfidfVectorizer\n\nfrom nltk.tokenize import word_tokenize, sent_tokenize, TweetTokenizer \nfrom nltk.stem import WordNetLemmatizer, PorterStemmer\nfrom nltk.corpus import stopwords\nfrom string import punctuation\n\nimport re\nfrom functools import reduce\n\nimport bokeh.plotting as bp\nfrom bokeh.models import HoverTool, BoxSelectTool\nfrom bokeh.models import ColumnDataSource\nfrom bokeh.plotting import figure, show, output_notebook, reset_output\nfrom bokeh.palettes import d3\nimport bokeh.models as bmo\nfrom bokeh.io import save, output_file\n\n# init_notebook_mode(connected = True)\n# color = sns.color_palette(\"Set2\")\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n%matplotlib inline\n\npd.options.mode.chained_assignment = None\npd.options.display.max_columns = 999\npd.options.display.max_rows = 999","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"Reading in the train and test datasets and inspecting the first 5 rows of the dataset, we see that each Quora question text comes with a label (the \"target\" column) consisting of the binary values of either 1 - insincere or 0 - sincere questions."},{"metadata":{"trusted":true,"_uuid":"15923429f129fc13c2522748ce14e9370e07a968"},"cell_type":"code","source":"train = pd.read_csv(\"../input/train.csv\")\ntest = pd.read_csv(\"../input/test.csv\")\ndisplay(train[train.target == 0].head(3))\ndisplay(train[train.target == 1].head(3))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"339b656a3eaac6cd612b263bbffa7eab4e3504b9"},"cell_type":"markdown","source":"Now, inspecting the class distributions we see that the provided training set and labels are rather imbalanced with about 6% of the data points being insincere, while the remainder 94% being sincere questions."},{"metadata":{"trusted":true,"_uuid":"df9aecc95168e03d7615523992af9c3d3c12205c"},"cell_type":"code","source":"train.target.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"77eede4ec536696d3a6d2e65082029f49ac2dd07"},"cell_type":"markdown","source":"For the purposes of speeding up code execution and modeling in the latter stages, I shall re-balance (naively via random sampling) the dataset set by with equal rows of sincere and insincere data points:"},{"metadata":{"trusted":true,"_uuid":"b21951ebd8571d30adcd529c57902a3bfe74162b"},"cell_type":"code","source":"# Full number of insincere data points\nsample_size = 80_810 \n\n# halved the original size as rendering all the data points was causing lag issues\n# sample_size = int(sample_size/2)\n\n# Rebalancing the training set\ntrain_rebal = train[train.target == 1].sample(sample_size).append(train[train.target == 0].sample(sample_size)).reset_index()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"81156c2485bdff2f80b796fb8314c89853b63262"},"cell_type":"markdown","source":"----\n## 1. Text Processing\n\nIn this section, we arrive at the pre-processing of the question text contained within the training data. The processing applied here are some of the standard NLP steps that one would implement in a text based problem, consisting of:\n\n* Tokenization\n* Stemming or Lemmatization"},{"metadata":{"trusted":true,"_uuid":"2e129d8ebc277aa88eecc5f766a2d413e2608230"},"cell_type":"code","source":"def remove_stopwords(words):\n    \"\"\"\n    Function to remove stopwords from the question text\n    \"\"\"\n    stop_words = set(stopwords.words(\"english\"))\n    return [word for word in words if word not in stop_words]\n\ndef remove_punctuation(text):\n    \"\"\"\n    Function to remove punctuation from the question text\n    \"\"\"\n    return re.sub(r'[^\\w\\s]', '', text)\n\ndef lemmatize_text(words):\n    \"\"\"\n    Function to lemmatize the question text\n    \"\"\"\n    lemmatizer = WordNetLemmatizer()\n    return [lemmatizer.lemmatize(word) for word in words]\n\ndef stem_text(words):\n    \"\"\"\n    Function to stem the question text\n    \"\"\"\n    ps = PorterStemmer()\n    return [ps.stem(word) for word in words]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cb7b3b96be9f2d5e957265c19e0959566729b79e"},"cell_type":"markdown","source":"**Updates - 28/12/18**\n\nBased on the many effective public LSTM/GRU-variant kernels which employ text pre-processing that cleans numbers, add spaces to punctuation and odd characters as well as expanding out contractions, let us employ these methods as well and see how that affects the T-SNE and Doc2Vec plots."},{"metadata":{"_uuid":"e1b91e45620c409d0e67b5e3f3c42c19f606c271"},"cell_type":"markdown","source":"In the cell below, I've defined a whole bunch of punctuation, odd characters and contractions so feel free to unhide it should you wish to see the terms explicitly."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"92f080daf20280fc3a3e28cb03ed42efbb9b0d90"},"cell_type":"code","source":"puncts=['☹', 'Ź', 'Ż', 'ἰ', 'ή', 'Š', '＞', 'ξ','ฉ', 'ั', 'น', 'จ', 'ะ', 'ท', 'ำ', 'ใ', 'ห', '้', 'ด', 'ี', '่', 'ส', 'ุ', 'Π', 'प', 'ऊ', 'Ö', 'خ', 'ب', 'ஜ', 'ோ', 'ட', '「', 'ẽ', '½', '△', 'É', 'ķ', 'ï', '¿', 'ł', '북', '한', '¼', '∆', '≥', '⇒', '¬', '∨', 'č', 'š', '∫', 'ḥ', 'ā', 'ī', 'Ñ', 'à', '▾', 'Ω', '＾', 'ý', 'µ', '?', '!', '.', ',', '\"', '#', '$', '%', '\\\\', \"'\", '(', ')', '*', '+', '-', '/', ':', ';', '<', '=', '>', '@', '[', ']', '^', '_', '`', '{', '|', '}', '~', '“', '”', '’', 'é', 'á', '′', '…', 'ɾ', '̃', 'ɖ', 'ö', '–', '‘', 'ऋ', 'ॠ', 'ऌ', 'ॡ', 'ò', 'è', 'ù', 'â', 'ğ', 'म', 'ि', 'ल', 'ग', 'ई', 'क', 'े', 'ज', 'ो', 'ठ', 'ं', 'ड', 'Ž', 'ž', 'ó', '®', 'ê', 'ạ', 'ệ', '°', 'ص', 'و', 'ر', 'ü', '²', '₹', 'ú', '√', 'α', '→', 'ū', '—', '£', 'ä', '️', 'ø', '´', '×', 'í', 'ō', 'π', '÷', 'ʿ', '€', 'ñ', 'ç', 'へ', 'の', 'と', 'も', '↑', '∞', 'ʻ', '℅''ι', '•', 'ì', '−', 'л', 'я', 'д', 'ل', 'ك', 'م', 'ق', 'ا', '∈', '∩', '⊆', 'ã', 'अ', 'न', 'ु', 'स', '्', 'व', 'ा', 'र', 'त', '§', '℃', 'θ', '±', '≤', 'उ', 'द', 'य', 'ब', 'ट', '͡', '͜', 'ʖ', '⁴', '™', 'ć', 'ô', 'с', 'п', 'и', 'б', 'о', 'г', '≠', '∂', 'आ', 'ह', 'भ', 'ी', '³', 'च', '...', '⌚', '⟨', '⟩', '∖', '˂', 'ⁿ', '⅔', 'న', 'ీ', 'క', 'ె', 'ం', 'ద', 'ు', 'ా', 'గ', 'ర', 'ి', 'చ', 'র', 'ড়', 'ঢ়', 'સ', 'ં', 'ઘ', 'ર', 'ા', 'જ', '્', 'ય', 'ε', 'ν', 'τ', 'σ', 'ş', 'ś', 'س', 'ت', 'ط', 'ي', 'ع', 'ة', 'د', 'Å', '☺', 'ℇ', '❤', '♨', '✌', 'ﬁ', 'て', '„', 'Ā', 'ត', 'ើ', 'ប', 'ង', '្', 'អ', 'ូ', 'ន', 'ម', 'ា', 'ធ', 'យ', 'វ', 'ី', 'ខ', 'ល', 'ះ', 'ដ', 'រ', 'ក', 'ឃ', 'ញ', 'ឯ', 'ស', 'ំ', 'ព', 'ិ', 'ៃ', 'ទ', 'គ', '¢', 'つ', 'や', 'ค', 'ณ', 'ก', 'ล', 'ง', 'อ', 'ไ', 'ร', 'į', 'ی', 'ю', 'ʌ', 'ʊ', 'י', 'ה', 'ו', 'ד', 'ת', 'ᠠ', 'ᡳ', 'ᠰ', 'ᠨ', 'ᡤ', 'ᡠ', 'ᡵ', 'ṭ', 'ế', 'ध', 'ड़', 'ß', '¸', 'ч',  'ễ', 'ộ', 'फ', 'μ', '⧼', '⧽', 'ম', 'হ', 'া', 'ব', 'ি', 'শ', '্', 'প', 'ত', 'ন', 'য়', 'স', 'চ', 'ছ', 'ে', 'ষ', 'য', '়', 'ট', 'উ', 'থ', 'ক', 'ῥ', 'ζ', 'ὤ', 'Ü', 'Δ', '내', '제', 'ʃ', 'ɸ', 'ợ', 'ĺ', 'º', 'ष', '♭', '़', '✅', '✓', 'ě', '∘', '¨', '″', 'İ', '⃗', '̂', 'æ', 'ɔ', '∑', '¾', 'Я', 'х', 'О', 'з', 'ف', 'ن', 'ḵ', 'Č', 'П', 'ь', 'В', 'Φ', 'ỵ', 'ɦ', 'ʏ', 'ɨ', 'ɛ', 'ʀ', 'ċ', 'օ', 'ʍ', 'ռ', 'ք', 'ʋ', '兰', 'ϵ', 'δ', 'Ľ', 'ɒ', 'î', 'Ἀ', 'χ', 'ῆ', 'ύ', 'ኤ', 'ል', 'ሮ', 'ኢ', 'የ', 'ኝ', 'ን', 'አ', 'ሁ', '≅', 'ϕ', '‑', 'ả', '￼', 'ֿ', 'か', 'く', 'れ', 'ő', '－', 'ș', 'ן', 'Γ', '∪', 'φ', 'ψ', '⊨', 'β', '∠', 'Ó', '«', '»', 'Í', 'க', 'வ', 'ா', 'ம', '≈', '⁰', '⁷', 'ấ', 'ũ', '눈', '치', 'ụ', 'å', '،', '＝', '（', '）', 'ə', 'ਨ', 'ਾ', 'ਮ', 'ੁ', '︠', '︡', 'ɑ', 'ː', 'λ', '∧', '∀', 'Ō', 'ㅜ', 'Ο', 'ς', 'ο', 'η', 'Σ', 'ण']\nodd_chars=[ '大','能', '化', '生', '水', '谷', '精', '微', 'ル', 'ー', 'ジ', 'ュ', '支', '那', '¹', 'マ', 'リ', '仲', '直', 'り', 'し', 'た', '主', '席', '血', '⅓', '漢', '髪', '金', '茶', '訓', '読', '黒', 'ř', 'あ', 'わ', 'る', '胡', '南', '수', '능', '广', '电', '总', 'ί', '서', '로', '가', '를', '행', '복', '하', '게', '기', '乡', '故', '爾', '汝', '言', '得', '理', '让', '骂', '野', '比', 'び', '太', '後', '宮', '甄', '嬛', '傳', '做', '莫', '你', '酱', '紫', '甲', '骨', '陳', '宗', '陈', '什', '么', '说', '伊', '藤', '長', 'ﷺ', '僕', 'だ', 'け', 'が', '街', '◦', '火', '团', '表',  '看', '他', '顺', '眼', '中', '華', '民', '國', '許', '自', '東', '儿', '臣', '惶', '恐', 'っ', '木', 'ホ', 'ج', '教', '官', '국', '고', '등', '학', '교', '는', '몇', '시', '간', '업', '니', '本', '語', '上', '手', 'で', 'ね', '台', '湾', '最', '美', '风', '景', 'Î', '≡', '皎', '滢', '杨', '∛', '簡', '訊', '短', '送', '發', 'お', '早', 'う', '朝', 'ش', 'ه', '饭', '乱', '吃', '话', '讲', '男', '女', '授', '受', '亲', '好', '心', '没', '报', '攻', '克', '禮', '儀', '統', '已', '經', '失', '存', '٨', '八', '‛', '字', '：', '别', '高', '兴', '还', '几', '个', '条', '件', '呢', '觀', '《', '》', '記', '宋', '楚', '瑜', '孫', '瀛', '枚', '无', '挑', '剔', '聖', '部', '頭', '合', '約', 'ρ', '油', '腻', '邋', '遢', 'ٌ', 'Ä', '射', '籍', '贯', '老', '常', '谈', '族', '伟', '复', '平', '天', '下', '悠', '堵', '阻', '愛', '过', '会', '俄', '罗', '斯', '茹', '西', '亚', '싱', '관', '없', '어', '나', '이', '키', '夢', '彩', '蛋', '鰹', '節', '狐', '狸', '鳳', '凰', '露', '王', '晓', '菲', '恋', 'に', '落', 'ち', 'ら', 'よ', '悲', '反', '清', '復', '明', '肉', '希', '望', '沒', '公', '病', '配', '信', '開', '始', '日', '商', '品', '発', '売', '分', '子', '创', '意', '梦', '工', '坊', 'ک', 'پ', 'ڤ', '蘭', '花', '羡', '慕', '和', '嫉', '妒', '是', '样', 'ご', 'め', 'な', 'さ', 'い', 'す', 'み', 'ま', 'せ', 'ん', '音', '红', '宝', '书', '封', '柏', '荣', '江', '青', '鸡', '汤', '文', '粵', '拼', '寧', '可', '錯', '殺', '千', '絕', '放', '過', '」', '之', '勢', '请', '国', '知', '识', '产', '权', '局', '標', '點', '符', '號', '新', '年', '快', '乐', '学', '业', '进', '步', '身', '体', '健', '康', '们', '读', '我', '的', '翻', '译', '篇', '章', '欢', '迎', '入', '坑', '有', '毒', '黎', '氏', '玉', '英', '啧', '您', '这', '口', '味', '奇', '特', '也', '就', '罢', '了', '非', '要', '以', '此', '为', '依', '据', '对', '人', '家', '批', '判', '一', '番', '不', '地', '道', '啊', '谢', '六', '佬']\ncontraction_mapping = {\"ain't\": \"is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"can't've\": \"cannot have\", \"'cause\": \"because\", \"could've\": \"could have\", \"couldn't\": \"could not\", \"couldn't've\": \"could not have\",\"didn't\": \"did not\",  \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hadn't've\": \"had not have\", \"hasn't\": \"has not\", \"haven't\": \"have not\",  \"he'd\": \"he would\", \"he'd've\": \"he would have\", \"he'll\": \"he will\", \"he'll've\": \"he will have\", \"he's\": \"he is\", \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\", \"how's\": \"how is\",  \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \"i'd've\": \"i would have\", \"i'll\": \"i will\", \"i'll've\": \"i will have\",\"i'm\": \"i am\", \"i've\": \"i have\", \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\", \"it'll\": \"it will\", \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \"ma'am\": \"madam\", \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \"mustn't've\": \"must not have\", \"needn't\": \"need not\", \"needn't've\": \"need not have\",\"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \"oughtn't've\": \"ought not have\", \"shan't\": \"shall not\",\"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \"she'll've\": \"she will have\", \"she's\": \"she is\", \"should've\": \"should have\", \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\"so's\": \"so as\", \"this's\": \"this is\",\"that'd\": \"that would\", \"that'd've\": \"that would have\",\"that's\": \"that is\", \"there'd\": \"there would\", \"there'd've\": \"there would have\",\"there's\": \"there is\", \"here's\": \"here is\",\"they'd\": \"they would\", \"they'd've\": \"they would have\", \"they'll\": \"they will\", \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \"to've\": \"to have\", \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \"we'll\": \"we will\", \"we'll've\": \"we will have\", \"we're\": \"we are\", \"we've\": \"we have\", \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \"what're\": \"what are\", \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\", \"when've\": \"when have\", \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \"who'll\": \"who will\", \"who'll've\": \"who will have\", \"who's\": \"who is\", \"who've\": \"who have\", \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \"y'all'd\": \"you all would\",\"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\"y'all've\": \"you all have\",\"you'd\": \"you would\", \"you'd've\": \"you would have\", \"you'll\": \"you will\", \"you'll've\": \"you will have\", \"you're\": \"you are\", \"you've\": \"you have\" } ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"59e55a0a2d44aa9e3dd5f38b70c7661beda1b540"},"cell_type":"code","source":"def clean_numbers(x):\n    x = re.sub('[0-9]{5,}', ' ##### ', x)\n    x = re.sub('[0-9]{4}', ' #### ', x)\n    x = re.sub('[0-9]{3}', ' ### ', x)\n    x = re.sub('[0-9]{2}', ' ## ', x)\n    return x\n\ndef punct_add_space(x):\n    x = str(x)\n    for punct in puncts:\n        x = x.replace(punct, f' {punct} ')\n    return x  \n\ndef odd_add_space(x):\n    x = str(x)\n    for odd in odd_chars:\n        x = x.replace(odd, f' {odd} ')\n    return x \n\ndef clean_contractions(text, mapping):\n    specials = [\"’\", \"‘\", \"´\", \"`\"]\n    for s in specials:\n        text = text.replace(s, \"'\")\n    text = ' '.join([mapping[t] if t in mapping else t for t in text.split(\" \")])\n    return text","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d686de1ac1ab35fcb19b850a19ad76eac8b6edd9"},"cell_type":"markdown","source":"Now applying these functions sequentially to the question text field:"},{"metadata":{"trusted":true,"_uuid":"70d07745bb163e14d593e55bf8c285fe5296ec04"},"cell_type":"code","source":"train_rebal[\"question_text\"] = train_rebal[\"question_text\"].apply(lambda x: clean_numbers(x))\ntrain_rebal[\"question_text\"] = train_rebal[\"question_text\"].apply(lambda x: punct_add_space(x))\ntrain_rebal[\"question_text\"] = train_rebal[\"question_text\"].apply(lambda x: odd_add_space(x))\ntrain_rebal[\"question_text\"] = train_rebal[\"question_text\"].apply(lambda x: clean_contractions(x, contraction_mapping))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ca4849af26d775bddcbc5ee0a879acff66674e68"},"cell_type":"code","source":"# With the new updated processing - let's leave out the removal of stopwords, punctuation and lemmatization functions.\n\n# Tokenizing the text\ntrain_rebal[\"question_text\"] = train_rebal[\"question_text\"].apply(lambda x: word_tokenize(x))\n\n# Removing stopwords - Leaving Stopwords\n# train_rebal[\"question_text\"] = train_rebal[\"question_text\"].apply(lambda x: remove_stopwords(x))\n\n# Lemmatizting\n# train_rebal[\"question_text\"] = train_rebal[\"question_text\"].apply(lambda x: lemmatize_text(x))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"13d5e1c154dd31f56f3cea779c6041789ea0532f"},"cell_type":"markdown","source":"Peeking into the newly processed training set, we can see that the resulting text in the \"question text\" columns are now nicely tokenized and lemmatized with all stopwords removed."},{"metadata":{"trusted":true,"_uuid":"1c9bb3697b35e490f3ca4826be210a6fc773fcee"},"cell_type":"code","source":"train_rebal.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5c01cf9a9d07a2c4716639624a00eb464ca2e2f7"},"cell_type":"markdown","source":"----\n## 2. T-SNE applied to Latent Semantic (LSA) space\n\nTo start off we look at the sparse representation of text documents via the Term frequency Inverse document frequency method. What this does is create a matrix representation that upweights locally prevalent but globally rare terms - therefore accounting for the occurence bias when using just term frequencies."},{"metadata":{"_uuid":"bbdcb02c4d266df6e1410d30feab6f65bf537d81"},"cell_type":"markdown","source":"**Tf-idf space**"},{"metadata":{"trusted":true,"_uuid":"598055ac0a301c3beb351c4a5daeafa4219dff4b"},"cell_type":"code","source":"tf_idf_vec = TfidfVectorizer(min_df=3,\n                             max_features = 60_000, #100_000,\n                             analyzer=\"word\",\n                             ngram_range=(1,3), # (1,6)\n                             stop_words=\"english\")\ntf_idf = tf_idf_vec.fit_transform(list(train_rebal[\"question_text\"].map(lambda tokens: \" \".join(tokens))))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0d2f0cab80e3cc59bc7f5e931fcdfba9885f6eb8"},"cell_type":"markdown","source":"Having obtained our tf-idf matrix - a sparse matrix object, we now apply the TruncatedSVD method to first reduce the dimensionality of the Tf-idf matrix to a decomposed feature space, referred to in the community as the LSA (Latent Semantic Analysis) method.\n\nLSA has been one of the classical methods in text that have existed for a while allowing \"concept\" searching of words whereby words which are semantically similar to each other (i.e. have more context) are closer to each other in this space and vice-versa."},{"metadata":{"trusted":true,"_uuid":"ceec50ac5ea16c886e9c612983ee4f80dc55aab0"},"cell_type":"code","source":"# Applying the Singular value decomposition\nfrom sklearn.decomposition import TruncatedSVD\nsvd = TruncatedSVD(n_components=50, random_state=2018)\nsvd_tfidf = svd.fit_transform(tf_idf)\nprint(\"Dimensionality of LSA space: {}\".format(svd_tfidf.shape))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6c1b06fd78f9cd02c441cd21c6682eb45704bb47"},"cell_type":"markdown","source":"Quickly plotting a scatter plot of the first 3 dimensions of the latent semantic space just to get an initial feel for how the target variables are distributed:"},{"metadata":{"trusted":true,"_uuid":"eb8c012dd5588ab20dd580549e032e2341c77b20"},"cell_type":"code","source":"# Showing scatter plots \nfrom mpl_toolkits.mplot3d import Axes3D\nfig = plt.figure(figsize=(16,12))\n\n# Plot models:\nax = Axes3D(fig) \nax.scatter(svd_tfidf[:,0],\n           svd_tfidf[:,1],\n           svd_tfidf[:,2],\n           c=train_rebal.target.values,\n           cmap=plt.cm.winter_r,\n           s=2,\n           edgecolor='none',\n           marker='o')\nplt.title(\"Semantic Tf-Idf-SVD reduced plot of Sincere-Insincere data distribution\")\nplt.xlabel(\"First dimension\")\nplt.ylabel(\"Second dimension\")\nplt.legend()\nplt.xlim(0.0, 0.20)\nplt.ylim(-0.2,0.4)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7f92f4f7483d408f84f1238ebab7c9b96998640a"},"cell_type":"markdown","source":"**Takeaways from the plot**\n\nFrom the above scatter plots, It is apparent that sincere and insincere question data points overlap quite significantly in the LSA semantic space. The data points visually appear to be evenly distributed and there does not seem to be any clear or obvious pattern in segregating the class labels. However, do keep in mind that this involves only the first three dimensions of the decomposed space whilst we haven't incorporated information from the other 48 dimensions yet and also the fact that the SVD is a linear decomposition technique."},{"metadata":{"_uuid":"df5cad545c11965c18a26b4405179030ccdd1211"},"cell_type":"markdown","source":"Perhaps a non-linear technique (T-SNE) yield more insights? This also helps to collapse the information from all 50 dimensions when we apply the T-SNE technique to this LSA reduced space . Here I've used the multicore implementation of T-SNE to speed things up instead of the plain vanilla sklearn version. "},{"metadata":{"trusted":true,"_uuid":"7645bb2cf10a4cf49face7d87a7048d0564b1b64"},"cell_type":"code","source":"# from sklearn.manifold import TSNE\n\n# Importing multicore version of TSNE\nfrom MulticoreTSNE import MulticoreTSNE as TSNE","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a20396d4abcadb72e475e6757b127ee53aff4aac"},"cell_type":"code","source":"tsne_model = TSNE(n_jobs=4,\n                  early_exaggeration=4, # Trying out exaggeration trick\n                  n_components=2,\n                  verbose=1,\n                  random_state=2018,\n                  n_iter=500)\ntsne_tfidf = tsne_model.fit_transform(svd_tfidf)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bc28987c91739421cce2dbb031cef30d86a90168"},"cell_type":"markdown","source":"### Visualization of target variable via Bokeh\n\nTurning to the target variable visualization in T-SNE reduced concept space, we will use the plotting library Bokeh to suit our purposes:"},{"metadata":{"trusted":true,"_uuid":"c799cd1f2e193c57e1529eccea8f5ecb876b914d"},"cell_type":"code","source":"# Putting the tsne information into a dataframe\ntsne_tfidf_df = pd.DataFrame(data=tsne_tfidf, columns=[\"x\", \"y\"])\ntsne_tfidf_df[\"qid\"] = train_rebal[\"qid\"].values\ntsne_tfidf_df[\"question_text\"] = train_rebal[\"question_text\"].values\ntsne_tfidf_df[\"target\"] = train_rebal[\"target\"].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dcef60c9295301c1fa5d2a9f0ec12dc270e8a0a2"},"cell_type":"code","source":"output_notebook()\nplot_tfidf = bp.figure(plot_width = 800, plot_height = 700, \n                       title = \"T-SNE applied to Tfidf_SVD space\",\n                       tools = \"pan, wheel_zoom, box_zoom, reset, hover, previewsave\",\n                       x_axis_type = None, y_axis_type = None, min_border = 1)\n\n# colormap = np.array([\"#6d8dca\", \"#d07d3c\"])\ncolormap = np.array([\"darkblue\", \"red\"])\n\n# palette = d3[\"Category10\"][len(tsne_tfidf_df[\"asset_name\"].unique())]\nsource = ColumnDataSource(data = dict(x = tsne_tfidf_df[\"x\"], \n                                      y = tsne_tfidf_df[\"y\"],\n                                      color = colormap[tsne_tfidf_df[\"target\"]],\n                                      question_text = tsne_tfidf_df[\"question_text\"],\n                                      qid = tsne_tfidf_df[\"qid\"],\n                                      target = tsne_tfidf_df[\"target\"]))\n\nplot_tfidf.scatter(x = \"x\", \n                   y = \"y\", \n                   color=\"color\",\n                   legend = \"target\",\n                   source = source,\n                   alpha = 0.7)\nhover = plot_tfidf.select(dict(type = HoverTool))\nhover.tooltips = {\"qid\": \"@qid\", \n                  \"question_text\": \"@question_text\", \n                  \"target\":\"@target\"}\n\nshow(plot_tfidf)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6eceee3994198e22c99d9ce70f354d2364571222"},"cell_type":"markdown","source":"**Takeaways from the plot**\n\n- It seems that the distribution of the sincere to insincere data points overlap quite substantially in certain regions of the T-SNE plots in concept space, which unfortunately does not allow easy visual discernment between the two classes (even using this non-linear method).\n- This begs the question of how easy therefore, is it in terms of semantic meaning (to a human) to distinguish between an insincere and a sincere question, when we see data from both class labels overlapping quite heavily across each other. Quite a few of the insincere and sincere questions when read aloud do share quite a lot of similarities as well.\n- There is also a popular thread going on in the Discussions/Forums on how the Sincere and Insincere labels were generated - potentially offering the argument that some of the classes could even have been wrongly applied.\n- There does appear to be one region in the T-SNE plot where the sincere class labels visually seem to be quite separated and exist in its own area of the non-linear semantic space. \n\nOne key point to note in T-SNE plots is that cluster sizes as well as distances from one cluster to another are not meaningful to the overall global geometry of the picture. This can be evinced from the variation in geometry when we for example change the perplexity value (defaulted to 30) to range from 5 all the way to 50 and note how this in turn affects the geometry of the manifold :\n"},{"metadata":{"trusted":true,"_uuid":"f12bf74a540d5ac0cd10911cdacf09a6c74fd90b"},"cell_type":"code","source":"# Perplexity = 5\ntsne_model_5 = TSNE(n_jobs=4, \n                    early_exaggeration=4,\n                  perplexity=5,\n                  n_components=2,\n                  verbose=1,\n                  random_state=2018,\n                  n_iter=500)\ntsne_tfidf_5 = tsne_model_5.fit_transform(svd_tfidf[:50_000,:])\n# Creating a Dataframe for Perplexity=5\ntsne_tfidf_df_5 = pd.DataFrame(data=tsne_tfidf_5, columns=[\"x5\", \"y5\"])\ntsne_tfidf_df_5[\"target\"] = train_rebal[\"target\"][:50_000].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1b6d0e0cc0c5f9ea047ad2f2201bcd5b3757ab4b"},"cell_type":"code","source":"# Perplexity = 25\ntsne_model_25 = TSNE(n_jobs=4, \n                     early_exaggeration=4,\n                  perplexity=25,\n                  n_components=2,\n                  verbose=1,\n                  random_state=2018,\n                  n_iter=500)\ntsne_tfidf_25 = tsne_model_25.fit_transform(svd_tfidf[:50_000,:])\n# Creating a Dataframe for Perplexity=5\ntsne_tfidf_df_25 = pd.DataFrame(data=tsne_tfidf_25, \n                             columns=[\"x25\", \"y25\"])\ntsne_tfidf_df_25[\"target\"] = train_rebal[\"target\"][:50_000].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8c0f5611f81b065ed8ce47f03e62136cd973a750"},"cell_type":"code","source":"# Perplexity = 50\ntsne_model_50 = TSNE(n_jobs=4, \n                     early_exaggeration=4,\n                  perplexity=200,\n                  n_components=2,\n                  verbose=1,\n                  random_state=2018,\n                  n_iter=500)\ntsne_tfidf_50 = tsne_model_50.fit_transform(svd_tfidf[:50_000,:])\n# Creating a Dataframe for Perplexity=50\ntsne_tfidf_df_50 = pd.DataFrame(data=tsne_tfidf_50, \n                                columns=[\"x50\", \"y50\"])\ntsne_tfidf_df_50[\"target\"] = train_rebal[\"target\"][:50_000].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"de7d86bf695dc424d42fc930a36659949572ad36"},"cell_type":"code","source":"# Showing scatter plots \nplt.figure(figsize=(14,8))\nplt.scatter(tsne_tfidf_df_5.x5, \n            tsne_tfidf_df_5.y5, \n            alpha=0.75,\n            c=tsne_tfidf_df_5.target,\n            cmap=plt.cm.coolwarm)\nplt.title(\"T-SNE plot in SVD space (perplexity=5)\")\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cfbb4365b787f29996c73644949cc6701f698b83"},"cell_type":"code","source":"plt.figure(figsize=(14,8))\nplt.scatter(tsne_tfidf_df_25.x25, \n            tsne_tfidf_df_25.y25, \n            c=tsne_tfidf_df_25.target,\n            cmap=plt.cm.coolwarm)\nplt.title(\"T-SNE plot in SVD space (perplexity=25)\")\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"782edb59e2c8f9ff641922b2d4451902b59e21c6"},"cell_type":"code","source":"plt.figure(figsize=(14,8))\nplt.scatter(tsne_tfidf_df_50.x50, \n            tsne_tfidf_df_50.y50, \n            c=tsne_tfidf_df_50.target,\n            cmap=plt.cm.coolwarm)\nplt.title(\"T-SNE plot in SVD space (perplexity=50)\")\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"86071684a2c14df6d9a20b57fc1618b3ad2b8b44"},"cell_type":"markdown","source":"----\n## 3. T-SNE applied on Doc2Vec embedding\n\nPushing forward with our T-SNE visual explorations, we next move away from semantic matrices into the realm of embeddings. Here we will use the Doc2Vec algorithm and much like its very well known counterpart Word2vec involves unsupervised learning of continuous representations for text. Unlike Word2vec which involves finding the representations for words (i.e. word embeddings), Doc2vec modifies the former method and extends it to  sentences and even documents.\n\nFor this notebook, we will be using gensim's Doc2Vec class which inherits from the base Word2Vec class where style of usage and parameters are similar. The only differences lie in the naming terminology of the training method used which are the “distributed memory” or “distributed bag of words” methods.\n\nAccording to the Gensim documentation, Doc2Vec requires the input to be an iterable object representing the sentences in the form of two lists, a list of the terms and a list of labels."},{"metadata":{"trusted":true,"_uuid":"badd8fb4b82113181ea1e9e4e8f36607800a1d14"},"cell_type":"code","source":"from gensim.test.utils import common_texts\nfrom gensim.models.doc2vec import Doc2Vec, TaggedDocument","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bbaf8bd9e50e1e55e0dce88f9fc176a629910ec1"},"cell_type":"code","source":"# Storing the question texts in a list\nquora_texts = list(train_rebal[\"question_text\"])\n\n# Creating a list of terms and a list of labels to go with it\ndocuments = [TaggedDocument(doc, tags=[str(i)]) for i, doc in enumerate(quora_texts)]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1936c235283bb3b722412a99ace57076de922904"},"cell_type":"markdown","source":"Finally we can implement the Doc2Vec model as follows"},{"metadata":{"trusted":true,"_uuid":"44123a9cb748a0fffc03c9eb611ab7569327081c"},"cell_type":"code","source":"max_epochs = 100\nalpha=0.025\nmodel = Doc2Vec(documents,\n                size=10, \n                min_alpha=0.00025,\n                alpha=alpha,\n                min_count=1,\n#                 window=2, \n                workers=4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4e04e8bf974bac00a2fe2779cfb2e817849d178c"},"cell_type":"code","source":"# model.build_vocab(documents)\n\n# for epoch in range(max_epochs):\n#     print('iteration {0}'.format(epoch))\n#     model.train(documents,\n#                 total_examples=model.corpus_count,\n#                 epochs=model.iter)\n#     # decrease the learning rate\n#     model.alpha -= 0.0002\n#     # fix the learning rate, no decay\n#     model.min_alpha = model.alpha","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f688a4a1cc9bc79d231fc415e99325bf529d8294"},"cell_type":"markdown","source":"Fitting a T-SNE model to the dense embeddings and overlaying that with the target visuals, we get:"},{"metadata":{"trusted":true,"_uuid":"d6d09a84f87e73732d382d2f6b3738a40ae618d4"},"cell_type":"code","source":"# Creating and fitting the tsne model to the document embeddings\ntsne_model = TSNE(n_jobs=4,\n                  early_exaggeration=4,\n                  n_components=2,\n                  verbose=1,\n                  random_state=2018,\n                  n_iter=300)\ntsne_d2v = tsne_model.fit_transform(model.docvecs.vectors_docs)\n\n# Putting the tsne information into sq\ntsne_d2v_df = pd.DataFrame(data=tsne_d2v, columns=[\"x\", \"y\"])\n# tsne_tfidf_df.columns = [\"x\", \"y\"]\ntsne_d2v_df[\"qid\"] = train_rebal[\"qid\"].values\ntsne_d2v_df[\"question_text\"] = train_rebal[\"question_text\"].values\ntsne_d2v_df[\"target\"] = train_rebal[\"target\"].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4aaf9dfddb71b4982916b44227ec5a39fe7daa7"},"cell_type":"code","source":"output_notebook()\nplot_d2v = bp.figure(plot_width = 800, plot_height = 700, \n                       title = \"T-SNE applied to Doc2vec document embeddings\",\n                       tools = \"pan, wheel_zoom, box_zoom, reset, hover, previewsave\",\n                       x_axis_type = None, y_axis_type = None, min_border = 1)\n\n# colormap = np.array([\"#6d8dca\", \"#d07d3c\"])\ncolormap = np.array([\"darkblue\", \"cyan\"])\n\n# palette = d3[\"Category10\"][len(tsne_tfidf_df[\"asset_name\"].unique())]\nsource = ColumnDataSource(data = dict(x = tsne_d2v_df[\"x\"], \n                                      y = tsne_d2v_df[\"y\"],\n                                      color = colormap[tsne_d2v_df[\"target\"]],\n                                      question_text = tsne_d2v_df[\"question_text\"],\n                                      qid = tsne_d2v_df[\"qid\"],\n                                      target = tsne_d2v_df[\"target\"]))\n\nplot_d2v.scatter(x = \"x\", \n                   y = \"y\", \n                   color=\"color\",\n                   legend = \"target\",\n                   source = source,\n                   alpha = 0.7)\nhover = plot_d2v.select(dict(type = HoverTool))\nhover.tooltips = {\"qid\": \"@qid\", \n                  \"question_text\": \"@question_text\", \n                  \"target\":\"@target\"}\n\nshow(plot_d2v)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"59ca26e92885b792b1c6ad83515d061cf3e03a3d"},"cell_type":"markdown","source":"**Takeaways from the plot**\n\nThe visual overlap between Sincere and Insincere labelled questions are even greater in the Doc2Vec plots - so much so that there doesn't seem to be any obvious manner to segragate the labels via eye-balling if going down the route of document embeddings."},{"metadata":{"trusted":true,"_uuid":"05b3f79296abb7954ce82e831918417ecef3a6a2"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}