{"cells":[{"metadata":{"_uuid":"3ea475d9869a8c8b48113e140ee9e9d012229c3a"},"cell_type":"markdown","source":"## Cleaning & Preprocessing Documents\n\nEssential tasks in NLP are cleaning and tokenizing a corpus. As I am building toward applying some deep learning models, I chose to clean to maximize coverage of the Fasttext embedding file. \n\nThe Fasttext (created by Facebook) embedding file appears to be the most recent of the embeddings available for this contest. Fasttext wiki-news-100d-1M.vec was trained on:\n* [Wikipedia 2017](https://meta.wikimedia.org/wiki/Data_dumps#Download)\n* [UMBC webbase corpus](https://ebiquity.umbc.edu/blogger/2013/05/01/umbc-webbase-corpus-of-3b-english-words/) - February 2007 web crawl of 100 million web pages \n* [statmt.org news dataset](http://www.casmacat.eu/corpus/news-commentary.html) - which consists of political and economic commentary crawled from the website Project Syndicate in 2016.\n\nSince a non-trivial number of the Quora questions relate to fairly recent events in US politics (e.g. Trump & Hillary) and these questions have a high rate of insincerity, it seemed a logical choice. Naturally, given time one would clean and preprocess for of the available embeddings and try all three, as one of the others may well have better performance or provide a worthwhile model to add to an ensemble. \n\nThis script accomplishes the following:\n1. Clean question text to remove contractions (don't = do not), correct frequenty seen misspellings, replace accronyms (ie D&D = Dungeons & Dragons), remove LaTeX and extraneous spaces.\n2. Tokenize cleaned documents using TF-IDF and CountVectorizer with stemming or lemmatization and bigrams or 4-gram encodings, resulting in several term-document matrices to test models on.\n3. Creation of engineered features\n4. Evaluation of the coverage of final cleaned documents in word embedding file (e.g. what proportion of tokens in the cleaned documents are exactly matched in the FastText word embedding file).\n\nFollowing document cleaning and tokenization, this kernel concludes by creating bag-of-words and TF-IDF bigram and 4-gram document-term matrices for use in next analysis steps.\n\nThe following sources were valueable in developing this code.\n\n* Text clean up. The following scripts build on one another (e.g. everyone is adapting and improving one another's code). I used some functions wholesale, most I at least made some fixes or tweaks**. I cited athorship in the docstrings (let me know if I missed anything):\n    * [KevinLiao159 (GitHub)](https://github.com/KevinLiao159/Quora/blob/master/kernels/submission_v50.py)\n    * [Heng Zheng (Kaggle)](https://www.kaggle.com/hengzheng/attention-capsule-why-not-both-lb-0-694)\n    * [Theo Viel (Kaggle)](https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2)\n* Evaluating embedding file coverage:\n     * [Dieter (Kaggle)](https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport re\nimport string\nimport unicodedata\nfrom collections import Counter\nimport os\nimport pickle\n\nfrom sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer\n\n# Important! NLTK lemmatizer needs a POS (parts of speech) tags.\n# https://www.kaggle.com/alvations/basic-nlp-with-nltk\nfrom nltk.stem import WordNetLemmatizer, PorterStemmer\nfrom nltk import pos_tag\nfrom nltk import word_tokenize\nimport nltk","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c5508a5833bc7266c6ede83ac9a66cf7fea066b5"},"cell_type":"markdown","source":"### Text Preprocessing Functions"},{"metadata":{"trusted":true,"_uuid":"52c1607c3ef60faad33a0c69418e2f09218749aa"},"cell_type":"code","source":"def load_data(file_prefix=''):\n    \"\"\"\n    Load test and train data from csv.\n    Parameters\n    __________\n    file_prefix: str\n        Optional prefix to add to \"train.csv\" or \"test.csv\" file names\n    \n    Returns\n    _______\n    df_train: DataFrame\n        Full raw training dataset\n    df_test: DataFrame\n        Full raw test dataset\n    \"\"\"\n    \n    # Select local path vs kaggle kernel\n    path = os.getcwd()\n    if 'data-projects/kaggle_quora/notebooks' in path:\n        data_dir = '../data/raw/'\n    else:\n        data_dir = '../input/'\n\n    df_train = pd.read_csv(data_dir + file_prefix +'train.csv')\n    df_test = pd.read_csv(data_dir + file_prefix +'test.csv')\n    return df_train, df_test\n\ndef normalize_unicode(text):\n    \"\"\"\n    author: Kevin Liao\n    unicode string normalization\n    \"\"\"\n    return unicodedata.normalize('NFKD', text)\n\n\ndef remove_newline(text):\n    \"\"\"\n    author: Kevin Liao\n    remove \\n and  \\t\n    \"\"\"\n    text = re.sub('\\n', ' ', text)\n    text = re.sub('\\t', ' ', text)\n    text = re.sub('\\b', ' ', text)\n    text = re.sub('\\r', ' ', text)\n    return text\n\ndef clean_latex(text):\n    \"\"\"\n    author: Kevin Liao\n    convert r\"[math]\\vec{x} + \\vec{y}\" to English\n    \"\"\"\n    # edge case\n    text = re.sub(r'\\[math\\]', ' LaTex math ', text)\n    text = re.sub(r'\\[\\/math\\]', ' LaTex math ', text)\n    text = re.sub(r'\\\\', ' LaTex ', text)\n\n    pattern_to_sub = {\n        r'\\\\mathrm': ' LaTex math mode ',\n        r'\\\\mathbb': ' LaTex math mode ',\n        r'\\\\boxed': ' LaTex equation ',\n        r'\\\\begin': ' LaTex equation ',\n        r'\\\\end': ' LaTex equation ',\n        r'\\\\left': ' LaTex equation ',\n        r'\\\\right': ' LaTex equation ',\n        r'\\\\(over|under)brace': ' LaTex equation ',\n        r'\\\\text': ' LaTex equation ',\n        r'\\\\vec': ' vector ',\n        r'\\\\var': ' variable ',\n        r'\\\\theta': ' theta ',\n        r'\\\\mu': ' average ',\n        r'\\\\min': ' minimum ',\n        r'\\\\max': ' maximum ',\n        r'\\\\sum': ' + ',\n        r'\\\\times': ' * ',\n        r'\\\\cdot': ' * ',\n        r'\\\\hat': ' ^ ',\n        r'\\\\frac': ' / ',\n        r'\\\\div': ' / ',\n        r'\\\\sin': ' Sine ',\n        r'\\\\cos': ' Cosine ',\n        r'\\\\tan': ' Tangent ',\n        r'\\\\infty': ' infinity ',\n        r'\\\\int': ' integer ',\n        r'\\\\in': ' in ',\n    }\n    # post process for look up\n    pattern_dict = {k.strip('\\\\'): v for k, v in pattern_to_sub.items()}\n    # init re\n    patterns = pattern_to_sub.keys()\n    pattern_re = re.compile('(%s)' % '|'.join(patterns))\n\n    def _replace(match):\n        \"\"\"\n        reference: https://www.kaggle.com/hengzheng/attention-capsule-why-not-both-lb-0-694 # noqa\n        \"\"\"\n        return pattern_dict.get(match.group(0).strip('\\\\'), match.group(0))\n    return pattern_re.sub(_replace, text)\n\ndef decontracted(text):\n    \"\"\"\n    author: Kevin Liao\n    de-contract the contraction\n    \"\"\"\n    try:\n        # specific\n        text = re.sub(r\"(W|w)on(\\'|\\’)t\", \"will not\", text)\n        text = re.sub(r\"(C|c)an(\\'|\\’)t\", \"can not\", text)\n        text = re.sub(r\"(Y|y)(\\'|\\’)all\", \"you all\", text)\n        text = re.sub(r\"(Y|y)a(\\'|\\’)ll\", \"you all\", text)\n\n        # general\n        text = re.sub(r\"(I|i)(\\'|\\’)m\", \"i am\", text)\n        text = re.sub(r\"(A|a)in(\\'|\\’)t\", \"is not\", text)\n        text = re.sub(r\"n(\\'|\\’)t\", \" not\", text)\n        text = re.sub(r\"(\\'|\\’)re\", \" are\", text)\n        text = re.sub(r\"(\\'|\\’)s\", \" is\", text)\n        text = re.sub(r\"(\\'|\\’)d\", \" would\", text)\n        text = re.sub(r\"(\\'|\\’)ll\", \" will\", text)\n        text = re.sub(r\"(\\'|\\’)t(?!h)\", \" not\", text)\n        text = re.sub(r\"(\\'|\\’)ve\", \" have\", text)\n    except:\n        print('error processing text:{}'.format(text))\n        \n    return text\n\ndef remove_string(text, string_to_omit=['']):\n    \"\"\"\n    author: Kevin Liao\n    Substrings to delete if present.\n    \"\"\"    \n    # light arg checking\n    if type(string_to_omit) == str:\n        string_to_omit = [string_to_omit]\n    \n    re_tok = re.compile(f'({string_to_omit})')\n    return re_tok.sub(r'', text)    \n\ndef spacing_digit(text):\n    \"\"\"\n    author: Kevin Liao\n    add space before and after digits\n    \"\"\"\n    re_tok = re.compile('([0-9])')\n    return re_tok.sub(r' \\1 ', text)\n\n\ndef spacing_number(text):\n    \"\"\"\n    author: Kevin Liao\n    add space before and after numbers\n    \"\"\"\n    re_tok = re.compile('([0-9]{1,})')\n    return re_tok.sub(r' \\1 ', text)\n\n\ndef remove_number(text):\n    \"\"\"\n    author: Kevin Liao\n    numbers are not toxic\n    \"\"\"\n    return re.sub('\\d+', ' ', text)\n\ndef remove_space(text):\n    \"\"\"\n    author: Kevin Liao\n    remove extra spaces and ending space if any\n    \"\"\"\n    text = re.sub('\\s+', ' ', text)\n    text = re.sub('\\s+$', '', text)\n    return text\n\ndef clean_misspell(text):\n    \"\"\"\n    adapted from: Kevin Liao\n    misspell list (quora vs. fasttext wiki-news-300d-1M)\n    \"\"\"\n    misspell_to_sub = {\n        '“': ' \" ',\n        '”': ' \" ',\n        '°C': 'degrees Celsius',\n        '&amp;': ' & ',\n        '2k17': '2017',\n        '2k18': '2018',\n        '9/11': 'terrorist attack',\n        'Aadhar': 'Indian identification number',\n        'aadhar': 'Indian identification number',\n        ' adhar': 'Indian identification number',\n        'Adityanath': 'Indian monk Yogi Adityanath',\n        'AFCAT': 'Indian air force recruitment exam',\n        'airhostess': 'air hostess',\n        'Ambedkarite': 'Dalit Buddhist movement ',\n        'AMCAT': 'Indian employment assessment examination',\n        'and/or': 'and or',\n        'antibrahmin': 'anti Brahminism',\n        'articleship': 'chartered accountant internship',\n        'Asifa': 'abduction rape murder case ',\n        'AT&T': 'telecommunication company',\n        'atrracted': 'attract',\n        'Awadesh': 'Indian engineer Awdhesh Singh',\n        'Awdhesh': 'Indian engineer Awdhesh Singh',\n        'Babchenko': 'Arkady Arkadyevich Babchenko faked death',\n        'Barracoon': 'Black slave',\n        'Bathla': 'Namit Bathla',\n        'bcom': 'bachelor of commerce',\n        'beyon´çe': 'Beyoncé',\n        'Bhakts': 'Bhakt',\n        'bhakts': 'Bhakt',\n        'bigdata': 'big data',\n        'biharis': 'Biharis',\n        'BIMARU': 'Bihar Madhya Pradesh Rajasthan Uttar Pradesh',\n        'BITSAT': 'Birla Institute of Technology entrance examination',\n        'BNBR': 'be nice be respectful',\n        'bodycams': 'body cams',\n        'bodyshame': 'body shaming',\n        'bodyshoppers': 'body shopping',\n        'Bolsonaro': 'Jair Bolsonaro',\n        'Boshniak': 'Bosniaks ',\n        'Boshniaks': 'Bosniaks',\n        'bremainer': 'anti Brexit',\n        'bremoaner': 'Brexit remainer',\n        'Brexiteer': 'Brexit supporter',\n        'Brexiteers': 'Brexit supporters',\n        'Brexiter': 'Brexit supporter',\n        'Brexiters': 'Brexit supporters',\n        'brexiters': 'Brexit supporters',\n        'Brexiting': 'Brexit',\n        'Brexitosis': 'Brexit disorder',\n        'Brexshit': 'Brexit bullshit',\n        'C#': 'computer programming language',\n        'c#': 'computer programming language',\n        'C++': 'computer programming language',\n        'c++': 'computer programming language',\n        'Cananybody': 'Can any body',\n        'cancelled': 'canceled',\n        'Castrater': 'castration',\n        'castrater': 'castration',\n        'centre': 'center',\n        'Chodu': 'fucker',\n        'Chutiya': 'Tibet people ',\n        'Chutiyas': 'Tibet people ',\n        'cishet': 'cisgender and heterosexual person',\n        'citicise': 'criticize',\n        'cliché': 'cliche',\n        'clichéd': 'cliche',\n        'clichés': 'cliche',\n        'Clickbait': 'click bait ',\n        'clickbait': 'click bait ',\n        'coinbase': 'bitcoin wallet',\n        'Coinbase': 'bitcoin wallet',\n        'colour': 'color',\n        'COMEDK': 'medical engineering and dental colleges of Karnataka entrance examination',\n        'counselling': 'counseling',\n        'Crimean': 'Crimea people ',\n        'currancies': 'currencies',\n        'currancy': 'currency',\n        'cybertrolling': 'cyber trolling',\n        'D&D': 'dungeons & dragons game',\n        'daesh': 'Islamic State of Iraq and the Levant',\n        'deadbody': 'dead body',\n        'deaddict': 'de addict',\n        'demcoratic': 'Democratic',\n        'demonetisation': 'demonetization',\n        'demonetisation': 'demonetization',\n        'Demonetization': 'demonetization',\n        'demonitisation': 'demonetization',\n        'demonitization': 'demonetization',\n        'deplorables': 'deplorable',\n        'doI': 'do I',\n        'Doklam': 'disputed Indian Chinese border area',\n        'Doklam': 'Tibet',\n        'Dönmeh': 'Islam',\n        'Dravidanadu': 'Dravida Nadu',\n        'dropshipping': 'drop shipping',\n        'Drumpf ': 'Donald Trump fool ',\n        'Drumpfs': 'Donald Trump fools',\n        'Dumbassistan': 'dumb ass Pakistan',\n        'emiratis': 'Emiratis',\n        'Eroupian': 'European',\n        'Etherium': 'Ethereum',\n        'Eurocentric': 'Eurocentrism ',\n        'exboyfriend': 'ex boyfriend',\n        'facetards': 'Facebook retards',\n        'Fadnavis': 'Indian politician Devendra Fadnavis',\n        'favourite': 'favorite',\n        'Fck': 'Fuck',\n        'fck': 'fuck',\n        'Feku': 'The Man of India ',\n        'feminazism': 'feminism nazi',\n        'FIITJEE': 'Indian tutoring service',\n        'fiitjee': 'Indian tutoring service',\n        'fortnite': 'Fortnite ',\n        'Fortnite': 'video game',\n        'Gixxer': 'motorcycle',\n        'Golang': 'computer programming language',\n        'golang': 'computer programming language',\n        'Gujratis': 'Gujarati',\n        'Gurmehar': 'Gurmehar Kaur Indian student activist',\n        'h1b': 'US work visa',\n        'H1B': 'US work visa',\n        'hairfall': 'hair loss',\n        'harrase': 'harass',\n        'he/she': 'he or she',\n        'healhtcare': 'healthcare',\n        'him/her': 'him or her',\n        'Hindians': 'North Indian who hate British',\n        'Hinduphobia': 'Hindu phobic',\n        'hinduphobia': 'Hindu phobic',\n        'Hinduphobic': 'Hindu phobic',\n        'hinduphobic': 'Hindu phobic',\n        'his/her': 'his or her',\n        'Hongkongese': 'HongKong people',\n        'hongkongese': 'HongKong people',\n        'howcan': 'how can',\n        'Howdo': 'How do',\n        'howdo': 'how do',\n        'howdoes': 'how does',\n        'howmany': 'how many',\n        'howmuch': 'how much',\n        'HYPS': ' Harvard Yale Princeton Stanford',\n        'HYPSM': ' Harvard Yale Princeton Stanford MIT',\n        'ICOs': 'cryptocurrencies initial coin offering',\n        'Idiotism': 'idiotism',\n        'IITian': 'Indian Institutes of Technology student',\n        'IITians': 'Indian Institutes of Technology students',\n        'IITJEE': 'Indian Institutes of Technology entrance examination',\n        ' incel': ' involuntary celibates',\n        ' incels': ' involuntary celibates',\n        'indans': 'Indian',\n        'jallikattu': 'Jallikattu',\n        'JEE MAINS': 'Indian university entrance examination',\n        'Jewdar': 'Jew dar',\n        'Jewism': 'Judaism',\n        'jewplicate': 'jewish replicate',\n        'JIIT': 'Jaypee Institute of Information Technology',\n        'Kalergi': 'Coudenhove-Kalergi',\n        'Kashmirians': 'Kashmirian',\n        'Khalistanis': 'Sikh separatist movement',\n        'Khazari': 'Khazars',\n        'kompromat': 'compromising material',\n        'koreaboo': 'Korea boo ',\n        'KVPY': 'entrance examination',\n        'labour': 'labor',\n        'langague': 'language',\n        'LGBTQ': 'lesbian  gay  bisexual  transgender queer',\n        'LGBT': 'lesbian  gay  bisexual  transgender',\n        'Machedo': 'Indian internet celebrity',\n        'madheshi': 'Madheshi',\n        'Madridiots': 'Real Madrid idiot supporters',\n        'mailbait': 'mail bait',\n        'MAINS': 'exam',\n        'marathis': 'Marathi',\n        'marksheet': 'university transcript',\n        'mastrubate': 'masturbate',\n        'mastrubating': 'masturbating',\n        'mastrubation': 'masturbation',\n        'mastuburate': 'masturbate',\n        'meninism': 'male feminism',\n        'MeToo': 'feminist activism campaign',\n        'Mewani': 'Indian politician Jignesh Mevani',\n        'MGTOWS': 'Men Going Their Own Way',\n        'micropenis': 'tiny penis',\n        'moeslim': 'Muslim',\n        'mongloid': 'Mongoloid',\n        'mtech': 'Master of Engineering',\n        'muhajirs': 'Muslim immigrant',\n        'Myeshia': 'widow of Green Beret killed in Niger',\n        'mysoginists': 'misogynists',\n        'naïve': 'naive',\n        'narcisist': 'narcissist',\n        'narcissit': 'narcissist',\n        'narcissit': 'narcissist',\n        'Naxali ': 'Naxalite ',\n        'Naxalities': 'Naxalites',\n        'NICMAR': 'Indian university',\n        'Niggeriah': 'Nigger',\n        'Niggerism': 'Nigger',\n        'NMAT': 'Indian MBA exam',\n        'Northindian': 'North Indian ',\n        'northindian': 'north Indian ',\n        'northkorea': 'North Korea',\n        'Novichok': 'Soviet Union agents',\n        'organisation': 'organization',\n        'Padmavat': 'Indian Movie Padmaavat',\n        'Pahul': 'Amrit Sanskar',\n        'penish': 'penis',\n        'pennis': 'penis',\n        'Pizzagate': 'Pizzagate conspiracy theory',\n        'Pribumi': 'Native Indonesian',\n        'qouta': 'quota',\n        'quorans': 'advice website user',\n        'quoran': 'advice website user',\n        'Quorans': 'advice website user',\n        'Quoran': 'advice website user',\n        'quoras': 'advice website',\n        'Qoura ': 'advice website ',\n        'Qoura': 'advice website',\n        'Quora': 'advice website',\n        'Quroa': 'advice website',\n        'QUORA': 'advice website',\n        'R&D': 'research and development',\n        'r&d': 'research and development',\n        'r-aping': 'raping',\n        'raaping': 'rape',\n        'rapefugees': 'rapist refugee',\n        'Rapistan': 'Pakistan rapist',\n        'rapistan': 'Pakistan rapist',\n        'Rejuvalex': 'hair growth formula',\n        'ReleaseTheMemo': 'cry for the right and Trump supporters',\n        'Remainers': 'anti Brexit',\n        'remainers': 'anti Brexit',\n        'remoaner': 'remainer ',\n        'rohingya': 'Rohingya ',\n        'sallary': 'salary',\n        'Sanghis': 'Sanghi',\n        'sh*t': 'shit',\n        'shithole': ' shithole ',\n        'shitlords': 'shit lords',\n        'shitpost': 'shit post',\n        'shitslam': 'shit Islam',\n        'sickular': 'India sick secular ',\n        'signuficance': 'significance',\n        'SJW': 'social justice warrior',\n        'SJWs': 'social justice warrior',\n        'Skripal': 'Sergei Skripal',\n        'Strzok': 'Hillary Clinton scandal',\n        'suckimg': 'sucking',\n        'superficious': 'superficial',\n        'Swachh': 'Swachh Bharat mission campaign ',\n        'Tambrahms': 'Tamil Brahmin',\n        'Tamilans': 'Tamils',\n        'Terroristan': 'terrorist Pakistan',\n        'terroristan': 'terrorist Pakistan',\n        'Tharki': 'pervert',\n        'tharki': 'pervert',\n        'theatre': 'theater',\n        'theBest': 'the best',\n        'thighing': 'masturbate',\n        'travelling': 'traveling',\n        'trollbots': 'troll bots',\n        'trollimg': 'trolling',\n        'trollled': 'trolled',\n        'Trumpers': 'Trump supporters',\n        'Trumpanzees': 'Trump chimpanzee fool',\n        'Turkified': 'Turkification',\n        'turkified': 'Turkification',\n        'UCEED': 'Indian Institute of Technology Bombay entrance examination',\n        'unacadamy': 'Indian online classroom',\n        'Unacadamy': 'Indian online classroom',\n        'unoin': 'Union',\n        'unsincere': 'insincere',\n        'UPES': 'Indian university',\n        'UPSEE': 'Indian university entrance examination',\n        'vaxxer': 'vocal nationalist ',\n        'VITEEE': 'Vellore institute of technology',\n        'watsapp': 'Whatsapp',\n        'whattsapp': 'Whatsapp',\n        'WBJEE': 'West Bengal entrance examination',\n        'weatern': 'western',\n        'westernise': 'westernize',\n        'Whatare': 'What are',\n        'whatare': 'what are',\n        'whst': 'what',\n        'Whta': 'What',\n        'whydo': 'why do',\n        'Whykorean': 'Why Korean',\n        'Wjy': 'Why',\n        'WMAF': 'White male married Asian female',\n        'wumao ': 'cheap Chinese stuff',\n        'wumaos': 'cheap Chinese stuff',\n        'wwii': 'world war 2',\n        ' xender': ' gender',\n        'XXXTentacion': 'Tentacion',\n        'youtu ': 'youtube ',\n        'Zerodha': 'online stock brokerage',\n        'Žižek': 'Slovenian philosopher Slavoj Žižek',\n        'Zoë': 'Zoe',\n        '卐': 'Nazi Germany'\n    }\n\n    escape_cars = re.compile('(\\+|\\*)')\n    misspell = '|'.join([escape_cars.sub(r\"\\\\\\1\",i) for i in misspell_to_sub.keys()])\n    misspell_re = re.compile(misspell)\n    \n    def _replace(match):\n        return misspell_to_sub.get(match.group(0), match.group(0))\n    \n    return misspell_re.sub(_replace, text)\n\ndef space_chars(text, chars_to_space):\n    \"\"\"\n    Takes a string and list of characters, insert space before and after \n    characters that appear in text.\n    \n    Parameters\n    ----------\n    text : str\n        String to search\n    chars_to_space : list\n        list of characters to find and space\n        \n    Returns\n    -------\n    str\n        modified text string    \n    \"\"\"\n    \n    # light arg checking\n    if type(chars_to_space) == str:\n        chars_to_space = [chars_to_space]\n        \n    chars_to_space = set(chars_to_space)\n    chars_to_space = '|'.join(chars_to_space)\n    re_tok = re.compile('({})'.format(chars_to_space))\n    \n    return re_tok.sub(r' \\1 ', text)\n\ndef preprocess(text, remove_num=False):\n    \"\"\"\n    Text preprocessing pipeline\n    \n    Parameters\n    ----------\n    text : str\n        String to process\n        \n    Returns\n    -------\n    str\n        modified string  \n    \"\"\"    \n    # 1. Normalize \n    # normalize_unicode(text)\n    \n    # 2. Remove new-lines\n    # text = remove_newline(text)\n    \n    # 3. replace contractions (e.g. won't -> will not)\n    text = decontracted(text)\n\n    # 4. replace LateX with English\n    text = clean_latex(text)\n    \n    # 5. space characters\n    text = space_chars(text, ['\\?', ',', '\"', '\\(', '\\)', '%', ':', '\\$', \n                              '\\.', '\\+', '\\^', '/', '\\{', '\\}', '\\!', \n                              '#', '=', '-','\\|', '\\[', '\\]','\\.'])\n    \n    # 6. handle number\n    if remove_num:\n        text = remove_number(text)\n    else:\n        text = spacing_digit(text)\n    \n    # 7. fix typos and swap terms that are not recognized by embedding\n    text = clean_misspell(text)\n\n    # 8. remove space\n    text = remove_space(text)\n    \n    # 9. remove strings\n    text = remove_string(text, '_')\n    \n    return text","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e8afd9a02e06a3d1a60efbcb9b53a0e6dc09a9a9"},"cell_type":"markdown","source":"###  Feature Engineering Functions"},{"metadata":{"trusted":true,"_uuid":"09c72c38ce54ac23c06afbb414b2280b471e8ff8"},"cell_type":"code","source":"def count_uppercase_words(text):\n    \"\"\"\n    Count SHOUTY all-caps words more than 1 character long (eg not \"I\")\n    \n    Parameters\n    ----------\n    text : str\n        String to process\n        \n    Returns\n    -------\n    int\n        count of all-caps words in text  \n    \"\"\"   \n    tokens = text.split()\n    upper = [1 if u.isupper() and len(u) > 1 else 0 for u in tokens]\n    return(sum(upper))\n\ndef programming_related(text):\n    \"\"\"\n    Identify references to common programming languages, frameworks, \n    tools or databases in question text\n    \n    Parameters\n    ----------\n    text : str\n        String to process\n        \n    Returns\n    -------\n    bool\n        True if programming reference identified \n    \"\"\"\n    programming = ['javascript', 'html', 'css', 'sql', 'java', 'bash', 'python',\n                   'c#', 'c++', 'c language', 'c programming', 'c programing',\n                   'typescript', 'ruby', 'matlab', 'f#', 'clojure', 'haskell', \n                   'erlang', 'coffeescript', 'cobol', 'fortran', 'vba', '.net',\n                   'asp.net', 'scala', 'perl', 'php', 'kotlin', 'node.js', \n                   'react.js', 'angular', 'django', 'cordova', 'tensorflow', 'keras',\n                   'xamarin', 'hadoop', 'pytorch', 'mongo', 'redis', 'elasticsearch', \n                   'mariadb', 'azure', 'dynamodb', ' rds', 'redshift', 'cassandra',\n                   'apache hive', 'bigquery', 'hbase', 'linux', 'raspberry pi', \n                   'rpi ', 'arduino', 'heroku', 'drupal', 'visual studio', \n                   'sublime text', 'rstudio', 'jupyter', 'pycharm', 'netbeans',\n                   'emacs', 'vim ', 'komodo', 'graphql', 'golang']\n    \n    for word in text.split():\n        if word.lower() in programming:\n            return True\n    \n    return False\n\nclass FeatureEngineering():\n    def __init__(self, doc_column, max_words=10):\n        \"\"\"\n        Create features from text\n\n        Parameters\n        ----------\n        doc_column : str\n            column name of text to process\n        \n        max_words: int\n            maximum number of leading words (1st word in sentence) to count\n        \"\"\"\n        self._most_common = None\n        self._data = None\n        self._doc_column = doc_column\n        self._max_words = max_words\n    \n    def fit(self, df):\n        # save leading tokens (first word in sentences) information from fit dataset\n        # keep top max_word, convert to one-hot and append to dataframe\n        leading_tokens = df[self._doc_column].apply(lambda x: re.match('\\w+|\\d+|.', x)[0].lower())\n        leading_token_count = Counter(leading_tokens)\n        max_count = min(self._max_words, len(leading_token_count)-1)\n        self._most_common = [w for w,c in leading_token_count.most_common(max_count)]\n        self._data = df\n        \n    def transform(self, df):\n        # Leading tokens\n        # keep top max_word, convert to one-hot and append to dataframe\n        leading_tokens = df[self._doc_column].apply(lambda x: re.match('\\w+|\\d+|.', x)[0].lower())\n        df_leading_tokens = leading_tokens.apply(lambda x: x.lower() if x.lower() in self._most_common else 'other')\n        df_leading_tokens = pd.get_dummies(df_leading_tokens)\n        for token in self._most_common:\n            if token not in df_leading_tokens.columns:\n                df_leading_tokens[token] = 0\n        df_leading_tokens = df_leading_tokens.rename(columns = {c: 'leading_word_' + c for c in df_leading_tokens.columns})\n        # Using 'other' as reference category\n        if 'leading_word_other' in df_leading_tokens.columns:\n            df_leading_tokens = df_leading_tokens.drop('leading_word_other', axis=1)\n        df = pd.concat([df, df_leading_tokens], axis=1)\n        \n        # Word count\n        df['word_count'] = df[self._doc_column].apply(lambda x: len(re.findall(r'\\w+',x)))\n\n        # Character count\n        df['char_count'] = df[self._doc_column].apply(lambda x: len(x))\n\n        # How many question marks\n        df['question_mark_count'] = df[self._doc_column].apply(lambda x: len(re.findall(r'\\?',x)))\n\n        # LaTex or Math symbols\n        # Programming questions\n        df['programming'] = df[self._doc_column].apply(lambda x: programming_related(x))\n\n        # ALL CAPS Words\n        df['caps_count'] = df[self._doc_column].apply(lambda x: count_uppercase_words(x))\n        \n        return df","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7cf1344052e9eb2557c9df26402f44fd1eccbbda"},"cell_type":"markdown","source":"### Tokenizer"},{"metadata":{"trusted":true,"_uuid":"49819f81a7ddd5b92f4ee6ac4ea6aedbe1e8d5f4"},"cell_type":"code","source":"# Important! NLTK lemmatizer needs a POS (parts of speech) tags to work correctly\n# https://www.kaggle.com/alvations/basic-nlp-with-nltk\n# TL/DR otherwise all words are assumed to be nouns and\n# do -> doe\ndef penn2morphy(penntag):\n    \"\"\" \n    Author: Liling Tan https://www.kaggle.com/alvations/basic-nlp-with-nltk\n    Converts Penn Treebank tags to WordNet.\n    \"\"\"\n    morphy_tag = {'NN':'n', 'JJ':'a',\n                  'VB':'v', 'RB':'r'}\n    try:\n        return morphy_tag[penntag[:2]]\n    except:\n        return 'n' # if mapping isn't found, fall back to Noun.\n\n# Note: Lemmatizing rather than stemming takes much longer\ndef tokenize(text, stem = False, stop_words = None):\n    \"\"\"\n    Text tokenization pipeline\n\n    Parameters\n    ----------\n    test : str\n        document to tokenize.\n\n    stem: bool\n        If true use porter stemmer, else use WordNet lemmatizer.\n        Note that lemmatizing is a much more expensive operation than \n        stemming text.\n        \n    stop_words: list\n        optional list of stop words.\n    \n    Returns\n    -------\n    list\n        ordered list of tokens \n    \"\"\"\n    \n    if stop_words == None:\n        # Starting from a small stop-word list. TODO: build this out further.\n        stop_words = ('it', 'its', 'this', 'that', 'these', 'those', \n                      'a', 'an', 'the', 'and', 'but', 'if', 'or', \n                      'as', 'of', 'at', 'by', 'to', 'in', 'so')\n\n    if stem:\n        stemmer = PorterStemmer()\n    else:\n        lemmer = WordNetLemmatizer()\n    \n    # tokenize text into words\n    tokens = nltk.word_tokenize(text)\n    \n    # drop punctuation\n    tokens = [t for t in tokens if t.isalpha()]\n\n    # drop stop words\n    tokens = [t for t in tokens if t not in stop_words]\n\n    # lowercase\n    tokens  = [t.lower() for t in tokens]\n    \n    if stem:\n        tokens = [stemmer.stem(t) for t in tokens]\n    else:\n        # parts-of-speech tags\n        # required for nltk WordNetLemmatizer - if not supplied default is \"noun\"\n        tagged_tokens = [t for t in pos_tag(tokens)]\n\n        # lemmetize\n        tokens = [lemmer.lemmatize(t, pos=penn2morphy(tag)) for t,tag in tagged_tokens]\n    \n    return tokens","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"31c4c67e9ac923ec89fb5f1ac8a6f44ef915ccac"},"cell_type":"markdown","source":"> ### Process Docs & Save"},{"metadata":{"trusted":true,"_uuid":"f5138f7df49b09bdd8581633283fa753ce371331"},"cell_type":"code","source":"df_train, df_test = load_data()\n\ndf_train['question_text_pr'] = df_train['question_text'].apply(preprocess)\ndf_test['question_text_pr'] = df_test['question_text'].apply(preprocess)\n\nef = FeatureEngineering('question_text', 20)\nef.fit(df_train)\ndf_train_plus = ef.transform(df_train)\ndf_test_plus = ef.transform(df_test)\n\ndf_train_plus.to_csv('processed_train.csv')\ndf_test_plus.to_csv('processed_test.csv')\n\ndel df_train_plus\ndel df_test_plus\ndel ef","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1b874ae585ce450481bce50c32c17346afc21188"},"cell_type":"code","source":"%%time\n# Simple tokenizer without any pre-processing of text, using nltk built in tokenizer\n# and english stop words for comparison sake, as this comes out of the box\ntfidf = TfidfVectorizer(tokenizer=nltk.word_tokenize, \n                        ngram_range=(1,4),\n                        min_df=5,\n                        max_df=0.9,\n                        strip_accents='unicode',\n                        use_idf=True,\n                        smooth_idf=True,\n                        sublinear_tf=True)\n\nall_text = np.concatenate([df_train['question_text'], df_train['question_text']])\ntfidf.fit(all_text)\nX_train = tfidf.transform(df_train['question_text'])\nX_train_feats = tfidf.get_feature_names()\n\n# save results for use in further kernels, as this takes a long time to complete.\npickle.dump(X_train, open(\"tfidf_train_base.pickle\", \"wb\"))\npickle.dump(X_train_feats, open(\"tfidf_feats_base.pickle\", \"wb\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"da5fa7e2bdd95853bb983b251ccd596788ead0c9"},"cell_type":"code","source":"%%time\n# Simple tokenizer without any pre-processing of text, using nltk built in tokenizer\n# and english stop words for comparison sake, as this comes out of the box\ntfidf = TfidfVectorizer(tokenizer=nltk.word_tokenize, \n                        ngram_range=(1,4),\n                        min_df=5,\n                        max_df=0.9,\n                        strip_accents='unicode',\n                        use_idf=True,\n                        smooth_idf=True,\n                        sublinear_tf=True)\n\ntfidf.fit(all_text)\nX_train = tfidf.transform(df_train['question_text'])\nX_train_feats = tfidf.get_feature_names()\n\n# save results for use in further kernels, as this takes a long time to complete.\npickle.dump(X_train, open(\"tfidf_train_pr.pickle\", \"wb\"))\npickle.dump(X_train_feats, open(\"tfidf_feats_pr.pickle\", \"wb\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5519ef936a5cc45c976e66efa16a8a85a2217e4d"},"cell_type":"code","source":"%%time\n# Create TD-IDF vectorizer on whole corpus (train + test), n-gram\ncountv2 = CountVectorizer(tokenizer=tokenize,\n                          ngram_range=(1,2),\n                          min_df=5,\n                          max_df=0.9,\n                          strip_accents='unicode')\n\ncountv2.fit(all_text)\nX_train = countv2.transform(df_train['question_text_pr'])\nX_train_feats = countv2.get_feature_names()\n\n# save results for use in further kernels, as this takes a long time to complete.\npickle.dump(X_train, open(\"count_train_lem_ng2.pickle\", \"wb\"))\npickle.dump(X_train_feats, open(\"count_feats_lem_ng2.pickle\", \"wb\"))\n\ndel countv2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e046ef2ddcc99b28ab87944079bcfb2840c1e460"},"cell_type":"code","source":"%%time\n# Create TD-IDF vectorizer on whole corpus (train + test), n-gram\ntfidf = TfidfVectorizer(tokenizer=tokenize, \n                        ngram_range=(1,2),\n                        min_df=5,\n                        max_df=0.9,\n                        strip_accents='unicode',\n                        use_idf=True,\n                        smooth_idf=True,\n                        sublinear_tf=True)\n\nall_text = np.concatenate([df_train['question_text_pr'], df_train['question_text_pr']])\ntfidf.fit(all_text)\nX_train = tfidf.transform(df_train['question_text_pr'])\nX_train_feats = tfidf.get_feature_names()\n\n# save results for use in further kernels, as this takes a long time to complete.\npickle.dump(X_train, open(\"tfidf_train_lem_ng2.pickle\", \"wb\"))\npickle.dump(X_train_feats, open(\"tfidf_feats_lem_ng2.pickle\", \"wb\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"09da0202704d750a46fb91d2f7363944a50f6507"},"cell_type":"code","source":"%%time\n# Create TD-IDF vectorizer on whole corpus (train + test), n-gram\ntfidf = TfidfVectorizer(tokenizer=tokenize, \n                        ngram_range=(1,4),\n                        min_df=5,\n                        max_df=0.9,\n                        strip_accents='unicode',\n                        use_idf=True,\n                        smooth_idf=True,\n                        sublinear_tf=True)\n\nall_text = np.concatenate([df_train['question_text_pr'], df_train['question_text_pr']])\ntfidf.fit(all_text)\nX_train = tfidf.transform(df_train['question_text_pr'])\nX_test = tfidf.transform(df_test['question_text_pr'])\nX_train_feats = tfidf.get_feature_names()\n\n# save results for use in further kernels, as this takes a long time to complete.\npickle.dump(X_train, open(\"tfidf_train_lem_ng4.pickle\", \"wb\"))\npickle.dump(X_test, open(\"tfidf_test_lem_ng4.pickle\", \"wb\"))\npickle.dump(X_train_feats, open(\"tfidf_feats_lem_ng4.pickle\", \"wb\"))\n\ndel tfidf\ndel X_train\ndel X_test\ndel X_train_feats","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0b13c3fb54b024fe21a762896d594f6b48bb7594"},"cell_type":"markdown","source":"### Embedding Document Coverage\nLast thing we'll take a look at the token coverage in the embedding file. We'll see that processing the text improves coverage in the embedding file dramatically, increasing the vocabulary coverage from 30% to 67%.\n\nThis section in particular was informed by the excellent work of [Dieter (Kaggle)](https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings) and [Theo Viel (Kaggle)](https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2)."},{"metadata":{"trusted":true,"_uuid":"acfdf7a7497ebfb8321f8189510c42e4519dc54d"},"cell_type":"code","source":"def load_word_embedding(filepath, verbose=True):\n    \"\"\"\n    author: Theo Viel\n    given a filepath to embeddings file, return a word to vec\n    dictionary, in other words, word_embedding\n    E.g. {'word': array([0.1, 0.2, ...])}\n    \"\"\"\n    def _get_vec(word, *arr):\n        return word, np.asarray(arr, dtype='float32')\n\n    if verbose:\n        print('load word embedding ......')\n        \n    try:\n        word_embedding = dict(_get_vec(*w.split(' ')) for w in open(filepath))\n    except UnicodeDecodeError:\n        word_embedding = dict(_get_vec(*w.split(' ')) for w in open(\n            filepath, encoding=\"utf8\", errors='ignore'))\n\n    if verbose:\n        print('finished load of word embedding.')\n\n    return word_embedding\n\ndef build_vocab(docs):\n    \"\"\"\n    author: Theo Viel\n    given a list or np.array of strings create a dictionary of unique words with frequencies.\n    \n    Parameters\n    __________\n    docs: list or np.array\n        iterable of text\n    \n    Returns\n    _______\n    dict\n        unique words as keys, frequencies as values\n    \"\"\"\n    vocab = {}\n    \n    for doc in docs:\n        for word in doc.split():\n            vocab[word] = vocab.get(word, 0) + 1\n                \n    return vocab\n\ndef vocab_embedding_coverage(vocab, embedding, verbose = False):\n    \"\"\"\n    author: Theo Viel\n    \n    given a dict representing the word frequency of a corpus, \n    calculate the percentage of unique words and \n    the percentage of the corpus matched in the embedding dict.\n    \n    Parameters\n    __________\n    vocab: dict\n        word frequency of corpus\n    embedding: dict\n        embedding vector converted to dict\n    verbse: bool\n        print summary statistics\n    \n    Returns\n    _______\n    \n    perc_words : float\n        percentage of unique words identified in corpus\n    perc_corpus : float\n        percentage of corpus identified in corpus\n    words_in_embedding: dict\n        dictionary of unique words, frequency and whether found in embedding (true / false)\n    \"\"\"\n    \n    words_in_embedding = {}\n    word_found_count = 0\n    corpus_found_count = 0\n    corpus_count = 0 \n    \n    for word, freq in vocab.items():\n        corpus_count += freq\n        words_in_embedding[word] = {\n            'frequency': freq,\n            'embedding': (word in embedding)\n        }\n        if word in embedding:\n            word_found_count += 1\n            corpus_found_count += freq\n    \n    perc_words = word_found_count / len(vocab)\n    perc_corpus = corpus_found_count / corpus_count\n    \n    print('{}% of vocabulary words found in embedding files'.format(round(100*perc_words,2)))\n    print('{}% of corpus found in embedding files'.format(round(100*perc_corpus,2)))\n    \n    return perc_words, perc_corpus, words_in_embedding","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"29c923df3cb15aea018e6b33cceb58fca31c6abf"},"cell_type":"code","source":"fasttext = load_word_embedding('../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a6b39c6156bf0378ad4a77cf13c1d9a0a3c55f31"},"cell_type":"markdown","source":"#### Before"},{"metadata":{"trusted":true,"_uuid":"04aaba74aee7ae3656335b4943213d1f0353623d"},"cell_type":"code","source":"vocab = build_vocab(np.concatenate((df_train.question_text, df_test.question_text)))\nw,c,words_in_embedding = vocab_embedding_coverage(vocab, fasttext, True)\ndf_words = pd.DataFrame.from_dict(words_in_embedding, orient = 'index')\ndf_words = df_words.sort_values(by='frequency', ascending=False)\ndf_words[np.logical_not(df_words.embedding)].head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2e7cac9a14058f693b049a952e6aaf58e5fff564"},"cell_type":"markdown","source":"#### After"},{"metadata":{"trusted":true,"_uuid":"e94db3eeff0a8420a8c70b7a3c98cc0d462d7aa2"},"cell_type":"code","source":"vocab = build_vocab(np.concatenate((df_train.question_text_pr, df_test.question_text_pr)))\nw,c,words_in_embedding = vocab_embedding_coverage(vocab, fasttext, True)\ndf_words = pd.DataFrame.from_dict(words_in_embedding, orient = 'index')\ndf_words = df_words.sort_values(by='frequency', ascending=False)\ndf_words[np.logical_not(df_words.embedding)].head(10)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}