{"cells":[{"metadata":{"trusted":true,"_uuid":"7b00cb2ec0e1c5ca68a1ffee81b0bf7ed31ce5be","scrolled":true},"cell_type":"code","source":"%%javascript\n\nJupyter.keyboard_manager.command_shortcuts.add_shortcut('Q', {\n    help : 'run all cells',\n    help_index : 'zz',\n    handler : function (event) {\n        IPython.notebook.execute_all_cells();\n        return false;\n    }}\n);","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ccf29b1edf489249d5148f1f56fb9fb1055182c1"},"cell_type":"markdown","source":"## Overview\n\nMain focus: Generating synthetic data because only 6% of data is tagged insincere\n\nGreat place to start code review would be to find the classic C-style \"main\" function near the bottom.\n\nCode quality: \n* Lots of OOP, not seen commonly with ML projects, but no getters/setters for dev speed and bloat.\n* Primary reason for OOP would be readability and avoiding huge function signatures.  Another is reuse for future projects.\n* Naming matters to me for clarity of thought, and OOP indirections enable better naming.  \n* Lots of focus on \"i believe what I see\", and so, lots of print statements.  Brevity not a focus.\n* Lots of interim results written to filesystem .. saved dev. time.\n\n### Takebacks\n\n* Editor issues: Coding on Kaggle's Jupyter is painful primarily because lots of code means too much scrolling\n* I am new to Kaggle and AI competitions.  Explored multiprocessing on python first time. Did it by using filesystem.\n* Foresight suggested that I might be experimenting with many models and I wasn't clear how the limits on 2h GPU vs 6hr CPU, would pan out.\n* So multiprocessing seemed good to invest time and also for learning. What I learnt is applicable for multi-node processing in future.\n* The question generation takes a while, and would have been much faster on 4 CPU, but the training over 2DCNN is much faster on GPU.\n* I followed \"no external data sources\" so strictly that I didn't realize I could import intermediate results from other kernels\n\n\nMyNotes\n-------------\n\n* Improve q generation;; questions not used for generating q.\n* Mix embeddings: GLOVE and PARA seem popular\n* Add features by hand\n* Glove only, OOV hand fixes; explore words in different embeddings"},{"metadata":{"_uuid":"ea970328e34ce92f60be85eb156551dd0d6dd5e5"},"cell_type":"markdown","source":"# Acknowledgements\n\n1. https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2\n\n"},{"metadata":{"_uuid":"f80d164b407dceb8d1ea3eb70fb7082210319beb"},"cell_type":"markdown","source":"# Levers for different kernels\n\n* Many of these have the biggest impact on compute time."},{"metadata":{"trusted":true,"_uuid":"208e586ad367c2a29f5428eb80e0619df2f0ea87"},"cell_type":"code","source":"gFirstTime = True\ngNumEpochs = 8\ngTrainingBatchSize = 256\ngModelName = \"GPULSTM\"\ngSelectedEmbeddings = [\"GLOVE\",\"PARAGRAM\",\"WIKINEWS\"]\n#gSelectedEmbeddings = [\"GLOVE\"]\ngTrainableEmbeddings = False\ngWIP = False\ngInspect = False\ngLimit = False\ngExternalData = False\ngGenerateNewData = False\ngExecuteMain = False","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport string, time\nimport gc\nfrom collections import defaultdict, Counter\nimport numpy as np # linear algebra\nfrom IPython.core.display import display, HTML\nimport random\nimport re\nimport os, psutil, sys, pickle, operator\nimport csv\nimport multiprocessing as mpc\nimport textwrap\nimport nltk.data\nfrom shutil import copyfile, rmtree\n\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom gensim.test.utils import datapath, get_tmpfile\nfrom gensim.models import KeyedVectors\nfrom gensim.scripts.glove2word2vec import glove2word2vec\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import f1_score\nfrom tensorflow.keras.layers import *\nfrom tensorflow.keras.models import *\nfrom tensorflow.keras.optimizers import *\nfrom tensorflow.keras.callbacks import EarlyStopping, ModelCheckpoint, ReduceLROnPlateau\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom tensorflow.keras import backend as K\nfrom tensorflow.keras import initializers, regularizers, constraints, optimizers, layers\nfrom keras.utils import multi_gpu_model, Sequence\nfrom tensorflow.python.client import device_lib\nfrom keras import backend as K\nfrom keras.callbacks import Callback\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import metrics\nimport nltk\nimport spacy\nfrom spacy import displacy\nimport en_core_web_sm\nfrom nltk.corpus import stopwords\n\nfrom tqdm.autonotebook import tqdm\nfrom langdetect import detect, DetectorFactory\nDetectorFactory.seed = 0\n\npd.set_option('display.max_colwidth', -1)\npd.set_option('display.expand_frame_repr', False)\nsent_tokenizer = nltk.data.load('tokenizers/punkt/english.pickle')\nprocess = psutil.Process(os.getpid())\ngCurrentMemory = process.memory_info().rss\ngStopWords = set(stopwords.words('english'))\nSEED = 2019\nnp.random.seed(SEED)\n\ndef show_html(x):\n    display(HTML(x))\n    \nshow_html(\"<hr/><h1 align='center'>Levers of the Machine</h1><hr/>\")\nshow_html(f\"<big><big>Execute Main?   <b>{gExecuteMain}</b></big></big>\")\nshow_html(f\"<big><big>Logging Enabled?    <b>{gWIP}</b></big></big>\")\nshow_html(f\"<big><big>Data Limit Enabled?   <b>{gLimit}</b></big></big>\")\nshow_html(f\"<big><big>Inspection Enabled?    <b>{gInspect}</b></big></big>\")\nshow_html(f\"<big><big>Generate New Data?   <b>{gGenerateNewData}</b></big></big>\")\nshow_html(f\"<big><big>Trainable Embeddings?   <b>{gTrainableEmbeddings}</b></big></big>\")\nshow_html(\"<hr/>\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3b7b094c0a363faa00c7b72421631d291b24d8fd"},"cell_type":"markdown","source":"# Section: Utilities\n\n## Leverage HTML output for pretty printing logs"},{"metadata":{"trusted":true,"_uuid":"20106d191644e863c9a6ebc5f446ce1fb26b71de"},"cell_type":"code","source":"def log(x, *argv):\n    if gWIP == False:\n        return\n    x = str(x)\n    for arg in argv:\n        x = x + str(arg)\n    display(HTML(x))\n\ndef log_list(items, transpose=False, desc=None, limit=None, force=False):\n    if gWIP == False and not force:\n        return\n    \n    if items is None:\n        if desc is None:\n            desc = \"<no desc>\"\n        if gWIP:\n            log(\"For the list {}, items are None\".format(desc))\n        return\n    \n    html = []\n    html.append(\"<table>\")\n    if desc is not None:\n        if not transpose:\n            html.append(\"<th>No.</th>\")\n        html.append(\"<th colspan='3'>\" + desc + \"</th>\")\n    if transpose:\n        html.append(\"<tr>\")\n        \n    count = 0\n    if isinstance(items, dict):\n        for key, value in items.items():\n            if limit is not None and limit < count:\n                break\n            if transpose:\n                html.append(\"<td>{} : {}</td>\".format(key, value))\n            else:\n                html.append(\"<tr><td>{}</td>\".format(count))\n                html.append(\"<td>{}</td><td>{}</td></tr>\".format(key, value))\n            count += 1\n    else:\n        for item in items:\n            if limit is not None and limit < count:\n                break\n            if transpose:\n                html.append(\"<td>\" + str(item) + \"</td>\")\n            else:\n                if isinstance(item, tuple) or isinstance(item, list):\n                    html.append(\"<tr><td>{}</td>\".format(count))\n                    for col in item:\n                        html.append(\"<td>{}</td>\".format(col)) \n                else:\n                    html.append(\"<tr><td>{}</td><td>{}</td></tr>\".format(count, str(item)))\n                count += 1\n                \n    if transpose:\n        html.append(\"</tr>\")            \n\n    html.append(\"</table>\")\n    log(''.join(html))\n    \ndef log_current_memory(caption):\n    if gWIP == False:\n        return\n    \n    global gCurrentMemory\n    tmp = process.memory_info().rss\n    log(\"MEMORY -- {} : {:.4f} MB  ; delta: <b>{:.2f}</b>\".format(caption, tmp/(1024*1024), (tmp - gCurrentMemory)/(1024*1024)))\n    gCurrentMemory = tmp\n\ndef log_dir(path):\n    \n    dir_list = os.listdir(path)\n    pairs = []\n    for file in dir_list:\n        # Use join to get full file path.\n        location = os.path.join(path, file)\n\n        # Get size and add to list of tuples.\n        size = os.path.getsize(location)\n        modified = time.ctime(os.path.getmtime(location))\n        pairs.append((file, str(size/(1024)) + \" KB\", str(modified)))\n\n        # Sort list of tuples by the first element, size.\n        pairs.sort(key=lambda s: s[1])\n    \n    log_list(pairs, desc=\"Files from '{}'\".format(path))\n\ndef remove_file(file):\n    \n    if file is not None and os.path.isfile(file):\n        #log(\"Removed file...{}\".format(file))\n        os.remove(file)\n\ndef has_gpu():\n    return True\n    #return len(K.tensorflow_backend._get_available_gpus()) > 0 \n\ndef move_files(source_dir, target_dir):\n    files = os.listdir(source_dir)\n    files.sort()\n    for f in files:\n        src = source_dir+f\n        dst = target_dir+f\n        copyfile(src,dst)\n    \nif False:\n    log_dir(\"../input\")\n    log_dir(\"../working\")\n    log_current_memory(\"In the beginning\")\n    log(\"CPU Count:{}\".format(mpc.cpu_count()))\n    log(f\"Has GPU: {has_gpu()}, count: {len(K.tensorflow_backend._get_available_gpus())}\")\n    #rmtree('../working/sincerely-gpu-lstm')\n    \nif gExternalData:\n    if os.path.isfile('../working/MATRIX_GOOGLENEWS') == False:\n        log(\"Moving files from input to working\")\n        move_files('../input/sincerely-gpu-lstm/', '../working/')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"426f320eb100cd45ac05f08d4a9d53103fc1cb29"},"cell_type":"markdown","source":"## Profiler utility for performance monitoring\nPut  '@profile\" above any method, and performance is automatically tracked.\nSingle threaded only."},{"metadata":{"trusted":true,"_uuid":"c53e577318d2cf0b86fbe40bc291ebfc8d576847"},"cell_type":"code","source":"import time\nfrom functools import wraps\n\nPROF_DATA = {}\n\nclass_extract_regex = re.compile('^.*(main__.)(.*)\\'>$', re.IGNORECASE)\n\ndef profile(fn):\n    @wraps(fn)\n    def with_profiling(*args, **kwargs):\n        start_time = time.perf_counter()\n\n        ret = fn(*args, **kwargs)\n        \n        elapsed_time = time.perf_counter() - start_time\n\n        if len(args) > 0:\n            profile_key = str(type(args[0])) + \":::\" + fn.__name__\n        else:\n            profile_key = fn.__name__\n                \n        if profile_key not in PROF_DATA:\n            PROF_DATA[profile_key] = [0, []]\n        PROF_DATA[profile_key][0] += 1\n        PROF_DATA[profile_key][1].append(elapsed_time)\n    \n        return ret\n    \n    with_profiling.__wrapped__ = fn\n    \n    return with_profiling\n\ndef log_prof_data():\n    for fname, data in PROF_DATA.items():\n        max_time = max(data[1])\n        avg_time = sum(data[1]) / len(data[1])\n        print (\"Function %s called %d times. \" % (fname, data[0]))\n        print ('Execution time max: %.3f s, average: %.3f s' % (max_time, avg_time))\n\ndef get_prof_data():\n    headers = (\"Class\", \"Function\", \"Call Frequency\", \"Max Time (m)\", \"Avg Time (m)\")\n    contents = []\n    \n    for profile_key, data in PROF_DATA.items():\n        max_time = round(1000*max(data[1])/(1000*60), 3)\n        avg_time = round(1000*sum(data[1]) / (len(data[1])*1000*60), 3)\n                    \n        if ':::' in profile_key:\n            src = profile_key.split(':::')\n            if src is None:\n                className = src\n            elif class_extract_regex.match(src[0]) is None:\n                className = \"classmethod\" + str(src[0])\n            else:\n                className = class_extract_regex.match(src[0])[2]\n        else:\n            className = \"-\"\n            src = ['-', profile_key]\n            \n        contents.append((className, src[1], data[0], max_time, avg_time))\n        \n    return headers, contents\n\ndef pp_prof_data():\n    headers, contents = get_prof_data()\n    tmp = (headers,)\n    tmp += tuple(contents)\n    log_list(tmp)\n\ndef clear_prof_data():\n    global PROF_DATA\n    PROF_DATA = {}\n    \n\nclass Profiler:\n    def __init__(self):\n        clear_prof_data()\n    \n    def start(self):\n        clear_prof_data()\n    \n    def get_data(self):\n        return get_prof_data()\n    \ngProfiler = Profiler()\n\n@profile\ndef save_binary(obj, file):\n    log(f\"Saving file: {file}\")\n    pickle.dump(obj, open(file, \"wb\"), protocol=pickle.HIGHEST_PROTOCOL)\n\nclass TestProfiler:\n    def __init__(self):\n        pass\n    \n    @profile\n    def test_method(self, hello=1):\n        log(\"In Test Method:\" + str(hello))\n    \n    @classmethod\n    @profile\n    def class_method(cls, hello=\"hi\"):\n        log(\"In Class Method: \" + str(hello))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8e867e514bab90033949170af3293dbd7ba32ddc"},"cell_type":"code","source":"if False:\n    test = TestProfiler()\n    clear_prof_data()\n    test.test_method()\n    test.class_method(hello='hi')\n    pp_prof_data()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7c25fd781065b403a06d2acff6c9d89bc9657ee1"},"cell_type":"markdown","source":"## MultiProcessing Utility\n\nI am new to multiprocessing, especially on Python.\n\n* Most likely python has these facilities, but easier to implement my intentions here than learn new stuff due to time.\n* Assumes that each worker process will write output to a file (map out), and when they are all done, the outputs will be reduced\nand written out.\n* Reduction handles csv."},{"metadata":{"trusted":true,"_uuid":"d2a634f603a90242537c8376f7ce5b8d0fea9550"},"cell_type":"code","source":"class TestMPCWithFileCache:\n    def __init__(self, cache=None):\n        self.cache = cache\n    \n    def process(self, cpu_id, cpu_count, cache_file, v1, v2, kv3=[1, 2, 4]):\n        self.cache = {\"foo\":\"bar\", \"Hello\" : v1, \"Hey\": cpu_id*cpu_count}\n        save_binary(self.cache, cache_file)\n        f = open(cache_file+\".csv\", \"w\")\n        f.write(\"key,value\\r\\n\")\n        for k, v in self.cache.items():\n            f.write(\"{},{}\\r\\n\".format(k, v))\n        f.close()\n        \nclass MPHelper:\n    def __init__(self, file_prefix, verbose=False):\n        self.output_file = file_prefix\n        self.file_prefix = file_prefix + \"_{}\"\n        self.verbose = verbose\n        \n    @classmethod\n    def range_partition(cls, total_count, cpu_id):\n        cpu_count = mpc.cpu_count()\n        start = 0\n        end = 0\n        \n        n = int(total_count/cpu_count)\n        start = cpu_id*n\n        end = start + n\n        if cpu_id + 1 == cpu_count:\n            r = int(total_count % cpu_count)\n            end += r\n        \n        return range(start, end)\n    \n    # All target methods must have a signature : func(self, cpu_id, cpu_count, cache_file, ...)\n    def map_process(self, target, *args, **kwargs):\n        cpu_count = mpc.cpu_count()\n        jobs = []\n        for i in range(cpu_count):\n            mpc_args = ()\n            mpc_args += tuple([i])\n            mpc_args += tuple([cpu_count])\n            mpc_args += tuple([self.file_prefix.format(i)])\n            mpc_args += args\n            if self.verbose:\n                log_list(mpc_args, desc=\"Starting process on CPU {}\".format(i))\n            p = mpc.Process(target=target, args=mpc_args, kwargs=kwargs)\n            jobs.append(p)\n        \n        def spawn():\n            [j.start() for j in jobs]\n            [j.join() for j in jobs]\n            \n        \n        is_wrapped = hasattr(target, \"__wrapped__\")\n        if is_wrapped == False:\n            spawn()\n        else:\n            start_time = time.perf_counter()\n            spawn()\n            elapsed_time = time.perf_counter() - start_time\n        \n            class_name = str(target.__self__)\n            profile_key = class_name + \":::\" + target.__name__        \n            if profile_key not in PROF_DATA:\n                PROF_DATA[profile_key] = [0, []]\n            PROF_DATA[profile_key][0] += len(jobs)\n            PROF_DATA[profile_key][1].append(elapsed_time)\n\n        \n    # Helper function\n    def cpu_cache_name(self, cpu_id=None, is_csv=False):\n        \n        count = cpu_id\n        if count is None:\n            count = mpc.cpu_count()\n        for cpu_index in range(count):\n            retval = self.file_prefix.format(cpu_index)\n            if is_csv:\n                retval += \".csv\"\n            if os.path.isfile(retval):\n                yield retval\n            else:\n                #if gWIP:\n                #    log(f\"Not found {retval}\")\n                yield None\n                break\n    \n    # Reduces output files from worker processes to single file.\n    # Handles csv files, if they exist.\n    @profile\n    def reduce(self, clean=True):\n        \n        # Pickled binary object\n        retval = []\n        for cpu_cached_file in self.cpu_cache_name():\n            if cpu_cached_file is None:\n                continue\n            tmp = pickle.load(open(cpu_cached_file, \"rb\"))\n            if tmp is not None:\n                retval.append(tmp)\n        if len(retval) > 0:\n            save_binary(retval, self.output_file)\n        \n        # If there is a CSV file, then reduce.\n        merged_df = None\n        for cpu_cached_file in self.cpu_cache_name(is_csv=True):\n            if cpu_cached_file is None:\n                continue\n            cpu_df = pd.read_csv(cpu_cached_file)\n            if merged_df is None:\n                merged_df = cpu_df\n            else:\n                merged_df = merged_df.append(cpu_df, ignore_index=True)\n            dest_file = self.output_file + \".csv\"\n            merged_df.to_csv(dest_file, index=False)\n            \n        if retval is None:\n            retval = merged_df\n            \n        if clean:\n            self.clean_cache(verbose=False)\n            \n        return retval\n    \n    def clean_cache(self, verbose=True, remove_output=False):\n        if verbose:\n            log(\"Before cleanup...\")\n            log_dir('../working')\n            log(\"Starting Cleanup of MultiProcessing Cache\")\n            \n        for cache_file in self.cpu_cache_name():\n            remove_file(cache_file)\n        \n        for cache_file in self.cpu_cache_name(is_csv=True):\n            remove_file(cache_file)\n        \n        if remove_output:\n            remove_file(self.output_file)\n            remove_file(self.output_file+\".csv\")\n        \n        if verbose:\n            log(\"Finished cleanup of MultiProcessing Cache\")\n            log_dir('../working')\n            \n    def inspect(self):\n        if not gInspect:\n            return\n        \n        log(\"<h1>Inspect MPHelper</h1>\")\n        results = self.reduce(clean=False)\n        log(\"Merged results from pickle\")\n        log_list(results)\n        for item in results:\n            log(\"Results from CPU {}\".format(results.index(item)))\n            log_list(item)\n        \n        dest_file = self.output_file + \".csv\"\n        if os.path.isfile(dest_file):\n            log(\"CSV output from reduced file as pandas frame.\")\n            df = pd.read_csv(dest_file)\n            log(df.to_html())\n        \n    @classmethod\n    def testme(cls):\n        \n        clear_prof_data()\n        tmp = []\n        n = 23\n        for cpu_id in range(mpc.cpu_count()):\n            tmp.append(str(MPHelper.range_partition(n, cpu_id)))\n        log_list(tmp, desc=\"Test range partitioning by CPU count for length 23\")\n    \n        log_dir(\"../working\")\n        me = cls(\"Test_MPFileCache\", verbose=True)\n        test = TestMPCWithFileCache()\n        me.map_process(test.process, \"hello\", \"world\", kv3=[100, 300, 400])\n        me.inspect()\n        me.clean_cache(remove_output=True)\n        del me\n        del test\n        pp_prof_data()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f96e4756ac15521675641d21183a082c5155975d","scrolled":false},"cell_type":"code","source":"if False:\n    compute = MPHelper.testme()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67ab1ec5a8cf72d5f03cea1c5b4f2d5f706052c6"},"cell_type":"markdown","source":"## Embeddings Utility\n\n* Memory and performance efficient way to load embeddings."},{"metadata":{"trusted":true,"_uuid":"4d3635e553290a3734a1d4e221f75c2d1aed9a44"},"cell_type":"code","source":"gEmbeddingsSources = {\n    \"GLOVE\" : {\n        'path' : '../input/embeddings/glove.840B.300d/glove.840B.300d.txt',\n        'mean' : -0.005838498938828707,\n        'std' : 0.4878219664096832\n    },\n    \"WIKINEWS\" : {\n        'path' : '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec',\n        'mean' : -0.0033469984773546457,\n        'std' : 0.10985549539327621\n    },\n    \"PARAGRAM\" : {\n        'path' : '../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt',\n        'mean' : -0.005324783269315958,\n        'std' : 0.4934646189212799\n    },\n    \"GOOGLENEWS\" : {\n        'path' : '../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin',\n        'mean' : 0,\n        'std' : 0\n    },\n}\n        \nclass EmbeddingsControl():\n    \n    OUTPUT_CACHE_FILE_PREFIX = \"MATRIX\"\n    OOV_CACHE_FILE_PREFIX = \"OOV\"\n    \n    @classmethod\n    def output_cache_name(cls, prefix, source_name):\n        return prefix + \"_\" + source_name\n        \n    @classmethod\n    def load_embedding_index(cls, source_name):\n        \n        assert(source_name in gEmbeddingsSources)\n\n        log(\"Loading embeddings index {}\".format(source_name))\n        \n        def get_coefs(word,*arr): \n            return word, np.asarray(arr, dtype='float32')\n    \n        file = gEmbeddingsSources[source_name]['path']\n        retval = None\n        if source_name == \"GOOGLENEWS\":\n            retval = KeyedVectors.load_word2vec_format(file, binary=True, limit=500000)\n        elif source_name == \"PARAGRAM\":\n            retval = dict(get_coefs(*o.split(\" \")) for o in open(file, encoding='latin'))\n        elif source_name == \"WIKINEWS\" or source_name == \"GLOVE\":\n            retval = dict(get_coefs(*o.split(\" \")) for o in open(file) if len(o)>100)\n            \n        return retval\n                    \n    # Memory efficient loading, except for GOOGLENEWS (Word2Vec) which is loaded as a whole.\n    # Returns embeddings matrix for one source.\n    @classmethod\n    @profile\n    def load(cls, source_name, word2index, embed_size=300):   \n        \n        log_current_memory('Before {} processing, memory at:'.format(source_name))\n        \n        assert(source_name in gEmbeddingsSources)\n\n        embedding_matrix = None\n        oov = None\n\n        cache = EmbeddingsControl.output_cache_name(EmbeddingsControl.OUTPUT_CACHE_FILE_PREFIX, source_name)\n        if os.path.isfile(cache):\n            embedding_matrix = pickle.load(open(cache, 'rb'))\n        \n        cache = EmbeddingsControl.output_cache_name(EmbeddingsControl.OOV_CACHE_FILE_PREFIX, source_name)\n        if gWIP and os.path.isfile(cache):\n            oov = pickle.load(open(cache, 'rb'))\n            \n        if (not gWIP and embedding_matrix is not None) or (gWIP and embedding_matrix is not None and oov is not None):\n            return embedding_matrix, oov\n        \n        embeddings_params = gEmbeddingsSources[source_name]\n        nb_words = len(word2index) + 1\n        log(f\"<h3>Embeddings for {source_name}<h3><br> Matrix size: {nb_words}; <br>word2index length: {len(word2index)};<br> max_words : {nb_words}\")\n        found_vecs = set()\n        \n        if source_name == \"GOOGLENEWS\":\n            embeddings_index = KeyedVectors.load_word2vec_format(embeddings_params['path'], binary=True, limit=500000)\n            embedding_matrix = (np.random.rand(nb_words, embed_size) - 0.5)/5.0\n            \n            for word, index in word2index.items():\n                if len(found_vecs) >= nb_words:\n                    break\n                    \n                if word not in embeddings_index: \n                    if word.lower() in embeddings_index:\n                        # Upper-case word's index gets assigned to lower-case word's vector\n                        word = word.lower()\n                    else:\n                        continue\n                \n                embedding_vector = embeddings_index.get_vector(word)\n                if len(embedding_vector) == embed_size:\n                    embedding_matrix[index] = embedding_vector\n                    found_vecs.add(index)\n            \n            del embeddings_index\n        else:\n            # For paragram, let's give the vector for lowercase word to uppercase word\n            lowercase_word2index = {}\n            for w, index in word2index.items():\n                l_w = w.lower()\n                if l_w not in lowercase_word2index:\n                    lowercase_word2index[l_w] = [index]\n                else:\n                    lowercase_word2index[l_w].append(index)\n                    \n            words = ['Quora','Trump','Indian','US','Would','Google']\n            for w in words:\n                indices = lowercase_word2index[w.lower()]\n                log(f\"{w} as lower {w.lower()} : {indices}\")\n            words = [w.lower() for w in words]\n            log_list(words, desc=\"Samples of OOV\")\n            \n            mean = embeddings_params['mean']\n            std = embeddings_params['std']\n            embedding_matrix = np.random.normal(mean, std, (nb_words, embed_size))\n\n            def read_embeddings_file(f):\n                for line in f:                    \n                    if len(line) < 100:\n                        continue\n                    \n                    embedded_word, vec = line.split(' ', 1)\n                    \n                    indices = []\n                    # Allow any form of a word that is in the word2index,\n                    # and if not present, then find the lowercase, and\n                    # and if present, but PARAGRAM, then, find all indices\n                    # of all forms of the word in word2index.\n                    # PARAGRAM has all lowercase.\n                    if embedded_word in word2index:\n                        if source_name == \"PARAGRAM\":\n                            if embedded_word in lowercase_word2index:\n                                indices = lowercase_word2index[embedded_word]\n                            else:\n                                continue\n                        else:\n                            indices = [word2index[embedded_word]]\n                    else:\n                        embedded_word = embedded_word.lower()\n                        if embedded_word in lowercase_word2index:\n                            indices = lowercase_word2index[embedded_word]\n                        else:\n                            continue\n\n                    if embedded_word in words:\n                        log(f\" >>>> Adding indices for {embedded_word} : {indices}\")\n                                \n                    embedding_vector = np.asarray(vec.split(' '), dtype='float32')[:embed_size]\n                    if len(embedding_vector) == embed_size:\n                        for index in indices:\n                            embedding_matrix[index] = embedding_vector\n                            found_vecs.add(index)                    \n                    if len(found_vecs) == nb_words:\n                        break\n                                \n            if source_name == \"PARAGRAM\":\n                with open(embeddings_params['path'], 'r', encoding='utf8', errors='ignore') as f:\n                    read_embeddings_file(f)\n            else:\n                with open(embeddings_params['path']) as f:\n                    read_embeddings_file(f)\n        \n        gc.collect()\n        \n        if gWIP:\n            log_current_memory('After {} processing, memory at'.format(source_name))\n            oov = set()\n            vocab_indices = set(word2index.values())\n            oov = vocab_indices.difference(found_vecs)      \n            log(f\"Length of found vecs dictionary: {len(found_vecs)};\\nLength of oov: {len(oov)}\")\n            cache = EmbeddingsControl.output_cache_name(EmbeddingsControl.OOV_CACHE_FILE_PREFIX, source_name)\n            save_binary(oov, cache)\n        \n        cache = EmbeddingsControl.output_cache_name(EmbeddingsControl.OUTPUT_CACHE_FILE_PREFIX, source_name)\n        save_binary(embedding_matrix, cache)\n        \n        return embedding_matrix, oov\n    \n    @classmethod \n    def clean_cache(cls):\n        for source, _ in gEmbeddingsSources.items():\n            remove_file(EmbeddingsControl.output_cache_name(EmbeddingsControl.OUTPUT_CACHE_FILE_PREFIX, source))\n            remove_file(EmbeddingsControl.output_cache_name(EmbeddingsControl.OOV_CACHE_FILE_PREFIX, source))\n    \n    @classmethod\n    def inspect(cls, source, index2word, word2count, w2v, oov=None):\n        log(f\"<b>Embeddings for {source}</b>\")\n        log(\"Loaded Embeddings  \", len(w2v))\n        \n        if w2v is not None:\n            log(\"Shape of embeddings for {} is {}\".format(source, w2v.shape))\n        \n        if oov is not None and index2word is not None:\n            log(\"Number of words \", len(index2word))\n            log(f\"Number of OOV: {len(oov)}\")\n            oov_words = Counter()\n            for index in oov:\n                word = index2word[index]\n                oov_words[word] = word2count[word]\n            \n            log_list(oov_words.most_common(100), desc=\"OOV for {}\".format(source))\n\n    @classmethod\n    def testme(cls, source=\"GOOGLENEWS\"):\n        me = cls()\n        \n        use_test_samples = False\n        index2word = None\n        word2index = None\n        vocab2count = None\n        \n        if use_test_samples:\n            word2index = {'good': 0, 'bad': 1, 'sincere': 2, 'insincere': 3, 'lkjsdlfajsflkdj' : 4, 'sexual_intercourse':5, 'Donald_Trump' : 6, 'obama' : 7, 'f_**_king' : 8, '?' : 9, \"LOVE\":10, \"Why\":11, \"What\":12}\n            index2word = {}\n            vocab2count = Counter()\n            for k, v in word2index.items():\n                index2word[v] = k\n                if k in vocab2count:\n                    vocab2count[k] += 1\n                else:\n                    vocab2count[k] = 1\n        else:\n            word2index = pickle.load(open('w2index_BOTH', 'rb'))\n            index2word = pickle.load(open('index2w_BOTH','rb'))\n            vocab2count = pickle.load(open('vocab_BOTH', 'rb'))\n        \n        log_current_memory('Before embeddings processing')\n        for a_source, _ in gEmbeddingsSources.items():\n            if a_source == source or source == None:\n                w2v, oov = EmbeddingsControl.load(a_source, word2index)\n                EmbeddingsControl.inspect(a_source, index2word, vocab2count, w2v, oov)\n        del w2v\n        del oov\n        del me\n        gc.collect()\n        log_current_memory('After embeddings processing')\n                \n    @classmethod\n    def precompute_mean_std(cls):\n        \n        def calculate(source):\n            embeddings = EmbeddingsControl.load_embedding_index(source)\n            all_embs = None\n            all_embs = np.stack(embeddings.values())\n            emb_mean,emb_std = all_embs.mean(), all_embs.std()\n            log(\"{} Embeddings... mean: {} std:{}\".format(source, emb_mean, emb_std))\n            del all_embs\n            del embeddings\n            gc.collect()\n        \n        calculate(\"GLOVE\")\n        calculate(\"WIKINEWS\")\n        calculate(\"PARAGRAM\")\n        # GOOGLENEWS is done differently; doesn't use mean/std (why?? XX)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1250874a01a38d132f6ad0077c45fd805b0959e3","scrolled":false},"cell_type":"code","source":"if False:\n    log_current_memory('Start processing embeddings')\n    EmbeddingsControl.clean_cache()\n    log_dir(\"../working/\")\n    EmbeddingsControl.testme(source=None)\n    gc.collect()\n    log_current_memory('Finished processing embeddings')\n    \nif False:\n    log_current_memory('Start processing embeddings')\n    EmbeddingsControl.precompute_mean_std()\n    log_current_memory('Finished processing embeddings')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"610c7fe6e87f061f8afce1d71b6bda454fc3a25f"},"cell_type":"markdown","source":"# DataManager\n\nData loader"},{"metadata":{"trusted":true,"_uuid":"9df2be168f8b0548dcb23216a39eb297ef3da282"},"cell_type":"code","source":"class DataManager:\n    \n    INPUT_TRAINING_DATA = '../input/train.csv'\n    INPUT_TEST_DATA = '../input/test.csv'\n    INPUT_GENERATED_DATA = 'gen_q.csv'\n    OUTPUT_CACHE_FILE = 'orig_gen.csv'\n    gInstance = None\n    DEV_LIMIT = 5000\n    \n    def __init__(self):\n        self.training_data = None \n        self.test_data = None        \n        \n        if gExternalData:\n            DataManager.INPUT_TRAINING_DATA = '../input/quora-insincere-questions-classification/train.csv'\n            DataManager.INPUT_TEST_DATA = '../input/quora-insincere-questions-classification/test.csv'\n        \n    @classmethod\n    def instance(cls):\n        if DataManager.gInstance == None:\n            return cls()\n        else:\n            return gInstance\n        \n    def load(self, orig=True, combined=False, test=False, source_file=None):\n\n        if source_file is not None:\n            self.training_data = pd.read_csv(source_file)\n        elif combined and os.path.isfile(DataManager.OUTPUT_CACHE_FILE):\n            log(\"Found combined training data file\")\n            self.training_data = pd.read_csv(DataManager.OUTPUT_CACHE_FILE)\n        elif combined or orig:\n            self.training_data = pd.read_csv(DataManager.INPUT_TRAINING_DATA)\n\n        if gLimit and self.training_data is not None and combined == False:\n                self.training_data = self.training_data.sample(DataManager.DEV_LIMIT)\n        \n        if test:\n            self.test_data = pd.read_csv(DataManager.INPUT_TEST_DATA)\n            if gLimit:\n                self.test_data = self.test_data[:min(len(self.test_data), DataManager.DEV_LIMIT)]\n                \n    def save_combined_data(self):\n        \n        if os.path.isfile(DataManager.INPUT_GENERATED_DATA):\n            self.load(orig=True)\n            more_training_data = pd.read_csv(DataManager.INPUT_GENERATED_DATA)\n            combined = self.training_data.append(more_training_data, ignore_index=True)\n            combined.to_csv(DataManager.OUTPUT_CACHE_FILE)\n        \n    def memclean(self):\n        del self.training_data\n        self.training_data = None\n        del self.test_data\n        self.test_data = None\n        \n    @classmethod\n    def clean_cache(cls):\n        log(\"Cleaning combined original and generated questions file.\")\n        remove_file(DataManager.OUTPUT_CACHE_FILE)\n        \n    def inspect(self):\n        if not gInspect:\n            return\n        \n        if self.training_data is not None:\n            log(\"Shape of training data {}\".format(self.training_data.shape))\n            log(\"Length of unique q_id {}\".format(len(self.training_data.iloc[0,:]['qid'])))\n                  \n        if self.test_data is not None:\n            log(\"Shape of test data {}\".format(self.test_data.shape))                  \n    \n    @classmethod\n    @profile\n    def testme(cls):\n        DataManager.clean_cache()\n        me = DataManager.instance()\n        log(\"<hr>Original training nd test data\")\n        me.load(test=True)\n        me.inspect()\n        me.save_combined_data()\n        log(\"<hr>Combined training data\")\n        me.load(combined=True)\n        me.inspect()\n        me.memclean()\n        del me\n        gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4196df411567e90b9a46a608be9df40fa7895365"},"cell_type":"code","source":"if False:\n    DataManager.testme()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"beca47067b2b7f3b7a88ee8ae873dda0c880f380"},"cell_type":"markdown","source":"# Section: Preprocessing Text\n\n## Text Handling\n\n* So far, we have loaded up the files provided by the competition, and added functionality for their content.\n* Now, we need to :\n    * Build some intuition about the entire dataset (eyeballing stats and samples)\n    * Preprocess the text\n    * Generate synthetic data (see how below)\n    * Create word indices\n    * Create training sequences (each question is an array of indices, which refer to words in the embeddings matrix)\n    * To save developer time, save to file: embeddings matrix, word indices, training sequences\n*  Generate synthetic data\n     * Get external list of terms that are expletives, ethnicities, and politics using word2vec word similarities.\n     * For all insincere questions, if there is a term that matches any list above, create new questions by interchanging the terms from the appropriate list.\n     * For manual review as well as saving dev time, save generated questions to file."},{"metadata":{"trusted":true,"_uuid":"d2670f5b54fbe2835903be32e377540453799b63","_kg_hide-input":false},"cell_type":"code","source":"class TextualDataControl:\n    \n    PUNCTUATIONS = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n    '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n    '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n    '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n    '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]\n        \n    CONTRACTIONS_MAP = {\"ain't\": \"is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"'cause\": \"because\", \"could've\": \"could have\", \"couldn't\": \"could not\", \"didn't\": \"did not\",  \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hasn't\": \"has not\", \"haven't\": \"have not\", \"he'd\": \"he would\",\"he'll\": \"he will\", \"he's\": \"he is\", \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\", \"how's\": \"how is\",  \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \"i'd've\": \"i would have\", \"i'll\": \"i will\",  \"i'll've\": \"i will have\",\"i'm\": \"i am\", \"i've\": \"i have\", \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\", \"it'll\": \"it will\", \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \"ma'am\": \"madam\", \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \"mustn't've\": \"must not have\", \"needn't\": \"need not\", \"needn't've\": \"need not have\",\"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \"oughtn't've\": \"ought not have\", \"shan't\": \"shall not\", \"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \"she'll've\": \"she will have\", \"she's\": \"she is\", \"should've\": \"should have\", \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\"so's\": \"so as\", \"this's\": \"this is\",\"that'd\": \"that would\", \"that'd've\": \"that would have\", \"that's\": \"that is\", \"there'd\": \"there would\", \"there'd've\": \"there would have\", \"there's\": \"there is\", \"here's\": \"here is\",\"they'd\": \"they would\", \"they'd've\": \"they would have\", \"they'll\": \"they will\", \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \"to've\": \"to have\", \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \"we'll\": \"we will\", \"we'll've\": \"we will have\", \"we're\": \"we are\", \"we've\": \"we have\", \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \"what're\": \"what are\",  \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\", \"when've\": \"when have\", \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \"who'll\": \"who will\", \"who'll've\": \"who will have\", \"who's\": \"who is\", \"who've\": \"who have\", \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \"y'all'd\": \"you all would\",\"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\"y'all've\": \"you all have\",\"you'd\": \"you would\", \"you'd've\": \"you would have\", \"you'll\": \"you will\", \"you'll've\": \"you will have\", \"you're\": \"you are\", \"you've\": \"you have\" }\n    \n    MISPELL_MAP = {'to':None, 'a':None, 'of':None,'.':None,'and': 'plus','':None,'akistani':'Pakistani','Snapchat':'online social medium','WhatsApp':'messaging app','huminity':'humanity','motherfuckin':'motherfucking','motherfuckingg':'motherfucking','mujahiddin':'mujahideen','descpicable':'despicable','MOZLEMS':'Muslims','Quorans':'Quora users','Quoran':'Quora user','Niccaragua':'Nicaragua','Shivdharma':'Shiva dharma','colour': 'color', 'neighbour':'neighbor','behaviour':'behavior','favour':'favor','favoured':'favored','generalised':'generalized','realise':'realize','centre': 'center', 'favourite': 'favorite', 'travelling': 'traveling', 'counselling': 'counseling', 'theatre': 'theater', 'cancelled': 'canceled', 'labour': 'labor', 'organisation': 'organization', 'wwii': 'world war 2', 'citicise': 'criticize', 'youtu ': 'youtube ', 'Qoura': 'Quora', 'sallary': 'salary', 'Whta': 'What', 'narcisist': 'narcissist', 'howdo': 'how do', 'whatare': 'what are', 'howcan': 'how can', 'howmuch': 'how much', 'howmany': 'how many', 'whydo': 'why do', 'doI': 'do I', 'theBest': 'the best', 'howdoes': 'how does', 'mastrubation': 'masturbation', 'mastrubate': 'masturbate', \"mastrubating\": 'masturbating', 'pennis': 'penis', 'Etherium': 'Ethereum', 'narcissit': 'narcissist', 'bigdata': 'big data', '2k17': '2017', '2k18': '2018', 'qouta': 'quota', 'exboyfriend': 'ex boyfriend', 'airhostess': 'air hostess', \"whst\": 'what', 'watsapp': 'whatsapp', 'demonitisation': 'demonetization', 'demonitization': 'demonetization', 'demonetisation': 'demonetization', 'pokémon': 'pokemon'}\n    \n    SPECIAL_CHARS_MAP = {\"‘\": \"'\", \"₹\": \"e\", \"´\": \"'\", \"°\": \"\", \"€\": \"e\", \"™\": \"tm\", \"√\": \" sqrt \", \"×\": \"x\", \"²\": \"2\", \"—\": \"-\", \"–\": \"-\", \"’\": \"'\", \"`\": \"'\", '“': '\"', '”': '\"', '“': '\"', \"£\": \"e\", '∞': 'infinity', 'θ': 'theta', '÷': '/', 'α': 'alpha', '•': '.', 'à': 'a', '−': '-', 'β': 'beta', '∅': '', '³': '3', 'π': 'pi', '\\u200b': ' ', '…': ' ... ', '\\ufeff': '', 'करना': '', 'है': ''}\n    \n    SPACES_LIST = ['\\u200b', '\\u200e', '\\u202a', '\\u202c', '\\ufeff', '\\uf0d8', '\\u2061', '\\x10', '\\x7f', '\\x9d', '\\xad', '\\xa0']\n    \n    def __init__(self, data, target_column = 'target'):        \n        self.rawdata = data\n        self.label_column_name = target_column    \n        \n    def memclean(self):\n        del self.rawdata\n        self.rawdata = None\n        \n    @classmethod                \n    def find_string(cls, src, target):\n        \n        retval = False\n        if re.search(r\"\\b\" + re.escape(src.lower()) + r\"\\b\", target.lower()):\n          retval = True\n        \n        return retval\n    \n    @classmethod\n    def sent_capitalize(cls, src):\n        \n        retval = None\n        sentences = sent_tokenizer.tokenize(src)\n        sentences = [sent[:1].upper() + sent[1:] for sent in sentences]\n        retval = ' '.join(sentences)\n        \n        return retval\n    \n    @classmethod\n    def fix_special_chars(cls, phrase):\n        \n        if not re.match(\"^[a-zA-Z0-9 _]*$\", phrase):\n            all_special_chars = TextualDataControl.SPECIAL_CHARS_MAP.keys()\n            for s in all_special_chars:\n                if s in phrase:\n                    phrase = phrase.replace(s, TextualDataControl.SPECIAL_CHARS_MAP[s])\n                    \n        return phrase\n    \n    @classmethod\n    def decontract(cls, phrase):\n        \n        specials = [\"’\", \"‘\", \"´\", \"`\"]\n        for s in specials:\n            if s in phrase:\n                phrase = phrase.replace(s, \"'\")\n        phrase = ' '.join([TextualDataControl.CONTRACTIONS_MAP[t.lower()] if t.lower() in TextualDataControl.CONTRACTIONS_MAP else t for t in phrase.split(\" \")])\n        possessive = \"'s\"\n        if possessive in phrase:\n            phrase = phrase.replace(possessive, \" \")\n        \n        return phrase\n    \n    @classmethod\n    def despace(cls, phrase):\n        for space in TextualDataControl.SPACES_LIST:\n            if space in phrase:\n                phrase = phrase.replace(space, ' ')\n        \n        phrase = phrase.strip()\n        phrase = re.sub('\\s+', ' ', phrase)\n        \n        return phrase\n        \n        \n    @classmethod\n    def denumber(cls, phrase):\n        \n        if bool(re.search(r'\\d', phrase)):\n            phrase = re.sub('[0-9]{5,}', '#####', phrase)\n            phrase = re.sub('[0-9]{4}', '####', phrase)\n            phrase = re.sub('[0-9]{3}', '###', phrase)\n            phrase = re.sub('[0-9]{2}', '##', phrase)\n        \n        return phrase\n    \n    @classmethod\n    def fillna(cls, phrase, fill_value=\"_na_\"):\n        \n        if phrase is None or phrase.strip() is None:\n            phrase = \"_na_\"\n        return phrase\n    \n    @classmethod\n    def depunctuate(cls, phrase, remove=False):\n        \n        for punct in TextualDataControl.PUNCTUATIONS:\n            if punct in phrase:\n                if not remove:\n                    # phrase = phrase.replace(punct, f' {punct} ')\n                    phrase = re.sub('(['+punct+']{2,})', r' \\1 ', phrase)\n                else:\n                    phrase = phrase.replace(punct, f' ')                    \n        \n        phrase = phrase.replace('  ', ' ').strip()\n        \n        return phrase\n        \n    @classmethod\n    def respell(cls, a_word):\n\n        retval = a_word\n        if a_word in TextualDataControl.MISPELL_MAP:\n            retval = TextualDataControl.MISPELL_MAP[a_word]    \n        return retval\n        \n    def inspect(self):\n        if not gInspect:\n            return\n        \n        log(\"--------\" + type(self).__name__ + \"---------\")\n        log(\"Raw input data shape : \", self.rawdata.shape)\n        log(\"Number of unique labels: \", self.rawdata[self.label_column_name].nunique())\n        tmp = pd.DataFrame([self.rawdata[\"target\"].unique(), self.rawdata[\"target\"].value_counts()], index=[\"Label\", \"Count\"])\n        display(HTML(\"<h3>\" + \"Labels\" + \"</h3>\" + tmp.to_html(max_rows=10)))\n        #log_list(, desc=\"Raw Data Statistics\")\n        #log(\"Test shape : \",self.test.shape)\n        display(HTML(self.rawdata.to_html(max_rows=5)))\n        display(HTML(self.rawdata[self.rawdata[self.label_column_name] == 1].to_html(max_rows=5)))\n\n        f,ax=plt.subplots(1,2,figsize=(18,8))\n        self.rawdata[self.label_column_name].value_counts().plot.pie(explode=[0,0.1],autopct='%1.1f%%',ax=ax[0],shadow=True)\n        ax[0].set_title('By ' + self.label_column_name)\n        ax[0].set_ylabel('')\n\n        sns.countplot(self.label_column_name,data=self.rawdata,ax=ax[1])\n        ax[1].set_title('By ' + self.label_column_name)\n        plt.show()\n        \n    @classmethod\n    def testme(cls):\n        app = DataManager.instance()\n        app.load(orig=True)\n        td = cls(app.training_data)\n        td.inspect()\n        td.memclean()\n        del td\n        app.memclean()\n        del app\n        gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b1c695164c4e27dcaf863a5178f6a6ca8abdabb7","scrolled":false},"cell_type":"code","source":"if False:\n    TextualDataControl.testme()\n\nif False:\n    log(TextualDataControl.decontract(\"shouldn't\"))\n    log(TextualDataControl.decontract(\"child's\"))\n    log(TextualDataControl.denumber(\"Is it advisable to take up 20 (4 credits and 4 non credits) courses at Harvard Summer School in 3 weeks?\"))\n    log(TextualDataControl.fix_special_chars(\"∞ करना `foo_bar` … right? f_**_king\"))\n    log(str(TextualDataControl.respell('colour')))\n    log(\"Test \\u200b Test\")\n    tmp = \"Is it advisable to >>\\u200b<< take up 20 (4 credits and 4 non credits) courses at Harvard Summer School in 3 weeks?\"\n    log(tmp)\n    log(str(TextualDataControl.despace(tmp)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"69f932cd0fcd33833d32b184aa7f62e16c9c91fb"},"cell_type":"markdown","source":"## QuoraPreprocessor\n\nAssimilation of lots of good and partially good advice on the Kaggle kernels for this competition."},{"metadata":{"trusted":true,"_uuid":"c2c3dfa4ceb24613d7f858d5adf127ccbe00947c"},"cell_type":"code","source":"class QuoraPreprocessor(TextualDataControl):\n    \n    CLEANED_OUTPUT_FILE = \"cleaned_q\"\n    VOCAB_OUTPUT_FILE = \"vocab\"\n    W2INDEX_OUTPUT_FILE = \"w2index\"\n    INDEX2W_OUTPUT_FILE = \"index2w\"\n    SEQUENCES_OUTPUT_FILE = \"sequences\"\n    MAX_WORDS_IN_QUESTION = 45\n    MAX_VOCAB_SIZE = 60000\n        \n    def __init__(self, training_data=None, test_data=None, is_gen=True):\n        TextualDataControl.__init__(self, None)\n        \n        self.datasource = None\n        self.rawdata = None\n        \n        self.training_data = training_data\n        self.test_data = test_data\n        self.is_generated_data = is_gen\n\n        self.word2count = None\n        self.word2index = None\n        self.index2word = None    \n        self.sample2sequence = None\n    \n    def output_cache_name(self, file_prefix, datasource=None, is_csv=False):\n        \n        if datasource is None:\n            datasource = self.datasource\n        retval = file_prefix + \"_\"  + datasource\n        if is_csv:\n            retval += \".csv\"\n        return retval\n    \n    @classmethod\n    @profile\n    def make_tokens(cls, text):\n    \n        words = list(text.split())\n        multi_words = {}\n        punct_to_split = \"[a-zA-Z0-9']+|[.,!?;]\"\n        for word_index in range(len(words)):\n            words[word_index] = words[word_index].strip(string.punctuation)\n            found_multi_words = re.findall(r\"{}\".format(punct_to_split), words[word_index])\n            if len(found_multi_words) > 1:\n                multi_words[word_index] = found_multi_words\n\n        if len(multi_words) > 0:\n            offset = 0\n            for index, replacements in multi_words.items():\n                before = words[:index+offset]\n                after = words[index+offset+1:]\n                #log(\"Before:{}<br>Replacement:{}<br>After:{}\".format(before, replacements, after))\n                words = before + replacements + after\n                offset += len(replacements)-1\n\n        # Spelling errors - tokenwise\n        retval = []\n        for w in words: \n            if w is not None and len(w) > 0:\n                replacement = TextualDataControl.respell(w)\n                if replacement is not None and len(replacement.strip()) > 0:\n                    tmp = replacement.split()\n                    for item in tmp:\n                        if len(item.strip()) > 0 and item not in gStopWords:\n                            retval.append(item)\n\n        return retval\n    \n    # Does more than just split the phrase.\n    # Returns an array of words\n    @profile\n    def tokenize(self, text):\n        return QuoraPreprocessor.make_tokens(text)\n        \n    \n    @profile\n    def cleaning_proc(self, cpu_id, cpu_count, cachefile):\n        \n         # RawData\n        q_range = MPHelper.range_partition(len(self.rawdata), cpu_id)\n        rawdata = self.rawdata[q_range.start:q_range.stop]\n        \n        # Pre-processing\n        preprocessed_text = rawdata[\"question_text\"].apply(lambda x: self.fix_special_chars(x))\n        preprocessed_text = preprocessed_text.apply(lambda x: self.decontract(x))\n        preprocessed_text = preprocessed_text.apply(lambda x: self.denumber(x))\n        preprocessed_text = preprocessed_text.apply(lambda x: self.despace(x))\n        preprocessed_text = preprocessed_text.apply(lambda x: self.tokenize(x))\n        counts = preprocessed_text.apply(lambda x: len(x))\n        preprocessed_text = preprocessed_text.apply(lambda x: ','.join(x))\n        \n        rawdata = self.rawdata.iloc[q_range.start:q_range.stop].assign(cleaned_text=preprocessed_text,num_words=counts)\n        rawdata.to_csv(cachefile+\".csv\")\n    \n    # Word:Count\n    @classmethod\n    def make_word2count(cls, results, words):\n        words = words.split(\",\")\n        for word in words:\n            if word in results:\n                results[word] += 1\n            else:\n                results[word] = 1\n                \n        return len(words)\n\n    def vocab_proc(self, cpu_id, cpu_count, cachefile):\n        \n        # Raw, cleaned data\n        cleaned_q_csv_file = self.output_cache_name(QuoraPreprocessor.CLEANED_OUTPUT_FILE, is_csv=True)\n        rawdata = pd.read_csv(cleaned_q_csv_file)\n        q_range = MPHelper.range_partition(len(rawdata), cpu_id)\n        \n        # Produce vocab dictionary\n        results = {}\n        maxlen = 0\n        maxqid = -1\n        num_words = []\n        for i in q_range:\n            cleaned_text = rawdata.loc[i][\"cleaned_text\"]\n            word_count = QuoraPreprocessor.make_word2count(results, str(cleaned_text))\n            num_words.append(word_count)\n            maxlen = max(maxlen, word_count)\n            if maxlen == word_count:\n                maxqid = i\n        \n        results = {'maxlen' : maxlen, 'maxq' : maxqid, 'vocab' : results }\n        save_binary(results, cachefile)\n        \n    # The maxlength of questions at 60 arrived at after counting words in training set.\n    def sequences_proc(self, cpu_id, cpu_count, cachefile, maxlen=60):\n        # Raw, cleaned data\n        cleaned_q_csv_file = self.output_cache_name(QuoraPreprocessor.CLEANED_OUTPUT_FILE, is_csv=True)\n        rawdata = pd.read_csv(cleaned_q_csv_file)\n        q_range = MPHelper.range_partition(len(rawdata), cpu_id)\n        \n        word2index = pickle.load(open(self.output_cache_name(QuoraPreprocessor.W2INDEX_OUTPUT_FILE, datasource=\"BOTH\"), \"rb\"))\n        \n        results = []\n        for i in q_range:\n            cleaned_text = str(rawdata.loc[i][\"cleaned_text\"])\n            cleaned_tokens = cleaned_text.split(\",\")\n            a_sequence = [word2index[token] for token in cleaned_tokens if token in word2index]\n            results.append(a_sequence)\n        \n        results = pad_sequences(results, maxlen=maxlen)\n        save_binary(results, cachefile)\n        \n    @profile\n    def make_clean_questions(self):\n        \n        source_file = self.output_cache_name(QuoraPreprocessor.CLEANED_OUTPUT_FILE, is_csv=True)\n        if os.path.isfile(source_file):\n            self.rawdata = pd.read_csv(source_file)\n            return\n        \n        # Clean Question Text\n        mp = MPHelper(self.output_cache_name(QuoraPreprocessor.CLEANED_OUTPUT_FILE, is_csv=False))\n        mp.map_process(self.cleaning_proc)\n        mp.reduce(clean=True)\n\n        # load reduced csv\n        self.rawdata = pd.read_csv(source_file)\n        \n    @profile\n    def make_vocab(self, lower=False):\n        \n        cachefile = self.output_cache_name(QuoraPreprocessor.VOCAB_OUTPUT_FILE)\n        if os.path.isfile(cachefile):\n            word2count = pickle.load(open(cachefile, 'rb'))\n            return word2count\n        \n        # Collect unique words (vocab)\n        mp = MPHelper(QuoraPreprocessor.VOCAB_OUTPUT_FILE)\n        mp.map_process(self.vocab_proc)\n        merged_results = mp.reduce(clean=True)\n        \n        word2count = Counter()\n        for cpu_result in merged_results:\n            cpu_vocab = cpu_result['vocab']            \n            for word, count in cpu_vocab.items():\n                if lower:\n                    word = word.lower()\n                if word in word2count:\n                    word2count[word] += count\n                else:\n                    word2count[word] = count\n        \n        save_binary(word2count, cachefile)\n        return word2count\n\n    # Combined training and test vocab into 1.\n    def combine_vocab(self):\n        \n        cachefile = self.output_cache_name(QuoraPreprocessor.VOCAB_OUTPUT_FILE, datasource=\"BOTH\")\n        if os.path.isfile(cachefile):\n            self.word2count = pickle.load(open(cachefile, \"rb\"))\n            return\n\n        vocab_file = self.output_cache_name(QuoraPreprocessor.VOCAB_OUTPUT_FILE)\n        test_file = self.output_cache_name(QuoraPreprocessor.VOCAB_OUTPUT_FILE, datasource='test')\n        # word2count is a Counter dictionary\n        training_vocab = pickle.load(open(vocab_file, 'rb'))\n        test_vocab = pickle.load(open(test_file, 'rb'))\n        \n        for wrd, count in test_vocab.items():\n            if wrd in training_vocab:\n                training_vocab[wrd] += count\n            else:\n                training_vocab[wrd] = count\n            \n        self.word2count = training_vocab\n        remove_file(vocab_file)\n        remove_file(test_file)\n        save_binary(self.word2count, cachefile)\n        \n    @profile\n    def make_word2index(self):\n        \n        cachefile1 = self.output_cache_name(QuoraPreprocessor.W2INDEX_OUTPUT_FILE, datasource=\"BOTH\")\n        cachefile2 = self.output_cache_name(QuoraPreprocessor.INDEX2W_OUTPUT_FILE, datasource=\"BOTH\")\n        \n        if os.path.isfile(cachefile1) and os.path.isfile(cachefile2):\n            self.word2index = pickle.load(open(cachefile1, \"rb\"))\n            self.index2word = pickle.load(open(cachefile2, \"rb\"))\n            return\n        \n        if self.word2count is None:\n            return\n\n        self.word2index = {}\n        self.index2word = {}\n        \n        # Starting index at 1, so that index=0 is reserved.\n        index = 1\n        items = self.word2count.most_common()\n        \n        for (word, _) in items:\n            self.word2index[word] = index\n            if gInspect:\n                self.index2word[index] = word\n            index += 1\n        \n        save_binary(self.word2index, cachefile1)\n        if gInspect:\n            save_binary(self.index2word, cachefile2)\n\n    @profile\n    def make_sequences(self):\n        \n        cachefile = self.output_cache_name(QuoraPreprocessor.SEQUENCES_OUTPUT_FILE)\n        if os.path.isfile(cachefile):\n            self.sample2sequence = pickle.load(open(cachefile, \"rb\"))\n            return\n        \n        # Collect unique words (vocab)\n        mp = MPHelper(QuoraPreprocessor.SEQUENCES_OUTPUT_FILE)\n        mp.map_process(self.sequences_proc, maxlen=QuoraPreprocessor.MAX_WORDS_IN_QUESTION)\n        merged_results = mp.reduce(clean=True)\n        \n        self.sample2sequence = []\n        count = 0\n        for cpu_result in merged_results:\n            self.sample2sequence.extend(cpu_result)\n        \n        save_binary(self.sample2sequence, cachefile)\n\n    # This class was originally written to process some rawdata and write the results to file without focus on \n    # training vs test.This method changes the underlying rawdata and file names for processing training vs test data.\n    def set_mode(self, is_training=False, is_test=False):\n        \n        assert(is_training or is_test)\n        \n        if is_training:\n            self.rawdata = self.training_data\n            if self.is_generated_data:\n                self.datasource = 'gen'\n            else:\n                self.datasource = 'orig'\n        else:\n            self.rawdata = self.test_data\n            self.datasource = 'test'\n        \n    @profile\n    def process(self, sequences=True):        \n\n        self.set_mode(is_training=True)\n        self.make_clean_questions()\n        self.make_vocab()\n        \n        self.set_mode(is_test=True)\n        self.make_clean_questions()\n        self.make_vocab()\n        \n        # Important insight - in many published kernels, the vocabulary is created from \n        # training data only.  This isn't correct, because the model for training and predictions \n        # has the same embeddings matrix.\n        self.set_mode(is_training=True)\n        self.combine_vocab()\n        self.make_word2index()\n        \n        if sequences:\n            self.set_mode(is_test=True)\n            self.make_sequences()                \n\n            self.set_mode(is_training=True)\n            self.make_sequences()\n    \n    @classmethod\n    def clean_cache(cls, cleaned_q=False, vocab=False, word2index=False, sequences=False):\n        \n        # can't use member function output_cache_name\n        def make_path(file_prefix, datasource):\n            return file_prefix + \"_\" + datasource\n        \n        for datasource in ['orig', 'gen', 'test', 'BOTH']:\n            if cleaned_q:\n                remove_file(make_path(QuoraPreprocessor.CLEANED_OUTPUT_FILE, datasource)+\".csv\")\n                remove_file(QuoraPreprocessor.CLEANED_OUTPUT_FILE+\".csv\")\n\n            if vocab:\n                remove_file(make_path(QuoraPreprocessor.VOCAB_OUTPUT_FILE, datasource))\n                remove_file(QuoraPreprocessor.VOCAB_OUTPUT_FILE)\n\n            if sequences:\n                remove_file(make_path(QuoraPreprocessor.SEQUENCES_OUTPUT_FILE, datasource))\n\n        if word2index:\n            remove_file(make_path(QuoraPreprocessor.W2INDEX_OUTPUT_FILE, \"BOTH\"))\n            remove_file(make_path(QuoraPreprocessor.INDEX2W_OUTPUT_FILE, \"BOTH\"))\n\n\n    @classmethod\n    def testme(cls):\n        \n        clear_prof_data()\n        log_current_memory(caption=\"Start Quora Preprocessing...\")\n        app = DataManager.instance()\n        app.load(orig=True, combined=False, test=True)\n        \n        me = QuoraPreprocessor(training_data=app.training_data, test_data=app.test_data, is_gen=False)\n        me.process()\n        texts = [\"The fool of a Took went off on his own!\"]\n        for text in texts:\n            log(f'{me.tokenize(text)}')\n        me.inspect()\n        log_current_memory(caption=\"End Quora Preprocessing...\")\n        me.memclean()        \n        del me\n        \n        app.memclean()\n        del app\n        \n        gc.collect()        \n        pp_prof_data()\n    \n    @classmethod\n    def testme2(cls):\n        app = DataManager.instance()\n        app.load(orig=True, combined=False, test=False)\n        clear_prof_data()\n        me = QuoraPreprocessor(training_data=app.training_data, test_data=None, is_gen=False)\n        me.set_mode(is_training=True)\n        me.make_clean_questions()\n        word2count = me.make_vocab(lower=True)\n        log(f\"Number of words in uncleaned question_text {len(word2count)}\")\n        pp_prof_data()\n        del word2count\n        del me\n        del app\n        gc.collect()\n    \n    def inspect(self):\n        super().inspect()\n        \n        if not gInspect:\n            return\n\n        log(\"--------\" + type(self).__name__ + \"---------\")\n        \n        log_list(self.word2count.most_common(10), desc=\"Vocab2Count - Top 10\")\n        log_list(self.word2index, limit=10, desc=\"Word2Index - Limit 10\")\n        log_list(self.sample2sequence, limit=10, desc=\"Question Sequences - Limit 10\")\n        log(f\"Number of sequences: {len(self.sample2sequence)}\")\n        log(f\"Number of questions:{len(self.rawdata)}\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0edfada9159e5ff4c8be7755c37e9485e69aa4f0","scrolled":true},"cell_type":"code","source":"if False:\n    QuoraPreprocessor.clean_cache(cleaned_q=True, vocab=True, word2index=True, sequences=True)\n    QuoraPreprocessor.testme2()\n    log_dir(\"../working\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cb3ff095080adf0f2e3d74a2116633ca027a6ae9"},"cell_type":"code","source":"if False:\n    questions = ['If men are roughly 40% of all domestic violence victims, why are there no resources to help them? Why is there so much resistance from feminists/women when this topic/topical/tear is brought up? Why are men!women not able to freely speak about these real problems?',\n             'Why don\\'t liberal progressive_voices oppose Amazon.com\\'s potential takeover of state authority as the company ponders 2345 locations for its new HQ?', '\"All countries support Indian Army to occupy Chinese land in Doklan,\" Indian FM claimed. \"Can you name one,\" asked a reporter. \"Hiiimmmmmm,\" replied the Indian FM. Why has India been good at nothing, but false claiming for the past 70 years?']\n    qp = QuoraPreprocessor()\n    for q in questions:\n        tq = TextualDataControl.decontract(q)\n        tq = TextualDataControl.denumber(tq)\n        tq = TextualDataControl.fillna(tq)\n        tq = qp.lowering(tq)\n        log_list(qp.tokenize(tq), desc=q)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd49440a06d7bd7b2e4e4f747838506202e01865"},"cell_type":"markdown","source":"# Section: Synthetic Data Generator\n\n## Overview\n\nObjective:  There are only 6% insincere questions. For training, more data is better.  \nHow can we generate more insincere questions?\n\n In essence: \n Bad things said about one group of people, are likely to be equally bad for another group. And similarly with other taboo topics.\n An expletive can probably be replaced with another similar expletive.\n We don't need to teach model human history. Just that bad things said about a group of people and what is said...can make a question insincere.\n \n* Taboo topics are - \n   - Religion/Ethnicity : group of people\n   - Politics\n   - Sexuality\n   - Words with negative intense emotion (expletives)\n * Generate taboo words from similarities to some set of 'seed' words in word2vec for each topic above.\n * The similar words may fall under 3 kinds: \n    1. words that partially match the seeds, and \n    1.  words that relate the most to the list of partially matched\n    1. words that are similar to the seed, but not the words in the other 2 kinds. These words can be used to expand the seed.\n* Additionally, words that belong in a category should match the type of \"Named Entity\" of the category.\n * For each insincere question:\n     - If a word/phrase in the question is in the list of taboo words that are similar, then\n         - Replace that word/phrase with other similar words or phrases."},{"metadata":{"_uuid":"e1427bdfeb07bb092eeb4d5335bb604c9f0954cb"},"cell_type":"markdown","source":"## TopicalWords\n\nUsed GOOGLENEWS Word2Vec to generate words that are 'similar' to each other, such that they are substitutable.\nSimilarity does not mean synonymity.  It means 'substitutable' for the purposes of insincerity."},{"metadata":{"trusted":true,"_uuid":"96b6fae29a35b913c19b117996642d37ad5a74ff"},"cell_type":"code","source":"class TopicalWords: \n     \n    embeddings_index = None\n    vocab2count = None\n    _nlp = None\n    punctuations = ''.join(TextualDataControl.PUNCTUATIONS)\n    \n    MIN_SIMILARITY_SCORE = 0.5\n\n    def __init__(self, topical_word):\n        self.source_word = topical_word\n        self.substitutes = None\n        self.related_sources = None\n                \n        self.my_NER_type = TopicalWords.get_NER_label(self.source_word)\n        self.my_POS_tag = TopicalWords.get_tag(self.source_word.replace(\"_\", \" \"))\n\n    @classmethod\n    def get_NER_label(cls, word):\n        retval = None\n        word = TextualDataControl.depunctuate(word, remove=True)\n        nlp = TopicalWords.nlp()\n        ents = nlp(word).ents\n        if len(ents) > 0:\n            retval = ents[0].label_\n        \n        return retval\n\n    @classmethod\n    def get_tag(cls, phrase):\n    \n        retval = None\n        nlp = TopicalWords.nlp()\n        \n        doc = nlp(u'{}'.format(phrase))\n        last_tag = None\n        for token in doc:\n            last_tag = token.tag_\n            if last_tag in [\"NNS\", \"NN\", \"NNP\"]:\n                retval = last_tag\n                break\n        \n        return retval\n    \n    \n    @classmethod\n    def nlp(cls):\n        \n        if TopicalWords._nlp is None:\n            TopicalWords._nlp = en_core_web_sm.load()\n\n        return TopicalWords._nlp    \n    \n    @classmethod\n    @profile\n    def embeddings(cls):\n        \n        if TopicalWords.embeddings_index is None:\n            glove_file = datapath(gEmbeddingsSources['GLOVE']['path'])\n            tmp_file = get_tmpfile(\"glove_word2vec.txt\")\n            _ = glove2word2vec(glove_file, tmp_file)\n            TopicalWords.embeddings_index = KeyedVectors.load_word2vec_format(tmp_file, binary=True, limit=500000)\n\n        return TopicalWords.embeddings_index\n    \n    @classmethod\n    @profile\n    def word2count(cls):\n        \n        if TopicalWords.vocab2count is None:\n            app = DataManager.instance()\n            app.load(orig=True, test=False)\n            orig_quora = QuoraPreprocessor(training_data=app.training_data[app.training_data.target == 1], test_data=None, is_gen=False)\n            orig_quora.set_mode(is_training=True)\n            orig_quora.make_clean_questions()\n            TopicalWords.vocab2count = orig_quora.make_vocab(lower=True)\n            log(f\"Number of words in original input's vocabulary:{len(TopicalWords.vocab2count)}\")\n            del orig_quora\n            del app\n            gc.collect()\n        \n        return TopicalWords.vocab2count\n    \n    @classmethod\n    def contains(cls, word1, word2):\n        return (word1 in word2 or word2 in word1) and len(word1) != len(word2)   \n    \n    @classmethod\n    def multi_word_in_vocab(cls, phrase):\n        # because GOOGLENEWS has compound words with underscore\n        # Allow underscore in single word, and check if each component is also\n        # in the vocab.\n        lower_phrase = phrase.lower()\n        word2count = TopicalWords.word2count()\n        retval = False\n        if lower_phrase in word2count:\n            retval = True\n        elif '_' in phrase:\n            punct_to_split = \"[a-zA-Z0-9']+|[.,!?;]\"\n            found_multi_words = re.findall(r\"{}\".format(punct_to_split), phrase)\n            if found_multi_words is not None and len(found_multi_words) > 1:\n                found = True\n                for w in found_multi_words:\n                    if  w.lower() not in word2count:\n                        found = False\n                        break\n                if found:\n                    phrase = \" \".join(found_multi_words)\n                #if not found and verbose:\n                #    log_list(found_multi_words, desc=\"Words of MultiWord not found.\")\n                retval = found\n                \n        return retval, phrase\n    \n    @classmethod\n    def is_vocab(cls, phrase, verbose=False):\n        retval, phrase = TopicalWords.multi_word_in_vocab(phrase)\n        if retval == False and verbose:\n            log(phrase + \" is not in vocab\")\n        return retval\n        \n    \n    def find_matching_phrases(self, word_list):\n        tmp = self.source_word.lower()\n        retval = []\n        \n        for (similar, _) in word_list:\n            if TopicalWords.contains(similar.lower(), tmp.lower()) == False:\n                continue\n            \n            found = False\n            for existing_word in retval:\n                if existing_word.lower() == similar.lower():\n                    found = True\n                    break\n\n            if not found:\n                if self.my_POS_tag is None or TopicalWords.get_tag(similar) == self.my_POS_tag:\n                    retval.append(similar)\n                \n        retval.insert(0, self.source_word)\n        \n        return retval\n    \n    def find_substitute_phrases(self, matching_word_list, similars):\n         # Combine all words that vertically belong to the source word (which is the topic)\n        retval = [(w, 1) for w in matching_word_list]\n        \n        matching_lower = set()\n        for w in matching_word_list:\n            matching_lower.add(w.lower())\n        # Topics that are more similar to textually matching words. They are \"vertically\" matching.\n        if len(matching_word_list) > 2:\n            substitutes = TopicalWords.embeddings().most_similar(positive=matching_word_list, topn=30)\n            substitutes = [(w, score) for (w, score) in substitutes if score > TopicalWords.MIN_SIMILARITY_SCORE  and TopicalWords.is_vocab(w) and w.lower() not in matching_lower]\n            retval.extend(substitutes)      \n        else:\n            for (w, score) in similars: \n                if score > 0.6 and TopicalWords.is_vocab(w):\n                    w_NER = TopicalWords.get_NER_label(w)\n                    if ((self.my_NER_type is None) or (self.my_NER_type == w_NER)):\n                        if self.my_POS_tag is None or TopicalWords.get_tag(w) == self.my_POS_tag:\n                            retval.append((w, score))\n            \n        return retval\n    \n    def filter_related_sources(self, similar_phrases):        \n        \n        unique_subs = set([w.lower() for (w, _) in self.substitutes])\n        retval = [(w, score) for (w, score) in similar_phrases if w.lower() not in unique_subs and score > TopicalWords.MIN_SIMILARITY_SCORE and TopicalWords.is_vocab(w) and ((self.my_NER_type is None) or (self.my_NER_type == TopicalWords.get_NER_label(w)))]\n        retval = (sorted(retval, key=lambda t:t[1], reverse=True))\n\n        return retval\n    \n    def filter_substitutes(self):\n        retval = []\n\n        for (w, score) in self.substitutes:\n            if score > TopicalWords.MIN_SIMILARITY_SCORE and TopicalWords.is_vocab(w):\n                NER_type = TopicalWords.get_NER_label(w)\n                if NER_type == self.my_NER_type:\n                    if self.my_POS_tag is None or TopicalWords.get_tag(w) == self.my_POS_tag:\n                        w = w.strip(TopicalWords.punctuations).strip()\n                        if len(w) > 2 and w not in ['ing']:\n                            retval.append((w, score, NER_type))\n\n        retval = (sorted(retval, key=lambda t:t[1], reverse=True))    \n        return retval\n    \n    @profile\n    def process(self, num_top = 10, verbose=True):\n        \n        if verbose:\n            log(\"Processing topical words for <b>{}</b>\".format(self.source_word))\n        \n        # horizontally relevant to topical word\n        other_topical_words = TopicalWords.embeddings().similar_by_word(self.source_word, topn=50)\n        if verbose:\n            log(\"Found other topical words: {}\".format(len(other_topical_words)))\n            \n        # Textually matching words\n        matching = self.find_matching_phrases(other_topical_words)        \n        if verbose and matching is not None and len(matching) > 0:\n            log(\"Found matching words: {}\".format(len(matching)))\n\n        # Words that could be substituted for each other, and the source word.\n        self.substitutes = self.find_substitute_phrases(matching, other_topical_words)\n        if verbose and self.substitutes is not None and len(self.substitutes) > 0:\n            log(\"Found initial set of substitutes: {}\".format(len(self.substitutes)))\n                            \n        # cleanup\n        # related words that are not in word_list are newly discovered topics.\n        other_topical_words = self.filter_related_sources(other_topical_words)\n        \n        self.substitutes = self.filter_substitutes()[:min(num_top, len(self.substitutes))]        \n        self.related_sources = other_topical_words[:min(num_top, len(other_topical_words))]\n    \n    @classmethod\n    def memclean(cls):\n        del TopicalWords.embeddings_index\n        TopicalWords.embeddings_index = None\n        del TopicalWords._nlp\n        TopicalWords._nlp = None\n        del TopicalWords.vocab2count\n        TopicalWords.vocab2count = None\n    \n    def inspect(self):\n        if not gInspect:\n            return\n        \n        log_list(self.substitutes, desc=\"Substitutes to {} with NER type:{}\".format(self.source_word, self.my_NER_type))\n        log_list(self.related_sources, transpose=True, desc=\"Source Words to {} with NER type:{}\".format(self.source_word, self.my_NER_type))\n        log(f\"Number of words in vocabulary:{len(TopicalWords.word2count())}\")\n    \n    @classmethod \n    def testme(cls):\n        sources = ['f_ing','Pakistanis','ass', \"Orthodox_Jews\", \"Hinduism\", 'Muslim', 'Islamic_fundamentalist', 'Muslims', 'fuck', 'Barack_Obama', 'Christian', 'blacks']\n\n        for source in sources:\n            me = cls(source)\n            me.process(num_top=10, verbose=True)\n            me.inspect()\n            if TopicalWords.is_vocab(source, verbose=True) == False:\n                log(f\"This word NOT found in data:{source}\")\n            del me\n        \n        TopicalWords.memclean()\n        gc.collect() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0613101dffa04afd739cde8a15d7624fa8b86f2c","scrolled":true},"cell_type":"code","source":"if False:\n    clear_prof_data()\n    InsincereVocabGenerator.clean_cache(bow=True, vocab=True)\n    TopicalWords.testme()\n    pp_prof_data()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"47b662d89ac04a1b71b0d7d463074995e66cfa3b"},"cell_type":"markdown","source":"## InsincereVocabGenerator\n\nUsing TopicalWords and an initial set of words, this class uses similarity function from Word2Vec to discover similar sets of words that could be substituted into a question with each other.\n* So, this complies with the rule of no external data sources.\n![](http://)* The final result is cached to file for use by the QuestionGenerator."},{"metadata":{"trusted":true,"_uuid":"d89806e61497f83c42deec01ae5f7e1e1c26364e"},"cell_type":"code","source":"# Find related words to source word.\n# Find matching words in related words to source word\n# Add all related words that are not matching words to insincere_vocab category\nclass InsincereVocabGenerator:\n    \n    InsincereBoW = \"InsincereBoW\"\n    InsincereVocab = \"InsincereVocab\"\n    MAX_DERIVED_WORDS_PER_TOPIC = 25\n    INSINCERE_VOCAB = {\n                \"PEOPLES\" : {\"Hindus\" : None, \"Muslims\" : None, \"Christians\" : None, \"Jews\" : None, \"Latinos\" : None, 'women':None, \"gay\" : None, \"blacks\" : None},\n                \"RELIGIONS\" : {\"Hinduism\":None, \"Islam\" : None, \"Christianity\" : None},\n                \"NATIONALITIES\" : {\"Arabs\":None, \"Chinese\":None, \"Americans\":None, \"Europeans\": None, \"Pakistanis\":None, 'Americans' : None},\n                \"PERSONS\" : {\"Donald_Trump\" : None, \"Bill_Clinton\" : None, \"Trump\" : None, \"Barack_Obama\" : None},\n                \"PLACES\"  : {\"India\" : None, \"Pakistan\": None, \"Korea\" : None, \"China\" : None, \"Florida\":None, 'Israel' : None, 'Iran':None},\n                \"SEXUAL\"  : {\"fuck\" : None, \"shit\" : None, \"asshole\" : None, \"ass\":None, \"pussy\" : None, \"penis\": None, \"rape\": None},\n                \"NEGATIVE\" : {\"bitch\" : None,\"terrorists\" : None, \"hate\" : None, \"stupid\":None },\n                \"POLITICAL_GROUPS\" : {\"Democrats\" : None, \"Republicans\" : None, \"liberals\" : None, \"conservatives\" : None}\n    }\n    SINCERE_VOCAB = [\"latina\",\"c'mon\",\"outta\",\"fanny\",\"love\", \"sex\" 'ear_lobes', 'ear_lobe', 'inner_thigh','thighs', 'lobe', 'gland','African', 'funny','woman','religions','undocumented_aliens','Hindu_statesman_Rajan_Zed','religious', 'dude', 'cuz','wanna','shiz','hey','lol','pee','short_sighted']\n    \n    def __init__(self):\n    \n        self.already_cached = False\n        if os.path.isfile(InsincereVocabGenerator.InsincereVocab) and os.path.isfile(InsincereVocabGenerator.InsincereBoW):\n            self.insincere_vocab = pickle.load(open(InsincereVocabGenerator.InsincereVocab, \"rb\"))\n            self.insincere_bow = pickle.load(open(InsincereVocabGenerator.InsincereBoW, \"rb\")) \n            self.already_cached = True\n        else:\n            # seed topical words by category\n            self.insincere_vocab = InsincereVocabGenerator.INSINCERE_VOCAB\n            self.insincere_bow = {}\n        \n    # Insincere_bow is an dictionary of categories. \n    # Each category is mapped to an array of dictionaries of words.\n    @profile\n    def generate_bow(self):\n        for category, category_topics in self.insincere_vocab.items():\n            self.insincere_bow[category] = []\n            for source, topical_words in category_topics.items():\n                if source in InsincereVocabGenerator.SINCERE_VOCAB:\n                    continue\n                \n                (first_word, score, _) = topical_words[0]\n                if first_word.lower() == source.lower() and score < 1.0:\n                    continue\n                    \n                word_dict = {}\n                for (phrase, score, _) in topical_words:                    \n                    _, phrase = TopicalWords.multi_word_in_vocab(phrase)\n                    word_dict[phrase] = score\n\n                if len(word_dict) > 1:\n                    self.insincere_bow[category].append(word_dict)\n        \n        if len(self.insincere_bow.items()) > 0:\n            save_binary(self.insincere_bow, InsincereVocabGenerator.InsincereBoW)\n    \n    def get_topical_words(self, source, verbose):\n\n        retval = ([], [])\n        try:\n            if verbose:\n                log(\"Processing source word: {}\".format(source))\n            derived_words = TopicalWords(source)\n            derived_words.process(num_top=InsincereVocabGenerator.MAX_DERIVED_WORDS_PER_TOPIC, verbose=verbose)\n            if verbose:\n                log_list(derived_words.substitutes, desc=\"Strongly Matched to {}\".format(source))\n                log_list(derived_words.related_sources, transpose=True, desc=\"Newly discovered based on {}\".format(source))\n            retval = (derived_words.substitutes, derived_words.related_sources)\n            del derived_words\n        except:\n            if verbose:\n                log(\"Source word {} not found\".format(source))\n        \n        return retval\n    \n    # XXX culling requires more attention.\n    def clean(self):\n\n        for category, category_topics in self.insincere_vocab.items():\n            to_delete = []\n            for source, topical_words in category_topics.items():\n                if topical_words is None or len(topical_words) < 3:\n                    to_delete.append(source)\n            for source in to_delete:\n                del category_topics[source] \n            \n    \n    @profile\n    def generate_similar_words(self, verbose=False):\n        \n        # Twice for newly discovered topical words\n        twice = 0\n        while twice < 2: \n            # Iterate categories\n            # \"PEOPLES\":{ \"Hindu\":None ... }\n            for category, category_topics in self.insincere_vocab.items():            \n                if verbose:\n                    log(\"<h1>{}</h1>\".format(category))\n                discovered_source_words = []\n                # Iterate topical words in each category\n                # \"Hindu\" : \"Sikhs\", ...\n                for source, topical_words in category_topics.items():\n                    if TopicalWords.is_vocab(source) == False:\n                        continue\n                        \n                    if topical_words is None:\n                        (substitutes, more_source_words) = self.get_topical_words(source, verbose)\n                        if len(substitutes) > 0:\n                            category_topics[source] = substitutes\n                        if len(more_source_words) > 0:\n                            discovered_source_words.extend(more_source_words)\n                \n                # Add newly discovered words to topical words for this category\n                if twice == 0:                    \n                    for (w, score) in discovered_source_words:\n                        found = False\n                        for _, category_topics_2 in self.insincere_vocab.items():\n                            if w in category_topics_2 or w.lower() in category_topics_2 or w in InsincereVocabGenerator.SINCERE_VOCAB:\n                                found = True\n                                break\n                        if not found:\n                            category_topics[w] = None\n                            \n            twice = twice + 1\n        \n        self.clean()\n        if len(self.insincere_vocab.items()) > 0:\n            save_binary(self.insincere_vocab, InsincereVocabGenerator.InsincereVocab)\n                            \n    def mislabeled_sincerity(self, selected_categories=['SEXUAL', 'NEGATIVE']):\n        \n        app = DataManager.instance()\n        app.load(orig=True)\n        raw_data = app.training_data\n        self.sincere = raw_data.index[raw_data.target == 0]        \n            \n        for i in random.sample(range(0, len(self.sincere)), 10000):\n            q_text = raw_data.iloc[ self.sincere[i],:]['question_text']\n            for category, category_lists in self.insincere_bow.items():\n                if category not in selected_categories:\n                    continue\n                    \n                # Find multiple word lists with the largest phrase \n                for words_list in category_lists:\n                    largest_phrase = None\n                    size = 0\n                    words_list = list(words_list.keys())\n\n                    for phrase in words_list:\n                        if TextualDataControl.find_string(phrase, q_text):\n                            if len(phrase) > size:\n                                size = len(phrase)\n                                largest_phrase = phrase\n\n                    if largest_phrase is not None:\n                        highlighted_q_text = re.sub(r\"\\b\" + re.escape(largest_phrase) + r\"\\b\", \"<b>{} ({})</b>\".format(largest_phrase, \"??\"), q_text, flags=re.IGNORECASE)\n                        log(highlighted_q_text)\n    @profile\n    def process(self, verbose=False):\n        if self.already_cached:\n            if verbose:\n                log(\"Similar words are already cached.\")\n        else:\n            self.generate_similar_words(verbose=verbose)\n            self.generate_bow()\n                \n    @classmethod\n    def clean_cache(cls, vocab=False, bow=False):        \n        try:\n            if vocab:\n                log(\"Deleting insincere vocabulary cache files...\")\n                remove_file(InsincereVocabGenerator.InsincereBoW)\n                remove_file(InsincereVocabGenerator.InsincereVocab)\n\n            if bow:\n                log(\"Deleting insincere vocabulaty BoW cache file...\")\n                remove_file(InsincereVocabGenerator.InsincereBoW)\n        except:\n            log(\"Some error during cleaning insincere vocab.\")\n\n    \n    def inspect(self, bow=True, vocab=False):        \n        if bow:\n            log(\"<h1>Insincere, Categorized Bag of Words</h1>\")\n            for category, category_dicts in self.insincere_bow.items():\n                word_lists = []\n                for a_dict in category_dicts:\n                    substitutes = []\n                    for word, score in a_dict.items():\n                        substitutes.append(\"{} - {}\".format(word, score))\n                    substitutes[0] = \"<b>{}</b>\".format(substitutes[0])\n                    word_lists.append(', '.join(substitutes))\n                log_list(word_lists, desc=\"<b>{}</b>\".format(category))\n        \n        if vocab:\n            log(\"<h1>Insincere Vocabulary - Basis of BoW</h1>\")\n            for category, category_topics in self.insincere_vocab.items():\n                log(\"<h2>{}</h2>\".format(category))\n                for source, topical_words in category_topics.items():\n                    log_list(topical_words, desc=\"Words within topic {}\".format(source))\n    \n    @classmethod\n    def testme(cls):\n        me = cls() \n        me.process(verbose=False)\n        me.inspect(vocab=False)\n        del me\n        gc.collect()\n        \n    @classmethod\n    def testme2(cls):\n        oov = []\n        for category, category_topics in InsincereVocabGenerator.INSINCERE_VOCAB.items():\n            for source, _ in category_topics.items():\n                if TopicalWords.is_vocab(source, verbose=True) == False:\n                    oov.append([category, source])\n        \n        log_list(oov, desc=\"Category Topics not in Vocab\")\n        \n        words = ['Barack', 'Obama', 'Donald', 'Trump','Bill','Clinton', \"Hinduism\"]\n        tmp = []\n        for w in words:\n            tmp.append([w, TopicalWords.is_vocab(w, verbose=True)])\n        \n        log_list(tmp, desc=\"Words in Vocab ??\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"623ffe6a6509baaac30e1a4e5aa36060ad24a1f7","scrolled":false},"cell_type":"code","source":"if False:\n    clear_prof_data()\n    #InsincereVocabGenerator.testme2()\n    InsincereVocabGenerator.testme()\n    pp_prof_data()\n    log_dir(\"../working\")\n\nif False:\n    log_dir(\"../working\")\n    InsincereVocabGenerator.clean_cache(vocab=True, bow=True)\n    log_dir(\"../working\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f708551b2d12ed3768f4a90635b6f7bce5210d9"},"cell_type":"markdown","source":"## InsincerityWithNER\n\n* This is used to eyeball the highest frequency, named entities that are involved in insincerity.\n* It is used for manually creating the seed for InsincerityVocabGenerator."},{"metadata":{"trusted":true,"_uuid":"15b4859875bc58eb4902c2382020a48c73169251"},"cell_type":"code","source":"class InsincerityWithNER:\n    \n    CACHED_FILE_PREFIX = \"NER\"\n    OUTPUT_FILE = \"NER\"\n    \n    def __init__(self):\n        self.q_NER = None        \n    \n    def gen_NER(self, cpu_id, cpu_count, cachefile, insincere=True):     \n        if os.path.isfile(cachefile):\n            return\n\n        app = DataManager.instance()\n        app.load(orig=True)\n        raw_data = app.training_data\n        if insincere:\n            raw_data = raw_data[raw_data.target == 1]\n\n        q_range = MPHelper.range_partition(cpu_id=cpu_id, total_count=len(raw_data))\n        NER_dict = Counter()\n        for i in q_range:\n            q_text = raw_data.iloc[i,:]['question_text']\n\n            ents = TopicalWords.nlp()(q_text).ents\n            if len(ents) == 0:\n                continue\n            else:\n                for X in ents:\n                    if X.label_ in ['GPE', 'NORP', 'PERSON', 'ORG']:\n                        NER_dict[X.text + \"/\" + X.label_] += 1    \n\n        save_binary(NER_dict, cachefile)\n    \n    @profile\n    def process(self):\n        app = DataManager.instance()\n        app.load(orig=True)\n        raw_data = app.training_data\n        my_insincere_q = raw_data.index[raw_data.target == 1]\n        \n        log(\"Generating NER for {} insincere questions.\".format(len(my_insincere_q)))\n\n        mp = MPHelper(InsincerityWithNER.CACHED_FILE_PREFIX, verbose=True)\n        mp.map_process(self.gen_NER)\n        results = mp.reduce()\n        merged = Counter()\n        for NER_dict in results:\n            merged += NER_dict\n        \n        merged = merged.most_common(len(merged.items()))\n        self.q_NER = merged\n        save_binary(self.q_NER, InsincerityWithNER.OUTPUT_FILE)\n        del mp\n        gc.collect()\n\n    def inspect(self):\n        if not gInspect:\n            return\n        \n        log_list(self.q_NER[:200], desc=\"Most frequent 200 NER in 20K questions\")        \n        \n\n    @classmethod\n    def testme(cls):\n        me = cls()\n        clear_prof_data()\n        me.process()\n        pp_prof_data()\n        me.inspect()\n        del me\n        \n    @classmethod\n    def display_cache(cls):\n        me = cls()\n        me.q_NER = pickle.load(open(InsincerityWithNER.OUTPUT_FILE, \"rb\"))\n        me.inspect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"643ec0a37c17025933c5fb474b27ed166feb9761"},"cell_type":"code","source":"if False:\n    InsincerityWithNER.testme()\n    \nif False:\n    InsincerityWithNER.display_cache()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"29fd19eb24245ea514d2032efdf52dca902a062b"},"cell_type":"markdown","source":"## QuestionsGenerator\n\nFinally, the multi-processor generation of questions using phrase substitutions from InsincereBoW above"},{"metadata":{"trusted":true,"_uuid":"d2342455e268ea5ec2dd4f7c1a34e2934938ef2d"},"cell_type":"code","source":"class QuestionsGenerator:\n    \n    CACHED_FILE_PREFIX = \"gen_q\"\n    NOT_GEN_FILE_PREFIX = \"no_gen_q\"\n    \n    def __init__(self):\n        self.gen_q = {}\n        self.non_gen_q = []\n        \n        if os.path.isfile(QuestionsGenerator.CACHED_FILE_PREFIX):\n            self.gen_q = pickle.load(open(QuestionsGenerator.CACHED_FILE_PREFIX, 'rb'))\n        if os.path.isfile(QuestionsGenerator.NOT_GEN_FILE_PREFIX):\n            self.non_gen_q = pickle.load(open(QuestionsGenerator.NOT_GEN_FILE_PREFIX, 'rb'))\n            \n        self.rawdata = None\n        self.word_substitutes = None\n        self.substitutes_set = None\n    \n    @classmethod\n    def find_substitute_candidates(cls, q_text, substitutions, limit=10):\n        retval = {}\n        \n        # Search for substitutes by matching the longest phrase in the question\n        # within each category.\n        for category, category_dicts in substitutions.insincere_bow.items():            \n            for a_dict in category_dicts:\n                largest_phrase = None\n                size = 0\n                for phrase, _ in a_dict.items():\n                    if TextualDataControl.find_string(phrase, q_text):\n                        if len(phrase) > size:\n                            size = len(phrase)\n                            largest_phrase = phrase\n                            \n                if largest_phrase is not None:\n                    if largest_phrase not in retval:\n                        retval[largest_phrase] = []\n                    retval[largest_phrase].append(a_dict)\n            \n            '''\n            substitutes_count = 1\n            for _, subs_dict_list in retval.items():\n                maxlen = max([len(d) for d in subs_dict_list])\n                substitutes_count *= maxlen\n            if substitutes_count > limit:\n                break\n            '''\n\n        # For each matching phrase, there could be multiple lists of substitutions.\n        return retval\n\n    @classmethod\n    def pos_phrase(cls, phrase, q_text):\n        tmp = q_text.split()\n        pos = nltk.pos_tag(tmp)\n        \n        retval = set()\n        retval.add(nltk.pos_tag([phrase])[0][1])\n        \n        for i in range(len(tmp)): \n            if re.search(r\"\\b\" + re.escape(phrase) + r\"\\b\", tmp[i]):\n                retval.add(pos[i][1])\n                break\n        \n        return retval\n        \n    @classmethod\n    def find_substitutes_list(cls, matching_phrase, substitutes_list):\n         # Find the right list of substitutes\n        retval = substitutes_list[0]\n        \n        if len(substitutes_list) > 1:\n            highest_score = 0\n            for a_dict in substitutes_list:\n                if (highest_score < a_dict[matching_phrase]) or ((highest_score == a_dict[matching_phrase]) and (len(a_dict) > len(retval))):\n                    retval = a_dict\n                    highest_score = a_dict[matching_phrase]\n\n        retval = list(retval.keys())\n        \n        return retval\n    \n    @classmethod\n    def do_substitutes(cls, q_text, phrase, all_substitutes, verbose=False):\n        retval = {}\n        \n        num_generated = 0\n        pos = QuestionsGenerator.pos_phrase(phrase, q_text)\n        q_tokens = QuoraPreprocessor.make_tokens(q_text)\n        q_tokens = set([t.lower() for t in q_tokens])\n        \n        for substitute in all_substitutes:\n            subs_pos = nltk.pos_tag([substitute])[0][1]\n            if substitute.lower() != phrase.lower() and subs_pos in pos and substitute.lower() not in q_tokens:\n                highlighted_q_text = re.sub(r\"\\b\" + re.escape(phrase) + r\"\\b\", \"__{}__\".format(phrase, pos), q_text, flags=re.IGNORECASE)\n                    \n                if highlighted_q_text not in retval:\n                    retval[highlighted_q_text] = set()\n                if len(retval[highlighted_q_text]) < 10:\n                    num_generated += 1\n                    \n                    if verbose:\n                        gen_q = re.sub(r\"\\b\" + re.escape(phrase) + r\"\\b\", \"<b>{} ({})</b>\".format(substitute, subs_pos), q_text, flags=re.IGNORECASE)\n                    else:\n                        gen_q = re.sub(r\"\\b\" + re.escape(phrase) + r\"\\b\", substitute, q_text, flags=re.IGNORECASE)                            \n                    \n                    gen_q = TextualDataControl.sent_capitalize(gen_q)\n                    retval[highlighted_q_text].add(gen_q)\n                    \n        return retval, num_generated\n    \n    @classmethod\n    @profile\n    def gen_from_one_q(cls, q_text, substitutions, limit=20, verbose=False):\n        \n        results = {}\n        total_generated = 0\n        \n        # For each substitutable phrase\n        for largest_phrase, subs_dict_list in substitutions.items():\n            if verbose:\n                log(\"Number of substitutable lists for phrase <b>{}</b> found:<b>{}</b>\".format(largest_phrase, len(subs_dict_list)))\n\n            substitutes = QuestionsGenerator.find_substitutes_list(largest_phrase, subs_dict_list)\n            new_q_dict, num_generated = QuestionsGenerator.do_substitutes(q_text, largest_phrase, substitutes, verbose=verbose)\n            total_generated += num_generated\n            results.update(new_q_dict)\n        \n        if verbose:\n            log_list(results, desc=\"New Questions Prior to Mixing\")\n        \n        mix_gen_set = set()\n        [mix_gen_set.update(q) for _, q in results.items()]\n        more_q = {}\n        for largest_phrase, subs_dict_list in substitutions.items():\n            for orig_q_text, new_q_list in results.items():\n                if not re.search(re.escape(\"__\" + largest_phrase + \"__\"), orig_q_text, flags=re.IGNORECASE):\n                    for new_q in new_q_list:\n                        if total_generated > limit:\n                            break\n                        substitutes = QuestionsGenerator.find_substitutes_list(largest_phrase, subs_dict_list)\n                        new_q_dict, num_generated = QuestionsGenerator.do_substitutes(new_q, largest_phrase, substitutes)\n                        found = False\n                        for _, q in new_q_dict.items():\n                            if q in mix_gen_set:\n                                found = True\n                                break\n                        if not found:\n                            more_q.update(new_q_dict)\n                            [mix_gen_set.update(q) for _, q in new_q_dict.items()]\n                            total_generated += len(new_q_dict.values())\n\n        results = {q_text : list(mix_gen_set)[-max(limit, len(mix_gen_set)):]}\n                                        \n        if verbose: \n            if total_generated == 0:\n                results[\"REPORT: \" + q_text] = set([\"<b><big>NO</big></b> substitutable word found.\"])\n            else:\n                results[\"REPORT: \" + q_text] = set([\"Number of questions generated: <b>{}</b><br>\".format(total_generated)])\n                \n        return results\n    \n    @profile\n    def gen_q_proc(self, cpu_id, cpu_count, cache_file):\n        \n        if os.path.isfile(cache_file):\n            return        \n        \n        q_range = MPHelper.range_partition(cpu_id=cpu_id, total_count=len(self.rawdata))\n        # Output        \n        results = {}    \n        max_len = q_range.stop - q_range.start\n        total_gen = 0\n        for i in q_range:\n            q_text = self.rawdata.iloc[i,:]['question_text']\n            if self.is_substitutable(q_text):\n                # Find multiple word lists with the largest phrase\n                substitutions = QuestionsGenerator.find_substitute_candidates(q_text, self.word_substitutes)            \n                new_q_dict = QuestionsGenerator.gen_from_one_q(q_text, substitutions, verbose=False, limit=10)\n                num_new_q = len(list(new_q_dict.values())[0])\n                total_gen += num_new_q\n                results.update(new_q_dict)\n                if total_gen > max_len:\n                    break\n\n        for q, a_set in results.items():\n            results[q] = list(a_set)\n        \n        self.save([results], cache_file)\n        \n    @profile\n    def nogen_q_proc(self, cpu_id, cpu_count, cache_file):\n        if os.path.isfile(cache_file):\n            return        \n        \n        q_range = MPHelper.range_partition(cpu_id=cpu_id, total_count=len(self.rawdata))\n        results = []\n        for i in q_range:\n            q_text = self.rawdata.iloc[i,:]['question_text']\n            if self.is_substitutable(q_text) == False:\n                results.append(q_text)\n                    \n        save_binary(results, cache_file)\n\n    def save(self, results, cache_file):\n        if gWIP:\n            save_binary(results[0], cache_file)\n        \n        # Create generated questions CSV that matches the train.csv format.\n        # XXX what assumptions can be made about train.csv format.\n        cache_file_csv = cache_file + \".csv\"\n        headers = ['qid', 'question_text', 'target']\n            \n        dest_f = open(cache_file_csv, 'w')\n        csv_writer = csv.writer(dest_f, delimiter=',', quotechar='\"', quoting=csv.QUOTE_MINIMAL)\n        csv_writer.writerow(headers)\n        for cpu_result in results:\n            for orig_q, gen_q_list in cpu_result.items():\n                for gen_q in gen_q_list:\n                    qid = ''.join(random.choice(string.hexdigits) for x in range(20)).lower()\n                    row = [qid, gen_q, str(1)]\n                    csv_writer.writerow(row)\n            \n        dest_f.close()\n        \n    def make_substitutes_bow(self):\n        self.substitutes_set = set()\n        for category, category_dicts in self.word_substitutes.insincere_bow.items():            \n            for a_dict in category_dicts:\n                for k, _ in a_dict.items():\n                    self.substitutes_set.add(k.lower())\n        \n        log(f\"Number of substitutes: {len(self.substitutes_set)}\")\n    \n    def is_substitutable(self, q_text):\n        tokens = QuoraPreprocessor.make_tokens(q_text)\n        tokens = set([t.lower() for t in tokens])\n        \n        retval = len(tokens & self.substitutes_set) > 0  \n        \n        return retval\n    \n    @profile\n    def process(self, debug_df=None):\n        \n        if len(self.gen_q) > 0:\n            return\n        \n        # RawData\n        if debug_df is None:\n            app = DataManager.instance()\n            app.load(orig=True)\n            self.rawdata = app.training_data[app.training_data.target == 1]        \n        else:\n            self.rawdata = debug_df\n        \n        # word substitutes\n        insincere_vocab = InsincereVocabGenerator()\n        insincere_vocab.process()\n        self.word_substitutes = insincere_vocab\n        self.make_substitutes_bow()\n        \n        if gWIP:\n            mp = MPHelper(QuestionsGenerator.NOT_GEN_FILE_PREFIX, verbose=True)\n            mp.map_process(self.nogen_q_proc)\n            results = mp.reduce(clean=True)\n            self.non_gen_q = [q for result in results for q in result]\n        \n        # Multi-process q generation\n        mp = MPHelper(QuestionsGenerator.CACHED_FILE_PREFIX, verbose=False)\n        mp.map_process(self.gen_q_proc)\n        self.gen_q = mp.reduce(clean=True)\n        \n                    \n    def inspect(self, limit=1000):\n        # To help improve generation algorithm.\n        q_count = self.rawdata.count()['qid']\n\n        gen_q_count = 0\n        for gen_q_i in self.gen_q:\n            for q_text, new_q in gen_q_i.items():\n                gen_q_count += len(new_q)\n\n        cpu_index = 0\n        gen_text = set()\n        for gen_q_i in self.gen_q:\n            log(\"<h2>Random question sets From CPU {}</h2>\".format(cpu_index))\n            q_range = MPHelper.range_partition(cpu_id=cpu_index, total_count=len(self.rawdata))\n\n            cpu_index += 1\n            num_iter = min(int(limit/mpc.cpu_count()), len(gen_q_i))\n            keys = list(gen_q_i.keys())\n            samples = np.random.permutation(keys)\n            for i in range(num_iter):\n                q_text = samples[i]\n                q_list = gen_q_i[q_text]\n                gen_text.update(q_list)\n                log_list(q_list, desc=q_text)\n\n        log(\"<h1>Number of new questions generated: {} from total {} insincere questions.\".format(gen_q_count, q_count))\n\n        log(\"<h2>Number of original, insincere questions <b>NOT</b> used for generating new questions: {}</h2>\".format(len(self.non_gen_q)))\n        log_list(self.non_gen_q[0:min(50, len(self.non_gen_q))], desc=\"50 insincere questions with no substitutable words\")\n\n        log(\"<h3>Prior to generated data</h3>\")\n        self.rawdata.at[0,'target'] = 0\n        \n        tdc = TextualDataControl(self.rawdata)\n        tdc.inspect()\n        del tdc\n        log(\"<h3>POST data generation</h3>\")\n        generated_data = pd.DataFrame()\n        log(f\"For limited display purposes, adding {len(gen_text)} new questionst to training data.\")\n        generated_data = generated_data.assign(qid=np.arange(len(gen_text)), question_text=gen_text, target=np.ones(len(gen_text), dtype=np.int8))\n        self.rawdata = self.rawdata.append(generated_data, ignore_index=False)\n        log(f\"<h4>Total training data size: {len(self.rawdata)}</h4>\")\n        tdc = TextualDataControl(self.rawdata)\n        tdc.inspect()\n        del tdc\n        gc.collect()\n        \n    @classmethod\n    def clean_cache(cls, remove_output=False):\n        try:\n            if remove_output:\n                remove_file(QuestionsGenerator.CACHED_FILE_PREFIX)\n                remove_file(QuestionsGenerator.NOT_GEN_FILE_PREFIX)\n                remove_file(QuestionsGenerator.CACHED_FILE_PREFIX + \".csv\")\n            \n            for index in range(mpc.cpu_count()):\n                remove_file(QuestionsGenerator.CACHED_FILE_PREFIX + \"_{}\".format(index))\n            for index in range(mpc.cpu_count()):\n                remove_file(QuestionsGenerator.CACHED_FILE_PREFIX + \"_{}.csv\".format(index))\n                \n        except:\n            log(\"Got an exception while cleaning gen_q cache.\")\n    \n    @classmethod\n    def testme(cls):\n       \n        use_test_data = False\n        debug_df = None\n        if use_test_data:\n            q_texts = [\"I would like to punish my son by legally changing his name to Rapist. How do I do this?\", \"This is not an insincere question.\",\"Is there anything wrong with ageism?\", \"What percentage of Americans are stupid, and why?\", \"Do moms have sex with their sons?\", \"Why women in Tinder are shit?\",\"If blacks support school choice and mandatory sentencing for criminals why don't they vote Republican?\", \"Why do Christians in India hate Brahmins so much?\", \"Why do Europeans say they're the superior race, when in fact it took them over 2,000 years until mid 19th century to surpass China's largest economy?\"]\n\n            debug_df = pd.DataFrame()\n            debug_df = debug_df.assign(qid=np.arange(len(q_texts)), question_text=q_texts, target=np.ones(len(q_texts), dtype=np.int8))\n            debug_df.at[1,'target'] = 0\n            debug_df.head()\n            display(debug_df.head())\n        \n        me = cls()\n        me.process(debug_df=debug_df)\n        me.inspect(limit=20)\n        QuestionsGenerator.clean_cache()\n        del me\n        gc.collect()\n        \n        if False:\n            insincere_vocab = InsincereVocabGenerator()\n            insincere_vocab.process()\n            for q_text in q_texts:\n                log(\"Generating new q from <h1>{}</h1>\".format(q_text))\n                gen_q = {}\n                results = QuestionsGenerator.find_substitute_candidates(q_text=q_text, substitutions=insincere_vocab)\n                log_list(results)\n                results = QuestionsGenerator.gen_from_one_q(q_text, results, verbose=False)\n                gen_q.update(results)\n                log_list(gen_q)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"402fd491ed000f913139879adc8732651d9c2a6f","scrolled":false},"cell_type":"code","source":"if False:\n    clear_prof_data()\n    #TopicalWords.memclean()\n    #InsincereVocabGenerator.clean_cache(vocab=True, bow=True)\n    QuestionsGenerator.clean_cache(remove_output=True)\n    QuestionsGenerator.testme()\n    log_dir('../working')\n    pp_prof_data()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7004e04bdaa3363b2f08b154ac05ed711d0e45f7"},"cell_type":"markdown","source":"# Section: Machine Learning\n## Attention Layer\n\nCopy-pasted from  https://www.kaggle.com/nikhilroxtomar/gru-with-kfold-lb-0-689"},{"metadata":{"trusted":true,"_uuid":"1e2597a52435cef23675e2e9efa5753886d2fcd6"},"cell_type":"code","source":"class AttentionLayer(Layer):\n    def __init__(self, step_dim=QuoraPreprocessor.MAX_WORDS_IN_QUESTION,\n                 W_regularizer=None, b_regularizer=None,\n                 W_constraint=None, b_constraint=None,\n                 bias=True, **kwargs):\n        self.supports_masking = True\n        self.init = initializers.get('glorot_uniform')\n\n        self.W_regularizer = regularizers.get(W_regularizer)\n        self.b_regularizer = regularizers.get(b_regularizer)\n\n        self.W_constraint = constraints.get(W_constraint)\n        self.b_constraint = constraints.get(b_constraint)\n\n        self.bias = bias\n        self.step_dim = step_dim\n        self.features_dim = 0\n        super(AttentionLayer, self).__init__(**kwargs)\n\n    def build(self, input_shape):\n        assert len(input_shape) == 3\n\n        self.W = self.add_weight('{}_W'.format(self.name), shape=(input_shape[-1].value,),\n                                 initializer=self.init,\n                                 regularizer=self.W_regularizer,\n                                 constraint=self.W_constraint)\n        self.features_dim = input_shape[-1]\n\n        if self.bias:\n            self.b = self.add_weight('{}_b'.format(self.name), shape=(input_shape[1].value,),\n                                     initializer='zero',\n                                     regularizer=self.b_regularizer,\n                                     constraint=self.b_constraint\n                                    )\n        else:\n            self.b = None\n\n        self.built = True\n\n    def compute_mask(self, input, input_mask=None):\n        return None\n\n    def call(self, x, mask=None):\n        features_dim = self.features_dim\n        step_dim = self.step_dim\n\n        eij = K.reshape(K.dot(K.reshape(x, (-1, features_dim)),\n                        K.reshape(self.W, (features_dim, 1))), (-1, step_dim))\n\n        if self.bias:\n            eij += self.b\n\n        eij = K.tanh(eij)\n\n        a = K.exp(eij)\n\n        if mask is not None:\n            a *= K.cast(mask, K.floatx())\n\n        a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())\n\n        a = K.expand_dims(a)\n        weighted_input = x * a\n        return K.sum(weighted_input, axis=1)\n\n    def compute_output_shape(self, input_shape):\n        return input_shape[0],  self.features_dim","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"026ed551ebcf303501c57aed2b7ddd3116ef655d"},"cell_type":"markdown","source":"## Model Factory\n\nTo try out different models, since inputs and outputs from the model are the same.\nOne model factory per word embedding.  Can explore merging word embeddings later."},{"metadata":{"trusted":true,"_uuid":"e161608cb367344ee205b38d7204ac585b68fa6b"},"cell_type":"code","source":"class ModelFactory:\n    def __init__(self, embeddings_matrix):\n        self.embeddings_matrix = embeddings_matrix\n        self.max_features = self.embeddings_matrix.shape[0]\n        self.embed_size = self.embeddings_matrix.shape[1]\n        self.maxlen = QuoraPreprocessor.MAX_WORDS_IN_QUESTION\n        \n        self.mask_zero = False\n        self.models = {}\n    \n    @classmethod\n    def multi_gpu_model(cls, model):        \n        retval = model\n        return retval\n        \n    def create_GPU_LSTM_Model(self, units=64):\n        \n        inp = Input(shape=(self.maxlen,))\n        x = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inp)\n        x = Bidirectional(CuDNNLSTM(units*2, return_sequences=True))(x)\n        x = Bidirectional(CuDNNLSTM(units, return_sequences=True))(x)\n        x = AttentionLayer(self.maxlen)(x)\n        x = Dense(units, activation='relu')(x)\n        outp= Dense(1, activation='sigmoid')(x)\n        model = ModelFactory.multi_gpu_model(Model(inputs=inp, outputs=outp))\n        model = ModelFactory.compilation(model)\n        \n        return model\n    \n    def create_LSTMGRU_Model(self, units=64):\n        \n        inp = Input(shape=(self.maxlen,))\n        x = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inp)\n        x = SpatialDropout1D(0.1)(x)\n        x = Bidirectional(CuDNNLSTM(units, return_sequences=True))(x)\n        y = Bidirectional(CuDNNGRU(units, return_sequences=True))(x)\n\n        atten_1 = AttentionLayer(self.maxlen)(x)\n        atten_2 = AttentionLayer(self.maxlen)(y)\n        avg_pool = GlobalAveragePooling1D()(y)\n        max_pool = GlobalMaxPool1D()(y)\n        \n        conc = concatenate([atten_1, atten_2, avg_pool, max_pool])\n        conc = Dense(int(units/4), activation='relu')(conc)\n        conc = Dropout(0.1)(conc)\n        outp= Dense(1, activation='sigmoid')(conc)\n        \n        model = ModelFactory.multi_gpu_model(Model(inputs=inp, outputs=outp))\n        model = ModelFactory.compilation(model)\n        \n        return model\n    \n    \n    def create_CPU_LSTM_Model(self, units=64):\n        \n        inp = Input(shape=(self.maxlen,))\n        x = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inp)\n        x = Bidirectional(LSTM(units*2, return_sequences=True))(x)\n        x = Bidirectional(LSTM(units, return_sequences=True))(x)\n        x = AttentionLayer(self.maxlen)(x)\n        x = Dense(units, activation='relu')(x)\n        outp= Dense(1, activation='sigmoid')(x)\n        model = Model(inputs=inp, outputs=outp)\n        model = ModelFactory.compilation(model)\n        \n        return model\n    \n    def create_GPU_GRU3_Model(self, units=64):\n        \n        inp = Input(shape=(self.maxlen,))\n        x = Embedding(self.max_features,  self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inp)\n        x = Bidirectional(CuDNNGRU(units*2, return_sequences=True))(x)\n        x = Bidirectional(CuDNNGRU(int(1.5*units), return_sequences=True))(x)\n        x = Bidirectional(CuDNNGRU(units, return_sequences=True))(x)\n        x = AttentionLayer(self.maxlen)(x)\n        outp = Dense(1, activation=\"sigmoid\")(x)\n\n        model = Model(inputs=inp, outputs=outp)\n        model = ModelFactory.compilation(model)\n    \n        return model\n\n    def create_2DCNN_Model(self, num_filters=64, filter_size=3):\n        \n        inp = Input(shape=(self.maxlen, ))\n        x = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inp)\n        #    x = SpatialDropout1D(0.4)(x)\n        x = Reshape((self.maxlen, self.embed_size, 1))(x)\n\n        x = Conv2D(num_filters, kernel_size=(filter_size, filter_size), \n                   kernel_initializer='he_normal', activation='elu')(x)\n        x = MaxPool2D(pool_size=(2,2))(x)\n        # BatchNorm XXX\n        x = Conv2D(num_filters, kernel_size=(filter_size, filter_size),\n                   kernel_initializer='he_normal', activation='elu')(x)\n        x = MaxPool2D(pool_size=(2,2))(x)\n        \n        z = Flatten()(x)\n        z = BatchNormalization()(z)\n\n        outp = Dense(1, activation=\"sigmoid\")(z)\n\n        model = ModelFactory.multi_gpu_model(Model(inputs=inp, outputs=outp))\n        model = ModelFactory.compilation(model)\n\n        return model\n    \n    def create_2DCNN_ConcatModel(self, num_filters=36, filter_sizes=[1,2,3,5]):\n        \n        inp = Input(shape=(self.maxlen,))\n        x = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inp)\n        x = Reshape((self.maxlen, self.embed_size, 1))(x)\n\n        maxpool_pool = []\n        for i in range(len(filter_sizes)):\n            conv = Conv2D(num_filters, kernel_size=(filter_sizes[i], self.embed_size),\n                                         kernel_initializer='he_normal', activation='elu')(x)\n            maxpool_pool.append(MaxPool2D(pool_size=(self.maxlen - filter_sizes[i] + 1, 1))(conv))\n\n        z = Concatenate(axis=1)(maxpool_pool)   \n        z = Flatten()(z)\n        z = Dropout(0.1)(z)\n\n        outp = Dense(1, activation=\"sigmoid\")(z)\n\n        model = ModelFactory.multi_gpu_model(Model(inputs=inp, outputs=outp))\n        model = ModelFactory.compilation(model)\n\n        return model\n    \n    def create_1DCNNNgram_Model(self, num_filters=32, kernel_sizes=[2,4,8], units=None):\n        \n        # channel 1\n        inputs1 = Input(shape=(self.maxlen,))\n        embedding1 = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inputs1)\n        conv1 = Conv1D(filters=num_filters, kernel_size=kernel_sizes[0], activation='relu')(embedding1)\n        drop1 = Dropout(0.5)(conv1)\n        pool1 = MaxPooling1D(pool_size=2)(drop1)\n        flat1 = Flatten()(pool1)\n        # channel 2\n        inputs2 = Input(shape=(self.maxlen,))\n        embedding2 = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inputs2)\n        conv2 = Conv1D(filters=num_filters, kernel_size=kernel_sizes[1], activation='relu')(embedding2)\n        drop2 = Dropout(0.5)(conv2)\n        pool2 = MaxPooling1D(pool_size=2)(drop2)\n        flat2 = Flatten()(pool2)\n        # channel 3\n        inputs3 = Input(shape=(self.maxlen,))\n        embedding3 = Embedding(self.max_features, self.embed_size, weights=[self.embeddings_matrix], trainable=gTrainableEmbeddings, mask_zero=self.mask_zero)(inputs3)\n        conv3 = Conv1D(filters=num_filters, kernel_size=kernel_sizes[2], activation='relu')(embedding3)\n        drop3 = Dropout(0.5)(conv3)\n        pool3 = MaxPooling1D(pool_size=2)(drop3)\n        flat3 = Flatten()(pool3)\n        # merge\n        merged = concatenate([flat1, flat2, flat3])\n        # interpretation\n        dense1 = Dense(10, activation='relu')(merged)\n        outputs = Dense(1, activation='sigmoid')(dense1)\n        model = ModelFactory.multi_gpu_model(Model(inputs=[inputs1, inputs2, inputs3], outputs=outputs))\n        model = ModelFactory.compilation(model)\n        \n        return model\n        \n\n    @classmethod\n    def compilation(cls, model):\n        model.compile(loss='binary_crossentropy',optimizer='adam', metrics=['accuracy'])        \n        return model\n    \n    def get_model(self, model_name, **kwargs):\n        retval = None\n        if model_name in self.models:\n            retval = self.models[model_name]\n        else:            \n            if model_name in ['2DCNNConcat']:\n                retval = self.create_2DCNN_ConcatModel(**kwargs)\n            elif model_name in ['GPUGRU3']:\n                retval = self.create_GPU_GRU3_Model(**kwargs)\n            elif model_name in ['GPULSTM']:\n                retval = self.create_GPU_LSTM_Model(**kwargs)\n            elif model_name in ['LSTMGRU']:\n                retval = self.create_LSTMGRU_Model(**kwargs)\n            elif model_name in ['CPULSTM']:\n                retval = self.create_CPU_LSTM_Model(**kwargs)\n            elif model_name in ['2DCNN']:\n                retval = self.create_2DCNN_Model(**kwargs)\n            elif model_name in ['1DCNNNgram']:\n                retval = self.create_1DCNNNgram_Model(**kwargs)\n            else:\n                assert True, \"No model of this name found.\"\n        \n            self.models[model_name] = retval\n        \n        return retval\n                \n    \n    def inspect(self):\n        if not gInspect:\n            return\n        \n        for name, model in self.models.items():\n            log(f\"<h1>Summary for {name}</h1>\")\n            model.summary()        \n    \n    @classmethod\n    def testme(cls):\n        sample_word2index = {'good': 0, 'bad': 1, 'sincere': 2, 'insincere': 3, 'sexual_intercourse':5, 'Donald_Trump' : 6}\n        index2word = {}\n        vocab2count = Counter()\n        for k, v in sample_word2index.items():\n            index2word[v] = k\n            if k in vocab2count:\n                vocab2count[k] += 1\n            else:\n                vocab2count[k] =1\n        \n        w2v, oov = EmbeddingsControl.load(gSelectedEmbeddings[0], sample_word2index)\n        me = cls(w2v)\n        me.get_model('GPULSTM')\n        me.get_model('1DCNNNgram')\n        me.get_model('GPUGRU3')\n        tmp = me.get_model('LSTMGRU')\n        tmp2 = clone_model(tmp)\n        tmp2.summary()\n        me.inspect()\n        del w2v\n        del oov\n        del me\n        gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ec3ff5d3b29f6c0df5322b88d5bc48f8eb49e6dc","scrolled":true},"cell_type":"code","source":"if False:\n    ModelFactory.testme()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"001de39ea420a7b23ade903a2601f39e6cf307ac"},"cell_type":"markdown","source":"\n##  NLPPatterns\n\n* The NLPPatterns only operates on numbers.  \n* It doesn't know anything about natural language.\n* It receives a model, training data as padded sequences."},{"metadata":{"trusted":true,"_uuid":"e324069b1e0f62a50898bcd8379d8a6d233ba06a"},"cell_type":"code","source":"class QuoraSequence(Sequence):\n    \n    def __init__(self, x_set, y_set, batch_size):\n        self.x, self.y = x_set, y_set\n        self.batch_size = batch_size\n\n    def __len__(self):\n        return int(np.ceil(self.x.shape[0] / float(self.batch_size)))\n\n    def __getitem__(self, idx):\n        batch_x = self.x[idx * self.batch_size:(idx + 1) * self.batch_size]\n        batch_y = self.y[idx * self.batch_size:(idx + 1) * self.batch_size]\n\n        return np.array(batch_x), np.array(batch_y)\n\nclass F1Evaluation(Callback):\n    def __init__(self, validation_data=(), interval=1, verbose=1):\n        super(Callback, self).__init__()\n\n        self.interval = interval\n        self.X_val, self.y_val = validation_data\n        self.verbose = verbose\n\n    def on_epoch_end(self, epoch, logs={}):\n        if epoch % self.interval == 0:\n            y_pred = self.model.predict(self.X_val, verbose=self.verbose)\n            score, threshold = NLPPatterns.find_prediction_threshold(self.y_val, y_pred)\n            #log(\"F1 Score - epoch: %d - threshold: %f - <b>score: %.6f</b>\" % (epoch+1, best_threshold, best_score))\n            log(\"\\n F1 Score - epoch: %d - score: %.6f \\n\" % (epoch+1, score))\n\nclass NLPPatterns:\n    \n    MODEL_CACHE_FILE = 'bestmodel.h5'\n    \n    def __init__(self, model, X, y, epochs=2):\n        assert(model is not None)\n        self.model = model\n        self.X_train = X\n        self.y_train = y\n        self.X_val = None\n        self.y_val = None\n        self.threshold = -1\n        self.score = 0\n        self.epochs = epochs\n\n        self.thresholds = []\n        self.scores = []        \n\n    @classmethod\n    def make_callbacks(cls, X_val, y_val, model_cachefile):\n            \n        # evaluator_cb = F1Evaluation(validation_data=(X_val, y_val), interval=1, verbose=1)\n        checkpoint_cb = ModelCheckpoint(filepath=model_cachefile, monitor='val_loss', save_best_only=True, verbose=2, mode='min')\n        earlystopping_cb = EarlyStopping(monitor='val_loss', min_delta=0.0001, patience=2, verbose=2, mode='auto')\n        reduce_lr_cb = ReduceLROnPlateau(monitor='val_loss', factor=0.6, patience=1, min_lr=0.0001, verbose=2)\n        \n        callbacks = [earlystopping_cb, checkpoint_cb, reduce_lr_cb]\n\n        return callbacks\n    \n    def data_partition(self, train_to_val_ratio=0.9, strategy=\"interleave\"):\n        \n        if strategy == \"interleave\":            \n            tmp_insincere = np.where(self.y_train == 1)[0]\n            tmp_sincere = np.where(self.y_train == 0)[0]\n            np.random.shuffle(tmp_sincere)\n            np.random.shuffle(tmp_insincere)\n\n            sincere_count = 0\n            insincere_count = 0\n            modulo = int(len(tmp_sincere)/len(tmp_insincere))+1\n            indices = np.zeros(len(self.y_train), dtype=np.int32)\n            \n            for i in range(len(indices)):\n                if i % modulo == 0 and insincere_count < len(tmp_insincere):\n                    indices[i] = tmp_insincere[insincere_count]\n                    insincere_count += 1\n                elif sincere_count < len(tmp_sincere):\n                    indices[i] = tmp_sincere[sincere_count]\n                    sincere_count += 1\n            \n            self.X_train = np.array(self.X_train)\n            indices = np.array(indices)\n            self.X_train = self.X_train[indices]\n            self.y_train = self.y_train[indices]\n            \n        \n        self.X_train, self.X_val, self.y_train, self.y_val = train_test_split(self.X_train, self.y_train, train_size=train_to_val_ratio, random_state=SEED, shuffle=True)\n        self.X_train = np.asarray(self.X_train)\n        self.y_train = np.asarray(self.y_train)\n        self.X_val = np.asarray(self.X_val)\n        self.y_val = np.asarray(self.y_val)\n        \n    def data_generator(self, train=False, val=False, batch_size=256):\n        X = None\n        y = None\n        \n        if train:\n            X = self.X_train\n            y = self.y_train\n        \n        if val:\n            X = self.X_val\n            y = self.y_val\n        \n        # Initialize a counter\n        counter = 0\n        num_batches = int(np.ceil(X.shape[0] / float(batch_size)))\n        while True:\n            for i in range(num_batches):\n                yield  X[i*batch_size:(i+1)*batch_size], y[i*batch_size:(i+1)*batch_size]\n\n    \n    @profile\n    def train(self, train_batch_size=256, validation_batch_size=1024, use_generator=False):\n        \n        callbacks = NLPPatterns.make_callbacks(self.X_val, self.y_val, NLPPatterns.MODEL_CACHE_FILE)\n        \n        if use_generator:\n            num_batches = int(np.ceil(self.X_train.shape[0] / float(train_batch_size)))\n            training_generator = self.data_generator(train=True, batch_size=train_batch_size)\n            \n            self.model.fit_generator(generator=training_generator, validation_data=(self.X_val, self.y_val), steps_per_epoch=num_batches, epochs=self.epochs, use_multiprocessing=True, workers=mpc.cpu_count(), callbacks=callbacks,verbose=1)\n        elif gModelName == \"1DCNNNgram\":\n            self.model.fit([self.X_train, self.X_train, self.X_train], self.y_train, validation_data=([self.X_val, self.X_val, self.X_val], self.y_val), batch_size=train_batch_size, epochs=self.epochs, callbacks=callbacks, verbose=1)\n        else:\n            self.model.fit(self.X_train, self.y_train, validation_data=(self.X_val, self.y_val), batch_size=train_batch_size, epochs=self.epochs, callbacks=callbacks, verbose=1)\n            \n        self.load()\n        self.validate(batch_size=validation_batch_size)\n        \n    @profile\n    def validate(self, batch_size=1024):        \n        \n        y_val_pred = None\n        if gModelName == \"1DCNNNgram\":\n            y_val_pred = self.model.predict([self.X_val, self.X_val, self.X_val], batch_size=batch_size)\n        else:\n            y_val_pred = self.model.predict(self.X_val, batch_size=batch_size)\n        \n        self.score, self.threshold = NLPPatterns.find_prediction_threshold(self.y_val, y_val_pred)      \n        log('optimal F1: {:.4f} at threshold: {:.4f}'.format(self.score, self.threshold))\n        \n    @profile\n    def predict(self, X_test, batch_size=1024):\n        if gModelName == \"1DCNNNgram\":\n            retval = self.model.predict([X_test, X_test, X_test], batch_size=batch_size, verbose=1)\n        else:\n            retval = self.model.predict(X_test, batch_size=batch_size, verbose=1)\n            \n        return retval\n    \n    def save(self):\n        if os.path.isfile(NLPPatterns.MODEL_CACHE_FILE):\n            return\n        \n        self.model.save(NLPPatterns.MODEL_CACHE_FILE)        \n        log(\"Saved model to disk\")   \n            \n    def load(self):\n        self.model = load_model(NLPPatterns.MODEL_CACHE_FILE, custom_objects={'AttentionLayer': AttentionLayer})\n        log(\"Loaded model from disk\")\n    \n    @classmethod\n    def find_prediction_threshold(cls, y_true, y_proba):\n        \n        best_threshold = 0.01\n        best_score = 0.0\n        for threshold in [i * 0.01 for i in range(1,100)]:\n            tmp = (y_proba > threshold).astype(int)\n            if np.count_nonzero(tmp) == 0:\n                continue\n            score = f1_score(y_true=y_true, y_pred=tmp)\n            if score > best_score:\n                best_threshold = threshold\n                best_score = score\n\n        return best_score, best_threshold\n    \n    @classmethod\n    def clean_cache(cls):\n        remove_file(NLPPatterns.MODEL_CACHE_FILE)\n    \n    def process(self):\n        \n        self.data_partition(strategy=\"interleave\")\n\n        if os.path.isfile(NLPPatterns.MODEL_CACHE_FILE):\n            self.load()\n            self.validate()\n            return\n        \n        if has_gpu():\n            self.train(train_batch_size=gTrainingBatchSize, use_generator=False)\n        else:\n            #self.train(use_sequence=True)\n            self.train(train_batch_size=gTrainingBatchSize, use_generator=True)\n    \n    def inspect(self):\n        if not gInspect:\n            return\n        \n        log(f\"Shape of Training input set {self.X_train.shape}\")\n        log(f\"Shape of Training output set {self.y_train.shape}\")\n        log(f\"Shape of Validation input set {self.X_val.shape}\")\n        log(f\"Shape of Validation output {self.y_val.shape}\")\n        log_list(self.y_val[0:25], desc=\"y validation samples (25)\")\n        log(f\"At threshold {self.threshold}, best F1 score: {self.score}\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"43758331488da480c077c90f6ce51071bd4bb40a"},"cell_type":"code","source":"if False:\n    NLPPatterns.testme()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d578320d50c63aa566a171c327b832aade8c1b5"},"cell_type":"markdown","source":"# The Flow"},{"metadata":{"trusted":true,"_uuid":"77c6ae45cd6c53bd2a685f37d9076dfe865b8485"},"cell_type":"code","source":"def initialize_cache():\n    if gFirstTime:\n        log(\"<hr><h2 align='center'>Initialize Cache</h2><hr>\")\n        log(\"Before cleanup...\")\n        log_dir('../working')\n        log(f\"Inspection is {gInspect}\")\n\n        DataManager.clean_cache()\n        InsincereVocabGenerator.clean_cache(bow=True, vocab=True)\n        QuestionsGenerator.clean_cache(remove_output=True)\n        for source, _ in gEmbeddingsSources.items():\n            QuoraPreprocessor.clean_cache(cleaned_q=True, vocab=True, word2index=True, sequences=True)\n        \n        EmbeddingsControl.clean_cache() \n        NLPPatterns.clean_cache()\n        log(\"After cleanup, cache contents:\")\n        log_dir('../working')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4a6b1e93d6cc207ad627b453e6314e5c735f433c"},"cell_type":"code","source":"def synthesize_data():    \n    if gGenerateNewData == False:\n        log(\"<hr><h2 align='center'>Using Original Training Data ONLY</h2><hr>\")\n        return\n    \n    log(\"<hr><h2 align='center'>Generate Insincere Vocab</h2><hr>\")\n\n    gen_vocab = InsincereVocabGenerator()\n    gen_vocab.process(verbose=False)\n    if gInspect:\n        gen_vocab.inspect(vocab=False)\n\n    log(\"<hr><h2 align='center'>Generate Questions</h2><hr>\")\n    \n    gen_q = QuestionsGenerator()\n    gen_q.process()\n    \n    # Combined generated cached file with input to create big training data cache.\n    app = DataManager.instance()\n    app.save_combined_data()\n    if gInspect:\n        gen_q.inspect(limit=10)\n        \n    if gLimit == False:\n        # Since we have created a combined file.\n        QuestionsGenerator.clean_cache(remove_output=True)\n        TopicalWords.memclean()\n        del gen_q\n        del gen_vocab\n        gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f75c1bea5b363cf6513ffbd88a34e43568dcd5c7"},"cell_type":"code","source":"def massage_data():\n    log(\"<hr><h2 align='center'>Pre-Processing Question Text</h2><hr>\")\n\n    app = DataManager.instance()\n    app.load(combined=(gGenerateNewData==True), orig=(gGenerateNewData == False), test=True)\n    \n    if gWIP:\n        stats = []\n        stats.append([\"Training Data Size\", str(app.training_data.shape)])\n        stats.append([\"Test Data Size\", str(app.test_data.shape)])\n        log_list(stats, desc=\"<big>Data Sizes</big>\")\n    \n    data2numbers = QuoraPreprocessor(training_data=app.training_data, test_data=app.test_data, is_gen=gGenerateNewData)\n    data2numbers.process()\n    data2numbers.inspect()\n    \n    all_w2v = {}\n    for source, _ in gEmbeddingsSources.items():        \n        if source not in gSelectedEmbeddings:\n            continue\n        a_w2v, a_oov = EmbeddingsControl.load(source, data2numbers.word2index)\n        all_w2v[source] = a_w2v\n        EmbeddingsControl.inspect(index2word=data2numbers.index2word, word2count=data2numbers.word2count, oov=a_oov, source=source, w2v=a_w2v)\n        del a_oov\n    \n    i=0\n    embeddings_matrix = None\n    for source in gSelectedEmbeddings:\n        if i == 0:\n            embeddings_matrix = all_w2v[source]\n        else:\n            embeddings_matrix += all_w2v[source]\n        i += 1\n            \n    embeddings_matrix = embeddings_matrix/len(gSelectedEmbeddings)\n    log(\"Combined embedding matrix shape:\" + str(embeddings_matrix.shape))\n\n    EmbeddingsControl.clean_cache()\n    save_binary(embeddings_matrix, \"MATRIX_ALL\")\n    del embeddings_matrix\n    \n    if gLimit == False:\n        for source, _ in gEmbeddingsSources.items():\n            QuoraPreprocessor.clean_cache(cleaned_q=False, vocab=True, word2index=True, sequences=False)\n    \n    del app\n    del data2numbers\n    del all_w2v\n    gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0995b46f6c69a8e26e4e4d6032fc6c759d5b4163"},"cell_type":"code","source":"@profile\ndef find_patterns():\n    log(\"<hr><h2 align='center'>Training</h2><hr>\")    \n    retval = {}\n    \n    gen_suffix = \"orig\"\n    if gGenerateNewData:\n        gen_suffix = \"gen\"\n\n    # Get X\n    sequences_file = QuoraPreprocessor.SEQUENCES_OUTPUT_FILE + \"_\" + gen_suffix\n    X = np.squeeze(pickle.load(open(sequences_file, 'rb')))\n\n    # Get y\n    cleaned_data_file = QuoraPreprocessor.CLEANED_OUTPUT_FILE + \"_\" + gen_suffix + \".csv\"\n    data = pd.read_csv(cleaned_data_file)\n    y = np.squeeze(data['target'].values)\n\n    # Get embeddings matrix\n    embeddings_matrix = pickle.load(open(\"MATRIX_ALL\", 'rb'))\n\n    # Select Model\n    model_factory = ModelFactory(embeddings_matrix=embeddings_matrix)\n    model = model_factory.get_model(gModelName, units=64)\n    model_factory.inspect()\n        \n    # Train\n    patterns = NLPPatterns(X=X,y=y,model=model,epochs=gNumEpochs)\n    patterns.process()\n    patterns.inspect()\n\n    return patterns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c05e0cff824b18a0fe3c6dfd91f7296660ca9d4c"},"cell_type":"code","source":"def apply_patterns(patterns):\n    log(\"<hr><h2 align='center'>Predict</h2><hr>\")\n    \n    y_predictions = {}\n    sequences = QuoraPreprocessor.SEQUENCES_OUTPUT_FILE + \"_\" + \"test\"\n    X_test = pickle.load(open(sequences, 'rb'))\n    log(f\"Number of test samples: {len(X_test)}\")\n    X_test = np.squeeze(X_test)\n    y_predictions = patterns.predict(X_test)\n    \n    y_predictions = y_predictions > patterns.threshold\n    \n    return y_predictions","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"180f47e247a0230d00d64059c3327d672c16b9c3"},"cell_type":"code","source":"def save(output):\n    log(\"<hr><h2 align='center'>Submission of Test Predictions</h2><hr>\")\n    \n    y_pred = output\n    y_pred = y_pred.reshape(-1).astype(int)\n    \n    app = DataManager.instance()\n    app.load(orig=False,combined=False, test=True)\n    test_data = app.test_data\n    log(f\"Shapes<br>test_data {test_data.shape}\")\n    log(f\"output {y_pred.shape}\")\n    \n    if gInspect:\n        display_frame = pd.DataFrame()\n        display_frame = display_frame.assign(qid=test_data['qid'],question_text=test_data['question_text'],prediction=y_pred)\n        display(display_frame[display_frame.prediction == 1].head(100))\n\n    submission = pd.DataFrame()\n    submission = submission.assign(qid=test_data['qid'],prediction=y_pred)\n    submission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1e5a4e6acabad4fdf865e218efb1e66176dec83f"},"cell_type":"markdown","source":"# Main: Where It Begins"},{"metadata":{"trusted":true,"_uuid":"48e0a237e57fff340713908c143901c556954a82"},"cell_type":"code","source":"@profile\ndef main():\n    global gFirstTime\n    initialize_cache()\n    synthesize_data()\n    massage_data()\n    NLPPatterns.clean_cache()\n    gFirstTime = False\n    patterns = find_patterns()\n    output = apply_patterns(patterns)\n    save(output)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"39b75bfd4bdf268b87c58cbd94206f9555835009","scrolled":false},"cell_type":"code","source":"if gExecuteMain:\n    log_current_memory(caption=\"Memory before starting Main\")\n    clear_prof_data()\n    main()\n    pp_prof_data()\n    log_current_memory(caption=\"Memory after Main\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8839e84d8e648fc90e96ccdf863e2a6c2a70d50e"},"cell_type":"markdown","source":"# Rough Work area"},{"metadata":{"trusted":true,"_uuid":"f91554d12c95f69b097032d83c719552c0f6a72a"},"cell_type":"code","source":"if False:\n    # initialize_cache()\n    # remove_file('sequences')\n    # remove_file('submission.csv')\n    log_dir('../working')\n    log_current_memory(caption=\"The End\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3799315e38bde1a87c9718521202e6c2f2674d10"},"cell_type":"code","source":"if False:\n    app = DataManager.instance()\n    app.load(source_file=\"cleaned_q.csv\")\n    data = app.training_data\n    log = []\n    for i in range(0, 200, 10):\n        df = data[(data.target==1) & (data.num_words > i)]\n        log.append([str(i), len(df)])\n    log_list(log, desc=\"Question Lengths Count\")\n    \n    word2index = pickle.load(open(QuoraPreprocessor.W2INDEX_OUTPUT_FILE+\"_GOOGLENEWS\", \"rb\"))\n    log(f'vocab_size: {len(word2index)}')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8151a1a59f7428dee23b617048f9f75fc5a6d471"},"cell_type":"code","source":"if False:\n    tmp = pickle.load(open('no_gen_q', 'rb'))\n    tmp = [q for result in tmp for q in result]\n    log_list(tmp[0:20])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"31968070d4dc54055237c4ae13e16e9f8be3a7cf"},"cell_type":"code","source":"if False:\n    y_train = np.asarray([0,0,0,1,1,0,1,0,0,0,1,0,0,1,0,0,0,0,1, 0, 0])\n    tmp_X = np.repeat([\"Hello\"], y_train.shape[0])\n    X = []\n    for i in range(y_train.shape[0]):\n        tmp = []\n        for j in range(10):\n            tmp.append(tmp_X[i] + str(i) + \"--\" + str(y_train[i]))\n        X.append(tmp)\n        \n    print(X)\n    \n    tmp_insincere = np.where(y_train == 1)[0]\n    tmp_sincere = np.where(y_train == 0)[0]\n    \n    np.random.shuffle(tmp_sincere)\n    np.random.shuffle(tmp_insincere)\n    \n    print(tmp_insincere)\n    print(tmp_sincere)\n    \n    print(str(len(tmp_insincere)))\n    print(str(len(tmp_sincere)))\n    modulo = int(len(tmp_sincere)/len(tmp_insincere))+1\n    print(modulo)\n    \n    sincere_count = 0\n    insincere_count = 0\n    indices = np.zeros(len(y_train),dtype=np.int8)\n    for i in range(len(indices)):\n        if i % modulo == 0 and insincere_count < len(tmp_insincere):\n            indices[i] = tmp_insincere[insincere_count]\n            insincere_count += 1\n        elif sincere_count < len(tmp_sincere):\n            indices[i] = tmp_sincere[sincere_count]\n            sincere_count += 1\n    \n    print(\"Before:\\n\", y_train)\n    indices = np.array(indices)\n    print(y_train[indices])\n    X = np.array(X)\n    print(X[indices])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cdaf76f258678e06ea56d49790422cd8f913722f"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dacc15c09c5604a0967a70be00afbcb4a9398b53"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}