{"cells":[{"metadata":{"_uuid":"01f361ddc47e0b386595316fe3d7f4dabbd260db"},"cell_type":"markdown","source":"In what follows I tried to SMOTE by using the bottleneck features, but it really did not work. I suspect that\nwhen the classification is applied, then the weights of the networ tend to project the data points in a region, \nremoving a lot of variance in the process. So smoting would basically give very little to no added value.\nWell, it was an attempt. Maybe you get an idea or two, below I also combined topic models with attention to see if there was\nany improvement.\n\n\n**Notebook Objective:**\n\nThe objective of this notebook is to try an attention model I worked on in the past on a similar task that seemed to behave rather well. The\nidea is basically is to make use of a topic model in addition to the standard deep learning network.\n\nBelow I initialize the environment.\n\nReferences:\n\n### https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\n### https://www.kaggle.com/hireme/fun-api-keras-f1-metric-cyclical-learning-rate/code\n### https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\n\n"},{"metadata":{"_uuid":"a25e470489124c2d93b1ef2b0681ba901a654118","trusted":true},"cell_type":"code","source":"import numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import metrics\nfrom sklearn.model_selection import GridSearchCV, StratifiedKFold\nfrom sklearn.metrics import f1_score, roc_auc_score\n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers import Dense, Input, CuDNNLSTM, Embedding, Dropout, Activation, CuDNNGRU, Conv1D\nfrom keras.layers import Bidirectional, GlobalMaxPool1D, GlobalMaxPooling1D, GlobalAveragePooling1D\nfrom keras.layers import Input, Embedding, Dense, Conv2D, MaxPool2D, concatenate\nfrom keras.layers import Reshape, Flatten, Concatenate, Dropout, SpatialDropout1D\nfrom keras.optimizers import Adam\nfrom keras.models import Model\nfrom keras import backend as K\nfrom keras.engine.topology import Layer\nfrom keras import initializers, regularizers, constraints, optimizers, layers\nfrom keras.layers import concatenate\nfrom keras.callbacks import *\n\niterations_learning = 24\nbatch_size_learning= 4096\nimportance = 0.9999\nsplits_kfold = 4\n\ntrain_batch_size = 2048\ntrain_epochs = 15\n\n\ndirectory_emb = '../input/embeddings/'\ndirectory_data = '../input/'\n\ndef load_glove(word_index):\n    EMBEDDING_FILE = directory_emb+'glove.840B.300d/glove.840B.300d.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = -0.005838499,0.48782197\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return embedding_matrix \n    \ndef load_fasttext(word_index):    \n    EMBEDDING_FILE = directory_emb+'wiki-news-300d-1M/wiki-news-300d-1M.vec'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE) if len(o)>100)\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n\n    return embedding_matrix\n\ndef load_para(word_index):\n    EMBEDDING_FILE = directory_emb + 'paragram_300_sl999/paragram_300_sl999.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore') if len(o)>100)\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = -0.0053247833,0.49346462\n    embed_size = all_embs.shape[1]\n    print(emb_mean,emb_std,\"para\")\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n    \n    return embedding_matrix\n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"import os\nimport time\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nimport math\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import metrics\n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers import Dense, Input, LSTM, Embedding, Dropout, Activation, CuDNNGRU, Conv1D, RepeatVector, Flatten\nfrom keras.layers import Bidirectional, GlobalMaxPool1D, Permute, Lambda, Layer, BatchNormalization, GRU\nfrom keras.models import Model\nfrom keras import initializers, regularizers, constraints, optimizers, layers, callbacks\nfrom keras.callbacks import Callback","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9e769511a2a4ab3225bd708fc3cfd4873e9f1eb2","trusted":true},"cell_type":"code","source":"import re\nimport time\nimport gc\nimport random\nimport os\n\nimport numpy as np\nimport pandas as pd\n\nmax_features = 95000 # how many unique words to use (i.e num rows in embedding vector)\nmaxlen = 72\n\npuncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]\n\ndef clean_text(x):\n    x = str(x)\n    for punct in puncts:\n        x = x.replace(punct, f' {punct} ')\n    return x\n\ndef clean_numbers(x):\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x\n\nmispell_dict = {\"aren't\" : \"are not\",\n\"can't\" : \"cannot\",\n\"couldn't\" : \"could not\",\n\"didn't\" : \"did not\",\n\"doesn't\" : \"does not\",\n\"don't\" : \"do not\",\n\"hadn't\" : \"had not\",\n\"hasn't\" : \"has not\",\n\"haven't\" : \"have not\",\n\"he'd\" : \"he would\",\n\"he'll\" : \"he will\",\n\"he's\" : \"he is\",\n\"i'd\" : \"I would\",\n\"i'd\" : \"I had\",\n\"i'll\" : \"I will\",\n\"i'm\" : \"I am\",\n\"isn't\" : \"is not\",\n\"it's\" : \"it is\",\n\"it'll\":\"it will\",\n\"i've\" : \"I have\",\n\"let's\" : \"let us\",\n\"mightn't\" : \"might not\",\n\"mustn't\" : \"must not\",\n\"shan't\" : \"shall not\",\n\"she'd\" : \"she would\",\n\"she'll\" : \"she will\",\n\"she's\" : \"she is\",\n\"shouldn't\" : \"should not\",\n\"that's\" : \"that is\",\n\"there's\" : \"there is\",\n\"they'd\" : \"they would\",\n\"they'll\" : \"they will\",\n\"they're\" : \"they are\",\n\"they've\" : \"they have\",\n\"we'd\" : \"we would\",\n\"we're\" : \"we are\",\n\"weren't\" : \"were not\",\n\"we've\" : \"we have\",\n\"what'll\" : \"what will\",\n\"what're\" : \"what are\",\n\"what's\" : \"what is\",\n\"what've\" : \"what have\",\n\"where's\" : \"where is\",\n\"who'd\" : \"who would\",\n\"who'll\" : \"who will\",\n\"who're\" : \"who are\",\n\"who's\" : \"who is\",\n\"who've\" : \"who have\",\n\"won't\" : \"will not\",\n\"wouldn't\" : \"would not\",\n\"you'd\" : \"you would\",\n\"you'll\" : \"you will\",\n\"you're\" : \"you are\",\n\"you've\" : \"you have\",\n\"'re\": \" are\",\n\"wasn't\": \"was not\",\n\"we'll\":\" will\",\n\"didn't\": \"did not\",\n\"tryin'\":\"trying\"}\n\ndef _get_mispell(mispell_dict):\n    mispell_re = re.compile('(%s)' % '|'.join(mispell_dict.keys()))\n    return mispell_dict, mispell_re\n\nmispellings, mispellings_re = _get_mispell(mispell_dict)\ndef replace_typical_misspell(text):\n    def replace(match):\n        return mispellings[match.group(0)]\n    return mispellings_re.sub(replace, text)\n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(directory_data+\"train.csv\")\ntest_df = pd.read_csv(directory_data+\"test.csv\")\n\n#train_X, test_X, train_y, word_index = load_and_prec()\n\nprint(\"Train shape : \",train_df.shape)\nprint(\"Test shape : \",test_df.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73301b57e449531aa251bdfc5889e172e55492ed","trusted":true},"cell_type":"code","source":"contraction_mapping = {\"ain't\": \"is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"'cause\": \"because\", \"could've\": \"could have\", \"couldn't\": \"could not\", \n                       \"didn't\": \"did not\",  \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hasn't\": \"has not\", \"haven't\": \"have not\", \n                       \"he'd\": \"he would\",\"he'll\": \"he will\", \"he's\": \"he is\", \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\",\n                       \"how's\": \"how is\",  \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\n                       \"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \n                       \"i'd've\": \"i would have\", \"i'll\": \"i will\",  \"i'll've\": \"i will have\",\"i'm\": \"i am\", \n                       \"i've\": \"i have\", \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\",\n                       \"it'll\": \"it will\", \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \n                       \"ma'am\": \"madam\", \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\n                       \"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \"mustn't've\": \"must not have\", \n                       \"needn't\": \"need not\", \"needn't've\": \"need not have\",\"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \n                       \"oughtn't've\": \"ought not have\", \"shan't\": \"shall not\", \"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \n                       \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \"she'll've\": \"she will have\",\n                       \"she's\": \"she is\", \"should've\": \"should have\", \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\"so's\": \"so as\", \n                       \"this's\": \"this is\",\"that'd\": \"that would\", \"that'd've\": \"that would have\", \"that's\": \"that is\", \"there'd\": \"there would\", \n                       \"there'd've\": \"there would have\", \"there's\": \"there is\", \"here's\": \"here is\",\"they'd\": \"they would\", \"they'd've\": \"they would have\", \n                       \"they'll\": \"they will\", \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \"to've\": \"to have\", \n                       \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \"we'll\": \"we will\", \"we'll've\": \"we will have\", \n                       \"we're\": \"we are\", \"we've\": \"we have\", \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \n                       \"what're\": \"what are\",  \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\", \"when've\": \"when have\", \n                       \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \"who'll\": \"who will\", \"who'll've\": \"who will have\", \n                       \"who's\": \"who is\", \"who've\": \"who have\", \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \n                       \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \n                       \"y'all'd\": \"you all would\",\"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\n                       \"y'all've\": \"you all have\",\"you'd\": \"you would\", \n                       \"you'd've\": \"you would have\", \"you'll\": \"you will\", \"you'll've\": \"you will have\", \n                       \"you're\": \"you are\", \"you've\": \"you have\" }\n\npunct = \"/-'?!.,#$%\\'()*+-/:;<=>@[\\\\]^_`{|}~\" + '\"\"“”’' + '∞θ÷α•à−β∅³π‘₹´°£€\\×™√²—–&'\n\npunct_mapping = {\"‘\": \"'\", \"₹\": \"e\", \"´\": \"'\", \"°\": \"\", \n                 \"€\": \"e\", \"™\": \"tm\", \"√\": \" sqrt \", \"×\": \"x\", \"²\": \"2\", \"—\": \"-\", \"–\": \"-\", \n                 \"’\": \"'\", \"_\": \"-\", \"`\": \"'\", '“': '\"', '”': '\"', '“': '\"', \n                 \"£\": \"e\", '∞': 'infinity', 'θ': 'theta', '÷': '/', 'α': 'alpha', \n                 '•': '.', 'à': 'a', '−': '-', 'β': 'beta', '∅': '', '³': '3', 'π': 'pi', }\n\n\n\nmispell_dict = {'colour': 'color', 'centre': 'center', 'favourite': 'favorite', 'travelling': 'traveling', 'counselling': 'counseling', \n                'theatre': 'theater', 'cancelled': 'canceled', 'labour': 'labor', 'organisation': 'organization', 'wwii': 'world war 2',\n                'citicise': 'criticize', 'youtu ': 'youtube ', 'Qoura': 'Quora', 'sallary': 'salary', \n                'Whta': 'What', 'narcisist': 'narcissist', 'howdo': 'how do', 'whatare': 'what are', \n                'howcan': 'how can', 'howmuch': 'how much', 'howmany': 'how many', 'whydo': 'why do', 'doI': 'do I', \n                'theBest': 'the best', 'howdoes': 'how does', 'mastrubation': 'masturbation', 'mastrubate': 'masturbate', \"mastrubating\": 'masturbating',\n                'pennis': 'penis', 'Etherium': 'Ethereum', 'narcissit': 'narcissist', 'bigdata': 'big data', '2k17': '2017', '2k18': '2018', 'qouta': 'quota', \n                'exboyfriend': 'ex boyfriend', 'airhostess': 'air hostess', \"whst\": 'what', 'watsapp': 'whatsapp', 'demonitisation': 'demonetization', \n                'demonitization': 'demonetization', 'demonetisation': 'demonetization', 'trump':'president', 'obama':'president', 'potus':'president'}\n\n\n\n\ndef add_lower(embedding, vocab):\n    count = 0\n    for word in vocab:\n        if word in embedding and word.lower() not in embedding:  \n            embedding[word.lower()] = embedding[word]\n            count += 1\n    print(f\"Added {count} words to embedding\")\n\n\ndef clean_contractions(text, mapping):\n    specials = [\"’\", \"‘\", \"´\", \"`\"]\n    for s in specials:\n        text = text.replace(s, \"'\")\n    text = ' '.join([mapping[t] if t in mapping else t for t in text.split(\" \")])\n    return text\n\n\ndef unknown_punct(embed, punct):\n    unknown = ''\n    for p in punct:\n        if p not in embed:\n            unknown += p\n            unknown += ' '\n    return unknown\n\n\ndef clean_special_chars(text, punct, mapping):\n    for p in mapping:\n        text = text.replace(p, mapping[p])\n    \n    for p in punct:\n        text = text.replace(p, f' {p} ')\n    \n    specials = {'\\u200b': ' ', '…': ' ... ', '\\ufeff': '', 'करना': '', 'है': ''}  # Other special characters that I have to deal with in last\n    for s in specials:\n        text = text.replace(s, specials[s])\n    \n    return text\n\ndef correct_spelling(x, dic):\n    for word in dic.keys():\n        x = x.replace(word, dic[word])\n    return x\n\n\ndef transform_df(dfs):\n    dfs['question_text'] = dfs['question_text'].apply(lambda x: x.lower())\n    dfs['question_text'] = dfs['question_text'].apply(lambda x: clean_contractions(x, contraction_mapping))\n    dfs['question_text'] = dfs['question_text'].apply(lambda x: clean_special_chars(x, punct, punct_mapping))\n    dfs['question_text'] = dfs['question_text'].apply(lambda x: correct_spelling(x, mispell_dict))\n\n    \ntransform_df(train_df)\ntransform_df(test_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ad1e9fdf97f3291a7d1797b25f0c7a0c5d1f1edd"},"cell_type":"markdown","source":"# As in other kernels, preprocessing steps.\n\nNext steps are as follows:\n * Split the training dataset into train and val sample. Cross validation is a time consuming process and so let us do simple train val split.\n * Fill up the missing values in the text column with '_na_'\n * Tokenize the text column and convert them to vector sequences\n * Pad the sequence as needed - if the number of words in the text is greater than 'max_len' trunacate them to 'max_len' or if the number of words in the text is lesser than 'max_len' add zeros for remaining values."},{"metadata":{"_uuid":"ba5a1b8109dee2c9fbc628d5da4a7c3447d42fb8","trusted":true},"cell_type":"code","source":"## split to train and val\ntrain_df, val_df = train_test_split(train_df, test_size=0.1, random_state=2018)\n\n## some config values \nembed_size = 300 # how big is each word vector\nmax_features = 95000 # how many unique words to use (i.e num rows in embedding vector)\nmaxlen = 72 # max number of words in a question to use\n\n## fill up the missing values\ntrain_X = train_df[\"question_text\"].fillna(\"_na_\").values\nval_X = val_df[\"question_text\"].fillna(\"_na_\").values\ntest_X = test_df[\"question_text\"].fillna(\"_na_\").values\n\n## Tokenize the sentences\ntokenizer = Tokenizer(num_words=max_features)\ntokenizer.fit_on_texts(list(train_X))\ntrain_X = tokenizer.texts_to_sequences(train_X)\nval_X = tokenizer.texts_to_sequences(val_X)\ntest_X = tokenizer.texts_to_sequences(test_X)\n\n## Pad the sentences \ntrain_X = pad_sequences(train_X, maxlen=maxlen)\nval_X = pad_sequences(val_X, maxlen=maxlen)\ntest_X = pad_sequences(test_X, maxlen=maxlen)\n\n## Get the target values\ntrain_y = train_df['target'].values\nval_y = val_df['target'].values","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2a5f324273d8e4726a6f0f9206170845d5ead890"},"cell_type":"markdown","source":"**My own model without Pretrained Embeddings:**\n\nFor the moment I am not using the pretrained embeddings, I am used an embedding trained in the domain of Quora\nand then I put attention on the topics of the message. "},{"metadata":{"_uuid":"abca54b8956820fddfc69dfda1bd0cc2dae120c5","trusted":true},"cell_type":"code","source":"class Attention(Layer):\n    def __init__(self, step_dim,\n                 W_regularizer=None, b_regularizer=None,\n                 W_constraint=None, b_constraint=None,\n                 bias=True, **kwargs):\n        self.supports_masking = True\n        self.init = initializers.get('glorot_uniform')\n\n        self.W_regularizer = regularizers.get(W_regularizer)\n        self.b_regularizer = regularizers.get(b_regularizer)\n\n        self.W_constraint = constraints.get(W_constraint)\n        self.b_constraint = constraints.get(b_constraint)\n\n        self.bias = bias\n        self.step_dim = step_dim\n        self.features_dim = 0\n        super(Attention, self).__init__(**kwargs)\n\n    def build(self, input_shape):\n        assert len(input_shape) == 3\n\n        self.W = self.add_weight((input_shape[-1],),\n                                 initializer=self.init,\n                                 name='{}_W'.format(self.name),\n                                 regularizer=self.W_regularizer,\n                                 constraint=self.W_constraint)\n        self.features_dim = input_shape[-1]\n\n        if self.bias:\n            self.b = self.add_weight((input_shape[1],),\n                                     initializer='zero',\n                                     name='{}_b'.format(self.name),\n                                     regularizer=self.b_regularizer,\n                                     constraint=self.b_constraint)\n        else:\n            self.b = None\n\n        self.built = True\n\n    def compute_mask(self, input, input_mask=None):\n        return None\n\n    def call(self, x, mask=None):\n        features_dim = self.features_dim\n        step_dim = self.step_dim\n\n        eij = K.reshape(K.dot(K.reshape(x, (-1, features_dim)),\n                        K.reshape(self.W, (features_dim, 1))), (-1, step_dim))\n\n        if self.bias:\n            eij += self.b\n\n        eij = K.tanh(eij)\n\n        a = K.exp(eij)\n\n        if mask is not None:\n            a *= K.cast(mask, K.floatx())\n\n        a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())\n\n        a = K.expand_dims(a)\n        weighted_input = x * a\n        return K.sum(weighted_input, axis=1)\n\n    def compute_output_shape(self, input_shape):\n        return input_shape[0],  self.features_dim","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a52c49481aa75e8b9b53efebef9dc8a461fe936a","trusted":true},"cell_type":"code","source":"def squash(x, axis=-1):\n    s_squared_norm = K.sum(K.square(x), axis, keepdims=True)\n    scale = K.sqrt(s_squared_norm + K.epsilon())\n    return x / scale\n\nclass Capsule(Layer):\n    def __init__(self, num_capsule, dim_capsule, routings=3, kernel_size=(9, 1), share_weights=True,\n                 activation='default', **kwargs):\n        super(Capsule, self).__init__(**kwargs)\n        self.num_capsule = num_capsule\n        self.dim_capsule = dim_capsule\n        self.routings = routings\n        self.kernel_size = kernel_size\n        self.share_weights = share_weights\n        if activation == 'default':\n            self.activation = squash\n        else:\n            self.activation = Activation(activation)\n\n    def build(self, input_shape):\n        super(Capsule, self).build(input_shape)\n        input_dim_capsule = input_shape[-1]\n        if self.share_weights:\n            self.W = self.add_weight(name='capsule_kernel',\n                                     shape=(1, input_dim_capsule,\n                                            self.num_capsule * self.dim_capsule),\n                                     # shape=self.kernel_size,\n                                     initializer='glorot_uniform',\n                                     trainable=True)\n        else:\n            input_num_capsule = input_shape[-2]\n            self.W = self.add_weight(name='capsule_kernel',\n                                     shape=(input_num_capsule,\n                                            input_dim_capsule,\n                                            self.num_capsule * self.dim_capsule),\n                                     initializer='glorot_uniform',\n                                     trainable=True)\n\n    def call(self, u_vecs):\n        if self.share_weights:\n            u_hat_vecs = K.conv1d(u_vecs, self.W)\n        else:\n            u_hat_vecs = K.local_conv1d(u_vecs, self.W, [1], [1])\n\n        batch_size = K.shape(u_vecs)[0]\n        input_num_capsule = K.shape(u_vecs)[1]\n        u_hat_vecs = K.reshape(u_hat_vecs, (batch_size, input_num_capsule,\n                                            self.num_capsule, self.dim_capsule))\n        u_hat_vecs = K.permute_dimensions(u_hat_vecs, (0, 2, 1, 3))\n        # final u_hat_vecs.shape = [None, num_capsule, input_num_capsule, dim_capsule]\n\n        b = K.zeros_like(u_hat_vecs[:, :, :, 0])  # shape = [None, num_capsule, input_num_capsule]\n        for i in range(self.routings):\n            b = K.permute_dimensions(b, (0, 2, 1))  # shape = [None, input_num_capsule, num_capsule]\n            c = K.softmax(b)\n            c = K.permute_dimensions(c, (0, 2, 1))\n            b = K.permute_dimensions(b, (0, 2, 1))\n            outputs = self.activation(tf.keras.backend.batch_dot(c, u_hat_vecs, [2, 2]))\n            if i < self.routings - 1:\n                b = tf.keras.backend.batch_dot(outputs, u_hat_vecs, [2, 3])\n\n        return outputs\n\n    def compute_output_shape(self, input_shape):\n        return (None, self.num_capsule, self.dim_capsule)\n\n# DropConnect\n# https://github.com/andry9454/KerasDropconnect\n\nfrom keras.layers import Wrapper\n\nclass DropConnect(Wrapper):\n    def __init__(self, layer, prob=1., **kwargs):\n        self.prob = prob\n        self.layer = layer\n        super(DropConnect, self).__init__(layer, **kwargs)\n        if 0. < self.prob < 1.:\n            self.uses_learning_phase = True\n\n    def build(self, input_shape):\n        if not self.layer.built:\n            self.layer.build(input_shape)\n            self.layer.built = True\n        super(DropConnect, self).build()\n\n    def compute_output_shape(self, input_shape):\n        return self.layer.compute_output_shape(input_shape)\n\n    def call(self, x):\n        if 0. < self.prob < 1.:\n            self.layer.kernel = K.in_train_phase(K.dropout(self.layer.kernel, self.prob), self.layer.kernel)\n            self.layer.bias = K.in_train_phase(K.dropout(self.layer.bias, self.prob), self.layer.bias)\n        return self.layer.call(x)\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"327b389420b7604bbd1790dd71905eba91aad362","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nfrom tqdm import tqdm\nimport math\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import metrics\nfrom sklearn.model_selection import GridSearchCV, StratifiedKFold\nfrom sklearn.metrics import f1_score, roc_auc_score\n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers import Dense, Input, CuDNNLSTM, Embedding, Dropout, Activation, CuDNNGRU, Conv1D\nfrom keras.layers import Bidirectional, GlobalMaxPool1D, GlobalMaxPooling1D, GlobalAveragePooling1D\nfrom keras.layers import Input, Embedding, Dense, Conv2D, MaxPool2D, concatenate, add, BatchNormalization\nfrom keras.layers import Reshape, Flatten, Concatenate, Dropout, SpatialDropout1D\nfrom keras.optimizers import Adam\nfrom keras.models import Model\nfrom keras import backend as K\nfrom keras.engine.topology import Layer\nfrom keras import initializers, regularizers, constraints, optimizers, layers\nfrom keras.initializers import glorot_normal, orthogonal\nfrom keras.layers import concatenate\nfrom keras.callbacks import *\n\nimport tensorflow as tf\n\nimport re\nimport time\nimport os\n\n\nfrom keras.layers.core import *\n\nimport tensorflow as tf\nimport keras.backend as K\nfrom keras.losses import binary_crossentropy\n\ndef f1(y_true, y_pred):\n    y_pred = K.round(y_pred)\n    tp = K.sum(K.cast(y_true*y_pred, 'float'), axis=0)\n    tn = K.sum(K.cast((1-y_true)*(1-y_pred), 'float'), axis=0)\n    fp = K.sum(K.cast((1-y_true)*y_pred, 'float'), axis=0)\n    fn = K.sum(K.cast(y_true*(1-y_pred), 'float'), axis=0)\n\n    p = tp / (tp + fp + K.epsilon())\n    r = tp / (tp + fn + K.epsilon())\n\n    f1 = 2*p*r / (p+r+K.epsilon())\n    f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n    return K.mean(f1)\n\ndef f1_loss(y_true, y_pred):\n    \n    tp = K.sum(K.cast(y_true*y_pred, 'float'), axis=0)\n    tn = K.sum(K.cast((1-y_true)*(1-y_pred), 'float'), axis=0)\n    fp = K.sum(K.cast((1-y_true)*y_pred, 'float'), axis=0)\n    fn = K.sum(K.cast(y_true*(1-y_pred), 'float'), axis=0)\n\n    p = tp / (tp + fp + K.epsilon())\n    r = tp / (tp + fn + K.epsilon())\n\n    f1 = 2*p*r / (p+r+K.epsilon())\n    f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n    return 1 - K.mean(f1)\n\n\n# the \ndef my_model_lstm(topic,CONTEXT_LENGTH, embedding_size, max_features, embedding_matrix):\n        L = CONTEXT_LENGTH\n        NEURONS = 128\n        \n        inp = Input(shape=(L, ))\n        encoding_input = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n        drop_out = Dropout(0.1, name='dropout')(encoding_input)\n        lstm_fwd = LSTM(NEURONS, return_sequences=True, name='lstm_fwd')(drop_out)\n        lstm_bwd = LSTM(NEURONS, return_sequences=True, go_backwards=True, name='lstm_bwd')(drop_out)\n        \n        import keras\n        bilstm = keras.layers.concatenate([lstm_fwd, lstm_bwd], name='bilstm')\n        \n        #bilstm = merge([lstm_fwd, lstm_bwd], name='bilstm', mode='concat')\n        drop_out = Dropout(0.1)(bilstm)\n        \n        topic_input = Input(shape=(topic,), name='top_in')\n        attention_topic = Dense(NEURONS, activation='linear')(topic_input)\n        attention_topic =  RepeatVector(L)(attention_topic) #bidirectional attention model\n        #attention_topic = Flatten()(attention_topic)\n        \n        # compute importance for each step\n        attention = Dense(NEURONS, activation='linear')(drop_out)\n        #attention = Flatten()(attention)\n        \n        attention = keras.layers.add([attention,attention_topic])\n        \n        \n        attention = Dense(1, activation='tanh')(attention) # sequence + topic\n        attention = Flatten()(attention)\n        \n        attention = Activation('softmax')(attention)\n        attention = RepeatVector(NEURONS*2)(attention) #bidirectional attention model\n        attention = Permute([2, 1])(attention)\n        \n        sent_representation = keras.layers.multiply([drop_out, attention])\n        sent_representation = Lambda(lambda xin: K.sum(xin, axis=-2), output_shape=(NEURONS*2,))(sent_representation)\n        sent_representation = BatchNormalization(axis=-1)(sent_representation)\n        bottleneck_features = Dense(NEURONS, activation='relu', name='bottleneck')(sent_representation) #the projection for smoting  \n        out = Dense(1, activation='sigmoid')(bottleneck_features)\n        #out2 = Dense(expected_length, activation='softmax')(sent_representation)\n        #output = [out,out2]\n        model = Model(input=[inp,topic_input], output=out)\n        model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy',f1])\n        print(model.summary())\n        return model  \n\n    \n       \n\ndef my_small_model(size_intermediate):\n    inp = Input(shape=(size_intermediate,))\n    bn = BatchNormalization(axis=-1)(inp)\n    interm = Dense(64, activation='relu')(bn)\n    drop_out = Dropout(0.1)(interm)\n    interm = Dense(32, activation='relu')(drop_out)\n    drop_out = Dropout(0.1)(interm)\n    out = Dense(1, activation='sigmoid')(drop_out)\n    model = Model(input=inp, output=out)\n    model.compile(loss=[binary_crossentropy], optimizer='adam', metrics=[f1])\n    return model      \n   \n\n\n\ndef build_my_model(embedding_matrix):\n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = SpatialDropout1D(rate=0.24)(x)\n    x = Bidirectional(CuDNNLSTM(80, \n                                return_sequences=True, \n                                kernel_initializer=glorot_normal(seed=1029), \n                                recurrent_initializer=orthogonal(gain=1.0, seed=1029)))(x)\n\n    x_1 = Attention(maxlen)(x)\n    x_1 = DropConnect(Dense(32, activation=\"relu\"), prob=0.1)(x_1)\n    \n    x_2 = Capsule(num_capsule=10, dim_capsule=10, routings=4, share_weights=True)(x)\n    x_2 = Flatten()(x_2)\n    x_2 = DropConnect(Dense(32, activation=\"relu\"), prob=0.1)(x_2)\n\n    conc = concatenate([x_1, x_2], name='bottleneck')\n    \n    # conc = add([x_1, x_2])\n    outp = Dense(1, activation=\"sigmoid\")(conc)\n    model = Model(inputs=inp, outputs=outp)\n    model.compile(loss='binary_crossentropy', optimizer=Adam(), metrics=[f1])\n    return model\n    \n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ad1915b1fda1c790da5b05719a6ab3615cd7f1f3","trusted":true},"cell_type":"code","source":"np.add([[1,2,3],[1,2,3]],[[4,5,6],[4,5,6]])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ef5f647f468235395ec512f9e59c41ce97aa6964","trusted":true},"cell_type":"code","source":"# https://www.kaggle.com/hireme/fun-api-keras-f1-metric-cyclical-learning-rate/code\n\nclass CyclicLR(Callback):\n    \"\"\"This callback implements a cyclical learning rate policy (CLR).\n    The method cycles the learning rate between two boundaries with\n    some constant frequency, as detailed in this paper (https://arxiv.org/abs/1506.01186).\n    The amplitude of the cycle can be scaled on a per-iteration or \n    per-cycle basis.\n    This class has three built-in policies, as put forth in the paper.\n    \"triangular\":\n        A basic triangular cycle w/ no amplitude scaling.\n    \"triangular2\":\n        A basic triangular cycle that scales initial amplitude by half each cycle.\n    \"exp_range\":\n        A cycle that scales initial amplitude by gamma**(cycle iterations) at each \n        cycle iteration.\n    For more detail, please see paper.\n    \n    # Example\n        ```python\n            clr = CyclicLR(base_lr=0.001, max_lr=0.006,\n                                step_size=2000., mode='triangular')\n            model.fit(X_train, Y_train, callbacks=[clr])\n        ```\n    \n    Class also supports custom scaling functions:\n        ```python\n            clr_fn = lambda x: 0.5*(1+np.sin(x*np.pi/2.))\n            clr = CyclicLR(base_lr=0.001, max_lr=0.006,\n                                step_size=2000., scale_fn=clr_fn,\n                                scale_mode='cycle')\n            model.fit(X_train, Y_train, callbacks=[clr])\n        ```    \n    # Arguments\n        base_lr: initial learning rate which is the\n            lower boundary in the cycle.\n        max_lr: upper boundary in the cycle. Functionally,\n            it defines the cycle amplitude (max_lr - base_lr).\n            The lr at any cycle is the sum of base_lr\n            and some scaling of the amplitude; therefore \n            max_lr may not actually be reached depending on\n            scaling function.\n        step_size: number of training iterations per\n            half cycle. Authors suggest setting step_size\n            2-8 x training iterations in epoch.\n        mode: one of {triangular, triangular2, exp_range}.\n            Default 'triangular'.\n            Values correspond to policies detailed above.\n            If scale_fn is not None, this argument is ignored.\n        gamma: constant in 'exp_range' scaling function:\n            gamma**(cycle iterations)\n        scale_fn: Custom scaling policy defined by a single\n            argument lambda function, where \n            0 <= scale_fn(x) <= 1 for all x >= 0.\n            mode paramater is ignored \n        scale_mode: {'cycle', 'iterations'}.\n            Defines whether scale_fn is evaluated on \n            cycle number or cycle iterations (training\n            iterations since start of cycle). Default is 'cycle'.\n    \"\"\"\n\n    def __init__(self, base_lr=0.001, max_lr=0.006, step_size=2000., mode='triangular',\n                 gamma=1., scale_fn=None, scale_mode='cycle'):\n        super(CyclicLR, self).__init__()\n\n        self.base_lr = base_lr\n        self.max_lr = max_lr\n        self.step_size = step_size\n        self.mode = mode\n        self.gamma = gamma\n        if scale_fn == None:\n            if self.mode == 'triangular':\n                self.scale_fn = lambda x: 1.\n                self.scale_mode = 'cycle'\n            elif self.mode == 'triangular2':\n                self.scale_fn = lambda x: 1/(2.**(x-1))\n                self.scale_mode = 'cycle'\n            elif self.mode == 'exp_range':\n                self.scale_fn = lambda x: gamma**(x)\n                self.scale_mode = 'iterations'\n        else:\n            self.scale_fn = scale_fn\n            self.scale_mode = scale_mode\n        self.clr_iterations = 0.\n        self.trn_iterations = 0.\n        self.history = {}\n\n        self._reset()\n\n    def _reset(self, new_base_lr=None, new_max_lr=None,\n               new_step_size=None):\n        \"\"\"Resets cycle iterations.\n        Optional boundary/step size adjustment.\n        \"\"\"\n        if new_base_lr != None:\n            self.base_lr = new_base_lr\n        if new_max_lr != None:\n            self.max_lr = new_max_lr\n        if new_step_size != None:\n            self.step_size = new_step_size\n        self.clr_iterations = 0.\n        \n    def clr(self):\n        cycle = np.floor(1+self.clr_iterations/(2*self.step_size))\n        x = np.abs(self.clr_iterations/self.step_size - 2*cycle + 1)\n        if self.scale_mode == 'cycle':\n            return self.base_lr + (self.max_lr-self.base_lr)*np.maximum(0, (1-x))*self.scale_fn(cycle)\n        else:\n            return self.base_lr + (self.max_lr-self.base_lr)*np.maximum(0, (1-x))*self.scale_fn(self.clr_iterations)\n        \n    def on_train_begin(self, logs={}):\n        logs = logs or {}\n\n        if self.clr_iterations == 0:\n            K.set_value(self.model.optimizer.lr, self.base_lr)\n        else:\n            K.set_value(self.model.optimizer.lr, self.clr())        \n            \n    def on_batch_end(self, epoch, logs=None):\n        \n        logs = logs or {}\n        self.trn_iterations += 1\n        self.clr_iterations += 1\n\n        self.history.setdefault('lr', []).append(K.get_value(self.model.optimizer.lr))\n        self.history.setdefault('iterations', []).append(self.trn_iterations)\n\n        for k, v in logs.items():\n            self.history.setdefault(k, []).append(v)\n        \n        K.set_value(self.model.optimizer.lr, self.clr())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5551f4fbd20fe8a68efc22cbdf030e7ffcd9a24a","trusted":true},"cell_type":"code","source":"#train_X, val_X, test_X, train_y, val_y, word_index = load_and_prec()\nword_index = tokenizer.word_index\nvocab = []\nfor w,k in word_index.items():\n    vocab.append(w)\n    if k >= max_features:\n        break\nembedding_matrix = load_glove(word_index)\nembedding_matrix = np.add(embedding_matrix,load_fasttext(word_index))\nembedding_matrix = np.add(embedding_matrix, load_para(word_index))\nembedding_matrix = embedding_matrix/3\n\n#embedding_matrix = np.mean([embedding_matrix_1, embedding_matrix_3], axis = 0)\nnp.shape(embedding_matrix)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7e5835588e942953819dbabe136848939dd56696","trusted":true},"cell_type":"code","source":"embedding_matrix_1 = None\nembedding_matrix_2 = None\nnp.shape(embedding_matrix)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c47de15ff76d2919cbe32d8947f122ae6a0594ba","trusted":true},"cell_type":"code","source":"'''EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nembeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))\n\n#add_lower(embeddings_index, vocab)\n\nall_embs = np.stack(embeddings_index.values())\nemb_mean,emb_std = all_embs.mean(), all_embs.std()\nembed_size = all_embs.shape[1]\n\nword_index = tokenizer.word_index\nnb_words = min(max_features, len(word_index))\nembedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\nfor word, i in word_index.items():\n    if i >= max_features: continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: embedding_matrix[i] = embedding_vector'''","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9feda29053d2a2bb5b4ca45010726bc9e301bc00","trusted":true},"cell_type":"code","source":"embedding_size = 300\n\n#model = my_model_gru_2(topic,CONTEXT_LENGTH, embedding_size, max_features, embedding_matrix)\nmodel = build_my_model(embedding_matrix)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"10f049dc1c10a00d25618252a86ac675aad1ff14","trusted":true},"cell_type":"code","source":"model.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"520cb96c89b1b69347fd7506c76bcc694a763034","trusted":true},"cell_type":"code","source":"bottleneck_layer = model.get_layer(name='bottleneck')\nmodel_extractor = Model(inputs=model.get_input_at(0), outputs= [bottleneck_layer.get_output_at(0)])\nmodel_extractor.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"328b2f7f79660eb70472283a893589f86c17fbf4"},"cell_type":"code","source":"model_extractor.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3ffdd36f4f42d6f0f3248c3209bb4c9e0350648a"},"cell_type":"markdown","source":"# Transform everything in a Numpy array"},{"metadata":{"_uuid":"f8d97efe2bef55df592b9f7d2d88ea719e299db2","trusted":true},"cell_type":"code","source":"clr = CyclicLR(base_lr=0.001, max_lr=0.002,\n               step_size=300., mode='exp_range',\n               gamma=0.99994)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ef1e1015e7c3ab5bc5d9774e49820c4b286d7847","trusted":true},"cell_type":"code","source":"## Train the model \nfrom keras.optimizers import rmsprop\nfrom keras.callbacks import EarlyStopping, ModelCheckpoint\ncheck_point = ModelCheckpoint('mymodel20.wgt', save_best_only=True)\n\n\nearly_stop = EarlyStopping(monitor='val_loss', patience=5, mode='min')\n#model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n\nmodel.fit([train_X], [train_y], batch_size=batch_size_learning,callbacks=[clr,check_point,early_stop], epochs=train_epochs, validation_data=([val_X], [val_y]))\n\n#model.fit([train_X,data_train_lsi], [train_y,train_y], batch_size=2048,callbacks=[clr,check_point,early_stop], epochs=9, validation_data=([val_X,data_val_lsi], [val_y,val_y]))\n#model.save_weights('mymodel20.wgt')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bbce36b31ba06d386e9f2e5e359f76ca72d3e253","trusted":true},"cell_type":"code","source":"model.load_weights('mymodel20.wgt')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a72ba82481de9f96c62c334cc40bd1d134d38e2d"},"cell_type":"markdown","source":"Now let us get the validation sample predictions and also get the best threshold for F1 score. "},{"metadata":{"_uuid":"996cd238b6d37bb6d6fa64edab3fcb88933561ab","trusted":true},"cell_type":"code","source":"'''\ndef threshold_search(y_true, y_proba):\n    best_threshold = 0\n    best_score = 0\n    for threshold in [i * 0.01 for i in range(100)]:\n        score = f1_score(y_true=y_true, y_pred=y_proba > threshold)\n        if score > best_score:\n            best_threshold = threshold\n            best_score = score\n    search_result = {'threshold': best_threshold, 'f1': best_score}\n    return search_result\n\ndef train_pred(model, train_X,split_data_train_lsi, train_y, X_val,split_data_val_lsi, y_val, epochs=2, callback=None):\n    model.fit([train_X,split_data_train_lsi], [train_y,train_y], batch_size=batch_size_learning,callbacks=[clr,early_stop], epochs=epochs, validation_data=([X_val,split_data_val_lsi], [y_val,y_val]))\n    #pred_val_y = model.predict([X_val,split_data_val_lsi])\n    #pred_val_y = importance*pred_val_y[0] + (1-importance)*pred_val_y[1]\n        #model.fit(train_X, train_y, batch_size=512, epochs=1, validation_data=(val_X, val_y), callbacks = callback, verbose=0)\n        #pred_val_y = model.predict([val_X], batch_size=1024, verbose=0)\n        #mythres = threshold_search(y_val,pred_val_y)\n        #thr = mythres['threshold']\n    \n    pred_val_y = model.predict([val_X,data_val_lsi], batch_size=1024, verbose=0)\n    pred_val_y = pred_val_y[0]*importance + (1-importance)*pred_val_y[1]\n    \n    mythres = threshold_search(val_y,pred_val_y)\n    thr = mythres['threshold']\n\n    best_score = metrics.f1_score(val_y, (pred_val_y > thr).astype(int))    \n    \n    print(\"SPLIT Val F1 Score: {:.4f}\".format(best_score))\n\n    pred_test_y = model.predict([test_X,data_test_lsi], batch_size=1024, verbose=0)\n    pred_test_y = pred_test_y[0]*(importance) + (1-importance)*pred_test_y[1]\n    print('=' * 60)\n    return pred_val_y, pred_test_y, best_score\n\n\n\nDATA_SPLIT_SEED = 2018\n\n\nvalid_meta = np.zeros(val_X.shape[0])\ntest_meta = np.zeros(test_X.shape[0])\nsplits = list(StratifiedKFold(n_splits=splits_kfold, shuffle=True, random_state=DATA_SPLIT_SEED).split(train_X, train_y))\n\ncount = 0\nfor idx, (train_idx, valid_idx) in enumerate(splits):\n        X_train = train_X[train_idx]\n        split_data_train_lsi = data_train_lsi[train_idx]\n        y_train = train_y[train_idx]\n        X_val = train_X[valid_idx]\n        split_data_val_lsi = data_train_lsi[valid_idx]\n        y_val = train_y[valid_idx]\n        #model = model_lstm_atten(embedding_matrix)\n        model = my_model_gru_2(topic,CONTEXT_LENGTH, embedding_size, max_features, embedding_matrix)\n        pred_val_y, pred_test_y, best_score = train_pred(model, X_train,split_data_train_lsi, y_train, X_val,split_data_val_lsi, y_val, epochs = iterations_learning, callback = [clr])\n        if best_score>0.67:\n            valid_meta += pred_val_y.reshape(-1) #/len(splits)\n            test_meta += pred_test_y.reshape(-1) #/ len(splits)\n            count = count +1\n            \nvalid_meta = valid_meta/count            \ntest_meta = test_meta/count\n\nthreshold_search(val_y,valid_meta)\n\nmythres = threshold_search(val_y,valid_meta)\nthr = mythres['threshold']\nprint('-'*60)\nprint('thr')\nprint(mythres)\n'''","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"89f93577133c92938a5f53daf949d820cce6aeeb","trusted":true},"cell_type":"code","source":"'''sub = pd.read_csv('../input/sample_submission.csv')\nsub.prediction = test_meta > thr\nsub.to_csv(\"submission.csv\", index=False)'''","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"47b63dca0247a08a808db7ae6eea33065c554948","trusted":true},"cell_type":"code","source":"thresh_saved = 0.1\nmax_f1 = 0\n\n\npred_noemb_val_y = model.predict([val_X], batch_size=1024, verbose=1)\n#pred_noemb_val_y_1 = model.predict([val_X,data_val_lsi], batch_size=1024, verbose=1)[1]\n\n#pred_noemb_val_y =pred_noemb_val_y #[0]*importance + (1-importance)*pred_noemb_val_y[1] #0.99*pred_noemb_val_y[0] + 0.01*pred_noemb_val_y[1]\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"383e51177bf33da0b8fee42dd6a093908b808f64"},"cell_type":"markdown","source":"Now let us get the test set predictions as well and save them"},{"metadata":{"_uuid":"befd6d416b934b5824f091f96658aceb21d722e0","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a88df747f43259bab84447b50e45aa9e978f2cee","trusted":true},"cell_type":"code","source":"pred_noemb_test_y = model.predict([test_X], batch_size=1024, verbose=1)\n\n#pred_noemb_test_y =  #importance*pred_noemb_test_y[0] + (1-importance)*pred_noemb_test_y[1]\n#pred_noemb_test_y = (pred_noemb_test_y>thresh_saved).astype(int)\n\n#out_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\n#out_df['prediction'] = pred_noemb_test_y\n#out_df.to_csv(\"submission.csv\", index=False)\n#'''","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3edf798761fe334e4c9bf76d40eb93632e57b740"},"cell_type":"code","source":"embedding_matrix = None\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aed37970454b52aa7b58227336a783174d652dce"},"cell_type":"markdown","source":"try:\n    del model, data_test_lsi, data_train_lsi, data_val_lsi, val_tf, test_tf, lsi\n    import gc; gc.collect()\n    time.sleep(10)\nexcept:\n    import gc; gc.collect()\n    time.sleep(10)\n    pass"},{"metadata":{"_uuid":"3f00d85bf15c9110e48b57702ed301cc019ee22e","trusted":true},"cell_type":"code","source":"from imblearn.ensemble import BalancedRandomForestClassifier\nfrom sklearn.preprocessing import StandardScaler\n\nnew_train_X = model_extractor.predict([train_X], batch_size=1024, verbose=1)\nnew_val_X = model_extractor.predict([val_X], batch_size=1024, verbose=1)\nnew_test_X= model_extractor.predict([test_X], batch_size=1024, verbose=1)\nscl = StandardScaler()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3f00d85bf15c9110e48b57702ed301cc019ee22e","trusted":true},"cell_type":"code","source":"new_train_X_scaled = scl.fit_transform(new_train_X)\nnew_val_X_scaled = scl.transform(new_val_X)\nnew_test_X_scaled = scl.transform(new_test_X)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4d293579a292e99ff1229f1d00d5ec7d998d217f"},"cell_type":"code","source":"print(np.shape(new_train_X))\nprint(np.shape(new_train_X_scaled))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"99c4fe4d6be4c04f3089d3b7a5fb12bd7a357945"},"cell_type":"code","source":"print(len(val_y))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fb15d229e0978929579633cf3e35fbb8478ea306"},"cell_type":"code","source":"from imblearn.over_sampling import SMOTE\nfrom sklearn.model_selection import train_test_split\nfrom random import choice, sample\n\n#X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)\nneg_sample = 500000\npos_sample = 40000\n\nX_smote = None\n\nX_smote_neg = np.array(sample(new_train_X_scaled[train_y==0,:].tolist(),neg_sample))\nX_smote_pos = np.array(sample(new_train_X_scaled[train_y==1,:].tolist(),pos_sample))\n\ny_neg = np.zeros(X_smote_neg.shape[0])\ny_pos = np.zeros(X_smote_pos.shape[0])+1\n\nto_smote = np.append(X_smote_neg, X_smote_pos,axis=0)\nto_smote_y = np.append(y_neg, y_pos)\n\nprint(np.shape(to_smote),np.shape(to_smote_y))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"00e4fa178fc5fb943d44342e074c66457bd5626b"},"cell_type":"code","source":"sm = SMOTE(random_state=2)\nX_train_res, y_train_res = sm.fit_sample(to_smote, to_smote_y)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"664d5feef827ee22dfd9d13dd4ef4a2afdd3351d"},"cell_type":"code","source":"#pos_sample = 40000\n\nnew_x_pos = X_train_res[y_train_res==1,:][pos_sample:]\nnew_y_pos = y_train_res[y_train_res==1][pos_sample:]\n\nnew_train_X_scaled_smoted=np.append(new_train_X_scaled,new_x_pos,axis=0)\ntrain_y_smoted = np.append(train_y,new_y_pos)\n\nprint(np.shape(new_train_X_scaled_smoted))\n\nprint(np.shape(train_y_smoted))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"af4f53d7b79b2eabd20dbe47a7d617ef9677147e","trusted":true},"cell_type":"code","source":"\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4bd9668ef552731a13dce48b64da1dc08f44d29a"},"cell_type":"code","source":"#clr2 = CyclicLR(base_lr=0.001, max_lr=0.002,\n#               step_size=300., mode='exp_range',\n#               gamma=0.99994)\n\n#mymod= my_small_model(64)\n#mymod.fit(new_train_X_scaled_smoted,train_y_smoted, callbacks=[clr2,early_stop], epochs=20, batch_size=batch_size_learning, validation_data=([new_val_X_scaled], [val_y]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"272dc3b8c7cc3aa334bf37e62ccfba35f3b2360e"},"cell_type":"code","source":"#predicted_val_smoted_y = mymod.predict(new_val_X_scaled)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5f9130f1d6eff4b0162aee8e32256841f36cb5c5"},"cell_type":"code","source":"#pred_test_smoted_y = mymod.predict(new_test_X_scaled) # ensemble.predict_proba(new_test_X_scaled)[:,1]*0.5 + pred_noemb_test_y*0.5 #svmoc.predict_proba(new_test_X_scaled)[:,1]\n#pred_extract_test_y = (pred_extract_test_y>thresh_saved2).astype(int)\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ea56406dab39195e74aad287b0534ec49a5e9df5"},"cell_type":"code","source":"check_point2 = ModelCheckpoint('intermediate.wgt', save_best_only=True)\n\n\n\n\ndef threshold_search(y_true, y_proba):\n    best_threshold = 0\n    best_score = 0\n    for threshold in [i * 0.01 for i in range(100)]:\n        score = f1_score(y_true=y_true, y_pred=y_proba > threshold)\n        if score > best_score:\n            best_threshold = threshold\n            best_score = score\n    search_result = {'threshold': best_threshold, 'f1': best_score}\n    return search_result\n\ndef train_pred(model, train_X, train_y, X_val, y_val, epochs=2, callback=None):\n    model.fit([train_X], [train_y], batch_size=batch_size_learning,callbacks=[clr,early_stop,check_point2], epochs=epochs, validation_data=([X_val], [y_val]))\n    model.load_weights('intermediate.wgt')\n    pred_val_y = model.predict([new_val_X_scaled], batch_size=1024, verbose=0)\n    \n    mythres = threshold_search(val_y,pred_val_y)\n    thr = mythres['threshold']\n\n    best_score = metrics.f1_score(val_y, (pred_val_y > thr).astype(int))    \n    \n    print(\"SPLIT Val F1 Score: {:.4f}\".format(best_score))\n\n    pred_test_y = model.predict([new_test_X_scaled], batch_size=1024, verbose=0)\n    #pred_test_y = pred_test_y #[0]*(importance) + (1-importance)*pred_test_y[1]\n    print('=' * 60)\n    return pred_val_y, pred_test_y, best_score\n\n\n\nDATA_SPLIT_SEED = 2018\n\n\nvalid_meta = np.zeros(val_X.shape[0])\ntest_meta = np.zeros(test_X.shape[0])\nsplits = list(StratifiedKFold(n_splits=splits_kfold, shuffle=True, random_state=DATA_SPLIT_SEED).split(new_train_X_scaled_smoted, train_y_smoted))\n\ncount = 0\nfor idx, (train_idx, valid_idx) in enumerate(splits):\n        X_train = new_train_X_scaled_smoted[train_idx]\n        y_train = train_y_smoted[train_idx]\n        X_val = new_train_X_scaled_smoted[valid_idx]\n        y_val = train_y_smoted[valid_idx]\n        #model = model_lstm_atten(embedding_matrix)\n        model = my_small_model(64)\n        pred_val_y, pred_test_y, best_score = train_pred(model, X_train, y_train, X_val, y_val, epochs = iterations_learning, callback = [clr])\n        #if best_score>0.675:\n        valid_meta += pred_val_y.reshape(-1) #/len(splits)\n        test_meta += pred_test_y.reshape(-1) #/ len(splits)\n        count = count +1\n            \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d0d91a4699ecfc825cc92585067e7e72413699f2"},"cell_type":"code","source":"valid_meta = valid_meta/count            \ntest_meta = test_meta/count\n\nthreshold_search(val_y,valid_meta)\n\nmythres = threshold_search(val_y,valid_meta)\nthr = mythres['threshold']\nprint('-'*60)\nprint('thr')\nprint(mythres)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6a0f99c737ed386bedd67d42e7b7ba426c936a82"},"cell_type":"code","source":"predict_final = (test_meta>thr).astype(int)\n\nout_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\nout_df['prediction'] = predict_final\nout_df.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0b3016d14476ea0cffe5dcfc190faca3858a386f"},"cell_type":"markdown","source":"from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import VotingClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.discriminant_analysis import QuadraticDiscriminantAnalysis\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\nfrom imblearn.ensemble import BalancedRandomForestClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom imblearn.over_sampling import SMOTE\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.naive_bayes import MultinomialNB\nfrom sklearn.discriminant_analysis import QuadraticDiscriminantAnalysis\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\nfrom imblearn.ensemble import BalancedRandomForestClassifier\n#svmoc.fit(new_train_X,train_y)\nfrom sklearn.ensemble import VotingClassifier\nfrom sklearn.ensemble import ExtraTreesClassifier\n\nestimators = []\n\nmodel1 = BalancedRandomForestClassifier(n_estimators=70, class_weight='balanced') \nmodel2 = ExtraTreesClassifier(n_estimators=70, class_weight='balanced')\nmodel3 = LinearDiscriminantAnalysis()\n#model4 = QuadraticDiscriminantAnalysis()\n#model5 = GaussianNB()\nmodel6 = LogisticRegression()\n\n\nestimators.append(('tree', model1))\nestimators.append(('cart', model2))\nestimators.append(('discriminant', model3))\n#estimators.append(('discriminant1', model4))\n#estimators.append(('gb', model5))\nestimators.append(('lr', model6))\n\n\nensemble = VotingClassifier(estimators,voting='soft')\nensemble.fit(new_train_X_scaled_smoted,train_y_smoted)\n\n\n"},{"metadata":{"trusted":true,"_uuid":"330279f67d770ab17f27e9f139393a30ce16a80c"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f8290c6090a7de381dfbbe15be4cbc3e05bdce2"},"cell_type":"markdown","source":"\n\npredicted_val_ensemble_y = ensemble.predict_proba(new_val_X_scaled)[:,1]\n\nfor i in range(0,len(predicted_val_ensemble_y)):\n    predicted_val_ensemble_y[i] = 0.33*predicted_val_smoted_y[i] + 0.33*pred_noemb_val_y[i] + 0.34*predicted_val_ensemble_y[i]\n\n\nthresh_saved2 = 0.1\nmax_f1 = 0\n\nfor thresh2 in np.arange(0.01, 1, 0.01):\n    thresh2 = np.round(thresh2, 2)\n    value = metrics.f1_score(val_y, (predicted_val_ensemble_y>thresh2).astype(int))\n    print(\"F1 score at threshold {0} is {1}\".format(thresh2, value))\n    if value> max_f1:\n                max_f1 = value\n                thresh_saved2=thresh2\nprint(thresh_saved2)    \n"},{"metadata":{"trusted":true,"_uuid":"1529d3ea422417c1cda89ef19dfdbae09762cdc4"},"cell_type":"markdown","source":"predict_test_ensemble_y = ensemble.predict_proba(new_test_X_scaled)[:,1]\n\nfor i in range(0,len(predict_test_ensemble_y)):\n    predict_test_ensemble_y[i]= 0.33*pred_test_smoted_y[i] + 0.33*pred_noemb_test_y[i] + 0.34*predict_test_ensemble_y[i] \n\n\npredict_final =   predict_test_ensemble_y\n\npredict_final = (predict_final>thresh_saved2).astype(int)\n\nout_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\nout_df['prediction'] = predict_final\nout_df.to_csv(\"submission.csv\", index=False)"},{"metadata":{"trusted":true,"_uuid":"f53e3f30c5f9ef8fb5038cbb7f05bb088e1efda1"},"cell_type":"code","source":"'''import numpy as np\nfrom sklearn.decomposition import PCA\n\n\npca = PCA(n_components=2)\nX_proj = pca.fit_transform(new_train_X_scaled_smoted)\nX_test_proj = pca.transform(new_val_X_scaled)\n\nimport matplotlib.pyplot as plt\n\n\nplt.plot(X_proj[train_y_smoted==0,0],X_proj[train_y_smoted==0,1],'*r')\n\nplt.plot(X_proj[train_y_smoted==1,0],X_proj[train_y_smoted==1,1],'*b')\n\nplt.plot(X_test_proj[val_y==0,0],X_test_proj[val_y==0,1],'*g')\n\nplt.plot(X_test_proj[val_y==1,0],X_test_proj[val_y==1,1],'*c')\n\n\nplt.show()\n'''","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8e8733d3ce71d83e8eb7d29e136f669d3ac21743","trusted":true},"cell_type":"code","source":"'''estimators = 500\nthresh_saved2 = 0.1\nmax_f1 = 0\n\n#balancing classifier\nsvmoc = BalancedRandomForestClassifier(n_estimators=estimators, class_weight='balanced') \n#svmoc.fit(new_train_X,train_y)'''","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d3db264dd8777ef833bb3d31851be5ba6d8f1466","trusted":true},"cell_type":"code","source":"'''\nimport pandas\nfrom sklearn import model_selection\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.svm import SVC\nfrom sklearn.ensemble import VotingClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\n\nestimators = []\nmodel1 = LogisticRegression()\nestimators.append(('logistic', model1))\nmodel2 = DecisionTreeClassifier()\nestimators.append(('cart', model2))\nmodel3 = SVC()\nestimators.append(('svm', model3))\n\nmodel4 = KNeighborsClassifier(n_neighbors=3)\n\nestimators.append(('knn', model4))\n\nmodel5 = svmoc\n\nestimators.append(('btree', model5))\n\n\n# create the ensemble model\nensemble = VotingClassifier(estimators)\nensemble.fit(new_train_X,train_y)\n'''\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3f00d85bf15c9110e48b57702ed301cc019ee22e","trusted":true},"cell_type":"code","source":"\n\n'''\npredicted_val_y = ensemble.predict_proba(new_val_X_scaled)[:,1]\n\n#predicted_val_y = 0.5*predicted_val_y + 0.5*pred_noemb_test_y\n\nthresh_saved2 = 0.1\nfor thresh2 in np.arange(0.1, 1, 0.01):\n    thresh2 = np.round(thresh2, 2)\n    value = metrics.f1_score(val_y, (predicted_val_y>thresh2).astype(int))\n    print(\"F1 score at threshold {0} is {1}\".format(thresh2, value))\n    if value> max_f1:\n                max_f1 = value\n                thresh_saved2=thresh2\nprint(thresh_saved2)    '''","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"309da562f38f57b35e38f16afc34b0be736f83fd","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7bd3ca4ff807a5a48fa1253ac015dccd930c821c"},"cell_type":"markdown","source":"# Embeddings"},{"metadata":{"_uuid":"c0c11cebf36eeff44776ae0743c7e58adbbeeb3c"},"cell_type":"markdown","source":"'''\nimport numpy as np\n\ndef load_embed(file):\n    def get_coefs(word,*arr): \n        return word, np.asarray(arr, dtype='float32')\n    \n    if file == '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec':\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(file) if len(o)>100)\n    else:\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(file, encoding='latin'))\n        \n    return embeddings_index\n\nglove = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\nparagram =  '../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt'\nwiki_news = '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec'\n\nprint(\"Extracting GloVe embedding\")\nembed_glove = load_embed(glove)\nprint(\"Extracting Paragram embedding\")\nembed_paragram = load_embed(paragram)\nprint(\"Extracting FastText embedding\")\nembed_fasttext = load_embed(wiki_news)\n'''\n\n\n\n"},{"metadata":{"_uuid":"9ecf4a039dd66f79478c85329f8ffe665a654fd9"},"cell_type":"markdown","source":"import operator \ndef build_vocab(texts):\n    sentences = texts.apply(lambda x: x.split()).values\n    vocab = {}\n    for sentence in sentences:\n        for word in sentence:\n            try:\n                vocab[word] += 1\n            except KeyError:\n                vocab[word] = 1\n    return vocab\n\ndef check_coverage(vocab, embeddings_index):\n    known_words = {}\n    unknown_words = {}\n    nb_known_words = 0\n    nb_unknown_words = 0\n    for word in vocab.keys():\n        try:\n            known_words[word] = embeddings_index[word]\n            nb_known_words += vocab[word]\n        except:\n            unknown_words[word] = vocab[word]\n            nb_unknown_words += vocab[word]\n            pass\n\n    print('Found embeddings for {:.2%} of vocab'.format(len(known_words) / len(vocab)))\n    print('Found embeddings for  {:.2%} of all text'.format(nb_known_words / (nb_known_words + nb_unknown_words)))\n    unknown_words = sorted(unknown_words.items(), key=operator.itemgetter(1))[::-1]\n\n    return unknown_words"},{"metadata":{"_uuid":"8f8c455f17c6cb65cfeb402d92809cda03d33998"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"b6b34fa244849de0bc5fdcf68c5e73ff746c8892"},"cell_type":"markdown","source":"# Coverage starting point"},{"metadata":{"_uuid":"43ac2fbb1128b6e2fe287c6cf79c8f194a2b00df"},"cell_type":"markdown","source":"del train_df, test_df, val_df, train_X, val_X, test_X #delete everything, start from scratch now.\nimport gc; gc.collect()\ntime.sleep(10)"},{"metadata":{"_uuid":"e2ec8170649314f97718fee08b2bbb8e1afd2294"},"cell_type":"markdown","source":"import pandas as pd\n\ntrain_frame = pd.read_csv(\"../input/train.csv\")\ntest_frame = pd.read_csv(\"../input/test.csv\")\n\ndf = train_frame.append(test_frame)\n\nvocab = build_vocab(df['question_text'])"},{"metadata":{"_uuid":"1574736c763f6718eadaf258c28f238101d6dadc","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"794aeed983cb7775fca364b9288a530769c71227","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d07e734b5fefaa25437a86e1abf387296f00175c","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"64349ab6b44e7308d18d96593501cca3150d39a0","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a0f51d16b5953035f730419347a9c65af7eea48f","trusted":true},"cell_type":"code","source":"#print(\"F1 score at threshold {0} is {1}\".format(thresh, metrics.f1_score(test_Y, (pred_noemb_test_y>0.28).astype(int))))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a36a071fb50f6c120e099b5fe27ad6ac977f1125"},"cell_type":"markdown","source":"## split to train and val\ntrain_df, val_df = train_test_split(train_frame, test_size=0.1, random_state=2018)\n\n## some config values \nembed_size = 300 # how big is each word vector\nmax_features = 95000 # how many unique words to use (i.e num rows in embedding vector)\nmaxlen = 100 # max number of words in a question to use\n\ntransform_df(train_df)\ntransform_df(val_df)\ntransform_df(test_frame)\n\n## fill up the missing values\ntrain_X = train_df[\"treated_question\"].fillna(\"_na_\").values\nval_X = val_df[\"treated_question\"].fillna(\"_na_\").values\ntest_X = test_frame[\"treated_question\"].fillna(\"_na_\").values\n\n## Tokenize the sentences\ntokenizer = Tokenizer(num_words=max_features)\ntokenizer.fit_on_texts(list(train_X))\ntrain_X = tokenizer.texts_to_sequences(train_X)\nval_X = tokenizer.texts_to_sequences(val_X)\ntest_X = tokenizer.texts_to_sequences(test_X)\n\n## Pad the sentences \ntrain_X = pad_sequences(train_X, maxlen=maxlen)\nval_X = pad_sequences(val_X, maxlen=maxlen)\ntest_X = pad_sequences(test_X, maxlen=maxlen)\n\n## Get the target values\ntrain_y = train_df['target'].values\nval_y = val_df['target'].values"},{"metadata":{"_uuid":"b9d263852f653e466e24f9827548d7d1a7ee7262","trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7c0010e518288bc7f588776c58610949140a139a"},"cell_type":"markdown","source":"We have four different types of embeddings.\n * GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n * glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n * paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n * wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html\n \n A very good explanation for different types of embeddings are given in this [kernel](https://www.kaggle.com/sbongo/do-pretrained-embeddings-give-you-the-extra-edge). Please refer the same for more details..\n\n**Glove Embeddings:**\n\nIn this section, let us use the Glove embeddings and rebuild the GRU model."},{"metadata":{"_uuid":"23f130e80159bb1701e449e2e91199dbfff1f1d4"},"cell_type":"markdown","source":"EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nembeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))\n\nadd_lower(embeddings_index, vocab)\n\nall_embs = np.stack(embeddings_index.values())\nemb_mean,emb_std = all_embs.mean(), all_embs.std()\nembed_size = all_embs.shape[1]\n\nword_index = tokenizer.word_index\nnb_words = min(max_features, len(word_index))\nembedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\nfor word, i in word_index.items():\n    if i >= max_features: continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n        \ninp = Input(shape=(maxlen,))\nx = Embedding(max_features, embed_size, weights=[embedding_matrix])(inp)\nx = Bidirectional(CuDNNGRU(64, return_sequences=True))(x)\nx = GlobalMaxPool1D()(x)\nx = Dense(16, activation=\"relu\")(x)\nx = Dropout(0.1)(x)\nx = Dense(1, activation=\"sigmoid\")(x)\nmodel = Model(inputs=inp, outputs=x)\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\nprint(model.summary()) "},{"metadata":{"_uuid":"a560ab0dbab9cf6fdbdae6721ec030e300f19d78"},"cell_type":"markdown","source":"model.fit(train_X, train_y, batch_size=512, callbacks=[clr], epochs=2) #validation_data=(val_X, val_y))"},{"metadata":{"_uuid":"ff43855164472de035a5a1d80b3db4838684701a"},"cell_type":"markdown","source":"thresh_saved = 0.1\nmax_f1 = 0\n\n\npred_glove_val_y = model.predict([val_X], batch_size=1024, verbose=1)\nfor thresh in np.arange(0.1, 0.501, 0.01):\n    thresh = np.round(thresh, 2)\n    value = metrics.f1_score(val_y, (pred_glove_val_y>thresh).astype(int))\n    print(\"F1 score at threshold {0} is {1}\".format(thresh, value))\n    if value> max_f1:\n                max_f1 = value\n                thresh_saved=thresh\nprint(thresh_saved)    \n"},{"metadata":{"_uuid":"d2a33c252f31fddcc65896053184226128562776"},"cell_type":"markdown","source":"Results seem to be better than the model without pretrained embeddings."},{"metadata":{"_uuid":"d51ff8ed6a87b488fec3ac84ca50df661d7c8193"},"cell_type":"markdown","source":"pred_glove_test_y = model.predict([test_X], batch_size=1024, verbose=1)\npred_glove_test_y = (pred_glove_test_y>thresh_saved).astype(int)\n"},{"metadata":{"_uuid":"39d4fedab4ac170863a0ee1ca3aa9be1ee58fe02"},"cell_type":"markdown","source":"del word_index, embeddings_index, all_embs, embedding_matrix, model, inp, x\nimport gc; gc.collect()\ntime.sleep(10)"},{"metadata":{"_uuid":"bc6bab22dd12a09378f4b8b159cb7a5d88a3e7c0"},"cell_type":"markdown","source":"**Wiki News FastText Embeddings:**\n\nNow let us use the FastText embeddings trained on Wiki News corpus in place of Glove embeddings and rebuild the model."},{"metadata":{"_uuid":"6f3d0fd28dd2b04eaccb732b96b872e5a223d962"},"cell_type":"markdown","source":"EMBEDDING_FILE = '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec'\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nembeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE) if len(o)>100)\n\nadd_lower(embeddings_index, vocab)\n\nall_embs = np.stack(embeddings_index.values())\nemb_mean,emb_std = all_embs.mean(), all_embs.std()\nembed_size = all_embs.shape[1]\n\nword_index = tokenizer.word_index\nnb_words = min(max_features, len(word_index))\nembedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\nfor word, i in word_index.items():\n    if i >= max_features: continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n        \ninp = Input(shape=(maxlen,))\nx = Embedding(max_features, embed_size, weights=[embedding_matrix])(inp)\nx = Bidirectional(CuDNNGRU(64, return_sequences=True))(x)\nx = GlobalMaxPool1D()(x)\nx = Dense(16, activation=\"relu\")(x)\nx = Dropout(0.1)(x)\nx = Dense(1, activation=\"sigmoid\")(x)\nmodel = Model(inputs=inp, outputs=x)\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])"},{"metadata":{"_uuid":"47238831a4701c8a67dc7ecb130ac1402baf7bb2"},"cell_type":"markdown","source":"model.fit(train_X, train_y, batch_size=512, epochs=2, callbacks=[clr], validation_data=(val_X, val_y))"},{"metadata":{"_uuid":"b7ab4100f723ad535528865b1edc7896bce80223"},"cell_type":"markdown","source":"thresh_saved = 0.1\nmax_f1 = 0\n\n\npred_fasttext_val_y = model.predict([val_X], batch_size=1024, verbose=1)\nfor thresh in np.arange(0.1, 0.501, 0.01):\n    thresh = np.round(thresh, 2)\n    value = metrics.f1_score(val_y, (pred_fasttext_val_y>thresh).astype(int))\n    print(\"F1 score at threshold {0} is {1}\".format(thresh, value))\n    if value> max_f1:\n                max_f1 = value\n                thresh_saved=thresh\nprint(thresh_saved)    "},{"metadata":{"_uuid":"3216362afb0f49579d287a06f13adf8cd7d8b0cf"},"cell_type":"markdown","source":"pred_fasttext_test_y = model.predict([test_X], batch_size=1024, verbose=1)\npred_fasttext_test_y = (pred_fasttext_test_y>thresh_saved).astype(int)\n"},{"metadata":{"_uuid":"f24f9753ff1d933fa4f75a0ba34df305632d6e93"},"cell_type":"markdown","source":"del word_index, embeddings_index, all_embs, embedding_matrix, model, inp, x\nimport gc; gc.collect()\ntime.sleep(10)"},{"metadata":{"_uuid":"4ca44ac68bf404b9c26e07fbcc9c8ac793e04510"},"cell_type":"markdown","source":"**Paragram Embeddings:**\n\nIn this section, we can use the paragram embeddings and build the model and make predictions."},{"metadata":{"_uuid":"25ec1aac4aedbf431a2d30de64030ce8e3203c18"},"cell_type":"markdown","source":"EMBEDDING_FILE = '../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt'\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nembeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore') if len(o)>100)\nadd_lower(embeddings_index, vocab)\n\nall_embs = np.stack(embeddings_index.values())\nemb_mean,emb_std = all_embs.mean(), all_embs.std()\nembed_size = all_embs.shape[1]\n\nword_index = tokenizer.word_index\nnb_words = min(max_features, len(word_index))\nembedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\nfor word, i in word_index.items():\n    if i >= max_features: continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n        \ninp = Input(shape=(maxlen,))\nx = Embedding(max_features, embed_size, weights=[embedding_matrix])(inp)\nx = Bidirectional(CuDNNGRU(64, return_sequences=True))(x)\nx = GlobalMaxPool1D()(x)\nx = Dense(16, activation=\"relu\")(x)\nx = Dropout(0.1)(x)\nx = Dense(1, activation=\"sigmoid\")(x)\nmodel = Model(inputs=inp, outputs=x)\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])"},{"metadata":{"_uuid":"cc188f2787ea7b98d3a40953a95a5fc09ff2764d"},"cell_type":"markdown","source":"model.fit(train_X, train_y, batch_size=512, epochs=2,callbacks = [clr]) # validation_data=(val_X, val_y))"},{"metadata":{"_uuid":"9abdfd1cf15257f2c0c2181a13327796e8d4584e"},"cell_type":"markdown","source":"thresh_saved = 0.1\nmax_f1 = 0\n\n\npred_paragram_val_y = model.predict([val_X], batch_size=1024, verbose=1)\nfor thresh in np.arange(0.1, 0.501, 0.01):\n    thresh = np.round(thresh, 2)\n    value = metrics.f1_score(val_y, (pred_paragram_val_y>thresh).astype(int))\n    print(\"F1 score at threshold {0} is {1}\".format(thresh, value))\n    if value> max_f1:\n                max_f1 = value\n                thresh_saved=thresh\nprint(thresh_saved)    "},{"metadata":{"_uuid":"99cb9f6145da909bd7436e46d47547efc097499d"},"cell_type":"markdown","source":"pred_paragram_test_y = model.predict([test_X], batch_size=1024, verbose=1)\npred_paragram_test_y = (pred_paragram_test_y>thresh_saved).astype(int)\n"},{"metadata":{"_uuid":"af087d21bdb4358701e31aded6b522accd5a8a64"},"cell_type":"markdown","source":"del word_index, embeddings_index, all_embs, embedding_matrix, model, inp, x\nimport gc; gc.collect()\ntime.sleep(10)"},{"metadata":{"_uuid":"e1312b7a4c3b67ca4ebd26fb083dbac3b6635dc2"},"cell_type":"markdown","source":"**Observations:**\n * Overall pretrained embeddings seem to give better results comapred to non-pretrained model. \n * The performance of the different pretrained embeddings are almost similar.\n \n**Final Blend:**\n\nThough the results of the models with different pre-trained embeddings are similar, there is a good chance that they might capture different type of information from the data. So let us do a blend of these three models by averaging their predictions."},{"metadata":{"_uuid":"449bc59fdc9a719aa0759ac51a4481df113604ca"},"cell_type":"markdown","source":"thresh_saved = 0.1\nmax_f1 = 0\n\npred_val_y = 0.25*pred_glove_val_y + 0.25*pred_fasttext_val_y + 0.25*pred_paragram_val_y +0.25*pred_noemb_val_y\n\nfor thresh in np.arange(0.1, 0.501, 0.01):\n    thresh = np.round(thresh, 2)\n    value = metrics.f1_score(val_y, (pred_val_y>thresh).astype(int))\n    print(\"F1 score at threshold {0} is {1}\".format(thresh, value))\n    if value> max_f1:\n                max_f1 = value\n                thresh_saved=thresh\nprint(thresh_saved)    \n\n"},{"metadata":{"_uuid":"4fdbeffc0f84643d2832eec49234bd9d6c6e216b"},"cell_type":"markdown","source":"The result seems to better than individual pre-trained models and so we let us create a submission file using this model blend."},{"metadata":{"_uuid":"c90fb4a4ef1b3b2ea06563a6901deac1b38822f3"},"cell_type":"markdown","source":"pred_test_y = 0.25*pred_glove_test_y + 0.25*pred_fasttext_test_y + 0.25*pred_paragram_test_y +0.25*pred_noemb_test_y\npred_test_y = (pred_test_y>thresh_saved).astype(int)\nout_df = pd.DataFrame({\"qid\":test_frame[\"qid\"].values})\nout_df['prediction'] = pred_test_y\nout_df.to_csv(\"submission.csv\", index=False)"},{"metadata":{"_uuid":"f6797ab73bdd5bdb8c8f6d80ec361c50a2b0f56f"},"cell_type":"markdown","source":"\n**References:**\n\nThanks to the below kernels which helped me with this one. \n1. https://www.kaggle.com/jhoward/improved-lstm-baseline-glove-dropout\n2. https://www.kaggle.com/sbongo/do-pretrained-embeddings-give-you-the-extra-edge"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}