{"cells":[{"metadata":{"_uuid":"b67e44b35341421fd12f2c1652530cf791746fb6"},"cell_type":"markdown","source":"In this competition, no correlations between cross validation score and leaderboard score has bothered many participators a long time. The goal of this notebook is to mimic the leaderboard. "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np, pandas as pd, random as rn, os, gc, re, time\nstart = time.time()\nseed = 32\nos.environ['PYTHONHASHSEED'] = str(seed)\nos.environ['OMP_NUM_THREADS'] = '4'\nnp.random.seed(seed)\nrn.seed(seed)\nimport tensorflow as tf\nsession_conf = tf.ConfigProto(intra_op_parallelism_threads = 1,\n                              inter_op_parallelism_threads = 1)\ntf.set_random_seed(seed)\nsess = tf.Session(graph = tf.get_default_graph(), config = session_conf)\nfrom keras import backend as K\nK.set_session(sess)\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.metrics import f1_score, precision_recall_curve, confusion_matrix\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import StratifiedKFold\n\nfrom keras.layers import Input, Dense, CuDNNLSTM, Bidirectional, Activation, Conv1D\nfrom keras.layers import Dropout, Embedding, GlobalMaxPooling1D, MaxPooling1D\nfrom keras.layers import Add, Flatten, BatchNormalization, GlobalAveragePooling1D\nfrom keras.layers import concatenate, SpatialDropout1D, CuDNNGRU, Lambda, GaussianDropout\nfrom keras.layers import PReLU, ReLU, ELU\nfrom keras import initializers, regularizers, constraints, optimizers, layers, callbacks\nfrom keras.callbacks import EarlyStopping, ModelCheckpoint, Callback\nfrom keras.initializers import he_normal, he_uniform,  glorot_normal\nfrom keras.initializers import glorot_uniform, zeros, orthogonal\nfrom keras.models import Model, load_model\nfrom keras.optimizers import Adam, RMSprop, SGD\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.callbacks import EarlyStopping, ModelCheckpoint\n\ntrain = pd.read_csv(\"../input/train.csv\").fillna(\"missing\")\ntest = pd.read_csv(\"../input/test.csv\").fillna(\"missing\")\n\nembedding_file1 = \"../input/embeddings/glove.840B.300d/glove.840B.300d.txt\"\nembedding_file2 = \"../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt\"\n\nembed_size = 300\nmax_features = 100000\nmax_len = 60","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dabe864fcee39494308a85a820b09a6d903d5811"},"cell_type":"code","source":"puncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', \n          '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n          '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  \n          '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', '“', '★', '”', \n          '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', \n          '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', '▒', '：', '¼', '⊕', '▼', \n          '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', \n          'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', '∙', '）', '↓', '、', '│', '（', '»', \n          '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', \n          '¹', '≤', '‡', '√', 'β', 'α', '∅', 'θ', '÷', '₹']\n\ndef clean_punct(x):\n    x = str(x)\n    for punct in puncts:\n        if punct in x:\n            x = x.replace(punct, f' {punct} ')\n    return x\n\ntrain[\"question_text\"] = train[\"question_text\"].apply(lambda x: clean_punct(x))\ntest[\"question_text\"] = test[\"question_text\"].apply(lambda x: clean_punct(x))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8519bf4ef09715d88b0359a9ed62f2419cef67a6"},"cell_type":"code","source":"test_shape = test.shape\nprint(test_shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"422348ab223831e078ed09e3912026b8e4656b01"},"cell_type":"code","source":"sincere = train[train[\"target\"] == 0]\ninsincere = train[train[\"target\"] == 1]\n\nprint(\"Sincere questions {}; Insincere questions {}\".format(sincere.shape[0], insincere.shape[0]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"131358f5c711f497fe4fdb0bb8fb7f85983994a5"},"cell_type":"markdown","source":"### Create fake testing data from given training data\n\nBased on this [kernel](https://www.kaggle.com/arthurtok/target-visualization-t-sne-and-doc2vec), we know that the distribution of given training data and testing data are similar. We assume that the given testing data has about $95\\%$ sincere questions $(53552)$ and $5\\%$ insincere questions $(2818)$."},{"metadata":{"trusted":true,"_uuid":"20d844ec3326264b7eff3045b290b4aafea71239"},"cell_type":"code","source":"temp1 = sincere.sample(53552)\ntemp2 = insincere.sample(2818)\nfake_test = pd.concat([temp1, temp2], sort = True).reset_index()\ntrain = train.drop(fake_test.index).reset_index()\ntarget = train[\"target\"].values\n\nprint(\"Fake test data shape {}\".format(fake_test.shape))\nprint(\"New train data shape {}\".format(train.shape))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7e3e69dbd98911ad572a53f975e4d59b75036d9b"},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cf7fb3e621ac5bdbcf3d9c57acb0b0703782c97f"},"cell_type":"code","source":"fake_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"19c6fe5b2e3f2336df9103a7514d513d10623e25"},"cell_type":"code","source":"def get_glove(embedding_file):\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(embedding_file))\n    \n    all_embs = np.stack(embeddings_index.values())\n    emb_mean, emb_std = all_embs.mean(), all_embs.std()\n    return embeddings_index, emb_mean, emb_std\n\ndef get_para(embedding_file):\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(embedding_file, \n                                                                   encoding=\"utf8\", \n                                                                   errors='ignore') if len(o)>100)\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean, emb_std = all_embs.mean(), all_embs.std()\n    return embeddings_index, emb_mean, emb_std\n\nglove_index, glove_mean, glove_std = get_glove(embedding_file1)\npara_index, para_mean, para_std = get_para(embedding_file2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"69a863735d35a7311f09e7d1bd02b36c9a233336"},"cell_type":"code","source":"def get_embed(tokenizer = None, embeddings_index = None, emb_mean = None, emb_std = None):\n    word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return nb_words, embedding_matrix","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"tokenizer = Tokenizer(num_words = max_features, lower = True)\ntokenizer.fit_on_texts(train[\"question_text\"])\n\ntrain_token = tokenizer.texts_to_sequences(train[\"question_text\"])\nfake_test_token = tokenizer.texts_to_sequences(fake_test[\"question_text\"])\ntest_token = tokenizer.texts_to_sequences(test[\"question_text\"])\n\ntrain_seq = pad_sequences(train_token, maxlen = max_len)\nfake_test_seq = pad_sequences(fake_test_token, maxlen = max_len)\nX_test = pad_sequences(test_token, maxlen = max_len)\ndel train_token, fake_test_token, test_token; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9e5f431dd643f660f246d5ef6afbba1ec2358838"},"cell_type":"code","source":"nb_words, embedding_matrix1 = get_embed(tokenizer = tokenizer, embeddings_index = glove_index, \n                                        emb_mean = glove_mean, \n                                        emb_std = glove_std)\nnb_words, embedding_matrix2 = get_embed(tokenizer = tokenizer, embeddings_index = para_index, \n                                        emb_mean = para_mean, \n                                        emb_std = para_std)\nembedding_matrix = np.mean([embedding_matrix1, embedding_matrix2], axis = 0)\ndel embedding_matrix1, embedding_matrix2; gc.collect()\nprint(\"Embedding matrix completed!\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1aa72211c42dca582caaa24145d9a57c96b2a935"},"cell_type":"code","source":"from keras.engine import Layer, InputSpec\nfrom keras.layers import K\n\nclass Attention(Layer):\n    def __init__(self, step_dim,\n                 W_regularizer=None, b_regularizer=None,\n                 W_constraint=None, b_constraint=None,\n                 bias=True, **kwargs):\n        self.supports_masking = True\n        self.init = initializers.get('glorot_uniform')\n\n        self.W_regularizer = regularizers.get(W_regularizer)\n        self.b_regularizer = regularizers.get(b_regularizer)\n\n        self.W_constraint = constraints.get(W_constraint)\n        self.b_constraint = constraints.get(b_constraint)\n\n        self.bias = bias\n        self.step_dim = step_dim\n        self.features_dim = 0\n        super(Attention, self).__init__(**kwargs)\n\n    def build(self, input_shape):\n        assert len(input_shape) == 3\n\n        self.W = self.add_weight((input_shape[-1],),\n                                 initializer=self.init,\n                                 name='{}_W'.format(self.name),\n                                 regularizer=self.W_regularizer,\n                                 constraint=self.W_constraint)\n        self.features_dim = input_shape[-1]\n\n        if self.bias:\n            self.b = self.add_weight((input_shape[1],),\n                                     initializer='zero',\n                                     name='{}_b'.format(self.name),\n                                     regularizer=self.b_regularizer,\n                                     constraint=self.b_constraint)\n        else:\n            self.b = None\n\n        self.built = True\n\n    def compute_mask(self, input, input_mask=None):\n        return None\n\n    def call(self, x, mask=None):\n        features_dim = self.features_dim\n        step_dim = self.step_dim\n\n        eij = K.reshape(K.dot(K.reshape(x, (-1, features_dim)),\n                        K.reshape(self.W, (features_dim, 1))), (-1, step_dim))\n\n        if self.bias:\n            eij += self.b\n\n        eij = K.tanh(eij)\n\n        a = K.exp(eij)\n\n        if mask is not None:\n            a *= K.cast(mask, K.floatx())\n\n        a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())\n\n        a = K.expand_dims(a)\n        weighted_input = x * a\n        return K.sum(weighted_input, axis=1)\n\n    def compute_output_shape(self, input_shape):\n        return input_shape[0],  self.features_dim","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0acf4b7fdec60cc54425c63e7919eb2f2b776f60"},"cell_type":"code","source":"def get_f1(true, val):\n    precision, recall, thresholds = precision_recall_curve(true, val)\n    thresholds = np.append(thresholds, 1.001) \n    F = 2 / (1/precision + 1/recall)\n    best_score = np.max(F)\n    best_threshold = thresholds[np.argmax(F)]\n    \n    return best_threshold, best_score    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fdb2df7e7a48cb6ed1a7a13d152ae4d5626568c8"},"cell_type":"code","source":"def build_model(units = 40, dr = 0.3):\n    inp = Input(shape = (max_len, ))\n    embed_layer = Embedding(nb_words, embed_size, input_length = max_len,\n                            weights = [embedding_matrix], trainable = False)(inp)\n    x = SpatialDropout1D(dr, seed = seed)(embed_layer)\n    x = Bidirectional(CuDNNLSTM(units, kernel_initializer = glorot_normal(seed = seed), \n                                recurrent_initializer = orthogonal(gain = 1.0, seed = seed), \n                                return_sequences = True))(x)\n    x = Bidirectional(CuDNNGRU(units, kernel_initializer = glorot_normal(seed = seed),\n                               recurrent_initializer = orthogonal(gain = 1.0, seed = seed),\n                               return_sequences = True))(x)    \n    att = Attention(max_len)(x)\n    avg_pool = GlobalAveragePooling1D()(x)\n    max_pool = GlobalMaxPooling1D()(x)\n    \n    main = concatenate([att, avg_pool, max_pool])\n    main = Dense(64, kernel_initializer = glorot_normal(seed = seed))(main)\n    main = Activation(\"relu\")(main)\n    main = Dropout(0.1, seed = seed)(main)\n    \n    out = Dense(1, activation = \"sigmoid\", \n                kernel_initializer = glorot_normal(seed = seed))(main)\n    model = Model(inputs = inp, outputs = out)\n    model.compile(loss = \"binary_crossentropy\",\n                  optimizer = Adam(), \n                  metrics = None)\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"225d98f0919544875c888e06450b61a690c1150c"},"cell_type":"code","source":"fold = 5\nbatch_size = 1024\nepochs = 5\noof_pred = np.zeros((train.shape[0], 1))\npred = np.zeros((test_shape[0], 1))\nfake_pred = np.zeros((test_shape[0], 1))\nthresholds = []\n\nk_fold = StratifiedKFold(n_splits = fold, random_state = seed, shuffle = True)\n\nfor i, (train_idx, val_idx) in enumerate(k_fold.split(train_seq, target)):\n    print(\"-\"*50)\n    print(\"Trainging fold {}/{}\".format(i+1, fold))\n    \n    X_train, y_train = train_seq[train_idx], target[train_idx]\n    X_val, y_val = train_seq[val_idx], target[val_idx]\n    \n    K.clear_session()\n    model = build_model(units = 60)\n    model.fit(X_train, y_train, batch_size = batch_size, epochs = epochs,\n              validation_data = (X_val, y_val), verbose = 2)\n    val_pred = model.predict(X_val, batch_size = batch_size)\n    oof_pred[val_idx] = val_pred\n    fake_pred += model.predict(fake_test_seq, batch_size = batch_size)/fold\n    pred += model.predict(X_test, batch_size = batch_size)/fold\n    \n    threshold, score = get_f1(y_val, val_pred)\n    print(\"F1 score at threshold {} is {}\".format(threshold, score))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cb58a8acbfc979327d420333aae49ab7805b88da"},"cell_type":"code","source":"threshold, score = get_f1(target, oof_pred)\nprint(\"F1 score after K fold at threshold {} is {}\".format(threshold, score))\nfake_test[\"pred\"] = (fake_pred > threshold).astype(int)\nprint(\"Fake test F1 score is {}\".format(f1_score(fake_test[\"target\"], \n                                                 (fake_test[\"pred\"]).astype(int))))\ntest[\"prediction\"] = (pred > threshold).astype(int)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ae362b1a255e38232e36ddac3aae192d1a39cf19"},"cell_type":"code","source":"submission = test[[\"qid\", \"prediction\"]]\nsubmission.to_csv(\"submission.csv\", index = False)\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e43a01c930939bcfb5614170975f414f71ac6da2"},"cell_type":"markdown","source":"* Fist submission: \n    * threshold = 0.4211253821849823, cv = 0.6824454496590792, fake test = 0.6862931307375753, lb = 0.688\n* Second submission: \n    * threshold = 0.4139455258846283, cv = 0.682632138928462, fake test = 0.686315118687271, lb = 0.689\n* Third submission:\n    * threshold = 0.3900025486946106, cv = 0.682267364499246, fake test = 0.684897451833437, lb = 0.687\n    \nWe see that the f1 scores of fake testes and lb are somewhat correlated. So, we assume that is related to the testing data size (~56K) and class imbalance, which bring high variances with different thresholds. \n\nWe are not sure if this assumption is true because we only look at three submissions.  "},{"metadata":{"trusted":true,"_uuid":"0fd5127db334877b06f37edbf2e578ef88366496"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}