{"cells":[{"metadata":{"_uuid":"786a1f030ad6ba3d13621dbdfbd663ebc37ff42f"},"cell_type":"markdown","source":"This kernel is a part of a [**post**](http://mlwhiz.com/blog/2018/12/17/text_classification/)  written by me on my [**blog**](http://mlwhiz.com/blog/2018/12/17/text_classification/) to try to learn text classification using Deep learning. I have tried to explain some of the models in an intutive way. Do take a look. \n\n### Import Libraries"},{"metadata":{"trusted":true,"_uuid":"452ccaa5f89f25121f86f54a6a6eb07c9134a567"},"cell_type":"code","source":"# Some imports, we are not gong to use all the imports in this workbook but in subsequent workbooks we surely will.\nimport os\nimport time\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nimport math\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import metrics\n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers import Dense, Input, CuDNNLSTM, Embedding, Dropout, Activation, CuDNNGRU, Conv1D\nfrom keras.layers import Bidirectional, GlobalMaxPool1D, GlobalMaxPooling1D, GlobalAveragePooling1D\nfrom keras.layers import Input, Embedding, Dense, Conv2D, MaxPool2D, concatenate\nfrom keras.layers import Reshape, Flatten, Concatenate, Dropout, SpatialDropout1D\nfrom keras.optimizers import Adam\nfrom keras.models import Model\nfrom keras import backend as K\nfrom keras.engine.topology import Layer\nfrom keras import initializers, regularizers, constraints, optimizers, layers\n\n\nfrom keras.layers import *\nfrom keras.models import *\nfrom keras import initializers, regularizers, constraints, optimizers, layers\nfrom keras.initializers import *\nfrom keras.optimizers import *\nimport keras.backend as K\nfrom keras.callbacks import *\nimport tensorflow as tf\nimport os\nimport time\nimport gc\nimport re\nimport glob","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"620eae354ed80c3af5559314ab8d1ccc768dbbe9"},"cell_type":"markdown","source":"### Data Preprocessing"},{"metadata":{"trusted":true,"_uuid":"1694357ed892df24447d2045570582142b7616f4"},"cell_type":"code","source":"# Define some Global Variables\nmax_features = 100000 # Maximum Number of words we want to include in our dictionary\nmaxlen = 72 # No of words in question we want to create a sequence with\nembed_size = 300# Size of word to vec embedding we are using","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ded9ec27c9625c9fd159f25dfd21ba3ba9a3643a"},"cell_type":"code","source":"# Some preprocesssing that will be common to all the text classification methods you will see. \npuncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]\ndef clean_text(x):\n    x = str(x)\n    for punct in puncts:\n        x = x.replace(punct, f' {punct} ')\n    return x\n\n# Loading the data\ndef load_and_prec():\n    train_df = pd.read_csv(\"../input/train.csv\")\n    test_df = pd.read_csv(\"../input/test.csv\")\n    \n    print(\"Train shape : \",train_df.shape)\n    print(\"Test shape : \",test_df.shape)\n    \n    train_df[\"question_text\"] = train_df[\"question_text\"].apply(lambda x: clean_text(x))\n    test_df[\"question_text\"] = test_df[\"question_text\"].apply(lambda x: clean_text(x))\n    \n    ## split to train and val\n    train_df, val_df = train_test_split(train_df, test_size=0.08, random_state=2018) # .08 since the datasize is large enough.\n\n    ## fill up the missing values\n    train_X = train_df[\"question_text\"].fillna(\"_##_\").values\n    val_X = val_df[\"question_text\"].fillna(\"_##_\").values\n    test_X = test_df[\"question_text\"].fillna(\"_##_\").values\n\n    ## Tokenize the sentences\n    '''\n    keras.preprocessing.text.Tokenizer tokenizes(splits) the texts into tokens(words).\n    Signature:\n    Tokenizer(num_words=None, filters='!\"#$%&()*+,-./:;<=>?@[\\\\]^_`{|}~\\t\\n', \n    lower=True, split=' ', char_level=False, oov_token=None, document_count=0, **kwargs)\n\n    The num_words parameter keeps a prespecified number of words in the text only. \n    It also filters some non wanted tokens by default and converts the text into lowercase.\n\n    It keeps an index of words(dictionary of words which we can use to assign a unique number to a word) \n    which can be accessed by tokenizer.word_index.\n    For example - For a text corpus the tokenizer word index might look like. \n    The words in the indexed dictionary are sort of ranked in order of frequencies,\n    {'the': 1,'what': 2,'is': 3, 'a': 4, 'to': 5, 'in': 6, 'of': 7, 'i': 8, 'how': 9}\n    \n    The texts_to_sequence function converts every word(token) to its respective index in the word_index\n    \n    So Lets say we started with \n    train_X as something like ['This is a sentence','This is another bigger sentence']\n    and after fitting our tokenizer we get the word_index as {'this':1,'is':2,'sentence':3,'a':4,'another':5,'bigger':6}\n    The texts_to_sequence function will tokenize the sentences and replace words with individual tokens to give us \n    train_X = [[1,2,4,3],[1,2,5,6,3]]\n    '''\n    tokenizer = Tokenizer(num_words=max_features)\n    tokenizer.fit_on_texts(list(train_X))\n    train_X = tokenizer.texts_to_sequences(train_X)\n    val_X = tokenizer.texts_to_sequences(val_X)\n    test_X = tokenizer.texts_to_sequences(test_X)\n\n    ## Pad the sentences. We need to pad the sequence with 0's to achieve consistent length across examples.\n    '''\n    We had train_X = [[1,2,4,3],[1,2,5,6,3]]\n    lets say maxlen=6\n        We will then get \n        train_X = [[1,2,4,3,0,0],[1,2,5,6,3,0]]\n    '''\n    train_X = pad_sequences(train_X, maxlen=maxlen)\n    val_X = pad_sequences(val_X, maxlen=maxlen)\n    test_X = pad_sequences(test_X, maxlen=maxlen)\n\n    ## Get the target values\n    train_y = train_df['target'].values\n    val_y = val_df['target'].values  \n    \n    #shuffling the data\n    np.random.seed(2018)\n    trn_idx = np.random.permutation(len(train_X))\n    val_idx = np.random.permutation(len(val_X))\n\n    train_X = train_X[trn_idx]\n    val_X = val_X[val_idx]\n    train_y = train_y[trn_idx]\n    val_y = val_y[val_idx]    \n    \n    return train_X, val_X, test_X, train_y, val_y, tokenizer.word_index","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f911779024fee43a5dd97cb1e7c4c60c960c7f28"},"cell_type":"code","source":"train_X, val_X, test_X, train_y, val_y, word_index = load_and_prec()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"def8373452faff54d665749e62bdad81166ceec0"},"cell_type":"code","source":"# Word 2 vec Embedding\n\ndef load_glove(word_index):\n    '''We want to create an embedding matrix in which we keep only the word2vec for words which are in our word_index\n    '''\n    EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = -0.005838499,0.48782197\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return embedding_matrix ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bb2816699bba88f7a97c611ba61c63b259cbdeb0"},"cell_type":"code","source":"embedding_matrix = load_glove(word_index)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"25b62fa53df10749b2b114886a24ef859cd9754c"},"cell_type":"code","source":"# https://www.kaggle.com/yekenot/2dcnn-textclassifier\ndef model_cnn(embedding_matrix):\n    filter_sizes = [1,2,3,5]\n    num_filters = 36\n\n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix])(inp)\n    x = Reshape((maxlen, embed_size, 1))(x)\n\n    maxpool_pool = []\n    for i in range(len(filter_sizes)):\n        conv = Conv2D(num_filters, kernel_size=(filter_sizes[i], embed_size),\n                                     kernel_initializer='he_normal', activation='elu')(x)\n        maxpool_pool.append(MaxPool2D(pool_size=(maxlen - filter_sizes[i] + 1, 1))(conv))\n\n    z = Concatenate(axis=1)(maxpool_pool)   \n    z = Flatten()(z)\n    z = Dropout(0.1)(z)\n\n    outp = Dense(1, activation=\"sigmoid\")(z)\n\n    model = Model(inputs=inp, outputs=outp)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"582f4442e773e54a9a4685fbe4792b52cc282674"},"cell_type":"code","source":"model = model_cnn(embedding_matrix)\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7a18102915344e787126ebba1f5c92587ffa2d54"},"cell_type":"code","source":"def train_pred(model, epochs=2):\n    filepath=\"weights_best.h5\"\n    checkpoint = ModelCheckpoint(filepath, monitor='val_loss', verbose=2, save_best_only=True, mode='min')\n    reduce_lr = ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=1, min_lr=0.0001, verbose=2)\n    earlystopping = EarlyStopping(monitor='val_loss', min_delta=0.0001, patience=2, verbose=2, mode='auto')\n    callbacks = [checkpoint, reduce_lr]\n    for e in range(epochs):\n        model.fit(train_X, train_y, batch_size=512, epochs=1, validation_data=(val_X, val_y),callbacks=callbacks)\n    model.load_weights(filepath)\n    pred_val_y = model.predict([val_X], batch_size=1024, verbose=0)\n    pred_test_y = model.predict([test_X], batch_size=1024, verbose=0)\n    return pred_val_y, pred_test_y","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c7e4b7948f12d43a61217139914202a67f5f0c5b"},"cell_type":"code","source":"pred_val_y, pred_test_y = train_pred(model, epochs=8)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"911e98d96a4b090f2148d6917bc708076ac08ffe"},"cell_type":"code","source":"'''\nA function specific to this competition since the organizers don't want probabilities \nand only want 0/1 classification maximizing the F1 score. This function computes the best F1 score by looking at val set predictions\n'''\n\ndef f1_smart(y_true, y_pred):\n    thresholds = []\n    for thresh in np.arange(0.1, 0.501, 0.01):\n        thresh = np.round(thresh, 2)\n        res = metrics.f1_score(y_true, (y_pred > thresh).astype(int))\n        thresholds.append([thresh, res])\n        print(\"F1 score at threshold {0} is {1}\".format(thresh, res))\n\n    thresholds.sort(key=lambda x: x[1], reverse=True)\n    best_thresh = thresholds[0][0]\n    best_f1 = thresholds[0][1]\n    print(\"Best threshold: \", best_thresh)\n    return  best_f1, best_thresh","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"56fb1ac393af08448f9539c8a9934619b50bdb3f"},"cell_type":"code","source":"pred_val_y","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ccfdd5f740f976aae465c5e304c1bce129edf198"},"cell_type":"code","source":"val_y","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fadb40378cce0d73f809fb92e4b3198b6b006ca6"},"cell_type":"code","source":"f1, threshold = f1_smart(val_y, pred_val_y)\nprint('Optimal F1: {} at threshold: {}'.format(f1, threshold))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e465c940632f88f46e0c50bc361af713a3fa598e"},"cell_type":"code","source":"pred_test_y = (pred_test_y >threshold).astype(int)\ntest_df = pd.read_csv(\"../input/test.csv\", usecols=[\"qid\"])\nout_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\nout_df['prediction'] = pred_test_y\nout_df.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a639045cf40f9d1fa07d134119774df52add4733"},"cell_type":"markdown","source":"## References\n\n* Based on SRK's kernel: https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\n* Vladimir Demidov's 2DCNN textClassifier: https://www.kaggle.com/yekenot/2dcnn-textclassifier\n* Shujian's https://www.kaggle.com/shujian/fork-of-mix-of-nn-models"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}