{"cells":[{"metadata":{"_uuid":"d65f57d963f9cbb340d5d3dcfa70b495734efc0a"},"cell_type":"markdown","source":"![](https://cdn-images-1.medium.com/max/1600/1*QEDiKhSaOgi6W52S5QUHGg.png)"},{"metadata":{"_uuid":"cf4886eae4db0ff1fce33c1f31ab466d887574ba"},"cell_type":"markdown","source":"## Goal of this kernel is to first give [theoretical introduction](https://drive.google.com/file/d/1Eva0e817NkyfMmpjrEO6APMuvpXLrerx/view?usp=sharing) to **Natural Language Processing**, than tackle the **[Quora Insincere Question](https://www.kaggle.com/c/quora-insincere-questions-classification)** problem."},{"metadata":{"_uuid":"b7137c1be99a3bae82f38574be23ad4723a34304"},"cell_type":"markdown","source":"# **NOTE** Goal of this kernel is learning, not only foundations but also some advanced things and tricks taken from great kernels. Because of that some additional examples will be introduced which will put a strain on memory and computation time (there is a limit 7200 seconds to be accepted in stage 2 of this competition and mine runtime is 150 Minutes). So if you want to use it as submission tamplate -> Do not, but rather extract the usefull parts"},{"metadata":{"_uuid":"5148e24452e20f0739b9277981f7fb5f1750ab49"},"cell_type":"markdown","source":"In the [theoretical part](https://drive.google.com/file/d/1Eva0e817NkyfMmpjrEO6APMuvpXLrerx/view?usp=sharing) we are going to give Introduction to Natural Language Processing (NLP for short),\nto elaborate different techniques that propose themselves as solutions to different problems in\ncontext of text processing, to explain potential pitfalls and drawbacks and finally to present\nsome real-world python code that should give us an practitioners overview of the situation. And finally we are going to discuss different neural network\narchitectures that have evolved over the years specifically to tackle text problems, like classifying\nwhether the question was sincere or not."},{"metadata":{"_uuid":"522d9790478f62193ea5c315372a2ab9cbe9b27f"},"cell_type":"markdown","source":"**Load packages and data**"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import time\nimport random\nimport pandas as pd\nimport numpy as np\nimport gc\nimport re\nimport torch\nfrom torchtext import data\nimport spacy\nfrom tqdm import tqdm_notebook, tnrange\nfrom tqdm.auto import tqdm\n\ntqdm.pandas(desc='Progress')\nfrom collections import Counter\nfrom textblob import TextBlob\nfrom nltk import word_tokenize\n\nimport torch.nn as nn\nimport torch.optim as optim\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.nn.utils.rnn import pack_padded_sequence, pad_packed_sequence\nfrom torch.autograd import Variable\nfrom torchtext.data import Example\nfrom sklearn.metrics import f1_score\nimport torchtext\nimport os \n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\n\n# cross validation and metrics\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import f1_score\nfrom torch.optim.optimizer import Optimizer\nfrom unidecode import unidecode\n\n","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"a5f38d2cbdd99e8e2cbd310f61cc23f2faeb7195"},"cell_type":"code","source":"def seed_everything(seed=1029):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\nseed_everything()\n\nSEED=12345","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"a0472ff6aea1e8f08168daf6fd28699f5f3fe706"},"cell_type":"code","source":"embed_size = 300 # how big is each word vector\nmax_features = 120000 # how many unique words to use (i.e num rows in embedding vector)\nmaxlen = 80 # max number of words in a question to use\nbatch_size = 256 # how many samples to process at once\nn_epochs = 5 # how many times to iterate over all samples\nn_splits = 5 # Number of K-fold Splits\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6a13e6115685b19048d23bc48af4eac331946bca"},"cell_type":"markdown","source":"# First step\n\nWe have our text and we want to mak it machine readable. After the reader of this kernel went trought the theoretical part in the pdf file, one should have a clue what we are going to do. Before jumping into word2vec and other distributed (not hot one=bag of words) embeddings we need to in a sense \"label encode\" the words. In other words, given our large corpus of text (each document is a sentence) we want to seperate each document/sentence into smaller parts = Tokens (words in this case, but not always) and proceede to quantify them."},{"metadata":{"_uuid":"944df2554dee4d2df54ab4bc47efd33bac4a4117"},"cell_type":"markdown","source":"We initialise the **Tokenizer** (which is a class from keras package), and we stop at the desired number of words that we want to (en)-tokenize. \"**fit_on_texts**\" than makes the representation based on the text we pass (in our case the Training set), than we call a method **\"texts_to_sequences\"** which performs \"label-encoding\" of our vocabulary (output of the function below is exactly \"tokenizer.word_index\") . In other words we have just represented the desired number of words in a machine readable form. But it is still not good in enough. We can do better that. Thats were the transition to word2vec and other embeddings comes to place. Remember from to theory, sparcity is really bad when our dictionary is large! One natural question that arises is after we used texts_to_sequences on train_X we used on test and validation data also. But we trained it on the train data set, so what happens if we do not come across a word that was in our train set? Let us test it:"},{"metadata":{"trusted":true,"_uuid":"bd72c2730da9ff8e3cd0cda80d55d838b8682c13"},"cell_type":"code","source":"token = Tokenizer()\ntoken.fit_on_texts([\"Let us learn on a example\"]) \nprint(token.texts_to_sequences([\"Let us learn on a example\"])) \nprint(token.texts_to_sequences([\"Let us hopefully learn on a example\"]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"45e42c56aedd77081d8eafa03f9183ec46c6ea6c"},"cell_type":"code","source":"del token","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4ded096c3b72c650f5748783d9042eb9eb59b0c4"},"cell_type":"markdown","source":"No problem here, we will just skip the word. Now let us turn to the Quora problem([read the problem introduction](https://www.kaggle.com/c/quora-insincere-questions-classification)). We have the text. Let us do the first stage and initialise the Tokenize class and apply all of the necessary methods discussed above to achieve the first machine-readable form of text."},{"metadata":{"_uuid":"ccb4a7e10e7d859e8e88a5d7f55066f729ae871d"},"cell_type":"markdown","source":"# BUT\nbefore we extend our small example on the actual problem let us discuss pre-processing. **What kind of words are we encoding anyways?** We do not really everything it should be problem-specific and on the other we should also make sure that there is a representation (well for the most part) in our embedded matrix. "},{"metadata":{"_uuid":"a8426cfebb4f04fb8cb899d834cf05cfa1cb3985"},"cell_type":"markdown","source":"For example *glove embeddings*** ---not all of the words in the universe will be embedded, hence we will surely not be able to have all of the words of our dictionary represented. But we can reduce the number of such words with some corpus specific analysis. As in [here](https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings) for this competition, and built on top of that is [this](https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2) kernel. In other words there it was searched for **OOV**-out of vocabulary (words). And we are trying to have less of those. Hence with some text specific analysis (like the fact that & is in google embedinngs and? is not we can get rid of such symbols and increase the coverage of our embedding matrix! But not only that in the following piece of code following things (among others) will be conducted:\n\n1. Build a vocabulary (all of the words)\n1. Than clean it up \n1. Clean it of (out of the place) numbers, special characters, contractions and incorrect words...\n"},{"metadata":{"trusted":true,"_uuid":"496a58dc3274aed0bec343ed1b148dbdba0bbafa"},"cell_type":"code","source":"df_train = pd.read_csv(\"../input/quora-insincere-questions-classification/train.csv\")\ndf_test = pd.read_csv(\"../input/quora-insincere-questions-classification/test.csv\")\ndf = pd.concat([df_train ,df_test],sort=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8e8eec224aca39a82e337fd00d5959d3769fdf61","_kg_hide-input":true},"cell_type":"code","source":"def build_vocab(texts):\n    sentences = texts.apply(lambda x: x.split()).values\n    vocab = {}\n    for sentence in sentences:\n        for word in sentence:\n            try:\n                vocab[word] += 1\n            except KeyError:\n                vocab[word] = 1\n    return vocab\n\ndef known_contractions(embed):\n    known = []\n    for contract in contraction_mapping:\n        if contract in embed:\n            known.append(contract)\n    return known\ndef clean_contractions(text, mapping):\n    specials = [\"’\", \"‘\", \"´\", \"`\"]\n    for s in specials:\n        text = text.replace(s, \"'\")\n    text = ' '.join([mapping[t] if t in mapping else t for t in text.split(\" \")])\n    return text\ndef correct_spelling(x, dic):\n    for word in dic.keys():\n        x = x.replace(word, dic[word])\n    return x\ndef unknown_punct(embed, punct):\n    unknown = ''\n    for p in punct:\n        if p not in embed:\n            unknown += p\n            unknown += ' '\n    return unknown\n\ndef clean_numbers(x):\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x\n\ndef clean_special_chars(text, punct, mapping):\n    for p in mapping:\n        text = text.replace(p, mapping[p])\n    \n    for p in punct:\n        text = text.replace(p, f' {p} ')\n    \n    specials = {'\\u200b': ' ', '…': ' ... ', '\\ufeff': '', 'करना': '', 'है': ''}  # Other special characters that I have to deal with in last\n    for s in specials:\n        text = text.replace(s, specials[s])\n    \n    return text\ndef add_lower(embedding, vocab):\n    count = 0\n    for word in vocab:\n        if word in embedding and word.lower() not in embedding:  \n            embedding[word.lower()] = embedding[word]\n            count += 1\n    print(f\"Added {count} words to embedding\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c0d485011afdf800f7756176aad9919888afc801"},"cell_type":"code","source":"puncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]\n\ndef clean_text(x):\n    x = str(x)\n    for punct in puncts:\n        x = x.replace(punct, f' {punct} ')\n    return x\n\ndef clean_numbers(x):\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x\n\nmispell_dict = {\"ain't\": \"is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"'cause\": \"because\", \"could've\": \"could have\", \"couldn't\": \"could not\", \"didn't\": \"did not\",  \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hasn't\": \"has not\", \"haven't\": \"have not\", \"he'd\": \"he would\",\"he'll\": \"he will\", \"he's\": \"he is\", \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\", \"how's\": \"how is\",  \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \"i'd've\": \"i would have\", \"i'll\": \"i will\",  \"i'll've\": \"i will have\",\"i'm\": \"i am\", \"i've\": \"i have\", \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\", \"it'll\": \"it will\", \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \"ma'am\": \"madam\", \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \"mustn't've\": \"must not have\", \"needn't\": \"need not\", \"needn't've\": \"need not have\",\"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \"oughtn't've\": \"ought not have\", \"shan't\": \"shall not\", \"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \"she'll've\": \"she will have\", \"she's\": \"she is\", \"should've\": \"should have\", \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\"so's\": \"so as\", \"this's\": \"this is\",\"that'd\": \"that would\", \"that'd've\": \"that would have\", \"that's\": \"that is\", \"there'd\": \"there would\", \"there'd've\": \"there would have\", \"there's\": \"there is\", \"here's\": \"here is\",\"they'd\": \"they would\", \"they'd've\": \"they would have\", \"they'll\": \"they will\", \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \"to've\": \"to have\", \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \"we'll\": \"we will\", \"we'll've\": \"we will have\", \"we're\": \"we are\", \"we've\": \"we have\", \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \"what're\": \"what are\",  \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\", \"when've\": \"when have\", \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \"who'll\": \"who will\", \"who'll've\": \"who will have\", \"who's\": \"who is\", \"who've\": \"who have\", \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \"y'all'd\": \"you all would\",\"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\"y'all've\": \"you all have\",\"you'd\": \"you would\", \"you'd've\": \"you would have\", \"you'll\": \"you will\", \"you'll've\": \"you will have\", \"you're\": \"you are\", \"you've\": \"you have\", 'colour': 'color', 'centre': 'center', 'favourite': 'favorite', 'travelling': 'traveling', 'counselling': 'counseling', 'theatre': 'theater', 'cancelled': 'canceled', 'labour': 'labor', 'organisation': 'organization', 'wwii': 'world war 2', 'citicise': 'criticize', 'youtu ': 'youtube ', 'Qoura': 'Quora', 'sallary': 'salary', 'Whta': 'What', 'narcisist': 'narcissist', 'howdo': 'how do', 'whatare': 'what are', 'howcan': 'how can', 'howmuch': 'how much', 'howmany': 'how many', 'whydo': 'why do', 'doI': 'do I', 'theBest': 'the best', 'howdoes': 'how does', 'mastrubation': 'masturbation', 'mastrubate': 'masturbate', \"mastrubating\": 'masturbating', 'pennis': 'penis', 'Etherium': 'Ethereum', 'narcissit': 'narcissist', 'bigdata': 'big data', '2k17': '2017', '2k18': '2018', 'qouta': 'quota', 'exboyfriend': 'ex boyfriend', 'airhostess': 'air hostess', \"whst\": 'what', 'watsapp': 'whatsapp', 'demonitisation': 'demonetization', 'demonitization': 'demonetization', 'demonetisation': 'demonetization'}\n\ndef _get_mispell(mispell_dict):\n    mispell_re = re.compile('(%s)' % '|'.join(mispell_dict.keys()))\n    return mispell_dict, mispell_re\n\nmispellings, mispellings_re = _get_mispell(mispell_dict)\ndef replace_typical_misspell(text):\n    def replace(match):\n        return mispellings[match.group(0)]\n    return mispellings_re.sub(replace, text)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"52354a8a1977daad5921146fa5b6ca202bc75efd"},"cell_type":"markdown","source":"**Feature engineering** There can always be done some [feature engineering,](https://github.com/wongchunghang/toxic-comment-challenge-lstm/blob/master/toxic_comment_9872_model.ipynb) even with text data\n\n1.  Capital letters\n1. Length\n1. Number of words, etc...\n\n"},{"metadata":{"trusted":true,"_uuid":"2086825ac8251c33855128d7fe84d0f296923ae1"},"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\n\ndef add_features(df):\n    \n    df['question_text'] = df['question_text'].progress_apply(lambda x:str(x))\n    df['total_length'] = df['question_text'].progress_apply(len)\n    df['capitals'] = df['question_text'].progress_apply(lambda comment: sum(1 for c in comment if c.isupper()))\n    df['caps_vs_length'] = df.progress_apply(lambda row: float(row['capitals'])/float(row['total_length']),\n                                axis=1)\n    df['num_words'] = df.question_text.str.count('\\S+')\n    df['num_unique_words'] = df['question_text'].progress_apply(lambda comment: len(set(w for w in comment.split())))\n    df['words_vs_unique'] = df['num_unique_words'] / df['num_words']  \n\n    return df","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f35f32c77ca936eb3b2643b6a61f122563836101"},"cell_type":"markdown","source":"NOW we can initialise the **Tokenizer** and transform our modified text. First we will read the text, apply the above defined functions (**pre-processing & feature engineering) and than Tokenise** it."},{"metadata":{"trusted":true,"_uuid":"5cdc95950037613c690c49b27930ae0f59eb23c3","_kg_hide-input":true},"cell_type":"code","source":"def load_and_prec():\n    train_df = pd.read_csv(\"../input/quora-insincere-questions-classification/train.csv\")\n    test_df = pd.read_csv(\"../input/quora-insincere-questions-classification/test.csv\")\n    print(\"Train shape : \",train_df.shape)\n    print(\"Test shape : \",test_df.shape)\n    \n    # lower\n    train_df[\"question_text\"] = train_df[\"question_text\"].apply(lambda x: x.lower())\n    test_df[\"question_text\"] = test_df[\"question_text\"].apply(lambda x: x.lower())\n\n    # Clean the text\n    train_df[\"question_text\"] = train_df[\"question_text\"].progress_apply(lambda x: clean_text(x))\n    test_df[\"question_text\"] = test_df[\"question_text\"].apply(lambda x: clean_text(x))\n    \n    # Clean numbers\n    train_df[\"question_text\"] = train_df[\"question_text\"].progress_apply(lambda x: clean_numbers(x))\n    test_df[\"question_text\"] = test_df[\"question_text\"].apply(lambda x: clean_numbers(x))\n    \n    # Clean speelings\n    train_df[\"question_text\"] = train_df[\"question_text\"].progress_apply(lambda x: replace_typical_misspell(x))\n    test_df[\"question_text\"] = test_df[\"question_text\"].apply(lambda x: replace_typical_misspell(x))\n    \n    ## fill up the missing values\n    train_X = train_df[\"question_text\"].fillna(\"_##_\").values\n    test_X = test_df[\"question_text\"].fillna(\"_##_\").values\n\n\n    \n    ###################### Add Features ###############################\n\n    train = add_features(train_df)\n    test = add_features(test_df)\n\n    features = train[['caps_vs_length', 'words_vs_unique']].fillna(0)\n    test_features = test[['caps_vs_length', 'words_vs_unique']].fillna(0)\n\n    ss = StandardScaler()\n    ss.fit(np.vstack((features, test_features)))\n    features = ss.transform(features)\n    test_features = ss.transform(test_features)\n    ###########################################################################\n\n    ## Tokenize the sentences, as in introductory example\n    tokenizer = Tokenizer(num_words=max_features)\n    tokenizer.fit_on_texts(list(train_X))\n    train_X = tokenizer.texts_to_sequences(train_X)\n    test_X = tokenizer.texts_to_sequences(test_X)\n\n    ## Pad the sentences \n    train_X = pad_sequences(train_X, maxlen=maxlen)\n    test_X = pad_sequences(test_X, maxlen=maxlen)\n\n    ## Get the target values\n    train_y = train_df['target'].values\n    \n\n    #shuffling the data\n    np.random.seed(123)\n    trn_idx = np.random.permutation(len(train_X))\n\n    train_X = train_X[trn_idx]\n    train_y = train_y[trn_idx]\n    features = features[trn_idx]\n    \n    return train_X, test_X, train_y, features, test_features, tokenizer.word_index\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9da2cf31f3f3129615a47d7bd695d7e0d4ec4ada","_kg_hide-output":true},"cell_type":"code","source":"x_train, x_test, y_train, features, test_features, word_index = load_and_prec() ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cf65dda9ca95edd9d0f00aabc50129c2aa951729"},"cell_type":"markdown","source":"Save some memory:\n"},{"metadata":{"trusted":true,"_uuid":"7c5cdda51c1d21cba27fd3cfd935a2b8e658922f"},"cell_type":"code","source":"np.save(\"x_train\",x_train)\nnp.save(\"x_test\",x_test)\nnp.save(\"y_train\",y_train)\n\nnp.save(\"features\",features)\nnp.save(\"test_features\",test_features)\nnp.save(\"word_index.npy\",word_index)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b38c53574c6a5d1727334a0066f2397108d674a8"},"cell_type":"code","source":"x_train = np.load(\"x_train.npy\")\nx_test = np.load(\"x_test.npy\")\ny_train = np.load(\"y_train.npy\")\nfeatures = np.load(\"features.npy\")\ntest_features = np.load(\"test_features.npy\")\nword_index = np.load(\"word_index.npy\").item()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fcf76f2a96a6bab945873f7769459dd3cf8e8fb7"},"cell_type":"markdown","source":"**Distributed representation** Ok now we want to represent our words in a distributed form, as in the paper. What are some of the concerns and things we need to do before we can achieve that?\n"},{"metadata":{"_uuid":"dcf9953e0e6615f84ddc1862a8d7c8cd02fc036a"},"cell_type":"markdown","source":"If one has read the paper and paid specific attention to the Fig. 7 he may ask:\nHow can it happen that the size of the embedded vector is smaller than the size of the dictionary? In other words (and usually that is concesus) the size of the word vector is around 300-500? Now in our dictionary we definitely have more than 300 words. If we procede as in Fig. 7 in the paper I wrote, i.e. and we start training the (recurrent) NN one would think that we need to start with the one-hot (label) encoded variant. And than slowly learn the relationships between the words to get to distributed representation. **NOT REALLY** when we specify as above \"embed_size = 300\" than we just said what the dimensionality of our word embeddings. Trnaslated into our paper and NN training this 300 is actually the number of columns of the embedding matrix. And the way we do not start of with the dimensionality as big as our dictionary is that the vectors are initiated randomly in the begininng but with the size 300, and with the objective (and sufficiently large dataset) to learn the correct values in their representation."},{"metadata":{"_uuid":"f787e17682e934b9286c32878d062a79dc609e84"},"cell_type":"markdown","source":"**COOL THING** Now the training is conducted with the embedded matrix. Number of rows V is the size of the dictionary and the number of columns D is the desired size. We can learn these representation from scratch **OR** we can already used pre-trained embeddings were words these values were already optimized to a certain CORPUS (underneath is the twitter one). Below we would see that the word \"out\" already has representation which we can build upon and use to optimise it regarding our own corpus. As we shall do in this kernel."},{"metadata":{"trusted":true,"_uuid":"162d5afc163108dd399151527215e7198140385b"},"cell_type":"code","source":"with open(\"../input/glove-wiki-twitter2550/glove.twitter.27B.50d.txt\") as f:\n    lines = f.readlines()\nlines = [line.rstrip().split() for line in lines]\n\nprint(len(lines))          # number of words (aka vocabulary size)\nprint(len(lines[0]))       # length of a line\nprint(lines[99][0])       # word 99\nprint(lines[99][1:])      # vector representation of word 99\nprint(len(lines[99][1:])) # dimensionality of word 99","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d58ced125b296efa7c3d61d4be953d4dc423a1b4"},"cell_type":"code","source":"del lines\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8e9f3a1262223f0410a683230e4b2f2a5c6a363d"},"cell_type":"markdown","source":"So I said that there are certaing pre-trained embeddings. Usually **glove, fasttext and paragram** are the ones used (as we did here)"},{"metadata":{"_uuid":"dba1893c267a1e7536bbf720636647d85c7e349c"},"cell_type":"markdown","source":"**Load embeddings** \n\nWe already discussed above whats the point of pre-trained embeddings. For different problems we would change the row \n\n*\"embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\"*\n\nthat is the **number of words** and the **size of the vector** and we are set to go.\n\n\nUnderneath there are functions which take the dictionary and return the matrix of words embedded in that particular settting."},{"metadata":{"trusted":true,"_uuid":"a662716cc5fbbcc0c84019a87c52332ed8912e8d","_kg_hide-input":true},"cell_type":"code","source":"def load_glove(word_index):\n    EMBEDDING_FILE = '../input/quora-insincere-questions-classification/embeddings/glove.840B.300d/glove.840B.300d.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return embedding_matrix \n    \ndef load_fasttext(word_index):    \n    EMBEDDING_FILE = '../input/quora-insincere-questions-classification/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE) if len(o)>100)\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n\n    return embedding_matrix\n\ndef load_para(word_index):\n    EMBEDDING_FILE = '../input/quora-insincere-questions-classification/embeddings/paragram_300_sl999/paragram_300_sl999.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore') if len(o)>100)\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n    \n    return embedding_matrix","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6fb99d544ef56579faa31c777092cb24f834c2a8"},"cell_type":"markdown","source":"So which embedding do we use? Well there are a couple of ways to combine different embeddings, one that comes naturally is taking the mean across all of them:"},{"metadata":{"trusted":true,"_uuid":"b5da2d33dd3ffac3fc70007869cc4a277188c590"},"cell_type":"code","source":"# missing entries in the embedding are set using np.random.normal so we have to seed here too\nseed_everything()\n\nglove_embeddings = load_glove(word_index)\nparagram_embeddings = load_para(word_index)\nfasttext_embeddings = load_fasttext(word_index)\n\nembedding_matrix = np.mean([glove_embeddings, paragram_embeddings, fasttext_embeddings], axis=0)\n\n# vocab = build_vocab(df['question_text'])\n# add_lower(embedding_matrix, vocab)\ndel glove_embeddings, paragram_embeddings, fasttext_embeddings\ngc.collect()\n\nnp.shape(embedding_matrix)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b606acacb19410b0ce9a4341678a2a433c668a59"},"cell_type":"markdown","source":"Now that we got our different matrices let us build our models. Before that we wan to use **cross-validation**, and not only that but **stratified cross validation**. In short CV enables us to use the entire dataset when training, and Stratification is the process of rearranging the data as to ensure each fold is a good representative of the whole. For example in a binary classification problem where each class comprises 50% of the data, it is best to arrange the data such that in every fold, each class comprises around half the instances."},{"metadata":{"trusted":true,"_uuid":"e79228ede929922e5c3ced39247193cfa97bc150"},"cell_type":"code","source":"splits = list(StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=SEED).split(x_train, y_train))\nsplits[:3]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1e56c529ee5ba202edd827658f162bc0bd422b8b"},"cell_type":"markdown","source":"# **Model Architecture**\n\n(there will be a lot of extensions to Binary LSTM (NN Variant) model...)"},{"metadata":{"_uuid":"5d62d3916fc31855284f19f50124eebcab09290e"},"cell_type":"markdown","source":"We are going to build NN (variants) hence the following Python class is the idea from [CLR](https://www.kaggle.com/dannykliu/lstm-with-attention-clr-in-pytorch).  **Cyclic CLR**. Original paper [Smith](https://arxiv.org/pdf/1506.01186.pdf). Whyeven modify the Learning Rate, it turns out its one of the most important hyper-parameters of NN and it is a very important in what fashion we update our weights when backpropagating in a NN"},{"metadata":{"trusted":true,"_uuid":"d7ad7f50f75a8581ee6b7ecc758121d6c358b7f1","_kg_hide-input":true},"cell_type":"code","source":"class CyclicLR(object):\n    def __init__(self, optimizer, base_lr=1e-3, max_lr=6e-3,\n                 step_size=2000, mode='triangular', gamma=1.,\n                 scale_fn=None, scale_mode='cycle', last_batch_iteration=-1):\n\n        if not isinstance(optimizer, Optimizer):\n            raise TypeError('{} is not an Optimizer'.format(\n                type(optimizer).__name__))\n        self.optimizer = optimizer\n\n        if isinstance(base_lr, list) or isinstance(base_lr, tuple):\n            if len(base_lr) != len(optimizer.param_groups):\n                raise ValueError(\"expected {} base_lr, got {}\".format(\n                    len(optimizer.param_groups), len(base_lr)))\n            self.base_lrs = list(base_lr)\n        else:\n            self.base_lrs = [base_lr] * len(optimizer.param_groups)\n\n        if isinstance(max_lr, list) or isinstance(max_lr, tuple):\n            if len(max_lr) != len(optimizer.param_groups):\n                raise ValueError(\"expected {} max_lr, got {}\".format(\n                    len(optimizer.param_groups), len(max_lr)))\n            self.max_lrs = list(max_lr)\n        else:\n            self.max_lrs = [max_lr] * len(optimizer.param_groups)\n\n        self.step_size = step_size\n\n        if mode not in ['triangular', 'triangular2', 'exp_range'] \\\n                and scale_fn is None:\n            raise ValueError('mode is invalid and scale_fn is None')\n\n        self.mode = mode\n        self.gamma = gamma\n\n        if scale_fn is None:\n            if self.mode == 'triangular':\n                self.scale_fn = self._triangular_scale_fn\n                self.scale_mode = 'cycle'\n            elif self.mode == 'triangular2':\n                self.scale_fn = self._triangular2_scale_fn\n                self.scale_mode = 'cycle'\n            elif self.mode == 'exp_range':\n                self.scale_fn = self._exp_range_scale_fn\n                self.scale_mode = 'iterations'\n        else:\n            self.scale_fn = scale_fn\n            self.scale_mode = scale_mode\n\n        self.batch_step(last_batch_iteration + 1)\n        self.last_batch_iteration = last_batch_iteration\n\n    def batch_step(self, batch_iteration=None):\n        if batch_iteration is None:\n            batch_iteration = self.last_batch_iteration + 1\n        self.last_batch_iteration = batch_iteration\n        for param_group, lr in zip(self.optimizer.param_groups, self.get_lr()):\n            param_group['lr'] = lr\n\n    def _triangular_scale_fn(self, x):\n        return 1.\n\n    def _triangular2_scale_fn(self, x):\n        return 1 / (2. ** (x - 1))\n\n    def _exp_range_scale_fn(self, x):\n        return self.gamma**(x)\n\n    def get_lr(self):\n        step_size = float(self.step_size)\n        cycle = np.floor(1 + self.last_batch_iteration / (2 * step_size))\n        x = np.abs(self.last_batch_iteration / step_size - 2 * cycle + 1)\n\n        lrs = []\n        param_lrs = zip(self.optimizer.param_groups, self.base_lrs, self.max_lrs)\n        for param_group, base_lr, max_lr in param_lrs:\n            base_height = (max_lr - base_lr) * np.maximum(0, (1 - x))\n            if self.scale_mode == 'cycle':\n                lr = base_lr + base_height * self.scale_fn(cycle)\n            else:\n                lr = base_lr + base_height * self.scale_fn(self.last_batch_iteration)\n            lrs.append(lr)\n        return lrs","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73d68544af4c48bf9ee37492ecd05feb0b494351"},"cell_type":"markdown","source":"**Convolutional Neural Network** Why? Why not just a normal neural network. And what is so special about cNN and different in the architecture of a cNN. I also disected this black box method in another [kernel](https://www.kaggle.com/zikazika/theory-and-practice-of-image-processing-cnn). There the reader can find the original use of cNN on images but in short (and in my opinion) the fact that layers are not fully connected and the network \"only sees a small piece of an image\" or in our case we \"see/observe only a small peace of text around the desired word\" can get us better results. Read the theory!"},{"metadata":{"_uuid":"cfd7ffd9618ecad2a2846cd3e0be91dd074f2d8b"},"cell_type":"markdown","source":"Some other components (modifications) to the NN architecture:"},{"metadata":{"trusted":true,"_uuid":"626bb4a8c51a4f73d4af8981c403c70abeb49bfa"},"cell_type":"markdown","source":"# Attention Layer\n**So what is attention layer in short:** [Exposition](http://www.wildml.com/2016/01/attention-and-memory-in-deep-learning-and-nlp/)\nAttention Mechanism can be viewed as a method for making the RNN work better by letting the network know where to look as it is performing its task. Or as described in the task of generating an image caption *\"Let’s say you are trying to generate a caption from an image. Each input could be part of an image fed into the attention model. The memory layer would feed in the words already generated, the context for future word predictions. The attention model would help the algorithm decide which parts of the image to focus on as it generated each new word...\"*"},{"metadata":{"_uuid":"5a676c3a275514a3351edf306e02d832a5f39317"},"cell_type":"markdown","source":"**Attention layer**:"},{"metadata":{"trusted":true,"_uuid":"84e00df2c7b94205f5588af503f62412c48f46f3","_kg_hide-input":true},"cell_type":"code","source":"class Attention(nn.Module):\n    def __init__(self, feature_dim, step_dim, bias=True, **kwargs):\n        super(Attention, self).__init__(**kwargs)\n        \n        self.supports_masking = True\n\n        self.bias = bias\n        self.feature_dim = feature_dim\n        self.step_dim = step_dim\n        self.features_dim = 0\n        \n        weight = torch.zeros(feature_dim, 1)\n        nn.init.xavier_uniform_(weight)\n        self.weight = nn.Parameter(weight)\n        \n        if bias:\n            self.b = nn.Parameter(torch.zeros(step_dim))\n        \n    def forward(self, x, mask=None):\n        feature_dim = self.feature_dim\n        step_dim = self.step_dim\n\n        eij = torch.mm(\n            x.contiguous().view(-1, feature_dim), \n            self.weight\n        ).view(-1, step_dim)\n        \n        if self.bias:\n            eij = eij + self.b\n            \n        eij = torch.tanh(eij)\n        a = torch.exp(eij)\n        \n        if mask is not None:\n            a = a * mask\n\n        a = a / torch.sum(a, 1, keepdim=True) + 1e-10\n\n        weighted_input = x * torch.unsqueeze(a, -1)\n        return torch.sum(weighted_input, 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2cb930c54d1e7a04727b6566b59b282948d5c747"},"cell_type":"markdown","source":"* Additional layers to the **Bianry LSTM** (taken from [baseline](https://www.kaggle.com/ziliwang/baseline-pytorch-bilstm)) (Pytorch is a library which allows to be more flexible with building one owns network, but requires more expertise. It is not user friendly as keras)\nSome of the components to be found in classes underneath:\n1. [Capsules](https://www.kaggle.com/fizzbuzz/beginner-s-guide-to-capsule-networks) **Capsules** are group of neurons that \"fixate\" on a certain entity (let us say object when working with images) Powerfull thing that capsules give is **a)** probability that the entitiy exists, **b)** properties of that entity in an image for example\n2.  **GRU layer** I already discuss it in mine paper towards the end."},{"metadata":{"trusted":true,"_uuid":"d29dcc9cfab017cb8ef1797dd4b5c5a067289338","_kg_hide-input":true},"cell_type":"code","source":"import torch as t\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nembedding_dim = 300\nembedding_path = '../save/embedding_matrix.npy'  # or False, not use pre-trained-matrix\nuse_pretrained_embedding = True\n\nhidden_size = 60\ngru_len = hidden_size\n\nRoutings = 4 #5\nNum_capsule = 5\nDim_capsule = 5#16\ndropout_p = 0.25\nrate_drop_dense = 0.28\nLR = 0.001\nT_epsilon = 1e-7\nnum_classes = 30\n\n\nclass Embed_Layer(nn.Module):\n    def __init__(self, embedding_matrix=None, vocab_size=None, embedding_dim=300):\n        super(Embed_Layer, self).__init__()\n        self.encoder = nn.Embedding(vocab_size + 1, embedding_dim)\n        if use_pretrained_embedding:\n            # self.encoder.weight.data.copy_(t.from_numpy(np.load(embedding_path))) # 方法一，加载np.save的npy文件\n            self.encoder.weight.data.copy_(t.from_numpy(embedding_matrix))  # 方法二\n\n    def forward(self, x, dropout_p=0.25):\n        return nn.Dropout(p=dropout_p)(self.encoder(x))\n\n\nclass GRU_Layer(nn.Module):\n    def __init__(self):\n        super(GRU_Layer, self).__init__()\n        self.gru = nn.GRU(input_size=300,\n                          hidden_size=gru_len,\n                          bidirectional=True)\n        '''\n        自己修改GRU里面的激活函数及加dropout和recurrent_dropout\n        如果要使用，把rnn_revised import进来，但好像是使用cpu跑的，比较慢\n       '''\n        # # if you uncomment /*from rnn_revised import * */, uncomment following code aswell\n        # self.gru = RNNHardSigmoid('GRU', input_size=300,\n        #                           hidden_size=gru_len,\n        #                           bidirectional=True)\n\n    # 这步很关键，需要像keras一样用glorot_uniform和orthogonal_uniform初始化参数\n    def init_weights(self):\n        ih = (param.data for name, param in self.named_parameters() if 'weight_ih' in name)\n        hh = (param.data for name, param in self.named_parameters() if 'weight_hh' in name)\n        b = (param.data for name, param in self.named_parameters() if 'bias' in name)\n        for k in ih:\n            nn.init.xavier_uniform_(k)\n        for k in hh:\n            nn.init.orthogonal_(k)\n        for k in b:\n            nn.init.constant_(k, 0)\n\n    def forward(self, x):\n        return self.gru(x)\n\n\n# core caps_layer with squash func\nclass Caps_Layer(nn.Module):\n    def __init__(self, input_dim_capsule=gru_len * 2, num_capsule=Num_capsule, dim_capsule=Dim_capsule, \\\n                 routings=Routings, kernel_size=(9, 1), share_weights=True,\n                 activation='default', **kwargs):\n        super(Caps_Layer, self).__init__(**kwargs)\n\n        self.num_capsule = num_capsule\n        self.dim_capsule = dim_capsule\n        self.routings = routings\n        self.kernel_size = kernel_size  # 暂时没用到\n        self.share_weights = share_weights\n        if activation == 'default':\n            self.activation = self.squash\n        else:\n            self.activation = nn.ReLU(inplace=True)\n\n        if self.share_weights:\n            self.W = nn.Parameter(\n                nn.init.xavier_normal_(t.empty(1, input_dim_capsule, self.num_capsule * self.dim_capsule)))\n        else:\n            self.W = nn.Parameter(\n                t.randn(BATCH_SIZE, input_dim_capsule, self.num_capsule * self.dim_capsule))  # 64即batch_size\n\n    def forward(self, x):\n\n        if self.share_weights:\n            u_hat_vecs = t.matmul(x, self.W)\n        else:\n            print('add later')\n\n        batch_size = x.size(0)\n        input_num_capsule = x.size(1)\n        u_hat_vecs = u_hat_vecs.view((batch_size, input_num_capsule,\n                                      self.num_capsule, self.dim_capsule))\n        u_hat_vecs = u_hat_vecs.permute(0, 2, 1, 3)  # 转成(batch_size,num_capsule,input_num_capsule,dim_capsule)\n        b = t.zeros_like(u_hat_vecs[:, :, :, 0])  # (batch_size,num_capsule,input_num_capsule)\n\n        for i in range(self.routings):\n            b = b.permute(0, 2, 1)\n            c = F.softmax(b, dim=2)\n            c = c.permute(0, 2, 1)\n            b = b.permute(0, 2, 1)\n            outputs = self.activation(t.einsum('bij,bijk->bik', (c, u_hat_vecs)))  # batch matrix multiplication\n            # outputs shape (batch_size, num_capsule, dim_capsule)\n            if i < self.routings - 1:\n                b = t.einsum('bik,bijk->bij', (outputs, u_hat_vecs))  # batch matrix multiplication\n        return outputs  # (batch_size, num_capsule, dim_capsule)\n\n    # text version of squash, slight different from original one\n    def squash(self, x, axis=-1):\n        s_squared_norm = (x ** 2).sum(axis, keepdim=True)\n        scale = t.sqrt(s_squared_norm + T_epsilon)\n        return x / scale\n    \nclass Capsule_Main(nn.Module):\n    def __init__(self, embedding_matrix=None, vocab_size=None):\n        super(Capsule_Main, self).__init__()\n        self.embed_layer = Embed_Layer(embedding_matrix, vocab_size)\n        self.gru_layer = GRU_Layer()\n        # 【重要】初始化GRU权重操作，这一步非常关键，acc上升到0.98，如果用默认的uniform初始化则acc一直在0.5左右\n        self.gru_layer.init_weights()\n        self.caps_layer = Caps_Layer()\n        self.dense_layer = Dense_Layer()\n\n    def forward(self, content):\n        content1 = self.embed_layer(content)\n        content2, _ = self.gru_layer(\n            content1)  # 这个输出是个tuple，一个output(seq_len, batch_size, num_directions * hidden_size)，一个hn\n        content3 = self.caps_layer(content2)\n        output = self.dense_layer(content3)\n        return output\n    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d96793d88c22274d985436e192f62970c227c324"},"cell_type":"markdown","source":"**Binary LSTM itself**"},{"metadata":{"trusted":true,"_uuid":"05164d541a0c35cae727d0338548d156efe21427","_kg_hide-input":true},"cell_type":"code","source":"class NeuralNet(nn.Module):\n    def __init__(self):\n        super(NeuralNet, self).__init__()\n        \n        fc_layer = 16\n        fc_layer1 = 16\n\n        self.embedding = nn.Embedding(max_features, embed_size)\n        self.embedding.weight = nn.Parameter(torch.tensor(embedding_matrix, dtype=torch.float32))\n        self.embedding.weight.requires_grad = False\n        \n        self.embedding_dropout = nn.Dropout2d(0.1)\n        self.lstm = nn.LSTM(embed_size, hidden_size, bidirectional=True, batch_first=True)\n        self.gru = nn.GRU(hidden_size * 2, hidden_size, bidirectional=True, batch_first=True)\n\n        self.lstm2 = nn.LSTM(hidden_size * 2, hidden_size, bidirectional=True, batch_first=True)\n\n        self.lstm_attention = Attention(hidden_size * 2, maxlen)\n        self.gru_attention = Attention(hidden_size * 2, maxlen)\n        self.bn = nn.BatchNorm1d(16, momentum=0.5)\n        self.linear = nn.Linear(hidden_size*8+3, fc_layer1) #643:80 - 483:60 - 323:40\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(0.1)\n        self.fc = nn.Linear(fc_layer**2,fc_layer)\n        self.out = nn.Linear(fc_layer, 1)\n        self.lincaps = nn.Linear(Num_capsule * Dim_capsule, 1)\n        self.caps_layer = Caps_Layer()\n    \n    def forward(self, x):\n        \n#         Capsule(num_capsule=10, dim_capsule=10, routings=4, share_weights=True)(x)\n\n        h_embedding = self.embedding(x[0])\n        h_embedding = torch.squeeze(\n            self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))\n        \n        h_lstm, _ = self.lstm(h_embedding)\n        h_gru, _ = self.gru(h_lstm)\n\n        ##Capsule Layer        \n        content3 = self.caps_layer(h_gru)\n        content3 = self.dropout(content3)\n        batch_size = content3.size(0)\n        content3 = content3.view(batch_size, -1)\n        content3 = self.relu(self.lincaps(content3))\n\n        ##Attention Layer\n        h_lstm_atten = self.lstm_attention(h_lstm)\n        h_gru_atten = self.gru_attention(h_gru)\n        \n        # global average pooling\n        avg_pool = torch.mean(h_gru, 1)\n        # global max pooling\n        max_pool, _ = torch.max(h_gru, 1)\n        \n        f = torch.tensor(x[1], dtype=torch.float).cuda()\n\n                #[512,160]\n        conc = torch.cat((h_lstm_atten, h_gru_atten,content3, avg_pool, max_pool,f), 1)\n        conc = self.relu(self.linear(conc))\n        conc = self.bn(conc)\n        conc = self.dropout(conc)\n\n        out = self.out(conc)\n        \n        return out","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2bf953e2b5b6d9363eda89971ecc0ac416e2ddd0"},"cell_type":"markdown","source":"Model is ready, now only thing left to do is to train and make predictions:\nOne peculiar thing is that with Pytorch is that it wont go as smoothly as with keras. WE have some manual work do to. Idea [from](https://www.kaggle.com/hengzheng/pytorch-starter)"},{"metadata":{"_uuid":"9761f6d10dbca6d29a2d2e3c1a6a7dda1314591e","trusted":true},"cell_type":"code","source":"class MyDataset(Dataset):\n    def __init__(self,dataset):\n        self.dataset = dataset\n\n    def __getitem__(self, index):\n        data, target = self.dataset[index]\n\n        return data, target, index\n    def __len__(self):\n        return len(self.dataset)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0e597338552325edf82aeeda408f809575c61133","trusted":true},"cell_type":"code","source":"\n# matrix for the out-of-fold predictions\ntrain_preds = np.zeros((len(x_train)))\n# matrix for the predictions on the test set\ntest_preds = np.zeros((len(df_test)))\n\n# always call this before training for deterministic results\nseed_everything()\n\nx_test_cuda = torch.tensor(x_test, dtype=torch.long).cuda()\ntest = torch.utils.data.TensorDataset(x_test_cuda)\ntest_loader = torch.utils.data.DataLoader(test, batch_size=batch_size, shuffle=False)\n\navg_losses_f = []\navg_val_losses_f = []\n\ndef sigmoid(x):\n    return 1 / (1 + np.exp(-x))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5fdbb11f19f169855f3723b6cc4d40e7b77f5004","trusted":true},"cell_type":"code","source":"for i, (train_idx, valid_idx) in enumerate(splits):    \n    # split data in train / validation according to the KFold indeces\n    # also, convert them to a torch tensor and store them on the GPU (done with .cuda())\n    x_train = np.array(x_train)\n    y_train = np.array(y_train)\n    features = np.array(features)\n\n    x_train_fold = torch.tensor(x_train[train_idx.astype(int)], dtype=torch.long).cuda()\n    y_train_fold = torch.tensor(y_train[train_idx.astype(int), np.newaxis], dtype=torch.float32).cuda()\n    \n    kfold_X_features = features[train_idx.astype(int)]\n    kfold_X_valid_features = features[valid_idx.astype(int)]\n    x_val_fold = torch.tensor(x_train[valid_idx.astype(int)], dtype=torch.long).cuda()\n    y_val_fold = torch.tensor(y_train[valid_idx.astype(int), np.newaxis], dtype=torch.float32).cuda()\n    \n#     model = BiLSTM(lstm_layer=2,hidden_dim=40,dropout=DROPOUT).cuda()\n    model = NeuralNet()\n\n    # make sure everything in the model is running on the GPU\n    model.cuda()\n\n    # define binary cross entropy loss\n    # note that the model returns logit to take advantage of the log-sum-exp trick \n    # for numerical stability in the loss\n    loss_fn = torch.nn.BCEWithLogitsLoss(reduction='sum')\n\n    step_size = 400\n    base_lr, max_lr = 0.001, 0.004   \n    optimizer = torch.optim.Adam(filter(lambda p: p.requires_grad, model.parameters()), \n                             lr=max_lr)\n    \n    ################################################################################################\n    scheduler = CyclicLR(optimizer, base_lr=base_lr, max_lr=max_lr,\n               step_size=step_size, mode='exp_range',\n               gamma=0.99994)\n    ###############################################################################################\n\n    train = torch.utils.data.TensorDataset(x_train_fold, y_train_fold)\n    valid = torch.utils.data.TensorDataset(x_val_fold, y_val_fold)\n    \n    train = MyDataset(train)\n    valid = MyDataset(valid)\n\n    ##No need to shuffle the data again here. Shuffling happens when splitting for kfolds.\n    train_loader = torch.utils.data.DataLoader(train, batch_size=batch_size, shuffle=True)\n    \n    valid_loader = torch.utils.data.DataLoader(valid, batch_size=batch_size, shuffle=False)\n\n    print(f'Fold {i + 1}')\n    for epoch in range(n_epochs):\n        # set train mode of the model. This enables operations which are only applied during training like dropout\n        start_time = time.time()\n        model.train()\n\n        avg_loss = 0.  \n        for i, (x_batch, y_batch, index) in enumerate(train_loader):\n            # Forward pass: compute predicted y by passing x to the model.\n            ################################################################################################            \n            f = kfold_X_features[index]\n            y_pred = model([x_batch,f])\n            ################################################################################################\n\n            ################################################################################################\n\n            if scheduler:\n                scheduler.batch_step()\n            ################################################################################################\n\n\n            # Compute and print loss.\n            loss = loss_fn(y_pred, y_batch)\n\n            # Before the backward pass, use the optimizer object to zero all of the\n            # gradients for the Tensors it will update (which are the learnable weights\n            # of the model)\n            optimizer.zero_grad()\n\n            # Backward pass: compute gradient of the loss with respect to model parameters\n            loss.backward()\n\n            # Calling the step function on an Optimizer makes an update to its parameters\n            optimizer.step()\n            avg_loss += loss.item() / len(train_loader)\n            \n        # set evaluation mode of the model. This disabled operations which are only applied during training like dropout\n        model.eval()\n        \n        # predict all the samples in y_val_fold batch per batch\n        valid_preds_fold = np.zeros((x_val_fold.size(0)))\n        test_preds_fold = np.zeros((len(df_test)))\n        \n        avg_val_loss = 0.\n        for i, (x_batch, y_batch, index) in enumerate(valid_loader):\n            f = kfold_X_valid_features[index]\n            y_pred = model([x_batch,f]).detach()\n            \n            avg_val_loss += loss_fn(y_pred, y_batch).item() / len(valid_loader)\n            valid_preds_fold[i * batch_size:(i+1) * batch_size] = sigmoid(y_pred.cpu().numpy())[:, 0]\n        \n        elapsed_time = time.time() - start_time \n        print('Epoch {}/{} \\t loss={:.4f} \\t val_loss={:.4f} \\t time={:.2f}s'.format(\n            epoch + 1, n_epochs, avg_loss, avg_val_loss, elapsed_time))\n    avg_losses_f.append(avg_loss)\n    avg_val_losses_f.append(avg_val_loss) \n    # predict all samples in the test set batch per batch\n    for i, (x_batch,) in enumerate(test_loader):\n        f = test_features[i * batch_size:(i+1) * batch_size]\n        y_pred = model([x_batch,f]).detach()\n\n        test_preds_fold[i * batch_size:(i+1) * batch_size] = sigmoid(y_pred.cpu().numpy())[:, 0]\n        \n    train_preds[valid_idx] = valid_preds_fold\n    test_preds += test_preds_fold / len(splits)\n\nprint('All \\t loss={:.4f} \\t val_loss={:.4f} \\t '.format(np.average(avg_losses_f),np.average(avg_val_losses_f)))\n\n# x_train, x_test_f, y_train, y_test_f","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bb661e3c8d8ce6b65582d0bc18fad9f76eab12ff"},"cell_type":"markdown","source":"**Cut-off level.** One needs to remember that we predict essentially probabilites and really 1 and 0. What we need to do when finishing up with any classification is find the cut-off level that minimizes the loss function (Binary it is usually F1 metric) and than using that cut-off level make predictions on the test set. Just gor trough the cut-off levels and measure the F1 score. The best cut-off level is the best F1 score. "},{"metadata":{"_uuid":"00202e47c9799253f443d88b93a554d20cdf1abf","trusted":true},"cell_type":"code","source":"def bestThresshold(y_train,train_preds):\n    tmp = [0,0,0] # idx, cur, max\n    delta = 0\n    for tmp[0] in tqdm(np.arange(0.1, 0.501, 0.01)):\n        tmp[1] = f1_score(y_train, np.array(train_preds)>tmp[0])\n        if tmp[1] > tmp[2]:\n            delta = tmp[0]\n            tmp[2] = tmp[1]\n    print('best threshold is {:.4f} with F1 score: {:.4f}'.format(delta, tmp[2]))\n    return delta\ndelta = bestThresshold(y_train,train_preds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"25345e06192cde8967dd704e7c595f6714cd1c6d"},"cell_type":"code","source":"submission = df_test[['qid']].copy()\nsubmission['prediction'] = (test_preds > delta).astype(int)\nsubmission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}