{"cells":[{"metadata":{"_uuid":"f5334d3fee899a7efff85cf767bd889c0150f21a"},"cell_type":"markdown","source":"# Comparison of different Models\n"},{"metadata":{"_uuid":"068757b458895f8a96b5492b609a9cc47504b90d"},"cell_type":"markdown","source":"## 1. Loading & Analyzing\n\n### Loading\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"#Load General packages\nimport os\nimport numpy as np\nimport pandas as pd # for data processing\nimport matplotlib.pyplot as plt #for vizualization of data\nfrom tqdm import tqdm #for progress information\ntqdm.pandas()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"#Load \"train\" dataset\ndata = pd.read_csv(\"../input/train.csv\")\n\n#Load \"test\" dataset\nsubmission_data = pd.read_csv(\"../input/test.csv\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2efcb09d1ba79069d6f44c2aa8e75c64110c044f"},"cell_type":"markdown","source":"In the following cells, the Glove Embedding layer is loaded into a dictionary. "},{"metadata":{"trusted":true,"_uuid":"94cc0bbb87d225771f1eb99242f9322e02bba883"},"cell_type":"code","source":"#Load \"glove.840B.300d\" data\nembeddings_index = {}    #creates empty list\nglove = open('../input/embeddings/glove.840B.300d/glove.840B.300d.txt') #opens the test document for reading\nfor line in tqdm(glove): #for every line in this text do the following        (tqdm: and show the progress)\n    values = line.split(\" \")  #splits the string every time there is a space into seperate strings\n    word = values[0] #the first string in this text file is always the word\n    coefs = np.asarray(values[1:], dtype='float32') # the following strings are the \"explanation\"\n    embeddings_index[word] = coefs #the list is now filled with entries consisting of the word and the respective \"explanations\" (word vectors)\nglove.close() #closes the file such that is not possible to read it anymore\n\nprint('The dictionary contains %s word vectors.' % len(embeddings_index))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dd20492c4520b633f9f639e6b1c5ae6b139b3d67"},"cell_type":"code","source":"#Example on how Glove represents the word \"kitchen\"\nword = \"kitchen\"\nprint(\"The vector of\", word, \"in the dictionary is\", embeddings_index[word])\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2a1f44a966be3311e6779dfe461b0373356ba985"},"cell_type":"markdown","source":"### Short Data Analysis\nTo understand the Data and get a grab on how to build the Neural Network, a short analyzis of the data set seems necessary. At first the structure of the dataset should be seen."},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"666c15a7430b95c1bf5db68d1080ae34ebac3439"},"cell_type":"code","source":"data.head() # shows the first 5 rows of a dataset","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f43c719da4a45181665997aefb27eea2d4fb1da4"},"cell_type":"markdown","source":"The table shows the first rows of a set as a table with the following coloumns: quid (unique question identifier, question_text (guess what) and target (the whether it is an sincere question (0) or not (1)."},{"metadata":{"trusted":true,"_uuid":"b78e0e10150c0bcc50c8e8a3ce80d603c032f60c"},"cell_type":"code","source":"data.info() #presents general information to the dataset","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ac9256715fc7ca763e23461ca55a2fd04dc370df"},"cell_type":"markdown","source":"The info function shows amongst other whether there are empty entries (which would require some feature engineering). Here all entries are suitable. Therefore no engineering is necessary. Till now."},{"metadata":{"_uuid":"761d6d932a53991dab73649ed027a55ebce243dd"},"cell_type":"markdown","source":"At second it seems interesting how sincere questions and how insincere questions look like. "},{"metadata":{"trusted":true,"_uuid":"91fc9bd7312ce6afa1495ac29a97a6294f703ea6","scrolled":true},"cell_type":"code","source":"pd.options.display.max_colwidth = 300 # for setting the width of the table longer\n\n#Sincere questions\ndata.loc[data['target'] == 0].head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9c9a1c6ac8a792b1ce5c958c4554bfcfd9d0a044"},"cell_type":"code","source":"#Insincere questions\ndata.loc[data['target'] == 1].head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"842f2e8e90c06c1dc078aaeb7927d6571477d73a"},"cell_type":"markdown","source":"At third a short look on the distribution of sincere and insincere questions should be done. This shows that only a minor fraction of the data has a lable 1. This could lead to aproblem later, because if the system would optimize only according to accuracy (percantage of \"right\" choices), the model would reach a high accuracy with predicting only Sincere (93.8%). Fortunately accuracy is not used in the loss function. "},{"metadata":{"trusted":true,"_uuid":"4c501c526556f89f9fda6d3e38c301f16ea42971"},"cell_type":"code","source":"fig1, ax1 = plt.subplots()\nax1.pie(data[\"target\"].value_counts(), explode=(0, 0.3), labels= [\"Sincere\", \"Insincere\"], autopct='%1.1f%%',\n        shadow=True, startangle=45)\nax1.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle.\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"015cf90bf78a8b332a6862f4cd4066a72a881802"},"cell_type":"markdown","source":"## 2. Preprocessing data and building the embedding layer\n### Preprocessing\nIn Machine Learning it is common to train the model on one dataset and to test in on another dataset, to ensure that the model is not only reproducing the datra but learning the \"function\" behind the data. This Train-Test_split is done next."},{"metadata":{"trusted":true,"_uuid":"a82b459cb118b2fcd968b1c52937e5a211defcda","scrolled":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ntrain, test = train_test_split(data, #performing the split\n                 test_size = 0.3)\ntrain = train.reset_index(drop=True) #thus the df counts from 0 to x and does not spring from one number to another number\ntest = test.reset_index(drop = True)\nprint (\"The training dataset has the shape:\" , train.shape)\nprint (\"The test dataset has the shape:\", test.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"645c6d0b48ebebb3d6f6c868a602d392322a118c"},"cell_type":"markdown","source":"Following this Training and Test Sets get divided into X and Y sets. Where X contains the information used to predict the class and Y contains the information which class the question belongs to. The result will be 4 Datasets with one column."},{"metadata":{"trusted":true,"_uuid":"25bd68d79f013b9c09151ac2526b975434b29ae6"},"cell_type":"code","source":"X_train = train.iloc[:,1] #Takes all rows of the first column as new dataset\nY_train = np.array(train.iloc[:, 2]) #Takes all rows of the second column as new dataset\nX_test = test.iloc[:,1]\nY_test = np.array(test.iloc[:, 2])\n\nprint(X_train.shape)\nprint(Y_train.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7b0621d71ac6f9fa44eeb577c772ddc8da126583"},"cell_type":"markdown","source":"Because the computer can not compute strings, the words will encoded as numbers. This is done through the Keras Tokenizer. How that works is explained [here](https://machinelearningmastery.com/prepare-text-data-deep-learning-keras/)."},{"metadata":{"trusted":true,"_uuid":"d7d2ac717a408a9d6983821cff420f49c57550f4"},"cell_type":"code","source":"from keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences \nfrom keras.preprocessing import sequence \n\n#features\ntokenizer = Tokenizer(filters='', lower=False) #To ensure that no preprocessing is done at all\ntokenizer.fit_on_texts(list(data[\"question_text\"]))\nword_index = tokenizer.word_index\nprint('Found %s unique tokens.' % len(word_index))\n\nsequences = tokenizer.texts_to_sequences(data[\"question_text\"])\nmaxlen = len(max(sequences, key = len)) #max number of words in a question (längste sequenz aus tokenisierten wörtern)\n\nX_train_seq = tokenizer.texts_to_sequences(X_train)\nX_test_seq =tokenizer.texts_to_sequences(X_test)\nX_train_seq = pad_sequences(X_train_seq, maxlen=maxlen)\nX_test_seq = pad_sequences(X_test_seq, maxlen=maxlen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0b9716434f9ffb0f4bf5a8c14e483d8813da23bf"},"cell_type":"code","source":"word = \"kitchen\"\nprint(\"The index of\", word, \"in the vocabulary is\", word_index[word], \".\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d4531b037ffd6b8a205234086137917529f48a31"},"cell_type":"markdown","source":"### Fitting the dataset to the embedding layer\nIt is sayed that the text should be preprocessed especially to make it as similar as possible to the Embedding layer (for example [here](https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings)). This will be tested in the following, using some of the methods shown in the kaggle just linked. The following code helps in testing on similarities:"},{"metadata":{"trusted":true,"_uuid":"32cd347f30fe08c6d38dbb11fcf0c216b3d3a4ec"},"cell_type":"code","source":"import operator \n\ndef check_coverage(vocab,embeddings_index):\n    a = {}\n    oov = {}\n    k = 0\n    i = 0\n    for word in tqdm(vocab):\n        try:\n            a[word] = embeddings_index[word]\n            k += vocab[word]\n        except:\n\n            oov[word] = vocab[word]\n            i += vocab[word]\n            pass\n\n    print('Found embeddings for {:.2%} of vocab'.format(len(a) / len(vocab)))\n    print('Found embeddings for  {:.2%} of all text'.format(k / (k + i)))\n    sorted_x = sorted(oov.items(), key=operator.itemgetter(1))[::-1]\n\n    return sorted_x","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6e9aca8c46f1b3e77a353ac185559c2a5d30596f"},"cell_type":"markdown","source":"#### Normal Keras tokenization\n\nSome of the preprocessing is already done by the tokenizer which was used above (e.g. the text is set lowecase and and some of the punctuation is taken out of the text). The following code shows how equal the text is to the embedding layer after this sort of preprocessing."},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"1d5e2b00cf6c3361cff14e7f30e8f7e75b681f35"},"cell_type":"code","source":"oov = check_coverage(tokenizer.word_counts,embeddings_index)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4a6975d33540ff87632e4a7ffec344b04804fedc"},"cell_type":"markdown","source":"An enquiry into the most used words, not represented in the embedding layer, shows that the deviations are largely due to vernacular. "},{"metadata":{"trusted":true,"_uuid":"1362ff47c1ff15f8917902e60ba8f470a2cca983"},"cell_type":"code","source":"oov[:10]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6cb0e7cf70556eeb77d0d761aa3f284781535c9f"},"cell_type":"markdown","source":"#### Correction of vernacular\nIn [PyTorch Starter](https://www.kaggle.com/hung96ad/pytorch-starter) the authors already builded a dictionary with words which need to be replaced and wrote the text to use it. However with the glove embedding this is not necessary, as it recognizes this words, when the points are removed. Glove know's words like arent, couldnt or didnt. Therefore it is only important to remove al this interpunctuation.[](http://)"},{"metadata":{"_uuid":"c440e682515f30bf9cf3442561cb11cfefaca0b4"},"cell_type":"markdown","source":"#### Removal of Interpunctiation\nThe keras tokenizer automatically removes [\"all punctuation, plus tabs and line breaks, minus the ' character\"](http://https://keras.io/preprocessing/text/). But as shown above, this is not enough. Especially because the ‘ character gets not removed. Here again [PyTorch Starter](https://www.kaggle.com/hung96ad/pytorch-starter) offers a wider range of possibilities. If the dictionary of this authors is used, the difference to the tokenizer filter is quite large."},{"metadata":{"trusted":true,"_uuid":"3ab2568bd72743aaa1a2f7728715d7b6010d9784"},"cell_type":"code","source":"#dictionary\npuncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"63b89ede0179071506fee1335f29340791320b90","scrolled":false},"cell_type":"code","source":"#Creation of the improved data\nt1 = Tokenizer(filters = puncts,lower = False)\nt1.fit_on_texts(list(data[\"question_text\"]))\noov = check_coverage(t1.word_counts,embeddings_index)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aadd11f3220d933ba11588def1d6383062f419a9"},"cell_type":"markdown","source":"However there are still some words unknown to the embedding. Yet these words seem to be names or rather young concepts, therefore it does not seem as if anything could be made against this words."},{"metadata":{"trusted":true,"_uuid":"9f665796efb140a69f240f27daddbe807ed2d630"},"cell_type":"code","source":"oov[:10]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6d8e93ef271e7946da23e5f7beb40dab26a6ce59"},"cell_type":"markdown","source":"#### Lower case/ Upper case\nIn reviewing the keras tokenizer, it was noticable that the tokenizer automatically sets all words in lower case. This might be helpful if a small embedding only knows words in lower case, however it also can be argued that usefull information about these words get lost through the lower casing. Consider a text writing about a hotel called \"Pearl\". If this word would be lowercased, the network could think that a pearl would be meant. That Lowercasing worsenes the score can be seen next"},{"metadata":{"trusted":true,"_uuid":"0472d0a01c380be555919c329086a936abcfb025"},"cell_type":"code","source":"t2 = Tokenizer(filters = puncts, lower = True)\nt2.fit_on_texts(list(data[\"question_text\"]))\noov = check_coverage(t2.word_counts,embeddings_index)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4e0356c4a95c3630161f206cb92265eb3e5f843f"},"cell_type":"markdown","source":"Apparently lowecasing worsened and not improved the text quality. As seen in the next box, some words, like *Etherum*, where only not known, because they were lowecased but are names."},{"metadata":{"trusted":true,"_uuid":"658422dc5012e9ef3ac224384ebe95602d4539c2"},"cell_type":"code","source":"oov[:10]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3f471566f8f4c9672c2494e30bf237f8eca04dfe"},"cell_type":"markdown","source":"In the end there are now two kinds of datasets which will be used. The dataset which resulted only from the keras tokenizer and the dataset which results from the tweaked tokenizer. The second kind of datasets will be created in the next box"},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"0d38b7a4bad74fbf0abb90e16ee74c3134f49781"},"cell_type":"code","source":"tokenizer_2 = Tokenizer(filters = puncts, lower = False)\ntokenizer_2.fit_on_texts(list(data[\"question_text\"]))\nword_index_2 = tokenizer_2.word_index\nprint('Found %s unique tokens.' % len(word_index_2))\n\nsequences_2 = tokenizer_2.texts_to_sequences(data[\"question_text\"])\nmaxlen_2 = len(max(sequences_2, key = len)) #max number of words in a question (längste sequenz aus tokenisierten wörtern)\n\nX_train_seq_2 = tokenizer_2.texts_to_sequences(X_train)\nX_test_seq_2 =tokenizer_2.texts_to_sequences(X_test)\nX_train_seq_2 = pad_sequences(X_train_seq_2, maxlen=maxlen_2)\nX_test_seq_2 = pad_sequences(X_test_seq_2, maxlen=maxlen_2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"81dbe3a14f72d649d5efed4babb80c9c5d446e4e"},"cell_type":"markdown","source":"### Creating embedding matrix\nNow the two dictionaries word_index and embeddings_index get  joined together into a numpy array, so that it can be loaded into the keras framework. Throught this the  the data mass gets reduced."},{"metadata":{"trusted":true,"_uuid":"d7c4d2125cb6536debfb8cf763f01f21a7ba8a09"},"cell_type":"code","source":"#Compute Embedding Matrix for the old data\nembed_dim = 300 #da glove.840B.300d.txt bedeutet, dass 300d. vektor\nembedding_matrix = np.zeros((len(word_index) + 1, embed_dim)) #creation of the numpy array\nfor word, i in tqdm(word_index.items()): #loop going through each word in word_index\n    embedding_vector = embeddings_index.get(word) #for each word the programm takes the respective vector and calls it embedding_vector\n    if embedding_vector is not None:\n        # words not found in embedding index will be all-zeros.\n        embedding_matrix[i] = embedding_vector # vector gets inserted into the numpy array at the place where the word would stand according to the index.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"041c77255611f5d5d1e4352171564cc5f4f2a656"},"cell_type":"code","source":"#Compute Embedding Matrix for the tweaked data\nembed_dim = 300 #da glove.840B.300d.txt bedeutet, dass 300d. vektor\nembedding_matrix_2 = np.zeros((len(word_index_2) + 1, embed_dim)) #creation of the numpy array\nfor word, i in tqdm(word_index_2.items()): #loop going through each word in word_index\n    embedding_vector = embeddings_index.get(word) #for each word the programm takes the respective vector and calls it embedding_vector\n    if embedding_vector is not None:\n        # words not found in embedding index will be all-zeros.\n        embedding_matrix_2[i] = embedding_vector # vector gets inserted into the numpy array at the place where the word would stand according to the index.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"35290781c32f4f281bd6c18192534d870971477c"},"cell_type":"markdown","source":"This array will now be transfered into a Keras Embedding layer"},{"metadata":{"trusted":true,"_uuid":"ed8215ae32c4a228cb74d0e23d6bdf4472c2a4ee"},"cell_type":"code","source":"from keras.layers.embeddings import Embedding\n\n#Load into Keras Embedding layer\nembedding_layer = Embedding(len(word_index) + 1,\n                            embed_dim,\n                            weights=[embedding_matrix],\n                            input_length=maxlen,\n                            trainable=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6f3aad33a9b0eda1880168ec7ea1131e4afb9d1c"},"cell_type":"code","source":"#Load into Keras Embedding layer\nembedding_layer_2 = Embedding(len(word_index_2) + 1,\n                            embed_dim,\n                            weights=[embedding_matrix_2],\n                            input_length=maxlen_2,\n                            trainable=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec6291fc908ffc7a1d2ca5fa44f76927ed6ab4a8"},"cell_type":"markdown","source":"## 3. The Models\n### Definitions"},{"metadata":{"_uuid":"9659e1f56bc55f7fe4d25bcc2608336dd15daa54"},"cell_type":"markdown","source":"The definition of the metric put forward in the paper gets translated into code"},{"metadata":{"trusted":true,"_uuid":"4b4c5079f880073458c423198b2e18c9949e0934"},"cell_type":"code","source":"#Metric: F1 score: F1: wikipedia, umsetzung https://github.com/keras-team/keras/blob/53e541f7bf55de036f4f5641bd2947b96dd8c4c3/keras/metrics.py\n\ndef fmeasure (y_true, y_pred):\n    \n    true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n    predicted_positives = K.sum(K.round(K.clip(y_pred, 0, 1)))\n    possible_positives = K.sum(K.round(K.clip(y_true, 0, 1)))\n    \n    precision = true_positives / (predicted_positives + K.epsilon())\n    recall = true_positives / (possible_positives + K.epsilon())\n    \n    f1 = 5 * (precision*recall) / (4*precision+recall+K.epsilon())\n    \n    return f1","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"be35e21837110822c4a823df4b9a8b5360dd3aab"},"cell_type":"markdown","source":"The above definition does not work as a loss for a neural network. Therefore it must be tweaked a bit"},{"metadata":{"trusted":true,"_uuid":"16c0b8e875dd0652bf5865fd5e44912952ae2472"},"cell_type":"code","source":"#Metric: F1 score: F1: wikipedia, umsetzung https://github.com/keras-team/keras/blob/53e541f7bf55de036f4f5641bd2947b96dd8c4c3/keras/metrics.py\n\ndef loss_f1 (y_true, y_pred):\n    \n    true_positives = K.sum(y_true * y_pred)\n    predicted_positives = K.sum(y_pred)\n    possible_positives = K.sum(y_true)\n    \n    precision = true_positives / (predicted_positives + K.epsilon())\n    recall = true_positives / (possible_positives + K.epsilon())\n    \n    f1 = 5 * (precision*recall) / (4*precision+recall+K.epsilon())\n    \n    return 1-f1","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d566e20e53ad60348a6d4372bbb9f455be9454d2"},"cell_type":"markdown","source":"1. The code for the attention layer was taken from [here](https://www.kaggle.com/qqgeogor/keras-lstm-attention-glove840b-lb-0-043)"},{"metadata":{"trusted":true,"_uuid":"d92eeae0efc2bc9423d8ba903d2b8db9130d110a"},"cell_type":"code","source":"#Attention layer \nfrom keras import initializers, regularizers, constraints\nfrom keras.engine.topology import Layer\nimport keras.backend as K #to use math functions like \"keras.backend.sum\"\n\nclass Attention(Layer):\n    def __init__(self, step_dim,\n                 W_regularizer=None, b_regularizer=None,\n                 W_constraint=None, b_constraint=None,\n                 bias=True, **kwargs):\n        \n        self.supports_masking = True\n        self.init = initializers.get('glorot_uniform')\n\n        self.W_regularizer = regularizers.get(W_regularizer)\n        self.b_regularizer = regularizers.get(b_regularizer)\n\n        self.W_constraint = constraints.get(W_constraint)\n        self.b_constraint = constraints.get(b_constraint)\n\n        self.bias = bias\n        self.step_dim = step_dim\n        self.features_dim = 0\n        super(Attention, self).__init__(**kwargs)\n\n    def build(self, input_shape):\n        assert len(input_shape) == 3\n\n        self.W = self.add_weight((input_shape[-1],),\n                                 initializer=self.init,\n                                 name='{}_W'.format(self.name),\n                                 regularizer=self.W_regularizer,\n                                 constraint=self.W_constraint)\n        self.features_dim = input_shape[-1]\n\n        if self.bias:\n            self.b = self.add_weight((input_shape[1],),\n                                     initializer='zero',\n                                     name='{}_b'.format(self.name),\n                                     regularizer=self.b_regularizer,\n                                     constraint=self.b_constraint)\n        else:\n            self.b = None\n\n        self.built = True\n\n    def compute_mask(self, input, input_mask=None):\n        return None\n\n    def call(self, x, mask=None):\n        features_dim = self.features_dim\n        step_dim = self.step_dim\n\n        eij = K.reshape(K.dot(K.reshape(x, (-1, features_dim)),\n                        K.reshape(self.W, (features_dim, 1))), (-1, step_dim))\n\n        if self.bias:\n            eij += self.b\n\n        eij = K.tanh(eij)\n\n        a = K.exp(eij)\n\n        if mask is not None:\n            a *= K.cast(mask, K.floatx())\n\n        a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())\n\n        a = K.expand_dims(a)\n        weighted_input = x * a\n        return K.sum(weighted_input, axis=1)\n\n    def compute_output_shape(self, input_shape):\n        return input_shape[0],  self.features_dim","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"83dfab8b38c696a3f65c69d34ff8186475e1ec71"},"cell_type":"markdown","source":"### Building the Structure of the Models"},{"metadata":{"trusted":true,"_uuid":"7179ace1e0f75219be7db9ee0890e48e5bb853a4"},"cell_type":"code","source":"from keras.models import Model #to build the Model\nfrom keras.layers import Dense, Input, Dropout, LSTM, Activation,Bidirectional, CuDNNGRU, CuDNNLSTM # the layers we will use","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"66a8bb7103d4c740ded146f305011ddd1bc3ad30"},"cell_type":"code","source":"# First Model: Unprocessed Data, Basic Structure\nsequence_input = Input(shape=(maxlen,), dtype='int32')\nembedded_sequences = embedding_layer(sequence_input)\nX = LSTM(128, return_sequences=True)(embedded_sequences)\nX = Dropout(0.5)(X)\nX = LSTM(128)(X)\nX = Dropout(0.5)(X)\nX = Dense(1, activation = \"sigmoid\")(X)\nX = Activation('sigmoid')(X)\n\nmodel = Model(inputs=sequence_input, outputs=X)\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2e0d78f9e8cc55a3dafefcc1ca189611fb1909b5"},"cell_type":"code","source":"# Second Model: Unprocessed Data, Improved Structure\nsequence_input_2 = Input(shape=(maxlen,), dtype='int32')\nembedded_sequences_2 = embedding_layer(sequence_input_2)\nX_2 = Bidirectional(CuDNNGRU(128, return_sequences=True))(embedded_sequences_2)\nX_2 = Dropout(0.5)(X_2)\nX_2 = Bidirectional(CuDNNGRU(128, return_sequences=False))(X_2)\nX_2 = Dropout(0.5)(X_2)\nX_2 = Dense(1)(X_2)\nX_2 = Activation('sigmoid')(X_2)\n\nmodel_2 = Model(inputs = sequence_input_2, outputs=X_2)\nmodel_2.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3a8ff7e802b8733e9272d578edfd7b1ca6693cbb"},"cell_type":"code","source":"# Third Model: Preprocessed Data, Improved Structure\nsequence_input_3 = Input(shape=(maxlen_2,), dtype='int32')\nembedded_sequences_3 = embedding_layer_2(sequence_input_3)\nX_3 = Bidirectional(CuDNNGRU(128, return_sequences=True))(embedded_sequences_3)\nX_3 = Dropout(0.5)(X_3)\nX_3 = Bidirectional(CuDNNGRU(128, return_sequences=False))(X_3)\nX_3 = Dropout(0.5)(X_3)\nX_3 = Dense(1)(X_3)\nX_3 = Activation('sigmoid')(X_3)\n\nmodel_3 = Model(inputs = sequence_input_3, outputs=X_3)\nmodel_3.summary()\n\nmodel_5 = Model(inputs = sequence_input_3, outputs=X_3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"75cfe224ebe0c8194dffcc9fc7dc1365dee0b620"},"cell_type":"code","source":"# Fourth Model: Preprocessed Data, With Attention Layer\nsequence_input_4 = Input(shape=(maxlen_2,), dtype='int32')\nembedded_sequences_4 = embedding_layer_2(sequence_input_4)\nX_4 = Bidirectional(CuDNNGRU(128, return_sequences=True))(embedded_sequences_4)\nX_4 = Bidirectional(CuDNNGRU(64, return_sequences=True))(X_4)\nX_4 = Attention(maxlen_2)(X_4)\nX_4 = Dense(64, activation = \"relu\")(X_4)\nX_4 = Dense(1, activation = \"sigmoid\")(X_4)\n\nmodel_4 = Model(inputs = sequence_input_4, outputs = X_4)\nmodel_4.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"59a2c0ed7abd2faaacdef110f01c69568d01f70e"},"cell_type":"markdown","source":"### Finalising the Model"},{"metadata":{"_uuid":"fa8287671dff395c95123dc102290c5014088b71"},"cell_type":"markdown","source":"To further improve performance Early stopping is used. When the validation measure does not improve anymore, the model stops learning (thos overfitting is prevented)"},{"metadata":{"trusted":true,"_uuid":"307bad6759bb09de0cfe9e61868f8e1acb66f155"},"cell_type":"code","source":"from keras.callbacks import EarlyStopping\nearly_stopping = EarlyStopping(monitor='val_fmeasure', min_delta=0.0001, patience=2, mode='max')\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f871567833df91ec635117293913dfe534f6489"},"cell_type":"markdown","source":"The following code starts the learning procedure the respective models"},{"metadata":{"trusted":true,"_uuid":"233331ba62ce6849676ffde49bfb6685d4b74df0"},"cell_type":"code","source":"model.compile(loss= \"binary_crossentropy\", optimizer='adam', metrics= [fmeasure])\nhistory = model.fit(X_train_seq, Y_train, validation_data=(X_test_seq, Y_test),\n          epochs= 8, batch_size = 512, shuffle = True, callbacks = [early_stopping], verbose = 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fb9a4a84cfee33cd91cee0220bd930f85fcce054","scrolled":true},"cell_type":"code","source":"model_2.compile(loss= \"binary_crossentropy\", optimizer='adam', metrics= [fmeasure])\nhistory_2 = model_2.fit(X_train_seq, Y_train, validation_data=(X_test_seq, Y_test),\n          epochs= 8, batch_size = 512, shuffle = True, callbacks = [early_stopping], verbose = 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1e6e828c362cb088ec9a7efa9c04f942bc619a81"},"cell_type":"code","source":"model_3.compile(loss= \"binary_crossentropy\", optimizer='adam', metrics= [fmeasure])\nhistory_3 = model_3.fit(X_train_seq_2, Y_train, validation_data=(X_test_seq_2, Y_test),\n          epochs= 8, batch_size = 512, shuffle = True, callbacks = [early_stopping], verbose = 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f2ae96e22343bdbd0bbb3d5336e3ae3e98f52be1"},"cell_type":"code","source":"model_4.compile(loss= \"binary_crossentropy\", optimizer='adam', metrics= [fmeasure])\nhistory_4 = model_4.fit(X_train_seq_2, Y_train, validation_data=(X_test_seq_2, Y_test),\n          epochs= 8, batch_size = 512, shuffle = True, callbacks = [early_stopping], verbose = 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"20628b51946f29b103b002a5bd8102d83c02773f"},"cell_type":"code","source":"# F1-Measure as loss-function\nmodel_5.compile(loss= loss_f1, optimizer='adam', metrics= [fmeasure])\nhistory_5 = model_5.fit(X_train_seq_2, Y_train, validation_data=(X_test_seq_2, Y_test),\n          epochs= 8, batch_size = 512, shuffle = True, callbacks = [early_stopping], verbose = 2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"273427e0131baa1cf1eb90f88192d0a7656237a8"},"cell_type":"markdown","source":"## 4. The Results\n\n1. How did the f1_score change on the training data and how did it change on the test data"},{"metadata":{"trusted":true,"_uuid":"d996c5974e1a76da206b8b5b95edd8a3c7b2b532"},"cell_type":"code","source":"print(history.history.keys())\nplt.plot(history.history[\"fmeasure\"])\nplt.plot(history.history[\"val_fmeasure\"])\nplt.title('Model 1')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1894795231fce877a3342e3a17a18c88e39da818"},"cell_type":"code","source":"print(history_2.history.keys())\nplt.plot(history_2.history[\"fmeasure\"])\nplt.plot(history_2.history[\"val_fmeasure\"])\nplt.title('Model 2 -  with improved structure')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d04e6634c2fcd5eb66b83d2f6c06bff5106a07b2"},"cell_type":"code","source":"print(history_3.history.keys())\nplt.plot(history_3.history[\"fmeasure\"])\nplt.plot(history_3.history[\"val_fmeasure\"])\nplt.title('Model 3 -  with improved structure and data preparation')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"62ae71a22e2b191aa9d5fdb1ce2eed533c7df321"},"cell_type":"code","source":"print(history_4.history.keys())\nplt.plot(history_4.history[\"fmeasure\"])\nplt.plot(history_4.history[\"val_fmeasure\"])\nplt.title('Model 4 -  with data preparation and attention')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e8d50696b9a8e4d1d51ecafe5089a32b600ac18"},"cell_type":"code","source":"print(history_5.history.keys())\nplt.plot(history_5.history[\"fmeasure\"])\nplt.plot(history_5.history[\"val_fmeasure\"])\nplt.title('Model 5 -  with improved structure, data preparation and f-measure as loss function')\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}