{"cells":[{"metadata":{"id":"1pnaFCv8uD3j","colab_type":"text","_uuid":"7d0f74a66080fcb7f36357bae3b8dae5dcd00054"},"cell_type":"markdown","source":"# Challenge Kaggle - Quora Insincere Questions Classification - Simple Version"},{"metadata":{"id":"VVr5u1YO4IdF","colab_type":"text","_uuid":"5d7d37b2bc7187160ea643458400afb1da52b9f5"},"cell_type":"markdown","source":"This notebook presents a simple solution for the Kaggle Challenge: [Quora Insincere Questions Classification](https://www.kaggle.com/c/quora-insincere-questions-classification).\n\nThis problem can be understood as a sentiment analysis problem, one of the most common downstream tasks in NLP.\n\nThe objective is to build a simple RNN model step-by-step in order to understand how such problem can be tackled.\n\nI am going to use the PyTorch library as deep-learning framework.\n\nIn this solution, I am not going to use one of pre-trained embedding made available by the challenge. In order to present a more didactic solution, I am going to hand-craft the tokenization and we are going to add a embedding layer for training in the model. Please note that this approach will increase the training time, as the model will need to learn the paramaters of the embeddings, which quantity can be huge given the size of the vocabulary."},{"metadata":{"id":"h-the-mXuRy1","colab_type":"text","_uuid":"545a5d254f783d0faeaed201b08951291290f2ee"},"cell_type":"markdown","source":"## Import of useful libraries"},{"metadata":{"id":"VilPlNhrwCwb","colab_type":"code","colab":{},"trusted":true,"_uuid":"36051c0e5c772fc1ff0321d142ad944fbc33e6bf"},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import f1_score\nimport torch\nfrom torch import nn\nfrom torch import optim\nimport torch.nn.functional as F\nfrom torch.utils.data import TensorDataset, DataLoader","execution_count":null,"outputs":[]},{"metadata":{"id":"oENx8XRpvI3g","colab_type":"text","_uuid":"b2beec1b7c8e43d02904d9f6127ab0bd7fcc39e2"},"cell_type":"markdown","source":"\n## Loading the data"},{"metadata":{"id":"2_4RYe_w36Mf","colab_type":"text","_uuid":"c39435cff1afa467ad2859675a07d956d2a13eef"},"cell_type":"markdown","source":"### Loading data and visualizing data"},{"metadata":{"id":"lrAEZp29ufAM","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":51},"outputId":"b4740296-441d-47ee-ce67-407708b4433e","trusted":true,"_uuid":"9f66ced7a4b20eb36528e2ba3ad4872040bd3681"},"cell_type":"code","source":"train_df = pd.read_csv(\"../input/train.csv\")\ntest_df = pd.read_csv(\"../input/test.csv\")\nprint(\"Train shape : \",train_df.shape)\nprint(\"Test shape : \",test_df.shape)","execution_count":null,"outputs":[]},{"metadata":{"id":"DyEkCFDKBk9T","colab_type":"text","_uuid":"d3d0439e93e0a2e527b0add5d682dae5b1f54ae2"},"cell_type":"markdown","source":"Let's visualizing the structure of the dataframes and the Quora's questions"},{"metadata":{"id":"DeCuml9lBt45","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":204},"outputId":"b876498a-6cd0-4261-ec72-4c11dc69f11f","trusted":true,"_uuid":"84d50488ab57831e4f8a7fa00544e1eb1573d7e7"},"cell_type":"code","source":"train_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"id":"0MkaVW8EBy4P","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":102},"outputId":"13a8d3e4-c0ed-46b3-8c33-2e7c3b09eae6","trusted":true,"_uuid":"b05f3274614beb52dacacb6399b4092820fd52ac"},"cell_type":"code","source":"for question in train_df.question_text[:5]:\n    print(question)","execution_count":null,"outputs":[]},{"metadata":{"id":"a6RDhnNlOfha","colab_type":"text","_uuid":"2f1f77730631569f4eef9b501d32e219b35a361f"},"cell_type":"markdown","source":"Some statitistics of the datasets:"},{"metadata":{"id":"N6Lo_GceNq6f","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":68},"outputId":"52da436e-9c79-4a57-8e12-210c00fb03aa","trusted":true,"_uuid":"e792bc8dcefdf3562ab61067ce5d920a4bf9f69f"},"cell_type":"code","source":"labels = np.array(train_df.target)\n\nprint('Number of training samples:', len(train_df))\nprint('Percentage of insincere questions: {:.2%}'.format(labels.sum()/len(train_df)))\nprint('Number of test samples:', len(test_df))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67e5ece30ab19b20990c2f97df1a93f6b9c29db4"},"cell_type":"markdown","source":"It can be noted that the data set is very unballanced (only 6.19% of the training samples are labeled 1), therefore accuracy is not a good metric to evalutate performance. The metric chosen by the competition, F1-score, is a good metric for evluation of unbalaced datasets."},{"metadata":{"colab_type":"text","id":"hhrrCT-DaHw_","_uuid":"93cdd4a94874225a08a6aae073a107eeabcfdc6f"},"cell_type":"markdown","source":"Visualizing a few examples of insincere questions:"},{"metadata":{"id":"2Kzu0wiyQOSI","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":122},"outputId":"ab5b76c7-f4f1-4ed6-e663-bae3b469c69d","trusted":true,"_uuid":"df1445deb67cfd503cd28cf256701987068b5ccc"},"cell_type":"code","source":"ins_idx = np.where(np.array(train_df.target)==1)[0]\nfor question in train_df.question_text[ins_idx[:5]]:\n    print(question)","execution_count":null,"outputs":[]},{"metadata":{"id":"XoPwK8r_PDuF","colab_type":"text","_uuid":"0a92de91db8c575e5f326c1a5712d3db0d5d247c"},"cell_type":"markdown","source":"## Data preprocessing"},{"metadata":{"id":"ZzOCWgzAP_-8","colab_type":"text","_uuid":"c4d34b29ce40e93075d4daeac2ffa9ba4e6c943b"},"cell_type":"markdown","source":"### Eliminate punctuation and lower the sentences\n\nFor this simpler notebook, I am only going to deal with the words in the sentences, without taking into account the punctuation. It is however important to notice that ponctuation might play a big role in this task, as insincere questions are likely to have some particular patterns of punctuation."},{"metadata":{"id":"HfSNahd7P-00","colab_type":"code","colab":{},"trusted":true,"_uuid":"b5d402d13bdab51b442ba7fee3cb6eaf1cc220c7"},"cell_type":"code","source":"from string import punctuation\n\ndef lower_eliminate_punctuation(sentence):\n    '''\n    Function that takes an string as input, lower it and get rid of its punctuation\n    '''\n    filtered_sentence = ''.join([c for c in sentence if c not in punctuation])\n    return filtered_sentence.lower()\n\n# Getting rid of punctuation in both datasets\ntrain_df.question_text = train_df.question_text.apply(lower_eliminate_punctuation)\ntest_df.question_text = test_df.question_text.apply(lower_eliminate_punctuation)","execution_count":null,"outputs":[]},{"metadata":{"id":"gK68PyINggsm","colab_type":"text","_uuid":"e863dc19a9fcff9e3281c670a52960d8c1ad79a1"},"cell_type":"markdown","source":"### Dealing with outliers\n\nIf there are empty sentences in the datasets, we should not take them into account"},{"metadata":{"id":"tUrDT69Xgj2l","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":34},"outputId":"aa47ee2e-4dec-4a42-a9d5-dea6adfc5032","trusted":true,"_uuid":"04a85129902a02f9bb4fb1955a4bb2335249f98a"},"cell_type":"code","source":"# Training set\nempty_idx_train = []\nfor i, question in enumerate(train_df.question_text):\n    if question == '':\n        print('Empty question at index', i, 'with label ', train_df.target[i])\n        empty_idx_train.append(i)","execution_count":null,"outputs":[]},{"metadata":{"id":"DJfNzQ7qhhaX","colab_type":"code","colab":{},"trusted":true,"_uuid":"6f17d9141e4fbad83fa75cf5677624b3eff9c917"},"cell_type":"code","source":"# Test set\nempty_idx_test = []\nfor i, question in enumerate(test_df.question_text):\n    if question == '':\n        print('Empty question at index', i, 'with label ', test_df.target[i])\n        empty_idx_test.append(i)","execution_count":null,"outputs":[]},{"metadata":{"id":"FGvTLOEoihIF","colab_type":"code","colab":{},"trusted":true,"_uuid":"95957babdb7167e6396a6c922b9632e1d5aa04d0"},"cell_type":"code","source":"# Eliminating sample with empty question in the training set\ntrain_df = train_df.drop(empty_idx_train, axis=0)\nlabels = np.delete(labels, empty_idx_train)","execution_count":null,"outputs":[]},{"metadata":{"id":"-Mq3e_9vUT5x","colab_type":"text","_uuid":"3bf73c156f02056fd60b67fbb17d9f968c5222e9"},"cell_type":"markdown","source":"### Creating vocabulary\n\nIn other to tokenize the sentences, we need to create a vocabulary. I will use a dictionnaire to map word and integer. \n\nNote: the first number of the vocabulary will be a 1, as we will use 0 for the padding of the sentences."},{"metadata":{"id":"miSqPZ2lWtYN","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":34},"outputId":"2eaad21d-eb9b-47fc-d255-0865c8544156","trusted":true,"_uuid":"64502fd0bd480bb261a50d6d3b6719e4100ca44f"},"cell_type":"code","source":"# Getting all words in both samples\nall_text_list = list(train_df.question_text) + list(test_df.question_text)\nall_text = ' '.join(all_text_list)\nwords = set(all_text.split())\n\nprint('Number of unique words:', len(words))","execution_count":null,"outputs":[]},{"metadata":{"id":"7z4C5yN6XJZD","colab_type":"code","colab":{},"trusted":true,"_uuid":"ac767d5bcd4600fa554a53f22723331196257325"},"cell_type":"code","source":"# Dictionary that maps words to integers\nword_to_int = {word: i for i, word in enumerate(words, 1)}","execution_count":null,"outputs":[]},{"metadata":{"id":"YVuYzSkGXijv","colab_type":"text","_uuid":"1bd771ea24922c8d1213079e78db89f2581741d1"},"cell_type":"markdown","source":"### Tokenizing and padding"},{"metadata":{"id":"B_FynwPVYOq5","colab_type":"text","_uuid":"506db22a434a7bd321baa64652c0105a80c6263b"},"cell_type":"markdown","source":"Tokeninzing questions in both datasets:"},{"metadata":{"id":"j3iLj_hhYNbb","colab_type":"code","colab":{},"trusted":true,"_uuid":"0af44ce32e6614de2f6339b8529ba5c576e96a1b"},"cell_type":"code","source":"def tokenize(sentence):\n    '''\n    Function that tokenize a sentence using the word_to_int dictionnaire and \n    return a list of tokens\n    '''\n    tokens = []\n    for word in sentence.split():\n        tokens.append(word_to_int[word])\n    return tokens\n\n\ntrain_tokens = train_df.question_text.apply(tokenize)\ntest_tokens = test_df.question_text.apply(tokenize)","execution_count":null,"outputs":[]},{"metadata":{"id":"CKDHZIvIat9W","colab_type":"text","_uuid":"4071eb84b2af927e1fabecf727fbbafe68d6693e"},"cell_type":"markdown","source":"Checking the size of the longest question in both datasets"},{"metadata":{"id":"zbx48UvsaNlL","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":68},"outputId":"55eb7ffc-819a-4694-8290-5c4fef8de5c0","trusted":true,"_uuid":"35f4c2e13f47e9b3e91e298cbe84de3fdd83f3c0"},"cell_type":"code","source":"print('Size of longest question:')\nprint('Training set:', max(train_tokens.apply(len)))\nprint('Test set:', max(test_tokens.apply(len)))","execution_count":null,"outputs":[]},{"metadata":{"id":"qiUP2soQbTw9","colab_type":"text","_uuid":"f3b900b7d2aa0f040c48035357c4f28eddfb1cca"},"cell_type":"markdown","source":"As the questions are not too long, we can set the sequence length of the samples to the highest value and we will not need to deal with truncating. \n\nI will then pad the questions at the left using the token 0."},{"metadata":{"id":"m1uQabhramxD","colab_type":"code","colab":{},"trusted":true,"_uuid":"851fb4b47424ed69eb76b934faa56321e474df0c"},"cell_type":"code","source":"seq_length = max(train_tokens.apply(len)) # == 132","execution_count":null,"outputs":[]},{"metadata":{"id":"8qUaFnnpbukl","colab_type":"code","colab":{},"trusted":true,"_uuid":"43719e0892c6abf297cc993821e1f1a70b4d7d62"},"cell_type":"code","source":"def pad(questions, seq_length):\n    '''\n    This function pad the questions fed as series of tokens with 0 at left\n    and returns a numpy array\n    '''\n    \n    features = np.zeros((len(questions), seq_length), dtype=int)\n    for i, sentence in enumerate(questions):\n        features[i, -len(sentence):] = sentence\n    \n    return features","execution_count":null,"outputs":[]},{"metadata":{"id":"SB9Eh0bcdMLk","colab_type":"code","colab":{},"trusted":true,"_uuid":"374a77ddfd8d57142556387ec15e10fa4d58524d"},"cell_type":"code","source":"train = pad(train_tokens, seq_length)\ntest = pad(test_tokens, seq_length)","execution_count":null,"outputs":[]},{"metadata":{"id":"SoKvgVOmdWNa","colab_type":"text","_uuid":"8ef6a770c7bc4b6f730a2a3ffc7e4b4bf5a69dd5"},"cell_type":"markdown","source":"### Splitting training data and creating dataloaders"},{"metadata":{"id":"EYXCGuYGlVam","colab_type":"text","_uuid":"d7628f531bdfc7f9dacc9f2f92f267d25ce51fc9"},"cell_type":"markdown","source":"In order to avoid overfitting, we need to use some data as validation during the training phase.\n\nI am going to use a 90/10 ratio for the split of training and validation sets, in order to get a training dataset the closest as possible to the original training dataset"},{"metadata":{"id":"8TRLy3h3k2CH","colab_type":"code","colab":{},"trusted":true,"_uuid":"ac62dcc443c87ee060b40f47db8933363f0452a6"},"cell_type":"code","source":"x_train, x_val, label_train, label_val = train_test_split(train, labels, test_size=0.1, random_state=0) ","execution_count":null,"outputs":[]},{"metadata":{"id":"og2tweaDnhX2","colab_type":"text","_uuid":"a2380e9ca0021a8ba8913ae3e63d4be3e2cf27bd"},"cell_type":"markdown","source":"Now, let's create Dataloaders for the datasets that will help us with batch iteration."},{"metadata":{"id":"vLaQsod_ng5l","colab_type":"code","colab":{},"trusted":true,"_uuid":"85d5d41590029575b31ecae4e45b3644ae12a403"},"cell_type":"code","source":"# Create Tensor datasets\ntrain_data = TensorDataset(torch.from_numpy(x_train), torch.from_numpy(label_train))\nvalid_data = TensorDataset(torch.from_numpy(x_val), torch.from_numpy(label_val))\ntest_data = TensorDataset(torch.from_numpy(test))\n\n# Create Dataloaders\nbatch_size = 56\n\ntrain_loader = DataLoader(train_data, shuffle=True, batch_size=batch_size)\nvalid_loader = DataLoader(valid_data, shuffle=True, batch_size=batch_size)\ntest_loader = DataLoader(test_data, shuffle=False, batch_size=batch_size)","execution_count":null,"outputs":[]},{"metadata":{"id":"H94B9OH7o8ja","colab_type":"text","_uuid":"df8f2fe14bc39fe58593bba77a2d2db2ab13a6a0"},"cell_type":"markdown","source":"## Building the model"},{"metadata":{"id":"12g32tGpo_T3","colab_type":"text","_uuid":"9a24c55389b8ea27bc5ea363e2f21dca1d4a2120"},"cell_type":"markdown","source":"Now that we have processesed the data, I am going to build a RNN model using 1-Layer GRUs. Such RNN structure is pretty simple and it usually give good results. One can choose to use LSTM instead, I chose GRU because it will have to train less parameters while keeping good performance."},{"metadata":{"id":"mCuevokFpKXS","colab_type":"text","_uuid":"b238b08715ec5e4d01e8a8f9f976c7d8f295d5ff"},"cell_type":"markdown","source":"### Training on GPU or CPU"},{"metadata":{"id":"pN18m6aioTed","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":34},"outputId":"6ce2e259-2c99-426d-8f87-97c3b9affedb","trusted":true,"_uuid":"42c9d65e3f2c148b7a560b44f571978b99e3a1bc"},"cell_type":"code","source":"# Checking if GPU is available\ntrain_on_gpu=torch.cuda.is_available()\n\nif train_on_gpu:\n    print('Training on GPU.')\nelse:\n    print('No GPU available, training on CPU.')","execution_count":null,"outputs":[]},{"metadata":{"id":"v2AJneG-pV7g","colab_type":"text","_uuid":"74c23f5d8a5e599c161a47c7667c515ac75c431b"},"cell_type":"markdown","source":"### The model\n\nThe structure of the model is simple:\n1. Embedding layer. As said above, we are not going to use pre-trained embeddings.\n2. 1-Layer GRU\n3. Dropout Layer to avoid overfitting\n4. Fully connected layer followed by the application of a sigmoid\n5. Use the output of the last position of the setence as prediction probability"},{"metadata":{"id":"yYzc2i7Po7tG","colab_type":"code","colab":{},"trusted":true,"_uuid":"b47409c3e384d49829df8a729279f5ffec81bc33"},"cell_type":"code","source":"class RNN_model(nn.Module):\n    \"\"\"\n    The RNN model that will be used for our classification task\n    \"\"\"\n\n    def __init__(self, vocab_size, output_size, embedding_dim, hidden_dim, n_layers, drop_prob=0.5):\n        \"\"\"\n        Initialize the model by setting up the layers\n        \"\"\"\n        super(RNN_model, self).__init__()\n\n        self.output_size = output_size\n        self.n_layers = n_layers\n        self.hidden_dim = hidden_dim       \n        \n        # Embedding layer\n        self.embedding = nn.Embedding(vocab_size, embedding_dim)\n        \n        # GRU layer\n        self.gru = nn.GRU(embedding_dim, hidden_dim, n_layers, batch_first=True, dropout = drop_prob)\n        \n        # Dropout layer\n        self.dropout = nn.Dropout(p=drop_prob)\n        \n        # Fully-connected layer\n        self.fc = nn.Linear(hidden_dim, output_size)\n        \n        # Sigmoid layer\n        self.sigmoid = nn.Sigmoid()\n        \n\n    def forward(self, x, hidden):\n        \"\"\"\n        Perform a forward pass of our model on some input and hidden state.\n        \"\"\"\n        \n        batch_size = x.size(0)\n        \n        # Deal with cases were the current batch_size is different from general batch_size\n        # It occurrs at the end of iteration with the Dataloaders\n        if hidden.size(1) != batch_size:\n            hidden = hidden[:, :batch_size, :].contiguous()\n        \n        # Apply embedding\n        x = self.embedding(x)\n        \n        # GRU Layer\n        out, hidden = self.gru(x, hidden)\n        \n        # Stack up GRU outputs --> preparation for the fully-connected layer\n        out = out.contiguous().view(-1, self.hidden_dim)\n        \n        # Dropout and fully-connected layers\n        out = self.dropout(out)\n        sig_out = self.sigmoid(self.fc(out))\n        \n        # Unstack outputs to come back to correct dimensions per sample (batch_size, seq_length)\n        sig_out = sig_out.contiguous().view(batch_size, -1)\n        \n        # return last sigmoid output and hidden state\n        return sig_out[:, -1], hidden\n    \n    \n    def init_hidden(self, batch_size):\n        ''' Initializes hidden state '''\n        # Create a new tensor with sizes n_layers x batch_size x hidden_dim,\n        # initialized to zero\n        \n        weight = next(self.parameters()).data\n        \n        if train_on_gpu:\n            hidden = weight.new(self.n_layers, batch_size, self.hidden_dim).zero_().cuda()\n            \n        else:\n            hidden = weight.new(self.n_layers, batch_size, self.hidden_dim).zero_()\n        \n        return hidden\n        ","execution_count":null,"outputs":[]},{"metadata":{"id":"OeB7NH2aqWxt","colab_type":"text","_uuid":"f3c0886f91864ada945711abf39d3449a852f2cf"},"cell_type":"markdown","source":"### Defining hyperparameters and initiating the model"},{"metadata":{"id":"2es53fY8qi1C","colab_type":"code","colab":{},"trusted":true,"_uuid":"6545cf9c1c481eb807928d20de323e6eb2c6b24c"},"cell_type":"code","source":"vocab_size = len(word_to_int) + 1 # including token 0\noutput_size = 1 # binary classification task \nembedding_dim = 256\nhidden_dim = 256\nn_layers = 1\n\n# Initiating the model\nmodel = RNN_model(vocab_size, output_size, embedding_dim, hidden_dim, n_layers)","execution_count":null,"outputs":[]},{"metadata":{"id":"9TZSnvaLq5Dl","colab_type":"text","_uuid":"2b17e139a758cee06d7c3c04fcb09d4f31a64a3e"},"cell_type":"markdown","source":"## Training\n\nWe are going to use binary cross-entropy loss (BCELoss()) as loss function and the ADAM optimizer.\n\nWe are algo going to print the F1-score on valuation set each 1000 steps to follow the progress of the training. Please note that, for simplicity, we are considering a threshold of 0.5 for prediction of a positive label. This is also a hyperparameter that can be learned by cross-validation.\n\nWe are also going to clip the gradient whenever its norm is higher than 5 in order to avoid the explosion gradient effect that can happen often in RNNs."},{"metadata":{"id":"_A2Y4w7brSvt","colab_type":"code","colab":{},"trusted":true,"_uuid":"48b2dcc1de20373b523520b42c7522af1154c132"},"cell_type":"code","source":"# Training parameters\n\nepochs = 4\n\nprint_every = 1000\nclip = 5 # gradient clipping - to avoid gradient explosion\n\nlr=0.001\n\n# Defining loss and optimization functions\n\ncriterion = nn.BCELoss()\noptimizer = torch.optim.Adam(model.parameters(), lr=lr)","execution_count":null,"outputs":[]},{"metadata":{"id":"172DNky7rrdJ","colab_type":"code","colab":{"base_uri":"https://localhost:8080/","height":646},"outputId":"ad38d3b5-91b4-4026-bd65-972bb8a311bc","trusted":true,"_uuid":"77aa1d29cccd103ce293a46ab800d2afd906cfc6"},"cell_type":"code","source":"def train_model(model, train_loader, valid_loader, batch_size, epochs, optimizer, criterion, print_every, clip):\n    \n    # move model to GPU, if available\n    if(train_on_gpu):\n        model.cuda()\n    \n    counter = 0\n    \n    # Model in training mode\n    model.train()\n    breaker = False\n    for e in range(epochs):\n\n        # Batch loop\n        for inputs, labels in train_loader:\n            counter += 1\n\n            # move data to GPU, if available\n            if(train_on_gpu):\n                inputs, labels = inputs.cuda(), labels.cuda()\n\n            # Initialize hidden state\n            h = model.init_hidden(batch_size)\n\n            # Setting accumulated gradients to zero before backward step\n            model.zero_grad()\n\n            # Output from the model\n            output, _ = model(inputs, h)\n\n            # Calculate the loss and perform backprop\n            loss = criterion(output.squeeze(), labels.float())\n            loss.backward()\n\n            # Clipping the gradient to avoid explosion\n            nn.utils.clip_grad_norm_(model.parameters(), clip)\n\n            # Backpropagation step\n            optimizer.step()\n\n            # Validation stats\n            if counter % print_every == 0:\n\n                with torch.no_grad():\n\n                    # Get validation loss and F1-score on validation set\n\n                    val_losses = []\n                    all_val_labels = []\n                    all_val_preds = []\n\n                    # Model in evaluation mode\n                    model.eval()\n                    for inputs, labels in valid_loader:\n\n                        all_val_labels += list(labels)\n\n                        # Sending data to GPU\n                        if(train_on_gpu):\n                            inputs, labels = inputs.cuda(), labels.cuda()\n\n                        # Initiating hidden state for the validation set\n                        val_h = model.init_hidden(batch_size)\n\n                        output, _ = model(inputs, val_h)\n\n                        # Computing validation loss\n                        val_loss = criterion(output.squeeze(), labels.float())\n\n                        val_losses.append(val_loss.item())\n\n                        # Computing validation F1-score\n\n                        preds = torch.round(output.squeeze())  # 1 if output probability >= 0.5\n                        preds = np.squeeze(preds.numpy()) if not train_on_gpu else np.squeeze(preds.cpu().numpy())\n                        all_val_preds += list(preds)\n\n                current_loss = np.mean(val_losses)\n                \n                print(\"Epoch: {}/{}...\".format(e+1, epochs),\n                      \"Step: {}...\".format(counter),\n                      \"Loss: {:.6f}...\".format(loss.item()),\n                      \"Val Loss: {:.6f}...\".format(current_loss),\n                      \"F1-score: {:.3%}\".format(f1_score(all_val_labels, all_val_preds)))\n                \n                # Saving the best model and stopping if there is no improvement after 10 evaluations\n                \n                if counter == print_every: # first evaluation\n                    best_loss = current_loss\n                    counter_eval = 0  \n                    \n                if current_loss < best_loss:\n                    best_loss = current_loss\n                    torch.save(model.state_dict(), 'checkpoint.pth')\n                    counter_eval = 0 \n                    \n                counter_eval += 1\n                if counter_eval == 10:\n                    breaker = True\n                    break\n\n                # Put model back to training mode\n                model.train()\n        \n        # breaking outer loop on epochs\n        if breaker:\n            break\n    \n    # Loading best model\n    state_dict = torch.load('checkpoint.pth')\n    model.load_state_dict(state_dict)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"feb41187955c711f3bc8a8018399e574b39632a4"},"cell_type":"code","source":"train_model(model, train_loader, valid_loader, batch_size, epochs, optimizer, criterion, print_every, clip)","execution_count":null,"outputs":[]},{"metadata":{"id":"Tp-IwnJqtdxW","colab_type":"text","_uuid":"b2eba4e5ef794ecd32b121dbf3a7fa8af89878ae"},"cell_type":"markdown","source":"## Predictions on test set"},{"metadata":{"id":"BtbIXn2gggvK","colab_type":"code","colab":{},"trusted":true,"_uuid":"40e99d64d21183b657df46c610933d8e8ad6dda7"},"cell_type":"code","source":"# Model in evaluation mode\nmodel.eval()\n\nwith torch.no_grad():\n    all_test_preds = []\n\n    for inputs in test_loader:\n        inputs = inputs[0]\n        \n        # Sending data to GPU\n        if(train_on_gpu):\n            inputs = inputs.cuda()\n            \n        test_h = model.init_hidden(batch_size)\n        output, _ = model(inputs, test_h)\n        \n        preds = torch.round(output.squeeze())  # 1 if output probability >= 0.5\n        preds = np.squeeze(preds.numpy()) if not train_on_gpu else np.squeeze(preds.cpu().numpy())\n        all_test_preds += list(preds.astype(int))","execution_count":null,"outputs":[]},{"metadata":{"id":"k_wd_dwHhorl","colab_type":"code","colab":{},"trusted":true,"_uuid":"21d1100c117b2fef47f0f9080d3483562f4c6413"},"cell_type":"code","source":"sub = pd.DataFrame({\n    'qid': test_df.qid,\n    'prediction': all_test_preds\n})\n\n# Make sure the columns are in the correct order\nsub = sub[['qid', 'prediction']]","execution_count":null,"outputs":[]},{"metadata":{"id":"0JGmr2_SkNQZ","colab_type":"code","colab":{},"trusted":true,"_uuid":"3361d19322e160bdf7c3e6c2fbc81530bcab431a"},"cell_type":"code","source":"sub.to_csv('submission.csv', index=False, sep=',')","execution_count":null,"outputs":[]}],"metadata":{"colab":{"name":"-Kaggle-Challenge_Quora_Simple.ipynb","version":"0.3.2","provenance":[],"collapsed_sections":["ZzY2qDH930Nw"],"toc_visible":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"accelerator":"GPU","language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}