{"cells":[{"metadata":{"_uuid":"9a549cb734be7cb00689ea4829bd0f852c1b5ce4"},"cell_type":"markdown","source":"**Overview**\n\nThe overview of notebook is to create a basic attention based LSTM for text classification. There are some advanced  strategies like adding embeddings to this process is yet to be done. It is still at it's raw form. This kernel follows PEP-8 convention.\n\n**Upvote this kernel if you feel it is useful, It also makes me motivated. Any comments on this kernel are welcome I will take that feedback constructively and go ahead **\n\n**Business Understanding:**\n\nAn existential problem for any major website today is how to handle toxic and divisive content. Quora wants to tackle this problem head-on to keep their platform a place where users can feel safe sharing their knowledge with the world.\n\n"},{"metadata":{"_uuid":"394bf61f67d5d94f185817336b2321c42246123e"},"cell_type":"markdown","source":"#### Import Dependencies"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\nfrom tensorflow.python.ops.rnn import bidirectional_dynamic_rnn as bi_rnn\nfrom tensorflow.contrib.rnn import BasicLSTMCell\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.sequence import pad_sequences\nfrom sklearn.utils import shuffle\nimport time\nfrom sklearn.metrics import f1_score\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"28d17971bb90aa16bc1687e74468ea1cf2d383a2"},"cell_type":"markdown","source":"**Attention model for text classification:**\n\nText classification is one of the principal tasks of machine learning. It aims to design proper algorithms to enable computers to extract features and classify texts automatically.\n\nThis classification uses an attention mechanism to learn weighting for each word. Under the setting, key words will have a higher weight, and common words will have lower weight. Therefore, the representation of texts not only considers all words, but also pays more attention to key words. Then we feed the feature vector to a softmax classifier. \n\n![Attention](https://www.researchgate.net/publication/323130660/figure/fig1/AS:593383479324672@1518485054048/Attention-based-bidirectional-RNN-structure_W840.jpg)\n\n**I have added this refernce at the bottom**"},{"metadata":{"_uuid":"da9f2f16669359e10c687f63f3600bdedf1b9571"},"cell_type":"markdown","source":"#### Template for Attention Block"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"class ABLSTM(object):\n    def __init__(self, config):\n        self.max_len = config[\"max_len\"]\n        self.hidden_size = config[\"hidden_size\"]\n        self.vocab_size = config[\"vocab_size\"]\n        self.embedding_size = config[\"embedding_size\"]\n        self.n_class = config[\"n_class\"]\n        self.learning_rate = config[\"learning_rate\"]\n\n        # placeholder\n        self.x = tf.placeholder(tf.int32, [None, self.max_len])\n        self.label = tf.placeholder(tf.int32, [None])\n        self.keep_prob = tf.placeholder(tf.float32)\n\n    def build_graph(self):\n        print(\"building graph\")\n        # Word embedding\n        embeddings_var = tf.Variable(tf.random_uniform([self.vocab_size, self.embedding_size], -1.0, 1.0),\n                                     trainable=True)\n        batch_embedded = tf.nn.embedding_lookup(embeddings_var, self.x)\n\n        rnn_outputs, _ = bi_rnn(BasicLSTMCell(self.hidden_size),\n                                BasicLSTMCell(self.hidden_size),\n                                inputs=batch_embedded,dtype=tf.float32)\n\n        fw_outputs, bw_outputs = rnn_outputs\n\n        W = tf.Variable(tf.random_normal([self.hidden_size], stddev=0.1))\n        H = fw_outputs + bw_outputs  # (batch_size, seq_len, HIDDEN_SIZE)\n        M = tf.tanh(H)  # M = tanh(H)  (batch_size, seq_len, HIDDEN_SIZE)\n\n        self.alpha = tf.nn.softmax(tf.reshape(tf.matmul(tf.reshape(M, [-1, self.hidden_size]),\n                                                        tf.reshape(W, [-1, 1])),\n                                              (-1, self.max_len)))  # batch_size x seq_len\n        r = tf.matmul(tf.transpose(H, [0, 2, 1]),\n                      tf.reshape(self.alpha, [-1, self.max_len, 1]))\n        r = tf.squeeze(r)\n        h_star = tf.tanh(r)  # (batch , HIDDEN_SIZE\n\n        h_drop = tf.nn.dropout(h_star, self.keep_prob)\n\n        # Fully connected layer（dense layer)\n        FC_W = tf.Variable(tf.truncated_normal([self.hidden_size, self.n_class], stddev=0.1))\n        FC_b = tf.Variable(tf.constant(0., shape=[self.n_class]))\n        y_hat = tf.nn.xw_plus_b(h_drop, FC_W, FC_b)\n\n        self.loss = tf.reduce_mean(\n            tf.nn.sparse_softmax_cross_entropy_with_logits(logits=y_hat, labels=self.label))\n\n        # prediction\n        self.prediction = tf.argmax(tf.nn.sigmoid(y_hat), 1)\n\n        # optimization\n        loss_to_minimize = self.loss\n        tvars = tf.trainable_variables()\n        gradients = tf.gradients(loss_to_minimize, tvars, aggregation_method=tf.AggregationMethod.EXPERIMENTAL_TREE)\n        grads, global_norm = tf.clip_by_global_norm(gradients, 1.0)\n\n        self.global_step = tf.Variable(0, name=\"global_step\", trainable=False)\n        self.optimizer = tf.train.AdamOptimizer(learning_rate=self.learning_rate)\n        self.train_op = self.optimizer.apply_gradients(zip(grads, tvars), global_step=self.global_step,\n                                                       name='train_step')\n        print(\"graph built successfully!\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"68140fb787c686c4919922ec1ae45b8f7448db98"},"cell_type":"markdown","source":"#### Template for preprocessing block:"},{"metadata":{"trusted":true,"_uuid":"d982c10a6d61012d8dcd0c79b1f599c43d165dc9"},"cell_type":"code","source":"\nnames = [\"qid\", \"question_text\", \"target\"]\n\ndef load_data(file_name, sample_ratio=1, names=names):\n    '''load data from .csv file'''\n    csv_file = pd.read_csv(file_name, names=names)\n    shuffle_csv = csv_file.sample(frac=sample_ratio)\n    return shuffle_csv[\"question_text\"], shuffle_csv[\"target\"]\n\ndef data_preprocessing(train, test, max_len):\n    \"\"\"transform to one-hot idx vector by VocabularyProcessor\"\"\"\n    \"\"\"VocabularyProcessor is deprecated, use v2 instead\"\"\"\n    vocab_processor = tf.contrib.learn.preprocessing.VocabularyProcessor(max_len)\n    x_transform_train = vocab_processor.fit_transform(train)\n    x_transform_test = vocab_processor.transform(test)\n    vocab = vocab_processor.vocabulary_\n    vocab_size = len(vocab)\n    x_train_list = list(x_transform_train)\n    x_test_list = list(x_transform_test)\n    x_train = np.array(x_train_list)\n    x_test = np.array(x_test_list)\n\n    return x_train, x_test, vocab, vocab_size\n\n\ndef data_preprocessing_v2(train, max_len, max_words=50000):\n    tokenizer = tf.keras.preprocessing.text.Tokenizer(num_words=max_words)\n    tokenizer.fit_on_texts(train)\n    train_idx = tokenizer.texts_to_sequences(train)\n    train_padded = pad_sequences(train_idx, maxlen=max_len, padding='post', truncating='post')\n    # vocab size = len(word_docs) + 2  (<UNK>, <PAD>)\n    return train_padded, max_words + 2\n\n\ndef data_preprocessing_with_dict(train, test, max_len):\n    tokenizer = tf.keras.preprocessing.text.Tokenizer(oov_token='<UNK>')\n    tokenizer.fit_on_texts(train)\n    train_idx = tokenizer.texts_to_sequences(train)\n    test_idx = tokenizer.texts_to_sequences(test)\n    train_padded = pad_sequences(train_idx, maxlen=max_len, padding='post', truncating='post')\n    test_padded = pad_sequences(test_idx, maxlen=max_len, padding='post', truncating='post')\n    # vocab size = len(word_docs) + 2  (<UNK>, <PAD>)\n    return train_padded, test_padded, tokenizer.word_docs, tokenizer.word_index, len(tokenizer.word_docs) + 2\n\n\ndef split_dataset(x_test, y_test, dev_ratio):\n    \"\"\"split test dataset to test and dev set with ratio \"\"\"\n    test_size = len(x_test)\n    print(test_size)\n    dev_size = (int)(test_size * dev_ratio)\n    print(dev_size)\n    x_dev = x_test[:dev_size]\n    x_test = x_test[dev_size:]\n    y_dev = y_test[:dev_size]\n    y_test = y_test[dev_size:]\n    return x_test, x_dev, y_test, y_dev, dev_size, test_size - dev_size\n\n\ndef fill_feed_dict(data_X, data_Y, batch_size):\n    \"\"\"Generator to yield batches\"\"\"\n    shuffled_X, shuffled_Y = shuffle(data_X, data_Y)\n    for idx in range(data_X.shape[0] // batch_size):\n        x_batch = shuffled_X[batch_size * idx: batch_size * (idx + 1)]\n        y_batch = shuffled_Y[batch_size * idx: batch_size * (idx + 1)]\n        yield x_batch, y_batch","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"07daf2d7da0c2e0518d5c3360a2cdde891054698"},"cell_type":"markdown","source":"#### Model Helper Functions"},{"metadata":{"trusted":true,"_uuid":"957d7a492d0ef42190b66eca945a0a0544666346"},"cell_type":"code","source":"def make_train_feed_dict(model, batch):\n    \"\"\"make train feed dict for training\"\"\"\n    feed_dict = {model.x: batch[0],\n                 model.label: batch[1],\n                 model.keep_prob: .5}\n    return feed_dict\n\n\ndef make_test_feed_dict(model, batch):\n    feed_dict = {model.x: batch[0],\n                 model.label: batch[1],\n                 model.keep_prob: 1.0}\n    return feed_dict\n\n\ndef run_train_step(model, sess, batch):\n    feed_dict = make_train_feed_dict(model, batch)\n    to_return = {\n        'train_op': model.train_op,\n        'loss': model.loss,\n        'global_step': model.global_step,\n    }\n    return sess.run(to_return, feed_dict)\n\n\ndef run_eval_step(model, sess, batch):\n    feed_dict = make_test_feed_dict(model, batch)\n    prediction = sess.run(model.prediction, feed_dict)\n    return prediction\n\n\ndef get_attn_weight(model, sess, batch):\n    feed_dict = make_train_feed_dict(model, batch)\n    return sess.run(model.alpha, feed_dict)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a9d0638026739175e8340a14023c2f86b4a8432"},"cell_type":"markdown","source":"#### Read the train dataset:\n\n**I am facing resource exhausted issue so at start I am using just one percent of train data**"},{"metadata":{"trusted":true,"_uuid":"33c0dd0eb0ed0be9205ea3484acd7d0751439067"},"cell_type":"code","source":"train=pd.read_csv('../input/train.csv')\nx_train, y_train = train['question_text'],train['target']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"718fe79ebae79e8c2748a0b121972545de047157"},"cell_type":"markdown","source":"**Preprocessing and Data splitting**"},{"metadata":{"trusted":true,"_uuid":"f53aef8d36570dbc8f4311373a20f21075b2b86e"},"cell_type":"code","source":"x_train, vocab_size = data_preprocessing_v2(x_train, max_len=32)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dd486b55f404081429bd402f41a80eca082e0006"},"cell_type":"code","source":"x_train, x_dev, y_train, y_dev, dev_size, train_size = split_dataset(x_train, y_train, 0.1)\nprint(\"Validation Size: \", dev_size)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6b3003a1577f00ed2378355c42e52395614dea09"},"cell_type":"markdown","source":"**Hyper Parameters**"},{"metadata":{"trusted":true,"_uuid":"80af7b55cb874a3617f677e42c47c0b0f8891352"},"cell_type":"code","source":"config = {\n        \"max_len\": 32,\n        \"hidden_size\": 64,\n        \"vocab_size\": vocab_size,\n        \"embedding_size\": 128,\n        \"n_class\": 2,\n        \"learning_rate\": 1e-3,\n        \"batch_size\": 32,\n        \"train_epoch\": 2\n}","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d727e47bcc0365f885c88c3c377d143e088cbb1e"},"cell_type":"markdown","source":"**Tensorflow session**"},{"metadata":{"trusted":true,"_uuid":"4028cf32777de8be2f7d104c01d8faae12535bad"},"cell_type":"code","source":"tf.reset_default_graph()\nclassifier = ABLSTM(config)\nclassifier.build_graph()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7eb2813ade2869b577bcdf73be1b9ab720a46354"},"cell_type":"code","source":"sess = tf.Session()\nsess.run(tf.global_variables_initializer())\ndev_batch = (x_dev, y_dev)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fdb116a79b56258abad21cde686e50ede95dfeef"},"cell_type":"markdown","source":"**Training**"},{"metadata":{"trusted":true,"_uuid":"317ebac7eff55d810bd00c865ea11d2598cdc6ea"},"cell_type":"code","source":"start = time.time()\npredictions=[]\nfor e in range(config[\"train_epoch\"]):\n   t0 = time.time()\n   print(\"Epoch %d start !\" % (e + 1))\n   for x_batch, y_batch in fill_feed_dict(x_train, y_train, config[\"batch_size\"]):\n    return_dict = run_train_step(classifier, sess, (x_batch, y_batch))\n    attn = get_attn_weight(classifier, sess, (x_batch, y_batch))\n   t1 = time.time()\n   print(\"Train Epoch time:  %.3f s\" % (t1 - t0))\n   dev_preds = run_eval_step(classifier, sess, dev_batch)\n   print(\"validation F1: %.3f \" % f1_score(y_dev.values,dev_preds))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9ace885e6058bf4df60c50d73ac6df2a81b9ee56"},"cell_type":"markdown","source":"**Submission**"},{"metadata":{"trusted":true,"_uuid":"63c34b9e2bca8058819e8bcbd5d0febf2a484f42"},"cell_type":"code","source":"test_df=pd.read_csv('../input/test.csv')\n\ntarget=[0 for i in range(test_df.shape[0])]\ntest_df['target']=target","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9f640e70a2d3937a3281ebbc74135650c3349d65"},"cell_type":"code","source":"x_text, vocab_size = data_preprocessing_v2(test_df['question_text'], max_len=32)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b6ea164066cfb77c2ff19215b620dd30bbdaedf9"},"cell_type":"code","source":"test_batch=(x_text,test_df['target'].values)\ntest_preds = run_eval_step(classifier, sess, test_batch)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5de1402185a82d03566b8261da8d761bbf77d9b0"},"cell_type":"code","source":"sub=pd.read_csv('../input/sample_submission.csv')\nsub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4ad42aa4891358abb127b3641ef45fe34f39c170"},"cell_type":"code","source":"sub.prediction=test_preds\nsub.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"da6f5cd65bd94f3f0669c314d5a749c8667e5210"},"cell_type":"markdown","source":"### References:\n\nhttps://github.com/TobiasLee/Text-Classification/blob/master/models/attn_bi_lstm.py\n\nhttp://univagora.ro/jour/index.php/ijccc/article/download/3142/pdf"},{"metadata":{"trusted":true,"_uuid":"5a2ee52e385cebe0540efa13b3af936f10a06b0f"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"97d0c9a00fe0fb4cc28ef78e08435803df69c4b6"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}