{"cells":[{"metadata":{"_uuid":"31371f2922505cff97f8526260f809d169d0ae30"},"cell_type":"markdown","source":"In this kernel we try different models to complete a [Quora Toxic Comment Classification Challenge](https://www.kaggle.com/c/quora-insincere-questions-classification).  \n\nIn this competition, Kagglers will develop models that identify and flag insincere questions. To date, Quora has employed both machine learning and manual review to address this problem. we develop more scalable methods to detect toxic and misleading content. Here we try various models inside this kernel.\n\n# Reference\n\n"},{"metadata":{"_uuid":"24e1642f6d4dabecca6b0724e18880f9f1a3f3cf"},"cell_type":"markdown","source":"# Importing Python Library"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"from keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\n\nimport os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nimport math\nfrom sklearn.model_selection import train_test_split\n\n# training library\nfrom keras.models import Sequential\nfrom keras.layers import CuDNNLSTM, Dense, Bidirectional\nfrom sklearn.metrics import f1_score","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"21f85dfd5b9bf3d8b27ba29149d52253e5d64049"},"cell_type":"markdown","source":"# Setup"},{"metadata":{"trusted":true,"_uuid":"78578eab64a477d0a5ad6b1c917ae154868a44df"},"cell_type":"code","source":"# load training data\ntrain_df = pd.read_csv(\"../input/train.csv\")\n# slipt data into train data frame and test data frame\ntrain_df, test_df = train_test_split(train_df, test_size=0.2)\n# slipt training into train data frame and validation data frame\ntrain_df, val_df = train_test_split(train_df, test_size=0.2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"92ffbf2ef35d2dc5863ee27c61eadd1869ca5440"},"cell_type":"code","source":"# embdedding setup\n# Source https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html\nembeddings_index = {}\nf = open('../input/embeddings/glove.840B.300d/glove.840B.300d.txt')\nfor line in tqdm(f):\n    values = line.split(\" \")\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs\nf.close()\n\n# print('Found %s word vectors.' % len(embeddings_index))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"728b0b913396d70fda00b789a9dc4ce0ce361b11"},"cell_type":"markdown","source":"convert question text sequence into embedding and target columns"},{"metadata":{"trusted":true,"_uuid":"9cc8b0bcd285225ce651d024f506615d4656b9f7"},"cell_type":"code","source":"# Convert values to embeddings\ndef text_to_array(text):\n    empyt_emb = np.zeros(300) # a list of zeros\n    text = text[:-1].split()[:30] # split the string by space and take the first 30 chars\n    embeds = [embeddings_index.get(x, empyt_emb) for x in text] \n    embeds+= [empyt_emb] * (30 - len(embeds))\n    return np.array(embeds)\n\n\n# validation set\nval_vects = np.array([text_to_array(X_text) for X_text in tqdm(val_df[\"question_text\"][:3000])])\nval_y = np.array(val_df[\"target\"][:3000])\n\n# training set\ntrain_vects = np.array([text_to_array(X_text) for X_text in tqdm(train_df[\"question_text\"][:3000])])\ntrain_y = np.array(train_df[\"target\"][:3000])\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0c950448e0717eaebb920d93cfdc6b4561e21853"},"cell_type":"code","source":"# Data providers\nbatch_size = 128\n\n# create an generator for the training set\ndef batch_gen(train_df):\n    # number of batches\n    n_batches = math.ceil(len(train_df) / batch_size)\n    while True: \n        train_df = train_df.sample(frac=1.)  # Shuffle the data.\n        for i in range(n_batches):\n            texts = train_df.iloc[i*batch_size:(i+1)*batch_size, 1]\n            text_arr = np.array([text_to_array(text) for text in texts]) \n            yield text_arr, np.array(train_df[\"target\"][i*batch_size:(i+1)*batch_size]) # generator: embbeding and target\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"706c0224b6112e5a8f00ad25f35e90fbb9519a5f"},"cell_type":"markdown","source":"# Baseline Training\nA very simple network to be improved"},{"metadata":{"trusted":true,"_uuid":"9ad418c95b31ab5691d49b72c7c9622ef9ea42cf"},"cell_type":"code","source":"# building nerual network layers\nmodel = Sequential()\nmodel.add(Bidirectional(CuDNNLSTM(64, return_sequences=True),\n                        input_shape=(30, 300)))\nmodel.add(Bidirectional(CuDNNLSTM(64)))\nmodel.add(Dense(1, activation=\"sigmoid\"))\n\nmodel.compile(loss='binary_crossentropy',\n              optimizer='adam',\n              metrics=['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3f0b703ded691293450beb4ddf2d903d0e08c737"},"cell_type":"code","source":"# mg is a combination of word embbeding and target value. \nmg = batch_gen(train_df)\n# training on mg, tuples of embbeding and target\nmodel.fit_generator(mg, epochs=20,\n                    steps_per_epoch=1000,\n                    validation_data=(val_vects, val_y),\n                    verbose=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f1c7228a91deb969c79c58271d947c1b847fb33d"},"cell_type":"code","source":"batch_size = 256\ndef batch_gen(test_df):\n    n_batches = math.ceil(len(test_df) / batch_size)\n    for i in range(n_batches):\n        texts = test_df.iloc[i*batch_size:(i+1)*batch_size, 1]\n        text_arr = np.array([text_to_array(text) for text in texts])\n        yield text_arr\n\n# making prediction using test data frame\nall_preds = []\nfor x in tqdm(batch_gen(test_df)):\n    all_preds.extend(model.predict(x).flatten())\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5154a620ce757521c539e6756cc86fae9c3f08f9"},"cell_type":"code","source":"# try various thretholds and print the f1 score using the prediction. \nfor thresh in np.arange(0.1, 0.501, 0.01):\n    thresh = np.round(thresh, 2)\n    test_y = np.array(test_df[\"target\"]) # ground truth\n    print(\"F1 score at threshold {0} is {1}\".format(thresh, f1_score(test_y, (np.array(all_preds)>thresh).astype(int))))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b18906f81600a2517152c877accaa546770f8c1"},"cell_type":"markdown","source":"**Repeat the same pattern above, but try the following**\n1.  add more network layers\n2. change training settings \n3. combine other kernels, if find anything useful\n4. try different test set\n\n----Di"},{"metadata":{"_uuid":"d6cdda0e301be0b84e826e29f0f3a85c41aa6da9"},"cell_type":"markdown","source":"# Submission: final output \nwe use the best trained network to predict the submission file.\nLet's ignore this for now."},{"metadata":{"trusted":true,"_uuid":"b58cd95254f41e5002a17de0c3feab54a5fc3c67"},"cell_type":"code","source":"# prediction part\n# batch_size = 256\n# def batch_gen(test_df):\n#     n_batches = math.ceil(len(test_df) / batch_size)\n#     for i in range(n_batches):\n#         texts = test_df.iloc[i*batch_size:(i+1)*batch_size, 1]\n#         text_arr = np.array([text_to_array(text) for text in texts])\n#         yield text_arr\n\n# test_df = pd.read_csv(\"../input/test.csv\")\n\n# all_preds = []\n# for x in tqdm(batch_gen(test_df)):\n#     all_preds.extend(model.predict(x).flatten())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e6ed54def110c881f401c6f8a844752baedcfbd"},"cell_type":"code","source":"# y_te = (np.array(all_preds) > 0.5).astype(np.int)\n\n# submit_df = pd.DataFrame({\"qid\": test_df[\"qid\"], \"prediction\": y_te})\n# submit_df.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"809b70a1ce8cabab7a505c409191eeafeab43301"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}