{"cells":[{"metadata":{"_uuid":"521178d5e7c0e427d5aff12df5d6b64aacc495ec"},"cell_type":"markdown","source":"# About the dataset\nAn existential problem for any major website today is how to handle toxic and divisive content. Quora wants to tackle this problem head-on to keep their platform a place where users can feel safe sharing their knowledge with the world.\n\nQuora is a platform that empowers people to learn from each other. On Quora, people can ask questions and connect with others who contribute unique insights and quality answers. A key challenge is to weed out insincere questions -- those founded upon false premises, or that intend to make a statement rather than look for helpful answers.\n\nIn this competition, Kagglers will develop models that identify and flag insincere questions. To date, Quora has employed both machine learning and manual review to address this problem. With your help, they can develop more scalable methods to detect toxic and misleading content.\n\nHere's your chance to combat online trolls at scale. Help Quora uphold their policy of “Be Nice, Be Respectful” and continue to be a place for sharing and growing the world’s knowledge."},{"metadata":{"_uuid":"f655f1d9fbfcfd335f6bff2a0c4cffe4d540877f"},"cell_type":"markdown","source":"This kernel is highly inspired from the following article: [https://machinelearningmastery.com/develop-n-gram-multichannel-convolutional-neural-network-sentiment-analysis/](https://machinelearningmastery.com/develop-n-gram-multichannel-convolutional-neural-network-sentiment-analysis/)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Usual imports\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport string\nimport matplotlib.pyplot as plt\nfrom sklearn.decomposition import NMF, LatentDirichletAllocation, TruncatedSVD\nfrom sklearn.feature_extraction.text import CountVectorizer\nimport concurrent.futures\nimport time\nimport pyLDAvis.sklearn\nfrom pylab import bone, pcolor, colorbar, plot, show, rcParams, savefig\nimport warnings\nwarnings.filterwarnings('ignore')\n\n%matplotlib inline\nimport os\nprint(os.listdir(\"../input\"))\nprint(os.listdir(\"../input/embeddings\"))\n\n# Plotly based imports for visualization\nfrom plotly import tools\nimport plotly.plotly as py\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.figure_factory as ff\n\n# spaCy based imports\nimport spacy\nfrom spacy.lang.en.stop_words import STOP_WORDS\nfrom spacy.lang.en import English\n\n# Keras based imports\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.utils.vis_utils import plot_model\nfrom keras.models import Model\nfrom keras.layers import Input\nfrom keras.layers import Dense\nfrom keras.layers import Flatten\nfrom keras.layers import Dropout\nfrom keras.layers import Embedding\nfrom keras.layers import CuDNNLSTM, CuDNNGRU, Bidirectional\nfrom keras.layers.convolutional import Conv1D\nfrom keras.layers.convolutional import MaxPooling1D\nfrom keras.layers.merge import concatenate","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0e2da97e5cd6e24ce093f3e6ccfa5bb514b5fb49"},"cell_type":"code","source":"quora_train = pd.read_csv(\"../input/train.csv\")\nquora_test = pd.read_csv(\"../input/test.csv\")\nquora_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"239f186671cd2dcea82cf3e34b994eb29fab5279"},"cell_type":"code","source":"punctuations = string.punctuation\n\ndef punct_remover(my_str):\n    my_str = my_str.lower()\n    no_punct = \"\"\n    for char in my_str:\n       if char not in punctuations:\n           no_punct = no_punct + char\n    return no_punct\n\npunctuations","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"57656d248a9e4c3972c812c61d68c8d215704732"},"cell_type":"code","source":"tqdm.pandas()\nquestions = quora_train[\"question_text\"].progress_apply(punct_remover)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8f951b4a7eb99d1d2bc2707108e0debc4539229e"},"cell_type":"code","source":"test = quora_test[\"question_text\"].progress_apply(punct_remover)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fe74e1bf021317562c45ac9b2a8b4c8e685216e2"},"cell_type":"code","source":"# fit a tokenizer\ndef create_tokenizer(lines):\n    tokenizer = Tokenizer()\n    tokenizer.fit_on_texts(lines)\n    return tokenizer\n\n# calculate the maximum document length\ndef max_length(lines):\n    return max([len(s.split()) for s in lines])\n \n# encode a list of lines\ndef encode_text(tokenizer, lines, length):\n    # integer encode\n    encoded = tokenizer.texts_to_sequences(lines)\n    # pad encoded sequences\n    padded = pad_sequences(encoded, maxlen=length, padding='post')\n    return padded","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ffcd6b61fcf221f9f952fab6d6a149beb2b14d55"},"cell_type":"code","source":"embeddings_index = dict()\nf = open('../input/embeddings/glove.840B.300d/glove.840B.300d.txt',encoding='utf8')\nfor line in f:\n    values = line.split(\" \")\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs\nf.close()\n\nembed_token = create_tokenizer(questions)\nvocabulary_size = 90000\n\nembedding_matrix = np.zeros((vocabulary_size, 300))\nfor word, index in embed_token.word_index.items():\n    if index > vocabulary_size - 1:\n        break\n    else:\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None:\n            embedding_matrix[index] = embedding_vector","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0d61a0b77d0e41743ab18e8910152d5a4107aafa"},"cell_type":"code","source":"# define the model\ndef define_model(length, vocab_size):\n    # channel 1\n    inputs1 = Input(shape=(length,))\n    embedding1 = Embedding(vocabulary_size, 300, weights=[embedding_matrix])(inputs1)\n    conv1 = Conv1D(filters=16, kernel_size=4, activation='relu')(embedding1)\n    drop1 = Dropout(0.5)(conv1)\n    lstm1 = Bidirectional(CuDNNLSTM(10, return_sequences = True))(drop1)\n    gru1 = Bidirectional(CuDNNGRU(10, return_sequences = True))(lstm1)\n    pool1 = MaxPooling1D(pool_size=2)(gru1)\n    flat1 = Flatten()(pool1)\n    # channel 2\n    inputs2 = Input(shape=(length,))\n    embedding2 = Embedding(vocabulary_size, 300, weights=[embedding_matrix])(inputs2)\n    conv2 = Conv1D(filters=16, kernel_size=6, activation='relu')(embedding2)\n    drop2 = Dropout(0.5)(conv2)\n    lstm2 = Bidirectional(CuDNNLSTM(10, return_sequences = True))(drop2)\n    gru2 = Bidirectional(CuDNNLSTM(10, return_sequences = True))(lstm2)\n    pool2 = MaxPooling1D(pool_size=2)(gru2)\n    flat2 = Flatten()(pool2)\n    # channel 3\n    inputs3 = Input(shape=(length,))\n    embedding3 = Embedding(vocabulary_size, 300, weights=[embedding_matrix])(inputs3)\n    conv3 = Conv1D(filters=16, kernel_size=8, activation='relu')(embedding3)\n    drop3 = Dropout(0.5)(conv3)\n    lstm3 = Bidirectional(CuDNNLSTM(10, return_sequences = True))(drop3)\n    gru3 = Bidirectional(CuDNNGRU(10, return_sequences = True))(lstm3)\n    pool3 = MaxPooling1D(pool_size=2)(gru3)\n    flat3 = Flatten()(pool3)\n    # merge\n    merged = concatenate([flat1, flat2, flat3])\n    # interpretation\n    dense1 = Dense(10, activation='relu')(merged)\n    outputs = Dense(1, activation='sigmoid')(dense1)\n    model = Model(inputs=[inputs1, inputs2, inputs3], outputs=outputs)\n    # compile\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    # summarize\n    print(model.summary())\n    plot_model(model, show_shapes=True, to_file='multichannel.png')\n    return model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e70dfaa1214fd461b5723e1251b16f6d5814c407"},"cell_type":"code","source":"%%time\n# Preprocess data\n\n# create tokenizer\ntokenizer = create_tokenizer(questions)\n# calculate max document length\nlength = max_length(questions)\n# calculate vocabulary size\nvocab_size = len(tokenizer.word_index) + 1\nprint('Max document length: %d' % length)\nprint('Vocabulary size: %d' % vocab_size)\n# encode data\ntrainX = encode_text(tokenizer, questions, length)\ntestX = encode_text(tokenizer, test, length)\nprint(trainX.shape, testX.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a19351de02bb70d36d64b0f8afcfb724e16d13c0"},"cell_type":"code","source":"# define model\nmodel = define_model(length, vocab_size)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"de32673c94d5f2fbb08b99286c9ba679c6462581"},"cell_type":"code","source":"# fit model\nmodel.fit([trainX,trainX,trainX], quora_train[\"target\"].values, epochs=5, batch_size=4096)\n\n# save the model\nmodel.save('model.h5')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f38da68a59617d7038fb2c36f81234f819591a02"},"cell_type":"code","source":"preds = model.predict([testX,testX,testX])\npreds = (preds[:,0] > 0.5).astype(np.int)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cf0173678838131f1b6e3114871520fed7eaba37"},"cell_type":"code","source":"submission = pd.DataFrame.from_dict({'qid': quora_test['qid']})\nsubmission['prediction'] = preds\nsubmission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}