{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Description**\n\nNotebook to load the quora sincere/insincere questions data set and figure out some initial impressions.\n\nThe dataset contains the questions posted but no information about time of post or possibility to extract other posts made by the same user. So predictions will be made just based on the text content of each single question.\n\nLooking through quora is a guilty pleasure of mine, and I definitely notice the inflamatory fake looking posts. Usually I'm impressed with the time and thought the responders put in to responding about the question, and it's a huge waste of their time. I have to admit, it can be hard to tell some of the sincere but perhaps ill-informed questions from fake ones, so I think this will be a fun thing to look at!\n\nThis is my first quora kernel and a work in progress! I'll go through some simple analysis try to choose one of the suggested word embedding tools and work on incorporating that into the kernel.","metadata":{"_uuid":"e2694e753fd4b6af60f0a663f385699f145adb42"}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nprint(os.listdir(\"../input\"))\n\n## Adding in some more useful packages here\nimport matplotlib.pyplot as plt\n\nfrom wordcloud import WordCloud\nfrom nltk.corpus import stopwords\nstop = set(stopwords.words('english'))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-06-07T10:45:39.366877Z","iopub.execute_input":"2023-06-07T10:45:39.367336Z","iopub.status.idle":"2023-06-07T10:45:40.368219Z","shell.execute_reply.started":"2023-06-07T10:45:39.367266Z","shell.execute_reply":"2023-06-07T10:45:40.367370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll start by loading the training data and taking a look at the first few entries","metadata":{"_uuid":"7fffed58382952cfead158fdd6fb43d1811734d5"}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/train.csv')","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.status.busy":"2023-06-07T10:45:40.372117Z","iopub.execute_input":"2023-06-07T10:45:40.372377Z","iopub.status.idle":"2023-06-07T10:45:44.300441Z","shell.execute_reply.started":"2023-06-07T10:45:40.372329Z","shell.execute_reply":"2023-06-07T10:45:44.299500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"_uuid":"61524e90b23a760dd4e2649b6f59c3ccdf4be4b8","execution":{"iopub.status.busy":"2023-06-07T10:45:44.302673Z","iopub.execute_input":"2023-06-07T10:45:44.302985Z","iopub.status.idle":"2023-06-07T10:45:44.333080Z","shell.execute_reply.started":"2023-06-07T10:45:44.302920Z","shell.execute_reply":"2023-06-07T10:45:44.332256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This dataset contains just 3 values for each of the categories: \n\n    'qid' \n    'question text' \n    'target' value of 0 or 1\n\nThe first thing I see is that all of the entries we see so fat have a target value of 0 - so they are categorized as sincere. \n\nLets see what fraction of the data is labeled insincere (has target value 1) and take a look at some example data that was categorized as insincere.","metadata":{"_uuid":"9cf346f149b90d8780828e1942b80095b7e32d3d"}},{"cell_type":"code","source":"fig,ax = plt.subplots(1,1)\ntrain.hist(column = 'target', ax = ax)\nax.set_title('Number of entries classified as sincere vs insincere')\nax.set_xticks([0,1])\nprint('Percent of insincere entries %.3f %%'%(100*(sum(train['target'])/len(train))))\n\ntrain[train['target']==1].head()","metadata":{"_uuid":"4b6f1f2f8d9fb985b9f55996ee89ba460d4eac81","execution":{"iopub.status.busy":"2023-06-07T10:45:44.335553Z","iopub.execute_input":"2023-06-07T10:45:44.336060Z","iopub.status.idle":"2023-06-07T10:45:44.741706Z","shell.execute_reply.started":"2023-06-07T10:45:44.336002Z","shell.execute_reply":"2023-06-07T10:45:44.740799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only a little over 6% of the questions are insincere.\n\nAt first glance, in the 5 displayed insincere titles, there some words related to race and politics. Some sincere questions could also have these words but maybe some top words could be used as initial flags for insincere posts that warrant further examination.\n_______________________________________________________________________________________________________________________________\n* **Next step** - start parsing through the word data and find the initial trends in word usage amongst the insincere posts.\nI'll use a word cloud to visualize top words used in the questions. To speed up processing I'll choose just 1000 posts of each category\nI'm using the standard stop words from nltk.corpus (loaded above).","metadata":{"_uuid":"c400fb68001fe163b81353fd1a7c12a60028960e"}},{"cell_type":"code","source":"n_posts = 1000\nq_S = ' '.join(train[train['target'] == 0]['question_text'].str.lower().values[:n_posts])\nq_I = ' '.join(train[train['target'] == 1]['question_text'].str.lower().values[:n_posts])\n\nwordcloud_S = WordCloud(max_font_size=None, stopwords=stop,scale = 2,colormap = 'Dark2').generate(q_S)\nwordcloud_I = WordCloud(max_font_size=None, stopwords=stop,scale = 2,colormap = 'Dark2').generate(q_I)\n\nfig, ax = plt.subplots(1,2, figsize=(20, 5))\nax[0].imshow(wordcloud_S)\nax[0].set_title('Top words sincere posts',fontsize = 20)\nax[0].axis(\"off\")\n\nax[1].imshow(wordcloud_I)\nax[1].set_title('Top words INsincere posts',fontsize = 20)\nax[1].axis(\"off\")\n\nplt.show()","metadata":{"_uuid":"75b150187931b446aa91b68826acc3c1a4f0b85b","execution":{"iopub.status.busy":"2023-06-07T10:45:44.742954Z","iopub.execute_input":"2023-06-07T10:45:44.743429Z","iopub.status.idle":"2023-06-07T10:45:46.763257Z","shell.execute_reply.started":"2023-06-07T10:45:44.743377Z","shell.execute_reply":"2023-06-07T10:45:46.762470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There certainly looks to be a difference in the words used in 'insincere' posts, but again many of these words can appear in legitimate questions as well.\n\n____________________________________________________________________________________________________________________________________________________________\n\n**Embeddings** I'll move on to loading and using the embeddings tools.\n\nThere are 4 options with links provided in the dataset description.\nI chose to start with 'glove' This is new stuff for me, so I used https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html and https://medium.com/@japneet121/word-vectorization-using-glove-76919685ee0b as references and modified their code to work with this dataset (other referenced to be added).\n\nGloVe is feature description dataset built on a large corpus of words that represent words based on their co-occcurence with other words. In the file provided, each line lists one word that is followed by a vector of numbers that represents the word.\n\nFirst we read in the embeddings file into a dictionary - each entry is a word, followed by the vector of numbers to represent its values\n\n\n\n","metadata":{"_uuid":"99fabd761b0ca932a907de55b09989603a62e7a9"}},{"cell_type":"code","source":"embeddings_index = {}\nf = open('/kaggle/input/glove-embeddings/glove.6B.300d.txt')\nfor line in f:\n    values = line.split(' ')\n    word = values[0] ## The first entry is the word\n    coefs = np.asarray(values[1:], dtype='float32') ## These are the vecotrs representing the embedding for the word\n    embeddings_index[word] = coefs\nf.close()\n\nprint('GloVe data loaded')","metadata":{"_uuid":"9ab8ed5ba0c94ba9f877ce78f61fcddd8ec60ef8","execution":{"iopub.status.busy":"2023-06-07T10:45:46.764325Z","iopub.execute_input":"2023-06-07T10:45:46.764595Z","iopub.status.idle":"2023-06-07T10:46:33.131709Z","shell.execute_reply.started":"2023-06-07T10:45:46.764553Z","shell.execute_reply":"2023-06-07T10:46:33.130864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now preprocess the question text\n\nAs with the work cloud, I initially start by working on a subset of the posts.\n6% of 10,000 posts will be 600, so I won't go lower than total of 10k posts.\nCurrent code includes all posts in analysis","metadata":{"_uuid":"9cde0089ed5dc551754e7526f2d461fbaf93432d"}},{"cell_type":"code","source":"import re\n\n## Iterate over the data to preprocess by removing stopwords\nlines_without_stopwords=[] \nfor line in train['question_text'].values: \n    line = line.lower()\n    line_by_words = re.findall(r'(?:\\w+)', line, flags = re.UNICODE) # remove punctuation ans split\n    new_line=[]\n    for word in line_by_words:\n        if word not in stop:\n            new_line.append(word)\n    lines_without_stopwords.append(new_line)\ntexts = lines_without_stopwords\n\nprint(texts[0:5])","metadata":{"_uuid":"16bdcabe15c12557e9180fd374a1002633ae5516","execution":{"iopub.status.busy":"2023-06-07T10:46:33.132612Z","iopub.execute_input":"2023-06-07T10:46:33.132872Z","iopub.status.idle":"2023-06-07T10:46:48.216124Z","shell.execute_reply.started":"2023-06-07T10:46:33.132823Z","shell.execute_reply":"2023-06-07T10:46:48.215220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following code uses some tools from keras to 'Tokenize' the questions text - ie assign numbers to every word\n\nlabels get converted to a 2-column array (1,0) for 0 and (0,1) for 1 in the target column.\n\nThis is the first preprocessing step to be able to use the embedding.","metadata":{"_uuid":"52f6e9cb058b9994e9c0c5cd9b4d2f3324d5c77c"}},{"cell_type":"code","source":"## Code adapted from (https://github.com/keras-team/keras/blob/master/examples/pretrained_word_embeddings.py)\n# Vectorize the text samples\n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.utils import to_categorical\n\nMAX_NUM_WORDS = 1000\nMAX_SEQUENCE_LENGTH = 100\ntokenizer = Tokenizer(num_words=MAX_NUM_WORDS)\ntokenizer.fit_on_texts(texts)\nsequences = tokenizer.texts_to_sequences(texts)\n\nword_index = tokenizer.word_index\nprint('Found %s unique tokens.' % len(word_index))\n\ndata = pad_sequences(sequences, maxlen=MAX_SEQUENCE_LENGTH)\n\nlabels = to_categorical(np.asarray(train['target']))\nprint(data.shape)\nprint(labels.shape)","metadata":{"_uuid":"4dfff8d6729e4d307466f4ce0d201fd681ffd847","execution":{"iopub.status.busy":"2023-06-07T10:46:48.216979Z","iopub.execute_input":"2023-06-07T10:46:48.217238Z","iopub.status.idle":"2023-06-07T10:47:18.675026Z","shell.execute_reply.started":"2023-06-07T10:46:48.217191Z","shell.execute_reply":"2023-06-07T10:47:18.674193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## More code adapted from the keras reference (https://github.com/keras-team/keras/blob/master/examples/pretrained_word_embeddings.py)\n# prepare embedding matrix \nfrom keras.layers import Embedding\nfrom keras.initializers import Constant\n\n## EMBEDDING_DIM =  ## seems to need to match the embeddings_index dimension\nEMBEDDING_DIM = embeddings_index.get('a').shape[0]\nnum_words = min(MAX_NUM_WORDS, len(word_index)) + 1\nembedding_matrix = np.zeros((num_words, EMBEDDING_DIM))\nfor word, i in word_index.items():\n    if i > MAX_NUM_WORDS:\n        continue\n    embedding_vector = embeddings_index.get(word) ## This references the loaded embeddings dictionary\n    if embedding_vector is not None:\n        # words not found in embedding index will be all-zeros.\n        embedding_matrix[i] = embedding_vector\n\n# load pre-trained word embeddings into an Embedding layer\n# note that we set trainable = False so as to keep the embeddings fixed\nembedding_layer = Embedding(num_words,\n                            EMBEDDING_DIM,\n                            embeddings_initializer=Constant(embedding_matrix),\n                            input_length=MAX_SEQUENCE_LENGTH,\n                            trainable=False)\n","metadata":{"_uuid":"d231771e5658010612c2cee07d30ba9c4fb295c9","execution":{"iopub.status.busy":"2023-06-07T10:47:18.676152Z","iopub.execute_input":"2023-06-07T10:47:18.676422Z","iopub.status.idle":"2023-06-07T10:47:18.759377Z","shell.execute_reply.started":"2023-06-07T10:47:18.676372Z","shell.execute_reply":"2023-06-07T10:47:18.758229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Peeking at the embedding matrix values\nprint(embedding_matrix.shape)\nplt.plot(embedding_matrix[16])\nplt.plot(embedding_matrix[37])\nplt.plot(embedding_matrix[18])\nplt.title('example vectors')","metadata":{"_uuid":"c688f7cd8e6cd18450e36ddad9ad4e6ac7dcad57","execution":{"iopub.status.busy":"2023-06-07T10:47:18.760654Z","iopub.execute_input":"2023-06-07T10:47:18.760982Z","iopub.status.idle":"2023-06-07T10:47:19.039870Z","shell.execute_reply.started":"2023-06-07T10:47:18.760906Z","shell.execute_reply":"2023-06-07T10:47:19.038721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now try applying a keras sequential model - this isn't a great model yet - just copied over from the source\n\nI'll update it as I make it better","metadata":{"_uuid":"2c1db12a9658b5dfa6eda8195f38970fdafe1b75"}},{"cell_type":"code","source":"## Code from: https://medium.com/@sabber/classifying-yelp-review-comments-using-cnn-lstm-and-pre-trained-glove-word-embeddings-part-3-53fcea9a17fa\n## To create and visualize a model\n\nfrom keras.models import Sequential, Model\nimport tensorflow as tf\nfrom keras.layers import Dense, Flatten, LSTM, Conv1D, MaxPooling1D, Dropout, Activation, Input\n\nglove_input = Input(shape = (100,), dtype = tf.float32, name = 'gloveinput')\nglove_outputs = Embedding(num_words, 300, input_length=100, weights= [embedding_matrix], trainable=False)(glove_input)","metadata":{"_uuid":"2a289ce63aaa7ac8f99bd939dcdbec2afc2d130b","execution":{"iopub.status.busy":"2023-06-07T10:47:19.052704Z","iopub.execute_input":"2023-06-07T10:47:19.053510Z","iopub.status.idle":"2023-06-07T10:47:20.552122Z","shell.execute_reply.started":"2023-06-07T10:47:19.053418Z","shell.execute_reply":"2023-06-07T10:47:20.551381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install --upgrade tensorflow_hub","metadata":{"execution":{"iopub.status.busy":"2023-06-07T10:47:20.556090Z","iopub.execute_input":"2023-06-07T10:47:20.556375Z","iopub.status.idle":"2023-06-07T10:47:20.561977Z","shell.execute_reply.started":"2023-06-07T10:47:20.556323Z","shell.execute_reply":"2023-06-07T10:47:20.560003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow_hub as hub\nbert_preprocess = hub.KerasLayer(\"https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3\")\nbert_encoder = hub.KerasLayer(\"https://tfhub.dev/tensorflow/bert_en_uncased_L-12_H-768_A-12/4\")","metadata":{"execution":{"iopub.status.busy":"2023-06-07T10:47:20.563466Z","iopub.execute_input":"2023-06-07T10:47:20.563942Z","iopub.status.idle":"2023-06-07T10:47:20.954650Z","shell.execute_reply.started":"2023-06-07T10:47:20.563725Z","shell.execute_reply":"2023-06-07T10:47:20.952985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_input = Input(shape = (X_train.shape[0],), dtype=tf.string, name='BERT input')\npreprocessed_text = bert_preprocess(text_input)\noutputs = bert_encoder(preprocessed_text)\n\nbert_outputs = tf.keras.layers.BatchNormalization()(outputs['pooled_output'])\nbert_outputs = tf.keras.layers.Dense(768,activation='relu')(bert_outputs)\nbert_outputs = tf.keras.layers.Dense(300,activation='relu')(bert_outputs)","metadata":{"execution":{"iopub.status.busy":"2023-06-07T10:47:20.955998Z","iopub.status.idle":"2023-06-07T10:47:20.956715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"outputs = Dropout(0.2)(glove_outputs + bert_outputs)\noutputs = Conv1D(64, 5, activation='relu')(outputs)\noutputs = MaxPooling1D(pool_size=4)(outputs)\noutputs = LSTM(100)(outputs)\noutputs = Dense(2, activation='sigmoid')(outputs)","metadata":{"_uuid":"2a289ce63aaa7ac8f99bd939dcdbec2afc2d130b","execution":{"iopub.status.busy":"2023-06-07T10:47:20.958885Z","iopub.status.idle":"2023-06-07T10:47:20.959610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = Model(inputs = [glove_input, bert_input], outputs = [glove_embedding])","metadata":{"_uuid":"2a289ce63aaa7ac8f99bd939dcdbec2afc2d130b","execution":{"iopub.status.busy":"2023-06-07T10:47:20.960700Z","iopub.status.idle":"2023-06-07T10:47:20.961466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])","metadata":{"_uuid":"2a289ce63aaa7ac8f99bd939dcdbec2afc2d130b","execution":{"iopub.status.busy":"2023-06-07T10:47:20.962606Z","iopub.status.idle":"2023-06-07T10:47:20.963343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-06-07T10:47:20.964406Z","iopub.status.idle":"2023-06-07T10:47:20.965147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\n","metadata":{"execution":{"iopub.status.busy":"2023-06-07T10:47:20.966226Z","iopub.status.idle":"2023-06-07T10:47:20.966949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.utils.plot_model(model, show_shapes = True)","metadata":{"execution":{"iopub.status.busy":"2023-06-07T10:47:20.968000Z","iopub.status.idle":"2023-06-07T10:47:20.968765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data","metadata":{"execution":{"iopub.status.busy":"2023-06-07T10:47:20.969834Z","iopub.status.idle":"2023-06-07T10:47:20.970569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Fit train data\nprint(labels.shape)\nmodel.fit([data,train['question_text']], np.array(labels), validation_split=0.1, epochs = 1)","metadata":{"_uuid":"f6ca38f490cddea51a9a1826b5c4f197e9cf3f34","execution":{"iopub.status.busy":"2023-06-07T10:47:20.971629Z","iopub.status.idle":"2023-06-07T10:47:20.972336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Model visualization code adapted from: https://medium.com/@sabber/classifying-yelp-review-comments-using-cnn-lstm-and-pre-trained-glove-word-embeddings-part-3-53fcea9a17fa\n\nfrom sklearn.manifold import TSNE\n## Get weights\nembds = model.layers[0].get_weights()[0]\n## Plotting function\n## Visualize words in two dimensions \ntsne_embds = TSNE(n_components=2).fit_transform(embds)\n\nplt.plot(tsne_embds[:,0],tsne_embds[:,1],'.')","metadata":{"_uuid":"c2cb610aadab005f063d037a967f2d2df8b47af3","execution":{"iopub.status.busy":"2023-06-07T10:47:20.973399Z","iopub.status.idle":"2023-06-07T10:47:20.974120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* results visualization to be improved..","metadata":{"_uuid":"be45f8981e94ad536b9996a0a94c7b3411bb283a","trusted":true}},{"cell_type":"markdown","source":"Apply the model to predict values","metadata":{"_uuid":"a37e60f069382bed3445cca57a5ed603406dd896"}},{"cell_type":"code","source":"#test = pd.read_csv('../input/test.csv')\n#test.head()\n#pred = model.predict()\n#pred = np.round(pred)","metadata":{"_uuid":"2df0f0aa75d795afbf5f3d1bf16d388ca359cc11","execution":{"iopub.status.busy":"2023-06-07T10:47:20.975185Z","iopub.status.idle":"2023-06-07T10:47:20.975898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df = pd.DataFrame({\"qid\": test_df[\"qid\"], \"prediction\": pred})\n#df.to_csv(\"submission.csv\", index=False)","metadata":{"_uuid":"abf0bbc9e6ddd3d66c16b5abf09a5ee873f75251","execution":{"iopub.status.busy":"2023-06-07T10:47:20.976971Z","iopub.status.idle":"2023-06-07T10:47:20.977682Z"},"trusted":true},"execution_count":null,"outputs":[]}]}