{"cells":[{"metadata":{"_uuid":"6a008194c547387dc17116ce8edeb52a62a84f6a"},"cell_type":"markdown","source":"**Introduction** \n\nThe aim of this notebook is to leverage insights from several public kernels to eventually formalize a workflow for NLP beginners on kaggle, like me.  Any comments, recommendations and insights would thus be very appreciated!\n\n**Credits**\n\nEDA code was heavily sourced from: \n \nhttps://www.kaggle.com/arunsankar/key-insights-from-quora-insincere-questions\n\nData Pre-processing code adapted from:\n\nhttps://www.kaggle.com/enerrio/scary-nlp-with-spacy-and-keras\n\nWord Embedding model was heavily adapted from:\n\nhttps://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\n\nLSTM architecture was adapted from:\n\nhttps://www.kaggle.com/mihaskalic/lstm-is-all-you-need-well-maybe-embeddings-also\n\n\n**Analysis Sections**\n* Data Understanding\n    * EDA plot 1 - Word Cloud\n    * EDA plot 2 - side by side plot comparison using N-gram\n    * EDA Plot 3 - Word count distribution, Character Length Distribution, Stop words, Punctuation, Upper case\n    * Overview of EDA results\n* Data Preparation \n* Modelling\n* Evaluation\n\n**Other helpful kernels and links**\n\n*Kernels* :\n\n* https://www.kaggle.com/mjbahmani/a-data-science-framework-for-quora\n* https://www.kaggle.com/shujian/test-the-difficulty-of-this-classification-tasks\n* https://www.kaggle.com/enerrio/scary-nlp-with-spacy-and-keras\n* https://www.kaggle.com/wakamezake/visualizing-word-vectors\n\n*Links*\n* https://www.kdnuggets.com/2017/02/natural-language-processing-key-terms-explained.html\n\n**Next Steps**\n\n* Investigate how dimensionality reduction techniques affect the model\n* Implement oversampling to cater for unbalanced positive classes\n\n**Thanks for reading!**\n"},{"metadata":{"_uuid":"df4071ea75b66abdd5a84824878893e0932a4e8d"},"cell_type":"markdown","source":"# **Data Understanding**\n \nIn order to start the anlaysis, the nature of the datasets and the files provided will first need to be examined. This normally involves examining the shape of the datasets,  followed by an EDA, with the objective to understand what defines a question as insincere or sincere.\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Input data files are available in the \"../input/\" directory.\n# For example, running this will list the files in the input directory\nimport os\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"792011fe9c2a66b547d0080ea01d406ff11e5861"},"cell_type":"markdown","source":"Word Embeddings allow words that are used in similar ways to result in having similar vector representations, naturally capturing their meaning.\n\nFurther details at: https://machinelearningmastery.com/what-are-word-embeddings/"},{"metadata":{"trusted":true,"_uuid":"3360d149d30554ecb057f9713738c17680214d41"},"cell_type":"code","source":"#Verify which embeddings are provided\n!ls ../input/embeddings\n#there are 4 embeddings provided with the dataset","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"#import packages\nimport numpy as np \nimport pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport random\nimport spacy\nimport nltk\nfrom nltk.tokenize.toktok import ToktokTokenizer\nimport re\nfrom bs4 import BeautifulSoup\nimport unicodedata\nfrom collections import defaultdict\nimport string\n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\n\nfrom sklearn.model_selection import train_test_split\n\nfrom nltk.corpus import stopwords\nfrom sklearn.metrics import log_loss\nfrom tqdm import tqdm\nstopwords = stopwords.words('english')\nsns.set_context('notebook')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"88e591ca6c42a1d9e4d9e45c7e291855c3b53658"},"cell_type":"code","source":"#Code adapted from: https://www.kaggle.com/arunsankar/key-insights-from-quora-insincere-questions\n#import the different datasets and print the characteristics of each\ntrain = pd.read_csv('../input/train.csv')\ntest = pd.read_csv('../input/test.csv')\nsub = pd.read_csv('../input/sample_submission.csv')\n\n#Print the different statistics of the different files\nprint('Train data: \\nRows: {}\\nCols: {}'.format(train.shape[0],train.shape[1]))\nprint(train.columns)\n\nprint('\\nTest data: \\nRows: {}\\nCols: {}'.format(test.shape[0],test.shape[1]))\nprint(test.columns)\n\nprint('\\nSubmission data: \\nRows: {}\\nCols: {}'.format(sub.shape[0],sub.shape[1]))\nprint(sub.columns)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b566a970e62c24807b8928138f3f4861575ef15a"},"cell_type":"code","source":"#View the first 5 entries of the training data\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"257ed835e13d45c88ad8a99182bed26d9af107ee"},"cell_type":"code","source":"#View information about the train dataset\ntrain.info()\n##1306122 observations and 3 columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f0dcb6104bc8407f1256c02440ffcb3a47e9060a"},"cell_type":"code","source":"#check for the number of positive and negative classes\npd.crosstab(index = train.target, columns = \"count\" )\n#There seems to be unbalanced classes in the dataset [first issue]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"07576a1448ac4e9e73789272996384ee205d809c"},"cell_type":"markdown","source":"**EDA plot 1 - Word Cloud**\n\nWord clouds can identify trends and patterns that would otherwise be unclear or difficult to see in a tabular format. Frequently used keywords stand out better in a word cloud. Common words that might be overlooked in tabular form are highlighted in larger text making them pop out when displayed in a word cloud."},{"metadata":{"trusted":true,"_uuid":"85caaf1086b2c752cd0e3b8cf5bac1fb00e4846a"},"cell_type":"code","source":"#Code sourced from : https://www.kaggle.com/sudalairajkumar/simple-exploration-notebook-qiqc\n\n#import the wordcloud package\nfrom wordcloud import WordCloud, STOPWORDS\n\n#Define the word cloud function with a max of 200 words\ndef plot_wordcloud(text, mask=None, max_words=200, max_font_size=100, figure_size=(24.0,16.0), \n                   title = None, title_size=40, image_color=False):\n    stopwords = set(STOPWORDS)\n    #define additional stop words that are not contained in the dictionary\n    more_stopwords = {'one', 'br', 'Po', 'th', 'sayi', 'fo', 'Unknown'}\n    stopwords = stopwords.union(more_stopwords)\n    #Generate the word cloud\n    wordcloud = WordCloud(background_color='black',\n                    stopwords = stopwords,\n                    max_words = max_words,\n                    max_font_size = max_font_size, \n                    random_state = 42,\n                    width=800, \n                    height=400,\n                    mask = mask)\n    wordcloud.generate(str(text))\n    #set the plot parameters\n    plt.figure(figsize=figure_size)\n    if image_color:\n        image_colors = ImageColorGenerator(mask);\n        plt.imshow(wordcloud.recolor(color_func=image_colors), interpolation=\"bilinear\");\n        plt.title(title, fontdict={'size': title_size,  \n                                  'verticalalignment': 'bottom'})\n    else:\n        plt.imshow(wordcloud);\n        plt.title(title, fontdict={'size': title_size, 'color': 'black', \n                                  'verticalalignment': 'bottom'})\n    plt.axis('off');\n    plt.tight_layout()  ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"38739d1534833e955a5cf3864751922c5624e76f"},"cell_type":"code","source":"#Select insincere questions from training dataset\ninsincere = train.loc[train['target'] == 1]\n#run the function on the insincere questions\nplot_wordcloud(insincere[\"question_text\"], title=\"Word Cloud of Insincere Questions\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0f7cc5643c46e1add2b0d82c31ec2bfc937bcb2a"},"cell_type":"code","source":"#Select sincere questions from training dataset\nsincere = train.loc[train['target'] == 0]\n#run the function on the insincere questions\nplot_wordcloud(sincere[\"question_text\"], title=\"Word Cloud of Sincere Questions\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"35dcd6fdad32f164bdde1f0251eccf79ff45cb2b"},"cell_type":"markdown","source":"**EDA plot 2 - side by side plot comparison using N-gram**\n\nAn n-gram is a contiguous sequence of n items from a given sample of text or speech. Different definitions of n-grams will allow for the identification of the most prevalent words/sentences in the training data and thus help distinguish what comprises insincere and sincere questions.\n\nIt should be noted that prior to displaying individual words or sentences, the text will first be tokenized (based on a desired integer) and then put into a dataframe which will be used to construct side by side plots. \n\nTokenization is, generally, an early step in the NLP process, a step which splits longer strings of text into smaller pieces, or tokens. Larger chunks of text can be tokenized into sentences, sentences can be tokenized into words, etc. "},{"metadata":{"trusted":true,"_uuid":"0d92e5df7e84bb4c4452a0f2df3f12fb7454bad5"},"cell_type":"code","source":"def ngram_extractor(text, n_gram):\n    token = [token for token in text.lower().split(\" \") if token != \"\" if token not in STOPWORDS]\n    ngrams = zip(*[token[i:] for i in range(n_gram)])\n    return [\" \".join(ngram) for ngram in ngrams]\n\n# Function to generate a dataframe with n_gram and top max_row frequencies\ndef generate_ngrams(df, col, n_gram, max_row):\n    temp_dict = defaultdict(int)\n    for question in df[col]:\n        for word in ngram_extractor(question, n_gram):\n            temp_dict[word] += 1\n    temp_df = pd.DataFrame(sorted(temp_dict.items(), key=lambda x: x[1])[::-1]).head(max_row)\n    temp_df.columns = [\"word\", \"wordcount\"]\n    return temp_df\n\n#Function to construct side by side comparison plots\ndef comparison_plot(df_1,df_2,col_1,col_2, space):\n    fig, ax = plt.subplots(1, 2, figsize=(20,10))\n    \n    sns.barplot(x=col_2, y=col_1, data=df_1, ax=ax[0], color=\"royalblue\")\n    sns.barplot(x=col_2, y=col_1, data=df_2, ax=ax[1], color=\"royalblue\")\n\n    ax[0].set_xlabel('Word count', size=14)\n    ax[0].set_ylabel('Words', size=14)\n    ax[0].set_title('Top words in sincere questions', size=18)\n\n    ax[1].set_xlabel('Word count', size=14)\n    ax[1].set_ylabel('Words', size=14)\n    ax[1].set_title('Top words in insincere questions', size=18)\n\n    fig.subplots_adjust(wspace=space)\n    \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"111d7eaac4ddad6e6cbf6cece4aa434a5a58ead5"},"cell_type":"code","source":"#Obtain sincere and insincere ngram based on 1 gram (top 20)\nsincere_1gram = generate_ngrams(train[train[\"target\"]==0], 'question_text', 1, 20)\ninsincere_1gram = generate_ngrams(train[train[\"target\"]==1], 'question_text', 1, 20)\n#compare the bar plots\ncomparison_plot(sincere_1gram,insincere_1gram,'word','wordcount', 0.25)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"28bf6e9ac67f221ce05475a378ca48d659a6d08d"},"cell_type":"code","source":"#Obtain sincere and insincere ngram based on 2 gram (top 20)\nsincere_2gram = generate_ngrams(train[train[\"target\"]==0], 'question_text', 2, 20)\ninsincere_2gram = generate_ngrams(train[train[\"target\"]==1], 'question_text', 2, 20)\n#compare the bar plots\ncomparison_plot(sincere_2gram,insincere_2gram,'word','wordcount', 0.25)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9ef660c2243e2d669c283d8130a1036f8fcfa146"},"cell_type":"code","source":"#Obtain sincere and insincere ngram based on 3 gram (top 20)\nsincere_3gram = generate_ngrams(train[train[\"target\"]==0], 'question_text', 3, 20)\ninsincere_3gram = generate_ngrams(train[train[\"target\"]==1], 'question_text', 3, 20)\n#compare the bar plots\ncomparison_plot(sincere_3gram,insincere_3gram,'word','wordcount', 0.25)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c1d0c81535f3be251b221af0907bbe2eddde08bb"},"cell_type":"markdown","source":"**EDA Plot 3 - Word count distribution, Character Length Distribution, Stop words, Punctuation, Upper case **"},{"metadata":{"trusted":true,"_uuid":"030f5f480047065328ac341e4e848f048473c2ff"},"cell_type":"code","source":"# Number of words in the questions\ntrain[\"word_count\"] = train[\"question_text\"].apply(lambda x: len(str(x).split()))\ntest[\"word_count\"] = test[\"question_text\"].apply(lambda x: len(str(x).split()))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"word_count\", y=\"target\", data=train, ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Word Count', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Word Count distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9a49afcc8a9b58438d2b060e6229e26d3477b9b4"},"cell_type":"code","source":"# Number of characters in the questions\ntrain[\"char_length\"] = train[\"question_text\"].apply(lambda x: len(str(x)))\ntest[\"char_length\"] = test[\"question_text\"].apply(lambda x: len(str(x)))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"char_length\", y=\"target\", data=train, ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Character Length', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Character Length distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"25f65d85776486c2103a6597bc5673bdf80a8303"},"cell_type":"code","source":"# Number of stop words in the questions\ntrain[\"stop_words_count\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\ntest[\"stop_words_count\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in STOPWORDS]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"stop_words_count\", y=\"target\", data=train, ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of stop words', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Number of Stop Words distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e0b4a4aa4b2701289d953bd159b3f6578aba0afe"},"cell_type":"code","source":"# Number of punctuations in the questions\ntrain[\"punc_count\"] = train[\"question_text\"].apply(lambda x: len([c for c in str(x) if c in string.punctuation]))\ntest[\"punc_count\"] = test[\"question_text\"].apply(lambda x: len([c for c in str(x) if c in string.punctuation]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"punc_count\", y=\"target\", data=train[train['punc_count']<train['punc_count'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of punctuations', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Punctuation distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"44b2f6e778ab9f526e7a99fad47b6cc4fd2ac44b"},"cell_type":"code","source":"# Number of upper case words in the questions\ntrain[\"upper_words\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\ntest[\"upper_words\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"upper_words\", y=\"target\", data=train[train['upper_words']<train['upper_words'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of Upper case words', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Upper case words distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3eb4ea5ae43a154d60655bf448ae49415460fd29"},"cell_type":"code","source":"# Number of title words in the questions\ntrain[\"title_words\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\ntest[\"title_words\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"title_words\", y=\"target\", data=train[train['title_words']<train['title_words'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Number of Title words', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Title words distribution', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"edbc57ac2c583e222088103cee1606f4015b55b3"},"cell_type":"code","source":"# Mean word length in the questions\ntrain[\"word_length\"] = train[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ntest[\"word_length\"] = test[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\n\nfig, ax = plt.subplots(figsize=(15,2))\nsns.boxplot(x=\"word_length\", y=\"target\", data=train[train['word_length']<train['word_length'].quantile(.99)], ax=ax, palette=sns.color_palette(\"RdYlGn_r\", 10), orient='h')\nax.set_xlabel('Mean word length', size=10, color=\"#0D47A1\")\nax.set_ylabel('Target', size=10, color=\"#0D47A1\")\nax.set_title('[Horizontal Box Plot] Distribution of mean word length', size=12, color=\"#0D47A1\")\nplt.gca().xaxis.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"33ab3f7c9c30d2ea21f56789f20b0402b0d6826f"},"cell_type":"markdown","source":"*Overview of EDA results*\n\nThe EDA has helped outline a few characteristics which define insincere questions as such:\n\n* Insincere questions are mostly focused at politics, religion and can contain profanity.\n* Insincere questions (Class 1) are generally more lengthy, with the exceptions of certain outliers for sincere questions. They thus have more stop words, punctions, characters, a higher average word count and title words.\n* Insincere questions are mostly lower case.\n* There are no missing cases.\n\nThe EDA has also revealed a few issues about the dataset which can be regrouped as:\n\n* Unbalanced classes : a model trained on the current split of classes for the target variable will create a model more apt at predicting the '0' class rather than the '1' class, leading to false negatives in the predictions.\n* Uneven length of questions: the questions asked do not have a standard length and can thus lead to some questions being longer or shorter than others. \n* Unstandardized letter cases\n* Punctuations\n* Stop Words\n* Outliers"},{"metadata":{"_uuid":"839e80af57f2d6d567a24c41d1fe2b63e9edced6"},"cell_type":"markdown","source":"# Data Preparation \n\n"},{"metadata":{"_uuid":"a2f39d176de0f0e25816f30802de57bdb420e407"},"cell_type":"markdown","source":"There are two ways of fitting an NLP model:\n\n* Without the use of Word Embeddings (this will include steps such as lemmentization, removal of stopwords, punctuation and standardizing the characters)\n* With Word Embeddings (this should involve limited pre-processing steps as compared to the above)\n\nThis version of the kernel will focus on the use of Word Embeddings for sentiment analysis, with some sample code for text pre-processing should word embeddings not be present  (as shown below).\n"},{"metadata":{"trusted":true,"_uuid":"5104fc1aca1be22479a3582060a40ff666826ab4"},"cell_type":"code","source":"# nlp = spacy.load('en_core_web_sm')\n# # Clean text before feeding it to model\n# punctuations = string.punctuation\n\n# # Define function to cleanup text by removing personal pronouns, stopwords, puncuation and reducing all characters to lowercase \n# def cleanup_text(docs, logging=False):\n#     texts = []\n#     for doc in tqdm(docs):\n#         doc = nlp(doc, disable=['parser', 'ner'])\n#         tokens = [tok.lemma_.lower().strip() for tok in doc if tok.lemma_ != '-PRON-']\n#         #remove stopwords and punctuations\n#         tokens = [tok for tok in tokens if tok not in stopwords and tok not in punctuations]\n#         tokens = ' '.join(tokens)\n#         texts.append(tokens)\n#     return pd.Series(texts)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bc71652099729d957c6078118506eb1bd243bb6b"},"cell_type":"code","source":"# # Cleanup text and make sure it retains original shape\n# print('Original training data shape: ', train['question_text'].shape)\n# train_cleaned = cleanup_text(train['question_text'], logging=True)\n# print('Cleaned up training data shape: ', train_cleaned.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"33d491e0b8f59cc3ed039743b6bfd8b0172972bf"},"cell_type":"code","source":"#use 90-10 split for validation dataset\ntrain, val_df = train_test_split(train, test_size=0.1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"728e5b5c91221ef8c2cbad12fd9d3e01e52c1616"},"cell_type":"code","source":"# embdedding setup\n# Source https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html\n#Based on https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\n#GloVe is the most comprehensive word embedding\n\nembeddings_index = {}\nf = open('../input/embeddings/glove.840B.300d/glove.840B.300d.txt')\nfor line in tqdm(f):\n    values = line.split(\" \")\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs\nf.close()\n\nprint('Found %s word vectors.' % len(embeddings_index))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ad864642e77263264c747b936202e74d85e63add"},"cell_type":"code","source":"# Convert values to embeddings\ndef text_to_array(text):\n    empyt_emb = np.zeros(300)\n    text = text[:-1].split()[:30]\n    embeds = [embeddings_index.get(x, empyt_emb) for x in text]\n    embeds+= [empyt_emb] * (30 - len(embeds))\n    return np.array(embeds)\n\n# train_vects = [text_to_array(X_text) for X_text in tqdm(train[\"question_text\"])]\nval_vects = np.array([text_to_array(X_text) for X_text in tqdm(val_df[\"question_text\"][:3000])])\nval_y = np.array(val_df[\"target\"][:3000])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"78d8dd5d10dcaf0ec5279bdbb9311de1cd502a6e"},"cell_type":"code","source":"# Data providers\nbatch_size = 128\n\ndef batch_gen(train):\n    n_batches = math.ceil(len(train) / batch_size)\n    while True: \n        train = train.sample(frac=1.)  # Shuffle the data.\n        for i in range(n_batches):\n            texts = train.iloc[i*batch_size:(i+1)*batch_size, 1]\n            text_arr = np.array([text_to_array(text) for text in texts])\n            yield text_arr, np.array(train[\"target\"][i*batch_size:(i+1)*batch_size])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"86f7836192d466e5a7e0d00f5affeb2f06cdfe85"},"cell_type":"markdown","source":"# Modelling"},{"metadata":{"trusted":true,"_uuid":"5afa576991afc9efde1d7880e5e289e495e523e9"},"cell_type":"code","source":"#import Bi-Directional LSTM \nfrom keras.models import Sequential\nfrom keras.layers import CuDNNLSTM, Dense, Bidirectional\nimport math","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4559b7c4d2b448b6562a6dd08dcb76beb17d81ff"},"cell_type":"code","source":"#Define the model architecture\nmodel = Sequential()\nmodel.add(Bidirectional(CuDNNLSTM(64, return_sequences=True),\n                        input_shape=(30, 300)))\nmodel.add(Bidirectional(CuDNNLSTM(64)))\nmodel.add(Dense(1, activation=\"sigmoid\"))\n\nmodel.compile(loss='binary_crossentropy',\n              optimizer='adam',\n              metrics=['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"96455c016100c4f0f502f75a787440e722b6049b"},"cell_type":"code","source":"mg = batch_gen(train)\n#remember to change the number of epochs\nmodel.fit_generator(mg, epochs=10,\n                    steps_per_epoch=1000,\n                    validation_data=(val_vects, val_y),\n                    verbose= True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"95d2f40c4c6ce4c26de69f9c737a661083b65303"},"cell_type":"code","source":"# prediction part\nbatch_size = 256\ndef batch_gen(test):\n    n_batches = math.ceil(len(test) / batch_size)\n    for i in range(n_batches):\n        texts = test.iloc[i*batch_size:(i+1)*batch_size, 1]\n        text_arr = np.array([text_to_array(text) for text in texts])\n        yield text_arr\n\ntest = pd.read_csv(\"../input/test.csv\")\n\nall_preds = []\nfor x in tqdm(batch_gen(test)):\n    all_preds.extend(model.predict(x).flatten())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"790e8463d1fa60d55ee305930f5a0618372b6c04"},"cell_type":"code","source":"#Submit predictions\ny_te = (np.array(all_preds) > 0.5).astype(np.int)\n\nsubmit_df = pd.DataFrame({\"qid\": test[\"qid\"], \"prediction\": y_te})\nsubmit_df.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d90ae1cdd33e528a9a164c6d5324917f5e6738b5"},"cell_type":"markdown","source":"# Evaluation\n\nThe current model configuration leads to a score of around 0.547 \n\nThe next steps for this kernel will be:\n\n* Investigate how dimensionality reduction techniques affect the model\n* Implement oversampling to cater for unbalanced positive classes\n* Implement the callback history in the model defintion to allow for the mapping of training and testing error (overfitting detection)\n* Apply rules of thumb for LSTM architecture defintion (sourced from journals)\n* Implement text treatment as per: https://www.kaggle.com/theoviel/improve-your-score-with-text-preprocessing-v2\n\n**Thanks for reading so far!**\n\n**Any comments, advice, recommendations and upvotes would be much appreciated!**"},{"metadata":{"_uuid":"0d93f7ec07bab4b411628598bffd510e7c37d7b0"},"cell_type":"markdown","source":""}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}