{"cells":[{"metadata":{"_uuid":"4674c9b86ac96014bd6fa27282142e80126cc3df"},"cell_type":"markdown","source":"Most of the time when looking at a Kaggle competition (especially when trying to learn something new) it can be extremely intimidating to see kernels that are tagged as 'beginner' or have 'simple starter' pasted across the headline and when exploring these kernels,  there are ensembled neural nets with bidirectional LSTM's, pretrained embeddings etc. \n\nI also find that on kaggle it is extremely satisfying getting from loading the data to a useable submission as fast as possible because then one has a \"working model\", thus reducing the intimidation barrier regarding the goal of producing something functional.\n\nSo I thought I would take a shot at creating a starter kernel that I would like to read when starting a competition as a beginner, in the case of the Quora Insincere Questions Classification competition that would be one that would allow someone to train a neural net with embeddings as fast and easily as possible, using best practice. Word embeddings are not necessarily a \"beginner\" concept, however the competition will no doubtedly depend heavily on them. Here it goes:"},{"metadata":{"_uuid":"a56c862411568a5990781d33702d181296075ea6"},"cell_type":"markdown","source":"# Contents\n\n* Import libraries and read in data\n* Turn text data into a vector\n* Make vectors uniform length\n* Split training data into training and validation sets\n* Create simple multi-layer perceptron (MLP) neural network with embedding layer\n* What is the F1 score?\n* Optimize F1 score for submission\n* Make submission file using optimized predictions\n* Going Further"},{"metadata":{"_uuid":"3ca8415cefd9a69fba38cf01c8d5f812a197153a"},"cell_type":"markdown","source":"# Import libraries and read in data"},{"metadata":{"_uuid":"ece84d9645369d0873bbcd4714ccd3d1c33f14de"},"cell_type":"markdown","source":"First we import libraries that we are likely to use:"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n#plot figures in the notebook wihtout the need to call plt.show()\n%matplotlib inline \nplt.style.use(\"seaborn-ticks\") #set default plotting style for matplotlib\n\nimport time\nimport os\nprint(os.listdir(\"../input\")) #Print directories/folders in the directory: current_working_directory/input/embeddings","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"be93d3e400838de1d6126bb42448c8089d603043"},"cell_type":"markdown","source":"Next let us read in the train and test data from the above directories using pandas. The training data has three columns: 'qid' , 'question_text' and 'target'. The qid is a unique identifier for each question, the question_text is a string of a question/sentence and the target is a signifier of insincerity where 0 is not insincere and 1 is insincere. "},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train=pd.read_csv(\"../input/train.csv\")\ntest=pd.read_csv(\"../input/test.csv\")\n\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b69b235fbcdd6a4310b1e4a1ef44740ab0005c10"},"cell_type":"markdown","source":"The test data only has the qid and question_text columns as the goal is to predict the target for the questions in the test set:"},{"metadata":{"trusted":true,"_uuid":"1f6acfe768383f8648d108d5078729c0906be817"},"cell_type":"code","source":"test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2a8402528b7385494eefa727d3c6de2103b5d676"},"cell_type":"markdown","source":"We can confirm that the qid's are unique by comparing the number of unique values to that of the size of the dataframe in both the training data and the test data. We find that they are the same so we will focus only on the question_text column:"},{"metadata":{"trusted":true,"_uuid":"38e9756f73fce85be50fd84621a6f9b7a41ee7b8"},"cell_type":"code","source":"print(train.shape,train.qid.nunique())\nprint(test.shape,test.qid.nunique())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"36722c8117b2d1fdddebb7fbcd68ff7d61328f79"},"cell_type":"markdown","source":"# Turn text data into vector"},{"metadata":{"_uuid":"aada49a498755052238e96458952d89facfdd136"},"cell_type":"markdown","source":"Now we need to turn the text from each question into something that a neural network will understand, specifically a vector for each question where each word has a unique integer identifier based on the number of words in a text corpus, here being all the words in our training set. To do this we will use the Tokenizer class from keras. Specifially we will make use of the tensorflow.keras implementation:"},{"metadata":{"trusted":true,"_uuid":"8e27d67fdd910bf66034b39afa6bdf76df1d55f8"},"cell_type":"code","source":"from tensorflow.keras.preprocessing.text import Tokenizer\nprint(Tokenizer.__doc__)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0621b162560da889b6006107d2260914c2a88be1"},"cell_type":"markdown","source":"So we will no choose a number of possible words/tokens to keep amongst all of our possible words. This will govern the ultimate dimensionality of our training matrix. We must fit the Tokenizer class to our training data and then transform both our training data and our test data:"},{"metadata":{"trusted":true,"_uuid":"77c83474207fd9cf4e4c7ce9b7ad06c3e78da31c"},"cell_type":"code","source":"num_possible_tokens=10000 #At this stage this was chosen arbitrarily\n\ntokenizer=Tokenizer(num_words=num_possible_tokens) #Instantiate tokenizer class with number of possible tokens\ntokenizer.fit_on_texts(train.question_text) #Fit the tokenizer to training data\nsequences_train=tokenizer.texts_to_sequences(train.question_text) #Convert training data to vectors\nsequences_test=tokenizer.texts_to_sequences(test.question_text) #Convert test data to vectors","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"96671c2d6efdd622d1085ef7f77ff1001dd066ab"},"cell_type":"code","source":"sequences_train[0:5]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ac31f834af8ad9b19e6b3b670e7a954908ef45eb"},"cell_type":"markdown","source":"The only problem now before going into the neural network is that the network requires a matrix (equal length vectors), we will cover that next."},{"metadata":{"_uuid":"5bd2bd90dc36aef8944715c6c7bd874c2b822693"},"cell_type":"markdown","source":"# Make vectors uniform length"},{"metadata":{"_uuid":"695acc9fed91f5028f93a1cb41999b0a9358336f"},"cell_type":"markdown","source":"To convert our vectors all to the same length, let us use a list comprehension to get the length of the longest vector amongst all vectors in the training and test sets:"},{"metadata":{"trusted":true,"_uuid":"f56944cc4458f7b3961d0d72f765d747c8328fd1"},"cell_type":"code","source":"max_len=np.max([len(i) for i in sequences_train]+[len(i) for i in sequences_test])\nprint(max_len)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c17c9f0f1760654b17478a1d30c54ebd1f08eba"},"cell_type":"markdown","source":"Next let us use the pad_sequences function from keras to either add zeros to the beginning of each vector or add trailing zeros to each vector to make the all the maximu length calculated above. It can be seen in the docstring below that the default is to pre-pad the input. We need to pass the maximum length required to the function to know how far to pad:"},{"metadata":{"trusted":true,"_uuid":"666352e179f42f95929a83c2f10dc6cc496a2252"},"cell_type":"code","source":"from tensorflow.keras.preprocessing.sequence import pad_sequences\nprint(pad_sequences.__doc__)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6c96fae13b1f4bf434ea9f4c9ed37a3bfd794656"},"cell_type":"code","source":"X=pad_sequences(sequences_train,maxlen=max_len) #Pad the training data, later to be split into a smaller training set and a validation set\nX_test=pad_sequences(sequences_test,maxlen=max_len) #Pad the test data\n\ny=train.target.values #Make and independent target variable from the training target. Also to be split into a smaller training set and validation set.\n\nprint(X[0:10,:]) # Print first ten rows of the training data\nprint(X.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"03ed29e16d443351d3c77f1cdae3e226fbdb26b2"},"cell_type":"markdown","source":"We now see that our input is a matrix with a dimension equal to the length of the longest vector in the training and test data. The leading zeros (pre-padded) can also be seen."},{"metadata":{"_uuid":"d634b0259cacef6af8ec7dc4794913ee5fe7cd0f"},"cell_type":"markdown","source":"# Split training data into training and validation sets"},{"metadata":{"_uuid":"768caec6bded1aa914f911d7d6f5d58bf66b089b"},"cell_type":"markdown","source":"Seeing as we do not have the targets for the test data (obviously) we need a validation set. If the dataset is too small one can use cross validation, however with approximately 1.3 million questions let us use a single validation set. We split the training data and training target above (variable X) into smaller training sets and validation sets. The test size is arbitrarily chosen to be 20 percent."},{"metadata":{"trusted":true,"_uuid":"fa16240136eeeeb4fc8f722782e928c2835786a5"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nX_train,X_val,y_train,y_val=train_test_split(X,y, test_size=0.2, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"966d0315a08ea64510cebc9eff5a0591c3dcc746"},"cell_type":"markdown","source":"# Create simple multi-layer perceptron (MLP) neural network with embedding layer"},{"metadata":{"_uuid":"902c19c040c2f098a1a9fec02837cd30b76bc40b"},"cell_type":"markdown","source":"Before getting into the network let us discuss the embedding layer for the neural network. The embedding layer is much like a normal layer of a neural network that transforms an input matrix into a learned representation at each node, in this case each node in the layer is a word and the parameters of the layer will depict where that word sits in vector space relative to other words.\n\nFrom the docstring below we can see that the embedding layer receives a 2D vector with the shape (sample size,input length) where the input length is the length of our questions modified above. The output of the embedding layer is then (sample size,input length, embedding dimension).  From the docstring, it is important to note that if we are to use a MLP, the input_length must be specified in order to flatten the embedding layer as input for a dense layer"},{"metadata":{"trusted":true,"_uuid":"55731c4093b4d801a0216b61b2a3a495d714a8f2"},"cell_type":"code","source":"import tensorflow.keras as keras\nfrom tensorflow.keras import layers\n\nprint(layers.Embedding.__doc__)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"290196a77f811129883cc49b7fdd9d609b7cfbd1"},"cell_type":"markdown","source":"For our embedding layer:\n* input_dim is the size of our vocabulary + 1 for unknowns\n* output dim is a chosen dimensionality\n* input_length is the dimension of the input vector specified as the length of the longest question in our input vector\n\n Let us make a small 3 layer MLP where the first layer is the embedding layer and the remaing two layers are conventional Dense layers:"},{"metadata":{"trusted":true,"_uuid":"19886df2b77dabf81916f34da9fb7bb68013315c"},"cell_type":"code","source":"embedding_dimension=32 # Arbitraily choose an embedding dimension,the 157 dimension input vector will be compressed down to this dimension\n\nmodel=keras.models.Sequential() # Instantiate the Sequential class\n\nmodel.add(layers.Embedding(num_possible_tokens+1,embedding_dimension,input_length=max_len)) # Creat embedding layer as described above\nmodel.add(layers.Flatten()) #Flatten the embedding layer as input to a Dense layer\nmodel.add(layers.Dense(32, activation='relu')) # Dense layer with relu activation\nmodel.add(layers.Dense(1,activation='sigmoid')) # Dense layer with sigmoid activation for binary target\nmodel.compile(optimizer='rmsprop',loss='binary_crossentropy',metrics=['accuracy']) #binary cross entropy is used as the loss function and accuracy as the metric \nmodel.summary() # print out summary of the network","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4a820df60fc512103b3f7d6b68c684aa9d75cc6"},"cell_type":"code","source":"batch_size=1024 # Choose a batch size\nepochs=3 #Choose number of epochs to train\n\nhistory=model.fit(X_train,y_train,epochs=epochs,batch_size=batch_size,validation_data=[X_val,y_val])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"466e78d77d4201b341b0c3aff692a7bea93a34b8"},"cell_type":"markdown","source":"# What is the F1 score?"},{"metadata":{"_uuid":"954b50f9f2298043ea9d17ad813d71b222daabbb"},"cell_type":"markdown","source":"The F1 score is defined as:\n\n*  2 * (precision x recall) / (precision + recall)\n\nwhich is effectively a weighted average of precision and recall where:\n* precision is true positive/(true positive+false positive)\n* recall is true positive/(true positive+false negative)\n\nWe can caluclate it directly using ScikitLearn, where we first need to convert our predicted values to a binary target on order to calculate the precision and recall, initially chosen to be zero below and equal to 0.5 and 1 above 0.5. Remember that the neural network outputs a probability distribution, so we can choose the probability threshold at which our binary target is split.\n\nThe f1 score varies between 0 (worst) and 1 (best)."},{"metadata":{"trusted":true,"_uuid":"c1a2e3e256c530b46d35d85407466821c7c783ce"},"cell_type":"code","source":"from sklearn.metrics import f1_score\n\nval_pred=model.predict(X_val,batch_size=batch_size).ravel() # predict the values in the validation set which the neural net has not seen\n\nf1_score(y_val,val_pred>0.5) #Predict the f1 score at a threshold of 50%, the point at which our binary target is split in our neural networks output probability distribution","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7542d273efdab685811f70df462cb5976df1e9a0"},"cell_type":"markdown","source":"# Optimize F1 score for submission"},{"metadata":{"_uuid":"6ba26670c30389c09fd6042918134debbe2a18c9"},"cell_type":"markdown","source":"Above we chose a probability threshold of 50% to split our binary target, but is this always the best predictor? We can sweep a range of thresholds between 0 and 50% and recalculate the f1 score based on each threshold:"},{"metadata":{"trusted":true,"_uuid":"b75e3cb8e85aa76bdde917e9403dd0e087d8535a"},"cell_type":"code","source":"Threshold=[] # List ot store tested thresholds\nf1=[] # List to store associated f1 score for threshold\n\nfor i in np.arange(0.1, 0.501, 0.01):\n    Threshold.append(i)\n    temp_val_pred=val_pred>i # convert to True or False Boolean based on threshold\n    temp_val_pred=temp_val_pred.astype(int) # Convert Boolean to integer\n    score=f1_score(y_val,temp_val_pred) #Calculate f1 score at threshold\n    f1.append(score) #store f1 score\n    print(\"Threshold: {} \\t F1 Score: {}\".format(np.round(i,2),score))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"95d3debee06febbfc6245eac651a5b70a42c5475"},"cell_type":"markdown","source":"From the stored thresholds and lists calculate the optimum threshold:"},{"metadata":{"trusted":true,"_uuid":"c9e480901eeba47d6727c43483913ee0e62e2675"},"cell_type":"code","source":"best_threshold=Threshold[np.argmax(f1)] #Get threshold at index of largest f1 score.\nbest_threshold","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9c474e03c690f053e260cd2ef424b67eca09025d"},"cell_type":"markdown","source":"# Make submission file using optimized predictions"},{"metadata":{"_uuid":"b3ff06bc1f516f3548c1fd2cc5f983a5b20a7fea"},"cell_type":"markdown","source":"Finally we calculate our predictions on the test data and put the data into a dataframe in the required format, where our predictions are converted to a binary target based on the best threshold calculated above. "},{"metadata":{"trusted":true,"_uuid":"726ea0d74625fa401251bfbdc0f835f7da68f4b8"},"cell_type":"code","source":"test_pred=model.predict(X_test,batch_size=4096).ravel() #Predict test data\n\ndf=pd.DataFrame({'qid':test.qid.values,'prediction':test_pred}) #Create dataframe of unique id's and predicted target \ndf.prediction=(df.prediction>best_threshold).astype(int) #Convert target to binary based on best f1 threshold\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"91d5f3808eccff87276d9cf2d04592241b6fc39b"},"cell_type":"markdown","source":"Write csv for submission!"},{"metadata":{"trusted":true,"_uuid":"600dece9333205be7e9d4b27178df9e826602f46"},"cell_type":"code","source":"df.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d063f87142ade8ad1100c1eb3997ef2100cea08"},"cell_type":"markdown","source":"# Going Further"},{"metadata":{"_uuid":"0d9715c1171dcabb1c73721ad5d9f597b7fa3fc5"},"cell_type":"markdown","source":"Obvioulsy this is an extremely naive model having never looked at the data but its purpose was to get from input to working submission in a minimal manner in order to understand the process. Now with a working model one can quite easily start to:\n\n* Perform Exploratory Data Analysis (EDA) knowing how to convert it into something useable\n* Clean the data knowing the required format\n* Vary the network structure by making small changes to a working network\n* Change network parameters etc,\n\nwithout the looming intimidation of getting through the process.  Let me know if this helps if you are new to keras and embeddings :) \n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}