{"cells":[{"metadata":{"trusted":true,"_uuid":"875f43d8fd1ee65143f5d153f657b64f62fbb88f"},"cell_type":"markdown","source":"In Machine Learning models, we can know how they make predictions by visualizing features' weight. However, it is difficult to understand them when we apply them on text questions. In addition, when people try to use Neural Network (Deep Learning) models, predictions are mysterious. Feature weights are in the black box. \n\nIn this notebook, I will apply LIME (Local Interpretable Model-agnostic Explanations), which was introduced in 2016 in a paper called [\"Why Should I Trust You?\": Explaining the Predictions of Any Classifier](https://arxiv.org/abs/1602.04938), on a simple Logistic mode and a simple NN model. The purpose of LIME is to explain a model prediction for a specific sample in a human-interpretable way."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np, pandas as pd, random as rn, time, gc, string, warnings\n\nseed = 32\nnp.random.seed(seed)\nrn.seed(seed)\nimport tensorflow as tf\nsession_conf = tf.ConfigProto(intra_op_parallelism_threads = 1,\n                              inter_op_parallelism_threads = 1)\ntf.set_random_seed(seed) \nsess = tf.Session(graph = tf.get_default_graph(), config = session_conf)\nfrom keras import backend as K\nK.set_session(sess)\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.linear_model import LogisticRegression\n\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom nltk.corpus import stopwords\n\nwarnings.filterwarnings(\"ignore\")\n%matplotlib inline\n\ntrain = pd.read_csv(\"../input/train.csv\")\ntest = pd.read_csv(\"../input/test.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d85cd024492796a3b79462f4d46bd937fd8cd6a0"},"cell_type":"code","source":"sincere = train[train[\"target\"] == 0]\ninsincere = train[train[\"target\"] == 1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bfcc804ac297f3d214e08b514e99a622bc704757"},"cell_type":"code","source":"print(sincere.shape)\nprint(insincere.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"23739b5681290cd2480472f8544ae981757a184c"},"cell_type":"code","source":"train = pd.concat([sincere[:int(len(sincere)*0.9)], insincere[:int(len(insincere)*0.9)]])\nval = pd.concat([sincere[int(len(sincere)*0.9):], insincere[int(len(insincere)*0.9):]])","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"tfidf_vc = TfidfVectorizer(min_df = 10,\n                          max_features = 100000,\n                          analyzer = \"word\",\n                          ngram_range = (1, 2),\n                          stop_words = \"english\",\n                          lowercase = True)\n\ntrain_vc = tfidf_vc.fit_transform(train[\"question_text\"])\nval_vc = tfidf_vc.transform(val[\"question_text\"])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f1c984543d4901ee98d222e3eaf0cb15323d821b"},"cell_type":"markdown","source":"## Machine Learning model: Logistic Regression"},{"metadata":{"trusted":true,"_uuid":"6056f04451375c2556d7077c2210c2e7c128570a"},"cell_type":"code","source":"model = LogisticRegression(C = 0.5, solver = \"sag\")\nmodel = model.fit(train_vc, train.target)\nval_pred = model.predict(val_vc)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"397d7fefe035d7d4b9430a9649a7f3ae59549f88"},"cell_type":"code","source":"from sklearn.metrics import f1_score\n\nval_cv = f1_score(val.target, val_pred, average = \"binary\")\nprint(val_cv)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"445fe74fa2d413c5b8354bd445594282876c2378"},"cell_type":"code","source":"from lime import lime_text\nfrom sklearn.pipeline import make_pipeline\nfrom lime.lime_text import LimeTextExplainer\nfrom lime import submodular_pick\nfrom collections import OrderedDict\n\nidx = val.index[32]\nc = make_pipeline(tfidf_vc, model)\nclass_names = [\"sincere\", \"insincere\"]\nexplainer = LimeTextExplainer(class_names = class_names)\nexp = explainer.explain_instance(val[\"question_text\"][idx], c.predict_proba, num_features = 50)\n\nprint(\"Question: \\n\", val[\"question_text\"][idx])\nprint(\"Probability (Insincere) =\", c.predict_proba([val[\"question_text\"][idx]])[0, 1])\nprint(\"True Class is:\", class_names[val[\"target\"][idx]])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8584224ee9d09ab5c326c0696cc11c5598445327"},"cell_type":"code","source":"exp.show_in_notebook()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"aef304ae1e71a1feee903075bdf8790523456a3f"},"cell_type":"code","source":"weights = OrderedDict(exp.as_list())\nlime_weights = pd.DataFrame({\"words\": list(weights.keys()), \n                             \"weights\": list(weights.values())})\n\nsns.barplot(x = \"words\", y = \"weights\", data = lime_weights)\nplt.xticks(rotation = 45)\nplt.title(\"Sample {} features weights given by LIME\".format(idx))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3347c5881ee825de99c50eac2ba5207e9f5c97a8"},"cell_type":"markdown","source":"As we see from above, our model consider \"engineering\" and \"software\" are negative features."},{"metadata":{"trusted":true,"_uuid":"0340bfebd63296b35c5d5911cba297b04ea6797c"},"cell_type":"code","source":"sp_obj = submodular_pick.SubmodularPick(explainer, val[\"question_text\"].values, \n                                        c.predict_proba, sample_size = 10, \n                                        num_features = 50, num_exps_desired = 6,\n                                        top_labels = 3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"92a84c585ce79f34001d372d5583f84c705c6393"},"cell_type":"markdown","source":"We list $50$ words from our model to see how they effect the prediction."},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"0a427da3712955f59e1be1e5c8ffed5a2fdca5a8"},"cell_type":"code","source":"[exp.as_pyplot_figure(label = 0) for exp in sp_obj.sp_explanations]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c6c9868e55cea7c4765db02d483d4930a08547bb"},"cell_type":"markdown","source":"## Deep Learning model: CNN"},{"metadata":{"trusted":true,"_uuid":"a672c59e77008e08702e80afb3df8164e7097316"},"cell_type":"code","source":"from keras.layers import Input, Embedding, Dense, Conv1D, MaxPool1D, BatchNormalization\nfrom keras.layers import Flatten, Concatenate, Dropout, SpatialDropout1D\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.models import Model\n\nEMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\nmax_features = 50000\nmax_len = 70\nembed_size = 300","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4c39377d69b4ed12ea57f7c3ff65ae09dc30372c"},"cell_type":"code","source":"X_train, y_train = train[\"question_text\"], train[\"target\"]\nX_val, y_val = val[\"question_text\"], val[\"target\"]\n\ntk = Tokenizer(num_words = max_features)\ntk.fit_on_texts(X_train)\n# X_train = tk.texts_to_sequences(X_train)\n# X_val = tk.texts_to_sequences(X_val)\n# X_train = pad_sequences(X_train, maxlen = max_len)\n# X_val = pad_sequences(X_val, maxlen = max_len)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"87e8f8acd72237dcad7f0c2511d21ea36ecb708a"},"cell_type":"code","source":"from sklearn.pipeline import TransformerMixin\nfrom sklearn.base import BaseEstimator\n\nclass TextsToSequences(Tokenizer, BaseEstimator, TransformerMixin):\n    \"\"\" Sklearn transformer to convert texts to indices list \n    (e.g. [[\"the cute cat\"], [\"the dog\"]] -> [[1, 2, 3], [1, 4]])\"\"\"\n    def __init__(self,  **kwargs):\n        super().__init__(**kwargs)\n        \n    def fit(self, texts, y=None):\n        self.fit_on_texts(texts)\n        return self\n    \n    def transform(self, texts, y=None):\n        return np.array(self.texts_to_sequences(texts))\n        \nsequencer = TextsToSequences(num_words = max_features)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"22fcfbd8996b6e01bf15e876a35861761d19440d"},"cell_type":"code","source":"class Padder(BaseEstimator, TransformerMixin):\n    \"\"\" Pad and crop uneven lists to the same length. \n    Only the end of lists longernthan the maxlen attribute are\n    kept, and lists shorter than maxlen are left-padded with zeros\n    \n    Attributes\n    ----------\n    maxlen: int\n        sizes of sequences after padding\n    max_index: int\n        maximum index known by the Padder, if a higher index is met during \n        transform it is transformed to a 0\n    \"\"\"\n    def __init__(self, maxlen=500):\n        self.maxlen = maxlen\n        self.max_index = None\n        \n    def fit(self, X, y=None):\n        self.max_index = pad_sequences(X, maxlen=self.maxlen).max()\n        return self\n    \n    def transform(self, X, y=None):\n        X = pad_sequences(X, maxlen=self.maxlen)\n        X[X > self.max_index] = 0\n        return X\n\npadder = Padder(max_len)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9239a702dc30d1f2e6dc55d964a8c3ca47104d0a"},"cell_type":"code","source":"def get_coefs(word, *arr): return word, np.asarray(arr, dtype='float32')\nembeddings_index = dict(get_coefs(*o.rstrip().rsplit(' ')) for o in open(EMBEDDING_FILE))\n\nword_index = tk.word_index\nnb_words = min(max_features, len(word_index))\nembedding_matrix = np.zeros((nb_words, embed_size))\nfor word, i in word_index.items():\n    if i >= max_features: continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: embedding_matrix[i] = embedding_vector","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"16253ed4c6ae86710e5b144e787671fd65f2a9c8"},"cell_type":"code","source":"from keras.callbacks import Callback\n\nclass F1Evaluation(Callback):\n    def __init__(self, validation_data=(), interval=1):\n        super(Callback, self).__init__()\n\n        self.interval = interval\n        self.X_val, self.y_val = validation_data\n\n    def on_epoch_end(self, epoch, logs={}):\n        if epoch % self.interval == 0:\n            y_pred = self.model.predict(self.X_val, verbose=0)\n            y_pred = (y_pred > 0.35).astype(int)\n            score = f1_score(self.y_val, y_pred)\n            print(\"\\n F1 Score - epoch: %d - score: %.6f \\n\" % (epoch+1, score))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"33cecdb2829b118b402b8ec14ebc196de2a1a502"},"cell_type":"code","source":"from keras.wrappers.scikit_learn import KerasClassifier\nfrom sklearn.pipeline import make_pipeline\n\ndef build_model():\n    inp = Input(shape = (max_len,))\n    x = Embedding(max_features, embed_size, weights = [embedding_matrix])(inp)\n    x = SpatialDropout1D(0.5)(x)\n    x = Conv1D(32, kernel_size = 2, activation = \"relu\")(x)\n    x = Conv1D(32, kernel_size = 2, activation = \"relu\")(x)\n    x = MaxPool1D(pool_size = 3)(x)\n    \n    x = Flatten()(x)\n    x = BatchNormalization()(x)\n    out = Dense(1, activation = \"sigmoid\")(x)\n    \n    model = Model(inputs = inp, outputs = out)\n    model.compile(loss = \"binary_crossentropy\", optimizer = \"adam\", metrics = [\"accuracy\"])\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4525ba7307d7d5c09befcc91e01574b857945bc3"},"cell_type":"code","source":"batch_size = 512\nsklearn_cnn = KerasClassifier(build_fn = build_model, epochs = 5, \n                              batch_size = batch_size, verbose = 2)\n\npipeline = make_pipeline(sequencer, padder, sklearn_cnn)\npipeline.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d14ce755c9eb90a3106c8ca6603d476afedf5dbc"},"cell_type":"code","source":"val_pred = pipeline.predict(X_val)\nval_cv = f1_score(y_val, val_pred, average = \"binary\")\nprint(\"Local CV is {}\".format(val_cv))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3d932fd6c49e81aab20f723968a0b7ee04e2a581"},"cell_type":"code","source":"sp_obj = submodular_pick.SubmodularPick(explainer, val[\"question_text\"].values, \n                                        pipeline.predict_proba, sample_size = 10, \n                                        num_features = 100, num_exps_desired = 6,\n                                        top_labels = 3)\n\n[exp.as_pyplot_figure(label = 0) for exp in sp_obj.sp_explanations]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"97c3ceba27be00c1f7321abc04e6b934bc306016"},"cell_type":"code","source":"def explain(idx):\n    print(\"Sample question:\\n\", X_val.iloc[idx])\n    print(\"-\"*50)\n    print(\"Probability (Insincere): {}\".format(pipeline.predict_proba([X_val.iloc[idx]])[0, 1]))\n    print(\"True Class is {}\".format(class_names[y_val.iloc[idx]]))\n    explanation = explainer.explain_instance(X_val.iloc[idx], pipeline.predict_proba, \n                                             num_features = 20)\n    explanation.show_in_notebook()\n    weights = OrderedDict(explanation.as_list())\n    lime_weights = pd.DataFrame({\"words\": list(weights.keys()), \n                                 \"weights\": list(weights.values())})\n\n    sns.barplot(x = \"words\", y = \"weights\", data = lime_weights)\n    plt.xticks(rotation = 45)\n    plt.title(\"Sample {} features weights given by LIME\".format(idx))\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6fa44bd67ff3616aa72fbe0ceccd957e95f94331"},"cell_type":"code","source":"explain(32)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d82516333f6acf53347d603817cda00fe771f142"},"cell_type":"markdown","source":"Our NN model considers most words in the question are negative to insincere, which is good! "},{"metadata":{"trusted":true,"_uuid":"1739367ed8eaa908421840e7402de6d975cea7c4"},"cell_type":"code","source":"Class = pd.DataFrame({\"true\": y_val}).reset_index()\nClass[\"pred\"] = val_pred","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aa23b49d8918e1d1da566e4960abe72990daf004"},"cell_type":"markdown","source":"## Examples\n#### True Positive example"},{"metadata":{"trusted":true,"_uuid":"4252f0e6653f581bd8e14c2c4298df52235bf969"},"cell_type":"code","source":"temp = Class[(Class[\"true\"] == 1) & (Class[\"pred\"] > 0.5)]\nexplain(temp.index[0])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7024af3db7be81df6a1d2c622a0d4d3d2f12b01b"},"cell_type":"markdown","source":"#### True Negative"},{"metadata":{"trusted":true,"_uuid":"92c03fd32a5ef15bf5466c3d487ec2c8d5811520"},"cell_type":"code","source":"temp = Class[(Class[\"true\"] == 1) & (Class[\"pred\"] < 0.5)]\nexplain(temp.index[0])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"088dbb288c7c22f7abf64f48d54a3ebb670af4c5"},"cell_type":"markdown","source":"#### False Positive"},{"metadata":{"trusted":true,"_uuid":"d41859e83cac1968de18e6e97b3695c3b6c2a6b5"},"cell_type":"code","source":"temp = Class[(Class[\"true\"] == 0) & (Class[\"pred\"] < 0.5)]\nexplain(temp.index[0])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"784bca0196a2798537e76e8dc7ebd349346390cd"},"cell_type":"markdown","source":"#### False negative"},{"metadata":{"trusted":true,"_uuid":"071f73af1b935b3a42f2f668688a2d2a05b75a21"},"cell_type":"code","source":"temp = Class[(Class[\"true\"] == 0) & (Class[\"pred\"] > 0.5)]\nexplain(temp.index[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"08feda7ebbf94c975f10aae461f02f0b86982fc4"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}