{"cells":[{"metadata":{"_uuid":"e189ac8e0df6135ff5ccb16abf1d896d2ce47461"},"cell_type":"markdown","source":"## **Interpreting Machine Learning models is no more a lavishness but a requirement given the acute compliance of AI in the industry.**\n\n---\n\n### **Hands-on Guide of explain potential black-box machine learning models**\n\n* Kaggle FIFA Dataset\n* **Objective of this notebook to give hands-on experience of Machine Learning model intepretation.**\n\n### **Frame Work Going to Cover**\n* [**1.ELI5**](#1.ELI5)\n* [**2.Skater**](#2.Skater)\n* [**3.SHAP**](#3.SHAP)\n* [**References**](#References)"},{"metadata":{"_uuid":"83f3880851e6b33c1a8f8eee66a64a77967a881c"},"cell_type":"markdown","source":"---\n\n# **1.ELI5**\n\n---\n # **Permutation Feature Importance**\n* ***Permutation Feature Importance is an algorithm that ascertains hugeness scores for every one of the characteristic factors in a dataset. Proportions of importance are controlled by figuring the affectability of a model to arbitrary permutations of characteristic values. As such, an importance score measures the commitment of a specific characteristic to the prescient execution of a model as far as the amount of a picked assessment metric that is avoided in the wake of changing the estimations of that characteristic.***\n\n* ***The instinct behind permutation importance is that on the off chance that a feature isn't helpful for forecasting a result, at that point adjusting or permuting its qualities won't result in a critical decrease in a model's execution. This procedure is regularly utilized in [random forests](http://oz.berkeley.edu/~breiman/randomforests-rev.pdf) and is portrayed by Breiman in his original paper Random Forests. The methodology, nonetheless, can be summed up and adjusted to other order and relapse models.***\n \n ***Reading : https://eli5.readthedocs.io/en/latest/overview.html***"},{"metadata":{"_uuid":"ccaf6726220c51a8d5d9e1035b93394a233523b4"},"cell_type":"markdown","source":"### **1.Load Packages**"},{"metadata":{"trusted":true,"_uuid":"b4d881c3d7d6e823a101eb88c3095f3ffda5c8f9"},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1a98266acfcdc81a53205fab89565b4bb533cd21"},"cell_type":"markdown","source":"### **2.Read Data**"},{"metadata":{"trusted":true,"_uuid":"0f011e2a5ecbc75e5b68e8590fa8e9bf73b1edfa"},"cell_type":"code","source":"data = pd.read_csv('../input/FIFA 2018 Statistics.csv')\ny = (data['Man of the Match'] == \"Yes\")  # Convert from string \"Yes\"/\"No\" to binary\n\nfeature_names = [i for i in data.columns if data[i].dtype in [np.int64]]\nX = data[feature_names]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"27760770797636ea45f36f17f16f62bd0beca062"},"cell_type":"markdown","source":"### **3.Train test split**"},{"metadata":{"trusted":true,"_uuid":"a7dae3bb376a5409b8eba2b6b821c3dcf057b3cc"},"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0543b9015a533e2bf4c36e5d0b6c3b736a0c7d2c"},"cell_type":"markdown","source":"### **4.Model Training**"},{"metadata":{"trusted":true,"_uuid":"7cd6a44a3ff66897d1ed117c2cf71daf81387a23"},"cell_type":"code","source":"my_model = RandomForestClassifier(random_state=0).fit(train_X, train_y)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"11a1a27b78dbe286b65f525a4a5e639c88595286"},"cell_type":"markdown","source":"### **5.Permutation Importance**"},{"metadata":{"trusted":true,"_uuid":"780e2b0ef7e855925089845caae62da5610408e1"},"cell_type":"code","source":"import eli5\nfrom eli5.sklearn import PermutationImportance","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fa0cad2c4c51e9d740b4fe931d6a2bd46d8f560e"},"cell_type":"code","source":"perm = PermutationImportance(my_model, random_state = 1).fit(X_test, y_test)\neli5.show_weights(perm, feature_names = val_X.columns.tolist())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7fbadf5e8c40ff43c4da6dd67f067c8763784e40"},"cell_type":"markdown","source":"### **Interpreting Permutation Importances**\n* The qualities towards the best are the most essential features, and those towards the base issue least.\n* The primary number in each row demonstrates how much model execution diminished with an random shuffling (for this situation, utilizing \"accuracy\" as the execution metric).\n* Like most things in the stream of data science, there is some randomness to the definite execution change from a shuffling a data. We measure the measure of haphazardness in our permutation importance estimation by repeating the procedure with numerous mixes. The number after the ± measures how execution differed starting with one-reshuffling then onto the next.\n* We'll sporadically observe negative values for permutation importances. In those cases, the forecasts on the shuffled (or noisy) information happened to be more exact than the genuine data. This happens when the feature didn't make a difference (ought to have had an importance near 0), yet arbitrary shot made the forecasts on shuffled data be progressively precise. This is increasingly basic with little datasets, similar to the one in this model, in light of the fact that there is more space for good fortune/possibility. \n* In our precedent, the most imperative feature was Goals scored. That appears to be reasonable. Soccer fans may have some instinct about whether the orderings of different variables(factor) are astounding or not."},{"metadata":{"_uuid":"0931afb127155d1152da6769895cb48fb9e636ad"},"cell_type":"markdown","source":"---\n\n# **2.Skater**\n\n---\n### **Reference : https://datascienceinc.github.io/Skater/tutorial.html**\n* ***Skater is a unified framework that allows the interpretation of models for all forms of models to help build an interpretable machine learning system that is often needed for real world use cases using an agnostic approach.*** The **Python library** is designed to ***demystify the structures learned from a black box model globally*** (inference based on a complete data set) and locally (inference about individual prediction).\n\n![](https://cdn-images-1.medium.com/max/720/0*JwGzVZvQ6yWNODLM.png)\n\n* The skater initially began as a part of LIME and was then developed as a independent structure with an assortment of features and abilities to interpretation independent models for all black box. The undertaking started as an research idea to discover courses for better interpretability (ideally human interpretability) to predict \"Black Boxes\" for scientists and researchers.\n\nDocumentation for Reading\n\n|||\n|---------------------------------------------------------------------------------------|----------------------------------------------|\n| [**Overview**](https://datascienceinc.github.io/Skater/overview.html) | Introduction to the Skater library |\n| [**Installing**](https://datascienceinc.github.io/Skater/install.html) | How to install the Skater library |\n| [**Tutorial**](https://datascienceinc.github.io/Skater/tutorial.html) | Steps to use Skater effectively. |\n| [**API Reference**](https://datascienceinc.github.io/Skater/api.html) | The detailed reference for Skater's API. |\n| [**Contributing**](https://github.com/datascienceinc/Skater/blob/master/CONTRIBUTING.rst) | Guide to contributing to the Skater project. |\n\n### **Skater Technique**\n\n![](https://cdn-images-1.medium.com/max/900/1*qrdvdL1aIElWDw58ghFK4Q.png)"},{"metadata":{"_uuid":"bf9aa0b42a37da6d6720a03a8aa0446d93cf0c56"},"cell_type":"markdown","source":"### **Installation**"},{"metadata":{"trusted":true,"_uuid":"b62ac7a2598cab496b20cffef09aeecdd5882434","_kg_hide-output":true},"cell_type":"code","source":"!conda install -c conda-forge Skater -y","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b3b66408e642185700dcfa24cbd992a032704f7c"},"cell_type":"markdown","source":"### **1.Load Packages**"},{"metadata":{"trusted":true,"_uuid":"5f0aa5e98cc1d1645963ea8c146e8c3cf254d94e"},"cell_type":"code","source":"%matplotlib inline\nimport warnings\nwarnings.filterwarnings('ignore')\nimport matplotlib.pyplot as plt\nimport pandas as pd\n# Reference for customizing matplotlib: https://matplotlib.org/users/style_sheets.html\nplt.style.use('ggplot')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6323bebf3507f965dfe5a3a960c527aa90b985d0"},"cell_type":"code","source":"from sklearn.datasets import load_breast_cancer\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.ensemble import RandomForestClassifier, VotingClassifier\n\nfrom skater.core.explanations import Interpretation\nfrom skater.model import InMemoryModel","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d17e645ebdae9ea6e87f3c9541e860903e80dd5d"},"cell_type":"markdown","source":"### **2.Load Data**"},{"metadata":{"trusted":true,"_uuid":"0500391d6a2d44396718638e5be790c2766c9839","_kg_hide-output":true},"cell_type":"code","source":"data = load_breast_cancer()\n# Description of the data\nprint(data.DESCR)\npd.DataFrame(data.target_names)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"62b852ce9b3df90df9a4bcc21d1e1219b0d3cfe5"},"cell_type":"markdown","source":"### **3.Train Test Split**"},{"metadata":{"trusted":true,"_uuid":"69db76152dfb2423e2a5fe0df2334b931b57536d"},"cell_type":"code","source":"X = data.data\ny = data.target\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bc547a1b2f829ae2192e70bafd349813308290d8"},"cell_type":"markdown","source":"### **4.Model Training**"},{"metadata":{"trusted":true,"_uuid":"fe980655a784d685c0387ead78e2f829d90ffc8c"},"cell_type":"code","source":"def model_training(X_train, y_train):\n    clf1 = LogisticRegression(random_state=1)\n    clf2 = RandomForestClassifier(random_state=1)\n    clf3 = GaussianNB()\n    eclf = VotingClassifier(estimators=[('lr', clf1), ('rf', clf2), ('gnb', clf3)], voting='soft')\n    eclf = eclf.fit(X_train, y_train)\n    return (clf1,clf2,clf3, eclf)\n\nclf1,clf2,clf3, eclf = model_training(X_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"948727665afc5528e06f6d584851e0dc47b4e2ef"},"cell_type":"code","source":"def train_all_model(clf1,clf2,clf3, X_train,y_train):\n    clf1 = clf1.fit(X_train, y_train)\n    clf2 = clf2.fit(X_train, y_train)\n    clf3 = clf3.fit(X_train, y_train)\n    models = {'lr':clf1, 'rf':clf2, 'gnb':clf3, 'ensemble':eclf}\n    return (clf1,clf2,clf3, models)\n\nclf1,clf2,clf3, models = train_all_model(clf1,clf2,clf3,X_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"850bbf5e8f9cadc0713d857375c95b43bc581fff"},"cell_type":"markdown","source":"### **5.Model Interpretation**"},{"metadata":{"trusted":true,"_uuid":"42131be7296fc5aaf4a3180b6f0ae054d1776e83"},"cell_type":"code","source":"# Ensemble Classifier does not have feature importance enabled by default\nf, axes = plt.subplots(2, 2, figsize = (26, 18))\n\nax_dict = {'lr':axes[0][0],'rf':axes[1][0],'gnb':axes[0][1],'ensemble':axes[1][1]}\ninterpreter = Interpretation(X_test, feature_names=data.feature_names)\n\nfor model_key in models:\n    pyint_model = InMemoryModel(models[model_key].predict_proba, examples=X_test)\n    ax = ax_dict[model_key]\n    interpreter.feature_importance.plot_feature_importance(pyint_model, ascending=True, ax=ax)\n    ax.set_title(model_key)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1f932aef3c847c56c5879cd703768d0d7f5e44e7"},"cell_type":"code","source":"# Before interpreting, lets check on the accuracy of all the models\nfrom sklearn.metrics import f1_score\nfor model_key in models:\n        print(\"Model Type: {0} -> F1 Score: {1}\".\n              format(model_key, f1_score(y_test, models[model_key].predict(X_test))))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd58abf440316f42943b01a80fb907f2a0df8b10"},"cell_type":"markdown","source":"### **6.Decision Boundaries**"},{"metadata":{"trusted":true,"_uuid":"097dc328dcdf6ee8a310d00522d78a8683e05549"},"cell_type":"code","source":"%matplotlib inline\nX_train = pd.DataFrame(X_train)\nX_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b3102ca20fa429c83ec41bbbfd6420a3d34751dc"},"cell_type":"markdown","source":"### **7.Partial Dependence Plots with Interactive slider for controlling grid resolution**"},{"metadata":{"trusted":true,"_uuid":"604a512583306428c29e52e763da7bbfb32c1e0e"},"cell_type":"code","source":"def understanding_interaction():\n    pyint_model = InMemoryModel(eclf.predict_proba, examples=X_test, target_names=data.target_names)\n    # ['worst area', 'mean perimeter'] --> list(feature_selection.value)\n    interpreter.partial_dependence.plot_partial_dependence(list(feature_selection.value),\n                                                                    pyint_model, \n                                                                    grid_resolution=grid_resolution.value, \n                                                                    with_variance=True)\n        \n    # Lets understand interaction using 2-way interaction using the same covariates\n    # feature_selection.value --> ('worst area', 'mean perimeter')\n    axes_list = interpreter.partial_dependence.plot_partial_dependence([feature_selection.value],\n                                                                       pyint_model, \n                                                                       grid_resolution=grid_resolution.value, \n                                                                       with_variance=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"22f0176d6509825777a5b90d2cd40d43452b0684"},"cell_type":"markdown","source":"### **8.Understanding interaction using interactive widgets**"},{"metadata":{"_kg_hide-output":true,"trusted":true,"_uuid":"6f20d788d2da0a82b99f17e57101b66b7c83db80"},"cell_type":"code","source":"!conda install ipywidgets --yes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"89a7beef30ab0cbc1ee79bbc9016fbe0712cce3b"},"cell_type":"code","source":"# One could further improve this by setting up an event callback using\n# asynchronous widgets\nimport ipywidgets as widgets\nfrom ipywidgets import Layout\nfrom IPython.display import display\nfrom IPython.display import clear_output\ngrid_resolution = widgets.IntSlider(description=\"GR\", \n                                    value=10, min=10, max=100)\ndisplay(grid_resolution)\n\n# dropdown to select relevant features from the dataset\nfeature_selection = widgets.SelectMultiple(\n    options=tuple(data.feature_names),\n    value=['worst area', 'mean perimeter'],\n    description='Features',\n    layout=widgets.Layout(display=\"flex\", flex_flow='column', align_items = 'stretch'),\n    disabled=False,\n    multiple=True\n)\ndisplay(feature_selection)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"be707426654018c907c2637b275ddb1f1119a9db"},"cell_type":"code","source":"# Reference: http://ipywidgets.readthedocs.io/en/latest/examples/Widget%20Events.html\nbutton = widgets.Button(description=\"Generate Interactions\")\ndisplay(button)\n\ndef on_button_clicked(button_func_ref):\n    clear_output()\n    understanding_interaction()\n\nbutton.on_click(on_button_clicked)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"df081feb04b27f0847ca1989bc2b533b9358b8ff"},"cell_type":"markdown","source":"### **9.Evaluate a point locally, lets apply Local Interpretation using an interactive slider**"},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"eb4fbae4f1f477e8dd6446cd13f910fc2cebc06a"},"cell_type":"code","source":"from skater.core.local_interpretation.lime.lime_tabular import LimeTabularExplainer\nfrom IPython.display import display, HTML, clear_output\nint_range = widgets.IntSlider(description=\"Index Selector\", value=9, min=0, max=100)\ndisplay(int_range)\n\ndef on_value_change(change):\n    index = change['new']\n    exp = LimeTabularExplainer(X_test, \n                           feature_names=data.feature_names, \n                           discretize_continuous=False, \n                           class_names=['p(Cancer)-malignant', 'p(No Cancer)-benign'])\n    print(\"Model behavior at row: {}\".format(index))\n    # Lets evaluate the prediction from the model and actual target label\n    print(\"prediction from the model:{}\".format(eclf.predict(X_test[index].reshape(1, -1))))\n    print(\"Target Label on the row: {}\".format(y_test.reshape(1,-1)[0][index]))\n    clear_output()\n    display(HTML(exp.explain_instance(X_test[index], models['ensemble'].predict_proba).as_html()))\n    \nint_range.observe(on_value_change, names='value')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"16f32edbd134876ff96a03121cc478b9e4e22303"},"cell_type":"markdown","source":"Through global and local interpretation, one can understand the interaction between independent (input features) and dependent variables (P (cancer) / P (no cancer) by interrogating the behavior of the model, and the importance of features helps us understand better The weights of the variables used are understood by the predictive model: partial dependency graphs and LIME help to understand the interactions between variables that influence prediction."},{"metadata":{"_uuid":"115bd9f5ecb13268785f065ba7bcde3628956614"},"cell_type":"markdown","source":"---\n\n# **3.SHAP**\n\n---\n\n**More Example for Learning: https://github.com/slundberg/shap**\n\n* Shap values show how much a given feature changed our prediction (compared to if we made that prediction at some baseline value of that feature).\n* For example, consider an ultra-simple model: $$y = 4 * x1 + 2 * x2$$\n\n* If $x1$ takes the value 2, instead of a baseline value of 0, then our SHAP value for $x1$ would be 8 (from 4 times 2).\n* These are harder to calculate with the sophisticated models we use in practice. But through some algorithmic cleverness, Shap values allow us to decompose any prediction into the sum of effects of each feature value, yielding a graph like this:\n\n![](https://camo.githubusercontent.com/06cf2db4ee53c00baa7d20c8b1ccfcccdfd964f1/68747470733a2f2f692e696d6775722e636f6d2f4a56443255376b2e706e67)\n\n\n### **1.Read Packages**"},{"metadata":{"trusted":true,"_uuid":"05cfafeccc1a894a5ff30c551bdc0e70e22a08fa"},"cell_type":"code","source":"from IPython.display import Image\nimport numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split\nimport xgboost\nimport matplotlib.pyplot as plt\nfrom IPython.display import Image","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8ea8760a770bceaa5e3da6290a2376c5675f96a5"},"cell_type":"markdown","source":"### **2.Load Dataset**"},{"metadata":{"trusted":true,"_uuid":"67a616876427614788226039d065337c36a93db5"},"cell_type":"code","source":"data = pd.read_csv(\"../input/FIFA 2018 Statistics.csv\")\ny = (data['Man of the Match'] == \"Yes\")  # Convert from string \"Yes\"/\"No\" to binary\nfeature_names = [i for i in data.columns if data[i].dtype in [np.int64, np.int64]]\nX = data[feature_names]\nX_train, X_val, y_train, y_val = train_test_split(X, y, random_state=9487)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"de47631f1994759fe6218c5839cb79f7e66b2767"},"cell_type":"markdown","source":"### **3.Parameter Tuning**"},{"metadata":{"trusted":true,"_uuid":"d5848b0a6d4103f89c89966c171c5e6a96c4f4f1"},"cell_type":"code","source":"params = {'base_score': 0.5,\n         'booster': 'gbtree',\n         'colsample_bylevel': 1,\n         'colsample_bytree': 1,\n         'gamma': 0,\n         'learning_rate': 0.05,\n         'max_delta_step': 0,\n         'max_depth': 3,\n         'min_child_weight': 1,\n         'missing': None,\n         'n_estimators': 400,\n         'n_jobs': 1,\n         'objective': 'binary:logistic',\n         'random_state': 0,\n         'reg_alpha': 0,\n         'reg_lambda': 1,\n         'scale_pos_weight': 1,\n         'seed': 0,\n         'silent': True,\n         'subsample': 1}\n\nparams['eval_metric'] = 'auc'","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fcf2f79f11848c7ceb33991d5ce6a6c45496f047"},"cell_type":"markdown","source":"### **4.Model Training**"},{"metadata":{"trusted":true,"_uuid":"6d5209ef471252a092262a9565c5af0a4600704a"},"cell_type":"code","source":"d_train = xgboost.DMatrix(X_train, y_train)\nd_val = xgboost.DMatrix(X_val, y_val)\nwatchlist = [(d_train, \"train\"), (d_val, \"valid\")]\n\n#train model\n\nmodel = xgboost.train(params, d_train, num_boost_round=2000, evals=watchlist, early_stopping_rounds=100, verbose_eval=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0218a76f0f4aeb6fd0ad451181802a0f848ce57d"},"cell_type":"code","source":"# Simple check what model ran out of things\ndata_for_prediction = xgboost.DMatrix(X_train.iloc[[83],:])  # use 1 row of data here. Could use multiple rows if desired\nmodel.predict(data_for_prediction)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7a1f2f013cc61849fcd5af1574288b465b6aa07"},"cell_type":"markdown","source":"### **5.Model Interpretability**"},{"metadata":{"trusted":true,"_uuid":"ec1f6378d5599f43bb4f74349f4228eac05ab186"},"cell_type":"code","source":"import shap  # package used to calculate Shap values\n\n# Create object that can calculate shap values\nexplainer = shap.TreeExplainer(model)\n# Calculate Shap values\nshap_values = explainer.shap_values(X_train)\nshap.initjs()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d3bfffe2a28960f0768e262f717d70da88c961e4"},"cell_type":"markdown","source":"### **6.Global Intepretability**"},{"metadata":{"trusted":true,"_uuid":"f2ab2d2fde6d4f5a6e3bda99d1b179b77f30c474"},"cell_type":"code","source":"shap.summary_plot(shap_values, X_train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2ec1c859f8d1625f9655b36f49a34e049dcb4ac0"},"cell_type":"markdown","source":"### **7.ForcePlot**"},{"metadata":{"trusted":true,"_uuid":"0062cdbde70a644b6121e3b26af4c2eac6091aa7"},"cell_type":"code","source":"data_for_prediction = xgboost.DMatrix(X_train.iloc[[10],:])  # use 1 row of data here. Could use multiple rows if desired\nprint(f\"The 85th data is predicted to be True's probability: {model.predict(data_for_prediction)}\")\nshap.force_plot(explainer.expected_value, shap_values[10,:], X_train.iloc[10,:])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4b8888b0193526d98abb0c323af634dbb0bbdc9c"},"cell_type":"code","source":"data_for_prediction = xgboost.DMatrix(X_train.iloc[[83],:])  # use 1 row of data here. Could use multiple rows if desired\nprint(f\"The 83rd data is predicted to be True's probability: {model.predict(data_for_prediction)}\")\nshap.force_plot(explainer.expected_value, shap_values[83,:], X_train.iloc[83,:])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e8674be7941e5b87c3907685e8d510ae4bb70a66"},"cell_type":"markdown","source":"### **8.Output Value **"},{"metadata":{"trusted":true,"_uuid":"76bc67f67a726e9c50cf64aa4b6fdd5cabad9e15"},"cell_type":"code","source":"plt.figure(figsize=(20,8))\nxs = np.linspace(-5,5,100)\nplt.xlabel(\"Log odds of winning\")\nplt.ylabel(\"Probability of winning\")\nplt.title(\"Log odds & prob of winning convert\")\nplt.plot(xs, 1/(1+np.exp(-xs)))\n\nnew_ticks = np.linspace(-5, 5, 11)\nplt.xticks(new_ticks)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"719820f9f1755c3b5749256b3fbdd0940f4476e2"},"cell_type":"markdown","source":"### **9.Aggregated force_plot**"},{"metadata":{"trusted":true,"_uuid":"117840f48b28d062d4228b2b6c86ebdfb8412438"},"cell_type":"code","source":"shap.force_plot(explainer.expected_value, shap_values, X_train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c4c1c2a0c7888ee2cd98003be533915571cac37b"},"cell_type":"markdown","source":"### **10.Dependence_plot**"},{"metadata":{"trusted":true,"_uuid":"94902e92a0859f7dfce7fec46ec0cd67caee4d14"},"cell_type":"code","source":"shap.dependence_plot('Ball Possession %', shap_values, X_train, interaction_index=\"Goal Scored\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"198373cd77ffa2b49d5188f4bf110d2d1c091c70"},"cell_type":"markdown","source":"**Note :** This tutorial made for learning and all content with proper references listed below. still found any feed back on comment.\n\n### **References**\n\n1. https://github.com/slundberg/shap\n2. https://github.com/datascienceinc/Skater\n3. https://eli5.readthedocs.io/en/latest/"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}