{"cells":[{"metadata":{},"cell_type":"markdown","source":"<h1> Brief Introduction </h1>\nIn this project, my main aim is to show ways to go deep into the data story-telling even though the dataset is small. Also, I will work on a model that could give us an approximation as to what will be the charges of the patients. Nevertheless, we must go deeply into what factors influenced the charge of a specific patient. In order to do this we must look for patterns in our data analysis and gain extensive insight of what the data is telling us.  Lastly, we will go step by step to understand the story behind the patients in this dataset only through this way we could have a better understanding of what features will help our model have a closer accuracy to the true patient charge. \n\n<h4>Things to Notice</h4>\nI will importing the library bassed upon the requirement, so that it will easy for you to understand which library is used where and for what purpose. "},{"metadata":{},"cell_type":"markdown","source":"## Data Exploration\n\nHere in this section, I will try to explore the data as much as I can. In other words I will try to find hidden patterns. I will also try to give plausible explaination for each of the steps and interpret the graph as much as possible. "},{"metadata":{"_cell_guid":"","_uuid":"","trusted":true},"cell_type":"code","source":"import pandas as pd\ninsurance = pd.read_csv(\"../input/insurance/insurance.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Finding the missing values\n\n- It seems that the data does not have any missing values. "},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.isna().sum()/len(insurance)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Describe** helps us to find statistical information about the dataset that we are working with. "},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The important thing to notice is the standard deviation. I believe that the standard deviation lies at the core of statistics. And if we somehow manage to get this thing we have won the battle. \n\n**Note**: A good standard deviation is somewhere between 0 to 1. Even if it is 1 or bit higher than it, it is manageable."},{"metadata":{},"cell_type":"markdown","source":"### Visualisation\n\nIn order to get a good picture of what is happen under the hood we need to take help of the visualisation tools. With matplotlib and seaborn we can manage to do that. "},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#just a trail\nsns.distplot(insurance.age)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### A note about Logs or Logarithms\n- Using Logarithms: Logarithms helps us have a normal distribution which could help us in a number of different ways such as outlier detection, implementation of statistical concepts based on the central limit theorem and for our predictive model in the foreseen future. \n- Here we will observe that the stardard deviation after using log is below 1 standard deviation. This will help us to find the error while testing the ML model so keep that in mind. \n- Below is the example of how Logs can transform the distribution to normal distribution. "},{"metadata":{"trusted":true},"cell_type":"code","source":"print(np.std(np.log(insurance.charges)))\nsns.distplot(np.log(insurance.charges))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Age Analysis:\n\nTurning Age into Categorical Variables:\n- Young Adult: from 18 - 35\n- Senior Adult: from 36 - 55\n- Elder: 56 or older\n- Share of each Category: Young Adults (42.9%), Senior Adults (41%) and Elder (16.1%)"},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance['age_cat'] = np.nan\nlst = [insurance]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lst","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for col in lst:\n    col.loc[(col['age'] >= 18) & (col['age'] <= 35), 'age_cat'] = 'Young Adult'\n    col.loc[(col['age'] > 35) & (col['age'] <= 55), 'age_cat'] = 'Senior Adult'\n    col.loc[col['age'] > 55, 'age_cat'] = 'Elder'\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(lst)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"age_cat = insurance.age_cat.map({'Young Adult':0, \n 'Senior Adult':1,\n 'Elder':2})\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels = insurance[\"age_cat\"].unique()\namount = insurance[\"age_cat\"].value_counts().tolist()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"my_circle=plt.Circle( (0,0), 0.7, color='white')\n\nplt.figure(figsize=(10,10))\nplt.pie(amount, labels=labels, colors=['red','green','blue'])\n\np=plt.gcf()\np.gca().add_artist(my_circle)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.distplot(insurance.bmi)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Is there a Relationship between BMI and Age\n- BMI frequency: Most of the BMI frequency is concentrated between 28 - 32.\n- Correlations Age and charges have a correlation of 0.29 while bmi and charges have a correlation of 0.19\n- Relationship betweem BMI and Age: The correlation for these two variables is 0.10 which is not that great. Therefore, we can disregard that age has a huge influence on BMI.\n- Also, the influence of BMI and Age is very little. Which means these two factors does effect charges as much as we wanted."},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.heatmap(insurance.corr())\nplt.show()\n\nprint('*'*100)\nprint(insurance.corr())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"young_adults = insurance[\"bmi\"].loc[insurance[\"age_cat\"] == \"Young Adult\"].values\nsenior_adult = insurance[\"bmi\"].loc[insurance[\"age_cat\"] == \"Senior Adult\"].values\nelders = insurance[\"bmi\"].loc[insurance[\"age_cat\"] == \"Elder\"].values","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**observations**: Young adults have extreme outliers. We need to deal with it. "},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.boxplot(data= [young_adults, senior_adult, elders])\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import statsmodels.api as sm\nfrom statsmodels.formula.api import ols\n\n\nmoore_lm = ols(\"bmi ~ age_cat\", data=insurance).fit()\nprint(moore_lm.summary())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label = LabelEncoder()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#sex\nlabel.fit(insurance.sex.drop_duplicates())\ninsurance.sex = label.transform(insurance.sex)\ninsurance.sex.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#smoker or non-smoker\ninsurance.smoker = label.fit_transform(insurance.smoker)\ninsurance.smoker.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#region\n\ninsurance.region = label.fit_transform(insurance.region)\ninsurance.region.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.corr()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"- Smoker shows the strong correlation with charges. That means that the smokers pay more treatment charges than anyone else. \n- Strong correlation suggest that as the independent variable increases it potential the dependent also get affected from the potential."},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.heatmap(insurance.corr())\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"moore_lm = ols(\"charges ~ smoker\", data=insurance).fit()\nprint(moore_lm.summary())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.distplot(insurance.charges)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.loc[(insurance.smoker == 1)].charges","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f = plt.figure(figsize=(20,10))\n\nax = f.add_subplot(121)\nsns.distplot(insurance.loc[(insurance.smoker == 1)].charges, ax=ax)\nax.set_title('Smokers')\n\n\nax = f.add_subplot(122)\nsns.distplot(insurance.loc[(insurance.smoker == 0)].charges, color='r', ax = ax)\nax.set_title('Non-Smokers')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Smoking patients spends much on treatment"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.catplot(x='smoker', kind='count', hue = 'sex', palette='PuBuGn_r', data=insurance)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f = plt.figure(figsize=(20,20))\n\nax = f.add_subplot(211)\nsns.boxenplot(x = 'age', y='charges', hue='sex', data=insurance, ax=ax)\n\nax = f.add_subplot(212)\nsns.scatterplot(x = 'charges', y='age', hue='smoker', data=insurance, ax=ax)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.distplot(insurance.age, color='r')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance[\"weight_condition\"] = np.nan\nlst = [insurance]\n\nfor col in lst:\n    col.loc[col[\"bmi\"] < 18.5, \"weight_condition\"] = \"Underweight\"\n    col.loc[(col[\"bmi\"] >= 18.5) & (col[\"bmi\"] < 24.986), \"weight_condition\"] = \"Normal Weight\"\n    col.loc[(col[\"bmi\"] >= 25) & (col[\"bmi\"] < 29.926), \"weight_condition\"] = \"Overweight\"\n    col.loc[col[\"bmi\"] >= 30, \"weight_condition\"] = \"Obese\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2, ax3) = plt.subplots(ncols=3, figsize=(18,8))\n\n# I wonder if the cluster that is on the top is from obese people\nsns.stripplot(x=\"age_cat\", y=\"charges\", data=insurance, ax=ax1, linewidth=1, palette=\"Reds\")\nax1.set_title(\"Relationship between Charges and Age\")\n\n\nsns.stripplot(x=\"age_cat\", y=\"charges\", hue=\"weight_condition\", data=insurance, ax=ax2, linewidth=1, palette=\"Set2\")\nax2.set_title(\"Relationship of Weight Condition, Age and Charges\")\n\nsns.stripplot(x=\"smoker\", y=\"charges\", hue=\"weight_condition\", data=insurance, ax=ax3, linewidth=1, palette=\"Set2\")\nax3.legend_.remove()\nax3.set_title(\"Relationship between Smokers and Charges\")\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import seaborn as sns\nsns.set(style=\"ticks\")\npal = [\"#FA5858\", \"#58D3F7\"]\n\nsns.pairplot(insurance, hue=\"smoker\", palette=pal)\nplt.title(\"Smokers\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2) = plt.subplots(ncols=2, figsize=(18,8))\nsns.scatterplot(x=\"bmi\", y=\"charges\", hue=\"weight_condition\", data=insurance, palette=\"Set1\", ax=ax1)\nax1.set_title(\"Relationship between Charges and BMI by Weight Condition\")\n\nsns.scatterplot(x=\"bmi\", y=\"charges\", hue=\"smoker\", data=insurance, palette=\"Set1\", ax=ax2)\nax2.set_title(\"Relationship between Charges and BMI by Smoking Condition\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.scatterplot(x='children', y='age', data=insurance, hue='charges')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.children.unique()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.hist(insurance.children)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.barplot(insurance.children, insurance.charges)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.violinplot(x='children', y='charges', data=insurance)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.boxplot(insurance.children)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.children.std()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Unsupervised Learning"},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.cluster import KMeans","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cluster = KMeans(n_clusters=3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X = insurance.drop(['age_cat', 'weight_condition' ], axis=1)\ny = insurance.charges","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cluster.fit(X)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cluster.cluster_centers_","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X.values[:,0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(12,8))\n\nplt.scatter(X.values[:,2], X.values[:,6], c=cluster.labels_, cmap=\"Set1_r\", s=25)\nplt.scatter(cluster.cluster_centers_[:,2] ,cluster.cluster_centers_[:,6], color='black', marker=\"o\", s=250)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Feature Engineering"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.heatmap(insurance.corr())\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X = insurance.drop('region', axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.heatmap(X.corr())\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.scatterplot(x='children', y='bmi', hue='weight_condition', data=insurance)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X.std()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10, 12))\nplt.boxplot(insurance.bmi)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Removing Outliers"},{"metadata":{"trusted":true},"cell_type":"code","source":"new_bmi = X.bmi.values\nq25, q75 = np.percentile(new_bmi, 25), np.percentile(new_bmi, 75)\nprint(f'Quartile 25: {q25} | Quartile 75: {q75}')\nnew_bmi_iqr = q75 - q25\nprint(f'iqr: {new_bmi_iqr}')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"new_bmi_cutoff = new_bmi_iqr * 1.5\nnew_bmi_lower, new_bmi_upper = q25 - new_bmi_cutoff, q75 + new_bmi_cutoff\nprint('Lower: ', new_bmi_lower)\nprint('Upper :', new_bmi_upper)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"outliers = [x for x in new_bmi if x<new_bmi_lower or x>new_bmi_upper]\noutliers, len(outliers)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"final_df = X.drop(X[(X.bmi>new_bmi_upper) | (X.bmi<new_bmi_lower)].index)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,15))\nplt.boxplot(final_df.bmi)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"new_age = X.age.values\nq25, q75 = np.percentile(new_age, 25), np.percentile(new_age, 75)\nprint(f'Quartile 25: {q25} | Quartile 75: {q75}')\nnew_age_iqr = q75 - q25\nprint(f'iqr: {new_age_iqr}')\n\nnew_age_cutoff = new_age_iqr * 1.5\nnew_age_lower, new_age_upper = q25 - new_age_cutoff, q75 + new_age_cutoff\nprint('Lower: ', new_age_lower)\nprint('Upper :', new_age_upper)\n\noutliers = [x for x in new_age if x<new_age_lower or x>new_age_upper]\noutliers, len(outliers)\n\nfinal_df = X.drop(X[(X.age>new_age_upper) | (X.age<new_age_lower)].index)\n\nplt.figure(figsize=(10,15))\nplt.boxplot(final_df.age)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Scale = StandardScaler()\nfinal_df.bmi = Scale.fit_transform(final_df.bmi.values.reshape(-1,1))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"final_df.std()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Modeling \n## Linear Regression"},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LinearRegression","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import math\ndef rmse(x,y): return math.sqrt(((x-y)**2).mean())\n\ndef print_score(m):\n    res = [rmse(m.predict(X_train), y_train), rmse(m.predict(X_test), y_test),\n                m.score(X_train, y_train), m.score(X_test, y_test)]\n    if hasattr(m, 'oob_score_'): res.append(m.oob_score_)\n    print(res)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X = final_df.drop(['charges', 'age_cat', 'weight_condition'], axis=1)\ny = np.log(final_df.charges)\n\n\nX_train, X_test, y_train, y_test = train_test_split(X,y, random_state = 23, test_size=0.3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model=LinearRegression()\nmodel.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print_score(model)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.intercept_","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.coef_","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nplt.plot(X_train, y_train, 'ro')\nplt.plot(X_train,model.coef_[0]*X_train + model.intercept_)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Random Forest"},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nmodel=RandomForestRegressor()\nmodel.fit(X_train, y_train)\nmodel.score(X_test, y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model=RandomForestRegressor(n_estimators=25, n_jobs=-1, max_depth=6, max_features=0.5)\nmodel.fit(X_train, y_train)\nmodel.score(X_test, y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rmse(model.predict(X_test), y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.tree import export_graphviz\nfrom IPython import display\nfrom io import StringIO\nimport re","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import graphviz\nimport IPython\n\ndef draw_tree(t, df, size=10, ratio=0.6, precision=0):\n    \"\"\" Draws a representation of a random forest in IPython.\n    Parameters:\n    -----------\n    t: The tree you wish to draw\n    df: The data used to train the tree. This is used to get the names of the features.\n    \"\"\"\n    s=export_graphviz(t, out_file=None, feature_names=df.columns, filled=True,\n                      special_characters=True, rotate=True, precision=precision)\n    IPython.display.display(graphviz.Source(re.sub('Tree {',\n       f'Tree {{ size={size}; ratio={ratio}', s)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"draw_tree(model.estimators_[0], X, precision=5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print_score(model)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"np.exp(model.predict([[30, 0, 0.4, 3, 0]]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"insurance.loc[(insurance.age == 30) & (insurance.bmi<=20)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Feature Importance"},{"metadata":{"trusted":true},"cell_type":"code","source":"feature_importances = pd.DataFrame(model.feature_importances_,\n                                   index = X.columns,\n                                    columns=['importance']).sort_values('importance',  ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"feature_importances.plot.barh(figsize=(15,8))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}