{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nI copied [Interactive Porto Insights - A Plot.ly Tutorial](https://www.kaggle.com/arthurtok/interactive-porto-insights-a-plot-ly-tutorial)\n\nThis completition is hosted by the third largest insurance company in Brazil: [Porto Seguro](https://en.wikipedia.org/wiki/Porto_Seguro_S.A.) with the task of predicting the probability that a driver will initiate an insurance claim in the next year.\n\nThis notebook will aim to provides some interactive charts and analysis of the competition data by way of the Python cisualisation library Plot.ly and hopefully bring some insights and beautiful plots that others can take replicate. Plot.ly is one of the main products offered by the software company - Plotly which specializes in providing online graphical and statistical visualisations (charts and dashboards) as wll as providing an API to a whole rich suite of programming languages and tools such as Pyhons, R, Matlab, Node.js etc.\n\n----\nListed below for easy convenience are links to the various Plotly plots in this notebook:\n\n- Simple horizontal bar plot - Used to inspect the Target variable distribution\n- Correlation Heatmap plot - Inspect the correlation between the different features\n- Scatter plot - Compare the feature importances generated by Random Forest and Gradient_Boosted mode\n- Vertical bar blot - List in Decending order, the importances of the various features\n- 3D Scatter plot\n\nThe themes in this notebook can be briefly summarized follows:\n\n**[1. Data Quality Checks](https://www.kaggle.com/code/arthurtok/interactive-porto-insights-a-plot-ly-tutorial/notebook#quality)** - Visualising and evaluating all missing/Null values (values that are -1)\n\n**2. Feature inspection and filtering** - Correlation and feature Mutual information plots against the target variable. Inspection of the Binary, categorical and other variables.\n\n**3. Feature importance ranking via learning models** /n Building a Random Forest and Gradient Boosted model to help us rank features based off the learning process.\n\nLet's Go","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"# Let us load in the relevant Python modules\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport plotly.offline as py\npy.init_notebook_mode(connected = True)\nimport plotly.graph_objs as go\nimport plotly.tools as tls\nimport warnings\nwarnings.filterwarnings('ignore')\nfrom collections import Counter\nfrom sklearn.feature_selection import mutual_info_classif","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:15:55.139796Z","iopub.execute_input":"2022-07-08T05:15:55.140173Z","iopub.status.idle":"2022-07-08T05:15:56.078426Z","shell.execute_reply.started":"2022-07-08T05:15:55.140143Z","shell.execute_reply":"2022-07-08T05:15:56.077276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us load in traing data provied using Pandas","metadata":{}},{"cell_type":"code","source":"path = '../input/porto-seguro-safe-driver-prediction/'\ntrain = pd.read_csv(path+ 'train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:15:56.081024Z","iopub.execute_input":"2022-07-08T05:15:56.081978Z","iopub.status.idle":"2022-07-08T05:16:00.361517Z","shell.execute_reply.started":"2022-07-08T05:15:56.081924Z","shell.execute_reply":"2022-07-08T05:16:00.360387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Taking a look at how many rows and columns the train dataset contains\nrows = train.shape[0]\ncolumns = train.shape[1]\nprint('The train dataset contains {} rows and {} columns'.format(rows, columns))","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:16:00.362872Z","iopub.execute_input":"2022-07-08T05:16:00.363245Z","iopub.status.idle":"2022-07-08T05:16:00.369415Z","shell.execute_reply.started":"2022-07-08T05:16:00.363193Z","shell.execute_reply":"2022-07-08T05:16:00.368241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. Data Quality checks\n\n**Null or missing values check**\n\nAs part of our quality checks, let us quick look at whether are any null values in the train dataset as follows:","metadata":{}},{"cell_type":"code","source":"# any() applied twice to check run the isnull check across all columns.\ntrain.isnull().any().any()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:16:00.372106Z","iopub.execute_input":"2022-07-08T05:16:00.372547Z","iopub.status.idle":"2022-07-08T05:16:00.402843Z","shell.execute_reply.started":"2022-07-08T05:16:00.372506Z","shell.execute_reply":"2022-07-08T05:16:00.401719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Our null values check returns False but however, this does not really mean that this caes has been closed as the data is also decirbed as [\"Values of -1 indicate that the feature was missing from the observation.\"](https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/data) Therefore I take it that Porto Seguro gas simply conducted a blanket replacement of all null values in the data with the value of -1. Let us now inspect if there any missing values in the data.","metadata":{}},{"cell_type":"markdown","source":"Here we can see that which columns contained -1 in their values so we could easily for example make a blanket repalacement of all -1 with nulls first as follows:","metadata":{}},{"cell_type":"code","source":"train_copy = train\ntrain_copy = train_copy.replace(-1, np.NaN)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:16:00.403999Z","iopub.execute_input":"2022-07-08T05:16:00.404350Z","iopub.status.idle":"2022-07-08T05:16:00.551730Z","shell.execute_reply.started":"2022-07-08T05:16:00.404321Z","shell.execute_reply":"2022-07-08T05:16:00.550789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we can use resident Kaggler's [Alekssy Bilougur](https://www.kaggle.com/residentmario) - creator of the \"Missingno\" package which is a most useful and convenient tool in visualising missing values in the dataset, so check it out.","metadata":{}},{"cell_type":"code","source":"import missingno as msno\n# Nullity or missing values by columns\nmsno.matrix(df= train_copy.iloc[:,2:39], figsize= (20, 14), color= (0.42, 0.1, 0.05))","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:16:00.554402Z","iopub.execute_input":"2022-07-08T05:16:00.554783Z","iopub.status.idle":"2022-07-08T05:16:06.567873Z","shell.execute_reply.started":"2022-07-08T05:16:00.554742Z","shell.execute_reply":"2022-07-08T05:16:06.566785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, the missing values now become much mor apparent and clear when we visualise it, where the empty white bands (data that is missing) superposed on the vertical dark red bands (non-missing data) reflect the nullity of  the data in that particular column. In this instance, we can observe that there are 7 features out of the 59 total features ( although as rightly pointed out by Justin Nafe in the comments section there are really a grand total fo 30 columns with missing vlaues) that actually contained null vales. This is due to the fact that the missingno matrix plot can only comfortable fit in approximately 40 odd features to one plot after which some columns may be excluded, and hence the remaining 5 null columns have been excluded. To visualize all nulls, try changing the figsize argument as wll as tweaking how we slice the dataframe.\n\nFor the 7 null columns thatt we are able to observe, they are hece listed here as follows:\n\n**ps_ind_05_cat | ps_reg_03 | ps_car_03_cat | ps_car_05_cat | ps_car_07_cat | ps_car_09_cat | ps_car_14**\n\nMost of the missing values occur in the columns suffixed with \\_cat. One should really take further note of the columns ps_reg_03, ps_car_03_cat and ps_car_05_cat. Evinced from the ratio of white to dark bands, it is very apparent that a big majority of values are missing from these 3 columns, and therefore a blanket replacement of -1 for the nulls might not be a very good strategy.","metadata":{}},{"cell_type":"markdown","source":"**Target variable inspection**\n\nAnother standard check normally conducted on the data is with regards to our target variable, where in this case, the columns in conveniently titled \"target\". The target value also comes by the moniker of class/label/correct answer and is used in superviesd learning models along with the corresponding data that is given (in our case all our train data except the id column) to learn the function that best maps the data to our target in the hope that this learned function can generalize and predict well with new unseen data.","metadata":{}},{"cell_type":"code","source":"data = [go.Bar(x = train['target'].value_counts().index.values,\n               y = train['target'].value_counts().values,\n               text = 'Distribution of target variable')]\nlayout = go.Layout(title = 'Target variable distribution')\n\nfig = go.Figure(data= data, layout= layout)\npy.iplot(fig, filename = 'basic-bar')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:22:03.082688Z","iopub.execute_input":"2022-07-08T05:22:03.083104Z","iopub.status.idle":"2022-07-08T05:22:03.852291Z","shell.execute_reply.started":"2022-07-08T05:22:03.083068Z","shell.execute_reply":"2022-07-08T05:22:03.851125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Hmmn, the target vairable is rather imbalanced so it night be something to keep in mind. An imbalanced target will prove quite","metadata":{}},{"cell_type":"markdown","source":"**Datatype check**\n\nThis check is carried out to see what kind of datatypes the train set is comprised of : integers or characters or floats just to gain a better overview of the data we were provided with. One trick to obtain counts of the unique types in a python sequenc is to use the Counter method, when you import the Collections module as follows:","metadata":{}},{"cell_type":"code","source":"Counter(train.dtypes.values)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:26:50.138788Z","iopub.execute_input":"2022-07-08T05:26:50.139602Z","iopub.status.idle":"2022-07-08T05:26:50.146970Z","shell.execute_reply.started":"2022-07-08T05:26:50.139562Z","shell.execute_reply":"2022-07-08T05:26:50.145924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As alluded to above, there are a total of 59 columns that make up the train dataset and as we can observe from this check, the features/columns consist of only two datatypes - integer and floats.\n\nAnother point to note is that Porto Seguro has actually provided us data with headers that come suffixed with abbreviations such as \"\\_bin\", \"\\_cat\", \"\\_reg\", where they have given us a rough explanation that \\_bin indicates binary features while \\_cat indicates categorical features whilst the rest are either continous or ordinal features. Here I shall simplify this a vit further just bt looking at float values (probably only the continuous features) and integer datatypes (binary, categorical and ordinal features).","metadata":{}},{"cell_type":"code","source":"train_float = train.select_dtypes(include= ['float64'])\ntrain_int = train.select_dtypes(include= ['int64'])","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:31:51.160759Z","iopub.execute_input":"2022-07-08T05:31:51.161184Z","iopub.status.idle":"2022-07-08T05:31:51.274424Z","shell.execute_reply.started":"2022-07-08T05:31:51.161151Z","shell.execute_reply":"2022-07-08T05:31:51.273239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_float","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:31:56.685342Z","iopub.execute_input":"2022-07-08T05:31:56.685726Z","iopub.status.idle":"2022-07-08T05:31:56.710641Z","shell.execute_reply.started":"2022-07-08T05:31:56.685691Z","shell.execute_reply":"2022-07-08T05:31:56.709511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_int","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:32:04.754348Z","iopub.execute_input":"2022-07-08T05:32:04.755270Z","iopub.status.idle":"2022-07-08T05:32:04.815706Z","shell.execute_reply.started":"2022-07-08T05:32:04.755227Z","shell.execute_reply":"2022-07-08T05:32:04.814856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Correlation plots\n\nAs a starter, let us generate some linear correlation plots just to have a quick look at how a feature is linearly correlated to the next and perhaps start gaining some insights from here. At this juncture, I will use the seaborn statisical visualisation package to plot a heatmap of the correlation values. Coonveniently, Pandas dataframes come with the corr() method inbuilt, which calcuates the Pearson correlation. Also as convenient is Seaborn's way of invoking a correlation plot. Just literally the word \"heatmap\"","metadata":{}},{"cell_type":"code","source":"colormap = plt.cm.magma\nplt.figure(figsize= (16, 12))\nplt.title('Pearson correlation of continuous features', y= 1.05, size= 15)\nsns.heatmap(train_float.corr(), linewidth= 0.1, vmax= 1.0, square= True, \n           cmap= colormap, linecolor= 'white', annot= True)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T05:38:06.290172Z","iopub.execute_input":"2022-07-08T05:38:06.290613Z","iopub.status.idle":"2022-07-08T05:38:07.259901Z","shell.execute_reply.started":"2022-07-08T05:38:06.290577Z","shell.execute_reply":"2022-07-08T05:38:07.258680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the correlation plot, we can see that the majority of the features display zero or no correlation to one another. This is quite an interesting observation that will warrant our further investigation later down. For now, the paired features that display a positive linear correlation are listed as follows:\n\n**(ps_reg_01, ps_reg_03)**\n\n**(ps_reg_02, ps_reg_03)**\n\n**(ps_car_12, ps_car_13)**\n\n**(ps_car_13, ps_car_15)**","metadata":{}},{"cell_type":"markdown","source":"**Correlation of integer features**\n\nFor the columns of integer datatype, I shall now switch to using the Plotly library to show how one cna also generate a heatmap of correlation values interactively. Much like our earlier Plotly plot, we generate a heatmap object by simply invoking the \"go.Heatmap\". Here we have to provide values to three different axes, where x and y axes take in the column names while the correlation values is provided by the z-axis. The colorscale attribute takes in keywords that correspond to different color palettes that you will see in the heatmap where in this example, I have used the Grays colorscale (other include Portland and Viridis - try it for yourself).","metadata":{}},{"cell_type":"code","source":"# train_int = train_int.drop(['id', 'target'], axis= 1)\n# colormap = plt.cm.bone\n# plt.figure(figsize= (21, 16))\n# plt.title('Pearson correlation of categorical features', y= 1.05, size= 15)\n# sns.heatmap(train_cat.corr(), linewidth= 0.1, vamax= 1.0, square= True,\n#            cmap= colormap, linecolor= 'white', annot= False)\ndata= [go.Heatmap(z= train_int.corr().values,\n                 x= train_int.columns.values,\n                 y= train_int.columns.values,\n                 colorscale= 'Viridis',\n                 reversescale= False,\n                 text= None,\n                 opacity= 1.0)]\n\nlayout = go.Layout(title= 'Pearson Correlation of Integer-type features',\n                  xaxis= dict(ticks= '', nticks= 36),\n                  yaxis= dict(ticks= ''),\n                  width= 900, height= 700)\n\nfig= go.Figure(data= data, layout= layout)\npy.iplot(fig, filename= 'labelled-heatmap')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T06:02:16.095770Z","iopub.execute_input":"2022-07-08T06:02:16.096269Z","iopub.status.idle":"2022-07-08T06:02:20.237916Z","shell.execute_reply.started":"2022-07-08T06:02:16.096233Z","shell.execute_reply":"2022-07-08T06:02:20.236777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Similarly, we can observe that there are a huge number of columns that are not linearly correlated with each other at all, evident from the fact that we observe quite a lot of 0 value cells in our correlation plot. This is quite a useful observation to us, especially if we are trying to perform dimensionality reduction transformations such as Principal Component Analysis (PCA), this would require a certain degree of correlation. We can note some features of interest are as follows:\n\n***Negatively correlated features*** : ps_ind_06_bin, ps_ind_07_bin, ps_ind_08_bin, ps_ind_09_bin\n\nOne interesting aspect to note is that in our earlier analysis on nullity, ps_car_03_cat and ps_car_05_cat were found to contain many missing or null valuse. Therefore it should come as no surprise that both these features show quite a strong positive linear correlation to each other on this basis, albeit one that may not really reflect the underlying truth for the data.","metadata":{}},{"cell_type":"markdown","source":"## Mutual Information plots\n\nMutual information is another useful tool as it allows one to inspect the mutual information between the target variable and the corresponding feature it is calculated against. For classification problems, we can conveniently call Sklearn's mutual_info_classif method which measures the dependency between two random variables and ranges from zero (where the random variables are independent of each other) to higher values (indicate some dependency). This therefore will help give us an idea of how musch information from the garget may be contained within the features.\n\nThe sklearn implementation of the mutual_info_classif function tells us that it \"relies on nonparametric methods based on entropy estimation from k-nearest neighbors distances\", where you can go into more detail on the official sklearn page in the [link here](http://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.mutual_info_classif.html#sklearn.feature_selection.mutual_info_classif)","metadata":{}},{"cell_type":"code","source":"mf = mutual_info_classif(train_float.values, train.target.values, n_neighbors= 3, random_state= 17)\nprint(mf)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T07:17:35.128020Z","iopub.execute_input":"2022-07-08T07:17:35.128445Z","iopub.status.idle":"2022-07-08T07:18:22.450785Z","shell.execute_reply.started":"2022-07-08T07:17:35.128403Z","shell.execute_reply":"2022-07-08T07:18:22.449597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Binary features inspection\n\nAnother aspect of the data that we may want to inspect would be the columns that only contain binary values, i.e where values take on only either of the two values 1 or 0. Proceeding, we store all columns that contain these binary values and then generate a vertical plotly barplot of these binary values as follows:","metadata":{}},{"cell_type":"code","source":"bin_col = [col for col in train.columns if '_bin' in col]\nzero_list= []\none_list= []\nfor col in bin_col:\n    zero_list.append((train[col]== 0).sum())\n    one_list.append((train[col]== 1).sum())","metadata":{"execution":{"iopub.status.busy":"2022-07-08T07:29:34.070850Z","iopub.execute_input":"2022-07-08T07:29:34.071285Z","iopub.status.idle":"2022-07-08T07:29:34.129278Z","shell.execute_reply.started":"2022-07-08T07:29:34.071248Z","shell.execute_reply":"2022-07-08T07:29:34.128075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trace1 = go.Bar(x= bin_col,\n               y= zero_list,\n               name= 'Zero count')\ntrace2 = go.Bar(x= bin_col,\n               y= one_list,\n               name= 'One count')\n\ndata= [trace1, trace2]\nlayout= go.Layout(barmode= 'stack',\n                 title= 'Count of 1 and 0 in binary variables')\n\nfig= go.Figure(data= data, layout= layout)\npy.iplot(fig, filename= 'stacked-bar')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T07:35:53.022128Z","iopub.execute_input":"2022-07-08T07:35:53.022520Z","iopub.status.idle":"2022-07-08T07:35:53.070150Z","shell.execute_reply.started":"2022-07-08T07:35:53.022486Z","shell.execute_reply":"2022-07-08T07:35:53.069187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we observe that there are 4 features: **ps_ind_10_bin, ps_ind_11_bin, ps_ind_12_bin, ps_ind_13_bin** which are completely dominated by zeros. This begs the question of whether these features are useful at all as they do not contain much information about the other class vis-a-vis-the target.","metadata":{}},{"cell_type":"markdown","source":"## Categorical and Ordinal feature inspection\n\nLet us first take a look at the features that are termed categorical as per their suffix \"\\_cat\".","metadata":{}},{"cell_type":"markdown","source":"## Feature importance via Random Forest\n\nLet us now implement a Random Forest model where we fit the training data with a Random Forest Classifier and look at the ranking of the features after the model has finished training. This is a quick way of using an ensemble model (ensemble of weak decision  tree learners applied under Bootstrap aggregated) which does not require much parameter tuning in obtaining useful feature importances and is alos pretty robust to target imbalances. We call the Random Forest as follows:","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nrf = RandomForestClassifier(n_estimators= 150, max_depth= 8, min_samples_leaf= 4,\n                           max_features= 0.2, n_jobs= -1, random_state= 0)\nrf.fit(train.drop(['id', 'target'], axis= 1), train.target)\nfeatures = train.drop(['id', 'target'], axis= 1).columns.values\nprint('---- Training Done ----')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T07:50:42.700693Z","iopub.execute_input":"2022-07-08T07:50:42.701082Z","iopub.status.idle":"2022-07-08T07:52:16.372562Z","shell.execute_reply.started":"2022-07-08T07:50:42.701050Z","shell.execute_reply":"2022-07-08T07:52:16.371276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Plot.ly Scatter Plot of feature importances**\n\nHaving trained the Random Forest, we can obtain the list of feature importances by invoking the attribute \"feature_importances_\" and plot our next Plotly plot, the Scatter plot.\n\nHere we invoke the command Scatter and as per the previous Plotly plots, we have to define our y and x-axes. However the one thing that we pay ateention to in scatter plots is the marker attribute. It is the marker attribute where we define and hence control the size, color and scale of  the scatter points embedded.","metadata":{}},{"cell_type":"code","source":"# Scatter plot\ntrace = go.Scatter(y = rf.feature_importances_,\n                  x = features,\n                  mode = 'markers',\n                  marker= dict(sizemode = 'diameter',\n                              sizeref= 1,\n                              size= 13,\n                              # size= rf.feature_importances_,\n                              # color= np.random.randn(500),  # set color equal to a variable\n                              color = rf.feature_importances_,\n                              colorscale = 'Portland',\n                              showscale = True),\n                  text = features)\ndata = [trace]\n\nlayout= go.Layout(autosize= True,\n                 title= 'Random Forest Feature Importance',\n                 hovermode= 'closest',\n                 xaxis= dict(ticklen= 5,\n                            showgrid= False,\n                            zeroline= False,\n                            showline= False),\n                 yaxis= dict(title= 'Feature Importance',\n                            showgrid = False,\n                            zeroline = False,\n                            ticklen = 5,\n                            gridwidth = 2),\n                 showlegend = False)\nfig = go.Figure(data= data, layout= layout)\npy.iplot(fig, filename= 'scatter2010')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T08:04:08.790824Z","iopub.execute_input":"2022-07-08T08:04:08.791246Z","iopub.status.idle":"2022-07-08T08:04:09.083748Z","shell.execute_reply.started":"2022-07-08T08:04:08.791197Z","shell.execute_reply":"2022-07-08T08:04:09.082900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Furthermore we could also display a sorted list of all the features ranked by order of their importance, from highest to lowest via the same plotly barplots as follows:","metadata":{}},{"cell_type":"code","source":"x, y = (list(x) for x in zip(*sorted(zip(rf.feature_importances_, features), reverse = False)))\n\ntrace2 = go.Bar(x = x,\n               y = y,\n               marker = dict(color= x,\n                            colorscale= 'Viridis',\n                            reversescale= True),\n               name= 'Random Forest Feature Importance',\n               orientation= 'h')\n\nlayout= dict(title= 'Barplot of Feature Importance',\n            width= 900,\n            height= 2000,\n            yaxis= dict(showgrid= False,\n                       showline= False,\n                       showticklabels= True,\n                       ))\n\nfig1 = go.Figure(data= [trace2])\nfig1['layout'].update(layout)\npy.iplot(fig1, filename= 'plots')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T08:14:40.539318Z","iopub.execute_input":"2022-07-08T08:14:40.539685Z","iopub.status.idle":"2022-07-08T08:14:40.694811Z","shell.execute_reply.started":"2022-07-08T08:14:40.539655Z","shell.execute_reply":"2022-07-08T08:14:40.693661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Decision Tree visualisation**\n\nOne other interesting trick or technique oft used would be to visualize the tree branches or decisions made by the model. For simplicity, I fit a decision tree (of max_depth = 3) and hence you only see 3 levels in the decision branch, use the export to graph visualization attribute in sklearn \"export_graphviz\" and then export and import the tree image for visualization in this notebook.","metadata":{}},{"cell_type":"code","source":"from sklearn import tree\nfrom IPython.display import Image as PImage\nfrom subprocess import check_call\nfrom PIL import Image, ImageDraw, ImageFont\nimport re\n\ndecision_tree = tree.DecisionTreeClassifier(max_depth = 3)\ndecision_tree.fit(train.drop(['id', 'target'],axis=1), train.target)\n\n# Export our trained model as a .dot file\nwith open(\"tree1.dot\", 'w') as f:\n     f = tree.export_graphviz(decision_tree,\n                              out_file=f,\n                              max_depth = 4,\n                              impurity = False,\n                              feature_names = train.drop(['id', 'target'],axis=1).columns.values,\n                              class_names = ['No', 'Yes'],\n                              rounded = True,\n                              filled= True )\n        \n#Convert .dot to .png to allow display in web notebook\ncheck_call(['dot','-Tpng','tree1.dot','-o','tree1.png'])\n\n# Annotating chart with PIL\nimg = Image.open(\"tree1.png\")\ndraw = ImageDraw.Draw(img)\nimg.save('sample-out.png')\nPImage(\"sample-out.png\",)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T08:44:50.227314Z","iopub.execute_input":"2022-07-08T08:44:50.227741Z","iopub.status.idle":"2022-07-08T08:44:53.438952Z","shell.execute_reply.started":"2022-07-08T08:44:50.227701Z","shell.execute_reply":"2022-07-08T08:44:53.437624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Importance via Gradient Boosting model\n\nJust for curiosity, let us try another learning method in getting our feature importances. This time, we use a Gradient Boosting classifier to fit to the training data. Gradient Boosting proceeds in a forward stage-wise fashion, where at each stage regression trees are fitted on the gradient of the loss function (which defaults to the deviance in Sklearn implementation).","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import GradientBoostingClassifier\ngb= GradientBoostingClassifier(n_estimators= 100, max_depth= 3, min_samples_leaf= 4,\n                              max_features= 0.2, random_state= 0)\ngb.fit(train.drop(['id', 'target'], axis= 1), train.target)\nfeatures = train.drop(['id', 'target'], axis= 1).columns.values\nprint('---- Training Done ----')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T08:51:06.254307Z","iopub.execute_input":"2022-07-08T08:51:06.254840Z","iopub.status.idle":"2022-07-08T08:52:13.117276Z","shell.execute_reply.started":"2022-07-08T08:51:06.254793Z","shell.execute_reply":"2022-07-08T08:52:13.116300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Scatter plot\ntrace = go.Scatter(y= gb.feature_importances_,\n                  x= features,\n                  mode= 'markers',\n                  marker= dict(sizemode= 'diameter',\n                              sizeref= 1,\n                              size= 13,\n                              color= gb.feature_importances_,\n                              colorscale= 'Portland',\n                              showscale= True),\n                  text= features)\n\ndata= [trace]\n\nlayout= go.Layout(autosize= True,\n                 title= 'Gradient Boosting Machine Feature Importance',\n                 hovermode= 'closest',\n                 xaxis= dict(ticklen= 5,\n                            showgrid= False,\n                            zeroline= False,\n                            showline= False),\n                 yaxis= dict(title= 'Feature Importance',\n                            showgrid= False,\n                            zeroline= False,\n                            ticklen= 5,\n                            gridwidth= 2),\n                 showlegend= False)\n\nfig = go.Figure(data= data, layout= layout)\npy.iplot(fig, filename= 'scatter2010')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T08:58:16.433613Z","iopub.execute_input":"2022-07-08T08:58:16.434446Z","iopub.status.idle":"2022-07-08T08:58:16.488384Z","shell.execute_reply.started":"2022-07-08T08:58:16.434410Z","shell.execute_reply":"2022-07-08T08:58:16.487223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x, y = (list(x) for x in zip(*sorted(zip(gb.feature_importances_, features), reverse= False)))\ntrace2 = go.Bar(x= x,\n               y= y,\n               marker= dict(color= x,\n                           colorscale= 'Viridis',\n                           reversescale= True),\n               name= 'Gradient Boosting Classifier Feature Importance',\n               orientation= 'h')\n\nlayout= dict(title= 'Barplot of Feature Importance',\n            width= 900,\n            height= 2000,\n            yaxis= dict(showgrid= False,\n                       showline= False,\n                       showticklabels= True))\n\nfig1= go.Figure(data= [trace2])\nfig1['layout'].update(layout)\npy.iplot(fig1, filename= 'plots')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T09:05:23.540245Z","iopub.execute_input":"2022-07-08T09:05:23.540678Z","iopub.status.idle":"2022-07-08T09:05:23.596261Z","shell.execute_reply.started":"2022-07-08T09:05:23.540644Z","shell.execute_reply":"2022-07-08T09:05:23.595162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interestingly we observe that in both Random Forest and Gradient Boosted learning models, the most important feature that both models picked out was the column : **ps_car_13**\n\nThis particular feature warrants further investigation so let us conduct a deep-dive into it.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\n\nWe have perfomed quite an extensive inspection of the Porto Seguro dataset by inspecting for null values and data quality, investigated linear correlations between features, inspected some of t he feature disributions as well as implemented a couple of learning models (Random Forest and Gradient Boosting classifier) so as to identify features that the models deemed important.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}