{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":1.667179,"end_time":"2022-07-06T13:38:37.035119","exception":false,"start_time":"2022-07-06T13:38:35.36794","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-07T06:12:16.759533Z","iopub.execute_input":"2022-07-07T06:12:16.760012Z","iopub.status.idle":"2022-07-07T06:12:16.770223Z","shell.execute_reply.started":"2022-07-07T06:12:16.759977Z","shell.execute_reply":"2022-07-07T06:12:16.768878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center>Fast.ai⚡+ Random Forests 🌳</center>","metadata":{"papermill":{"duration":0.008262,"end_time":"2022-07-06T13:38:37.052042","exception":false,"start_time":"2022-07-06T13:38:37.04378","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"In this notebook, we'll be using `Fast.ai` along with Scikit learn's Random Forests to predict the Titanic Passenger Survival chances.","metadata":{"papermill":{"duration":0.008384,"end_time":"2022-07-06T13:38:37.070655","exception":false,"start_time":"2022-07-06T13:38:37.062271","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"0\"></a>\n# **<center><span style=\"color:#00cc00;\">Libraries</span></center>**","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nfrom fastai.tabular.all import *","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"| S.No. | Content                   |\n|-------|---------------------------|\n| 1     | [Setting up Data](#1)     |\n| 2     | [Using Fast.ai ⚡](#2)    |\n| 3     | [Scoring Metrics](#3)     |\n| 4     | [Back to Sklearn](#4)     | \n| 5     | [Exploring the Model](#5) |","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# **<center><span style=\"color:#00cc00;\">1. Setting up Data</span></center>**","metadata":{}},{"cell_type":"markdown","source":"In order to use `Fast.ai`, first we need to convert our pandas `DataFrame` into fastai's `TabularPandas`. \n\nSo, first let's import the train and test sets as usual using pandas `read_csv`. After that, we'll concatenate both `train` and `test` in one DataFrame and we'll name it say `df`.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('../input/titanic/train.csv')\ntest = pd.read_csv('../input/titanic/test.csv')","metadata":{"papermill":{"duration":0.05285,"end_time":"2022-07-06T13:38:37.133384","exception":false,"start_time":"2022-07-06T13:38:37.080534","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-07T06:12:17.138345Z","iopub.execute_input":"2022-07-07T06:12:17.139000Z","iopub.status.idle":"2022-07-07T06:12:17.155135Z","shell.execute_reply.started":"2022-07-07T06:12:17.138955Z","shell.execute_reply":"2022-07-07T06:12:17.154211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.concat([train, test], ignore_index=True)\ndf.drop(['Name', 'Ticket', 'Cabin', 'PassengerId'], axis=1, inplace=True)\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:17.455876Z","iopub.execute_input":"2022-07-07T06:12:17.456302Z","iopub.status.idle":"2022-07-07T06:12:17.488283Z","shell.execute_reply.started":"2022-07-07T06:12:17.456269Z","shell.execute_reply":"2022-07-07T06:12:17.487417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"papermill":{"duration":0.041007,"end_time":"2022-07-06T13:38:37.203274","exception":false,"start_time":"2022-07-06T13:38:37.162267","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-07T06:12:17.664904Z","iopub.execute_input":"2022-07-07T06:12:17.665787Z","iopub.status.idle":"2022-07-07T06:12:17.672879Z","shell.execute_reply.started":"2022-07-07T06:12:17.665748Z","shell.execute_reply":"2022-07-07T06:12:17.671915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n# **<center><span style=\"color:#00cc00;\">2. Using Fast.ai ⚡</span></center>**\n","metadata":{}},{"cell_type":"markdown","source":"Now, in order to convert pandas `DataFrame` to fastai's `TabularPandas`, we'll use a function by the same name i.e. `TabularPandas` and pass in some arguments. Those are:\n\n- `df`: the pandas Dataframe object\n- `procs`: a transformation object\n- `cat_names`: categorical columns\n- `cont_names`: continuous/numerical columns\n- `y_names`: target variable\n- `y_block`: whether the target variable is for Categorization or Regression: `CategoryBlock` for categorical target; `RegressionBlock` for continuous target.\n\n## About TabularProcs\n\nTabularProc is used for variable transformation which may include categorization, filling of missing values, normalization. Two of it's features are:\n- It returns the exact same object that's passed to it, after modifying the object in place; and\n- It runs transform once, when the data is passed in, rather than lazily as the data is accessed.\n\n### About the list `procs` (below)\n\nWe'll use three TabularProcs:\n- `Categorify`: which will replace a column with a numerical categorical column\n- `FillMissing`: which will fill missing values (default is median)\n- `Normalize`: which will normalize the numerical features","metadata":{}},{"cell_type":"code","source":"procs = [Categorify, FillMissing, Normalize]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:17.885342Z","iopub.execute_input":"2022-07-07T06:12:17.885788Z","iopub.status.idle":"2022-07-07T06:12:17.891424Z","shell.execute_reply.started":"2022-07-07T06:12:17.885754Z","shell.execute_reply":"2022-07-07T06:12:17.890148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Getting categorical and continuous columns automatically\n### About `cont_cat_split`\n\n`cont_cat_split` automatically tries to find the categorical or continuous columns from the dataset which is given. This function works by determining if a column is continuous or categorical based on the cardinality of its values. If it is above the `max_card` parameter (default is `20`) (or a `float` datatype) then it will be added to the `cont_names` else `cat_names`.","metadata":{}},{"cell_type":"code","source":"target = 'Survived'\n\ncont, cat = cont_cat_split(df, 1, dep_var=target)\n\nto = TabularPandas(df, procs, cat, cont, y_names=target, y_block=CategoryBlock)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:18.080306Z","iopub.execute_input":"2022-07-07T06:12:18.080916Z","iopub.status.idle":"2022-07-07T06:12:18.124500Z","shell.execute_reply.started":"2022-07-07T06:12:18.080872Z","shell.execute_reply":"2022-07-07T06:12:18.123321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(to.train), len(to.valid)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:18.261999Z","iopub.execute_input":"2022-07-07T06:12:18.262614Z","iopub.status.idle":"2022-07-07T06:12:18.272814Z","shell.execute_reply.started":"2022-07-07T06:12:18.262566Z","shell.execute_reply":"2022-07-07T06:12:18.271631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to show the dataframe now, we can use `.show()` method.","metadata":{}},{"cell_type":"code","source":"to.show(5)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:18.466074Z","iopub.execute_input":"2022-07-07T06:12:18.466701Z","iopub.status.idle":"2022-07-07T06:12:18.491763Z","shell.execute_reply.started":"2022-07-07T06:12:18.466654Z","shell.execute_reply":"2022-07-07T06:12:18.490603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To check the numerical columns, we can use `.items` method.","metadata":{}},{"cell_type":"code","source":"to.items.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:18.671048Z","iopub.execute_input":"2022-07-07T06:12:18.671474Z","iopub.status.idle":"2022-07-07T06:12:18.686679Z","shell.execute_reply.started":"2022-07-07T06:12:18.671440Z","shell.execute_reply":"2022-07-07T06:12:18.685317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Also, if we want to know how the mapping for categorical variables was done, we can find out by using `.classes[Variable_Name]` method.","metadata":{}},{"cell_type":"code","source":"to.classes['Age_na']","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:18.955287Z","iopub.execute_input":"2022-07-07T06:12:18.956106Z","iopub.status.idle":"2022-07-07T06:12:18.962730Z","shell.execute_reply.started":"2022-07-07T06:12:18.956057Z","shell.execute_reply":"2022-07-07T06:12:18.961528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's get our training and testing sets back from `TabularPandas` back. In this scenario, our combined dataset is stored in `.train` method of the `TabularPandas`. Hence, we get it back by the following:","metadata":{}},{"cell_type":"code","source":"train = to.train[:train.shape[0]]\ntest = to.train[train.shape[0]:]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:19.081639Z","iopub.execute_input":"2022-07-07T06:12:19.082449Z","iopub.status.idle":"2022-07-07T06:12:19.090472Z","shell.execute_reply.started":"2022-07-07T06:12:19.082407Z","shell.execute_reply":"2022-07-07T06:12:19.089423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n# **<center><span style=\"color:#00cc00;\">3. Scoring Metrics</span></center>**\n","metadata":{}},{"cell_type":"code","source":"def rmse(pred, y):\n    return round(math.sqrt(((pred-y)**2).mean()), 6)\n\ndef mrmse(model, x_train, y_train):\n    return rmse(model.predict(x_train), y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:19.218332Z","iopub.execute_input":"2022-07-07T06:12:19.219202Z","iopub.status.idle":"2022-07-07T06:12:19.225091Z","shell.execute_reply.started":"2022-07-07T06:12:19.219165Z","shell.execute_reply":"2022-07-07T06:12:19.223845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n# **<center><span style=\"color:#00cc00;\">4. Back to Sklearn</span></center>**\n","metadata":{}},{"cell_type":"markdown","source":"Now, we are back to our `Sklearn`. We can proceed usually from here now.","metadata":{}},{"cell_type":"code","source":"X = train.drop('Survived', axis=1)\ny = train['Survived']\n\nx_train, x_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=101)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:19.353994Z","iopub.execute_input":"2022-07-07T06:12:19.354877Z","iopub.status.idle":"2022-07-07T06:12:19.364824Z","shell.execute_reply.started":"2022-07-07T06:12:19.354836Z","shell.execute_reply":"2022-07-07T06:12:19.363696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\ndef rf(x_train, y_train, n_estimators=400, max_features=0.5, min_samples_leaf=5, **kwargs):\n    return RandomForestClassifier(n_jobs=-1, n_estimators=n_estimators,\n                                  max_features=max_features, min_samples_leaf=min_samples_leaf, oob_score=True).fit(x_train, y_train)\n\nmodel = rf(x_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:19.477566Z","iopub.execute_input":"2022-07-07T06:12:19.478506Z","iopub.status.idle":"2022-07-07T06:12:20.799547Z","shell.execute_reply.started":"2022-07-07T06:12:19.478455Z","shell.execute_reply":"2022-07-07T06:12:20.798153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = model.predict(x_test)\npred","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:20.801568Z","iopub.execute_input":"2022-07-07T06:12:20.801908Z","iopub.status.idle":"2022-07-07T06:12:20.914821Z","shell.execute_reply.started":"2022-07-07T06:12:20.801879Z","shell.execute_reply":"2022-07-07T06:12:20.913667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"error = mrmse(model, x_train, y_train)\nacc = accuracy_score(y_test, pred)\npd.DataFrame({'RMSE': error, 'Accuracy Score': acc}, index=['Value'])","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:20.916329Z","iopub.execute_input":"2022-07-07T06:12:20.917068Z","iopub.status.idle":"2022-07-07T06:12:21.133180Z","shell.execute_reply.started":"2022-07-07T06:12:20.917025Z","shell.execute_reply":"2022-07-07T06:12:21.132262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n# **<center><span style=\"color:#00cc00;\">5. Submission</span></center>**\n","metadata":{}},{"cell_type":"code","source":"test.drop('Survived', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:21.135856Z","iopub.execute_input":"2022-07-07T06:12:21.136318Z","iopub.status.idle":"2022-07-07T06:12:21.143956Z","shell.execute_reply.started":"2022-07-07T06:12:21.136287Z","shell.execute_reply":"2022-07-07T06:12:21.142794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = model.predict(test)\npreds","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:21.145227Z","iopub.execute_input":"2022-07-07T06:12:21.145508Z","iopub.status.idle":"2022-07-07T06:12:21.363542Z","shell.execute_reply.started":"2022-07-07T06:12:21.145482Z","shell.execute_reply":"2022-07-07T06:12:21.362421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"idx = list(test.index)\nsubs = pd.read_csv('../input/titanic/gender_submission.csv')\nsubs['Survived'] = preds\nsubs.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:12:21.367150Z","iopub.execute_input":"2022-07-07T06:12:21.367554Z","iopub.status.idle":"2022-07-07T06:12:21.379070Z","shell.execute_reply.started":"2022-07-07T06:12:21.367523Z","shell.execute_reply":"2022-07-07T06:12:21.378063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n# **<center><span style=\"color:#00cc00;\">5. Exploring the Model</span></center>**\n","metadata":{}},{"cell_type":"markdown","source":"## Impact of `n_estimators`\n\nOne most important properties of random forests is that they aren't very sensitive to the hyperparameter choices such as `max_features`. We can set `n_estimators` to as high a number as we want: the more trees we have, the more the accurate the  model will be.\n\nBut still let's see the impact of that on our model, whether increasing it does good?","metadata":{}},{"cell_type":"code","source":"preds = np.stack([t.predict(x_test) for t in model.estimators_])\n\nplt.plot([rmse(preds[:i+1].mean(0), y_test) for i in range(100)])\nplt.xlabel('n_estimators')\nplt.ylabel('RMSE')","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:13:09.292952Z","iopub.execute_input":"2022-07-07T06:13:09.293394Z","iopub.status.idle":"2022-07-07T06:13:10.021097Z","shell.execute_reply.started":"2022-07-07T06:13:09.293360Z","shell.execute_reply":"2022-07-07T06:13:10.019732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Importance\n\nOur model can make predictions that's good; but we also want to know how it's making those predictions. There is an attribute `feature_importances_` of Random Forest that can help us with that.","metadata":{}},{"cell_type":"code","source":"def rf_feat_importance(m ,df):\n    return pd.DataFrame({'cols': df.columns, 'imp':model.feature_importances_}).sort_values('imp', ascending=False)\n\ndef plot_fi(fi):\n    return fi.plot('cols', 'imp', 'barh', figsize=(12, 7), legend=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:15:07.392499Z","iopub.execute_input":"2022-07-07T06:15:07.393211Z","iopub.status.idle":"2022-07-07T06:15:07.400245Z","shell.execute_reply.started":"2022-07-07T06:15:07.393175Z","shell.execute_reply":"2022-07-07T06:15:07.399088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following table shows the most important columns moving down to the least important columns (their contribution to the model prediction).","metadata":{}},{"cell_type":"code","source":"fi = rf_feat_importance(model, x_train)\nfi","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:15:56.835038Z","iopub.execute_input":"2022-07-07T06:15:56.835415Z","iopub.status.idle":"2022-07-07T06:15:56.951463Z","shell.execute_reply.started":"2022-07-07T06:15:56.835387Z","shell.execute_reply":"2022-07-07T06:15:56.950297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The plot helps us easily visualize and understand it.","metadata":{}},{"cell_type":"code","source":"plot_fi(fi)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:16:13.646984Z","iopub.execute_input":"2022-07-07T06:16:13.648263Z","iopub.status.idle":"2022-07-07T06:16:13.852949Z","shell.execute_reply.started":"2022-07-07T06:16:13.648223Z","shell.execute_reply":"2022-07-07T06:16:13.851803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I hope you liked this notebook! If so, please don't forget to upvote this notebook!\n\nAlso, if you have any suggestions, I'll be happy to hear those in the comments!","metadata":{"papermill":{"duration":0.015767,"end_time":"2022-07-06T13:39:10.127362","exception":false,"start_time":"2022-07-06T13:39:10.111595","status":"completed"},"tags":[]}}]}