{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Stay Alert! The Ford Challenge using Classification\n\nThe following topics are covered in this colab :\n\n- Downloading a real-world dataset\n- Preparing a dataset for training\n- Training and interpreting decision trees\n- Training and interpreting random forests\n- Overfitting, hyperparameter tuning & regularization\n- Making predictions on single inputs\n","metadata":{"id":"7Vxy3OgU9Tgx"}},{"cell_type":"markdown","source":"# Problem Statement\n\n\n\n> **QUESTION**:Driving while distracted, fatigued or drowsy may lead to accidents. Activities that divert the driver's attention from the road ahead, such as engaging in a conversation with other passengers in the car, making or receiving phone calls, sending or receiving text messages, eating while driving or events outside the car may cause driver distraction. Fatigue and drowsiness can result from driving long hours or from lack of sleep.\n\n>The objective of this challenge is to design a detector/classifier that will detect whether the driver is alert or not alert, employing any combination of vehicular, environmental and driver physiological data that are acquired while driving.","metadata":{"id":"sG6-L1a49ZcM"}},{"cell_type":"markdown","source":"##Step 1 - Download and Explore the Data\n\nThe dataset is available as a ZIP file at the following url:","metadata":{"id":"ojBFx6ti9lbH"}},{"cell_type":"code","source":"import os\ndata_dir='../input/stayalert'\nos.listdir(data_dir)","metadata":{"id":"U4jMfvdq9yr7","outputId":"8c1c65dc-b620-4e11-931e-f5dea1785e63","execution":{"iopub.status.busy":"2022-08-06T17:49:47.751438Z","iopub.execute_input":"2022-08-06T17:49:47.752488Z","iopub.status.idle":"2022-08-06T17:49:47.764838Z","shell.execute_reply.started":"2022-08-06T17:49:47.752448Z","shell.execute_reply":"2022-08-06T17:49:47.763563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\npd.set_option('display.max_columns', None)","metadata":{"id":"eyldmE1P91dJ","execution":{"iopub.status.busy":"2022-08-06T17:49:47.850590Z","iopub.execute_input":"2022-08-06T17:49:47.851694Z","iopub.status.idle":"2022-08-06T17:49:47.857553Z","shell.execute_reply.started":"2022-08-06T17:49:47.851646Z","shell.execute_reply":"2022-08-06T17:49:47.856669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path= data_dir +\"/fordTrain.csv\"","metadata":{"id":"cv7Z48c894tt","execution":{"iopub.status.busy":"2022-08-06T17:49:47.928783Z","iopub.execute_input":"2022-08-06T17:49:47.929420Z","iopub.status.idle":"2022-08-06T17:49:47.934468Z","shell.execute_reply.started":"2022-08-06T17:49:47.929382Z","shell.execute_reply":"2022-08-06T17:49:47.933216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **QUESTION 1**: Load the data from the file `train.csv` into a Pandas data frame.","metadata":{"id":"tqxkSq_-xaYF"}},{"cell_type":"code","source":"ford=pd.read_csv(path)","metadata":{"id":"9OP3PQR9xiyC","execution":{"iopub.status.busy":"2022-08-06T17:49:47.995325Z","iopub.execute_input":"2022-08-06T17:49:47.996398Z","iopub.status.idle":"2022-08-06T17:49:49.788750Z","shell.execute_reply.started":"2022-08-06T17:49:47.996355Z","shell.execute_reply":"2022-08-06T17:49:49.787441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ford","metadata":{"id":"mHG2aEk1-Y1l","outputId":"ddb9f353-f64c-4afc-e592-762c18fd3434","execution":{"iopub.status.busy":"2022-08-06T17:49:49.790646Z","iopub.execute_input":"2022-08-06T17:49:49.791127Z","iopub.status.idle":"2022-08-06T17:49:49.834263Z","shell.execute_reply.started":"2022-08-06T17:49:49.791091Z","shell.execute_reply":"2022-08-06T17:49:49.833235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataset contains 604329 rows and 33 columns. Each row of the dataset contains information about one customer. \n\nThe objective of this challenge is to design a detector/classifier that will detect whether the driver is alert or not alert, employing any combination of vehicular, environmental and driver physiological data that are acquired while driving.\n\nLet's check the data type for each column.","metadata":{"id":"yJ5tgOSxxtTo"}},{"cell_type":"code","source":"ford.info()","metadata":{"id":"K9_jXt1F-f91","outputId":"0caa04b6-ca3a-4dc6-85b8-d971ee6f90c5","execution":{"iopub.status.busy":"2022-08-06T17:49:49.835783Z","iopub.execute_input":"2022-08-06T17:49:49.836137Z","iopub.status.idle":"2022-08-06T17:49:49.899228Z","shell.execute_reply.started":"2022-08-06T17:49:49.836104Z","shell.execute_reply":"2022-08-06T17:49:49.898075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **QUESTION 2**: How many rows and columns does the dataset contain? ","metadata":{"id":"C0DMkjx5x0rV"}},{"cell_type":"code","source":"n_rows = ford.shape[0]","metadata":{"id":"pOnPw7cdx4-2","execution":{"iopub.status.busy":"2022-08-06T17:49:49.902185Z","iopub.execute_input":"2022-08-06T17:49:49.902810Z","iopub.status.idle":"2022-08-06T17:49:49.907879Z","shell.execute_reply.started":"2022-08-06T17:49:49.902762Z","shell.execute_reply":"2022-08-06T17:49:49.906652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_cols = ford.shape[1]","metadata":{"id":"rRz3sqzSx4yE","execution":{"iopub.status.busy":"2022-08-06T17:49:49.910054Z","iopub.execute_input":"2022-08-06T17:49:49.910915Z","iopub.status.idle":"2022-08-06T17:49:49.917278Z","shell.execute_reply.started":"2022-08-06T17:49:49.910870Z","shell.execute_reply":"2022-08-06T17:49:49.916406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('The dataset contains {} rows and {} columns.'.format(n_rows, n_cols))","metadata":{"id":"su7es1Vpx4mU","outputId":"ee1b3667-5e94-4851-a5d4-57b2e6176dee","execution":{"iopub.status.busy":"2022-08-06T17:49:49.918333Z","iopub.execute_input":"2022-08-06T17:49:49.919020Z","iopub.status.idle":"2022-08-06T17:49:49.927348Z","shell.execute_reply.started":"2022-08-06T17:49:49.918988Z","shell.execute_reply":"2022-08-06T17:49:49.926438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Exploratory Analysis and Visualization\n\nLet's explore the data by visualizing the distribution of values in some columns of the dataset, and the relationships between \"charges\" and other columns.\n","metadata":{"id":"Agbs6-ljyGgr"}},{"cell_type":"markdown","source":"* libraries that we are going to use in this collab ","metadata":{"id":"cNIR7Kk0zHjn"}},{"cell_type":"code","source":"import seaborn as sns\nimport plotly.express as px\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n#The following settings will improve the default style and font sizes for our charts\nsns.set_style('darkgrid')\nmatplotlib.rcParams['font.size'] = 14\nmatplotlib.rcParams['figure.figsize'] = (10, 6)\nmatplotlib.rcParams['figure.facecolor'] = '#00000000'","metadata":{"id":"N2J8If1ozGR4","execution":{"iopub.status.busy":"2022-08-06T17:49:49.928552Z","iopub.execute_input":"2022-08-06T17:49:49.929290Z","iopub.status.idle":"2022-08-06T17:49:51.638568Z","shell.execute_reply.started":"2022-08-06T17:49:49.929257Z","shell.execute_reply":"2022-08-06T17:49:51.637446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **QUESTION 3**: How many `missing values` does the dataset contain in percentage? ","metadata":{"id":"n0UAQgDZyQRa"}},{"cell_type":"code","source":"# Here we will check the percentage of nan values present in each feature\n# 1 -step make the list of features which has missing values\nfeature_with_na=[feature for feature in ford.columns if ford[feature].isnull().sum()>1]\n# 2- step print the feature name and the percentage of missing values\nfor feature in feature_with_na:\n  print(feature, np.round(ford[feature].isnull().mean(), 4)*100,  \" % missing values\")","metadata":{"id":"14IzWrXlFLKt","execution":{"iopub.status.busy":"2022-08-06T17:49:51.641017Z","iopub.execute_input":"2022-08-06T17:49:51.642227Z","iopub.status.idle":"2022-08-06T17:49:51.685670Z","shell.execute_reply.started":"2022-08-06T17:49:51.642176Z","shell.execute_reply":"2022-08-06T17:49:51.684599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"there is `no null` value in the dataset that reduce our preprocessing time","metadata":{"id":"yKpu2xcMFwXO"}},{"cell_type":"code","source":"# 0 represent the driver is not aler\n# 1 represent the driver is alert\nnumerical = [feature for feature in ford.columns if ford[feature].dtype in ['int64', 'float64']]\ndf = ford[numerical]\n\nfig = plt.figure(figsize = (25, 35))\n\ni=1\nfor n in df.columns:\n    plt.subplot(7, 5, i)\n    ax = sns.histplot(x = ford[n],hue = ford['IsAlert'], palette = ['#676FA3', '#FF5959'], bins = 40)\n    ax.set(xlabel = None, ylabel = None)\n    plt.title(str(n), loc = 'center')\n    plt.xticks(rotation = 20, fontsize = 10)\n    i += 1","metadata":{"id":"8mpbFdfaGw-G","outputId":"921be309-9c53-4431-c827-c654cd8e9c7a","execution":{"iopub.status.busy":"2022-08-06T17:49:51.686963Z","iopub.execute_input":"2022-08-06T17:49:51.687691Z","iopub.status.idle":"2022-08-06T17:50:10.131125Z","shell.execute_reply.started":"2022-08-06T17:49:51.687656Z","shell.execute_reply":"2022-08-06T17:50:10.129858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" the above graph represents that the data is `skewed`.","metadata":{"id":"0VYFrqDmHi0c"}},{"cell_type":"markdown","source":"## Step 2 - Prepare the Dataset for Training\n\n\nBefore we can train the model, we need to prepare the dataset. Here are the steps we'll follow:\n\n1. Identify the input and target column(s) for training the model.\n2. Identify numeric and categorical input columns.\n3. [Impute](https://scikit-learn.org/stable/modules/impute.html) (fill) missing values in numeric columns\n4. [Scale](https://scikit-learn.org/stable/modules/preprocessing.html#scaling-features-to-a-range) values in numeric columns to a $(0,1)$ range.\n5. [Encode](https://scikit-learn.org/stable/modules/preprocessing.html#encoding-categorical-features) categorical data into one-hot vectors.\n6. Split the dataset into training and validation sets.\n","metadata":{"id":"0WwkH6uC9_EO"}},{"cell_type":"markdown","source":"### Identify Inputs and Targets\n\nWhile the dataset contains `33` columns, not all of them are useful for modeling. Note the following:\n\n- The first column and second column- `Trial ID`and `ObsNum` is a unique ID and observation number and isn't useful for training the model.\n- The third column `IsAlert` contains the value we need to predict i.e. it's the target column.\n- Data from all the other columns (except the first,second and the third column) can be used as inputs to the model.\n ","metadata":{"id":"s2qnXMj_h2z0"}},{"cell_type":"markdown","source":"> **QUESTION 4**: Create a list `input_cols` of column names containing data that can be used as input to train the model, and identify the target column as the variable `target_col`.","metadata":{"id":"zvWhINVt0egw"}},{"cell_type":"code","source":"# Identify the input columns (a list of column names)\ninput_cols = list(ford.columns)[3:]","metadata":{"id":"7T7AeTcxhvzW","execution":{"iopub.status.busy":"2022-08-06T17:50:10.135185Z","iopub.execute_input":"2022-08-06T17:50:10.135868Z","iopub.status.idle":"2022-08-06T17:50:10.140258Z","shell.execute_reply.started":"2022-08-06T17:50:10.135829Z","shell.execute_reply":"2022-08-06T17:50:10.139123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(input_cols)","metadata":{"id":"fc-DatjX04qx","outputId":"f9b63fe0-c82f-4c7e-dc43-442e2997f354","execution":{"iopub.status.busy":"2022-08-06T17:50:10.141830Z","iopub.execute_input":"2022-08-06T17:50:10.142218Z","iopub.status.idle":"2022-08-06T17:50:10.154687Z","shell.execute_reply.started":"2022-08-06T17:50:10.142182Z","shell.execute_reply":"2022-08-06T17:50:10.153581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Identify the name of the target column (a single string, not a list)\ntarget_col =list(ford.columns)[2]","metadata":{"id":"2w8Fp_HB0mMu","execution":{"iopub.status.busy":"2022-08-06T17:50:10.156003Z","iopub.execute_input":"2022-08-06T17:50:10.156605Z","iopub.status.idle":"2022-08-06T17:50:10.168301Z","shell.execute_reply.started":"2022-08-06T17:50:10.156566Z","shell.execute_reply":"2022-08-06T17:50:10.167135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(target_col)","metadata":{"id":"7HHkRWNm03dP","outputId":"18bf2214-55f8-4bdf-fc81-f5544208bcb6","execution":{"iopub.status.busy":"2022-08-06T17:50:10.170149Z","iopub.execute_input":"2022-08-06T17:50:10.170542Z","iopub.status.idle":"2022-08-06T17:50:10.180726Z","shell.execute_reply.started":"2022-08-06T17:50:10.170504Z","shell.execute_reply":"2022-08-06T17:50:10.179543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Make sure that the `Id` and `SalePrice` columns are not included in `input_cols`.\n\nNow that we've identified the input and target columns, we can separate input & target data.","metadata":{"id":"JsdUeloo0wJx"}},{"cell_type":"code","source":"inputs_df = ford[input_cols]\ntargets = ford[target_col]","metadata":{"id":"TQJmCsQo095e","execution":{"iopub.status.busy":"2022-08-06T17:50:10.182611Z","iopub.execute_input":"2022-08-06T17:50:10.182981Z","iopub.status.idle":"2022-08-06T17:50:10.264867Z","shell.execute_reply.started":"2022-08-06T17:50:10.182947Z","shell.execute_reply":"2022-08-06T17:50:10.263464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inputs_df","metadata":{"id":"WLlMcG6aI1JT","outputId":"ee7d9350-3910-42c3-9fd0-956563664428","execution":{"iopub.status.busy":"2022-08-06T17:50:10.266809Z","iopub.execute_input":"2022-08-06T17:50:10.267253Z","iopub.status.idle":"2022-08-06T17:50:10.311934Z","shell.execute_reply.started":"2022-08-06T17:50:10.267215Z","shell.execute_reply":"2022-08-06T17:50:10.310745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets","metadata":{"id":"ymaJraGmI3WG","outputId":"e54e21cf-0c14-4d20-cd96-36e3b78ddac3","execution":{"iopub.status.busy":"2022-08-06T17:50:10.313509Z","iopub.execute_input":"2022-08-06T17:50:10.313845Z","iopub.status.idle":"2022-08-06T17:50:10.322789Z","shell.execute_reply.started":"2022-08-06T17:50:10.313812Z","shell.execute_reply":"2022-08-06T17:50:10.321911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###Identify Numeric and Categorical Data\nThe next step in data preparation is to identify numeric and categorical columns. We can do this by looking at the data type of each column.","metadata":{"id":"EwazBRPDiNap"}},{"cell_type":"markdown","source":"> **QUESTION 5**: Crate two lists `numeric_cols` and `categorical_cols` containing names of numeric and categorical input columns within the dataframe respectively. Numeric columns have data types `int64` and `float64`, whereas categorical columns have the data type `object`.\n>\n> *Hint*: See this [StackOverflow question](https://stackoverflow.com/questions/25039626/how-do-i-find-numeric-columns-in-pandas). ","metadata":{"id":"d9ocavp91cdT"}},{"cell_type":"code","source":"#numerical=medical.select_dtypes(include=np.number).columns.tolist()\nnumeric_cols = inputs_df.select_dtypes(include=['int64', 'float64']).columns.tolist()\ncategorical_cols = inputs_df.select_dtypes(include=[object]).columns.tolist()","metadata":{"id":"Vr6XCGsPiNAJ","execution":{"iopub.status.busy":"2022-08-06T17:50:10.324915Z","iopub.execute_input":"2022-08-06T17:50:10.325298Z","iopub.status.idle":"2022-08-06T17:50:10.412060Z","shell.execute_reply.started":"2022-08-06T17:50:10.325265Z","shell.execute_reply":"2022-08-06T17:50:10.410280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numeric_cols","metadata":{"id":"_mFGA7xfI8bi","outputId":"a43168e5-c553-491d-b0f6-e4388037a7fe","execution":{"iopub.status.busy":"2022-08-06T17:50:10.414563Z","iopub.execute_input":"2022-08-06T17:50:10.415233Z","iopub.status.idle":"2022-08-06T17:50:10.423459Z","shell.execute_reply.started":"2022-08-06T17:50:10.415141Z","shell.execute_reply":"2022-08-06T17:50:10.422465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_cols","metadata":{"id":"-bMky-j_I79r","outputId":"8152bc3b-47b0-4763-b2b3-0f81dd206f80","execution":{"iopub.status.busy":"2022-08-06T17:50:10.424897Z","iopub.execute_input":"2022-08-06T17:50:10.425372Z","iopub.status.idle":"2022-08-06T17:50:10.438559Z","shell.execute_reply.started":"2022-08-06T17:50:10.425336Z","shell.execute_reply":"2022-08-06T17:50:10.437406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###Scale Numerical Values\nThe numeric columns in our dataset have varying ranges.","metadata":{"id":"K0xV7nPui-VH"}},{"cell_type":"code","source":"inputs_df[numeric_cols].describe().loc[['min', 'max']]","metadata":{"id":"KkHP1jBNjCoM","outputId":"6656f875-ea2f-4782-f258-83a12cd64f29","execution":{"iopub.status.busy":"2022-08-06T17:50:10.440141Z","iopub.execute_input":"2022-08-06T17:50:10.440851Z","iopub.status.idle":"2022-08-06T17:50:11.130641Z","shell.execute_reply.started":"2022-08-06T17:50:10.440801Z","shell.execute_reply":"2022-08-06T17:50:11.129527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A good practice is to [scale numeric features](https://scikit-learn.org/stable/modules/preprocessing.html#scaling-features-to-a-range) to a small range of values e.g. $(0,1)$. Scaling numeric features ensures that no particular feature has a disproportionate impact on the model's loss. Optimization algorithms also work better in practice with smaller numbers.\n","metadata":{"id":"RfHBHiPI2JU2"}},{"cell_type":"markdown","source":"> **QUESTION 6**: Scale numeric values to the $(0, 1)$ range using `MinMaxScaler` from `sklearn.preprocessing`.","metadata":{"id":"SEZvAGJt2NRN"}},{"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler\n\n# Create the scaler\nscaler = MinMaxScaler()\n\n# Fit the scaler to the numeric columns\nscaler.fit(inputs_df[numeric_cols])\n\n# Transform and replace the numeric columns\ninputs_df[numeric_cols] = scaler.transform(inputs_df[numeric_cols])","metadata":{"id":"xRcn3CHs2F8d","outputId":"be042c31-7515-4e8e-e88a-080fd7748c96","execution":{"iopub.status.busy":"2022-08-06T17:50:11.132134Z","iopub.execute_input":"2022-08-06T17:50:11.132498Z","iopub.status.idle":"2022-08-06T17:50:11.768755Z","shell.execute_reply.started":"2022-08-06T17:50:11.132464Z","shell.execute_reply":"2022-08-06T17:50:11.767356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After scaling, the ranges of all numeric columns should be (0, 1).","metadata":{"id":"o-l3QJBm2ccj"}},{"cell_type":"code","source":"inputs_df[numeric_cols].describe().loc[['min', 'max']]","metadata":{"id":"UioP8mNB2gq1","outputId":"abb2e0eb-2075-4de9-8dfc-fa84570f5d67","execution":{"iopub.status.busy":"2022-08-06T17:50:11.770427Z","iopub.execute_input":"2022-08-06T17:50:11.770834Z","iopub.status.idle":"2022-08-06T17:50:12.642218Z","shell.execute_reply.started":"2022-08-06T17:50:11.770801Z","shell.execute_reply":"2022-08-06T17:50:12.641312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###Training and Validation Set\nFinally, let's split the dataset into a training and validation set. We'll use a randomly select 25% subset of the data for validation. Also, we'll use just the numeric and encoded columns, since the inputs to our model must be numbers.","metadata":{"id":"uVS7HjBzjKAp"}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\ntrain_inputs, val_inputs, train_targets, val_targets = train_test_split(inputs_df[numeric_cols], \n                                                                        targets, \n                                                                        test_size=0.25, \n                                                                        random_state=42)\n","metadata":{"id":"Fz1MG0IFjPWs","execution":{"iopub.status.busy":"2022-08-06T17:50:12.643538Z","iopub.execute_input":"2022-08-06T17:50:12.644414Z","iopub.status.idle":"2022-08-06T17:50:12.993583Z","shell.execute_reply.started":"2022-08-06T17:50:12.644378Z","shell.execute_reply":"2022-08-06T17:50:12.992618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_inputs","metadata":{"id":"GHJr-XzE3P47","outputId":"0d8c7e77-7039-495f-8b0d-026f60ab1736","execution":{"iopub.status.busy":"2022-08-06T17:50:12.995195Z","iopub.execute_input":"2022-08-06T17:50:12.995827Z","iopub.status.idle":"2022-08-06T17:50:13.039317Z","shell.execute_reply.started":"2022-08-06T17:50:12.995789Z","shell.execute_reply":"2022-08-06T17:50:13.038104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_targets","metadata":{"id":"Beb7sLS83SiV","outputId":"18d04c32-d76d-45af-f1bf-569ec9f4d75d","execution":{"iopub.status.busy":"2022-08-06T17:50:13.040674Z","iopub.execute_input":"2022-08-06T17:50:13.041049Z","iopub.status.idle":"2022-08-06T17:50:13.051622Z","shell.execute_reply.started":"2022-08-06T17:50:13.040994Z","shell.execute_reply":"2022-08-06T17:50:13.050616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_inputs","metadata":{"id":"UAVwiPms3W-l","outputId":"a8a13912-4979-467d-a061-67b2cdd6880c","execution":{"iopub.status.busy":"2022-08-06T17:50:13.052923Z","iopub.execute_input":"2022-08-06T17:50:13.054249Z","iopub.status.idle":"2022-08-06T17:50:13.102110Z","shell.execute_reply.started":"2022-08-06T17:50:13.054209Z","shell.execute_reply":"2022-08-06T17:50:13.101281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_targets","metadata":{"id":"TTPqItJq3Yyw","outputId":"71fbc4ca-5eb6-4528-bd3a-7045e17b162d","execution":{"iopub.status.busy":"2022-08-06T17:50:13.103289Z","iopub.execute_input":"2022-08-06T17:50:13.104293Z","iopub.status.idle":"2022-08-06T17:50:13.112885Z","shell.execute_reply.started":"2022-08-06T17:50:13.104253Z","shell.execute_reply":"2022-08-06T17:50:13.111378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model 1 - Training a Logistic Regression Model\n\n Training a Logistic Regression Model\n\nLogistic regression is a commonly used technique for solving binary classification problems. In a logistic regression model: \n\n- we take linear combination (or weighted sum of the input features) \n- we apply the sigmoid function to the result to obtain a number between 0 and 1\n- this number represents the probability of the input being classified as \"Yes\"\n- instead of RMSE, the cross entropy loss function is used to evaluate the results\n\n\nHere's a visual summary of how a logistic regression model is structured ([source](http://datahacker.rs/005-pytorch-logistic-regression-in-pytorch/)):\n\n\n<img src=\"https://i.imgur.com/YMaMo5D.png\" width=\"480\">\n\nThe sigmoid function applied to the linear combination of inputs has the following formula:\n\n<img src=\"https://i.imgur.com/sAVwvZP.png\" width=\"400\">\n\nTo train a logistic regression model, we can use the `LogisticRegression` class from Scikit-learn.","metadata":{"id":"IY7dqLjyk4ZQ"}},{"cell_type":"markdown","source":"> **QUESTION 7**: Create and train a Logistic regression  from `sklearn.linear_model`.","metadata":{"id":"KW5zJVpP3myr"}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\n\n# Create the model\nmodel = LogisticRegression(solver='liblinear')\n\n# Fit the model using inputs and targets\nmodel.fit(train_inputs,train_targets)","metadata":{"id":"cO4lE9kDk6HZ","outputId":"c4ce6290-0d34-448b-bd5b-72bb2896a662","execution":{"iopub.status.busy":"2022-08-06T17:50:13.121734Z","iopub.execute_input":"2022-08-06T17:50:13.122318Z","iopub.status.idle":"2022-08-06T17:50:29.265800Z","shell.execute_reply.started":"2022-08-06T17:50:13.122271Z","shell.execute_reply":"2022-08-06T17:50:29.264613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`model.fit` uses the following strategy for training the model (source):\n\n1. We initialize a model with random parameters (weights & biases).\n2. We pass some inputs into the model to obtain predictions.\n3. We compare the model's predictions with the actual targets using the loss function.\n4. We use an optimization technique (like least squares, gradient descent etc.) to reduce the loss by adjusting the weights & biases of the model\n5. We repeat steps 1 to 4 till the predictions from the model are good enough.\n\n<img src=\"https://www.deepnetts.com/blog/wp-content/uploads/2019/02/SupervisedLearning.png\" width=\"480\">","metadata":{"id":"qJx5kKiOlCTr"}},{"cell_type":"markdown","source":"## Make Predictions and Evaluate Your Model\n\nThe model is now trained, and we can use it to generate predictions for the training and validation inputs. We can evaluate the model's performance using the RMSE (root mean squared error) loss function.","metadata":{"id":"8sPyXjNLlIKz"}},{"cell_type":"markdown","source":"> **QUESTION 8**: Generate predictions and compute the RMSE loss for the training and validation sets. \n> \n> *Hint*: Use the `mean_squared_error` with the argument `squared=False` to compute RMSE loss.","metadata":{"id":"S_KzvC0O37pZ"}},{"cell_type":"code","source":"train_preds =model.predict(train_inputs)","metadata":{"id":"uxi9ItKPlLF3","execution":{"iopub.status.busy":"2022-08-06T17:50:29.268294Z","iopub.execute_input":"2022-08-06T17:50:29.268658Z","iopub.status.idle":"2022-08-06T17:50:29.294343Z","shell.execute_reply.started":"2022-08-06T17:50:29.268627Z","shell.execute_reply":"2022-08-06T17:50:29.293080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_preds","metadata":{"id":"LaAqj2id4GGy","outputId":"10269271-94e1-4406-d3d9-534bcd90bbce","execution":{"iopub.status.busy":"2022-08-06T17:50:29.295812Z","iopub.execute_input":"2022-08-06T17:50:29.296977Z","iopub.status.idle":"2022-08-06T17:50:29.305113Z","shell.execute_reply.started":"2022-08-06T17:50:29.296921Z","shell.execute_reply":"2022-08-06T17:50:29.303921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error,r2_score\n# Root mean square error\ntrain_rmse=mean_squared_error(train_targets,train_preds,squared=False)\nprint('The RMSE loss for the training set is {}.'.format(train_rmse))\n# r2_score\ntrain_r2=r2_score(train_targets,train_preds)*100\nprint('The r2_score for the training set is {} %'.format(train_r2))","metadata":{"id":"c4es-Pmk4F2u","outputId":"d9f58195-2f33-4669-d923-e2dd2df04567","execution":{"iopub.status.busy":"2022-08-06T17:50:29.306807Z","iopub.execute_input":"2022-08-06T17:50:29.307619Z","iopub.status.idle":"2022-08-06T17:50:29.330187Z","shell.execute_reply.started":"2022-08-06T17:50:29.307573Z","shell.execute_reply":"2022-08-06T17:50:29.328974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_preds = model.predict(val_inputs)","metadata":{"id":"MUqCQZvh4QYV","execution":{"iopub.status.busy":"2022-08-06T17:50:29.331630Z","iopub.execute_input":"2022-08-06T17:50:29.332687Z","iopub.status.idle":"2022-08-06T17:50:29.349548Z","shell.execute_reply.started":"2022-08-06T17:50:29.332640Z","shell.execute_reply":"2022-08-06T17:50:29.348271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_preds","metadata":{"id":"9WEMo0sD4RC4","outputId":"d98ae22b-42ca-41e9-d91d-d8bd49e8db5a","execution":{"iopub.status.busy":"2022-08-06T17:50:29.351457Z","iopub.execute_input":"2022-08-06T17:50:29.352278Z","iopub.status.idle":"2022-08-06T17:50:29.360077Z","shell.execute_reply.started":"2022-08-06T17:50:29.352230Z","shell.execute_reply":"2022-08-06T17:50:29.358883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_rmse =mean_squared_error(val_targets,val_preds,squared=False)\nprint('The RMSE loss for the validation set is {}.'.format(val_rmse))\ntrain_r2=r2_score(val_targets,val_preds)\nprint('The r2_score for the validation set is {} %.'.format(train_r2))","metadata":{"id":"g9GodvSu4V9u","outputId":"e1ada089-83ad-4187-d564-94ca9261a449","execution":{"iopub.status.busy":"2022-08-06T17:50:29.361968Z","iopub.execute_input":"2022-08-06T17:50:29.362832Z","iopub.status.idle":"2022-08-06T17:50:29.374849Z","shell.execute_reply.started":"2022-08-06T17:50:29.362786Z","shell.execute_reply":"2022-08-06T17:50:29.373684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model 2 -Training and Visualizing Decision Trees\n\nA decision tree in general parlance represents a hierarchical series of binary decisions:\n\n<img src=\"https://i.imgur.com/qSH4lqz.png\" width=\"480\">\n\nA decision tree in machine learning works in exactly the same way, and except that we let the computer figure out the optimal structure & hierarchy of decisions, instead of coming up with criteria manually.","metadata":{"id":"p0jz6w6hfo_9"}},{"cell_type":"markdown","source":"## Training\n\nWe can use `DecisionTreeClassifier` from `sklearn.tree` to train a decision tree.","metadata":{"id":"vqGbbfjPf9Mw"}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\n\n# Create the model\nmodel = DecisionTreeClassifier(random_state=42)\n\n# Fit the model\nmodel.fit(train_inputs, train_targets)","metadata":{"id":"kO40lLfifzUD","outputId":"970165c3-6f9f-44b7-8f6d-e709c651be90","execution":{"iopub.status.busy":"2022-08-06T17:50:29.376973Z","iopub.execute_input":"2022-08-06T17:50:29.377770Z","iopub.status.idle":"2022-08-06T17:50:47.118712Z","shell.execute_reply.started":"2022-08-06T17:50:29.377721Z","shell.execute_reply":"2022-08-06T17:50:47.117131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"An optimal decision tree has now been created using the training data.","metadata":{"id":"SnYtlWkxhsry"}},{"cell_type":"markdown","source":"##Evaluation\n\nLet's evaluate the decision tree using the accuracy score.","metadata":{"id":"qvnCslvqgXhO"}},{"cell_type":"code","source":"train_preds = model.predict(train_inputs)","metadata":{"id":"RB76kWs0f7CW","execution":{"iopub.status.busy":"2022-08-06T17:50:47.120545Z","iopub.execute_input":"2022-08-06T17:50:47.121349Z","iopub.status.idle":"2022-08-06T17:50:47.290414Z","shell.execute_reply.started":"2022-08-06T17:50:47.121301Z","shell.execute_reply":"2022-08-06T17:50:47.289136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_preds","metadata":{"id":"7h4uS5dHg7aK","outputId":"d7aa7720-a9cd-4576-e037-5efe60795ad0","execution":{"iopub.status.busy":"2022-08-06T17:50:47.292101Z","iopub.execute_input":"2022-08-06T17:50:47.292853Z","iopub.status.idle":"2022-08-06T17:50:47.301572Z","shell.execute_reply.started":"2022-08-06T17:50:47.292804Z","shell.execute_reply":"2022-08-06T17:50:47.300301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The decision tree also returns probabilities for each prediction.","metadata":{"id":"xhtOEM7ghQxG"}},{"cell_type":"code","source":"train_probs = model.predict_proba(train_inputs)\ntrain_probs","metadata":{"id":"JwGs87JJhUGW","outputId":"958b6b78-498a-4251-f399-b3f7d1b6d10d","execution":{"iopub.status.busy":"2022-08-06T17:50:47.302803Z","iopub.execute_input":"2022-08-06T17:50:47.303196Z","iopub.status.idle":"2022-08-06T17:50:47.473042Z","shell.execute_reply.started":"2022-08-06T17:50:47.303152Z","shell.execute_reply":"2022-08-06T17:50:47.471945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Seems like the decision tree is quite confident about its predictions.\n\nLet's check the accuracy of its predictions.","metadata":{"id":"yk_d3Bfth7_V"}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, confusion_matrix\n\naccuracy_score(train_targets, train_preds)","metadata":{"id":"J1QY_Y3vg9me","outputId":"efd6e09d-2154-4ff0-c259-fcbb76ad883e","execution":{"iopub.status.busy":"2022-08-06T17:50:47.474285Z","iopub.execute_input":"2022-08-06T17:50:47.475365Z","iopub.status.idle":"2022-08-06T17:50:47.533457Z","shell.execute_reply.started":"2022-08-06T17:50:47.475331Z","shell.execute_reply":"2022-08-06T17:50:47.532269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The training set accuracy is close to 100%! But we can't rely solely on the training set accuracy, we must evaluate the model on the validation set too. \n\nWe can make predictions and compute accuracy in one step using `model.score`","metadata":{"id":"Y9umuDGMhPz_"}},{"cell_type":"code","source":"model.score(val_inputs, val_targets)","metadata":{"id":"DlVX6VG0iBRR","outputId":"8ad52a4a-f8f2-4999-c090-3e2bcc0af388","execution":{"iopub.status.busy":"2022-08-06T17:50:47.534636Z","iopub.execute_input":"2022-08-06T17:50:47.534977Z","iopub.status.idle":"2022-08-06T17:50:47.606194Z","shell.execute_reply.started":"2022-08-06T17:50:47.534943Z","shell.execute_reply":"2022-08-06T17:50:47.604949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although the training accuracy is 100%, the accuracy on the validation set is  about 98%, which is the best result.\n ","metadata":{"id":"GgxSvGVYiMr2"}},{"cell_type":"markdown","source":"## Visualization\n\nWe can visualize the decision tree _learned_ from the training data.","metadata":{"id":"4vkp9yYViZrR"}},{"cell_type":"code","source":"from sklearn.tree import plot_tree, export_text\nplt.figure(figsize=(80,20))\nplot_tree(model, feature_names=train_inputs.columns, max_depth=2, filled=True);","metadata":{"id":"u_ZvV61viXBA","outputId":"87d0749e-fbea-43a0-c369-e023a9854285","execution":{"iopub.status.busy":"2022-08-06T17:50:47.607826Z","iopub.execute_input":"2022-08-06T17:50:47.608768Z","iopub.status.idle":"2022-08-06T17:50:48.898987Z","shell.execute_reply.started":"2022-08-06T17:50:47.608728Z","shell.execute_reply":"2022-08-06T17:50:48.897834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Can you see how the model classifies a given input as a series of decisions? The tree is truncated here, but following any path from the root node down to a leaf will result in \"Yes\" or \"No\". Do you see how a decision tree differs from a logistic regression model?\n\n\n**How a Decision Tree is Created**\n\nNote the `gini` value in each box. This is the loss function used by the decision tree to decide which column should be used for splitting the data, and at what point the column should be split. A lower Gini index indicates a better split. A perfect split (only one class on each side) has a Gini index of 0. \n\nFor a mathematical discussion of the Gini Index, watch this video: https://www.youtube.com/watch?v=-W0DnxQK1Eo . It has the following formula:\n\n<img src=\"https://i.imgur.com/CSC0gAo.png\" width=\"240\">\n\nConceptually speaking, while training the models evaluates all possible splits across all possible columns and picks the best one. Then, it recursively performs an optimal split for the two portions. In practice, however, it's very inefficient to check all possible splits, so the model uses a heuristic (predefined strategy) combined with some randomization.\n\nThe iterative approach of the machine learning workflow in the case of a decision tree involves growing the tree layer-by-layer:\n\n<img src=\"https://www.deepnetts.com/blog/wp-content/uploads/2019/02/SupervisedLearning.png\" width=\"480\">\n\n\nLet's check the depth of the tree that was created.","metadata":{"id":"1Vznv8keikyt"}},{"cell_type":"code","source":"model.tree_.max_depth","metadata":{"id":"H9v0BIjpilfj","outputId":"9cf2e2dc-4294-4ddd-cc6a-25e1b5c235d7","execution":{"iopub.status.busy":"2022-08-06T17:50:48.900762Z","iopub.execute_input":"2022-08-06T17:50:48.901114Z","iopub.status.idle":"2022-08-06T17:50:48.906994Z","shell.execute_reply.started":"2022-08-06T17:50:48.901081Z","shell.execute_reply":"2022-08-06T17:50:48.905793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can also display the tree as text, which can be easier to follow for deeper trees.","metadata":{"id":"csuRXS-vjFhw"}},{"cell_type":"code","source":"tree_text = export_text(model, max_depth=10, feature_names=list(train_inputs.columns))\nprint(tree_text[:5000])","metadata":{"id":"5sJrkR6GjIP_","outputId":"da40bc62-c3db-4a3e-aa62-992e4c4446ab","execution":{"iopub.status.busy":"2022-08-06T17:50:48.908565Z","iopub.execute_input":"2022-08-06T17:50:48.909530Z","iopub.status.idle":"2022-08-06T17:50:48.965060Z","shell.execute_reply.started":"2022-08-06T17:50:48.909493Z","shell.execute_reply":"2022-08-06T17:50:48.963841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Importance\n\nBased on the gini index computations, a decision tree assigns an \"importance\" value to each feature. These values can be used to interpret the results given by a decision tree.","metadata":{"id":"89UALm00jMk2"}},{"cell_type":"code","source":"model.feature_importances_","metadata":{"id":"bMmZzrwGjNlQ","outputId":"8bd636d2-0828-415c-9650-073892aa5739","execution":{"iopub.status.busy":"2022-08-06T17:50:48.966670Z","iopub.execute_input":"2022-08-06T17:50:48.967057Z","iopub.status.idle":"2022-08-06T17:50:48.975088Z","shell.execute_reply.started":"2022-08-06T17:50:48.967002Z","shell.execute_reply":"2022-08-06T17:50:48.974157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's turn this into a dataframe and visualize the most important features.","metadata":{"id":"nJqarmmLjW2f"}},{"cell_type":"code","source":"importance_df = pd.DataFrame({\n    'feature': train_inputs.columns,\n    'importance': model.feature_importances_\n}).sort_values('importance', ascending=False)","metadata":{"id":"x5BObc-6jULf","execution":{"iopub.status.busy":"2022-08-06T17:50:48.976658Z","iopub.execute_input":"2022-08-06T17:50:48.977795Z","iopub.status.idle":"2022-08-06T17:50:48.988422Z","shell.execute_reply.started":"2022-08-06T17:50:48.977753Z","shell.execute_reply":"2022-08-06T17:50:48.987300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importance_df.head(10)","metadata":{"id":"1Fges4NFjaBE","outputId":"b8e9710d-7cf6-428f-a2ec-143be45d6ec0","execution":{"iopub.status.busy":"2022-08-06T17:50:48.990087Z","iopub.execute_input":"2022-08-06T17:50:48.990773Z","iopub.status.idle":"2022-08-06T17:50:49.006762Z","shell.execute_reply.started":"2022-08-06T17:50:48.990725Z","shell.execute_reply":"2022-08-06T17:50:49.005406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.title('Feature Importance')\nsns.barplot(data=importance_df.head(10), x='importance', y='feature');","metadata":{"id":"PWmpT3PjjcBW","outputId":"d1565456-215f-4a45-ed22-b37857fef612","execution":{"iopub.status.busy":"2022-08-06T17:50:49.008644Z","iopub.execute_input":"2022-08-06T17:50:49.009106Z","iopub.status.idle":"2022-08-06T17:50:49.287400Z","shell.execute_reply.started":"2022-08-06T17:50:49.009058Z","shell.execute_reply":"2022-08-06T17:50:49.286290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model 3-Training a Random Forest\n\nWhile tuning the hyperparameters of a single decision tree may lead to some improvements, a much more effective strategy is to combine the results of several decision trees trained with slightly different parameters. This is called a random forest model. \n\nThe key idea here is that each decision tree in the forest will make different kinds of errors, and upon averaging, many of their errors will cancel out. This idea is also commonly known as the \"wisdom of the crowd\":\n\n<img src=\"https://i.imgur.com/4Dg0XK4.png\" width=\"480\">","metadata":{"id":"hq4HsthJl3KM"}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nmodel = RandomForestClassifier(n_jobs=-1, random_state=42)\n\nmodel.fit(train_inputs, train_targets)","metadata":{"id":"d7eDojEMlxAs","outputId":"b2eb1f93-49fa-4e16-c17b-112660706b97","execution":{"iopub.status.busy":"2022-08-06T17:50:49.288691Z","iopub.execute_input":"2022-08-06T17:50:49.289362Z","iopub.status.idle":"2022-08-06T17:52:11.168156Z","shell.execute_reply.started":"2022-08-06T17:50:49.289320Z","shell.execute_reply":"2022-08-06T17:52:11.166943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`n_jobs` allows the random forest to use mutiple parallel workers to train decision trees, and `random_state=42` ensures that the we get the same results for each execution.","metadata":{"id":"FR_O2M3QYJIi"}},{"cell_type":"code","source":"model.score(train_inputs, train_targets)","metadata":{"id":"IMnIG71wlw72","outputId":"f1e8ae17-3981-4f6c-d1eb-08f4e095cc04","execution":{"iopub.status.busy":"2022-08-06T17:52:11.169674Z","iopub.execute_input":"2022-08-06T17:52:11.170380Z","iopub.status.idle":"2022-08-06T17:52:15.884161Z","shell.execute_reply.started":"2022-08-06T17:52:11.170341Z","shell.execute_reply":"2022-08-06T17:52:15.882952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.score(val_inputs, val_targets)","metadata":{"id":"iof8JdpIlw28","outputId":"5a7dd903-7d75-4616-997d-e49ed10afe74","execution":{"iopub.status.busy":"2022-08-06T17:52:15.887237Z","iopub.execute_input":"2022-08-06T17:52:15.887638Z","iopub.status.idle":"2022-08-06T17:52:17.530313Z","shell.execute_reply.started":"2022-08-06T17:52:15.887603Z","shell.execute_reply":"2022-08-06T17:52:17.529199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once again, the training accuracy is almost 100%, but this time the validation accuracy is much better. In fact, it is better than the best single decision tree we had trained so far. Do you see the power of random forests?\n\nThis general technique of combining the results of many models is called \"ensembling\", it works because most errors of individual models cancel out on averaging. Here's what it looks like visually:\n\n<img src=\"https://i.imgur.com/qJo8D8b.png\" width=\"640\">\n\n\nWe can also look at the probabilities for the predictions. The probability of a class is simply the fraction of trees which that predicted the given class.","metadata":{"id":"rfpuDHi-YV_D"}},{"cell_type":"code","source":"train_probs = model.predict_proba(train_inputs)\ntrain_probs","metadata":{"id":"Qqlh--RrYX_D","outputId":"f072969a-4664-48f8-a6e0-8798eaeebdad","execution":{"iopub.status.busy":"2022-08-06T17:52:17.533003Z","iopub.execute_input":"2022-08-06T17:52:17.533841Z","iopub.status.idle":"2022-08-06T17:52:22.194164Z","shell.execute_reply.started":"2022-08-06T17:52:17.533792Z","shell.execute_reply":"2022-08-06T17:52:22.192271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Importance","metadata":{"id":"HkZZqqoMYvRd"}},{"cell_type":"markdown","source":"Just like decision tree, random forests also assign an \"importance\" to each feature, by combining the importance values from individual trees.","metadata":{"id":"MJtij0PNYpIN"}},{"cell_type":"code","source":"importance_df = pd.DataFrame({\n    'feature': train_inputs.columns,\n    'importance': model.feature_importances_\n}).sort_values('importance', ascending=False)","metadata":{"id":"WIHlVX6Olwwz","execution":{"iopub.status.busy":"2022-08-06T17:52:22.195779Z","iopub.execute_input":"2022-08-06T17:52:22.196198Z","iopub.status.idle":"2022-08-06T17:52:22.306105Z","shell.execute_reply.started":"2022-08-06T17:52:22.196160Z","shell.execute_reply":"2022-08-06T17:52:22.305075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.title('Feature Importance')\nsns.barplot(data=importance_df.head(10), x='importance', y='feature');","metadata":{"id":"Z0kKGRQdY1Lm","outputId":"55e8007b-ef60-481e-fa91-534b895f0765","execution":{"iopub.status.busy":"2022-08-06T17:52:22.307342Z","iopub.execute_input":"2022-08-06T17:52:22.307675Z","iopub.status.idle":"2022-08-06T17:52:22.567343Z","shell.execute_reply.started":"2022-08-06T17:52:22.307645Z","shell.execute_reply":"2022-08-06T17:52:22.566120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nNotice that the distribution is a lot less skewed than that for a single decision tree.","metadata":{"id":"7UUkBbJYY6ca"}},{"cell_type":"markdown","source":"Finally, let's also compute the accuracy of our model on the test set.","metadata":{"id":"oO4vIBf17YKu"}},{"cell_type":"markdown","source":"## Making Predictions on the Test Set\n\nLet's make predictions on the test set provided with the data.","metadata":{"id":"9E00qXMM7c6n"}},{"cell_type":"code","source":"test_df = pd.read_csv('../input/stayalert/fordTest.csv')","metadata":{"id":"sJ2_AiMMOlzo","execution":{"iopub.status.busy":"2022-08-06T17:54:54.192483Z","iopub.execute_input":"2022-08-06T17:54:54.192933Z","iopub.status.idle":"2022-08-06T17:54:54.515579Z","shell.execute_reply.started":"2022-08-06T17:54:54.192895Z","shell.execute_reply":"2022-08-06T17:54:54.514433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{"id":"4GsoJ0DdOltU","outputId":"6dc13d1f-dfce-499f-8a19-eeaa203ec6d9","execution":{"iopub.status.busy":"2022-08-06T17:54:56.526188Z","iopub.execute_input":"2022-08-06T17:54:56.526617Z","iopub.status.idle":"2022-08-06T17:54:56.575091Z","shell.execute_reply.started":"2022-08-06T17:54:56.526581Z","shell.execute_reply":"2022-08-06T17:54:56.574047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, we need to reapply all the preprocessing steps.","metadata":{"id":"JhOaXBdsOwhq"}},{"cell_type":"code","source":"test_df[numeric_cols] = scaler.transform(test_df[numeric_cols])","metadata":{"id":"3DIxu5k6OlnQ","execution":{"iopub.status.busy":"2022-08-06T17:55:00.523802Z","iopub.execute_input":"2022-08-06T17:55:00.524375Z","iopub.status.idle":"2022-08-06T17:55:00.628285Z","shell.execute_reply.started":"2022-08-06T17:55:00.524308Z","shell.execute_reply":"2022-08-06T17:55:00.627197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_preds = model.predict(test_df[numeric_cols])","metadata":{"id":"J9sK305WO7Ns","execution":{"iopub.status.busy":"2022-08-06T17:55:02.061901Z","iopub.execute_input":"2022-08-06T17:55:02.062308Z","iopub.status.idle":"2022-08-06T17:55:02.711237Z","shell.execute_reply.started":"2022-08-06T17:55:02.062273Z","shell.execute_reply":"2022-08-06T17:55:02.709708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.read_csv('../input/stayalert/example_submission.csv')","metadata":{"id":"Ltm_G1XJPLbB","execution":{"iopub.status.busy":"2022-08-06T17:55:03.725914Z","iopub.execute_input":"2022-08-06T17:55:03.726367Z","iopub.status.idle":"2022-08-06T17:55:03.768147Z","shell.execute_reply.started":"2022-08-06T17:55:03.726331Z","shell.execute_reply":"2022-08-06T17:55:03.766142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df","metadata":{"id":"OE8wQJGaPYOf","outputId":"82f3a77b-7526-48f2-a8a3-e7967e4ea14e","execution":{"iopub.status.busy":"2022-08-06T17:55:05.414406Z","iopub.execute_input":"2022-08-06T17:55:05.415218Z","iopub.status.idle":"2022-08-06T17:55:05.428559Z","shell.execute_reply.started":"2022-08-06T17:55:05.415112Z","shell.execute_reply":"2022-08-06T17:55:05.427233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's replace the values of the `Predictipn` column with our predictions.","metadata":{"id":"g3s-KcCOPeu0"}},{"cell_type":"code","source":"submission_df['Prediction'] = test_preds","metadata":{"id":"PiuS4TXjPkMT","execution":{"iopub.status.busy":"2022-08-06T17:55:10.030129Z","iopub.execute_input":"2022-08-06T17:55:10.030799Z","iopub.status.idle":"2022-08-06T17:55:10.036538Z","shell.execute_reply.started":"2022-08-06T17:55:10.030747Z","shell.execute_reply":"2022-08-06T17:55:10.035348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's save it as a CSV file and download it.","metadata":{"id":"Ys5c-NQLPs47"}},{"cell_type":"code","source":"submission_df.to_csv('submission.csv', index=False)","metadata":{"id":"n-HX6BVPPp7E","execution":{"iopub.status.busy":"2022-08-06T17:53:53.357972Z","iopub.execute_input":"2022-08-06T17:53:53.358401Z","iopub.status.idle":"2022-08-06T17:53:55.196801Z","shell.execute_reply.started":"2022-08-06T17:53:53.358366Z","shell.execute_reply":"2022-08-06T17:53:55.195575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"solution = pd.read_csv('../input/stayalert/Solution.csv')","metadata":{"id":"1s6RJFFBQHUo","execution":{"iopub.status.busy":"2022-08-06T17:53:55.201586Z","iopub.execute_input":"2022-08-06T17:53:55.201990Z","iopub.status.idle":"2022-08-06T17:53:55.262265Z","shell.execute_reply.started":"2022-08-06T17:53:55.201952Z","shell.execute_reply":"2022-08-06T17:53:55.261094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"solution","metadata":{"id":"7J5Ldb9BQOA0","outputId":"3ca5e833-1cf4-43b8-8470-06b0141487fb","execution":{"iopub.status.busy":"2022-08-06T17:53:55.264288Z","iopub.execute_input":"2022-08-06T17:53:55.264675Z","iopub.status.idle":"2022-08-06T17:53:55.280093Z","shell.execute_reply.started":"2022-08-06T17:53:55.264639Z","shell.execute_reply":"2022-08-06T17:53:55.278782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Summary and References\n\nThe following topics were covered in this tutorial:\n\n- Downloading a real-world dataset\n- Preparing a dataset for training\n- Training and interpreting decision trees\n- Training and interpreting random forests\n- Overfitting, hyperparameter tuning & regularization\n- Making predictions on single inputs\n\n\n\nWe also introduced the following terms:\n\n* Decision tree\n* Random forest\n* Overfitting\n* Hyperparameter\n* Hyperparameter tuning\n* Regularization\n* Ensembling\n* Generalization\n* Bootstrapping\n\n\nCheck out the following resources to learn more: \n\n- https://scikit-learn.org/stable/modules/tree.html\n- https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html\n- https://www.kaggle.com/willkoehrsen/start-here-a-gentle-introduction\n- https://www.kaggle.com/willkoehrsen/introduction-to-manual-feature-engineering\n- https://www.kaggle.com/willkoehrsen/intro-to-model-tuning-grid-and-random-search","metadata":{"id":"aMIT3DQ48A9e"}},{"cell_type":"code","source":"","metadata":{"id":"KsGs5gse8BhB"},"execution_count":null,"outputs":[]}]}