---
title: "TalkingData with Breaking Bad Feature Engineering (°ロ°)☝ "
author: "Pranav Pandya"
output:
  html_document:
    number_sections: false
    code_folding: hide
    toc: true
    toc_depth: 6
    fig_width: 10
    highlight: tango
    theme: cosmo
    smart: true
editor_options: 
  chunk_output_type: console
---

Remember Walter White and train scene from Breaking Bad?

![Nothing stops this model training](https://cloud.lovindublin.com/images/uploads/2015/12/_blogWide/nothing-stops-this-train.gif)


# Introduction


Cool thing about R is that there are thousands of packages available for variety of purpose. 
A simple search on binning in Google led me to this awesome breaking bad package **woeBinning** which implements an automated binning with respect to a dichotomous target variable. 
It uses similar weight of evidence and information value criteria to merge and split bins. It is possible to use it for single variables or an entire data frame. 
I found this technique much useful and as an **early warning system** to detect useless or suspicious features. Later I found out that credit risk modelling is a whole new chapter and there are many packages like this. 
Luckily, the package documentation was quite short and easy to understand. In this report, it's my effort to make use of such techniques to see how feature engnineering can be improved. 

# What is Binning?

Binning is the term used in scoring modeling for what is also known in Machine Learning as Discretization, the process of transforming a continuous characteristic into a finite number of intervals (the bins), which allows for a better understanding of its distribution and its relationship with a binary variable.
The bins generated by the this process will eventually become the attributes of a predictive characteristic.

Supervised Discretization divides a continuous feature into groups (bins) mapped to a target variable. 
The central idea is to find those cutpoints that maximize the difference between the groups.


## What are the benefits?
- Efficient and organized way to significantly reduce the time spent on data exploration, feature engineering, variable selection and other recurrent tasks.
- Detect and eliminate useless or suspicious features (important for this dataset)
- IV (Information value) criteria as shown in the later parts may be able to answer the most asked question about **whether or not to use future features** i.e. frequency count features. 

In other words, let's say I have created huge number of new features but not all of them will be useful. Knowing predictive power of those features is useful for variable selection during model building. We will assess the preditive power with information value and weight of evidence criteria.

# Data preparation

## Load training data

We will use 1 million observations of training data that are close to the date and time of test set for this example. 
```{r message=FALSE}
if (!require("pacman")) install.packages("pacman")
pacman::p_load(knitr, kableExtra, DT, pryr, tidyverse, data.table, fasttime, woeBinning, lubridate, tictoc, DescTools)
set.seed(84)               
options(scipen = 9999, warn = -1, digits= 4)

#Shorten the code by declaring known values
train_path <- "../input/train.csv"
vars <- c("ip", "app", "device", "os", "channel", "click_time", "attributed_time", "is_attributed")
most_freq_hours_in_test_data <- c("4","5","9","10","13","14")

total_rows <- 184903890
chunk_rows <- 1000000
skip1_rows <- total_rows - chunk_rows 

train <- fread(train_path, skip=skip1_rows, nrow=chunk_rows, showProgress = FALSE,
               col.names = vars) %>% select(-c(attributed_time))

datatable(head(train))
```

## Basic feature engineering
We will analyse predictive power of default and custom features to make an educated guess on using them. For this example, I have chosen time based features and frequency count features with various interactions.

### Create functions to avoid repititive code

- fun_add_basic_features function is used for feature engineering 
- prepare_dataset function is to apply feature engineering on selected dataframe i.e train in this example
```{r message=FALSE}
# for feature engineering
fun_add_basic_features <- function(dataframe) {
  cat("- Adding basic features... ", "\n")
  dataframe <- dataframe %>% 
    mutate(wday = Weekday(click_time), 
           hour = as.numeric(hour(click_time)),
           min  = as.numeric(minute(click_time)),
           sec  = as.numeric(second(click_time)),
           test_hour_group = ifelse(hour %in% most_freq_hours_in_test_data, 1, 2), 
           numeric_click_time = as.numeric(fastPOSIXct(format(as.character(click_time))))) %>%
    add_count(ip) %>% rename("nip" = n)  %>%
    add_count(ip, wday) %>% rename("nip_d" = n)   %>%
    add_count(ip, wday, hour) %>% rename("nip_d_h" = n)  %>%
    add_count(device, hour) %>% rename("ndev_h" = n)  %>%
    add_count(ip, device, wday, hour) %>% rename("nip_dev_d_h" = n)  %>%
    add_count(channel, hour, device) %>% rename("nchan_dev_h" = n)  %>%
    add_count(ip, channel, app) %>% rename("nip_chan_app" = n) 
  return(dataframe)
} 

# to prepare dataset
prepare_dataset <- function(dataframe){
  cat(">>>>---Piping it down--->>>>", "\n")
  tic("Total processing time --->")
  dataframe <- fun_add_basic_features(dataframe)
  print("finished preparing data...")
  toc()
  return(dataframe)
}

train <- as.data.frame(prepare_dataset(train))
datatable(head(train))
char_vars <- c("ip", "app", "device", "os", "channel", "numeric_click_time", "wday", "hour", "test_hour_group", "min", "sec")
freq_vars <- c("nip", "nip_d", "nip_d_h", "ndev_h", "nip_dev_d_h", "nchan_dev_h", "nip_chan_app")
```

Categorical features
```{r message=FALSE}
char_vars
```

Frequency count features
```{r message=FALSE}
freq_vars
```

# Let's get straight into this **Weight of Evidence Binning** thingy thing!

In our case, most of our default variables are categorical. We use woe.binning function which generates a supervised fine and coarse classing of numeric variables and factors with respect to a dichotomous target variable. 

This function provides flexibility in finding a binning that fits specific data characteristics and practical needs.

## Example with single feature (**IP address**)

```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "ip")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E")
  
```

Let's understand what this table represents. 

Output provides nice summary of count and distribution of **binned IP address** with respect to our target vaiable i.e is_attributed. The one we are most interested in is weight of evidence (WOE) and information values (IV) columns. 

We can see that total IV for IP feature is **1.323 which means suspicious predictor**. Notice that even if we use the bins, both have IV of > 0.6. Let’s understand what does it mean by IV and what it represents. In the next section, we will discover how to find optimal bins. 

From left to right, this table contains following columns:

- the final bin labels
- total counts
- total distribution (column percentages)
- counts for the first and the second target class (in our case 0 for app not downloaded and 1 for app downloaded)
- distribution of the first and the second target class (column percentages)
- rate (row percentages) of the target event 
- weight of evidence (WOE) 
- information values (IV).

## Key terms to remember (useful for later part)

- **IV** : Information Value
- **WOE**: Weight of Evidence

In credit scoring, Information Value (IV) is frequently used to compare predictive power among variables. When developing new scorecards using logistic regression, variables are often binned and recoded using WoE concept. 

## How to interpret value of IV?

    Information Value     Predictive Power
      < 0.02              useless for prediction
        0.02 to 0.1       Weak predictor
        0.1 to 0.3        Medium predictor
        0.3 to 0.5        Strong predictor
      > 0.5               Suspicious or too good to be true

# Procedure for optimal binning to avoid overfitting
Although this may be lengthy read but I hope it would be helpful for quick reference:

## Formula 

	woe.binning(df, target.var, pred.var, min.perc.total, min.perc.class, 
				stop.limit, abbrev.fact.levels, event.class)


## Description of important arguments: 

df	

	Name of data frame with input data.

target.var	

	Name of dichotomous target variable in quotes. Only target variables with two distinct values (e.g. 0, 1 or “Y”, “N”) are accepted; cases with NAs in the target variable will be ignored.

pred.var	

	Name of predictor variable(s) to be binned in quotes. A single variable name can be provided, e.g. “varname1”, or a list of variable names, e.g. c(“varname1”, “varname2”). Alternatively one can repeat the name of the input data frame; the function will be applied to all its variables apart from the target variable then. Numeric variables and factors are supported and may contain NAs.

**min.perc.total** (Important)

	For numeric variables this parameter defines the number of initial classes before any merging is applied. For example min.perc.total=0.05 (5%) will result in 20 initial classes. For factors the original levels with a percentage below this limit are collected in a ‘miscellaneous’ level before the merging based on the min.perc.class and on the WOE starts. Increasing the min.perc.total parameter will avoid sparse bins. Accepted range: 0.01-0.2; default: 0.05.

**min.perc.class**	(Important)

	If a column percentage of one of the target classes within a bin is below this limit (e.g. below 0.01=1%) then the respective bin will be joined with others. In case of numeric variables adjacent predictor classes are merged. For factors respective levels (including sparse NAs) are assigned to a ‘miscellaneous’ level. Setting min.perc.class>0 may provide more reliable WOE values. Accepted range: 0-0.2; default: 0, i.e. no merging with respect to sparse target classes is applied.

**stop.limit**	(Important)

	Stops WOE based merging of the predictor's classes/levels in case the resulting information value (IV) decreases more than x% (e.g. 0.05 = 5%) compared to the preceding binning step. stop.limit=0 will skip any WOE based merging. Increasing the stop.limit will simplify the binning solution and may avoid overfitting. Accepted range: 0-0.5; default: 0.1.

abbrev.fact.levels	

	Abbreviates the names of new (merged) factor levels via the base R abbreviate function in case the specified number of characters is exceeded. Accepted range: 0-1000; default: 200. 0 will prevent applying any abbreviation, i.e. only factor levels with more than 1000 characters will be truncated then. This option is particularly relevant in case one wants to generate dummy variables via the woe.binning.deploy function, because the factor levels will be part of the dummy variable names then.

event.class	

  Optional parameter for specifying the class of the target event. This class typically indicates a negative event like a loan default or a disease. Use integers (e.g. 1) or characters in quotes (e.g. “bad”). This class will be represented by negative WOE values then.





# Tabulate WOE and IV for all categorical variables

**Note** : I'm using default parameters for this example and for illustration purpose. Tuning above mentioned additional parameters is highly recommended. 

## Device
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "device")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#0f8b39")
  
```

Notice that device 0 and 1 contributes 83.5% of distribution where app was downloaded but IV is low. Using device feature as it is has medium prediction power of 0.116

## App
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "app")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3:5, bold = T, color = "white", background = "#0f8b39")
  
```


Using app feature could be bad idea depending on total IV. However using discrete bins shows good predictive power. 

## OS
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "os")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(2:3, bold = T, color = "white", background = "#0f8b39")
```

## Channel
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "channel")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(9, bold = T, color = "white", background = "#0f8b39")
```

## Click time (numeric)
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "numeric_click_time")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E")
```

## wday
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "wday")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E")
  
```

As we can see that IV is 0 for weekday, hour and test_hour_group and can be viewed as **useless predictors**. 
Using them in model will probably contribute some noise in my opinion.


## hour
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "hour")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E") 
  
```

## test_hour_group
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "test_hour_group")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E")
```

## min
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "min")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(4, bold = T, color = "white", background = "#D7261E")
```

## sec
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "sec")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(8, bold = T, color = "white", background = "#D7261E")
```

# Tabulate WOE and IV for all numeric variables 

For this example, I have genrated few numeric features as shown in feature engineering part. We use the same woe.binning function which generates a supervised fine and coarse classing of numeric variables and factors with respect to a dichotomous target variable. 

## IP counts (nip)
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "nip")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E")
  
```

Turns out that ip count, ip-device-hour count and ip-device-hour count are either suspicious or too good to be true kind of predictors depending on their total IV.

## IP-day count
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "nip_d")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E")
```

## ip-device-hour count
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "nip_d_h")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E") 
```

## device-hour count
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "ndev_h")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(2, bold = T, color = "white", background = "#0f8b39")
```

## IP-device-day-hour count
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "nip_dev_d_h")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#D7261E")
```

## IP-channel-app count
```{r message=FALSE}
dt <- as.data.frame(woe.binning.table(woe.binning(train, 'is_attributed', "nip_chan_app")))
names(dt) <- sub("^WOE.Table.for.", "", names(dt))
kable(dt, "html") %>%
  kable_styling("responsive", full_width = T, position = "left") %>%
  column_spec(9:10, bold = T) %>%
  row_spec(3, bold = T, color = "white", background = "#0f8b39")
```


# Everything together
Instead of separate evaluation for each variables, it is also possible to process everything together for selected variables and add binning features into main dataframe as shown in the code chunk below:

```{r message=FALSE}
bin_vars <- c("ip", "app", "device", "os", "channel", "numeric_click_time", "wday", "hour", "min", "sec", "test_hour_group", 
              "nip", "nip_d", "nip_d_h", "ndev_h", "nip_dev_d_h", "nchan_dev_h", "nip_chan_app")
tic("Total processing time for woe binning--->")
binning <- woe.binning(train, 'is_attributed', bin_vars)
#woe.binning.table(binning) # Tabulate the binned variables (same like above)
# woe.binning.plot(binning) # problem displaying pop up Plots on kaggle kernel
train <- woe.binning.deploy(train, binning, add.woe.or.dum.var = "woe") %>% select_if(~ !is.factor(.)) 
str(train)

```

Note that the idea is not to have as much new features as possible but to have features with strong prection power as per IV. Illustration for so many new features above is just to give an idea that it's possible to merge binned features with main dataframe. 

# Visualization

## Variable ranking by IV

**Important note** : Following ranking doesn't mean the top variables are the best ones. Please follow the thumb rule as shown below for educated guess:


- **IV** means Information Value and **WOE** means Weight of Evidence. **Any variable with IV > 0.5 is suspicious.** 


        Information Value     Predictive Power
        <0.02               useless for prediction
        0.02 to 0.1         Weak predictor
        0.1 to 0.3          Medium predictor
        0.3 to 0.5          Strong predictor
        >0.5                Suspicious or too good to be true



![vars_by_IV](https://github.com/pranavpandya84/Kaggle_kernels/blob/master/TalkingData_plots/var_ranking_by_IV.png)

In case, image is not shown, [here][https://github.com/pranavpandya84/Kaggle_kernels/blob/master/TalkingData_plots/var_ranking_by_IV.png] is the link.




# Final thoughts
- Using features without evaluating them is time consuming process. With IV and WOE, we can make educated guess on whether or not to use them even before modelling.
- Variables with medium and strong predictive powers are indeed the best choice for model development. (at least theoritically.)
- Using customized bins for weak predictor may turn them into good predictor.
- Considering the huge size of training data, few features with strong prediction power (**between 0.3 to 0.5 IV**) may result into comparatively better score. 


Note: Things that are proven theoritically may go wrong practically depending on the type of data, sampling from complete data and indirect objectives. 
It's advisable to have a look at some of the past competitions with abnormal data (mainly discussion section) to better understand best and proven practices.
    
# References

- https://cran.r-project.org/web/packages/woeBinning/woeBinning.pdf
- http://blog.revolutionanalytics.com/2015/03/r-package-smbinning-optimal-binning-for-scoring-modeling.html
- http://ucanalytics.com/blogs/information-value-and-weight-of-evidencebanking-case/

**Related work**
---

- **R**:

    - [TalkingData: EDA to Model Evaluation](https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683/)
    - [LightGBM in R with 75 mln rows](https://www.kaggle.com/pranav84/single-lightgbm-in-r-with-75-mln-rows-lb-0-9690?scriptVersionId=2989011)
    - [Single Histogram optimised XGBoost in R](https://www.kaggle.com/pranav84/single-xgboost-hist-hitting-0-9686-on-lb?scriptVersionId=2968554)

- **Python**:

    - [LightGBM (Fixing unbalanced data)](https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-lb-0-9680)
    - [XGBoost : Histogram Optimized Version](https://www.kaggle.com/pranav84/xgboost-histogram-optimized-version?scriptVersionId=2794247)