{"cells":[{"metadata":{"_uuid":"fd8cb28c947feaaa379dbc6180be371671f37769","_execution_state":"idle","trusted":false},"cell_type":"markdown","source":"**Introduction**\n\nXGBoost (which stands for eXtreme Gradient Boosting) is an especialy efficent implimentation of gradient boosting. In practice, XGBoost is a very powerful tool for classification and regression. In particular, it has proven to be very powerful in Kaggle competitions, and winning submissions will often incorporate it. In this tutorial, you'll learn how to take a new dataset and use XGBoost to make predictions.\n\n**Why xgboost?**\n\nIn kaggle competition the most popular winning algorithm was a Random Forest. However, this has changed over the last six months. A new algorithm XGboost is becoming a winner, it is taking over practically every competition for structured data.\n\n**What will I see here?**\n\nBy the time you finish this tutorial, you will see:\n\n**What XGBoost is**\n\nHow to prepare your data\nHow to train and tune a model using XGBoost\nHow to visualize & explore your model\n\nFor the exploration of the data I will not come back because I made a tutorial on subject using the model of the logistic regression which you can consult it on this link: https://www.kaggle.com/mahmoud86/credit-defaut-risk-binary-logistic-regression\n\n\n**What is XGBOOST?**\n\nXGBoost is an implementation of the Gradient Boosted Decision Trees algorithm.\n\n"},{"metadata":{"trusted":true,"_uuid":"2d1f6db050640278bf6f2e65c0c8887cffe66aa1"},"cell_type":"code","source":"library(xgboost) # for xgboost\nlibrary(tidyverse)# general utility functions\nlibrary(corrplot)\nlibrary(knitr) \nlibrary(ggplot2) # Data visualization\nlibrary(readr) # CSV file I/O, e.g. the read_csv function\nlibrary(caret)\nlibrary(DMwR) # smote\nlibrary(Matrix)\nlibrary(reshape) #melt\nlibrary(pROC) # AUC\nlibrary(gridExtra)\nlibrary(scales)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6e813ec9f5089185d74d5418a3edf0c876ac8269"},"cell_type":"code","source":"test <- read.csv(\"../input/application_test.csv\", stringsAsFactors = F)\ntrain <- read.csv(\"../input/application_train.csv\", stringsAsFactors = F)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"74790aec1e296d35b978f18528e6e3b4a7b16933"},"cell_type":"code","source":"# And change ID in test to character\ntest$SK_ID_CURR <- as.numeric(test$SK_ID_CURR)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a7d5cb7f7701db29d290ba24d46c62f85a6bc53c"},"cell_type":"code","source":"#A glance on our data\nhead(train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"14f8c1bdd531312c343d6b01e4fc6d9ca40943a3"},"cell_type":"markdown","source":"Are there any missing data?"},{"metadata":{"trusted":true,"_uuid":"72f8a5b9c8add5957c95d9980641b64cda817610"},"cell_type":"code","source":"NAcol <- which(colSums(is.na(train)) > 0)\nsort(colSums(sapply(train[NAcol], is.na)), decreasing = TRUE)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"974f766da214d11f8d589287b125a6f1e219c190"},"cell_type":"code","source":"miss_pct <- map_dbl(train, function(x) { round((sum(is.na(x)) / length(x)) * 100, 1) })\n\nmiss_pct <- miss_pct[miss_pct > 0]\n\ndata.frame(miss=miss_pct, var=names(miss_pct), row.names=NULL) %>%\n    ggplot(aes(x=reorder(var, -miss), y=miss)) + \n    geom_bar(stat='identity', fill='red') +\n    labs(x='', y='% missing', title='Percent missing data by feature') +\n    theme(axis.text.x=element_text(angle=90, hjust=1))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3d4a483bb017420b25e92697c9802e541a0f6449"},"cell_type":"markdown","source":"\nFocus the missing value variable less than or equal to 5%\nWhy missing value less thant or equal 5%?\nIn missing data theory, the method of comprehensive data analysis is the simplest and most common (this is\nthe default method of many software, R included). It consists in keeping only\nobservations that do not contain any missing data. Statisticians admit\nthat this method is acceptable when individuals with missing values\nrepresent less than 5% of the population. Otherwise, it quickly becomes dangerous because\nmany observations can disappear from the dataset (indeed, the\nproportion of complete observations may be low even though, for each variable, the\nprobability of a data being observed is large). Even if it is acceptable, this\nmethod is unbiased only for missing MCAR data, but will have\ndespite all tendencies to decrease the accuracy of modeling.\n\n\nIn our Dtrain and Dtest we will maintain that missing data less than 5%"},{"metadata":{"trusted":true,"_uuid":"2cf47b843123d3e1e776da4502fcaa20e792c524"},"cell_type":"code","source":"missing_data <- as.data.frame(sort(sapply(train, function(x) sum(is.na(x))),decreasing = T))\n\nmissing_data <- (missing_data/nrow(train))*100\n\ncolnames(missing_data)[1] <- \"missingvaluesPercentage\"\nmissing_data$features <- rownames(missing_data)\nggplot(missing_data[missing_data$missingvaluesPercentage <=5,],aes(reorder(features,-missingvaluesPercentage),missingvaluesPercentage,fill= features)) +\n  geom_bar(stat=\"identity\") +theme_minimal() +\n  theme(axis.text.x = element_text(angle = 90, hjust = 1), legend.position = \"none\") + ylab(\"Percentage of missingvalues\") +\n  xlab(\"Feature\") + ggtitle(\"Understanding Missing Data\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"aa342fe5639c57cdd0b486843257ee0b8a9862bd"},"cell_type":"code","source":"\noptions(repr.plot.height=4)\nNAcol <- which(colSums(is.na(train)) > 0)\nNAcount <- sort(colSums(sapply(train[NAcol], is.na)), decreasing = TRUE)\nNADF <- data.frame(variable=names(NAcount), missing=NAcount)\nNADF$PctMissing <- round(((NADF$missing/nrow(train))*100),1)\n\nNADF %>%\n    ggplot(aes(x=reorder(variable, PctMissing), y=PctMissing)) +\n    geom_bar(stat='identity', fill='blue') + coord_flip(y=c(0,110)) +\n    labs(x=\"\", y=\"Percent missing\") +\n    geom_text(aes(label=paste0(NADF$PctMissing, \"%\"), hjust=-0.1))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2dde4c6a02d323fa4195a4a0202973c06703fa80"},"cell_type":"markdown","source":"Here are the variables that can be in the predictive model\n\nExploring some of the most important variables\n\nThe response variable; TARGET"},{"metadata":{"trusted":true,"_uuid":"42c8c8797b65f1fff2ca58bd8c4bb9b99015787d"},"cell_type":"code","source":"train$TARGET<-as.factor(train$TARGET)\nggplot(train[!is.na(train$TARGET),], aes(x = TARGET, fill = TARGET)) +\n  geom_bar(stat='count') +\n  labs(x = 'TARGET') +\n        geom_label(stat='count',aes(label=..count..), size=7) +\n        theme_grey(base_size = 18)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"8d780fc7e512c2494467a9120f78ce7ac0cfba84"},"cell_type":"code","source":"str(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"92ea1bd62c3806462196bcd19f4518bb9e042fda"},"cell_type":"code","source":"dtrain<-train[,c(\"TARGET\",\"OBS_30_CNT_SOCIAL_CIRCLE\",\"DEF_30_CNT_SOCIAL_CIRCLE\",\"OBS_60_CNT_SOCIAL_CIRCLE\",\"DEF_60_CNT_SOCIAL_CIRCLE\",\"EXT_SOURCE_2\",\"AMT_GOODS_PRICE\",\"DAYS_LAST_PHONE_CHANGE\",\"DAYS_BIRTH\",\"DAYS_EMPLOYED\",\"DAYS_REGISTRATION\",\"AMT_ANNUITY\",\"AMT_CREDIT\",\"AMT_INCOME_TOTAL\",\"REGION_POPULATION_RELATIVE\",\"CNT_CHILDREN\",\"CNT_FAM_MEMBERS\",\"FLAG_MOBIL\",\"FLAG_EMP_PHONE\",\"FLAG_WORK_PHONE\",\"FLAG_PHONE\",\"CODE_GENDER\",\"FLAG_OWN_CAR\",\"FLAG_OWN_REALTY\")]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8ed5a840ca2880bb4e08845d9aa0b91a91a7db9f"},"cell_type":"code","source":"numericVars <- which(sapply(dtrain, is.numeric)) #index vector numeric variables\ndat_numVar <- dtrain[, numericVars]\ncor_numVar <- cor(dat_numVar, use=\"pairwise.complete.obs\") #correlations of all numeric variables\n#sort on decreasing correlations with degree_spondylolisthesis\ncor_sorted <- as.matrix(sort(cor_numVar[,'EXT_SOURCE_2'], decreasing = TRUE))\n#select only high corelations\nCorHigh <- names(which(apply(cor_sorted, 1, function(x) abs(x)>0)))\ncor_numVar <- cor_numVar[CorHigh, CorHigh]\ncorrplot.mixed(cor_numVar, tl.col=\"black\", tl.pos = \"lt\", tl.cex = 0.7,cl.cex = .7, number.cex=.7)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"05b1e8ff83bab2b391a518d184c2cf19bc21be53","scrolled":true},"cell_type":"code","source":"p1<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$OBS_30_CNT_SOCIAL_CIRCLE)),], aes(x =OBS_30_CNT_SOCIAL_CIRCLE  , fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and OBS_30_CNT_SOCIAL_CIRCLE \") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np2<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$AMT_GOODS_PRICE)),], aes(x = AMT_GOODS_PRICE, fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and AMT_GOODS_PRICE\") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np3<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$DEF_30_CNT_SOCIAL_CIRCLE)),], aes(x = DEF_30_CNT_SOCIAL_CIRCLE , fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and DEF_30_CNT_SOCIAL_CIRCLE\") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np4<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$OBS_60_CNT_SOCIAL_CIRCLE)),], aes(x =OBS_60_CNT_SOCIAL_CIRCLE, fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and OBS_60_CNT_SOCIAL_CIRCLE\") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np5<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$DEF_60_CNT_SOCIAL_CIRCLE )),], aes(x = DEF_60_CNT_SOCIAL_CIRCLE , fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and DEF_60_CNT_SOCIAL_CIRCLE \") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np6<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$EXT_SOURCE_2)),], aes(x =EXT_SOURCE_2 , fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and EXT_SOURCE_2\") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np7<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$AMT_ANNUITY)),], aes(x = AMT_ANNUITY , fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and AMT_ANNUITY\") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np8<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$CNT_FAM_MEMBERS)),], aes(x =CNT_FAM_MEMBERS, fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and CNT_FAM_MEMBERS \") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\np9<-ggplot(dtrain[(!is.na(dtrain$TARGET) & !is.na(dtrain$DAYS_LAST_PHONE_CHANGE)),], aes(x = DAYS_LAST_PHONE_CHANGE , fill = TARGET)) +\ngeom_density(alpha=0.5, aes(fill=factor(TARGET))) + labs(title=\"Target density and DAYS_LAST_PHONE_CHANGE \") +\nscale_x_continuous(breaks = scales::pretty_breaks(n = 10)) + theme_grey()\ngrid.arrange(p1, p2,p3,p4,p5,p6,p7,p8,p9)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3fa2b13965f6592112786cb14b3fb9a1515fc53d"},"cell_type":"markdown","source":"Traitment of missing value \n\nHow to treat missing values?\nAbove all, spend some time describing the missing data in your game\ndata: what is the percentage of missing data per variable? By group of\npopulation? Are there potential predictors of missing data? etc. this will\nwill then allow to choose the most appropriate method. Regarding the methods of\ntreatment, know that there is a lot of it (it is not for nothing that Little and Rubin\nhave devoted a book on this subject). In practice, it may be necessary to test several\ncorrection approaches. We will check over the corrections the impact of treatments\nback and forth between the raw data and the corrected data. The full data analysis method is the simplest and most\nthe default method of many software, R included). It consists in keeping only\nobservations that do not contain any missing data. Statisticians admit\nthat this method is acceptable when individuals with missing values\nrepresent less than 5% of the population. Otherwise, it quickly becomes dangerous because\nmany observations can disappear from the dataset (indeed, the\nproportion of complete observations may be low even though, for each variable, the\nprobability of a data being observed is large). Even if it is acceptable, this\nmethod is unbiased only for missing MCAR data, but will have\ndespite all tendencies to decrease the accuracy of modeling.\nLet's talk about a richer method: imputation. Its goal is to\nreplace the missing value with an artificially generated value from the others\ndata available. There are many variants that are valid, depending on the case,\nfor missing MCAR and / or MAR data. There are two main approaches:\nsimple imputation and multiple imputation. Many possibilities of imputation\nsimple exist, for example:\n• imputation by rule: if we know it, we can define a rule that allows us to\ncalculate the missing value from other available data;\n• the replacement of the missing values ​​by the average value of all the\nresponses (variants: median, mode for qualitative data,\netc.);\n• imputation by ratio or by regression: one defines the missing value by the value\npredicted by a regression model including one (ratio) or several (regression)\nother variables available2. There is a stochastic variant: we draw a value\naround the prediction of the model;\n• the hot-deck and the cold-deck: one draws randomly with discount a value among\nall individuals with an observed value for the variable of interest, in order to\nreplace the missing value. For the hot version, this draw is done directly\nin the analyzed dataset; for the cold version, it is done in a source of\nexternal data (this variant is therefore rarely applicable in machine learning);\n• the nearest neighbor method: the principle is the same as hot and cold-deck,\nbut the draw is done in individuals with the same characteristics. For that, we\nmust therefore define a function of distance which makes it possible to characterize the proximity between\nthese individuals (see chapter on classification for more details on this notion of\ndistance).\nAttention, even in the case where the assumption of missing data MAR or MCAR is\njustified, these methods are not neutral and can significantly alter the\ndata. For example, they can modify the relationships between variables. In addition,\noften considers that these methods underestimate uncertainty. Indeed, it is difficult to\nbelieve that a single imputed value will be able to represent all the uncertainty\nrelating to the value to be imputed.\nTo overcome this problem, multiple imputation methods can be used. Their\ngeneral principle is as follows: • Step 1: Replace each missing value with M (> 1) values ​​from a\nappropriate distribution.\nFigure 16-1 - Multiple Imputation: Step 1\n• Step 2: independent analyzes are performed, but with the same method, from M bases\ncharged.\nFigure 16-2 - Multiple Imputation: Step 2\n• Step 3: Combine the results of these analyzes to reflect the variability\nadditional information due to missing data. Figure 16-3 - Multiple Imputation: Step 3\nThe bigger the M, the more precise the results, but we generally consider that we have\ngood results from M = 5. Remember that the goal is not to predict the data\nmissing with the utmost precision, but to best reflect the uncertainty of\nmissing values, distributions and relationships between variables.\nTo perform this multiple imputation, several algorithms exist. The best known is\nalgorithm.\n\n\nIn our case we will replace the missing data by the median\n\n"},{"metadata":{"trusted":true,"_uuid":"2d5a940836b578bb5a97fbcd87db42e968c0daa7"},"cell_type":"code","source":"NAcol <- which(colSums(is.na(dtrain)) > 0)\nsort(colSums(sapply(dtrain[NAcol], is.na)), decreasing = TRUE)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"75c7f79a7651d84d4ecd59f587839c96b0822435","scrolled":true},"cell_type":"code","source":"str(dtrain)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"87c061d4ccefe3b214a22d6b6fab87a0cd720663"},"cell_type":"code","source":"dtrain$AMT_GOODS_PRICE[is.na(dtrain$AMT_GOODS_PRICE)]<-450000 \ndtrain$AMT_ANNUITY[is.na(dtrain$AMT_ANNUITY)]<-24903 \ndtrain$CNT_FAM_MEMBERS[is.na(dtrain$CNT_FAM_MEMBERS)]<-2\ndtrain$OBS_30_CNT_SOCIAL_CIRCLE[is.na(dtrain$OBS_30_CNT_SOCIAL_CIRCLE)]<-0\ndtrain$DEF_30_CNT_SOCIAL_CIRCLE[is.na(dtrain$DEF_30_CNT_SOCIAL_CIRCLE)]<-0\ndtrain$OBS_60_CNT_SOCIAL_CIRCLE[is.na(dtrain$OBS_60_CNT_SOCIAL_CIRCLE)]<-0\ndtrain$DEF_60_CNT_SOCIAL_CIRCLE[is.na(dtrain$DEF_60_CNT_SOCIAL_CIRCLE)]<-0\ndtrain$EXT_SOURCE_2[is.na(dtrain$EXT_SOURCE_2)]<-0.5660 \ndtrain$DAYS_LAST_PHONE_CHANGE[is.na(dtrain$DAYS_LAST_PHONE_CHANGE)]<-757.0\ndtrain$DAYS_BIRTH<-dtrain[,'DAYS_BIRTH']/ -365\ndtrain$DAYS_EMPLOYED<-dtrain[,'DAYS_EMPLOYED']*-1\ndtrain$DAYS_REGISTRATION<-dtrain[,'DAYS_REGISTRATION']*-1\ndtrain$DAYS_LAST_PHONE_CHANGE<-dtrain[,'DAYS_LAST_PHONE_CHANGE']*-1","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ae978b616113a90ab1fca39226dba29971e8aac7"},"cell_type":"markdown","source":"What is XGBoost?\n\nXGBoost is an implementation of the Gradient Boosted Decision Trees algorithm. What is Gradient Boosted Decision Trees? We'll walk through a diagram."},{"metadata":{"trusted":true,"_uuid":"9afb16e2c3294312920d65d438f9d2c616fffe9f"},"cell_type":"code","source":"dtrain_numeric <- dtrain %>%\n    select(-TARGET) %>% # the case id shouldn't contain useful information\n    select_if(is.numeric) # select remaining numeric columns\n# make sure that our dataframe is all numeric\nstr(dtrain_numeric)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"487e12a7f4c99e1be0e2de167e1771f4d7fb9e74"},"cell_type":"markdown","source":"Convert categorical information  to a numeric format\n\n\nAlright, so now we've got all the numeric values we need. But what about non-numeric variables? For example, we have a column \"code gender\" that tells us which gender an observation is from."},{"metadata":{"trusted":true,"_uuid":"b595cf4d1d766e00f9ff6993bd359e77a8e1931a"},"cell_type":"code","source":"# one-hot matrix for just the first few rows of the \"gender\" column\nmodel.matrix(~CODE_GENDER-1,head(dtrain))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3c7f0ed5e8e322a5ac9f29cd8e03e1e57ff9f042"},"cell_type":"markdown","source":"How  we are convert these categories to a matrix? One way to do this is using one-hot encoding. One-hot encoding takes each category and makes it its own column. Then, for each observation, it puts a \"0\" in that column if that observation doesn't belong to that column and \"1\" if it does. In R, we can convert a column with a categorical variable in it to a one-hot matrix like so:"},{"metadata":{"_uuid":"507799e3dec192156ab1f9d909e5e95a01b1f2cb"},"cell_type":"markdown","source":"We will do the same for FLAG_OWN_CAR and FLAG_OWN_REALTY"},{"metadata":{"trusted":true,"_uuid":"b9e34cb81e3c2460ec6f1c242efc9ad5aa5d1228"},"cell_type":"code","source":"Car<-model.matrix(~FLAG_OWN_CAR-1,dtrain)\nRealty<-model.matrix(~FLAG_OWN_REALTY-1,dtrain)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"91a19b1ddc0335f06fa73cc87ae6ce6c7a8f971f"},"cell_type":"code","source":"dtrain_numeric1 <- cbind(dtrain_numeric,Car,Realty)\ndtrain_matrix <- data.matrix(dtrain_numeric1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0c7df9fb73346fc1c8413da636743be9eb5d96aa"},"cell_type":"markdown","source":"Target"},{"metadata":{"trusted":true,"_uuid":"25c870296cd67eeef134e1e64e39a4c9557c8a45"},"cell_type":"code","source":"train_labels<-dtrain[,1]\ntrain_labels<-as.numeric(as.factor(train_labels))-1","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"111d568920cc2b62f54b3ef669d0103fa004e9d4"},"cell_type":"markdown","source":"For test"},{"metadata":{"trusted":true,"_uuid":"c7ca5862397de43f4cf37aac5ed12f254b3ee571"},"cell_type":"code","source":"dtest<-test[,c(\"OBS_30_CNT_SOCIAL_CIRCLE\",\"DEF_30_CNT_SOCIAL_CIRCLE\",\"OBS_60_CNT_SOCIAL_CIRCLE\",\"DEF_60_CNT_SOCIAL_CIRCLE\",\"EXT_SOURCE_2\",\"AMT_GOODS_PRICE\",\"DAYS_LAST_PHONE_CHANGE\",\"DAYS_BIRTH\",\"DAYS_EMPLOYED\",\"DAYS_REGISTRATION\",\"AMT_ANNUITY\",\"AMT_CREDIT\",\"AMT_INCOME_TOTAL\",\"REGION_POPULATION_RELATIVE\",\"CNT_CHILDREN\",\"CNT_FAM_MEMBERS\",\"FLAG_MOBIL\",\"FLAG_EMP_PHONE\",\"FLAG_WORK_PHONE\",\"FLAG_PHONE\",\"CODE_GENDER\",\"FLAG_OWN_CAR\",\"FLAG_OWN_REALTY\")]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6aab3051c7b6259cc3aaea0089a4527918bcecfa"},"cell_type":"code","source":"NAcol <- which(colSums(is.na(dtest)) > 0)\nsort(colSums(sapply(dtest[NAcol], is.na)), decreasing = TRUE)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fe8bab0b5ac5aac419ce08e447b490109ef2d2da"},"cell_type":"markdown","source":"Treatment Missing values"},{"metadata":{"trusted":true,"_uuid":"507891106119c0156b25e4fab7db47c3b483d334"},"cell_type":"code","source":"\ndtest$AMT_ANNUITY[is.na(dtest$AMT_ANNUITY)]<-26199 \ndtest$OBS_30_CNT_SOCIAL_CIRCLE[is.na(dtest$OBS_30_CNT_SOCIAL_CIRCLE)]<-0\ndtest$DEF_30_CNT_SOCIAL_CIRCLE[is.na(dtest$DEF_30_CNT_SOCIAL_CIRCLE)]<-0\ndtest$OBS_60_CNT_SOCIAL_CIRCLE[is.na(dtest$OBS_60_CNT_SOCIAL_CIRCLE)]<-0\ndtest$DEF_60_CNT_SOCIAL_CIRCLE[is.na(dtest$DEF_60_CNT_SOCIAL_CIRCLE)]<-0\ndtest$EXT_SOURCE_2[is.na(dtest$EXT_SOURCE_2)]<-0.558758 \ndtest$DAYS_BIRTH<-dtest[,'DAYS_BIRTH']/ -365\ndtest$DAYS_EMPLOYED<-dtest[,'DAYS_EMPLOYED']*-1\ndtest$DAYS_REGISTRATION<-dtest[,'DAYS_REGISTRATION']*-1\ndtest$DAYS_LAST_PHONE_CHANGE<-dtest[,'DAYS_LAST_PHONE_CHANGE']*-1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e2722c72e1a683f53428101e817422543df3f5ff"},"cell_type":"code","source":"dtest_numeric <- dtest %>%\n    select_if(is.numeric) # select remaining numeric columns\n# make sure that our dataframe is all numeric\nstr(dtest_numeric)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"76eb778a94ce73001a5e1e3605bb91f0834036a7"},"cell_type":"code","source":"# one-hot matrix for just the first few rows of the \"gender\" column\nmodel.matrix(~CODE_GENDER-1,head(dtest))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"28bf0b0637444483f4cf89bb9516841c89e0023a"},"cell_type":"code","source":"Car<-model.matrix(~FLAG_OWN_CAR-1,dtest)\nRealty<-model.matrix(~FLAG_OWN_REALTY-1,dtest)\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"778891e1a8647cc7590dbf53e87365084a639d3e"},"cell_type":"code","source":"\ndtest_numeric1 <- cbind(dtest_numeric,Car,Realty)\ndtest_matrix <- data.matrix(dtest_numeric1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c64b1607bc0dc974f33a070246cfaa2a53ac58bc"},"cell_type":"markdown","source":"Test Label\n"},{"metadata":{"trusted":true,"_uuid":"22115d0af8383d535f1bd06df657309d97b1170b"},"cell_type":"code","source":"test_labels<-train_labels[1:48744]\ntest_labels<-as.numeric(test_labels)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f58d3623c638d09819f3c79d7e352b710ee65b24"},"cell_type":"markdown","source":"Convert the cleaned dataframe to a dmatrix\n\n\nThe very final step is to convert our matrixes into dmatrix objects. This step isn't absolutely necessary, but it will help our model train move more quickly, and you'll need to to this if you ever want to train a model on multiple cores."},{"metadata":{"trusted":true,"_uuid":"a79a31e3640f14c78e0b20f3c8bb73aa7c6ba51e"},"cell_type":"code","source":"# put our testing & training data into two seperates Dmatrixs objects\ndtrain2 <- xgb.DMatrix(data =dtrain_matrix, label= train_labels)\ndtest2 <- xgb.DMatrix(data = dtest_matrix, label= test_labels)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"24cf7b531af64051c2104628eb8e00a8db3bf38c"},"cell_type":"markdown","source":"Training our model"},{"metadata":{"trusted":true,"_uuid":"a5d9bb4a10734b1d55921b878da9ebfa70d41a0e"},"cell_type":"code","source":"set.seed(1)\n# train a model using our training data\nmodel <- xgboost(data = dtrain2, # the data   \n                 nround = 2, # max number of boosting iterations\n                 objective = \"binary:logistic\")  # the objective function","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"42c9ea9db749878524c96123f62301b7cbc858ed"},"cell_type":"markdown","source":"We can see looking out the output of our model that for both the first and second rounds, we had the same error on the training data. This means that we didn't see an improvement in the second round of training.\n\nThat said, we're not really interested in how accurate we are on the training data. We're more interested in how accurate we are on the testing data, which our model hasn't seen before."},{"metadata":{"trusted":true,"_uuid":"8789f97b121d67fe6d0ffe3798f259a2a68677de"},"cell_type":"code","source":"# generate predictions for our held-out testing data\npred <- predict(model,dtest2)\n# get & print the classification error\nerr <- mean(as.numeric(pred > 0.5) != test_labels)\nprint(paste(\"test-error=\", err))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1d0d1b0cba94ae2cedd6b2987536d4941cbee5a9"},"cell_type":"markdown","source":"Now that we've got a basic model, we can try our hand at parameter tuning.\n\n"},{"metadata":{"_uuid":"e25ecfd6eb063fd658a4feee0b24e77761ace51f"},"cell_type":"markdown","source":"Tuning our model"},{"metadata":{"trusted":true,"_uuid":"14b390735360e5ea4fbd4e8ab6c99686820caf48"},"cell_type":"code","source":"set.seed(1)\n# train an xgboost model\nmodel_tuned <- xgboost(data = dtrain2, # the data           \n                 max.depth = 3, # the maximum depth of each decision tree\n                 nround = 2, # max number of boosting iterations\n                 objective = \"binary:logistic\") # the objective function\n# generate predictions for our held-out testing data\npred <- predict(model_tuned, dtest2)\n\n# get & print the classification error\nerr <- mean(as.numeric(pred > 0.5) != test_labels)\nprint(paste(\"test-error=\", err))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bd80a456081f693c55d6d6af97729afa4905a552"},"cell_type":"markdown","source":"Let's try re-training our model with these tweaks."},{"metadata":{"trusted":true,"_uuid":"2523732b97d1e54db021e7a12ff701c2e2a45730"},"cell_type":"code","source":"set.seed(1)\n# get the number of negative & positive cases in our data\nnegative_cases <- sum(train_labels == FALSE)\npostive_cases <- sum(train_labels == TRUE)\n\n# train a model using our training data\nmodel_tuned <- xgboost(data = dtrain2, # the data           \n                 max.depth = 3, # the maximum depth of each decision tree\n                 nround = 10, # number of boosting rounds\n                 early_stopping_rounds = 3, # if we dont see an improvement in this many rounds, stop\n                 objective = \"binary:logistic\", # the objective function\n                 scale_pos_weight = negative_cases/postive_cases) # control for imbalanced classes\n\n# generate predictions for our held-out testing data\npred <- predict(model_tuned, dtest2)\n\n# get & print the classification error\nerr <- mean(as.numeric(pred > 0.5) != test_labels)\nprint(paste(\"test-error=\", err))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d530c36e8191da326b9fa1a4cb807cd3d256c22"},"cell_type":"markdown","source":"There are a couple things to notice here."},{"metadata":{"_uuid":"74c7461bc5c95b07b40f812ff608ce5b6282f217"},"cell_type":"markdown","source":"Here, I'll set it to one, which is fairly high. (By default gamma is zero.)"},{"metadata":{"trusted":true,"_uuid":"29c3f30acb5dba61809d4af7e773822825df4902"},"cell_type":"code","source":"set.seed(1)\n# train a model using our training data\nmodel_tuned <- xgboost(data = dtrain2, # the data           \n                 max.depth = 3, # the maximum depth of each decision tree\n                 nround = 10, # number of boosting rounds\n                 early_stopping_rounds = 3, # if we dont see an improvement in this many rounds, stop\n                 objective = \"binary:logistic\", # the objective function\n                 scale_pos_weight = negative_cases/postive_cases, # control for imbalanced classes\n                 gamma = 1) # add a regularization term\n\n# generate predictions for our held-out testing data\npred <- predict(model_tuned, dtest2)\n\n# get & print the classification error\nerr <- mean(as.numeric(pred > 0.5) != test_labels)\nprint(paste(\"test-error=\", err))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"42b73914af8de44bc6bbf1df47473df9c3a8bbcb"},"cell_type":"markdown","source":"Adding a regularization terms makes our model more conservative, so it doesn't end up adding the models which were reducing our accuracy."},{"metadata":{"_uuid":"eb8c775117d88f9ee451bd712cffc8681244ece5"},"cell_type":"markdown","source":"Your turn!"},{"metadata":{"_uuid":"cbd6e558d3af5b86cdb3b8f096eaedb964f9ad02"},"cell_type":"markdown","source":"Examining our model\nSo far, we've:\n\ncleaned & prepared our data\ntrained our model\ntuned our model (not strictly necessary in this case, but generally it will help!)\nNow we can spend some time examing and interpreting our model. One of the really nice things about xgboost is that is has a lot of built-in functions to help us figure out why our model is making the distictions it's making.\n\nOne way that we can examine our model is by looking at a representation of the combination of all the decision trees in our model. Since all the trees have the same depth (remember that we set that with a parameter!) we can stack them all on top of one another and pick the things that show up most often in each node."},{"metadata":{"trusted":true,"_uuid":"852d6f19bb33f7bba8db9f17117b17bb5bc4b8cd"},"cell_type":"code","source":"# plot them features! what's contributing most to our model?\nxgb.plot.multi.trees(feature_names = names(dtrain_matrix), \n                     model = model)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b36266e8c999ee775c1660281384d244265f50ad"},"cell_type":"markdown","source":"we want a quick way to see which features are most important? We can do that using by creating and then plotting the importance matrix, like so:"},{"metadata":{"trusted":true,"_uuid":"6e1bd7b4b775925230c9ed5e9a59be0a4db9a444"},"cell_type":"code","source":"# get information on how important each feature is\nimportance_matrix <- xgb.importance(names(dtrain_matrix), model = model)\n\n# and plot it!\nxgb.plot.importance(importance_matrix)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"65a0fd8c64bb227a939c3da03c8a04e9e4d757b2"},"cell_type":"markdown","source":"We notice the most important variables for this model is EXT-SOURCE2"},{"metadata":{"_uuid":"b36e3f9d4854a6b35ee440ff6faeadb82f5365c1"},"cell_type":"markdown","source":"Conclusion\n\nNow we have seen the importance of tune a model using XGBoost and for very large data XGBOOST and logistic regression are faster to be completed to get results\n\nFor more details on XGBOOST for the novice you can watch Racheal Tatman tutorial"},{"metadata":{"trusted":true,"_uuid":"69f10a086e806085778a499e910b5ea2dad39651"},"cell_type":"markdown","source":"**If you liked this job, please click on I like it**"},{"metadata":{"trusted":true,"_uuid":"e58cac7ab6fc8610de2f49b74222a2a1831c0d26"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}