{"cells":[{"metadata":{"_uuid":"814d500196c954e37e68d42ad4e5e3cb7a74df08","_cell_guid":"900e0541-e9ed-4a9d-96f4-18478973fe6e"},"cell_type":"markdown","source":"**About the Analysis: **\n\nTalkingData , China’s largest independent Big Data service platform, covers over 70% of active mobile devices nationwide. They handle 3 billion clicks per day, of which 90% are potentially fraudulent. They want us to build an algorithm that predicts whether a user will download an app after clicking a mobile app ad. \n\n**About the Data:**\n\nThe data is provided as follow,\n\ntrain.csv - this is the training set(on which we are going to train and validate our model)\n\ntest.csv - the test set(We have to make predictions on this set, this will not have a target variable(obiviously))\n\nsampleSubmission.csv - a sample submission file in the correct format(the format in which we have to submit our predictions)\n\nApart from the above, we are also given with the train_sample data which is randomly-selected rows of training data, to inspect data before downloading full set and test_supplement, which is a larger test set. The official test data is a subset of this data."},{"metadata":{"_uuid":"88f3ef6e56451fae59599aa363fb16292abc33a0","_cell_guid":"fa941cba-484a-40ec-8daf-629a4169a6f5"},"cell_type":"markdown","source":"**Load Libraries & Data:**\n\n\nSo, let's get going!!! First we will have to know more about the data, and we can do that by doing some Exploratory Data Analysis(EDA) and we cannot do it without loading the required libraries and the data. So, let's do it right away!!! "},{"metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"_uuid":"524bab1248c5856ad98c058303f196c153aba354","trusted":true,"_cell_guid":"883fa21d-77cd-42c5-9e33-29a5ddd0dfdd"},"cell_type":"code","source":"library('readr') \nlibrary('data.table')\nlibrary('tidyr') \nlibrary('dplyr') \nlibrary('ggplot2') \nlibrary('ggthemes')\nlibrary('corrplot') \nlibrary('lubridate') \nlibrary('purrr') \nlibrary('cowplot')\nlibrary('graphics')\nlibrary('IRdisplay')\nlibrary('viridis')\nlibrary('plotly')\nlibrary('ROSE')\nlibrary('DMwR')\nlibrary('MASS')\nlibrary(car)\nlibrary(e1071)\nlibrary(caret)\nlibrary(caTools)\nlibrary(randomForest)\noptions(scipen = 999)\noptions(warn = -1)","execution_count":1,"outputs":[]},{"metadata":{"_uuid":"2f10707d60f2720d88b5a6071495c3b25dcfc530","trusted":true,"_cell_guid":"b02d6980-bb59-48f7-91ef-c2efe34d8347"},"cell_type":"code","source":"#Load the train data\ntrain <-fread('../input/train_sample.csv', stringsAsFactors = FALSE, data.table = FALSE)\n","execution_count":2,"outputs":[]},{"metadata":{"_uuid":"c1198c1e0268d802dec2b528cf29080d266f457a","trusted":true,"_cell_guid":"d0a36a60-1ab8-4ce3-877e-3f7ce55e43e5"},"cell_type":"code","source":"#load the test data\ntest <-fread('../input/test.csv', stringsAsFactors = FALSE, data.table = FALSE)","execution_count":3,"outputs":[]},{"metadata":{"_uuid":"8e0391aa4b3d5eba0c12654a8acd9d771849d11e","_cell_guid":"8c45fac0-9459-4e06-a54f-af9a108b3a8e"},"cell_type":"markdown","source":"Inspect the structure and the summary of the test and train data"},{"metadata":{"_uuid":"c4b462ad12d8bf4560b65b0f7c29f6470d16cb86","trusted":true,"_cell_guid":"c576cee4-2499-4668-84f4-522a932650a1"},"cell_type":"code","source":"str(train)","execution_count":4,"outputs":[]},{"metadata":{"_uuid":"90be506e95360457962053c1aacb2c2d5fc6de5e","trusted":true,"_cell_guid":"a8eac27c-6049-4431-aece-ee795e3079f0"},"cell_type":"code","source":"str(test)","execution_count":5,"outputs":[]},{"metadata":{"_uuid":"a583ac836c9c44f9cd2cb3ff156e634f11d43f34","_cell_guid":"3911c6d7-cc04-4f28-a0ee-4419b8875cf3"},"cell_type":"markdown","source":"Check for NAs in the train and test dataset"},{"metadata":{"_uuid":"c9048cf234e79e083266c2fa411aee97b7acad0d","trusted":true,"_cell_guid":"8c6a7376-d96f-423a-b0f9-e124c9f96a35"},"cell_type":"code","source":"colSums(is.na(train))","execution_count":6,"outputs":[]},{"metadata":{"_uuid":"3b9c63594260bdd6054645b894479494340345ee","trusted":true,"_cell_guid":"65405371-8a46-4eb1-9914-b16a69be6222"},"cell_type":"code","source":"colSums(is.na(test))","execution_count":7,"outputs":[]},{"metadata":{"_uuid":"72de4c566d0f7b809f8cb7bcb9dbd08c73269fd9","trusted":true,"_cell_guid":"43e5f417-f1b8-48ab-808f-f34a5012f234"},"cell_type":"code","source":"head(train)","execution_count":8,"outputs":[]},{"metadata":{"_uuid":"55737861109828e14f0d147d5e6342609fdc1f23","_cell_guid":"d974ddd1-b54c-44a6-815b-02ff8d120997"},"cell_type":"markdown","source":"**About the Data fields**\n\nEach row of the training data contains a click record, with the following features.\n\n**ip**: ip address of click.\n\n**app:** app id for marketing.\n\n**device:** device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n\n**os:** os version id of user mobile phone\n\n**channel:** channel id of mobile ad publisher\n\n**click_time:** timestamp of click (UTC)\n\n**attributed_time:** if user download the app for after clicking an ad, this is the time of the app download\n\n**is_attributed:** the target that is to be predicted, indicating the app was downloaded. This feature will be blank in the data if a person has not downloaded the app.\n\nSome of the features such as  ip, app, device, os, and channel are encoded.\n\n\nThe test data is similar, with the following differences:\n\n**click_id:** reference for making predictions\n\n**is_attributed:** not included"},{"metadata":{"_uuid":"38cc57612399de76be2ad5f2aed58beb1edeaa0e","trusted":true,"_cell_guid":"be3046ff-73f8-456a-8962-76957879d50c"},"cell_type":"code","source":"head(test)","execution_count":9,"outputs":[]},{"metadata":{"_uuid":"2c3d765e1c77e711224a2b2fd096ccfd5682b391","_cell_guid":"7442d71f-614d-4780-9238-1add67439f1b"},"cell_type":"markdown","source":"In the train and test data, we see timestamps as click_time and attributed_time(only in train data). This data can be split into different features such as year, month, weekday and hour of the day. So, we will be doing that now and extract some important features from the timestamp variables. "},{"metadata":{"_uuid":"6b57b56defe1b1dab3d60a9ceee6f86be89b963b","trusted":true,"_cell_guid":"13687b2d-5db1-4392-9254-fcc453954787"},"cell_type":"code","source":"train$click_time <-  parse_date_time(x = train$click_time,\n                orders = c(\"%Y-%m-%d %H:%M:%S\"),\n                locale = \"eng\")\n\ntrain$attributed_time <-  parse_date_time(x = train$attributed_time,\n                orders = c(\"%Y-%m-%d %H:%M:%S\"),\n                locale = \"eng\")\n\ntrain$click_year <- year(train$click_time)\ntrain$click_month <- month(train$click_time)\ntrain$click_day <- weekdays(train$click_time)\ntrain$click_hour <- hour(train$click_time)\n\nhead(train)","execution_count":10,"outputs":[]},{"metadata":{"_uuid":"fc8791cad798492712b0851e646a850122ed89db","_cell_guid":"1feb8546-daa5-4d6b-83cc-7348e9389c6f"},"cell_type":"markdown","source":"We know that all the variables above are categorical in nature, so we can convert them to factor type for better EDA."},{"metadata":{"_uuid":"65ea3d6906d988aece6719f8ccb17e02bc626c73","trusted":true,"_cell_guid":"e62189df-e1da-452e-900f-cf2e8533c362"},"cell_type":"code","source":"nms <- c(\"ip\", \"app\", \"os\", \"channel\", \"is_attributed\", \"click_year\", \"click_month\", \"click_day\", \"click_hour\") \ntrain1 <- train\ntrain1[nms] <- lapply(train1[nms], as.factor) ","execution_count":11,"outputs":[]},{"metadata":{"_uuid":"caba605a25d14e38dc998109f5dc990b8fcfed39","_cell_guid":"b87ee037-48ca-47bd-846f-2f34d3afa70f"},"cell_type":"markdown","source":"Univariate analysis of our features shows the top 10 ip addresses, os, channel and apps. We can also see that, despite the huge traffic in just few days only a fraction of people actually downloaded the app and most of the clicks are actually fraudulent clicks. This is the business problem, because it raises the cost of advertisement for Talking data."},{"metadata":{"_kg_hide-input":true,"_uuid":"d7c525f32d1ff502e28d3cbd3f25aba12a5d43dc","trusted":true,"_cell_guid":"b569b193-e4bf-4d52-8b0c-ba27c0d33f3c"},"cell_type":"code","source":"options(repr.plot.width=10, repr.plot.height=12)\n\np1 <- ggplotly(train1  %>% group_by(ip) %>% summarise(count = length(ip))%>%top_n(10, wt = count)%>%\nggplot(aes(x = reorder(ip, -count), y = count, fill = count )) + xlab(\"Top 10 IP addresses\")+ylab(\"\")+\n  geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5)+\n               scale_fill_viridis(option = \"magma\",direction = -1)+theme_bw()+\n                theme(legend.position=\"none\"), tooltip = c(\"y\"))\n\np2 <- ggplotly(train1  %>% group_by(app) %>% summarise(count = length(app))%>%top_n(10, wt = count)%>%\nggplot(aes(x = reorder(app, -count), y = count, fill = count )) + xlab(\"Top 10 App\")+ylab(\"\")+\n  geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5)+\n               scale_fill_viridis(option = \"magma\",direction = -1)+\n                  theme_bw()+theme(legend.position=\"none\"), tooltip = c(\"y\"))\n\np3 <- ggplotly(train1  %>% group_by(os) %>% summarise(count = length(os))%>%top_n(10, wt = count)%>%\nggplot(aes(x = reorder(os, -count), y = count , fill = count)) + xlab(\"Top 10 OS\")+ylab(\"\")+\n  geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5)+ \n               scale_fill_viridis(option = \"magma\",direction = -1)+theme_bw()+\n                theme(legend.position=\"none\"), tooltip = c(\"y\"))\n\n\np4 <- ggplotly(train1  %>% group_by(channel) %>% summarise(count = length(channel))%>%top_n(10, wt = count)%>%\nggplot(aes(x = reorder(channel, -count), y = count , fill = count)) + xlab(\"Top 10 Channel\")+ylab(\"\")+\n  geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5)+\n               scale_fill_viridis(option = \"magma\",direction = -1)+ theme_bw()+\n                theme(legend.position=\"none\"), tooltip = c(\"y\"))\n\n\np5 <- ggplotly(train1  %>% group_by(is_attributed) %>% summarise(Count = length(is_attributed))%>%top_n(10, wt = Count)%>%\nggplot(aes(x = reorder(is_attributed, Count), y = Count, fill = Count )) + xlab(\"is_attributed\")+ylab(\"\")+\n  geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5)+ \n                theme(legend.position=\"none\"), tooltip = c(\"y\"))\n\nhtmlwidgets::saveWidget(p1, \"p1.html\")\nhtmlwidgets::saveWidget(p2, \"p2.html\")\nhtmlwidgets::saveWidget(p3, \"p3.html\")\nhtmlwidgets::saveWidget(p4, \"p4.html\")\nhtmlwidgets::saveWidget(p5, \"p5.html\")\ndisplay_html('<iframe src=\"p1.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p2.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p3.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p4.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p5.html\" width=100% height=450></iframe>')\n\n\n#plot_grid(p1,p2,p3,p4,p5, labels = \"AUTO\", hjust = 0, vjust = 1,\n         # scale = c(1., 1., 0.9, 0.9))","execution_count":12,"outputs":[]},{"metadata":{"_uuid":"b2ffb5533f729356e3cd5ff0ae664641fa134c5f","_cell_guid":"dbfce798-72e1-43a4-8eb7-8c1826ac10e2"},"cell_type":"markdown","source":"Now let's see how are derived variables(the click_day and click_hour) perform. Since, the year(2017) and month(11) is same for all the data, we are not going to plot them.\n\nFrom the plots below, we can see that, most of the clicks happen on Wednesday, followed by Tuesday and Thursday. And the morning 4 am has the maximum number of clicks, which is quite puzzling."},{"metadata":{"_kg_hide-input":true,"_uuid":"6a894237dad8356ec1cab4141f37c606c9995a9d","trusted":true,"_cell_guid":"92f33b45-a603-41fd-9655-bfff8572eb91"},"cell_type":"code","source":"options(repr.plot.width=12, repr.plot.height=4)\np6<- ggplotly(train1  %>% group_by(click_day) %>% summarise(Count = length(click_day))%>%\nggplot(aes(x = reorder(click_day, -Count), y = Count, fill = click_day )) + xlab(\"click_day\")+ylab(\"\")+\n  geom_bar(stat = 'identity',aes(color = I('black')), size = 0.5)+ theme_bw()+\ntheme(legend.position=\"none\"), tooltip = c(\"Count\"))\n\np7 <- ggplotly(train1  %>% group_by(click_hour) %>% summarise(Count = length(click_hour))%>%\nggplot(aes(x = reorder(click_hour, -Count), y = Count, fill = Count)) +scale_fill_viridis(option = \"magma\",direction = -1)+xlab(\"click_hour\")+ylab(\"\")+\n  geom_bar(stat = 'identity',aes(color = I('black')), size = 0.5)+ theme_bw()+\ntheme(legend.position=\"none\"), tooltip = c(\"y\"))\n\nhtmlwidgets::saveWidget(p6, \"p6.html\")\nhtmlwidgets::saveWidget(p7, \"p7.html\")\n\ndisplay_html('<iframe src=\"p6.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p7.html\" width=100% height=450></iframe>')\n\n#plot_grid(p6,p7, labels = \"AUTO\", hjust = 0, vjust = 1,\n          #scale = c(1., 1., 0.9, 0.9))","execution_count":13,"outputs":[]},{"metadata":{"_uuid":"c4835f79e082ac5da99fd60d5e35a44c22125391","_cell_guid":"bf2582da-06a8-4370-b54b-e2aaedc1e939"},"cell_type":"markdown","source":"Now, lets filter out the ip addresses that have actually downloaded the app and try to analyse their behaiviour(like, which channel, os and app they are using and also their preferred day of the week and time of day.\n\nThe top 3 Apps are 19, 35 and 10\n\nThe top 3 OS are 19, 13, 24\n\nThe top 3 channels are 213, 113 and 21"},{"metadata":{"_kg_hide-input":true,"_uuid":"ea92ba194f0f8a0c1ffe7bfc4a9d617b9a3823db","trusted":true,"_cell_guid":"bd0ad36e-9863-4a16-b697-c7a37bdb3982"},"cell_type":"code","source":"options(repr.plot.width=10, repr.plot.height=6)\n\np8 <- ggplotly(train1 %>% filter(is_attributed == 1) %>% group_by(app) %>% summarise(Count = length(app)) %>% top_n(10, wt = Count)%>%\nggplot(aes(x = reorder(app, -Count), y = Count, fill = Count)) + geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5) \n               +scale_fill_viridis(option = \"magma\",direction = -1)+\n     theme_bw()+ xlab(\"Apps leading to download\") + ylab(\"\")+theme(legend.position=\"none\"), tooltip = c(\"app\", \"y\"))\n\np9 <- ggplotly(train1 %>% filter(is_attributed == 1) %>% group_by(os) %>% summarise(Count = length(os)) %>% top_n(10, wt = Count)%>%\nggplot(aes(x = reorder(os, -Count), y = Count, fill = Count)) + geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5) \n               +scale_fill_viridis(option = \"magma\",direction = -1)+\n     theme_bw()+ xlab(\"OS leading to download\") + ylab(\"\")+theme(legend.position=\"none\"), tooltip = c(\"os\", \"y\"))\n\np10 <- ggplotly(train1 %>% filter(is_attributed == 1) %>% group_by(channel) %>% summarise(Count = length(channel)) %>% top_n(10, wt = Count)%>%\nggplot(aes(x = reorder(channel, -Count), y = Count, fill = Count)) + geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5) \n                +scale_fill_viridis(option = \"magma\",direction = -1)+\n     theme_bw()+ xlab(\"Channels leading to download\") + ylab(\"\")+theme(legend.position=\"none\"), tooltip = c(\"channel\", \"y\"))\n\n\nhtmlwidgets::saveWidget(p8, \"p8.html\")\nhtmlwidgets::saveWidget(p9, \"p9.html\")\nhtmlwidgets::saveWidget(p10, \"p10.html\")\ndisplay_html('<iframe src=\"p8.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p9.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p10.html\" width=100% height=450></iframe>')\n#plot_grid(p8,p9,p10, labels = \"AUTO\", ncol = 1, align = 'v')","execution_count":14,"outputs":[]},{"metadata":{"_uuid":"50d7b3bba326e048f37a17741d3ee1a4c7f9e369","_cell_guid":"37eb9b97-fd68-4d2c-9b6f-1ba28d5366e2"},"cell_type":"markdown","source":"Now let's find out the time and day pattern of those who have attributed\n\nWe can see from the plots below that, the wee hours of the day i.e. hours after midnight(3 am, 2 am etc.) get more successful downloads as compared to the afternoon and evening hours. 4 am which is the most clicked hour(as per the plot above), doesn't lead to successful download. "},{"metadata":{"_kg_hide-input":true,"_uuid":"c9e65693d8603c212dec562715a59a2e66ac53f4","trusted":true,"_cell_guid":"a5779b46-469c-4e3b-9a39-a81ea7f9fb1d"},"cell_type":"code","source":"options(repr.plot.width=10, repr.plot.height=6)\n\np11 <- ggplotly(train1 %>% filter(is_attributed == 1) %>% group_by(click_hour) %>% summarise(Count = length(click_hour)) %>% \nggplot(aes(x = reorder(click_hour, -Count), y = Count, fill = Count)) + geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5) + \n                scale_fill_viridis(option = \"magma\",direction = -1)+\n     theme_bw()+ xlab(\"Downloaded hours\") + ylab(\"\")+theme(legend.position=\"none\"), tooltip = c(\"click_hour\", \"y\"))\n\np12 <- ggplotly(train1 %>% filter(is_attributed == 1) %>% group_by(click_day) %>% summarise(Count = length(click_day)) %>% \nggplot(aes(x = reorder(click_day, -Count), y = Count, fill = click_day)) + geom_bar(stat = 'identity', aes(color = I('black')), size = 0.5) + \n     xlab(\"click_day\")+ geom_bar(stat = 'identity',aes(color = I('black')), size = 0.5)+ theme_bw()+\n        theme(legend.position=\"none\"), tooltip = c(\"Count\"))\n\nhtmlwidgets::saveWidget(p11, \"p11.html\")\nhtmlwidgets::saveWidget(p12, \"p12.html\")\ndisplay_html('<iframe src=\"p11.html\" width=100% height=450></iframe>')\ndisplay_html('<iframe src=\"p12.html\" width=100% height=450></iframe>')\n#plot_grid(p11,p12, labels = \"AUTO\", ncol = 1, align = 'v')","execution_count":15,"outputs":[]},{"metadata":{"_uuid":"c4b89f6ccea5873d8ca9c0ca9d553a0aa34f7310","_cell_guid":"5163542c-71f9-494b-98cf-5a9c8a1d1d9c"},"cell_type":"markdown","source":"Time series of number of clicks during the four days of the data, shows that the maximum clicks happen in all four days are during 4 am, 1 pm and 2 pm and minimum clicks happend during 8 pm in a day. (All timing as per UTC tz). The data for monday and thursday is not for the full day."},{"metadata":{"_kg_hide-input":true,"_uuid":"3f3d88837bae7d06b771398be55ae12649cfa394","trusted":true,"_cell_guid":"04394943-865e-440e-b7a4-a2e2489e69e6"},"cell_type":"code","source":"options(repr.plot.width=8, repr.plot.height=6)\n\np13 <- ggplotly(train %>% group_by(click_day, click_hour) %>% summarise(count = length(click_hour))%>%\nggplot(aes(x = click_hour, y = count, shape = click_day, colour = click_day)) +\n    geom_point() +geom_line() +ylab(\"\")+\n    theme_few()+\n    facet_grid(facets = click_day ~ .),tooltip = c(\"click_hour\",\"count\" ))\n\nhtmlwidgets::saveWidget(p13, \"p13.html\")\ndisplay_html('<iframe src=\"p13.html\" width=100% height=450></iframe>')","execution_count":16,"outputs":[]},{"metadata":{"_uuid":"74a1e7bd299f160c50df19e0c0ac3c9f763ff360","_cell_guid":"aecae5c5-b780-41f0-ae24-c818d1fe2e45"},"cell_type":"markdown","source":"**Correlation between the variables:**\n\nAlthough the variables ip, app, os, channel etc. are categorical in nature, but let's try to see if there exists any correlation between them using a correlation matrix(they are in 'int' format in the train data, and we will try using that). \n\nSo from the correlation matrix below, we can see that our target variable is_attributed slightly positively correlated with app and ip.  App has a positive correlation with os and device, while device has a strong positive correlation with os. There are no negative correlations though."},{"metadata":{"_kg_hide-input":true,"_uuid":"81cb9ac42e76fbff48d33533b4e0e9f163f474d9","trusted":true,"_cell_guid":"a68f2cec-919d-492a-a218-6cb5c30d0c15"},"cell_type":"code","source":"options(repr.plot.width=6, repr.plot.height=6)\nlibrary(ggcorrplot)\ntrain_cor <- round(cor(train[,c(1:5,8,12)]), 1)\n#p14 <- ggplotly(\nggcorrplot(train_cor,  title = \"Correlation\")+theme(plot.title = element_text(hjust = 0.5))\n#, tooltip = c(\"y\", \"value\"))\n\n#htmlwidgets::saveWidget(p14, \"p14.html\")\n#display_html('<iframe src=\"p14.html\" width=100% height=450></iframe>')","execution_count":17,"outputs":[]},{"metadata":{"_uuid":"7d1704140368668e9467d056320d22f0bc9025e3","_cell_guid":"56f66c36-ee64-4df3-9726-f196fcaec96c"},"cell_type":"markdown","source":"**Modelling:**\n\nNow, comes the Modelling part, for which we need to prepare our training dataset, on which we will train our model. We did our EDA on the train_sample data provided to us, and we will try and do the modelling in the same dataset. But before that we have to treat this highly imbalanced data. From the EDA above we came to know that, out of 100000 random observations, only 0.2% of people have actually dowloaded the app. So, the number of observation belonging to class 1 is significantly low as compared to class 0, and this will lead to biased prediction by our machine learning model. "},{"metadata":{"_uuid":"2956d0fd75531036008ee6a66bbe6413b834aaae","trusted":true,"_cell_guid":"91ffff73-7a8a-4f15-a79b-cfecd759fa2c"},"cell_type":"code","source":"#Let's prepare our train data accordingly, \n#by removing the unwanted features such as month , year, click_time(from which we have already extracted few features)\n\ntrain <- train[, -c(6,7,9,10)]\n\n#We will now convert our target variable  and click day to factor type\ntrain$is_attributed <- as.factor(train$is_attributed)\ntrain$click_day <- as.factor(train$click_day)\n\n\ntable(train$is_attributed)\n","execution_count":18,"outputs":[]},{"metadata":{"_uuid":"1bce0c3cbb598573b6196a049e8944ecd9ba3973","trusted":true,"_cell_guid":"1e7a0287-88e5-4333-8099-9f91f42616b8"},"cell_type":"code","source":"# splitting the data between train and validation sets in 70:30 \nset.seed(100)\n\nindices = sample.split(train$is_attributed, SplitRatio = 0.7)\n\ntrainf = train[indices,]\n\nvalidation = train[!(indices),]","execution_count":19,"outputs":[]},{"metadata":{"_uuid":"6b66e07e6b1b3ddfa10bf21164e230ac71b397dc","_cell_guid":"2a226136-58f6-4e0e-a4ac-634f413d7e15"},"cell_type":"markdown","source":"I am going to try two balancing techniques smote and rose and see which one works better with respect to auc."},{"metadata":{"_uuid":"b2faa7a1e65468e804a047b695a799ceb7b7c048","trusted":true,"_cell_guid":"17631946-3914-4bce-bf89-89be7e24b437"},"cell_type":"code","source":"prop.table(table(trainf$is_attributed))\nprop.table(table(validation$is_attributed))","execution_count":20,"outputs":[]},{"metadata":{"_uuid":"5f4ffaef0b53a8d17a66701a242c8e8a03c525d0","trusted":true,"_cell_guid":"23299625-7b8d-45f5-afd2-b683ddde9d33"},"cell_type":"code","source":"set.seed(1234)\n#Data balacing with ROSE\ndata.rose <- ROSE(is_attributed ~ ., data = trainf, seed = 1)$data\ntable(data.rose$is_attributed)\n\n\n#Data balancing with SMOTE\ndata.smote <- SMOTE(is_attributed ~ ., data = trainf)\ntable(data.smote$is_attributed)","execution_count":21,"outputs":[]},{"metadata":{"_uuid":"9f39d37b53ac0962e201a3d033eef348dc5e97da","_cell_guid":"5a88620f-b9f6-4086-9947-558d3607f5b5"},"cell_type":"markdown","source":"Now, i will use the basic random forest to train our two balanced train dataset arrived after data balancing using ROSE and SMOTE"},{"metadata":{"_uuid":"73244fc8b9b6cef9d4266833cb83093d88d16847","trusted":true,"_cell_guid":"f270a839-8b87-48d4-ae9b-693657b3d039"},"cell_type":"code","source":"rose.rf <- randomForest(is_attributed ~ ., data=data.rose, proximity=FALSE, importance=TRUE,\n                        ntree=1000, mtry=3)\n\nsmote.rf <- randomForest(is_attributed ~ ., data=data.smote, proximity=FALSE, importance=TRUE,\n                        ntree=1000, mtry=3)\n\nrose.rf\n\nsmote.rf","execution_count":22,"outputs":[]},{"metadata":{"_uuid":"5d52e75d757480225b6313c3156fa2acfb2b0fc2","trusted":true,"_cell_guid":"b4160ed8-442d-4990-87f1-6b1ad41d869d"},"cell_type":"code","source":"options(repr.plot.width=6, repr.plot.height=4)\n#Prediction using the train data that has been balanced using ROSE technique\n\ntestPred.rose <- predict(rose.rf, newdata=validation[,-6])\ntable(testPred.rose, validation$is_attributed)\nroc.curve(validation$is_attributed, testPred.rose)","execution_count":23,"outputs":[]},{"metadata":{"_uuid":"5b8baec692a7f6a228ff63723f32f61d0d387c47","trusted":true,"_cell_guid":"a3a5cf3f-ce17-4e5b-9fdc-c6ca1072d031"},"cell_type":"code","source":"#Prediction using the train data that has been balanced using SMOTE technique\n\ntestPred.smote <- predict(smote.rf, newdata=validation[,-6])\ntable(testPred.smote, validation$is_attributed)\nroc.curve(validation$is_attributed, testPred.smote)","execution_count":24,"outputs":[]},{"metadata":{"_uuid":"94e419048b3b3e5c1e840476c4d87d0c5e7a6cd6","_cell_guid":"5cd69b79-5ffa-4935-871c-ca63207e6c3e"},"cell_type":"markdown","source":"So we got way better accuracy by using the smote method for data balancing as compared to rose. So, we will go ahead with smote method for our further prediction and parameter tuning."},{"metadata":{"_uuid":"f9218760f83f1d597c73de5fa8727ae1331641e7","_cell_guid":"9350b7b2-bc7e-4071-87d3-1b1c3622004c"},"cell_type":"markdown","source":"Now, let's check the variable importance by making a variable importance plot arrived from our random forest model"},{"metadata":{"_uuid":"b942c48197d7f241972737449cfa75cbbc0f84b8","trusted":true,"_cell_guid":"eab8755a-e670-4e2e-93a6-173fcdeb3f28"},"cell_type":"code","source":"smote.rf","execution_count":25,"outputs":[]},{"metadata":{"_uuid":"f3bdd1470250a5970905b14f74577da63bccd0e5","trusted":true,"_cell_guid":"dcbdb286-9b15-4edc-80f2-c4e15baa0f67"},"cell_type":"code","source":"varImpPlot(smote.rf)","execution_count":26,"outputs":[]},{"metadata":{"_uuid":"0669523e3dcbdd531c91d5848ab2dc66f7a7401f","_cell_guid":"8c5952dd-a784-4ef0-acec-5649fda51dd3"},"cell_type":"markdown","source":"So we have two types of importance measures as shown above. The Mean decrease accuracy shows how worse the model performs without each variable. The Mean decrease Gini  measures how pure the nodes are at the end of the tree. The mean decrease in Gini coefficient is a measure of how each variable contributes to the homogeneity of the nodes( this is very important and we should try to split the nodes such that the resulting nodes are as homogenous as possible). \n\nFrom the above plots we know the important variables which are app, device and ip, also we know the not so important variable and in this case it is click day."},{"metadata":{"_uuid":"8ab98a201ee5644067e172d0cc1718970ee3a438","_cell_guid":"310a82b2-7ecb-4458-8c9f-1987ba89ca33"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"41f853c9276e2c1a4f85c41795ebef7588fb7c51","trusted":true,"_cell_guid":"24cc8de4-cdf8-4f96-8ce2-8b0ae948e82c"},"cell_type":"code","source":"set.seed(1234)\nbest.rf <- randomForest(is_attributed ~ ., data=data.smote, proximity=FALSE, importance=TRUE,\n                        ntree=500, mtry=4)\n\ntestPred <- predict(best.rf, newdata=validation[,-6])\ntable(testPred, validation$is_attributed)\nroc.curve(validation$is_attributed, testPred, col = 'green')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}