{"cells":[{"metadata":{"_uuid":"26b7c9d72380551ebdc925f8e675533f4205fd9f","_cell_guid":"d099e485-313b-4033-9350-3c68a183119a"},"cell_type":"markdown","source":"![](http://res.cloudinary.com/dqlrb47d0/image/upload/v1524755006/sell.jpg)"},{"metadata":{"_uuid":"6f4b78f8f5039099261477d8e6a6c1a4537e2a9c","_cell_guid":"6e17da21-9039-4be7-bdc4-72c7e45d5c45"},"cell_type":"markdown","source":"* <a href='#1'>1. Introduction</a>  \n* <a href='#2'>2. Loading Libraries</a>\n* <a href='#3'>3. Glimpse of Data</a>\n*  <a href='#4'>4. Check for unique values</a>\n*  <a href='#5'>5. Overlapp of test and train data</a>\n* <a href='#6'>6. Histogram of the target Variable </a>\n* <a href='#7'>7. Quantile distribution of target Variable</a>\n* <a href='#8'>8.Most Popular parent category  (count and quantile distribution pattern)</a>\n *      <a href='#81'>8.1 Inference from the plot</a> \n* <a href='#9'>9. Analysis of Missing data patterns</a>\n> *      <a href='#91'>9.1 Patterns with and without param_1</a>\n> *       <a href='#92'>9.2 Patterns with and without param_2</a>\n> *        <a href='#93'>9.3 Patterns with and without param_3</a>\n> *         <a href='#94'>9.4 Patterns with and without description</a>\n> *        <a href='#95'>9.5 Patterns with and without price</a>\n> *          <a href='#96'>9.5 Patterns with and without image</a>\n* <a href='#10'>10.Most Popular Region  (count and quantile distribution pattern)</a>\n* <a href='#11'>11. Corretaltion between text length and  Deal Probablity</a> \n* <a href='#12'>12. Binning the Target Variable</a>\n* <a href='#13'>13. Zero Probablities by the sequence of user attempts with in a category </a>\n* <a href='#14'>14. Zero Probablities by the sequence of user attempts  </a>\n* <a href='#15'>15. Alluvial plot betweeen user type and deal Probablity  </a>\n* <a href='#16'>16.Usage of English letters   </a>\n* <a href='#17'>17.Venn Diagram of top 100 words from title under each category    </a>\n* <a href='#18'>18.Venn Diagram of top 100 words from description under each category    </a>\n* <a id='#19'>19.Model Building    </a>\n\n\n"},{"metadata":{"_uuid":"4dbf3dc33746f420c2c142a3fb106f45268dfca0","_cell_guid":"8588d21c-876f-47bc-bb32-0f2f9bfc5023"},"cell_type":"markdown","source":"**<a id='1'>1. Introduction</a>**\n\" Selling used goods online, a combination of tiny, nuanced details in a product description can make a big difference in drumming up interest \"\n\nI am a beginner hoping to know some interesting story thats hidden in the data."},{"metadata":{"_uuid":"44c930f1a02082472c9e84b6bc21e9890bb17258","_cell_guid":"07f622ee-ff08-48c8-81c0-ba19b52a11f4"},"cell_type":"markdown","source":"**<a id='2'>2. Loading Libraries</a>**\n"},{"metadata":{"scrolled":true,"_kg_hide-output":true,"_execution_state":"idle","_kg_hide-input":true,"_uuid":"50917d64ce39e3807fa85b55a7be836b1f9391c0","trusted":false,"_cell_guid":"86538ff7-c5ce-41ee-8451-83fe4068457b"},"cell_type":"code","source":"library(ggplot2) # Data visualization\nlibrary(knitr)\nlibrary(readr) # CSV file I/O, e.g. the read_csv function\nlibrary(data.table)\nlibrary(dplyr)\nlibrary(tidytext)\nlibrary(stringr)\nlibrary(quanteda)\nlibrary(knitr)\nlibrary(reshape2)\nlibrary(base)\nlibrary(gridExtra)\nlibrary(tm)\nlibrary(wordcloud) \nlibrary(SnowballC)\nlibrary(text2vec)\nlibrary(ggalluvial)\nlibrary(caret)\nlibrary(stringr)\nlibrary(tokenizers)\nlibrary(RAM)\nlist.files(\"../input\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"704420f93e302ac6709f2a057bde008a2fd1c869","_cell_guid":"ca43bae3-4760-4d82-a330-96f0fbeab2cb"},"cell_type":"markdown","source":"Lets start with the training and test data  "},{"metadata":{"_kg_hide-output":true,"_uuid":"198b44d031986a32da0188a140d1e0986937b060","trusted":false,"_cell_guid":"ce472826-eaf4-4799-9970-0a708b55f6ce"},"cell_type":"code","source":"\ntrain = read_csv(\"../input/avito-demand-prediction/train.csv\",locale = locale(encoding = stringi::stri_enc_get()))\nindex_train<-nrow(train)\ntest=read_csv(\"../input/avito-demand-prediction/test.csv\",locale = locale(encoding = stringi::stri_enc_get()))\nnames(train)\ntest$deal_probability<-NA\n\ncateg <- c('Личные вещи' = \"Personal_things\", 'Для дома и дачи' = \"For_home_and_cottage\", 'Бытовая электроника' ='Consumer_electronics', \n               'Недвижимость' = 'The_property', 'Хобби и отдых' = 'Hobbies_and_Recreation', 'Транспорт' = 'Transport', 'Услуги'='Services',\n               'Животные'='Animals', 'Для бизнеса'= 'For_business')\nconversion <- c('Свердловская область'= 'Sverdlovsk oblast','Самарская область'= 'Samara oblast','Ростовская область'= 'Rostov oblast','Татарстан' = 'Tatarstan',\n'Волгоградская область'= 'Volgograd oblast','Нижегородская область' ='Nizhny Novgorod oblast','Пермский край'= 'Perm Krai',\n'Оренбургская область'= 'Orenburg oblast','Ханты-Мансийский АО'= 'Khanty-Mansi Autonomous Okrug','Тюменская область'= 'Tyumen oblast',\n'Башкортостан'= 'Bashkortostan','Краснодарский край'= 'Krasnodar Krai','Новосибирская область'= 'Novosibirsk oblast','Омская область'= 'Omsk oblast',\n'Белгородская область'= 'Belgorod oblast','Челябинская область'= 'Chelyabinsk oblast','Воронежская область'= 'Voronezh oblast','Кемеровская область'= 'Kemerovo oblast',\n'Саратовская область'= 'Saratov oblast','Владимирская область'= 'Vladimir oblast','Калининградская область'= 'Kaliningrad oblast','Красноярский край'= 'Krasnoyarsk Krai',\n'Ярославская область'= 'Yaroslavl oblast','Удмуртия'= 'Udmurtia','Алтайский край'= 'Altai Krai','Иркутская область'= 'Irkutsk oblast','Ставропольский край'= 'Stavropol Krai',\n'Тульская область'= 'Tula oblast')\ntrain$parent_category_name<-(plyr::revalue(train$parent_category_name, categ))\ntrain$region<-(plyr::revalue(train$region, conversion))\ntrain$activation_date<-as.Date(train$activation_date)\ntest$parent_category_name<-(plyr::revalue(test$parent_category_name, categ))\ntest$region<-(plyr::revalue(test$region, conversion))\ntest$activation_date<-as.Date(test$activation_date)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f33623fa2dc29c82d249c8453f4d5845fbdab16d","_cell_guid":"0a076f67-3585-42ab-8656-d6de322fa657"},"cell_type":"markdown","source":"**<a id='3'>3. Glimpse of Data</a>**\nThe train data has the item ids user ids and all other attributes of the item. Just to remember,  the user is selling the product online. And the only difference between the test and the train data is the deal probablity. Lets dig deeper in the data to see the unique count of all different categories that are present in the data. My guess is that the item id would be unique and just that lets check. "},{"metadata":{"_uuid":"fb2fd4391032cdd12e69aac8b475bdb5689fbe8b","trusted":false,"_cell_guid":"858ff158-0cbf-44f0-99cc-06ab4203e4da"},"cell_type":"code","source":"glimpse(train)\nglimpse(test)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fabef67c00db3f7af905e569a1e86f257c69e11f","_cell_guid":"ddeff71e-96c1-467a-bafd-c605cd932281"},"cell_type":"markdown","source":"**<a id='4'>4. Check for unique values</a>**\nThe item id is unique and not the user id. So my assumption is right. The items are grouped in 47 categories and its a good point to start our analysis."},{"metadata":{"_kg_hide-input":true,"_uuid":"95ece71fcc7b4c1bf8cbbe2edc0d82934ece1113","trusted":false,"_cell_guid":"df316baa-0547-4da0-9e6f-9dd26c15dccf"},"cell_type":"code","source":"print(paste(\"Total no of rows in train data\", nrow(train)))  \nt(lapply(train, function (x) length(unique(x))))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f87fbb6917e95c26dd002b50620c1f24586e3e96","_cell_guid":"d8ee9a60-484f-42c1-9abe-b3d2f10f89a9"},"cell_type":"markdown","source":"**<a id='5'>5. Overlapp of test and train data</a>**\nA few user ids are common to both test and train whereas the item ids are found to be unique."},{"metadata":{"_uuid":"7357a1afd66016d2fc82350dd9c72619624cbff1","trusted":false,"_cell_guid":"e205468f-04b7-48ea-a5db-c5c6ab0242f1"},"cell_type":"code","source":"lista <- unique(test$user_id)\nlistb <- unique(train$user_id)\ngroup.venn(list(Test=lista, Train=listb), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1d242a88907361686d2a8aa615a4f1a5509d1a17","_cell_guid":"6a86950f-fd5e-499b-852c-8ed5e780ae51"},"cell_type":"markdown","source":"Overlap of test and train item category_name-\nWe can see that there are no new categories in the test set. Its a complete overlapp"},{"metadata":{"_uuid":"dec4a58a0fb96acbf20d47b52c1181886981e1ee","_cell_guid":"5160d719-c5cc-4b15-a487-8516b24b26b0"},"cell_type":"markdown","source":"**<a id='6'>6. Histogram of the target Variable </a>**\n\nOn checking the distribution we can infer a few things:\nGetting a deal probablity of 1 is going to be extremely difficult but 0 is most likely. So it shows how difficult it is for an user to sell a used product onl"},{"metadata":{"_uuid":"7c8468163cc5fbfcf9409c55af525c0649dd96ee","trusted":false,"_cell_guid":"fc0191ef-7434-4c8d-805b-e10e06a96d02"},"cell_type":"code","source":"train%>%ggplot(aes(x=train$deal_probability, fill=as.factor(parent_category_name)))+geom_histogram()+labs(x = 'Deal Probablity', \n       y = 'Count', \n       title = 'Distribution of deal Probability') ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"db427a50dba4c139b1b570bfcd67e4c6a7f8cfe3","_cell_guid":"9347e132-87df-4563-b4d2-19822d07c6b0"},"cell_type":"markdown","source":"**<a id='7'>7. Quantile distribution of target Variable </a>**\nOn looking at the Quantile distribution we can infer the following points;\n* More than 50% of the deals has a zero probablity. So not everyone wins in selling a used good\n* approx between 50% to 75% (ie 25% of the deals should have a probablity between 0.15 to 0)\n* less than 1% of the deals has a probablity  of 1. (But I am just curious to know how and why)\n"},{"metadata":{"_uuid":"57bc4b90464f2050ab39c7e70fa1a35a81399834","trusted":false,"_cell_guid":"d4afe655-2b4d-4d04-bcf7-05eece86a0a7"},"cell_type":"code","source":"probs = c(0.05, 0.1, 0.25, 0.5, 0.75, 0.8, 0.85, 0.9, 0.95, 0.97, 0.99, 1)\nt(quantile(train$deal_probability, probs = probs))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"118a5325301ea148f3f59983f2dd50fa9a4e78d6","_cell_guid":"5bf43043-13bb-4e61-845a-42fc341cdf3d"},"cell_type":"markdown","source":"**<a id='8'>8.Most Popular parent category  (count and quantile distribution pattern)</a>**"},{"metadata":{"scrolled":true,"_kg_hide-input":true,"_uuid":"b20fa010dafa694333a80371300b686498eb5226","trusted":false,"_cell_guid":"8cc051e4-ae54-43cb-9401-9d17829c11d0"},"cell_type":"code","source":"\nlength(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\n\nd1<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%mutate(parent_category_name=factor(parent_category_name, levels=Parent_cat_name))%>%  ggplot(aes(x=parent_category_name, y=count))+geom_bar(stat=\"identity\")+\ngeom_text(aes(x = parent_category_name, y = 1, label = count),\n            hjust=0, vjust=.5, size = 4, colour = 'black',\n            fontface = 'bold') +\ncoord_flip()+theme_bw()+theme(legend.position = 'bottom')\n\nd2<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'Percentile distribution by Parent category')+\ntheme_bw()+theme(legend.position = 'bottom')\n\ngrid.arrange(d1, d2, ncol = 1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a9a522d4e737d304a3af2723c58d575e820cf44","_cell_guid":"f14b442a-634b-4f05-9fa9-9fecb0a027eb"},"cell_type":"markdown","source":"**<a id='81'>8.1 Inference from the plot</a>**\n\n1. As we can see Personal things is the most popular category followed by home_cottage category and others.\n\n\n2.On looking at the quantile distribution for services category we can see that proporationately a larger number of items are having a higher deal probablity.\n\n\n3.Personal things, Property and Business category is on the lower side of the deal probablity. More than 90% of the items has a deal probablity of less that 0.25. "},{"metadata":{"_uuid":"2f0f4cf9f31931d62a3b7334e76e36fe331fc45b","_cell_guid":"fe04cca5-a64a-4b48-9961-7fb6d27f525a"},"cell_type":"markdown","source":"**<a id='9'>9. Analysis of Missing data patterns</a>**"},{"metadata":{"_uuid":"0d9e35bd0effd3c383d2b994b8554d1c36e5113e","trusted":false,"_cell_guid":"b5632231-017c-4d6c-afaf-3767c9fedf11"},"cell_type":"code","source":"t(lapply(train, function(x) sum(is.na(x))))\nhead(train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9aa3b726793cf411b30ea38fdd494f0b0af56307","_cell_guid":"5a6c120a-dbb9-4c4d-97e7-9fac43978455"},"cell_type":"markdown","source":"**Influence of missing Values**\n\nI guess that for this business case the missing data will provide us a lots of information.\nThe reason behind is if i post an add for a used item people wont buy if I miss out certain details or if i dont post the images. Lets take time in analyzing all the missing attributes one by one."},{"metadata":{"_uuid":"7304a4e7a8f7aca4bc223c6aa4ee344e5175c279","_cell_guid":"637ef758-d7ec-4837-b60a-65e7966af467"},"cell_type":"markdown","source":"** <a id='9.1'>91 Patterns with and without param_1</a>**"},{"metadata":{"_kg_hide-input":true,"_uuid":"ddbcb9afadb82cc57e250af13d30563165284612","trusted":false,"_cell_guid":"a754ca60-0cc1-43fd-9038-b209a54b68b9"},"cell_type":"code","source":"length(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&is.na(train$param_1)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\n\n\nd2<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'without param1')+\ntheme_bw()+theme(legend.position = 'bottom')\n\n\nlength(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&!is.na(train$param_1)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\nd1<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'With param1')+\ntheme_bw()+theme(legend.position = 'bottom')\n\ngrid.arrange(d1, d2, ncol = 1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"89d2f9ab19dba33264a2641b4a5554f8d7c59486","_cell_guid":"19234f62-994f-47ee-b546-017c509c1dd0"},"cell_type":"markdown","source":"** <a id='92'>9.2 Patterns with and without param_2</a>**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"1dd9f5118b228781ea0f62456883647e1704f6da","trusted":false,"_cell_guid":"7e00eb91-8b99-4d3a-a77a-3e7b1deff7b5"},"cell_type":"code","source":"length(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&is.na(train$param_2)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\n\n\nd2<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'without param2')+\ntheme_bw()+theme(legend.position = 'bottom')\n\n\nlength(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&!is.na(train$param_2)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\nd1<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'With param2')+\ntheme_bw()+theme(legend.position = 'bottom')\n\ngrid.arrange(d1, d2, ncol = 1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"14aa1e4e8d6cc7e61e22f782be43a7f033ffc8f0","_cell_guid":"967094f0-437c-4e75-be7f-b30872479726"},"cell_type":"markdown","source":"** <a id='93'>9.3 Patterns with and without param_3</a>**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"fc34ae17cd30cdd1383a17dd94613d14f83f983a","trusted":false,"_cell_guid":"f11d0083-220a-41e9-928c-ab29c71f09b8"},"cell_type":"code","source":"length(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&is.na(train$param_3)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\n\n\nd2<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'without param3')+\ntheme_bw()+theme(legend.position = 'bottom')\n\n\nlength(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&!is.na(train$param_3)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\nd1<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'With param3')+\ntheme_bw()+theme(legend.position = 'bottom')\n\ngrid.arrange(d1, d2, ncol = 1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3185a5e8b00200dacb96567e52abeff2768d14af","_cell_guid":"68a75940-d01e-4c56-8c6f-24f7d1c99dd3"},"cell_type":"markdown","source":"** <a id='94'>9.4 Patterns with and without item description</a>**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"cd5db194250cf7d3d320ff45b29868fa80b10eff","trusted":false,"_cell_guid":"c4dfc2d9-2190-4190-9aac-81017119b659"},"cell_type":"code","source":"length(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&is.na(train$description)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\n\n\nd2<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'without description')+\ntheme_bw()+theme(legend.position = 'bottom')\n\n\nlength(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&!is.na(train$description)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\nd1<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'With description')+\ntheme_bw()+theme(legend.position = 'bottom')\n\ngrid.arrange(d1, d2, ncol = 1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f6f1aece54e1e06bf1825b8c41870df19eca524d","_cell_guid":"9e504701-a0e7-4cad-91cf-0840efafd8c0"},"cell_type":"markdown","source":" **<a id='95'>9.5 Patterns with and without price</a>**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"a4f6d9716f4dcd6d10f072500eb622e5856aed99","trusted":false,"_cell_guid":"009236ed-cd51-498d-b6e4-8d600ed14db4"},"cell_type":"code","source":"length(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&is.na(train$price)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\n\n\nd2<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'without price')+\ntheme_bw()+theme(legend.position = 'bottom')\n\n\nlength(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&!is.na(train$price)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\nd1<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'With price')+\ntheme_bw()+theme(legend.position = 'bottom')\n\ngrid.arrange(d1, d2, ncol = 1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ddf7f65630ab86901a959b1541dfc86f98131c6b","_cell_guid":"ef3189a0-595b-4be2-8563-94d00638155f"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"4478b9c1aa7d23c3ddf4e5c8e8d3afed17ab48bf","_cell_guid":"3f0b9f27-4f7e-4e84-b215-80f7bcd33e05"},"cell_type":"markdown","source":"** <a id='96'>9.6 Patterns with and without image</a>**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"9bb180289ad349e54e219d149df8245a3ef2e607","trusted":false,"_cell_guid":"620566a9-416e-44d8-9d4f-f10fcf835ec0"},"cell_type":"code","source":"length(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&is.na(train$image)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\n\n\nd2<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'without image')+\ntheme_bw()+theme(legend.position = 'bottom')\n\n\nlength(probs)\nD<-matrix(, nrow = length(unique(train$parent_category_name)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(parent_category_name)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nParent_cat_name<-category_count$parent_category_name\nfor (i in (1:length(Parent_cat_name))){\n\nD[i,]<- quantile(train$deal_probability[train$parent_category_name==Parent_cat_name[i]&!is.na(train$image)], probs = probs)  }\n\nD<-as.data.frame(D)\ncolnames(D)<-probs\nD$parent_category_name<-Parent_cat_name\n\nx1=melt(D, id=\"parent_category_name\")\nd1<-x1%>%ggplot(aes(y =factor(parent_category_name,levels=Parent_cat_name), x = variable, fill = value)) +scale_fill_gradient(low = 'lightblue', high = 'cyan4')+\n    geom_tile()+  labs(x = 'Percentiles', y = 'Parent Category name', fill = 'Deal Probabilities', title = 'With image')+\ntheme_bw()+theme(legend.position = 'bottom')\n\ngrid.arrange(d1, d2, ncol = 1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"03fc15c995f95267a8987d7789b05ae7dbfed215","_cell_guid":"875b922c-8a89-4601-9733-7e61ca0a711a"},"cell_type":"markdown","source":"1. **<a id='10'>10.Most Popular Region  (count and quantile distribution pattern)</a>**"},{"metadata":{"scrolled":true,"_uuid":"f0259dd99df73dde0fb863e0a9b72353c965dc42","trusted":false,"_cell_guid":"6468578a-c309-41db-bea3-06b0f1f5c8b3"},"cell_type":"code","source":"D<-matrix(, nrow = length(unique(train$region)), ncol =length(probs) )\n#D[1,]<-train%>%select(deal_probability) %>%quantile(, probs = probs)\ncategory_count<-train%>% group_by(region)%>%summarize(count=n())%>%\nungroup%>%arrange(desc(count))\nregion_data<-category_count$region\n\ntrain%>% group_by(region)%>%summarize(count=n())%>%mutate(region=factor(region, levels=region_data))%>%  ggplot(aes(x=region, y=count))+geom_bar(stat=\"identity\")+\ngeom_text(aes(x = region, y = 1, label = count),\n            hjust=0, vjust=.5, size = 4, colour = 'black',\n            fontface = 'bold') +\ncoord_flip()+theme_bw()+theme(legend.position = 'bottom')\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7f908de2cd8474ecb8f06b3c36aeb5dc9f50171","_cell_guid":"f4bd9311-2171-4fac-85f8-c8e2de70b130"},"cell_type":"markdown","source":"**<a id='11'>11. Corretaltion between text length and  Deal Probablity</a>**"},{"metadata":{"_uuid":"c86e357642bd16b7e68cc7bc64806c0d2ace7fa3","trusted":false,"_cell_guid":"e283c97e-49ac-4ca2-864f-73147462232a"},"cell_type":"code","source":"train$no_of_char<- nchar(train$description)\nprint(paste(\"the Correlation between the description length and deal probablity is \" ,  cor(train$deal_probability[!is.na(train$description)],train$no_of_char[!is.na(train$description)])))\nprint(paste(\"the Correlation between the description length and price is \" ,  cor(train$price[!is.na(train$description)&!is.na(train$price)],train$no_of_char[!is.na(train$description)&!is.na(train$price)])))\n","execution_count":null,"outputs":[]},{"metadata":{"scrolled":false,"_kg_hide-input":true,"_uuid":"1b086eb90f23967ab8a40e349cd31b9dc6cb6329","trusted":false,"_cell_guid":"b843c3b9-91b8-4b96-8402-5f3199dd99da"},"cell_type":"code","source":"users<-train%>%group_by(user_id)%>%summarize(count=n())%>%arrange(desc(count))%>%head(40)\ntop_40_user<-users$user_id","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"18420fa0a19fba382eae268954ae8e4180cc521d","_cell_guid":"8d421649-99ba-41e5-ad13-648a2797e362"},"cell_type":"markdown","source":"**Characterization of Long term users  and short term users**"},{"metadata":{"_uuid":"84238958c7688f397c63ab92fcf206bbe14b690b","_cell_guid":"712ee839-63ed-49c7-b054-fe445d09e213"},"cell_type":"markdown","source":"Long term users are actually the ones who has learnt the art of selling products online. Its an interesting feature to have in the model. "},{"metadata":{"_kg_hide-input":true,"_uuid":"01b78206c6f6a1d5b48e1b0f9a6084af7713bdb6","trusted":false,"_cell_guid":"86d2bd53-7b7c-4e22-8911-368e2e535b52"},"cell_type":"code","source":"train<-data.table(train)\ntrain[, Usrattempt:=1:.N,  by=list(user_id)]\ntrain[, Usrattempt_by_category:=1:.N,  by=list(user_id,category_name)]\ng1<-train%>%group_by(Usrattempt)%>%summarize(median=median(deal_probability))%>%ggplot(aes(x=Usrattempt, y=median, color=median))+geom_line()+labs(x=\"sequence of attempts by the user\", y=\"Deal_probablity\", title=\"The long term users and short term users\")\ng2<-train%>%group_by(Usrattempt_by_category)%>%summarize(median=median(deal_probability))%>%ggplot(aes(x=Usrattempt_by_category, y=median, color=median))+geom_line()+labs(x=\"sequence of attempts by the user\", y=\"Deal_probablity\", title=\"The long term users and short term users by category\")\ngrid.arrange(g1,g2)\n#test\ntest<-data.table(test)\ntest[, Usrattempt:=1:.N,  by=list(user_id)]\ntest[, Usrattempt_by_category:=1:.N,  by=list(user_id,category_name)]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"411fca130507ae72028edbe23c265a4bb0e03a77","_cell_guid":"ae9d9e9a-98f4-4027-98fd-391f56dc4191"},"cell_type":"markdown","source":"**<a id='12'>12. Binning the Target Variable</a>**"},{"metadata":{"_kg_hide-input":true,"_uuid":"5d35631dca7192f13a3a6625acf2d2ee59ba9d93","trusted":false,"_cell_guid":"00318e4f-214f-4b64-a61f-a2940be852e0"},"cell_type":"code","source":"train$Zero_or_not<-NA\ntest$Zero_or_not<-NA\ntrain$Zero_or_not[train$deal_probability==0]<-'Zero_deal_P'\ntrain$Zero_or_not[train$deal_probability!=0]<-'not_Zero_deal_P'\ntable(train$Zero_or_not)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"46b32e4b090246db1536de73beea4756bbcbc607","_cell_guid":"c70402c0-7be5-4f65-88af-bc284914f54f"},"cell_type":"markdown","source":"**<a id='13'>13. Zero Probablities by the sequence of user attempts within a category </a>**"},{"metadata":{"_kg_hide-input":true,"_uuid":"8d756bc96f0185c64f7d2415c3acc3ca2e6bada0","trusted":false,"_cell_guid":"b83cd874-0705-4d8a-85c8-749d6429a2cf"},"cell_type":"code","source":"g1<-train%>%filter(Usrattempt_by_category%in%seq(1000))%>%filter(user_type=='Private')%>% group_by(Zero_or_not,Usrattempt_by_category)%>%summarize(count=n())%>%  ggplot(aes(x=as.factor(Usrattempt_by_category),y=count, fill=Zero_or_not))+\ngeom_bar(stat=\"identity\", position='fill')+labs(x=\"The Sequence of user attempts by Category \", y=\"Percentage % of Zero_or_not\", title=\"For Private users\")+theme_minimal()\n\n\ng2<- train%>%filter(Usrattempt_by_category%in%seq(1000))%>%filter(user_type=='Company')%>% group_by(Zero_or_not,Usrattempt_by_category)%>%summarize(count=n())%>%  ggplot(aes(x=as.factor(Usrattempt_by_category),y=count, fill=Zero_or_not))+\ngeom_bar(stat=\"identity\", position='fill')+labs(x=\"The Sequence of user attempts by Category\", y=\"Percentage % of Zero_or_not\", title=\"For Companies\")+theme_minimal()\n\ng3<- train%>%filter(Usrattempt_by_category%in%seq(1000))%>%filter(user_type=='Shop')%>% group_by(Zero_or_not,Usrattempt_by_category)%>%summarize(count=n())%>%  ggplot(aes(x=as.factor(Usrattempt_by_category),y=count, fill=Zero_or_not))+\ngeom_bar(stat=\"identity\", position='fill')+labs(x=\"The Sequence of user attempts by Category\", y=\"Percentage % of Zero_or_not\", title=\"For Shop\")+theme_minimal()\n\n\ngrid.arrange(g1,g2,g3)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fe8db75d8d0e1fa131be34648b4078040e8341ed","_cell_guid":"aecb76f6-3d6c-445f-a9b8-f8054fe243df"},"cell_type":"markdown","source":"**Inference:**\n* For shops with more attempts the user learns the art and trying to excel in selling.\n* For Companies it does not look like a good sign they should have short term users intead of long term. May be the customers are getting bored of their usual style of posts. Who knows?\n* The Private users are smart this might be because they are not working for anyone like companies and they are not getting bored of selling their own goods at a good price.  Because they are earning money. But how much ever hard they try they cannot sustain selling their used goods. The customers are becoming smart too. When you become a familiar user you will start getting reviews and customers tend to choose only the users who are authentic."},{"metadata":{"_uuid":"a449b9c1451ba413536b24790c6039b7b01f1e8a","_cell_guid":"df7f1182-7bcd-4155-bbc6-d0c6405f3968"},"cell_type":"markdown","source":"**<a id='14'>14. Zero Probablities by the sequence of user attempts </a>**"},{"metadata":{"_kg_hide-input":true,"_uuid":"ce835aa56375e12a65401996976177e1525e215a","trusted":false,"_cell_guid":"a2654176-cf5f-429c-9744-8f26164a9a21"},"cell_type":"code","source":"g1<-train%>%filter(Usrattempt%in%seq(100))%>%filter(user_type=='Private')%>% group_by(Zero_or_not,Usrattempt)%>%\nsummarize(count=n())%>%  ggplot(aes(x=as.factor(Usrattempt),y=count, fill=Zero_or_not))+\ngeom_bar(stat=\"identity\", position='fill')+labs(x=\"The Sequence of user attempts\", y=\"Percentage % of Zero_or_not\", title=\"For Private users\")+\ntheme_minimal()\n\n\ng2<- train%>%filter(Usrattempt%in%seq(1000))%>%filter(user_type=='Company')%>% \ngroup_by(Zero_or_not,Usrattempt)%>%summarize(count=n())%>%  ggplot(aes(x=as.factor(Usrattempt),y=count, fill=Zero_or_not))+\ngeom_bar(stat=\"identity\", position='fill')+labs(x=\"The Sequence of user attempts\", y=\"Percentage % of Zero_or_not\", title=\"For Companies\")+\ntheme_minimal()\n\ng3<- train%>%filter(Usrattempt%in%seq(1000))%>%filter(user_type=='Shop')%>%\ngroup_by(Zero_or_not,Usrattempt)%>%summarize(count=n())%>%  ggplot(aes(x=as.factor(Usrattempt),y=count, fill=Zero_or_not))+\ngeom_bar(stat=\"identity\", position='fill')+labs(x=\"The Sequence of user attempts\", y=\"Percentage % of Zero_or_not\", title=\"For Shop\")+theme_minimal()\n\n\ngrid.arrange(g1,g2,g3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a76db32d937aba79a7437259fb4c67376171705d","_cell_guid":"77c031c7-2c27-4c3b-af00-bfa7a1a79b19"},"cell_type":"markdown","source":"**Inference:**\n* For shops with more attempts the user learns the art and trying to excel in selling.\n* For Companies it does not look like a good sign they should have short term users intead of long term. May be the customers are getting bored of their usual style of posts. Who knows?\n* The Private users are smart this might be because they are not working for anyone like companies and they are not getting bored of selling their own goods at a good price.  Because they are earning money. But how much ever hard they try they cannot sustain selling their used goods. The customers are becoming smart too. When you become a familiar user you will start getting reviews and customers tend to choose only the users who are authentic."},{"metadata":{"_uuid":"296b961f8d03ddee220134d8b62d99182a683379","_cell_guid":"91d80333-a2fb-48a1-acbf-4fbe909f0f47"},"cell_type":"markdown","source":"** <a id='15'>15. Alluvial plot betweeen user type and deal Probablity  </a>**"},{"metadata":{"_kg_hide-input":true,"_uuid":"a760dca5054e8a4ba2d62e0ba37542fbb90f014f","trusted":false,"_cell_guid":"1990a944-e4d0-465b-9549-a20ce67731db"},"cell_type":"code","source":"train[, flag_param_1:=as.numeric(is.na(param_1))][, flag_param_2:=as.numeric(is.na(param_2))][, flag_param_3:=as.numeric(is.na(param_3))][, flag_description:=as.numeric(is.na(description))][, flag_price:=as.numeric(is.na(price))] [, flag_image:=as.numeric(is.na(image))]\n\ntest[, flag_param_1:=as.numeric(is.na(param_1))][, flag_param_2:=as.numeric(is.na(param_2))][, flag_param_3:=as.numeric(is.na(param_3))][, flag_description:=as.numeric(is.na(description))][, flag_price:=as.numeric(is.na(price))] [, flag_image:=as.numeric(is.na(image))]\n\ntrain%>%group_by(Zero_or_not,user_type )%>%summarise(count=n())%>%  ggplot(\n       aes(weight = count, axis1 = Zero_or_not, axis2 = user_type)) +\n  geom_alluvium(aes(fill = Zero_or_not), width = 1/12) +\n  geom_stratum(width = 1/12, fill = \"black\", color = \"grey\") +\n  geom_label(stat = \"stratum\", label.strata = TRUE) +\n  scale_x_continuous(breaks = 1:2, labels = c(\"Zero Deal Probablity or not\", \"User_type\")) +\n  scale_fill_brewer(type = \"qual\", palette = \"Set1\") +\n  ggtitle(\"Zero Deal Probablities by user type\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bbb45d2d5df906dc39aedac3ad844484c7030e89","_cell_guid":"2333f7b9-1783-4af6-a2d8-c4f8e60480a0"},"cell_type":"markdown","source":"**Adding an additional  Probablity bin  0, 0-0.5 and >0.5**"},{"metadata":{"_uuid":"ac0fd5bb351fac585310508c987f8f8fb9ef6c9e","trusted":false,"_cell_guid":"2bd356ca-ff58-46e7-8425-2f2ff900576a"},"cell_type":"code","source":"train[, Probablity_bins:=ifelse( deal_probability==0,\"Zero\",ifelse(deal_probability<=0.5,\"0 to 0.5\", \"0.5 to 1\"))]\ntest[, Probablity_bins:=NA]\ntrain$Probablity_bins<-factor(train$Probablity_bins, levels=c(\"Zero\",\"0 to 0.5\",\"0.5 to 1\"))\ntrain%>%group_by(Probablity_bins,user_type )%>%summarise(count=n())%>%  ggplot(\n       aes(weight = count, axis1 = Probablity_bins, axis2 = user_type)) +\n  geom_alluvium(aes(fill = Probablity_bins), width = 1/12) +\n  geom_stratum(width = 1/12, fill = \"black\", color = \"grey\") +\n  geom_label(stat = \"stratum\", label.strata = TRUE) +\n  scale_x_continuous(breaks = 1:2, labels = c(\"Probablity bins\", \"User_type\")) +\n  scale_fill_brewer(type = \"qual\", palette = \"Set1\") +\n  ggtitle(\"Probablity bins by user type\")","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"e7629780bd9f1edeac9f3d664a3ffdf921aa65d6","trusted":false,"_cell_guid":"0d0a96a7-2f86-46b8-bdaf-564e841055f1"},"cell_type":"code","source":"train[, complete_cas:=as.numeric(complete.cases(train))]\ntest[, complete_cas:=as.numeric(complete.cases(test))]\ntrain%>%group_by(Probablity_bins,complete_cas )%>%summarise(count=n())%>%  ggplot(\n       aes(weight = count, axis1 = Probablity_bins, axis2 = complete_cas)) +\n  geom_alluvium(aes(fill = Probablity_bins), width = 1/12) +\n  geom_stratum(width = 1/12, fill = \"black\", color = \"grey\") +\n  geom_label(stat = \"stratum\", label.strata = TRUE) +\n  scale_x_continuous(breaks = 1:2, labels = c(\"Probablity bins\", \"Rows with and without missing values 1 and 0\")) +\n  scale_fill_brewer(type = \"qual\", palette = \"Set1\") +\n  ggtitle(\"Probablity bins by  rows with and without missing values\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d22d3920c86d64f9ce61a455531282eccfcd5b20","_cell_guid":"70650ecf-6d90-4375-a25e-948703eda2dd"},"cell_type":"markdown","source":"Image top_1 "},{"metadata":{"_uuid":"200712bcb710338fdf4dc370602ef4c3e7e3e5d0","trusted":false,"_cell_guid":"7630a82f-3568-4747-8759-7eed8af552aa"},"cell_type":"code","source":"train%>% ggplot(aes(x=image_top_1, fill=Probablity_bins)) +geom_histogram()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3e85bd92308d06764e52910e959bd9130dcab996","_cell_guid":"a0e0cda1-a2f6-4a7d-833e-223066db9c8f"},"cell_type":"markdown","source":"**Count of Characters in title**"},{"metadata":{"_uuid":"eb0399fe8ad964668657d4ed199ba85bf88af845","trusted":false,"_cell_guid":"738574de-276a-4438-9c16-c55871dbfa10"},"cell_type":"code","source":"train[,nchar_title:=nchar(title)]\ntest[,nchar_title:=nchar(title)]\ntrain%>% ggplot(aes(x=nchar_title, fill=Probablity_bins)) +geom_histogram()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e344372886ae4371d364ef18d18bfa855f50aae9","_cell_guid":"efc45caf-1702-4606-ab6c-cdcd44dd5a8f"},"cell_type":"markdown","source":"**Count of numbers in Title**"},{"metadata":{"_uuid":"930d93c218a9cf944c8ee0e6ce69eaac592b0a95","trusted":false,"_cell_guid":"6a01a6ef-73b4-4cfc-a075-7762878843ac"},"cell_type":"code","source":"train[,no_numbers_title:=str_count(title, '\\\\d')]\ntest[,no_numbers_title:=str_count(title, '\\\\d')]\ntrain%>% ggplot(aes(x=no_numbers_title, fill=Probablity_bins)) +geom_histogram()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b54b60e279be962108a8287aaa117c5e43b77eec","_cell_guid":"74d6eec9-a563-401d-8471-c2175db3474f"},"cell_type":"markdown","source":"**Count of number of characters in title**"},{"metadata":{"_uuid":"a269c49031171cd10dab6fc517369d010e51b72a","trusted":false,"_cell_guid":"1735b5bc-cf0d-4fcf-8e2b-a6d5feb12c26"},"cell_type":"code","source":"train[,nchar_title:=nchar(title)]\ntest[,nchar_title:=nchar(title)]\ntrain%>% ggplot(aes(x=nchar_title, fill=Probablity_bins)) +geom_histogram()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"60566427460d73e4699d905585a543178858cd47","_cell_guid":"4863c4d1-554e-4945-966d-87083f0b01ea"},"cell_type":"markdown","source":"**Count of number of numbers in description**"},{"metadata":{"_uuid":"682edc6a0b678344fbe9b9c4fcbd9fdc63852746","trusted":false,"_cell_guid":"b9126e59-f70f-41ce-8531-c2d7faa9db40"},"cell_type":"code","source":"train[,no_numbers_desc:=str_count(description, '\\\\d')]\ntest[,no_numbers_desc:=str_count(description, '\\\\d')]\ntrain%>% ggplot(aes(x=no_numbers_desc, fill=Probablity_bins)) +geom_histogram()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ea5fc8d866d9110d767afbe0a654306728cdc81a","_cell_guid":"3d713a69-61c2-4372-a450-4a8506edafb3"},"cell_type":"markdown","source":"** <a id='16'>16.Usage of English letters   </a>**"},{"metadata":{"_uuid":"190eb6dce69e4f704ff8e31f3ec6b0d733647775","_cell_guid":"b2ade32a-5e1c-43aa-af8a-5d34aae5b1b2"},"cell_type":"markdown","source":"I saw some English words during my exploration and I was curious to know the pattern of usage. The plots show an interesting information.  As we could guess, in the consumer electronics category the English word usage is more frequent. And another noticeable category is Trasport category. \n* It shows the importance of international brand value in both those category.  Intuitively I feel those brands will be a good predictor of the deal probability.\n* In Property we can see a minimal usage of English words. There should be other drivers.. "},{"metadata":{"_uuid":"55860df5a6a6c90742ef1f900036a1b497f0fec7","trusted":false,"_cell_guid":"28e9d65c-4b5d-4528-a9e1-75cf8beca36d"},"cell_type":"code","source":"train[,no_English_words:=str_count(title, '[a-zA-Z]')]\ntrain[,no_English_words:=ifelse(is.na(no_English_words),0,no_English_words)]\ntrain[,no_of_letters_title:=nchar(title)]\ng1<-train%>%group_by(parent_category_name)%>%summarize(count=n())%>% ggplot(aes(y=count, x=parent_category_name))+geom_bar(stat=\"identity\", fill = \"blue\")+\nlabs(x = 'Parent Category name', \n       y = 'Count of items', \n       title = 'Items count')+coord_flip()\ng2<-train%>%group_by(parent_category_name)%>%summarize(sum=sum(no_English_words))%>%ggplot(aes(y=sum, x=parent_category_name))+geom_bar(stat=\"identity\", fill=\"blue\")+\nlabs(x = 'Parent Category name', \n       y = 'Count of English letters in Title', \n       title = 'Title English letters ')+coord_flip()\ngrid.arrange(g1,g2,ncol=2)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8b180a3508efaa544eff98b8d514346e7e9ba61f","_cell_guid":"8b1feb47-0ea2-4bbb-bd6d-4c4b4af1d0fa"},"cell_type":"markdown","source":"**No of English characters in description**"},{"metadata":{"_kg_hide-input":true,"_uuid":"b14713dcc378ce9fd0a583d7d380ebe347625838","trusted":false,"_cell_guid":"1d792f37-d9e2-4c81-af15-24eea06a8840"},"cell_type":"code","source":"train[,no_English_letters_description:=str_count(description, '[a-zA-Z]')]\ntrain[,no_English_letters_description:=ifelse(is.na(no_English_letters_description),0,no_English_letters_description)]\ng1<-train%>%group_by(parent_category_name)%>%summarize(count=n())%>% ggplot(aes(y=count, x=parent_category_name))+geom_bar(stat=\"identity\", fill = \"blue\")+\nlabs(x = 'Parent Category name', \n       y = 'Count of items', \n       title = 'Items count')+coord_flip()\ng2<-train%>%group_by(parent_category_name)%>%summarize(sum=sum(no_English_letters_description))%>%ggplot(aes(y=sum, x=parent_category_name))+geom_bar(stat=\"identity\", fill=\"blue\")+\nlabs(x = 'Parent Category name', \n       y = 'Count of English letters', \n       title = 'Desc English letters ')+coord_flip()\ngrid.arrange(g1,g2,ncol=2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f4af8251c9156fb2c6fa675ed679aef03f35e8f1","_cell_guid":"e66b54d1-3576-4d47-a027-1d9bce2cad7f"},"cell_type":"markdown","source":"** <a id='17'>17.Venn Diagram of top  words in each category in Title   </a>**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"d3d6c5f4b9998bc58fd62f72e19f227544e70d34","trusted":false,"_cell_guid":"a63fb61c-9627-4413-9dd5-1c482c232e88"},"cell_type":"code","source":"list.files(\"../input/fork-of-word-vector-category-wise\")\nAnimals<-fread(\"../input/fork-of-word-vector-category-wise/Animals.csv\")\nConsumer_electronics<-fread(\"../input/fork-of-word-vector-category-wise/Consumer_electronics.csv\")\nFor_business<-fread(\"../input/fork-of-word-vector-category-wise/For_business.csv\")\nFor_home_and_cottage<-fread(\"../input/fork-of-word-vector-category-wise/For_home_and_cottage.csv\")\nHobbies_and_Recreation<-fread(\"../input/fork-of-word-vector-category-wise/Hobbies_and_Recreation.csv\")\nPersonal_things<-fread(\"../input/fork-of-word-vector-category-wise/Personal_things.csv\")\nServices<-fread(\"../input/fork-of-word-vector-category-wise/Services.csv\")\nThe_property<-fread(\"../input/fork-of-word-vector-category-wise/The_property.csv\")\nTransport<-fread(\"../input/fork-of-word-vector-category-wise/Transport.csv\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a81143a2dfc0c175777475145d3f51d4d314c8da","_cell_guid":"f9e6f47f-6c45-4186-8567-7531be87211c"},"cell_type":"markdown","source":"**Venn diagram with Top 100 Words from each category in Title**"},{"metadata":{"_uuid":"a7829e570cba7b62ee9e84ffe08b09d7405b2e41","_cell_guid":"f7520f0d-cdba-48cb-80ab-824baae2a02f"},"cell_type":"markdown","source":"The first 100 frequent words are almost unique to all categories. In title there is a noticeable overlap between property and business  almost 20 words are common to them. And also there is a good overlapp between the Personal things and hobbies and recreation."},{"metadata":{"_kg_hide-input":true,"_uuid":"2461148c794faea9fd06d2dde7f275fb217ec15e","trusted":false,"_cell_guid":"4798bc53-b5ef-4e6c-9719-dde07953b67c"},"cell_type":"code","source":"group.venn(list(Transport=Transport$word[1:100],Consumer_electronics=Consumer_electronics$word[1:100],\n           For_business=For_business$word[1:100],Personal_things=Personal_things$word[1:100],\n           The_property=The_property$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list( Consumer_electronics=Consumer_electronics$word[1:100],\n          For_home_and_cottage=For_home_and_cottage$word[1:100],\n                    Hobbies_and_Recreation=Hobbies_and_Recreation$word[1:100],Animals=Animals$word[1:100],\n                   Personal_things=Personal_things$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list(\n           For_business=For_business$word[1:100],\n           The_property=The_property$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n\ngroup.venn(list( Hobbies_and_Recreation=Hobbies_and_Recreation$word[1:100],\n                   Personal_things=Personal_things$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"714f2dfd98bdf4e74b29fbdb158e939f7b91ed06","_cell_guid":"3cfa79bc-7f89-4d66-9c5d-41079a823351"},"cell_type":"markdown","source":"**Venn diagram with Top 100-500 Words from each category in Title**"},{"metadata":{"_uuid":"cbb2717fbe01325779b1503e0078bd74a7cd556a","_cell_guid":"94ca5cc5-d8fe-4352-a2ad-3c6e00249f93"},"cell_type":"markdown","source":"The overlapp between the most frequent words shows a similar pattern btween 101-500 ranked words.\nI guess it will be a more interesting to look at the words in the descriptions. "},{"metadata":{"_kg_hide-input":true,"_uuid":"f72c59f95f8ebc94fa6d2b9caa682b3875e15666","trusted":false,"_cell_guid":"0dc0a760-4a78-4ed7-a6bc-6d4e3cc99b6e"},"cell_type":"code","source":"group.venn(list(Transport=Transport$word[101:500],Consumer_electronics=Consumer_electronics$word[101:500],\n           For_business=For_business$word[101:500],Personal_things=Personal_things$word[101:500],\n           The_property=The_property$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list( Consumer_electronics=Consumer_electronics$word[101:500],\n          For_home_and_cottage=For_home_and_cottage$word[101:500],\n                    Hobbies_and_Recreation=Hobbies_and_Recreation$word[101:500],Animals=Animals$word[101:500],\n                   Personal_things=Personal_things$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list(\n           For_business=For_business$word[101:500],\n           The_property=The_property$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n\ngroup.venn(list( Hobbies_and_Recreation=Hobbies_and_Recreation$word[101:500],\n                   Personal_things=Personal_things$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"828fca5b1a7648adbd0e10744c9f7dc59a0a8b79","_cell_guid":"9e4bbd49-e887-421c-8ccf-9c817770cc30"},"cell_type":"markdown","source":"** <a id='18'>18.Venn Diagram of top  words in each category in Description   </a>**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"d17946281367e8a8aae1ccda17e3f84fa68074cf","trusted":false,"_cell_guid":"477e0df8-8cef-4f3c-90d9-321a7496c8ba"},"cell_type":"code","source":"list.files(\"../input/description-word-vector-category-wise\")\nAnimals<-fread(\"../input/description-word-vector-category-wise/Animals.csv\")\nConsumer_electronics<-fread(\"../input/description-word-vector-category-wise/Consumer_electronics.csv\")\nFor_business<-fread(\"../input/description-word-vector-category-wise/For_business.csv\")\nFor_home_and_cottage<-fread(\"../input/description-word-vector-category-wise/For_home_and_cottage.csv\")\nHobbies_and_Recreation<-fread(\"../input/description-word-vector-category-wise/Hobbies_and_Recreation.csv\")\nPersonal_things<-fread(\"../input/description-word-vector-category-wise/Personal_things.csv\")\nServices<-fread(\"../input/description-word-vector-category-wise/Services.csv\")\nThe_property<-fread(\"../input/description-word-vector-category-wise/The_property.csv\")\nTransport<-fread(\"../input/description-word-vector-category-wise/Transport.csv\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f1ecba41d4514ece8dde3f986038890154bafb54","_cell_guid":"a2497eae-350d-45e1-8b66-2a004a84a5e1"},"cell_type":"markdown","source":"There is a good overlapp between the owrds in the descriptions in each category. And itrs a just a thought these words can be grouped in two categories one are common words and others are specific words.  We adjust the weights between these two groups to build a good model."},{"metadata":{"_uuid":"abb0754cf50705c7ec4b861f1ef0cffad02b82a5","_cell_guid":"20266f43-c243-46db-a4fd-866fdc0eb06b"},"cell_type":"markdown","source":"**Venn diagram with Top 100 Words from each category in Description**"},{"metadata":{"_uuid":"a6e0039174a7b06b5501aeb29f11abc608a06a48","trusted":false,"_cell_guid":"a5db3d09-d159-45a3-8d78-0fecd14e4875"},"cell_type":"code","source":"group.venn(list(Transport=Transport$word[1:100],Consumer_electronics=Consumer_electronics$word[1:100],\n           For_business=For_business$word[1:100],Personal_things=Personal_things$word[1:100],\n           The_property=The_property$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list( Consumer_electronics=Consumer_electronics$word[1:100],\n          For_home_and_cottage=For_home_and_cottage$word[1:100],\n                    Hobbies_and_Recreation=Hobbies_and_Recreation$word[1:100],Animals=Animals$word[1:100],\n                   Personal_things=Personal_things$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list(\n           For_business=For_business$word[1:100],\n           The_property=The_property$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n\ngroup.venn(list( Hobbies_and_Recreation=Hobbies_and_Recreation$word[1:100],\n                   Personal_things=Personal_things$word[1:100]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6343b9461eedafb252631ab3433a0fee894a9a57","_cell_guid":"93f3c2e8-49ca-4f83-b2f0-4affc0431da0"},"cell_type":"markdown","source":"**Venn diagram with Top 100-500 Words from each category in Description**"},{"metadata":{"_uuid":"2a58f4661e6f006266239b9e642b84abdb0493cd","trusted":false,"_cell_guid":"1d035d34-9fff-4cd2-82f4-896c42056fcc"},"cell_type":"code","source":"group.venn(list(Transport=Transport$word[101:500],Consumer_electronics=Consumer_electronics$word[101:500],\n           For_business=For_business$word[101:500],Personal_things=Personal_things$word[101:500],\n           The_property=The_property$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list( Consumer_electronics=Consumer_electronics$word[101:500],\n          For_home_and_cottage=For_home_and_cottage$word[101:500],\n                    Hobbies_and_Recreation=Hobbies_and_Recreation$word[101:500],Animals=Animals$word[101:500],\n                   Personal_things=Personal_things$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\",\"green\",\"red\",\"yellow\"),\n    cat.pos = c(0, 0, 1,0,0),\n    lab.cex=1.1)\n\n\ngroup.venn(list(\n           For_business=For_business$word[101:500],\n           The_property=The_property$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n\ngroup.venn(list( Hobbies_and_Recreation=Hobbies_and_Recreation$word[101:500],\n                   Personal_things=Personal_things$word[101:500]), label=FALSE, \n    fill = c(\"orange\", \"blue\"),\n    cat.pos = c(0, 0),\n    lab.cex=1.1)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f52dec06b51821e86fe02f2a66d58a527a933269","_cell_guid":"1e4eb58f-e10d-467b-a4ed-d95a9b7df02b"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"60a6e33eac7c7425480ef3036070357ee3f8ff1b","_cell_guid":"0e7929e7-fa1c-4e89-afbf-53721130934a"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"1569a7c5e2850f23fef647712b6fc5596c34b204","_cell_guid":"11ab2dfc-78cb-458a-9b6d-60b15293d09e"},"cell_type":"markdown","source":"*** <a id='19'>19.Model Building    </a>**"},{"metadata":{"_uuid":"3832891504a0069021d839d40db7ea9f9699a73c","trusted":false,"_cell_guid":"fa881e9a-90b3-46d9-a497-d6cf98782f68"},"cell_type":"code","source":"rm(list=ls()) #removing all stored dataframes\ntrain = read_csv(\"../input/avito-demand-prediction/train.csv\",locale = locale(encoding = stringi::stri_enc_get()))\nindex_train<-nrow(train)\ntest=read_csv(\"../input/avito-demand-prediction/test.csv\",locale = locale(encoding = stringi::stri_enc_get()))\nnames(train)\ntest$deal_probability<-NA\n\ncateg <- c('Личные вещи' = \"Personal_things\", 'Для дома и дачи' = \"For_home_and_cottage\", 'Бытовая электроника' ='Consumer_electronics', \n               'Недвижимость' = 'The_property', 'Хобби и отдых' = 'Hobbies_and_Recreation', 'Транспорт' = 'Transport', 'Услуги'='Services',\n               'Животные'='Animals', 'Для бизнеса'= 'For_business')\nconversion <- c('Свердловская область'= 'Sverdlovsk oblast','Самарская область'= 'Samara oblast','Ростовская область'= 'Rostov oblast','Татарстан' = 'Tatarstan',\n'Волгоградская область'= 'Volgograd oblast','Нижегородская область' ='Nizhny Novgorod oblast','Пермский край'= 'Perm Krai',\n'Оренбургская область'= 'Orenburg oblast','Ханты-Мансийский АО'= 'Khanty-Mansi Autonomous Okrug','Тюменская область'= 'Tyumen oblast',\n'Башкортостан'= 'Bashkortostan','Краснодарский край'= 'Krasnodar Krai','Новосибирская область'= 'Novosibirsk oblast','Омская область'= 'Omsk oblast',\n'Белгородская область'= 'Belgorod oblast','Челябинская область'= 'Chelyabinsk oblast','Воронежская область'= 'Voronezh oblast','Кемеровская область'= 'Kemerovo oblast',\n'Саратовская область'= 'Saratov oblast','Владимирская область'= 'Vladimir oblast','Калининградская область'= 'Kaliningrad oblast','Красноярский край'= 'Krasnoyarsk Krai',\n'Ярославская область'= 'Yaroslavl oblast','Удмуртия'= 'Udmurtia','Алтайский край'= 'Altai Krai','Иркутская область'= 'Irkutsk oblast','Ставропольский край'= 'Stavropol Krai',\n'Тульская область'= 'Tula oblast')\ntrain$parent_category_name<-(plyr::revalue(train$parent_category_name, categ))\ntrain$region<-(plyr::revalue(train$region, conversion))\ntrain$activation_date<-as.Date(train$activation_date)\ntest$parent_category_name<-(plyr::revalue(test$parent_category_name, categ))\ntest$region<-(plyr::revalue(test$region, conversion))\ntest$activation_date<-as.Date(test$activation_date)\nsample<-read_csv(\"../input/avito-demand-prediction/sample_submission.csv\",locale = locale(encoding = stringi::stri_enc_get()))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"59bebaed3fd36aecb574e5e9d6ae5df104f3158f","trusted":false,"_cell_guid":"734ad727-5f7d-42ce-9adc-05e56f11e11d"},"cell_type":"code","source":"full_data<-rbind(train,test)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1af677941cc12b66379f6c9c6c46c2075151ea87","_cell_guid":"26becfff-ea28-4d78-84d2-92a5553a29f9"},"cell_type":"markdown","source":"Adding all selected features"},{"metadata":{"_uuid":"a92742908c401d136d22d023716b9f3e9da0eb31","trusted":false,"_cell_guid":"f2aa8de0-eaad-4d96-85cc-c922c4dd9b73"},"cell_type":"code","source":"full_data<-data.table(full_data)\nfull_data[, Usrattempt:=1:.N,  by=list(user_id)]\nfull_data[, Usrno_of_attempt:=.N,  by=list(user_id)]\nfull_data[,activation_date:=as.Date(activation_date)]\nfull_data[,month:=month(activation_date)]\nfull_data[,month:=wday(activation_date)]\nfull_data[,nchar_title:=nchar(title)]\nfull_data[,no_numbers_title:=str_count(title, '\\\\d')]\nfull_data[,nchar_title:=nchar(title)]\nfull_data[,no_numbers_desc:=str_count(description, '\\\\d')]\nfull_data[,no_numbers_desc:=ifelse(is.na(no_numbers_desc),0,no_numbers_desc)]\nfull_data[,no_English_words:=str_count(title, '[a-zA-Z]')]\nfull_data[,no_English_words:=ifelse(is.na(no_English_words),0,no_English_words)]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e6ed9f6dfa404186d773e8584e066c508eabce69","_cell_guid":"389bdbc4-3b64-47b0-8965-d4249a1fc9b4"},"cell_type":"markdown","source":"Importing the words feature matrix after pca"},{"metadata":{"_uuid":"11c7b1bcdd30287bf025d97ba9b44dfa60543a16","trusted":false,"_cell_guid":"b50fed32-0e4c-4c01-bdab-81f3b754fa2c"},"cell_type":"code","source":"Animals<-fread(\"../input/pca-for-words-in-title/Animals.csv\")\nAnimals$word<-\"Animals\"\nConsumer_electronics<-fread(\"../input/pca-for-words-in-title/Consumer_electronics.csv\")\nConsumer_electronics$word<-\"Consumer_electronics\"\nFor_business<-fread(\"../input/pca-for-words-in-title/For_business.csv\")\nFor_business$word<-\"For_business\"\nFor_home_and_cottage<-fread(\"../input/pca-for-words-in-title/For_home_and_cottage.csv\")\nFor_home_and_cottage$word<-\"For_home_and_cottage\"\nHobbies_and_Recreation<-fread(\"../input/pca-for-words-in-title/Hobbies_and_Recreation.csv\")\nHobbies_and_Recreation$word<-\"Hobbies_and_Recreation\"\nPersonal_things<-fread(\"../input/pca-for-words-in-title/Personal_things.csv\")\nPersonal_things$word<-\"Personal_things\"\nServices<-fread(\"../input/pca-for-words-in-title/Services.csv\")\nServices$word<-\"Services\"\nThe_property<-fread(\"../input/pca-for-words-in-title/The_property.csv\")\nThe_property$word<-\"The_property\"\nTransport<-fread(\"../input/pca-for-words-in-title/Transport.csv\")\nTransport$word<-\"Transport\"","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"1bed531d6cd644a4660f746614bca88d454c816c","trusted":false,"_cell_guid":"b790c0a4-e256-4b0a-98bf-59ddfb8946ac"},"cell_type":"code","source":"new_table<-rbind(Animals,Consumer_electronics,For_business,For_home_and_cottage,Hobbies_and_Recreation,Personal_things,Services,The_property,Transport)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"46875b23aff611f8415f6cd9925e57eb160f05de","trusted":false,"_cell_guid":"c62860b5-126b-486e-a8e4-97ffe0d058a7"},"cell_type":"code","source":"full_data<-data.table(full_data)\nfull_data <- full_data[order(full_data$parent_category_name),]\nids_test<-full_data%>%filter(is.na(deal_probability))%>%select(item_id)\nids_test<-ids_test$item_id\nfull_data<-cbind(full_data,new_table)\nhead(full_data)\nfull_data<-full_data%>%select(-c(item_id, user_id,region,city,param_1,Usrattempt,param_2,param_3,title,description,activation_date,image,word ))\nnames(full_data)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"f0497fcea6a746e0bb73f17c8a0de18a8ab34d8d","trusted":false,"_cell_guid":"44120fb9-ca59-4300-9a7e-9b3d7a76ef83"},"cell_type":"code","source":"full_data[,parent_category_name:=as.factor(parent_category_name)][,category_name:=as.factor(category_name)][,category_name:=as.factor(user_type)][,month:=as.factor(month)]\nlibrary (caret)\ndmy <- dummyVars(\" ~ .\",  data = full_data)\ntdd<-  data.frame(predict(dmy, newdata = full_data))\ntest<-tdd%>%filter(is.na(deal_probability))\ntrain<-tdd%>%filter(!is.na(deal_probability))\nlibrary(xgboost)\nindex<-createDataPartition(train$deal_probability, p = 0.9, list = F) \ndtrain <- xgb.DMatrix( data = as.matrix(train[index,colnames(train) != \"deal_probability\" ]), label = train[index,\"deal_probability\"])\ndval <- xgb.DMatrix( data =as.matrix(  train[-index, colnames(train) != \"deal_probability\"]), label = train[-index,\"deal_probability\"])\ndtest <- xgb.DMatrix( data =as.matrix(  test[, colnames(test) != \"deal_probability\"]))\n\ncols <- colnames(train)\ncols<-cols[cols != \"deal_probability\"]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9162c1e3d4648b11eb4e925aa925a2407ba467b4","trusted":false,"_cell_guid":"e2eedd5d-fb5e-4b68-8dca-bb408f5b63b0"},"cell_type":"code","source":"params <- list(objective = \"reg:logistic\",\n          booster = \"gbtree\",\n          eval_metric = \"rmse\",\n          nthread = 8,\n          eta = 0.05,\n          max_depth = 7,\n          min_child_weight = 1,\n          subsample = 0.7,\n          colsample_bytree = 0.7,\n          nrounds = 2000)\nxgb <- xgb.train(params, dtrain, params$nrounds, list(val = dval), print_every_n = 50, early_stopping_rounds = 50)\nxgb.importance(cols, model=xgb) %>%\n  xgb.plot.importance(top_n = 15)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"988bb26a7cbc996836689150bf2f27ed1a679fd8","trusted":false,"_cell_guid":"b69edbe1-0a70-4775-8fc6-8f7cf4508929"},"cell_type":"code","source":"pred<- predict(xgb,dtest)\nsample_pred<-data.table(item_id=ids_test,deal_probability_1=pred)\n\nsubmit<-read_csv(\"../input/avito-demand-prediction/sample_submission.csv\") %>% left_join(sample_pred, by='item_id')%>%mutate(deal_probability=deal_probability_1)%>%select(-deal_probability_1)\nfwrite(submit, \"final_pred.csv\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cae9b8a3839da0b6a26296d612a1600db4ba0b2a","_cell_guid":"c658bc03-9417-45bb-9568-aaadbc5350c2"},"cell_type":"markdown","source":"**More analysis to come!! Thanks for looking at my kernel. Its  encouraging for beginners like me. I welcome your suggestions abd feedbac**"}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}