---
title: "EDA and XGB Avito"
author: "Bukun"
output:
  html_document:
    number_sections: true
    toc: true
    fig_width: 10
    code_folding: hide
    fig_height: 4.5
    theme: cosmo
    highlight: tango
---

#Introduction                 

From the Competiton Page 


> In their fourth Kaggle competition, Avito is challenging you to predict demand for an online advertisement based on its full description (title, description, images, etc.), its context (geographically where it was posted, similar ads already posted) and historical demand for similar ads in similar contexts. With this information, Avito can inform sellers on how to best optimize their listing and provide some indication of how much interest they should realistically expect to receive.   

<br/>

<hr/>

**Description of Data**

<hr/>

* `item_id` - Ad id.          

* `user_id` - User id.                     

* `region` - Ad region.          

* `city` - Ad city.              

* `parent_category_name` - Top level ad category as classified by Avito's ad model.           

* `category_name` - Fine grain ad category as classified by Avito's ad model.             

* `param_1` - Optional parameter from Avito's ad model.               

* `param_2` - Optional parameter from Avito's ad model.            

* `param_3` - Optional parameter from Avito's ad model.             

* `title` - Ad title.                

* `description` - Ad description.         

* `price` - Ad price.         

* `item_seq_number` - Ad sequential number for user.             

* `activation_date`- Date ad was placed.               

* `user_type` - User type.             

* `image` - Id code of image. Ties to a jpg file in train_jpg. Not every ad has an image.           

* `image_top_1` - Avito's classification code for the image.           

* `deal_probability` - The target variable. This is the likelihood that an ad actually sold something. It's not possible to verify every transaction with certainty, so this column's value can be any float from zero to one.             


<br/>

<hr/>

**Summary**        

<hr/>

* The Most Popular Regions are `Krasnodar` , `Sverdlovsk` , `Rostov` , `Tatarstan` and `Chelyabinsk`           

* The Most Popular Cities are `Krasnodar` , `Ekaterinburg` , `Novosibirsk` , `Rostov-na-Donu` and `Nizhny Novgorod`      

* The Most Popular Category is `Clothes, shoes, accessories` , `Children's clothing and footwear` , `Goods for children and toys` , `Apartments` , `Phones`               

* The Most Popular Parent Categories are `Personal things` , `home and cottages` , `Consumer electronics` , `Property` and `Hobbies and Recreation`              

* The Most Popular Parameter 1 are `Womens clothing` , `For Boys` , `For Girls`,`Selling` and `With Mileage`                         
* The Most Popular Parameter 2 are `Footwear` , `Outerwear` , `Dresses and skirts`,`Other` and `Knitwear`   

* The most Popular Titles are `Dress`,`Shoes`,`Jacket`,`Coat` and `Jeans`         

* The Deal Probability has maximum number of rows with probability `zero`   

* No real pattern appears for Activation on a day of week                

* The Activation days are usually seen to be in the later part of the month               

* The `median` price is `1300` and `mean` price is `316K`           

* The regions with the highest `median` prices are `Krasnodar` , `Stavropol` , `Khanty-Mansiysk Autonomous Okrug` , `Voronezh` and `Irkutsk`     


* `Property Abroad` , `Houses and Cottages` , `Apartments` , `Land` and `Rooms` are the highest prices among Categories            

* `Property` , `Transport` , `Business` , `Consumer Electronics` , `Home and Cottages` are the are the highest prices among Parent Categories           

* The **Costly items** appear at the beginning of the month                

* The Median Price seems to be the same on all Weekdays       

* The Mean Price is the highest on `Wednesday`         

* Median length of the Titles is `20`       

* Median length of the Description is `56`        

* `72%` of the users are private            


<hr/>


#Preparation{.tabset .tabset-fade .tabset-pills}

       
##Load Libraries

```{r,message=FALSE,warning=FALSE}

library(tidyverse)
library(tidytext)
library(stringr)
library(knitr)
library(lubridate)
library(caret)

library(tidytext)
library(wordcloud)

library(text2vec)
library(stopwords)
library(Matrix)

```

##Read the data

```{r,message=FALSE,warning=FALSE}

rm(list=ls())

fillColor = "#FFA07A"
fillColor2 = "#F1C40F"

train = read_csv("../input/train.csv",locale = locale(encoding = stringi::stri_enc_get()))
test = read_csv("../input/test.csv",locale = locale(encoding = stringi::stri_enc_get()))
periods_train <- read_csv("../input/periods_train.csv")

TotalNumberOfRows = nrow(train)

train <-train %>%
  mutate(title_len = str_count(title)) %>%
  mutate(description_len = str_count(description))

test <-test %>%
  mutate(title_len = str_count(title)) %>%
  mutate(description_len = str_count(description))

```

#Glimpse of Data{.tabset .tabset-fade .tabset-pills}

##Train dataset

```{r,message=FALSE,warning=FALSE}

glimpse(train)

```

#Region Analysis{.tabset .tabset-fade .tabset-pills}

##Most Popular Region

The Most Popular Regions are `Krasnodar region` , `Sverdlovsk` , `Rostov` , `Tatarstan` and `Chelyabinsk`            

```{r,message=FALSE,warning=FALSE}

region <- c("Краснодарский край","Свердловская область","Ростовская область",
"Татарстан","Челябинская область","Нижегородская область","Самарская область",
"Башкортостан","Пермский край","Новосибирская область","Ставропольский край",
"Ханты-Мансийский АО","Воронежская область","Иркутская область","Тульская область","Тюменская область",
"Белгородская область")

region_en <- c("Krasnodar","Sverdlovsk","Rostov","Tatarstan","Chelyabinsk",
"Nizhny Novgorod","Samara","Bashkortostan","Perm","Novosibirsk","Stavropol",
"Khanty-Mansiysk Autonomous Okrug","Voronezh","Irkutsk","Tula","Tyumen",
"Belgorod")


df_regions_en <- as.data.frame(cbind(region,region_en) )

train %>%
  filter(!is.na(region)) %>%
  group_by(region) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  left_join(df_regions_en) %>%
   mutate(region_en = reorder(region_en,Count)) %>%
  
  ggplot(aes(x = region_en,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = region_en, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'region', 
       y = 'Percentage', 
       title = 'Most popular region') +
  coord_flip() + 
  theme_bw()

```

##Region Data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(region)) %>%
  group_by(region) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(region = reorder(region,Count)) %>%
  head(10) %>%
  kable()

```


##Region and Deal Probability

```{r,message=FALSE,warning=FALSE}

dataset <- train %>%
  filter(!is.na(region)) %>%
  group_by(region) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(region = reorder(region,Count)) %>%
  head(10)

train %>%
  filter(region %in% dataset$region) %>%
  mutate( region = as.factor(region)) %>%
  ggplot(aes(x = region, y= deal_probability, fill = region)) +
  geom_boxplot() +
  labs(x= 'Region',y = 'Deal Probablity', 
       title = paste("Distribution of", 'Deal Probablity ')) +
  theme_bw() + theme(axis.text.x = element_text(angle = 90, hjust = 1))


```


#Most Popular City{.tabset .tabset-fade .tabset-pills}

##Bar Plot

The Most Popular Cities are `Krasnodar` , `Ekaterinburg` , `Novosibirsk` , `Rostov-na-Donu` and `Nizhny Novgorod`

```{r,message=FALSE,warning=FALSE}
city <-c("Краснодар","Екатеринбург","Новосибирск","Ростов-на-Дону","Нижний Новгород",
"Челябинск","Пермь","Казань","Самара","Омск")

city_en <-c("Krasnodar","Ekaterinburg","Novosibirsk","Rostov-na-Donu","Nizhny Novgorod",
"Chelyabinsk","Permian","Kazan","Samara","Omsk")

df_city_en <- as.data.frame(cbind(city,city_en) )

train %>%
  filter(!is.na(city)) %>%
  group_by(city) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  left_join(df_city_en) %>%
  mutate(city_en = reorder(city_en,Count)) %>%
  
  ggplot(aes(x = city_en,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = city_en, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'city', 
       y = 'Percentage', 
       title = 'Most popular city') +
  coord_flip() + 
  theme_bw()

```


##City Data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(city)) %>%
  group_by(city) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(city = reorder(city,Count)) %>%
  head(10) %>%
 left_join(df_city_en) %>%
 kable() 
 
```



#Most Popular Category{.tabset .tabset-fade .tabset-pills}

The Most Popular Category is `Clothes, shoes, accessories` , `Children's clothing and footwear` , `Goods for children and toys` , `Apartments` , `Phones`               

##Bar Plot

```{r,message=FALSE,warning=FALSE}

category_name <- c("Одежда, обувь, аксессуары","Детская одежда и обувь","Товары для детей и игрушки",
"Квартиры","Телефоны","Мебель и интерьер","Предложение услуг","Автомобили","Ремонт и строительство",
"Бытовая техника","Недвижимость за рубежом","Квартиры","Дома, дачи, коттеджи",
"Земельные участки","Комнаты","Грузовики и спецтехника","Готовый бизнес","Автомобили",
"Гаражи и машиноместа","Коммерческая недвижимость")

category_name_en <- c("Clothes,shoes accessories" , "Children's clothing and footwear" ,
"Goods for children and toys" , "Apartments" , "Phones","Furniture and interior","Offer of services","Cars",
"Repair and construction","Appliances","Property Abroad",
"Houses, cottages, cottages","Land","Rooms","Trucks and special equipment","Ready business",
"Garages and parking places","Commercial Property")

df_category_en <- as.data.frame(cbind(category_name,category_name_en ) )


train %>%
  filter(!is.na(category_name)) %>%
  group_by(category_name) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  left_join(df_category_en) %>%
  mutate(category_name_en = reorder(category_name_en,Count)) %>%
  
  ggplot(aes(x = category_name_en,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = category_name_en, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Category', 
       y = 'Percentage', 
       title = 'Most popular Category') +
  coord_flip() + 
  theme_bw()

```

##Most Popular Category data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(category_name)) %>%
  group_by(category_name) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(category_name = reorder(category_name,Count)) %>%
  head(10) %>%
  left_join(df_category_en) %>%
  kable() 
  
  
```


#Most Popular Parent Category{.tabset .tabset-fade .tabset-pills}

The Most Popular Parent Categories are `Personal things` , `home and cottages` , `Consumer electronics` , `Property` and `Hobbies and Recreation`              

##Bar Plot


```{r,message=FALSE,warning=FALSE}

parent_category_name <- c("Личные вещи","Для дома и дачи","Бытовая электроника","Недвижимость",
"Хобби и отдых","Транспорт","Услуги","Животные","Для бизнеса")

parent_category_name_en <- c("Personal things","home and cottages","Consumer electronics","Property",
"Hobbies and Recreation","Transport","services","Animals","business")

df_parentcategory_en <- as.data.frame(cbind(parent_category_name,parent_category_name_en ) )


train %>%
  filter(!is.na(parent_category_name)) %>%
  group_by(parent_category_name) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  left_join(df_parentcategory_en) %>%
  mutate(parent_category_name_en = reorder(parent_category_name_en,Count)) %>%
  
  
  ggplot(aes(x = parent_category_name_en,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = parent_category_name_en, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Parent Category', 
       y = 'Percentage', 
       title = 'Most popular Parent Category') +
  coord_flip() + 
  theme_bw()

```

##Most Popular Category data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(parent_category_name)) %>%
  group_by(parent_category_name) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(parent_category_name = reorder(parent_category_name,Count)) %>%
  head(10) %>%
  left_join(df_parentcategory_en) %>%
  kable()


```


#Most Popular Image Code 

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(image_top_1)) %>%
  group_by(image_top_1) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  mutate(image_top_1 = reorder(image_top_1,Count)) %>%
  
  ggplot(aes(x = image_top_1,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = image_top_1, y = 0.25, label = paste0("(",round(Count,2)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Image Code', 
       y = 'Percentage', 
       title = 'Most popular Image Code') +
  coord_flip() + 
  theme_bw()



```

##Most Popular Image Code with Deal Probability greater than 0.25

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(image_top_1)) %>%
  filter(deal_probability > 0.25) %>%
  group_by(image_top_1) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  mutate(image_top_1 = reorder(image_top_1,Count)) %>%
  
  ggplot(aes(x = image_top_1,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = image_top_1, y = 0.25, label = paste0("(",round(Count,2)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Image Code', 
       y = 'Percentage', 
       title = 'Most popular Image Code with Deal Probability greater than 0.25') +
  coord_flip() + 
  theme_bw()

```


#Most Popular Param1{.tabset .tabset-fade .tabset-pills}

The Most Popular Parameter 1 are `Womens clothing` , `For Boys` , `For Girls`,`Selling` and `With Mileage`      

##Bar plot

```{r,message=FALSE,warning=FALSE}

param_1 <- c("Женская одежда","Для девочек","Для мальчиков",
"Продам","С пробегом","Аксессуары","Мужская одежда","Другое",
"Игрушки","Детские коляски")

param_1_en <- c("Women's clothing","For girls","For boys",
"Selling","With mileage","Accessories","Men's clothing","Other",
"Toys","Baby carriages")

df_param_1_en <- as.data.frame(cbind(param_1,param_1_en ) )


train %>%
  filter(!is.na(param_1)) %>%
  group_by(param_1) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  left_join(df_param_1_en) %>%
  mutate(param_1_en = reorder(param_1_en,Count)) %>%
  
  
  ggplot(aes(x = param_1_en,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = param_1_en, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Param 1', 
       y = 'Percentage', 
       title = 'Most popular Param 1') +
  coord_flip() + 
  theme_bw()

```

##Most Popular Param 1

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(param_1)) %>%
  group_by(param_1) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(param_1 = reorder(param_1,Count)) %>%
  head(10) %>%
  left_join(df_param_1_en) %>%
  kable() 


```

#Most Popular Param2{.tabset .tabset-fade .tabset-pills}

The Most Popular Parameter 2 are `Footwear` , `Outerwear` , `Dresses and skirts`,
`Other` and `Knitwear`     

##Bar Plot

```{r,message=FALSE,warning=FALSE}

param_2 <- c("Обувь","Верхняя одежда","Платья и юбки","Другое",
"Трикотаж","Брюки","1","2","На длительный срок","Дом")

param_2_en <- c("Footwear","Outerwear","Dresses and skirts","Other",
"Knitwear","Pants","1","2","For a long time","House")

df_param_2_en <- as.data.frame(cbind(param_2,param_2_en ) )


train %>%
  filter(!is.na(param_2)) %>%
  group_by(param_2) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  left_join(df_param_2_en) %>%
  mutate(param_2_en = reorder(param_2_en,Count)) %>%
  
  
  ggplot(aes(x = param_2_en,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = param_2_en, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Param 2', 
       y = 'Percentage', 
       title = 'Most popular Param 2') +
  coord_flip() + 
  theme_bw()

```

##Most Popular Param 2

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(param_2)) %>%
  group_by(param_2) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(param_2 = reorder(param_2,Count)) %>%
  head(10) %>%
  left_join(df_param_2_en) %>%
  kable()

```

#Most Popular Param3{.tabset .tabset-fade .tabset-pills}   


##Bar Plot

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(param_3)) %>%
  group_by(param_3) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  mutate(param_3 = reorder(param_3,Count)) %>%
  
  
  ggplot(aes(x = param_3,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = param_3, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Param 3', 
       y = 'Percentage', 
       title = 'Most popular Param 3') +
  coord_flip() + 
  theme_bw()

```

##Data

```{r,message=FALSE, warning=FALSE}

train %>%
  filter(!is.na(param_3)) %>%
  group_by(param_3) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(param_3 = reorder(param_3,Count)) %>%
  head(10) %>%
  
  kable()

```

#Popular Titles

The most Popular Titles are `Dress`,`Shoes`,`Jacket`,`Coat` and `Jeans`

##Bar Plot

```{r,message=FALSE,warning=FALSE}

title <- c("Платье","Туфли","Куртка",
"Пальто","Джинсы","Комбинезон",
"Кроссовки","Костюм","Ботинки",
"Босоножки")

title_en <- c("Dress","Shoes","Jacket",
"Coat","Jeans","Overalls",
"Sneakers","Costume","Boots",
"Sandals")

df_title_en <- as.data.frame(cbind(title,title_en) )

train %>%
  filter(!is.na(title)) %>%
  group_by(title) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  head(10) %>%
  left_join(df_title_en) %>%
  mutate(title_en = reorder(title_en,Count)) %>%
  
  ggplot(aes(x = title_en,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = title_en, y = 1, label = paste0("(",round(Count,2)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Title', 
       y = 'Percentage', 
       title = 'Most popular Title') +
  coord_flip() + 
  theme_bw()

```

##Title Data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(title)) %>%
  group_by(title) %>%
  summarise(Count = n()) %>%
  arrange(desc(Count)) %>%
  mutate(title = reorder(title,Count)) %>%
  head(10) %>%
  left_join(df_title_en) %>%
  kable()

```


#Image in Data

The plot shows the percentage of data having `images`. Most of the ads have images associated with it.    


```{r,message=FALSE,warning=FALSE}

getImageIndicator <- function(dataset)
{
  
  if ( !is.na(dataset))
  {
    return(1)
  }
  else
  {
    return(0)
  }
  
}

train$image_indicator = NULL
train$image_indicator = sapply(train$image,getImageIndicator)

test$image_indicator = NULL
test$image_indicator = sapply(test$image,getImageIndicator)


train %>%
  filter(!is.na(image_indicator)) %>%
  group_by(image_indicator) %>%
  summarise(Count = n()/TotalNumberOfRows *100 )%>%
  arrange(desc(Count)) %>%
  mutate(image_indicator = reorder(image_indicator,Count)) %>%
  head(10) %>%
  
  ggplot(aes(x = image_indicator,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = image_indicator, y = 1, label = paste0("(",round(Count,2)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Image Indicator ', 
       y = 'Percentage', 
       title = 'Image Indicator and Count') +
  coord_flip() + 
  theme_bw()

```

#Distribution of Duration of Ad

The number of days the ad runs the most is **13**             


```{r,message=FALSE,warning=FALSE}

periods_train <- periods_train %>%
  mutate(duration_ad = interval(date_from, date_to) /ddays(1))

periods_train %>%
  ggplot(aes(x = duration_ad)) +
  geom_histogram(bins = 30,fill = fillColor2) +
  labs(x= 'Duration of Ad',y = 'Count', title = paste("Distribution of", ' Duration of Ad ')) +
  theme_bw()

summary(periods_train$duration_ad)

```


#Deal Probablity Analysis


##Distribution of Deal Probablity

The Deal Probability has maximum number of rows with probability `zero`

```{r,message=FALSE,warning=FALSE}

train %>%
  ggplot(aes(x = deal_probability)) +
  geom_histogram(bins = 30,fill = fillColor2) +
  labs(x= 'Deal Probablity',y = 'Count', title = paste("Distribution of", ' Deal Probablity ')) +
  theme_bw()

summary(train$deal_probability)

```

##Distribution of Deal Probablity with 10 bins

```{r,message=FALSE,warning=FALSE}

train %>%
  ggplot(aes(x = deal_probability)) +
  geom_histogram(bins = 10,fill = fillColor2) +
  scale_x_continuous(breaks = seq(0 , 1 , 0.2 )) +
  labs(x= 'Deal Probablity',y = 'Count', title = paste("Distribution of", ' Deal Probablity ')) +
  theme_bw()

```


##Distribution of Deal Probablity with 2 bins

```{r,message=FALSE,warning=FALSE}

train %>%
  ggplot(aes(x = deal_probability)) +
  geom_histogram(bins = 2,fill = fillColor) +
  labs(x= 'Deal Probablity',y = 'Count', title = paste("Distribution of", ' Deal Probablity ')) +
  theme_bw()

```

#Distribution of User Type            

`72%` of the users are private            


```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(user_type)) %>%
  group_by(user_type) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  mutate(user_type = reorder(user_type,Count)) %>%
  head(10) %>%
  
  ggplot(aes(x = user_type,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = user_type, y = 1, label = paste0("(",round(Count)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'User Type', 
       y = 'Count', 
       title = 'Most popular User Type') +
  coord_flip() + 
  theme_bw()

```

#Most Popular Activation Day of Week 

No real pattern appears for Activation on a day of week                    

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(activation_date)) %>%
  mutate(WeekDayName = wday(ymd(activation_date),label = TRUE)) %>%
  group_by(WeekDayName) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  mutate(WeekDayName = reorder(WeekDayName,Count)) %>%
  head(10) %>%
  
  ggplot(aes(x = WeekDayName,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = WeekDayName, y = 1, label = paste0("(",round(Count,0)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Activation WeekDayName', 
       y = 'Percentage', 
       title = 'Activation WeekDayName and Count') +
  coord_flip() + 
  theme_bw()
  

```

#Most Popular Activation Day of Month 

The Activation days are usually seen to be in the later part of the month            


```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(activation_date)) %>%
  mutate(WeekDayName = day(ymd(activation_date))) %>%
  group_by(WeekDayName) %>%
  summarise(Count = n()/TotalNumberOfRows *100) %>%
  arrange(desc(Count)) %>%
  mutate(WeekDayName = reorder(WeekDayName,Count)) %>%
  head(10) %>%
  
  ggplot(aes(x = WeekDayName,y = Count) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = WeekDayName, y = 1, label = paste0("(",round(Count,2)," %)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Activation Day of the Month', 
       y = 'Percentage', 
       title = 'Activation Day of the Month and Count') +
  coord_flip() + 
  theme_bw()
  

```


#Distribution of Price

The `median` price is `1300` and `mean` price is `316K`         


```{r,message=FALSE,warning=FALSE}

train %>%
    ggplot(aes(x = price) )+
    scale_x_log10(
      breaks = scales::trans_breaks("log10", function(x) 10^x),
      labels = scales::trans_format("log10", scales::math_format(10^.x))
    ) +
    scale_y_log10(
      breaks = scales::trans_breaks("log10", function(x) 10^x),
      labels = scales::trans_format("log10", scales::math_format(10^.x))
    ) + 
    geom_histogram(fill = fillColor2,bins=50) +
    labs(x = 'Price' ,y = 'Count', title = paste("Distribution of", "Price")) +
    theme_bw()

summary(train$price)

```

##Median Price and Region Analysis{.tabset .tabset-fade .tabset-pills}     

The regions with the highest `median` prices are `Krasnodar` , `Stavropol` , `Khanty-Mansiysk Autonomous Okrug` , `Voronezh` and `Irkutsk`           


###Bar Plot 

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(region)) %>%
  group_by(region) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  head(10) %>%
  left_join(df_regions_en) %>%
  mutate(region_en = reorder(region_en,MedianPrice)) %>%
  
  
  ggplot(aes(x = region_en,y = MedianPrice) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = region_en, y = 1, label = paste0("(",round(MedianPrice/1e3,0),"K )",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black', 
            fontface = 'bold') +
  labs(x = 'Region ', 
       y = 'Median Price', 
       title = 'Region and Median Price') +
  coord_flip() + 
  theme_bw()
  

```


###Data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(region)) %>%
  group_by(region) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  mutate(region = reorder(region,MedianPrice)) %>%
  head(10) %>%
  kable()

```



##Median Price and Category Analysis{.tabset .tabset-fade .tabset-pills}     

`Property Abroad` , `Houses and Cottages` , `Apartments` , `Land` and `Rooms` are the highest prices among Categories            


###Bar Plot 

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(category_name)) %>%
  group_by(category_name) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  head(10) %>%
  left_join(df_category_en) %>%
  mutate(category_name_en = reorder(category_name_en,MedianPrice)) %>%
  
  
  ggplot(aes(x = category_name_en,y = MedianPrice) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = category_name_en, y = 1, label = paste0("(",round(MedianPrice/1e3,0),"K )",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black', 
            fontface = 'bold') +
  labs(x = 'Category ', 
       y = 'Median Price', 
       title = 'Category and Median Price') +
  coord_flip() + 
  theme_bw()
  

```


###Data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(category_name)) %>%
  group_by(category_name) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  mutate(category_name = reorder(category_name,MedianPrice)) %>%
  head(10) %>%
  kable()

```


##Median Price and Parent Category Analysis{.tabset .tabset-fade .tabset-pills}    

`Property` , `Transport` , `Business` , `Consumer Electronics` , `Home and Cottages` are the are the highest prices among Parent Categories            

###Bar Plot 

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(parent_category_name)) %>%
  group_by(parent_category_name) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  head(10) %>%
  left_join(df_parentcategory_en) %>%
  mutate(parent_category_name_en = reorder(parent_category_name_en,MedianPrice)) %>%
  
  
  ggplot(aes(x = parent_category_name_en,y = MedianPrice) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = parent_category_name_en, y = 1, label = paste0("(",round(MedianPrice,0)," )",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Parent Category ', 
       y = 'Median Price', 
       title = 'Parent Category and Median Price') +
  coord_flip() + 
  theme_bw()
  

```


###Data

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(parent_category_name)) %>%
  group_by(parent_category_name) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  mutate(parent_category_name = reorder(parent_category_name,MedianPrice)) %>%
  head(10) %>%
  kable()

```


##Median Price on Week of the Month             

The **Costly items** appear at the beginning of the month                


```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(activation_date)) %>%
  mutate(WeekDayName = day(ymd(activation_date))) %>%
  group_by(WeekDayName) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  mutate(WeekDayName = reorder(WeekDayName,MedianPrice)) %>%
  head(10) %>%
  
  ggplot(aes(x = WeekDayName,y = MedianPrice) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = WeekDayName, y = 1, label = paste0("(",round(MedianPrice,0)," )",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Activation Day of Month ', 
       y = 'Median Price', 
       title = 'Activation Day of Month and Median Price') +
  coord_flip() + 
  theme_bw()
  

```

##Median Price on Weekday 

The Median Price seems to be the same on all Weekdays

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(activation_date)) %>%
  mutate(WeekDayName = wday(ymd(activation_date),label = TRUE)) %>%
  group_by(WeekDayName) %>%
  summarise(MedianPrice = median(price,na.rm = TRUE) )%>%
  arrange(desc(MedianPrice)) %>%
  mutate(WeekDayName = reorder(WeekDayName,MedianPrice)) %>%
  head(10) %>%
  
  ggplot(aes(x = WeekDayName,y = MedianPrice) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = WeekDayName, y = 1, label = paste0("(",round(MedianPrice,0)/1e3," K)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Activation WeekDayName ', 
       y = 'Median Price', 
       title = 'Activation WeekDayName and Median Price') +
  coord_flip() + 
  theme_bw()
  

```

##Mean Price on Weekday 

The Mean Price is the highest on `Wednesday`            


```{r,message=FALSE,warning=FALSE}

train %>%
  filter(!is.na(activation_date)) %>%
  mutate(WeekDayName = wday(ymd(activation_date),label = TRUE)) %>%
  group_by(WeekDayName) %>%
  summarise(MeanPrice = mean(price,na.rm = TRUE) )%>%
  arrange(desc(MeanPrice)) %>%
  mutate(WeekDayName = reorder(WeekDayName,MeanPrice)) %>%
  head(10) %>%
  
  ggplot(aes(x = WeekDayName,y = MeanPrice) ) +
  geom_bar(stat='identity',colour="white", fill = fillColor2) +
  geom_text(aes(x = WeekDayName, y = 1, label = paste0("(",round(MeanPrice,0)/1e3," K)",sep="")),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Activation WeekDayName ', 
       y = 'Mean Price', 
       title = 'Activation WeekDayName and Mean Price') +
  coord_flip() + 
  theme_bw()
  

```

#Distribution of Title length

Median length of the Titles is `20`                  

```{r,message=FALSE,warning=FALSE}

train %>%
  ggplot(aes(x = title_len)) +
  geom_histogram(bins = 30,fill = fillColor) +
  labs(x= 'Title Length',y = 'Count', title = paste("Distribution of", ' Title Length ')) +
  theme_bw()

summary(train$title_len)

```

#Distribution of Description length

Median length of the Description is `56`                  

```{r,message=FALSE,warning=FALSE}

train %>%
  filter(description_len <1e3) %>%
  ggplot(aes(x = description_len)) +
  geom_histogram(bins = 30,fill = fillColor2) +
  labs(x= 'Description Length',y = 'Count', title = paste("Distribution of", ' Description Length ')) +
  theme_bw()

summary(train$description_len)

```


#Feature Engineering

Adding `Day of Month` and `Day Name` for the dates            


```{r,message=FALSE,warning=FALSE}

train <- train %>%
  mutate(DayOfMonth = day(ymd(activation_date)))

train <- train %>%
  filter(!is.na(activation_date)) %>%
  mutate(DayName = wday(ymd(activation_date)))

test <- test %>%
  mutate(DayOfMonth = day(ymd(activation_date)))

test <- test %>%
  mutate(DayName = wday(ymd(activation_date)))


```

#XGBoost Modelling

##Selecting Columns

```{r,message=FALSE,warning=FALSE}

train2 <- train %>%
  select(-user_id,-item_id,-title,-description,-image,-item_seq_number,
         -deal_probability,-activation_date)

test2 <- test %>%
  select(-user_id,-item_id,-title,-description,-image,-item_seq_number,-activation_date)

train2 <- train2 %>%
  replace_na(list(image_top_1 = -1, price = -1))

test2 <- test2 %>%
  replace_na(list(image_top_1 = -1, price = -1))


colnames(train2)

```

##Transform to Numeric

The following code snippet transforms all the variables to **numeric**.

```{r,message=FALSE,warning=FALSE}

features <- colnames(train2)

for (f in features) {
  if ((class(train2[[f]])=="factor") || (class(train2[[f]])=="character")) {
    levels <- unique(train2[[f]])
    train2[[f]] <- as.numeric(factor(train2[[f]], levels=levels))
  }
}

train2$deal_probability = NULL
train2$deal_probability = train$deal_probability
train2$price = log(train2$price)

features <- colnames(test2)

for (f in features) {
  if ((class(test2[[f]])=="factor") || (class(test2[[f]])=="character")) {
    levels <- unique(test2[[f]])
    test2[[f]] <- as.numeric(factor(test2[[f]], levels=levels))
  }
}

test2$price = log1p(test2$price)

```

##Model

```{r,message=FALSE,warning=FALSE}

formula = deal_probability ~ .

fitControl <- trainControl(method="none",number = 3)

xgbGrid <- expand.grid(nrounds = 5,
                       max_depth = 7,
                       eta = .05,
                       gamma = 0,
                       colsample_bytree = .8,
                       min_child_weight = 1,
                       subsample = 1)

set.seed(13)

DealProbablityXGB = train(formula, data = train2,
                        method = "xgbTree",trControl = fitControl,
                        tuneGrid = xgbGrid,na.action = na.pass,
                        metric = "RMSE")

DealProbablityXGB

```


##Variable Importance

```{r,message=FALSE,warning=FALSE}

importance = varImp(DealProbablityXGB)

varImportance <- data.frame(Variables = row.names(importance[[1]]), 
                            Importance = round(importance[[1]]$Overall,2))

# Create a rank variable based on importance
rankImportance <- varImportance %>%
  mutate(Rank = paste0('#',dense_rank(desc(Importance)))) %>%
  head(10)

rankImportancefull = rankImportance

ggplot(rankImportance, aes(x = reorder(Variables, Importance), 
                           y = Importance)) +
  geom_bar(stat='identity',colour="white", fill = fillColor) +
  geom_text(aes(x = Variables, y = 1, label = Rank),
            hjust=0, vjust=.5, size = 4, colour = 'black',
            fontface = 'bold') +
  labs(x = 'Variables', title = 'Relative Variable Importance') +
  coord_flip() + 
  theme_bw()

```


#Predictions

```{r,message=FALSE,warning=FALSE}

predictions = predict(DealProbablityXGB,test2,na.action=na.pass)

predictions = abs(predictions)

# Save the solution to a dataframe
solution <- data.frame('item_id' = test$item_id, 'deal_probability' = predictions)

head(solution)

# Write it to file
write.csv(solution, 'XGB.csv', row.names = F)

```


