---
title: "Feature engineering 2: Opportunity?"
author: "Roberto Ruiz"
date: "March 15, 2017"
output: html_document
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE)
library(jsonlite)
library(dplyr)
library(purrr)
library(stringr)
library(DT)
data <- fromJSON("../input/train.json")

# unlist every variable except `photos` and `features` and convert to tibble
vars <- setdiff(names(data), c("photos", "features"))
data <- map_at(data, vars, unlist) %>% tibble::as_tibble(.)
train=data
```

<font size="4">Probably you have though about that, but I would like to put it on the table… What about if a big part of the market are people looking for business opportunities to subrent the flats?

I would like to try to answer this question. For that issue I will try to recognize repeated announcements or similar descriptions in the train dataset and see if there is a variation in the price followed by a variation in the interest level. </font>


## Repeated announces:


<font size="4">In order to make a first new variable we can join the street_adress and the number of bedrooms:</font>


```{r, warning=FALSE, message=FALSE}

train$repeated1<-str_c(train$street_address, train$bedrooms)

rep1<-as.data.frame(table(train$repeated1))

dff<-subset(rep1, rep1$Freq>=6)
colnames(dff)<-c("New_variable","Appearances")

datatable(dff)
```


<font size="4">So we can see that for example that only the street adress “200 Water Street” with 2 bedrooms have 69 appareances in the train data set. Furthemore 2372 street adresses have more than 5 appareances, being the total of announcements 22050, while 8113 street adresses have appeared at least two times, being the total number of announcements 36900.</font>


<font size="4">Is there a relationship between the number of appearances and the interest level? We can see if there is an interesting correlation running a linear model:</font>



```{r, warning=FALSE, message=FALSE}

df<-merge(train, rep1, by.x="repeated1", by.y="Var1", all.x = T)
df$interest<-ifelse(df$interest_level=="low",0,ifelse(df$interest_level=="medium",1,2))

mod<-lm(df$interest~df$Freq)
summary(mod)


```


<font size="4">According to the model there is some negative relationship between the interest level and how many times a street address with its number of rooms appears.</font>

<font size="6">Opportunity?</font>

<font size="4">At this point we can try to proof our first hypothesis. Is there a market looking for business opportunities?</font>

<font size="4">We can try to improvise a new variable. This could be the difference between the mean price of a flat with similar characteristics in the same street address and its current price.  </font>

```{r,warning=FALSE, message=FALSE}

average<-aggregate(df$price, by=list(df$repeated1), mean)
colnames(average)<-c("repeated1", "average_price")

df<-merge(df, average, by="repeated1", all.x = T)

df$bussiness_oportunity<-((df$price-df$average_price)/df$average_price)*100

mod<-lm(df$interest~df$bussiness_oportunity)
plot(df$interest~df$bussiness_oportunity)
summary(mod)
```

<font size="4">So, there is an interesting relationship between the difference of the mean price and the current price (Bussines Oportunity) and the interest_level.</font>

<font size="4">I have used this concept in my ML model and my LB score imroved some points....But maybe is only me. </font>

<font size="4">If you find this notebook usefull, give me an up!</font>
