{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "52923875-9752-9293-7f26-3f0128c8be6a"
      },
      "source": [
        "**Managers mistakes ?**  \n",
        "At first I tried data wrangling with **\"photos\", \"bedrooms\", \"bathrooms\" and \"price\".**   \n",
        "I found some mysterious data. Well, you might think it is not fundamental issues ....   \n",
        "\n",
        "I am sorry that I cannot show photo images here in kaggle environment, so I commented-out all \"saveShow_img()\".\n",
        "if you copy this scripts, uncomment all \"saveShow_img()\", and execute in normal local computing,  you can see related photos.\n",
        "\n",
        "My native language is Japanese, sorry for my poor English.  "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "c8d9195e-3521-ee02-23e6-2a81e8889e17"
      },
      "outputs": [],
      "source": [
        "packages <- c(\"jsonlite\", \"dplyr\", \"purrr\", \"knitr\", \"jpeg\", \"tidyverse\")\n",
        "purrr::walk(packages, library, character.only = TRUE, warn.conflicts = FALSE)\n",
        "\n",
        "data_dir <- \"../input/\"\n",
        "\n",
        "# -----------------------------------------------------------------------------\n",
        "# Load data and convert to tibble\n",
        "# -----------------------------------------------------------------------------\n",
        "train_org <- fromJSON(paste0(data_dir, \"train.json\"))\n",
        "test <- fromJSON(paste0(data_dir, \"test.json\"))\n",
        "\n",
        "# unlist every variable except `photos` and `features` and convert to tibble\n",
        "vars <- setdiff(names(train_org), c(\"photos\", \"features\"))\n",
        "\n",
        "train_org <- map_at(train_org, vars, unlist) %>% tibble::as_tibble(.)\n",
        "test <- map_at(test, vars, unlist) %>% tibble::as_tibble(.)\n",
        "\n",
        "# Convert interest_level to factor\n",
        "train_org$interest_level <- factor(train_org$interest_level, c(\"low\", \"medium\", \"high\"))\n",
        "\n",
        "# bind train and test data\n",
        "test$interest_level <- NA\n",
        "all <- rbind(train_org, test)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "634fcc8f-79da-168f-25c6-676111fed15f"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# FUNCTION to save and show images\n",
        "# YOU CAN USE THIS FUNCTION IN NORMAL LOCAL COMPUTING, BUT NOT WORK IN THIS KAGGLE KERNEL ENVIRONMENT.\n",
        "# SO HERE I COMMENTED-OUT ALL THE SCRIPTS USING THIS FUNCTION.\n",
        "# -----------------------------------------------------------------------------\n",
        "saveShow_img <- function(data){\n",
        "  photo_url <- lapply(1:length(data$photos), function(x) data$photos[[x]] %>% strsplit(., split=\" \")) %>% unlist()\n",
        "  print(paste0(\"reading and saving \", length(photo_url), \" images\"))\n",
        "  sapply(1:length(photo_url), function(x) download.file(photo_url[x], paste0(\"data\",x,\".jpg\"), mode = \"wb\"))\n",
        "  print(paste0(\"showing \", length(photo_url), \" images\"))\n",
        "  sapply(1:length(photo_url), function(x){\n",
        "    data_img <- readJPEG(paste0(\"data\", x, \".jpg\"), native=TRUE)\n",
        "    dev.new()\n",
        "    plot(0:1, 0:1, type=\"n\", ann=FALSE, axes=FALSE)\n",
        "    rasterImage(data_img, 0, 0, 1, 1)})\n",
        "}"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "3214fa4d-5326-01ce-b3e4-a19f037bf10b"
      },
      "source": [
        "**Photos**  \n",
        "While data wrangling with **number of \"photos\"** by listing, I found that there are some data where **building_id and manager_id is different , but almost all photos are same .... price is different ...**, What is happening ...?\n",
        "\n",
        "Anyway, number of photos seem to contribute a little to increasing interest_level.  But too many photos is bad."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "aa2d0089-232e-f599-6422-62b46bf032a3"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"photos\"\n",
        "# - number of photos\n",
        "# -----------------------------------------------------------------------------\n",
        "# number of photos by each listing\n",
        "all$numOfPhotos <- sapply(1:nrow(all), function(x) length(unlist(all$photos[x])))\n",
        "train_org$numOfPhotos <- sapply(1:nrow(train_org), function(x) length(unlist(train_org$photos[x])))\n",
        "\n",
        "\n",
        "# Distribution of number of photos of TEST data is not very much different from that of TRAIN data\n",
        "# Mode is 5 and most listings have 3-8 photos\n",
        "# Maximum number of photos in TRAIN data is 68, in TEST data is 50\n",
        "\n",
        "# par(mfrow=c(3,1))\n",
        "# plot(table(all$numOfPhotos), xlim=c(0,70), ylim=c(0,20000), xlab=\"num of photos\", ylab=\"num of listings\", main=\"TRAIN+TEST: dist of num of photos\")\n",
        "# plot(table(all[1:nrow(train_org),]$numOfPhotos), xlim=c(0,70), ylim=c(0,20000), xlab=\"num of photos\", ylab=\"num of listings\", main=\"TRAIN: dist of num of photos\")\n",
        "# plot(table(all[1+nrow(train_org):nrow(all),]$numOfPhotos), xlim=c(0,70), ylim=c(0,20000), xlab=\"num of photos\", ylab=\"num of listings\", main=\"TEST: dist of num of photos\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "fb792cb5-eebc-393c-2a26-87fa30b28ef5"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"photos\"\n",
        "# - Does number of photos impact to interest level ?\n",
        "# -----------------------------------------------------------------------------\n",
        "# It seems the listings with zero photos has high probability to be \"LOW\" interest_level\n",
        "par(mfrow=c(3,1))\n",
        "all[1:nrow(train_org),] %>% filter(interest_level==\"high\") %>% select(numOfPhotos) %>% unlist() %>% hist(., breaks=seq(0,70,by=1), xlim=c(0,70), ylim=c(0,6000), xlab=\"num of photos\", ylab=\"num of listings\", main=\"HIGH listings\")\n",
        "all[1:nrow(train_org),] %>% filter(interest_level==\"medium\") %>% select(numOfPhotos) %>% unlist() %>% hist(., breaks=seq(0,70,by=1), xlim=c(0,70), ylim=c(0,6000), xlab=\"num of photos\", ylab=\"num of listings\", main=\"MEDIUM listings\")\n",
        "all[1:nrow(train_org),] %>% filter(interest_level==\"low\") %>% select(numOfPhotos) %>% unlist() %>% hist(., breaks=seq(0,70,by=1), xlim=c(0,70), ylim=c(0,6000), xlab=\"num of photos\", ylab=\"num of listings\", main=\"LOW listings\")\n",
        "\n",
        "# table of number of photos and number of listings by interest_level\n",
        "# numOfPhotos_tbl <- spread(data.frame(xtabs(data=all[1:nrow(train_org),], ~ interest_level + numOfPhotos)), key=interest_level, val=Freq) %>% arrange(numOfPhotos)\n",
        "numOfPhotos_tbl <- spread(data.frame(xtabs(data=train_org, ~ interest_level + numOfPhotos)), key=interest_level, val=Freq) %>% arrange(numOfPhotos)\n",
        "numOfPhotos_tbl <- numOfPhotos_tbl %>%\n",
        "  mutate(sum=low+medium+high) %>% mutate(lowRatio=low/sum, medRatio=medium/sum, highRatio=high/sum) %>%\n",
        "  mutate(medHigh=medium+high, medHighRatio=medHigh/sum)\n",
        "numOfPhotos_tbl"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "0723c300-13cc-d9cf-293f-914a383ed291"
      },
      "outputs": [],
      "source": [
        "# if number of photos = 0, high probability of \"LOW\" interest_level\n",
        "# if number of photos = 1 - 15, the ratio of \"Medium\" or \"High\" interest_level is more than 20%\n",
        "# if number of photos >= 16, it seems that ratio of \"LOW\" interest_level is increased.\n",
        "\n",
        "# par(mfrow=c(1,1))\n",
        "# plot(numOfPhotos_tbl$lowRatio, type=\"l\", ylim=c(0, 1.0), lty=2, col=\"red\", xlab=\"num of photos\", ylab=\"ratio of listings\", main=\"ratio of listings by num of photos\", xaxt=\"n\")\n",
        "# par(new=T); plot(numOfPhotos_tbl$medRatio, type=\"l\", ylim=c(0, 1.0), lty=2, col=\"blue\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(numOfPhotos_tbl$highRatio, type=\"l\", ylim=c(0, 1.0), lty=2, lwd=2, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(numOfPhotos_tbl$medHighRatio, type=\"l\", ylim=c(0, 1.0), lty=1, lwd=3, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "9d2c8f79-e110-bc9e-71ce-2429958fdc14"
      },
      "source": [
        "Red:  \"Low\" interest_level  \n",
        "Blue:  \"Medium\"     \n",
        "Black:  \"High\"     \n",
        "Black Thick:  \"Medium\" + \"High\"\n",
        "![][1]  \n",
        "[1]: https://www.kaggle.io/svf/1003625/2472063f1bf3f756e117193bf0ee369e/Rplot003.png"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "3b0099c7-6a89-7010-ea79-6f4993e898cc"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"photos\"\n",
        "# - Check photos of the listing with many photos >= 16\n",
        "# -----------------------------------------------------------------------------\n",
        "# Show the data of number of photos = 22\n",
        "# Can see some same managers, so sort by manager_id\n",
        "var <- c(\"numOfPhotos\", \"manager_id\", \"building_id\", \"interest_level\")\n",
        "tmp_ <- all %>% filter(numOfPhotos==22) %>% arrange(manager_id)\n",
        "tmp_[,var] %>% head(20)\n",
        "\n",
        "# Check the photos of some managers...\n",
        "# building_id and manager_is id different, but It seems those are almost same set of photos !!??\n",
        "# --- PLEASE UNCOMMENT ---------\n",
        "# ------------------------------\n",
        "# saveShow_img(data=tmp_[7,])\n",
        "# saveShow_img(data=tmp_[8,])\n",
        "# saveShow_img(data=tmp_[15,])\n",
        "# graphics.off()\n",
        "# ------------------------------"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "9915c307-d77b-2ea9-0a93-ecd422c873f6"
      },
      "source": [
        "**Bathrooms**  \n",
        "**What does mean \"bathrooms\" = 112 ??  Is it area ?  or number of bathrooms ?**   \n",
        "I checked photos for the data. **It does not seem that this have so many bathrooms** ... and price is also relatively cheap.  \n",
        "**How about the listing with bathrooms = 20 ??**\n",
        "\n",
        "Anyway, listing with bathrooms = 0 seem to have high probability to be \"low\" interest_level."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "389c3db7-eb90-3d35-2547-a59f1e5944c0"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"bathrooms\"\n",
        "# - number of bathrooms\n",
        "# -----------------------------------------------------------------------------\n",
        "# Most of listings have 1 bathroom\n",
        "# Maximum number of bathrooms in TRAIN data is 10, but in TEST data is 112 !!!\n",
        "table(all$bathrooms)\n",
        "\n",
        "summary(all$bathrooms)\n",
        "summary(train_org$bathrooms)\n",
        "summary(test$bathrooms)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "4d99f340-6a78-c272-84ab-f38138d7a207"
      },
      "outputs": [],
      "source": [
        "var <- c(\"bathrooms\", \"bedrooms\", \"numOfPhotos\", \"price\", \"interest_level\", \"features\")\n",
        "\n",
        "# Show the data of bathrooms = 112\n",
        "# Checked the photos of this listings, but it does not seem that this have so many bathrooms !! or so large bathroom !!\n",
        "# Price is also relatively cheap\n",
        "tmp_ <- all %>% filter(bathrooms==112)\n",
        "tmp_[,var]\n",
        "# --- PLEASE UNCOMMENT ---------\n",
        "# ------------------------------\n",
        "# saveShow_img(data=tmp_)\n",
        "# graphics.off()\n",
        "# ------------------------------"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "b916d1a8-b67b-434f-15ff-12cd278c5e6a"
      },
      "outputs": [],
      "source": [
        "# How about the data of bathrooms = 20 ?\n",
        "tmp_ <- all %>% filter(bathrooms==20)\n",
        "tmp_[,var]\n",
        "# --- PLEASE UNCOMMENT ---------\n",
        "# ------------------------------\n",
        "# saveShow_img(data=tmp_[1,])\n",
        "# graphics.off()\n",
        "\n",
        "# saveShow_img(data=tmp_[2,])\n",
        "# graphics.off()\n",
        "# ------------------------------"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ceb2213f-dcf6-9e54-20a2-b5564242b7d1"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"bathrooms\"\n",
        "# - Does \"bathrooms\" impact to interest level ?\n",
        "# -----------------------------------------------------------------------------\n",
        "# bathtooms and number of listings by interest_level\n",
        "# bathrooms_tbl <- spread(data.frame(xtabs(data=all[1:nrow(train_org),], ~ bathrooms + interest_level)), key=interest_level, val=Freq)\n",
        "bathrooms_tbl <- spread(data.frame(xtabs(data=train_org, ~ bathrooms + interest_level)), key=interest_level, val=Freq)\n",
        "bathrooms_tbl <- bathrooms_tbl %>%\n",
        "  mutate(sum=low+medium+high) %>% mutate(lowRatio=low/sum, medRatio=medium/sum, highRatio=high/sum) %>%\n",
        "  mutate(medHigh=medium+high, medHighRatio=medHigh/sum)\n",
        "bathrooms_tbl\n",
        "\n",
        "# if number of bathrooms = 0, high probability of \"LOW\" interest_level\n",
        "# if number of bathrooms >= 2.5, it seems that high ratio of \"LOW\" interest_level....\n",
        "# par(mfrow=c(1,1))\n",
        "# plot(bathrooms_tbl$lowRatio, type=\"l\", ylim=c(0, 1.0), lty=2, col=\"red\", xlab=\"bathrooms\", ylab=\"ratio of listings\", main=\"ratio of listings by bathrooms\", xaxt=\"n\")\n",
        "# par(new=T); plot(bathrooms_tbl$medRatio, type=\"l\", ylim=c(0, 1.0), lty=2, col=\"blue\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(bathrooms_tbl$highRatio, type=\"l\", ylim=c(0, 1.0), lty=2, lwd=2, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(bathrooms_tbl$medHighRatio, type=\"l\", ylim=c(0, 1.0), lty=1, lwd=3, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "14201a9a-3753-a5f4-09f6-8359d463313f"
      },
      "source": [
        "Red:  \"Low\" interest_level  \n",
        "Blue:  \"Medium\"     \n",
        "Black:   \"High\"     \n",
        "Black Thick:   \"Medium\" + \"High\"\n",
        "![][1]  \n",
        "[1]: https://www.kaggle.io/svf/1003625/2472063f1bf3f756e117193bf0ee369e/Rplot004.png"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "31931b33-e198-0262-4cdc-37b396830140"
      },
      "source": [
        "**Bedrooms**  \n",
        "Most of listings have 0 - 3 bedrooms. **But is it really zero bedroom** ?   \n",
        "I checked some data with bedrooms = 0, **I found some same managers**.  Need train them to input correctly the number of bedrooms ?\n",
        "\n",
        "Anyway, number of bedrooms (1 to 4) seem to contribute a little to increasing interest_level."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "dad5e498-0a38-fd50-605d-ce1550aa8630"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"bedrooms\"\n",
        "# - number of bedrooms\n",
        "# -----------------------------------------------------------------------------\n",
        "# Most of listings have 0-3 bedrooms\n",
        "# Maximum number of bathrooms in TRAIN data is 8, and in TEST data is 7\n",
        "table(all$bedrooms)\n",
        "\n",
        "summary(all$bedrooms)\n",
        "summary(train_org$bedrooms)\n",
        "summary(test$bedrooms)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "10b4f089-f6b7-01dc-37cd-e602712d7acb"
      },
      "outputs": [],
      "source": [
        "var <- c(\"bathrooms\", \"bedrooms\", \"numOfPhotos\", \"price\", \"interest_level\", \"features\")\n",
        "\n",
        "# Show the data of bedrooms = 8\n",
        "# first one has no photo, but second one has.\n",
        "# Second one is moderately expensive: price = 9995\n",
        "tmp_ <- all %>% filter(bedrooms==8)\n",
        "tmp_[,var]\n",
        "# --- PLEASE UNCOMMENT ---------\n",
        "# ------------------------------\n",
        "# saveShow_img(data=tmp_[2,])\n",
        "# graphics.off()\n",
        "# ------------------------------"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "1f313bb4-725b-c44e-d967-5d19af8368a1"
      },
      "outputs": [],
      "source": [
        "# # How about the data of bedrooms = 0 ?\n",
        "# Can see some same manager, so sort by manager_id\n",
        "# fourth one: really bedrooms = 0 ?\n",
        "tmp_ <- all %>% filter(bedrooms==0) %>% arrange(manager_id)\n",
        "tmp_[1:30,var]\n",
        "# --- PLEASE UNCOMMENT ---------\n",
        "# ------------------------------\n",
        "# saveShow_img(data=tmp_[4,])\n",
        "# graphics.off()\n",
        "# ------------------------------"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "fdc83453-a54c-8f7e-5d5a-356a80302a43"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"bedrooms\"\n",
        "# - Does \"bedrooms\" impact to interest level ?\n",
        "# -----------------------------------------------------------------------------\n",
        "# bathtooms and number of listings by interest_level\n",
        "# bedrooms_tbl <- spread(data.frame(xtabs(data=all[1:nrow(train_org),], ~ bedrooms + interest_level)), key=interest_level, val=Freq)\n",
        "bedrooms_tbl <- spread(data.frame(xtabs(data=train_org, ~ bedrooms + interest_level)), key=interest_level, val=Freq)\n",
        "bedrooms_tbl <- bedrooms_tbl %>%\n",
        "  mutate(sum=low+medium+high) %>% mutate(lowRatio=low/sum, medRatio=medium/sum, highRatio=high/sum) %>%\n",
        "  mutate(medHigh=medium+high, medHighRatio=medHigh/sum)\n",
        "bedrooms_tbl\n",
        "\n",
        "# if number of bedrooms >= 6, it seems that high ratio of \"LOW\" interest_level....\n",
        "# par(mfrow=c(1,1))\n",
        "# plot(bedrooms_tbl$lowRatio, type=\"l\", ylim=c(0, 1.0), lty=2, col=\"red\", xlab=\"bedrooms\", ylab=\"ratio of listings\", main=\"ratio of listings by bedrooms\", xaxt=\"n\")\n",
        "# par(new=T); plot(bedrooms_tbl$medRatio, type=\"l\", ylim=c(0, 1.0), lty=2, col=\"blue\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(bedrooms_tbl$highRatio, type=\"l\", ylim=c(0, 1.0), lty=2, lwd=2, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(bedrooms_tbl$medHighRatio, type=\"l\", ylim=c(0, 1.0), lty=1, lwd=3, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "52318fad-cd94-3c5d-f46e-996d12a5ecc9"
      },
      "source": [
        "Red:  \"Low\" interest_level  \n",
        "Blue:  \"Medium\"     \n",
        "Black:   \"High\"     \n",
        "Black Thick:   \"Medium\" + \"High\"\n",
        "![][1]  \n",
        "[1]: https://www.kaggle.io/svf/1003625/2472063f1bf3f756e117193bf0ee369e/Rplot005.png"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "c0ac637a-1b5d-0527-56b2-9fab0590a343"
      },
      "source": [
        ".. and **Price**  \n",
        "There is one data with price = 1.  \n",
        "I checked photos and it says \"seal\" and photos URL is \"http://sealagreements.com/leasing/\"  \n",
        "**It seems that price information is just entered for some other reasons.**  \n",
        "I also checkd the photos of the data with price <= 100. **This looks some sharing offices, or student condos.**  \n",
        "\n",
        "Anyway, some range of price seem to have relatively high probability to be \"medium\" or \"high\" ratio other than other price ranges."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "45f713a4-8055-2ec7-77f1-5c01dd6865f0"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"price\"\n",
        "# - Distribution of price and range\n",
        "# -----------------------------------------------------------------------------\n",
        "# Price has large range but Median is same in TRAIN and TEST data = 3150, Mean is also close.\n",
        "# Minimum price in TEST data is 1, maximum price in TRAIN is 4490000 !!!\n",
        "summary(all$price)\n",
        "summary(train_org$price)\n",
        "summary(test$price)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a51dfacc-37dc-1ac3-ee48-521a83e1320f"
      },
      "outputs": [],
      "source": [
        "var <- c(\"bathrooms\", \"bedrooms\", \"numOfPhotos\", \"price\", \"interest_level\", \"features\")\n",
        "\n",
        "# Show the data of price = 4490000\n",
        "# price is very expensive, but bedrooms = 2 and bathrooms = 1, and no photos\n",
        "tmp_ <- all %>% filter(price==4490000)\n",
        "tmp_[,var]"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "22ee6fb3-db75-b4cc-c8d3-720215c71ef0"
      },
      "outputs": [],
      "source": [
        "# How about the data of price = 1 ?\n",
        "# only 1 record, almost no information.\n",
        "# photos says \"seal\" and photos URL is \"http://sealagreements.com/leasing/\"\n",
        "# It seems that price information is just entered for some reasons\n",
        "tmp_ <- all %>% filter(price==1)\n",
        "tmp_[,var]\n",
        "# --- PLEASE UNCOMMENT ---------\n",
        "# ------------------------------\n",
        "# saveShow_img(data=tmp_[1,])\n",
        "# graphics.off()\n",
        "# ------------------------------"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "7e897f8b-90b7-9a15-efe7-a0a423dc6cd5"
      },
      "outputs": [],
      "source": [
        "# Check the price is price <= 100\n",
        "# 3 records including price = 1\n",
        "# Check the second one. This seems to be some (sharing) offices.\n",
        "tmp_ <- all %>% filter(price<=100)\n",
        "tmp_[,var]\n",
        "# saveShow_img(data=tmp_[2,])\n",
        "# graphics.off()"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "f56c900b-b818-05f8-2b52-986655466326"
      },
      "outputs": [],
      "source": [
        "# -----------------------------------------------------------------------------\n",
        "# Basics: \"price\"\n",
        "# - Does \"price\" impact to interest level ?\n",
        "# -----------------------------------------------------------------------------\n",
        "# Discretize price\n",
        "all$priceCat <- cut(all$price, breaks=c(0, 2^(5:23)))\n",
        "train_org$priceCat <- cut(train_org$price, breaks=c(0, 2^(5:23)))\n",
        "\n",
        "# bathtooms and number of listings by interest_level\n",
        "# priceCat_tbl <- spread(data.frame(xtabs(data=all[1:nrow(train_org),], ~ priceCat + interest_level)), key=interest_level, val=Freq)\n",
        "priceCat_tbl <- spread(data.frame(xtabs(data=train_org, ~ priceCat + interest_level)), key=interest_level, val=Freq)\n",
        "priceCat_tbl <- priceCat_tbl %>%\n",
        "  mutate(sum=low+medium+high) %>% mutate(lowRatio=low/sum, medRatio=medium/sum, highRatio=high/sum) %>%\n",
        "  mutate(medHigh=medium+high, medHighRatio=medHigh/sum)\n",
        "priceCat_tbl\n",
        "\n",
        "# if price between 512 and 2048 have high probability to be \"High\" interest_level\n",
        "# if price betwenn 2048 and 8192 have high probability to be \"Medium\" interest_level\n",
        "# par(mfrow=c(1,1))\n",
        "# plot(priceCat_tbl$lowRatio, type=\"b\", ylim=c(0, 1.0), lty=2, col=\"red\", xlab=\"priceCat\", ylab=\"ratio of listings\", main=\"ratio of listings by priceCat\", xaxt=\"n\")\n",
        "# par(new=T); plot(priceCat_tbl$medRatio, type=\"b\", ylim=c(0, 1.0), lty=2, col=\"blue\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(priceCat_tbl$highRatio, type=\"b\", ylim=c(0, 1.0), lty=2, lwd=2, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");\n",
        "# par(new=T); plot(priceCat_tbl$medHighRatio, type=\"b\", ylim=c(0, 1.0), lty=1, lwd=3, col=\"black\", xlab=\"\", ylab=\"\", xaxt=\"n\", yaxt=\"n\");"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "377ce149-4eda-cae8-4e13-64b7481fa2c2"
      },
      "source": [
        "Red:  \"Low\" interest_level  \n",
        "Blue:  \"Medium\"     \n",
        "Black:  \"High\"     \n",
        "Black Thick:  \"Medium\" + \"High\"\n",
        "![][1]  \n",
        "[1]: https://www.kaggle.io/svf/1003625/2472063f1bf3f756e117193bf0ee369e/Rplot006.png"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "R",
      "language": "R",
      "name": "ir"
    },
    "language_info": {
      "codemirror_mode": "r",
      "file_extension": ".r",
      "mimetype": "text/x-r-source",
      "name": "R",
      "pygments_lexer": "r",
      "version": "3.3.3"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}