---
title: "Exploring Data"
author: "Orange81"
date: "7 november 2015"
output: html_document
---

The goal of this KAGGLE competion is to predict the category of the crime 
incident in the test dataset using information from the training set. 

The train dataset contains incidents derived from the SFPD Crime incident 
Reporting system. The data ranges from 1/1/2003 to 5/13/2015. The training set 
and the test set rotate every week, meaning week 1, 3, 5, 7, ... belong to the 
testset week 2, 4, 6, 8,... belong to the training set.

The libraries needed
```{r}
library(dplyr)
library(ggplot2)
```

Load the data
```{r}
train <- read.csv("../input/train.csv", header = T)
test  <- read.csv("../input/test.csv", header = T)
train <- tbl_df(train)
test  <- tbl_df(test)
```

Lets have a first look at both the train and the test dataset.
```{r}
train
test
```

The train set consists of 878049 observations of 9 variables. These 9 variables
are **Dates**, **Category**, **Descript**, **DayOfWeek**, **PdDistric**, 
**Resolution**, **Address**, **X** and **Y**.

The test set consists of 884262 observations of 7 variables. These 7 variables
are **Id**, **Dates**, **DayOfWeek**, **PdDistric**, **Address**, **X** and 
**Y**.

The variable **Category** needs to be predicted in the test set. The variables 
**Descript** and **Resolution** can only be known when the variable **Category** 
is known therefore not be used to predict the variable **Category**. These 
variables are also not present in the test set. 

As shown in the figure below, the variable **Category** contains 39 
different categories, of which *LARENCY/THEFT* is the most common and *TREA* is 
the least common.

```{r, echo=FALSE}
category <- as.data.frame(table(train$Category))
names(category) <- c("Category", "Frequency")
attach(category)
nd <- category[order(-Frequency),]
barplot(nd$Frequency, names.arg=nd$Category, las=2, cex.names=0.8,
        main="Number of incidents per Category",
        ylim=c(0,200000))
```

The variable **Dates** is present in both the training and the test set and
can therefore potentially be used to predict the variable **Category**. The
variable **Dates** contains information about the date and the time the incident
happened. Therefore I split the variable **Dates** into the variable **Date**
representing the date and the variable **Time** representing the time the 
incident happened.

```{r}
#convert the variable "Dates" to class Date
train$Dates <- as.Date(train$Dates)
test$Dates  <- as.Date(test$Dates)
train$Time <- format(as.POSIXct(train$Dates, format="%Y-%m-%d %H:%M:%S"), format="%H:%M:%S")
test$Time  <- format(as.POSIXct(test$Dates, format="%Y-%m-%d %H:%M:%S"), format="%H:%M:%S")
```

As shown below, the training set starts at the sixth of January 2003 (a monday 
and the second week of the year) and ends at 13 may 2015 (a wednesday and the 
20th week of the year). The test set starts art the first of January 2003 (a 
wednesday and the first week of the year) and ends at 10 may 2015 (a sunday and 
the 19th week of the year).

```{r}
summary(train$Date)
summary(test$Date)
```

It was mentioned that the trainings set only contains the even weeks and that the
test set contains only the uneven weeks. Lets check that assumption.

```{r}
#get the weeknumbers. Week of the year as decimal number (01–53) as defined in 
#ISO 8601.If the week (starting on Monday) containing 1 January has four or more 
#days in the new year, then it is considered week 1. Otherwise, it is the last 
#week of the previous year, and the next week is week 1. 
train$WeekNumber <-  strftime(as.POSIXlt(train$Date),format="%V") 
test$WeekNumber <-  strftime(as.POSIXlt(test$Date),format="%V") 
table(train$WeekNumber)
table(test$WeekNumber)
```

As shown above, the training set contains the even weeknumbers and the test 
the uneven weeknumbers.

The variable **Descript** gives a more detailed description of the variable
"Category" and contains 879 different levels. This variable cannot
be used to predict Category since it is not present in the test set. 

```{r}
length(unique(train$Descript))
```

The variable **DayOfWeek** is present in both the training and the test set and 
can thereforebe used to potentially predict the variable "Category".

As shown in the figure below, most incidents happen on a friday, on sunday 
the least incidents happen in both the training and the test set. 
When comparing the train and the test set it is observed that in the training 
set more incidents happen on Thursday and in the test set more incident happen 
on a Tuesday.

```{r, echo=FALSE}
#training set
dayofweek <- as.data.frame(table(train$DayOfWeek))
names(dayofweek) <- c("DayOfWeek", "Frequency")
attach(dayofweek)
nd <- dayofweek[order(-Frequency),]
barplot(nd$Frequency, names.arg=nd$DayOfWeek, las=2, cex.names=0.8,
        main="Number of incidents per weekday in the training set",
        ylim=c(0,140000))

#test set
dayofweek <- as.data.frame(table(test$DayOfWeek))
names(dayofweek) <- c("DayOfWeek", "Frequency")
attach(dayofweek)
nd <- dayofweek[order(-Frequency),]
barplot(nd$Frequency, names.arg=nd$DayOfWeek, las=2, cex.names=0.8,
        main="Number of incidents per weekday in the test set",
        ylim=c(0,140000))

```

The variable **PdDistrict** is present in both the training and the test set 
and can therefore potentially be used to predict the Category.

The variable PdDistrict contains 10 different districts. As shown in the 
figure below, the most incidents happen in the Southern PdDistrict and the 
least incidents happen in the Richmond PdDistrict for both the training and 
the test set. Also both graphs look similar for both the test and the training 
set.


```{r, echo=FALSE}
#training set
pddistrict <- as.data.frame(table(train$PdDistrict))
names(pddistrict) <- c("PdDistrict", "Frequency")
attach(pddistrict)
nd <- pddistrict[order(-Frequency),]
barplot(nd$Frequency, names.arg=nd$PdDistrict, las=2, cex.names=0.8,
        main="Number of incidents per PdDistrict in the training set",
        ylim=c(0,200000))

#test set
pddistrict <- as.data.frame(table(test$PdDistrict))
names(pddistrict) <- c("PdDistrict", "Frequency")
attach(pddistrict)
nd <- pddistrict[order(-Frequency),]
barplot(nd$Frequency, names.arg=nd$PdDistrict, las=2, cex.names=0.8,
        main="Number of incidents per PdDistrict in the test set",
        ylim=c(0,200000))

```

The variable **Resolution** is only present in the train set and therefore 
cannot be used to predict the Category in the test set.

This variable **Address** contains the exact address were the incident 
happened and is present in both the training and the test set. This variable
contains 23228 different levels in the training set and 23184 levels for the test
set. At this point I do not know how to use this data to predict the Category.

```{r}
length(unique(train$Address))
length(unique(test$Address))
```

The variable **X** is the first part of the X-Y coordinate and gives 
information about the location were the incident happened. This variable is 
present in both the training and the test set.

The variable **Y** is the second part of the X-Y coordinate and gives information about
the location were the incident happened. This variable is present in both the 
training and the test set and can therefore potentially be used to predict the
category.

