---
title: "PLAsTiCC EDA"
author: "Mitchell O'Brien"
date: "October 8, 2018"
output:
  html_document:
    number_sections: true
    toc: true
    fig_width: 7
    fig_height: 4.5
    theme: cosmo
    highlight: tango
    code_folding: hide
---

**If you find this kernel interesting or you learned something from it, I would really appreciate an upvote. Additionally, if you have any comments, I would really like the constructive criticism. Thank you!**


```{r, eval=TRUE, include=FALSE}
library(readr)
library(dplyr)
library(ggplot2)
library(gridExtra)


train<-data.frame(read_csv("../input/training_set_metadata.csv"))
test<-data.frame(read_csv("../input/test_set_metadata.csv"))

#submission is at the object_id level
train<-data.frame(apply(train, 2, function(x){ifelse(x=="NaN", NA, x)}))
test<-data.frame(apply(test, 2, function(x){ifelse(x=="NaN", NA, x)}))

train$target<-as.factor(train$target)
```

#Introduction
I have been wanting to work on this submission for awhile now. Unfortunately, the Google Analytics competition has been taking up most of my time. Needing a break from the Google Analytics competition, I decided I would create and EDA focused on the difference between the train and test set. The last few competitions I have participated in have had very different train and test sets, so I thought I would get ahead of it in this one!

#Target Analysis
```{r}
ggplot(train, aes(target))+
  geom_bar(fill = "blue", col = "black")+
  geom_text(stat = "count", aes(label=..count..), vjust = -.5, fontface = "bold")+
  ggtitle("Count of Target")+
  theme(plot.title = element_text(hjust = .5))


ggplot(train, aes(gal_l, gal_b, col = target))+
  geom_point()+
  facet_wrap(~target)+
  ggtitle("Latitude and Longitude by Target")+
  theme(plot.title = element_text(hjust = .5))

ggplot(train, aes(ra, decl, col = target))+
  geom_point()+
  facet_wrap(~target)+
  ggtitle("Right Ascension and Declination by Target")+
  theme(plot.title = element_text(hjust = .5))
  
```

Before I analyze differences between train and test set, I wanted to investigate the target variable a bit. In classification problems, it is good to know the counts of classes in the target variable that you will eventually predict. Having severely unbalanced classes can lead to issues with machine learning algorithms identifying instances of minority classes in the test set. Additionally, I wanted to look at scatter plots of Galactic Latitude and Galactic Longitude, as well as Right Ascension and Declination to see if there is any apparent pattern.

\vspace{12pt}

What I find:

* There is certainly some class imbalance. For example, target class 53 has 30 observations in the training set while 90 has 2313. In binary classification problems, I have used up-sampling with great success. I do not know of a generalization to multi-classification problems, but, it would be worth looking into

* I do not see any clear pattern differentiating classes by latitude and longitude. However, there does appear to be some classes with sparse observations from latitude 100 to 200. For example, target class 16 has very few observations from 100 to about 150 latitude. Compare this to target class 42 and there is a clear difference. Note they have similar counts

* In some classes, there appears to be V-shape of sparseness. Examples of this are target class 15, 62, and 90. Other classes only have sparseness on the right side of the V. As a matter of fact, class 16 is dense on the left side of the V where as some classes are very sparse. 

#Train vs Test
##Null Values
```{r}
sort(apply(train, 2, function(x){sum(is.na(x))/length(x)}),decreasing = TRUE)
sort(apply(test, 2, function(x){sum(is.na(x))/length(x)}), decreasing = TRUE)
```

Generating percentages of missing values from each column reveals the following:

* In the train set, the only variable with missing values is distmod
* In the test set, we are missing hostgal_specz and distmod
* hostgal_specz may be a big loss when modeling since it is described as very accurate measure of red shift

##Hostgal_specz
```{r}
train_hist_specz<-ggplot(train, aes(hostgal_specz))+
  geom_histogram(fill = "blue")+
  ggtitle("Training Distribution of hostgal_specz")+
  theme(plot.title = element_text(hjust = .5))+
  scale_x_continuous(limits=c(0, 3))
  

test_hist_specz<-ggplot(na.omit(test), aes(hostgal_specz))+
  geom_histogram(fill = "red")+
  ggtitle("Test Distribution of hostgal_specz")+
  theme(plot.title = element_text(hjust = .5))+
  scale_x_continuous(limits=c(0, 3))

grid.arrange(train_hist_specz, test_hist_specz)

```

Hostgal_specz is defined as  the spectroscopic red shift of the source. This is an extremely accurate measure of red shift, available for the training set and a small fraction of the test set. Float32. To visualize the differences in distributions, I created histograms for training and test set. I find: 

* There appears to be noticeable differences between train and test set
* Train set is right skewed with a center of roughly .2
* Test set is still a little right skewed though not as much as the train set
* The test set distribution is shifted further right than train set with a center of roughly .3

##Hostgal_photoz

Hostgal_photoz is defined as The photometric red shift of the host galaxy of the astronomical source. While this is meant to be a proxy for hostgal_specz, there can be large differences between the two and should be regarded as a far less accurate version of hostgal_specz. Float32. 

\vspace{12pt}

Due to the correlation between hostgal_photoz and hostgal_specz, I wanted to see how similar they are. Although hostgal_photoz is a less accurate version of hostgal_specz, we would hope that they are similar. I will now analyze the differences between hostgal_specz and hostgal_photoz for train and test set.

###Train Difference hostgal_specz and hostgal_photoz
```{r}
train_hist_photoz<-ggplot(train, aes(hostgal_photoz))+
  geom_histogram(fill = "blue")+
  ggtitle("Training Distribution of hostgal_photoz")+
  theme(plot.title = element_text(hjust = .5))

grid.arrange(train_hist_specz, train_hist_photoz)

myText<-paste("r =", round(cor(train$hostgal_photoz, train$hostgal_specz, method = "pearson"),2))

ggplot(train, aes(hostgal_specz, hostgal_photoz))+
  geom_point()+
  ggtitle("hostgal_specz vs hostgal_photoz")+
  theme(plot.title = element_text(hjust = .5))+
  annotate("text", label = myText, x = 3, y = 1, size = 5, col = "red")

```

As stated earlier, I would hope that the distributions are similar for hostgal_photoz and hostgal_specz. Additionally, I would hope that they are strongly linearly correlated. I find:

* The distributions have the same general shape. hostgal_photoz has a more elongated tail and a cluster around 0

* Based on the scatter plot, there is certainly linear correlation. However, the cluster in the upper left hand corner disturbs the Pearson correlation coefficient. This results in a lower Pearson correlation coefficient than evidenced by the plot. 

###Train Difference hostgal_specz and hostgal_photoz
```{r}
test_hist_photoz<-ggplot(na.omit(test), aes(hostgal_photoz))+
  geom_histogram(fill = "red")+
  ggtitle("Test Distribution of hostgal_photoz")+
  theme(plot.title = element_text(hjust = .5))


grid.arrange(test_hist_specz, test_hist_photoz)

myText<-paste("r =", round(cor(na.omit(test)$hostgal_specz, na.omit(test)$hostgal_photoz),2))

ggplot(na.omit(test), aes(hostgal_specz, hostgal_photoz))+
  geom_point()+
  ggtitle("hostgal_specz vs hostgal_photoz")+
  theme(plot.title = element_text(hjust = .5))+
  annotate("text", label = myText, x = 1, y = 2, size = 5, col = "red")


```

The test set presents an interesting issue. Most of the values for hostgal_specz are gone. Thus, it is important that we feel confident in the ability of hostgal_photoz to accurately represent the red shift.  

\vspace{12pt}

Since so many values are missing, comparisons are tough. However, the test set is very large, so I filtered for missing values and made my plots/comparisons. I find that:  

* The distributions for hostgal_specz and hostgal_photoz are similar in shape
* hostgal_photoz distribution has a more elongated tail with a slightly right shifted center than hostgal_specz
* The scatter plot shows a linear association. However, the cluster in the upper hand corner is bigger than the one in the training set. We see a low Pearson correlation as a result.
* It seems that in the test set, hostgal_photoz can seriously over estimate red shift

###Train vs Test hostgal_photoz
```{r}
grid.arrange(train_hist_photoz, test_hist_photoz)
```

Now to analyze the difference between train and test set for hostgal_photoz. I find that: 

* The train set distribution has right skew. But, a concentration of samples around 0

* The test set is also right skewed. But, without the concentration of samples at 0

* The test set distribution is centered further right than the training set

**More to come soon!**
