---
title: "Jaccard Index of Baseline on Training Data and Public LB"
output: html_document
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE, message=FALSE, warning=FALSE)
library(rgeos)
library(dplyr)
library(readr)
train.path <- '../input/train_wkt_v4.csv'
grid.path <- '../input/grid_sizes.csv'
```

## Training Data

First, we compute the average Jaccard index over the training data 
for the 'sample submission model'.
That is, we predict all of the area of each image for each class.
First, we need the area of each image:

```{r prep.grid}
# Compute whole image area
gr <- read_csv(grid.path, col_types = 'cdd')
names(gr) <- c('ImageId', 'Xmax', 'Ymin')
gr$GridArea <- abs(gr$Xmax * gr$Ymin)
gr <- select(gr, ImageId, GridArea)
```

Next we need the area of each of the training polygons.
We apply `rgeos::gArea` and `rgeos::readWKT` to all of the polygon text in train to get their areas.

```{r prep.train}
# Compute ground truth area for each training class/image
tr <- read_csv(train.path, col_types='cic')
tr$TrainArea <- sapply(tr$MultipolygonWKT, 
                       function(x) ifelse(x=='MULTIPOLYGON EMPTY', 
                                          0, gArea(readWKT(x))))
tr <- select(tr, ImageId, ClassType, TrainArea)
```

We combine the data and group by class. 
These are the Jaccard index values by class.

```{r combine}
areas <- inner_join(tr, gr) %>% group_by(ClassType)
out <- summarise(areas, trArea=sum(TrainArea), 
                 grArea=sum(GridArea), jacIdx=trArea / grArea)
out
```

Class 5 is trees, and this model gets about 0.1. 
Class 6 is crops and it gets about 0.27.
Everything else is pretty low.
For these full-image predictions, the Jaccard index for a class is equal to the fraction of the image areas occupied by that class.
So trees fill about 10% of the training images and crops fill about 27%.
Finally, we take the mean over the classes.

```{r result}
c(AvgJacIdx=mean(out$jacIdx))
```

So the average Jaccard index is 0.04624. Note that trees and crops account for around 0.038 by themselves. 
That score is a fair bit lower than what we get on the public test set (0.07524). 
And *that* suggests that the training data is not that representative of the test set.

## Public LB

I made a sequence of submissions using the sample submission, but with one class zeroed out.
One can show that the area fraction of the removed class over the images that make up the public LB 
is 10 times the score delta between these submissions and the sample submission.
Here is what those results look like:

| Class Removed | LB Score | Delta  | Coverage |
|:-------------:|---------:|-------:|---------:|
| none          | 0.07524  |   0    |    x     |
| 1             | 0.07492  |0.00032 |  0.0032  |
| 2             | 0.07500  | 0.00024|  0.0024  |
| 3             | 0.07483  | 0.00041| 0.0041   |
| 4             | 0.07324  | 0.00200| 0.0200   |
| 5             | 0.06963  | 0.00561|  0.0561  |
| 6             | 0.01523  | 0.06001|  0.6001  |
| 7             |  0.06894 | 0.00630|  0.0630  |
| 8             | 0.07488  | 0.00036|  0.0036  |
| 9             | 0.07523  | 0.00001|  0.0001  |
| 10            | 0.07524  | 0.00000| 0.0000   |

So on the public LB, trees cover about 5.6% of the image area, down from about 10% on train.
Crops cover about 60%, up from 27% on train. Waterways are at 6%, up from 0.5% on train.
Really, this shows that the public LB images and the training images are pretty different.





