---
title: 'MLB Player Digital Engagement Forecasting'
author: 'Long Le'
date: '`r Sys.Date()`'
output:
  html_document:
    number_sections: false
    toc: true
    toc_depth: 2
    toc_float: false
    fig_caption: true
    fig_width: 7
    fig_height: 4.5
    fig_align: center
    theme: cosmo
    highlight: tango
    code_folding: show
    df_print: paged
editor_options: 
  markdown: 
    wrap: sentence
---

Hi reader, this is the notebook that I made to share how I unpacked the train.csv data. 
I had a bit of problem with this data format at first (still does :D) so data processing is likely not optimized.
Feel free to leave a comment if you have suggestions to optimize the process or have a different approach you want to share. I would be very happy to check it out!

# Preparations

This chapter includes loading in required packages and importing raw data, we will do data processing in a later chapter.

## Importing packages



```{r Libraries, message=F}
# Data import
library(readr)
library(jsonlite)

# Data manipulation
library(dplyr)
library(tidyr)

# Graph
library(ggplot2)
library(lattice)
library(plotly)

# Model building
library(caret)
library(tseries)

# Mass loading lib
library(pacman)
p_load(tidyverse, fs, vroom, glue, janitor, lubridate, ggridges, viridis) # Neat function, it will load all the packages we specified. If the packages is not yet installed in your Rstudio it will download it automatically.
```

## Data import

```{r Data import}
# This is an very nice approach to data importing that I saw on Ringa_hyj's notebook 
# check it out here: https://www.kaggle.com/hiroshihiroshi/mlb-engagement-first-load-and-check-data
file_list <- dir_info("../input/mlb-player-digital-engagement-forecasting") %>% 
  select(path,type,size)

file_path <- file_list %>% filter(type=="file") %>% pull(path)

train <- read_csv(file_path[7],col_types = cols(date = col_date(format = "%Y%m%d"))) %>% clean_names()
example_sample_submission <- read_csv(file_path[2],col_types = cols()) %>% clean_names()
example_test <- read_csv(file_path[3],col_types = cols()) %>% clean_names()
```

# Expanding datasets from train.csv

```{r glimpse train}
str(train)
```

The dataset train.csv is in nested JSON format and require "unpacking" to get data from it. 

Here we separate train.csv to 11 different datasets. Although the sets are not yet usable since they are still in JSON format. Na.omit is required here since in my testing, the fromJSON function that do the JSON unpacking does not work with NA values.

```{r Define dataframes}
next_day_player_engagement <- na.omit(train$next_day_player_engagement)
games <- na.omit(train$games)
rosters <- na.omit(train$rosters)
player_box_scores <- na.omit(train$player_box_scores)
team_box_scores <- na.omit(train$team_box_scores)
transactions <- na.omit(train$transactions)
standings <- na.omit(train$standings)
awards <- na.omit(train$awards)
events <- na.omit(train$events)
player_twitter_followers <- na.omit(train$player_twitter_followers)
team_twitter_followers <- na.omit(train$team_twitter_followers)
```

Here we define the function used to unpack the datasets to its tidy form.

```{r data unnest fuction}
unnest <- function(df){
    new <- fromJSON(df[1])

    for(i in 2:length(df)){
    new.2 <- fromJSON(df[i])
    new <- rbind(new,new.2)
    }
    return(new)
}

```

## The next_day_player_engagement dataset

```{r unnest1}
# expand next_day_player_engagement data
next_day_player_engagement.df <- unnest(next_day_player_engagement)
str(next_day_player_engagement.df)
```

## The games dataset

```{r unnest2}
# expand games data
games.df <- unnest(games)
str(games.df)
```

## The rosters dataset

```{r unnest3}
# expand rosters data
rosters.df <- unnest(rosters)
str(rosters.df)
```

## The player_box_scores dataset

```{r unnest4}
# expand player_box_scores data
player_box_scores.df <- unnest(player_box_scores)
str(player_box_scores.df)
```

## The team_box_scores dataset

```{r unnest5}
# expand team_box_scores data
team_box_scores.df <- unnest(team_box_scores)
str(team_box_scores.df)
```

## The transactions dataset

```{r unnest6}
# expand transactions data
transactions.df <- unnest(transactions)
str(transactions.df)
```

## The standings dataset

```{r unnest7}
# expand standings data
standings.df <- unnest(standings)
str(standings.df)
```

## The awards dataset

```{r unnest8}
# expand awards data
awards.df <- unnest(awards)
str(awards.df)
```

## The events dataset

```{r unnest9}
# expand events data
events.df <- unnest(events)
str(events.df)
```

## The player_twitter_followers dataset

```{r unnest10}
# expand player_twitter_followers data
player_twitter_followers.df <- unnest(player_twitter_followers)
str(player_twitter_followers.df)
```

## The team_twitter_followers dataset

```{r unnest11}
# expand team_twitter_followers data
team_twitter_followers.df <- unnest(team_twitter_followers)
str(team_twitter_followers.df)
```































