{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I'm going to try use R and participate in this competition over the coming weeks.  Initially my focus will be to replicate some of the good work on the python side, and the pioneering work on the R side.  I've borrowed liberally from Long Le (especially https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r), and others.  I welcome feedback on breaches in protocol!","metadata":{}},{"cell_type":"code","source":"# I like the way Long Le organizes package import and the p_load function (from the pacman package)\n\n# Data import\nlibrary(readr)\nlibrary(jsonlite)\n\n# Data manipulation\nlibrary(dplyr)\nlibrary(tidyr)\n\n# Graph\nlibrary(ggplot2)\nlibrary(lattice)\nlibrary(plotly)\n\n# Model building\nlibrary(caret)\nlibrary(tseries)\n\n# Mass loading lib\nlibrary(pacman) # Lots of nice package manipulation functions\np_load(tidyverse, fs, vroom, glue, janitor, lubridate, ggridges, viridis) # Part of the pacman package","metadata":{"execution":{"iopub.status.busy":"2021-07-12T19:58:37.582048Z","iopub.execute_input":"2021-07-12T19:58:37.584276Z","iopub.status.idle":"2021-07-12T19:58:41.484906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Long Le's method along with his acknowledgements, and my in-line commentary.  (SB:  Wow, I'm realizing this is like the development of the Talmud!)\n# This is an very nice approach to data importing that I saw on Ringa_hyj's notebook \n# check it out here: https://www.kaggle.com/hiroshihiroshi/mlb-engagement-first-load-and-check-data\nfile_list <- dir_info(\"../input/mlb-player-digital-engagement-forecasting\") %>% \n  select(path,type,size)\n\nfile_path <- file_list %>% filter(type==\"file\") %>% pull(path)  #pull() seems similar to select(<single column>)\n\n# Useful to use 'read_csv' rather than 'read.csv' as is my habit.  Good commentary at https://medium.com/r-tutorials/r-functions-daily-read-csv-3c418c25cba4\n# The below is still JSON.  It is a clever way to separate the 'train' dataset from the other files in the directory\ntrain <- read_csv(file_path[7],col_types = cols(date = col_date(format = \"%Y%m%d\"))) %>% clean_names()\n# Not yet sure what we are accomplishing here except that we are pulling out the example sample submission file.\nexample_sample_submission <- read_csv(file_path[2],col_types = cols()) %>% clean_names()\n# Not yet sure what we are accomplishing here except that we are pulling out the example test file.\nexample_test <- read_csv(file_path[3],col_types = cols()) %>% clean_names()","metadata":{"execution":{"iopub.status.busy":"2021-07-12T19:58:41.487352Z","iopub.execute_input":"2021-07-12T19:58:41.528339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Learn about the 'train' dataset.  This is interesting because essentially, \n# most cells in the data set are themselves giant json files.  str below helps us know the different components of the\n# JSON\nstr(train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Again, borrowed whole-cloth from Long Je.  These are still JSON!\nnext_day_player_engagement <- na.omit(train$next_day_player_engagement)\ngames <- na.omit(train$games)\nrosters <- na.omit(train$rosters)\nplayer_box_scores <- na.omit(train$player_box_scores)\nteam_box_scores <- na.omit(train$team_box_scores)\ntransactions <- na.omit(train$transactions)\nstandings <- na.omit(train$standings)\nawards <- na.omit(train$awards)\nevents <- na.omit(train$events)\nplayer_twitter_followers <- na.omit(train$player_twitter_followers)\nteam_twitter_followers <- na.omit(train$team_twitter_followers)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is a thing of beauty from Long Je. It writes the function that unpacks the JSON and writes it into a \n# more familiar dataframe.  This will be used later as we dig into each of the component datasets. I guess you start\n# at line 2 in order to skip the names row?\nunnest <- function(df){\n    new <- fromJSON(df[1])\n\n    for(i in 2:length(df)){\n    new.2 <- fromJSON(df[i])\n    new <- rbind(new,new.2)\n    }\n    return(new)\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# expand next_day_player_engagement data.  \nnext_day_player_engagement.df <- unnest(next_day_player_engagement)\nstr(next_day_player_engagement.df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For fun I will include all of the datasets in the following sections.  I'm interested in knowing how long it will take\n# to load and run this entire thing in Kaggle Notebooks.  In other words, should I do the work somewhere else, \n# then publish it here?\n\ngames.df <- unnest(games)\nrosters.df <- unnest(rosters)\nplayer_box_scores.df <- unnest(player_box_scores)\nteam_box_scores.df <- unnest(team_box_scores)\ntransactions.df <- unnest(transactions)\nstandings.df <- unnest(standings)\nawards.df <- unnest(awards)\nevents.df <- unnest(events)\nplayer_twitter_followers.df <- unnest(player_twitter_followers)\nteam_twitter_followers.df <- unnest(team_twitter_followers)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}