{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "adece2e1-0bd9-abfa-d054-260b4a8b85eb"
      },
      "source": [
        "**Plotting Elo Ratings for each team from 1985 to 2016.** \n",
        "\n",
        "*Note: If you'd like to use these Elo ratings in your March Madness predictions, you can simply download the output file. The output will provide Elo ratings at the end of each regular season in a melted format for joining/merging.* "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "5a72f91e-337d-cf5b-1470-3ff307944ca6"
      },
      "outputs": [],
      "source": [
        "# contains R code to:\n",
        "# -read in Kaggle's March Machine Learning Challenge data csv files\n",
        "# -clean, analyze and standardize dataset\n",
        "# -implement Elo ratings system package \n",
        "# -plot the Elo ratings history time series for selected teams\n",
        "\n",
        "# read data\n",
        "teams.df <- read.csv(\"../input/Teams.csv\", header = TRUE)\n",
        "seasons.df <- read.csv(\"../input/Seasons.csv\", header = TRUE)\n",
        "rcr.df <- read.csv(\"../input/RegularSeasonCompactResults.csv\", header = TRUE)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "48f9aa2c-4e93-297b-c03c-e0fb205f00a1"
      },
      "source": [
        "**Load libraries.** Core package is the [PlayerRatings][1] package. Provides an implementation for computing Elo ratings. Also includes rating implementations for Fide, Glicko and Stephenson as alternatives. \n",
        "\n",
        "\n",
        "  [1]: https://cran.r-project.org/web/packages/PlayerRatings/PlayerRatings.pdf"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "d7c13461-e02b-36b1-de3f-b3fa24117028"
      },
      "outputs": [],
      "source": [
        "# load libraries\n",
        "suppressPackageStartupMessages({\n",
        "    library(dplyr)  # set for data wrangling\n",
        "    library(reshape2)  # set for melting data\n",
        "    library(PlayerRatings)  # set for ELO ranking system\n",
        "    library(ggplot2)  # set for plotting\n",
        "})"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "d7d62bcd-edb2-fe4d-e0fa-9b5e8529eb74"
      },
      "source": [
        "**Adjust data to proper form.** [PlayerRatings'][1] implementation of Elo has specific argument structure for input variables; requires dataframe containing vectors for Time, Player 1, Player 2, and Player 1 Win/Loss.  Code below establishes Elo dataframe and sets gamma vector for home court advantage.\n",
        "\n",
        "\n",
        "  [1]: https://cran.r-project.org/web/packages/PlayerRatings/PlayerRatings.pdf"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "809a630e-b2e4-b9f3-f6e2-6bc7876fbe52"
      },
      "outputs": [],
      "source": [
        "##############################\n",
        "# Wrangle data into proper form\n",
        "##############################\n",
        "\n",
        "# create dataframe to feed Elo package\n",
        "regSeason.df <- rcr.df %>% \n",
        "                select(Season, Daynum, Wteam, Lteam) %>%  # select necessary variables\n",
        "                mutate(Outcome = rep(1, nrow(rcr.df)))  # add Outcome vector\n",
        "\n",
        "# join Dayzero value and create Date index for inter-season time series\n",
        "regSeason.df <- left_join(regSeason.df, \n",
        "                          select(seasons.df, Season, Dayzero),  # only join Dayzero value\n",
        "                          by = c(\"Season\" = \"Season\"))  \n",
        "\n",
        "regSeason.df <- regSeason.df %>% \n",
        "                mutate(Date = as.Date(regSeason.df$Dayzero, format = \"%m/%d/%Y\") + \n",
        "                                      regSeason.df$Daynum) %>%\n",
        "                mutate(Date = as.numeric(Date)) %>%  # convert to numeric class for Elo \n",
        "                select(Season, Date, Wteam, Lteam, Outcome)  # select variables for Elo\n",
        "\n",
        "# set gamma (home court advantage) vector\n",
        "gammaVect <- as.matrix(as.numeric(rcr.df$Wloc == \"H\"))  "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "c24f1641-d397-40f4-caa6-f4c467758182"
      },
      "source": [
        "**Implement Elo computation.** Set arbitrary initialization at 1500, and start initialization in 1985 season. Set gamma vector to account for home court advantage, and set k-factor = 20. Note that FiveThrityEight found optimal K-factor for NBA games to be 20, which was higher than they expected."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "fdb27ec7-7b1b-f678-6324-c661d598daf1"
      },
      "outputs": [],
      "source": [
        "##############################\n",
        "# Run Elo package\n",
        "##############################\n",
        "\n",
        "# initalize Elo in 1985 and at value = 1500\n",
        "# set kfac = 20; FiveThrityEight found 20 as optimal NBA parameter\n",
        "eloRatings <- elo(select(regSeason.df, -Season),  # using date index, so excluding Season\n",
        "                  init = 1500, gamma = gammaVect, kfac = 20, history = TRUE)\n",
        "\n",
        "# return Elo ratings history & clean up output \n",
        "eloHistory.df <- data.frame(eloRatings$history[ , , 1])  # return history\n",
        "eloHistory.df <- tibble::rownames_to_column(eloHistory.df, var = \"Team\")  # adj row.names\n",
        "eloHistory.df$Team <- as.numeric(eloHistory.df$Team)  # adjust Team class for join\n",
        "\n",
        "# join actual team names to dataframe\n",
        "eloHistory.df <- inner_join(eloHistory.df, teams.df, by = c(\"Team\" = \"Team_Id\"))   \n",
        "eloHistory.df <- eloHistory.df %>% \n",
        "                 select(Team, Team_Name, everything()) %>%  # fold in actual team names\n",
        "                 select(-Team)  # remove Team Id for melt"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "a6783c0c-f63e-97dd-8e23-1ba2bff9740f"
      },
      "source": [
        "**Plot Elo rating evolution from 1985 to 2016.** Filtered for plotting of 2016 Final 4 teams, which shows Villanova surging to the top rating spot through the regular season. Output seems reasonable and consistent with tourney outcomes."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ea51e67b-8e87-fa65-fbd6-d4770dbb15ba"
      },
      "outputs": [],
      "source": [
        "##############################\n",
        "# Plot the Elo ratings\n",
        "##############################\n",
        "\n",
        "# melt Elo history output for plotting\n",
        "eloHistoryMelted.df <- melt(eloHistory.df, id = \"Team_Name\", \n",
        "                            variable.name = \"Date\", \n",
        "                            value.name = \"Rating\")\n",
        "\n",
        "# plot melted Elo time series for selected teams\n",
        "ggplot(eloHistoryMelted.df %>% filter(Team_Name %in%  \n",
        "                                      c(\"Villanova\", \n",
        "                                        \"North Carolina\", \n",
        "                                        \"Syracuse\", \n",
        "                                        \"Oklahoma\")),\n",
        "       aes(x = Date, y = Rating, colour = Team_Name, group = Team_Name)) + \n",
        "  geom_line() + \n",
        "  scale_x_discrete(breaks = c(\"X1\", \"X586\", \"X1172\", \"X1758\", \"X2344\", \"X2930\", \"X3516\"),\n",
        "                   labels = c(\"1985\", \"1990\", \"1995\", \"2000\", \"2005\", \"2010\", \"2015\")) + \n",
        "  scale_y_continuous(breaks = seq(1000, 2200, 50)) + \n",
        "  ggtitle(\"Elo Rating Evolution from 1985 to 2016: The 2016 Final 4 Tourney Teams\") +\n",
        "  labs(x = \"Year\", y = \"Elo Rating\") + \n",
        "  theme(plot.title = element_text(size = 8, face = \"bold\"),\n",
        "        axis.text = element_text(size = 8), \n",
        "        axis.title = element_text(size = 8, face = \"bold\"),\n",
        "        legend.title = element_blank(),\n",
        "        legend.text = element_text(size = 8),\n",
        "        legend.position = \"bottom\", \n",
        "        legend.key = element_rect(fill = \"white\"),\n",
        "        legend.key.width = unit(0.75, \"cm\")) +\n",
        "   coord_fixed(ratio = 1.618)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "9a3c5140-073d-c3d8-8be1-00a0912fbaef"
      },
      "source": [
        "**Extract Elo ratings at end of each season and save data to csv.** Elo ratings data can be pulled into a core data set as another feature for predicting tourney outcomes. It consistently ranks as one of the top 2 most important explanatory variables (along with Seed Rank)."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "8de7e9a6-d2a6-facf-f919-6ea55ac04c0b"
      },
      "outputs": [],
      "source": [
        "##############################\n",
        "# Get Elo at end of each regaular season\n",
        "##############################\n",
        "\n",
        "# wrangle date time-series into workable format\n",
        "# create lookup to crosswalk Elo time index with actual dates\n",
        "dateSeqLookup.df <- data.frame(as.numeric(levels(as.factor(regSeason.df$Date))))  \n",
        "dateSeqLookup.df <- plyr::rename(dateSeqLookup.df, \n",
        "                            c(\"as.numeric.levels.as.factor.regSeason.df.Date...\" = \"Date\"))\n",
        "dateSeqLookup.df <- mutate(dateSeqLookup.df, Seq = seq(1:nrow(dateSeqLookup.df)))  \n",
        "\n",
        "# create table with max date values to pull end of season Elo\n",
        "endSeasonLookup.df <- regSeason.df %>% \n",
        "                      group_by(Season) %>%\n",
        "                      summarize(SeasonEnd = max(Date)) %>%\n",
        "                      select(SeasonEnd, Season)\n",
        "\n",
        "# join date with index sequence for dplyr column downselect on Season\n",
        "endSeasonLookup.df <- inner_join(endSeasonLookup.df, \n",
        "                                 dateSeqLookup.df,  \n",
        "                                 by = c(\"SeasonEnd\" = \"Date\")) \n",
        "\n",
        "# down select elo rating columns representing end of each Season\n",
        "dateSeqVect <- as.vector(endSeasonLookup.df$Seq) \n",
        "eloSeasonRatings.df <- eloHistory.df %>% \n",
        "                       select(Team_Name, dateSeqVect)\n",
        "\n",
        "# adjust colnames to reflect Season year\n",
        "dateSeasonVect <- as.vector(endSeasonLookup.df$Season)\n",
        "colnames(eloSeasonRatings.df)[2:ncol(eloSeasonRatings.df)] <- dateSeasonVect\n",
        "\n",
        "# melt the Elo history output for lookup\n",
        "eloSeasonMelted.df <- melt(eloSeasonRatings.df, id = \"Team_Name\", \n",
        "                           variable.name = \"Season\", \n",
        "                           value.name = \"Rating\")\n",
        "\n",
        "# join the Elo ratings for each season with Team One and Team Two\n",
        "eloSeasonMelted.df <- left_join(eloSeasonMelted.df, \n",
        "                                teams.df, \n",
        "                                by = c(\"Team_Name\" = \"Team_Name\"))  # assign team id \n",
        "eloSeasonMelted.df$TeamId <- paste0(eloSeasonMelted.df$Season, \"_\", \n",
        "                                    eloSeasonMelted.df$Team_Id)  # create Season/TeamOne id\n",
        "eloSeasonMelted.df <- plyr::rename(eloSeasonMelted.df, c(\"Team_Id\" = \"Team\"))\n",
        "\n",
        "##############################\n",
        "# Save Elo per Season data to csv\n",
        "##############################\n",
        "\n",
        "# save csv file for external pull\n",
        "write.csv(eloSeasonMelted.df, file = \"eloSeasonMelted.csv\", row.names=FALSE)"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "R",
      "language": "R",
      "name": "ir"
    },
    "language_info": {
      "codemirror_mode": "r",
      "file_extension": ".r",
      "mimetype": "text/x-r-source",
      "name": "R",
      "pygments_lexer": "r",
      "version": "3.3.3"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}