{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Kyle Guy finds his Cinderella "},{"metadata":{},"cell_type":"markdown","source":"March is madness because of the upsets. If the tournament was all chalk every year, there would be no reason to watch, or even hold the tournament in the first place. Unfortunately, this year showed us just how convoluted this logic is. We didn't get the madness this year, and we missed every second of it. Pundits can crown pseudo-champions, claiming this team or that team would have danced all the way to a title, yet it's not the same. \n\nWe may have been some of the only fans in the country rooting for 1-seeded UVA to pull out their 2018 first round matchup against the 16-seeded Retrievers of UMBC - we are UVA students, after all. Nevertheless, we certainly could sense the intense excitement (or, in our case, dread) associated with that game. An upset of that proportion was unprecedented, and completely unpredicted. Therein lies our question - how could this historic feat happen? What factors existed going into that game that suggested maybe, just maybe, UMBC actually had a chance? \n\nThese are the questions we set out to answer. Welcome to our NCAA Tournament EDA.\n \n\n"},{"metadata":{},"cell_type":"markdown","source":"# Problem\n\n"},{"metadata":{},"cell_type":"markdown","source":"Before diving into our own analysis, we wanted to explore the current ranking system landscape to see if this question had already been answered. In evaluating some of the top ranking systems (defined by popularity and data comprehensiveness), we determined that none of them provide a particular advantage or ‘edge’ in predicting tournament game outcomes compared to others or the designated regional seeds. We evaluated all 10 rankings systems that have been active from 2003 through 2019 and found that they all had a 70-72% success rate in predicting tournament game outcomes (Pomeroy and Sagarin success rates as well as the average of the ten systems are shown below). There was a clear gap in current ranking systems in their ability to predict upsets.\n\nHow might we be able to identify certain matchup characteristics that increase this percentage? What stats could make us think twice before automatically advancing the favorite to the next round of our bracket? What aspects of a game could have made those UVA students, so confident that their team would coast to an easy victory before tipoff, a little more prepared for the nightmare that followed?\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"setwd(\"../\")\ntourney <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneyCompactResults.csv\")\nrankings <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MMasseyOrdinals.csv\")\n\nnames <- unique(rankings$SystemName)\n\nni <- data.frame(\"System\" = character(1), \"SeasonCount\" = numeric(1),stringsAsFactors=FALSE)\ncolnames(ni) <- c(\"System\",\"SeasonCount\")\nfor (name in names) {\n  nr <- rankings[rankings$SystemName == name,]\n  yc <- length(unique(nr$Season))\n  ng <- c(name,yc)\n  ni <- rbind(ni,ng)\n}\nni <- ni[-1,]\n\nfull <- ni[ni$SeasonCount == 18,]\nnames <- full$System\nnumnames <- length(names)\nrankings <- rankings[is.element(rankings$SystemName,names),]\n\ntourneyranks <- data.frame(matrix(ncol=numnames+3,nrow=0))\ncolnames(tourneyranks) <- c(\"Year\",\"Team\",names,\"Average\")\nfor (year in 2003:2019) {\n  ygames <- tourney[tourney$Season==year,]\n  teams <- sort(unique(c(unique(ygames$WTeamID),unique(ygames$LTeamID))))\n  for (team in teams) {\n    ytranks <- rankings[rankings$Season==year & rankings$TeamID==team,]\n    set <- data.frame(matrix(0,1,numnames+3))\n    set[1,1] <- year\n    set[1,2] <- team\n    colnames(set) <- c(\"Year\",\"Team\",names,\"Average\")\n    n = 0\n    counter = 0\n    for (i in 1:numnames) {\n      name <- names[i]\n      nextr <- ytranks[ytranks$SystemName==name,]\n      nextr <- tail(nextr,1)\n      if (length(nextr$OrdinalRank > 0)) {\n        set[1,i+2] <- nextr$OrdinalRank\n        n = n + 1\n        counter = counter + nextr$OrdinalRank\n      } else {\n        set[1,i+2] <- NA\n      }\n    }\n    set[1,3+numnames] <- counter/n\n    tourneyranks <- rbind(tourneyranks,set)\n  }\n}\n\nwrite.csv(tourneyranks,\"TourneyRanksFull.csv\")\n\ntourneysub <- tourney[tourney$Season >= 2003,]\nfor (i in 1:nrow(tourneysub)) {\n  game <- tourneysub[i,]\n  wid <- game$WTeamID\n  lid <- game$LTeamID\n  year <- game$Season\n  for (col in 1:(ncol(tourneyranks)-2)) {\n    wseed <- tourneyranks[tourneyranks$Team == wid & tourneyranks$Year == year,col+2]\n    lseed <- tourneyranks[tourneyranks$Team == lid & tourneyranks$Year == year,col+2]\n    if (is.na(wseed) & is.na(lseed)){\n      res <- NA\n    }\n    else if (is.na(lseed)) {\n      res <- 1\n    }\n    else if (is.na(wseed)) {\n      res <- 0\n    }\n    else if (wseed < lseed) {\n      res <- 1\n    }\n    else {\n      res <- 0\n    }\n    tourneysub[i,8+col] <- res\n  }\n}\n\naccuracy <- data.frame(matrix(0,ncol=2,nrow=length(names(tourneyranks)[-c(1,2)])))\ncolnames(accuracy) <- c(\"Ranking System\", \"Percent Accuracy\")\nrn <- names(tourneyranks)[-c(1,2)]\nfor (rs in 1:length(rn)){\n  accuracy[rs,1] <- rn[rs]\n  col <- tourneysub[,8+rs]\n  col <- col[!is.na(col)]\n  accuracy[rs,2] <- mean(col)\n}\naccuracy","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Libraries"},{"metadata":{"trusted":true},"cell_type":"code","source":"knitr::opts_chunk$set(echo = TRUE)\n\nlibrary(dplyr)\nlibrary(ggmap)\nlibrary(tidyverse)\nlibrary(plyr)\nlibrary(gdata)\nlibrary(ggplot2)\nlibrary(mlbench)\nlibrary(MASS)\nlibrary(pROC)\nlibrary(BAS)\nlibrary(readr)\nlibrary(geosphere)\nlibrary(caret)\nlibrary(randomForest)\nlibrary(pscl)\n\nregister_google(key = \"AIzaSyBKQ2BQOIVQW5XTQC4U0aNFCnHafmFYw3g\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data read in"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n\n#Starting Data Set\n\ntournament_data <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneyDetailedResults.csv\")\n\n#City Data\ncity_names <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/Cities.csv\")\n\nMRegularGamesCompactResult <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MRegularSeasonCompactResults.csv\")\n\nMGameCities <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MGameCities.csv\")\n\n# Regular Season Games\nMRegularGames <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MRegularSeasonDetailedResults.csv\")\n\n# Seed Data\nMTournamentSeeds <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneySeeds.csv\")\n\n# Coaching Data\nCoachingYears <- read.csv(\"../kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MTeamCoaches.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"To be perfectly clear, we aren't identifying any in-game strategies - we don't pretend to be expert basketball tacticians, so we'll leave those to the experts. What we are trying to accomplish is looking at the data that exists before the tournament game even starts to see if the underdog has a little more of a fighting chance than they're being given credit for.\n\nThis first step in the EDA is thus to compile all of the data that might be relevant to an underdog's chance of winning, and a large part of this data is comprised of the season's statistical averages for the team. Things like better rebounding, fewer turnovers, or more steals could all theoretically make an underdog a little more likely to win, right? We thought so, so we included stats like those.\n\nWe also made sure to test and correct the quality of our data by ensuring that there was enough data points for a statistically signficant conclusion as well as checking important underlying assumptions such as multicollinearity in our regression models.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"winning_team_info = MRegularGames[,c(\"Season\",\"WTeamID\",\"WScore\", \"WFGM\", \"WFGA\", \"WFGM3\", \"WFGA3\", \"WFTM\", \"WFTA\", \"WOR\", \"WDR\", \"WAst\", \"WTO\", \"WStl\", \"WBlk\", \"WPF\")]\nnames(winning_team_info) = c(\"Season\",\"TeamID\",\"Score\", \"FGM\", \"FGA\", \"FGM3\", \"FGA3\", \"FTM\", \"FTA\", \"OR\", \"DR\", \"Ast\", \"TO\", \"Stl\", \"Blk\", \"PF\")\n\nlosing_team_info = MRegularGames[,c(\"Season\",\"LTeamID\",\"LScore\", \"LFGM\", \"LFGA\", \"LFGM3\", \"LFGA3\", \"LFTM\", \"LFTA\", \"LOR\", \"LDR\", \"LAst\", \"LTO\", \"LStl\", \"LBlk\", \"LPF\")]\nnames(losing_team_info) = c(\"Season\",\"TeamID\",\"Score\", \"FGM\", \"FGA\", \"FGM3\", \"FGA3\", \"FTM\", \"FTA\", \"OR\", \"DR\", \"Ast\", \"TO\", \"Stl\", \"Blk\", \"PF\")\nMRegularGames = rbind(winning_team_info,losing_team_info)\n\nteam_season_level_avg_summary = MRegularGames %>% group_by(Season, TeamID) %>% summarise_at(vars(\"Score\", \"FGM\", \"FGA\", \"FGM3\", \"FGA3\", \"FTM\", \"FTA\", \"OR\", \"DR\", \"Ast\", \"TO\", \"Stl\", \"Blk\", \"PF\"), mean)\n\n\n#Winning Team Join\ntournament_data <- left_join(tournament_data, team_season_level_avg_summary, by=c(\"WTeamID\"=\"TeamID\", \"Season\"))\n\nnames(tournament_data)[names(tournament_data)==\"Score\"] <- \"Season_AVG_WScore\"\nnames(tournament_data)[names(tournament_data)==\"FGM\"] <- \"Season_AVG_WFGM\"\nnames(tournament_data)[names(tournament_data)==\"FGA\"] <- \"Season_AVG_WFGA\"\nnames(tournament_data)[names(tournament_data)==\"FGM3\"] <- \"Season_AVG_WFGM3\"\nnames(tournament_data)[names(tournament_data)==\"FGA3\"] <- \"Season_AVG_WFGA3\"\nnames(tournament_data)[names(tournament_data)==\"FTM\"] <- \"Season_AVG_WFTM\"\nnames(tournament_data)[names(tournament_data)==\"FTA\"] <- \"Season_AVG_WFTA\"\nnames(tournament_data)[names(tournament_data)==\"OR\"] <- \"Season_AVG_WOR\"\nnames(tournament_data)[names(tournament_data)==\"DR\"] <- \"Season_AVG_WDR\"\nnames(tournament_data)[names(tournament_data)==\"Ast\"] <- \"Season_AVG_WAst\"\nnames(tournament_data)[names(tournament_data)==\"TO\"] <- \"Season_AVG_WTO\"\nnames(tournament_data)[names(tournament_data)==\"Stl\"] <- \"Season_AVG_WStl\"\nnames(tournament_data)[names(tournament_data)==\"Blk\"] <- \"Season_AVG_WBlk\"\nnames(tournament_data)[names(tournament_data)==\"PF\"] <- \"Season_AVG_WPF\"\n\n#Losing Team Join\ntournament_data <- left_join(tournament_data, team_season_level_avg_summary, by=c(\"LTeamID\"=\"TeamID\", \"Season\"))\nnames(tournament_data)[names(tournament_data)==\"Score\"] <- \"Season_AVG_LScore\"\nnames(tournament_data)[names(tournament_data)==\"FGM\"] <- \"Season_AVG_LFGM\"\nnames(tournament_data)[names(tournament_data)==\"FGA\"] <- \"Season_AVG_LFGA\"\nnames(tournament_data)[names(tournament_data)==\"FGM3\"] <- \"Season_AVG_LFGM3\"\nnames(tournament_data)[names(tournament_data)==\"FGA3\"] <- \"Season_AVG_LFGA3\"\nnames(tournament_data)[names(tournament_data)==\"FTM\"] <- \"Season_AVG_LFTM\"\nnames(tournament_data)[names(tournament_data)==\"FTA\"] <- \"Season_AVG_LFTA\"\nnames(tournament_data)[names(tournament_data)==\"OR\"] <- \"Season_AVG_LOR\"\nnames(tournament_data)[names(tournament_data)==\"DR\"] <- \"Season_AVG_LDR\"\nnames(tournament_data)[names(tournament_data)==\"Ast\"] <- \"Season_AVG_LAst\"\nnames(tournament_data)[names(tournament_data)==\"TO\"] <- \"Season_AVG_LTO\"\nnames(tournament_data)[names(tournament_data)==\"Stl\"] <- \"Season_AVG_LStl\"\nnames(tournament_data)[names(tournament_data)==\"Blk\"] <- \"Season_AVG_LBlk\"\nnames(tournament_data)[names(tournament_data)==\"PF\"] <- \"Season_AVG_LPF\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Next, we merged in seed data so that we could organize our stats based on higher and lower seeds instead of winning and losing team."},{"metadata":{"trusted":true},"cell_type":"code","source":"\nMTournamentSeeds <- mutate(MTournamentSeeds, Seed = Seed)\n\nMTournamentSeeds$Seed <- as.numeric(gsub(\"\\\\D\",\"\", MTournamentSeeds$Seed))\n\ntournament_data <- left_join(tournament_data, MTournamentSeeds, by=c(\"Season\", \"WTeamID\"=\"TeamID\"))\nnames(tournament_data)[names(tournament_data)==\"Seed\"] <- \"WSeed\"\n\ntournament_data <- left_join(tournament_data, MTournamentSeeds, by=c(\"Season\", \"LTeamID\"=\"TeamID\"))\nnames(tournament_data)[names(tournament_data)==\"Seed\"] <- \"LSeed\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Another important factor we thought we should consider is location. NCAA tournament games are considered neutral-site contests, since the games aren't played at either team's home arena. However, it's plausible that the distance the team has to travel to get to the game site influences their performance, for two reasons. First, we thought it possible that longer distances might present physical challenges to the players, whose routines are interrupted. They might have to play at an unnatural time or spend time traveling that could have been put to better use. Second, contests played far away from campus could make it much less likely a member of that team's fan base makes the trip to watch - making it a more difficult environment for the team to play in. As such, we decided to include location data in the model, as well."},{"metadata":{"trusted":true},"cell_type":"code","source":"\n\n\n#Create TeamID -> CityName Table\ncity_games <- left_join(MGameCities, city_names)\ncity_games_winners <- left_join(city_games, MRegularGamesCompactResult)\ncity_games_winners <- city_games_winners[c(-5,-1,-2,-12,-10,-9)] %>% filter(WLoc!=\"N\")\ncity_games_winners$WCityID <- ifelse(city_games_winners$WLoc==\"H\", city_games_winners$CityID, 0)\n\n\ncity_table <- data.frame(\"TeamID\"=(city_games_winners$WTeamID), \"CityID\"=city_games_winners$WCityID) %>% filter(CityID!=0)\n\n#Only keep one combination of TeamID and CityID\ncity_table <- unique(city_table)\ncity_table <- left_join(city_table, city_names)\n\n#Load Winning Locations into Tournament Data\ntournament_data <- left_join(tournament_data, city_table, by=c(\"WTeamID\"=\"TeamID\"))\nnames(tournament_data)[names(tournament_data)==\"CityID\"] = \"WCityID\"\nnames(tournament_data)[names(tournament_data)==\"City\"] = \"WCity\"\nnames(tournament_data)[names(tournament_data)==\"State\"] = \"WState\"\n\ntournament_data <- left_join(tournament_data, city_table, by=c(\"LTeamID\"=\"TeamID\"))\nnames(tournament_data)[names(tournament_data)==\"CityID\"] = \"LCityID\"\nnames(tournament_data)[names(tournament_data)==\"City\"] = \"LCity\"\nnames(tournament_data)[names(tournament_data)==\"State\"] = \"LState\"\n\n#Load game locations into tournament data\ntournament_data <-left_join(tournament_data, city_games[-5])\n\n\n\nfor (i in 1:length(tournament_data$WTeamID))\n{\n  miles <- NA\n  if (is.na(tournament_data$City[i])==FALSE)\n  {\n    if (tournament_data$WState[i]==\"HI\" || tournament_data$State[i]==\"HI\")\n    {\n      miles <- NA\n    }\n    else\n      {\n    dist <- mapdist(paste(tournament_data$WCity[i], tournament_data$WState[i], sep=\", \"), paste(tournament_data$City[i], tournament_data$State[i], sep=\", \"), mode=\"driving\")\n    miles <- dist[1,5]\n    }\n  }\n  tournament_data$WMiles[i] <- as.numeric(miles)\n}\n\n\n#Losing Team Distance\nfor (i in 1:length(tournament_data$LTeamID))\n{\n  miles <- NA\n  if (is.na(tournament_data$City[i])==FALSE)\n  {\n    if (tournament_data$LState[i]==\"HI\" || tournament_data$State[i]==\"HI\")\n    {\n      miles <- NA\n    }\n    else\n    {\n      dist <- mapdist(paste(tournament_data$LCity[i], tournament_data$LState[i], sep=\", \"), paste(tournament_data$City[i], tournament_data$State[i], sep=\", \"), mode=\"driving\")\n      miles <- dist[1,5]\n    }\n  }\n  tournament_data$LMiles[i] <- as.numeric(miles)\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Finally, as fans of the best coach in college basketball, UVA's own Tony Bennett, we knew coaching data also had to play a part in our model. It's feasible that teams with more experienced coaches have a better chance of pulling out the close games that upsets tend to be. It's also possible that mid-season coaching changes could affect tournament performance. Accordingly, we included both of these measures."},{"metadata":{"trusted":true},"cell_type":"code","source":"\nmid_season_changes = CoachingYears[CoachingYears$LastDayNum != 154,]\n\nfor (i in (nrow(CoachingYears[CoachingYears$Season == 1985,])+1):nrow(CoachingYears)){\n  \n  row <- CoachingYears[i,]\n  season = as.integer(row$Season)\n  season_subsetted_data = CoachingYears[CoachingYears$Season < season,]\n  coach_subsetted_data = season_subsetted_data[season_subsetted_data$CoachName == row$CoachName,]\n  CoachingYears[i,6] = nrow(coach_subsetted_data)\n}\n\n#Merge Coaching Years\ntournament_data <- left_join(tournament_data, CoachingYears[c(\"Season\", \"TeamID\", \"V6\")], by=c(\"Season\", \"WTeamID\"=\"TeamID\"))\n\nnames(tournament_data)[names(tournament_data)==\"V6\"] <- \"WCoachingYears\"\n\ntournament_data <- left_join(tournament_data, CoachingYears[c(\"Season\", \"TeamID\", \"V6\")], by=c(\"Season\", \"LTeamID\"=\"TeamID\"))\n\nnames(tournament_data)[names(tournament_data)==\"V6\"] <- \"LCoachingYears\"\n\n# Merging Coaching Changes\nmid_season_changes$change <- 1\n\ntournament_data <- left_join(tournament_data, mid_season_changes[c(\"Season\",\"TeamID\",\"change\")], by=c(\"Season\", \"WTeamID\"=\"TeamID\"))\nnames(tournament_data)[names(tournament_data)==\"change\"] <- \"WChange\"\ntournament_data$WChange <- ifelse(is.na(tournament_data$WChange), 0, 1)\n\n\ntournament_data <- left_join(tournament_data, mid_season_changes[c(\"Season\",\"TeamID\",\"change\")], by=c(\"Season\", \"LTeamID\"=\"TeamID\"))\nnames(tournament_data)[names(tournament_data)==\"change\"] <- \"LChange\"\ntournament_data$LChange <- ifelse(is.na(tournament_data$LChange), 0, 1)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"After we compiled our data set, we had to decide which games should be included in our analysis. To properly answer our question, we needed to determine what factors might make an underdog more likely to pull off the upset. In some games, an underdog doesn't really exist, as the teams are evenly matched and a win from either team wouldn't be a surprise.\n\nWhat, then, counts as an upset? A 5-12 matchup is rife with potential, whereas an 8-9 matchup pits pretty similar teams against each other. We decided to split the difference, and count any game with opponents more than three seed lines apart as a potential upset. For example, a first round game between a 7-seed and 10-seed wouldn't qualify as a potential upset, and neither would a 1-4 matchup in the Sweet 16. A 6-11 game does count though, as does a 2-7 contest in the second round. There's some room for debate here (is a 1-4 matchup really that different from the 1-seed playing a 5-seed?), but we had to draw the line somewhere.\n\nAdditionally, at this stage in our process, we realized it might be interesting to look at potential cinderellas as well. Is there any difference between a team who might just get lucky with their first round matchup to pull an upset, and a team with real potential to advance past the first weekend? What factors might indicate the difference between two such teams?\n\nAs such, we decided to analyze games that had cinderella potential as well - here defined as any team seed eighth or below that makes it past the second weekend into the Sweet 16.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"\ntournament_data$Upset_potential <- ifelse(abs(tournament_data$WSeed-tournament_data$LSeed)>3, 1, 0)\n\ntournament_data$Cinderella_potential <- ifelse(((tournament_data$WSeed>7 | tournament_data$LSeed>7) & tournament_data$DayNum>137), 1, 0)\n\ntournament_data$Upset <- ifelse((tournament_data$WSeed-tournament_data$LSeed)>3, 1, 0)\n\ntournament_data$Cinderella <- ifelse(tournament_data$WSeed>7 & tournament_data$DayNum>137, 1, 0)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Renaming variables to higher and lower seeds rather than winners and losers"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Do this AFTER everything else\n\nupsets <- filter(tournament_data, tournament_data$WSeed>tournament_data$LSeed)\nnonupsets <- filter(tournament_data, tournament_data$WSeed<tournament_data$LSeed)\n\n# Renaming Losers to Higher Seed\n\nnames(upsets)[names(upsets)==\"LTeamID\"] <- \"HTeamID\"\nnames(upsets)[names(upsets)==\"LScore\"] <- \"HScore\"\nnames(upsets)[names(upsets)==\"LFGM\"] <- \"HFGM\"\nnames(upsets)[names(upsets)==\"LFGA\"] <- \"HFGA\"\nnames(upsets)[names(upsets)==\"LFGM3\"] <- \"HFGM3\"\nnames(upsets)[names(upsets)==\"LFGA3\"] <- \"HFGA3\"\nnames(upsets)[names(upsets)==\"LFTM\"] <- \"HFTM\"\nnames(upsets)[names(upsets)==\"LFTA\"] <- \"HFTA\"\nnames(upsets)[names(upsets)==\"LOR\"] <- \"HOR\"\nnames(upsets)[names(upsets)==\"LDR\"] <- \"HDR\"\nnames(upsets)[names(upsets)==\"LAst\"] <- \"HAst\"\nnames(upsets)[names(upsets)==\"LTO\"] <- \"HTO\"\nnames(upsets)[names(upsets)==\"LStl\"] <- \"HStl\"\nnames(upsets)[names(upsets)==\"LBlk\"] <- \"HBlk\"\nnames(upsets)[names(upsets)==\"LPF\"] <- \"HPF\"\nnames(upsets)[names(upsets)==\"LSeed\"] <- \"HSeed\"\nnames(upsets)[names(upsets)==\"Season_AVG_LScore\"] <- \"Season_AVG_HScore\"\nnames(upsets)[names(upsets)==\"Season_AVG_LFGM\"] <- \"Season_AVG_HFGM\"\nnames(upsets)[names(upsets)==\"Season_AVG_LFGA\"] <- \"Season_AVG_HFGA\"\nnames(upsets)[names(upsets)==\"Season_AVG_LFGM3\"] <- \"Season_AVG_HFGM3\"\nnames(upsets)[names(upsets)==\"Season_AVG_LFGA3\"] <- \"Season_AVG_HFGA3\"\nnames(upsets)[names(upsets)==\"Season_AVG_LFTM\"] <- \"Season_AVG_HFTM\"\nnames(upsets)[names(upsets)==\"Season_AVG_LFTA\"] <- \"Season_AVG_HFTA\"\nnames(upsets)[names(upsets)==\"Season_AVG_LOR\"] <- \"Season_AVG_HOR\"\nnames(upsets)[names(upsets)==\"Season_AVG_LDR\"] <- \"Season_AVG_HDR\"\nnames(upsets)[names(upsets)==\"Season_AVG_LAst\"] <- \"Season_AVG_HAst\"\nnames(upsets)[names(upsets)==\"Season_AVG_LTO\"] <- \"Season_AVG_HTO\"\nnames(upsets)[names(upsets)==\"Season_AVG_LStl\"] <- \"Season_AVG_HStl\"\nnames(upsets)[names(upsets)==\"Season_AVG_LBlk\"] <- \"Season_AVG_HBlk\"\nnames(upsets)[names(upsets)==\"Season_AVG_LPF\"] <- \"Season_AVG_HPF\"\nnames(upsets)[names(upsets)==\"LCityID\"] <- \"HCityID\"\nnames(upsets)[names(upsets)==\"LCity\"] <- \"HCity\"\nnames(upsets)[names(upsets)==\"LState\"] <- \"HState\"\nnames(upsets)[names(upsets)==\"LMiles\"] <- \"HMiles\"\nnames(upsets)[names(upsets)==\"LChange\"] <- \"HChange\"\nnames(upsets)[names(upsets)==\"LCoachingYears\"] <- \"HCoachingYears\"\n\n# Renaming Winners to Lower Seed\n\nnames(upsets)[names(upsets)==\"WTeamID\"] <- \"LTeamID\"\nnames(upsets)[names(upsets)==\"WScore\"] <- \"LScore\"\nnames(upsets)[names(upsets)==\"WFGM\"] <- \"LFGM\"\nnames(upsets)[names(upsets)==\"WFGA\"] <- \"LFGA\"\nnames(upsets)[names(upsets)==\"WFGM3\"] <- \"LFGM3\"\nnames(upsets)[names(upsets)==\"WFGA3\"] <- \"LFGA3\"\nnames(upsets)[names(upsets)==\"WFTM\"] <- \"LFTM\"\nnames(upsets)[names(upsets)==\"WFTA\"] <- \"LFTA\"\nnames(upsets)[names(upsets)==\"WOR\"] <- \"LOR\"\nnames(upsets)[names(upsets)==\"WDR\"] <- \"LDR\"\nnames(upsets)[names(upsets)==\"WAst\"] <- \"LAst\"\nnames(upsets)[names(upsets)==\"WTO\"] <- \"LTO\"\nnames(upsets)[names(upsets)==\"WStl\"] <- \"LStl\"\nnames(upsets)[names(upsets)==\"WBlk\"] <- \"LBlk\"\nnames(upsets)[names(upsets)==\"WPF\"] <- \"LPF\"\nnames(upsets)[names(upsets)==\"WSeed\"] <- \"LSeed\"\nnames(upsets)[names(upsets)==\"Season_AVG_WScore\"] <- \"Season_AVG_LScore\"\nnames(upsets)[names(upsets)==\"Season_AVG_WFGM\"] <- \"Season_AVG_LFGM\"\nnames(upsets)[names(upsets)==\"Season_AVG_WFGM\"] <- \"Season_AVG_LFGM\"\nnames(upsets)[names(upsets)==\"Season_AVG_WFGA\"] <- \"Season_AVG_LFGA\"\nnames(upsets)[names(upsets)==\"Season_AVG_WFGM3\"] <- \"Season_AVG_LFGM3\"\nnames(upsets)[names(upsets)==\"Season_AVG_WFGA3\"] <- \"Season_AVG_LFGA3\"\nnames(upsets)[names(upsets)==\"Season_AVG_WFTM\"] <- \"Season_AVG_LFTM\"\nnames(upsets)[names(upsets)==\"Season_AVG_WFTA\"] <- \"Season_AVG_LFTA\"\nnames(upsets)[names(upsets)==\"Season_AVG_WOR\"] <- \"Season_AVG_LOR\"\nnames(upsets)[names(upsets)==\"Season_AVG_WDR\"] <- \"Season_AVG_LDR\"\nnames(upsets)[names(upsets)==\"Season_AVG_WAst\"] <- \"Season_AVG_LAst\"\nnames(upsets)[names(upsets)==\"Season_AVG_WTO\"] <- \"Season_AVG_LTO\"\nnames(upsets)[names(upsets)==\"Season_AVG_WStl\"] <- \"Season_AVG_LStl\"\nnames(upsets)[names(upsets)==\"Season_AVG_WBlk\"] <- \"Season_AVG_LBlk\"\nnames(upsets)[names(upsets)==\"Season_AVG_WPF\"] <- \"Season_AVG_LPF\"\nnames(upsets)[names(upsets)==\"WSeed\"] <- \"LSeed\"\nnames(upsets)[names(upsets)==\"WCityID\"] <- \"LCityID\"\nnames(upsets)[names(upsets)==\"WCity\"] <- \"LCity\"\nnames(upsets)[names(upsets)==\"WState\"] <- \"LState\"\nnames(upsets)[names(upsets)==\"WMiles\"] <- \"LMiles\"\nnames(upsets)[names(upsets)==\"WChange\"] <- \"LChange\"\nnames(upsets)[names(upsets)==\"WCoachingYears\"] <- \"LCoachingYears\"\n\n\n# Renaming Winners to Higher Seeds\n\nnames(nonupsets)[names(nonupsets)==\"WTeamID\"] <- \"HTeamID\"\nnames(nonupsets)[names(nonupsets)==\"WScore\"] <- \"HScore\"\nnames(nonupsets)[names(nonupsets)==\"WFGM\"] <- \"HFGM\"\nnames(nonupsets)[names(nonupsets)==\"WFGA\"] <- \"HFGA\"\nnames(nonupsets)[names(nonupsets)==\"WFGM3\"] <- \"HFGM3\"\nnames(nonupsets)[names(nonupsets)==\"WFGA3\"] <- \"HFGA3\"\nnames(nonupsets)[names(nonupsets)==\"WFTM\"] <- \"HFTM\"\nnames(nonupsets)[names(nonupsets)==\"WFTA\"] <- \"HFTA\"\nnames(nonupsets)[names(nonupsets)==\"WOR\"] <- \"HOR\"\nnames(nonupsets)[names(nonupsets)==\"WDR\"] <- \"HDR\"\nnames(nonupsets)[names(nonupsets)==\"WAst\"] <- \"HAst\"\nnames(nonupsets)[names(nonupsets)==\"WTO\"] <- \"HTO\"\nnames(nonupsets)[names(nonupsets)==\"WStl\"] <- \"HStl\"\nnames(nonupsets)[names(nonupsets)==\"WBlk\"] <- \"HBlk\"\nnames(nonupsets)[names(nonupsets)==\"WPF\"] <- \"HPF\"\nnames(nonupsets)[names(nonupsets)==\"WSeed\"] <- \"HSeed\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WScore\"] <- \"Season_AVG_HScore\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WFGM\"] <- \"Season_AVG_HFGM\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WFGA\"] <- \"Season_AVG_HFGA\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WFGM3\"] <- \"Season_AVG_HFGM3\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WFGA3\"] <- \"Season_AVG_HFGA3\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WFTM\"] <- \"Season_AVG_HFTM\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WFTA\"] <- \"Season_AVG_HFTA\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WOR\"] <- \"Season_AVG_HOR\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WDR\"] <- \"Season_AVG_HDR\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WAst\"] <- \"Season_AVG_HAst\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WTO\"] <- \"Season_AVG_HTO\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WStl\"] <- \"Season_AVG_HStl\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WBlk\"] <- \"Season_AVG_HBlk\"\nnames(nonupsets)[names(nonupsets)==\"Season_AVG_WPF\"] <- \"Season_AVG_HPF\"\nnames(nonupsets)[names(nonupsets)==\"WSeed\"] <- \"HSeed\"\nnames(nonupsets)[names(nonupsets)==\"WCityID\"] <- \"HCityID\"\nnames(nonupsets)[names(nonupsets)==\"WCity\"] <- \"HCity\"\nnames(nonupsets)[names(nonupsets)==\"WState\"] <- \"HState\"\nnames(nonupsets)[names(nonupsets)==\"WMiles\"] <- \"HMiles\"\nnames(nonupsets)[names(nonupsets)==\"WChange\"] <- \"HChange\"\nnames(nonupsets)[names(nonupsets)==\"WCoachingYears\"] <- \"HCoachingYears\"\n\nfinal <- rbind(upsets, nonupsets)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Separating data into upset and cinderella sets for analysis"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n\nregression_data <- final \n\nnames(regression_data)[names(regression_data)==\"Upset\"] <- \"Upset_result\"\nnames(regression_data)[names(regression_data)==\"Cinderella\"] <- \"Cinderella.\"\nnames(regression_data)[names(regression_data)==\"HCoachingYears\"] <- \"HCoach_Exp\"\nnames(regression_data)[names(regression_data)==\"LCoachingYears\"] <- \"LCoach_Exp\"\nnames(regression_data)[names(regression_data)==\"HChange\"] <- \"HCoach_Change\"\nnames(regression_data)[names(regression_data)==\"LChange\"] <- \"LCoach_Change\"\n\n\nupset_regression_data = regression_data[regression_data$Upset_potential==1,]\nupset_regression_data_cleaned = na.omit(upset_regression_data)\n\ncinderella_regression_data = regression_data[regression_data$Cinderella_potential==1,]\ncinderella_regression_data_cleaned = na.omit(cinderella_regression_data)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Logistic models"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n\nupset_logit_model <- glm(Upset_result ~ HMiles + LMiles + HCoach_Exp + LCoach_Exp + HCoach_Change + LCoach_Change + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = upset_regression_data_cleaned, family = \"binomial\")\nupset_logit_model_stepwise <- stepAIC(upset_logit_model)\nsummary(upset_logit_model_stepwise)\npR2(upset_logit_model_stepwise)\n\ncinderella_logit_model <- glm(Cinderella. ~ HMiles + LMiles + HCoach_Exp + LCoach_Exp + HCoach_Change + LCoach_Change + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = cinderella_regression_data_cleaned, family = \"binomial\")\ncinderella_logit_model_stepwise <- stepAIC(cinderella_logit_model)\nsummary(cinderella_logit_model_stepwise)\npR2(cinderella_logit_model_stepwise)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Random Forest"},{"metadata":{"trusted":true},"cell_type":"code","source":"\nupset_forest_data = upset_regression_data\nupset_forest_data$Upset_result = as.factor(upset_forest_data$Upset_result)\n\nupset_forest_data_split <- sample(nrow(upset_forest_data), 0.7*nrow(upset_forest_data), replace = FALSE)\nupset_forest_training_data <- upset_forest_data[upset_forest_data_split,]\nupset_forest_validation_data <- upset_forest_data[-upset_forest_data_split,]\n\nupset_forest_model <- randomForest(Upset_result ~ HMiles + LMiles + HCoach_Exp + LCoach_Exp + HCoach_Change + LCoach_Change + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = upset_forest_training_data, importance = TRUE,na.action=na.exclude)\nupset_forest_model\n\nvarImpPlot(upset_forest_model)\n\n\ncinderella_forest_data = cinderella_regression_data\ncinderella_forest_data$Cinderella. = as.factor(cinderella_forest_data$Cinderella.)\n\ncinderella_forest_data_split <- sample(nrow(cinderella_forest_data), 0.7*nrow(cinderella_forest_data), replace = FALSE)\ncinderella_forest_training_data <- cinderella_forest_data[cinderella_forest_data_split,]\ncinderella_forest_validation_data <- cinderella_forest_data[-cinderella_forest_data_split,]\n\ncinderella_forest_model <- randomForest(Cinderella. ~ HMiles + LMiles + HCoach_Exp + LCoach_Exp + HCoach_Change + LCoach_Change + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = cinderella_forest_training_data, importance = TRUE,na.action=na.exclude)\ncinderella_forest_model\n\nvarImpPlot(cinderella_forest_model)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Bayesian multiple linear regression"},{"metadata":{"_uuid":"77434788-f2df-40d7-987a-42debfb371f9","_cell_guid":"df5e0228-21df-45e8-a9c8-b3895135e2b5","trusted":true},"cell_type":"code","source":"upset_bas <- bas.lm(Upset_result ~ HMiles + LMiles + HCoach_Exp + LCoach_Exp + HCoach_Change + LCoach_Change + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF,\n                     data = upset_regression_data,\n                     method = \"MCMC\",\n                     prior = \"ZS-null\",\n                     modelprior = uniform())\nsummary(upset_bas)\nimage(upset_bas, rotate=F,drop.always.included=TRUE, top.models=5)\n\n\ncinderella_bas <- bas.lm(Cinderella. ~ HMiles + LMiles + HCoach_Exp + LCoach_Exp + HCoach_Change + LCoach_Change + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF,\n                    data = cinderella_regression_data,\n                    method = \"MCMC\",\n                    prior = \"ZS-null\",\n                    modelprior = uniform())\nsummary(cinderella_bas)\nimage(cinderella_bas, rotate=F, drop.always.included=TRUE, top.models=5)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analysis of results\n\nAs we can see from the outputs above, we ran a couple of different models on both the upset and cinderella data - a random forest analysis as well as Bayesian multiple linear regression. We didn't want poor interaction between a model and any one variable to disrupt the validity of our test. If both analyses determine that a particular factor is important, then the evidence is robust enough for us. \n\nFour variables in the upset data were ranked as quite important to both models: the coaching experience of the higher-seeded coach, the distance the higher seed had to travel to get to the game, the season average field goals made in a game by the higher seed, and the lower seed's season average of blocks per game. While the presence of blocks as so important is a bit surprising, it could be chalked up to the low numbers of blocks that occur in an average game and the relatively small sample size of our data set. The other factors, however, make a lot of sense. Inexperienced coaches are more likely to make poor decisions during the flow of a game, giving underdogs a better chance of pulling off the upset. Higher-seeded teams that travel further to play the game might be impacted by the interruption to their normal schedule, or the absence of die-hard fans that they normally have cheering them on. Fewer shots made on average throughout the season could indicate just a poor shooting team, who would obviously be more prone to an upset. However, as UVA fans, we can say with confidence that higher-seeded teams with fewer shots made per game likely score low in that category because of their slow pace of play, not poor shooting ability. Yet this pace, too, might make a higher seed more susceptible to an upset. A slower pace means fewer possessions per game, which increases the variability and randomness of the outcome. \n\nAnother interesting point to note from these models is the relative importance of data from the higher-seeded team over that from the lower seed. Maybe pulling an upset isn't about an underdog that was seeded lower than they should have been. Maybe, instead, the weaknesses of the higher seed are what truly matters, and the particular underdog they're facing isn't quite as important. \n\nThe cinderella data yields fairly similar results to the upset models. A couple more variables that our analysis rates as important are the average number of free throws per game attempted by the higher seed and the distance the lower seed has to travel to get to the game. The distance particularly makes sense in this instance - fans of the lower-seeded team might decide to flock to the game once they see their team starting the tournament well. A contest held closer to home makes it more likely these fans will make the trip. The number of free throws attempted by the higher seed can be seen as a representation of a team's aggressiveness - if a team sits back and doesn't attack the basket (or take advantage of their superior strength/height/physicality), then they won't shoot as many free throws. Teams like these will be more susceptible to upsets, especially against cinderellas in later rounds that have proven to be worthy contenders. \n\nOnce again, the cinderella data demonstrates that most of the power in the model is derived from data about the higher seed. Even in this situation, where a lower seed has a track record of winning, the data suggests that weaknesses of their opponent affect their probability of pulling an upset more than their own strengths. The repeated successes may just be a factor of getting lucky with multiple opponents who aren't actually as good as they're expected to be. "},{"metadata":{},"cell_type":"markdown","source":"# Women's NCAA"},{"metadata":{},"cell_type":"markdown","source":"After looking at the men's data, we realized that we could apply the same process to analyzing upsets and cinderellas in the women's tournament as well. Could there be some factors that are important for men or women but not the other? Upsets tend to be more rare in the women's tournament, making the identification of characteristics important to upsets all the more valuable. "},{"metadata":{"trusted":true},"cell_type":"code","source":"Wtournament_data <- read.csv(\"../kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/WNCAATourneyDetailedResults.csv\")\n\n#City Data\nWcity_names <- read.csv(\"../kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/Cities.csv\")\n\nWRegularGamesCompactResult <- read.csv(\"../kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/WRegularSeasonCompactResults.csv\")\n\nWGameCities <- read.csv(\"../kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/WGameCities.csv\")\n\n# Regular Season Games\nWRegularGames <- read.csv(\"../kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/WRegularSeasonDetailedResults.csv\")\n\n# Seed Data\nWTournamentSeeds <- read.csv(\"../kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/WNCAATourneySeeds.csv\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"winning_team_info = WRegularGames[,c(\"Season\",\"WTeamID\",\"WScore\", \"WFGM\", \"WFGA\", \"WFGM3\", \"WFGA3\", \"WFTM\", \"WFTA\", \"WOR\", \"WDR\", \"WAst\", \"WTO\", \"WStl\", \"WBlk\", \"WPF\")]\nnames(winning_team_info) = c(\"Season\",\"TeamID\",\"Score\", \"FGM\", \"FGA\", \"FGM3\", \"FGA3\", \"FTM\", \"FTA\", \"OR\", \"DR\", \"Ast\", \"TO\", \"Stl\", \"Blk\", \"PF\")\n\nlosing_team_info = WRegularGames[,c(\"Season\",\"LTeamID\",\"LScore\", \"LFGM\", \"LFGA\", \"LFGM3\", \"LFGA3\", \"LFTM\", \"LFTA\", \"LOR\", \"LDR\", \"LAst\", \"LTO\", \"LStl\", \"LBlk\", \"LPF\")]\nnames(losing_team_info) = c(\"Season\",\"TeamID\",\"Score\", \"FGM\", \"FGA\", \"FGM3\", \"FGA3\", \"FTM\", \"FTA\", \"OR\", \"DR\", \"Ast\", \"TO\", \"Stl\", \"Blk\", \"PF\")\nWRegularGames = rbind(winning_team_info,losing_team_info)\n\nWteam_season_level_avg_summary = WRegularGames %>% group_by(Season, TeamID) %>% summarise_at(vars(\"Score\", \"FGM\", \"FGA\", \"FGM3\", \"FGA3\", \"FTM\", \"FTA\", \"OR\", \"DR\", \"Ast\", \"TO\", \"Stl\", \"Blk\", \"PF\"), mean)\n\n\n#Winning Team Join\nWtournament_data <- left_join(Wtournament_data, Wteam_season_level_avg_summary, by=c(\"WTeamID\"=\"TeamID\", \"Season\"))\n\nnames(Wtournament_data)[names(Wtournament_data)==\"Score\"] <- \"Season_AVG_WScore\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGM\"] <- \"Season_AVG_WFGM\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGA\"] <- \"Season_AVG_WFGA\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGM3\"] <- \"Season_AVG_WFGM3\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGA3\"] <- \"Season_AVG_WFGA3\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FTM\"] <- \"Season_AVG_WFTM\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FTA\"] <- \"Season_AVG_WFTA\"\nnames(Wtournament_data)[names(Wtournament_data)==\"OR\"] <- \"Season_AVG_WOR\"\nnames(Wtournament_data)[names(Wtournament_data)==\"DR\"] <- \"Season_AVG_WDR\"\nnames(Wtournament_data)[names(Wtournament_data)==\"Ast\"] <- \"Season_AVG_WAst\"\nnames(Wtournament_data)[names(Wtournament_data)==\"TO\"] <- \"Season_AVG_WTO\"\nnames(Wtournament_data)[names(Wtournament_data)==\"Stl\"] <- \"Season_AVG_WStl\"\nnames(Wtournament_data)[names(Wtournament_data)==\"Blk\"] <- \"Season_AVG_WBlk\"\nnames(Wtournament_data)[names(Wtournament_data)==\"PF\"] <- \"Season_AVG_WPF\"\n\n#Losing Team Join\nWtournament_data <- left_join(Wtournament_data, Wteam_season_level_avg_summary, by=c(\"LTeamID\"=\"TeamID\", \"Season\"))\nnames(Wtournament_data)[names(Wtournament_data)==\"Score\"] <- \"Season_AVG_LScore\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGM\"] <- \"Season_AVG_LFGM\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGA\"] <- \"Season_AVG_LFGA\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGM3\"] <- \"Season_AVG_LFGM3\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FGA3\"] <- \"Season_AVG_LFGA3\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FTM\"] <- \"Season_AVG_LFTM\"\nnames(Wtournament_data)[names(Wtournament_data)==\"FTA\"] <- \"Season_AVG_LFTA\"\nnames(Wtournament_data)[names(Wtournament_data)==\"OR\"] <- \"Season_AVG_LOR\"\nnames(Wtournament_data)[names(Wtournament_data)==\"DR\"] <- \"Season_AVG_LDR\"\nnames(Wtournament_data)[names(Wtournament_data)==\"Ast\"] <- \"Season_AVG_LAst\"\nnames(Wtournament_data)[names(Wtournament_data)==\"TO\"] <- \"Season_AVG_LTO\"\nnames(Wtournament_data)[names(Wtournament_data)==\"Stl\"] <- \"Season_AVG_LStl\"\nnames(Wtournament_data)[names(Wtournament_data)==\"Blk\"] <- \"Season_AVG_LBlk\"\nnames(Wtournament_data)[names(Wtournament_data)==\"PF\"] <- \"Season_AVG_LPF\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"WTournamentSeeds <- mutate(WTournamentSeeds, Seed = Seed)\n\nWTournamentSeeds$Seed <- as.numeric(gsub(\"\\\\D\",\"\", WTournamentSeeds$Seed))\n\nWtournament_data <- left_join(Wtournament_data, WTournamentSeeds, by=c(\"Season\", \"WTeamID\"=\"TeamID\"))\nnames(Wtournament_data)[names(Wtournament_data)==\"Seed\"] <- \"WSeed\"\n\nWtournament_data <- left_join(Wtournament_data, WTournamentSeeds, by=c(\"Season\", \"LTeamID\"=\"TeamID\"))\nnames(Wtournament_data)[names(Wtournament_data)==\"Seed\"] <- \"LSeed\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Create TeamID -> CityName Table\nWcity_games <- left_join(WGameCities, Wcity_names)\nWcity_games_winners <- left_join(Wcity_games, WRegularGamesCompactResult)\nWcity_games_winners <- Wcity_games_winners[c(-5,-1,-2,-12,-10,-9)] %>% filter(WLoc!=\"N\")\nWcity_games_winners$WCityID <- ifelse(Wcity_games_winners$WLoc==\"H\", Wcity_games_winners$CityID, 0)\n\n\nWcity_table <- data.frame(\"TeamID\"=(Wcity_games_winners$WTeamID), \"CityID\"=Wcity_games_winners$WCityID) %>% filter(CityID!=0)\n\n#Only keep one combination of TeamID and CityID\nWcity_table <- unique(Wcity_table)\nWcity_table <- left_join(Wcity_table, Wcity_names)\n\n#Load Winning Locations into Tournament Data\nWtournament_data <- left_join(Wtournament_data, Wcity_table, by=c(\"WTeamID\"=\"TeamID\"))\nnames(Wtournament_data)[names(Wtournament_data)==\"CityID\"] = \"WCityID\"\nnames(Wtournament_data)[names(Wtournament_data)==\"City\"] = \"WCity\"\nnames(Wtournament_data)[names(Wtournament_data)==\"State\"] = \"WState\"\n\nWtournament_data <- left_join(Wtournament_data, Wcity_table, by=c(\"LTeamID\"=\"TeamID\"))\nnames(Wtournament_data)[names(Wtournament_data)==\"CityID\"] = \"LCityID\"\nnames(Wtournament_data)[names(Wtournament_data)==\"City\"] = \"LCity\"\nnames(Wtournament_data)[names(Wtournament_data)==\"State\"] = \"LState\"\n\n#Load game locations into tournament data\nWtournament_data <-left_join(Wtournament_data, Wcity_games[-5])\n\nWtournament_data\n\nfor (i in 1:length(Wtournament_data$WTeamID))\n{\n  miles <- NA\n  if (is.na(Wtournament_data$City[i])==FALSE)\n  {\n    if (Wtournament_data$WState[i]==\"HI\" || Wtournament_data$State[i]==\"HI\" || Wtournament_data$WState[i]==\"PR\" || Wtournament_data$State[i]==\"PR\" || Wtournament_data$WState[i]==\"BA\" || Wtournament_data$State[i]==\"BA\" || Wtournament_data$WState[i]==\"VI\" || Wtournament_data$State[i]==\"VI\")\n    {\n      miles <- NA\n    }\n    else\n      {\n    dist <- mapdist(paste(Wtournament_data$WCity[i], Wtournament_data$WState[i], sep=\", \"), paste(Wtournament_data$City[i], Wtournament_data$State[i], sep=\", \"), mode=\"driving\")\n    miles <- dist[1,5]\n    }\n  }\n  Wtournament_data$WMiles[i] <- as.numeric(miles)\n}\n\n\n#Losing Team Distance\nfor (i in 1:length(Wtournament_data$LTeamID))\n{\n  miles <- NA\n  if (is.na(Wtournament_data$City[i])==FALSE)\n  {\n    if (Wtournament_data$LState[i]==\"HI\" || Wtournament_data$State[i]==\"HI\" || Wtournament_data$LState[i]==\"PR\" || Wtournament_data$State[i]==\"PR\" || Wtournament_data$LState[i]==\"BA\" || Wtournament_data$State[i]==\"BA\" || Wtournament_data$LState[i]==\"VI\" || Wtournament_data$State[i]==\"VI\")\n    {\n      miles <- NA\n    }\n    else\n    {\n      dist <- mapdist(paste(Wtournament_data$LCity[i], Wtournament_data$LState[i], sep=\", \"), paste(Wtournament_data$City[i], Wtournament_data$State[i], sep=\", \"), mode=\"driving\")\n      miles <- dist[1,5]\n    }\n  }\n  Wtournament_data$LMiles[i] <- as.numeric(miles)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Wtournament_data$Upset_potential <- ifelse(abs(Wtournament_data$WSeed-Wtournament_data$LSeed)>3, 1, 0)\n\nWtournament_data$Cinderella_potential <- ifelse(((Wtournament_data$WSeed>7 | Wtournament_data$LSeed>7) & Wtournament_data$DayNum>137), 1, 0)\n\nWtournament_data$Upset <- ifelse((Wtournament_data$WSeed-Wtournament_data$LSeed)>3, 1, 0)\n\nWtournament_data$Cinderella <- ifelse(Wtournament_data$WSeed>7 & Wtournament_data$DayNum>137, 1, 0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nWupsets <- filter(Wtournament_data, Wtournament_data$WSeed>Wtournament_data$LSeed)\nWnonupsets <- filter(Wtournament_data, Wtournament_data$WSeed<Wtournament_data$LSeed)\n\n# Renaming Losers to Higher Seed\n\nnames(Wupsets)[names(Wupsets)==\"LTeamID\"] <- \"HTeamID\"\nnames(Wupsets)[names(Wupsets)==\"LScore\"] <- \"HScore\"\nnames(Wupsets)[names(Wupsets)==\"LFGM\"] <- \"HFGM\"\nnames(Wupsets)[names(Wupsets)==\"LFGA\"] <- \"HFGA\"\nnames(Wupsets)[names(Wupsets)==\"LFGM3\"] <- \"HFGM3\"\nnames(Wupsets)[names(Wupsets)==\"LFGA3\"] <- \"HFGA3\"\nnames(Wupsets)[names(Wupsets)==\"LFTM\"] <- \"HFTM\"\nnames(Wupsets)[names(Wupsets)==\"LFTA\"] <- \"HFTA\"\nnames(Wupsets)[names(Wupsets)==\"LOR\"] <- \"HOR\"\nnames(Wupsets)[names(Wupsets)==\"LDR\"] <- \"HDR\"\nnames(Wupsets)[names(Wupsets)==\"LAst\"] <- \"HAst\"\nnames(Wupsets)[names(Wupsets)==\"LTO\"] <- \"HTO\"\nnames(Wupsets)[names(Wupsets)==\"LStl\"] <- \"HStl\"\nnames(Wupsets)[names(Wupsets)==\"LBlk\"] <- \"HBlk\"\nnames(Wupsets)[names(Wupsets)==\"LPF\"] <- \"HPF\"\nnames(Wupsets)[names(Wupsets)==\"LSeed\"] <- \"HSeed\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LScore\"] <- \"Season_AVG_HScore\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LFGM\"] <- \"Season_AVG_HFGM\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LFGA\"] <- \"Season_AVG_HFGA\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LFGM3\"] <- \"Season_AVG_HFGM3\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LFGA3\"] <- \"Season_AVG_HFGA3\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LFTM\"] <- \"Season_AVG_HFTM\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LFTA\"] <- \"Season_AVG_HFTA\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LOR\"] <- \"Season_AVG_HOR\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LDR\"] <- \"Season_AVG_HDR\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LAst\"] <- \"Season_AVG_HAst\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LTO\"] <- \"Season_AVG_HTO\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LStl\"] <- \"Season_AVG_HStl\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LBlk\"] <- \"Season_AVG_HBlk\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_LPF\"] <- \"Season_AVG_HPF\"\nnames(Wupsets)[names(Wupsets)==\"LCityID\"] <- \"HCityID\"\nnames(Wupsets)[names(Wupsets)==\"LCity\"] <- \"HCity\"\nnames(Wupsets)[names(Wupsets)==\"LState\"] <- \"HState\"\nnames(Wupsets)[names(Wupsets)==\"LMiles\"] <- \"HMiles\"\n\n# Renaming Winners to Lower Seed\n\nnames(Wupsets)[names(Wupsets)==\"WTeamID\"] <- \"LTeamID\"\nnames(Wupsets)[names(Wupsets)==\"WScore\"] <- \"LScore\"\nnames(Wupsets)[names(Wupsets)==\"WFGM\"] <- \"LFGM\"\nnames(Wupsets)[names(Wupsets)==\"WFGA\"] <- \"LFGA\"\nnames(Wupsets)[names(Wupsets)==\"WFGM3\"] <- \"LFGM3\"\nnames(Wupsets)[names(Wupsets)==\"WFGA3\"] <- \"LFGA3\"\nnames(Wupsets)[names(Wupsets)==\"WFTM\"] <- \"LFTM\"\nnames(Wupsets)[names(Wupsets)==\"WFTA\"] <- \"LFTA\"\nnames(Wupsets)[names(Wupsets)==\"WOR\"] <- \"LOR\"\nnames(Wupsets)[names(Wupsets)==\"WDR\"] <- \"LDR\"\nnames(Wupsets)[names(Wupsets)==\"WAst\"] <- \"LAst\"\nnames(Wupsets)[names(Wupsets)==\"WTO\"] <- \"LTO\"\nnames(Wupsets)[names(Wupsets)==\"WStl\"] <- \"LStl\"\nnames(Wupsets)[names(Wupsets)==\"WBlk\"] <- \"LBlk\"\nnames(Wupsets)[names(Wupsets)==\"WPF\"] <- \"LPF\"\nnames(Wupsets)[names(Wupsets)==\"WSeed\"] <- \"LSeed\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WScore\"] <- \"Season_AVG_LScore\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WFGM\"] <- \"Season_AVG_LFGM\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WFGM\"] <- \"Season_AVG_LFGM\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WFGA\"] <- \"Season_AVG_LFGA\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WFGM3\"] <- \"Season_AVG_LFGM3\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WFGA3\"] <- \"Season_AVG_LFGA3\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WFTM\"] <- \"Season_AVG_LFTM\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WFTA\"] <- \"Season_AVG_LFTA\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WOR\"] <- \"Season_AVG_LOR\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WDR\"] <- \"Season_AVG_LDR\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WAst\"] <- \"Season_AVG_LAst\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WTO\"] <- \"Season_AVG_LTO\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WStl\"] <- \"Season_AVG_LStl\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WBlk\"] <- \"Season_AVG_LBlk\"\nnames(Wupsets)[names(Wupsets)==\"Season_AVG_WPF\"] <- \"Season_AVG_LPF\"\nnames(Wupsets)[names(Wupsets)==\"WSeed\"] <- \"LSeed\"\nnames(Wupsets)[names(Wupsets)==\"WCityID\"] <- \"LCityID\"\nnames(Wupsets)[names(Wupsets)==\"WCity\"] <- \"LCity\"\nnames(Wupsets)[names(Wupsets)==\"WState\"] <- \"LState\"\nnames(Wupsets)[names(Wupsets)==\"WMiles\"] <- \"LMiles\"\n\n\n# Renaming Winners to Higher Seeds\n\nnames(Wnonupsets)[names(Wnonupsets)==\"WTeamID\"] <- \"HTeamID\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WScore\"] <- \"HScore\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WFGM\"] <- \"HFGM\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WFGA\"] <- \"HFGA\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WFGM3\"] <- \"HFGM3\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WFGA3\"] <- \"HFGA3\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WFTM\"] <- \"HFTM\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WFTA\"] <- \"HFTA\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WOR\"] <- \"HOR\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WDR\"] <- \"HDR\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WAst\"] <- \"HAst\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WTO\"] <- \"HTO\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WStl\"] <- \"HStl\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WBlk\"] <- \"HBlk\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WPF\"] <- \"HPF\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WSeed\"] <- \"HSeed\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WScore\"] <- \"Season_AVG_HScore\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WFGM\"] <- \"Season_AVG_HFGM\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WFGA\"] <- \"Season_AVG_HFGA\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WFGM3\"] <- \"Season_AVG_HFGM3\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WFGA3\"] <- \"Season_AVG_HFGA3\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WFTM\"] <- \"Season_AVG_HFTM\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WFTA\"] <- \"Season_AVG_HFTA\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WOR\"] <- \"Season_AVG_HOR\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WDR\"] <- \"Season_AVG_HDR\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WAst\"] <- \"Season_AVG_HAst\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WTO\"] <- \"Season_AVG_HTO\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WStl\"] <- \"Season_AVG_HStl\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WBlk\"] <- \"Season_AVG_HBlk\"\nnames(Wnonupsets)[names(Wnonupsets)==\"Season_AVG_WPF\"] <- \"Season_AVG_HPF\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WSeed\"] <- \"HSeed\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WCityID\"] <- \"HCityID\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WCity\"] <- \"HCity\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WState\"] <- \"HState\"\nnames(Wnonupsets)[names(Wnonupsets)==\"WMiles\"] <- \"HMiles\"\n\n\n\nWfinal <- rbind(Wupsets, Wnonupsets)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nWregression_data <- Wfinal \n\nnames(Wregression_data)[names(Wregression_data)==\"Upset\"] <- \"Upset_result\"\nnames(Wregression_data)[names(Wregression_data)==\"Cinderella\"] <- \"Cinderella.\"\n\n\nWupset_regression_data = Wregression_data[Wregression_data$Upset_potential==1,]\nWupset_regression_data_cleaned = na.omit(Wupset_regression_data)\n\nWcinderella_regression_data = Wregression_data[Wregression_data$Cinderella_potential==1,]\nWcinderella_regression_data_cleaned = na.omit(Wcinderella_regression_data)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Wupset_logit_model <- glm(Upset_result ~ HMiles + LMiles + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = Wupset_regression_data_cleaned, family = \"binomial\")\nWupset_logit_model_stepwise <- stepAIC(Wupset_logit_model)\nsummary(Wupset_logit_model_stepwise)\npR2(Wupset_logit_model_stepwise)\n\nWcinderella_logit_model <- glm(Cinderella. ~ HMiles + LMiles + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = Wcinderella_regression_data_cleaned, family = \"binomial\")\nWcinderella_logit_model_stepwise <- stepAIC(Wcinderella_logit_model)\nsummary(Wcinderella_logit_model_stepwise)\npR2(Wcinderella_logit_model_stepwise)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Wupset_forest_data = Wupset_regression_data\nWupset_forest_data$Upset_result = as.factor(Wupset_forest_data$Upset_result)\n\nWupset_forest_data_split <- sample(nrow(Wupset_forest_data), 0.7*nrow(Wupset_forest_data), replace = FALSE)\nWupset_forest_training_data <- Wupset_forest_data[Wupset_forest_data_split,]\nWupset_forest_validation_data <- Wupset_forest_data[-Wupset_forest_data_split,]\n\nWupset_forest_model <- randomForest(Upset_result ~ HMiles + LMiles + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = Wupset_forest_training_data, importance = TRUE,na.action=na.exclude)\nWupset_forest_model\n\nvarImpPlot(Wupset_forest_model)\n\n\nWcinderella_forest_data = Wcinderella_regression_data\nWcinderella_forest_data$Cinderella. = as.factor(Wcinderella_forest_data$Cinderella.)\n\nWcinderella_forest_data_split <- sample(nrow(Wcinderella_forest_data), 0.7*nrow(Wcinderella_forest_data), replace = FALSE)\nWcinderella_forest_training_data <- Wcinderella_forest_data[Wcinderella_forest_data_split,]\nWcinderella_forest_validation_data <- Wcinderella_forest_data[-Wcinderella_forest_data_split,]\n\nWcinderella_forest_model <- randomForest(Cinderella. ~ HMiles + LMiles + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF, data = Wcinderella_forest_training_data, importance = TRUE,na.action=na.exclude)\nWcinderella_forest_model\n\nvarImpPlot(Wcinderella_forest_model)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Wupset_bas <- bas.lm(Upset_result ~ HMiles + LMiles + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF,\n                     data = Wupset_regression_data,\n                     method = \"MCMC\",\n                     prior = \"ZS-null\",\n                     modelprior = uniform())\nsummary(Wupset_bas)\nimage(Wupset_bas, rotate=F,drop.always.included=TRUE, top.models=5)\n\n\nWcinderella_bas <- bas.lm(Cinderella. ~ HMiles + LMiles + Season_AVG_HScore + Season_AVG_HFGM + Season_AVG_HFGA + Season_AVG_HFGM3 + Season_AVG_HFGA3 + Season_AVG_HFTM + Season_AVG_HFTA + Season_AVG_HOR + Season_AVG_HDR + Season_AVG_HAst + Season_AVG_HTO + Season_AVG_HStl + Season_AVG_HBlk + Season_AVG_HPF + Season_AVG_LScore + Season_AVG_LFGM + Season_AVG_LFGA + Season_AVG_LFGM3 + Season_AVG_LFGA3 + Season_AVG_LFTM + Season_AVG_LFTA + Season_AVG_LOR + Season_AVG_LDR + Season_AVG_LAst + Season_AVG_LTO + Season_AVG_LStl + Season_AVG_LBlk + Season_AVG_LPF,\n                    data = Wcinderella_regression_data,\n                    method = \"MCMC\",\n                    prior = \"ZS-null\",\n                    modelprior = uniform())\nsummary(Wcinderella_bas)\nimage(Wcinderella_bas, rotate=F, drop.always.included=TRUE, top.models=5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Women's results analysis & comparison to men's\n\nAs with the men's data, we used both random forests and Bayesian multiple linear regression models to analyze upsets and cinderellas from the women's competition. One immediate takeaway we noticed was that the error rates from the random forest were much lower (almost 0) for the women's data compared to the ~25% error rates for men's in both the upset and cinderella models. This coupled with the high posterior inclusion probabilities for the variables in the women's Bayesian analysis leads us to believe that the women's NCAA tournament is much more predictable given the data we have been tracking.\n\nOne statistic that could be valuable in predicting both upsets and cinderellas is the number of assists per game by the lower seed, suggesting that underdogs who play as a team and don't rely on any one star player are more capable of pulling an upset. Talent in the women's game is more concentrated within the top tier of teams than it is on the men's side of things - it's less likely that an underdog in the women's tournament could have a star player to take over the contest than for a lower seeded team in the men's bracket to do the same. Thus, it's important for the team to play together, as a team, facilitating more assists. It's also worth noting that distance doesn't play as important of a role as it does in the men's game - this could be attributed to the smaller number of fans that follow women's basketball. \n\nFurther, we noticed that while blocks were very important for both men and women in predicting upsets and cinderellas, distance was only important for men. we hypothesize that this could be because the fans follow the men's team much more intensely and changing locations will mean a much larger change in atmosphere and energy for the men's teams. Finally, we noticed, like the men's, the women's tournament upsets and cinderellas were much more dependent on the higher seeded teams stats, revealing the importance of a vulnerable high seeded team rather than an unlikely performance by the lower seeded team. "},{"metadata":{},"cell_type":"markdown","source":"# Conclusion"},{"metadata":{},"cell_type":"markdown","source":"For next steps, we propose accounting for a few key adjustments. First, in order to have a full women’s dataset that can be comparable to the men’s data, we would need access to women’s coaching data. This will give us insight into specific team’s coaching experience and any coaching disruptions during the season.\n\nSecond, most of our analysis was done at a game-level. We can certainly get more granular by exploring play-by-play data and seeing if different teams perform at different levels throughout a single game.\n\nThird, we can and should include player data where available. Being able to analyze a team’s potential matchup at a player and position level could certainly offer insights into the potential for an upset.\n\nFourth, we can take a step back and explore information related to the school itself, such as fan base or other programs’ success at the same school. This can give us a proxy for how much the school cares about their sports and invests in their successes.\n\nMoving forward, we should future test this model in future seasons to get an accurate understanding of its performance in an out-of-sample dataset.\n\nWe should also keep in mind that this is a time-stamped analysis. As the game and rules of the game change, so will these factors. As time passes, different teams will emphasize different strategies that will favor different skills in the end. As such, we should consider our analysis in the context of the game as it currently stands and be prepared to adjust it in the future.\n\nWith that in mind, our analysis can be applied to the game in three ways: one, it shows the gaps in ranking quality when it comes to predicting upsets and Cinderella events; two, it can help coaches whose teams may be unfavored understand what factors could lead to a win in the tournament; three, help women’s coaches specifically focus on specific skills in-season to develop an edge during the tournament.\n\nAs UVA students, thinking back to when we all watched the semi-finals in Chirag’s apartment last year, we can find a new level of appreciation for the game. A block was not just something to cheer about, but a serious indicator of where we were headed. We can now view a failed three-point attempts not as the end of the world, but rather just a risk with high upside. Ultimately, this research has supplemented our emotional connection to March Madness with a data-backed understanding of the game. "}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":4}