{"cells":[{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"#Libraries to use.\nsuppressMessages(library(janitor))\nsuppressMessages(library(tidyverse,warn.conflicts = FALSE))\nsuppressMessages(library(xgboost))\nsuppressMessages(library(MLmetrics))\noptions(warn=-1)\nsuppressMessages(library(dplyr, warn.conflicts = FALSE))\nsuppressMessages(library(plotly))\n\n#Helper function for figures\nfig <- function(width, heigth){\n     options(repr.plot.width = width, repr.plot.height = heigth)\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Competitiveness in NCAA and European Competitions"},{"metadata":{},"cell_type":"markdown","source":""},{"metadata":{},"cell_type":"markdown","source":"# 0. Executive Summary\n\nThis notebook compares competitiveness between NCAA Regular Season, NCAA Tournament, Euroleague and Eurocup competitions at man's basketball between 2012 and 2019 using approximately 70000 games from NCAA Regular season, 800 from the NCAA-Tournament, 3000 from the Euroleague and 1700 from the Eurocup. I defined competitiveness is having teams in a given league that are close to each other in key metrics from bottom to top ranks. I used two different methods to quantify this, simply averaging key metrics and looking their differences between 10th and 90th quartiles is my first method. Secondly, I used Dean Oliver's four factors which mainly focusing on scoring, second chances (offensive rebounds), not losing the ball (turnovers) and free throws with different weights in order to quantify competitiveness. I looked teams difference at the given leagues from top to bottom. 10th and 90th quartiles are selected intuitively in order to avoid outliers.\n\nNCAA data is provided by the Kaggle and the Euroleague data is scraped by me with help of several github repos which are credited at the dataset.\n\nAs result of my analysis, I have seen that, regular season of the NCAA is the most competitive tournament comparing to other leagues. Only at 2018, Euroleague became more competitive at both metrics. European competitions which consist professional players have better performance in key metrics, nevertheless have bigger margin between top and bottom teams.\n\nAfter analyzing competitiveness between leagues, I tried to generalize results by predicting Euroleague games with NCAA games and the other way around. NCAA games were to some extent helpful to predict European competitions with 66% of the accuracy while it was much harder to predict NCAA games using only the European competitions data with no better performance than random guessing.\n\n# 1. Warm-up:Introduction\n\nBasketball is one of the most popular sports in Europe and two of the top competition is Euroleague and Eurocup. Every year, several NCAA and NBA players continue their career at Europe's top clubs and top European players moves to NBA (e.g. Luka Doncic, Bogdan Bogdanovic). Even though competitiveness between college basketball and top European competitions which consists fully professional players are completely different, I think it is still interesting to make a comparison.\n\nMy analysis will focus on three aspects:\n\n- Competitiveness within these leagues and aiming to extract insights on their competitiveness.\n- Made an exploratory analysis on players who pursue their career in Europe after college and check what are the common aspects of the ones who started a successful career at old continent.\n- Create a baseline model to predict results of the games and see if the same features generalize to predict the other competition.\n\n## 1.1. What is the NCAA?\n\nThe National Collegiate Athletic Association is a member-led organization dedicated to the well-being and lifelong success of college athletes according to it [website](http://www.ncaa.org/about/resources/media-center/ncaa-101/what-ncaa). Even though it is mostly known by basketball tournament and March Madness at least in some parts of the world, the NCAA consists 24 sport branches.\n\n## 1.2. What is Turkish Airlines Euroleague and 7 Days Eurocup? Who Plays There?\n\nMany people knew about the NCAA due to this competition, their personal interest and amazing notebooks from the Kaggle community however European basketball could be an uncharted territory for many non-Europeans. Before moving on my analysis let me introduce you the Turkish Airlines Euroleague and the 7 Days Eurocup, Europe's most prestigious basketball tournaments. Every year top teams from every European countries joined the league with different licence types. 11 clubs have A Licence which gives them long term participation to the tournament due to their investment to the game and previous track racord. These teams are the followings at the moment: Barcelona (Spain), Olympiacos (Greece), Anadolu Efes (Turkey), Žalgiris (Lithuania),Kirolbet Baskonia (Spain), Panathinaikos OPAP (Greece), Fenerbahçe Beko (Turkey),\tCSKA Moscow (Russia), Real Madrid (Spain), Maccabi FOX Tel Aviv (Israel), A|X Armani Exchange Milan (Italy). In addition to them, every year 2-3 teams join with Wild Card from Euroleague Basketball management along with champion of the Adriatic League, VTB League (top Russian competition) and Eurocup champion. If winner of those leagues have already A Licence, the runner-up joins the competition.\n\nAt Euroleague, currently 18 team competes and top 8 reaches to the playoffs, however format changes every couple of years. At the moment after the league, top-8 teams play best-of-5 series between each other and then reach to the Final-4 weekend at a major European city. Final-4 is on single match elimination format and winner became European champion similar to the end of March Madness.\n\nOn the other hand, Eurocup is the competition where top 24 second-tier teams participate from various European countries in four different groups of six. Best of four advance through the top 16 and compete to reach top two position at their group of four. Then Quarter-finals, Semi-finals and finals are played on best-of-three series to get the championship and Euroleague wild card.\n\n## 1.3. Datasets\n\nTwo different data sources are used at the analysis. First one is the NCAA dataset provided by the competition. \n\nSecond one is for Euroleague and Eurocup games, I scraped play by play statistics from Euroleague website (over 2 million rows) and uploaded to [Kaggle](https://www.kaggle.com/efehandanisman/euroleague-play-by-play-data-20072020]). You can also find the short script on how I scraped the [dataset](https://www.kaggle.com/efehandanisman/diy-scraping-euroleague-games).\n \nIn order to align the two different data sources, I used games from 2012 to 2018 where full seasons are exist for year by year analysis. Since the Eurocup data starts from 2012, for the aggregate analysis I started from 2012 till today."},{"metadata":{},"cell_type":"markdown","source":""},{"metadata":{},"cell_type":"markdown","source":"# 2. Setting-up the Court for Play: Data Preparation\n\n## 2.1. Data Prep For Aggregate Stats"},{"metadata":{},"cell_type":"markdown","source":"Here we prepare compact results for preliminary analysis. Since the Eurocup data starts from 2012, I will take my starting point for NCAA from 2012 as well. If you would like to see results of my insights immediately I suggest you to move past to exploratory analysis part from here."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Get NCAA games and filter between 2012 and 2019 which is our base.\n\nreg_season_stats <- read.csv(\"../input/march-madness-analytics-2020/MDataFiles_Stage2/MRegularSeasonCompactResults.csv\", stringsAsFactors = FALSE)\nteamnames <- read.csv(\"../input/march-madness-analytics-2020/MDataFiles_Stage2/MTeams.csv\", stringsAsFactors = FALSE)\ntourney_stats_sec <- read.csv(\"../input/march-madness-analytics-2020/MDataFiles_Stage2/MSecondaryTourneyCompactResults.csv\", stringsAsFactors = FALSE)\ntourney_stats_pri <- read.csv(\"../input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneyCompactResults.csv\", stringsAsFactors = FALSE)\nreg_season_detstats <- read.csv(\"../input/march-madness-analytics-2020/MDataFiles_Stage2/MRegularSeasonDetailedResults.csv\", stringsAsFactors = FALSE)\ntourney_stats_detpri <- read.csv(\"../input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneyDetailedResults.csv\", stringsAsFactors = FALSE)\n\nreg_season_stats$Season_type <- \"Regular\" \ntourney_stats_pri$Season_type <- \"Tournament\" \ncompact_results_ncaa <- rbind(reg_season_stats,tourney_stats_pri)\ncompact_results_ncaa <- compact_results_ncaa %>% filter(Season >= 2012 & Season < 2019)\nreg_season_detstats$Season_type <- \"Regular\" \ntourney_stats_detpri$Season_type <- \"Tournament\" \ndetailed_res_ncaa <- rbind(reg_season_detstats,tourney_stats_detpri)\ndetailed_res_ncaa <- detailed_res_ncaa %>% filter(Season >= 2012 & Season < 2019)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"#NCAA Opponent Rebounds Year By Year. A feature we need for four factor analysis.\n\nopp_rebounds_ncaa1 <- detailed_res_ncaa %>% select(WTeamID,Season_type,Season,LTeamID,WOR,WDR,LOR,LDR)  %>% group_by(WTeamID,Season_type,Season) %>% summarize(Opp_Off_Reb = sum(LOR),Opp_Def_Reb = sum(LDR)) %>% rename(\"TeamID\"=\"WTeamID\")\nopp_rebounds_ncaa2<-  detailed_res_ncaa %>% select(WTeamID,Season_type,Season,LTeamID,WOR,WDR,LOR,LDR)  %>% group_by(LTeamID,Season_type,Season) %>% summarize(Opp_Off_Reb = sum(WOR),Opp_Def_Reb = sum(WDR)) %>% rename(\"TeamID\"=\"LTeamID\")\nopp_rebounds_ncaa_yby <- rbind(opp_rebounds_ncaa1,opp_rebounds_ncaa2)\nopp_rebounds_ncaa_yby <- opp_rebounds_ncaa_yby %>% group_by(TeamID,Season_type,Season) %>% summarize(opp_offreb = sum(Opp_Off_Reb),opp_defreb = sum(Opp_Def_Reb))\n\n# Create features from detailed statistics from the NCAA data. We create them for teams and do not need match-up's for the moment.\nL <- grepl('L', colnames(detailed_res_ncaa))\nwin_stats <- detailed_res_ncaa[!L]\nnames(win_stats) <- gsub(\"W\", \"\", names(win_stats), fixed = TRUE)\nwin_stats <- win_stats %>% select(-DayNum,-NumOT)\n\nW <- grepl('W', colnames(detailed_res_ncaa))\nlose_stats <- detailed_res_ncaa[!W]\nnames(lose_stats) <- gsub(\"L\", \"\", names(lose_stats), fixed = TRUE)\nlose_stats <- lose_stats %>% select (-DayNum,-NumOT)\nncaa_stats <- rbind(win_stats,lose_stats) %>% select(-Season)\nncaa_totalgames <- ncaa_stats %>% group_by(TeamID,Season_type) %>%   summarize(TOTAL_GAMES=n()) %>% arrange(desc(TOTAL_GAMES))\n\nncaa_stats <- ncaa_stats %>% group_by(TeamID,Season_type) %>% summarise_all(funs(sum)) %>% mutate(TWOPTS_PERC = round(FGM / (FGM+ FGA),2),THRPTS_PERC = round(FGM3 / (FGM3 + FGA3),2), FT_PERC = round(FTM /  (FTM + FTA),2)) %>% left_join(ncaa_totalgames, by = c(\"TeamID\"=\"TeamID\",\"Season_type\"=\"Season_type\")) %>% \nmutate(SCORE_PG = round(Score / TOTAL_GAMES,2),TWOPTS_ATTMPT_PG = round((FGA - FGA3) / TOTAL_GAMES,2), TWOPTS_MAKE_PG= round((FGM - FGM3) / TOTAL_GAMES,2),THR_MAKE_PG = round(FGM3 / TOTAL_GAMES,2),\n       THR_ATTMPT_PG= round(FGA3 /TOTAL_GAMES,2),FT_ATTEMPS_PG = round(FTA / TOTAL_GAMES,2 ), FT_MAKE_PG = round(FTM / TOTAL_GAMES,2), ASST_PG = round(Ast / TOTAL_GAMES,2),DRBND_PG = round (DR / TOTAL_GAMES,2), ORBND_PG = round(OR / TOTAL_GAMES,2),FOUL_CMMT_PG = round(PF / TOTAL_GAMES,2),TO_PG = round(TO/TOTAL_GAMES,2), \n       ST_PG = round(Stl / TOTAL_GAMES,2), FG_ATTEMP_PG = round(FT_ATTEMPS_PG +TWOPTS_ATTMPT_PG + THR_ATTMPT_PG, 3), BLCK_PG = round(Blk /TOTAL_GAMES,2)) %>% left_join(teamnames, by=c(\"TeamID\"=\"TeamID\")) %>% select(-FirstD1Season,-LastD1Season)\n\n# NCAA Opponent Rebounds. A feature we need for four factor analysis.\n\nopp_rebounds_ncaa1 <- detailed_res_ncaa %>% select(WTeamID,Season_type,LTeamID,WOR,WDR,LOR,LDR)  %>% group_by(WTeamID,Season_type) %>% summarize(Opp_Off_Reb = sum(LOR),Opp_Def_Reb = sum(LDR)) %>% rename(\"TeamID\"=\"WTeamID\")\nopp_rebounds_ncaa2<-  detailed_res_ncaa %>% select(WTeamID,Season_type,LTeamID,WOR,WDR,LOR,LDR)  %>% group_by(LTeamID,Season_type) %>% summarize(Opp_Off_Reb = sum(WOR),Opp_Def_Reb = sum(WDR)) %>% rename(\"TeamID\"=\"LTeamID\")\nopp_rebounds_ncaa <- rbind(opp_rebounds_ncaa1,opp_rebounds_ncaa2)\nopp_rebounds_ncaa <- opp_rebounds_ncaa %>% group_by(TeamID,Season_type) %>% summarize(opp_offreb = sum(Opp_Off_Reb),opp_defreb = sum(Opp_Def_Reb))\n\n# Opponent Rebound feature joined\nncaa_stats <- ncaa_stats %>% left_join(opp_rebounds_ncaa,by=c(\"Season_type\"=\"Season_type\",\"TeamID\"=\"TeamID\")) %>% mutate(OPP_DREB_PG = round(opp_defreb / TOTAL_GAMES,2), OPP_OREB_PG =round(opp_offreb / TOTAL_GAMES,2))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Have look to the features! \nHere you can see aggregated statistics between 2012 and 2018 for each NCAA team with key metrics . We have free throw, two points and three points attemps, successful and failed shots, assists, rebounds, blocks,steals,committed fouls, turnovers, opponent rebounds, per game statistics here. Preparing this we will prepare the same for European competitions."},{"metadata":{"trusted":true},"cell_type":"code","source":"head(ncaa_stats)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Get European games and merge them into one dataframe.\neurocup <- read.csv(\"../input/euroleague-play-by-play-data-20072020/eurocup.csv\", stringsAsFactors = FALSE)\neuroleague <- read.csv(\"../input/euroleague-play-by-play-data-20072020/euroleague.csv\", stringsAsFactors = FALSE)\neuroleague$Tournament <- \"Euroleague\"\neurocup$Tournament <- \"Eurocup\"\neuroleague <- rbind(euroleague,eurocup)\neuroleague <- euroleague %>% filter(yer >=2012 & yer < 2019)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Due to the sponsorship changes, name changes is frequent at European basketball. So names are standardized below for the TEAM, TeamA and TeamB variables. Usually city of the team stays the same however the sponsor before the name changes frequently. This requires small research and some experience of following up European basketball."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Make all team names with capital letters.\neuroleague$TeamA <- str_to_title(euroleague$TeamA)\neuroleague$TeamB <- str_to_title(euroleague$TeamB)\neuroleague$TEAM <- str_to_title(euroleague$TEAM)\n\n#Fix names.\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Bamberg\",ignore_case=T)), \"Brose Bamberg\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Brose\",ignore_case=T)), \"Brose Bamberg\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Banvit\",ignore_case=T)), \"Bandirma BK\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Ulm\",ignore_case=T)), \"Ratiopharm Ulm\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Sevill\",ignore_case=T)), \"Real Betis Baloncesto\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Elan Chalon\",ignore_case=T)), \"Elan Chalon\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Kuban\",ignore_case=T)), \"Lokomotiv Kuban\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Roanne\",ignore_case=T)), \"Roanne\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Zagreb\",ignore_case=T)), \"KK Zagreb\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Olimpija\",ignore_case=T)), \"Union Olimpija\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Gran Canaria\",ignore_case=T)), \"Gran Canaria\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Fb\",ignore_case=T)),\"Fenerbahce Istanbul\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Artland\",ignore_case=T)), \"Artland Dragons\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Gescrap\",ignore_case=T)), \"Bilbao Basket\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Bilbao\",ignore_case=T)), \"Bilbao Basket\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Oldenburg\",ignore_case=T)), \"Ewe Baskets Oldenburg\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Valencia\",ignore_case=T)), \"Valencia Basket\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Spirou\",ignore_case=T)), \"Spirou Charleroi\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Spartak\",ignore_case=T)), \"Spartak Saint Petersburg\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Laboral\",ignore_case=T)), \"Baskonia\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Cantu\",ignore_case=T)), \"Cantu\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Kiev\",ignore_case=T)), \"Budivielnik Kiev\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Telenet\",ignore_case=T)), \"Oostende\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Nanterre\",ignore_case=T)), \"Nanterre\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Jerusalem\",ignore_case=T)), \"Hapoel Jerusalem\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Rom\",ignore_case=T)), \"Virtus Rome\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Panionios\",ignore_case=T)), \"Panionios Athens\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Darus\",ignore_case=T)), \"Darussafaka Istanbul\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Cibona\",ignore_case=T)), \"Cibona Zagreb\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Buducnost\",ignore_case=T)), \"Buducnost Podgorica\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Aris\",ignore_case=T)), \"Aris Thessaloniki\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Rytas\",ignore_case=T)), \"Lietuvos Rytas\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Besik\",ignore_case=T)), \"Besiktas Istanbul\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Maccabi\",ignore_case=T)), \"Maccabi Tel Aviv\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Barcelona\",ignore_case=T)), \"FC Barcelona\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Fener\",ignore_case=T)),\"Fenerbahce Istanbul\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Panat\",ignore_case=T)), \"Panathinaikos Athens\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"EFES\",ignore_case=T)), \"Anadolu Efes Istanbul\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"MILAN\",ignore_case=T)), \"Milano\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"VITORIA\",ignore_case=T)), \"Baskonia\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"CSKA\",ignore_case=T)), \"CSKA Moscow\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"TAU\",ignore_case=T)), \"Baskonia\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Berlin\",ignore_case=T)), \"Alba Berlin\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"REAL\",ignore_case=T)), \"Real Madrid\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"ASSECO\",ignore_case=T)), \"Arka Gdynia\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Prokom\",ignore_case=T)), \"Arka Gdynia\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"SLUC\",ignore_case=T)), \"Nancy\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"BASKON\",ignore_case=T)), \"Baskonia\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"ZVEZDA\",ignore_case=T)), \"Red Star Belgrade\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, \"Gala\"), \"Galatasaray Istanbul\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"Zal\",ignore_case=T)), \"Zalgiris Kaunas\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"PARTI\",ignore_case=T)), \"Partizan Belgrade\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"KHIMK\",ignore_case=T)), \"Khimki Moscow\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"UNION OLIMP\",ignore_case=T)), \"Union Olimpija\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"OLYMP\",ignore_case=T)), \"Olympiacos Piraeus\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"MONTE\",ignore_case=T)), \"Montepaschi Siena\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"ASVEL\",ignore_case=T)), \"Asvel\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"UNICA\",ignore_case=T)), \"Malaga\", euroleague$TEAM)\neuroleague$TEAM <- ifelse(str_detect(euroleague$TEAM, fixed(\"UNICS\",ignore_case=T)), \"Unics Kazan\", euroleague$TEAM)\n\n###\n\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Bamberg\",ignore_case=T)), \"Brose Bamberg\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Bilbao\",ignore_case=T)), \"Bilbao Basket\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Oldenburg\",ignore_case=T)), \"Ewe Baskets Oldenburg\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Valencia\",ignore_case=T)), \"Valencia Basket\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Spirou\",ignore_case=T)), \"Spirou Charleroi\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Spartak\",ignore_case=T)), \"Spartak Saint Petersburg\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Laboral\",ignore_case=T)), \"Baskonia\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Cantu\",ignore_case=T)), \"Cantu\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Ulm\",ignore_case=T)), \"Ratiopharm Ulm\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Sevill\",ignore_case=T)), \"Real Betis Baloncesto\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Elan Chalon\",ignore_case=T)), \"Elan Chalon\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Kuban\",ignore_case=T)), \"Lokomotiv Kuban\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Roanne\",ignore_case=T)), \"Roanne\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Zagreb\",ignore_case=T)), \"KK Zagreb\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Olimpija\",ignore_case=T)), \"Union Olimpija\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Gran Canaria\",ignore_case=T)), \"Gran Canaria\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Fb\",ignore_case=T)),\"Fenerbahce Istanbul\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Gescrap\",ignore_case=T)), \"Bilbao Basket\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Kiev\",ignore_case=T)), \"Budivielnik Kiev\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Telenet\",ignore_case=T)), \"Oostende\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Nanterre\",ignore_case=T)), \"Nanterre\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Jerusalem\",ignore_case=T)), \"Hapoel Jerusalem\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Rom\",ignore_case=T)), \"Virtus Rome\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Panionios\",ignore_case=T)), \"Panionios Athens\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Darus\",ignore_case=T)), \"Darussafaka Istanbul\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Cibona\",ignore_case=T)), \"Cibona Zagreb\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Buducnost\",ignore_case=T)), \"Buducnost Podgorica\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Aris\",ignore_case=T)), \"Aris Thessaloniki\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Berlin\",ignore_case=T)), \"Alba Berlin\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Artland\",ignore_case=T)), \"Artland Dragons\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"REAL\",ignore_case=T)), \"Real Madrid\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Brose\",ignore_case=T)), \"Brose Bamberg\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Banvit\",ignore_case=T)), \"Bandirma BK\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Rytas\",ignore_case=T)), \"Lietuvos Rytas\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Besik\",ignore_case=T)), \"Besiktas Istanbul\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Maccabi\",ignore_case=T)), \"Maccabi Tel Aviv\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Barcelona\",ignore_case=T)), \"FC Barcelona\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Fener\",ignore_case=T)),\"Fenerbahce Istanbul\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Panat\",ignore_case=T)), \"Panathinaikos Athens\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"EFES\",ignore_case=T)), \"Anadolu Efes Istanbul\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"MILAN\",ignore_case=T)), \"Milano\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"VITORIA\",ignore_case=T)), \"Baskonia\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"CSKA\",ignore_case=T)), \"CSKA Moscow\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"TAU\",ignore_case=T)), \"Baskonia\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"ASSECO\",ignore_case=T)), \"Arka Gdynia\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Prokom\",ignore_case=T)), \"Arka Gdynia\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"SLUC\",ignore_case=T)), \"Nancy\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"BASKON\",ignore_case=T)), \"Baskonia\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"ZVEZDA\",ignore_case=T)), \"Red Star Belgrade\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, \"Gala\"), \"Galatasaray Istanbul\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"Zal\",ignore_case=T)), \"Zalgiris Kaunas\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"PARTI\",ignore_case=T)), \"Partizan Belgrade\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"KHIMK\",ignore_case=T)), \"Khimki Moscow\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"UNION OLIMP\",ignore_case=T)), \"Union Olimpija\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"OLYMP\",ignore_case=T)), \"Olympiacos Piraeus\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"MONTE\",ignore_case=T)), \"Montepaschi Siena\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"ASVEL\",ignore_case=T)), \"Asvel\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"UNICA\",ignore_case=T)), \"Malaga\", euroleague$TeamA)\neuroleague$TeamA <- ifelse(str_detect(euroleague$TeamA, fixed(\"UNICS\",ignore_case=T)), \"Unics Kazan\", euroleague$TeamA)\n\n\n###\n\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Bamberg\",ignore_case=T)), \"Brose Bamberg\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Bilbao\",ignore_case=T)), \"Bilbao Basket\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Oldenburg\",ignore_case=T)), \"Ewe Baskets Oldenburg\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Valencia\",ignore_case=T)), \"Valencia Basket\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Ulm\",ignore_case=T)), \"Ratiopharm Ulm\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Sevill\",ignore_case=T)), \"Real Betis Baloncesto\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Elan Chalon\",ignore_case=T)), \"Elan Chalon\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Kuban\",ignore_case=T)), \"Lokomotiv Kuban\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Roanne\",ignore_case=T)), \"Roanne\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Zagreb\",ignore_case=T)), \"KK Zagreb\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Olimpija\",ignore_case=T)), \"Union Olimpija\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Gran Canaria\",ignore_case=T)), \"Gran Canaria\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Fb\",ignore_case=T)),\"Fenerbahce Istanbul\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Gescrap\",ignore_case=T)), \"Bilbao Basket\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Spirou\",ignore_case=T)), \"Spirou Charleroi\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Spartak\",ignore_case=T)), \"Spartak Saint Petersburg\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Laboral\",ignore_case=T)), \"Baskonia\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Cantu\",ignore_case=T)), \"Cantu\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Berlin\",ignore_case=T)), \"Alba Berlin\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Artland\",ignore_case=T)), \"Artland Dragons\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"REAL\",ignore_case=T)), \"Real Madrid\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Brose\",ignore_case=T)), \"Brose Bamberg\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Banvit\",ignore_case=T)), \"Bandirma BK\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Kiev\",ignore_case=T)), \"Budivielnik Kiev\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Telenet\",ignore_case=T)), \"Oostende\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Nanterre\",ignore_case=T)), \"Nanterre\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Jerusalem\",ignore_case=T)), \"Hapoel Jerusalem\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Rom\",ignore_case=T)), \"Virtus Rome\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Panionios\",ignore_case=T)), \"Panionios Athens\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Darus\",ignore_case=T)), \"Darussafaka Istanbul\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Cibona\",ignore_case=T)), \"Cibona Zagreb\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Buducnost\",ignore_case=T)), \"Buducnost Podgorica\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Aris\",ignore_case=T)), \"Aris Thessaloniki\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Rytas\",ignore_case=T)), \"Lietuvos Rytas\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Besik\",ignore_case=T)), \"Besiktas Istanbul\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Maccabi\",ignore_case=T)), \"Maccabi Tel Aviv\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Barcelona\",ignore_case=T)), \"FC Barcelona\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Fener\",ignore_case=T)),\"Fenerbahce Istanbul\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Panat\",ignore_case=T)), \"Panathinaikos Athens\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"EFES\",ignore_case=T)), \"Anadolu Efes Istanbul\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"MILAN\",ignore_case=T)), \"Milano\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"VITORIA\",ignore_case=T)), \"Baskonia\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"CSKA\",ignore_case=T)), \"CSKA Moscow\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"TAU\",ignore_case=T)), \"Baskonia\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"ASSECO\",ignore_case=T)), \"Arka Gdynia\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Prokom\",ignore_case=T)), \"Arka Gdynia\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"SLUC\",ignore_case=T)), \"Nancy\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"BASKON\",ignore_case=T)), \"Baskonia\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"ZVEZDA\",ignore_case=T)), \"Red Star Belgrade\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, \"Gala\"), \"Galatasaray Istanbul\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"Zal\",ignore_case=T)), \"Zalgiris Kaunas\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"PARTI\",ignore_case=T)), \"Partizan Belgrade\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"KHIMK\",ignore_case=T)), \"Khimki Moscow\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"UNION OLIMP\",ignore_case=T)), \"Union Olimpija\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"OLYMP\",ignore_case=T)), \"Olympiacos Piraeus\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"MONTE\",ignore_case=T)), \"Montepaschi Siena\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"ASVEL\",ignore_case=T)), \"Asvel\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"UNICA\",ignore_case=T)), \"Malaga\", euroleague$TeamB)\neuroleague$TeamB <- ifelse(str_detect(euroleague$TeamB, fixed(\"UNICS\",ignore_case=T)), \"Unics Kazan\", euroleague$TeamB)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From here I will be creating the required features like the NCAA."},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# Create aggregate stats like the NCAA. \neuro_aggstats <- euroleague %>% group_by(TEAM,PLAYTYPE,Tournament) %>% summarize(counts=n()) %>% filter(PLAYTYPE != \"\") %>% pivot_wider(names_from = \"PLAYTYPE\",values_from = \"counts\") %>% mutate_if(is.numeric, funs(replace_na(., 0))) %>% mutate(TWOPTS_PERC = round((`2FGM` +DUNK + LAYUPMD) / (`2FGM` + `2FGAB` + `2FGA` +DUNK+ LAYUPMD +LAYUPATT ),2),\nTHRPTS_PERC = round(`3FGM` / (`3FGM` + `3FGAB` + `3FGA`),2), FT_PERC = round(FTM / (FTM + FTA),2 )) %>% arrange(FT_PERC)\n\n# Counting opponents rebounds for Four Factors Index.\nopponent_rebounds_euro1 <- euroleague %>% group_by(TeamA,TeamB,TEAM,PLAYTYPE,Tournament) %>% filter((PLAYTYPE == \"O\" |PLAYTYPE == \"D\") & TEAM != TeamA) %>% summarize(count= n()) %>% group_by(TeamA,Tournament) %>% summarize(opp_offreb = sum(count[PLAYTYPE==\"O\"]),opp_defreb = sum(count[PLAYTYPE==\"D\"]))\nopponent_rebounds_euro2 <- euroleague %>% group_by(TeamA,TeamB,TEAM,PLAYTYPE,Tournament) %>% filter((PLAYTYPE == \"O\" |PLAYTYPE == \"D\") & TEAM != TeamB) %>% summarize(count= n()) %>% group_by(TeamA,Tournament) %>% summarize(opp_offreb = sum(count[PLAYTYPE==\"O\"]),opp_defreb = sum(count[PLAYTYPE==\"D\"]))\nopponent_rebounds_euro <- rbind(opponent_rebounds_euro1,opponent_rebounds_euro2)\nopponent_rebounds_euro <- opponent_rebounds_euro %>%group_by(TeamA,Tournament) %>% summarise(opp_offreb= sum(opp_offreb),opp_defreb= sum(opp_defreb))\n\n# Counting opponents rebounds year by year for Four Factors Index. \nopponent_rebounds_euro1 <- euroleague %>% group_by(TeamA,TeamB,TEAM,PLAYTYPE,Tournament,yer) %>% filter((PLAYTYPE == \"O\" |PLAYTYPE == \"D\") & TEAM != TeamA) %>% summarize(count= n()) %>% group_by(TeamA,Tournament,yer) %>% summarize(opp_offreb = sum(count[PLAYTYPE==\"O\"]),opp_defreb = sum(count[PLAYTYPE==\"D\"]))\nopponent_rebounds_euro2 <- euroleague %>% group_by(TeamA,TeamB,TEAM,PLAYTYPE,Tournament,yer) %>% filter((PLAYTYPE == \"O\" |PLAYTYPE == \"D\") & TEAM != TeamB) %>% summarize(count= n()) %>% group_by(TeamB,Tournament,yer) %>% summarize(opp_offreb = sum(count[PLAYTYPE==\"O\"]),opp_defreb = sum(count[PLAYTYPE==\"D\"]))\nopponent_rebounds_euro_yby <- rbind(opponent_rebounds_euro1,opponent_rebounds_euro2)\nopponent_rebounds_euro_yby <- opponent_rebounds_euro_yby %>%group_by(TeamA,Tournament,yer) %>% summarise(opp_offreb= sum(opp_offreb),opp_defreb= sum(opp_defreb))\n\n# Get number of games teams played at home and away and join with the dataframe.\nhome_games <- euroleague %>% filter(PLAYTYPE == \"EG\") %>% group_by(TeamA,Tournament) %>% summarize(HOME_GAMES=n())\naway_games <- euroleague %>% filter(PLAYTYPE == \"EG\") %>% group_by(TeamB,Tournament) %>% summarize(AWAY_GAMES=n())\neuro_aggstats <- euro_aggstats %>% left_join(home_games, by=c(\"TEAM\" = \"TeamA\",\"Tournament\"=\"Tournament\")) %>%  left_join(away_games, by=c(\"TEAM\" = \"TeamB\",\"Tournament\"=\"Tournament\"))\n\n# Join opponent rebounds\neuro_aggstats <- euro_aggstats %>% left_join(opponent_rebounds_euro,by=c(\"TEAM\"=\"TeamA\",\"Tournament\"=\"Tournament\"))\n\n# Create a compact results dataframe, \ncompact_results_euro <- euroleague %>% group_by(TeamA,TeamB,gamenumber,yer,Tournament) %>% summarize(Home_score = max(POINTS_A,na.rm=TRUE),Away_score=max(POINTS_B,na.rm=TRUE)) %>% rename(Home=TeamA) %>% rename(Away=TeamB) %>% rename(year=yer)\n\n# Remove unscrapeable games. There are two games whose play by play details were not avaiable at the json format. I omitted them.\ncompact_results_euro[compact_results_euro == -Inf] <- NA\ncompact_results_euro <- na.omit(compact_results_euro)\n\n# Calculating total scores to reach score per game.\nscore_home <- compact_results_euro %>% group_by(Home,Tournament) %>% summarize(SCORE = sum(Home_score)) %>% rename(Team = Home)\nscore_away <- compact_results_euro %>% group_by(Away,Tournament) %>% summarize(SCORE = sum(Away_score))%>% rename(Team = Away)\nscores <- rbind(score_home,score_away)\nscores <-scores %>% group_by(Team,Tournament) %>% summarize(TOTAL_SCORE = sum(SCORE,na.rm=TRUE))\n\n# Preparing per game stats for Euroleague\n\neuro_aggstats <- euro_aggstats  %>% mutate_if(is.numeric, funs(replace_na(., 0))) %>% left_join(scores,by=c(\"Tournament\"=\"Tournament\",\"TEAM\"=\"Team\")) %>% mutate(TOTAL_GAMES = HOME_GAMES + AWAY_GAMES,SCORE_PG = round(TOTAL_SCORE / TOTAL_GAMES,2), TWOPTS_ATTMPT_PG = ((`2FGM` + `2FGAB` + `2FGA` +DUNK+ LAYUPMD +LAYUPATT )  / TOTAL_GAMES ), TWOPTS_MAKE_PG = ((`2FGM` +DUNK + LAYUPMD)/ TOTAL_GAMES ),\nTHR_MAKE_PG = round(`3FGM` / TOTAL_GAMES,2), THR_ATTMPT_PG= round((`3FGM` + `3FGAB` + `3FGA`) /TOTAL_GAMES,2),  FT_ATTEMPS_PG = round((FTA + FTM) / TOTAL_GAMES,2 ), FT_MAKE_PG = round(FTM / TOTAL_GAMES,2), ASST_PG = round(AS / TOTAL_GAMES,2), DRBND_PG = round (D / TOTAL_GAMES,2), ORBND_PG = round(O / TOTAL_GAMES,2),FOUL_CMMT_PG = round(CM / TOTAL_GAMES,2), FOUL_RCVE_PG = round(RV / TOTAL_GAMES,2),\nTO_PG = round(TO/TOTAL_GAMES,2), ST_PG = round(ST / TOTAL_GAMES,2), DUNK_PG = round(DUNK / TOTAL_GAMES,2), BLCK_PG = round(FV /TOTAL_GAMES,2),BLCKED_PG = round(AG /TOTAL_GAMES,2), UNSPRTFL_PG = round(CMU / TOTAL_GAMES,2),OFOUL_PG = round(OF/TOTAL_GAMES,2),TCHFOUL_PG = round (CMT/TOTAL_GAMES,2), BENCHFOUL_PG =round(B/TOTAL_GAMES,2),COACHFOUL_PG = round(C/TOTAL_GAMES,2),\nFG_ATTEMP_PG = round(TWOPTS_ATTMPT_PG + THR_ATTMPT_PG + FT_ATTEMPS_PG, 2), OPP_OREB_PG =  round(opp_offreb/TOTAL_GAMES,2), OPP_DREB_PG = round(opp_defreb/TOTAL_GAMES,2)\n) %>% filter(!grepl('Media', TEAM) & !is.na(TEAM))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now our aggregate stats for European competitions are ready. Let's have a look!"},{"metadata":{"trusted":true},"cell_type":"code","source":"head(euro_aggstats)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Not all columns are identical and there are features that are not exist in both of the datasets. Hence, I selected common stats here."},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"euro_aggstats <- euro_aggstats %>% select(TEAM,Tournament,TOTAL_GAMES,TWOPTS_PERC,THRPTS_PERC,FT_PERC,SCORE_PG,TWOPTS_ATTMPT_PG,TWOPTS_MAKE_PG,THR_MAKE_PG,THR_ATTMPT_PG,FT_ATTEMPS_PG,FT_MAKE_PG,ASST_PG,DRBND_PG,ORBND_PG,FOUL_CMMT_PG,TO_PG,ST_PG,BLCK_PG,FG_ATTEMP_PG,OPP_DREB_PG,OPP_OREB_PG)\nncaa_aggstats <- ncaa_stats %>% select(TeamName,Season_type,TOTAL_GAMES,TWOPTS_PERC,THRPTS_PERC,FT_PERC,SCORE_PG,TWOPTS_ATTMPT_PG,TWOPTS_MAKE_PG,THR_MAKE_PG,THR_ATTMPT_PG,FT_ATTEMPS_PG,FT_MAKE_PG,ASST_PG,DRBND_PG,ORBND_PG,FOUL_CMMT_PG,TO_PG,ST_PG,BLCK_PG,FG_ATTEMP_PG,OPP_DREB_PG,OPP_OREB_PG) %>% select(-TeamID)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Calculating score margin between teams as a feature.\ncompact_results_euro$AbsMargin <- abs(compact_results_euro$Home_score - compact_results_euro$Away_score)\ncompact_results_ncaa$AbsMargin <- abs(compact_results_ncaa$WScore - compact_results_ncaa$LScore)\navg_margin_euro <- compact_results_euro %>% group_by(year,Tournament) %>% summarize(AVG_MARGIN = mean(AbsMargin),Margin_Q90 = quantile(AbsMargin,0.9), Margin_Q10 = quantile(AbsMargin,0.1), MARGIN_QDIFF = Margin_Q90 - Margin_Q10)\navg_margin_ncaa <- compact_results_ncaa  %>% filter(Season_type != \"Secondary Tournament\") %>% group_by(Season,Season_type) %>% rename(\"year\"=\"Season\") %>% summarize(AVG_MARGIN = mean(AbsMargin),Margin_Q90 = quantile(AbsMargin,0.9), Margin_Q10 = quantile(AbsMargin,0.1), MARGIN_QDIFF = Margin_Q90 - Margin_Q10)\navg_margin_ncaa$Season_type[avg_margin_ncaa$Season_type == \"Regular\" ] <- \"NCAA-Regular\"\navg_margin_ncaa$Season_type[avg_margin_ncaa$Season_type == \"Tournament\" ] <- \"NCAA-Tournament\"\ncolnames(avg_margin_ncaa)[2] <- \"Tournament\"\nmargin <- rbind(avg_margin_euro,avg_margin_ncaa)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 2.2. Data Prep For Year by Year Stats"},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# Data preparation to compare competitiveness year by year\n\n# Data prep for Euro\n\neuro_aggstats_yby <- euroleague %>% group_by(TEAM,PLAYTYPE,Tournament,yer) %>% summarize(counts=n()) %>% filter(PLAYTYPE != \"\") %>% pivot_wider(names_from = \"PLAYTYPE\",values_from = \"counts\") %>% mutate_if(is.numeric, funs(replace_na(., 0))) %>% mutate(TWOPTS_PERC = round((`2FGM` +DUNK + LAYUPMD) / (`2FGM` + `2FGAB` + `2FGA` +DUNK+ LAYUPMD +LAYUPATT ),2),\nTHRPTS_PERC = round(`3FGM` / (`3FGM` + `3FGAB` + `3FGA`),2), FT_PERC = round(FTM / (FTM + FTA),2 )) %>% arrange(FT_PERC)\nscore_home <- compact_results_euro %>% group_by(Home,Tournament,year) %>% summarize(SCORE = sum(Home_score)) %>% rename(Team = Home)\nscore_away <- compact_results_euro %>% group_by(Away,Tournament,year) %>% summarize(SCORE = sum(Away_score))%>% rename(Team = Away)\nscores <- rbind(score_home,score_away)\nscores <-scores %>% group_by(Team,Tournament,year) %>% summarize(TOTAL_SCORE = sum(SCORE,na.rm=TRUE))\n\n# Get number of games teams played at home and away and join with the dataframe.\n\nhome_games <- compact_results_euro %>% group_by(Home,Tournament,year) %>% summarize(HOME_GAMES=n())\naway_games <- compact_results_euro %>% group_by(Away,Tournament,year) %>% summarize(AWAY_GAMES=n())\neuro_aggstats_yby <- euro_aggstats_yby %>% left_join(home_games, by=c(\"TEAM\" = \"Home\",\"Tournament\"=\"Tournament\",\"yer\"=\"year\")) %>%  left_join(away_games, by=c(\"TEAM\" = \"Away\",\"Tournament\"=\"Tournament\",\"yer\"=\"year\")) %>% left_join(opponent_rebounds_euro_yby,by=c(\"Tournament\"=\"Tournament\",\"yer\"=\"yer\",\"TEAM\"=\"TeamA\"))\neuro_aggstats_yby <- euro_aggstats_yby  %>% mutate_if(is.numeric, funs(replace_na(., 0))) %>% left_join(scores,by=c(\"Tournament\"=\"Tournament\",\"TEAM\"=\"Team\",\"yer\"=\"year\")) %>% mutate(TOTAL_GAMES = HOME_GAMES + AWAY_GAMES,SCORE_PG = round(TOTAL_SCORE / TOTAL_GAMES,2), TWOPTS_ATTMPT_PG = ((`2FGM` + `2FGAB` + `2FGA` +DUNK+ LAYUPMD +LAYUPATT )  / TOTAL_GAMES ), TWOPTS_MAKE_PG = ((`2FGM` +DUNK + LAYUPMD)/ TOTAL_GAMES ),\nTHR_MAKE_PG = round(`3FGM` / TOTAL_GAMES,2), THR_ATTMPT_PG= round((`3FGM` + `3FGAB` + `3FGA`) /TOTAL_GAMES,2),  FT_ATTEMPS_PG = round((FTA + FTM) / TOTAL_GAMES,2 ), FT_MAKE_PG = round(FTM / TOTAL_GAMES,2), ASST_PG = round(AS / TOTAL_GAMES,2), DRBND_PG = round (D / TOTAL_GAMES,2), ORBND_PG = round(O / TOTAL_GAMES,2),FOUL_CMMT_PG = round(CM / TOTAL_GAMES,2), FOUL_RCVE_PG = round(RV / TOTAL_GAMES,2),\nTO_PG = round(TO/TOTAL_GAMES,2), ST_PG = round(ST / TOTAL_GAMES,2), DUNK_PG = round(DUNK / TOTAL_GAMES,2), BLCK_PG = round(FV /TOTAL_GAMES,2),BLCKED_PG = round(AG /TOTAL_GAMES,2), UNSPRTFL_PG = round(CMU / TOTAL_GAMES,2),OFOUL_PG = round(OF/TOTAL_GAMES,2),TCHFOUL_PG = round (CMT/TOTAL_GAMES,2), BENCHFOUL_PG =round(B/TOTAL_GAMES,2),COACHFOUL_PG = round(C/TOTAL_GAMES,2),\nFG_ATTEMP_PG = round(TWOPTS_ATTMPT_PG + THR_ATTMPT_PG + FT_ATTEMPS_PG, 2), OPP_OREB_PG =  round(opp_offreb/TOTAL_GAMES,2), OPP_DREB_PG = round(opp_defreb/TOTAL_GAMES,2)  \n) %>% filter(!grepl('Media', TEAM) & !is.na(TEAM))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"## Year by year preparation for the NCAA\nncaa_stats_yby <- rbind(win_stats,lose_stats)\nncaa_stats_yby <- ncaa_stats_yby %>% left_join(opp_rebounds_ncaa_yby,by=c(\"Season_type\"=\"Season_type\",\"TeamID\"=\"TeamID\",\"Season\"=\"Season\")) \nncaa_totalgames <- ncaa_stats_yby %>% group_by(TeamID,Season,Season_type) %>%   summarize(TOTAL_GAMES=n()) %>% arrange(desc(TOTAL_GAMES))\nncaa_stats_yby <- ncaa_stats_yby %>% group_by(TeamID,Season_type,Season) %>% summarise_all(funs(sum)) %>% mutate(TWOPTS_PERC = round(FGM / (FGM+ FGA),2),THRPTS_PERC = round(FGM3 / (FGM3 + FGA3),2), FT_PERC = round(FTM /  (FTM + FTA),2)) %>% left_join(ncaa_totalgames, by = c(\"TeamID\"=\"TeamID\",\"Season\"=\"Season\",\"Season_type\"=\"Season_type\")) %>% \nmutate(SCORE_PG = round(Score / TOTAL_GAMES,2),TWOPTS_ATTMPT_PG = round((FGA - FGA3) / TOTAL_GAMES,2), TWOPTS_MAKE_PG= round((FGM - FGM3) / TOTAL_GAMES,2),THR_MAKE_PG = round(FGM3 / TOTAL_GAMES,2),\n       THR_ATTMPT_PG= round(FGA3 /TOTAL_GAMES,2),FT_ATTEMPS_PG = round(FTA / TOTAL_GAMES,2 ), FT_MAKE_PG = round(FTM / TOTAL_GAMES,2), ASST_PG = round(Ast / TOTAL_GAMES,2),\n      DRBND_PG = round (DR / TOTAL_GAMES,2), ORBND_PG = round(OR / TOTAL_GAMES,2),FOUL_CMMT_PG = round(PF / TOTAL_GAMES,2),TO_PG = round(TO/TOTAL_GAMES,2), \n       ST_PG = round(Stl / TOTAL_GAMES,2), FG_ATTEMP_PG = round(FT_ATTEMPS_PG +TWOPTS_ATTMPT_PG + THR_ATTMPT_PG, 3), BLCK_PG = round(Blk /TOTAL_GAMES,2),\n       OPP_OREB_PG =  round(opp_offreb/TOTAL_GAMES,2), OPP_DREB_PG = round(opp_defreb/TOTAL_GAMES,2)) %>% left_join(teamnames, by=c(\"TeamID\"=\"TeamID\")) %>% select(-FirstD1Season,-LastD1Season)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Not all columns are identical and there are features that are not exist in both of the datasets. Hence, I selected common stats here."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"## Binding dataframes\n\neuro_aggstats_yby <- euro_aggstats_yby %>% select(TEAM,yer,Tournament,TOTAL_GAMES,TWOPTS_PERC,THRPTS_PERC,FT_PERC,SCORE_PG,TWOPTS_ATTMPT_PG,TWOPTS_MAKE_PG,THR_MAKE_PG,THR_ATTMPT_PG,FT_ATTEMPS_PG,FT_MAKE_PG,ASST_PG,DRBND_PG,ORBND_PG,FOUL_CMMT_PG,TO_PG,ST_PG,BLCK_PG,FG_ATTEMP_PG,OPP_DREB_PG,OPP_OREB_PG)\nncaa_stats_yby <- ncaa_stats_yby %>% select(TeamName,Season,Season_type,TOTAL_GAMES,TWOPTS_PERC,THRPTS_PERC,FT_PERC,SCORE_PG,TWOPTS_ATTMPT_PG,TWOPTS_MAKE_PG,THR_MAKE_PG,THR_ATTMPT_PG,FT_ATTEMPS_PG,FT_MAKE_PG,ASST_PG,DRBND_PG,ORBND_PG,FOUL_CMMT_PG,TO_PG,ST_PG,BLCK_PG,FG_ATTEMP_PG,OPP_DREB_PG,OPP_OREB_PG) %>% select(-TeamID)\n# Last touches then adding them together.\neuro_aggstats_yby <- as_tibble(euro_aggstats_yby) %>% rename(Year = \"yer\")\nncaa_stats_yby <- ncaa_stats_yby %>% select(-TeamID) %>% rename(Tournament=\"Season_type\") %>% rename(TEAM = \"TeamName\") %>% rename(Year = \"Season\") %>% mutate(League=\"NCAA\")\neuro_aggstats_yby <- euro_aggstats_yby %>% mutate(League=\"Euroleague\")\nncaa_stats_yby <- ncaa_stats_yby %>% ungroup() %>%\n  select(-TeamID)\nagg_df_yby <- rbind(euro_aggstats_yby,ncaa_stats_yby)\nagg_df_yby$Tournament[agg_df_yby$Tournament == \"Regular\" ] <- \"NCAA-Regular\"\nagg_df_yby$Tournament[agg_df_yby$Tournament == \"Tournament\" ] <- \"NCAA-Tournament\"\nhead(agg_df_yby)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Last touches then adding them together.\neuro_aggstats <- as_tibble(euro_aggstats)\nncaa_aggstats <- ncaa_aggstats %>% select(-TeamID) %>% rename(Tournament=\"Season_type\") %>% rename(TEAM = \"TeamName\") %>% mutate(League=\"NCAA\")\neuro_aggstats <- euro_aggstats %>% mutate(League=\"Euroleague\")\nncaa_aggstats <- ncaa_aggstats %>% ungroup() %>%\n  select(-TeamID)\nagg_df <- rbind(euro_aggstats,ncaa_aggstats)\nagg_df$Tournament[agg_df$Tournament == \"Regular\" ] <- \"NCAA-Regular\"\nagg_df$Tournament[agg_df$Tournament == \"Tournament\" ] <- \"NCAA-Tournament\"\nhead(agg_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 3. Game-On:Exploratory Analysis and Competitiveness of the Leagues\n\nSince our data is ready, now I will briefly make an exploratory analysis for the features I created with scatterplots. After that I will analyze competitiveness within the NCAA and the Euroleague for aggregate and year by year stats. To do this I selected two different approaches that compares team strengths and will compare them with each other. Here is the brief description of those approaches.\n\n**1- Simple Averaging:** Here I will compare the leagues within themselves to see how top and bottom teams differ in key statistics. Note that, I am not comparing two leagues with each other since this would be unfair to the future stars of the basketball at NCAA. However teams are still comparable in terms of competitiveness within the league. As a metric, I choose each teams average statistics per game since number of games they play each year differently. \n\n**2- Four Factor Approach: **[Dean Oliver](http://www.basketballonpaper.com)  developed this approach in order to quantify how a team win the games in 2004. He focused on shooting (40%), turnovers(25%), rebounds (20%) and free throws (15%) as key metrics with varying weights. I will be looking for average four factors between teams in each competition and then again look for difference between 90th and 10th quartiles.\n\nIn order to clear up impact of outliers, I looked into difference between 10th and 90th quartiles for all approaches and checked the differences."},{"metadata":{},"cell_type":"markdown","source":"## 3.1. Exploratory Analysis\n\nBelow I will be using similar scatterplots to compare key features. Before starting, there are a few things to note:\n\n* Size of the dots represent number of games team is played. \n* Color of the dots represent specific team.\n* All values are per game statistics for teams. \n\n### 3.1.1. Overall Exploration\n\nI start my work with plots on success percentage per game and number of attemps for two-pointers, three-pointers and free throws.\n\nHere we go with the two-pointers. Below the plots above shows the European competitions and below the different parts of the season at the NCAA.  \n\n* During regular season, teams succeed at similar rates (0.27 to 0.32) in large number of games (many teams play 200+ games since 2012 at NCAA regular season) unlike tournament games or Euroleague games. This sounds like law of large numbers here due to number of games played. As number of games increases, stats tend to normalize. \n\n* The gap between teams at NCAA regular season much lower than the tournament and the European competitions.\n\n* I have already states that it would be unfair to compare future stars with professional players in Europe, however even top NCAA teams can not get closer to the success percentages of European teams at the bottom. \n\n"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# height and width of plots (feel free to adjust, if plots are too small/big on your screen)\nHEIGHT = 750\nWIDTH = 1300\nplt1 = ggplot(agg_df, aes(x =TWOPTS_PERC, y = TWOPTS_ATTMPT_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Two Points Success Percentage: ', TWOPTS_PERC,\n                                              '<br>Two Points Attempt Per Game', TWOPTS_ATTMPT_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ .)+\n    theme_bw()+\n    scale_x_continuous(limits = c(0.2,0.6),\n                     labels = scales::percent_format(accuracy = 5L))+\n    xlab(\"Success Percentage of Two Point Shoots\")+\n    ylab(\"Two Point Shoots Per Game\")+\n    labs(title= \"Two Point Shoots Success vs. Two Point Shoots Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), panel.grid.minor = element_blank(),strip.text = element_text(size=16,face=\"bold\", colour = \"orange\"),strip.background =element_rect(fill=\"black\"),\n         axis.title.y = element_text(face='bold', colour='saddlebrown', size=14), plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),plot.margin = unit(c(2, 0, 0, 0), \"cm\"),   \n        plot.subtitle = element_text(hjust = 0.5, vjust=1,size=18),axis.title.x = element_text(face='bold', colour='saddlebrown', size=14),          \n         axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\npltly = ggplotly(plt1, tooltip = c('text')) %>%   layout(title = list(text = paste0('Two Point Shoots Success vs. Two Point Shoots Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot1.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot1.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH/1.2)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":""},{"metadata":{},"cell_type":"markdown","source":"* For three-pointers some teams at the top of their performance reached average Euroleague and Eurocup teams during the Tournament but number of games those played are really low, so it is not convincing to generalize any insights on this. However, teams performance between NCAA and European teams are slightly closer to each other than two-pointers. \n\n* Top 3-pointers of the NCAA regular season are also close to the teams at the bottom of Eurocup performance. Not bad for college teams!\n\n* Number of attempts per game are similar for different tournaments."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt2 = ggplot(agg_df, aes(x =THRPTS_PERC, y = THR_ATTMPT_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Three Points Success Percentage: ', THRPTS_PERC,\n                                              '<br>Three Points Attempt Per Game', THR_ATTMPT_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ .)+\n    theme_bw()+\n    scale_x_continuous(limits = c(0.2,0.6),\n                     labels = scales::percent_format(accuracy = 5L))+\n    xlab(\"Success Percentage of Three Point Shoots\")+\n    ylab(\"Three Point Shoots Per Game\")+\n    labs(title= \"Three Point Shoots Success vs. Three Point Shoots Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=16,face=\"bold\", colour = \"orange\"),\n          strip.background =element_rect(fill=\"black\"),\n         axis.title.y = element_text(face='bold', colour='saddlebrown', size=14,margin = margin(t = 0, r = 10, b = 0, l = 0)), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 0, 0, 0), \"cm\"),   \n        plot.subtitle = element_text(hjust = 0.5, vjust=1,size=18),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=14),\n         axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\n\n\npltly = ggplotly(plt2, tooltip = c('text')) %>%   layout(title = list(text = paste0('Three Point Shoots Success vs. Three Point Shoots Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot2.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot2.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH/1.2)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Free throws are super easy for some players while a nightmare for others. What are factors that affect success in free throws? Practice (or experience) and mental preparedness are the two most important aspect I think. Below we see the sharpest difference between European competitions and the NCAA. Like in other metrics, NCAA teams at regular season are clustered together."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt3 = ggplot(agg_df, aes(x =FT_PERC, y = FT_ATTEMPS_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Free Throw Success Percentage: ', FT_PERC,\n                                              '<br>Free Throw Attempt Per Game', FT_ATTEMPS_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ .)+\n    theme_bw()+\n    scale_x_continuous(limits = c(0.2,1),\n                     labels = scales::percent_format(accuracy = 5L))+\n    xlab(\"Success Percentage of Free Throws\")+\n    ylab(\"Free Throw Shoots Per Game\")+\n    labs(title= \"Free Throw Success vs. Free Throw Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=16,face=\"bold\", colour = \"orange\"),\n          strip.background =element_rect(fill=\"black\"),\n         axis.title.y = element_text(face='bold', colour='saddlebrown', size=14,margin = margin(t = 0, r = 10, b = 0, l = 0)), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 0, 0, 0), \"cm\"),   \n        plot.subtitle = element_text(hjust = 0.5, vjust=1,size=18),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=14),\n         axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\n\n\npltly = ggplotly(plt3, tooltip = c('text')) %>%   layout(title = list(text = paste0('Free Throw Success vs. Free Throw Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot3.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot3.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH/1.2)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For score per game and number of field goal attempts, there is a clear linear relationship and similar distribution between tournaments. Teams at the bottom averaged around 60 points per game while top performers score around 80. Tournament part of the NCAA has some outliers at the bottom."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt4 = ggplot(agg_df, aes(x =SCORE_PG, y = FG_ATTEMP_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Score Per Game: ', SCORE_PG,\n                                              '<br>Field Goal Attempts Per Game', FG_ATTEMP_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ .)+\n    theme_bw()+\n    #scale_x_continuous(limits = c(0.2,1),\n                     #labels = scales::percent_format(accuracy = 5L))+\n    xlab(\"Score Per Game\")+\n    ylab(\"Field Goal Attempts Per Game\")+\n    labs(title= \"Score Per Game vs. Field Goal Attempt Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=16,face=\"bold\", colour = \"orange\"),\n          strip.background =element_rect(fill=\"black\"),\n         axis.title.y = element_text(face='bold', colour='saddlebrown', size=14,margin = margin(t = 0, r = 10, b = 0, l = 0)), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 0, 0, 0), \"cm\"),   \n        plot.subtitle = element_text(hjust = 0.5, vjust=1,size=18),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=14),\n         axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\n\n\npltly = ggplotly(plt4, tooltip = c('text')) %>%   layout(title = list(text = paste0('Score Per Game vs. Field Goal Attempt Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot4.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot4.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH/1.2)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":""},{"metadata":{},"cell_type":"markdown","source":"There is a slight positive relationship with assists per game and score per game. What is more striking here is, when teams make more assists they tend to advance into later stages into tournaments hence play more game."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt5 = ggplot(agg_df, aes(x =ASST_PG, y = SCORE_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Assist Per Game: ', ASST_PG,\n                                              '<br>Score Per Game', SCORE_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ .)+\n    theme_bw()+\n    #scale_x_continuous(limits = c(0.2,1),\n                     #labels = scales::percent_format(accuracy = 5L))+\n    xlab(\"Assist Per Game\")+\n    ylab(\"Score Per Game\")+\n    labs(title= \"Assist Per Game vs. Score Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=16,face=\"bold\", colour = \"orange\"),\n          strip.background =element_rect(fill=\"black\"),\n         axis.title.y = element_text(face='bold', colour='saddlebrown', size=14,margin = margin(t = 0, r = 10, b = 0, l = 0)), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 0, 0, 0), \"cm\"),   \n        plot.subtitle = element_text(hjust = 0.5, vjust=1,size=18),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=14),\n         axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\n\n\npltly = ggplotly(plt5, tooltip = c('text')) %>%   layout(title = list(text = paste0('Assist Per Game vs. Score Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot5.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot5.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH/1.2)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Fouls and blocks are about getting physical so I plotted them with the intuition to have some relationship. However, it looks like there is not.  When I compared number of fouls committed per game in different tournaments, they distributed similar except the tournament section of the NCAA where there are lots of outliers."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt6 = ggplot(agg_df, aes(x =FOUL_CMMT_PG, y = BLCK_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Foul Committed Per Game: ', FOUL_CMMT_PG,\n                                              '<br>Block Per Game', BLCK_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ .)+\n    theme_bw()+\n    #scale_x_continuous(limits = c(0.2,1),\n                     #labels = scales::percent_format(accuracy = 5L))+\n    xlab(\"Foul Committed Per Game\")+\n    ylab(\"Blocks Per Game\")+\n    labs(title= \"Foul Commited Per Game vs. Block Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=16,face=\"bold\", colour = \"orange\"),\n          strip.background =element_rect(fill=\"black\"),\n         axis.title.y = element_text(face='bold', colour='saddlebrown', size=14,margin = margin(t = 0, r = 10, b = 0, l = 0)), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 0, 0, 0), \"cm\"),   \n        plot.subtitle = element_text(hjust = 0.5, vjust=1,size=18),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=14),\n         axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\n\n\npltly = ggplotly(plt6, tooltip = c('text')) %>%   layout(title = list(text = paste0('Foul Committed Per Game vs. Block Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot6.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot6.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH/1.2)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 3.1.2. Year by Year Exploration\n\nAt this section, I will be looking for the changes in tournaments key features year by year before moving on to calculate competitiveness. Hence we can see the change in the game over the years and tournaments which can leads to us smart questions to consider.\n\nAt first look, I see that two points attempts at NCAA regular season slightly decreased from 2012 towards today. Particularly at the first years of my analysis, there are more two pointers attempt at NCAA than European competitions.\n\nNote that since NCAA tournament is played in knock-out format, most of the teams playing only 1 or 2 games, hence we should look for the team stats with a pinch of salt for the tournament.\n\n"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"HEIGHT = 1400\nWIDTH = 1600\nplt7 = ggplot(agg_df_yby, aes(x =TWOPTS_PERC, y = TWOPTS_ATTMPT_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Two Points Success Percentage: ', TWOPTS_PERC,\n                                              '<br>Two Points Attempt Per Game', TWOPTS_ATTMPT_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ Year,ncol=7,labeller = label_wrap_gen())+\n    theme_bw()+\n    scale_x_continuous(limits = c(0.2,0.7),breaks=c(0.2,0.35,0.50,0.65),\n                     labels = scales::percent_format(accuracy = 5L))+\n    xlab(\"Success Percentage of Two Point Shoots\")+\n    ylab(\"Two Point Shoots Per Game\")+\n    labs(title= \"Two Point Shoots Success vs. Two Point Shoots Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=12,face=\"bold\", colour = \"black\", hjust = -1),\n          strip.background =element_rect(fill=\"white\",color=\"white\",size=4),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=12), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          plot.subtitle = element_text(hjust = 0.5, vjust=1,size=16),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=12),          \n          axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\npltly = ggplotly(plt7, tooltip = c('text')) %>%   layout(title = list(text = paste0('Two Point Shoots Success vs. Two Point Shoots Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\nhtmlwidgets::saveWidget(pltly, \"plot7.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot7.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"[There is an increasing number of 3-pointers at basketball particularly at NBA over the years.](https://shottracker.com/articles/the-3-point-revolution) which is a widely discussed topic in basketball analytics. We see that this is also true for the NCAA and the Eurocup, however number of 3-pointers attempted does not change much for the Euroleague. Moreover we see smaller variation at European competition in terms of performance comparing to the NCAA."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# height and width of plots (feel free to adjust, if plots are too small/big on your screen)\nplt8 = ggplot(agg_df_yby, aes(x =THRPTS_PERC, y = THR_ATTMPT_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Three Points Success Percentage: ', THRPTS_PERC,\n                                              '<br>Three Points Attempt Per Game', THR_ATTMPT_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ Year,ncol=7)+\n    theme_bw()+\n    scale_x_continuous(limits = c(0.2,0.6),breaks=c(0.2,0.35,0.5),\n                     labels = scales::percent_format(accuracy = 5L))+\n    scale_y_continuous(limits = c(0,50),breaks=c(10,20,30,40,50))+\n    xlab(\"Success Percentage of Three Point Shoots\")+\n    ylab(\"Three Point Shoots Per Game\")+\n    labs(title= \"Three Point Shoots Success vs. Three Point Shoots Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=12,face=\"bold\", colour = \"black\", hjust = -1),\n          strip.background =element_rect(fill=\"white\",color=\"white\",size=4),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=12), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          plot.subtitle = element_text(hjust = 0.5, vjust=1,size=16),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=12),          \n          axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\npltly = ggplotly(plt8, tooltip = c('text')) %>%   layout(title = list(text = paste0('Three Point Shoots Success vs. Three Point Shoots Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot8.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot8.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# height and width of plots (feel free to adjust, if plots are too small/big on your screen)\nplt9 = ggplot(agg_df_yby, aes(x =FT_PERC, y = FT_ATTEMPS_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Free Throw Success Percentage: ', FT_PERC,\n                                              '<br>Free Throw Attempt Per Game', FT_ATTEMPS_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ Year,ncol=7)+\n    theme_bw()+\n    xlab(\"Success Percentage of Free Throws\")+\n    ylab(\"Free Throws Per Game\")+\n    scale_x_continuous(limits = c(0.2,0.9),breaks=c(0.2,0.4,0.6,0.8),\n                     labels = scales::percent_format(accuracy = 5L))+\n    labs(title= \"Free Throw Success vs. Free Throw Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=12,face=\"bold\", colour = \"black\", hjust = -1),\n          strip.background =element_rect(fill=\"white\",color=\"white\",size=4),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=12), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          plot.subtitle = element_text(hjust = 0.5, vjust=1,size=16),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=12),          \n          axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\npltly = ggplotly(plt9, tooltip = c('text')) %>%   layout(title = list(text = paste0('Free Throw Success vs. Free Throws Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot9.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot9.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"On score per game and field goal attemps per game, we see that variation at the Euroleague is diminishing, hence team scores are getting closer to each other, which can be a sign of competitiveness. At NCAA, variation is bigger between teams for score per game. Unlike in other tournaments, number of field goal attemps are the closest at Euroleague. Moreover, score per games are increasing at all tournaments. Game is getting more efficient or is it getting faster?"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt10 = ggplot(agg_df_yby, aes(x =SCORE_PG, y = FG_ATTEMP_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Score Per Game: ', SCORE_PG,\n                                              '<br>Field Goal Attempt Per Game', FG_ATTEMP_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ Year,ncol=7)+\n    theme_bw()+\n    scale_y_continuous(limits = c(50,110),breaks=c(55,75,95))+\n    scale_x_continuous(limits = c(40,100),breaks=c(55,75,95))+\n    xlab(\"Score Per Game\")+\n    ylab(\"Field Goal Attemps Per Game\")+\n    labs(title= \"Free Throw Success vs. Free Throw Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=12,face=\"bold\", colour = \"black\", hjust = -1),\n          strip.background =element_rect(fill=\"white\",color=\"white\",size=4),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=12), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          plot.subtitle = element_text(hjust = 0.5, vjust=1,size=16),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=12),          \n          axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\npltly = ggplotly(plt10, tooltip = c('text')) %>%   layout(title = list(text = paste0('Score Per Game vs. Field Goal Attemps Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot10.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot10.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"With scores assists per games are increasing as well. Are they increasing because of the increase in scores or teams are just playing more collaboratively? This could be another research question!"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt11 = ggplot(agg_df_yby, aes(x =SCORE_PG, y = ASST_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Score Per Game:', SCORE_PG,\n                                              '<br>Assist Per Game:', FG_ATTEMP_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ Year,ncol=7)+\n    theme_bw()+\n    xlab(\"Score Per Game\")+\n    ylab(\"Assist Per Game\")+\n    labs(title= \"Assist Per Game vs. Score Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=12,face=\"bold\", colour = \"black\", hjust = -1),\n          strip.background =element_rect(fill=\"white\",color=\"white\",size=4),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=12), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          plot.subtitle = element_text(hjust = 0.5, vjust=1,size=16),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=12),          \n          axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\npltly = ggplotly(plt11, tooltip = c('text')) %>%   layout(title = list(text = paste0('Assist Per Game vs. Score Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot11.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot11.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"On fouls and blocks teams are getting closer to each other at all competitions with some decrease in blocks per games. "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt12 = ggplot(agg_df_yby, aes(x =FOUL_CMMT_PG, y = BLCK_PG,color=TEAM,size=TOTAL_GAMES,text = paste('Foul Committed Per Game:', FOUL_CMMT_PG,\n                                              '<br>Block Per Game:', BLCK_PG, \n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_point(alpha=0.6)+\n    facet_wrap(Tournament ~ Year,ncol=7)+\n    theme_bw()+\n    xlab(\"Foul Committed Per Game\")+\n    ylab(\"Block Per Game\")+\n    labs(title= \"Foul Commited Per Game vs. Block Per Game\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=12,face=\"bold\", colour = \"black\", hjust = -1),\n          strip.background =element_rect(fill=\"white\",color=\"white\",size=4),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=12), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          plot.subtitle = element_text(hjust = 0.5, vjust=1,size=16),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=12),          \n          axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=12))\n\npltly = ggplotly(plt12, tooltip = c('text')) %>%   layout(title = list(text = paste0('Foul Committed Per Game vs. Block Per Game',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot12.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot12.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 3.2. Competitiveness\n\nWe are done with the exploratory analysis for now, we can extend it much more however our main is quantifying competitiveness. First thing is first, our definition of comeptitiveness is having teams in a given league close to each other at key metrics using our Simple Averaging Method and Four Factors Method. We will calculate key statistics for each team aggreagately and year by year. \n\nNow we will look to the competitiveness between the leagues. While the plots above give some clues, trying to quantify them would be necessary. Below code calculates difference for the key statistics between 10th and 90th percentile for the metrics below. Before diving into the differences, I will also calculate winning margin between teams.\n\n### 3.2.1. Simple Averaging Method\n#### 3.2.1.1 Aggregate Stats\n\nAs stated above, I simply took average of the featured I created and look for difference between them at 10th and 90th quartiles."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Calculating score margin between teams as a feature.\ncompact_results_euro$AbsMargin <- abs(compact_results_euro$Home_score - compact_results_euro$Away_score)\ncompact_results_ncaa$AbsMargin <- abs(compact_results_ncaa$WScore - compact_results_ncaa$LScore)\navg_margin_euro <- compact_results_euro %>% group_by(Tournament) %>% summarize(AVG_MARGIN = mean(AbsMargin),Margin_Q90 = quantile(AbsMargin,0.9), Margin_Q10 = quantile(AbsMargin,0.1), MARGIN_QDIFF = Margin_Q90 - Margin_Q10)\navg_margin_ncaa <- compact_results_ncaa %>% filter(Season_type != \"Secondary Tournament\") %>% group_by(Season_type) %>% summarize(AVG_MARGIN = mean(AbsMargin),Margin_Q90 = quantile(AbsMargin,0.9), Margin_Q10 = quantile(AbsMargin,0.1), MARGIN_QDIFF = Margin_Q90 - Margin_Q10)\navg_margin_ncaa$Season_type[avg_margin_ncaa$Season_type == \"Regular\" ] <- \"NCAA-Regular\"\navg_margin_ncaa$Season_type[avg_margin_ncaa$Season_type == \"Tournament\" ] <- \"NCAA-Tournament\"\ncolnames(avg_margin_ncaa)[1] <- \"Tournament\"\nmargin <- rbind(avg_margin_ncaa,avg_margin_euro)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-output":true,"_kg_hide-input":true},"cell_type":"code","source":"# Percentage of tight games. \ncompact_results_ncaa %>% mutate(TIGHT_GAME = ifelse(AbsMargin <= 5,1,0)) %>% filter(Season_type != \"Secondary Tournament\") %>% group_by(Season_type,TIGHT_GAME) %>% summarize(count=n()) %>% mutate(freq= round(count / sum(count),2))\ncompact_results_euro %>% mutate(TIGHT_GAME = ifelse(AbsMargin <= 5,1,0)) %>% group_by(Tournament,TIGHT_GAME) %>% summarize(count=n()) %>% mutate(freq= round(count / sum(count),2))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# For the calculation of quartiles inspired by quick solution here. https://tbradley1013.github.io/2018/10/01/calculating-quantiles-for-groups-with-dplyr-summarize-and-purrr-partial/\n\np <- c(0.1, 0.5, 0.9,1)\np_names <- map_chr(p, ~paste0(.x*100, \"%\"))\n\np_funs <- map(p, ~partial(quantile, probs = .x, na.rm = TRUE)) %>% \n  set_names(nm = p_names)\n\nquartile_diff <- agg_df  %>%\n  group_by(Tournament) %>% left_join(margin,by=c(\"Tournament\"=\"Tournament\")) %>%\n  summarize_at(vars(TWOPTS_PERC,THRPTS_PERC,FT_PERC,ASST_PG,TO_PG,ST_PG,BLCK_PG,DRBND_PG,ORBND_PG), funs(!!!p_funs)) %>% mutate(TWOPTS_QDIFF = `TWOPTS_PERC_90%` - `TWOPTS_PERC_10%`,\n                                                                                                                                THRPTS_QDIFF = `THRPTS_PERC_90%` - `THRPTS_PERC_10%`, \n                                                                                                                                FTPTS_QDIFF = `FT_PERC_90%` - `FT_PERC_10%`,\n                                                                                                                                ASST_QDIFF = `ASST_PG_90%` - `ASST_PG_10%`, \n                                                                                                                                TO_QDIFF = `TO_PG_90%` - `TO_PG_10%`, \n                                                                                                                                ST_QDIFF = `ST_PG_90%` - `ST_PG_10%`,\n                                                                                                                               BLCK_QDIFF = `BLCK_PG_90%` - `BLCK_PG_10%`,\n                                                                                                                                DRBND_QDIFF = `DRBND_PG_90%` - `DRBND_PG_10%`,\n                                                                                                                                ORBND_QDIFF = `ORBND_PG_90%` - `ORBND_PG_10%`) %>%  select(Tournament,matches('QDIFF'))\nquartile_diff <- quartile_diff %>% left_join(margin,by=c(\"Tournament\"=\"Tournament\")) %>% select(-Margin_Q90,-Margin_Q10,-AVG_MARGIN)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Below you can see my results for the first method. I selected 10 features and looked for difference between teams at 10th and 90th quartiles in order to neutralize outliers affects:\n\n- 90th and 10th quartiles difference  in two points success percentage difference.\n- 90th and 10th quartiles difference  in three points success percantage difference.\n- 90th and 10th quartiles difference  in free throw success percentage difference.\n- 90th and 10th quartiles difference in number of assists per game.\n- 90th and 10th quartiles difference in number of turnovers per game.\n- 90th and 10th quartiles difference in number of steals per game.\n- 90th and 10th quartiles difference in number of blocks per game.\n- 90th and 10th quartiles difference in number of defensive rebounds per game.\n- 90th and 10th quartiles difference in number of offensive rebounds per game.\n- 90th and 10th quartiles difference in winning margin between teams per game.\n\nMetrics we have built above are not in the range. So in order to be able to compare them I will normalize them between 0 and 1. After that I summed them up and reached our **Competitiveness Index with Simple Averaging**. Here is my insights on competitiveness of the leagues.\n\n- Most competitive tournament is the NCAA-Regular season in overall while NCAA-Tournament is the one with most prone to surprises. On almost all aspects, difference in key metrics are bigger than other teams.\n- European competitions are similar to each other in terms of competitiveness however not as competitive as NCAA. The main reason could be change in teams strength at the NCAA since students come and go which varies team strengths. On the other hand, at Euroleague certain powerhouses mostly advances through next rounds. For example, number of different teams that reach to the Final Four at the last 5 years only 8. \n- Average margin is highly similar at all competitions with slighly more margin at the NCAA regular season. \n- Between each other, Euroleague is much more professional with higher success percentage in shoots, however it looks like NCAA is more open to surprises and Cindrella stories.\n- I created a new metric, TIGHT_GAMES, games that are finished with less than 5 points. In all tournaments around 28-30% of the games finished tightly, so two continents are similar on this matter too. Hence I did not even add it as a metric to my calculation.\n- Lastly note that number of games which we based on our analysis are 70K for NCAA regular season, 800 for tournament, 1700 for the Eurocup and 2983 for the Euroleague. \n\nNote that above I derived another metric from margins, **TIGHT_GAMES**, *games that are finished with less than 5 points*. In all tournaments around 28-30% of the games finished tightly, so two continents are similar on this matter too.\n"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"range01 <- function(x){(x-min(x))/(max(x)-min(x))}\nnon_normalized_results <- quartile_diff %>% mutate(comp  = rowSums(.[2:11]))\nnormalized <- quartile_diff %>% mutate_at(c(\"TWOPTS_QDIFF\",\"THRPTS_QDIFF\",\"FTPTS_QDIFF\",\"ASST_QDIFF\",\"TO_QDIFF\",\"ST_QDIFF\",\"BLCK_QDIFF\",\"DRBND_QDIFF\",\"ORBND_QDIFF\",\"MARGIN_QDIFF\"), ~(range01(.) %>% as.vector))%>% mutate(Comp_Index  = round(rowSums(.[2:11]),2))\nnormalized","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 2.2.1.2 Year by Year\n\nNow we will do the same analysis for each year separately to see if the competitiveness varies between leagues year by year. First we created features for the differences between quartiles."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"avg_margin_euro <- compact_results_euro %>% group_by(year,Tournament) %>% summarize(AVG_MARGIN = mean(AbsMargin),Margin_Q90 = quantile(AbsMargin,0.9), Margin_Q10 = quantile(AbsMargin,0.1), MARGIN_QDIFF = Margin_Q90 - Margin_Q10)\navg_margin_ncaa <- compact_results_ncaa %>% filter(Season_type != \"Secondary Tournament\") %>% group_by(Season,Season_type) %>% summarize(AVG_MARGIN = mean(AbsMargin),Margin_Q90 = quantile(AbsMargin,0.9), Margin_Q10 = quantile(AbsMargin,0.1), MARGIN_QDIFF = Margin_Q90 - Margin_Q10)\navg_margin_ncaa$Season_type[avg_margin_ncaa$Season_type == \"Regular\" ] <- \"NCAA-Regular\"\navg_margin_ncaa$Season_type[avg_margin_ncaa$Season_type == \"Tournament\" ] <- \"NCAA-Tournament\"\ncolnames(avg_margin_ncaa)[2] <- \"Tournament\"\ncolnames(avg_margin_euro)[1] <- \"Season\"\nmargin <- rbind(avg_margin_ncaa,avg_margin_euro)\nhead(margin)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"quartile_diff <- agg_df_yby  %>%\n  group_by(Year,Tournament) %>%\n  summarize_at(vars(TWOPTS_PERC,THRPTS_PERC,FT_PERC,ASST_PG,TO_PG,ST_PG,BLCK_PG,DRBND_PG,ORBND_PG), funs(!!!p_funs)) %>% mutate(TWOPTS_QDIFF = `TWOPTS_PERC_90%` - `TWOPTS_PERC_10%`,\n                                                                                                                                THRPTS_QDIFF = `THRPTS_PERC_90%` - `THRPTS_PERC_10%`, \n                                                                                                                                FTPTS_QDIFF = `FT_PERC_90%` - `FT_PERC_10%`,\n                                                                                                                                ASST_QDIFF = `ASST_PG_90%` - `ASST_PG_10%`, \n                                                                                                                                TO_QDIFF = `TO_PG_90%` - `TO_PG_10%`, \n                                                                                                                                ST_QDIFF = `ST_PG_90%` - `ST_PG_10%`,\n                                                                                                                               BLCK_QDIFF = `BLCK_PG_90%` - `BLCK_PG_10%`,\n                                                                                                                                DRBND_QDIFF = `DRBND_PG_90%` - `DRBND_PG_10%`,\n                                                                                                                                ORBND_QDIFF = `ORBND_PG_90%` - `ORBND_PG_10%`) %>%  select(Tournament,Year,matches('QDIFF'))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"quartile_diff <- quartile_diff %>% left_join(margin,by=c(\"Tournament\"=\"Tournament\",\"Year\"=\"Season\")) %>% select(-Margin_Q90,-Margin_Q10,-AVG_MARGIN)\nquartile_diff <- quartile_diff %>% mutate_at(c(\"TWOPTS_QDIFF\",\"THRPTS_QDIFF\",\"FTPTS_QDIFF\",\"ASST_QDIFF\",\"TO_QDIFF\",\"ST_QDIFF\",\"BLCK_QDIFF\",\"DRBND_QDIFF\",\"ORBND_QDIFF\",\"MARGIN_QDIFF\"), ~(range01(.) %>% as.vector))\ncomp<- quartile_diff %>%  subset(select=3:12) %>% mutate(Competitiveness=rowSums(.))\nquartile_diff <- quartile_diff %>% left_join(comp[,c(1,2,4,5,9,10,11)],by=c(\"TWOPTS_QDIFF\"=\"TWOPTS_QDIFF\",\"THRPTS_QDIFF\"=\"THRPTS_QDIFF\",\"ASST_QDIFF\"=\"ASST_QDIFF\",\"MARGIN_QDIFF\"=\"MARGIN_QDIFF\",\"TO_QDIFF\"=\"TO_QDIFF\",\"ORBND_QDIFF\"=\"ORBND_QDIFF\"))\nquartile_diff <- quartile_diff %>% arrange(Year, Competitiveness) %>% mutate(Comp_Rank = rank(Competitiveness)) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Above we calculated quartiles for the same metrics for each year and tournament. The plot below shows the changes in the competitiveness ranking of the different competitions. It looks like until recent years regular season of the NCAA has the teams with the closest metrics to each other. However Euroleague able to overcome it twice in 2016 and 2018.  When I dive in to see which metrics (*see the non-normalized version of the metrics below for the years where the Euroleague became most competitive tournament*) caused this change from the European basketball, I see that NCAA-Regular season is better than the Euroleague in terms of shootings however, Euroleague is significantly better on competitiveness where difference in number of turovers, steals, blocks, rebounds and most importantly margin between teams much less. It is very much similar story for the 2018 as well."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"agg_df_yby  %>%\n  group_by(Year,Tournament) %>%\n  summarize_at(vars(TWOPTS_PERC,THRPTS_PERC,FT_PERC,ASST_PG,TO_PG,ST_PG,BLCK_PG,DRBND_PG,ORBND_PG), funs(!!!p_funs)) %>% mutate(TWOPTS_QDIFF = `TWOPTS_PERC_90%` - `TWOPTS_PERC_10%`,\n                                                                                                                                THRPTS_QDIFF = `THRPTS_PERC_90%` - `THRPTS_PERC_10%`, \n                                                                                                                                FTPTS_QDIFF = `FT_PERC_90%` - `FT_PERC_10%`,\n                                                                                                                                ASST_QDIFF = `ASST_PG_90%` - `ASST_PG_10%`, \n                                                                                                                                TO_QDIFF = `TO_PG_90%` - `TO_PG_10%`, \n                                                                                                                                ST_QDIFF = `ST_PG_90%` - `ST_PG_10%`,\n                                                                                                                               BLCK_QDIFF = `BLCK_PG_90%` - `BLCK_PG_10%`,\n                                                                                                                                DRBND_QDIFF = `DRBND_PG_90%` - `DRBND_PG_10%`,\n                                                                                                                                ORBND_QDIFF = `ORBND_PG_90%` - `ORBND_PG_10%`) %>%  select(Tournament,Year,matches('QDIFF'))%>% left_join(margin,by=c(\"Tournament\"=\"Tournament\",\"Year\"=\"Season\")) %>% select(-Margin_Q90,-Margin_Q10,-AVG_MARGIN) %>% filter(Year==2016 | Year == 2018)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig(20,20)\nggplot(data = quartile_diff, aes(x = Year, y = Comp_Rank, fill = Tournament)) +\ngeom_line(aes(col=Tournament),size=2) + \ntheme_bw()+\n xlab(\"Season\")+\n    ylab(\"Competitiveness Rank Between Different Competitions\")+\n    labs(title= \"Competitiveness Rank of the Competitions Over the Years By Simple Averaging\")+\n    theme(panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=24), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 24),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=24),          \n          axis.text.x = element_text(color=\"black\",face=\"bold\", \n                           size=18),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=18),\n         legend.text=element_text(size=26),legend.title=element_blank(),legend.key.size = unit(3,\"line\"))+\n         guides(colour = guide_legend(override.aes = list(size=3)))\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 3.2.2. Four Factors\n#### 3.2.2.1 Overall\n\n\nThis approach focuses on four key factors at success in basketball with varying weights. We will use this approach and look for difference between team scores at 90th and 10th quartiles. Roughly these [four factors are calculated by](https://squared2020.com/2017/09/05/introduction-to-olivers-four-factors/) equations below. Let's do it!\n\n**Shooting- Effective Field Goal Percentage\n**\n\neFG% = (Field Goals + 0.5 * 3-pointers) / Field Goal Attempts\n\n**Turnovers- Turnover Rate\n**\n\nTOV% = Turnovers / (Field Goal Attemps + 0.44 * Free Throws + Turnovers)\n\n**Rebounding- Offensive and Defensive Rebounding Percentage\n**\n\nORB% = ORB / (ORB + Opp DRB)\n\n**Free Throws- Free Throw Rate\n**\nFTR = FTA/FGA\n\n**Overall Formula**\n\n(eFG% * 0.475) - (TOV% * 0.25) + (0.2 * ORB%)  + (0.15 * FT_PERC)"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Calculating four factors. Since we already have free throw success rate, I do not calculate it.\nagg_df <- agg_df %>% mutate(EFG = (TWOPTS_MAKE_PG + (THR_MAKE_PG * 0.5)) / FG_ATTEMP_PG, ORB = ORBND_PG / (ORBND_PG + OPP_OREB_PG ), TOV = TO_PG / (FG_ATTEMP_PG + (0.475* FT_ATTEMPS_PG) + TO_PG ),FOURFACTORINDEX = 0.4* EFG + 0.2*ORB + 0.15* FT_PERC - 0.25 * TOV)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"Before looking for four factor index, I would like to see where teams are these factors. First I plotted field goal efficiency and to be able to plot efficiently selected only teams that played at least 50 games for European competitions, 20 games for tournament part of the NCAA and at least 375 games for the NCAA regular season. As a result, we have seen that teams are really close to each other with small margins particularly at Europe."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"HEIGHT = 1200\nWIDTH = 900\nplt13 = agg_df %>% filter((Tournament != \"NCAA-Tournament\" & Tournament !=\"NCAA-Regular\"& TOTAL_GAMES > 50) | (Tournament ==\"NCAA-Regular\" & TOTAL_GAMES > 380) | (Tournament ==\"NCAA-Tournament\" & TOTAL_GAMES > 20)) %>% mutate(TEAM = fct_reorder(TEAM,EFG)) %>% ggplot(., aes(x =EFG,y=TEAM,fill=TOTAL_GAMES,text = paste('Field Goal Efficiency', EFG,\n                                              '<br>Team Name:', TEAM,\n                                              '<br>Number of Games Played:', TOTAL_GAMES))) +\n    geom_bar(stat=\"identity\")+\n    facet_wrap(Tournament~.,scales = \"free_y\",nrow=3)+\n    theme_bw()+\n scale_x_continuous(limits = c(0,0.35),breaks=c(0,0.25,0.35),\n                     labels = scales::percent_format(accuracy = 5L))+\nxlab(\"Field Goal Efficiency\")+\n    labs(title= \"Field Goal Efficiency\", subtitle = \"By Different Competitions\")+\n    theme(legend.position = \"none\",panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          strip.text = element_text(size=15,face=\"bold\", colour = \"black\"),\n          strip.background =element_rect(fill=\"white\",color=\"white\",size=4),\n          axis.title.y = element_blank(), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 18),\n          plot.margin = unit(c(2, 0, 0, 0), \"cm\"),   \n          plot.subtitle = element_text(hjust = 0.5, vjust=1,size=16),\n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=12),          \n          axis.text.x = element_text(face=\"italic\", color=\"black\", \n                           size=12),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=9))\n\npltly = ggplotly(plt13, tooltip = c('text')) %>%   layout(title = list(text = paste0('Field Goal Efficiency',\n                                    '<br>',\n                                    '<sup>',\n                                    'By Different Competitions',\n                                    '</sup>')))\n\n# this is just a workaround to make the kaggle kernel show the plot correctly\nhtmlwidgets::saveWidget(pltly, \"plot13.html\")\nIRdisplay::display_html(sprintf('<iframe src=\"plot13.html\" height=%d width=%d></iframe>', HEIGHT, as.integer(WIDTH)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Below we see the difference between tournaments in terms of four factor index. As in first method, NCAA Regular season resulted the most competitive over the years while European competitions following them slightly behind. Looking at individual factors, we have seen that teams are closer to each other at almost all metrics except Free Throws with almost 0.1 point difference."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"agg_df %>% group_by(Tournament) %>% summarize_at(vars(FOURFACTORINDEX,TOV,EFG,ORB,FT_PERC), funs(!!!p_funs)) %>% mutate(FOURFACTORINDEX_QDIFF = `FOURFACTORINDEX_90%` - `FOURFACTORINDEX_10%`,EFG_QDIFF = `EFG_90%` - `EFG_10%`,TOV_QDIFF = `TOV_90%` - `TOV_10%`,\n                                    ORB_QDIFF = `ORB_90%` - `ORB_10%`,FT_QDIFF = `FT_PERC_90%` - `FT_PERC_10%`) %>% select(Tournament,FOURFACTORINDEX_QDIFF,EFG_QDIFF,TOV_QDIFF,ORB_QDIFF,FT_QDIFF)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 3.2.1.1 Year by Year\n\nAfter analyzing overall competitiveness using Four Factor, now I looked for year by year difference in competitiveness by four factor index."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Calculating four factors year by year.\nagg_df_yby <- agg_df_yby %>% mutate(EFG = (TWOPTS_MAKE_PG + (THR_MAKE_PG * 0.5)) / FG_ATTEMP_PG, ORB = ORBND_PG / (ORBND_PG + OPP_OREB_PG ), TOV = TO_PG / (FG_ATTEMP_PG + (0.475* FT_ATTEMPS_PG) + TO_PG ),FOURFACTORINDEX = 0.4* EFG + 0.2*ORB + 0.15* FT_PERC - 0.25 * TOV)\nfour_fac_yby <- agg_df_yby %>% group_by(Year,Tournament) %>% summarize_at(vars(FOURFACTORINDEX,TOV,EFG,ORB,FT_PERC), funs(!!!p_funs)) %>% mutate(FOURFACTORINDEX_QDIFF = `FOURFACTORINDEX_90%` - `FOURFACTORINDEX_10%`,EFG_QDIFF = `EFG_90%` - `EFG_10%`,TOV_QDIFF = `TOV_90%` - `TOV_10%`,ORB_QDIFF = `ORB_90%` - `ORB_10%`,FT_QDIFF = `FT_PERC_90%` - `FT_PERC_10%`) %>% select(Tournament,FOURFACTORINDEX_QDIFF,EFG_QDIFF,TOV_QDIFF,ORB_QDIFF,FT_QDIFF) %>%  arrange(Year, FOURFACTORINDEX_QDIFF) %>% mutate(Comp_Rank = rank(FOURFACTORINDEX_QDIFF)) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here is the year by year change of the competitiveness ranking of the leagues. We see that this time regular season of the NCAA is more competitive than others while Euroleague able to overcome NCAA once at 2018. It looks like using different methods do not change the result except for the 2016."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig(20,20)\nggplot(data = four_fac_yby, aes(x = Year, y = Comp_Rank, fill = Tournament)) +\ngeom_line(aes(col=Tournament),size=2) + \ntheme_bw()+\n xlab(\"Season\")+\n    ylab(\"Competitiveness Rank Between Different Competitions\")+\n    labs(title= \"Competitiveness Rank of the Competitions Over the Years By Four Factors\")+\n    theme(panel.grid.major = element_blank(), \n          panel.grid.minor = element_blank(),\n          axis.title.y = element_text(face='bold', colour='saddlebrown', size=32), \n          plot.title = element_text(colour='saddlebrown',face=\"bold\",hjust = 0.5,vjust=1,size = 32),\n          plot.margin = unit(c(2, 3, 3, 3), \"cm\"),   \n          axis.title.x = element_text(face='bold', colour='saddlebrown', size=32),          \n          axis.text.x = element_text(color=\"black\",face=\"bold\", \n                           size=18),\n          axis.text.y = element_text(face=\"bold\", color=\"black\", \n                           size=18),\n         legend.text=element_text(size=26),legend.title=element_blank(),legend.key.size = unit(3,\"line\"))+\n         guides(colour = guide_legend(override.aes = list(size=3)))\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 4. Clutch Time: Generalization Challenge\n\nAim of this challenge is using the features we have created previously to predict the results of the games from another continent. I will be doing two different cases, first will try to predict European games with the NCAA results and then do the opposite. Below there data preparation for train and test sets which has per game statistics and four factor indexes year by year for teams. \n\nFirst I tried to predict the Euroleague results with the NCAA games and only succeeded for **65%** of the accuracy.  When I do the same for the NCAA with the European competitions data, result was only **53%**. \n\nNote that even though this result can be much higher with some more feature engineering and hyperparemeter tuning, my aim was only to show a baseline and whether we can predict results of the NCAA results with European competitions and vice versa. Result is that, it is slightly easier to predict European competitions while it is harder to predict the NCAA games with basic features as a baseline. This is another sign of the NCAA's competitiveness. "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"win <- detailed_res_ncaa %>% select(WTeamID,LTeamID,Season) %>% mutate(Target = 1 )  %>% left_join(teamnames, by=c(\"WTeamID\"=\"TeamID\")) %>% rename(\"WTeamName\"=\"TeamName\") %>% select(-FirstD1Season,-LastD1Season,-WTeamID) %>% left_join(teamnames, by=c(\"LTeamID\"=\"TeamID\")) %>% rename(\"LTeamName\"=\"TeamName\")  %>% select(-FirstD1Season,-LastD1Season,-LTeamID) %>% left_join(agg_df_yby,by=c(\"WTeamName\"=\"TEAM\",\"Season\"=\"Year\")) %>% left_join(agg_df_yby,by=c(\"LTeamName\"=\"TEAM\",\"Season\"=\"Year\"))\nlose <- detailed_res_ncaa %>% select(WTeamID,LTeamID,Season) %>% mutate(Target = 0 )  %>% left_join(teamnames, by=c(\"WTeamID\"=\"TeamID\")) %>% rename(\"WTeamName\"=\"TeamName\") %>% select(-FirstD1Season,-LastD1Season,-WTeamID) %>% left_join(teamnames, by=c(\"LTeamID\"=\"TeamID\")) %>% rename(\"LTeamName\"=\"TeamName\")  %>% select(-FirstD1Season,-LastD1Season,-LTeamID) %>% left_join(agg_df_yby,by=c(\"LTeamName\"=\"TEAM\",\"Season\"=\"Year\")) %>% left_join(agg_df_yby,by=c(\"WTeamName\"=\"TEAM\",\"Season\"=\"Year\"))\ntrain <- rbind(win,lose)\ntrain <- train %>% select(TOTAL_GAMES.x,TOTAL_GAMES.y,TWOPTS_PERC.x,TWOPTS_PERC.y,THRPTS_PERC.x,THRPTS_PERC.y,FT_PERC.x,FT_PERC.y,SCORE_PG.x,SCORE_PG.y,TWOPTS_ATTMPT_PG.x,TWOPTS_ATTMPT_PG.y,TWOPTS_MAKE_PG.x,TWOPTS_MAKE_PG.y,\nTHR_MAKE_PG.x,THR_MAKE_PG.y,THR_ATTMPT_PG.x,THR_ATTMPT_PG.y,FT_ATTEMPS_PG.x,FT_ATTEMPS_PG.y,FT_MAKE_PG.x,FT_MAKE_PG.y,ASST_PG.x,ASST_PG.y,DRBND_PG.x,DRBND_PG.y,ORBND_PG.x,ORBND_PG.y,TO_PG.x,TO_PG.y,ST_PG.x,ST_PG.y,BLCK_PG.x,BLCK_PG.y,OPP_DREB_PG.x,OPP_DREB_PG.y,EFG.x,EFG.y,ORB.x,ORB.y,TOV.x,TOV.y,FOURFACTORINDEX.x,FOURFACTORINDEX.y,Target)\n\ntest <- euroleague %>% group_by(TeamA,TeamB,gamenumber,yer) %>% summarize(POINTA = max(POINTS_A,na.rm=TRUE),POINTB = max(POINTS_B,na.rm=TRUE)) %>% mutate(Target = ifelse(POINTA > POINTB,1, 0)) %>% select(TeamA,TeamB,yer,Target) %>% ungroup() %>% select(-gamenumber)  %>% left_join(agg_df_yby,by=c(\"TeamA\"=\"TEAM\",\"yer\"=\"Year\")) %>% left_join(agg_df_yby,by=c(\"TeamB\"=\"TEAM\",\"yer\"=\"Year\"))\ntest <- test %>% select(TOTAL_GAMES.x,TOTAL_GAMES.y,TWOPTS_PERC.x,TWOPTS_PERC.y,THRPTS_PERC.x,THRPTS_PERC.y,FT_PERC.x,FT_PERC.y,SCORE_PG.x,SCORE_PG.y,TWOPTS_ATTMPT_PG.x,TWOPTS_ATTMPT_PG.y,TWOPTS_MAKE_PG.x,TWOPTS_MAKE_PG.y,\nTHR_MAKE_PG.x,THR_MAKE_PG.y,THR_ATTMPT_PG.x,THR_ATTMPT_PG.y,FT_ATTEMPS_PG.x,FT_ATTEMPS_PG.y,FT_MAKE_PG.x,FT_MAKE_PG.y,ASST_PG.x,ASST_PG.y,DRBND_PG.x,DRBND_PG.y,ORBND_PG.x,ORBND_PG.y,TO_PG.x,TO_PG.y,ST_PG.x,ST_PG.y,BLCK_PG.x,BLCK_PG.y,OPP_DREB_PG.x,OPP_DREB_PG.y,EFG.x,EFG.y,ORB.x,ORB.y,TOV.x,TOV.y,FOURFACTORINDEX.x,FOURFACTORINDEX.y,Target)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's have a look at train and test sets. "},{"metadata":{"trusted":true},"cell_type":"code","source":"head(train)\nhead(test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"targetTrain <- train$Target\ntargetTest <- test$Target\ntrainxgb <- train %>% select(-Target)\ntestxgb <- test %>% select(-Target)\ndtrain <- xgb.DMatrix(data = as.matrix(trainxgb),label = targetTrain)\ndtest <- xgb.DMatrix(data = as.matrix(testxgb))\nparams <- list(booster = \"gbtree\", objective = \"binary:hinge\", max_depth=4, min_child_weight=5, learning_rate=0.1,\n               subsample=0.3, colsample_bytree=0.9,eval_metric = \"auc\",lambda=100)\nxgb <- xgboost(params = params, data = dtrain, nrounds =250, \n                 showsd = T, \n                print_every_n = 50, maximize = F)\npred <- predict(xgb, dtest)\nprint(Accuracy(pred,targetTest))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"importance_matrix <- xgb.importance(colnames(dtrain), model = xgb)\nfig(20,20)\nxgb.ggplot.importance(importance_matrix, rel_to_first = TRUE, xlab = \"Relative Imp\")+\ntheme( text = element_text(size = 20), axis.text.x = element_text(size = 15, angle = 45, hjust = 1))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Swap train-test so that we can use Euroleague to predict the NCAA games.\n\ntargetTrain <- test$Target\ntargetTest <- train$Target\ntrainxgb <- test %>% select(-Target)\ntestxgb <- train %>% select(-Target)\ndtrain <- xgb.DMatrix(data = as.matrix(trainxgb),label = targetTrain)\ndtest <- xgb.DMatrix(data = as.matrix(testxgb))\nparams <- list(booster = \"gbtree\", objective = \"binary:hinge\", max_depth=4, min_child_weight=5, learning_rate=0.1,\n               subsample=0.3, colsample_bytree=0.9,eval_metric = \"auc\",lambda=100)\nxgb <- xgboost(params = params, data = dtrain, nrounds =250, \n                 showsd = T, \n                print_every_n = 50, maximize = F)\npred <- predict(xgb, dtest)\nprint(Accuracy(pred,targetTest))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 4. Overtime: Concluding Remarks\n\nThank you very much if you come this far in my notebook! Here I would like to summarize my findings and complete my comparative analysis.\n\n* Even though European competitions are professional, competitiveness is slightly higher at the NCAA. We tried two different methods, one is a simple method and another method is well-known four factor method at basketball analytics. Both gave us similar results.\n\n* Possible reasons for that are, in Europe teams who invest heavily do not change year by year and certain teams are always at the top. While at NCAA, players or in another word students, change every year which can affect strength of the team heavily.\n\n* At most key statistics, average of the teams at European competitions are much better than the NCAA teams however difference between the NCAA teams are smaller. At the recent years European competitions are started to get more competitive particularly by some metrics.\n\n\n\n## Limitations and Future Work\n\nWhile there are many ways to quantify competitiveness, if i had more time to dive into this data I would do those to dive deeper:\n\n- For players and teams efficiency, as an alternative to the Four Factors, we can look for Player Efficiency Rating (PER) by [Hollinger](https://bleacherreport.com/articles/1813902-advanced-nba-stats-for-dummies-how-to-understand-the-new-hoops-math) which takes into account minutes played by player and teams pace. Also I would look for Performance Index Rating: [PIR](https://en.wikipedia.org/wiki/Performance_Index_Rating)\n\n- Include women's games to the analysis. I would like to do the same for Euroleague Women and the Women's NCAA competition and see the results. This could be a challenge for me for 2021 season.\n\n- Also, working on whether teams are playing more collaboratively or stars are leading the way is another interesting research question that can be worked on.\n\n## Miscellanous Questions That May Pop-up\n\n* Why did you divide tournaments into Euroleague and Eurocup on the one hand and NCAA-Regular Season and NCAA-Tournament on the other hand? Why Euroleague play-off are not included? \n\nAt Euroleague playoff's only a few games played every year compared to the NCAA and comparing little number of games do not sound intuitive to me. Hence, I decided to divide competitions in this way.\n\n* Why did you select 10th and 90th quartiles? \n\nI chose these numbers intuitively with some trials and errors in order to avoid outliers and quantify competitiveness.\n\n* How did you select model parameters?\n\nI did couple of trial and error in order to optimize results. Since main aim is not modelling, I would like to reach a baseline with the best possible parameters.\n\nThank you all for reading, wish everyone healthy days!"},{"metadata":{},"cell_type":"markdown","source":"**Code Credits**\n\n[Subtitles with ggplotly](https://datascott.com/blog/subtitles-with-ggplotly/)\n\n[Tooltip formatting with ggplotly](https://stackoverflow.com/questions/50222764/how-do-i-format-the-names-of-the-variables-in-the-r-plotly-tooltip)\n\n[Quick fix to calculate quartiles](https://tbradley1013.github.io/2018/10/01/calculating-quantiles-for-groups-with-dplyr-summarize-and-purrr-partial)\n \n **Basketball Analytics Credits**\n\n[3 Point Revolution](https://shottracker.com/articles/the-3-point-revolution)\n\n[Four Factors](https://squared2020.com/2017/09/05/introduction-to-olivers-four-factors/)\n\n[Four Factors-2](http://www.basketballonpaper.com/)\n\n[Hollinger](https://bleacherreport.com/articles/1813902-advanced-nba-stats-for-dummies-how-to-understand-the-new-hoops-math) \n\n[NCAA Basic Info](http://www.ncaa.org/about/resources/media-center/ncaa-101/what-ncaa)\n\n[Performance Index Rating](https://en.wikipedia.org/wiki/Performance_Index_Rating)\n\n"}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":4}