{"cells":[{"metadata":{},"cell_type":"markdown","source":"<div class=\"h1\">The subtle art of winning March Madness</div>"},{"metadata":{},"cell_type":"markdown","source":"<img src=\"https://github.com/javiferran/ncaa-analytics-2020/blob/master/kyle_guy.jpg?raw=true[](http://)\" alt=\"zeke\" height=\"400\" width=\"720\"/>"},{"metadata":{},"cell_type":"markdown","source":"This notebook tries to find some sense in (March) Madness. My approach is focused on the study of top teams and 'weak' a priori teams that become cinderellas. In this work I try to understand what makes teams perform well in the end-of season tournament and what makes some 'underdogs' join that category. Note than I consider top performers those that made it to the Elite Eight round. In order to search for answers to this problem I have mainly focused on the data provided by the regular season. Since every year teams change a lot, I discarded data further from the same season where the tournament takes place because I assume it would add more noise than signal."},{"metadata":{"id":"Q1mWub3Ob_2j","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# Import Libraries\nimport pandas as pd\nimport numpy as np\nfrom sklearn.linear_model import LogisticRegression\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nfrom sklearn.utils import shuffle\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import KFold\nimport lightgbm as lgb\n\nimport xgboost as xgb\nfrom xgboost import XGBClassifier\nimport gc\nimport copy\nfrom sklearn.linear_model import LogisticRegression\n\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.decomposition import PCA\nfrom sklearn.cluster import KMeans\nimport plotly.express as px\nfrom matplotlib.transforms import Affine2D\nfrom sklearn.metrics.pairwise import cosine_similarity\n#from adjustText import adjust_text\n\nfrom sklearn.cluster import AgglomerativeClustering\nimport scipy.cluster.hierarchy as sch\n\nfrom advanced_stats import *\nfrom elo_points import *\nfrom matplotlib.pylab import rcParams\nimport warnings \nwarnings.filterwarnings('ignore')\n\nimport seaborn as sns\nfrom IPython.core.display import HTML\nHTML(\"\"\"\n    <style>\n    .output_png {\n        display: table-cell;\n        text-align: center;\n        margin:auto;\n    }\n    .prompt \n        display:none;\n    }  \n    </style>\n\"\"\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The sections in which I divide my analysis are:\n  -  [Do cinderellas go far?](#Do-cinderellas-go-far?)\n  -  [What do competitive teams have in common?](#What-do-competitive-teams-have-in-common?)\n  -  [Predictive Model](#Predictive-Model)\n  -  [Where do predictive models fail?](#Where-do-predictive-models-fail?)"},{"metadata":{"id":"6NdTIrnl59WL","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"#Delete leaked games in training\ndef concat_row(r):\n    if r['WTeamID'] < r['LTeamID']:\n        res = str(r['Season'])+\"_\"+str(r['WTeamID'])+\"_\"+str(r['LTeamID'])\n    else:\n        res = str(r['Season'])+\"_\"+str(r['LTeamID'])+\"_\"+str(r['WTeamID'])\n    return res\n\n# Delete leaked from train\ndef delete_leaked_from_df_train(df_train, df_test):\n    df_train['Concats'] = df_train.apply(concat_row, axis=1)\n    df_train_duplicates = df_train[df_train['Concats'].isin(df_test['ID'].unique())]\n    df_train_idx = df_train_duplicates.index.values\n    df_train = df_train.drop(df_train_idx)\n    df_train = df_train.drop('Concats', axis=1)\n    \n    return df_train \n\ndef read_data(inFile, sep=','):\n    df_op = pd.read_csv(filepath_or_buffer=inFile, low_memory=False, encoding='utf-8', sep=sep)\n    return df_op\n\n\n#ID calculator\ndef id_calculator(team_name_guess,Teams):\n    identifier = Teams[(Teams.TeamName.str.startswith(team_name_guess))]['TeamID'].to_numpy()\n    if len(identifier) != 1:\n        print('Several teams with those characters')\n        return int(identifier[0])\n    else:\n        return int(identifier)\n\n#Team names claculator\ndef name_calculator(id,Teams):\n    names_list_from_id = []\n    for identifier in id:\n        names = Teams[Teams.TeamID == identifier].TeamName.to_numpy()[0]\n        names_list_from_id.append(names)\n\n    return(names_list_from_id)\n\n#Rivals calculator\ndef season_rivals(team_name,season,season_compact):\n    team_id = id_calculator(team_name)\n    loser_rivals = season_compact[(season_compact.WTeamID==team_id)& (season_compact.Season==season)]['LTeamName'].to_list()\n    winner_rivals = season_compact[(season_compact.LTeamID==team_id)& (season_compact.Season==season)]['WTeamName'].to_list()\n\n    return loser_rivals, winner_rivals\n\ndef tournament_rivals(team_name,season,tourney_df):\n    team_id = id_calculator(team_name,MTeams)\n    loser_rivals = tourney_df[(tourney_df.WTeamID==team_id)& (tourney_df.Season==season)]['LTeamID'].to_list()\n    winner_rivals = tourney_df[(tourney_df.LTeamID==team_id)& (tourney_df.Season==season)]['WTeamID'].to_list()\n\n    return loser_rivals, winner_rivals\n\n#Elite Eight\ndef elite_eight(season_year,tourney_df,Teams):\n    # check if men or women data\n    if str(tourney_df.WTeamID.to_numpy()[0])[0] is '1':\n        Wfinal_four_ids = tourney_df[((tourney_df.DayNum==145) | (tourney_df.DayNum==146)) & (tourney_df.Season==season_year)]['WTeamID'].to_list()\n        Lfinal_four_ids = tourney_df[((tourney_df.DayNum==145) | (tourney_df.DayNum==146)) & (tourney_df.Season==season_year)]['LTeamID'].to_list()\n    else:\n        if season_year>=2003 | season_year<=2016:\n            Wfinal_four_ids = tourney_df[((tourney_df.DayNum==147) | (tourney_df.DayNum==148)) & (tourney_df.Season==season_year)]['WTeamID'].to_list()\n            Lfinal_four_ids = tourney_df[((tourney_df.DayNum==147) | (tourney_df.DayNum==148)) & (tourney_df.Season==season_year)]['LTeamID'].to_list()\n        else:\n            Wfinal_four_ids = tourney_df[((tourney_df.DayNum==146) | (tourney_df.DayNum==147)) & (tourney_df.Season==season_year)]['WTeamID'].to_list()\n            Lfinal_four_ids = tourney_df[((tourney_df.DayNum==146) | (tourney_df.DayNum==147)) & (tourney_df.Season==season_year)]['LTeamID'].to_list()\n\n    elite_eight_ids = Wfinal_four_ids + Lfinal_four_ids\n\n    elite_eight_names = []\n\n    for id in elite_eight_ids:\n        elite_eight_name = Teams[Teams.TeamID == id]['TeamName'].to_numpy()[0]\n        elite_eight_names.append(elite_eight_name)\n\n    return elite_eight_names\n\n#Final Four\n#Days 145 and 146 Elite Eight games are played (winners teams play the Final Four). Day 152, Final Four games and Day 154 the final.\ndef final_four(season_year,tourney_df,Teams):\n    if str(tourney_df.WTeamID.to_numpy()[0])[0] is '1':\n        final_four_ids = tourney_df[((tourney_df.DayNum==145) | (tourney_df.DayNum==146)) & (tourney_df.Season==season_year)]['WTeamID'].to_list()\n    else:\n        if season_year>=2003 & season_year<=2016:\n            final_four_ids = tourney_df[((tourney_df.DayNum==147) | (tourney_df.DayNum==148)) & (tourney_df.Season==season_year)]['WTeamID'].to_list()\n        else:\n            final_four_ids = tourney_df[((tourney_df.DayNum==146) | (tourney_df.DayNum==147)) & (tourney_df.Season==season_year)]['WTeamID'].to_list()\n    \n    final_four_names = []\n    for id in final_four_ids:\n        final_four_name = Teams[Teams.TeamID == id]['TeamName'].to_numpy()[0]\n        final_four_names.append(final_four_name)\n    return final_four_names\n\n# Upsets calculator\ndef upsets_calculator(season_year,tourney_df):\n    upset_df = tourney_df[(tourney_df.Seed_diff >= 5) & (tourney_df.Season == season_year)]\n    upset_df = upset_df[['WTeamID','LTeamID']]\n    return upset_df\n\n# Calculate teams which played in the tournament\ndef march_madness_teams(season_year,tourney_df):\n    winner_teams = list(set(tourney_df[tourney_df.Season == season_year]['WTeamID'].to_list()))\n    loser_teams = list(set(tourney_df[tourney_df.Season == season_year]['LTeamID'].to_list()))\n    mm_teams = winner_teams + loser_teams\n    return mm_teams\n\n#Distance between 2 points on Earth\nfrom math import radians, cos, sin, asin, sqrt \ndef distance(lat1, lat2, lon1, lon2): \n      \n    # The math module contains a function named \n    # radians which converts from degrees to radians. \n    lon1 = radians(lon1) \n    lon2 = radians(lon2) \n    lat1 = radians(lat1) \n    lat2 = radians(lat2) \n       \n    # Haversine formula  \n    dlon = lon2 - lon1  \n    dlat = lat2 - lat1 \n    a = sin(dlat / 2)**2 + cos(lat1) * cos(lat2) * sin(dlon / 2)**2\n  \n    c = 2 * asin(sqrt(a))  \n     \n    # Radius of earth in kilometers. Use 3956 for miles \n    r = 6371\n       \n    # calculate the result \n    return(c * r) ","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","id":"YpoHl17fb_2o","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"#Data loading\ntourney_result = pd.read_csv('../input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneyCompactResults.csv')\nWtourney_result = pd.read_csv('../input/march-madness-analytics-2020/WDataFiles_Stage2/WNCAATourneyCompactResults.csv')\ntourney_seed = pd.read_csv('../input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneySeeds.csv')\nWtourney_seed = pd.read_csv('../input/march-madness-analytics-2020/WDataFiles_Stage2/WNCAATourneySeeds.csv')\n#submission_df = pd.read_csv(path_to_drive + '/input/google-cloud-ncaa-march-madness-2020-division-1-mens-tournament/MSampleSubmissionStage1_2020.csv')\nMRegularSeasonCompactResults = pd.read_csv('../input/march-madness-analytics-2020/MDataFiles_Stage2/MRegularSeasonCompactResults.csv')\nMTeams = pd.read_csv('../input/march-madness-analytics-2020/MDataFiles_Stage2/MTeams.csv')\nWTeams = pd.read_csv('../input/march-madness-analytics-2020/WDataFiles_Stage2/WTeams.csv')\n\nMEvents_2018 = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020DataFiles/2020-Mens-Data/MEvents2018.csv')\nMEvents_2017 = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020DataFiles/2020-Mens-Data/MEvents2017.csv')\nMEvents_2016 = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020DataFiles/2020-Mens-Data/MEvents2016.csv')\n\n# game cities data\ngame_cities = pd.read_csv('../input/march-madness-analytics-2020/MDataFiles_Stage2/MGameCities.csv')\n# city names\ncity_names = pd.read_csv('../input/march-madness-analytics-2020/MDataFiles_Stage2/Cities.csv')\nlocation_data = pd.read_csv('/kaggle/input/location-data/location_data.csv')\n# Adding Team Names to SeasonCompactResults\nMRegularSeasonCompactResults = pd.merge(MRegularSeasonCompactResults,MTeams[['TeamID','TeamName']],left_on=['WTeamID'], right_on=['TeamID'], how='left')\nMRegularSeasonCompactResults = pd.merge(MRegularSeasonCompactResults,MTeams[['TeamID','TeamName']],left_on=['LTeamID'], right_on=['TeamID'], how='left')\n\nMRegularSeasonCompactResults = MRegularSeasonCompactResults.drop('TeamID_x', axis=1)\nMRegularSeasonCompactResults = MRegularSeasonCompactResults.drop('TeamID_y', axis=1)\nMRegularSeasonCompactResults.rename(columns={'TeamName_x':'WTeamName'}, inplace=True)\nMRegularSeasonCompactResults.rename(columns={'TeamName_y':'LTeamName'}, inplace=True)\n\nWRegularSeasonCompactResults = pd.read_csv('../input/march-madness-analytics-2020/WDataFiles_Stage2/WRegularSeasonCompactResults.csv')\n#season_enriched = create_advanced_stats()\nseason_enriched = pd.read_csv('../input/mncaa-enriched-season-data/MNCAASeasonDetailedResultsEnriched.csv')\nWseason_enriched = pd.read_csv('../input/mncaa-enriched-season-data/WNCAASeasonDetailedResultsEnriched.csv')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"elo_points = get_elo_points(MRegularSeasonCompactResults)\nWelo_points = get_elo_points(WRegularSeasonCompactResults)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"tourney_result[(tourney_result.Season==2017) & (tourney_result.LTeamID==1376)]\n#id_calculator('South Car')\ngame_cities[(game_cities.Season==2017) & (game_cities.CRType=='NCAA') & (game_cities.WTeamID==1376)]\ncity_names[city_names.CityID==4139]","execution_count":null,"outputs":[]},{"metadata":{"id":"km1J1f2TRA1d","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# SEASON DATA\n# Number of wins in last 10 games calculation\ndef get_season_data(season_enriched, elo_points,Teams_name_df):\n    Wins = season_enriched[['Season','WTeamID','DayNum']]\n    Wins.rename(columns={'WTeamID':'TeamID'}, inplace=True)\n    Wins['Result'] = 1\n    Loses = season_enriched[['Season','LTeamID','DayNum']]\n    Loses.rename(columns={'LTeamID':'TeamID'}, inplace=True)\n    Loses['Result'] = 0\n    season_results = pd.concat((Wins,Loses)).reset_index(drop=True)\n    season_results = season_results.sort_values(by=['Season', 'TeamID','DayNum'])\n    teams = season_results['TeamID'].unique()\n    years = season_results['Season'].unique()\n    last_10_df = pd.DataFrame(columns = season_results.columns)\n    for year in years:\n      for team in teams:\n      #for index, row in a.iterrows():\n        last_10 = season_results[(season_results['Season']== year) & (season_results['TeamID']==team)][-10:]\n        last_10_df = pd.concat((last_10_df,last_10))\n    last_10_df_sum = last_10_df.groupby(['Season', 'TeamID'])['Result'].sum().reset_index()\n    last_10_df_sum.rename(columns={'Result':'L10wins'}, inplace=True)\n\n    # Regular season aggregated data\n    style_metrics = ['Pos','AstR','Stl','Blk','FTAR','ORP','TOR','TSP']\n    performance_metrics = ['Score','eFGP','NetRtg','PIE']\n    other_metrics = ['FGA3','DR']\n    metrics = performance_metrics + style_metrics + other_metrics\n\n    winner_metrics = []\n    loser_metrics = []\n\n    for metric in metrics:\n        winner_metrics.append('W' + metric)\n        loser_metrics.append('L' + metric)\n\n    season_enrichedW = season_enriched[['Season','WTeamID'] + winner_metrics]\n    season_enrichedW.rename(columns={'WTeamID':'TeamID'}, inplace=True)\n\n    for variable in winner_metrics:\n\n        season_enrichedW.rename(columns={variable:variable[1:]}, inplace=True)\n\n\n    season_enrichedL = season_enriched[['Season','LTeamID'] + loser_metrics]\n    season_enrichedL.rename(columns={'LTeamID':'TeamID'}, inplace=True)\n    for variable in loser_metrics:\n        season_enrichedL.rename(columns={variable:variable[1:]}, inplace=True)\n\n    season_enriched_concatenation = pd.concat((season_enrichedW, season_enrichedL)).reset_index(drop=True)\n    season_enriched_concatenation = season_enriched_concatenation.groupby(['Season', 'TeamID'])[metrics].mean().reset_index()\n\n\n    # Merge season_enriched_concat with Elo table, Teams Names and last 10 games wins\n    season_enriched_concatenation = pd.merge(season_enriched_concatenation, elo_points, left_on=['Season', 'TeamID'], right_on=['season', 'team_id'], how='left')\n    #season_enriched_concat = season_enriched_concat.drop('Season', axis=1)\n    season_enriched_concatenation = season_enriched_concatenation.drop('season', axis=1)\n    season_enriched_concatenation = season_enriched_concatenation.drop('team_id', axis=1)\n    #Merge with last_10\n    season_enriched_concatenation = pd.merge(season_enriched_concatenation,last_10_df_sum,left_on=['Season','TeamID'], right_on=['Season','TeamID'], how='left')\n    #Merge with Names\n    season_enriched_concatenation = pd.merge(season_enriched_concatenation,Teams_name_df[['TeamID','TeamName']],left_on=['TeamID'], right_on=['TeamID'], how='left')\n    return season_enriched_concatenation, last_10_df_sum\n\nWseason_enriched_concat, Wlast_10_df_sum = get_season_data(Wseason_enriched,Welo_points,WTeams)\nseason_enriched_concat, Mlast_10_df_sum = get_season_data(season_enriched,elo_points,MTeams)","execution_count":null,"outputs":[]},{"metadata":{"id":"EA4xu83WR0cv","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# TOURNAMENT DATA\n# Adding Seed\ndef get_seed(x):\n        return int(x[1:3])\n    \ndef get_tournament_data(tourney_data,tourney_seed_data):\n    tourney_data = pd.merge(tourney_data, tourney_seed_data, left_on=['Season', 'WTeamID'], right_on=['Season', 'TeamID'], how='left')\n    tourney_data.rename(columns={'Seed':'WSeed'}, inplace=True)\n    tourney_data = tourney_data.drop('TeamID', axis=1)\n    tourney_data = pd.merge(tourney_data, tourney_seed_data, left_on=['Season', 'LTeamID'], right_on=['Season', 'TeamID'], how='left')\n    tourney_data.rename(columns={'Seed':'LSeed'}, inplace=True)\n    tourney_data = tourney_data.drop('TeamID', axis=1)    \n\n    tourney_data['WSeed'] = tourney_data['WSeed'].map(lambda x: get_seed(x))\n    tourney_data['LSeed'] = tourney_data['LSeed'].map(lambda x: get_seed(x))\n    tourney_data['Seed_diff'] = tourney_data['WSeed'] - tourney_data['LSeed']\n    Wins_tourney = tourney_data[['Season','WTeamID','DayNum']]\n    Wins_tourney.rename(columns={'WTeamID':'TeamID'}, inplace=True)\n    Wins_tourney['Result'] = 1\n    Loses_tourney = tourney_data[['Season','LTeamID','DayNum']]\n    Loses_tourney.rename(columns={'LTeamID':'TeamID'}, inplace=True)\n    Loses_tourney['Result'] = 0\n    tourney_results_concatenation = pd.concat((Wins_tourney,Loses_tourney)).reset_index(drop=True)\n    tourney_results_concatenation = tourney_results_concatenation[tourney_results_concatenation.Season>=2003]\n    \n    return tourney_data, tourney_results_concatenation\n\nWtourney_data, Wtourney_results_concat = get_tournament_data(Wtourney_result,Wtourney_seed)\nMtourney_data, tourney_results_concat = get_tournament_data(tourney_result,tourney_seed)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Do cinderellas go far?"},{"metadata":{},"cell_type":"markdown","source":"A question that arises when an underdog becomes a _cinderella_ is if it has just been a matter of luck or actually it is an underrated competitive team with high chances of making into the final rounds. To help clarify this doubt we can have a look at how many _cinderellas_ (teams that have beaten a team with 5 points less seeding) have made it to the Elite Eight or Final Four rounds in the past years."},{"metadata":{},"cell_type":"markdown","source":"### Men cinderellas"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"year_list = list(range(2003, 2020))\n#d = np.empty((2, 0)).tolist()\ndicts = {'year','count'}\nlist_of_tuples_four = []\nlist_of_tuples_eight = []\npd.DataFrame(columns = ['Season','count'])\nfor year in year_list:\n    upsets = upsets_calculator(year,Mtourney_data)\n    cinderellas = upsets.WTeamID.to_list()\n    conderellas_list = name_calculator(cinderellas,MTeams)\n    final_four_list = final_four(year,Mtourney_data,MTeams)\n    elite_eight_list = elite_eight(year,Mtourney_data,MTeams)\n    count_four = len(set(conderellas_list) & set(final_four_list))\n    count_eight = len(set(conderellas_list) & set(elite_eight_list))\n    pair_four = (year,count_four)\n    pair_eight = (year,count_eight)\n    list_of_tuples_four.append(pair_four)\n    list_of_tuples_eight.append(pair_eight)\n  #dicts[year] = count\n\ncinderellas_df_four = pd.DataFrame(data=list_of_tuples_four,columns = ['Season','cinderellas_count'])\ncinderellas_df_eight = pd.DataFrame(data=list_of_tuples_eight,columns = ['Season','cinderellas_count'])\n#set(names_lista).intersection(set(names_list))\n\n\nfig, ax = plt.subplots(1,2,figsize=(11,6))\nplot_eight = sns.barplot(x=\"Season\", y=\"cinderellas_count\", data=cinderellas_df_eight,ax=ax[0], color= 'blue')\nplot_four = sns.barplot(x=\"Season\", y=\"cinderellas_count\", data=cinderellas_df_four,ax=ax[1], color= 'blue')\naxes = plot_four.axes\nax[0].set_ylim(0,8)\nax[1].set_ylim(0,4)\nax[0].set_title('Number of cinderellas in Elite Eight')\nax[1].set_title('Number of cinderellas in Final Four')\nplot_four.set_xticklabels(plot_four.get_xticklabels(), rotation=45)\nax[0].set_ylabel('')\nax[1].set_xlabel\nplot_eight.set_xticklabels(plot_eight.get_xticklabels(), rotation=45)\nax[1].annotate('South Carolina', xy=(14, 1), xytext=(15, 2),\n            arrowprops=dict(facecolor='black', shrink=0.05),\n            )\nax[1].annotate('Loyola Chicago', xy=(15, 1), xytext=(17, 1.6),\n            arrowprops=dict(facecolor='black', shrink=0.05),\n            )\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There has been 16 _cinderellas_ in the Elite Eight rounds in the past 17 years and 10 _cinderellas_ in the Final Four rounds. So, on average there has almost been one _cinderella_ in Elite Eight round every year since 2003. That makes me ask: \"Are _cinderellas_ really _cinderellas_, or are those teams just underrated?\"\n\nThis drives me to the next study, that tries to understand which are the shared characteristics of Elite Eight teams throughout the different years."},{"metadata":{},"cell_type":"markdown","source":"### Women cinderellas"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"year_list = list(range(2003, 2020))\n#d = np.empty((2, 0)).tolist()\ndicts = {'year','count'}\nlist_of_tuples_four = []\nlist_of_tuples_eight = []\npd.DataFrame(columns = ['Season','count'])\nfor year in year_list:\n    upsets = upsets_calculator(year,Wtourney_data)\n    cinderellas = upsets.WTeamID.to_list()\n    conderellas_list = name_calculator(cinderellas,WTeams)\n    final_four_list = final_four(year,Wtourney_data,WTeams)\n    elite_eight_list = elite_eight(year,Wtourney_data,WTeams)\n    count_four = len(set(conderellas_list) & set(final_four_list))\n    count_eight = len(set(conderellas_list) & set(elite_eight_list))\n    pair_four = (year,count_four)\n    pair_eight = (year,count_eight)\n    list_of_tuples_four.append(pair_four)\n    list_of_tuples_eight.append(pair_eight)\n  #dicts[year] = count\n\ncinderellas_df_four = pd.DataFrame(data=list_of_tuples_four,columns = ['Season','cinderellas_count'])\ncinderellas_df_eight = pd.DataFrame(data=list_of_tuples_eight,columns = ['Season','cinderellas_count'])\n#set(names_lista).intersection(set(names_list))\n\n\nfig, ax = plt.subplots(1,2,figsize=(11,6))\nplot_eight = sns.barplot(x=\"Season\", y=\"cinderellas_count\", data=cinderellas_df_eight,ax=ax[0], color= 'blue')\nplot_four = sns.barplot(x=\"Season\", y=\"cinderellas_count\", data=cinderellas_df_four,ax=ax[1], color= 'blue')\naxes = plot_four.axes\nax[0].set_ylim(0,8)\nax[1].set_ylim(0,4)\nax[0].set_title('Number of cinderellas in Elite Eight')\nax[1].set_title('Number of cinderellas in Final Four')\nplot_four.set_xticklabels(plot_four.get_xticklabels(), rotation=45)\nax[0].set_ylabel('')\nax[1].set_xlabel\nplot_eight.set_xticklabels(plot_eight.get_xticklabels(), rotation=45)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Just 12 cinderellas entered the Elite Eight in Women tournaments since 2003 and just one achieved the Final Four. This can be explained by the inquality in Women's March Madness teams, in terms of final rounds it is less of a madness when compared with Men."},{"metadata":{"id":"DC1_jAFW8GaS"},"cell_type":"markdown","source":"# What do competitive teams have in common?"},{"metadata":{},"cell_type":"markdown","source":"To answer this question I have aggreagated the regular season data for each year and team by averaging the statistics for every game played during the regular season. Furthermore, I have included elo ratings [source](https://www.kaggle.com/lpkirwin/fivethirtyeight-elo-ratings) and advanced statistics [source](https://www.kaggle.com/lnatml/feature-engineering-with-advanced-stats) to make a more complete analysis."},{"metadata":{},"cell_type":"markdown","source":"Firstly, I want to know how does the last 10 games of the regular season record affect the performance of the teams during March Madness. One may think this could have a big impact in one team success in the final tournament."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Last 10 wins before tournament\nsum_tourney_wins = (tourney_results_concat.groupby(['Season', 'TeamID'])['Result']\n                     .sum()\n                     .rename('tourney_wins')\n                     .reset_index())\nsum_tourney_wins = pd.merge(sum_tourney_wins,Mlast_10_df_sum,left_on=['Season','TeamID'], right_on=['Season','TeamID'], how='left')\ncounts_foundation = (sum_tourney_wins.groupby(['L10wins'])['tourney_wins']\n                     .value_counts(normalize=True)\n                     .mul(100)\n                     .rename('pct of wins')\n                     .reset_index()\n                     .sort_values('L10wins'))\n\nsum_tourney_wins_divided = sum_tourney_wins.copy()\ndef made_final_four (row):\n    if row['tourney_wins'] >= 5:\n      return True\n    else:\n        return False\nsum_tourney_wins_divided['final_four'] = sum_tourney_wins_divided.apply (lambda row: made_final_four(row), axis=1)\n\ndef made_elite_eight (row):\n    if row['tourney_wins'] >= 4:\n      return True\n    else:\n        return False\nsum_tourney_wins_divided['elite_eight'] = sum_tourney_wins_divided.apply (lambda row: made_elite_eight(row), axis=1)\n\n# Plotting distribution differences Final four vs not final four based on last 10 reg season wins\ncolor_palette = ['green','red']\ni=0\nfor elite_eight_bool in [True,False]:\n    subset = sum_tourney_wins_divided[sum_tourney_wins_divided['elite_eight'] == elite_eight_bool]\n    sns.distplot(subset['L10wins'], hist=False, kde=True, color = color_palette[i], \n             #bins=int(180/5),\n             #hist_kws={'edgecolor':'black'},\n             kde_kws={'linewidth': 1, \"shade\": True, \"bw\":0.25},\n            label= str(elite_eight_bool))\n    i+=1\n# Put the legend out of the figure\nplt.title('Tournament performance based on last 10 Regular season games (2003-2019)')\nplt.ylabel('')\nplt.xlabel('Number of wins in last 10 regular season games')\nplt.annotate('Elite Eight teams more likely\\nto come from end of regular\\nseason winning records', xy=(9, 0.33), xytext=(11.8, 0.2),\n            arrowprops=dict(facecolor='black', shrink=0.025,width=1),\n            )\n\nplt.legend(title='Arrived to Elite Eight Round', bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the above figure we can observe that those teams which get to the quarter finals tend to come from high winning records in the last 10 games."},{"metadata":{"id":"q8IUZqvKTflY","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# Last 10 wins before tournament\nsum_tourney_wins = (tourney_results_concat.groupby(['Season', 'TeamID'])['Result']\n                     .sum()\n                     .rename('tourney_wins')\n                     .reset_index())\nsum_tourney_wins = pd.merge(sum_tourney_wins,Mlast_10_df_sum,left_on=['Season','TeamID'], right_on=['Season','TeamID'], how='left')\ncounts_foundation = (sum_tourney_wins.groupby(['L10wins'])['tourney_wins']\n                     .value_counts(normalize=True)\n                     .mul(100)\n                     .rename('pct of wins')\n                     .reset_index()\n                     .sort_values('L10wins'))\nsns.barplot(x=\"L10wins\", y=\"pct of wins\", hue=\"tourney_wins\", data=counts_foundation)\n# Put the legend out of the figure\nplt.title('Percentage of wins in tournament based on last 10 games performance (2003-2019)')\nplt.ylabel('')\nplt.xlabel('Number of wins in last 10 regular season games')\nplt.annotate('50% of teams with 10 wins streak lose in first round', xy=(7.55, 52.5), xytext=(10, 33),\n            arrowprops=dict(facecolor='black', shrink=0.05),\n            )\nplt.legend(title='Number of wins in tournament', bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"However, the average number of wins in the last 10 regular season games is 7.32 for teams entering March Madness. This suggests that teams making it to the NCCA final tournament are teams that have a good record at the end of the season. Below we can observe the distribution differences between teams coming from 7 (average March Madness team) wins out of last 10 vs teams with 10 wins streak. Now we can see that the difference is not significant, in fact approximately the same number of teams fail to even win one game."},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":false},"cell_type":"code","source":"# Last 10 games, 7 wins vs 10 wins (small difference!!), average NCAA teams has 7\nfor last_number_wins in [7,10]:\n    subset = sum_tourney_wins[sum_tourney_wins['L10wins'] == last_number_wins]\n    sns.distplot(subset['tourney_wins'], hist=False, kde=True, \n             #bins=int(180/5), #color = 'darkblue',\n             #hist_kws={'edgecolor':'black'},\n             kde_kws={'linewidth': 1, \"shade\": True, \"bw\":0.25},\n            label= str(last_number_wins) + ' wins')\n# Put the legend out of the figure\nplt.title('Wins in tournament distribution based on last 10 games performance (2003-2019)')\nplt.ylabel('')\nplt.xlabel('Tournament wins')\n\nplt.legend(title='Number of wins in last 10 regular season games', bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"h4\">A strong predictor</div>"},{"metadata":{},"cell_type":"markdown","source":"Elo rating is a method for calculating the relative skill levels of players or teams in zero-sum games, originally used in chess. This feature brings information about the current state of a single team. Every team in the competition starts with an Elo rating of 1500 points and, each game its value gets updated based on game scores and home advantage factor as follows:\n   "},{"metadata":{},"cell_type":"markdown","source":"<script src=\"//yihui.org/js/math-code.js\"></script>\n<!-- Just one possible MathJax CDN below. You may use others. -->\n<script async\n  src=\"//mathjax.rstudio.com/latest/MathJax.js?config=TeX-MML-AM_CHTML\">\n</script>\n\n$$update = K * \\frac{margin+3}{7.5 + 0.06 * (elo_{winner_{t-1}}-elo_{loser_{t-1}})}*(1-\\frac{1}{10-\\frac{(-elo_{winner_{t-1}}-elo_{loser_{t-1}})}{400}+1})$$\n\n$$elo_{winner_{t}} = elo_{winner_{t-1}} + update$$\n\n$$elo_{loser_{t}} = elo_{loser_{t-1}} + update$$"},{"metadata":{},"cell_type":"markdown","source":"Where margin represents the game score difference, $elo_{winner_{t-1}}$ and $elo_{loser_{t-1}}$ the elo points before the game adjusted for home advantage (home team receives 100 Elo points, while the visitor gets substracted 100) and $K$ the update factor, in charge of adjusting the speed at which Elo ratings are updated from game to game (K = 20 for NBA).\n\nHere we can see Virginia vs Duke Elo points since 1985:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"virginia_elo = elo_points[elo_points.team_id==1438]#id_calculator('Virginia',MTeams)\nduke_elo = elo_points[elo_points.team_id==id_calculator('Duke',MTeams)]\n\nplt.plot(virginia_elo.season,virginia_elo.season_elo,label= 'Virginia')\nplt.plot(duke_elo.season,duke_elo.season_elo,label= 'Duke')\nplt.title(\"Elo rating history\")\nplt.legend(title='Team', bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"To better understand how Elo ratings can correlate to a good performance in March Madness I plot the distributions of Elo rating for Elite Eight teams and the rest of the teams:"},{"metadata":{"_kg_hide-output":false,"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"#Elo points with tourney results (Final_four boolean) merge\ntourney_elo = pd.merge(sum_tourney_wins_divided,elo_points,left_on=['Season','TeamID'],right_on=['season', 'team_id'],how='left')\ntourney_elo = tourney_elo.drop(['tourney_wins','L10wins','team_id','season'],axis=1)\n\n#Plotting distribution difference between Elo points\nrcParams['figure.figsize'] = 5,5\n\ncolor_palette = ['green','red']\n##### Elite Eight\ni=0\nfor elite_eight_bool in [True,False]:\n    subset = tourney_elo[tourney_elo['elite_eight'] == elite_eight_bool]\n    sns.distplot(subset['season_elo'], hist=False, kde=True, color = color_palette[i], \n             #bins=int(180/5),\n             #hist_kws={'edgecolor':'black'},\n             kde_kws={'linewidth': 1, \"shade\": True, \"bw\":10},\n            label= str(elite_eight_bool))\n    i+=1\n# Put the legend out of the figure\nplt.title('Elo points difference of Final Four vs Not Elite Eight teams  (2003-2019)')\nplt.ylabel('')\nplt.xlabel('Elo points')\n#plt.annotate('Final Four teams more likely\\nto come from regular season\\nwinning streaks', xy=(9, 0.4), xytext=(11.8, 0.2),\n#            arrowprops=dict(facecolor='black', shrink=0.025,width=1),\n#            )\nplt.legend(title='Arrived to Elite Eight', bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.);\n\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This graph suggests that it could be a strong indicator of the competitivenes of teams since there is a clear difference in the distributions."},{"metadata":{"id":"PPY7CiuGIVvu","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"def pca_season(df,season_year,features,n_components):\n    #df = df[df['Season']==season_year].reset_index()\n    # Separating out the features\n    x = df.loc[:, features].values\n    # Separating out the target\n    #y = df.loc[:,['target']].values\n    # Standardizing the features\n    x = StandardScaler().fit_transform(x)\n\n    pca = PCA(n_components=n_components)\n    pca.fit(x)\n    principalComponents = pca.transform(x)\n    if n_components == 2:\n        principalComponentsvalues = pd.DataFrame(data = principalComponents\n                    , columns = ['principal component 1', 'principal component 2'])\n    elif n_components == 3:\n        principalComponentsvalues = pd.DataFrame(data = principalComponents\n                    , columns = ['principal component 1', 'principal component 2', 'principal component 3'])\n    return pca, principalComponentsvalues","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Next, I show a dimensionality reduction method applied to the aggregated data for every team and season. Principal Complonent Analysis projects the original cloud of data into othogonal axis by maintining as much information as possible. In this case, I take the aggregated data with several factors and apply PCA. In the figures below the first two principal components are shown, which retain more than 70\\% of the inital information. This allows us to observe which are the differences between teams in terms of regular season performance. If the factors used are able to determine similarities between team we will expect similar teams to lay closely in the 2D representation."},{"metadata":{},"cell_type":"markdown","source":"<script src=\"//yihui.org/js/math-code.js\"></script>\n<!-- Just one possible MathJax CDN below. You may use others. -->\n<script async\n  src=\"//mathjax.rstudio.com/latest/MathJax.js?config=TeX-MML-AM_CHTML\">\n</script>\n\nThe used variables are:\n\n- NetRtg: shows a balanced indicator of a team offensive and defensive capabilities. It is calculated as the difference between scored and allowed points per 100 possesions:\n\n$$NetRtg = OffRtg - DeffRtg = 100 \\cdot \\frac{Points Made}{Poss} - 100 \\cdot \\frac{Points Allowed}{Poss}$$\n\n$\\text{Where } Poss = 0.96 \\cdot (FGA + TOV + 0.44 \\cdot FTA -OR)$ aproximates the number of possesions during a game.\n\n- FTAR: the team's ability of drawing fouls, proportion of shots made in the free line:\n\n$$FTAR = \\frac{FTA}{FGA}$$\n\n- eFG: effective Field Goal takes into account that 3-pointers are worth 1.5 more than 2-pointers,\n\n$$eFG = FGM + \\frac{0.5 \\cdot 3PM}{FGA}$$\n\n- PIE: Team (or Player) Impact Estimate groups several indicators into a single metric that represent how 'impactful' has the team been:\n    \n    $$PIE = PTS + FGM + FTM - FGA - FTA + DR + 0.5*OR + AST + STL + 0.5*BLK - PF - TOV$$\n    \nAlso end of the season Elo points (season_elo) and the number of wins during the last 10 regular season games (L10wins). Note that since this an 'objective' summary of each team I haven't included the experts pre-tournament seeding."},{"metadata":{},"cell_type":"markdown","source":"Each data point showed in the following plots represents the projection on the first to principal components plane of a TeamID in a Season with above explained features:"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"#Chossing features in PCA\nseason=2017\nstyle_metrics = ['Pos','AstR','Stl','ORP','TOR','TSP']\n#style_metrics = ['Pos','TOR','AstR','TSP']\nperformance_metrics = ['NetRtg','L10wins','FTAR','PIE','eFGP']\n#random_metrics = ['TOR','Stl'] + other_metrics\nperformance_metrics = performance_metrics + ['season_elo']\nfeatures = performance_metrics\n#####","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":false,"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"display(season_enriched_concat[season_enriched_concat.Season==2017][['Season','TeamID']+features].sample(1))","execution_count":null,"outputs":[]},{"metadata":{"id":"wECwBlORlINc","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"from matplotlib.patches import Ellipse\n\ndef PCA_matchup(teams_list, df,features,season_year,every_team,Teams,draw_ellipse=(0,0)):\n    #names_list can be either names or ids\n    df = df[df['Season']==season_year].reset_index()\n    \n    #To do PCA of only teams in tournament\n    if every_team == False:\n        df = df[df['March_mad_team']==0].reset_index()\n    labels_march = df['March_mad_team'].to_numpy()\n    #df = df.drop(['March_mad_team'],axis=1).reset_index()\n\n    pca, principalComponents_df = pca_season(df,season_year,features,2)\n    \n\n    team_id_list = []\n    team_name_list = []\n    if isinstance(teams_list[0], str):\n        for name in teams_list:\n            team_id = Teams[(Teams.TeamName.str.startswith(name))]['TeamID'].to_numpy()[0]\n            team_id_list.append(team_id)\n            team_name = Teams[(Teams.TeamName.str.startswith(name))]['TeamName'].to_numpy()[0]\n            team_name_list.append(team_name)\n    else:\n        team_id_list = teams_list.copy()\n        team_name_list = name_calculator(team_id_list,Teams)\n    # Index in dataframe\n    reset_indeces = []\n    for id in team_id_list:\n        index = df.index[df['TeamID']==id][0]\n        reset_indeces.append(index)\n\n    # Data values from PCA (projection on axes)\n    pca_data_projections = []\n    for index in reset_indeces:\n        projection = principalComponents_df.iloc[index]\n        pca_data_projections.append(projection)\n\n    kmeans = KMeans(n_clusters=5, random_state=0)\n    kmeans.fit(principalComponents_df.values)\n    centroids = kmeans.cluster_centers_\n    labels_kmeans = kmeans.labels_\n\n    fig = plt.figure(figsize = (5,5))\n    ax = fig.add_subplot(1,1,1)\n    \n    if draw_ellipse == (0,0):\n        pass\n    else:\n        e = Ellipse(xy=draw_ellipse, width=1, height=1,facecolor='none',edgecolor='red')\n        ax.add_artist(e)\n    \n    ax.scatter(centroids[:, 0], centroids[:, 1],\n            marker='', s=169, linewidths=3,\n            color='r', zorder=10)\n    \n\n    #cmap, norm = mpl.colors.from_levels_and_colors([0, 1], ['blue', 'black'])\n    ax.set_xlabel('PC 1 ' + str(round(pca.explained_variance_ratio_[0]*100,2)) + '%', fontsize = 15)\n    ax.set_ylabel('PC 2 '+ str(round(pca.explained_variance_ratio_[1]*100,2)) + '%', fontsize = 15)\n    ax.set_title(str(season_year) + ' Season PCA', fontsize = 20)\n    ax.scatter(principalComponents_df.loc[:,'principal component 1'],\n            principalComponents_df.loc[:,'principal component 2'],\n             c=labels_march.astype(np.float), cmap='coolwarm'#, norm=norm\n             )\n    coeff = pca.components_\n    n = coeff.shape[1]\n    texts = []\n    #for i in range(n):\n        #plt.arrow(0, 0, coeff[0,i]*3, coeff[1,i]*3,color = 'r',alpha = 0.5)\n        #plt.text(coeff[0,i]* 3.15, coeff[1,i] * 3.15, features[i], color = 'r', ha = 'center', va = 'center',fontsize = 13)\n        #adjust_text(texts, only_move='y', arrowprops=dict(arrowstyle=\"->\", color='r', lw=0.5))\n    #if isinstance(names_list[0], str):\n    for i, data_point in enumerate(pca_data_projections):\n        ax.annotate(team_name_list[i],data_point,\n                    xytext = data_point + 0.5,\n                    arrowprops=dict(arrowstyle=\"->\",\n                                    connectionstyle=\"arc3\"))\n    return ax, coeff, principalComponents_df\n  #ax.grid()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"Auburn_road = ['Auburn', 'New Mexico St', 'Kansas' ,'North Car', 'Kentucky', 'Virginia']\npair_analysis = ['Auburn', 'Kentucky']\nupsets = upsets_calculator(season,Mtourney_data)\ncinderellas = upsets.WTeamID.to_list()\n#names_list = winners + pair_analysis\n#names_list = name_calculator(cinderellas)\n#names_list = final_four(season,tourney_result)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Men's PCA"},{"metadata":{},"cell_type":"markdown","source":"I start by applying this method to the regular season data from 2017 (Elite Eight teams are marked with arrows). In the first plot we can observe that teams which classified to March Madness (blue) are clustered towards the right. That demostrates that out of 351 teams, the 64 selected to participate share some common facors.\nAfter applying PCA just taking into account the March Madness teams we can observe that South Carolina and Xavier seem to fall apart from the rest of the teams. Indeed, South Carolina (7) and Xavier (11), as I mentioned [earlier](#Do-cinderellas-go-far?) were the two cinderellas in Elite Eight that year."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true},"cell_type":"code","source":"season = 2017\nmmt_list = march_madness_teams(season,tourney_result)\nmmt_df = pd.DataFrame(columns = ['TeamID','March_mad_team'])\nmmt_df['TeamID'] = mmt_list\nmmt_df['March_mad_team'] = 0.0\nseason_enriched_mmt_list = pd.merge(mmt_df, season_enriched_concat[season_enriched_concat.Season==season],left_on=[\"TeamID\"],right_on=\"TeamID\", how=\"outer\")\nseason_enriched_mmt_list.fillna(1, inplace=True)\n\nnames_list = elite_eight(season,tourney_result,MTeams)\n#names_list=cinderellas\n\n\n\nplot_final_four1, pca_coefficients1, pca_df1= PCA_matchup(names_list, season_enriched_mmt_list, features, season,True,MTeams,draw_ellipse=(0,0))\nplot_final_four2, pca_coefficients2, pca_df2= PCA_matchup(names_list, season_enriched_mmt_list, features, season,False,MTeams,draw_ellipse=(-2,-0.827))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"If we see the correlation of the Principal Components with original variables we can conclude what makes South Carolina and Xavier different from the rest of the Elite Eight participants. The Ox axis (PC1) is higly correlated with NetRtg and PIE, so since Xavier and South Carolina are projected in the left side of the plot they have low values of NetRtg and PIE. They are also projected towards the bottom (PC2), which represents low winning records at end of regular season and poor ability of drawing fouls."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(7,3))\nplt.matshow(pca_coefficients2,cmap='viridis',fignum=1)\nplt.yticks([0,1],['1st Comp','2nd Comp'],fontsize=10)\nplt.colorbar()\nplt.xticks(range(len(features)),features,rotation=65,ha='left')\nax = plt.gca()\ne = Ellipse(xy=(0,0), width=1, height=1,facecolor='none',edgecolor='red')\nax.add_artist(e)\ne = Ellipse(xy=(3,0), width=1, height=1,facecolor='none',edgecolor='red')\nax.add_artist(e)\ne = Ellipse(xy=(1,1), width=1, height=1,facecolor='none',edgecolor='red')\nax.add_artist(e)\ne = Ellipse(xy=(2,1), width=1, height=1,facecolor='none',edgecolor='red')\nax.add_artist(e)\n#plt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In 2018 season we can Loyola-Chicago stays far the rest of the teams."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true},"cell_type":"code","source":"season = 2018\nmmt_list = march_madness_teams(season,tourney_result)\nmmt_df = pd.DataFrame(columns = ['TeamID','March_mad_team'])\nmmt_df['TeamID'] = mmt_list\nmmt_df['March_mad_team'] = 0.0\nseason_enriched_mmt_list = pd.merge(mmt_df, season_enriched_concat[season_enriched_concat.Season==season],left_on=[\"TeamID\"],right_on=\"TeamID\", how=\"outer\")\nseason_enriched_mmt_list.fillna(1, inplace=True)\n\nnames_list = elite_eight(season,tourney_result,MTeams)\n#names_list=cinderellas\n\n\n\n#plot_final_four1, pca_coefficients1, pca_df1= PCA_matchup(names_list, season_enriched_mmt_list, features, season,True,draw_ellipse=(0,0))\nplot_final_four3, pca_coefficients3, pca_df3= PCA_matchup(names_list, season_enriched_mmt_list, features, season,False,MTeams,draw_ellipse=(1.84,-1.8))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In this year, Elo points seems to be the major factor determining the success of teams."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(7,3))\nplt.matshow(pca_coefficients3,cmap='viridis',fignum=1)\nplt.yticks([0,1],['1st Comp','2nd Comp'],fontsize=10)\nplt.colorbar()\nplt.xticks(range(len(features)),features,rotation=65,ha='left')\n#plt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the case of 2019, every Elite Eight seem to fall closely. That year, no cinderella got to that round as we will have seen in the previous section."},{"metadata":{"id":"jPBRoSwSpQ1B","outputId":"4524c389-8053-4066-d7c5-327d3446e92e","trusted":true,"_kg_hide-input":true,"_kg_hide-output":false},"cell_type":"code","source":"season = 2019\nmmt_list = march_madness_teams(season,tourney_result)\nmmt_df = pd.DataFrame(columns = ['TeamID','March_mad_team'])\nmmt_df['TeamID'] = mmt_list\nmmt_df['March_mad_team'] = 0.0\nseason_enriched_mmt_list = pd.merge(mmt_df, season_enriched_concat[season_enriched_concat.Season==season],left_on=[\"TeamID\"],right_on=\"TeamID\", how=\"outer\")\nseason_enriched_mmt_list.fillna(1, inplace=True)\n\nnames_list = elite_eight(season,tourney_result,MTeams)\n#names_list=cinderellas\n\n#plot_final_four1, pca_coefficients1, pca_df1= PCA_matchup(names_list, season_enriched_mmt_list, features, season,True)\n\nplot_final_four2, pca_coefficients2, pca_df2= PCA_matchup(names_list, season_enriched_mmt_list, features, season,False,MTeams,draw_ellipse=(0,0))\n#plt.gca().invert_xaxis()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"If we plot the 2019 Final Four teams, even smaller is the distance between those teams."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true},"cell_type":"code","source":"season = 2019\nmmt_list = march_madness_teams(season,tourney_result)\nmmt_df = pd.DataFrame(columns = ['TeamID','March_mad_team'])\nmmt_df['TeamID'] = mmt_list\nmmt_df['March_mad_team'] = 0.0\nseason_enriched_mmt_list = pd.merge(mmt_df, season_enriched_concat[season_enriched_concat.Season==season],left_on=[\"TeamID\"],right_on=\"TeamID\", how=\"outer\")\nseason_enriched_mmt_list.fillna(1, inplace=True)\n\nnames_list = final_four(season,tourney_result,MTeams)\n#names_list=cinderellas\n\n#plot_final_four1, pca_coefficients1, pca_df1= PCA_matchup(names_list, season_enriched_mmt_list, features, season,True)\n\nplot_final_four2, pca_coefficients2, pca_df2= PCA_matchup(names_list, season_enriched_mmt_list, features, season,False,MTeams,draw_ellipse=(0,0))\n#plt.gca().invert_xaxis()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Women's PCA"},{"metadata":{},"cell_type":"markdown","source":"Women data shows similar patterns, in 2017 for example Oregon was the only team to get to the Elite Eight round and PCA shows that team's projection is separate towards the left (weak) part of the plot, while Connecticut, clear winner that season lays at the right end of the plot."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"season = 2017\nmmt_list = march_madness_teams(season,Wtourney_result)\nmmt_df = pd.DataFrame(columns = ['TeamID','March_mad_team'])\nmmt_df['TeamID'] = mmt_list\nmmt_df['March_mad_team'] = 0.0\nseason_enriched_mmt_list = pd.merge(mmt_df, Wseason_enriched_concat[Wseason_enriched_concat.Season==season],left_on=[\"TeamID\"],right_on=\"TeamID\", how=\"outer\")\nseason_enriched_mmt_list.fillna(1, inplace=True)\n\nnames_list = elite_eight(season,Wtourney_result,WTeams)\n#names_list=cinderellas\n\n\nplot_final_four1, pca_coefficients1, pca_df1= PCA_matchup(names_list, season_enriched_mmt_list, features, season,True,WTeams,draw_ellipse=(0,0))\nplot_final_four2, pca_coefficients2, pca_df2= PCA_matchup(names_list, season_enriched_mmt_list, features, season,False,WTeams,draw_ellipse=(-1.15,0.1))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"With the previous representations we have seen that competitive teams have common charcateristics that can be visualized by the PCA. This give us information about the ability of these variables to measure the competitiveness of teams."},{"metadata":{},"cell_type":"markdown","source":"# Predictive Model"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"location_ncaa_games = pd.merge(game_cities,city_names,left_on=['CityID'],right_on=['CityID'],how='left')\nlocation_ncaa_games = location_ncaa_games[location_ncaa_games.CRType=='NCAA']\nlocation_ncaa_games_extended = pd.merge(location_ncaa_games,location_data[['CityID','lat','lng']],left_on=['CityID'],right_on=['CityID'],how='left').drop(['City',\n'State'],axis = 1)\nlocation_ncaa_games_extended.rename(columns={'lat':'game_lat'}, inplace=True)\nlocation_ncaa_games_extended.rename(columns={'lng':'game_lng'}, inplace=True)\nlocation_ncaa_games_winner = pd.merge(location_ncaa_games_extended,location_data[['TeamID','lat','lng']],left_on=['WTeamID'],right_on=['TeamID'],how='left').drop(['TeamID'],axis=1)\nlocation_ncaa_games_winner.rename(columns={'lat':'W_lat'}, inplace=True)\nlocation_ncaa_games_winner.rename(columns={'lng':'W_lng'}, inplace=True)\nlocation_ncaa_games_winner_loser = pd.merge(location_ncaa_games_winner,location_data[['TeamID','lat','lng']],left_on=['LTeamID'],right_on=['TeamID'],how='left').drop(['TeamID'],axis=1)\nlocation_ncaa_games_winner_loser.rename(columns={'lat':'L_lat'}, inplace=True)\nlocation_ncaa_games_winner_loser.rename(columns={'lng':'L_lng'}, inplace=True)\nloc_df = location_ncaa_games_winner_loser\nloc_df['W_dist'] = np.vectorize(distance)(loc_df['game_lat'], loc_df['W_lat'], loc_df['game_lng'], loc_df['W_lng'])\nloc_df['L_dist'] = np.vectorize(distance)(loc_df['game_lat'], loc_df['L_lat'], loc_df['game_lng'], loc_df['L_lng'])\nloc_df = loc_df.drop(['W_lat','W_lng','L_lat','L_lng','CRType','game_lat','game_lng'],axis=1)\nloc_df['Dist_diff'] = loc_df['L_dist']-loc_df['W_dist']\n#loc_df = loc_df.drop(['L_dist','W_dist'],axis=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I trained a XGB model with men's season data from 2003 to 2019 with every variable from the aggregated dataset:"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true},"cell_type":"code","source":"#tourney_result = pd.read_csv('../input/march-madness-analytics-2020/MDataFiles_Stage2/MNCAATourneyCompactResults.csv')\ntourney_result_train = Mtourney_data[Mtourney_data.Season>=2003]\nseason_enriched_concat_train = season_enriched_concat[season_enriched_concat.Season<2020]\n\n# Winning side\nWtrain_df = pd.merge(tourney_result_train, season_enriched_concat_train, left_on=['Season', 'WTeamID'], right_on=['Season', 'TeamID'], how='left')\nWtrain_df = Wtrain_df.drop(['WScore','LScore','WLoc','NumOT','DayNum','TeamName','TeamID','Seed_diff','LSeed'],axis=1)\n# We keep LTeamID to merge in the future\nWtrain_df.rename(columns={\"WTeamID\": \"TeamID\"},inplace=True)\nWtrain_df.rename(columns={\"WSeed\": \"Seed\"},inplace=True)\ntrain_df_2 = Wtrain_df.copy()\n\nfor column in Wtrain_df.columns[1:]:\n        Wtrain_df.rename(columns={column: column + \"_1\"},inplace=True)\n\n## Losing side\nLtrain_df = pd.merge(tourney_result_train, season_enriched_concat_train, left_on=['Season', 'LTeamID'], right_on=['Season', 'TeamID'], how='left')\nLtrain_df = Ltrain_df.drop(['WScore','LScore','WLoc','NumOT','DayNum','TeamName','TeamID','WTeamID','Seed_diff','WSeed'],axis=1)\nLtrain_df.rename(columns={\"LTeamID\": \"TeamID\"},inplace=True)\nLtrain_df.rename(columns={\"LSeed\": \"Seed\"},inplace=True)\ntrain_df_3 = Ltrain_df.copy()\n\nfor column in Ltrain_df.columns[1:]:\n    if \"1\" in column[-2:]:\n        pass\n    else:\n        Ltrain_df.rename(columns={column: column + \"_2\"},inplace=True)\n        \nupper_train = pd.merge(Wtrain_df,Ltrain_df, left_on=['Season', 'LTeamID_1'], right_on=['Season', 'TeamID_2'], how='left')\n#train_df['result'] = 1\nupper_train = upper_train.drop(['LTeamID_1'],axis=1)\nupper_train['result'] = 1\n\n# Adding distance upper\n# comment if no distance used\nupper_train = pd.merge(upper_train,loc_df,left_on=['Season','TeamID_1','TeamID_2'],right_on=['Season','WTeamID','LTeamID'],how=\"left\")\nupper_train = upper_train.drop(['DayNum','WTeamID','LTeamID','CityID','CityID', 'Dist_diff'],axis=1)\nupper_train.rename(columns={\"W_dist\": \"Dist_1\"},inplace=True)\nupper_train.rename(columns={\"L_dist\": \"Dist_2\"},inplace=True)\n\n# Duplicate upper_train and reverse it to add result = 0\nfor column in train_df_2.columns[1:]:\n    train_df_2.rename(columns={column: column + \"_2\"},inplace=True)\n\nfor column in train_df_3.columns[1:]:\n    train_df_3.rename(columns={column: column + \"_1\"},inplace=True)\ntrain_df_2.head()\n\nlower_train = pd.merge(train_df_3,train_df_2, left_on=['Season', 'TeamID_1'], right_on=['Season', 'LTeamID_2'], how='left')\nlower_train = lower_train.drop(['LTeamID_2'],axis=1)\nlower_train['result'] = 0\n\n# Adding distance\n# comment if no distance used\nlower_train = pd.merge(lower_train,loc_df,left_on=['Season','TeamID_2','TeamID_1'],right_on=['Season','WTeamID','LTeamID'],how=\"left\")\nlower_train = lower_train.drop(['DayNum','WTeamID','LTeamID','CityID','CityID', 'Dist_diff'],axis=1)\nlower_train.rename(columns={\"W_dist\": \"Dist_2\"},inplace=True)\nlower_train.rename(columns={\"L_dist\": \"Dist_1\"},inplace=True)\n\ntrain = pd.concat((upper_train,lower_train)).reset_index(drop=True)\ntrain['season_elo_1'] = train['season_elo_1'].astype(float)\ntrain['season_elo_2'] = train['season_elo_2'].astype(float)\n#train['L10wins_2'].describe()\ntest = train.copy()\ntest = test[test.Season==2019]\n\n# Following line just for distance\n#train = train[train.Season>=2010]\n\ntrain = train[train.Season!=2019]\n#display(train.head(5))\n\n\n\ntrain.pop('Dist_1')\ntrain.pop('Dist_2')\ny = train.pop('result')\ntrain = train.drop(['Season','TeamID_1','TeamID_2'],axis=1)\nX_initial = train\n\n\nfinal_cols = train.columns\nX = X_initial[final_cols].copy()\n\ndisplay(X.head())\n\ny_test = test.pop('result')\nX_test = test[final_cols].copy()\n\n\nxgbdata = xgb.DMatrix(data=X,label=y.to_numpy())\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":" I use this approach to see what variables does the model consider as important. To do so I get the [SHAP values](https://towardsdatascience.com/explain-your-model-with-the-shap-values-bc36aac4de3d) of the model, which outputs:"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import log_loss\nimport xgboost as xgb\nfrom xgboost import plot_tree\nimport shap\n\nparams = {'objective': 'binary:logistic',\n          'eval_metric': 'logloss',\n          'eta': 0.03,\n          'max_depth': 3\n          }\n\ncv_scores = xgb.cv(dtrain=xgbdata,\n                    params=params,\n                    nfold=5,\n                    num_boost_round=100,\n                    early_stopping_rounds=60,\n                    verbose_eval=False,\n                    #as_pandas=True,\n                    seed=135\n                    )\nbest_score = cv_scores['test-logloss-mean'].min()\nbest_round = cv_scores['test-logloss-mean'].idxmin()\nprint(f'Model Score: {best_score:.2f} log loss after {best_round} iterations.')\n\n#preds = model_final.predict(dval, ntree_limit=cv_scores.best_ntree_limit)\n#oof_preds[val_idx] = preds\n#score = log_loss(y, oof_preds)\n","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"params = {'objective': 'binary:logistic',\n          'eval_metric': 'logloss',\n          'eta': 0.03,\n          'max_depth': 3\n          }\n\nskf = StratifiedKFold(n_splits=10, random_state=2233)\noof_preds = np.zeros((len(train)))\nfor train_idx, val_idx in skf.split(X, y):\n    X_train, X_val = X.loc[train_idx], X.loc[val_idx]\n    y_train, y_val = y.to_numpy()[train_idx], y.to_numpy()[val_idx]\n    dtrain = xgb.DMatrix(data=X_train, label=y_train)\n    dval = xgb.DMatrix(data=X_val, label=y_val)\n    watchlist = [(dtrain, 'train'), (dval, 'eval')]\n    model_final = xgb.train(params,\n                            dtrain,\n                            num_boost_round=100,\n                            evals=watchlist,\n                            early_stopping_rounds=60,\n                            verbose_eval=False\n                            )\n    preds = model_final.predict(dval, ntree_limit=model_final.best_ntree_limit)\n    oof_preds[val_idx] = preds\nscore = log_loss(y, oof_preds)\n#print(f'Optimized Model Score: {score:.2f} log loss.')\n\n## Test data (one season tournament)\ntest_xgb_data = xgb.DMatrix(data=X_test, label=y_test)\npreds = model_final.predict(test_xgb_data, ntree_limit=model_final.best_ntree_limit)\n#preds_test = np.zeros((len(X_test)))\n#preds_test[] = preds\nscore = log_loss(y_test, preds)\nprint(f'Optimized Model Score: {score:.2f} log loss.')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"explainer = shap.TreeExplainer(model_final, data=X,\n                               model_output='probability')\nshap_values_final = explainer.shap_values(X)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"display(shap.summary_plot(shap_values_final, X,plot_type=\"bar\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The model agrees with the PCA analysis, Elo rating points, NetRG and FTAR are detected as the most important features together with the human seeding. After iterating once again but with the 8 most important features, the model reduces the logloss score. However we have seen through PCA that these factors alone are good measuring competitiviness but not enough to detect _cinderellas_. Then, how can upsets be predicted?\n"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"X_initial = train\n\n\nfinal_cols = ['NetRtg_1', 'FTAR_1',\n                'season_elo_1', 'Seed_1',\n              'NetRtg_2', 'FTAR_2',\n                'season_elo_2', 'Seed_2']\n\n#final_cols = train.columns\nX = X_initial[final_cols].copy()\nX_test = test[final_cols].copy()\n\nparams = {'objective': 'binary:logistic',\n          'eval_metric': 'logloss',\n          'eta': 0.03,\n          'max_depth': 3\n          }\n\nskf = StratifiedKFold(n_splits=10, random_state=2233)\noof_preds = np.zeros((len(train)))\nfor train_idx, val_idx in skf.split(X, y):\n    X_train, X_val = X.loc[train_idx], X.loc[val_idx]\n    y_train, y_val = y.to_numpy()[train_idx], y.to_numpy()[val_idx]\n    dtrain = xgb.DMatrix(data=X_train, label=y_train)\n    dval = xgb.DMatrix(data=X_val, label=y_val)\n    watchlist = [(dtrain, 'train'), (dval, 'eval')]\n    model_final = xgb.train(params,\n                            dtrain,\n                            num_boost_round=100,\n                            evals=watchlist,\n                            early_stopping_rounds=60,\n                            verbose_eval=False\n                            )\n    preds = model_final.predict(dval, ntree_limit=model_final.best_ntree_limit)\n    oof_preds[val_idx] = preds\nscore = log_loss(y, oof_preds)\n#print(f'Optimized Model Score: {score:.2f} log loss.')\n\n## Test data (one season tournament)\ntest_xgb_data = xgb.DMatrix(data=X_test, label=y_test)\npreds = model_final.predict(test_xgb_data, ntree_limit=model_final.best_ntree_limit)\n#preds_test = np.zeros((len(X_test)))\n#preds_test[] = preds\nscore = log_loss(y_test, preds)\nprint(f'Optimized Model Score: {score:.2f} log loss.')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"#explainer = shap.TreeExplainer(model_final, data=X,\n#                               model_output='probability')\n#shap_values_final = explainer.shap_values(X)\n#display(shap.summary_plot(shap_values_final, X,plot_type=\"bar\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Where do predictive models fail?"},{"metadata":{},"cell_type":"markdown","source":"<div class=\"h4\">Loyola Chicago 2018 case study</div>"},{"metadata":{},"cell_type":"markdown","source":"Loyola Chicago was seeded 11 in the 2018 tournament, however they managed to get through the South Regional conference bracket reaching the Final Four participation. If you watch this team games one can observe a play that is repeated over and over again. In the offensive end one of the centers does a screen in the top of the 3-pointer line (see below footage). A pick&roll, pick&pop or penetration opportunities arise from there, giving Loyola the opportunity to score an easy layout/dunk under the basket or a 3-point shot. Seems that rivals struggled to deffend that play, which could be a key of the succesful tournament run."},{"metadata":{},"cell_type":"markdown","source":"<img src=\"https://github.com//javiferran/ncaa-analytics-2020/blob/master/loyola_2018.gif?raw=true[](http://)\" alt=\"zeke\" height=\"400\" width=\"720\"/>"},{"metadata":{},"cell_type":"markdown","source":"I analysed the play-by-play dataset to obtain the average number of dunks or layouts per game and 3-pointers made during the regular season. The Loyola results are:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"Loyola_below_basket = MEvents_2018[(MEvents_2018.EventTeamID==id_calculator('Loyola-Chicago',MTeams)) & ((MEvents_2018.EventType=='made3') | ((MEvents_2018.EventType=='made2') & ((MEvents_2018.EventSubType=='lay') | (MEvents_2018.EventSubType=='dunk'))))]\n# made 2\nLoyola_below_basket_agg = (Loyola_below_basket[Loyola_below_basket.EventType=='made2'].groupby(['Season', 'DayNum'])['EventSubType']\n                     .count()\n                     .rename('under_basket_per_game')\n                     .reset_index())\n## Aggreageted regular season\nLoyola_below_basket_agg_season = (Loyola_below_basket_agg[Loyola_below_basket_agg.DayNum<136].groupby(['Season'])['under_basket_per_game']\n                     .mean()\n                     .rename('Season Dunk/Layups avg')\n                     .reset_index())\n# made3\nLoyola_made3_agg = (Loyola_below_basket[Loyola_below_basket.EventType=='made3'].groupby(['Season', 'DayNum'])['EventSubType']\n                     .count()\n                     .rename('made3_per_game')\n                     .reset_index())\n## Aggreageted regular season\nLoyola_made3_agg_season = (Loyola_made3_agg[Loyola_made3_agg.DayNum<136].groupby(['Season'])['made3_per_game']\n                     .mean()\n                     .rename('Season 3PM avg')\n                     .reset_index())\n\nLoyola_made3_agg_season\nLoyola_style = pd.merge(Loyola_below_basket_agg_season,Loyola_made3_agg_season,left_on=[\"Season\"],right_on=[\"Season\"],how=\"left\")\ndisplay(Loyola_style)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can observe that Loyola probably played this way during the regular season since a high average of those metrics is obtained.\n\nNow, in order to see how Loyola March Madness rivals have managed to defend a similar offense during regular season, I compute the mean of dunks/layups and 3-PM (and those metrics combined in _Season Received Combined_) __received__ during the 2018 regular season and check the difference with what they received from Loyola (_Loyola Combined Damage_) in the tournament."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"losers, winners = tournament_rivals('Loyola-Chicago',2018,Mtourney_data)\n#print(losers + winners)\nrivals = losers + winners\n\nname_calculator(rivals,MTeams)\nfor i, team_identifier in enumerate(rivals):\n    #team_identifier = id_calculator(team)\n    rival_reveived_basket = MEvents_2018[((MEvents_2018.LTeamID==team_identifier) | (MEvents_2018.WTeamID==team_identifier)) & (MEvents_2018.EventTeamID!=team_identifier) & ((MEvents_2018.EventType=='made3') | ((MEvents_2018.EventType=='made2') & ((MEvents_2018.EventSubType=='lay') | (MEvents_2018.EventSubType=='dunk'))))]\n    # made 2\n    rival_reveived_below_basket_agg = (rival_reveived_basket[rival_reveived_basket.EventType=='made2'].groupby(['Season', 'DayNum'])['EventSubType']\n                         .count()\n                         .rename('under_basket_per_game')\n                         .reset_index())\n    ## Aggreageted regular season\n    rival_reveived_below_basket_agg_season = (rival_reveived_below_basket_agg[rival_reveived_below_basket_agg.DayNum<136].groupby(['Season'])['under_basket_per_game']\n                         .mean()\n                         .rename('Season Received Dunk/Layups avg')\n                         .reset_index())\n    # made 3\n    rival_made3_agg = (rival_reveived_basket[rival_reveived_basket.EventType=='made3'].groupby(['Season', 'DayNum'])['EventSubType']\n                         .count()\n                         .rename('made3_per_game')\n                         .reset_index())\n    ## Aggreageted regular season\n    rival_made3_agg_season = (rival_made3_agg[rival_made3_agg.DayNum<136].groupby(['Season'])['made3_per_game']\n                         .mean()\n                         .rename('Season Received 3PM avg')\n                         .reset_index())\n\n    rival_style = pd.merge(rival_reveived_below_basket_agg_season,rival_made3_agg_season,left_on=[\"Season\"],right_on=[\"Season\"],how=\"left\")\n    rival_style['rival'] = name_calculator([team_identifier],MTeams)\n    if i == 0:\n        rival_style_df = rival_style\n    else:\n        rival_style_df = pd.concat((rival_style_df,rival_style)).reset_index(drop=True)\n\n\nsum_list = [x + y for x, y in zip(Loyola_made3_agg[-5:]['made3_per_game'].to_list(), Loyola_below_basket_agg[-5:]['under_basket_per_game'].to_list())]\nrival_style_df['Season Received Combined'] = rival_style_df['Season Received Dunk/Layups avg'] + rival_style_df['Season Received 3PM avg']\nrival_style_df['Loyola Combined Damage'] = sum_list\nrival_style_df['Round'] = [64,32,16,8,4]\nrival_style_df['Result'] = ['W','W','W','W','L']\nrival_style_df = rival_style_df.round(2)\ndisplay(rival_style_df[['Round','rival','Season Received Dunk/Layups avg','Season Received 3PM avg','Season Received Combined','Loyola Combined Damage','Result']])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It can be observed that Loyola produced more easy baskets and 3-pointers than what opponents where used to during the regular season. Loyola obtained 4 wins, even beating Tennesse (3), all of them where teams that during that regular season allowed on average a moderate number of baskets where Loyola was proven to be very good at. During the tournament they received a huge amount of those kind of baskets so it makes me suspicious about them not having previously faced similar offenses like Loyola's. However, when they faced Michigan, a team with a very strong defense in under basket and 3-point shots, things didn't turn out well for the Loyola offense system, scoring just eight dunks/layups and one 3-point shot."},{"metadata":{},"cell_type":"markdown","source":"<div class=\"h3\">South Carolina 2017 case study</div>"},{"metadata":{},"cell_type":"markdown","source":"South Carolina team was seeded 7. However, its Elo rating had 1801.42 points at the end of the season. In fact, it was placed the 41º best team in Elo rating classification. So based on a predictive model that heavily weights this factor it is very difficult to determine that South Carolina could become a cinderella."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig = px.bar(season_enriched_concat[season_enriched_concat.Season==2017].sort_values('Score',ascending=False)[0:43], \n             x=\"TeamID\", \n             y=\"Score\",\n             title='Score average in season')\nseason_enriched_concat_name = season_enriched_concat.set_index(['TeamName'])\nseason_enriched_concat_name[season_enriched_concat_name.Season==2017]['season_elo'] \\\n    .sort_values(ascending=False)[0:43] \\\n    .plot(kind='bar',\n          figsize=(15, 5),\n         title='Elo ratings 2017')\nplt.xticks(rotation=45)\nplt.ylim(1700,2150)\nplt.annotate('South Carolina Elo Rating', xy=(41, 1803), xytext=(42, 1850),\n            arrowprops=dict(facecolor='black', shrink=0.025,width=1),\n            )\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"MEvents_2017[MEvents_2017.EventType=='block']\nseason_enriched_concat[(season_enriched_concat.Season==2017) & (season_enriched_concat.TeamID==id_calculator('South Car',MTeams))]","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"#id_calculator('South Car') -> 1376\n#loc_df[(loc_df.Season==2017) & (loc_df.WTeamID==1376) & (loc_df.DayNum == 137)]['W_dist']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"But South Carolina played the first two games in Greenville (South Carolina State), just around 160 miles from University of South Carolina. So, coming from a 71.58 points/game during regular season, they scored 93 and 88 points against Marquette (10) and Duke (2). And, from a regular season average of 7.9 steals/game and 3.67 blocks, they overperformed with 11 steals and 5 blocks agains Duke. Bear in mind that Duke was seeded 2 with only 4.74 steals and 2.77 blocks received per game."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Gets avg of season under basket and made3\nSouth_Car_below_basket = MEvents_2017[(MEvents_2017.EventTeamID==id_calculator('South Car',MTeams)) & ((MEvents_2017.EventType=='steal') | (MEvents_2017.EventType=='block'))]\n# made 2\nSouth_Car_steal_agg = (South_Car_below_basket[South_Car_below_basket.EventType=='steal'].groupby(['Season', 'DayNum'])['EventType']\n                     .count()\n                     .rename('steal_per_game')\n                     .reset_index())\n## Aggreageted regular season\nSouth_Car_steal_agg_season = (South_Car_steal_agg[South_Car_steal_agg.DayNum<136].groupby(['Season'])['steal_per_game']\n                     .mean()\n                     .rename('Season Steals avg')\n                     .reset_index())\n# made3\nSouth_Car_blocks_agg = (South_Car_below_basket[South_Car_below_basket.EventType=='block'].groupby(['Season', 'DayNum'])['EventType']\n                     .count()\n                     .rename('blocks_per_game')\n                     .reset_index())\n## Aggreageted regular season\nSouth_Car_blocks_agg_season = (South_Car_blocks_agg[South_Car_blocks_agg.DayNum<136].groupby(['Season'])['blocks_per_game']\n                     .mean()\n                     .rename('Season Blocks avg')\n                     .reset_index())\n\nSouth_car_style = pd.merge(South_Car_steal_agg_season,South_Car_blocks_agg_season,left_on=[\"Season\"],right_on=[\"Season\"],how=\"left\")\ndisplay(South_car_style)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"losers, winners = tournament_rivals('South Car',2017,Mtourney_data)\n#print(losers + winners)\nrivals = losers + winners\n#print(rivals)\nname_calculator(rivals,MTeams)\nfor i, team_identifier in enumerate(rivals):\n    #team_identifier = id_calculator(team)\n    rival_reveived_basket = MEvents_2017[((MEvents_2017.LTeamID==team_identifier) | (MEvents_2017.WTeamID==team_identifier)) & (MEvents_2017.EventTeamID!=team_identifier) & ((MEvents_2017.EventType=='steal') | ((MEvents_2017.EventType=='block')))]\n    # made 2\n    rival_reveived_below_basket_agg = (rival_reveived_basket[rival_reveived_basket.EventType=='steal'].groupby(['Season', 'DayNum'])['EventType']\n                         .count()\n                         .rename('steal_per_game')\n                         .reset_index())\n    ## Aggreageted regular season\n    rival_reveived_below_basket_agg_season = (rival_reveived_below_basket_agg[rival_reveived_below_basket_agg.DayNum<136].groupby(['Season'])['steal_per_game']\n                         .mean()\n                         .rename('Season Received Steals avg')\n                         .reset_index())\n    # made 3\n    rival_made3_agg = (rival_reveived_basket[rival_reveived_basket.EventType=='block'].groupby(['Season', 'DayNum'])['EventType']\n                         .count()\n                         .rename('blocks_per_game')\n                         .reset_index())\n    ## Aggreageted regular season\n    rival_made3_agg_season = (rival_made3_agg[rival_made3_agg.DayNum<136].groupby(['Season'])['blocks_per_game']\n                         .mean()\n                         .rename('Season Received Blocks avg')\n                         .reset_index())\n\n    rival_style = pd.merge(rival_reveived_below_basket_agg_season,rival_made3_agg_season,left_on=[\"Season\"],right_on=[\"Season\"],how=\"left\")\n    rival_style['rival'] = name_calculator([team_identifier],MTeams)\n    if i == 0:\n        rival_style_df = rival_style\n    else:\n        rival_style_df = pd.concat((rival_style_df,rival_style)).reset_index(drop=True)\n\nrival_style_df = rival_style_df[-4:-3]\nrival_style_df['Round'] = [32]\n\nsum_list = [x + y for x, y in zip(rival_reveived_below_basket_agg[-4:-3]['steal_per_game'].to_list(), rival_made3_agg[-4:-3]['blocks_per_game'].to_list())]\n#rival_style_df['Received steals by S. Carolina'] = rival_made3_agg[['blocks_per_game']][-5:].to_numpy()\nrival_style_df['Season Received Combined Steals&Blocks'] = rival_style_df['Season Received Steals avg'] + rival_style_df['Season Received Blocks avg']\nrival_style_df['S.Carolina Combined Steals&Blocks'] = sum_list\nrival_style_df['Results'] = ['W']\nrival_style_df = rival_style_df.round(2)\ndisplay(rival_style_df[['Round','rival','Season Received Steals avg','Season Received Blocks avg','Season Received Combined Steals&Blocks','S.Carolina Combined Steals&Blocks','Results']])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Distance to the arena is a variable which teams have nothing to do about. Locations are chosen across the United States, so I wonder how much that random process affect the development of the tournament. After studying the distance from the city [source](https://www.kaggle.com/kidnaps/space-time-advantage-march-madness-eda/data) where the university is located and the place where the game is played, the mathcup output distribution for the loser teams seem to be moved towards higher travelled distances. Although it may not be significant, there is evidence that travelling distance somehow affects outcome of the game. Note that since game location data is only available from 2010, this variable is not included in the previous shown predictive model, however I have introduced it for a 2010-2019 training and this variable was discarded by the model."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"sns.distplot(loc_df['W_dist'], hist=False, kde=True, \n             bins=int(180/5), color = 'darkblue',\n             hist_kws={'edgecolor':'black'},\n             kde_kws={'linewidth': 1, \"shade\": True},\n            label= 'Winner')\nsns.distplot(loc_df['L_dist'], hist=False, kde=True, \n             bins=int(180/5), color = 'red',\n             hist_kws={'edgecolor':'black'},\n             kde_kws={'linewidth': 1, \"shade\": True},\n            label= 'Loser')\nplt.legend(title='Distance to arena', bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0.);\nplt.title('Tournament winners and losers distribution based on distance to Arena')\nplt.ylabel('')\nplt.xlabel('Distance to Arena (km)')\nplt.legend();","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From then onwards South Carolina won two more games, but how can the psycological impact of playing at home and beating a top seed team like Duke be measured before even starting the tournament?\n\nLoyola University Chicago 2018 and University of South Carolina 2017 are some examples of teams that didn't shared the common characteristics of the rest of the 'Elite' teams. It was some subtle peculiarities that made them progress through the tournament. But these peculiarities appear few times in history, thus few times in the data. This makes predictive model fail to find them. It is after a detailed analysis on specific matchups that one can come up with a stronger prediction based not only on the mentioned regular season performance metrics, but also based on any of a large number of features that predictive models discard. This is a huge challenge for common machine learning algorithms and without more data such as visual/tracking data there is a limit in the ability to predict these upsets."},{"metadata":{"_kg_hide-input":true,"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"%matplotlib inline\nfrom xgboost import plot_tree\n\n##set up the parameters\nrcParams['figure.figsize'] = 50,100\nplot_tree(model_final,num_trees=2);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"rcParams['figure.figsize'] = 5,5\nplt.hist(oof_preds, bins = 10) \nplt.title(\"histogram\") \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"# load JS visualization code to notebook\nshap.initjs()\ndisplay(shap.force_plot(explainer.expected_value, shap_values_final[0,:], X.iloc[0,:]))\nshap.summary_plot(shap_values_final, X)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":4}