{"cells":[{"metadata":{},"cell_type":"markdown","source":"### What keeps your eyes glued to the screen while you are watching a basketball game? \n\nArguably, the answer is in the lines of\n* your love for the game\n* your support for one of the contending teams\n* how hard-fought the game is\n\nIn this notebook, we will try to quantify the last point or, by abusing the terminology of the competition, define **how competitive a college basketball game can be**.\n\nTrying to describe in numbers the emotion rushing through the crowd while the ball is flying towards the basket and ends in a buzzer-beater might be too ambitious. However, the core assumption of this notebook is that **a score for competitiveness can be extracted by observing how the game unfolded**. To do so, we will make extensive use of play-by-play data that tells us what happened during the game and when. For the past 6 Seasons, we have this information for both the Men's and the Women's tournaments and, while some [differences between the two tournaments can be found](https://www.kaggle.com/lucabasa/are-men-s-and-women-s-tournaments-different), we found that the way the game can unfold, either by turning out to be competitive or not, is the same for both. We will thus model them together as we are interested in finding out how to create such a score for a generic basketball game. \n\nThis notebook is organized as follows\n\n* **Data preparation**: everything can be found in [this utility script](https://www.kaggle.com/lucabasa/mm-data-manipulation) and the key concept will be summarized in this section.\n* **A hard definition for a competitive game**: what conditions a game should satisfy to be considered competitive.\n* **A score for competitiveness**: where a Machine Learning model will guide us towards the definition of our score.\n* **Finding the Madness**: here we will analyse how the score relates to the key characteristics of a game.\n* **Limits and next steps**: where we will conclude our journey.\n\n\n# Data preparation\n\nThe functions that create the data we are going to use can be found in [this utility script](https://www.kaggle.com/lucabasa/mm-data-manipulation), where the non-trivial statistics were generated by following [these definitions](https://stats.nba.com/help/glossary/). In summary, \n\n* By using the scoring events, we get the game result during the game.\n* We focus on 3 moments: the full game, the second half, the last 3 minutes.\n* We get how many times the lead of the game changed in each of these (overlapping) periods.\n* We aggregate and count the events in each period, for example the number of TO in the last 3 minutes.\n\nHere a sample of the final result"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"!pip install pandas==1.0.3  # some conflicts with the matplotib version\n\nimport numpy as np\nimport pandas as pd\n\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import log_loss, accuracy_score, roc_auc_score\nfrom sklearn.model_selection import KFold\nfrom sklearn.inspection import partial_dependence\n\nimport gc\n\n# https://www.kaggle.com/lucabasa/mm-data-manipulation\nfrom mm_data_manipulation import *\n# https://www.kaggle.com/lucabasa/mm-plots\nfrom mm_plots import *\n\npd.set_option(\"max_columns\", 300)\n\nkfolds = KFold(n_splits=5, shuffle=True, random_state=345)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def prepare_competitive(league, events):\n    if league == 'women':\n        main_loc = '/kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/'\n        regular_season = main_loc + 'WRegularSeasonDetailedResults.csv'\n        playoff = main_loc + 'WNCAATourneyDetailedResults.csv'\n        rank = None\n        season_info = main_loc + 'WSeasons.csv'\n    else:\n        main_loc = '/kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/'\n        regular_season = main_loc + 'MRegularSeasonDetailedResults.csv'\n        playoff = main_loc + 'MNCAATourneyDetailedResults.csv'\n        seed = main_loc + 'MNCAATourneySeeds.csv'\n        rank = main_loc + 'MMasseyOrdinals.csv'\n        season_info = main_loc + 'MSeasons.csv'\n        \n    reg = pd.read_csv(regular_season)\n    reg = process_details(reg, rank)\n    play = pd.read_csv(playoff)\n    play = process_details(play)\n    full = pd.concat([reg, play])\n    \n    to_use = [col for col in events if not col.endswith('_game') and \n              'FinalScore' not in col and \n              'n_OT' not in col and \n              '_difference' not in col]\n    full = pd.merge(full, events[to_use], on=['Season', 'DayNum', 'WTeamID', 'LTeamID'])\n    \n    rolling = rolling_stats(full, season_info)\n    \n    \n    competitive = events[['Season', 'DayNum', 'WTeamID', 'LTeamID', \n                          'tourney', 'Final_difference', 'Halftime_difference', '3mins_difference', \n                          'game_lc', 'half2_lc', 'crunchtime_lc', 'competitive']].copy()\n    \n    tmp = rolling.copy()\n    tmp.columns = ['Season'] + \\\n                ['W'+col for col in tmp.columns if col not in ['Season', 'DayNum']] + ['DayNum']\n    \n    competitive = pd.merge(competitive, tmp, on=['Season', 'DayNum', 'WTeamID'])\n    \n    tmp = rolling.copy()\n    tmp.columns = ['Season'] + \\\n                ['L'+col for col in tmp.columns if col not in ['Season', 'DayNum']] + ['DayNum']\n    \n    competitive = pd.merge(competitive, tmp, on=['Season', 'DayNum', 'LTeamID'])\n    \n    return competitive, full","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"all_events_m = []\nall_events_m_plot = []\nall_events_w = []\nall_events_w_plot = []\n\nfor year in np.arange(2015, 2021):\n    df = pd.read_csv(f'/kaggle/input/march-madness-analytics-2020/MPlayByPlay_Stage2/MEvents{year}.csv')\n    df = make_scores(df)\n    df = quarter_score(df)\n    df = lead_changes(df)\n    all_events_m_plot.append(df)\n    df = event_count(df)\n    all_events_m.append(df)\n    gc.collect()\n    df = pd.read_csv(f'/kaggle/input/march-madness-analytics-2020/WPlayByPlay_Stage2/WEvents{year}.csv')\n    df = make_scores(df)\n    df = quarter_score(df, men=False)\n    df = lead_changes(df)\n    all_events_w_plot.append(df)\n    df = event_count(df)\n    all_events_w.append(df)\n    gc.collect()\n\nall_events_m = pd.concat(all_events_m, ignore_index=True)\nall_events_m_plot = pd.concat(all_events_m_plot, ignore_index=True)\nall_events_w = pd.concat(all_events_w, ignore_index=True)\nall_events_w_plot = pd.concat(all_events_w_plot, ignore_index=True)\n\nall_events_m = make_competitive(all_events_m)\nall_events_m_plot = make_competitive(all_events_m_plot)\nall_events_w = make_competitive(all_events_w)\nall_events_w_plot = make_competitive(all_events_w_plot)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"comp_m, events_ext_m = prepare_competitive('men', all_events_m)\ncomp_w, events_ext_w = prepare_competitive('women', all_events_w)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"all_events_m.sample(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We are ready to use these information about the over **63000 games played in the two leagues since 2015** to find out what *competitiveness* means.\n\n# A hard definition of competitive games\n\nIn this section, we want to follow our intuition of how a competitive game looks like. The resulting label will be the key player in the next section's analysis. As we said before, we want to capture the effect of two teams going back and forth for the entire game, or maybe turning up the heat in the later stages. We don't want to rely solely on the final result as a game can be very hard-fought but one of the two teams can give up in the last few minutes and the final score would not reflect the competitiveness of the game.\n\nWe thus define a competitive game if it satisfies at least one of the following conditions:\n\n* The two teams left the floor with at most 3 points of difference.\n* The game was within 2 points 3 minutes to the end of regular time. \n* The game went to overtime.\n* The game saw more than 20 lead changes in total, or more than 10 in the second half of the game, or more than 2 in the last 3 minutes.\n\nThe distribution of some of these characteristics in the games of the past 6 years is the following."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"hardcuts_comp(all_events_m, \"Hard criteria for competitiveness - Men's tournament\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"hardcuts_comp(all_events_w, \"Hard criteria for competitiveness - Women's tournament\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Where in red we find the fraction of the games that satisfy the conditions above. We can see those games as the most competitive according to that particular criterium.\n\nWe see that the fraction of games satisfying these conditions is a bit higher in the Men's tournament, with a visible high point in the 2018 NCAA tourney. Looking at the [Kaggle competition of that year](https://www.kaggle.com/c/mens-machine-learning-competition-2018/leaderboard), it is also evident how harder it was to predict the outcome of the games of the tourney with respect to other years."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(2,1, figsize=(14, 12), facecolor='#f7f7f7')\nfig.subplots_adjust(top=0.92)\nfig.suptitle('Proportion of Competitive Games', fontsize=18)\n\ntmp_m = all_events_m.copy()\ntmp_m.tourney = tmp_m.tourney.map({0: 'Regular Season', 1: 'NCAA Tourney'})\ntmp_w = all_events_w.copy()\ntmp_w.tourney = tmp_w.tourney.map({0: 'Regular Season', 1: 'NCAA Tourney'})\n\ntmp_m.groupby(['Season', \n               'tourney']).competitive.mean().unstack()[['Regular Season', \n                                                         'NCAA Tourney']].plot(kind='bar', \n                                                                               color=['g', 'r'], \n                                                                               alpha=0.6, \n                                                                               ax=ax[0])\ntmp_w.groupby(['Season', \n               'tourney']).competitive.mean().unstack()[['Regular Season', \n                                                         'NCAA Tourney']].plot(kind='bar', \n                                                                               color=['g', 'r'], \n                                                                               alpha=0.6, \n                                                                               ax=ax[1])\nax[0].set_title(\"Men's Competition\", fontsize=14)\nax[1].set_title(\"Women's Competition\", fontsize=14)\n\n\nfor axes in ax:\n    axes.set_xticklabels(axes.get_xticklabels(), rotation=0, fontsize=12)\n    axes.set_ylim((0,1))\n    axes.set_yticklabels(['{:,.0%}'.format(x) for x in axes.get_yticks()])\n    axes.set_xlabel('')\n    axes.legend(title='', fancybox=True, fontsize=10)\n    axes.grid(axis='y')\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Interesting to notice how the fraction of competitive games stays about the same every regular season.\n\nAccording to this definition, we can see how a non-competitive game looks like"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"names_m = pd.read_csv('/kaggle/input/march-madness-analytics-2020/MDataFiles_Stage2/MTeams.csv')\nnames_w = pd.read_csv('/kaggle/input/march-madness-analytics-2020/WDataFiles_Stage2/WTeams.csv')\n\nall_names = pd.concat([names_m[['TeamID', 'TeamName']], names_w[['TeamID', 'TeamName']]])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"get_game(all_events_w_plot, all_names, final_score=20, final_smaller=False, use_competitive=True, competitive=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"While a competitive game in the men's tournament looks like this"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"get_game(all_events_m_plot, all_names, half_score=3, final_smaller=True, use_competitive=True, competitive=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Out of curiosity, we can also find the game with most lead changes (56, with 5 OT) in the women's tournament, which naturally was very competitive."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"get_game(all_events_w_plot, all_names, game_lc=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"All that being said, we can see a **limit in our definition of competitiveness**. For example, let's have a look at following game"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"get_game(all_events_m_plot, all_names, half_score=2, half_smaller=True, crunch_score=3, crunch_smaller=True, use_competitive=True, competitive=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This game was **close at half time, close 3 minutes to the end, and yet not labeled as competitive** according to our definition.\n\n# A score for competitiveness\n\nIn this section, we will train a Machine Learning model to learn the relation between game characteristics and the label for competitiveness we defined above. Including the features used to define the label, for example the point difference at the end of the game, would make this a fairly trivial exercise. Therefore we are not going to use them to train the model.\n\nWe will focus instead on features that might be directly related to these key characteristics (for example the difference between the Field Goal Made by the two teams is directly related to the final score) or with more indirect relation (like the number of Rebounds). The model is trained on 3 types of features:\n\n* **Differences between the two teams in each statistic** (for example the difference in the number of Rebounds). We also take the absolute value of this difference as the data are provided by Winning and Losing teams, therefore the signs would also inform the model on who won the game and this score should be independent on that.\n* **Total count of each statistic** (for example the total points made in the last 3 minutes).\n* **Proportions**. This can be either the total proportion of shots that went in or features like the proportions of rebounds that happened in the second half.\n\nTo have a score on each game played in the past 6 years, we use a 5-fold validation scheme so that each score is predicted after using the other 4/5 of the available data, ensuring that each prediction is not done on the same data used for training. The algorithm used is an XGBoost Classifier and the final score is given by the probability of being competitive (according to our hard definition). In other words, **a very competitive game will have a score close to 1, while a non competitive one will have a score around 0**.\n\nAfter training our model, we can have a look at the most relevant features and at how the score relates to the label we defined in the previous section."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def get_pdp(clf, feature, data, fold):\n    val, exes = partial_dependence(clf, features=feature, \n                                   X=data, grid_resolution=50)\n    fold_tmp = pd.DataFrame({'x': exes[0], 'y': val[0]})\n    fold_tmp['feat'] = feature\n    fold_tmp['fold'] = fold + 1\n    \n    return fold_tmp\n\n\ndef xgb_train(train, target, kfolds):\n\n    oof = np.zeros(len(train))\n    pd.options.mode.chained_assignment = None \n    \n    feat_df = pd.DataFrame()\n    feat_pdp = pd.DataFrame()\n\n    for fold_, (trn_idx, val_idx) in enumerate(kfolds.split(train.values, target.values)):\n        trn_data = train.iloc[trn_idx].copy()\n        val_data = train.iloc[val_idx].copy()\n        \n        trn_target = target.iloc[trn_idx]\n        val_target = target.iloc[val_idx]\n\n        \n        clf = XGBClassifier(objective='binary:logistic', \n                                       subsample=0.8,\n                                       n_jobs=5, \n                                       max_depth=8, \n                                       learning_rate=0.05, \n                                       n_estimators=10000).fit(trn_data, \n                                                             trn_target,\n                                                             eval_set=[(val_data, val_target)], \n                                                             eval_metric='logloss', \n                                                             early_stopping_rounds=100, \n                                                             verbose=False)\n        # for each split, predict on the remaining fold\n        oof[val_idx] = clf.predict_proba(val_data, ntree_limit=clf.best_iteration)[:,1]\n        # For each split, calculate the pdp\n        for feat in ['impact_diff', 'points_made_half2_diff', 'reb_crunchtime_tot', 'points_made_crunchtime_tot']:\n            fold_tmp = get_pdp(clf, feat, trn_data, fold_)\n            feat_pdp = pd.concat([feat_pdp, fold_tmp], axis=0)\n        # For each split, store the feature importance\n        fold_df = pd.DataFrame()\n        fold_df[\"feat\"] = trn_data.columns\n        fold_df[\"score\"] = clf.feature_importances_     \n        fold_df['fold'] = fold_ + 1\n        feat_df = pd.concat([feat_df, fold_df], axis=0)\n       \n\n    feat_df = feat_df.groupby('feat')['score'].agg(['mean', 'std'])\n    feat_df['abs_sco'] = (abs(feat_df['mean']))\n    feat_df = feat_df.sort_values(by=['abs_sco'],ascending=False)\n    del feat_df['abs_sco']\n\n    print(\"CV log loss: \\t {:<8.5f}\".format(log_loss(y_true=target, y_pred=oof)))\n    print('CV accuracy: \\t {:<8.5f}'.format(accuracy_score(y_true=target, y_pred=(oof>0.5).astype(int))))\n    print('CV ROC_AUC: \\t {:<8.5f}'.format(roc_auc_score(y_true=target, y_score=oof)))\n    pd.options.mode.chained_assignment = 'warn'\n    \n    return oof, feat_df, feat_pdp","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"events_m = make_feats(events_ext_m)\nevents_w = make_feats(events_ext_w)\n\nevents_m['male'] = 1\nevents_w['male'] = 0\n\nevents_t = pd.concat([events_m, events_w])\n\n# set of important features: totals, differences, and percentages\nfull_set = (['Season', 'DayNum', 'tourney', 'competitive', 'game_lc', 'half2_lc', 'crunchtime_lc', 'male', 'NumOT'] + \n            [col for col in events_m if '_tot' in col or '_diff' in col] + \n            ['OR_perc', 'points_half2_perc', 'points_crunchtime_perc', 'Shooting_perc', 'Ast_perc', 'Stl_TO', \n             'reb_half2_perc', 'reb_crunchtime_perc', 'block_half2_perc', \n             'block_crunchtime_perc', 'steal_half2_perc', 'steal_crunchtime_perc'])\n\nnot_use = ['game_lc', 'half2_lc', 'crunchtime_lc', 'Score_diff', 'NumOT', 'off_rating_diff']  # features used to define the label explictly\n# offensive rating is too directly related to the final score\n\ncustom_set = [col for col in full_set if col not in not_use]\n\n# model training and predictions\noof, feat_imp, pdps = xgb_train(events_t[custom_set].drop('competitive', axis=1), events_t.competitive, kfolds)\n\n\n# Plot feature importance and prediction vs label\nfig, ax = plt.subplots(1,2, figsize=(15, 6), facecolor='#f7f7f7')\n    \ndf = pd.DataFrame()\ndf['true'] = events_t.competitive\ndf['Prediction'] = oof\n\ndf[df.true==1]['Prediction'].hist(bins=50, ax=ax[1], alpha=0.6, color='g', label='Competitive')\ndf[df.true==0]['Prediction'].hist(bins=50, ax=ax[1], alpha=0.5, color='r', label='Not Competitive')\n\nax[1].axvline(0.5, color='k', linestyle='--')\n\nsns.barplot(x=\"mean\", y=\"feat\", ax=ax[0],\n            data=feat_imp.head(12).reset_index(), \n            xerr=feat_imp.head(12)['std'])\n\nax[1].set_title('Score vs Hard Label', fontsize=14)\nax[1].set_xlabel('Score', fontsize=12)\nax[1].grid(False)\nax[1].legend()\n\nax[0].set_xlabel('')\nax[0].set_ylabel('')\nax[0].set_title('Top Features Importance', fontsize=14)\n\n\nfig.suptitle('Competiveness Score', fontsize=18)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that the model is learning the data fairly well, although it misclassifies about 13% of the games. In the next section, we will have a look at some of these games to understand if they are indeed misclassified or rather mislabeled.\n\nThe most important feature, here generically called `impact_diff` (being the difference in impact of the 2 teams), is the team version of the Player Impact Estimate and it is defined by\n$$\nimpact_{diff} = PTS_{diff} + FGM_{diff} + FTM_{diff} - FGA_{diff} - FTA_{diff} + DREB_{diff} + \\frac{1}{2} OREB_{diff} + \\\\ AST_{diff} + STL_{diff} + \\frac{1}{2} BLK_{diff} - PF_{diff} - TO_{diff}\\,.\n$$\n\nTo better understand how the model creates the final score, we have a look at the **partial dependence plots**, which show the marginal effect of each feature on the predicted outcome. Some interesting patterns are"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def plot_pdp(data, feature, title, axes):\n    data[data.feat==feature].plot(ax=axes, x='x', y='mean', color='k')\n    axes.fill_between(data[data.feat==feature].x, (data[data.feat==feature]['mean'] - data[data.feat==feature]['std']).astype(float),\n                                (data[data.feat==feature]['mean'] + data[data.feat==feature]['std']).astype(float), alpha=0.3, color='r')\n    axes.set_title(title, fontsize=14)\n    axes.legend().set_visible(False)\n    axes.set_xlabel('')\n    return axes\n\ntmp = pdps.groupby(['feat', 'x']).y.agg(['mean', 'std']).reset_index() # mean and std across the 5 folds\n\nfig, ax = plt.subplots(2,2, figsize=(15, 12), facecolor='#f7f7f7')\nfig.subplots_adjust(top=0.93)\nfig.suptitle('Partial Dependence Plots', fontsize=18)\n\nax[0][0] = plot_pdp(tmp, 'impact_diff', 'Difference in Impact', ax[0][0])\nax[0][1] = plot_pdp(tmp, 'points_made_half2_diff', 'Difference in Points made (2nd half)', ax[0][1])\nax[1][0] = plot_pdp(tmp, 'reb_crunchtime_tot', 'Total Game Rebound', ax[1][0])\nax[1][1] = plot_pdp(tmp, 'points_made_crunchtime_tot', 'Total Points in Crunchtime', ax[1][1])\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The way the difference in points made in the second half is influencing the competitiveness score is particularly interesting. When the two teams score about the same number of points in the second half, thus in the lower end of the graph, we have that this feature contributes very little to the score. This can be interpreted as **a game already decided in the first half**, a game where the two teams are simply cruising to the final buzzer. We see how the competitiveness is increasing the more this difference increases, which can indicate a team making a **great comeback** after a disappointing first half. Then, when the difference gets above 15, the competitiveness drops again, maybe indicating a **blowout**.\n\nOther features play a more obvious role. For example, we see the negative effect of having more and more rebounds in the game because this means that we simply have **more and more missed shots**. Or how having a large number of points scored after the 37th minute mark is increasing the level of competitiveness, most likely indicating a game going to overtime.\n\nAt last, it becomes clear why Impact is our most important feature. It contains all the important aspects of team performance, including one of the features used to define the target label. We decided to keep it in the model because the presence of several statistics together was letting the model learn more complex relations while maintaining the general intuition that the **closer two teams are, the more the game is competitive**.\n\n# Finding the Madness\n\nWe have now a score that aims to capture aspects of competitiveness more complex than the intuitive but simple hard cuts we made in the first place. Naturally, this score is severely affected by the choice of these hard cuts. Therefore, it makes sense to have a look at how the competitiveness score changes when these statistics, which were *not* used to train the model, change."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# adding the score to the games\nids = events_t[['Season', 'DayNum', 'WTeamID', 'LTeamID', 'tourney', 'male', 'game_lc', 'half2_lc', 'crunchtime_lc', 'competitive', 'Score_diff', 'NumOT']].copy()\nids['competitive_score'] = oof\n\nfig, ax = plt.subplots(2,2, figsize=(15, 12), facecolor='#f7f7f7')\nfig.subplots_adjust(top=0.94)\n\nids.groupby('game_lc').competitive_score.mean().plot(ax=ax[0][0], color='r')\nax[0][0].axvline(21, color='k', linestyle='--')\nax[0][0].set_xlim((0,70))\nax[0][0].set_ylim((-0.05,1.05))\nids.groupby('half2_lc').competitive_score.mean().plot(ax=ax[0][1], color='g')\nax[0][1].axvline(11, color='k', linestyle='--')\nax[0][1].set_xlim((0,30))\nax[0][1].set_ylim((-0.05,1.05))\nids.groupby('NumOT').competitive_score.mean().plot(ax=ax[1][0], color='b')\nax[1][0].axvline(1, color='k', linestyle='--')\nax[1][0].set_xlim((0,5))\nax[1][0].set_ylim((-0.05,1.05))\nids.groupby('Score_diff').competitive_score.mean().plot(ax=ax[1][1], color='orange')\nax[1][1].axvline(3, color='k', linestyle='--')\nax[1][1].set_xlim((0,40))\nax[1][1].set_ylim((-0.05,1.05))\n\nax[0][0].set_xlabel('Lead changes in the game', fontsize=12)\nax[0][1].set_xlabel('Lead changes in the second half', fontsize=12)\nax[1][0].set_xlabel('Number of OT', fontsize=12)\nax[1][1].set_xlabel('Final score difference', fontsize=12)\nax[0][0].set_ylabel('Mean Score', fontsize=12)\nax[0][1].set_ylabel('Mean Score', fontsize=12)\nax[1][0].set_ylabel('Mean Score', fontsize=12)\nax[1][1].set_ylabel('Mean Score', fontsize=12)\n\nfig.suptitle('Mean competitiveness Score', fontsize=18)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We called *competitive* any game that had at least 20 lead changes and we see now that our score is passing the 0.5 mark around the 20th lead change. This is a very interesting consequence of our model's predictions as it had no direct information about how many times the two teams went back and forth and yet its score has a very intuitive relation with the lead changes. Moreover, the score is giving a **nuanced description** of the difference between a game with, say, 10 lead changes and one with 15, something that was not possible with the hard-cut definition. We also observe some fluctuations at the high end of the graph but this is due to the fact that there are not many games with more than 40 lead changes and the plot is displaying an average taken on very few data points.\n\nSecondly, our model is now describing the difference between games with zero, one, or more OT, following the logic that the big difference is between having or not an overtime.\n\nAt last, we see how much smoother is the curve for the final score difference that is, rightfully, saying that there is no chance of having a competitive game if it ended with more than 20 points of difference. \n\nWe can see the distribution of the score difference and lead changes and see how the 10000 most competitive games compare to the least 10000 ones."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig, ax = plt.subplots(2,1, figsize=(14, 12), facecolor='#f7f7f7')\nfig.subplots_adjust(top=0.92)\nfig.suptitle('High vs Low Competitiveness', fontsize=18)\n\nhigh = ids.sort_values('competitive_score', ascending=False).head(10000)\nlow = ids.sort_values('competitive_score', ascending=False).tail(10000)\n\nsns.kdeplot(high.Score_diff, ax=ax[0], label='High Competitiveness', color='r')\nsns.kdeplot(low.Score_diff, ax=ax[0], label='Low Competitiveness', color='b')\n\nsns.kdeplot(high.game_lc, ax=ax[1], label='High Competitiveness', color='r')\nsns.kdeplot(low.game_lc, ax=ax[1], label='Low Competitiveness', color='b')\n\nax[0].set_title('Point difference distribution', fontsize=14)\nax[1].set_title('Lead changes distribution', fontsize=14)\nax[0].get_yaxis().set_ticks([])\nax[1].get_yaxis().set_ticks([])\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Again, we see how this score is indicating as competitive or not also games that do not check all the boxes, which was the purpose of creating the score rather than a label. Indeed, we see how some of the least competitive games had the two teams within less than 10 points. Or how some of the most competitive games had no lead changes at all!\n\nWe can then have a look at one of the most competitive games in the **women's tournament** (a score above 0.99)"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_game(all_events_w_plot, all_names, 2019, 120, 3269, 3206)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Similarly, one of the most competitive games (a score above 0.99) in the **men's tournament** is"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_game(all_events_m_plot, all_names, 2018, 111, 1284, 1373)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"At last, we want to address some examples of games that got *misclassified* by our model. The first example, is a game that was **not** supposed to be competitive according to our definition, but ended up with **high competitiveness score** (0.98, the highest scoring *false positive*). It is Georgia Tech vs Florida State of 2015 men's tournament"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_game(all_events_m_plot, all_names, 2015, 103, 1199, 1210)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see how Georgia Tech came back in the second half from a 9 points deficit, firing up a back-and-forth game in which Florida State got the upper hand only in the last few minutes.\n\nSimilarly, we can have a look at a game that was supposed to be competitive, but ended up with a **low competitive score** (0, the lowest scoring *false negative*). It is Oregon-Stanford of the 2020 women's tournament."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_game(all_events_w_plot, all_names, 2020, 125, 3332, 3390)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that indeed the game had a lot of lead changes (which is why it was labeled as competitive in the first place) but then Oregon took over, never look back, and won by a mile.\n\nNaturally, there are some games for which deciding if it was a case of misclassification or mislabeling is more difficult but we can say that at the extremes we mostly find **mislabeled** games, i.e. games for which the hard-cuts definition of *competitive* was inadequate and for which the predicted score is a much better descriptor of the games' *competitiveness*.\n\n# Conclusions, limits, and next steps\n\nIn this notebook, we started with an intuitive definition of *competitiveness* and used machine learning to find the relation between our intuition and the core statistics of a basketball game. Doing so, we are now able to better describe how a hard-fought game looks like and our model showed us examples of how our intuition was sometimes wrong.\n\nNaturally, this process is highly dependent on the initial definition and this can be seen by plotting the average score across seasons, which very much looks like the average target levels we have seen at the beginning of our analysis."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(2,1, figsize=(14, 12), facecolor='#f7f7f7')\nfig.subplots_adjust(top=0.92)\nfig.suptitle('Mean Competitive Score', fontsize=18)\n\ntmp_m = ids[ids.male==1].copy()\ntmp_m.tourney = tmp_m.tourney.map({0: 'Regular Season', 1: 'NCAA Tourney'})\ntmp_w = ids[ids.male==0].copy()\ntmp_w.tourney = tmp_w.tourney.map({0: 'Regular Season', 1: 'NCAA Tourney'})\n\ntmp_m.groupby(['Season', \n               'tourney']).competitive.mean().unstack()[['Regular Season', \n                                                         'NCAA Tourney']].plot(kind='bar', \n                                                                               color=['g', 'r'], \n                                                                               alpha=0.6, \n                                                                               ax=ax[0])\ntmp_w.groupby(['Season', \n               'tourney']).competitive.mean().unstack()[['Regular Season', \n                                                         'NCAA Tourney']].plot(kind='bar', \n                                                                               color=['g', 'r'], \n                                                                               alpha=0.6, \n                                                                               ax=ax[1])\nax[0].set_title(\"Men's Competition\", fontsize=14)\nax[1].set_title(\"Women's Competition\", fontsize=14)\n\n\nfor axes in ax:\n    axes.set_xticklabels(axes.get_xticklabels(), rotation=0, fontsize=12)\n    axes.set_ylim((0,1))\n    axes.set_xlabel('')\n    axes.legend(title='', fancybox=True, fontsize=10)\n    axes.grid(axis='y')\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Therefore, more experienced sports enthusiasts can find better hard-cut definitions of *competitive* to get even richer results. However, it is interesting to notice how also the score we generated is indicating the 2018 Men's NCAA Tourney as a particularly competitive one, reflected for example into the somewhat lower performance of the [Kaggle community](https://www.kaggle.com/c/mens-machine-learning-competition-2018/leaderboard) as a whole in predicting that year's game with respect to other editions. \n\nAnother limitation that we should address is that our model is trained by using features that indirectly relate to the ones used to generate the label (for example, the difference in the number of Assists is very related to the final score difference which was used in defining the target). Being more strict in the feature selection would lead to a model that misclassifies about 10% more games (thus about 20-25% of misclassified in total). Manual inspection has revealed that most of them were genuine misclassified instances rather than mislabeled ones, as it is in the case discussed above. Therefore, more can be done in this direction to generate better features.\n\nLooking ahead, we could try to **predict the competitiveness score before the game starts**. The idea is to use the statistics about 2 teams in the (arbitrary chosen) 30 days prior to the game date and use them to predict how competitive the game between them is going to be. The value in this is that people tune in to watch a game because the game is important and how the game was promoted. However, what keeps people watching the game is also how much it is engaging (or competitive). If we could predict ahead how much a game is going to be competitive, we could predict how long people will watch it and thus be exposed to the advertisement on the court or during the breaks. This can have great value both for who sells the commercial slots and for the advertisers.\n\nThe first attempts to build a reliable model for this were unsuccessful, but if the reader forks this notebook they will find an implementation for the processing of rolling statistics for each of the available games and give it a shot.\n\nOn the other hand, one could argue that embedded in *competitiveness* there is an element of **unpredictability**. \n\nThere is the player coming out of nowhere for a chase down block. \n\nThere is the lucky bounce of the ball on the rim which determines if we will scream of joy when the buzzer goes off. \n\nThere is the feeling that no matter how bad the game is, you keep watching because the next play could be the best one in basketball history. \n\nThere is the love for the game.\n\n\n# References\n\n* [The NBA advanced stats glossary](https://stats.nba.com/help/glossary/) has been the main source of inspiration for exploring multiple possibilities and create new features\n* [Data manipulation utility script](https://www.kaggle.com/lucabasa/mm-data-manipulation)\n* [Utility script for the plots](https://www.kaggle.com/lucabasa/mm-plots)\n* [Insights on differences between Men's and Women's tournaments](https://www.kaggle.com/lucabasa/are-men-s-and-women-s-tournaments-different)\n"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}