{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In this kernel, I try to find good measurements from league parameters to evaluate the final tourney win count. This count can show us the who wins the tourney at the end. However, when I analyze the league, I see that general statistics of team is not related with final performance. Final performance has something inclusive inside. There can be several methods that can be created. Nonetheless, while I am improvasing my method, I finally thought that again a total metric score for final measurement of wins could be benefical. Nonetheless, I took mostly samples from data and compare it with tourney win results. Then ranking the score and win rate can give us the tournement result in the end. I do not know about NCAA or its process. However, we can say that some parameters of team performance in some time period could be the projector of final tourney performance as a win rate. Thus, this analysis is based on these assumptions."},{"metadata":{"trusted":true},"cell_type":"code","source":"import seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"women_players = pd.read_csv('/kaggle/input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WPlayers.csv')\nman_teams = pd.read_csv('/kaggle/input/march-madness-analytics-2020/2020DataFiles/2020-Mens-Data/MDataFiles_Stage1/MTeams.csv')\nwoman_comp_results = pd.read_csv('/kaggle/input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WRegularSeasonCompactResults.csv')\nwoman_detailed_results = pd.read_csv('/kaggle/input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WRegularSeasonDetailedResults.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"woman_detailed_results","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"woman_comp_results","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analysis\n\nI want to find how the difference between home and away lose and win ratio. Another thing that I am curios is that how much the score difference between home and away team."},{"metadata":{"trusted":true},"cell_type":"code","source":"woman_comp_results['score_diff'] = woman_comp_results['WScore'] - woman_comp_results['LScore']\nwoman_detailed_results['score_diff'] = woman_detailed_results['WScore'] - woman_detailed_results['LScore']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = sns.barplot(x=\"WLoc\", y=\"score_diff\", data=woman_comp_results)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = sns.barplot(x=\"WLoc\", y=\"score_diff\", data=woman_detailed_results)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = sns.scatterplot(x=\"WLoc\", y=\"score_diff\", data=woman_comp_results)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_best_scoring_teams = woman_comp_results.groupby('WTeamID').mean().sort_values('score_diff',ascending =False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_best_scoring_teams","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"woman_detailed_results.corr().sort_values('score_diff',ascending =False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_total_score = woman_detailed_results.groupby('WTeamID').mean().sort_values('score_diff',ascending =False) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_total_score['total_score'] = the_total_score['score_diff'] +  the_total_score['WFGM']* 0.5 + the_total_score['WAst'] * 0.5+ the_total_score['WStl'] * 0.4 ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_total_score","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"woman_detailed_results.groupby('WTeamID').count()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table = pd.merge(woman_detailed_results.groupby('WTeamID').count()['Season'],the_total_score['total_score'],left_index = True,right_index = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table.columns = ['win_count','total_score']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table.sort_values('win_count',ascending =False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table.corr().iloc[1][0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I found that there is highly correlation with win count and score of important features. That is said, There are lots of analysis to make for total understanding"},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_detailed_results = pd.read_csv('/kaggle/input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WNCAATourneyDetailedResults.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_detailed_results.groupby('WTeamID').count()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have 81 different teams on the final dataset. So I will get some samples from league and control how it match with final season matches."},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_detailed_results['score_diff'] = tourney_detailed_results['WScore'] - tourney_detailed_results['LScore']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_total_score = tourney_detailed_results.groupby('WTeamID').mean().sort_values('score_diff',ascending =False) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_total_score.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_total_score.corr().sort_values('win',ascending = False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_total_score['total_score_tourney'] = tourney_total_score['score_diff'] + tourney_total_score['DayNum'] * 0.6 + tourney_total_score['WFGM']* 0.5 + tourney_total_score['WAst'] * 0.3 ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table_tourney = pd.merge(tourney_detailed_results.groupby('WTeamID').count()['Season'],tourney_total_score['total_score_tourney'],left_index = True,right_index = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table_tourney.columns = ['win_count_tourney','total_score_tourney']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def take_samples(number):\n    the_list_of_corr = []\n    sample = woman_detailed_results.loc[515*number:515*(number+1)]\n    the_total_score = sample.groupby('WTeamID').mean().sort_values('score_diff',ascending =False)\n    the_total_score['total_score'] = the_total_score['score_diff']  + the_total_score['WScore'] * 0.5 + the_total_score['WFGM']* 0.5 + the_total_score['WAst'] * 0.3+ the_total_score['WDR'] * 0.2 + the_total_score['WBlk'] * 0.2\n    the_final_table = pd.merge(sample.groupby('WTeamID').count()['Season'],the_total_score['total_score'],left_index = True,right_index = True)\n    the_final_table.columns = ['win_count_league','total_score_league']\n    the_final_table = pd.merge(the_final_table_tourney,the_final_table,left_index = True,right_index = True)\n    the_final_table['win_count_tourney'] = the_final_table['win_count_tourney'].rank(pct=True)\n    the_final_table['total_score_league'] = the_final_table['total_score_league'].rank(pct=True)\n    the_list_of_corr.append(the_final_table.corr().iloc[0][3])\n    return the_list_of_corr","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_list_of_similarity = [take_samples(i)[0] for i in range(100)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"take_samples(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\nplt.plot(the_list_of_similarity)\nplt.ylabel('similartiy correlations')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We take 515 number of samples from league data and collect the score correlation changes from samples to final win rate at tourney. There is no significant relation in the graph and performance has a trend increases and decreases over time. Maybe if we can test all the performances of different teams, with time series analysis, with performance metric, we can predict the final performance."},{"metadata":{},"cell_type":"markdown","source":"# How to optimize this correlation to close the 1?"},{"metadata":{},"cell_type":"markdown","source":"Maybe we cannot optimize this metric because it is a different mathematic issue but we can predict final win number with machine learning"},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table_tourney_league_match = pd.merge(woman_detailed_results,the_final_table_tourney,left_index = True,right_index = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table_tourney_league_match['total_score'] = the_final_table_tourney_league_match['score_diff']  + the_final_table_tourney_league_match['WScore'] * 0.5 + the_final_table_tourney_league_match['WFGM']* 0.5 + the_final_table_tourney_league_match['WAst'] * 0.3+ the_final_table_tourney_league_match['WDR'] * 0.2 + the_final_table_tourney_league_match['WBlk'] * 0.2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table_tourney_league_match.loc[the_final_table_tourney_league_match.WLoc == \"H\",\"WLoc\"] = 1\nthe_final_table_tourney_league_match.loc[the_final_table_tourney_league_match.WLoc == \"A\",\"WLoc\"] = 2\nthe_final_table_tourney_league_match.loc[the_final_table_tourney_league_match.WLoc == \"N\",\"WLoc\"] = 3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn import preprocessing\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.linear_model import Lasso\nfrom sklearn.linear_model import ElasticNet\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.ensemble import GradientBoostingRegressor\nfrom sklearn import linear_model\nfrom sklearn import svm\nfrom sklearn import tree\nimport xgboost as xgb\nfrom sklearn.ensemble import BaggingRegressor\nimport numpy as np \nimport pandas as pd \nimport random\nfrom sklearn.decomposition import PCA","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X = the_final_table_tourney_league_match.drop('win_count_tourney',axis = 1)\ny = the_final_table_tourney_league_match['win_count_tourney']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"regr = RandomForestRegressor()\nregr.fit(X_train, y_train)\n\npredictions = regr.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\npca = PCA(n_components=3)\nprincipalComponents_train = pca.fit_transform(X)\nsum(pca.explained_variance_ratio_)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X['component_1'] = [i[0] for i in principalComponents_train]\nX['component_2'] = [i[1] for i in principalComponents_train]\nX['component_3'] = [i[2] for i in principalComponents_train]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X = the_final_table_tourney_league_match.drop('win_count_tourney',axis = 1)\ny = the_final_table_tourney_league_match['win_count_tourney']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"regr = RandomForestRegressor(n_estimators = 400,min_samples_split = 2,min_samples_leaf = 1,max_features= 'sqrt',max_depth =None,bootstrap= False)\nregr.fit(X_train, y_train)\n\npredictions = regr.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"mean_squared_error(predictions.round(), y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ml_df = pd.DataFrame(predictions.round(),y_test).reset_index()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ml_df.corr()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that machine learning in this case is futile attempt and old score is better than this"},{"metadata":{},"cell_type":"markdown","source":"# Creating Metric for Final Win Rate"},{"metadata":{"trusted":true},"cell_type":"code","source":"the_final_table_tourney_league_match.corr().sort_values('win_count_tourney',ascending = False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is no sufficient correlation values for final win rate so I will try different method"},{"metadata":{"trusted":true},"cell_type":"code","source":"def take_samples(number):\n    the_list_of_corr = []\n    the_list_of_parameters = []\n    sample = woman_detailed_results.loc[200*number:200*(number+1)]\n    random_num_list = [random.randint(-100,100) for i in range(11)]\n    the_total_score = sample.groupby('WTeamID').mean().sort_values('score_diff',ascending =False)\n    the_total_score['total_score'] = the_total_score['score_diff'] * random_num_list[0] + the_total_score['WScore'] * random_num_list[1]+ the_total_score['WFGM']* random_num_list[2] + the_total_score['WAst'] * random_num_list[3] + the_total_score['WDR'] * random_num_list[4] + the_total_score['WBlk'] * random_num_list[5] + the_total_score['WFGM3'] * random_num_list[6] + the_total_score['WFGA3'] * random_num_list[7] + the_total_score['WFTA'] * random_num_list[8] + the_total_score['WOR'] * random_num_list[9] + the_total_score['WFTM'] * random_num_list[10]\n   \n    the_final_table = pd.merge(sample.groupby('WTeamID').count()['Season'],the_total_score['total_score'],left_index = True,right_index = True)\n    the_final_table.columns = ['win_count_league','total_score_league']\n    the_final_table = pd.merge(the_final_table_tourney,the_final_table,left_index = True,right_index = True)\n    the_final_table['win_count_tourney'] = the_final_table['win_count_tourney'].rank(pct=True)\n    the_final_table['total_score_league'] = the_final_table['total_score_league'].rank(pct=True)\n    the_list_of_corr.append(the_final_table.corr().iloc[0][3])\n    the_list_of_parameters.append(random_num_list)\n    return the_list_of_corr,the_list_of_parameters","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"take_samples(0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_list_of_similarity_parameters = [take_samples(i) for i in range(100)]\nthe_similarity_list = [i[0] for i in the_list_of_similarity_parameters]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\nplt.plot(the_similarity_list)\nplt.ylabel('similartiy correlations')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"the_max_value = 0\nparameter_list = []\nfor i in range(100):\n    the_list_of_similarity_parameters = [take_samples(i) for i in range(200)]\n    the_similarity_list = [i[0] for i in the_list_of_similarity_parameters]\n    the_parameter_list  = [i[1] for i in the_list_of_similarity_parameters]\n    if max(the_similarity_list)[0] > the_max_value:\n        the_max_value = max(the_similarity_list)[0]\n        parameter_list = the_parameter_list[the_similarity_list.index(max(the_similarity_list)[0])]\n        print(the_max_value)\n        print(parameter_list)\n        \n    ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This parameter list shows the coefficients of the equation that ==> the_total_score['score_diff'] * random_num_list[0] + the_total_score['WScore'] * random_num_list[1]+ the_total_score['WFGM']* random_num_list[2] + the_total_score['WAst'] * random_num_list[3] + the_total_score['WDR'] * random_num_list[4] + the_total_score['WBlk'] * random_num_list[5] + the_total_score['WFGM3'] * random_num_list[6] + the_total_score['WFGA3'] * random_num_list[7] + the_total_score['WFTA'] * random_num_list[8] + the_total_score['WOR'] * random_num_list[9] + the_total_score['WFTM'] * random_num_list[10]"},{"metadata":{},"cell_type":"markdown","source":"I took the parameters from table that includes league statistics and final win rate"},{"metadata":{},"cell_type":"markdown","source":"You can change the parameters of equation that can increase the correlation factor. Good luck"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}