{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"\n# Part 1 : Understanding the data\n\n![](https://cdn-images-1.medium.com/max/1200/1*wdGUB-2bxIMCzbUJOGw65w.jpeg)\n\nMarch brings one of the most awaited events for sports fans in the US — The **NCAA Women’s and Men’s Division 1 Tournament** aka **March madness**. This brings the sport of basketball into the spotlight and many basketball fanatics get into the work of predicting the winners and rooting for their favorites. Basketball is a fairly popular game in the US and is [ranked second to](https://en.wikipedia.org/wiki/Basketball_in_the_United_States) American football. However, for a person like me, who is born in India, my familiarity with March madness is on a lower side. Things would have been different if this were a competition for predicting [**IPL**](https://www.iplt20.com/) winners. The **Indian Premier League** (**IPL**) is a professional [**Twenty20 cricket**](https://en.wikipedia.org/wiki/Twenty20_cricket \"Twenty20 cricket\") league in [India](https://en.wikipedia.org/wiki/India \"India\") contested during March or April and May of every year by eight teams representing eight different cities in India. So, before exploring the dataset, I shall first explain the whole concept of NCAA March Madness and how the format is designed. Hopefully, this will help the people to actually understand a large amount of dataset and not be daunted by it.\n\n# **NCAA Division I Basketball Tournament**\n\n \n![](https://cdn-images-1.medium.com/max/800/1*2fU29HRgk4ySJ_l1N5NaJw.jpeg)\n\nThis tournament is a knockout tournament where the loser is immediately eliminated from the tournament. Since it is mostly played in march, hence it has been accorded the title of **March Madness**. The first edition took place in 1939 and has been regularly held since then. the Women’s Championship was inaugurated in the 1981–82 season.\n\n# Format\n\nThe male edition tournament comprises of **68** teams that compete in **7** rounds for the National Championship Title. However, the number of Teams in the Women’s edition is **64**.\n\n![](https://cdn-images-1.medium.com/max/800/1*TaaEJ3zTwhuU67QPqqrkaA.png)\n\n---\n\n# Selection\n\nThe selection procedure takes place by two methods:\n\n![](https://cdn-images-1.medium.com/max/800/1*s7gpAnvzL-mQ0lKlzc8xXQ.png)\n\n## 1. Automatic\n\n32 Teams get selected in this way.\n\n-   Men’s Division 1 Team comprises of **353** Teams.\n\n![](https://cdn-images-1.medium.com/max/800/1*DBT72cUKGLIvXmjO7mBgyQ.png)\n\n-   Each one of those teams belongs to **32** [conferences](https://en.wikipedia.org/wiki/List_of_NCAA_conferences).\n\n![](https://cdn-images-1.medium.com/max/800/1*rq4HBtMnQeGsiI7hOmBfIA.png)\n\n-   Each of those conferences conducts a tournament and if a time wins the tournament, they get selected for the NCAA.\n\n  \n\n## 2. At Large\n\nThe second selection process is called ‘At Large’ where The NCAA selection committee convenes at the final days of the regular season and decides which 36 teams which are not the Automatic qualifiers can be sent to the playoffs. This selection is based on multiple stats and rankings.\n\n---\n\n## Selection Sunday\n\nThese “at-large” teams are announced in a nationally televised event on the Sunday preceding the [“First Four” play-in games](https://en.wikipedia.org/wiki/NCAA_Men%27s_Division_I_Basketball_Opening_Round_game \"NCAA Men's Division I Basketball Opening Round game\"). This Sunday is called ‘Selection Sunday and is on March 15.\n\n## Seeding\n\nAfter all the 68(64 in case of Women), have been decided, the selection committee ranks them in a process called seeding where each team gets a ranking from 1 to 68. Then **First Four** play-in games are contested between teams holding the four lowest-seeded automatic bids and the four lowest-seeded at-large bids.\n\nThe Teams are then split into 4 regions of 16 Teams each. Each team is now ranked from 1 to 16 in each region. After the [First Four](https://en.wikipedia.org/wiki/First_Four \"First Four\"), the tournament occurs during the course of three weekends, at pre-selected neutral sites across the United States. Here, the first round matches are determined by pitting the top team in the region with the lowest-seeded team in that region and so on. This ranking is the team’s seed.\n\n# March Madness Begins\n![](https://www.ncaa.com/sites/default/files/public/styles/original/public-s3/images/2020/02/12/kelly-campbell-depaul-2020-ncaa.jpg?itok=suRzKxqw)\n\n## First Round\n\nThe First round consisting of 64 teams playing in 32 games over the course of a week. From here 32 teams emerge as winners and go on to the second round.\n\n## Sweet Sixteen\n\nNext, the sweet sixteen round takes place, which sees the elimination of 16 teams. Rest of the 16 teams move forward.\n\n## Elite Eight\n\nThe next fight is for the Elite Eight as only 8 teams remain in the competition.\n\n## Final Four\n![](https://media.giphy.com/media/2fMOp0fPmvwwCgLXUK/giphy.gif)\nThe penultimate round of the tournament where the 4 teams contest to reserve a place in the finals. Four teams, one from each region (East, South, Midwest, and West), compete in a preselected location for the national championship.\n \n ---\n\n# Who are Cinderellas?\n\nUpsets do happen in the tournament and sometimes the underdogs, who are seeded low, deliver an unexpected. They are called Cinderellas.\n\n\nSo, this was a background behind the NCAA Baskerball tournament. Now let's have a look at the datasets provided.I shall be analysing the NCAA Division I Women's Basketball Tournament data. I assume the Men's tournament data should also be on the same lines."},{"metadata":{},"cell_type":"markdown","source":"---\n# Part 2\n### Analysing NCAA Division I Women's Basketball Tournament Data\n\nOur goal is to use the historical data to understand \"*what dictates the ability of a team to “stay in the game” and increase their chance to win late in the contest*?\""},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"# Importing the libraries\nimport pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nfrom sklearn import model_selection\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import log_loss\n\nimport seaborn as sns\nfrom IPython.display import display","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Methodology"},{"metadata":{},"cell_type":"markdown","source":"The data that has been provided has been grouped under different sections. This is useful due to the large amount of data.The various groups are:\n>>\n* Basics\n* Team Box Scores\n* Geography\n* Public Rankings\n* Play by Play\n* Other Supplementary data\n\nIt'll be prudent to go over every section to understand the nature of the data provided. In the coming days, I shall be doing exactly this and analysing which factors contribute towards better performance."},{"metadata":{"trusted":true},"cell_type":"markdown","source":"## Data Section 1- The Basics\n\nThis includes the details about the Team, Seasons, Seeds Information,Game Results. \n\n### 1. The Team"},{"metadata":{"trusted":true},"cell_type":"code","source":"\nWteams = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WTeams.csv')\nWteams.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# No of Teams\n\nWteams['TeamID'].nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> Team ID referes to the unique ID which identifies every team. There are 365 participating Women 's Teams."},{"metadata":{},"cell_type":"markdown","source":"### 2. Seasons\n\nThe year in which the tournament was played.The current season counts as 2020."},{"metadata":{"trusted":true},"cell_type":"code","source":"Wseason = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WSeasons.csv')\nWseason.tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Total held seasons including the current\nWseason['Season'].count()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> There are 4 regions in the final tournament- X, W,X,Y and Z."},{"metadata":{},"cell_type":"markdown","source":"### 3. Seed Data\nThis file identifies the seeds for all teams in each NCAA® tournament, for all seasons of historical data"},{"metadata":{"trusted":true},"cell_type":"code","source":"Wseeds = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WNCAATourneySeeds.csv')\nWseeds.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The Seed value consists of a 3 character Identifier. The first character denotes the region and the last two denote the seed in that region. Let's merge the Team's name from the Wteams file."},{"metadata":{"trusted":true},"cell_type":"code","source":"Wseeds = pd.merge(Wseeds, Wteams,on='TeamID')\nWseeds.head()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Separating the regions from the Seeds\n\nWseeds['Region'] = Wseeds['Seed'].apply(lambda x: x[0][:1])\nWseeds['Seed'] = Wseeds['Seed'].apply(lambda x: int(x[1:3]))\nprint(Wseeds.head())\nprint(Wseeds.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Teams with maximum top seeds\nfig = plt.gcf()\nfig.set_size_inches(10, 6)\ncolors = ['dodgerblue', 'plum', '#F0A30A','#8c564b','orange','green','yellow'] \n\nWseeds[Wseeds['Seed'] ==1]['TeamName'].value_counts()[:10].plot(kind='bar',color=colors,linewidth=2,edgecolor='black')\nplt.xlabel('Number of times in Top seeded positions')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Connecticut/UConn has been the top seeded team for the maximum no of times"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Teams with maximum lowest seeds\nfig = plt.gcf()\nfig.set_size_inches(10, 6)\n\nWseeds[Wseeds['Seed'] ==16]['TeamName'].value_counts()[:10].plot(kind='bar',color=colors,edgecolor='black',linewidth=1)\nplt.xlabel('Number of times in bottom seeded positions')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Does being Top/ Lower seeding affect the tournament results? This is a question to be looked upon."},{"metadata":{"trusted":true},"cell_type":"markdown","source":"### 4. Regular Season Compact results\nThis file identifies the game-by-game results for many seasons of historical data, starting with the 1998 season. There are 0 to 132 day numbers for selection of 64 teams.\n\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"rg_season_compact_results = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WRegularSeasonCompactResults.csv')\nrg_season_compact_results.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"where \n* WScore -  the numberof points scored by the winning team.\n* WTeamID - the id number of the team that won the game\n* LTeamID - the id number of the team that lost the game.\n* LScore - the number of points scored by the losing team. \n"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Winning and Losing score Average over the years\nx = rg_season_compact_results.groupby('Season')[['WScore','LScore']].mean()\n\nfig = plt.gcf()\nfig.set_size_inches(14, 6)\nplt.plot(x.index,x['WScore'],marker='o', markerfacecolor='green', markersize=12, color='green', linewidth=4)\nplt.plot(x.index,x['LScore'],marker=7, markerfacecolor='red', markersize=12, color='red', linewidth=4)\nplt.legend()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 5.Tourney Compact Results\nThis file identifies the game-by-game tournament results for all seasons of historical data. "},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_compact_results = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WNCAATourneyCompactResults.csv')\ntourney_compact_results .tail()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This file is pretty similar to the previous file except that there are 63 games listed in all the seasons.Whereas for the regular season, it displays all the games played.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"games_played = tourney_compact_results.groupby('Season')['DayNum'].count().to_frame().merge(rg_season_compact_results.groupby('Season')['DayNum'].count().to_frame(),on='Season')\ngames_played.rename(columns={\"DayNum_x\": \"Tournament Games\", \"DayNum_y\": \"Regular season games\"})\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Is there a home team advantage?"},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = sns.countplot(x=tourney_compact_results['WLoc'])\nax.set_title(\"Win Locations\")\nax.set_xlabel(\"Location\")\nax.set_ylabel(\"Frequency\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data Section 2- Team Box Scores\n\nThis section provides game-by-game stats at a team level (free throws attempted, defensive rebounds, turnovers, etc.) for all regular season, conference tournament, and NCAA® tournament games since the 2009-10 season."},{"metadata":{},"cell_type":"markdown","source":"### 1.WNCAA Tourney Detailed Results.\nThis file provides team-level box scores for many NCAA® tournaments, starting with the 2010 season"},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_detailed_results = pd.read_csv('../input/march-madness-analytics-2020/2020DataFiles/2020-Womens-Data/WDataFiles_Stage1/WNCAATourneyDetailedResults.csv')\ntourney_detailed_results.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tourney_detailed_results.columns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Again let's checkput if there is a home team advantage in the tournaments?"},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = sns.countplot(x=tourney_detailed_results['WLoc'])\nax.set_title(\"Win Locations\")\nax.set_xlabel(\"Location\")\nax.set_ylabel(\"Frequency\");\n","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"\ngames_stats = []\n\nfor row in tourney_detailed_results.to_dict('records'):\n    game = {}\n    game['Season'] =  row['Season']\n    game['DayNum'] = row['DayNum']\n    game['TeamID'] = row['WTeamID']\n    game['OpponentID'] = row['LTeamID']\n    game['FGM'] = row['WFGM']\n    game['Loc'] = row['WLoc']\n    game['Won'] = 1\n    game['Score'] = row['WScore']\n    game['FGA'] = row['WFGA']\n    game['FGM3'] = row['WFGM3']\n    game['FGA3'] = row['WFGA3']\n    game['FTM'] = row['WFTM']\n    game['FTA'] = row['WFTA']\n    game['OR'] = row['WOR']\n    game['DR'] = row['WDR']\n    game['AST'] = row['WAst']\n    game['TO'] = row['WTO']\n    game['STL'] = row['WStl']\n    game['BLK'] = row['WBlk']\n    game['PF'] = row['WPF']\n    games_stats.append(game)\n    game = {}\n    game['Season'] = row['Season']\n    game['DayNum'] = row['DayNum']\n    game['TeamID'] = row['LTeamID']\n    game['OpponentID'] = row['WTeamID']\n    game['FGM'] = row['LFGM']\n    game['Loc'] = row['WLoc']\n    game['Won']= 0\n    game['Score'] = row['LScore']\n    game['FGA'] = row['LFGA']\n    game['FGM3'] = row['LFGM3']\n    game['FGA3'] = row['LFGA3']\n    game['FTM'] = row['LFTM']\n    game['FTA'] = row['LFTA']\n    game['OR'] = row['LOR']\n    game['DR'] = row['LDR']\n    game['AST'] = row['LAst']\n    game['TO'] = row['LTO']\n    game['STL'] = row['LStl']\n    game['BLK'] = row['LBlk']\n    game['PF'] = row['LPF']\n    games_stats.append(game)\n\n    \n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Separating winners and losers using Won Column which is set to 1 for winner and 0 for loser\n\ntournament = pd.DataFrame(games_stats)\ntournament.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's add the Team Seed and ID from the `Wseed` dataset. We shall include the seed for both the current and the opponent's team\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"tournament_df = pd.merge(tournament , Wseeds, on= ['Season','TeamID'])\ntournament_df.rename(columns={'Seed': 'Team_Seed'}, inplace=True)\ntournament_df[:2]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tournament_df2 = pd.merge(tournament_df , Wseeds.rename(columns={'TeamID':'OpponentID'}), on= ['Season','OpponentID'])\ntournament_df2 .rename(columns={'Seed': 'OpponentSeed',\n                                'TeamName_x':'Team',\n                                'TeamName_y':'Opponents',\n                                 'Region_x':'Team_Region',\n                                 'Region_y':'Opponent_Region'}, inplace=True)\ntournament_df2 .head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Thus we have a database that has winners and losers clearly marked along with their seeds and regions."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Winning_Teams\n\nwinning_Teams = tournament_df2[tournament_df2['Won'] == 1]\n\n# Losing_Teams\n\nlosing_Teams = tournament_df2[tournament_df2['Won'] == 0]\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Analysis of the Winning Teams"},{"metadata":{"trusted":true},"cell_type":"code","source":"winning_Teams.head().T","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Most successful teams\nfig = plt.gcf()\nfig.set_size_inches(10, 6)\n\ncolors = ['dodgerblue', 'plum', '#F0A30A','#8c564b','orange','green','yellow'] \nwinning_Teams['Team'].value_counts()[:10].plot(kind='bar',color=colors,edgecolor='black',linewidth=1 )\n\nplt.title('Most successful Teams')\nplt.tight_layout(h_pad=2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# How seed ranking affect game stats for winners\nsns.pairplot(winning_Teams[['FGA3','FGM3','AST','BLK','DR','FTA','FTM','OR','Team_Seed',]], hue='Team_Seed',kind=\"scatter\",plot_kws=dict(s=80, edgecolor=\"white\", linewidth=2.5))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Features Correlated with Wins\nf,ax = plt.subplots(figsize=(20,15))\ncorr = tournament_df2.corr()\nsns.heatmap(corr, cmap='inferno', annot=True)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Baseline Logistic Regression Model\n\nThe `tournament_df2` dataset will be the data used for creating our model. Let's time split it to create a validation dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Converting Loc column to numeric\n\ntournament_df2 = pd.get_dummies(tournament_df2, columns=['Loc'])\ntournament_df2.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = tournament_df2[tournament_df2['Season'] < 2015]\nvalidation = tournament_df2[tournament_df2['Season'] >= 2015]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.columns\nvalidation.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nX = ['DayNum', 'TeamID', 'OpponentID', 'FGM','FGA', 'FGM3', 'FGA3', 'FTM', 'FTA', 'OR', 'DR', 'AST', 'TO', 'STL',\n       'BLK', 'PF', 'Team_Seed','Loc_A', 'Loc_H','Loc_N']\ny = 'Won'\n\nmodel = LogisticRegression(solver='liblinear',C=1.0)\nmodel.fit(train[X],train[y])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predictions = mode.predict_proba(validation1[X])[:, 1]\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"validation.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# submit predicitions\n\nsubmission = pd.read_csv('/kaggle/input/google-cloud-ncaa-march-madness-2020-division-1-womens-tournament/WSampleSubmissionStage1_2020.csv')\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"validation['Pred'] = predictions\nvalidation['ID'] = validation.apply(lambda row: '{}_{}_{}'.format(int(row['Season']), int(row['TeamID']), int(row['OpponentID'])), axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\npredictions = pd.merge(submission.drop('Pred', axis = 1), validation[['ID', 'Pred']], how='left', on=['ID']).fillna(0.5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predictions.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predictions.to_csv('submision_baseline.csv',index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}