{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-25T11:41:05.742276Z","iopub.execute_input":"2022-10-25T11:41:05.742920Z","iopub.status.idle":"2022-10-25T11:41:05.751518Z","shell.execute_reply.started":"2022-10-25T11:41:05.742886Z","shell.execute_reply":"2022-10-25T11:41:05.750221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Introduction\n\nThis notebook presents an analysis/prediction of the TPS10.22 competition using a resource-frugal approach given the large size of the dataset and the limited memory resources available.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:41:15.456220Z","iopub.execute_input":"2022-10-25T11:41:15.456597Z","iopub.status.idle":"2022-10-25T11:41:16.411064Z","shell.execute_reply.started":"2022-10-25T11:41:15.456567Z","shell.execute_reply":"2022-10-25T11:41:16.409422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n    <b>Begin Optional Run</b>\n</div>","metadata":{}},{"cell_type":"code","source":"dtypes_df = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/train_dtypes.csv')\ndtypes = {k: v for (k, v) in zip(dtypes_df.column, dtypes_df.dtype)}","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:41:26.398154Z","iopub.execute_input":"2022-10-25T11:41:26.398780Z","iopub.status.idle":"2022-10-25T11:41:26.410234Z","shell.execute_reply.started":"2022-10-25T11:41:26.398715Z","shell.execute_reply":"2022-10-25T11:41:26.408502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\ntrain0_df = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/train_0.csv', dtype=dtypes)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:41:27.627510Z","iopub.execute_input":"2022-10-25T11:41:27.627925Z","iopub.status.idle":"2022-10-25T11:41:51.806059Z","shell.execute_reply.started":"2022-10-25T11:41:27.627881Z","shell.execute_reply":"2022-10-25T11:41:51.804876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#What does data on an event look like?\ntrain0_df[train0_df['event_id'] == 1002].head(25)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:41:55.218089Z","iopub.execute_input":"2022-10-25T11:41:55.218422Z","iopub.status.idle":"2022-10-25T11:41:55.270895Z","shell.execute_reply.started":"2022-10-25T11:41:55.218398Z","shell.execute_reply":"2022-10-25T11:41:55.269919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#Delete dataframe to free up memory.\ndel train0_df","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:42:00.210304Z","iopub.execute_input":"2022-10-25T11:42:00.210677Z","iopub.status.idle":"2022-10-25T11:42:00.223285Z","shell.execute_reply.started":"2022-10-25T11:42:00.210649Z","shell.execute_reply":"2022-10-25T11:42:00.222227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Understanding the data\n\nThe dataset contains data on Rocket League series matches which in turn contain\na series of events which culminate in one of the three possibilities:\n- Team A scoring(when player 0,1,2 score)\n- Team B scoring(when player 3,4,5 score)\n- No team scoring\n\nThe events themselves are timestamped snapshots specified by sequential sets of\nplayer positions, ball positions and velocities, etc.  In other words, these\nare multiple time-series running simultaneously culminating in an event which\nis one of team A scoring, team B scoring or no one scoring.  \n\n**Target label description**\nThe target label to be predicted is the team scoring within the next 10 seconds\nof the current snapshot.  The training data already has the player scoring for\nthat particular event and also the last 10 seconds of the event are highlighted\nwith a positive target label of 1 for team A or team B or none of them to\nindicate team A, team B or no one scoring respectively.\n\n**Prediction requirements**\nGiven a snapshot in the test data, predict the probability of individual team\nscoring within the next 10 seconds. \n\n## Approach\nExamining the test data, we find that no timestamp is given to help with the\nprediction.  This means we cannot treat the training data as a set of time\nseries, to make our predictions. Rather, we should make a point prediction based\non the given test data snapshot(i.e. one row).\n\nCan we then exploit the time series nature of the data in some way to our\nadvantage?\n\nSince this is undeniably a time series kind of data, the corresponding observations\nare consecutively correlated. We are going to make the simplifying Markov\nassumption that the next state(snapshot) depends only on the current state.\nWe can thus delete all snapshots earlier than 10.1 seconds(a snapshot exists\nfor every 0.1 second) from the culminating event. This prunes our dataset\nconsiderably.\n\nThis leads us to the following approach:\n\n1. **Markovian assumption:** Delete all rows with snapshot values less than -10.1\nseconds. This will considerably prune the dataframes which could even perhaps\nbe concatenated.\n1. **Missing value mitigation:** We need to look at missing values after dropping\nthe rows in the previous step. The preferable option might be to drop the columns\nwith missing values. Dropping rows will not be viable because the last\n10 seconds of our event are crucial to our prediction model and we cannot drop\nany rows in this time segment.\n3. **Target label engineering** We create a new target label with values 0,1 and\n2 respectively for team A, team B and no one scoring in the next 10 seconds.\n\nThis processed data has both positive and negative cases for all the three target\nlabels 0,1 and 2.","metadata":{}},{"cell_type":"markdown","source":"## Data prep","metadata":{}},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#Incrementally grow a processed dataframe using steps 1 and 3 above.\nprocd_df = pd.DataFrame()  #Empty dataframe to grow\n\nfor i in range(10):\n    tmpdf = pd.read_csv(f\"/kaggle/input/tabular-playground-series-oct-2022/train_{i}.csv\", dtype=dtypes)\n    #Filter according to Markovian assumption\n    tmpdf = tmpdf[tmpdf.event_time > -10.1]\n    procd_df = pd.concat([procd_df,tmpdf],axis=0)\n    print(f'Processed training datafile: {i}')\n    \ndel tmpdf  #Free up memory","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:42:16.119938Z","iopub.execute_input":"2022-10-25T11:42:16.120656Z","iopub.status.idle":"2022-10-25T11:46:03.415849Z","shell.execute_reply.started":"2022-10-25T11:42:16.120620Z","shell.execute_reply":"2022-10-25T11:46:03.414138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Memory usage: {procd_df.memory_usage().sum()/(1024*1024)} MB; Shape: {procd_df.shape}')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:47:26.824238Z","iopub.execute_input":"2022-10-25T11:47:26.824667Z","iopub.status.idle":"2022-10-25T11:47:26.841798Z","shell.execute_reply.started":"2022-10-25T11:47:26.824631Z","shell.execute_reply":"2022-10-25T11:47:26.840458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The new dataframe size is quite tractable. We have reduced it from about 10 GB to just 600+ MB.\n\n**Missing value mitigation**","metadata":{}},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#Missing value analysis\npd.set_option('display.max_rows', None)\nprint(procd_df.isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:47:35.795134Z","iopub.execute_input":"2022-10-25T11:47:35.796166Z","iopub.status.idle":"2022-10-25T11:47:36.109229Z","shell.execute_reply.started":"2022-10-25T11:47:35.796130Z","shell.execute_reply":"2022-10-25T11:47:36.107989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we see above we have quite a few columns with quite a number of missing values. We will drop these columns.","metadata":{}},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n# Column names to be dropped\nto_be_dropped = procd_df.columns[procd_df.isna().sum()>0]","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:47:49.254434Z","iopub.execute_input":"2022-10-25T11:47:49.255617Z","iopub.status.idle":"2022-10-25T11:47:49.576164Z","shell.execute_reply.started":"2022-10-25T11:47:49.255577Z","shell.execute_reply":"2022-10-25T11:47:49.574940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Col names to be dropped\nlist(to_be_dropped)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:47:52.907324Z","iopub.execute_input":"2022-10-25T11:47:52.907696Z","iopub.status.idle":"2022-10-25T11:47:52.916011Z","shell.execute_reply.started":"2022-10-25T11:47:52.907668Z","shell.execute_reply":"2022-10-25T11:47:52.914791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#Now actually drop these columns\nprocd_df.drop(list(to_be_dropped),axis=1, inplace=True)\nprocd_df.reset_index(inplace=True, drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:05.122987Z","iopub.execute_input":"2022-10-25T11:48:05.123434Z","iopub.status.idle":"2022-10-25T11:48:05.348832Z","shell.execute_reply.started":"2022-10-25T11:48:05.123397Z","shell.execute_reply":"2022-10-25T11:48:05.347267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#Let's examine the shape of this pruned dataframe\nprocd_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:11.913965Z","iopub.execute_input":"2022-10-25T11:48:11.914360Z","iopub.status.idle":"2022-10-25T11:48:11.921748Z","shell.execute_reply.started":"2022-10-25T11:48:11.914330Z","shell.execute_reply":"2022-10-25T11:48:11.920678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\nprocd_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:16.472753Z","iopub.execute_input":"2022-10-25T11:48:16.473162Z","iopub.status.idle":"2022-10-25T11:48:16.494782Z","shell.execute_reply.started":"2022-10-25T11:48:16.473127Z","shell.execute_reply":"2022-10-25T11:48:16.493199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Missing value mitigation**\nWe will use the columns `team_A_scoring_within_10sec, team_B_scoring_within_10sec` to construct the column `team_scoring_within_10sec` which will\nhave a value of 0,1,2 for team A, team B or no one scoring respectively.","metadata":{}},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#It is assumed that mutual exclusivity of team A, B or no one scoring in the dataset is guaranteed\n#First fill the entire column with 2 -- no one scoring\nprocd_df['team_scoring_within_10sec'] = 2\n\n#Next, fill only those rows in this column with 0 where team_A_scoring_within_10sec == 1\n#Generate an index\nidx = procd_df[procd_df['team_A_scoring_within_10sec'] == 1].index\nprocd_df.loc[idx, 'team_scoring_within_10sec'] = 0\n\n#Next, fill only those rows in this column with 1 where team_B_scoring_within_10sec == 1\n#Generate an index\nidx1 = procd_df[procd_df['team_B_scoring_within_10sec'] == 1].index\nprocd_df.loc[idx1, 'team_scoring_within_10sec'] = 1\n","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:21.711674Z","iopub.execute_input":"2022-10-25T11:48:21.712395Z","iopub.status.idle":"2022-10-25T11:48:22.026772Z","shell.execute_reply.started":"2022-10-25T11:48:21.712359Z","shell.execute_reply":"2022-10-25T11:48:22.024910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\nprocd_df['team_scoring_within_10sec'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:25.697537Z","iopub.execute_input":"2022-10-25T11:48:25.697907Z","iopub.status.idle":"2022-10-25T11:48:25.731764Z","shell.execute_reply.started":"2022-10-25T11:48:25.697875Z","shell.execute_reply":"2022-10-25T11:48:25.731020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\nprocd_df[procd_df.team_B_scoring_within_10sec == 1].head(10)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:30.091588Z","iopub.execute_input":"2022-10-25T11:48:30.093022Z","iopub.status.idle":"2022-10-25T11:48:30.226636Z","shell.execute_reply.started":"2022-10-25T11:48:30.092915Z","shell.execute_reply":"2022-10-25T11:48:30.225417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Drop three more columns because they are redundant now.\nprocd_df.drop(['team_A_scoring_within_10sec','team_B_scoring_within_10sec'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:52.416848Z","iopub.execute_input":"2022-10-25T11:48:52.417221Z","iopub.status.idle":"2022-10-25T11:48:52.465309Z","shell.execute_reply.started":"2022-10-25T11:48:52.417194Z","shell.execute_reply":"2022-10-25T11:48:52.464382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#OPTIONAL RUN. JUMP TO DIRECTLY READING THE PROCESS DATA FILE\n#Write this processed data frame to /kaggle/working.\nprocd_df.to_csv('/kaggle/working/procd_df.csv', index=False)\n!ls -l /kaggle/working/","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:48:57.991303Z","iopub.execute_input":"2022-10-25T11:48:57.992356Z","iopub.status.idle":"2022-10-25T11:49:20.496738Z","shell.execute_reply.started":"2022-10-25T11:48:57.992306Z","shell.execute_reply":"2022-10-25T11:49:20.495226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n    <b>End Optional Run</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n    <b>You can directly jump to this point after running the earlier non-optional cells</b>\n</div>","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/working","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:49:40.109563Z","iopub.execute_input":"2022-10-25T11:49:40.110199Z","iopub.status.idle":"2022-10-25T11:49:40.387108Z","shell.execute_reply.started":"2022-10-25T11:49:40.110166Z","shell.execute_reply":"2022-10-25T11:49:40.385953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Directly read the prepped data.\nprocd_df = pd.read_csv('/kaggle/working/procd_df.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:49:49.332131Z","iopub.execute_input":"2022-10-25T11:49:49.332558Z","iopub.status.idle":"2022-10-25T11:49:52.273245Z","shell.execute_reply.started":"2022-10-25T11:49:49.332520Z","shell.execute_reply":"2022-10-25T11:49:52.270958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"procd_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:50:14.824954Z","iopub.execute_input":"2022-10-25T11:50:14.825295Z","iopub.status.idle":"2022-10-25T11:50:14.852798Z","shell.execute_reply.started":"2022-10-25T11:50:14.825269Z","shell.execute_reply":"2022-10-25T11:50:14.851799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n    <b>You can skip generating the graphs repeatitively and jump to model building.</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## Exploratory Data Analysis\nUnivariate Data Analysis is already available with the details of the dataset on Kaggle. We are going to concentrate here on multivariate relationships that will help us build our prediction model.","metadata":{}},{"cell_type":"markdown","source":"### Relationship plots","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:50:26.702560Z","iopub.execute_input":"2022-10-25T11:50:26.702951Z","iopub.status.idle":"2022-10-25T11:50:26.709617Z","shell.execute_reply.started":"2022-10-25T11:50:26.702921Z","shell.execute_reply":"2022-10-25T11:50:26.708222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the wallpaper-worthy scatterplots below, we can note a number of things:\n- In the events where either team scores, we observe the ball positions clustering around the opponents goal.\n- In the events where no one scores, the ball positions are not near any goal in particular.\n\nNote that the origin is at the center of the playfield and each dot is the actual ball position in the playfield.","metadata":{}},{"cell_type":"code","source":"#Explore relationship between ball position and team scoring within 10 seconds(our prediction label)\n\ng = sns.relplot(data=procd_df, x='ball_pos_x', y='ball_pos_y', size='ball_pos_z',hue='event_time',col='team_scoring_within_10sec', kind='scatter',alpha=0.3)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:50:31.216829Z","iopub.execute_input":"2022-10-25T11:50:31.217222Z","iopub.status.idle":"2022-10-25T11:53:19.301718Z","shell.execute_reply.started":"2022-10-25T11:50:31.217192Z","shell.execute_reply":"2022-10-25T11:53:19.300299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the velocity vector as profiled against event time, we can make the following observations:\n- When either team is scoring, the velocity vector points increasingly towards the opponents goal as the event time draws to a close.\n- When no one is scoring, the velocity vector points all over the place all through the event.\n\nNote that each dot is not a position but the terminal point of the velocity vector starting at the origin which is at the center of the field.","metadata":{}},{"cell_type":"code","source":"g1 = sns.relplot(data=procd_df, x='ball_vel_x', y='ball_vel_y', hue='event_time', size='ball_vel_z',col='team_scoring_within_10sec', kind='scatter',alpha=0.3)\ng1.map(plt.axhline, y=0, color=\".7\", dashes=(2, 1), zorder=0)\ng1.map(plt.axvline, x=0, color=\".7\", dashes=(2, 1), zorder=0)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T11:53:31.960597Z","iopub.execute_input":"2022-10-25T11:53:31.961384Z","iopub.status.idle":"2022-10-25T11:56:33.306988Z","shell.execute_reply.started":"2022-10-25T11:53:31.961348Z","shell.execute_reply":"2022-10-25T11:56:33.305336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Plotting boost timer distribution for each player stratified by the team scoring next, we see the following:\n- corresponding team member(s) of the team scoring within 10 seconds gets the most 0 seconds(i.e. immediate) boosts.\n- when no team is to score in the next 10 seconds, players get relatively fewer boosts.\n","metadata":{}},{"cell_type":"code","source":"#Let us now look at the distribution of boost timers.\ngd = sns.displot(data=procd_df,x='boost0_timer',hue='team_scoring_within_10sec',kind='kde')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:00:19.505027Z","iopub.execute_input":"2022-10-25T12:00:19.505418Z","iopub.status.idle":"2022-10-25T12:00:29.241495Z","shell.execute_reply.started":"2022-10-25T12:00:19.505391Z","shell.execute_reply":"2022-10-25T12:00:29.240386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gd = sns.displot(data=procd_df,x='boost3_timer',hue='team_scoring_within_10sec',kind='kde')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:00:33.153067Z","iopub.execute_input":"2022-10-25T12:00:33.153909Z","iopub.status.idle":"2022-10-25T12:00:42.431890Z","shell.execute_reply.started":"2022-10-25T12:00:33.153845Z","shell.execute_reply":"2022-10-25T12:00:42.430757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Distribution plot\nWhat is the distribution of x and y positions as stratified by who is scoring? What do we observe?\n\nIn the two sets of plots below, we observe the following:\n- For the ball_pos_x, the value 0 shows a very high peak indicating an outlier, and the steepness flattens on both sides. This shows that ball_pos_x at position 0 is more suitable for scoring goals within 10 seconds.\n- For the ball_pos_y, the value 0 again shows a very high peak and unlike the previous graph, there are slopes increasing from left to right and decreasing from right to left. These indicate proximity to either goal positions.\n- When no team is scoring, the x and y positions are equally distributed.\n\n\n","metadata":{}},{"cell_type":"code","source":"xdis = sns.displot(data=procd_df, x='ball_pos_x', col='team_scoring_within_10sec', kind='hist')\nxdis.fig.subplots_adjust(top=.9)\nxdis.fig.suptitle('Faceted Distribution of ball position x')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:01:08.863956Z","iopub.execute_input":"2022-10-25T12:01:08.865233Z","iopub.status.idle":"2022-10-25T12:01:11.968714Z","shell.execute_reply.started":"2022-10-25T12:01:08.865169Z","shell.execute_reply":"2022-10-25T12:01:11.967317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ydis = sns.displot(data=procd_df, x='ball_pos_y', col='team_scoring_within_10sec', kind='hist')\nydis.fig.subplots_adjust(top=.9)\nydis.fig.suptitle('Faceted Distribution of ball position y')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:01:17.633263Z","iopub.execute_input":"2022-10-25T12:01:17.634066Z","iopub.status.idle":"2022-10-25T12:01:19.725079Z","shell.execute_reply.started":"2022-10-25T12:01:17.634021Z","shell.execute_reply":"2022-10-25T12:01:19.723767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.displot(data=procd_df, x='ball_pos_x', y='ball_pos_y', col='team_scoring_within_10sec', fill=True, kind='kde')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:01:44.204980Z","iopub.execute_input":"2022-10-25T12:01:44.205388Z","iopub.status.idle":"2022-10-25T12:25:36.884818Z","shell.execute_reply.started":"2022-10-25T12:01:44.205357Z","shell.execute_reply":"2022-10-25T12:25:36.883119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\">\n    <b>Jump here directly to start modeling and prediction.</b>\n</div>","metadata":{}},{"cell_type":"markdown","source":"The PMF(Probability Mass Function) of our priors(unconditional probability of the labels) are equally distributed -- we have near equal probabilities of team A, team B or no one scoring. \n\nFrom the above distribution/contour plots, we can clearly infer the conditional likelihood of the ball positions given each of the three labels. These likelihoods are continuously distributed but are not gaussian. We could consider two options for our classification model(our classification model will be used to predict probabilities):\n1. A categorical Naive Bayes after discretizing the position and velocity features.\n2. A neural network with a softmax output activation.\n\nWe will choose option 2.\n\n## Data modelling","metadata":{}},{"cell_type":"code","source":"#Select the columns we want for modelling.\ncols_for_X = list(procd_df)\n\nfor name in ['game_num','event_id','event_time','player_scoring_next','team_scoring_within_10sec']:\n    cols_for_X.remove(name)\n\ncols_for_X","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:27:26.903597Z","iopub.execute_input":"2022-10-25T12:27:26.904004Z","iopub.status.idle":"2022-10-25T12:27:26.913204Z","shell.execute_reply.started":"2022-10-25T12:27:26.903965Z","shell.execute_reply":"2022-10-25T12:27:26.911921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Create the X and y vector.\nX = procd_df[cols_for_X]\ny = procd_df['team_scoring_within_10sec']","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:27:30.718657Z","iopub.execute_input":"2022-10-25T12:27:30.719052Z","iopub.status.idle":"2022-10-25T12:27:30.848871Z","shell.execute_reply.started":"2022-10-25T12:27:30.719026Z","shell.execute_reply":"2022-10-25T12:27:30.846944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.shape, y.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:28:14.387146Z","iopub.execute_input":"2022-10-25T12:28:14.388614Z","iopub.status.idle":"2022-10-25T12:28:14.396812Z","shell.execute_reply.started":"2022-10-25T12:28:14.388549Z","shell.execute_reply":"2022-10-25T12:28:14.394853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\nimport matplotlib.pyplot as plt\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\n\nfrom sklearn.neural_network import MLPClassifier\n\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.metrics import plot_confusion_matrix\nfrom sklearn.metrics import classification_report\n\nfrom sklearn.model_selection import GridSearchCV","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:28:19.116359Z","iopub.execute_input":"2022-10-25T12:28:19.116778Z","iopub.status.idle":"2022-10-25T12:28:19.328404Z","shell.execute_reply.started":"2022-10-25T12:28:19.116740Z","shell.execute_reply":"2022-10-25T12:28:19.326968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We will do a train_test_split with 30% test split\nX_train, X_test, y_train, y_test = train_test_split(X,y,test_size=0.3)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:28:21.611836Z","iopub.execute_input":"2022-10-25T12:28:21.612244Z","iopub.status.idle":"2022-10-25T12:28:22.462255Z","shell.execute_reply.started":"2022-10-25T12:28:21.612212Z","shell.execute_reply":"2022-10-25T12:28:22.460172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Create the MLP classifier.\nmlp_clf = MLPClassifier(hidden_layer_sizes=(48,),max_iter=100,activation='relu',verbose=True)\nmlp_clf","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:28:25.400706Z","iopub.execute_input":"2022-10-25T12:28:25.401095Z","iopub.status.idle":"2022-10-25T12:28:25.413209Z","shell.execute_reply.started":"2022-10-25T12:28:25.401069Z","shell.execute_reply":"2022-10-25T12:28:25.411807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mlp_clf.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:28:36.599873Z","iopub.execute_input":"2022-10-25T12:28:36.600287Z","iopub.status.idle":"2022-10-25T12:34:03.224439Z","shell.execute_reply.started":"2022-10-25T12:28:36.600255Z","shell.execute_reply":"2022-10-25T12:34:03.223146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Make sure that the output activation is softmax\nprint(mlp_clf.out_activation_)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:34:22.835530Z","iopub.execute_input":"2022-10-25T12:34:22.835935Z","iopub.status.idle":"2022-10-25T12:34:22.841652Z","shell.execute_reply.started":"2022-10-25T12:34:22.835901Z","shell.execute_reply":"2022-10-25T12:34:22.840454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Let us examine the quality of the convergence\nplt.plot(mlp_clf.loss_curve_)\nplt.title(\"Loss Curve\", fontsize=14)\nplt.xlabel('Iterations')\nplt.ylabel('Loss')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:34:28.757458Z","iopub.execute_input":"2022-10-25T12:34:28.757831Z","iopub.status.idle":"2022-10-25T12:34:28.894428Z","shell.execute_reply.started":"2022-10-25T12:34:28.757805Z","shell.execute_reply":"2022-10-25T12:34:28.893342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Although we don't need a confusion matrix, we will plot one to see how good our classification is.\n#The quality of our classifier is a direct function of the probabilities we need to predict.\n\n\nfig = plot_confusion_matrix(mlp_clf, X_test, y_test, display_labels=mlp_clf.classes_,cmap='Blues')\nfig.figure_.suptitle(\"Confusion matrix\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T12:34:45.551896Z","iopub.execute_input":"2022-10-25T12:34:45.552258Z","iopub.status.idle":"2022-10-25T12:34:46.419674Z","shell.execute_reply.started":"2022-10-25T12:34:45.552230Z","shell.execute_reply":"2022-10-25T12:34:46.418755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are interested only in the prediction of labels 0 and 1(team A scoring and team B scoring) in the three-way confusion matrix above. \nThe confusion matrix shows us some important points:\n- There is an almost equal misclassification in the two classes 0 and 1.\n- For class 0(team A scoring), precision is 60% and recall is 74%(using one versus rest - OVR)\n\nLet us now predict probabilities of classes 0 and 1 which is what we are interested in.","metadata":{}},{"cell_type":"code","source":"probs_df = pd.DataFrame(mlp_clf.predict_proba(X_test)[:,:2],columns=['team_A_scoring_within_10sec','team_B_scoring_within_10sec'])","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:20:33.665975Z","iopub.execute_input":"2022-10-25T13:20:33.666337Z","iopub.status.idle":"2022-10-25T13:20:34.398368Z","shell.execute_reply.started":"2022-10-25T13:20:33.666307Z","shell.execute_reply":"2022-10-25T13:20:34.396242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have the probabilities of our two classes(team A scoring, team B scoring), let us take another look at how well our classifier is doing. This is more intuitive than looking at a confusion matrix.","metadata":{}},{"cell_type":"code","source":"#Create a dataframe with probabilities of two classes and the true class, one in each column.\nprobs_plot_df = pd.concat([probs_df,y_test.reset_index(drop=True)],axis=1)\nprobs_plot_df","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**How are our probability predictions distributed?**  \nBelow is a plot of the distribution of probabilities of team A scoring within 10 seconds but stratified by the actual team scoring. Ideally, the curve 0(i.e when team A actually scores) should lie to the right of 0.5 and the other two should lie to the left of 0.5 indicating a clear separation of categories. However, we don't live in an ideal world and we can observe a considerable overlap in the curves. This is not surprising given the high density of positions at the center of the field in both kinds of events(i.e. team A and B scoring) as seen from the contour plots above.","metadata":{}},{"cell_type":"code","source":"sns.displot(data=probs_plot_df, x='team_A_scoring_within_10sec',hue='team_scoring_within_10sec',kind='kde',palette=sns.color_palette(\"tab10\")[:3])","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:22:06.056620Z","iopub.execute_input":"2022-10-25T13:22:06.057010Z","iopub.status.idle":"2022-10-25T13:22:09.249985Z","shell.execute_reply.started":"2022-10-25T13:22:06.056982Z","shell.execute_reply":"2022-10-25T13:22:09.249025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below is a similar set of curves for team B probability predictions. Ideally, the curve 1 should lie as much to the right of 0.5 as possible and the other curves to the left of it. The amount of overlap is inversely proportional to the quality of the probability prediction.\n","metadata":{}},{"cell_type":"code","source":"sns.displot(data=probs_plot_df, x='team_B_scoring_within_10sec',hue='team_scoring_within_10sec',kind='kde',palette=sns.color_palette(\"tab10\")[-3:])","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:22:44.109236Z","iopub.execute_input":"2022-10-25T13:22:44.109631Z","iopub.status.idle":"2022-10-25T13:22:47.195136Z","shell.execute_reply.started":"2022-10-25T13:22:44.109598Z","shell.execute_reply":"2022-10-25T13:22:47.194054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prediction with Test data","metadata":{}},{"cell_type":"code","source":"#Read in the test df\ntest_df = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:22:52.871955Z","iopub.execute_input":"2022-10-25T13:22:52.872376Z","iopub.status.idle":"2022-10-25T13:23:00.447332Z","shell.execute_reply.started":"2022-10-25T13:22:52.872335Z","shell.execute_reply":"2022-10-25T13:23:00.445645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Prune the testdf\ntest_df = test_df[cols_for_X]\ntest_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:05.950267Z","iopub.execute_input":"2022-10-25T13:23:05.950826Z","iopub.status.idle":"2022-10-25T13:23:05.973850Z","shell.execute_reply.started":"2022-10-25T13:23:05.950767Z","shell.execute_reply":"2022-10-25T13:23:05.972680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_probs_df = pd.DataFrame(mlp_clf.predict_proba(test_df)[:,:2],columns=['team_A_scoring_within_10sec','team_B_scoring_within_10sec'])","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:13.456223Z","iopub.execute_input":"2022-10-25T13:23:13.456624Z","iopub.status.idle":"2022-10-25T13:23:13.864063Z","shell.execute_reply.started":"2022-10-25T13:23:13.456588Z","shell.execute_reply":"2022-10-25T13:23:13.863060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_probs_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:17.496064Z","iopub.execute_input":"2022-10-25T13:23:17.497457Z","iopub.status.idle":"2022-10-25T13:23:17.507714Z","shell.execute_reply.started":"2022-10-25T13:23:17.497384Z","shell.execute_reply":"2022-10-25T13:23:17.506637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission","metadata":{}},{"cell_type":"code","source":"#Read in the sample submission file\nsubmission_df = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:24.215786Z","iopub.execute_input":"2022-10-25T13:23:24.217222Z","iopub.status.idle":"2022-10-25T13:23:24.395183Z","shell.execute_reply.started":"2022-10-25T13:23:24.217157Z","shell.execute_reply":"2022-10-25T13:23:24.394042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.tail(5)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:30.346306Z","iopub.execute_input":"2022-10-25T13:23:30.347642Z","iopub.status.idle":"2022-10-25T13:23:30.358756Z","shell.execute_reply.started":"2022-10-25T13:23:30.347577Z","shell.execute_reply":"2022-10-25T13:23:30.357108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df['team_A_scoring_within_10sec'] = test_probs_df['team_A_scoring_within_10sec']\nsubmission_df['team_B_scoring_within_10sec'] = test_probs_df['team_B_scoring_within_10sec']","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:36.281693Z","iopub.execute_input":"2022-10-25T13:23:36.282166Z","iopub.status.idle":"2022-10-25T13:23:36.292333Z","shell.execute_reply.started":"2022-10-25T13:23:36.282130Z","shell.execute_reply":"2022-10-25T13:23:36.290942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.tail(5)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:43.978367Z","iopub.execute_input":"2022-10-25T13:23:43.978830Z","iopub.status.idle":"2022-10-25T13:23:43.992280Z","shell.execute_reply.started":"2022-10-25T13:23:43.978786Z","shell.execute_reply":"2022-10-25T13:23:43.991035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Write the submission file\nsubmission_df.to_csv(\"/kaggle/working/submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:48.575774Z","iopub.execute_input":"2022-10-25T13:23:48.576229Z","iopub.status.idle":"2022-10-25T13:23:50.451308Z","shell.execute_reply.started":"2022-10-25T13:23:48.576197Z","shell.execute_reply":"2022-10-25T13:23:50.450105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T13:23:57.153195Z","iopub.execute_input":"2022-10-25T13:23:57.153581Z","iopub.status.idle":"2022-10-25T13:23:58.989336Z","shell.execute_reply.started":"2022-10-25T13:23:57.153552Z","shell.execute_reply":"2022-10-25T13:23:58.987561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n    <b>End of notebook</b>\n</div>","metadata":{}}]}