{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Tabular Challenge Oct 2022 - EDA","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Introduction","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we will explore the data set within the [Kaggle Tabular Challenge Oct 2022](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022). The objective of the problem is to predict the probability of goal for each team that could score within 10 seconds. To this end, we will explore the training data set and review some of the features that could be explored in the feature generation step.","metadata":{}},{"cell_type":"markdown","source":"## Data Analysis","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport os\nimport seaborn as sns\nimport gc","metadata":{"execution":{"iopub.status.busy":"2022-10-01T16:06:15.877449Z","iopub.execute_input":"2022-10-01T16:06:15.878184Z","iopub.status.idle":"2022-10-01T16:06:16.588323Z","shell.execute_reply.started":"2022-10-01T16:06:15.878073Z","shell.execute_reply":"2022-10-01T16:06:16.586771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DAT_DIR = '../input/tabular-playground-series-oct-2022'","metadata":{"execution":{"iopub.status.busy":"2022-10-01T16:06:16.590129Z","iopub.execute_input":"2022-10-01T16:06:16.590673Z","iopub.status.idle":"2022-10-01T16:06:16.596827Z","shell.execute_reply.started":"2022-10-01T16:06:16.590631Z","shell.execute_reply":"2022-10-01T16:06:16.595425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will use the data types specified in the train_dtypes.csv file to save memory footprint.","metadata":{}},{"cell_type":"code","source":"train_dtypes = pd.read_csv(os.path.join(DAT_DIR, 'train_dtypes.csv'))\ntrain_dtypes","metadata":{"execution":{"iopub.status.busy":"2022-10-01T16:06:16.598897Z","iopub.execute_input":"2022-10-01T16:06:16.599335Z","iopub.status.idle":"2022-10-01T16:06:16.631499Z","shell.execute_reply.started":"2022-10-01T16:06:16.599296Z","shell.execute_reply":"2022-10-01T16:06:16.630028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For convenience of presentation, we will just use the first training file for this work.","metadata":{}},{"cell_type":"code","source":"train_0 = pd.read_csv(os.path.join(DAT_DIR, f'train_{0}.csv'), dtype=dict(zip(train_dtypes.column.tolist(), train_dtypes.dtype.tolist())))","metadata":{"execution":{"iopub.status.busy":"2022-10-01T16:06:16.634799Z","iopub.execute_input":"2022-10-01T16:06:16.635224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will impute the column team_scoring_next so that the non-goal events do not get accidentally dropped within data preprocessing and visualization.","metadata":{}},{"cell_type":"code","source":"train_0['team_scoring_next'] = np.where(train_0['team_scoring_next'].isna(), \"NA\", train_0['team_scoring_next'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To understand the duration of events, we extract the start and end time of each event and visualize the distributions","metadata":{}},{"cell_type":"code","source":"event_time = train_0.groupby(['game_num', 'event_id'])['event_time'].agg(['min', 'max']).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(event_time, x='min')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(event_time, x='max')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that event_time is coded as seconds before end and the duration of each event has quite some variance. While the mode of the distribution is probably around 30 seconds, we have some outliers that last more than 10 minutes.","metadata":{}},{"cell_type":"code","source":"(train_0['event_time'] - train_0.groupby(['game_num', 'event_id'])['event_time'].shift(1)).mean()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Results of the query above shows the data set is sampled about 0.1 seconds on average. Therefore, a duration of 10 seconds will probably take 100 snapshots.","metadata":{}},{"cell_type":"markdown","source":"Let us also insert columns for the time of event start and end so as to faciliate subsequent analysis.","metadata":{}},{"cell_type":"code","source":"train_0['event_time_start'] = train_0.groupby(['game_num', 'event_id'])['event_time'].transform('min')\ntrain_0['event_time_end'] = train_0.groupby(['game_num', 'event_id'])['event_time'].transform('max')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will want to explore the locations of the ball and players next. To this end, we will examine:\n\n1. Positions of ball and player at the end of event (goal or not).\n2. Position of ball and players 10 seconds before the end.","metadata":{}},{"cell_type":"markdown","source":"To properly sample the data, we will make use of the following helper function.","metadata":{}},{"cell_type":"code","source":"def get_train_snapshot(train, dt):\n    train_snap = train.groupby(['game_num', 'event_id']).apply(lambda x: np.max(np.where(x['event_time'] - x['event_time_end'] + dt >0,\n                                                                                     - np.inf,\n                                                                                     x['event_time'] - x['event_time_end'] + dt) +\n                                                                              -dt + x['event_time_end'])).reset_index().rename(columns={0: 'snapshot'})\n    train = pd.merge(train, train_snap, on=['game_num', 'event_id'])\n    train_at_dt = train[train['event_time'] == train['snapshot']]\n    del train['snapshot']\n    return train_at_dt","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The function get_train_snapshot() returns a subset of the train data set at the (dt) seconds prior to the end. To save memory (for subsequent full train data set), we add required snapshot column for filtering within the function but drop it before returning the value. We do keep this column for the returned data set.","metadata":{}},{"cell_type":"markdown","source":"## At the End of Events","metadata":{}},{"cell_type":"markdown","source":"Let us first extract the snapshot data at the end.","metadata":{}},{"cell_type":"code","source":"train_at_0 = get_train_snapshot(train_0, dt=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we can view the distributions of the ball positions on each coordiate.","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(train_at_0, col=\"team_scoring_next\")\ng.map(sns.histplot, \"ball_pos_x\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.FacetGrid(train_at_0, col=\"team_scoring_next\")\ng.map(sns.histplot, \"ball_pos_y\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.FacetGrid(train_at_0, col=\"team_scoring_next\")\ng.map(sns.histplot, \"ball_pos_z\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from above:\n\n1. When Team A scores a goal:\n\n    1.1. The $x$ coordinate of the ball position at the end is approximately within $[-20, 20]$. This is probably the width of the goal dimension on the $x$ axis.\n    \n    1.2. The $y$ coordinate of the ball position at the end is approximately very narrowly around $100$. This is probably the location of the Goal line of the opposite team (i.e., team B).\n    \n    1.3. The $z$ coordinate of the ball position at the end largely resides within $[0, 12]$. This probably suggests some of the goals are made by the vehicles in the air.\n    \n2. When Team B scores a goal, the distribution of the coordiantes are largely similar to the counterparts of the team A's, with the only exceptin that the $y$ coordinates concentrate on $-100$, which suggests the opposite Goal line (i.e., team A's).\n\n3. When neither team scores a goal:\n\n    3.1. Coordinates $x$ and $y$ shows a crest shape. This is likely due to the activity more likely to happen within one team's side than the center.\n    \n    3.2. The distribution of $z$ coordinate is more homogeneous and wider. This probably suggests that boost items are used more often when neither team has clear advantage.","metadata":{}},{"cell_type":"markdown","source":"We can also review the position of the scoring player at the end of event.","metadata":{}},{"cell_type":"code","source":"goals = train_at_0[train_at_0.player_scoring_next != -1].reset_index(drop=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goals","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"players_cols = ['game_num', 'event_id', 'player_scoring_next'] + [f'p{i}_pos_x' for i in range(6)] + [f'p{i}_pos_y' for i in range(6)] + [f'p{i}_pos_z' for i in range(6)]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will make use the function find_player_pos_coord() to extract the position of the player who scores the goal.","metadata":{}},{"cell_type":"code","source":"def find_player_pos_coord(goals, coord='x'):\n    score_player = goals[['game_num', 'event_id', 'player_scoring_next'] + [f'p{i}_pos_{coord}' for i in range(6)]]\n    score_player = score_player.rename(columns=dict(zip([f'p{i}_pos_{coord}' for i in range(6)], [f'p{i}' for i in range(6)])))\n    score_player_long = pd.wide_to_long(score_player, stubnames=['p'], i = ['game_num', 'event_id', 'player_scoring_next'], j = 'player').reset_index()\n    return score_player_long[score_player_long.player == score_player_long.player_scoring_next].reset_index(drop=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in ['x', 'y', 'z']:\n    goals[f'score_player_{i}'] = find_player_pos_coord(goals, coord=i)['p']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goals","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=goals, x='score_player_x', y='score_player_y', hue='team_scoring_next', alpha=0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=goals, x='score_player_x', y='score_player_z', hue='team_scoring_next', alpha=0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=goals, x='score_player_y', y='score_player_z', hue='team_scoring_next', alpha=0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The scatterplots above shows the positions of the scoring player, while more likely to be close to opposing team's goal, has quite some variances.","metadata":{}},{"cell_type":"markdown","source":"### 10 Seconds Before the End","metadata":{}},{"cell_type":"code","source":"train_at_10 = get_train_snapshot(train_0, dt=10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_at_10","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.FacetGrid(train_at_10, col=\"team_scoring_next\")\ng.map(sns.histplot, \"ball_pos_x\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.FacetGrid(train_at_10, col=\"team_scoring_next\")\ng.map(sns.histplot, \"ball_pos_y\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.FacetGrid(train_at_10, col=\"team_scoring_next\")\ng.map(sns.histplot, \"ball_pos_z\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The ball distribution at 10 seconds is much less informative. It is however still observable through the $y$-coordinate that the team that scores a goal is more likely to be in the opposite team's field.","metadata":{}},{"cell_type":"code","source":"goals = train_at_10[train_at_10.player_scoring_next != -1].reset_index(drop=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in ['x', 'y', 'z']:\n    goals[f'score_player_{i}'] = find_player_pos_coord(goals, coord=i)['p']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goals","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=goals, x='score_player_x', y='score_player_y', hue='team_scoring_next', alpha=0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=goals, x='score_player_x', y='score_player_z', hue='team_scoring_next', alpha=0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=goals, x='score_player_y', y='score_player_z', hue='team_scoring_next', alpha=0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Again, the location of the scoring player is also less informative and we cannot see an apparent patterns.","metadata":{}},{"cell_type":"markdown","source":"## Conclusions","metadata":{}},{"cell_type":"markdown","source":"In this work, we have explored the training data set and extract useful insights:\n\n1. The center of field is at $y=0$. Team A's goal is at $y=-100$ and team B's goal is at $y=100$.\n\n2. When goal is made, both the ball and the scoring player are more likely to locate near the losing team's goal. While this is intuitive, we do see quite some variance from the goal position.\n\n3. 10 seconds prior to the goal, both the ball and player position tend to be much more volatile. However, the distribution of $y$ coordinate tends to suggest that ball is more likely in the losing team's field.\n\n4. Since the correlation of #3 is quite weak, it probably makes sense to explore other features by leveraging analogies to soccer. It will be interesting to see how some of the soccer metrics, e.g., ball possession, turnover, shootings, etc, will be good predictors.","metadata":{}}]}