{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Where are most goals scored from?\n\nWell, the answer is obvious: near the goal posts. But I wanted to visualize it anyway because I like heatmaps :D","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-02T05:31:24.274772Z","iopub.execute_input":"2022-10-02T05:31:24.275441Z","iopub.status.idle":"2022-10-02T05:31:25.533295Z","shell.execute_reply.started":"2022-10-02T05:31:24.275273Z","shell.execute_reply":"2022-10-02T05:31:25.531761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading in the data\nI will use the [parquet files of the original TPS Oct 2022 data](https://www.kaggle.com/datasets/reymaster/tps-oct-2022-compressed-parquet-files) kindly prepared by reybahl, because this allows me to pick which columns to load to memory rather than having to load the whole dataframe as you would with csv files. In the cell below, I just load the 'ball_pos_x', 'ball_pos_y', 'team_A_scoring_within_10sec', 'team_B_scoring_within_10sec' columns from each table (from train_1.csv to train_9.csv) and then concatenate them together.","metadata":{}},{"cell_type":"code","source":"parquet_folder = '/kaggle/input/tps-oct-2022-compressed-parquet-files'\nd = {\"t0\":None,\"t1\":None,\"t2\":None,\"t3\":None,\"t4\":None,\n     \"t5\":None,\"t6\":None,\"t7\":None,\"t8\":None,\"t9\":None}\nfor i, k in enumerate(d):\n    d[k] = pd.read_parquet(os.path.join(parquet_folder, f\"train_{i}.parquet.gzip\"),\n                           columns=['ball_pos_x', 'ball_pos_y',\n                                    'team_A_scoring_within_10sec',\n                                    'team_B_scoring_within_10sec'])\n\ndf = pd.concat([d['t0'], d['t1'], d['t2'], d['t3'], d['t4'],\n                d['t5'], d['t6'], d['t7'], d['t8'], d['t9']]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-10-02T05:45:06.104652Z","iopub.execute_input":"2022-10-02T05:45:06.105283Z","iopub.status.idle":"2022-10-02T05:45:12.121899Z","shell.execute_reply.started":"2022-10-02T05:45:06.105240Z","shell.execute_reply":"2022-10-02T05:45:12.120179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And then I divide the x and y positions into discrete chunks in order to be able to visualize them as a grid later (choice of the number of chunks - 100 and 129 for x and y respectively - is roughly representative of the dimensions of the rocket league playing arena).","metadata":{}},{"cell_type":"code","source":"df[\"x_cut\"] = pd.cut(df.ball_pos_x, bins=100, right=False)\ndf[\"y_cut\"] = pd.cut(df.ball_pos_y, bins=129, right=False)\ndf","metadata":{"execution":{"iopub.status.busy":"2022-10-02T05:46:06.277858Z","iopub.execute_input":"2022-10-02T05:46:06.280256Z","iopub.status.idle":"2022-10-02T05:46:08.538321Z","shell.execute_reply.started":"2022-10-02T05:46:06.280200Z","shell.execute_reply":"2022-10-02T05:46:08.536958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Team A's Goals ⚽🎉\nFirst, I'll just use the total number of goals scored from each x,y cell. This method has the flaw that since some locations occur more often in the game than others - most notably the centre of the field (presumably because games always start at the centre of the field?) - more goals appear to happen from those locations just because they occur more often in the game. We'll remedy this later.","metadata":{}},{"cell_type":"code","source":"# Getting the locations where Team A's goals happened\nA = pd.DataFrame(df.groupby(['x_cut', 'y_cut'])['team_A_scoring_within_10sec'].sum()).reset_index()\nA_grid = A.pivot(index='y_cut', columns='x_cut', values='team_A_scoring_within_10sec')\nA_grid","metadata":{"execution":{"iopub.status.busy":"2022-10-02T05:57:16.915805Z","iopub.execute_input":"2022-10-02T05:57:16.916346Z","iopub.status.idle":"2022-10-02T05:57:18.399239Z","shell.execute_reply.started":"2022-10-02T05:57:16.916304Z","shell.execute_reply":"2022-10-02T05:57:18.397939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing where Team A's goals occurred from (1)\n##### (just based on number of goals from each location, regardless of how often the location occurs in the game)\nie, locations from which Team A scored goals within the next 10 seconds","metadata":{}},{"cell_type":"code","source":"# First make the number of goals scored from the centre of the field 0\n# so it doesn't mess up the heatmap\nx, y = A_grid.stack().index[np.argmax(A_grid.values)]\nA_grid.loc[x][y] = 0\n\n# Visualize using seaborn's heatmap\nfig, ax = plt.subplots(figsize=(13, 10))\nsns.heatmap(A_grid)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-02T06:06:51.167320Z","iopub.execute_input":"2022-10-02T06:06:51.167839Z","iopub.status.idle":"2022-10-02T06:06:52.869138Z","shell.execute_reply.started":"2022-10-02T06:06:51.167799Z","shell.execute_reply":"2022-10-02T06:06:52.867751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that the dimensions of the arena are not displayed proportionately here, because the y dimension is actually about twice as long as the x dimension. See below \"XY plane\" section for visualization closer to the actual dimensions of the arena.","metadata":{}},{"cell_type":"markdown","source":"# Visualizing where Team A's goals occurred from (2)\n##### (taking into account the number of times each location occurs in the game)","metadata":{}},{"cell_type":"code","source":"# Counting the number of times each location occurs in the game and thus in the dataset\ncounts = pd.DataFrame(df[['x_cut', 'y_cut']].value_counts())\ncounts = counts.rename(columns={0:'count'})\ncounts = counts.reset_index()\n\ncounts_grid = counts.pivot(index='y_cut', columns='x_cut', values='count')\ncounts_grid = counts_grid.fillna(0.00001) # fill locations that do not appear at all with 0.00001\n# because while we want to fill with zeros, we need to divide using these values later","metadata":{"execution":{"iopub.status.busy":"2022-10-02T06:12:40.198567Z","iopub.execute_input":"2022-10-02T06:12:40.199001Z","iopub.status.idle":"2022-10-02T06:12:41.433598Z","shell.execute_reply.started":"2022-10-02T06:12:40.198967Z","shell.execute_reply":"2022-10-02T06:12:41.432220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Divide the number of goals scored in each xy cell by\n# the number of times each xy cell occurs\nA = pd.DataFrame(df.groupby(['x_cut', 'y_cut'])['team_A_scoring_within_10sec'].sum()).reset_index()\nA_grid = A.pivot(index='y_cut', columns='x_cut', values='team_A_scoring_within_10sec')\nA_frac_grid = A_grid/counts_grid\n\n# Visualize\nfig, ax = plt.subplots(figsize=(13, 10))\nsns.heatmap(A_frac_grid)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-02T06:13:53.677446Z","iopub.execute_input":"2022-10-02T06:13:53.677909Z","iopub.status.idle":"2022-10-02T06:13:56.563039Z","shell.execute_reply.started":"2022-10-02T06:13:53.677873Z","shell.execute_reply":"2022-10-02T06:13:56.561611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As to be expected, the area closer to the goal post is where goals are more likely to happen in the next 10 seconds.\n\nTo-do:\n- Find out the relationship between distance to goal post and likelihood of goal in 10 sec (linear? exponential?)","metadata":{}},{"cell_type":"markdown","source":"# ⚽🎉 Team B's Goals \nWe can do the same visualization for Team B.","metadata":{}},{"cell_type":"code","source":"B = pd.DataFrame(df.groupby(['x_cut', 'y_cut'])['team_B_scoring_within_10sec'].sum()).reset_index()\nB_grid = B.pivot(index='y_cut', columns='x_cut', values='team_B_scoring_within_10sec')\nB_frac_grid = B_grid/counts_grid\n\nfig, ax = plt.subplots(1, 2, figsize=(20, 7))\nx, y = B_grid.stack().index[np.argmax(B_grid.values)]\nB_grid.loc[x][y] = 0\nsns.heatmap(B_grid, ax=ax[0])\nsns.heatmap(B_frac_grid, ax=ax[1])\n\nax[0].set_title(\"Just based on number of goals from each location in the data\")\nax[1].set_title(\"Adjusted for the number of times each location occurs in the data\")\nfig.suptitle(\"Locations where Team B is likely to score in the next 10s\", fontsize=30)","metadata":{"execution":{"iopub.status.busy":"2022-10-02T06:28:41.525127Z","iopub.execute_input":"2022-10-02T06:28:41.525557Z","iopub.status.idle":"2022-10-02T06:28:45.273233Z","shell.execute_reply.started":"2022-10-02T06:28:41.525523Z","shell.execute_reply":"2022-10-02T06:28:45.271740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig.savefig('team_b_xy.jpg', bbox_inches='tight', pad_inches=0.3)","metadata":{"execution":{"iopub.status.busy":"2022-10-02T06:32:49.415021Z","iopub.execute_input":"2022-10-02T06:32:49.415535Z","iopub.status.idle":"2022-10-02T06:32:50.589765Z","shell.execute_reply.started":"2022-10-02T06:32:49.415496Z","shell.execute_reply":"2022-10-02T06:32:50.588333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Applied to other planes (YZ, XZ)","metadata":{}},{"cell_type":"code","source":"parquet_folder = '/kaggle/input/tps-oct-2022-compressed-parquet-files'\nd = {\"t0\":None,\"t1\":None,\"t2\":None,\"t3\":None,\"t4\":None,\n     \"t5\":None,\"t6\":None,\"t7\":None,\"t8\":None,\"t9\":None}\nfor i, k in enumerate(d):\n    d[k] = pd.read_parquet(os.path.join(parquet_folder, f\"train_{i}.parquet.gzip\"),\n                           columns=['ball_pos_x', 'ball_pos_y', 'ball_pos_z',\n                                    'team_A_scoring_within_10sec',\n                                    'team_B_scoring_within_10sec'])\n\ndf = pd.concat([d['t0'], d['t1'], d['t2'], d['t3'], d['t4'],\n                d['t5'], d['t6'], d['t7'], d['t8'], d['t9']]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:22:34.132993Z","iopub.execute_input":"2022-10-02T07:22:34.133478Z","iopub.status.idle":"2022-10-02T07:22:37.713540Z","shell.execute_reply.started":"2022-10-02T07:22:34.133439Z","shell.execute_reply":"2022-10-02T07:22:37.711902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def goals_heatmap(df, team, dim1, dim2):\n    '''\n    df: DataFrame containing ball_pos_[x|y|z] (2 of these) and team_[A|B]_scoring_within_10sec\n    team: A or B (str)\n    dim1: x or y or z - to be placed on horizontal axis of heatmap (str)\n    dim2: x or y or z - to be placed on vertical axis of heatmap (str)\n    '''\n    \n    dim1_length = df[f\"ball_pos_{dim1}\"].describe()['max'] - df[f\"ball_pos_{dim1}\"].describe()['min']\n    dim2_length = df[f\"ball_pos_{dim2}\"].describe()['max'] - df[f\"ball_pos_{dim2}\"].describe()['min']\n    \n    df[dim1] = pd.cut(df[f\"ball_pos_{dim1}\"], bins=int(dim1_length/2), right=False)\n    df[dim2] = pd.cut(df[f\"ball_pos_{dim2}\"], bins=int(dim2_length/2), right=False)\n    \n    goals_df = pd.DataFrame(df.groupby([dim1, dim2])[f'team_{team}_scoring_within_10sec'].sum()).reset_index()\n    goals_grid = goals_df.pivot(index=dim2, columns=dim1, values=f'team_{team}_scoring_within_10sec')\n    if dim2 == \"z\":\n        goals_grid = goals_grid[::-1]\n    \n    counts = pd.DataFrame(df[[dim1, dim2]].value_counts())\n    counts = counts.rename(columns={0:'count'})\n    counts = counts.reset_index()\n    counts_grid = counts.pivot(index=dim2, columns=dim1, values='count')\n    counts_grid = counts_grid.fillna(0.00001)\n    if dim2 == 'z':\n        counts_grid = counts_grid[::-1]\n    \n    goals_frac_grid = goals_grid/counts_grid\n    \n    if dim2 == 'z':\n        fig, ax = plt.subplots(2, 1, figsize=(int(dim1_length/10), int(dim2_length/7)), sharex=True)\n    else:\n        fig, ax = plt.subplots(1, 2, figsize=(int(dim1_length/10), int(dim2_length/10)), sharey=True)\n\n    x, y = goals_grid.stack().index[np.argmax(goals_grid.values)]\n    goals_grid.loc[x][y] = 0\n    sns.heatmap(goals_grid, ax=ax[0])\n    sns.heatmap(goals_frac_grid, ax=ax[1])     \n        \n    ax[0].set_title(\"Just based on number of goals from each location\")\n    ax[1].set_title(\"Adjusted for the number of times each location occurs\")\n    fig.suptitle(f\"Locations where Team {team} is likely to score in the next 10s\", fontsize=15)","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:47:44.176498Z","iopub.execute_input":"2022-10-02T07:47:44.176995Z","iopub.status.idle":"2022-10-02T07:47:44.193211Z","shell.execute_reply.started":"2022-10-02T07:47:44.176958Z","shell.execute_reply":"2022-10-02T07:47:44.191896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# YZ plane","metadata":{}},{"cell_type":"code","source":"goals_heatmap(df, \"A\", \"y\", \"z\")","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:47:47.075354Z","iopub.execute_input":"2022-10-02T07:47:47.075793Z","iopub.status.idle":"2022-10-02T07:47:57.846701Z","shell.execute_reply.started":"2022-10-02T07:47:47.075756Z","shell.execute_reply":"2022-10-02T07:47:57.845214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goals_heatmap(df, \"B\", \"y\", \"z\")","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:47:57.849385Z","iopub.execute_input":"2022-10-02T07:47:57.850518Z","iopub.status.idle":"2022-10-02T07:48:08.776100Z","shell.execute_reply.started":"2022-10-02T07:47:57.850448Z","shell.execute_reply":"2022-10-02T07:48:08.774849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks like it's very hard to score from directly above the goal post - which makes sense intuitively.","metadata":{}},{"cell_type":"markdown","source":"# XZ plane","metadata":{}},{"cell_type":"code","source":"goals_heatmap(df, \"A\", \"x\", \"z\")","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:48:08.778331Z","iopub.execute_input":"2022-10-02T07:48:08.778954Z","iopub.status.idle":"2022-10-02T07:48:18.632025Z","shell.execute_reply.started":"2022-10-02T07:48:08.778895Z","shell.execute_reply":"2022-10-02T07:48:18.630794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goals_heatmap(df, \"B\", \"x\", \"z\")","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:48:18.634336Z","iopub.execute_input":"2022-10-02T07:48:18.634758Z","iopub.status.idle":"2022-10-02T07:48:28.578784Z","shell.execute_reply.started":"2022-10-02T07:48:18.634722Z","shell.execute_reply":"2022-10-02T07:48:28.577138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# XY plane\nVisualized closer to the actual scale of arena","metadata":{}},{"cell_type":"code","source":"goals_heatmap(df, \"A\", \"x\", \"y\")","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:45:34.984162Z","iopub.execute_input":"2022-10-02T07:45:34.984632Z","iopub.status.idle":"2022-10-02T07:45:46.518607Z","shell.execute_reply.started":"2022-10-02T07:45:34.984589Z","shell.execute_reply":"2022-10-02T07:45:46.517068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goals_heatmap(df, \"B\", \"x\", \"y\")","metadata":{"execution":{"iopub.status.busy":"2022-10-02T07:48:43.790248Z","iopub.execute_input":"2022-10-02T07:48:43.791271Z","iopub.status.idle":"2022-10-02T07:48:56.485761Z","shell.execute_reply.started":"2022-10-02T07:48:43.791228Z","shell.execute_reply":"2022-10-02T07:48:56.484675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}