{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nFor each player and the ball, there are 3 coordinates, resulting 7x3= 21 columns. In this notebook, I tried to reduce this information to **Euclidean distances** and **Direction**. Later I used this information to analyze the score of both teams. \n\n## Euclidean distance\n\\begin{equation}\n    d(p,q) = \\sqrt{\\sum_{i=0}^{n} ( p_i - q_i )^2}\n\\end{equation}\n\nhere `n` is the number of dimention. \nSince we have 3 dimention (x,y,z). The formula for our case would be: \n\n\\begin{equation}\n    d(p,q) = \\sqrt{ ( p_x - q_x )^2 + ( p_y - q_y )^2 + ( p_z - q_z )^2}\n\\end{equation}\n\n## Rocket League Field:\n<div style='text-align:center'>\n<img src=\"https://external-preview.redd.it/vdVj6_r4_bRb8qyzYPaeNjqgpNg8a_dO7eulrL7qM6c.png?auto=webp&s=44f2c0184ce142f33ea792b21f9761f9bfa89c93\" alt=\"drawing\" width=\"200\"/>\n</div>\n\nAccording to a Reddit post from `r/RocketLeague` the dimension of this field is 100 unit by 80 unit. In this scale, the goal post should be about 6.8 unit high","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\nsns.set_context('paper', font_scale=1.4)\npd.set_option('mode.chained_assignment', None)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-04T21:53:04.919880Z","iopub.execute_input":"2022-10-04T21:53:04.921404Z","iopub.status.idle":"2022-10-04T21:53:06.121802Z","shell.execute_reply.started":"2022-10-04T21:53:04.921243Z","shell.execute_reply":"2022-10-04T21:53:06.120409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dtypes_df = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/train_dtypes.csv')\ndtypes = {k: v for (k, v) in zip(dtypes_df.column, dtypes_df.dtype)}\ntrain0_df = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/train_0.csv', dtype=dtypes)\nX_train = train0_df","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:53:06.133908Z","iopub.execute_input":"2022-10-04T21:53:06.135590Z","iopub.status.idle":"2022-10-04T21:53:41.824990Z","shell.execute_reply.started":"2022-10-04T21:53:06.135536Z","shell.execute_reply":"2022-10-04T21:53:41.823546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.columns","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:53:41.827254Z","iopub.execute_input":"2022-10-04T21:53:41.827620Z","iopub.status.idle":"2022-10-04T21:53:41.837003Z","shell.execute_reply.started":"2022-10-04T21:53:41.827588Z","shell.execute_reply":"2022-10-04T21:53:41.836143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# work with a fraction of data for quick-run or whole dataframe\n# sub_df = X_train.loc[(X_train['game_num']<=50)].copy()\nsub_df = X_train.copy()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:54:59.955648Z","iopub.execute_input":"2022-10-04T22:54:59.956104Z","iopub.status.idle":"2022-10-04T22:55:00.015669Z","shell.execute_reply.started":"2022-10-04T22:54:59.956063Z","shell.execute_reply":"2022-10-04T22:55:00.014480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:53:42.019319Z","iopub.execute_input":"2022-10-04T21:53:42.019685Z","iopub.status.idle":"2022-10-04T21:53:42.062042Z","shell.execute_reply.started":"2022-10-04T21:53:42.019651Z","shell.execute_reply":"2022-10-04T21:53:42.060791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:53:42.063499Z","iopub.execute_input":"2022-10-04T21:53:42.063838Z","iopub.status.idle":"2022-10-04T21:53:48.908509Z","shell.execute_reply.started":"2022-10-04T21:53:42.063808Z","shell.execute_reply":"2022-10-04T21:53:48.907293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distribution of Balls x,y,z coordinates","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(3,2, figsize=(12,12))\nsns.kdeplot(data = sub_df, x= 'ball_pos_y', hue = 'team_B_scoring_within_10sec', ax = axes[0][0])\nsns.kdeplot(data = sub_df, x= 'ball_pos_y', hue = 'team_A_scoring_within_10sec',ax = axes[0][1])\n\nsns.kdeplot(data = sub_df, x= 'ball_pos_x', hue = 'team_B_scoring_within_10sec', ax = axes[1][0])\nsns.kdeplot(data = sub_df, x= 'ball_pos_x', hue = 'team_A_scoring_within_10sec',ax = axes[1][1])\n\nsns.kdeplot(data = sub_df, x= 'ball_pos_z', hue = 'team_B_scoring_within_10sec', ax = axes[2][0])\nsns.kdeplot(data = sub_df, x= 'ball_pos_z', hue = 'team_A_scoring_within_10sec',ax = axes[2][1])\n\nplt.legend('',frameon=False)","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:53:48.910254Z","iopub.execute_input":"2022-10-04T21:53:48.910579Z","iopub.status.idle":"2022-10-04T21:54:42.037998Z","shell.execute_reply.started":"2022-10-04T21:53:48.910549Z","shell.execute_reply":"2022-10-04T21:54:42.037038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(figsize=(10,6))\nsns.violinplot(data = sub_df, x= 'ball_pos_z', ax = axes)\naxes.set_title('Distribution of balls height')","metadata":{"execution":{"iopub.status.busy":"2022-10-04T23:15:26.506221Z","iopub.execute_input":"2022-10-04T23:15:26.506655Z","iopub.status.idle":"2022-10-04T23:15:27.029817Z","shell.execute_reply.started":"2022-10-04T23:15:26.506618Z","shell.execute_reply":"2022-10-04T23:15:27.028066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Insight: </strong><br>\n1) The ball_y coordinate has a multimodal distribution. Ball has stayed mostly in the middle with two distribution peaks on both sides. But, in cases when team B scored, the ball had a right skewed distribution peaking near -100 meters. Similarly, in cases when team A scored, the ball had a skewed distribution near 100 meters. <br>\n2) The ball_x coordinate also has a multimodal distribution. Ball has stayed mostly in the middle (0m) and on both sides. (-80m and 80m). In both cases, where team A and team B scored, the ball has mostly stayed in the middle. <br>\n3) The ball_z coordinate mostly follows same distribution for both cases. It has a peak near 2-3m. <br>  \n    <strong>Conclusion: </strong><br>\nball_pos_y is a vital feature for prediction\n</blockquote>","metadata":{}},{"cell_type":"markdown","source":"# Formula for calculating Distances","metadata":{}},{"cell_type":"code","source":"import math\ndef calculate_distance_1(x1,y1,z1,x2,y2,z2):\n    d = 0.0\n    d+= (x1-x2)**2\n    d+= (y1-y2)**2\n    d+= (z1-z2)**2\n    return math.sqrt(d)","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:54:42.039677Z","iopub.execute_input":"2022-10-04T21:54:42.040320Z","iopub.status.idle":"2022-10-04T21:54:42.045845Z","shell.execute_reply.started":"2022-10-04T21:54:42.040284Z","shell.execute_reply":"2022-10-04T21:54:42.045028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Balls distance from Goal post","metadata":{}},{"cell_type":"code","source":"sub_df.loc[:,'ball_dist_B_goal'] = sub_df.apply(lambda x: calculate_distance_1(x.ball_pos_x, x.ball_pos_y, x.ball_pos_z, 0, -100, 6.8), axis=1)\nsub_df.loc[:,'ball_dist_A_goal'] = sub_df.apply(lambda x: calculate_distance_1(x.ball_pos_x, x.ball_pos_y, x.ball_pos_z, 0, 100, 6.8), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:54:42.049108Z","iopub.execute_input":"2022-10-04T21:54:42.050132Z","iopub.status.idle":"2022-10-04T21:57:14.246176Z","shell.execute_reply.started":"2022-10-04T21:54:42.050090Z","shell.execute_reply":"2022-10-04T21:57:14.245136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(1,2, figsize=(10,6))\nsns.violinplot(data = sub_df, y= 'ball_dist_B_goal',x='team_B_scoring_within_10sec', ax = axes[0])\nsns.violinplot(data = sub_df, y= 'ball_dist_A_goal',x='team_A_scoring_within_10sec', ax = axes[1])","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:57:14.247854Z","iopub.execute_input":"2022-10-04T21:57:14.248600Z","iopub.status.idle":"2022-10-04T21:57:24.010532Z","shell.execute_reply.started":"2022-10-04T21:57:14.248550Z","shell.execute_reply":"2022-10-04T21:57:24.009562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Insight: </strong><br>\n1) Balls distance from goal_post significantly varies (Have different distribution) when the team scores or doesn't score. This trend is uniform for both teams. <br>\n    <strong>Conclusion: </strong><br>\n    \"ball_dist_B_goal\" and \"ball_dist_A_goal\" both are important feature for final prediction\n</blockquote>\n","metadata":{}},{"cell_type":"markdown","source":"# Players Distance from Goal post","metadata":{}},{"cell_type":"code","source":"for p in ['p0','p1','p2','p3','p4','p5']:\n    col1 = p+'_dist_B_goal'\n    col2 = p+'_dist_A_goal'\n    p_x = p+'_pos_x'\n    p_y = p+'_pos_y'\n    p_z = p+'_pos_z'\n    sub_df[col1] = sub_df.apply(lambda x: calculate_distance_1(x[p_x], x[p_y], x[p_z], 0, -100, 6.8), axis=1)\n    sub_df[col2] = sub_df.apply(lambda x: calculate_distance_1(x[p_x], x[p_y], x[p_z], 0, 100, 6.8), axis=1)    ","metadata":{"execution":{"iopub.status.busy":"2022-10-04T21:57:24.011763Z","iopub.execute_input":"2022-10-04T21:57:24.012656Z","iopub.status.idle":"2022-10-04T22:07:51.271818Z","shell.execute_reply.started":"2022-10-04T21:57:24.012618Z","shell.execute_reply":"2022-10-04T22:07:51.270498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(3,4, figsize=(20,15), sharex=True)\naxes = axes.flatten()\nax_no = 0\nfor p in ['p0','p3','p1','p4','p2','p5']:\n    col1 = p+'_dist_B_goal'\n    col2 = p+'_dist_A_goal'\n    palette = 'tab10' if p in ['p0','p1','p2'] else 'rocket'\n    sns.violinplot(data = sub_df, y= col1,x='team_B_scoring_within_10sec', ax = axes[ax_no], palette=palette)\n    sns.violinplot(data = sub_df, y= col2,x='team_A_scoring_within_10sec', ax = axes[ax_no+1], palette=palette)\n    axes[ax_no].set_title(p)\n    axes[ax_no+1].set_title(p)\n    ax_no+=2","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:07:51.273217Z","iopub.execute_input":"2022-10-04T22:07:51.273583Z","iopub.status.idle":"2022-10-04T22:08:48.788416Z","shell.execute_reply.started":"2022-10-04T22:07:51.273551Z","shell.execute_reply":"2022-10-04T22:08:48.787260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Insight: </strong><br>\n1) Players distance from goal_post significantly varies (Have different distribution) when the team scores or doesn't score. This trend is uniform for both teams. <br>\n    <strong>Conclusion: </strong><br>\n    \"px_dist_B_goal\" and \"px_dist_A_goal\" both are important feature for final prediction\n</blockquote>\n","metadata":{}},{"cell_type":"markdown","source":"# Ball Direction","metadata":{}},{"cell_type":"code","source":"# sub_df['ball_direction'] = pd.Series(np.where(X_train['ball_vel_y']>=0,'up','down'))\nsub_df.loc[:,'ball_direction'] = pd.Series(np.where(X_train['ball_vel_y']>=0,'up','down'))\nsub_df[['ball_direction']].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:08:48.790034Z","iopub.execute_input":"2022-10-04T22:08:48.790421Z","iopub.status.idle":"2022-10-04T22:08:49.495760Z","shell.execute_reply.started":"2022-10-04T22:08:48.790386Z","shell.execute_reply":"2022-10-04T22:08:49.494522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(1,2, figsize=(12,3))\nsns.barplot(data= sub_df, y='ball_direction', x= 'team_A_scoring_within_10sec', ax=axes[0])\nsns.barplot(data= sub_df, y='ball_direction', x= 'team_B_scoring_within_10sec', ax=axes[1])","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:08:49.498309Z","iopub.execute_input":"2022-10-04T22:08:49.500227Z","iopub.status.idle":"2022-10-04T22:09:44.929936Z","shell.execute_reply.started":"2022-10-04T22:08:49.500147Z","shell.execute_reply":"2022-10-04T22:09:44.928709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Insight: </strong><br>\n1) Balls direction significantly determines the chances of scoring for a team. So, \"ball_direction\" should be a feature.  <br>\n</blockquote>\n","metadata":{}},{"cell_type":"markdown","source":"# Player Direction","metadata":{}},{"cell_type":"code","source":"for p in ['p0','p1','p2','p3','p4','p5']:\n    col1 = p+'_direction'\n    col2 = p+ '_vel_y'\n    sub_df[col1] = pd.Series(np.where(X_train[col2]>=0,'up','down'))","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:09:44.931626Z","iopub.execute_input":"2022-10-04T22:09:44.932794Z","iopub.status.idle":"2022-10-04T22:09:46.177267Z","shell.execute_reply.started":"2022-10-04T22:09:44.932745Z","shell.execute_reply":"2022-10-04T22:09:46.176047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(3,4, figsize=(20,20), sharey= True)\naxes = axes.flatten()\nax_no = 0\nfor p in ['p0','p3','p1','p4','p2','p5']:\n    col1 = p+'_direction'\n    palette = 'tab10' if p in ['p0','p1','p2'] else 'rocket'\n    sns.barplot(data= sub_df, x=col1, y= 'team_A_scoring_within_10sec', ax=axes[ax_no], palette=palette)\n    sns.barplot(data= sub_df, x=col1, y= 'team_B_scoring_within_10sec', ax=axes[ax_no+1], palette=palette)\n    ax_no+=2\nplt.xticks([])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:13:15.732689Z","iopub.execute_input":"2022-10-04T22:13:15.733120Z","iopub.status.idle":"2022-10-04T22:18:58.534471Z","shell.execute_reply.started":"2022-10-04T22:13:15.733080Z","shell.execute_reply":"2022-10-04T22:18:58.533060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Insight: </strong><br>\n1) The trend from p_0 and p_2 is consistent, but p_1 varies for team A.  <br>\n2) The trend from p_3 and p_5 is consistent, but p_4 varies for team B.  <br>\n    <strong>Conclusion: </strong><br>\nThere might be a reason why the data is inconsistent for the middle player for both teams. But until we find the reason. Players direction should not be included as a feature. \n</blockquote>","metadata":{}},{"cell_type":"markdown","source":"# Missing player","metadata":{}},{"cell_type":"code","source":"sub_df['missing_A'] = pd.Series(np.where((sub_df.p0_pos_x.isnull() | sub_df.p1_pos_x.isnull() | sub_df.p2_pos_x.isnull()), 'y','n'))\nsub_df['missing_A'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:55:08.971891Z","iopub.execute_input":"2022-10-04T22:55:08.972355Z","iopub.status.idle":"2022-10-04T22:55:08.995532Z","shell.execute_reply.started":"2022-10-04T22:55:08.972315Z","shell.execute_reply":"2022-10-04T22:55:08.994727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df['missing_B'] = pd.Series(np.where((sub_df.p3_pos_x.isnull() | sub_df.p4_pos_x.isnull() | sub_df.p5_pos_x.isnull()), 'y','n'))\nsub_df['missing_B'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T22:55:11.663347Z","iopub.execute_input":"2022-10-04T22:55:11.663796Z","iopub.status.idle":"2022-10-04T22:55:11.697228Z","shell.execute_reply.started":"2022-10-04T22:55:11.663751Z","shell.execute_reply":"2022-10-04T22:55:11.695968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(2,2, figsize=(12,9), sharex=True)\naxes = axes.flatten()\nsns.barplot(data= sub_df, y='missing_A', x= 'team_A_scoring_within_10sec', ax=axes[0])\nsns.barplot(data= sub_df, y='missing_A', x= 'team_B_scoring_within_10sec', ax=axes[1])\nsns.barplot(data= sub_df, y='missing_B', x= 'team_A_scoring_within_10sec', ax=axes[2])\nsns.barplot(data= sub_df, y='missing_B', x= 'team_B_scoring_within_10sec', ax=axes[3])","metadata":{"execution":{"iopub.status.busy":"2022-10-04T23:08:48.848066Z","iopub.execute_input":"2022-10-04T23:08:48.848565Z","iopub.status.idle":"2022-10-04T23:08:55.732987Z","shell.execute_reply.started":"2022-10-04T23:08:48.848527Z","shell.execute_reply":"2022-10-04T23:08:55.731724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<blockquote style=\"margin-right:auto; margin-left:auto; background-color: #ebf9ff; padding: 1em; margin:24px;\">\n    <strong>Insight: </strong><br>\n1) Missing a player from any team affects scoring for both teams. So, \"missing_A\" and \"missing_B\" both should be a feature.\n</blockquote>","metadata":{}},{"cell_type":"markdown","source":"This is my original idea. Please let me know if this idea was helpful or not. I appreciate your feedback. ","metadata":{}},{"cell_type":"markdown","source":"# References:\n1) Image link: [field image](https://external-preview.redd.it/vdVj6_r4_bRb8qyzYPaeNjqgpNg8a_dO7eulrL7qM6c.png?auto=webp&s=44f2c0184ce142f33ea792b21f9761f9bfa89c93)","metadata":{}}]}