{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!wget https://raw.githubusercontent.com/JoseCaliz/dotfiles/main/css/gruvbox.css 2>/dev/null 1>&2\n!pip install feature_engine 2>/dev/null 1>&2\n!pip install fastparquet 2>/dev/null 1>&2\n    \nfrom IPython.core.display import HTML\nwith open('./gruvbox.css', 'r') as file:\n    custom_css = file.read()\n\nHTML(custom_css)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:51:25.262296Z","iopub.execute_input":"2022-10-04T19:51:25.262763Z","iopub.status.idle":"2022-10-04T19:51:58.112153Z","shell.execute_reply.started":"2022-10-04T19:51:25.262664Z","shell.execute_reply":"2022-10-04T19:51:58.110619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Introduction\n\nRocket league is a multiplaform game similar to soccer but with control remote cars. Each team is composed by 2 to 8 players. Cars can only hit the ball, they cannot grab. The goal of the competition is to predict either if Team A | B score within the next 10 seconds, or the third case in which none of them is about to score.\n\nI'm creating this EDA to grab some ideas on which to put effort on feature engineering.\n\nThe stadium reference:\n\n<img src='https://i.imgur.com/bnfkLHC.jpeg'/>","metadata":{"execution":{"iopub.status.busy":"2022-10-01T05:10:52.821938Z","iopub.execute_input":"2022-10-01T05:10:52.823027Z","iopub.status.idle":"2022-10-01T05:10:52.827607Z","shell.execute_reply.started":"2022-10-01T05:10:52.822988Z","shell.execute_reply":"2022-10-01T05:10:52.826297Z"}}},{"cell_type":"markdown","source":"# Toc\n\n1. [Introduction](#Introduction)\n1. [Toc](#Toc)\n1. [Library Import](#Library-Import)\n1. [Read Competition Data](#Read-Competition-Data)\n1. [Events per game](#Events-per-game)\n1. [Target Relationships](#Target-Relationships)\n1. [Bi-variate](#Bi-variate)\n1. [Nan Values](#Nan-Values)\n    1. [Nan Values have predictive value](#Nan-Values-have-predictive-value)\n    1. [Agregate counts and do the test](#Agregate-counts-and-do-the-test)\n1. [Label](#Label)\n1. [Preprocessing:](#Preprocessing:)\n    1. [Save To CSV](#Save-To-CSV)\n1. [Keras Baseline](#Keras-Baseline)\n    1. [How to load your model](#How-to-load-your-model)\n    1. [Generate Submission File](#Generate-Submission-File)\n1. [How to improve upon this Baseline?](#How-to-improve-upon-this-Baseline?)\n\n","metadata":{}},{"cell_type":"markdown","source":"# Library Import ","metadata":{}},{"cell_type":"code","source":"from rich import print\nfrom plotly.subplots import make_subplots\nfrom feature_engine.wrappers import SklearnTransformerWrapper as SKWrapper\nfrom sklearn.preprocessing import StandardScaler\n\nimport matplotlib as mpl\nimport matplotlib.cm as cmap\nimport matplotlib.colors as mpl_colors\nimport matplotlib.pyplot as plt\nimport matplotlib.ticker as ticker\nimport numpy as np\nimport pandas as pd\nfrom rich.progress import track\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport plotly.io as pio\nimport warnings\nimport gc\nimport seaborn as sns\nimport random\nimport tensorflow as tf\nimport os\nfrom cycler import cycler\nimport plotly.io as pio\n\ntf.random.set_seed(42)\nrandom.seed(42)\ndef hex_to_rgb(h):\n    h = h.lstrip('#')\n    return tuple(int(h[i:i+2], 16)/255 for i in (0, 2, 4))\n\ncolor_palette = [\"#ad7a99\",\"#5c7457\",\"#6dc0d5\",\"#ceec97\",\"#5a716a\"]\ncolor_palette_rgb = [hex_to_rgb(x) for x in color_palette]\ncmap = mpl_colors.ListedColormap(color_palette_rgb)\ncolors = cmap.colors\n\nbg_color = '#2c2c2e'\ntick_color = '#e5e5ea'\n\ncustom_params = {\n    \"axes.spines.right\": False,\n    \"axes.spines.top\": False,\n    'grid.alpha':0.3,\n    'figure.figsize': (16, 5),\n    'axes.titlesize': 'Large',\n    'axes.labelsize': 'Large',\n    'figure.facecolor': bg_color,\n    'axes.facecolor': bg_color,\n    'text.color': tick_color,\n    'axes.labelcolor': tick_color,\n    'axes.edgecolor': tick_color,\n    'xtick.color': tick_color,\n    'ytick.color': tick_color,\n    'axes.prop_cycle': cycler('color',color_palette_rgb)\n}\n\nsns.set_theme(\n    style='whitegrid',\n    palette=sns.color_palette(color_palette),\n    rc=custom_params\n)\n\n\npio.templates.default = \"plotly_dark\"\npio.templates['plotly_dark'].layout.colorway = color_palette\npio.templates['plotly_dark'].layout.plot_bgcolor = bg_color\npio.templates['plotly_dark'].layout.paper_bgcolor = bg_color\npio.templates['plotly_dark'].layout.paper_bgcolor = bg_color\n\nprint('Notebooks Color Palette')\nsns.color_palette(color_palette)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:51:58.115227Z","iopub.execute_input":"2022-10-04T19:51:58.116098Z","iopub.status.idle":"2022-10-04T19:52:06.425977Z","shell.execute_reply.started":"2022-10-04T19:51:58.116038Z","shell.execute_reply":"2022-10-04T19:52:06.424800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read Competition Data\n\nThere are just too many records to process (2149381 in just one file!!). One can just take a sample and expect that distributions won't change. One can just grab a couple of games and use those as a starting point to understand the data.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_parquet('../input/tps-oct22-parquet-datasets/train_0.parquet.gzip', engine='fastparquet')\ntest_df = pd.read_csv('../input/tabular-playground-series-oct-2022/test.csv', index_col=0)\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-04T19:52:06.427558Z","iopub.execute_input":"2022-10-04T19:52:06.428570Z","iopub.status.idle":"2022-10-04T19:52:27.961045Z","shell.execute_reply.started":"2022-10-04T19:52:06.428530Z","shell.execute_reply":"2022-10-04T19:52:27.960061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"games = random.sample(list(train_df.game_num.unique()), 300)\ntrain_df = train_df[train_df.game_num.isin(games)]\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-04T19:52:27.963690Z","iopub.execute_input":"2022-10-04T19:52:27.964061Z","iopub.status.idle":"2022-10-04T19:52:28.274052Z","shell.execute_reply.started":"2022-10-04T19:52:27.964028Z","shell.execute_reply":"2022-10-04T19:52:28.272540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Events per game\n\n**Insight**\n\n* There are several events per game, each event is either a score by one team or a null event where the game just ended.\n* Games can have no scores.\n* Chunks follows the same distribution","metadata":{"execution":{"iopub.status.busy":"2022-10-01T05:20:48.230946Z","iopub.execute_input":"2022-10-01T05:20:48.231408Z","iopub.status.idle":"2022-10-01T05:20:48.239224Z","shell.execute_reply.started":"2022-10-01T05:20:48.231348Z","shell.execute_reply":"2022-10-01T05:20:48.237477Z"}}},{"cell_type":"code","source":"counts = []\nfor i in range(10):\n    chunk_df = pd.read_parquet(\n        f'../input/tps-oct22-parquet-datasets/train_{i}.parquet.gzip', engine='fastparquet'\n    )\n    \n    counts.append(\n        chunk_df.groupby('game_num')['event_id'].nunique().\n        to_frame('event_count').assign(chunk=i)\n    )\n\n\nax = sns.countplot(data=pd.concat(counts, axis=0), x='event_count', hue='chunk', palette=color_palette)\nplt.title('Events Per Game');","metadata":{"execution":{"iopub.status.busy":"2022-10-04T19:52:28.277105Z","iopub.execute_input":"2022-10-04T19:52:28.277523Z","iopub.status.idle":"2022-10-04T19:53:30.931361Z","shell.execute_reply.started":"2022-10-04T19:52:28.277453Z","shell.execute_reply":"2022-10-04T19:53:30.929519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Animation of events","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\nimport gc\nfrom plotly.offline import init_notebook_mode\n\ninit_notebook_mode()\n\nevent = train_df[train_df.event_id.eq(1003)]\nframes = []\nfor i, (index, row) in enumerate(event.iterrows()):\n    if i == 0:\n        continue\n    data = [\n        go.Scatter(y=[row.p0_pos_x], x=[row.p0_pos_y], mode='text', text=['🏎'], textfont=dict(size=40)),\n        go.Scatter(y=[row.p1_pos_x], x=[row.p1_pos_y], mode='text', text=['🏎'], textfont=dict(size=40)),\n        go.Scatter(y=[row.p2_pos_x], x=[row.p2_pos_y], mode='text', text=['🏎'], textfont=dict(size=40)),\n        go.Scatter(y=[row.p3_pos_x], x=[row.p3_pos_y], mode='text', text=['🚙'], textfont=dict(size=40)),\n        go.Scatter(y=[row.p4_pos_x], x=[row.p4_pos_y], mode='text', text=['🚙'], textfont=dict(size=40)),\n        go.Scatter(y=[row.p5_pos_x], x=[row.p5_pos_y], mode='text', text=['🚙'], textfont=dict(size=40)),\n        go.Scatter(y=[row.ball_pos_x], x=[row.ball_pos_y], mode='text', text=['⚽️'], textfont=dict(size=80))\n    ]\n    \n    frames.append(\n        go.Frame(data=data)\n    )\n    \nfig = go.Figure(\n    data=[\n        go.Scatter(y=[event.iloc[0].p0_pos_x], x=[event.iloc[0].p0_pos_y], mode='text', text=['🏎'], textfont=dict(size=40)),\n        go.Scatter(y=[event.iloc[0].p1_pos_x], x=[event.iloc[0].p1_pos_y], mode='text', text=['🏎'], textfont=dict(size=40)),\n        go.Scatter(y=[event.iloc[0].p2_pos_x], x=[event.iloc[0].p2_pos_y], mode='text', text=['🏎'], textfont=dict(size=40)),\n        go.Scatter(y=[event.iloc[0].p3_pos_x], x=[event.iloc[0].p3_pos_y], mode='text', text=['🚙'], textfont=dict(size=40)),\n        go.Scatter(y=[event.iloc[0].p4_pos_x], x=[event.iloc[0].p4_pos_y], mode='text', text=['🚙'], textfont=dict(size=40)),\n        go.Scatter(y=[event.iloc[0].p5_pos_x], x=[event.iloc[0].p5_pos_y], mode='text', text=['🚙'], textfont=dict(size=40)),\n        go.Scatter(y=[event.iloc[0].ball_pos_x], x=[event.iloc[0].ball_pos_y], mode='text', text=['⚽️'], textfont=dict(size=80))\n    ],\n    layout=go.Layout(\n        xaxis=dict(range=[-120, 120], autorange=False),\n        yaxis=dict(range=[-85, 85], autorange=False),\n        title=\"Team B scoring\",\n        updatemenus=[\n            dict(buttons = [\n                dict(\n                    args = [None, {\"frame\": {\"duration\": 20, \"redraw\": False},\n                                   \"fromcurrent\": False, \n                                   \"transition\": {\"duration\": 100}}],\n                    label = \"Play\",\n                    method = \"animate\")\n            ],\n            type='buttons',\n            showactive=False,\n            y=1.2,\n            x=0.5,\n            xanchor='right',\n            yanchor='top')\n        ]\n    ),\n    frames=frames\n)\n\nfig.update_layout(\n    plot_bgcolor=bg_color,\n    xaxis=dict(showgrid=False, zeroline=False, showticklabels=False),\n    yaxis=dict(showgrid=False, zeroline=False, showticklabels=False),\n    showlegend=False,\n    height=700,\n)\n\nfig.add_layout_image(\n    dict(\n        source=\"https://pbs.twimg.com/media/CIZ_-2vUwAAthQq.jpg\",\n        xref=\"paper\",\n        yref=\"paper\",\n        x=-0.05,\n        y=1.05,\n        sizex=1.1,\n        sizey=1.1,\n        sizing=\"stretch\",\n        opacity=0.5,\n        layer=\"below\"\n    )\n)\n\nfig","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:53:30.934359Z","iopub.execute_input":"2022-10-04T19:53:30.935855Z","iopub.status.idle":"2022-10-04T19:53:39.879879Z","shell.execute_reply.started":"2022-10-04T19:53:30.935803Z","shell.execute_reply":"2022-10-04T19:53:39.878548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target Relationships\n\n**Insight**\n\n* Only ball position on Y axes carries information whether a score is about to happen.\n* Most of the games are near the mid of the field see `ball_pos_x`","metadata":{"execution":{"iopub.status.busy":"2022-10-01T05:24:51.826614Z","iopub.execute_input":"2022-10-01T05:24:51.826996Z","iopub.status.idle":"2022-10-01T05:24:51.909503Z","shell.execute_reply.started":"2022-10-01T05:24:51.826964Z","shell.execute_reply":"2022-10-01T05:24:51.907254Z"}}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=3, ncols=2, figsize=(16, 12))\n\ndf_near_to_score = train_df[\n    train_df.team_A_scoring_within_10sec.gt(0) |\n    train_df.team_B_scoring_within_10sec.gt(0)\n]\n\ndf_other_time = train_df[~(\n    train_df.team_A_scoring_within_10sec.gt(0) |\n    train_df.team_B_scoring_within_10sec.gt(0)\n)]\n\ncols = ['ball_pos_x', 'ball_pos_y', 'ball_pos_z']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_near_to_score,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[0],\n        stat='density'\n    )\n    ax[0].set_title(cols[i] + ' near to score')\n\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_other_time,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[1],\n        stat='density'\n    )\n    ax[1].set_title(cols[i] + ' entire game')\n    \nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:53:39.881314Z","iopub.execute_input":"2022-10-04T19:53:39.881691Z","iopub.status.idle":"2022-10-04T19:53:47.283510Z","shell.execute_reply.started":"2022-10-04T19:53:39.881649Z","shell.execute_reply":"2022-10-04T19:53:47.282450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What about velocity, I suspect ball's velocity should increase when a score is about to happen\n\n**Insight**\n\n* Indeed. Velocity increases on the y axes when a team is about to score.","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=3, ncols=2, figsize=(16, 12))\n\ncols = ['ball_vel_x', 'ball_vel_y', 'ball_vel_z']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_near_to_score,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[0],\n        stat='density'\n    )\n    ax[0].set_title(cols[i] + ' near to score')\n\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_other_time,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[1],\n        stat='density'\n    )\n    ax[1].set_title(cols[i] + ' not near to score')\n    \nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:53:47.285110Z","iopub.execute_input":"2022-10-04T19:53:47.286357Z","iopub.status.idle":"2022-10-04T19:53:58.559375Z","shell.execute_reply.started":"2022-10-04T19:53:47.286299Z","shell.execute_reply":"2022-10-04T19:53:58.558019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Insight**\n\n* It's pretty common that positions of your team are related to the score, I'm only anlyzing `y` as previos plots suggested.\n* Distributions are different when your team is attacking than when your team is defending\n* p0, p1 and p2 belong to one team and p3, p4 and p5 belong to the other team.","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=3, ncols=2, figsize=(16, 12))\n\ncols = ['p0_pos_y', 'p1_pos_y', 'p2_pos_y']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_near_to_score,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[0],\n        stat='density'\n    )\n    ax[0].set_title(cols[i] + ' near to score')\n\ncols = ['p3_pos_y', 'p4_pos_y', 'p5_pos_y']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_near_to_score,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[1],\n        stat='density'\n    )\n    ax[1].set_title(cols[i] + ' near to score')\n    \nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:53:58.561286Z","iopub.execute_input":"2022-10-04T19:53:58.562492Z","iopub.status.idle":"2022-10-04T19:54:02.890273Z","shell.execute_reply.started":"2022-10-04T19:53:58.562436Z","shell.execute_reply":"2022-10-04T19:54:02.888864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Insight**\n\n* Velocities shows a different distribution than position but have the same interpretation.\n* We should model the interaction between velocity and position for the `y` axes.","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=3, ncols=2, figsize=(16, 12))\n\ndf_near_to_score = train_df[\n    train_df.team_A_scoring_within_10sec.gt(0) |\n    train_df.team_B_scoring_within_10sec.gt(0)\n]\n\ncols = ['p0_vel_y', 'p1_vel_y', 'p2_vel_y']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_near_to_score,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[0],\n        stat='density'\n    )\n    ax[0].set_title(cols[i] + ' near to score')\n\ncols = ['p3_vel_y', 'p4_vel_y', 'p5_vel_y']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_other_time,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[1],\n        stat='density'\n    )\n    ax[1].set_title(cols[i] + ' not near to score')\n    \nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:54:02.895089Z","iopub.execute_input":"2022-10-04T19:54:02.895659Z","iopub.status.idle":"2022-10-04T19:54:10.739101Z","shell.execute_reply.started":"2022-10-04T19:54:02.895605Z","shell.execute_reply":"2022-10-04T19:54:10.738118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nLet's compare what are the current boost reserve in a general situation vs a near-to-score, only plotting one team as therer are high chances that the other team also have the same distribution. TODO: validate this using t-test\n\n**Insight**\n\n* No clear difference","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=3, ncols=2, figsize=(16, 12))\n\ndf_near_to_score = train_df[\n    train_df.team_A_scoring_within_10sec.gt(0) |\n    train_df.team_B_scoring_within_10sec.gt(0)\n]\n\ncols = ['p0_boost', 'p1_boost', 'p2_boost']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_near_to_score,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[0],\n        bins=50,\n        stat='density'\n    )\n    ax[0].set_title(cols[i] + ' near to score')\n\ncols = ['p0_boost', 'p1_boost', 'p2_boost']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_other_time,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[1],\n        bins=50,\n        stat='density'\n    )\n    ax[1].set_title(cols[i] + ' not near to score')\n    \nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:54:10.740410Z","iopub.execute_input":"2022-10-04T19:54:10.741353Z","iopub.status.idle":"2022-10-04T19:54:17.111562Z","shell.execute_reply.started":"2022-10-04T19:54:10.741314Z","shell.execute_reply":"2022-10-04T19:54:17.109741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Did it occurr to you that boost spawning time is different when the team is about to score that when nothing is happening?, Apparently they follow the same distribution","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=3, ncols=2, figsize=(16, 12))\n\ndf_near_to_score = train_df[\n    train_df.team_A_scoring_within_10sec.gt(0) |\n    train_df.team_B_scoring_within_10sec.gt(0)\n]\n\ndf_other_time = train_df[~(\n    train_df.team_A_scoring_within_10sec.gt(0) |\n    train_df.team_B_scoring_within_10sec.gt(0)\n)]\n\ncols = ['boost0_timer', 'boost1_timer', 'boost2_timer']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_near_to_score,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[0],\n        bins=10,\n        stat='density'\n    )\n    ax[0].set_title(cols[i] + ' near to score')\n\ncols = ['boost0_timer', 'boost1_timer', 'boost2_timer']\nfor i, ax in enumerate(axs):\n    sns.histplot(\n        data=df_other_time,\n        x=cols[i],\n        hue='team_scoring_next',\n        ax=ax[1],\n        bins=10,\n        stat='density'\n    )\n    ax[1].set_title(cols[i] + ' not near to score')\n    \nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:54:17.113367Z","iopub.execute_input":"2022-10-04T19:54:17.113831Z","iopub.status.idle":"2022-10-04T19:54:22.255128Z","shell.execute_reply.started":"2022-10-04T19:54:17.113779Z","shell.execute_reply":"2022-10-04T19:54:22.253706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Bi-variate\n\nFor the sake of completeness let's review the bivarite distribution of ball positions.\n\n**insight**\n* The closests you are to the goal, the probability increases.\n\n<font color='darkblue'> Feature Engineering idea </font>: Create a distance variable from ball_position to goal A or goal B.","metadata":{}},{"cell_type":"code","source":"train_df['label'] = train_df.team_A_scoring_within_10sec + train_df.team_B_scoring_within_10sec.replace(1, 2)\nwarnings.simplefilter(\"ignore\", UserWarning)\nfig, ax = plt.subplots(figsize=(16, 6))\nax = sns.scatterplot(\n    data=train_df.sample(frac=0.3),\n    y='ball_pos_x',\n    x='ball_pos_y',\n    marker='s',\n    s=40,\n    alpha=0.2,\n    hue='label',\n    palette=color_palette[:3],\n    legend='brief'\n);\n\nhandles, labels  =  ax.get_legend_handles_labels();\nax.legend(handles, ['Nobody Scores', 'Team A Scores', 'Team B Scores']);\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-04T19:54:22.256424Z","iopub.execute_input":"2022-10-04T19:54:22.256817Z","iopub.status.idle":"2022-10-04T19:54:37.181172Z","shell.execute_reply.started":"2022-10-04T19:54:22.256782Z","shell.execute_reply":"2022-10-04T19:54:37.179785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Nan Values\n\n\n**insight**\n* The position and the boost of some players can have nan value, why? Because [extermination](#https://earlygame.com/rocket-league/how-to-get-rocket-league-extermination-what-is-extermination) where a players hits another so fast that takes them out of the game.\n* Nan % is really slow. Filling won't have a critical Role\n","metadata":{}},{"cell_type":"code","source":"import matplotlib.patches as mpatches\n\nnan_count = []\ntotal_records = 0\nfor i in track(range(10)):\n    chunk_df = pd.read_parquet(\n        f'../input/tps-oct22-parquet-datasets/train_{i}.parquet.gzip', engine='fastparquet'\n    )\n    \n    nan_count.append(\n        chunk_df.isna().sum(axis=0)\n    )\n    \n    total_records += chunk_df.shape[0]\n\nfig, ax = plt.subplots(figsize=(16, 15))\ntotal_nan_count = sum(nan_count)/total_records\ntotal_nan_count = total_nan_count.to_frame('nan%')\ntotal_nan_count['non_nan%'] = 1 - total_nan_count['nan%']\ntotal_nan_count['total'] = 1 \ntotal_nan_count.reset_index(inplace=True)\n\nbar1 = sns.barplot(y=\"index\",  x=\"total\", data=total_nan_count, orient='h', color=color_palette[0])\nbar2 = sns.barplot(y=\"index\",  x=\"nan%\", data=total_nan_count, orient='h', color=color_palette[1])\n\ntop_bar = mpatches.Patch(color=color_palette[0], label='NonNan %')\nbottom_bar = mpatches.Patch(color=color_palette[1], label='Nan %')\nplt.legend(handles=[top_bar, bottom_bar])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:54:37.182799Z","iopub.execute_input":"2022-10-04T19:54:37.183359Z","iopub.status.idle":"2022-10-04T19:55:26.530324Z","shell.execute_reply.started":"2022-10-04T19:54:37.183293Z","shell.execute_reply":"2022-10-04T19:55:26.528386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Nan Values have predictive value\n\nIf a demolition happened, then one team has less players. Let's do a proportion test to validate that. To do so, the following code cycles to the entire data and extracts:\n\n1. Number of observations with demolition \n2. Counts of observations with demolition and score within 10 seconds.\n3. Number of observations without demolition\n4. Counts of observations without demolition and score within 10 seconds.\n\n\nGiven the amuont of data, we can asume that mean is normally distributed the central limit theorem and calculate z_scores","metadata":{}},{"cell_type":"code","source":"players_cols = [\n    'p0_pos_x',\n    'p0_pos_y', 'p0_pos_z', 'p0_vel_x', 'p0_vel_y', 'p0_vel_z', 'p0_boost',\n    'p1_pos_x', 'p1_pos_y', 'p1_pos_z', 'p1_vel_x', 'p1_vel_y', 'p1_vel_z',\n    'p1_boost', 'p2_pos_x', 'p2_pos_y', 'p2_pos_z', 'p2_vel_x', 'p2_vel_y',\n    'p2_vel_z', 'p2_boost', 'p3_pos_x', 'p3_pos_y', 'p3_pos_z', 'p3_vel_x',\n    'p3_vel_y', 'p3_vel_z', 'p3_boost', 'p4_pos_x', 'p4_pos_y', 'p4_pos_z',\n    'p4_vel_x', 'p4_vel_y', 'p4_vel_z', 'p4_boost', 'p5_pos_x', 'p5_pos_y',\n    'p5_pos_z', 'p5_vel_x', 'p5_vel_y', 'p5_vel_z', 'p5_boost'\n]\n\npop1_counts_nob = {}\npop2_counts_nob = {}\ntotal_population = {}\n\ntotal_records = 0\nfor i in range(10):\n    chunk_df = pd.read_parquet(\n        f'../input/tps-oct22-parquet-datasets/train_{i}.parquet.gzip', engine='fastparquet'\n    )\n    \n    chunk_df['near_score'] = chunk_df.team_A_scoring_within_10sec + chunk_df.team_B_scoring_within_10sec\n    score_maks = chunk_df.near_score.eq(1)\n    demolition_mask = chunk_df[players_cols].isnull().any(axis=1)\n    \n    pop1 = chunk_df.loc[demolition_mask]\n    pop2 = chunk_df.loc[~demolition_mask]\n    \n    pop1_counts_nob[i]={'counts': pop1.near_score.sum(), 'nobs':pop1.shape[0]}\n    pop2_counts_nob[i]={'counts': pop2.near_score.sum(), 'nobs':pop2.shape[0]}\n    total_population[i]={'counts': chunk_df.near_score.sum(), 'nobs': chunk_df.shape[0]}\n    \ndel chunk_df\ngc.collect()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:55:26.532288Z","iopub.execute_input":"2022-10-04T19:55:26.532943Z","iopub.status.idle":"2022-10-04T19:56:16.306108Z","shell.execute_reply.started":"2022-10-04T19:55:26.532869Z","shell.execute_reply":"2022-10-04T19:56:16.304090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Agregate counts and do the test\n\n**Insight** With the number of observations that big, we can conclue that probability of having a record with score within 10 seconds it's totally different when a players was demolished.","metadata":{}},{"cell_type":"code","source":"import scipy\n\ntotal_nobs_sum = sum([x['nobs'] for x in total_population.values()])\ntotal_counts_sum = sum([x['counts'] for x in total_population.values()])\nE = total_counts_sum / total_nobs_sum\n\nlabels = ['demolition_yes', 'demolition_no']\nprint(f'Global mean: {E:.3f}')\nprint()\n\nprint('label\\t\\t\\tz_score\\t\\tcount/nobs (mean)\\t\\tp_value')\n\nfor i, population in enumerate([pop1_counts_nob, pop2_counts_nob]):\n    p_nobs_sum = sum([x['nobs'] for x in population.values()])\n    p_counts_sum = sum([x['counts'] for x in population.values()])\n    p_mean = p_counts_sum / p_nobs_sum\n\n    z_value = (p_mean - E) / (np.sqrt(E * (1 - E)) / np.sqrt(p_nobs_sum))\n    p_value = 2*scipy.stats.norm.cdf(-abs(z_value))\n    print(\n        labels[i],\n        f'\\t\\t{z_value:.2f}',\n        f'\\t\\t{p_counts_sum}/{p_nobs_sum} : {p_counts_sum/p_nobs_sum:.3f}  ',\n        f'\\t{p_value:.4f}', sep=''\n    )\n    \ndel pop1\ndel pop2","metadata":{"execution":{"iopub.status.busy":"2022-10-04T19:56:16.308125Z","iopub.execute_input":"2022-10-04T19:56:16.309104Z","iopub.status.idle":"2022-10-04T19:56:16.354601Z","shell.execute_reply.started":"2022-10-04T19:56:16.309050Z","shell.execute_reply":"2022-10-04T19:56:16.353323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Adversarial Validation\n\n\nThe following piece of code is large, what it does is to read the entire train data and compares each columns with the test set using Kolmogorov Smirnov test.\n\n**Insight** Even though KS shows high statistics for `z` axes related variables. Visual exploration shows that the results are due to having strong precisions for floats.","metadata":{}},{"cell_type":"code","source":"# Read Data\nchunks = []\nfor i in track(range(10)):\n    chunk_df = pd.read_parquet(\n        f'../input/tps-oct22-parquet-datasets/train_{i}.parquet.gzip', engine='fastparquet'\n    )\n    \n    chunks.append(chunk_df[test_df.columns])\n    \ntotal_df = pd.concat(chunks)\n\ndel chunk_df\ndel chunks\ngc.collect()\n\n\n# Helper function to get cumualtive distribution function\ndef get_cdf(train_serie, test_serie):\n    bins = np.linspace(train_serie.min(), train_serie.max(), 200000)\n    train_serie = train_serie.dropna()\n    test_serie = test_serie.dropna()\n    \n    train_groups = pd.Series(pd.cut(train_serie, bins=bins, labels=bins[1:])).astype(float)\n    test_groups = pd.Series(pd.cut(test_serie, bins=bins, labels=bins[1:])).astype(float)\n    \n    train_cdf = train_serie.groupby(train_groups).size()\n    test_cdf = test_serie.groupby(test_groups).size()\n    \n    train_cdf.sort_index(inplace=True)\n    test_cdf.sort_index(inplace=True)\n    \n    train_cdf = (train_cdf / train_cdf.sum()).cumsum()\n    test_cdf = (test_cdf / test_cdf.sum()).cumsum()\n    \n    return train_cdf, test_cdf\n\n\n# # Calculate KS using scipy\nks_ = {}\nfor column in test_df.columns:\n    ks_[column] = scipy.stats.ks_2samp(total_df[column], test_df[column])\n    \ndisplay(pd.DataFrame.from_dict(ks_, orient='index').sort_values('statistic', ascending=False).head())\n\n# Plot columns with large KS-statistic\nks_columns = ['p0_pos_z', 'p1_pos_z', 'p2_pos_z', 'p3_pos_z', 'p4_pos_z', 'p5_pos_z']\nfigure = make_subplots(\n    1, 6, shared_xaxes=True,\n    subplot_titles=['cudf_' + x for x in ks_columns]\n)\n\nfor j, column in enumerate(ks_columns):\n    train_cdf, test_cdf = get_cdf(total_df[column], test_df[column])\n\n    figure.add_trace(\n        go.Scatter(\n            x=train_cdf.index, y=train_cdf, line=dict(color=color_palette[0]),\n            legendgroup='train',\n            name='train',\n            showlegend=j==0\n        ),\n        row=1, col=j+1\n    )\n    \n    figure.add_trace(\n        go.Scatter(\n            x=test_cdf.index, y=test_cdf, line=dict(color=color_palette[1]),\n            legendgroup='test',\n            name='test',\n            showlegend=j==0\n        ),\n        row=1, col=j+1,\n    )\n\nfigure.update_layout(\n    hovermode='x',\n    xaxis1=dict(range=(0.3395, 0.341)),\n    xaxis2=dict(range=(0.3395, 0.341)),\n    xaxis3=dict(range=(0.3395, 0.341)),\n    xaxis4=dict(range=(0.3395, 0.341)),\n    xaxis5=dict(range=(0.3395, 0.341)),\n    xaxis6=dict(range=(0.3395, 0.341)),\n)\n\nfigure.update_annotations(font_size=13)\n\ndel total_df\ngc.collect()\nfigure","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-10-04T19:56:16.356828Z","iopub.execute_input":"2022-10-04T19:56:16.357244Z","iopub.status.idle":"2022-10-04T20:02:00.729784Z","shell.execute_reply.started":"2022-10-04T19:56:16.357208Z","shell.execute_reply":"2022-10-04T20:02:00.728774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Label\n\n\nThis model can be seen as a multilable classification with three labels. 1. `nobody_scores`, `team_A_scores`, `team_B_scores`, coded an [0, 1, 2] respectively.\n\n**Insights**:\n- This as imabalaced dataset, we can use oversampling or undersampling techniques, but due to the high amount of data, I think is better to use undersampling.\n- class weights also can help too. [check this link](https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#train_a_model_with_class_weights)\n\nI'll be using class_weight to don't bother modifying the dataset.","metadata":{}},{"cell_type":"code","source":"train_df['label'] = train_df.team_A_scoring_within_10sec + train_df.team_B_scoring_within_10sec.replace(1, 2)\ntrain_df.label.value_counts(True).to_frame(name='label proportion')","metadata":{"execution":{"iopub.status.busy":"2022-10-04T18:10:45.652018Z","iopub.execute_input":"2022-10-04T18:10:45.652460Z","iopub.status.idle":"2022-10-04T18:10:45.689966Z","shell.execute_reply.started":"2022-10-04T18:10:45.652421Z","shell.execute_reply":"2022-10-04T18:10:45.688601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocessing:\n\nNN has the cool property of partial training this is going to be useful when improving upon a model or by loading different datasets in ram and process them one by one\n\nBased on the EDA I'm only going to keep:\n- `ball_pos_y`\n- `ball_pos_x`\n- `ball_vel_y`\n- `p{x}_pos_y`\n\n\nVariable to create:\n- `distance_from_goal` distance variable from ball position to the goal_a or goal_b\n","metadata":{}},{"cell_type":"code","source":"columns_to_keep = [\n    'ball_pos_y', 'ball_pos_x', 'ball_vel_y',\n    'p0_pos_y', 'p1_pos_y', 'p2_pos_y',\n    'p3_pos_y', 'p4_pos_y', 'p5_pos_y', \n]\n\ntrain_df_ = train_df[columns_to_keep].copy()\ntest_df_ = test_df[columns_to_keep].copy()\ntarget = pd.get_dummies(train_df['label'])\ntarget.columns = ['nobody_scores', 'team_A_scores', 'team_b_scores']\n\ngoal_A_location = (0, -100)\ngoal_B_location = (0, 100)\n\ndef get_distances(df):\n    df['ball_distance_to_goal_A'] = np.sqrt(\n        (df.ball_pos_x)**2 + (df.ball_pos_y + 100)**2\n    )\n\n    df['ball_distance_to_goal_B'] = np.sqrt(\n        (df.ball_pos_x)**2 + (df.ball_pos_y - 100)**2\n    )\n    \nget_distances(train_df_)\nget_distances(test_df_)","metadata":{"execution":{"iopub.status.busy":"2022-10-04T18:10:45.691584Z","iopub.execute_input":"2022-10-04T18:10:45.691992Z","iopub.status.idle":"2022-10-04T18:10:45.830334Z","shell.execute_reply.started":"2022-10-04T18:10:45.691955Z","shell.execute_reply":"2022-10-04T18:10:45.829055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Save To CSV\n\nYou can save this to parquet too, CSV is human readable","metadata":{}},{"cell_type":"code","source":"scaler = SKWrapper(StandardScaler(), variables=train_df_.columns.tolist())\ntrain_df_ = scaler.fit_transform(train_df_)\ntest_df_ = scaler.transform(test_df_)\ntrain_df_.fillna(0, inplace=True)\ntest_df_.fillna(0, inplace=True)\n\npd.concat([train_df_, target], axis=1).to_csv('./train_0_processed_and_scaled.csv', index=False)\ntest_df_.to_csv('./test_df_processed_and_scaled.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-10-04T18:10:45.831912Z","iopub.execute_input":"2022-10-04T18:10:45.832326Z","iopub.status.idle":"2022-10-04T18:11:09.423852Z","shell.execute_reply.started":"2022-10-04T18:10:45.832283Z","shell.execute_reply":"2022-10-04T18:11:09.422048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Keras Baseline\n\nSimple sequential model that predicts multi-class label using categorical cross_entropy loss","metadata":{"execution":{"iopub.status.busy":"2022-10-01T14:35:00.543402Z","iopub.execute_input":"2022-10-01T14:35:00.543742Z","iopub.status.idle":"2022-10-01T14:35:00.563927Z","shell.execute_reply.started":"2022-10-01T14:35:00.543715Z","shell.execute_reply":"2022-10-01T14:35:00.563321Z"}}},{"cell_type":"code","source":"os.mkdir('checkpoint')\ntf.random.set_seed(42)\n\ndef create_model():\n    model = tf.keras.models.Sequential([\n        tf.keras.layers.Dense(32, activation='relu'),\n        tf.keras.layers.BatchNormalization(),\n        tf.keras.layers.Dense(32, activation='relu'),\n        tf.keras.layers.BatchNormalization(),\n        tf.keras.layers.Dense(3, activation='softmax')\n    ])\n\n    model.compile(\n        tf.keras.optimizers.Adam(learning_rate=0.0001),\n        loss='categorical_crossentropy'\n    )\n    return model\n    \n\ncp_callback = tf.keras.callbacks.ModelCheckpoint(\n    filepath='./checkpoint/',\n    save_weights_only=True,\n    verbose=0\n)\n\nes_callback = tf.keras.callbacks.EarlyStopping(\n    monitor='val_loss',\n    mode='min',\n    patience=10,\n    restore_best_weights=True\n)\n\nmodel = create_model()\nhistory = model.fit(\n    train_df_, target,\n    epochs=20, batch_size=512,\n    callbacks=[cp_callback, es_callback,  tf.keras.callbacks.TerminateOnNaN()],\n    class_weight = {\n        0: 0.5,\n        1: 2,\n        2: 2\n    },\n    validation_split=0.2\n)\n\nmodel.save(\"baseline_model\")","metadata":{"execution":{"iopub.status.busy":"2022-10-04T18:11:09.426060Z","iopub.execute_input":"2022-10-04T18:11:09.427303Z","iopub.status.idle":"2022-10-04T18:12:56.330024Z","shell.execute_reply.started":"2022-10-04T18:11:09.427248Z","shell.execute_reply":"2022-10-04T18:12:56.328835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['loss'], label='loss', linewidth=1.5, marker='.')\nplt.plot(history.history['val_loss'], label='val_loss', linewidth=1.5, marker='.')\nplt.legend()\nplt.title('Train and Val Loss')","metadata":{"execution":{"iopub.status.busy":"2022-10-04T18:12:56.334586Z","iopub.execute_input":"2022-10-04T18:12:56.334963Z","iopub.status.idle":"2022-10-04T18:12:56.693077Z","shell.execute_reply.started":"2022-10-04T18:12:56.334929Z","shell.execute_reply":"2022-10-04T18:12:56.691471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## How to load your model\n\nAlways keep the latest version of your model, starting from a checkpoint saves a lot of time","metadata":{}},{"cell_type":"code","source":"model = tf.keras.models.load_model(\"baseline_model\")","metadata":{"execution":{"iopub.status.busy":"2022-10-04T18:12:56.694513Z","iopub.execute_input":"2022-10-04T18:12:56.694887Z","iopub.status.idle":"2022-10-04T18:12:57.152214Z","shell.execute_reply.started":"2022-10-04T18:12:56.694839Z","shell.execute_reply":"2022-10-04T18:12:57.150918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generate Submission File","metadata":{}},{"cell_type":"code","source":"preds = pd.DataFrame(\n    model.predict(test_df_),\n    columns=['nobody_scores', 'team_A_scoring_within_10sec', 'team_B_scoring_within_10sec'],\n    index=test_df.index\n)\n\npreds[['team_A_scoring_within_10sec', 'team_B_scoring_within_10sec']].to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-04T18:12:57.156016Z","iopub.execute_input":"2022-10-04T18:12:57.156398Z","iopub.status.idle":"2022-10-04T18:13:21.599072Z","shell.execute_reply.started":"2022-10-04T18:12:57.156352Z","shell.execute_reply":"2022-10-04T18:13:21.597785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# How to improve upon this Baseline?\n\n\n1. More Featue Engineering, I did not consider anything related to the number of players in the game (see null values on px_pos)\n2. Train with the rest of the data\n3. Do feature selection, for example Lasso + Logistic Regression can give ideas on where most important features are.\n4. Watch videos of Rocket League games. But keep in mind that pro players tend to play slightly different.","metadata":{}}]}