{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nsns.set(style='darkgrid', font_scale=1.6)\nimport matplotlib.pyplot as plt\n%matplotlib inline\npd.options.display.max_columns = None\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-29T20:28:14.082184Z","iopub.execute_input":"2022-10-29T20:28:14.082792Z","iopub.status.idle":"2022-10-29T20:28:14.825405Z","shell.execute_reply.started":"2022-10-29T20:28:14.082669Z","shell.execute_reply":"2022-10-29T20:28:14.824004Z"},"jupyter":{"outputs_hidden":true,"source_hidden":true},"collapsed":true,"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Introduction\n\nThe objective of this notebook is to present some first insights and comments about **October 22 Tabular Playground Series**. The challenge this month is about [**Rocket League**](http://https://www.rocketleague.com/es-es/) game. \n\nThis challenge presents a big dataset (**10 GB**) consisting of sequences of snapshots of different Rocket League matches. These snapshots include position and velocity of all players and the ball. Besides, some extra information is given like players remaining boost.\n\nMain goal is to predict, from a given snapshot, for each team, the probability that they will score within the next 10 secondos of game time. \n\n## Disclaimer\n\nThe starting point of this notebook is the EDA made by [Samuel Cortinhas](https://www.kaggle.com/code/samuelcortinhas/tps-oct-22-rocket-league-eda/notebook). His work has great insights, very well organized information and nice graph presentations👍 ","metadata":{}},{"cell_type":"markdown","source":"# 2. Dataset\n\n## Preview\n\n- Train set is divided into 10 files of aroung 1 GB each.\n- Test set have a total of 701143 entries. \n- Dtypes are provided in order to optimize memory.\n- A tool for compressing data will be needed since dataset is very big. Loading time is slow for only one .csv file\n- In order to save memory and make a faster analysis we only use 200 thousand samples from the first train file.","metadata":{}},{"cell_type":"code","source":"# Data types\ndtypes_dict_train = dict(pd.read_csv('../input/tabular-playground-series-oct-2022/train_dtypes.csv').values)\ndtypes_dict_test = dict(pd.read_csv('../input/tabular-playground-series-oct-2022/test_dtypes.csv').values)\n\n# Data (only 10% of train set)\ntrain = pd.read_csv(\"/kaggle/input/tabular-playground-series-oct-2022/train_0.csv\", dtype=dtypes_dict_train)\ntest = pd.read_csv(\"../input/tabular-playground-series-oct-2022/test.csv\", dtype=dtypes_dict_test)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T20:28:22.886521Z","iopub.execute_input":"2022-10-29T20:28:22.887887Z","iopub.status.idle":"2022-10-29T20:29:10.567951Z","shell.execute_reply.started":"2022-10-29T20:28:22.887842Z","shell.execute_reply":"2022-10-29T20:29:10.566541Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Shape and preview\nprint('Train set shape:', train.shape)\ndisplay(train.head(3))\n\nprint('Test set shape:', test.shape)\ndisplay(test.head(3))","metadata":{"execution":{"iopub.status.busy":"2022-10-29T20:29:43.607299Z","iopub.execute_input":"2022-10-29T20:29:43.607776Z","iopub.status.idle":"2022-10-29T20:29:43.718829Z","shell.execute_reply.started":"2022-10-29T20:29:43.607736Z","shell.execute_reply":"2022-10-29T20:29:43.717489Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# To speed up analysis, we only use a subset of the data\ntrain = train.iloc[:200000]","metadata":{"execution":{"iopub.status.busy":"2022-10-29T20:29:54.279453Z","iopub.execute_input":"2022-10-29T20:29:54.279978Z","iopub.status.idle":"2022-10-29T20:29:54.285778Z","shell.execute_reply.started":"2022-10-29T20:29:54.279937Z","shell.execute_reply":"2022-10-29T20:29:54.284466Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Comments\n\n- Each team have three players. Players 0, 1 and 2 belong to team A and Players 3, 4 and 5 belong to team B.\n- For each player we have the following data:\n\n - Player position (x, y, z)\n - Player velocity (x, y, z)\n - Remaining boost \n \n \n- Train set is grouped by **game number** and **event id**. Events are a sequences of consecutive snapshots which end with a team scoring or with neither team scoring. Events with neither team scoring were truncated due to non-competitive situation (eg one team winning by 3+ goals with litle time remaining) or nearing end of regulation (game ending when ball touches the ground).\n- For each event id, snapshots are consecutive and ordered.\n- Test set does not have information about game number, event id or event time. Snapshots are shuffled. ","metadata":{}},{"cell_type":"markdown","source":"# 3. Analysis\n\n## Missing Values\n\n- There are two sources of missing values. \n- Most of them appear on **'team_scoring_next'** column. Missing values in this column represents that neither team will score during that event. \n- Other missing values appears when a certain player is **demolished** and needs to respawn in a few seconds. Each time a player is demolished all of his data is filled with missing values (position, velocity and booster). \n- Demolitions could lead to some advantage over the team with less active players.","metadata":{}},{"cell_type":"code","source":"train.isnull().sum().sort_values(ascending=False)[:30]","metadata":{"execution":{"iopub.status.busy":"2022-10-29T20:30:38.159347Z","iopub.execute_input":"2022-10-29T20:30:38.159828Z","iopub.status.idle":"2022-10-29T20:30:38.208844Z","shell.execute_reply.started":"2022-10-29T20:30:38.159792Z","shell.execute_reply":"2022-10-29T20:30:38.207657Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.isnull().sum().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T20:32:25.791791Z","iopub.execute_input":"2022-10-29T20:32:25.792295Z","iopub.status.idle":"2022-10-29T20:32:25.890742Z","shell.execute_reply.started":"2022-10-29T20:32:25.792256Z","shell.execute_reply":"2022-10-29T20:32:25.889520Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Demolitions\n\nWe can define two new variables to keep track of the number of active players in each team by using missing values in players variables. This function was created on Samuel Cortinhas notebook. ","metadata":{}},{"cell_type":"code","source":"def demolitions(df):\n    df['active_players_A'] = 3-(df['p0_pos_x'].isna()).astype(int)-(df['p1_pos_x'].isna()).astype(int)-(df['p2_pos_x'].isna()).astype(int)\n    df['active_players_B'] = 3-(df['p3_pos_x'].isna()).astype(int)-(df['p4_pos_x'].isna()).astype(int)-(df['p5_pos_x'].isna()).astype(int)\n    return df\n\ntrain = demolitions(train)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T21:02:38.124375Z","iopub.execute_input":"2022-10-29T21:02:38.124856Z","iopub.status.idle":"2022-10-29T21:02:38.143207Z","shell.execute_reply.started":"2022-10-29T21:02:38.124822Z","shell.execute_reply":"2022-10-29T21:02:38.141770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Generate new column: 'next_score'. \nfilters = [\n    train['team_A_scoring_within_10sec'] == train['team_B_scoring_within_10sec'],\n    train['team_A_scoring_within_10sec'] == 1,\n   train['team_B_scoring_within_10sec'] == 1\n]\nvalues = [\"None\", \"A\", \"B\"]\n\ntrain['next_score'] = np.select(filters, values)\n\nplt.figure(figsize=(10,5))\nsns.histplot(train, x='active_players_A', hue='next_score', multiple='dodge')\nplt.yscale('log')\nplt.xticks([1, 2, 3])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T21:10:39.078219Z","iopub.execute_input":"2022-10-29T21:10:39.078654Z","iopub.status.idle":"2022-10-29T21:10:40.139867Z","shell.execute_reply.started":"2022-10-29T21:10:39.078624Z","shell.execute_reply":"2022-10-29T21:10:40.138538Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If a team has less active players the other team has greater probability of scoring next. This effect is more evident when a team has only 1 active player.","metadata":{}},{"cell_type":"code","source":"pattern = '{:,.3f}'.format\nprint ('Active A players: \\n', train.active_players_A.value_counts(normalize=True).to_string(float_format=pattern))\nprint ('Active B players: \\n', train.active_players_B.value_counts(normalize=True).to_string(float_format=pattern))","metadata":{"execution":{"iopub.status.busy":"2022-10-29T21:22:51.187538Z","iopub.execute_input":"2022-10-29T21:22:51.188073Z","iopub.status.idle":"2022-10-29T21:22:51.201890Z","shell.execute_reply.started":"2022-10-29T21:22:51.188032Z","shell.execute_reply":"2022-10-29T21:22:51.200053Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"However, we can see that in only a 2% of the dataset teams are with two active players. Only 1 active player hardly ever happens. ","metadata":{}},{"cell_type":"markdown","source":"## Target variables\n\nFor both teams, around **6%** of the samples have a target of 1 (meaning the team will score within 10 seconds). Since events end with a team scoring, target column have value of 1 only for samples representing the last ten seconds of that event. Because of that, there are a much more entries with target 0 than target 1. This results in an **unbalanced dataset**.  ","metadata":{}},{"cell_type":"code","source":"#Number of 0 and 1 for column: 'team_A_scoring_whithin_10sec'\nteam_A_count = train['team_A_scoring_within_10sec'].value_counts(normalize=True)\n\nteam_A_data = [\n    {'score': value, 'percentage': percentage*100} for \n    value, percentage in dict(team_A_count).items()\n]\n\ndf_team_A = pd.DataFrame(team_A_data)\ndisplay(df_team_A.head(3))\n\n#Number of 0 and 1 for column: 'team_B_scoring_whithin_10sec'\nteam_B_count = train['team_B_scoring_within_10sec'].value_counts(normalize=True)\n\nteam_B_data = [\n    {'score': value, 'percentage': percentage*100} for \n    value, percentage in dict(team_B_count).items()\n]\n\ndf_team_B = pd.DataFrame(team_B_data)\ndisplay(df_team_B.head(3))\n\nplt.figure(figsize=(15, 6))\nplt.subplot(1,2,1)\nsns.barplot(data=df_team_A, x='score', y='percentage')\nplt.title('Target A')\nplt.ylabel('%')\n\nplt.subplot(1,2,2)\nsns.barplot(data=df_team_B, x='score', y='percentage')\nplt.title('Target B')\nplt.ylabel('%')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T22:18:57.617902Z","iopub.execute_input":"2022-10-29T22:18:57.618409Z","iopub.status.idle":"2022-10-29T22:18:57.871770Z","shell.execute_reply.started":"2022-10-29T22:18:57.618361Z","shell.execute_reply":"2022-10-29T22:18:57.870456Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unbalance becomes more evident by plotting a three category plot.","metadata":{}},{"cell_type":"code","source":"score_count = train['next_score'].value_counts(normalize=True)\n\nscore_data = [\n    {'team_scoring': value, 'percentage': percentage*100} for \n    value, percentage in dict(score_count).items()\n]\n\ndf_score_data = pd.DataFrame(score_data)\n\nplt.figure(figsize=(15, 6))\nsns.barplot(data=df_score_data, x='team_scoring', y='percentage')\nplt.ylabel('%')\nplt.xlabel('Team scoring Next')\nplt.title('Target')","metadata":{"execution":{"iopub.status.busy":"2022-10-29T22:19:18.080244Z","iopub.execute_input":"2022-10-29T22:19:18.080709Z","iopub.status.idle":"2022-10-29T22:19:18.346113Z","shell.execute_reply.started":"2022-10-29T22:19:18.080674Z","shell.execute_reply":"2022-10-29T22:19:18.344604Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Three different proposals obtain a more balanced dataset are presented.\n\n### First\n\nSelect events where A or B score and keep only the last ten seconds of that records.","metadata":{}},{"cell_type":"code","source":"#Temporal Filter\ntrain_filtered = train.drop(index=train[(train['team_scoring_next'].notnull()) & (train['event_time'] < -10)].index)\n\nscore_count = train_filtered['next_score'].value_counts(normalize=True)\n\nscore_data = [\n    {'team_scoring': value, 'percentage': percentage*100} for \n    value, percentage in dict(score_count).items()\n]\n\ndf_score_data = pd.DataFrame(score_data)\n\nplt.figure(figsize=(15, 6))\nsns.barplot(data=df_score_data, x='team_scoring', y='percentage')\nplt.ylabel('%')\nplt.xlabel('Team scoring Next')\nplt.title('Target')\n\ndisplay(df_score_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T22:19:31.407714Z","iopub.execute_input":"2022-10-29T22:19:31.409036Z","iopub.status.idle":"2022-10-29T22:19:31.791940Z","shell.execute_reply.started":"2022-10-29T22:19:31.408984Z","shell.execute_reply":"2022-10-29T22:19:31.790996Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Second\n\nTake all events and filter them to the last 10 seconds.","metadata":{}},{"cell_type":"code","source":"#Temporal Filter\ntrain_filtered = train.drop(index=train[train['event_time'] < -10].index)\n\nscore_count = train_filtered['next_score'].value_counts(normalize=True)\n\nscore_data = [\n    {'team_scoring': value, 'percentage': percentage*100} for \n    value, percentage in dict(score_count).items()\n]\n\ndf_score_data = pd.DataFrame(score_data)\n\nplt.figure(figsize=(15, 6))\nsns.barplot(data=df_score_data, x='team_scoring', y='percentage')\nplt.ylabel('%')\nplt.xlabel('Team scoring Next')\nplt.title('Target')\n\ndisplay(df_score_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T22:19:37.590025Z","iopub.execute_input":"2022-10-29T22:19:37.590455Z","iopub.status.idle":"2022-10-29T22:19:37.966551Z","shell.execute_reply.started":"2022-10-29T22:19:37.590424Z","shell.execute_reply":"2022-10-29T22:19:37.965332Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Third\n\nEvents with team A or B scoring are reduced to the final 10 seconds.\n\nEvents with no team scoring are reduced to the last 30 seconds.","metadata":{}},{"cell_type":"code","source":"#Temporal Filter\ntrain_filtered = train.drop(index=train[(train['team_scoring_next'].notnull()) & (train['event_time'] < -10)].index)\ntrain_filtered.drop(index=train[(train['team_scoring_next'].isnull()) & (train['event_time'] < -30)].index, inplace=True)\n\nscore_count = train_filtered['next_score'].value_counts(normalize=True)\n\nscore_data = [\n    {'team_scoring': value, 'percentage': percentage*100} for \n    value, percentage in dict(score_count).items()\n]\n\ndf_score_data = pd.DataFrame(score_data)\n\nplt.figure(figsize=(15, 6))\nsns.barplot(data=df_score_data, x='team_scoring', y='percentage')\nplt.ylabel('%')\nplt.xlabel('Team scoring Next')\nplt.title('Target')\n\ndisplay(df_score_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T22:19:44.747814Z","iopub.execute_input":"2022-10-29T22:19:44.748333Z","iopub.status.idle":"2022-10-29T22:19:45.158901Z","shell.execute_reply.started":"2022-10-29T22:19:44.748292Z","shell.execute_reply":"2022-10-29T22:19:45.157448Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]}]}