{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Final submission for **Tabular Playground Series - Oct 2022** \n\n# 1. Introduction\n\nThis notebook documents my work throughout Tabular Playground Series competition. The main objective is to collect ideas, different approaches, difficulties faced and briefly explain my thought process around this challenge. It is also an invitation to review the challenge and think about what was done correctly, what kind of mistakes were made and what could have been done in a different way in order to get better results. \n\n\n### Libraries","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n\nimport os\nimport gc\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport seaborn as sns\nsns.set(style='darkgrid', font_scale=1.6)\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\npd.options.display.max_columns = None\nfrom sklearn.model_selection import train_test_split\nimport random\nfrom sklearn.preprocessing import StandardScaler\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-10-31T18:20:25.816785Z","iopub.execute_input":"2022-10-31T18:20:25.817344Z","iopub.status.idle":"2022-10-31T18:20:33.847695Z","shell.execute_reply.started":"2022-10-31T18:20:25.817231Z","shell.execute_reply":"2022-10-31T18:20:33.846674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. EDA\n\nComments and dataset first impressions are presented in this notebook: [TPS-Oct22-Dataset First Impressions](https://www.kaggle.com/code/dariopaez/tps-oct22-dataset-first-impressions).\n\nBesides, these works have great information and insights about this challenge data:\n- [🚀 TPS Oct 22 - Rocket League EDA](https://www.kaggle.com/code/samuelcortinhas/tps-oct-22-rocket-league-eda/notebook#3.-EDA)  by Samuel Cortinhas\n- [TPS OCT 2022 - EDA + Simple Models](https://www.kaggle.com/code/craigmthomas/tps-oct-2022-eda-simple-models/notebook) by Craig Thomas \n\nThese presents a great exploratory analysis. \n\nFrom EDA we can extract some general conclusions:\n\n- Dataset needs to be compressed\n- Target is unbalanced (90% for no team scoring, 5% for team A scoring and 5% for team B scoring)\n- Missing values that needs to be filled are associated to players position, velocity and booster remaining\n\n\nAfter these first impressions I tried different approaches to balance the data. These are shown [here](https://www.kaggle.com/code/dariopaez/tps-oct22-dataset-first-impressions). Then, I spent quite some time trying different models and parameters using the new balanced dataset. However, these models gave me bad results despite changing their structure and parameters. Since competition submission are evaluated by LogLoss, better results were obtained without modifying the data distribution and, probably, the test set distribution is similar to the train set. This explanation is better explained and detailed in this work: [How to kill all your efforts? ](https://www.kaggle.com/code/alexryzhkov/how-to-kill-all-your-efforts) by Alexander Ryzhkov. Thanks to his work I finally understood that I was wasting time and needed to take a different approach. ","metadata":{}},{"cell_type":"markdown","source":"# 3. Data Processing\n\nFeature engineering and data compression were handled together: [TPS-Oct22-Data compression and FE](https://www.kaggle.com/code/dariopaez/tps-oct22-data-compression-and-fe). Through an iteration over the training set and saving each file into the compressed feather format. \n\n### Feature Engineering\nSome added features included:\n\n- Distance between Ball and both goals\n- Ball speed\n- Players speed\n- Distance between team A players and team B goal\n- Distance between team B players and team A goal\n- Players to ball distance\n\nDropped columns:\n\n- 'team_scoring_next'\n- 'player_scoring_next'\n- 'event_time'\n- 'event_id'\n- 'game_num'\n\n### Missing Values \n\nAfter trying different ideas to fill missing values I ended up filling all values with 0. Other approaches did not present noticeable better results and replacing with zeros was easier. \n\nIn brief, when a player is demolished, his position, speed and remaining booster is set to 0 as if the player is in the center of the field. \n\n### Useful notebooks\n\n- Another compressing solution using Dask: [🚀How to load 21M rows in 1 minute using 2 lines](https://www.kaggle.com/code/donatoriccio/how-to-load-21m-rows-in-1-minute-using-2-lines) by Donatio Riccio\n\n- FE and compressing with feather: [TPS-OCT-2022 - Simple TF](https://www.kaggle.com/code/eavelardev/tps-oct-2022-simple-tf/notebook) by Eduardo Avelar","metadata":{}},{"cell_type":"markdown","source":"# 4. Model\n\nModel with best result achieved has the following features:\n\n- Compressed dataset in feather format\n- Added computed distances and speed features \n- Missing values replaced by 0s \n- Data transformation to make each column distribution to have mean 0 and standard deviation 1\n- 4 Dense layers with Relu activation\n- Added Normalization and Dropout layers\n- 2 Output dense neurons with sigmoid activation\n\nDespite **compressing** the data, I could not train a model with the whole dataset without running **out of memory**. Because of that, I made a loop to mix 5% samples from each train file and use that **mixed file** to train each model. With another loop, I generated different versions of that mixed train file and train several models. After training for one epoch I used the model to make predictions on the test set and saved results into csv files. The idea was to compute the mean of several predictions and analyze if that process gave better results than training only one model for more epochs. ","metadata":{}},{"cell_type":"markdown","source":"### Cells below are to pre-process the test set and iteration to obtain several predictions","metadata":{}},{"cell_type":"code","source":"data_path = '../input/tps-oct-22-feather'\nsubmissions_path = '../input/rocket-league-submissions'","metadata":{"execution":{"iopub.status.busy":"2022-10-31T18:20:33.850412Z","iopub.execute_input":"2022-10-31T18:20:33.851290Z","iopub.status.idle":"2022-10-31T18:20:33.860370Z","shell.execute_reply.started":"2022-10-31T18:20:33.851245Z","shell.execute_reply":"2022-10-31T18:20:33.856283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#List of functions\n\ndef demolitions(df):\n    df['active_players_A'] = 3-(df['p0_pos_x'].isna()).astype(int)-(df['p1_pos_x'].isna()).astype(int)-(df['p2_pos_x'].isna()).astype(int)\n    df['active_players_B'] = 3-(df['p3_pos_x'].isna()).astype(int)-(df['p4_pos_x'].isna()).astype(int)-(df['p5_pos_x'].isna()).astype(int)\n    return df\n\n#Computes distances, speeds and also calculates the maximum distance to ball in order to replace NaN values in players positions\ndef compute_df(df):\n    # Estimates\n    goalA_coord = (0,-103.5, 6)\n    goalB_coord = (0,103.5, 6)\n    df.fillna(0, inplace=True)\n    \n    # Euclidean distance. Ball to goals distances\n    df['ball_dist_to_goalA'] = np.sqrt((df['ball_pos_x']-goalA_coord[0])**2 + (df['ball_pos_y']-goalA_coord[1])**2 + (df['ball_pos_z']-goalA_coord[2])**2)\n    df['ball_dist_to_goalB'] = np.sqrt((df['ball_pos_x']-goalB_coord[0])**2 + (df['ball_pos_y']-goalB_coord[1])**2 + (df['ball_pos_z']-goalB_coord[2])**2)\n    \n    #Ball speed\n    df['ball_speed'] = np.sqrt(df['ball_vel_x']**2 + df['ball_vel_y']**2 + df['ball_vel_z']**2) \n    \n    #Players speeds\n    for i in range(6):\n        df[f'p{i}_speed'] = np.sqrt(df[f'p{i}_vel_x']**2 + df[f'p{i}_vel_y']**2 + df[f'p{i}_vel_z']**2)\n        \n        \n    #Team A distance to Goal B\n    for i in range(3):\n        df[f'p{i}_dist_to_goal'] = np.sqrt((df[f'p{i}_pos_x'] - goalB_coord[0])**2 +\n                                           (df[f'p{i}_pos_y'] - goalB_coord[1])**2 +\n                                           (df[f'p{i}_pos_z'] - goalB_coord[2])**2)\n        \n        \n    #Team B distance to Goal A\n    for i in range(3, 6):\n        df[f'p{i}_dist_to_goal'] = np.sqrt((df[f'p{i}_pos_x'] - goalA_coord[0])**2 +\n                                           (df[f'p{i}_pos_y'] - goalA_coord[1])**2 +\n                                           (df[f'p{i}_pos_z'] - goalA_coord[2])**2)\n        \n        \n    #Player to Ball distance\n    \n    for i in range(6):\n        df[f'p{i}_dist_to_ball'] = np.sqrt((df[f'p{i}_pos_x']-df['ball_pos_x'])**2 + (df[f'p{i}_pos_y']-df['ball_pos_y'])**2 + (df[f'p{i}_pos_z']-df['ball_pos_z'])**2)\n        \n        \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-10-31T18:20:33.862506Z","iopub.execute_input":"2022-10-31T18:20:33.863568Z","iopub.status.idle":"2022-10-31T18:20:33.888771Z","shell.execute_reply.started":"2022-10-31T18:20:33.863514Z","shell.execute_reply":"2022-10-31T18:20:33.887329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# #Pre-Process test set\n# test_df = pd.read_feather('../input/tps-oct-22-feather/test.feather')\n# test_df = demolitions(test_df)\n# test_dt = compute_df(test_df)\n# test_df.drop(columns=['id'], inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T18:20:33.891014Z","iopub.execute_input":"2022-10-31T18:20:33.892221Z","iopub.status.idle":"2022-10-31T18:20:33.907908Z","shell.execute_reply.started":"2022-10-31T18:20:33.892172Z","shell.execute_reply":"2022-10-31T18:20:33.906492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # 10 times loop in order to obtain ten submission files from different models and different mixed data\n\n# for i in range(10):\n    \n#     #Get a train mix file with a random 5% of each train file\n#     train_dfs = []\n#     for j in range(10):\n#         print(f'############## Writing file {j} ##############')\n#         train_file = os.path.join(data_path, f'train_{j}.feather')\n#         train = pd.read_feather(train_file).sample(frac=0.05)\n\n#         #Append each file\n#         train_dfs.append(train)\n\n#     pd.concat(train_dfs).reset_index(drop=True).to_feather('train_mix.feather')\n#     del train_dfs\n#     gc.collect()\n    \n#     #Load train mix\n#     train_df = pd.read_feather('./train_mix.feather')\n    \n#     #Drop game_num\n#     train_df.drop(columns=['game_num'], inplace=True)\n    \n#     #Define targets\n#     targets_df = pd.DataFrame()\n#     targets_df['team_A'] = train_df.pop('team_A_scoring_within_10sec')\n#     targets_df['team_B'] = train_df.pop('team_B_scoring_within_10sec')\n#     print('team_A mean: ', targets_df['team_A'].mean())\n#     print('team_B mean: ', targets_df['team_B'].mean())\n    \n#     #Standarization\n#     scaler = StandardScaler()\n#     scaler_cols = [x for x in train_df.columns if x not in ['active_players_A','active_players_B']]\n\n#     train_estandar = pd.DataFrame(scaler.fit_transform(train_df[scaler_cols]))\n#     train_estandar.columns = scaler_cols\n#     train_estandar.index = train_df.index\n\n#     train_active_players = train_df.drop(columns=scaler_cols)\n#     train_final = pd.concat([train_active_players, train_estandar], axis=1)\n    \n#     test_estandar = pd.DataFrame(scaler.transform(test_df[scaler_cols]))\n#     test_estandar.columns = scaler_cols\n\n#     test_active_players = test_df.drop(columns=scaler_cols)\n#     test_final = pd.concat([test_active_players, test_estandar], axis=1)\n    \n#     del train_active_players, train_estandar, train_df, test_estandar, test_active_players\n#     gc.collect()\n    \n#     #Model Training\n#     model = Sequential([\n#     tf.keras.layers.BatchNormalization(input_shape = [len(train_final.columns)]),\n#     tf.keras.layers.Dense(128, activation='relu'),\n    \n#     tf.keras.layers.BatchNormalization(),\n#     tf.keras.layers.Dropout(0.05),\n#     tf.keras.layers.Dense(256, activation='relu'),\n    \n#     tf.keras.layers.BatchNormalization(),\n#     tf.keras.layers.Dropout(0.1),\n#     tf.keras.layers.Dense(256, activation='relu'),\n \n#     tf.keras.layers.BatchNormalization(),\n#     tf.keras.layers.Dropout(0.15),\n#     tf.keras.layers.Dense(128, activation='relu'),    \n    \n#     tf.keras.layers.BatchNormalization(),\n#     tf.keras.layers.Dropout(0.2),\n#     tf.keras.layers.Dense(2, activation='sigmoid')\n#     ])\n    \n#     model.compile(optimizer=tf.keras.optimizers.Adam(),\n#              loss='binary_crossentropy',\n#              metrics=['accuracy'])\n    \n#     print(f'############## Training model number {i} ##############')\n#     history = model.fit(train_final, \n#                     targets_df, \n#                     epochs=1, \n#                     batch_size=32)\n    \n#     #model prediction\n#     model.trainable = False\n#     model.compile(optimizer=tf.keras.optimizers.Adam(),\n#              loss='binary_crossentropy',\n#              metrics=['accuracy'])\n    \n#     preds = model.predict(test_final)\n    \n#     ss = pd.read_csv('../input/tabular-playground-series-oct-2022/sample_submission.csv')\n#     ss['team_A_scoring_within_10sec'] = preds[:,0]\n#     ss['team_B_scoring_within_10sec'] = preds[:,1]\n#     ss.to_csv(f'Submission_{i}.csv', index=False)\n    \n#     print(ss['team_A_scoring_within_10sec'].mean())\n#     print(ss['team_B_scoring_within_10sec'].mean())\n    \n#     del model, preds, test_final, train_final, targets_df\n#     gc.collect()\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T18:20:33.909844Z","iopub.execute_input":"2022-10-31T18:20:33.910703Z","iopub.status.idle":"2022-10-31T18:20:33.922605Z","shell.execute_reply.started":"2022-10-31T18:20:33.910655Z","shell.execute_reply":"2022-10-31T18:20:33.920985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Final submission","metadata":{}},{"cell_type":"code","source":"submission_df = pd.DataFrame()\n\nfor i in range(60):\n    sub = pd.read_csv(f'../input/rocket-league-submissions/Submission_{i}.csv', encoding='utf-8')\n    \n    submission_df = pd.concat([submission_df, sub.iloc[:,1:]], axis=1)\n\n#submission_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T18:20:33.925474Z","iopub.execute_input":"2022-10-31T18:20:33.926632Z","iopub.status.idle":"2022-10-31T18:21:02.033500Z","shell.execute_reply.started":"2022-10-31T18:20:33.926586Z","shell.execute_reply":"2022-10-31T18:21:02.030864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ss = pd.read_csv('../input/tabular-playground-series-oct-2022/sample_submission.csv')\nss['team_A_scoring_within_10sec'] = submission_df['team_A_scoring_within_10sec'].mean(axis=1)\nss['team_B_scoring_within_10sec'] = submission_df['team_B_scoring_within_10sec'].mean(axis=1)\nss.to_csv(f'Submission.csv', index=False)\n\nprint(ss['team_A_scoring_within_10sec'].mean())\nprint(ss['team_B_scoring_within_10sec'].mean())","metadata":{"execution":{"iopub.status.busy":"2022-10-31T18:21:02.035347Z","iopub.execute_input":"2022-10-31T18:21:02.036867Z","iopub.status.idle":"2022-10-31T18:21:04.751658Z","shell.execute_reply.started":"2022-10-31T18:21:02.036815Z","shell.execute_reply":"2022-10-31T18:21:04.750049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Final Comments\n\n- I didn´t try **Data augmentation** solutions like:\n - Shuffle players in the same team\n - Flipping X axis (symmetrical)\n - Mirror Team A and Team B\n \n \n Some of these solutions could be applied with a custom Data Generator from Keras in this notebook: [TPS October 2022 - Keras/TF Neural Network](https://www.kaggle.com/code/fabianbong/tps-october-2022-keras-tf-neural-network) by Fabian Bong\n \n- Since training data was grouped by game and event, a correct way to split into train and validation sets would be using Group K Fold. This is explained here: [➡️Understanding the data + why you need GroupKFold](https://www.kaggle.com/code/donatoriccio/understanding-the-data-why-you-need-groupkfold) by Donato Riccio.\n\n- Calibration to correct the effect of sampling data:\n[Calibration is all you need!](https://www.kaggle.com/code/alexryzhkov/calibration-is-all-you-need/notebook) by Alexander Ryzhkov\n","metadata":{}}]}