{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This month's challenge is to predict the probability of each team scoring within the next 10 seconds of the game given a snapshot from a Rocket League match. Sounds awesome, right?\n\nWell, it's not that simple. The training data is fairly large; trying to read and model it in a single go might pose some challenges. The purpose of this month's competition is for you to explore ways you can take a big dataset and make it manageable within the time and resources you have. For most people, typical brute force approaches aren't going to work well.\n\n- Can you scale down the dataset?\n- Can you use, e.g., online learning methods that allow you to train from the data one row at a time? (FTLR is a great place to start if you're not familiar with online learning! e.g., this notebook)\n- Can you figure out a nice set of features to reduce the dataset down to?\n\nIn addition to that challenge, while your predictions must be made pointwise, the training data is made up of timeseries—maybe you can use that temporal information to improve your model? This competition also has plenty of opportunity for data visualizations. Let's see some pretty graphs!\n\nSo, share your ideas about tackling this beast of a dataset and have a great time!","metadata":{}},{"cell_type":"markdown","source":"## Import Packages ","metadata":{}},{"cell_type":"code","source":"import numpy as np  # linear algebra\nimport pandas as pd  # data manipulation\nimport os  # file navigation\nimport gc  # garbage collection\n\n# visualization\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly import subplots\n\nfrom sklearn.model_selection import cross_validate  # k-fold Cross Validation\nfrom sklearn.preprocessing import LabelEncoder  # output binary encoding\n\nfrom xgboost import XGBClassifier  # Gradient Boosted Tree (XGBoost)\n\nfrom tensorflow.config import list_physical_devices  # check if GPU is available","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-22T16:10:18.505877Z","iopub.execute_input":"2022-10-22T16:10:18.506530Z","iopub.status.idle":"2022-10-22T16:10:26.501611Z","shell.execute_reply.started":"2022-10-22T16:10:18.506427Z","shell.execute_reply":"2022-10-22T16:10:26.500650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Config","metadata":{}},{"cell_type":"code","source":"# training and cross validation\ngpu = list_physical_devices('GPU') != []\nn_estimators = 1000\nmax_depth = 5\nlearning_rate = 0.001\nfolds = 5\n\n# data loading\ndebug = False\nsample = 0.3\nseed = 42","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:10:43.731929Z","iopub.execute_input":"2022-10-22T16:10:43.733019Z","iopub.status.idle":"2022-10-22T16:10:43.811331Z","shell.execute_reply.started":"2022-10-22T16:10:43.732979Z","shell.execute_reply":"2022-10-22T16:10:43.810237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset","metadata":{}},{"cell_type":"markdown","source":"Because of memory limitations, we are not gonna use all the data from the train.csv file. Instead, we'll get a random sample of each file (aprox 20-33% sample size), and combine these samples into a unique training dataset.","metadata":{}},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-22T16:10:52.232614Z","iopub.execute_input":"2022-10-22T16:10:52.233053Z","iopub.status.idle":"2022-10-22T16:10:52.247429Z","shell.execute_reply.started":"2022-10-22T16:10:52.233016Z","shell.execute_reply":"2022-10-22T16:10:52.246514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ncol_dtypes = {\n    'game_num': 'int8', 'event_id': 'int8', 'event_time': 'float16',\n    'ball_pos_x': 'float16', 'ball_pos_y': 'float16', 'ball_pos_z': 'float16',\n    'ball_vel_x': 'float16', 'ball_vel_y': 'float16', 'ball_vel_z': 'float16',\n    'p0_pos_x': 'float16', 'p0_pos_y': 'float16', 'p0_pos_z': 'float16',\n    'p0_vel_x': 'float16', 'p0_vel_y': 'float16', 'p0_vel_z': 'float16',\n    'p0_boost': 'float16', 'p1_pos_x': 'float16', 'p1_pos_y': 'float16',\n    'p1_pos_z': 'float16', 'p1_vel_x': 'float16', 'p1_vel_y': 'float16',\n    'p1_vel_z': 'float16', 'p1_boost': 'float16', 'p2_pos_x': 'float16',\n    'p2_pos_y': 'float16', 'p2_pos_z': 'float16', 'p2_vel_x': 'float16',\n    'p2_vel_y': 'float16', 'p2_vel_z': 'float16', 'p2_boost': 'float16',\n    'p3_pos_x': 'float16', 'p3_pos_y': 'float16', 'p3_pos_z': 'float16',\n    'p3_vel_x': 'float16', 'p3_vel_y': 'float16', 'p3_vel_z': 'float16',\n    'p3_boost': 'float16', 'p4_pos_x': 'float16', 'p4_pos_y': 'float16',\n    'p4_pos_z': 'float16', 'p4_vel_x': 'float16', 'p4_vel_y': 'float16',\n    'p4_vel_z': 'float16', 'p4_boost': 'float16', 'p5_pos_x': 'float16',\n    'p5_pos_y': 'float16', 'p5_pos_z': 'float16', 'p5_vel_x': 'float16',\n    'p5_vel_y': 'float16', 'p5_vel_z': 'float16', 'p5_boost': 'float16',\n    'boost0_timer': 'float16', 'boost1_timer': 'float16', 'boost2_timer': 'float16',\n    'boost3_timer': 'float16', 'boost4_timer': 'float16', 'boost5_timer': 'float16',\n    'player_scoring_next': 'O', 'team_scoring_next': 'O', 'team_A_scoring_within_10sec': 'O',\n    'team_B_scoring_within_10sec': 'O'\n}\ncols = list(col_dtypes.keys())\n\npath_to_data = '../input/tabular-playground-series-oct-2022'\ndf = pd.DataFrame({}, columns=cols)\nfor i in range(10):\n    df_tmp = pd.read_csv(f'{path_to_data}/train_{i}.csv', dtype=col_dtypes)\n    if sample < 1:\n        df_tmp = df_tmp.sample(frac=sample, random_state=seed)\n        \n    df = pd.concat([df, df_tmp])\n    del df_tmp\n    gc.collect()\n    if debug:\n        break\n        \ndf","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:10:55.884258Z","iopub.execute_input":"2022-10-22T16:10:55.884642Z","iopub.status.idle":"2022-10-22T16:15:31.614698Z","shell.execute_reply.started":"2022-10-22T16:10:55.884609Z","shell.execute_reply":"2022-10-22T16:15:31.613588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Quick EDA","metadata":{}},{"cell_type":"code","source":"input_cols = [\n    'ball_pos_x', 'ball_pos_y', 'ball_pos_z', 'ball_vel_x', 'ball_vel_y', 'ball_vel_z', \n    'p0_pos_x', 'p0_pos_y', 'p0_pos_z', 'p0_vel_x', 'p0_vel_y', 'p0_vel_z', \n    'p1_pos_x', 'p1_pos_y', 'p1_pos_z', 'p1_vel_x', 'p1_vel_y', 'p1_vel_z',\n    'p2_pos_x', 'p2_pos_y', 'p2_pos_z', 'p2_vel_x', 'p2_vel_y', 'p2_vel_z',\n    'p3_pos_x', 'p3_pos_y', 'p3_pos_z', 'p3_vel_x', 'p3_vel_y', 'p3_vel_z',\n    'p4_pos_x', 'p4_pos_y', 'p4_pos_z', 'p4_vel_x', 'p4_vel_y', 'p4_vel_z',\n    'p5_pos_x', 'p5_pos_y', 'p5_pos_z', 'p5_vel_x', 'p5_vel_y', 'p5_vel_z',\n    'p0_boost', 'p1_boost',  'p2_boost', 'p3_boost', 'p4_boost', 'p5_boost',\n    'boost0_timer', 'boost1_timer', 'boost2_timer', 'boost3_timer', 'boost4_timer', 'boost5_timer'\n]\n\noutput_cols = ['team_A_scoring_within_10sec', 'team_B_scoring_within_10sec']","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:04.748892Z","iopub.execute_input":"2022-10-22T16:16:04.749258Z","iopub.status.idle":"2022-10-22T16:16:04.756008Z","shell.execute_reply.started":"2022-10-22T16:16:04.749227Z","shell.execute_reply":"2022-10-22T16:16:04.755004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Input Variables","metadata":{}},{"cell_type":"code","source":"def int_to_grid_coord(k, n):\n    return (k // n) + 1, (k % n) + 1","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:11.100318Z","iopub.execute_input":"2022-10-22T16:16:11.100745Z","iopub.status.idle":"2022-10-22T16:16:11.106471Z","shell.execute_reply.started":"2022-10-22T16:16:11.100711Z","shell.execute_reply":"2022-10-22T16:16:11.105256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_distributions(df, row_count, col_count, title, height):\n    features = df.columns\n    fig = subplots.make_subplots(\n        rows=row_count, cols=col_count,\n        subplot_titles=features\n    )\n\n    for k, col in enumerate(features):\n        i, j = int_to_grid_coord(k, col_count)\n\n        fig.add_trace(\n            go.Histogram(\n                x=df[col].astype('float32'),\n                name=col\n            ),\n            row=i, col=j\n        )\n\n    fig.update_layout(\n        title=title,\n        height=row_count * height,\n        showlegend=False\n    )\n\n    return fig","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:13.690825Z","iopub.execute_input":"2022-10-22T16:16:13.691500Z","iopub.status.idle":"2022-10-22T16:16:13.698457Z","shell.execute_reply.started":"2022-10-22T16:16:13.691461Z","shell.execute_reply":"2022-10-22T16:16:13.697354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_distributions(df[input_cols].sample(frac=0.0005), 9, 6, \"Input Variables Distributions\", 300)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:22.274793Z","iopub.execute_input":"2022-10-22T16:16:22.275180Z","iopub.status.idle":"2022-10-22T16:16:25.936394Z","shell.execute_reply.started":"2022-10-22T16:16:22.275146Z","shell.execute_reply":"2022-10-22T16:16:25.935225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Output Variables","metadata":{}},{"cell_type":"code","source":"plot_distributions(df[output_cols].sample(frac=0.0005), 1, 2, \"Output Variables Distributions\", height=600)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:46.749823Z","iopub.execute_input":"2022-10-22T16:16:46.750223Z","iopub.status.idle":"2022-10-22T16:16:47.139766Z","shell.execute_reply.started":"2022-10-22T16:16:46.750189Z","shell.execute_reply":"2022-10-22T16:16:47.138743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"Shoutout [this post by samuelcortinhas](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356852) for the idea.","metadata":{}},{"cell_type":"markdown","source":"Let's derive 2 new features:\n* For each player (and the ball) let's get their velocity's magnitude.\n* For each player, let's get their distance from the ball.","metadata":{}},{"cell_type":"markdown","source":"## Euclidian Norm","metadata":{}},{"cell_type":"markdown","source":"For these 2 new features, we'll need to get the 3D euclidian norm of a vector:\n$$ \\| \\overrightarrow{v} \\| = \\sqrt{x^2 + y^2 + z^2} $$\nWe'll use numpy's linalg.norm() method for that.","metadata":{}},{"cell_type":"code","source":"def euclidian_norm(x):\n    return np.linalg.norm(x, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:52.834694Z","iopub.execute_input":"2022-10-22T16:16:52.835248Z","iopub.status.idle":"2022-10-22T16:16:52.840269Z","shell.execute_reply.started":"2022-10-22T16:16:52.835210Z","shell.execute_reply":"2022-10-22T16:16:52.839123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let's group the x, y and z variables by player and ball categories\n# this simplifies the code for the euclidian norm calculation\nvel_groups = {\n    f\"{el}_vel\": [f'{el}_vel_x', f'{el}_vel_y', f'{el}_vel_z']\n    for el in ['ball'] + [f'p{i}' for i in range(6)]\n}\npos_groups = {\n    f\"{el}_pos\": [f'{el}_pos_x', f'{el}_pos_y', f'{el}_pos_z']\n    for el in ['ball'] + [f'p{i}' for i in range(6)]\n}\npos_groups","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:54.577241Z","iopub.execute_input":"2022-10-22T16:16:54.577596Z","iopub.status.idle":"2022-10-22T16:16:54.586058Z","shell.execute_reply.started":"2022-10-22T16:16:54.577567Z","shell.execute_reply":"2022-10-22T16:16:54.584942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# velocity magnitude\nfor col, vec in vel_groups.items():\n    df[col] = euclidian_norm(df[vec])","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:16:58.882932Z","iopub.execute_input":"2022-10-22T16:16:58.883941Z","iopub.status.idle":"2022-10-22T16:17:09.239777Z","shell.execute_reply.started":"2022-10-22T16:16:58.883888Z","shell.execute_reply":"2022-10-22T16:17:09.238605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# distance from ball\nfor col, vec in pos_groups.items():\n    df[col + \"_ball_dist\"] = euclidian_norm(df[vec].values - df[pos_groups[\"ball_pos\"]].values)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:17:26.088779Z","iopub.execute_input":"2022-10-22T16:17:26.089186Z","iopub.status.idle":"2022-10-22T16:17:33.069260Z","shell.execute_reply.started":"2022-10-22T16:17:26.089153Z","shell.execute_reply":"2022-10-22T16:17:33.068251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Cleaning","metadata":{}},{"cell_type":"markdown","source":"We drop the columns below because they should not influence the results.  \n- game_num, event_id and event_time are irrelevant.  \n- player_scoring_next and team_scoring next are a form of data leakage, as in they're synonymous with the output variable.  \n- ball_pos_ball_dist is always 0, the distance between the ball and itself.","metadata":{}},{"cell_type":"markdown","source":"## Dropping columns","metadata":{}},{"cell_type":"code","source":"cols_to_drop = [\n    'game_num', 'event_id', 'event_time', 'player_scoring_next', 'team_scoring_next', 'ball_pos_ball_dist'\n]\n\ndf = df.drop(columns=cols_to_drop)\ndf","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:17:44.046499Z","iopub.execute_input":"2022-10-22T16:17:44.046887Z","iopub.status.idle":"2022-10-22T16:17:46.913058Z","shell.execute_reply.started":"2022-10-22T16:17:44.046854Z","shell.execute_reply":"2022-10-22T16:17:46.912123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dropping rows containing NaN","metadata":{}},{"cell_type":"code","source":"has_na = {}\nfor col in df.columns:\n    has_na[col] = df[col].isnull().values.any()\n\nprint(\"Columns that contain null values:\")\nfor col in has_na:\n    if has_na[col]:\n        print(col)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:17:49.868980Z","iopub.execute_input":"2022-10-22T16:17:49.869693Z","iopub.status.idle":"2022-10-22T16:17:51.171036Z","shell.execute_reply.started":"2022-10-22T16:17:49.869657Z","shell.execute_reply":"2022-10-22T16:17:51.169909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"null_p0_pos_x_count = df['p0_pos_x'].isna().sum()\nnull_p0_pos_x_perc = null_p0_pos_x_count / df.shape[0]\nprint(f\"Missing {null_p0_pos_x_count} values ({null_p0_pos_x_perc:.2%})\")","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:17:54.430915Z","iopub.execute_input":"2022-10-22T16:17:54.431694Z","iopub.status.idle":"2022-10-22T16:17:54.459272Z","shell.execute_reply.started":"2022-10-22T16:17:54.431657Z","shell.execute_reply":"2022-10-22T16:17:54.458239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To keep things simple, let's just drop all null values.","metadata":{}},{"cell_type":"code","source":"df = df.dropna(axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:18:00.523938Z","iopub.execute_input":"2022-10-22T16:18:00.524836Z","iopub.status.idle":"2022-10-22T16:18:06.565062Z","shell.execute_reply.started":"2022-10-22T16:18:00.524781Z","shell.execute_reply":"2022-10-22T16:18:06.564038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"has_na = {}\nfor col in df.columns:\n    has_na[col] = df[col].isnull().values.any()\n\nprint(\"Columns that contain null values:\")\nfor col in has_na:\n    if has_na[col]:\n        print(col)\n        \ndf","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:18:09.100639Z","iopub.execute_input":"2022-10-22T16:18:09.101699Z","iopub.status.idle":"2022-10-22T16:18:11.390377Z","shell.execute_reply.started":"2022-10-22T16:18:09.101658Z","shell.execute_reply":"2022-10-22T16:18:11.389311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Training","metadata":{}},{"cell_type":"markdown","source":"We'll be predicting the probability of team A scoring and team B scoring with 2 separate models.","metadata":{}},{"cell_type":"markdown","source":"## Model 1","metadata":{}},{"cell_type":"code","source":"# used to encode the binary classes\nle_1 = LabelEncoder()","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:18:16.304466Z","iopub.execute_input":"2022-10-22T16:18:16.305330Z","iopub.status.idle":"2022-10-22T16:18:16.309947Z","shell.execute_reply.started":"2022-10-22T16:18:16.305292Z","shell.execute_reply":"2022-10-22T16:18:16.308750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_1 = XGBClassifier(\n    n_estimators=n_estimators,\n    max_depth=max_depth,\n    learning_rate=learning_rate,\n    objective='binary:logistic',\n    tree_method='gpu_hist' if gpu else 'hist'\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:18:24.304675Z","iopub.execute_input":"2022-10-22T16:18:24.305787Z","iopub.status.idle":"2022-10-22T16:18:24.310982Z","shell.execute_reply.started":"2022-10-22T16:18:24.305739Z","shell.execute_reply":"2022-10-22T16:18:24.309972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv_1 = cross_validate(\n    model_1, \n    X=df.drop(columns=['team_A_scoring_within_10sec', 'team_B_scoring_within_10sec']).values,\n    y=le_1.fit_transform(df['team_A_scoring_within_10sec'].values),\n    scoring=\"neg_log_loss\",\n    cv=folds,\n    verbose=2,\n    return_estimator=True\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:18:37.079908Z","iopub.execute_input":"2022-10-22T16:18:37.080302Z","iopub.status.idle":"2022-10-22T16:25:10.030559Z","shell.execute_reply.started":"2022-10-22T16:18:37.080260Z","shell.execute_reply":"2022-10-22T16:25:10.029454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model 2","metadata":{}},{"cell_type":"code","source":"# used to encode the binary classes\nle_2 = LabelEncoder()","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:25:10.032452Z","iopub.execute_input":"2022-10-22T16:25:10.033422Z","iopub.status.idle":"2022-10-22T16:25:10.039842Z","shell.execute_reply.started":"2022-10-22T16:25:10.033381Z","shell.execute_reply":"2022-10-22T16:25:10.036694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_2 = XGBClassifier(\n    n_estimators=n_estimators,\n    max_depth=max_depth,\n    learning_rate=learning_rate,\n    objective='binary:logistic',\n    tree_method='gpu_hist' if gpu else 'hist'\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:25:10.042700Z","iopub.execute_input":"2022-10-22T16:25:10.042981Z","iopub.status.idle":"2022-10-22T16:25:10.049258Z","shell.execute_reply.started":"2022-10-22T16:25:10.042956Z","shell.execute_reply":"2022-10-22T16:25:10.048273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv_2 = cross_validate(\n    model_2, \n    X=df.drop(columns=['team_A_scoring_within_10sec', 'team_B_scoring_within_10sec']).values,\n    y=le_2.fit_transform(df['team_B_scoring_within_10sec'].values),\n    scoring=\"neg_log_loss\",\n    cv=folds,\n    verbose=2,\n    return_estimator=True\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:28:13.542308Z","iopub.execute_input":"2022-10-22T16:28:13.543317Z","iopub.status.idle":"2022-10-22T16:34:41.264708Z","shell.execute_reply.started":"2022-10-22T16:28:13.543271Z","shell.execute_reply":"2022-10-22T16:34:41.263673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Evaluation","metadata":{}},{"cell_type":"markdown","source":"## Cross Validation Test Score","metadata":{}},{"cell_type":"markdown","source":"Let's visualize the log loss score on each of our folds, for both our models","metadata":{}},{"cell_type":"code","source":"df_cv_1 = pd.DataFrame(\n    {\n        \"model\": \"Model 1\",\n        \"fold\": list(range(folds)),\n        \"test_log_loss\": - cv_1[\"test_score\"]\n    }\n)\ndf_cv_2 = pd.DataFrame(\n    {\n        \"model\": \"Model 2\",\n        \"fold\": list(range(folds)),\n        \"test_log_loss\": - cv_2[\"test_score\"]\n    }\n)\ndf_cv = pd.concat([df_cv_1, df_cv_2])\n\ndel df_cv_1\ndel df_cv_2\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:34:47.916582Z","iopub.execute_input":"2022-10-22T16:34:47.916965Z","iopub.status.idle":"2022-10-22T16:34:48.123592Z","shell.execute_reply.started":"2022-10-22T16:34:47.916930Z","shell.execute_reply":"2022-10-22T16:34:48.122367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.bar(\n    df_cv, x='fold', y='test_log_loss', color='model', \n    barmode='group', title='Cross Validation Log Loss'\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:34:57.183552Z","iopub.execute_input":"2022-10-22T16:34:57.183921Z","iopub.status.idle":"2022-10-22T16:34:57.811398Z","shell.execute_reply.started":"2022-10-22T16:34:57.183889Z","shell.execute_reply":"2022-10-22T16:34:57.810506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/test.csv')\ndf_test","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:35:06.776504Z","iopub.execute_input":"2022-10-22T16:35:06.776973Z","iopub.status.idle":"2022-10-22T16:35:15.562409Z","shell.execute_reply.started":"2022-10-22T16:35:06.776929Z","shell.execute_reply":"2022-10-22T16:35:15.561314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def preprocess(df):\n    # velocity magnitude\n    for col, vec in vel_groups.items():\n        df[col] = euclidian_norm(df[vec])\n    \n    # ball distance\n    for col, vec in pos_groups.items():\n        df[col + \"_ball_dist\"] = euclidian_norm(df[vec].values - df[pos_groups[\"ball_pos\"]].values)\n    \n    df = df.drop(columns=['ball_pos_ball_dist'])\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:35:49.982872Z","iopub.execute_input":"2022-10-22T16:35:49.983920Z","iopub.status.idle":"2022-10-22T16:35:49.990847Z","shell.execute_reply.started":"2022-10-22T16:35:49.983870Z","shell.execute_reply":"2022-10-22T16:35:49.989540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = preprocess(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:35:54.211106Z","iopub.execute_input":"2022-10-22T16:35:54.211781Z","iopub.status.idle":"2022-10-22T16:35:57.673581Z","shell.execute_reply.started":"2022-10-22T16:35:54.211744Z","shell.execute_reply":"2022-10-22T16:35:57.672557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# take the mean of the predictions made by the k models gotten out of the k-fold cross validation\npred_1 = np.zeros(df_test.shape[0])\nfor estimator in cv_1['estimator']:\n    pred_1 += estimator.predict_proba(df_test.drop(columns=['id']).values)[:, 1]\n\npred_1 /= folds","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:36:04.657616Z","iopub.execute_input":"2022-10-22T16:36:04.657978Z","iopub.status.idle":"2022-10-22T16:36:57.205881Z","shell.execute_reply.started":"2022-10-22T16:36:04.657946Z","shell.execute_reply":"2022-10-22T16:36:57.205132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# take the mean of the predictions made by the k models gotten out of the k-fold cross validation\npred_2 = np.zeros(df_test.shape[0])\nfor estimator in cv_2['estimator']:\n    pred_2 += estimator.predict_proba(df_test.drop(columns=['id']).values)[:, 1]\n\npred_2 /= folds","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:38:10.369368Z","iopub.execute_input":"2022-10-22T16:38:10.369741Z","iopub.status.idle":"2022-10-22T16:39:06.129710Z","shell.execute_reply.started":"2022-10-22T16:38:10.369709Z","shell.execute_reply":"2022-10-22T16:39:06.128938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission = pd.DataFrame(\n    {\n        \"id\": df_test['id'],\n        \"team_A_scoring_within_10sec\": pred_1,\n        \"team_B_scoring_within_10sec\": pred_2\n    }\n)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:40:06.843095Z","iopub.execute_input":"2022-10-22T16:40:06.843471Z","iopub.status.idle":"2022-10-22T16:40:06.852633Z","shell.execute_reply.started":"2022-10-22T16:40:06.843437Z","shell.execute_reply":"2022-10-22T16:40:06.851170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:41:05.09012Z","iopub.execute_input":"2022-10-05T08:41:05.090481Z","iopub.status.idle":"2022-10-05T08:41:05.107317Z","shell.execute_reply.started":"2022-10-05T08:41:05.090424Z","shell.execute_reply":"2022-10-05T08:41:05.106319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-10-22T16:40:33.631458Z","iopub.execute_input":"2022-10-22T16:40:33.631809Z","iopub.status.idle":"2022-10-22T16:40:35.671548Z","shell.execute_reply.started":"2022-10-22T16:40:33.631779Z","shell.execute_reply":"2022-10-22T16:40:35.670561Z"},"trusted":true},"execution_count":null,"outputs":[]}]}