{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"library(tidyverse) # metapackage of all tidyverse packages\nlibrary(arrow)\nlibrary(ggplot2)\n\nlibrary(glue)\n\nfig <- function(width, heigth){\n    # borrowed from https://www.kaggle.com/getting-started/105201\n    options(repr.plot.width = width, repr.plot.height = heigth)\n}\nset.seed(42)\n\n# load random train sample - with augmentation (reflection, flips etc)\ndf_train <- read_parquet('../input/tps-2022-10-create-data-sample/train_fe_0.parquet') %>%\n    sample_frac(0.3)\n\ndf_val <- read_parquet('../input/tps-2022-10-create-data-sample/val_fe.parquet')","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-10-12T09:39:54.263149Z","iopub.execute_input":"2022-10-12T09:39:54.264954Z","iopub.status.idle":"2022-10-12T09:39:56.445044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Introduction\n\nIn this notebook I fit a logistic regression additive model to understand the goal scoring \"probability surface\".\n\nThe data used has been augmented with reflections and flips so there is no player or \"handed-ness\" bias.\n\n## Model 0\n\nPredicting a constant value we'd expect a logloss ~ 0.22","metadata":{}},{"cell_type":"code","source":"fit.0 <- glm(\n    team_A_scoring_within_10sec ~ 1,\n    family = 'binomial', \n    data = df_train)\n\ncat('\\n')\nll <- Metrics::logLoss(\n        df_val$team_A_scoring_within_10sec, \n        predict(fit.0, df_val, type = 'response'))\ncat(glue('logloss {round(ll,3)}'))","metadata":{"execution":{"iopub.status.busy":"2022-10-12T09:42:02.229636Z","iopub.execute_input":"2022-10-12T09:42:02.230989Z","iopub.status.idle":"2022-10-12T09:42:02.979513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model 1\n\nThe model is made up of the features: ball position & velocity and the \"boost advantage\" to team A.\n\nThe oof log-loss is ~0.2, a little better than predicting a constant value","metadata":{}},{"cell_type":"code","source":"fit.1 <- mgcv::bam(\n    I(team_A_scoring_within_10sec) ~ \n         te(ball_pos_x, ball_pos_y, ball_pos_z) \n         + te(ball_vel_x, ball_vel_y, ball_vel_z)\n         + s(boost_advantage)\n         ,\n    family = 'binomial', \n    data = df_train %>% mutate(\n        boost_advantage = p0_boost + p1_boost + p2_boost \n                        - (p3_boost + p4_boost + p5_boost)), \n    discrete = TRUE, \n    gamma = 1.2)\n\nsummary(fit.1)\n\ncat('\\n')\nll <- Metrics::logLoss(\n        df_val$team_A_scoring_within_10sec, \n        predict(\n            fit.1, \n            df_val %>% mutate(\n                boost_advantage = p0_boost + p1_boost + p2_boost \n                                - (p3_boost + p4_boost + p5_boost)), \n            type = 'response'))\ncat(glue('logloss {round(ll,3)}'))","metadata":{"execution":{"iopub.status.busy":"2022-10-12T09:37:05.873413Z","iopub.execute_input":"2022-10-12T09:37:05.874914Z","iopub.status.idle":"2022-10-12T09:37:20.096268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the closer the ball is to the opponent's goal (x, y and z), the greater the probability of scoring within the next 10 seconds (no shock there), and that an advantage in boosts is linearly related to the probability of scoring.","metadata":{}},{"cell_type":"code","source":"fig(8, 8)\nplot(fit.1, pages = 1, scale = 0)","metadata":{"execution":{"iopub.status.busy":"2022-10-12T09:09:11.742116Z","iopub.execute_input":"2022-10-12T09:09:11.743381Z","iopub.status.idle":"2022-10-12T09:09:12.388207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# position vectors\n\nfig(12,7)\npar(mfrow = c(1, 2))\nmgcv::vis.gam(fit.1, view=c(\"ball_pos_x\",\"ball_pos_y\"), \n    color = 'terrain', theta = 45, type = 'link')\nmgcv::vis.gam(fit.1, view=c(\"ball_pos_z\",\"ball_pos_y\"), \n    color = \"terrain\", theta = 45,  type = 'link')\n\n# least chance of scoring when the ball is in our own goal :D","metadata":{"execution":{"iopub.status.busy":"2022-10-12T09:09:12.390975Z","iopub.execute_input":"2022-10-12T09:09:12.392734Z","iopub.status.idle":"2022-10-12T09:09:12.731242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# movement vectors\n\nfig(12,7)\npar(mfrow = c(1, 2))\nmgcv::vis.gam(fit.1, view=c(\"ball_vel_x\",\"ball_vel_y\"), \n    color = 'terrain', theta = 45, type = 'link')\nmgcv::vis.gam(fit.1, view=c(\"ball_vel_z\",\"ball_vel_y\"), \n    color = \"terrain\", theta = 45,  type = 'link')\n\n# we need the ball to be travelling in the Y direction toward the goal (faster is better)\n# less X and Z movement is better.","metadata":{"execution":{"iopub.status.busy":"2022-10-12T09:09:12.733108Z","iopub.execute_input":"2022-10-12T09:09:12.734152Z","iopub.status.idle":"2022-10-12T09:09:13.056986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model 2\n\nLet's add a feature for the average team y-position. This is a slightly better model - there is some explanatory power to the new feature.\n\n- There is more chance of scoring when arranged centrally\n- Even better, if player B is close to their goal line\n- There is advantage to being closer to the ball than the opponent.","metadata":{}},{"cell_type":"code","source":"fit.2 <- mgcv::bam(\n    team_A_scoring_within_10sec ~ \n         te(ball_pos_x, ball_pos_y, ball_pos_z) \n         + te(ball_vel_x, ball_vel_y, ball_vel_z)\n         + s(boost_advantage)\n         + te(pA_pos_y, pB_pos_y, ball_pos_y)\n         ,\n    family = 'binomial', \n    data = df_train %>% mutate(\n        boost_advantage = (\n            p0_boost + p1_boost + p2_boost \n            - (p3_boost + p4_boost + p5_boost)),\n        pA_pos_y = (p0_pos_y + p1_pos_y + p2_pos_y) / 3,\n        pB_pos_y = (p3_pos_y + p4_pos_y + p5_pos_y) / 3),\n    discrete = TRUE, \n    gamma = 1.2)\n\nsummary(fit.2)\n\ncat('\\n')\nll <- Metrics::logLoss(\n        df_val$team_A_scoring_within_10sec, \n        predict(\n            fit.2, \n            df_val %>% mutate(\n                boost_advantage = (\n                    p0_boost + p1_boost + p2_boost \n                    - (p3_boost + p4_boost + p5_boost)),\n                pA_pos_y = (p0_pos_y + p1_pos_y + p2_pos_y) / 3,\n                pB_pos_y = (p3_pos_y + p4_pos_y + p5_pos_y) / 3),\n            type = 'response'))\ncat(glue('logloss {round(ll,3)}'))","metadata":{"execution":{"iopub.status.busy":"2022-10-12T09:09:13.058838Z","iopub.execute_input":"2022-10-12T09:09:13.060000Z","iopub.status.idle":"2022-10-12T09:12:19.191320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(15,7)\npar(mfrow = c(1, 3))\nmgcv::vis.gam(fit.2, view=c(\"pB_pos_y\",\"pA_pos_y\"), \n    color = 'terrain', theta = 135, type = 'link')\n\nmgcv::vis.gam(fit.2, view=c(\"pA_pos_y\",\"ball_pos_y\"), \n    color = 'terrain', theta = 35, type = 'link')\n\nmgcv::vis.gam(fit.2, view=c(\"pB_pos_y\",\"ball_pos_y\"), \n    color = 'terrain', theta = 35, type = 'link')","metadata":{"execution":{"iopub.status.busy":"2022-10-12T09:14:25.177067Z","iopub.execute_input":"2022-10-12T09:14:25.178386Z","iopub.status.idle":"2022-10-12T09:14:25.629046Z"},"trusted":true},"execution_count":null,"outputs":[]}]}