{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1 Introduction\n\nThis EDA explores the data available for the Tabular Playground Series - October 2022 competition. Simple data exploration is performed, as well as preliminary modeling.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport gc","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1.1 Overall Dataset Impressions\n\nLet's start by looking only at one training file first. As indicated, we'll use the dataframe type file to format the Pandas dataframe so that we don't waste memory using `float64` values for every numeric we see.","metadata":{}},{"cell_type":"code","source":"dtypes = pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_dtypes.csv\")\ndtype_dict = {k: v for (k, v) in zip(dtypes.column, dtypes.dtype)}\ntrain = pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_0.csv\", dtype=dtype_dict)\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_1.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_2.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_3.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_4.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_5.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_6.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_7.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_8.csv\", dtype=dtype_dict))\ntrain = train.append(pd.read_csv(\"../input/tabular-playground-series-oct-2022/train_9.csv\", dtype=dtype_dict))\ntest = pd.read_csv(\"../input/tabular-playground-series-oct-2022/test.csv\", dtype=dtype_dict)\n\ntrain","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It can be helpful to see what kind of shape we're dealing with for the data itself. This is already described in the competition rules and data description, but sometimes it helps to see the raw numbers.","metadata":{}},{"cell_type":"code","source":"def cat_column_info(column):\n    num_categories = train[column].nunique()\n    print(\"------> {} <------\".format(column))\n    print(\"--: train - type {}\".format(train[column].dtype))\n    print(\"--: test  - type {}\".format(test[column].dtype))\n    print(\"--: train - # categories {}\".format(train[column].nunique()))\n    print(\"--: test  - # categories {}\".format(test[column].nunique()))\n    if num_categories < 10:\n        if train[column].dtype == \"int64\":\n            print(\"--: train - values {}\".format(np.sort(train[column].unique())))\n            print(\"--: test  - values {}\".format(np.sort(test[column].unique())))\n        else:\n            print(\"--: train - values {}\".format(train[column].unique()))\n            print(\"--: test  - values {}\".format(test[column].unique()))\n    print(\"--: train - NaN count {}\".format(train[column].isnull().values.sum()))\n    print(\"--: test  - NaN count {}\".format(test[column].isnull().values.sum()))\n    print(\"\")\n\ndef cont_column_info(column):\n    print(\"------> {} <------\".format(column))\n    print(\"--: train - type {}\".format(train[column].dtype))\n    print(\"--: test  - type {}\".format(test[column].dtype))\n    print(\"--: train - min {}\".format(train[column].min()))\n    print(\"--: test  - min {}\".format(test[column].min()))\n    print(\"--: train - max {}\".format(train[column].max()))\n    print(\"--: test  - max {}\".format(test[column].max()))    \n    print(\"--: train - NaN count {}\".format(train[column].isnull().values.sum()))\n    print(\"--: test  - NaN count {}\".format(test[column].isnull().values.sum()))\n    print(\"\")\n    \nprint(\": Train set shape {}\".format(train.shape))\nprint(\": Test set shape {}\".format(test.shape))\nprint(\"\")","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our input data consists of:\n\n* `train.csv` - 9 GB in size, containing 61 columns and 21,198,036 rows\n* `test.csv` - 287 MB in size, containing 55 columns and 701,143 rows\n\nOne main observation here is the sheer size of the data we are looking at. According to Jupyter notebook memory usage, the full training dataset exerts a memory pressure of 7.10 GB when loaded with the correct data types. While this fits in Kaggle's CPU memory allotment, model training will exert more pressure on the Kaggle 16 GB CPU memory and GPU memory limitations. We should definitely explore what column formats are at play, and whether we can reduce memory pressure by compressing or dropping data as required.\n\n---","metadata":{}},{"cell_type":"markdown","source":"# 2 Features\n\nLet's take a deeper dive on some of the particulars relating to the features in the dataset.","metadata":{}},{"cell_type":"markdown","source":"# 2.1 Null Values\n\nLet's explore the issue of missing values in the dataset to see if there are systemic problems with data representation.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nsns.set_style(\"whitegrid\")\nsns_params = {\"palette\": \"dark\"}\n\nfig, axs = plt.subplots(nrows=1, ncols=2, figsize=(15, 5))\n\ntrain[\"null_count\"] = train.isnull().sum(axis=1)\ncounts = train.groupby(\"null_count\")[\"event_id\"].count().to_dict()\nnull_data = {\"{} Null Value(s)\".format(k) : v for k, v in counts.items() if k < 8}\nnull_data[\"More Than 8 Null Value(s)\"] = sum([v for k, v in counts.items() if k >= 8])\n\n_ = axs[0].pie(\n    x=list(null_data.values()), \n    autopct=\"%.2f%%\", \n    explode=[0.05] * len(null_data.keys()), \n    labels=null_data.keys(), \n    pctdistance=0.5, \n    colors=sns.color_palette(\"Paired\")[0:5]\n)\n_ = axs[0].set_title(\"Training Data\")\n\ntest[\"null_count\"] = test.isnull().sum(axis=1)\ncounts = test.groupby(\"null_count\")[\"p5_vel_x\"].count().to_dict()\nnull_data = {\"{} Null Value(s)\".format(k) : v for k, v in counts.items() if k < 8}\nnull_data[\"More Than 8 Null Value(s)\"] = sum([v for k, v in counts.items() if k >= 8])\n\n_ = axs[1].pie(\n    x=list(null_data.values()), \n    autopct=\"%.2f%%\", \n    explode=[0.05] * len(null_data.keys()), \n    labels=null_data.keys(), \n    pctdistance=0.5, \n    colors=sns.color_palette(\"Paired\")[0:5]\n)\n_ = axs[1].set_title(\"Test Data\")","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A few observations:\n\n* In the training set, there are occurrences of single null value features that appear roughly 22% of the data set. This same instance of a single null value does not occur in the testing set. We need to dig into the null value appearances to find out why there is a difference.\n* Otherwise, it looks like we have null values that impact around 5% of the records we have available for training and testing.\n\nLet's dig into the training set and see where the null values are occurring.","metadata":{}},{"cell_type":"code","source":"# Look at the rows that have NaN values, and count up how many occur\nnull_rows = train[(train[\"null_count\"] > 0)]\nnulls = {}\nfor feature in null_rows.columns:\n    if null_rows[feature].isnull().any():\n        nulls[feature] = np.count_nonzero(null_rows[feature].isnull().values)\n        \n# Plot out the counts of each \nkeys = [key for key in nulls.keys()]\nvalues = [int(nulls[key]) for key in keys]\n\nf, ax = plt.subplots(figsize=(20, 20))\nax = sns.barplot(x=values, y=keys, palette=sns.color_palette(\"Paired\")[0:len(keys)])\n_ = ax.set_title(\"Null Value Counts by Feature\", fontsize=15)\n_ = ax.set_ylabel(\"Number of Records\", fontsize=15)\n_ = ax.set_xlabel(\"Feature\", fontsize=15)\n_ = plt.ticklabel_format(style='plain', axis='x')","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A few things appear to be going on here.  \n\n* As described by the competition, `team_scoring_next` will be a null value if neither team ends up scoring within the bounds of the game sequence. So somewhere around 22% of our data available for training represents game sequences where neither team scores. Overall, this may end up skewing our data slightly, however, the impact may be negligible when we look at the number of records that a team scores within the next 10 seconds versus the number of records where no scoring is occurring (more on this later).\n* The `team_scoring_next` field is only part of the training set (i.e. it is a target feature). Since the testing set does not have this field, this explains the null count disparity between the training and testing sets. \n\nThere is an interesting phenomenon occurring with player position, velocity, and boost time features. Let's isolate a single player and look at the number of null values across positions and velocities. Let's look at player 5:","metadata":{}},{"cell_type":"code","source":"# Plot out the nulls for player 5 only \nkeys = [key for key in nulls.keys() if key.startswith(\"p5\")]\nvalues = [int(nulls[key]) for key in keys]\n\nf, ax = plt.subplots(figsize=(15, 15))\nax = sns.barplot(y=values, x=keys, palette=sns.color_palette(\"Paired\")[0:len(keys)])\nfor p in ax.patches:\n    ax.text(x=p.get_x()+(p.get_width()/2), y=p.get_height(), s=\"{:,d}\".format(round(p.get_height())), ha=\"center\")\n_ = ax.set_title(\"Player 5 Null Row Counts\", fontsize=15)\n_ = ax.set_ylabel(\"Number of Rows\", fontsize=15)\n_ = ax.set_xlabel(\"Player 5 Feature\", fontsize=15)","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As observed, we can see the number of rows impacted by null values is the same for each player 5 feature. The question is whether or not these features are null at the same (i.e. position, velocity, and boost are all null at the same time), or if the nulls can appear in a single feature independently of the other null values (i.e. can `p5_vel_x` be null but `p5_vel_y` be not null in the same row). A simple query will tell us. We'll isolate player 5 statistics, and select any row that has a null in any column.","metadata":{}},{"cell_type":"code","source":"player_5_features = [feature for feature in train.columns if feature.startswith(\"p5\")]\nplayer_5 = train[player_5_features]\nplayer_5[player_5.isnull().any(axis=1)]","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above query, we see there are 193,658 rows that have at least 1 null value in it. As you can see from the resultant dataframe, all the player statistics are null at the same time. Since there are only 193,658 rows in the dataframe, and each statistic shows null 193,658 times, we can conclude that when one statistic value is missing, then all of them are missing at the same time.\n\nThe question is why? The answer is of course related to certain _gameplay mechanics_. In Rocket League, ramming a player at supersonic speeds will temporarily demolish that player. In essence, the player is removed from gameplay for a certain amount of time. The null values indicate moments in time when player 5 has been demolished. So, the null values are not due to missing data, but rather a player that has been removed from the active board.\n\nFor the purposes of our analysis, we'll have to fill in missing values with something. With player positions, we can fill in missing values with a value that exists outside the scope of the playing field. For velocities, we'll set them to zero.\n\n### Key Observations About Null Values\n\n* Null values in the `team_scoring_next` feature are expected, since there are training examples where no team scores a point. It may make sense to one-hot encode the field into `team_a_scores_next` and `team_b_scores_next` so that the nulls are translated into `0` values for both fields when a null is encountered.\n* Null values in the player statistic features happen when a player has been demolished from the field. We should probably create a new binary feature called `pX_demolished` to indicate that player `X` has been removed from gameplay, and then fill the missing statistic values with a large or small numbers to indicate that they are no longer in the play area. This will stop problems that may arise with nulls in our dataset and various machine learning algorithms.","metadata":{}},{"cell_type":"code","source":"# Clean up unused variables to save RAM\nimport gc\n\ndel(keys)\ndel(values)\ndel(nulls)\ndel(null_data)\ndel(counts)\ndel(test)\ndel(player_5_features)\ndel(player_5)\n_ = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2.2 Playing Field Definition\n\nOne of the things we need to do is determine where the goal areas are located. We can do so by looking at event times in the training set and see where the ball position is at the end of the event sequence. If we combine this with the `team_scoring_next` feature, we can determine whether a goal was scored. This should hopefully show us where in three dimensional space the goals exist.One of the things we need to do is determine where the goal areas are located. We can do so by looking at event times in the training set and see where the ball position is at the end of the event sequence. If we combine this with the `team_scoring_next` feature, we can determine whether a goal was scored. This should hopefully show us where in three dimensional space the goals exist.\n\nThere are a number of fields that refer to the position of both the players, and the ball within the play field area. Let's take a look at how the field is laid out.","metadata":{}},{"cell_type":"code","source":"y_features = [feature for feature in train.columns if feature.endswith(\"pos_y\")]\nx_features = [feature for feature in train.columns if feature.endswith(\"pos_x\")]\nz_features = [feature for feature in train.columns if feature.endswith(\"pos_z\")]\n\nx_max = -999\nx_min = 999\nfor x_feature in x_features:\n    max_val = train[x_feature].max()\n    min_val = train[x_feature].min()\n    x_max = max_val if max_val > x_max else x_max\n    x_min = min_val if min_val < x_min else x_min\n\ny_max = -999\ny_min = 999\nfor y_feature in y_features:\n    max_val = train[y_feature].max()\n    min_val = train[y_feature].min()\n    y_max = max_val if max_val > y_max else y_max\n    y_min = min_val if min_val < y_min else y_min\n\nz_max = -999\nz_min = 999\nfor z_feature in z_features:\n    max_val = train[z_feature].max()\n    min_val = train[z_feature].min()\n    z_max = max_val if max_val > z_max else z_max\n    z_min = min_val if min_val < z_min else z_min\n\nprint(\": X coordinate ranges from {} to {}\".format(x_min, x_max))\nprint(\": Y coordinate ranges from {} to {}\".format(y_min, y_max))\nprint(\": Z coordinate ranges from {} to {}\".format(z_min, z_max))","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This tells us that we have a play area that is 164 units wide, 240 units long, and 40 units tall. The question is where the goal areas are located. Let's generate a heatmap of the are and plot where the players and balls can end up. For the sake of the heatmap, we'll translate the floats into integers, since we aren't too concerned where within a cubic unit of area the ball or player is, just that it ends up there. We'll stick to X and Y positions for now. We'll also adjust the heatmap such that it plots X coordinates from 0 to 164, and Y coordinates from 0 to 240.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\n\nfor feature in y_features:\n    train[feature] = train[feature].fillna(250).astype(np.int16)\n    train[feature] = train[feature].apply(lambda x: x + 119 if x != 250 else x)\n    train[feature] = train[feature].astype(np.uint8)\n\nfor feature in x_features:\n    train[feature] = train[feature].fillna(250).astype(np.int16)\n    train[feature] = train[feature].apply(lambda x: x + 82 if x != 250 else x)\n    train[feature] = train[feature].astype(np.uint8)\n    \nfor feature in z_features:\n    train[feature] = train[feature].fillna(250).astype(np.int16)\n    train[feature] = train[feature].astype(np.uint8)\n    \nvalid_positions = set()\n\nxy_features = [\n    \"p0_pos_x\", \"p0_pos_y\", \"p1_pos_x\", \"p1_pos_y\", \"p2_pos_x\", \"p2_pos_y\", \"p3_pos_x\", \"p3_pos_y\", \n    \"p4_pos_x\", \"p4_pos_y\", \"p5_pos_x\", \"p5_pos_y\", \"ball_pos_x\", \"ball_pos_y\"\n]\n\nfor (_, p0_x, p0_y, p1_x, p1_y, p2_x, p2_y, p3_x, p3_y, p4_x, p4_y, p5_x, p5_y, ball_x, ball_y) in train[xy_features].itertuples():\n    valid_positions.add((p0_x, p0_y))\n    valid_positions.add((p1_x, p1_y))\n    valid_positions.add((p2_x, p2_y))\n    valid_positions.add((p3_x, p3_y))\n    valid_positions.add((p4_x, p4_y))\n    valid_positions.add((p5_x, p5_y))\n    valid_positions.add((ball_x, ball_y))\n\nnp_array = np.zeros(shape=(255, 255))\nfor (x, y) in valid_positions:\n    np_array[y, x] += 1\n    \nf, ax = plt.subplots(figsize=(10, 10))\nax.imshow(np_array, cmap='hot', interpolation='nearest')\n_ = ax.grid(False)\n_ = ax.set_title(\"Playing Field Definition (X, Y)\", fontweight=\"bold\", size=15)\n\n# Clean up unused variables to save on RAM\ndel(valid_positions)\ndel(xy_features)\ndel(np_array)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this particular map, we can see quite well where the two goal areas are located - on the north and south ends of the field. Let's take a look at a 3-dimensional representation of the playing field.","metadata":{}},{"cell_type":"code","source":"import gc\n\nvalid_positions = set()\ngc.collect()\n\nxyz_features = [\n    \"p0_pos_x\", \"p0_pos_y\", \"p0_pos_z\", \"p1_pos_x\", \"p1_pos_y\", \"p1_pos_z\", \"p2_pos_x\", \"p2_pos_y\", \"p2_pos_z\", \n    \"p3_pos_x\", \"p3_pos_y\", \"p3_pos_z\", \"p4_pos_x\", \"p4_pos_y\", \"p4_pos_z\", \"p5_pos_x\", \"p5_pos_y\", \"p5_pos_z\", \n    \"ball_pos_x\", \"ball_pos_y\", \"ball_pos_z\"\n]\n\nfor (_, p0_x, p0_y, p0_z, p1_x, p1_y, p1_z, p2_x, p2_y, p2_z, p3_x, p3_y, p3_z, p4_x, p4_y, p4_z, p5_x, p5_y, p5_z, ball_x, ball_y, ball_z) in train[xyz_features].itertuples():\n    valid_positions.add((p0_x, p0_y, p0_z))\n    valid_positions.add((p1_x, p1_y, p1_z))\n    valid_positions.add((p2_x, p2_y, p2_z))\n    valid_positions.add((p3_x, p3_y, p3_z))\n    valid_positions.add((p4_x, p4_y, p4_z))\n    valid_positions.add((p5_x, p5_y, p5_z))\n    valid_positions.add((ball_x, ball_y, ball_z))\n    \nfrom mpl_toolkits.mplot3d import Axes3D\nimport matplotlib.pyplot as plt\nimport numpy as np\nfrom pylab import *\n\nx = []\ny = []\nz = []\n\nxyz_max_data = dict()\nfor (px, py, pz) in valid_positions:\n    if px not in xyz_max_data:\n        xyz_max_data[px] = dict()\n    if py not in xyz_max_data[px]:\n        xyz_max_data[px][py] = pz\n    elif pz > xyz_max_data[px][py]:\n        xyz_max_data[px][py] = pz\n\nfor px in range(164):\n    for py in range(240):\n        x.append(px)\n        y.append(py)\n        z.append(xyz_max_data[px][py] if px in xyz_max_data and py in xyz_max_data[px] else 0)\n\nfig = plt.figure(figsize=(10, 10))\nax = fig.add_subplot(111, projection='3d')\nsurface = ax.plot_trisurf(x, y, z, cmap=cm.jet, linewidth=0)\n  \nax.set_title(\"3D Playing Field Definition (X, Y, Z)\")\nax.set_xlabel(\"X-axis\")\nax.set_ylabel(\"Y-axis\")\nax.set_zlabel(\"Z-axis\")\n\n# Clean up unused variables to save on RAM\ndel(valid_positions)\ndel(xyz_features)\n_  = gc.collect()\n\nplt.show()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we can clearly see the goal areas defined at the north and south ends of the field. The goal openings themselves are quite narrow, and taper off. Their Z position opening occurs at roughly 13 units high, and is a narrow opening of roughly 20 units. \n\n### Key Observations About Playing Field\n\n* The goal areas are to the north and south of the field.\n* The opening of the goal area appears to be roughly 20 units wide, by 13 units high.\n* Sides of the playing area are all curved.","metadata":{}},{"cell_type":"markdown","source":"# 2.3 Player Positions\n\nNow that we know what the field looks like, we need to know which end of the field belongs to which team. In other words, is team A trying to score goals on the net to the north or the south? We find out this information by looking at the ball position when team A scores, and when team B scores. Most of the time the players will be near the net which is under their control when they are defending. Let's take a closer look.","metadata":{}},{"cell_type":"code","source":"xy_features = [\n    \"p0_pos_x\", \"p0_pos_y\", \"p1_pos_x\", \"p1_pos_y\", \"p2_pos_x\", \"p2_pos_y\", \"p3_pos_x\", \"p3_pos_y\", \n    \"p4_pos_x\", \"p4_pos_y\", \"p5_pos_x\", \"p5_pos_y\", \"ball_pos_x\", \"ball_pos_y\"\n]\n\ntrain_a_near_scoring_net = train[(train[\"team_A_scoring_within_10sec\"] == 1) & (train[\"event_time\"] > -1.)]\n\nvalid_positions = []\n\nfor (_, p0_x, p0_y, p1_x, p1_y, p2_x, p2_y, p3_x, p3_y, p4_x, p4_y, p5_x, p5_y, ball_x, ball_y) in train_a_near_scoring_net[xy_features].itertuples():\n    if p0_x < 250:\n        valid_positions.append((p0_x, p0_y))\n    if p1_x < 250:\n        valid_positions.append((p1_x, p1_y))\n    if p2_x < 250:\n        valid_positions.append((p2_x, p2_y))\n    # valid_positions.append((ball_x, ball_y))\n\nnp_array = np.zeros(shape=(255, 170))\nfor (x, y) in valid_positions:\n    np_array[y, x] += 1.\n\nfig, axs = plt.subplots(nrows=1, ncols=2, figsize=(20, 10), sharey=True)\n\n_ = sns.heatmap(np_array, ax=axs[0])\n_ = axs[0].grid(False)\n_ = axs[0].set_title(\"Team A Scoring Positions\", fontweight=\"bold\", size=15)\n\n# Clean up unused variables to save on RAM\ndel(valid_positions)\ndel(np_array)\ndel(train_a_near_scoring_net)\n_  = gc.collect()\n\ntrain_b_near_scoring_net = train[(train[\"team_B_scoring_within_10sec\"] == 1) & (train[\"event_time\"] > -1.)]\n\nvalid_positions = []\n\nfor (_, p0_x, p0_y, p1_x, p1_y, p2_x, p2_y, p3_x, p3_y, p4_x, p4_y, p5_x, p5_y, ball_x, ball_y) in train_b_near_scoring_net[xy_features].itertuples():\n    if p3_x < 250:\n        valid_positions.append((p3_x, p3_y))\n    if p4_x < 250:\n        valid_positions.append((p4_x, p4_y))\n    if p5_x < 250:\n        valid_positions.append((p5_x, p5_y))\n    # valid_positions.append((ball_x, ball_y))\n\nnp_array = np.zeros(shape=(255, 170))\nfor (x, y) in valid_positions:\n    np_array[y, x] += 1.\n    \n_ = sns.heatmap(np_array, ax=axs[1])\n_ = axs[1].grid(False)\n_ = axs[1].set_title(\"Team B Scoring Positions\", fontweight=\"bold\", size=15)\n\n# Clean up unused variables to save on RAM\ndel(valid_positions)\ndel(xy_features)\ndel(np_array)\ndel(train_b_near_scoring_net)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above, we can see that team A is trying to score on the net closest to the 250 Y position. Team B is trying to score on the net closest to the 0 Y position. We can also take a look at the opposite - where team B is defending, and where team A is defending right before scoring on either side.","metadata":{}},{"cell_type":"code","source":"xy_features = [\n    \"p0_pos_x\", \"p0_pos_y\", \"p1_pos_x\", \"p1_pos_y\", \"p2_pos_x\", \"p2_pos_y\", \"p3_pos_x\", \"p3_pos_y\", \n    \"p4_pos_x\", \"p4_pos_y\", \"p5_pos_x\", \"p5_pos_y\", \"ball_pos_x\", \"ball_pos_y\"\n]\n\ntrain_a_near_scoring_net = train[(train[\"team_A_scoring_within_10sec\"] == 1) & (train[\"event_time\"] > -1.)]\n\nvalid_positions = []\n\nfor (_, p0_x, p0_y, p1_x, p1_y, p2_x, p2_y, p3_x, p3_y, p4_x, p4_y, p5_x, p5_y, ball_x, ball_y) in train_a_near_scoring_net[xy_features].itertuples():\n    if p3_x < 250:\n        valid_positions.append((p3_x, p3_y))\n    if p4_x < 250:\n        valid_positions.append((p4_x, p4_y))\n    if p5_x < 250:\n        valid_positions.append((p5_x, p5_y))\n    # valid_positions.append((ball_x, ball_y))\n\nnp_array = np.zeros(shape=(255, 170))\nfor (x, y) in valid_positions:\n    np_array[y, x] += 1.\n\nfig, axs = plt.subplots(nrows=1, ncols=2, figsize=(20, 10), sharey=True)\n\n_ = sns.heatmap(np_array, ax=axs[0])\n_ = axs[0].grid(False)\n_ = axs[0].set_title(\"Team B Defending Positions\", fontweight=\"bold\", size=15)\n\n# Clean up unused variables to save on RAM\ndel(valid_positions)\ndel(np_array)\ndel(train_a_near_scoring_net)\n_  = gc.collect()\n\ntrain_b_near_scoring_net = train[(train[\"team_B_scoring_within_10sec\"] == 1) & (train[\"event_time\"] > -1.)]\n\nvalid_positions = []\n\nfor (_, p0_x, p0_y, p1_x, p1_y, p2_x, p2_y, p3_x, p3_y, p4_x, p4_y, p5_x, p5_y, ball_x, ball_y) in train_b_near_scoring_net[xy_features].itertuples():\n    if p0_x < 250:\n        valid_positions.append((p0_x, p0_y))\n    if p1_x < 250:\n        valid_positions.append((p1_x, p1_y))\n    if p2_x < 250:\n        valid_positions.append((p2_x, p2_y))\n    # valid_positions.append((ball_x, ball_y))\n\nnp_array = np.zeros(shape=(255, 170))\nfor (x, y) in valid_positions:\n    np_array[y, x] += 1.\n    \n_ = sns.heatmap(np_array, ax=axs[1])\n_ = axs[1].grid(False)\n_ = axs[1].set_title(\"Team A Defending Positions\", fontweight=\"bold\", size=15)\n\n# Clean up unused variables to save on RAM\ndel(valid_positions)\ndel(xy_features)\ndel(np_array)\ndel(train_b_near_scoring_net)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Key Observations About Player Positions\n\n* Team A is trying to score on the net located near the 250 Y position.\n* Team B is trying to score on the net located near the 0 Y position.\n* Team A and B both tend to cluster near the openings of their nets when defending against potential shots on the goal.","metadata":{}},{"cell_type":"markdown","source":"# 2.4 Statistical Breakdown\n\nLet's take a closer look at some of the statistical properties of the various features. Since trying to analyze 21+ million rows is likely going to cause memory problems, we'll examine a random sample of the dataset.\n\nBefore we go too far, we'll have to first deal with the null values in the velocity statistics (remember, we already dealt with null values in player and ball positions).Before we go too far, we'll have to first deal with the null values in the velocity statistics (remember, we already dealt with null values in player and ball positions).","metadata":{}},{"cell_type":"code","source":"velocity_features = [feature for feature in train.columns if feature.endswith(\"vel_z\") or feature.endswith(\"vel_x\") or feature.endswith(\"vel_y\")]\nfor feature in velocity_features:\n    train[feature] = train[feature].fillna(0.0)\n    \nboost_features = [feature for feature in train.columns if feature.endswith(\"boost\")]\nfor feature in boost_features:\n    train[feature] = train[feature].fillna(0.0)","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can look at statistical properties.","metadata":{}},{"cell_type":"code","source":"features = [\n    \"ball_pos_x\", \"ball_pos_y\", \"ball_pos_z\", \"ball_vel_x\", \"ball_vel_y\", \"ball_vel_z\", \n    \"p0_pos_x\", \"p0_pos_y\", \"p0_pos_z\", \"p0_vel_x\", \"p0_vel_y\", \"p0_vel_z\", \n    \"p1_pos_x\", \"p1_pos_y\", \"p1_pos_z\", \"p1_vel_x\", \"p1_vel_y\", \"p1_vel_z\", \n    \"p2_pos_x\", \"p2_pos_y\", \"p2_pos_z\", \"p2_vel_x\", \"p2_vel_y\", \"p2_vel_z\",\n    \"p3_pos_x\", \"p3_pos_y\", \"p3_pos_z\", \"p3_vel_x\", \"p3_vel_y\", \"p3_vel_z\", \n    \"p4_pos_x\", \"p4_pos_y\", \"p4_pos_z\", \"p4_vel_x\", \"p4_vel_y\", \"p4_vel_z\",\n    \"p5_pos_x\", \"p5_pos_y\", \"p5_pos_z\", \"p5_vel_x\", \"p5_vel_y\", \"p5_vel_z\",\n    \"p0_boost\", \"p1_boost\", \"p2_boost\", \"p3_boost\", \"p4_boost\", \"p5_boost\",\n    \"boost0_timer\", \"boost1_timer\", \"boost2_timer\", \"boost3_timer\", \"boost4_timer\", \"boost5_timer\",\n]\n\ntrain_subset = train.sample(frac=0.1, random_state=2022)\ntrain_subset = train_subset[(train_subset[\"p0_pos_x\"] != 250) & (train_subset[\"p1_pos_x\"] != 250) & (train_subset[\"p2_pos_x\"] != 250) & (train_subset[\"p3_pos_x\"] != 250) & (train_subset[\"p4_pos_x\"] != 250) & (train_subset[\"p5_pos_x\"] != 250) & (train_subset[\"ball_pos_x\"] != 250)]\ntrain_subset[features].describe().T.style.bar(subset=['mean'], color='#7BCC70')\\\n    .background_gradient(subset=['std'], cmap='Reds')\\\n    .background_gradient(subset=['50%'], cmap='coolwarm')","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Clean up unused variables to save on RAM\ndel(train_subset)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Key Observations About Statistical Information\n\n* The X position of each player has a mean that centers around the center of the field. This makes sense, since we assume that a random sample of players will tend to play both the right and left hand sides of the field somewhat equally. No big spikes exist, which suggests that there are no anomalies in the play field that may force players to favor one side over the other.\n* The Y position of the players has a mean depending on their team. Players on team A tend to have a mean closer to their end of the field, while players on team B have a mean closer to their end instead. This makes sense, since after scoring, player positions get reset to positions on their own half of the field before a kickoff occurs.\n* The Z position of players is interesting. On average, players remain on the ground most of the time. In fact, players remain firmly on the ground even within the 75% percentile, suggesting that jumps are somewhat infrequent or short in duration. Surprisingly, player data maxes out to the top of the play field in the 100% percentile, suggesting that while players don't spend a lot of time there, it is possible that they can reach and possibly drive on the ceiling.\n* Ball position appears to be similar to that of player position with respect to X and Y axis. In the Z axis however, we see the ball is even distributed across the vertical space. This makes sense since the ball when hit may have more variable trajectories, and therefore may travel through a much wider range of Z positions.","metadata":{}},{"cell_type":"markdown","source":"# 2.5 P-Value Test\n\nWhile looking at features visually will tell us some interesting information, we can also use p-value testing to see if a feature has a net impact on a simple regression model. This method is controversial in that it likely doesn't provide a correct look at what features are informative. Our null hypothesis is that the feature impacts the target variable of `team_A_scoring_within_10sec` or `team_B_scoring_within_10sec`. In this case, anything with a p-value greater than 0.05 means we reject that hypothesis, and can potentially flag it for removal. Again, due to memory pressure, we'll go with the subset of our training data for analysis.","metadata":{}},{"cell_type":"markdown","source":"### P-Value Test for `team_A_scoring_within_10sec`","metadata":{}},{"cell_type":"code","source":"from statsmodels.regression.linear_model import OLS\nfrom statsmodels.tools.tools import add_constant\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\ntrain_subset = train.sample(frac=0.05, random_state=2022)\n\nx = add_constant(train_subset[features])\nmodel = OLS(train_subset[\"team_A_scoring_within_10sec\"], x).fit()\n\npvalues = pd.DataFrame(model.pvalues)\npvalues.reset_index(inplace=True)\npvalues.rename(columns={0: \"pvalue\", \"index\": \"feature\"}, inplace=True)\npvalues.style.background_gradient(cmap='YlOrRd')","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some interesting observations:\n\n* The P-Value test reveals that any X velocity or X position is likely uninformative with respect to whether team A scores within the next 10 seconds. This reflects the nature of the playing field. The goals are oriented at either end of the Y axis, so the ball's Y position and the player's Y positions are very much related to whether the team can score. The height of the net in the Z axis determines how high up the ball can go before it goes over the net, so again, the ball's Z position and other player's Z positions relate to whether the team can score. The X axis however, is largely unimportant. This suggests the ball can originate from either the left or right of the goal and still have a very good chance of going into the goal. ","metadata":{}},{"cell_type":"code","source":"# Clean up unused variables to save on RAM\ndel(train_subset)\ndel(pvalues)\ndel(model)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### P-Value Test for `team_B_scoring_within_10sec`","metadata":{}},{"cell_type":"code","source":"train_subset = train.sample(frac=0.05, random_state=2022)\n\nx = add_constant(train_subset[features])\nmodel = OLS(train_subset[\"team_B_scoring_within_10sec\"], x).fit()\n\npvalues = pd.DataFrame(model.pvalues)\npvalues.reset_index(inplace=True)\npvalues.rename(columns={0: \"pvalue\", \"index\": \"feature\"}, inplace=True)\npvalues.style.background_gradient(cmap='YlOrRd')","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some interesting observations:\n\n* The same applies to the team B scoring data as to team A scoring within 10 seconds. The X position and X velocity of the ball and players do not matter. Again, it's the Y and Z position and velocities that matter. ","metadata":{}},{"cell_type":"code","source":"# Clean up unused variables to save on RAM\ndel(train_subset)\ndel(pvalues)\ndel(model)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Key Observations About P-Value Tests\n\n* We should experiment with removing the following features:\n * `p0_pos_x`\n * `p0_vel_x`\n * `p1_pos_x`\n * `p1_vel_x`\n * `p2_pos_x`\n * `p2_vel_x`\n * `p3_pos_x`\n * `p3_vel_x`\n * `p4_pos_x`\n * `p4_vel_x`\n * `p5_pos_x`\n * `p5_vel_x`\n * `ball_pos_x`\n * `ball_vel_x`","metadata":{}},{"cell_type":"markdown","source":"# 2.6 Spearman Correlation\n\nWe should also check to see what variables are correlated to one another. Spearman correlation does not make assumptions about distribution types or linearity. With Spearman correlation, we have values that range from -1 to +1. Values around either extreme end mean a neagative or positive correlation, while those around 0 mean no correlation exists.","metadata":{}},{"cell_type":"markdown","source":"### Spearman Correlation for `team_A_scoring_within_10sec`","metadata":{}},{"cell_type":"code","source":"columns_to_check = features.copy()\ncolumns_to_check.append(\"team_A_scoring_within_10sec\")\ntrain_subset = train.sample(frac=0.05, random_state=2022)\ncorrelation_matrix = train_subset[columns_to_check].corr(method=\"spearman\")\n\nfrom matplotlib.colors import SymLogNorm\n\nf, ax = plt.subplots(figsize=(20, 20))\ncmap = sns.diverging_palette(230, 20, as_cmap=True)\n_ = sns.heatmap(\n    correlation_matrix, \n    mask=np.triu(np.ones_like(correlation_matrix, dtype=bool)), \n    cmap=sns.diverging_palette(230, 20, as_cmap=True), \n    center=0,\n    square=True, \n    linewidths=.1, \n    cbar_kws={\"shrink\": .2},\n    norm=SymLogNorm(linthresh=0.03, linscale=0.03, vmin=-1.0, vmax=1.0, base=10),\n)\n\n# Clean up unused variables to save on RAM\ndel(train_subset)\ndel(columns_to_check)\ndel(correlation_matrix)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some observations:\n\n* Each of the player's X, Y, and Z positions have a strong positive correlation with the ball's X, Y, and Z positions. This makes sense since players typically follow the ball around the field in order to take possession of it to take shots on the goal.\n* The player's X and Y positions are also positively correlated with one another. Again, this makes sense. Players on the same team will tend to remain close to one another, since they can help each other gain control of the ball or defend their goal. \n* The player's Z positions are not positively or negatively correlated. This means that movement in the Z axis by any player is pretty much of their own volition, and is independent of what other players do.\n* There is some weak positive correlation between the Y positions of the ball and players in conjunction with the target variable. This makes sense, since most of the players will be close to the goal when it happens.","metadata":{}},{"cell_type":"markdown","source":"### Spearman Correlation for `team_B_scoring_within_10sec`","metadata":{}},{"cell_type":"code","source":"columns_to_check = features.copy()\ncolumns_to_check.append(\"team_B_scoring_within_10sec\")\ntrain_subset = train.sample(frac=0.05, random_state=2022)\ncorrelation_matrix = train_subset[columns_to_check].corr(method=\"spearman\")\n\nfrom matplotlib.colors import SymLogNorm\n\nf, ax = plt.subplots(figsize=(20, 20))\ncmap = sns.diverging_palette(230, 20, as_cmap=True)\n_ = sns.heatmap(\n    correlation_matrix, \n    mask=np.triu(np.ones_like(correlation_matrix, dtype=bool)), \n    cmap=sns.diverging_palette(230, 20, as_cmap=True), \n    center=0,\n    square=True, \n    linewidths=.1, \n    cbar_kws={\"shrink\": .2},\n    norm=SymLogNorm(linthresh=0.03, linscale=0.03, vmin=-1.0, vmax=1.0, base=10),\n)\n\n# Clean up unused variables to save on RAM\ndel(train_subset)\ndel(columns_to_check)\ndel(correlation_matrix)\n_  = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Again, we have the same observations as above with regards to correlation between variables and targets.","metadata":{}},{"cell_type":"markdown","source":"### Key Observations About Spearman Correlations\n\n* There is strong positive correlation between ball position and player positions along the X, Y, and Z axis. This makes sense since the players will be close to where the ball is.\n* There is weakly positive correlation between the target variable and the Y position of both the ball and the players. This suggests that the Y coordinates may be slightly more informative than other features.","metadata":{}},{"cell_type":"markdown","source":"# 2.7 Class Balance\n\nAs always, it is a good idea to take a look at class balance to see whether there is a skew in the target variable that we are trying to predict. ","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(15, 5))\n\ncounts_a = pd.DataFrame(train[\"team_A_scoring_within_10sec\"].value_counts())\n_ = sns.barplot(x=counts_a.index, y=counts_a.team_A_scoring_within_10sec, ax=axs[0], **sns_params)\nfor p in axs[0].patches:\n    axs[0].text(x=p.get_x()+(p.get_width()/2), y=p.get_height(), s=\"{:,d}\".format(round(p.get_height())), ha=\"center\")\n_ = axs[0].set_title(\"Class Balance for Team A Scoring\", fontsize=15)\n_ = axs[0].set_ylabel(\"Number of Records\", fontsize=15)\n_ = axs[0].set_xlabel(\"Class\", fontsize=15)\n\ncounts_b = pd.DataFrame(train[\"team_B_scoring_within_10sec\"].value_counts())\n_ = sns.barplot(x=counts_b.index, y=counts_b.team_B_scoring_within_10sec, ax=axs[1], **sns_params)\nfor p in axs[1].patches:\n    axs[1].text(x=p.get_x()+(p.get_width()/2), y=p.get_height(), s=\"{:,d}\".format(round(p.get_height())), ha=\"center\")\n_ = axs[1].set_title(\"Class Balance for Team B Scoring\", fontsize=15)\n_ = axs[1].set_ylabel(\"Number of Records\", fontsize=15)\n_ = axs[1].set_xlabel(\"Class\", fontsize=15)\n\ndel(counts_a)\ndel(counts_b)\n_ = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Key Observations about Class Balance\n\n* Both scoring metrics are very imbalanced with respect to the negative class. In both cases, the positive examples account for approximately 5% of the data available. \n* In order to help with the class balance problem, we can perform _data augmentation_. In slightly more detail, to add more examples for team A, we can take all the positive examples from team B, swap the players from team B to team A, flip the X and Y coordinates, and add the result to the training set. ","metadata":{}},{"cell_type":"markdown","source":"# 3 Simple Models\n\nNow that we have examined the data in more detail, we should look at building some simple models to understand how various features may impact our ability to make predictions.","metadata":{}},{"cell_type":"markdown","source":"# 3.1 Unmodified LightGBM\n\nWe'll start by building a simple LightGBM model without any feature engineering. We're using LightGBM here as a decently performing GBDT that is CPU bound so that we don't have to use any of our allocated GPU time just for simple EDA purposes.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import classification_report\nfrom lightgbm import LGBMClassifier\nfrom sklearn.metrics import log_loss\nfrom sklearn.metrics import confusion_matrix\nfrom lightgbm import log_evaluation\nfrom lightgbm import early_stopping\n\n# Follow the competition's metric calculation\ndef rocket_logloss(target, probas):\n    observation_loss_a = sum([(1 if int(truth) == 1 else 0)*log(proba[0]) + (1-(1 if int(truth) == 1 else 0))*log(1-proba[0]) for truth, proba in zip(target, probas)]) / len(target)\n    observation_loss_b = sum([(1 if int(truth) == 2 else 0)*log(proba[1]) + (1-(1 if int(truth) == 2 else 0))*log(1-proba[1]) for truth, proba in zip(target, probas)]) / len(target)\n    return -(0.5 * (observation_loss_a + observation_loss_b))\n\ntrain[\"combined_target\"] = train[[\"team_A_scoring_within_10sec\", \"team_B_scoring_within_10sec\"]].apply(lambda x: \"{}{}\".format(str(x.team_A_scoring_within_10sec), str(x.team_B_scoring_within_10sec)), axis=1)\ntrain[\"combined_target\"] = train[\"combined_target\"].apply(lambda x: \"1\" if x == \"10\" else x)\ntrain[\"combined_target\"] = train[\"combined_target\"].apply(lambda x: \"2\" if x == \"01\" else x)\ntrain[\"combined_target\"] = train[\"combined_target\"].apply(lambda x: \"3\" if x == \"00\" else x)\ntrain[\"combined_target\"] = train[\"combined_target\"].apply(lambda x: int(x))\ntrain[\"combined_target\"] = train[\"combined_target\"].astype(np.int8)\n\ntrain_subset = train.sample(frac=0.05, random_state=2022)\ntarget = train_subset[\"combined_target\"]\n\nclass_labels = [\"A Scores\", \"B Scores\", \"No Score\"]\n\ncv_rounds = 3\n\nk_fold = StratifiedKFold(\n    n_splits=cv_rounds,\n    random_state=2021,\n    shuffle=True,\n)\n\ntrain_preds = np.zeros(len(train_subset.index), )\ntrain_probas = np.zeros((len(train_subset.index), 3), )\n\nfor fold, (train_index, test_index) in enumerate(k_fold.split(train_subset[features], target)):\n    print(\"-- Fold {}:\".format(fold+1))\n\n    x_train = train_subset[features].iloc[train_index]\n    y_train = target.iloc[train_index]\n\n    x_valid = train_subset[features].iloc[test_index]\n    y_valid = target.iloc[test_index]\n\n    model = LGBMClassifier(\n        random_state=2022,\n        n_estimators=2000,\n        verbose=-1,\n        metric=\"multi_logloss\",\n        objective=\"multiclass\",\n    )\n    model.fit(\n        x_train,\n        y_train,\n        eval_set=[(x_valid, y_valid)],\n        callbacks=[log_evaluation(period=0), early_stopping(50, verbose=False)],\n    )\n    train_oof_preds = model.predict(x_valid)\n    train_oof_probas = model.predict_proba(x_valid)\n    train_preds[test_index] = train_oof_preds\n    train_probas[test_index] = train_oof_probas\n    \n    del(model)\n    gc.collect()\n    \n    print(\"{}\".format(classification_report(y_valid, train_oof_preds, target_names=class_labels)))\n    \nunmodified_rocket_logloss = rocket_logloss(target, train_probas)\n\nprint(\"-- Overall:\")\nprint(\"{}\".format(classification_report(target, train_preds, target_names=class_labels)))\nprint(\"-- Model log loss: {}\".format(log_loss(target, train_probas)))\nprint(\"-- Competition log loss: {}\".format(unmodified_rocket_logloss))\n\n# Show the confusion matrix\nconfusion = confusion_matrix(target, train_preds)\nax = sns.heatmap(confusion, annot=True, fmt=\",d\", xticklabels=class_labels, yticklabels=class_labels)\n_ = ax.set_title(\"Confusion Matrix for LGB Classifier (Unmodified Dataset)\", fontsize=15)\n_ = ax.set_ylabel(\"Actual Class\")\n_ = ax.set_xlabel(\"Predicted Class\")\n\n# Delete unused variables to save RAM\ndel(train_preds)\ndel(train_probas)\ndel(confusion)\n_ = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3.2 Removing X Positions\n\nAs indicated by the P-Value tests above, it may be possible to remove X position information without impacting model performance. This would purely be a memory saving feature.","metadata":{}},{"cell_type":"code","source":"features.remove(\"p0_pos_x\")\nfeatures.remove(\"p1_pos_x\")\nfeatures.remove(\"p2_pos_x\")\nfeatures.remove(\"p3_pos_x\")\nfeatures.remove(\"p4_pos_x\")\nfeatures.remove(\"p5_pos_x\")\n\nk_fold = StratifiedKFold(\n    n_splits=cv_rounds,\n    random_state=2021,\n    shuffle=True,\n)\n\ntrain_preds = np.zeros(len(train_subset.index), )\ntrain_probas = np.zeros((len(train_subset.index), 3), )\n\nfor fold, (train_index, test_index) in enumerate(k_fold.split(train_subset[features], target)):\n    print(\"-- Fold {}:\".format(fold+1))\n\n    x_train = train_subset[features].iloc[train_index]\n    y_train = target.iloc[train_index]\n\n    x_valid = train_subset[features].iloc[test_index]\n    y_valid = target.iloc[test_index]\n\n    model = LGBMClassifier(\n        random_state=2022,\n        n_estimators=2000,\n        verbose=-1,\n        metric=\"multi_logloss\",\n        objective=\"multiclass\",\n    )\n    model.fit(\n        x_train,\n        y_train,\n        eval_set=[(x_valid, y_valid)],\n        callbacks=[log_evaluation(period=0), early_stopping(50, verbose=False)],\n    )\n    train_oof_preds = model.predict(x_valid)\n    train_oof_probas = model.predict_proba(x_valid)\n    train_preds[test_index] = train_oof_preds\n    train_probas[test_index] = train_oof_probas\n    \n    del(model)\n    gc.collect()\n    \n    print(\"{}\".format(classification_report(y_valid, train_oof_preds, target_names=class_labels)))\n    \nremoved_x_pos_rocket_logloss = rocket_logloss(target, train_probas)\n\nprint(\"-- Overall:\")\nprint(\"{}\".format(classification_report(target, train_preds, target_names=class_labels)))\nprint(\"-- Model log loss: {}\".format(log_loss(target, train_probas)))\nprint(\"-- Competition log loss: {}\".format(removed_x_pos_rocket_logloss))\n\n# Show the confusion matrix\nconfusion = confusion_matrix(target, train_preds)\nax = sns.heatmap(confusion, annot=True, fmt=\",d\", xticklabels=class_labels, yticklabels=class_labels)\n_ = ax.set_title(\"Confusion Matrix for LGB Classifier (Float to Int)\", fontsize=15)\n_ = ax.set_ylabel(\"Actual Class\")\n_ = ax.set_xlabel(\"Predicted Class\")\n\n# Delete unused variables to save RAM\ndel(train_preds)\ndel(train_probas)\ndel(confusion)\n_ = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Overall, removing X positions from the dataset increases log loss, thus we should leave in the values regardless of what the P-Test suggests.","metadata":{}},{"cell_type":"markdown","source":"# 3.3 Counting Players in Play\n\nKeeping track of what players are still in play may also help with keeping track of score information. We can keep a simple count of how many players are in play at any given moment based on their substituted X position when they were null (demolished).","metadata":{}},{"cell_type":"code","source":"train_subset[\"team_A_num_players\"] = train_subset[[\"p0_pos_x\", \"p1_pos_x\", \"p2_pos_x\"]].apply(\n    lambda x: (1 if x.p0_pos_x != 250 else 0) + (1 if x.p1_pos_x != 250 else 0) + (1 if x.p2_pos_x != 250 else 0)\n, axis=1)\n\ntrain_subset[\"team_B_num_players\"] = train_subset[[\"p3_pos_x\", \"p4_pos_x\", \"p5_pos_x\"]].apply(\n    lambda x: (1 if x.p3_pos_x != 250 else 0) + (1 if x.p4_pos_x != 250 else 0) + (1 if x.p5_pos_x != 250 else 0)\n, axis=1)","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features.append(\"p0_pos_x\")\nfeatures.append(\"p1_pos_x\")\nfeatures.append(\"p2_pos_x\")\nfeatures.append(\"p3_pos_x\")\nfeatures.append(\"p4_pos_x\")\nfeatures.append(\"p5_pos_x\")\nfeatures.append(\"team_A_num_players\")\nfeatures.append(\"team_B_num_players\")\n\nk_fold = StratifiedKFold(\n    n_splits=cv_rounds,\n    random_state=2021,\n    shuffle=True,\n)\n\ntrain_preds = np.zeros(len(train_subset.index), )\ntrain_probas = np.zeros((len(train_subset.index), 3), )\n\nfor fold, (train_index, test_index) in enumerate(k_fold.split(train_subset[features], target)):\n    print(\"-- Fold {}:\".format(fold+1))\n\n    x_train = train_subset[features].iloc[train_index]\n    y_train = target.iloc[train_index]\n\n    x_valid = train_subset[features].iloc[test_index]\n    y_valid = target.iloc[test_index]\n\n    model = LGBMClassifier(\n        random_state=2022,\n        n_estimators=2000,\n        verbose=-1,\n        metric=\"multi_logloss\",\n        objective=\"multiclass\",\n    )\n    model.fit(\n        x_train,\n        y_train,\n        eval_set=[(x_valid, y_valid)],\n        callbacks=[log_evaluation(period=0), early_stopping(50, verbose=False)],\n    )\n    train_oof_preds = model.predict(x_valid)\n    train_oof_probas = model.predict_proba(x_valid)\n    train_preds[test_index] = train_oof_preds\n    train_probas[test_index] = train_oof_probas\n    \n    del(model)\n    gc.collect()\n    \n    print(\"{}\".format(classification_report(y_valid, train_oof_preds, target_names=class_labels)))\n    \nplayer_count_rocket_logloss = rocket_logloss(target, train_probas)\n\nprint(\"-- Overall:\")\nprint(\"{}\".format(classification_report(target, train_preds, target_names=class_labels)))\nprint(\"-- Model log loss: {}\".format(log_loss(target, train_probas)))\nprint(\"-- Competition log loss: {}\".format(player_count_rocket_logloss))\n\n# Show the confusion matrix\nconfusion = confusion_matrix(target, train_preds)\nax = sns.heatmap(confusion, annot=True, fmt=\",d\", xticklabels=class_labels, yticklabels=class_labels)\n_ = ax.set_title(\"Confusion Matrix for LGB Classifier (Player Count)\", fontsize=15)\n_ = ax.set_ylabel(\"Actual Class\")\n_ = ax.set_xlabel(\"Predicted Class\")\n\n# Delete unused variables to save RAM\ndel(train_preds)\ndel(train_probas)\ndel(confusion)\n_ = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Counting the number of players in play has a slightly lower log loss than our baseline of an unmodified LGBM model.","metadata":{}},{"cell_type":"markdown","source":"# 3.4 Indicating Player Positions\n\nAn additional piece of information that could be useful is to indicate where in the field the players are located. For example, we can record if the defender is in their own goal posts. This may give an indication as to whether players are in a position to strike or defend.","metadata":{}},{"cell_type":"code","source":"features.remove(\"team_A_num_players\")\nfeatures.remove(\"team_B_num_players\")\n\ntrain_subset['team_b_players_in_goal'] = train_subset.apply(\n    lambda x: (1 if x.p3_pos_y > 222 else 0) + (1 if x.p4_pos_y > 222 else 0) + (1 if x.p5_pos_y > 222 else 0)\n, axis=1)\ntrain_subset['team_a_players_in_goal'] = train_subset.apply(\n    lambda x: (1 if x.p0_pos_y < 13 else 0) + (1 if x.p1_pos_y < 13 else 0) + (1 if x.p2_pos_y < 13 else 0)\n, axis=1)\ntrain_subset['team_b_players_defending'] = train_subset.apply(\n    lambda x: (1 if x.p3_pos_y <= 222 and x.p3_pos_y > 185 else 0) + (1 if x.p4_pos_y <= 222 and x.p4_pos_y > 185 else 0) + (1 if x.p5_pos_y <= 222 and x.p5_pos_y > 185 else 0)\n, axis=1)\ntrain_subset['team_a_players_defending'] = train_subset.apply(\n    lambda x: (1 if x.p0_pos_y >= 13 and x.p0_pos_y < 75 else 0) + (1 if x.p1_pos_y >= 13 and x.p1_pos_y < 75 else 0) + (1 if x.p2_pos_y >= 13 and x.p2_pos_y < 75 else 0)\n, axis=1)\ntrain_subset['team_b_players_striking'] = train_subset.apply(\n    lambda x: (1 if x.p3_pos_y >= 13 and x.p3_pos_y < 75 else 0) + (1 if x.p4_pos_y >= 13 and x.p4_pos_y < 75 else 0) + (1 if x.p5_pos_y >= 13 and x.p5_pos_y < 75 else 0)\n, axis=1)\ntrain_subset['team_a_players_striking'] = train_subset.apply(\n    lambda x: (1 if x.p0_pos_y <= 222 and x.p0_pos_y > 185 else 0) + (1 if x.p1_pos_y <= 222 and x.p1_pos_y > 185 else 0) + (1 if x.p2_pos_y <= 222 and x.p2_pos_y > 185 else 0)\n, axis=1)\n\nfeatures.append(\"team_b_players_in_goal\")\nfeatures.append(\"team_a_players_in_goal\")\nfeatures.append(\"team_b_players_defending\")\nfeatures.append(\"team_a_players_defending\")\nfeatures.append(\"team_b_players_striking\")\nfeatures.append(\"team_a_players_striking\")\n\nk_fold = StratifiedKFold(\n    n_splits=cv_rounds,\n    random_state=2021,\n    shuffle=True,\n)\n\ntrain_preds = np.zeros(len(train_subset.index), )\ntrain_probas = np.zeros((len(train_subset.index), 3), )\n\nfor fold, (train_index, test_index) in enumerate(k_fold.split(train_subset[features], target)):\n    print(\"-- Fold {}:\".format(fold+1))\n\n    x_train = train_subset[features].iloc[train_index]\n    y_train = target.iloc[train_index]\n\n    x_valid = train_subset[features].iloc[test_index]\n    y_valid = target.iloc[test_index]\n\n    model = LGBMClassifier(\n        random_state=2022,\n        n_estimators=2000,\n        verbose=-1,\n        metric=\"multi_logloss\",\n        objective=\"multiclass\",\n    )\n    model.fit(\n        x_train,\n        y_train,\n        eval_set=[(x_valid, y_valid)],\n        callbacks=[log_evaluation(period=0), early_stopping(50, verbose=False)],\n    )\n    train_oof_preds = model.predict(x_valid)\n    train_oof_probas = model.predict_proba(x_valid)\n    train_preds[test_index] = train_oof_preds\n    train_probas[test_index] = train_oof_probas\n    \n    del(model)\n    gc.collect()\n    \n    print(\"{}\".format(classification_report(y_valid, train_oof_preds, target_names=class_labels)))\n    \nposition_rocket_logloss = rocket_logloss(target, train_probas)\n\nprint(\"-- Overall:\")\nprint(\"{}\".format(classification_report(target, train_preds, target_names=class_labels)))\nprint(\"-- Model log loss: {}\".format(log_loss(target, train_probas)))\nprint(\"-- Competition log loss: {}\".format(position_rocket_logloss))\n\n# Show the confusion matrix\nconfusion = confusion_matrix(target, train_preds)\nax = sns.heatmap(confusion, annot=True, fmt=\",d\", xticklabels=class_labels, yticklabels=class_labels)\n_ = ax.set_title(\"Confusion Matrix for LGB Classifier (Player Positions)\", fontsize=15)\n_ = ax.set_ylabel(\"Actual Class\")\n_ = ax.set_xlabel(\"Predicted Class\")\n\n# Delete unused variables to save RAM\ndel(train_preds)\ndel(train_probas)\ndel(confusion)\n_ = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3.5 Extrapolating Positions\n\nGiven that we have velocity information relating to each player and the ball, we can extrapolate the next move (or next few moves) for the player. This may help give us clues as to what we think is going to happen next. This extrapolation is necessarily limited however, given the fact that we don't have acceleration information, or much in the way (currently) of calculating the physical forces of the game world. This means that our extrapolation will become wildly inaccurate the further into the future we attempt to make predictions. However, it may provide an immediate boost one or two steps into the future.","metadata":{}},{"cell_type":"code","source":"print(\": Extrapolating new player 0 positions\")\ntrain_subset[\"p0_pos_x_step1\"] = train_subset.apply(lambda x: int(x.p0_pos_x + (x.p0_vel_x / 10.)), axis=1)\ntrain_subset[\"p0_pos_y_step1\"] = train_subset.apply(lambda x: int(x.p0_pos_y + (x.p0_vel_y / 10.)), axis=1)\ntrain_subset[\"p0_pos_z_step1\"] = train_subset.apply(lambda x: int(x.p0_pos_z + (x.p0_vel_z / 10.)), axis=1)\n\nprint(\": Extrapolating new player 1 positions\")\ntrain_subset[\"p1_pos_x_step1\"] = train_subset.apply(lambda x: int(x.p1_pos_x + (x.p1_vel_x / 10.)), axis=1)\ntrain_subset[\"p1_pos_y_step1\"] = train_subset.apply(lambda x: int(x.p1_pos_y + (x.p1_vel_y / 10.)), axis=1)\ntrain_subset[\"p1_pos_z_step1\"] = train_subset.apply(lambda x: int(x.p1_pos_z + (x.p1_vel_z / 10.)), axis=1)\n\nprint(\": Extrapolating new player 2 positions\")\ntrain_subset[\"p2_pos_x_step1\"] = train_subset.apply(lambda x: int(x.p2_pos_x + (x.p2_vel_x / 10.)), axis=1)\ntrain_subset[\"p2_pos_y_step1\"] = train_subset.apply(lambda x: int(x.p2_pos_y + (x.p2_vel_y / 10.)), axis=1)\ntrain_subset[\"p2_pos_z_step1\"] = train_subset.apply(lambda x: int(x.p2_pos_z + (x.p2_vel_z / 10.)), axis=1)\n\nprint(\": Extrapolating new player 3 positions\")\ntrain_subset[\"p3_pos_x_step1\"] = train_subset.apply(lambda x: int(x.p3_pos_x + (x.p3_vel_x / 10.)), axis=1)\ntrain_subset[\"p3_pos_y_step1\"] = train_subset.apply(lambda x: int(x.p3_pos_y + (x.p3_vel_y / 10.)), axis=1)\ntrain_subset[\"p3_pos_z_step1\"] = train_subset.apply(lambda x: int(x.p3_pos_z + (x.p3_vel_z / 10.)), axis=1)\n\nprint(\": Extrapolating new player 4 positions\")\ntrain_subset[\"p4_pos_x_step1\"] = train_subset.apply(lambda x: int(x.p4_pos_x + (x.p4_vel_x / 10.)), axis=1)\ntrain_subset[\"p4_pos_y_step1\"] = train_subset.apply(lambda x: int(x.p4_pos_y + (x.p4_vel_y / 10.)), axis=1)\ntrain_subset[\"p4_pos_z_step1\"] = train_subset.apply(lambda x: int(x.p4_pos_z + (x.p4_vel_z / 10.)), axis=1)\n\nprint(\": Extrapolating new player 5 positions\")\ntrain_subset[\"p5_pos_x_step1\"] = train_subset.apply(lambda x: int(x.p5_pos_x + (x.p5_vel_x / 10.)), axis=1)\ntrain_subset[\"p5_pos_y_step1\"] = train_subset.apply(lambda x: int(x.p5_pos_y + (x.p5_vel_y / 10.)), axis=1)\ntrain_subset[\"p5_pos_z_step1\"] = train_subset.apply(lambda x: int(x.p5_pos_z + (x.p5_vel_z / 10.)), axis=1)\n\nprint(\": Extrapolating new ball positions\")\ntrain_subset[\"ball_pos_x_step1\"] = train_subset.apply(lambda x: int(x.ball_pos_x + (x.ball_vel_x / 10.)), axis=1)\ntrain_subset[\"ball_pos_y_step1\"] = train_subset.apply(lambda x: int(x.ball_pos_y + (x.ball_vel_y / 10.)), axis=1)\ntrain_subset[\"ball_pos_z_step1\"] = train_subset.apply(lambda x: int(x.ball_pos_z + (x.ball_vel_z / 10.)), axis=1)","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features.remove(\"team_b_players_in_goal\")\nfeatures.remove(\"team_a_players_in_goal\")\nfeatures.remove(\"team_b_players_defending\")\nfeatures.remove(\"team_a_players_defending\")\nfeatures.remove(\"team_b_players_striking\")\nfeatures.remove(\"team_a_players_striking\")\n\nfeatures.append(\"p0_pos_x_step1\")\nfeatures.append(\"p0_pos_y_step1\")\nfeatures.append(\"p0_pos_z_step1\")\nfeatures.append(\"p1_pos_x_step1\")\nfeatures.append(\"p1_pos_y_step1\")\nfeatures.append(\"p1_pos_z_step1\")\nfeatures.append(\"p2_pos_x_step1\")\nfeatures.append(\"p2_pos_y_step1\")\nfeatures.append(\"p2_pos_z_step1\")\nfeatures.append(\"p3_pos_x_step1\")\nfeatures.append(\"p3_pos_y_step1\")\nfeatures.append(\"p3_pos_z_step1\")\nfeatures.append(\"p4_pos_x_step1\")\nfeatures.append(\"p4_pos_y_step1\")\nfeatures.append(\"p4_pos_z_step1\")\nfeatures.append(\"p5_pos_x_step1\")\nfeatures.append(\"p5_pos_y_step1\")\nfeatures.append(\"p5_pos_z_step1\")\nfeatures.append(\"ball_pos_x_step1\")\nfeatures.append(\"ball_pos_y_step1\")\nfeatures.append(\"ball_pos_z_step1\")","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"k_fold = StratifiedKFold(\n    n_splits=cv_rounds,\n    random_state=2021,\n    shuffle=True,\n)\n\ntrain_preds = np.zeros(len(train_subset.index), )\ntrain_probas = np.zeros((len(train_subset.index), 3), )\n\nfor fold, (train_index, test_index) in enumerate(k_fold.split(train_subset[features], target)):\n    print(\"-- Fold {}:\".format(fold+1))\n\n    x_train = train_subset[features].iloc[train_index]\n    y_train = target.iloc[train_index]\n\n    x_valid = train_subset[features].iloc[test_index]\n    y_valid = target.iloc[test_index]\n\n    model = LGBMClassifier(\n        random_state=2022,\n        n_estimators=2000,\n        verbose=-1,\n        metric=\"multi_logloss\",\n        objective=\"multiclass\",\n    )\n    model.fit(\n        x_train,\n        y_train,\n        eval_set=[(x_valid, y_valid)],\n        callbacks=[log_evaluation(period=0), early_stopping(50, verbose=False)],\n    )\n    train_oof_preds = model.predict(x_valid)\n    train_oof_probas = model.predict_proba(x_valid)\n    train_preds[test_index] = train_oof_preds\n    train_probas[test_index] = train_oof_probas\n    \n    del(model)\n    gc.collect()\n    \n    print(\"{}\".format(classification_report(y_valid, train_oof_preds, target_names=class_labels)))\n    \nextra_rocket_logloss = rocket_logloss(target, train_probas)\n\nprint(\"-- Overall:\")\nprint(\"{}\".format(classification_report(target, train_preds, target_names=class_labels)))\nprint(\"-- Model log loss: {}\".format(log_loss(target, train_probas)))\nprint(\"-- Competition log loss: {}\".format(extra_rocket_logloss))\n\n# Show the confusion matrix\nconfusion = confusion_matrix(target, train_preds)\nax = sns.heatmap(confusion, annot=True, fmt=\",d\", xticklabels=class_labels, yticklabels=class_labels)\n_ = ax.set_title(\"Confusion Matrix for LGB Classifier (Extrapolation)\", fontsize=15)\n_ = ax.set_ylabel(\"Actual Class\")\n_ = ax.set_xlabel(\"Predicted Class\")\n\n# Delete unused variables to save RAM\ndel(train_preds)\ndel(train_probas)\ndel(confusion)\n_ = gc.collect()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3.6 Comparison of Approaches\n\nOverall, we can take a look at what the difference is between the approaches in terms of log loss.","metadata":{}},{"cell_type":"code","source":"bar, ax = plt.subplots(figsize=(10, 5))\nax = sns.barplot(\n    x=[\"Unmodified\", \"Removing X Features\", \"Counting Players\", \"Player Positions\", \"Extrapolation\"],\n    y=[\n        unmodified_rocket_logloss,\n        removed_x_pos_rocket_logloss,\n        player_count_rocket_logloss,\n        position_rocket_logloss,\n        extra_rocket_logloss,\n    ],\n    ax=ax,\n    **sns_params\n)\n_ = ax.set_title(\"Competition Log Loss by Approach (Smaller = Better)\", fontsize=15)\n_ = ax.set_xlabel(\"Approach\")\n_ = ax.set_ylabel(\"Competition Log Loss\")\n_ = ax.set(ylim=(0.186, 0.188))\nfor p in ax.patches:\n    height = p.get_height()\n    ax.text(\n        x=p.get_x()+(p.get_width()/2),\n        y=height,\n        s=\"{:.4f}\".format(height),\n        ha=\"center\"\n    )","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4 Future Work\n\nFrom the data provided, there is plenty of opportunities for some future work. Specifically:\n\n* A better model for future predictions of positions is warranted. Given that it provided an advantage to the model, a newer way of calculating positions in relation to the actual definition of the field is needed. Additionally, future increments of player positions more than a single step ahead could provide further boost to the model. \n* Player collisions should also be modeled. Collisions could indicate a change in the player's direction or heading and may be important information to capture.\n* Chains of related rows could be determined by extrapolating player positions, and then finding rows in the data that match what those future positions should be. This would provide us a really good opportunity to use LSTM neural networks to take into account say 10 time slices from the board. This would likely greatly increase the accuracy of our machine learning models, at the expense of having to do a lot of pre-processing to match up current frames with future frames. \n* Stacking models is likely to provide increases in performance. \n* Model hyper-parameter tuning is likely to provide increases in performance.\n\nIf you found this EDA useful, please consider upvoting!","metadata":{}}]}