{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <span style=\"color:green; opacity: 0.70\"> TPS-OCT22, Gradient Boosted Desicion Trees. 🍃 </span>\n### <span style=\"color:gray; opacity: 0.8\"> Predict the Probability of Each Team Scoring Within the Next 10 Seconds of the Game. Sounds Awesome, Right? </span>\n&nbsp;\n\n<img src='https://cdn.mos.cms.futurecdn.net/CxLvbQNp2Y4BQkTkpW5m7b-1920-80.jpg.webp' width = 750>\n...\n\n### ♟️ **Strategy**\nHello, In this Notebook I will follow a detail explantion on how to implement a GDBT model, in this case using XGBoost, because of the speed by using GPUs.\n* Start simple building a baseline model not features at the start. \n* Focus on strategies to maximize the utilization of the information available as all we know there is quite a lot of data.\n* Move to feature engineering.\n* Add a cross validation loop.\n* Incorporate optimization of hyper-parameters.\n* Other ideas...\n\nA simple description or essential steps that we need to consider during the developement of the **Notebook**\n* We need to predict, **team_A_scoring_within_10sec** & **team_B_scoring_within_10sec**\n* We will build two models one from the perspective of **team a** and another from the perspective of **team b**\n* We need to be aware of the impacts on **team a** by **team b** and the impacts on **team b** by **team a** (I will use all the variables all the time)\n* I will start with a simple desicion tree (to slow so I skip it) and then move to something more complex like **Extreme Gradient Boost Trees** or **Light Gradient Boosted Trees** (I can use GPU so it's faster, I startet here)\n* We need to predict the probability of scoring in the next 10 seconds for each team\n\n&nbsp;\n\n---\n\n### 🗄️ **Dataset Description (From Kaggle)**\n\nThe dataset consists of sequences of snapshots of the state of a Rocket League match, including position and velocity of all players and the ball, as well as extra information. The goal of the competition is to predict -- from a given snapshot in the game -- for each team, the probability that they will score within the next 10 seconds of game time.\n\nThe data was taken from professional Rocket League matches. Each event consists of a chronological series of frames recorded at 10 frames per second. All events begin with a kickoff, and most end in one team scoring a goal, but some are truncated and end with no goal scored due to circumstances which can cause gameplay strategies to shift, for example 1) nearing end of regulation (where the game continues until the ball touches the ground) or 2) becoming non-competitive, eg one team winning by 3+ goals with little time remaining.\n\n#### **Files:**\n\n* **train_[0-9].csv:** Train set split into 10 files. Rows are sorted by game_num, event_id, and event_time, and each event is entirely contained in one file.\n* **test.csv:** Test set. Unlike the train set, the rows are scrambled.\n* **[train|test]_dtypes.csv:** pandas dtypes for the columns in the train / test set, which can be pulled and passed to pd.read_csv() on the full set to read it with correct types since by default, pd.read_csv() will use 64-bit types which will waste memory. See below for example code.\n* **sample_submission.csv:** A sample submission in the correct format.\n\n#### **Columns:**\n\n* **game_num (train only):** Unique identifier for the game from which the event was taken.\n\n* **event_id (train only):** Unique identifier for the sequence of consecutive frames.\n\n* **event_time (train only):** Time in seconds before the event ended, either by a goal being scored or simply when we decided to truncate the timeseries if a goal was not scored.\n\n* **ball_pos_[xyz]:** Ball's position as a 3d vector.\n\n* **ball_vel_[xyz]:** Ball's velocity as a 3d vector.\n\n* **For i in [0, 6):**\n\n    * p{i}_pos_[xyz]: Player i's position as a 3d vector.\n    * p{i}_vel_[xyz]: Player i's velocity as a 3d vector.\n    * p{i}_boost: Player i's boost remaining, in [0, 100]. A player can consume boost to substantially increase their speed, and is required to fly up into the z dimension (besides driving up a wall, or the small air gained by a jump).\n    * All p{i} columns will be NaN if and only if the player is demolished (destroyed by an enemy player; will respawn within a few seconds).\n    * Players 0, 1, and 2 make up team A and players 3, 4, and 5 make up team B.\n    * The orientation vector of the player's car (which way the car is facing) does not necessarily match the player's velocity vector, and this dataset does not capture orientation data.\n    * boost{i}_timer: Time in seconds until big boost orb i respawns, or 0 if it's available. Big boost orbs grant a full 100 boost to a player driving over it. The orb (x, y) locations are roughly [ (-61.4, -81.9), (61.4, -81.9), (-71.7, 0), (71.7, 0), (-61.4, 81.9), (61.4, 81.9) ] with z = 0. (Players can also gain boost from small boost pads across the map, but we do not capture those pads in this dataset).\n \n\n* **player_scoring_next (train only):** Which player scores at the end of the current event, in [0, 6), or -1 if the event does not end in a goal.\n\n* **team_scoring_next (train only):** Which team scores at the end of the current event (A or B), or NaN if the event does not end in a goal.\n\n* **team_[A|B]_scoring_within_10sec (train only):** [Target columns] Value of 1 if team_scoring_next == [A|B] and time_before_event is in [-10, 0], otherwise 0.\n\n* **id (test and submission only):** Unique identifier for each test row. Your submission should be a pair of team_A_scoring_within_10sec and team_B_scoring_within_10sec probability predictions for each id, where your predictions can range the real numbers from [0, 1].","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 📚 1.0 Loading Required Libraries...\nLoading what's needed, I try to avoid copy pasting a massive list of libraries that will not be used as is the case of many **Notebooks** in the wild...","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-08T16:36:14.456029Z","iopub.execute_input":"2022-10-08T16:36:14.456444Z","iopub.status.idle":"2022-10-08T16:36:14.471824Z","shell.execute_reply.started":"2022-10-08T16:36:14.456355Z","shell.execute_reply":"2022-10-08T16:36:14.470798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Utility function and other libraries...\nimport warnings # Import the system warnings\nimport gc # Import the python garbage collector \n\n# Import sklearn modules to pre-process some of the data...\nfrom sklearn.model_selection import train_test_split # Generates train and validations datasets.\nfrom sklearn.preprocessing import StandardScaler # Standarize the dataset values to avoid training issues.\n\n\n# Machine learning model libraries...\nfrom xgboost import XGBClassifier # Import a machine learning classifier","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:14.473120Z","iopub.execute_input":"2022-10-08T16:36:14.473659Z","iopub.status.idle":"2022-10-08T16:36:15.004526Z","shell.execute_reply.started":"2022-10-08T16:36:14.473621Z","shell.execute_reply":"2022-10-08T16:36:15.003553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 🔧 2.0 Configuring the Noteboook\nI like to disable warnings, set the maximin numbers of decimals for pandas and other nice parameters to make the **Notebook** look good, and increase readibility...","metadata":{}},{"cell_type":"code","source":"%%time\n# I like to disable my Notebook warnings to reduce noice.\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:15.005838Z","iopub.execute_input":"2022-10-08T16:36:15.006212Z","iopub.status.idle":"2022-10-08T16:36:15.014404Z","shell.execute_reply.started":"2022-10-08T16:36:15.006175Z","shell.execute_reply":"2022-10-08T16:36:15.013222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Notebook Configuration.\n\n# Amount of data we want to load into the model from Pandas.\nDATA_ROWS = None\n\n# Memory and replicability\nDATA_PCT = 0.35 # Load only 50% of the dataset to avoid memory issues...\nSEED = 777 # This will be the seed utilized across the notebook\n\n# Dataframe, the amount of rows and cols to visualize.\nNROWS = 30\nNCOLS = 25\n\n# Main data location base path.\nBASE_PATH = '...'\n\n# Model development parameters\nTEST_PCT = 0.20","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:15.018244Z","iopub.execute_input":"2022-10-08T16:36:15.019109Z","iopub.status.idle":"2022-10-08T16:36:15.027194Z","shell.execute_reply.started":"2022-10-08T16:36:15.019071Z","shell.execute_reply":"2022-10-08T16:36:15.025939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Configure notebook display settings to only use 2 decimal places, tables look nicer and compressed.\npd.options.display.float_format = '{:,.3f}'.format\npd.set_option('display.max_columns', NCOLS) \npd.set_option('display.max_rows', NROWS)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:15.028882Z","iopub.execute_input":"2022-10-08T16:36:15.029321Z","iopub.status.idle":"2022-10-08T16:36:15.039106Z","shell.execute_reply.started":"2022-10-08T16:36:15.029283Z","shell.execute_reply":"2022-10-08T16:36:15.037983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 🗃️ 3.0 Loading the Train Dataset\nIn this section we load the train dataset, in this case we use a custom dataset that I created to avoid problems with the **Notebook** memory\nif you want to learn more how it was created take a look to this:\n\n\n**Notebook**\nhttps://www.kaggle.com/code/cv13j0/tps-oct22-sequential-dataset-loader\n\n\n**Final Kaggle Dataset**\nhttps://www.kaggle.com/datasets/cv13j0/tpsoct22-memory-optimized-dataset","metadata":{}},{"cell_type":"code","source":"%%time\n# Read an optimized version of the dataset.\n# >>> Final Kaggle Dataset https://www.kaggle.com/datasets/cv13j0/tpsoct22-memory-optimized-dataset\n\ntrn_data = pd.read_pickle('../input/tpsoct22-memory-optimized-dataset/merged_train_dataset.pkl')","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:15.040668Z","iopub.execute_input":"2022-10-08T16:36:15.041243Z","iopub.status.idle":"2022-10-08T16:36:29.396464Z","shell.execute_reply.started":"2022-10-08T16:36:15.041206Z","shell.execute_reply":"2022-10-08T16:36:29.395220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Remove features that will not be used...\n# Select a subset of the dataset to save some memory...\n\nskip = ['game_num', 'event_id', 'event_time', 'player_scoring_next', 'team_scoring_next']\ntrn_data.drop(columns = skip, inplace = True)\n\nz_values = [var for var in trn_data.columns if 'z' in var]\ntrn_data.drop(columns = z_values, inplace = True)\n\ntrn_data = trn_data.sample(frac = DATA_PCT, replace = True, random_state = SEED)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:29.398134Z","iopub.execute_input":"2022-10-08T16:36:29.398808Z","iopub.status.idle":"2022-10-08T16:36:44.457696Z","shell.execute_reply.started":"2022-10-08T16:36:29.398767Z","shell.execute_reply":"2022-10-08T16:36:44.456426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display some basic dataset information, the most important is the memory usage...\n\ntrn_data.info(verbose = False)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:44.459357Z","iopub.execute_input":"2022-10-08T16:36:44.459954Z","iopub.status.idle":"2022-10-08T16:36:44.474815Z","shell.execute_reply.started":"2022-10-08T16:36:44.459913Z","shell.execute_reply":"2022-10-08T16:36:44.473419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 🗃️ 4.0 Loading the Test Dataset","metadata":{}},{"cell_type":"code","source":"%%time\n# Load the Test Datset utilizing the provided dtype arrays for memory eficiency\n\ndtypes_df = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/test_dtypes.csv')\ndtypes = {k: v for (k, v) in zip(dtypes_df.column, dtypes_df.dtype)}\n\ntst_data = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/test.csv', dtype = dtypes)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:44.476701Z","iopub.execute_input":"2022-10-08T16:36:44.477473Z","iopub.status.idle":"2022-10-08T16:36:49.445559Z","shell.execute_reply.started":"2022-10-08T16:36:44.477430Z","shell.execute_reply":"2022-10-08T16:36:49.444285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display some basic dataset information, the most important is the memory usage...\n\ntst_data.info(verbose = False)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:49.447417Z","iopub.execute_input":"2022-10-08T16:36:49.447830Z","iopub.status.idle":"2022-10-08T16:36:49.460474Z","shell.execute_reply.started":"2022-10-08T16:36:49.447785Z","shell.execute_reply":"2022-10-08T16:36:49.459268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display information of the first 5 rows in the datset...\n\ntst_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:49.462403Z","iopub.execute_input":"2022-10-08T16:36:49.463097Z","iopub.status.idle":"2022-10-08T16:36:49.496562Z","shell.execute_reply.started":"2022-10-08T16:36:49.462996Z","shell.execute_reply":"2022-10-08T16:36:49.495031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 🗺️ 5.0 Exploring the Loaded Dataset...","metadata":{}},{"cell_type":"code","source":"%%time\n# Display some basic dataset information, the most important is the memory usage...\n\ntrn_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:49.498566Z","iopub.execute_input":"2022-10-08T16:36:49.499203Z","iopub.status.idle":"2022-10-08T16:36:49.510901Z","shell.execute_reply.started":"2022-10-08T16:36:49.499165Z","shell.execute_reply":"2022-10-08T16:36:49.509618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display information of the first 5 rows in the datset...\n\ntrn_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:49.516907Z","iopub.execute_input":"2022-10-08T16:36:49.517369Z","iopub.status.idle":"2022-10-08T16:36:49.539979Z","shell.execute_reply.started":"2022-10-08T16:36:49.517343Z","shell.execute_reply":"2022-10-08T16:36:49.539066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display advanced statistics of the dataset...\n#trn_data.describe()  ","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:49.541412Z","iopub.execute_input":"2022-10-08T16:36:49.542112Z","iopub.status.idle":"2022-10-08T16:36:49.549480Z","shell.execute_reply.started":"2022-10-08T16:36:49.542071Z","shell.execute_reply":"2022-10-08T16:36:49.548434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Display the number of NaNs in each column of the dataset...\n\ntrn_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:49.550917Z","iopub.execute_input":"2022-10-08T16:36:49.552089Z","iopub.status.idle":"2022-10-08T16:36:50.577801Z","shell.execute_reply.started":"2022-10-08T16:36:49.552023Z","shell.execute_reply":"2022-10-08T16:36:50.576765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Selected specifically the columns that are missing values...\n# Display the amount of NaNs in each of this columns...\n\ntrn_data[trn_data.columns[trn_data.isnull().any()]].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:50.579298Z","iopub.execute_input":"2022-10-08T16:36:50.579927Z","iopub.status.idle":"2022-10-08T16:36:52.744294Z","shell.execute_reply.started":"2022-10-08T16:36:50.579887Z","shell.execute_reply":"2022-10-08T16:36:52.743204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 💡 6. Feature Engineering\nIn this sections I will create multiple features from what I belive can add significan signal to the model\nI will keep a list of ideas and as they are implemented I will add comments of their impacts...\n\n**Ideas for Features...**\n* Magnitude of velocity\n* Each player distances to the ball\n* Number of players alive during the game section (10s before scoring)\n* Total team average distance from the ball\n* Total team average velocity\n* Ball distance from the scoring location","metadata":{}},{"cell_type":"code","source":"%%time\n# Calculates something...\n\ndef total_team_boost(df):\n    '''\n    Calculates the total boost available in the team...\n    '''\n    df['team_a_boost'] = df['p0_boost'] + df['p1_boost']+ df['p2_boost']\n    df['team_b_boost'] = df['p3_boost'] + df['p4_boost']+ df['p5_boost']\n    #df['team_a_advantage'] = df['team_a_boost'] - df['team_b_boost']\n    \n    df['boost_timer_a'] = df['boost0_timer'] + df['boost1_timer']+ df['boost2_timer']\n    df['boost_timer_b'] = df['boost3_timer'] + df['boost4_timer']+ df['boost5_timer']\n    \n    return df\n\ntrn_data = total_team_boost(trn_data)\ntst_data = total_team_boost(tst_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:52.745851Z","iopub.execute_input":"2022-10-08T16:36:52.746493Z","iopub.status.idle":"2022-10-08T16:36:53.817051Z","shell.execute_reply.started":"2022-10-08T16:36:52.746454Z","shell.execute_reply":"2022-10-08T16:36:53.815868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Calculates something...\n\ndef calculates_magnitudes(df):\n    '''\n    \n    '''\n    for i in range(0,6):\n        df[f'p{i}_pos'] = np.sqrt(df[f'p{i}_pos_x'] ** 2 + df[f'p{i}_pos_y'] ** 2 + df[f'p{i}_pos_z'] ** 2)\n        df[f'p{i}_vel'] = np.sqrt(df[f'p{i}_vel_x'] ** 2 + df[f'p{i}_vel_y'] ** 2 + df[f'p{i}_vel_z'] ** 2)\n    \n    df['ball_vel'] = np.sqrt(df['ball_vel_x'] ** 2 + df['ball_vel_y'] ** 2 + df['ball_vel_z'] ** 2)\n    df['ball_pos'] = np.sqrt(df['ball_pos_x'] ** 2 + df['ball_pos_y'] ** 2 + df['ball_pos_z'] ** 2)\n        \n    return df\n\n#trn_data = calculates_magnitudes(trn_data)\n#tst_data = calculates_magnitudes(tst_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:53.818739Z","iopub.execute_input":"2022-10-08T16:36:53.819140Z","iopub.status.idle":"2022-10-08T16:36:53.826844Z","shell.execute_reply.started":"2022-10-08T16:36:53.819101Z","shell.execute_reply":"2022-10-08T16:36:53.825749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Build features around NaNs\ndef missing_values_feat(df):\n    '''\n    \n    '''\n    team_a = ['p0_pos_x','p1_pos_x','p2_pos_x']\n    team_b = ['p3_pos_x','p4_pos_x','p5_pos_x']\n    \n    df['missing_team_A_players'] = df[team_a].isnull().sum(axis = 1)\n    df['missing_team_B_players'] = df[team_b].isnull().sum(axis = 1)\n\n    return df\n\ntrn_data = missing_values_feat(trn_data)\ntst_data = missing_values_feat(tst_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:53.828363Z","iopub.execute_input":"2022-10-08T16:36:53.829582Z","iopub.status.idle":"2022-10-08T16:36:54.561928Z","shell.execute_reply.started":"2022-10-08T16:36:53.829543Z","shell.execute_reply":"2022-10-08T16:36:54.560480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trn_data.sample(10)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:54.563493Z","iopub.execute_input":"2022-10-08T16:36:54.564115Z","iopub.status.idle":"2022-10-08T16:36:54.954026Z","shell.execute_reply.started":"2022-10-08T16:36:54.564073Z","shell.execute_reply":"2022-10-08T16:36:54.952997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🍳 7. Preparing the Model Data","metadata":{}},{"cell_type":"code","source":"%%time\n# ...\n\n# List of featuare that needs to be avoided...\nskip = ['game_num', 'event_id', 'event_time', 'player_scoring_next', 'team_scoring_next', 'team_A_scoring_within_10sec','team_B_scoring_within_10sec']\n\n# Using a list comprehension we generate a list of features to train the model...\nfeatures = [feat for feat in trn_data.columns if feat not in skip]\nlabel_a = 'team_A_scoring_within_10sec'\nlabel_b = 'team_B_scoring_within_10sec'","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:54.955372Z","iopub.execute_input":"2022-10-08T16:36:54.956508Z","iopub.status.idle":"2022-10-08T16:36:54.963676Z","shell.execute_reply.started":"2022-10-08T16:36:54.956468Z","shell.execute_reply":"2022-10-08T16:36:54.962669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# The type of ML model I'm currenthly using can deal with NaNs\n\ntrn_data[features] = trn_data[features].fillna(value = 0)\ntst_data[features] = tst_data[features].fillna(value = 0)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:54.965152Z","iopub.execute_input":"2022-10-08T16:36:54.965776Z","iopub.status.idle":"2022-10-08T16:36:58.009467Z","shell.execute_reply.started":"2022-10-08T16:36:54.965737Z","shell.execute_reply":"2022-10-08T16:36:58.008344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Scale the train and test datasets to improve the learning capabilities of the model.\n# During the training stage.\n\nscaler = StandardScaler()\ntrn_data[features] = scaler.fit_transform(trn_data[features])\ntst_data[features] = scaler.transform(tst_data[features])","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:36:58.010896Z","iopub.execute_input":"2022-10-08T16:36:58.011327Z","iopub.status.idle":"2022-10-08T16:37:13.514094Z","shell.execute_reply.started":"2022-10-08T16:36:58.011286Z","shell.execute_reply":"2022-10-08T16:37:13.512960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Converts the train dataset into a train and validations datasets.\n# We will use only 20% of the data for validation.\n\nX_train, X_val, y_train_a, y_val_a = train_test_split(trn_data[features], trn_data[label_a], test_size = TEST_PCT, random_state = SEED)\nX_train, X_val, y_train_b, y_val_b = train_test_split(trn_data[features], trn_data[label_b], test_size = TEST_PCT, random_state = SEED)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:37:13.515831Z","iopub.execute_input":"2022-10-08T16:37:13.516262Z","iopub.status.idle":"2022-10-08T16:37:31.629802Z","shell.execute_reply.started":"2022-10-08T16:37:13.516221Z","shell.execute_reply":"2022-10-08T16:37:31.627993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# We need to release some memory to avoid to run into Kernel memory issues...\n# We will use the garbage collector...\n\ndel trn_data\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:37:31.631401Z","iopub.execute_input":"2022-10-08T16:37:31.631780Z","iopub.status.idle":"2022-10-08T16:37:31.783808Z","shell.execute_reply.started":"2022-10-08T16:37:31.631726Z","shell.execute_reply":"2022-10-08T16:37:31.782406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 👾 8. A Baseline Model Using XGBoost...","metadata":{}},{"cell_type":"code","source":"%%time\n# In this cell we define simple model parameters...\n# Instanciate a classifier using XGBoost...\n\n\n# Definition of the XGboost model parameters...\nxgb_params = {'n_estimators'    : 256,\n              'max_depth'       : 16,\n              'learning_rate'   : 0.1,\n              'subsample'       : 0.95,\n              'colsample_bytree': 0.90,\n              'reg_lambda'      : 1.50,\n              'reg_alpha'       : 6.10,\n              'gamma'           : 1.40,\n              'random_state'    : SEED,\n              'objective'       : 'binary:logistic',\n              'tree_method'     : 'gpu_hist',\n             }\n\n# Instanciate a classifier...\nclf = XGBClassifier(**xgb_params)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:37:31.786343Z","iopub.execute_input":"2022-10-08T16:37:31.787142Z","iopub.status.idle":"2022-10-08T16:37:31.801788Z","shell.execute_reply.started":"2022-10-08T16:37:31.787086Z","shell.execute_reply":"2022-10-08T16:37:31.798766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8.1. Train a Model for Team A...","metadata":{}},{"cell_type":"code","source":"%%time \n# Train a ML model from the perspective of Team A\n\nclf.fit(X_train, y_train_a, eval_set = [(X_val, y_val_a)], eval_metric = ['logloss'], early_stopping_rounds = 128, verbose = 32)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:37:31.805885Z","iopub.execute_input":"2022-10-08T16:37:31.807966Z","iopub.status.idle":"2022-10-08T16:40:25.824805Z","shell.execute_reply.started":"2022-10-08T16:37:31.807869Z","shell.execute_reply":"2022-10-08T16:40:25.823578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Generate model predictions utilizing the training ML model for team A\n\npredictions_a = clf.predict_proba(tst_data[features])[:,1]","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:40:25.827985Z","iopub.execute_input":"2022-10-08T16:40:25.828310Z","iopub.status.idle":"2022-10-08T16:40:45.144525Z","shell.execute_reply.started":"2022-10-08T16:40:25.828280Z","shell.execute_reply":"2022-10-08T16:40:45.143561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Delete some of the dataset not needed anymore to save some memory\n\ndel y_val_a, y_train_a\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:40:45.148269Z","iopub.execute_input":"2022-10-08T16:40:45.149655Z","iopub.status.idle":"2022-10-08T16:40:45.297540Z","shell.execute_reply.started":"2022-10-08T16:40:45.149614Z","shell.execute_reply":"2022-10-08T16:40:45.296283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8.2. Train a Model for Team B...","metadata":{}},{"cell_type":"code","source":"%%time \n# Train a ML model from the perspective of Team B\n\nclf.fit(X_train, y_train_b, eval_set = [(X_val, y_val_b)], eval_metric = ['logloss'], early_stopping_rounds = 128, verbose = 32)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:40:45.299103Z","iopub.execute_input":"2022-10-08T16:40:45.299991Z","iopub.status.idle":"2022-10-08T16:43:38.859793Z","shell.execute_reply.started":"2022-10-08T16:40:45.299949Z","shell.execute_reply":"2022-10-08T16:43:38.858614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Generate model predictions utilizing the training ML model for team B\n\npredictions_b = clf.predict_proba(tst_data[features])[:,1]","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:38.861168Z","iopub.execute_input":"2022-10-08T16:43:38.862226Z","iopub.status.idle":"2022-10-08T16:43:57.948696Z","shell.execute_reply.started":"2022-10-08T16:43:38.862184Z","shell.execute_reply":"2022-10-08T16:43:57.947697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Delete some of the dataset not needed anymore to save some memory\n\ndel y_val_b, y_train_b\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:57.953142Z","iopub.execute_input":"2022-10-08T16:43:57.955965Z","iopub.status.idle":"2022-10-08T16:43:58.148633Z","shell.execute_reply.started":"2022-10-08T16:43:57.955905Z","shell.execute_reply":"2022-10-08T16:43:58.147517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Capture some of the model scores for guidance, I should use Neptune.\n# Probably I will implemented in the future...\n\n# Experiments...\n# LogLoss: 0.12593 ... (first plain runs...) LR = 0.05\n# LogLoss: 0.10568 ... (first plain runs...) LR = 0.1\n# LogLoss: 0.09632 ... (first plain runs...) LR = 0.1, Estimators = 512","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:58.152758Z","iopub.execute_input":"2022-10-08T16:43:58.155017Z","iopub.status.idle":"2022-10-08T16:43:58.161264Z","shell.execute_reply.started":"2022-10-08T16:43:58.154964Z","shell.execute_reply":"2022-10-08T16:43:58.160237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8.3 Feature Importance","metadata":{}},{"cell_type":"code","source":"%%time\n# ...\ndef plot_feature_importance(importance, names, model_type, max_features = 10):\n    '''\n    \n    '''\n    \n    # Create arrays from feature importance and feature names\n    feature_importance = np.array(importance)\n    feature_names = np.array(names)\n\n    # Create a DataFrame using a Dictionary\n    data={'feature_names':feature_names,'feature_importance':feature_importance}\n    fi_df = pd.DataFrame(data)\n\n    # Sort the DataFrame in order decreasing feature importance\n    fi_df.sort_values(by=['feature_importance'], ascending=False,inplace=True)\n    fi_df = fi_df.head(max_features)\n\n    # Define size of bar plot\n    plt.figure(figsize=(8,15))\n    \n    # Plot Searborn bar chart\n    sns.barplot(x=fi_df['feature_importance'], y=fi_df['feature_names'])\n    \n    # Add chart labels\n    plt.title(model_type + 'FEATURE IMPORTANCE')\n    plt.xlabel('FEATURE IMPORTANCE')\n    plt.ylabel('FEATURE NAMES')\n    \n    return None","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:58.165795Z","iopub.execute_input":"2022-10-08T16:43:58.166985Z","iopub.status.idle":"2022-10-08T16:43:58.180653Z","shell.execute_reply.started":"2022-10-08T16:43:58.166950Z","shell.execute_reply":"2022-10-08T16:43:58.179570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = {'feature_names':X_train.columns,'feature_importance':clf.feature_importances_}\nfi_df = pd.DataFrame(data)\nfi_df.sort_values(by=['feature_importance'], ascending = False, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:58.183364Z","iopub.execute_input":"2022-10-08T16:43:58.185875Z","iopub.status.idle":"2022-10-08T16:43:58.242625Z","shell.execute_reply.started":"2022-10-08T16:43:58.185839Z","shell.execute_reply":"2022-10-08T16:43:58.241652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fi_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:58.247596Z","iopub.execute_input":"2022-10-08T16:43:58.250079Z","iopub.status.idle":"2022-10-08T16:43:58.260817Z","shell.execute_reply.started":"2022-10-08T16:43:58.250041Z","shell.execute_reply":"2022-10-08T16:43:58.259637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fi_df.head(30)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:58.265982Z","iopub.execute_input":"2022-10-08T16:43:58.268504Z","iopub.status.idle":"2022-10-08T16:43:58.288964Z","shell.execute_reply.started":"2022-10-08T16:43:58.268465Z","shell.execute_reply":"2022-10-08T16:43:58.286333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fi_df.tail(30)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:58.290445Z","iopub.execute_input":"2022-10-08T16:43:58.290781Z","iopub.status.idle":"2022-10-08T16:43:58.303369Z","shell.execute_reply.started":"2022-10-08T16:43:58.290746Z","shell.execute_reply":"2022-10-08T16:43:58.302064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# ...\nplot_feature_importance(clf.feature_importances_, X_train.columns,'XG BOOST ', max_features = 45)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:58.305319Z","iopub.execute_input":"2022-10-08T16:43:58.306110Z","iopub.status.idle":"2022-10-08T16:43:59.298040Z","shell.execute_reply.started":"2022-10-08T16:43:58.306074Z","shell.execute_reply":"2022-10-08T16:43:59.296965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 📁 9. Generate Model Submission File...\nIn this section we load the **Submission** dataset and populate it with the predictions of the models for **Teams A** and **Teams B**","metadata":{}},{"cell_type":"code","source":"%%time\n# Load the submission file into a pandas dataframe...\n\nsubmission_data = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/sample_submission.csv', dtype = dtypes)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:59.304699Z","iopub.execute_input":"2022-10-08T16:43:59.305668Z","iopub.status.idle":"2022-10-08T16:43:59.477328Z","shell.execute_reply.started":"2022-10-08T16:43:59.305628Z","shell.execute_reply":"2022-10-08T16:43:59.476318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Review the first 5 rows of the submission file, gives general idea what to populate...\n\nsubmission_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:59.479791Z","iopub.execute_input":"2022-10-08T16:43:59.480683Z","iopub.status.idle":"2022-10-08T16:43:59.493912Z","shell.execute_reply.started":"2022-10-08T16:43:59.480641Z","shell.execute_reply":"2022-10-08T16:43:59.492696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Write back the model predictions to the submission datastet...\n# This file will be used in Kaggle...\n\nsubmission_data['team_A_scoring_within_10sec'] = predictions_a\nsubmission_data['team_B_scoring_within_10sec'] = predictions_b","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:59.495495Z","iopub.execute_input":"2022-10-08T16:43:59.496133Z","iopub.status.idle":"2022-10-08T16:43:59.508919Z","shell.execute_reply.started":"2022-10-08T16:43:59.496093Z","shell.execute_reply":"2022-10-08T16:43:59.507812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Export the submission dataset to a CSV...\n\nsubmission_data.to_csv('submission_xgboost_10012022.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T16:43:59.510315Z","iopub.execute_input":"2022-10-08T16:43:59.510770Z","iopub.status.idle":"2022-10-08T16:44:01.135513Z","shell.execute_reply.started":"2022-10-08T16:43:59.510731Z","shell.execute_reply":"2022-10-08T16:44:01.134165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}}]}