{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://www.esports.net/wp-content/uploads/2021/05/best-goals-in-rocket-league.jpg)\n# The dataset structure","metadata":{}},{"cell_type":"markdown","source":"The first step of building any machine learning model is understanding the data. I won't go into any details of variable description or make any visualization, since you'll find plenty of them in the notebook section.\nInstead, let's focus on the difference between train and test set.","metadata":{}},{"cell_type":"markdown","source":"# Difference between train and test data\nYou'll immediately notice that train and test sets have different columns.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ntrain_cols = pd.read_csv('../input/tabular-playground-series-oct-2022/train_0.csv', nrows = 0).columns\ntest_cols = pd.read_csv('../input/tabular-playground-series-oct-2022/test.csv', nrows = 0).columns\n\ncols = [i for i in train_cols if i not in test_cols]\ncols","metadata":{"execution":{"iopub.status.busy":"2022-10-23T10:10:03.182376Z","iopub.execute_input":"2022-10-23T10:10:03.183434Z","iopub.status.idle":"2022-10-23T10:10:03.241876Z","shell.execute_reply.started":"2022-10-23T10:10:03.183393Z","shell.execute_reply":"2022-10-23T10:10:03.240628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the dataset description page:  \n* `game_num`: Unique identifier for the game from which the event was taken.\n\n* `event_id`: Unique identifier for the sequence of consecutive frames.\n\n* `event_time`: Time in seconds before the event ended, either by a goal being scored or simply when we decided to truncate the timeseries if a goal was not scored.\n\n* `player_scoring_next` (train only): Which player scores at the end of the current event.\n\n* `team_scoring_next (train only)`: Which team scores at the end of the current event (A or B), or NaN if the event does not end in a goal.\n\n* `team_[A|B]_scoring_within_10sec` (train only): (Target columns) Value of 1 if `team_scoring_next == [A|B]` and time_before_event is in (-10, 0), otherwise 0.\n\n","metadata":{}},{"cell_type":"markdown","source":"Every row is a snapshot of the game in a given moment in time. The samples are taken every 0.3 seconds.","metadata":{}},{"cell_type":"markdown","source":"# Comparing distributions\nLet's compare now the distributions of train and test data.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('../input/tabular-playground-series-oct-2022/train_0.csv',usecols = ['ball_pos_x'])\ntest = pd.read_csv('../input/tabular-playground-series-oct-2022/test.csv',usecols = ['ball_pos_x'])","metadata":{"execution":{"iopub.status.busy":"2022-10-23T09:35:39.768615Z","iopub.execute_input":"2022-10-23T09:35:39.769033Z","iopub.status.idle":"2022-10-23T09:35:47.340300Z","shell.execute_reply.started":"2022-10-23T09:35:39.769000Z","shell.execute_reply":"2022-10-23T09:35:47.339242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.hist(bins = 40)","metadata":{"execution":{"iopub.status.busy":"2022-10-23T09:35:47.341891Z","iopub.execute_input":"2022-10-23T09:35:47.342266Z","iopub.status.idle":"2022-10-23T09:35:47.689853Z","shell.execute_reply.started":"2022-10-23T09:35:47.342229Z","shell.execute_reply":"2022-10-23T09:35:47.688796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.hist(bins = 40)","metadata":{"execution":{"iopub.status.busy":"2022-10-23T09:35:47.692942Z","iopub.execute_input":"2022-10-23T09:35:47.693554Z","iopub.status.idle":"2022-10-23T09:35:47.989518Z","shell.execute_reply.started":"2022-10-23T09:35:47.693514Z","shell.execute_reply":"2022-10-23T09:35:47.988589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Two-sample Kolmogorov-Smirnov (KS) test  \nWe take two samples of n = 10000, one from train and one from test. The null hypotesis of this test states that both samples come from the same distribution.","metadata":{}},{"cell_type":"code","source":"train_sample = train.sample(n = 10000)\ntest_sample = test.sample(n = 10000)\n\nfrom scipy.stats import kstest,norm\nkstest(train_sample.to_numpy().ravel(),test_sample.to_numpy().ravel())","metadata":{"execution":{"iopub.status.busy":"2022-10-23T09:35:47.990998Z","iopub.execute_input":"2022-10-23T09:35:47.991620Z","iopub.status.idle":"2022-10-23T09:35:48.082812Z","shell.execute_reply.started":"2022-10-23T09:35:47.991584Z","shell.execute_reply":"2022-10-23T09:35:48.081751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since the pvalue > 0.05, we don't have enough evidence to reject the null hypotesis.\nThus, the two samples come from from the same distribution.  \nThis means that the test set is probably collected in the same way: snapshots of the game in random instants.\nHowever, in the test test there are no additional information on the game. It's reasonable to think that the test rows are collected from different games.\n","metadata":{}},{"cell_type":"markdown","source":"# The problem: target leakage","metadata":{}},{"cell_type":"markdown","source":"What does this mean? That training on train set and predicting on the given test set is fine.\nBut what happens if you want to add cross validation to the train set? If you split the folds with a standard KFold CV, the groups will be split randomly.\nRows from the same game will end up in both train and test set, causing target leakage, that in the end would make your model overfit. Basically, you are trying to predict data that the model has already seen: past snapshots of the same game.\n\n\nWe need to train and test on **different games**. The following code has been posted from @sergiosaharovskiy in the [discussion section](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/360070). This issue has been pointed out in the following discussion: [Validation - how not to OVERFIT here!](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714)\n\nIn the end, I'll add more making it parallelizable easy to use with a model.  ","metadata":{}},{"cell_type":"code","source":"def create_kf_groups(game_num, n_folds=5):\n    import string\n    \"\"\"\n    Creates groups for the GroupKFold based on the game_num.\n    We don't want to let train data spill over the validation set.\n    :param game_num: pandas.core.series.Series\n    :param n_folds: int (default 10)\n    :return: pandas.core.series.Series (alpha list)\n    \"\"\"\n    count_unique_games = game_num.unique().shape[0]\n    group_size = count_unique_games // n_folds\n    bins = [group_size * i for i in range(n_folds)]\n    bins += [count_unique_games]\n    labels = [letter for letter in string.ascii_letters[:n_folds]]\n    kf_groups = pd.cut(game_num, bins=bins, labels=labels)\n    return kf_groups","metadata":{"execution":{"iopub.status.busy":"2022-10-23T09:35:48.084324Z","iopub.execute_input":"2022-10-23T09:35:48.084791Z","iopub.status.idle":"2022-10-23T09:35:48.091782Z","shell.execute_reply.started":"2022-10-23T09:35:48.084752Z","shell.execute_reply":"2022-10-23T09:35:48.090362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = pd.read_csv('../input/tabular-playground-series-oct-2022/train_0.csv', usecols = ['game_num', 'event_id', 'event_time','ball_pos_x'])\ny = pd.read_csv('../input/tabular-playground-series-oct-2022/train_0.csv', usecols = ['team_A_scoring_within_10sec'])\nX['group'] = create_kf_groups(X.game_num, n_folds=5)\n\nfrom sklearn.model_selection import GroupKFold\ngroups = create_kf_groups(X.game_num, n_folds=5)\ngkf = GroupKFold(n_splits=5)\nfor fold, (idx_train, idx_val) in enumerate(gkf.split(train, y, groups=groups)):\n    X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n    X_val, y_val = X.iloc[idx_val], y.iloc[idx_val]\n    print('============================')\n    print(f'Fold {fold}')\n    print('============================')\n    print(f'       TRAIN SET')\n    print(X_train['group'].value_counts())\n    print(f'     VALIDATION SET')\n    print(X_val['group'].value_counts())\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-10-23T09:36:27.481636Z","iopub.execute_input":"2022-10-23T09:36:27.482022Z","iopub.status.idle":"2022-10-23T09:36:29.081638Z","shell.execute_reply.started":"2022-10-23T09:36:27.481990Z","shell.execute_reply":"2022-10-23T09:36:29.080562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Implementation of GroupKFold","metadata":{}},{"cell_type":"markdown","source":"As you can see, now train and validation sets are taken from different games, divided in groups. Below is my implementation that makes it equivalent to `cross_val_score` from sklearn.  \nYou can set the number of cores used with the n_jobs parameter.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GroupKFold\nfrom joblib import Parallel, delayed\nfrom sklearn.metrics import log_loss\ndef cross_val_score_gkf(model, X,y, n_jobs=-1, cv = 5):\n    import string\n    count_unique_games = X.game_num.unique().shape[0]\n    group_size = count_unique_games // cv\n    bins = [group_size * i for i in range(cv)]\n    bins += [count_unique_games]\n    labels = [letter for letter in string.ascii_letters[:cv]]\n    groups = pd.cut(X.game_num, bins=bins, labels=labels)\n    from joblib import Parallel, delayed\n    n_folds = 5\n    scores = []\n    X = X.drop('game_num', axis=1)\n    gkf = GroupKFold(n_splits=cv)\n    \n    def train_folds(X, model, idx_train, idx_val):\n        X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n        X_train = X_train.to_numpy()\n        y_train = y_train.to_numpy()\n        model.fit(X_train, y_train)\n        X_val, y_val = X.iloc[idx_val], y.iloc[idx_val]\n        y_pred = model.predict_proba(X_val)\n        score = log_loss(y_val, y_pred)\n        scores.append(score)\n        return scores\n    \n    out = Parallel(n_jobs=n_jobs, verbose=0)\\\n    (delayed(train_folds)(X, model, idx_train, idx_val) for idx_train, idx_val in gkf.split(X, y, groups=groups))\n    scores = [item for sublist in out for item in sublist]\n    return scores","metadata":{"execution":{"iopub.status.busy":"2022-10-23T10:10:30.124298Z","iopub.execute_input":"2022-10-23T10:10:30.124713Z","iopub.status.idle":"2022-10-23T10:10:30.136316Z","shell.execute_reply.started":"2022-10-23T10:10:30.124676Z","shell.execute_reply":"2022-10-23T10:10:30.135300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model training with GroupKFold","metadata":{}},{"cell_type":"code","source":"from lightgbm import LGBMClassifier\nfrom sklearn.model_selection import cross_val_score\nmodel = LGBMClassifier(device = 'gpu')\nX = pd.read_csv('../input/tabular-playground-series-oct-2022/train_0.csv')\nX.drop(cols, axis =1, inplace = True)\nX['game_num'] = pd.read_csv('../input/tabular-playground-series-oct-2022/train_0.csv', usecols = ['game_num'])\ny = pd.read_csv('../input/tabular-playground-series-oct-2022/train_0.csv', usecols = ['team_A_scoring_within_10sec'])\n\nscores = cross_val_score_gkf(model, X,y)\n","metadata":{},"execution_count":null,"outputs":[]}]}