{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Tabular Playground Oct 2022 - Baseline","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we present a baseline model via lightgbm. While there are a few notebooks with similar approaches (e.g., [this one](https://www.kaggle.com/code/tracyporter/oct-22-tabular-simple-lgbm/notebook?scriptVersionId=106985125), [this one](https://www.kaggle.com/code/chal1ce/lgbm-baseline-with-half-of-data-and-no-fold) and [this one](https://www.kaggle.com/code/infrarosso/tps-oct-2022-eda-lgbm-model-submit/notebook?scriptVersionId=106987513)), the contributions of this notebook are the following:\n\n1. Instead of merging all the training datasets prior to fitting, we train a lightgbm model by batch. This significantly reduces the burdens on the memory requirement. Our results show this approach produces results with quality similar to or better than the similar models trained with full set of data at once.\n\n2. We also analyze the variable importance and come up with some interesting observations.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport gc\nfrom itertools import chain\nimport lightgbm\nfrom sklearn.metrics import log_loss","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:45:59.442697Z","iopub.execute_input":"2022-10-01T22:45:59.443166Z","iopub.status.idle":"2022-10-01T22:46:01.824611Z","shell.execute_reply.started":"2022-10-01T22:45:59.443067Z","shell.execute_reply":"2022-10-01T22:46:01.823390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DAT_DIR = '../input/tabular-playground-series-oct-2022'","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:46:03.072883Z","iopub.execute_input":"2022-10-01T22:46:03.074956Z","iopub.status.idle":"2022-10-01T22:46:03.082408Z","shell.execute_reply.started":"2022-10-01T22:46:03.074838Z","shell.execute_reply":"2022-10-01T22:46:03.080454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(os.path.join(DAT_DIR, \"train_0.csv\"))","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:46:06.144747Z","iopub.execute_input":"2022-10-01T22:46:06.145158Z","iopub.status.idle":"2022-10-01T22:46:38.770826Z","shell.execute_reply.started":"2022-10-01T22:46:06.145124Z","shell.execute_reply":"2022-10-01T22:46:38.769680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we will train the model by batch, we will divide the training files into categories of training, validation, and (out-of-sample) testing. The last category is to assess the model performance free of biases as they have never been seen by the model before. ","metadata":{}},{"cell_type":"code","source":"train_files = list(range(7))\nvalid_files = list(range(7,8))\ntest_files = list(range(8, 10))","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:46:38.773254Z","iopub.execute_input":"2022-10-01T22:46:38.773769Z","iopub.status.idle":"2022-10-01T22:46:38.780519Z","shell.execute_reply.started":"2022-10-01T22:46:38.773723Z","shell.execute_reply":"2022-10-01T22:46:38.779234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'train_files = {train_files}')\nprint(f'valid_files = {valid_files}')\nprint(f'test_files = {test_files}')","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:46:38.781781Z","iopub.execute_input":"2022-10-01T22:46:38.782073Z","iopub.status.idle":"2022-10-01T22:46:38.795446Z","shell.execute_reply.started":"2022-10-01T22:46:38.782047Z","shell.execute_reply":"2022-10-01T22:46:38.794194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid = pd.read_csv(os.path.join(DAT_DIR, f\"train_{valid_files[0]}.csv\"))","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:46:38.798128Z","iopub.execute_input":"2022-10-01T22:46:38.798872Z","iopub.status.idle":"2022-10-01T22:47:15.789113Z","shell.execute_reply.started":"2022-10-01T22:46:38.798836Z","shell.execute_reply":"2022-10-01T22:47:15.786954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:15.792015Z","iopub.execute_input":"2022-10-01T22:47:15.792640Z","iopub.status.idle":"2022-10-01T22:47:16.214757Z","shell.execute_reply.started":"2022-10-01T22:47:15.792571Z","shell.execute_reply":"2022-10-01T22:47:16.213829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.columns","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:16.215960Z","iopub.execute_input":"2022-10-01T22:47:16.216510Z","iopub.status.idle":"2022-10-01T22:47:16.224588Z","shell.execute_reply.started":"2022-10-01T22:47:16.216475Z","shell.execute_reply":"2022-10-01T22:47:16.223262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ball_pos_cols = [f'ball_pos_{i}' for i in ['x', 'y', 'z']]\nball_vel_cols = [f'ball_vel_{i}' for i in ['x', 'y', 'z']]\nplayer_pos_cols = list(chain.from_iterable([[f'p{i}_pos_{j}' for j in ['x', 'y', 'z']] for i in range(6)]))\nplayer_vel_cols = list(chain.from_iterable([[f'p{i}_vel_{j}' for j in ['x', 'y', 'z']] for i in range(6)]))\nboost_cols = [f'p{i}_boost' for i in range(6)]\nboost_timer_cols = [f'boost{i}_timer' for i in range(6)]","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:16.237155Z","iopub.execute_input":"2022-10-01T22:47:16.237710Z","iopub.status.idle":"2022-10-01T22:47:16.250960Z","shell.execute_reply.started":"2022-10-01T22:47:16.237664Z","shell.execute_reply":"2022-10-01T22:47:16.249615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'ball_pos_cols = {ball_pos_cols}')\nprint(f'ball_vel_cols = {ball_vel_cols}')\nprint(f'player_pos_cols = {player_pos_cols}')\nprint(f'boost_cols = {boost_cols}')\nprint(f'boost_timer_cols = {boost_timer_cols}')","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:16.252599Z","iopub.execute_input":"2022-10-01T22:47:16.253344Z","iopub.status.idle":"2022-10-01T22:47:16.265053Z","shell.execute_reply.started":"2022-10-01T22:47:16.253273Z","shell.execute_reply":"2022-10-01T22:47:16.263908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_cols = ball_pos_cols + ball_vel_cols + player_pos_cols + player_vel_cols + boost_cols + boost_timer_cols\nprint(f'x_cols = {x_cols}')","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:16.268754Z","iopub.execute_input":"2022-10-01T22:47:16.269112Z","iopub.status.idle":"2022-10-01T22:47:16.277312Z","shell.execute_reply.started":"2022-10-01T22:47:16.269082Z","shell.execute_reply":"2022-10-01T22:47:16.275947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_cols = ['team_A_scoring_within_10sec', 'team_B_scoring_within_10sec']\nprint(f'y_cols = {y_cols}')","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:16.278558Z","iopub.execute_input":"2022-10-01T22:47:16.278880Z","iopub.status.idle":"2022-10-01T22:47:16.289979Z","shell.execute_reply.started":"2022-10-01T22:47:16.278851Z","shell.execute_reply":"2022-10-01T22:47:16.288644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fitting and Predicting Team A","metadata":{}},{"cell_type":"markdown","source":"We will use the first training file to train the model from scratch. Subsequent files are fed to the fitted model from the previous iteration, i.e., boosting. We keep all the hyperparameter constant from each iterations.  We set max iteration high and rely on early termination to terminate boosting.","metadata":{}},{"cell_type":"code","source":"MAX_ITER = 1000\nPATIENCE = 100\nDISPLAY_FREQ = 100","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:51.594297Z","iopub.execute_input":"2022-10-01T22:47:51.595719Z","iopub.status.idle":"2022-10-01T22:47:51.601195Z","shell.execute_reply.started":"2022-10-01T22:47:51.595662Z","shell.execute_reply":"2022-10-01T22:47:51.600158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"MODEL_PARAMS = {'random_state': 1234,    \n                'learning_rate': 0.1,                \n                'n_estimators': MAX_ITER}\n\nmdl_A = lightgbm.LGBMClassifier(**MODEL_PARAMS)\n\nmdl_A.fit(X=train[x_cols], y=train[y_cols[0]],\n          eval_set=[(valid[x_cols], valid[y_cols[0]])],\n          eval_names=['valid'],\n          callbacks=[lightgbm.log_evaluation(period=DISPLAY_FREQ, show_stdv=True),\n                     lightgbm.early_stopping(stopping_rounds=PATIENCE)])","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:47:54.576406Z","iopub.execute_input":"2022-10-01T22:47:54.576903Z","iopub.status.idle":"2022-10-01T22:49:30.112193Z","shell.execute_reply.started":"2022-10-01T22:47:54.576860Z","shell.execute_reply":"2022-10-01T22:49:30.110964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in train_files[1:]:\n    print(f'Iteration {i}, Reading train_{i}.csv')\n    train_new = pd.read_csv(os.path.join(DAT_DIR, f'train_{i}.csv'))\n    print('Fitting model')\n    mdl_A.fit(X=train_new[x_cols], \n              y=train_new[y_cols[0]],\n              eval_set=[(valid[x_cols], valid[y_cols[0]])],\n              eval_names=['valid'],\n              callbacks=[lightgbm.log_evaluation(period=DISPLAY_FREQ, show_stdv=True),\n                         lightgbm.early_stopping(stopping_rounds=PATIENCE)],\n             init_model=mdl_A)","metadata":{"execution":{"iopub.status.busy":"2022-10-01T22:56:53.301405Z","iopub.execute_input":"2022-10-01T22:56:53.302424Z","iopub.status.idle":"2022-10-01T23:05:49.011671Z","shell.execute_reply.started":"2022-10-01T22:56:53.302379Z","shell.execute_reply":"2022-10-01T23:05:49.010551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in test_files:\n    test = pd.read_csv(os.path.join(DAT_DIR, f\"train_{i}.csv\"))\n    pred_probs = mdl_A.predict_proba(test[x_cols])[:, 1]\n    test['pred_probs'] = pred_probs\n    print(f'LogLoss of test on train_{i}.csv: {round(log_loss(test[\"team_A_scoring_within_10sec\"], test[\"pred_probs\"]), 3)}')","metadata":{"execution":{"iopub.status.busy":"2022-10-01T23:07:27.558957Z","iopub.execute_input":"2022-10-01T23:07:27.559490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us check the importance of variables.","metadata":{}},{"cell_type":"code","source":"lightgbm.plot_importance(mdl_A, max_num_features=20)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is quite interesting to note that the feature importances are not what I have assumed in my [EDA notebook](https://www.kaggle.com/code/stautxie/tabular-challenge-oct-2022-eda). Therefore, I had assumed $y$-coordinates should be the most relevant. However, the lightgbm model seems to suggest the $x$-coordinates of ball as well as boost items are the most important.","metadata":{}},{"cell_type":"markdown","source":"### Fitting and Predicting Team B","metadata":{}},{"cell_type":"markdown","source":"We apply the same approach to predict scoring probability of team B in the next 10 seconds.","metadata":{}},{"cell_type":"code","source":"MODEL_PARAMS = {'random_state': 1234,    \n                'learning_rate': 0.1,                \n                'n_estimators': MAX_ITER}\n\nmdl_B = lightgbm.LGBMClassifier(**MODEL_PARAMS)\n\nmdl_B.fit(X=train[x_cols], y=train[y_cols[1]],\n          eval_set=[(valid[x_cols], valid[y_cols[1]])],\n          eval_names=['valid'],\n          callbacks=[lightgbm.log_evaluation(period=DISPLAY_FREQ, show_stdv=True),\n                     lightgbm.early_stopping(stopping_rounds=PATIENCE)])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in train_files[1:]:\n    print(f'Iteration {i}, Reading train_{i}.csv')\n    train_new = pd.read_csv(os.path.join(DAT_DIR, f'train_{i}.csv'))\n    print('Fitting model')\n    mdl_B.fit(X=train_new[x_cols], \n              y=train_new[y_cols[1]],\n              eval_set=[(valid[x_cols], valid[y_cols[1]])],\n              eval_names=['valid'],\n              callbacks=[lightgbm.log_evaluation(period=DISPLAY_FREQ, show_stdv=True),\n                         lightgbm.early_stopping(stopping_rounds=PATIENCE)],\n             init_model=mdl_B)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in test_files:\n    test = pd.read_csv(os.path.join(DAT_DIR, f\"train_{i}.csv\"))\n    pred_probs = mdl_B.predict_proba(test[x_cols])[:, 1]\n    test['pred_probs'] = pred_probs\n    print(f'LogLoss of test on train_{i}.csv: {round(log_loss(test[\"team_B_scoring_within_10sec\"], test[\"pred_probs\"]), 3)}')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lightgbm.plot_importance(mdl_A, max_num_features=20)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that the prediction quality and variable importance of the team B's are very similar to that of the team As.","metadata":{}},{"cell_type":"markdown","source":"### Predicting on Test","metadata":{}},{"cell_type":"code","source":"final_test = pd.read_csv(os.path.join(DAT_DIR, \"test.csv\"))\nsample_submit = pd.read_csv(os.path.join(DAT_DIR, \"sample_submission.csv\"))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_test","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submit[\"team_A_scoring_within_10sec\"] = mdl_A.predict_proba(final_test[x_cols])[:,1]\nsample_submit[\"team_B_scoring_within_10sec\"] = mdl_B.predict_proba(final_test[x_cols])[:,1]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submit","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submit.to_csv(\"submission.csv\", index=None)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusions","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we present a boosting approach to apply lightgbm to the problem. By doing that, we demonstrate such an approach delivers similar or better results than using all data together. Meanwhile, training data by batch significantly reduces the memory requirement and permits much larger model to be experimented.\n\nThe variable importance analysis suggests insights quite contrary to our previous expectation. Taking into consideration the marginal improvement of lightgbm model over a [naive single-variable model](https://www.kaggle.com/code/nigelhenry/single-feature-baseline), it is possible that the native features are not very strong.  Therefore, more sophisticated features are needed to significantly improve the model performance.","metadata":{}}]}