{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🎓[Training] Student Perf from Gameplay\n---\nIn this notebook, I used the baseline of *Chris Deotte* based on the idea proposed by *DataManyo*. You can check out there respective works:\n- [XGBoost Baseline - [0.676]](https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/notebook)\n- [🥇 LightGBM baseline with aggregated log data](https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data)\n\nI am also using **RAPIDS cuDF** instead of pandas to take advantage of GPU acceleration for reading the data, feature engineering, and training. The modifications are showed in this notebook by *Shashwat Raman*.\n- [GPU XGB Baseline (Using RAPIDS cuDF) - Train](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train/notebook)\n\n\nI will keep track of my progress during the start of the competition. My goal now is to improve this baseline.\n\n## ✨Progress\n- ***Baseline: LB Score:* 0.676**\n - Training using 18 XGBoost models for individual questions.\n - The data is grouped into 3 groups of questions and statistics of the features are calculated.\n - Cross validation on 5 folds and training on 80% of the data (last fold).\n- ***Increasing number of CV folds:* LB Score: 0.677**\n - Using 20 folds for cross-validation and training on 95% of the data [(Chris Deotte suggestion)](https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676/comments#2146075).\n - Tried with 10 folds - LB Score: 0.676\n- ***Training on full data:* LB Score: 0.667** (with 5 folds)\n - Averaging the `best_ntree_limit` along the folds for each question.\n - Try with 5, 10, 20 folds for the average\n - The performance was slightly better by averaging 5 folds\n- **LB Score:**\n - Added min/max features\n - Added room counts\n\n# Parameters\nHere are the parameters related to the model strategy.\n- `xgb_params`: You can check the different parameters available for XGBoost. Check the documentation [here](https://xgboost.readthedocs.io/en/stable/).\n- `SAVE`: Save the 'models' dictionnary as json files.\n- `N_FOLDS`: Number of folds for the cross-validation.\n- `RETRAIN_WITH_ALL_DATA`: Retrain or not a final model using all of the data, with the number of estimators for each model determined by the averages obtained during cross-validation. If set to False, the final model is trained using the exact numbers of estimators but on the last fold of cross-validation.\n- `INFERENCE`: Choose to use this notebook for inference. Use CPU for it.\n- `MODEL_PATH`: The path to the dictionary of models stored as a dataset.","metadata":{}},{"cell_type":"code","source":"PATH = '/kaggle/input/predict-student-performance-from-game-play'\n\n# Model params\nxgb_params = {'objective': 'binary:logistic',\n              'eval_metric': 'logloss',\n              'learning_rate': 0.05,\n              'max_depth': 4,\n              'n_estimators': 1000,\n              'early_stopping_rounds': 50,\n              'tree_method': 'gpu_hist',\n              'subsample': 0.8,\n              'colsample_bytree': 0.4,\n              'use_label_encoder': False}\n\n# Save the models as json files\nSAVE = True\n\n# Cross validation\nN_FOLDS = 5\n\n# Retrain using all data\nRETAIN_WITH_ALL_DATA = True\n\n# Use as inference notebook (Use with CPU)\nINFERENCE = False\nMODELS_DIR = '/kaggle/input/student-perf-from-gameplay-models'","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:52:13.776809Z","iopub.execute_input":"2023-02-26T14:52:13.777228Z","iopub.status.idle":"2023-02-26T14:52:13.785202Z","shell.execute_reply.started":"2023-02-26T14:52:13.777185Z","shell.execute_reply":"2023-02-26T14:52:13.784202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport json\n\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nfrom sklearn.model_selection import KFold, GroupKFold\nimport xgboost as xgb\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score\n\nfrom tqdm.notebook import tqdm\nimport zipfile\nfrom IPython.display import FileLink\n\nif INFERENCE:\n    import pandas as pd\n    print('Pandas version:', pd.__version__)\nelse:\n    import cudf\n    print('RAPIDS cuDF version:', cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:52:14.323765Z","iopub.execute_input":"2023-02-26T14:52:14.325635Z","iopub.status.idle":"2023-02-26T14:52:18.915743Z","shell.execute_reply.started":"2023-02-26T14:52:14.325594Z","shell.execute_reply":"2023-02-26T14:52:18.914560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read the training data","metadata":{}},{"cell_type":"code","source":"if not INFERENCE:\n    train = cudf.read_csv(f'{PATH}/train.csv')\n    labels = cudf.read_csv(f'{PATH}/train_labels.csv')\n    labels['session'] = cudf.to_numeric(labels.session_id.str.split('_').list.get(0))\n    labels['question'] = cudf.to_numeric(labels.session_id.str.split('_q').list.get(1))\n    print(\"Training set loaded.\")\nelse:\n    print(\"Inference. No training set needed.\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:52:18.917658Z","iopub.execute_input":"2023-02-26T14:52:18.918314Z","iopub.status.idle":"2023-02-26T14:53:28.863083Z","shell.execute_reply.started":"2023-02-26T14:52:18.918274Z","shell.execute_reply":"2023-02-26T14:53:28.861862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature engineering\nI will use the idea of aggreagating the data by `session_id` and `level_group`. Check out the references at the beginning of the notebook.\n\nThe aim is to create interesting features from these groups that we can use to train one model per question level.","metadata":{}},{"cell_type":"code","source":"def feature_engineer(train):\n    \n    # Prepare features\n    GROUP_COLS = ['session_id', 'level_group']\n    CATS = ['event_name', 'name', 'text', 'fqid', 'room_fqid', 'text_fqid']\n    NUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n    EVENT_NAMES = ['navigate_click', 'person_click', 'cutscene_click',\n               'object_click', 'map_hover', 'notification_click',\n               'map_click', 'observation_click', 'checkpoint']\n    ROOMS = ['tunic.capitol_0.hall', 'tunic.capitol_1.hall',\n             'tunic.capitol_2.hall', 'tunic.drycleaner.frontdesk',\n             'tunic.flaghouse.entry', 'tunic.historicalsociety.basement',\n             'tunic.historicalsociety.cage', 'tunic.historicalsociety.closet',\n             'tunic.historicalsociety.closet_dirty',\n             'tunic.historicalsociety.collection',\n             'tunic.historicalsociety.collection_flag',\n             'tunic.historicalsociety.entry',\n             'tunic.historicalsociety.frontdesk',\n             'tunic.historicalsociety.stacks', 'tunic.humanecology.frontdesk',\n             'tunic.kohlcenter.halloffame', 'tunic.library.frontdesk',\n             'tunic.library.microfiche', 'tunic.wildlife.center']\n    \n    dfs = []\n    \n    # N unique\n    for c in CATS:\n        tmp = train.groupby(GROUP_COLS)[c].agg('nunique')\n        tmp.name = tmp.name + '_nunique'\n        dfs.append(tmp)\n\n    # Mean\n    for c in NUMS:\n        tmp = train.groupby(GROUP_COLS)[c].agg('mean')\n        tmp.name = tmp.name + '_mean'\n        dfs.append(tmp)\n\n    # Standard deviation\n    for c in NUMS:\n        tmp = train.groupby(GROUP_COLS)[c].agg('std')\n        tmp.name = tmp.name + '_std'\n        dfs.append(tmp)\n    \n    # Maximum\n    for c in NUMS:\n        tmp = train.groupby(GROUP_COLS)[c].agg('max')\n        tmp.name = tmp.name + '_max'\n        dfs.append(tmp)\n    \n    # Minimum\n    for c in NUMS:\n        tmp = train.groupby(GROUP_COLS)[c].agg('min')\n        tmp.name = tmp.name + '_min'\n        dfs.append(tmp)\n        \n    # Count event names\n    for c in EVENT_NAMES:\n        train[c] = (train['event_name'] == c).astype('int8')\n        tmp = train.groupby(GROUP_COLS)[c].agg('sum')\n        tmp.name = tmp.name + '_count'\n        dfs.append(tmp)\n        train = train.drop(c, axis=1)\n        \n    # Count rooms\n    for c in ROOMS:\n        train[c] = (train['room_fqid'] == c).astype('int8')\n        tmp = train.groupby(GROUP_COLS)[c].agg('sum')\n        tmp.name = tmp.name + '_count'\n        dfs.append(tmp)\n        train = train.drop(c, axis=1)\n        \n    # Concatenate everything\n    try: df = cudf.concat(dfs, axis=1)\n    except: df = pd.concat(dfs, axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:53:28.865340Z","iopub.execute_input":"2023-02-26T14:53:28.866022Z","iopub.status.idle":"2023-02-26T14:53:28.882075Z","shell.execute_reply.started":"2023-02-26T14:53:28.865980Z","shell.execute_reply":"2023-02-26T14:53:28.880780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have defined the feature engineering function, we can apply it to the raw dataframe and examine the resulting aggregated dataframe, which we will use for training. Note that we have not normalized the features, because the XGBoost classifier we are using is a tree-based model. Typically, tree-based models do not require normalization, as their decision boundaries are based on the relative values of the features.","metadata":{}},{"cell_type":"code","source":"if not INFERENCE:\n    %time df = feature_engineer(train).to_pandas()\n    FEATURES = [c for c in df.columns if c != 'level_group']\n    SESSIONS = df.index.unique()\n    print(\"Nb of features used for training:\", len(FEATURES))\n    print(\"Nb of different sessions:\", len(SESSIONS))\nelse:\n    print(\"Inference. Feature engineering on the test set only.\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:53:28.885091Z","iopub.execute_input":"2023-02-26T14:53:28.885611Z","iopub.status.idle":"2023-02-26T14:53:37.681497Z","shell.execute_reply.started":"2023-02-26T14:53:28.885571Z","shell.execute_reply":"2023-02-26T14:53:37.680412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Cross validation\nThe purpose of cross-validation is to **evaluate the performance of our model** and to assess its generalization capability. The idea behind cross-validation is to split the data into several folds, and to use each fold as a validation set while the remaining data is used for training. This process is repeated several times, with each fold being used as the validation set once. It also helps to avoid overfitting, which occurs when a model is too complex and fits the training data too closely, leading to poor performance on new data.","metadata":{}},{"cell_type":"code","source":"if not INFERENCE:\n    \n    gkf = GroupKFold(n_splits=N_FOLDS)\n    oof = cudf.DataFrame(data=np.zeros((len(SESSIONS), 18)), index=SESSIONS)\n    models = {}\n    best_ntree_limits = {} # to keep track of it in case of retraining\n\n    # Compute CV score with 5-group K-fold\n    for i, (train_index, valid_index) in enumerate(gkf.split(X=df, groups=df.index)):\n        print('\\033[1m\\033[94m-' * 80, f\"\\n{' ' * 35}FOLD {i + 1}\")\n        print('-' * 80, '\\033[0m')\n\n        # Iterate from question 1 to 18\n        for q in range(1, 19):\n\n            # Select the question group\n            if q <= 3: grp = '0-4'\n            elif q <=13: grp = '5-12'\n            elif q <= 22: grp = '13-22'\n\n            # Train data\n            train_x = df.iloc[train_index]\n            train_x = train_x[train_x.level_group == grp]\n            train_sessions = train_x.index.values\n            train_y = labels[labels.question == q]\\\n                .set_index('session').loc[train_sessions]\n\n            # Valid data\n            valid_x = df.iloc[valid_index]\n            valid_x = valid_x[valid_x.level_group == grp]\n            valid_sessions = valid_x.index.values\n            valid_y = labels[labels.question == q]\\\n                .set_index('session').loc[valid_sessions]\n\n            # Train model\n            clf = XGBClassifier(**xgb_params)\n            clf.fit(train_x[FEATURES].astype('float32'),\n                    train_y['correct'],\n                    eval_set=[(valid_x[FEATURES].astype('float32'),\n                               valid_y['correct'])],\n                    verbose=0)\n\n            # Best_ntree_limit\n            best_ntree_limits[f'fold{i+1}_q{q}'] = clf.best_ntree_limit\n            print(f'Q{q}({clf.best_ntree_limit}), ', end='')\n\n            # Save and predict valid\n            models[f'{grp}_{q}'] = clf\n            oof.loc[valid_sessions, q - 1] = clf.predict_proba(valid_x[FEATURES])[:, 1]\n\n        print()\n        \n    print(\"\\nCross validation finished.\")\n    \nelse:\n    print(\"Inference. No cross validation.\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:53:37.683393Z","iopub.execute_input":"2023-02-26T14:53:37.684062Z","iopub.status.idle":"2023-02-26T14:54:30.107357Z","shell.execute_reply.started":"2023-02-26T14:53:37.684025Z","shell.execute_reply":"2023-02-26T14:54:30.106452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The models predicted values on the different validation folds are stored in the **out-of-fold** (oof) dataframe. These predictions represent the confidence that a given `session_id` and `question` are correct, on a scale from 0 to 1. We must now decide when to consider a question as correct or not by finding the threshold that gives the best final score. While we could arbitrarily take 0.5, **this value is not necessarily the best one**, and it is smarter to try different values to find the optimal threshold.","metadata":{}},{"cell_type":"code","source":"if not INFERENCE:\n    \n    # Put true labels into dataframe with 18 columns\n    true = oof.copy()\n    for q in range(1, 19):\n        tmp = labels.loc[labels.question == q].set_index('session')\n        true[q - 1] = tmp['correct'].values\n\n    # Find the optimal threshold for the final preditions\n    scores, thresholds = [], []\n    best_score, best_threshold = 0, 0\n    for thr in tqdm(np.arange(0.550, 0.705, 0.005)):\n        preds = (oof.to_pandas().values > thr).astype(int).reshape(-1)\n        m = f1_score(true.to_pandas().values.reshape(-1), preds, average='macro')\n        scores.append(m)\n        thresholds.append(thr)\n        if m > best_score:\n            best_score = m\n            best_threshold = round(thr, 3)\n    models['threshold'] = best_threshold\n\n    # Display threshold curve\n    plt.figure(figsize=(16, 5))\n    plt.scatter(best_threshold, best_score, s=200,\n                marker='x', color='red', lw=5, zorder=2)\n    plt.plot(thresholds, scores, zorder=0)\n    plt.scatter(thresholds, scores, color='black')\n    plt.legend(['Best F1-score'], loc='upper left')\n    plt.xlabel(\"Threshold\", size=12)\n    plt.ylabel(\"Validation F1-score\", size=12)\n    plt.title(f\"Best F1 = {best_score:.4f} | Threshold = {best_threshold}\", size=16)\n    plt.show()\n\n    # Detail for each model\n    print(f\"Using optimal threshold of {best_threshold}\")\n    print('-' * 35)\n    for k in range(18):\n    \n        # Compute F1 score per question\n        preds = (oof[k].to_pandas().values > best_threshold).astype(int)\n        m = f1_score(true[k].to_pandas().values, preds, average='macro')\n        print(f\" Q{k+1}: F1 = {m:.6f}\")\n    \n    preds = (oof.to_pandas().values.reshape((-1)) > best_threshold).astype('int')\n    m = f1_score(true.to_pandas().values.reshape((-1)), preds, average='macro')\n    print('-' * 35)\n    print(f\"Overall F1 = {m:.6f}\")\n\nelse:\n    print(\"Inference.\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:54:30.109046Z","iopub.execute_input":"2023-02-26T14:54:30.109800Z","iopub.status.idle":"2023-02-26T14:54:33.945538Z","shell.execute_reply.started":"2023-02-26T14:54:30.109758Z","shell.execute_reply":"2023-02-26T14:54:33.944352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After trying different thresholds in the range of [0.40, 0.80], we observe that the validation F1-score follows a bell curve. The optimal threshold is the one that **maximizes this curve**, and we should use this value when reusing this model for inference.\n\n# Retraining\nThe model can be retrained on **all the data**  after cross-validation. To do this, we take the average of the `best_ntree_limit` values across the folds for each of the 18 models, since the stopping criterion is based on a validation set. However, we do not have access to a validation set when retraining the model on all the data.","metadata":{}},{"cell_type":"code","source":"if INFERENCE:\n    print(\"Inference mode. Use models from the dataset.\")\n\nelif RETAIN_WITH_ALL_DATA:\n    print('\\033[1m\\033[94m-' * 80, f\"\\n{' ' * 35}RETRAINING\")\n    print('-' * 80, '\\033[0m')\n    \n    # Average best trees\n    estimators = []\n    for q in range(1, 19):\n        avg = 0\n        for fold in range(1, N_FOLDS + 1):\n            avg += best_ntree_limits[f'fold{fold}_q{q}'] / N_FOLDS\n        estimators.append(round(avg))\n        \n    # Disable early stopping\n    xgb_params['early_stopping_rounds'] = None\n\n    models = {}\n    # Iterate from question 1 to 18\n    for q in range(1, 19):   \n        # Select the number of estimators\n        xgb_params['n_estimators'] = estimators[q - 1]\n        \n        # Select the question group\n        if q <= 3: grp = '0-4'\n        elif q <=13: grp = '5-12'\n        elif q <= 22: grp = '13-22'\n\n        # Train data\n        train_x = df[df.level_group == grp]\n        train_y = labels[labels.question == q].set_index('session')\n\n        # Train model\n        clf = XGBClassifier(**xgb_params)\n        clf.fit(train_x[FEATURES].astype('float32'),\n                train_y['correct'],\n                verbose=0)\n\n        print(f'Q{q}({clf.best_ntree_limit}), ', end='')\n\n        # Save model\n        models[f'{grp}_{q}'] = clf\n\n    print(\"\\n\\nThe models have been retrained using all the data.\")\n\nelse:\n    print(\"No retraining. Models trained with last fold.\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:54:33.948829Z","iopub.execute_input":"2023-02-26T14:54:33.949129Z","iopub.status.idle":"2023-02-26T14:54:38.252269Z","shell.execute_reply.started":"2023-02-26T14:54:33.949103Z","shell.execute_reply":"2023-02-26T14:54:38.251295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save models\nI save the dictionary of models in **json format** in order to reuse it in another inference notebook.","metadata":{}},{"cell_type":"code","source":"if INFERENCE:\n    print(\"Inference mode.\")\nelif SAVE:\n    os.makedirs('models', exist_ok=True)\n    # Save threshold\n    thr_dict = {'threshold': best_threshold}\n    with open(\"models/threshold.json\", \"w\") as f:\n        json.dump(thr_dict, f)\n    # Save models\n    for name, model in models.items():\n        model.save_model(f'models/{name}.json')\n    # Zip the folder\n    folder_name = 'models'\n    with zipfile.ZipFile('models.zip', mode=\"w\") as my_zip:\n        for root, _, files in os.walk(folder_name):\n            for file in files:\n                file_path = os.path.join(root, file)\n                arcname = os.path.relpath(file_path, folder_name)\n                my_zip.write(file_path, arcname=arcname)\nelse:\n    print(\"Models not saved.\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:54:38.253594Z","iopub.execute_input":"2023-02-26T14:54:38.253942Z","iopub.status.idle":"2023-02-26T14:54:38.398397Z","shell.execute_reply.started":"2023-02-26T14:54:38.253905Z","shell.execute_reply":"2023-02-26T14:54:38.397472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference\nInference **should use the CPU**","metadata":{}},{"cell_type":"code","source":"if INFERENCE:\n    \n    # Use the competition env\n    import jo_wilder\n    env = jo_wilder.make_env()\n    iter_test = env.iter_test()\n\n    # Predict and create submission file\n    limits = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}\n    with open(f'{MODELS_DIR}/threshold.json', 'r') as f:\n        threshold = json.load(f)['threshold']\n    for (sample_submission, test) in iter_test:\n        df = feature_engineer(test)\n        FEATURES = [c for c in df.columns if c != 'level_group']\n        grp = test.level_group.values[0]\n        a, b = limits[grp]\n        for q in range(a, b):\n            clf = XGBClassifier()\n            clf.load_model(f'{MODELS_DIR}/{grp}_{q}.json')\n            p = clf.predict_proba(df[FEATURES])[:, 1]\n            mask = sample_submission.session_id.str.contains(f'q{q}')\n            sample_submission.loc[mask, 'correct'] = int(p.item() > threshold)\n        env.predict(sample_submission)\n\n    # Check submission\n    df = pd.read_csv('submission.csv')\n    print(df.shape)\n\nelse:\n    print(\"No inference.\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T14:54:38.399885Z","iopub.execute_input":"2023-02-26T14:54:38.400281Z","iopub.status.idle":"2023-02-26T14:54:38.410162Z","shell.execute_reply.started":"2023-02-26T14:54:38.400243Z","shell.execute_reply":"2023-02-26T14:54:38.409001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Thank you for reading me, good luck in the competition! 🙂**","metadata":{}}]}