{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# INTRODUCTION\nCalculating statistics on text-based features such as \"text\" or \"fqid\" may be effective in predicting student performance.  However, since these features contain over 100 unique values (over 500 for \"text\"), using all of them as features is impractical because of the computational cost.\n\nTherefore, we thought it would be better to calculate and compare the feature importance of each value, and use the top few values as the features.  In a previous post, linked below, we determined the importance of the \"text\", \"fqid\", and \"text_fqid\" features.  The prediction accuracy (CV-Score) using only each feature was 0.688, 0.692, and 0.684, respectively, suggesting that the importance of \"fqid\" was relatively high.  In this study, we generated more features for \"fqid\" and investigated the feature importance in more detail.  This notebook is the Train part of a series of investigations.","metadata":{}},{"cell_type":"markdown","source":"## Previous post\n- About \"text\" feature: https://www.kaggle.com/code/tsuyoshifujii/xgboost-using-only-text-which-text-is-important\n- About \"fqid\" feature: https://www.kaggle.com/code/tsuyoshifujii/xgboost-using-only-fqid-feature-importance\n- About \"text_fqid\" feature: https://www.kaggle.com/code/tsuyoshifujii/xgboost-using-only-text-fqid-feature-importance","metadata":{}},{"cell_type":"markdown","source":"## Reference\nThis code is based on the following amazing notebooks.\n- https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680\n- https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\n- https://www.kaggle.com/code/pourchot/simple-xgb\n\nThe idea of training on GPU and Inference on CPU was based on the following discussion.\n- https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218","metadata":{}},{"cell_type":"markdown","source":"# IMPORTS","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom sklearn.metrics import f1_score\nfrom xgboost import XGBClassifier\n\nfrom tqdm.notebook import tqdm\nfrom collections import defaultdict\nfrom itertools import combinations\nimport pickle\nimport warnings\nimport gc\n\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:40:10.857523Z","iopub.execute_input":"2023-04-12T02:40:10.858470Z","iopub.status.idle":"2023-04-12T02:40:11.482771Z","shell.execute_reply.started":"2023-04-12T02:40:10.858413Z","shell.execute_reply":"2023-04-12T02:40:11.481761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LOAD TRAIN DATA AND LABELS","metadata":{}},{"cell_type":"code","source":"# READ USER ID ONLY\ntmp = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', usecols=[0])\ntmp = tmp.groupby('session_id').session_id.agg('count')\nprint(f'all users: {len(tmp)}')\n\n# COMPUTE READS AND SKIPS\nPIECES = 10\nCHUNK = int(np.ceil(len(tmp)/PIECES))\n\nreads = []\nskips = [0]\nfor k in range(PIECES):\n    a = k * CHUNK\n    b = (k + 1) * CHUNK\n    if b > len(tmp):\n        b = len(tmp)\n    r = tmp.iloc[a:b].sum()\n    reads.append(r)\n    skips.append(skips[-1] + r)\n    \nprint(f'To avoid memory error, we will read train in {PIECES} pieces of sizes:')\nprint(reads)","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:40:11.486637Z","iopub.execute_input":"2023-04-12T02:40:11.486938Z","iopub.status.idle":"2023-04-12T02:41:25.968158Z","shell.execute_reply.started":"2023-04-12T02:40:11.486911Z","shell.execute_reply":"2023-04-12T02:41:25.966118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', nrows=reads[0])\nprint('Train size of first piece: ', train.shape)\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:41:25.969464Z","iopub.execute_input":"2023-04-12T02:41:25.970161Z","iopub.status.idle":"2023-04-12T02:41:31.172165Z","shell.execute_reply.started":"2023-04-12T02:41:25.970131Z","shell.execute_reply":"2023-04-12T02:41:31.170864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets['session_id'].apply(lambda x: int(x.split('_')[0]))\ntargets['q'] = targets['session_id'].apply(lambda x: int(x.split('_')[-1][1:]))\nprint(targets.shape)\ntargets.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:41:31.175299Z","iopub.execute_input":"2023-04-12T02:41:31.176057Z","iopub.status.idle":"2023-04-12T02:41:32.145541Z","shell.execute_reply.started":"2023-04-12T02:41:31.176012Z","shell.execute_reply":"2023-04-12T02:41:32.144423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:41:32.147149Z","iopub.execute_input":"2023-04-12T02:41:32.148227Z","iopub.status.idle":"2023-04-12T02:41:32.246473Z","shell.execute_reply.started":"2023-04-12T02:41:32.148188Z","shell.execute_reply":"2023-04-12T02:41:32.245259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# FEATURE ENGINEER\nPerform feature engineering only for \"fqid\".\n\nStatistics (mean, median, std, sum, min, max) are calculated for each element of \"fqid\" for the following five types of features, and these are used as new feature values.\n- Change in \"elapsed_time\"\n- Change in \"room_coor_x\"\n- Change in \"room_coor_y\"\n- Change in \"screen_coor_x\"\n- Change in \"screen_coor_y\"","metadata":{}},{"cell_type":"code","source":"tmp = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', usecols=['session_id', 'fqid'])\nFQIDS = tmp['fqid'].unique()\n\ndel tmp\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:41:32.248471Z","iopub.execute_input":"2023-04-12T02:41:32.249410Z","iopub.status.idle":"2023-04-12T02:41:56.687490Z","shell.execute_reply.started":"2023-04-12T02:41:32.249368Z","shell.execute_reply":"2023-04-12T02:41:56.686486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_engineer(x, grp):\n    # CALCULATE INTERVAL FOR EACH ACTION\n    x['time_diff'] = x['elapsed_time'] - x.groupby('session_id')['elapsed_time'].shift(1)\n    mask = x['time_diff'] < 0\n    x.loc[mask, 'time_diff'] = 0\n    x['time_diff'] = x['time_diff'] / 1000\n    \n    # ROOM LOCATION X CHANGED FOR CLICK\n    x['room_x_change'] = x['room_coor_x'] - x.groupby('session_id')['room_coor_x'].shift(1)\n    \n    # ROOM LOCATION Y CHANGED FOR CLICK\n    x['room_y_change'] = x['room_coor_y'] - x.groupby('session_id')['room_coor_y'].shift(1)\n    \n    # LOCATION X CHANGED FOR CLICK\n    x['screen_x_change'] = x['screen_coor_x'] - x.groupby('session_id')['screen_coor_x'].shift(1)\n    \n    # LOCATION Y CHANGED FOR CLICK\n    x['screen_y_change'] = x['screen_coor_y'] - x.groupby('session_id')['screen_coor_y'].shift(1)\n    \n    df_final = x.groupby('session_id')['index'].agg('count')\n    df_final.name = 'num_events'\n    df_final = df_final.reset_index()\n    df_final = df_final.set_index('session_id')\n    \n    for c in FQIDS:\n        x[c] = (x['fqid'] == c).astype('int8')\n    for c in FQIDS:\n        df_final[f'{c}_cnSum'] = x.groupby('session_id')[c].agg('sum')\n    x.drop(FQIDS, axis=1, inplace=True)\n        \n    for c in FQIDS:\n        df_final[f'{c}_etAve'] = x[x['fqid'] == c].groupby('session_id')['time_diff'].mean()\n        df_final[f'{c}_etMed'] = x[x['fqid'] == c].groupby('session_id')['time_diff'].median()\n        df_final[f'{c}_etStd'] = x[x['fqid'] == c].groupby('session_id')['time_diff'].std()\n        df_final[f'{c}_etSum'] = x[x['fqid'] == c].groupby('session_id')['time_diff'].sum()\n        df_final[f'{c}_etMin'] = x[x['fqid'] == c].groupby('session_id')['time_diff'].min()\n        df_final[f'{c}_etMax'] = x[x['fqid'] == c].groupby('session_id')['time_diff'].max()\n        \n        df_final[f'{c}_rxAve'] = x[x['fqid'] == c].groupby('session_id')['room_x_change'].mean()\n        df_final[f'{c}_rxMed'] = x[x['fqid'] == c].groupby('session_id')['room_x_change'].median()\n        df_final[f'{c}_rxStd'] = x[x['fqid'] == c].groupby('session_id')['room_x_change'].std()\n        df_final[f'{c}_rxSum'] = x[x['fqid'] == c].groupby('session_id')['room_x_change'].sum()\n        df_final[f'{c}_rxMin'] = x[x['fqid'] == c].groupby('session_id')['room_x_change'].min()\n        df_final[f'{c}_rxMax'] = x[x['fqid'] == c].groupby('session_id')['room_x_change'].max()\n        \n        df_final[f'{c}_ryAve'] = x[x['fqid'] == c].groupby('session_id')['room_y_change'].mean()\n        df_final[f'{c}_ryMed'] = x[x['fqid'] == c].groupby('session_id')['room_y_change'].median()\n        df_final[f'{c}_ryStd'] = x[x['fqid'] == c].groupby('session_id')['room_y_change'].std()\n        df_final[f'{c}_rySum'] = x[x['fqid'] == c].groupby('session_id')['room_y_change'].sum()\n        df_final[f'{c}_ryMin'] = x[x['fqid'] == c].groupby('session_id')['room_y_change'].min()\n        df_final[f'{c}_ryMax'] = x[x['fqid'] == c].groupby('session_id')['room_y_change'].max()\n        \n        df_final[f'{c}_sxAve'] = x[x['fqid'] == c].groupby('session_id')['screen_x_change'].mean()\n        df_final[f'{c}_sxMed'] = x[x['fqid'] == c].groupby('session_id')['screen_x_change'].median()\n        df_final[f'{c}_sxStd'] = x[x['fqid'] == c].groupby('session_id')['screen_x_change'].std()\n        df_final[f'{c}_sxSum'] = x[x['fqid'] == c].groupby('session_id')['screen_x_change'].sum()\n        df_final[f'{c}_sxMin'] = x[x['fqid'] == c].groupby('session_id')['screen_x_change'].min()\n        df_final[f'{c}_sxMax'] = x[x['fqid'] == c].groupby('session_id')['screen_x_change'].max()\n        \n        df_final[f'{c}_syAve'] = x[x['fqid'] == c].groupby('session_id')['screen_y_change'].mean()\n        df_final[f'{c}_syMed'] = x[x['fqid'] == c].groupby('session_id')['screen_y_change'].median()\n        df_final[f'{c}_syStd'] = x[x['fqid'] == c].groupby('session_id')['screen_y_change'].std()\n        df_final[f'{c}_sySum'] = x[x['fqid'] == c].groupby('session_id')['screen_y_change'].sum()\n        df_final[f'{c}_syMin'] = x[x['fqid'] == c].groupby('session_id')['screen_y_change'].min()\n        df_final[f'{c}_syMax'] = x[x['fqid'] == c].groupby('session_id')['screen_y_change'].max()\n        \n    df_final = df_final.fillna(0)\n    df_final = df_final.drop(columns='num_events')\n    \n    return df_final","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:41:56.689039Z","iopub.execute_input":"2023-04-12T02:41:56.689584Z","iopub.status.idle":"2023-04-12T02:41:56.711979Z","shell.execute_reply.started":"2023-04-12T02:41:56.689547Z","shell.execute_reply":"2023-04-12T02:41:56.710997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# PROCESS TRAIN DATA IN PIECES\nall_pieces_df1 = []\nall_pieces_df2 = []\nall_pieces_df3 = []\n\nfor k in range(PIECES):\n    SKIPS = 0\n    if k > 0:\n        SKIPS = range(1, skips[k] + 1)\n    train = pd.read_csv(\n        '/kaggle/input/predict-student-performance-from-game-play/train.csv',\n        usecols=['session_id', 'index', 'elapsed_time', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'fqid', 'level_group'],\n        nrows=reads[k],\n        skiprows=SKIPS\n    )\n    train1 = train[train['level_group'] == '0-4']\n    train2 = train[train['level_group'] == '5-12']\n    train3 = train[train['level_group'] == '13-22']\n    \n    df1 = feature_engineer(train1, grp='0-4')\n    df2 = feature_engineer(train2, grp='5-12')\n    df3 = feature_engineer(train3, grp='13-22')\n    \n    all_pieces_df1.append(df1)\n    all_pieces_df2.append(df2)\n    all_pieces_df3.append(df3)\n    \n    print(f'# PIECE {k + 1} COMPLETED')\n    \n# CONCATENATE ALL PIECES\ndel train, train1, train2, train3; gc.collect()\n\ndf1 = pd.concat(all_pieces_df1, axis=0)\ndf2 = pd.concat(all_pieces_df2, axis=0)\ndf3 = pd.concat(all_pieces_df3, axis=0)\nprint(f'df1: {df1.shape}')\nprint(f'df2: {df2.shape}')\nprint(f'df3: {df3.shape}')","metadata":{"execution":{"iopub.status.busy":"2023-04-12T02:41:56.713442Z","iopub.execute_input":"2023-04-12T02:41:56.713820Z","iopub.status.idle":"2023-04-12T04:09:43.131182Z","shell.execute_reply.started":"2023-04-12T02:41:56.713770Z","shell.execute_reply":"2023-04-12T04:09:43.130026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"null1 = df1.isnull().sum().sort_values(ascending=False) / len(df1)\nnull2 = df2.isnull().sum().sort_values(ascending=False) / len(df2)\nnull3 = df3.isnull().sum().sort_values(ascending=False) / len(df3)\n\ndrop1 = list(null1[null1 > 0.9].index)\ndrop2 = list(null2[null2 > 0.9].index)\ndrop3 = list(null3[null3 > 0.9].index)\n\nprint(f'drop1: {len(drop1)}')\nprint(f'drop2: {len(drop2)}')\nprint(f'drop3: {len(drop3)}')\n\nfor col in tqdm(df1.columns):\n    if df1[col].nunique() == 1:\n        drop1.append(col)\nprint('**********df1 DONE**********')\n\nfor col in tqdm(df2.columns):\n    if df2[col].nunique() == 1:\n        drop2.append(col)\nprint('**********df2 DONE**********')\n\nfor col in tqdm(df3.columns):\n    if df3[col].nunique() == 1:\n        drop3.append(col)\nprint('**********df3 DONE**********')","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:09:43.132963Z","iopub.execute_input":"2023-04-12T04:09:43.133353Z","iopub.status.idle":"2023-04-12T04:09:48.188952Z","shell.execute_reply.started":"2023-04-12T04:09:43.133305Z","shell.execute_reply":"2023-04-12T04:09:48.187851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:09:48.193966Z","iopub.execute_input":"2023-04-12T04:09:48.194439Z","iopub.status.idle":"2023-04-12T04:09:48.378567Z","shell.execute_reply.started":"2023-04-12T04:09:48.194409Z","shell.execute_reply":"2023-04-12T04:09:48.377104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# TRAIN XGBOOST MODEL","metadata":{}},{"cell_type":"code","source":"FEATURES1 = [c for c in df1.columns if c not in drop1 + ['level_group']]\nFEATURES2 = [c for c in df2.columns if c not in drop2 + ['level_group', 'first_index']]\nFEATURES3 = [c for c in df3.columns if c not in drop3 + ['level_group']]\n\nprint(f'We will train with df1: {len(FEATURES1)}, df2: {len(FEATURES2)}, df3: {len(FEATURES3)} features')\n\nALL_USERS = df1.index.unique()\n\nprint(f'We will train with {len(ALL_USERS)} users info')","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:09:48.383687Z","iopub.execute_input":"2023-04-12T04:09:48.386050Z","iopub.status.idle":"2023-04-12T04:09:48.811307Z","shell.execute_reply.started":"2023-04-12T04:09:48.386001Z","shell.execute_reply":"2023-04-12T04:09:48.810109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"oof_base = pd.DataFrame(data=np.zeros((len(ALL_USERS), 18)), index=ALL_USERS)\ntrue = oof_base.copy()\noof = oof_base.copy()\nfor k in range(18):\n    tmp = targets.loc[targets['q'] == k + 1].set_index('session').loc[ALL_USERS]\n    true[k] = tmp['correct'].values\n    \ntrue.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:09:48.813212Z","iopub.execute_input":"2023-04-12T04:09:48.813557Z","iopub.status.idle":"2023-04-12T04:09:48.923341Z","shell.execute_reply.started":"2023-04-12T04:09:48.813522Z","shell.execute_reply":"2023-04-12T04:09:48.922237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_fold = 5\ngkf = GroupKFold(n_splits=n_fold)\nbest_iteration_xgb = defaultdict(list)\nfeature_importance_dict = {}\n\nprint('# START TRAINING')\n\nfor t in range(1, 19):\n    if t <= 3:\n        grp = '0-4'\n        df = df1\n        FEATURES = FEATURES1\n    elif t <= 13:\n        grp = '5-12'\n        df = df2\n        FEATURES = FEATURES2\n    elif t <= 22:\n        grp = '13-22'\n        df = df3\n        FEATURES = FEATURES3\n        \n    print(f'## QUESTION {t} WITH {len(FEATURES)} FEATURES')\n    \n    params = {\n        'booster': 'gbtree',\n        'objective': 'binary:logistic',\n        'tree_method': 'gpu_hist',\n        'eval_metric': 'logloss',\n        'learning_rate': 0.02,\n        'alpha': 8,\n        'max_depth': 4,\n        'n_estimators': 9999,\n        'early_stopping_rounds': 90,\n        'subsample': 0.8,\n        'colsample_bytree': 0.5,\n        'seed': 42\n    }\n    \n    feature_importance_list = []\n    \n    # COMPUTE CV SCORE WITH 5 GROUP K FOLD\n    for i, (train_index, test_index) in enumerate(gkf.split(X=df, groups=df.index)):\n        \n        # TRAIN DATA\n        train_x = df.iloc[train_index]\n        train_users = train_x.index.values\n        train_y = targets.loc[targets['q'] == t].set_index('session').loc[train_users]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index]\n        valid_users = valid_x.index.values\n        valid_y = targets.loc[targets['q'] == t].set_index('session').loc[valid_users]\n        \n        # TRAIN MODEL\n        clf = XGBClassifier(**params)\n        clf.fit(train_x[FEATURES].astype('float32'), train_y['correct'], eval_set=[(valid_x[FEATURES].astype('float32'), valid_y['correct'])], verbose=0)\n        print(f'### FOLD {i + 1} COMPLETED')\n        best_iteration_xgb[str(t)].append(clf.best_ntree_limit)\n        \n        fold_importance_df = pd.DataFrame()\n        fold_importance_df['feature'] = FEATURES\n        fold_importance_df['importance'] = clf.feature_importances_\n        fold_importance_df['fold'] = i + 1\n        \n        feature_importance_list.append(fold_importance_df)\n        \n        # SAVE MODEL, PREDICT VALID OOF\n        oof.loc[valid_users, t-1] = clf.predict_proba(valid_x[FEATURES].astype('float32'))[:, 1]\n        \n    feature_importance_df = pd.concat(feature_importance_list)\n    feature_importance_df = feature_importance_df.groupby(['feature'])['importance'].agg(['mean']).sort_values(by='mean', ascending=False)\n    feature_importance_dict[str(t)] = feature_importance_df\n    \nf_save = open('feature_importance_dict.pkl', 'wb')\npickle.dump(feature_importance_dict, f_save)\nf_save.close()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:09:48.924921Z","iopub.execute_input":"2023-04-12T04:09:48.925576Z","iopub.status.idle":"2023-04-12T04:35:34.258051Z","shell.execute_reply.started":"2023-04-12T04:09:48.925536Z","shell.execute_reply":"2023-04-12T04:35:34.256849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CHECK RESULTS","metadata":{}},{"cell_type":"code","source":"feature_importance_dict['1'].head(10)","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:35:34.259808Z","iopub.execute_input":"2023-04-12T04:35:34.260439Z","iopub.status.idle":"2023-04-12T04:35:34.271350Z","shell.execute_reply.started":"2023-04-12T04:35:34.260404Z","shell.execute_reply":"2023-04-12T04:35:34.270119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_importance_dict['18'].head(10)","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:35:34.273077Z","iopub.execute_input":"2023-04-12T04:35:34.273803Z","iopub.status.idle":"2023-04-12T04:35:34.288366Z","shell.execute_reply.started":"2023-04-12T04:35:34.273748Z","shell.execute_reply":"2023-04-12T04:35:34.287545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EVALUATE MODELS","metadata":{}},{"cell_type":"code","source":"# FIND BEST THRESHOLD TO CONVERT PROBS INTO 1s AND 0s\nscores = []\nthresholds = []\nbest_score = 0\nbest_threshold = 0\n\nfor threshold in np.arange(0.4, 0.81, 0.005):\n    trues = true.values.reshape((-1))\n    preds = (oof.values.reshape((-1)) > threshold).astype('int')\n    f1 = f1_score(trues, preds, average='macro')\n    scores.append(f1)\n    thresholds.append(threshold)\n    \n    if f1 > best_score:\n        best_score = f1\n        best_threshold = threshold\n        \nprint(f'Best score: {best_score}')\nprint(f'Best threshold: {best_threshold}')","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:35:34.289749Z","iopub.execute_input":"2023-04-12T04:35:34.290428Z","iopub.status.idle":"2023-04-12T04:35:43.928428Z","shell.execute_reply.started":"2023-04-12T04:35:34.290391Z","shell.execute_reply":"2023-04-12T04:35:43.927193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# PLOT THRESHOLDS VS. F1 SCORES\nplt.figure(figsize=(20, 5))\nplt.plot(thresholds, scores, '-o', color='blue')\nplt.scatter([best_threshold], [best_score], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold', size=14)\nplt.ylabel('Validation F1 Score', size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score:.3f} at Best Threshold = {best_threshold:.3}', size=18)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:35:43.929902Z","iopub.execute_input":"2023-04-12T04:35:43.930514Z","iopub.status.idle":"2023-04-12T04:35:44.241893Z","shell.execute_reply.started":"2023-04-12T04:35:43.930472Z","shell.execute_reply":"2023-04-12T04:35:44.240857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('When using optimal threshold...')\nfor k in range(18):\n    # F1 SCORE PER QUESTION\n    true_k = true[k].values\n    oof_k = (oof[k].values > best_threshold).astype('int')\n    f1_k = f1_score(true_k, oof_k, average='macro')\n    \n    print(f'Q{k+1}: F1 = {f1_k:.3f}')\n\n# F1 SCORE OVERALL\ntrue_all = true.values.reshape((-1))\noof_all = (oof.values.reshape((-1)) > best_threshold).astype(int)\nf1_all = f1_score(true_all, oof_all, average='macro')\n\nprint('-'*30)\nprint(f'Overall F1 = {f1_all:.3f}')","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:35:44.242908Z","iopub.execute_input":"2023-04-12T04:35:44.243266Z","iopub.status.idle":"2023-04-12T04:35:44.487953Z","shell.execute_reply.started":"2023-04-12T04:35:44.243216Z","shell.execute_reply":"2023-04-12T04:35:44.486771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# RE-TRAIN\nIt may time out because it takes a lot of time for feature engineering, but retrain the model for inference just in case.\n\nIn this case, we will use all the generated features, but you may choose some of the more important features to use.","metadata":{}},{"cell_type":"code","source":"%%time\nfeatures_dict = {}\n\n# ITERATE FROM Q1 TO Q18\nfor t in range(1, 19):\n    if t <= 3:\n        grp = '0-4'\n        df = df1\n    elif t <= 13:\n        grp = '5-12'\n        df = df2\n    elif t <= 22:\n        grp = '13-22'\n        df = df3\n        \n    # USE THE 150 MOST IMPORTANT FEATURES\n    feature_importance = feature_importance_dict[str(t)]\n    feature_importance = feature_importance.reset_index()\n    FEATURES = feature_importance['feature']  # TO CHANGE THE NUMBER OF FEATURES USED, CHANGE HERE.\n    features_dict[str(t)] = FEATURES\n    \n    n_estimators = int(np.median(best_iteration_xgb[str(t)]) + 1)\n    xgb_params = {\n        'objective': 'binary:logistic',\n        'tree_method': 'gpu_hist',\n        'eval_metric': 'logloss',\n        'learning_rate': 0.02,\n        'alpha': 8,\n        'max_depth': 4,\n        'n_estimators': n_estimators,\n        'subsample': 0.8,\n        'colsample_bytree': 0.5\n    }\n    \n    print(f'## QUESTION {t} WITH {len(FEATURES)} FEATURES')\n    \n    # TRAIN DATA\n    train_users = df.index.values\n    train_y = targets.loc[targets['q'] == t].set_index('session').loc[train_users]\n    \n    # TRAIN MODEL\n    clf = XGBClassifier(**xgb_params)\n    clf.fit(df[FEATURES].astype('float32'), train_y['correct'], verbose=0)\n    clf.save_model(f'XGB_question{t}.xgb')","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:35:44.489677Z","iopub.execute_input":"2023-04-12T04:35:44.490111Z","iopub.status.idle":"2023-04-12T04:39:32.030632Z","shell.execute_reply.started":"2023-04-12T04:35:44.490070Z","shell.execute_reply":"2023-04-12T04:39:32.028820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f_save = open('features_dict.pkl', 'wb')\npickle.dump(features_dict, f_save)\nf_save.close()","metadata":{"execution":{"iopub.status.busy":"2023-04-12T04:39:32.032761Z","iopub.execute_input":"2023-04-12T04:39:32.033952Z","iopub.status.idle":"2023-04-12T04:39:32.043368Z","shell.execute_reply.started":"2023-04-12T04:39:32.033909Z","shell.execute_reply":"2023-04-12T04:39:32.042584Z"},"trusted":true},"execution_count":null,"outputs":[]}]}