{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Using GPU - Part 1\n## Feature Engineering and XGBoost Training using GPU\n\nThis is Part 1 of the implementation of Chris's idea to use GPU Kaggle notebook for feature engineering during training and then CPU for inference from the discussion [here][1]. Part 2 Notebook is used to make submission using CPU.\n\nChanges made-\n\n* I modified Chris's XGBoost Baseline [notebook][2] to use RAPIDS cuDF. Feature Engineering gets really fast with the help of RAPIDS cuDF. In the CPU notebook, feature engineering takes about 1 min but utilizing the power of GPU takes the time down to like **3 seconds**!!!\n* Changed the XGBoost `tree_method` from `'hist'` to `'gpu hist'`. This makes training faster too.\n* Added some features to get a better score.\n\nI wasn't able to import cudf in the latest Kaggle Notebook environment (don't know why), so I copied [this][3] notebook and used it's environment instead.\n\nI have written a comment wherever I made a change.\n\n## RAPIDS cuDF\n\ncuDF is a Python GPU DataFrame library which provides a pandas-like API. So just importing it as pd allows us to use cuDF without changing the pandas code much. \n\nBut to use scikit learn functions on cudf dataframes, we have to convert the cudf df to pandas df by calling `to_pandas()`.\n\nTo learn more about how cuDF work, check out cuDF’s documentation [here][4]. \n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218\n[2]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\n[3]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575?scriptVersionId=111214204\n[4]: https://docs.rapids.ai/api/cudf/stable/","metadata":{}},{"cell_type":"markdown","source":"## Version Updates:\n\n**Version 4:**\nChanged the structure of the notebook and added features with the help of [this][1] amazing notebook by @takanashihumbert\n\n[1]: https://www.kaggle.com/code/takanashihumbert/magic-bingo-train-part-lb-0-687","metadata":{}},{"cell_type":"code","source":"import cudf as pd #Change1\nimport numpy as np\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score\nimport matplotlib.pyplot as plt\nfrom tqdm.notebook import tqdm\nfrom collections import defaultdict\nimport warnings\nfrom itertools import combinations\nimport gc\nimport pickle\n\nprint('We will use RAPIDS version',pd.__version__)","metadata":{"papermill":{"duration":3.036143,"end_time":"2022-11-10T16:03:24.014816","exception":false,"start_time":"2022-11-10T16:03:20.978673","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-06T05:44:39.795888Z","iopub.execute_input":"2023-03-06T05:44:39.796380Z","iopub.status.idle":"2023-03-06T05:44:42.952095Z","shell.execute_reply.started":"2023-03-06T05:44:39.796294Z","shell.execute_reply":"2023-03-06T05:44:42.950977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load Train Data and Labels","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv')\nprint( train.shape )\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:44:42.954193Z","iopub.execute_input":"2023-03-06T05:44:42.955229Z","iopub.status.idle":"2023-03-06T05:45:11.385191Z","shell.execute_reply.started":"2023-03-06T05:44:42.955189Z","shell.execute_reply":"2023-03-06T05:45:11.383863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) ) #Change2\ntargets['q'] = pd.to_numeric( targets.session_id.str.split('_q').list.get(1) ) #Change3","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:45:11.386859Z","iopub.execute_input":"2023-03-06T05:45:11.388154Z","iopub.status.idle":"2023-03-06T05:45:11.490561Z","shell.execute_reply.started":"2023-03-06T05:45:11.388102Z","shell.execute_reply":"2023-03-06T05:45:11.489415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(targets.shape)\ntargets.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:45:11.493383Z","iopub.execute_input":"2023-03-06T05:45:11.493748Z","iopub.status.idle":"2023-03-06T05:45:11.518546Z","shell.execute_reply.started":"2023-03-06T05:45:11.493712Z","shell.execute_reply":"2023-03-06T05:45:11.517708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Engineer","metadata":{}},{"cell_type":"code","source":"#Calculate Elapsed Time Difference Column\ntrainp = train.to_pandas()\ntrainp['time_diff'] = (trainp['elapsed_time'] - trainp.groupby(['session_id','level_group'])['elapsed_time'].shift(1)).clip(0,1e7)\ntrain = pd.DataFrame(trainp)\ntrain.loc[train['time_diff']<0] = 0","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:45:11.519699Z","iopub.execute_input":"2023-03-06T05:45:11.520176Z","iopub.status.idle":"2023-03-06T05:45:24.924584Z","shell.execute_reply.started":"2023-03-06T05:45:11.520130Z","shell.execute_reply":"2023-03-06T05:45:24.923619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train1 = train[train[\"level_group\"]=='0-4']\ntrain2 = train[train[\"level_group\"]=='5-12']\ntrain3 = train[train[\"level_group\"]=='13-22']","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:45:24.925811Z","iopub.execute_input":"2023-03-06T05:45:24.926198Z","iopub.status.idle":"2023-03-06T05:45:25.097356Z","shell.execute_reply.started":"2023-03-06T05:45:24.926142Z","shell.execute_reply":"2023-03-06T05:45:25.096041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CATS = ['event_name','name','fqid','room_fqid','text_fqid']\n\nNUMS = ['page', 'room_coor_x','room_coor_y','screen_coor_x','screen_coor_y','hover_duration','time_diff']\n\nEVENTS = ['cutscene_click', 'person_click', 'navigate_click',\n       'observation_click', 'notification_click', 'object_click',\n       'object_hover', 'map_hover', 'map_click', 'checkpoint',\n       'notebook_click']\n\nNAMES = ['basic', 'undefined', 'close', 'open', 'prev', 'next']","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:45:25.098872Z","iopub.execute_input":"2023-03-06T05:45:25.099230Z","iopub.status.idle":"2023-03-06T05:45:25.105763Z","shell.execute_reply.started":"2023-03-06T05:45:25.099195Z","shell.execute_reply":"2023-03-06T05:45:25.104808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_engineer(x, grp):\n    \n    x['elapsed_time'] = x['elapsed_time'] / 1000\n    x['time_diff'] = x['time_diff'] / 1000\n    \n    #session duration\n    df_final = x.groupby('session_id')['index'].agg('count')\n    df_final.name = 'num_events'\n    df_final = df_final.reset_index()\n    df_final = df_final.set_index('session_id')\n    \n    #Bingo Features\n    if grp == '5-12':\n        \n        df_final['logbingo-logbook'] = x[(x['fqid']=='logbook.page.bingo')&(x['event_name']=='object_click')].groupby('session_id')['index'].agg('first') - x[x['fqid']=='logbook'].groupby('session_id')['index'].agg('first')\n        df_final['readerbingo-reader'] = x[(x['fqid']=='reader.paper2.bingo')&(x['event_name']=='object_click')].groupby('session_id')['index'].agg('first') - x[x['fqid']=='reader'].groupby('session_id')['index'].agg('first')\n        df_final['jourbingo-journalspic'] = x[(x['fqid']=='journals.pic_2.bingo')&(x['event_name']=='object_click')].groupby('session_id')['index'].agg('first') - x[x['fqid']=='journals.pic_0.next'].groupby('session_id')['index'].agg('first')\n        \n        df_final['logbingo-logbook_time'] = x[(x['fqid']=='logbook.page.bingo')&(x['event_name']=='object_click')].groupby('session_id')['elapsed_time'].agg('first') - x[x['fqid']=='logbook'].groupby('session_id')['elapsed_time'].agg('first')\n        df_final['readerbingo-reader_time'] = x[(x['fqid']=='reader.paper2.bingo')&(x['event_name']=='object_click')].groupby('session_id')['elapsed_time'].agg('first') - x[x['fqid']=='reader'].groupby('session_id')['elapsed_time'].agg('first')\n        df_final['jourbingo-journalspic_time'] = x[(x['fqid']=='journals.pic_2.bingo')&(x['event_name']=='object_click')].groupby('session_id')['elapsed_time'].agg('first') - x[x['fqid']=='journals.pic_0.next'].groupby('session_id')['elapsed_time'].agg('first')\n        \n    if grp=='13-22':\n        \n        df_final['readerbingo-reader_flag'] = x[(x['fqid']=='reader_flag.paper2.bingo')&(x['event_name']=='object_click')].groupby('session_id')['index'].agg('first') - x[x['fqid']=='reader_flag'].groupby('session_id')['index'].agg('first')\n        df_final['journalbingo-journals_flag'] = x[(x['fqid']=='journals_flag.pic_0.bingo')&(x['event_name']=='object_click')].groupby('session_id')['index'].agg('first') - x[x['fqid']=='journals_flag'].groupby('session_id')['index'].agg('first')\n        \n        df_final['readerbingo-reader_flag_time'] = x[(x['fqid']=='reader_flag.paper2.bingo')&(x['event_name']=='object_click')].groupby('session_id')['elapsed_time'].agg('first') - x[x['fqid']=='reader_flag'].groupby('session_id')['elapsed_time'].agg('first')\n        df_final['journalbingo-journals_flag_time'] = x[(x['fqid']=='journals_flag.pic_0.bingo')&(x['event_name']=='object_click')].groupby('session_id')['elapsed_time'].agg('first') - x[x['fqid']=='journals_flag'].groupby('session_id')['elapsed_time'].agg('first')\n        \n    df_final['first_elapsed_time'] = x.groupby('session_id')['elapsed_time'].agg('first')\n    df_final['elapsed_time'] = x.groupby('session_id')['elapsed_time'].agg('last') - df_final['first_elapsed_time']\n    \n    for c in CATS:\n        df_final[f'{c}_nuniques'] = x.groupby('session_id')[c].agg('nunique')\n    \n    for c in NUMS:\n        df_final[f'{c}_mean'] = x.groupby('session_id')[c].agg('mean')\n        df_final[f'{c}_min'] = x.groupby('session_id')[c].agg('min')\n        df_final[f'{c}_max'] = x.groupby('session_id')[c].agg('max')\n        \n    for c in EVENTS:\n        x[c] = (x.event_name == c).astype('int8')\n    for c in EVENTS:\n        df_final[f'{c}_sum'] = x.groupby('session_id')[c].agg('sum')\n    x.drop(EVENTS, axis=1, inplace=True)\n    \n    for c in EVENTS:\n        df_final[f'{c}_time_mean'] = x[x['event_name']==c].groupby('session_id')['time_diff'].mean()\n        df_final[f'{c}_time_min'] = x[x['event_name']==c].groupby('session_id')['time_diff'].min()\n        df_final[f'{c}_time_max'] = x[x['event_name']==c].groupby('session_id')['time_diff'].max()\n    \n    for c in NAMES:\n        x[c] = (x.name == c).astype('int8')\n    for c in NAMES:\n        df_final[f'{c}_sum'] = x.groupby('session_id')[c].agg('sum')\n    x.drop(NAMES, axis=1, inplace=True)\n    \n    for c in NAMES:\n        df_final[f'{c}_time_mean'] = x[x['name']==c].groupby('session_id')['time_diff'].mean()\n        df_final[f'{c}_time_min'] = x[x['name']==c].groupby('session_id')['time_diff'].min()\n        df_final[f'{c}_time_max'] = x[x['name']==c].groupby('session_id')['time_diff'].max()\n\n    return df_final","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:46:32.281234Z","iopub.execute_input":"2023-03-06T05:46:32.281617Z","iopub.status.idle":"2023-03-06T05:46:32.309467Z","shell.execute_reply.started":"2023-03-06T05:46:32.281585Z","shell.execute_reply":"2023-03-06T05:46:32.307340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf1 = feature_engineer(train1.copy(), grp='0-4')\nprint('df1 done')\ndf2 = feature_engineer(train2.copy(), grp='5-12')\nprint('df2 done')\ndf3 = feature_engineer(train3.copy(), grp='13-22')\nprint('df3 done')","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:46:32.490768Z","iopub.execute_input":"2023-03-06T05:46:32.491210Z","iopub.status.idle":"2023-03-06T05:46:42.989986Z","shell.execute_reply.started":"2023-03-06T05:46:32.491175Z","shell.execute_reply":"2023-03-06T05:46:42.988854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"null1 = df1.isnull().sum().sort_values(ascending=False) / len(df1)\nnull2 = df2.isnull().sum().sort_values(ascending=False) / len(df1)\nnull3 = df3.isnull().sum().sort_values(ascending=False) / len(df1)\n\ndrop1 = list(null1[null1>0.9].index.to_pandas())\ndrop2 = list(null2[null2>0.9].index.to_pandas())\ndrop3 = list(null3[null3>0.9].index.to_pandas())\nprint(len(drop1), len(drop2), len(drop3))\n\nfor col in tqdm(df1.columns):\n    if df1[col].nunique()==1:\n        print(col)\n        drop1.append(col)\nprint(\"*********df1 DONE*********\")\nfor col in tqdm(df2.columns):\n    if df2[col].nunique()==1:\n        print(col)\n        drop2.append(col)\nprint(\"*********df2 DONE*********\")\nfor col in tqdm(df3.columns):\n    if df3[col].nunique()==1:\n        print(col)\n        drop3.append(col)\nprint(\"*********df3 DONE*********\")","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:46:48.499969Z","iopub.execute_input":"2023-03-06T05:46:48.500325Z","iopub.status.idle":"2023-03-06T05:46:48.888820Z","shell.execute_reply.started":"2023-03-06T05:46:48.500296Z","shell.execute_reply":"2023-03-06T05:46:48.887808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train XGBoost Model","metadata":{}},{"cell_type":"code","source":"FEATURES1 = [c for c in df1.columns if c not in drop1+['level_group']]\nFEATURES2 = [c for c in df2.columns if c not in drop2+['level_group','first_index']]\nFEATURES3 = [c for c in df3.columns if c not in drop3+['level_group']]\nprint('We will train with', len(FEATURES1), len(FEATURES2), len(FEATURES3) ,'features')\nALL_USERS = df1.index.unique()\nprint('We will train with', len(ALL_USERS) ,'users info')","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:46:51.060989Z","iopub.execute_input":"2023-03-06T05:46:51.061352Z","iopub.status.idle":"2023-03-06T05:46:51.070392Z","shell.execute_reply.started":"2023-03-06T05:46:51.061321Z","shell.execute_reply":"2023-03-06T05:46:51.069305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gkf = GroupKFold(n_splits=5)\noof = pd.DataFrame(data=np.zeros((len(ALL_USERS), 18)),index=ALL_USERS)","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:46:51.293034Z","iopub.execute_input":"2023-03-06T05:46:51.293371Z","iopub.status.idle":"2023-03-06T05:46:51.843359Z","shell.execute_reply.started":"2023-03-06T05:46:51.293344Z","shell.execute_reply":"2023-03-06T05:46:51.842344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ngkf = GroupKFold(n_splits=5)\noof_xgb = pd.DataFrame(data=np.zeros((len(ALL_USERS),18)), index=ALL_USERS, columns=[f'meta_{i}' for i in range(1, 19)])\n#models = {}\nbest_iteration_xgb = defaultdict(list)\nimportance_dict = {}\n\n# ITERATE THRU QUESTIONS 1 THRU 18\nfor t in range(1,19):\n\n    # USE THIS TRAIN DATA WITH THESE QUESTIONS\n    if t<=3: \n        grp = '0-4'\n        df = df1\n        FEATURES = FEATURES1\n    elif t<=13: \n        grp = '5-12'\n        df = df2\n        FEATURES = FEATURES2\n    elif t<=22: \n        grp = '13-22'\n        df = df3\n        FEATURES = FEATURES3\n        \n    print('#'*25)\n    print('### question', t, 'with features', len(FEATURES))\n    print('#'*25)\n    \n    xgb_params = {\n        'booster': 'gbtree',\n        'objective': 'binary:logistic',\n        'tree_method': 'gpu_hist', #Change4\n        'eval_metric':'logloss',\n        'learning_rate': 0.02,\n        'alpha': 8,\n        'max_depth': 4,\n        'n_estimators': 9999,\n        'early_stopping_rounds': 90,\n        'subsample':0.8,\n        'colsample_bytree': 0.5,\n        'seed': 42\n    }\n\n    feature_importance_df = pd.DataFrame()\n    # COMPUTE CV SCORE WITH 5 GROUP K FOLD\n    P1 = df.iloc[:,0].to_pandas() #Change5\n    P2 = df.index.to_pandas() #Change6\n    for i, (train_index, test_index) in enumerate(gkf.split(X=P1, groups=P2)): #Change7\n        \n        # TRAIN DATA\n        train_x = df.iloc[train_index]\n        train_users = train_x.index.values\n        train_y = targets.loc[targets.q==t].set_index('session').loc[train_users]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index]\n        valid_users = valid_x.index.values\n        valid_y = targets.loc[targets.q==t].set_index('session').loc[valid_users]\n        \n        # TRAIN MODEL        \n        clf =  XGBClassifier(**xgb_params)\n        clf.fit(train_x[FEATURES].astype('float32'), train_y['correct'],\n                eval_set=[(valid_x[FEATURES].astype('float32'), valid_y['correct'])],\n                verbose=0)\n        print(i+1, ', ', end='')\n        best_iteration_xgb[str(t)].append(clf.best_ntree_limit)\n        \n        fold_importance_df = pd.DataFrame()\n        fold_importance_df[\"feature\"] = FEATURES\n        fold_importance_df[\"importance\"] = clf.feature_importances_\n        fold_importance_df[\"fold\"] = i + 1\n        feature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n        \n        # SAVE MODEL, PREDICT VALID OOF\n        oof_xgb.loc[valid_users, f'meta_{t}'] = clf.predict_proba(valid_x[FEATURES].astype('float32'))[:,1]\n            \n    print()\n    feature_importance_df = feature_importance_df.groupby(['feature'])['importance'].agg(['mean']).sort_values(by='mean', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:46:51.845823Z","iopub.execute_input":"2023-03-06T05:46:51.846221Z","iopub.status.idle":"2023-03-06T05:49:00.942506Z","shell.execute_reply.started":"2023-03-06T05:46:51.846185Z","shell.execute_reply":"2023-03-06T05:49:00.941197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Compute CV Score","metadata":{}},{"cell_type":"code","source":"true = oof_xgb.copy()\nfor i in range(1, 19):\n    # GET TRUE LABELS\n    tmp = targets.loc[targets.q==i].set_index('session').loc[ALL_USERS]\n    true[f'meta_{i}'] = tmp.correct.values\n\n# FIND BEST THRESHOLD TO CONVERT PROBS INTO 1s AND 0s\nscores = []; thresholds = []\nbest_score_xgb = 0; best_threshold_xgb = 0\n\nfor threshold in np.arange(0.4,0.81,0.005):\n    print(f'{threshold:.03f}, ',end='')\n    preds = (oof_xgb.to_pandas().values.reshape((-1))>threshold).astype('int') #Change8\n    m = f1_score(true.to_pandas().values.reshape((-1)), preds, average='macro') #Change9\n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score_xgb:\n        best_score_xgb = m\n        best_threshold_xgb = threshold\n\n# PLOT THRESHOLD VS. F1_SCORE\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='blue')\nplt.scatter([best_threshold_xgb], [best_score_xgb], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score_xgb:.5f} at Best Threshold = {best_threshold_xgb:.4}',size=18)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:49:00.944647Z","iopub.execute_input":"2023-03-06T05:49:00.947468Z","iopub.status.idle":"2023-03-06T05:49:08.500698Z","shell.execute_reply.started":"2023-03-06T05:49:00.947387Z","shell.execute_reply":"2023-03-06T05:49:08.499744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('When using optimal threshold...')\nfor k in range(18):\n    m = f1_score(true[f'meta_{k+1}'].to_pandas().values, (oof_xgb[f'meta_{k+1}'].to_pandas().values>best_threshold_xgb).astype('int'), average='macro') #Change10\n    print(f'Q{k}: F1 =',m)\n    \nm = f1_score(true.to_pandas().values.reshape((-1)), (oof_xgb.to_pandas().values.reshape((-1))>best_threshold_xgb).astype('int'), average='macro') #Change11\nprint('==> Overall F1 =',m)","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:49:08.502366Z","iopub.execute_input":"2023-03-06T05:49:08.503458Z","iopub.status.idle":"2023-03-06T05:49:08.696799Z","shell.execute_reply.started":"2023-03-06T05:49:08.503412Z","shell.execute_reply":"2023-03-06T05:49:08.695496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# ITERATE THRU QUESTIONS 1 THRU 18\nfor t in range(1,19):\n\n    # USE THIS TRAIN DATA WITH THESE QUESTIONS\n    if t<=3: \n        grp = '0-4'\n        df = df1\n        FEATURES = FEATURES1\n    elif t<=13: \n        grp = '5-12'\n        df = df2\n        FEATURES = FEATURES2\n    elif t<=22: \n        grp = '13-22'\n        df = df3\n        FEATURES = FEATURES3\n    \n    n_estimators = int(np.median(best_iteration_xgb[str(t)]) + 1)\n    xgb_params = {\n        'objective': 'binary:logistic',\n        'tree_method': 'gpu_hist',\n        'eval_metric':'logloss',\n        'learning_rate': 0.02,\n        'alpha': 8,\n        'max_depth': 4,\n        'n_estimators': n_estimators,\n        'subsample':0.8,\n        'colsample_bytree': 0.5,\n    }\n    \n    print('#'*25)\n    print(f'### question {t} features {len(FEATURES)}')\n        \n    # TRAIN DATA\n    train_users = df.index.values\n    train_y = targets.loc[targets.q==t].set_index('session').loc[train_users]\n\n    # TRAIN MODEL        \n    clf =  XGBClassifier(**xgb_params)\n    clf.fit(df[FEATURES].astype('float32'), train_y['correct'], verbose=0)\n    clf.save_model(f'XGB_question{t}.xgb')\n    \n    print()","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:45:25.397066Z","iopub.status.idle":"2023-03-06T05:45:25.397980Z","shell.execute_reply.started":"2023-03-06T05:45:25.397698Z","shell.execute_reply":"2023-03-06T05:45:25.397722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importance_dict = {}\nfor t in range(1, 19):\n    if t<=3: \n        importance_dict[str(t)] = FEATURES1\n    elif t<=13: \n        importance_dict[str(t)] = FEATURES2\n    elif t<=22:\n        importance_dict[str(t)] = FEATURES3\n\nf_save = open('importance_dict.pkl', 'wb')\npickle.dump(importance_dict, f_save)\nf_save.close()","metadata":{"execution":{"iopub.status.busy":"2023-03-06T05:45:25.399509Z","iopub.status.idle":"2023-03-06T05:45:25.400664Z","shell.execute_reply.started":"2023-03-06T05:45:25.400345Z","shell.execute_reply":"2023-03-06T05:45:25.400368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *Check the Part 2 - Inference Notebook to make a submission!!*","metadata":{}}]}