{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# XGBoost Baseline - LB 0.678\nIn this notebook we present a XGBoost baseline. We train GroupKFold models for each of the 18 questions. Our CV score is 0.678. We infer test using one of our KFold models. We can improve our CV and LB by engineering more features for our xgboost and/or trying different models (like other ML models and/or RNN and/or Transformer). Also we can improve our LB by using more KFold models OR training one model using all data (and the hyperparameters that we found from our KFold cross validation).\n\n**UPDATE** On March 20 2023, Kaggle doubled the size of train data. Therefore we updated this notebook to avoid memory error. We accomplish this by reading train data in chunks and feature engineering in chunks. Note that another way to avoid memory error is to use two notebooks. Train models in one notebook that has 32GB RAM (and save models), and then submit the required 8GB RAM notebook (with loaded models) as a second notebook. (Discussion [here][1]).\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218","metadata":{"papermill":{"duration":0.005932,"end_time":"2023-02-07T00:59:58.147501","exception":false,"start_time":"2023-02-07T00:59:58.141569","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pandas as pd, numpy as np, gc\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score","metadata":{"papermill":{"duration":1.027875,"end_time":"2023-02-07T00:59:59.180261","exception":false,"start_time":"2023-02-07T00:59:58.152386","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T18:24:53.909770Z","iopub.execute_input":"2023-06-07T18:24:53.910109Z","iopub.status.idle":"2023-06-07T18:24:54.971313Z","shell.execute_reply.started":"2023-06-07T18:24:53.910015Z","shell.execute_reply":"2023-06-07T18:24:54.970515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Train Data and Labels\nOn March 20 2023, Kaggle doubled the size of train data (discussion [here][1]). The train data is now 4.7GB! To avoid memory error, we will read the train data in as 10 pieces and feature engineer each piece before reading the next piece. This works because feature engineering shrinks the size of each piece.\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202","metadata":{"papermill":{"duration":0.004542,"end_time":"2023-02-07T00:59:59.189777","exception":false,"start_time":"2023-02-07T00:59:59.185235","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# READ USER ID ONLY\ntmp = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\",usecols=[0])\ntmp = tmp.groupby('session_id').session_id.agg('count')\n\n# COMPUTE READS AND SKIPS\nPIECES = 12\nCHUNK = int( np.ceil(len(tmp)/PIECES) )\n\nreads = []\nskips = [0]\nfor k in range(PIECES):\n    a = k*CHUNK\n    b = (k+1)*CHUNK\n    if b>len(tmp): b=len(tmp)\n    r = tmp.iloc[a:b].sum()\n    reads.append(r)\n    skips.append(skips[-1]+r)\n    \n    \nprint(f'To avoid memory error, we will read train in {PIECES} pieces of sizes:')\nprint(reads)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-06-07T18:24:54.975945Z","iopub.execute_input":"2023-06-07T18:24:54.977755Z","iopub.status.idle":"2023-06-07T18:26:09.976999Z","shell.execute_reply.started":"2023-06-07T18:24:54.977726Z","shell.execute_reply":"2023-06-07T18:26:09.976215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dtypes={\n    'elapsed_time':np.int32,\n    'event_name':'category',\n    'name':'category',\n    'level':np.uint8,\n    'room_coor_x':np.float32,\n    'room_coor_y':np.float32,\n    'screen_coor_x':np.float32,\n    'screen_coor_y':np.float32,\n    'hover_duration':np.float32,\n    'text':'category',\n    'fqid':'category',\n    'room_fqid':'category',\n    'text_fqid':'category',\n    'fullscreen':'category',\n    'hq':'category',\n    'music':'category',\n    'level_group':'category'}\n\ntrain = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', nrows=reads[0], dtype=dtypes)\nprint('Train size of first piece:', train.shape )\ntrain.head(3)","metadata":{"papermill":{"duration":59.284316,"end_time":"2023-02-07T01:00:58.478743","exception":false,"start_time":"2023-02-07T00:59:59.194427","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T18:26:09.978258Z","iopub.execute_input":"2023-06-07T18:26:09.978830Z","iopub.status.idle":"2023-06-07T18:26:15.436202Z","shell.execute_reply.started":"2023-06-07T18:26:09.978800Z","shell.execute_reply":"2023-06-07T18:26:15.435379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )\ntargets['q'] = targets.session_id.apply(lambda x: int(x.split('_')[-1][1:]))\ntargets['q_gr'] = targets.session_id.apply(lambda x: x.split('_')[-1])\nprint( targets.shape )\ntargets['full_correct'] = targets.groupby('session')['correct'].transform('sum')\ntargets.head()","metadata":{"papermill":{"duration":0.598155,"end_time":"2023-02-07T01:00:59.082015","exception":false,"start_time":"2023-02-07T01:00:58.48386","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T18:26:15.438422Z","iopub.execute_input":"2023-06-07T18:26:15.438903Z","iopub.status.idle":"2023-06-07T18:26:16.988835Z","shell.execute_reply.started":"2023-06-07T18:26:15.438876Z","shell.execute_reply":"2023-06-07T18:26:16.987980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Переведем в wide формат\ntarget_cols = ['session'] +[f'q{i}' for i in range(0, 18)]\nwide_targets = targets[['correct', 'session', 'q_gr']].pivot(values='correct',\n                                                          index=['session'],\n                                                          columns=['q_gr']).reset_index()#[target_cols]\n\nwide_targets.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-06-07T18:26:16.989975Z","iopub.execute_input":"2023-06-07T18:26:16.990340Z","iopub.status.idle":"2023-06-07T18:26:17.182004Z","shell.execute_reply.started":"2023-06-07T18:26:16.990312Z","shell.execute_reply":"2023-06-07T18:26:17.181281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineer\nWe create basic aggregate features. Try creating more features to boost CV and LB! The idea for EVENTS feature is from [here][1]\n\n[1]: https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data","metadata":{"papermill":{"duration":0.005196,"end_time":"2023-02-07T01:00:59.092865","exception":false,"start_time":"2023-02-07T01:00:59.087669","status":"completed"},"tags":[]}},{"cell_type":"code","source":"CATS = ['event_name', 'fqid', 'room_fqid', 'text']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\nSPECIFIC = ['fullscreen', 'hq', 'music']\n# https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']\n\ngrs2qus = {\n    '0-4':   [f'q{i}' for i in range(1, 4)],\n    '5-12':  [f'q{i}' for i in range(4, 14)],\n    '13-22': [f'q{i}' for i in range(14, 19)]\n          }\ngrs2qus","metadata":{"papermill":{"duration":0.014685,"end_time":"2023-02-07T01:00:59.112856","exception":false,"start_time":"2023-02-07T01:00:59.098171","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T18:26:17.182987Z","iopub.execute_input":"2023-06-07T18:26:17.183658Z","iopub.status.idle":"2023-06-07T18:26:17.193498Z","shell.execute_reply.started":"2023-06-07T18:26:17.183630Z","shell.execute_reply":"2023-06-07T18:26:17.192842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":0.017716,"end_time":"2023-02-07T01:00:59.136021","exception":false,"start_time":"2023-02-07T01:00:59.118305","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" ","metadata":{"papermill":{"duration":0.017716,"end_time":"2023-02-07T01:00:59.136021","exception":false,"start_time":"2023-02-07T01:00:59.118305","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_engineer(train):\n    \n    f = lambda x: x.diff()\n    train = train.sort_values(['session_id', 'level_group', 'index'])\n    train['diff_index'] = train.groupby(['session_id', 'level_group'])['index'].transform(f)\n    train['diff_time'] = train.groupby(['session_id', 'level_group'])['elapsed_time'].transform(f)\n    ##################################\n    target_col = 'session_id'\n    dfs = train[['session_id', 'level_group']].drop_duplicates().copy()\n    for col in ['fqid', 'event_name', 'name', 'text_fqid', 'room_fqid', 'page', 'level']: \n        dfs = (\n            train[['session_id', col, 'index', 'level_group']]\n            .pivot_table(values=target_col, index=['session_id', 'level_group'], aggfunc = 'count', columns=[col])\n            .reset_index()\n            .merge(dfs, on = ['session_id', 'level_group'], how = 'right')\n            .sort_values(['session_id', 'level_group'])\n        )\n    dfs.rename(columns = {c:f'{target_col}_count_{c}' for c in dfs.columns if c not in ['session_id', 'level_group']}, inplace=True)\n    all_dfs = dfs.copy().drop(fs2drop_dynam['fs_list'], axis=1, errors = 'ignore')\n    \n    #################################\n    target_col = 'diff_time'\n    dfs = train[['session_id', 'level_group']].drop_duplicates().copy()\n    for col in ['fqid', 'event_name', 'name', 'text_fqid', 'room_fqid', 'page', 'level']: \n        dfs = (\n            train[['session_id', col, 'index', 'level_group', target_col]]\n            .pivot_table(values=target_col,index=['session_id', 'level_group'], aggfunc = 'median', columns=[col])\n            .reset_index()\n            .merge(dfs, on = ['session_id', 'level_group'], how = 'right')\n            .sort_values(['session_id', 'level_group']))\n    dfs.rename(columns = {c:f'{target_col}_median_{c}' for c in dfs.columns if c not in ['session_id', 'level_group']}, inplace=True)\n    all_dfs_2 = dfs.copy().drop(fs2drop_dynam['fs_list'], axis=1, errors = 'ignore')\n    ##################################\n    target_col = 'hover_duration' \n    dfs = train[['session_id', 'level_group', ]].drop_duplicates().copy()\n    for col in ['fqid', 'event_name', 'name', 'text_fqid', 'room_fqid', 'page', 'level']: \n        dfs = (\n            train[['session_id', col, 'index', 'level_group', target_col]]\n            .pivot_table(values=target_col,index=['session_id', 'level_group'], aggfunc = 'median', columns=[col])\n            .reset_index()\n            .merge(dfs, on = ['session_id', 'level_group'], how = 'right')\n            .sort_values(['session_id', 'level_group']))\n    dfs.rename(columns = {c:f'{c}_{target_col}_median' for c in dfs.columns if c not in ['session_id', 'level_group']}, inplace=True)\n    all_dfs_3 = dfs.copy().drop(fs2drop_dynam['fs_list'], axis=1, errors = 'ignore')\n    ##################################\n#     target_col = 'level'\n#     dfs = train[['session_id', 'level_group']].drop_duplicates().copy()\n#     for col in ['fqid', 'event_name', 'name', 'text_fqid', 'room_fqid', 'page']: \n#         dfs = (\n#             train[['session_id', col, 'index', 'level_group', target_col]]\n#             .pivot_table(values=target_col, index=['session_id', 'level_group'], aggfunc = 'mean', columns=[col])\n#             .reset_index()\n#             .merge(dfs, on = ['session_id', 'level_group'], how = 'right')\n#             .sort_values(['session_id', 'level_group']))\n#     dfs.rename(columns = {c:f'{c}_{target_col}_mean' for c in dfs.columns if c not in ['session_id', 'level_group']}, inplace=True)\n#     all_dfs_4 = dfs.copy().drop(fs2drop_dynam['fs_list'], axis=1, errors = 'ignore')\n    ##################################\n#     target_col = 'level'\n#     dfs = train[['session_id', 'level_group']].drop_duplicates().copy()\n#     for col in ['fqid', 'event_name', 'name', 'text_fqid', 'room_fqid', 'page']: \n#         dfs = (\n#             train[['session_id', col, 'index', 'level_group', target_col]]\n#             .pivot_table(values=target_col, index=['session_id', 'level_group'], aggfunc = 'std', columns=[col])\n#             .reset_index()\n#             .merge(dfs, on = ['session_id', 'level_group'], how = 'right')\n#             .sort_values(['session_id', 'level_group']))\n#     dfs.rename(columns = {c:f'{c}_{target_col}_std' for c in dfs.columns if c not in ['session_id', 'level_group']}, inplace=True)\n#     all_dfs_5 = dfs.copy().drop(fs2drop_dynam['fs_list'], axis=1, errors = 'ignore')\n    ##################################\n    target_col = 'diff_time'\n    dfs = train[['session_id', 'level_group']].drop_duplicates().copy()\n    for col in ['fqid', 'event_name', 'name', 'text_fqid', 'room_fqid', 'page']: \n        dfs = (\n            train[['session_id', col, 'index', 'level_group', target_col]]\n            .pivot_table(values=target_col, index=['session_id', 'level_group'], aggfunc = 'max', columns=[col])\n            .reset_index()\n            .merge(dfs, on = ['session_id', 'level_group'], how = 'right')\n            .sort_values(['session_id', 'level_group']))\n    dfs.rename(columns = {c:f'{c}_{target_col}_std' for c in dfs.columns if c not in ['session_id', 'level_group']}, inplace=True)\n    all_dfs_6 = dfs.copy().drop(fs2drop_dynam['fs_list'], axis=1, errors = 'ignore')\n    ##################################\n    \n    def score_game(row):\n        text_visited = []; score = 0\n        for text in row:\n            if text not in ['undefined']:\n                if text in text_visited:score -=1\n                else:\n                    score+=1\n                    text_visited.append(text)\n        return score\n    \n    ##################################\n    \n    def score_time(row):\n        score = 0\n        for time in row:\n            if time > 3200: score -=1\n            else: score+=1\n        return score\n    \n    def score_time_short(row):\n        score = 0\n        for time in row:\n            if time > 50: score -=1\n            else: score+=1\n        return score\n    \n    def score_time_600(row):\n        score = 0\n        for time in row:\n            if time > 600: score -=1\n            else: score+=1\n        return score\n\n    ############################\n    def score_event_name(row):\n        text_visited = []; score = 0\n        for text in row:\n            if text in text_visited: score -=1\n            else:\n                score+=1\n                text_visited.append(text)\n                if len(text_visited) > 20: text_visited = text_visited[-20:]\n        return score\n    ############################\n    def score_event_name_short(row):\n        text_visited = []; score = 0\n        for text in row:\n            if text in text_visited: score -=1\n            else:\n                score+=1\n                text_visited.append(text)\n                if len(text_visited) > 10: text_visited = text_visited[-10:]\n        return score\n    ###################\n    \n    def score_hover(row):\n        score = 0\n        for time in row:\n            if abs(time) > 500: score -=1\n            else: score+=1\n        return score\n    \n    def smart_hower_score(row):\n        score = 0\n        row = row[-100:] # Последняя сотня\n        for i, time in enumerate(row):\n            if abs(time) > 300:score -=1*(i+1)\n            else:score+=1\n        return score\n    \n    def smart_score(row):\n        score = 0\n        row = row[-100:] # Последняя сотня\n        for i, time in enumerate(row):\n            if abs(time) > 500:score -=1*(i+1)\n            else:score+=1\n        return score\n    \n    def smart_argmax_score(row):\n        row = row[-100:] # Последняя сотня\n        return np.argmax(row)\n    \n    def random_choice_score(row):\n        row = row[-100:] # Последняя сотня\n        return np.argmax(row)\n    \n    def diff_index_score(row):\n        score = 0\n        for time in row:\n            if abs(time) > 1: score +=1\n        return score\n    \n    def smart_diff_index_score(row):\n        score = 0\n        row = row[-50:] # Последняя сотня\n        for i, time in enumerate(row):\n            if abs(time) > 1:score +=1*(i+1)\n        return score\n    \n    # ['fqid', 'event_name', 'name', 'text_fqid', 'room_fqid', 'page', 'level']: \n    arg_max_f = lambda x: np.argmax(x)\n    tmp_scores = train.groupby(['session_id', 'level_group'], as_index=False).agg(\n        \n       text_score =('text', score_game),\n       event_name_score =('event_name', score_game),\n       text_fqid_score =('text_fqid', score_game),\n       room_fqid_score =('room_fqid', score_game),    \n       event_name_score_10 =('event_name', score_event_name_short),\n       room_fqid_score_10 =('room_fqid', score_event_name_short),\n\n       text_fqid_score_20 =('text_fqid', score_event_name),\n       event_name_score_20 =('event_name', score_event_name),\n       room_fqid_short_score =('room_fqid', score_event_name_short),\n\n       hover_score =('hover_duration', score_hover),\n       smart_hower_score_score =('hover_duration', smart_hower_score),\n\n       diff_time_score =('diff_time', score_time),\n       diff_time_score_50 =('diff_time', score_time_short),\n       diff_time_score_600 =('diff_time', score_time_600),  \n\n       diff_index_score =('diff_index', diff_index_score),  \n       smart_diff_index_score_score =('diff_index', smart_diff_index_score),  \n       diff_index_random_choice_score =('diff_index', random_choice_score),  \n\n       diff_time_arg_max_score =('diff_time', arg_max_f),\n       diff_time_smart_argmax_score =('diff_time', smart_argmax_score),\n       diff_time_smart_score =('diff_time', smart_score),\n       length =('index', 'count')\n    )\n  \n    ######################################## \n    \n    \n    dfs = []\n    \n    # train = train.sort_values(['session_id', 'level_group', 'index'])\n    f = lambda x: x.diff()\n    train['diff_time'] = train.groupby(['session_id', 'level_group'])['elapsed_time'].transform(f)\n    \n    for c in CATS + ['page', 'diff_index']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('nunique')\n        tmp.name = tmp.name + '_nunique_score'\n        dfs.append(tmp)\n    for c in CATS + ['page']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('count')\n        tmp.name = tmp.name + '_count_score'\n        dfs.append(tmp)\n    for c in NUMS + ['diff_time', 'diff_index']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('mean')\n        tmp.name = tmp.name + '_mean'\n        dfs.append(tmp.fillna(0))\n    for c in NUMS +  ['diff_time']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('std')\n        tmp.name = tmp.name + '_std'\n        dfs.append(tmp.fillna(0))\n    for c in EVENTS: \n        train[c] = (train.event_name == c).astype('int8')\n    for c in EVENTS + ['elapsed_time']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('sum')\n        tmp.name = tmp.name + '_sum_score'\n        dfs.append(tmp)    \n    for c in ['index', 'diff_time', 'diff_index']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('max')\n        tmp.name = tmp.name + '_max'\n        dfs.append(tmp)  \n        \n    train = train.drop(EVENTS,axis=1)\n        \n        \n    df = pd.concat(dfs, axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    \n    df = df.merge(all_dfs, on=['session_id','level_group'], how = 'left')\n    df = df.merge(all_dfs_2, on=['session_id','level_group'], how = 'left')\n    df = df.merge(all_dfs_3, on=['session_id','level_group'], how = 'left')\n#     df = df.merge(all_dfs_4, on=['session_id','level_group'], how = 'left')\n#     df = df.merge(all_dfs_5, on=['session_id','level_group'], how = 'left')\n    df = df.merge(all_dfs_6, on=['session_id','level_group'], how = 'left')\n    # df = df.merge(all_dfs_7, on=['session_id','level_group'], how = 'left')\n    df = df.merge(tmp_scores, on=['session_id','level_group'], how = 'left')\n    \n    not2norm = ['diff_time_smart_argmax_score', 'diff_time_smart_score']\n    not2norm += [ c for c in df.columns if 'score_10' in c or 'score_20' in c ]\n    cols2norm = [col for col in df.columns if 'score' in col and col not in not2norm]\n    for c in cols2norm:\n        df[c] = df[c]/df['length'] \n    \n    df = df.set_index('session_id').drop(fs2drop_dynam['fs_list'], axis=1, errors = 'ignore')\n    \n    \n    return df","metadata":{"papermill":{"duration":0.017716,"end_time":"2023-02-07T01:00:59.136021","exception":false,"start_time":"2023-02-07T01:00:59.118305","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T18:26:17.194911Z","iopub.execute_input":"2023-06-07T18:26:17.195509Z","iopub.status.idle":"2023-06-07T18:26:17.272369Z","shell.execute_reply.started":"2023-06-07T18:26:17.195481Z","shell.execute_reply":"2023-06-07T18:26:17.271555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fs_zero_importance_2drop = ['diff_time_median_tunic.historicalsociety.collection_flag',\n     'diff_time_median_tunic.historicalsociety.collection_flag',\n     'session_id_count_retirement_letter',\n     'session_id_count_tunic.historicalsociety.closet_dirty.gramps.helpclean',\n     'tunic.historicalsociety.basement.savedteddy_level_mean',\n     'tunic.historicalsociety.stacks.journals_flag.pic_2.bingo_level_mean',\n     'session_id_count_tunic.historicalsociety.frontdesk.archivist.newspaper_recap',\n     'session_id_count_tunic.historicalsociety.closet.gramps.intro_0_cs_0',\n     'tunic.wildlife.center.expert.removed_cup_level_mean',\n     'reader_flag.paper0.prev_diff_time_std',\n     'report_diff_time_std',\n     'tunic.historicalsociety.collection_flag.gramps.flag_level_mean',\n     'tunic.historicalsociety.closet_dirty.what_happened_diff_time_std',\n     'tunic.wildlife.center.remove_cup_level_std',\n     'session_id_count_tunic.historicalsociety.closet.retirement_letter.hub',\n     'tunic.wildlife.center.wells.animals_level_mean',\n     'savedteddy_level_std',\n     'savedteddy_level_mean',\n     'session_id_count_tunic.historicalsociety.collection.gramps.look_0',\n     'seescratches_level_mean',\n     'room_fqid_count_score',\n     'event_name_count_score',\n     'tunic.wildlife.center.tracks.hub.deer_level_mean',\n     'tunic.historicalsociety.collection.cs_level_mean',\n     'tunic.historicalsociety.collection.gramps.lost_level_std',\n     'trigger_scarf_level_mean',\n     'chap4_finale_c_level_mean',\n     'what_happened_level_std',\n     'tunic.flaghouse.entry.colorbook_level_mean',\n     'tunic.historicalsociety.entry.groupconvo_flag_level_mean',\n     'session_id_count_tunic.historicalsociety.entry.boss.flag_recap',\n     'session_id_count_tunic.kohlcenter.halloffame.togrampa',\n     'ch3start_level_std',\n     'seescratches_level_std',\n     'tunic.historicalsociety.basement.seescratches_level_std',\n     'tunic.historicalsociety.entry.groupconvo_flag_level_std',\n     'chap4_finale_c_level_std',\n     'chap1_finale_c_level_mean',\n     'trigger_scarf_level_std',\n     'tunic.historicalsociety.closet_dirty.trigger_coffee_level_std',\n     'tunic.historicalsociety.closet_dirty.what_happened_level_std',\n     'what_happened_level_mean',\n     'tunic.drycleaner.frontdesk.worker.done_level_std',\n     'tunic.historicalsociety.closet_dirty.trigger_scarf_level_mean',\n     'tunic.historicalsociety.closet_dirty.what_happened_level_mean',\n     'wellsbadge_level_std',\n     'magnify_level_mean',\n     'chap1_finale_level_mean',\n     'wellsbadge_level_mean',\n     'tunic.drycleaner.frontdesk.worker.done_level_mean',\n     'togrampa_level_mean',\n     'groupconvo_level_mean']\n                            \nfs2drop_dynam = {}\nfs2drop_dynam['fs_list'] = fs_zero_importance_2drop","metadata":{"execution":{"iopub.status.busy":"2023-06-07T18:26:17.274346Z","iopub.execute_input":"2023-06-07T18:26:17.274929Z","iopub.status.idle":"2023-06-07T18:26:17.287162Z","shell.execute_reply.started":"2023-06-07T18:26:17.274899Z","shell.execute_reply":"2023-06-07T18:26:17.286286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n    \ndef reduce_mem_usage(df):\n    \"\"\" iterate through all the columns of a dataframe and modify the data type\n        to reduce memory usage.\n    \"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n\n    for col in df.columns:\n        col_type = df[col].dtype.name\n\n        if col_type not in ['object', 'category', 'datetime64[ns, UTC]']:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n\n    return df\n\n# PROCESS TRAIN DATA IN PIECES\nall_pieces = []\nprint(f'Processing train as {PIECES} pieces to avoid memory error... ')\nfor k in range(PIECES):\n    print(k,', ',end='')\n    SKIPS = 0\n    if k>0: SKIPS = range(1,skips[k]+1)\n    path = '/kaggle/input/predict-student-performance-from-game-play/train.csv'\n    train = pd.read_csv(path, nrows=reads[k], skiprows=SKIPS, dtype=dtypes)\n    df = feature_engineer(train)\n    df.drop(fs_zero_importance_2drop, axis=1, inplace=True, errors = 'ignore')\n    df = reduce_mem_usage(df)\n    all_pieces.append(df)\n    \n# CONCATENATE ALL PIECES\nprint('\\n')\ndel train; gc.collect()\ndf = pd.concat(all_pieces, axis=0)\nprint('Shape of all train data after feature engineering:', df.shape )\ndf.head()","metadata":{"papermill":{"duration":34.516494,"end_time":"2023-02-07T01:01:33.658043","exception":false,"start_time":"2023-02-07T01:00:59.141549","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T18:26:17.288394Z","iopub.execute_input":"2023-06-07T18:26:17.288735Z","iopub.status.idle":"2023-06-07T19:15:00.123708Z","shell.execute_reply.started":"2023-06-07T18:26:17.288708Z","shell.execute_reply":"2023-06-07T19:15:00.120986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import joblib\n\nfeatures2drop4grp = {}\n\nfor grp in ['0-4', '5-12', '13-22']:\n    tmp = df[df.level_group == grp]\n    t = tmp.isnull().sum() / tmp.shape[0]\n    \n    features2drop4grp[grp] = []\n    for c in tmp.columns:\n        if t[c] == 1:\n            # print(c, t[c])\n            features2drop4grp[grp].append(c)\n            \n    print('grp', len(features2drop4grp[grp]))\n    \njoblib.dump(features2drop4grp, f\"./features2drop4grp.pkl\")\nfeatures2drop4grp = joblib.load(f\"./features2drop4grp.pkl\")","metadata":{"execution":{"iopub.status.busy":"2023-06-07T19:15:00.132962Z","iopub.execute_input":"2023-06-07T19:15:00.134390Z","iopub.status.idle":"2023-06-07T19:15:04.629416Z","shell.execute_reply.started":"2023-06-07T19:15:00.134351Z","shell.execute_reply":"2023-06-07T19:15:04.628349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train XGBoost Model\nWe train one model for each of 18 questions. Furthermore, we use data from `level_groups = '0-4'` to train model for questions 1-3, and `level groups '5-12'` to train questions 4 thru 13 and `level groups '13-22'` to train questions 14 thru 18. Because this is the data we get (to predict corresponding questions) from Kaggle's inference API during test inference. We can improve our model by saving a user's previous data from earlier `level_groups` and using that to predict future `level_groups`.","metadata":{"papermill":{"duration":0.00565,"end_time":"2023-02-07T01:01:33.669525","exception":false,"start_time":"2023-02-07T01:01:33.663875","status":"completed"},"tags":[]}},{"cell_type":"code","source":" ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fs2drop4 = ['observation_click_sum_score','map_click_sum_score',\n            'session_id_count_0_x','session_id_count_3_x',\n            'room_coor_y_std','session_id_count_4_x','page_count_score',\n            'diff_index_nunique_score','room_fqid_nunique_score']\n\nfs2drop8 = ['session_id_count_8','observation_click_sum_score','map_click_sum_score',\n            'session_id_count_0_x','session_id_count_3_x','room_coor_y_std','session_id_count_4_x',\n            'session_id_count_7','page_count_score','diff_index_nunique_score','room_fqid_nunique_score',\n            'session_id_count_1_x','session_id_count_2_x','index_max','session_id_count_13',\n            'screen_coor_y_std','page_std','session_id_count_6_x','level_mean','elapsed_time_std']\n\nfeature2drop = fs2drop4 #+ fs2drop8","metadata":{"execution":{"iopub.status.busy":"2023-06-07T19:15:04.630934Z","iopub.execute_input":"2023-06-07T19:15:04.631256Z","iopub.status.idle":"2023-06-07T19:15:04.639739Z","shell.execute_reply.started":"2023-06-07T19:15:04.631227Z","shell.execute_reply":"2023-06-07T19:15:04.637092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filtered_features = [c for c in df.columns if c != 'level_group' and c not in fs_zero_importance_2drop]# and c not in feature2drop]\n\njoblib.dump(filtered_features, f\"./filtered_features.pkl\")\nfiltered_features = joblib.load(f\"./filtered_features.pkl\")\n\nprint('We will train with', len(filtered_features) ,'features')\nALL_USERS = df.index.unique()\nprint('We will train with', len(ALL_USERS) ,'users info')","metadata":{"papermill":{"duration":0.014699,"end_time":"2023-02-07T01:01:33.689953","exception":false,"start_time":"2023-02-07T01:01:33.675254","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T19:15:04.642501Z","iopub.execute_input":"2023-06-07T19:15:04.642841Z","iopub.status.idle":"2023-06-07T19:15:04.729892Z","shell.execute_reply.started":"2023-06-07T19:15:04.642807Z","shell.execute_reply":"2023-06-07T19:15:04.728861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from catboost import CatBoostClassifier, CatBoostRegressor, Pool\nfrom sklearn.model_selection import KFold  # k-фолдная валидация","metadata":{"execution":{"iopub.status.busy":"2023-06-07T19:15:04.732350Z","iopub.execute_input":"2023-06-07T19:15:04.732702Z","iopub.status.idle":"2023-06-07T19:15:05.141695Z","shell.execute_reply.started":"2023-06-07T19:15:04.732654Z","shell.execute_reply":"2023-06-07T19:15:05.140849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gkf = GroupKFold(n_splits=3)\noof = pd.DataFrame(data=np.zeros((len(ALL_USERS), 18)), index=ALL_USERS)\nmodels = {}; scores = []\n# COMPUTE CV SCORE WITH 5 GROUP K FOLD\nfor i, (train_index, test_index) in enumerate(gkf.split(X=df, groups=df.index)):\n    print('#'*25)\n    print('### Fold',i+1)\n    print('#'*25)\n    \n    cb_params = {\n        'depth': 6,\n        'iterations': 100_000,\n        'learning_rate': 0.029999999329447743,\n        'loss_function': \"MultiLogloss\",  # MultiClass\n        # 'eval_metric' :['F1:macro'],# 'Precision',  F1:macro / AUC:hints=skip_train~false\n        'custom_metric':[\"F1\"],  # 'AUC / Accuracy,\n        'cat_features': None,\n        \n        # Регуляризация и ускорение\n#         'colsample_bylevel':0.1/2,\n        'subsample': 0.9,  # 0.9 / 0.5\n        'l2_leaf_reg': 22,\n        'min_data_in_leaf': 250,\n        'max_bin': 170,\n        'random_strength':1,\n        \n        # 'max_leaves' : 100,\n        \n        'task_type':\"CPU\",    \n        'thread_count': 2,\n        'bootstrap_type':\"Bernoulli\", \n        \n        'random_seed':7575,\n        # 'auto_class_weights':\"SqrtBalanced\",\n        'early_stopping_rounds': 100\n    }\n    \n    \n    # ITERATE THRU QUESTIONS 1 THRU 18\n    # for t in range(1,19):\n    limits = {'0-4':(1, 4), '5-12':(4,14), '13-22':(14,19)}\n    for grp in ['0-4', '5-12', '13-22']:\n        \n         \n        # USE THIS TRAIN DATA WITH THESE QUESTIONS\n        # if t<=3: grp = '0-4'\n        # elif t<=13: grp = '5-12'\n        # elif t<=22: grp = '13-22'\n        qs = grs2qus[grp]\n            \n        # TRAIN DATA\n        train_x = df.iloc[train_index].copy()\n        train_x = train_x.loc[train_x.level_group == grp]\n        train_x.drop(features2drop4grp[grp], inplace=True, axis=1)\n        train_users = train_x.index.values\n        train_y = wide_targets.set_index('session').loc[train_users][qs]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index].copy()\n        valid_x = valid_x.loc[valid_x.level_group == grp]\n        valid_x.drop(features2drop4grp[grp], inplace=True, axis=1)\n        valid_users = valid_x.index.values\n        valid_y = wide_targets.set_index('session').loc[valid_users][qs]\n        \n        filtered_features_ = [c for c in filtered_features if c not in features2drop4grp[grp]]\n        # print(len(filtered_features_))\n        train_dataset = Pool(data=train_x[filtered_features_], label=train_y)#, cat_features=cat_features)\n        eval_dataset = Pool(data=valid_x[filtered_features_], label=valid_y)#, cat_features=cat_features)\n        \n        # TRAIN MODEL        \n        clf = CatBoostClassifier(**cb_params)\n        \n        # clf.fit(train_x[FEATURES].astype('float32'), train_y['correct'],\n                # eval_set=[ (valid_x[FEATURES].astype('float32'), valid_y['correct']) ],\n                # verbose=0)\n        \n        clf.fit(\n            train_dataset,\n            eval_set=eval_dataset,\n            verbose=False,\n            use_best_model=True,\n            plot=False)\n        \n        # print(f'{t}({clf.best_ntree_limit}), ',end='')\n        \n        # SAVE MODEL, PREDICT VALID OOF\n        models[f'fold_{i+1}_{grp}'] = clf\n        \n        loss = round([v for k, v in clf.best_score_[\"validation\"].items() if \"MultiLogloss\" in k][0], 4)\n        f1_mean = np.mean([v for k, v in clf.best_score_[\"validation\"].items() if \"F1\" in k], dtype='float16')\n        f1_std = np.std([v for k, v in clf.best_score_[\"validation\"].items() if \"F1\" in k], dtype='float16')\n        print('### Fold', i+1, 'grp', grp, 'f1', f1_mean,'/', f1_mean - f1_std, 'loss', loss)\n        scores.append(f1_mean - f1_std)\n        \n        # scores.append(clf.best_score_['validation']['MultiClass'])\n        predicts = clf.predict_proba(valid_x[filtered_features_].astype('float32'))\n        a, b = limits[grp]\n        for e, q_num in enumerate(range(a, b)):\n            p = predicts[:, e]\n            oof.loc[valid_users, q_num-1] = p\n        \n        # for key, value in clf.get_all_params().items():\n                # if key in ['depth', 'learning_rate']:\n                    # print(\"{}, {}\".format(key, value))\n    \njoblib.dump(models, f\"./catboost_models.pkl\")\nprint(f'over all folds {np.mean(scores)}, {np.mean(scores) - np.std(scores):.4f}')","metadata":{"papermill":{"duration":69.877213,"end_time":"2023-02-07T01:02:43.57299","exception":false,"start_time":"2023-02-07T01:01:33.695777","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-07T19:15:05.143906Z","iopub.execute_input":"2023-06-07T19:15:05.144261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute CV Score\nWe need to convert prediction probabilities into `1s` and `0s`. The competition metric is F1 Score which is the harmonic mean of precision and recall. Let's find the optimal threshold for `p > threshold` when to predict `1` and when to predict `0` to maximize F1 Score.","metadata":{"papermill":{"duration":0.011241,"end_time":"2023-02-07T01:02:43.59638","exception":false,"start_time":"2023-02-07T01:02:43.585139","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# PUT TRUE LABELS INTO DATAFRAME WITH 18 COLUMNS\ntrue = oof.copy()\nfor k in range(18):\n    # GET TRUE LABELS\n    tmp = targets.loc[targets.q == k+1].set_index('session').loc[ALL_USERS]\n    true[k] = tmp.correct.values","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# FIND BEST THRESHOLD TO CONVERT PROBS INTO 1s AND 0s\nscores = []; thresholds = []\nbest_score = 0; best_threshold = 0\n\nfor threshold in np.arange(0.4,0.81,0.01):\n    print(f'{threshold:.02f}, ',end='')\n    preds = (oof.values.reshape((-1))>threshold).astype('int')\n    m = f1_score(true.values.reshape((-1)), preds, average='macro')   \n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score:\n        best_score = m\n        best_threshold = threshold\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# PLOT THRESHOLD VS. F1_SCORE\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='blue')\nplt.scatter([best_threshold], [best_score], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score:.3f} at Best Threshold = {best_threshold:.3}',size=18)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('When using optimal threshold...')\nfor k in range(18):\n        \n    # COMPUTE F1 SCORE PER QUESTION\n    m = f1_score(true[k].values, (oof[k].values>best_threshold).astype('int'), average='macro')\n    print(f'Q{k}: F1 =',m)\n    \n# COMPUTE F1 SCORE OVERALL\nm = f1_score(true.values.reshape((-1)), (oof.values.reshape((-1))>best_threshold).astype('int'), average='macro')\nprint('==> Overall F1 =',m)","metadata":{"papermill":{"duration":0.771134,"end_time":"2023-02-07T01:02:44.378465","exception":false,"start_time":"2023-02-07T01:02:43.607331","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Infer Test Data","metadata":{"papermill":{"duration":0.011075,"end_time":"2023-02-07T01:02:44.400918","exception":false,"start_time":"2023-02-07T01:02:44.389843","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# IMPORT KAGGLE API\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n# CLEAR MEMORY\nimport gc\n# del targets, df, oof, true\n_ = gc.collect()","metadata":{"papermill":{"duration":0.052132,"end_time":"2023-02-07T01:02:44.464739","exception":false,"start_time":"2023-02-07T01:02:44.412607","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models\n# fold_{i+1}_{grp}\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"limits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n# limits = {'0-4':(1,4), '0-4_5-12':(4,14), '0-4_5-12_13-22':(14,19)}\n# limits_convert = {'0-4': '0-4', '5-12': '0-4_5-12', '13-22':'0-4_5-12_13-22'}\n\ndtypes_dict = {'event_name':'category',\n               'level_group':'category',\n               'session_id':'int64',\n               'fqid':'category',\n               'room_fqid':'category',\n               'text':'category',\n               'name':'category'}\n\n# for _ in range(1):\nfor (test, sample_submission) in iter_test:\n    \n    # FEATURE ENGINEER TEST DATA    \n    test = test.astype(dtypes_dict)        \n    df = feature_engineer(test)\n    \n    grp = test.level_group.values[0]\n    \n    for f in filtered_features:\n        if f not in df.columns and f not in features2drop4grp[grp]:\n            if 'count' in f:\n                df[f] = 0\n            else: df[f] = None\n            \n    # INFER TEST DATA\n    df.drop(features2drop4grp[grp], inplace=True, axis=1, errors = 'ignore')\n    \n#     grp = limits_convert[grp]\n    a,b = limits[grp]      \n    filtered_features_ = [c for c in filtered_features if c not in features2drop4grp[grp]]\n    pred1 = models[f'fold_{1}_{grp}'].predict_proba(df[filtered_features_]) \n    pred2 = models[f'fold_{2}_{grp}'].predict_proba(df[filtered_features_]) \n    pred3 = models[f'fold_{3}_{grp}'].predict_proba(df[filtered_features_]) \n    predicts = (pred1 + pred2 + pred3)/3\n    predicts = predicts[0]\n    \n    for i, t in enumerate(range(a, b)):\n        p = predicts[i]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n    \n    env.predict(sample_submission)\n    ","metadata":{"papermill":{"duration":1.002014,"end_time":"2023-02-07T01:02:45.47927","exception":false,"start_time":"2023-02-07T01:02:44.477256","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA submission.csv","metadata":{"papermill":{"duration":0.011427,"end_time":"2023-02-07T01:02:45.502331","exception":false,"start_time":"2023-02-07T01:02:45.490904","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df = pd.read_csv('submission.csv')\nprint( df.shape )\ndf.head()","metadata":{"papermill":{"duration":0.027432,"end_time":"2023-02-07T01:02:45.541022","exception":false,"start_time":"2023-02-07T01:02:45.51359","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df.correct.mean())","metadata":{"papermill":{"duration":0.020233,"end_time":"2023-02-07T01:02:45.57314","exception":false,"start_time":"2023-02-07T01:02:45.552907","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}