{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### <center style=\"background-color:Gainsboro; width:60%;\">Load Data</center>\nThe very first steps are to load the data and get a quick understanding of what it is and how we can use it to fulfill out goal:\n\n\"For each `<session_id>_<question #>`, you are predicting the `correct` column, identifying whether you believe the user for this particular session will answer this question correctly, using only the previous information for the session.\n\nThe timeseries API presents the questions and data to you in order of levels - level segments 0-4, 5-12, and 13-22 are each provided in sequence, and you will be predicting the correctness of each segment's questions as they are presented.\"","metadata":{}},{"cell_type":"code","source":"#=========================================================================\n# Load up the libraries\n#=========================================================================\nimport pandas as pd\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)\nimport numpy as np\nimport gc\nimport matplotlib.pyplot as plt\nimport xgboost as xgb\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score\nfrom sklearn.model_selection import KFold, GroupKFold","metadata":{"execution":{"iopub.status.busy":"2023-07-18T23:46:15.615227Z","iopub.execute_input":"2023-07-18T23:46:15.615661Z","iopub.status.idle":"2023-07-18T23:46:16.462306Z","shell.execute_reply.started":"2023-07-18T23:46:15.615629Z","shell.execute_reply":"2023-07-18T23:46:16.460561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#=========================================================================\n# Load in the data\n#=========================================================================\ntest  = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/test.csv')\nlabels = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2023-07-18T23:46:16.464424Z","iopub.execute_input":"2023-07-18T23:46:16.464815Z","iopub.status.idle":"2023-07-18T23:46:17.134947Z","shell.execute_reply.started":"2023-07-18T23:46:16.464781Z","shell.execute_reply":"2023-07-18T23:46:17.1339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note: I have been having extreme issues with RAM given the size of the dataset, and in transferring this to Kaggle I found I had to adopt code from another notebook or else reduce the dataset I used. It seems like the LHL-provided examples on how to do XGBoost classification do not work on my system with this dataset (the competition appears to have been changed only a few months ago), so some of this notebook uses code from a prior submission/baseline project that breaks it up into chunks that are processed individually. Otherwise, my features, effort, and outcome are dependent on my work from before finding this notebook.","metadata":{}},{"cell_type":"code","source":"# READ USER ID ONLY\ntmp = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\",usecols=[0])\ntmp = tmp.groupby('session_id').session_id.agg('count')\n\n# COMPUTE READS AND SKIPS\nPIECES = 12\nCHUNK = int( np.ceil(len(tmp)/PIECES) )\n\nreads = []\nskips = [0]\nfor k in range(PIECES):\n    a = k*CHUNK\n    b = (k+1)*CHUNK\n    if b>len(tmp): b=len(tmp)\n    r = tmp.iloc[a:b].sum()\n    reads.append(r)\n    skips.append(skips[-1]+r)\n    \nprint(f'To avoid memory error, we will read train in {PIECES} pieces of sizes:')\nprint(reads)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2023-07-18T23:46:17.136509Z","iopub.execute_input":"2023-07-18T23:46:17.13731Z","iopub.status.idle":"2023-07-18T23:47:47.802461Z","shell.execute_reply.started":"2023-07-18T23:46:17.137273Z","shell.execute_reply":"2023-07-18T23:47:47.801577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <center style=\"background-color:Gainsboro; width:60%;\">Feature Selection & Engineering</center>\nStarting off, `session_id` in the labels DF has a `_qX` suffix, indicating which question was answered. This can be split into two columns, although the `session_id` column is also being treated as the index column so it needs to be reset first.","metadata":{}},{"cell_type":"code","source":"#=========================================================================\n# reset index, split session_id\n#=========================================================================\nlabels['session'] = labels.session_id.apply(lambda x: int(x.split('_')[0]) )\nlabels['q_answered'] = labels.session_id.apply(lambda x: int(x.split('_')[-1][1:]) )\nprint( labels.shape )\nlabels.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T23:47:47.804858Z","iopub.execute_input":"2023-07-18T23:47:47.805851Z","iopub.status.idle":"2023-07-18T23:47:48.942431Z","shell.execute_reply.started":"2023-07-18T23:47:47.805817Z","shell.execute_reply":"2023-07-18T23:47:48.941603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I also want my final df to have features that would be missing, such as `elapsed_time`, so the `numbers` list has been used for that reason. ","metadata":{}},{"cell_type":"code","source":"#=========================================================================\n# define typed lists\n#=========================================================================\n# For Engineering\ncategories = ['event_name', 'fqid', 'room_fqid', 'text']\n\nevents = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']\n\n# For retaining features\nnumbers = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n\n#=========================================================================\n# create engineered dataframe\n#=========================================================================\ndef feature_engineer(train):\n    # Create list of new dfs to later merge together\n    eng_df = []\n\n    # We don't really need the exact data for these columns\n    for cat in categories:\n        temp = train.groupby(['session_id','level_group'])[cat].agg('nunique')\n        temp.name = temp.name + '_nunique'\n        eng_df.append(temp)\n\n    # Simplify events data\n    for step in events:\n        train[step] = (train.event_name == step).astype('int8')\n    for step in events + ['elapsed_time']:\n        temp = train.groupby(['session_id','level_group'])[step].agg('sum')\n        temp.name = temp.name + '_sum'\n        eng_df.append(temp)\n\n    # Keep numerical data\n    for step in numbers:\n        temp = train.groupby(['session_id','level_group'])[step].agg('sum')\n        eng_df.append(temp)\n\n    # Drop \n    train = train.drop(events,axis=1)\n\n    df = pd.concat(eng_df,axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-07-18T23:47:48.9435Z","iopub.execute_input":"2023-07-18T23:47:48.944367Z","iopub.status.idle":"2023-07-18T23:47:48.954967Z","shell.execute_reply.started":"2023-07-18T23:47:48.944335Z","shell.execute_reply":"2023-07-18T23:47:48.953944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# PROCESS TRAIN DATA IN PIECES\nall_pieces = []\nprint(f'Processing train as {PIECES} pieces to avoid memory error... ')\nfor k in range(PIECES):\n    print(k,', ',end='')\n    SKIPS = 0\n    if k>0: SKIPS = range(1,skips[k]+1)\n    train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv',\n                        nrows=reads[k], skiprows=SKIPS)\n    df = feature_engineer(train)\n    all_pieces.append(df)\n    \n# CONCATENATE ALL PIECES\nprint('\\n')\ndel train; gc.collect()\ndf = pd.concat(all_pieces, axis=0)\nprint('Shape of all train data after feature engineering:', df.shape )\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T23:47:48.956211Z","iopub.execute_input":"2023-07-18T23:47:48.957219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = labels.reset_index()\nprint(f\"Length = {len(labels)}\")\nlabels = labels.drop(\"index\", axis=1)\nlabels.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <center style=\"background-color:Gainsboro; width:60%;\">XGBoost Model</center>\nNow that I have my cleaned training data in `df`, I need to build he model. I have the following considerations to keep in mind.\n1. The `labels` df has a column for individual questions (i.e. level within a group), with the primary key being session_id.\n2. The new training `df` has a column for level group, but not for individual questions.","metadata":{}},{"cell_type":"code","source":"features = [c for c in df.columns if c != 'level_group']\nprint(len(features))\nuser_list = df.index.unique()\nprint(len(user_list))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gkf = GroupKFold(n_splits=5)\noof = pd.DataFrame(data=np.zeros((len(user_list),18)), index=user_list)\nmodels = {}\n\n# CV score with 5 group kfold\nfor i, (train_index, test_index) in enumerate(gkf.split(X=df, groups=df.index)):\n    print('#'*25)\n    print('### Fold',i+1)\n    print('#'*25)\n    \n    xgb_params = {\n    'objective' : 'binary:logistic',\n    'eval_metric':'logloss',\n    'learning_rate': 0.05,\n    'max_depth': 4,\n    'n_estimators': 1000,\n    'early_stopping_rounds': 50,\n    'tree_method':'hist',\n    'subsample':0.8,\n    'colsample_bytree': 0.4}\n    \n    # Iterate through questions in labels df\n    for t in range(1,19):\n        \n        # Use this to train data based on groupings that match training set\n        if t<=3: grp = '0-4'\n        elif t<=13: grp = '5-12'\n        elif t<=22: grp = '13-22'\n            \n        # TRAIN DATA\n        train_x = df.iloc[train_index]\n        train_x = train_x.loc[train_x.level_group == grp]\n        train_users = train_x.index.values\n        train_y = labels.loc[labels.q_answered==t].set_index('session').loc[train_users]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index]\n        valid_x = valid_x.loc[valid_x.level_group == grp]\n        valid_users = valid_x.index.values\n        valid_y = labels.loc[labels.q_answered==t].set_index('session').loc[valid_users]\n        \n        # Train Model        \n        clf =  XGBClassifier(**xgb_params)\n        clf.fit(train_x[features].astype('float32'), train_y['correct'],\n                eval_set=[ (valid_x[features].astype('float32'), valid_y['correct']) ],\n                verbose=0)\n        print(f'{t}({clf.best_ntree_limit}), ',end='')\n        \n        # Save model, predict valid OOF\n        models[f'{grp}_{t}'] = clf\n        oof.loc[valid_users, t-1] = clf.predict_proba(valid_x[features].astype('float32'))[:,1]\n        \n    print()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <center style=\"background-color:Gainsboro; width:60%;\">Compute CV Score</center>\nAgain borrowing from a baseline submission on Kaggle given RAM limitations...","metadata":{}},{"cell_type":"code","source":"# Create DF with 18 columns\ntrue = oof.copy()\nfor k in range(18):\n    tmp = labels.loc[labels.q_answered == k+1].set_index('session').loc[user_list]\n    true[k] = tmp.correct.values","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Determine best threshold for converting probabilities into 1s and 0s\nscores = []; thresholds = []\nbest_score = 0; best_threshold = 0\n\nfor threshold in np.arange(0.50,0.78,0.01):\n    print(f'{threshold:.02f}, ',end='')\n    preds = (oof.values.reshape((-1))>threshold).astype('int')\n    m = f1_score(true.values.reshape((-1)), preds, average='macro')   \n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score:\n        best_score = m\n        best_threshold = threshold","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='cyan')\nplt.scatter([best_threshold], [best_score], color='black', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score:.3f} at Best Threshold = {best_threshold:.3}',size=18)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('When using optimal threshold...')\nfor k in range(18):\n        \n    # COMPUTE F1 SCORE PER QUESTION\n    m = f1_score(true[k].values, (oof[k].values>best_threshold).astype('int'), average='macro')\n    print(f'Q{k}: F1 =',m)\n    \n# COMPUTE F1 SCORE OVERALL\nm = f1_score(true.values.reshape((-1)), (oof.values.reshape((-1))>best_threshold).astype('int'), average='macro')\nprint('==> Overall F1 =',m)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <center style=\"background-color:Gainsboro; width:60%;\">Submission</center>","metadata":{}},{"cell_type":"code","source":"true.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# IMPORT KAGGLE API\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n# CLEAR MEMORY\nimport gc\ndel targets, df, oof, true\n_ = gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}