{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Saving Predictions from Previous Levels - 2 Ways!\n\nOne of the challenges of this competition is how to deal with kaggle's time series API when making predictions. Unlike other competitions, we aren't given one file of test data to predict on, but rather presented the data sequentially, where we are provided with their data for one level group at a time, in order, one student at a time.\n\nThis unfortunately limits the models that we can make, because we don't have the full set of information when making predictions, but this is a really valuable skill to learn in data science because this is much more similar to real life - you can never know what happens in the future! \n\nThat being said, we do know what happens in the past, and using this information in the right way can give our models more information to work with to make more accurate predictions, and boost your model scores on the leaderboard. In this notebook, we will cover two different methods to use information from previous levels to improve our predictions at inference time. The steps taken to train your models with this extra information are very similar!\n\n# How does the Time Series API work?\n\nFirst, it's very important to understand how the API for this competition works. At inference time, we will be presented with the data for one student (or one `session_id`) at a time, in the order of the level groups.\n\nFor example, if we just print out the first few records the API presents us with, we will find:\n\n| Session ID | Level Group | \n| ----- | ----- |\n| 20090109393214576 | 0-4 | \n| 20090312143683264 | 0-4 | \n| 20090312331414616 | 0-4 | \n| 20090109393214576 | 5-12 |\n| 20090312143683264 | 5-12 | \n| 20090312331414616 | 5-12 | \n| 20090109393214576 | 13-22 |\n| 20090312143683264 | 13-22 | \n| 20090312331414616 | 13-22 | \n\nSo we're looping over our session IDs for one level group (20090109393214576 -> 20090312143683264 -> 20090312331414616), and then the same session IDs for the next level groups (0-4 -> 5-12 -> 13-22).  \n\n\nLet's print that out to prove it!","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom IPython.display import display","metadata":{"execution":{"iopub.status.busy":"2023-03-13T20:06:46.280642Z","iopub.execute_input":"2023-03-13T20:06:46.281084Z","iopub.status.idle":"2023-03-13T20:06:46.286071Z","shell.execute_reply.started":"2023-03-13T20:06:46.281048Z","shell.execute_reply":"2023-03-13T20:06:46.285147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-03-13T19:32:41.519606Z","iopub.execute_input":"2023-03-13T19:32:41.520992Z","iopub.status.idle":"2023-03-13T19:32:41.526869Z","shell.execute_reply.started":"2023-03-13T19:32:41.520938Z","shell.execute_reply":"2023-03-13T19:32:41.525849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"counter = 0\nfor (sample_submission, test) in iter_test:\n    session_id = test['session_id'].unique()\n    level_group = test['level_group'].unique()\n    \n    print()\n    print(f'Session ID: {session_id}')\n    print(f'Level Group: {level_group}')\n    print()\n    print('='*30)\n    \n    ## Make a dummy submission so we can move on \n    sample_submission['correct'] = 0\n    env.predict(sample_submission)\n    counter += 1","metadata":{"execution":{"iopub.status.busy":"2023-03-13T19:32:41.769264Z","iopub.execute_input":"2023-03-13T19:32:41.770102Z","iopub.status.idle":"2023-03-13T19:32:41.816223Z","shell.execute_reply.started":"2023-03-13T19:32:41.770060Z","shell.execute_reply":"2023-03-13T19:32:41.815313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we better understand how the API works, let's explore how we can use this ordering of the data to feed some extra information into our models!","metadata":{}},{"cell_type":"markdown","source":"# Method 1: Saving predictions from previous level groups\nOne of the easiest ways to save past information is to leverage the fact that the data is presented in the order of level groups. We can make predictions for the first level group for a session_id, store this, and then add these predictions as a feature when we want to make predictions for the next level group! That way we know how well the student has performed (or how well our models think they have performed) on previous quesions, and our model can use this as extra information. \n\nLet's see how we can do this!      \n(We won't use an actual model to make predictions for now, but the code to do so has been left commented out. The focus of this notebook is on how to save the predictions at inference time.)","metadata":{}},{"cell_type":"code","source":"# First let's reset the API\njo_wilder.make_env.__called__ = False\nenv.__called__ = False\ntype(env)._state = type(type(env)._state).__dict__['INIT']\n\n# And reinitialise it\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-03-13T20:06:53.466640Z","iopub.execute_input":"2023-03-13T20:06:53.467062Z","iopub.status.idle":"2023-03-13T20:06:53.473579Z","shell.execute_reply.started":"2023-03-13T20:06:53.467027Z","shell.execute_reply":"2023-03-13T20:06:53.472627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define some basic feature engineering, courtesy of Chris Deotte\nCATS = ['event_name', 'fqid', 'room_fqid', 'text']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n\n# https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']\n\n# https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\ndef feature_engineer(train):\n    \n    dfs = []\n    for c in CATS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('nunique')\n        tmp.name = tmp.name + '_nunique'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('mean')\n        tmp.name = tmp.name + '_mean'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('std')\n        tmp.name = tmp.name + '_std'\n        dfs.append(tmp)\n    for c in EVENTS: \n        train[c] = (train.event_name == c).astype('int8')\n    for c in EVENTS + ['elapsed_time']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('sum')\n        tmp.name = tmp.name + '_sum'\n        dfs.append(tmp)\n    train = train.drop(EVENTS,axis=1)\n        \n    df = pd.concat(dfs,axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-03-13T20:06:53.641974Z","iopub.execute_input":"2023-03-13T20:06:53.642809Z","iopub.status.idle":"2023-03-13T20:06:53.654810Z","shell.execute_reply.started":"2023-03-13T20:06:53.642738Z","shell.execute_reply":"2023-03-13T20:06:53.653563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Saving predictions from previous level groups\npreds = {}\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nbest_threshold = 0.63\n\n# LOAD MODELS FOR EACH QUESTION (excluded here for simplicity)\n# models = {}\n# for t in range(1, 19):\n#     clf = XGBClassifier()\n#     clf.load_model(f'/kaggle/input/.../question_{t}.xgb')\n#     models[t] = clf\n\n# The kaggle API will present data in order of level_groups\nfor (sample_submission, test) in iter_test:\n    \n    # Figure out our level group and the limits of that level group\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n\n    # FEATURE ENGINEERING\n    df = feature_engineer(test)\n\n    # Initialise preds for this session_id\n    session_id = sample_submission.iloc[0, :]['session_id'].split('_')[0]\n    if session_id not in preds.keys():\n        preds[session_id] = {}\n\n    # ADD PREDS FROM PREVIOUS LEVEL GROUPS TO PREDICT!\n    if grp == '5-12':\n        for i in range(1,4):\n            df[f'preds_{i}'] = preds[session_id][i]\n    elif grp == '13-22':\n        for i in range(1,14):\n            df[f'preds_{i}'] = preds[session_id][i]\n            \n    # Go in order of questions for each level group\n    for t in range(a,b):\n            \n        # Make a prediction (excluded here for simplicity)\n        # p = clf.predict_proba(df)[:,1].item()\n        p = np.random.rand() # (Replace with the above when making actual predictions)\n\n        # SAVE PREDICTION TO DICT\n        preds[session_id][t] = p\n\n        # Make a submission \n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask, 'correct'] = int(p>best_threshold)\n        \n    env.predict(sample_submission)\n    \n    # Print out the predictions we save after each level group!\n    display(pd.DataFrame.from_dict(preds).T)","metadata":{"execution":{"iopub.status.busy":"2023-03-13T20:06:53.979269Z","iopub.execute_input":"2023-03-13T20:06:53.979685Z","iopub.status.idle":"2023-03-13T20:06:54.674122Z","shell.execute_reply.started":"2023-03-13T20:06:53.979648Z","shell.execute_reply":"2023-03-13T20:06:54.672998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Method 2: Saving predictions from previous levels\nAnother way in which we can perform this knowledge distillation from the past, is by not only saving the predictions from previous level groups, but from *all* previous levels. e.g. Using predictions from level 1 when making predictions for level 2, and from levels 1-17 when predictions for level 18. This should give our model even more information to work with, and our model may be able to pick up patters where if a student performs well on one question then they might do well on another!\n\n(Again, we won't use an actual model to make predictions for now, but the code to do so has been left commented out. The focus of this notebook is on how to save the predictions at inference time.)","metadata":{}},{"cell_type":"code","source":"# Let's reset the API again\njo_wilder.make_env.__called__ = False\nenv.__called__ = False\ntype(env)._state = type(type(env)._state).__dict__['INIT']\n\n# And reinitialise it\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-03-13T20:12:05.496158Z","iopub.execute_input":"2023-03-13T20:12:05.496573Z","iopub.status.idle":"2023-03-13T20:12:05.503103Z","shell.execute_reply.started":"2023-03-13T20:12:05.496536Z","shell.execute_reply":"2023-03-13T20:12:05.502137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This time we will save predictions from previous levels\npreds = {}\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nbest_threshold = 0.63\n\n# LOAD MODELS FOR EACH QUESTION (excluded here for simplicity)\n# models = {}\n# for t in range(1, 19):\n#     clf = XGBClassifier()\n#     clf.load_model(f'/kaggle/input/.../question_{t}.xgb')\n#     models[t] = clf\n\n# The kaggle API will present data in order of level_groups\nfor (sample_submission, test) in iter_test:\n    \n    # Figure out our level group and the limits of that level group\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n\n    # FEATURE ENGINEERING\n    df = feature_engineer(test)\n\n    # Initialise preds for this session_id\n    session_id = sample_submission.iloc[0, :]['session_id'].split('_')[0]\n    if session_id not in preds.keys():\n        preds[session_id] = {}\n            \n    # Go in order of questions for each level group\n    for t in range(a,b):\n        \n        # ADD PREDS FROM PREVIOUS LEVELS TO PREDICT!\n        if t > 1:\n            for i in range(1, t):\n                df[i-1] = preds[session_id][i]\n            \n        # Make a prediction (excluded here for simplicity)\n        # p = clf.predict_proba(df)[:,1].item()\n        p = np.random.rand() # (Replace with the above when making actual predictions)\n\n        # SAVE PREDICTION TO DICT\n        preds[session_id][t] = p\n\n        # Make a submission \n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask, 'correct'] = int(p>best_threshold)\n        \n        # Print out the predictions we save after each level!\n        display(pd.DataFrame.from_dict(preds).T)\n        \n    env.predict(sample_submission)","metadata":{"execution":{"iopub.status.busy":"2023-03-13T20:12:05.852741Z","iopub.execute_input":"2023-03-13T20:12:05.853551Z","iopub.status.idle":"2023-03-13T20:12:07.288471Z","shell.execute_reply.started":"2023-03-13T20:12:05.853510Z","shell.execute_reply":"2023-03-13T20:12:07.287107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And now this time you can see how we add predictions one question at a time for each student!\n\nIf you found this notebook useful, please don't forget to upvote it. Hopefully this helps you to squeeze even more out of this competition and climb the leaderboard. Happy Kaggling!","metadata":{}}]}