{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**What are you trying to do in this notebook?**\n\nIn this competition, I'm predicting student performance during game-based learning in real-time. I'll develop a model trained on one of the largest open datasets of game logs.\n\nMy work will help advance research into knowledge-tracing methods for game-based learning. I'll be supporting developers of educational games to create more effective learning experiences for students.\n\nLearning is meant to be fun, which is where game-based learning comes in. This educational approach allows students to engage with educational content inside a game framework, making it enjoyable and dynamic. Although game-based learning is being used in a growing number of educational settings, there are still a limited number of open datasets available to apply data science and learning analytic principles to improve game-based learning.\n\n**Why are you trying it?**\n\nTo help advance research into knowledge-tracing methods for game-based learning. I'm using Kaggle's time series API. Test data will be delivered in groupings that do not allow access to future data. The objective of this competition is to use time series data generated by an online educational game to determine whether players will answer questions correctly. There are three question checkpoints (level 4, level 12, and level 22), each with a number of questions. At each checkpoint, I will have access to all previous test data for that section.\n\n**What should I expect the data format to be?**\n\nThe training columns are as listed below. The label rows are identified with <session_id>_<question #>. Each session will have 18 rows, representing 18 questions.\n\n**What am I predicting?**\n\nFor each <session_id>_<question #>, I'm predicting the correct column, identifying whether you believe the user for this particular session will answer this question correctly, using only the previous information for the session.\n\nNote that the hidden test set is roughly as large as the training set; you should expect it will take much longer to run on than the three test samples provided.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:55.059199Z","iopub.execute_input":"2023-02-21T05:43:55.059900Z","iopub.status.idle":"2023-02-21T05:43:55.219785Z","shell.execute_reply.started":"2023-02-21T05:43:55.059806Z","shell.execute_reply":"2023-02-21T05:43:55.218776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport gc\nimport sys\nimport polars as pl\n\nfrom catboost import CatBoostClassifier, Pool\n\nlevel_groups = [\"0-4\", \"5-12\", \"13-22\"]\nlevel_groups_reverse = {'0-4': 0, '5-12': 1, '13-22': 2}\nlevels = {'0-4': (0, 5), '5-12': (5, 13), '13-22': (13, 23)}\nquestions = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}\nEVENTS = ['checkpoint', 'cutscene_click', 'map_click', 'map_hover', 'navigate_click', 'notebook_click', 'notification_click', 'object_click', 'object_hover', 'observation_click', 'person_click']\nCATS = ['event_name', 'name','fqid', 'room_fqid', 'text_fqid']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'hover_duration', \"time_past\"]","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:55.221647Z","iopub.execute_input":"2023-02-21T05:43:55.222274Z","iopub.status.idle":"2023-02-21T05:43:56.990907Z","shell.execute_reply.started":"2023-02-21T05:43:55.222239Z","shell.execute_reply":"2023-02-21T05:43:56.989950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns = [\n    (\n        (pl.col(\"elapsed_time\") - pl.col(\"elapsed_time\").shift(1))\n         .fill_null(0)\n         .clip(0, 1e9)\n         .over([\"session_id\", \"level_group\"])\n         .alias(\"time_past\")\n    ),\n]\naggs = [\n    *[pl.col(c).drop_nulls().n_unique().alias(f\"{c}_unique\") for c in CATS],\n    *[pl.col(c).mean().alias(f\"{c}_mean\") for c in NUMS],\n    *[pl.col(c).std().alias(f\"{c}_std\") for c in NUMS],\n    *[(pl.col(\"event_name\") == c).sum().alias(f\"{c}_sum\") for c in EVENTS],\n]\nmodels_list = [[CatBoostClassifier().load_model(\n    f\"/kaggle/input/cpu-catboost-baseline-using-polars-train/fold{fold}_q{q}.cbm\"\n) for fold in range(5)] for q in range(1, 19)]","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:56.996097Z","iopub.execute_input":"2023-02-21T05:43:56.998397Z","iopub.status.idle":"2023-02-21T05:43:57.920996Z","shell.execute_reply.started":"2023-02-21T05:43:56.998351Z","shell.execute_reply":"2023-02-21T05:43:57.919759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def trans(test):\n    return (pl.from_pandas(test)    \n            .with_columns(columns)\n            .select(aggs)\n            .fill_null(-1)\n            .to_pandas()\n            )","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:57.928910Z","iopub.execute_input":"2023-02-21T05:43:57.929714Z","iopub.status.idle":"2023-02-21T05:43:57.935904Z","shell.execute_reply.started":"2023-02-21T05:43:57.929673Z","shell.execute_reply":"2023-02-21T05:43:57.934671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:57.940481Z","iopub.execute_input":"2023-02-21T05:43:57.941323Z","iopub.status.idle":"2023-02-21T05:43:57.976619Z","shell.execute_reply.started":"2023-02-21T05:43:57.941280Z","shell.execute_reply":"2023-02-21T05:43:57.975257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for (sample_submission, test) in iter_test:\n    target_level_group = level_groups_reverse[test.level_group.iloc[0]]\n    df = trans(test)\n    \n    fold = 0\n    preds = []\n    for q in range(*questions[level_groups[target_level_group]]):\n        model = models_list[q - 1][fold]\n        feature_cols = model.feature_names_\n        pred = model.predict_proba(df[feature_cols].astype(np.float32))[0,1]\n        preds.append(int(pred > 0.63))\n\n    sample_submission[\"correct\"] = preds\n\n    env.predict(sample_submission)","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:57.981185Z","iopub.execute_input":"2023-02-21T05:43:57.981746Z","iopub.status.idle":"2023-02-21T05:43:58.947131Z","shell.execute_reply.started":"2023-02-21T05:43:57.981706Z","shell.execute_reply":"2023-02-21T05:43:58.945794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.read_csv('submission.csv')\nprint(sub.shape, sub.correct.mean())","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:58.950032Z","iopub.execute_input":"2023-02-21T05:43:58.950408Z","iopub.status.idle":"2023-02-21T05:43:58.961233Z","shell.execute_reply.started":"2023-02-21T05:43:58.950374Z","shell.execute_reply":"2023-02-21T05:43:58.960283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T05:43:58.965474Z","iopub.execute_input":"2023-02-21T05:43:58.967769Z","iopub.status.idle":"2023-02-21T05:43:58.985510Z","shell.execute_reply.started":"2023-02-21T05:43:58.967726Z","shell.execute_reply":"2023-02-21T05:43:58.984005Z"},"trusted":true},"execution_count":null,"outputs":[]}]}