{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this competition is to predict student performance during game-based learning in real-time. I'll develop a model trained on one of the largest open datasets of game logs.\n\nMy work will help advance research into knowledge-tracing methods for game-based learning. I'll be supporting developers of educational games to create more effective learning experiences for students.\n\nLearning is meant to be fun, which is where game-based learning comes in. This educational approach allows students to engage with educational content inside a game framework, making it enjoyable and dynamic. Although game-based learning is being used in a growing number of educational settings, there are still a limited number of open datasets available to apply data science and learning analytic principles to improve game-based learning.\n\nMost game-based learning platforms do not sufficiently make use of knowledge tracing to support individual students. Knowledge tracing methods have been developed and studied in the context of online learning environments and intelligent tutoring systems. But there has been less focus on knowledge tracing in educational games.\n\nCompetition host Field Day Lab is a publicly-funded research lab at the Wisconsin Center for Educational Research. They design games for many subjects and age groups that bring contemporary research to the public, making use of the game data to understand how people learn. Field Day Lab's commitment to accessibility ensures all of its games are free and available to anyone. The lab also partners with nonprofits like The Learning Agency Lab, which is focused on developing science of learning-based tools and programs for the social good.\n\nIf successful, I'll enable game developers to improve educational games and further support the educators who use these games with dashboards and analytic tools. In turn, we might see broader support for game-based learning platforms.\n\nIn this competition, I am using time series data which is generated by an online educational game to determine whether players will answer questions correctly. Developing a model trained on one of the largest open datasets of game logs. Transforms the labels into a dataframe with multi labels for training and inference.\n- Use the sample notebooks to iterate over the test data, which is split as described and served up as Pandas dataframes.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:02:47.899155Z","iopub.execute_input":"2023-02-15T13:02:47.899554Z","iopub.status.idle":"2023-02-15T13:02:47.910700Z","shell.execute_reply.started":"2023-02-15T13:02:47.899522Z","shell.execute_reply":"2023-02-15T13:02:47.909704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import KFold, GroupKFold\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import f1_score","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:02:47.913086Z","iopub.execute_input":"2023-02-15T13:02:47.913836Z","iopub.status.idle":"2023-02-15T13:02:47.924367Z","shell.execute_reply.started":"2023-02-15T13:02:47.913799Z","shell.execute_reply":"2023-02-15T13:02:47.923302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from vowpalwabbit.sklearn_vw import VWClassifier","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:02:47.926151Z","iopub.execute_input":"2023-02-15T13:02:47.926945Z","iopub.status.idle":"2023-02-15T13:02:47.937896Z","shell.execute_reply.started":"2023-02-15T13:02:47.926910Z","shell.execute_reply":"2023-02-15T13:02:47.936820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv')\nprint( train.shape )\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:02:47.939681Z","iopub.execute_input":"2023-02-15T13:02:47.940503Z","iopub.status.idle":"2023-02-15T13:03:26.571775Z","shell.execute_reply.started":"2023-02-15T13:02:47.940461Z","shell.execute_reply":"2023-02-15T13:03:26.570744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )\ntargets['q'] = targets.session_id.apply(lambda x: int(x.split('_')[-1][1:]) )\nprint( targets.shape )\ntargets.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:03:26.574434Z","iopub.execute_input":"2023-02-15T13:03:26.575139Z","iopub.status.idle":"2023-02-15T13:03:27.224354Z","shell.execute_reply.started":"2023-02-15T13:03:26.575095Z","shell.execute_reply":"2023-02-15T13:03:27.223288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CATS = ['event_name', 'name','fqid', 'room_fqid', 'text_fqid']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:03:27.225694Z","iopub.execute_input":"2023-02-15T13:03:27.226037Z","iopub.status.idle":"2023-02-15T13:03:27.231497Z","shell.execute_reply.started":"2023-02-15T13:03:27.226008Z","shell.execute_reply":"2023-02-15T13:03:27.230259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_engineer(train):\n    dfs = []\n    for c in CATS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('nunique')\n        tmp.name = tmp.name + '_nunique'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('mean')\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('std')\n        tmp.name = tmp.name + '_std'\n        dfs.append(tmp)\n    df = pd.concat(dfs,axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:03:27.232903Z","iopub.execute_input":"2023-02-15T13:03:27.233311Z","iopub.status.idle":"2023-02-15T13:03:27.243133Z","shell.execute_reply.started":"2023-02-15T13:03:27.233281Z","shell.execute_reply":"2023-02-15T13:03:27.242031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf = feature_engineer(train)\nprint( df.shape )\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:03:27.245010Z","iopub.execute_input":"2023-02-15T13:03:27.245363Z","iopub.status.idle":"2023-02-15T13:04:12.001537Z","shell.execute_reply.started":"2023-02-15T13:03:27.245319Z","shell.execute_reply":"2023-02-15T13:04:12.000477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FEATURES = [c for c in df.columns if c != 'level_group']\nprint('We will train with', len(FEATURES) ,'features')\nALL_USERS = df.index.unique()\nprint('We will train with', len(ALL_USERS) ,'users info')","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:04:12.002950Z","iopub.execute_input":"2023-02-15T13:04:12.003665Z","iopub.status.idle":"2023-02-15T13:04:12.013495Z","shell.execute_reply.started":"2023-02-15T13:04:12.003625Z","shell.execute_reply":"2023-02-15T13:04:12.012454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gkf = GroupKFold(n_splits=5)\noof = pd.DataFrame(data=np.zeros((len(ALL_USERS),18)), index=ALL_USERS)\nmodels = {}\n\n# COMPUTE CV SCORE WITH 5 GROUP K FOLD\nfor i, (train_index, test_index) in enumerate(gkf.split(X=df, groups=df.index)):\n    print('#'*25)\n    print('### Fold',i+1)\n    print('#'*25)\n    \n    # ITERATE THRU QUESTIONS 1 THRU 18\n    for t in range(1,19):\n        print(t,', ',end='')\n        \n        # USE THIS TRAIN DATA WITH THESE QUESTIONS\n        if t<=3: grp = '0-4'\n        elif t<=13: grp = '5-12'\n        elif t<=22: grp = '13-22'\n            \n        # TRAIN DATA\n        train_x = df.iloc[train_index]\n        train_x = train_x.loc[train_x.level_group == grp]\n        train_users = train_x.index.values\n        train_y = targets.loc[targets.q==t].set_index('session').loc[train_users]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index]\n        valid_x = valid_x.loc[valid_x.level_group == grp]\n        valid_users = valid_x.index.values\n        valid_y = targets.loc[targets.q==t].set_index('session').loc[valid_users]\n        \n        # TRAIN MODEL\n        clf = VWClassifier()\n        clf.fit(train_x[FEATURES].astype('float32'), train_y['correct'])\n        \n        # SAVE MODEL, PREDICT VALID OOF\n        models[f'{grp}_{t}'] = clf\n        oof.loc[valid_users, t-1] = clf.predict_proba(valid_x[FEATURES].astype('float32'))[:,1]\n        \n    print()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:04:12.015480Z","iopub.execute_input":"2023-02-15T13:04:12.016334Z","iopub.status.idle":"2023-02-15T13:06:02.111143Z","shell.execute_reply.started":"2023-02-15T13:04:12.016173Z","shell.execute_reply":"2023-02-15T13:06:02.109822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"true = oof.copy()\nfor k in range(18):\n    tmp = targets.loc[targets.q == k+1].set_index('session').loc[ALL_USERS]\n    true[k] = tmp.correct.values","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:02.112568Z","iopub.execute_input":"2023-02-15T13:06:02.112923Z","iopub.status.idle":"2023-02-15T13:06:02.194954Z","shell.execute_reply.started":"2023-02-15T13:06:02.112892Z","shell.execute_reply":"2023-02-15T13:06:02.193933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scores = []; thresholds = []\nbest_score = 0; best_threshold = 0\n\nfor threshold in np.arange(0.4,0.81,0.01):\n    print(f'{threshold:.02f}, ',end='')\n    preds = (oof.values.reshape((-1))>threshold).astype('int')\n    m = f1_score(true.values.reshape((-1)), preds, average='macro')   \n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score:\n        best_score = m\n        best_threshold = threshold","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:02.196136Z","iopub.execute_input":"2023-02-15T13:06:02.196470Z","iopub.status.idle":"2023-02-15T13:06:05.289982Z","shell.execute_reply.started":"2023-02-15T13:06:02.196440Z","shell.execute_reply":"2023-02-15T13:06:05.288809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='blue')\nplt.scatter([best_threshold], [best_score], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score:.3f} at Best Threshold = {best_threshold:.3}',size=18)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:05.291533Z","iopub.execute_input":"2023-02-15T13:06:05.291869Z","iopub.status.idle":"2023-02-15T13:06:05.540233Z","shell.execute_reply.started":"2023-02-15T13:06:05.291840Z","shell.execute_reply":"2023-02-15T13:06:05.539421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('When using optimal threshold...')\nfor k in range(18):\n        \n    m = f1_score(true[k].values, (oof[k].values>best_threshold).astype('int'), average='macro')\n    print(f'Q{k}: F1 =',m)\n    \nm = f1_score(true.values.reshape((-1)), (oof.values.reshape((-1))>best_threshold).astype('int'), average='macro')\nprint('==> Overall F1 =',m)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:05.543363Z","iopub.execute_input":"2023-02-15T13:06:05.543889Z","iopub.status.idle":"2023-02-15T13:06:05.704253Z","shell.execute_reply.started":"2023-02-15T13:06:05.543857Z","shell.execute_reply":"2023-02-15T13:06:05.702943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:05.705771Z","iopub.execute_input":"2023-02-15T13:06:05.706273Z","iopub.status.idle":"2023-02-15T13:06:05.860939Z","shell.execute_reply.started":"2023-02-15T13:06:05.706091Z","shell.execute_reply":"2023-02-15T13:06:05.859620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"limits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (sample_submission, test) in iter_test:\n    \n    df = feature_engineer(test)\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[:,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int(p.item()>best_threshold)\n    \n    env.predict(sample_submission)","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:05.862019Z","iopub.status.idle":"2023-02-15T13:06:05.862969Z","shell.execute_reply.started":"2023-02-15T13:06:05.862731Z","shell.execute_reply":"2023-02-15T13:06:05.862755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv('submission.csv')\nprint( df.shape )\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:05.864369Z","iopub.status.idle":"2023-02-15T13:06:05.865035Z","shell.execute_reply.started":"2023-02-15T13:06:05.864827Z","shell.execute_reply":"2023-02-15T13:06:05.864847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df.correct.mean())","metadata":{"execution":{"iopub.status.busy":"2023-02-15T13:06:05.866391Z","iopub.status.idle":"2023-02-15T13:06:05.867427Z","shell.execute_reply.started":"2023-02-15T13:06:05.867186Z","shell.execute_reply":"2023-02-15T13:06:05.867208Z"},"trusted":true},"execution_count":null,"outputs":[]}]}