{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"The goal of this competition is to predict student performance during game-based learning in real-time. You'll develop a model trained on one of the largest open datasets of game logs.\n\nYour work will help advance research into knowledge-tracing methods for game-based learning. You'll be supporting developers of educational games to create more effective learning experiences for students.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h1>Data Descriptions </h1>\n\nLet's jump straight into the data analysis process and see what we got here, But remember, there are no full resources so we have to be aware of it.","metadata":{}},{"cell_type":"code","source":"from collections import Counter\nimport pandas as pd\nimport numpy as np \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nplt.rcParams['figure.dpi'] = 100","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:11:24.347603Z","iopub.execute_input":"2023-02-06T21:11:24.348047Z","iopub.status.idle":"2023-02-06T21:11:24.353726Z","shell.execute_reply.started":"2023-02-06T21:11:24.348011Z","shell.execute_reply":"2023-02-06T21:11:24.352736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Train Dataset </h1>\n","metadata":{}},{"cell_type":"code","source":"from typing import List, Dict, Union\nfrom typing import Any, TypeVar\n\nroot_dir = '/kaggle/input/predict-student-performance-from-game-play/'\n\ntrain = pd.read_csv(root_dir + 'train.csv')\n#test = pd.read_csv(root_dir + 'test.csv')\n#train_labels ","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:00:15.955449Z","iopub.execute_input":"2023-02-06T21:00:15.955862Z","iopub.status.idle":"2023-02-06T21:01:20.868778Z","shell.execute_reply.started":"2023-02-06T21:00:15.955815Z","shell.execute_reply":"2023-02-06T21:01:20.867796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h1> Surprise </h1>\n\nWow, almost 13kk observations and 20 features. I feel like we came back in time to 2010, especially with our VM performance ;)","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:24:26.096299Z","iopub.execute_input":"2023-02-06T21:24:26.096731Z","iopub.status.idle":"2023-02-06T21:24:26.104207Z","shell.execute_reply.started":"2023-02-06T21:24:26.096692Z","shell.execute_reply":"2023-02-06T21:24:26.102883Z"}}},{"cell_type":"markdown","source":"To properly perform this analysis I decided to sample 20% of observations from train dataset. ","metadata":{}},{"cell_type":"code","source":"import random\n_train_obs: int = train.shape[0]\n_n_obs: int = int(_train_obs * 0.2)\nrandom.sample\nixes: List = random.sample(range(_train_obs), _n_obs)\nsample_train = train.iloc[ixes, :]","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:02:43.968786Z","iopub.execute_input":"2023-02-06T21:02:43.969275Z","iopub.status.idle":"2023-02-06T21:02:53.338848Z","shell.execute_reply.started":"2023-02-06T21:02:43.969240Z","shell.execute_reply":"2023-02-06T21:02:53.337886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc; \ngc.collect();\ndel train","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:05:59.638519Z","iopub.execute_input":"2023-02-06T21:05:59.638943Z","iopub.status.idle":"2023-02-06T21:06:00.014529Z","shell.execute_reply.started":"2023-02-06T21:05:59.638889Z","shell.execute_reply":"2023-02-06T21:06:00.012200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_train.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:04:24.050243Z","iopub.execute_input":"2023-02-06T21:04:24.050800Z","iopub.status.idle":"2023-02-06T21:04:24.059096Z","shell.execute_reply.started":"2023-02-06T21:04:24.050766Z","shell.execute_reply":"2023-02-06T21:04:24.058040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:05:39.193109Z","iopub.execute_input":"2023-02-06T21:05:39.193451Z","iopub.status.idle":"2023-02-06T21:05:39.218368Z","shell.execute_reply.started":"2023-02-06T21:05:39.193420Z","shell.execute_reply":"2023-02-06T21:05:39.217062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_train.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:08:05.384587Z","iopub.execute_input":"2023-02-06T21:08:05.384970Z","iopub.status.idle":"2023-02-06T21:08:05.397610Z","shell.execute_reply.started":"2023-02-06T21:08:05.384938Z","shell.execute_reply":"2023-02-06T21:08:05.396673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Variables","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h1>2.2 Elapsed Time ","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (8, 8))\nsns.distplot(sample_train.elapsed_time[::100])\n\nplt.grid()\nplt.title(\"Elapsed time - distribution plot\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:13:49.720553Z","iopub.execute_input":"2023-02-06T21:13:49.720955Z","iopub.status.idle":"2023-02-06T21:13:50.109183Z","shell.execute_reply.started":"2023-02-06T21:13:49.720923Z","shell.execute_reply":"2023-02-06T21:13:50.108194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n\n<h3>\n    Lets look at some quantile of this variables\n</h3>\n\n","metadata":{}},{"cell_type":"code","source":"\nnp.quantile(sample_train.elapsed_time, [0.05, 0.1, 0.15, 0.25, 0.5, 0.95])","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:18:39.744603Z","iopub.execute_input":"2023-02-06T21:18:39.745142Z","iopub.status.idle":"2023-02-06T21:18:39.819826Z","shell.execute_reply.started":"2023-02-06T21:18:39.745084Z","shell.execute_reply":"2023-02-06T21:18:39.818946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\nSo we can see that first 5% of observation is less or equal than 73557, and last 5% of observation is greater than 5.137kk. thats a lot o variance to catch! ^^\n","metadata":{}},{"cell_type":"code","source":"print(f\"Maximum value {np.max(sample_train.elapsed_time)}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:20:42.962722Z","iopub.execute_input":"2023-02-06T21:20:42.963409Z","iopub.status.idle":"2023-02-06T21:20:42.972366Z","shell.execute_reply.started":"2023-02-06T21:20:42.963371Z","shell.execute_reply":"2023-02-06T21:20:42.971459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h4>\nDistribution without first and last 5% of values \n    </h4>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (8, 8))\n\nsns.displot(sample_train.elapsed_time[(sample_train.elapsed_time > 73557) & (sample_train.elapsed_time <5137548.7 )][::100]) \n\nplt.grid()\nplt.title(\"Distribution plot of v > q5 and v < q95\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:23:46.767743Z","iopub.execute_input":"2023-02-06T21:23:46.768630Z","iopub.status.idle":"2023-02-06T21:23:47.186706Z","shell.execute_reply.started":"2023-02-06T21:23:46.768587Z","shell.execute_reply":"2023-02-06T21:23:47.184382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h4> Elapsed time - interpretation </h4>\nSo we can observe that elapsed time is mostly focus on 0 values, but what is more concerning is that there exis\nobservations with values greater than 1e9.\nThat's very long game session ;)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h2>2.3 Event Name","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n\nNext variable, name of the event is an ```object``` type variable, which can be easy transformed to string variable. \nAt the very first insight, we can just check for number of levels, and frequency for every level of this variable. \n\nIt is worth to mention that at this stage of the analysis I dont focus on discrepancy of distribution between \ndifferent levels of response variable. \nThat kind of analysis will be covered (hope so ^^) in next iterations.","metadata":{}},{"cell_type":"code","source":"unique_event_name: np.ndarray = np.unique(sample_train.event_name)\nlen(unique_event_name)","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:39:51.236708Z","iopub.execute_input":"2023-02-06T21:39:51.237133Z","iopub.status.idle":"2023-02-06T21:39:53.884042Z","shell.execute_reply.started":"2023-02-06T21:39:51.237101Z","shell.execute_reply":"2023-02-06T21:39:53.882880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So we have 11 unique event names, now we can just check for its frequency :)","metadata":{}},{"cell_type":"code","source":"Counter(sample_train.event_name.iloc[::100].values)","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:41:47.325723Z","iopub.execute_input":"2023-02-06T21:41:47.326413Z","iopub.status.idle":"2023-02-06T21:41:47.340116Z","shell.execute_reply.started":"2023-02-06T21:41:47.326376Z","shell.execute_reply":"2023-02-06T21:41:47.339009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (6, 4))\n\nsns.countplot(sample_train.event_name[::100])\nplt.grid()\nplt.title(\"Event name, 1% of population count plot\")\nplt.xticks(rotation=45)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:43:16.856559Z","iopub.execute_input":"2023-02-06T21:43:16.856985Z","iopub.status.idle":"2023-02-06T21:43:17.097235Z","shell.execute_reply.started":"2023-02-06T21:43:16.856949Z","shell.execute_reply":"2023-02-06T21:43:17.096450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n    <h4>Event Name -  Interpretation </h4>\n    \n   So we can observe that wast amount of observation is taken from 2 levels \n    \n<li> - navigate click </li>\n<li> - object hover </li>\n    \nThe most informative part will be to see its distribution and strength of different level on response variable","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h1>2.4 Name","metadata":{}},{"cell_type":"code","source":"unique_name: np.ndarray = np.unique(sample_train.name)\nlen(unique_event_name)","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:56:00.073277Z","iopub.execute_input":"2023-02-06T21:56:00.073749Z","iopub.status.idle":"2023-02-06T21:56:02.652512Z","shell.execute_reply.started":"2023-02-06T21:56:00.073715Z","shell.execute_reply":"2023-02-06T21:56:02.651512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (6, 4))\n\nsns.countplot(sample_train.name)\nplt.grid()\nplt.title(\"Name, 1% of population count plot\")\nplt.xticks(rotation=45)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-06T21:57:35.312707Z","iopub.execute_input":"2023-02-06T21:57:35.313119Z","iopub.status.idle":"2023-02-06T21:57:37.294056Z","shell.execute_reply.started":"2023-02-06T21:57:35.313086Z","shell.execute_reply":"2023-02-06T21:57:37.293060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n    <h4>Name -  Interpretation </h4>\n    \n   So we can observe that wast amount of observation is taken from 2 levels \n    \n<li> - undefine </li>\n<li> - basic </li>\n    \nThe most informative part will be to see its distribution and strength of different level on response variable\n    \nAlso, it is interesting to observe and deduce, why ```undefined``` level is the most frequent, and if it has some \nadditive value.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h1>2.5 Page","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:08:04.588720Z","iopub.execute_input":"2023-02-06T22:08:04.589226Z","iopub.status.idle":"2023-02-06T22:08:04.595440Z","shell.execute_reply.started":"2023-02-06T22:08:04.589190Z","shell.execute_reply":"2023-02-06T22:08:04.594090Z"}}},{"cell_type":"markdown","source":"There exist a huge amount of na in this datasets, so let's look at this!\n","metadata":{}},{"cell_type":"code","source":"np.sum(sample_train.page.isna())/sample_train.shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:16:05.638041Z","iopub.execute_input":"2023-02-06T22:16:05.638538Z","iopub.status.idle":"2023-02-06T22:16:05.651560Z","shell.execute_reply.started":"2023-02-06T22:16:05.638501Z","shell.execute_reply":"2023-02-06T22:16:05.650207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"sample_train.page.value_counts()/sample_train.shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:14:33.271742Z","iopub.execute_input":"2023-02-06T22:14:33.272937Z","iopub.status.idle":"2023-02-06T22:14:33.288936Z","shell.execute_reply.started":"2023-02-06T22:14:33.272874Z","shell.execute_reply":"2023-02-06T22:14:33.287997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (6, 4))\n\nsns.countplot(sample_train.page)\nplt.grid()\nplt.title(\"Page, count plot\")\nplt.xticks(rotation=45)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:10:20.588259Z","iopub.execute_input":"2023-02-06T22:10:20.588692Z","iopub.status.idle":"2023-02-06T22:10:20.902907Z","shell.execute_reply.started":"2023-02-06T22:10:20.588659Z","shell.execute_reply":"2023-02-06T22:10:20.902036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n    <h4>Page -  Interpretation </h4>\n    \n  97% of observations are empty values. And there is one question which bothered me, 'Why' ? ;d\n\nAt the very first sight, there is not much of informative value in this characteristics, but, we'll extend it into discrepancy between response levels. ","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:17:48.181488Z","iopub.execute_input":"2023-02-06T22:17:48.181989Z","iopub.status.idle":"2023-02-06T22:17:48.189627Z","shell.execute_reply.started":"2023-02-06T22:17:48.181948Z","shell.execute_reply":"2023-02-06T22:17:48.188190Z"}}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_: np.ndarray = np.unique(sample_train.name)\nlen(unique_event_name)","metadata":{"execution":{"iopub.status.busy":"2023-02-06T20:56:12.764905Z","iopub.execute_input":"2023-02-06T20:56:12.765252Z","iopub.status.idle":"2023-02-06T20:56:12.773335Z","shell.execute_reply.started":"2023-02-06T20:56:12.765223Z","shell.execute_reply":"2023-02-06T20:56:12.772501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h1> Metric </h1>\n\nWe have to deal with binary classification task, while the main metric is the F1 Score, which can be interpreted as below: \n\n![F1 Score](https://wikimedia.org/api/rest_v1/media/math/render/svg/f5c869c51dba6f1df65a6e6630c516de161632d4)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-family:verdana; word-spacing:1.5px;\">\n<h1>2.6 Hover Time","metadata":{"execution":{"iopub.status.busy":"2023-02-06T20:50:45.590793Z","iopub.execute_input":"2023-02-06T20:50:45.591250Z","iopub.status.idle":"2023-02-06T20:50:46.251942Z","shell.execute_reply.started":"2023-02-06T20:50:45.591162Z","shell.execute_reply":"2023-02-06T20:50:46.251076Z"}}},{"cell_type":"code","source":"plt.figure(figsize = (6, 4))\n\nsns.displot(sample_train.hover_duration)[::500]\n\nplt.grid()\nplt.title(\"Hover Time distribution plot\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:33:53.999500Z","iopub.status.idle":"2023-02-06T22:33:54.000113Z","shell.execute_reply.started":"2023-02-06T22:33:53.999900Z","shell.execute_reply":"2023-02-06T22:33:53.999922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect();","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:34:29.948225Z","iopub.execute_input":"2023-02-06T22:34:29.949331Z","iopub.status.idle":"2023-02-06T22:34:34.593997Z","shell.execute_reply.started":"2023-02-06T22:34:29.949287Z","shell.execute_reply":"2023-02-06T22:34:34.592970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#del unique_name\ndel unique_event_name","metadata":{"execution":{"iopub.status.busy":"2023-02-06T22:36:04.176546Z","iopub.execute_input":"2023-02-06T22:36:04.176977Z","iopub.status.idle":"2023-02-06T22:36:04.182254Z","shell.execute_reply.started":"2023-02-06T22:36:04.176941Z","shell.execute_reply":"2023-02-06T22:36:04.180902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Todo \n\nDifferences between distribution on difference predictors will be nice. (kullback-leibler, wasserstein, correlation, ml, desc stats)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}