{"cells":[{"metadata":{},"cell_type":"markdown","source":"This notebook augments the Riiid input dataset with new columns: `this_question_had_explanation` and `this_question_elapsed_time`. These are essentially shifted and ffilled versions of `prior_question_had_explanation` and `prior_question_elapsed_time` from the orinal dataset.\nThey tracks whether **this** question provided feedback to the user and how long did **this** question take, rather than the previous one.\n\nI'm using alternative input/output formats based on info in [this notebook](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets/data)"},{"metadata":{},"cell_type":"markdown","source":"# Now less wrong!\n\nI've switched from using Pandas to datatable, because Pandas can't do what needs to be done without running out of memory.\n\nThis enabled me to fix two bugs:\n\n- Group by user before shifting prev_* to make this_*\n- Actually join this_* information properly so that it applies to each row, not just the first row per (user_id, task_container_id) combination."},{"metadata":{},"cell_type":"markdown","source":"## Read in training data from .jay file to datatable\n\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install datatable==0.11.0 > /dev/null\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport dask.dataframe as dd\nimport gc\nimport pyarrow.parquet as pq\nimport pyarrow\nimport datatable as dt\nfrom datatable import f\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\ndt_data = dt.fread(\"../input/riiid-train-data-multiple-formats/riiid_train.jay\")\ndt_data.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Compute question_elapsed_time and question_had_explanation for *this* question \n\n\nThis value is the same within each bundle of questions(set of questions asked of a given user with a given `task_container_id`). So, the whole bundle either got explanations or didn't; and the elapsed time is averaged over the bundle. To get explanation/elapsed time for *this* bundle we get the applicable bundles (those that represent groups of questions, not lectures) in the right order, then shift `prev_*` backwards. \n\nThen, for `had_explanation`, ffill the NaNs. (the only time that NaNs need to be filled is for the last question, beause there was no `prev_*` to get the data from; ffill makes sense in this context for had_explanation, since the only time users generally *dont't* get an explanation is at the beginning, when they are being asked diagnostic questions.)\n\n\nJudging by user 115 (for whom the order of `timestamp` and `task_container_id` do not match), it is `timestamp` which determines what the \"previous\" task container was: the first task container should have `prior_question_had_explanation == None`, and this is true for `timestamp` 0, not `task_container_id` 0."},{"metadata":{"trusted":true},"cell_type":"code","source":"# make a single column which contains a unique id for each [user_id, task_container_id] combination\ndt_data['tc_id'] = dt_data[:,dt.str32(f['user_id'])+'_'+f['task_container_id']]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\n# make a separate frame with just questions, and one row per tc_id\nquestions = dt_data[f[\"content_type_id\"]==0,:]\nq_task_containers = questions[\n    (f['tc_id']!=dt.shift(f['tc_id'])) , :]\nq_task_containers.shape  # expected shape: (76483597, 11)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# this_* is prior_* shifted once by tc_id, within user\nq_task_containers['this_question_elapsed_time'] = q_task_containers[:,dt.shift(f['prior_question_elapsed_time'], -1),dt.by(f['user_id'])]['prior_question_elapsed_time']\nq_task_containers['this_question_had_explanation'] = q_task_containers[:,dt.shift(f['prior_question_had_explanation'], -1),dt.by(f['user_id'])]['prior_question_had_explanation']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# sanity check - this_* values are null in the last row for a user\nq_task_containers[f['user_id']==115, :]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Put computed columns back into training data"},{"metadata":{"trusted":true},"cell_type":"code","source":"# drop all columns except the newest ones\nq_task_containers = q_task_containers[:,'tc_id':]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\n# key by tc_id for joining\nq_task_containers.key='tc_id'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\n# left outer join into original data table\ndt_data = dt_data[:,:,dt.join(q_task_containers)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Sanity check: last 3 task_container_id bundles for a particular user\ndt_data[(f['user_id']==2147012157) & (f['task_container_id'] > 4675),:]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Write back out "},{"metadata":{"trusted":true},"cell_type":"code","source":"# Make room in RAM\ndel q_task_containers\ngc.collect()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\ndt_data.to_csv('riiid_train_with_qdata.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\ndt_data.to_jay('riiid_train_with_qdata.jay')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}