{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction:\n\nHere's a technique that I'm using to simulate how my pipeline manages the new test data provided during submission. I've been using it to fine tune the feature engineering (FE) processess to ensure that the piple (including model prediction) on test groups don't take too long. \n\nYou can use this to check how much time your pipeline and model takes to parse the groups during sumbmission.\n\nFeel free to share techniques that have improved your FE pipeline in the comments. \n"},{"metadata":{},"cell_type":"markdown","source":"## Importing libraries"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport random\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm.notebook import  tqdm","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Reading Data\n\nI use only the first 10 million rows to create the simulation"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"%%time\n\ndata_type_dict = {'row_id': 'int64',\n                  'timestamp': 'int64',\n                  'user_id': 'int32',\n                  'content_id': 'int16',\n                  'content_type_id': 'int8',\n                  'task_container_id': 'int16',\n                  'user_answer': 'int8',\n                  'answered_correctly': 'int8',\n                  'prior_question_elapsed_time': 'float32', \n                  'prior_question_had_explanation': 'boolean'}\n\ntrain_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv',\n                       low_memory=True,\n                       dtype=data_type_dict,\n                       nrows = 10**7)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/questions.csv',index_col=0)\nquestions_df = questions_df.fillna(value={'tags':'-1'})\nquestions_df['correct_answer']=questions_df['correct_answer'].astype(np.int8)\nquestions_df['part']=questions_df['part'].astype(np.int8)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Spliting Data for simulation\n\nHere I truncate the the last 25th percentile of task_container_id for creating simulation data "},{"metadata":{"trusted":true},"cell_type":"code","source":"test_split = True\ntraining_set_ratio = 0.75","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"if test_split==True:\n    train_df = train_df.merge(pd.DataFrame(train_df.groupby('user_id')['task_container_id'].agg('max')).rename(columns={'task_container_id':'task_container_id_max'}),\n                          left_on='user_id',right_index=True)\n\n\n    train_df['for_training'] = train_df['task_container_id'].values<=training_set_ratio*train_df['task_container_id_max'].values\n\n    train_df = train_df.drop('task_container_id_max',axis=1)\n\n\n    test_df = train_df.loc[~train_df['for_training']]\n    train_df = train_df.loc[train_df['for_training']]\n\n    train_df = train_df.drop('for_training',axis=1)\n    test_df = test_df.drop('for_training',axis=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The code below assigns group numbers to the testing data while ensuring that each group only contains a single bundle for any user_id and that the task_container_id increase chronologically with group_no for each user.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['group_no']=0\n\nout_col_dict = {name:ii for ii,name in enumerate(test_df.columns)}\n\nfor ii in tqdm(range(1,test_df.shape[0])):\n    if (test_df.iat[ii,out_col_dict['user_id']]==test_df.iat[ii-1,out_col_dict['user_id']]):\n        if (test_df.iat[ii,out_col_dict['task_container_id']]!=test_df.iat[ii-1,out_col_dict['task_container_id']]):\n           test_df.iat[ii,out_col_dict['group_no']] = 1 + test_df.iat[ii-1,out_col_dict['group_no']]\n        else:\n            test_df.iat[ii,out_col_dict['group_no']] = test_df.iat[ii-1,out_col_dict['group_no']]\n\n    else:\n        test_df.iat[ii,out_col_dict['group_no']] = (ii//50_000)*100","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"columns_index_expected = ['row_id','group_no','timestamp', 'user_id', 'content_id', 'content_type_id',\n                           'task_container_id', 'prior_question_elapsed_time',\n                           'prior_question_had_explanation', 'prior_group_responses',\n                           'prior_group_answers_correct']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"groups = test_df.groupby('group_no')\n\ngroup_lengths = []\ntest_groups_list = []\nfor ii,frame in tqdm(groups):\n    group_lengths.append([ii,len(frame['user_id'].unique())])\n    test_groups_list.append(frame)\n    \nfor ii in tqdm(range(len(test_groups_list))):\n    test_groups_list[ii]['prior_group_answers_correct'] = np.nan\n    test_groups_list[ii]['prior_group_responses'] = np.nan \n    if ii>0:\n        test_groups_list[ii].loc[test_groups_list[ii].index[0],'prior_group_responses'] = str(list(test_groups_list[ii-1]['user_answer'].values.astype(np.int8)))\n        test_groups_list[ii].loc[test_groups_list[ii].index[0],'prior_group_answers_correct'] = str(list(test_groups_list[ii-1]['answered_correctly'].values.astype(np.int8)))\n        test_groups_list[ii-1].drop(['user_answer','answered_correctly'],axis=1,inplace=True)\n        test_groups_list[ii-1] = test_groups_list[ii-1][columns_index_expected]\n    else:\n        test_groups_list[ii].loc[test_groups_list[ii].index[0],'prior_group_responses'] = str('[]')\n        test_groups_list[ii].loc[test_groups_list[ii].index[0],'prior_group_answers_correct'] = str('[]')\n    \n    \ntest_groups_list[-1] = test_groups_list[-1][columns_index_expected]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_group_counts_df = pd.DataFrame(group_lengths,columns=['group_no','counts'])\n\nfig,ax=plt.subplots(1,1,figsize=(16,8))\nsns.distplot(test_group_counts_df['counts'],ax=ax,kde_kws={\"cut\":3});","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The compeitition host said that the number of bundles in one group of test data would range from 1 to 1000. the simulated data has a few groups that are a bit larger then 1000 but I believe that this wouldn't be much of an issue."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Intial groups have lots of bundles\n\ntemp = random.choice(test_groups_list[:10])\nprint(f'This Group has {len(temp[\"user_id\"].unique())} bundle/bundles')\ndisplay(temp)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Last groups have only a few bundles\n\ntemp = random.choice(test_groups_list[-10:])\nprint(f'This Group has {len(temp[\"user_id\"].unique())} bundle/bundles')\ndisplay(temp)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Testing Pipeline with Simulation\n\nYou can check how your pipline works on the test groups by editing the content of the for loop below."},{"metadata":{"trusted":true},"cell_type":"code","source":"for test_group_single in tqdm(test_groups_list):\n        test_group_single = test_group_single.merge(questions_df,how='left',left_on='content_id',right_index=True)\n        test_group_single['timestamp'] = (test_group_single['timestamp']//1000)\n\n        test_group_single['part'] =  test_group_single['part'].fillna(0).astype(dtype=np.int8)\n        test_group_single['bundle_id'] = test_group_single['bundle_id'].fillna(0).astype(dtype=np.uint16)\n        \n        for ii in range(test_group_single.shape[0]):\n            # make predictions here for each user_id or row_id\n            pass","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}