{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Investigation into some potential properties of test set.\nInspired by [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899) discussion.\n* are there any new questions in test (NO)\n* are there any new users in test (YES - new users with timestamp 0)\n* timeframe of test set? (FOLLOWING TRAIN for any given user)"},{"metadata":{"trusted":true},"cell_type":"code","source":"import riiideducation\nimport pandas as pd\n\nenv = riiideducation.make_env()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Training data is in the competition dataset as usual\nIt's larger than will fit in memory with default settings, so we'll specify more efficient datatypes and only load a subset of the data for now."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', low_memory=False, nrows=10**5, \n                       dtype={'row_id': 'int64', 'timestamp': 'int64', 'user_id': 'int32', 'content_id': 'int16', 'content_type_id': 'int8',\n                              'task_container_id': 'int16', 'user_answer': 'int8', 'answered_correctly': 'int8', 'prior_question_elapsed_time': 'float32', \n                             'prior_question_had_explanation': 'boolean',\n                             }\n                      )\ntrain_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"users = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', sep=',', usecols=['user_id', 'timestamp'], squeeze=True)\n#this takes few minutes - reading in the entire set of users","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"users_with_latest_ts = users.groupby('user_id')['timestamp'].max()\n#get the latest timestamp for all train users","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#create set for comparision to the test set\nuser_set = set(users.user_id.unique())\nlen(user_set)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 393656 unique users in train set. We will later compare if test API returns any new users, not already present in train."},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/questions.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df.question_id.max() + 1 == questions_df.shape[0] \nquestions_df.shape[0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 13523 unique questions in questions.cvs. We will later check if test API returns any new questions. "},{"metadata":{},"cell_type":"markdown","source":"## Iterate through example test set. \n\nFollowing example notebook, getting the example test set. "},{"metadata":{"trusted":true},"cell_type":"code","source":"iter_test = env.iter_test()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's get the data for the first test batch and check it out."},{"metadata":{"trusted":true},"cell_type":"code","source":"(test_df, sample_prediction_df) = next(iter_test)\ntest_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#get users and timestamps\ntest_users_and_ts = test_df[['user_id','timestamp']]\ntest_users_and_ts.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#work with sets to create a set of unique users and questions returned by test API\nquestion_ids = set(test_df.content_id.unique())\nnew_ids = set(test_df.content_id.unique())\nquestion_ids = question_ids.union(new_ids)\n\nuser_ids = set(test_df.user_id.unique())\nnew_users = set(test_df.user_id.unique())\nuser_ids = user_ids.union(new_users)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"env.predict(sample_prediction_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Main Loop\nLet's loop through all the remaining batches in the test set generator and make the default prediction for each. \n\nLet's store all users, timestamps, content_id (questions) and check them for novelty."},{"metadata":{"trusted":true},"cell_type":"code","source":"for (test_df, sample_prediction_df) in iter_test:\n    new_ids = set(test_df.content_id.unique())\n    question_ids = question_ids.union(new_ids)\n    \n    new_users = set(test_df.user_id.unique())\n    user_ids = user_ids.union(new_users)\n    \n    print(\"Length of test set {}, unique users {}\".format(len(test_df), len(new_ids)))\n    \n    test_users_and_ts_i = test_df[['user_id','timestamp']]\n    test_users_and_ts = pd.concat([test_users_and_ts,test_users_and_ts_i])\n    #print(test_users_and_ts.shape)\n    \n    test_df['answered_correctly'] = 0.5\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_set_min_ts = test_users_and_ts.groupby('user_id')['timestamp'].min().reset_index()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.merge(users_with_latest_ts,test_set_min_ts, on = 'user_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"if any(df['timestamp_y'] < df['timestamp_x']): \n    print(\"USER INTERACTION IN TEST SET HAS HAPPENED _BEFORE_ THE LATEST INTERACTION IN TRAIN SET. TIME MIXUP DETECTED\")\nelse:\n    print(\"ALL CLEAR, TEST SET ACTIONS FOLLOWED TRAIN SET ACTIONS FOR ANY GIVEN USER WHO WAS PRESENT IN BOTH\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(user_ids - user_set, \"these users are new\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(question_ids - set(questions_df.question_id), \"these questions are new\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"new_users_are_really_new = test_users_and_ts[test_users_and_ts.user_id.isin(user_ids - user_set)].groupby('user_id')['timestamp'].min().reset_index()\nif new_users_are_really_new.timestamp.max() > 0:\n    print(\"new user detected in test who is not really new! (timestamp is not 0)\")\n    print(new_users_are_really_new[new_users_are_really_new.timestamp>0])\nelse:\n    print(\"ALL CLEAR. NEW USERS IN TEST ARE INDEED NEW - test contains their first interaction and possibly more\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"new_users_are_really_new","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}