{"cells":[{"metadata":{},"cell_type":"markdown","source":"The method introduced in [CV strategy](https://www.kaggle.com/its7171/cv-strategy) is very helpful, but it seems that the users are slightly biased."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import gc\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ntrain = pd.read_pickle('../input/riiid-cross-validation-files/cv1_train.pickle')\nvalid = pd.read_pickle('../input/riiid-cross-validation-files/cv1_valid.pickle')\ntrain = pd.concat([train, valid])\ndel valid\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"interval = 10000000\nuser_count = []\nfor i in range(10):\n    start = i * interval\n    user_count.append(train[start:start+interval].user_id.nunique())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(list(range(10)), user_count)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The bar plot above shows how many users exist in each of the 10 million. It seems that there are many users in the first and last 10 million. Something is wrong."},{"metadata":{"trusted":true},"cell_type":"code","source":"last10m = train[-interval:]\ninterval = 1000000\nuser_count = []\nfor i in range(10):\n    start = i * interval\n    user_count.append(last10m[start:start+interval].user_id.nunique())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(list(range(10)), user_count)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The bar plot above shows how many users exist in the last 10 million, separated by 1 million. It seems that there are many users in the last million. Why is this happening?"},{"metadata":{"trusted":true},"cell_type":"code","source":"max_timestamp_u = train[['user_id','timestamp']].groupby(['user_id']).agg(['max']).reset_index()\nmax_timestamp_u.columns = ['user_id', 'max_time_stamp']\nmax_timestamp_u.set_index('user_id', inplace=True)\nmax_timestamp_u.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"last100m = train[-100000000:]\ninterval = 2500000\nmean_max_timestamp = []\nfor i in range(40):\n    start = i * interval\n    user_list = last100m[start:start+interval].user_id.unique()\n    mean_max_timestamp.append(max_timestamp_u.reindex(user_list).mean())\nmean_max_timestamp = pd.concat(mean_max_timestamp)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(list(range(40)), mean_max_timestamp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The bar plot above is the average of the user's maximum timestamps, separated by 2.5 million of the last 100 million. It seems that the sooner and the later users have shorter time stamps."},{"metadata":{},"cell_type":"markdown","source":"Many users tackle many problems at the beginning, but the number of problems tends to decrease after that. Maybe that's the cause."},{"metadata":{},"cell_type":"markdown","source":"You should be aware that there are differences between users of training data and validation data."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}