{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Time Series EDA - Users and Real Sessions\nIn Kaggle's Otto competition the word \"session\" actually means \"user\". In this notebook, we display users and their real sessions time series EDA. We observe that users exhibit regular patterns of session behavior. These observations can help us characterize and engineer features for users. These observations can also give us insight into predicting future `click`, `cart`, and `order` behavior. We will use `RAPIDS cuDF` to process dataframes and `matplotlib` to display EDA. There is a Kaggle discussion about this notebook [here][1]\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138","metadata":{}},{"cell_type":"markdown","source":"# Load Libraries and Train Data","metadata":{}},{"cell_type":"code","source":"# LOAD LIBRARIES\nimport pandas as pd, numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as mpatches\nimport cudf, cupy\nprint('Using RAPIDS version',cudf.__version__)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-15T09:19:18.459061Z","iopub.execute_input":"2022-11-15T09:19:18.459728Z","iopub.status.idle":"2022-11-15T09:19:18.466002Z","shell.execute_reply.started":"2022-11-15T09:19:18.459692Z","shell.execute_reply":"2022-11-15T09:19:18.464949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# LOAD TRAIN DATA. RANDOM SAMPLE 10%\ntrain = cudf.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\nsessions = train.session.unique()\nsample = cupy.random.choice(sessions,len(sessions)//10,replace=False)\ntrain = train.loc[train.session.isin(sample)]\nprint('We are using random 1/10 of users. Truncated train data has shape', train.shape )\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-15T09:19:18.467853Z","iopub.execute_input":"2022-11-15T09:19:18.468831Z","iopub.status.idle":"2022-11-15T09:19:21.167907Z","shell.execute_reply.started":"2022-11-15T09:19:18.468795Z","shell.execute_reply":"2022-11-15T09:19:21.166849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# MIN AND MAX TRAIN DATES\n# IF USING ORIGINAL CSV, USE \"TS * 1e6\" BELOW\ntrain.ts = cudf.to_datetime(train.ts * 1e9)\nprint('Train min date and max date are:', train.ts.min(),'and', train.ts.max() )\nprint('We will truncate train data to begin Aug 1st, 2022')\ntrain = train.loc[train.ts >= cudf.to_datetime('2022-08-01')]","metadata":{"execution":{"iopub.status.busy":"2022-11-15T09:19:21.169631Z","iopub.execute_input":"2022-11-15T09:19:21.170028Z","iopub.status.idle":"2022-11-15T09:19:21.293440Z","shell.execute_reply.started":"2022-11-15T09:19:21.169990Z","shell.execute_reply":"2022-11-15T09:19:21.292452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# COMPUTE DAY AND HOUR OF ACTIVITY\ntrain['day'] = train.ts.dt.day\ntrain['hour'] = train.ts.dt.hour\ntrain = train.reset_index(drop=True)\n# THE NEXT TWO LINES REPLICATE GROUPBY TRANSFORM\ntmp = train.groupby('session').aid.agg('count').rename('n')\ntrain = train.merge(tmp,on='session')\nfrequent_users = cupy.asnumpy( train.loc[train.n>40,'session'].unique() )\nprint(f\"There are {len(frequent_users)} users whom each have over 40 item interactions in our truncated train data sample.\")\nprint(\"We will display 128 of these most active users' behavior below.\")","metadata":{"execution":{"iopub.status.busy":"2022-11-15T09:19:21.295718Z","iopub.execute_input":"2022-11-15T09:19:21.296309Z","iopub.status.idle":"2022-11-15T09:19:21.573016Z","shell.execute_reply.started":"2022-11-15T09:19:21.296272Z","shell.execute_reply":"2022-11-15T09:19:21.571936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute the Real User Sessions\nKaggle's version of RAPIDS cuDF does not contain `groupby` `diff` so we replicate it below. We find real sessions by finding gaps of 2 hours or more between activitiy. Then we label these gaps as `1` and everything else as `0`. Lastly we perform a groupby `cumsum` to locate sessions.","metadata":{}},{"cell_type":"code","source":"# COMPUTE USER REAL SESSIONS\ntrain.ts = train.ts.astype('int64')/1e9\n# THE NEXT THREE LINES REPLICATE GROUPBY DIFF\ntrain = train.sort_values(['session','ts']).reset_index(drop=True)\ntrain['d'] = train.ts.diff()\ntrain.loc[ train.session.diff()!=0, 'd'] = 0\n# IDENTITY REAL USER SESSIONS WHEN WE SEE 2 HOUR PAUSE IN ACTIVITY\ntrain.d = (train.d > 60*60*2).astype('int8').fillna(0)\ntrain['d'] = train.groupby('session').d.cumsum()","metadata":{"execution":{"iopub.status.busy":"2022-11-15T09:19:21.574408Z","iopub.execute_input":"2022-11-15T09:19:21.574848Z","iopub.status.idle":"2022-11-15T09:19:21.946904Z","shell.execute_reply.started":"2022-11-15T09:19:21.574813Z","shell.execute_reply":"2022-11-15T09:19:21.945925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(train.d.to_array(), bins=100)\nplt.title(\"Histogram of Train Users' Real Session Count\")\nm = train.d.mean()\nprint(f'The mean session count per train user is {m:0.1f} with right skewed distribution below')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-15T09:19:21.948199Z","iopub.execute_input":"2022-11-15T09:19:21.949217Z","iopub.status.idle":"2022-11-15T09:19:22.738162Z","shell.execute_reply.started":"2022-11-15T09:19:21.949175Z","shell.execute_reply":"2022-11-15T09:19:22.737253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Display User and Sessions Time Series\nBelow we display a scatter plot with jitter. The x axis is day of the month August 2022. And the y axis is hour of the day. Many dots would fall on top of each other, so we add random x and y jitter. Also we color the `clicks` blue, `carts` orange, and `orders` red. We plot the `clicks` first, then `carts`, then `orders`. This guarentees that the `orders` and `carts` (when present) will always be visible and not be obscured by `click` dots.","metadata":{}},{"cell_type":"code","source":"# DISPLAY USER ACTIVITY\ncolors = np.array( [(0,0,1),(1,0.5,0),(1,0,0)] )\n\nfor k in range(128):\n    u = np.random.choice(frequent_users)\n    tmp = train.loc[train.session==u].to_pandas()\n    \n    ss = tmp.d.max()+1\n    ii = len(tmp)\n    \n    plt.figure(figsize=(20,5))\n    for j in [0,1,2]:\n        s = 25\n        if j==1: s=50\n        elif j==2: s=100\n        tmp2 = tmp.loc[tmp['type']==j]\n        xx = np.random.uniform(-0.3,0.3,len(tmp2))\n        yy = np.random.uniform(-0.5,0.5,len(tmp2))\n        plt.scatter(tmp2.day.values+xx, tmp2.hour.values+yy, s=s, c=colors[tmp2['type'].values])\n    plt.ylim((0,24))\n    plt.xlim((0,30))\n    c1 = mpatches.Patch(color=colors[0], label='Click')\n    c2 = mpatches.Patch(color=colors[1], label='Cart')\n    c3 = mpatches.Patch(color=colors[2], label='Order')\n    plt.plot([0,30],[6-0.5,6-0.5],'--',color='gray')\n    plt.plot([0,30],[21+0.5,21+0.5],'--',color='gray')\n    for k in range(0,30):\n        plt.plot([k+0.5,k+0.5],[0,24],'--',color='gray')\n    for k in range(1,5):\n        plt.plot([7*k+0.5,7*k+0.5],[0,24],'--',color='black')\n    plt.legend(handles=[c1,c2,c3])\n    plt.xlabel('Day of August 2022',size=16)\n    plt.xticks([1,5,10,15,20,25,29],['Mon\\nAug 1st','Fri\\nAug 5th','Wed\\nAug 10th','Mon\\nAug 15th','Sat\\nAug 20th','Thr\\nAug 25th','Mon\\nAug 29th'])\n    plt.ylabel('Hour of Day',size=16)\n    plt.yticks([0,4,8,12,16,20,24],['midnight','4am','8am','noon','4pm','8pm','midnight'])\n    plt.title(f'User {u} has {ss} real sessions with {ii} item interactions',size=18)\n    plt.show()\n    print('\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2022-11-15T09:19:22.740439Z","iopub.execute_input":"2022-11-15T09:19:22.741165Z","iopub.status.idle":"2022-11-15T09:20:12.547924Z","shell.execute_reply.started":"2022-11-15T09:19:22.741125Z","shell.execute_reply":"2022-11-15T09:20:12.546988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Observations\nWe observe many patterns above. Most users exhibit regular behavior. They `click`, `cart` and `order` at the same hours each day. Also most users like to shop on the same days of the each week. Most users are active during the waking hours of day but some users like to shop during the night while others are sleeping. We also notice that users shop in clusters of activity. Our challenge in this competition is that we must both predict the remainder of the last cluster (provided in test data) and predict new clusters (after last timestamp in test). Furthermore all users in test data (not displayed in this notebook) have less than 1 week data, so we must predict user behavior given little user history information (i.e. the RecSys \"cold start\" problem). Understanding users and their behavior will help us predict test users' future behavior!","metadata":{}}]}