{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Visualize User Behavior with RAPIDS TSNE and Item Matrix Factorization\nIn this notebook, we will use RAPIDS TSNE together with item embeddings from item matrix factorization (thank you [CPMP][1] and [Radek][2] for your notebooks) to visualize users' activity behavior.\n\nIn Kaggle's Otto Recommender System competition, the item ids are anonymized. So we do not know which each item id refers to. However, with item matrix factorization we can convert the anonymus item ids into meaningful embeddings. Then similar embeddings will be similar items. If we project the embeddings into the 2D plane (with TSNE, UMAP, PCA etc) and plot them, we can see clusters of similar items. Then one cluster may be clothing and another cluster could be electronics, etc etc.\n\nUsing the 2D plane of item embeddings, we can draw dots and lines showing the progression of each users' activity. This allows us to understand how users shop. Do users browse similar items, then view different items, then return to buy the original items. Do users mainly explore similar items? Or do users explore all sorts of items? The plots below will illuminate user behavior and help us understand the data and build better features and models!\n\n**NOTE** The code for this notebook's item matrix factorization comes from CPMP's notebook [here][1] and Radek's notebook [here][2]. If you find my notebook helpful, please upvote their notebooks too!\n\nEnjoy!\n\n[1]: https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\n[2]: https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader","metadata":{}},{"cell_type":"markdown","source":"# Data Preprocessing\nWe will load and process data with RAPIDS cuDF","metadata":{}},{"cell_type":"code","source":"import cudf\nprint('RAPIDS cuDF version',cudf.__version__)\n\ntrain = cudf.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\ntest = cudf.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')\n\ntrain_pairs = cudf.concat([train, test])[['session', 'aid']]\ndel train, test\n\ntrain_pairs['aid_next'] = train_pairs.groupby('session').aid.shift(-1)\ntrain_pairs = train_pairs[['aid', 'aid_next']].dropna().reset_index(drop=True)\n\ncardinality_aids = max(train_pairs['aid'].max(), train_pairs['aid_next'].max())\nprint('Cardinality of items is',cardinality_aids)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-12-11T04:05:43.380045Z","iopub.execute_input":"2022-12-11T04:05:43.380441Z","iopub.status.idle":"2022-12-11T04:06:10.938452Z","shell.execute_reply.started":"2022-12-11T04:05:43.380364Z","shell.execute_reply":"2022-12-11T04:06:10.936516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install Merlin Dataloader!\nWe will feed our PyTorch model with Merlin dataloader!","metadata":{}},{"cell_type":"code","source":"!pip install merlin-dataloader==0.0.2\nfrom merlin.loader.torch import Loader \n\ntrain_pairs.to_pandas().to_parquet('train_pairs.parquet') # TRAIN WITH ALL DATA\ntrain_pairs[-10_000_000:].to_pandas().to_parquet('valid_pairs.parquet')\n\nfrom merlin.loader.torch import Loader \nfrom merlin.io import Dataset\n\ntrain_ds = Dataset('train_pairs.parquet')\ntrain_dl_merlin = Loader(train_ds, 65536, True)","metadata":{"_kg_hide-output":true,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-12-11T04:16:32.602818Z","iopub.execute_input":"2022-12-11T04:16:32.603425Z","iopub.status.idle":"2022-12-11T04:16:46.718120Z","shell.execute_reply.started":"2022-12-11T04:16:32.603382Z","shell.execute_reply":"2022-12-11T04:16:46.717087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Learn Item Embeddings with PyTorch Matrix Factorization Model\nWe will build a PyTorch model to generate item embeddings via matrix factorization.","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch import nn\n\nclass MatrixFactorization(nn.Module):\n    def __init__(self, n_aids, n_factors):\n        super().__init__()\n        self.aid_factors = nn.Embedding(n_aids, n_factors, sparse=True)\n        \n    def forward(self, aid1, aid2):\n        aid1 = self.aid_factors(aid1)\n        aid2 = self.aid_factors(aid2)\n        \n        return (aid1 * aid2).sum(dim=1)\n    \nclass AverageMeter(object):\n    \"\"\"Computes and stores the average and current value\"\"\"\n    def __init__(self, name, fmt=':f'):\n        self.name = name\n        self.fmt = fmt\n        self.reset()\n\n    def reset(self):\n        self.val = 0\n        self.avg = 0\n        self.sum = 0\n        self.count = 0\n\n    def update(self, val, n=1):\n        self.val = val\n        self.sum += val * n\n        self.count += n\n        self.avg = self.sum / self.count\n\n    def __str__(self):\n        fmtstr = '{name} {val' + self.fmt + '} ({avg' + self.fmt + '})'\n        return fmtstr.format(**self.__dict__)\n\nvalid_ds = Dataset('valid_pairs.parquet')\nvalid_dl_merlin = Loader(valid_ds, 65536, True)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-12-11T04:16:52.423489Z","iopub.execute_input":"2022-12-11T04:16:52.424190Z","iopub.status.idle":"2022-12-11T04:16:52.574838Z","shell.execute_reply.started":"2022-12-11T04:16:52.424156Z","shell.execute_reply":"2022-12-11T04:16:52.573881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from torch.optim import SparseAdam\n\nnum_epochs = 10\nlr=0.1\n\nmodel = MatrixFactorization(cardinality_aids+1, 32)\noptimizer = SparseAdam(model.parameters(), lr=lr)\ncriterion = nn.BCEWithLogitsLoss()\n\nmodel.to('cuda')\nfor epoch in range(num_epochs):\n    for batch, _ in train_dl_merlin:\n        model.train()\n        losses = AverageMeter('Loss', ':.4e')\n            \n        aid1, aid2 = batch['aid'], batch['aid_next']\n        aid1 = aid1.to('cuda')\n        aid2 = aid2.to('cuda')\n        output_pos = model(aid1, aid2)\n        output_neg = model(aid1, aid2[torch.randperm(aid2.shape[0])])\n        \n        output = torch.cat([output_pos, output_neg])\n        targets = torch.cat([torch.ones_like(output_pos), torch.zeros_like(output_pos)])\n        loss = criterion(output, targets)\n        losses.update(loss.item())\n        \n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        \n    model.eval()\n    \n    with torch.no_grad():\n        accuracy = AverageMeter('accuracy')\n        for batch, _ in valid_dl_merlin:\n            aid1, aid2 = batch['aid'], batch['aid_next']\n            output_pos = model(aid1, aid2)\n            output_neg = model(aid1, aid2[torch.randperm(aid2.shape[0])])\n            accuracy_batch = torch.cat([output_pos.sigmoid() > 0.5, output_neg.sigmoid() < 0.5]).float().mean()\n            accuracy.update(accuracy_batch, aid1.shape[0])\n            \n    print(f'{epoch+1:02d}: * TrainLoss {losses.avg:.3f}  * Accuracy {accuracy.avg:.3f}')","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-12-11T04:16:56.712476Z","iopub.execute_input":"2022-12-11T04:16:56.713597Z","iopub.status.idle":"2022-12-11T04:21:35.418295Z","shell.execute_reply.started":"2022-12-11T04:16:56.713530Z","shell.execute_reply":"2022-12-11T04:21:35.417243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Extract Item Embeddings\nWe extract item embeddings from our model's embedding table.","metadata":{}},{"cell_type":"code","source":"# EXTRACT EMBEDDINGS FROM MODEL EMBEDDING TABLE\nembeddings = model.aid_factors.weight.detach().cpu().numpy()\nprint('Item Matrix Factorization embeddings have shape',embeddings.shape)","metadata":{"execution":{"iopub.status.busy":"2022-12-11T04:21:35.420440Z","iopub.execute_input":"2022-12-11T04:21:35.421132Z","iopub.status.idle":"2022-12-11T04:21:35.618219Z","shell.execute_reply.started":"2022-12-11T04:21:35.421093Z","shell.execute_reply":"2022-12-11T04:21:35.617035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize User Behavior with RAPIDS TSNE\nWe will use RAPIDS cuML's TSNE to reduce embedding dimension to 2 dimensions so we can plot in x-y-plane.","metadata":{}},{"cell_type":"code","source":"# IMPORT RAPIDS TSNE\nfrom cuml import UMAP, TSNE, PCA\nimport matplotlib.pyplot as plt, numpy as np\nimport matplotlib.patches as mpatches, cuml\nprint('RAPIDS cuML version',cuml.__version__)\n\n# FIT TRANSFORM TSNE\nem_2d = TSNE(n_components=2).fit_transform(embeddings)\nprint('TSNE embeddings have shape',em_2d.shape)","metadata":{"execution":{"iopub.status.busy":"2022-12-11T04:21:35.619890Z","iopub.execute_input":"2022-12-11T04:21:35.620287Z","iopub.status.idle":"2022-12-11T04:24:07.568543Z","shell.execute_reply.started":"2022-12-11T04:21:35.620250Z","shell.execute_reply":"2022-12-11T04:24:07.567543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# LOAD TEST DATA\ntest = cudf.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')\ntmp = test.groupby('session').aid.agg('count').rename('n')\ntest = test.merge(tmp, left_on='session', right_index=True, how='left')\nactive_users = test.loc[test.n>20,'session'].unique().to_array()\ntest = test.sort_values(['session','ts'])\nprint('Test data shape:', test.shape )\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T04:24:07.571993Z","iopub.execute_input":"2022-12-11T04:24:07.573114Z","iopub.status.idle":"2022-12-11T04:24:08.000838Z","shell.execute_reply.started":"2022-12-11T04:24:07.573047Z","shell.execute_reply":"2022-12-11T04:24:07.999883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Display User Activity with TSNE Item Embeddings\nFor 50 random users (who have 20 or more activity), we will plot their activity over time. The top plots show the category of items interacted with and the bottom plot show the times of item interaction. When dots are close to each other in the top plot, then the items are similar. For example, two items of clothing will be close while an item of clothing will be far from an item of electronics.\n\nThe numbers in the plot are sorted by time. Number 1 is the first item a user interactied with while number 10 is the 10th item etc. We only display some numbers so that the numbers are readable. Orange dots are clicks, green dots are carts, and red dots are buys. The 1.8 million blue dots are all unique items. We observe clusters of blue dots which represent different categories of items.\n\nThe x-y-plane indicates different item categories. If users stay in the same area then they are shopping similar items like clothing department. When a user's plot changes from one region of x-y-plane to another, then the user changes to a different category of item like moving from clothing shopping to electronics shopping. We observe different types of users. Some users explore one category of item while other users explore a variety of item categories.","metadata":{}},{"cell_type":"code","source":"# DISPLAY EDA FOR 50 USERS\nfor k in range(50):\n    \n    # SELECT ONE USER WITH 20+ CLICKS\n    u = np.random.choice(active_users)\n    dff = test.loc[test.session==u].to_pandas().reset_index(drop=True)\n    tmp = dff.aid.values\n    clicks = test.loc[(test.session==u)&(test['type']==0)].to_pandas().aid.values\n    carts = test.loc[(test.session==u)&(test['type']==1)].to_pandas().aid.values\n    orders = test.loc[(test.session==u)&(test['type']==2)].to_pandas().aid.values\n\n    ############\n    ## PLOT HISTORY BY ITEM CATEGORY\n    ############\n    \n    # PLOT CLICKS, CARTS, ORDERS OVER TSNE ITEM EMBEDDING PLOT\n    plt.figure(figsize=(15,15))\n    plt.scatter(em_2d[::25,0],em_2d[::25,1],s=1,label='All 1.8M items!')\n    plt.plot(em_2d[tmp][:,0],em_2d[tmp][:,1],'-',color='orange')\n    plt.scatter(em_2d[tmp][:,0],em_2d[tmp][:,1],color='orange',s=25,label='Click')\n    plt.scatter(em_2d[carts][:,0],em_2d[carts][:,1],color='green',s=100,label='Cart')\n    plt.scatter(em_2d[orders][:,0],em_2d[orders][:,1],color='red',s=250,label='Order')\n    \n    # PLOT NUMBERS OF ORDER VISITED\n    old_xy = []; pos = []\n    for i,(x,y) in enumerate(zip(em_2d[tmp][:,0],em_2d[tmp][:,1])):\n        new_location = True\n        for j in old_xy:\n            if (np.abs(x-j[0])<5) & (np.abs(y-j[1])<5):\n                new_location = False\n        if new_location:\n            plt.text(x,y,f'{i+1}',size=18)\n            old_xy.append( (x,y) ); pos.append(i)\n            \n    # LABEL PLOT\n    plt.legend()\n    plt.title(f'Test User {u} - {len(clicks)} clicks, {len(carts)} carts, {len(orders)} orders:',size=18)\n    #plt.xlabel('Item category',size=16)\n    plt.ylabel('\\n\\nItem category',size=16)\n    plt.xticks([], [])\n    plt.yticks([], [])\n    plt.show()\n    \n    ############\n    ## PLOT HISTORY BY DAY AND HOUR\n    ############\n    \n    mn = test.ts.min()\n    dff['day'] = (dff.ts - mn) // (60*60*24)\n    dff['hour'] = ((dff.ts - mn) % (60*60*24)) // (60*60)\n    \n    plt.figure(figsize=(15,3))\n    xx = np.random.uniform(-0.2,0.2,len(dff))\n    yy = np.random.uniform(-0.5,0.5,len(dff))\n    plt.scatter(dff.day.values+xx, dff.hour.values+yy, s=25, color='orange')\n    cidx = dff.loc[dff['type']==1].index.values\n    oidx = dff.loc[dff['type']==2].index.values\n    plt.scatter(dff.day.values[cidx]+xx[cidx], dff.hour.values[cidx]+yy[cidx], s=50, color='green')\n    plt.scatter(dff.day.values[oidx]+xx[oidx], dff.hour.values[oidx]+yy[oidx], s=100, color='red')\n    old_xy = []\n    for i in range(len(dff)):\n        if 1: #i in pos:\n            x = dff.day.values[i]+xx[i]\n            y = dff.hour.values[i]+yy[i]\n            new_location = True\n            for j in old_xy:\n                if (np.abs(x-j[0])<0.5) & (np.abs(y-j[1])<4):\n                    new_location = False\n            if new_location:\n                plt.text(x, y, f'{i+1}', size=18)\n                old_xy.append( (x,y) )\n    plt.ylim((-1,25))\n    plt.xlim((-1,7))\n    plt.ylabel('Hour of Day',size=16)\n    plt.xlabel('Day of Month',size=16)\n    plt.yticks([0,4,8,12,16,20,24],['12am','4am','8am','noon','4pm','8pm','12am'])\n    plt.xticks([0,1,2,3,4,5,6],['Mon\\nAug 29th','Tue\\nAug 30rd','Wed\\nAug 31st',\n                            'Thr\\nSep 1st','Fri\\nSep 2nd','Sat\\nSep 3rd','Sun\\nSep 4th'])\n    c1 = mpatches.Patch(color='orange', label='Click')\n    c2 = mpatches.Patch(color='green', label='Cart')\n    c3 = mpatches.Patch(color='red', label='Order')\n    plt.legend(handles=[c1,c2,c3])\n    plt.show()\n    \n    print('\\n\\n\\n\\n\\n\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2022-12-11T04:24:08.002033Z","iopub.execute_input":"2022-12-11T04:24:08.002547Z","iopub.status.idle":"2022-12-11T04:24:55.206878Z","shell.execute_reply.started":"2022-12-11T04:24:08.002504Z","shell.execute_reply":"2022-12-11T04:24:55.205863Z"},"trusted":true},"execution_count":null,"outputs":[]}]}