{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Using GPU\n\nThis notebook is a GPU version of the [excellent matrix factoization notebook](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader) by @radek1 .  Most of the text is from his notebook except when I discuss GPU.  Main differences between Radek's notebook and this notebook are:\n- Radek used polar to handle dataframes, we will use cudf. \n- Radek used annoy for nearest neighbors, we will use cuml. \n- Radek trained his factpoization model on CPU, we will use GPU.\n\nLet's go!\n\n## Original introduction by Radek\n\nA co-visitation matrix is essentially an \"analog\" approximation to matrix factorization! I talk a bit more about this idea here: [💡 What is the co-visitation matrix, really?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358).\n\nBut matrix factorization has a lot of advantages as compared to co-visitation matrices. First of all, it can make better use of data -- it operates on the notion of similarity between categories. We can construct a more powerful representation if our model understands that aid `1` is similar to aid `142` as opposed to it treating each aid as an atomic entity (this is the jump from unigram/bigram/trigram models to word2vec in NLP).\n\nLet us thus train a matrix factorization model and replace the co-visitation matrices with it!\n\nNow, I don't expect that the first version of the model will be particularly well tuned. There has already been a lot of work put into co-visitation matrices and in the later versions we work off 3 different matrices, one for each category of actions! A similar progression can and will happen with matrix factorization 🙂 This notebook hopefully will enable us to jumpstart this type of exploration 🙂\n\nTo streamline the work, we will use data in `parquet` format. (Here is the notebook [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files) and here is [the most up-to-date version of the dataset](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint), no need for dealing with `jasonl` files and the associated mess any longer! Please upvote if you find this useful!)\n\nFor data processing we will use [polars](https://www.pola.rs/). `Polars` has a much smaller memory footprint than `pandas` and is quite fast. Plus it has really clean, intuitive API.\n\nLet's get to work! 🙂\n\n\n## Other resources you might find useful:\n\n* [💡 Training an XGBoost Ranker on the GPU with Merlin Models 🔥🔥🔥](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368848)\n* [How to train a Word2Vec model 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368384)\n* [💡 Can you beat static rules with a ranker model without additional features?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474)\n* [🐘 the elephant in the room -- high cardinality of targets and what to do about this](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722)\n* [📖 What are some good resources to learn about how gradient-boosted tree ranking models work?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366477)\n* [💡How to ensemble predictions -- a key component to every strong solution 🏅](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368747)\n* [from zero to 60 in 2 seconds or less 🏎️🚓🚓🚓](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367058)\n* [💡What is a good initial goal in the competition? How to improve beyond it? 📈](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368685)\n* [💡How to improve the results of your Approximate Nearest Neighbor search! (annoy)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368385)\n* [📅 Dataset for local validation created using organizer's repository (parquet files)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534)\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Data Preprocessing","metadata":{}},{"cell_type":"code","source":"import cudf\n\ntrain = cudf.read_parquet('../input/otto-train-and-test-data-for-local-validation/train.parquet')\ntest = cudf.read_parquet('../input/otto-train-and-test-data-for-local-validation/test.parquet')\n\n!pip install pickle5\nimport pickle5 as pickle\n\nwith open('../input/otto-train-and-test-data-for-local-validation/id2type.pkl', \"rb\") as fh:\n    id2type = pickle.load(fh)\nwith open('../input/otto-train-and-test-data-for-local-validation/type2id.pkl', \"rb\") as fh:\n    type2id = pickle.load(fh)","metadata":{"execution":{"iopub.status.busy":"2023-06-05T07:43:26.257273Z","iopub.execute_input":"2023-06-05T07:43:26.257662Z","iopub.status.idle":"2023-06-05T07:44:01.586272Z","shell.execute_reply.started":"2023-06-05T07:43:26.257579Z","shell.execute_reply":"2023-06-05T07:44:01.585018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test","metadata":{"execution":{"iopub.status.busy":"2023-06-05T07:44:01.588952Z","iopub.execute_input":"2023-06-05T07:44:01.589389Z","iopub.status.idle":"2023-06-05T07:44:01.679123Z","shell.execute_reply.started":"2023-06-05T07:44:01.589343Z","shell.execute_reply":"2023-06-05T07:44:01.678159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We need to create `aid-aid` pairs to train our matrix factorization model!\n\nLet's us grab the pairs both from the train and test set.","metadata":{}},{"cell_type":"code","source":"%%time\n\ntrain_pairs = cudf.concat([train, test])[['session', 'aid']]\ndel train, test\n\ntrain_pairs['aid_next'] = train_pairs.groupby('session').aid.shift(-1)\ntrain_pairs = train_pairs[['aid', 'aid_next']].dropna().reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:32:55.431561Z","iopub.execute_input":"2023-06-04T13:32:55.434039Z","iopub.status.idle":"2023-06-04T13:32:57.150297Z","shell.execute_reply.started":"2023-06-04T13:32:55.43399Z","shell.execute_reply":"2023-06-04T13:32:57.149207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The running time is 15x better than when using polar with CPU!","metadata":{}},{"cell_type":"code","source":"train_pairs.shape[0] / 1_000_000","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:32:57.152396Z","iopub.execute_input":"2023-06-04T13:32:57.153456Z","iopub.status.idle":"2023-06-04T13:32:57.163117Z","shell.execute_reply.started":"2023-06-04T13:32:57.153416Z","shell.execute_reply":"2023-06-04T13:32:57.161933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That is 209 million pairs created in 40 seconds without running out of RAM! 🙂 Not too bad","metadata":{}},{"cell_type":"code","source":"train_pairs.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-04T12:26:16.770908Z","iopub.execute_input":"2023-06-04T12:26:16.771212Z","iopub.status.idle":"2023-06-04T12:26:16.82475Z","shell.execute_reply.started":"2023-06-04T12:26:16.77118Z","shell.execute_reply":"2023-06-04T12:26:16.823792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see what is the cardinality of our aids -- we will need this to create the embedding layer.","metadata":{}},{"cell_type":"code","source":"cardinality_aids = max(train_pairs['aid'].max(), train_pairs['aid_next'].max())\ncardinality_aids","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:32:57.165993Z","iopub.execute_input":"2023-06-04T13:32:57.166425Z","iopub.status.idle":"2023-06-04T13:32:57.205842Z","shell.execute_reply.started":"2023-06-04T13:32:57.166389Z","shell.execute_reply":"2023-06-04T13:32:57.204428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will have up to `1855602` -- that is a lot! But our matrix factorization model will be able to handle this.\n","metadata":{}},{"cell_type":"markdown","source":"Oh dear, that took forever! Mind you, were are not doing anything here, apart from iterating over the dataset for a single epoch (and that is without validation!).\n\nThe reason this is taking so long is that indexing into the the arrays and collating results into batches is very computationally expensive.\n\nThere are ways to work around this but they require writing a lot of code (you could use the iterable-style dataset). And still our solution wouldn't be particularly well optimized.\n\nLet us do something else instead!\n\nWe will use a brand new [Merlin Dataloader](https://github.com/NVIDIA-Merlin/dataloader). It is a library that my team launched just a couple of days ago 🙂\n\nNow this library shines when you have a GPU, which is what you generally want when training DL models. But, alas, Kaggle gives you only 13 GB of RAM on a kernel with a GPU, and that wouldn't allow us to process our dataset!\n\nLet's see how far we can get with CPU only.","metadata":{}},{"cell_type":"code","source":"!pip install merlin-dataloader==0.0.2","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:32:57.208109Z","iopub.execute_input":"2023-06-04T13:32:57.208464Z","iopub.status.idle":"2023-06-04T13:34:36.798906Z","shell.execute_reply.started":"2023-06-04T13:32:57.208428Z","shell.execute_reply":"2023-06-04T13:34:36.797696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from merlin.loader.torch import Loader ","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:41:38.82199Z","iopub.execute_input":"2023-06-04T13:41:38.822375Z","iopub.status.idle":"2023-06-04T13:41:38.827825Z","shell.execute_reply.started":"2023-06-04T13:41:38.822337Z","shell.execute_reply":"2023-06-04T13:41:38.826623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can read data directly from the disk -- even better!\n\nLet's write our datasets to disk.","metadata":{}},{"cell_type":"code","source":"train_pairs[:-10_000_000].to_pandas().to_parquet('train_pairs.parquet')\ntrain_pairs[-10_000_000:].to_pandas().to_parquet('valid_pairs.parquet')","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:34:36.800778Z","iopub.execute_input":"2023-06-04T13:34:36.801521Z","iopub.status.idle":"2023-06-04T13:34:43.466616Z","shell.execute_reply.started":"2023-06-04T13:34:36.801474Z","shell.execute_reply":"2023-06-04T13:34:43.465339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Try to divide the train and validate set","metadata":{}},{"cell_type":"code","source":"#if DO_LOCAL_VALIDATION:\n #   seven_days = 7*24*60*60\n\n  #  train_cutoff = train.ts.max() - seven_days\n  # test = train[train.ts > train_cutoff]\n  # train = train[train.ts <= train_cutoff]\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from merlin.loader.torch import Loader \nfrom merlin.io import Dataset\n\ntrain_ds = Dataset('train_pairs.parquet')\ntrain_dl_merlin = Loader(train_ds, 65536, True)","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:34:43.468234Z","iopub.execute_input":"2023-06-04T13:34:43.468821Z","iopub.status.idle":"2023-06-04T13:34:48.326387Z","shell.execute_reply.started":"2023-06-04T13:34:43.468782Z","shell.execute_reply":"2023-06-04T13:34:48.324873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nfor batch in train_dl_merlin:\n    aid1, aid2 = batch[0], batch[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That is much better 🙂. Let's train our matrix factorization model!","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch import nn\n\nclass MatrixFactorization(nn.Module):\n    def __init__(self, n_aids, n_factors):\n        super().__init__()\n        self.aid_factors = nn.Embedding(n_aids, n_factors, sparse=True)\n        \n    def forward(self, aid1, aid2):\n        aid1 = self.aid_factors(aid1)\n        aid2 = self.aid_factors(aid2)\n        \n        return (aid1 * aid2).sum(dim=1)\n    \nclass AverageMeter(object):\n    \"\"\"Computes and stores the average and current value\"\"\"\n    def __init__(self, name, fmt=':f'):\n        self.name = name\n        self.fmt = fmt\n        self.reset()\n\n    def reset(self):\n        self.val = 0\n        self.avg = 0\n        self.sum = 0\n        self.count = 0\n\n    def update(self, val, n=1):\n        self.val = val\n        self.sum += val * n\n        self.count += n\n        self.avg = self.sum / self.count\n\n    def __str__(self):\n        fmtstr = '{name} {val' + self.fmt + '} ({avg' + self.fmt + '})'\n        return fmtstr.format(**self.__dict__)\n\nvalid_ds = Dataset('valid_pairs.parquet')\nvalid_dl_merlin = Loader(valid_ds, 65536, True)","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:34:48.343407Z","iopub.execute_input":"2023-06-04T13:34:48.344073Z","iopub.status.idle":"2023-06-04T13:34:48.609691Z","shell.execute_reply.started":"2023-06-04T13:34:48.344035Z","shell.execute_reply":"2023-06-04T13:34:48.608316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a= (aid1 * aid2).sum(dim=1)\na","metadata":{"execution":{"iopub.status.busy":"2023-06-04T12:37:14.03871Z","iopub.execute_input":"2023-06-04T12:37:14.039103Z","iopub.status.idle":"2023-06-04T12:37:14.047839Z","shell.execute_reply.started":"2023-06-04T12:37:14.039071Z","shell.execute_reply":"2023-06-04T12:37:14.046673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from torch.optim import SparseAdam\n\nnum_epochs=1\nlr=0.1\n\nmodel = MatrixFactorization(cardinality_aids+1, 32)\noptimizer = SparseAdam(model.parameters(), lr=lr)\ncriterion = nn.BCEWithLogitsLoss()","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:34:48.620698Z","iopub.execute_input":"2023-06-04T13:34:48.624143Z","iopub.status.idle":"2023-06-04T13:34:49.526185Z","shell.execute_reply.started":"2023-06-04T13:34:48.624086Z","shell.execute_reply":"2023-06-04T13:34:49.525169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"criterion: This variable defines the loss function used to compute the loss between the model's predictions and the target values. In this case, nn.BCEWithLogitsLoss() is used, which combines a sigmoid activation function with the binary cross-entropy loss. It is commonly used for binary classification problems.\n\n32 is a number of factors","metadata":{}},{"cell_type":"code","source":"%%time\nmodel.to('cuda')\nfor epoch in range(num_epochs):\n    for batch, _ in train_dl_merlin:\n        model.train()\n        losses = AverageMeter('Loss', ':.4e')\n            \n        aid1, aid2 = batch['aid'], batch['aid_next']\n        aid1 = aid1.to('cuda')\n        aid2 = aid2.to('cuda')\n        output_pos = model(aid1, aid2)\n        output_neg = model(aid1, aid2[torch.randperm(aid2.shape[0])])\n        \n        output = torch.cat([output_pos, output_neg])\n        targets = torch.cat([torch.ones_like(output_pos), torch.zeros_like(output_pos)])\n        loss = criterion(output, targets)\n        losses.update(loss.item())\n        \n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        \n    #evaluation\n    model.eval()\n    \n    with torch.no_grad():\n        accuracy = AverageMeter('accuracy')\n        for batch, _ in valid_dl_merlin:\n            aid1, aid2 = batch['aid'], batch['aid_next']\n            output_pos = model(aid1, aid2)\n            output_neg = model(aid1, aid2[torch.randperm(aid2.shape[0])])\n            accuracy_batch = torch.cat([output_pos.sigmoid() > 0.5, output_neg.sigmoid() < 0.5]).float().mean()\n            accuracy.update(accuracy_batch, aid1.shape[0])\n            \n    print(f'{epoch+1:02d}: * TrainLoss {losses.avg:.3f}  * Accuracy {accuracy.avg:.3f}')\n    \n    ","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:34:49.5276Z","iopub.execute_input":"2023-06-04T13:34:49.527976Z","iopub.status.idle":"2023-06-04T13:35:09.415296Z","shell.execute_reply.started":"2023-06-04T13:34:49.52794Z","shell.execute_reply":"2023-06-04T13:35:09.414193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"calculate loss\nloss = -[targets * log(sigmoid(predictions)) + (1 - targets) * log(1 - sigmoid(predictions))]\nPyTorch's nn.BCEWithLogitsLoss() combines the sigmoid activation and binary cross-entropy loss in a single operation for numerical stability. So, you don't need to apply the sigmoid function separately in this case.","metadata":{}},{"cell_type":"markdown","source":"output_neg\noutput_pos\nThe purpose of computing both positive and negative outputs is usually for training a recommendation or ranking model, where the goal is to maximize the similarity or relevance between positive pairs (aid1, aid2) and minimize the similarity or relevance between negative pairs (aid1, negative_aid2).\nevaluation: 0.5. Values above the threshold are considered positive predictions, while values below the threshold are considered negative predictions. The accuracy is calculated as the mean of the binary values.\nThe initialization of the weights is handled internally by the nn.Embedding module. By default, the weights are randomly initialized using a uniform distribution. During training, these initial weights are adjusted based on the computed gradients and the optimization algorithm used (in this case, SparseAdam). The purpose of training is to update the weights in a way that minimizes the defined loss function and improves the model's ability to make accurate predictions.","metadata":{}},{"cell_type":"markdown","source":"# Visualization the training process","metadata":{}},{"cell_type":"code","source":"aid1","metadata":{"execution":{"iopub.status.busy":"2023-06-04T12:36:23.80022Z","iopub.execute_input":"2023-06-04T12:36:23.80062Z","iopub.status.idle":"2023-06-04T12:36:23.809416Z","shell.execute_reply.started":"2023-06-04T12:36:23.800584Z","shell.execute_reply":"2023-06-04T12:36:23.808269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code ran about 60x faster than the cpu code form Radek's notebook. And we have not tuned the batch size! Using GPU larger batch size is possible which would reduce the running time further.","metadata":{}},{"cell_type":"markdown","source":"Let's grab the embeddings!","metadata":{}},{"cell_type":"code","source":"embeddings = model.aid_factors.weight.detach().cpu().numpy()","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:35:09.416803Z","iopub.execute_input":"2023-06-04T13:35:09.417881Z","iopub.status.idle":"2023-06-04T13:35:09.617049Z","shell.execute_reply.started":"2023-06-04T13:35:09.417839Z","shell.execute_reply":"2023-06-04T13:35:09.615816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And construct create the index for approximate nearest neighbor search. We will use cuml for that.","metadata":{}},{"cell_type":"code","source":"from cuml.neighbors import NearestNeighbors","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:35:09.618367Z","iopub.execute_input":"2023-06-04T13:35:09.618749Z","iopub.status.idle":"2023-06-04T13:35:10.649528Z","shell.execute_reply.started":"2023-06-04T13:35:09.618712Z","shell.execute_reply":"2023-06-04T13:35:10.648495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will compute 21 nearest neighbors as in Radek's notebook. ","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:40:01.459632Z","iopub.execute_input":"2022-12-08T09:40:01.460041Z","iopub.status.idle":"2022-12-08T09:40:01.467068Z","shell.execute_reply.started":"2022-12-08T09:40:01.460008Z","shell.execute_reply":"2022-12-08T09:40:01.465616Z"}}},{"cell_type":"code","source":"%%time\n\nknn = NearestNeighbors(n_neighbors=21, metric='euclidean')\nknn.fit(embeddings)","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:35:10.651664Z","iopub.execute_input":"2023-06-04T13:35:10.65204Z","iopub.status.idle":"2023-06-04T13:35:12.414025Z","shell.execute_reply.started":"2023-06-04T13:35:10.651997Z","shell.execute_reply":"2023-06-04T13:35:12.412996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom sklearn.manifold import TSNE\n\n# Perform dimensionality reduction using t-SNE\ntsne = TSNE(n_components=2)\nembeddings_2d = tsne.fit_transform(embeddings)\n\n# Plot the 2D embeddings\nplt.scatter(embeddings_2d[:, 0], embeddings_2d[:, 1])\nplt.xlabel('Dimension 1')\nplt.ylabel('Dimension 2')\nplt.title('Visualization of Embeddings')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:07:49.330527Z","iopub.execute_input":"2023-06-04T13:07:49.330898Z","iopub.status.idle":"2023-06-04T13:07:52.862968Z","shell.execute_reply.started":"2023-06-04T13:07:49.330866Z","shell.execute_reply":"2023-06-04T13:07:52.859869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now for any `aid`, we can find its nearest neighbor! cuml let you do this in parallel for all input at once.","metadata":{}},{"cell_type":"code","source":"%%time\n\n_, aid_nns = knn.kneighbors(embeddings)","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:35:12.415932Z","iopub.execute_input":"2023-06-04T13:35:12.416655Z","iopub.status.idle":"2023-06-04T13:36:26.955587Z","shell.execute_reply.started":"2023-06-04T13:35:12.416615Z","shell.execute_reply":"2023-06-04T13:36:26.954561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check we get 21 neighbors for each aid:","metadata":{}},{"cell_type":"code","source":"aid_nns.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:36:26.956937Z","iopub.execute_input":"2023-06-04T13:36:26.959784Z","iopub.status.idle":"2023-06-04T13:36:26.967684Z","shell.execute_reply.started":"2023-06-04T13:36:26.959745Z","shell.execute_reply":"2023-06-04T13:36:26.966618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aid_nns[0]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can get rid of the first neigbor directly.","metadata":{}},{"cell_type":"code","source":"aid_nns = aid_nns[:, 1:]","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:36:26.969307Z","iopub.execute_input":"2023-06-04T13:36:26.969685Z","iopub.status.idle":"2023-06-04T13:36:26.980344Z","shell.execute_reply.started":"2023-06-04T13:36:26.969649Z","shell.execute_reply":"2023-06-04T13:36:26.979294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create a submission! 🙂","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nfrom collections import defaultdict\n\n#sample_sub = pd.read_csv('../input/otto-recommender-system//sample_submission.csv')\ntest = cudf.read_parquet('../input/otto-train-and-test-data-for-local-validation/test.parquet')\n\nsession_types = ['clicks', 'carts', 'orders']\ngr = test.reset_index(drop=True).to_pandas().groupby('session')\ntest_session_AIDs = gr['aid'].apply(list)\ntest_session_types = gr['type'].apply(list)\n\nlabels = []\n\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\nfor AIDs, types in zip(test_session_AIDs, test_session_types):\n    if len(AIDs) >= 20:\n        # if we have enough aids (over equals 20) we don't need to look for candidates! we just use the old logic\n        weights=np.logspace(0.1,1,len(AIDs),base=2, endpoint=True)-1\n        aids_temp=defaultdict(lambda: 0)\n        for aid,w,t in zip(AIDs,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers[t]\n            \n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n        labels.append(sorted_aids[:20])\n    else:\n        # here we don't have 20 aids to output -- we will use approximate nearest neighbor search and our embeddings\n        # to generate candidates!\n        AIDs = list(dict.fromkeys(AIDs[::-1]))\n        \n        # let's grab the most recent aid\n        most_recent_aid = AIDs[0]\n        \n        # and look for some neighbors!\n        nns = list(aid_nns[most_recent_aid])\n                        \n        labels.append((AIDs+nns)[:20])","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:36:26.981856Z","iopub.execute_input":"2023-06-04T13:36:26.982218Z","iopub.status.idle":"2023-06-04T13:38:10.853157Z","shell.execute_reply.started":"2023-06-04T13:36:26.982181Z","shell.execute_reply":"2023-06-04T13:38:10.851849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AIDs = list(dict.fromkeys(AIDs[::-1]))\nprint(AIDs)","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:49:46.427977Z","iopub.execute_input":"2023-06-04T13:49:46.428346Z","iopub.status.idle":"2023-06-04T13:49:46.435229Z","shell.execute_reply.started":"2023-06-04T13:49:46.428315Z","shell.execute_reply":"2023-06-04T13:49:46.434057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AIDs[0]","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:50:15.693565Z","iopub.execute_input":"2023-06-04T13:50:15.694601Z","iopub.status.idle":"2023-06-04T13:50:15.702385Z","shell.execute_reply.started":"2023-06-04T13:50:15.694533Z","shell.execute_reply":"2023-06-04T13:50:15.701211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"raw","source":"assign weight with log\nIn the context of the code snippet you provided, these weights are multiplied with the type weight multipliers (type_weight_multipliers[t]) and used to assign higher weights to aids that appear earlier in the session. This reflects the assumption that earlier interactions or clicks are more significant or indicative of user preferences compared to later ones.","metadata":{}},{"cell_type":"markdown","source":"cuDF is a GPU-accelerated library for handling tabular data, similar to pandas but optimized for GPU computation.","metadata":{}},{"cell_type":"code","source":"test_session_AIDs","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:13:23.00313Z","iopub.execute_input":"2023-06-04T13:13:23.00353Z","iopub.status.idle":"2023-06-04T13:13:23.013705Z","shell.execute_reply.started":"2023-06-04T13:13:23.003497Z","shell.execute_reply":"2023-06-04T13:13:23.012706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_session_types","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:13:42.726471Z","iopub.execute_input":"2023-06-04T13:13:42.726847Z","iopub.status.idle":"2023-06-04T13:13:42.737419Z","shell.execute_reply.started":"2023-06-04T13:13:42.726814Z","shell.execute_reply":"2023-06-04T13:13:42.736437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now pull it all together and write to a file,","metadata":{}},{"cell_type":"code","source":"labels_as_strings = [' '.join([str(l) for l in lls]) for lls in labels]\n\npredictions = pd.DataFrame(data={'session_type': test_session_AIDs.index, 'labels': labels_as_strings})\n\nprediction_dfs = []\n\nfor st in session_types:\n    modified_predictions = predictions.copy()\n    modified_predictions.session_type = modified_predictions.session_type.astype('str') + f'_{st}'\n    prediction_dfs.append(modified_predictions)\n\nsubmission = pd.concat(prediction_dfs).reset_index(drop=True)\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:38:10.855276Z","iopub.execute_input":"2023-06-04T13:38:10.8558Z","iopub.status.idle":"2023-06-04T13:39:17.448427Z","shell.execute_reply.started":"2023-06-04T13:38:10.855747Z","shell.execute_reply":"2023-06-04T13:39:17.44735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:39:17.450149Z","iopub.execute_input":"2023-06-04T13:39:17.450576Z","iopub.status.idle":"2023-06-04T13:39:17.466396Z","shell.execute_reply.started":"2023-06-04T13:39:17.450511Z","shell.execute_reply":"2023-06-04T13:39:17.465376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Local Validation","metadata":{}},{"cell_type":"code","source":"test_labels = pd.read_parquet('../input/otto-train-and-test-data-for-local-validation/test_labels.parquet')\ntest_labels","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:29:26.211622Z","iopub.execute_input":"2023-06-04T13:29:26.211993Z","iopub.status.idle":"2023-06-04T13:29:27.790477Z","shell.execute_reply.started":"2023-06-04T13:29:26.211962Z","shell.execute_reply":"2023-06-04T13:29:27.789227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# COMPUTE METRIC\nscore = 0\nweights = {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}\nfor t in ['clicks','carts','orders']:\n    sub = submission.loc[submission.session_type.str.contains(t)].copy()\n    sub['session'] = sub.session_type.apply(lambda x: int(x.split('_')[0]))\n    sub.labels = sub.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n    test_labels = pd.read_parquet('../input/otto-train-and-test-data-for-local-validation/test_labels.parquet')\n    test_labels = test_labels.loc[test_labels['type']==t]\n    test_labels = test_labels.merge(sub, how='left', on=['session'])\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n    recall = test_labels['hits'].sum() / test_labels['gt_count'].sum()\n    score += weights[t]*recall\n    print(f'{t} recall =',recall)\n    \nprint('=============')\nprint('Overall Recall =',score)\nprint('=============')","metadata":{"execution":{"iopub.status.busy":"2023-06-04T13:39:17.467749Z","iopub.execute_input":"2023-06-04T13:39:17.468201Z","iopub.status.idle":"2023-06-04T13:41:38.82029Z","shell.execute_reply.started":"2023-06-04T13:39:17.468164Z","shell.execute_reply":"2023-06-04T13:41:38.819098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we are done!\n\n\n**If you like this notebook, please smash the upvote button! Thank you! 😊**\n\nThere are many ways in which this can be expanded:\n* we can train with larger batch sizes\n* we can train for longer\n* maybe we would get better results if we were to filter our train data by type?\n* should we train only on adjacent aids? maybe we should expand the neighborhood we train on\n\nWe can keep asking ourselves many questions like this 🙂 Now we have a framework to start answering them!\n\nThank you for reading! Happy Kaggling! 🙌","metadata":{"execution":{"iopub.status.busy":"2022-11-18T02:49:02.940358Z","iopub.execute_input":"2022-11-18T02:49:02.940858Z","iopub.status.idle":"2022-11-18T02:49:02.973867Z","shell.execute_reply.started":"2022-11-18T02:49:02.94076Z","shell.execute_reply":"2022-11-18T02:49:02.972223Z"}}},{"cell_type":"markdown","source":"# CPMP Addendum\n\nWe kept the logic used in Radeks notbook to create the submission. Better handcrafted ways to build submissions have been shared by Radek, @cdeotte, and others. Nothing prevents you from combining the GPU code from this notebook with these better submission logic.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}