{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Using GPU\n\nThis notebook is a GPU version of the [excellent matrix factoization notebook](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader) by @radek1 .  Most of the text is from his notebook except when I discuss GPU.  Main differences between Radek's notebook and this notebook are:\n- Radek used polar to handle dataframes, we will use cudf. \n- Radek used annoy for nearest neighbors, we will use cuml. \n- Radek trained his factpoization model on CPU, we will use GPU.\n\nLet's go!\n\n## Original introduction by Radek\n\nA co-visitation matrix is essentially an \"analog\" approximation to matrix factorization! I talk a bit more about this idea here: [💡 What is the co-visitation matrix, really?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358).\n\nBut matrix factorization has a lot of advantages as compared to co-visitation matrices. First of all, it can make better use of data -- it operates on the notion of similarity between categories. We can construct a more powerful representation if our model understands that aid `1` is similar to aid `142` as opposed to it treating each aid as an atomic entity (this is the jump from unigram/bigram/trigram models to word2vec in NLP).\n\nLet us thus train a matrix factorization model and replace the co-visitation matrices with it!\n\nNow, I don't expect that the first version of the model will be particularly well tuned. There has already been a lot of work put into co-visitation matrices and in the later versions we work off 3 different matrices, one for each category of actions! A similar progression can and will happen with matrix factorization 🙂 This notebook hopefully will enable us to jumpstart this type of exploration 🙂\n\nTo streamline the work, we will use data in `parquet` format. (Here is the notebook [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files) and here is [the most up-to-date version of the dataset](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint), no need for dealing with `jasonl` files and the associated mess any longer! Please upvote if you find this useful!)\n\nFor data processing we will use [polars](https://www.pola.rs/). `Polars` has a much smaller memory footprint than `pandas` and is quite fast. Plus it has really clean, intuitive API.\n\nLet's get to work! 🙂\n\n\n## Other resources you might find useful:\n\n* [💡 Training an XGBoost Ranker on the GPU with Merlin Models 🔥🔥🔥](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368848)\n* [How to train a Word2Vec model 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368384)\n* [💡 Can you beat static rules with a ranker model without additional features?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366474)\n* [🐘 the elephant in the room -- high cardinality of targets and what to do about this](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722)\n* [📖 What are some good resources to learn about how gradient-boosted tree ranking models work?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366477)\n* [💡How to ensemble predictions -- a key component to every strong solution 🏅](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368747)\n* [from zero to 60 in 2 seconds or less 🏎️🚓🚓🚓](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367058)\n* [💡What is a good initial goal in the competition? How to improve beyond it? 📈](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368685)\n* [💡How to improve the results of your Approximate Nearest Neighbor search! (annoy)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368385)\n* [📅 Dataset for local validation created using organizer's repository (parquet files)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534)\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Data Preprocessing","metadata":{}},{"cell_type":"code","source":"import cudf\n\ntrain = cudf.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\ntest = cudf.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:19:23.231262Z","iopub.execute_input":"2022-12-08T09:19:23.231728Z","iopub.status.idle":"2022-12-08T09:19:48.359168Z","shell.execute_reply.started":"2022-12-08T09:19:23.231621Z","shell.execute_reply":"2022-12-08T09:19:48.358053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We need to create `aid-aid` pairs to train our matrix factorization model!\n\nLet's us grab the pairs both from the train and test set.","metadata":{}},{"cell_type":"code","source":"%%time\n\ntrain_pairs = cudf.concat([train, test])[['session', 'aid']]\ndel train, test\n\ntrain_pairs['aid_next'] = train_pairs.groupby('session').aid.shift(-1)\ntrain_pairs = train_pairs[['aid', 'aid_next']].dropna().reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:19:48.361391Z","iopub.execute_input":"2022-12-08T09:19:48.362002Z","iopub.status.idle":"2022-12-08T09:19:50.338699Z","shell.execute_reply.started":"2022-12-08T09:19:48.361961Z","shell.execute_reply":"2022-12-08T09:19:50.337355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The running time is 15x better than when using polar with CPU!","metadata":{}},{"cell_type":"code","source":"train_pairs.shape[0] / 1_000_000","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:19:50.343253Z","iopub.execute_input":"2022-12-08T09:19:50.346035Z","iopub.status.idle":"2022-12-08T09:19:50.362772Z","shell.execute_reply.started":"2022-12-08T09:19:50.345988Z","shell.execute_reply":"2022-12-08T09:19:50.359408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That is 209 million pairs created in 40 seconds without running out of RAM! 🙂 Not too bad","metadata":{}},{"cell_type":"code","source":"train_pairs.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:19:50.367962Z","iopub.execute_input":"2022-12-08T09:19:50.369149Z","iopub.status.idle":"2022-12-08T09:19:50.427827Z","shell.execute_reply.started":"2022-12-08T09:19:50.369111Z","shell.execute_reply":"2022-12-08T09:19:50.426909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see what is the cardinality of our aids -- we will need this to create the embedding layer.","metadata":{}},{"cell_type":"code","source":"cardinality_aids = max(train_pairs['aid'].max(), train_pairs['aid_next'].max())\ncardinality_aids","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:19:50.430741Z","iopub.execute_input":"2022-12-08T09:19:50.431114Z","iopub.status.idle":"2022-12-08T09:19:50.466290Z","shell.execute_reply.started":"2022-12-08T09:19:50.431072Z","shell.execute_reply":"2022-12-08T09:19:50.465207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will have up to `1855602` -- that is a lot! But our matrix factorization model will be able to handle this.\n","metadata":{}},{"cell_type":"markdown","source":"Oh dear, that took forever! Mind you, were are not doing anything here, apart from iterating over the dataset for a single epoch (and that is without validation!).\n\nThe reason this is taking so long is that indexing into the the arrays and collating results into batches is very computationally expensive.\n\nThere are ways to work around this but they require writing a lot of code (you could use the iterable-style dataset). And still our solution wouldn't be particularly well optimized.\n\nLet us do something else instead!\n\nWe will use a brand new [Merlin Dataloader](https://github.com/NVIDIA-Merlin/dataloader). It is a library that my team launched just a couple of days ago 🙂\n\nNow this library shines when you have a GPU, which is what you generally want when training DL models. But, alas, Kaggle gives you only 13 GB of RAM on a kernel with a GPU, and that wouldn't allow us to process our dataset!\n\nLet's see how far we can get with CPU only.","metadata":{}},{"cell_type":"code","source":"!pip install merlin-dataloader==0.0.2","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:19:50.467696Z","iopub.execute_input":"2022-12-08T09:19:50.468167Z","iopub.status.idle":"2022-12-08T09:21:20.025182Z","shell.execute_reply.started":"2022-12-08T09:19:50.468130Z","shell.execute_reply":"2022-12-08T09:21:20.023965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from merlin.loader.torch import Loader ","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:21:20.028721Z","iopub.execute_input":"2022-12-08T09:21:20.029195Z","iopub.status.idle":"2022-12-08T09:21:22.137243Z","shell.execute_reply.started":"2022-12-08T09:21:20.029161Z","shell.execute_reply":"2022-12-08T09:21:22.136155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can read data directly from the disk -- even better!\n\nLet's write our datasets to disk.","metadata":{}},{"cell_type":"code","source":"train_pairs[:-10_000_000].to_pandas().to_parquet('train_pairs.parquet')\ntrain_pairs[-10_000_000:].to_pandas().to_parquet('valid_pairs.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:21:22.138802Z","iopub.execute_input":"2022-12-08T09:21:22.139175Z","iopub.status.idle":"2022-12-08T09:21:34.226301Z","shell.execute_reply.started":"2022-12-08T09:21:22.139138Z","shell.execute_reply":"2022-12-08T09:21:34.225212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from merlin.loader.torch import Loader \nfrom merlin.io import Dataset\n\ntrain_ds = Dataset('train_pairs.parquet')\ntrain_dl_merlin = Loader(train_ds, 65536, True)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:21:34.227763Z","iopub.execute_input":"2022-12-08T09:21:34.228166Z","iopub.status.idle":"2022-12-08T09:21:35.098436Z","shell.execute_reply.started":"2022-12-08T09:21:34.228127Z","shell.execute_reply":"2022-12-08T09:21:35.097358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nfor batch in train_dl_merlin:\n    aid1, aid2 = batch[0], batch[1]","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:21:35.099809Z","iopub.execute_input":"2022-12-08T09:21:35.100234Z","iopub.status.idle":"2022-12-08T09:21:38.341848Z","shell.execute_reply.started":"2022-12-08T09:21:35.100193Z","shell.execute_reply":"2022-12-08T09:21:38.340816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That is much better 🙂. Let's train our matrix factorization model!","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch import nn\n\nclass MatrixFactorization(nn.Module):\n    def __init__(self, n_aids, n_factors):\n        super().__init__()\n        self.aid_factors = nn.Embedding(n_aids, n_factors, sparse=True)\n        \n    def forward(self, aid1, aid2):\n        aid1 = self.aid_factors(aid1)\n        aid2 = self.aid_factors(aid2)\n        \n        return (aid1 * aid2).sum(dim=1)\n    \nclass AverageMeter(object):\n    \"\"\"Computes and stores the average and current value\"\"\"\n    def __init__(self, name, fmt=':f'):\n        self.name = name\n        self.fmt = fmt\n        self.reset()\n\n    def reset(self):\n        self.val = 0\n        self.avg = 0\n        self.sum = 0\n        self.count = 0\n\n    def update(self, val, n=1):\n        self.val = val\n        self.sum += val * n\n        self.count += n\n        self.avg = self.sum / self.count\n\n    def __str__(self):\n        fmtstr = '{name} {val' + self.fmt + '} ({avg' + self.fmt + '})'\n        return fmtstr.format(**self.__dict__)\n\nvalid_ds = Dataset('valid_pairs.parquet')\nvalid_dl_merlin = Loader(valid_ds, 65536, True)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:23:51.967216Z","iopub.execute_input":"2022-12-08T09:23:51.967661Z","iopub.status.idle":"2022-12-08T09:23:52.129883Z","shell.execute_reply.started":"2022-12-08T09:23:51.967624Z","shell.execute_reply":"2022-12-08T09:23:52.128862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from torch.optim import SparseAdam\n\nnum_epochs=1\nlr=0.1\n\nmodel = MatrixFactorization(cardinality_aids+1, 32)\noptimizer = SparseAdam(model.parameters(), lr=lr)\ncriterion = nn.BCEWithLogitsLoss()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:25:02.536730Z","iopub.execute_input":"2022-12-08T09:25:02.537511Z","iopub.status.idle":"2022-12-08T09:25:03.019582Z","shell.execute_reply.started":"2022-12-08T09:25:02.537469Z","shell.execute_reply":"2022-12-08T09:25:03.018534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel.to('cuda')\nfor epoch in range(num_epochs):\n    for batch, _ in train_dl_merlin:\n        model.train()\n        losses = AverageMeter('Loss', ':.4e')\n            \n        aid1, aid2 = batch['aid'], batch['aid_next']\n        aid1 = aid1.to('cuda')\n        aid2 = aid2.to('cuda')\n        output_pos = model(aid1, aid2)\n        output_neg = model(aid1, aid2[torch.randperm(aid2.shape[0])])\n        \n        output = torch.cat([output_pos, output_neg])\n        targets = torch.cat([torch.ones_like(output_pos), torch.zeros_like(output_pos)])\n        loss = criterion(output, targets)\n        losses.update(loss.item())\n        \n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        \n    model.eval()\n    \n    with torch.no_grad():\n        accuracy = AverageMeter('accuracy')\n        for batch, _ in valid_dl_merlin:\n            aid1, aid2 = batch['aid'], batch['aid_next']\n            output_pos = model(aid1, aid2)\n            output_neg = model(aid1, aid2[torch.randperm(aid2.shape[0])])\n            accuracy_batch = torch.cat([output_pos.sigmoid() > 0.5, output_neg.sigmoid() < 0.5]).float().mean()\n            accuracy.update(accuracy_batch, aid1.shape[0])\n            \n    print(f'{epoch+1:02d}: * TrainLoss {losses.avg:.3f}  * Accuracy {accuracy.avg:.3f}')","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:25:05.687807Z","iopub.execute_input":"2022-12-08T09:25:05.688207Z","iopub.status.idle":"2022-12-08T09:25:30.337578Z","shell.execute_reply.started":"2022-12-08T09:25:05.688176Z","shell.execute_reply":"2022-12-08T09:25:30.336476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code ran about 60x faster than the cpu code form Radek's notebook. And we have not tuned the batch size! Using GPU larger batch size is possible which would reduce the running time further.","metadata":{}},{"cell_type":"markdown","source":"Let's grab the embeddings!","metadata":{}},{"cell_type":"code","source":"embeddings = model.aid_factors.weight.detach().cpu().numpy()","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:27:06.091705Z","iopub.execute_input":"2022-12-08T09:27:06.092101Z","iopub.status.idle":"2022-12-08T09:27:06.285784Z","shell.execute_reply.started":"2022-12-08T09:27:06.092067Z","shell.execute_reply":"2022-12-08T09:27:06.284836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And construct create the index for approximate nearest neighbor search. We will use cuml for that.","metadata":{}},{"cell_type":"code","source":"from cuml.neighbors import NearestNeighbors","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:28:57.120190Z","iopub.execute_input":"2022-12-08T09:28:57.121322Z","iopub.status.idle":"2022-12-08T09:28:57.511365Z","shell.execute_reply.started":"2022-12-08T09:28:57.121281Z","shell.execute_reply":"2022-12-08T09:28:57.510356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will compute 21 nearest neighbors as in Radek's notebook. ","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:40:01.459632Z","iopub.execute_input":"2022-12-08T09:40:01.460041Z","iopub.status.idle":"2022-12-08T09:40:01.467068Z","shell.execute_reply.started":"2022-12-08T09:40:01.460008Z","shell.execute_reply":"2022-12-08T09:40:01.465616Z"}}},{"cell_type":"code","source":"%%time\n\nknn = NearestNeighbors(n_neighbors=21, metric='euclidean')\nknn.fit(embeddings)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:42:55.307730Z","iopub.execute_input":"2022-12-08T09:42:55.308583Z","iopub.status.idle":"2022-12-08T09:42:55.367333Z","shell.execute_reply.started":"2022-12-08T09:42:55.308542Z","shell.execute_reply":"2022-12-08T09:42:55.366297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now for any `aid`, we can find its nearest neighbor! cuml let you do this in parallel for all input at once.","metadata":{}},{"cell_type":"code","source":"%%time\n\n_, aid_nns = knn.kneighbors(embeddings)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:42:56.949102Z","iopub.execute_input":"2022-12-08T09:42:56.950161Z","iopub.status.idle":"2022-12-08T09:44:09.780045Z","shell.execute_reply.started":"2022-12-08T09:42:56.950121Z","shell.execute_reply":"2022-12-08T09:44:09.778951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check we get 21 neighbors for each aid:","metadata":{}},{"cell_type":"code","source":"aid_nns.shape","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:44:09.783431Z","iopub.execute_input":"2022-12-08T09:44:09.783728Z","iopub.status.idle":"2022-12-08T09:44:09.794025Z","shell.execute_reply.started":"2022-12-08T09:44:09.783693Z","shell.execute_reply":"2022-12-08T09:44:09.793033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can get rid of the first neigbor directly.","metadata":{}},{"cell_type":"code","source":"aid_nns = aid_nns[:, 1:]","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:44:09.804175Z","iopub.execute_input":"2022-12-08T09:44:09.804614Z","iopub.status.idle":"2022-12-08T09:44:09.809459Z","shell.execute_reply.started":"2022-12-08T09:44:09.804577Z","shell.execute_reply":"2022-12-08T09:44:09.808265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create a submission! 🙂","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nfrom collections import defaultdict\n\nsample_sub = pd.read_csv('../input/otto-recommender-system//sample_submission.csv')\ntest = cudf.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')\n\nsession_types = ['clicks', 'carts', 'orders']\ngr = test.reset_index(drop=True).to_pandas().groupby('session')\ntest_session_AIDs = gr['aid'].apply(list)\ntest_session_types = gr['type'].apply(list)\n\nlabels = []\n\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\nfor AIDs, types in zip(test_session_AIDs, test_session_types):\n    if len(AIDs) >= 20:\n        # if we have enough aids (over equals 20) we don't need to look for candidates! we just use the old logic\n        weights=np.logspace(0.1,1,len(AIDs),base=2, endpoint=True)-1\n        aids_temp=defaultdict(lambda: 0)\n        for aid,w,t in zip(AIDs,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers[t]\n            \n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n        labels.append(sorted_aids[:20])\n    else:\n        # here we don't have 20 aids to output -- we will use approximate nearest neighbor search and our embeddings\n        # to generate candidates!\n        AIDs = list(dict.fromkeys(AIDs[::-1]))\n        \n        # let's grab the most recent aid\n        most_recent_aid = AIDs[0]\n        \n        # and look for some neighbors!\n        nns = list(aid_nns[most_recent_aid])\n                        \n        labels.append((AIDs+nns)[:20])","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:44:18.451861Z","iopub.execute_input":"2022-12-08T09:44:18.453090Z","iopub.status.idle":"2022-12-08T09:45:32.465032Z","shell.execute_reply.started":"2022-12-08T09:44:18.453041Z","shell.execute_reply":"2022-12-08T09:45:32.463877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now pull it all together and write to a file,","metadata":{}},{"cell_type":"code","source":"labels_as_strings = [' '.join([str(l) for l in lls]) for lls in labels]\n\npredictions = pd.DataFrame(data={'session_type': test_session_AIDs.index, 'labels': labels_as_strings})\n\nprediction_dfs = []\n\nfor st in session_types:\n    modified_predictions = predictions.copy()\n    modified_predictions.session_type = modified_predictions.session_type.astype('str') + f'_{st}'\n    prediction_dfs.append(modified_predictions)\n\nsubmission = pd.concat(prediction_dfs).reset_index(drop=True)\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-12-08T09:45:32.467441Z","iopub.execute_input":"2022-12-08T09:45:32.468379Z","iopub.status.idle":"2022-12-08T09:46:08.946202Z","shell.execute_reply.started":"2022-12-08T09:45:32.468324Z","shell.execute_reply":"2022-12-08T09:46:08.945070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we are done!\n\n\n**If you like this notebook, please smash the upvote button! Thank you! 😊**\n\nThere are many ways in which this can be expanded:\n* we can train with larger batch sizes\n* we can train for longer\n* maybe we would get better results if we were to filter our train data by type?\n* should we train only on adjacent aids? maybe we should expand the neighborhood we train on\n\nWe can keep asking ourselves many questions like this 🙂 Now we have a framework to start answering them!\n\nThank you for reading! Happy Kaggling! 🙌","metadata":{"execution":{"iopub.status.busy":"2022-11-18T02:49:02.940358Z","iopub.execute_input":"2022-11-18T02:49:02.940858Z","iopub.status.idle":"2022-11-18T02:49:02.973867Z","shell.execute_reply.started":"2022-11-18T02:49:02.94076Z","shell.execute_reply":"2022-11-18T02:49:02.972223Z"}}},{"cell_type":"markdown","source":"# CPMP Addendum\n\nWe kept the logic used in Radeks notbook to create the submission. Better handcrafted ways to build submissions have been shared by Radek, @cdeotte, and others. Nothing prevents you from combining the GPU code from this notebook with these better submission logic.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}