{
  "id": 365369,
  "title": "30x Faster Co-Visitation Matrices using RAPIDS cuDF!",
  "url": "/competitions/otto-recommender-system/discussion/365369",
  "author_name": "Chris Deotte",
  "post_date": "2022-11-11T00:47:57.854000",
  "votes": 88,
  "comment_count": 32,
  "views": 0,
  "content": "<p>I published a notebook <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\" target=\"_blank\">here</a> that uses RAPIDS cuDF to compute 30x faster co-visitation matrices. Using co-visitation matrices helps provide our models with \"candidates\". Then our models (and/or human logic) and \"rerank\" these \"candidates\" and select our final 20 predictions for submission CSV. </p>\n<p><strong>UPDATE</strong> 🔥 latest version computes co-visitation matrices in <strong>3min</strong> each 🔥</p>\n<p>For more information about \"candidate rerank\" models, see Ravi's discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 2025071,
      "postDate": "2022-11-11T00:47:57.853Z",
      "content": "<p>I published a notebook <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\" target=\"_blank\">here</a> that uses RAPIDS cuDF to compute 30x faster co-visitation matrices. Using co-visitation matrices helps provide our models with \"candidates\". Then our models (and/or human logic) and \"rerank\" these \"candidates\" and select our final 20 predictions for submission CSV. </p>\n<p><strong>UPDATE</strong> 🔥 latest version computes co-visitation matrices in <strong>3min</strong> each 🔥</p>\n<p>For more information about \"candidate rerank\" models, see Ravi's discussion <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I published a notebook [here][1] that uses RAPIDS cuDF to compute 30x faster co-visitation matrices. Using co-visitation matrices helps provide our models with \"candidates\". Then our models (and/or human logic) and \"rerank\" these \"candidates\" and select our final 20 predictions for submission CSV. \n\n**UPDATE** 🔥 latest version computes co-visitation matrices in **3min** each 🔥\n\nFor more information about \"candidate rerank\" models, see Ravi's discussion [here][2]\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\n[2]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721",
      "votes": 88
    },
    {
      "id": 2028397,
      "postDate": "2022-11-13T21:32:00.943Z",
      "content": "<p>UPDATE: I updated my code to create co-visitation matrices 1.5x faster. It now creates co-visitation matrices in Kaggle notebooks at 6 minutes! (And offline using 1xV100 in 1.5 minutes!)</p>",
      "rawMarkdown": "UPDATE: I updated my code to create co-visitation matrices 1.5x faster. It now creates co-visitation matrices in Kaggle notebooks at 6 minutes! (And offline using 1xV100 in 1.5 minutes!)",
      "votes": 7,
      "replies": [
        {
          "id": 2032873,
          "postDate": "2022-11-16T21:01:18.247Z",
          "content": "<p>UPDATE: 🔥 I updated my code again to create co-visitation matrices 2x faster. It now creates co-visitation matrices in Kaggle notebooks at <strong>3 minutes</strong>! (This is <strong>3x</strong> faster than version 1 of the notebook. And <strong>30x</strong> faster than public notebooks using Pandas CPU) 🔥</p>",
          "rawMarkdown": "UPDATE: 🔥 I updated my code again to create co-visitation matrices 2x faster. It now creates co-visitation matrices in Kaggle notebooks at **3 minutes**! (This is **3x** faster than version 1 of the notebook. And **30x** faster than public notebooks using Pandas CPU) 🔥",
          "votes": 4
        }
      ]
    },
    {
      "id": 2035237,
      "postDate": "2022-11-18T17:59:08.493Z",
      "content": "<p>Hi Chris,</p>\n<p>Thanks for showing us this wonderful solution 👍. </p>\n<p>I find it hard to improve on this simple heuristic approach combining information on IN_SESSION flag, TOP most popular items for given item in the session and TOP most popular items overall.</p>\n<p>I was wondering, how did you come up with this heuristic? Is it something that you tried out before and it worked, or you tried out bunch of different techniques and this one worked the best or simply you had a lucky/intuitive guess?</p>\n<p>Best </p>\n<p>Andrej</p>",
      "rawMarkdown": "Hi Chris,\n\nThanks for showing us this wonderful solution 👍. \n\nI find it hard to improve on this simple heuristic approach combining information on IN_SESSION flag, TOP most popular items for given item in the session and TOP most popular items overall.\n\nI was wondering, how did you come up with this heuristic? Is it something that you tried out before and it worked, or you tried out bunch of different techniques and this one worked the best or simply you had a lucky/intuitive guess?\n\nBest \n\nAndrej",
      "votes": 1,
      "replies": [
        {
          "id": 2036565,
          "postDate": "2022-11-19T22:22:38.557Z",
          "content": "<p>There are two ways to improve CV and LB. Find more <strong>candidates</strong> and find better <strong>rerank</strong> logic. I discover both by using intuition and thinking about how users behave. Then i evaluate my ideas by observing if local validation score increases or not.</p>\n<p>The easiest way to improve my public notebook is to create more co-visitation matrices. It is possible to improve my notebook by changing the rerank logic, but it will be easier to just make more co-visitation matrices and add their weight in.</p>\n<p>I suggest that you focus on the case of <code>suggest_buys</code> for <code>len(unique_aids)&lt;20</code>. This has the largest effect on CV score and LB score. Change the code to allow for weights. The current code is</p>\n<pre><code># USE \"CART ORDER\" CO-VISITATION MATRIX\naids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n# USE \"BUY2BUY\" CO-VISITATION MATRIX\naids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n# RERANK CANDIDATES\ntop_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n</code></pre>\n<p>Change this to</p>\n<pre><code>from collections import Counter\naids_temp = Counter() \n# USE \"CART ORDER\" CO-VISITATION MATRIX\naids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\nfor aid in aids2: aids_temp[aid] += 1\n# USE \"BUY2BUY\" CO-VISITATION MATRIX\naids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\nfor aid in aids3: aids_temp[aid] += 1\n# RERANK CANDIDATES\ntop_aids2 = [k for k,v in aids_temp.most_common(20) if k not in unique_aids]\n</code></pre>\n<p>Now we can create new co-visitation matrices and include them with</p>\n<pre><code>aids4 = list(itertools.chain(*[NEW_MATRIX[aid] for aid in SUBSET_OF_USER_HISTORY_ITEMS if aid in NEW_MATRIX]))\nfor aid in aids4: aids_temp[aid] += 1\n</code></pre>\n<p>You need to be creative when making co-visitation matrices. The matrices makes pairs of items. If a <strong>certain type of user</strong> interacts in a <strong>certain type of way with item</strong>. Then we predict they will click and order a certain item in the future. Then when you include this new co-visitation matrix in your inference, we only apply it to the same <strong>certain type of user</strong> when they interact with an item in a <strong>certain type of way</strong>.</p>\n<p>Also note that once you change the code to include <code>for aid in aids2: aids_temp[aid] += 1</code>, you can now assign different situations different amount of weights during inference. This is how we change the <code>rerank</code> logic. </p>\n<pre><code>weights = LOGIC_TO_MAKES_WEIGHTS\nfor aid,w in zip(aids2,weights):\n    aids_temp[aid] += w\n</code></pre>",
          "rawMarkdown": "There are two ways to improve CV and LB. Find more **candidates** and find better **rerank** logic. I discover both by using intuition and thinking about how users behave. Then i evaluate my ideas by observing if local validation score increases or not.\n\nThe easiest way to improve my public notebook is to create more co-visitation matrices. It is possible to improve my notebook by changing the rerank logic, but it will be easier to just make more co-visitation matrices and add their weight in.\n\nI suggest that you focus on the case of `suggest_buys` for `len(unique_aids)<20`. This has the largest effect on CV score and LB score. Change the code to allow for weights. The current code is\n\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n\nChange this to\n\n    from collections import Counter\n    aids_temp = Counter() \n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    for aid in aids2: aids_temp[aid] += 1\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    for aid in aids3: aids_temp[aid] += 1\n    # RERANK CANDIDATES\n    top_aids2 = [k for k,v in aids_temp.most_common(20) if k not in unique_aids]\n\nNow we can create new co-visitation matrices and include them with\n\n    aids4 = list(itertools.chain(*[NEW_MATRIX[aid] for aid in SUBSET_OF_USER_HISTORY_ITEMS if aid in NEW_MATRIX]))\n    for aid in aids4: aids_temp[aid] += 1\n\nYou need to be creative when making co-visitation matrices. The matrices makes pairs of items. If a **certain type of user** interacts in a **certain type of way with item**. Then we predict they will click and order a certain item in the future. Then when you include this new co-visitation matrix in your inference, we only apply it to the same **certain type of user** when they interact with an item in a **certain type of way**.\n\nAlso note that once you change the code to include `for aid in aids2: aids_temp[aid] += 1`, you can now assign different situations different amount of weights during inference. This is how we change the `rerank` logic. \n\n    weights = LOGIC_TO_MAKES_WEIGHTS\n    for aid,w in zip(aids2,weights):\n        aids_temp[aid] += w",
          "votes": 31
        }
      ]
    },
    {
      "id": 2032899,
      "postDate": "2022-11-16T21:21:55.620Z",
      "content": "<p>this is so valuable and lets us experiment with different weighting approaches in an acceptable time! tnx chris!</p>",
      "rawMarkdown": "this is so valuable and lets us experiment with different weighting approaches in an acceptable time! tnx chris!",
      "votes": 1
    },
    {
      "id": 2027772,
      "postDate": "2022-11-13T07:20:54.380Z",
      "content": "<p>Very awesome <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, not sure how to implement but curious if the session was treated as a corpus where each paring is a word and the cart and order events are treated as words in the sequence then a generator can just simulate the next 20 words which can include a cart or order event in the sequence, guessing they are not too accurate yet as they are randomized for creativity but that can also add some flavor for the buyers although in this simulation we also have to predict the behavior of current suggestion models. Back to your model, great work!</p>",
      "rawMarkdown": "Very awesome @cdeotte, not sure how to implement but curious if the session was treated as a corpus where each paring is a word and the cart and order events are treated as words in the sequence then a generator can just simulate the next 20 words which can include a cart or order event in the sequence, guessing they are not too accurate yet as they are randomized for creativity but that can also add some flavor for the buyers although in this simulation we also have to predict the behavior of current suggestion models. Back to your model, great work!",
      "votes": 1
    },
    {
      "id": 2026387,
      "postDate": "2022-11-12T03:02:50.447Z",
      "content": "<p>Great kernel, as always <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>",
      "rawMarkdown": "Great kernel, as always @cdeotte.\n\n",
      "votes": 1
    },
    {
      "id": 2030572,
      "postDate": "2022-11-15T14:23:26.880Z",
      "content": "<p>How about the modin library it is bit faster in some areas</p>",
      "rawMarkdown": "How about the modin library it is bit faster in some areas",
      "votes": 2,
      "replies": [
        {
          "id": 2030951,
          "postDate": "2022-11-15T18:29:55.410Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/satyaprakashshukl\" target=\"_blank\">@satyaprakashshukl</a> for the suggestion. I was not aware of Modin library. I'll check it out. Another option to accelerate Pandas (groupby functions) on CPU is <code>pandarallel</code> <a href=\"https://pypi.org/project/pandarallel/\" target=\"_blank\">here</a>. I tried out <code>pandarallel</code> on my public notebook inference and it boosted inference offline Kaggle from 30 minutes to 5 minutes using 20 CPU! (discussion <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2026011\" target=\"_blank\">here</a>)</p>\n<p>Eventually i will rewrite my inference using RAPIDS cuDF and expect even more speed boost than pandarallel.</p>",
          "rawMarkdown": "Thanks @satyaprakashshukl for the suggestion. I was not aware of Modin library. I'll check it out. Another option to accelerate Pandas (groupby functions) on CPU is `pandarallel` [here][1]. I tried out `pandarallel` on my public notebook inference and it boosted inference offline Kaggle from 30 minutes to 5 minutes using 20 CPU! (discussion [here][2])\n\nEventually i will rewrite my inference using RAPIDS cuDF and expect even more speed boost than pandarallel.\n\n[1]: https://pypi.org/project/pandarallel/\n[2]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2026011",
          "votes": 5
        },
        {
          "id": 2031356,
          "postDate": "2022-11-16T03:05:34.503Z",
          "content": "<p>Thank you chris,i will explore too  these option in my kernel  🙏🏽🙂</p>",
          "rawMarkdown": "Thank you chris,i will explore too  these option in my kernel  🙏🏽🙂"
        },
        {
          "id": 2032901,
          "postDate": "2022-11-16T21:27:09.273Z",
          "content": "<p>that would be great chris! than nothing is stopping us from trying even more experiments!</p>",
          "rawMarkdown": "that would be great chris! than nothing is stopping us from trying even more experiments!"
        }
      ]
    },
    {
      "id": 2025082,
      "postDate": "2022-11-11T01:00:44.913Z",
      "content": "<p>Those speed gains are insane! 🙂 Didn't assume this workload could be transformed to run fast on the GPU… and yet here we are! 😄 </p>\n<p>Have to look more closely at how you are creating the co-visitation matrix dicts. That is some real magic happening there and something that has been an extremely slow part of the process thus far (when we went <code>matrix[aid_x][aid_y] + w</code>)!</p>",
      "rawMarkdown": "Those speed gains are insane! 🙂 Didn't assume this workload could be transformed to run fast on the GPU... and yet here we are! 😄 \n\nHave to look more closely at how you are creating the co-visitation matrix dicts. That is some real magic happening there and something that has been an extremely slow part of the process thus far (when we went `matrix[aid_x][aid_y] + w`)!",
      "votes": 2,
      "replies": [
        {
          "id": 2025091,
          "postDate": "2022-11-11T01:05:38.060Z",
          "content": "<p>Thanks Radek. To run many experiments, we need to increase the speed of our entire pipeline. Using 1xV100 32GB, we can compute co-visitation matrices in 2 minutes using that code. Now we need to speed up the test dataframe inference code. Then we can perform 100's of experiments quickly (in a few hours)!</p>",
          "rawMarkdown": "Thanks Radek. To run many experiments, we need to increase the speed of our entire pipeline. Using 1xV100 32GB, we can compute co-visitation matrices in 2 minutes using that code. Now we need to speed up the test dataframe inference code. Then we can perform 100's of experiments quickly (in a few hours)!",
          "votes": 3
        },
        {
          "id": 2025475,
          "postDate": "2022-11-11T08:41:59.943Z",
          "content": "<p>That sounds really awesome, on a completely another level! 🙌 Wonder what hyperparameters you'll be able to find with this approach! 🙂</p>\n<p>I didn't have very high hopes <code>cudf</code> could do well here, but I was mistaken 😄 GPUs strike again! Learned a lot from your code, some really awesome use of <code>cumcount</code> and sorting to emulate <code>head</code> and <code>tail</code>! And of course that trick where you are using series to replace the Counter/defaultdict combo! Really awesome stuff 🙂</p>\n<p>I have not tested it fully yet, still have a bit more code to change, but I wanted to operate on complete files instead of chunks. Was able to speed up calculating a single co-visitation matrix to 1 min 8 seconds 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F15ce231219e39c8f2f78d4b97fa8df60%2Fco_visitation_matrix.png?generation=1668155980264714&amp;alt=media\" alt=\"\"> </p>\n<p>Here is the code in case it can be of use:</p>\n<pre><code>%%time\n\nchunk_size = 2_000_000\n\ntmp = None\nfor i in range(0, sessions.shape[0], chunk_size):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n&lt;30].drop('n',axis=1)\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()&lt; 24 * 60 * 60) &amp; (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n\n\n# CONVERT MATRIX TO DICTIONARY\ntmp = tmp.reset_index()\ntmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n# SAVE TOP 40\ntmp = tmp.reset_index(drop=True)\ntmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\ntmp = tmp.loc[tmp.n&lt;40].drop('n',axis=1)\n# SAVE TO DISK\ndf = tmp.to_pandas().groupby('aid_x').aid_y.apply(list)\nwith open(f'{PATH}/top_40_carts_orders_v{VER}.pkl', 'wb') as f:\n    pickle.dump(df.to_dict(), f)\n</code></pre>\n<p>I ran this on a Quadro with 48GB of RAM, but you lose only a couple of seconds if you make the <code>chunk_size</code> much smaller.</p>\n<p>Anyhow, will only know if I didn't screw anything up once I change the rest of the notebook and run it all 🙂</p>",
          "rawMarkdown": "That sounds really awesome, on a completely another level! 🙌 Wonder what hyperparameters you'll be able to find with this approach! 🙂\n\nI didn't have very high hopes `cudf` could do well here, but I was mistaken 😄 GPUs strike again! Learned a lot from your code, some really awesome use of `cumcount` and sorting to emulate `head` and `tail`! And of course that trick where you are using series to replace the Counter/defaultdict combo! Really awesome stuff 🙂\n\nI have not tested it fully yet, still have a bit more code to change, but I wanted to operate on complete files instead of chunks. Was able to speed up calculating a single co-visitation matrix to 1 min 8 seconds 🙂\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F15ce231219e39c8f2f78d4b97fa8df60%2Fco_visitation_matrix.png?generation=1668155980264714&alt=media) \n\nHere is the code in case it can be of use:\n\n```\n%%time\n\nchunk_size = 2_000_000\n\ntmp = None\nfor i in range(0, sessions.shape[0], chunk_size):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n<30].drop('n',axis=1)\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n    \n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n        \n        \n# CONVERT MATRIX TO DICTIONARY\ntmp = tmp.reset_index()\ntmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n# SAVE TOP 40\ntmp = tmp.reset_index(drop=True)\ntmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\ntmp = tmp.loc[tmp.n<40].drop('n',axis=1)\n# SAVE TO DISK\ndf = tmp.to_pandas().groupby('aid_x').aid_y.apply(list)\nwith open(f'{PATH}/top_40_carts_orders_v{VER}.pkl', 'wb') as f:\n    pickle.dump(df.to_dict(), f)\n```\n\nI ran this on a Quadro with 48GB of RAM, but you lose only a couple of seconds if you make the `chunk_size` much smaller.\n\nAnyhow, will only know if I didn't screw anything up once I change the rest of the notebook and run it all 🙂",
          "votes": 1
        },
        {
          "id": 2025491,
          "postDate": "2022-11-11T08:49:26.860Z",
          "content": "<p>👀</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F8e2f2a1ad7eaf8aafb31b8f3035f0bfd%2Ffast.png?generation=1668156556627991&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "👀\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F8e2f2a1ad7eaf8aafb31b8f3035f0bfd%2Ffast.png?generation=1668156556627991&alt=media)\n",
          "votes": 1
        },
        {
          "id": 2025540,
          "postDate": "2022-11-11T09:17:23.727Z",
          "content": "<p>I just realized I am running this on my validation set, so that is nearly a quarter less of data… that would also contribute to this being quite a bit faster 🤦‍♂️</p>",
          "rawMarkdown": "I just realized I am running this on my validation set, so that is nearly a quarter less of data... that would also contribute to this being quite a bit faster 🤦‍♂️"
        },
        {
          "id": 2025901,
          "postDate": "2022-11-11T15:50:17.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/Radek\" target=\"_blank\">@Radek</a>, i notice that your code removes the inner and outer loops. Using nested loops is for improved speed not memory management. First we pick a comfortable chunk size that the GPU can handle. Afterward say that we have 12 chunks (i.e. first level in diagram below). Instead of merging all the chunks sequentially, it is faster to merge the chunks into 3 larger pieces then merge those 3 pieces into our final dataframe. Only the first level is using groupby and processing, so the GPU can handle the larger second level pieces for the simple operation of merge. To maximize speed, you need to experiment with level 1 and level 2 sizes on your GPU:</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/merge.png\" alt=\"\"></p>",
          "rawMarkdown": "@Radek, i notice that your code removes the inner and outer loops. Using nested loops is for improved speed not memory management. First we pick a comfortable chunk size that the GPU can handle. Afterward say that we have 12 chunks (i.e. first level in diagram below). Instead of merging all the chunks sequentially, it is faster to merge the chunks into 3 larger pieces then merge those 3 pieces into our final dataframe. Only the first level is using groupby and processing, so the GPU can handle the larger second level pieces for the simple operation of merge. To maximize speed, you need to experiment with level 1 and level 2 sizes on your GPU:\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/merge.png)",
          "votes": 8
        },
        {
          "id": 2026289,
          "postDate": "2022-11-11T22:30:01.780Z",
          "content": "<p>wow, thank you very much for this, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!!! I completely didn't realize this, that this reasoning would apply here. I thought it was all for memory management, but it all makes sense 🙂</p>\n<p>I have encountered something similar with DL models where beyond a certain batch size the computation is slower, time per epoch grows. But I have been wondering to what extent it is because of memory transfer issues vs how much work the GPU needs to do with bigger chunks. But here this is a clear example where operating on smaller pieces of data is faster! Amazing! 😄</p>\n<p>Thank you very much for your explanation 🙏</p>",
          "rawMarkdown": "wow, thank you very much for this, @cdeotte!!! I completely didn't realize this, that this reasoning would apply here. I thought it was all for memory management, but it all makes sense 🙂\n\nI have encountered something similar with DL models where beyond a certain batch size the computation is slower, time per epoch grows. But I have been wondering to what extent it is because of memory transfer issues vs how much work the GPU needs to do with bigger chunks. But here this is a clear example where operating on smaller pieces of data is faster! Amazing! 😄\n\nThank you very much for your explanation 🙏"
        },
        {
          "id": 2026343,
          "postDate": "2022-11-12T01:00:13.760Z",
          "content": "<p>You two deserve more upvotes 😄</p>",
          "rawMarkdown": "You two deserve more upvotes 😄",
          "votes": 2
        },
        {
          "id": 2026350,
          "postDate": "2022-11-12T01:19:08.673Z",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I get an error that this method doesn't exist for DataFrames <code>.to_pandas()</code>, since <code>tmp</code> is already a DataFrame. Is that expected?</p>",
          "rawMarkdown": "@radek1 I get an error that this method doesn't exist for DataFrames `.to_pandas()`, since `tmp` is already a DataFrame. Is that expected?"
        },
        {
          "id": 2026367,
          "postDate": "2022-11-12T02:18:07.620Z",
          "content": "<p>I didn't get any errors when running this code. Not at my computer ATM, can't check, but the df here in this code are cudf dataframes (not pandas) and IIRC tmp is a cudf series, maybe this can help</p>",
          "rawMarkdown": "I didn't get any errors when running this code. Not at my computer ATM, can't check, but the df here in this code are cudf dataframes (not pandas) and IIRC tmp is a cudf series, maybe this can help",
          "votes": 1
        },
        {
          "id": 2027353,
          "postDate": "2022-11-12T18:47:53.953Z",
          "content": "<p>Radek, I have some questions about your code, if you don't mind helping me.</p>\n<p>For example, what is the variable \"sessions\"? The dataframe \"data\" is indexed by sessions? If now wouldn't some sessions divided in different chunks?</p>\n<p>And the \"tmp.add\" code, this tmp has a fixed shape, so I don't understand how it is getting the information of all aid pairs.</p>\n<p>Thanks!                  </p>",
          "rawMarkdown": "Radek, I have some questions about your code, if you don't mind helping me.\n\nFor example, what is the variable \"sessions\"? The dataframe \"data\" is indexed by sessions? If now wouldn't some sessions divided in different chunks?\n\nAnd the \"tmp.add\" code, this tmp has a fixed shape, so I don't understand how it is getting the information of all aid pairs.\n\nThanks!                  "
        },
        {
          "id": 2027537,
          "postDate": "2022-11-12T22:31:50.797Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/gabrielmoraesbarros\" target=\"_blank\">@gabrielmoraesbarros</a>!</p>\n<p>I am including the rest of the code here, might be easiest for you to step through it and see how things work 🙂</p>\n<pre><code>VER = 1\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)\n\nPATH='valid'\n\ntrain = cudf.read_parquet(f'{PATH}/train.parquet')\ntest = cudf.read_parquet(f'{PATH}/test.parquet')\n\ndata = cudf.concat([train, test])\n\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\ntype_weight = {0:1, 1:6, 2:3}\n\ndata = data.set_index('session')\nsessions = data.index.unique()\n\nchunk_size = 2_000_000\n\ntmp = None\nfor i in range(0, sessions.shape[0], chunk_size):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    # USE TAIL OF SESSION\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n&lt;30].drop('n',axis=1)\n\n    # CREATE PAIRS\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()&lt; 24 * 60 * 60) &amp; (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n\n\n# CONVERT MATRIX TO DICTIONARY\ntmp = tmp.reset_index()\ntmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n# SAVE TOP 40\ntmp = tmp.reset_index(drop=True)\ntmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\ntmp = tmp.loc[tmp.n&lt;40].drop('n',axis=1)\n# SAVE TO DISK\ndf = tmp.to_pandas().groupby('aid_x').aid_y.apply(list)\nwith open(f'{PATH}/top_40_carts_orders_v{VER}.pkl', 'wb') as f:\n    pickle.dump(df.to_dict(), f)\n</code></pre>\n<p>This is the unoptimized code, without the two loops from Chris's code. Meaning, it is not going to be as fast as it could be, but is possibly easier to reason about.</p>\n<p>Hope this helps! 🙂</p>",
          "rawMarkdown": "Hey @gabrielmoraesbarros!\n\nI am including the rest of the code here, might be easiest for you to step through it and see how things work 🙂\n\n```\nVER = 1\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)\n\nPATH='valid'\n\ntrain = cudf.read_parquet(f'{PATH}/train.parquet')\ntest = cudf.read_parquet(f'{PATH}/test.parquet')\n\ndata = cudf.concat([train, test])\n\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\ntype_weight = {0:1, 1:6, 2:3}\n\ndata = data.set_index('session')\nsessions = data.index.unique()\n\nchunk_size = 2_000_000\n\ntmp = None\nfor i in range(0, sessions.shape[0], chunk_size):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n    \n    # USE TAIL OF SESSION\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n<30].drop('n',axis=1)\n    \n    # CREATE PAIRS\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n    \n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n        \n        \n# CONVERT MATRIX TO DICTIONARY\ntmp = tmp.reset_index()\ntmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n# SAVE TOP 40\ntmp = tmp.reset_index(drop=True)\ntmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\ntmp = tmp.loc[tmp.n<40].drop('n',axis=1)\n# SAVE TO DISK\ndf = tmp.to_pandas().groupby('aid_x').aid_y.apply(list)\nwith open(f'{PATH}/top_40_carts_orders_v{VER}.pkl', 'wb') as f:\n    pickle.dump(df.to_dict(), f)\n```\n\nThis is the unoptimized code, without the two loops from Chris's code. Meaning, it is not going to be as fast as it could be, but is possibly easier to reason about.\n\nHope this helps! 🙂"
        },
        {
          "id": 2027568,
          "postDate": "2022-11-13T01:12:08.027Z",
          "content": "<p>Thanks, man, appreciated.</p>\n<p>The \"sessions\" part was exactly what I was thinking, but I still suspicious about the \"tmp.add\" thing.</p>\n<p>Here is some snippet with your parquets <strong>otto-full-optimized-memory-footprint</strong>.</p>\n<pre><code>VER = 1\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)\n\nPATH='valid'\n\n# train = cudf.read_parquet(f'/kaggle/input/train.parquet')\ntest = cudf.read_parquet(f'/kaggle/input/otto-full-optimized-memory-footprint/test.parquet')\n\n# data = cudf.concat([train, test])\ndata = test\n\n\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\ntype_weight = {0:1, 1:6, 2:3}\n\ndata = data.set_index('session')\nsessions = data.index.unique()\n\nchunk_size = 10_000\n\nprint(data.shape)\n\ntmp = None\nlist_of_dataframes = []\nfor it, i in enumerate(range(0, sessions.shape[0], chunk_size)):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    # USE TAIL OF SESSION\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n&lt;30].drop('n',axis=1)\n\n    # CREATE PAIRS\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()&lt; 24 * 60 * 60) &amp; (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n    list_of_dataframes.append(df)\n    print('Iteration = {1}, tmp shape = {1} and pd.concat shape {2}'.format(it,tmp.shape,cudf.concat(list_of_dataframes, axis=0).shape))\n    print()\n    if it &gt; 10:\n          break\n</code></pre>",
          "rawMarkdown": "Thanks, man, appreciated.\n\nThe \"sessions\" part was exactly what I was thinking, but I still suspicious about the \"tmp.add\" thing.\n\nHere is some snippet with your parquets **otto-full-optimized-memory-footprint**.\n\n\n```\nVER = 1\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)\n\nPATH='valid'\n\n# train = cudf.read_parquet(f'/kaggle/input/train.parquet')\ntest = cudf.read_parquet(f'/kaggle/input/otto-full-optimized-memory-footprint/test.parquet')\n\n# data = cudf.concat([train, test])\ndata = test\n\n\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\ntype_weight = {0:1, 1:6, 2:3}\n\ndata = data.set_index('session')\nsessions = data.index.unique()\n\nchunk_size = 10_000\n\nprint(data.shape)\n\ntmp = None\nlist_of_dataframes = []\nfor it, i in enumerate(range(0, sessions.shape[0], chunk_size)):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    # USE TAIL OF SESSION\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n<30].drop('n',axis=1)\n\n    # CREATE PAIRS\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n    list_of_dataframes.append(df)\n    print('Iteration = {1}, tmp shape = {1} and pd.concat shape {2}'.format(it,tmp.shape,cudf.concat(list_of_dataframes, axis=0).shape))\n    print()\n    if it > 10:\n          break\n        \n```\n\n"
        },
        {
          "id": 2027570,
          "postDate": "2022-11-13T01:15:01.453Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2027587,
          "postDate": "2022-11-13T01:25:54.367Z",
          "content": "<p>The command <code>tmp.add()</code> uses both dataframe indexes to perform the add (and the column being added is <code>wgt</code>). The previous command is <code>df = df.groupby(['aid_x','aid_y']).wgt.sum()</code> this makes the index a multilevel index with <code>['aid_x','aid_y']</code>. So when we <code>tmp.add()</code> it  adds to the weights of existing pairs of <code>aid_x</code> and <code>aid_y</code>. And if the pair does not exist, it makes a new pair.</p>",
          "rawMarkdown": "The command `tmp.add()` uses both dataframe indexes to perform the add (and the column being added is `wgt`). The previous command is `df = df.groupby(['aid_x','aid_y']).wgt.sum()` this makes the index a multilevel index with `['aid_x','aid_y']`. So when we `tmp.add()` it  adds to the weights of existing pairs of `aid_x` and `aid_y`. And if the pair does not exist, it makes a new pair.",
          "votes": 3
        },
        {
          "id": 2027611,
          "postDate": "2022-11-13T02:12:09.813Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2027629,
          "postDate": "2022-11-13T03:12:38.830Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>\n<p>Now I understand my confusion:</p>\n<pre><code> if tmp is None:\n     tmp = df\nelse:\n      tmp.add(df, fill_value=0)\n</code></pre>\n<p>The last line should be: </p>\n<pre><code>tmp = tmp.add(df, fill_value=0)\n</code></pre>\n<p>Which is equivalent to:<br>\n<code>tmp = (tmp + df).fillna(0)</code></p>\n<p>Thanks for the replies.</p>",
          "rawMarkdown": "@cdeotte and @radek1 \n\nNow I understand my confusion:\n\n```\n if tmp is None:\n     tmp = df\nelse:\n      tmp.add(df, fill_value=0)\n```\n\nThe last line should be: \n\n```\ntmp = tmp.add(df, fill_value=0)\n```\n\nWhich is equivalent to:\n`tmp = (tmp + df).fillna(0)`\n\nThanks for the replies.",
          "votes": 3
        },
        {
          "id": 2028426,
          "postDate": "2022-11-13T22:39:51.030Z",
          "content": "<p>Ah so there is a bug! Well spotted <a href=\"https://www.kaggle.com/gabrielmoraesbarros\" target=\"_blank\">@gabrielmoraesbarros</a>! 🙂</p>",
          "rawMarkdown": "Ah so there is a bug! Well spotted @gabrielmoraesbarros! 🙂",
          "votes": 1
        }
      ]
    },
    {
      "id": 2056638,
      "postDate": "2022-12-06T10:17:03.190Z",
      "content": "<p>Hey Chris!</p>\n<p>Have you ever encountered this error: MemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:70: cudaErrorMemoryAllocation out of memory?</p>\n<p>What would you do then?</p>\n<p>Thanks. Cheers!</p>",
      "rawMarkdown": "Hey Chris!\n\nHave you ever encountered this error: MemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:70: cudaErrorMemoryAllocation out of memory?\n\nWhat would you do then?\n\nThanks. Cheers!",
      "replies": [
        {
          "id": 2056813,
          "postDate": "2022-12-06T13:05:46.847Z",
          "content": "<p>In this case, increase the variable <code>DISK_PIECES</code>. This will break the processing into more chunks with each chunk using less memory. And avoid memory error.</p>",
          "rawMarkdown": "In this case, increase the variable `DISK_PIECES`. This will break the processing into more chunks with each chunk using less memory. And avoid memory error.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2031150,
      "postDate": "2022-11-15T22:39:19.343Z",
      "content": "<p>Just start learning RAPIDS, thanks!</p>",
      "rawMarkdown": "Just start learning RAPIDS, thanks!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2028397,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-11-13T21:32:00.943000",
      "content": "<p>UPDATE: I updated my code to create co-visitation matrices 1.5x faster. It now creates co-visitation matrices in Kaggle notebooks at 6 minutes! (And offline using 1xV100 in 1.5 minutes!)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2032873,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-16T21:01:18.247000",
          "content": "<p>UPDATE: 🔥 I updated my code again to create co-visitation matrices 2x faster. It now creates co-visitation matrices in Kaggle notebooks at <strong>3 minutes</strong>! (This is <strong>3x</strong> faster than version 1 of the notebook. And <strong>30x</strong> faster than public notebooks using Pandas CPU) 🔥</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2035237,
      "author_name": "Andrej Zubaľ",
      "author_url": "",
      "post_date": "2022-11-18T17:59:08.493000",
      "content": "<p>Hi Chris,</p>\n<p>Thanks for showing us this wonderful solution 👍. </p>\n<p>I find it hard to improve on this simple heuristic approach combining information on IN_SESSION flag, TOP most popular items for given item in the session and TOP most popular items overall.</p>\n<p>I was wondering, how did you come up with this heuristic? Is it something that you tried out before and it worked, or you tried out bunch of different techniques and this one worked the best or simply you had a lucky/intuitive guess?</p>\n<p>Best </p>\n<p>Andrej</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2036565,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-19T22:22:38.557000",
          "content": "<p>There are two ways to improve CV and LB. Find more <strong>candidates</strong> and find better <strong>rerank</strong> logic. I discover both by using intuition and thinking about how users behave. Then i evaluate my ideas by observing if local validation score increases or not.</p>\n<p>The easiest way to improve my public notebook is to create more co-visitation matrices. It is possible to improve my notebook by changing the rerank logic, but it will be easier to just make more co-visitation matrices and add their weight in.</p>\n<p>I suggest that you focus on the case of <code>suggest_buys</code> for <code>len(unique_aids)&lt;20</code>. This has the largest effect on CV score and LB score. Change the code to allow for weights. The current code is</p>\n<pre><code># USE \"CART ORDER\" CO-VISITATION MATRIX\naids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n# USE \"BUY2BUY\" CO-VISITATION MATRIX\naids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n# RERANK CANDIDATES\ntop_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n</code></pre>\n<p>Change this to</p>\n<pre><code>from collections import Counter\naids_temp = Counter() \n# USE \"CART ORDER\" CO-VISITATION MATRIX\naids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\nfor aid in aids2: aids_temp[aid] += 1\n# USE \"BUY2BUY\" CO-VISITATION MATRIX\naids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\nfor aid in aids3: aids_temp[aid] += 1\n# RERANK CANDIDATES\ntop_aids2 = [k for k,v in aids_temp.most_common(20) if k not in unique_aids]\n</code></pre>\n<p>Now we can create new co-visitation matrices and include them with</p>\n<pre><code>aids4 = list(itertools.chain(*[NEW_MATRIX[aid] for aid in SUBSET_OF_USER_HISTORY_ITEMS if aid in NEW_MATRIX]))\nfor aid in aids4: aids_temp[aid] += 1\n</code></pre>\n<p>You need to be creative when making co-visitation matrices. The matrices makes pairs of items. If a <strong>certain type of user</strong> interacts in a <strong>certain type of way with item</strong>. Then we predict they will click and order a certain item in the future. Then when you include this new co-visitation matrix in your inference, we only apply it to the same <strong>certain type of user</strong> when they interact with an item in a <strong>certain type of way</strong>.</p>\n<p>Also note that once you change the code to include <code>for aid in aids2: aids_temp[aid] += 1</code>, you can now assign different situations different amount of weights during inference. This is how we change the <code>rerank</code> logic. </p>\n<pre><code>weights = LOGIC_TO_MAKES_WEIGHTS\nfor aid,w in zip(aids2,weights):\n    aids_temp[aid] += w\n</code></pre>",
          "votes": 31,
          "replies": []
        }
      ]
    },
    {
      "id": 2032899,
      "author_name": "Simon Veitner",
      "author_url": "",
      "post_date": "2022-11-16T21:21:55.620000",
      "content": "<p>this is so valuable and lets us experiment with different weighting approaches in an acceptable time! tnx chris!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2027772,
      "author_name": "jazivxt",
      "author_url": "",
      "post_date": "2022-11-13T07:20:54.380000",
      "content": "<p>Very awesome <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, not sure how to implement but curious if the session was treated as a corpus where each paring is a word and the cart and order events are treated as words in the sequence then a generator can just simulate the next 20 words which can include a cart or order event in the sequence, guessing they are not too accurate yet as they are randomized for creativity but that can also add some flavor for the buyers although in this simulation we also have to predict the behavior of current suggestion models. Back to your model, great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2026387,
      "author_name": "GabrielMoraesBarros",
      "author_url": "",
      "post_date": "2022-11-12T03:02:50.447000",
      "content": "<p>Great kernel, as always <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2030572,
      "author_name": "Satya",
      "author_url": "",
      "post_date": "2022-11-15T14:23:26.880000",
      "content": "<p>How about the modin library it is bit faster in some areas</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2030951,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-15T18:29:55.410000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/satyaprakashshukl\" target=\"_blank\">@satyaprakashshukl</a> for the suggestion. I was not aware of Modin library. I'll check it out. Another option to accelerate Pandas (groupby functions) on CPU is <code>pandarallel</code> <a href=\"https://pypi.org/project/pandarallel/\" target=\"_blank\">here</a>. I tried out <code>pandarallel</code> on my public notebook inference and it boosted inference offline Kaggle from 30 minutes to 5 minutes using 20 CPU! (discussion <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2026011\" target=\"_blank\">here</a>)</p>\n<p>Eventually i will rewrite my inference using RAPIDS cuDF and expect even more speed boost than pandarallel.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2031356,
          "author_name": "Satya",
          "author_url": "",
          "post_date": "2022-11-16T03:05:34.503000",
          "content": "<p>Thank you chris,i will explore too  these option in my kernel  🙏🏽🙂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2032901,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-11-16T21:27:09.273000",
          "content": "<p>that would be great chris! than nothing is stopping us from trying even more experiments!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2025082,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-11-11T01:00:44.913000",
      "content": "<p>Those speed gains are insane! 🙂 Didn't assume this workload could be transformed to run fast on the GPU… and yet here we are! 😄 </p>\n<p>Have to look more closely at how you are creating the co-visitation matrix dicts. That is some real magic happening there and something that has been an extremely slow part of the process thus far (when we went <code>matrix[aid_x][aid_y] + w</code>)!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2025091,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-11T01:05:38.060000",
          "content": "<p>Thanks Radek. To run many experiments, we need to increase the speed of our entire pipeline. Using 1xV100 32GB, we can compute co-visitation matrices in 2 minutes using that code. Now we need to speed up the test dataframe inference code. Then we can perform 100's of experiments quickly (in a few hours)!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2025475,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-11T08:41:59.943000",
          "content": "<p>That sounds really awesome, on a completely another level! 🙌 Wonder what hyperparameters you'll be able to find with this approach! 🙂</p>\n<p>I didn't have very high hopes <code>cudf</code> could do well here, but I was mistaken 😄 GPUs strike again! Learned a lot from your code, some really awesome use of <code>cumcount</code> and sorting to emulate <code>head</code> and <code>tail</code>! And of course that trick where you are using series to replace the Counter/defaultdict combo! Really awesome stuff 🙂</p>\n<p>I have not tested it fully yet, still have a bit more code to change, but I wanted to operate on complete files instead of chunks. Was able to speed up calculating a single co-visitation matrix to 1 min 8 seconds 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F15ce231219e39c8f2f78d4b97fa8df60%2Fco_visitation_matrix.png?generation=1668155980264714&amp;alt=media\" alt=\"\"> </p>\n<p>Here is the code in case it can be of use:</p>\n<pre><code>%%time\n\nchunk_size = 2_000_000\n\ntmp = None\nfor i in range(0, sessions.shape[0], chunk_size):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n&lt;30].drop('n',axis=1)\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()&lt; 24 * 60 * 60) &amp; (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n\n\n# CONVERT MATRIX TO DICTIONARY\ntmp = tmp.reset_index()\ntmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n# SAVE TOP 40\ntmp = tmp.reset_index(drop=True)\ntmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\ntmp = tmp.loc[tmp.n&lt;40].drop('n',axis=1)\n# SAVE TO DISK\ndf = tmp.to_pandas().groupby('aid_x').aid_y.apply(list)\nwith open(f'{PATH}/top_40_carts_orders_v{VER}.pkl', 'wb') as f:\n    pickle.dump(df.to_dict(), f)\n</code></pre>\n<p>I ran this on a Quadro with 48GB of RAM, but you lose only a couple of seconds if you make the <code>chunk_size</code> much smaller.</p>\n<p>Anyhow, will only know if I didn't screw anything up once I change the rest of the notebook and run it all 🙂</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2025491,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-11T08:49:26.860000",
          "content": "<p>👀</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F8e2f2a1ad7eaf8aafb31b8f3035f0bfd%2Ffast.png?generation=1668156556627991&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2025540,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-11T09:17:23.727000",
          "content": "<p>I just realized I am running this on my validation set, so that is nearly a quarter less of data… that would also contribute to this being quite a bit faster 🤦‍♂️</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2025901,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-11T15:50:17.677000",
          "content": "<p><a href=\"https://www.kaggle.com/Radek\" target=\"_blank\">@Radek</a>, i notice that your code removes the inner and outer loops. Using nested loops is for improved speed not memory management. First we pick a comfortable chunk size that the GPU can handle. Afterward say that we have 12 chunks (i.e. first level in diagram below). Instead of merging all the chunks sequentially, it is faster to merge the chunks into 3 larger pieces then merge those 3 pieces into our final dataframe. Only the first level is using groupby and processing, so the GPU can handle the larger second level pieces for the simple operation of merge. To maximize speed, you need to experiment with level 1 and level 2 sizes on your GPU:</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/merge.png\" alt=\"\"></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 2026289,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-11T22:30:01.780000",
          "content": "<p>wow, thank you very much for this, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!!! I completely didn't realize this, that this reasoning would apply here. I thought it was all for memory management, but it all makes sense 🙂</p>\n<p>I have encountered something similar with DL models where beyond a certain batch size the computation is slower, time per epoch grows. But I have been wondering to what extent it is because of memory transfer issues vs how much work the GPU needs to do with bigger chunks. But here this is a clear example where operating on smaller pieces of data is faster! Amazing! 😄</p>\n<p>Thank you very much for your explanation 🙏</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2026343,
          "author_name": "dpalbrecht",
          "author_url": "",
          "post_date": "2022-11-12T01:00:13.760000",
          "content": "<p>You two deserve more upvotes 😄</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2026350,
          "author_name": "dpalbrecht",
          "author_url": "",
          "post_date": "2022-11-12T01:19:08.673000",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I get an error that this method doesn't exist for DataFrames <code>.to_pandas()</code>, since <code>tmp</code> is already a DataFrame. Is that expected?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2026367,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-12T02:18:07.620000",
          "content": "<p>I didn't get any errors when running this code. Not at my computer ATM, can't check, but the df here in this code are cudf dataframes (not pandas) and IIRC tmp is a cudf series, maybe this can help</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2027353,
          "author_name": "GabrielMoraesBarros",
          "author_url": "",
          "post_date": "2022-11-12T18:47:53.953000",
          "content": "<p>Radek, I have some questions about your code, if you don't mind helping me.</p>\n<p>For example, what is the variable \"sessions\"? The dataframe \"data\" is indexed by sessions? If now wouldn't some sessions divided in different chunks?</p>\n<p>And the \"tmp.add\" code, this tmp has a fixed shape, so I don't understand how it is getting the information of all aid pairs.</p>\n<p>Thanks!                  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2027537,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-12T22:31:50.797000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/gabrielmoraesbarros\" target=\"_blank\">@gabrielmoraesbarros</a>!</p>\n<p>I am including the rest of the code here, might be easiest for you to step through it and see how things work 🙂</p>\n<pre><code>VER = 1\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)\n\nPATH='valid'\n\ntrain = cudf.read_parquet(f'{PATH}/train.parquet')\ntest = cudf.read_parquet(f'{PATH}/test.parquet')\n\ndata = cudf.concat([train, test])\n\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\ntype_weight = {0:1, 1:6, 2:3}\n\ndata = data.set_index('session')\nsessions = data.index.unique()\n\nchunk_size = 2_000_000\n\ntmp = None\nfor i in range(0, sessions.shape[0], chunk_size):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    # USE TAIL OF SESSION\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n&lt;30].drop('n',axis=1)\n\n    # CREATE PAIRS\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()&lt; 24 * 60 * 60) &amp; (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n\n\n# CONVERT MATRIX TO DICTIONARY\ntmp = tmp.reset_index()\ntmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n# SAVE TOP 40\ntmp = tmp.reset_index(drop=True)\ntmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\ntmp = tmp.loc[tmp.n&lt;40].drop('n',axis=1)\n# SAVE TO DISK\ndf = tmp.to_pandas().groupby('aid_x').aid_y.apply(list)\nwith open(f'{PATH}/top_40_carts_orders_v{VER}.pkl', 'wb') as f:\n    pickle.dump(df.to_dict(), f)\n</code></pre>\n<p>This is the unoptimized code, without the two loops from Chris's code. Meaning, it is not going to be as fast as it could be, but is possibly easier to reason about.</p>\n<p>Hope this helps! 🙂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2027568,
          "author_name": "GabrielMoraesBarros",
          "author_url": "",
          "post_date": "2022-11-13T01:12:08.027000",
          "content": "<p>Thanks, man, appreciated.</p>\n<p>The \"sessions\" part was exactly what I was thinking, but I still suspicious about the \"tmp.add\" thing.</p>\n<p>Here is some snippet with your parquets <strong>otto-full-optimized-memory-footprint</strong>.</p>\n<pre><code>VER = 1\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)\n\nPATH='valid'\n\n# train = cudf.read_parquet(f'/kaggle/input/train.parquet')\ntest = cudf.read_parquet(f'/kaggle/input/otto-full-optimized-memory-footprint/test.parquet')\n\n# data = cudf.concat([train, test])\ndata = test\n\n\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\ntype_weight = {0:1, 1:6, 2:3}\n\ndata = data.set_index('session')\nsessions = data.index.unique()\n\nchunk_size = 10_000\n\nprint(data.shape)\n\ntmp = None\nlist_of_dataframes = []\nfor it, i in enumerate(range(0, sessions.shape[0], chunk_size)):\n    df = data.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index()\n\n    df = df.sort_values(['session','ts'],ascending=[True,False])\n\n    # USE TAIL OF SESSION\n    df = df.reset_index(drop=True)\n    df['n'] = df.groupby('session').cumcount()\n    df = df.loc[df.n&lt;30].drop('n',axis=1)\n\n    # CREATE PAIRS\n    df = df.merge(df,on='session')\n    df = df.loc[ ((df.ts_x - df.ts_y).abs()&lt; 24 * 60 * 60) &amp; (df.aid_x != df.aid_y) ]\n\n    # ASSIGN WEIGHTS\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    df['wgt'] = df.type_y.map(type_weight)\n    df = df[['aid_x','aid_y','wgt']]\n    df.wgt = df.wgt.astype('float32')\n    df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n    if tmp is None: tmp = df\n    else: tmp.add(df, fill_value=0)\n    list_of_dataframes.append(df)\n    print('Iteration = {1}, tmp shape = {1} and pd.concat shape {2}'.format(it,tmp.shape,cudf.concat(list_of_dataframes, axis=0).shape))\n    print()\n    if it &gt; 10:\n          break\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2027570,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-13T01:15:01.453000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2027587,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-13T01:25:54.367000",
          "content": "<p>The command <code>tmp.add()</code> uses both dataframe indexes to perform the add (and the column being added is <code>wgt</code>). The previous command is <code>df = df.groupby(['aid_x','aid_y']).wgt.sum()</code> this makes the index a multilevel index with <code>['aid_x','aid_y']</code>. So when we <code>tmp.add()</code> it  adds to the weights of existing pairs of <code>aid_x</code> and <code>aid_y</code>. And if the pair does not exist, it makes a new pair.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2027611,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-13T02:12:09.813000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2027629,
          "author_name": "GabrielMoraesBarros",
          "author_url": "",
          "post_date": "2022-11-13T03:12:38.830000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>\n<p>Now I understand my confusion:</p>\n<pre><code> if tmp is None:\n     tmp = df\nelse:\n      tmp.add(df, fill_value=0)\n</code></pre>\n<p>The last line should be: </p>\n<pre><code>tmp = tmp.add(df, fill_value=0)\n</code></pre>\n<p>Which is equivalent to:<br>\n<code>tmp = (tmp + df).fillna(0)</code></p>\n<p>Thanks for the replies.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2028426,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-13T22:39:51.030000",
          "content": "<p>Ah so there is a bug! Well spotted <a href=\"https://www.kaggle.com/gabrielmoraesbarros\" target=\"_blank\">@gabrielmoraesbarros</a>! 🙂</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2056638,
      "author_name": "enizzzz",
      "author_url": "",
      "post_date": "2022-12-06T10:17:03.190000",
      "content": "<p>Hey Chris!</p>\n<p>Have you ever encountered this error: MemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:70: cudaErrorMemoryAllocation out of memory?</p>\n<p>What would you do then?</p>\n<p>Thanks. Cheers!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2056813,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-06T13:05:46.847000",
          "content": "<p>In this case, increase the variable <code>DISK_PIECES</code>. This will break the processing into more chunks with each chunk using less memory. And avoid memory error.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2031150,
      "author_name": "Peach9",
      "author_url": "",
      "post_date": "2022-11-15T22:39:19.343000",
      "content": "<p>Just start learning RAPIDS, thanks!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2025071": "I published a notebook [here][1] that uses RAPIDS cuDF to compute 30x faster co-visitation matrices. Using co-visitation matrices helps provide our models with \"candidates\". Then our models (and/or human logic) and \"rerank\" these \"candidates\" and select our final 20 predictions for submission CSV. \n\n**UPDATE** 🔥 latest version computes co-visitation matrices in **3min** each 🔥\n\nFor more information about \"candidate rerank\" models, see Ravi's discussion [here][2]\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\n[2]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721",
    "2028397": "UPDATE: I updated my code to create co-visitation matrices 1.5x faster. It now creates co-visitation matrices in Kaggle notebooks at 6 minutes! (And offline using 1xV100 in 1.5 minutes!)",
    "2035237": "Hi Chris,\n\nThanks for showing us this wonderful solution 👍. \n\nI find it hard to improve on this simple heuristic approach combining information on IN_SESSION flag, TOP most popular items for given item in the session and TOP most popular items overall.\n\nI was wondering, how did you come up with this heuristic? Is it something that you tried out before and it worked, or you tried out bunch of different techniques and this one worked the best or simply you had a lucky/intuitive guess?\n\nBest \n\nAndrej",
    "2032899": "this is so valuable and lets us experiment with different weighting approaches in an acceptable time! tnx chris!",
    "2027772": "Very awesome @cdeotte, not sure how to implement but curious if the session was treated as a corpus where each paring is a word and the cart and order events are treated as words in the sequence then a generator can just simulate the next 20 words which can include a cart or order event in the sequence, guessing they are not too accurate yet as they are randomized for creativity but that can also add some flavor for the buyers although in this simulation we also have to predict the behavior of current suggestion models. Back to your model, great work!",
    "2026387": "Great kernel, as always @cdeotte.\n\n",
    "2030572": "How about the modin library it is bit faster in some areas",
    "2025082": "Those speed gains are insane! 🙂 Didn't assume this workload could be transformed to run fast on the GPU... and yet here we are! 😄 \n\nHave to look more closely at how you are creating the co-visitation matrix dicts. That is some real magic happening there and something that has been an extremely slow part of the process thus far (when we went `matrix[aid_x][aid_y] + w`)!",
    "2056638": "Hey Chris!\n\nHave you ever encountered this error: MemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:70: cudaErrorMemoryAllocation out of memory?\n\nWhat would you do then?\n\nThanks. Cheers!",
    "2031150": "Just start learning RAPIDS, thanks!"
  }
}