{
  "id": 372945,
  "title": "GPU Inference by RAPIDS",
  "url": "/competitions/otto-recommender-system/discussion/372945",
  "author_name": "Bilzard",
  "post_date": "2022-12-18T20:55:53.799000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi.</p>\n<p>I implemented <a href=\"https://www.kaggle.com/code/tatamikenn/otto-gpu-inference-lb-0-555\" target=\"_blank\">a notebook with GPU inference</a> based on <a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">Chris Deotte's notebook</a> using CUDF.<br>\nEvaluation time becomes from 27min to few sec, which means hundreds times faster than the original notebook. That will boost your experimental turn-around time.</p>\n<p>Note: due to memory limitation of Kaggle notebook (16GB), I had to merge carts_orders &amp; buy2buy co-visitation matrix. It scored LB=0.555 (-0.02 from original score). In my local machine, where I fed carts_orders and buy2buy co-visitation matrix separately, it scored LB=0.572 which is very close to Chris's original notebook's score LB=0.575. If you had enough GPU memory, you can try below code that would reproduce LB=0.572:</p>\n<pre><code>\n\nsuggest_buy = suggest(test_df, [covis_carts_orders, covis_buy2buy], **params)\n</code></pre>\n<p>Also, the evaluation mode doesn't work in current Kaggle's RAPIDS version (MultiIndex.intersection isn't implemented). If you want to evaluate in Kaggle notebook, you have to fix evaluation code using pandas DataFrame (it would take ~1 min to evaluate). </p>\n<p>Enjoy!</p>\n<hr>\n<h3>Update</h3>\n<p>Merging two co-visitation matrix beforehand turned out to be a valid approach.<br>\nThe below code in my local environment scores LB=0.573.<br>\nSo, the score dropdown written above seems to be a problem of inconsistent LAPIDS version between my local environment and Kaggle notebook.</p>\n<pre><code>covis_carts_orders = cudf.concat([covis_carts_orders, covis_buy2buy])\ncovis_carts_orders = covis_carts_orders.groupby([, ])[].().reset_index()\nsuggest_buy = suggest(test_df, [covis_carts_orders], **params)\n</code></pre>\n<hr>\n<p>Changing <code>drop_duplicates(['session', 'aid_x', 'aid_y'])</code> to <code>drop_duplicates(['session', 'aid_x', 'aid_y', 'type_y'])</code></p>\n<p>This trick <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2044545\" target=\"_blank\">^1</a> boosts +0.001 score up on both CV(0.562 -&gt; 0.563) and LB (0.573 -&gt; 0.574).</p>",
  "messages": [
    {
      "id": 2069313,
      "postDate": "2022-12-18T20:55:53.800Z",
      "content": "<p>Hi.</p>\n<p>I implemented <a href=\"https://www.kaggle.com/code/tatamikenn/otto-gpu-inference-lb-0-555\" target=\"_blank\">a notebook with GPU inference</a> based on <a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">Chris Deotte's notebook</a> using CUDF.<br>\nEvaluation time becomes from 27min to few sec, which means hundreds times faster than the original notebook. That will boost your experimental turn-around time.</p>\n<p>Note: due to memory limitation of Kaggle notebook (16GB), I had to merge carts_orders &amp; buy2buy co-visitation matrix. It scored LB=0.555 (-0.02 from original score). In my local machine, where I fed carts_orders and buy2buy co-visitation matrix separately, it scored LB=0.572 which is very close to Chris's original notebook's score LB=0.575. If you had enough GPU memory, you can try below code that would reproduce LB=0.572:</p>\n<pre><code>\n\nsuggest_buy = suggest(test_df, [covis_carts_orders, covis_buy2buy], **params)\n</code></pre>\n<p>Also, the evaluation mode doesn't work in current Kaggle's RAPIDS version (MultiIndex.intersection isn't implemented). If you want to evaluate in Kaggle notebook, you have to fix evaluation code using pandas DataFrame (it would take ~1 min to evaluate). </p>\n<p>Enjoy!</p>\n<hr>\n<h3>Update</h3>\n<p>Merging two co-visitation matrix beforehand turned out to be a valid approach.<br>\nThe below code in my local environment scores LB=0.573.<br>\nSo, the score dropdown written above seems to be a problem of inconsistent LAPIDS version between my local environment and Kaggle notebook.</p>\n<pre><code>covis_carts_orders = cudf.concat([covis_carts_orders, covis_buy2buy])\ncovis_carts_orders = covis_carts_orders.groupby([, ])[].().reset_index()\nsuggest_buy = suggest(test_df, [covis_carts_orders], **params)\n</code></pre>\n<hr>\n<p>Changing <code>drop_duplicates(['session', 'aid_x', 'aid_y'])</code> to <code>drop_duplicates(['session', 'aid_x', 'aid_y', 'type_y'])</code></p>\n<p>This trick <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2044545\" target=\"_blank\">^1</a> boosts +0.001 score up on both CV(0.562 -&gt; 0.563) and LB (0.573 -&gt; 0.574).</p>",
      "rawMarkdown": "Hi.\n\nI implemented [a notebook with GPU inference] based on [Chris Deotte's notebook] using CUDF.\nEvaluation time becomes from 27min to few sec, which means hundreds times faster than the original notebook. That will boost your experimental turn-around time.\n\nNote: due to memory limitation of Kaggle notebook (16GB), I had to merge carts_orders & buy2buy co-visitation matrix. It scored LB=0.555 (-0.02 from original score). In my local machine, where I fed carts_orders and buy2buy co-visitation matrix separately, it scored LB=0.572 which is very close to Chris's original notebook's score LB=0.575. If you had enough GPU memory, you can try below code that would reproduce LB=0.572:\n\n```python\n# covis_carts_orders = cudf.concat([covis_carts_orders, covis_clicks])\n# covis_carts_orders = covis_carts_orders.groupby([\"aid_x\", \"aid_y\"])[\"wgt\"].sum().reset_index()\nsuggest_buy = suggest(test_df, [covis_carts_orders, covis_buy2buy], **params)\n```\n\nAlso, the evaluation mode doesn't work in current Kaggle's RAPIDS version (MultiIndex.intersection isn't implemented). If you want to evaluate in Kaggle notebook, you have to fix evaluation code using pandas DataFrame (it would take ~1 min to evaluate). \n\nEnjoy!\n\n[a notebook with GPU inference]: https://www.kaggle.com/code/tatamikenn/otto-gpu-inference-lb-0-555\n[Chris Deotte's notebook]: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\n\n---\n\n### Update\n\nMerging two co-visitation matrix beforehand turned out to be a valid approach.\nThe below code in my local environment scores LB=0.573.\nSo, the score dropdown written above seems to be a problem of inconsistent LAPIDS version between my local environment and Kaggle notebook.\n\n```python\ncovis_carts_orders = cudf.concat([covis_carts_orders, covis_buy2buy])\ncovis_carts_orders = covis_carts_orders.groupby([\"aid_x\", \"aid_y\"])[\"wgt\"].sum().reset_index()\nsuggest_buy = suggest(test_df, [covis_carts_orders], **params)\n```\n\n---\n\nChanging `drop_duplicates(['session', 'aid_x', 'aid_y'])` to `drop_duplicates(['session', 'aid_x', 'aid_y', 'type_y'])`\n\nThis trick [^1] boosts +0.001 score up on both CV(0.562 -> 0.563) and LB (0.573 -> 0.574).\n\n[^1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2044545\n",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2069313": "Hi.\n\nI implemented [a notebook with GPU inference] based on [Chris Deotte's notebook] using CUDF.\nEvaluation time becomes from 27min to few sec, which means hundreds times faster than the original notebook. That will boost your experimental turn-around time.\n\nNote: due to memory limitation of Kaggle notebook (16GB), I had to merge carts_orders & buy2buy co-visitation matrix. It scored LB=0.555 (-0.02 from original score). In my local machine, where I fed carts_orders and buy2buy co-visitation matrix separately, it scored LB=0.572 which is very close to Chris's original notebook's score LB=0.575. If you had enough GPU memory, you can try below code that would reproduce LB=0.572:\n\n```python\n# covis_carts_orders = cudf.concat([covis_carts_orders, covis_clicks])\n# covis_carts_orders = covis_carts_orders.groupby([\"aid_x\", \"aid_y\"])[\"wgt\"].sum().reset_index()\nsuggest_buy = suggest(test_df, [covis_carts_orders, covis_buy2buy], **params)\n```\n\nAlso, the evaluation mode doesn't work in current Kaggle's RAPIDS version (MultiIndex.intersection isn't implemented). If you want to evaluate in Kaggle notebook, you have to fix evaluation code using pandas DataFrame (it would take ~1 min to evaluate). \n\nEnjoy!\n\n[a notebook with GPU inference]: https://www.kaggle.com/code/tatamikenn/otto-gpu-inference-lb-0-555\n[Chris Deotte's notebook]: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\n\n---\n\n### Update\n\nMerging two co-visitation matrix beforehand turned out to be a valid approach.\nThe below code in my local environment scores LB=0.573.\nSo, the score dropdown written above seems to be a problem of inconsistent LAPIDS version between my local environment and Kaggle notebook.\n\n```python\ncovis_carts_orders = cudf.concat([covis_carts_orders, covis_buy2buy])\ncovis_carts_orders = covis_carts_orders.groupby([\"aid_x\", \"aid_y\"])[\"wgt\"].sum().reset_index()\nsuggest_buy = suggest(test_df, [covis_carts_orders], **params)\n```\n\n---\n\nChanging `drop_duplicates(['session', 'aid_x', 'aid_y'])` to `drop_duplicates(['session', 'aid_x', 'aid_y', 'type_y'])`\n\nThis trick [^1] boosts +0.001 score up on both CV(0.562 -> 0.563) and LB (0.573 -> 0.574).\n\n[^1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/comments#2044545\n"
  }
}