{
  "id": 371111,
  "title": "💡 ANN -> NN: get better results and run faster with NN search on the GPU 🔥🔥🔥",
  "url": "/competitions/otto-recommender-system/discussion/371111",
  "author_name": "Radek Osmulski",
  "post_date": "2022-12-08T02:26:12.600000",
  "votes": 19,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>UPDATE</strong>: I seem to have a tendency for \"rediscovering\" solutions to problems that my colleagues already have implemented a much better solution to 😄 For running NN search on the GPU please see: [Matrix Factorization with GPU: 6.5x faster!] by <a href=\"https://www.kaggle.com/cpmpm\" target=\"_blank\">@cpmpm</a>.</p>\n<p>Using <code>cuml</code> is a very elegant solution to this problem 🙂</p>\n<p>Plus you get many algorithms that do NN search on the GPU for free!</p>\n<p>If you would like to read a justification of why that is important, please read the post below.</p>\n<p>Thank you for stopping by 🙂</p>\n<hr>\n<p>Hey,</p>\n<p>ANN (approximate nearest neighbor search)  helps us deal with a situation where there are more items to retrieve than our compute can handle.</p>\n<p>I initially jumped to using ANN because that is generally what people do with millions of items to recommend, but I didn't realize how far you can get when running NN (nearest neighbor search) on the GPU!</p>\n<p>Here are the results (the first are runtime on the CPU and the scores with ANN, the second are on the GPU):</p>\n<p><strong>CPU ANN:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fc143f9c3a84bb7767f957142dd3a5cfd%2Fcpu_with_annoy.png?generation=1670472787700069&amp;alt=media\" alt=\"\"></p>\n<p><strong>GPU NN:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F24bbd1d97f66938dcc577976877b0051%2FGPU_with_faiss.png?generation=1670472817907642&amp;alt=media\" alt=\"\"></p>\n<p>In this relatively simple example, ANN gets worse results and takes longer to run than NN on the GPU!</p>\n<p>But for the benchmarking to be complete, I attempted to run NN both on the GPU and the CPU. The CPU is quite powerful (24-core AMD Ryzen 9 3900X)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F7fa1b89c3a12d39b8a6d68af91c5732d%2FScreenshot%202022-12-08%20135420.png?generation=1670471752759475&amp;alt=media\" alt=\"\"></p>\n<p>NN search on the GPU is 72x faster!</p>\n<p>I was very excited when I stumbled into this 🙂 This not only helps you get a better result in much shorter time, but also (which I really appreciate) takes away the overhead of having to find the number of trees for your ANN index to get a good trade-off between run time and performance, etc (in fact, once you move to more complex and powerful ANN implementations than <code>annoy</code> there are even more hyperparameters to tune!)</p>\n<p>Of course, beyond a certain point even the sheer speed of the GPU will not save you and you will have to switch to ANN. But it was quite amazing to me that with 1.7 million aids and a 32-dimensional vector this executed so quickly 🙂That GPUs push the boundaries of what is possible in this domain by that much (NN search is a key component of many RecSys solutions).</p>\n<p>Wanted to share this with you as this can definitely streamline the solutions and lead to an improved result. Happy Kaggling! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2058538,
      "postDate": "2022-12-08T02:26:12.600Z",
      "content": "<p><strong>UPDATE</strong>: I seem to have a tendency for \"rediscovering\" solutions to problems that my colleagues already have implemented a much better solution to 😄 For running NN search on the GPU please see: [Matrix Factorization with GPU: 6.5x faster!] by <a href=\"https://www.kaggle.com/cpmpm\" target=\"_blank\">@cpmpm</a>.</p>\n<p>Using <code>cuml</code> is a very elegant solution to this problem 🙂</p>\n<p>Plus you get many algorithms that do NN search on the GPU for free!</p>\n<p>If you would like to read a justification of why that is important, please read the post below.</p>\n<p>Thank you for stopping by 🙂</p>\n<hr>\n<p>Hey,</p>\n<p>ANN (approximate nearest neighbor search)  helps us deal with a situation where there are more items to retrieve than our compute can handle.</p>\n<p>I initially jumped to using ANN because that is generally what people do with millions of items to recommend, but I didn't realize how far you can get when running NN (nearest neighbor search) on the GPU!</p>\n<p>Here are the results (the first are runtime on the CPU and the scores with ANN, the second are on the GPU):</p>\n<p><strong>CPU ANN:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fc143f9c3a84bb7767f957142dd3a5cfd%2Fcpu_with_annoy.png?generation=1670472787700069&amp;alt=media\" alt=\"\"></p>\n<p><strong>GPU NN:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F24bbd1d97f66938dcc577976877b0051%2FGPU_with_faiss.png?generation=1670472817907642&amp;alt=media\" alt=\"\"></p>\n<p>In this relatively simple example, ANN gets worse results and takes longer to run than NN on the GPU!</p>\n<p>But for the benchmarking to be complete, I attempted to run NN both on the GPU and the CPU. The CPU is quite powerful (24-core AMD Ryzen 9 3900X)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F7fa1b89c3a12d39b8a6d68af91c5732d%2FScreenshot%202022-12-08%20135420.png?generation=1670471752759475&amp;alt=media\" alt=\"\"></p>\n<p>NN search on the GPU is 72x faster!</p>\n<p>I was very excited when I stumbled into this 🙂 This not only helps you get a better result in much shorter time, but also (which I really appreciate) takes away the overhead of having to find the number of trees for your ANN index to get a good trade-off between run time and performance, etc (in fact, once you move to more complex and powerful ANN implementations than <code>annoy</code> there are even more hyperparameters to tune!)</p>\n<p>Of course, beyond a certain point even the sheer speed of the GPU will not save you and you will have to switch to ANN. But it was quite amazing to me that with 1.7 million aids and a 32-dimensional vector this executed so quickly 🙂That GPUs push the boundaries of what is possible in this domain by that much (NN search is a key component of many RecSys solutions).</p>\n<p>Wanted to share this with you as this can definitely streamline the solutions and lead to an improved result. Happy Kaggling! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "**UPDATE**: I seem to have a tendency for \"rediscovering\" solutions to problems that my colleagues already have implemented a much better solution to 😄 For running NN search on the GPU please see: [Matrix Factorization with GPU: 6.5x faster!] by @cpmpm.\n\nUsing `cuml` is a very elegant solution to this problem 🙂\n\nPlus you get many algorithms that do NN search on the GPU for free!\n\nIf you would like to read a justification of why that is important, please read the post below.\n\nThank you for stopping by 🙂\n\n------------------------------\n\nHey,\n\nANN (approximate nearest neighbor search)  helps us deal with a situation where there are more items to retrieve than our compute can handle.\n\nI initially jumped to using ANN because that is generally what people do with millions of items to recommend, but I didn't realize how far you can get when running NN (nearest neighbor search) on the GPU!\n\nHere are the results (the first are runtime on the CPU and the scores with ANN, the second are on the GPU):\n\n**CPU ANN:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fc143f9c3a84bb7767f957142dd3a5cfd%2Fcpu_with_annoy.png?generation=1670472787700069&alt=media)\n\n**GPU NN:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F24bbd1d97f66938dcc577976877b0051%2FGPU_with_faiss.png?generation=1670472817907642&alt=media)\n\nIn this relatively simple example, ANN gets worse results and takes longer to run than NN on the GPU!\n\nBut for the benchmarking to be complete, I attempted to run NN both on the GPU and the CPU. The CPU is quite powerful (24-core AMD Ryzen 9 3900X)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F7fa1b89c3a12d39b8a6d68af91c5732d%2FScreenshot%202022-12-08%20135420.png?generation=1670471752759475&alt=media)\n\nNN search on the GPU is 72x faster!\n\nI was very excited when I stumbled into this 🙂 This not only helps you get a better result in much shorter time, but also (which I really appreciate) takes away the overhead of having to find the number of trees for your ANN index to get a good trade-off between run time and performance, etc (in fact, once you move to more complex and powerful ANN implementations than `annoy` there are even more hyperparameters to tune!)\n\nOf course, beyond a certain point even the sheer speed of the GPU will not save you and you will have to switch to ANN. But it was quite amazing to me that with 1.7 million aids and a 32-dimensional vector this executed so quickly 🙂That GPUs push the boundaries of what is possible in this domain by that much (NN search is a key component of many RecSys solutions).\n\nWanted to share this with you as this can definitely streamline the solutions and lead to an improved result. Happy Kaggling! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": 19
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2058538": "**UPDATE**: I seem to have a tendency for \"rediscovering\" solutions to problems that my colleagues already have implemented a much better solution to 😄 For running NN search on the GPU please see: [Matrix Factorization with GPU: 6.5x faster!] by @cpmpm.\n\nUsing `cuml` is a very elegant solution to this problem 🙂\n\nPlus you get many algorithms that do NN search on the GPU for free!\n\nIf you would like to read a justification of why that is important, please read the post below.\n\nThank you for stopping by 🙂\n\n------------------------------\n\nHey,\n\nANN (approximate nearest neighbor search)  helps us deal with a situation where there are more items to retrieve than our compute can handle.\n\nI initially jumped to using ANN because that is generally what people do with millions of items to recommend, but I didn't realize how far you can get when running NN (nearest neighbor search) on the GPU!\n\nHere are the results (the first are runtime on the CPU and the scores with ANN, the second are on the GPU):\n\n**CPU ANN:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fc143f9c3a84bb7767f957142dd3a5cfd%2Fcpu_with_annoy.png?generation=1670472787700069&alt=media)\n\n**GPU NN:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F24bbd1d97f66938dcc577976877b0051%2FGPU_with_faiss.png?generation=1670472817907642&alt=media)\n\nIn this relatively simple example, ANN gets worse results and takes longer to run than NN on the GPU!\n\nBut for the benchmarking to be complete, I attempted to run NN both on the GPU and the CPU. The CPU is quite powerful (24-core AMD Ryzen 9 3900X)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F7fa1b89c3a12d39b8a6d68af91c5732d%2FScreenshot%202022-12-08%20135420.png?generation=1670471752759475&alt=media)\n\nNN search on the GPU is 72x faster!\n\nI was very excited when I stumbled into this 🙂 This not only helps you get a better result in much shorter time, but also (which I really appreciate) takes away the overhead of having to find the number of trees for your ANN index to get a good trade-off between run time and performance, etc (in fact, once you move to more complex and powerful ANN implementations than `annoy` there are even more hyperparameters to tune!)\n\nOf course, beyond a certain point even the sheer speed of the GPU will not save you and you will have to switch to ANN. But it was quite amazing to me that with 1.7 million aids and a 32-dimensional vector this executed so quickly 🙂That GPUs push the boundaries of what is possible in this domain by that much (NN search is a key component of many RecSys solutions).\n\nWanted to share this with you as this can definitely streamline the solutions and lead to an improved result. Happy Kaggling! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)"
  }
}