{
  "id": 368385,
  "title": "💡How to improve the results of your Approximate Nearest Neighbor search! (annoy)",
  "url": "/competitions/otto-recommender-system/discussion/368385",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-25T01:41:49.779000",
  "votes": 29,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>So I have been sharing a bunch of code that uses <code>annoy</code> to perform an approximate nearest neighbor search in the embedding space.</p>\n<p>But there is one very vital piece of information (and I have been quite surprised as to the effect of this!). Please see below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F37d7b67144fb9f7f7989cfeadeb7695f%2Fn_trees_vs_recall.png?generation=1669340255319961&amp;alt=media\" alt=\"\"></p>\n<p>The above tracks recall with a bunch of embeddings AND index created with various tree counts! As the tree count increases, so does the performance.</p>\n<p>But the extent to which this happens is insane! I didn't expect you could squeeze that much more performance out of the algo with more trees in the index.</p>\n<p>Unfortunately, with more trees in the index, the operation to create it (and in particular, the lookups) becomes much more costly, takes much longer, so realistically probably using between 30 and 50 trees is the sweet spot.</p>\n<p>Thanks for reading! 🙌 Hope this helps! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2042763,
      "postDate": "2022-11-25T01:41:49.780Z",
      "content": "<p>Hey!</p>\n<p>So I have been sharing a bunch of code that uses <code>annoy</code> to perform an approximate nearest neighbor search in the embedding space.</p>\n<p>But there is one very vital piece of information (and I have been quite surprised as to the effect of this!). Please see below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F37d7b67144fb9f7f7989cfeadeb7695f%2Fn_trees_vs_recall.png?generation=1669340255319961&amp;alt=media\" alt=\"\"></p>\n<p>The above tracks recall with a bunch of embeddings AND index created with various tree counts! As the tree count increases, so does the performance.</p>\n<p>But the extent to which this happens is insane! I didn't expect you could squeeze that much more performance out of the algo with more trees in the index.</p>\n<p>Unfortunately, with more trees in the index, the operation to create it (and in particular, the lookups) becomes much more costly, takes much longer, so realistically probably using between 30 and 50 trees is the sweet spot.</p>\n<p>Thanks for reading! 🙌 Hope this helps! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nSo I have been sharing a bunch of code that uses `annoy` to perform an approximate nearest neighbor search in the embedding space.\n\nBut there is one very vital piece of information (and I have been quite surprised as to the effect of this!). Please see below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F37d7b67144fb9f7f7989cfeadeb7695f%2Fn_trees_vs_recall.png?generation=1669340255319961&alt=media)\n\nThe above tracks recall with a bunch of embeddings AND index created with various tree counts! As the tree count increases, so does the performance.\n\nBut the extent to which this happens is insane! I didn't expect you could squeeze that much more performance out of the algo with more trees in the index.\n\nUnfortunately, with more trees in the index, the operation to create it (and in particular, the lookups) becomes much more costly, takes much longer, so realistically probably using between 30 and 50 trees is the sweet spot.\n\nThanks for reading! 🙌 Hope this helps! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": 29
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2042763": "Hey!\n\nSo I have been sharing a bunch of code that uses `annoy` to perform an approximate nearest neighbor search in the embedding space.\n\nBut there is one very vital piece of information (and I have been quite surprised as to the effect of this!). Please see below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F37d7b67144fb9f7f7989cfeadeb7695f%2Fn_trees_vs_recall.png?generation=1669340255319961&alt=media)\n\nThe above tracks recall with a bunch of embeddings AND index created with various tree counts! As the tree count increases, so does the performance.\n\nBut the extent to which this happens is insane! I didn't expect you could squeeze that much more performance out of the algo with more trees in the index.\n\nUnfortunately, with more trees in the index, the operation to create it (and in particular, the lookups) becomes much more costly, takes much longer, so realistically probably using between 30 and 50 trees is the sweet spot.\n\nThanks for reading! 🙌 Hope this helps! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)"
  }
}