{
  "id": 316903,
  "title": "Recommender System using turicreate library",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/316903",
  "author_name": "Andrada",
  "post_date": "2022-04-04T14:58:47.941000",
  "votes": 21,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As this was my first time handling Recommender Systems, I naturally went to Google to see a few approaches that I might learn and apply on this competition. However most of them were used for rating-based data. </p>\n<p>These examples work something like: if you have customer A and B that are very similar in their preferences towards movies, you can use what customer A has watched new to also recommend to customer B.</p>\n<p><img src=\"https://i.imgur.com/pJWTLzx.png\"></p>\n<p>However in this competition we are working with something different - we are trying to recommend what a customer will buy based on previous purchases, which at quantity level are different than ratings (which are within a strict interval of [1, 5] or [1, 10]).</p>\n<p>After some digging I have found <a href=\"https://medium.datadriveninvestor.com/how-to-build-a-recommendation-system-for-purchase-data-step-by-step-d6d7a78800b6\" target=\"_blank\">this medium post</a> that helped me get started with this competition.</p>\n<h1>Turicreate library</h1>\n<p>Turicreate has very easy functions (<a href=\"https://www.kaggle.com/code/andradaolteanu/h-m-eda-rapids-and-market-basket-analysis\" target=\"_blank\">you can check out my notebook</a> or the article mentioned above for examples) that can be used to create a recommender system.</p>\n<blockquote>\n  <p><strong>Note</strong>: to install <code>turicreate</code> in your notebook use <code>!pip install turicreate --user</code> instead of the simple <code>!pip install turicreate</code> - as the simple version has dependency issues.</p>\n</blockquote>\n<h2>Popularity Recommender</h2>\n<p>The easiest fastest way is to create a popularity reccomendation - meaning what are the items that are sold the most across all customers (I have found this model to be very low in performance).</p>\n<p>Example code: <code>tc.popularity_recommender.create(train_data, user_id, item_id, target)</code></p>\n<h2>Similarity Recommender</h2>\n<p>This methodology looks a lot in my opinion like Market Basket Analysis - where if you see that customers A buys milk, eggs, butter and flour, then customer D with milk, eggs, butter in the cart might also need flour (because, you guessed it, they're making pancakes).</p>\n<p><img src=\"https://i.imgur.com/m8T6MhL.png\"></p>\n<p>There are multiple algorithms that calculate <em>how similar is a one product to another</em>. The idea is to compute for all milk, eggs, butter and flour how many times they were bought together and individually for each customer. Then, 4 vectors for all 4 items are computed in space.</p>\n<p>The way to compute how close together (similar) are these 2 vectors can be done in various ways, but within the article and my notebook 2 methodologies are used: <code>cosine</code> distance and <code>pearson</code> distance.</p>\n<p><code>cosine</code> : <code>tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='cosine')</code></p>\n<p><code>pearson</code> : <code>tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='pearson')</code></p>\n<p>After we have assessed that, for example, these vectors are VERY similat to eachother we can compute how likely is that Jane will buy the flour.</p>\n<h3>Cosine Similarity Formula</h3>\n<p><img src=\"https://i.imgur.com/giRHV6J.png\"></p>\n<h3>Pearson Similarity Formula</h3>\n<p><img src=\"https://i.imgur.com/TjVVKpV.png\"></p>",
  "messages": [
    {
      "id": 1745059,
      "postDate": "2022-04-04T14:58:47.940Z",
      "content": "<p>As this was my first time handling Recommender Systems, I naturally went to Google to see a few approaches that I might learn and apply on this competition. However most of them were used for rating-based data. </p>\n<p>These examples work something like: if you have customer A and B that are very similar in their preferences towards movies, you can use what customer A has watched new to also recommend to customer B.</p>\n<p><img src=\"https://i.imgur.com/pJWTLzx.png\"></p>\n<p>However in this competition we are working with something different - we are trying to recommend what a customer will buy based on previous purchases, which at quantity level are different than ratings (which are within a strict interval of [1, 5] or [1, 10]).</p>\n<p>After some digging I have found <a href=\"https://medium.datadriveninvestor.com/how-to-build-a-recommendation-system-for-purchase-data-step-by-step-d6d7a78800b6\" target=\"_blank\">this medium post</a> that helped me get started with this competition.</p>\n<h1>Turicreate library</h1>\n<p>Turicreate has very easy functions (<a href=\"https://www.kaggle.com/code/andradaolteanu/h-m-eda-rapids-and-market-basket-analysis\" target=\"_blank\">you can check out my notebook</a> or the article mentioned above for examples) that can be used to create a recommender system.</p>\n<blockquote>\n  <p><strong>Note</strong>: to install <code>turicreate</code> in your notebook use <code>!pip install turicreate --user</code> instead of the simple <code>!pip install turicreate</code> - as the simple version has dependency issues.</p>\n</blockquote>\n<h2>Popularity Recommender</h2>\n<p>The easiest fastest way is to create a popularity reccomendation - meaning what are the items that are sold the most across all customers (I have found this model to be very low in performance).</p>\n<p>Example code: <code>tc.popularity_recommender.create(train_data, user_id, item_id, target)</code></p>\n<h2>Similarity Recommender</h2>\n<p>This methodology looks a lot in my opinion like Market Basket Analysis - where if you see that customers A buys milk, eggs, butter and flour, then customer D with milk, eggs, butter in the cart might also need flour (because, you guessed it, they're making pancakes).</p>\n<p><img src=\"https://i.imgur.com/m8T6MhL.png\"></p>\n<p>There are multiple algorithms that calculate <em>how similar is a one product to another</em>. The idea is to compute for all milk, eggs, butter and flour how many times they were bought together and individually for each customer. Then, 4 vectors for all 4 items are computed in space.</p>\n<p>The way to compute how close together (similar) are these 2 vectors can be done in various ways, but within the article and my notebook 2 methodologies are used: <code>cosine</code> distance and <code>pearson</code> distance.</p>\n<p><code>cosine</code> : <code>tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='cosine')</code></p>\n<p><code>pearson</code> : <code>tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='pearson')</code></p>\n<p>After we have assessed that, for example, these vectors are VERY similat to eachother we can compute how likely is that Jane will buy the flour.</p>\n<h3>Cosine Similarity Formula</h3>\n<p><img src=\"https://i.imgur.com/giRHV6J.png\"></p>\n<h3>Pearson Similarity Formula</h3>\n<p><img src=\"https://i.imgur.com/TjVVKpV.png\"></p>",
      "rawMarkdown": "As this was my first time handling Recommender Systems, I naturally went to Google to see a few approaches that I might learn and apply on this competition. However most of them were used for rating-based data. \n\nThese examples work something like: if you have customer A and B that are very similar in their preferences towards movies, you can use what customer A has watched new to also recommend to customer B.\n\n<img src=\"https://i.imgur.com/pJWTLzx.png\">\n\nHowever in this competition we are working with something different - we are trying to recommend what a customer will buy based on previous purchases, which at quantity level are different than ratings (which are within a strict interval of [1, 5] or [1, 10]).\n\nAfter some digging I have found [this medium post](https://medium.datadriveninvestor.com/how-to-build-a-recommendation-system-for-purchase-data-step-by-step-d6d7a78800b6) that helped me get started with this competition.\n\n# Turicreate library\n\nTuricreate has very easy functions ([you can check out my notebook](https://www.kaggle.com/code/andradaolteanu/h-m-eda-rapids-and-market-basket-analysis) or the article mentioned above for examples) that can be used to create a recommender system.\n\n> **Note**: to install `turicreate` in your notebook use `!pip install turicreate --user` instead of the simple `!pip install turicreate` - as the simple version has dependency issues.\n\n## Popularity Recommender\n\nThe easiest fastest way is to create a popularity reccomendation - meaning what are the items that are sold the most across all customers (I have found this model to be very low in performance).\n\nExample code: `tc.popularity_recommender.create(train_data, user_id, item_id, target)`\n\n## Similarity Recommender\n\nThis methodology looks a lot in my opinion like Market Basket Analysis - where if you see that customers A buys milk, eggs, butter and flour, then customer D with milk, eggs, butter in the cart might also need flour (because, you guessed it, they're making pancakes).\n\n<img src=\"https://i.imgur.com/m8T6MhL.png\">\n\nThere are multiple algorithms that calculate *how similar is a one product to another*. The idea is to compute for all milk, eggs, butter and flour how many times they were bought together and individually for each customer. Then, 4 vectors for all 4 items are computed in space.\n\nThe way to compute how close together (similar) are these 2 vectors can be done in various ways, but within the article and my notebook 2 methodologies are used: `cosine` distance and `pearson` distance.\n\n`cosine` : `tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='cosine')`\n\n`pearson` : `tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='pearson')`\n\nAfter we have assessed that, for example, these vectors are VERY similat to eachother we can compute how likely is that Jane will buy the flour.\n\n### Cosine Similarity Formula\n\n<center><img src=\"https://i.imgur.com/giRHV6J.png\" width=400></center>\n\n### Pearson Similarity Formula\n\n<center><img src=\"https://i.imgur.com/TjVVKpV.png\" width=250></center>",
      "votes": 21
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1745059": "As this was my first time handling Recommender Systems, I naturally went to Google to see a few approaches that I might learn and apply on this competition. However most of them were used for rating-based data. \n\nThese examples work something like: if you have customer A and B that are very similar in their preferences towards movies, you can use what customer A has watched new to also recommend to customer B.\n\n<img src=\"https://i.imgur.com/pJWTLzx.png\">\n\nHowever in this competition we are working with something different - we are trying to recommend what a customer will buy based on previous purchases, which at quantity level are different than ratings (which are within a strict interval of [1, 5] or [1, 10]).\n\nAfter some digging I have found [this medium post](https://medium.datadriveninvestor.com/how-to-build-a-recommendation-system-for-purchase-data-step-by-step-d6d7a78800b6) that helped me get started with this competition.\n\n# Turicreate library\n\nTuricreate has very easy functions ([you can check out my notebook](https://www.kaggle.com/code/andradaolteanu/h-m-eda-rapids-and-market-basket-analysis) or the article mentioned above for examples) that can be used to create a recommender system.\n\n> **Note**: to install `turicreate` in your notebook use `!pip install turicreate --user` instead of the simple `!pip install turicreate` - as the simple version has dependency issues.\n\n## Popularity Recommender\n\nThe easiest fastest way is to create a popularity reccomendation - meaning what are the items that are sold the most across all customers (I have found this model to be very low in performance).\n\nExample code: `tc.popularity_recommender.create(train_data, user_id, item_id, target)`\n\n## Similarity Recommender\n\nThis methodology looks a lot in my opinion like Market Basket Analysis - where if you see that customers A buys milk, eggs, butter and flour, then customer D with milk, eggs, butter in the cart might also need flour (because, you guessed it, they're making pancakes).\n\n<img src=\"https://i.imgur.com/m8T6MhL.png\">\n\nThere are multiple algorithms that calculate *how similar is a one product to another*. The idea is to compute for all milk, eggs, butter and flour how many times they were bought together and individually for each customer. Then, 4 vectors for all 4 items are computed in space.\n\nThe way to compute how close together (similar) are these 2 vectors can be done in various ways, but within the article and my notebook 2 methodologies are used: `cosine` distance and `pearson` distance.\n\n`cosine` : `tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='cosine')`\n\n`pearson` : `tc.item_similarity_recommender.create(train_data, user_id, item_id, target, similarity_type='pearson')`\n\nAfter we have assessed that, for example, these vectors are VERY similat to eachother we can compute how likely is that Jane will buy the flour.\n\n### Cosine Similarity Formula\n\n<center><img src=\"https://i.imgur.com/giRHV6J.png\" width=400></center>\n\n### Pearson Similarity Formula\n\n<center><img src=\"https://i.imgur.com/TjVVKpV.png\" width=250></center>"
  }
}