{
  "id": 365589,
  "title": "Co-Visitation Matrices and Matrix Factorization",
  "url": "/competitions/otto-recommender-system/discussion/365589",
  "author_name": "Ravi Shah",
  "post_date": "2022-11-12T00:10:29.001000",
  "votes": 25,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Co-Visitation Matrices</h1>\n<p>As you might have seen in several notebooks, creating features based on similarly viewed items (co-visited items) is an effective technique for recommendation systems.</p>\n<p><strong>How it's being computed:</strong></p>\n<ol>\n<li>For a given id_x, find all other products that were also clicked on within a day of id_x</li>\n<li>The actual values inside the matrix are usually the number of times id_x and id_y were visited together, but the value may also be weights (see later)</li>\n<li>For each id_x, take the top 20 or 30 id_y</li>\n<li>The matrix is often stored in a dictionary format</li>\n</ol>\n<p><strong>How it's being stored:</strong></p>\n<ul>\n<li>This type of matrix is often stored in a dictionary style data structure for each use. </li>\n<li>The item you are trying to find the related items for (id_x) are the keys in the dictionary.</li>\n<li>The values of the dictionary are lists of item ids (id_y). </li>\n<li>Sometimes, these are the lists also contain a data structure such as a Counter to hold weights (wgt).</li>\n<li>Sometimes the lists are sorted rather than storing weights</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fef0a3d96c4cdc74c1cd6cf5271c1284c%2Fco-visit.PNG?generation=1668210251805605&amp;alt=media\" alt=\"\"></p>\n<p><strong>Weighted Co-Visitation Matrix:</strong></p>\n<p>The weights can be decided in a variety of ways:</p>\n<ul>\n<li>Simply the number of times the items were looked at together</li>\n<li>The difference in time the items were looked at together (shorter time gap, larger weight)</li>\n<li>Whether the items were clicks, carts, or orders</li>\n</ul>\n<p><strong>Example Notebooks</strong></p>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\" target=\"_blank\">Candidate ReRank Model by Chris</a><br>\n<a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix by Radek</a></p>\n<h1>Matrix Factorization</h1>\n<ul>\n<li>Matrix Factorization is a very popular technique for recommendation systems. </li>\n<li>Matrix factorization is a collaborative filtering technique in which the items (ids) are the rows and the users (sessions) are the columns. </li>\n<li>The actual values in the matrix are typically the level of preference a given user has for that item (in our case: clicks, carts, or orders)</li>\n<li>After creating the matrix, use some sort of similarity metric and algorithm such as nearest neighbors.</li>\n<li>Since we have a lot of data, this technique could be effective; however, the matrix can be hard to compute given the size of the data.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F71464c8729f135bc64a88866840521e6%2Fmat_fact.PNG?generation=1668211406072561&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2026327,
      "postDate": "2022-11-12T00:10:29Z",
      "content": "<h1>Co-Visitation Matrices</h1>\n<p>As you might have seen in several notebooks, creating features based on similarly viewed items (co-visited items) is an effective technique for recommendation systems.</p>\n<p><strong>How it's being computed:</strong></p>\n<ol>\n<li>For a given id_x, find all other products that were also clicked on within a day of id_x</li>\n<li>The actual values inside the matrix are usually the number of times id_x and id_y were visited together, but the value may also be weights (see later)</li>\n<li>For each id_x, take the top 20 or 30 id_y</li>\n<li>The matrix is often stored in a dictionary format</li>\n</ol>\n<p><strong>How it's being stored:</strong></p>\n<ul>\n<li>This type of matrix is often stored in a dictionary style data structure for each use. </li>\n<li>The item you are trying to find the related items for (id_x) are the keys in the dictionary.</li>\n<li>The values of the dictionary are lists of item ids (id_y). </li>\n<li>Sometimes, these are the lists also contain a data structure such as a Counter to hold weights (wgt).</li>\n<li>Sometimes the lists are sorted rather than storing weights</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fef0a3d96c4cdc74c1cd6cf5271c1284c%2Fco-visit.PNG?generation=1668210251805605&amp;alt=media\" alt=\"\"></p>\n<p><strong>Weighted Co-Visitation Matrix:</strong></p>\n<p>The weights can be decided in a variety of ways:</p>\n<ul>\n<li>Simply the number of times the items were looked at together</li>\n<li>The difference in time the items were looked at together (shorter time gap, larger weight)</li>\n<li>Whether the items were clicks, carts, or orders</li>\n</ul>\n<p><strong>Example Notebooks</strong></p>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573\" target=\"_blank\">Candidate ReRank Model by Chris</a><br>\n<a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix by Radek</a></p>\n<h1>Matrix Factorization</h1>\n<ul>\n<li>Matrix Factorization is a very popular technique for recommendation systems. </li>\n<li>Matrix factorization is a collaborative filtering technique in which the items (ids) are the rows and the users (sessions) are the columns. </li>\n<li>The actual values in the matrix are typically the level of preference a given user has for that item (in our case: clicks, carts, or orders)</li>\n<li>After creating the matrix, use some sort of similarity metric and algorithm such as nearest neighbors.</li>\n<li>Since we have a lot of data, this technique could be effective; however, the matrix can be hard to compute given the size of the data.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F71464c8729f135bc64a88866840521e6%2Fmat_fact.PNG?generation=1668211406072561&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "# Co-Visitation Matrices\n\nAs you might have seen in several notebooks, creating features based on similarly viewed items (co-visited items) is an effective technique for recommendation systems.\n\n**How it's being computed:**\n1. For a given id_x, find all other products that were also clicked on within a day of id_x\n2. The actual values inside the matrix are usually the number of times id_x and id_y were visited together, but the value may also be weights (see later)\n3. For each id_x, take the top 20 or 30 id_y\n4. The matrix is often stored in a dictionary format\n\n**How it's being stored:**\n\n- This type of matrix is often stored in a dictionary style data structure for each use. \n- The item you are trying to find the related items for (id_x) are the keys in the dictionary.\n- The values of the dictionary are lists of item ids (id_y). \n- Sometimes, these are the lists also contain a data structure such as a Counter to hold weights (wgt).\n- Sometimes the lists are sorted rather than storing weights\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fef0a3d96c4cdc74c1cd6cf5271c1284c%2Fco-visit.PNG?generation=1668210251805605&alt=media)\n\n**Weighted Co-Visitation Matrix:**\n\nThe weights can be decided in a variety of ways:\n- Simply the number of times the items were looked at together\n- The difference in time the items were looked at together (shorter time gap, larger weight)\n- Whether the items were clicks, carts, or orders\n\n**Example Notebooks**\n\n[Candidate ReRank Model by Chris](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573)\n[co-visitation matrix by Radek](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n\n# Matrix Factorization\n\n- Matrix Factorization is a very popular technique for recommendation systems. \n- Matrix factorization is a collaborative filtering technique in which the items (ids) are the rows and the users (sessions) are the columns. \n- The actual values in the matrix are typically the level of preference a given user has for that item (in our case: clicks, carts, or orders)\n- After creating the matrix, use some sort of similarity metric and algorithm such as nearest neighbors.\n- Since we have a lot of data, this technique could be effective; however, the matrix can be hard to compute given the size of the data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F71464c8729f135bc64a88866840521e6%2Fmat_fact.PNG?generation=1668211406072561&alt=media)",
      "votes": 25
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2026327": "# Co-Visitation Matrices\n\nAs you might have seen in several notebooks, creating features based on similarly viewed items (co-visited items) is an effective technique for recommendation systems.\n\n**How it's being computed:**\n1. For a given id_x, find all other products that were also clicked on within a day of id_x\n2. The actual values inside the matrix are usually the number of times id_x and id_y were visited together, but the value may also be weights (see later)\n3. For each id_x, take the top 20 or 30 id_y\n4. The matrix is often stored in a dictionary format\n\n**How it's being stored:**\n\n- This type of matrix is often stored in a dictionary style data structure for each use. \n- The item you are trying to find the related items for (id_x) are the keys in the dictionary.\n- The values of the dictionary are lists of item ids (id_y). \n- Sometimes, these are the lists also contain a data structure such as a Counter to hold weights (wgt).\n- Sometimes the lists are sorted rather than storing weights\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fef0a3d96c4cdc74c1cd6cf5271c1284c%2Fco-visit.PNG?generation=1668210251805605&alt=media)\n\n**Weighted Co-Visitation Matrix:**\n\nThe weights can be decided in a variety of ways:\n- Simply the number of times the items were looked at together\n- The difference in time the items were looked at together (shorter time gap, larger weight)\n- Whether the items were clicks, carts, or orders\n\n**Example Notebooks**\n\n[Candidate ReRank Model by Chris](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-573)\n[co-visitation matrix by Radek](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n\n# Matrix Factorization\n\n- Matrix Factorization is a very popular technique for recommendation systems. \n- Matrix factorization is a collaborative filtering technique in which the items (ids) are the rows and the users (sessions) are the columns. \n- The actual values in the matrix are typically the level of preference a given user has for that item (in our case: clicks, carts, or orders)\n- After creating the matrix, use some sort of similarity metric and algorithm such as nearest neighbors.\n- Since we have a lot of data, this technique could be effective; however, the matrix can be hard to compute given the size of the data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F71464c8729f135bc64a88866840521e6%2Fmat_fact.PNG?generation=1668211406072561&alt=media)"
  }
}