{
  "id": 365358,
  "title": "💡 What is the co-visitation matrix, really?",
  "url": "/competitions/otto-recommender-system/discussion/365358",
  "author_name": "",
  "post_date": "2022-11-10T23:20:21.262308200Z",
  "votes": 66,
  "comment_count": 18,
  "views": 0,
  "content": "<p>It is very interesting to think of modern techniques in the context of their roots.</p>\n<p>For instance, when thinking about RNNs we should consider unigram, bigram, and trigram models.</p>\n<p>What were they?</p>\n<p>They estimated the probability of a word given the words that came before.</p>\n<p>We see \"Radek is a <strong><em><em>_</em></em></strong>\".</p>\n<p>In a trigram model, we would consider the chain \"Radek\", \"is\", and \"a\" and could find the most likely word to come next.</p>\n<p>The easiest way would be to count the occurrences of \"Radek is a\" and the words that came after and pick the one most common.</p>\n<p>So why RNNs are such a great improvement?</p>\n<p>Well, there are not that many \"Radek is a <strong><em><em>_</em></em></strong>\" in most text corpora!</p>\n<p>With RNNs (or word2vec) we could operate on embeddings and look at many more words in sequence.</p>\n<p>Instead of looking at only \"Radek is a <strong><em><strong></strong></em></strong>\" we can use \"Radek was an <strong><em><em>_</em></em></strong>\", \"Tommy is a <strong><em><em>__</em></em></strong>\" as examples to learn from, etc.</p>\n<p>Embeddings take values of a variable of high cardinality and project them to a representation that ideally captures similarities.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F70e2a7f1b9621cd0afb2b815a7b3e01b%2Fyou_shall_know_a_word.png?generation=1668122115555809&amp;alt=media\" alt=\"\"></p>\n<p>So how does this relate to the co-visitation matrix?</p>\n<p>A co-visitation matrix counts the co-occurrence of two actions in close proximity.</p>\n<p>If a user bought A and shortly after bought B, we store these values together.</p>\n<p>We calculate counts and use them to estimate the probability of future actions based on recent history.</p>\n<p>It is quite important to understand what is happening in the co-visitation matrix approach…</p>\n<p>Since it suffers from the same issues as our trigram example!</p>\n<p>Plus what does the co-visitation matrix resemble?</p>\n<p>You are right, it is akin to doing Matrix Factorization by counting!</p>\n<p>It is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2024992",
      "postDate": "11/10/2022 23:20:21",
      "content": "<p>It is very interesting to think of modern techniques in the context of their roots.</p>\n<p>For instance, when thinking about RNNs we should consider unigram, bigram, and trigram models.</p>\n<p>What were they?</p>\n<p>They estimated the probability of a word given the words that came before.</p>\n<p>We see \"Radek is a <strong><em><em>_</em></em></strong>\".</p>\n<p>In a trigram model, we would consider the chain \"Radek\", \"is\", and \"a\" and could find the most likely word to come next.</p>\n<p>The easiest way would be to count the occurrences of \"Radek is a\" and the words that came after and pick the one most common.</p>\n<p>So why RNNs are such a great improvement?</p>\n<p>Well, there are not that many \"Radek is a <strong><em><em>_</em></em></strong>\" in most text corpora!</p>\n<p>With RNNs (or word2vec) we could operate on embeddings and look at many more words in sequence.</p>\n<p>Instead of looking at only \"Radek is a <strong><em><strong></strong></em></strong>\" we can use \"Radek was an <strong><em><em>_</em></em></strong>\", \"Tommy is a <strong><em><em>__</em></em></strong>\" as examples to learn from, etc.</p>\n<p>Embeddings take values of a variable of high cardinality and project them to a representation that ideally captures similarities.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F70e2a7f1b9621cd0afb2b815a7b3e01b%2Fyou_shall_know_a_word.png?generation=1668122115555809&amp;alt=media\" alt=\"\"></p>\n<p>So how does this relate to the co-visitation matrix?</p>\n<p>A co-visitation matrix counts the co-occurrence of two actions in close proximity.</p>\n<p>If a user bought A and shortly after bought B, we store these values together.</p>\n<p>We calculate counts and use them to estimate the probability of future actions based on recent history.</p>\n<p>It is quite important to understand what is happening in the co-visitation matrix approach…</p>\n<p>Since it suffers from the same issues as our trigram example!</p>\n<p>Plus what does the co-visitation matrix resemble?</p>\n<p>You are right, it is akin to doing Matrix Factorization by counting!</p>\n<p>It is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "It is very interesting to think of modern techniques in the context of their roots.\n\nFor instance, when thinking about RNNs we should consider unigram, bigram, and trigram models.\n\nWhat were they?\n\nThey estimated the probability of a word given the words that came before.\n\nWe see \"Radek is a _________\".\n\nIn a trigram model, we would consider the chain \"Radek\", \"is\", and \"a\" and could find the most likely word to come next.\n\nThe easiest way would be to count the occurrences of \"Radek is a\" and the words that came after and pick the one most common.\n\nSo why RNNs are such a great improvement?\n\nWell, there are not that many \"Radek is a _________\" in most text corpora!\n\nWith RNNs (or word2vec) we could operate on embeddings and look at many more words in sequence.\n\nInstead of looking at only \"Radek is a ________\" we can use \"Radek was an ___________\", \"Tommy is a __________\" as examples to learn from, etc.\n\nEmbeddings take values of a variable of high cardinality and project them to a representation that ideally captures similarities.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F70e2a7f1b9621cd0afb2b815a7b3e01b%2Fyou_shall_know_a_word.png?generation=1668122115555809&alt=media)\n\nSo how does this relate to the co-visitation matrix?\n\nA co-visitation matrix counts the co-occurrence of two actions in close proximity.\n\nIf a user bought A and shortly after bought B, we store these values together.\n\nWe calculate counts and use them to estimate the probability of future actions based on recent history.\n\nIt is quite important to understand what is happening in the co-visitation matrix approach...\n\nSince it suffers from the same issues as our trigram example!\n\nPlus what does the co-visitation matrix resemble?\n\nYou are right, it is akin to doing Matrix Factorization by counting!\n\nIt is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2024995",
      "postDate": "11/10/2022 23:23:57",
      "content": "<p>Nice explanation….Really helpful…Thanks👍</p>",
      "rawMarkdown": "Nice explanation....Really helpful...Thanks👍",
      "votes": null
    },
    {
      "id": "2025025",
      "postDate": "11/11/2022 00:03:24",
      "content": "<p>How funny that I was playing with the Word2Vec and its variants and then saw this post :D </p>",
      "rawMarkdown": "How funny that I was playing with the Word2Vec and its variants and then saw this post :D",
      "votes": null
    },
    {
      "id": "2025036",
      "postDate": "11/11/2022 00:15:38",
      "content": "<p>Yeah, it would be really interesting to see what word2vec can do in this competition 🙂 Might be it could be quite nice for candidate generation or some sort of similarity scoring for reranking 🤔</p>\n<p>Anyhow, cool that we were both circulating around the same topic! 🙌</p>",
      "rawMarkdown": "Yeah, it would be really interesting to see what word2vec can do in this competition 🙂 Might be it could be quite nice for candidate generation or some sort of similarity scoring for reranking 🤔\n\nAnyhow, cool that we were both circulating around the same topic! 🙌",
      "votes": null
    },
    {
      "id": "2025045",
      "postDate": "11/11/2022 00:27:43",
      "content": "<p>Co-visitation matrices are the natural statistical approach. It is a way to compute conditional probabilities. Given a user has interacted with item A, it computes the most likely future items that this user will interact with.</p>\n<p>And we can make many more specific variants such as given a user has <strong>ordered</strong> item A, what is the most likely item that this user will <strong>click, cart, order</strong> etc. Or given a user has visited this item <strong>in the morning</strong>, what is the most likely item that this user will <strong>order</strong>. Etc etc.</p>\n<p>Of course if we train a reranker model it will learn all these varieties of conditional probabilities on its own.</p>",
      "rawMarkdown": "Co-visitation matrices are the natural statistical approach. It is a way to compute conditional probabilities. Given a user has interacted with item A, it computes the most likely future items that this user will interact with.\n\nAnd we can make many more specific variants such as given a user has **ordered** item A, what is the most likely item that this user will **click, cart, order** etc. Or given a user has visited this item **in the morning**, what is the most likely item that this user will **order**. Etc etc.\n\nOf course if we train a reranker model it will learn all these varieties of conditional probabilities on its own.",
      "votes": null
    },
    {
      "id": "2025050",
      "postDate": "11/11/2022 00:32:37",
      "content": "<p>Training Word2Vec would be very helpful. That would produce embeddings such that similar items (with co-visitation) would have distance similar embeddings. It would help us understand what these anonymous item ids represent.</p>\n<p>Then we could do lots of EDA like how do the items that users purchase over time change? How do items purchased in the morning differ from items purchased in the evening. Etc etc. For each of these analysis, we could plot 2D pictures with clusters of dots (using their Word2Vec embedding 2D projection)</p>",
      "rawMarkdown": "Training Word2Vec would be very helpful. That would produce embeddings such that similar items (with co-visitation) would have distance similar embeddings. It would help us understand what these anonymous item ids represent.\n\nThen we could do lots of EDA like how do the items that users purchase over time change? How do items purchased in the morning differ from items purchased in the evening. Etc etc. For each of these analysis, we could plot 2D pictures with clusters of dots (using their Word2Vec embedding 2D projection)",
      "votes": null
    },
    {
      "id": "2025058",
      "postDate": "11/11/2022 00:41:37",
      "content": "<p>Well explained. Thanks for sharing👍</p>",
      "rawMarkdown": "Well explained. Thanks for sharing👍",
      "votes": null
    },
    {
      "id": "2025072",
      "postDate": "11/11/2022 00:49:31",
      "content": "<p>my pleasure, <a href=\"https://www.kaggle.com/oscarm524\" target=\"_blank\">@oscarm524</a>! very glad you found this useful! 🙏</p>",
      "rawMarkdown": "my pleasure, @oscarm524! very glad you found this useful! 🙏",
      "votes": null
    },
    {
      "id": "2025074",
      "postDate": "11/11/2022 00:51:14",
      "content": "<p>That is some amazing information <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, as always! 🙂 Thank you so very much for sharing all this with us 🙏</p>",
      "rawMarkdown": "That is some amazing information @cdeotte, as always! 🙂 Thank you so very much for sharing all this with us 🙏",
      "votes": null
    },
    {
      "id": "2025075",
      "postDate": "11/11/2022 00:51:40",
      "content": "<p>Yes, exactly! I was also thinking of it for the candidate generation. 🤜🤛</p>",
      "rawMarkdown": "Yes, exactly! I was also thinking of it for the candidate generation. 🤜🤛",
      "votes": null
    },
    {
      "id": "2025093",
      "postDate": "11/11/2022 01:08:02",
      "content": "<p>Word2Vec would also be great for EDA. It (with 2D projection via UMAP or TSNE) would help us visualize these anonymous items.</p>",
      "rawMarkdown": "Word2Vec would also be great for EDA. It (with 2D projection via UMAP or TSNE) would help us visualize these anonymous items.",
      "votes": null
    },
    {
      "id": "2025750",
      "postDate": "11/11/2022 13:07:07",
      "content": "<p>Great idea(s), Chris! Also as you have mentioned in your other comments, visualizations based on time would be insightful! I have the model, let's see if I can come up with something. </p>",
      "rawMarkdown": "Great idea(s), Chris! Also as you have mentioned in your other comments, visualizations based on time would be insightful! I have the model, let's see if I can come up with something.",
      "votes": null
    },
    {
      "id": "2060868",
      "postDate": "12/10/2022 14:20:39",
      "content": "<p>Thank you for both Radek and Chris; nowThank you for both Radek and Chris, now I am clear with the intuition of co-visitation matrices</p>",
      "rawMarkdown": "Thank you for both Radek and Chris; nowThank you for both Radek and Chris, now I am clear with the intuition of co-visitation matrices",
      "votes": null
    },
    {
      "id": "2061273",
      "postDate": "12/10/2022 22:27:19",
      "content": "<p>Wonderful! that is great to hear <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>! 🙂</p>",
      "rawMarkdown": "Wonderful! that is great to hear @leiwong! 🙂",
      "votes": null
    },
    {
      "id": "2105430",
      "postDate": "01/18/2023 13:52:08",
      "content": "<p>Hi Radek, I didn't appreciate this post of yours enough when I read it last month, now what you said about co-visitation matrix makes more sense to me now as I was writing down my reflection of co-visitation matrix these two days. </p>\n<p>Here is a question I have about comparing co-visitation matrix and word2vect.  Since you said co-visitation matrix suffers the same thing as trigram, then word2vec would in theory be a better model for generating candidates right? </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F27930a141e9e30ced711ae9a91e518ad%2FScreen%20Shot%202023-01-18%20at%2021.49.24.png?generation=1674049779499957&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi Radek, I didn't appreciate this post of yours enough when I read it last month, now what you said about co-visitation matrix makes more sense to me now as I was writing down my reflection of co-visitation matrix these two days. \n\nHere is a question I have about comparing co-visitation matrix and word2vect.  Since you said co-visitation matrix suffers the same thing as trigram, then word2vec would in theory be a better model for generating candidates right? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F27930a141e9e30ced711ae9a91e518ad%2FScreen%20Shot%202023-01-18%20at%2021.49.24.png?generation=1674049779499957&alt=media)",
      "votes": null
    },
    {
      "id": "2105932",
      "postDate": "01/18/2023 20:42:11",
      "content": "<p>In theory, matrix factorization (and word2vec) should do better! But it is important to check how they work in practice. Here, I only played a little bit around with candidate generation using MF and didn't get better results than using the co-visitation matrix. But there could be many things at play.</p>\n<p>In general, yes, matrix factorization and word2vec should be superior methods to the co-visitation matrix, that is my belief. But all of this is very situational dependant and depending on the dataset a heuristic such as the co-visitation matrix might work best 🙂 (for instances, the co-visitation matrix might better capture short-term trends, or can work really well if there is a long-tail of items purchased).</p>\n<p>It is fun and useful to think about these methods, what they do, how they work, etc but I wouldn't draw too definite conclusions, always best to test your hypotheses out on the dataset you are working with 🙂</p>",
      "rawMarkdown": "In theory, matrix factorization (and word2vec) should do better! But it is important to check how they work in practice. Here, I only played a little bit around with candidate generation using MF and didn't get better results than using the co-visitation matrix. But there could be many things at play.\n\nIn general, yes, matrix factorization and word2vec should be superior methods to the co-visitation matrix, that is my belief. But all of this is very situational dependant and depending on the dataset a heuristic such as the co-visitation matrix might work best 🙂 (for instances, the co-visitation matrix might better capture short-term trends, or can work really well if there is a long-tail of items purchased).\n\nIt is fun and useful to think about these methods, what they do, how they work, etc but I wouldn't draw too definite conclusions, always best to test your hypotheses out on the dataset you are working with 🙂",
      "votes": null
    },
    {
      "id": "2106172",
      "postDate": "01/19/2023 02:13:52",
      "content": "<p>Thank you so much Radek! What an illuminating and insightful answer! I am looking forward to test them out!</p>",
      "rawMarkdown": "Thank you so much Radek! What an illuminating and insightful answer! I am looking forward to test them out!",
      "votes": null
    },
    {
      "id": "2106250",
      "postDate": "01/19/2023 03:42:48",
      "content": "<p>My pleasure, <a href=\"https://www.kaggle.com/saha8631\" target=\"_blank\">@saha8631</a>! 🙂 Thank you for your kind comment! 🙌 </p>",
      "rawMarkdown": "My pleasure, @saha8631! 🙂 Thank you for your kind comment! 🙌",
      "votes": null
    },
    {
      "id": "2173215",
      "postDate": "03/08/2023 07:45:01",
      "content": "<p>Can the same principle of co-visitation matrix work when dealing with datasets of restaurants, for recommending a restaurant for a user?</p>",
      "rawMarkdown": "Can the same principle of co-visitation matrix work when dealing with datasets of restaurants, for recommending a restaurant for a user?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2024995,
      "author_name": "saha8631",
      "author_url": "",
      "post_date": "11/10/2022 23:23:57",
      "content": "<p>Nice explanation….Really helpful…Thanks👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2106250,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "01/19/2023 03:42:48",
          "content": "<p>My pleasure, <a href=\"https://www.kaggle.com/saha8631\" target=\"_blank\">@saha8631</a>! 🙂 Thank you for your kind comment! 🙌 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2025025,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "11/11/2022 00:03:24",
      "content": "<p>How funny that I was playing with the Word2Vec and its variants and then saw this post :D </p>",
      "votes": null,
      "replies": [
        {
          "id": 2025036,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/11/2022 00:15:38",
          "content": "<p>Yeah, it would be really interesting to see what word2vec can do in this competition 🙂 Might be it could be quite nice for candidate generation or some sort of similarity scoring for reranking 🤔</p>\n<p>Anyhow, cool that we were both circulating around the same topic! 🙌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2025075,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "11/11/2022 00:51:40",
          "content": "<p>Yes, exactly! I was also thinking of it for the candidate generation. 🤜🤛</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2025093,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/11/2022 01:08:02",
          "content": "<p>Word2Vec would also be great for EDA. It (with 2D projection via UMAP or TSNE) would help us visualize these anonymous items.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2025750,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "11/11/2022 13:07:07",
          "content": "<p>Great idea(s), Chris! Also as you have mentioned in your other comments, visualizations based on time would be insightful! I have the model, let's see if I can come up with something. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2025045,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/11/2022 00:27:43",
      "content": "<p>Co-visitation matrices are the natural statistical approach. It is a way to compute conditional probabilities. Given a user has interacted with item A, it computes the most likely future items that this user will interact with.</p>\n<p>And we can make many more specific variants such as given a user has <strong>ordered</strong> item A, what is the most likely item that this user will <strong>click, cart, order</strong> etc. Or given a user has visited this item <strong>in the morning</strong>, what is the most likely item that this user will <strong>order</strong>. Etc etc.</p>\n<p>Of course if we train a reranker model it will learn all these varieties of conditional probabilities on its own.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2025074,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/11/2022 00:51:14",
          "content": "<p>That is some amazing information <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, as always! 🙂 Thank you so very much for sharing all this with us 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2060868,
          "author_name": "leiwong",
          "author_url": "",
          "post_date": "12/10/2022 14:20:39",
          "content": "<p>Thank you for both Radek and Chris; nowThank you for both Radek and Chris, now I am clear with the intuition of co-visitation matrices</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2061273,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/10/2022 22:27:19",
          "content": "<p>Wonderful! that is great to hear <a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">@leiwong</a>! 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2173215,
          "author_name": "edoziemenyinnaya",
          "author_url": "",
          "post_date": "03/08/2023 07:45:01",
          "content": "<p>Can the same principle of co-visitation matrix work when dealing with datasets of restaurants, for recommending a restaurant for a user?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2025050,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/11/2022 00:32:37",
      "content": "<p>Training Word2Vec would be very helpful. That would produce embeddings such that similar items (with co-visitation) would have distance similar embeddings. It would help us understand what these anonymous item ids represent.</p>\n<p>Then we could do lots of EDA like how do the items that users purchase over time change? How do items purchased in the morning differ from items purchased in the evening. Etc etc. For each of these analysis, we could plot 2D pictures with clusters of dots (using their Word2Vec embedding 2D projection)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2025058,
      "author_name": "oscarm524",
      "author_url": "",
      "post_date": "11/11/2022 00:41:37",
      "content": "<p>Well explained. Thanks for sharing👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2025072,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/11/2022 00:49:31",
          "content": "<p>my pleasure, <a href=\"https://www.kaggle.com/oscarm524\" target=\"_blank\">@oscarm524</a>! very glad you found this useful! 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2105430,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "01/18/2023 13:52:08",
      "content": "<p>Hi Radek, I didn't appreciate this post of yours enough when I read it last month, now what you said about co-visitation matrix makes more sense to me now as I was writing down my reflection of co-visitation matrix these two days. </p>\n<p>Here is a question I have about comparing co-visitation matrix and word2vect.  Since you said co-visitation matrix suffers the same thing as trigram, then word2vec would in theory be a better model for generating candidates right? </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F27930a141e9e30ced711ae9a91e518ad%2FScreen%20Shot%202023-01-18%20at%2021.49.24.png?generation=1674049779499957&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2105932,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "01/18/2023 20:42:11",
          "content": "<p>In theory, matrix factorization (and word2vec) should do better! But it is important to check how they work in practice. Here, I only played a little bit around with candidate generation using MF and didn't get better results than using the co-visitation matrix. But there could be many things at play.</p>\n<p>In general, yes, matrix factorization and word2vec should be superior methods to the co-visitation matrix, that is my belief. But all of this is very situational dependant and depending on the dataset a heuristic such as the co-visitation matrix might work best 🙂 (for instances, the co-visitation matrix might better capture short-term trends, or can work really well if there is a long-tail of items purchased).</p>\n<p>It is fun and useful to think about these methods, what they do, how they work, etc but I wouldn't draw too definite conclusions, always best to test your hypotheses out on the dataset you are working with 🙂</p>",
          "votes": null,
          "replies": [
            {
              "id": 2106172,
              "author_name": "danielliao",
              "author_url": "",
              "post_date": "01/19/2023 02:13:52",
              "content": "<p>Thank you so much Radek! What an illuminating and insightful answer! I am looking forward to test them out!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2024992": "It is very interesting to think of modern techniques in the context of their roots.\n\nFor instance, when thinking about RNNs we should consider unigram, bigram, and trigram models.\n\nWhat were they?\n\nThey estimated the probability of a word given the words that came before.\n\nWe see \"Radek is a _________\".\n\nIn a trigram model, we would consider the chain \"Radek\", \"is\", and \"a\" and could find the most likely word to come next.\n\nThe easiest way would be to count the occurrences of \"Radek is a\" and the words that came after and pick the one most common.\n\nSo why RNNs are such a great improvement?\n\nWell, there are not that many \"Radek is a _________\" in most text corpora!\n\nWith RNNs (or word2vec) we could operate on embeddings and look at many more words in sequence.\n\nInstead of looking at only \"Radek is a ________\" we can use \"Radek was an ___________\", \"Tommy is a __________\" as examples to learn from, etc.\n\nEmbeddings take values of a variable of high cardinality and project them to a representation that ideally captures similarities.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F70e2a7f1b9621cd0afb2b815a7b3e01b%2Fyou_shall_know_a_word.png?generation=1668122115555809&alt=media)\n\nSo how does this relate to the co-visitation matrix?\n\nA co-visitation matrix counts the co-occurrence of two actions in close proximity.\n\nIf a user bought A and shortly after bought B, we store these values together.\n\nWe calculate counts and use them to estimate the probability of future actions based on recent history.\n\nIt is quite important to understand what is happening in the co-visitation matrix approach...\n\nSince it suffers from the same issues as our trigram example!\n\nPlus what does the co-visitation matrix resemble?\n\nYou are right, it is akin to doing Matrix Factorization by counting!\n\nIt is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2024995": "Nice explanation....Really helpful...Thanks👍",
    "2025025": "How funny that I was playing with the Word2Vec and its variants and then saw this post :D",
    "2025036": "Yeah, it would be really interesting to see what word2vec can do in this competition 🙂 Might be it could be quite nice for candidate generation or some sort of similarity scoring for reranking 🤔\n\nAnyhow, cool that we were both circulating around the same topic! 🙌",
    "2025045": "Co-visitation matrices are the natural statistical approach. It is a way to compute conditional probabilities. Given a user has interacted with item A, it computes the most likely future items that this user will interact with.\n\nAnd we can make many more specific variants such as given a user has **ordered** item A, what is the most likely item that this user will **click, cart, order** etc. Or given a user has visited this item **in the morning**, what is the most likely item that this user will **order**. Etc etc.\n\nOf course if we train a reranker model it will learn all these varieties of conditional probabilities on its own.",
    "2025050": "Training Word2Vec would be very helpful. That would produce embeddings such that similar items (with co-visitation) would have distance similar embeddings. It would help us understand what these anonymous item ids represent.\n\nThen we could do lots of EDA like how do the items that users purchase over time change? How do items purchased in the morning differ from items purchased in the evening. Etc etc. For each of these analysis, we could plot 2D pictures with clusters of dots (using their Word2Vec embedding 2D projection)",
    "2025058": "Well explained. Thanks for sharing👍",
    "2025072": "my pleasure, @oscarm524! very glad you found this useful! 🙏",
    "2025074": "That is some amazing information @cdeotte, as always! 🙂 Thank you so very much for sharing all this with us 🙏",
    "2025075": "Yes, exactly! I was also thinking of it for the candidate generation. 🤜🤛",
    "2025093": "Word2Vec would also be great for EDA. It (with 2D projection via UMAP or TSNE) would help us visualize these anonymous items.",
    "2025750": "Great idea(s), Chris! Also as you have mentioned in your other comments, visualizations based on time would be insightful! I have the model, let's see if I can come up with something.",
    "2060868": "Thank you for both Radek and Chris; nowThank you for both Radek and Chris, now I am clear with the intuition of co-visitation matrices",
    "2061273": "Wonderful! that is great to hear @leiwong! 🙂",
    "2105430": "Hi Radek, I didn't appreciate this post of yours enough when I read it last month, now what you said about co-visitation matrix makes more sense to me now as I was writing down my reflection of co-visitation matrix these two days. \n\nHere is a question I have about comparing co-visitation matrix and word2vect.  Since you said co-visitation matrix suffers the same thing as trigram, then word2vec would in theory be a better model for generating candidates right? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F27930a141e9e30ced711ae9a91e518ad%2FScreen%20Shot%202023-01-18%20at%2021.49.24.png?generation=1674049779499957&alt=media)",
    "2105932": "In theory, matrix factorization (and word2vec) should do better! But it is important to check how they work in practice. Here, I only played a little bit around with candidate generation using MF and didn't get better results than using the co-visitation matrix. But there could be many things at play.\n\nIn general, yes, matrix factorization and word2vec should be superior methods to the co-visitation matrix, that is my belief. But all of this is very situational dependant and depending on the dataset a heuristic such as the co-visitation matrix might work best 🙂 (for instances, the co-visitation matrix might better capture short-term trends, or can work really well if there is a long-tail of items purchased).\n\nIt is fun and useful to think about these methods, what they do, how they work, etc but I wouldn't draw too definite conclusions, always best to test your hypotheses out on the dataset you are working with 🙂",
    "2106172": "Thank you so much Radek! What an illuminating and insightful answer! I am looking forward to test them out!",
    "2106250": "My pleasure, @saha8631! 🙂 Thank you for your kind comment! 🙌",
    "2173215": "Can the same principle of co-visitation matrix work when dealing with datasets of restaurants, for recommending a restaurant for a user?"
  },
  "source": "meta"
}