{
  "id": 372585,
  "title": "Difference between Co-visitation Matrix v.s. Matrix Factorization",
  "url": "/competitions/otto-recommender-system/discussion/372585",
  "author_name": "Bilzard",
  "post_date": "2022-12-16T18:08:29.855000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>From looking at the scores of the publicly posted notebooks so far, it seems that the approach of the co-visitation matrix appears to score higher than the matrix decomposition approach. I will share one of the possible reasons for this.<br>\nHere is a point I would like to mention: the co-visitation matrix approach <a href=\"https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\" target=\"_blank\">1</a> and the matrix decomposition approach <a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">2</a> in the publicly posted notebooks seem to be interchangeable approaches at first glance, but there is one decisive difference between them. That is the dependence of events. It could be called the context in consideration.</p>\n<p>First, let me explain the common points of both approaches. Both approaches model $P(F|C)$. Here, $C$ represents the context given as a query, i.e. the event history, and $F$ represents the expected set of future events, i.e. the correct labels.</p>\n<p>Next, let me explain the differences between the two. The co-visitation matrix takes events that occurred during the past and future of the last day (i.e. clicks, carts, orders) as the context, while the matrix decomposition approach only gives the previous event as the context (refer to the figure). The learning method of the co-visitation matrix is similar to BERT, while the matrix decomposition is similar to the bigram. In short, the current matrix decomposition approach has less information given as context compared to the co-visitation matrix approach.</p>\n<p>Since there is no evidence to support this, it is only a hypothesis at this point.</p>\n<p>One way to improve the matrix decomposition approach is to add pairs of the previous and next N events. If the score improves as a result, it will be one piece of evidence to support the hypothesis.</p>\n<p><a href=\"https://ibb.co/2MZ5KX8\"><img src=\"https://i.ibb.co/kDmMcY3/Screen-Shot-2022-12-17-at-2-43-27.png\" alt=\"Screen-Shot-2022-12-17-at-2-43-27\"></a></p>",
  "messages": [
    {
      "id": 2067454,
      "postDate": "2022-12-16T18:08:29.857Z",
      "content": "<p>From looking at the scores of the publicly posted notebooks so far, it seems that the approach of the co-visitation matrix appears to score higher than the matrix decomposition approach. I will share one of the possible reasons for this.<br>\nHere is a point I would like to mention: the co-visitation matrix approach <a href=\"https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\" target=\"_blank\">1</a> and the matrix decomposition approach <a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">2</a> in the publicly posted notebooks seem to be interchangeable approaches at first glance, but there is one decisive difference between them. That is the dependence of events. It could be called the context in consideration.</p>\n<p>First, let me explain the common points of both approaches. Both approaches model $P(F|C)$. Here, $C$ represents the context given as a query, i.e. the event history, and $F$ represents the expected set of future events, i.e. the correct labels.</p>\n<p>Next, let me explain the differences between the two. The co-visitation matrix takes events that occurred during the past and future of the last day (i.e. clicks, carts, orders) as the context, while the matrix decomposition approach only gives the previous event as the context (refer to the figure). The learning method of the co-visitation matrix is similar to BERT, while the matrix decomposition is similar to the bigram. In short, the current matrix decomposition approach has less information given as context compared to the co-visitation matrix approach.</p>\n<p>Since there is no evidence to support this, it is only a hypothesis at this point.</p>\n<p>One way to improve the matrix decomposition approach is to add pairs of the previous and next N events. If the score improves as a result, it will be one piece of evidence to support the hypothesis.</p>\n<p><a href=\"https://ibb.co/2MZ5KX8\"><img src=\"https://i.ibb.co/kDmMcY3/Screen-Shot-2022-12-17-at-2-43-27.png\" alt=\"Screen-Shot-2022-12-17-at-2-43-27\"></a></p>",
      "rawMarkdown": "From looking at the scores of the publicly posted notebooks so far, it seems that the approach of the co-visitation matrix appears to score higher than the matrix decomposition approach. I will share one of the possible reasons for this.\nHere is a point I would like to mention: the co-visitation matrix approach [1] and the matrix decomposition approach [2] in the publicly posted notebooks seem to be interchangeable approaches at first glance, but there is one decisive difference between them. That is the dependence of events. It could be called the context in consideration.\n\nFirst, let me explain the common points of both approaches. Both approaches model $P(F|C)$. Here, $C$ represents the context given as a query, i.e. the event history, and $F$ represents the expected set of future events, i.e. the correct labels.\n\nNext, let me explain the differences between the two. The co-visitation matrix takes events that occurred during the past and future of the last day (i.e. clicks, carts, orders) as the context, while the matrix decomposition approach only gives the previous event as the context (refer to the figure). The learning method of the co-visitation matrix is similar to BERT, while the matrix decomposition is similar to the bigram. In short, the current matrix decomposition approach has less information given as context compared to the co-visitation matrix approach.\n\nSince there is no evidence to support this, it is only a hypothesis at this point.\n\nOne way to improve the matrix decomposition approach is to add pairs of the previous and next N events. If the score improves as a result, it will be one piece of evidence to support the hypothesis.\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\n\n<a href=\"https://ibb.co/2MZ5KX8\"><img src=\"https://i.ibb.co/kDmMcY3/Screen-Shot-2022-12-17-at-2-43-27.png\" alt=\"Screen-Shot-2022-12-17-at-2-43-27\" border=\"0\"></a>\n",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2067454": "From looking at the scores of the publicly posted notebooks so far, it seems that the approach of the co-visitation matrix appears to score higher than the matrix decomposition approach. I will share one of the possible reasons for this.\nHere is a point I would like to mention: the co-visitation matrix approach [1] and the matrix decomposition approach [2] in the publicly posted notebooks seem to be interchangeable approaches at first glance, but there is one decisive difference between them. That is the dependence of events. It could be called the context in consideration.\n\nFirst, let me explain the common points of both approaches. Both approaches model $P(F|C)$. Here, $C$ represents the context given as a query, i.e. the event history, and $F$ represents the expected set of future events, i.e. the correct labels.\n\nNext, let me explain the differences between the two. The co-visitation matrix takes events that occurred during the past and future of the last day (i.e. clicks, carts, orders) as the context, while the matrix decomposition approach only gives the previous event as the context (refer to the figure). The learning method of the co-visitation matrix is similar to BERT, while the matrix decomposition is similar to the bigram. In short, the current matrix decomposition approach has less information given as context compared to the co-visitation matrix approach.\n\nSince there is no evidence to support this, it is only a hypothesis at this point.\n\nOne way to improve the matrix decomposition approach is to add pairs of the previous and next N events. If the score improves as a result, it will be one piece of evidence to support the hypothesis.\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\n\n<a href=\"https://ibb.co/2MZ5KX8\"><img src=\"https://i.ibb.co/kDmMcY3/Screen-Shot-2022-12-17-at-2-43-27.png\" alt=\"Screen-Shot-2022-12-17-at-2-43-27\" border=\"0\"></a>\n"
  }
}