{
  "id": 382254,
  "title": "Discussion: If Collaborative Filtering Algorithm is suitable for this challenge or not?",
  "url": "/competitions/otto-recommender-system/discussion/382254",
  "author_name": "liuxudong1986",
  "post_date": "2023-01-30T09:50:09.450000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I just learnt a classic recommendation algorithm, collaborative filtering algorithm, which use the user's rate of items and feature matrix to train a recommendation system. And want to discuss here how suitable it is to this challenge.<br>\nIn my experiment, the items in collaborative filtering algorithm are the aids in this challenge, users are the sessions, although there is no rate of each item, but still can treat it as binary collaborative filtering application, for example click/not click, so for each aid and each session, you can mark it as 1 (click) or 0 (not click). It will be like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2Fcebca88007f3f537fea2e2c484c82cb2%2FY.PNG?generation=1675071369741715&amp;alt=media\" alt=\"\"><br>\nThere is no feature matrix but can assume there are 100 features and they can be trained using collaborative filtering algorithm. Because this is binary recommendation system, so the R matrix in collaborative filtering algorithm will be one.<br>\nIn this way, I can get all the necessary inputs for collaborative filtering algorithm, then use the data and gradient descent to train the model like this:</p>\n<pre><code>Training loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \n...\nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nIteration stops at iteration \n</code></pre>\n<p>After training, I can get a trained feature matrix, and it has learned the features of each iteam (aid), so I can cluster aid using the features. The recommendation will be based on the aids in same cluster. <br>\nIf use cluster algorithm like KMeans to cluster the aids, I can get the cluster visualization like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2F122f4bd4da99ef330d60380c2c0b964f%2Faid_cluster.PNG?generation=1675071906313809&amp;alt=media\" alt=\"\"></p>\n<p>So want to discuss here, is such kind of proposal is a possible way?</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": 2121410,
      "postDate": "2023-01-30T09:50:09.450Z",
      "content": "<p>I just learnt a classic recommendation algorithm, collaborative filtering algorithm, which use the user's rate of items and feature matrix to train a recommendation system. And want to discuss here how suitable it is to this challenge.<br>\nIn my experiment, the items in collaborative filtering algorithm are the aids in this challenge, users are the sessions, although there is no rate of each item, but still can treat it as binary collaborative filtering application, for example click/not click, so for each aid and each session, you can mark it as 1 (click) or 0 (not click). It will be like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2Fcebca88007f3f537fea2e2c484c82cb2%2FY.PNG?generation=1675071369741715&amp;alt=media\" alt=\"\"><br>\nThere is no feature matrix but can assume there are 100 features and they can be trained using collaborative filtering algorithm. Because this is binary recommendation system, so the R matrix in collaborative filtering algorithm will be one.<br>\nIn this way, I can get all the necessary inputs for collaborative filtering algorithm, then use the data and gradient descent to train the model like this:</p>\n<pre><code>Training loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \n...\nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nTraining loss at iteration : \nIteration stops at iteration \n</code></pre>\n<p>After training, I can get a trained feature matrix, and it has learned the features of each iteam (aid), so I can cluster aid using the features. The recommendation will be based on the aids in same cluster. <br>\nIf use cluster algorithm like KMeans to cluster the aids, I can get the cluster visualization like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2F122f4bd4da99ef330d60380c2c0b964f%2Faid_cluster.PNG?generation=1675071906313809&amp;alt=media\" alt=\"\"></p>\n<p>So want to discuss here, is such kind of proposal is a possible way?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "I just learnt a classic recommendation algorithm, collaborative filtering algorithm, which use the user's rate of items and feature matrix to train a recommendation system. And want to discuss here how suitable it is to this challenge.\nIn my experiment, the items in collaborative filtering algorithm are the aids in this challenge, users are the sessions, although there is no rate of each item, but still can treat it as binary collaborative filtering application, for example click/not click, so for each aid and each session, you can mark it as 1 (click) or 0 (not click). It will be like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2Fcebca88007f3f537fea2e2c484c82cb2%2FY.PNG?generation=1675071369741715&alt=media)\nThere is no feature matrix but can assume there are 100 features and they can be trained using collaborative filtering algorithm. Because this is binary recommendation system, so the R matrix in collaborative filtering algorithm will be one.\nIn this way, I can get all the necessary inputs for collaborative filtering algorithm, then use the data and gradient descent to train the model like this:\n```python\nTraining loss at iteration 0: 1431177.3\nTraining loss at iteration 1: 1155055.0\nTraining loss at iteration 2: 904407.1\nTraining loss at iteration 3: 665228.5\nTraining loss at iteration 4: 454217.2\n...\nTraining loss at iteration 97: 9518.7\nTraining loss at iteration 98: 9507.3\nTraining loss at iteration 99: 9496.4\nTraining loss at iteration 100: 9486.1\nTraining loss at iteration 101: 9476.2\nIteration stops at iteration 101\n```\nAfter training, I can get a trained feature matrix, and it has learned the features of each iteam (aid), so I can cluster aid using the features. The recommendation will be based on the aids in same cluster. \nIf use cluster algorithm like KMeans to cluster the aids, I can get the cluster visualization like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2F122f4bd4da99ef330d60380c2c0b964f%2Faid_cluster.PNG?generation=1675071906313809&alt=media)\n\nSo want to discuss here, is such kind of proposal is a possible way?\n\nThank you!",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2121410": "I just learnt a classic recommendation algorithm, collaborative filtering algorithm, which use the user's rate of items and feature matrix to train a recommendation system. And want to discuss here how suitable it is to this challenge.\nIn my experiment, the items in collaborative filtering algorithm are the aids in this challenge, users are the sessions, although there is no rate of each item, but still can treat it as binary collaborative filtering application, for example click/not click, so for each aid and each session, you can mark it as 1 (click) or 0 (not click). It will be like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2Fcebca88007f3f537fea2e2c484c82cb2%2FY.PNG?generation=1675071369741715&alt=media)\nThere is no feature matrix but can assume there are 100 features and they can be trained using collaborative filtering algorithm. Because this is binary recommendation system, so the R matrix in collaborative filtering algorithm will be one.\nIn this way, I can get all the necessary inputs for collaborative filtering algorithm, then use the data and gradient descent to train the model like this:\n```python\nTraining loss at iteration 0: 1431177.3\nTraining loss at iteration 1: 1155055.0\nTraining loss at iteration 2: 904407.1\nTraining loss at iteration 3: 665228.5\nTraining loss at iteration 4: 454217.2\n...\nTraining loss at iteration 97: 9518.7\nTraining loss at iteration 98: 9507.3\nTraining loss at iteration 99: 9496.4\nTraining loss at iteration 100: 9486.1\nTraining loss at iteration 101: 9476.2\nIteration stops at iteration 101\n```\nAfter training, I can get a trained feature matrix, and it has learned the features of each iteam (aid), so I can cluster aid using the features. The recommendation will be based on the aids in same cluster. \nIf use cluster algorithm like KMeans to cluster the aids, I can get the cluster visualization like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6376068%2F122f4bd4da99ef330d60380c2c0b964f%2Faid_cluster.PNG?generation=1675071906313809&alt=media)\n\nSo want to discuss here, is such kind of proposal is a possible way?\n\nThank you!"
  }
}