{
  "id": 21014,
  "title": "Collaborative Filtering",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21014",
  "author_name": "",
  "post_date": "2016-05-17T13:38:27.547Z",
  "votes": 1,
  "comment_count": 5,
  "views": 1508,
  "content": "<p>Collaborative filtering is a classical solution in recommender systems. \nI have had a look at many of scripts, but few of the solutions are based on collaborative filtering. Why is this the case?\nDo you think collaborative filtering can handle the huge data set?</p>\n\n<p>I guess forming a User-Item-Rating matrix is not a big deal!\nThanks</p>",
  "messages": [
    {
      "id": "120329",
      "postDate": "05/17/2016 13:38:27",
      "content": "<p>Collaborative filtering is a classical solution in recommender systems. \nI have had a look at many of scripts, but few of the solutions are based on collaborative filtering. Why is this the case?\nDo you think collaborative filtering can handle the huge data set?</p>\n\n<p>I guess forming a User-Item-Rating matrix is not a big deal!\nThanks</p>",
      "rawMarkdown": "Collaborative filtering is a classical solution in recommender systems. \r\nI have had a look at many of scripts, but few of the solutions are based on collaborative filtering. Why is this the case?\r\nDo you think collaborative filtering can handle the huge data set?\r\n\r\nI guess forming a User-Item-Rating matrix is not a big deal!\r\nThanks",
      "votes": null
    },
    {
      "id": "120493",
      "postDate": "05/18/2016 16:45:26",
      "content": "<p>I'll take a stab at this but anyone else please correct me where I'm wrong. I think collaborative filtering is usually (if not always?) used as an unsupervised method of machine learning. If you tried to apply collaborative filtering to this problem, you would need to obtain clustering of the hotels in exactly the way Expedia has done it. I don't know of a method that allows you search for the best way of doing this. Instead, most people have approached this as a supervised learning problem where they need only classify new user activity based on previously existing user activity.</p>",
      "rawMarkdown": "I'll take a stab at this but anyone else please correct me where I'm wrong. I think collaborative filtering is usually (if not always?) used as an unsupervised method of machine learning. If you tried to apply collaborative filtering to this problem, you would need to obtain clustering of the hotels in exactly the way Expedia has done it. I don't know of a method that allows you search for the best way of doing this. Instead, most people have approached this as a supervised learning problem where they need only classify new user activity based on previously existing user activity.",
      "votes": null
    },
    {
      "id": "120575",
      "postDate": "05/19/2016 07:02:16",
      "content": "<p>Collaborative filtering predicts users' ratings on unrated items. This problem is not directly about predicting rating on unrated items. Your user-item-rating matrix might be useful in ranking hotel clusters based on users' preference. But there is more to it. Timing of check-in, location specific info, etc. </p>\n\n<p>From my own experiment, user's preference contributed about 0.006~0.007 to the score, when combined with a strong model. I didn't use the collaborative filtering approach though. So definitely, users' preference is important for this problem. How to squeeze the score out of it is the challenge.</p>",
      "rawMarkdown": "Collaborative filtering predicts users' ratings on unrated items. This problem is not directly about predicting rating on unrated items. Your user-item-rating matrix might be useful in ranking hotel clusters based on users' preference. But there is more to it. Timing of check-in, location specific info, etc. \r\n\r\nFrom my own experiment, user's preference contributed about 0.006~0.007 to the score, when combined with a strong model. I didn't use the collaborative filtering approach though. So definitely, users' preference is important for this problem. How to squeeze the score out of it is the challenge.",
      "votes": null
    },
    {
      "id": "120602",
      "postDate": "05/19/2016 12:21:02",
      "content": "<p>hi @Andaman</p>\n\n<p>What do you mean by user preference? Do you mean that we need to track each of them with their user_id and check the clusters they usually pick?</p>\n\n<p>Thanks in advance.</p>",
      "rawMarkdown": "hi @Andaman\r\n\r\nWhat do you mean by user preference? Do you mean that we need to track each of them with their user_id and check the clusters they usually pick?\r\n\r\nThanks in advance.",
      "votes": null
    },
    {
      "id": "120613",
      "postDate": "05/19/2016 13:48:44",
      "content": "<p>hi zyazzy,</p>\n\n<p>For example, look at the public scripts that yield 0.50xx. They use 'user_id' as part of the hash key. That's users' preference. Yes, you can track each one of them and check the clusters they picked. Those scripts also have other contexts in the keys. So it's much more than tracking individual user id.</p>",
      "rawMarkdown": "hi zyazzy,\r\n\r\nFor example, look at the public scripts that yield 0.50xx. They use 'user_id' as part of the hash key. That's users' preference. Yes, you can track each one of them and check the clusters they picked. Those scripts also have other contexts in the keys. So it's much more than tracking individual user id.",
      "votes": null
    },
    {
      "id": "122897",
      "postDate": "06/08/2016 05:39:15",
      "content": "<p>Can somebody share with me the link of the script of the solution where collaborative filtering is implemented</p>",
      "rawMarkdown": "Can somebody share with me the link of the script of the solution where collaborative filtering is implemented",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 120493,
      "author_name": "kennmyers",
      "author_url": "",
      "post_date": "05/18/2016 16:45:26",
      "content": "<p>I'll take a stab at this but anyone else please correct me where I'm wrong. I think collaborative filtering is usually (if not always?) used as an unsupervised method of machine learning. If you tried to apply collaborative filtering to this problem, you would need to obtain clustering of the hotels in exactly the way Expedia has done it. I don't know of a method that allows you search for the best way of doing this. Instead, most people have approached this as a supervised learning problem where they need only classify new user activity based on previously existing user activity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120575,
      "author_name": "pichai",
      "author_url": "",
      "post_date": "05/19/2016 07:02:16",
      "content": "<p>Collaborative filtering predicts users' ratings on unrated items. This problem is not directly about predicting rating on unrated items. Your user-item-rating matrix might be useful in ranking hotel clusters based on users' preference. But there is more to it. Timing of check-in, location specific info, etc. </p>\n\n<p>From my own experiment, user's preference contributed about 0.006~0.007 to the score, when combined with a strong model. I didn't use the collaborative filtering approach though. So definitely, users' preference is important for this problem. How to squeeze the score out of it is the challenge.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120602,
      "author_name": "zyazzy",
      "author_url": "",
      "post_date": "05/19/2016 12:21:02",
      "content": "<p>hi @Andaman</p>\n\n<p>What do you mean by user preference? Do you mean that we need to track each of them with their user_id and check the clusters they usually pick?</p>\n\n<p>Thanks in advance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120613,
      "author_name": "pichai",
      "author_url": "",
      "post_date": "05/19/2016 13:48:44",
      "content": "<p>hi zyazzy,</p>\n\n<p>For example, look at the public scripts that yield 0.50xx. They use 'user_id' as part of the hash key. That's users' preference. Yes, you can track each one of them and check the clusters they picked. Those scripts also have other contexts in the keys. So it's much more than tracking individual user id.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122897,
      "author_name": "",
      "author_url": "",
      "post_date": "06/08/2016 05:39:15",
      "content": "<p>Can somebody share with me the link of the script of the solution where collaborative filtering is implemented</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "120329": "Collaborative filtering is a classical solution in recommender systems. \r\nI have had a look at many of scripts, but few of the solutions are based on collaborative filtering. Why is this the case?\r\nDo you think collaborative filtering can handle the huge data set?\r\n\r\nI guess forming a User-Item-Rating matrix is not a big deal!\r\nThanks",
    "120493": "I'll take a stab at this but anyone else please correct me where I'm wrong. I think collaborative filtering is usually (if not always?) used as an unsupervised method of machine learning. If you tried to apply collaborative filtering to this problem, you would need to obtain clustering of the hotels in exactly the way Expedia has done it. I don't know of a method that allows you search for the best way of doing this. Instead, most people have approached this as a supervised learning problem where they need only classify new user activity based on previously existing user activity.",
    "120575": "Collaborative filtering predicts users' ratings on unrated items. This problem is not directly about predicting rating on unrated items. Your user-item-rating matrix might be useful in ranking hotel clusters based on users' preference. But there is more to it. Timing of check-in, location specific info, etc. \r\n\r\nFrom my own experiment, user's preference contributed about 0.006~0.007 to the score, when combined with a strong model. I didn't use the collaborative filtering approach though. So definitely, users' preference is important for this problem. How to squeeze the score out of it is the challenge.",
    "120602": "hi @Andaman\r\n\r\nWhat do you mean by user preference? Do you mean that we need to track each of them with their user_id and check the clusters they usually pick?\r\n\r\nThanks in advance.",
    "120613": "hi zyazzy,\r\n\r\nFor example, look at the public scripts that yield 0.50xx. They use 'user_id' as part of the hash key. That's users' preference. Yes, you can track each one of them and check the clusters they picked. Those scripts also have other contexts in the keys. So it's much more than tracking individual user id.",
    "122897": "Can somebody share with me the link of the script of the solution where collaborative filtering is implemented"
  },
  "source": "meta"
}