{
  "id": 20767,
  "title": "How useful will these recommendations be?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20767",
  "author_name": "",
  "post_date": "2016-05-06T13:15:40.347Z",
  "votes": 1,
  "comment_count": 7,
  "views": 1661,
  "content": "<p>In the test data set we are provided with the records corresponding to booked data\nbut in the real time implementation of this recommendation engine we need to give recommendation to a customer while he is just clicking and browsing AND NOT WHILE HE IS MAKING AN ACTUAL BOOKING. So i wonder how useful will this recommendation engine be? </p>",
  "messages": [
    {
      "id": "118974",
      "postDate": "05/06/2016 13:15:40",
      "content": "<p>In the test data set we are provided with the records corresponding to booked data\nbut in the real time implementation of this recommendation engine we need to give recommendation to a customer while he is just clicking and browsing AND NOT WHILE HE IS MAKING AN ACTUAL BOOKING. So i wonder how useful will this recommendation engine be? </p>",
      "rawMarkdown": "In the test data set we are provided with the records corresponding to booked data\r\nbut in the real time implementation of this recommendation engine we need to give recommendation to a customer while he is just clicking and browsing AND NOT WHILE HE IS MAKING AN ACTUAL BOOKING. So i wonder how useful will this recommendation engine be?",
      "votes": null
    },
    {
      "id": "118977",
      "postDate": "05/06/2016 13:48:34",
      "content": "<p>Good observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier. </p>",
      "rawMarkdown": "Good observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier.",
      "votes": null
    },
    {
      "id": "119026",
      "postDate": "05/06/2016 18:52:21",
      "content": "<p>I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.</p>\n\n<p>I understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.</p>",
      "rawMarkdown": "I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.\r\n\r\nI understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.",
      "votes": null
    },
    {
      "id": "119045",
      "postDate": "05/06/2016 22:17:03",
      "content": "<p>Well, at the very least you have the clicks that the user made before each booking. And his/her previous bookings...</p>",
      "rawMarkdown": "Well, at the very least you have the clicks that the user made before each booking. And his/her previous bookings...",
      "votes": null
    },
    {
      "id": "119403",
      "postDate": "05/09/2016 23:55:50",
      "content": "<p>[quote=GreatDataAnalyst;119026]</p>\n\n<p>I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.</p>\n\n<p>I understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.</p>\n\n<p>[/quote]</p>\n\n<p>user_id itself MIGHT not be important, but the fact the it reflects the user doing to booking it is obviously an importante thing to pay attention to.\nThe  more you know about a specific user the more you can target him/her with specific ads.</p>\n\n<p>Thats why we have all other click history, to make up how the user behavior on the past impacts its current booking.</p>",
      "rawMarkdown": "[quote=GreatDataAnalyst;119026]\r\n\r\nI am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.\r\n\r\nI understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.\r\n\r\n[/quote]\r\n\r\nuser_id itself MIGHT not be important, but the fact the it reflects the user doing to booking it is obviously an importante thing to pay attention to.\r\nThe  more you know about a specific user the more you can target him/her with specific ads.\r\n\r\nThats why we have all other click history, to make up how the user behavior on the past impacts its current booking.",
      "votes": null
    },
    {
      "id": "120338",
      "postDate": "05/17/2016 15:00:41",
      "content": "<p>[quote=Adam;118977]</p>\n\n<p>Good observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier. </p>\n\n<p>[/quote]</p>\n\n<p>@Admin: If i understand correctly, is the idea to first predict the cluster, and then the top hotels for that cluster through some in-house algorithm? Also, since you said the realtime implementation would involve recommendations when the user visits, prior to any clicks, what exactly are we trying to achieve in this prediction exercise where we have results from user's clicks and searches like hotel market, search destination id ? Or is it so that the recommendations can be made at any level or the user's browsing process?</p>",
      "rawMarkdown": "[quote=Adam;118977]\r\n\r\nGood observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier. \r\n\r\n\r\n[/quote]\r\n\r\n@Admin: If i understand correctly, is the idea to first predict the cluster, and then the top hotels for that cluster through some in-house algorithm? Also, since you said the realtime implementation would involve recommendations when the user visits, prior to any clicks, what exactly are we trying to achieve in this prediction exercise where we have results from user's clicks and searches like hotel market, search destination id ? Or is it so that the recommendations can be made at any level or the user's browsing process?",
      "votes": null
    },
    {
      "id": "120617",
      "postDate": "05/19/2016 13:59:32",
      "content": "<p>Yes, we would like to make hotel recommendations at any level of the browsing process. The problem is that at the beginning of the &quot;conversion funnel&quot; there is usually not much click data we can rely on. </p>\n\n<p>Predicting hotel cluster can indeed be used as a per-filtering step which will limit the number of hotels; filtered hotels can be then sorted by our in-house algorithm that uses information on live pricing and promotions.</p>",
      "rawMarkdown": "Yes, we would like to make hotel recommendations at any level of the browsing process. The problem is that at the beginning of the \"conversion funnel\" there is usually not much click data we can rely on. \r\n\r\nPredicting hotel cluster can indeed be used as a per-filtering step which will limit the number of hotels; filtered hotels can be then sorted by our in-house algorithm that uses information on live pricing and promotions.",
      "votes": null
    },
    {
      "id": "120738",
      "postDate": "05/20/2016 07:22:31",
      "content": "<p>[quote=GreatDataAnalyst;119026]</p>\n\n<p>I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.</p>\n\n<p>I understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.</p>\n\n<p>[/quote]</p>\n\n<p>According to my experiments, using just user_id and the history of  his/her  hotel_cluster booking without any other context, one can obtain MAP@5~0.107. It is very low value, if we compare it with e.g. most frequent benchmark  MAP@5~0.06 or the best script MAP@5~0.50x. However, we should know, that a great part (~1/3) of the results MAP@5~0.50x states almost perfect prediction of data with leakage. The real prediction strength of more or less sophisticated algorithms is in the range 0.2-0.25.  In that context I do not know if we may say that the history of user's booking has low predictive value.</p>",
      "rawMarkdown": "[quote=GreatDataAnalyst;119026]\r\n\r\nI am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.\r\n\r\nI understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.\r\n\r\n[/quote]\r\n\r\nAccording to my experiments, using just user_id and the history of  his/her  hotel_cluster booking without any other context, one can obtain MAP@5~0.107. It is very low value, if we compare it with e.g. most frequent benchmark  MAP@5~0.06 or the best script MAP@5~0.50x. However, we should know, that a great part (~1/3) of the results MAP@5~0.50x states almost perfect prediction of data with leakage. The real prediction strength of more or less sophisticated algorithms is in the range 0.2-0.25.  In that context I do not know if we may say that the history of user's booking has low predictive value.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118977,
      "author_name": "adamwoz",
      "author_url": "",
      "post_date": "05/06/2016 13:48:34",
      "content": "<p>Good observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119026,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "05/06/2016 18:52:21",
      "content": "<p>I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.</p>\n\n<p>I understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119045,
      "author_name": "carrdelling",
      "author_url": "",
      "post_date": "05/06/2016 22:17:03",
      "content": "<p>Well, at the very least you have the clicks that the user made before each booking. And his/her previous bookings...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119403,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "05/09/2016 23:55:50",
      "content": "<p>[quote=GreatDataAnalyst;119026]</p>\n\n<p>I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.</p>\n\n<p>I understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.</p>\n\n<p>[/quote]</p>\n\n<p>user_id itself MIGHT not be important, but the fact the it reflects the user doing to booking it is obviously an importante thing to pay attention to.\nThe  more you know about a specific user the more you can target him/her with specific ads.</p>\n\n<p>Thats why we have all other click history, to make up how the user behavior on the past impacts its current booking.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120338,
      "author_name": "piyushjaiswal14",
      "author_url": "",
      "post_date": "05/17/2016 15:00:41",
      "content": "<p>[quote=Adam;118977]</p>\n\n<p>Good observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier. </p>\n\n<p>[/quote]</p>\n\n<p>@Admin: If i understand correctly, is the idea to first predict the cluster, and then the top hotels for that cluster through some in-house algorithm? Also, since you said the realtime implementation would involve recommendations when the user visits, prior to any clicks, what exactly are we trying to achieve in this prediction exercise where we have results from user's clicks and searches like hotel market, search destination id ? Or is it so that the recommendations can be made at any level or the user's browsing process?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120617,
      "author_name": "adamwoz",
      "author_url": "",
      "post_date": "05/19/2016 13:59:32",
      "content": "<p>Yes, we would like to make hotel recommendations at any level of the browsing process. The problem is that at the beginning of the &quot;conversion funnel&quot; there is usually not much click data we can rely on. </p>\n\n<p>Predicting hotel cluster can indeed be used as a per-filtering step which will limit the number of hotels; filtered hotels can be then sorted by our in-house algorithm that uses information on live pricing and promotions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120738,
      "author_name": "sionek",
      "author_url": "",
      "post_date": "05/20/2016 07:22:31",
      "content": "<p>[quote=GreatDataAnalyst;119026]</p>\n\n<p>I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.</p>\n\n<p>I understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.</p>\n\n<p>[/quote]</p>\n\n<p>According to my experiments, using just user_id and the history of  his/her  hotel_cluster booking without any other context, one can obtain MAP@5~0.107. It is very low value, if we compare it with e.g. most frequent benchmark  MAP@5~0.06 or the best script MAP@5~0.50x. However, we should know, that a great part (~1/3) of the results MAP@5~0.50x states almost perfect prediction of data with leakage. The real prediction strength of more or less sophisticated algorithms is in the range 0.2-0.25.  In that context I do not know if we may say that the history of user's booking has low predictive value.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118974": "In the test data set we are provided with the records corresponding to booked data\r\nbut in the real time implementation of this recommendation engine we need to give recommendation to a customer while he is just clicking and browsing AND NOT WHILE HE IS MAKING AN ACTUAL BOOKING. So i wonder how useful will this recommendation engine be?",
    "118977": "Good observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier.",
    "119026": "I am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.\r\n\r\nI understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.",
    "119045": "Well, at the very least you have the clicks that the user made before each booking. And his/her previous bookings...",
    "119403": "[quote=GreatDataAnalyst;119026]\r\n\r\nI am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.\r\n\r\nI understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.\r\n\r\n[/quote]\r\n\r\nuser_id itself MIGHT not be important, but the fact the it reflects the user doing to booking it is obviously an importante thing to pay attention to.\r\nThe  more you know about a specific user the more you can target him/her with specific ads.\r\n\r\nThats why we have all other click history, to make up how the user behavior on the past impacts its current booking.",
    "120338": "[quote=Adam;118977]\r\n\r\nGood observation! It can be useful because during a user visit, and prior to any clicks, we can show hotels that are likely to be booked on top of hotel sort, on expedia home page, etc. We can also include hotel recommendation in emails we send to customers. The other reason why we didn't release clicks from booking sessions in the holdout data is that with this information the prediction problem will become significantly easier. \r\n\r\n\r\n[/quote]\r\n\r\n@Admin: If i understand correctly, is the idea to first predict the cluster, and then the top hotels for that cluster through some in-house algorithm? Also, since you said the realtime implementation would involve recommendations when the user visits, prior to any clicks, what exactly are we trying to achieve in this prediction exercise where we have results from user's clicks and searches like hotel market, search destination id ? Or is it so that the recommendations can be made at any level or the user's browsing process?",
    "120617": "Yes, we would like to make hotel recommendations at any level of the browsing process. The problem is that at the beginning of the \"conversion funnel\" there is usually not much click data we can rely on. \r\n\r\nPredicting hotel cluster can indeed be used as a per-filtering step which will limit the number of hotels; filtered hotels can be then sorted by our in-house algorithm that uses information on live pricing and promotions.",
    "120738": "[quote=GreatDataAnalyst;119026]\r\n\r\nI am failing to understand, how user_id has any predictive power, we do not have any other attribute related to user. How is user_id impacting target, can someone please let me know.\r\n\r\nI understand, it is useful, if we have browsing history of the user, but how is it useful otherwise.\r\n\r\n[/quote]\r\n\r\nAccording to my experiments, using just user_id and the history of  his/her  hotel_cluster booking without any other context, one can obtain MAP@5~0.107. It is very low value, if we compare it with e.g. most frequent benchmark  MAP@5~0.06 or the best script MAP@5~0.50x. However, we should know, that a great part (~1/3) of the results MAP@5~0.50x states almost perfect prediction of data with leakage. The real prediction strength of more or less sophisticated algorithms is in the range 0.2-0.25.  In that context I do not know if we may say that the history of user's booking has low predictive value."
  },
  "source": "meta"
}