{
  "id": 21055,
  "title": "Data wide tends vs. user preferences",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21055",
  "author_name": "",
  "post_date": "2016-05-19T00:56:36.597Z",
  "votes": 1,
  "comment_count": 8,
  "views": 1180,
  "content": "<p>As rightly noted by <a href=\"https://www.kaggle.com/zfturbo/expedia-hotel-recommendations/leakage-solution\">ZFTurbo in the data leak solution</a>, inclusion of user specific trends would most likely enhance results. Actually, I strongly believe the winning solution will be the one which thoughtfully combines data wide and user specific trends. That being said, I myself haven't really made progress in systematically incorporating personal preferences and so I'm wondering if anyone has tried this approach and are willing to share some pointers. </p>",
  "messages": [
    {
      "id": "120554",
      "postDate": "05/19/2016 00:56:36",
      "content": "<p>As rightly noted by <a href=\"https://www.kaggle.com/zfturbo/expedia-hotel-recommendations/leakage-solution\">ZFTurbo in the data leak solution</a>, inclusion of user specific trends would most likely enhance results. Actually, I strongly believe the winning solution will be the one which thoughtfully combines data wide and user specific trends. That being said, I myself haven't really made progress in systematically incorporating personal preferences and so I'm wondering if anyone has tried this approach and are willing to share some pointers. </p>",
      "rawMarkdown": "As rightly noted by [ZFTurbo in the data leak solution][1], inclusion of user specific trends would most likely enhance results. Actually, I strongly believe the winning solution will be the one which thoughtfully combines data wide and user specific trends. That being said, I myself haven't really made progress in systematically incorporating personal preferences and so I'm wondering if anyone has tried this approach and are willing to share some pointers. \r\n\r\n  [1]: https://www.kaggle.com/zfturbo/expedia-hotel-recommendations/leakage-solution",
      "votes": null
    },
    {
      "id": "120729",
      "postDate": "05/20/2016 05:35:47",
      "content": "<p>I have a 5 tier approach, and these are 2 of them in my modeling.</p>",
      "rawMarkdown": "I have a 5 tier approach, and these are 2 of them in my modeling.",
      "votes": null
    },
    {
      "id": "120751",
      "postDate": "05/20/2016 09:49:15",
      "content": "<p>@eipiplus1, no wonder you have a high score, congratulations.  Does any of your tiers include some classification algorithm(e.g., xgboost)  or its all frequency based?</p>",
      "rawMarkdown": "eipiplus1, no wonder you have a high score, congratulations.  Does any of your tiers include some classification algorithm(e.g., xgboost)  or its all frequency based?",
      "votes": null
    },
    {
      "id": "120859",
      "postDate": "05/21/2016 02:01:00",
      "content": "<p>I've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.</p>",
      "rawMarkdown": "I've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.",
      "votes": null
    },
    {
      "id": "120934",
      "postDate": "05/21/2016 19:43:51",
      "content": "<p>I'm curious about this topic too. I'm new to Kaggle (and a lot of other things), but when I looked at the contest description, I leaned more toward a user-focused approach, weighted by global trends. I noticed that the forums seemed to go the other way, but since this is a learning experience, I've been pursuing the user approach.  I'm aware of the data leak, but have thus far ignored it because it seems like a known component. </p>\n\n<p>I'd welcome feedback from others on what i have done... it hasn't worked very well.. particularly appreciate comments about why it's done much better in testing on the training set than on the actual test set. </p>\n\n<p>My original thoughts were that since the hotel clusters were int'l concept, not tied to locations, i hypothesized that users might be reasonably consistent in their preferences. Of course, they might vary for any number of reasons (vacation vs work vs advance planning, etc.), but thought the user history was important. </p>\n\n<p>thought of lots of fancy things to do and then got hit with the reality of the data size.  decided to go for the simplest user-based approach i could think of, which was to take the most common user-clicks. for this first experiment, i used clicks instead of books for a range of reasons. then, i randomly sampled the books in the training data for testing purposes.  this does mean that my predictions are constant for each user, independent of all other factors, and that it might not predict five things. (based on the instructions, it seems fine to list less than five predictions) but it seemed like a start. </p>\n\n<p>anyhow, on my laptop, with a wide range of random sampling, this gave me a score of ~0.35 consistently. however, when i created predictions for the full test dataset, the leaderboard score is <em>way</em> lower, at ~0.08. </p>\n\n<p>i'd love any thoughts about why there would be such a large difference (I would expect some, but that seems very large diff).... </p>\n\n<p>and also any thoughts as to why the user-focused might not be effective... it seems, as you say, a mix is best. </p>",
      "rawMarkdown": "I'm curious about this topic too. I'm new to Kaggle (and a lot of other things), but when I looked at the contest description, I leaned more toward a user-focused approach, weighted by global trends. I noticed that the forums seemed to go the other way, but since this is a learning experience, I've been pursuing the user approach.  I'm aware of the data leak, but have thus far ignored it because it seems like a known component. \r\n\r\nI'd welcome feedback from others on what i have done... it hasn't worked very well.. particularly appreciate comments about why it's done much better in testing on the training set than on the actual test set. \r\n\r\nMy original thoughts were that since the hotel clusters were int'l concept, not tied to locations, i hypothesized that users might be reasonably consistent in their preferences. Of course, they might vary for any number of reasons (vacation vs work vs advance planning, etc.), but thought the user history was important. \r\n\r\nthought of lots of fancy things to do and then got hit with the reality of the data size.  decided to go for the simplest user-based approach i could think of, which was to take the most common user-clicks. for this first experiment, i used clicks instead of books for a range of reasons. then, i randomly sampled the books in the training data for testing purposes.  this does mean that my predictions are constant for each user, independent of all other factors, and that it might not predict five things. (based on the instructions, it seems fine to list less than five predictions) but it seemed like a start. \r\n\r\nanyhow, on my laptop, with a wide range of random sampling, this gave me a score of ~0.35 consistently. however, when i created predictions for the full test dataset, the leaderboard score is *way* lower, at ~0.08. \r\n\r\ni'd love any thoughts about why there would be such a large difference (I would expect some, but that seems very large diff).... \r\n\r\nand also any thoughts as to why the user-focused might not be effective... it seems, as you say, a mix is best.",
      "votes": null
    },
    {
      "id": "120935",
      "postDate": "05/21/2016 19:51:19",
      "content": "<p>One thing about clicks is, though the user may have been shown the HC as indicated, the combination of search vars, say, (CI - DT, nbr_adults, nbr_children) may possibly be proven to be impossible to book for the given HC.  That is, the user is looking for a room +1 days later, for 2 adults and 2 children, and the HC=X is listed in the data.  Then for (1,2,2,X) you may find 0 booking ever exist.  So in that case, is the click data point worthwhile?</p>",
      "rawMarkdown": "One thing about clicks is, though the user may have been shown the HC as indicated, the combination of search vars, say, (CI - DT, nbr_adults, nbr_children) may possibly be proven to be impossible to book for the given HC.  That is, the user is looking for a room +1 days later, for 2 adults and 2 children, and the HC=X is listed in the data.  Then for (1,2,2,X) you may find 0 booking ever exist.  So in that case, is the click data point worthwhile?",
      "votes": null
    },
    {
      "id": "120942",
      "postDate": "05/21/2016 21:51:47",
      "content": "<p>Also - and in addition to what @elpiplus1 stated - click data is usually much less indicative of a conversion (booking an hotel) than previous conversions from the same user. For example, you might click on a search for a variety of reasons (checking prices, look for availability of dates, planning trips with your partner ...), whereas you only generate booking data for a reason - because you like that kind of hotel.</p>\n\n<p>I had similar problems with a problem in a company I was working at some time ago, and at the the end the best solution turned out to be something somewhat close to what the Public Scripts are doing in this competition: Click data is useful, but previous conversions (and other signals) should have more weight in the final models. But still, it is not a bad thing to use, specially for when you lack of other data....</p>",
      "rawMarkdown": "Also - and in addition to what @elpiplus1 stated - click data is usually much less indicative of a conversion (booking an hotel) than previous conversions from the same user. For example, you might click on a search for a variety of reasons (checking prices, look for availability of dates, planning trips with your partner ...), whereas you only generate booking data for a reason - because you like that kind of hotel.\r\n\r\nI had similar problems with a problem in a company I was working at some time ago, and at the the end the best solution turned out to be something somewhat close to what the Public Scripts are doing in this competition: Click data is useful, but previous conversions (and other signals) should have more weight in the final models. But still, it is not a bad thing to use, specially for when you lack of other data....",
      "votes": null
    },
    {
      "id": "120981",
      "postDate": "05/22/2016 11:43:11",
      "content": "<p>@eipiplus1 and @carrdelling, thanks for the replies. </p>\n\n<p>completely agree in the weaknesses of click information. i hadn't expected much from it alone, but was more surprised by the difference in local testing versus the real test data. I think i've figured that out, as i had been randomly sampling across all the data, and when i instead sharply divide by time, the cv score also goes way down. </p>\n\n<p>while the weakness of click information is intuitive to me, the strength of global trends to an individual's choice is less so to me. though from my reckoning, about 30% of the test users have no history of booking (only clicks), so it makes sense that you'll have to leverage that global data... and maybe it weighs much more heavily than i'd expect. </p>\n\n<p>the insight that whether a given combination of parameters is <em>ever</em> associated with a booked cluster is interesting, eipiplus1... i'll look at that.  so that independent of user preferences, only certain clusters might be available in given locations and search parameter combinations.  i did a bunch of data exploration first, but not that specifically. thanks. </p>",
      "rawMarkdown": "eipiplus1 and @carrdelling, thanks for the replies. \r\n\r\ncompletely agree in the weaknesses of click information. i hadn't expected much from it alone, but was more surprised by the difference in local testing versus the real test data. I think i've figured that out, as i had been randomly sampling across all the data, and when i instead sharply divide by time, the cv score also goes way down. \r\n\r\nwhile the weakness of click information is intuitive to me, the strength of global trends to an individual's choice is less so to me. though from my reckoning, about 30% of the test users have no history of booking (only clicks), so it makes sense that you'll have to leverage that global data... and maybe it weighs much more heavily than i'd expect. \r\n\r\nthe insight that whether a given combination of parameters is _ever_ associated with a booked cluster is interesting, eipiplus1... i'll look at that.  so that independent of user preferences, only certain clusters might be available in given locations and search parameter combinations.  i did a bunch of data exploration first, but not that specifically. thanks.",
      "votes": null
    },
    {
      "id": "121022",
      "postDate": "05/23/2016 01:23:37",
      "content": "<p>[quote=eipiplus1;120859]</p>\n\n<p>I've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for these insights, I pretty much had/have a similar idea,namely to focus on the non-leak events but I haven't implemented it yet, but its encouraging to know that I will get positive results once I get to do it. </p>",
      "rawMarkdown": "[quote=eipiplus1;120859]\r\n\r\nI've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.\r\n\r\n[/quote]\r\n\r\nThanks for these insights, I pretty much had/have a similar idea,namely to focus on the non-leak events but I haven't implemented it yet, but its encouraging to know that I will get positive results once I get to do it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 120729,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "05/20/2016 05:35:47",
      "content": "<p>I have a 5 tier approach, and these are 2 of them in my modeling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120751,
      "author_name": "dmatekenya",
      "author_url": "",
      "post_date": "05/20/2016 09:49:15",
      "content": "<p>@eipiplus1, no wonder you have a high score, congratulations.  Does any of your tiers include some classification algorithm(e.g., xgboost)  or its all frequency based?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120859,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "05/21/2016 02:01:00",
      "content": "<p>I've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120934,
      "author_name": "knitcode",
      "author_url": "",
      "post_date": "05/21/2016 19:43:51",
      "content": "<p>I'm curious about this topic too. I'm new to Kaggle (and a lot of other things), but when I looked at the contest description, I leaned more toward a user-focused approach, weighted by global trends. I noticed that the forums seemed to go the other way, but since this is a learning experience, I've been pursuing the user approach.  I'm aware of the data leak, but have thus far ignored it because it seems like a known component. </p>\n\n<p>I'd welcome feedback from others on what i have done... it hasn't worked very well.. particularly appreciate comments about why it's done much better in testing on the training set than on the actual test set. </p>\n\n<p>My original thoughts were that since the hotel clusters were int'l concept, not tied to locations, i hypothesized that users might be reasonably consistent in their preferences. Of course, they might vary for any number of reasons (vacation vs work vs advance planning, etc.), but thought the user history was important. </p>\n\n<p>thought of lots of fancy things to do and then got hit with the reality of the data size.  decided to go for the simplest user-based approach i could think of, which was to take the most common user-clicks. for this first experiment, i used clicks instead of books for a range of reasons. then, i randomly sampled the books in the training data for testing purposes.  this does mean that my predictions are constant for each user, independent of all other factors, and that it might not predict five things. (based on the instructions, it seems fine to list less than five predictions) but it seemed like a start. </p>\n\n<p>anyhow, on my laptop, with a wide range of random sampling, this gave me a score of ~0.35 consistently. however, when i created predictions for the full test dataset, the leaderboard score is <em>way</em> lower, at ~0.08. </p>\n\n<p>i'd love any thoughts about why there would be such a large difference (I would expect some, but that seems very large diff).... </p>\n\n<p>and also any thoughts as to why the user-focused might not be effective... it seems, as you say, a mix is best. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120935,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "05/21/2016 19:51:19",
      "content": "<p>One thing about clicks is, though the user may have been shown the HC as indicated, the combination of search vars, say, (CI - DT, nbr_adults, nbr_children) may possibly be proven to be impossible to book for the given HC.  That is, the user is looking for a room +1 days later, for 2 adults and 2 children, and the HC=X is listed in the data.  Then for (1,2,2,X) you may find 0 booking ever exist.  So in that case, is the click data point worthwhile?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120942,
      "author_name": "carrdelling",
      "author_url": "",
      "post_date": "05/21/2016 21:51:47",
      "content": "<p>Also - and in addition to what @elpiplus1 stated - click data is usually much less indicative of a conversion (booking an hotel) than previous conversions from the same user. For example, you might click on a search for a variety of reasons (checking prices, look for availability of dates, planning trips with your partner ...), whereas you only generate booking data for a reason - because you like that kind of hotel.</p>\n\n<p>I had similar problems with a problem in a company I was working at some time ago, and at the the end the best solution turned out to be something somewhat close to what the Public Scripts are doing in this competition: Click data is useful, but previous conversions (and other signals) should have more weight in the final models. But still, it is not a bad thing to use, specially for when you lack of other data....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120981,
      "author_name": "knitcode",
      "author_url": "",
      "post_date": "05/22/2016 11:43:11",
      "content": "<p>@eipiplus1 and @carrdelling, thanks for the replies. </p>\n\n<p>completely agree in the weaknesses of click information. i hadn't expected much from it alone, but was more surprised by the difference in local testing versus the real test data. I think i've figured that out, as i had been randomly sampling across all the data, and when i instead sharply divide by time, the cv score also goes way down. </p>\n\n<p>while the weakness of click information is intuitive to me, the strength of global trends to an individual's choice is less so to me. though from my reckoning, about 30% of the test users have no history of booking (only clicks), so it makes sense that you'll have to leverage that global data... and maybe it weighs much more heavily than i'd expect. </p>\n\n<p>the insight that whether a given combination of parameters is <em>ever</em> associated with a booked cluster is interesting, eipiplus1... i'll look at that.  so that independent of user preferences, only certain clusters might be available in given locations and search parameter combinations.  i did a bunch of data exploration first, but not that specifically. thanks. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121022,
      "author_name": "dmatekenya",
      "author_url": "",
      "post_date": "05/23/2016 01:23:37",
      "content": "<p>[quote=eipiplus1;120859]</p>\n\n<p>I've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for these insights, I pretty much had/have a similar idea,namely to focus on the non-leak events but I haven't implemented it yet, but its encouraging to know that I will get positive results once I get to do it. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "120554": "As rightly noted by [ZFTurbo in the data leak solution][1], inclusion of user specific trends would most likely enhance results. Actually, I strongly believe the winning solution will be the one which thoughtfully combines data wide and user specific trends. That being said, I myself haven't really made progress in systematically incorporating personal preferences and so I'm wondering if anyone has tried this approach and are willing to share some pointers. \r\n\r\n  [1]: https://www.kaggle.com/zfturbo/expedia-hotel-recommendations/leakage-solution",
    "120729": "I have a 5 tier approach, and these are 2 of them in my modeling.",
    "120751": "eipiplus1, no wonder you have a high score, congratulations.  Does any of your tiers include some classification algorithm(e.g., xgboost)  or its all frequency based?",
    "120859": "I've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.",
    "120934": "I'm curious about this topic too. I'm new to Kaggle (and a lot of other things), but when I looked at the contest description, I leaned more toward a user-focused approach, weighted by global trends. I noticed that the forums seemed to go the other way, but since this is a learning experience, I've been pursuing the user approach.  I'm aware of the data leak, but have thus far ignored it because it seems like a known component. \r\n\r\nI'd welcome feedback from others on what i have done... it hasn't worked very well.. particularly appreciate comments about why it's done much better in testing on the training set than on the actual test set. \r\n\r\nMy original thoughts were that since the hotel clusters were int'l concept, not tied to locations, i hypothesized that users might be reasonably consistent in their preferences. Of course, they might vary for any number of reasons (vacation vs work vs advance planning, etc.), but thought the user history was important. \r\n\r\nthought of lots of fancy things to do and then got hit with the reality of the data size.  decided to go for the simplest user-based approach i could think of, which was to take the most common user-clicks. for this first experiment, i used clicks instead of books for a range of reasons. then, i randomly sampled the books in the training data for testing purposes.  this does mean that my predictions are constant for each user, independent of all other factors, and that it might not predict five things. (based on the instructions, it seems fine to list less than five predictions) but it seemed like a start. \r\n\r\nanyhow, on my laptop, with a wide range of random sampling, this gave me a score of ~0.35 consistently. however, when i created predictions for the full test dataset, the leaderboard score is *way* lower, at ~0.08. \r\n\r\ni'd love any thoughts about why there would be such a large difference (I would expect some, but that seems very large diff).... \r\n\r\nand also any thoughts as to why the user-focused might not be effective... it seems, as you say, a mix is best.",
    "120935": "One thing about clicks is, though the user may have been shown the HC as indicated, the combination of search vars, say, (CI - DT, nbr_adults, nbr_children) may possibly be proven to be impossible to book for the given HC.  That is, the user is looking for a room +1 days later, for 2 adults and 2 children, and the HC=X is listed in the data.  Then for (1,2,2,X) you may find 0 booking ever exist.  So in that case, is the click data point worthwhile?",
    "120942": "Also - and in addition to what @elpiplus1 stated - click data is usually much less indicative of a conversion (booking an hotel) than previous conversions from the same user. For example, you might click on a search for a variety of reasons (checking prices, look for availability of dates, planning trips with your partner ...), whereas you only generate booking data for a reason - because you like that kind of hotel.\r\n\r\nI had similar problems with a problem in a company I was working at some time ago, and at the the end the best solution turned out to be something somewhat close to what the Public Scripts are doing in this competition: Click data is useful, but previous conversions (and other signals) should have more weight in the final models. But still, it is not a bad thing to use, specially for when you lack of other data....",
    "120981": "eipiplus1 and @carrdelling, thanks for the replies. \r\n\r\ncompletely agree in the weaknesses of click information. i hadn't expected much from it alone, but was more surprised by the difference in local testing versus the real test data. I think i've figured that out, as i had been randomly sampling across all the data, and when i instead sharply divide by time, the cv score also goes way down. \r\n\r\nwhile the weakness of click information is intuitive to me, the strength of global trends to an individual's choice is less so to me. though from my reckoning, about 30% of the test users have no history of booking (only clicks), so it makes sense that you'll have to leverage that global data... and maybe it weighs much more heavily than i'd expect. \r\n\r\nthe insight that whether a given combination of parameters is _ever_ associated with a booked cluster is interesting, eipiplus1... i'll look at that.  so that independent of user preferences, only certain clusters might be available in given locations and search parameter combinations.  i did a bunch of data exploration first, but not that specifically. thanks.",
    "121022": "[quote=eipiplus1;120859]\r\n\r\nI've tried Naive Bayes and xgboost and ranger and have settled on some combination of those.  I use perl to do mass scoring for, say, the test file.  I'll point out one tier to consider - the final ~100K test rows which are not properly addressed by leak or any of your other models.  I've done a little bit of work on those final straggler rows, and I anticipate in the final week I'll really hunker down on them to eke out the final improvements that I can.  For now, there are bigger gains in the 2/3 non-leak data.\r\n\r\n[/quote]\r\n\r\nThanks for these insights, I pretty much had/have a similar idea,namely to focus on the non-leak events but I haven't implemented it yet, but its encouraging to know that I will get positive results once I get to do it."
  },
  "source": "meta"
}