{
  "id": 20727,
  "title": "does it make sense to keep only is_booking=1",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20727",
  "author_name": "",
  "post_date": "2016-05-05T07:53:48.257Z",
  "votes": null,
  "comment_count": 8,
  "views": 1489,
  "content": "<p>Hello Kagglers,</p>\n\n<p>I am wondering does it make sense to keep people with is_booking=0 to predict behaviors of those with is_booking=1 in the test set?</p>\n\n<p>Cheers.</p>",
  "messages": [
    {
      "id": "118761",
      "postDate": "05/05/2016 07:53:48",
      "content": "<p>Hello Kagglers,</p>\n\n<p>I am wondering does it make sense to keep people with is_booking=0 to predict behaviors of those with is_booking=1 in the test set?</p>\n\n<p>Cheers.</p>",
      "rawMarkdown": "Hello Kagglers,\r\n\r\nI am wondering does it make sense to keep people with is_booking=0 to predict behaviors of those with is_booking=1 in the test set?\r\n\r\nCheers.",
      "votes": null
    },
    {
      "id": "118769",
      "postDate": "05/05/2016 08:11:11",
      "content": "<p>If clicks are a strong predictor of bookings then the answer is yes</p>\n\n<p>To find out - try a model on clicks and booking and then try a model just on bookings</p>\n\n<p>You might want to do a similar approach on all users and just users in the test set </p>",
      "rawMarkdown": "If clicks are a strong predictor of bookings then the answer is yes\r\n\r\nTo find out - try a model on clicks and booking and then try a model just on bookings\r\n\r\nYou might want to do a similar approach on all users and just users in the test set",
      "votes": null
    },
    {
      "id": "118771",
      "postDate": "05/05/2016 08:25:07",
      "content": "<p>People have tried that only is_booking=1 data give bad results.</p>",
      "rawMarkdown": "People have tried that only is_booking=1 data give bad results.",
      "votes": null
    },
    {
      "id": "118772",
      "postDate": "05/05/2016 08:40:13",
      "content": "<p>Thank you guys. I think it makes sense. You will click and select the hotels that correspond to your need only, not random ones. Just the matter of booking or going to other sites.</p>",
      "rawMarkdown": "Thank you guys. I think it makes sense. You will click and select the hotels that correspond to your need only, not random ones. Just the matter of booking or going to other sites.",
      "votes": null
    },
    {
      "id": "118805",
      "postDate": "05/05/2016 12:45:14",
      "content": "<p>Keep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.</p>",
      "rawMarkdown": "Keep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.",
      "votes": null
    },
    {
      "id": "118811",
      "postDate": "05/05/2016 13:57:52",
      "content": "<p>[quote=Richard Giles;118805]</p>\n\n<p>Keep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.</p>\n\n<p>[/quote]</p>\n\n<p>It is true, I have noticed that. I think the problem would be more challenging to predict the clicks.</p>",
      "rawMarkdown": "[quote=Richard Giles;118805]\r\n\r\nKeep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.\r\n\r\n[/quote]\r\n\r\nIt is true, I have noticed that. I think the problem would be more challenging to predict the clicks.",
      "votes": null
    },
    {
      "id": "119402",
      "postDate": "05/09/2016 23:53:47",
      "content": "<p>Think in this way. In some region, the test set only have data with booking == 0, how to get a better guess then. Of course check the booking == 1 data, and it is at least better than random guess.</p>",
      "rawMarkdown": "Think in this way. In some region, the test set only have data with booking == 0, how to get a better guess then. Of course check the booking == 1 data, and it is at least better than random guess.",
      "votes": null
    },
    {
      "id": "120060",
      "postDate": "05/15/2016 02:38:11",
      "content": "<p>[quote=Li Li;118771]</p>\n\n<p>People have tried that only is_booking=1 data give bad results.</p>\n\n<p>[/quote]</p>\n\n<p>I'm one of those. I tried using only is_booking on training to train the model, and, while doing cross validation, or even train on is_booking_2013  val in is_booking_2014, give some amazing results (reached almost 0.7 in some cases). My LB score (running on all 3M is_booking ==1 train samples) ~ 0.07 :~(</p>\n\n<p>An early model of mine using all training samples bot ~0.19 with 300K training samples.\nOne thing I find funny is that I get such high scores in local CV. One thing that I can imagine is that on the training set there is a huge class imbalance (that my modle might be taking advantage of). While on the test set, the imbalance should be smaller and hence my model sucks on it...</p>\n\n<p>Any more experienced kaggler, that have seen this behavior, and is willing to comment on it is more than welcome.</p>",
      "rawMarkdown": "[quote=Li Li;118771]\r\n\r\nPeople have tried that only is_booking=1 data give bad results.\r\n\r\n[/quote]\r\n\r\nI'm one of those. I tried using only is_booking on training to train the model, and, while doing cross validation, or even train on is_booking_2013  val in is_booking_2014, give some amazing results (reached almost 0.7 in some cases). My LB score (running on all 3M is_booking ==1 train samples) ~ 0.07 :~(\r\n\r\nAn early model of mine using all training samples bot ~0.19 with 300K training samples.\r\nOne thing I find funny is that I get such high scores in local CV. One thing that I can imagine is that on the training set there is a huge class imbalance (that my modle might be taking advantage of). While on the test set, the imbalance should be smaller and hence my model sucks on it...\r\n\r\nAny more experienced kaggler, that have seen this behavior, and is willing to comment on it is more than welcome.",
      "votes": null
    },
    {
      "id": "120067",
      "postDate": "05/15/2016 03:28:36",
      "content": "<p>@Antonio Augusto Santos</p>\n\n<p>What features did you use?</p>\n\n<p>What I did is grouping the data based on destinations, using Random Forest and pick the highest 5 probability location to replace the local top5 guess. I didn't have this kind of issue as long as I didn't use the distance feature, as it is strongly related to data leak and can ruin your model.</p>\n\n<p>However, I cannot train for all destinations, cause there is not enough data. </p>",
      "rawMarkdown": "Antonio Augusto Santos\r\n\r\nWhat features did you use?\r\n\r\nWhat I did is grouping the data based on destinations, using Random Forest and pick the highest 5 probability location to replace the local top5 guess. I didn't have this kind of issue as long as I didn't use the distance feature, as it is strongly related to data leak and can ruin your model.\r\n\r\nHowever, I cannot train for all destinations, cause there is not enough data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118769,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "05/05/2016 08:11:11",
      "content": "<p>If clicks are a strong predictor of bookings then the answer is yes</p>\n\n<p>To find out - try a model on clicks and booking and then try a model just on bookings</p>\n\n<p>You might want to do a similar approach on all users and just users in the test set </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118771,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/05/2016 08:25:07",
      "content": "<p>People have tried that only is_booking=1 data give bad results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118772,
      "author_name": "linhvannguyen",
      "author_url": "",
      "post_date": "05/05/2016 08:40:13",
      "content": "<p>Thank you guys. I think it makes sense. You will click and select the hotels that correspond to your need only, not random ones. Just the matter of booking or going to other sites.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118805,
      "author_name": "rrgiles",
      "author_url": "",
      "post_date": "05/05/2016 12:45:14",
      "content": "<p>Keep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118811,
      "author_name": "linhvannguyen",
      "author_url": "",
      "post_date": "05/05/2016 13:57:52",
      "content": "<p>[quote=Richard Giles;118805]</p>\n\n<p>Keep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.</p>\n\n<p>[/quote]</p>\n\n<p>It is true, I have noticed that. I think the problem would be more challenging to predict the clicks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119402,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "05/09/2016 23:53:47",
      "content": "<p>Think in this way. In some region, the test set only have data with booking == 0, how to get a better guess then. Of course check the booking == 1 data, and it is at least better than random guess.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120060,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "05/15/2016 02:38:11",
      "content": "<p>[quote=Li Li;118771]</p>\n\n<p>People have tried that only is_booking=1 data give bad results.</p>\n\n<p>[/quote]</p>\n\n<p>I'm one of those. I tried using only is_booking on training to train the model, and, while doing cross validation, or even train on is_booking_2013  val in is_booking_2014, give some amazing results (reached almost 0.7 in some cases). My LB score (running on all 3M is_booking ==1 train samples) ~ 0.07 :~(</p>\n\n<p>An early model of mine using all training samples bot ~0.19 with 300K training samples.\nOne thing I find funny is that I get such high scores in local CV. One thing that I can imagine is that on the training set there is a huge class imbalance (that my modle might be taking advantage of). While on the test set, the imbalance should be smaller and hence my model sucks on it...</p>\n\n<p>Any more experienced kaggler, that have seen this behavior, and is willing to comment on it is more than welcome.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120067,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "05/15/2016 03:28:36",
      "content": "<p>@Antonio Augusto Santos</p>\n\n<p>What features did you use?</p>\n\n<p>What I did is grouping the data based on destinations, using Random Forest and pick the highest 5 probability location to replace the local top5 guess. I didn't have this kind of issue as long as I didn't use the distance feature, as it is strongly related to data leak and can ruin your model.</p>\n\n<p>However, I cannot train for all destinations, cause there is not enough data. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118761": "Hello Kagglers,\r\n\r\nI am wondering does it make sense to keep people with is_booking=0 to predict behaviors of those with is_booking=1 in the test set?\r\n\r\nCheers.",
    "118769": "If clicks are a strong predictor of bookings then the answer is yes\r\n\r\nTo find out - try a model on clicks and booking and then try a model just on bookings\r\n\r\nYou might want to do a similar approach on all users and just users in the test set",
    "118771": "People have tried that only is_booking=1 data give bad results.",
    "118772": "Thank you guys. I think it makes sense. You will click and select the hotels that correspond to your need only, not random ones. Just the matter of booking or going to other sites.",
    "118805": "Keep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.",
    "118811": "[quote=Richard Giles;118805]\r\n\r\nKeep in mind that the Admin for the competition has already stated that only bookings are included in the test set. It's up to you to decide whether including only bookings in the training set make sense or not.\r\n\r\n[/quote]\r\n\r\nIt is true, I have noticed that. I think the problem would be more challenging to predict the clicks.",
    "119402": "Think in this way. In some region, the test set only have data with booking == 0, how to get a better guess then. Of course check the booking == 1 data, and it is at least better than random guess.",
    "120060": "[quote=Li Li;118771]\r\n\r\nPeople have tried that only is_booking=1 data give bad results.\r\n\r\n[/quote]\r\n\r\nI'm one of those. I tried using only is_booking on training to train the model, and, while doing cross validation, or even train on is_booking_2013  val in is_booking_2014, give some amazing results (reached almost 0.7 in some cases). My LB score (running on all 3M is_booking ==1 train samples) ~ 0.07 :~(\r\n\r\nAn early model of mine using all training samples bot ~0.19 with 300K training samples.\r\nOne thing I find funny is that I get such high scores in local CV. One thing that I can imagine is that on the training set there is a huge class imbalance (that my modle might be taking advantage of). While on the test set, the imbalance should be smaller and hence my model sucks on it...\r\n\r\nAny more experienced kaggler, that have seen this behavior, and is willing to comment on it is more than welcome.",
    "120067": "Antonio Augusto Santos\r\n\r\nWhat features did you use?\r\n\r\nWhat I did is grouping the data based on destinations, using Random Forest and pick the highest 5 probability location to replace the local top5 guess. I didn't have this kind of issue as long as I didn't use the distance feature, as it is strongly related to data leak and can ruin your model.\r\n\r\nHowever, I cannot train for all destinations, cause there is not enough data."
  },
  "source": "meta"
}