{
  "id": 20281,
  "title": "Clicks and bookings / need some advice",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20281",
  "author_name": "MightyBird",
  "post_date": "2016-04-20T12:12:38.370000",
  "votes": 0,
  "comment_count": 0,
  "views": 433,
  "content": "<p>Hi,</p>\n\n<p>in a first model for memory reasons I used only train data where is_booking == 1. This was a nice small data set but i got stuck at MAP@5 around 0.2x. So I conclude that we need to look at click data as well. Like it has been done in one popular script for example. So I added another column to my train set where I summed up clicks and bookings, weighted.</p>\n\n<p>This looks promising in cross validation. However, we do not have the click information in the test set, only is_booking == 1 (implied), resulting in a constant column - which is of no use.</p>\n\n<p>Does anyone have an idea how to close this gap?</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
  "messages": [
    {
      "id": 115834,
      "postDate": "2016-04-20T12:12:38.370Z",
      "content": "<p>Hi,</p>\n\n<p>in a first model for memory reasons I used only train data where is_booking == 1. This was a nice small data set but i got stuck at MAP@5 around 0.2x. So I conclude that we need to look at click data as well. Like it has been done in one popular script for example. So I added another column to my train set where I summed up clicks and bookings, weighted.</p>\n\n<p>This looks promising in cross validation. However, we do not have the click information in the test set, only is_booking == 1 (implied), resulting in a constant column - which is of no use.</p>\n\n<p>Does anyone have an idea how to close this gap?</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nin a first model for memory reasons I used only train data where is_booking == 1. This was a nice small data set but i got stuck at MAP@5 around 0.2x. So I conclude that we need to look at click data as well. Like it has been done in one popular script for example. So I added another column to my train set where I summed up clicks and bookings, weighted.\r\n\r\nThis looks promising in cross validation. However, we do not have the click information in the test set, only is_booking == 1 (implied), resulting in a constant column - which is of no use.\r\n\r\nDoes anyone have an idea how to close this gap?\r\n\r\nThanks\r\n\r\nGerhard\r\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "115834": "Hi,\r\n\r\nin a first model for memory reasons I used only train data where is_booking == 1. This was a nice small data set but i got stuck at MAP@5 around 0.2x. So I conclude that we need to look at click data as well. Like it has been done in one popular script for example. So I added another column to my train set where I summed up clicks and bookings, weighted.\r\n\r\nThis looks promising in cross validation. However, we do not have the click information in the test set, only is_booking == 1 (implied), resulting in a constant column - which is of no use.\r\n\r\nDoes anyone have an idea how to close this gap?\r\n\r\nThanks\r\n\r\nGerhard\r\n"
  }
}