{
  "id": 21563,
  "title": "How are people choosing coefficients of append_1",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21563",
  "author_name": "",
  "post_date": "2016-06-10T01:29:22.593Z",
  "votes": 1,
  "comment_count": 6,
  "views": 1005,
  "content": "<p>Can anyone explain what procedure are you guys using to choose various coefficients for append_1. At least tell me where did append_1 came from in first place. append_0 makes some sense but append_1 is crazy and people are actually tuning it..</p>",
  "messages": [
    {
      "id": "123159",
      "postDate": "06/10/2016 01:29:22",
      "content": "<p>Can anyone explain what procedure are you guys using to choose various coefficients for append_1. At least tell me where did append_1 came from in first place. append_0 makes some sense but append_1 is crazy and people are actually tuning it..</p>",
      "rawMarkdown": "Can anyone explain what procedure are you guys using to choose various coefficients for append_1. At least tell me where did append_1 came from in first place. append_0 makes some sense but append_1 is crazy and people are actually tuning it..",
      "votes": null
    },
    {
      "id": "123188",
      "postDate": "06/10/2016 06:49:15",
      "content": "<p>I programmed a grid search. I used either 2013 for training and 2014 for validation, or, 12 months of 2013 and 2014 for training and the final three months of 2014 for validation. You would think that the validation should be booking only, but when I did that, I always got coefficients highly skewed towards booking. And, that did not perform well on the leader board. A self-selected 4/16 to start, then noticed that 3/17 was used frequently. I adopted that briefly, figuring it was the wisdom of the masses, but I can assure you that my coefficients are nowhere near that now. I then built in power and log based recency multipliers and addedt those to my grid search. Hope this helps a little bit.</p>",
      "rawMarkdown": "I programmed a grid search. I used either 2013 for training and 2014 for validation, or, 12 months of 2013 and 2014 for training and the final three months of 2014 for validation. You would think that the validation should be booking only, but when I did that, I always got coefficients highly skewed towards booking. And, that did not perform well on the leader board. A self-selected 4/16 to start, then noticed that 3/17 was used frequently. I adopted that briefly, figuring it was the wisdom of the masses, but I can assure you that my coefficients are nowhere near that now. I then built in power and log based recency multipliers and addedt those to my grid search. Hope this helps a little bit.",
      "votes": null
    },
    {
      "id": "123214",
      "postDate": "06/10/2016 11:37:16",
      "content": "<p>I was unable to detect much advantage to tweaking these values, the features that I grouped by was much more influential. But a grid search would have been nice - and especially one per feature combination that I grouped on. But that would have taken forever.</p>",
      "rawMarkdown": "I was unable to detect much advantage to tweaking these values, the features that I grouped by was much more influential. But a grid search would have been nice - and especially one per feature combination that I grouped on. But that would have taken forever.",
      "votes": null
    },
    {
      "id": "123257",
      "postDate": "06/10/2016 17:48:43",
      "content": "<p>It smells of overfitting the public LB to me...</p>",
      "rawMarkdown": "It smells of overfitting the public LB to me...",
      "votes": null
    },
    {
      "id": "123269",
      "postDate": "06/10/2016 18:39:55",
      "content": "<p>I fiddled with CI-DT=advance booking time and CO-CI=duration for quite a bit, figuring that there had to be something there.  I use those, but the grid search on them was pretty wild.  Month recency, using log or power functions, and package and childbin have always been front and center for me.  site_name I didn't tune; is_mobile never gave me reliable signal.  I was doing grid search on 2013 and 2014 data, so it doesn't strike me as overfitting the LB, per happycube's comment.  We shall see this evening!</p>",
      "rawMarkdown": "I fiddled with CI-DT=advance booking time and CO-CI=duration for quite a bit, figuring that there had to be something there.  I use those, but the grid search on them was pretty wild.  Month recency, using log or power functions, and package and childbin have always been front and center for me.  site_name I didn't tune; is_mobile never gave me reliable signal.  I was doing grid search on 2013 and 2014 data, so it doesn't strike me as overfitting the LB, per happycube's comment.  We shall see this evening!",
      "votes": null
    },
    {
      "id": "123296",
      "postDate": "06/10/2016 22:37:26",
      "content": "<p>And has anyone tried booking_month or checkin_month along with each hotel. According to Admin, hotels change cluster based on season. So there must be some information there. I wasn't able to check it out because of low ram. Even running script on kaggle gave memory limit error. </p>\n\n<p>I think all top scripts must be using some kind of seasonal information to differentiate between same hotels in different cluster.</p>",
      "rawMarkdown": "And has anyone tried booking_month or checkin_month along with each hotel. According to Admin, hotels change cluster based on season. So there must be some information there. I wasn't able to check it out because of low ram. Even running script on kaggle gave memory limit error. \r\n\r\nI think all top scripts must be using some kind of seasonal information to differentiate between same hotels in different cluster.",
      "votes": null
    },
    {
      "id": "123486",
      "postDate": "06/12/2016 07:17:52",
      "content": "<p>Yes, I used month and even season. I'll post something.</p>",
      "rawMarkdown": "Yes, I used month and even season. I'll post something.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 123188,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 06:49:15",
      "content": "<p>I programmed a grid search. I used either 2013 for training and 2014 for validation, or, 12 months of 2013 and 2014 for training and the final three months of 2014 for validation. You would think that the validation should be booking only, but when I did that, I always got coefficients highly skewed towards booking. And, that did not perform well on the leader board. A self-selected 4/16 to start, then noticed that 3/17 was used frequently. I adopted that briefly, figuring it was the wisdom of the masses, but I can assure you that my coefficients are nowhere near that now. I then built in power and log based recency multipliers and addedt those to my grid search. Hope this helps a little bit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123214,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 11:37:16",
      "content": "<p>I was unable to detect much advantage to tweaking these values, the features that I grouped by was much more influential. But a grid search would have been nice - and especially one per feature combination that I grouped on. But that would have taken forever.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123257,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "06/10/2016 17:48:43",
      "content": "<p>It smells of overfitting the public LB to me...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123269,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 18:39:55",
      "content": "<p>I fiddled with CI-DT=advance booking time and CO-CI=duration for quite a bit, figuring that there had to be something there.  I use those, but the grid search on them was pretty wild.  Month recency, using log or power functions, and package and childbin have always been front and center for me.  site_name I didn't tune; is_mobile never gave me reliable signal.  I was doing grid search on 2013 and 2014 data, so it doesn't strike me as overfitting the LB, per happycube's comment.  We shall see this evening!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123296,
      "author_name": "shahnawazakhtar",
      "author_url": "",
      "post_date": "06/10/2016 22:37:26",
      "content": "<p>And has anyone tried booking_month or checkin_month along with each hotel. According to Admin, hotels change cluster based on season. So there must be some information there. I wasn't able to check it out because of low ram. Even running script on kaggle gave memory limit error. </p>\n\n<p>I think all top scripts must be using some kind of seasonal information to differentiate between same hotels in different cluster.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123486,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/12/2016 07:17:52",
      "content": "<p>Yes, I used month and even season. I'll post something.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "123159": "Can anyone explain what procedure are you guys using to choose various coefficients for append_1. At least tell me where did append_1 came from in first place. append_0 makes some sense but append_1 is crazy and people are actually tuning it..",
    "123188": "I programmed a grid search. I used either 2013 for training and 2014 for validation, or, 12 months of 2013 and 2014 for training and the final three months of 2014 for validation. You would think that the validation should be booking only, but when I did that, I always got coefficients highly skewed towards booking. And, that did not perform well on the leader board. A self-selected 4/16 to start, then noticed that 3/17 was used frequently. I adopted that briefly, figuring it was the wisdom of the masses, but I can assure you that my coefficients are nowhere near that now. I then built in power and log based recency multipliers and addedt those to my grid search. Hope this helps a little bit.",
    "123214": "I was unable to detect much advantage to tweaking these values, the features that I grouped by was much more influential. But a grid search would have been nice - and especially one per feature combination that I grouped on. But that would have taken forever.",
    "123257": "It smells of overfitting the public LB to me...",
    "123269": "I fiddled with CI-DT=advance booking time and CO-CI=duration for quite a bit, figuring that there had to be something there.  I use those, but the grid search on them was pretty wild.  Month recency, using log or power functions, and package and childbin have always been front and center for me.  site_name I didn't tune; is_mobile never gave me reliable signal.  I was doing grid search on 2013 and 2014 data, so it doesn't strike me as overfitting the LB, per happycube's comment.  We shall see this evening!",
    "123296": "And has anyone tried booking_month or checkin_month along with each hotel. According to Admin, hotels change cluster based on season. So there must be some information there. I wasn't able to check it out because of low ram. Even running script on kaggle gave memory limit error. \r\n\r\nI think all top scripts must be using some kind of seasonal information to differentiate between same hotels in different cluster.",
    "123486": "Yes, I used month and even season. I'll post something."
  },
  "source": "meta"
}