{
  "id": 21588,
  "title": "Any interesting features created?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21588",
  "author_name": "",
  "post_date": "2016-06-11T07:12:46.883Z",
  "votes": 4,
  "comment_count": 23,
  "views": 2733,
  "content": "<p>I am curious about what features did all those who rank 30 or higher created.</p>\n\n<p>Mine is pretty simple: use a pivot table to find out all the clusters a user have clicked, and merge the pivot table with both the training set and testing set. So for each click, we can find a history of what cluster the user have viewed.</p>\n\n<p>Then I use XGB to train for each destination with data points more than 500 (? I forget the exact number). And for the rest destinations, I use the local popular model.</p>\n\n<p>This feature improve my prediction for about 0.005 and I guess it's somehow like collaborative filtering. Of course not exactly the same, but the idea is to conclude from what cluster would an other similar person have clicked.</p>\n\n<p>So, what interesting features did you created? And how do you come up with it?</p>",
  "messages": [
    {
      "id": "123345",
      "postDate": "06/11/2016 07:12:46",
      "content": "<p>I am curious about what features did all those who rank 30 or higher created.</p>\n\n<p>Mine is pretty simple: use a pivot table to find out all the clusters a user have clicked, and merge the pivot table with both the training set and testing set. So for each click, we can find a history of what cluster the user have viewed.</p>\n\n<p>Then I use XGB to train for each destination with data points more than 500 (? I forget the exact number). And for the rest destinations, I use the local popular model.</p>\n\n<p>This feature improve my prediction for about 0.005 and I guess it's somehow like collaborative filtering. Of course not exactly the same, but the idea is to conclude from what cluster would an other similar person have clicked.</p>\n\n<p>So, what interesting features did you created? And how do you come up with it?</p>",
      "rawMarkdown": "I am curious about what features did all those who rank 30 or higher created.\r\n\r\nMine is pretty simple: use a pivot table to find out all the clusters a user have clicked, and merge the pivot table with both the training set and testing set. So for each click, we can find a history of what cluster the user have viewed.\r\n\r\nThen I use XGB to train for each destination with data points more than 500 (? I forget the exact number). And for the rest destinations, I use the local popular model.\r\n\r\nThis feature improve my prediction for about 0.005 and I guess it's somehow like collaborative filtering. Of course not exactly the same, but the idea is to conclude from what cluster would an other similar person have clicked.\r\n\r\nSo, what interesting features did you created? And how do you come up with it?",
      "votes": null
    },
    {
      "id": "123357",
      "postDate": "06/11/2016 09:18:55",
      "content": "<p>Interesting question!</p>\n\n<p>I counted bookings/clicks in each cluster for all levels of most variables to create features - I guess that is a lot like your pivot - and then:</p>\n\n<ul>\n<li>have the last 4 months of 2014 as a holdout set that is not used for counting bookings/clicks </li>\n<li>from the holdout set, create a dataset with 100 rows for each booking: one for every cluster, with the cluster specific count for all variables as the features, and a 0/1 (booked) target</li>\n<li>scale the counts of each feature to sum to 1 over all clusters</li>\n<li>train an xgboost model (maxdepth 12!) on the holdout set - including only 10% of 0 targets</li>\n<li>build a testset like the holdout set, using counts over all months including the holdout set (note that this step makes the scaling necessary) and apply the model to the resulting 250,000,000 rows</li>\n</ul>\n\n<p>Apart from the leak variable (city || distance), I also built some interactions with hotel market and a 'user type' variable (usercountry || adultcount || childcount || hotelmarket) that seems to help the model a bit.</p>\n\n<p>During the last days, inspired by the unbelievable score of idle_speculation, I tried to exploit the leak further, by:</p>\n\n<ul>\n<li>identifying hotels that were <em>not</em> booked because the distance did <em>not</em> match the leak feature</li>\n<li>'widening the leak' for cities on a straight line from a destination, as suggested by Mattias</li>\n</ul>",
      "rawMarkdown": "Interesting question!\r\n\r\nI counted bookings/clicks in each cluster for all levels of most variables to create features - I guess that is a lot like your pivot - and then:\r\n\r\n - have the last 4 months of 2014 as a holdout set that is not used for counting bookings/clicks \r\n - from the holdout set, create a dataset with 100 rows for each booking: one for every cluster, with the cluster specific count for all variables as the features, and a 0/1 (booked) target\r\n - scale the counts of each feature to sum to 1 over all clusters\r\n - train an xgboost model (maxdepth 12!) on the holdout set - including only 10% of 0 targets\r\n - build a testset like the holdout set, using counts over all months including the holdout set (note that this step makes the scaling necessary) and apply the model to the resulting 250,000,000 rows\r\n\r\nApart from the leak variable (city || distance), I also built some interactions with hotel market and a 'user type' variable (usercountry || adultcount || childcount || hotelmarket) that seems to help the model a bit.\r\n\r\nDuring the last days, inspired by the unbelievable score of idle_speculation, I tried to exploit the leak further, by:\r\n \r\n - identifying hotels that were *not* booked because the distance did *not* match the leak feature\r\n - 'widening the leak' for cities on a straight line from a destination, as suggested by Mattias",
      "votes": null
    },
    {
      "id": "123366",
      "postDate": "06/11/2016 10:16:00",
      "content": "<p>While not in the top30, I made more or less rule based prediction. I was thinking 100 classes is too many to predict with ml, but in the last two weeks realized this must be wrong. </p>\n\n<ul>\n<li>One neat thing was to add the leaked test rows (particularly\nrows where only one hotel_cluster was found in the leak) on top of the training\nset before starting the modelling. This worked very well using 2013 as train, and 2014 as test, but\nnot so well on the LB (about .002). I guess this was because the train and\ntest size is less imbalanced there. </li>\n<li>using bayes mean to weight book vs click per a few different features, as a\nopposed to global book vs click rate </li>\n<li>ranking with layers of weights for different features, bins of stay time, booking lead time, season... </li>\n</ul>\n\n<p>Script is <a href=\"https://www.dropbox.com/s/tgs8cjx66ikcebf/sub-bmleak-1006-0.50745.R?dl=1\">here</a> if its of use to anyone. I won't list out the things that did not work as I guess Kaggle have set some limitation on how long a single comment can be :) </p>",
      "rawMarkdown": "While not in the top30, I made more or less rule based prediction. I was thinking 100 classes is too many to predict with ml, but in the last two weeks realized this must be wrong. \r\n\r\n - One neat thing was to add the leaked test rows (particularly\r\n   rows where only one hotel_cluster was found in the leak) on top of the training\r\n   set before starting the modelling. This worked very well using 2013 as train, and 2014 as test, but\r\n   not so well on the LB (about .002). I guess this was because the train and\r\n   test size is less imbalanced there. \r\n - using bayes mean to weight book vs click per a few different features, as a\r\n   opposed to global book vs click rate \r\n - ranking with layers of weights for different features, bins of stay time, booking lead time, season... \r\n\r\nScript is [here][1] if its of use to anyone. I won't list out the things that did not work as I guess Kaggle have set some limitation on how long a single comment can be :) \r\n\r\n\r\n  [1]: https://www.dropbox.com/s/tgs8cjx66ikcebf/sub-bmleak-1006-0.50745.R?dl=1",
      "votes": null
    },
    {
      "id": "123399",
      "postDate": "06/11/2016 16:53:41",
      "content": "<p>@YijunMa Could you elaborate your pivot table ? Is it user_id based? The data pints larger than 500 means there're at least 500 points each cluster? Thank you.</p>",
      "rawMarkdown": "YijunMa Could you elaborate your pivot table ? Is it user_id based? The data pints larger than 500 means there're at least 500 points each cluster? Thank you.",
      "votes": null
    },
    {
      "id": "123401",
      "postDate": "06/11/2016 17:12:06",
      "content": "<p>@Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .</p>",
      "rawMarkdown": "Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .",
      "votes": null
    },
    {
      "id": "123403",
      "postDate": "06/11/2016 17:49:29",
      "content": "<p>[quote=FengLi;123401]</p>\n\n<p>@Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .</p>\n\n<p>[/quote]</p>\n\n<p>around 0.518; &quot;identifying hotels that were not booked&quot; improved my score by about 0.001 and &quot;widening the leak&quot; gave my last minute improvement from 0.519 to 0.521 (public LB)</p>",
      "rawMarkdown": "[quote=FengLi;123401]\r\n\r\n@Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .\r\n\r\n[/quote]\r\n\r\naround 0.518; \"identifying hotels that were not booked\" improved my score by about 0.001 and \"widening the leak\" gave my last minute improvement from 0.519 to 0.521 (public LB)",
      "votes": null
    },
    {
      "id": "123404",
      "postDate": "06/11/2016 17:55:34",
      "content": "<p>@ Gert It seems a very large dataset you created. What kind of machine you used? </p>",
      "rawMarkdown": "Gert It seems a very large dataset you created. What kind of machine you used?",
      "votes": null
    },
    {
      "id": "123453",
      "postDate": "06/12/2016 00:25:23",
      "content": "<p>sure, the code is like this:</p>\n\n<blockquote>\n  <p>user_cluster_clicks = train.pivot_table(index='user_id', columns='cluster', values='count') \n  train = train.merge(user_cluster_clicks, on = 'user_id')</p>\n</blockquote>",
      "rawMarkdown": "sure, the code is like this:\r\n\r\n> user_cluster_clicks = train.pivot_table(index='user_id', columns='cluster', values='count') \r\n> train = train.merge(user_cluster_clicks, on = 'user_id')",
      "votes": null
    },
    {
      "id": "123481",
      "postDate": "06/12/2016 07:02:58",
      "content": "<p>[quote=FengLi;123404]</p>\n\n<p>@ Gert It seems a very large dataset you created. What kind of machine you used? </p>\n\n<p>[/quote]</p>\n\n<p>a macbook laptop; the trick was to take only 10% of not-booked clusters for training; and to process the test set in batches!</p>",
      "rawMarkdown": "[quote=FengLi;123404]\r\n\r\n@ Gert It seems a very large dataset you created. What kind of machine you used? \r\n\r\n[/quote]\r\n\r\na macbook laptop; the trick was to take only 10% of not-booked clusters for training; and to process the test set in batches!",
      "votes": null
    },
    {
      "id": "123484",
      "postDate": "06/12/2016 07:14:44",
      "content": "<p>In our team we focused on couple of questions:</p>\n\n<ol>\n<li><p>What should be the relative importance of clicks vs bookings? \nThe optimal weight for click samples relative to book samples was 0.95. So, the clicks samples were only slightly less important than book samples in the xgb model. We used 100 multiclass classification model.</p></li>\n<li><p>What should be the relative importance of &quot;fresh data&quot; vs &quot;old data&quot;? Sergey was testing different weighting schemes to maximize the CV. This brought a significant improvement to the popular script based on local most popular hotels.</p></li>\n<li><p>In which distances should we use XGB vs simple local most popular? Here we determined that above some threshold of destinations sample count, the XGB performs better. I think it was around 5000.</p></li>\n<li><p>Which features added a lot of value to XGB? The strongest one was\nthe time different between moment of booking and moment of stay. My\nintuition is that when book a lot in advance, you are a more cost-conscious traveler.</p></li>\n<li><p>How to use the information coming from the leak to the maximum? We added the leaked samples from test set into the training set for XGB model, since they were almost 100% accurate.  This gave only a small uplift, however.</p></li>\n</ol>",
      "rawMarkdown": "In our team we focused on couple of questions:\r\n\r\n 1. What should be the relative importance of clicks vs bookings? \r\nThe optimal weight for click samples relative to book samples was 0.95. So, the clicks samples were only slightly less important than book samples in the xgb model. We used 100 multiclass classification model.\r\n\r\n 2. What should be the relative importance of \"fresh data\" vs \"old data\"? Sergey was testing different weighting schemes to maximize the CV. This brought a significant improvement to the popular script based on local most popular hotels.\r\n\r\n 3. In which distances should we use XGB vs simple local most popular? Here we determined that above some threshold of destinations sample count, the XGB performs better. I think it was around 5000.\r\n\r\n 4. Which features added a lot of value to XGB? The strongest one was\r\n    the time different between moment of booking and moment of stay. My\r\n    intuition is that when book a lot in advance, you are a more cost-conscious traveler.\r\n\r\n 5. How to use the information coming from the leak to the maximum? We added the leaked samples from test set into the training set for XGB model, since they were almost 100% accurate.  This gave only a small uplift, however.",
      "votes": null
    },
    {
      "id": "123533",
      "postDate": "06/12/2016 15:02:12",
      "content": "<p>@Gert. Thanks for sharing. One of my models was very similar to yours. Unfortunately, I probably went to a wrong way, trying to build 100 logistic regression models, one for each cluster. The model + obvious leaks got 0.498 on public board and added about 0.01 to my final ensemble. I guess on this problem, one xgb &gt; 100 logistic regression.</p>",
      "rawMarkdown": "Gert. Thanks for sharing. One of my models was very similar to yours. Unfortunately, I probably went to a wrong way, trying to build 100 logistic regression models, one for each cluster. The model + obvious leaks got 0.498 on public board and added about 0.01 to my final ensemble. I guess on this problem, one xgb > 100 logistic regression.",
      "votes": null
    },
    {
      "id": "123535",
      "postDate": "06/12/2016 15:32:32",
      "content": "<p>@narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.</p>",
      "rawMarkdown": "narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.",
      "votes": null
    },
    {
      "id": "123548",
      "postDate": "06/12/2016 16:23:06",
      "content": "<p>@eipiplus1 How do you apply the ratio into xgb model? </p>",
      "rawMarkdown": "eipiplus1 How do you apply the ratio into xgb model?",
      "votes": null
    },
    {
      "id": "123549",
      "postDate": "06/12/2016 16:24:28",
      "content": "<p>@narsil You mean that you just train xgb model for the samples whose org_destination_distance &gt;5000? </p>",
      "rawMarkdown": "narsil You mean that you just train xgb model for the samples whose org_destination_distance >5000?",
      "votes": null
    },
    {
      "id": "123616",
      "postDate": "06/13/2016 05:53:53",
      "content": "<p>[quote=eipiplus1;123535]</p>\n\n<p>@narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.</p>\n\n<p>[/quote]</p>\n\n<p>Thank You\n0.95 ratio was just slightly above the 1.0 ratio in the CV, so probably for your CV partition you could have gotten 1.0 as optimal. </p>",
      "rawMarkdown": "[quote=eipiplus1;123535]\r\n\r\n@narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.\r\n\r\n[/quote]\r\n\r\nThank You\r\n0.95 ratio was just slightly above the 1.0 ratio in the CV, so probably for your CV partition you could have gotten 1.0 as optimal.",
      "votes": null
    },
    {
      "id": "123617",
      "postDate": "06/13/2016 05:55:30",
      "content": "<p>[quote=FengLi;123549]</p>\n\n<p>@narsil You mean that you just train xgb model for the samples whose org_destination_distance &gt;5000? </p>\n\n<p>[/quote]</p>\n\n<p>I use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).</p>",
      "rawMarkdown": "[quote=FengLi;123549]\r\n\r\n@narsil You mean that you just train xgb model for the samples whose org_destination_distance >5000? \r\n\r\n[/quote]\r\n\r\nI use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).",
      "votes": null
    },
    {
      "id": "123618",
      "postDate": "06/13/2016 05:55:30",
      "content": "<p>[quote=FengLi;123549]</p>\n\n<p>@narsil You mean that you just train xgb model for the samples whose org_destination_distance &gt;5000? </p>\n\n<p>[/quote]</p>\n\n<p>I use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).</p>",
      "rawMarkdown": "[quote=FengLi;123549]\r\n\r\n@narsil You mean that you just train xgb model for the samples whose org_destination_distance >5000? \r\n\r\n[/quote]\r\n\r\nI use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).",
      "votes": null
    },
    {
      "id": "123660",
      "postDate": "06/13/2016 10:59:13",
      "content": "<p>@Gert, cool to hear you increased your score by &quot;widening the leak&quot; - that's quite a jump! Could you outline the method you used, was it similar to what me and eipiplus1 were trying?</p>",
      "rawMarkdown": "Gert, cool to hear you increased your score by \"widening the leak\" - that's quite a jump! Could you outline the method you used, was it similar to what me and eipiplus1 were trying?",
      "votes": null
    },
    {
      "id": "123701",
      "postDate": "06/13/2016 15:36:19",
      "content": "<p>@Mattias, I guess I did not exploit it very well because I had to code it in a single day after I realized how much you gained from this method.</p>\n\n<p>I created a set of orig_destination_distances to each hotel (srch_destination_id,hotel_cluster) from every user_location_city, separately for every region (user_location_country,user_location_region)</p>\n\n<p>Then for all combinations of region and destination, I looked at every pair of cities to see whether their distance to at least three hotels differed by only a constant (neglecting sets with len&gt;1).</p>\n\n<p>I applied the result as a post processing step to my best submission, inserting the best wide-leak-match (if any) in position 1 if there was no original-leak-match, and position 2 otherwise.</p>\n\n<p>At least I tried to achieve something similar to what you did ;)\nthanks again for sharing, and congratulations on your nice result!</p>",
      "rawMarkdown": "Mattias, I guess I did not exploit it very well because I had to code it in a single day after I realized how much you gained from this method.\r\n\r\nI created a set of orig_destination_distances to each hotel (srch_destination_id,hotel_cluster) from every user_location_city, separately for every region (user_location_country,user_location_region)\r\n\r\nThen for all combinations of region and destination, I looked at every pair of cities to see whether their distance to at least three hotels differed by only a constant (neglecting sets with len>1).\r\n\r\nI applied the result as a post processing step to my best submission, inserting the best wide-leak-match (if any) in position 1 if there was no original-leak-match, and position 2 otherwise.\r\n\r\nAt least I tried to achieve something similar to what you did ;)\r\nthanks again for sharing, and congratulations on your nice result!",
      "votes": null
    },
    {
      "id": "123864",
      "postDate": "06/14/2016 08:23:52",
      "content": "<p>@Gert, thanks for the info. I came up with the idea 48 hours before the deadline so I was also pressed for time, but for me coding to a hard deadline can be lots of fun. </p>",
      "rawMarkdown": "Gert, thanks for the info. I came up with the idea 48 hours before the deadline so I was also pressed for time, but for me coding to a hard deadline can be lots of fun.",
      "votes": null
    },
    {
      "id": "124088",
      "postDate": "06/15/2016 13:19:06",
      "content": "<p>@Gert, congrats, could you please elaborate more on this count features? do you mean for each cluster, you transform each user_id  as frequency counts and you do the same for all the variables? do you do the same for the interaction variables as well?  i always struggle dealing with the high cardinality features like user_id. </p>\n\n<p>how do you decide which 10% to use? random or stratified? </p>\n\n<p>if you use the last four month as holdoutset, how do you do cv? 3+1?</p>\n\n<p>thank you</p>",
      "rawMarkdown": "Gert, congrats, could you please elaborate more on this count features? do you mean for each cluster, you transform each user_id  as frequency counts and you do the same for all the variables? do you do the same for the interaction variables as well?  i always struggle dealing with the high cardinality features like user_id. \r\n\r\nhow do you decide which 10% to use? random or stratified? \r\n\r\nif you use the last four month as holdoutset, how do you do cv? 3+1?\r\n\r\nthank you",
      "votes": null
    },
    {
      "id": "124098",
      "postDate": "06/15/2016 14:16:49",
      "content": "<p>Hi Shubin, I think of it as the frequency of the cluster in the presence of the same user - but it should be the same as the user_id frequency in the presence of the same cluster ;)</p>\n\n<p>Yes for all variables, including interaction variables and even the leak (usercity||distance)</p>\n\n<p>The 10% is random, but only within the not-booked clusters (so it is kind of stratified). I did cv using a random 3-fold within the holdout set.</p>",
      "rawMarkdown": "Hi Shubin, I think of it as the frequency of the cluster in the presence of the same user - but it should be the same as the user_id frequency in the presence of the same cluster ;)\r\n\r\nYes for all variables, including interaction variables and even the leak (usercity||distance)\r\n\r\nThe 10% is random, but only within the not-booked clusters (so it is kind of stratified). I did cv using a random 3-fold within the holdout set.",
      "votes": null
    },
    {
      "id": "124256",
      "postDate": "06/16/2016 14:29:36",
      "content": "<p>@Gert, thanks i got it. just to confirm when you scale the features, you are scaling within the 100 samples you created for one booking row right?</p>\n\n<p>your dataset setup reminds me in previous expedia, the lambdamart model. have you tried that to the dataset you created? </p>\n\n<p>thanks a lot.</p>",
      "rawMarkdown": "Gert, thanks i got it. just to confirm when you scale the features, you are scaling within the 100 samples you created for one booking row right?\r\n\r\nyour dataset setup reminds me in previous expedia, the lambdamart model. have you tried that to the dataset you created? \r\n\r\nthanks a lot.",
      "votes": null
    },
    {
      "id": "124269",
      "postDate": "06/16/2016 16:40:51",
      "content": "<p>@Shubin you are right about the scaling; yes I tried lambdamart, but it was outperformed by XGBoost</p>",
      "rawMarkdown": "Shubin you are right about the scaling; yes I tried lambdamart, but it was outperformed by XGBoost",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 123357,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/11/2016 09:18:55",
      "content": "<p>Interesting question!</p>\n\n<p>I counted bookings/clicks in each cluster for all levels of most variables to create features - I guess that is a lot like your pivot - and then:</p>\n\n<ul>\n<li>have the last 4 months of 2014 as a holdout set that is not used for counting bookings/clicks </li>\n<li>from the holdout set, create a dataset with 100 rows for each booking: one for every cluster, with the cluster specific count for all variables as the features, and a 0/1 (booked) target</li>\n<li>scale the counts of each feature to sum to 1 over all clusters</li>\n<li>train an xgboost model (maxdepth 12!) on the holdout set - including only 10% of 0 targets</li>\n<li>build a testset like the holdout set, using counts over all months including the holdout set (note that this step makes the scaling necessary) and apply the model to the resulting 250,000,000 rows</li>\n</ul>\n\n<p>Apart from the leak variable (city || distance), I also built some interactions with hotel market and a 'user type' variable (usercountry || adultcount || childcount || hotelmarket) that seems to help the model a bit.</p>\n\n<p>During the last days, inspired by the unbelievable score of idle_speculation, I tried to exploit the leak further, by:</p>\n\n<ul>\n<li>identifying hotels that were <em>not</em> booked because the distance did <em>not</em> match the leak feature</li>\n<li>'widening the leak' for cities on a straight line from a destination, as suggested by Mattias</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123366,
      "author_name": "darraghdog",
      "author_url": "",
      "post_date": "06/11/2016 10:16:00",
      "content": "<p>While not in the top30, I made more or less rule based prediction. I was thinking 100 classes is too many to predict with ml, but in the last two weeks realized this must be wrong. </p>\n\n<ul>\n<li>One neat thing was to add the leaked test rows (particularly\nrows where only one hotel_cluster was found in the leak) on top of the training\nset before starting the modelling. This worked very well using 2013 as train, and 2014 as test, but\nnot so well on the LB (about .002). I guess this was because the train and\ntest size is less imbalanced there. </li>\n<li>using bayes mean to weight book vs click per a few different features, as a\nopposed to global book vs click rate </li>\n<li>ranking with layers of weights for different features, bins of stay time, booking lead time, season... </li>\n</ul>\n\n<p>Script is <a href=\"https://www.dropbox.com/s/tgs8cjx66ikcebf/sub-bmleak-1006-0.50745.R?dl=1\">here</a> if its of use to anyone. I won't list out the things that did not work as I guess Kaggle have set some limitation on how long a single comment can be :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123399,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/11/2016 16:53:41",
      "content": "<p>@YijunMa Could you elaborate your pivot table ? Is it user_id based? The data pints larger than 500 means there're at least 500 points each cluster? Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123401,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/11/2016 17:12:06",
      "content": "<p>@Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123403,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/11/2016 17:49:29",
      "content": "<p>[quote=FengLi;123401]</p>\n\n<p>@Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .</p>\n\n<p>[/quote]</p>\n\n<p>around 0.518; &quot;identifying hotels that were not booked&quot; improved my score by about 0.001 and &quot;widening the leak&quot; gave my last minute improvement from 0.519 to 0.521 (public LB)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123404,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/11/2016 17:55:34",
      "content": "<p>@ Gert It seems a very large dataset you created. What kind of machine you used? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123453,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "06/12/2016 00:25:23",
      "content": "<p>sure, the code is like this:</p>\n\n<blockquote>\n  <p>user_cluster_clicks = train.pivot_table(index='user_id', columns='cluster', values='count') \n  train = train.merge(user_cluster_clicks, on = 'user_id')</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123481,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/12/2016 07:02:58",
      "content": "<p>[quote=FengLi;123404]</p>\n\n<p>@ Gert It seems a very large dataset you created. What kind of machine you used? </p>\n\n<p>[/quote]</p>\n\n<p>a macbook laptop; the trick was to take only 10% of not-booked clusters for training; and to process the test set in batches!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123484,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "06/12/2016 07:14:44",
      "content": "<p>In our team we focused on couple of questions:</p>\n\n<ol>\n<li><p>What should be the relative importance of clicks vs bookings? \nThe optimal weight for click samples relative to book samples was 0.95. So, the clicks samples were only slightly less important than book samples in the xgb model. We used 100 multiclass classification model.</p></li>\n<li><p>What should be the relative importance of &quot;fresh data&quot; vs &quot;old data&quot;? Sergey was testing different weighting schemes to maximize the CV. This brought a significant improvement to the popular script based on local most popular hotels.</p></li>\n<li><p>In which distances should we use XGB vs simple local most popular? Here we determined that above some threshold of destinations sample count, the XGB performs better. I think it was around 5000.</p></li>\n<li><p>Which features added a lot of value to XGB? The strongest one was\nthe time different between moment of booking and moment of stay. My\nintuition is that when book a lot in advance, you are a more cost-conscious traveler.</p></li>\n<li><p>How to use the information coming from the leak to the maximum? We added the leaked samples from test set into the training set for XGB model, since they were almost 100% accurate.  This gave only a small uplift, however.</p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123533,
      "author_name": "pichai",
      "author_url": "",
      "post_date": "06/12/2016 15:02:12",
      "content": "<p>@Gert. Thanks for sharing. One of my models was very similar to yours. Unfortunately, I probably went to a wrong way, trying to build 100 logistic regression models, one for each cluster. The model + obvious leaks got 0.498 on public board and added about 0.01 to my final ensemble. I guess on this problem, one xgb &gt; 100 logistic regression.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123535,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/12/2016 15:32:32",
      "content": "<p>@narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123548,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/12/2016 16:23:06",
      "content": "<p>@eipiplus1 How do you apply the ratio into xgb model? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123549,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/12/2016 16:24:28",
      "content": "<p>@narsil You mean that you just train xgb model for the samples whose org_destination_distance &gt;5000? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123616,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "06/13/2016 05:53:53",
      "content": "<p>[quote=eipiplus1;123535]</p>\n\n<p>@narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.</p>\n\n<p>[/quote]</p>\n\n<p>Thank You\n0.95 ratio was just slightly above the 1.0 ratio in the CV, so probably for your CV partition you could have gotten 1.0 as optimal. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123617,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "06/13/2016 05:55:30",
      "content": "<p>[quote=FengLi;123549]</p>\n\n<p>@narsil You mean that you just train xgb model for the samples whose org_destination_distance &gt;5000? </p>\n\n<p>[/quote]</p>\n\n<p>I use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123618,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "06/13/2016 05:55:30",
      "content": "<p>[quote=FengLi;123549]</p>\n\n<p>@narsil You mean that you just train xgb model for the samples whose org_destination_distance &gt;5000? </p>\n\n<p>[/quote]</p>\n\n<p>I use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123660,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/13/2016 10:59:13",
      "content": "<p>@Gert, cool to hear you increased your score by &quot;widening the leak&quot; - that's quite a jump! Could you outline the method you used, was it similar to what me and eipiplus1 were trying?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123701,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/13/2016 15:36:19",
      "content": "<p>@Mattias, I guess I did not exploit it very well because I had to code it in a single day after I realized how much you gained from this method.</p>\n\n<p>I created a set of orig_destination_distances to each hotel (srch_destination_id,hotel_cluster) from every user_location_city, separately for every region (user_location_country,user_location_region)</p>\n\n<p>Then for all combinations of region and destination, I looked at every pair of cities to see whether their distance to at least three hotels differed by only a constant (neglecting sets with len&gt;1).</p>\n\n<p>I applied the result as a post processing step to my best submission, inserting the best wide-leak-match (if any) in position 1 if there was no original-leak-match, and position 2 otherwise.</p>\n\n<p>At least I tried to achieve something similar to what you did ;)\nthanks again for sharing, and congratulations on your nice result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123864,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/14/2016 08:23:52",
      "content": "<p>@Gert, thanks for the info. I came up with the idea 48 hours before the deadline so I was also pressed for time, but for me coding to a hard deadline can be lots of fun. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124088,
      "author_name": "lishubin",
      "author_url": "",
      "post_date": "06/15/2016 13:19:06",
      "content": "<p>@Gert, congrats, could you please elaborate more on this count features? do you mean for each cluster, you transform each user_id  as frequency counts and you do the same for all the variables? do you do the same for the interaction variables as well?  i always struggle dealing with the high cardinality features like user_id. </p>\n\n<p>how do you decide which 10% to use? random or stratified? </p>\n\n<p>if you use the last four month as holdoutset, how do you do cv? 3+1?</p>\n\n<p>thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124098,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/15/2016 14:16:49",
      "content": "<p>Hi Shubin, I think of it as the frequency of the cluster in the presence of the same user - but it should be the same as the user_id frequency in the presence of the same cluster ;)</p>\n\n<p>Yes for all variables, including interaction variables and even the leak (usercity||distance)</p>\n\n<p>The 10% is random, but only within the not-booked clusters (so it is kind of stratified). I did cv using a random 3-fold within the holdout set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124256,
      "author_name": "lishubin",
      "author_url": "",
      "post_date": "06/16/2016 14:29:36",
      "content": "<p>@Gert, thanks i got it. just to confirm when you scale the features, you are scaling within the 100 samples you created for one booking row right?</p>\n\n<p>your dataset setup reminds me in previous expedia, the lambdamart model. have you tried that to the dataset you created? </p>\n\n<p>thanks a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124269,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/16/2016 16:40:51",
      "content": "<p>@Shubin you are right about the scaling; yes I tried lambdamart, but it was outperformed by XGBoost</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "123345": "I am curious about what features did all those who rank 30 or higher created.\r\n\r\nMine is pretty simple: use a pivot table to find out all the clusters a user have clicked, and merge the pivot table with both the training set and testing set. So for each click, we can find a history of what cluster the user have viewed.\r\n\r\nThen I use XGB to train for each destination with data points more than 500 (? I forget the exact number). And for the rest destinations, I use the local popular model.\r\n\r\nThis feature improve my prediction for about 0.005 and I guess it's somehow like collaborative filtering. Of course not exactly the same, but the idea is to conclude from what cluster would an other similar person have clicked.\r\n\r\nSo, what interesting features did you created? And how do you come up with it?",
    "123357": "Interesting question!\r\n\r\nI counted bookings/clicks in each cluster for all levels of most variables to create features - I guess that is a lot like your pivot - and then:\r\n\r\n - have the last 4 months of 2014 as a holdout set that is not used for counting bookings/clicks \r\n - from the holdout set, create a dataset with 100 rows for each booking: one for every cluster, with the cluster specific count for all variables as the features, and a 0/1 (booked) target\r\n - scale the counts of each feature to sum to 1 over all clusters\r\n - train an xgboost model (maxdepth 12!) on the holdout set - including only 10% of 0 targets\r\n - build a testset like the holdout set, using counts over all months including the holdout set (note that this step makes the scaling necessary) and apply the model to the resulting 250,000,000 rows\r\n\r\nApart from the leak variable (city || distance), I also built some interactions with hotel market and a 'user type' variable (usercountry || adultcount || childcount || hotelmarket) that seems to help the model a bit.\r\n\r\nDuring the last days, inspired by the unbelievable score of idle_speculation, I tried to exploit the leak further, by:\r\n \r\n - identifying hotels that were *not* booked because the distance did *not* match the leak feature\r\n - 'widening the leak' for cities on a straight line from a destination, as suggested by Mattias",
    "123366": "While not in the top30, I made more or less rule based prediction. I was thinking 100 classes is too many to predict with ml, but in the last two weeks realized this must be wrong. \r\n\r\n - One neat thing was to add the leaked test rows (particularly\r\n   rows where only one hotel_cluster was found in the leak) on top of the training\r\n   set before starting the modelling. This worked very well using 2013 as train, and 2014 as test, but\r\n   not so well on the LB (about .002). I guess this was because the train and\r\n   test size is less imbalanced there. \r\n - using bayes mean to weight book vs click per a few different features, as a\r\n   opposed to global book vs click rate \r\n - ranking with layers of weights for different features, bins of stay time, booking lead time, season... \r\n\r\nScript is [here][1] if its of use to anyone. I won't list out the things that did not work as I guess Kaggle have set some limitation on how long a single comment can be :) \r\n\r\n\r\n  [1]: https://www.dropbox.com/s/tgs8cjx66ikcebf/sub-bmleak-1006-0.50745.R?dl=1",
    "123399": "YijunMa Could you elaborate your pivot table ? Is it user_id based? The data pints larger than 500 means there're at least 500 points each cluster? Thank you.",
    "123401": "Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .",
    "123403": "[quote=FengLi;123401]\r\n\r\n@Gert. Before you exploit the leak further, how about your score using the feature you mentioned? Thanks .\r\n\r\n[/quote]\r\n\r\naround 0.518; \"identifying hotels that were not booked\" improved my score by about 0.001 and \"widening the leak\" gave my last minute improvement from 0.519 to 0.521 (public LB)",
    "123404": "Gert It seems a very large dataset you created. What kind of machine you used?",
    "123453": "sure, the code is like this:\r\n\r\n> user_cluster_clicks = train.pivot_table(index='user_id', columns='cluster', values='count') \r\n> train = train.merge(user_cluster_clicks, on = 'user_id')",
    "123481": "[quote=FengLi;123404]\r\n\r\n@ Gert It seems a very large dataset you created. What kind of machine you used? \r\n\r\n[/quote]\r\n\r\na macbook laptop; the trick was to take only 10% of not-booked clusters for training; and to process the test set in batches!",
    "123484": "In our team we focused on couple of questions:\r\n\r\n 1. What should be the relative importance of clicks vs bookings? \r\nThe optimal weight for click samples relative to book samples was 0.95. So, the clicks samples were only slightly less important than book samples in the xgb model. We used 100 multiclass classification model.\r\n\r\n 2. What should be the relative importance of \"fresh data\" vs \"old data\"? Sergey was testing different weighting schemes to maximize the CV. This brought a significant improvement to the popular script based on local most popular hotels.\r\n\r\n 3. In which distances should we use XGB vs simple local most popular? Here we determined that above some threshold of destinations sample count, the XGB performs better. I think it was around 5000.\r\n\r\n 4. Which features added a lot of value to XGB? The strongest one was\r\n    the time different between moment of booking and moment of stay. My\r\n    intuition is that when book a lot in advance, you are a more cost-conscious traveler.\r\n\r\n 5. How to use the information coming from the leak to the maximum? We added the leaked samples from test set into the training set for XGB model, since they were almost 100% accurate.  This gave only a small uplift, however.",
    "123533": "Gert. Thanks for sharing. One of my models was very similar to yours. Unfortunately, I probably went to a wrong way, trying to build 100 logistic regression models, one for each cluster. The model + obvious leaks got 0.498 on public board and added about 0.01 to my final ensemble. I guess on this problem, one xgb > 100 logistic regression.",
    "123535": "narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.",
    "123548": "eipiplus1 How do you apply the ratio into xgb model?",
    "123549": "narsil You mean that you just train xgb model for the samples whose org_destination_distance >5000?",
    "123616": "[quote=eipiplus1;123535]\r\n\r\n@narsil I came up with 1.0 click to book ratio, so it was interesting to know your 0.95 ratio. I worked independently for many weeks before I consulted the scripts directory to see what was there. I kept seeing that 3:20 ratio (3/17 all/booking), and I could never figure out where that came from. For one of my models, the optimum ratio was 5:7, never further from 1:1 than that. Congratulations on your result.\r\n\r\n[/quote]\r\n\r\nThank You\r\n0.95 ratio was just slightly above the 1.0 ratio in the CV, so probably for your CV partition you could have gotten 1.0 as optimal.",
    "123617": "[quote=FengLi;123549]\r\n\r\n@narsil You mean that you just train xgb model for the samples whose org_destination_distance >5000? \r\n\r\n[/quote]\r\n\r\nI use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).",
    "123618": "[quote=FengLi;123549]\r\n\r\n@narsil You mean that you just train xgb model for the samples whose org_destination_distance >5000? \r\n\r\n[/quote]\r\n\r\nI use all the samples possible in the train set. I train model also for all samples. But then, for destinations with less than 5000 samples in test set, I substitute XGB prediction with simple prediction coming from popular script (local Top 5 most popular).",
    "123660": "Gert, cool to hear you increased your score by \"widening the leak\" - that's quite a jump! Could you outline the method you used, was it similar to what me and eipiplus1 were trying?",
    "123701": "Mattias, I guess I did not exploit it very well because I had to code it in a single day after I realized how much you gained from this method.\r\n\r\nI created a set of orig_destination_distances to each hotel (srch_destination_id,hotel_cluster) from every user_location_city, separately for every region (user_location_country,user_location_region)\r\n\r\nThen for all combinations of region and destination, I looked at every pair of cities to see whether their distance to at least three hotels differed by only a constant (neglecting sets with len>1).\r\n\r\nI applied the result as a post processing step to my best submission, inserting the best wide-leak-match (if any) in position 1 if there was no original-leak-match, and position 2 otherwise.\r\n\r\nAt least I tried to achieve something similar to what you did ;)\r\nthanks again for sharing, and congratulations on your nice result!",
    "123864": "Gert, thanks for the info. I came up with the idea 48 hours before the deadline so I was also pressed for time, but for me coding to a hard deadline can be lots of fun.",
    "124088": "Gert, congrats, could you please elaborate more on this count features? do you mean for each cluster, you transform each user_id  as frequency counts and you do the same for all the variables? do you do the same for the interaction variables as well?  i always struggle dealing with the high cardinality features like user_id. \r\n\r\nhow do you decide which 10% to use? random or stratified? \r\n\r\nif you use the last four month as holdoutset, how do you do cv? 3+1?\r\n\r\nthank you",
    "124098": "Hi Shubin, I think of it as the frequency of the cluster in the presence of the same user - but it should be the same as the user_id frequency in the presence of the same cluster ;)\r\n\r\nYes for all variables, including interaction variables and even the leak (usercity||distance)\r\n\r\nThe 10% is random, but only within the not-booked clusters (so it is kind of stratified). I did cv using a random 3-fold within the holdout set.",
    "124256": "Gert, thanks i got it. just to confirm when you scale the features, you are scaling within the 100 samples you created for one booking row right?\r\n\r\nyour dataset setup reminds me in previous expedia, the lambdamart model. have you tried that to the dataset you created? \r\n\r\nthanks a lot.",
    "124269": "Shubin you are right about the scaling; yes I tried lambdamart, but it was outperformed by XGBoost"
  },
  "source": "meta"
}