{
  "id": 21526,
  "title": "Widening the leak",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21526",
  "author_name": "",
  "post_date": "2016-06-09T05:43:14.530Z",
  "votes": 9,
  "comment_count": 39,
  "views": 4342,
  "content": "<p>I was thinking about the leak, and I realized that there must be cities which are close enough to eachother that their distances to hotels on the other side of the world will be very similar. Different, but the difference ought to be constant.</p>\n\n<p>Take city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.</p>\n\n<p>How can this be exploited? If we've been able to find a constant distance between city A and city B, the leaked hotels for city A can be used for city B and vice versa.</p>\n\n<p>It does seem to work, but it's not the huge gain I was  hoping for ;)</p>",
  "messages": [
    {
      "id": "123008",
      "postDate": "06/09/2016 05:43:14",
      "content": "<p>I was thinking about the leak, and I realized that there must be cities which are close enough to eachother that their distances to hotels on the other side of the world will be very similar. Different, but the difference ought to be constant.</p>\n\n<p>Take city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.</p>\n\n<p>How can this be exploited? If we've been able to find a constant distance between city A and city B, the leaked hotels for city A can be used for city B and vice versa.</p>\n\n<p>It does seem to work, but it's not the huge gain I was  hoping for ;)</p>",
      "rawMarkdown": "I was thinking about the leak, and I realized that there must be cities which are close enough to eachother that their distances to hotels on the other side of the world will be very similar. Different, but the difference ought to be constant.\r\n\r\nTake city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.\r\n\r\nHow can this be exploited? If we've been able to find a constant distance between city A and city B, the leaked hotels for city A can be used for city B and vice versa.\r\n\r\nIt does seem to work, but it's not the huge gain I was  hoping for ;)",
      "votes": null
    },
    {
      "id": "123021",
      "postDate": "06/09/2016 07:25:22",
      "content": "<p>I tried something similar, but not quite the same.  For a given (user_country, user_city, hotel_country, destid) I looked at the pairs (train_distance, HC) in relation, for all the training data - call it D.  Then for test, I explored those keys (user_country, user_city, hotel_country, destid) at test_distance and noted all the matching D entries where abs(test_distance - train_distance)\n\n</p><p>These sets of HC at 1 mile are sometimes 60+!  At distance 0.1 mile, there are even some 10+ counts.  I focused in on cases where it was HC set sizes of 1 or 2 only.  I compared the HC sets thus found with my best submission (out to N places), and these D HC often were at positions 6+ in my lists.</p>\n\n<p>The reason to focus on set sizes 1 or 2 is that I wouldn't put much faith in the utility for a larger set of HC for, say, Chicago or New York City or Frankfurt or Beijing or wherever.  On the flip side, that means I may be limiting to HC sets for smaller destinations with fewer booked trips and, with less training data, higher variance if applied to predictions for test.</p>\n\n<p>I purposely left out hotel_market, but it may be worth trying.  </p>\n\n<p>There was one catch to all this for me - I saved data from R to file, and didn't set options(digits=6) or such.  So, the distances I used in my coding was rounded to 3 digits.  The problem with this is that I was awarding a huge weight to exact matches, where abs(x-y)&lt;0.000001 (float precision), and therefore though there were certainly many exact matches in my data, when the places were 3 or less, I definitely labeled many of the rows as INEXACT when they were exact.</p>\n\n<p>Does this idea have some merit?  It's not what you described, but on the same train of thought.</p>",
      "rawMarkdown": "I tried something similar, but not quite the same.  For a given (user_country, user_city, hotel_country, destid) I looked at the pairs (train_distance, HC) in relation, for all the training data - call it D.  Then for test, I explored those keys (user_country, user_city, hotel_country, destid) at test_distance and noted all the matching D entries where abs(test_distance - train_distance)<M (for M=1, 0.5, and 0.1=one city block).  The goal was to uncover instances in the test data where the set of candidate HC could be known exactly.\r\n\r\nThese sets of HC at 1 mile are sometimes 60+!  At distance 0.1 mile, there are even some 10+ counts.  I focused in on cases where it was HC set sizes of 1 or 2 only.  I compared the HC sets thus found with my best submission (out to N places), and these D HC often were at positions 6+ in my lists.\r\n\r\nThe reason to focus on set sizes 1 or 2 is that I wouldn't put much faith in the utility for a larger set of HC for, say, Chicago or New York City or Frankfurt or Beijing or wherever.  On the flip side, that means I may be limiting to HC sets for smaller destinations with fewer booked trips and, with less training data, higher variance if applied to predictions for test.\r\n\r\nI purposely left out hotel_market, but it may be worth trying.  \r\n\r\nThere was one catch to all this for me - I saved data from R to file, and didn't set options(digits=6) or such.  So, the distances I used in my coding was rounded to 3 digits.  The problem with this is that I was awarding a huge weight to exact matches, where abs(x-y)<0.000001 (float precision), and therefore though there were certainly many exact matches in my data, when the places were 3 or less, I definitely labeled many of the rows as INEXACT when they were exact.\r\n\r\nDoes this idea have some merit?  It's not what you described, but on the same train of thought.",
      "votes": null
    },
    {
      "id": "123025",
      "postDate": "06/09/2016 07:53:56",
      "content": "<p>I think that that would pick up on the leak, but weaken it. The leak is that any combination of [orig_destination_distance, srch_destination_id_base, user_location_city] typically has  a unique hotel cluster - because it identifies one exact hotel.</p>\n\n<p>You're looking at [user_country, user_city, hotel_country, destid]. But user_country is tied to user_city and hotel_country is mostly tied to destid. You're not grouping on the exact orig_destination_distance, but a soft grouping. That means that on top of the exact matches (the leak) you'll get a number of false matches.</p>\n\n<p>To tell the truth, I also tried this, looking at hotels that don't have exact distance matches, but matching up to some range. The larger I made the range, the weaker the feature became. Leaving the range at 0 (which then becomes the leak) worked best in my attempts.</p>\n\n<p>/m</p>",
      "rawMarkdown": "I think that that would pick up on the leak, but weaken it. The leak is that any combination of [orig_destination_distance, srch_destination_id_base, user_location_city] typically has  a unique hotel cluster - because it identifies one exact hotel.\r\n\r\nYou're looking at [user_country, user_city, hotel_country, destid]. But user_country is tied to user_city and hotel_country is mostly tied to destid. You're not grouping on the exact orig_destination_distance, but a soft grouping. That means that on top of the exact matches (the leak) you'll get a number of false matches.\r\n\r\nTo tell the truth, I also tried this, looking at hotels that don't have exact distance matches, but matching up to some range. The larger I made the range, the weaker the feature became. Leaving the range at 0 (which then becomes the leak) worked best in my attempts.\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "123055",
      "postDate": "06/09/2016 12:21:27",
      "content": "<p>I can now report that this technique improved my score by 0.00437, so; yay!</p>",
      "rawMarkdown": "I can now report that this technique improved my score by 0.00437, so; yay!",
      "votes": null
    },
    {
      "id": "123067",
      "postDate": "06/09/2016 13:50:39",
      "content": "<p>Nice going! So close cities with constant differences has some value, good find. If I have time, I might try that.</p>",
      "rawMarkdown": "Nice going! So close cities with constant differences has some value, good find. If I have time, I might try that.",
      "votes": null
    },
    {
      "id": "123135",
      "postDate": "06/09/2016 22:40:33",
      "content": "<p><em>Take city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.</em></p>\n\n<p>@Mattias Fagerlund Why I didn't find the pattern in this example.It's seems in each pair, the orig_destination_distance are all different and the difference between them is also not -5.061.</p>\n\n<p>@eipiplus1   Does you approach work? (abs(x-y)&lt;0.000001 ). So tiny difference will cause mismatch?</p>",
      "rawMarkdown": "*Take city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.*\r\n\r\n@Mattias Fagerlund Why I didn't find the pattern in this example.It's seems in each pair, the orig_destination_distance are all different and the difference between them is also not -5.061.\r\n\r\n@eipiplus1   Does you approach work? (abs(x-y)<0.000001 ). So tiny difference will cause mismatch?",
      "votes": null
    },
    {
      "id": "123153",
      "postDate": "06/10/2016 00:16:00",
      "content": "<p>@FengLi</p>\n\n<p>Here is the data per his example:</p>\n\n<blockquote>\n  <p>49178 v 55729 - 8745 - 64 - 10555.076 or 10550.015 --&gt; 5.0610 </p>\n  \n  <p>49178 v 55729 - 8745 - 64 - 10555.076 or 10550.2091 --&gt; 4.8669 </p>\n  \n  <p>49178 v 55729 - 8745 - 64 - 10555.076 or 10550.8242 --&gt; 4.2518 </p>\n  \n  <p>49178 v 55729 - 8745 - 8 - 10556.3596 or 10548.0258 --&gt; 8.3338 </p>\n  \n  <p>49178 v 55729 - 8745 - 97 - 10555.228 or 10550.167 --&gt; 5.0610</p>\n</blockquote>\n\n<p>This is data for country=77, region=824, showing c1 v c2.  All data is for destid=8745 as you can see.\nThe HCs are 64, 8, and 97.  Note that the distance changed for HC 64, for whatever reason.</p>\n\n<p>I can confirm the 5.0610 difference, seen exactly twice here.  </p>\n\n<p>I take this to mean that for a pair of cities (c1, c2), if there exists a distance difference DD which repeats 2+ times, this is a sign that they may be sister cities (as I'm calling them).  In this case, we have 2/1/1/1 when grouping by DD.</p>\n\n<p>@Mattias, I'm looking at this now in terms of 3 thresholds: (1) DD&gt;=N, some critical value, like N=2 or N=3 or higher - it turns out, there are a whole lot of sister cities with at least one DD=2; or, the density of non-1 DD counts amongst all DD counts is at least X% (2) DD&lt;=A miles, that is, the sister cities must be no more than A miles apart (3) leakcity to destid distance&gt;=B miles, that is, the destid must be at least B miles from the leakcity.  I have tried N=2, A=100, and B=1000 so far.  Perhaps sister cities should be considered as closer than that.  With B=1000, I'm trying to avoid &quot;local&quot; trips, some of which are hotel bookings within a 5 mile radius.  I haven't tuned N/A/B yet, these are just tester values.  What do you think?</p>",
      "rawMarkdown": "FengLi\r\n\r\nHere is the data per his example:\r\n\r\n> 49178 v 55729 - 8745 - 64 - 10555.076 or 10550.015 --> 5.0610 \r\n\r\n> 49178 v 55729 - 8745 - 64 - 10555.076 or 10550.2091 --> 4.8669 \r\n\r\n> 49178 v 55729 - 8745 - 64 - 10555.076 or 10550.8242 --> 4.2518 \r\n\r\n> 49178 v 55729 - 8745 - 8 - 10556.3596 or 10548.0258 --> 8.3338 \r\n\r\n> 49178 v 55729 - 8745 - 97 - 10555.228 or 10550.167 --> 5.0610\r\n\r\nThis is data for country=77, region=824, showing c1 v c2.  All data is for destid=8745 as you can see.\r\nThe HCs are 64, 8, and 97.  Note that the distance changed for HC 64, for whatever reason.\r\n\r\nI can confirm the 5.0610 difference, seen exactly twice here.  \r\n\r\nI take this to mean that for a pair of cities (c1, c2), if there exists a distance difference DD which repeats 2+ times, this is a sign that they may be sister cities (as I'm calling them).  In this case, we have 2/1/1/1 when grouping by DD.\r\n\r\n@Mattias, I'm looking at this now in terms of 3 thresholds: (1) DD>=N, some critical value, like N=2 or N=3 or higher - it turns out, there are a whole lot of sister cities with at least one DD=2; or, the density of non-1 DD counts amongst all DD counts is at least X% (2) DD<=A miles, that is, the sister cities must be no more than A miles apart (3) leakcity to destid distance>=B miles, that is, the destid must be at least B miles from the leakcity.  I have tried N=2, A=100, and B=1000 so far.  Perhaps sister cities should be considered as closer than that.  With B=1000, I'm trying to avoid \"local\" trips, some of which are hotel bookings within a 5 mile radius.  I haven't tuned N/A/B yet, these are just tester values.  What do you think?",
      "votes": null
    },
    {
      "id": "123154",
      "postDate": "06/10/2016 00:20:13",
      "content": "<p>@FengLi</p>\n\n<p>At this point in the game, I'm inclined to say my approach won't work.  I still hold hope for it, but with 1 day left, I won't have to time to further explore.  What Mattias is describing makes use of the leak heuristic, which is proven, while being exploratory about cities; what I'm doing is unproven and, while it may turn out to be the big secret, I'm sure it would require a lot more diligence to figure out exactly how to make use of the radii mapped destid/HC.</p>",
      "rawMarkdown": "FengLi\r\n\r\nAt this point in the game, I'm inclined to say my approach won't work.  I still hold hope for it, but with 1 day left, I won't have to time to further explore.  What Mattias is describing makes use of the leak heuristic, which is proven, while being exploratory about cities; what I'm doing is unproven and, while it may turn out to be the big secret, I'm sure it would require a lot more diligence to figure out exactly how to make use of the radii mapped destid/HC.",
      "votes": null
    },
    {
      "id": "123160",
      "postDate": "06/10/2016 02:00:25",
      "content": "<p>@Mattias Fagerlund - Are you saying that the processes is as follows:\n1. cluster rows by user_location_country, user_location_region, srch_destination_id.\n2. Given some example that i want to make a prediction on:\nTake it's orig_destination_distance, compare it to all other groups.\n3. You get a list of values. \n4. What is the next step? </p>\n\n<p>Thank you.</p>",
      "rawMarkdown": "Mattias Fagerlund - Are you saying that the processes is as follows:\r\n1. cluster rows by user_location_country, user_location_region, srch_destination_id.\r\n2. Given some example that i want to make a prediction on:\r\nTake it's orig_destination_distance, compare it to all other groups.\r\n3. You get a list of values. \r\n4. What is the next step? \r\n\r\nThank you.",
      "votes": null
    },
    {
      "id": "123178",
      "postDate": "06/10/2016 05:39:31",
      "content": "<p>@epiplus1 - that sounds about right. </p>\n\n<p>I find cities that are close and that are considered for &quot;sister city&quot; testing. I then find potential target search destinations that are far away (both sisters must have hotels there). From both sister cities. I then start to try to work out if they have any N hotels that have the same hotel cluster and the same distace - given some fixed constant between them. </p>\n\n<p>I'm not able to find many, here's one run on my validation set;\nTotal: 0.03619763(1976304), Specific: 0.4383184(163209), Count: 200585</p>\n\n<p>So <em>at best</em> the result would improve by 0.036 - but a lot of those would be in the regular leak and so there would be no additional benefits to using this method.</p>\n\n<p>For that run, near limit=30 miles, far limit=150 miles, N (min number of matched hotels) was 2, and the city comparison epsilon was 0.01.</p>\n\n<p>These values can be tweaked to get more hits but of lower quality. They can be tweaked the other way to give fewer hits, but of greater quality. This particular run found 200585 hits and got a score of 0.438 on those. Now, 0.438 is better than any other feature I've found with the exception of the leak. so that's good. But the fact that it creates so few votes means it's not super helpful.</p>",
      "rawMarkdown": "epiplus1 - that sounds about right. \r\n\r\nI find cities that are close and that are considered for \"sister city\" testing. I then find potential target search destinations that are far away (both sisters must have hotels there). From both sister cities. I then start to try to work out if they have any N hotels that have the same hotel cluster and the same distace - given some fixed constant between them. \r\n\r\nI'm not able to find many, here's one run on my validation set;\r\nTotal: 0.03619763(1976304), Specific: 0.4383184(163209), Count: 200585\r\n\r\nSo *at best* the result would improve by 0.036 - but a lot of those would be in the regular leak and so there would be no additional benefits to using this method.\r\n\r\nFor that run, near limit=30 miles, far limit=150 miles, N (min number of matched hotels) was 2, and the city comparison epsilon was 0.01.\r\n\r\nThese values can be tweaked to get more hits but of lower quality. They can be tweaked the other way to give fewer hits, but of greater quality. This particular run found 200585 hits and got a score of 0.438 on those. Now, 0.438 is better than any other feature I've found with the exception of the leak. so that's good. But the fact that it creates so few votes means it's not super helpful.",
      "votes": null
    },
    {
      "id": "123179",
      "postDate": "06/10/2016 06:00:26",
      "content": "<p>@kaukau, did that answer your question? There's 750 lines of code to do this, so it's not trivial, but I think it's an interesting idea!</p>",
      "rawMarkdown": "kaukau, did that answer your question? There's 750 lines of code to do this, so it's not trivial, but I think it's an interesting idea!",
      "votes": null
    },
    {
      "id": "123184",
      "postDate": "06/10/2016 06:15:38",
      "content": "<p>Thanks for the explanation and analysis.  I implemented two different versions of this, based on whether I strictly included the leakcity leakdist as part of my key, and here is what I found.  Starting with the 852K leak data rows, I expanded out the HC map and then reduced the final key list uniquely.  So that's the leak city search domain.  Then, I searched across the entire (country, region, destid) universe for any cities c1!=c2, computing the diff(dist) per HC, and then aggregating.  I ended up with 1.19M sister rows with leakdist, 964K sister rows without leakdist.  The advantage of the 964K row set is not having to compute sistercitydist+diff, I have the leakdist at hand.</p>\n\n<p>Using the 964K set, I tried twice and got worse LB scores.  I used the 1.19M data set and improved my LB score.  It requires calculating sistercitydist+diff, as I mentioned.  I got 3,100 test row diff in that submission, and it proved HIGHLY effective.</p>\n\n<p>I found that I missed many sister cities due to float imprecison.  Now that I'm using numeric(12,4) I'm getting toward 7,500 matches.</p>\n\n<p>To go one step further, I'm going to eventually test nearest neighbor cities constrained to a diff tolerance of 0.0005 or thereabouts.  Have you tried that?  It is not the leak heuristic, but it may prove out.  It's effectively going back to my radius idea, but at a WAY smaller value!</p>\n\n<p>By the way, when it came to crunching the 1.19M leak domain data set against the 11.5M row search city data set, I wrote in SQL.  That flew!  It was the most natural way to join and process this data, in my view.</p>\n\n<p>I'm still working on it and measuring the efficacy of my approach.  Thanks again for sharing ideas and congrats for your place on the LB.</p>",
      "rawMarkdown": "Thanks for the explanation and analysis.  I implemented two different versions of this, based on whether I strictly included the leakcity leakdist as part of my key, and here is what I found.  Starting with the 852K leak data rows, I expanded out the HC map and then reduced the final key list uniquely.  So that's the leak city search domain.  Then, I searched across the entire (country, region, destid) universe for any cities c1!=c2, computing the diff(dist) per HC, and then aggregating.  I ended up with 1.19M sister rows with leakdist, 964K sister rows without leakdist.  The advantage of the 964K row set is not having to compute sistercitydist+diff, I have the leakdist at hand.\r\n\r\nUsing the 964K set, I tried twice and got worse LB scores.  I used the 1.19M data set and improved my LB score.  It requires calculating sistercitydist+diff, as I mentioned.  I got 3,100 test row diff in that submission, and it proved HIGHLY effective.\r\n\r\nI found that I missed many sister cities due to float imprecison.  Now that I'm using numeric(12,4) I'm getting toward 7,500 matches.\r\n\r\nTo go one step further, I'm going to eventually test nearest neighbor cities constrained to a diff tolerance of 0.0005 or thereabouts.  Have you tried that?  It is not the leak heuristic, but it may prove out.  It's effectively going back to my radius idea, but at a WAY smaller value!\r\n\r\nBy the way, when it came to crunching the 1.19M leak domain data set against the 11.5M row search city data set, I wrote in SQL.  That flew!  It was the most natural way to join and process this data, in my view.\r\n\r\nI'm still working on it and measuring the efficacy of my approach.  Thanks again for sharing ideas and congrats for your place on the LB.",
      "votes": null
    },
    {
      "id": "123216",
      "postDate": "06/10/2016 11:59:09",
      "content": "<p>Thanks, I was able to improve my score some more using this technique (by 0.00048, meaning 0.00485 from this technique alone), but that was only enough to move up one position. </p>\n\n<p>As I was typing this, I made another submission that improved by another 0.00721, so that's 0,01206 from this technique. Currently at #11.</p>\n\n<p>I feel I'm running out of time here, very few submissions left to go. I'm thinking I could probably eek a bit more performance out of this, but I'm going to have to let it go.</p>\n\n<p>A further cool thing I would have liked to have done was to create a map given the city-hotel distances. This is surprisingly easy for 2d coordinates, you just keep moving the nodes (cities and hotels) around to satisfy their internal distance demands. Takes a number of iterations, but the math is trivial and fast. But on a sphere, it's not so easy, because you have to place the nodes using longitudes and latitudes and you have to handle wrap-around issues and so on. Very messy.</p>\n\n<p>But once you have a map, you don't even need sister cities to both index the same hotels, with the map, you can search the map for hotels with a distance of X from city Y. But alas, there isn't enough time left for that - and it's a fairly low probability that I would have been able to get sufficient precision on the map coordinates.</p>\n\n<p>/m</p>",
      "rawMarkdown": "Thanks, I was able to improve my score some more using this technique (by 0.00048, meaning 0.00485 from this technique alone), but that was only enough to move up one position. \r\n\r\nAs I was typing this, I made another submission that improved by another 0.00721, so that's 0,01206 from this technique. Currently at #11.\r\n\r\nI feel I'm running out of time here, very few submissions left to go. I'm thinking I could probably eek a bit more performance out of this, but I'm going to have to let it go.\r\n\r\nA further cool thing I would have liked to have done was to create a map given the city-hotel distances. This is surprisingly easy for 2d coordinates, you just keep moving the nodes (cities and hotels) around to satisfy their internal distance demands. Takes a number of iterations, but the math is trivial and fast. But on a sphere, it's not so easy, because you have to place the nodes using longitudes and latitudes and you have to handle wrap-around issues and so on. Very messy.\r\n\r\nBut once you have a map, you don't even need sister cities to both index the same hotels, with the map, you can search the map for hotels with a distance of X from city Y. But alas, there isn't enough time left for that - and it's a fairly low probability that I would have been able to get sufficient precision on the map coordinates.\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "123223",
      "postDate": "06/10/2016 12:38:27",
      "content": "<p>@Mattias Fagerlund - How did you find cities that are close?\nI mean if i group by country, region and dest, and after that check all the all possible paris of\norig_di between one city and all other in this group i will get a lot of combinations?\nHow did you handle that?</p>",
      "rawMarkdown": "Mattias Fagerlund - How did you find cities that are close?\r\nI mean if i group by country, region and dest, and after that check all the all possible paris of\r\norig_di between one city and all other in this group i will get a lot of combinations?\r\nHow did you handle that?",
      "votes": null
    },
    {
      "id": "123225",
      "postDate": "06/10/2016 13:08:39",
      "content": "<p>Actually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. </p>\n\n<p>You can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping &quot;near&quot; lists. </p>\n\n<p>If the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).</p>\n\n<p>But as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. </p>\n\n<p>For instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.</p>\n\n<p>So you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.</p>",
      "rawMarkdown": "Actually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. \r\n\r\nYou can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping \"near\" lists. \r\n\r\nIf the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).\r\n\r\nBut as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. \r\n\r\nFor instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.\r\n\r\nSo you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.",
      "votes": null
    },
    {
      "id": "123228",
      "postDate": "06/10/2016 13:25:44",
      "content": "<p>Very interesting Mattias, thanks for sharing. I would not be surprised if the solution of idle_speculation is related to what you described in your previous post - find all locations and extend the leak to all records with a distance in the test set. I have looked for ways to do that, but without external data, sufficient precision is quite a challenge like you state.</p>",
      "rawMarkdown": "Very interesting Mattias, thanks for sharing. I would not be surprised if the solution of idle_speculation is related to what you described in your previous post - find all locations and extend the leak to all records with a distance in the test set. I have looked for ways to do that, but without external data, sufficient precision is quite a challenge like you state.",
      "votes": null
    },
    {
      "id": "123230",
      "postDate": "06/10/2016 13:33:04",
      "content": "<p>I have been working on this for a few days as well - this is a tough challange to map all these distances - good work Mattias. Few more days and I think you would get into top3 :) I am highly certain idle_specaulation did this.</p>",
      "rawMarkdown": "I have been working on this for a few days as well - this is a tough challange to map all these distances - good work Mattias. Few more days and I think you would get into top3 :) I am highly certain idle_specaulation did this.",
      "votes": null
    },
    {
      "id": "123233",
      "postDate": "06/10/2016 14:01:23",
      "content": "<p>Yikes. One more submission to go, and I just reached #6.</p>\n\n<p><a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/leaderboard?submissionId=3150212\">https://www.kaggle.com/c/expedia-hotel-recommendations/leaderboard?submissionId=3150212</a></p>\n\n<p>Best placement ever for me ever on Kaggle, let's see if I can hang on. But this thread describes what I've been doing; I'm basically relaxing my requirements to increase the number of votes.</p>",
      "rawMarkdown": "Yikes. One more submission to go, and I just reached #6.\r\n\r\nhttps://www.kaggle.com/c/expedia-hotel-recommendations/leaderboard?submissionId=3150212\r\n\r\nBest placement ever for me ever on Kaggle, let's see if I can hang on. But this thread describes what I've been doing; I'm basically relaxing my requirements to increase the number of votes.",
      "votes": null
    },
    {
      "id": "123234",
      "postDate": "06/10/2016 14:06:19",
      "content": "<p>wish you best of luck. with a careful submission top3 is yours :)</p>",
      "rawMarkdown": "wish you best of luck. with a careful submission top3 is yours :)",
      "votes": null
    },
    {
      "id": "123253",
      "postDate": "06/10/2016 17:22:52",
      "content": "<p>Great work!  Congrats on reaching #6.  You were using &quot;strongly match&quot; as N&gt;=2 for number of same distances to HC in a single destination.  Are you using N&gt;=3 or more now?  My dist. is:</p>\n\n<blockquote>\n  <p>10679303  2  </p>\n  \n  <p>335235  3   </p>\n  \n  <p>25042  4    </p>\n  \n  <p>3840  5    </p>\n  \n  <p>1102  6</p>\n  \n  <p>507  7</p>\n  \n  <p>297  8</p>\n  \n  <p>217  9</p>\n  \n  <p>117  10</p>\n</blockquote>\n\n<p>so there is quite a drop from 2 to 3.  I am still using 2 as the qualify cutoff.  The thing is, I know that when I go through the search space of (HC, dist) for a given destid, where there can be and often are multiple dist for a given HC, that the probability of a match on dist+diff to leakcity dist is very, very small.  So using N&gt;=2 seems ok to me, or, in fact, may be missing out on lots of matches if I don't include it. </p>\n\n<p>[quote=Mattias Fagerlund;123225]</p>\n\n<p>Actually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. </p>\n\n<p>You can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping &quot;near&quot; lists. </p>\n\n<p>If the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).</p>\n\n<p>But as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. </p>\n\n<p>For instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.</p>\n\n<p>So you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Great work!  Congrats on reaching #6.  You were using \"strongly match\" as N>=2 for number of same distances to HC in a single destination.  Are you using N>=3 or more now?  My dist. is:\r\n\r\n> 10679303  2  \r\n\r\n> 335235  3   \r\n\r\n> 25042  4    \r\n\r\n> 3840  5    \r\n\r\n> 1102  6\r\n\r\n>  507  7\r\n\r\n>  297  8\r\n\r\n>  217  9\r\n\r\n>  117  10\r\n\r\n\r\nso there is quite a drop from 2 to 3.  I am still using 2 as the qualify cutoff.  The thing is, I know that when I go through the search space of (HC, dist) for a given destid, where there can be and often are multiple dist for a given HC, that the probability of a match on dist+diff to leakcity dist is very, very small.  So using N>=2 seems ok to me, or, in fact, may be missing out on lots of matches if I don't include it. \r\n\r\n\r\n\r\n[quote=Mattias Fagerlund;123225]\r\n\r\nActually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. \r\n\r\nYou can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping \"near\" lists. \r\n\r\nIf the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).\r\n\r\nBut as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. \r\n\r\nFor instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.\r\n\r\nSo you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "123261",
      "postDate": "06/10/2016 18:10:19",
      "content": "<p>For full transparancy, I'll reveal my &quot;final&quot; constants. There's no guarantee that these are the best, but they worked well for me in testing. Also, I'm afraid I won't be able to make a subission using these constants, because so far the system has been working for hours without progress. </p>\n\n<pre><code>    public float LimitForClose = 600;\n    public float LimitForFar = 0;\n    public int MaxDistanceDifferenceFactor = 200000;\n    public int ItemsToCheckPerCity = 40000;\n    public int MaxDistancesToCheck = 5000;\n    public int MinimumHits = 1;\n    public float Epsilon = 0.05f;//0.00001f;\n    public float MatchEpsilon = 0.05f;\n    public int MaxConstantsPerDestination = 750;\n    public float MaxSearchDestinationSize = 1000;\n</code></pre>\n\n<p>But if things holds true, and if I'm able to make a submission, it should be significantly better than my previous submissions.</p>\n\n<p>Testing for my best submission:\n  Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121 =&gt; 0.52171</p>\n\n<p>Testing for the submission I'm computing now:\n  Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530</p>\n\n<p>Now competition for #1, but it should result in better scores than my previous attempt. Problem though, with the previous entry it took 1:26 to create a submission (hour and a half). This submission looks to be about 3 times as complicated. So at 4 hours+, I may be out of time to even submit it. Oh well, that's the game, isn't it?</p>\n\n<p>/m</p>",
      "rawMarkdown": "For full transparancy, I'll reveal my \"final\" constants. There's no guarantee that these are the best, but they worked well for me in testing. Also, I'm afraid I won't be able to make a subission using these constants, because so far the system has been working for hours without progress. \r\n\r\n        public float LimitForClose = 600;\r\n        public float LimitForFar = 0;\r\n        public int MaxDistanceDifferenceFactor = 200000;\r\n        public int ItemsToCheckPerCity = 40000;\r\n        public int MaxDistancesToCheck = 5000;\r\n        public int MinimumHits = 1;\r\n        public float Epsilon = 0.05f;//0.00001f;\r\n        public float MatchEpsilon = 0.05f;\r\n        public int MaxConstantsPerDestination = 750;\r\n        public float MaxSearchDestinationSize = 1000;\r\n\r\nBut if things holds true, and if I'm able to make a submission, it should be significantly better than my previous submissions.\r\n\r\nTesting for my best submission:\r\n  Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121 => 0.52171\r\n\r\nTesting for the submission I'm computing now:\r\n  Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530\r\n\r\nNow competition for #1, but it should result in better scores than my previous attempt. Problem though, with the previous entry it took 1:26 to create a submission (hour and a half). This submission looks to be about 3 times as complicated. So at 4 hours+, I may be out of time to even submit it. Oh well, that's the game, isn't it?\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "123264",
      "postDate": "06/10/2016 18:24:43",
      "content": "<p>anyway, share your algo after the competition, very interesting to see what have you done :)</p>",
      "rawMarkdown": "anyway, share your algo after the competition, very interesting to see what have you done :)",
      "votes": null
    },
    {
      "id": "123265",
      "postDate": "06/10/2016 18:27:42",
      "content": "<p>@epiplus I'm no longer sure if I used 1 or 2 for my latest submission - probably 1 though. It decreases the quality but it increases the number of votes. On the whole, it seems better. But our mileages may vary though.</p>\n\n<p>/m</p>",
      "rawMarkdown": "epiplus I'm no longer sure if I used 1 or 2 for my latest submission - probably 1 though. It decreases the quality but it increases the number of votes. On the whole, it seems better. But our mileages may vary though.\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "123267",
      "postDate": "06/10/2016 18:30:08",
      "content": "<p>LimitForClose = max dist between city1 and city2? min dist from leakcity to destid?</p>\n\n<p>I didn't put a limit on max distances to check, that's an interesting one to consider for speeding things up.  I do have an early check for if sistercitydist+diff&lt;=0, in which case a diff was added to the base sister city and results in negative dist to destid, which of course is impossible.</p>\n\n<p>I exploded my city-city table to not be constrained by country/region, and now I'm running with that.  It was that, or create a key for (country, region, city) and do the Cartesian-ish product on that joining still by destid and HC.  So now my domain search space is 11.9M rows.</p>\n\n<p>I'm sitting on two right now with 34K row diff and 38K row diff, looking things over, code checking, comparing to previous submissions, checking how close sister cities are, distribution of distances to destid, etc.  You know how it goes!</p>\n\n<p>Good luck on your final submissions!  </p>",
      "rawMarkdown": "LimitForClose = max dist between city1 and city2? min dist from leakcity to destid?\r\n\r\nI didn't put a limit on max distances to check, that's an interesting one to consider for speeding things up.  I do have an early check for if sistercitydist+diff<=0, in which case a diff was added to the base sister city and results in negative dist to destid, which of course is impossible.\r\n\r\nI exploded my city-city table to not be constrained by country/region, and now I'm running with that.  It was that, or create a key for (country, region, city) and do the Cartesian-ish product on that joining still by destid and HC.  So now my domain search space is 11.9M rows.\r\n\r\nI'm sitting on two right now with 34K row diff and 38K row diff, looking things over, code checking, comparing to previous submissions, checking how close sister cities are, distribution of distances to destid, etc.  You know how it goes!\r\n\r\nGood luck on your final submissions!",
      "votes": null
    },
    {
      "id": "123268",
      "postDate": "06/10/2016 18:30:56",
      "content": "<p>Great work Mattias...  congrats on the strong finish, and thanks for your generosity in sharing your ideas!</p>\n\n<p>I wish I didn't have to work today so I could explore this idea before the deadline.  My exploitation of the leak is so basic and took 10 minutes and a dozen or so lines of code, and I figured 0.865 on affected rows was good enough and I'd focus work on features for a single model.  My weeks of work on incremental features yielded little, so it seems like a missed opportunity to have not spent more time looking at exploiting those distances instead.  Always so many &quot;should haves&quot; at the end of these contests...  for me anyway.  </p>\n\n<p>One of these days I'll feel good at the end of a kaggle competition...  but not this one.  :)</p>",
      "rawMarkdown": "Great work Mattias...  congrats on the strong finish, and thanks for your generosity in sharing your ideas!\r\n\r\nI wish I didn't have to work today so I could explore this idea before the deadline.  My exploitation of the leak is so basic and took 10 minutes and a dozen or so lines of code, and I figured 0.865 on affected rows was good enough and I'd focus work on features for a single model.  My weeks of work on incremental features yielded little, so it seems like a missed opportunity to have not spent more time looking at exploiting those distances instead.  Always so many \"should haves\" at the end of these contests...  for me anyway.  \r\n\r\nOne of these days I'll feel good at the end of a kaggle competition...  but not this one.  :)",
      "votes": null
    },
    {
      "id": "123271",
      "postDate": "06/10/2016 18:58:03",
      "content": "<p>@vtKMH , I tried boosted trees (using my own XGBoost simiar implementation), naive bayes categorization (using my own implementation) where I evolved weights in attempt to increse it's power, k-means clustering (using my own implementation), some other boosted tree attempts (using my own stupid implementations). And oh, I used PCA for some of my features. Using my own implementation, of course (You say why, you say why, you say why - don't ask me why! /Eurythmics ).</p>\n\n<p><em>Nothing</em> I tried worked. I did start fairly strong by looking at the public scripts and using them as a basis for evolving my own version (using Genetic Algorithms, my own implementation of course). That worked well and temporarily got me to #10. I was later pushed to #36 and tried a million things, as stated above, but nothing helped.</p>\n\n<p>Then idle_speculation showed up, demonstrating that no small improvements would cut it. That made me look at the leak again, thinking how I could exploit it. That took <em>weeks</em>. I mean, I thought about what I could do for a very, very long time. Seems obvious in retrospect, but that's the way the cookie crumbles.</p>\n\n<p>Anyway, every competition I've been in, post facto I've though &quot;Oh, I was so close to doing that!&quot; when the winners explain what they did. Truth is, I was close to hundreds of different solutions, so luck plays a fairly strong hand in this.</p>\n\n<p>Luck favors the prepared! Good luck today and in the future!</p>",
      "rawMarkdown": "vtKMH , I tried boosted trees (using my own XGBoost simiar implementation), naive bayes categorization (using my own implementation) where I evolved weights in attempt to increse it's power, k-means clustering (using my own implementation), some other boosted tree attempts (using my own stupid implementations). And oh, I used PCA for some of my features. Using my own implementation, of course (You say why, you say why, you say why - don't ask me why! /Eurythmics ).\r\n\r\n *Nothing* I tried worked. I did start fairly strong by looking at the public scripts and using them as a basis for evolving my own version (using Genetic Algorithms, my own implementation of course). That worked well and temporarily got me to #10. I was later pushed to #36 and tried a million things, as stated above, but nothing helped.\r\n\r\nThen idle_speculation showed up, demonstrating that no small improvements would cut it. That made me look at the leak again, thinking how I could exploit it. That took *weeks*. I mean, I thought about what I could do for a very, very long time. Seems obvious in retrospect, but that's the way the cookie crumbles.\r\n\r\nAnyway, every competition I've been in, post facto I've though \"Oh, I was so close to doing that!\" when the winners explain what they did. Truth is, I was close to hundreds of different solutions, so luck plays a fairly strong hand in this.\r\n\r\nLuck favors the prepared! Good luck today and in the future!",
      "votes": null
    },
    {
      "id": "123272",
      "postDate": "06/10/2016 19:03:25",
      "content": "<p>@epiplus, turns out that the maximum distance to allow for &quot;close cities&quot; is important, but 600 miles gave better results (in testing) than 500 miles. Not very close at all! 1200 miles did worse, so there's some logic to it. LimitForFar, how far cities should be apart, is set at 0. Also gave better results than a more &quot;reasonable&quot; value, say 1000 miles, but what the heck. I'll trust the results over the theory - just this once. If my computer ever delivers up a submission for this latest set of constants.</p>\n\n<p>As for the magnitude of this frigging thing, I found 102.825.935 different City/City/SearchDestination combinations to use. That's 103 MILLION combinations. Took a while. The file for storing these is 4.33GB large...</p>",
      "rawMarkdown": "epiplus, turns out that the maximum distance to allow for \"close cities\" is important, but 600 miles gave better results (in testing) than 500 miles. Not very close at all! 1200 miles did worse, so there's some logic to it. LimitForFar, how far cities should be apart, is set at 0. Also gave better results than a more \"reasonable\" value, say 1000 miles, but what the heck. I'll trust the results over the theory - just this once. If my computer ever delivers up a submission for this latest set of constants.\r\n\r\nAs for the magnitude of this frigging thing, I found 102.825.935 different City/City/SearchDestination combinations to use. That's 103 MILLION combinations. Took a while. The file for storing these is 4.33GB large...",
      "votes": null
    },
    {
      "id": "123274",
      "postDate": "06/10/2016 19:15:16",
      "content": "<p>I am using LimitForClose=800.  My first pass where I matched by (country, region) was 1.19M rows, with max abs distance 900 miles.  With full-on data, it's over 8,000 miles. Using 800 miles cut my 11.9M row count to 5.5M, about half.  For city dist, I'm using LimitForFar=5.  My reasoning is that I saw in the data there are times when the booking is &lt;1 mile away, and after mapping sister city to leak city, throwing out negative distances, there are some with city to destid dist &lt; 0.01.  For some degree of &quot;spread&quot; amongst the distances to destid, I chose 5.</p>\n\n<p>I have one small add-on to what we are doing, that is, I do look for close city matches within 0.005 after applying sistercitydist+diff (what you called constant at one point).  This is akin to what you were saying about rounding, and goes back to my idea of radii on destid.  My theory then was that the HC locations must be static, and therefore, putting a boundary around them would be something to anchor on strongly.  So now I'm effectively doing that in this context.</p>\n\n<p>I have about 2 more hours of run time on my full-on data set (without the add-on) and then it's go time!</p>",
      "rawMarkdown": "I am using LimitForClose=800.  My first pass where I matched by (country, region) was 1.19M rows, with max abs distance 900 miles.  With full-on data, it's over 8,000 miles. Using 800 miles cut my 11.9M row count to 5.5M, about half.  For city dist, I'm using LimitForFar=5.  My reasoning is that I saw in the data there are times when the booking is <1 mile away, and after mapping sister city to leak city, throwing out negative distances, there are some with city to destid dist < 0.01.  For some degree of \"spread\" amongst the distances to destid, I chose 5.\r\n\r\nI have one small add-on to what we are doing, that is, I do look for close city matches within 0.005 after applying sistercitydist+diff (what you called constant at one point).  This is akin to what you were saying about rounding, and goes back to my idea of radii on destid.  My theory then was that the HC locations must be static, and therefore, putting a boundary around them would be something to anchor on strongly.  So now I'm effectively doing that in this context.\r\n\r\nI have about 2 more hours of run time on my full-on data set (without the add-on) and then it's go time!",
      "votes": null
    },
    {
      "id": "123277",
      "postDate": "06/10/2016 19:28:47",
      "content": "<p>I was actually able to submit my last entry in the competition - and it did worse. Which was what I was expecting. I'll try to give you an explanation here.</p>\n\n<p>Once I realized that lowering my thresholds for &quot;wide leak&quot; entries was helping my results, I knew that lowering my thresholds too much would eventually damage my results. Finding the sweet spot was the game, and time was running out.</p>\n\n<p>Here's a breakdown of the last four submissions I did.</p>\n\n<ul>\n<li>Total: 0.01394861(600000), Specific: 0.6348937(13182), Count: 6159 =&gt;\n<strong>0.50566</strong>  </li>\n</ul>\n\n<p>I used a set of validation searches (600k) to test my strategy. I used a base submission that I'd created earlier using a counting strategy similar to the public scripts. This submission had a score of 0.50566 and it of course included the leak. My new wide leak solution was injected between the leak and before my previous solution and it improved my result to 0.51051. I knew I was on the right track - but this was during the last day of the competition! What to do? Well, I had to further relax the constraints and see if I could improve the results.</p>\n\n<p>The results indicate that the score of the votes placed by this strategy was 0.63 - meaning it was extremely good but it gave to few votes. (The leak was at 0.88, so that was much better). Relaxing the constraints would increase the number of votes but decrease the quality of those votes. At some point, any new votes added by this strategy would override better votes by my base strategy - and thus reduce the results.</p>\n\n<ul>\n<li>Total: 0.02265029(600000), Specific: 0.5906204(23010), Count: 14449\n=&gt; <strong>0.51772</strong></li>\n</ul>\n\n<p>First attempt for a second entry, I was waaay to conservative. I decreased the specific score (the score of actual votes cast, ignoring entries where no votes where cast) to 0.59. It significantly improved the result, but seeing as I had only a day to work with, I should have been more aggressive.</p>\n\n<ul>\n<li>Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121\n=&gt; <strong>0.52171</strong></li>\n</ul>\n\n<p>Next attempt, I decided to go for a specific score of 0.5 and try to maximize the total score I could get from that. Remember, the total score doesn't include the leak and it doesn't include my base strategy. It just gives a null vote for any search that the wide leak doesn't give a result for. That significantly increased my results again. That was a bold move in contrast to my second move.</p>\n\n<ul>\n<li>Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530\n=&gt; <strong>0.51960</strong></li>\n</ul>\n\n<p>Here I overshot at my last attempt. I should probably have gone for a specific score of 0.45 but I ended up with a specific score of .42. The total score looked nicer, but this strategy replaced too many better results from my base strategy that it did worse. It would have been enough (right now) for a #8 place, but no better. What would a specific score of .45 done? Who knows, I may have maximized the potential of my strategy on my previous entry, but 0.42 was clearly too aggressive.</p>\n\n<p>I had so much fun in this competition, and I learned so much. Thanks to all of you who've taken the time to discuss these things with me!</p>",
      "rawMarkdown": "I was actually able to submit my last entry in the competition - and it did worse. Which was what I was expecting. I'll try to give you an explanation here.\r\n\r\nOnce I realized that lowering my thresholds for \"wide leak\" entries was helping my results, I knew that lowering my thresholds too much would eventually damage my results. Finding the sweet spot was the game, and time was running out.\r\n\r\nHere's a breakdown of the last four submissions I did.\r\n\r\n- Total: 0.01394861(600000), Specific: 0.6348937(13182), Count: 6159 =>\r\n   **0.50566**  \r\n\r\nI used a set of validation searches (600k) to test my strategy. I used a base submission that I'd created earlier using a counting strategy similar to the public scripts. This submission had a score of 0.50566 and it of course included the leak. My new wide leak solution was injected between the leak and before my previous solution and it improved my result to 0.51051. I knew I was on the right track - but this was during the last day of the competition! What to do? Well, I had to further relax the constraints and see if I could improve the results.\r\n\r\nThe results indicate that the score of the votes placed by this strategy was 0.63 - meaning it was extremely good but it gave to few votes. (The leak was at 0.88, so that was much better). Relaxing the constraints would increase the number of votes but decrease the quality of those votes. At some point, any new votes added by this strategy would override better votes by my base strategy - and thus reduce the results.\r\n\r\n- Total: 0.02265029(600000), Specific: 0.5906204(23010), Count: 14449\r\n   => **0.51772**\r\n\r\nFirst attempt for a second entry, I was waaay to conservative. I decreased the specific score (the score of actual votes cast, ignoring entries where no votes where cast) to 0.59. It significantly improved the result, but seeing as I had only a day to work with, I should have been more aggressive.\r\n\r\n- Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121\r\n   => **0.52171**\r\n\r\nNext attempt, I decided to go for a specific score of 0.5 and try to maximize the total score I could get from that. Remember, the total score doesn't include the leak and it doesn't include my base strategy. It just gives a null vote for any search that the wide leak doesn't give a result for. That significantly increased my results again. That was a bold move in contrast to my second move.\r\n\r\n- Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530\r\n   => **0.51960**\r\n\r\nHere I overshot at my last attempt. I should probably have gone for a specific score of 0.45 but I ended up with a specific score of .42. The total score looked nicer, but this strategy replaced too many better results from my base strategy that it did worse. It would have been enough (right now) for a #8 place, but no better. What would a specific score of .45 done? Who knows, I may have maximized the potential of my strategy on my previous entry, but 0.42 was clearly too aggressive.\r\n\r\nI had so much fun in this competition, and I learned so much. Thanks to all of you who've taken the time to discuss these things with me!",
      "votes": null
    },
    {
      "id": "123278",
      "postDate": "06/10/2016 19:31:16",
      "content": "<p>@eipiplus1 The best of luck to you, I'm all out of submissions, but I hope this improves your results!</p>",
      "rawMarkdown": "eipiplus1 The best of luck to you, I'm all out of submissions, but I hope this improves your results!",
      "votes": null
    },
    {
      "id": "123280",
      "postDate": "06/10/2016 20:05:54",
      "content": "<p>Awesome work, and thanks for the writeup.  I whiteboarded a hier. solution on day #1 which had: is_mobile, channel, adult_cnt, child_cnt, rm_cnt, hotel_market.  6 straight features, then did some binning, and got surprisingly good results.  (Needless to say, I dropped is_mobile quickly! And channel, too, though site_name came into play later.)  I moved straight into NB, which seemed a natural fit with counts, but no tuning of laplace, eps or threshold produced good results.  Then one round of xgboost, and nothing.  So back to layering!  Then came the date columns DT, CI , CO, then recency weighting, then tuning.  Along the way the leak was announced.  I spent a few days on that trying to get to 0.90 on it, but I landed at 0.882.  Then I tried user recom. and found that worked, with lots of tuning on click and book counts.  (This may be the downfall for many of us, we shall see!)  There are 3 layers high up which do geometrical averaging of the top 5 highest variance features I found, of 5, then 4, then 3 features. Next came many days of grid search on 2013 and 2014 data, and improvements.  Now, like you, I have this sister-city layer wedged right under leak and above 27 other hier. layers.   I've picked up 40 ticks so far, and that was with a (co, reg, c1-c2, destid) approach.  Also, I didn't until today fully utilize the (c1-c2) pairs irrespective of destid; and lastly, the full city blowout without (co, reg).  Congrats on a STRONG finish and I hope to see you near the top if I am on target with my final 3 submissions.</p>",
      "rawMarkdown": "Awesome work, and thanks for the writeup.  I whiteboarded a hier. solution on day #1 which had: is_mobile, channel, adult_cnt, child_cnt, rm_cnt, hotel_market.  6 straight features, then did some binning, and got surprisingly good results.  (Needless to say, I dropped is_mobile quickly! And channel, too, though site_name came into play later.)  I moved straight into NB, which seemed a natural fit with counts, but no tuning of laplace, eps or threshold produced good results.  Then one round of xgboost, and nothing.  So back to layering!  Then came the date columns DT, CI , CO, then recency weighting, then tuning.  Along the way the leak was announced.  I spent a few days on that trying to get to 0.90 on it, but I landed at 0.882.  Then I tried user recom. and found that worked, with lots of tuning on click and book counts.  (This may be the downfall for many of us, we shall see!)  There are 3 layers high up which do geometrical averaging of the top 5 highest variance features I found, of 5, then 4, then 3 features. Next came many days of grid search on 2013 and 2014 data, and improvements.  Now, like you, I have this sister-city layer wedged right under leak and above 27 other hier. layers.   I've picked up 40 ticks so far, and that was with a (co, reg, c1-c2, destid) approach.  Also, I didn't until today fully utilize the (c1-c2) pairs irrespective of destid; and lastly, the full city blowout without (co, reg).  Congrats on a STRONG finish and I hope to see you near the top if I am on target with my final 3 submissions.",
      "votes": null
    },
    {
      "id": "123297",
      "postDate": "06/10/2016 22:53:18",
      "content": "<p>Thanks a lot for sharing all of this and your genetic algorithms too.  Unfortunately for me, I preferred to work on something completely different rather than reusing your ideas.</p>\n\n<p>I hope you'll be on time to submit your current work in progress.  I won't have the same chance, my current WIP is due to finish sometime this WE :(</p>",
      "rawMarkdown": "Thanks a lot for sharing all of this and your genetic algorithms too.  Unfortunately for me, I preferred to work on something completely different rather than reusing your ideas.\r\n\r\nI hope you'll be on time to submit your current work in progress.  I won't have the same chance, my current WIP is due to finish sometime this WE :(",
      "votes": null
    },
    {
      "id": "123448",
      "postDate": "06/12/2016 00:14:11",
      "content": "<p>@Mattias Fagerlund, @eipiplus1, congratulations! </p>\n\n<p>Any chance to share your code/script of widening the leak? Not full code, only the widening part. Although a full code would be the best ;)) </p>",
      "rawMarkdown": "Mattias Fagerlund, @eipiplus1, congratulations! \r\n\r\nAny chance to share your code/script of widening the leak? Not full code, only the widening part. Although a full code would be the best ;))",
      "votes": null
    },
    {
      "id": "123454",
      "postDate": "06/12/2016 00:47:01",
      "content": "<p>It'll be best to see code (or snippets) from Mattias, who moved up to 6th place.  I picked up 40+ ticks working on this, but I think he claimed 473 ticks (either real or potential) at one point.  In any case, here you go!</p>\n\n<ol>\n<li>Take whatever is your leak code or file, and produce the search set of known leak keys\n(country, region, city, dist, HC).  My leak was 852K rows, but some had multiple HC.  After expanding it to 1.0M+ rows, unique it.  I ended up with 813K defined leak rows keyed as indicated earlier.</li>\n</ol>\n\n<p>Sample data:</p>\n\n<pre><code>205,354,40193,11813,1845.0739,16\n205,354,40193,11813,1851.8713,83\n205,354,40193,11813,1859.7697,0\n205,354,40193,11827,54.911,42\n205,354,40193,11827,56.8564,59\n205,354,40193,11827,56.8564,95\n205,354,40193,11827,59.9279,42\n205,354,40193,11827,60.2396,4\n</code></pre>\n\n<ol start=\"2\">\n<li><p>Create a search domain of cities for which you hope to find sister cities.  I did this in R.  I ended up not using N, but I included it.  It's very important to set the precision before saving the file.</p>\n\n<p>lg &lt;- subset(train, !is.na(orig_destination_distance)) %&gt;% group_by(user_location_country, user_location_region, user_location_city, \n                     srch_destination_id, orig_destination_distance, hotel_cluster) %&gt;% summarise(N=n())</p>\n\n<p>options(digits=8)</p>\n\n<p>write.csv(lg, &quot;ud.csv&quot;, row.names = F)</p></li>\n</ol>\n\n<blockquote>\n  \n</blockquote>\n\n<ol start=\"3\">\n<li>Load these two data sets into database tables - I started to code a search algorithm and quickly realized that SQL would tackle it much more quickly!  It's resolving 800K and 11M row data sets.</li>\n</ol>\n\n<blockquote>\n  \n</blockquote>\n\n<ol start=\"4\">\n<li>Run SQL, such as below.  This was my first attempt, later refined.  Here, I am looking for sister cities in the same country and region - therefore, the diff values were no more than 980 miles.  I did not yet consider finding cities in line with leak cities.  My result set was limited to any tuples (destid, HC, diff) with diff for cities c1&lt;c2 is distc2-distc1, having count 2+.  That is, a common distance difference diff for 2+ (destid, HC) was found for that (co, reg, c1) and (co, reg, c2) pairing.  That it is signal that if c1 is the leak city, and c2 is a non-leak city, then taking (co, reg, c1, destid, distc2-diff) will convert city c2 into a c1 lookalike.  This is what I did to pick up 40+ ticks.  The sister city lookup file for me at this point was 1.19M lines.</li>\n</ol>\n\n<blockquote>\n<pre><code>select co, reg, leakcity, sistercity, destid, diff, count(*) AS CNT\nfrom (\nselect e.co, e.reg, e.city AS leakcity, d.city AS sistercity, d.destid, e.dist AS leakdist, d.dist, e.dist - d.dist AS diff\nfrom analytics_sandbox.ekey1 e, analytics_sandbox.eud d\nwhere e.co = d.co\nand e.reg = d.reg\nand e.destid = d.destid\nand e.hc = d.hc\n-- and e.co=1-- testing\n-- and e.reg=824 -- testing\n-- and e.city=49178 -- testing\nand e.city &lt; d.city\n)\ngroup by co, reg, leakcity, sistercity, destid, diff\nhaving count(*)&gt;1\norder by co, reg, leakcity, sistercity, destid;\n</code></pre>\n</blockquote>\n\n<p>The next levels of this can be addressed by Mattias.  But clearly to me, the first was to allow for sister cities without regard to (co, reg), e.g., San Francisco could be a sister city to Edmonton, Canada.</p>\n\n<p>The distribution of counts was about 90% value 2.  Those might be considered weak sister cities, but it turns out, given the decimal precision, they were very valuable.  Using only counts 3+ eliminates so many sister cities that, when compared to my best submission at point in time, I got 110 test row prediction changes - hardly worth the effort.  Again, Mattias can comment on this.</p>\n\n<p>Lastly, when I opened the floodgates to allow sister cities to exist anywhere, I was finding sister cities 9,000 miles apart!  I put in a hard limit of max distance between cities of 800 miles, which was an arbitrary choice.  I also found cases where distc2-diff created leak distances of &lt;0.01 miles, and I put in a hard 5 mile lower limit.  I believe Mattias used 600 miles and 0 miles for these thresholds.  My base file at this point was 11.9M lines; I think Mattias mentioned a file size of over 100M lines, so at this juncture, our approaches must have differed.  He got to 6th, so take it from him at this point!</p>",
      "rawMarkdown": "It'll be best to see code (or snippets) from Mattias, who moved up to 6th place.  I picked up 40+ ticks working on this, but I think he claimed 473 ticks (either real or potential) at one point.  In any case, here you go!\r\n\r\n1. Take whatever is your leak code or file, and produce the search set of known leak keys\r\n(country, region, city, dist, HC).  My leak was 852K rows, but some had multiple HC.  After expanding it to 1.0M+ rows, unique it.  I ended up with 813K defined leak rows keyed as indicated earlier.\r\n\r\nSample data:\r\n\r\n    205,354,40193,11813,1845.0739,16\r\n    205,354,40193,11813,1851.8713,83\r\n    205,354,40193,11813,1859.7697,0\r\n    205,354,40193,11827,54.911,42\r\n    205,354,40193,11827,56.8564,59\r\n    205,354,40193,11827,56.8564,95\r\n    205,354,40193,11827,59.9279,42\r\n    205,354,40193,11827,60.2396,4\r\n\r\n2. Create a search domain of cities for which you hope to find sister cities.  I did this in R.  I ended up not using N, but I included it.  It's very important to set the precision before saving the file.\r\n\r\n    lg <- subset(train, !is.na(orig_destination_distance)) %>% group_by(user_location_country, user_location_region, user_location_city, \r\n                         srch_destination_id, orig_destination_distance, hotel_cluster) %>% summarise(N=n())\r\n\r\n    options(digits=8)\r\n\r\n    write.csv(lg, \"ud.csv\", row.names = F)\r\n\r\n\r\n\r\n> \r\n\r\n3. Load these two data sets into database tables - I started to code a search algorithm and quickly realized that SQL would tackle it much more quickly!  It's resolving 800K and 11M row data sets.\r\n\r\n> \r\n\r\n4. Run SQL, such as below.  This was my first attempt, later refined.  Here, I am looking for sister cities in the same country and region - therefore, the diff values were no more than 980 miles.  I did not yet consider finding cities in line with leak cities.  My result set was limited to any tuples (destid, HC, diff) with diff for cities c1&lt;c2 is distc2-distc1, having count 2+.  That is, a common distance difference diff for 2+ (destid, HC) was found for that (co, reg, c1) and (co, reg, c2) pairing.  That it is signal that if c1 is the leak city, and c2 is a non-leak city, then taking (co, reg, c1, destid, distc2-diff) will convert city c2 into a c1 lookalike.  This is what I did to pick up 40+ ticks.  The sister city lookup file for me at this point was 1.19M lines.\r\n\r\n\r\n>     select co, reg, leakcity, sistercity, destid, diff, count(*) AS CNT\r\n>     from (\r\n>     select e.co, e.reg, e.city AS leakcity, d.city AS sistercity, d.destid, e.dist AS leakdist, d.dist, e.dist - d.dist AS diff\r\n>     from analytics_sandbox.ekey1 e, analytics_sandbox.eud d\r\n>     where e.co = d.co\r\n>     and e.reg = d.reg\r\n>     and e.destid = d.destid\r\n>     and e.hc = d.hc\r\n>     -- and e.co=1-- testing\r\n>     -- and e.reg=824 -- testing\r\n>     -- and e.city=49178 -- testing\r\n>     and e.city < d.city\r\n>     )\r\n>     group by co, reg, leakcity, sistercity, destid, diff\r\n>     having count(*)>1\r\n>     order by co, reg, leakcity, sistercity, destid;\r\n\r\n\r\nThe next levels of this can be addressed by Mattias.  But clearly to me, the first was to allow for sister cities without regard to (co, reg), e.g., San Francisco could be a sister city to Edmonton, Canada.\r\n\r\nThe distribution of counts was about 90% value 2.  Those might be considered weak sister cities, but it turns out, given the decimal precision, they were very valuable.  Using only counts 3+ eliminates so many sister cities that, when compared to my best submission at point in time, I got 110 test row prediction changes - hardly worth the effort.  Again, Mattias can comment on this.\r\n\r\nLastly, when I opened the floodgates to allow sister cities to exist anywhere, I was finding sister cities 9,000 miles apart!  I put in a hard limit of max distance between cities of 800 miles, which was an arbitrary choice.  I also found cases where distc2-diff created leak distances of <0.01 miles, and I put in a hard 5 mile lower limit.  I believe Mattias used 600 miles and 0 miles for these thresholds.  My base file at this point was 11.9M lines; I think Mattias mentioned a file size of over 100M lines, so at this juncture, our approaches must have differed.  He got to 6th, so take it from him at this point!",
      "votes": null
    },
    {
      "id": "123456",
      "postDate": "06/12/2016 01:23:02",
      "content": "<p>@eipiplus1, thank you so much. Looking forward to learn Mattias' detailed approach :) </p>",
      "rawMarkdown": "eipiplus1, thank you so much. Looking forward to learn Mattias' detailed approach :)",
      "votes": null
    },
    {
      "id": "123483",
      "postDate": "06/12/2016 07:08:45",
      "content": "<p>Hi, my best set of constants brought my base submission at 0.50566 (in PuLB) to 0.52171 (in PuLB) meaning an increase of 0.01605. In the PriLB I ended up with 0.51881 and I was very far from the #5 position, so I probably couldn't have made it there.</p>\n\n<p>Anyway, my last set of constants, as outlined in a previous post, did significantly worse. I stupidly didn't tag that version of my code - you should always tag your code once you make a better submission so you can easily revert to that once you do worse. Turns out that the exact constants were lost, but I was able to approximate the result.</p>\n\n<p>You won't be able to run it, because it's missing large parts. One really important part of this, I think, is that I was able to evaluate a new set of constants in 14 seconds. @epiplus1, I'm guessing your SQL version didn't allow for that? Using 600k rows instead of the full dataset I computed how well the wide leak would all on its own. And with a 14 second turnaround, that allowed me to continuously tweak parameters  trying to find a good set.</p>\n\n<p>Here are the parameters;</p>\n\n<pre><code>    public static int MaxRows = 600 * 1000;\n    //public static int MaxRows = int.MaxValue;\n    public float LimitForClose = 230;\n    public float LimitForFar = 30;\n    public int MinimumHits = 1;\n    public float Epsilon = 0.03f;\n    public float MatchEpsilon = 0.031f;\n</code></pre>\n\n<p>I'll try to figure out how to publish the code. It's not pretty...</p>",
      "rawMarkdown": "Hi, my best set of constants brought my base submission at 0.50566 (in PuLB) to 0.52171 (in PuLB) meaning an increase of 0.01605. In the PriLB I ended up with 0.51881 and I was very far from the #5 position, so I probably couldn't have made it there.\r\n\r\nAnyway, my last set of constants, as outlined in a previous post, did significantly worse. I stupidly didn't tag that version of my code - you should always tag your code once you make a better submission so you can easily revert to that once you do worse. Turns out that the exact constants were lost, but I was able to approximate the result.\r\n\r\nYou won't be able to run it, because it's missing large parts. One really important part of this, I think, is that I was able to evaluate a new set of constants in 14 seconds. @epiplus1, I'm guessing your SQL version didn't allow for that? Using 600k rows instead of the full dataset I computed how well the wide leak would all on its own. And with a 14 second turnaround, that allowed me to continuously tweak parameters  trying to find a good set.\r\n\r\nHere are the parameters;\r\n\r\n        public static int MaxRows = 600 * 1000;\r\n        //public static int MaxRows = int.MaxValue;\r\n        public float LimitForClose = 230;\r\n        public float LimitForFar = 30;\r\n        public int MinimumHits = 1;\r\n        public float Epsilon = 0.03f;\r\n        public float MatchEpsilon = 0.031f;\r\n\r\nI'll try to figure out how to publish the code. It's not pretty...",
      "votes": null
    },
    {
      "id": "123485",
      "postDate": "06/12/2016 07:15:58",
      "content": "<p>I couldn't actually post the code. Very strange, I got an error message. Maybe the code was too long. I've posted it on my blog instead. I doubt very much you'll get much from it, but there it is. It uses a ton of proprietary software that isn't public, but the algorithm is all in there.</p>\n\n<p><a href=\"https://lotsacode.wordpress.com/2016/06/12/code-from-expedia-hotel-recommendations/\">https://lotsacode.wordpress.com/2016/06/12/code-from-expedia-hotel-recommendations/</a></p>\n\n<p>/m</p>",
      "rawMarkdown": "I couldn't actually post the code. Very strange, I got an error message. Maybe the code was too long. I've posted it on my blog instead. I doubt very much you'll get much from it, but there it is. It uses a ton of proprietary software that isn't public, but the algorithm is all in there.\r\n\r\nhttps://lotsacode.wordpress.com/2016/06/12/code-from-expedia-hotel-recommendations/\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "123522",
      "postDate": "06/12/2016 14:26:08",
      "content": "<p>@Mattias, thank you very much for the code. </p>",
      "rawMarkdown": "Mattias, thank you very much for the code.",
      "votes": null
    },
    {
      "id": "123527",
      "postDate": "06/12/2016 14:38:34",
      "content": "<p>The SQL itself ran relatively quickly, under one minute, but you're right, I didn't consider thresholds until back in my code.  So I didn't do tuning in SQL. I fixed my mind on a confirmation constant difference for each pair of cities, but now that I think about it more, even one match apparently is really valuable. I think going from N=2+ (giving me 11.9M before constraints) to N=1+ would get to 100M+ resultant rows to match on and would, based on what you said and your results, do better than other algorithms. Since it was the last day, and my code was taking about 2 hours to check through 5 million rows, after applying constraints, I would have had no hope to try N=1 anyhow.</p>",
      "rawMarkdown": "The SQL itself ran relatively quickly, under one minute, but you're right, I didn't consider thresholds until back in my code.  So I didn't do tuning in SQL. I fixed my mind on a confirmation constant difference for each pair of cities, but now that I think about it more, even one match apparently is really valuable. I think going from N=2+ (giving me 11.9M before constraints) to N=1+ would get to 100M+ resultant rows to match on and would, based on what you said and your results, do better than other algorithms. Since it was the last day, and my code was taking about 2 hours to check through 5 million rows, after applying constraints, I would have had no hope to try N=1 anyhow.",
      "votes": null
    },
    {
      "id": "123612",
      "postDate": "06/13/2016 05:26:06",
      "content": "<p>Thanks guys for sharing the journey with me, it was really fun to work on this at the very last couple of hours and bouncing ideas back and forth!</p>",
      "rawMarkdown": "Thanks guys for sharing the journey with me, it was really fun to work on this at the very last couple of hours and bouncing ideas back and forth!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 123021,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/09/2016 07:25:22",
      "content": "<p>I tried something similar, but not quite the same.  For a given (user_country, user_city, hotel_country, destid) I looked at the pairs (train_distance, HC) in relation, for all the training data - call it D.  Then for test, I explored those keys (user_country, user_city, hotel_country, destid) at test_distance and noted all the matching D entries where abs(test_distance - train_distance)\n\n</p><p>These sets of HC at 1 mile are sometimes 60+!  At distance 0.1 mile, there are even some 10+ counts.  I focused in on cases where it was HC set sizes of 1 or 2 only.  I compared the HC sets thus found with my best submission (out to N places), and these D HC often were at positions 6+ in my lists.</p>\n\n<p>The reason to focus on set sizes 1 or 2 is that I wouldn't put much faith in the utility for a larger set of HC for, say, Chicago or New York City or Frankfurt or Beijing or wherever.  On the flip side, that means I may be limiting to HC sets for smaller destinations with fewer booked trips and, with less training data, higher variance if applied to predictions for test.</p>\n\n<p>I purposely left out hotel_market, but it may be worth trying.  </p>\n\n<p>There was one catch to all this for me - I saved data from R to file, and didn't set options(digits=6) or such.  So, the distances I used in my coding was rounded to 3 digits.  The problem with this is that I was awarding a huge weight to exact matches, where abs(x-y)&lt;0.000001 (float precision), and therefore though there were certainly many exact matches in my data, when the places were 3 or less, I definitely labeled many of the rows as INEXACT when they were exact.</p>\n\n<p>Does this idea have some merit?  It's not what you described, but on the same train of thought.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123025,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/09/2016 07:53:56",
      "content": "<p>I think that that would pick up on the leak, but weaken it. The leak is that any combination of [orig_destination_distance, srch_destination_id_base, user_location_city] typically has  a unique hotel cluster - because it identifies one exact hotel.</p>\n\n<p>You're looking at [user_country, user_city, hotel_country, destid]. But user_country is tied to user_city and hotel_country is mostly tied to destid. You're not grouping on the exact orig_destination_distance, but a soft grouping. That means that on top of the exact matches (the leak) you'll get a number of false matches.</p>\n\n<p>To tell the truth, I also tried this, looking at hotels that don't have exact distance matches, but matching up to some range. The larger I made the range, the weaker the feature became. Leaving the range at 0 (which then becomes the leak) worked best in my attempts.</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123055,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/09/2016 12:21:27",
      "content": "<p>I can now report that this technique improved my score by 0.00437, so; yay!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123067,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/09/2016 13:50:39",
      "content": "<p>Nice going! So close cities with constant differences has some value, good find. If I have time, I might try that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123135,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/09/2016 22:40:33",
      "content": "<p><em>Take city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.</em></p>\n\n<p>@Mattias Fagerlund Why I didn't find the pattern in this example.It's seems in each pair, the orig_destination_distance are all different and the difference between them is also not -5.061.</p>\n\n<p>@eipiplus1   Does you approach work? (abs(x-y)&lt;0.000001 ). So tiny difference will cause mismatch?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123153,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 00:16:00",
      "content": "<p>@FengLi</p>\n\n<p>Here is the data per his example:</p>\n\n<blockquote>\n  <p>49178 v 55729 - 8745 - 64 - 10555.076 or 10550.015 --&gt; 5.0610 </p>\n  \n  <p>49178 v 55729 - 8745 - 64 - 10555.076 or 10550.2091 --&gt; 4.8669 </p>\n  \n  <p>49178 v 55729 - 8745 - 64 - 10555.076 or 10550.8242 --&gt; 4.2518 </p>\n  \n  <p>49178 v 55729 - 8745 - 8 - 10556.3596 or 10548.0258 --&gt; 8.3338 </p>\n  \n  <p>49178 v 55729 - 8745 - 97 - 10555.228 or 10550.167 --&gt; 5.0610</p>\n</blockquote>\n\n<p>This is data for country=77, region=824, showing c1 v c2.  All data is for destid=8745 as you can see.\nThe HCs are 64, 8, and 97.  Note that the distance changed for HC 64, for whatever reason.</p>\n\n<p>I can confirm the 5.0610 difference, seen exactly twice here.  </p>\n\n<p>I take this to mean that for a pair of cities (c1, c2), if there exists a distance difference DD which repeats 2+ times, this is a sign that they may be sister cities (as I'm calling them).  In this case, we have 2/1/1/1 when grouping by DD.</p>\n\n<p>@Mattias, I'm looking at this now in terms of 3 thresholds: (1) DD&gt;=N, some critical value, like N=2 or N=3 or higher - it turns out, there are a whole lot of sister cities with at least one DD=2; or, the density of non-1 DD counts amongst all DD counts is at least X% (2) DD&lt;=A miles, that is, the sister cities must be no more than A miles apart (3) leakcity to destid distance&gt;=B miles, that is, the destid must be at least B miles from the leakcity.  I have tried N=2, A=100, and B=1000 so far.  Perhaps sister cities should be considered as closer than that.  With B=1000, I'm trying to avoid &quot;local&quot; trips, some of which are hotel bookings within a 5 mile radius.  I haven't tuned N/A/B yet, these are just tester values.  What do you think?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123154,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 00:20:13",
      "content": "<p>@FengLi</p>\n\n<p>At this point in the game, I'm inclined to say my approach won't work.  I still hold hope for it, but with 1 day left, I won't have to time to further explore.  What Mattias is describing makes use of the leak heuristic, which is proven, while being exploratory about cities; what I'm doing is unproven and, while it may turn out to be the big secret, I'm sure it would require a lot more diligence to figure out exactly how to make use of the radii mapped destid/HC.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123160,
      "author_name": "kaukau",
      "author_url": "",
      "post_date": "06/10/2016 02:00:25",
      "content": "<p>@Mattias Fagerlund - Are you saying that the processes is as follows:\n1. cluster rows by user_location_country, user_location_region, srch_destination_id.\n2. Given some example that i want to make a prediction on:\nTake it's orig_destination_distance, compare it to all other groups.\n3. You get a list of values. \n4. What is the next step? </p>\n\n<p>Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123178,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 05:39:31",
      "content": "<p>@epiplus1 - that sounds about right. </p>\n\n<p>I find cities that are close and that are considered for &quot;sister city&quot; testing. I then find potential target search destinations that are far away (both sisters must have hotels there). From both sister cities. I then start to try to work out if they have any N hotels that have the same hotel cluster and the same distace - given some fixed constant between them. </p>\n\n<p>I'm not able to find many, here's one run on my validation set;\nTotal: 0.03619763(1976304), Specific: 0.4383184(163209), Count: 200585</p>\n\n<p>So <em>at best</em> the result would improve by 0.036 - but a lot of those would be in the regular leak and so there would be no additional benefits to using this method.</p>\n\n<p>For that run, near limit=30 miles, far limit=150 miles, N (min number of matched hotels) was 2, and the city comparison epsilon was 0.01.</p>\n\n<p>These values can be tweaked to get more hits but of lower quality. They can be tweaked the other way to give fewer hits, but of greater quality. This particular run found 200585 hits and got a score of 0.438 on those. Now, 0.438 is better than any other feature I've found with the exception of the leak. so that's good. But the fact that it creates so few votes means it's not super helpful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123179,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 06:00:26",
      "content": "<p>@kaukau, did that answer your question? There's 750 lines of code to do this, so it's not trivial, but I think it's an interesting idea!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123184,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 06:15:38",
      "content": "<p>Thanks for the explanation and analysis.  I implemented two different versions of this, based on whether I strictly included the leakcity leakdist as part of my key, and here is what I found.  Starting with the 852K leak data rows, I expanded out the HC map and then reduced the final key list uniquely.  So that's the leak city search domain.  Then, I searched across the entire (country, region, destid) universe for any cities c1!=c2, computing the diff(dist) per HC, and then aggregating.  I ended up with 1.19M sister rows with leakdist, 964K sister rows without leakdist.  The advantage of the 964K row set is not having to compute sistercitydist+diff, I have the leakdist at hand.</p>\n\n<p>Using the 964K set, I tried twice and got worse LB scores.  I used the 1.19M data set and improved my LB score.  It requires calculating sistercitydist+diff, as I mentioned.  I got 3,100 test row diff in that submission, and it proved HIGHLY effective.</p>\n\n<p>I found that I missed many sister cities due to float imprecison.  Now that I'm using numeric(12,4) I'm getting toward 7,500 matches.</p>\n\n<p>To go one step further, I'm going to eventually test nearest neighbor cities constrained to a diff tolerance of 0.0005 or thereabouts.  Have you tried that?  It is not the leak heuristic, but it may prove out.  It's effectively going back to my radius idea, but at a WAY smaller value!</p>\n\n<p>By the way, when it came to crunching the 1.19M leak domain data set against the 11.5M row search city data set, I wrote in SQL.  That flew!  It was the most natural way to join and process this data, in my view.</p>\n\n<p>I'm still working on it and measuring the efficacy of my approach.  Thanks again for sharing ideas and congrats for your place on the LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123216,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 11:59:09",
      "content": "<p>Thanks, I was able to improve my score some more using this technique (by 0.00048, meaning 0.00485 from this technique alone), but that was only enough to move up one position. </p>\n\n<p>As I was typing this, I made another submission that improved by another 0.00721, so that's 0,01206 from this technique. Currently at #11.</p>\n\n<p>I feel I'm running out of time here, very few submissions left to go. I'm thinking I could probably eek a bit more performance out of this, but I'm going to have to let it go.</p>\n\n<p>A further cool thing I would have liked to have done was to create a map given the city-hotel distances. This is surprisingly easy for 2d coordinates, you just keep moving the nodes (cities and hotels) around to satisfy their internal distance demands. Takes a number of iterations, but the math is trivial and fast. But on a sphere, it's not so easy, because you have to place the nodes using longitudes and latitudes and you have to handle wrap-around issues and so on. Very messy.</p>\n\n<p>But once you have a map, you don't even need sister cities to both index the same hotels, with the map, you can search the map for hotels with a distance of X from city Y. But alas, there isn't enough time left for that - and it's a fairly low probability that I would have been able to get sufficient precision on the map coordinates.</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123223,
      "author_name": "kaukau",
      "author_url": "",
      "post_date": "06/10/2016 12:38:27",
      "content": "<p>@Mattias Fagerlund - How did you find cities that are close?\nI mean if i group by country, region and dest, and after that check all the all possible paris of\norig_di between one city and all other in this group i will get a lot of combinations?\nHow did you handle that?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123225,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 13:08:39",
      "content": "<p>Actually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. </p>\n\n<p>You can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping &quot;near&quot; lists. </p>\n\n<p>If the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).</p>\n\n<p>But as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. </p>\n\n<p>For instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.</p>\n\n<p>So you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123228,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/10/2016 13:25:44",
      "content": "<p>Very interesting Mattias, thanks for sharing. I would not be surprised if the solution of idle_speculation is related to what you described in your previous post - find all locations and extend the leak to all records with a distance in the test set. I have looked for ways to do that, but without external data, sufficient precision is quite a challenge like you state.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123230,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/10/2016 13:33:04",
      "content": "<p>I have been working on this for a few days as well - this is a tough challange to map all these distances - good work Mattias. Few more days and I think you would get into top3 :) I am highly certain idle_specaulation did this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123233,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 14:01:23",
      "content": "<p>Yikes. One more submission to go, and I just reached #6.</p>\n\n<p><a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/leaderboard?submissionId=3150212\">https://www.kaggle.com/c/expedia-hotel-recommendations/leaderboard?submissionId=3150212</a></p>\n\n<p>Best placement ever for me ever on Kaggle, let's see if I can hang on. But this thread describes what I've been doing; I'm basically relaxing my requirements to increase the number of votes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123234,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/10/2016 14:06:19",
      "content": "<p>wish you best of luck. with a careful submission top3 is yours :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123253,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 17:22:52",
      "content": "<p>Great work!  Congrats on reaching #6.  You were using &quot;strongly match&quot; as N&gt;=2 for number of same distances to HC in a single destination.  Are you using N&gt;=3 or more now?  My dist. is:</p>\n\n<blockquote>\n  <p>10679303  2  </p>\n  \n  <p>335235  3   </p>\n  \n  <p>25042  4    </p>\n  \n  <p>3840  5    </p>\n  \n  <p>1102  6</p>\n  \n  <p>507  7</p>\n  \n  <p>297  8</p>\n  \n  <p>217  9</p>\n  \n  <p>117  10</p>\n</blockquote>\n\n<p>so there is quite a drop from 2 to 3.  I am still using 2 as the qualify cutoff.  The thing is, I know that when I go through the search space of (HC, dist) for a given destid, where there can be and often are multiple dist for a given HC, that the probability of a match on dist+diff to leakcity dist is very, very small.  So using N&gt;=2 seems ok to me, or, in fact, may be missing out on lots of matches if I don't include it. </p>\n\n<p>[quote=Mattias Fagerlund;123225]</p>\n\n<p>Actually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. </p>\n\n<p>You can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping &quot;near&quot; lists. </p>\n\n<p>If the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).</p>\n\n<p>But as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. </p>\n\n<p>For instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.</p>\n\n<p>So you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123261,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 18:10:19",
      "content": "<p>For full transparancy, I'll reveal my &quot;final&quot; constants. There's no guarantee that these are the best, but they worked well for me in testing. Also, I'm afraid I won't be able to make a subission using these constants, because so far the system has been working for hours without progress. </p>\n\n<pre><code>    public float LimitForClose = 600;\n    public float LimitForFar = 0;\n    public int MaxDistanceDifferenceFactor = 200000;\n    public int ItemsToCheckPerCity = 40000;\n    public int MaxDistancesToCheck = 5000;\n    public int MinimumHits = 1;\n    public float Epsilon = 0.05f;//0.00001f;\n    public float MatchEpsilon = 0.05f;\n    public int MaxConstantsPerDestination = 750;\n    public float MaxSearchDestinationSize = 1000;\n</code></pre>\n\n<p>But if things holds true, and if I'm able to make a submission, it should be significantly better than my previous submissions.</p>\n\n<p>Testing for my best submission:\n  Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121 =&gt; 0.52171</p>\n\n<p>Testing for the submission I'm computing now:\n  Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530</p>\n\n<p>Now competition for #1, but it should result in better scores than my previous attempt. Problem though, with the previous entry it took 1:26 to create a submission (hour and a half). This submission looks to be about 3 times as complicated. So at 4 hours+, I may be out of time to even submit it. Oh well, that's the game, isn't it?</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123264,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/10/2016 18:24:43",
      "content": "<p>anyway, share your algo after the competition, very interesting to see what have you done :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123265,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 18:27:42",
      "content": "<p>@epiplus I'm no longer sure if I used 1 or 2 for my latest submission - probably 1 though. It decreases the quality but it increases the number of votes. On the whole, it seems better. But our mileages may vary though.</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123267,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 18:30:08",
      "content": "<p>LimitForClose = max dist between city1 and city2? min dist from leakcity to destid?</p>\n\n<p>I didn't put a limit on max distances to check, that's an interesting one to consider for speeding things up.  I do have an early check for if sistercitydist+diff&lt;=0, in which case a diff was added to the base sister city and results in negative dist to destid, which of course is impossible.</p>\n\n<p>I exploded my city-city table to not be constrained by country/region, and now I'm running with that.  It was that, or create a key for (country, region, city) and do the Cartesian-ish product on that joining still by destid and HC.  So now my domain search space is 11.9M rows.</p>\n\n<p>I'm sitting on two right now with 34K row diff and 38K row diff, looking things over, code checking, comparing to previous submissions, checking how close sister cities are, distribution of distances to destid, etc.  You know how it goes!</p>\n\n<p>Good luck on your final submissions!  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123268,
      "author_name": "kevinhinson",
      "author_url": "",
      "post_date": "06/10/2016 18:30:56",
      "content": "<p>Great work Mattias...  congrats on the strong finish, and thanks for your generosity in sharing your ideas!</p>\n\n<p>I wish I didn't have to work today so I could explore this idea before the deadline.  My exploitation of the leak is so basic and took 10 minutes and a dozen or so lines of code, and I figured 0.865 on affected rows was good enough and I'd focus work on features for a single model.  My weeks of work on incremental features yielded little, so it seems like a missed opportunity to have not spent more time looking at exploiting those distances instead.  Always so many &quot;should haves&quot; at the end of these contests...  for me anyway.  </p>\n\n<p>One of these days I'll feel good at the end of a kaggle competition...  but not this one.  :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123271,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 18:58:03",
      "content": "<p>@vtKMH , I tried boosted trees (using my own XGBoost simiar implementation), naive bayes categorization (using my own implementation) where I evolved weights in attempt to increse it's power, k-means clustering (using my own implementation), some other boosted tree attempts (using my own stupid implementations). And oh, I used PCA for some of my features. Using my own implementation, of course (You say why, you say why, you say why - don't ask me why! /Eurythmics ).</p>\n\n<p><em>Nothing</em> I tried worked. I did start fairly strong by looking at the public scripts and using them as a basis for evolving my own version (using Genetic Algorithms, my own implementation of course). That worked well and temporarily got me to #10. I was later pushed to #36 and tried a million things, as stated above, but nothing helped.</p>\n\n<p>Then idle_speculation showed up, demonstrating that no small improvements would cut it. That made me look at the leak again, thinking how I could exploit it. That took <em>weeks</em>. I mean, I thought about what I could do for a very, very long time. Seems obvious in retrospect, but that's the way the cookie crumbles.</p>\n\n<p>Anyway, every competition I've been in, post facto I've though &quot;Oh, I was so close to doing that!&quot; when the winners explain what they did. Truth is, I was close to hundreds of different solutions, so luck plays a fairly strong hand in this.</p>\n\n<p>Luck favors the prepared! Good luck today and in the future!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123272,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 19:03:25",
      "content": "<p>@epiplus, turns out that the maximum distance to allow for &quot;close cities&quot; is important, but 600 miles gave better results (in testing) than 500 miles. Not very close at all! 1200 miles did worse, so there's some logic to it. LimitForFar, how far cities should be apart, is set at 0. Also gave better results than a more &quot;reasonable&quot; value, say 1000 miles, but what the heck. I'll trust the results over the theory - just this once. If my computer ever delivers up a submission for this latest set of constants.</p>\n\n<p>As for the magnitude of this frigging thing, I found 102.825.935 different City/City/SearchDestination combinations to use. That's 103 MILLION combinations. Took a while. The file for storing these is 4.33GB large...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123274,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 19:15:16",
      "content": "<p>I am using LimitForClose=800.  My first pass where I matched by (country, region) was 1.19M rows, with max abs distance 900 miles.  With full-on data, it's over 8,000 miles. Using 800 miles cut my 11.9M row count to 5.5M, about half.  For city dist, I'm using LimitForFar=5.  My reasoning is that I saw in the data there are times when the booking is &lt;1 mile away, and after mapping sister city to leak city, throwing out negative distances, there are some with city to destid dist &lt; 0.01.  For some degree of &quot;spread&quot; amongst the distances to destid, I chose 5.</p>\n\n<p>I have one small add-on to what we are doing, that is, I do look for close city matches within 0.005 after applying sistercitydist+diff (what you called constant at one point).  This is akin to what you were saying about rounding, and goes back to my idea of radii on destid.  My theory then was that the HC locations must be static, and therefore, putting a boundary around them would be something to anchor on strongly.  So now I'm effectively doing that in this context.</p>\n\n<p>I have about 2 more hours of run time on my full-on data set (without the add-on) and then it's go time!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123277,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 19:28:47",
      "content": "<p>I was actually able to submit my last entry in the competition - and it did worse. Which was what I was expecting. I'll try to give you an explanation here.</p>\n\n<p>Once I realized that lowering my thresholds for &quot;wide leak&quot; entries was helping my results, I knew that lowering my thresholds too much would eventually damage my results. Finding the sweet spot was the game, and time was running out.</p>\n\n<p>Here's a breakdown of the last four submissions I did.</p>\n\n<ul>\n<li>Total: 0.01394861(600000), Specific: 0.6348937(13182), Count: 6159 =&gt;\n<strong>0.50566</strong>  </li>\n</ul>\n\n<p>I used a set of validation searches (600k) to test my strategy. I used a base submission that I'd created earlier using a counting strategy similar to the public scripts. This submission had a score of 0.50566 and it of course included the leak. My new wide leak solution was injected between the leak and before my previous solution and it improved my result to 0.51051. I knew I was on the right track - but this was during the last day of the competition! What to do? Well, I had to further relax the constraints and see if I could improve the results.</p>\n\n<p>The results indicate that the score of the votes placed by this strategy was 0.63 - meaning it was extremely good but it gave to few votes. (The leak was at 0.88, so that was much better). Relaxing the constraints would increase the number of votes but decrease the quality of those votes. At some point, any new votes added by this strategy would override better votes by my base strategy - and thus reduce the results.</p>\n\n<ul>\n<li>Total: 0.02265029(600000), Specific: 0.5906204(23010), Count: 14449\n=&gt; <strong>0.51772</strong></li>\n</ul>\n\n<p>First attempt for a second entry, I was waaay to conservative. I decreased the specific score (the score of actual votes cast, ignoring entries where no votes where cast) to 0.59. It significantly improved the result, but seeing as I had only a day to work with, I should have been more aggressive.</p>\n\n<ul>\n<li>Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121\n=&gt; <strong>0.52171</strong></li>\n</ul>\n\n<p>Next attempt, I decided to go for a specific score of 0.5 and try to maximize the total score I could get from that. Remember, the total score doesn't include the leak and it doesn't include my base strategy. It just gives a null vote for any search that the wide leak doesn't give a result for. That significantly increased my results again. That was a bold move in contrast to my second move.</p>\n\n<ul>\n<li>Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530\n=&gt; <strong>0.51960</strong></li>\n</ul>\n\n<p>Here I overshot at my last attempt. I should probably have gone for a specific score of 0.45 but I ended up with a specific score of .42. The total score looked nicer, but this strategy replaced too many better results from my base strategy that it did worse. It would have been enough (right now) for a #8 place, but no better. What would a specific score of .45 done? Who knows, I may have maximized the potential of my strategy on my previous entry, but 0.42 was clearly too aggressive.</p>\n\n<p>I had so much fun in this competition, and I learned so much. Thanks to all of you who've taken the time to discuss these things with me!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123278,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/10/2016 19:31:16",
      "content": "<p>@eipiplus1 The best of luck to you, I'm all out of submissions, but I hope this improves your results!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123280,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/10/2016 20:05:54",
      "content": "<p>Awesome work, and thanks for the writeup.  I whiteboarded a hier. solution on day #1 which had: is_mobile, channel, adult_cnt, child_cnt, rm_cnt, hotel_market.  6 straight features, then did some binning, and got surprisingly good results.  (Needless to say, I dropped is_mobile quickly! And channel, too, though site_name came into play later.)  I moved straight into NB, which seemed a natural fit with counts, but no tuning of laplace, eps or threshold produced good results.  Then one round of xgboost, and nothing.  So back to layering!  Then came the date columns DT, CI , CO, then recency weighting, then tuning.  Along the way the leak was announced.  I spent a few days on that trying to get to 0.90 on it, but I landed at 0.882.  Then I tried user recom. and found that worked, with lots of tuning on click and book counts.  (This may be the downfall for many of us, we shall see!)  There are 3 layers high up which do geometrical averaging of the top 5 highest variance features I found, of 5, then 4, then 3 features. Next came many days of grid search on 2013 and 2014 data, and improvements.  Now, like you, I have this sister-city layer wedged right under leak and above 27 other hier. layers.   I've picked up 40 ticks so far, and that was with a (co, reg, c1-c2, destid) approach.  Also, I didn't until today fully utilize the (c1-c2) pairs irrespective of destid; and lastly, the full city blowout without (co, reg).  Congrats on a STRONG finish and I hope to see you near the top if I am on target with my final 3 submissions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123297,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/10/2016 22:53:18",
      "content": "<p>Thanks a lot for sharing all of this and your genetic algorithms too.  Unfortunately for me, I preferred to work on something completely different rather than reusing your ideas.</p>\n\n<p>I hope you'll be on time to submit your current work in progress.  I won't have the same chance, my current WIP is due to finish sometime this WE :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123448,
      "author_name": "aldente",
      "author_url": "",
      "post_date": "06/12/2016 00:14:11",
      "content": "<p>@Mattias Fagerlund, @eipiplus1, congratulations! </p>\n\n<p>Any chance to share your code/script of widening the leak? Not full code, only the widening part. Although a full code would be the best ;)) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123454,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/12/2016 00:47:01",
      "content": "<p>It'll be best to see code (or snippets) from Mattias, who moved up to 6th place.  I picked up 40+ ticks working on this, but I think he claimed 473 ticks (either real or potential) at one point.  In any case, here you go!</p>\n\n<ol>\n<li>Take whatever is your leak code or file, and produce the search set of known leak keys\n(country, region, city, dist, HC).  My leak was 852K rows, but some had multiple HC.  After expanding it to 1.0M+ rows, unique it.  I ended up with 813K defined leak rows keyed as indicated earlier.</li>\n</ol>\n\n<p>Sample data:</p>\n\n<pre><code>205,354,40193,11813,1845.0739,16\n205,354,40193,11813,1851.8713,83\n205,354,40193,11813,1859.7697,0\n205,354,40193,11827,54.911,42\n205,354,40193,11827,56.8564,59\n205,354,40193,11827,56.8564,95\n205,354,40193,11827,59.9279,42\n205,354,40193,11827,60.2396,4\n</code></pre>\n\n<ol start=\"2\">\n<li><p>Create a search domain of cities for which you hope to find sister cities.  I did this in R.  I ended up not using N, but I included it.  It's very important to set the precision before saving the file.</p>\n\n<p>lg &lt;- subset(train, !is.na(orig_destination_distance)) %&gt;% group_by(user_location_country, user_location_region, user_location_city, \n                     srch_destination_id, orig_destination_distance, hotel_cluster) %&gt;% summarise(N=n())</p>\n\n<p>options(digits=8)</p>\n\n<p>write.csv(lg, &quot;ud.csv&quot;, row.names = F)</p></li>\n</ol>\n\n<blockquote>\n  \n</blockquote>\n\n<ol start=\"3\">\n<li>Load these two data sets into database tables - I started to code a search algorithm and quickly realized that SQL would tackle it much more quickly!  It's resolving 800K and 11M row data sets.</li>\n</ol>\n\n<blockquote>\n  \n</blockquote>\n\n<ol start=\"4\">\n<li>Run SQL, such as below.  This was my first attempt, later refined.  Here, I am looking for sister cities in the same country and region - therefore, the diff values were no more than 980 miles.  I did not yet consider finding cities in line with leak cities.  My result set was limited to any tuples (destid, HC, diff) with diff for cities c1&lt;c2 is distc2-distc1, having count 2+.  That is, a common distance difference diff for 2+ (destid, HC) was found for that (co, reg, c1) and (co, reg, c2) pairing.  That it is signal that if c1 is the leak city, and c2 is a non-leak city, then taking (co, reg, c1, destid, distc2-diff) will convert city c2 into a c1 lookalike.  This is what I did to pick up 40+ ticks.  The sister city lookup file for me at this point was 1.19M lines.</li>\n</ol>\n\n<blockquote>\n<pre><code>select co, reg, leakcity, sistercity, destid, diff, count(*) AS CNT\nfrom (\nselect e.co, e.reg, e.city AS leakcity, d.city AS sistercity, d.destid, e.dist AS leakdist, d.dist, e.dist - d.dist AS diff\nfrom analytics_sandbox.ekey1 e, analytics_sandbox.eud d\nwhere e.co = d.co\nand e.reg = d.reg\nand e.destid = d.destid\nand e.hc = d.hc\n-- and e.co=1-- testing\n-- and e.reg=824 -- testing\n-- and e.city=49178 -- testing\nand e.city &lt; d.city\n)\ngroup by co, reg, leakcity, sistercity, destid, diff\nhaving count(*)&gt;1\norder by co, reg, leakcity, sistercity, destid;\n</code></pre>\n</blockquote>\n\n<p>The next levels of this can be addressed by Mattias.  But clearly to me, the first was to allow for sister cities without regard to (co, reg), e.g., San Francisco could be a sister city to Edmonton, Canada.</p>\n\n<p>The distribution of counts was about 90% value 2.  Those might be considered weak sister cities, but it turns out, given the decimal precision, they were very valuable.  Using only counts 3+ eliminates so many sister cities that, when compared to my best submission at point in time, I got 110 test row prediction changes - hardly worth the effort.  Again, Mattias can comment on this.</p>\n\n<p>Lastly, when I opened the floodgates to allow sister cities to exist anywhere, I was finding sister cities 9,000 miles apart!  I put in a hard limit of max distance between cities of 800 miles, which was an arbitrary choice.  I also found cases where distc2-diff created leak distances of &lt;0.01 miles, and I put in a hard 5 mile lower limit.  I believe Mattias used 600 miles and 0 miles for these thresholds.  My base file at this point was 11.9M lines; I think Mattias mentioned a file size of over 100M lines, so at this juncture, our approaches must have differed.  He got to 6th, so take it from him at this point!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123456,
      "author_name": "aldente",
      "author_url": "",
      "post_date": "06/12/2016 01:23:02",
      "content": "<p>@eipiplus1, thank you so much. Looking forward to learn Mattias' detailed approach :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123483,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/12/2016 07:08:45",
      "content": "<p>Hi, my best set of constants brought my base submission at 0.50566 (in PuLB) to 0.52171 (in PuLB) meaning an increase of 0.01605. In the PriLB I ended up with 0.51881 and I was very far from the #5 position, so I probably couldn't have made it there.</p>\n\n<p>Anyway, my last set of constants, as outlined in a previous post, did significantly worse. I stupidly didn't tag that version of my code - you should always tag your code once you make a better submission so you can easily revert to that once you do worse. Turns out that the exact constants were lost, but I was able to approximate the result.</p>\n\n<p>You won't be able to run it, because it's missing large parts. One really important part of this, I think, is that I was able to evaluate a new set of constants in 14 seconds. @epiplus1, I'm guessing your SQL version didn't allow for that? Using 600k rows instead of the full dataset I computed how well the wide leak would all on its own. And with a 14 second turnaround, that allowed me to continuously tweak parameters  trying to find a good set.</p>\n\n<p>Here are the parameters;</p>\n\n<pre><code>    public static int MaxRows = 600 * 1000;\n    //public static int MaxRows = int.MaxValue;\n    public float LimitForClose = 230;\n    public float LimitForFar = 30;\n    public int MinimumHits = 1;\n    public float Epsilon = 0.03f;\n    public float MatchEpsilon = 0.031f;\n</code></pre>\n\n<p>I'll try to figure out how to publish the code. It's not pretty...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123485,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/12/2016 07:15:58",
      "content": "<p>I couldn't actually post the code. Very strange, I got an error message. Maybe the code was too long. I've posted it on my blog instead. I doubt very much you'll get much from it, but there it is. It uses a ton of proprietary software that isn't public, but the algorithm is all in there.</p>\n\n<p><a href=\"https://lotsacode.wordpress.com/2016/06/12/code-from-expedia-hotel-recommendations/\">https://lotsacode.wordpress.com/2016/06/12/code-from-expedia-hotel-recommendations/</a></p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123522,
      "author_name": "aldente",
      "author_url": "",
      "post_date": "06/12/2016 14:26:08",
      "content": "<p>@Mattias, thank you very much for the code. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123527,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/12/2016 14:38:34",
      "content": "<p>The SQL itself ran relatively quickly, under one minute, but you're right, I didn't consider thresholds until back in my code.  So I didn't do tuning in SQL. I fixed my mind on a confirmation constant difference for each pair of cities, but now that I think about it more, even one match apparently is really valuable. I think going from N=2+ (giving me 11.9M before constraints) to N=1+ would get to 100M+ resultant rows to match on and would, based on what you said and your results, do better than other algorithms. Since it was the last day, and my code was taking about 2 hours to check through 5 million rows, after applying constraints, I would have had no hope to try N=1 anyhow.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123612,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/13/2016 05:26:06",
      "content": "<p>Thanks guys for sharing the journey with me, it was really fun to work on this at the very last couple of hours and bouncing ideas back and forth!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "123008": "I was thinking about the leak, and I realized that there must be cities which are close enough to eachother that their distances to hotels on the other side of the world will be very similar. Different, but the difference ought to be constant.\r\n\r\nTake city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.\r\n\r\nHow can this be exploited? If we've been able to find a constant distance between city A and city B, the leaked hotels for city A can be used for city B and vice versa.\r\n\r\nIt does seem to work, but it's not the huge gain I was  hoping for ;)",
    "123021": "I tried something similar, but not quite the same.  For a given (user_country, user_city, hotel_country, destid) I looked at the pairs (train_distance, HC) in relation, for all the training data - call it D.  Then for test, I explored those keys (user_country, user_city, hotel_country, destid) at test_distance and noted all the matching D entries where abs(test_distance - train_distance)<M (for M=1, 0.5, and 0.1=one city block).  The goal was to uncover instances in the test data where the set of candidate HC could be known exactly.\r\n\r\nThese sets of HC at 1 mile are sometimes 60+!  At distance 0.1 mile, there are even some 10+ counts.  I focused in on cases where it was HC set sizes of 1 or 2 only.  I compared the HC sets thus found with my best submission (out to N places), and these D HC often were at positions 6+ in my lists.\r\n\r\nThe reason to focus on set sizes 1 or 2 is that I wouldn't put much faith in the utility for a larger set of HC for, say, Chicago or New York City or Frankfurt or Beijing or wherever.  On the flip side, that means I may be limiting to HC sets for smaller destinations with fewer booked trips and, with less training data, higher variance if applied to predictions for test.\r\n\r\nI purposely left out hotel_market, but it may be worth trying.  \r\n\r\nThere was one catch to all this for me - I saved data from R to file, and didn't set options(digits=6) or such.  So, the distances I used in my coding was rounded to 3 digits.  The problem with this is that I was awarding a huge weight to exact matches, where abs(x-y)<0.000001 (float precision), and therefore though there were certainly many exact matches in my data, when the places were 3 or less, I definitely labeled many of the rows as INEXACT when they were exact.\r\n\r\nDoes this idea have some merit?  It's not what you described, but on the same train of thought.",
    "123025": "I think that that would pick up on the leak, but weaken it. The leak is that any combination of [orig_destination_distance, srch_destination_id_base, user_location_city] typically has  a unique hotel cluster - because it identifies one exact hotel.\r\n\r\nYou're looking at [user_country, user_city, hotel_country, destid]. But user_country is tied to user_city and hotel_country is mostly tied to destid. You're not grouping on the exact orig_destination_distance, but a soft grouping. That means that on top of the exact matches (the leak) you'll get a number of false matches.\r\n\r\nTo tell the truth, I also tried this, looking at hotels that don't have exact distance matches, but matching up to some range. The larger I made the range, the weaker the feature became. Leaving the range at 0 (which then becomes the leak) worked best in my attempts.\r\n\r\n/m",
    "123055": "I can now report that this technique improved my score by 0.00437, so; yay!",
    "123067": "Nice going! So close cities with constant differences has some value, good find. If I have time, I might try that.",
    "123135": "*Take city (country, region, city) 77,824,49178 and city 77,824,55729. For the search_destination_id 8745, my code tells me that there seems to be a constant difference of -5.061in orig_destination_distance.*\r\n\r\n@Mattias Fagerlund Why I didn't find the pattern in this example.It's seems in each pair, the orig_destination_distance are all different and the difference between them is also not -5.061.\r\n\r\n@eipiplus1   Does you approach work? (abs(x-y)<0.000001 ). So tiny difference will cause mismatch?",
    "123153": "FengLi\r\n\r\nHere is the data per his example:\r\n\r\n> 49178 v 55729 - 8745 - 64 - 10555.076 or 10550.015 --> 5.0610 \r\n\r\n> 49178 v 55729 - 8745 - 64 - 10555.076 or 10550.2091 --> 4.8669 \r\n\r\n> 49178 v 55729 - 8745 - 64 - 10555.076 or 10550.8242 --> 4.2518 \r\n\r\n> 49178 v 55729 - 8745 - 8 - 10556.3596 or 10548.0258 --> 8.3338 \r\n\r\n> 49178 v 55729 - 8745 - 97 - 10555.228 or 10550.167 --> 5.0610\r\n\r\nThis is data for country=77, region=824, showing c1 v c2.  All data is for destid=8745 as you can see.\r\nThe HCs are 64, 8, and 97.  Note that the distance changed for HC 64, for whatever reason.\r\n\r\nI can confirm the 5.0610 difference, seen exactly twice here.  \r\n\r\nI take this to mean that for a pair of cities (c1, c2), if there exists a distance difference DD which repeats 2+ times, this is a sign that they may be sister cities (as I'm calling them).  In this case, we have 2/1/1/1 when grouping by DD.\r\n\r\n@Mattias, I'm looking at this now in terms of 3 thresholds: (1) DD>=N, some critical value, like N=2 or N=3 or higher - it turns out, there are a whole lot of sister cities with at least one DD=2; or, the density of non-1 DD counts amongst all DD counts is at least X% (2) DD<=A miles, that is, the sister cities must be no more than A miles apart (3) leakcity to destid distance>=B miles, that is, the destid must be at least B miles from the leakcity.  I have tried N=2, A=100, and B=1000 so far.  Perhaps sister cities should be considered as closer than that.  With B=1000, I'm trying to avoid \"local\" trips, some of which are hotel bookings within a 5 mile radius.  I haven't tuned N/A/B yet, these are just tester values.  What do you think?",
    "123154": "FengLi\r\n\r\nAt this point in the game, I'm inclined to say my approach won't work.  I still hold hope for it, but with 1 day left, I won't have to time to further explore.  What Mattias is describing makes use of the leak heuristic, which is proven, while being exploratory about cities; what I'm doing is unproven and, while it may turn out to be the big secret, I'm sure it would require a lot more diligence to figure out exactly how to make use of the radii mapped destid/HC.",
    "123160": "Mattias Fagerlund - Are you saying that the processes is as follows:\r\n1. cluster rows by user_location_country, user_location_region, srch_destination_id.\r\n2. Given some example that i want to make a prediction on:\r\nTake it's orig_destination_distance, compare it to all other groups.\r\n3. You get a list of values. \r\n4. What is the next step? \r\n\r\nThank you.",
    "123178": "epiplus1 - that sounds about right. \r\n\r\nI find cities that are close and that are considered for \"sister city\" testing. I then find potential target search destinations that are far away (both sisters must have hotels there). From both sister cities. I then start to try to work out if they have any N hotels that have the same hotel cluster and the same distace - given some fixed constant between them. \r\n\r\nI'm not able to find many, here's one run on my validation set;\r\nTotal: 0.03619763(1976304), Specific: 0.4383184(163209), Count: 200585\r\n\r\nSo *at best* the result would improve by 0.036 - but a lot of those would be in the regular leak and so there would be no additional benefits to using this method.\r\n\r\nFor that run, near limit=30 miles, far limit=150 miles, N (min number of matched hotels) was 2, and the city comparison epsilon was 0.01.\r\n\r\nThese values can be tweaked to get more hits but of lower quality. They can be tweaked the other way to give fewer hits, but of greater quality. This particular run found 200585 hits and got a score of 0.438 on those. Now, 0.438 is better than any other feature I've found with the exception of the leak. so that's good. But the fact that it creates so few votes means it's not super helpful.",
    "123179": "kaukau, did that answer your question? There's 750 lines of code to do this, so it's not trivial, but I think it's an interesting idea!",
    "123184": "Thanks for the explanation and analysis.  I implemented two different versions of this, based on whether I strictly included the leakcity leakdist as part of my key, and here is what I found.  Starting with the 852K leak data rows, I expanded out the HC map and then reduced the final key list uniquely.  So that's the leak city search domain.  Then, I searched across the entire (country, region, destid) universe for any cities c1!=c2, computing the diff(dist) per HC, and then aggregating.  I ended up with 1.19M sister rows with leakdist, 964K sister rows without leakdist.  The advantage of the 964K row set is not having to compute sistercitydist+diff, I have the leakdist at hand.\r\n\r\nUsing the 964K set, I tried twice and got worse LB scores.  I used the 1.19M data set and improved my LB score.  It requires calculating sistercitydist+diff, as I mentioned.  I got 3,100 test row diff in that submission, and it proved HIGHLY effective.\r\n\r\nI found that I missed many sister cities due to float imprecison.  Now that I'm using numeric(12,4) I'm getting toward 7,500 matches.\r\n\r\nTo go one step further, I'm going to eventually test nearest neighbor cities constrained to a diff tolerance of 0.0005 or thereabouts.  Have you tried that?  It is not the leak heuristic, but it may prove out.  It's effectively going back to my radius idea, but at a WAY smaller value!\r\n\r\nBy the way, when it came to crunching the 1.19M leak domain data set against the 11.5M row search city data set, I wrote in SQL.  That flew!  It was the most natural way to join and process this data, in my view.\r\n\r\nI'm still working on it and measuring the efficacy of my approach.  Thanks again for sharing ideas and congrats for your place on the LB.",
    "123216": "Thanks, I was able to improve my score some more using this technique (by 0.00048, meaning 0.00485 from this technique alone), but that was only enough to move up one position. \r\n\r\nAs I was typing this, I made another submission that improved by another 0.00721, so that's 0,01206 from this technique. Currently at #11.\r\n\r\nI feel I'm running out of time here, very few submissions left to go. I'm thinking I could probably eek a bit more performance out of this, but I'm going to have to let it go.\r\n\r\nA further cool thing I would have liked to have done was to create a map given the city-hotel distances. This is surprisingly easy for 2d coordinates, you just keep moving the nodes (cities and hotels) around to satisfy their internal distance demands. Takes a number of iterations, but the math is trivial and fast. But on a sphere, it's not so easy, because you have to place the nodes using longitudes and latitudes and you have to handle wrap-around issues and so on. Very messy.\r\n\r\nBut once you have a map, you don't even need sister cities to both index the same hotels, with the map, you can search the map for hotels with a distance of X from city Y. But alas, there isn't enough time left for that - and it's a fairly low probability that I would have been able to get sufficient precision on the map coordinates.\r\n\r\n/m",
    "123223": "Mattias Fagerlund - How did you find cities that are close?\r\nI mean if i group by country, region and dest, and after that check all the all possible paris of\r\norig_di between one city and all other in this group i will get a lot of combinations?\r\nHow did you handle that?",
    "123225": "Actually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. \r\n\r\nYou can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping \"near\" lists. \r\n\r\nIf the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).\r\n\r\nBut as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. \r\n\r\nFor instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.\r\n\r\nSo you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.",
    "123228": "Very interesting Mattias, thanks for sharing. I would not be surprised if the solution of idle_speculation is related to what you described in your previous post - find all locations and extend the leak to all records with a distance in the test set. I have looked for ways to do that, but without external data, sufficient precision is quite a challenge like you state.",
    "123230": "I have been working on this for a few days as well - this is a tough challange to map all these distances - good work Mattias. Few more days and I think you would get into top3 :) I am highly certain idle_specaulation did this.",
    "123233": "Yikes. One more submission to go, and I just reached #6.\r\n\r\nhttps://www.kaggle.com/c/expedia-hotel-recommendations/leaderboard?submissionId=3150212\r\n\r\nBest placement ever for me ever on Kaggle, let's see if I can hang on. But this thread describes what I've been doing; I'm basically relaxing my requirements to increase the number of votes.",
    "123234": "wish you best of luck. with a careful submission top3 is yours :)",
    "123253": "Great work!  Congrats on reaching #6.  You were using \"strongly match\" as N>=2 for number of same distances to HC in a single destination.  Are you using N>=3 or more now?  My dist. is:\r\n\r\n> 10679303  2  \r\n\r\n> 335235  3   \r\n\r\n> 25042  4    \r\n\r\n> 3840  5    \r\n\r\n> 1102  6\r\n\r\n>  507  7\r\n\r\n>  297  8\r\n\r\n>  217  9\r\n\r\n>  117  10\r\n\r\n\r\nso there is quite a drop from 2 to 3.  I am still using 2 as the qualify cutoff.  The thing is, I know that when I go through the search space of (HC, dist) for a given destid, where there can be and often are multiple dist for a given HC, that the probability of a match on dist+diff to leakcity dist is very, very small.  So using N>=2 seems ok to me, or, in fact, may be missing out on lots of matches if I don't include it. \r\n\r\n\r\n\r\n[quote=Mattias Fagerlund;123225]\r\n\r\nActually, it turns out that being close isn't that important. To get more useful hits, I've relaxed that requirement. But finding close and far cities is quite easy. \r\n\r\nYou can determine the distance from a city to a particular search destination - f.i. you can use the max, min or average distance. So for each city, compute the distance to each search destination. For each city, create a list of near search destinations (where the distance is below X miles). For each city in a country (or region or continent), figure out which have overlapping \"near\" lists. \r\n\r\nIf the lists of near regions overlap, the cities can't be further apart than 2*X miles. If I know I'm within 2 miles of point A and I also know what you're within 2 miles of point A, I can determine that the furthest apart we can be is 4 miles (we're on totally opposite sides).\r\n\r\nBut as I said, this turned out to be a not necessary requirement. It helps cut down on pointless checks - cities that are far apart are unlikely to match up - so that saves time. But my algorithm doesn't take all that long, and sometimes cities that are far apart are good sister-city candidates anyway. \r\n\r\nFor instance, looking at cities A and B in relation to search destination X - if B is precisely on the line from A to X, you're likely to find that the distortions are small, even if A and B are a quite far apart. If, on the other hand, B is far away from the line between A and X, the distortions will be bigger.\r\n\r\nSo you need a test that only includes cities as sisters if they strongly match hotel distances and that works fast. Then you rely on that method to weed out the weak ones.\r\n\r\n[/quote]",
    "123261": "For full transparancy, I'll reveal my \"final\" constants. There's no guarantee that these are the best, but they worked well for me in testing. Also, I'm afraid I won't be able to make a subission using these constants, because so far the system has been working for hours without progress. \r\n\r\n        public float LimitForClose = 600;\r\n        public float LimitForFar = 0;\r\n        public int MaxDistanceDifferenceFactor = 200000;\r\n        public int ItemsToCheckPerCity = 40000;\r\n        public int MaxDistancesToCheck = 5000;\r\n        public int MinimumHits = 1;\r\n        public float Epsilon = 0.05f;//0.00001f;\r\n        public float MatchEpsilon = 0.05f;\r\n        public int MaxConstantsPerDestination = 750;\r\n        public float MaxSearchDestinationSize = 1000;\r\n\r\nBut if things holds true, and if I'm able to make a submission, it should be significantly better than my previous submissions.\r\n\r\nTesting for my best submission:\r\n  Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121 => 0.52171\r\n\r\nTesting for the submission I'm computing now:\r\n  Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530\r\n\r\nNow competition for #1, but it should result in better scores than my previous attempt. Problem though, with the previous entry it took 1:26 to create a submission (hour and a half). This submission looks to be about 3 times as complicated. So at 4 hours+, I may be out of time to even submit it. Oh well, that's the game, isn't it?\r\n\r\n/m",
    "123264": "anyway, share your algo after the competition, very interesting to see what have you done :)",
    "123265": "epiplus I'm no longer sure if I used 1 or 2 for my latest submission - probably 1 though. It decreases the quality but it increases the number of votes. On the whole, it seems better. But our mileages may vary though.\r\n\r\n/m",
    "123267": "LimitForClose = max dist between city1 and city2? min dist from leakcity to destid?\r\n\r\nI didn't put a limit on max distances to check, that's an interesting one to consider for speeding things up.  I do have an early check for if sistercitydist+diff<=0, in which case a diff was added to the base sister city and results in negative dist to destid, which of course is impossible.\r\n\r\nI exploded my city-city table to not be constrained by country/region, and now I'm running with that.  It was that, or create a key for (country, region, city) and do the Cartesian-ish product on that joining still by destid and HC.  So now my domain search space is 11.9M rows.\r\n\r\nI'm sitting on two right now with 34K row diff and 38K row diff, looking things over, code checking, comparing to previous submissions, checking how close sister cities are, distribution of distances to destid, etc.  You know how it goes!\r\n\r\nGood luck on your final submissions!",
    "123268": "Great work Mattias...  congrats on the strong finish, and thanks for your generosity in sharing your ideas!\r\n\r\nI wish I didn't have to work today so I could explore this idea before the deadline.  My exploitation of the leak is so basic and took 10 minutes and a dozen or so lines of code, and I figured 0.865 on affected rows was good enough and I'd focus work on features for a single model.  My weeks of work on incremental features yielded little, so it seems like a missed opportunity to have not spent more time looking at exploiting those distances instead.  Always so many \"should haves\" at the end of these contests...  for me anyway.  \r\n\r\nOne of these days I'll feel good at the end of a kaggle competition...  but not this one.  :)",
    "123271": "vtKMH , I tried boosted trees (using my own XGBoost simiar implementation), naive bayes categorization (using my own implementation) where I evolved weights in attempt to increse it's power, k-means clustering (using my own implementation), some other boosted tree attempts (using my own stupid implementations). And oh, I used PCA for some of my features. Using my own implementation, of course (You say why, you say why, you say why - don't ask me why! /Eurythmics ).\r\n\r\n *Nothing* I tried worked. I did start fairly strong by looking at the public scripts and using them as a basis for evolving my own version (using Genetic Algorithms, my own implementation of course). That worked well and temporarily got me to #10. I was later pushed to #36 and tried a million things, as stated above, but nothing helped.\r\n\r\nThen idle_speculation showed up, demonstrating that no small improvements would cut it. That made me look at the leak again, thinking how I could exploit it. That took *weeks*. I mean, I thought about what I could do for a very, very long time. Seems obvious in retrospect, but that's the way the cookie crumbles.\r\n\r\nAnyway, every competition I've been in, post facto I've though \"Oh, I was so close to doing that!\" when the winners explain what they did. Truth is, I was close to hundreds of different solutions, so luck plays a fairly strong hand in this.\r\n\r\nLuck favors the prepared! Good luck today and in the future!",
    "123272": "epiplus, turns out that the maximum distance to allow for \"close cities\" is important, but 600 miles gave better results (in testing) than 500 miles. Not very close at all! 1200 miles did worse, so there's some logic to it. LimitForFar, how far cities should be apart, is set at 0. Also gave better results than a more \"reasonable\" value, say 1000 miles, but what the heck. I'll trust the results over the theory - just this once. If my computer ever delivers up a submission for this latest set of constants.\r\n\r\nAs for the magnitude of this frigging thing, I found 102.825.935 different City/City/SearchDestination combinations to use. That's 103 MILLION combinations. Took a while. The file for storing these is 4.33GB large...",
    "123274": "I am using LimitForClose=800.  My first pass where I matched by (country, region) was 1.19M rows, with max abs distance 900 miles.  With full-on data, it's over 8,000 miles. Using 800 miles cut my 11.9M row count to 5.5M, about half.  For city dist, I'm using LimitForFar=5.  My reasoning is that I saw in the data there are times when the booking is <1 mile away, and after mapping sister city to leak city, throwing out negative distances, there are some with city to destid dist < 0.01.  For some degree of \"spread\" amongst the distances to destid, I chose 5.\r\n\r\nI have one small add-on to what we are doing, that is, I do look for close city matches within 0.005 after applying sistercitydist+diff (what you called constant at one point).  This is akin to what you were saying about rounding, and goes back to my idea of radii on destid.  My theory then was that the HC locations must be static, and therefore, putting a boundary around them would be something to anchor on strongly.  So now I'm effectively doing that in this context.\r\n\r\nI have about 2 more hours of run time on my full-on data set (without the add-on) and then it's go time!",
    "123277": "I was actually able to submit my last entry in the competition - and it did worse. Which was what I was expecting. I'll try to give you an explanation here.\r\n\r\nOnce I realized that lowering my thresholds for \"wide leak\" entries was helping my results, I knew that lowering my thresholds too much would eventually damage my results. Finding the sweet spot was the game, and time was running out.\r\n\r\nHere's a breakdown of the last four submissions I did.\r\n\r\n- Total: 0.01394861(600000), Specific: 0.6348937(13182), Count: 6159 =>\r\n   **0.50566**  \r\n\r\nI used a set of validation searches (600k) to test my strategy. I used a base submission that I'd created earlier using a counting strategy similar to the public scripts. This submission had a score of 0.50566 and it of course included the leak. My new wide leak solution was injected between the leak and before my previous solution and it improved my result to 0.51051. I knew I was on the right track - but this was during the last day of the competition! What to do? Well, I had to further relax the constraints and see if I could improve the results.\r\n\r\nThe results indicate that the score of the votes placed by this strategy was 0.63 - meaning it was extremely good but it gave to few votes. (The leak was at 0.88, so that was much better). Relaxing the constraints would increase the number of votes but decrease the quality of those votes. At some point, any new votes added by this strategy would override better votes by my base strategy - and thus reduce the results.\r\n\r\n- Total: 0.02265029(600000), Specific: 0.5906204(23010), Count: 14449\r\n   => **0.51772**\r\n\r\nFirst attempt for a second entry, I was waaay to conservative. I decreased the specific score (the score of actual votes cast, ignoring entries where no votes where cast) to 0.59. It significantly improved the result, but seeing as I had only a day to work with, I should have been more aggressive.\r\n\r\n- Total: 0.03969542(600000), Specific: 0.4976858(47856), Count: 121121\r\n   => **0.52171**\r\n\r\nNext attempt, I decided to go for a specific score of 0.5 and try to maximize the total score I could get from that. Remember, the total score doesn't include the leak and it doesn't include my base strategy. It just gives a null vote for any search that the wide leak doesn't give a result for. That significantly increased my results again. That was a bold move in contrast to my second move.\r\n\r\n- Total: 0.05186441(600000), Specific: 0.4211255(73894), Count: 411530\r\n   => **0.51960**\r\n\r\nHere I overshot at my last attempt. I should probably have gone for a specific score of 0.45 but I ended up with a specific score of .42. The total score looked nicer, but this strategy replaced too many better results from my base strategy that it did worse. It would have been enough (right now) for a #8 place, but no better. What would a specific score of .45 done? Who knows, I may have maximized the potential of my strategy on my previous entry, but 0.42 was clearly too aggressive.\r\n\r\nI had so much fun in this competition, and I learned so much. Thanks to all of you who've taken the time to discuss these things with me!",
    "123278": "eipiplus1 The best of luck to you, I'm all out of submissions, but I hope this improves your results!",
    "123280": "Awesome work, and thanks for the writeup.  I whiteboarded a hier. solution on day #1 which had: is_mobile, channel, adult_cnt, child_cnt, rm_cnt, hotel_market.  6 straight features, then did some binning, and got surprisingly good results.  (Needless to say, I dropped is_mobile quickly! And channel, too, though site_name came into play later.)  I moved straight into NB, which seemed a natural fit with counts, but no tuning of laplace, eps or threshold produced good results.  Then one round of xgboost, and nothing.  So back to layering!  Then came the date columns DT, CI , CO, then recency weighting, then tuning.  Along the way the leak was announced.  I spent a few days on that trying to get to 0.90 on it, but I landed at 0.882.  Then I tried user recom. and found that worked, with lots of tuning on click and book counts.  (This may be the downfall for many of us, we shall see!)  There are 3 layers high up which do geometrical averaging of the top 5 highest variance features I found, of 5, then 4, then 3 features. Next came many days of grid search on 2013 and 2014 data, and improvements.  Now, like you, I have this sister-city layer wedged right under leak and above 27 other hier. layers.   I've picked up 40 ticks so far, and that was with a (co, reg, c1-c2, destid) approach.  Also, I didn't until today fully utilize the (c1-c2) pairs irrespective of destid; and lastly, the full city blowout without (co, reg).  Congrats on a STRONG finish and I hope to see you near the top if I am on target with my final 3 submissions.",
    "123297": "Thanks a lot for sharing all of this and your genetic algorithms too.  Unfortunately for me, I preferred to work on something completely different rather than reusing your ideas.\r\n\r\nI hope you'll be on time to submit your current work in progress.  I won't have the same chance, my current WIP is due to finish sometime this WE :(",
    "123448": "Mattias Fagerlund, @eipiplus1, congratulations! \r\n\r\nAny chance to share your code/script of widening the leak? Not full code, only the widening part. Although a full code would be the best ;))",
    "123454": "It'll be best to see code (or snippets) from Mattias, who moved up to 6th place.  I picked up 40+ ticks working on this, but I think he claimed 473 ticks (either real or potential) at one point.  In any case, here you go!\r\n\r\n1. Take whatever is your leak code or file, and produce the search set of known leak keys\r\n(country, region, city, dist, HC).  My leak was 852K rows, but some had multiple HC.  After expanding it to 1.0M+ rows, unique it.  I ended up with 813K defined leak rows keyed as indicated earlier.\r\n\r\nSample data:\r\n\r\n    205,354,40193,11813,1845.0739,16\r\n    205,354,40193,11813,1851.8713,83\r\n    205,354,40193,11813,1859.7697,0\r\n    205,354,40193,11827,54.911,42\r\n    205,354,40193,11827,56.8564,59\r\n    205,354,40193,11827,56.8564,95\r\n    205,354,40193,11827,59.9279,42\r\n    205,354,40193,11827,60.2396,4\r\n\r\n2. Create a search domain of cities for which you hope to find sister cities.  I did this in R.  I ended up not using N, but I included it.  It's very important to set the precision before saving the file.\r\n\r\n    lg <- subset(train, !is.na(orig_destination_distance)) %>% group_by(user_location_country, user_location_region, user_location_city, \r\n                         srch_destination_id, orig_destination_distance, hotel_cluster) %>% summarise(N=n())\r\n\r\n    options(digits=8)\r\n\r\n    write.csv(lg, \"ud.csv\", row.names = F)\r\n\r\n\r\n\r\n> \r\n\r\n3. Load these two data sets into database tables - I started to code a search algorithm and quickly realized that SQL would tackle it much more quickly!  It's resolving 800K and 11M row data sets.\r\n\r\n> \r\n\r\n4. Run SQL, such as below.  This was my first attempt, later refined.  Here, I am looking for sister cities in the same country and region - therefore, the diff values were no more than 980 miles.  I did not yet consider finding cities in line with leak cities.  My result set was limited to any tuples (destid, HC, diff) with diff for cities c1&lt;c2 is distc2-distc1, having count 2+.  That is, a common distance difference diff for 2+ (destid, HC) was found for that (co, reg, c1) and (co, reg, c2) pairing.  That it is signal that if c1 is the leak city, and c2 is a non-leak city, then taking (co, reg, c1, destid, distc2-diff) will convert city c2 into a c1 lookalike.  This is what I did to pick up 40+ ticks.  The sister city lookup file for me at this point was 1.19M lines.\r\n\r\n\r\n>     select co, reg, leakcity, sistercity, destid, diff, count(*) AS CNT\r\n>     from (\r\n>     select e.co, e.reg, e.city AS leakcity, d.city AS sistercity, d.destid, e.dist AS leakdist, d.dist, e.dist - d.dist AS diff\r\n>     from analytics_sandbox.ekey1 e, analytics_sandbox.eud d\r\n>     where e.co = d.co\r\n>     and e.reg = d.reg\r\n>     and e.destid = d.destid\r\n>     and e.hc = d.hc\r\n>     -- and e.co=1-- testing\r\n>     -- and e.reg=824 -- testing\r\n>     -- and e.city=49178 -- testing\r\n>     and e.city < d.city\r\n>     )\r\n>     group by co, reg, leakcity, sistercity, destid, diff\r\n>     having count(*)>1\r\n>     order by co, reg, leakcity, sistercity, destid;\r\n\r\n\r\nThe next levels of this can be addressed by Mattias.  But clearly to me, the first was to allow for sister cities without regard to (co, reg), e.g., San Francisco could be a sister city to Edmonton, Canada.\r\n\r\nThe distribution of counts was about 90% value 2.  Those might be considered weak sister cities, but it turns out, given the decimal precision, they were very valuable.  Using only counts 3+ eliminates so many sister cities that, when compared to my best submission at point in time, I got 110 test row prediction changes - hardly worth the effort.  Again, Mattias can comment on this.\r\n\r\nLastly, when I opened the floodgates to allow sister cities to exist anywhere, I was finding sister cities 9,000 miles apart!  I put in a hard limit of max distance between cities of 800 miles, which was an arbitrary choice.  I also found cases where distc2-diff created leak distances of <0.01 miles, and I put in a hard 5 mile lower limit.  I believe Mattias used 600 miles and 0 miles for these thresholds.  My base file at this point was 11.9M lines; I think Mattias mentioned a file size of over 100M lines, so at this juncture, our approaches must have differed.  He got to 6th, so take it from him at this point!",
    "123456": "eipiplus1, thank you so much. Looking forward to learn Mattias' detailed approach :)",
    "123483": "Hi, my best set of constants brought my base submission at 0.50566 (in PuLB) to 0.52171 (in PuLB) meaning an increase of 0.01605. In the PriLB I ended up with 0.51881 and I was very far from the #5 position, so I probably couldn't have made it there.\r\n\r\nAnyway, my last set of constants, as outlined in a previous post, did significantly worse. I stupidly didn't tag that version of my code - you should always tag your code once you make a better submission so you can easily revert to that once you do worse. Turns out that the exact constants were lost, but I was able to approximate the result.\r\n\r\nYou won't be able to run it, because it's missing large parts. One really important part of this, I think, is that I was able to evaluate a new set of constants in 14 seconds. @epiplus1, I'm guessing your SQL version didn't allow for that? Using 600k rows instead of the full dataset I computed how well the wide leak would all on its own. And with a 14 second turnaround, that allowed me to continuously tweak parameters  trying to find a good set.\r\n\r\nHere are the parameters;\r\n\r\n        public static int MaxRows = 600 * 1000;\r\n        //public static int MaxRows = int.MaxValue;\r\n        public float LimitForClose = 230;\r\n        public float LimitForFar = 30;\r\n        public int MinimumHits = 1;\r\n        public float Epsilon = 0.03f;\r\n        public float MatchEpsilon = 0.031f;\r\n\r\nI'll try to figure out how to publish the code. It's not pretty...",
    "123485": "I couldn't actually post the code. Very strange, I got an error message. Maybe the code was too long. I've posted it on my blog instead. I doubt very much you'll get much from it, but there it is. It uses a ton of proprietary software that isn't public, but the algorithm is all in there.\r\n\r\nhttps://lotsacode.wordpress.com/2016/06/12/code-from-expedia-hotel-recommendations/\r\n\r\n/m",
    "123522": "Mattias, thank you very much for the code.",
    "123527": "The SQL itself ran relatively quickly, under one minute, but you're right, I didn't consider thresholds until back in my code.  So I didn't do tuning in SQL. I fixed my mind on a confirmation constant difference for each pair of cities, but now that I think about it more, even one match apparently is really valuable. I think going from N=2+ (giving me 11.9M before constraints) to N=1+ would get to 100M+ resultant rows to match on and would, based on what you said and your results, do better than other algorithms. Since it was the last day, and my code was taking about 2 hours to check through 5 million rows, after applying constraints, I would have had no hope to try N=1 anyhow.",
    "123612": "Thanks guys for sharing the journey with me, it was really fun to work on this at the very last couple of hours and bouncing ideas back and forth!"
  },
  "source": "meta"
}