{
  "id": 21506,
  "title": "Co-centered Search Destinations",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21506",
  "author_name": "",
  "post_date": "2016-06-08T05:44:13.103Z",
  "votes": 3,
  "comment_count": 5,
  "views": 918,
  "content": "<p>[A previous version of this post talked about co-centered search destinations, but it's in fact overlapping search destinations.]</p>\n\n<p>I've written some code to determine which search destinations overlap with other search destinations. The way I've worked this out is by looking at distances from one city to a search destination A. If there are trips, from the same city, to other search destination B that have the exact same distances <em>and</em> the exact same cluster, it would seem they're going to the same hotel. Which means that particular hotel is included in both search destination A &amp; B which is to be expected. The distance from a city to a hotel will be the same no matter what search destination they're included from.</p>\n\n<p>Take 8250 and 12206 for instance, I was able to find 37823 distance matching trips between these. Out of the matched trips, 37790 had the same clusters and 33 had different clusters. That is 33 with the same distance but different clusters - that could mean duplicate distances or that a hotel changed clusters (which they're supposed to do sometimes).</p>\n\n<p>Note that I am <em>not</em> saying they're the same search destinations, but that they're two destinations that overlap. Maybe it's Manhattan and New York City or Paris and France. Different searches, but partially the same hotels.</p>\n\n<p>Here are three examples of search destinations that have overlapping hotels.</p>\n\n<p>12391,21254,24620</p>\n\n<p>439,15941,38798</p>\n\n<p>61057,61060</p>\n\n<h1>Testing the feature</h1>\n\n<p>I haven't been able to really use this new knowledge yet, though the leak becomes slightly stronger with this information than without it. The leak can be used by grouping on the city a user came from and the distance to the hotel. If you find a hit, then use the hotel_cluster for that hotel.</p>\n\n<p>But sometimes there are duplicates that aren't actual hits, they may go to a different hotel at the exact same distance. You can fix this by also grouping on srch_destination_id - in which you have better precision on the votes cast, but ultimately you cast fewer votes and the result is worse.</p>\n\n<p>Using srch_destination_id_base (which is simply the lowest srch_destination_id in a pool of overlapping srch_destination_ids) allows you to 1) filter out the false positives and 2) keep on the true positives. Unfortunately, though, it didn't help that much.</p>\n\n<h2>Grouped on: user_location_city, orig_destination_distance</h2>\n\n<p>Getting leaked info 67.998s  </p>\n\n<p>Total: 0.2282762(1207263), Specific: 0.8851196(435533) </p>\n\n<blockquote>\n  <p>Many votes cast and score is high.</p>\n</blockquote>\n\n<h2>Grouped on: user_location_city, orig_destination_distance, srch_destination_id</h2>\n\n<p>Getting leaked info 61.883s</p>\n\n<p>Total:0.1758085(1144100), Specific: 0.8818949(342576)</p>\n\n<blockquote>\n  <p>Far fewer votes cast, but precision is better on the actual votes cast.</p>\n</blockquote>\n\n<h2>Grouped on: user_location_city, orig_destination_distance, srch_destination_id_<strong>base</strong></h2>\n\n<p>Getting leaked info 63.564s</p>\n\n<p>Total:0.2362686(1140054), Specific: 0.8861945(424267)</p>\n\n<blockquote>\n  <p>My method of finding the base search destination isn't flawless, but the number of votes is higher since it's able to match up more items. The number of cast votes is higher which leads to better overall result. Compared to the version that ignores the srch_destination_id, it performs slightly better (+0.001). </p>\n</blockquote>\n\n<h1>Conclusion</h1>\n\n<p>Hardly worth the effort for the boost above, but it might prove a useful feature elsewhere...</p>",
  "messages": [
    {
      "id": "122898",
      "postDate": "06/08/2016 05:44:13",
      "content": "<p>[A previous version of this post talked about co-centered search destinations, but it's in fact overlapping search destinations.]</p>\n\n<p>I've written some code to determine which search destinations overlap with other search destinations. The way I've worked this out is by looking at distances from one city to a search destination A. If there are trips, from the same city, to other search destination B that have the exact same distances <em>and</em> the exact same cluster, it would seem they're going to the same hotel. Which means that particular hotel is included in both search destination A &amp; B which is to be expected. The distance from a city to a hotel will be the same no matter what search destination they're included from.</p>\n\n<p>Take 8250 and 12206 for instance, I was able to find 37823 distance matching trips between these. Out of the matched trips, 37790 had the same clusters and 33 had different clusters. That is 33 with the same distance but different clusters - that could mean duplicate distances or that a hotel changed clusters (which they're supposed to do sometimes).</p>\n\n<p>Note that I am <em>not</em> saying they're the same search destinations, but that they're two destinations that overlap. Maybe it's Manhattan and New York City or Paris and France. Different searches, but partially the same hotels.</p>\n\n<p>Here are three examples of search destinations that have overlapping hotels.</p>\n\n<p>12391,21254,24620</p>\n\n<p>439,15941,38798</p>\n\n<p>61057,61060</p>\n\n<h1>Testing the feature</h1>\n\n<p>I haven't been able to really use this new knowledge yet, though the leak becomes slightly stronger with this information than without it. The leak can be used by grouping on the city a user came from and the distance to the hotel. If you find a hit, then use the hotel_cluster for that hotel.</p>\n\n<p>But sometimes there are duplicates that aren't actual hits, they may go to a different hotel at the exact same distance. You can fix this by also grouping on srch_destination_id - in which you have better precision on the votes cast, but ultimately you cast fewer votes and the result is worse.</p>\n\n<p>Using srch_destination_id_base (which is simply the lowest srch_destination_id in a pool of overlapping srch_destination_ids) allows you to 1) filter out the false positives and 2) keep on the true positives. Unfortunately, though, it didn't help that much.</p>\n\n<h2>Grouped on: user_location_city, orig_destination_distance</h2>\n\n<p>Getting leaked info 67.998s  </p>\n\n<p>Total: 0.2282762(1207263), Specific: 0.8851196(435533) </p>\n\n<blockquote>\n  <p>Many votes cast and score is high.</p>\n</blockquote>\n\n<h2>Grouped on: user_location_city, orig_destination_distance, srch_destination_id</h2>\n\n<p>Getting leaked info 61.883s</p>\n\n<p>Total:0.1758085(1144100), Specific: 0.8818949(342576)</p>\n\n<blockquote>\n  <p>Far fewer votes cast, but precision is better on the actual votes cast.</p>\n</blockquote>\n\n<h2>Grouped on: user_location_city, orig_destination_distance, srch_destination_id_<strong>base</strong></h2>\n\n<p>Getting leaked info 63.564s</p>\n\n<p>Total:0.2362686(1140054), Specific: 0.8861945(424267)</p>\n\n<blockquote>\n  <p>My method of finding the base search destination isn't flawless, but the number of votes is higher since it's able to match up more items. The number of cast votes is higher which leads to better overall result. Compared to the version that ignores the srch_destination_id, it performs slightly better (+0.001). </p>\n</blockquote>\n\n<h1>Conclusion</h1>\n\n<p>Hardly worth the effort for the boost above, but it might prove a useful feature elsewhere...</p>",
      "rawMarkdown": "[A previous version of this post talked about co-centered search destinations, but it's in fact overlapping search destinations.]\r\n\r\nI've written some code to determine which search destinations overlap with other search destinations. The way I've worked this out is by looking at distances from one city to a search destination A. If there are trips, from the same city, to other search destination B that have the exact same distances *and* the exact same cluster, it would seem they're going to the same hotel. Which means that particular hotel is included in both search destination A & B which is to be expected. The distance from a city to a hotel will be the same no matter what search destination they're included from.\r\n\r\nTake 8250 and 12206 for instance, I was able to find 37823 distance matching trips between these. Out of the matched trips, 37790 had the same clusters and 33 had different clusters. That is 33 with the same distance but different clusters - that could mean duplicate distances or that a hotel changed clusters (which they're supposed to do sometimes).\r\n\r\nNote that I am *not* saying they're the same search destinations, but that they're two destinations that overlap. Maybe it's Manhattan and New York City or Paris and France. Different searches, but partially the same hotels.\r\n\r\nHere are three examples of search destinations that have overlapping hotels.\r\n\r\n12391,21254,24620\r\n\r\n439,15941,38798\r\n\r\n61057,61060\r\n\r\nTesting the feature\r\n===================\r\n\r\nI haven't been able to really use this new knowledge yet, though the leak becomes slightly stronger with this information than without it. The leak can be used by grouping on the city a user came from and the distance to the hotel. If you find a hit, then use the hotel_cluster for that hotel.\r\n\r\nBut sometimes there are duplicates that aren't actual hits, they may go to a different hotel at the exact same distance. You can fix this by also grouping on srch_destination_id - in which you have better precision on the votes cast, but ultimately you cast fewer votes and the result is worse.\r\n\r\nUsing srch_destination_id_base (which is simply the lowest srch_destination_id in a pool of overlapping srch_destination_ids) allows you to 1) filter out the false positives and 2) keep on the true positives. Unfortunately, though, it didn't help that much.\r\n\r\n\r\nGrouped on: user_location_city, orig_destination_distance\r\n---------------------------------------------------------\r\n\r\nGetting leaked info 67.998s  \r\n\r\nTotal: 0.2282762(1207263), Specific: 0.8851196(435533) \r\n> Many votes cast and score is high.\r\n\r\n\r\nGrouped on: user_location_city, orig_destination_distance, srch_destination_id\r\n---------------------------------------------------------\r\n\r\nGetting leaked info 61.883s\r\n\r\nTotal:0.1758085(1144100), Specific: 0.8818949(342576)\r\n> Far fewer votes cast, but precision is better on the actual votes cast.\r\n\r\nGrouped on: user_location_city, orig_destination_distance, srch_destination_id_**base**\r\n---------------------------------------------------------\r\n\r\nGetting leaked info 63.564s\r\n\r\nTotal:0.2362686(1140054), Specific: 0.8861945(424267)\r\n\r\n> My method of finding the base search destination isn't flawless, but the number of votes is higher since it's able to match up more items. The number of cast votes is higher which leads to better overall result. Compared to the version that ignores the srch_destination_id, it performs slightly better (+0.001). \r\n\r\nConclusion\r\n==========\r\n\r\nHardly worth the effort for the boost above, but it might prove a useful feature elsewhere...",
      "votes": null
    },
    {
      "id": "122907",
      "postDate": "06/08/2016 07:47:55",
      "content": "<p>@Mattias, do you use any leakage information or purely use models in your method?</p>",
      "rawMarkdown": "Mattias, do you use any leakage information or purely use models in your method?",
      "votes": null
    },
    {
      "id": "122911",
      "postDate": "06/08/2016 08:51:32",
      "content": "<p>This was a pure leakage study - but my other models mix the two. I've been unable to find any ML solution that helps though, grouping and counting seems to beat everything else I try...</p>",
      "rawMarkdown": "This was a pure leakage study - but my other models mix the two. I've been unable to find any ML solution that helps though, grouping and counting seems to beat everything else I try...",
      "votes": null
    },
    {
      "id": "122974",
      "postDate": "06/08/2016 23:18:53",
      "content": "<p>I got a very minor LB improvement using the exact same approach.  It may due to chance?</p>",
      "rawMarkdown": "I got a very minor LB improvement using the exact same approach.  It may due to chance?",
      "votes": null
    },
    {
      "id": "122997",
      "postDate": "06/09/2016 03:49:24",
      "content": "<p>Hi, perhaps I am not interpreting your numbers correctly - you said:</p>\n\n<blockquote>\n  <p>(city, dist) - Total: 0.2282762(1207263), Specific: 0.8851196(435533) </p>\n  \n  <p>(city, destid, dist) - Total:0.1758085(1144100), Specific: 0.8818949(342576)</p>\n</blockquote>\n\n<p>I have higher cardinality when I add destid, 11930247 vs. 10105719 to be exact.\nYou show 0.885 vs 0.881 for precision - isn't that worse?  Maybe I am reading\nspecific/precision incorrectly.</p>\n\n<p>Then with dest_base:</p>\n\n<blockquote>\n  <p>(city, destid_base, dist) - Total:0.2362686(1140054), Specific: 0.8861945(424267)</p>\n</blockquote>\n\n<p>which is 0.886 and is higher than your first number by the 0.001.  Did you actually\nsee a +0.001 on the LB after applying your destid_base method?</p>\n\n<p>Great work!</p>",
      "rawMarkdown": "Hi, perhaps I am not interpreting your numbers correctly - you said:\r\n\r\n> (city, dist) - Total: 0.2282762(1207263), Specific: 0.8851196(435533) \r\n\r\n> (city, destid, dist) - Total:0.1758085(1144100), Specific: 0.8818949(342576)\r\n\r\nI have higher cardinality when I add destid, 11930247 vs. 10105719 to be exact.\r\nYou show 0.885 vs 0.881 for precision - isn't that worse?  Maybe I am reading\r\nspecific/precision incorrectly.\r\n\r\nThen with dest_base:\r\n\r\n> (city, destid_base, dist) - Total:0.2362686(1140054), Specific: 0.8861945(424267)\r\n\r\nwhich is 0.886 and is higher than your first number by the 0.001.  Did you actually\r\nsee a +0.001 on the LB after applying your destid_base method?\r\n\r\nGreat work!",
      "votes": null
    },
    {
      "id": "123006",
      "postDate": "06/09/2016 05:36:49",
      "content": "<p>Hi, you're right, the I got the specific values confused. </p>\n\n<p>I haven't submitted any entry with these changes; I found that the improvement they provide wouldn't be enough to change my position in any meaningful way.</p>",
      "rawMarkdown": "Hi, you're right, the I got the specific values confused. \r\n\r\nI haven't submitted any entry with these changes; I found that the improvement they provide wouldn't be enough to change my position in any meaningful way.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122907,
      "author_name": "masterliu",
      "author_url": "",
      "post_date": "06/08/2016 07:47:55",
      "content": "<p>@Mattias, do you use any leakage information or purely use models in your method?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122911,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/08/2016 08:51:32",
      "content": "<p>This was a pure leakage study - but my other models mix the two. I've been unable to find any ML solution that helps though, grouping and counting seems to beat everything else I try...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122974,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/08/2016 23:18:53",
      "content": "<p>I got a very minor LB improvement using the exact same approach.  It may due to chance?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122997,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/09/2016 03:49:24",
      "content": "<p>Hi, perhaps I am not interpreting your numbers correctly - you said:</p>\n\n<blockquote>\n  <p>(city, dist) - Total: 0.2282762(1207263), Specific: 0.8851196(435533) </p>\n  \n  <p>(city, destid, dist) - Total:0.1758085(1144100), Specific: 0.8818949(342576)</p>\n</blockquote>\n\n<p>I have higher cardinality when I add destid, 11930247 vs. 10105719 to be exact.\nYou show 0.885 vs 0.881 for precision - isn't that worse?  Maybe I am reading\nspecific/precision incorrectly.</p>\n\n<p>Then with dest_base:</p>\n\n<blockquote>\n  <p>(city, destid_base, dist) - Total:0.2362686(1140054), Specific: 0.8861945(424267)</p>\n</blockquote>\n\n<p>which is 0.886 and is higher than your first number by the 0.001.  Did you actually\nsee a +0.001 on the LB after applying your destid_base method?</p>\n\n<p>Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123006,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/09/2016 05:36:49",
      "content": "<p>Hi, you're right, the I got the specific values confused. </p>\n\n<p>I haven't submitted any entry with these changes; I found that the improvement they provide wouldn't be enough to change my position in any meaningful way.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122898": "[A previous version of this post talked about co-centered search destinations, but it's in fact overlapping search destinations.]\r\n\r\nI've written some code to determine which search destinations overlap with other search destinations. The way I've worked this out is by looking at distances from one city to a search destination A. If there are trips, from the same city, to other search destination B that have the exact same distances *and* the exact same cluster, it would seem they're going to the same hotel. Which means that particular hotel is included in both search destination A & B which is to be expected. The distance from a city to a hotel will be the same no matter what search destination they're included from.\r\n\r\nTake 8250 and 12206 for instance, I was able to find 37823 distance matching trips between these. Out of the matched trips, 37790 had the same clusters and 33 had different clusters. That is 33 with the same distance but different clusters - that could mean duplicate distances or that a hotel changed clusters (which they're supposed to do sometimes).\r\n\r\nNote that I am *not* saying they're the same search destinations, but that they're two destinations that overlap. Maybe it's Manhattan and New York City or Paris and France. Different searches, but partially the same hotels.\r\n\r\nHere are three examples of search destinations that have overlapping hotels.\r\n\r\n12391,21254,24620\r\n\r\n439,15941,38798\r\n\r\n61057,61060\r\n\r\nTesting the feature\r\n===================\r\n\r\nI haven't been able to really use this new knowledge yet, though the leak becomes slightly stronger with this information than without it. The leak can be used by grouping on the city a user came from and the distance to the hotel. If you find a hit, then use the hotel_cluster for that hotel.\r\n\r\nBut sometimes there are duplicates that aren't actual hits, they may go to a different hotel at the exact same distance. You can fix this by also grouping on srch_destination_id - in which you have better precision on the votes cast, but ultimately you cast fewer votes and the result is worse.\r\n\r\nUsing srch_destination_id_base (which is simply the lowest srch_destination_id in a pool of overlapping srch_destination_ids) allows you to 1) filter out the false positives and 2) keep on the true positives. Unfortunately, though, it didn't help that much.\r\n\r\n\r\nGrouped on: user_location_city, orig_destination_distance\r\n---------------------------------------------------------\r\n\r\nGetting leaked info 67.998s  \r\n\r\nTotal: 0.2282762(1207263), Specific: 0.8851196(435533) \r\n> Many votes cast and score is high.\r\n\r\n\r\nGrouped on: user_location_city, orig_destination_distance, srch_destination_id\r\n---------------------------------------------------------\r\n\r\nGetting leaked info 61.883s\r\n\r\nTotal:0.1758085(1144100), Specific: 0.8818949(342576)\r\n> Far fewer votes cast, but precision is better on the actual votes cast.\r\n\r\nGrouped on: user_location_city, orig_destination_distance, srch_destination_id_**base**\r\n---------------------------------------------------------\r\n\r\nGetting leaked info 63.564s\r\n\r\nTotal:0.2362686(1140054), Specific: 0.8861945(424267)\r\n\r\n> My method of finding the base search destination isn't flawless, but the number of votes is higher since it's able to match up more items. The number of cast votes is higher which leads to better overall result. Compared to the version that ignores the srch_destination_id, it performs slightly better (+0.001). \r\n\r\nConclusion\r\n==========\r\n\r\nHardly worth the effort for the boost above, but it might prove a useful feature elsewhere...",
    "122907": "Mattias, do you use any leakage information or purely use models in your method?",
    "122911": "This was a pure leakage study - but my other models mix the two. I've been unable to find any ML solution that helps though, grouping and counting seems to beat everything else I try...",
    "122974": "I got a very minor LB improvement using the exact same approach.  It may due to chance?",
    "122997": "Hi, perhaps I am not interpreting your numbers correctly - you said:\r\n\r\n> (city, dist) - Total: 0.2282762(1207263), Specific: 0.8851196(435533) \r\n\r\n> (city, destid, dist) - Total:0.1758085(1144100), Specific: 0.8818949(342576)\r\n\r\nI have higher cardinality when I add destid, 11930247 vs. 10105719 to be exact.\r\nYou show 0.885 vs 0.881 for precision - isn't that worse?  Maybe I am reading\r\nspecific/precision incorrectly.\r\n\r\nThen with dest_base:\r\n\r\n> (city, destid_base, dist) - Total:0.2362686(1140054), Specific: 0.8861945(424267)\r\n\r\nwhich is 0.886 and is higher than your first number by the 0.001.  Did you actually\r\nsee a +0.001 on the LB after applying your destid_base method?\r\n\r\nGreat work!",
    "123006": "Hi, you're right, the I got the specific values confused. \r\n\r\nI haven't submitted any entry with these changes; I found that the improvement they provide wouldn't be enough to change my position in any meaningful way."
  },
  "source": "meta"
}