{
  "id": 20571,
  "title": "Terminology",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20571",
  "author_name": "",
  "post_date": "2016-04-30T14:27:39.380Z",
  "votes": null,
  "comment_count": 5,
  "views": 587,
  "content": "<p>I am still a bit confused about what the different terms mean. </p>\n\n<p>I understand that a user searches for a destination. There seem to be about 65k of those. Then, she will be offered a choice of hotel clusters, maximum 100. \nWe don't know anything about the hotel clusters.\nHowever, for the destinations we have some &quot;features&quot;. \nFor each destination ID, the hotel clusters are different. So the list of 100 hotel clusters is only relevant in the context of the destination.</p>\n\n<p>The description: &quot;srch_destination_id   ID of the destination where the hotel search was performed&quot; should really be:\nID of the destination FOR WHICH the hotel search was performed.</p>\n\n<p>orig_destination_distance   Physical distance between a hotel and a customer at the time of search. A null means the distance could not be calculated</p>\n\n<p>From what I said earlier, this should be:\ndistance between a customer and a DESTINATION.</p>\n\n<p>Is my understanding correct?</p>",
  "messages": [
    {
      "id": "117718",
      "postDate": "04/30/2016 14:27:39",
      "content": "<p>I am still a bit confused about what the different terms mean. </p>\n\n<p>I understand that a user searches for a destination. There seem to be about 65k of those. Then, she will be offered a choice of hotel clusters, maximum 100. \nWe don't know anything about the hotel clusters.\nHowever, for the destinations we have some &quot;features&quot;. \nFor each destination ID, the hotel clusters are different. So the list of 100 hotel clusters is only relevant in the context of the destination.</p>\n\n<p>The description: &quot;srch_destination_id   ID of the destination where the hotel search was performed&quot; should really be:\nID of the destination FOR WHICH the hotel search was performed.</p>\n\n<p>orig_destination_distance   Physical distance between a hotel and a customer at the time of search. A null means the distance could not be calculated</p>\n\n<p>From what I said earlier, this should be:\ndistance between a customer and a DESTINATION.</p>\n\n<p>Is my understanding correct?</p>",
      "rawMarkdown": "I am still a bit confused about what the different terms mean. \r\n\r\nI understand that a user searches for a destination. There seem to be about 65k of those. Then, she will be offered a choice of hotel clusters, maximum 100. \r\nWe don't know anything about the hotel clusters.\r\nHowever, for the destinations we have some \"features\". \r\nFor each destination ID, the hotel clusters are different. So the list of 100 hotel clusters is only relevant in the context of the destination.\r\n\r\nThe description: \"srch_destination_id \tID of the destination where the hotel search was performed\" should really be:\r\nID of the destination FOR WHICH the hotel search was performed.\r\n\r\norig_destination_distance \tPhysical distance between a hotel and a customer at the time of search. A null means the distance could not be calculated\r\n\r\nFrom what I said earlier, this should be:\r\ndistance between a customer and a DESTINATION.\r\n\r\nIs my understanding correct?",
      "votes": null
    },
    {
      "id": "117720",
      "postDate": "04/30/2016 14:34:16",
      "content": "<p>Hi,</p>\n\n<p>reading about the data leak your last statement seems not correct. It is the distance between customer and selected hotel. Although the same hotel can belong to different clusters depending on who knows what...</p>\n\n<p>And I don't think the customers are offered hotel clusters. They are offered hotels. Which belong to a cluster - or another - depending on who knows what.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nreading about the data leak your last statement seems not correct. It is the distance between customer and selected hotel. Although the same hotel can belong to different clusters depending on who knows what...\r\n\r\nAnd I don't think the customers are offered hotel clusters. They are offered hotels. Which belong to a cluster - or another - depending on who knows what.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117721",
      "postDate": "04/30/2016 14:38:51",
      "content": "<p>@Gerhard - thanks for your quick reaction. \nOf course, in real life customers are offered hotels, not clusters. But in the context of the data we have, we can only identify the cluster, right? We don't have any data about individual hotels, do we?</p>",
      "rawMarkdown": "Gerhard - thanks for your quick reaction. \r\nOf course, in real life customers are offered hotels, not clusters. But in the context of the data we have, we can only identify the cluster, right? We don't have any data about individual hotels, do we?",
      "votes": null
    },
    {
      "id": "117733",
      "postDate": "04/30/2016 16:04:54",
      "content": "<p>As I understood it we can most surely identify the cluster of a hotel if the distance is given. Although we don't know the specifics of the hotel then. This holds if there is only one hotel in a building. And if the cluster didn't change for some reason or other.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "As I understood it we can most surely identify the cluster of a hotel if the distance is given. Although we don't know the specifics of the hotel then. This holds if there is only one hotel in a building. And if the cluster didn't change for some reason or other.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117738",
      "postDate": "04/30/2016 16:22:59",
      "content": "<p>I am a bit slow :-) So although a place is usually defined by two coordinates (Lat, Lon) or perhaps (distance, angle), distance is unique enough to identify a particular hotel? \nI have 10,622,932 unique combinations of the tuple mentioned in the data leak article. Adding hotel_cluster increases the number to 14,346,455. So for one combination I get 1.4 hotels? Hotel clusters? </p>",
      "rawMarkdown": "I am a bit slow :-) So although a place is usually defined by two coordinates (Lat, Lon) or perhaps (distance, angle), distance is unique enough to identify a particular hotel? \r\nI have 10,622,932 unique combinations of the tuple mentioned in the data leak article. Adding hotel_cluster increases the number to 14,346,455. So for one combination I get 1.4 hotels? Hotel clusters?",
      "votes": null
    },
    {
      "id": "117741",
      "postDate": "04/30/2016 16:27:26",
      "content": "<p>I think that might be correct. Not that I noted the numbers. But you can check. Roughly 1/3 of the test set should match one of your many rows.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "I think that might be correct. Not that I noted the numbers. But you can check. Roughly 1/3 of the test set should match one of your many rows.\r\n\r\nGerhard",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117720,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/30/2016 14:34:16",
      "content": "<p>Hi,</p>\n\n<p>reading about the data leak your last statement seems not correct. It is the distance between customer and selected hotel. Although the same hotel can belong to different clusters depending on who knows what...</p>\n\n<p>And I don't think the customers are offered hotel clusters. They are offered hotels. Which belong to a cluster - or another - depending on who knows what.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117721,
      "author_name": "mafux777",
      "author_url": "",
      "post_date": "04/30/2016 14:38:51",
      "content": "<p>@Gerhard - thanks for your quick reaction. \nOf course, in real life customers are offered hotels, not clusters. But in the context of the data we have, we can only identify the cluster, right? We don't have any data about individual hotels, do we?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117733,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/30/2016 16:04:54",
      "content": "<p>As I understood it we can most surely identify the cluster of a hotel if the distance is given. Although we don't know the specifics of the hotel then. This holds if there is only one hotel in a building. And if the cluster didn't change for some reason or other.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117738,
      "author_name": "mafux777",
      "author_url": "",
      "post_date": "04/30/2016 16:22:59",
      "content": "<p>I am a bit slow :-) So although a place is usually defined by two coordinates (Lat, Lon) or perhaps (distance, angle), distance is unique enough to identify a particular hotel? \nI have 10,622,932 unique combinations of the tuple mentioned in the data leak article. Adding hotel_cluster increases the number to 14,346,455. So for one combination I get 1.4 hotels? Hotel clusters? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117741,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/30/2016 16:27:26",
      "content": "<p>I think that might be correct. Not that I noted the numbers. But you can check. Roughly 1/3 of the test set should match one of your many rows.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117718": "I am still a bit confused about what the different terms mean. \r\n\r\nI understand that a user searches for a destination. There seem to be about 65k of those. Then, she will be offered a choice of hotel clusters, maximum 100. \r\nWe don't know anything about the hotel clusters.\r\nHowever, for the destinations we have some \"features\". \r\nFor each destination ID, the hotel clusters are different. So the list of 100 hotel clusters is only relevant in the context of the destination.\r\n\r\nThe description: \"srch_destination_id \tID of the destination where the hotel search was performed\" should really be:\r\nID of the destination FOR WHICH the hotel search was performed.\r\n\r\norig_destination_distance \tPhysical distance between a hotel and a customer at the time of search. A null means the distance could not be calculated\r\n\r\nFrom what I said earlier, this should be:\r\ndistance between a customer and a DESTINATION.\r\n\r\nIs my understanding correct?",
    "117720": "Hi,\r\n\r\nreading about the data leak your last statement seems not correct. It is the distance between customer and selected hotel. Although the same hotel can belong to different clusters depending on who knows what...\r\n\r\nAnd I don't think the customers are offered hotel clusters. They are offered hotels. Which belong to a cluster - or another - depending on who knows what.\r\n\r\nGerhard",
    "117721": "Gerhard - thanks for your quick reaction. \r\nOf course, in real life customers are offered hotels, not clusters. But in the context of the data we have, we can only identify the cluster, right? We don't have any data about individual hotels, do we?",
    "117733": "As I understood it we can most surely identify the cluster of a hotel if the distance is given. Although we don't know the specifics of the hotel then. This holds if there is only one hotel in a building. And if the cluster didn't change for some reason or other.\r\n\r\nGerhard",
    "117738": "I am a bit slow :-) So although a place is usually defined by two coordinates (Lat, Lon) or perhaps (distance, angle), distance is unique enough to identify a particular hotel? \r\nI have 10,622,932 unique combinations of the tuple mentioned in the data leak article. Adding hotel_cluster increases the number to 14,346,455. So for one combination I get 1.4 hotels? Hotel clusters?",
    "117741": "I think that might be correct. Not that I noted the numbers. But you can check. Roughly 1/3 of the test set should match one of your many rows.\r\n\r\nGerhard"
  },
  "source": "meta"
}