{
  "id": 20703,
  "title": "Help to understand data.table",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20703",
  "author_name": "",
  "post_date": "2016-05-04T09:51:50.730Z",
  "votes": null,
  "comment_count": 4,
  "views": 564,
  "content": "<p>Hello All,</p>\n\n<p>This may seem stupid but I need help with the following:</p>\n\n<p>In this <a href=\"https://www.kaggle.com/signochastic/expedia-hotel-recommendations/apr-23/code\">code script</a>, I am not able to understand how the line below works:</p>\n\n<p><code>dest_id_hotel_cluster_count &lt;- expedia_train[,sum_and_count(is_booking),by=list(orig_destination_distance, hotel_cluster)]</code></p>\n\n<p>Does it first group the data by <code>orig_destination_distance</code> and  <code>hotel_cluster</code> and then apply the function on <code>is_booking</code> ?</p>\n\n<p>But then how the grouping works on <code>orig_destination_distance</code>? As it is a double and it could be any value.</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "118567",
      "postDate": "05/04/2016 09:51:50",
      "content": "<p>Hello All,</p>\n\n<p>This may seem stupid but I need help with the following:</p>\n\n<p>In this <a href=\"https://www.kaggle.com/signochastic/expedia-hotel-recommendations/apr-23/code\">code script</a>, I am not able to understand how the line below works:</p>\n\n<p><code>dest_id_hotel_cluster_count &lt;- expedia_train[,sum_and_count(is_booking),by=list(orig_destination_distance, hotel_cluster)]</code></p>\n\n<p>Does it first group the data by <code>orig_destination_distance</code> and  <code>hotel_cluster</code> and then apply the function on <code>is_booking</code> ?</p>\n\n<p>But then how the grouping works on <code>orig_destination_distance</code>? As it is a double and it could be any value.</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hello All,\r\n\r\nThis may seem stupid but I need help with the following:\r\n\r\nIn this [code script][1], I am not able to understand how the line below works:\r\n\r\n\r\n`dest_id_hotel_cluster_count <- expedia_train[,sum_and_count(is_booking),by=list(orig_destination_distance, hotel_cluster)]`\r\n\r\nDoes it first group the data by `orig_destination_distance` and  `hotel_cluster` and then apply the function on `is_booking` ?\r\n\r\nBut then how the grouping works on `orig_destination_distance`? As it is a double and it could be any value.\r\n\r\nThanks.\r\n\r\n\r\n  [1]: https://www.kaggle.com/signochastic/expedia-hotel-recommendations/apr-23/code",
      "votes": null
    },
    {
      "id": "118718",
      "postDate": "05/04/2016 23:14:42",
      "content": "<p>@Chintan Shah: though  <code>orig_destination_distance</code> is a double it has finite number of unique values which is &lt;&lt; the number of rows in the dataset. So, you can group them here. Though grouping on a continuous variable does not make much sense in the first place, the above was done in order to exploit the data leak between the train and test set.</p>",
      "rawMarkdown": "Chintan Shah: though  `orig_destination_distance` is a double it has finite number of unique values which is << the number of rows in the dataset. So, you can group them here. Though grouping on a continuous variable does not make much sense in the first place, the above was done in order to exploit the data leak between the train and test set.",
      "votes": null
    },
    {
      "id": "118781",
      "postDate": "05/05/2016 09:21:03",
      "content": "<p>@Bishwarup thank you for your reply! I understand it now. Can you please tell me what is Data Leak? I had a look on web but could not find a better explanation. </p>",
      "rawMarkdown": "Bishwarup thank you for your reply! I understand it now. Can you please tell me what is Data Leak? I had a look on web but could not find a better explanation.",
      "votes": null
    },
    {
      "id": "118792",
      "postDate": "05/05/2016 11:04:11",
      "content": "<p>Please visit the link <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20345/data-leak\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20345/data-leak</a> for more information on the data leak.</p>",
      "rawMarkdown": "Please visit the link https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20345/data-leak for more information on the data leak.",
      "votes": null
    },
    {
      "id": "118796",
      "postDate": "05/05/2016 11:21:34",
      "content": "<p>Thanks for the link!</p>",
      "rawMarkdown": "Thanks for the link!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118718,
      "author_name": "bishwarup",
      "author_url": "",
      "post_date": "05/04/2016 23:14:42",
      "content": "<p>@Chintan Shah: though  <code>orig_destination_distance</code> is a double it has finite number of unique values which is &lt;&lt; the number of rows in the dataset. So, you can group them here. Though grouping on a continuous variable does not make much sense in the first place, the above was done in order to exploit the data leak between the train and test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118781,
      "author_name": "thecpshah",
      "author_url": "",
      "post_date": "05/05/2016 09:21:03",
      "content": "<p>@Bishwarup thank you for your reply! I understand it now. Can you please tell me what is Data Leak? I had a look on web but could not find a better explanation. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118792,
      "author_name": "bishwarup",
      "author_url": "",
      "post_date": "05/05/2016 11:04:11",
      "content": "<p>Please visit the link <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20345/data-leak\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20345/data-leak</a> for more information on the data leak.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118796,
      "author_name": "thecpshah",
      "author_url": "",
      "post_date": "05/05/2016 11:21:34",
      "content": "<p>Thanks for the link!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118567": "Hello All,\r\n\r\nThis may seem stupid but I need help with the following:\r\n\r\nIn this [code script][1], I am not able to understand how the line below works:\r\n\r\n\r\n`dest_id_hotel_cluster_count <- expedia_train[,sum_and_count(is_booking),by=list(orig_destination_distance, hotel_cluster)]`\r\n\r\nDoes it first group the data by `orig_destination_distance` and  `hotel_cluster` and then apply the function on `is_booking` ?\r\n\r\nBut then how the grouping works on `orig_destination_distance`? As it is a double and it could be any value.\r\n\r\nThanks.\r\n\r\n\r\n  [1]: https://www.kaggle.com/signochastic/expedia-hotel-recommendations/apr-23/code",
    "118718": "Chintan Shah: though  `orig_destination_distance` is a double it has finite number of unique values which is << the number of rows in the dataset. So, you can group them here. Though grouping on a continuous variable does not make much sense in the first place, the above was done in order to exploit the data leak between the train and test set.",
    "118781": "Bishwarup thank you for your reply! I understand it now. Can you please tell me what is Data Leak? I had a look on web but could not find a better explanation.",
    "118792": "Please visit the link https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20345/data-leak for more information on the data leak.",
    "118796": "Thanks for the link!"
  },
  "source": "meta"
}