{
  "id": 20646,
  "title": "New to Kaggle - question on feature engineering",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20646",
  "author_name": "dchan",
  "post_date": "2016-05-02T23:55:48.003000",
  "votes": 1,
  "comment_count": 0,
  "views": 501,
  "content": "<p>Hello all,</p>\n\n<p>I've been mainly participating as an observer Kaggle for the past year, but decided to commit to the Expedia challenge so that I can learn by actively participating in the challenging and engaging with the community.</p>\n\n<p>One question that I struggle with conceptually is how people incorporate the benchmarks scripts into their machine learning workflow.</p>\n\n<p>Take for example the benchmark where people work out the most popular local hotel clusters per search destination (i.e. &quot;the most popular local hotel&quot; benchmark, <a href=\"https://www.kaggle.com/ccccat/expedia-hotel-recommendations/r-version-of-most-popular-local-hotel\">https://www.kaggle.com/ccccat/expedia-hotel-recommendations/r-version-of-most-popular-local-hotel</a>), and index match it back to the test dataset to generate a set of benchmark predictions. I struggle to understand conceptually how information from this benchmark can be used in building features for machine learning, as we don't have any information on hotel_cluster in our test set.</p>\n\n<p>I understand that people may not be willing to give away their competitive advantage on public forums, but any pointers to guide my thinking would be greatly appreciated.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 118162,
      "postDate": "2016-05-02T23:55:48.003Z",
      "content": "<p>Hello all,</p>\n\n<p>I've been mainly participating as an observer Kaggle for the past year, but decided to commit to the Expedia challenge so that I can learn by actively participating in the challenging and engaging with the community.</p>\n\n<p>One question that I struggle with conceptually is how people incorporate the benchmarks scripts into their machine learning workflow.</p>\n\n<p>Take for example the benchmark where people work out the most popular local hotel clusters per search destination (i.e. &quot;the most popular local hotel&quot; benchmark, <a href=\"https://www.kaggle.com/ccccat/expedia-hotel-recommendations/r-version-of-most-popular-local-hotel\">https://www.kaggle.com/ccccat/expedia-hotel-recommendations/r-version-of-most-popular-local-hotel</a>), and index match it back to the test dataset to generate a set of benchmark predictions. I struggle to understand conceptually how information from this benchmark can be used in building features for machine learning, as we don't have any information on hotel_cluster in our test set.</p>\n\n<p>I understand that people may not be willing to give away their competitive advantage on public forums, but any pointers to guide my thinking would be greatly appreciated.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hello all,\r\n\r\nI've been mainly participating as an observer Kaggle for the past year, but decided to commit to the Expedia challenge so that I can learn by actively participating in the challenging and engaging with the community.\r\n\r\nOne question that I struggle with conceptually is how people incorporate the benchmarks scripts into their machine learning workflow.\r\n\r\nTake for example the benchmark where people work out the most popular local hotel clusters per search destination (i.e. \"the most popular local hotel\" benchmark, https://www.kaggle.com/ccccat/expedia-hotel-recommendations/r-version-of-most-popular-local-hotel), and index match it back to the test dataset to generate a set of benchmark predictions. I struggle to understand conceptually how information from this benchmark can be used in building features for machine learning, as we don't have any information on hotel_cluster in our test set.\r\n\r\nI understand that people may not be willing to give away their competitive advantage on public forums, but any pointers to guide my thinking would be greatly appreciated.\r\n\r\nThanks!",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "118162": "Hello all,\r\n\r\nI've been mainly participating as an observer Kaggle for the past year, but decided to commit to the Expedia challenge so that I can learn by actively participating in the challenging and engaging with the community.\r\n\r\nOne question that I struggle with conceptually is how people incorporate the benchmarks scripts into their machine learning workflow.\r\n\r\nTake for example the benchmark where people work out the most popular local hotel clusters per search destination (i.e. \"the most popular local hotel\" benchmark, https://www.kaggle.com/ccccat/expedia-hotel-recommendations/r-version-of-most-popular-local-hotel), and index match it back to the test dataset to generate a set of benchmark predictions. I struggle to understand conceptually how information from this benchmark can be used in building features for machine learning, as we don't have any information on hotel_cluster in our test set.\r\n\r\nI understand that people may not be willing to give away their competitive advantage on public forums, but any pointers to guide my thinking would be greatly appreciated.\r\n\r\nThanks!"
  }
}