{
  "id": 20926,
  "title": "Will I be inserting a leak with summary features",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20926",
  "author_name": "",
  "post_date": "2016-05-13T06:59:19.480Z",
  "votes": null,
  "comment_count": 1,
  "views": 537,
  "content": "<p>Hello</p>\n\n<p>If I took statistics about user interests from training data for example, number of times user i has clicked cluster j in month k(same for booking), total number of clicks, bookings on cluster j in month k and insert them as new features along with existing training data. </p>\n\n<p>Will this create a new data leak in this model? </p>",
  "messages": [
    {
      "id": "119856",
      "postDate": "05/13/2016 06:59:19",
      "content": "<p>Hello</p>\n\n<p>If I took statistics about user interests from training data for example, number of times user i has clicked cluster j in month k(same for booking), total number of clicks, bookings on cluster j in month k and insert them as new features along with existing training data. </p>\n\n<p>Will this create a new data leak in this model? </p>",
      "rawMarkdown": "Hello\r\n\r\nIf I took statistics about user interests from training data for example, number of times user i has clicked cluster j in month k(same for booking), total number of clicks, bookings on cluster j in month k and insert them as new features along with existing training data. \r\n\r\nWill this create a new data leak in this model?",
      "votes": null
    },
    {
      "id": "119960",
      "postDate": "05/14/2016 04:10:10",
      "content": "<p>Yes, averages and sums are models, even if they are very basic ones.  Just calculate these values ignoring the current observation. These data are 37mil records though, so I would just use k-folds sampling to make features for the hold out sets. </p>",
      "rawMarkdown": "Yes, averages and sums are models, even if they are very basic ones.  Just calculate these values ignoring the current observation. These data are 37mil records though, so I would just use k-folds sampling to make features for the hold out sets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119960,
      "author_name": "nschneider",
      "author_url": "",
      "post_date": "05/14/2016 04:10:10",
      "content": "<p>Yes, averages and sums are models, even if they are very basic ones.  Just calculate these values ignoring the current observation. These data are 37mil records though, so I would just use k-folds sampling to make features for the hold out sets. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119856": "Hello\r\n\r\nIf I took statistics about user interests from training data for example, number of times user i has clicked cluster j in month k(same for booking), total number of clicks, bookings on cluster j in month k and insert them as new features along with existing training data. \r\n\r\nWill this create a new data leak in this model?",
    "119960": "Yes, averages and sums are models, even if they are very basic ones.  Just calculate these values ignoring the current observation. These data are 37mil records though, so I would just use k-folds sampling to make features for the hold out sets."
  },
  "source": "meta"
}