{
  "id": 53754,
  "title": "Are there missing clicks in Train set?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53754",
  "author_name": "",
  "post_date": "2018-04-04T17:45:48.432302400Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Train data ends at 2017-11-09 at 16:00:00.  Old test started at 2017-11-09 14:23:39.  That's over one and a half hours of overlap.  </p>\n\n<p>My question is:  was it a mistake in old test data and some of train made it into test or does it mean that there is data missing from train set, thus making counts for those hours understated? (which would be no good as they partially correspond to hour ranges in the actual test...)</p>",
  "messages": [
    {
      "id": "309138",
      "postDate": "04/04/2018 17:45:48",
      "content": "<p>Train data ends at 2017-11-09 at 16:00:00.  Old test started at 2017-11-09 14:23:39.  That's over one and a half hours of overlap.  </p>\n\n<p>My question is:  was it a mistake in old test data and some of train made it into test or does it mean that there is data missing from train set, thus making counts for those hours understated? (which would be no good as they partially correspond to hour ranges in the actual test...)</p>",
      "rawMarkdown": "Train data ends at 2017-11-09 at 16:00:00.  Old test started at 2017-11-09 14:23:39.  That's over one and a half hours of overlap.  \n\nMy question is:  was it a mistake in old test data and some of train made it into test or does it mean that there is data missing from train set, thus making counts for those hours understated? (which would be no good as they partially correspond to hour ranges in the actual test...)",
      "votes": null
    },
    {
      "id": "309934",
      "postDate": "04/06/2018 08:07:38",
      "content": "<p>I think it was stated somewhere that the train data is only a subset of the full data from those days, subsetted by times and ips. So hopefully for the present ips we have the full history, otherwise it would be indeed quite problematic for all kind of count features.</p>",
      "rawMarkdown": "I think it was stated somewhere that the train data is only a subset of the full data from those days, subsetted by times and ips. So hopefully for the present ips we have the full history, otherwise it would be indeed quite problematic for all kind of count features.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 309934,
      "author_name": "malten",
      "author_url": "",
      "post_date": "04/06/2018 08:07:38",
      "content": "<p>I think it was stated somewhere that the train data is only a subset of the full data from those days, subsetted by times and ips. So hopefully for the present ips we have the full history, otherwise it would be indeed quite problematic for all kind of count features.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "309138": "Train data ends at 2017-11-09 at 16:00:00.  Old test started at 2017-11-09 14:23:39.  That's over one and a half hours of overlap.  \n\nMy question is:  was it a mistake in old test data and some of train made it into test or does it mean that there is data missing from train set, thus making counts for those hours understated? (which would be no good as they partially correspond to hour ranges in the actual test...)",
    "309934": "I think it was stated somewhere that the train data is only a subset of the full data from those days, subsetted by times and ips. So hopefully for the present ips we have the full history, otherwise it would be indeed quite problematic for all kind of count features."
  },
  "source": "meta"
}