{"cells":[
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "library(data.table)"
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "expedia_train <- fread('../input/train.csv', header=TRUE, \n                       select= c(\"user_location_city\", \"hotel_market\", \"orig_destination_distance\",\n                                 \"is_booking\", \"date_time\", \"hotel_cluster\"),\n                       verbose = F)\nhead(expedia_train)"
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "expedia_test <- fread('../input/test.csv', header=TRUE,\n                     select= c(\"user_location_city\", \"hotel_market\", \"orig_destination_distance\",\n                               \"date_time\"),\n                     verbose = F)\nhead(expedia_test)"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "Let's see how many orig_destination_distances in the test set can be found in the training set:"
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "test_distances <- expedia_test$orig_destination_distance\nexpedia_train[, sum(test_distances %in% orig_destination_distance) / nrow(expedia_test)]"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "75 percent of the distances in the test set exist in the training set.\nThat's a lot!\n\nLet's see some of the cases..."
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "test_index  <- expedia_train[!is.na(orig_destination_distance), \n                             which(test_distances %in% orig_destination_distance)]\ncat(\"TEST DATA:\", fill = T)\nexpedia_test[test_index[1:10]]\n\ncat(\"\\n\\nTRAIN DATA:\", fill = T)\nexpedia_train[orig_destination_distance %in% expedia_test[test_index[1:10]]$orig_destination_distance]"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "And this is the data leak!\nAs we can see all the clicks and bookings the user did for a given location, \nand the hotel clusters he visited, we already have the data we had to predict."
 }
],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"}}, "nbformat": 4, "nbformat_minor": 0}