{"cells":[{"cell_type":"markdown","metadata":{"_cell_guid":"3bb141b4-56be-3e8a-11af-f359b57a4570"},"source":"We are told there is a time-based train/test split. Let's have a look at it, and a small explore of the time-structure of the data in general."},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"0024ede8-85e8-a023-ddd7-ef2b2c266e41"},"outputs":[],"source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\nevents = pd.read_csv(\"../input/events.csv\", dtype=np.int32, index_col=0, usecols=[0,3])\nevents.head()"},{"cell_type":"markdown","metadata":{"_cell_guid":"e2a1f8b7-9ed7-13d6-f4c5-81e68a370d60"},"source":"Timestamps are in milliseconds since 1970-01-01 - 1465876799998, so time zero approximately corresponds to 04:00 UTC, 14th June 2016. I'm not going to do any correction for timezones, but it's worth pointing out that this corresponds to 00:00 Eastern Daylight Time."},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"8f176aa9-ad88-d95e-9c20-0e59e15bcce8"},"outputs":[],"source":"train = pd.merge(pd.read_csv(\"../input/clicks_train.csv\", dtype=np.int32, index_col=0).sample(frac=0.1),\n                 events, left_index=True, right_index=True)\ntest = pd.merge(pd.read_csv(\"../input/clicks_test.csv\", dtype=np.int32, index_col=0).sample(frac=0.1),\n                events, left_index=True, right_index=True)"},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"fb21f0c3-a388-9cb7-49e1-5860709b4fa9"},"outputs":[],"source":"test[\"hour\"] = (test.timestamp // (3600 * 1000)) % 24\ntest[\"day\"] = test.timestamp // (3600 * 24 * 1000)\n\ntrain[\"hour\"] = (train.timestamp // (3600 * 1000)) % 24\ntrain[\"day\"] = train.timestamp // (3600 * 24 * 1000)\n\nplt.figure(figsize=(12,4))\ntrain.hour.hist(bins=np.linspace(-0.5, 23.5, 25), label=\"train\", alpha=0.7, normed=True)\ntest.hour.hist(bins=np.linspace(-0.5, 23.5, 25), label=\"test\", alpha=0.7, normed=True)\nplt.xlim(-0.5, 23.5)\nplt.legend(loc=\"best\")\nplt.xlabel(\"Hour of Day\")\nplt.ylabel(\"Fraction of Events\")"},{"cell_type":"markdown","metadata":{"_cell_guid":"34104a50-da0f-561a-302b-a9f0e1fb0c2d"},"source":"The time-distribution of clicks appears consistent with the majority of the dataset being in UTC-4 to UTC-6. There appears to be a slight shift from train to test."},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"89c33942-d9d0-1d77-2dc7-7e62946f19d9"},"outputs":[],"source":"plt.figure(figsize=(12,4))\ntrain.day.hist(bins=np.linspace(-.5, 14.5, 16), label=\"train\", alpha=0.7, normed=True)\ntest.day.hist(bins=np.linspace(-.5, 14.5, 16), label=\"test\", alpha=0.7, normed=True)\nplt.xlim(-0.5, 14.5)\nplt.legend(loc=\"best\")\nplt.xlabel(\"Days since June 14\")\nplt.ylabel(\"Fraction of Events\")"},{"cell_type":"markdown","metadata":{"_cell_guid":"15f4a54b-1971-ec73-7c56-8910e35520c0"},"source":"Now we can see the very interesting time-based train/test split. About half of the test data is sampled from the same time as the train set, with another half sampled from the two days immediately following.\n\nNext for a more detailed look at the distribution of the training data. Note again that no timezone correction has been applied, but 80% of the data are from the US."},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"7202bace-83d8-403e-a133-a00e60019844"},"outputs":[],"source":"plt.figure(figsize=(12,6))\nhour_day_counts = train.groupby([\"hour\", \"day\"]).count().ad_id.values.reshape(24,-1)\n# plot 2d hist in days and hours, with each day normalised to 1 \nplt.imshow((hour_day_counts / hour_day_counts.sum(axis=0)).T,\n           interpolation=\"none\", cmap=\"rainbow\")\nplt.xlabel(\"Hour of Day\")\nplt.ylabel(\"Days since June 14\")"},{"cell_type":"markdown","metadata":{"_cell_guid":"1cd7786d-e26c-7326-e1a9-04f245091672"},"source":"The weekend (days 4, 5 and 11 and 12) looks noticeably different from the weekdays -- greater proportion of late-night activity, no real peak in the early afternoon."},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"b4884d69-11db-71ea-9456-8f8a9fb95902"},"outputs":[],"source":"# for completeness, the test set too:\nplt.figure(figsize=(12,6))\nhour_day_counts = test.groupby([\"hour\", \"day\"]).count().ad_id.values.reshape(24,-1)\n# plot 2d hist in days and hours, with each day normalised to 1 \nplt.imshow((hour_day_counts / hour_day_counts.sum(axis=0)).T,\n           interpolation=\"none\", cmap=\"rainbow\")\nplt.xlabel(\"Hour of Day\")\nplt.ylabel(\"Days since June 14\")"},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"11597cc1-05c3-c6ac-9af2-66e4142979d0"},"outputs":[],"source":""}],"metadata":{"_change_revision":0,"_is_fork":false,"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.5.2"}},"nbformat":4,"nbformat_minor":0}