{"cells":[{"cell_type":"markdown","metadata":{"_cell_guid":"2f903da0-a9ed-9ab4-6e58-eea1e5e09d6f"},"source":"As I checked if access information for landing page of ad clicks is in page_views.csv for clicks_test.csv in [this script](https://www.kaggle.com/its7171/outbrain-click-prediction/leakage-solution/discussion), I tryed to check for clicks_train.csv to estimate this efects."},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"ec85085e-fc71-488e-f8fc-a0a242900712"},"outputs":[],"source":"df_traindf_train"},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"50849970-dc6f-f209-11ab-3babfb2333a1"},"outputs":[],"source":"df_train"},{"cell_type":"code","execution_count":null,"metadata":{"_cell_guid":"a59fcf76-6900-80e6-6511-0de2654dfc3d"},"outputs":[],"source":"\ntime_dict = df_train[['timestamp']].to_dict()['timestamp']\n# set page_views.csv for full data\n# f = open(\"../input/page_views.csv\", \"r\")\nf = open(\"../input/page_views_sample.csv\",\"r\")\nline = f.readline().strip()\nhead_arr = line.split(\",\")\nfld_index = dict(zip(head_arr,range(0,len(head_arr))))\ntotal = 0\nwhile 1:\n    line = f.readline().strip()\n    if nrows is not None and total == nrows:\n        break\n    total += 1\n    if line == '':\n        break\n    arr = line.split(\",\")\n    usr_doc = arr[fld_index['uuid']] + '_' + arr[fld_index['document_id']]\n    if usr_doc in time_dict:\n        #don't use timestamp yet.\n        #time_diff = time_dict[usr_doc] - int(arr[fld_index['timestamp']])\n        #if abs(time_diff) < 600:\n            # set -1 if found that this user sow this document\n            time_dict[usr_doc] = -1\n\ndf_train=df_train.reset_index()\ndf_train['fixed_timestamp'] = df_train['usr_doc'].apply(lambda x: time_dict[x])\nfound_in_page_views = set(df_train[df_train['fixed_timestamp'] < 0].index)\nclicked = set(df_train[df_train['clicked'] == 1].index)\nall_ids = set(df_train.index)\n\nTP = len(clicked & found_in_page_views)\nFP = len(found_in_page_views - clicked)\nFN = len(clicked - found_in_page_views)\nrecall = TP/float(TP+FN)\nprecision = TP/float(TP+FP)\nprint('TP:{}'.format(TP))\nprint('FP:{}'.format(FP))\nprint('FN:{}'.format(FN))\nprint('recall:{0:.1f}%'.format(recall*100))\nprint('precision:{0:.1f}%'.format(precision*100))"},{"cell_type":"markdown","metadata":{"_cell_guid":"19e25aa7-c3c9-1f1d-e745-1e6bedd5a70d"},"source":"For full data, this output would be:\n\n<pre>\nTP:724749\nFP:31813\nFN:16149844\nrecall:4.3%\nprecision:95.8%\n</pre>"},{"cell_type":"markdown","metadata":{"_cell_guid":"b8601b9e-cdc5-e75c-4fcd-810e36873375"},"source":"Only 4.3% of Click data is found in page_views.csv.\nThis percentage is far smaller than I thought.\nI thought that landing page log should be in page_views.csv, if some ad link is clicked.\nWhere is remaining 95.7% access?\nDoes page_views.csv includes only 4.3% sampling data?\nOr Outbrain does not have all page view data for landing page of ads?\n\nupdate: I guess that page_views.csv includes access logs for all the page which has ads in it. If ad landing page dose not have ads, the access for the landing page will not be recorded in page_views.csv. So only 4.3% of ad may have ads in the landing page.\n \nOn the other hand if access information for landing page of ad clicks is found in page_views.csv, 95.8% of them are clicked.\nSuppose test data has same high precision, this feature would be useful."}],"metadata":{"_change_revision":0,"_is_fork":false,"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.5.2"}},"nbformat":4,"nbformat_minor":0}