{"cells":[{"metadata":{"_uuid":"4f47cb6fdcc92f3972bb2c4cb313f76754afc42b","_cell_guid":"8c2014f0-b678-43dc-b958-0c0168285795"},"cell_type":"markdown","source":"I wondered how many prediction we might get wrong with duplicate values in the test data. These duplicate clicks might be diiffer in microsecond values. So I comapred day 9 train data duplicate click's target values which have different target with test data for percent of wrong predictions. "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","collapsed":true,"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":false},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib\n%matplotlib inline \nimport time\nimport gc\nfrom datetime import datetime\nimport os\n\nos.listdir(\"../input\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"code","source":"dtypes = {\n        'ip'            : 'uint32',\n        'app'           : 'uint16',\n        'device'        : 'uint16',\n        'os'            : 'uint16',\n        'channel'       : 'uint16',\n        'is_attributed' : 'uint8',\n        'click_id'      : 'uint32'\n        }\nprint('loading train data...')\ntrain_df = pd.read_csv(\"../input/train.csv\", dtype=dtypes,skiprows=range(1,131886954),\n                       usecols=['ip','app','device','os', 'channel', 'click_time', 'is_attributed'])\nlen_train = len(train_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b37a6d67a16d85e90d82a91b56cacc30838ff5d9","collapsed":true,"_cell_guid":"90a38c20-6ee3-4099-87d8-62ddeecc75b5","trusted":false},"cell_type":"code","source":"train_df['click_time'] = pd.to_datetime(train_df['click_time'])\ndup = train_df[train_df.duplicated([\"ip\", 'app', 'channel', 'device', 'os', 'click_time'], False)]\ngroup = dup.groupby(['ip', 'app', 'channel', 'device', 'os', 'click_time']).is_attributed.mean().reset_index().rename(index=str, columns={'is_attributed': 'mean'})\ndup = dup.merge(group, on=['ip', 'app', 'channel', 'device', 'os', 'click_time'], how='left')\ndel train_df\ndel group\ngc.collect()\nlen_dup = len(dup)\nprint('Number of Duplicate clicks in train data: ', len_dup)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f323da913b6dcd4b92ff29b432a139934d3003df","_cell_guid":"7b5d60c7-ffaf-4121-b723-c7f315d37833"},"cell_type":"markdown","source":"Get the dupplicate clicks with different target values. Cannot simply drop duplicate from 'dup' because there are some clicks with more than 2 duplicates."},{"metadata":{"_uuid":"9214515509d2bdf600c0dec56e240c44202163b8","collapsed":true,"_cell_guid":"c80ddab5-1d6c-4c43-a271-2db67ce16348","trusted":false},"cell_type":"code","source":"dup_diff_target = dup[(dup['mean']!=0.0) & (dup['mean'] !=1.0)]\nlen_dup_diff_target = len(dup_diff_target)\nprint('NUmber of duplicate clicks with different target values in train data: ', len_dup_diff_target)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"173f9b72dec77dc6129a73aca66975fb1c71436d","collapsed":true,"_cell_guid":"6194eb42-04f7-4334-937b-b8916c42c4f8","trusted":false},"cell_type":"code","source":"dup_diff_target.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"571d62de3c6914c9641b27d78594c9efd95d8033","_cell_guid":"f73e82a7-93c0-4f6d-bab9-054b8090c987"},"cell_type":"markdown","source":"Percent of duplicate clicks with different target values"},{"metadata":{"_uuid":"f51b73d785dd4ef09f698c11121d13f08a42f34c","collapsed":true,"_cell_guid":"1484e6c1-e2d9-4681-b201-11160bde95a2","trusted":false},"cell_type":"code","source":"len_dup_diff_target/len_dup","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fc98b774690c341592f8d14a5b678570c81de64a","collapsed":true,"_cell_guid":"5aea1672-ee09-4cda-9be5-d1980d24a9b7","trusted":false},"cell_type":"code","source":"print('loading test supplement data...')\ntest_sup_df = pd.read_csv(\"../input/test_supplement.csv\", dtype=dtypes, usecols=['ip','app','device','os', 'channel', 'click_time', 'click_id'])\ntest_sup_df['click_time'] = pd.to_datetime(test_sup_df['click_time'])\ntest_sup_dup = test_sup_df[test_sup_df.duplicated([\"ip\", 'app', 'channel', 'device', 'os', 'click_time'], False)]\nlen_test_sup = len(test_sup_df)\ndel test_sup_df\ngc.collect()\nlen_test_sup_dup = len(test_sup_dup)\nprint('Number of Duplicate clicks in test supplement data: ', len_test_sup_dup)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ca4d930c47d62f9de140f6df0f9c7774d26e528e","collapsed":true,"_cell_guid":"49a95682-6f0b-4efb-bf71-b692302aa39e","trusted":false},"cell_type":"code","source":"print('loading test data...')\ntest_df = pd.read_csv(\"../input/test.csv\", dtype=dtypes, usecols=['ip','app','device','os', 'channel', 'click_time', 'click_id'])\ntest_df['click_time'] = pd.to_datetime(test_df['click_time'])\ntest_dup = test_df[test_df.duplicated([\"ip\", 'app', 'channel', 'device', 'os', 'click_time'], False)]\ndel test_df\ngc.collect()\nlen_test_dup = len(test_dup)\nprint('Number of Duplicate clicks in test(For submission) data: ', len_test_dup)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"980690cbaf53c4ceca11e7cd541fed92cca65450","collapsed":true,"_cell_guid":"92cd21dd-da2b-4cdf-9723-1c5a644ecd4d","trusted":false},"cell_type":"code","source":"percent_train_dup = len_dup/len_train\nprint(\"percent of dupplicate in train data: \", percent_train_dup)\n\npercent_test_dup = len_test_sup_dup/len_test_sup\nprint(\"percent of dupplicate in test data: \", percent_test_dup)\n\npercent_of_test_csv_dup = len_test_dup/len_test_sup_dup\nprint(\"Percent of test.csv duplicates in test_supplement.csv: \", percent_of_test_csv_dup)\n\nmaybe_worng_preds = (len_test_sup_dup*len_dup_diff_target)/len_dup\nprint(\"Number of wrong(may be) predictions for test supplement: \", maybe_worng_preds)\n\nmaybe_worng_preds_for_submission = percent_of_test_csv_dup*maybe_worng_preds\nprint(\"Number of wrong(may be) predictions for submission: \", maybe_worng_preds_for_submission)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0810c2b2724b3d610631a9dbe0616dcb6c286ab2","collapsed":true,"_cell_guid":"8370876a-d90b-4e1e-b3c1-24a2e5d0c2ad","trusted":false},"cell_type":"code","source":"group = dup_diff_target.groupby([\"ip\", 'app', 'channel', 'device', 'os', 'click_time'])\nfirst = group.nth(0).is_attributed.value_counts()\nsecond = group.nth(1).is_attributed.value_counts()\nthird = group.nth(2).is_attributed.value_counts()\nprint('First click target value counts:\\n', first)\nprint('Second click target value counts:\\n', second)\nprint('Third click target value counts:\\n', third)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a33c3f5f0d8c462d30f3a33f32732df96b0e1413","_cell_guid":"2f30d15d-6555-4de0-a80e-a15f6b8fe59d"},"cell_type":"markdown","source":"About 2/3rd of the taget values of 'duplicate clicks with diffrent tagets' have there taget values as \"1\" for lower index. which means first click has more probability than next one which is microseconds appart.  We will have more chances if we can find those duplicates which are going to have different target values among all the duplicate clicks."},{"metadata":{"_uuid":"fe12cb513f1120fa269b614218241a8bb945d321","_cell_guid":"f00c2a6a-3f8a-4af1-8755-35dc8800876c"},"cell_type":"markdown","source":"I think I made so many assumptions here. One of it is distribution of duplicate clicks for test and test_supplement.\n\nI have done it for only day 9 of train data\n\nand i might be wrong somewhere, if so please let me know"},{"metadata":{"_uuid":"538c5d371ec97e84ba6f649059dff5fee4334c6f","collapsed":true,"_cell_guid":"650bd68e-4989-4694-8ba4-10cb44164197","trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}