{"cells":[{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"\n#Import necessary packages\nimport numpy as np \nimport pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nimport os\nprint(os.listdir(\"../input\"))\n\nfrom IPython.core.interactiveshell import InteractiveShell\nInteractiveShell.ast_node_interactivity = \"all\"","execution_count":15,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_spl = pd.read_csv(\"../input/train_sample.csv\")\ntrain_spl.info()","execution_count":7,"outputs":[]},{"metadata":{"_uuid":"9734fc4725777dab74e0036908aefb502331e329"},"cell_type":"markdown","source":"1. Overall information on sample training dataset:\n            *  100,000 records, 8 atrrtibutes in total\n            *  labels: 0 for not download the app; 1 for download the app\n2.  Two subset specifically for records which download the app (dld_train_spl) and not download the app(no_dld_train_spl)"},{"metadata":{"trusted":true,"_uuid":"15a15e5188cd4bbac7090e1c20dc9c198a451062"},"cell_type":"code","source":"#Create subset for download apps and NOT download apps\ndld_train_spl = train_spl[train_spl['is_attributed']==1]\ndld_train_spl.info()\n\nno_dld_train_spl = train_spl[train_spl['is_attributed']==0]\nno_dld_train_spl.info()","execution_count":8,"outputs":[]},{"metadata":{"_uuid":"3070bf27bf82a1347826ba52ee773b55555b84a1"},"cell_type":"markdown","source":"* Among 100,000 records in the train sample dataset:\n            251 records download apps\n            99,749 records NOT download apps"},{"metadata":{"_uuid":"95c3c4958dd48a354ee446e9abaa6c3db8a9ddd5"},"cell_type":"markdown","source":"#### Take a look of ip address for those not download app."},{"metadata":{"trusted":true,"_uuid":"ba8659d6faea4be91477df2214b11c1aafc7b7e4"},"cell_type":"code","source":"ip_no_dld = no_dld_train_spl[\"ip\"].value_counts()\nip_no_dld[:10]","execution_count":29,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1152336ccf19afec9a5339b00e9c9420acd6b10f"},"cell_type":"code","source":"#Plot the ip addresses which are not download apps\nsns.countplot(x = \"ip\",  data = no_dld_train_spl,\\\n             order = ip_no_dld[:10].index).set(\\\n            xlabel = \"ip address not download the app\")","execution_count":27,"outputs":[]},{"metadata":{"_uuid":"a270c477ce7dd362233fba7ba55d3c04bd1c1779"},"cell_type":"markdown","source":"* From the ip address amount, top 10 ip addresses clicked the ad but not download. they are: 5348, 5314, 73487, 73516, 53454, 114276, 26995, 95766, 17149, 100275\n* For the top 10 ip addresses which not download the app, find them from the dataset which download apps."},{"metadata":{"trusted":true,"_uuid":"1ce973f8f00ecdebada334c58093f0c364624b22"},"cell_type":"code","source":"#Extract the top 10 ip addresses(not download the app) from the download dataset\nip_in_dld = dld_train_spl.loc[dld_train_spl['ip'].isin(ip_no_dld[:10].index)]\nip_in_dld.shape\n\n#Group by ip\ngroupby_ip_dld = ip_in_dld.groupby(\"ip\").size()\ngroupby_ip_dld[:10]","execution_count":30,"outputs":[]},{"metadata":{"_uuid":"b35054d73f6acf61899c8d142907efe14f61f017"},"cell_type":"markdown","source":"*  Then we can take a look of the download rate(download amount/total click amount) for each top 10 ip addresses."},{"metadata":{"trusted":true,"_uuid":"80a94f4cf89d826160c7fc8a40dcfd10291e417a"},"cell_type":"code","source":"#In train sample set, extract the ip addresses whcih not download the app\n#In this way, we can find the total click amount\nip_no_dld_in_all = train_spl.loc[train_spl['ip'].isin(ip_no_dld[:10].index)]\nip_no_dld_in_all.shape\n\n#Group by ip\ngroupby_no_dld_ip=ip_no_dld_in_all.groupby(\"ip\").size()\ngroupby_no_dld_ip\n\n#Plot the total click amount by ip addresses\nsns.countplot(x = 'ip', data = ip_no_dld_in_all,\\\n              order = ip_no_dld[:10].index).set(\\\n            xlabel = \"ip address click amount\")","execution_count":31,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6b2a047e657ee1e8785d899f8739a91959f9ee75"},"cell_type":"code","source":"#Merge the download times and click times based on same ip addresses\nclick_dld_ = pd.merge(groupby_no_dld_ip.reset_index(),ip_no_dld[:10].reset_index(),\\\n                      left_on = \"ip\", right_on = \"index\").iloc[:,[0,1,3]]\nclick_dld_.columns = ['ip','click times', 'not download times']\nclick_dld_['download times'] = click_dld_['click times'] - click_dld_['not download times']\n\n#Calculate the download rate\nclick_dld_['download rate'] = click_dld_['download times']/ click_dld_['click times']\nclick_dld_\n\n#Average download rate for the 10 ip addresses\nprint(\"Average download rate is: \" , click_dld_[\"download rate\"].mean())","execution_count":32,"outputs":[]},{"metadata":{"_uuid":"73a5d4d6df9f9fb7e3c5f6c3e61926deff156280"},"cell_type":"markdown","source":"* For those top 10 ip addresses(not download the app), they have the highest click amount, but only have extremely less download amounts.\n* In other words, for instance, for ip \"5314\", this ip address has the click times of 640, but  download only 3 times,  it means the download rate is only 3/640 = 0.004687"},{"metadata":{"_uuid":"bf1e5d1e334a755628263c924b99ae3e798fc26c"},"cell_type":"markdown","source":"#### Take a look of ip address for those download app.\nIn order to have a better contrast of the click times and download times, we can use the download ip address as a contrast. "},{"metadata":{"trusted":true,"_uuid":"5b2aca4ac5e00c351dfb889307fdb052d86e8ea0"},"cell_type":"code","source":"#Top 10 ip addresses which download the apps\nip_dld = dld_train_spl[\"ip\"].value_counts()\nip_dld[:10]\n\nip_dld_in_all = train_spl.loc[train_spl['ip'].isin(ip_dld[:10].index)]\n\nsns.countplot(x = 'ip', data = ip_dld_in_all,\\\n              order = ip_dld[:10].index).set(\\\n            xlabel = \"ip address click amount\")","execution_count":43,"outputs":[]},{"metadata":{"_uuid":"76894177e4b36cb0e78860e6f1c7ecd20b7a87cc"},"cell_type":"markdown","source":"Ip addresses \"5314\" has a overwhelming amount click times compared with other ip addresses, but extremely less download times, so it'll affect the whole download rate for other ip addresses. So as ip address \"100275\". These two \"spetical\" ip address should be outlier.\n\nTherefore, the following analysis will not take these two ip addresses into account"},{"metadata":{"trusted":true,"_uuid":"2a97cb38f8b2f49e7e869495ea0570d32b91d1ca"},"cell_type":"code","source":"#In train sample set, extract the ip addresses whcih download the app\nip_dld_in_all = train_spl.loc[train_spl['ip'].isin(ip_dld[2:12].index)]\n\n#Group by ip\ngroupby_dld_ip=ip_dld_in_all.groupby(\"ip\").size()\ngroupby_dld_ip\n\n#Plot the total click amount by ip addresses\nsns.countplot(x = 'ip', data = ip_dld_in_all,\\\n              order = ip_dld[2:12].index).set(\\\n            xlabel = \"ip address click amount\")","execution_count":50,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0ae43b7c5c249d0f481a6ba433948c0cd9421446"},"cell_type":"code","source":"#Merge the download times and click times based on same ip addresses\nclick_dld = pd.merge(groupby_dld_ip.reset_index(),ip_dld[2:12].reset_index(),\\\n                      left_on = \"ip\", right_on = \"index\").iloc[:,[0,1,3]]\nclick_dld.columns = ['ip','click times', 'download times']\n\n#Calculate the download rate\nclick_dld['download rate'] = click_dld['download times']/ click_dld['click times']\nclick_dld\n\n#Average download rate for the 10 ip addresses\nprint(\"Average download rate is: \" , click_dld[\"download rate\"].mean())","execution_count":51,"outputs":[]},{"metadata":{"_uuid":"348fa2a04fe56f7a134a91c86eb01976681f13b8"},"cell_type":"markdown","source":"The  download rate for an average in the 10 ip addresses(download the app) is 0.5683. \n\nIn general,  each ip address usually click no more than 5 times, download only 1 time."},{"metadata":{"_uuid":"818601e02610b97e1fb8f8ead5b5066e5c25e650"},"cell_type":"markdown","source":"**Summary:**\n\nBased on the download rate comparing with download and NOT download the app, we assume a hypothesis that: if all the clicks are not fradulent, then we could expect a positive linear relationship between the click times and download times. \n\nIn other words, if an ip adress has a huge click amount but  download amount is extremely sparse, it's more likely to be frau**dulent than a genuine one."},{"metadata":{"collapsed":true,"trusted":true,"_uuid":"1cdade1976a4f05dd0d03402e6c1bee61c1a8ef3"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.6.4"}},"nbformat":4,"nbformat_minor":1}