{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Problem: \nFraud risk is everywhere, but for companies that advertise online, click fraud can happen at an overwhelming volume, resulting in misleading click data and wasted money. Ad channels can drive up costs by simply clicking on the ad at a large scale. With over 1 billion smart mobile devices in active use every month, China is the largest mobile market in the world and therefore suffers from huge volumes of fradulent traffic.\n\nTalkingData, China’s largest independent big data service platform, covers over 70% of active mobile devices nationwide. They handle 3 billion clicks per day, of which 90% are potentially fraudulent. Their current approach to prevent click fraud for app developers is to measure the journey of a user’s click across their portfolio, and flag IP addresses who produce lots of clicks, but never end up installing apps. With this information, they've built an IP blacklist and device blacklist.\n\nWhile successful, they want to always be one step ahead of fraudsters and have turned to the Kaggle community for help in further developing their solution. In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a user will download an app after clicking a mobile app ad. To support your modeling, they have provided a generous dataset covering approximately 200 million clicks over 4 days!","metadata":{}},{"cell_type":"markdown","source":"### Goal:\nPredicting the probabilities for different click_id's in the test set.**\n\nFor each click_id in the test set, we must predict a probability for the target is_attributed variable. The file should contain a header and have the following format:\n* click_id,is_attributed\n* 1,0.003\n* 2,0.001\n* 3,0.000\n* etc.","metadata":{}},{"cell_type":"markdown","source":"### Dataset Description\nEach row of the training data contains a click record, with the following features.\n\n- **ip**: ip address of click.\n- **app**: app id for marketing.\n- **device**: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n- **os**: os version id of user mobile phone\n- **channel**: channel id of mobile ad publisher\n- **click_time**: timestamp of click (UTC)\n- **attributed_time**: if user download the app for after clicking an ad, this is the time of the app download\n- **is_attributed**: the target that is to be predicted, indicating the app was downloaded\n    Note that ip, app, device, os, and channel are encoded.\n\nThe test data is similar, with the following differences:\n- **click_id**: reference for making predictions\n- **is_attributed**: not included","metadata":{}},{"cell_type":"markdown","source":"### Evaluation:\nSubmissions are evaluated on area under the ROC curve between the predicted probability and the observed target.","metadata":{}},{"cell_type":"markdown","source":"### TLDR; Summary\n- Total train samples are 184,903,890.\n- Total test samples are 18,790,469.\n- Total test supplement samples are 57,537,505. \n- Percentage of positive data: 0.2%\n- Used last 10M rows to down sample the data since the dataset is very big.\n- Used Kaggle kernel for training and generating the output for this problem \n- Used XGBoost for training and testing purposes","metadata":{}},{"cell_type":"markdown","source":"### Leaderboard Score:\n- **Public score**: 0.95613\n- **Private score**: 0.95558","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2021-08-27T01:06:55.930148Z","iopub.execute_input":"2021-08-27T01:06:55.930418Z","iopub.status.idle":"2021-08-27T01:06:55.936966Z","shell.execute_reply.started":"2021-08-27T01:06:55.930364Z","shell.execute_reply":"2021-08-27T01:06:55.936069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nimport matplotlib.pyplot as plt\nimport xgboost as xgb\nfrom xgboost import plot_importance\nimport gc\nimport time\nfrom datetime import datetime\n%config InlineBackend.figure_format = 'retina'\nplt.figure(figsize=(12,5))\nfrom sklearn.utils import resample","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.status.busy":"2021-08-27T01:06:56.465249Z","iopub.execute_input":"2021-08-27T01:06:56.465945Z","iopub.status.idle":"2021-08-27T01:06:57.691975Z","shell.execute_reply.started":"2021-08-27T01:06:56.465885Z","shell.execute_reply":"2021-08-27T01:06:57.691181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train dataset","metadata":{}},{"cell_type":"code","source":"train_path = \"../input/\" + 'train.csv'\ntrain_df = pd.read_csv(train_path, nrows = 10000)\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:06:57.997933Z","iopub.execute_input":"2021-08-27T01:06:57.998238Z","iopub.status.idle":"2021-08-27T01:06:58.063590Z","shell.execute_reply.started":"2021-08-27T01:06:57.998190Z","shell.execute_reply":"2021-08-27T01:06:58.062972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Test dataset","metadata":{}},{"cell_type":"code","source":"test_path = \"../input/\" + 'test.csv'\ntest_df = pd.read_csv(test_path, nrows = 10000)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:06:59.107526Z","iopub.execute_input":"2021-08-27T01:06:59.107799Z","iopub.status.idle":"2021-08-27T01:06:59.160736Z","shell.execute_reply.started":"2021-08-27T01:06:59.107742Z","shell.execute_reply":"2021-08-27T01:06:59.159812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Test supplement dataset","metadata":{}},{"cell_type":"code","source":"test_sup_path = \"../input/\" + 'test_supplement.csv'\ntest_sup_df = pd.read_csv(test_sup_path, nrows = 10000)\ntest_sup_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:07:00.246649Z","iopub.execute_input":"2021-08-27T01:07:00.246938Z","iopub.status.idle":"2021-08-27T01:07:00.294993Z","shell.execute_reply.started":"2021-08-27T01:07:00.246890Z","shell.execute_reply":"2021-08-27T01:07:00.294121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.dtypes","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:07:00.666954Z","iopub.execute_input":"2021-08-27T01:07:00.667227Z","iopub.status.idle":"2021-08-27T01:07:00.678840Z","shell.execute_reply.started":"2021-08-27T01:07:00.667178Z","shell.execute_reply":"2021-08-27T01:07:00.677735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_cols = ['ip','app','device','os','channel','click_time','is_attributed']\ntest_cols = ['ip','app','device','os','channel','click_time','click_id']\n\n# By default, pandas sets the dtype of integers to int64. In many cases, \n# this datatype takes up extra memory which is just not required.\n# Hence, memory reduction by changing the datatypes is very helpful.\n\ndtypes = {\n        'ip'            : 'uint32',\n        'app'           : 'uint16',\n        'device'        : 'uint16',\n        'os'            : 'uint16',\n        'channel'       : 'uint16',\n        'click_id'      : 'uint32',\n        'is_attributed' : 'uint8'\n        }","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:07:48.507018Z","iopub.execute_input":"2021-08-27T01:07:48.507292Z","iopub.status.idle":"2021-08-27T01:07:48.513987Z","shell.execute_reply.started":"2021-08-27T01:07:48.507240Z","shell.execute_reply":"2021-08-27T01:07:48.513021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading the last 10M rows for down sampling the data\ntrain_df = pd.read_csv( '../input/' + \"train.csv\", skiprows=range(1,123903891), nrows=61000000, usecols=train_cols, dtype=dtypes)\ntest_sup_df = pd.read_csv('../input/' + \"test_supplement.csv\", usecols=test_cols, dtype=dtypes)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:07:52.557335Z","iopub.execute_input":"2021-08-27T01:07:52.557601Z","iopub.status.idle":"2021-08-27T01:12:17.413866Z","shell.execute_reply.started":"2021-08-27T01:07:52.557555Z","shell.execute_reply":"2021-08-27T01:12:17.413151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Separate majority and minority classes\ndf_train_majority = train_df[train_df.is_attributed==0]\ndf_train_minority = train_df[train_df.is_attributed==1]","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:12:54.265612Z","iopub.execute_input":"2021-08-27T01:12:54.265926Z","iopub.status.idle":"2021-08-27T01:12:58.379661Z","shell.execute_reply.started":"2021-08-27T01:12:54.265872Z","shell.execute_reply":"2021-08-27T01:12:58.378001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(df_train_majority)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:12:58.383194Z","iopub.execute_input":"2021-08-27T01:12:58.383503Z","iopub.status.idle":"2021-08-27T01:12:58.388169Z","shell.execute_reply.started":"2021-08-27T01:12:58.383444Z","shell.execute_reply":"2021-08-27T01:12:58.387414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(df_train_minority)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:01.846671Z","iopub.execute_input":"2021-08-27T01:13:01.846969Z","iopub.status.idle":"2021-08-27T01:13:01.854760Z","shell.execute_reply.started":"2021-08-27T01:13:01.846916Z","shell.execute_reply":"2021-08-27T01:13:01.853785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Downsample majority class\ndf_majority_downsampled = resample(df_train_majority, \n                                 replace=False,    # sample without replacement\n                                 n_samples=10000000,     # to match minority class\n                                 random_state=123) # reproducible results","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:02.871227Z","iopub.execute_input":"2021-08-27T01:13:02.871598Z","iopub.status.idle":"2021-08-27T01:13:10.095427Z","shell.execute_reply.started":"2021-08-27T01:13:02.871542Z","shell.execute_reply":"2021-08-27T01:13:10.094619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Combine minority class with downsampled majority class\ndf_downsampled = pd.concat([df_majority_downsampled, df_train_minority])","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:12.241939Z","iopub.execute_input":"2021-08-27T01:13:12.242269Z","iopub.status.idle":"2021-08-27T01:13:12.597635Z","shell.execute_reply.started":"2021-08-27T01:13:12.242225Z","shell.execute_reply":"2021-08-27T01:13:12.596889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display new class counts\ndf_downsampled.is_attributed.value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:14.056372Z","iopub.execute_input":"2021-08-27T01:13:14.056638Z","iopub.status.idle":"2021-08-27T01:13:14.179218Z","shell.execute_reply.started":"2021-08-27T01:13:14.056592Z","shell.execute_reply":"2021-08-27T01:13:14.178226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Feature extraction using click_time column in the datasets\ndef feature_extraction(df):\n    df['date'] = pd.to_datetime(df['click_time'])\n    df['dayOfWeek'] = df['date'].dt.dayofweek.astype('uint16')\n    df['dayOfYear'] = df['date'].dt.dayofyear.astype('uint16')\n    df['hour'] = df['date'].dt.hour.astype('uint8')\n    df['min'] = df['date'].dt.minute.astype('uint8')\n    df['sec'] = df['date'].dt.second.astype('uint8')\n    df.drop(['date','click_time'], axis= 1, inplace=True)\n    return df","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:14.981806Z","iopub.execute_input":"2021-08-27T01:13:14.982080Z","iopub.status.idle":"2021-08-27T01:13:14.987344Z","shell.execute_reply.started":"2021-08-27T01:13:14.982030Z","shell.execute_reply":"2021-08-27T01:13:14.986637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df_train_minority, train_df, df_majority_downsampled\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:16.389628Z","iopub.execute_input":"2021-08-27T01:13:16.389904Z","iopub.status.idle":"2021-08-27T01:13:16.663279Z","shell.execute_reply.started":"2021-08-27T01:13:16.389857Z","shell.execute_reply":"2021-08-27T01:13:16.662280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_downsampled.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:17.365968Z","iopub.execute_input":"2021-08-27T01:13:17.366240Z","iopub.status.idle":"2021-08-27T01:13:17.397484Z","shell.execute_reply.started":"2021-08-27T01:13:17.366193Z","shell.execute_reply":"2021-08-27T01:13:17.396467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop the target values from train dataset \ny = df_downsampled['is_attributed']\ndf_downsampled.drop(['is_attributed'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:20.133196Z","iopub.execute_input":"2021-08-27T01:13:20.133471Z","iopub.status.idle":"2021-08-27T01:13:20.389213Z","shell.execute_reply.started":"2021-08-27T01:13:20.133424Z","shell.execute_reply":"2021-08-27T01:13:20.388468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop the click_time from the test data\ntest_sup_df.drop(['click_id'], axis=1, inplace=True)\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:20.590673Z","iopub.execute_input":"2021-08-27T01:13:20.590986Z","iopub.status.idle":"2021-08-27T01:13:22.613295Z","shell.execute_reply.started":"2021-08-27T01:13:20.590934Z","shell.execute_reply":"2021-08-27T01:13:22.612678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merging the supplement data with test data\nrows_train = df_downsampled.shape[0]\nmerge_df = pd.concat([df_downsampled, test_sup_df])\nmerge_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:22.614349Z","iopub.execute_input":"2021-08-27T01:13:22.614601Z","iopub.status.idle":"2021-08-27T01:13:24.640552Z","shell.execute_reply.started":"2021-08-27T01:13:22.614556Z","shell.execute_reply":"2021-08-27T01:13:24.639766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Length of combine dataset: ', len(merge_df))","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:24.641641Z","iopub.execute_input":"2021-08-27T01:13:24.641915Z","iopub.status.idle":"2021-08-27T01:13:24.647757Z","shell.execute_reply.started":"2021-08-27T01:13:24.641868Z","shell.execute_reply":"2021-08-27T01:13:24.645218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df_downsampled, test_sup_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:24.649281Z","iopub.execute_input":"2021-08-27T01:13:24.649602Z","iopub.status.idle":"2021-08-27T01:13:24.913453Z","shell.execute_reply.started":"2021-08-27T01:13:24.649526Z","shell.execute_reply":"2021-08-27T01:13:24.912714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Group by ip to count the number of clicks\nip_groups = merge_df.groupby(['ip'])['channel'].count().reset_index(name = 'clicks_by_ip')\nprint(ip_groups)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:24.915887Z","iopub.execute_input":"2021-08-27T01:13:24.916498Z","iopub.status.idle":"2021-08-27T01:13:29.024892Z","shell.execute_reply.started":"2021-08-27T01:13:24.916447Z","shell.execute_reply":"2021-08-27T01:13:29.023887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"merge_df = pd.merge(merge_df, ip_groups, on='ip', how='left', sort=False)\nprint(merge_df.head())","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:29.026103Z","iopub.execute_input":"2021-08-27T01:13:29.026673Z","iopub.status.idle":"2021-08-27T01:13:49.481409Z","shell.execute_reply.started":"2021-08-27T01:13:29.026527Z","shell.execute_reply":"2021-08-27T01:13:49.480663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"merge_df['clicks_by_ip'] = merge_df['clicks_by_ip'].astype('uint16')\nmerge_df.drop('ip', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:49.482521Z","iopub.execute_input":"2021-08-27T01:13:49.482951Z","iopub.status.idle":"2021-08-27T01:13:53.191906Z","shell.execute_reply.started":"2021-08-27T01:13:49.482903Z","shell.execute_reply":"2021-08-27T01:13:53.191111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = merge_df[:rows_train]\ntest_df = merge_df[rows_train:]","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:53.193050Z","iopub.execute_input":"2021-08-27T01:13:53.193305Z","iopub.status.idle":"2021-08-27T01:13:53.200117Z","shell.execute_reply.started":"2021-08-27T01:13:53.193264Z","shell.execute_reply":"2021-08-27T01:13:53.199280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del test_df,merge_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:53.201463Z","iopub.execute_input":"2021-08-27T01:13:53.201887Z","iopub.status.idle":"2021-08-27T01:13:53.259488Z","shell.execute_reply.started":"2021-08-27T01:13:53.201706Z","shell.execute_reply":"2021-08-27T01:13:53.258819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = feature_extraction(train_df)\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:13:53.261045Z","iopub.execute_input":"2021-08-27T01:13:53.261343Z","iopub.status.idle":"2021-08-27T01:14:00.075513Z","shell.execute_reply.started":"2021-08-27T01:13:53.261284Z","shell.execute_reply":"2021-08-27T01:14:00.074926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:14:00.076828Z","iopub.execute_input":"2021-08-27T01:14:00.077094Z","iopub.status.idle":"2021-08-27T01:14:00.099844Z","shell.execute_reply.started":"2021-08-27T01:14:00.077050Z","shell.execute_reply":"2021-08-27T01:14:00.098758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the params for xgboost model\nparams = {'eta': 0.3,\n          'tree_method': \"hist\",\n          'grow_policy': \"lossguide\",\n          'max_leaves': 1400,  \n          'max_depth': 0, \n          'subsample': 0.9, \n          'colsample_bytree': 0.7, \n          'colsample_bylevel':0.7,\n          'min_child_weight':0,\n          'alpha':4,\n          'objective': 'binary:logistic', \n          'scale_pos_weight':9,\n          'eval_metric': 'auc', \n          'nthread':8,\n          'random_state': 99, \n          'silent': True}","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:14:00.101495Z","iopub.execute_input":"2021-08-27T01:14:00.101946Z","iopub.status.idle":"2021-08-27T01:14:00.111016Z","shell.execute_reply.started":"2021-08-27T01:14:00.101747Z","shell.execute_reply":"2021-08-27T01:14:00.110110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(train_df, y, test_size=0.1, stratify=y, random_state=99)\ndtrain = xgb.DMatrix(X_train, y_train)\ndvalid = xgb.DMatrix(X_test, y_test)\ndel X_train, y_train\ngc.collect()\nwatchlist = [(dtrain, 'train'), (dvalid, 'valid')]\nmodel = xgb.train(params, dtrain, 200, watchlist, maximize=True, early_stopping_rounds = 20, verbose_eval=5)\ndel dvalid\n\nprint(\"Validating...\")\ncheck = model.predict(xgb.DMatrix(X_test), ntree_limit=model.best_iteration+1)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:14:00.112579Z","iopub.execute_input":"2021-08-27T01:14:00.113291Z","iopub.status.idle":"2021-08-27T01:19:51.824572Z","shell.execute_reply.started":"2021-08-27T01:14:00.113192Z","shell.execute_reply":"2021-08-27T01:19:51.823712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import roc_curve, auc\n# Compute micro-average ROC curve and ROC area\nfpr, tpr, _ = roc_curve(y_test.values, check)\nroc_auc = auc(fpr, tpr)\nplt.figure()\nlw = 2\nplt.plot(fpr, tpr, color='darkorange',\n         lw=lw, label='ROC curve (area = %0.2f)' % roc_auc)\nplt.plot([0, 1], [0, 1], color='navy', lw=lw, linestyle='--')\nplt.xlim([-0.02, 1.0])\nplt.ylim([0.0, 1.05])\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('ROC curve')\nplt.legend(loc=\"lower right\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:19:51.825659Z","iopub.execute_input":"2021-08-27T01:19:51.825949Z","iopub.status.idle":"2021-08-27T01:19:52.427806Z","shell.execute_reply.started":"2021-08-27T01:19:51.825904Z","shell.execute_reply":"2021-08-27T01:19:52.427000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the feature importances from xgboost\nplot_importance(model)\nplt.gcf().savefig('xgb_feature_importance_v2_downsample_roc.png')","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:19:52.428975Z","iopub.execute_input":"2021-08-27T01:19:52.429235Z","iopub.status.idle":"2021-08-27T01:19:53.173557Z","shell.execute_reply.started":"2021-08-27T01:19:52.429190Z","shell.execute_reply":"2021-08-27T01:19:53.172769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the test dataset for prediction\ntest_cols  = ['ip', 'app', 'device', 'os', 'channel', 'click_time', 'click_id']\ntest_df = pd.read_csv('../input/' +\"test.csv\", usecols=test_cols, dtype=dtypes)\ntest_df = pd.merge(test_df, ip_groups, on='ip', how='left', sort=False)\ndel ip_groups\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:20:18.725705Z","iopub.execute_input":"2021-08-27T01:20:18.726032Z","iopub.status.idle":"2021-08-27T01:20:48.553155Z","shell.execute_reply.started":"2021-08-27T01:20:18.725979Z","shell.execute_reply":"2021-08-27T01:20:48.552511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:20:48.554302Z","iopub.execute_input":"2021-08-27T01:20:48.554554Z","iopub.status.idle":"2021-08-27T01:20:48.578680Z","shell.execute_reply.started":"2021-08-27T01:20:48.554509Z","shell.execute_reply":"2021-08-27T01:20:48.578120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del check, X_test, y_test","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:20:48.579991Z","iopub.execute_input":"2021-08-27T01:20:48.580400Z","iopub.status.idle":"2021-08-27T01:20:48.584350Z","shell.execute_reply.started":"2021-08-27T01:20:48.580230Z","shell.execute_reply":"2021-08-27T01:20:48.583262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a dataframe for submission\nsubmission_df = pd.DataFrame()\nsubmission_df['click_id'] = test_df['click_id'].astype('int')\n\ntest_df['clicks_by_ip'] = test_df['clicks_by_ip'].astype('uint16')\ntest_df = feature_extraction(test_df)\ntest_df.drop(['click_id', 'ip'], axis=1, inplace=True)\ndtest = xgb.DMatrix(test_df)\ndel test_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:20:48.585722Z","iopub.execute_input":"2021-08-27T01:20:48.586182Z","iopub.status.idle":"2021-08-27T01:21:08.476574Z","shell.execute_reply.started":"2021-08-27T01:20:48.585976Z","shell.execute_reply":"2021-08-27T01:21:08.475884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:21:08.477633Z","iopub.execute_input":"2021-08-27T01:21:08.478092Z","iopub.status.idle":"2021-08-27T01:21:08.489248Z","shell.execute_reply.started":"2021-08-27T01:21:08.478043Z","shell.execute_reply":"2021-08-27T01:21:08.488309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:22:07.808967Z","iopub.execute_input":"2021-08-27T01:22:07.809241Z","iopub.status.idle":"2021-08-27T01:22:07.862941Z","shell.execute_reply.started":"2021-08-27T01:22:07.809192Z","shell.execute_reply":"2021-08-27T01:22:07.862143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get predictions from the best iteration with model.best_ntree_limit.\nsubmission_df['is_attributed'] = model.predict(dtest, ntree_limit=model.best_ntree_limit)\nsubmission_df.to_csv('xgb_submission_roc.csv', float_format='%.8f', index=False)","metadata":{"execution":{"iopub.status.busy":"2021-08-27T01:22:10.194703Z","iopub.execute_input":"2021-08-27T01:22:10.195006Z","iopub.status.idle":"2021-08-27T01:24:10.682452Z","shell.execute_reply.started":"2021-08-27T01:22:10.194957Z","shell.execute_reply":"2021-08-27T01:24:10.681603Z"},"trusted":true},"execution_count":null,"outputs":[]}]}