{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TalkingData: Fraudulent Click Prediction","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we will apply various boosting algorithms to solve an interesting classification problem from the domain of 'digital fraud'.\n\nThe analysis is divided into the following sections:\n- Understanding the business problem\n- Understanding and exploring the data\n- Feature engineering: Creating new features\n- Model building and evaluation: AdaBoost\n- Modelling building and evaluation: Gradient Boosting\n- Modelling building and evaluation: XGBoost\n","metadata":{}},{"cell_type":"markdown","source":"## Understanding the Business Problem\n\n<a href=\"https://www.talkingdata.com/\">TalkingData</a> is a Chinese big data company, and one of their areas of expertise is mobile advertisements.\n\nIn mobile advertisements, **click fraud** is a major source of losses. Click fraud is the practice of repeatedly clicking on an advertisement hosted on a website with the intention of generating revenue for the host website or draining revenue from the advertiser.\n\nIn this case, TalkingData happens to be serving the advertisers (their clients). TalkingData cover a whopping **approx. 70% of the active mobile devices in China**, of which 90% are potentially fraudulent (i.e. the user is actually not going to download the app after clicking).\n\nYou can imagine the amount of money they can help clients save if they are able to predict whether a given click is fraudulent (or equivalently, whether a given click will result in a download). \n\nTheir current approach to solve this problem is that they've generated a blacklist of IP addresses - those IPs which produce lots of clicks, but never install any apps. Now, they want to try some advanced techniques to predict the probability of a click being genuine/fraud.\n\nIn this problem, we will use the features associated with clicks, such as IP address, operating system, device type, time of click etc. to predict the probability of a click being fraud.","metadata":{}},{"cell_type":"markdown","source":"## Understanding and Exploring the Data\n\nThe data contains observations of about 240 million clicks, and whether a given click resulted in a download or not (1/0). \n\nOn Kaggle, the data is split into train.csv and train_sample.csv (100,000 observations). We'll use the smaller train_sample.csv in this notebook for speed, though while training the model for Kaggle submissions, the full training data will obviously produce better results.\n\nThe detailed data dictionary is mentioned here:\n- ```ip```: ip address of click.\n- ```app```: app id for marketing.\n- ```device```: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n- ```os```: os version id of user mobile phone\n- ```channel```: channel id of mobile ad publisher\n- ```click_time```: timestamp of click (UTC)\n- ```attributed_time```: if user download the app for after clicking an ad, this is the time of the app download\n- ```is_attributed```: the target that is to be predicted, indicating the app was downloaded\n\nLet's try finding some useful trends in the data.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport sklearn\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import KFold\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom sklearn import metrics\n\nimport xgboost as xgb\nfrom xgboost import XGBClassifier\nfrom xgboost import plot_importance\nimport gc # for deleting unused variables\n\n%matplotlib inline\n\nimport os\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:17.910070Z","iopub.execute_input":"2023-07-03T11:24:17.910459Z","iopub.status.idle":"2023-07-03T11:24:19.932196Z","shell.execute_reply.started":"2023-07-03T11:24:17.910414Z","shell.execute_reply":"2023-07-03T11:24:19.931498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Reading the Data  \n\nThe code below reads the train_sample.csv file if you set testing = True, else reads the full train.csv file. You can read the sample while tuning the model etc., and then run the model on the full data once done.\n\n#### Important Note: Save memory when the data is huge\n\nSince the training data is quite huge, the program will be quite slow if you don't consciously follow some best practices to save memory. This notebook demonstrates some of those practices. ","metadata":{}},{"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-07-03T11:24:19.933683Z","iopub.execute_input":"2023-07-03T11:24:19.934232Z","iopub.status.idle":"2023-07-03T11:24:19.941371Z","shell.execute_reply.started":"2023-07-03T11:24:19.934205Z","shell.execute_reply":"2023-07-03T11:24:19.940277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dtypes = {\n        'ip'            : 'uint16',\n        'app'           : 'uint16',\n        'device'        : 'uint16',\n        'os'            : 'uint16',\n        'channel'       : 'uint16',\n        'is_attributed' : 'uint8',\n        'click_id'      : 'uint32' # note that click_id is only in test data, not training data\n        }\n\ncolnames=['ip','app','device','os', 'channel', 'click_time', 'is_attributed']\n\ntrain_sample = pd.read_csv('/kaggle/input/talkingdata-adtracking-fraud-detection/train_sample.csv',dtype=dtypes,usecols=colnames)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:19.942501Z","iopub.execute_input":"2023-07-03T11:24:19.943278Z","iopub.status.idle":"2023-07-03T11:24:20.101486Z","shell.execute_reply.started":"2023-07-03T11:24:19.943245Z","shell.execute_reply":"2023-07-03T11:24:20.100748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_sample)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.103795Z","iopub.execute_input":"2023-07-03T11:24:20.104361Z","iopub.status.idle":"2023-07-03T11:24:20.110832Z","shell.execute_reply.started":"2023-07-03T11:24:20.104333Z","shell.execute_reply":"2023-07-03T11:24:20.110164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample.memory_usage()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.112106Z","iopub.execute_input":"2023-07-03T11:24:20.112646Z","iopub.status.idle":"2023-07-03T11:24:20.131194Z","shell.execute_reply.started":"2023-07-03T11:24:20.112618Z","shell.execute_reply":"2023-07-03T11:24:20.130359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# space used by training data\nprint('Training dataset uses {0} MB'.format(train_sample.memory_usage().sum()/1024**2))","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.132449Z","iopub.execute_input":"2023-07-03T11:24:20.132977Z","iopub.status.idle":"2023-07-03T11:24:20.142464Z","shell.execute_reply.started":"2023-07-03T11:24:20.132948Z","shell.execute_reply":"2023-07-03T11:24:20.141750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.143752Z","iopub.execute_input":"2023-07-03T11:24:20.144303Z","iopub.status.idle":"2023-07-03T11:24:20.174154Z","shell.execute_reply.started":"2023-07-03T11:24:20.144276Z","shell.execute_reply":"2023-07-03T11:24:20.173472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Exploring the Data - Univariate Analysis\n","metadata":{}},{"cell_type":"markdown","source":"Let's now understand and explore the data. Let's start with understanding the size and data types of the train_sample data.","metadata":{}},{"cell_type":"code","source":"# look at non-null values, number of entries etc.\n# there are no missing values\ntrain_sample.info()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.175934Z","iopub.execute_input":"2023-07-03T11:24:20.176997Z","iopub.status.idle":"2023-07-03T11:24:20.244728Z","shell.execute_reply.started":"2023-07-03T11:24:20.176952Z","shell.execute_reply":"2023-07-03T11:24:20.243346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Basic exploratory analysis \n\n# Number of unique values in each column\n\ndef fraction_unique(x):\n    return len(train_sample[x].unique())\n\nnumber_unique_vals = {x : fraction_unique(x) for x in train_sample.columns}\nnumber_unique_vals","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.246002Z","iopub.execute_input":"2023-07-03T11:24:20.246551Z","iopub.status.idle":"2023-07-03T11:24:20.282941Z","shell.execute_reply.started":"2023-07-03T11:24:20.246515Z","shell.execute_reply":"2023-07-03T11:24:20.282134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# All columns apart from click time are originally int type, \n# though note that they are all actually categorical \ntrain_sample.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.286286Z","iopub.execute_input":"2023-07-03T11:24:20.287611Z","iopub.status.idle":"2023-07-03T11:24:20.295261Z","shell.execute_reply.started":"2023-07-03T11:24:20.287550Z","shell.execute_reply":"2023-07-03T11:24:20.294484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are certain 'apps' which have quite high number of instances/rows (each row is a click). The plot below shows this. ","metadata":{}},{"cell_type":"code","source":"# distribution of 'app' \n# some 'apps' have a disproportionately high number of clicks (>15k), and some are very rare (3-4)\nplt.figure(figsize=(14,8))\nsns.countplot(x=\"app\",data=train_sample)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:20.296326Z","iopub.execute_input":"2023-07-03T11:24:20.297078Z","iopub.status.idle":"2023-07-03T11:24:21.845040Z","shell.execute_reply.started":"2023-07-03T11:24:20.297052Z","shell.execute_reply":"2023-07-03T11:24:21.843858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# distribution of 'device' \n# this is expected because a few popular devices are used heavily\nplt.figure(figsize=(14, 8))\nsns.countplot(x=\"device\", data=train_sample)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:21.846358Z","iopub.execute_input":"2023-07-03T11:24:21.846709Z","iopub.status.idle":"2023-07-03T11:24:22.694134Z","shell.execute_reply.started":"2023-07-03T11:24:21.846681Z","shell.execute_reply":"2023-07-03T11:24:22.693178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# channel: various channels get clicks in comparable quantities\nplt.figure(figsize=(14, 8))\nsns.countplot(x=\"channel\", data=train_sample)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:22.695377Z","iopub.execute_input":"2023-07-03T11:24:22.695743Z","iopub.status.idle":"2023-07-03T11:24:23.940358Z","shell.execute_reply.started":"2023-07-03T11:24:22.695711Z","shell.execute_reply":"2023-07-03T11:24:23.938916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# os: there are a couple commos OSes (android and ios?), though some are rare and can indicate suspicion \nplt.figure(figsize=(14, 8))\nsns.countplot(x=\"os\", data=train_sample)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:23.941772Z","iopub.execute_input":"2023-07-03T11:24:23.942454Z","iopub.status.idle":"2023-07-03T11:24:24.968658Z","shell.execute_reply.started":"2023-07-03T11:24:23.942399Z","shell.execute_reply":"2023-07-03T11:24:24.967591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now look at the distribution of the target  variable 'is_attributed'.","metadata":{}},{"cell_type":"code","source":"# target variable distribution\n100 * (train_sample['is_attributed'].astype('object').value_counts()/len(train_sample.index))","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:24.969824Z","iopub.execute_input":"2023-07-03T11:24:24.970092Z","iopub.status.idle":"2023-07-03T11:24:24.985064Z","shell.execute_reply.started":"2023-07-03T11:24:24.970070Z","shell.execute_reply":"2023-07-03T11:24:24.983924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only **about 0.2% of clicks are 'fraudulent'**, which is expected in a fraud detection problem. Such high class imbalance is probably going to be the toughest challenge of this problem.","metadata":{}},{"cell_type":"markdown","source":"### Exploring the Data - Segmented Univariate Analysis\n\nLet's now look at how the target variable varies with the various predictors.","metadata":{}},{"cell_type":"code","source":"# plot the average of 'is_attributed', or 'download rate'\n# with app (clearly this is non-readable)\n\napp_target = train_sample.groupby('app').is_attributed.agg(['mean','count'])\napp_target","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:24.986304Z","iopub.execute_input":"2023-07-03T11:24:24.986592Z","iopub.status.idle":"2023-07-03T11:24:25.008633Z","shell.execute_reply.started":"2023-07-03T11:24:24.986543Z","shell.execute_reply":"2023-07-03T11:24:25.007578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is clearly non-readable, so let's first get rid of all the apps that are very rare (say which comprise of less than 20% clicks) and plot the rest.","metadata":{}},{"cell_type":"code","source":"frequent_apps = train_sample.groupby('app').size().reset_index(name='count')\nfrequent_apps = frequent_apps[frequent_apps['count']>frequent_apps['count'].quantile(0.80)]\nfrequent_apps = frequent_apps.merge(train_sample,on='app',how='inner')\nfrequent_apps.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.011663Z","iopub.execute_input":"2023-07-03T11:24:25.011999Z","iopub.status.idle":"2023-07-03T11:24:25.058948Z","shell.execute_reply.started":"2023-07-03T11:24:25.011969Z","shell.execute_reply":"2023-07-03T11:24:25.057565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.countplot(y=\"app\", hue=\"is_attributed\", data=frequent_apps)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.060748Z","iopub.execute_input":"2023-07-03T11:24:25.061256Z","iopub.status.idle":"2023-07-03T11:24:25.767807Z","shell.execute_reply.started":"2023-07-03T11:24:25.061216Z","shell.execute_reply":"2023-07-03T11:24:25.766852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can do lots of other interesting ananlysis with the existing features. For now, let's create some new features which will probably improve the model.","metadata":{}},{"cell_type":"markdown","source":"## Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"Let's now derive some new features from the existing ones. There are a number of features one can extract from ```click_time``` itself, and by grouping combinations of IP with other features. ","metadata":{}},{"cell_type":"markdown","source":"### Datetime Based Features\n","metadata":{}},{"cell_type":"code","source":"# Creating datetime variables\n# takes in a df, adds date/time based columns to it, and returns the modified df\n\ndef timeFeatures(df):\n    # Derive new features using the click_time column\n    df['datetime'] = pd.to_datetime(df['click_time'])\n    df['day_of_week'] = df['datetime'].dt.dayofweek\n    df['day_of_year'] = df['datetime'].dt.dayofyear\n    df['month'] = df['datetime'].dt.month\n    df['hour'] =df['datetime'].dt.hour\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.771479Z","iopub.execute_input":"2023-07-03T11:24:25.772364Z","iopub.status.idle":"2023-07-03T11:24:25.777908Z","shell.execute_reply.started":"2023-07-03T11:24:25.772326Z","shell.execute_reply":"2023-07-03T11:24:25.776983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# creating new datetime variables and dropping the old ones\ntrain_sample = timeFeatures(train_sample)\ntrain_sample.drop(['click_time','datetime'], axis=1, inplace=True)\ntrain_sample.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.778992Z","iopub.execute_input":"2023-07-03T11:24:25.779327Z","iopub.status.idle":"2023-07-03T11:24:25.833337Z","shell.execute_reply.started":"2023-07-03T11:24:25.779272Z","shell.execute_reply":"2023-07-03T11:24:25.832485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# datatypes\n# note that by default the new datetime variables are int64\ntrain_sample.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.834295Z","iopub.execute_input":"2023-07-03T11:24:25.834561Z","iopub.status.idle":"2023-07-03T11:24:25.842287Z","shell.execute_reply.started":"2023-07-03T11:24:25.834535Z","shell.execute_reply":"2023-07-03T11:24:25.840854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# memory used by training data\nprint('Training dataset uses {0} MB'.format(train_sample.memory_usage().sum()/1024**2))","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.843925Z","iopub.execute_input":"2023-07-03T11:24:25.844525Z","iopub.status.idle":"2023-07-03T11:24:25.855013Z","shell.execute_reply.started":"2023-07-03T11:24:25.844496Z","shell.execute_reply":"2023-07-03T11:24:25.853998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# lets convert the variables back to lower dtype again\nint_vars = ['app', 'device', 'os', 'channel', 'day_of_week','day_of_year', 'month', 'hour']\ntrain_sample[int_vars] = train_sample[int_vars].astype('uint16')","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.856114Z","iopub.execute_input":"2023-07-03T11:24:25.857328Z","iopub.status.idle":"2023-07-03T11:24:25.876793Z","shell.execute_reply.started":"2023-07-03T11:24:25.857296Z","shell.execute_reply":"2023-07-03T11:24:25.875565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.877947Z","iopub.execute_input":"2023-07-03T11:24:25.878340Z","iopub.status.idle":"2023-07-03T11:24:25.885590Z","shell.execute_reply.started":"2023-07-03T11:24:25.878317Z","shell.execute_reply":"2023-07-03T11:24:25.884598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# space used by training data\nprint('Training dataset uses {0} MB'.format(train_sample.memory_usage().sum()/1024**2))","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.886774Z","iopub.execute_input":"2023-07-03T11:24:25.887465Z","iopub.status.idle":"2023-07-03T11:24:25.900677Z","shell.execute_reply.started":"2023-07-03T11:24:25.887439Z","shell.execute_reply":"2023-07-03T11:24:25.899707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### IP Grouping Based Features","metadata":{}},{"cell_type":"markdown","source":"Let's now create some important features by grouping IP addresses with features such as os, channel, hour, day etc. Also, count of each IP address will also be a feature.\n\nNote that though we are deriving new features by grouping IP addresses, using IP adress itself as a features is not a good idea. This is because (in the test data) if a new IP address is seen, the model will see a new 'category' and will not be able to make predictions (IP is a categorical variable, it has just been encoded with numbers).","metadata":{}},{"cell_type":"code","source":"# number of clicks by count of IP address\n# note that we are explicitly asking pandas to re-encode the aggregated features \n# as 'int16' to save memory\nip_count = train_sample.groupby('ip').size().reset_index(name='ip_count').astype('int16')\nip_count.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.901791Z","iopub.execute_input":"2023-07-03T11:24:25.902051Z","iopub.status.idle":"2023-07-03T11:24:25.926138Z","shell.execute_reply.started":"2023-07-03T11:24:25.902030Z","shell.execute_reply":"2023-07-03T11:24:25.925101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now merge this dataframe with the original training df. Similarly, we can create combinations of various features such as ip_day_hour (count of ip-day-hour combinations), ip_hour_channel, ip_hour_app, etc. \n\nThe following function takes in a dataframe and creates these features.","metadata":{}},{"cell_type":"code","source":"# creates groupings of IP addresses with other features and appends the new features to the df\n\ndef grouped_features(df):\n    # ip_count\n    ip_count = df.groupby('ip').size().reset_index(name='ip_count').astype('uint16')\n    ip_day_hour = df.groupby(['ip','day_of_week','hour']).size().reset_index(name='ip_day_hour').astype('uint16')\n    ip_hour_channel = df[['ip','hour','channel']].groupby(['ip','hour','channel']).size().reset_index(name='ip_hour_channel').astype('uint16')\n    ip_hour_os = df.groupby(['ip', 'hour', 'os']).channel.count().reset_index(name='ip_hour_os').astype('uint16')\n    ip_hour_app = df.groupby(['ip', 'hour', 'app']).channel.count().reset_index(name='ip_hour_app').astype('uint16')\n    ip_hour_device = df.groupby(['ip', 'hour', 'device']).channel.count().reset_index(name='ip_hour_device').astype('uint16')\n    \n    # merge the new aggregated features with the df\n    df = pd.merge(df, ip_count, on='ip', how='left')\n    del ip_count\n    df = pd.merge(df, ip_day_hour, on=['ip', 'day_of_week', 'hour'], how='left')\n    del ip_day_hour\n    df = pd.merge(df, ip_hour_channel, on=['ip', 'hour', 'channel'], how='left')\n    del ip_hour_channel\n    df = pd.merge(df, ip_hour_os, on=['ip', 'hour', 'os'], how='left')\n    del ip_hour_os\n    df = pd.merge(df, ip_hour_app, on=['ip', 'hour', 'app'], how='left')\n    del ip_hour_app\n    df = pd.merge(df, ip_hour_device, on=['ip', 'hour', 'device'], how='left')\n    del ip_hour_device\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.927814Z","iopub.execute_input":"2023-07-03T11:24:25.928187Z","iopub.status.idle":"2023-07-03T11:24:25.938877Z","shell.execute_reply.started":"2023-07-03T11:24:25.928156Z","shell.execute_reply":"2023-07-03T11:24:25.938159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample = grouped_features(train_sample)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:25.943836Z","iopub.execute_input":"2023-07-03T11:24:25.944326Z","iopub.status.idle":"2023-07-03T11:24:26.309238Z","shell.execute_reply.started":"2023-07-03T11:24:25.944296Z","shell.execute_reply":"2023-07-03T11:24:26.308036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:26.310383Z","iopub.execute_input":"2023-07-03T11:24:26.310746Z","iopub.status.idle":"2023-07-03T11:24:26.324374Z","shell.execute_reply.started":"2023-07-03T11:24:26.310715Z","shell.execute_reply":"2023-07-03T11:24:26.323340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Training dataset uses {0} MB'.format(train_sample.memory_usage().sum()/1024**2))","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:26.325640Z","iopub.execute_input":"2023-07-03T11:24:26.326413Z","iopub.status.idle":"2023-07-03T11:24:26.340729Z","shell.execute_reply.started":"2023-07-03T11:24:26.326386Z","shell.execute_reply":"2023-07-03T11:24:26.338490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# garbage collect (unused) object\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:26.342646Z","iopub.execute_input":"2023-07-03T11:24:26.343040Z","iopub.status.idle":"2023-07-03T11:24:26.524824Z","shell.execute_reply.started":"2023-07-03T11:24:26.343015Z","shell.execute_reply":"2023-07-03T11:24:26.524114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Modelling\n\nLet's now build models to predict the variable ```is_attributed``` (downloaded). We'll try the several variants of boosting (adaboost, gradient boosting and XGBoost), tune the hyperparameters in each model and choose the one which gives the best performance.\n\nIn the Kaggle competition, the metric for model evaluation is **area under the ROC curve**.\n","metadata":{}},{"cell_type":"code","source":"# create x and y train\nX = train_sample.drop('is_attributed',axis=1)\ny = train_sample[['is_attributed']]\n\n# split data into train and test/validation sets\nX_train, X_test, y_train, y_test = train_test_split(X,y,test_size=0.20,random_state=101)\nprint(X_train.shape)\nprint(y_train.shape)\nprint(X_test.shape)\nprint(y_test.shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:26.525804Z","iopub.execute_input":"2023-07-03T11:24:26.526323Z","iopub.status.idle":"2023-07-03T11:24:26.564742Z","shell.execute_reply.started":"2023-07-03T11:24:26.526298Z","shell.execute_reply":"2023-07-03T11:24:26.563960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check the average download rates in train and test data, should be comparable\nprint(y_train.mean())\nprint(y_test.mean())","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:26.565787Z","iopub.execute_input":"2023-07-03T11:24:26.566030Z","iopub.status.idle":"2023-07-03T11:24:26.573365Z","shell.execute_reply.started":"2023-07-03T11:24:26.566005Z","shell.execute_reply":"2023-07-03T11:24:26.571827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## AdaBoost","metadata":{}},{"cell_type":"code","source":"# adaboost classifier with max 600 decision tress of depth=2\n# learning_rate/shrinkage = 1.5\n\n# base estimator\ntree = DecisionTreeClassifier(max_depth=2)\n\n# adaboost with the tree as base estimator\nadaboost_model_1 = AdaBoostClassifier(base_estimator=tree,\n                                     n_estimators=600,\n                                     learning_rate=1.5,\n                                     algorithm=\"SAMME\")","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:26.575084Z","iopub.execute_input":"2023-07-03T11:24:26.575461Z","iopub.status.idle":"2023-07-03T11:24:26.589818Z","shell.execute_reply.started":"2023-07-03T11:24:26.575409Z","shell.execute_reply":"2023-07-03T11:24:26.588892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fit\nadaboost_model_1.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:24:26.590846Z","iopub.execute_input":"2023-07-03T11:24:26.591314Z","iopub.status.idle":"2023-07-03T11:25:28.882094Z","shell.execute_reply.started":"2023-07-03T11:24:26.591282Z","shell.execute_reply":"2023-07-03T11:25:28.881478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predictions\n# the second column represents the probability of click resulting in a download\n\npredictions = adaboost_model_1.predict_proba(X_test)\npredictions[:10]","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:25:28.883009Z","iopub.execute_input":"2023-07-03T11:25:28.883342Z","iopub.status.idle":"2023-07-03T11:25:29.328362Z","shell.execute_reply.started":"2023-07-03T11:25:28.883322Z","shell.execute_reply":"2023-07-03T11:25:29.326990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# metrics : AUC\nmetrics.roc_auc_score(y_test,predictions[:,1])","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:25:29.329508Z","iopub.execute_input":"2023-07-03T11:25:29.329774Z","iopub.status.idle":"2023-07-03T11:25:29.351344Z","shell.execute_reply.started":"2023-07-03T11:25:29.329753Z","shell.execute_reply":"2023-07-03T11:25:29.350354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### AdaBoost - Hyperparameter Tuning\n\nLet's now tune the hyperparameters of the AdaBoost classifier. In this case, we have two types of hyperparameters - those of the component trees (max_depth etc.) and those of the ensemble (n_estimators, learning_rate etc.). \n\n\nWe can tune both using the following technique - the keys of the form ```base_estimator_parameter_name``` belong to the trees (base estimator), and the rest belong to the ensemble.","metadata":{}},{"cell_type":"code","source":"# parameter grid\nparam_grid = {\"base_estimator__max_depth\" : [2,5],\n             \"n_estimators\" : [200,400,600]\n             }","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:25:29.352660Z","iopub.execute_input":"2023-07-03T11:25:29.352913Z","iopub.status.idle":"2023-07-03T11:25:29.357621Z","shell.execute_reply.started":"2023-07-03T11:25:29.352892Z","shell.execute_reply":"2023-07-03T11:25:29.356685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# base estimator\ntree = DecisionTreeClassifier()\n\n# adaboost with the tree as base estimator\n# learning rate is arbitrality set to 0.6 \n\nABC = AdaBoostClassifier(base_estimator=tree,\n                        learning_rate=0.6,\n                        algorithm=\"SAMME\")","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:25:29.359370Z","iopub.execute_input":"2023-07-03T11:25:29.359897Z","iopub.status.idle":"2023-07-03T11:25:29.370458Z","shell.execute_reply.started":"2023-07-03T11:25:29.359869Z","shell.execute_reply":"2023-07-03T11:25:29.369090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# run grid search\nfolds = 3\ngrid_search_ABC = GridSearchCV(ABC,\n                              cv=folds,\n                              param_grid=param_grid,\n                              scoring='roc_auc',\n                              return_train_score=True,\n                              verbose=1)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:25:29.371961Z","iopub.execute_input":"2023-07-03T11:25:29.372282Z","iopub.status.idle":"2023-07-03T11:25:29.383219Z","shell.execute_reply.started":"2023-07-03T11:25:29.372256Z","shell.execute_reply":"2023-07-03T11:25:29.382412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fit \ngrid_search_ABC.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:25:29.384227Z","iopub.execute_input":"2023-07-03T11:25:29.384733Z","iopub.status.idle":"2023-07-03T11:39:07.391519Z","shell.execute_reply.started":"2023-07-03T11:25:29.384706Z","shell.execute_reply":"2023-07-03T11:39:07.390300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# cv results\ncv_results = pd.DataFrame(grid_search_ABC.cv_results_)\ncv_results","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:07.392646Z","iopub.execute_input":"2023-07-03T11:39:07.392914Z","iopub.status.idle":"2023-07-03T11:39:07.415913Z","shell.execute_reply.started":"2023-07-03T11:39:07.392894Z","shell.execute_reply":"2023-07-03T11:39:07.414761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plotting AUC with hyperparameter combinations\n\nplt.figure(figsize=(16,6))\nfor n, depth in enumerate(param_grid['base_estimator__max_depth']):\n    \n\n    # subplot 1/n\n    plt.subplot(1,3, n+1)\n    depth_df = cv_results[cv_results['param_base_estimator__max_depth']==depth]\n\n    plt.plot(depth_df[\"param_n_estimators\"], depth_df[\"mean_test_score\"])\n    plt.plot(depth_df[\"param_n_estimators\"], depth_df[\"mean_train_score\"])\n    plt.xlabel('n_estimators')\n    plt.ylabel('AUC')\n    plt.title(\"max_depth={0}\".format(depth))\n    plt.ylim([0.60, 1])\n    plt.legend(['test score', 'train score'], loc='upper left')\n    plt.xscale('log')\n\n    \n","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:07.417264Z","iopub.execute_input":"2023-07-03T11:39:07.417663Z","iopub.status.idle":"2023-07-03T11:39:08.133673Z","shell.execute_reply.started":"2023-07-03T11:39:07.417632Z","shell.execute_reply":"2023-07-03T11:39:08.132543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The results above show that:\n- The ensemble with max_depth=5 is clearly overfitting (training auc is almost 1, while the test score is much lower)\n- At max_depth=2, the model performs slightly better (approx 95% AUC) with a higher test score \n\nThus, we should go ahead with ```max_depth=2``` and ```n_estimators=200```.\n\nNote that we haven't experimented with many other important hyperparameters till now, such as ```learning rate```, ```subsample``` etc., and the results might be considerably improved by tuning them. We'll next experiment with these hyperparameters.","metadata":{}},{"cell_type":"code","source":"# model performance on test data with chosen hyperparameters\n\n# base estimator\ntree = DecisionTreeClassifier(max_depth=2)\n\n# adaboost with the tree as base estimator\n# learning rate is arbitrarily set, we'll discuss learning_rate below\nABC = AdaBoostClassifier(\n    base_estimator=tree,\n    learning_rate=0.6,\n    n_estimators=200,\n    algorithm=\"SAMME\")\n\nABC.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:08.134829Z","iopub.execute_input":"2023-07-03T11:39:08.135158Z","iopub.status.idle":"2023-07-03T11:39:28.232711Z","shell.execute_reply.started":"2023-07-03T11:39:08.135128Z","shell.execute_reply":"2023-07-03T11:39:28.231562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict on test data\npredictions = ABC.predict_proba(X_test)\npredictions[:10]","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:28.234189Z","iopub.execute_input":"2023-07-03T11:39:28.234494Z","iopub.status.idle":"2023-07-03T11:39:28.385088Z","shell.execute_reply.started":"2023-07-03T11:39:28.234469Z","shell.execute_reply":"2023-07-03T11:39:28.383836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# roc auc\nmetrics.roc_auc_score(y_test, predictions[:, 1])","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:28.386451Z","iopub.execute_input":"2023-07-03T11:39:28.387104Z","iopub.status.idle":"2023-07-03T11:39:28.402895Z","shell.execute_reply.started":"2023-07-03T11:39:28.387075Z","shell.execute_reply":"2023-07-03T11:39:28.401548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Gradient Boosting Classifier\n\nLet's now try the gradient boosting classifier. We'll experiment with two main hyperparameters now - ```learning_rate``` (shrinkage) and ```subsample```. \n\nBy adjusting the learning rate to less than 1, we can regularize the model. A model with higher learning_rate learns fast, but is prone to overfitting; one with a lower learning rate learns slowly, but avoids overfitting.\n\nAlso, there's a trade-off between ```learning_rate``` and ```n_estimators``` - the higher the learning rate, the lesser trees the model needs (and thus we usually tune only one of them).\n\nAlso, by subsampling (setting ```subsample``` to less than 1), we can have the individual models built on random subsamples of size ```subsample```. That way, each tree will be trained on different subsets and reduce the model's variance.","metadata":{}},{"cell_type":"code","source":"# parameter grid\nparam_grid = {\"learning_rate\": [0.2, 0.6, 0.9],\n              \"subsample\": [0.3, 0.6, 0.9]\n             }","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:28.404633Z","iopub.execute_input":"2023-07-03T11:39:28.404945Z","iopub.status.idle":"2023-07-03T11:39:28.409796Z","shell.execute_reply.started":"2023-07-03T11:39:28.404919Z","shell.execute_reply":"2023-07-03T11:39:28.408432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# adaboost with the tree as base estimator\nGBC = GradientBoostingClassifier(max_depth=2, n_estimators=200)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:28.410993Z","iopub.execute_input":"2023-07-03T11:39:28.411234Z","iopub.status.idle":"2023-07-03T11:39:28.422886Z","shell.execute_reply.started":"2023-07-03T11:39:28.411214Z","shell.execute_reply":"2023-07-03T11:39:28.422102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# run grid search\nfolds = 3\ngrid_search_GBC = GridSearchCV(GBC, \n                               cv = folds,\n                               param_grid=param_grid, \n                               scoring = 'roc_auc', \n                               return_train_score=True,                         \n                               verbose = 1)\n\ngrid_search_GBC.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:39:28.423868Z","iopub.execute_input":"2023-07-03T11:39:28.424105Z","iopub.status.idle":"2023-07-03T11:42:45.737184Z","shell.execute_reply.started":"2023-07-03T11:39:28.424084Z","shell.execute_reply":"2023-07-03T11:42:45.735846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv_results = pd.DataFrame(grid_search_GBC.cv_results_)\ncv_results.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:42:45.740312Z","iopub.execute_input":"2023-07-03T11:42:45.740632Z","iopub.status.idle":"2023-07-03T11:42:45.763892Z","shell.execute_reply.started":"2023-07-03T11:42:45.740605Z","shell.execute_reply":"2023-07-03T11:42:45.762529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # plotting\nplt.figure(figsize=(16,6))\n\n\nfor n, subsample in enumerate(param_grid['subsample']):\n    \n\n    # subplot 1/n\n    plt.subplot(1,len(param_grid['subsample']), n+1)\n    df = cv_results[cv_results['param_subsample']==subsample]\n\n    plt.plot(df[\"param_learning_rate\"], df[\"mean_test_score\"])\n    plt.plot(df[\"param_learning_rate\"], df[\"mean_train_score\"])\n    plt.xlabel('learning_rate')\n    plt.ylabel('AUC')\n    plt.title(\"subsample={0}\".format(subsample))\n    plt.ylim([0.60, 1])\n    plt.legend(['test score', 'train score'], loc='upper left')\n    plt.xscale('log')\n","metadata":{"execution":{"iopub.status.busy":"2023-07-03T11:42:45.765119Z","iopub.execute_input":"2023-07-03T11:42:45.765395Z","iopub.status.idle":"2023-07-03T11:42:46.736350Z","shell.execute_reply.started":"2023-07-03T11:42:45.765370Z","shell.execute_reply":"2023-07-03T11:42:46.735646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is clear from the plot above that the model with a lower subsample ratio performs better, while those with higher subsamples tend to overfit. \n\nAlso, a lower learning rate results in less overfitting.","metadata":{}},{"cell_type":"markdown","source":"### XGBoost\n\nLet's finally try XGBoost. The hyperparameters are the same, some important ones being ```subsample```, ```learning_rate```, ```max_depth``` etc.\n","metadata":{}},{"cell_type":"code","source":"# fit model on training data with default hyperparameters\nmodel = XGBClassifier()\nmodel.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:15:52.624459Z","iopub.execute_input":"2023-07-03T12:15:52.624833Z","iopub.status.idle":"2023-07-03T12:15:56.285782Z","shell.execute_reply.started":"2023-07-03T12:15:52.624801Z","shell.execute_reply":"2023-07-03T12:15:56.284945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# make predictions for test data\n# use predict_proba since we need probabilities to compute auc\ny_pred = model.predict_proba(X_test)\ny_pred[:10]","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:17:06.703992Z","iopub.execute_input":"2023-07-03T12:17:06.704357Z","iopub.status.idle":"2023-07-03T12:17:06.735360Z","shell.execute_reply.started":"2023-07-03T12:17:06.704327Z","shell.execute_reply":"2023-07-03T12:17:06.734345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# evaluate predictions\nroc = metrics.roc_auc_score(y_test,y_pred[:,1])\nprint(\"AUC : %.2f%%\" %(roc * 100.0))","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:17:07.268781Z","iopub.execute_input":"2023-07-03T12:17:07.269612Z","iopub.status.idle":"2023-07-03T12:17:07.283956Z","shell.execute_reply.started":"2023-07-03T12:17:07.269579Z","shell.execute_reply":"2023-07-03T12:17:07.282843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The roc_auc in this case is about 0.95% with default hyperparameters. Let's try changing the hyperparameters","metadata":{}},{"cell_type":"markdown","source":"Let's now try tuning the hyperparameters using k-fold CV. We'll then use grid search CV to find the optimal values of hyperparameters.","metadata":{}},{"cell_type":"code","source":"# hyperparameter tuning with XGBoost\n# creating a KFold object\nfolds = 3\n# specify range of hyperparamaters\nparam_grid = {'learning_rate' : [0.2,0.6],\n             'subsample' : [0.3,0.6,0.9]\n             }\n# specify model\nxgb_model = XGBClassifier(max_depth=2,n_estimators=200)\n# set up GridSearchCV()\nmodel_cv = GridSearchCV(estimator = xgb_model,\n                       param_grid = param_grid,\n                       scoring = 'roc_auc',\n                       cv = folds,\n                       verbose = 1,\n                       return_train_score = True)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:17:09.354453Z","iopub.execute_input":"2023-07-03T12:17:09.355027Z","iopub.status.idle":"2023-07-03T12:17:09.361578Z","shell.execute_reply.started":"2023-07-03T12:17:09.354997Z","shell.execute_reply":"2023-07-03T12:17:09.359707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fit the model\nmodel_cv.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:17:09.704188Z","iopub.execute_input":"2023-07-03T12:17:09.704630Z","iopub.status.idle":"2023-07-03T12:17:53.341904Z","shell.execute_reply.started":"2023-07-03T12:17:09.704594Z","shell.execute_reply":"2023-07-03T12:17:53.340447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# cv results\ncv_results = pd.DataFrame(model_cv.cv_results_)\ncv_results","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:17:53.345406Z","iopub.execute_input":"2023-07-03T12:17:53.345715Z","iopub.status.idle":"2023-07-03T12:17:53.370343Z","shell.execute_reply.started":"2023-07-03T12:17:53.345691Z","shell.execute_reply":"2023-07-03T12:17:53.369078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv_results.columns","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:19:19.044043Z","iopub.execute_input":"2023-07-03T12:19:19.044442Z","iopub.status.idle":"2023-07-03T12:19:19.052621Z","shell.execute_reply.started":"2023-07-03T12:19:19.044394Z","shell.execute_reply":"2023-07-03T12:19:19.051187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convert parameters to int for plotting on x-axis\ncv_results['param_learning_rate'] = cv_results['param_learning_rate'].astype('float')\ncv_results.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:00:12.480415Z","iopub.execute_input":"2023-07-03T13:00:12.480803Z","iopub.status.idle":"2023-07-03T13:00:12.506025Z","shell.execute_reply.started":"2023-07-03T13:00:12.480777Z","shell.execute_reply":"2023-07-03T13:00:12.504733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # plotting\nplt.figure(figsize=(16,6))\n\nparam_grid = {'learning_rate': [0.2, 0.6], \n             'subsample': [0.3, 0.6, 0.9]} \n\n\nfor n, subsample in enumerate(param_grid['subsample']):\n    \n\n    # subplot 1/n\n    plt.subplot(1,len(param_grid['subsample']), n+1)\n    df = cv_results[cv_results['param_subsample']==subsample]\n\n    plt.plot(df[\"param_learning_rate\"], df[\"mean_test_score\"])\n    plt.plot(df[\"param_learning_rate\"], df[\"mean_train_score\"])\n    plt.xlabel('learning_rate')\n    plt.ylabel('AUC')\n    plt.title(\"subsample={0}\".format(subsample))\n    plt.ylim([0.60, 1])\n    plt.legend(['test score', 'train score'], loc='upper left')\n    plt.xscale('log')","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:00:17.877578Z","iopub.execute_input":"2023-07-03T13:00:17.877936Z","iopub.status.idle":"2023-07-03T13:00:19.312920Z","shell.execute_reply.started":"2023-07-03T13:00:17.877908Z","shell.execute_reply":"2023-07-03T13:00:19.311775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The results show that a subsample size of 0.6 and learning_rate of about 0.2 seems optimal. \nAlso, XGBoost has resulted in the highest ROC AUC obtained (across various hyperparameters). \n\n\nLet's build a final model with the chosen hyperparameters.","metadata":{}},{"cell_type":"code","source":"# chosen hyperparameters\n# 'objective':'binary:logistic' outputs probability rather than label, which we need for auc\nparams = {'learning_rate': 0.2,\n          'max_depth': 2, \n          'n_estimators':200,\n          'subsample':0.6,\n         'objective':'binary:logistic'}\n\n# fit model on training data\nmodel = XGBClassifier(params = params)\nmodel.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:00:40.961337Z","iopub.execute_input":"2023-07-03T13:00:40.962208Z","iopub.status.idle":"2023-07-03T13:00:44.658357Z","shell.execute_reply.started":"2023-07-03T13:00:40.962175Z","shell.execute_reply":"2023-07-03T13:00:44.657664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict\ny_pred = model.predict_proba(X_test)\ny_pred[:10]","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:01:10.536050Z","iopub.execute_input":"2023-07-03T13:01:10.536635Z","iopub.status.idle":"2023-07-03T13:01:10.565112Z","shell.execute_reply.started":"2023-07-03T13:01:10.536603Z","shell.execute_reply":"2023-07-03T13:01:10.564387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first column in y_pred is the P(0), i.e. P(not fraud), and the second column is P(1/fraud).","metadata":{}},{"cell_type":"code","source":"# roc_auc\nauc = sklearn.metrics.roc_auc_score(y_test,y_pred[:,1])\nauc","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:02:48.997879Z","iopub.execute_input":"2023-07-03T13:02:48.998242Z","iopub.status.idle":"2023-07-03T13:02:49.016644Z","shell.execute_reply.started":"2023-07-03T13:02:48.998216Z","shell.execute_reply":"2023-07-03T13:02:49.015313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, let's also look at the feature importances.","metadata":{}},{"cell_type":"code","source":"# feature importance\nimportance = dict(zip(X_train.columns, model.feature_importances_))\nimportance","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:03:56.876479Z","iopub.execute_input":"2023-07-03T13:03:56.877454Z","iopub.status.idle":"2023-07-03T13:03:56.885291Z","shell.execute_reply.started":"2023-07-03T13:03:56.877409Z","shell.execute_reply":"2023-07-03T13:03:56.883317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot\nplt.bar(range(len(model.feature_importances_)), model.feature_importances_)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:04:10.477406Z","iopub.execute_input":"2023-07-03T13:04:10.478083Z","iopub.status.idle":"2023-07-03T13:04:10.649306Z","shell.execute_reply.started":"2023-07-03T13:04:10.478051Z","shell.execute_reply":"2023-07-03T13:04:10.648472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read submission file\nsample_sub = pd.read_csv('/kaggle/input/talkingdata-adtracking-fraud-detection/sample_submission.csv')\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:30:26.726803Z","iopub.execute_input":"2023-07-03T13:30:26.727174Z","iopub.status.idle":"2023-07-03T13:30:28.608534Z","shell.execute_reply.started":"2023-07-03T13:30:26.727145Z","shell.execute_reply":"2023-07-03T13:30:28.607640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict probability of test data\ntest_final = pd.read_csv('/kaggle/input/talkingdata-adtracking-fraud-detection/test.csv')\ntest_final.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:30:30.436226Z","iopub.execute_input":"2023-07-03T13:30:30.436804Z","iopub.status.idle":"2023-07-03T13:30:38.532734Z","shell.execute_reply.started":"2023-07-03T13:30:30.436768Z","shell.execute_reply":"2023-07-03T13:30:38.531767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predictions on test data\ntest_final = timeFeatures(test_final)\ntest_final.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:30:39.945599Z","iopub.execute_input":"2023-07-03T13:30:39.946704Z","iopub.status.idle":"2023-07-03T13:30:44.254328Z","shell.execute_reply.started":"2023-07-03T13:30:39.946648Z","shell.execute_reply":"2023-07-03T13:30:44.252194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final.drop(['click_time', 'datetime'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:30:51.371271Z","iopub.execute_input":"2023-07-03T13:30:51.375106Z","iopub.status.idle":"2023-07-03T13:30:51.767802Z","shell.execute_reply.started":"2023-07-03T13:30:51.375009Z","shell.execute_reply":"2023-07-03T13:30:51.767073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:30:54.831064Z","iopub.execute_input":"2023-07-03T13:30:54.832486Z","iopub.status.idle":"2023-07-03T13:30:54.845608Z","shell.execute_reply.started":"2023-07-03T13:30:54.832413Z","shell.execute_reply":"2023-07-03T13:30:54.844019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test_final[categorical_cols]=test_final[categorical_cols].apply(lambda x: le.fit_transform(x))","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:26:15.010444Z","iopub.execute_input":"2023-07-03T13:26:15.010771Z","iopub.status.idle":"2023-07-03T13:26:15.015647Z","shell.execute_reply.started":"2023-07-03T13:26:15.010746Z","shell.execute_reply":"2023-07-03T13:26:15.014219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final.info()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:30:59.585916Z","iopub.execute_input":"2023-07-03T13:30:59.586332Z","iopub.status.idle":"2023-07-03T13:30:59.601027Z","shell.execute_reply.started":"2023-07-03T13:30:59.586300Z","shell.execute_reply":"2023-07-03T13:30:59.599299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# number of clicks by IP\nip_count = test_final.groupby('ip')['channel'].count().reset_index()\nip_count.columns = ['ip', 'count_by_ip']\nip_count.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:31:11.231272Z","iopub.execute_input":"2023-07-03T13:31:11.231693Z","iopub.status.idle":"2023-07-03T13:31:11.662803Z","shell.execute_reply.started":"2023-07-03T13:31:11.231663Z","shell.execute_reply":"2023-07-03T13:31:11.661225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merge this with the training data\ntest_final = pd.merge(test_final, ip_count, on='ip', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:31:15.706754Z","iopub.execute_input":"2023-07-03T13:31:15.707362Z","iopub.status.idle":"2023-07-03T13:31:18.652909Z","shell.execute_reply.started":"2023-07-03T13:31:15.707293Z","shell.execute_reply":"2023-07-03T13:31:18.651877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del ip_count","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:31:20.866278Z","iopub.execute_input":"2023-07-03T13:31:20.866640Z","iopub.status.idle":"2023-07-03T13:31:20.872054Z","shell.execute_reply.started":"2023-07-03T13:31:20.866615Z","shell.execute_reply":"2023-07-03T13:31:20.870861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final.info()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:31:22.386290Z","iopub.execute_input":"2023-07-03T13:31:22.387894Z","iopub.status.idle":"2023-07-03T13:31:22.399172Z","shell.execute_reply.started":"2023-07-03T13:31:22.387843Z","shell.execute_reply":"2023-07-03T13:31:22.397663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict on test data\n# y_pred_test = model.predict_proba(test_final.drop('click_id', axis=1))\n# y_pred_test[:9]","metadata":{"execution":{"iopub.status.busy":"2023-07-03T13:32:26.261454Z","iopub.execute_input":"2023-07-03T13:32:26.261902Z","iopub.status.idle":"2023-07-03T13:32:26.267526Z","shell.execute_reply.started":"2023-07-03T13:32:26.261832Z","shell.execute_reply":"2023-07-03T13:32:26.266603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}