{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TalkingData: Fraudulent Click Prediction\n\n\n\n\n","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a"}},{"cell_type":"markdown","source":"## Understanding and Exploring the Data\n.\n\nThe detailed data dictionary is mentioned here:\n- ```ip```: ip address of click.\n- ```app```: app id for marketing.\n- ```device```: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n- ```os```: os version id of user mobile phone\n- ```channel```: channel id of mobile ad publisher\n- ```click_time```: timestamp of click (UTC)\n- ```attributed_time```: if user download the app for after clicking an ad, this is the time of the app download\n- ```is_attributed```: the target that is to be predicted, indicating the app was downloaded\n\nLet's try finding some useful trends in the data.","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport sklearn\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import KFold\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom sklearn import metrics\n\nfrom xgboost import plot_importance\n\n%matplotlib inline\n\nimport os\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_cell_guid":"5ba86a36-b8be-44db-ad92-e3cddbb354d1","_uuid":"cd4d66f4ecc6daa0ea4cab6359fc5203112f245c","execution":{"iopub.status.busy":"2021-10-16T13:36:08.328054Z","iopub.execute_input":"2021-10-16T13:36:08.328409Z","iopub.status.idle":"2021-10-16T13:36:09.595879Z","shell.execute_reply.started":"2021-10-16T13:36:08.328307Z","shell.execute_reply":"2021-10-16T13:36:09.595182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Reading the Data  \n\nThe code below reads the train_sample.csv file if you set testing = True, else reads the full train.csv file. You can read the sample while tuning the model etc., and then run the model on the full data once done.\n\n#### Important Note: Save memory when the data is huge\n\nSince the training data is quite huge, the program will be quite slow if you don't consciously follow some best practices to save memory. This notebook demonstrates some of those practices. ","metadata":{"_cell_guid":"d408baf0-f640-45f3-8dec-141746a985a8","_uuid":"7a3dc1d7ebed1cdc9818338682bf5858e276681e"}},{"cell_type":"code","source":"# reading training data\n\n# specify column dtypes to save memory\ndtypes = {\n        'ip'            : 'uint16',\n        'app'           : 'uint16',\n        'device'        : 'uint16',\n        'os'            : 'uint16',\n        'channel'       : 'uint16',\n        'is_attributed' : 'uint8',\n        'click_id'      : 'uint32' \n        }\n\ntrain_path = \"../input/talkingdata-adtracking-fraud-detection/train_sample.csv\"\nskiprows = None\nnrows = None\ncolnames=['ip','app','device','os', 'channel', 'click_time', 'is_attributed']\n# read training data\ntrain_sample = pd.read_csv(train_path, skiprows=skiprows, nrows=nrows, dtype=dtypes, usecols=colnames)\n","metadata":{"_cell_guid":"54ed4833-efb3-4adf-971b-28a56ae19d01","_uuid":"bac7dea7334ff1fb925347f9302b3d06e189fbc7","execution":{"iopub.status.busy":"2021-10-16T13:46:04.248247Z","iopub.execute_input":"2021-10-16T13:46:04.248630Z","iopub.status.idle":"2021-10-16T13:46:04.393780Z","shell.execute_reply.started":"2021-10-16T13:46:04.248593Z","shell.execute_reply":"2021-10-16T13:46:04.392894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# length of training data\nlen(train_sample.index)","metadata":{"_cell_guid":"98d1b00a-8ee6-4959-9091-e8b56e4bcb15","_uuid":"a9f5559317e6cc476c47b166bd95118c9f29dbbe","execution":{"iopub.status.busy":"2021-10-16T13:36:10.806167Z","iopub.execute_input":"2021-10-16T13:36:10.806957Z","iopub.status.idle":"2021-10-16T13:36:10.813870Z","shell.execute_reply.started":"2021-10-16T13:36:10.806922Z","shell.execute_reply":"2021-10-16T13:36:10.813079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# training data top rows\ntrain_sample.head()","metadata":{"_cell_guid":"85dce024-c56b-412d-b230-7a42d1e61661","_uuid":"831d573b61c925fed93c55d4a4094e2f374e727b","scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:37:03.307139Z","iopub.execute_input":"2021-10-16T13:37:03.308200Z","iopub.status.idle":"2021-10-16T13:37:03.323784Z","shell.execute_reply.started":"2021-10-16T13:37:03.308156Z","shell.execute_reply":"2021-10-16T13:37:03.323022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Exploring the Data - Univariate Analysis\n","metadata":{"_cell_guid":"047cdedb-b9a4-4185-9079-31d19e418dd5","_uuid":"bdca59dbbab3ea6e7938ed6348bda8b022b54437"}},{"cell_type":"markdown","source":"Let's now understand and explore the data. Let's start with understanding the size and data types of the train_sample data.","metadata":{"_cell_guid":"f1924177-9395-4f17-bbec-6746cf8c360e","_uuid":"778a6efbfc8c7c9d82f6aa6c84af3e145777be13"}},{"cell_type":"code","source":"# look at non-null values, number of entries etc.\n# there are no missing values\ntrain_sample.info()","metadata":{"_cell_guid":"4539d9f9-fe8e-438b-9c6e-8d0f4d089e17","_uuid":"0aff4ad2d4abb7a49cfeccbd0d65f96a5256eab8","execution":{"iopub.status.busy":"2021-10-16T13:37:11.206883Z","iopub.execute_input":"2021-10-16T13:37:11.207185Z","iopub.status.idle":"2021-10-16T13:37:11.242267Z","shell.execute_reply.started":"2021-10-16T13:37:11.207158Z","shell.execute_reply":"2021-10-16T13:37:11.241399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Basic exploratory analysis \n\n# Number of unique values in each column\ndef fraction_unique(x):\n    return len(train_sample[x].unique())\n\nnumber_unique_vals = {x: fraction_unique(x) for x in train_sample.columns}\nnumber_unique_vals","metadata":{"_cell_guid":"dd930aa6-557f-4b59-9a36-8f940324f594","_uuid":"816535e82869da6601f9dd7227a201adda9843d1","execution":{"iopub.status.busy":"2021-10-16T13:37:15.776360Z","iopub.execute_input":"2021-10-16T13:37:15.776712Z","iopub.status.idle":"2021-10-16T13:37:15.814747Z","shell.execute_reply.started":"2021-10-16T13:37:15.776674Z","shell.execute_reply":"2021-10-16T13:37:15.813891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# All columns apart from click time are originally int type, \n# though note that they are all actually categorical \ntrain_sample.dtypes","metadata":{"_cell_guid":"408e795a-2d0d-4ae0-bf96-868721d9394d","_uuid":"c6bd82f7f439f73bf543d2313c82e4d3e99519e4","execution":{"iopub.status.busy":"2021-10-16T13:37:16.356201Z","iopub.execute_input":"2021-10-16T13:37:16.356474Z","iopub.status.idle":"2021-10-16T13:37:16.364391Z","shell.execute_reply.started":"2021-10-16T13:37:16.356447Z","shell.execute_reply":"2021-10-16T13:37:16.363368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # distribution of 'app' \n# # some 'apps' have a disproportionately high number of clicks (>15k), and some are very rare (3-4)\nplt.figure(figsize=(14, 8))\nsns.countplot(x=\"app\", data=train_sample)","metadata":{"_cell_guid":"3e9c25d8-31a7-4b43-b2ac-5ecc21921f26","_uuid":"b99fbb4955c7f05d78495d958d16d673479aca64","execution":{"iopub.status.busy":"2021-10-16T13:37:29.306077Z","iopub.execute_input":"2021-10-16T13:37:29.306714Z","iopub.status.idle":"2021-10-16T13:37:31.713301Z","shell.execute_reply.started":"2021-10-16T13:37:29.306672Z","shell.execute_reply":"2021-10-16T13:37:31.712324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # distribution of 'device' \n# # this is expected because a few popular devices are used heavily\nplt.figure(figsize=(14, 8))\nsns.countplot(x=\"device\", data=train_sample)","metadata":{"_cell_guid":"77e73dce-cf2f-43a9-b0a8-fe5956ff99a9","_uuid":"d3815fab7e31db39452c9fca0b3ade757124da25","execution":{"iopub.status.busy":"2021-10-16T13:37:37.666596Z","iopub.execute_input":"2021-10-16T13:37:37.666924Z","iopub.status.idle":"2021-10-16T13:37:39.192202Z","shell.execute_reply.started":"2021-10-16T13:37:37.666896Z","shell.execute_reply":"2021-10-16T13:37:39.191369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # channel: various channels get clicks in comparable quantities\nplt.figure(figsize=(14, 8))\nsns.countplot(x=\"channel\", data=train_sample)","metadata":{"_cell_guid":"098c5d37-5302-45e8-9845-eb3dfa7be87b","_uuid":"656c6cfa30d189febea9bb4c12b2f2519c0b0cef","execution":{"iopub.status.busy":"2021-10-16T13:37:39.193399Z","iopub.execute_input":"2021-10-16T13:37:39.194155Z","iopub.status.idle":"2021-10-16T13:37:41.759850Z","shell.execute_reply.started":"2021-10-16T13:37:39.194109Z","shell.execute_reply":"2021-10-16T13:37:41.758803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # os: there are a couple commos OSes (android and ios?), though some are rare and can indicate suspicion \nplt.figure(figsize=(14, 8))\nsns.countplot(x=\"os\", data=train_sample)","metadata":{"_cell_guid":"c39356d4-3f97-41be-944f-ce88db264365","_uuid":"8f33610186e195da39033b9a64cd2d11b244113a","execution":{"iopub.status.busy":"2021-10-16T13:37:41.762690Z","iopub.execute_input":"2021-10-16T13:37:41.762953Z","iopub.status.idle":"2021-10-16T13:37:43.651518Z","shell.execute_reply.started":"2021-10-16T13:37:41.762926Z","shell.execute_reply":"2021-10-16T13:37:43.650732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now look at the distribution of the target  variable 'is_attributed'.","metadata":{"_cell_guid":"a52d4651-f5ca-45e0-8b20-ec1f70d80edb","_uuid":"6604d330cee83328b4e28790829c82b40fe83ef1"}},{"cell_type":"code","source":"# # target variable distribution\n100*(train_sample['is_attributed'].astype('object').value_counts()/len(train_sample.index))","metadata":{"_cell_guid":"332d0c92-31f6-4109-91ba-5d1f38725ca8","_uuid":"755b0ff3de0028805b6e33fe5b6fb70e5a1247d5","scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:37:49.546663Z","iopub.execute_input":"2021-10-16T13:37:49.546977Z","iopub.status.idle":"2021-10-16T13:37:49.572722Z","shell.execute_reply.started":"2021-10-16T13:37:49.546948Z","shell.execute_reply":"2021-10-16T13:37:49.571965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only **about 0.2% of clicks are 'fraudulent'**, which is expected in a fraud detection problem. Such high class imbalance is probably going to be the toughest challenge of this problem.","metadata":{"_cell_guid":"030933bd-d06b-42c8-8d1e-882f70776a74","_uuid":"b604a87e4ca95bbcbe40e3192f7b66774c03d84d"}},{"cell_type":"code","source":"# plot the average of 'is_attributed', or 'download rate'\n# with app (clearly this is non-readable)\napp_target = train_sample.groupby('app').is_attributed.agg(['mean', 'count'])\napp_target","metadata":{"_cell_guid":"73089123-995e-4142-8aa9-784bbcddccb2","_uuid":"1a5d777dc7d2a003752eefd8af94c4433099bf44","scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:38:11.797499Z","iopub.execute_input":"2021-10-16T13:38:11.797804Z","iopub.status.idle":"2021-10-16T13:38:11.818999Z","shell.execute_reply.started":"2021-10-16T13:38:11.797752Z","shell.execute_reply":"2021-10-16T13:38:11.818186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is clearly non-readable, so let's first get rid of all the apps that are very rare (say which comprise of less than 20% clicks) and plot the rest.","metadata":{}},{"cell_type":"code","source":"frequent_apps = train_sample.groupby('app').size().reset_index(name='count')\nfrequent_apps = frequent_apps[frequent_apps['count']>frequent_apps['count'].quantile(0.80)]\nfrequent_apps = frequent_apps.merge(train_sample, on='app', how='inner')\nfrequent_apps.head()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:38:15.656269Z","iopub.execute_input":"2021-10-16T13:38:15.656587Z","iopub.status.idle":"2021-10-16T13:38:15.709157Z","shell.execute_reply.started":"2021-10-16T13:38:15.656553Z","shell.execute_reply":"2021-10-16T13:38:15.707637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.countplot(y=\"app\", hue=\"is_attributed\", data=frequent_apps);","metadata":{"execution":{"iopub.status.busy":"2021-10-16T13:38:18.756834Z","iopub.execute_input":"2021-10-16T13:38:18.757216Z","iopub.status.idle":"2021-10-16T13:38:19.421297Z","shell.execute_reply.started":"2021-10-16T13:38:18.757173Z","shell.execute_reply":"2021-10-16T13:38:19.420452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Engineering","metadata":{"_cell_guid":"40f6f8d9-bd39-4876-bca8-5dcaa14eed04","_uuid":"b76d4314871b8db705dc26b51f4ba07df2cd857b"}},{"cell_type":"code","source":"# Creating datetime variables\n# takes in a df, adds date/time based columns to it, and returns the modified df\ndef timeFeatures(df):\n    # Derive new features using the click_time column\n    df['datetime'] = pd.to_datetime(df['click_time'])\n    df['day_of_week'] = df['datetime'].dt.dayofweek\n    df[\"day_of_year\"] = df[\"datetime\"].dt.dayofyear\n    df[\"month\"] = df[\"datetime\"].dt.month\n    df[\"hour\"] = df[\"datetime\"].dt.hour\n    return df","metadata":{"_cell_guid":"fd611758-c8ec-4401-9875-1eb760ccddb7","_uuid":"a38688f6b07581964699b4a8c41100844f21090c","execution":{"iopub.status.busy":"2021-10-16T13:38:40.376711Z","iopub.execute_input":"2021-10-16T13:38:40.377530Z","iopub.status.idle":"2021-10-16T13:38:40.386140Z","shell.execute_reply.started":"2021-10-16T13:38:40.377480Z","shell.execute_reply":"2021-10-16T13:38:40.385253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# creating new datetime variables and dropping the old ones\ntrain_sample = timeFeatures(train_sample)\ntrain_sample.drop(['click_time', 'datetime'], axis=1, inplace=True)\ntrain_sample.head()","metadata":{"_cell_guid":"237df241-844b-4003-b906-298432b8d868","_uuid":"694177530d20a1e33801d6053a88d99f28cfd88c","scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:38:41.827384Z","iopub.execute_input":"2021-10-16T13:38:41.827669Z","iopub.status.idle":"2021-10-16T13:38:41.932797Z","shell.execute_reply.started":"2021-10-16T13:38:41.827642Z","shell.execute_reply":"2021-10-16T13:38:41.931970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# datatypes\n# note that by default the new datetime variables are int64\ntrain_sample.dtypes","metadata":{"_cell_guid":"4cfe36bc-b218-48ce-ac97-9b9da3758d2f","_uuid":"0eead49461cc6e85bf733b2682c95b098c6b4762","execution":{"iopub.status.busy":"2021-10-16T13:38:49.776939Z","iopub.execute_input":"2021-10-16T13:38:49.777257Z","iopub.status.idle":"2021-10-16T13:38:49.785360Z","shell.execute_reply.started":"2021-10-16T13:38:49.777223Z","shell.execute_reply":"2021-10-16T13:38:49.784499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# lets convert the variables back to lower dtype again\nint_vars = ['app', 'device', 'os', 'channel', 'day_of_week','day_of_year', 'month', 'hour']\ntrain_sample[int_vars] = train_sample[int_vars].astype('uint16')","metadata":{"_cell_guid":"630048e3-b363-4eb7-8620-fc7c0aed8794","_uuid":"f358d55b47fdd1f3458fc640a37b83f93abd01d9","execution":{"iopub.status.busy":"2021-10-16T13:38:55.776832Z","iopub.execute_input":"2021-10-16T13:38:55.777117Z","iopub.status.idle":"2021-10-16T13:38:55.793569Z","shell.execute_reply.started":"2021-10-16T13:38:55.777089Z","shell.execute_reply":"2021-10-16T13:38:55.792538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample.dtypes","metadata":{"_cell_guid":"bfb0e108-be87-4f1f-8690-b6a0dd253d5f","_uuid":"2ad5b5082c465d89a3749ef500121a0068fc7abe","scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:38:57.586037Z","iopub.execute_input":"2021-10-16T13:38:57.586339Z","iopub.status.idle":"2021-10-16T13:38:57.594816Z","shell.execute_reply.started":"2021-10-16T13:38:57.586307Z","shell.execute_reply":"2021-10-16T13:38:57.593828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### IP Grouping Based Features","metadata":{"_cell_guid":"f1a41e27-1c82-42a1-b06f-2b2d384cecee","_uuid":"1a82f07ac45c9a68893d104349971344ef7aa32b"}},{"cell_type":"code","source":"# number of clicks by count of IP address\n# note that we are explicitly asking pandas to re-encode the aggregated features \n# as 'int16' to save memory\nip_count = train_sample.groupby('ip').size().reset_index(name='ip_count').astype('int16')\nip_count.head()","metadata":{"_cell_guid":"b0cd1ca3-e3ef-414b-b9cc-add42aebf813","_uuid":"59690cdbb2b67074d81b9be15375e6ffb0b8f85e","execution":{"iopub.status.busy":"2021-10-16T13:39:10.126900Z","iopub.execute_input":"2021-10-16T13:39:10.127179Z","iopub.status.idle":"2021-10-16T13:39:10.147946Z","shell.execute_reply.started":"2021-10-16T13:39:10.127152Z","shell.execute_reply":"2021-10-16T13:39:10.147338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now merge this dataframe with the original training df. Similarly, we can create combinations of various features such as ip_day_hour (count of ip-day-hour combinations), ip_hour_channel, ip_hour_app, etc. \n\nThe following function takes in a dataframe and creates these features.","metadata":{}},{"cell_type":"code","source":"# creates groupings of IP addresses with other features and appends the new features to the df\ndef grouped_features(df):\n    # ip_count\n    ip_count = df.groupby('ip').size().reset_index(name='ip_count').astype('uint16')\n    ip_day_hour = df.groupby(['ip', 'day_of_week', 'hour']).size().reset_index(name='ip_day_hour').astype('uint16')\n    ip_hour_channel = df[['ip', 'hour', 'channel']].groupby(['ip', 'hour', 'channel']).size().reset_index(name='ip_hour_channel').astype('uint16')\n    ip_hour_os = df.groupby(['ip', 'hour', 'os']).channel.count().reset_index(name='ip_hour_os').astype('uint16')\n    ip_hour_app = df.groupby(['ip', 'hour', 'app']).channel.count().reset_index(name='ip_hour_app').astype('uint16')\n    ip_hour_device = df.groupby(['ip', 'hour', 'device']).channel.count().reset_index(name='ip_hour_device').astype('uint16')\n    \n    # merge the new aggregated features with the df\n    df = pd.merge(df, ip_count, on='ip', how='left')\n    del ip_count\n    df = pd.merge(df, ip_day_hour, on=['ip', 'day_of_week', 'hour'], how='left')\n    del ip_day_hour\n    df = pd.merge(df, ip_hour_channel, on=['ip', 'hour', 'channel'], how='left')\n    del ip_hour_channel\n    df = pd.merge(df, ip_hour_os, on=['ip', 'hour', 'os'], how='left')\n    del ip_hour_os\n    df = pd.merge(df, ip_hour_app, on=['ip', 'hour', 'app'], how='left')\n    del ip_hour_app\n    df = pd.merge(df, ip_hour_device, on=['ip', 'hour', 'device'], how='left')\n    del ip_hour_device\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2021-10-16T13:39:22.667102Z","iopub.execute_input":"2021-10-16T13:39:22.667408Z","iopub.status.idle":"2021-10-16T13:39:22.680225Z","shell.execute_reply.started":"2021-10-16T13:39:22.667366Z","shell.execute_reply":"2021-10-16T13:39:22.679441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample = grouped_features(train_sample)","metadata":{"execution":{"iopub.status.busy":"2021-10-16T13:39:24.015730Z","iopub.execute_input":"2021-10-16T13:39:24.016042Z","iopub.status.idle":"2021-10-16T13:39:24.451620Z","shell.execute_reply.started":"2021-10-16T13:39:24.016013Z","shell.execute_reply":"2021-10-16T13:39:24.450754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample.head()","metadata":{"_cell_guid":"dcb86f85-6418-40dd-8e4a-670f1df0b3fc","_uuid":"3fffc4f935664d861cc0556f78f654c0c7121dac","scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:39:25.716440Z","iopub.execute_input":"2021-10-16T13:39:25.716740Z","iopub.status.idle":"2021-10-16T13:39:25.731476Z","shell.execute_reply.started":"2021-10-16T13:39:25.716706Z","shell.execute_reply":"2021-10-16T13:39:25.730587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Modelling\n","metadata":{}},{"cell_type":"code","source":"# create x and y train\nX = train_sample.drop('is_attributed', axis=1)\ny = train_sample[['is_attributed']]\n\n# split data into train and test/validation sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=101)\nprint(X_train.shape)\nprint(y_train.shape)\nprint(X_test.shape)\nprint(y_test.shape)","metadata":{"_cell_guid":"2b85ad1e-932b-4570-99ad-df89d8e09eee","_uuid":"fb851915b9cd57bb5fd85a0d08b317c695db015b","execution":{"iopub.status.busy":"2021-10-16T13:40:06.756976Z","iopub.execute_input":"2021-10-16T13:40:06.757429Z","iopub.status.idle":"2021-10-16T13:40:06.785501Z","shell.execute_reply.started":"2021-10-16T13:40:06.757399Z","shell.execute_reply":"2021-10-16T13:40:06.784890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2021-10-16T13:40:07.616137Z","iopub.execute_input":"2021-10-16T13:40:07.616712Z","iopub.status.idle":"2021-10-16T13:40:07.631193Z","shell.execute_reply.started":"2021-10-16T13:40:07.616674Z","shell.execute_reply":"2021-10-16T13:40:07.630380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check the average download rates in train and test data, should be comparable\nprint(y_train.mean())\nprint(y_test.mean())","metadata":{"_cell_guid":"d0e93c90-38df-486b-b7e5-52c82b21d821","_uuid":"116f681214fa50bf5710df325d37203a10d66df2","scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:40:08.746265Z","iopub.execute_input":"2021-10-16T13:40:08.746597Z","iopub.status.idle":"2021-10-16T13:40:08.756844Z","shell.execute_reply.started":"2021-10-16T13:40:08.746562Z","shell.execute_reply":"2021-10-16T13:40:08.755871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### AdaBoost","metadata":{}},{"cell_type":"code","source":"# adaboost classifier with max 600 decision trees of depth=2\n# learning_rate/shrinkage=1.5\n\n# base estimator\ntree = DecisionTreeClassifier(max_depth=2)\n\n# adaboost with the tree as base estimator\nadaboost_model_1 = AdaBoostClassifier(\n    base_estimator=tree,\n    n_estimators=600,\n    learning_rate=1.5,\n    algorithm=\"SAMME\")","metadata":{"execution":{"iopub.status.busy":"2021-10-16T13:40:28.806139Z","iopub.execute_input":"2021-10-16T13:40:28.806441Z","iopub.status.idle":"2021-10-16T13:40:28.811515Z","shell.execute_reply.started":"2021-10-16T13:40:28.806407Z","shell.execute_reply":"2021-10-16T13:40:28.810698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fit\nadaboost_model_1.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2021-10-16T13:40:30.266631Z","iopub.execute_input":"2021-10-16T13:40:30.267324Z","iopub.status.idle":"2021-10-16T13:41:18.638258Z","shell.execute_reply.started":"2021-10-16T13:40:30.267274Z","shell.execute_reply":"2021-10-16T13:41:18.637634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predictions\n# the second column represents the probability of a click resulting in a download\npredictions = adaboost_model_1.predict_proba(X_test)\npredictions[:10]","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2021-10-16T13:41:18.639426Z","iopub.execute_input":"2021-10-16T13:41:18.640212Z","iopub.status.idle":"2021-10-16T13:41:19.161184Z","shell.execute_reply.started":"2021-10-16T13:41:18.640178Z","shell.execute_reply":"2021-10-16T13:41:19.160540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# metrics: AUC\nmetrics.roc_auc_score(y_test, predictions[:,1])","metadata":{"execution":{"iopub.status.busy":"2021-10-16T13:41:19.162056Z","iopub.execute_input":"2021-10-16T13:41:19.162869Z","iopub.status.idle":"2021-10-16T13:41:19.177347Z","shell.execute_reply.started":"2021-10-16T13:41:19.162825Z","shell.execute_reply":"2021-10-16T13:41:19.176609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}