{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Fraud risk is everywhere, but for companies that advertise online, click fraud can happen at an overwhelming volume, resulting in misleading click data and wasted money. Ad channels can drive up costs by simply clicking on the ad at a large scale. With over 1 billion smart mobile devices in active use every month, China is the largest\nmobile market in the world and therefore suffers from huge volumes of fradulent traffic.\n\nTalkingData, China’s largest independent big data service platform, covers over 70% of active mobile devices nationwide. They handle 3 billion clicks per day, of which 90% are potentially fraudulent. Their current approach to prevent click fraud for app developers is to measure the journey of a user’s click across their portfolio, and flag IP addresses who produce lots of clicks, but never end up installing apps. With this information, they've built an IP blacklist and device blacklist.\n\nWhile successful, they want to always be one step ahead of fraudsters and have turned to the Kaggle community for help in further developing their solution. In their 2nd competition with Kaggle,\n\n**your mission Jim, should you choose to accept it**\n\n\nyou’re challenged to build an algorithm that predicts whether a user will download an app after clicking a mobile app ad. To support your modeling, they have provided a generous dataset covering approximately 200 million clicks over 4 days!","metadata":{}},{"cell_type":"markdown","source":"## A simple solution attempt","metadata":{}},{"cell_type":"markdown","source":"Data fields\n\nEach row of the training data contains a click record, with the following features.\n\n    ip: ip address of click.\n    app: app id for marketing.\n    device: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n    os: os version id of user mobile phone\n    channel: channel id of mobile ad publisher\n    click_time: timestamp of click (UTC)\n    attributed_time: if user download the app for after clicking an ad, this is the time of the app download\n    is_attributed: the target that is to be predicted, indicating the app was downloaded\n\nNote that ip, app, device, os, and channel are encoded.\n\nThe test data is similar, with the following differences:\n\n    click_id: reference for making predictions\n    is_attributed: not included\n","metadata":{}},{"cell_type":"markdown","source":"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/overview","metadata":{}},{"cell_type":"code","source":"# basics\nimport os\nimport numpy as np\nimport pandas as pd\nimport datetime as dt\n\n#graphs\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n#models\nfrom sklearn.cluster import KMeans\nfrom xgboost import XGBClassifier\nimport lightgbm as lgb\n\n#intermediary tools\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix, classification_report\nfrom imblearn.over_sampling import SMOTE\nfrom sklearn.metrics import silhouette_score","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:12:01.101741Z","iopub.execute_input":"2021-07-09T13:12:01.102243Z","iopub.status.idle":"2021-07-09T13:12:04.408183Z","shell.execute_reply.started":"2021-07-09T13:12:01.102161Z","shell.execute_reply":"2021-07-09T13:12:04.406866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#set globals and constants\nrandom_state = 42","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:12:04.409839Z","iopub.execute_input":"2021-07-09T13:12:04.410190Z","iopub.status.idle":"2021-07-09T13:12:04.414541Z","shell.execute_reply.started":"2021-07-09T13:12:04.410161Z","shell.execute_reply":"2021-07-09T13:12:04.413133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# run on all/ any path\ndef get_data_files(filePaths, hdr = 'infer'):\n    data=pd.DataFrame()\n    for csvfile in filePaths:\n        df = pd.read_csv(csvfile, header = hdr)\n        data=pd.concat([df,data],ignore_index=True)\n    return data","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:12:04.416955Z","iopub.execute_input":"2021-07-09T13:12:04.417325Z","iopub.status.idle":"2021-07-09T13:12:04.435147Z","shell.execute_reply.started":"2021-07-09T13:12:04.417285Z","shell.execute_reply":"2021-07-09T13:12:04.433532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def x_elbow(df,range0=np.arange(2,10)):\n    distortions = []\n    silhuettes = []\n\n#K = range(1,10)\n    for k in range0:\n        x_cluster = KMeans(n_clusters=k,init='k-means++', n_init=20, random_state=random_state,max_iter=400)\n        x_cluster.fit(df)\n        distortions.append(x_cluster.inertia_)\n        silhuettes.append(silhouette_score(df, x_cluster.labels_, metric='euclidean'))\n\n    #https://matplotlib.org/2.2.5/gallery/api/two_scales.html\n    fig, ax1 = plt.subplots()\n\n    color = 'tab:red'\n    ax1.set_xlabel('k')\n    ax1.set_ylabel('Distortion', color=color)\n    ax1.plot(range0, distortions, 'bx-')\n    ax1.tick_params(axis='y', labelcolor=color)\n\n    ax2 = ax1.twinx()  # instantiate a second axes that shares the same x-axis\n\n    color = 'tab:blue'\n    ax2.set_ylabel('silhuette score', color=color)  # we already handled the x-label with ax1\n    ax2.plot(range0, silhuettes, color=color)\n    ax2.tick_params(axis='y', labelcolor=color)\n    \n    return distortions, silhuettes","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:18:53.123337Z","iopub.execute_input":"2021-07-09T13:18:53.123694Z","iopub.status.idle":"2021-07-09T13:18:53.134284Z","shell.execute_reply.started":"2021-07-09T13:18:53.123652Z","shell.execute_reply":"2021-07-09T13:18:53.133364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def x_hist_stats(series0,title = ''):\n    num_bins = 50\n    fig, ax = plt.subplots(1,1, tight_layout = True)\n    ax.hist(series0, num_bins)\n    fig.tight_layout()\n    plt.title(title)\n    plt.show()\n    \n    print('mean ' + title + ': ' + str(series0.mean()))\n    print('std ' + title + ': ' + str(series0.std()))\n    print('median ' + title + ': ' + str(series0.median()))","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:35:49.306837Z","iopub.execute_input":"2021-07-09T13:35:49.307118Z","iopub.status.idle":"2021-07-09T13:35:49.313033Z","shell.execute_reply.started":"2021-07-09T13:35:49.307096Z","shell.execute_reply":"2021-07-09T13:35:49.311842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The files pool\n<pre>\n/kaggle/input/talkingdata-adtracking-fraud-detection/sample_submission.csv\n/kaggle/input/talkingdata-adtracking-fraud-detection/train_sample.csv\n/kaggle/input/talkingdata-adtracking-fraud-detection/test_supplement.csv\n/kaggle/input/talkingdata-adtracking-fraud-detection/train.csv\n/kaggle/input/talkingdata-adtracking-fraud-detection/test.csv\n</pre>\n","metadata":{}},{"cell_type":"code","source":"filePaths = ['../input/talkingdata-adtracking-fraud-detection/train_sample.csv']\nbase = get_data_files(filePaths)\n\nfilePaths = []\n#get the extra couple of samples\nfor dirname, _, filenames in os.walk('/kaggle/input/adtracking-click-for-app-250k-samples-from-total'):\n    for filename in filenames:\n        filePaths.append(os.path.join(dirname, filename))\n\nextra_train = get_data_files(filePaths, None)\nextra_train.columns = base.columns","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:18:53.460749Z","iopub.execute_input":"2021-07-09T13:18:53.461209Z","iopub.status.idle":"2021-07-09T13:18:54.157447Z","shell.execute_reply.started":"2021-07-09T13:18:53.461169Z","shell.execute_reply":"2021-07-09T13:18:54.156768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:18:54.158797Z","iopub.execute_input":"2021-07-09T13:18:54.159174Z","iopub.status.idle":"2021-07-09T13:18:54.164388Z","shell.execute_reply.started":"2021-07-09T13:18:54.159137Z","shell.execute_reply":"2021-07-09T13:18:54.163095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:18:54.166612Z","iopub.execute_input":"2021-07-09T13:18:54.166997Z","iopub.status.idle":"2021-07-09T13:18:54.192410Z","shell.execute_reply.started":"2021-07-09T13:18:54.166958Z","shell.execute_reply":"2021-07-09T13:18:54.190750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base[base['is_attributed'] == 1].tail(100)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:18:54.194355Z","iopub.execute_input":"2021-07-09T13:18:54.194771Z","iopub.status.idle":"2021-07-09T13:18:54.217991Z","shell.execute_reply.started":"2021-07-09T13:18:54.194724Z","shell.execute_reply":"2021-07-09T13:18:54.217400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *Feature engineering, Time analysis*","metadata":{}},{"cell_type":"code","source":"base['click_time'] = pd.to_datetime(base['click_time'])\nbase['hr'] = base['click_time'].dt.hour\nbase['day'] = base['click_time'].dt.day\nbase['weekday'] = base['click_time'].dt.weekday\nbase['month'] = base['click_time'].dt.month\nbase['attributed_time'] = pd.to_datetime(base['attributed_time'])\nbase['hr_at'] = base['attributed_time'].dt.hour","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:18:56.284551Z","iopub.execute_input":"2021-07-09T13:18:56.284858Z","iopub.status.idle":"2021-07-09T13:18:56.346785Z","shell.execute_reply.started":"2021-07-09T13:18:56.284833Z","shell.execute_reply":"2021-07-09T13:18:56.345599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"attributed_time = pd.DataFrame(base[['click_time','attributed_time']][~base['attributed_time'].isnull()])\nattributed_time['attr_hr'] = pd.to_datetime(attributed_time['attributed_time']).dt.hour\nattributed_time['click_hr'] = pd.to_datetime(attributed_time['click_time']).dt.hour\nattributed_time['click2attr'] = (pd.to_datetime(attributed_time['attributed_time'])-pd.to_datetime(attributed_time['click_time'])).astype('timedelta64[m]')\nattributed_time.head(200)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:19:59.364390Z","iopub.execute_input":"2021-07-09T13:19:59.364700Z","iopub.status.idle":"2021-07-09T13:19:59.389918Z","shell.execute_reply.started":"2021-07-09T13:19:59.364654Z","shell.execute_reply":"2021-07-09T13:19:59.389189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_hist_stats(attributed_time['attr_hr'])","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:35:56.948358Z","iopub.execute_input":"2021-07-09T13:35:56.948697Z","iopub.status.idle":"2021-07-09T13:35:57.206899Z","shell.execute_reply.started":"2021-07-09T13:35:56.948655Z","shell.execute_reply":"2021-07-09T13:35:57.206010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_hist_stats(attributed_time['click_hr'])","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:35:59.356567Z","iopub.execute_input":"2021-07-09T13:35:59.356883Z","iopub.status.idle":"2021-07-09T13:35:59.613735Z","shell.execute_reply.started":"2021-07-09T13:35:59.356857Z","shell.execute_reply":"2021-07-09T13:35:59.612907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_hist_stats(base['hr'][base['is_attributed'] == 0])","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:36:01.611805Z","iopub.execute_input":"2021-07-09T13:36:01.612108Z","iopub.status.idle":"2021-07-09T13:36:01.870184Z","shell.execute_reply.started":"2021-07-09T13:36:01.612081Z","shell.execute_reply":"2021-07-09T13:36:01.869334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_hist_stats(attributed_time['click2attr'])","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:37:08.236626Z","iopub.execute_input":"2021-07-09T13:37:08.236942Z","iopub.status.idle":"2021-07-09T13:37:08.518645Z","shell.execute_reply.started":"2021-07-09T13:37:08.236917Z","shell.execute_reply":"2021-07-09T13:37:08.517262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can see that there are slight differences between clicks for attributed customers and none attributec customer.\nwe can also see that most click are attributed immediately\nnext we can check the attribution time per hour, to see if in certain hours the attibution is immediate and in certain hours not\nthe best way to look at the data is the median, since the long tail affects the average","metadata":{}},{"cell_type":"code","source":"attributed_time.groupby('click_hr').agg({'click2attr':'median'})","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:42:19.430852Z","iopub.execute_input":"2021-07-09T13:42:19.431209Z","iopub.status.idle":"2021-07-09T13:42:19.448336Z","shell.execute_reply.started":"2021-07-09T13:42:19.431179Z","shell.execute_reply":"2021-07-09T13:42:19.447353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can use that as some sort of a probability estimator for attribution, which we'll be able to apply on any test data. <br>\nwe'll need to save the results in a table, merge it on the train and on any test/ validation data <br>\nin hours case, we can afford to add it to the data before splitting it to train and test","metadata":{}},{"cell_type":"code","source":"click2attr_per_hour = attributed_time.groupby('click_hr').agg({'click2attr':'median'})\nbase = pd.merge(base,click2attr_per_hour,left_on='hr',right_index=True)\nbase.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:47:17.487853Z","iopub.execute_input":"2021-07-09T13:47:17.488247Z","iopub.status.idle":"2021-07-09T13:47:17.526395Z","shell.execute_reply.started":"2021-07-09T13:47:17.488221Z","shell.execute_reply":"2021-07-09T13:47:17.525042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base['click_time']=base['click_time'].map(dt.datetime.toordinal)\nbase['attributed_time']=base['attributed_time'].map(dt.datetime.toordinal)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:51:09.585560Z","iopub.execute_input":"2021-07-09T13:51:09.585883Z","iopub.status.idle":"2021-07-09T13:51:10.691045Z","shell.execute_reply.started":"2021-07-09T13:51:09.585859Z","shell.execute_reply":"2021-07-09T13:51:10.689642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base.columns","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:51:10.692581Z","iopub.execute_input":"2021-07-09T13:51:10.692981Z","iopub.status.idle":"2021-07-09T13:51:10.699940Z","shell.execute_reply.started":"2021-07-09T13:51:10.692944Z","shell.execute_reply":"2021-07-09T13:51:10.698784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corrdf = base\nsns.set_theme(style=\"white\")\ncorr = corrdf.corr()\nf, ax = plt.subplots(figsize=(11, 9))\ncolormap = sns.diverging_palette(230, 20, as_cmap=True)\nmask = np.triu(np.ones_like(corr, dtype=bool))\nsns.heatmap(corr, mask=mask, cmap=colormap, vmax=.3, center=0,\n            square=True, linewidths=.5, cbar_kws={\"shrink\": .5})","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:51:10.701950Z","iopub.execute_input":"2021-07-09T13:51:10.702290Z","iopub.status.idle":"2021-07-09T13:51:11.366694Z","shell.execute_reply.started":"2021-07-09T13:51:10.702227Z","shell.execute_reply":"2021-07-09T13:51:11.365225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"so, our y variable, is_attributed, is mostly correlated with ip and app, and the app is very much correlated (for obvious reasons, with os and device)\nthe hour of the attribution (hr_at) is, in turn correlated with the ip","metadata":{}},{"cell_type":"markdown","source":"Looking for clusters ofs apps, OSs, devices and channels in order to find a common pattern and use that as a predictor","metadata":{}},{"cell_type":"code","source":"base['hr_at'].fillna(-1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:51:38.753267Z","iopub.execute_input":"2021-07-09T13:51:38.753816Z","iopub.status.idle":"2021-07-09T13:51:38.759922Z","shell.execute_reply.started":"2021-07-09T13:51:38.753775Z","shell.execute_reply":"2021-07-09T13:51:38.758098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Sampling for KMeans\nKMtrain = pd.DataFrame()\nKMtrain = base.sample(n = 10000, replace=False, random_state = random_state) # 10% of 100,000","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:51:40.079326Z","iopub.execute_input":"2021-07-09T13:51:40.079701Z","iopub.status.idle":"2021-07-09T13:51:40.095586Z","shell.execute_reply.started":"2021-07-09T13:51:40.079651Z","shell.execute_reply":"2021-07-09T13:51:40.093927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dists, sils = x_elbow(KMtrain[['app','device','os','channel','ip','click2attr']],np.arange(2,10))","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:53:53.508527Z","iopub.execute_input":"2021-07-09T13:53:53.508904Z","iopub.status.idle":"2021-07-09T13:54:34.365797Z","shell.execute_reply.started":"2021-07-09T13:53:53.508874Z","shell.execute_reply":"2021-07-09T13:54:34.364549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the elbow shows us 4 clusters (dustortion) will be the best, the silhuette score shows us 4 as well. 4 it is","metadata":{}},{"cell_type":"code","source":"kmeans = KMeans(init=\"k-means++\", n_clusters=8, n_init=6,random_state=random_state,max_iter=300).fit(base[['app','device','os','channel','ip','click2attr']])\nbase['cluster'] = kmeans.predict(base[['app','device','os','channel','ip','click2attr']])\n#kmeans.fit(KMtrain) #,'channel'","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:55:32.332149Z","iopub.execute_input":"2021-07-09T13:55:32.332467Z","iopub.status.idle":"2021-07-09T13:55:34.701404Z","shell.execute_reply.started":"2021-07-09T13:55:32.332441Z","shell.execute_reply":"2021-07-09T13:55:34.700464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train, test = train_test_split(base, test_size=0.3)\n\nY_train = train['is_attributed']\nX_train = train.drop(['is_attributed','attributed_time','hr_at'], axis = 1, inplace=False)\nY_test = test['is_attributed']\nX_test = test.drop(['is_attributed','attributed_time','hr_at'], axis = 1, inplace=False)\nX_test.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:55:34.702874Z","iopub.execute_input":"2021-07-09T13:55:34.703169Z","iopub.status.idle":"2021-07-09T13:55:34.736442Z","shell.execute_reply.started":"2021-07-09T13:55:34.703138Z","shell.execute_reply":"2021-07-09T13:55:34.735356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sm = SMOTE(random_state = random_state)\nX_train, Y_train = sm.fit_resample(X_train, Y_train.ravel())","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:55:34.737920Z","iopub.execute_input":"2021-07-09T13:55:34.738146Z","iopub.status.idle":"2021-07-09T13:55:34.783098Z","shell.execute_reply.started":"2021-07-09T13:55:34.738122Z","shell.execute_reply":"2021-07-09T13:55:34.782137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('total cases: ' + str(Y_train.size))\nprint('total attributed: ' + str(sum(Y_train)))\nprint('ratio: ' + str(sum(Y_train) / Y_train.size))\nprint('well, after smoting, what did we expect')","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:55:34.784639Z","iopub.execute_input":"2021-07-09T13:55:34.785008Z","iopub.status.idle":"2021-07-09T13:55:34.842107Z","shell.execute_reply.started":"2021-07-09T13:55:34.784977Z","shell.execute_reply":"2021-07-09T13:55:34.841017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = XGBClassifier(\n            random_state = random_state, \n            #scale_pos_weight = 30,\n            learning_rate = 0.1,\n            max_depth= 4,\n            min_child_weight= 4,\n            subsample = 0.9,\n            colsample_bytree = 0.8,\n            colsample_bylevel = 0.8,\n            reg_lambda = 0.6\n\n)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:55:36.254840Z","iopub.execute_input":"2021-07-09T13:55:36.255145Z","iopub.status.idle":"2021-07-09T13:55:36.260899Z","shell.execute_reply.started":"2021-07-09T13:55:36.255120Z","shell.execute_reply":"2021-07-09T13:55:36.259595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(X_train, Y_train)\nprediction = model.predict(X_test)\nprobabilities = model.predict_proba(X_test)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:55:44.181941Z","iopub.execute_input":"2021-07-09T13:55:44.182219Z","iopub.status.idle":"2021-07-09T13:55:47.916055Z","shell.execute_reply.started":"2021-07-09T13:55:44.182196Z","shell.execute_reply":"2021-07-09T13:55:47.915080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(Y_test, prediction))","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:55:47.917371Z","iopub.execute_input":"2021-07-09T13:55:47.917820Z","iopub.status.idle":"2021-07-09T13:55:47.955494Z","shell.execute_reply.started":"2021-07-09T13:55:47.917777Z","shell.execute_reply":"2021-07-09T13:55:47.954776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = lgb.LGBMClassifier(\n    num_leaves = 10,\n    max_bin = 45\n)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:56:12.424334Z","iopub.execute_input":"2021-07-09T13:56:12.424610Z","iopub.status.idle":"2021-07-09T13:56:12.429406Z","shell.execute_reply.started":"2021-07-09T13:56:12.424586Z","shell.execute_reply":"2021-07-09T13:56:12.428139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(X_train, Y_train)\nprediction = model.predict(X_test)\nprobabilities = model.predict_proba(X_test)","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:56:12.717829Z","iopub.execute_input":"2021-07-09T13:56:12.718141Z","iopub.status.idle":"2021-07-09T13:56:13.718369Z","shell.execute_reply.started":"2021-07-09T13:56:12.718114Z","shell.execute_reply":"2021-07-09T13:56:13.717604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(Y_test, prediction))","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:56:13.719583Z","iopub.execute_input":"2021-07-09T13:56:13.720126Z","iopub.status.idle":"2021-07-09T13:56:13.760058Z","shell.execute_reply.started":"2021-07-09T13:56:13.720093Z","shell.execute_reply":"2021-07-09T13:56:13.759333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# vars used\n\nplt.figure(figsize=(12,6))\nfeat_importances = pd.Series(model.feature_importances_, index=X_train.columns)\nfeat_importances.nlargest(25).sort_values().plot(kind='barh')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:56:15.988770Z","iopub.execute_input":"2021-07-09T13:56:15.990004Z","iopub.status.idle":"2021-07-09T13:56:16.242933Z","shell.execute_reply.started":"2021-07-09T13:56:15.989967Z","shell.execute_reply":"2021-07-09T13:56:16.241922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#https://www.kaggle.com/ravikishore/titanic-survival-prediction\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.metrics import roc_curve\n\nlogit_roc_auc = roc_auc_score(Y_test, prediction)\nfpr, tpr, thresholds = roc_curve(Y_test, prediction)\nplt.figure()\nplt.plot(fpr, tpr, label='Light GBM (area = %0.2f)' % logit_roc_auc)\nplt.plot([0, 1], [0, 1],'r--')\nplt.xlim([0.0, 1.0])\nplt.ylim([0.0, 1.05])\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Receiver operating characteristic')\nplt.legend(loc=\"lower right\")\nplt.savefig('Log_ROC')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-07-09T13:58:19.665455Z","iopub.execute_input":"2021-07-09T13:58:19.665859Z","iopub.status.idle":"2021-07-09T13:58:19.948784Z","shell.execute_reply.started":"2021-07-09T13:58:19.665832Z","shell.execute_reply":"2021-07-09T13:58:19.947535Z"},"trusted":true},"execution_count":null,"outputs":[]}]}