{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TalkingData AdTracking Fraud Detection Challenge\n<!-- ![](https://i.ytimg.com/vi/0V8CSoO23tw/maxresdefault.jpg) -->\n\nTalkingData is back with another competition: This time, our task is to predict where a click on some advertising is fraudlent given a few basic attributes about the device that made the click. What sets this competition apart is the sheer scale of the dataset: with 240 million rows it might be the biggest one I've seen on Kaggle so far.\n\nThere are some similarities with the last competition TalkingData launched: https://www.kaggle.com/c/talkingdata-mobile-user-demographics - that competition was about predicting the demographics of a user given their activity, and you can view this as a similar problem (predicting whether a user is real or not given their activity). However, that competition was plagued by a [leak](https://www.kaggle.com/wiki/Leakage) where the dataset wasn't sorted properly and certain portions of the dataset had different demographic distribtions. This meant that by adding the row ID as a feature you could get a huge boost in performance. Let's hope TalkingData have learnt their lesson this time around. 😉\n\nLooking at the evaluation page, we can see that the evaluation metric used is** ROC-AUC** (the area under a curve on a Receiver Operator Characteristic graph).\nIn english, this means a few important things:\n* This competition is a **binary classification** problem - i.e. our target variable is a binary attribute (Is the user making the click fraudlent or not?) and our goal is to classify users into \"fraudlent\" or \"not fraudlent\" as well as possible\n* Unlike metrics such as [LogLoss](http://www.exegetic.biz/blog/2015/12/making-sense-logarithmic-loss/), the AUC score only depends on **how well you well you can separate the two classes**. In practice, this means that only the order of your predictions matter,\n    * As a result of this, any rescaling done to your model's output probabilities will have no effect on your score. In some other competitions, adding a constant or multiplier to your predictions to rescale it to the distribution can help but that doesn't apply here.\n  ","metadata":{"_uuid":"80f731850042d5be24eb6c28273cc25f678ed0c8"}},{"cell_type":"markdown","source":"### Looking at the columns\n\nAccording to the data page, our data contains:\n\n* `ip`: ip address of click\n* `app`: app id for marketing\n* `device`: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n* `os`: os version id of user mobile phone\n* `channel`: channel id of mobile ad publisher\n* `click_time`: timestamp of click (UTC)\n* `attributed_time`: if user download the app for after clicking an ad, this is the time of the app download\n* `is_attributed`: the target that is to be predicted, indicating the app was downloaded\n\n**A few things of note:**\n* If you look at the data samples above, you'll notice that all these variables are encoded - meaning we don't know what the actual value corresponds to - each value has instead been assigned an ID which we're given. This has likely been done because data such as IP addresses are sensitive, although it does unfortunately reduce the amount of feature engineering we can do on these.\n* The `attributed_time` variable is only available in the training set - it's not immediately useful for classification but it could be used for some interesting analysis (for example, one could fill in the variable in the test set by building a model to predict it).\n\nFor each of our encoded values, let's look at the number of unique values:","metadata":{}},{"cell_type":"markdown","source":"# Imports\nWe are using a typical data science stack: ``numpy``, ``pandas``, ``sklearn``, ``matplotlib``.","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport mlcrate as mlc\nimport os\nimport gc\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport time\nimport csv\nimport random\n\nfrom sklearn import preprocessing\nfrom matplotlib import pyplot as plt\nimport matplotlib as mpl\nimport scipy.stats as st\nfrom sklearn import ensemble, tree, linear_model\nimport missingno as msno\nimport math\nimport copy\nfrom sklearn.svm import SVC, LinearSVC\nfrom sklearn.ensemble import RandomForestClassifier\n\nfrom imblearn.ensemble import BalancedBaggingClassifier\nfrom sklearn.tree import DecisionTreeClassifier\n\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\n\nfrom matplotlib import pyplot\nfrom sklearn.metrics import make_scorer, accuracy_score\nfrom sklearn.model_selection import train_test_split \nimport gc\nfrom numpy import loadtxt\nfrom xgboost import XGBClassifier\nimport xgboost as xgb\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\nfrom sklearn import datasets, linear_model\nfrom sklearn.model_selection import train_test_split\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:29:41.773266Z","iopub.execute_input":"2023-06-11T12:29:41.773617Z","iopub.status.idle":"2023-06-11T12:29:42.826466Z","shell.execute_reply.started":"2023-06-11T12:29:41.773561Z","shell.execute_reply":"2023-06-11T12:29:42.825667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pal = sns.color_palette()\n\nprint('# File sizes')\nfor f in os.listdir('../input'):\n    if 'zip' not in f:\n        print(f.ljust(30) + str(round(os.path.getsize('../input/' + f) / 1000000, 2)) + 'MB')","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2023-06-11T12:29:42.827412Z","iopub.execute_input":"2023-06-11T12:29:42.827650Z","iopub.status.idle":"2023-06-11T12:29:42.842173Z","shell.execute_reply.started":"2023-06-11T12:29:42.827608Z","shell.execute_reply":"2023-06-11T12:29:42.841550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Wow, that is some really big data. Unfortunately we don't have enough kernel memory to load the full dataset into memory; however we can get a glimpse at some of the statistics:","metadata":{"_uuid":"33c5eec493e287aae60a9226d58fa61545d3a77a"}},{"cell_type":"code","source":"import subprocess\nprint('# Line count:')\nfor file in ['train.csv', 'test.csv', 'train_sample.csv']:\n    lines = subprocess.run(['wc', '-l', '../input/{}'.format(file)], stdout=subprocess.PIPE).stdout.decode('utf-8')\n    print(lines, end='', flush=True)","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_kg_hide-input":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.status.busy":"2023-06-11T12:29:42.843192Z","iopub.execute_input":"2023-06-11T12:29:42.843602Z","iopub.status.idle":"2023-06-11T12:30:19.430247Z","shell.execute_reply.started":"2023-06-11T12:29:42.843541Z","shell.execute_reply":"2023-06-11T12:30:19.429461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That makes **185 million rows** in the training set and ** 19 million** in the test set. Handily the organisers have provided a `train_sample.csv` which contains 100K rows in case you don't want to download the full data\n\nFor this analysis, I'm going to use the first 1M rows of the training and test datasets.\n\n## Glimpse of Data","metadata":{"_uuid":"b1ec574cb675150940b453bb9fbd648a5285390f"}},{"cell_type":"code","source":"df_train_full = pd.read_csv('../input/train.csv', nrows=1000000, parse_dates=['click_time'])\ndf_train = pd.read_csv('../input/train_sample.csv',  parse_dates=['click_time'])\ndf_test = pd.read_csv('../input/test.csv', parse_dates=['click_time'])","metadata":{"_uuid":"67335c75ba85c06463285be0cf78ccdc68e38609","execution":{"iopub.status.busy":"2023-06-11T12:30:19.431930Z","iopub.execute_input":"2023-06-11T12:30:19.432293Z","iopub.status.idle":"2023-06-11T12:30:50.941820Z","shell.execute_reply.started":"2023-06-11T12:30:19.432203Z","shell.execute_reply":"2023-06-11T12:30:50.941005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Training set","metadata":{}},{"cell_type":"code","source":"print('Training set:')\ndf_train.head()","metadata":{"_kg_hide-input":true,"_uuid":"e44fbc483dc19be1b44f2aba7d4356ac183534e9","execution":{"iopub.status.busy":"2023-06-11T12:30:50.943281Z","iopub.execute_input":"2023-06-11T12:30:50.943575Z","iopub.status.idle":"2023-06-11T12:30:50.965262Z","shell.execute_reply.started":"2023-06-11T12:30:50.943531Z","shell.execute_reply":"2023-06-11T12:30:50.964490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Test set","metadata":{}},{"cell_type":"code","source":"print('Test set:')\ndf_test.head()","metadata":{"_kg_hide-input":true,"_uuid":"65a54cab2e0492a5f984d761f7efbc7ed5377aee","execution":{"iopub.status.busy":"2023-06-11T12:30:50.966602Z","iopub.execute_input":"2023-06-11T12:30:50.967115Z","iopub.status.idle":"2023-06-11T12:30:50.984001Z","shell.execute_reply.started":"2023-06-11T12:30:50.967061Z","shell.execute_reply":"2023-06-11T12:30:50.983190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features = df_train.columns.values[0:30]\nunique_max_train = []\nunique_max_test = []\nfor feature in features:\n    values = df_train[feature].value_counts()\n    perc = values.max() / ( df_train.shape[0]*100 )\n    unique_max_train.append([feature, values.max(), values.idxmax(), perc])\n    \nnp.transpose((pd.DataFrame(unique_max_train, columns=['Feature', 'Max duplicados', 'Valor', 'Percentage'])).\\\n            sort_values(by = 'Max duplicados', ascending=False).head(15))","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:50.985317Z","iopub.execute_input":"2023-06-11T12:30:50.985778Z","iopub.status.idle":"2023-06-11T12:30:51.051921Z","shell.execute_reply.started":"2023-06-11T12:30:50.985725Z","shell.execute_reply":"2023-06-11T12:30:51.051253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target Distribution","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 8))\ncols = ['ip', 'app', 'device', 'os', 'channel']\nuniques = [len(df_train[col].unique()) for col in cols]\nsns.set(font_scale=1.2)\nax = sns.barplot(cols, uniques, palette=pal, log=True)\nax.set(xlabel='Feature', ylabel='log(unique count)', title='Number of unique values per feature')\nfor p, uniq in zip(ax.patches, uniques):\n    height = p.get_height()\n    ax.text(p.get_x()+p.get_width()/2.,\n            height + 10,\n            uniq,\n            ha=\"center\") \n# for col, uniq in zip(cols, uniques):\n#     ax.text(col, uniq, uniq, color='black', ha=\"center\")","metadata":{"_kg_hide-input":true,"_uuid":"8eb35b5d9e773650c20b43484f598a3dea5ecbb0","execution":{"iopub.status.busy":"2023-06-11T12:30:51.053141Z","iopub.execute_input":"2023-06-11T12:30:51.053456Z","iopub.status.idle":"2023-06-11T12:30:51.588409Z","shell.execute_reply.started":"2023-06-11T12:30:51.053406Z","shell.execute_reply":"2023-06-11T12:30:51.587648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Examine Missing Value\nNext we can look at the number and percentage of missing values in each column","metadata":{}},{"cell_type":"code","source":"# checking missing data\ntotal = df_train.isnull().sum().sort_values(ascending = False)\npercent = (df_train.isnull().sum()/df_train.isnull().count()*100).sort_values(ascending = False)\nmissing__train_data  = pd.concat([total, percent], axis=1, keys=['Total', 'Percent'])\nmissing__train_data.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:51.589633Z","iopub.execute_input":"2023-06-11T12:30:51.590097Z","iopub.status.idle":"2023-06-11T12:30:51.653667Z","shell.execute_reply.started":"2023-06-11T12:30:51.590046Z","shell.execute_reply":"2023-06-11T12:30:51.652868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Package called missingno (https://github.com/ResidentMario/missingno) ``!pip install quilt``","metadata":{}},{"cell_type":"code","source":"# !pip install quilt","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:51.655046Z","iopub.execute_input":"2023-06-11T12:30:51.655375Z","iopub.status.idle":"2023-06-11T12:30:51.659189Z","shell.execute_reply.started":"2023-06-11T12:30:51.655311Z","shell.execute_reply":"2023-06-11T12:30:51.658422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nullity Matrix\nThe msno.matrix nullity matrix is a data-dense display which lets you quickly visually analyse data completion\n\n","metadata":{}},{"cell_type":"code","source":"# import missingno as msno\n# msno.matrix(df_train.head(20000))","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:51.660561Z","iopub.execute_input":"2023-06-11T12:30:51.660872Z","iopub.status.idle":"2023-06-11T12:30:51.668686Z","shell.execute_reply.started":"2023-06-11T12:30:51.660804Z","shell.execute_reply":"2023-06-11T12:30:51.667879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Heatmap\nThe missingno correlation heatmap measures nullity correlation: how strongly the presence or absence of one variable affects the presence of another:","metadata":{}},{"cell_type":"code","source":"# msno.heatmap(df_train)","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:51.670196Z","iopub.execute_input":"2023-06-11T12:30:51.670580Z","iopub.status.idle":"2023-06-11T12:30:51.679256Z","shell.execute_reply.started":"2023-06-11T12:30:51.670516Z","shell.execute_reply":"2023-06-11T12:30:51.678608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"anually dealing with missing values will often improve model performance\n\n","metadata":{}},{"cell_type":"markdown","source":"##  Encoded variables statistics\n\nAlthough the actual values of these variables aren't helpful for us, it can still be useful to know what their distributions are. Note these statistics are computed on 1M samples, and so will be higher for the full dataset.","metadata":{"_uuid":"922f91e941b7b268c1fcd0aedb777ad4d24b7fe1"}},{"cell_type":"code","source":"# !pip uninstall plotly \n\n# !python -m pip install plotly","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:51.680524Z","iopub.execute_input":"2023-06-11T12:30:51.680830Z","iopub.status.idle":"2023-06-11T12:30:51.688101Z","shell.execute_reply.started":"2023-06-11T12:30:51.680760Z","shell.execute_reply":"2023-06-11T12:30:51.687452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from plotly import tools\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\n# from plotly.subplots import make_subplots\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:51.689497Z","iopub.execute_input":"2023-06-11T12:30:51.689799Z","iopub.status.idle":"2023-06-11T12:30:51.915572Z","shell.execute_reply.started":"2023-06-11T12:30:51.689743Z","shell.execute_reply":"2023-06-11T12:30:51.914809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col, uniq in zip(cols, uniques):\n    counts = df_train[col].value_counts()\n\n    sorted_counts = np.sort(counts.values)\n    fig = plt.figure()\n    ax = fig.add_subplot(1, 1, 1)\n    line, = ax.plot(sorted_counts, color='red')\n    ax.set_yscale('log')\n    plt.title(\"Distribution of value counts for {}\".format(col))\n    plt.ylabel('log(Occurence count)')\n    plt.xlabel('Index')\n    plt.show()\n    \n    fig = plt.figure()\n    ax = fig.add_subplot(1, 1, 1)\n    plt.hist(sorted_counts, bins=50)\n    ax.set_yscale('log', nonposy='clip')\n    plt.title(\"Histogram of value counts for {}\".format(col))\n    plt.ylabel('Number of IDs')\n    plt.xlabel('Occurences of value for ID')\n    plt.show()\n    \n    max_count = np.max(counts)\n    min_count = np.min(counts)\n    gt = [10, 100, 1000]\n    prop_gt = []\n    for value in gt:\n        prop_gt.append(round((counts > value).mean()*100, 2))\n    print(\"Variable '{}': | Unique values: {} | Count of most common: {} | Count of least common: {} | count>10: {}% | count>100: {}% | count>1000: {}%\".format(col, uniq, max_count, min_count, *prop_gt))\n    ","metadata":{"_kg_hide-input":true,"_uuid":"8d3bbdefa393b9365b2b58aeab8d2677c817a971","execution":{"iopub.status.busy":"2023-06-11T12:30:51.917750Z","iopub.execute_input":"2023-06-11T12:30:51.918008Z","iopub.status.idle":"2023-06-11T12:30:55.381816Z","shell.execute_reply.started":"2023-06-11T12:30:51.917963Z","shell.execute_reply":"2023-06-11T12:30:55.381006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis\nExploratory Data Analysis (EDA) is an open-ended process where we calculate statistics and make figures to find trends, anomalies, patterns, or relationships within the data","metadata":{}},{"cell_type":"markdown","source":"## What we're trying to predict","metadata":{"_uuid":"f46b1c067e51ba664e1c8a22434919456d7f4774"}},{"cell_type":"code","source":"plt.figure(figsize=(8, 8))\nsns.set(font_scale=1.2)\nmean = (df_train.is_attributed.values == 1).mean()\nax = sns.barplot(['Fraudulent (1)', 'Not Fradulent (0)'], [mean, 1-mean], palette=pal)\nax.set(xlabel='Target Value', ylabel='Probability', title='Target value distribution')\nfor p, uniq in zip(ax.patches, [mean, 1-mean]):\n    height = p.get_height()\n    ax.text(p.get_x()+p.get_width()/2.,\n            height+0.01,\n            '{}%'.format(round(uniq * 100, 2)),\n            ha=\"center\") ","metadata":{"_kg_hide-input":true,"_uuid":"efe6fde36b993a865e7a3f592e01061b1b49caa2","execution":{"iopub.status.busy":"2023-06-11T12:30:55.383272Z","iopub.execute_input":"2023-06-11T12:30:55.383787Z","iopub.status.idle":"2023-06-11T12:30:55.535065Z","shell.execute_reply.started":"2023-06-11T12:30:55.383734Z","shell.execute_reply":"2023-06-11T12:30:55.534229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is unbalanced dataset. Only 0.2% of the dataset is made up of fradulent clicks. This means that any models we run on the data will either need to be robust against class imbalance or will require some data resampling.","metadata":{"_uuid":"219351441d94b83e4e0ede20dcd7ac01973b7251"}},{"cell_type":"markdown","source":"# Correlation\nNow that we have dealt with the categorical variables and the outliers, let's continue with the EDA. One way to try and understand the data is by looking for correlations between the features and the target. We can calculate the Pearson correlation coefficient between every variable and the target using the .corr dataframe method.\n\nThe correlation coefficient is not the greatest method to represent \"relevance\" of a feature, but it does give us an idea of possible relationships within the data. [Some general interpretations](http://www.statstutor.ac.uk/resources/uploaded/pearsons.pdf) of the absolute value of the correlation coefficent are:\n\n- .00-.19 “very weak”\n- .20-.39 “weak”\n- .40-.59 “moderate”\n- .60-.79 “strong”\n- .80-1.0 “very strong”","metadata":{}},{"cell_type":"code","source":"print(\"Train Data Heatmap of Correlations\")\nplt.figure(figsize = (20, 8))\n\n# Heatmap of correlations\nsns.heatmap(df_train.corr(), cmap = plt.cm.RdYlBu_r, vmin = -0.25, annot = True, vmax = 0.6)\nplt.title('Correlation Heatmap');","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:55.536520Z","iopub.execute_input":"2023-06-11T12:30:55.537038Z","iopub.status.idle":"2023-06-11T12:30:55.873650Z","shell.execute_reply.started":"2023-06-11T12:30:55.536879Z","shell.execute_reply":"2023-06-11T12:30:55.872869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def Parse_time(df):\n    df['day'] = df['click_time'].dt.day.astype('uint8')\n    df['hour'] = df['click_time'].dt.hour.astype('uint8')\n    df['minute'] = df['click_time'].dt.minute.astype('uint8')\n    df['second'] = df['click_time'].dt.second.astype('uint8')","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:55.874859Z","iopub.execute_input":"2023-06-11T12:30:55.875307Z","iopub.status.idle":"2023-06-11T12:30:55.884477Z","shell.execute_reply.started":"2023-06-11T12:30:55.875255Z","shell.execute_reply":"2023-06-11T12:30:55.883660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Parse_time(df_train)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:55.885795Z","iopub.execute_input":"2023-06-11T12:30:55.886369Z","iopub.status.idle":"2023-06-11T12:30:55.924038Z","shell.execute_reply.started":"2023-06-11T12:30:55.886314Z","shell.execute_reply":"2023-06-11T12:30:55.923291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Parse_time(df_test)\ndf_test.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:55.925465Z","iopub.execute_input":"2023-06-11T12:30:55.925771Z","iopub.status.idle":"2023-06-11T12:30:59.714960Z","shell.execute_reply.started":"2023-06-11T12:30:55.925709Z","shell.execute_reply":"2023-06-11T12:30:59.714229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def Drop_cols(df, x):\n    Num_of_line = 100\n    print(Num_of_line*'=')\n    print('Before drop =\\n', df.head(3))\n    print(Num_of_line*'=')\n    df.drop(labels = x, axis = 1, inplace = True)\n    print('After drop =\\n', df.head(3))\n    return df\n\ndef Normalized(df):\n    df_col_names = df.columns\n    x = df.values \n    min_max_scaler = preprocessing.MinMaxScaler()\n    x_scaled = min_max_scaler.fit_transform(x)\n    df = pd.DataFrame(x_scaled)\n    df.columns = df_col_names","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:59.716245Z","iopub.execute_input":"2023-06-11T12:30:59.716722Z","iopub.status.idle":"2023-06-11T12:30:59.736485Z","shell.execute_reply.started":"2023-06-11T12:30:59.716667Z","shell.execute_reply":"2023-06-11T12:30:59.735710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Drop and normalize \nprint('Drop colum and normalize, training data...!');\ncolmn_names = ['attributed_time']; \ndf_train = Drop_cols(df_train, colmn_names)\n# df_train = Normalized(df_train)\nprint('Drop colum and normalize, training data, Done!'); ","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:59.737650Z","iopub.execute_input":"2023-06-11T12:30:59.737966Z","iopub.status.idle":"2023-06-11T12:30:59.764871Z","shell.execute_reply.started":"2023-06-11T12:30:59.737904Z","shell.execute_reply":"2023-06-11T12:30:59.764064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# #Drop and normalize \n# print('Drop colum and normalize, testing data...!');\n# colmn_names = ['click_time']; \n# # df_test = Drop_cols(df_test, colmn_names)\n# # df_train = Normalized(df_train)\n# print('Drop colum and normalize, testing data, Done!'); ","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:59.766307Z","iopub.execute_input":"2023-06-11T12:30:59.766608Z","iopub.status.idle":"2023-06-11T12:30:59.771305Z","shell.execute_reply.started":"2023-06-11T12:30:59.766547Z","shell.execute_reply":"2023-06-11T12:30:59.770526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# Not the best solution to ValueError: y contains previously unseen labels: [0, 1, 2,...\nunknown_value = -1 #Make sure this is int (as other labels) or you will not be able to predict in the end ⚠️\n\nfrom sklearn import preprocessing\n\ncat_features = ['ip', 'app', 'device', 'os', 'channel']\n\n#encoder = preprocessing.LabelEncoder() - Incorrect, we need a label encoder for each feature\n# Create new columns in clicks using preprocessing.LabelEncoder()\n\nfor feature in cat_features:\n    #New encoder for each feature\n    encoder = preprocessing.LabelEncoder()\n    #Fit on all possible values of this feature\n    encoder.fit(df_train[feature])\n    #Create LabelEncoder of input to output\n    le_dict = dict(zip(encoder.classes_, encoder.transform(encoder.classes_)))\n    #Encode unseen values to the unknown_value label\n    encoded = df_train[feature].apply(lambda x: le_dict.get(x, unknown_value))\n    df_train[feature+'_labels'] = encoded\n    \n    #Competition submission\n    competition_encoded = df_test[feature].apply(lambda x: le_dict.get(x, unknown_value))\n    #ValueError: y contains previously unseen labels: [0, 2, 3, 4, 5,\n    df_test[feature+'_labels'] = competition_encoded","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:30:59.772820Z","iopub.execute_input":"2023-06-11T12:30:59.773134Z","iopub.status.idle":"2023-06-11T12:32:47.300279Z","shell.execute_reply.started":"2023-06-11T12:30:59.773068Z","shell.execute_reply":"2023-06-11T12:32:47.299519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:32:47.301722Z","iopub.execute_input":"2023-06-11T12:32:47.302177Z","iopub.status.idle":"2023-06-11T12:32:47.325823Z","shell.execute_reply.started":"2023-06-11T12:32:47.302121Z","shell.execute_reply":"2023-06-11T12:32:47.325084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:32:47.327076Z","iopub.execute_input":"2023-06-11T12:32:47.327624Z","iopub.status.idle":"2023-06-11T12:32:47.352303Z","shell.execute_reply.started":"2023-06-11T12:32:47.327490Z","shell.execute_reply":"2023-06-11T12:32:47.351621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ip_labels_unknowns = sum(df_train['ip_labels'] == unknown_value)\ntrain_ip_labels_unknowns","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:32:47.353784Z","iopub.execute_input":"2023-06-11T12:32:47.354079Z","iopub.status.idle":"2023-06-11T12:32:47.725737Z","shell.execute_reply.started":"2023-06-11T12:32:47.354006Z","shell.execute_reply":"2023-06-11T12:32:47.724917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"compet_test_ip_labels_unknowns = sum(df_test['ip_labels'] == unknown_value)\ncompet_test_ip_labels_unknowns","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:32:47.727126Z","iopub.execute_input":"2023-06-11T12:32:47.727477Z","iopub.status.idle":"2023-06-11T12:33:58.934594Z","shell.execute_reply.started":"2023-06-11T12:32:47.727413Z","shell.execute_reply":"2023-06-11T12:33:58.932986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Read only first limit rows\nlimit = 20_000_000\nmy_own_metrics={'limit': min(limit, df_train.shape[0]),\n                'df_test':df_test.shape[0],\n                'train ip_labels unknowns': train_ip_labels_unknowns,\n                'compet_test ip_labels unknowns':compet_test_ip_labels_unknowns}\nmy_own_metrics","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:33:58.935975Z","iopub.execute_input":"2023-06-11T12:33:58.938397Z","iopub.status.idle":"2023-06-11T12:33:58.946929Z","shell.execute_reply.started":"2023-06-11T12:33:58.938341Z","shell.execute_reply":"2023-06-11T12:33:58.946193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_cols = ['day', 'hour', 'minute', 'second', \n                'ip_labels', 'app_labels', 'device_labels',\n                'os_labels', 'channel_labels']\n\nvalid_fraction = 0.1\nclicks_srt = df_train.sort_values('click_time')\nvalid_rows = int(len(clicks_srt) * valid_fraction)\ntrain = clicks_srt[:-valid_rows * 2]\n# valid size == test size, last two sections of the data\nvalid = clicks_srt[-valid_rows * 2:-valid_rows]\ntest = clicks_srt[-valid_rows:]","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:33:58.948301Z","iopub.execute_input":"2023-06-11T12:33:58.948752Z","iopub.status.idle":"2023-06-11T12:33:58.999439Z","shell.execute_reply.started":"2023-06-11T12:33:58.948605Z","shell.execute_reply":"2023-06-11T12:33:58.998740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import lightgbm as lgb\n\ndtrain = lgb.Dataset(train[feature_cols], label=train['is_attributed'])\ndvalid = lgb.Dataset(valid[feature_cols], label=valid['is_attributed'])\ndtest = lgb.Dataset(test[feature_cols], label=test['is_attributed'])\n\nparam = {'num_leaves': 64, 'objective': 'binary'}\nparam['metric'] = 'auc'\nnum_round = 1000","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:33:59.000728Z","iopub.execute_input":"2023-06-11T12:33:59.001050Z","iopub.status.idle":"2023-06-11T12:33:59.040415Z","shell.execute_reply.started":"2023-06-11T12:33:59.000985Z","shell.execute_reply":"2023-06-11T12:33:59.039748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Record eval results for plotting\nvalidation_metrics = {}  \n\nbst = lgb.train(param, \n                dtrain, \n                num_round, \n                valid_sets=[dvalid], \n                early_stopping_rounds=10,\n                evals_result=validation_metrics,\n                verbose_eval=10)","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:33:59.041785Z","iopub.execute_input":"2023-06-11T12:33:59.042099Z","iopub.status.idle":"2023-06-11T12:33:59.277954Z","shell.execute_reply.started":"2023-06-11T12:33:59.042034Z","shell.execute_reply":"2023-06-11T12:33:59.277279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nplt.rcParams[\"figure.figsize\"] = [10,5]\n\nax = lgb.plot_metric(validation_metrics, metric='auc');\n#plt.show();","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:33:59.281463Z","iopub.execute_input":"2023-06-11T12:33:59.283276Z","iopub.status.idle":"2023-06-11T12:33:59.424951Z","shell.execute_reply.started":"2023-06-11T12:33:59.283207Z","shell.execute_reply":"2023-06-11T12:33:59.424195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Plot feature importances...')\nax = lgb.plot_importance(bst, max_num_features=15)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:33:59.426176Z","iopub.execute_input":"2023-06-11T12:33:59.426661Z","iopub.status.idle":"2023-06-11T12:33:59.596684Z","shell.execute_reply.started":"2023-06-11T12:33:59.426607Z","shell.execute_reply":"2023-06-11T12:33:59.595955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tree_index = 0\nprint('Plot '+str(tree_index)+'th tree...')  # one tree use categorical feature to split\nax = lgb.plot_tree(bst, tree_index=tree_index, figsize=(64, 36), show_info=['split_gain'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:33:59.597976Z","iopub.execute_input":"2023-06-11T12:33:59.598456Z","iopub.status.idle":"2023-06-11T12:34:03.784996Z","shell.execute_reply.started":"2023-06-11T12:33:59.598403Z","shell.execute_reply":"2023-06-11T12:34:03.784194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print('Plot'+str(tree_index)+'th tree with graphviz...')\n# graph = lgb.create_tree_digraph(bst, tree_index=tree_index, name='Tree'+str(tree_index))\n# # graph.render(view=True)","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:35:20.402157Z","iopub.execute_input":"2023-06-11T12:35:20.402519Z","iopub.status.idle":"2023-06-11T12:35:20.406590Z","shell.execute_reply.started":"2023-06-11T12:35:20.402464Z","shell.execute_reply":"2023-06-11T12:35:20.405725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn import metrics\n\nypred = bst.predict(test[feature_cols])\nscore = metrics.roc_auc_score(test['is_attributed'], ypred)\nprint(f\"Test score: {score}\")","metadata":{"execution":{"iopub.status.busy":"2023-06-11T12:35:32.723058Z","iopub.execute_input":"2023-06-11T12:35:32.723387Z","iopub.status.idle":"2023-06-11T12:35:32.752035Z","shell.execute_reply.started":"2023-06-11T12:35:32.723328Z","shell.execute_reply":"2023-06-11T12:35:32.751269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The end","metadata":{}}]}