{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img width=\"100%\" src=\"https://images.creativemarket.com/0.1.0/ps/2930173/910/607/m1/fpnw/wm0/credit-card-design-by-jan-baca-.jpg?1499199821&s=9dbd168ee67bf12b45e533800f1f9aed&fmt=webp?1499199821&s=772da0dd5276e72af0dc92f36c2f8149?auto=compress&cs=tinysrgb&w=200&h=200&dpr=1\">\n<figcaption>Pop–Art Diamond Bank Card Design</figcaption>","metadata":{"execution":{"iopub.status.busy":"2022-10-04T14:07:14.628014Z","iopub.execute_input":"2022-10-04T14:07:14.628486Z","iopub.status.idle":"2022-10-04T14:07:14.636875Z","shell.execute_reply.started":"2022-10-04T14:07:14.628447Z","shell.execute_reply":"2022-10-04T14:07:14.635234Z"}}},{"cell_type":"markdown","source":"# Comments\nThis Notebook is very similar to my previous notebook ([https://www.kaggle.com/code/schopenhacker75/lightgbm-catboost-0-8-end2end-study?scriptVersionId=106234992](https://www.kaggle.com/code/schopenhacker75/lightgbm-catboost-0-8-end2end-study?scriptVersionId=106234992) whose the feture engineering part is inspired from [Martin's gret notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977). In this notebook I used different feature engineering.\n\n> **Table of Contents:**\n> * [🚫 Introduction on Payment Default](#1)\n> * [🙏 Credits](#2)\n> * [🛠 Feature Engineering](#3)\n> * [🥢 Feature Selection](#4)\n> * [🤖 Model Training](#5)\n> * [🏅 Feature Importance](#6)\n> * [📈 Cross-Val Roc Curves](#7)\n> * [🤞 Predict Test data & Submission](#8)\n> ---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# <div style='display:fill;color:#2d3a41;background-color:#a3d3eb;padding:20px'>   🚫  <b> Introduction on Payment Default </b> </div>","metadata":{"execution":{"iopub.status.busy":"2022-08-15T05:13:08.754535Z","iopub.execute_input":"2022-08-15T05:13:08.754974Z","iopub.status.idle":"2022-08-15T05:13:08.762661Z","shell.execute_reply.started":"2022-08-15T05:13:08.754942Z","shell.execute_reply":"2022-08-15T05:13:08.760955Z"}}},{"cell_type":"markdown","source":"Even though credit cards presents many advantages such us avoiding carrying a bulky wallet in your pocket, tracking the spending behaviour, fraud detection... On the other hand the major downside for credit card usage is the increasing tendency to default on their payments. Aggressive marketing strategies can encourage credit card use beyond payment capacity, thus increasing the bearer’s credit risk and resulting in defaults and losses that might have not been properly anticipated","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n# <div style='display:fill;color:#2d3a41;background-color:#a3d3eb;padding:20px'>   🙏 <b> Credits </b> </div>","metadata":{}},{"cell_type":"markdown","source":"##### In this notebook we build and train a LightGBM model using [@raddar'dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) (discussion [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514) and my notebook explanation [here](https://www.kaggle.com/code/schopenhacker75/data-deanonymization) )\n\nThe feature engineering part  is very inspired from [this notbook](https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created) and the design aspect s from [my previous notebook](https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda/notebook) whose it self ispired from [this notebook](https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline/notebook). The LightGBM implementation is inspired from [jayjay's notebook](https://www.kaggle.com/code/jayjay75/wids2020-lgb-starter-adversarial-validation)","metadata":{}},{"cell_type":"code","source":"## ESSENTIALS\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport gc\nimport sys\nimport pickle\nimport glob\nimport pandas as pd\nfrom sklearn.preprocessing import LabelEncoder\n\npd.set_option('display.max_columns', None)\nimport random\nrandom.seed(75)\nfrom tqdm.notebook import tqdm_notebook \nfrom functools import partial, reduce\n\n### warnings setting\nimport sys\nimport warnings\nif not sys.warnoptions:\n    warnings.simplefilter(\"ignore\")\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning)\n\n#### model\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import roc_auc_score, roc_curve, auc\nimport catboost\nfrom catboost import Pool, CatBoostClassifier \nimport lightgbm as lgb\nimport joblib\nimport pickle\nfrom tqdm.notebook import tqdm_notebook \nimport uuid\n\n##### LOGGING Stettings #####\nimport logging\n# Create logger\nlogger = logging.getLogger()\nlogger.setLevel(logging.INFO)\n# Create STDERR handler\nhandler = logging.StreamHandler(sys.stderr)\n# Create formatter and add it to the handler\nformatter = logging.Formatter('%(asctime)s [%(levelname)s] %(name)s - %(message)s', datefmt='%Y-%m-%d %H:%M:%S',)\nhandler.setFormatter(formatter)\n# Set STDERR handler as the only handler \nlogger.handlers = [handler]\n\n#### plots\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.colors\n\nsns.set(rc={'axes.facecolor':'#f9ecec', 'figure.facecolor':'#f9ecec'})\n\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode\n\n### Plotly settings\ntheme_palette={\n    'base': '#a3d3eb',\n    'complementary':'#ebbba3',\n    'triadic' : '#eba3d3', \n    'backgound' : '#f6fbfd' \n}\n\ntemp=dict(layout=go.Layout(font=dict(family=\"Ubuntu\", size=14), \n                           height=600, \n                         legend=dict(#traceorder='reversed',\n                            orientation=\"v\",\n                            y=1.15,\n                            x=0.9),\n                    plot_bgcolor = theme_palette['backgound'],\n                      paper_bgcolor = theme_palette['backgound']))\n\n\nSAVED=True\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-05T07:47:01.058866Z","iopub.execute_input":"2022-10-05T07:47:01.059732Z","iopub.status.idle":"2022-10-05T07:47:04.379172Z","shell.execute_reply.started":"2022-10-05T07:47:01.059629Z","shell.execute_reply":"2022-10-05T07:47:04.377873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n# <div style='display:fill;color:#2d3a41;background-color:#a3d3eb;padding:20px'>   🛠 <b>  Feature Engineering </b> </div>\n\n<h4 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐    <b> a - Extract date based features </b> </h4>\n\nThe `S_2` is obviously a date type feature we'll extract the date related features such as the **month** and the **day of week**\n\n<h4 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐    <b> b - Row Rise aggregation based features </b> </h4>\n\nFor each group of variables (Delinquency, Spend, Payment, Balance, Risk variables) we apply agg functions `[mean, sum]` at each row, as well as **count of missing values** by row\n\n<h4 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐    <b> c - Column Rise aggregation based features</b> </h4>\n\nGroup by `custom_ID` and apply `['mean', 'std', 'min', 'max', 'last']` function for each group of variables\n\n<h4 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐    <b> c -  Difference based features</b> </h4>\n\nIt [has been shown](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977/notebook) that these features carry a powerful predictive signal:\n\n* the difference between the last transaction and the second of last\n* the difference between the last transaction and the mean","metadata":{}},{"cell_type":"markdown","source":"<h3 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    <b>-- Get Train Data --</b> </h3>","metadata":{}},{"cell_type":"code","source":"if not(SAVED):\n    train = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet')\n    print(f\"Shape = {train.shape}, number of customers = {train['customer_ID'].nunique()}\")\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:23.088148Z","iopub.execute_input":"2022-10-05T08:35:23.089743Z","iopub.status.idle":"2022-10-05T08:35:23.095751Z","shell.execute_reply.started":"2022-10-05T08:35:23.089703Z","shell.execute_reply":"2022-10-05T08:35:23.094879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    <b>-- Feature Engineering --</b> </h3>","metadata":{}},{"cell_type":"markdown","source":"Lets define features by category types:\n* **Delinquency variables** : features starting with **D_**\n* **Spend variables** : features starting with **S_**\n* **Payment variables** : features starting with **P_**\n* **Balance variables** : features starting with **B_**\n* **Risk variables** : features starting with **R_**\n\nWith the following categorical features = `['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']`","metadata":{}},{"cell_type":"code","source":"if not(SAVED):\n    ## Define some features by category\n    features = train.drop(['customer_ID', 'S_2'], axis = 1).columns.to_list()\n    cat_vars = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68', 'month', 'day_of_week' ]\n    num_vars = list(filter(lambda x : x not in cat_vars, features))\n\n    # devide nums vars by AmEx Categories\n    delequincy_vars = filter(lambda x:x.startswith('D') and x not in cat_vars, features)\n    spend_vars = filter(lambda x:(x.startswith('S')) and (x not in cat_vars),features)\n    payment_vars = filter(lambda x:x.startswith('P') and x not in cat_vars, features)\n    balance_vars = filter(lambda x:x.startswith('B') and x not in cat_vars, features)\n    risk_vars = filter(lambda x:x.startswith('R') and x not in cat_vars, features)\n\n    with open('features.pkl', 'wb') as f:\n        pickle.dump(features, f)\n\n    with open('cat_vars.pkl', 'wb') as f:\n        pickle.dump(cat_vars, f)\n\n    with open('num_vars.pkl', 'wb') as f:\n        pickle.dump(num_vars, f)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:35.793947Z","iopub.execute_input":"2022-10-05T08:35:35.794363Z","iopub.status.idle":"2022-10-05T08:35:35.807270Z","shell.execute_reply.started":"2022-10-05T08:35:35.794329Z","shell.execute_reply":"2022-10-05T08:35:35.805974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To address the **OOM issues**, I inplemented a batch generator to apply the data processing batch by batch then cancatenate it all. Please note that most of the used aggregation functions are **grouped by customer_ID**, so **all the statements of a given client must be grouped on a single batch.**","metadata":{}},{"cell_type":"code","source":"if not SAVED:\n    class BatchGenerator:\n        def __init__(self, df, batch_feature='customer_ID', keep_features=features, n_batchs=750):\n            self.df = df\n            self.batch_feature = batch_feature\n            self.keep_feature = list(set([batch_feature]+keep_features))\n            self.n_batchs = n_batchs\n\n        def __iter__(self):\n            unique_vals = self.df[self.batch_feature].unique()\n            batch_size = int(np.ceil(len(unique_vals) / self.n_batchs))\n            groups = self.df.groupby(self.batch_feature).groups\n            n_batchs= min(self.n_batchs, int(np.ceil(len(unique_vals) / batch_size)))\n            for i in range(n_batchs):\n                keys = unique_vals[i*batch_size:(i+1)*batch_size]\n                idx=[i for s in keys for i in groups[s] ]\n                if i == n_batchs-1:\n                    keys = unique_vals[(i+1)*batch_size:]\n                    idx = idx + [i for s in keys for i in groups[s] ]\n                yield self.df.loc[idx, self.keep_feature]\n\n\n\n    week_days ={1: 'Mon', 2: 'Tue', 3: 'Wen', 4: 'Thu', 5: 'Fri', 6: 'Sat', 7: 'Sun'}\n\n    def extract_date_vars(df, date_var='S_2', sort_by=['customer_ID','S_2'], week_days=week_days):\n        # change to datetime\n        df[date_var] = pd.to_datetime(df[date_var])\n        # sort by custoner ther by date \n        df = df.sort_values(by=sort_by)\n        # extract some date characteristics\n        # year has not a very\n        # month\n        df['month'] = df[date_var].dt.month\n        # day of week\n        df['day_of_week'] = df[date_var].apply(lambda x : x.isocalendar()[-1])\n        return df\n\n    group_names = [\"delequincy_vars\", \"spend_vars\", \"payment_vars\", \"balance_vars\", \"risk_vars\"]\n    # row rise aggregation\n    def row_rise_aggregation(df, \n                             group_vars=[delequincy_vars, spend_vars, payment_vars, balance_vars, risk_vars],\n                            group_names=group_names,\n                            save=True):\n        print('shape before row_rise_aggregation', df.shape )\n        for group_name, group_var in zip(group_names, group_vars):\n            df[group_name+'_sum'] = df[group_var].sum(axis=1)\n            df[group_name+'_mean'] = df[group_var].mean(axis=1)\n            df[group_name+'_missing'] = df.isnull().sum(axis=1)\n        print('shape after row_rise_aggregation', df.shape )\n        if save:    \n            df.reset_index(drop=False).to_feather(f\"row_agg_{str(uuid.uuid4())}.ftr\")\n            return df['customer_ID'].nunique()\n        return df\n\n    def column_rise_aggregation(df, num_vars=num_vars, cat_vars=cat_vars, save=True):\n        print('shape before column_rise_aggregation', df.shape )\n        group_names = filter(lambda x: '_vars' in x, df.columns)\n        num_agg = df.groupby(\"customer_ID\")[list(set(list(num_vars) + list(group_names)))].agg(['mean', 'std', 'min', 'max', 'last'])\n        num_agg.columns = ['_'.join(x) for x in num_agg.columns]\n\n        cat_agg = df.groupby(\"customer_ID\")[list(set(list(cat_vars)+['month', 'day_of_week']))].agg(['count', 'last', 'nunique', pd.Series.mode])\n        cat_agg.columns = ['_'.join(x) for x in cat_agg.columns]\n\n        mode_cols = filter(lambda x:x.endswith('_mode'), cat_agg.columns)\n        for col in mode_cols:\n            cat_agg[col] = cat_agg[col].apply(lambda x: random.choice(str(x).strip('[]').split()))\n        #concat the two dataframes\n        df = pd.concat([num_agg, cat_agg], axis=1)\n        del num_agg, cat_agg\n\n        gc.collect()\n        print('shape after column_rise_aggregation', df.shape )\n        if save:\n            df.reset_index(drop=False).to_feather(f\"col_agg_{str(uuid.uuid4())}.ftr\")\n            return len(df)#df['customer_ID'].nunique()\n        return df\n\n\n    # from https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n    def get_difference(df, num_features):\n        res = []\n        customer_ids = []\n        for customer_id, df in tqdm_notebook(df.groupby(['customer_ID'])):\n            # Get the differences\n            diff_df = df[num_features].diff(1).iloc[[-1]].values.astype(np.float32)\n            # Append to lists\n            res.append(diff_df)\n            customer_ids.append(customer_id)\n        # Concatenate\n        res = np.concatenate(res, axis = 0)\n        # Transform to dataframe\n        res = pd.DataFrame(res, columns = [col + '_diff1' for col in df[num_features].columns])\n        # Add customer id\n        res['customer_ID'] = customer_ids\n        print('final shape', res.shape)\n    #       df = df.merge(res, on='customer_ID', how='inner')\n    #       df.reset_index(drop=False).to_feather(f\"diff_{str(uuid.uuid4())}.ftr\")\n        return res#df['customer_ID'].nunique()\n\n    c=0\n    def save_partition(df, prefix='train'):\n        global c\n        df.reset_index(drop=True).to_feather(f'{prefix}_{c}.ftr')\n        c=c+1\n        return df['customer_ID'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:36.514626Z","iopub.execute_input":"2022-10-05T08:35:36.515244Z","iopub.status.idle":"2022-10-05T08:35:36.543412Z","shell.execute_reply.started":"2022-10-05T08:35:36.515206Z","shell.execute_reply":"2022-10-05T08:35:36.542173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N_BATCHS = 100\nc=0\nif not SAVED:\n    print(train.shape, train.customer_ID.nunique())\n    samples_df = BatchGenerator(train, batch_feature='customer_ID', keep_features=features+['S_2'], n_batchs=N_BATCHS)\n    processed_elements = sum(map(partial(save_partition, prefix='train'), tqdm_notebook(samples_df, total=N_BATCHS)))\n    \n    del train\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:36.744427Z","iopub.execute_input":"2022-10-05T08:35:36.744822Z","iopub.status.idle":"2022-10-05T08:35:36.752709Z","shell.execute_reply.started":"2022-10-05T08:35:36.744791Z","shell.execute_reply":"2022-10-05T08:35:36.751475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"processed = []\n\nif not SAVED:       \n    n_paths = 0\n    for path in tqdm_notebook(glob.glob('train_*.ftr')):\n        # apply on train\n        sample_df = pd.read_feather(path)\n        diff_df = get_difference(sample_df, num_features=num_vars)\n        sample_df = sample_df.merge(diff_df, on='customer_ID', how='inner')\n        \n     #   if sample_df.shape[1]<300:\n        sample_df = extract_date_vars(sample_df)\n        all_num_vars = num_vars + list(map(lambda x: x+'_diff1', num_vars))\n        sample_df = row_rise_aggregation(sample_df, save=False)\n        sample_df = column_rise_aggregation(sample_df, num_vars=all_num_vars, save=False)\n        \n        print(\"diff between last and mean transaction\")\n        for col in tqdm_notebook(num_vars):\n            try:\n                sample_df[f'{col}_last_mean_diff'] = sample_df[f'{col}_last'] - sample_df[f'{col}_mean']\n            except:\n                pass\n\n        sample_df = sample_df.reset_index()\n        \n        sample_df.reset_index(drop=True).to_feather(path)\n        print(\"save processed\", path)\n\n        \n    train = pd.concat(map(lambda sample_df: pd.read_feather(sample_df), tqdm_notebook(glob.glob('train_*.ftr'))))\n        \n        \n    ","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:37.469205Z","iopub.execute_input":"2022-10-05T08:35:37.469574Z","iopub.status.idle":"2022-10-05T08:35:37.479567Z","shell.execute_reply.started":"2022-10-05T08:35:37.469546Z","shell.execute_reply":"2022-10-05T08:35:37.478427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    <b>-- Join with Labels--</b> </h3>","metadata":{}},{"cell_type":"code","source":"if not(SAVED):\n    ## Left join with labels:\n    labels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')\n    print(labels.shape, labels['customer_ID'].nunique())\n    labels = labels.set_index('customer_ID')\n    train = train.set_index('customer_ID')\n    train['target'] = labels['target']\n    # del labels\n    gc.collect()\n    # save result\n    train.reset_index().to_feather('feat_eng_agg_train_with_diff.ftr')\n\n","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:38.237740Z","iopub.execute_input":"2022-10-05T08:35:38.238276Z","iopub.status.idle":"2022-10-05T08:35:38.244820Z","shell.execute_reply.started":"2022-10-05T08:35:38.238243Z","shell.execute_reply":"2022-10-05T08:35:38.243728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    -- Apply preprocess pipeline on Test Dataset -- </h2>\n","metadata":{}},{"cell_type":"code","source":"N_BATCHS = 300\nc=0\nif not SAVED:\n    test = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet')\n    n_cid = test.customer_ID.nunique()\n    print(test.shape, n_cid )\n    samples_df = BatchGenerator(test, batch_feature='customer_ID', keep_features=features+['S_2'], n_batchs=N_BATCHS)\n    processed_elements = sum(map(partial(save_partition, prefix='test'), tqdm_notebook(samples_df, total=N_BATCHS)))\n    \n    del test\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:39.048320Z","iopub.execute_input":"2022-10-05T08:35:39.048737Z","iopub.status.idle":"2022-10-05T08:35:39.055443Z","shell.execute_reply.started":"2022-10-05T08:35:39.048701Z","shell.execute_reply":"2022-10-05T08:35:39.054247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"processed = 0\n\nif not SAVED:       \n    for path in tqdm_notebook(glob.glob('test_*.ftr')):\n        # apply on train\n        sample_df = pd.read_feather(path)\n        diff_df = get_difference(sample_df, num_features=num_vars)\n        sample_df = sample_df.merge(diff_df, on='customer_ID', how='inner')\n        \n     #   if sample_df.shape[1]<300:\n        sample_df = extract_date_vars(sample_df)\n        all_num_vars = num_vars + list(map(lambda x: x+'_diff1', num_vars))\n        sample_df = row_rise_aggregation(sample_df, save=False)\n        sample_df = column_rise_aggregation(sample_df, num_vars=all_num_vars, save=False)\n        \n        print(\"diff between last and mean transaction\")\n        for col in tqdm_notebook(num_vars):\n            try:\n                sample_df[f'{col}_last_mean_diff'] = sample_df[f'{col}_last'] - sample_df[f'{col}_mean']\n            except:\n                pass\n\n        sample_df = sample_df.reset_index()\n        \n        sample_df.reset_index(drop=True).to_feather(path)\n        print(\"save processed\", path)\n        \n        processed += sample_df.customer_ID.nunique()\n        \n    test = pd.concat(map(lambda sample_df: pd.read_feather(sample_df), tqdm_notebook(glob.glob('test_*.ftr'))))\n    test.reset_index(drop=True).to_feather('feat_eng_agg_test_with_diff.ftr')","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:39.557034Z","iopub.execute_input":"2022-10-05T08:35:39.557825Z","iopub.status.idle":"2022-10-05T08:35:39.568665Z","shell.execute_reply.started":"2022-10-05T08:35:39.557780Z","shell.execute_reply":"2022-10-05T08:35:39.567515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# if saved extract the data directly the prepared data anad save time\nif SAVED:\n    train = pd.read_feather('../input/feat-eng-with-diff/feat_eng_agg_train_with_diff.ftr').set_index('customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:35:40.196804Z","iopub.execute_input":"2022-10-05T08:35:40.197327Z","iopub.status.idle":"2022-10-05T08:36:17.535376Z","shell.execute_reply.started":"2022-10-05T08:35:40.197290Z","shell.execute_reply":"2022-10-05T08:36:17.532312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def reduce_size(df):\n# Transform float64 columns to float32\n    print(\"reduce float data size\")\n    cols = list(df.dtypes[df.dtypes == 'float64'].index)\n    for col in tqdm_notebook(cols):\n        df[col] = df[col].astype(np.float32)\n    # Transform int64 columns to int32\n    print(\"reduce cat data size\")\n    cols = list(df.dtypes[df.dtypes == 'int64'].index)\n    for col in tqdm_notebook(cols):\n        df[col] = df[col].astype(np.int32)\n    return df\n        \ntrain = reduce_size(train)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n# <div style='display:fill;color:#2d3a41;background-color:#a3d3eb;padding:20px'>   🥢 <b> Feature Selection </b> </div>\nFor the feature selection part I used some **filter based techniques** because these methods are faster and less computationally expensive than other feature selection methods such as wrapper methods. Basically:\n\n* Drop Features with high missing values rate\n* Keep the TOP target correlated features\n\n\n<h2 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    -- Drop Features with high missing rate -- </h2>\n\nAll features having **more than 80% of missing values are ommited**","metadata":{}},{"cell_type":"code","source":"def missing_values_table(df):\n    # Total missing values by column\n    mis_val = df.isnull().sum()\n\n    # Percentage of missing values by column\n    mis_val_percent = 100 * df.isnull().sum() / len(df)\n\n    # build a table with the thw columns\n    mis_val_table = pd.concat([mis_val, mis_val_percent], axis=1)\n\n    # Rename the columns\n    mis_val_table_ren_columns = mis_val_table.rename(\n    columns = {0 : 'Missing Values', 1 : '% of Total Values'})\n\n    # Sort the table by percentage of missing descending\n    mis_val_table_ren_columns = mis_val_table_ren_columns[\n        mis_val_table_ren_columns.iloc[:,1] != 0].sort_values(\n    '% of Total Values', ascending=False).round(1)\n\n    # Print some summary information\n    print (\"Your selected dataframe has \" + str(df.shape[1]) + \" columns.\\n\"      \n        \"There are \" + str(mis_val_table_ren_columns.shape[0]) +\n          \" columns that have missing values.\")\n\n    # Return the dataframe with missing information\n    return mis_val_table_ren_columns\n\n# Missing values for training data\nmissing_values_train = missing_values_table(train)\n#cm = sns.color_palette('Set2', as_cmap=True)\n#missing_values_train[:20]#.style.background_gradient(cmap=cm)\nTHRESHOLD = 80\nprint(train.shape)\ndrop_cols = missing_values_train[missing_values_train['% of Total Values']>THRESHOLD].index.to_list()\nprint(f\"Drop {len(drop_cols)} features with more than {THRESHOLD}% of missing values\")\ntrain = train.drop(drop_cols, axis=1)\nprint(\"Training data shape after dropping highly missing values columns\", train.shape)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:36:29.636474Z","iopub.execute_input":"2022-10-05T08:36:29.637304Z","iopub.status.idle":"2022-10-05T08:36:29.690771Z","shell.execute_reply.started":"2022-10-05T08:36:29.637264Z","shell.execute_reply":"2022-10-05T08:36:29.689891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h2 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    -- Select top correlated features with the target-- </h2>\n\nWith this method we assume that high predictive features are **highly correlated with the target**\n\n","metadata":{}},{"cell_type":"code","source":"corr = train.corrwith(train['target'], axis=0)\ncorr = corr[corr.notna()].sort_values(key=abs, ascending=False)\nTHRESHOLD = 0.15\nCORR_SELECTION=True\nif CORR_SELECTION:\n    selected_feats = corr[corr.abs()>THRESHOLD].index\n    train = train[list(selected_feats)]\n    print(f\"Training data shape after dropping uncorrelated features\"\n          f\"(threshold Pearson correlation = {THRESHOLD})\", \n          train.shape)\n\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:36:35.462128Z","iopub.execute_input":"2022-10-05T08:36:35.462978Z","iopub.status.idle":"2022-10-05T08:36:36.123219Z","shell.execute_reply.started":"2022-10-05T08:36:35.462934Z","shell.execute_reply":"2022-10-05T08:36:36.121970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:36:36.125067Z","iopub.execute_input":"2022-10-05T08:36:36.125403Z","iopub.status.idle":"2022-10-05T08:36:36.440301Z","shell.execute_reply.started":"2022-10-05T08:36:36.125373Z","shell.execute_reply":"2022-10-05T08:36:36.438852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's display the top correlated features to the target ","metadata":{}},{"cell_type":"code","source":"sorted_corr = corr.sort_values(key=abs, ascending=False)[:11] # top but we have to  drop corr=1\n\n\npos = sorted_corr[(sorted_corr>0) & (sorted_corr<1)]\nneg = sorted_corr[(sorted_corr<0)].sort_values(ascending=False)\n\nfig = go.Figure()\nfig.add_trace(go.Bar(x=pos.index, y= pos.values,\n                     orientation='v',\n                     name='Positive',\n                     marker=dict(color=theme_palette['base'],line=dict(color=theme_palette['complementary'],width=0)),\n                     text = [\"%.2f\" %(round(v ,2) *100) + '%' for v in pos.values],\n                     textposition = 'outside',\n                     textfont_color = '#212a2f'))\n\nfig.add_trace(go.Bar(x=neg.index, y= neg.values,\n                     orientation='v',\n                     name='Negative',\n                     marker=dict(color=theme_palette['complementary'],line=dict(color=theme_palette['base'],width=0)),\n                     text = [\"%.2f\" %(round(v ,2) *100) + '%' for v in neg.values],\n                     textposition = 'outside',\n                     textfont_color = '#212a2f'))\n\n#theme_palette['2']\nfig.update_layout(template = temp,\n                  title={\n                      \"text\": \"<b>Top-10 Correlated Features with the Payment Default Feature</b> <BR />Pearson Values > 0.5<br> <br> \",\n                      \"x\":0.035,\n                      \"font_size\": 18,\n                      \n                  },\n                 plot_bgcolor = theme_palette['backgound'],\n                      paper_bgcolor = theme_palette['backgound'],\n                 legend=dict(\n                            y=1.15,\n                            x=0.88))","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:36:37.690760Z","iopub.execute_input":"2022-10-05T08:36:37.691371Z","iopub.status.idle":"2022-10-05T08:36:37.841222Z","shell.execute_reply.started":"2022-10-05T08:36:37.691329Z","shell.execute_reply":"2022-10-05T08:36:37.840127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    -- Store Interim Data-- </h2>\n\nDue to the redundant crashs we ought to store the used columns to use it later","metadata":{}},{"cell_type":"code","source":"if True:\n    types = train.dtypes\n    target_col = 'target'\n\n    cat_cols = list(types[types.apply(lambda x:not(str(x).startswith('float')))].index)\n    cat_cols = list(filter(lambda x:x!=target_col, cat_cols))\n    features = list(train.drop(target_col, axis=1).columns)\n    gc.collect()\n    print('len cat_col', len(cat_cols))\n    print('len features', len(features))\n    \n    with open('features.pkl', 'wb') as f:\n        pickle.dump(features, f)\n\n    with open('cat_cols.pkl', 'wb') as f:\n        pickle.dump(cat_cols, f)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:37:35.794765Z","iopub.execute_input":"2022-10-05T08:37:35.795296Z","iopub.status.idle":"2022-10-05T08:37:36.152158Z","shell.execute_reply.started":"2022-10-05T08:37:35.795255Z","shell.execute_reply":"2022-10-05T08:37:36.150909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n# <div style='display:fill;color:#2d3a41;background-color:#a3d3eb;padding:20px'>   🤖 <b> Model Training </b> </div>\nI implemented a gradient boosting model Wrapper **BaseModel** (inspired from [jayjay's notebook](https://www.kaggle.com/code/jayjay75/wids2020-lgb-starter-adversarial-validation)) in which I defined most of in common methods and attributes of gradient boosting based models, such as cv training, feature importances, prediction ... **LgbModel** and **CatBoost** would inherit from the BaseModel, Later I can add XGboost, hence we get the overall Gradient Boosting models comparison.\n<h2 style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉    -- Compettion Metric -- </h2>","metadata":{}},{"cell_type":"code","source":"def amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n    #https://www.kaggle.com/code/inversion/amex-competition-metric-python\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x == 0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n\n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x == 0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n    \n    y_pred=pd.DataFrame(data={'prediction':y_pred})\n    y_true=pd.DataFrame(data={'target':y_true.reset_index(drop=True)})\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:37:41.400002Z","iopub.execute_input":"2022-10-05T08:37:41.400483Z","iopub.status.idle":"2022-10-05T08:37:41.415323Z","shell.execute_reply.started":"2022-10-05T08:37:41.400446Z","shell.execute_reply":"2022-10-05T08:37:41.414009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# a wrapper class  that we can have the same ouput whatever the model we choose\ndef get_partial_pred_df(fold, y_val, pred):\n    df = pd.DataFrame()\n    df['y_true'] = y_val\n    df['y_pred'] = pred\n    df['fold'] = fold\n    return df[['fold','y_true', 'y_pred']]\n\nclass BaseModel:\n    RAND_SEED = 75\n    # TODO make a longer list of colors\n    COLORS =['#a3d3eb', '#dfa3eb', '#ebbba3', '#afeba3', '#ebe6a3']\n    def __init__(self, features, categoricals=[], n_splits=5, verbose=True, ps=None, target_col='target'):\n        self.features = features\n        self.n_splits = n_splits\n        self.categoricals = categoricals\n        self.target = target_col\n        self.cv_models = []\n        self.verbose = verbose\n        self.model=None\n        if not ps:\n            self.params = self.get_default_params()\n        else:\n            self.params = ps\n            \n        \n    def train_model(self, train_set, val_set, eval_metric=[]):\n        raise NotImplementedError\n        \n    def get_cv(self, train_df):\n        cv = StratifiedKFold(n_splits=self.n_splits, shuffle=True, random_state=self.RAND_SEED)\n        return cv.split(train_df, train_df[self.target])\n    \n    def get_default_params(self):\n        raise NotImplementedError\n        \n    def convert_dataset(self, x_train, y_train):\n        raise NotImplementedError\n        \n    def convert_x(self, x):\n        return x\n            \n    def fit_cv(self, train_df, save_cv=False):\n        self.oof_pred = np.zeros((len(train_df), ))\n    \n        cv = self.get_cv(train_df)\n        self.partial_oof_scores_=[]\n        self.cv_df=pd.DataFrame()\n        \n        for fold, (train_idx, val_idx) in enumerate(cv):\n            \n            \n            print(\"*-\"*100)\n            print(f\"*** FOLD == {fold} **\")\n            print(\"*-\"*100)\n            x_train, x_val = train_df[self.features].iloc[train_idx], train_df[self.features].iloc[val_idx]\n            y_train, y_val = train_df[self.target][train_idx], train_df[self.target][val_idx]\n            train_set = self.convert_dataset(x_train, y_train)\n            val_set = self.convert_dataset(x_val, y_val)\n            \n            model = self.train_model(train_set, val_set)\n            self.cv_models.append(model)\n            \n            conv_x_val = self.convert_x(x_val)\n            self.oof_pred[val_idx] = model.predict(conv_x_val).reshape(self.oof_pred[val_idx].shape)\n            \n            partial_oof_score = amex_metric(y_val, self.oof_pred[val_idx])\n            self.partial_oof_scores_.append(partial_oof_score)\n            print('Partial score of fold {} is: {}'.format(fold,  partial_oof_score))\n            if save_cv:                \n                # save model\n                joblib.dump(model, f'lgb_{fold}.pkl')\n                \n            partial_pred_df = get_partial_pred_df(fold, y_val, self.oof_pred[val_idx])\n            self.cv_df = pd.concat([self.cv_df, partial_pred_df])\n\n        self.oof_score_ = amex_metric(train_df[self.target], self.oof_pred) \n        if self.verbose:\n                print('Our oof score is: ', self.oof_score_)\n                \n    def fit(self, train):\n        x_train = train[self.features]\n        y_train = train[self.target]\n        train_set = self.convert_dataset(x_train, y_train)\n        self.model = self.train_model(train_set)\n        print(\"model trained in all training dataset\")\n        \n    def predict_cv(self, test_df):\n        y_pred = np.zeros((len(test_df), ))\n        x_test = self.convert_x(test_df[self.features])\n        for model in self.cv_models:\n            y_pred += model.predict(x_test).reshape(y_pred.shape) / self.n_splits\n        return y_pred\n    \n    def predict(self, test_df):\n        x_test = self.convert_x(test_df[self.features])\n        y_pred = self.model.predict(x_test).reshape((len(test_df), ))\n        return y_pred\n    \n    def get_cv_feature_importance(self):\n        raise NotImplementedError\n    \n    def plot_cv_roc(self): \n        fig=go.Figure()\n        fig.add_trace(go.Scatter(x=np.linspace(0,1,11), y=np.linspace(0,1,11), \n                                 name='Random Chance',mode='lines', showlegend=False,\n                                 line=dict(color=\"Black\", width=1, dash=\"dot\")))\n        for fold, sample_df in self.cv_df.groupby('fold'):\n            fpr, tpr, _ = roc_curve(sample_df.y_true, sample_df.y_pred)\n            roc_auc = auc(fpr,tpr)\n            fig.add_trace(go.Scatter(x=fpr, y=tpr, line=dict(color=self.COLORS[fold], width=3), \n                                     hovertemplate = 'True positive rate = %{y:.3f}<br>False positive rate = %{x:.3f}',\n                                     name='Fold {}: AUC = {:.3f}'.format(fold+1, roc_auc)))\n\n        fig.update_layout(template = temp,\n                      yaxis_automargin=True,\n                      height = 800,\n                      plot_bgcolor = theme_palette['backgound'],\n                      paper_bgcolor = theme_palette['backgound'],\n                      title={\n                          \"text\": \"<b>Cross-Validation ROC Curves</b> \",\n                          \"x\":0.045,\n                          \"font_size\": 20,                    \n                      },\n                      margin={'pad':5},\n                          xaxis_title='False Positive Rate (1 - Specificity)',\n                          yaxis_title='True Positive Rate (Sensitivity)',\n                          legend=dict(orientation='v', y=.07, x=1, xanchor=\"right\",\n                                      bordercolor=\"black\", borderwidth=.5)\n        )\n\n        return fig\n    \n    def plot_cv_feature_importance(self, top=20):\n        feat_imp = self.get_cv_feature_importance()\n        data = feat_imp[:top]\n        threshold = data.iloc[5]['importance']\n        fig = go.Figure()\n        fig.add_trace(go.Bar(y=data.index, x= data['importance'],\n                             orientation='h',\n                             width=[0.6]*len(data),\n                             marker=dict(color=(data['importance'] < threshold).astype('int'),\n                                         colorscale=[[0, theme_palette['complementary']], [1, theme_palette['base']]], ),\n\n                             text = [\"<b>%.2f\"%(round(v ,2))+'</b>' for v in data['importance']],\n                             textposition = 'inside',\n                             textfont_color = theme_palette['backgound']))\n\n        fig.update_layout(template = temp,\n                          yaxis_automargin=True,\n                          height = 800,\n                          plot_bgcolor = theme_palette['backgound'],\n                          paper_bgcolor = theme_palette['backgound'],\n                          yaxis=dict(autorange=\"reversed\"),\n                          title={\n                              \"text\": \"<b>Features Importances</b> \",#\"— Gain - Threshold = 5.2e+04<BR />the last payment 2 is farway in the top of list<br> <br> \",\n                              \"x\":0.045,\n                              \"font_size\": 20,                    \n                          },\n                          margin={'pad':5},\n        )\n        return fig\n\n","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:37:43.733723Z","iopub.execute_input":"2022-10-05T08:37:43.734202Z","iopub.status.idle":"2022-10-05T08:37:43.770152Z","shell.execute_reply.started":"2022-10-05T08:37:43.734159Z","shell.execute_reply":"2022-10-05T08:37:43.768924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#we choose to try a LightGbM using the Base_Model class\nclass LgbModel(BaseModel):\n    \n    def train_model(self, train_set, val_set=None):\n        verbosity = 100 if self.verbose else 0\n        valid_sets=[train_set]\n        if val_set:\n            valid_sets = [train_set, val_set]\n        return lgb.train(self.params, train_set, \n                         valid_sets=valid_sets, \n                         verbose_eval=verbosity)\n    \n    def convert_dataset(self, x_train, y_train):\n        train_set = lgb.Dataset(x_train, y_train, categorical_feature=self.categoricals)\n        return train_set\n    \n    \n    def get_default_params(self):\n        params = {'n_estimators':50000,\n                    'boosting_type': 'gbdt',\n                    'objective': 'binary',\n                    'metric': 'auc',\n                    'subsample': 0.75,\n                    'subsample_freq': 1,\n                    'learning_rate': 0.1,\n                    'feature_fraction': 0.9,\n                    'max_depth': 15,\n                    'lambda_l1': 1,  \n                    'lambda_l2': 1,\n                    'early_stopping_rounds': 100,\n                    #'is_unbalance' : True ,\n                    'scale_pos_weight' : 3\n                  \n                    }\n        return params\n    \n    def get_cv_feature_importance(self):\n        imp_df = pd.DataFrame(index=self.features)\n        imp_df['importance'] = 0\n        for model in self.cv_models:      \n            imp_df['importance'] = (imp_df['importance'] + pd.Series(model.feature_importance(), index=model.feature_name()))/self.n_splits\n        return imp_df.sort_values('importance', ascending=False)\n                                    \n    \n\n#we choose to try a LightGbM using the Base_Model class\nclass CatBoost(BaseModel):\n    def train_model(self, train_set, val_set=None):\n#        eval_set=[val_set] if val_set else None            \n        verbosity = 100 if self.verbose else 0   \n        return catboost.train(pool=train_set, params=self.params, \n                         eval_set=val_set, verbose_eval=verbosity)   \n    \n    def convert_dataset(self, x_train, y_train):\n        train_set = Pool(data=x_train, label=y_train, \n                         cat_features=self.categoricals)\n        return train_set\n    \n    \n    def get_default_params(self):\n        params = {'iterations' : 5000,\n                  'random_seed':self.RAND_SEED\n#                  'metric': ['auc','binary_logloss'],\n                    }\n        return params\n    \n    def get_cv_feature_importance(self):\n        imp_df = pd.DataFrame(index=self.features)\n        imp_df['importance'] = 0\n        for model in self.cv_models:      \n            imp_df['importance'] = (imp_df['importance'] + pd.Series(model.feature_importances_, index=model.feature_names_))/self.n_splits\n        return imp_df.sort_values('importance', ascending=False)\n                                    \n    ","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:37:44.558250Z","iopub.execute_input":"2022-10-05T08:37:44.559233Z","iopub.status.idle":"2022-10-05T08:37:44.574599Z","shell.execute_reply.started":"2022-10-05T08:37:44.559187Z","shell.execute_reply":"2022-10-05T08:37:44.573689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['target'] = train[\"target\"].astype('int')\nprint('Transform all String features to category.\\n')\nos.makedirs('label_encoders')\n\nfor usecol in tqdm_notebook(cat_cols):\n#    print(usecol)\n    train[usecol] = train[usecol].astype('str')\n#    test[usecol] = test[usecol].astype('str')\n\n    #Fit LabelEncoder\n    le = LabelEncoder().fit(\n            np.unique(train[usecol].unique().tolist()))#+\n#                      test[usecol].unique().tolist()))\n\n    #At the end 0 will be used for null values so we start at 1 \n    train[usecol] = le.transform(train[usecol])+1\n#    test[usecol]  = le.transform(test[usecol])+1\n\n    train[usecol] = train[usecol].replace(np.nan, 0).astype('int').astype('category')\n#    test[usecol]  = test[usecol].replace(np.nan, 0).astype('int').astype('category')\n\n    joblib.dump(le, f'label_encoders/{usecol}_label_encoder.pkl')\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:37:47.038465Z","iopub.execute_input":"2022-10-05T08:37:47.039593Z","iopub.status.idle":"2022-10-05T08:37:48.101654Z","shell.execute_reply.started":"2022-10-05T08:37:47.039543Z","shell.execute_reply":"2022-10-05T08:37:48.100331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#! zip -r label_encoders.zip label_encoders","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:37:51.675879Z","iopub.execute_input":"2022-10-05T08:37:51.677255Z","iopub.status.idle":"2022-10-05T08:37:52.928779Z","shell.execute_reply.started":"2022-10-05T08:37:51.677190Z","shell.execute_reply":"2022-10-05T08:37:52.927535Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'> 👉  -- LightGBM Training  -- </h2>\n","metadata":{}},{"cell_type":"code","source":"#https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977#kln-129\nparams = {\n        'objective': 'binary',\n        'metric': 'binary_logloss',\n        'boosting': 'dart',\n        'seed': 75,\n        'num_leaves': 100,\n        'learning_rate': 0.01,\n        'feature_fraction': 0.20,\n        'bagging_freq': 10,\n        'bagging_fraction': 0.50,\n        'n_jobs': -1,\n        'lambda_l2': 2,\n        'min_data_in_leaf': 40,\n        }\n\n\nmodel_path = \"../input/modelsv2/lgb_model_with_diff.sav\"\nif not model_path:\n    lgb_model = LgbModel(features=features, n_splits=5, categoricals=cat_cols, ps=params, verbose=None)\n    lgb_model.fit_cv(train, save_cv=True)\n    joblib.dump(lgb_model, \"lgb_model_with_diff.sav\")\nelse:\n    lgb_model = joblib.load(model_path)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:38:44.731363Z","iopub.execute_input":"2022-10-05T08:38:44.731846Z","iopub.status.idle":"2022-10-05T08:38:47.504439Z","shell.execute_reply.started":"2022-10-05T08:38:44.731806Z","shell.execute_reply":"2022-10-05T08:38:47.503245Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:38:57.397481Z","iopub.execute_input":"2022-10-05T08:38:57.398029Z","iopub.status.idle":"2022-10-05T08:38:57.773718Z","shell.execute_reply.started":"2022-10-05T08:38:57.397980Z","shell.execute_reply":"2022-10-05T08:38:57.772373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgb_model.oof_score_","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:38:57.775537Z","iopub.execute_input":"2022-10-05T08:38:57.775876Z","iopub.status.idle":"2022-10-05T08:38:57.788822Z","shell.execute_reply.started":"2022-10-05T08:38:57.775846Z","shell.execute_reply":"2022-10-05T08:38:57.787386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'> 👉  -- CatBoost Training  -- </h2>\n\nFor the Catboost that takes longer time to train, we will apply correlation based feature selection","metadata":{}},{"cell_type":"code","source":"# very heavy with all feature => we will select more \nprint(THRESHOLD)\ndeeper_feat_sel = True\nif deeper_feat_sel:\n    THRESHOLD += 0.05\n    print(\"new threshold\", THRESHOLD)\n    \ncorr = train.corrwith(train['target'], axis=0)\ncorr = corr[corr.notna()].sort_values(key=abs, ascending=False)\nselected_feats = corr[corr.abs()>THRESHOLD].index\ntrain = train[list(selected_feats)]\n\ntypes = train.dtypes\ntarget_col = 'target'\n\nselected_cat_cols = list(types[types.apply(lambda x:not(str(x).startswith('float')))].index)\ncb_cat_cols = list(filter(lambda x:x!=target_col, selected_cat_cols))\ncb_features = list(train.drop(target_col, axis=1).columns)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:41:28.902926Z","iopub.execute_input":"2022-10-05T08:41:28.903413Z","iopub.status.idle":"2022-10-05T08:41:29.486283Z","shell.execute_reply.started":"2022-10-05T08:41:28.903377Z","shell.execute_reply":"2022-10-05T08:41:29.485133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_path = \"../input/modelsv2/cb_model_with_diff.sav\"\nif not model_path:\n    cb_model = CatBoost(features=cb_features, categoricals=cb_cat_cols, n_splits=2)\n    cb_model.fit_cv(train, save_cv=True)\n    joblib.dump(cb_model, \"cb_model_with_diff.sav\")\nelse:\n    cb_model = joblib.load(model_path)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:42:52.025288Z","iopub.execute_input":"2022-10-05T08:42:52.025775Z","iopub.status.idle":"2022-10-05T08:44:04.438325Z","shell.execute_reply.started":"2022-10-05T08:42:52.025738Z","shell.execute_reply":"2022-10-05T08:44:04.437192Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n# <div style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  🏅  Feature Importance </div>\n\nLet's explore the feature importance of each model. One of the advantages of gradient boosting based models is that the feature importances can be directly extracted from the splitting gain of tree algorithm\n\n\n<h2 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐   <b>   LightGBM Feature Importance      </b>  </h2>\n","metadata":{}},{"cell_type":"code","source":"fig = lgb_model.plot_cv_feature_importance()\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:45:35.607977Z","iopub.execute_input":"2022-10-05T08:45:35.609125Z","iopub.status.idle":"2022-10-05T08:45:35.670661Z","shell.execute_reply.started":"2022-10-05T08:45:35.609046Z","shell.execute_reply":"2022-10-05T08:45:35.669403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(features)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:45:40.090183Z","iopub.execute_input":"2022-10-05T08:45:40.090656Z","iopub.status.idle":"2022-10-05T08:45:40.099348Z","shell.execute_reply.started":"2022-10-05T08:45:40.090622Z","shell.execute_reply":"2022-10-05T08:45:40.098088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐   <b>   CatBoost Feature Importance      </b>  </h2>","metadata":{}},{"cell_type":"code","source":"fig = cb_model.plot_cv_feature_importance()\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:45:41.225699Z","iopub.execute_input":"2022-10-05T08:45:41.226939Z","iopub.status.idle":"2022-10-05T08:45:41.254331Z","shell.execute_reply.started":"2022-10-05T08:45:41.226893Z","shell.execute_reply":"2022-10-05T08:45:41.253345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"></a>\n# <div style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  📈   CV Roc Curves </div>\nLets display Cross validation Roc curves for each model\n\n\n<h2 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐   <b>   LightGBM CV roc curves  </b>  </h2>\n","metadata":{}},{"cell_type":"code","source":"fig = lgb_model.plot_cv_roc()\n#fig = plot_roc(lgb_model)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:45:46.579612Z","iopub.execute_input":"2022-10-05T08:45:46.580062Z","iopub.status.idle":"2022-10-05T08:45:46.672745Z","shell.execute_reply.started":"2022-10-05T08:45:46.580029Z","shell.execute_reply":"2022-10-05T08:45:46.671871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐   <b>   CatBoost CV roc curves  </b>  </h2>\n","metadata":{}},{"cell_type":"code","source":"fig = cb_model.plot_cv_roc()\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:45:48.582830Z","iopub.execute_input":"2022-10-05T08:45:48.583649Z","iopub.status.idle":"2022-10-05T08:45:48.613952Z","shell.execute_reply.started":"2022-10-05T08:45:48.583602Z","shell.execute_reply":"2022-10-05T08:45:48.612813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:45:54.564578Z","iopub.execute_input":"2022-10-05T08:45:54.565046Z","iopub.status.idle":"2022-10-05T08:45:54.929258Z","shell.execute_reply.started":"2022-10-05T08:45:54.565011Z","shell.execute_reply":"2022-10-05T08:45:54.928336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8\"></a>\n# <div style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  👉 Infer (or Predict Test data)</div>\n\nGenerate the test predictions ","metadata":{}},{"cell_type":"code","source":"from functools import partial, reduce\n\nmodels_folder = '../input/modelsv2'\n\ndef get_sample(path):\n    test = pd.read_feather(path).set_index('customer_ID')\n\n    test = reduce_size(test)\n\n    test = test[lgb_model.features]\n    gc.collect()\n\n    for usecol in tqdm_notebook(cat_cols):\n        le = joblib.load(os.path.join(models_folder,f'{usecol}_label_encoder.pkl'))\n        test[usecol] = test[usecol].astype('str').apply(lambda x:x.split('.')[0])\n        test[usecol] = test[usecol].map(lambda s: '<unknown>' if s not in le.classes_ else s)\n        le.classes_ = np.append(le.classes_, '<unknown>')\n        #At the end 0 will be used for null values so we start at 1 \n        test[usecol]  = le.transform(test[usecol])+1\n        test[usecol] = test[usecol].replace(np.nan, 0).astype('int').astype('category')\n    return test\n\n\ndef get_batchs(df, batch_size, keep_features):\n    n_batchs = int(len(df)/batch_size)\n    cid = list(df.index)\n    for i in range(n_batchs):\n        idx = cid[i*batch_size:(i+1)*batch_size]\n        if i == n_batchs-1:\n            idx = idx + cid[(i+1)*batch_size:]\n        yield df.loc[idx, keep_features]\n        \ndef predict(sample_df, model):\n    y_pred = model.predict_cv(sample_df)\n    sub = pd.Series(y_pred, index=sample_df.index, name='prediction')\n#    sub.to_frame().to_csv(outfile)\n    return sub\n\n","metadata":{"execution":{"iopub.status.busy":"2022-10-05T08:48:09.928553Z","iopub.execute_input":"2022-10-05T08:48:09.929140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2022-09-28T10:06:29.265944Z","iopub.execute_input":"2022-09-28T10:06:29.266361Z","iopub.status.idle":"2022-09-28T10:06:29.276470Z","shell.execute_reply.started":"2022-09-28T10:06:29.266327Z","shell.execute_reply":"2022-09-28T10:06:29.275234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐   <b>   LightGBM submission </b>  </h3>","metadata":{}},{"cell_type":"code","source":"n_batchs=50\nsub_lg = pd.Series()\nfor path in glob.glob(\"../input/feat-eng-with-diff/feat_eng_agg_test_with_diff_*.ftr\"):\n    print(path)\n    test = get_sample(path)\n    samples_df = get_batchs(test,batch_size=n_batchs, keep_features=lgb_model.features)\n    sample_sub_lg = pd.concat(map(partial(predict, model=lgb_model), tqdm_notebook(samples_df)))\n    sub_lg = pd.concat([sub_lg, sample_sub_lg], ignore_index=True)\n    del test\n    gc.collect()\n    \nprint(len(sub_lg))\n\nsub_lg.to_frame().to_csv('submission_lgb_with_diff.csv')","metadata":{"execution":{"iopub.status.busy":"2022-09-28T10:06:35.030666Z","iopub.execute_input":"2022-09-28T10:06:35.031081Z","iopub.status.idle":"2022-09-28T10:11:46.682349Z","shell.execute_reply.started":"2022-09-28T10:06:35.031048Z","shell.execute_reply":"2022-09-28T10:11:46.681211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐   <b>   CatBoost Submission  </b>  </h3>","metadata":{}},{"cell_type":"code","source":"n_batchs=50\nsub_cb = pd.Series()\nfor path in glob.glob(\"../input/feat-eng-with-diff/feat_eng_agg_test_with_diff_*.ftr\"):\n    print(path)\n    test = get_sample(path)\n    samples_df = get_batchs(test,batch_size=n_batchs, keep_features=cb_model.features)\n    sample_sub_cb = pd.concat(map(partial(predict, model=cb_model), tqdm_notebook(samples_df)))\n    sub_cb = pd.concat([sub_cb, sample_sub_cb], ignore_index=True)\n    del test\n    gc.collect()\nprint(len(sub_cb))\n\nsub_cb.to_frame().to_csv('submission_cb_with_diff.csv')\n","metadata":{"execution":{"iopub.status.busy":"2022-09-13T20:45:34.279534Z","iopub.execute_input":"2022-09-13T20:45:34.280609Z","iopub.status.idle":"2022-09-13T20:46:14.834367Z","shell.execute_reply.started":"2022-09-13T20:45:34.280564Z","shell.execute_reply":"2022-09-13T20:46:14.833162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#2d3a41;background-color:#c8e5f3;padding:10px'>  🫐   <b>   Weighted submission  </b>  </h3>","metadata":{}},{"cell_type":"code","source":"# weighted submission\nweighted_sub = (sub_lg*lgb_model.oof_score_ + sub_cb*cb_model.oof_score_)/(lgb_model.oof_score_+cb_model.oof_score_)\nweighted_sub.to_frame().to_csv('weighted_sub.csv')","metadata":{"execution":{"iopub.status.busy":"2022-09-13T20:47:20.227198Z","iopub.execute_input":"2022-09-13T20:47:20.227708Z","iopub.status.idle":"2022-09-13T20:47:23.112647Z","shell.execute_reply.started":"2022-09-13T20:47:20.227668Z","shell.execute_reply":"2022-09-13T20:47:23.111501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8\"></a>\n# <div style='color:#2d3a41;background-color:#a3d3eb;padding:10px'>  📊 Predictions Distribution</div>","metadata":{}},{"cell_type":"code","source":"def get_dirtibution(sub,threshold=0.5):\n    target = (sub>threshold).astype(int)\n    target = target.value_counts(normalize=True)\n    target.rename(index={1:'Default',0:'Paid'},inplace=True)\n    fig=go.Figure()\n    fig.add_trace(go.Pie(labels=target.index, values=target*100,# hole=.45, \n                         showlegend=True,#sort=True, \n                         marker=dict(colors=list(theme_palette.values())),\n    #                     marker=dict(colors=color,line=dict(color=pal,width=2.5)),\n                         hovertemplate = \"%{label} Amex Acoounts: <b>%{value:.2f}</b>%<extra></extra>\"))\n    fig.update_layout(template=temp, \n                      title={\n                          \"text\":'<b>Default VS Paid Transaction</b><BR />Unballanced dataset',\n                          \"x\":0.035,\n                          \"font_size\": 20,\n\n                      },\n\n                      uniformtext_minsize=15,# width=700,\n    #                  height=800,\n                      margin={'t':150, 'l':5})\n    return fig","metadata":{"execution":{"iopub.status.busy":"2022-09-23T14:06:59.320029Z","iopub.execute_input":"2022-09-23T14:06:59.320537Z","iopub.status.idle":"2022-09-23T14:06:59.332546Z","shell.execute_reply.started":"2022-09-23T14:06:59.320499Z","shell.execute_reply":"2022-09-23T14:06:59.331226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = get_dirtibution(sub_lg)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-23T14:06:59.882337Z","iopub.execute_input":"2022-09-23T14:06:59.883846Z","iopub.status.idle":"2022-09-23T14:06:59.957919Z","shell.execute_reply.started":"2022-09-23T14:06:59.883782Z","shell.execute_reply":"2022-09-23T14:06:59.956301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = get_dirtibution(sub_cb)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-13T20:51:15.559244Z","iopub.execute_input":"2022-09-13T20:51:15.560189Z","iopub.status.idle":"2022-09-13T20:51:15.595226Z","shell.execute_reply.started":"2022-09-13T20:51:15.560139Z","shell.execute_reply":"2022-09-13T20:51:15.593393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style='color:#016CC9;text-align:center;font-size:100%'>*** WORK IN PROGRESS ***</div>\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}