{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style='background:#31317C; border:0; color: #FFFFFF'><center>💳 AMEX Default Prediction Top10% Solution 🎉 🥉</center></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# 💳 AMEX - Default Prediction Top10% Solution 🎉 🥉\n\n### **It's My First Medal in Competition Thanks for Kagglers!**\n\n**I'll study hard in the furture and share very useful information!!**\n\n**A lot of people shared Kernel, So I think, I got Bronze Medal.**\n\n**Thank you everyone. This medal is mine and yours!!**\n\n![banner](https://storage.googleapis.com/kaggle-competitions/kaggle/35332/logos/header.png?t=2022-03-23-01-05-50)\n\n\n<h1 style='background:#31317C; border:0; color:#FFFFFF'><center>TABLE OF CONTENTS</center></h1>\n\n[1. Import Libraries and Load Dataset](#1)\n    \n[2. Get Difference](#2)    \n\n[3. Processing Data](#3)        \n\n[4. AMEX - METRIC](#4)  \n\n[5. Configuration](#5)    \n\n[6. KFOLD - Training & Evalutate](#6)  \n\n[7. Ensemble - Rank Ensemble](#7)\n\n[8. Reference](#8)\n\n<h1 style='background:#31317C; border:0; color:#FFFFFF'><center>START</center></h1>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# **<span style=\"color:#4686C8;\">Import Libraries and Load Dataset</span>**\n\n### Whats RAPIDS?\n\n**The RAPIDS suite of open source software libraries and APIs gives you the ability to execute end-to-end data science and analytics pipelines entirely on GPUs**\n\n**for Detail : <a href = \"https://rapids.ai/index.html\">LINK</a>**","metadata":{}},{"cell_type":"code","source":"import gc\nimport os\nimport warnings\nimport joblib\nimport glob\n\nimport cudf\nimport cupy\nimport pandas as pd\nimport numpy as np\n\nfrom tqdm.auto import tqdm\nimport itertools\n\nimport scipy as sp\nfrom scipy.stats import rankdata\nfrom sklearn.model_selection import StratifiedKFold, train_test_split\nfrom sklearn.preprocessing import LabelEncoder\nimport lightgbm as lgb\nfrom itertools import combinations\n\nimport matplotlib.pyplot as plt\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:43.998167Z","iopub.execute_input":"2022-08-29T10:57:43.998958Z","iopub.status.idle":"2022-08-29T10:57:44.005444Z","shell.execute_reply.started":"2022-08-29T10:57:43.998914Z","shell.execute_reply":"2022-08-29T10:57:44.004441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n# **<span style=\"color:#4686C8;\">Get Difference</span>**\n**This Function add new features and improved**\n\n**I don't know detail of why this function improved performance, But it's useful**\n\n**Therefore I'm using it!**","metadata":{}},{"cell_type":"code","source":"def get_difference(data, num_features):\n    df1 = []\n    customer_ids = []\n    \n    for customer_id, df in tqdm(data.groupby(['customer_ID'])):\n        diff_df1 = df[num_features].diff(1).iloc[[-1]].values.astype(np.float32)\n        df1.append(diff_df1)\n        customer_ids.append(customer_id)\n        \n    df1 = np.concatenate(df1, axis = 0)\n    df1 = cudf.DataFrame(df1, columns = [col + '_diff1' for col in df[num_features].columns])\n    # Add customer id\n    df1['customer_ID'] = customer_ids\n    return df1","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:45.720643Z","iopub.execute_input":"2022-08-29T10:57:45.721004Z","iopub.status.idle":"2022-08-29T10:57:45.727872Z","shell.execute_reply.started":"2022-08-29T10:57:45.720974Z","shell.execute_reply":"2022-08-29T10:57:45.726634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n# **<span style=\"color:#4686C8;\">Processing Data</span>**\n\n**Since the amount of data is large and takes up quite a lot of memory,**\n\n**you should always use gc.collect or Del**","metadata":{}},{"cell_type":"code","source":"def process_parquet_data(df, istrain = True):\n    cat_features = [\"B_30\", \"B_38\", \"D_114\", \"D_116\", \"D_117\", \"D_120\", \"D_126\", \"D_63\", \"D_64\", \"D_66\", \"D_68\"]\n    features = df.drop(['customer_ID', 'S_2'], axis = 1).columns.to_list()\n    num_features = [col for col in features if col not in cat_features]\n    \n    df_num_agg = df.groupby(\"customer_ID\")[num_features].agg(['mean', 'std', 'min', 'max', 'last'])\n    df_num_agg.columns = ['_'.join(x) for x in train_num_agg.columns]\n    df_num_agg.rest_index(inplace = True)\n    \n    df_cat_agg = df.groupby(\"customer_ID\")[cat_features].agg(['count', 'last', 'nunique'])\n    df_cat_agg.columns = ['_'.join(x) for x in df_cat_agg.columns]\n    df_cat_agg.reset_index(inplace = True)\n    \n    cols = list(df_num_agg.dtypes[df_num_agg.dtypes == 'float64'].index)\n    for col in tqdm(cols):\n        df_num_agg[col] = df_num_agg[col].astype(np.float32)\n        \n    cols = list(df_cat_agg.dtypes[df_cat_agg.dtypes == 'int64'].index)  \n    for col in tqdm(cols):\n        df_cat_agg[col] = df_cat_agg[col].astype(np.int32)\n        \n    df_diff = get_difference(df, num_features)\n    \n    if istrain:\n        train_labels = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\n        df = df_num_agg.merge(df_cat_agg, how = 'inner', on = 'customer_ID').merge(df_diff, how = 'inner', on = 'customer_ID').merge(train_labels, how = 'inner', on = 'customer_ID')\n        del df_num_agg, df_cat_agg, train_labels, df_diff\n    else:\n        df = df_num_agg.merge(df_cat_agg, how = 'inner', on = 'customer_ID').merge(df_dfif, how = 'inner', on = 'customer_ID')\n        del df_num_agg, df_cat_agg, df_diff\n    \n    gc.collect()\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:48.066266Z","iopub.execute_input":"2022-08-29T10:57:48.067071Z","iopub.status.idle":"2022-08-29T10:57:48.078489Z","shell.execute_reply.started":"2022-08-29T10:57:48.067034Z","shell.execute_reply":"2022-08-29T10:57:48.077442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_process_parquet_data():\n    print(\"Staring train Feature Engineering...\")\n    train = cudf.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')\n    train = process_parquet_data(train, istrain = True)\n    \n    print(\"Staring test Feature Engineering...\")\n    test = cudf.read_parquet('../input/amex-data/integer-dtypes-parquet-format/test.parquet')\n    test = process_parquet_data(test, istrain = False)\n    \n    print(\"Saving Train & Test to Parquet...\")\n    train.to_parquet('data/train_fe.parquet')\n    test.to_parquet('data/test_fe.parquet')\n    \n    del train, test\n    gc.collect()\n\n# read_process_parquet_data()","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:49.249812Z","iopub.execute_input":"2022-08-29T10:57:49.250876Z","iopub.status.idle":"2022-08-29T10:57:49.257293Z","shell.execute_reply.started":"2022-08-29T10:57:49.250832Z","shell.execute_reply":"2022-08-29T10:57:49.256234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n# **<span style=\"color:#4686C8;\">AMEX - METRIC</span>**\n\n**You can see AMEX - METRIC in Competition main page.**","metadata":{}},{"cell_type":"code","source":"def amex_metric(y_true, y_pred):\n    labels = np.transpose(np.array([y_true, y_pred]))\n    labels = labels[labels[:, 1].argsort()[::-1]]\n    weights = np.where(labels[:,0]==0, 20, 1)\n    cut_vals = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n    gini = [0,0]\n    for i in [1,0]:\n        labels = np.transpose(np.array([y_true, y_pred]))\n        labels = labels[labels[:, i].argsort()[::-1]]\n        weight = np.where(labels[:,0]==0, 20, 1)\n        weight_random = np.cumsum(weight / np.sum(weight))\n        total_pos = np.sum(labels[:, 0] *  weight)\n        cum_pos_found = np.cumsum(labels[:, 0] * weight)\n        lorentz = cum_pos_found / total_pos\n        gini[i] = np.sum((lorentz - weight_random) * weight)\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:50.424430Z","iopub.execute_input":"2022-08-29T10:57:50.425503Z","iopub.status.idle":"2022-08-29T10:57:50.435379Z","shell.execute_reply.started":"2022-08-29T10:57:50.425450Z","shell.execute_reply":"2022-08-29T10:57:50.434086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def lgb_amex_metric(y_pred, y_true):\n    y_true = y_true.get_label()\n    return 'amex_metric', amex_metric(y_true, y_pred), True","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:51.118053Z","iopub.execute_input":"2022-08-29T10:57:51.118436Z","iopub.status.idle":"2022-08-29T10:57:51.123770Z","shell.execute_reply.started":"2022-08-29T10:57:51.118382Z","shell.execute_reply":"2022-08-29T10:57:51.122524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n# **<span style=\"color:#4686C8;\">Configuration</span>**\n\n**We have limited Memory. Therefore, We need to keep saving the datasets we worked on.**\n\n**MARTIN KOVACEVIC BUVINIC's kernel say seed blend(42, 52, 62) make LB boost nicely!!**","metadata":{}},{"cell_type":"code","source":"class CFG:\n    input_dir = 'data/'\n    seed = 42  # 52, 62\n    n_folds = 5\n    target = 'target'\n    boosting_type = 'dart'\n    metric = 'binary_logloss'\n\ndef seed_everything(seed):\n    random.seed(seed)\n    np.random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n\ndef read_data():\n    train.to_parquet(input_dir + 'train_fe.parquet')\n    test.to_parquet(input_dir + 'test_fe.parquet')\n    return train, test","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:52.195563Z","iopub.execute_input":"2022-08-29T10:57:52.195986Z","iopub.status.idle":"2022-08-29T10:57:52.203770Z","shell.execute_reply.started":"2022-08-29T10:57:52.195950Z","shell.execute_reply":"2022-08-29T10:57:52.202666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n# **<span style=\"color:#4686C8;\">KFOLD - Training & Evalutate</span>**\n\n**If you using Optuna. Don't using Regulation params  (reg_alpha, lambda_l1, lambda_l2, reg_lambda, min_split)**\n\n**Since we know most Models don't overfit for dataset, Optimizing Regularization featues is unlikely to important**\n\n**Reducing the parameter search makes a lot of sence to get result in a reasonable time period.!!**\n\n**I recommend to apply Regularization after optimizing remaining Parameters.**","metadata":{}},{"cell_type":"code","source":"def train_and_evaluate(train, test):\n    cat_features = [\"B_30\", \"B_38\", \"D_114\", \"D_116\", \"D_117\", \"D_120\", \"D_126\", \"D_63\", \"D_64\", \"D_66\", \"D_68\"]\n    cat_features = [f\"{cf}_last\" for cf in cat_features]\n    \n    for cat_col in cat_features:\n        encoder = LabelEncoder()\n        train[cat_col] = encoder.fit_transfrom(train[cat_col])\n        test[cat_col] = encoder.transform(test[cat_col])\n    \n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    for col in num_cols:\n        train[col + '_round2'] = train[col].round(2)\n        test[col + '_round2'] = test[col].round(2)\n    \n    num_cols = [col for col in train.columns if 'last' in col]\n    num_cols = [col[:-5] for col in num_cols if 'round' not in col]\n    \n    for col in num_cols:\n        try:\n            train[f'{col}_last_mean_diff'] = train[f'{col}_last'] - train[f'{col}_mean']\n            test[f'{col}_last_mean_diff'] = test[f'{col}_last'] - test[f'{col}_mean']\n        except:\n            pass\n        \n    num_cols = list(train.dtypes[(train.dtypes == 'float32') | (train.dtypes == 'float64')].index)\n    \n    for col in tqdm(num_cols):\n        train[col] = train[col].astype(np.float16)\n        test[col] = test[col].astype(np.float16)\n        \n    features = [col for col in train.columns if col not in ['customer_ID', CFG.target]]\n    \n    params = {\n        'objective': 'binary',\n        'metric': CFG.metric,\n        'boosting': CFG.boosting_type,\n        'seed': CFG.seed,\n        'num_leaves': 100,\n        'learning_rate': 0.01,\n        'feature_fraction': 0.20,\n        'bagging_freq': 10,\n        'bagging_fraction': 0.50,\n        'n_jobs': -1,\n        'lambda_l2': 2,\n        'min_data_in_leaf': 40,\n        }\n    \n    test_predictions = np.zeros(len(test))\n    oof_predictions = np.zeros(len(train))\n    kfold = StratifiedKFold(n_splits = CFG.n_folds, shuffle = True, random_state = CFG.seed)\n    \n    for fold, (trn_ind, val_ind) in enumerate(kfold.split(train, train[CFG.target])):\n        \n        print(' ')\n        print('-'*50)\n        print(f'Training fold {fold} with {len(features)} features...')\n        \n        x_train, x_val = train[features].iloc[trn_ind], train[features].iloc[val_ind]\n        y_train, y_val = train[CFG.target].iloc[trn_ind], train[CFG.target].iloc[val_ind]\n        \n        lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n        lgb_valid = lgb.Dataset(x_val, y_val, categorical_feature = cat_features)\n        \n        model = lgb.train(\n            params = params,\n            train_set = lgb_train,\n            num_boost_round = 10500,\n            valid_sets = [lgb_train, lgb_valid],\n            early_stopping_rounds = 1500,\n            verbose_eval = 500,\n            feval = lgb_amex_metric\n            )\n        \n        joblib.dump(model,  f'Models/lgbm_{CFG.boosting_type}_fold{fold}_seed{CFG.seed}.pkl')\n        val_pred = model.predict(x_val)\n        oof_predictions[val_ind] = val_pred\n        \n        test_pred = model.predict(test[features])\n        test_predictions += test_pred / CFG.n_folds\n        score = amex_metric(y_val, val_pred)\n        \n        print(f'Our fold {fold} CV score is {score}')\n        del x_train, x_val, y_train, y_val, lgb_train, lgb_valid\n        gc.collect()\n    \n    score = amex_metric(train[CFG.target], oof_predictions)\n    print(f'Our out of folds CV score is {score}')\n    oof_df = pd.DataFrame({'customer_ID': train['customer_ID'], 'target': train[CFG.target], \n                           'prediction': oof_predictions})\n    \n    oof_df.to_csv(f'OOF/oof_lgbm_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n    test_df = pd.DataFrame({'customer_ID': test['customer_ID'], \n                            'prediction': test_predictions})\n    test_df.to_csv(f'Predictions/test_lgbm_{CFG.boosting_type}_baseline_{CFG.n_folds}fold_seed{CFG.seed}.csv', index = False)\n\n# seed_everything(CFG.seed)\n# train, test = read_data()\n# train_and_evaluate(train, test)    ","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:53.638087Z","iopub.execute_input":"2022-08-29T10:57:53.638852Z","iopub.status.idle":"2022-08-29T10:57:53.658298Z","shell.execute_reply.started":"2022-08-29T10:57:53.638815Z","shell.execute_reply":"2022-08-29T10:57:53.657111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"></a>\n# **<span style=\"color:#4686C8;\">Ensemble - Rank Ensemble</span>**\n\n**If you know about detail Rank Ensemble, <a href=\"https://www.analyticsvidhya.com/blog/2021/03/basic-ensemble-technique-in-machine-learning/\">Click LINK</a>**\n\n**Rank Eensemble is great performance than Average Ensemble.**\n\n**But, I think when I using more Model, improve Private score**","metadata":{}},{"cell_type":"code","source":"def AMEX_Rank_Ensemble():\n    paths = [x for x in glob.glob('../input/*/*.csv') if 'amex-default-prediction' not in x if 'xgboost' not in x]\n    \n    df_all = [pd.read_csv(x) for x in paths]\n    df_sort_all = [x.sort_values(by='customer_ID') for x in df_all]\n    weights = [0.5, 0.9, 0.9, 0.5, 1, 0.8]\n    \n    for df in df_sort_all:\n        df['prediction'] = np.clip(df['prediction'], 0, 1)\n        \n    submit = pd.read_csv('../input/amex-default-prediction/sample_submission.csv')\n    submit['prediction'] = 0\n    \n    for df, weight in zip(df_sort_all, weights):\n        submit['prediction'] += (rankdata(df['prediction'])/df.shape[0]) * weight\n        \n    submit['prediction'] /= 5\n    submit.to_csv('submission_5.csv', index=None)\n\nAMEX_Rank_Ensemble()","metadata":{"execution":{"iopub.status.busy":"2022-08-29T10:57:54.853799Z","iopub.execute_input":"2022-08-29T10:57:54.854151Z","iopub.status.idle":"2022-08-29T10:58:07.580446Z","shell.execute_reply.started":"2022-08-29T10:57:54.854120Z","shell.execute_reply":"2022-08-29T10:58:07.579433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8\"></a>\n# **<span style=\"color:#4686C8;\">Reference</span>**\n\n- <a href = \"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\">AMEX data - integer dtypes - parquet format</a>\n- <a href = \"https://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds\">Amex Features: The best of both worlds</a>\n- <a href = \"https://www.kaggle.com/code/thedevastator/the-fine-art-of-hyperparameter-tuning/notebook\">The Fine Art of Hyperparameter Tuning</a>\n- <a href = \"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\">Amex LGBM Dart CV 0.7977</a>\n- <a href = \"https://www.kaggle.com/code/slowlearnermack/amex-lgbm-dart-cv-0-7963-improved\">Amex LGBM Dart CV 0.7963|Improved</a>\n- <a href = \"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\">AMEX data - integer dtypes - parquet format</a>\n- <a href = \"https://www.kaggle.com/code/songseungwon/xgboost-tutorial\">🇰🇷한국어🇰🇷 XGBoost Tutorial</a>","metadata":{}}]}