{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Reference:\n- [amex-lightgbm-quickstart](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart).\n- [time-to-explore-lce](https://towardsdatascience.com/random-forest-or-xgboost-it-is-time-to-explore-lce-2fed913eafb8).\n","metadata":{"execution":{"iopub.execute_input":"2022-05-27T15:44:46.297514Z","iopub.status.busy":"2022-05-27T15:44:46.296976Z","iopub.status.idle":"2022-05-27T15:44:46.3242Z","shell.execute_reply":"2022-05-27T15:44:46.322853Z","shell.execute_reply.started":"2022-05-27T15:44:46.297404Z"},"trusted":true}},{"cell_type":"markdown","source":"**Local Cascade Ensemble (LCE)** is a machine learning method that further enhances the prediction performance of the state-of-the-art Random Forest and XGBoost. LCE combines their strengths and adopts a complementary diversification approach to obtain a better generalizing predictor. Specifically, LCE is a hybrid ensemble method that combines an explicit boosting-bagging approach to handle the bias-variance trade-off faced by machine learning models and an implicit divide-and-conquer approach to individualize classifier errors on different parts of the training data. LCE has been evaluated on a public benchmark and published in the journal Data Mining and Knowledge Discovery.\nLCE package is compatible with scikit-learn; it passes the check_estimator. Therefore, it can interact with scikit-learn pipelines and model selection tools.","metadata":{}},{"cell_type":"markdown","source":"### Installation\nLCE is available in a Python package (Python ≥ 3.7). It can be installed using","metadata":{}},{"cell_type":"code","source":"#! conda install -c conda-forge lcensemble && yes\n# ! pip install lcensemble","metadata":{"execution":{"iopub.status.busy":"2022-05-27T19:05:47.451328Z","iopub.execute_input":"2022-05-27T19:05:47.451736Z","iopub.status.idle":"2022-05-27T19:05:47.455789Z","shell.execute_reply.started":"2022-05-27T19:05:47.451704Z","shell.execute_reply":"2022-05-27T19:05:47.454953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install ../input/lcensemble/numpy-1.21.5-cp37-cp37m-manylinux_2_12_x86_64.manylinux2010_x86_64.whl","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:11:49.020369Z","iopub.execute_input":"2022-05-30T02:11:49.021491Z","iopub.status.idle":"2022-05-30T02:12:18.27245Z","shell.execute_reply.started":"2022-05-30T02:11:49.021443Z","shell.execute_reply":"2022-05-30T02:12:18.271366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install ../input/lcensemble/xgboost-1.5.0-py3-none-manylinux2014_x86_64.whl","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:12:18.274313Z","iopub.execute_input":"2022-05-30T02:12:18.274665Z","iopub.status.idle":"2022-05-30T02:12:47.282395Z","shell.execute_reply.started":"2022-05-30T02:12:18.274625Z","shell.execute_reply":"2022-05-30T02:12:47.281138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install ../input/lcensemble/lcensemble-0.2.3/lcensemble-0.2.3.tar","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:12:47.284249Z","iopub.execute_input":"2022-05-30T02:12:47.284776Z","iopub.status.idle":"2022-05-30T02:13:17.861224Z","shell.execute_reply.started":"2022-05-30T02:12:47.284724Z","shell.execute_reply":"2022-05-30T02:13:17.859887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom matplotlib.ticker import MaxNLocator\nfrom matplotlib.colors import ListedColormap\nimport seaborn as sns\nfrom cycler import cycler\nfrom IPython.display import display\nimport datetime\nimport scipy.stats\nimport warnings\nimport gc\nfrom sklearn import preprocessing\n\n\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.calibration import CalibrationDisplay\nfrom lce import LCEClassifier\n\nplt.rcParams['axes.facecolor'] = '#0057b8' # blue\nplt.rcParams['axes.prop_cycle'] = cycler(color=['#ffd700'] +\n                                         plt.rcParams['axes.prop_cycle'].by_key()['color'][1:])\nplt.rcParams['text.color'] = 'w'","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:14:26.974423Z","iopub.execute_input":"2022-05-30T02:14:26.974881Z","iopub.status.idle":"2022-05-30T02:14:26.98298Z","shell.execute_reply.started":"2022-05-30T02:14:26.974841Z","shell.execute_reply":"2022-05-30T02:14:26.98206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### Reading and preprocessing the data\n\nWe read the data from @munumbutt's AMEX-Feather-Dataset. Then we reduce the amount of data by keeping only the most recent statement for every customer, as suggested by @inversion here.\n","metadata":{}},{"cell_type":"code","source":"# From https://www.kaggle.com/code/inversion/amex-competition-metric-python\ndef amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        \"\"\"Almost equal to 2 * auc - 1\"\"\"\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n    #print(f\"{g:.5f} {d:.5f}\")\n\n    return 0.5 * (g + d)\n","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:14:30.524435Z","iopub.execute_input":"2022-05-30T02:14:30.52485Z","iopub.status.idle":"2022-05-30T02:14:30.537849Z","shell.execute_reply.started":"2022-05-30T02:14:30.524814Z","shell.execute_reply":"2022-05-30T02:14:30.536693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nFPATH = '../input/amexfeather/'\ntrain = pd.read_feather(FPATH+'train_data.ftr')\ntest = pd.read_feather(FPATH+'test_data.ftr')","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:14:32.153614Z","iopub.execute_input":"2022-05-30T02:14:32.154492Z","iopub.status.idle":"2022-05-30T02:15:31.274456Z","shell.execute_reply.started":"2022-05-30T02:14:32.154441Z","shell.execute_reply":"2022-05-30T02:15:31.273025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train =  (train\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .drop(['S_2'], axis='columns'))\n\ntest =  (test\n            .groupby('customer_ID')\n            .tail(1)\n            .set_index('customer_ID', drop=True)\n            .sort_index()\n            .drop(['S_2'], axis='columns'))","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:15:31.277286Z","iopub.execute_input":"2022-05-30T02:15:31.277789Z","iopub.status.idle":"2022-05-30T02:15:40.589672Z","shell.execute_reply.started":"2022-05-30T02:15:31.277742Z","shell.execute_reply":"2022-05-30T02:15:40.588435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:15:40.591172Z","iopub.execute_input":"2022-05-30T02:15:40.591534Z","iopub.status.idle":"2022-05-30T02:15:40.776933Z","shell.execute_reply.started":"2022-05-30T02:15:40.591503Z","shell.execute_reply":"2022-05-30T02:15:40.775837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:18:18.705637Z","iopub.execute_input":"2022-05-30T02:18:18.706157Z","iopub.status.idle":"2022-05-30T02:18:18.740767Z","shell.execute_reply.started":"2022-05-30T02:18:18.706114Z","shell.execute_reply":"2022-05-30T02:18:18.739819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Select variables ","metadata":{}},{"cell_type":"code","source":"variables= ['D_42', 'D_49', 'B_10', 'B_12', 'D_63', 'D_64', 'D_66', 'D_68', 'D_69',\n       'D_73', 'D_76', 'R_7', 'B_26', 'R_9', 'R_14', 'B_29', 'B_30', 'D_88',\n       'B_31', 'S_23', 'D_106', 'R_26', 'B_38', 'D_108', 'D_110', 'D_111',\n       'B_39', 'B_40', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'B_42',\n       'D_132', 'D_134', 'D_135', 'D_136', 'D_137', 'D_138', 'target']\n\ntrain = train[variables]\n","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:19:35.581355Z","iopub.execute_input":"2022-05-30T02:19:35.581818Z","iopub.status.idle":"2022-05-30T02:19:35.627395Z","shell.execute_reply.started":"2022-05-30T02:19:35.581783Z","shell.execute_reply":"2022-05-30T02:19:35.626576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"variables= ['D_42', 'D_49', 'B_10', 'B_12', 'D_63', 'D_64', 'D_66', 'D_68', 'D_69',\n       'D_73', 'D_76', 'R_7', 'B_26', 'R_9', 'R_14', 'B_29', 'B_30', 'D_88',\n       'B_31', 'S_23', 'D_106', 'R_26', 'B_38', 'D_108', 'D_110', 'D_111',\n       'B_39', 'B_40', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'B_42',\n       'D_132', 'D_134', 'D_135', 'D_136', 'D_137', 'D_138']\n\ntest = test[variables]","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:20:07.066224Z","iopub.execute_input":"2022-05-30T02:20:07.067133Z","iopub.status.idle":"2022-05-30T02:20:07.199272Z","shell.execute_reply.started":"2022-05-30T02:20:07.067092Z","shell.execute_reply":"2022-05-30T02:20:07.198535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Encoding Categorical variables","metadata":{}},{"cell_type":"code","source":"le_D_63 = preprocessing.LabelEncoder()\nle_D_64 = preprocessing.LabelEncoder()\n\nle_D_63.fit(train.loc[:,\"D_63\"].to_list() + [\"NA\"])\nle_D_64.fit(train.loc[:,\"D_64\"].to_list() + [\"NA\"])","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:20:18.78095Z","iopub.execute_input":"2022-05-30T02:20:18.781499Z","iopub.status.idle":"2022-05-30T02:20:19.021341Z","shell.execute_reply.started":"2022-05-30T02:20:18.781465Z","shell.execute_reply":"2022-05-30T02:20:19.020537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.loc[:,\"D_63\"] = le_D_63.transform(train.loc[:,\"D_63\"])\ntrain.loc[:,\"D_64\"] = le_D_64.transform(train.loc[:,\"D_64\"])\n\n\ntest.loc[:,\"D_63\"] = le_D_63.transform(test.loc[:,\"D_63\"])\ntest.loc[:,\"D_64\"] = le_D_64.transform(test.loc[:,\"D_64\"])","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:20:23.210637Z","iopub.execute_input":"2022-05-30T02:20:23.211068Z","iopub.status.idle":"2022-05-30T02:20:23.797439Z","shell.execute_reply.started":"2022-05-30T02:20:23.211027Z","shell.execute_reply":"2022-05-30T02:20:23.796435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features = [f for f in train.columns if f != 'customer_ID' and f != 'target']","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:20:27.306734Z","iopub.execute_input":"2022-05-30T02:20:27.307163Z","iopub.status.idle":"2022-05-30T02:20:27.312083Z","shell.execute_reply.started":"2022-05-30T02:20:27.307129Z","shell.execute_reply":"2022-05-30T02:20:27.311183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### Cross-validation\n\nWe cross-validate with a five-fold StratifiedKFold.","metadata":{}},{"cell_type":"code","source":"#%%time\n# Cross-validation of the classifier\n\nINFERENCE = True\n\nfeatures = [f for f in train.columns if f != 'customer_ID' and f != 'target']\n\n\nprint(f\"{len(features)} features\")\nscore_list = []\ny_pred_list = []\nkf = StratifiedKFold(n_splits=5)\nfor fold, (idx_tr, idx_va) in enumerate(kf.split(train, train.target)):\n    start_time = datetime.datetime.now()\n    X_tr = train.iloc[idx_tr][features]\n    X_va = train.iloc[idx_va][features]\n    y_tr = train.iloc[idx_tr].target\n    y_va = train.iloc[idx_va].target\n    \n    # Train LCEClassifier with default parameters\n    clf = LCEClassifier(n_jobs=-1, random_state=123)\n    \n    with warnings.catch_warnings():\n        warnings.filterwarnings('ignore', category=UserWarning)\n        clf.fit(X_tr, y_tr)\n        \n    y_va_pred = clf.predict_proba(X_va)[:,1]\n    score = amex_metric(pd.DataFrame({'target': y_va.values}), pd.Series(y_va_pred, name='prediction'))\n    \n    print(f\"Fold {fold} | {str(datetime.datetime.now() - start_time)[-12:-7]} |\"\n          f\"                Score = {score:.5f}\")\n    score_list.append(score)\n\n    gc.collect()\n    \n    if INFERENCE:\n        y_pred_list.append(clf.predict_proba(test[features])[:,1])\n        \n    # break # we only want the first fold\n    \nprint(f\"OOF Score:                       {np.mean(score_list):.5f}\")\n\n","metadata":{"execution":{"iopub.status.busy":"2022-05-30T02:20:40.339638Z","iopub.execute_input":"2022-05-30T02:20:40.340305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### Calibration diagram\n\nThe calibration diagram shows how the model predicts the default probability of customers:\n","metadata":{}},{"cell_type":"code","source":"\nplt.figure(figsize=(12, 4))\nCalibrationDisplay.from_predictions(y_va, y_va_pred, n_bins=50, strategy='quantile', ax=plt.gca())\nplt.title('Probability calibration')\nplt.show()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Inference","metadata":{"execution":{"iopub.execute_input":"2022-05-27T15:02:04.198228Z","iopub.status.busy":"2022-05-27T15:02:04.197806Z","iopub.status.idle":"2022-05-27T15:02:04.203666Z","shell.execute_reply":"2022-05-27T15:02:04.202644Z","shell.execute_reply.started":"2022-05-27T15:02:04.198184Z"},"trusted":true}},{"cell_type":"code","source":"if INFERENCE:\n    sub = pd.DataFrame({'customer_ID': test.index,\n                        'prediction': np.mean(y_pred_list, axis=0)})\n    sub.to_csv('submission.csv', index=False)\n    sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### To be continued","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}