{
  "id": 328565,
  "title": "Let's catchup with all the learnings so far",
  "url": "/competitions/amex-default-prediction/discussion/328565",
  "author_name": "Mohsin",
  "post_date": "2022-06-01T19:32:56.532000",
  "votes": 143,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>Let's catchup with all the learnings so far!</h1>\n<p>I started 7 days after this competition launched and trying to catch up with all the learnings. If you are also just starting fresh, then this discussion might be useful to you. For those of you are already contributing can ignore this. The purpose is to save us all some time by not repeating the same EDA, or same baseline model building. </p>\n<p>This post is based on compilation of information from various posts. I tried to cite the actual contributors, please let me know if I miss any key contribution by mistake. Also please let me know if I am missing anything. I hope you also understand that I was not able to go over all the posts and all the notebooks. I hope you will find this useful! </p>\n<p>This post has this following sections. </p>\n<ol>\n<li>Memory optimization. </li>\n<li>Metrics calculation. </li>\n<li>Base line models. </li>\n<li>High Level stats. </li>\n</ol>\n<p><strong>Memory Optimization:</strong> <br>\nAs you already noticed that this project requires us to download and process almost 50GB of data. To process this large amount of data we either need huge RAM or we need to optimize it to analyze with existing hardware we have.  Here are some tricks I found in the discussions useful, </p>\n<ol>\n<li>Start with feather data created by Ruchi. <a href=\"https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\" target=\"_blank\">feather_data</a></li>\n<li>Encode customer ID column. It is ~600MB on RAM. Encoding this column will suppress it to ~40MB on RAM. </li>\n</ol>\n<p>Use following techniques for achieving it. </p>\n<ol>\n<li><p>From  <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">Chris Deotte</a><br>\n<code>train['customer_id'] =    train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')</code></p></li>\n<li><p>Process other columns to optimize data. Most of these are to lower the precision. Might not impact initial model building but might impact in later stages of competition. Thanks, Chris Deotte, for this nice discussion. <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">optimize_other_cols</a></p></li>\n</ol>\n<p><strong>Cross Validation:</strong> <br>\nSince this is a example of imbalanced data we can use stratifiedkfold for cross validation. </p>\n<p><strong>Metrics Calculation:</strong> <br>\nThere is a nice explanation of the metrics here: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">graphical explanation of evaluation metrics </a><br>\nFor implementation Rohan suggested multiple code snippet with their performance <a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\">here</a>. </p>\n<p><strong>Base line models:</strong> <br>\nAs of 06/01/2022 I have compiled the list of outstanding models with their public score. I am acknowledging authors by directly sharing their notebook URLs. Few outstanding models as of 6/1/2022: (Not in any order). </p>\n<table>\n<thead>\n<tr>\n<th>Technique/Tools</th>\n<th>Public Score</th>\n<th>URL</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Catboost</td>\n<td>0.783</td>\n<td><a href=\"https://www.kaggle.com/code/aninda/first-submission-using-catboost\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/aninda/first-submission-using-catboost\" target=\"_blank\">https://www.kaggle.com/code/aninda/first-submission-using-catboost</a> </td>\n</tr>\n<tr>\n<td>Catboost</td>\n<td>0.793</td>\n<td><a href=\"https://www.kaggle.com/code/huseyincot/amex-catboost-0-793\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/huseyincot/amex-catboost-0-793\" target=\"_blank\">https://www.kaggle.com/code/huseyincot/amex-catboost-0-793</a> </td>\n</tr>\n<tr>\n<td>Ensemble weighted average</td>\n<td>0.796</td>\n<td><a href=\"https://www.kaggle.com/code/beezus666/ensemble-weighted-average\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/beezus666/ensemble-weighted-average\" target=\"_blank\">https://www.kaggle.com/code/beezus666/ensemble-weighted-average</a> </td>\n</tr>\n<tr>\n<td>GradientBoosting (py-boost)</td>\n<td>0.791</td>\n<td><a href=\"https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline\" target=\"_blank\">https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline</a> </td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.786</td>\n<td><a href=\"https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng</a> </td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.792</td>\n<td><a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a></td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.765</td>\n<td><a href=\"https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model\" target=\"_blank\">https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model</a> </td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.783</td>\n<td><a href=\"https://www.kaggle.com/code/munumbutt/simple-lgbm-starter\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/munumbutt/simple-lgbm-starter\" target=\"_blank\">https://www.kaggle.com/code/munumbutt/simple-lgbm-starter</a> </td>\n</tr>\n<tr>\n<td>LightAutoML</td>\n<td>0.794</td>\n<td><a href=\"https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter\" target=\"_blank\">https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter</a> </td>\n</tr>\n<tr>\n<td>NN + skip connection + Dropout layer. Keras</td>\n<td>0.79</td>\n<td><a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training</a> </td>\n</tr>\n<tr>\n<td>NN using Keras</td>\n<td>0.783</td>\n<td><a href=\"https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter\" target=\"_blank\">https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter</a> </td>\n</tr>\n<tr>\n<td>Randdom Forest</td>\n<td>0.715</td>\n<td><a href=\"https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation\" target=\"_blank\">https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation</a> </td>\n</tr>\n<tr>\n<td>Random Forest with only two predictors ([['P_2', 'D_48']) and only for last statement.</td>\n<td>0.595</td>\n<td><a href=\"https://www.kaggle.com/code/bgmello/the-platinum-solution\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/bgmello/the-platinum-solution\" target=\"_blank\">https://www.kaggle.com/code/bgmello/the-platinum-solution</a> </td>\n</tr>\n<tr>\n<td>TensorFlow GRU</td>\n<td>0.79</td>\n<td><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790</a> </td>\n</tr>\n<tr>\n<td>TensorFlow Transformer</td>\n<td>0.789</td>\n<td><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790</a> </td>\n</tr>\n<tr>\n<td>Weight of Evidence (WOE)</td>\n<td>0.7</td>\n<td><a href=\"https://www.kaggle.com/code/lucasmorin/amex-woe-baseline\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/lucasmorin/amex-woe-baseline\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-woe-baseline</a> </td>\n</tr>\n<tr>\n<td>Weight of Evidence (WOE) + OptBinning</td>\n<td>0.756</td>\n<td><a href=\"https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model\" target=\"_blank\">https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model</a></td>\n</tr>\n<tr>\n<td>XGBoost</td>\n<td></td>\n<td><a href=\"https://www.kaggle.com/code/datajmcn/baseline-model-xgboost\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/datajmcn/baseline-model-xgboost\" target=\"_blank\">https://www.kaggle.com/code/datajmcn/baseline-model-xgboost</a> </td>\n</tr>\n<tr>\n<td>XGBoost with <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> data</td>\n<td>0.793</td>\n<td><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></td>\n</tr>\n</tbody>\n</table>\n<p><strong>High Level Stats:</strong> <br>\nFollowing are some useful numbers and findings I though might be helpful. </p>\n<table>\n<thead>\n<tr>\n<th>Key</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Total Unique Customers in Train</td>\n<td>458913</td>\n</tr>\n<tr>\n<td>Total Unique Customers in Test</td>\n<td>924621</td>\n</tr>\n<tr>\n<td>Non-Defaulted (Negative Class) Customers in Training Set</td>\n<td>340000 (74%), Down sampled at 5% from Population</td>\n</tr>\n<tr>\n<td>Defaulted Customers in Training Set</td>\n<td>119000 (26 %)</td>\n</tr>\n<tr>\n<td>Class imbalance in Training Set</td>\n<td>74%/26%</td>\n</tr>\n<tr>\n<td>Class imbalance in Population</td>\n<td>98%/2%</td>\n</tr>\n<tr>\n<td>Min(date), max(date) in Training Set</td>\n<td>2017-03-01, 2018-03-31</td>\n</tr>\n<tr>\n<td>Min(date), max(date) in Test Set</td>\n<td>2018-04-01, 2019-10-31</td>\n</tr>\n<tr>\n<td>Data Dimension in Training Set</td>\n<td>458913 Customers * ~13 statements * 190ish Columns</td>\n</tr>\n<tr>\n<td>Binary Features</td>\n<td>B_31 is always 0 or 1 and D_87 is always 1 or missing.</td>\n</tr>\n<tr>\n<td>Categorical Features</td>\n<td>Categorical features are defined as follows.<br><code>['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']</code></td>\n</tr>\n</tbody>\n</table>\n<p>There are few more columns can be added to the list of categorical (ordinal), e.g. <br>\ntrain_df['B_31'].unique()<br>\narray([1, 0], dtype=int64)<br>\ntrain_df['D_87'].unique()<br>\narray([nan,  1.], dtype=float16)</p>\n<p>Thank you all for sharing all those nice thoughts, ideas  in various discussions and notebooks. I will try to catchup more of your works and will try to summarize here. Looking forward to engaging more in coming days!</p>",
  "messages": [
    {
      "id": 1808414,
      "postDate": "2022-06-01T19:32:56.533Z",
      "content": "<h1>Let's catchup with all the learnings so far!</h1>\n<p>I started 7 days after this competition launched and trying to catch up with all the learnings. If you are also just starting fresh, then this discussion might be useful to you. For those of you are already contributing can ignore this. The purpose is to save us all some time by not repeating the same EDA, or same baseline model building. </p>\n<p>This post is based on compilation of information from various posts. I tried to cite the actual contributors, please let me know if I miss any key contribution by mistake. Also please let me know if I am missing anything. I hope you also understand that I was not able to go over all the posts and all the notebooks. I hope you will find this useful! </p>\n<p>This post has this following sections. </p>\n<ol>\n<li>Memory optimization. </li>\n<li>Metrics calculation. </li>\n<li>Base line models. </li>\n<li>High Level stats. </li>\n</ol>\n<p><strong>Memory Optimization:</strong> <br>\nAs you already noticed that this project requires us to download and process almost 50GB of data. To process this large amount of data we either need huge RAM or we need to optimize it to analyze with existing hardware we have.  Here are some tricks I found in the discussions useful, </p>\n<ol>\n<li>Start with feather data created by Ruchi. <a href=\"https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\" target=\"_blank\">feather_data</a></li>\n<li>Encode customer ID column. It is ~600MB on RAM. Encoding this column will suppress it to ~40MB on RAM. </li>\n</ol>\n<p>Use following techniques for achieving it. </p>\n<ol>\n<li><p>From  <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">Chris Deotte</a><br>\n<code>train['customer_id'] =    train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')</code></p></li>\n<li><p>Process other columns to optimize data. Most of these are to lower the precision. Might not impact initial model building but might impact in later stages of competition. Thanks, Chris Deotte, for this nice discussion. <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">optimize_other_cols</a></p></li>\n</ol>\n<p><strong>Cross Validation:</strong> <br>\nSince this is a example of imbalanced data we can use stratifiedkfold for cross validation. </p>\n<p><strong>Metrics Calculation:</strong> <br>\nThere is a nice explanation of the metrics here: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">graphical explanation of evaluation metrics </a><br>\nFor implementation Rohan suggested multiple code snippet with their performance <a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\">here</a>. </p>\n<p><strong>Base line models:</strong> <br>\nAs of 06/01/2022 I have compiled the list of outstanding models with their public score. I am acknowledging authors by directly sharing their notebook URLs. Few outstanding models as of 6/1/2022: (Not in any order). </p>\n<table>\n<thead>\n<tr>\n<th>Technique/Tools</th>\n<th>Public Score</th>\n<th>URL</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Catboost</td>\n<td>0.783</td>\n<td><a href=\"https://www.kaggle.com/code/aninda/first-submission-using-catboost\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/aninda/first-submission-using-catboost\" target=\"_blank\">https://www.kaggle.com/code/aninda/first-submission-using-catboost</a> </td>\n</tr>\n<tr>\n<td>Catboost</td>\n<td>0.793</td>\n<td><a href=\"https://www.kaggle.com/code/huseyincot/amex-catboost-0-793\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/huseyincot/amex-catboost-0-793\" target=\"_blank\">https://www.kaggle.com/code/huseyincot/amex-catboost-0-793</a> </td>\n</tr>\n<tr>\n<td>Ensemble weighted average</td>\n<td>0.796</td>\n<td><a href=\"https://www.kaggle.com/code/beezus666/ensemble-weighted-average\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/beezus666/ensemble-weighted-average\" target=\"_blank\">https://www.kaggle.com/code/beezus666/ensemble-weighted-average</a> </td>\n</tr>\n<tr>\n<td>GradientBoosting (py-boost)</td>\n<td>0.791</td>\n<td><a href=\"https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline\" target=\"_blank\">https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline</a> </td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.786</td>\n<td><a href=\"https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng</a> </td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.792</td>\n<td><a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a></td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.765</td>\n<td><a href=\"https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model\" target=\"_blank\">https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model</a> </td>\n</tr>\n<tr>\n<td>LGBM</td>\n<td>0.783</td>\n<td><a href=\"https://www.kaggle.com/code/munumbutt/simple-lgbm-starter\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/munumbutt/simple-lgbm-starter\" target=\"_blank\">https://www.kaggle.com/code/munumbutt/simple-lgbm-starter</a> </td>\n</tr>\n<tr>\n<td>LightAutoML</td>\n<td>0.794</td>\n<td><a href=\"https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter\" target=\"_blank\">https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter</a> </td>\n</tr>\n<tr>\n<td>NN + skip connection + Dropout layer. Keras</td>\n<td>0.79</td>\n<td><a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training</a> </td>\n</tr>\n<tr>\n<td>NN using Keras</td>\n<td>0.783</td>\n<td><a href=\"https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter\" target=\"_blank\">https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter</a> </td>\n</tr>\n<tr>\n<td>Randdom Forest</td>\n<td>0.715</td>\n<td><a href=\"https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation\" target=\"_blank\">https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation</a> </td>\n</tr>\n<tr>\n<td>Random Forest with only two predictors ([['P_2', 'D_48']) and only for last statement.</td>\n<td>0.595</td>\n<td><a href=\"https://www.kaggle.com/code/bgmello/the-platinum-solution\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/bgmello/the-platinum-solution\" target=\"_blank\">https://www.kaggle.com/code/bgmello/the-platinum-solution</a> </td>\n</tr>\n<tr>\n<td>TensorFlow GRU</td>\n<td>0.79</td>\n<td><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790</a> </td>\n</tr>\n<tr>\n<td>TensorFlow Transformer</td>\n<td>0.789</td>\n<td><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790</a> </td>\n</tr>\n<tr>\n<td>Weight of Evidence (WOE)</td>\n<td>0.7</td>\n<td><a href=\"https://www.kaggle.com/code/lucasmorin/amex-woe-baseline\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/lucasmorin/amex-woe-baseline\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-woe-baseline</a> </td>\n</tr>\n<tr>\n<td>Weight of Evidence (WOE) + OptBinning</td>\n<td>0.756</td>\n<td><a href=\"https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model\" target=\"_blank\">https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model</a></td>\n</tr>\n<tr>\n<td>XGBoost</td>\n<td></td>\n<td><a href=\"https://www.kaggle.com/code/datajmcn/baseline-model-xgboost\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/datajmcn/baseline-model-xgboost\" target=\"_blank\">https://www.kaggle.com/code/datajmcn/baseline-model-xgboost</a> </td>\n</tr>\n<tr>\n<td>XGBoost with <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> data</td>\n<td>0.793</td>\n<td><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></td>\n</tr>\n</tbody>\n</table>\n<p><strong>High Level Stats:</strong> <br>\nFollowing are some useful numbers and findings I though might be helpful. </p>\n<table>\n<thead>\n<tr>\n<th>Key</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Total Unique Customers in Train</td>\n<td>458913</td>\n</tr>\n<tr>\n<td>Total Unique Customers in Test</td>\n<td>924621</td>\n</tr>\n<tr>\n<td>Non-Defaulted (Negative Class) Customers in Training Set</td>\n<td>340000 (74%), Down sampled at 5% from Population</td>\n</tr>\n<tr>\n<td>Defaulted Customers in Training Set</td>\n<td>119000 (26 %)</td>\n</tr>\n<tr>\n<td>Class imbalance in Training Set</td>\n<td>74%/26%</td>\n</tr>\n<tr>\n<td>Class imbalance in Population</td>\n<td>98%/2%</td>\n</tr>\n<tr>\n<td>Min(date), max(date) in Training Set</td>\n<td>2017-03-01, 2018-03-31</td>\n</tr>\n<tr>\n<td>Min(date), max(date) in Test Set</td>\n<td>2018-04-01, 2019-10-31</td>\n</tr>\n<tr>\n<td>Data Dimension in Training Set</td>\n<td>458913 Customers * ~13 statements * 190ish Columns</td>\n</tr>\n<tr>\n<td>Binary Features</td>\n<td>B_31 is always 0 or 1 and D_87 is always 1 or missing.</td>\n</tr>\n<tr>\n<td>Categorical Features</td>\n<td>Categorical features are defined as follows.<br><code>['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']</code></td>\n</tr>\n</tbody>\n</table>\n<p>There are few more columns can be added to the list of categorical (ordinal), e.g. <br>\ntrain_df['B_31'].unique()<br>\narray([1, 0], dtype=int64)<br>\ntrain_df['D_87'].unique()<br>\narray([nan,  1.], dtype=float16)</p>\n<p>Thank you all for sharing all those nice thoughts, ideas  in various discussions and notebooks. I will try to catchup more of your works and will try to summarize here. Looking forward to engaging more in coming days!</p>",
      "rawMarkdown": "# Let's catchup with all the learnings so far!\n\nI started 7 days after this competition launched and trying to catch up with all the learnings. If you are also just starting fresh, then this discussion might be useful to you. For those of you are already contributing can ignore this. The purpose is to save us all some time by not repeating the same EDA, or same baseline model building. \n\nThis post is based on compilation of information from various posts. I tried to cite the actual contributors, please let me know if I miss any key contribution by mistake. Also please let me know if I am missing anything. I hope you also understand that I was not able to go over all the posts and all the notebooks. I hope you will find this useful! \n\nThis post has this following sections. \n\n 1. Memory optimization. \n 2. Metrics calculation. \n 3. Base line models. \n 4. High Level stats. \n\n**Memory Optimization:** \nAs you already noticed that this project requires us to download and process almost 50GB of data. To process this large amount of data we either need huge RAM or we need to optimize it to analyze with existing hardware we have.  Here are some tricks I found in the discussions useful, \n\n1.\tStart with feather data created by Ruchi. [feather_data](https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction)\n2.\tEncode customer ID column. It is ~600MB on RAM. Encoding this column will suppress it to ~40MB on RAM. \n\nUse following techniques for achieving it. \n\n1. From  [Chris Deotte](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635 )\n`train['customer_id'] =    train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')`\n\n2. Process other columns to optimize data. Most of these are to lower the precision. Might not impact initial model building but might impact in later stages of competition. Thanks, Chris Deotte, for this nice discussion. [optimize_other_cols](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054)\n\n**Cross Validation:** \nSince this is a example of imbalanced data we can use stratifiedkfold for cross validation. \n\n**Metrics Calculation:** \nThere is a nice explanation of the metrics here: [graphical explanation of evaluation metrics ](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464)\nFor implementation Rohan suggested multiple code snippet with their performance [here](https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations). \n\n**Base line models:** \nAs of 06/01/2022 I have compiled the list of outstanding models with their public score. I am acknowledging authors by directly sharing their notebook URLs. Few outstanding models as of 6/1/2022: (Not in any order). \n\n| Technique/Tools                                                                             | Public Score | URL                                                                                                                                                                     |\n| ------------------------------------------------------------------------------------------- | ------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Catboost                                                                                    | 0.783        | [https://www.kaggle.com/code/aninda/first-submission-using-catboost ](https://www.kaggle.com/code/aninda/first-submission-using-catboost)                               |\n| Catboost                                                                                    | 0.793        | [https://www.kaggle.com/code/huseyincot/amex-catboost-0-793 ](https://www.kaggle.com/code/huseyincot/amex-catboost-0-793)                                               |\n| Ensemble weighted average                                                                   | 0.796        | [https://www.kaggle.com/code/beezus666/ensemble-weighted-average ](https://www.kaggle.com/code/beezus666/ensemble-weighted-average)                                     |\n| GradientBoosting (py-boost)                                                                 | 0.791        | [https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline ](https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline)                       |\n| LGBM                                                                                        | 0.786        | [https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng ](https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng)                                         |\n| LGBM                                                                                        | 0.792        | [https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart)                                            |\n| LGBM                                                                                        | 0.765        | [https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model ](https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model)                   |\n| LGBM                                                                                        | 0.783        | [https://www.kaggle.com/code/munumbutt/simple-lgbm-starter ](https://www.kaggle.com/code/munumbutt/simple-lgbm-starter)                                                 |\n| LightAutoML                                                                                 | 0.794        | [https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter ](https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter)                                   |\n| NN + skip connection + Dropout layer. Keras                                                 | 0.79         | [https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training ](https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training)                           |\n| NN using Keras                                                                              | 0.783        | [https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter ](https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter)                   |\n| Randdom Forest                                                                              | 0.715        | [https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation ](https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation) |\n| Random Forest with only two predictors (\\[\\['P\\_2', 'D\\_48'\\]) and only for last statement. | 0.595        | [https://www.kaggle.com/code/bgmello/the-platinum-solution ](https://www.kaggle.com/code/bgmello/the-platinum-solution)                                                 |\n| TensorFlow GRU                                                                              | 0.79         | [https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790 ](https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790)                                   |\n| TensorFlow Transformer                                                                      | 0.789        | [https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790 ](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790)                                   |\n| Weight of Evidence (WOE)                                                                    | 0.7          | [https://www.kaggle.com/code/lucasmorin/amex-woe-baseline ](https://www.kaggle.com/code/lucasmorin/amex-woe-baseline)                                                   |\n| Weight of Evidence (WOE) + OptBinning                                                       | 0.756        | https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model                                                                                                     |\n| XGBoost                                                                                     |              | [https://www.kaggle.com/code/datajmcn/baseline-model-xgboost ](https://www.kaggle.com/code/datajmcn/baseline-model-xgboost)                                             |\n| XGBoost with @raddar data                                                                   | 0.793        | https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793                                                                                                               |                                                     \n\n\n**High Level Stats:** \nFollowing are some useful numbers and findings I though might be helpful. \n\n| Key                                                      | Value                                                                                                                                                                                                                    |\n| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| Total Unique Customers in Train                          | 458913                                                                                                                                                                                                                   |\n| Total Unique Customers in Test                           | 924621                                                                                                                                                                                                                   |\n| Non-Defaulted (Negative Class) Customers in Training Set | 340000 (74%), Down sampled at 5% from Population                                                                                                                                                                         |\n| Defaulted Customers in Training Set                      | 119000 (26 %)                                                                                                                                                                                                            |\n| Class imbalance in Training Set                          | 74%/26%                                                                                                                                                                                                                  |\n| Class imbalance in Population                            | 98%/2%                                                                                                                                                                                                                   |\n| Min(date), max(date) in Training Set                     | 2017-03-01, 2018-03-31                                                                                                                                                                                                   |\n| Min(date), max(date) in Test Set                         | 2018-04-01, 2019-10-31                                                                                                                                                                                                   |\n| Data Dimension in Training Set                           | 458913 Customers \\* ~13 statements \\* 190ish Columns                                                                                                                                                                     |\n| Binary Features                                          | B\\_31 is always 0 or 1 and D\\_87 is always 1 or missing.                                                                                                                                                            |\n| Categorical Features                                     | Categorical features are defined as follows.<br>`['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']`                                                                          |\n\nThere are few more columns can be added to the list of categorical (ordinal), e.g. \ntrain_df['B_31'].unique()\narray([1, 0], dtype=int64)\ntrain_df['D_87'].unique()\narray([nan,  1.], dtype=float16)\n\nThank you all for sharing all those nice thoughts, ideas  in various discussions and notebooks. I will try to catchup more of your works and will try to summarize here. Looking forward to engaging more in coming days!",
      "votes": 139
    },
    {
      "id": 1808489,
      "postDate": "2022-06-01T21:39:39.053Z",
      "content": "<p>train_df['B_31'].unique()<br>\narray([1, 0], dtype=int64)<br>\ntrain_df['D_87'].unique()<br>\narray([nan,  1.], dtype=float16)</p>\n<p>Looks like more column can be treated as int8. </p>",
      "rawMarkdown": "train_df['B_31'].unique()\narray([1, 0], dtype=int64)\ntrain_df['D_87'].unique()\narray([nan,  1.], dtype=float16)\n\nLooks like more column can be treated as int8. ",
      "votes": 1
    },
    {
      "id": 1808464,
      "postDate": "2022-06-01T20:56:49.007Z",
      "content": "<p>Nice one, thanks for compiling all the info!</p>",
      "rawMarkdown": "Nice one, thanks for compiling all the info!",
      "votes": 1
    },
    {
      "id": 1911767,
      "postDate": "2022-08-24T09:49:33.563Z",
      "content": "<p>thx a loooot <a href=\"https://www.kaggle.com/kmmohsin\" target=\"_blank\">@kmmohsin</a> for this great recap :)</p>",
      "rawMarkdown": "thx a loooot @kmmohsin for this great recap :)\n"
    },
    {
      "id": 1911425,
      "postDate": "2022-08-24T05:04:28.307Z",
      "content": "<p>If you are working on a huge dataset like this, then unless absolutely necessary, <strong>DO NOT LOOP THORUGH EACH ROW</strong> to calculate a new feature!! Try using NumPy arrays.</p>",
      "rawMarkdown": "If you are working on a huge dataset like this, then unless absolutely necessary, **DO NOT LOOP THORUGH EACH ROW** to calculate a new feature!! Try using NumPy arrays."
    },
    {
      "id": 1814606,
      "postDate": "2022-06-08T05:12:40.053Z",
      "content": "<p>Thanks for compiling this, super helpful for catching up. Cheers!</p>",
      "rawMarkdown": "Thanks for compiling this, super helpful for catching up. Cheers!"
    },
    {
      "id": 1810963,
      "postDate": "2022-06-04T05:39:16.747Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kmmohsin\" target=\"_blank\">@kmmohsin</a> thank you for sharing.</p>",
      "rawMarkdown": "Great work @kmmohsin thank you for sharing."
    },
    {
      "id": 1810421,
      "postDate": "2022-06-03T14:44:41.310Z",
      "content": "<p><a href=\"https://www.kaggle.com/elijahflorence\" target=\"_blank\">@elijahflorence</a>, to answer your question on how I got the class imbalance. In the competition overview tab&gt;&gt; Evaluation section competition host mentioned \"the negative labels are given a weight of 20 to adjust for downsampling\". That means negative samples are 20 times higher than what we are seeing in training set. I hope this answer your question.</p>",
      "rawMarkdown": "@elijahflorence, to answer your question on how I got the class imbalance. In the competition overview tab>> Evaluation section competition host mentioned \"the negative labels are given a weight of 20 to adjust for downsampling\". That means negative samples are 20 times higher than what we are seeing in training set. I hope this answer your question."
    },
    {
      "id": 1808890,
      "postDate": "2022-06-02T08:23:55.507Z",
      "content": "<p>Nice one. Thanks for bringing this all together!</p>",
      "rawMarkdown": "Nice one. Thanks for bringing this all together!"
    },
    {
      "id": 1811684,
      "postDate": "2022-06-05T02:25:39.877Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1809394,
      "postDate": "2022-06-02T16:46:45.610Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1869609,
      "postDate": "2022-07-24T22:02:36.363Z",
      "content": "<p>Awesome, thanks for sharing!</p>",
      "rawMarkdown": "Awesome, thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 1808489,
      "author_name": "Mohsin",
      "author_url": "",
      "post_date": "2022-06-01T21:39:39.053000",
      "content": "<p>train_df['B_31'].unique()<br>\narray([1, 0], dtype=int64)<br>\ntrain_df['D_87'].unique()<br>\narray([nan,  1.], dtype=float16)</p>\n<p>Looks like more column can be treated as int8. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808464,
      "author_name": "crayola",
      "author_url": "",
      "post_date": "2022-06-01T20:56:49.007000",
      "content": "<p>Nice one, thanks for compiling all the info!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1911767,
      "author_name": "SchopenHacker75",
      "author_url": "",
      "post_date": "2022-08-24T09:49:33.563000",
      "content": "<p>thx a loooot <a href=\"https://www.kaggle.com/kmmohsin\" target=\"_blank\">@kmmohsin</a> for this great recap :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1911425,
      "author_name": "Aakash Yadav",
      "author_url": "",
      "post_date": "2022-08-24T05:04:28.307000",
      "content": "<p>If you are working on a huge dataset like this, then unless absolutely necessary, <strong>DO NOT LOOP THORUGH EACH ROW</strong> to calculate a new feature!! Try using NumPy arrays.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1814606,
      "author_name": "Anuj Sindgi",
      "author_url": "",
      "post_date": "2022-06-08T05:12:40.053000",
      "content": "<p>Thanks for compiling this, super helpful for catching up. Cheers!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1810963,
      "author_name": "Hassan Shehzad",
      "author_url": "",
      "post_date": "2022-06-04T05:39:16.747000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kmmohsin\" target=\"_blank\">@kmmohsin</a> thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1810421,
      "author_name": "Mohsin",
      "author_url": "",
      "post_date": "2022-06-03T14:44:41.310000",
      "content": "<p><a href=\"https://www.kaggle.com/elijahflorence\" target=\"_blank\">@elijahflorence</a>, to answer your question on how I got the class imbalance. In the competition overview tab&gt;&gt; Evaluation section competition host mentioned \"the negative labels are given a weight of 20 to adjust for downsampling\". That means negative samples are 20 times higher than what we are seeing in training set. I hope this answer your question.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1808890,
      "author_name": "James McNeill",
      "author_url": "",
      "post_date": "2022-06-02T08:23:55.507000",
      "content": "<p>Nice one. Thanks for bringing this all together!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1811684,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T02:25:39.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809394,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T16:46:45.610000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1869609,
      "author_name": "George Lôbo",
      "author_url": "",
      "post_date": "2022-07-24T22:02:36.363000",
      "content": "<p>Awesome, thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1808414": "# Let's catchup with all the learnings so far!\n\nI started 7 days after this competition launched and trying to catch up with all the learnings. If you are also just starting fresh, then this discussion might be useful to you. For those of you are already contributing can ignore this. The purpose is to save us all some time by not repeating the same EDA, or same baseline model building. \n\nThis post is based on compilation of information from various posts. I tried to cite the actual contributors, please let me know if I miss any key contribution by mistake. Also please let me know if I am missing anything. I hope you also understand that I was not able to go over all the posts and all the notebooks. I hope you will find this useful! \n\nThis post has this following sections. \n\n 1. Memory optimization. \n 2. Metrics calculation. \n 3. Base line models. \n 4. High Level stats. \n\n**Memory Optimization:** \nAs you already noticed that this project requires us to download and process almost 50GB of data. To process this large amount of data we either need huge RAM or we need to optimize it to analyze with existing hardware we have.  Here are some tricks I found in the discussions useful, \n\n1.\tStart with feather data created by Ruchi. [feather_data](https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction)\n2.\tEncode customer ID column. It is ~600MB on RAM. Encoding this column will suppress it to ~40MB on RAM. \n\nUse following techniques for achieving it. \n\n1. From  [Chris Deotte](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635 )\n`train['customer_id'] =    train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')`\n\n2. Process other columns to optimize data. Most of these are to lower the precision. Might not impact initial model building but might impact in later stages of competition. Thanks, Chris Deotte, for this nice discussion. [optimize_other_cols](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054)\n\n**Cross Validation:** \nSince this is a example of imbalanced data we can use stratifiedkfold for cross validation. \n\n**Metrics Calculation:** \nThere is a nice explanation of the metrics here: [graphical explanation of evaluation metrics ](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464)\nFor implementation Rohan suggested multiple code snippet with their performance [here](https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations). \n\n**Base line models:** \nAs of 06/01/2022 I have compiled the list of outstanding models with their public score. I am acknowledging authors by directly sharing their notebook URLs. Few outstanding models as of 6/1/2022: (Not in any order). \n\n| Technique/Tools                                                                             | Public Score | URL                                                                                                                                                                     |\n| ------------------------------------------------------------------------------------------- | ------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Catboost                                                                                    | 0.783        | [https://www.kaggle.com/code/aninda/first-submission-using-catboost ](https://www.kaggle.com/code/aninda/first-submission-using-catboost)                               |\n| Catboost                                                                                    | 0.793        | [https://www.kaggle.com/code/huseyincot/amex-catboost-0-793 ](https://www.kaggle.com/code/huseyincot/amex-catboost-0-793)                                               |\n| Ensemble weighted average                                                                   | 0.796        | [https://www.kaggle.com/code/beezus666/ensemble-weighted-average ](https://www.kaggle.com/code/beezus666/ensemble-weighted-average)                                     |\n| GradientBoosting (py-boost)                                                                 | 0.791        | [https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline ](https://www.kaggle.com/code/btbpanda/fast-metric-and-py-boost-baseline)                       |\n| LGBM                                                                                        | 0.786        | [https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng ](https://www.kaggle.com/code/lucasmorin/amex-lgbm-features-eng)                                         |\n| LGBM                                                                                        | 0.792        | [https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart)                                            |\n| LGBM                                                                                        | 0.765        | [https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model ](https://www.kaggle.com/code/thebratattack/amex-ensemble-prediction-model)                   |\n| LGBM                                                                                        | 0.783        | [https://www.kaggle.com/code/munumbutt/simple-lgbm-starter ](https://www.kaggle.com/code/munumbutt/simple-lgbm-starter)                                                 |\n| LightAutoML                                                                                 | 0.794        | [https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter ](https://www.kaggle.com/code/alexryzhkov/amex-lightautoml-starter)                                   |\n| NN + skip connection + Dropout layer. Keras                                                 | 0.79         | [https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training ](https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training)                           |\n| NN using Keras                                                                              | 0.783        | [https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter ](https://www.kaggle.com/code/cv13j0/amex-default-prediction-keras-starter)                   |\n| Randdom Forest                                                                              | 0.715        | [https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation ](https://www.kaggle.com/code/bgmello/best-correlated-features-with-low-correlation) |\n| Random Forest with only two predictors (\\[\\['P\\_2', 'D\\_48'\\]) and only for last statement. | 0.595        | [https://www.kaggle.com/code/bgmello/the-platinum-solution ](https://www.kaggle.com/code/bgmello/the-platinum-solution)                                                 |\n| TensorFlow GRU                                                                              | 0.79         | [https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790 ](https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790)                                   |\n| TensorFlow Transformer                                                                      | 0.789        | [https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790 ](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790)                                   |\n| Weight of Evidence (WOE)                                                                    | 0.7          | [https://www.kaggle.com/code/lucasmorin/amex-woe-baseline ](https://www.kaggle.com/code/lucasmorin/amex-woe-baseline)                                                   |\n| Weight of Evidence (WOE) + OptBinning                                                       | 0.756        | https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model                                                                                                     |\n| XGBoost                                                                                     |              | [https://www.kaggle.com/code/datajmcn/baseline-model-xgboost ](https://www.kaggle.com/code/datajmcn/baseline-model-xgboost)                                             |\n| XGBoost with @raddar data                                                                   | 0.793        | https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793                                                                                                               |                                                     \n\n\n**High Level Stats:** \nFollowing are some useful numbers and findings I though might be helpful. \n\n| Key                                                      | Value                                                                                                                                                                                                                    |\n| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| Total Unique Customers in Train                          | 458913                                                                                                                                                                                                                   |\n| Total Unique Customers in Test                           | 924621                                                                                                                                                                                                                   |\n| Non-Defaulted (Negative Class) Customers in Training Set | 340000 (74%), Down sampled at 5% from Population                                                                                                                                                                         |\n| Defaulted Customers in Training Set                      | 119000 (26 %)                                                                                                                                                                                                            |\n| Class imbalance in Training Set                          | 74%/26%                                                                                                                                                                                                                  |\n| Class imbalance in Population                            | 98%/2%                                                                                                                                                                                                                   |\n| Min(date), max(date) in Training Set                     | 2017-03-01, 2018-03-31                                                                                                                                                                                                   |\n| Min(date), max(date) in Test Set                         | 2018-04-01, 2019-10-31                                                                                                                                                                                                   |\n| Data Dimension in Training Set                           | 458913 Customers \\* ~13 statements \\* 190ish Columns                                                                                                                                                                     |\n| Binary Features                                          | B\\_31 is always 0 or 1 and D\\_87 is always 1 or missing.                                                                                                                                                            |\n| Categorical Features                                     | Categorical features are defined as follows.<br>`['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']`                                                                          |\n\nThere are few more columns can be added to the list of categorical (ordinal), e.g. \ntrain_df['B_31'].unique()\narray([1, 0], dtype=int64)\ntrain_df['D_87'].unique()\narray([nan,  1.], dtype=float16)\n\nThank you all for sharing all those nice thoughts, ideas  in various discussions and notebooks. I will try to catchup more of your works and will try to summarize here. Looking forward to engaging more in coming days!",
    "1808489": "train_df['B_31'].unique()\narray([1, 0], dtype=int64)\ntrain_df['D_87'].unique()\narray([nan,  1.], dtype=float16)\n\nLooks like more column can be treated as int8. ",
    "1808464": "Nice one, thanks for compiling all the info!",
    "1911767": "thx a loooot @kmmohsin for this great recap :)\n",
    "1911425": "If you are working on a huge dataset like this, then unless absolutely necessary, **DO NOT LOOP THORUGH EACH ROW** to calculate a new feature!! Try using NumPy arrays.",
    "1814606": "Thanks for compiling this, super helpful for catching up. Cheers!",
    "1810963": "Great work @kmmohsin thank you for sharing.",
    "1810421": "@elijahflorence, to answer your question on how I got the class imbalance. In the competition overview tab>> Evaluation section competition host mentioned \"the negative labels are given a weight of 20 to adjust for downsampling\". That means negative samples are 20 times higher than what we are seeing in training set. I hope this answer your question.",
    "1808890": "Nice one. Thanks for bringing this all together!",
    "1811684": "",
    "1809394": "",
    "1869609": "Awesome, thanks for sharing!"
  }
}