{
  "id": 347908,
  "title": "15th Place Solution Meta features ,FE, DART, CAT, XG , Tabnet , MLP , ensemble 😊",
  "url": "/competitions/amex-default-prediction/writeups/hungry-for-gold-medal-15th-place-solution-meta-fea",
  "author_name": "",
  "post_date": "2022-08-29T20:44:44.777Z",
  "votes": 38,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks first to Kaggle for hosting this interesting competition  , I was personally interested in this as it cynosure with the financial domain which I spend majority of my time working on . We had a great team with <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> <a href=\"https://www.kaggle.com/liji11\" target=\"_blank\">@liji11</a> <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> and <a href=\"https://www.kaggle.com/hanzhou0315\" target=\"_blank\">@hanzhou0315</a>  who each brought there unique skills to the competition.. A big thank you to all of them .</p>\n<h1>Feature Engineering</h1>\n<p>Our feature engineering structure was influenced by this great notebook <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">rapids-cudf-feature-engineering-xgb</a> from <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> </p>\n<h2>Meta Features</h2>\n<p>Our meta features were the differentiations to boost our ensembles .These are similar and well described here <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786\" target=\"_blank\">12th Place Solution</a> and a big thanks for <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> to come up with those . As Sirius also mentioned in that post we used them a s 13 numerical features per CID after flattening .</p>\n<h2>For models without meta features</h2>\n<ul>\n<li>Great features suggested by Ragnar helped like last-mean features and last-min/max features for some columns, also we did apply diff 1,2 lags as well .After pay features were also helpful .</li>\n<li>We did usual aggregates std,mean,min,max for categoricals (nunique,count,first,mean). We also tried MAD (Mean Absolute deviation) and it did help our cv though Public LB was bit less so we didn't include the model based on mad in final scores though it could have helped in hindsight ,</li>\n<li>We also did add percentage change features , which basically do the percentage change on numerical features and thought they would work better than just a numeric diff . They did help us in some models .</li>\n<li>Also one feature that was there in most of our good models was keeping last, first and middle component of the statements as we hypothesized that will help us cover most variation for a customer as we had aggregated features .</li>\n<li>Another feature we tried was to say calculate spend/balance or spend_sum/balance_sum ratios , It did help in some models but not most .</li>\n<li>Trimming of non impactful features also did help to reduce number of features in 1k-2k range .</li>\n</ul>\n<h2>Hybrid approach (meta+aggs) :</h2>\n<p>We also gave a shot at mixing the flattened meta features with aggregates and it did help our models specially LGBM CAT XG and Tabnet.<br>\nCause of meta features this usually converged faster (fewer rounds with early stopping in XG) so we had to <strong>Lower LR</strong>  for this approach .</p>\n<h1>Models</h1>\n<p>All models were trained on meta features and/or engineered aggregates . NN were scaled for the most part with GaussianScalar</p>\n<p><strong>LGBM (Max cv 0.79932 Private 0.80731)</strong>: This was as with most based on notebook by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> . We did do little tweaks to the hyperparams but most were similar .Here as noted in discussions lowering the LR certainly helped in the case of DART and also us adding meta features (around <em>0.0075 LR</em>).<br>\n<strong>XG (Max CV 0.7984 Priv. 0.80687)</strong>: XG was based on <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>  great notebook<br>\n<strong>CAT (Max CV 0.7972)</strong>: Cat we did tune a bit with below params and it helped us a lot .</p>\n<pre><code>CatBoostClassifier( random_state=CFG.seed,  \n                                bootstrap_type='Bernoulli',\n                                task_type=\"GPU\",\n                                devices='0:1',  \n                                use_best_model = True, \n                                iterations = 11500,\n                                num_leaves =64,\n                                subsample = 0.74,\n                                grow_policy = 'Lossguide', \n                                depth = 9) \n</code></pre>\n<p><strong>Tabnet (MAX CV 0.793133 Private 0.80447)</strong>: Tabnet used mostly standard params and did give better results with meta features .<br>\n<strong>MLP (Private 0.80109).</strong> We tried models with tabnet using cat embeddings and without (by onehotencoding). Also <strong>GaussianScaler</strong> helped with scaling of numerical features in case of NN and tabnet<br>\n<strong>AutoML (MAX CV 0.79714)</strong>: <a href=\"https://www.kaggle.com/liji11\" target=\"_blank\">@liji11</a> tried automl which did help our ensemble <br>\n<strong>Tabformer ( MAX CV 0.79507 Private 0.80480)</strong> : This mostly based on the public <a href=\"https://www.kaggle.com/code/gauravbrills/tabtransformer-training\" target=\"_blank\">notebook</a>  with more heads as per the paper . This we noticed was better than tabnet in private LB but we didnt choose it as was not working well in ensemble 😑. Hyperparams below </p>\n<pre><code>NUM_TRANSFORMER_BLOCKS = 6  # Number of transformer blocks. 6 paper recommends\nNUM_HEADS = 8  # Number of attention heads. 8 Heads paper recommends\nEMBEDDING_DIMS = 16#10  # Embedding dimensions of the categorical features. check 16,32\nDROPOUT_RATE = 0.1\nMLP_HIDDEN_UNITS_FACTORS = [\n    4,\n    2,\n]  # MLP hidden layer units, as factors of the number of inputs. =&gt;(4,2)\nNUM_MLP_BLOCKS = 4  # Number of MLP blocks in the baseline model. Paper 4\nMLP_ACTIVATION = keras.activations.selu\n</code></pre>\n<p><strong>TCN MLP</strong>: TCN over 13 sequences , Not chosen Max CV 0.78<br>\nGoes like this inspired by the amex paper </p>\n<pre><code>   embeddings = []\n    for k in range(11): \n      vocabulary = CATEGORICAL_FEATURES_WITH_VOCABULARY[cat_cols[k]]\n      #print(f\"cat {cat_cols[k]} index {k} len {len(vocabulary)}\")\n      emb = tf.keras.layers.Embedding(len(vocabulary),EMBEDDING_DIMS)\n      embeddings.append(emb(inputs[:,:,k]))\n    in_ = tf.keras.layers.Concatenate()([inputs[:,:,11:]]+embeddings)  \n    activation = 'swish'\n    l1 = 1e-7\n    l2 = 4e-4\n    reg = 4e-4\n    # SIMPLE Wavenet TCN BACKBONE  \n    _x = BatchNormalization()(in_)\n    x = TCN(nb_filters=256, kernel_size=4,return_sequences=False, dropout_rate=0.0, dilations=[2 ** i for i in range(9)])(_x)\n    #x = TCN(nb_filters=128, kernel_size=3,return_sequences=False,use_layer_norm=True,dropout_rate=0.05,dilations=[2 ** i for i in range(7)])(x)\n    x0 = Dense(128, \n               kernel_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n#                activity_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n              activation=activation,\n             )(x)\n    x0 = Dropout(0.1)(x0)\n... MORE OF MLP ....\nx_output = Dense(1,\n              activation='sigmoid',\n             )(x)\n</code></pre>\n<h1>Ensemble techniques</h1>\n<p>All ensembles were done on log odds as we found it worked better when ensembling NN models . We had also done ensemble with rank method but that didnt work well when ensembling with NN .</p>\n<p>We tried a bunch of meta ensemble techniques with are many models .The ones that worked were <a href=\"https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">forward selection</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , Optuna weights and Elasticnet on top of oof preds . We also tried stacking and voting techniques but they weren't very successful.</p>\n<h1>What did not work or DID ?</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/gauravbrills/tabtransformer-training\" target=\"_blank\">Tabformer</a> : We finally made tabformer reach 0.795 cv but it was not part of our final submission. This I will say <em>did work</em> but did not ensemble well though later we realized was our best Private LB scoring NN .</li>\n<li>TCN wavenet : We tried to use 13 customer history sequences in a tcn wavenet + MLP implementation but results were only around 0.78 ish so we dropped the idea .</li>\n<li>Saint and widedeep models : We also gave a stab on wide deep package to try out  <a href=\"https://arxiv.org/abs/2106.01342\" target=\"_blank\">SAINT</a>  but were not able to get far with that . reference <a href=\"https://github.com/jrzaurin/pytorch-widedeep\" target=\"_blank\">pytorch-widedeep</a></li>\n<li>Some features as described above and some dropped based on permutation or zero feature importance .</li>\n<li>We tried to do multi seed ensembles for all our good models . This somehow did not give us good results at the end compared to sticking with 42 seed.</li>\n<li>Sirius also tried Target encoding at the end but we had really very little time to check on these in ensemble .</li>\n<li>We also tried GAN techniques to impute missing values but dropped them at the end .</li>\n</ul>\n<p>Finally luckily we selected a good enough final submission for 🏅(There was as usual a lot of confusion thanks <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> for voting for this 😄), though seems some with lower LB were better so should have trusted cv a bit more :) </p>\n<p><strong>✅  Best selected submission Private LB 0.80838 ( Optuna weighted models , xg, tabnet.mlp,lgbm,cat and automl)</strong><br>\n<strong>📮 Best submission : Private LB 0.80852 ( Optuna xg,cat,tabnet,mlp,automl and lgbm)</strong></p>\n<p>Thanks for reading 😄</p>",
  "messages": [
    {
      "id": "1914272",
      "postDate": "08/25/2022 23:32:35",
      "content": "<p>Thanks first to Kaggle for hosting this interesting competition  , I was personally interested in this as it cynosure with the financial domain which I spend majority of my time working on . We had a great team with <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> <a href=\"https://www.kaggle.com/liji11\" target=\"_blank\">@liji11</a> <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> and <a href=\"https://www.kaggle.com/hanzhou0315\" target=\"_blank\">@hanzhou0315</a>  who each brought there unique skills to the competition.. A big thank you to all of them .</p>\n<h1>Feature Engineering</h1>\n<p>Our feature engineering structure was influenced by this great notebook <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">rapids-cudf-feature-engineering-xgb</a> from <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> </p>\n<h2>Meta Features</h2>\n<p>Our meta features were the differentiations to boost our ensembles .These are similar and well described here <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786\" target=\"_blank\">12th Place Solution</a> and a big thanks for <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> to come up with those . As Sirius also mentioned in that post we used them a s 13 numerical features per CID after flattening .</p>\n<h2>For models without meta features</h2>\n<ul>\n<li>Great features suggested by Ragnar helped like last-mean features and last-min/max features for some columns, also we did apply diff 1,2 lags as well .After pay features were also helpful .</li>\n<li>We did usual aggregates std,mean,min,max for categoricals (nunique,count,first,mean). We also tried MAD (Mean Absolute deviation) and it did help our cv though Public LB was bit less so we didn't include the model based on mad in final scores though it could have helped in hindsight ,</li>\n<li>We also did add percentage change features , which basically do the percentage change on numerical features and thought they would work better than just a numeric diff . They did help us in some models .</li>\n<li>Also one feature that was there in most of our good models was keeping last, first and middle component of the statements as we hypothesized that will help us cover most variation for a customer as we had aggregated features .</li>\n<li>Another feature we tried was to say calculate spend/balance or spend_sum/balance_sum ratios , It did help in some models but not most .</li>\n<li>Trimming of non impactful features also did help to reduce number of features in 1k-2k range .</li>\n</ul>\n<h2>Hybrid approach (meta+aggs) :</h2>\n<p>We also gave a shot at mixing the flattened meta features with aggregates and it did help our models specially LGBM CAT XG and Tabnet.<br>\nCause of meta features this usually converged faster (fewer rounds with early stopping in XG) so we had to <strong>Lower LR</strong>  for this approach .</p>\n<h1>Models</h1>\n<p>All models were trained on meta features and/or engineered aggregates . NN were scaled for the most part with GaussianScalar</p>\n<p><strong>LGBM (Max cv 0.79932 Private 0.80731)</strong>: This was as with most based on notebook by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> . We did do little tweaks to the hyperparams but most were similar .Here as noted in discussions lowering the LR certainly helped in the case of DART and also us adding meta features (around <em>0.0075 LR</em>).<br>\n<strong>XG (Max CV 0.7984 Priv. 0.80687)</strong>: XG was based on <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>  great notebook<br>\n<strong>CAT (Max CV 0.7972)</strong>: Cat we did tune a bit with below params and it helped us a lot .</p>\n<pre><code>CatBoostClassifier( random_state=CFG.seed,  \n                                bootstrap_type='Bernoulli',\n                                task_type=\"GPU\",\n                                devices='0:1',  \n                                use_best_model = True, \n                                iterations = 11500,\n                                num_leaves =64,\n                                subsample = 0.74,\n                                grow_policy = 'Lossguide', \n                                depth = 9) \n</code></pre>\n<p><strong>Tabnet (MAX CV 0.793133 Private 0.80447)</strong>: Tabnet used mostly standard params and did give better results with meta features .<br>\n<strong>MLP (Private 0.80109).</strong> We tried models with tabnet using cat embeddings and without (by onehotencoding). Also <strong>GaussianScaler</strong> helped with scaling of numerical features in case of NN and tabnet<br>\n<strong>AutoML (MAX CV 0.79714)</strong>: <a href=\"https://www.kaggle.com/liji11\" target=\"_blank\">@liji11</a> tried automl which did help our ensemble <br>\n<strong>Tabformer ( MAX CV 0.79507 Private 0.80480)</strong> : This mostly based on the public <a href=\"https://www.kaggle.com/code/gauravbrills/tabtransformer-training\" target=\"_blank\">notebook</a>  with more heads as per the paper . This we noticed was better than tabnet in private LB but we didnt choose it as was not working well in ensemble 😑. Hyperparams below </p>\n<pre><code>NUM_TRANSFORMER_BLOCKS = 6  # Number of transformer blocks. 6 paper recommends\nNUM_HEADS = 8  # Number of attention heads. 8 Heads paper recommends\nEMBEDDING_DIMS = 16#10  # Embedding dimensions of the categorical features. check 16,32\nDROPOUT_RATE = 0.1\nMLP_HIDDEN_UNITS_FACTORS = [\n    4,\n    2,\n]  # MLP hidden layer units, as factors of the number of inputs. =&gt;(4,2)\nNUM_MLP_BLOCKS = 4  # Number of MLP blocks in the baseline model. Paper 4\nMLP_ACTIVATION = keras.activations.selu\n</code></pre>\n<p><strong>TCN MLP</strong>: TCN over 13 sequences , Not chosen Max CV 0.78<br>\nGoes like this inspired by the amex paper </p>\n<pre><code>   embeddings = []\n    for k in range(11): \n      vocabulary = CATEGORICAL_FEATURES_WITH_VOCABULARY[cat_cols[k]]\n      #print(f\"cat {cat_cols[k]} index {k} len {len(vocabulary)}\")\n      emb = tf.keras.layers.Embedding(len(vocabulary),EMBEDDING_DIMS)\n      embeddings.append(emb(inputs[:,:,k]))\n    in_ = tf.keras.layers.Concatenate()([inputs[:,:,11:]]+embeddings)  \n    activation = 'swish'\n    l1 = 1e-7\n    l2 = 4e-4\n    reg = 4e-4\n    # SIMPLE Wavenet TCN BACKBONE  \n    _x = BatchNormalization()(in_)\n    x = TCN(nb_filters=256, kernel_size=4,return_sequences=False, dropout_rate=0.0, dilations=[2 ** i for i in range(9)])(_x)\n    #x = TCN(nb_filters=128, kernel_size=3,return_sequences=False,use_layer_norm=True,dropout_rate=0.05,dilations=[2 ** i for i in range(7)])(x)\n    x0 = Dense(128, \n               kernel_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n#                activity_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n              activation=activation,\n             )(x)\n    x0 = Dropout(0.1)(x0)\n... MORE OF MLP ....\nx_output = Dense(1,\n              activation='sigmoid',\n             )(x)\n</code></pre>\n<h1>Ensemble techniques</h1>\n<p>All ensembles were done on log odds as we found it worked better when ensembling NN models . We had also done ensemble with rank method but that didnt work well when ensembling with NN .</p>\n<p>We tried a bunch of meta ensemble techniques with are many models .The ones that worked were <a href=\"https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">forward selection</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , Optuna weights and Elasticnet on top of oof preds . We also tried stacking and voting techniques but they weren't very successful.</p>\n<h1>What did not work or DID ?</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/gauravbrills/tabtransformer-training\" target=\"_blank\">Tabformer</a> : We finally made tabformer reach 0.795 cv but it was not part of our final submission. This I will say <em>did work</em> but did not ensemble well though later we realized was our best Private LB scoring NN .</li>\n<li>TCN wavenet : We tried to use 13 customer history sequences in a tcn wavenet + MLP implementation but results were only around 0.78 ish so we dropped the idea .</li>\n<li>Saint and widedeep models : We also gave a stab on wide deep package to try out  <a href=\"https://arxiv.org/abs/2106.01342\" target=\"_blank\">SAINT</a>  but were not able to get far with that . reference <a href=\"https://github.com/jrzaurin/pytorch-widedeep\" target=\"_blank\">pytorch-widedeep</a></li>\n<li>Some features as described above and some dropped based on permutation or zero feature importance .</li>\n<li>We tried to do multi seed ensembles for all our good models . This somehow did not give us good results at the end compared to sticking with 42 seed.</li>\n<li>Sirius also tried Target encoding at the end but we had really very little time to check on these in ensemble .</li>\n<li>We also tried GAN techniques to impute missing values but dropped them at the end .</li>\n</ul>\n<p>Finally luckily we selected a good enough final submission for 🏅(There was as usual a lot of confusion thanks <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> for voting for this 😄), though seems some with lower LB were better so should have trusted cv a bit more :) </p>\n<p><strong>✅  Best selected submission Private LB 0.80838 ( Optuna weighted models , xg, tabnet.mlp,lgbm,cat and automl)</strong><br>\n<strong>📮 Best submission : Private LB 0.80852 ( Optuna xg,cat,tabnet,mlp,automl and lgbm)</strong></p>\n<p>Thanks for reading 😄</p>",
      "rawMarkdown": "Thanks first to Kaggle for hosting this interesting competition  , I was personally interested in this as it cynosure with the financial domain which I spend majority of my time working on . We had a great team with @sirius81 @liji11 @tonymarkchris and @hanzhou0315  who each brought there unique skills to the competition.. A big thank you to all of them .\n\n# Feature Engineering\nOur feature engineering structure was influenced by this great notebook [rapids-cudf-feature-engineering-xgb](https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb) from @jiweiliu \n\n## Meta Features\nOur meta features were the differentiations to boost our ensembles .These are similar and well described here [12th Place Solution](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786) and a big thanks for @sirius81 to come up with those . As Sirius also mentioned in that post we used them a s 13 numerical features per CID after flattening .\n\n## For models without meta features\n- Great features suggested by Ragnar helped like last-mean features and last-min/max features for some columns, also we did apply diff 1,2 lags as well .After pay features were also helpful .\n- We did usual aggregates std,mean,min,max for categoricals (nunique,count,first,mean). We also tried MAD (Mean Absolute deviation) and it did help our cv though Public LB was bit less so we didn't include the model based on mad in final scores though it could have helped in hindsight ,\n- We also did add percentage change features , which basically do the percentage change on numerical features and thought they would work better than just a numeric diff . They did help us in some models .\n- Also one feature that was there in most of our good models was keeping last, first and middle component of the statements as we hypothesized that will help us cover most variation for a customer as we had aggregated features .\n- Another feature we tried was to say calculate spend/balance or spend_sum/balance_sum ratios , It did help in some models but not most .\n- Trimming of non impactful features also did help to reduce number of features in 1k-2k range .\n\n## Hybrid approach (meta+aggs) : \nWe also gave a shot at mixing the flattened meta features with aggregates and it did help our models specially LGBM CAT XG and Tabnet.\nCause of meta features this usually converged faster (fewer rounds with early stopping in XG) so we had to **Lower LR**  for this approach .\n\n#Models\nAll models were trained on meta features and/or engineered aggregates . NN were scaled for the most part with GaussianScalar\n\n**LGBM (Max cv 0.79932 Private 0.80731)**: This was as with most based on notebook by @ragnar123 . We did do little tweaks to the hyperparams but most were similar .Here as noted in discussions lowering the LR certainly helped in the case of DART and also us adding meta features (around *0.0075 LR*).\n**XG (Max CV 0.7984 Priv. 0.80687)**: XG was based on @jiweiliu  great notebook\n**CAT (Max CV 0.7972)**: Cat we did tune a bit with below params and it helped us a lot .\n```python\nCatBoostClassifier( random_state=CFG.seed,  \n                                bootstrap_type='Bernoulli',\n                                task_type=\"GPU\",\n                                devices='0:1',  \n                                use_best_model = True, \n                                iterations = 11500,\n                                num_leaves =64,\n                                subsample = 0.74,\n                                grow_policy = 'Lossguide', \n                                depth = 9) \n```\n**Tabnet (MAX CV 0.793133 Private 0.80447)**: Tabnet used mostly standard params and did give better results with meta features .\n**MLP (Private 0.80109).** We tried models with tabnet using cat embeddings and without (by onehotencoding). Also **GaussianScaler** helped with scaling of numerical features in case of NN and tabnet\n**AutoML (MAX CV 0.79714)**: @liji11 tried automl which did help our ensemble \n**Tabformer ( MAX CV 0.79507 Private 0.80480)** : This mostly based on the public [notebook](https://www.kaggle.com/code/gauravbrills/tabtransformer-training)  with more heads as per the paper . This we noticed was better than tabnet in private LB but we didnt choose it as was not working well in ensemble 😑. Hyperparams below \n```python\nNUM_TRANSFORMER_BLOCKS = 6  # Number of transformer blocks. 6 paper recommends\nNUM_HEADS = 8  # Number of attention heads. 8 Heads paper recommends\nEMBEDDING_DIMS = 16#10  # Embedding dimensions of the categorical features. check 16,32\nDROPOUT_RATE = 0.1\nMLP_HIDDEN_UNITS_FACTORS = [\n    4,\n    2,\n]  # MLP hidden layer units, as factors of the number of inputs. =>(4,2)\nNUM_MLP_BLOCKS = 4  # Number of MLP blocks in the baseline model. Paper 4\nMLP_ACTIVATION = keras.activations.selu\n```\n\n**TCN MLP**: TCN over 13 sequences , Not chosen Max CV 0.78\nGoes like this inspired by the amex paper \n```python\n   embeddings = []\n    for k in range(11): \n      vocabulary = CATEGORICAL_FEATURES_WITH_VOCABULARY[cat_cols[k]]\n      #print(f\"cat {cat_cols[k]} index {k} len {len(vocabulary)}\")\n      emb = tf.keras.layers.Embedding(len(vocabulary),EMBEDDING_DIMS)\n      embeddings.append(emb(inputs[:,:,k]))\n    in_ = tf.keras.layers.Concatenate()([inputs[:,:,11:]]+embeddings)  \n    activation = 'swish'\n    l1 = 1e-7\n    l2 = 4e-4\n    reg = 4e-4\n    # SIMPLE Wavenet TCN BACKBONE  \n    _x = BatchNormalization()(in_)\n    x = TCN(nb_filters=256, kernel_size=4,return_sequences=False, dropout_rate=0.0, dilations=[2 ** i for i in range(9)])(_x)\n    #x = TCN(nb_filters=128, kernel_size=3,return_sequences=False,use_layer_norm=True,dropout_rate=0.05,dilations=[2 ** i for i in range(7)])(x)\n    x0 = Dense(128, \n               kernel_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n#                activity_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n              activation=activation,\n             )(x)\n    x0 = Dropout(0.1)(x0)\n... MORE OF MLP ....\nx_output = Dense(1,\n              activation='sigmoid',\n             )(x)\n```\n\n# Ensemble techniques\n\nAll ensembles were done on log odds as we found it worked better when ensembling NN models . We had also done ensemble with rank method but that didnt work well when ensembling with NN .\n\nWe tried a bunch of meta ensemble techniques with are many models .The ones that worked were [forward selection](https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private) by @cdeotte , Optuna weights and Elasticnet on top of oof preds . We also tried stacking and voting techniques but they weren't very successful.\n\n# What did not work or DID ?\n\n- [Tabformer](https://www.kaggle.com/code/gauravbrills/tabtransformer-training) : We finally made tabformer reach 0.795 cv but it was not part of our final submission. This I will say *did work* but did not ensemble well though later we realized was our best Private LB scoring NN .\n- TCN wavenet : We tried to use 13 customer history sequences in a tcn wavenet + MLP implementation but results were only around 0.78 ish so we dropped the idea .\n- Saint and widedeep models : We also gave a stab on wide deep package to try out  [SAINT](https://arxiv.org/abs/2106.01342)  but were not able to get far with that . reference [pytorch-widedeep](https://github.com/jrzaurin/pytorch-widedeep)\n- Some features as described above and some dropped based on permutation or zero feature importance .\n- We tried to do multi seed ensembles for all our good models . This somehow did not give us good results at the end compared to sticking with 42 seed.\n- Sirius also tried Target encoding at the end but we had really very little time to check on these in ensemble .\n- We also tried GAN techniques to impute missing values but dropped them at the end .\n\nFinally luckily we selected a good enough final submission for 🏅(There was as usual a lot of confusion thanks @tonymarkchris for voting for this 😄), though seems some with lower LB were better so should have trusted cv a bit more :) \n\n**✅  Best selected submission Private LB 0.80838 ( Optuna weighted models , xg, tabnet.mlp,lgbm,cat and automl)**\n**📮 Best submission : Private LB 0.80852 ( Optuna xg,cat,tabnet,mlp,automl and lgbm)**\n\nThanks for reading 😄",
      "votes": null
    },
    {
      "id": "1914353",
      "postDate": "08/26/2022 03:06:32",
      "content": "<p>This is my first time to join a team and the teamwork is smooth. Thank my teammates for your interesting ideas and hard working! <br>\nAs the making detail of <strong>meta features</strong> has beed described in this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786\" target=\"_blank\">thread</a> and I didn't know the so-called meta feature before, here I just give some motivation of coming up with these features.</p>\n<ul>\n<li>In the most of the public notebooks, features are based on agg methods, such as min, max, mean and std.</li>\n<li>Though these agg features are expressive, some information is losing during the 13 statements being compressed into limited statistics.</li>\n<li>I just thought about <strong>how can I utilise the information of each one statement instead of aggregating them.</strong></li>\n<li>Flatting the data from [N_USER, 13, 188] into [N_USER, 13*188] is a straight method, but I feel it will increase the feature dim too much. </li>\n<li>Then I came up with the idea of predicting a default score of each statement, <strong>which can be regarded as a distillation of each statement</strong>.</li>\n<li>These 13 predicted scores are the so-called meta features , which are used as numeric features of each customer in the following models.</li>\n<li>Finally, the meta features show a great improvement in my experiments, like<ul>\n<li>tabnet cv +0.003, 0.791-&gt;0.794 (lb 0.796)</li>\n<li>xgb cv +0.0015, 0.7953-&gt;0.7968 (lb 0.798)</li>\n<li>lgb cv +0.000x (not sure)</li></ul></li>\n</ul>",
      "rawMarkdown": "This is my first time to join a team and the teamwork is smooth. Thank my teammates for your interesting ideas and hard working! \nAs the making detail of **meta features** has beed described in this [thread](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786) and I didn't know the so-called meta feature before, here I just give some motivation of coming up with these features.\n- In the most of the public notebooks, features are based on agg methods, such as min, max, mean and std.\n- Though these agg features are expressive, some information is losing during the 13 statements being compressed into limited statistics.\n- I just thought about **how can I utilise the information of each one statement instead of aggregating them.**\n- Flatting the data from [N_USER, 13, 188] into [N_USER, 13*188] is a straight method, but I feel it will increase the feature dim too much. \n- Then I came up with the idea of predicting a default score of each statement, **which can be regarded as a distillation of each statement**.\n- These 13 predicted scores are the so-called meta features , which are used as numeric features of each customer in the following models.\n- Finally, the meta features show a great improvement in my experiments, like\n\t- tabnet cv +0.003, 0.791->0.794 (lb 0.796)\n\t- xgb cv +0.0015, 0.7953->0.7968 (lb 0.798)\n\t- lgb cv +0.000x (not sure)",
      "votes": null
    },
    {
      "id": "1915387",
      "postDate": "08/27/2022 00:07:06",
      "content": "<p>Great teammates👍</p>",
      "rawMarkdown": "Great teammates👍",
      "votes": null
    },
    {
      "id": "1915404",
      "postDate": "08/27/2022 00:24:22",
      "content": "<p>Yes <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> your features were the 🪄 for us .. Thanks 👍</p>",
      "rawMarkdown": "Yes @sirius81 your features were the 🪄 for us .. Thanks 👍",
      "votes": null
    },
    {
      "id": "1915405",
      "postDate": "08/27/2022 00:24:52",
      "content": "<p>Yess 🙌 <a href=\"https://www.kaggle.com/liji11\" target=\"_blank\">@liji11</a> </p>",
      "rawMarkdown": "Yess 🙌 @liji11",
      "votes": null
    },
    {
      "id": "1915863",
      "postDate": "08/27/2022 12:45:17",
      "content": "<p>👌It's great</p>",
      "rawMarkdown": "👌It's great",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1914353,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "08/26/2022 03:06:32",
      "content": "<p>This is my first time to join a team and the teamwork is smooth. Thank my teammates for your interesting ideas and hard working! <br>\nAs the making detail of <strong>meta features</strong> has beed described in this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786\" target=\"_blank\">thread</a> and I didn't know the so-called meta feature before, here I just give some motivation of coming up with these features.</p>\n<ul>\n<li>In the most of the public notebooks, features are based on agg methods, such as min, max, mean and std.</li>\n<li>Though these agg features are expressive, some information is losing during the 13 statements being compressed into limited statistics.</li>\n<li>I just thought about <strong>how can I utilise the information of each one statement instead of aggregating them.</strong></li>\n<li>Flatting the data from [N_USER, 13, 188] into [N_USER, 13*188] is a straight method, but I feel it will increase the feature dim too much. </li>\n<li>Then I came up with the idea of predicting a default score of each statement, <strong>which can be regarded as a distillation of each statement</strong>.</li>\n<li>These 13 predicted scores are the so-called meta features , which are used as numeric features of each customer in the following models.</li>\n<li>Finally, the meta features show a great improvement in my experiments, like<ul>\n<li>tabnet cv +0.003, 0.791-&gt;0.794 (lb 0.796)</li>\n<li>xgb cv +0.0015, 0.7953-&gt;0.7968 (lb 0.798)</li>\n<li>lgb cv +0.000x (not sure)</li></ul></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1915404,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "08/27/2022 00:24:22",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> your features were the 🪄 for us .. Thanks 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915387,
      "author_name": "liji11",
      "author_url": "",
      "post_date": "08/27/2022 00:07:06",
      "content": "<p>Great teammates👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915405,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "08/27/2022 00:24:52",
          "content": "<p>Yess 🙌 <a href=\"https://www.kaggle.com/liji11\" target=\"_blank\">@liji11</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915863,
      "author_name": "kushasabzevari",
      "author_url": "",
      "post_date": "08/27/2022 12:45:17",
      "content": "<p>👌It's great</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914272": "Thanks first to Kaggle for hosting this interesting competition  , I was personally interested in this as it cynosure with the financial domain which I spend majority of my time working on . We had a great team with @sirius81 @liji11 @tonymarkchris and @hanzhou0315  who each brought there unique skills to the competition.. A big thank you to all of them .\n\n# Feature Engineering\nOur feature engineering structure was influenced by this great notebook [rapids-cudf-feature-engineering-xgb](https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb) from @jiweiliu \n\n## Meta Features\nOur meta features were the differentiations to boost our ensembles .These are similar and well described here [12th Place Solution](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786) and a big thanks for @sirius81 to come up with those . As Sirius also mentioned in that post we used them a s 13 numerical features per CID after flattening .\n\n## For models without meta features\n- Great features suggested by Ragnar helped like last-mean features and last-min/max features for some columns, also we did apply diff 1,2 lags as well .After pay features were also helpful .\n- We did usual aggregates std,mean,min,max for categoricals (nunique,count,first,mean). We also tried MAD (Mean Absolute deviation) and it did help our cv though Public LB was bit less so we didn't include the model based on mad in final scores though it could have helped in hindsight ,\n- We also did add percentage change features , which basically do the percentage change on numerical features and thought they would work better than just a numeric diff . They did help us in some models .\n- Also one feature that was there in most of our good models was keeping last, first and middle component of the statements as we hypothesized that will help us cover most variation for a customer as we had aggregated features .\n- Another feature we tried was to say calculate spend/balance or spend_sum/balance_sum ratios , It did help in some models but not most .\n- Trimming of non impactful features also did help to reduce number of features in 1k-2k range .\n\n## Hybrid approach (meta+aggs) : \nWe also gave a shot at mixing the flattened meta features with aggregates and it did help our models specially LGBM CAT XG and Tabnet.\nCause of meta features this usually converged faster (fewer rounds with early stopping in XG) so we had to **Lower LR**  for this approach .\n\n#Models\nAll models were trained on meta features and/or engineered aggregates . NN were scaled for the most part with GaussianScalar\n\n**LGBM (Max cv 0.79932 Private 0.80731)**: This was as with most based on notebook by @ragnar123 . We did do little tweaks to the hyperparams but most were similar .Here as noted in discussions lowering the LR certainly helped in the case of DART and also us adding meta features (around *0.0075 LR*).\n**XG (Max CV 0.7984 Priv. 0.80687)**: XG was based on @jiweiliu  great notebook\n**CAT (Max CV 0.7972)**: Cat we did tune a bit with below params and it helped us a lot .\n```python\nCatBoostClassifier( random_state=CFG.seed,  \n                                bootstrap_type='Bernoulli',\n                                task_type=\"GPU\",\n                                devices='0:1',  \n                                use_best_model = True, \n                                iterations = 11500,\n                                num_leaves =64,\n                                subsample = 0.74,\n                                grow_policy = 'Lossguide', \n                                depth = 9) \n```\n**Tabnet (MAX CV 0.793133 Private 0.80447)**: Tabnet used mostly standard params and did give better results with meta features .\n**MLP (Private 0.80109).** We tried models with tabnet using cat embeddings and without (by onehotencoding). Also **GaussianScaler** helped with scaling of numerical features in case of NN and tabnet\n**AutoML (MAX CV 0.79714)**: @liji11 tried automl which did help our ensemble \n**Tabformer ( MAX CV 0.79507 Private 0.80480)** : This mostly based on the public [notebook](https://www.kaggle.com/code/gauravbrills/tabtransformer-training)  with more heads as per the paper . This we noticed was better than tabnet in private LB but we didnt choose it as was not working well in ensemble 😑. Hyperparams below \n```python\nNUM_TRANSFORMER_BLOCKS = 6  # Number of transformer blocks. 6 paper recommends\nNUM_HEADS = 8  # Number of attention heads. 8 Heads paper recommends\nEMBEDDING_DIMS = 16#10  # Embedding dimensions of the categorical features. check 16,32\nDROPOUT_RATE = 0.1\nMLP_HIDDEN_UNITS_FACTORS = [\n    4,\n    2,\n]  # MLP hidden layer units, as factors of the number of inputs. =>(4,2)\nNUM_MLP_BLOCKS = 4  # Number of MLP blocks in the baseline model. Paper 4\nMLP_ACTIVATION = keras.activations.selu\n```\n\n**TCN MLP**: TCN over 13 sequences , Not chosen Max CV 0.78\nGoes like this inspired by the amex paper \n```python\n   embeddings = []\n    for k in range(11): \n      vocabulary = CATEGORICAL_FEATURES_WITH_VOCABULARY[cat_cols[k]]\n      #print(f\"cat {cat_cols[k]} index {k} len {len(vocabulary)}\")\n      emb = tf.keras.layers.Embedding(len(vocabulary),EMBEDDING_DIMS)\n      embeddings.append(emb(inputs[:,:,k]))\n    in_ = tf.keras.layers.Concatenate()([inputs[:,:,11:]]+embeddings)  \n    activation = 'swish'\n    l1 = 1e-7\n    l2 = 4e-4\n    reg = 4e-4\n    # SIMPLE Wavenet TCN BACKBONE  \n    _x = BatchNormalization()(in_)\n    x = TCN(nb_filters=256, kernel_size=4,return_sequences=False, dropout_rate=0.0, dilations=[2 ** i for i in range(9)])(_x)\n    #x = TCN(nb_filters=128, kernel_size=3,return_sequences=False,use_layer_norm=True,dropout_rate=0.05,dilations=[2 ** i for i in range(7)])(x)\n    x0 = Dense(128, \n               kernel_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n#                activity_regularizer=tf.keras.regularizers.L1L2(l1=l1,l2=l2),\n              activation=activation,\n             )(x)\n    x0 = Dropout(0.1)(x0)\n... MORE OF MLP ....\nx_output = Dense(1,\n              activation='sigmoid',\n             )(x)\n```\n\n# Ensemble techniques\n\nAll ensembles were done on log odds as we found it worked better when ensembling NN models . We had also done ensemble with rank method but that didnt work well when ensembling with NN .\n\nWe tried a bunch of meta ensemble techniques with are many models .The ones that worked were [forward selection](https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private) by @cdeotte , Optuna weights and Elasticnet on top of oof preds . We also tried stacking and voting techniques but they weren't very successful.\n\n# What did not work or DID ?\n\n- [Tabformer](https://www.kaggle.com/code/gauravbrills/tabtransformer-training) : We finally made tabformer reach 0.795 cv but it was not part of our final submission. This I will say *did work* but did not ensemble well though later we realized was our best Private LB scoring NN .\n- TCN wavenet : We tried to use 13 customer history sequences in a tcn wavenet + MLP implementation but results were only around 0.78 ish so we dropped the idea .\n- Saint and widedeep models : We also gave a stab on wide deep package to try out  [SAINT](https://arxiv.org/abs/2106.01342)  but were not able to get far with that . reference [pytorch-widedeep](https://github.com/jrzaurin/pytorch-widedeep)\n- Some features as described above and some dropped based on permutation or zero feature importance .\n- We tried to do multi seed ensembles for all our good models . This somehow did not give us good results at the end compared to sticking with 42 seed.\n- Sirius also tried Target encoding at the end but we had really very little time to check on these in ensemble .\n- We also tried GAN techniques to impute missing values but dropped them at the end .\n\nFinally luckily we selected a good enough final submission for 🏅(There was as usual a lot of confusion thanks @tonymarkchris for voting for this 😄), though seems some with lower LB were better so should have trusted cv a bit more :) \n\n**✅  Best selected submission Private LB 0.80838 ( Optuna weighted models , xg, tabnet.mlp,lgbm,cat and automl)**\n**📮 Best submission : Private LB 0.80852 ( Optuna xg,cat,tabnet,mlp,automl and lgbm)**\n\nThanks for reading 😄",
    "1914353": "This is my first time to join a team and the teamwork is smooth. Thank my teammates for your interesting ideas and hard working! \nAs the making detail of **meta features** has beed described in this [thread](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347786) and I didn't know the so-called meta feature before, here I just give some motivation of coming up with these features.\n- In the most of the public notebooks, features are based on agg methods, such as min, max, mean and std.\n- Though these agg features are expressive, some information is losing during the 13 statements being compressed into limited statistics.\n- I just thought about **how can I utilise the information of each one statement instead of aggregating them.**\n- Flatting the data from [N_USER, 13, 188] into [N_USER, 13*188] is a straight method, but I feel it will increase the feature dim too much. \n- Then I came up with the idea of predicting a default score of each statement, **which can be regarded as a distillation of each statement**.\n- These 13 predicted scores are the so-called meta features , which are used as numeric features of each customer in the following models.\n- Finally, the meta features show a great improvement in my experiments, like\n\t- tabnet cv +0.003, 0.791->0.794 (lb 0.796)\n\t- xgb cv +0.0015, 0.7953->0.7968 (lb 0.798)\n\t- lgb cv +0.000x (not sure)",
    "1915387": "Great teammates👍",
    "1915404": "Yes @sirius81 your features were the 🪄 for us .. Thanks 👍",
    "1915405": "Yess 🙌 @liji11",
    "1915863": "👌It's great"
  },
  "source": "meta"
}