{
  "id": 508124,
  "title": "57th place, 0.528. Without hack.",
  "url": "/competitions/home-credit-credit-risk-model-stability/writeups/andrey-nesterov-57th-place-0-528-without-hack",
  "author_name": "",
  "post_date": "2024-06-05T12:18:26.263Z",
  "votes": 24,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Here is my efforts and final ensemble for this competition.</p>\n<table>\n<thead>\n<tr>\n<th>model*</th>\n<th>Features Qt.</th>\n<th>CV</th>\n<th>PublicLB</th>\n<th>PrivateLB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>catboost</td>\n<td>551</td>\n<td>0.688972</td>\n<td>0.589</td>\n<td>0.509</td>\n</tr>\n<tr>\n<td>lgbm</td>\n<td>697</td>\n<td>0.692120</td>\n<td>0.587</td>\n<td>0.505</td>\n</tr>\n<tr>\n<td>xgb</td>\n<td>718</td>\n<td>0.702661</td>\n<td>0.581</td>\n<td>0.504</td>\n</tr>\n<tr>\n<td>lgbm_no-dates</td>\n<td>601</td>\n<td>0.648829</td>\n<td>0.568</td>\n<td>0.494</td>\n</tr>\n<tr>\n<td>catboost_no-dates</td>\n<td>320</td>\n<td>0.632668</td>\n<td>0.566</td>\n<td>0.490</td>\n</tr>\n<tr>\n<td>lama_mlp</td>\n<td>433</td>\n<td>0.671007</td>\n<td>0.573</td>\n<td>0.488</td>\n</tr>\n<tr>\n<td>1dcnn</td>\n<td>570</td>\n<td>0.687478</td>\n<td>0.562</td>\n<td>0.476</td>\n</tr>\n<tr>\n<td>L3 ensemble</td>\n<td>-</td>\n<td>0.689244</td>\n<td>0.605</td>\n<td>0.528</td>\n</tr>\n</tbody>\n</table>\n<p>* All models are 5-fold ensembles using the same StratifiedGroupKFold.<br>\n<br><br>\nMy initial sub-goal was to create a good enough diversity of models and datasets for the final ensemble. It's a common approach and I worked on it during the whole competition. Unfortunately, I didn't spend enough time on metric hacking, so my final position is not so good (only 62nd), although I was in top-10 before starting metric hacking.<br>\nMy <a href=\"https://www.kaggle.com/code/andreynesterov/home-credit-inference-final/notebook\" target=\"_blank\">final ensemble</a> without metric hacking has a score of 0.528. The public metric hacking approach with adv classifier simply didn't work for PrivateLB.</p>\n<p><br><br>\nWhat helped improve the models:</p>\n<ol>\n<li>Sorting by numgroups before aggregations - at least it worked for \"first\" and \"last\" aggregators. For some models it was used with combination of aggregation by numgroup2 and then by numgroup1 (after that some tables can also be concatenated, for example applprev, credit_bureau_a, credit_bureau_b).</li>\n<li>Using a 10-fold meta-classifier (LinearRegression) instead of a simple weighted ensemble improved the score to about +0.002. Also, while training meta_cls, I added some noise to the OOF predictions of some models to match the LB scores.</li>\n<li>Models using datasets without date columns - the scores of these models weren't as big, but they complemented the ensemble well enough.</li>\n<li>Using \"Coefficient of Variations\" aggregator - it helped a lot to improve the score of the CatBoost model.</li>\n<li>LAMA-MLP model improved initial ensemble in early stages of competition, and I spent some time to select features to overcome OOM errors and overfitting.</li>\n<li>One place where I directly incorporated the stability approach was using the metric for ReduceLRonPlateau for training MLP and 1DCNN models.</li>\n<li>Count encoding of categorical columns (from public notebooks)</li>\n<li>Filtering of correlated columns (from public notebooks).</li>\n<li>Collapse rare categories in one group.</li>\n<li>To overcome OOM errors, creating separate datasets and then merging them helped. Also memory issue was a big problem for 1DCNN model, but I overcame it by training model on TPU and then inferencing on GPU. Also TPU was used for permutation importance.</li>\n<li>CatBoost model has best private and public scores among other models. Also in one of the experiments 5-fold CatBoost achieved score 0.520 (difference was on using early stopping 100 and lambda 10) but unfortunately I didn't include that model in the final ensemble (LB was the same and CV I didn't know because that was inference training).</li>\n<li>Tuning some models using Optuna.</li>\n<li>The inclusion of the XGB model was relatively important. I also had problems training this model, because (as I understand later) it is not so good at categorical features (unlike LGBM and CatBoost), and preprocessing was critical for this model. More specifically, after using TargetEncoder (and then CatBoostEncoder), the XGB score increased to good enough values. The same approach was also used for 1DCNN model.</li>\n<li>Using slightly different features for all the models (some tables, aggregators, types of columns, filtering approaches were not included for all the models).</li>\n<li>Replacing NaN in num columns with 0 helped for 1DCNN model.</li>\n<li>Aggregators used: max, min, first, last, mean, var, sum, std plus coef_var (distributed differently among all models, column types, or used only for certain columns).</li>\n</ol>\n<p><br><br>\nWhat didn't work (at least in my experiments):</p>\n<ol>\n<li>Feature selection based on adversarial validation (within train ds), permutation importance (including stability metric).</li>\n<li>I tried to create some loss functions (based on functions simpler than initial metric - from topics about metrics; incorporating weights and so on) for GBDT models, but without improvements.</li>\n<li>CV strategies based on splitting train ds on week numbers using distributions of some statistics.</li>\n<li>Tuning GBDT models using some specific parameters like dart for LGBM, ordered boosting_type for CatBoost and so on.</li>\n<li>Using IsolationForest with SMOTENC</li>\n<li>Different scalers for 1DCNN model like PowerTransformer, QuantileTransformer. Finally I used Log1p transformation only for columns with max_min-mean ratio greater than 100.</li>\n<li>Replace nums NaN with mean (or constant values like -99). BTW LAMA pipeline uses this approach, and I tried to use 0s, but LB didn't improve and I rejected this idea, but after deadline I saw that private score was improved for that submission.</li>\n<li>Compress some categories (e.g. by template P0_111_121 to P0 or P0_111)</li>\n<li>Some non-standard, custom aggregators such as entropy, mean-to-max, zero frequency, normalization within group, and so on. The goal was also to find aggragators with relative, not absolute, values.</li>\n<li>Additional feature engineering like ratios between some columns didn't improve models. As I understand after first competition was created very good features and squeezing out from dataset additional gains was not so simple task.</li>\n<li>By the results of previous competition I tried Regularized Greedy Forest (FastRGF version) but it has very poor performance and worsened final ensemble. Another idea from that competition was to train MLP only on some group of tables, but again without any results, so I discarded this idea as well. I also experimented with other models like DAE, TabTransformer, NN-cat_embeddings, packages like pytorch_tabular, but also I can't train them to good enough results.</li>\n<li>Combining some categorical features (similar to how CatBoost works).</li>\n<li>Balancing classes using parameters like scale_pos_weight.</li>\n<li>Include 1DCNN in the final ensemble. I struggled a lot to overcome the overfitting of this model, but eventually it is worsening the final ensemble from 0.529 to 0.528. Some variations of this model (with another ds) has relatively good private score (0.485) but LB was not so good (only 0.555), so I didn't include the best and maybe was digging in the wrong direction. Besides, without metric hacking this model will not contribute much to the final ensemble/position.</li>\n<li>Include some additional data such as FedFunds.</li>\n<li>Filter some weeks.</li>\n</ol>\n<p><br><br>\nOf course ideas from the community, discussions, public notebooks helped a lot.</p>\n<p><br><br>\nUpdate. 29.05.2024<br>\nJust out of curiosity, I tried adding the CatBoost model with score 0.520 to the final ensemble and the total private score increased to 0.536.</p>\n<p>Update. 5.06.2024<br>\nFeature importance by gain I used in selecting top-20 features from credit_bureau_a tables for lama-mlp model. This helped to avoid overfitting this model.</p>",
  "messages": [
    {
      "id": "2840979",
      "postDate": "05/28/2024 10:39:34",
      "content": "<p>Here is my efforts and final ensemble for this competition.</p>\n<table>\n<thead>\n<tr>\n<th>model*</th>\n<th>Features Qt.</th>\n<th>CV</th>\n<th>PublicLB</th>\n<th>PrivateLB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>catboost</td>\n<td>551</td>\n<td>0.688972</td>\n<td>0.589</td>\n<td>0.509</td>\n</tr>\n<tr>\n<td>lgbm</td>\n<td>697</td>\n<td>0.692120</td>\n<td>0.587</td>\n<td>0.505</td>\n</tr>\n<tr>\n<td>xgb</td>\n<td>718</td>\n<td>0.702661</td>\n<td>0.581</td>\n<td>0.504</td>\n</tr>\n<tr>\n<td>lgbm_no-dates</td>\n<td>601</td>\n<td>0.648829</td>\n<td>0.568</td>\n<td>0.494</td>\n</tr>\n<tr>\n<td>catboost_no-dates</td>\n<td>320</td>\n<td>0.632668</td>\n<td>0.566</td>\n<td>0.490</td>\n</tr>\n<tr>\n<td>lama_mlp</td>\n<td>433</td>\n<td>0.671007</td>\n<td>0.573</td>\n<td>0.488</td>\n</tr>\n<tr>\n<td>1dcnn</td>\n<td>570</td>\n<td>0.687478</td>\n<td>0.562</td>\n<td>0.476</td>\n</tr>\n<tr>\n<td>L3 ensemble</td>\n<td>-</td>\n<td>0.689244</td>\n<td>0.605</td>\n<td>0.528</td>\n</tr>\n</tbody>\n</table>\n<p>* All models are 5-fold ensembles using the same StratifiedGroupKFold.<br>\n<br><br>\nMy initial sub-goal was to create a good enough diversity of models and datasets for the final ensemble. It's a common approach and I worked on it during the whole competition. Unfortunately, I didn't spend enough time on metric hacking, so my final position is not so good (only 62nd), although I was in top-10 before starting metric hacking.<br>\nMy <a href=\"https://www.kaggle.com/code/andreynesterov/home-credit-inference-final/notebook\" target=\"_blank\">final ensemble</a> without metric hacking has a score of 0.528. The public metric hacking approach with adv classifier simply didn't work for PrivateLB.</p>\n<p><br><br>\nWhat helped improve the models:</p>\n<ol>\n<li>Sorting by numgroups before aggregations - at least it worked for \"first\" and \"last\" aggregators. For some models it was used with combination of aggregation by numgroup2 and then by numgroup1 (after that some tables can also be concatenated, for example applprev, credit_bureau_a, credit_bureau_b).</li>\n<li>Using a 10-fold meta-classifier (LinearRegression) instead of a simple weighted ensemble improved the score to about +0.002. Also, while training meta_cls, I added some noise to the OOF predictions of some models to match the LB scores.</li>\n<li>Models using datasets without date columns - the scores of these models weren't as big, but they complemented the ensemble well enough.</li>\n<li>Using \"Coefficient of Variations\" aggregator - it helped a lot to improve the score of the CatBoost model.</li>\n<li>LAMA-MLP model improved initial ensemble in early stages of competition, and I spent some time to select features to overcome OOM errors and overfitting.</li>\n<li>One place where I directly incorporated the stability approach was using the metric for ReduceLRonPlateau for training MLP and 1DCNN models.</li>\n<li>Count encoding of categorical columns (from public notebooks)</li>\n<li>Filtering of correlated columns (from public notebooks).</li>\n<li>Collapse rare categories in one group.</li>\n<li>To overcome OOM errors, creating separate datasets and then merging them helped. Also memory issue was a big problem for 1DCNN model, but I overcame it by training model on TPU and then inferencing on GPU. Also TPU was used for permutation importance.</li>\n<li>CatBoost model has best private and public scores among other models. Also in one of the experiments 5-fold CatBoost achieved score 0.520 (difference was on using early stopping 100 and lambda 10) but unfortunately I didn't include that model in the final ensemble (LB was the same and CV I didn't know because that was inference training).</li>\n<li>Tuning some models using Optuna.</li>\n<li>The inclusion of the XGB model was relatively important. I also had problems training this model, because (as I understand later) it is not so good at categorical features (unlike LGBM and CatBoost), and preprocessing was critical for this model. More specifically, after using TargetEncoder (and then CatBoostEncoder), the XGB score increased to good enough values. The same approach was also used for 1DCNN model.</li>\n<li>Using slightly different features for all the models (some tables, aggregators, types of columns, filtering approaches were not included for all the models).</li>\n<li>Replacing NaN in num columns with 0 helped for 1DCNN model.</li>\n<li>Aggregators used: max, min, first, last, mean, var, sum, std plus coef_var (distributed differently among all models, column types, or used only for certain columns).</li>\n</ol>\n<p><br><br>\nWhat didn't work (at least in my experiments):</p>\n<ol>\n<li>Feature selection based on adversarial validation (within train ds), permutation importance (including stability metric).</li>\n<li>I tried to create some loss functions (based on functions simpler than initial metric - from topics about metrics; incorporating weights and so on) for GBDT models, but without improvements.</li>\n<li>CV strategies based on splitting train ds on week numbers using distributions of some statistics.</li>\n<li>Tuning GBDT models using some specific parameters like dart for LGBM, ordered boosting_type for CatBoost and so on.</li>\n<li>Using IsolationForest with SMOTENC</li>\n<li>Different scalers for 1DCNN model like PowerTransformer, QuantileTransformer. Finally I used Log1p transformation only for columns with max_min-mean ratio greater than 100.</li>\n<li>Replace nums NaN with mean (or constant values like -99). BTW LAMA pipeline uses this approach, and I tried to use 0s, but LB didn't improve and I rejected this idea, but after deadline I saw that private score was improved for that submission.</li>\n<li>Compress some categories (e.g. by template P0_111_121 to P0 or P0_111)</li>\n<li>Some non-standard, custom aggregators such as entropy, mean-to-max, zero frequency, normalization within group, and so on. The goal was also to find aggragators with relative, not absolute, values.</li>\n<li>Additional feature engineering like ratios between some columns didn't improve models. As I understand after first competition was created very good features and squeezing out from dataset additional gains was not so simple task.</li>\n<li>By the results of previous competition I tried Regularized Greedy Forest (FastRGF version) but it has very poor performance and worsened final ensemble. Another idea from that competition was to train MLP only on some group of tables, but again without any results, so I discarded this idea as well. I also experimented with other models like DAE, TabTransformer, NN-cat_embeddings, packages like pytorch_tabular, but also I can't train them to good enough results.</li>\n<li>Combining some categorical features (similar to how CatBoost works).</li>\n<li>Balancing classes using parameters like scale_pos_weight.</li>\n<li>Include 1DCNN in the final ensemble. I struggled a lot to overcome the overfitting of this model, but eventually it is worsening the final ensemble from 0.529 to 0.528. Some variations of this model (with another ds) has relatively good private score (0.485) but LB was not so good (only 0.555), so I didn't include the best and maybe was digging in the wrong direction. Besides, without metric hacking this model will not contribute much to the final ensemble/position.</li>\n<li>Include some additional data such as FedFunds.</li>\n<li>Filter some weeks.</li>\n</ol>\n<p><br><br>\nOf course ideas from the community, discussions, public notebooks helped a lot.</p>\n<p><br><br>\nUpdate. 29.05.2024<br>\nJust out of curiosity, I tried adding the CatBoost model with score 0.520 to the final ensemble and the total private score increased to 0.536.</p>\n<p>Update. 5.06.2024<br>\nFeature importance by gain I used in selecting top-20 features from credit_bureau_a tables for lama-mlp model. This helped to avoid overfitting this model.</p>",
      "rawMarkdown": "Here is my efforts and final ensemble for this competition.\n\n| model* | Features Qt.\t| CV\t| PublicLB\t| PrivateLB |\n| ---| --- | --- | --- | --- |\n| catboost\t| 551\t| 0.688972\t| 0.589\t| 0.509 |\n| lgbm\t| 697\t\t| 0.692120\t| 0.587\t| 0.505 |\n| xgb\t\t| 718\t\t| 0.702661\t| 0.581\t| 0.504 |\n| lgbm_no-dates\t| 601\t| 0.648829\t| 0.568\t| 0.494 |\n| catboost_no-dates\t| 320\t\t| 0.632668\t| 0.566\t | 0.490 |\n| lama_mlp\t| 433\t\t| 0.671007\t| 0.573\t| 0.488 |\n| 1dcnn\t| 570\t\t| 0.687478\t| 0.562\t| 0.476 |\n| L3 ensemble\t\t| -\t| 0.689244\t| 0.605\t| 0.528 |\n\n\\* All models are 5-fold ensembles using the same StratifiedGroupKFold.\n<br>\nMy initial sub-goal was to create a good enough diversity of models and datasets for the final ensemble. It's a common approach and I worked on it during the whole competition. Unfortunately, I didn't spend enough time on metric hacking, so my final position is not so good (only 62nd), although I was in top-10 before starting metric hacking.\nMy [final ensemble](https://www.kaggle.com/code/andreynesterov/home-credit-inference-final/notebook) without metric hacking has a score of 0.528. The public metric hacking approach with adv classifier simply didn't work for PrivateLB.\n\n<br>\nWhat helped improve the models:\n1. Sorting by numgroups before aggregations - at least it worked for \"first\" and \"last\" aggregators. For some models it was used with combination of aggregation by numgroup2 and then by numgroup1 (after that some tables can also be concatenated, for example applprev, credit_bureau_a, credit_bureau_b).\n1. Using a 10-fold meta-classifier (LinearRegression) instead of a simple weighted ensemble improved the score to about +0.002. Also, while training meta_cls, I added some noise to the OOF predictions of some models to match the LB scores.\n1. Models using datasets without date columns - the scores of these models weren't as big, but they complemented the ensemble well enough.\n1. Using \"Coefficient of Variations\" aggregator - it helped a lot to improve the score of the CatBoost model.\n1. LAMA-MLP model improved initial ensemble in early stages of competition, and I spent some time to select features to overcome OOM errors and overfitting.\n1. One place where I directly incorporated the stability approach was using the metric for ReduceLRonPlateau for training MLP and 1DCNN models.\n1. Count encoding of categorical columns (from public notebooks)\n1. Filtering of correlated columns (from public notebooks).\n1. Collapse rare categories in one group.\n1. To overcome OOM errors, creating separate datasets and then merging them helped. Also memory issue was a big problem for 1DCNN model, but I overcame it by training model on TPU and then inferencing on GPU. Also TPU was used for permutation importance.\n1. CatBoost model has best private and public scores among other models. Also in one of the experiments 5-fold CatBoost achieved score 0.520 (difference was on using early stopping 100 and lambda 10) but unfortunately I didn't include that model in the final ensemble (LB was the same and CV I didn't know because that was inference training).\n1. Tuning some models using Optuna.\n1. The inclusion of the XGB model was relatively important. I also had problems training this model, because (as I understand later) it is not so good at categorical features (unlike LGBM and CatBoost), and preprocessing was critical for this model. More specifically, after using TargetEncoder (and then CatBoostEncoder), the XGB score increased to good enough values. The same approach was also used for 1DCNN model.\n1. Using slightly different features for all the models (some tables, aggregators, types of columns, filtering approaches were not included for all the models).\n1. Replacing NaN in num columns with 0 helped for 1DCNN model.\n1. Aggregators used: max, min, first, last, mean, var, sum, std plus coef_var (distributed differently among all models, column types, or used only for certain columns).\n\n<br>\nWhat didn't work (at least in my experiments):\n1. Feature selection based on adversarial validation (within train ds), permutation importance (including stability metric).\n1. I tried to create some loss functions (based on functions simpler than initial metric - from topics about metrics; incorporating weights and so on) for GBDT models, but without improvements.\n1. CV strategies based on splitting train ds on week numbers using distributions of some statistics.\n1. Tuning GBDT models using some specific parameters like dart for LGBM, ordered boosting_type for CatBoost and so on.\n1. Using IsolationForest with SMOTENC\n1. Different scalers for 1DCNN model like PowerTransformer, QuantileTransformer. Finally I used Log1p transformation only for columns with max_min-mean ratio greater than 100.\n1. Replace nums NaN with mean (or constant values like -99). BTW LAMA pipeline uses this approach, and I tried to use 0s, but LB didn't improve and I rejected this idea, but after deadline I saw that private score was improved for that submission.\n1. Compress some categories (e.g. by template P0_111_121 to P0 or P0_111)\n1. Some non-standard, custom aggregators such as entropy, mean-to-max, zero frequency, normalization within group, and so on. The goal was also to find aggragators with relative, not absolute, values.\n1. Additional feature engineering like ratios between some columns didn't improve models. As I understand after first competition was created very good features and squeezing out from dataset additional gains was not so simple task.\n1. By the results of previous competition I tried Regularized Greedy Forest (FastRGF version) but it has very poor performance and worsened final ensemble. Another idea from that competition was to train MLP only on some group of tables, but again without any results, so I discarded this idea as well. I also experimented with other models like DAE, TabTransformer, NN-cat_embeddings, packages like pytorch_tabular, but also I can't train them to good enough results.\n1. Combining some categorical features (similar to how CatBoost works).\n1. Balancing classes using parameters like scale_pos_weight.\n1. Include 1DCNN in the final ensemble. I struggled a lot to overcome the overfitting of this model, but eventually it is worsening the final ensemble from 0.529 to 0.528. Some variations of this model (with another ds) has relatively good private score (0.485) but LB was not so good (only 0.555), so I didn't include the best and maybe was digging in the wrong direction. Besides, without metric hacking this model will not contribute much to the final ensemble/position.\n1. Include some additional data such as FedFunds.\n1. Filter some weeks.\n\n<br>\nOf course ideas from the community, discussions, public notebooks helped a lot.\n\n<br>\nUpdate. 29.05.2024\nJust out of curiosity, I tried adding the CatBoost model with score 0.520 to the final ensemble and the total private score increased to 0.536.\n\nUpdate. 5.06.2024\nFeature importance by gain I used in selecting top-20 features from credit_bureau_a tables for lama-mlp model. This helped to avoid overfitting this model.",
      "votes": null
    },
    {
      "id": "2840985",
      "postDate": "05/28/2024 10:50:06",
      "content": "<p>What aggregates did you use for numeric columns?</p>",
      "rawMarkdown": "What aggregates did you use for numeric columns?",
      "votes": null
    },
    {
      "id": "2841016",
      "postDate": "05/28/2024 11:04:38",
      "content": "<p>I used pretty common aggregators: max, min, first, last, mean, var, sum, std plus coef_var. But again, each dataset has slightly different aggregators, for example, I didn't include var method to ds with coef_var. Sum aggregator I used only in some datasets and only for certain columns (with enough diversity - I selected them manually).</p>",
      "rawMarkdown": "I used pretty common aggregators: max, min, first, last, mean, var, sum, std plus coef_var. But again, each dataset has slightly different aggregators, for example, I didn't include var method to ds with coef_var. Sum aggregator I used only in some datasets and only for certain columns (with enough diversity - I selected them manually).",
      "votes": null
    },
    {
      "id": "2841041",
      "postDate": "05/28/2024 11:24:44",
      "content": "<p>Thanks! Have you experimented with diff features, namely last-mean, last-first?</p>",
      "rawMarkdown": "Thanks! Have you experimented with diff features, namely last-mean, last-first?",
      "votes": null
    },
    {
      "id": "2841062",
      "postDate": "05/28/2024 11:34:01",
      "content": "<p>No, unfortunately.</p>",
      "rawMarkdown": "No, unfortunately.",
      "votes": null
    },
    {
      "id": "2841208",
      "postDate": "05/28/2024 12:57:24",
      "content": "<p>Thank u for share .  You mentioned \"Feature selection based on importance by gain\"  didn't work.  Did importance by split work?</p>",
      "rawMarkdown": "Thank u for share .  You mentioned \"Feature selection based on importance by gain\"  didn't work.  Did importance by split work?",
      "votes": null
    },
    {
      "id": "2841214",
      "postDate": "05/28/2024 13:05:21",
      "content": "<p>Yes, I selected features for MLP based on importance by split of LGBM, but importance by gain in my cases led to overfitting.</p>",
      "rawMarkdown": "Yes, I selected features for MLP based on importance by split of LGBM, but importance by gain in my cases led to overfitting.",
      "votes": null
    },
    {
      "id": "2841222",
      "postDate": "05/28/2024 13:16:01",
      "content": "<p>Hi Andrey, congrats with the 62nd place! A lot of work is done!</p>\n<p>There are a few things that helped me to boost my score quite a bit are the following:</p>\n<ul>\n<li>Aggregation of numerical columns: last 3 rows tail(3).mean(); as well as diff() with mean and/or std</li>\n<li>Many features created by hand using logical assumptions</li>\n<li>The addition of stability metric had a minor effect</li>\n<li>Max aggregations of string columns had a big effect. However, I rejected this due to a problem of interpretability with the hope that this is an artifact of the public LB test sample.</li>\n</ul>\n<p>Just curious, have you tried/considered such tail/diff aggregations to add to your data?</p>\n<p>I am still very puzzled to why the class balance does not work in this competition. When training my models, I even balanced the classes versus week number trying to exclude the possibility for models to learn any discrimination from this dimension ( as the WEEK_NUM is correlated with some inputs) with the hope this can enhance the model stability. </p>",
      "rawMarkdown": "Hi Andrey, congrats with the 62nd place! A lot of work is done!\n\nThere are a few things that helped me to boost my score quite a bit are the following:\n - Aggregation of numerical columns: last 3 rows tail(3).mean(); as well as diff() with mean and/or std\n - Many features created by hand using logical assumptions\n - The addition of stability metric had a minor effect\n - Max aggregations of string columns had a big effect. However, I rejected this due to a problem of interpretability with the hope that this is an artifact of the public LB test sample.\n\nJust curious, have you tried/considered such tail/diff aggregations to add to your data?\n\nI am still very puzzled to why the class balance does not work in this competition. When training my models, I even balanced the classes versus week number trying to exclude the possibility for models to learn any discrimination from this dimension ( as the WEEK_NUM is correlated with some inputs) with the hope this can enhance the model stability.",
      "votes": null
    },
    {
      "id": "2841270",
      "postDate": "05/28/2024 13:48:16",
      "content": "<p>Hi Oleh. Thanks.<br>\nI was thinking about some kind of tail(n).some_agg() (I also read mention about it in previous competition solutions and discussions), but I think this idea got lost somewhere in my notes. I didn't think about diff.<br>\nAbout max agg for string columns - it may work because it contains some hierarchy in the some categories encoding (e.g. P0_111_121).</p>",
      "rawMarkdown": "Hi Oleh. Thanks.\nI was thinking about some kind of tail(n).some_agg() (I also read mention about it in previous competition solutions and discussions), but I think this idea got lost somewhere in my notes. I didn't think about diff.\nAbout max agg for string columns - it may work because it contains some hierarchy in the some categories encoding (e.g. P0_111_121).",
      "votes": null
    },
    {
      "id": "2841285",
      "postDate": "05/28/2024 13:55:45",
      "content": "<p>You gain a lot from ensembling. <br>\nI have 0.598 single model and only 0.6 ensemble</p>",
      "rawMarkdown": "You gain a lot from ensembling. \nI have 0.598 single model and only 0.6 ensemble",
      "votes": null
    },
    {
      "id": "2841298",
      "postDate": "05/28/2024 14:06:42",
      "content": "<p>Yes, that was a kind of exercise from my previous competitions. Of course, that's only one of the approaches, and a strong enough single model can beat a good ensemble.</p>",
      "rawMarkdown": "Yes, that was a kind of exercise from my previous competitions. Of course, that's only one of the approaches, and a strong enough single model can beat a good ensemble.",
      "votes": null
    },
    {
      "id": "2841342",
      "postDate": "05/28/2024 14:36:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/andreynesterov\" target=\"_blank\">@andreynesterov</a>, thank you for sharing. I read that DART did not work well, but in my case it has improved the result and reach almost the catboost performance.</p>",
      "rawMarkdown": "Hi @andreynesterov, thank you for sharing. I read that DART did not work well, but in my case it has improved the result and reach almost the catboost performance.",
      "votes": null
    },
    {
      "id": "2841364",
      "postDate": "05/28/2024 14:56:47",
      "content": "<p>That's great. I also thought dart was very promising, maybe I didn't pay much attention to tuning it or preparing features.</p>",
      "rawMarkdown": "That's great. I also thought dart was very promising, maybe I didn't pay much attention to tuning it or preparing features.",
      "votes": null
    },
    {
      "id": "2841461",
      "postDate": "05/28/2024 15:31:22",
      "content": "<p>Thank you for the detailed insights on your findings. I have been following your work from the beginning of this competition and a lot of my work is inspired from your notebooks. Thank you for generously sharing those. Beginners like me learn a lot from your notebooks.</p>\n<p>I especially referred to the notebook where you ensembled xgb, cat, lgb model. I took the lessons from that notebook and tried ensembling 2 lgb and 2 cat - one of each with class balancing and other without the balancing and that did well. But, my best one which i did not submit was the one which i added a post-processing layer of manual implementation of isotonic calibration. it was a 0.593 on public LB and 0.527 on private LB. </p>",
      "rawMarkdown": "Thank you for the detailed insights on your findings. I have been following your work from the beginning of this competition and a lot of my work is inspired from your notebooks. Thank you for generously sharing those. Beginners like me learn a lot from your notebooks.\n\nI especially referred to the notebook where you ensembled xgb, cat, lgb model. I took the lessons from that notebook and tried ensembling 2 lgb and 2 cat - one of each with class balancing and other without the balancing and that did well. But, my best one which i did not submit was the one which i added a post-processing layer of manual implementation of isotonic calibration. it was a 0.593 on public LB and 0.527 on private LB.",
      "votes": null
    },
    {
      "id": "2841501",
      "postDate": "05/28/2024 15:42:12",
      "content": "<p>That's interesting.</p>",
      "rawMarkdown": "That's interesting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2840985,
      "author_name": "andreasbis",
      "author_url": "",
      "post_date": "05/28/2024 10:50:06",
      "content": "<p>What aggregates did you use for numeric columns?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2841016,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "05/28/2024 11:04:38",
          "content": "<p>I used pretty common aggregators: max, min, first, last, mean, var, sum, std plus coef_var. But again, each dataset has slightly different aggregators, for example, I didn't include var method to ds with coef_var. Sum aggregator I used only in some datasets and only for certain columns (with enough diversity - I selected them manually).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2841041,
              "author_name": "andreasbis",
              "author_url": "",
              "post_date": "05/28/2024 11:24:44",
              "content": "<p>Thanks! Have you experimented with diff features, namely last-mean, last-first?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2841062,
                  "author_name": "andreynesterov",
                  "author_url": "",
                  "post_date": "05/28/2024 11:34:01",
                  "content": "<p>No, unfortunately.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2841208,
      "author_name": "sequoia12",
      "author_url": "",
      "post_date": "05/28/2024 12:57:24",
      "content": "<p>Thank u for share .  You mentioned \"Feature selection based on importance by gain\"  didn't work.  Did importance by split work?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2841214,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "05/28/2024 13:05:21",
          "content": "<p>Yes, I selected features for MLP based on importance by split of LGBM, but importance by gain in my cases led to overfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2841222,
      "author_name": "olehkivernyk",
      "author_url": "",
      "post_date": "05/28/2024 13:16:01",
      "content": "<p>Hi Andrey, congrats with the 62nd place! A lot of work is done!</p>\n<p>There are a few things that helped me to boost my score quite a bit are the following:</p>\n<ul>\n<li>Aggregation of numerical columns: last 3 rows tail(3).mean(); as well as diff() with mean and/or std</li>\n<li>Many features created by hand using logical assumptions</li>\n<li>The addition of stability metric had a minor effect</li>\n<li>Max aggregations of string columns had a big effect. However, I rejected this due to a problem of interpretability with the hope that this is an artifact of the public LB test sample.</li>\n</ul>\n<p>Just curious, have you tried/considered such tail/diff aggregations to add to your data?</p>\n<p>I am still very puzzled to why the class balance does not work in this competition. When training my models, I even balanced the classes versus week number trying to exclude the possibility for models to learn any discrimination from this dimension ( as the WEEK_NUM is correlated with some inputs) with the hope this can enhance the model stability. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2841270,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "05/28/2024 13:48:16",
          "content": "<p>Hi Oleh. Thanks.<br>\nI was thinking about some kind of tail(n).some_agg() (I also read mention about it in previous competition solutions and discussions), but I think this idea got lost somewhere in my notes. I didn't think about diff.<br>\nAbout max agg for string columns - it may work because it contains some hierarchy in the some categories encoding (e.g. P0_111_121).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2841285,
      "author_name": "skrrydg",
      "author_url": "",
      "post_date": "05/28/2024 13:55:45",
      "content": "<p>You gain a lot from ensembling. <br>\nI have 0.598 single model and only 0.6 ensemble</p>",
      "votes": null,
      "replies": [
        {
          "id": 2841298,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "05/28/2024 14:06:42",
          "content": "<p>Yes, that was a kind of exercise from my previous competitions. Of course, that's only one of the approaches, and a strong enough single model can beat a good ensemble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2841342,
      "author_name": "pourchot",
      "author_url": "",
      "post_date": "05/28/2024 14:36:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/andreynesterov\" target=\"_blank\">@andreynesterov</a>, thank you for sharing. I read that DART did not work well, but in my case it has improved the result and reach almost the catboost performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2841364,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "05/28/2024 14:56:47",
          "content": "<p>That's great. I also thought dart was very promising, maybe I didn't pay much attention to tuning it or preparing features.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2841461,
      "author_name": "varuniraothumsi",
      "author_url": "",
      "post_date": "05/28/2024 15:31:22",
      "content": "<p>Thank you for the detailed insights on your findings. I have been following your work from the beginning of this competition and a lot of my work is inspired from your notebooks. Thank you for generously sharing those. Beginners like me learn a lot from your notebooks.</p>\n<p>I especially referred to the notebook where you ensembled xgb, cat, lgb model. I took the lessons from that notebook and tried ensembling 2 lgb and 2 cat - one of each with class balancing and other without the balancing and that did well. But, my best one which i did not submit was the one which i added a post-processing layer of manual implementation of isotonic calibration. it was a 0.593 on public LB and 0.527 on private LB. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2841501,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "05/28/2024 15:42:12",
          "content": "<p>That's interesting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2840979": "Here is my efforts and final ensemble for this competition.\n\n| model* | Features Qt.\t| CV\t| PublicLB\t| PrivateLB |\n| ---| --- | --- | --- | --- |\n| catboost\t| 551\t| 0.688972\t| 0.589\t| 0.509 |\n| lgbm\t| 697\t\t| 0.692120\t| 0.587\t| 0.505 |\n| xgb\t\t| 718\t\t| 0.702661\t| 0.581\t| 0.504 |\n| lgbm_no-dates\t| 601\t| 0.648829\t| 0.568\t| 0.494 |\n| catboost_no-dates\t| 320\t\t| 0.632668\t| 0.566\t | 0.490 |\n| lama_mlp\t| 433\t\t| 0.671007\t| 0.573\t| 0.488 |\n| 1dcnn\t| 570\t\t| 0.687478\t| 0.562\t| 0.476 |\n| L3 ensemble\t\t| -\t| 0.689244\t| 0.605\t| 0.528 |\n\n\\* All models are 5-fold ensembles using the same StratifiedGroupKFold.\n<br>\nMy initial sub-goal was to create a good enough diversity of models and datasets for the final ensemble. It's a common approach and I worked on it during the whole competition. Unfortunately, I didn't spend enough time on metric hacking, so my final position is not so good (only 62nd), although I was in top-10 before starting metric hacking.\nMy [final ensemble](https://www.kaggle.com/code/andreynesterov/home-credit-inference-final/notebook) without metric hacking has a score of 0.528. The public metric hacking approach with adv classifier simply didn't work for PrivateLB.\n\n<br>\nWhat helped improve the models:\n1. Sorting by numgroups before aggregations - at least it worked for \"first\" and \"last\" aggregators. For some models it was used with combination of aggregation by numgroup2 and then by numgroup1 (after that some tables can also be concatenated, for example applprev, credit_bureau_a, credit_bureau_b).\n1. Using a 10-fold meta-classifier (LinearRegression) instead of a simple weighted ensemble improved the score to about +0.002. Also, while training meta_cls, I added some noise to the OOF predictions of some models to match the LB scores.\n1. Models using datasets without date columns - the scores of these models weren't as big, but they complemented the ensemble well enough.\n1. Using \"Coefficient of Variations\" aggregator - it helped a lot to improve the score of the CatBoost model.\n1. LAMA-MLP model improved initial ensemble in early stages of competition, and I spent some time to select features to overcome OOM errors and overfitting.\n1. One place where I directly incorporated the stability approach was using the metric for ReduceLRonPlateau for training MLP and 1DCNN models.\n1. Count encoding of categorical columns (from public notebooks)\n1. Filtering of correlated columns (from public notebooks).\n1. Collapse rare categories in one group.\n1. To overcome OOM errors, creating separate datasets and then merging them helped. Also memory issue was a big problem for 1DCNN model, but I overcame it by training model on TPU and then inferencing on GPU. Also TPU was used for permutation importance.\n1. CatBoost model has best private and public scores among other models. Also in one of the experiments 5-fold CatBoost achieved score 0.520 (difference was on using early stopping 100 and lambda 10) but unfortunately I didn't include that model in the final ensemble (LB was the same and CV I didn't know because that was inference training).\n1. Tuning some models using Optuna.\n1. The inclusion of the XGB model was relatively important. I also had problems training this model, because (as I understand later) it is not so good at categorical features (unlike LGBM and CatBoost), and preprocessing was critical for this model. More specifically, after using TargetEncoder (and then CatBoostEncoder), the XGB score increased to good enough values. The same approach was also used for 1DCNN model.\n1. Using slightly different features for all the models (some tables, aggregators, types of columns, filtering approaches were not included for all the models).\n1. Replacing NaN in num columns with 0 helped for 1DCNN model.\n1. Aggregators used: max, min, first, last, mean, var, sum, std plus coef_var (distributed differently among all models, column types, or used only for certain columns).\n\n<br>\nWhat didn't work (at least in my experiments):\n1. Feature selection based on adversarial validation (within train ds), permutation importance (including stability metric).\n1. I tried to create some loss functions (based on functions simpler than initial metric - from topics about metrics; incorporating weights and so on) for GBDT models, but without improvements.\n1. CV strategies based on splitting train ds on week numbers using distributions of some statistics.\n1. Tuning GBDT models using some specific parameters like dart for LGBM, ordered boosting_type for CatBoost and so on.\n1. Using IsolationForest with SMOTENC\n1. Different scalers for 1DCNN model like PowerTransformer, QuantileTransformer. Finally I used Log1p transformation only for columns with max_min-mean ratio greater than 100.\n1. Replace nums NaN with mean (or constant values like -99). BTW LAMA pipeline uses this approach, and I tried to use 0s, but LB didn't improve and I rejected this idea, but after deadline I saw that private score was improved for that submission.\n1. Compress some categories (e.g. by template P0_111_121 to P0 or P0_111)\n1. Some non-standard, custom aggregators such as entropy, mean-to-max, zero frequency, normalization within group, and so on. The goal was also to find aggragators with relative, not absolute, values.\n1. Additional feature engineering like ratios between some columns didn't improve models. As I understand after first competition was created very good features and squeezing out from dataset additional gains was not so simple task.\n1. By the results of previous competition I tried Regularized Greedy Forest (FastRGF version) but it has very poor performance and worsened final ensemble. Another idea from that competition was to train MLP only on some group of tables, but again without any results, so I discarded this idea as well. I also experimented with other models like DAE, TabTransformer, NN-cat_embeddings, packages like pytorch_tabular, but also I can't train them to good enough results.\n1. Combining some categorical features (similar to how CatBoost works).\n1. Balancing classes using parameters like scale_pos_weight.\n1. Include 1DCNN in the final ensemble. I struggled a lot to overcome the overfitting of this model, but eventually it is worsening the final ensemble from 0.529 to 0.528. Some variations of this model (with another ds) has relatively good private score (0.485) but LB was not so good (only 0.555), so I didn't include the best and maybe was digging in the wrong direction. Besides, without metric hacking this model will not contribute much to the final ensemble/position.\n1. Include some additional data such as FedFunds.\n1. Filter some weeks.\n\n<br>\nOf course ideas from the community, discussions, public notebooks helped a lot.\n\n<br>\nUpdate. 29.05.2024\nJust out of curiosity, I tried adding the CatBoost model with score 0.520 to the final ensemble and the total private score increased to 0.536.\n\nUpdate. 5.06.2024\nFeature importance by gain I used in selecting top-20 features from credit_bureau_a tables for lama-mlp model. This helped to avoid overfitting this model.",
    "2840985": "What aggregates did you use for numeric columns?",
    "2841016": "I used pretty common aggregators: max, min, first, last, mean, var, sum, std plus coef_var. But again, each dataset has slightly different aggregators, for example, I didn't include var method to ds with coef_var. Sum aggregator I used only in some datasets and only for certain columns (with enough diversity - I selected them manually).",
    "2841041": "Thanks! Have you experimented with diff features, namely last-mean, last-first?",
    "2841062": "No, unfortunately.",
    "2841208": "Thank u for share .  You mentioned \"Feature selection based on importance by gain\"  didn't work.  Did importance by split work?",
    "2841214": "Yes, I selected features for MLP based on importance by split of LGBM, but importance by gain in my cases led to overfitting.",
    "2841222": "Hi Andrey, congrats with the 62nd place! A lot of work is done!\n\nThere are a few things that helped me to boost my score quite a bit are the following:\n - Aggregation of numerical columns: last 3 rows tail(3).mean(); as well as diff() with mean and/or std\n - Many features created by hand using logical assumptions\n - The addition of stability metric had a minor effect\n - Max aggregations of string columns had a big effect. However, I rejected this due to a problem of interpretability with the hope that this is an artifact of the public LB test sample.\n\nJust curious, have you tried/considered such tail/diff aggregations to add to your data?\n\nI am still very puzzled to why the class balance does not work in this competition. When training my models, I even balanced the classes versus week number trying to exclude the possibility for models to learn any discrimination from this dimension ( as the WEEK_NUM is correlated with some inputs) with the hope this can enhance the model stability.",
    "2841270": "Hi Oleh. Thanks.\nI was thinking about some kind of tail(n).some_agg() (I also read mention about it in previous competition solutions and discussions), but I think this idea got lost somewhere in my notes. I didn't think about diff.\nAbout max agg for string columns - it may work because it contains some hierarchy in the some categories encoding (e.g. P0_111_121).",
    "2841285": "You gain a lot from ensembling. \nI have 0.598 single model and only 0.6 ensemble",
    "2841298": "Yes, that was a kind of exercise from my previous competitions. Of course, that's only one of the approaches, and a strong enough single model can beat a good ensemble.",
    "2841342": "Hi @andreynesterov, thank you for sharing. I read that DART did not work well, but in my case it has improved the result and reach almost the catboost performance.",
    "2841364": "That's great. I also thought dart was very promising, maybe I didn't pay much attention to tuning it or preparing features.",
    "2841461": "Thank you for the detailed insights on your findings. I have been following your work from the beginning of this competition and a lot of my work is inspired from your notebooks. Thank you for generously sharing those. Beginners like me learn a lot from your notebooks.\n\nI especially referred to the notebook where you ensembled xgb, cat, lgb model. I took the lessons from that notebook and tried ensembling 2 lgb and 2 cat - one of each with class balancing and other without the balancing and that did well. But, my best one which i did not submit was the one which i added a post-processing layer of manual implementation of isotonic calibration. it was a 0.593 on public LB and 0.527 on private LB.",
    "2841501": "That's interesting."
  },
  "source": "meta"
}