{
  "id": 347637,
  "title": "2nd place solution - team JuneHomes (writeup)",
  "url": "/competitions/amex-default-prediction/discussion/347637",
  "author_name": "Konstantin Yakovlev",
  "post_date": "2022-08-25T00:19:56.767000",
  "votes": 180,
  "comment_count": 57,
  "views": 0,
  "content": "<p><em>Note: Full code to retrain single model will be shared here in a 2 weeks.</em></p>\n<p>I would like to say thank you to competition hosts and Kaggle - it was a great pleasure to participate in tabular data competition after many months and years without one. </p>\n<p>Thank you to all participants (and of course winners - <a href=\"https://www.kaggle.com/daishu\" target=\"_blank\">@daishu</a> huge jump during last 3 days and fantastic solo result) - our success is your success - you forced us to try harder - without all of you It would be impossible to learn so many new things and achieve such result.</p>\n<p>And special \"thank you\" goes to my fantastic teammates:</p>\n<ul>\n<li>Danila</li>\n<li>Alexey</li>\n<li>Igor</li>\n</ul>\n<p>No words would be enough to say how much each of you contributed to the end result.<br>\n\"Do Data Scientists have hobby? Yes -&gt; DS competition\".  </p>\n<hr>\n<h3>Competition announcement</h3>\n<p>We were very hyped by the new tabular data competition release (sorry for the external link:  <a href=\"https://www.linkedin.com/posts/konstantin-yakovlev-010b46125_american-express-default-prediction-kaggle-activity-6935250015619592194-hqnM\" target=\"_blank\">link</a>) and immediately decided to participate. Slack notification -&gt; rules and perspective advertisement (100% chance too lose summer holidays and all free time) -&gt; and here we are - four members. Only one member of the team had previous experience in DS competitions participation.</p>\n<p>Few rules were established from the beginning:</p>\n<ul>\n<li>Only free time for competition</li>\n<li>No \"second\" accounts on Kaggle (even wife/friends to exclude any cheating suspicion)</li>\n<li>No competition discussion outside of the team</li>\n<li>We are here to learn and try our best</li>\n</ul>\n<hr>\n<h3>Infrastructure and pipelines:</h3>\n<p>Each of us had own machines / resources (GCP/AWS/local). We used Kaggle platform just for a few times. So the first thing we wanted to solve - unified machine to save all artefacts / experiments. We decided to go with AWS. I would say that it is possible to achieve the same result that we have just with Kaggle resources, but it would be bit more stressful for team management. We didn't want to spend a lot of money on AWS but sometimes (during very hot hours) RAM spikes were 500GB+ to permit simultaneous work.</p>\n<p>We tried to use neptune.ai for ML tracking but from July it was not very effective as we entered in brute force zone. </p>\n<p><em>Advise: Resources management is very critical - find bottleneck and remove it to make your team most effective. At the same time don't burn money recklessly - limit your budged. If any optimization possible - do it as soon as possible to save time and resources.</em></p>\n<hr>\n<h3>Project Structure</h3>\n<p>Each run was internally versioned (ex. v1.1.1 - (major version).(fe version).(model version))<br>\nOverall project structure:</p>\n<ul>\n<li>Initial preprocess -&gt; artifact cleaned and joined df</li>\n<li>FE -&gt; Many aligned (by uid) dfs with separated features </li>\n<li>Features selection -&gt; dictionary we selection metadata</li>\n<li>Holdout Model (fe check and tuning) -&gt; Local validation oof preds / holdout preds/ model / model metadata</li>\n<li>Full model run -&gt; Model / Model metadata</li>\n<li>Prediction -&gt; each fold oof predictions / cv split metadata / test predictions</li>\n</ul>\n<p>All these permitted us to go back and forward and check what worked well and what did not and restore experiments in each particular step.</p>\n<hr>\n<h3>Initial preprocess</h3>\n<p>We wanted to achieve several things with this step:</p>\n<ul>\n<li>Join Train and Test -&gt; due to many people involved I was afraid that some missed transformation on private test part will be unnoticed. So we sacrifice memory and speed optimization for overall stability and security.</li>\n<li>Remove detected noise -&gt; (we had options here but ended with unified single one)</li>\n<li>Transform Customer ID to unified uid </li>\n<li>Create internal subset feature -&gt; Train / Public / Private</li>\n<li>Create unified kfold and holdout split -&gt; To align all experiments</li>\n<li>Separate columns by type and store them separately to minify memory use and load time</li>\n</ul>\n<h4>Remove detected noise</h4>\n<p>We didn't use public notebooks for cleaning. Radar's Dataset is fantastic and it is 99% similar to our own transformations.<br>\nWe used \"isle\" identification without any pre-build coefficients.<br>\ndummy code is something like this:</p>\n<pre><code>    for col in process_columns: \n\n        df = temp_df[[col]].sort_values(by=[col]) \n        df = df[df[col].notna()].drop_duplicates(subset=[col]).reset_index(drop=True)\n\n        df['temp'] = np.floor(df[col] * 100000)\n        df['group'] = ((df['temp'] - df['temp'].shift()).abs() &gt;= 100).cumsum()\n\n        i = 0\n        while True:\n            min_val = df[df['group']==i]['temp'].min()\n            if min_val&gt;0:\n                break\n            i += 1\n\n        df['temp2'] = np.where(df['temp']&gt;=0, \n                                np.floor(df['temp']/min_val).astype(np.int32),\n                                np.round(df['temp']/min_val).astype(np.int32))\n\n        mapping = dict(zip(df[col],df['temp2']))\n        temp_df[col] = temp_df[col].map(mapping)\n\n        print(col, df['group'].nunique(), df[col].nunique())\n        print(df.groupby(['group'])['temp','temp2'].agg(['min','max','count','nunique']).head(40))\n</code></pre>\n<h4>Create internal subset feature</h4>\n<p>We used last statement month to create 0/1/2 feature and store in in \"index\" df</p>\n<h4>Create unified kfold and holdout split</h4>\n<p>Fixed random seed (of course 42) to make spits and then took 20% of customers to holdout group (to test stacking / blending / etc)</p>\n<h4>Separate columns by type</h4>\n<p>After cleaning we had several columns \"groups\".</p>\n<pre><code>all_files = [\n'p_columns', -&gt; just p columns as we thought that they are very different (and P_2 is internal amex \"scoring\" model)\n'objects_radar_columns', -&gt; order encoding (we were checking where out cleaning differs from public approaches and here was the unique place) \n'objects_columns', -&gt; onehot encoding\n'categorical_cleaned__D__columns', -&gt; no noise categoricals\n'categorical_binary__S__columns', -&gt; cleaned binary\n'categorical_binary__R__columns', -&gt; cleaned binary\n'categorical_binary__D__columns', -&gt; cleaned binary\n'categorical_binary__B__columns', -&gt; cleaned binary\n'categorical__D__columns', -&gt; removed noise categoricals\n'categorical__B__columns', -&gt; removed noise categoricals\n'cleaned__B__columns', -&gt; removed noise continuous \n'cleaned__D__columns', -&gt; removed noise continuous \n'cleaned__R__columns', -&gt; removed noise continuous \n'cleaned__S__columns', -&gt; removed noise continuous \n'rest__B__columns', -&gt; have no idea what to do with it -&gt; floor \n'rest__D__columns', -&gt; have no idea what to do with it -&gt; floor \n'rest__R__columns', -&gt; have no idea what to do with it -&gt; floor \n'rest__S__columns', -&gt; have no idea what to do with it -&gt; floor \n]\n</code></pre>\n<p>Thanks again to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> we always used your preprocess as a baseline.</p>\n<p>We were able to load just portion of data -&gt; do fe -&gt; concat to \"index\" as all dfs were aligned by index. Also such split permitted us to do fe by feature type to accelerate process and see more statistically valuable metric change. </p>\n<hr>\n<h3>FE</h3>\n<p>We started with careful fe column by column or small subset and it worked well until 1xx features and then any metric improvement or degradation was not statistically significant and many features \"overlapped\" on importance and significance. </p>\n<p><em>Note: I believe that it is possible to build silver zone robust model with only 3xx features</em></p>\n<p>So we started from scratch with brute force))) Of course there was no need to apply \"nunique\" (for example to binary features) and our previous step helped us to limit fe.</p>\n<pre><code>all_aggregations = {\n   'agg_func': ['last','mean','std','median','min','max','nunique'],\n   'diff_func': ['first','mean','std','median','min','max'],\n   'ratio_func': ['first','mean','std','median','min','max'],\n   'lags': [1,2,3,6,11],\n   'special': ['ewm','count_month_enc','monotonic_increase','diff_mean','major_class',\n              'normalization','top_outlier','bottom_outlier','normalization_mean','top_outlier_mean','top_outlier_mean']\n}    \n</code></pre>\n<ul>\n<li>pca (horizonal and vertical) + horizontal combinations + horizontal aggregations.</li>\n</ul>\n<p>agg_func -&gt; normal aggregations by uid<br>\ndiff_func -&gt; diff last - xxx -&gt; std diff worked better than any other<br>\nratio_func -&gt; ratio_func last/xxx<br>\nlags -&gt; diff last - Nx<br>\nspecial -&gt; some special transformations -&gt; count_month_enc worked well for categorical / emw for continous </p>\n<p>We ended up with about 7k features (stored file by group and by agg type for faster loading). <br>\nNext thing was to figure out what works and what not -&gt; this topic was the most challenging for us. </p>\n<h4>Normalizations</h4>\n<p>It's better to call it Standardization (x - m) / s -&gt; as we had also normalization test the name became constant \"normalization\")))</p>\n<pre><code>df.groupby(['dt_month','subset'])[col].agg(['mean','std'])\n</code></pre>\n<p>dt_month -&gt; month of the statement<br>\nsubset -&gt; train / public / private<br>\nand mean and std from clients that had full statement history.</p>\n<p>We have to have temporal shift to make it work. So we did a \"trick\" removed last statement for each client and applied exactly same transformation  for each client and merged appropriate labels. So we had 2 lines in training set for almost each client BUT validated results only on last statement during CV runs and Holdout checks. It more or less same as adding noised data but we had temporal drift and model was able to work better on unknown future data with  \"possible\" data drift.</p>\n<hr>\n<h3>Features selection</h3>\n<p>Ooohh that was really fun. </p>\n<p>We used gbdt boosting type during experiments as it was very aligned with dart mode but was significantly faster.<br>\nAlso, we used ROC AUC score during our experiments as we believed that due to amex instability we can't use it for decision making  (of course we tracked log loss and amex).</p>\n<p>In previous step we brute forced many features and now is time to clean them out.<br>\nAll feature selection was done with 5 CV folds training + independent check on 20% holdout data.</p>\n<ol>\n<li><p>Zero importance -&gt; Right from the start we were able to through away 1.5k features that had exactly 0 importance (lgbm importance). That means that with 250 bins and 2**10 data in leaf those features are not participating in any split.</p></li>\n<li><p>Stepped hierarchical permutation importance -&gt; we defined 300 initial features and looped over all other features subsets (600+) - was very time consuming but very stable.<br>\nNote: we shuffled order of the subset to force model try different combinations.<br>\nAdd features subset -&gt; train model -&gt; permutate -&gt; drop negative features (negative mean over 5 seeds) -&gt; add new subset -&gt; …<br>\nDuring this part that took almost 3 days we limited features to 3k -&gt; 0.800 lb</p></li>\n<li><p>Stepped permutation importance.<br>\nTake all features -&gt; train model -&gt; permutate -&gt; drop 20% of worst performed features (only negative) -&gt; repeat. Final subset was 25xx features (and different from previous step) -&gt; 0.800 lb</p></li>\n<li><p>Forward feature selection.<br>\nWe defined 300 initial features and simply added subset by subset and compared ROC AUC if metric change was &gt; 0.0003 we kept the subset. -&gt; 0.800 lb</p></li>\n<li><p>Time series CV.<br>\nFor very doubtful features as PCA and Normilized values we used to different validation stratagies:</p></li>\n</ol>\n<ul>\n<li>Train on first 6 month values (last statement of the first 6 months went to train set) and validate on last 6 (also just last statement of the last 6 months). We trained model without temporal feature and then with if result was better on CV and on holdout we added to final features subset.</li>\n<li>We used P_2, B_1, B_2 as a proxy target and MSE loss with combined Train and Test to see if we did right transformation and result did not degrade.</li>\n</ul>\n<p>Many other options we tried but result was not stable.</p>\n<p>Final subset came from \"Forward feature selection\" plus overlapped features from other technics minus overlapped negative combination.  -&gt; lb 0.801 single model.</p>\n<p>We tried to blend many models with different subset as we believed that it should give huge LB boost (based on holdout blending tests) but it didn't work well for lb. </p>\n<hr>\n<h3>Model</h3>\n<p>In my own experience, DART never worked better and here we have proof that in DS \"all depends.\" We did experiments with DART in the beginning and it did not show any metric improvement with our params and baseline model features subset. Later we found <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> notebook and gave it one more try and it worked marvelously.</p>\n<p>From the beginning, we tried to build a more complex model with 2**7+ leaves and 0.7+ features but failed. It still puzzles me why a simple model with a very low number of features works here. </p>\n<p>I saw such behaviour mostly on synthetic data and stacking - so we tried to find out if data is syntetic (at least partly) and deanonimize internal scoring values - but didn't make it.</p>\n<p>Our best single lgbm model was trained on 29xx features. 5 folds CV - no stratification by any option. Training data - 2 last staements for each client (transformed independently). Params:</p>\n<pre><code>lgb_params = {\n    'boosting_type': 'dart',\n    'objective': 'cross_entropy', \n    'metric': ['AUC'],\n    'subsample': 0.8,  \n    'subsample_freq': 1,\n    'learning_rate': 0.01, \n    'num_leaves': 2 ** 6, \n    'min_data_in_leaf': 2 ** 11, \n    'feature_fraction': 0.2, \n    'feature_fraction_bynode':0.3,\n    'first_metric_only': True,\n    'n_estimators': 17001,  # -&gt; 5000 for gbdt \n    'boost_from_average': False,\n    'early_stopping_rounds': 300,\n    'verbose': -1,\n    'num_threads': -1,\n    'seed': SEED,\n}\n</code></pre>\n<p>Blend -&gt; Power (2) rank blend of Dart lgbm (0.801 public) / GBDT lgbm  (0.799 public) / Catboost models (0.799 public)</p>\n<p>Single lgbm with 3 last statements showed even better CV by we didn't have enough time to retrain it (full DART run for 5 folds took 12+ hours there).</p>\n<p>It was obvious that clients with a little number of statements will not get benefit from all 2k features. So we created a special model that was trained only on 300 features with custom params (also dart). Predictions for clients with &lt;=2 statements came exclusively from such model and were not blended with other models. </p>\n<p>How did we combine the result from 2 independent models to not destroy the final ranking?</p>\n<table>\n<thead>\n<tr>\n<th>Client id</th>\n<th>Number of statements</th>\n<th>Basic ranking</th>\n<th>&lt;=2 prediction</th>\n<th>Final ranking</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>5</td>\n<td>…</td>\n<td>5</td>\n</tr>\n<tr>\n<td>1</td>\n<td>2</td>\n<td>4</td>\n<td>0.1</td>\n<td>3</td>\n</tr>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>8</td>\n<td>…</td>\n<td>8</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>3</td>\n<td>0.5</td>\n<td>4</td>\n</tr>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>7</td>\n<td>…</td>\n<td>7</td>\n</tr>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>2</td>\n<td>…</td>\n<td>2</td>\n</tr>\n</tbody>\n</table>\n<p>we kept ranking for &gt;2 statements and for rest resorted within initial ranking group (hope it's clear enough))). Was is perfect - no, but it was very stable with really tiny improvement (because of number of such clients in public and private test parts).</p>\n<p>What also worked well:</p>\n<ul>\n<li>Train on all data without folds splitting and stop just 1000 rounds further then CV showed. </li>\n</ul>\n<p>What didn't work:</p>\n<ul>\n<li>Stacking by any mean</li>\n<li>Many models with different seed and fe order</li>\n<li>Massive blend of different models with different types and features (blend worked well till 0.799 and then any low performed model 0.795- made public score worse - we use public CV as additional holdout set and used approaches that worked well on local CV and LB - anything that worked partially was not used in the end).</li>\n</ul>\n<p>Again, nothing really fancy here. The main thing that helped us align Train / CV with LB was very high  'min_data_in_leaf'.<br>\nNo optuna used -&gt; just manual old school tuning based on data feeling.<br>\nWe did many experiments with weights and loss functions but none of them worked.</p>\n<p>Due to AMEX metric specification it was obvious that focal loss should work but it didn't. We tried several times to switch loss function during the competition period and the result was the same. </p>\n<p>Error analysis showed that model makes errors without any \"pattern\" -&gt; stacking didn't work for holdout set (25% of data)  and we had doubts that it will work on private/public test parts. We kept only LR/Lasso(0.02) for blending options to choose submissions.</p>\n<p>Cross validation -&gt; standard 5folds CV split by client ID. The unique thing that we did here is \"prespliting\" to align all CV between team members to be able to compare results directly.</p>\n<hr>\n<p>What left without mentions:</p>\n<ul>\n<li>EDA on data</li>\n<li>Denoising experiments</li>\n<li>Data deanonymization -&gt; didn't manage to make it</li>\n<li>Features pairs and triples combinations -&gt; that didn't work well</li>\n<li>NaN filling -&gt; didn't work</li>\n<li>Clusterization -&gt; didn't work</li>\n<li>Hundreds of experiments with features selection process and internal discussions about it.</li>\n<li>Adding noised data (noise / swap noise) that leaded to interesting but doubtful results</li>\n<li>Model tuning </li>\n<li>Removing absolute values and keep only diff or ratios -&gt; should be more stable for future data but we saw some lb degradation and didn't proceed</li>\n<li>pseudo labeling</li>\n</ul>\n<p>What we always wanted but didn't found time to do:</p>\n<ul>\n<li>NN - we have no NN in our final blend</li>\n<li>P_2 or any other column prediction (1/2/3/4 months ahead) with combined data and use it as meta information for lgbm main model</li>\n<li>11 / 12 / 13 statements joined training on different subsets (df was too large and training was slow)</li>\n</ul>\n<hr>\n<h3>Internal initial plan</h3>\n<pre><code>########################### Data preprocessing and Data evaluation\n#################################################################################\n\n## Added noise removal -&gt; GOOD2DO\n# There is no doubt that some Noise was injected in data\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\n# We need to find a way to remove it \n# the best option to not follow public approach\n# At least with columns where columns have overlaped population\n\n## Data minification for FE -&gt; GOOD2DO\n# Datatype downcasting\n# Pickle/Parquet/Feather \n# Be careful with floats16 as it may lead to bad agg results\n# Also float16 may lead to some signal degradation due to precision and values changes\n\n## Evaluate values distributions and NaNs -&gt; GOOD2DO\n# Full 13 months history\n# Train against Test Public and Test Private\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\n#\n# Need to try:\n# Kolmogorov–Smirnov test -&gt; GOOD2DO\n#\n# Adversarial validation -&gt; GOOD2DO\n# https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation/notebook (just as simple example)\n#\n# Entropy / Distances / etc....\n#\n# Visual checks)))\n#\n# We need to find if ANY feature has very different distribution in PRIVATE test set\n# If that feature works for Public part it doesn't mean that it will work for Private\n\n########################### Targets\n#################################################################################\n\n# We need to find a way to get more targets -&gt; GOOD2DO\n# as we currently training on a single point by client we could greatly improve results\n# by extending our training set with new targets\n#\n# Find default periods in current client history and make appropriate labeling -&gt; GOOD2DO\n# Make 2 level model -&gt; predict p_2 values as normal time-series model and feed it to 2nd level GBT\n\n########################### Separate Models\n#################################################################################\n\n# Probably it's a good idea to make separate models for each subdatasets\n# Full history (13 months)\n# Less than 13 months\n\n########################### External Data\n#################################################################################\n# We can try to add \"Consumer index\" or any other independent temporal feature\n# Will not work if we will not be able to expand targets and add temporal feature\n\n########################### FE\n#################################################################################\n\n# We didn't make anything special here -&gt; July\n# AGGS (Stats by client)\n# Rollings\n# History length feature (not sure if it will help with Private Test)\n# Should we correct statements dates and add NaNs?\n# ReRanking categorical features by P_2 or Target\n# Clusterization (4+ groups feature by feature)\n# Count and Mean encodings for categorical features\n# Features combinations (sum/prod/power) -&gt; bruteforce\n# PCA or any other dimension reduction by features groups\n\n# We need to find if there is \"connection\" between clients in Train -&gt; Public Test -&gt; Private Test\n# we have 458913 + 924621 -&gt; 1383534 If I were AMEX I would export 1M clients (or other round number)\n# so may be 384 534 Clients are overlaps\n\n# Clip by 5 - 95 percentile\n\n########################### Features Selection\n#################################################################################\n# Permutation importance (use all fold only!!!) -&gt; recursive elimination (because of quantity of features -&gt; 3-4 rounds with 0 and 50% negative drop) \n# SHAP\n# Highly correlated features (.98+?)\n# Forward selection (may take ages and due aggs may be not effective - probably by feature block) \n# Backward elimination (may take ages and due aggs may be not effective - probably by feature block) \n\n########################### CV\n#################################################################################\n# Mean Target differs my \"history length\" -&gt; could be wise to do GroupedStratifeidFolds by history length\n# For sure Splits should be done by client\n# Target stratification to balance folds\n\n########################### Loss function / Metric\n#################################################################################\n# Clean and fast np/torch metric\n# Now it's in helper (need to cleanup that)\n# https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\n#\n# I don't believe that we will have better results with different loss function\n# But it worth to try at least focal loss\n# https://maxhalford.github.io/blog/lightgbm-focal-loss/\n#\n# Weights -&gt; we should try change weights there\n# weights by class\n# weights by some history length\n# weights by internal fe group\n#\n# We need custom metric for catboost\n# example https://catboost.ai/en/docs/concepts/python-usages-examples#logloss1\n\n########################### Models\n#################################################################################\n\n## First level choice\n# LGB/XGB/CTB -&gt; our main models here for sure\n# After stabilizing the baseline model and base feature we need to make 1st round tuning\n\n## Catboos specials\n# Categorical features\n# Embeding features\n\n## NN (GPU/TPU) -&gt; RNN / LSTM / Transformer\n# TPU -&gt; tensorflow (as it works better there)\n\n## NN -&gt; AE / VAE / DAE -&gt; as a denoising model hidden layer as input for GBT models\n# No need complex approach - just fast check the idea and in case of success move to big model\n\n########################### Blending\n#################################################################################\n# Weighted Average\n# Power Average\n# Weighted Rank Average\n# Linear/SVM\n# Postprocessing?\n</code></pre>",
  "messages": [
    {
      "id": 1912721,
      "postDate": "2022-08-25T00:19:56.767Z",
      "content": "<p><em>Note: Full code to retrain single model will be shared here in a 2 weeks.</em></p>\n<p>I would like to say thank you to competition hosts and Kaggle - it was a great pleasure to participate in tabular data competition after many months and years without one. </p>\n<p>Thank you to all participants (and of course winners - <a href=\"https://www.kaggle.com/daishu\" target=\"_blank\">@daishu</a> huge jump during last 3 days and fantastic solo result) - our success is your success - you forced us to try harder - without all of you It would be impossible to learn so many new things and achieve such result.</p>\n<p>And special \"thank you\" goes to my fantastic teammates:</p>\n<ul>\n<li>Danila</li>\n<li>Alexey</li>\n<li>Igor</li>\n</ul>\n<p>No words would be enough to say how much each of you contributed to the end result.<br>\n\"Do Data Scientists have hobby? Yes -&gt; DS competition\".  </p>\n<hr>\n<h3>Competition announcement</h3>\n<p>We were very hyped by the new tabular data competition release (sorry for the external link:  <a href=\"https://www.linkedin.com/posts/konstantin-yakovlev-010b46125_american-express-default-prediction-kaggle-activity-6935250015619592194-hqnM\" target=\"_blank\">link</a>) and immediately decided to participate. Slack notification -&gt; rules and perspective advertisement (100% chance too lose summer holidays and all free time) -&gt; and here we are - four members. Only one member of the team had previous experience in DS competitions participation.</p>\n<p>Few rules were established from the beginning:</p>\n<ul>\n<li>Only free time for competition</li>\n<li>No \"second\" accounts on Kaggle (even wife/friends to exclude any cheating suspicion)</li>\n<li>No competition discussion outside of the team</li>\n<li>We are here to learn and try our best</li>\n</ul>\n<hr>\n<h3>Infrastructure and pipelines:</h3>\n<p>Each of us had own machines / resources (GCP/AWS/local). We used Kaggle platform just for a few times. So the first thing we wanted to solve - unified machine to save all artefacts / experiments. We decided to go with AWS. I would say that it is possible to achieve the same result that we have just with Kaggle resources, but it would be bit more stressful for team management. We didn't want to spend a lot of money on AWS but sometimes (during very hot hours) RAM spikes were 500GB+ to permit simultaneous work.</p>\n<p>We tried to use neptune.ai for ML tracking but from July it was not very effective as we entered in brute force zone. </p>\n<p><em>Advise: Resources management is very critical - find bottleneck and remove it to make your team most effective. At the same time don't burn money recklessly - limit your budged. If any optimization possible - do it as soon as possible to save time and resources.</em></p>\n<hr>\n<h3>Project Structure</h3>\n<p>Each run was internally versioned (ex. v1.1.1 - (major version).(fe version).(model version))<br>\nOverall project structure:</p>\n<ul>\n<li>Initial preprocess -&gt; artifact cleaned and joined df</li>\n<li>FE -&gt; Many aligned (by uid) dfs with separated features </li>\n<li>Features selection -&gt; dictionary we selection metadata</li>\n<li>Holdout Model (fe check and tuning) -&gt; Local validation oof preds / holdout preds/ model / model metadata</li>\n<li>Full model run -&gt; Model / Model metadata</li>\n<li>Prediction -&gt; each fold oof predictions / cv split metadata / test predictions</li>\n</ul>\n<p>All these permitted us to go back and forward and check what worked well and what did not and restore experiments in each particular step.</p>\n<hr>\n<h3>Initial preprocess</h3>\n<p>We wanted to achieve several things with this step:</p>\n<ul>\n<li>Join Train and Test -&gt; due to many people involved I was afraid that some missed transformation on private test part will be unnoticed. So we sacrifice memory and speed optimization for overall stability and security.</li>\n<li>Remove detected noise -&gt; (we had options here but ended with unified single one)</li>\n<li>Transform Customer ID to unified uid </li>\n<li>Create internal subset feature -&gt; Train / Public / Private</li>\n<li>Create unified kfold and holdout split -&gt; To align all experiments</li>\n<li>Separate columns by type and store them separately to minify memory use and load time</li>\n</ul>\n<h4>Remove detected noise</h4>\n<p>We didn't use public notebooks for cleaning. Radar's Dataset is fantastic and it is 99% similar to our own transformations.<br>\nWe used \"isle\" identification without any pre-build coefficients.<br>\ndummy code is something like this:</p>\n<pre><code>    for col in process_columns: \n\n        df = temp_df[[col]].sort_values(by=[col]) \n        df = df[df[col].notna()].drop_duplicates(subset=[col]).reset_index(drop=True)\n\n        df['temp'] = np.floor(df[col] * 100000)\n        df['group'] = ((df['temp'] - df['temp'].shift()).abs() &gt;= 100).cumsum()\n\n        i = 0\n        while True:\n            min_val = df[df['group']==i]['temp'].min()\n            if min_val&gt;0:\n                break\n            i += 1\n\n        df['temp2'] = np.where(df['temp']&gt;=0, \n                                np.floor(df['temp']/min_val).astype(np.int32),\n                                np.round(df['temp']/min_val).astype(np.int32))\n\n        mapping = dict(zip(df[col],df['temp2']))\n        temp_df[col] = temp_df[col].map(mapping)\n\n        print(col, df['group'].nunique(), df[col].nunique())\n        print(df.groupby(['group'])['temp','temp2'].agg(['min','max','count','nunique']).head(40))\n</code></pre>\n<h4>Create internal subset feature</h4>\n<p>We used last statement month to create 0/1/2 feature and store in in \"index\" df</p>\n<h4>Create unified kfold and holdout split</h4>\n<p>Fixed random seed (of course 42) to make spits and then took 20% of customers to holdout group (to test stacking / blending / etc)</p>\n<h4>Separate columns by type</h4>\n<p>After cleaning we had several columns \"groups\".</p>\n<pre><code>all_files = [\n'p_columns', -&gt; just p columns as we thought that they are very different (and P_2 is internal amex \"scoring\" model)\n'objects_radar_columns', -&gt; order encoding (we were checking where out cleaning differs from public approaches and here was the unique place) \n'objects_columns', -&gt; onehot encoding\n'categorical_cleaned__D__columns', -&gt; no noise categoricals\n'categorical_binary__S__columns', -&gt; cleaned binary\n'categorical_binary__R__columns', -&gt; cleaned binary\n'categorical_binary__D__columns', -&gt; cleaned binary\n'categorical_binary__B__columns', -&gt; cleaned binary\n'categorical__D__columns', -&gt; removed noise categoricals\n'categorical__B__columns', -&gt; removed noise categoricals\n'cleaned__B__columns', -&gt; removed noise continuous \n'cleaned__D__columns', -&gt; removed noise continuous \n'cleaned__R__columns', -&gt; removed noise continuous \n'cleaned__S__columns', -&gt; removed noise continuous \n'rest__B__columns', -&gt; have no idea what to do with it -&gt; floor \n'rest__D__columns', -&gt; have no idea what to do with it -&gt; floor \n'rest__R__columns', -&gt; have no idea what to do with it -&gt; floor \n'rest__S__columns', -&gt; have no idea what to do with it -&gt; floor \n]\n</code></pre>\n<p>Thanks again to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> we always used your preprocess as a baseline.</p>\n<p>We were able to load just portion of data -&gt; do fe -&gt; concat to \"index\" as all dfs were aligned by index. Also such split permitted us to do fe by feature type to accelerate process and see more statistically valuable metric change. </p>\n<hr>\n<h3>FE</h3>\n<p>We started with careful fe column by column or small subset and it worked well until 1xx features and then any metric improvement or degradation was not statistically significant and many features \"overlapped\" on importance and significance. </p>\n<p><em>Note: I believe that it is possible to build silver zone robust model with only 3xx features</em></p>\n<p>So we started from scratch with brute force))) Of course there was no need to apply \"nunique\" (for example to binary features) and our previous step helped us to limit fe.</p>\n<pre><code>all_aggregations = {\n   'agg_func': ['last','mean','std','median','min','max','nunique'],\n   'diff_func': ['first','mean','std','median','min','max'],\n   'ratio_func': ['first','mean','std','median','min','max'],\n   'lags': [1,2,3,6,11],\n   'special': ['ewm','count_month_enc','monotonic_increase','diff_mean','major_class',\n              'normalization','top_outlier','bottom_outlier','normalization_mean','top_outlier_mean','top_outlier_mean']\n}    \n</code></pre>\n<ul>\n<li>pca (horizonal and vertical) + horizontal combinations + horizontal aggregations.</li>\n</ul>\n<p>agg_func -&gt; normal aggregations by uid<br>\ndiff_func -&gt; diff last - xxx -&gt; std diff worked better than any other<br>\nratio_func -&gt; ratio_func last/xxx<br>\nlags -&gt; diff last - Nx<br>\nspecial -&gt; some special transformations -&gt; count_month_enc worked well for categorical / emw for continous </p>\n<p>We ended up with about 7k features (stored file by group and by agg type for faster loading). <br>\nNext thing was to figure out what works and what not -&gt; this topic was the most challenging for us. </p>\n<h4>Normalizations</h4>\n<p>It's better to call it Standardization (x - m) / s -&gt; as we had also normalization test the name became constant \"normalization\")))</p>\n<pre><code>df.groupby(['dt_month','subset'])[col].agg(['mean','std'])\n</code></pre>\n<p>dt_month -&gt; month of the statement<br>\nsubset -&gt; train / public / private<br>\nand mean and std from clients that had full statement history.</p>\n<p>We have to have temporal shift to make it work. So we did a \"trick\" removed last statement for each client and applied exactly same transformation  for each client and merged appropriate labels. So we had 2 lines in training set for almost each client BUT validated results only on last statement during CV runs and Holdout checks. It more or less same as adding noised data but we had temporal drift and model was able to work better on unknown future data with  \"possible\" data drift.</p>\n<hr>\n<h3>Features selection</h3>\n<p>Ooohh that was really fun. </p>\n<p>We used gbdt boosting type during experiments as it was very aligned with dart mode but was significantly faster.<br>\nAlso, we used ROC AUC score during our experiments as we believed that due to amex instability we can't use it for decision making  (of course we tracked log loss and amex).</p>\n<p>In previous step we brute forced many features and now is time to clean them out.<br>\nAll feature selection was done with 5 CV folds training + independent check on 20% holdout data.</p>\n<ol>\n<li><p>Zero importance -&gt; Right from the start we were able to through away 1.5k features that had exactly 0 importance (lgbm importance). That means that with 250 bins and 2**10 data in leaf those features are not participating in any split.</p></li>\n<li><p>Stepped hierarchical permutation importance -&gt; we defined 300 initial features and looped over all other features subsets (600+) - was very time consuming but very stable.<br>\nNote: we shuffled order of the subset to force model try different combinations.<br>\nAdd features subset -&gt; train model -&gt; permutate -&gt; drop negative features (negative mean over 5 seeds) -&gt; add new subset -&gt; …<br>\nDuring this part that took almost 3 days we limited features to 3k -&gt; 0.800 lb</p></li>\n<li><p>Stepped permutation importance.<br>\nTake all features -&gt; train model -&gt; permutate -&gt; drop 20% of worst performed features (only negative) -&gt; repeat. Final subset was 25xx features (and different from previous step) -&gt; 0.800 lb</p></li>\n<li><p>Forward feature selection.<br>\nWe defined 300 initial features and simply added subset by subset and compared ROC AUC if metric change was &gt; 0.0003 we kept the subset. -&gt; 0.800 lb</p></li>\n<li><p>Time series CV.<br>\nFor very doubtful features as PCA and Normilized values we used to different validation stratagies:</p></li>\n</ol>\n<ul>\n<li>Train on first 6 month values (last statement of the first 6 months went to train set) and validate on last 6 (also just last statement of the last 6 months). We trained model without temporal feature and then with if result was better on CV and on holdout we added to final features subset.</li>\n<li>We used P_2, B_1, B_2 as a proxy target and MSE loss with combined Train and Test to see if we did right transformation and result did not degrade.</li>\n</ul>\n<p>Many other options we tried but result was not stable.</p>\n<p>Final subset came from \"Forward feature selection\" plus overlapped features from other technics minus overlapped negative combination.  -&gt; lb 0.801 single model.</p>\n<p>We tried to blend many models with different subset as we believed that it should give huge LB boost (based on holdout blending tests) but it didn't work well for lb. </p>\n<hr>\n<h3>Model</h3>\n<p>In my own experience, DART never worked better and here we have proof that in DS \"all depends.\" We did experiments with DART in the beginning and it did not show any metric improvement with our params and baseline model features subset. Later we found <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> notebook and gave it one more try and it worked marvelously.</p>\n<p>From the beginning, we tried to build a more complex model with 2**7+ leaves and 0.7+ features but failed. It still puzzles me why a simple model with a very low number of features works here. </p>\n<p>I saw such behaviour mostly on synthetic data and stacking - so we tried to find out if data is syntetic (at least partly) and deanonimize internal scoring values - but didn't make it.</p>\n<p>Our best single lgbm model was trained on 29xx features. 5 folds CV - no stratification by any option. Training data - 2 last staements for each client (transformed independently). Params:</p>\n<pre><code>lgb_params = {\n    'boosting_type': 'dart',\n    'objective': 'cross_entropy', \n    'metric': ['AUC'],\n    'subsample': 0.8,  \n    'subsample_freq': 1,\n    'learning_rate': 0.01, \n    'num_leaves': 2 ** 6, \n    'min_data_in_leaf': 2 ** 11, \n    'feature_fraction': 0.2, \n    'feature_fraction_bynode':0.3,\n    'first_metric_only': True,\n    'n_estimators': 17001,  # -&gt; 5000 for gbdt \n    'boost_from_average': False,\n    'early_stopping_rounds': 300,\n    'verbose': -1,\n    'num_threads': -1,\n    'seed': SEED,\n}\n</code></pre>\n<p>Blend -&gt; Power (2) rank blend of Dart lgbm (0.801 public) / GBDT lgbm  (0.799 public) / Catboost models (0.799 public)</p>\n<p>Single lgbm with 3 last statements showed even better CV by we didn't have enough time to retrain it (full DART run for 5 folds took 12+ hours there).</p>\n<p>It was obvious that clients with a little number of statements will not get benefit from all 2k features. So we created a special model that was trained only on 300 features with custom params (also dart). Predictions for clients with &lt;=2 statements came exclusively from such model and were not blended with other models. </p>\n<p>How did we combine the result from 2 independent models to not destroy the final ranking?</p>\n<table>\n<thead>\n<tr>\n<th>Client id</th>\n<th>Number of statements</th>\n<th>Basic ranking</th>\n<th>&lt;=2 prediction</th>\n<th>Final ranking</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>5</td>\n<td>…</td>\n<td>5</td>\n</tr>\n<tr>\n<td>1</td>\n<td>2</td>\n<td>4</td>\n<td>0.1</td>\n<td>3</td>\n</tr>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>8</td>\n<td>…</td>\n<td>8</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>3</td>\n<td>0.5</td>\n<td>4</td>\n</tr>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>7</td>\n<td>…</td>\n<td>7</td>\n</tr>\n<tr>\n<td>1</td>\n<td>13</td>\n<td>2</td>\n<td>…</td>\n<td>2</td>\n</tr>\n</tbody>\n</table>\n<p>we kept ranking for &gt;2 statements and for rest resorted within initial ranking group (hope it's clear enough))). Was is perfect - no, but it was very stable with really tiny improvement (because of number of such clients in public and private test parts).</p>\n<p>What also worked well:</p>\n<ul>\n<li>Train on all data without folds splitting and stop just 1000 rounds further then CV showed. </li>\n</ul>\n<p>What didn't work:</p>\n<ul>\n<li>Stacking by any mean</li>\n<li>Many models with different seed and fe order</li>\n<li>Massive blend of different models with different types and features (blend worked well till 0.799 and then any low performed model 0.795- made public score worse - we use public CV as additional holdout set and used approaches that worked well on local CV and LB - anything that worked partially was not used in the end).</li>\n</ul>\n<p>Again, nothing really fancy here. The main thing that helped us align Train / CV with LB was very high  'min_data_in_leaf'.<br>\nNo optuna used -&gt; just manual old school tuning based on data feeling.<br>\nWe did many experiments with weights and loss functions but none of them worked.</p>\n<p>Due to AMEX metric specification it was obvious that focal loss should work but it didn't. We tried several times to switch loss function during the competition period and the result was the same. </p>\n<p>Error analysis showed that model makes errors without any \"pattern\" -&gt; stacking didn't work for holdout set (25% of data)  and we had doubts that it will work on private/public test parts. We kept only LR/Lasso(0.02) for blending options to choose submissions.</p>\n<p>Cross validation -&gt; standard 5folds CV split by client ID. The unique thing that we did here is \"prespliting\" to align all CV between team members to be able to compare results directly.</p>\n<hr>\n<p>What left without mentions:</p>\n<ul>\n<li>EDA on data</li>\n<li>Denoising experiments</li>\n<li>Data deanonymization -&gt; didn't manage to make it</li>\n<li>Features pairs and triples combinations -&gt; that didn't work well</li>\n<li>NaN filling -&gt; didn't work</li>\n<li>Clusterization -&gt; didn't work</li>\n<li>Hundreds of experiments with features selection process and internal discussions about it.</li>\n<li>Adding noised data (noise / swap noise) that leaded to interesting but doubtful results</li>\n<li>Model tuning </li>\n<li>Removing absolute values and keep only diff or ratios -&gt; should be more stable for future data but we saw some lb degradation and didn't proceed</li>\n<li>pseudo labeling</li>\n</ul>\n<p>What we always wanted but didn't found time to do:</p>\n<ul>\n<li>NN - we have no NN in our final blend</li>\n<li>P_2 or any other column prediction (1/2/3/4 months ahead) with combined data and use it as meta information for lgbm main model</li>\n<li>11 / 12 / 13 statements joined training on different subsets (df was too large and training was slow)</li>\n</ul>\n<hr>\n<h3>Internal initial plan</h3>\n<pre><code>########################### Data preprocessing and Data evaluation\n#################################################################################\n\n## Added noise removal -&gt; GOOD2DO\n# There is no doubt that some Noise was injected in data\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\n# We need to find a way to remove it \n# the best option to not follow public approach\n# At least with columns where columns have overlaped population\n\n## Data minification for FE -&gt; GOOD2DO\n# Datatype downcasting\n# Pickle/Parquet/Feather \n# Be careful with floats16 as it may lead to bad agg results\n# Also float16 may lead to some signal degradation due to precision and values changes\n\n## Evaluate values distributions and NaNs -&gt; GOOD2DO\n# Full 13 months history\n# Train against Test Public and Test Private\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\n#\n# Need to try:\n# Kolmogorov–Smirnov test -&gt; GOOD2DO\n#\n# Adversarial validation -&gt; GOOD2DO\n# https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation/notebook (just as simple example)\n#\n# Entropy / Distances / etc....\n#\n# Visual checks)))\n#\n# We need to find if ANY feature has very different distribution in PRIVATE test set\n# If that feature works for Public part it doesn't mean that it will work for Private\n\n########################### Targets\n#################################################################################\n\n# We need to find a way to get more targets -&gt; GOOD2DO\n# as we currently training on a single point by client we could greatly improve results\n# by extending our training set with new targets\n#\n# Find default periods in current client history and make appropriate labeling -&gt; GOOD2DO\n# Make 2 level model -&gt; predict p_2 values as normal time-series model and feed it to 2nd level GBT\n\n########################### Separate Models\n#################################################################################\n\n# Probably it's a good idea to make separate models for each subdatasets\n# Full history (13 months)\n# Less than 13 months\n\n########################### External Data\n#################################################################################\n# We can try to add \"Consumer index\" or any other independent temporal feature\n# Will not work if we will not be able to expand targets and add temporal feature\n\n########################### FE\n#################################################################################\n\n# We didn't make anything special here -&gt; July\n# AGGS (Stats by client)\n# Rollings\n# History length feature (not sure if it will help with Private Test)\n# Should we correct statements dates and add NaNs?\n# ReRanking categorical features by P_2 or Target\n# Clusterization (4+ groups feature by feature)\n# Count and Mean encodings for categorical features\n# Features combinations (sum/prod/power) -&gt; bruteforce\n# PCA or any other dimension reduction by features groups\n\n# We need to find if there is \"connection\" between clients in Train -&gt; Public Test -&gt; Private Test\n# we have 458913 + 924621 -&gt; 1383534 If I were AMEX I would export 1M clients (or other round number)\n# so may be 384 534 Clients are overlaps\n\n# Clip by 5 - 95 percentile\n\n########################### Features Selection\n#################################################################################\n# Permutation importance (use all fold only!!!) -&gt; recursive elimination (because of quantity of features -&gt; 3-4 rounds with 0 and 50% negative drop) \n# SHAP\n# Highly correlated features (.98+?)\n# Forward selection (may take ages and due aggs may be not effective - probably by feature block) \n# Backward elimination (may take ages and due aggs may be not effective - probably by feature block) \n\n########################### CV\n#################################################################################\n# Mean Target differs my \"history length\" -&gt; could be wise to do GroupedStratifeidFolds by history length\n# For sure Splits should be done by client\n# Target stratification to balance folds\n\n########################### Loss function / Metric\n#################################################################################\n# Clean and fast np/torch metric\n# Now it's in helper (need to cleanup that)\n# https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\n#\n# I don't believe that we will have better results with different loss function\n# But it worth to try at least focal loss\n# https://maxhalford.github.io/blog/lightgbm-focal-loss/\n#\n# Weights -&gt; we should try change weights there\n# weights by class\n# weights by some history length\n# weights by internal fe group\n#\n# We need custom metric for catboost\n# example https://catboost.ai/en/docs/concepts/python-usages-examples#logloss1\n\n########################### Models\n#################################################################################\n\n## First level choice\n# LGB/XGB/CTB -&gt; our main models here for sure\n# After stabilizing the baseline model and base feature we need to make 1st round tuning\n\n## Catboos specials\n# Categorical features\n# Embeding features\n\n## NN (GPU/TPU) -&gt; RNN / LSTM / Transformer\n# TPU -&gt; tensorflow (as it works better there)\n\n## NN -&gt; AE / VAE / DAE -&gt; as a denoising model hidden layer as input for GBT models\n# No need complex approach - just fast check the idea and in case of success move to big model\n\n########################### Blending\n#################################################################################\n# Weighted Average\n# Power Average\n# Weighted Rank Average\n# Linear/SVM\n# Postprocessing?\n</code></pre>",
      "rawMarkdown": "*Note: Full code to retrain single model will be shared here in a 2 weeks.*\n\nI would like to say thank you to competition hosts and Kaggle - it was a great pleasure to participate in tabular data competition after many months and years without one. \n\nThank you to all participants (and of course winners - @daishu huge jump during last 3 days and fantastic solo result) - our success is your success - you forced us to try harder - without all of you It would be impossible to learn so many new things and achieve such result.\n\nAnd special \"thank you\" goes to my fantastic teammates:\n- Danila\n- Alexey\n- Igor\n\nNo words would be enough to say how much each of you contributed to the end result.\n\"Do Data Scientists have hobby? Yes -> DS competition\".  \n\n---\n\n### Competition announcement\n\nWe were very hyped by the new tabular data competition release (sorry for the external link:  [link](https://www.linkedin.com/posts/konstantin-yakovlev-010b46125_american-express-default-prediction-kaggle-activity-6935250015619592194-hqnM)) and immediately decided to participate. Slack notification -> rules and perspective advertisement (100% chance too lose summer holidays and all free time) -> and here we are - four members. Only one member of the team had previous experience in DS competitions participation.\n\nFew rules were established from the beginning:\n- Only free time for competition\n- No \"second\" accounts on Kaggle (even wife/friends to exclude any cheating suspicion)\n- No competition discussion outside of the team\n- We are here to learn and try our best\n\n---\n\n### Infrastructure and pipelines:\n\nEach of us had own machines / resources (GCP/AWS/local). We used Kaggle platform just for a few times. So the first thing we wanted to solve - unified machine to save all artefacts / experiments. We decided to go with AWS. I would say that it is possible to achieve the same result that we have just with Kaggle resources, but it would be bit more stressful for team management. We didn't want to spend a lot of money on AWS but sometimes (during very hot hours) RAM spikes were 500GB+ to permit simultaneous work.\n\nWe tried to use neptune.ai for ML tracking but from July it was not very effective as we entered in brute force zone. \n\n*Advise: Resources management is very critical - find bottleneck and remove it to make your team most effective. At the same time don't burn money recklessly - limit your budged. If any optimization possible - do it as soon as possible to save time and resources.*\n\n---\n\n### Project Structure\nEach run was internally versioned (ex. v1.1.1 - (major version).(fe version).(model version))\nOverall project structure:\n- Initial preprocess -> artifact cleaned and joined df\n- FE -> Many aligned (by uid) dfs with separated features \n- Features selection -> dictionary we selection metadata\n- Holdout Model (fe check and tuning) -> Local validation oof preds / holdout preds/ model / model metadata\n- Full model run -> Model / Model metadata\n- Prediction -> each fold oof predictions / cv split metadata / test predictions\n\nAll these permitted us to go back and forward and check what worked well and what did not and restore experiments in each particular step.\n\n---\n\n### Initial preprocess \nWe wanted to achieve several things with this step:\n- Join Train and Test -> due to many people involved I was afraid that some missed transformation on private test part will be unnoticed. So we sacrifice memory and speed optimization for overall stability and security.\n- Remove detected noise -> (we had options here but ended with unified single one)\n- Transform Customer ID to unified uid \n- Create internal subset feature -> Train / Public / Private\n- Create unified kfold and holdout split -> To align all experiments\n- Separate columns by type and store them separately to minify memory use and load time\n\n#### Remove detected noise\n\nWe didn't use public notebooks for cleaning. Radar's Dataset is fantastic and it is 99% similar to our own transformations.\nWe used \"isle\" identification without any pre-build coefficients.\ndummy code is something like this:\n\n```\n    for col in process_columns: \n\n        df = temp_df[[col]].sort_values(by=[col]) \n        df = df[df[col].notna()].drop_duplicates(subset=[col]).reset_index(drop=True)\n\n        df['temp'] = np.floor(df[col] * 100000)\n        df['group'] = ((df['temp'] - df['temp'].shift()).abs() >= 100).cumsum()\n\n        i = 0\n        while True:\n            min_val = df[df['group']==i]['temp'].min()\n            if min_val>0:\n                break\n            i += 1\n\n        df['temp2'] = np.where(df['temp']>=0, \n                                np.floor(df['temp']/min_val).astype(np.int32),\n                                np.round(df['temp']/min_val).astype(np.int32))\n        \n        mapping = dict(zip(df[col],df['temp2']))\n        temp_df[col] = temp_df[col].map(mapping)\n\n        print(col, df['group'].nunique(), df[col].nunique())\n        print(df.groupby(['group'])['temp','temp2'].agg(['min','max','count','nunique']).head(40))\n```\n\n#### Create internal subset feature\nWe used last statement month to create 0/1/2 feature and store in in \"index\" df\n\n#### Create unified kfold and holdout split\nFixed random seed (of course 42) to make spits and then took 20% of customers to holdout group (to test stacking / blending / etc)\n\n#### Separate columns by type\nAfter cleaning we had several columns \"groups\".\n```\nall_files = [\n'p_columns', -> just p columns as we thought that they are very different (and P_2 is internal amex \"scoring\" model)\n'objects_radar_columns', -> order encoding (we were checking where out cleaning differs from public approaches and here was the unique place) \n'objects_columns', -> onehot encoding\n'categorical_cleaned__D__columns', -> no noise categoricals\n'categorical_binary__S__columns', -> cleaned binary\n'categorical_binary__R__columns', -> cleaned binary\n'categorical_binary__D__columns', -> cleaned binary\n'categorical_binary__B__columns', -> cleaned binary\n'categorical__D__columns', -> removed noise categoricals\n'categorical__B__columns', -> removed noise categoricals\n'cleaned__B__columns', -> removed noise continuous \n'cleaned__D__columns', -> removed noise continuous \n'cleaned__R__columns', -> removed noise continuous \n'cleaned__S__columns', -> removed noise continuous \n'rest__B__columns', -> have no idea what to do with it -> floor \n'rest__D__columns', -> have no idea what to do with it -> floor \n'rest__R__columns', -> have no idea what to do with it -> floor \n'rest__S__columns', -> have no idea what to do with it -> floor \n]\n```\nThanks again to @raddar we always used your preprocess as a baseline.\n\nWe were able to load just portion of data -> do fe -> concat to \"index\" as all dfs were aligned by index. Also such split permitted us to do fe by feature type to accelerate process and see more statistically valuable metric change. \n\n---\n\n### FE\n\nWe started with careful fe column by column or small subset and it worked well until 1xx features and then any metric improvement or degradation was not statistically significant and many features \"overlapped\" on importance and significance. \n\n*Note: I believe that it is possible to build silver zone robust model with only 3xx features*\n\nSo we started from scratch with brute force))) Of course there was no need to apply \"nunique\" (for example to binary features) and our previous step helped us to limit fe.\n\n```\nall_aggregations = {\n   'agg_func': ['last','mean','std','median','min','max','nunique'],\n   'diff_func': ['first','mean','std','median','min','max'],\n   'ratio_func': ['first','mean','std','median','min','max'],\n   'lags': [1,2,3,6,11],\n   'special': ['ewm','count_month_enc','monotonic_increase','diff_mean','major_class',\n              'normalization','top_outlier','bottom_outlier','normalization_mean','top_outlier_mean','top_outlier_mean']\n}    \n```\n+ pca (horizonal and vertical) + horizontal combinations + horizontal aggregations.\n\nagg_func -> normal aggregations by uid\ndiff_func -> diff last - xxx -> std diff worked better than any other\nratio_func -> ratio_func last/xxx\nlags -> diff last - Nx\nspecial -> some special transformations -> count_month_enc worked well for categorical / emw for continous \n\nWe ended up with about 7k features (stored file by group and by agg type for faster loading). \nNext thing was to figure out what works and what not -> this topic was the most challenging for us. \n\n#### Normalizations\n\nIt's better to call it Standardization (x - m) / s -> as we had also normalization test the name became constant \"normalization\")))\n```\ndf.groupby(['dt_month','subset'])[col].agg(['mean','std'])\n```\ndt_month -> month of the statement\nsubset -> train / public / private\nand mean and std from clients that had full statement history.\n\nWe have to have temporal shift to make it work. So we did a \"trick\" removed last statement for each client and applied exactly same transformation  for each client and merged appropriate labels. So we had 2 lines in training set for almost each client BUT validated results only on last statement during CV runs and Holdout checks. It more or less same as adding noised data but we had temporal drift and model was able to work better on unknown future data with  \"possible\" data drift.\n\n---\n\n### Features selection\n\nOoohh that was really fun. \n\nWe used gbdt boosting type during experiments as it was very aligned with dart mode but was significantly faster.\nAlso, we used ROC AUC score during our experiments as we believed that due to amex instability we can't use it for decision making  (of course we tracked log loss and amex).\n\nIn previous step we brute forced many features and now is time to clean them out.\nAll feature selection was done with 5 CV folds training + independent check on 20% holdout data.\n\n1. Zero importance -> Right from the start we were able to through away 1.5k features that had exactly 0 importance (lgbm importance). That means that with 250 bins and 2**10 data in leaf those features are not participating in any split.\n\n2. Stepped hierarchical permutation importance -> we defined 300 initial features and looped over all other features subsets (600+) - was very time consuming but very stable.\nNote: we shuffled order of the subset to force model try different combinations.\nAdd features subset -> train model -> permutate -> drop negative features (negative mean over 5 seeds) -> add new subset -> ...\nDuring this part that took almost 3 days we limited features to 3k -> 0.800 lb\n\n3. Stepped permutation importance.\nTake all features -> train model -> permutate -> drop 20% of worst performed features (only negative) -> repeat. Final subset was 25xx features (and different from previous step) -> 0.800 lb\n\n4. Forward feature selection.\nWe defined 300 initial features and simply added subset by subset and compared ROC AUC if metric change was > 0.0003 we kept the subset. -> 0.800 lb\n\n5. Time series CV.\nFor very doubtful features as PCA and Normilized values we used to different validation stratagies:\n- Train on first 6 month values (last statement of the first 6 months went to train set) and validate on last 6 (also just last statement of the last 6 months). We trained model without temporal feature and then with if result was better on CV and on holdout we added to final features subset.\n- We used P_2, B_1, B_2 as a proxy target and MSE loss with combined Train and Test to see if we did right transformation and result did not degrade.\n\nMany other options we tried but result was not stable.\n\nFinal subset came from \"Forward feature selection\" plus overlapped features from other technics minus overlapped negative combination.  -> lb 0.801 single model.\n\nWe tried to blend many models with different subset as we believed that it should give huge LB boost (based on holdout blending tests) but it didn't work well for lb. \n\n--- \n\n### Model \n\nIn my own experience, DART never worked better and here we have proof that in DS \"all depends.\" We did experiments with DART in the beginning and it did not show any metric improvement with our params and baseline model features subset. Later we found @ragnar123 notebook and gave it one more try and it worked marvelously.\n\nFrom the beginning, we tried to build a more complex model with 2**7+ leaves and 0.7+ features but failed. It still puzzles me why a simple model with a very low number of features works here. \n\nI saw such behaviour mostly on synthetic data and stacking - so we tried to find out if data is syntetic (at least partly) and deanonimize internal scoring values - but didn't make it.\n\nOur best single lgbm model was trained on 29xx features. 5 folds CV - no stratification by any option. Training data - 2 last staements for each client (transformed independently). Params:\n\n```\nlgb_params = {\n    'boosting_type': 'dart',\n    'objective': 'cross_entropy', \n    'metric': ['AUC'],\n    'subsample': 0.8,  \n    'subsample_freq': 1,\n    'learning_rate': 0.01, \n    'num_leaves': 2 ** 6, \n    'min_data_in_leaf': 2 ** 11, \n    'feature_fraction': 0.2, \n    'feature_fraction_bynode':0.3,\n    'first_metric_only': True,\n    'n_estimators': 17001,  # -> 5000 for gbdt \n    'boost_from_average': False,\n    'early_stopping_rounds': 300,\n    'verbose': -1,\n    'num_threads': -1,\n    'seed': SEED,\n}\n```\n\nBlend -> Power (2) rank blend of Dart lgbm (0.801 public) / GBDT lgbm  (0.799 public) / Catboost models (0.799 public)\n\nSingle lgbm with 3 last statements showed even better CV by we didn't have enough time to retrain it (full DART run for 5 folds took 12+ hours there).\n\nIt was obvious that clients with a little number of statements will not get benefit from all 2k features. So we created a special model that was trained only on 300 features with custom params (also dart). Predictions for clients with <=2 statements came exclusively from such model and were not blended with other models. \n\nHow did we combine the result from 2 independent models to not destroy the final ranking?\n\n| Client id | Number of statements | Basic ranking | <=2 prediction | Final ranking |\n| --- | --- | --- | --- | --- |\n| 1 | 13 | 5 | ... | 5 |\n| 1 | 2 | 4 | 0.1 | 3 |\n| 1 | 13 | 8 | ... | 8 |\n| 1 | 1 | 3 | 0.5 | 4 |\n| 1 | 13 | 7 | ... | 7 |\n| 1 | 13 | 2 | ... | 2 |\n\nwe kept ranking for >2 statements and for rest resorted within initial ranking group (hope it's clear enough))). Was is perfect - no, but it was very stable with really tiny improvement (because of number of such clients in public and private test parts).\n\nWhat also worked well:\n- Train on all data without folds splitting and stop just 1000 rounds further then CV showed. \n\nWhat didn't work:\n-  Stacking by any mean\n-  Many models with different seed and fe order\n- Massive blend of different models with different types and features (blend worked well till 0.799 and then any low performed model 0.795- made public score worse - we use public CV as additional holdout set and used approaches that worked well on local CV and LB - anything that worked partially was not used in the end).\n\nAgain, nothing really fancy here. The main thing that helped us align Train / CV with LB was very high  'min_data_in_leaf'.\nNo optuna used -> just manual old school tuning based on data feeling.\nWe did many experiments with weights and loss functions but none of them worked.\n\nDue to AMEX metric specification it was obvious that focal loss should work but it didn't. We tried several times to switch loss function during the competition period and the result was the same. \n\nError analysis showed that model makes errors without any \"pattern\" -> stacking didn't work for holdout set (25% of data)  and we had doubts that it will work on private/public test parts. We kept only LR/Lasso(0.02) for blending options to choose submissions.\n\nCross validation -> standard 5folds CV split by client ID. The unique thing that we did here is \"prespliting\" to align all CV between team members to be able to compare results directly.\n \n---\n\nWhat left without mentions:\n- EDA on data\n- Denoising experiments\n- Data deanonymization -> didn't manage to make it\n- Features pairs and triples combinations -> that didn't work well\n- NaN filling -> didn't work\n- Clusterization -> didn't work\n- Hundreds of experiments with features selection process and internal discussions about it.\n- Adding noised data (noise / swap noise) that leaded to interesting but doubtful results\n- Model tuning \n- Removing absolute values and keep only diff or ratios -> should be more stable for future data but we saw some lb degradation and didn't proceed\n- pseudo labeling\n\nWhat we always wanted but didn't found time to do:\n- NN - we have no NN in our final blend\n- P_2 or any other column prediction (1/2/3/4 months ahead) with combined data and use it as meta information for lgbm main model\n- 11 / 12 / 13 statements joined training on different subsets (df was too large and training was slow)\n\n--- \n\n### Internal initial plan\n\n```\n########################### Data preprocessing and Data evaluation\n#################################################################################\n\n## Added noise removal -> GOOD2DO\n# There is no doubt that some Noise was injected in data\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\n# We need to find a way to remove it \n# the best option to not follow public approach\n# At least with columns where columns have overlaped population\n\n## Data minification for FE -> GOOD2DO\n# Datatype downcasting\n# Pickle/Parquet/Feather \n# Be careful with floats16 as it may lead to bad agg results\n# Also float16 may lead to some signal degradation due to precision and values changes\n\n## Evaluate values distributions and NaNs -> GOOD2DO\n# Full 13 months history\n# Train against Test Public and Test Private\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\n#\n# Need to try:\n# Kolmogorov–Smirnov test -> GOOD2DO\n#\n# Adversarial validation -> GOOD2DO\n# https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation/notebook (just as simple example)\n#\n# Entropy / Distances / etc....\n#\n# Visual checks)))\n#\n# We need to find if ANY feature has very different distribution in PRIVATE test set\n# If that feature works for Public part it doesn't mean that it will work for Private\n\n########################### Targets\n#################################################################################\n\n# We need to find a way to get more targets -> GOOD2DO\n# as we currently training on a single point by client we could greatly improve results\n# by extending our training set with new targets\n#\n# Find default periods in current client history and make appropriate labeling -> GOOD2DO\n# Make 2 level model -> predict p_2 values as normal time-series model and feed it to 2nd level GBT\n\n########################### Separate Models\n#################################################################################\n\n# Probably it's a good idea to make separate models for each subdatasets\n# Full history (13 months)\n# Less than 13 months\n\n########################### External Data\n#################################################################################\n# We can try to add \"Consumer index\" or any other independent temporal feature\n# Will not work if we will not be able to expand targets and add temporal feature\n\n########################### FE\n#################################################################################\n\n# We didn't make anything special here -> July\n# AGGS (Stats by client)\n# Rollings\n# History length feature (not sure if it will help with Private Test)\n# Should we correct statements dates and add NaNs?\n# ReRanking categorical features by P_2 or Target\n# Clusterization (4+ groups feature by feature)\n# Count and Mean encodings for categorical features\n# Features combinations (sum/prod/power) -> bruteforce\n# PCA or any other dimension reduction by features groups\n\n# We need to find if there is \"connection\" between clients in Train -> Public Test -> Private Test\n# we have 458913 + 924621 -> 1383534 If I were AMEX I would export 1M clients (or other round number)\n# so may be 384 534 Clients are overlaps\n\n# Clip by 5 - 95 percentile\n\n########################### Features Selection\n#################################################################################\n# Permutation importance (use all fold only!!!) -> recursive elimination (because of quantity of features -> 3-4 rounds with 0 and 50% negative drop) \n# SHAP\n# Highly correlated features (.98+?)\n# Forward selection (may take ages and due aggs may be not effective - probably by feature block) \n# Backward elimination (may take ages and due aggs may be not effective - probably by feature block) \n\n########################### CV\n#################################################################################\n# Mean Target differs my \"history length\" -> could be wise to do GroupedStratifeidFolds by history length\n# For sure Splits should be done by client\n# Target stratification to balance folds\n\n########################### Loss function / Metric\n#################################################################################\n# Clean and fast np/torch metric\n# Now it's in helper (need to cleanup that)\n# https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\n#\n# I don't believe that we will have better results with different loss function\n# But it worth to try at least focal loss\n# https://maxhalford.github.io/blog/lightgbm-focal-loss/\n#\n# Weights -> we should try change weights there\n# weights by class\n# weights by some history length\n# weights by internal fe group\n#\n# We need custom metric for catboost\n# example https://catboost.ai/en/docs/concepts/python-usages-examples#logloss1\n\n########################### Models\n#################################################################################\n\n## First level choice\n# LGB/XGB/CTB -> our main models here for sure\n# After stabilizing the baseline model and base feature we need to make 1st round tuning\n\n## Catboos specials\n# Categorical features\n# Embeding features\n\n## NN (GPU/TPU) -> RNN / LSTM / Transformer\n# TPU -> tensorflow (as it works better there)\n\n## NN -> AE / VAE / DAE -> as a denoising model hidden layer as input for GBT models\n# No need complex approach - just fast check the idea and in case of success move to big model\n\n########################### Blending\n#################################################################################\n# Weighted Average\n# Power Average\n# Weighted Rank Average\n# Linear/SVM\n# Postprocessing?\n```",
      "votes": 180
    },
    {
      "id": 1912746,
      "postDate": "2022-08-25T00:32:24.170Z",
      "content": "<p><a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a> Congratulations Konstantin and team. I am so happy for you. You are the best tabular data feature engineer that I know. I'm glad to see your skills achieved great success! I'm looking forward to learning about your solution!</p>",
      "rawMarkdown": "@kyakovlev Congratulations Konstantin and team. I am so happy for you. You are the best tabular data feature engineer that I know. I'm glad to see your skills achieved great success! I'm looking forward to learning about your solution!",
      "votes": 13,
      "replies": [
        {
          "id": 1912768,
          "postDate": "2022-08-25T00:52:04.023Z",
          "content": "<p>I think you groom him well from IEEE competition 3 years back :)  Congratulations to you both ! </p>",
          "rawMarkdown": "I think you groom him well from IEEE competition 3 years back :)  Congratulations to you both ! ",
          "votes": 1
        },
        {
          "id": 1912775,
          "postDate": "2022-08-25T00:57:00.003Z",
          "content": "<p>haha, he groomed me well. I was new to Kaggle then and he invited me to his team when I was like 30th place and he was 1st place! He taught me so much about tabular data and feature engineering. Many groupby aggregate tricks that I use today I learned from Konstantin. Thanks Konstantin!</p>",
          "rawMarkdown": "haha, he groomed me well. I was new to Kaggle then and he invited me to his team when I was like 30th place and he was 1st place! He taught me so much about tabular data and feature engineering. Many groupby aggregate tricks that I use today I learned from Konstantin. Thanks Konstantin!",
          "votes": 11
        },
        {
          "id": 1912845,
          "postDate": "2022-08-25T02:04:13.323Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> it's a huge honor to hear such words from you.</p>\n<p>Right from the start we decided to make \"internal\" team for this competition but I was so many times close to break that rule and ask you to join us - I was sure that you had fantastic and a very unique NN approach (and you had it) - congrats with solo gold with transformers. </p>",
          "rawMarkdown": "Thank you @cdeotte it's a huge honor to hear such words from you.\n\nRight from the start we decided to make \"internal\" team for this competition but I was so many times close to break that rule and ask you to join us - I was sure that you had fantastic and a very unique NN approach (and you had it) - congrats with solo gold with transformers. ",
          "votes": 6
        },
        {
          "id": 1912964,
          "postDate": "2022-08-25T04:28:29.393Z",
          "content": "<p>Well <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I think a small notebook about groupby tricks would be your best gift for us after this competition :)</p>",
          "rawMarkdown": "Well @cdeotte I think a small notebook about groupby tricks would be your best gift for us after this competition :)",
          "votes": 3
        },
        {
          "id": 1913157,
          "postDate": "2022-08-25T07:27:17.260Z",
          "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>  he already made an elaborate post about feature engineering:<br>\n<a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575\" target=\"_blank\">https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575</a><br>\nIt's probably the highest voted discussion post of all times!<br>\nGroupby is under \"aggregation\"</p>",
          "rawMarkdown": "@mohammad2012191  he already made an elaborate post about feature engineering:\nhttps://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575\nIt's probably the highest voted discussion post of all times!\nGroupby is under \"aggregation\"",
          "votes": 3
        },
        {
          "id": 1913165,
          "postDate": "2022-08-25T07:34:41.150Z",
          "content": "<p>Thanks, gonna check it.</p>",
          "rawMarkdown": "Thanks, gonna check it.",
          "votes": 1
        },
        {
          "id": 1913550,
          "postDate": "2022-08-25T11:06:44.673Z",
          "content": "<p>Yes, the link Yukiya provides has a list of good engineer techniques. </p>\n<p>And here is a link which explains why groupby works <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111453\" target=\"_blank\">here</a>. When we create new columns using groupby from other columns, we give our GBT new ways to classify rows that belong to different groups.</p>",
          "rawMarkdown": "Yes, the link Yukiya provides has a list of good engineer techniques. \n\nAnd here is a link which explains why groupby works [here][1]. When we create new columns using groupby from other columns, we give our GBT new ways to classify rows that belong to different groups.\n\n[1]: https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111453",
          "votes": 2
        },
        {
          "id": 1914175,
          "postDate": "2022-08-25T19:53:43.670Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1920789,
      "postDate": "2022-08-31T12:02:27.490Z",
      "content": "<p>A huge congratulations and thanks for such a detail explanation. As Chris said, the best tabular feature engineer! </p>\n<p>Some questions from my side:</p>\n<p><strong>1.</strong> Based on your experience, doing brute force feature engineering (i.e. create a ton of potential new features and then perform feature selection) gives better results than going column by column?<br>\n<strong>2.</strong> How did you choose the initial LGBM hyperparameters before starting the feature selection phase? <br>\n<strong>3.</strong> Did you retune the hyperparameters at some point of the feature selection phase (e.g. when you added/removed a significant number of features) or kept them fixed all the time?<br>\n<strong>4.</strong> Could you give some examples of what type of feature subsets you include at each step of the feature selection phase (e.g. one subset could be the mean of <code>cleaned__B__columns</code>)?<br>\n<strong>5.</strong> Your final selected features came from \"Forward feature selection\". I tried it a couple of times at work and saw that tends to overfit, do you have the same experience?<br>\n<strong>6.</strong> How did you realize that increasing the value of <code>min_data_in_leaf</code> favour CV/LB alignment?</p>\n<p>Also waiting for the code 👀</p>",
      "rawMarkdown": "A huge congratulations and thanks for such a detail explanation. As Chris said, the best tabular feature engineer! \n\nSome questions from my side:\n\n**1.** Based on your experience, doing brute force feature engineering (i.e. create a ton of potential new features and then perform feature selection) gives better results than going column by column?\n**2.** How did you choose the initial LGBM hyperparameters before starting the feature selection phase? \n**3.** Did you retune the hyperparameters at some point of the feature selection phase (e.g. when you added/removed a significant number of features) or kept them fixed all the time?\n**4.** Could you give some examples of what type of feature subsets you include at each step of the feature selection phase (e.g. one subset could be the mean of `cleaned__B__columns`)?\n**5.** Your final selected features came from \"Forward feature selection\". I tried it a couple of times at work and saw that tends to overfit, do you have the same experience?\n**6.** How did you realize that increasing the value of `min_data_in_leaf` favour CV/LB alignment?\n\nAlso waiting for the code 👀",
      "votes": 6,
      "replies": [
        {
          "id": 1920865,
          "postDate": "2022-08-31T13:01:08.950Z",
          "content": "<p>Thank you for your kind words and very nice questions.</p>\n<p>1 | Brute force is a dead end by all meanings - it will not work on production, it will not provide any interpretability or explainability, it is not reproducible result. Column by column and distilled carefully created features are always better and preferable. Even in this competition, manual feature engineering would work better if we could have access to raw data and columns description (meaning).</p>\n<hr>\n<p>2 | Initial params always depend on data, but I start most of times with 0.7/0.7 sampling/features sampling, 2^7 leaves and 2^8 min data in leaf (all other params as default). <br>\nFirst baseline run normally gives you a lot of information how to tune it further:</p>\n<ul>\n<li>if train score is much higher than CV and holdout -&gt; try regularizations / lower subsampling / less complex model</li>\n<li>if naive score is better than lgbm or difference is minimal -&gt; bad features or/and need more complex model (2**7 is complex enough for majority of tasks)</li>\n<li>if train score is low but very aligned with CV/holdout -&gt; underfitting -&gt; more complex model with less data in leaf<br>\nThen try to \"tweak\" loss function that is more appropriate for your task.<br>\nThen try to fit goss/dart/extra_trees -&gt; extra_trees works very nice sometimes -&gt; DART never worked better before)))</li>\n</ul>\n<hr>\n<p>3 | Params tuning normally takes several iterations during model creation - for competitions it means a slight tuning every 2 weeks (after significant feature sets changes or data structure changes) and final deep tuning round 2 weeks before competition ends.</p>\n<hr>\n<p>4 | Not much to add here - you are right with your assumption:<br>\nmean + cleaned__B__columns produces:<br>\nagg_func__mean__cleaned__B__columns|__B_4__mean<br>\nagg_func__mean__cleaned__B__columns|__B_16__mean<br>\nagg_func__mean__cleaned__B__columns|__B_20__mean<br>\nagg_func__mean__cleaned__B__columns|__B_22__mean<br>\nagg_func__mean__cleaned__B__columns|__B_41__mean<br>\n…</p>\n<hr>\n<p>5 | It works really nice for the start but with more features it becomes less informative and slower. It didn't show any overfitting but we didn't trust it fully as we were training and validating on the same time period (and didn't have any aligned validation set in the “future”).</p>\n<hr>\n<p>6 | We saw that ROC AUC score on train data went up to 0.9999 and CV/Holdout had cap of 0.96xxx -&gt; many teams decided to use regularization (l1 - punish features with low importance but we would like to keep even small signals, l2 - adds regularization to features with high importance -&gt; that is good move here but we didn't see any statistically significant CV score boost). We decided to make less complex model with 2^6 leaves and increase min data in leaf to force our model generalize better and found that increasing this param to 2^9 - 2^11 works very well on CV/holdout and also worked well on LB. </p>",
          "rawMarkdown": "Thank you for your kind words and very nice questions.\n\n\n1 | Brute force is a dead end by all meanings - it will not work on production, it will not provide any interpretability or explainability, it is not reproducible result. Column by column and distilled carefully created features are always better and preferable. Even in this competition, manual feature engineering would work better if we could have access to raw data and columns description (meaning).\n\n---\n\n2 | Initial params always depend on data, but I start most of times with 0.7/0.7 sampling/features sampling, 2^7 leaves and 2^8 min data in leaf (all other params as default). \nFirst baseline run normally gives you a lot of information how to tune it further:\n- if train score is much higher than CV and holdout -> try regularizations / lower subsampling / less complex model\n- if naive score is better than lgbm or difference is minimal -> bad features or/and need more complex model (2**7 is complex enough for majority of tasks)\n- if train score is low but very aligned with CV/holdout -> underfitting -> more complex model with less data in leaf\nThen try to \"tweak\" loss function that is more appropriate for your task.\nThen try to fit goss/dart/extra_trees -> extra_trees works very nice sometimes -> DART never worked better before)))\n\n---\n\n3 | Params tuning normally takes several iterations during model creation - for competitions it means a slight tuning every 2 weeks (after significant feature sets changes or data structure changes) and final deep tuning round 2 weeks before competition ends.\n\n---\n\n4 | Not much to add here - you are right with your assumption:\nmean + cleaned__B__columns produces:\nagg_func__mean__cleaned__B__columns|__B_4__mean\nagg_func__mean__cleaned__B__columns|__B_16__mean\nagg_func__mean__cleaned__B__columns|__B_20__mean\nagg_func__mean__cleaned__B__columns|__B_22__mean\nagg_func__mean__cleaned__B__columns|__B_41__mean\n...\n\n---\n\n5 | It works really nice for the start but with more features it becomes less informative and slower. It didn't show any overfitting but we didn't trust it fully as we were training and validating on the same time period (and didn't have any aligned validation set in the “future”).\n\n---\n\n6 | We saw that ROC AUC score on train data went up to 0.9999 and CV/Holdout had cap of 0.96xxx -> many teams decided to use regularization (l1 - punish features with low importance but we would like to keep even small signals, l2 - adds regularization to features with high importance -> that is good move here but we didn't see any statistically significant CV score boost). We decided to make less complex model with 2^6 leaves and increase min data in leaf to force our model generalize better and found that increasing this param to 2^9 - 2^11 works very well on CV/holdout and also worked well on LB. \n \n ",
          "votes": 8
        },
        {
          "id": 1923496,
          "postDate": "2022-09-02T09:16:25.547Z",
          "content": "<p>Thanks for such a complete answer!</p>",
          "rawMarkdown": "Thanks for such a complete answer!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1913704,
      "postDate": "2022-08-25T13:06:04.490Z",
      "content": "<p>It's becoming a long writeup))) will need bit more time</p>",
      "rawMarkdown": "It's becoming a long writeup))) will need bit more time",
      "votes": 5,
      "replies": [
        {
          "id": 1914039,
          "postDate": "2022-08-25T17:15:31.977Z",
          "content": "<p>Take your time! Great details already</p>",
          "rawMarkdown": "Take your time! Great details already",
          "votes": 1
        }
      ]
    },
    {
      "id": 1915111,
      "postDate": "2022-08-26T17:11:51.047Z",
      "content": "<p>Huge congrats Konstantin <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a>, this is incredible commitment your team had! I subscribed 2 months of Colab Pro+ for this comp and that was it…<br>\nCurious if you know besides <code>diff_mean</code> which of your <code>'special'</code> agg ideas worked best for you?</p>",
      "rawMarkdown": "Huge congrats Konstantin @kyakovlev, this is incredible commitment your team had! I subscribed 2 months of Colab Pro+ for this comp and that was it...\nCurious if you know besides `diff_mean` which of your `'special'` agg ideas worked best for you?",
      "votes": 3,
      "replies": [
        {
          "id": 1915119,
          "postDate": "2022-08-26T17:16:48.673Z",
          "content": "<p>btw diff_mean is: </p>\n<pre><code>df[f'{col}__diff_mean] = df[col] - df.groupby(['uid'])[col].shift()\ndf[f'{col}__diff_mean] = df.groupby(['uid'])[f'{col}__diff_mean]].transform('mean')\n\ndf = df.drop_duplicates(subset=['uid'], keep='last').reset_index(drop=True)\n</code></pre>\n<p>std features worked better than others - diff with std for example</p>\n<p>normal diff with client mean was in <br>\n'diff_func': ['first','mean','std','median','min','max']</p>\n<p>and my bad - which of specials.<br>\nmonotonic functions worked really well</p>",
          "rawMarkdown": "btw diff_mean is: \n```\ndf[f'{col}__diff_mean] = df[col] - df.groupby(['uid'])[col].shift()\ndf[f'{col}__diff_mean] = df.groupby(['uid'])[f'{col}__diff_mean]].transform('mean')\n\ndf = df.drop_duplicates(subset=['uid'], keep='last').reset_index(drop=True)\n```\n\nstd features worked better than others - diff with std for example\n\n\nnormal diff with client mean was in \n'diff_func': ['first','mean','std','median','min','max']\n\nand my bad - which of specials.\nmonotonic functions worked really well\n",
          "votes": 2
        },
        {
          "id": 1915208,
          "postDate": "2022-08-26T18:36:20.317Z",
          "content": "<p>Thanks for the reply! <br>\n<code>monotonic functions</code> - Did you create a T/F binary feature or did you calculate for example spearman correlation or similar to quantify monotonicity?  </p>",
          "rawMarkdown": "Thanks for the reply! \n`monotonic functions` - Did you create a T/F binary feature or did you calculate for example spearman correlation or similar to quantify monotonicity?  "
        },
        {
          "id": 1915211,
          "postDate": "2022-08-26T18:38:15.713Z",
          "content": "<p>Binary if current more or equal to previous and then took mean.</p>\n<p>We wanted to pass to the model information if last statements increase or decrease is normal for the client or not. If due payments are constantly increasing for us it was a sign that risk is growing.</p>\n<p>And such information is not based on absolute values - we constantly wanted to avoid absolute values to be able to generalize well on future data and data with drifts.</p>",
          "rawMarkdown": "Binary if current more or equal to previous and then took mean.\n\nWe wanted to pass to the model information if last statements increase or decrease is normal for the client or not. If due payments are constantly increasing for us it was a sign that risk is growing.\n\nAnd such information is not based on absolute values - we constantly wanted to avoid absolute values to be able to generalize well on future data and data with drifts.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1912725,
      "postDate": "2022-08-25T00:21:14.897Z",
      "content": "<p>Congrats guys .. cant wait for your solution to reach 0.809 .. </p>",
      "rawMarkdown": "Congrats guys .. cant wait for your solution to reach 0.809 .. ",
      "votes": 3
    },
    {
      "id": 1918704,
      "postDate": "2022-08-29T18:58:54.670Z",
      "content": "<p>very insightful tank youu</p>",
      "rawMarkdown": "very insightful tank youu",
      "votes": 1
    },
    {
      "id": 1917682,
      "postDate": "2022-08-29T01:25:08.577Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 1917052,
      "postDate": "2022-08-28T11:47:50.033Z",
      "content": "<p>Here is example how we augmented training Data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F7db49977cfebf34a15d87e5b4c95ec3f%2F2022-08-28%20%2012.41.40.png?generation=1661686959865932&amp;alt=media\" alt=\"\"></p>\n<p>We were removing last statement per client as it would never exists and then applied exactly the same transformations as we did for full history.</p>\n<ul>\n<li>Last values are different but within distributions </li>\n<li>Agg values are different but within distributions</li>\n<li>Model has temporal component </li>\n</ul>\n<p>We wanted to completely avoid absolute values and with data augmentation the CV and LB were almost on the same level as with absolute values - of course it would not work with clients with 1-2 statements but we had a separate model for them. I believe that that would give much more robust results but on kaggle even slight metric increase is important and we kept \"lasts\".</p>\n<p>Also, we were planning to go in -2 max depth and used only 1,2,3,6,11 (most common time-series lags) diff lags to have fe aligned. </p>\n<p>Combining original data with -1 statement or with -2 statement gave significant CV boost on ROC AUC score (our main metric) and tiny boost on AMEX metric. Combination of 3 sets gave just a tiny boost on ROC AUC but was significantly more memory consuming and slower to train model.</p>\n<p>Also, CV splits per original client uid to exclude leaks and validation was performed only on original data to exclude metric degradation tracking for shorten history predictions.</p>",
      "rawMarkdown": "Here is example how we augmented training Data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F7db49977cfebf34a15d87e5b4c95ec3f%2F2022-08-28%20%2012.41.40.png?generation=1661686959865932&alt=media)\n\nWe were removing last statement per client as it would never exists and then applied exactly the same transformations as we did for full history.\n\n- Last values are different but within distributions \n- Agg values are different but within distributions\n- Model has temporal component \n\nWe wanted to completely avoid absolute values and with data augmentation the CV and LB were almost on the same level as with absolute values - of course it would not work with clients with 1-2 statements but we had a separate model for them. I believe that that would give much more robust results but on kaggle even slight metric increase is important and we kept \"lasts\".\n\nAlso, we were planning to go in -2 max depth and used only 1,2,3,6,11 (most common time-series lags) diff lags to have fe aligned. \n\nCombining original data with -1 statement or with -2 statement gave significant CV boost on ROC AUC score (our main metric) and tiny boost on AMEX metric. Combination of 3 sets gave just a tiny boost on ROC AUC but was significantly more memory consuming and slower to train model.\n\nAlso, CV splits per original client uid to exclude leaks and validation was performed only on original data to exclude metric degradation tracking for shorten history predictions.",
      "votes": 1
    },
    {
      "id": 1917023,
      "postDate": "2022-08-28T11:13:19.733Z",
      "content": "<p>Few examples how noise cleaning works:<br>\nFor the most features nunique groups within 998 -&gt; 1001 is a sign of noised cleaning possibility.</p>\n<p>Simple feature cleaning - in this case group is our new value:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2Fea9cd72dbd9666335a1194331479fbc8%2F2022-08-28%20%2012.05.28.png?generation=1661684893433642&amp;alt=media\" alt=\"\"></p>\n<p>More complex feature with negative \"categories\" we keep negative values (temp2 is a new value):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F652007f94ba3b47bc300dfe32da85ce1%2F2022-08-28%20%2012.06.17.png?generation=1661684841167050&amp;alt=media\" alt=\"\"></p>\n<p>Very \"noised\" feature (we find initial coefficient and then apply it to groups with low population) - temp2 is a new value:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F2d811e27de9544d6eaeca47cf268b650%2F2022-08-28%20%2012.11.55.png?generation=1661685137975362&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Few examples how noise cleaning works:\nFor the most features nunique groups within 998 -> 1001 is a sign of noised cleaning possibility.\n\nSimple feature cleaning - in this case group is our new value:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2Fea9cd72dbd9666335a1194331479fbc8%2F2022-08-28%20%2012.05.28.png?generation=1661684893433642&alt=media)\n\nMore complex feature with negative \"categories\" we keep negative values (temp2 is a new value):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F652007f94ba3b47bc300dfe32da85ce1%2F2022-08-28%20%2012.06.17.png?generation=1661684841167050&alt=media)\n\nVery \"noised\" feature (we find initial coefficient and then apply it to groups with low population) - temp2 is a new value:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F2d811e27de9544d6eaeca47cf268b650%2F2022-08-28%20%2012.11.55.png?generation=1661685137975362&alt=media)",
      "votes": 1
    },
    {
      "id": 1915380,
      "postDate": "2022-08-26T23:54:33.063Z",
      "content": "<p>Thanks for sharing your solution! Very insightful.</p>",
      "rawMarkdown": "Thanks for sharing your solution! Very insightful.",
      "votes": 1
    },
    {
      "id": 1913980,
      "postDate": "2022-08-25T16:05:05.120Z",
      "content": "<p>Normalizations. Sorry, I'm having trouble understanding. For all EXCEPT last month you took mean and std for all customers fort that exact col, month and subset (train/pub/pri)? What about last? I'm actually super confused what the trick was, and why? Maybe try just restating/rephrasing it and maybe I'll figure it or at least have more precise questions…</p>",
      "rawMarkdown": "Normalizations. Sorry, I'm having trouble understanding. For all EXCEPT last month you took mean and std for all customers fort that exact col, month and subset (train/pub/pri)? What about last? I'm actually super confused what the trick was, and why? Maybe try just restating/rephrasing it and maybe I'll figure it or at least have more precise questions...",
      "votes": 1,
      "replies": [
        {
          "id": 1920719,
          "postDate": "2022-08-31T11:00:00Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1913512,
      "postDate": "2022-08-25T10:49:07.533Z",
      "content": "<p>Congratulation! Looking forward to dive into your solution!</p>",
      "rawMarkdown": "Congratulation! Looking forward to dive into your solution!",
      "votes": 1
    },
    {
      "id": 1913260,
      "postDate": "2022-08-25T09:03:27.457Z",
      "content": "<p>Good Job! well deserved</p>",
      "rawMarkdown": "Good Job! well deserved",
      "votes": 1
    },
    {
      "id": 1913210,
      "postDate": "2022-08-25T08:18:53.740Z",
      "content": "<p>Congrats to the team. One more waiting for the extended explanation</p>",
      "rawMarkdown": "Congrats to the team. One more waiting for the extended explanation",
      "votes": 1
    },
    {
      "id": 1913202,
      "postDate": "2022-08-25T08:11:16.213Z",
      "content": "<p>Congratulations ! :)</p>",
      "rawMarkdown": "Congratulations ! :)",
      "votes": 1
    },
    {
      "id": 1912995,
      "postDate": "2022-08-25T04:56:41.570Z",
      "content": "<p>Congratuation!</p>",
      "rawMarkdown": "Congratuation!",
      "votes": 1
    },
    {
      "id": 1912963,
      "postDate": "2022-08-25T04:28:18.733Z",
      "content": "<p>Glad to see you back on Kaggle competitions with a bang. Congratulations, amazing finish. </p>",
      "rawMarkdown": "Glad to see you back on Kaggle competitions with a bang. Congratulations, amazing finish. ",
      "votes": 1
    },
    {
      "id": 1912935,
      "postDate": "2022-08-25T03:57:12.200Z",
      "content": "<p>Congratulations! Thank you for sharing model params!! Very useful for next competitions!</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing model params!! Very useful for next competitions!",
      "votes": 1
    },
    {
      "id": 1912867,
      "postDate": "2022-08-25T02:34:02.830Z",
      "content": "<p>Congratulations! 🎉  Not a lot of experience on Kaggle so not sure if this is normal but was surprised to see that 3/4 of your team is novices and contributors. Gives me hope 😅</p>",
      "rawMarkdown": "Congratulations! 🎉  Not a lot of experience on Kaggle so not sure if this is normal but was surprised to see that 3/4 of your team is novices and contributors. Gives me hope 😅",
      "votes": 1,
      "replies": [
        {
          "id": 1915182,
          "postDate": "2022-08-26T18:14:18.977Z",
          "content": "<p><a href=\"https://www.kaggle.com/shahilap96\" target=\"_blank\">@shahilap96</a> our team was formed at the very early stage of the competition and exclusively from my company DS department - many other DS colleagues wanted to participate but were not able to made such 3month commitment and decided not to join. Kaggle experience is nice to have but it’s not an obligatory condition - the most important is a “fire in the eyes” and passion to learn and try something new.</p>\n<p>Yes, my teammates are novice on kaggle but not novice in analytics and DS. </p>\n<p>One thing that I can say - all my teammates work with the same passion with regular work tasks and trying really hard to achieve goals. I feel very proud of them.</p>",
          "rawMarkdown": "@shahilap96 our team was formed at the very early stage of the competition and exclusively from my company DS department - many other DS colleagues wanted to participate but were not able to made such 3month commitment and decided not to join. Kaggle experience is nice to have but it’s not an obligatory condition - the most important is a “fire in the eyes” and passion to learn and try something new.\n\nYes, my teammates are novice on kaggle but not novice in analytics and DS. \n\nOne thing that I can say - all my teammates work with the same passion with regular work tasks and trying really hard to achieve goals. I feel very proud of them.",
          "votes": 4
        },
        {
          "id": 1916523,
          "postDate": "2022-08-28T00:39:14.290Z",
          "content": "<p>Aha, that explains the novice part 😅 People with passion and fire in their eyes are always great to work with. The energy and drive always pushes me to do better. Congrats once again.</p>",
          "rawMarkdown": "Aha, that explains the novice part 😅 People with passion and fire in their eyes are always great to work with. The energy and drive always pushes me to do better. Congrats once again."
        }
      ]
    },
    {
      "id": 1912798,
      "postDate": "2022-08-25T01:14:05.613Z",
      "content": "<p>Congrats!<br>\nI'm waiting for your notebook to knwo what had been updated from IEEE competition :)</p>",
      "rawMarkdown": "Congrats!\nI'm waiting for your notebook to knwo what had been updated from IEEE competition :)",
      "votes": 1
    },
    {
      "id": 1912764,
      "postDate": "2022-08-25T00:49:01.710Z",
      "content": "<p>Congratulations！</p>",
      "rawMarkdown": "Congratulations！",
      "votes": 1
    },
    {
      "id": 1912741,
      "postDate": "2022-08-25T00:29:36.860Z",
      "content": "<p>Congratulations, guys, you deserved it! Really well done reaching this score</p>",
      "rawMarkdown": "Congratulations, guys, you deserved it! Really well done reaching this score",
      "votes": 1
    },
    {
      "id": 1912740,
      "postDate": "2022-08-25T00:28:09.160Z",
      "content": "<p>Congratulations 🎉🎉💥</p>",
      "rawMarkdown": "Congratulations 🎉🎉💥",
      "votes": 1
    },
    {
      "id": 1912732,
      "postDate": "2022-08-25T00:23:39.650Z",
      "content": "<p>Congratulations！</p>",
      "rawMarkdown": "Congratulations！",
      "votes": 1
    },
    {
      "id": 1913268,
      "postDate": "2022-08-25T09:05:23.537Z",
      "content": "<p>Congratulations 🎉🎉 and Thanks for the solution</p>",
      "rawMarkdown": "Congratulations 🎉🎉 and Thanks for the solution",
      "votes": 2,
      "replies": [
        {
          "id": 1913318,
          "postDate": "2022-08-25T09:21:55.303Z",
          "content": "<p>Solution is still in process. Was posted just initial plan and some information about model. Sorry for the delay, but will need 3-4h more to finish it. Hope you'll find it interesting.</p>",
          "rawMarkdown": "Solution is still in process. Was posted just initial plan and some information about model. Sorry for the delay, but will need 3-4h more to finish it. Hope you'll find it interesting.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1912814,
      "postDate": "2022-08-25T01:25:26.207Z",
      "content": "<p>Congratulations！🎉🎉</p>",
      "rawMarkdown": "Congratulations！🎉🎉",
      "votes": 2
    },
    {
      "id": 1914802,
      "postDate": "2022-08-26T12:41:23.430Z",
      "content": "<p>Here is a dataset with 2 submissions:</p>\n<p>single_lgbm_model_v12.0.1.csv -&gt; best single lgbm (mean over folds) -&gt; 0.80891 private / 0.80101 public<br>\n2nd_place_submission_ranked.csv -&gt; ranked 2nd place -&gt; 0.80938 private / 0.80134 public</p>\n<p>interesting that single lgbm could give 6th place on lb</p>\n<p>feel free to blend it and check results on lb<br>\n<a href=\"https://www.kaggle.com/datasets/kyakovlev/amex-submissions\" target=\"_blank\">https://www.kaggle.com/datasets/kyakovlev/amex-submissions</a></p>",
      "rawMarkdown": "Here is a dataset with 2 submissions:\n\nsingle_lgbm_model_v12.0.1.csv -> best single lgbm (mean over folds) -> 0.80891 private / 0.80101 public\n2nd_place_submission_ranked.csv -> ranked 2nd place -> 0.80938 private / 0.80134 public\n\ninteresting that single lgbm could give 6th place on lb\n\nfeel free to blend it and check results on lb\nhttps://www.kaggle.com/datasets/kyakovlev/amex-submissions"
    },
    {
      "id": 2245760,
      "postDate": "2023-05-04T15:08:19.577Z",
      "content": "<p>Is there a way to see the code ? </p>",
      "rawMarkdown": "Is there a way to see the code ? "
    },
    {
      "id": 2010286,
      "postDate": "2022-10-30T16:47:07.660Z",
      "content": "<p>thanks for sharing such insights, as a newbie, i loved the part on how you collaborated on cloud, recorded all the trained models to go fwd and back and the way you divided the data to apply ensemble techniques. lot to learn.. thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing such insights, as a newbie, i loved the part on how you collaborated on cloud, recorded all the trained models to go fwd and back and the way you divided the data to apply ensemble techniques. lot to learn.. thanks for sharing!"
    },
    {
      "id": 1946819,
      "postDate": "2022-09-20T05:51:59.280Z",
      "content": "<p>Congratulations！</p>",
      "rawMarkdown": "Congratulations！"
    },
    {
      "id": 1936425,
      "postDate": "2022-09-12T18:34:58.703Z",
      "content": "<p>can you share code, as was done for 1st place?</p>",
      "rawMarkdown": "can you share code, as was done for 1st place?"
    },
    {
      "id": 1936351,
      "postDate": "2022-09-12T17:18:45.667Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a> and team-mates. Thank you for posting such detailed approach for us to learn from.</p>",
      "rawMarkdown": "Congratulations @kyakovlev and team-mates. Thank you for posting such detailed approach for us to learn from."
    },
    {
      "id": 1925646,
      "postDate": "2022-09-04T07:04:35.363Z",
      "content": "<p>A huge congratulations! Thanks for sharing the techniques and the best practice, really learn a lot from you! In particular, I am interested in how you organize different versions of your experiments, it is very important and useful for complex research like you mentioned. </p>\n<p>Can you elaborate more on the details? much appreciated!</p>",
      "rawMarkdown": "A huge congratulations! Thanks for sharing the techniques and the best practice, really learn a lot from you! In particular, I am interested in how you organize different versions of your experiments, it is very important and useful for complex research like you mentioned. \n\nCan you elaborate more on the details? much appreciated!"
    },
    {
      "id": 1921781,
      "postDate": "2022-09-01T04:07:45.547Z",
      "content": "<p>Really great！</p>",
      "rawMarkdown": "Really great！"
    },
    {
      "id": 1918844,
      "postDate": "2022-08-29T22:47:28.043Z",
      "content": "<p>can you share code pls</p>",
      "rawMarkdown": "can you share code pls",
      "replies": [
        {
          "id": 1918856,
          "postDate": "2022-08-29T23:02:36.033Z",
          "content": "<p>Yes, we will share it for sure. Our plan is to make cleanup during the weekend (3-4 September), few days for review and share everything. Last weeks of the competition were very intense and we are taking small “break”. Sorry for the delay but this is the fastest possible timeline for us.</p>",
          "rawMarkdown": "Yes, we will share it for sure. Our plan is to make cleanup during the weekend (3-4 September), few days for review and share everything. Last weeks of the competition were very intense and we are taking small “break”. Sorry for the delay but this is the fastest possible timeline for us.",
          "votes": 7
        }
      ]
    },
    {
      "id": 1914257,
      "postDate": "2022-08-25T22:35:16.320Z",
      "content": "<p>Congratulations！Thanks for sharing!</p>",
      "rawMarkdown": "Congratulations！Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1913065,
      "postDate": "2022-08-25T05:46:17.893Z",
      "content": "<p>Thank You!<br>\nWaiting for more ))</p>",
      "rawMarkdown": "Thank You!\nWaiting for more ))"
    },
    {
      "id": 1920011,
      "postDate": "2022-08-30T19:42:37.180Z",
      "content": "<p>very useful, thanks for sharing!</p>",
      "rawMarkdown": "very useful, thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 1912746,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-25T00:32:24.170000",
      "content": "<p><a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a> Congratulations Konstantin and team. I am so happy for you. You are the best tabular data feature engineer that I know. I'm glad to see your skills achieved great success! I'm looking forward to learning about your solution!</p>",
      "votes": 13,
      "replies": [
        {
          "id": 1912768,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2022-08-25T00:52:04.023000",
          "content": "<p>I think you groom him well from IEEE competition 3 years back :)  Congratulations to you both ! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1912775,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T00:57:00.003000",
          "content": "<p>haha, he groomed me well. I was new to Kaggle then and he invited me to his team when I was like 30th place and he was 1st place! He taught me so much about tabular data and feature engineering. Many groupby aggregate tricks that I use today I learned from Konstantin. Thanks Konstantin!</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1912845,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-08-25T02:04:13.323000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> it's a huge honor to hear such words from you.</p>\n<p>Right from the start we decided to make \"internal\" team for this competition but I was so many times close to break that rule and ask you to join us - I was sure that you had fantastic and a very unique NN approach (and you had it) - congrats with solo gold with transformers. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1912964,
          "author_name": "Mohamed Eltayeb",
          "author_url": "",
          "post_date": "2022-08-25T04:28:29.393000",
          "content": "<p>Well <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I think a small notebook about groupby tricks would be your best gift for us after this competition :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1913157,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2022-08-25T07:27:17.260000",
          "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>  he already made an elaborate post about feature engineering:<br>\n<a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575\" target=\"_blank\">https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575</a><br>\nIt's probably the highest voted discussion post of all times!<br>\nGroupby is under \"aggregation\"</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1913165,
          "author_name": "Mohamed Eltayeb",
          "author_url": "",
          "post_date": "2022-08-25T07:34:41.150000",
          "content": "<p>Thanks, gonna check it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1913550,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T11:06:44.673000",
          "content": "<p>Yes, the link Yukiya provides has a list of good engineer techniques. </p>\n<p>And here is a link which explains why groupby works <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111453\" target=\"_blank\">here</a>. When we create new columns using groupby from other columns, we give our GBT new ways to classify rows that belong to different groups.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1914175,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T19:53:43.670000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1920789,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2022-08-31T12:02:27.490000",
      "content": "<p>A huge congratulations and thanks for such a detail explanation. As Chris said, the best tabular feature engineer! </p>\n<p>Some questions from my side:</p>\n<p><strong>1.</strong> Based on your experience, doing brute force feature engineering (i.e. create a ton of potential new features and then perform feature selection) gives better results than going column by column?<br>\n<strong>2.</strong> How did you choose the initial LGBM hyperparameters before starting the feature selection phase? <br>\n<strong>3.</strong> Did you retune the hyperparameters at some point of the feature selection phase (e.g. when you added/removed a significant number of features) or kept them fixed all the time?<br>\n<strong>4.</strong> Could you give some examples of what type of feature subsets you include at each step of the feature selection phase (e.g. one subset could be the mean of <code>cleaned__B__columns</code>)?<br>\n<strong>5.</strong> Your final selected features came from \"Forward feature selection\". I tried it a couple of times at work and saw that tends to overfit, do you have the same experience?<br>\n<strong>6.</strong> How did you realize that increasing the value of <code>min_data_in_leaf</code> favour CV/LB alignment?</p>\n<p>Also waiting for the code 👀</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1920865,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-08-31T13:01:08.950000",
          "content": "<p>Thank you for your kind words and very nice questions.</p>\n<p>1 | Brute force is a dead end by all meanings - it will not work on production, it will not provide any interpretability or explainability, it is not reproducible result. Column by column and distilled carefully created features are always better and preferable. Even in this competition, manual feature engineering would work better if we could have access to raw data and columns description (meaning).</p>\n<hr>\n<p>2 | Initial params always depend on data, but I start most of times with 0.7/0.7 sampling/features sampling, 2^7 leaves and 2^8 min data in leaf (all other params as default). <br>\nFirst baseline run normally gives you a lot of information how to tune it further:</p>\n<ul>\n<li>if train score is much higher than CV and holdout -&gt; try regularizations / lower subsampling / less complex model</li>\n<li>if naive score is better than lgbm or difference is minimal -&gt; bad features or/and need more complex model (2**7 is complex enough for majority of tasks)</li>\n<li>if train score is low but very aligned with CV/holdout -&gt; underfitting -&gt; more complex model with less data in leaf<br>\nThen try to \"tweak\" loss function that is more appropriate for your task.<br>\nThen try to fit goss/dart/extra_trees -&gt; extra_trees works very nice sometimes -&gt; DART never worked better before)))</li>\n</ul>\n<hr>\n<p>3 | Params tuning normally takes several iterations during model creation - for competitions it means a slight tuning every 2 weeks (after significant feature sets changes or data structure changes) and final deep tuning round 2 weeks before competition ends.</p>\n<hr>\n<p>4 | Not much to add here - you are right with your assumption:<br>\nmean + cleaned__B__columns produces:<br>\nagg_func__mean__cleaned__B__columns|__B_4__mean<br>\nagg_func__mean__cleaned__B__columns|__B_16__mean<br>\nagg_func__mean__cleaned__B__columns|__B_20__mean<br>\nagg_func__mean__cleaned__B__columns|__B_22__mean<br>\nagg_func__mean__cleaned__B__columns|__B_41__mean<br>\n…</p>\n<hr>\n<p>5 | It works really nice for the start but with more features it becomes less informative and slower. It didn't show any overfitting but we didn't trust it fully as we were training and validating on the same time period (and didn't have any aligned validation set in the “future”).</p>\n<hr>\n<p>6 | We saw that ROC AUC score on train data went up to 0.9999 and CV/Holdout had cap of 0.96xxx -&gt; many teams decided to use regularization (l1 - punish features with low importance but we would like to keep even small signals, l2 - adds regularization to features with high importance -&gt; that is good move here but we didn't see any statistically significant CV score boost). We decided to make less complex model with 2^6 leaves and increase min data in leaf to force our model generalize better and found that increasing this param to 2^9 - 2^11 works very well on CV/holdout and also worked well on LB. </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1923496,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "2022-09-02T09:16:25.547000",
          "content": "<p>Thanks for such a complete answer!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1913704,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2022-08-25T13:06:04.490000",
      "content": "<p>It's becoming a long writeup))) will need bit more time</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1914039,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-08-25T17:15:31.977000",
          "content": "<p>Take your time! Great details already</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1915111,
      "author_name": "Tonghui Li",
      "author_url": "",
      "post_date": "2022-08-26T17:11:51.047000",
      "content": "<p>Huge congrats Konstantin <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a>, this is incredible commitment your team had! I subscribed 2 months of Colab Pro+ for this comp and that was it…<br>\nCurious if you know besides <code>diff_mean</code> which of your <code>'special'</code> agg ideas worked best for you?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1915119,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-08-26T17:16:48.673000",
          "content": "<p>btw diff_mean is: </p>\n<pre><code>df[f'{col}__diff_mean] = df[col] - df.groupby(['uid'])[col].shift()\ndf[f'{col}__diff_mean] = df.groupby(['uid'])[f'{col}__diff_mean]].transform('mean')\n\ndf = df.drop_duplicates(subset=['uid'], keep='last').reset_index(drop=True)\n</code></pre>\n<p>std features worked better than others - diff with std for example</p>\n<p>normal diff with client mean was in <br>\n'diff_func': ['first','mean','std','median','min','max']</p>\n<p>and my bad - which of specials.<br>\nmonotonic functions worked really well</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1915208,
          "author_name": "Tonghui Li",
          "author_url": "",
          "post_date": "2022-08-26T18:36:20.317000",
          "content": "<p>Thanks for the reply! <br>\n<code>monotonic functions</code> - Did you create a T/F binary feature or did you calculate for example spearman correlation or similar to quantify monotonicity?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1915211,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-08-26T18:38:15.713000",
          "content": "<p>Binary if current more or equal to previous and then took mean.</p>\n<p>We wanted to pass to the model information if last statements increase or decrease is normal for the client or not. If due payments are constantly increasing for us it was a sign that risk is growing.</p>\n<p>And such information is not based on absolute values - we constantly wanted to avoid absolute values to be able to generalize well on future data and data with drifts.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1912725,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-08-25T00:21:14.897000",
      "content": "<p>Congrats guys .. cant wait for your solution to reach 0.809 .. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1918704,
      "author_name": "Ahmed Taha",
      "author_url": "",
      "post_date": "2022-08-29T18:58:54.670000",
      "content": "<p>very insightful tank youu</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917682,
      "author_name": "Runzhong Xu",
      "author_url": "",
      "post_date": "2022-08-29T01:25:08.577000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917052,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2022-08-28T11:47:50.033000",
      "content": "<p>Here is example how we augmented training Data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F7db49977cfebf34a15d87e5b4c95ec3f%2F2022-08-28%20%2012.41.40.png?generation=1661686959865932&amp;alt=media\" alt=\"\"></p>\n<p>We were removing last statement per client as it would never exists and then applied exactly the same transformations as we did for full history.</p>\n<ul>\n<li>Last values are different but within distributions </li>\n<li>Agg values are different but within distributions</li>\n<li>Model has temporal component </li>\n</ul>\n<p>We wanted to completely avoid absolute values and with data augmentation the CV and LB were almost on the same level as with absolute values - of course it would not work with clients with 1-2 statements but we had a separate model for them. I believe that that would give much more robust results but on kaggle even slight metric increase is important and we kept \"lasts\".</p>\n<p>Also, we were planning to go in -2 max depth and used only 1,2,3,6,11 (most common time-series lags) diff lags to have fe aligned. </p>\n<p>Combining original data with -1 statement or with -2 statement gave significant CV boost on ROC AUC score (our main metric) and tiny boost on AMEX metric. Combination of 3 sets gave just a tiny boost on ROC AUC but was significantly more memory consuming and slower to train model.</p>\n<p>Also, CV splits per original client uid to exclude leaks and validation was performed only on original data to exclude metric degradation tracking for shorten history predictions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917023,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2022-08-28T11:13:19.733000",
      "content": "<p>Few examples how noise cleaning works:<br>\nFor the most features nunique groups within 998 -&gt; 1001 is a sign of noised cleaning possibility.</p>\n<p>Simple feature cleaning - in this case group is our new value:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2Fea9cd72dbd9666335a1194331479fbc8%2F2022-08-28%20%2012.05.28.png?generation=1661684893433642&amp;alt=media\" alt=\"\"></p>\n<p>More complex feature with negative \"categories\" we keep negative values (temp2 is a new value):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F652007f94ba3b47bc300dfe32da85ce1%2F2022-08-28%20%2012.06.17.png?generation=1661684841167050&amp;alt=media\" alt=\"\"></p>\n<p>Very \"noised\" feature (we find initial coefficient and then apply it to groups with low population) - temp2 is a new value:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F2d811e27de9544d6eaeca47cf268b650%2F2022-08-28%20%2012.11.55.png?generation=1661685137975362&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1915380,
      "author_name": "Oscar Aguilar",
      "author_url": "",
      "post_date": "2022-08-26T23:54:33.063000",
      "content": "<p>Thanks for sharing your solution! Very insightful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913980,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-08-25T16:05:05.120000",
      "content": "<p>Normalizations. Sorry, I'm having trouble understanding. For all EXCEPT last month you took mean and std for all customers fort that exact col, month and subset (train/pub/pri)? What about last? I'm actually super confused what the trick was, and why? Maybe try just restating/rephrasing it and maybe I'll figure it or at least have more precise questions…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1920719,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-31T11:00:00",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913512,
      "author_name": "Alexander Guldbrand",
      "author_url": "",
      "post_date": "2022-08-25T10:49:07.533000",
      "content": "<p>Congratulation! Looking forward to dive into your solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913260,
      "author_name": "Mert Burabak",
      "author_url": "",
      "post_date": "2022-08-25T09:03:27.457000",
      "content": "<p>Good Job! well deserved</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913210,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2022-08-25T08:18:53.740000",
      "content": "<p>Congrats to the team. One more waiting for the extended explanation</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913202,
      "author_name": "Vadim Imaev",
      "author_url": "",
      "post_date": "2022-08-25T08:11:16.213000",
      "content": "<p>Congratulations ! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912995,
      "author_name": "iliiiiiili",
      "author_url": "",
      "post_date": "2022-08-25T04:56:41.570000",
      "content": "<p>Congratuation!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912963,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2022-08-25T04:28:18.733000",
      "content": "<p>Glad to see you back on Kaggle competitions with a bang. Congratulations, amazing finish. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912935,
      "author_name": "Daisy",
      "author_url": "",
      "post_date": "2022-08-25T03:57:12.200000",
      "content": "<p>Congratulations! Thank you for sharing model params!! Very useful for next competitions!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912867,
      "author_name": "shahilpravind",
      "author_url": "",
      "post_date": "2022-08-25T02:34:02.830000",
      "content": "<p>Congratulations! 🎉  Not a lot of experience on Kaggle so not sure if this is normal but was surprised to see that 3/4 of your team is novices and contributors. Gives me hope 😅</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1915182,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-08-26T18:14:18.977000",
          "content": "<p><a href=\"https://www.kaggle.com/shahilap96\" target=\"_blank\">@shahilap96</a> our team was formed at the very early stage of the competition and exclusively from my company DS department - many other DS colleagues wanted to participate but were not able to made such 3month commitment and decided not to join. Kaggle experience is nice to have but it’s not an obligatory condition - the most important is a “fire in the eyes” and passion to learn and try something new.</p>\n<p>Yes, my teammates are novice on kaggle but not novice in analytics and DS. </p>\n<p>One thing that I can say - all my teammates work with the same passion with regular work tasks and trying really hard to achieve goals. I feel very proud of them.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1916523,
          "author_name": "shahilpravind",
          "author_url": "",
          "post_date": "2022-08-28T00:39:14.290000",
          "content": "<p>Aha, that explains the novice part 😅 People with passion and fire in their eyes are always great to work with. The energy and drive always pushes me to do better. Congrats once again.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1912798,
      "author_name": "Kim Keonho",
      "author_url": "",
      "post_date": "2022-08-25T01:14:05.613000",
      "content": "<p>Congrats!<br>\nI'm waiting for your notebook to knwo what had been updated from IEEE competition :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912764,
      "author_name": "Sixian Chen",
      "author_url": "",
      "post_date": "2022-08-25T00:49:01.710000",
      "content": "<p>Congratulations！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912741,
      "author_name": "Leandro Destefani",
      "author_url": "",
      "post_date": "2022-08-25T00:29:36.860000",
      "content": "<p>Congratulations, guys, you deserved it! Really well done reaching this score</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912740,
      "author_name": "Neha",
      "author_url": "",
      "post_date": "2022-08-25T00:28:09.160000",
      "content": "<p>Congratulations 🎉🎉💥</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912732,
      "author_name": "cong wang",
      "author_url": "",
      "post_date": "2022-08-25T00:23:39.650000",
      "content": "<p>Congratulations！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913268,
      "author_name": "Gaju Ahmed",
      "author_url": "",
      "post_date": "2022-08-25T09:05:23.537000",
      "content": "<p>Congratulations 🎉🎉 and Thanks for the solution</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1913318,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-08-25T09:21:55.303000",
          "content": "<p>Solution is still in process. Was posted just initial plan and some information about model. Sorry for the delay, but will need 3-4h more to finish it. Hope you'll find it interesting.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1912814,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T01:25:26.207000",
      "content": "<p>Congratulations！🎉🎉</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1914802,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2022-08-26T12:41:23.430000",
      "content": "<p>Here is a dataset with 2 submissions:</p>\n<p>single_lgbm_model_v12.0.1.csv -&gt; best single lgbm (mean over folds) -&gt; 0.80891 private / 0.80101 public<br>\n2nd_place_submission_ranked.csv -&gt; ranked 2nd place -&gt; 0.80938 private / 0.80134 public</p>\n<p>interesting that single lgbm could give 6th place on lb</p>\n<p>feel free to blend it and check results on lb<br>\n<a href=\"https://www.kaggle.com/datasets/kyakovlev/amex-submissions\" target=\"_blank\">https://www.kaggle.com/datasets/kyakovlev/amex-submissions</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2245760,
      "author_name": "Mohit Batra",
      "author_url": "",
      "post_date": "2023-05-04T15:08:19.577000",
      "content": "<p>Is there a way to see the code ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2010286,
      "author_name": "Aman",
      "author_url": "",
      "post_date": "2022-10-30T16:47:07.660000",
      "content": "<p>thanks for sharing such insights, as a newbie, i loved the part on how you collaborated on cloud, recorded all the trained models to go fwd and back and the way you divided the data to apply ensemble techniques. lot to learn.. thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1946819,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-20T05:51:59.280000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1936425,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-12T18:34:58.703000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1936351,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-12T17:18:45.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1925646,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-04T07:04:35.363000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1921781,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-01T04:07:45.547000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1918844,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-29T22:47:28.043000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1918856,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-29T23:02:36.033000",
          "content": "",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1914257,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T22:35:16.320000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913065,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T05:46:17.893000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1920011,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-30T19:42:37.180000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912721": "*Note: Full code to retrain single model will be shared here in a 2 weeks.*\n\nI would like to say thank you to competition hosts and Kaggle - it was a great pleasure to participate in tabular data competition after many months and years without one. \n\nThank you to all participants (and of course winners - @daishu huge jump during last 3 days and fantastic solo result) - our success is your success - you forced us to try harder - without all of you It would be impossible to learn so many new things and achieve such result.\n\nAnd special \"thank you\" goes to my fantastic teammates:\n- Danila\n- Alexey\n- Igor\n\nNo words would be enough to say how much each of you contributed to the end result.\n\"Do Data Scientists have hobby? Yes -> DS competition\".  \n\n---\n\n### Competition announcement\n\nWe were very hyped by the new tabular data competition release (sorry for the external link:  [link](https://www.linkedin.com/posts/konstantin-yakovlev-010b46125_american-express-default-prediction-kaggle-activity-6935250015619592194-hqnM)) and immediately decided to participate. Slack notification -> rules and perspective advertisement (100% chance too lose summer holidays and all free time) -> and here we are - four members. Only one member of the team had previous experience in DS competitions participation.\n\nFew rules were established from the beginning:\n- Only free time for competition\n- No \"second\" accounts on Kaggle (even wife/friends to exclude any cheating suspicion)\n- No competition discussion outside of the team\n- We are here to learn and try our best\n\n---\n\n### Infrastructure and pipelines:\n\nEach of us had own machines / resources (GCP/AWS/local). We used Kaggle platform just for a few times. So the first thing we wanted to solve - unified machine to save all artefacts / experiments. We decided to go with AWS. I would say that it is possible to achieve the same result that we have just with Kaggle resources, but it would be bit more stressful for team management. We didn't want to spend a lot of money on AWS but sometimes (during very hot hours) RAM spikes were 500GB+ to permit simultaneous work.\n\nWe tried to use neptune.ai for ML tracking but from July it was not very effective as we entered in brute force zone. \n\n*Advise: Resources management is very critical - find bottleneck and remove it to make your team most effective. At the same time don't burn money recklessly - limit your budged. If any optimization possible - do it as soon as possible to save time and resources.*\n\n---\n\n### Project Structure\nEach run was internally versioned (ex. v1.1.1 - (major version).(fe version).(model version))\nOverall project structure:\n- Initial preprocess -> artifact cleaned and joined df\n- FE -> Many aligned (by uid) dfs with separated features \n- Features selection -> dictionary we selection metadata\n- Holdout Model (fe check and tuning) -> Local validation oof preds / holdout preds/ model / model metadata\n- Full model run -> Model / Model metadata\n- Prediction -> each fold oof predictions / cv split metadata / test predictions\n\nAll these permitted us to go back and forward and check what worked well and what did not and restore experiments in each particular step.\n\n---\n\n### Initial preprocess \nWe wanted to achieve several things with this step:\n- Join Train and Test -> due to many people involved I was afraid that some missed transformation on private test part will be unnoticed. So we sacrifice memory and speed optimization for overall stability and security.\n- Remove detected noise -> (we had options here but ended with unified single one)\n- Transform Customer ID to unified uid \n- Create internal subset feature -> Train / Public / Private\n- Create unified kfold and holdout split -> To align all experiments\n- Separate columns by type and store them separately to minify memory use and load time\n\n#### Remove detected noise\n\nWe didn't use public notebooks for cleaning. Radar's Dataset is fantastic and it is 99% similar to our own transformations.\nWe used \"isle\" identification without any pre-build coefficients.\ndummy code is something like this:\n\n```\n    for col in process_columns: \n\n        df = temp_df[[col]].sort_values(by=[col]) \n        df = df[df[col].notna()].drop_duplicates(subset=[col]).reset_index(drop=True)\n\n        df['temp'] = np.floor(df[col] * 100000)\n        df['group'] = ((df['temp'] - df['temp'].shift()).abs() >= 100).cumsum()\n\n        i = 0\n        while True:\n            min_val = df[df['group']==i]['temp'].min()\n            if min_val>0:\n                break\n            i += 1\n\n        df['temp2'] = np.where(df['temp']>=0, \n                                np.floor(df['temp']/min_val).astype(np.int32),\n                                np.round(df['temp']/min_val).astype(np.int32))\n        \n        mapping = dict(zip(df[col],df['temp2']))\n        temp_df[col] = temp_df[col].map(mapping)\n\n        print(col, df['group'].nunique(), df[col].nunique())\n        print(df.groupby(['group'])['temp','temp2'].agg(['min','max','count','nunique']).head(40))\n```\n\n#### Create internal subset feature\nWe used last statement month to create 0/1/2 feature and store in in \"index\" df\n\n#### Create unified kfold and holdout split\nFixed random seed (of course 42) to make spits and then took 20% of customers to holdout group (to test stacking / blending / etc)\n\n#### Separate columns by type\nAfter cleaning we had several columns \"groups\".\n```\nall_files = [\n'p_columns', -> just p columns as we thought that they are very different (and P_2 is internal amex \"scoring\" model)\n'objects_radar_columns', -> order encoding (we were checking where out cleaning differs from public approaches and here was the unique place) \n'objects_columns', -> onehot encoding\n'categorical_cleaned__D__columns', -> no noise categoricals\n'categorical_binary__S__columns', -> cleaned binary\n'categorical_binary__R__columns', -> cleaned binary\n'categorical_binary__D__columns', -> cleaned binary\n'categorical_binary__B__columns', -> cleaned binary\n'categorical__D__columns', -> removed noise categoricals\n'categorical__B__columns', -> removed noise categoricals\n'cleaned__B__columns', -> removed noise continuous \n'cleaned__D__columns', -> removed noise continuous \n'cleaned__R__columns', -> removed noise continuous \n'cleaned__S__columns', -> removed noise continuous \n'rest__B__columns', -> have no idea what to do with it -> floor \n'rest__D__columns', -> have no idea what to do with it -> floor \n'rest__R__columns', -> have no idea what to do with it -> floor \n'rest__S__columns', -> have no idea what to do with it -> floor \n]\n```\nThanks again to @raddar we always used your preprocess as a baseline.\n\nWe were able to load just portion of data -> do fe -> concat to \"index\" as all dfs were aligned by index. Also such split permitted us to do fe by feature type to accelerate process and see more statistically valuable metric change. \n\n---\n\n### FE\n\nWe started with careful fe column by column or small subset and it worked well until 1xx features and then any metric improvement or degradation was not statistically significant and many features \"overlapped\" on importance and significance. \n\n*Note: I believe that it is possible to build silver zone robust model with only 3xx features*\n\nSo we started from scratch with brute force))) Of course there was no need to apply \"nunique\" (for example to binary features) and our previous step helped us to limit fe.\n\n```\nall_aggregations = {\n   'agg_func': ['last','mean','std','median','min','max','nunique'],\n   'diff_func': ['first','mean','std','median','min','max'],\n   'ratio_func': ['first','mean','std','median','min','max'],\n   'lags': [1,2,3,6,11],\n   'special': ['ewm','count_month_enc','monotonic_increase','diff_mean','major_class',\n              'normalization','top_outlier','bottom_outlier','normalization_mean','top_outlier_mean','top_outlier_mean']\n}    \n```\n+ pca (horizonal and vertical) + horizontal combinations + horizontal aggregations.\n\nagg_func -> normal aggregations by uid\ndiff_func -> diff last - xxx -> std diff worked better than any other\nratio_func -> ratio_func last/xxx\nlags -> diff last - Nx\nspecial -> some special transformations -> count_month_enc worked well for categorical / emw for continous \n\nWe ended up with about 7k features (stored file by group and by agg type for faster loading). \nNext thing was to figure out what works and what not -> this topic was the most challenging for us. \n\n#### Normalizations\n\nIt's better to call it Standardization (x - m) / s -> as we had also normalization test the name became constant \"normalization\")))\n```\ndf.groupby(['dt_month','subset'])[col].agg(['mean','std'])\n```\ndt_month -> month of the statement\nsubset -> train / public / private\nand mean and std from clients that had full statement history.\n\nWe have to have temporal shift to make it work. So we did a \"trick\" removed last statement for each client and applied exactly same transformation  for each client and merged appropriate labels. So we had 2 lines in training set for almost each client BUT validated results only on last statement during CV runs and Holdout checks. It more or less same as adding noised data but we had temporal drift and model was able to work better on unknown future data with  \"possible\" data drift.\n\n---\n\n### Features selection\n\nOoohh that was really fun. \n\nWe used gbdt boosting type during experiments as it was very aligned with dart mode but was significantly faster.\nAlso, we used ROC AUC score during our experiments as we believed that due to amex instability we can't use it for decision making  (of course we tracked log loss and amex).\n\nIn previous step we brute forced many features and now is time to clean them out.\nAll feature selection was done with 5 CV folds training + independent check on 20% holdout data.\n\n1. Zero importance -> Right from the start we were able to through away 1.5k features that had exactly 0 importance (lgbm importance). That means that with 250 bins and 2**10 data in leaf those features are not participating in any split.\n\n2. Stepped hierarchical permutation importance -> we defined 300 initial features and looped over all other features subsets (600+) - was very time consuming but very stable.\nNote: we shuffled order of the subset to force model try different combinations.\nAdd features subset -> train model -> permutate -> drop negative features (negative mean over 5 seeds) -> add new subset -> ...\nDuring this part that took almost 3 days we limited features to 3k -> 0.800 lb\n\n3. Stepped permutation importance.\nTake all features -> train model -> permutate -> drop 20% of worst performed features (only negative) -> repeat. Final subset was 25xx features (and different from previous step) -> 0.800 lb\n\n4. Forward feature selection.\nWe defined 300 initial features and simply added subset by subset and compared ROC AUC if metric change was > 0.0003 we kept the subset. -> 0.800 lb\n\n5. Time series CV.\nFor very doubtful features as PCA and Normilized values we used to different validation stratagies:\n- Train on first 6 month values (last statement of the first 6 months went to train set) and validate on last 6 (also just last statement of the last 6 months). We trained model without temporal feature and then with if result was better on CV and on holdout we added to final features subset.\n- We used P_2, B_1, B_2 as a proxy target and MSE loss with combined Train and Test to see if we did right transformation and result did not degrade.\n\nMany other options we tried but result was not stable.\n\nFinal subset came from \"Forward feature selection\" plus overlapped features from other technics minus overlapped negative combination.  -> lb 0.801 single model.\n\nWe tried to blend many models with different subset as we believed that it should give huge LB boost (based on holdout blending tests) but it didn't work well for lb. \n\n--- \n\n### Model \n\nIn my own experience, DART never worked better and here we have proof that in DS \"all depends.\" We did experiments with DART in the beginning and it did not show any metric improvement with our params and baseline model features subset. Later we found @ragnar123 notebook and gave it one more try and it worked marvelously.\n\nFrom the beginning, we tried to build a more complex model with 2**7+ leaves and 0.7+ features but failed. It still puzzles me why a simple model with a very low number of features works here. \n\nI saw such behaviour mostly on synthetic data and stacking - so we tried to find out if data is syntetic (at least partly) and deanonimize internal scoring values - but didn't make it.\n\nOur best single lgbm model was trained on 29xx features. 5 folds CV - no stratification by any option. Training data - 2 last staements for each client (transformed independently). Params:\n\n```\nlgb_params = {\n    'boosting_type': 'dart',\n    'objective': 'cross_entropy', \n    'metric': ['AUC'],\n    'subsample': 0.8,  \n    'subsample_freq': 1,\n    'learning_rate': 0.01, \n    'num_leaves': 2 ** 6, \n    'min_data_in_leaf': 2 ** 11, \n    'feature_fraction': 0.2, \n    'feature_fraction_bynode':0.3,\n    'first_metric_only': True,\n    'n_estimators': 17001,  # -> 5000 for gbdt \n    'boost_from_average': False,\n    'early_stopping_rounds': 300,\n    'verbose': -1,\n    'num_threads': -1,\n    'seed': SEED,\n}\n```\n\nBlend -> Power (2) rank blend of Dart lgbm (0.801 public) / GBDT lgbm  (0.799 public) / Catboost models (0.799 public)\n\nSingle lgbm with 3 last statements showed even better CV by we didn't have enough time to retrain it (full DART run for 5 folds took 12+ hours there).\n\nIt was obvious that clients with a little number of statements will not get benefit from all 2k features. So we created a special model that was trained only on 300 features with custom params (also dart). Predictions for clients with <=2 statements came exclusively from such model and were not blended with other models. \n\nHow did we combine the result from 2 independent models to not destroy the final ranking?\n\n| Client id | Number of statements | Basic ranking | <=2 prediction | Final ranking |\n| --- | --- | --- | --- | --- |\n| 1 | 13 | 5 | ... | 5 |\n| 1 | 2 | 4 | 0.1 | 3 |\n| 1 | 13 | 8 | ... | 8 |\n| 1 | 1 | 3 | 0.5 | 4 |\n| 1 | 13 | 7 | ... | 7 |\n| 1 | 13 | 2 | ... | 2 |\n\nwe kept ranking for >2 statements and for rest resorted within initial ranking group (hope it's clear enough))). Was is perfect - no, but it was very stable with really tiny improvement (because of number of such clients in public and private test parts).\n\nWhat also worked well:\n- Train on all data without folds splitting and stop just 1000 rounds further then CV showed. \n\nWhat didn't work:\n-  Stacking by any mean\n-  Many models with different seed and fe order\n- Massive blend of different models with different types and features (blend worked well till 0.799 and then any low performed model 0.795- made public score worse - we use public CV as additional holdout set and used approaches that worked well on local CV and LB - anything that worked partially was not used in the end).\n\nAgain, nothing really fancy here. The main thing that helped us align Train / CV with LB was very high  'min_data_in_leaf'.\nNo optuna used -> just manual old school tuning based on data feeling.\nWe did many experiments with weights and loss functions but none of them worked.\n\nDue to AMEX metric specification it was obvious that focal loss should work but it didn't. We tried several times to switch loss function during the competition period and the result was the same. \n\nError analysis showed that model makes errors without any \"pattern\" -> stacking didn't work for holdout set (25% of data)  and we had doubts that it will work on private/public test parts. We kept only LR/Lasso(0.02) for blending options to choose submissions.\n\nCross validation -> standard 5folds CV split by client ID. The unique thing that we did here is \"prespliting\" to align all CV between team members to be able to compare results directly.\n \n---\n\nWhat left without mentions:\n- EDA on data\n- Denoising experiments\n- Data deanonymization -> didn't manage to make it\n- Features pairs and triples combinations -> that didn't work well\n- NaN filling -> didn't work\n- Clusterization -> didn't work\n- Hundreds of experiments with features selection process and internal discussions about it.\n- Adding noised data (noise / swap noise) that leaded to interesting but doubtful results\n- Model tuning \n- Removing absolute values and keep only diff or ratios -> should be more stable for future data but we saw some lb degradation and didn't proceed\n- pseudo labeling\n\nWhat we always wanted but didn't found time to do:\n- NN - we have no NN in our final blend\n- P_2 or any other column prediction (1/2/3/4 months ahead) with combined data and use it as meta information for lgbm main model\n- 11 / 12 / 13 statements joined training on different subsets (df was too large and training was slow)\n\n--- \n\n### Internal initial plan\n\n```\n########################### Data preprocessing and Data evaluation\n#################################################################################\n\n## Added noise removal -> GOOD2DO\n# There is no doubt that some Noise was injected in data\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\n# We need to find a way to remove it \n# the best option to not follow public approach\n# At least with columns where columns have overlaped population\n\n## Data minification for FE -> GOOD2DO\n# Datatype downcasting\n# Pickle/Parquet/Feather \n# Be careful with floats16 as it may lead to bad agg results\n# Also float16 may lead to some signal degradation due to precision and values changes\n\n## Evaluate values distributions and NaNs -> GOOD2DO\n# Full 13 months history\n# Train against Test Public and Test Private\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\n#\n# Need to try:\n# Kolmogorov–Smirnov test -> GOOD2DO\n#\n# Adversarial validation -> GOOD2DO\n# https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation/notebook (just as simple example)\n#\n# Entropy / Distances / etc....\n#\n# Visual checks)))\n#\n# We need to find if ANY feature has very different distribution in PRIVATE test set\n# If that feature works for Public part it doesn't mean that it will work for Private\n\n########################### Targets\n#################################################################################\n\n# We need to find a way to get more targets -> GOOD2DO\n# as we currently training on a single point by client we could greatly improve results\n# by extending our training set with new targets\n#\n# Find default periods in current client history and make appropriate labeling -> GOOD2DO\n# Make 2 level model -> predict p_2 values as normal time-series model and feed it to 2nd level GBT\n\n########################### Separate Models\n#################################################################################\n\n# Probably it's a good idea to make separate models for each subdatasets\n# Full history (13 months)\n# Less than 13 months\n\n########################### External Data\n#################################################################################\n# We can try to add \"Consumer index\" or any other independent temporal feature\n# Will not work if we will not be able to expand targets and add temporal feature\n\n########################### FE\n#################################################################################\n\n# We didn't make anything special here -> July\n# AGGS (Stats by client)\n# Rollings\n# History length feature (not sure if it will help with Private Test)\n# Should we correct statements dates and add NaNs?\n# ReRanking categorical features by P_2 or Target\n# Clusterization (4+ groups feature by feature)\n# Count and Mean encodings for categorical features\n# Features combinations (sum/prod/power) -> bruteforce\n# PCA or any other dimension reduction by features groups\n\n# We need to find if there is \"connection\" between clients in Train -> Public Test -> Private Test\n# we have 458913 + 924621 -> 1383534 If I were AMEX I would export 1M clients (or other round number)\n# so may be 384 534 Clients are overlaps\n\n# Clip by 5 - 95 percentile\n\n########################### Features Selection\n#################################################################################\n# Permutation importance (use all fold only!!!) -> recursive elimination (because of quantity of features -> 3-4 rounds with 0 and 50% negative drop) \n# SHAP\n# Highly correlated features (.98+?)\n# Forward selection (may take ages and due aggs may be not effective - probably by feature block) \n# Backward elimination (may take ages and due aggs may be not effective - probably by feature block) \n\n########################### CV\n#################################################################################\n# Mean Target differs my \"history length\" -> could be wise to do GroupedStratifeidFolds by history length\n# For sure Splits should be done by client\n# Target stratification to balance folds\n\n########################### Loss function / Metric\n#################################################################################\n# Clean and fast np/torch metric\n# Now it's in helper (need to cleanup that)\n# https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\n#\n# I don't believe that we will have better results with different loss function\n# But it worth to try at least focal loss\n# https://maxhalford.github.io/blog/lightgbm-focal-loss/\n#\n# Weights -> we should try change weights there\n# weights by class\n# weights by some history length\n# weights by internal fe group\n#\n# We need custom metric for catboost\n# example https://catboost.ai/en/docs/concepts/python-usages-examples#logloss1\n\n########################### Models\n#################################################################################\n\n## First level choice\n# LGB/XGB/CTB -> our main models here for sure\n# After stabilizing the baseline model and base feature we need to make 1st round tuning\n\n## Catboos specials\n# Categorical features\n# Embeding features\n\n## NN (GPU/TPU) -> RNN / LSTM / Transformer\n# TPU -> tensorflow (as it works better there)\n\n## NN -> AE / VAE / DAE -> as a denoising model hidden layer as input for GBT models\n# No need complex approach - just fast check the idea and in case of success move to big model\n\n########################### Blending\n#################################################################################\n# Weighted Average\n# Power Average\n# Weighted Rank Average\n# Linear/SVM\n# Postprocessing?\n```",
    "1912746": "@kyakovlev Congratulations Konstantin and team. I am so happy for you. You are the best tabular data feature engineer that I know. I'm glad to see your skills achieved great success! I'm looking forward to learning about your solution!",
    "1920789": "A huge congratulations and thanks for such a detail explanation. As Chris said, the best tabular feature engineer! \n\nSome questions from my side:\n\n**1.** Based on your experience, doing brute force feature engineering (i.e. create a ton of potential new features and then perform feature selection) gives better results than going column by column?\n**2.** How did you choose the initial LGBM hyperparameters before starting the feature selection phase? \n**3.** Did you retune the hyperparameters at some point of the feature selection phase (e.g. when you added/removed a significant number of features) or kept them fixed all the time?\n**4.** Could you give some examples of what type of feature subsets you include at each step of the feature selection phase (e.g. one subset could be the mean of `cleaned__B__columns`)?\n**5.** Your final selected features came from \"Forward feature selection\". I tried it a couple of times at work and saw that tends to overfit, do you have the same experience?\n**6.** How did you realize that increasing the value of `min_data_in_leaf` favour CV/LB alignment?\n\nAlso waiting for the code 👀",
    "1913704": "It's becoming a long writeup))) will need bit more time",
    "1915111": "Huge congrats Konstantin @kyakovlev, this is incredible commitment your team had! I subscribed 2 months of Colab Pro+ for this comp and that was it...\nCurious if you know besides `diff_mean` which of your `'special'` agg ideas worked best for you?",
    "1912725": "Congrats guys .. cant wait for your solution to reach 0.809 .. ",
    "1918704": "very insightful tank youu",
    "1917682": "Congratulations!",
    "1917052": "Here is example how we augmented training Data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F7db49977cfebf34a15d87e5b4c95ec3f%2F2022-08-28%20%2012.41.40.png?generation=1661686959865932&alt=media)\n\nWe were removing last statement per client as it would never exists and then applied exactly the same transformations as we did for full history.\n\n- Last values are different but within distributions \n- Agg values are different but within distributions\n- Model has temporal component \n\nWe wanted to completely avoid absolute values and with data augmentation the CV and LB were almost on the same level as with absolute values - of course it would not work with clients with 1-2 statements but we had a separate model for them. I believe that that would give much more robust results but on kaggle even slight metric increase is important and we kept \"lasts\".\n\nAlso, we were planning to go in -2 max depth and used only 1,2,3,6,11 (most common time-series lags) diff lags to have fe aligned. \n\nCombining original data with -1 statement or with -2 statement gave significant CV boost on ROC AUC score (our main metric) and tiny boost on AMEX metric. Combination of 3 sets gave just a tiny boost on ROC AUC but was significantly more memory consuming and slower to train model.\n\nAlso, CV splits per original client uid to exclude leaks and validation was performed only on original data to exclude metric degradation tracking for shorten history predictions.",
    "1917023": "Few examples how noise cleaning works:\nFor the most features nunique groups within 998 -> 1001 is a sign of noised cleaning possibility.\n\nSimple feature cleaning - in this case group is our new value:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2Fea9cd72dbd9666335a1194331479fbc8%2F2022-08-28%20%2012.05.28.png?generation=1661684893433642&alt=media)\n\nMore complex feature with negative \"categories\" we keep negative values (temp2 is a new value):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F652007f94ba3b47bc300dfe32da85ce1%2F2022-08-28%20%2012.06.17.png?generation=1661684841167050&alt=media)\n\nVery \"noised\" feature (we find initial coefficient and then apply it to groups with low population) - temp2 is a new value:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2405813%2F2d811e27de9544d6eaeca47cf268b650%2F2022-08-28%20%2012.11.55.png?generation=1661685137975362&alt=media)",
    "1915380": "Thanks for sharing your solution! Very insightful.",
    "1913980": "Normalizations. Sorry, I'm having trouble understanding. For all EXCEPT last month you took mean and std for all customers fort that exact col, month and subset (train/pub/pri)? What about last? I'm actually super confused what the trick was, and why? Maybe try just restating/rephrasing it and maybe I'll figure it or at least have more precise questions...",
    "1913512": "Congratulation! Looking forward to dive into your solution!",
    "1913260": "Good Job! well deserved",
    "1913210": "Congrats to the team. One more waiting for the extended explanation",
    "1913202": "Congratulations ! :)",
    "1912995": "Congratuation!",
    "1912963": "Glad to see you back on Kaggle competitions with a bang. Congratulations, amazing finish. ",
    "1912935": "Congratulations! Thank you for sharing model params!! Very useful for next competitions!",
    "1912867": "Congratulations! 🎉  Not a lot of experience on Kaggle so not sure if this is normal but was surprised to see that 3/4 of your team is novices and contributors. Gives me hope 😅",
    "1912798": "Congrats!\nI'm waiting for your notebook to knwo what had been updated from IEEE competition :)",
    "1912764": "Congratulations！",
    "1912741": "Congratulations, guys, you deserved it! Really well done reaching this score",
    "1912740": "Congratulations 🎉🎉💥",
    "1912732": "Congratulations！",
    "1913268": "Congratulations 🎉🎉 and Thanks for the solution",
    "1912814": "Congratulations！🎉🎉",
    "1914802": "Here is a dataset with 2 submissions:\n\nsingle_lgbm_model_v12.0.1.csv -> best single lgbm (mean over folds) -> 0.80891 private / 0.80101 public\n2nd_place_submission_ranked.csv -> ranked 2nd place -> 0.80938 private / 0.80134 public\n\ninteresting that single lgbm could give 6th place on lb\n\nfeel free to blend it and check results on lb\nhttps://www.kaggle.com/datasets/kyakovlev/amex-submissions",
    "2245760": "Is there a way to see the code ? ",
    "2010286": "thanks for sharing such insights, as a newbie, i loved the part on how you collaborated on cloud, recorded all the trained models to go fwd and back and the way you divided the data to apply ensemble techniques. lot to learn.. thanks for sharing!",
    "1946819": "Congratulations！",
    "1936425": "can you share code, as was done for 1st place?",
    "1936351": "Congratulations @kyakovlev and team-mates. Thank you for posting such detailed approach for us to learn from.",
    "1925646": "A huge congratulations! Thanks for sharing the techniques and the best practice, really learn a lot from you! In particular, I am interested in how you organize different versions of your experiments, it is very important and useful for complex research like you mentioned. \n\nCan you elaborate more on the details? much appreciated!",
    "1921781": "Really great！",
    "1918844": "can you share code pls",
    "1914257": "Congratulations！Thanks for sharing!",
    "1913065": "Thank You!\nWaiting for more ))",
    "1920011": "very useful, thanks for sharing!"
  }
}