{
  "id": 348014,
  "title": "13th Place Gold Solution",
  "url": "/competitions/amex-default-prediction/writeups/ci-2-k-uncle-giba-13th-place-gold-solution",
  "author_name": "",
  "post_date": "2022-09-06T15:49:29.303Z",
  "votes": 76,
  "comment_count": 20,
  "views": 0,
  "content": "<p>First of all, we would like to thank kaggle and the staff for hosting such a great competition. We really miss interesting tabular challenges these days.</p>\n<h1>1. Summary</h1>\n<p>Our solution is based on extensive data cleaning and multi-model weighted average and stacked ensembles using ranked probs. Using extensive data cleaning, our single model was boosted to cv 0.0004~0.0008 compared to the public Raddar's dataset.The final solution (PVT 0.80842/PUB 0.80105) which shaked up to gold zone  is the average of 3 ensemble models:</p>\n<ul>\n<li>LGBM stack with 61 models</li>\n<li>CMA(Covariance Matrix Adaptation) weighted average with 54 models</li>\n<li>Weighted average with 55 models</li>\n</ul>\n<h1>2. Extensive data cleaning</h1>\n<p>Starting from Raddar’s (thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>) preprocessed dataset, some other features were cleaned, added and modified in order to build a second version of the dataset. The idea here was to clean the data even more and add some diversity to the ensemble. Some clusters of clients were added based on the pattern of missing variables for each set of features: B, D, P, R and S. Also, some continuous variables showed a curious noise pattern. Looks like a uniform noise was added in the range of (0 - 0.01] of some variables. Take a look at feature B_11 distribution. White noise was clearly added just in the 0.01 range.</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/AMEX/main/giba2.jpg\" alt=\"\"></p>\n<p>Spent some time trying to figure out a way to remove that noise and found some combinations of filters using other features that worked pretty well. For example, the B_11 feature can be cleaned by using a filter in B_1. Taking indices when B_1 is in range of 0-0.01 and inverting the signal of B_11 based on these indices, the new histogram becomes:</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/AMEX/main/giba3.jpg\" alt=\"\"></p>\n<p>This kind of cleaning was made with some other features like: B_1, B_5, B_7, B_11, B_15, B_17, B_18, B_21, B_23, B_24, B_26, B_27, B_29, B_36, B_37, D_58, D_60, D_69, D_71, D_102, D_133, D_144, R_1, R_6, S_16, S_17, S_19, S_22 and S_27. Extra feature cleaning helped to boost GBDT scores and added some diversity when stacking with other models.</p>\n<h2>2.1 Example of boosting cv using cleaning data</h2>\n<ul>\n<li><p>LGBM : 0.7976 → 0.7983 ( +0.0007 )</p></li>\n<li><p>XGB : 0.7978 → 0.7986 ( + 0.0008 )</p></li>\n<li><p>CatBoost :    0.7964 → 0.7968 ( + 0.0004 ) </p>\n<p>※ clean data plus minor change </p>\n<ul>\n<li>LGBM 5kfold → 15kfold (+0.0013)</li>\n<li>CatBoost longer earlystop (+0.002 )</li></ul></li>\n</ul>\n<h1>3. Single model</h1>\n<h2>3.1 Features and modeling</h2>\n<p>Basically, we used the features and the models of public notebooks.<br>\nThank you for <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a></p>\n<p>Note: some of the 1st level models don’t include features (B_29, R_1, D_59) from adversarial validation.</p>\n<h2>3.2 Representative best single model in each model type</h2>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>features</th>\n<th>kfold</th>\n<th>cv</th>\n<th>public lb</th>\n<th>private lb</th>\n<th>comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM</td>\n<td>1784</td>\n<td>15</td>\n<td>0.7991</td>\n<td>0.79961</td>\n<td>0.80679</td>\n<td>dart with early stop</td>\n</tr>\n<tr>\n<td>XGB</td>\n<td>3413</td>\n<td>5</td>\n<td>0.7986</td>\n<td>0.79873</td>\n<td>0.80662</td>\n<td>pyramid <a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-cv-0-7968\" target=\"_blank\">REF</a></td>\n</tr>\n<tr>\n<td>CATBoost</td>\n<td>2290</td>\n<td>5</td>\n<td>0.7968</td>\n<td>0.79832</td>\n<td>0.80629</td>\n<td></td>\n</tr>\n<tr>\n<td>MLP</td>\n<td>446</td>\n<td>5</td>\n<td>0.7919</td>\n<td>0.79102</td>\n<td>0.79993</td>\n<td>KD using LGBM oof prediction</td>\n</tr>\n<tr>\n<td>GRU</td>\n<td>188</td>\n<td>5</td>\n<td>0.7948</td>\n<td>0.79418</td>\n<td>0.80363</td>\n<td>KD using LGBM oof prediction</td>\n</tr>\n<tr>\n<td>Transformer</td>\n<td>188</td>\n<td>5</td>\n<td>0.7932</td>\n<td>0.79498</td>\n<td>0.80379</td>\n<td>KD using LGBM oof prediction</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Ensemble</h1>\n<p>We used the two methods for ensemble with ranked probs. One is the LGBM stacking, the other is the CMA (Covariance Matrix Adaptation) Evolution Strategy<br>\n<a href=\"https://www.scm.com/doc/params/python/optimizers/cmaes.html\" target=\"_blank\">REF</a></p>\n<p>And final submission is the average of the following 3 ensemble models(case1～3) ::</p>\n<table>\n<thead>\n<tr>\n<th>Case-Method</th>\n<th>CV</th>\n<th>public</th>\n<th>private</th>\n<th>#GBDT</th>\n<th>#1/2dcnn</th>\n<th>#MLP</th>\n<th>#TCN</th>\n<th>#GRU</th>\n<th>#Transf</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1-LGBM stacking</td>\n<td>0.8010</td>\n<td>0.80084</td>\n<td>0.80882</td>\n<td>48</td>\n<td>2</td>\n<td>5</td>\n<td>2</td>\n<td>3</td>\n<td>1</td>\n</tr>\n<tr>\n<td>2- cma</td>\n<td>0.8032</td>\n<td>0.80086</td>\n<td>0.80794</td>\n<td>41</td>\n<td>2</td>\n<td>5</td>\n<td>2</td>\n<td>3</td>\n<td>1</td>\n</tr>\n<tr>\n<td>3- cma</td>\n<td>0.8028</td>\n<td>0.80068</td>\n<td>0.8079</td>\n<td>49</td>\n<td>0</td>\n<td>3</td>\n<td>0</td>\n<td>3</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>As you can see, the public lb is almost the same, but the private lb shows much better LGBM stacking. We would have been hard pressed to get the gold medal without LGBM stacking.</p>\n<h2>4.1. Final sub (ensemble results of case 1 ~ 3)</h2>\n<p>The ensemble result of the above case 1-3 is our final submission.<br>\npublic lb 0.80105, private lb 0.80842 (13th).</p>\n<p>※ Note that : LGBM stacking alone (0.80882) has the capability of ranking 6th.</p>\n<h1>5. Relationship of cv and lb</h1>\n<p>As for reference, we share the relationship of cv and lb. It is easy to see without looking at the plot with cma, but the lgbm stacking is a clean extension of the cv/lb straight line of the single model to some extend (Maybe cma is overfitting…)</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/AMEX/main/AMEX_cv_lb.jpg\" alt=\"\"></p>\n<h1>6. Other tips that worked well</h1>\n<ul>\n<li>knowledge distillation</li>\n<li>longer early stop (3000～10000)</li>\n<li>large kfold</li>\n<li>full training</li>\n<li>pseudo labeling</li>\n</ul>\n<h1>7. Didn't work well</h1>\n<ul>\n<li>post process</li>\n<li>using dow average data</li>\n<li>changing weights of samples in LGBM</li>\n<li>adding focal loss to LGBM</li>\n<li>optimizing 0.4G+0.6D, 0.8G+0.2D, etc</li>\n</ul>\n<h1>8. Late submission result</h1>\n<p>After the end of competition we realized that we had 2 ensembles with PVT LB &gt; 0.80920 (3-4th place) - however, were not considered due to lower CV scores.</p>\n<p>1) avg rank of 3 stack models (XGB+LGB+Catboost) - PVT 0.80928, PUB 0.80067, CV 0.8008</p>\n<p>2) LGB stack with 61 models + meta features (top 50 eng. features) - PVT 0.80926, PUB 0.80072, CV 0.80097</p>\n<h1>9. Team organization</h1>\n<p>We used github for code storage, wandb for experiments tracking and kaggle datasets for OOF storage and sharing with the team.</p>\n<p>Team members: <br>\n<a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">https://www.kaggle.com/titericz</a><br>\n<a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">https://www.kaggle.com/yamsam</a><br>\n<a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">https://www.kaggle.com/imeintanis</a><br>\n<a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">https://www.kaggle.com/chumajin</a><br>\n<a href=\"https://www.kaggle.com/kurokurob\" target=\"_blank\">https://www.kaggle.com/kurokurob</a></p>",
  "messages": [
    {
      "id": "1914748",
      "postDate": "08/26/2022 11:35:01",
      "content": "<p>First of all, we would like to thank kaggle and the staff for hosting such a great competition. We really miss interesting tabular challenges these days.</p>\n<h1>1. Summary</h1>\n<p>Our solution is based on extensive data cleaning and multi-model weighted average and stacked ensembles using ranked probs. Using extensive data cleaning, our single model was boosted to cv 0.0004~0.0008 compared to the public Raddar's dataset.The final solution (PVT 0.80842/PUB 0.80105) which shaked up to gold zone  is the average of 3 ensemble models:</p>\n<ul>\n<li>LGBM stack with 61 models</li>\n<li>CMA(Covariance Matrix Adaptation) weighted average with 54 models</li>\n<li>Weighted average with 55 models</li>\n</ul>\n<h1>2. Extensive data cleaning</h1>\n<p>Starting from Raddar’s (thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>) preprocessed dataset, some other features were cleaned, added and modified in order to build a second version of the dataset. The idea here was to clean the data even more and add some diversity to the ensemble. Some clusters of clients were added based on the pattern of missing variables for each set of features: B, D, P, R and S. Also, some continuous variables showed a curious noise pattern. Looks like a uniform noise was added in the range of (0 - 0.01] of some variables. Take a look at feature B_11 distribution. White noise was clearly added just in the 0.01 range.</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/AMEX/main/giba2.jpg\" alt=\"\"></p>\n<p>Spent some time trying to figure out a way to remove that noise and found some combinations of filters using other features that worked pretty well. For example, the B_11 feature can be cleaned by using a filter in B_1. Taking indices when B_1 is in range of 0-0.01 and inverting the signal of B_11 based on these indices, the new histogram becomes:</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/AMEX/main/giba3.jpg\" alt=\"\"></p>\n<p>This kind of cleaning was made with some other features like: B_1, B_5, B_7, B_11, B_15, B_17, B_18, B_21, B_23, B_24, B_26, B_27, B_29, B_36, B_37, D_58, D_60, D_69, D_71, D_102, D_133, D_144, R_1, R_6, S_16, S_17, S_19, S_22 and S_27. Extra feature cleaning helped to boost GBDT scores and added some diversity when stacking with other models.</p>\n<h2>2.1 Example of boosting cv using cleaning data</h2>\n<ul>\n<li><p>LGBM : 0.7976 → 0.7983 ( +0.0007 )</p></li>\n<li><p>XGB : 0.7978 → 0.7986 ( + 0.0008 )</p></li>\n<li><p>CatBoost :    0.7964 → 0.7968 ( + 0.0004 ) </p>\n<p>※ clean data plus minor change </p>\n<ul>\n<li>LGBM 5kfold → 15kfold (+0.0013)</li>\n<li>CatBoost longer earlystop (+0.002 )</li></ul></li>\n</ul>\n<h1>3. Single model</h1>\n<h2>3.1 Features and modeling</h2>\n<p>Basically, we used the features and the models of public notebooks.<br>\nThank you for <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a></p>\n<p>Note: some of the 1st level models don’t include features (B_29, R_1, D_59) from adversarial validation.</p>\n<h2>3.2 Representative best single model in each model type</h2>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>features</th>\n<th>kfold</th>\n<th>cv</th>\n<th>public lb</th>\n<th>private lb</th>\n<th>comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM</td>\n<td>1784</td>\n<td>15</td>\n<td>0.7991</td>\n<td>0.79961</td>\n<td>0.80679</td>\n<td>dart with early stop</td>\n</tr>\n<tr>\n<td>XGB</td>\n<td>3413</td>\n<td>5</td>\n<td>0.7986</td>\n<td>0.79873</td>\n<td>0.80662</td>\n<td>pyramid <a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-cv-0-7968\" target=\"_blank\">REF</a></td>\n</tr>\n<tr>\n<td>CATBoost</td>\n<td>2290</td>\n<td>5</td>\n<td>0.7968</td>\n<td>0.79832</td>\n<td>0.80629</td>\n<td></td>\n</tr>\n<tr>\n<td>MLP</td>\n<td>446</td>\n<td>5</td>\n<td>0.7919</td>\n<td>0.79102</td>\n<td>0.79993</td>\n<td>KD using LGBM oof prediction</td>\n</tr>\n<tr>\n<td>GRU</td>\n<td>188</td>\n<td>5</td>\n<td>0.7948</td>\n<td>0.79418</td>\n<td>0.80363</td>\n<td>KD using LGBM oof prediction</td>\n</tr>\n<tr>\n<td>Transformer</td>\n<td>188</td>\n<td>5</td>\n<td>0.7932</td>\n<td>0.79498</td>\n<td>0.80379</td>\n<td>KD using LGBM oof prediction</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Ensemble</h1>\n<p>We used the two methods for ensemble with ranked probs. One is the LGBM stacking, the other is the CMA (Covariance Matrix Adaptation) Evolution Strategy<br>\n<a href=\"https://www.scm.com/doc/params/python/optimizers/cmaes.html\" target=\"_blank\">REF</a></p>\n<p>And final submission is the average of the following 3 ensemble models(case1～3) ::</p>\n<table>\n<thead>\n<tr>\n<th>Case-Method</th>\n<th>CV</th>\n<th>public</th>\n<th>private</th>\n<th>#GBDT</th>\n<th>#1/2dcnn</th>\n<th>#MLP</th>\n<th>#TCN</th>\n<th>#GRU</th>\n<th>#Transf</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1-LGBM stacking</td>\n<td>0.8010</td>\n<td>0.80084</td>\n<td>0.80882</td>\n<td>48</td>\n<td>2</td>\n<td>5</td>\n<td>2</td>\n<td>3</td>\n<td>1</td>\n</tr>\n<tr>\n<td>2- cma</td>\n<td>0.8032</td>\n<td>0.80086</td>\n<td>0.80794</td>\n<td>41</td>\n<td>2</td>\n<td>5</td>\n<td>2</td>\n<td>3</td>\n<td>1</td>\n</tr>\n<tr>\n<td>3- cma</td>\n<td>0.8028</td>\n<td>0.80068</td>\n<td>0.8079</td>\n<td>49</td>\n<td>0</td>\n<td>3</td>\n<td>0</td>\n<td>3</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>As you can see, the public lb is almost the same, but the private lb shows much better LGBM stacking. We would have been hard pressed to get the gold medal without LGBM stacking.</p>\n<h2>4.1. Final sub (ensemble results of case 1 ~ 3)</h2>\n<p>The ensemble result of the above case 1-3 is our final submission.<br>\npublic lb 0.80105, private lb 0.80842 (13th).</p>\n<p>※ Note that : LGBM stacking alone (0.80882) has the capability of ranking 6th.</p>\n<h1>5. Relationship of cv and lb</h1>\n<p>As for reference, we share the relationship of cv and lb. It is easy to see without looking at the plot with cma, but the lgbm stacking is a clean extension of the cv/lb straight line of the single model to some extend (Maybe cma is overfitting…)</p>\n<p><img src=\"https://raw.githubusercontent.com/chumajin/AMEX/main/AMEX_cv_lb.jpg\" alt=\"\"></p>\n<h1>6. Other tips that worked well</h1>\n<ul>\n<li>knowledge distillation</li>\n<li>longer early stop (3000～10000)</li>\n<li>large kfold</li>\n<li>full training</li>\n<li>pseudo labeling</li>\n</ul>\n<h1>7. Didn't work well</h1>\n<ul>\n<li>post process</li>\n<li>using dow average data</li>\n<li>changing weights of samples in LGBM</li>\n<li>adding focal loss to LGBM</li>\n<li>optimizing 0.4G+0.6D, 0.8G+0.2D, etc</li>\n</ul>\n<h1>8. Late submission result</h1>\n<p>After the end of competition we realized that we had 2 ensembles with PVT LB &gt; 0.80920 (3-4th place) - however, were not considered due to lower CV scores.</p>\n<p>1) avg rank of 3 stack models (XGB+LGB+Catboost) - PVT 0.80928, PUB 0.80067, CV 0.8008</p>\n<p>2) LGB stack with 61 models + meta features (top 50 eng. features) - PVT 0.80926, PUB 0.80072, CV 0.80097</p>\n<h1>9. Team organization</h1>\n<p>We used github for code storage, wandb for experiments tracking and kaggle datasets for OOF storage and sharing with the team.</p>\n<p>Team members: <br>\n<a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">https://www.kaggle.com/titericz</a><br>\n<a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">https://www.kaggle.com/yamsam</a><br>\n<a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">https://www.kaggle.com/imeintanis</a><br>\n<a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">https://www.kaggle.com/chumajin</a><br>\n<a href=\"https://www.kaggle.com/kurokurob\" target=\"_blank\">https://www.kaggle.com/kurokurob</a></p>",
      "rawMarkdown": "First of all, we would like to thank kaggle and the staff for hosting such a great competition. We really miss interesting tabular challenges these days.\n\n# 1. Summary \n Our solution is based on extensive data cleaning and multi-model weighted average and stacked ensembles using ranked probs. Using extensive data cleaning, our single model was boosted to cv 0.0004~0.0008 compared to the public Raddar's dataset.The final solution (PVT 0.80842/PUB 0.80105) which shaked up to gold zone  is the average of 3 ensemble models:\n- LGBM stack with 61 models\n- CMA(Covariance Matrix Adaptation) weighted average with 54 models\n- Weighted average with 55 models\n\n# 2. Extensive data cleaning \nStarting from Raddar’s (thanks @raddar) preprocessed dataset, some other features were cleaned, added and modified in order to build a second version of the dataset. The idea here was to clean the data even more and add some diversity to the ensemble. Some clusters of clients were added based on the pattern of missing variables for each set of features: B, D, P, R and S. Also, some continuous variables showed a curious noise pattern. Looks like a uniform noise was added in the range of (0 - 0.01] of some variables. Take a look at feature B_11 distribution. White noise was clearly added just in the 0.01 range.\n\n\n![](https://raw.githubusercontent.com/chumajin/AMEX/main/giba2.jpg)\n\n\nSpent some time trying to figure out a way to remove that noise and found some combinations of filters using other features that worked pretty well. For example, the B_11 feature can be cleaned by using a filter in B_1. Taking indices when B_1 is in range of 0-0.01 and inverting the signal of B_11 based on these indices, the new histogram becomes:\n\n![](https://raw.githubusercontent.com/chumajin/AMEX/main/giba3.jpg)\n\nThis kind of cleaning was made with some other features like: B_1, B_5, B_7, B_11, B_15, B_17, B_18, B_21, B_23, B_24, B_26, B_27, B_29, B_36, B_37, D_58, D_60, D_69, D_71, D_102, D_133, D_144, R_1, R_6, S_16, S_17, S_19, S_22 and S_27. Extra feature cleaning helped to boost GBDT scores and added some diversity when stacking with other models.\n\n\n\n\n\n## 2.1 Example of boosting cv using cleaning data\n- LGBM : 0.7976 → 0.7983 ( +0.0007 )\n- XGB : 0.7978 → 0.7986 ( + 0.0008 )\n- CatBoost :    0.7964 → 0.7968 ( + 0.0004 ) \n\n ※ clean data plus minor change \n   - LGBM 5kfold → 15kfold (+0.0013)\n   - CatBoost longer earlystop (+0.002 )\n\n\n# 3. Single model \n\n## 3.1 Features and modeling\n\nBasically, we used the features and the models of public notebooks.\nThank you for @thedevastator, @ragnar123, @ambrosm, @cdeotte, @roberthatch\n\nNote: some of the 1st level models don’t include features (B_29, R_1, D_59) from adversarial validation.\n\n## 3.2 Representative best single model in each model type \n\n| model       | features | kfold | cv     | public lb | private lb | comment                      |\n|-------------|----------|-------|--------|-----------|------------|------------------------------|\n| LGBM        | 1784     | 15    | 0.7991 | 0.79961   | 0.80679    | dart with early stop         |\n| XGB         | 3413     | 5     | 0.7986 | 0.79873   | 0.80662    | pyramid [REF](https://www.kaggle.com/code/roberthatch/xgboost-pyramid-cv-0-7968)                      |\n| CATBoost    | 2290     | 5     | 0.7968 | 0.79832   | 0.80629    |                              |\n| MLP         | 446      | 5     | 0.7919 | 0.79102   | 0.79993    | KD using LGBM oof prediction |\n| GRU         | 188      | 5     | 0.7948 | 0.79418   | 0.80363    | KD using LGBM oof prediction |\n| Transformer | 188      | 5     | 0.7932 | 0.79498   | 0.80379    | KD using LGBM oof prediction |\n\n\n\n\n# 4. Ensemble \n\nWe used the two methods for ensemble with ranked probs. One is the LGBM stacking, the other is the CMA (Covariance Matrix Adaptation) Evolution Strategy\n[REF](https://www.scm.com/doc/params/python/optimizers/cmaes.html)\n\n\nAnd final submission is the average of the following 3 ensemble models(case1～3) ::\n|Case-Method|CV|public|private|#GBDT|#1/2dcnn|#MLP|#TCN|#GRU|#Transf|\n|------------|----------|-----------|-----------|-----------|----------|----------|--------|------------|\n|1-LGBM stacking| 0.8010 | 0.80084 | 0.80882    | 48 |2            | 5          | 2          | 3          | 1|\n|2- cma             | 0.8032 | 0.80086 | 0.80794    |  41 | 2            | 5          | 2          | 3          | 1|\n|3- cma             | 0.8028 | 0.80068 | 0.8079     |  49 | 0            | 3          | 0          | 3          | 0|\n\nAs you can see, the public lb is almost the same, but the private lb shows much better LGBM stacking. We would have been hard pressed to get the gold medal without LGBM stacking.\n\n## 4.1. Final sub (ensemble results of case 1 ~ 3)\nThe ensemble result of the above case 1-3 is our final submission.\npublic lb 0.80105, private lb 0.80842 (13th).\n\n※ Note that : LGBM stacking alone (0.80882) has the capability of ranking 6th.\n\n# 5. Relationship of cv and lb \n\nAs for reference, we share the relationship of cv and lb. It is easy to see without looking at the plot with cma, but the lgbm stacking is a clean extension of the cv/lb straight line of the single model to some extend (Maybe cma is overfitting…)\n\n![](https://raw.githubusercontent.com/chumajin/AMEX/main/AMEX_cv_lb.jpg)\n\n# 6. Other tips that worked well \n\n- knowledge distillation\n- longer early stop (3000～10000)\n- large kfold\n- full training\n- pseudo labeling\n\n\n# 7. Didn't work well \n- post process\n- using dow average data\n- changing weights of samples in LGBM\n- adding focal loss to LGBM\n- optimizing 0.4G+0.6D, 0.8G+0.2D, etc\n\n\n# 8. Late submission result\n After the end of competition we realized that we had 2 ensembles with PVT LB > 0.80920 (3-4th place) - however, were not considered due to lower CV scores.\n\n 1) avg rank of 3 stack models (XGB+LGB+Catboost) - PVT 0.80928, PUB 0.80067, CV 0.8008\n\n 2) LGB stack with 61 models + meta features (top 50 eng. features) - PVT 0.80926, PUB 0.80072, CV 0.80097\n\n# 9. Team organization \nWe used github for code storage, wandb for experiments tracking and kaggle datasets for OOF storage and sharing with the team.\n\nTeam members: \nhttps://www.kaggle.com/titericz\nhttps://www.kaggle.com/yamsam\nhttps://www.kaggle.com/imeintanis\nhttps://www.kaggle.com/chumajin\nhttps://www.kaggle.com/kurokurob",
      "votes": null
    },
    {
      "id": "1914799",
      "postDate": "08/26/2022 12:39:01",
      "content": "<p>Thank you Giba! Learned so much! Some of the tricks I honestly wouldn't have thought of trying myself, such as training 15-fold Dart LGBM. 5-fold already takes too long…Now that I know it works in some cases, it makes this idea a lot less daunting to try.</p>",
      "rawMarkdown": "Thank you Giba! Learned so much! Some of the tricks I honestly wouldn't have thought of trying myself, such as training 15-fold Dart LGBM. 5-fold already takes too long...Now that I know it works in some cases, it makes this idea a lot less daunting to try.",
      "votes": null
    },
    {
      "id": "1914877",
      "postDate": "08/26/2022 13:51:14",
      "content": "<p>Great work! Excited to see that my public model got leveraged at all by a gold medal team :)</p>",
      "rawMarkdown": "Great work! Excited to see that my public model got leveraged at all by a gold medal team :)",
      "votes": null
    },
    {
      "id": "1914879",
      "postDate": "08/26/2022 13:53:57",
      "content": "<p>And still only 5 fold on XGB? Did more not work?</p>",
      "rawMarkdown": "And still only 5 fold on XGB? Did more not work?",
      "votes": null
    },
    {
      "id": "1914899",
      "postDate": "08/26/2022 14:17:51",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> Dart XGB takes way too long. I mainly used plain XGB for diversity (with pruned features). Guess I should have tried your pyramid approach!</p>",
      "rawMarkdown": "roberthatch Dart XGB takes way too long. I mainly used plain XGB for diversity (with pruned features). Guess I should have tried your pyramid approach!",
      "votes": null
    },
    {
      "id": "1914915",
      "postDate": "08/26/2022 14:30:54",
      "content": "<p>Yeah, wondering if Giba tried more folds, since he did use the pyramid approach. </p>",
      "rawMarkdown": "Yeah, wondering if Giba tried more folds, since he did use the pyramid approach.",
      "votes": null
    },
    {
      "id": "1914966",
      "postDate": "08/26/2022 15:08:05",
      "content": "<p>good work buddy</p>",
      "rawMarkdown": "good work buddy",
      "votes": null
    },
    {
      "id": "1914993",
      "postDate": "08/26/2022 15:41:24",
      "content": "<p>Congratulations Giba and team! Well done.</p>\n<p>Could you explain your use of test data more? How did you do pseudo labeling. Did you use all test rows, or just confident predictions. And did you use soft or hard labels. Which models' test predictions did you use. And which models did you train with pseudo?</p>",
      "rawMarkdown": "Congratulations Giba and team! Well done.\n\nCould you explain your use of test data more? How did you do pseudo labeling. Did you use all test rows, or just confident predictions. And did you use soft or hard labels. Which models' test predictions did you use. And which models did you train with pseudo?",
      "votes": null
    },
    {
      "id": "1915030",
      "postDate": "08/26/2022 16:05:37",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> </p>",
      "rawMarkdown": "congrats @titericz",
      "votes": null
    },
    {
      "id": "1915209",
      "postDate": "08/26/2022 18:37:41",
      "content": "<p>Just confident predictions like &lt;0.1, &gt;0.8<br>\nTrain LGB/XGB models with soft labels from my best LGB (CV 0.7969, LB 0.799) at that time, however in the final ensembles only 1 or 2 models with pseudo are included.<br>\nPS: I will do some LB checks and give you more details on scores if you like. </p>\n<p>BTW any pseudo experiments tried with GRU failed to improve score, maybe I did something wrong, did you had any luck with that?</p>",
      "rawMarkdown": "Just confident predictions like <0.1, >0.8\nTrain LGB/XGB models with soft labels from my best LGB (CV 0.7969, LB 0.799) at that time, however in the final ensembles only 1 or 2 models with pseudo are included.\nPS: I will do some LB checks and give you more details on scores if you like. \n\nBTW any pseudo experiments tried with GRU failed to improve score, maybe I did something wrong, did you had any luck with that?",
      "votes": null
    },
    {
      "id": "1915242",
      "postDate": "08/26/2022 19:38:11",
      "content": "<p>No, pseudo labels did not help me. I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets. That made a huge boost like <code>+0.001 public and private test</code> and is my secret sauce.</p>\n<p>But all uses of pseudo labels failed. By pseudo labels, i mean training with test predictions during the finetuning (i.e. final epochs as opposed to pretraining). I tried using different proportions based on confidence, i tried hard and soft. I tried using just public test or just private test.</p>\n<p>I had a 100 model GBT nested 10 folds inside 10 folds that created CV 0.798 LB 0.799 leak free test preds, so I could check an accurate leak-free CV score when applying these pseudo labels to other models. The CV score of other models never improved and the LB never improved.</p>",
      "rawMarkdown": "No, pseudo labels did not help me. I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets. That made a huge boost like `+0.001 public and private test` and is my secret sauce.\n\nBut all uses of pseudo labels failed. By pseudo labels, i mean training with test predictions during the finetuning (i.e. final epochs as opposed to pretraining). I tried using different proportions based on confidence, i tried hard and soft. I tried using just public test or just private test.\n\nI had a 100 model GBT nested 10 folds inside 10 folds that created CV 0.798 LB 0.799 leak free test preds, so I could check an accurate leak-free CV score when applying these pseudo labels to other models. The CV score of other models never improved and the LB never improved.",
      "votes": null
    },
    {
      "id": "1915258",
      "postDate": "08/26/2022 20:03:38",
      "content": "<blockquote>\n  <p>I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets.</p>\n</blockquote>\n<p>Nice, I haven't thought of that, in Tabnet I tried smth similar, pretrain with train+test data using the unsupervised API but faced a bug and couldn't finished it. </p>\n<blockquote>\n  <p>The CV score of other models never improved and the LB never improved.</p>\n</blockquote>\n<p>Similar observations here, that is why we quit that path, however I remember we did precisely some checks in ensembles, with vs without pseudo models and both ensemble CV, Pub LB improved a bit.</p>",
      "rawMarkdown": "> I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets.\n\nNice, I haven't thought of that, in Tabnet I tried smth similar, pretrain with train+test data using the unsupervised API but faced a bug and couldn't finished it. \n\n> The CV score of other models never improved and the LB never improved.\n\nSimilar observations here, that is why we quit that path, however I remember we did precisely some checks in ensembles, with vs without pseudo models and both ensemble CV, Pub LB improved a bit.",
      "votes": null
    },
    {
      "id": "1915364",
      "postDate": "08/26/2022 23:08:45",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, excellent summary; it will help me to improve in the future</p>",
      "rawMarkdown": "Thanks for sharing @titericz, excellent summary; it will help me to improve in the future",
      "votes": null
    },
    {
      "id": "1915864",
      "postDate": "08/27/2022 12:46:58",
      "content": "<p>I tried training models using pseudo label and in general it got worse Public but better Private scores. So I didn't used it. </p>",
      "rawMarkdown": "I tried training models using pseudo label and in general it got worse Public but better Private scores. So I didn't used it.",
      "votes": null
    },
    {
      "id": "1916414",
      "postDate": "08/27/2022 21:21:43",
      "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>,</p>\n<p>Thank you so much for sharing. It is very useful. </p>",
      "rawMarkdown": "titericz,\n\nThank you so much for sharing. It is very useful.",
      "votes": null
    },
    {
      "id": "1920974",
      "postDate": "08/31/2022 14:02:25",
      "content": "<p>Congratulations!<br>\nThanks, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> for sharing the notebook from the competition so that rookies like me can also learn!</p>",
      "rawMarkdown": "Congratulations!\nThanks, @titericz for sharing the notebook from the competition so that rookies like me can also learn!",
      "votes": null
    },
    {
      "id": "1929986",
      "postDate": "09/07/2022 13:19:15",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> <a href=\"https://www.kaggle.com/kurokurob\" target=\"_blank\">@kurokurob</a> !</p>\n<p>How did you decide which models to include in each one of the ensembles? Hill climbing or something similar?<br>\nIn addition, I would like to know how did you ensemble oof predictions generated with different number of folds (e.g. 15 folds vs 5 folds).</p>\n<p>Thanks for sharing!</p>",
      "rawMarkdown": "Congrats @titericz @yamsam @imeintanis @chumajin @kurokurob !\n\nHow did you decide which models to include in each one of the ensembles? Hill climbing or something similar?\nIn addition, I would like to know how did you ensemble oof predictions generated with different number of folds (e.g. 15 folds vs 5 folds).\n\nThanks for sharing!",
      "votes": null
    },
    {
      "id": "1936431",
      "postDate": "09/12/2022 18:37:06",
      "content": "<p>can you share code, like  done for 1st place?</p>",
      "rawMarkdown": "can you share code, like  done for 1st place?",
      "votes": null
    },
    {
      "id": "2030199",
      "postDate": "11/15/2022 09:16:51",
      "content": "<p>excuse me, may i ask what's the meaning of KD using LGBM oof prediction, i know the LGBM oof prediction, but what's the meaning of KD? could you please tell something about it?</p>",
      "rawMarkdown": "excuse me, may i ask what's the meaning of KD using LGBM oof prediction, i know the LGBM oof prediction, but what's the meaning of KD? could you please tell something about it?",
      "votes": null
    },
    {
      "id": "2381367",
      "postDate": "08/09/2023 06:57:57",
      "content": "<p>maybe late，it is knowledge distillation</p>",
      "rawMarkdown": "maybe late，it is knowledge distillation",
      "votes": null
    },
    {
      "id": "2798689",
      "postDate": "05/07/2024 11:08:18",
      "content": "<p>Could you please explain what \"full training\" means? Thank you for your help.</p>",
      "rawMarkdown": "Could you please explain what \"full training\" means? Thank you for your help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1914799,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "08/26/2022 12:39:01",
      "content": "<p>Thank you Giba! Learned so much! Some of the tricks I honestly wouldn't have thought of trying myself, such as training 15-fold Dart LGBM. 5-fold already takes too long…Now that I know it works in some cases, it makes this idea a lot less daunting to try.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1914879,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/26/2022 13:53:57",
          "content": "<p>And still only 5 fold on XGB? Did more not work?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914899,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "08/26/2022 14:17:51",
          "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> Dart XGB takes way too long. I mainly used plain XGB for diversity (with pruned features). Guess I should have tried your pyramid approach!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914915,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/26/2022 14:30:54",
          "content": "<p>Yeah, wondering if Giba tried more folds, since he did use the pyramid approach. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1914877,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/26/2022 13:51:14",
      "content": "<p>Great work! Excited to see that my public model got leveraged at all by a gold medal team :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914966,
      "author_name": "ltrahul",
      "author_url": "",
      "post_date": "08/26/2022 15:08:05",
      "content": "<p>good work buddy</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914993,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/26/2022 15:41:24",
      "content": "<p>Congratulations Giba and team! Well done.</p>\n<p>Could you explain your use of test data more? How did you do pseudo labeling. Did you use all test rows, or just confident predictions. And did you use soft or hard labels. Which models' test predictions did you use. And which models did you train with pseudo?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915209,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "08/26/2022 18:37:41",
          "content": "<p>Just confident predictions like &lt;0.1, &gt;0.8<br>\nTrain LGB/XGB models with soft labels from my best LGB (CV 0.7969, LB 0.799) at that time, however in the final ensembles only 1 or 2 models with pseudo are included.<br>\nPS: I will do some LB checks and give you more details on scores if you like. </p>\n<p>BTW any pseudo experiments tried with GRU failed to improve score, maybe I did something wrong, did you had any luck with that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915242,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/26/2022 19:38:11",
          "content": "<p>No, pseudo labels did not help me. I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets. That made a huge boost like <code>+0.001 public and private test</code> and is my secret sauce.</p>\n<p>But all uses of pseudo labels failed. By pseudo labels, i mean training with test predictions during the finetuning (i.e. final epochs as opposed to pretraining). I tried using different proportions based on confidence, i tried hard and soft. I tried using just public test or just private test.</p>\n<p>I had a 100 model GBT nested 10 folds inside 10 folds that created CV 0.798 LB 0.799 leak free test preds, so I could check an accurate leak-free CV score when applying these pseudo labels to other models. The CV score of other models never improved and the LB never improved.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915258,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "08/26/2022 20:03:38",
          "content": "<blockquote>\n  <p>I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets.</p>\n</blockquote>\n<p>Nice, I haven't thought of that, in Tabnet I tried smth similar, pretrain with train+test data using the unsupervised API but faced a bug and couldn't finished it. </p>\n<blockquote>\n  <p>The CV score of other models never improved and the LB never improved.</p>\n</blockquote>\n<p>Similar observations here, that is why we quit that path, however I remember we did precisely some checks in ensembles, with vs without pseudo models and both ensemble CV, Pub LB improved a bit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915864,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "08/27/2022 12:46:58",
          "content": "<p>I tried training models using pseudo label and in general it got worse Public but better Private scores. So I didn't used it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915030,
      "author_name": "vandittyagi0909",
      "author_url": "",
      "post_date": "08/26/2022 16:05:37",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1915364,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "08/26/2022 23:08:45",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, excellent summary; it will help me to improve in the future</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1916414,
      "author_name": "gsigalaev",
      "author_url": "",
      "post_date": "08/27/2022 21:21:43",
      "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>,</p>\n<p>Thank you so much for sharing. It is very useful. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1920974,
      "author_name": "shehrozeshahzad",
      "author_url": "",
      "post_date": "08/31/2022 14:02:25",
      "content": "<p>Congratulations!<br>\nThanks, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> for sharing the notebook from the competition so that rookies like me can also learn!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1929986,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "09/07/2022 13:19:15",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> <a href=\"https://www.kaggle.com/kurokurob\" target=\"_blank\">@kurokurob</a> !</p>\n<p>How did you decide which models to include in each one of the ensembles? Hill climbing or something similar?<br>\nIn addition, I would like to know how did you ensemble oof predictions generated with different number of folds (e.g. 15 folds vs 5 folds).</p>\n<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1936431,
      "author_name": "sandystep",
      "author_url": "",
      "post_date": "09/12/2022 18:37:06",
      "content": "<p>can you share code, like  done for 1st place?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2030199,
      "author_name": "wcqglhf",
      "author_url": "",
      "post_date": "11/15/2022 09:16:51",
      "content": "<p>excuse me, may i ask what's the meaning of KD using LGBM oof prediction, i know the LGBM oof prediction, but what's the meaning of KD? could you please tell something about it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2381367,
          "author_name": "binaryhumble",
          "author_url": "",
          "post_date": "08/09/2023 06:57:57",
          "content": "<p>maybe late，it is knowledge distillation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2798689,
      "author_name": "cjzccc",
      "author_url": "",
      "post_date": "05/07/2024 11:08:18",
      "content": "<p>Could you please explain what \"full training\" means? Thank you for your help.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914748": "First of all, we would like to thank kaggle and the staff for hosting such a great competition. We really miss interesting tabular challenges these days.\n\n# 1. Summary \n Our solution is based on extensive data cleaning and multi-model weighted average and stacked ensembles using ranked probs. Using extensive data cleaning, our single model was boosted to cv 0.0004~0.0008 compared to the public Raddar's dataset.The final solution (PVT 0.80842/PUB 0.80105) which shaked up to gold zone  is the average of 3 ensemble models:\n- LGBM stack with 61 models\n- CMA(Covariance Matrix Adaptation) weighted average with 54 models\n- Weighted average with 55 models\n\n# 2. Extensive data cleaning \nStarting from Raddar’s (thanks @raddar) preprocessed dataset, some other features were cleaned, added and modified in order to build a second version of the dataset. The idea here was to clean the data even more and add some diversity to the ensemble. Some clusters of clients were added based on the pattern of missing variables for each set of features: B, D, P, R and S. Also, some continuous variables showed a curious noise pattern. Looks like a uniform noise was added in the range of (0 - 0.01] of some variables. Take a look at feature B_11 distribution. White noise was clearly added just in the 0.01 range.\n\n\n![](https://raw.githubusercontent.com/chumajin/AMEX/main/giba2.jpg)\n\n\nSpent some time trying to figure out a way to remove that noise and found some combinations of filters using other features that worked pretty well. For example, the B_11 feature can be cleaned by using a filter in B_1. Taking indices when B_1 is in range of 0-0.01 and inverting the signal of B_11 based on these indices, the new histogram becomes:\n\n![](https://raw.githubusercontent.com/chumajin/AMEX/main/giba3.jpg)\n\nThis kind of cleaning was made with some other features like: B_1, B_5, B_7, B_11, B_15, B_17, B_18, B_21, B_23, B_24, B_26, B_27, B_29, B_36, B_37, D_58, D_60, D_69, D_71, D_102, D_133, D_144, R_1, R_6, S_16, S_17, S_19, S_22 and S_27. Extra feature cleaning helped to boost GBDT scores and added some diversity when stacking with other models.\n\n\n\n\n\n## 2.1 Example of boosting cv using cleaning data\n- LGBM : 0.7976 → 0.7983 ( +0.0007 )\n- XGB : 0.7978 → 0.7986 ( + 0.0008 )\n- CatBoost :    0.7964 → 0.7968 ( + 0.0004 ) \n\n ※ clean data plus minor change \n   - LGBM 5kfold → 15kfold (+0.0013)\n   - CatBoost longer earlystop (+0.002 )\n\n\n# 3. Single model \n\n## 3.1 Features and modeling\n\nBasically, we used the features and the models of public notebooks.\nThank you for @thedevastator, @ragnar123, @ambrosm, @cdeotte, @roberthatch\n\nNote: some of the 1st level models don’t include features (B_29, R_1, D_59) from adversarial validation.\n\n## 3.2 Representative best single model in each model type \n\n| model       | features | kfold | cv     | public lb | private lb | comment                      |\n|-------------|----------|-------|--------|-----------|------------|------------------------------|\n| LGBM        | 1784     | 15    | 0.7991 | 0.79961   | 0.80679    | dart with early stop         |\n| XGB         | 3413     | 5     | 0.7986 | 0.79873   | 0.80662    | pyramid [REF](https://www.kaggle.com/code/roberthatch/xgboost-pyramid-cv-0-7968)                      |\n| CATBoost    | 2290     | 5     | 0.7968 | 0.79832   | 0.80629    |                              |\n| MLP         | 446      | 5     | 0.7919 | 0.79102   | 0.79993    | KD using LGBM oof prediction |\n| GRU         | 188      | 5     | 0.7948 | 0.79418   | 0.80363    | KD using LGBM oof prediction |\n| Transformer | 188      | 5     | 0.7932 | 0.79498   | 0.80379    | KD using LGBM oof prediction |\n\n\n\n\n# 4. Ensemble \n\nWe used the two methods for ensemble with ranked probs. One is the LGBM stacking, the other is the CMA (Covariance Matrix Adaptation) Evolution Strategy\n[REF](https://www.scm.com/doc/params/python/optimizers/cmaes.html)\n\n\nAnd final submission is the average of the following 3 ensemble models(case1～3) ::\n|Case-Method|CV|public|private|#GBDT|#1/2dcnn|#MLP|#TCN|#GRU|#Transf|\n|------------|----------|-----------|-----------|-----------|----------|----------|--------|------------|\n|1-LGBM stacking| 0.8010 | 0.80084 | 0.80882    | 48 |2            | 5          | 2          | 3          | 1|\n|2- cma             | 0.8032 | 0.80086 | 0.80794    |  41 | 2            | 5          | 2          | 3          | 1|\n|3- cma             | 0.8028 | 0.80068 | 0.8079     |  49 | 0            | 3          | 0          | 3          | 0|\n\nAs you can see, the public lb is almost the same, but the private lb shows much better LGBM stacking. We would have been hard pressed to get the gold medal without LGBM stacking.\n\n## 4.1. Final sub (ensemble results of case 1 ~ 3)\nThe ensemble result of the above case 1-3 is our final submission.\npublic lb 0.80105, private lb 0.80842 (13th).\n\n※ Note that : LGBM stacking alone (0.80882) has the capability of ranking 6th.\n\n# 5. Relationship of cv and lb \n\nAs for reference, we share the relationship of cv and lb. It is easy to see without looking at the plot with cma, but the lgbm stacking is a clean extension of the cv/lb straight line of the single model to some extend (Maybe cma is overfitting…)\n\n![](https://raw.githubusercontent.com/chumajin/AMEX/main/AMEX_cv_lb.jpg)\n\n# 6. Other tips that worked well \n\n- knowledge distillation\n- longer early stop (3000～10000)\n- large kfold\n- full training\n- pseudo labeling\n\n\n# 7. Didn't work well \n- post process\n- using dow average data\n- changing weights of samples in LGBM\n- adding focal loss to LGBM\n- optimizing 0.4G+0.6D, 0.8G+0.2D, etc\n\n\n# 8. Late submission result\n After the end of competition we realized that we had 2 ensembles with PVT LB > 0.80920 (3-4th place) - however, were not considered due to lower CV scores.\n\n 1) avg rank of 3 stack models (XGB+LGB+Catboost) - PVT 0.80928, PUB 0.80067, CV 0.8008\n\n 2) LGB stack with 61 models + meta features (top 50 eng. features) - PVT 0.80926, PUB 0.80072, CV 0.80097\n\n# 9. Team organization \nWe used github for code storage, wandb for experiments tracking and kaggle datasets for OOF storage and sharing with the team.\n\nTeam members: \nhttps://www.kaggle.com/titericz\nhttps://www.kaggle.com/yamsam\nhttps://www.kaggle.com/imeintanis\nhttps://www.kaggle.com/chumajin\nhttps://www.kaggle.com/kurokurob",
    "1914799": "Thank you Giba! Learned so much! Some of the tricks I honestly wouldn't have thought of trying myself, such as training 15-fold Dart LGBM. 5-fold already takes too long...Now that I know it works in some cases, it makes this idea a lot less daunting to try.",
    "1914877": "Great work! Excited to see that my public model got leveraged at all by a gold medal team :)",
    "1914879": "And still only 5 fold on XGB? Did more not work?",
    "1914899": "roberthatch Dart XGB takes way too long. I mainly used plain XGB for diversity (with pruned features). Guess I should have tried your pyramid approach!",
    "1914915": "Yeah, wondering if Giba tried more folds, since he did use the pyramid approach.",
    "1914966": "good work buddy",
    "1914993": "Congratulations Giba and team! Well done.\n\nCould you explain your use of test data more? How did you do pseudo labeling. Did you use all test rows, or just confident predictions. And did you use soft or hard labels. Which models' test predictions did you use. And which models did you train with pseudo?",
    "1915030": "congrats @titericz",
    "1915209": "Just confident predictions like <0.1, >0.8\nTrain LGB/XGB models with soft labels from my best LGB (CV 0.7969, LB 0.799) at that time, however in the final ensembles only 1 or 2 models with pseudo are included.\nPS: I will do some LB checks and give you more details on scores if you like. \n\nBTW any pseudo experiments tried with GRU failed to improve score, maybe I did something wrong, did you had any luck with that?",
    "1915242": "No, pseudo labels did not help me. I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets. That made a huge boost like `+0.001 public and private test` and is my secret sauce.\n\nBut all uses of pseudo labels failed. By pseudo labels, i mean training with test predictions during the finetuning (i.e. final epochs as opposed to pretraining). I tried using different proportions based on confidence, i tried hard and soft. I tried using just public test or just private test.\n\nI had a 100 model GBT nested 10 folds inside 10 folds that created CV 0.798 LB 0.799 leak free test preds, so I could check an accurate leak-free CV score when applying these pseudo labels to other models. The CV score of other models never improved and the LB never improved.",
    "1915258": "> I used soft test predictions from LGBM to pretrain my NN Transformer and then fine tuned on only train targets.\n\nNice, I haven't thought of that, in Tabnet I tried smth similar, pretrain with train+test data using the unsupervised API but faced a bug and couldn't finished it. \n\n> The CV score of other models never improved and the LB never improved.\n\nSimilar observations here, that is why we quit that path, however I remember we did precisely some checks in ensembles, with vs without pseudo models and both ensemble CV, Pub LB improved a bit.",
    "1915364": "Thanks for sharing @titericz, excellent summary; it will help me to improve in the future",
    "1915864": "I tried training models using pseudo label and in general it got worse Public but better Private scores. So I didn't used it.",
    "1916414": "titericz,\n\nThank you so much for sharing. It is very useful.",
    "1920974": "Congratulations!\nThanks, @titericz for sharing the notebook from the competition so that rookies like me can also learn!",
    "1929986": "Congrats @titericz @yamsam @imeintanis @chumajin @kurokurob !\n\nHow did you decide which models to include in each one of the ensembles? Hill climbing or something similar?\nIn addition, I would like to know how did you ensemble oof predictions generated with different number of folds (e.g. 15 folds vs 5 folds).\n\nThanks for sharing!",
    "1936431": "can you share code, like  done for 1st place?",
    "2030199": "excuse me, may i ask what's the meaning of KD using LGBM oof prediction, i know the LGBM oof prediction, but what's the meaning of KD? could you please tell something about it?",
    "2381367": "maybe late，it is knowledge distillation",
    "2798689": "Could you please explain what \"full training\" means? Thank you for your help."
  },
  "source": "meta"
}