{
  "id": 347858,
  "title": "[16th place solution] Features Diversity and Ensemble",
  "url": "/competitions/amex-default-prediction/writeups/www-tryninjastudy-com-16th-place-solution-features",
  "author_name": "",
  "post_date": "2022-09-01T09:13:20.063Z",
  "votes": 27,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We would like to thank the organizers and the Kaggle community for providing such a great competition. <br>\nI would like to thank <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> and <a href=\"https://www.kaggle.com/eventhorizon28\" target=\"_blank\">@eventhorizon28</a> for their support and contribution, our team's collective hard work helped us achieve this position. <br>\nSpecial thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>, <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>  for the awesome work they published without their analytics, This competition would have a different direction from what  it is now.</p>\n<p>Here brief Explanation of our solution.</p>\n<p><em>FEATURE ENGINEERING</em><br>\n<em>DIVERSITY IN MODELS</em></p>\n<h4>Feature Engineering</h4>\n<p>We used different features for various model training (mean, std, and last features were common). We trained our models in three ways.</p>\n<ol>\n<li>using only HMA(hull moving average) features</li>\n<li>using only diff(features)</li>\n<li>using HMA + diff features ( worked only with cat boost)</li>\n</ol>\n<p>Using all the diff features was not the right call for some of the models like NN and XG Boost as they were introducing some leakage while training and CV and LB  didn't correlate at all. For us, HMA features proved to be much better featured than diff features.</p>\n<h4>Models</h4>\n<p>We used various models that include, LGBM, XG Boost, Cat Boost, and Neural Networks with 2 different architectures and TABNET.</p>\n<p>Here are our best single model scores</p>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>Cross Val</th>\n<th>private LB</th>\n<th>public LB</th>\n<th>Description</th>\n<th>Core Features</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM</td>\n<td>.7973</td>\n<td>0.80687</td>\n<td>0.79906</td>\n<td>3 models with different seed+ 2 public models</td>\n<td>Diff+Last</td>\n</tr>\n<tr>\n<td>XG Boost</td>\n<td>.7972</td>\n<td>0.80639</td>\n<td>0.79718</td>\n<td>3 models with different seed+ 1 public model</td>\n<td>HMA+Last</td>\n</tr>\n<tr>\n<td>CAT Boost</td>\n<td>.7952</td>\n<td>0.80468</td>\n<td>0.79614</td>\n<td>3 models with different seed</td>\n<td>HMA+diff+Last</td>\n</tr>\n<tr>\n<td>NN-1</td>\n<td>.7923</td>\n<td>0.80190</td>\n<td>0.79240</td>\n<td>3 models with different seed</td>\n<td>HMA+Last</td>\n</tr>\n<tr>\n<td>NN-2</td>\n<td>.7921</td>\n<td>0.80186</td>\n<td>0.79188</td>\n<td>2 models with different seed</td>\n<td>diff + Last</td>\n</tr>\n<tr>\n<td>TABNET</td>\n<td>.7933</td>\n<td>-</td>\n<td>-</td>\n<td>single model trained last day to introduce diversity</td>\n<td>HMA + Last</td>\n</tr>\n</tbody>\n</table>\n<h4>ENSEMBLE</h4>\n<p>Since we had such a great CV LB correlation, we used Optuna to choose the ensemble weights and performed rank ensemble for our submissions. However, since some good public models didn't have oofs predictions we had to give weights to them manually.</p>\n<p>Things that we were not able to try:</p>\n<ol>\n<li>Using B_29  predicting the missing values and then using them as features</li>\n<li>LGBM with HMA feature that would be our best public model</li>\n<li>pseudo-labelling using just private LB data</li>\n</ol>\n<p>Notebooks that helped us:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></li>\n<li><a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></li>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training</a></li>\n<li><a href=\"https://www.kaggle.com/code/werus23/amex-keras-with-tpu\" target=\"_blank\">https://www.kaggle.com/code/werus23/amex-keras-with-tpu</a></li>\n<li><a href=\"https://www.kaggle.com/code/raddar/understanding-na-values-in-amex-competition\" target=\"_blank\">https://www.kaggle.com/code/raddar/understanding-na-values-in-amex-competition</a></li>\n<li><a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added</a></li>\n<li><a href=\"https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></li>\n</ul>",
  "messages": [
    {
      "id": "1914007",
      "postDate": "08/25/2022 16:39:20",
      "content": "<p>We would like to thank the organizers and the Kaggle community for providing such a great competition. <br>\nI would like to thank <a href=\"https://www.kaggle.com/shivamcyborg\" target=\"_blank\">@shivamcyborg</a> and <a href=\"https://www.kaggle.com/eventhorizon28\" target=\"_blank\">@eventhorizon28</a> for their support and contribution, our team's collective hard work helped us achieve this position. <br>\nSpecial thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>, <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>  for the awesome work they published without their analytics, This competition would have a different direction from what  it is now.</p>\n<p>Here brief Explanation of our solution.</p>\n<p><em>FEATURE ENGINEERING</em><br>\n<em>DIVERSITY IN MODELS</em></p>\n<h4>Feature Engineering</h4>\n<p>We used different features for various model training (mean, std, and last features were common). We trained our models in three ways.</p>\n<ol>\n<li>using only HMA(hull moving average) features</li>\n<li>using only diff(features)</li>\n<li>using HMA + diff features ( worked only with cat boost)</li>\n</ol>\n<p>Using all the diff features was not the right call for some of the models like NN and XG Boost as they were introducing some leakage while training and CV and LB  didn't correlate at all. For us, HMA features proved to be much better featured than diff features.</p>\n<h4>Models</h4>\n<p>We used various models that include, LGBM, XG Boost, Cat Boost, and Neural Networks with 2 different architectures and TABNET.</p>\n<p>Here are our best single model scores</p>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>Cross Val</th>\n<th>private LB</th>\n<th>public LB</th>\n<th>Description</th>\n<th>Core Features</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM</td>\n<td>.7973</td>\n<td>0.80687</td>\n<td>0.79906</td>\n<td>3 models with different seed+ 2 public models</td>\n<td>Diff+Last</td>\n</tr>\n<tr>\n<td>XG Boost</td>\n<td>.7972</td>\n<td>0.80639</td>\n<td>0.79718</td>\n<td>3 models with different seed+ 1 public model</td>\n<td>HMA+Last</td>\n</tr>\n<tr>\n<td>CAT Boost</td>\n<td>.7952</td>\n<td>0.80468</td>\n<td>0.79614</td>\n<td>3 models with different seed</td>\n<td>HMA+diff+Last</td>\n</tr>\n<tr>\n<td>NN-1</td>\n<td>.7923</td>\n<td>0.80190</td>\n<td>0.79240</td>\n<td>3 models with different seed</td>\n<td>HMA+Last</td>\n</tr>\n<tr>\n<td>NN-2</td>\n<td>.7921</td>\n<td>0.80186</td>\n<td>0.79188</td>\n<td>2 models with different seed</td>\n<td>diff + Last</td>\n</tr>\n<tr>\n<td>TABNET</td>\n<td>.7933</td>\n<td>-</td>\n<td>-</td>\n<td>single model trained last day to introduce diversity</td>\n<td>HMA + Last</td>\n</tr>\n</tbody>\n</table>\n<h4>ENSEMBLE</h4>\n<p>Since we had such a great CV LB correlation, we used Optuna to choose the ensemble weights and performed rank ensemble for our submissions. However, since some good public models didn't have oofs predictions we had to give weights to them manually.</p>\n<p>Things that we were not able to try:</p>\n<ol>\n<li>Using B_29  predicting the missing values and then using them as features</li>\n<li>LGBM with HMA feature that would be our best public model</li>\n<li>pseudo-labelling using just private LB data</li>\n</ol>\n<p>Notebooks that helped us:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></li>\n<li><a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></li>\n<li><a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training</a></li>\n<li><a href=\"https://www.kaggle.com/code/werus23/amex-keras-with-tpu\" target=\"_blank\">https://www.kaggle.com/code/werus23/amex-keras-with-tpu</a></li>\n<li><a href=\"https://www.kaggle.com/code/raddar/understanding-na-values-in-amex-competition\" target=\"_blank\">https://www.kaggle.com/code/raddar/understanding-na-values-in-amex-competition</a></li>\n<li><a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added</a></li>\n<li><a href=\"https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></li>\n</ul>",
      "rawMarkdown": "We would like to thank the organizers and the Kaggle community for providing such a great competition. \nI would like to thank @shivamcyborg and @eventhorizon28 for their support and contribution, our team's collective hard work helped us achieve this position. \nSpecial thanks to @raddar, @roberthatch, @cdeotte, @jiweiliu, @ragnar123  for the awesome work they published without their analytics, This competition would have a different direction from what  it is now.\n\nHere brief Explanation of our solution.\n\n*FEATURE ENGINEERING*\n*DIVERSITY IN MODELS*\n\n#### Feature Engineering\nWe used different features for various model training (mean, std, and last features were common). We trained our models in three ways.\n1. using only HMA(hull moving average) features\n2. using only diff(features)\n3. using HMA + diff features ( worked only with cat boost)\n\nUsing all the diff features was not the right call for some of the models like NN and XG Boost as they were introducing some leakage while training and CV and LB  didn't correlate at all. For us, HMA features proved to be much better featured than diff features.\n\n#### Models\nWe used various models that include, LGBM, XG Boost, Cat Boost, and Neural Networks with 2 different architectures and TABNET.\n\nHere are our best single model scores\n| models | Cross Val | private LB | public LB | Description | Core Features | \n| --- | --- | --- | --- |\n| LGBM |.7973 |0.80687 |0.79906|3 models with different seed+ 2 public models |Diff+Last|\n| XG Boost |.7972 |0.80639|0.79718|3 models with different seed+ 1 public model |HMA+Last|\n| CAT Boost |.7952 |0.80468|0.79614|3 models with different seed |HMA+diff+Last|\n| NN-1 |.7923 |0.80190|0.79240|3 models with different seed |HMA+Last|\n| NN-2 |.7921 |0.80186|0.79188|2 models with different seed |diff + Last|\n|TABNET |.7933 |-|-|single model trained last day to introduce diversity |HMA + Last|\n\n#### ENSEMBLE\nSince we had such a great CV LB correlation, we used Optuna to choose the ensemble weights and performed rank ensemble for our submissions. However, since some good public models didn't have oofs predictions we had to give weights to them manually.\n\nThings that we were not able to try:\n1. Using B_29  predicting the missing values and then using them as features\n2. LGBM with HMA feature that would be our best public model\n3. pseudo-labelling using just private LB data\n\n\nNotebooks that helped us:\n\n- https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n- https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n- https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\n- https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\n- https://www.kaggle.com/code/werus23/amex-keras-with-tpu\n- https://www.kaggle.com/code/raddar/understanding-na-values-in-amex-competition\n- https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\n- https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick.\n- https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
      "votes": null
    },
    {
      "id": "1914037",
      "postDate": "08/25/2022 17:11:29",
      "content": "<p>Huge congrats on the gold finish! <br>\nJust curious, were your final submissions among your highest scoring submissions on the private leaderboard?</p>",
      "rawMarkdown": "Huge congrats on the gold finish! \nJust curious, were your final submissions among your highest scoring submissions on the private leaderboard?",
      "votes": null
    },
    {
      "id": "1914054",
      "postDate": "08/25/2022 17:25:40",
      "content": "<p>Thanks a ton, we made our final subs yesterday. our best sub was 0.80839 which we didn't choose. We also tried one late submission it turned out that if we would have given more weight to TABNET and LGBM we would be on the 13th spot with 0.80846. However, I am still happy with what we achieved.<br>\nRegards</p>",
      "rawMarkdown": "Thanks a ton, we made our final subs yesterday. our best sub was 0.80839 which we didn't choose. We also tried one late submission it turned out that if we would have given more weight to TABNET and LGBM we would be on the 13th spot with 0.80846. However, I am still happy with what we achieved.\nRegards",
      "votes": null
    },
    {
      "id": "1914263",
      "postDate": "08/25/2022 23:05:14",
      "content": "<p>HMA! Awesome to see it do so well, though I think it had more to do with all the hard work and great ensemble your team built!</p>\n<p>Congrats!</p>",
      "rawMarkdown": "HMA! Awesome to see it do so well, though I think it had more to do with all the hard work and great ensemble your team built!\n\nCongrats!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1914037,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "08/25/2022 17:11:29",
      "content": "<p>Huge congrats on the gold finish! <br>\nJust curious, were your final submissions among your highest scoring submissions on the private leaderboard?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1914054,
          "author_name": "chaudharypriyanshu",
          "author_url": "",
          "post_date": "08/25/2022 17:25:40",
          "content": "<p>Thanks a ton, we made our final subs yesterday. our best sub was 0.80839 which we didn't choose. We also tried one late submission it turned out that if we would have given more weight to TABNET and LGBM we would be on the 13th spot with 0.80846. However, I am still happy with what we achieved.<br>\nRegards</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1914263,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/25/2022 23:05:14",
      "content": "<p>HMA! Awesome to see it do so well, though I think it had more to do with all the hard work and great ensemble your team built!</p>\n<p>Congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914007": "We would like to thank the organizers and the Kaggle community for providing such a great competition. \nI would like to thank @shivamcyborg and @eventhorizon28 for their support and contribution, our team's collective hard work helped us achieve this position. \nSpecial thanks to @raddar, @roberthatch, @cdeotte, @jiweiliu, @ragnar123  for the awesome work they published without their analytics, This competition would have a different direction from what  it is now.\n\nHere brief Explanation of our solution.\n\n*FEATURE ENGINEERING*\n*DIVERSITY IN MODELS*\n\n#### Feature Engineering\nWe used different features for various model training (mean, std, and last features were common). We trained our models in three ways.\n1. using only HMA(hull moving average) features\n2. using only diff(features)\n3. using HMA + diff features ( worked only with cat boost)\n\nUsing all the diff features was not the right call for some of the models like NN and XG Boost as they were introducing some leakage while training and CV and LB  didn't correlate at all. For us, HMA features proved to be much better featured than diff features.\n\n#### Models\nWe used various models that include, LGBM, XG Boost, Cat Boost, and Neural Networks with 2 different architectures and TABNET.\n\nHere are our best single model scores\n| models | Cross Val | private LB | public LB | Description | Core Features | \n| --- | --- | --- | --- |\n| LGBM |.7973 |0.80687 |0.79906|3 models with different seed+ 2 public models |Diff+Last|\n| XG Boost |.7972 |0.80639|0.79718|3 models with different seed+ 1 public model |HMA+Last|\n| CAT Boost |.7952 |0.80468|0.79614|3 models with different seed |HMA+diff+Last|\n| NN-1 |.7923 |0.80190|0.79240|3 models with different seed |HMA+Last|\n| NN-2 |.7921 |0.80186|0.79188|2 models with different seed |diff + Last|\n|TABNET |.7933 |-|-|single model trained last day to introduce diversity |HMA + Last|\n\n#### ENSEMBLE\nSince we had such a great CV LB correlation, we used Optuna to choose the ensemble weights and performed rank ensemble for our submissions. However, since some good public models didn't have oofs predictions we had to give weights to them manually.\n\nThings that we were not able to try:\n1. Using B_29  predicting the missing values and then using them as features\n2. LGBM with HMA feature that would be our best public model\n3. pseudo-labelling using just private LB data\n\n\nNotebooks that helped us:\n\n- https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n- https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n- https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\n- https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\n- https://www.kaggle.com/code/werus23/amex-keras-with-tpu\n- https://www.kaggle.com/code/raddar/understanding-na-values-in-amex-competition\n- https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\n- https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick.\n- https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
    "1914037": "Huge congrats on the gold finish! \nJust curious, were your final submissions among your highest scoring submissions on the private leaderboard?",
    "1914054": "Thanks a ton, we made our final subs yesterday. our best sub was 0.80839 which we didn't choose. We also tried one late submission it turned out that if we would have given more weight to TABNET and LGBM we would be on the 13th spot with 0.80846. However, I am still happy with what we achieved.\nRegards",
    "1914263": "HMA! Awesome to see it do so well, though I think it had more to do with all the hard work and great ensemble your team built!\n\nCongrats!"
  },
  "source": "meta"
}