{
  "id": 347757,
  "title": "82-nd place (Silver) solution",
  "url": "/competitions/amex-default-prediction/writeups/georgi-pamukov-82-nd-place-silver-solution",
  "author_name": "",
  "post_date": "2022-08-25T10:19:16.307Z",
  "votes": 20,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey Dear Kagglers :)</p>\n<p>First of all, congratulations to all winners and participants!<br>\nAnd many thanks to all contributors for their amazing kernels and topics!</p>\n<p>A brief overview of my solution below:</p>\n<h3>Feature engineering:</h3>\n<ul>\n<li>Simple agg features</li>\n<li>Feature interactions - lagg/diff</li>\n<li>Trend features</li>\n<li>Correlations between normalized feature and time index</li>\n<li>Autocorrelations</li>\n<li>Moving averages – and then aggregations like first/last/range</li>\n<li>Target encoding for categorical</li>\n<li>OHE</li>\n<li>Cluster based features</li>\n<li>PCA reduced features</li>\n<li>Latent vectors exported from deep autoencoders (heavy regularization + data augmentation)</li>\n<li>Other</li>\n</ul>\n<h3>Data sets</h3>\n<p>Total of 12 training sets:</p>\n<ul>\n<li>Combinations of features above</li>\n<li>Re-balanced sets (downsampling majority)</li>\n<li>Reduced sets</li>\n<li>Pseudo labeling</li>\n</ul>\n<h3>Modeling/Architecture – 4 levels:</h3>\n<p><strong>1) Base models trained:</strong></p>\n<ul>\n<li>LGB/XGB/Catboost</li>\n<li>GLM</li>\n<li>MLP</li>\n<li>DCN - Deep and cross neural networks (residual) with entity embeddings for the categoricals</li>\n</ul>\n<p><strong>2) Some 2-nd level models trained on the best data set but with base models predictions added – it resulted in very high performance models</strong></p>\n<p><strong>3) 3-rd level models trained on predictions from 1-st and 2-nd level models + PCA + Cluster features:</strong></p>\n<ul>\n<li>XGB</li>\n<li>NN with entity embeddings for the cluster features</li>\n</ul>\n<p><strong>4) Final blend of 3-rd level model predictions</strong></p>\n<p>CV amex score maximization was pursued which worked well – managed to choose (almost) the best model; Model performs better on private than on public lb.</p>\n<p>This was all done on a single Core I5/16GB/geforce 1060 6gb - so with limited resources…</p>\n<p>Let me know if you have any questions in the comments.</p>",
  "messages": [
    {
      "id": "1913419",
      "postDate": "08/25/2022 10:14:21",
      "content": "<p>Hey Dear Kagglers :)</p>\n<p>First of all, congratulations to all winners and participants!<br>\nAnd many thanks to all contributors for their amazing kernels and topics!</p>\n<p>A brief overview of my solution below:</p>\n<h3>Feature engineering:</h3>\n<ul>\n<li>Simple agg features</li>\n<li>Feature interactions - lagg/diff</li>\n<li>Trend features</li>\n<li>Correlations between normalized feature and time index</li>\n<li>Autocorrelations</li>\n<li>Moving averages – and then aggregations like first/last/range</li>\n<li>Target encoding for categorical</li>\n<li>OHE</li>\n<li>Cluster based features</li>\n<li>PCA reduced features</li>\n<li>Latent vectors exported from deep autoencoders (heavy regularization + data augmentation)</li>\n<li>Other</li>\n</ul>\n<h3>Data sets</h3>\n<p>Total of 12 training sets:</p>\n<ul>\n<li>Combinations of features above</li>\n<li>Re-balanced sets (downsampling majority)</li>\n<li>Reduced sets</li>\n<li>Pseudo labeling</li>\n</ul>\n<h3>Modeling/Architecture – 4 levels:</h3>\n<p><strong>1) Base models trained:</strong></p>\n<ul>\n<li>LGB/XGB/Catboost</li>\n<li>GLM</li>\n<li>MLP</li>\n<li>DCN - Deep and cross neural networks (residual) with entity embeddings for the categoricals</li>\n</ul>\n<p><strong>2) Some 2-nd level models trained on the best data set but with base models predictions added – it resulted in very high performance models</strong></p>\n<p><strong>3) 3-rd level models trained on predictions from 1-st and 2-nd level models + PCA + Cluster features:</strong></p>\n<ul>\n<li>XGB</li>\n<li>NN with entity embeddings for the cluster features</li>\n</ul>\n<p><strong>4) Final blend of 3-rd level model predictions</strong></p>\n<p>CV amex score maximization was pursued which worked well – managed to choose (almost) the best model; Model performs better on private than on public lb.</p>\n<p>This was all done on a single Core I5/16GB/geforce 1060 6gb - so with limited resources…</p>\n<p>Let me know if you have any questions in the comments.</p>",
      "rawMarkdown": "Hey Dear Kagglers :)\n\nFirst of all, congratulations to all winners and participants!\nAnd many thanks to all contributors for their amazing kernels and topics!\n\nA brief overview of my solution below:\n\n### Feature engineering:\n- Simple agg features\n- Feature interactions - lagg/diff\n- Trend features\n- Correlations between normalized feature and time index\n- Autocorrelations\n- Moving averages – and then aggregations like first/last/range\n- Target encoding for categorical\n- OHE\n- Cluster based features\n- PCA reduced features\n- Latent vectors exported from deep autoencoders (heavy regularization + data augmentation)\n- Other\n### Data sets\nTotal of 12 training sets:\n- Combinations of features above\n- Re-balanced sets (downsampling majority)\n- Reduced sets\n- Pseudo labeling\n### Modeling/Architecture – 4 levels:\n**1) Base models trained:**\n- LGB/XGB/Catboost\n- GLM\n- MLP\n- DCN - Deep and cross neural networks (residual) with entity embeddings for the categoricals\n\n**2) Some 2-nd level models trained on the best data set but with base models predictions added – it resulted in very high performance models**\n\n**3) 3-rd level models trained on predictions from 1-st and 2-nd level models + PCA + Cluster features:**\n- XGB\n- NN with entity embeddings for the cluster features\n\n**4) Final blend of 3-rd level model predictions**\n\nCV amex score maximization was pursued which worked well – managed to choose (almost) the best model; Model performs better on private than on public lb.\n\nThis was all done on a single Core I5/16GB/geforce 1060 6gb - so with limited resources...\n\nLet me know if you have any questions in the comments.",
      "votes": null
    },
    {
      "id": "1913842",
      "postDate": "08/25/2022 14:33:50",
      "content": "<p>I appreciate the approach with your hardware! Hearty congratulations and wishing you the best <a href=\"https://www.kaggle.com/gpamoukoff\" target=\"_blank\">@gpamoukoff</a> !!</p>",
      "rawMarkdown": "I appreciate the approach with your hardware! Hearty congratulations and wishing you the best @gpamoukoff !!",
      "votes": null
    },
    {
      "id": "1913973",
      "postDate": "08/25/2022 16:00:46",
      "content": "<p>Thanks, mate - all the best to you!</p>",
      "rawMarkdown": "Thanks, mate - all the best to you!",
      "votes": null
    },
    {
      "id": "1915491",
      "postDate": "08/27/2022 02:40:26",
      "content": "<p>I like this approach <a href=\"https://www.kaggle.com/gpamoukoff\" target=\"_blank\">@gpamoukoff</a> </p>",
      "rawMarkdown": "I like this approach @gpamoukoff",
      "votes": null
    },
    {
      "id": "1915924",
      "postDate": "08/27/2022 13:32:07",
      "content": "<p>nice map brother..</p>",
      "rawMarkdown": "nice map brother..",
      "votes": null
    },
    {
      "id": "1917496",
      "postDate": "08/28/2022 19:20:38",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1917499",
      "postDate": "08/28/2022 19:22:02",
      "content": "<p>Cheers mate :)</p>",
      "rawMarkdown": "Cheers mate :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913842,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/25/2022 14:33:50",
      "content": "<p>I appreciate the approach with your hardware! Hearty congratulations and wishing you the best <a href=\"https://www.kaggle.com/gpamoukoff\" target=\"_blank\">@gpamoukoff</a> !!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913973,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "08/25/2022 16:00:46",
          "content": "<p>Thanks, mate - all the best to you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915491,
      "author_name": "tarundalal",
      "author_url": "",
      "post_date": "08/27/2022 02:40:26",
      "content": "<p>I like this approach <a href=\"https://www.kaggle.com/gpamoukoff\" target=\"_blank\">@gpamoukoff</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1917496,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "08/28/2022 19:20:38",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915924,
      "author_name": "srijan70",
      "author_url": "",
      "post_date": "08/27/2022 13:32:07",
      "content": "<p>nice map brother..</p>",
      "votes": null,
      "replies": [
        {
          "id": 1917499,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "08/28/2022 19:22:02",
          "content": "<p>Cheers mate :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1913419": "Hey Dear Kagglers :)\n\nFirst of all, congratulations to all winners and participants!\nAnd many thanks to all contributors for their amazing kernels and topics!\n\nA brief overview of my solution below:\n\n### Feature engineering:\n- Simple agg features\n- Feature interactions - lagg/diff\n- Trend features\n- Correlations between normalized feature and time index\n- Autocorrelations\n- Moving averages – and then aggregations like first/last/range\n- Target encoding for categorical\n- OHE\n- Cluster based features\n- PCA reduced features\n- Latent vectors exported from deep autoencoders (heavy regularization + data augmentation)\n- Other\n### Data sets\nTotal of 12 training sets:\n- Combinations of features above\n- Re-balanced sets (downsampling majority)\n- Reduced sets\n- Pseudo labeling\n### Modeling/Architecture – 4 levels:\n**1) Base models trained:**\n- LGB/XGB/Catboost\n- GLM\n- MLP\n- DCN - Deep and cross neural networks (residual) with entity embeddings for the categoricals\n\n**2) Some 2-nd level models trained on the best data set but with base models predictions added – it resulted in very high performance models**\n\n**3) 3-rd level models trained on predictions from 1-st and 2-nd level models + PCA + Cluster features:**\n- XGB\n- NN with entity embeddings for the cluster features\n\n**4) Final blend of 3-rd level model predictions**\n\nCV amex score maximization was pursued which worked well – managed to choose (almost) the best model; Model performs better on private than on public lb.\n\nThis was all done on a single Core I5/16GB/geforce 1060 6gb - so with limited resources...\n\nLet me know if you have any questions in the comments.",
    "1913842": "I appreciate the approach with your hardware! Hearty congratulations and wishing you the best @gpamoukoff !!",
    "1913973": "Thanks, mate - all the best to you!",
    "1915491": "I like this approach @gpamoukoff",
    "1915924": "nice map brother..",
    "1917496": "Thank you!",
    "1917499": "Cheers mate :)"
  },
  "source": "meta"
}