{
  "id": 348118,
  "title": "5th Place Solution - Team 💳VISA💳(Patrick's part)",
  "url": "/competitions/amex-default-prediction/discussion/348118",
  "author_name": "",
  "post_date": "2022-08-27T02:10:49.057823Z",
  "votes": 51,
  "comment_count": 11,
  "views": 0,
  "content": "<p>URL to the Summary&amp;zakopuro part: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097</a></p>\n<p>To begin with, I would like to thank Amex for hosting the competition and my teammates ( <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a>). Congrats to <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> and <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> for being a Kaggle competition master!</p>\n<h1><strong>Pretrain + Finetune approach</strong></h1>\n<p>LightGBM + Feature engineering is very successful in this competition, and they outperform NN + raw features most of the time. Because of this, I believe the features used in LGBM models are very powerful, so I decided to first let the NN learn how to do feature engineering in the pretrain stage, then finetune the model with the target after that.<br>\nIn this way, we provide much more guidance to train the model by using thousands of features, and most importantly we can include test data in the pretrain stage.<br>\nWith pretraining, the model can have <strong>+0.002</strong> boost in public LB compared with training from scratch, and we can train a larger model (6 layers transformer) without any problem.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2F044ba388c80c6653df48f7238b3a8698%2FAmex_transformer.png?generation=1661564543157828&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fe1857b02f2ae097535b695d6fce903a3%2FAmex_Features_encoder.png?generation=1661564557974134&amp;alt=media\" alt=\"\"></p>\n<h1><strong>Model architecture</strong></h1>\n<p>We use different MLP layers to handle different types of inputs (Delinquency, Spend, Payment, Balance, Risk variables), then concatenate them and pass them to the transformer encoder. After the encoding part, we take the latest node and get the outputs through a linear layer.</p>\n<h1><strong>Pretrain stage</strong></h1>\n<p>In the pretrain stage, our target is tabular features. We use Huber loss to train the standardized target and won't pass any loss if the feature is nan. The number of epochs is around 200 in this stage.</p>\n<h1><strong>Finetune stage</strong></h1>\n<p>In the finetune stage, we train the model with the label. With pretraining, the converging speed for the model is very fast and we only need less than 5 epochs per fold in this stage!</p>\n<h1><strong>Model performance</strong></h1>\n<p>Our best Transformer model without meta features has <strong>0.794</strong> CV, <strong>0.796</strong> public LB, and <strong>0.804</strong> private LB.<br>\nWith pseudo labeling (We use soft predictions from our best ensemble model), a single transformer model can have <strong>0.800</strong> public LB and <strong>0.808</strong> private LB. (There is leaking because we didn't use Nested K-fold CV to generate pseudo label).</p>",
  "messages": [
    {
      "id": "1915468",
      "postDate": "08/27/2022 02:10:49",
      "content": "<p>URL to the Summary&amp;zakopuro part: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097</a></p>\n<p>To begin with, I would like to thank Amex for hosting the competition and my teammates ( <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a>). Congrats to <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> and <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> for being a Kaggle competition master!</p>\n<h1><strong>Pretrain + Finetune approach</strong></h1>\n<p>LightGBM + Feature engineering is very successful in this competition, and they outperform NN + raw features most of the time. Because of this, I believe the features used in LGBM models are very powerful, so I decided to first let the NN learn how to do feature engineering in the pretrain stage, then finetune the model with the target after that.<br>\nIn this way, we provide much more guidance to train the model by using thousands of features, and most importantly we can include test data in the pretrain stage.<br>\nWith pretraining, the model can have <strong>+0.002</strong> boost in public LB compared with training from scratch, and we can train a larger model (6 layers transformer) without any problem.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2F044ba388c80c6653df48f7238b3a8698%2FAmex_transformer.png?generation=1661564543157828&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fe1857b02f2ae097535b695d6fce903a3%2FAmex_Features_encoder.png?generation=1661564557974134&amp;alt=media\" alt=\"\"></p>\n<h1><strong>Model architecture</strong></h1>\n<p>We use different MLP layers to handle different types of inputs (Delinquency, Spend, Payment, Balance, Risk variables), then concatenate them and pass them to the transformer encoder. After the encoding part, we take the latest node and get the outputs through a linear layer.</p>\n<h1><strong>Pretrain stage</strong></h1>\n<p>In the pretrain stage, our target is tabular features. We use Huber loss to train the standardized target and won't pass any loss if the feature is nan. The number of epochs is around 200 in this stage.</p>\n<h1><strong>Finetune stage</strong></h1>\n<p>In the finetune stage, we train the model with the label. With pretraining, the converging speed for the model is very fast and we only need less than 5 epochs per fold in this stage!</p>\n<h1><strong>Model performance</strong></h1>\n<p>Our best Transformer model without meta features has <strong>0.794</strong> CV, <strong>0.796</strong> public LB, and <strong>0.804</strong> private LB.<br>\nWith pseudo labeling (We use soft predictions from our best ensemble model), a single transformer model can have <strong>0.800</strong> public LB and <strong>0.808</strong> private LB. (There is leaking because we didn't use Nested K-fold CV to generate pseudo label).</p>",
      "rawMarkdown": "URL to the Summary&zakopuro part: https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\n\nTo begin with, I would like to thank Amex for hosting the competition and my teammates ( @zakopur0 @scumufeng @baosenguo). Congrats to @zakopur0 and @scumufeng for being a Kaggle competition master!\n\n# **Pretrain + Finetune approach**\nLightGBM + Feature engineering is very successful in this competition, and they outperform NN + raw features most of the time. Because of this, I believe the features used in LGBM models are very powerful, so I decided to first let the NN learn how to do feature engineering in the pretrain stage, then finetune the model with the target after that.\nIn this way, we provide much more guidance to train the model by using thousands of features, and most importantly we can include test data in the pretrain stage.\nWith pretraining, the model can have **+0.002** boost in public LB compared with training from scratch, and we can train a larger model (6 layers transformer) without any problem.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2F044ba388c80c6653df48f7238b3a8698%2FAmex_transformer.png?generation=1661564543157828&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fe1857b02f2ae097535b695d6fce903a3%2FAmex_Features_encoder.png?generation=1661564557974134&alt=media)\n\n# **Model architecture**\nWe use different MLP layers to handle different types of inputs (Delinquency, Spend, Payment, Balance, Risk variables), then concatenate them and pass them to the transformer encoder. After the encoding part, we take the latest node and get the outputs through a linear layer.\n\n# **Pretrain stage**\nIn the pretrain stage, our target is tabular features. We use Huber loss to train the standardized target and won't pass any loss if the feature is nan. The number of epochs is around 200 in this stage.\n\n# **Finetune stage**\nIn the finetune stage, we train the model with the label. With pretraining, the converging speed for the model is very fast and we only need less than 5 epochs per fold in this stage!\n\n# **Model performance**\nOur best Transformer model without meta features has **0.794** CV, **0.796** public LB, and **0.804** private LB.\nWith pseudo labeling (We use soft predictions from our best ensemble model), a single transformer model can have **0.800** public LB and **0.808** private LB. (There is leaking because we didn't use Nested K-fold CV to generate pseudo label).",
      "votes": null
    },
    {
      "id": "1915508",
      "postDate": "08/27/2022 03:00:45",
      "content": "<p>Nice solution! Congrats <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> and team on results. One more gold to reaching GM 💪</p>",
      "rawMarkdown": "Nice solution! Congrats @wimwim and team on results. One more gold to reaching GM 💪",
      "votes": null
    },
    {
      "id": "1915514",
      "postDate": "08/27/2022 03:10:41",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> !</p>",
      "rawMarkdown": "Thank you @duykhanh99 !",
      "votes": null
    },
    {
      "id": "1915539",
      "postDate": "08/27/2022 04:13:55",
      "content": "<p>Thanks for sharing your brilliant solution!<br>\nI just want to ensure that I understand your flow:<br>\nInput to Features Encoder is <code>BSx13xM</code> (where <code>BS</code> is a batch size, 13 is a time dimension, <code>M</code> is a feature dimension)<br>\nOutput is <code>BSx13xX</code> where <code>X</code> is a embedding size of each customer's payment.<br>\nThen you pass it into Transformer Encoder which use an attention to adjust each payment's embedding according to other payments.<br>\nOutput of the transformer has the same size as input : <code>BSx13xX</code><br>\nThen you take only last payments's embedding <code>BSx1xX</code>  then squeeze to <code>BSxX</code><br>\nThen you use Linear Layer to transform <code>BSxX</code> into <code>BSxL</code> where <code>L</code> is a target size</p>\n<p>Is there any particular reason to input feature groups independently? Why just not input all features and output payments embedding of length <code>X</code>? Features from different groups are correlated. Sounds like removing their connections will force the model to work harder to extract the signal.</p>\n<p>Can you please elaborate on this?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Thanks for sharing your brilliant solution!\nI just want to ensure that I understand your flow:\nInput to Features Encoder is `BSx13xM` (where `BS` is a batch size, 13 is a time dimension, `M` is a feature dimension)\nOutput is `BSx13xX` where `X` is a embedding size of each customer's payment.\nThen you pass it into Transformer Encoder which use an attention to adjust each payment's embedding according to other payments.\nOutput of the transformer has the same size as input : `BSx13xX`\nThen you take only last payments's embedding `BSx1xX`  then squeeze to `BSxX`\nThen you use Linear Layer to transform `BSxX` into `BSxL` where `L` is a target size\n\nIs there any particular reason to input feature groups independently? Why just not input all features and output payments embedding of length `X`? Features from different groups are correlated. Sounds like removing their connections will force the model to work harder to extract the signal.\n\nCan you please elaborate on this?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1915560",
      "postDate": "08/27/2022 04:52:03",
      "content": "<p>Your understanding of my flow is correct!</p>\n<p>The reason that I input feature groups independently is that we cannot make full use of the fact that features can be divided into 5 categories if we input all features together. By using different MLP layers to process different feature groups, the model can learn the local information easier. After all, the model still has the chance to learn cross-group features because we will concatenate them after the feature encoder (payments embedding is able to store the information for a particular feature if it is important), so the connections will not be removed. I find this approach improves both my cv and lb.</p>",
      "rawMarkdown": "Your understanding of my flow is correct!\n\nThe reason that I input feature groups independently is that we cannot make full use of the fact that features can be divided into 5 categories if we input all features together. By using different MLP layers to process different feature groups, the model can learn the local information easier. After all, the model still has the chance to learn cross-group features because we will concatenate them after the feature encoder (payments embedding is able to store the information for a particular feature if it is important), so the connections will not be removed. I find this approach improves both my cv and lb.",
      "votes": null
    },
    {
      "id": "1915905",
      "postDate": "08/27/2022 13:11:42",
      "content": "<p>The idea  pretrain and finetune is so cool👍 Thanks for sharing!!<br>\nI wonder if the hidden dim，embedding dim important to model's performance? Do you spend much time on hidden dim, embedding dim?<br>\nI'd appreciate it if you could reply me😃Thanks!</p>",
      "rawMarkdown": "The idea  pretrain and finetune is so cool👍 Thanks for sharing!!\nI wonder if the hidden dim，embedding dim important to model's performance? Do you spend much time on hidden dim, embedding dim?\nI'd appreciate it if you could reply me😃Thanks!",
      "votes": null
    },
    {
      "id": "1915922",
      "postDate": "08/27/2022 13:30:30",
      "content": "<p>The model's performance is not sensitive to hyperparameters if we pretrain the model. I tried hidden dim = [512, 768, 1024], embedding dim = [8, 16], and the difference is small.</p>",
      "rawMarkdown": "The model's performance is not sensitive to hyperparameters if we pretrain the model. I tried hidden dim = [512, 768, 1024], embedding dim = [8, 16], and the difference is small.",
      "votes": null
    },
    {
      "id": "1921641",
      "postDate": "09/01/2022 01:10:50",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @wimwim May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    },
    {
      "id": "1924658",
      "postDate": "09/03/2022 09:19:14",
      "content": "<p>Congrats and thanks for sharing! A couple of questions from muy side: </p>\n<p><strong>1.</strong> Just to be sure, in the pretrain stage you predict all the lgbm features at once, right? The output layer of the transformer is <code>nn.Linear(in_features, #_of_lgbm_features)</code><br>\n<strong>2.</strong> Did you use the pseudolabels for the entire test data?<br>\n<strong>3.</strong> Was there any tricky part when pretraining/training the transformer (e.g. the learning rate, scheduler, etc)</p>",
      "rawMarkdown": "Congrats and thanks for sharing! A couple of questions from muy side: \n\n**1.** Just to be sure, in the pretrain stage you predict all the lgbm features at once, right? The output layer of the transformer is `nn.Linear(in_features, #_of_lgbm_features)`\n**2.** Did you use the pseudolabels for the entire test data?\n**3.** Was there any tricky part when pretraining/training the transformer (e.g. the learning rate, scheduler, etc)",
      "votes": null
    },
    {
      "id": "1926001",
      "postDate": "09/04/2022 13:36:08",
      "content": "<p>Nicely presented and straight to the point :)</p>\n<p>By any chance can you share he code section that generates the meta-features?</p>\n<p>Thanks </p>",
      "rawMarkdown": "Nicely presented and straight to the point :)\n\nBy any chance can you share he code section that generates the meta-features?\n\nThanks",
      "votes": null
    },
    {
      "id": "1926612",
      "postDate": "09/05/2022 01:53:49",
      "content": "<p>Thanks for your question and sorry for the late reply.</p>\n<ol>\n<li>Yes you are correct, the model will predict all the lgbm features at once, and I will replace the output layer with <code>nn.Linear(in_features, 1)</code> in the finetune stage.</li>\n<li>Yes I use the pseudo labels for the entire test data.</li>\n<li>I use a relatively high learning rate (0.01) in pretraining stage, and a low learning rate (0.001) in finetune stage. For the scheduler, I use CosineAnnealingLR in both pretraining and finetune stages.</li>\n</ol>",
      "rawMarkdown": "Thanks for your question and sorry for the late reply.\n\n1. Yes you are correct, the model will predict all the lgbm features at once, and I will replace the output layer with `nn.Linear(in_features, 1)` in the finetune stage.\n2. Yes I use the pseudo labels for the entire test data.\n3. I use a relatively high learning rate (0.01) in pretraining stage, and a low learning rate (0.001) in finetune stage. For the scheduler, I use CosineAnnealingLR in both pretraining and finetune stages.",
      "votes": null
    },
    {
      "id": "1926614",
      "postDate": "09/05/2022 01:55:28",
      "content": "<p>For the meta-features part, please refer to the Summary&amp;zakopuro part: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097</a><br>\nThanks</p>",
      "rawMarkdown": "For the meta-features part, please refer to the Summary&zakopuro part: https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\nThanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1915508,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "08/27/2022 03:00:45",
      "content": "<p>Nice solution! Congrats <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> and team on results. One more gold to reaching GM 💪</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915514,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "08/27/2022 03:10:41",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915539,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "08/27/2022 04:13:55",
      "content": "<p>Thanks for sharing your brilliant solution!<br>\nI just want to ensure that I understand your flow:<br>\nInput to Features Encoder is <code>BSx13xM</code> (where <code>BS</code> is a batch size, 13 is a time dimension, <code>M</code> is a feature dimension)<br>\nOutput is <code>BSx13xX</code> where <code>X</code> is a embedding size of each customer's payment.<br>\nThen you pass it into Transformer Encoder which use an attention to adjust each payment's embedding according to other payments.<br>\nOutput of the transformer has the same size as input : <code>BSx13xX</code><br>\nThen you take only last payments's embedding <code>BSx1xX</code>  then squeeze to <code>BSxX</code><br>\nThen you use Linear Layer to transform <code>BSxX</code> into <code>BSxL</code> where <code>L</code> is a target size</p>\n<p>Is there any particular reason to input feature groups independently? Why just not input all features and output payments embedding of length <code>X</code>? Features from different groups are correlated. Sounds like removing their connections will force the model to work harder to extract the signal.</p>\n<p>Can you please elaborate on this?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915560,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "08/27/2022 04:52:03",
          "content": "<p>Your understanding of my flow is correct!</p>\n<p>The reason that I input feature groups independently is that we cannot make full use of the fact that features can be divided into 5 categories if we input all features together. By using different MLP layers to process different feature groups, the model can learn the local information easier. After all, the model still has the chance to learn cross-group features because we will concatenate them after the feature encoder (payments embedding is able to store the information for a particular feature if it is important), so the connections will not be removed. I find this approach improves both my cv and lb.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915905,
      "author_name": "xiafire",
      "author_url": "",
      "post_date": "08/27/2022 13:11:42",
      "content": "<p>The idea  pretrain and finetune is so cool👍 Thanks for sharing!!<br>\nI wonder if the hidden dim，embedding dim important to model's performance? Do you spend much time on hidden dim, embedding dim?<br>\nI'd appreciate it if you could reply me😃Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915922,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "08/27/2022 13:30:30",
          "content": "<p>The model's performance is not sensitive to hyperparameters if we pretrain the model. I tried hidden dim = [512, 768, 1024], embedding dim = [8, 16], and the difference is small.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1921641,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 01:10:50",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1924658,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "09/03/2022 09:19:14",
      "content": "<p>Congrats and thanks for sharing! A couple of questions from muy side: </p>\n<p><strong>1.</strong> Just to be sure, in the pretrain stage you predict all the lgbm features at once, right? The output layer of the transformer is <code>nn.Linear(in_features, #_of_lgbm_features)</code><br>\n<strong>2.</strong> Did you use the pseudolabels for the entire test data?<br>\n<strong>3.</strong> Was there any tricky part when pretraining/training the transformer (e.g. the learning rate, scheduler, etc)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1926612,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "09/05/2022 01:53:49",
          "content": "<p>Thanks for your question and sorry for the late reply.</p>\n<ol>\n<li>Yes you are correct, the model will predict all the lgbm features at once, and I will replace the output layer with <code>nn.Linear(in_features, 1)</code> in the finetune stage.</li>\n<li>Yes I use the pseudo labels for the entire test data.</li>\n<li>I use a relatively high learning rate (0.01) in pretraining stage, and a low learning rate (0.001) in finetune stage. For the scheduler, I use CosineAnnealingLR in both pretraining and finetune stages.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1926001,
      "author_name": "jonimatix",
      "author_url": "",
      "post_date": "09/04/2022 13:36:08",
      "content": "<p>Nicely presented and straight to the point :)</p>\n<p>By any chance can you share he code section that generates the meta-features?</p>\n<p>Thanks </p>",
      "votes": null,
      "replies": [
        {
          "id": 1926614,
          "author_name": "wimwim",
          "author_url": "",
          "post_date": "09/05/2022 01:55:28",
          "content": "<p>For the meta-features part, please refer to the Summary&amp;zakopuro part: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097</a><br>\nThanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1915468": "URL to the Summary&zakopuro part: https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\n\nTo begin with, I would like to thank Amex for hosting the competition and my teammates ( @zakopur0 @scumufeng @baosenguo). Congrats to @zakopur0 and @scumufeng for being a Kaggle competition master!\n\n# **Pretrain + Finetune approach**\nLightGBM + Feature engineering is very successful in this competition, and they outperform NN + raw features most of the time. Because of this, I believe the features used in LGBM models are very powerful, so I decided to first let the NN learn how to do feature engineering in the pretrain stage, then finetune the model with the target after that.\nIn this way, we provide much more guidance to train the model by using thousands of features, and most importantly we can include test data in the pretrain stage.\nWith pretraining, the model can have **+0.002** boost in public LB compared with training from scratch, and we can train a larger model (6 layers transformer) without any problem.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2F044ba388c80c6653df48f7238b3a8698%2FAmex_transformer.png?generation=1661564543157828&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1904629%2Fe1857b02f2ae097535b695d6fce903a3%2FAmex_Features_encoder.png?generation=1661564557974134&alt=media)\n\n# **Model architecture**\nWe use different MLP layers to handle different types of inputs (Delinquency, Spend, Payment, Balance, Risk variables), then concatenate them and pass them to the transformer encoder. After the encoding part, we take the latest node and get the outputs through a linear layer.\n\n# **Pretrain stage**\nIn the pretrain stage, our target is tabular features. We use Huber loss to train the standardized target and won't pass any loss if the feature is nan. The number of epochs is around 200 in this stage.\n\n# **Finetune stage**\nIn the finetune stage, we train the model with the label. With pretraining, the converging speed for the model is very fast and we only need less than 5 epochs per fold in this stage!\n\n# **Model performance**\nOur best Transformer model without meta features has **0.794** CV, **0.796** public LB, and **0.804** private LB.\nWith pseudo labeling (We use soft predictions from our best ensemble model), a single transformer model can have **0.800** public LB and **0.808** private LB. (There is leaking because we didn't use Nested K-fold CV to generate pseudo label).",
    "1915508": "Nice solution! Congrats @wimwim and team on results. One more gold to reaching GM 💪",
    "1915514": "Thank you @duykhanh99 !",
    "1915539": "Thanks for sharing your brilliant solution!\nI just want to ensure that I understand your flow:\nInput to Features Encoder is `BSx13xM` (where `BS` is a batch size, 13 is a time dimension, `M` is a feature dimension)\nOutput is `BSx13xX` where `X` is a embedding size of each customer's payment.\nThen you pass it into Transformer Encoder which use an attention to adjust each payment's embedding according to other payments.\nOutput of the transformer has the same size as input : `BSx13xX`\nThen you take only last payments's embedding `BSx1xX`  then squeeze to `BSxX`\nThen you use Linear Layer to transform `BSxX` into `BSxL` where `L` is a target size\n\nIs there any particular reason to input feature groups independently? Why just not input all features and output payments embedding of length `X`? Features from different groups are correlated. Sounds like removing their connections will force the model to work harder to extract the signal.\n\nCan you please elaborate on this?\n\nThanks!",
    "1915560": "Your understanding of my flow is correct!\n\nThe reason that I input feature groups independently is that we cannot make full use of the fact that features can be divided into 5 categories if we input all features together. By using different MLP layers to process different feature groups, the model can learn the local information easier. After all, the model still has the chance to learn cross-group features because we will concatenate them after the feature encoder (payments embedding is able to store the information for a particular feature if it is important), so the connections will not be removed. I find this approach improves both my cv and lb.",
    "1915905": "The idea  pretrain and finetune is so cool👍 Thanks for sharing!!\nI wonder if the hidden dim，embedding dim important to model's performance? Do you spend much time on hidden dim, embedding dim?\nI'd appreciate it if you could reply me😃Thanks!",
    "1915922": "The model's performance is not sensitive to hyperparameters if we pretrain the model. I tried hidden dim = [512, 768, 1024], embedding dim = [8, 16], and the difference is small.",
    "1921641": "Hi @wimwim May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1924658": "Congrats and thanks for sharing! A couple of questions from muy side: \n\n**1.** Just to be sure, in the pretrain stage you predict all the lgbm features at once, right? The output layer of the transformer is `nn.Linear(in_features, #_of_lgbm_features)`\n**2.** Did you use the pseudolabels for the entire test data?\n**3.** Was there any tricky part when pretraining/training the transformer (e.g. the learning rate, scheduler, etc)",
    "1926001": "Nicely presented and straight to the point :)\n\nBy any chance can you share he code section that generates the meta-features?\n\nThanks",
    "1926612": "Thanks for your question and sorry for the late reply.\n\n1. Yes you are correct, the model will predict all the lgbm features at once, and I will replace the output layer with `nn.Linear(in_features, 1)` in the finetune stage.\n2. Yes I use the pseudo labels for the entire test data.\n3. I use a relatively high learning rate (0.01) in pretraining stage, and a low learning rate (0.001) in finetune stage. For the scheduler, I use CosineAnnealingLR in both pretraining and finetune stages.",
    "1926614": "For the meta-features part, please refer to the Summary&zakopuro part: https://www.kaggle.com/competitions/amex-default-prediction/discussion/348097\nThanks"
  },
  "source": "meta"
}