{
  "id": 348097,
  "title": "5th Place Solution - Team 💳VISA💳(Summary&zakopuro's part)",
  "url": "/competitions/amex-default-prediction/writeups/visa-5th-place-solution-team-visa-summary-zakopuro",
  "author_name": "",
  "post_date": "2022-08-27T02:25:36.200Z",
  "votes": 61,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I would like to thank the organizers for organizing this competition, the participants for sharing their many insights, and my teammates( <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> ). Now <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> and I are promoted to kaggle master and <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> is reaching for GM.<br>\nThis competition will be unforgettable for me:)</p>\n<h1>Summary</h1>\n<p>Our solution is the result of ensembling several GBDT models , Transfomr, 2d-CNN, and GRU.<br>\nWe noticed that ensemble weights are determined based on Public LB and overfit if based on CV.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F868d637a78b88773048c6f2c9b615402%2F2022-08-27%20063723.png?generation=1661549864910264&amp;alt=media\" alt=\"\"></p>\n<h1>Features</h1>\n<p>We are using the dataset shared with us by raddar. The features are based on those shared by ragnar.(Thanks to both of you.)</p>\n<h3>meta feature</h3>\n<p>I did not know this is called a meta feature.<br>\nThis feature was useful not only in GBDT, but also in Transformer.<br>\nIf added to <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790l\" target=\"_blank\">chris's Transformer</a>, the LB will increase from 0.790 to 793.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F9bc12bb44eee04b66decedd22ae185f1%2F2022-08-27%20055050.png?generation=1661547064786565&amp;alt=media\" alt=\"\"></p>\n<h3>Pivot</h3>\n<p>Combine all features horizontally.</p>\n<pre><code>P_2_0 , P_2_1 , P_2_3 , P_2_4 , ... , P_2_12 , B_30_0 , ... \n XXX  ,  XXX  ,  XXX  ,  XXX  , ... ,   XXX  ,   YYY  , ...\n</code></pre>\n<h1>Model</h1>\n<h3>GBDT</h3>\n<ul>\n<li>LightGBM<ul>\n<li>Almost no change from <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">this Notebook</a></li>\n<li>stratfiedKfold: 10</li></ul></li>\n<li>Catboost<ul>\n<li>Use GPU(I was surprised at how fast it was.)</li>\n<li>parameter : default</li></ul></li>\n</ul>\n<h3>Transformer</h3>\n<h4>zakopuro</h4>\n<ul>\n<li>Based on <a target=\"_blank\">chris's Notebook</a></li>\n<li>Some additional features.(Mainly meta features)</li>\n</ul>\n<h4>Patrick Yam</h4>\n<p>This is his solution.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118</a></p>\n<h4>mufeng</h4>\n<ul>\n<li>Add meta feature to Patrick's transformer.<ul>\n<li>LB increased from 0.795 to 0.798.</li></ul></li>\n</ul>\n<h1>Ensemble</h1>\n<ul>\n<li>Use 21 models.</li>\n<li>Ensemble weight<ul>\n<li>Determined based on Public LB</li>\n<li>In our case, the LB score will be lower if based on CV.(CV is 0.8016 or higher.)</li>\n<li>We trusted Public LB more than CV because it is close to Private and has a sufficient amount of data.</li></ul></li>\n<li>Weights are not complicated. (For example, 0.1,0.2,… etc.)</li>\n</ul>\n<h1>Select Submit</h1>\n<ul>\n<li>Best LB<ul>\n<li>Public : 0.80199(2nd)</li>\n<li>Priavte : 0.80881(6th)</li></ul></li>\n<li>Best LB*0.5 + Best CV *0.5<ul>\n<li>Public : 0.80154</li>\n<li>Private : 0.80862(Gold zone)</li></ul></li>\n<li>Correlation check<ul>\n<li>Check the Public and Private correlation values for all predictions used in the ensemble to see that there are no significant differences.</li></ul></li>\n</ul>\n<p>All posts above 0.801 in Public LB were in the Gold zone in Private LB.(I prayed on the last day not to Shake down😣)</p>\n<p>Let's enjoy kaggle! Thank you!!!!</p>",
  "messages": [
    {
      "id": "1915338",
      "postDate": "08/26/2022 22:29:13",
      "content": "<p>I would like to thank the organizers for organizing this competition, the participants for sharing their many insights, and my teammates( <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> ). Now <a href=\"https://www.kaggle.com/scumufeng\" target=\"_blank\">@scumufeng</a> and I are promoted to kaggle master and <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> is reaching for GM.<br>\nThis competition will be unforgettable for me:)</p>\n<h1>Summary</h1>\n<p>Our solution is the result of ensembling several GBDT models , Transfomr, 2d-CNN, and GRU.<br>\nWe noticed that ensemble weights are determined based on Public LB and overfit if based on CV.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F868d637a78b88773048c6f2c9b615402%2F2022-08-27%20063723.png?generation=1661549864910264&amp;alt=media\" alt=\"\"></p>\n<h1>Features</h1>\n<p>We are using the dataset shared with us by raddar. The features are based on those shared by ragnar.(Thanks to both of you.)</p>\n<h3>meta feature</h3>\n<p>I did not know this is called a meta feature.<br>\nThis feature was useful not only in GBDT, but also in Transformer.<br>\nIf added to <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790l\" target=\"_blank\">chris's Transformer</a>, the LB will increase from 0.790 to 793.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F9bc12bb44eee04b66decedd22ae185f1%2F2022-08-27%20055050.png?generation=1661547064786565&amp;alt=media\" alt=\"\"></p>\n<h3>Pivot</h3>\n<p>Combine all features horizontally.</p>\n<pre><code>P_2_0 , P_2_1 , P_2_3 , P_2_4 , ... , P_2_12 , B_30_0 , ... \n XXX  ,  XXX  ,  XXX  ,  XXX  , ... ,   XXX  ,   YYY  , ...\n</code></pre>\n<h1>Model</h1>\n<h3>GBDT</h3>\n<ul>\n<li>LightGBM<ul>\n<li>Almost no change from <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">this Notebook</a></li>\n<li>stratfiedKfold: 10</li></ul></li>\n<li>Catboost<ul>\n<li>Use GPU(I was surprised at how fast it was.)</li>\n<li>parameter : default</li></ul></li>\n</ul>\n<h3>Transformer</h3>\n<h4>zakopuro</h4>\n<ul>\n<li>Based on <a target=\"_blank\">chris's Notebook</a></li>\n<li>Some additional features.(Mainly meta features)</li>\n</ul>\n<h4>Patrick Yam</h4>\n<p>This is his solution.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118</a></p>\n<h4>mufeng</h4>\n<ul>\n<li>Add meta feature to Patrick's transformer.<ul>\n<li>LB increased from 0.795 to 0.798.</li></ul></li>\n</ul>\n<h1>Ensemble</h1>\n<ul>\n<li>Use 21 models.</li>\n<li>Ensemble weight<ul>\n<li>Determined based on Public LB</li>\n<li>In our case, the LB score will be lower if based on CV.(CV is 0.8016 or higher.)</li>\n<li>We trusted Public LB more than CV because it is close to Private and has a sufficient amount of data.</li></ul></li>\n<li>Weights are not complicated. (For example, 0.1,0.2,… etc.)</li>\n</ul>\n<h1>Select Submit</h1>\n<ul>\n<li>Best LB<ul>\n<li>Public : 0.80199(2nd)</li>\n<li>Priavte : 0.80881(6th)</li></ul></li>\n<li>Best LB*0.5 + Best CV *0.5<ul>\n<li>Public : 0.80154</li>\n<li>Private : 0.80862(Gold zone)</li></ul></li>\n<li>Correlation check<ul>\n<li>Check the Public and Private correlation values for all predictions used in the ensemble to see that there are no significant differences.</li></ul></li>\n</ul>\n<p>All posts above 0.801 in Public LB were in the Gold zone in Private LB.(I prayed on the last day not to Shake down😣)</p>\n<p>Let's enjoy kaggle! Thank you!!!!</p>",
      "rawMarkdown": "I would like to thank the organizers for organizing this competition, the participants for sharing their many insights, and my teammates( @baosenguo @wimwim @scumufeng ). Now @scumufeng and I are promoted to kaggle master and @wimwim is reaching for GM.\nThis competition will be unforgettable for me:)\n\n# Summary\nOur solution is the result of ensembling several GBDT models , Transfomr, 2d-CNN, and GRU.\nWe noticed that ensemble weights are determined based on Public LB and overfit if based on CV.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F868d637a78b88773048c6f2c9b615402%2F2022-08-27%20063723.png?generation=1661549864910264&alt=media)\n\n\n\n# Features\nWe are using the dataset shared with us by raddar. The features are based on those shared by ragnar.(Thanks to both of you.)\n\n### meta feature\nI did not know this is called a meta feature.\nThis feature was useful not only in GBDT, but also in Transformer.\nIf added to [chris's Transformer](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790l), the LB will increase from 0.790 to 793.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F9bc12bb44eee04b66decedd22ae185f1%2F2022-08-27%20055050.png?generation=1661547064786565&alt=media)\n\n### Pivot\nCombine all features horizontally.\n\n```\nP_2_0 , P_2_1 , P_2_3 , P_2_4 , ... , P_2_12 , B_30_0 , ... \n XXX  ,  XXX  ,  XXX  ,  XXX  , ... ,   XXX  ,   YYY  , ...\n```\n\n# Model\n### GBDT\n- LightGBM\n - Almost no change from [this Notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977)\n - stratfiedKfold: 10\n- Catboost\n - Use GPU(I was surprised at how fast it was.)\n - parameter : default\n\n### Transformer\n#### zakopuro\n- Based on [chris's Notebook]((https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790l)\n- Some additional features.(Mainly meta features)\n\n#### Patrick Yam\nThis is his solution.\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/348118\n\n#### mufeng\n- Add meta feature to Patrick's transformer.\n - LB increased from 0.795 to 0.798.\n\n# Ensemble\n- Use 21 models.\n- Ensemble weight\n - Determined based on Public LB\n - In our case, the LB score will be lower if based on CV.(CV is 0.8016 or higher.)\n - We trusted Public LB more than CV because it is close to Private and has a sufficient amount of data.\n- Weights are not complicated. (For example, 0.1,0.2,... etc.)\n\n# Select Submit\n- Best LB\n - Public : 0.80199(2nd)\n - Priavte : 0.80881(6th)\n- Best LB*0.5 + Best CV *0.5\n - Public : 0.80154\n - Private : 0.80862(Gold zone)\n- Correlation check\n - Check the Public and Private correlation values for all predictions used in the ensemble to see that there are no significant differences.\n\nAll posts above 0.801 in Public LB were in the Gold zone in Private LB.(I prayed on the last day not to Shake down😣)\n\nLet's enjoy kaggle! Thank you!!!!",
      "votes": null
    },
    {
      "id": "1915343",
      "postDate": "08/26/2022 22:38:32",
      "content": "<p>Congratulations team. Great solution. I didn't know about these \"meta features\" but i see that many top teams used them. I will use them in my future competitions. Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congratulations team. Great solution. I didn't know about these \"meta features\" but i see that many top teams used them. I will use them in my future competitions. Thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "1915353",
      "postDate": "08/26/2022 22:54:53",
      "content": "<p>Thank you very much.<br>\nI have learned a lot from you in this competition as well!</p>",
      "rawMarkdown": "Thank you very much.\nI have learned a lot from you in this competition as well!",
      "votes": null
    },
    {
      "id": "1915359",
      "postDate": "08/26/2022 23:00:22",
      "content": "<p>In your diagram, i see that you pretrain Transformer then finetune. What data did you pretrain with and what data did you finetune with?</p>",
      "rawMarkdown": "In your diagram, i see that you pretrain Transformer then finetune. What data did you pretrain with and what data did you finetune with?",
      "votes": null
    },
    {
      "id": "1915362",
      "postDate": "08/26/2022 23:04:55",
      "content": "<p>Patrick will write for more details.<br>\nMy understanding is that the pretrain for GBDT features and finetune using target.</p>",
      "rawMarkdown": "Patrick will write for more details.\nMy understanding is that the pretrain for GBDT features and finetune using target.",
      "votes": null
    },
    {
      "id": "1915470",
      "postDate": "08/27/2022 02:14:26",
      "content": "<p>This is his solution.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118</a></p>",
      "rawMarkdown": "This is his solution.\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/348118",
      "votes": null
    },
    {
      "id": "1916004",
      "postDate": "08/27/2022 15:09:42",
      "content": "<p>Congrats on the gold medal. Nice solution and nice write-up. The \"meta-features\" approach is very interesting. If I understand correctly you did a number of rounds of adding these meta-features -- taking the average of the previous ones and keeping the last one intact. After how many rounds did CV/LB improvements stagnate? </p>",
      "rawMarkdown": "Congrats on the gold medal. Nice solution and nice write-up. The \"meta-features\" approach is very interesting. If I understand correctly you did a number of rounds of adding these meta-features -- taking the average of the previous ones and keeping the last one intact. After how many rounds did CV/LB improvements stagnate?",
      "votes": null
    },
    {
      "id": "1916545",
      "postDate": "08/28/2022 01:28:31",
      "content": "<p>Thank you.<br>\nFour \"meta-features\" were used. However, in the case of GBDT, the use of \"meta-features\" does not improve the score of the single model much. The model using this feature is more effective for ensembles.<br>\nIn Transformer's case, increasing the number by more than four had no effect.</p>",
      "rawMarkdown": "Thank you.\nFour \"meta-features\" were used. However, in the case of GBDT, the use of \"meta-features\" does not improve the score of the single model much. The model using this feature is more effective for ensembles.\nIn Transformer's case, increasing the number by more than four had no effect.",
      "votes": null
    },
    {
      "id": "1918329",
      "postDate": "08/29/2022 13:57:53",
      "content": "<p>Congratulations. I found very interesting the process of ensembling that wide variety of models and it is a good reference for future competitions. Thank you so much for this answer.</p>",
      "rawMarkdown": "Congratulations. I found very interesting the process of ensembling that wide variety of models and it is a good reference for future competitions. Thank you so much for this answer.",
      "votes": null
    },
    {
      "id": "1924901",
      "postDate": "09/03/2022 14:01:36",
      "content": "<p>Congrats for all of you that were promoted and thanks for sharing! I would like to ask a couple of questions:</p>\n<p><strong>1.</strong> The \"pivot\" part is only for the Transformers, correct?<br>\n<strong>2.</strong> From picture I interpret that each team member developed his own features, how different are from Martin's features?<br>\n<strong>3.</strong> I'm curious about the 2d-CNN model, do you have more info?</p>",
      "rawMarkdown": "Congrats for all of you that were promoted and thanks for sharing! I would like to ask a couple of questions:\n\n**1.** The \"pivot\" part is only for the Transformers, correct?\n**2.** From picture I interpret that each team member developed his own features, how different are from Martin's features?\n**3.** I'm curious about the 2d-CNN model, do you have more info?",
      "votes": null
    },
    {
      "id": "1925530",
      "postDate": "09/04/2022 04:14:51",
      "content": "<p>Thank you!</p>\n<ol>\n<li><p>The \"pivot\" part is only for the Transformers, correct?<br>\n-&gt; This is a feature for the GBDT, not for the Transformer.</p></li>\n<li><p>From picture I interpret that each team member developed his own features, how different are from Martin's features?<br>\n-&gt; There are no major differences; diff features are added, rounding with float, etc.</p></li>\n<li><p>I'm curious about the 2d-CNN model, do you have more info?<br>\n-&gt; It is similar to this content.(<a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256#1106810\" target=\"_blank\">https://www.kaggle.com/competitions/lish-moa/discussion/202256#1106810</a>)</p></li>\n</ol>",
      "rawMarkdown": "Thank you!\n\n1. The \"pivot\" part is only for the Transformers, correct?\n-> This is a feature for the GBDT, not for the Transformer.\n\n2. From picture I interpret that each team member developed his own features, how different are from Martin's features?\n-> There are no major differences; diff features are added, rounding with float, etc.\n\n3. I'm curious about the 2d-CNN model, do you have more info?\n-> It is similar to this content.(https://www.kaggle.com/competitions/lish-moa/discussion/202256#1106810)",
      "votes": null
    },
    {
      "id": "1925746",
      "postDate": "09/04/2022 09:13:33",
      "content": "<p>Thanks for the answers! Then the GBDT used pivoted features (i.e. the 13 pivoted statements) instead of aggregations of them? Or maybe It used both?</p>",
      "rawMarkdown": "Thanks for the answers! Then the GBDT used pivoted features (i.e. the 13 pivoted statements) instead of aggregations of them? Or maybe It used both?",
      "votes": null
    },
    {
      "id": "1925755",
      "postDate": "09/04/2022 09:19:57",
      "content": "<p>It used both.</p>",
      "rawMarkdown": "It used both.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1915343,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/26/2022 22:38:32",
      "content": "<p>Congratulations team. Great solution. I didn't know about these \"meta features\" but i see that many top teams used them. I will use them in my future competitions. Thanks for sharing your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915353,
          "author_name": "zakopur0",
          "author_url": "",
          "post_date": "08/26/2022 22:54:53",
          "content": "<p>Thank you very much.<br>\nI have learned a lot from you in this competition as well!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915359,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/26/2022 23:00:22",
          "content": "<p>In your diagram, i see that you pretrain Transformer then finetune. What data did you pretrain with and what data did you finetune with?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915362,
          "author_name": "zakopur0",
          "author_url": "",
          "post_date": "08/26/2022 23:04:55",
          "content": "<p>Patrick will write for more details.<br>\nMy understanding is that the pretrain for GBDT features and finetune using target.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915470,
          "author_name": "zakopur0",
          "author_url": "",
          "post_date": "08/27/2022 02:14:26",
          "content": "<p>This is his solution.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348118</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1916004,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "08/27/2022 15:09:42",
      "content": "<p>Congrats on the gold medal. Nice solution and nice write-up. The \"meta-features\" approach is very interesting. If I understand correctly you did a number of rounds of adding these meta-features -- taking the average of the previous ones and keeping the last one intact. After how many rounds did CV/LB improvements stagnate? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1916545,
          "author_name": "zakopur0",
          "author_url": "",
          "post_date": "08/28/2022 01:28:31",
          "content": "<p>Thank you.<br>\nFour \"meta-features\" were used. However, in the case of GBDT, the use of \"meta-features\" does not improve the score of the single model much. The model using this feature is more effective for ensembles.<br>\nIn Transformer's case, increasing the number by more than four had no effect.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1918329,
      "author_name": "kevinmorgado",
      "author_url": "",
      "post_date": "08/29/2022 13:57:53",
      "content": "<p>Congratulations. I found very interesting the process of ensembling that wide variety of models and it is a good reference for future competitions. Thank you so much for this answer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1924901,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "09/03/2022 14:01:36",
      "content": "<p>Congrats for all of you that were promoted and thanks for sharing! I would like to ask a couple of questions:</p>\n<p><strong>1.</strong> The \"pivot\" part is only for the Transformers, correct?<br>\n<strong>2.</strong> From picture I interpret that each team member developed his own features, how different are from Martin's features?<br>\n<strong>3.</strong> I'm curious about the 2d-CNN model, do you have more info?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1925530,
          "author_name": "zakopur0",
          "author_url": "",
          "post_date": "09/04/2022 04:14:51",
          "content": "<p>Thank you!</p>\n<ol>\n<li><p>The \"pivot\" part is only for the Transformers, correct?<br>\n-&gt; This is a feature for the GBDT, not for the Transformer.</p></li>\n<li><p>From picture I interpret that each team member developed his own features, how different are from Martin's features?<br>\n-&gt; There are no major differences; diff features are added, rounding with float, etc.</p></li>\n<li><p>I'm curious about the 2d-CNN model, do you have more info?<br>\n-&gt; It is similar to this content.(<a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256#1106810\" target=\"_blank\">https://www.kaggle.com/competitions/lish-moa/discussion/202256#1106810</a>)</p></li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 1925746,
              "author_name": "delai50",
              "author_url": "",
              "post_date": "09/04/2022 09:13:33",
              "content": "<p>Thanks for the answers! Then the GBDT used pivoted features (i.e. the 13 pivoted statements) instead of aggregations of them? Or maybe It used both?</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1925755,
          "author_name": "zakopur0",
          "author_url": "",
          "post_date": "09/04/2022 09:19:57",
          "content": "<p>It used both.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1915338": "I would like to thank the organizers for organizing this competition, the participants for sharing their many insights, and my teammates( @baosenguo @wimwim @scumufeng ). Now @scumufeng and I are promoted to kaggle master and @wimwim is reaching for GM.\nThis competition will be unforgettable for me:)\n\n# Summary\nOur solution is the result of ensembling several GBDT models , Transfomr, 2d-CNN, and GRU.\nWe noticed that ensemble weights are determined based on Public LB and overfit if based on CV.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F868d637a78b88773048c6f2c9b615402%2F2022-08-27%20063723.png?generation=1661549864910264&alt=media)\n\n\n\n# Features\nWe are using the dataset shared with us by raddar. The features are based on those shared by ragnar.(Thanks to both of you.)\n\n### meta feature\nI did not know this is called a meta feature.\nThis feature was useful not only in GBDT, but also in Transformer.\nIf added to [chris's Transformer](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790l), the LB will increase from 0.790 to 793.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F9bc12bb44eee04b66decedd22ae185f1%2F2022-08-27%20055050.png?generation=1661547064786565&alt=media)\n\n### Pivot\nCombine all features horizontally.\n\n```\nP_2_0 , P_2_1 , P_2_3 , P_2_4 , ... , P_2_12 , B_30_0 , ... \n XXX  ,  XXX  ,  XXX  ,  XXX  , ... ,   XXX  ,   YYY  , ...\n```\n\n# Model\n### GBDT\n- LightGBM\n - Almost no change from [this Notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977)\n - stratfiedKfold: 10\n- Catboost\n - Use GPU(I was surprised at how fast it was.)\n - parameter : default\n\n### Transformer\n#### zakopuro\n- Based on [chris's Notebook]((https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790l)\n- Some additional features.(Mainly meta features)\n\n#### Patrick Yam\nThis is his solution.\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/348118\n\n#### mufeng\n- Add meta feature to Patrick's transformer.\n - LB increased from 0.795 to 0.798.\n\n# Ensemble\n- Use 21 models.\n- Ensemble weight\n - Determined based on Public LB\n - In our case, the LB score will be lower if based on CV.(CV is 0.8016 or higher.)\n - We trusted Public LB more than CV because it is close to Private and has a sufficient amount of data.\n- Weights are not complicated. (For example, 0.1,0.2,... etc.)\n\n# Select Submit\n- Best LB\n - Public : 0.80199(2nd)\n - Priavte : 0.80881(6th)\n- Best LB*0.5 + Best CV *0.5\n - Public : 0.80154\n - Private : 0.80862(Gold zone)\n- Correlation check\n - Check the Public and Private correlation values for all predictions used in the ensemble to see that there are no significant differences.\n\nAll posts above 0.801 in Public LB were in the Gold zone in Private LB.(I prayed on the last day not to Shake down😣)\n\nLet's enjoy kaggle! Thank you!!!!",
    "1915343": "Congratulations team. Great solution. I didn't know about these \"meta features\" but i see that many top teams used them. I will use them in my future competitions. Thanks for sharing your solution.",
    "1915353": "Thank you very much.\nI have learned a lot from you in this competition as well!",
    "1915359": "In your diagram, i see that you pretrain Transformer then finetune. What data did you pretrain with and what data did you finetune with?",
    "1915362": "Patrick will write for more details.\nMy understanding is that the pretrain for GBDT features and finetune using target.",
    "1915470": "This is his solution.\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/348118",
    "1916004": "Congrats on the gold medal. Nice solution and nice write-up. The \"meta-features\" approach is very interesting. If I understand correctly you did a number of rounds of adding these meta-features -- taking the average of the previous ones and keeping the last one intact. After how many rounds did CV/LB improvements stagnate?",
    "1916545": "Thank you.\nFour \"meta-features\" were used. However, in the case of GBDT, the use of \"meta-features\" does not improve the score of the single model much. The model using this feature is more effective for ensembles.\nIn Transformer's case, increasing the number by more than four had no effect.",
    "1918329": "Congratulations. I found very interesting the process of ensembling that wide variety of models and it is a good reference for future competitions. Thank you so much for this answer.",
    "1924901": "Congrats for all of you that were promoted and thanks for sharing! I would like to ask a couple of questions:\n\n**1.** The \"pivot\" part is only for the Transformers, correct?\n**2.** From picture I interpret that each team member developed his own features, how different are from Martin's features?\n**3.** I'm curious about the 2d-CNN model, do you have more info?",
    "1925530": "Thank you!\n\n1. The \"pivot\" part is only for the Transformers, correct?\n-> This is a feature for the GBDT, not for the Transformer.\n\n2. From picture I interpret that each team member developed his own features, how different are from Martin's features?\n-> There are no major differences; diff features are added, rounding with float, etc.\n\n3. I'm curious about the 2d-CNN model, do you have more info?\n-> It is similar to this content.(https://www.kaggle.com/competitions/lish-moa/discussion/202256#1106810)",
    "1925746": "Thanks for the answers! Then the GBDT used pivoted features (i.e. the 13 pivoted statements) instead of aggregations of them? Or maybe It used both?",
    "1925755": "It used both."
  },
  "source": "meta"
}