{
  "id": 347740,
  "title": "12th Place Gold (1/2) - lgbm+xgboost+FCN+Transformer",
  "url": "/competitions/amex-default-prediction/writeups/journey-to-801-12th-place-gold-1-2-lgbm-xgboost-fc",
  "author_name": "",
  "post_date": "2022-08-31T07:39:05.093Z",
  "votes": 43,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Thank you Amex and Kaggle for hosting! Very fun competition with a lot of insights. </p>\n<h2>Solution</h2>\n<p>This is the first of two posts of our ensemble <a href=\"https://www.kaggle.com/gandagorn\" target=\"_blank\">@gandagorn</a>  .  My leg of the ensemble consisted of taking different preprocessing methods and passing them through different architectures (Thank you <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>  ! for your excellent notebooks) <br>\nsecond part  is <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348058</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F7a02721f85aa7042adf4202cb9e7c624%2Fensemble.png?generation=1661418572999162&amp;alt=media\" alt=\"\"></p>\n<h2>Preprocessing</h2>\n<ol>\n<li>Used the <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>  integer features  and raw features and then created  separate pre-processing pipelines including standard aggregations (min, max means, last, diff) </li>\n<li>Added top 50 PCA Features. </li>\n<li>Globally ranked all the features and built products of features that shared the same trend. This boosted all the models. </li>\n<li>Used embedded layers for all categorical features. </li>\n</ol>\n<h2>PostProcessing</h2>\n<p>Sequentially removed features accordingly to split importance. </p>\n<h2>Models</h2>\n<h3>LGBM with Focal Loss + logloss</h3>\n<p>Standard lgbm trained with dart with one-sided focal loss added to the log-loss. The addition of the one-side focal loss was with the objective of having a better classification of the positive classes.</p>\n<h3>XGBoost</h3>\n<p>Nothing special.</p>\n<h3>FCN</h3>\n<p>Fully connected architecture using the same features that were used in the boosted trees. 20 folds CV.</p>\n<h3>FCN + Transformer Encoder</h3>\n<p>Probably the most interesting model.  A NN with double inputs: tabular features (from previous preprocessing) and full-time series of ranked raw features. The training process was as follows:</p>\n<ol>\n<li>Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]</li>\n<li>Freeze FCN weights and train the Transformer for ~30 epochs to produce 1d output [0,1]</li>\n<li>Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F31f1f8ded2cba706ad095ed7595847fc%2Farchitecture.png?generation=1661419863567301&amp;alt=media\" alt=\"\"></p>\n<p>For the Transformer, I used the standard TransformerEncoderBlock with 8 heads and 2 layers</p>",
  "messages": [
    {
      "id": "1913348",
      "postDate": "08/25/2022 09:35:37",
      "content": "<p>Thank you Amex and Kaggle for hosting! Very fun competition with a lot of insights. </p>\n<h2>Solution</h2>\n<p>This is the first of two posts of our ensemble <a href=\"https://www.kaggle.com/gandagorn\" target=\"_blank\">@gandagorn</a>  .  My leg of the ensemble consisted of taking different preprocessing methods and passing them through different architectures (Thank you <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>  ! for your excellent notebooks) <br>\nsecond part  is <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/348058</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F7a02721f85aa7042adf4202cb9e7c624%2Fensemble.png?generation=1661418572999162&amp;alt=media\" alt=\"\"></p>\n<h2>Preprocessing</h2>\n<ol>\n<li>Used the <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>  integer features  and raw features and then created  separate pre-processing pipelines including standard aggregations (min, max means, last, diff) </li>\n<li>Added top 50 PCA Features. </li>\n<li>Globally ranked all the features and built products of features that shared the same trend. This boosted all the models. </li>\n<li>Used embedded layers for all categorical features. </li>\n</ol>\n<h2>PostProcessing</h2>\n<p>Sequentially removed features accordingly to split importance. </p>\n<h2>Models</h2>\n<h3>LGBM with Focal Loss + logloss</h3>\n<p>Standard lgbm trained with dart with one-sided focal loss added to the log-loss. The addition of the one-side focal loss was with the objective of having a better classification of the positive classes.</p>\n<h3>XGBoost</h3>\n<p>Nothing special.</p>\n<h3>FCN</h3>\n<p>Fully connected architecture using the same features that were used in the boosted trees. 20 folds CV.</p>\n<h3>FCN + Transformer Encoder</h3>\n<p>Probably the most interesting model.  A NN with double inputs: tabular features (from previous preprocessing) and full-time series of ranked raw features. The training process was as follows:</p>\n<ol>\n<li>Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]</li>\n<li>Freeze FCN weights and train the Transformer for ~30 epochs to produce 1d output [0,1]</li>\n<li>Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F31f1f8ded2cba706ad095ed7595847fc%2Farchitecture.png?generation=1661419863567301&amp;alt=media\" alt=\"\"></p>\n<p>For the Transformer, I used the standard TransformerEncoderBlock with 8 heads and 2 layers</p>",
      "rawMarkdown": "Thank you Amex and Kaggle for hosting! Very fun competition with a lot of insights. \n\n## Solution\n\nThis is the first of two posts of our ensemble @gandagorn  .  My leg of the ensemble consisted of taking different preprocessing methods and passing them through different architectures (Thank you @jiweiliu @lucasmorin @raddar  ! for your excellent notebooks) \nsecond part  is [https://www.kaggle.com/competitions/amex-default-prediction/discussion/348058](url)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F7a02721f85aa7042adf4202cb9e7c624%2Fensemble.png?generation=1661418572999162&alt=media)\n\n\n## Preprocessing\n\n1. Used the @raddar  integer features  and raw features and then created  separate pre-processing pipelines including standard aggregations (min, max means, last, diff) \n2. Added top 50 PCA Features. \n3. Globally ranked all the features and built products of features that shared the same trend. This boosted all the models. \n4. Used embedded layers for all categorical features. \n\n## PostProcessing\n\nSequentially removed features accordingly to split importance. \n\n## Models\n\n### LGBM with Focal Loss + logloss\n\nStandard lgbm trained with dart with one-sided focal loss added to the log-loss. The addition of the one-side focal loss was with the objective of having a better classification of the positive classes.\n\n### XGBoost\n\nNothing special.\n\n### FCN\n\nFully connected architecture using the same features that were used in the boosted trees. 20 folds CV.\n\n### FCN + Transformer Encoder\n\nProbably the most interesting model.  A NN with double inputs: tabular features (from previous preprocessing) and full-time series of ranked raw features. The training process was as follows:\n\n1. Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]\n2. Freeze FCN weights and train the Transformer for ~30 epochs to produce 1d output [0,1]\n3. Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F31f1f8ded2cba706ad095ed7595847fc%2Farchitecture.png?generation=1661419863567301&alt=media)\n\n\nFor the Transformer, I used the standard TransformerEncoderBlock with 8 heads and 2 layers",
      "votes": null
    },
    {
      "id": "1913358",
      "postDate": "08/25/2022 09:38:31",
      "content": "<p>Muchas felicidades José Antonio (y a Daniel)</p>",
      "rawMarkdown": "Muchas felicidades José Antonio (y a Daniel)",
      "votes": null
    },
    {
      "id": "1913360",
      "postDate": "08/25/2022 09:40:07",
      "content": "<p>Gracias Santiago</p>",
      "rawMarkdown": "Gracias Santiago",
      "votes": null
    },
    {
      "id": "1913369",
      "postDate": "08/25/2022 09:51:10",
      "content": "<p>Congratulations on your achievement <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a>, it would be nice to get these notebooks</p>",
      "rawMarkdown": "Congratulations on your achievement @joseantonioalatorre, it would be nice to get these notebooks",
      "votes": null
    },
    {
      "id": "1913375",
      "postDate": "08/25/2022 09:54:52",
      "content": "<p>This is really really impressive! <br>\nDo you plan to release the code? </p>",
      "rawMarkdown": "This is really really impressive! \nDo you plan to release the code?",
      "votes": null
    },
    {
      "id": "1913437",
      "postDate": "08/25/2022 10:21:52",
      "content": "<p>I am afraid I cant because I use custom wrappers for the models that are not public so will take me looooong to re-write the code :/</p>",
      "rawMarkdown": "I am afraid I cant because I use custom wrappers for the models that are not public so will take me looooong to re-write the code :/",
      "votes": null
    },
    {
      "id": "1913507",
      "postDate": "08/25/2022 10:46:12",
      "content": "<p>Thanks a lot for the mention. Did you use more than base aggregate ? </p>",
      "rawMarkdown": "Thanks a lot for the mention. Did you use more than base aggregate ?",
      "votes": null
    },
    {
      "id": "1913532",
      "postDate": "08/25/2022 10:55:49",
      "content": "<p>No only min,max,mean last, and diff last-first</p>",
      "rawMarkdown": "No only min,max,mean last, and diff last-first",
      "votes": null
    },
    {
      "id": "1913639",
      "postDate": "08/25/2022 12:28:07",
      "content": "<p>Wonderful job! <br>\nUnfortunately in the end we failed to merge.</p>",
      "rawMarkdown": "Wonderful job! \nUnfortunately in the end we failed to merge.",
      "votes": null
    },
    {
      "id": "1913657",
      "postDate": "08/25/2022 12:35:22",
      "content": "<p>Maybe next time! </p>",
      "rawMarkdown": "Maybe next time!",
      "votes": null
    },
    {
      "id": "1913750",
      "postDate": "08/25/2022 13:39:22",
      "content": "<p>Wishing you hearty congratulations <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a> for the innovative approach and the fantastic result it gave you! </p>",
      "rawMarkdown": "Wishing you hearty congratulations @joseantonioalatorre for the innovative approach and the fantastic result it gave you!",
      "votes": null
    },
    {
      "id": "1913791",
      "postDate": "08/25/2022 14:05:25",
      "content": "<p>Wow！ good job！Thanks for sharing！</p>",
      "rawMarkdown": "Wow！ good job！Thanks for sharing！",
      "votes": null
    },
    {
      "id": "1914074",
      "postDate": "08/25/2022 17:48:03",
      "content": "<p>Congrats … THe hybrid approach is really interesting and seems the secret sauce we tried the two approaches separately and didn't bring fruits like the one you had merging the two as separate inputs. Hope if you could share the code for this Model architecture . </p>",
      "rawMarkdown": "Congrats ... THe hybrid approach is really interesting and seems the secret sauce we tried the two approaches separately and didn't bring fruits like the one you had merging the two as separate inputs. Hope if you could share the code for this Model architecture .",
      "votes": null
    },
    {
      "id": "1915083",
      "postDate": "08/26/2022 16:40:48",
      "content": "<p>Hi Gaurav, its on private wrappers so I cant share them </p>",
      "rawMarkdown": "Hi Gaurav, its on private wrappers so I cant share them",
      "votes": null
    },
    {
      "id": "1915120",
      "postDate": "08/26/2022 17:17:34",
      "content": "<p>np Thanks will try to make for my knowledge based on the notes you stated.. Think ensembles did at last achieve better scores ( but not selected I think this has been case for most :) ) ,but your model is great as it can be used in a more real world scenario .. 👍</p>",
      "rawMarkdown": "np Thanks will try to make for my knowledge based on the notes you stated.. Think ensembles did at last achieve better scores ( but not selected I think this has been case for most :) ) ,but your model is great as it can be used in a more real world scenario .. 👍",
      "votes": null
    },
    {
      "id": "1915328",
      "postDate": "08/26/2022 21:50:59",
      "content": "<p>Really impressive and innovative approach! Many congrats! </p>",
      "rawMarkdown": "Really impressive and innovative approach! Many congrats!",
      "votes": null
    },
    {
      "id": "1916950",
      "postDate": "08/28/2022 09:48:32",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a> I wanted to ask you a question, regarding the differences (diff) in feature engineering did you also use the differences in absolute value?</p>",
      "rawMarkdown": "Hi @joseantonioalatorre I wanted to ask you a question, regarding the differences (diff) in feature engineering did you also use the differences in absolute value?",
      "votes": null
    },
    {
      "id": "1917002",
      "postDate": "08/28/2022 10:29:26",
      "content": "<p>Hi Raimondo, no I did not</p>",
      "rawMarkdown": "Hi Raimondo, no I did not",
      "votes": null
    },
    {
      "id": "1928194",
      "postDate": "09/06/2022 10:11:40",
      "content": "<p>Happy to see a spanish kaggler in the gold zone, congratulations and thanks for sharing! </p>\n<p>In would like to ask some questions:</p>\n<p><strong>1.</strong> Over which features did you compute the PCA?<br>\n<strong>2.</strong> Regarding the globally ranked features, \"globally\" means that you ranked them considering all customers or per customer?<br>\n<strong>3.</strong> How was the performance of the FCN compared with the GBDTs? Could you roughly share the architecture?<br>\n<strong>4.</strong> Did you experiment a boost after sequentially removing features by split importance? <br>\n<strong>5.</strong> The hybrid FCN + Transformer is very interesting, what was its CV/LB score? However, I think I didn't fully get the training procedure. For example, at first, you freeze Transformer weights and train the FCN: does this mean that during this step the <code>0.5FCN + 0.5Transformer</code> is still carried out but with Transformer initial (not trained) weights?</p>",
      "rawMarkdown": "Happy to see a spanish kaggler in the gold zone, congratulations and thanks for sharing! \n\nIn would like to ask some questions:\n\n**1.** Over which features did you compute the PCA?\n**2.** Regarding the globally ranked features, \"globally\" means that you ranked them considering all customers or per customer?\n**3.** How was the performance of the FCN compared with the GBDTs? Could you roughly share the architecture?\n**4.** Did you experiment a boost after sequentially removing features by split importance? \n**5.** The hybrid FCN + Transformer is very interesting, what was its CV/LB score? However, I think I didn't fully get the training procedure. For example, at first, you freeze Transformer weights and train the FCN: does this mean that during this step the `0.5FCN + 0.5Transformer` is still carried out but with Transformer initial (not trained) weights?",
      "votes": null
    },
    {
      "id": "1928210",
      "postDate": "09/06/2022 10:20:07",
      "content": "<p>Hi </p>\n<ol>\n<li>All of them.</li>\n<li>I considered all consumers.</li>\n<li>FCN and all my DL approaches perform ~.002 worst than other methods.</li>\n<li>Yes.</li>\n<li>(CV like in 3)  Correct as you described it. </li>\n</ol>",
      "rawMarkdown": "Hi \n1. All of them.\n2.  I considered all consumers.\n3. FCN and all my DL approaches perform ~.002 worst than other methods.\n4. Yes.\n5. (CV like in 3)  Correct as you described it.",
      "votes": null
    },
    {
      "id": "1929711",
      "postDate": "09/07/2022 09:57:48",
      "content": "<p>Congrats ! The ensembling method is amzaing. By the way, I wonder the software you use to plot the first figrue.</p>",
      "rawMarkdown": "Congrats ! The ensembling method is amzaing. By the way, I wonder the software you use to plot the first figrue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913358,
      "author_name": "santiagomota",
      "author_url": "",
      "post_date": "08/25/2022 09:38:31",
      "content": "<p>Muchas felicidades José Antonio (y a Daniel)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913360,
          "author_name": "joseantonioalatorre",
          "author_url": "",
          "post_date": "08/25/2022 09:40:07",
          "content": "<p>Gracias Santiago</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913369,
      "author_name": "raimondomelis",
      "author_url": "",
      "post_date": "08/25/2022 09:51:10",
      "content": "<p>Congratulations on your achievement <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a>, it would be nice to get these notebooks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913375,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "08/25/2022 09:54:52",
      "content": "<p>This is really really impressive! <br>\nDo you plan to release the code? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1913437,
          "author_name": "joseantonioalatorre",
          "author_url": "",
          "post_date": "08/25/2022 10:21:52",
          "content": "<p>I am afraid I cant because I use custom wrappers for the models that are not public so will take me looooong to re-write the code :/</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913507,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "08/25/2022 10:46:12",
      "content": "<p>Thanks a lot for the mention. Did you use more than base aggregate ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1913532,
          "author_name": "joseantonioalatorre",
          "author_url": "",
          "post_date": "08/25/2022 10:55:49",
          "content": "<p>No only min,max,mean last, and diff last-first</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913639,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "08/25/2022 12:28:07",
      "content": "<p>Wonderful job! <br>\nUnfortunately in the end we failed to merge.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913657,
          "author_name": "joseantonioalatorre",
          "author_url": "",
          "post_date": "08/25/2022 12:35:22",
          "content": "<p>Maybe next time! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913750,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/25/2022 13:39:22",
      "content": "<p>Wishing you hearty congratulations <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a> for the innovative approach and the fantastic result it gave you! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1913791,
      "author_name": "shanggangli",
      "author_url": "",
      "post_date": "08/25/2022 14:05:25",
      "content": "<p>Wow！ good job！Thanks for sharing！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914074,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "08/25/2022 17:48:03",
      "content": "<p>Congrats … THe hybrid approach is really interesting and seems the secret sauce we tried the two approaches separately and didn't bring fruits like the one you had merging the two as separate inputs. Hope if you could share the code for this Model architecture . </p>",
      "votes": null,
      "replies": [
        {
          "id": 1915083,
          "author_name": "joseantonioalatorre",
          "author_url": "",
          "post_date": "08/26/2022 16:40:48",
          "content": "<p>Hi Gaurav, its on private wrappers so I cant share them </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915120,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "08/26/2022 17:17:34",
          "content": "<p>np Thanks will try to make for my knowledge based on the notes you stated.. Think ensembles did at last achieve better scores ( but not selected I think this has been case for most :) ) ,but your model is great as it can be used in a more real world scenario .. 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915328,
      "author_name": "alberteinsten",
      "author_url": "",
      "post_date": "08/26/2022 21:50:59",
      "content": "<p>Really impressive and innovative approach! Many congrats! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1916950,
      "author_name": "raimondomelis",
      "author_url": "",
      "post_date": "08/28/2022 09:48:32",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/joseantonioalatorre\" target=\"_blank\">@joseantonioalatorre</a> I wanted to ask you a question, regarding the differences (diff) in feature engineering did you also use the differences in absolute value?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1917002,
          "author_name": "joseantonioalatorre",
          "author_url": "",
          "post_date": "08/28/2022 10:29:26",
          "content": "<p>Hi Raimondo, no I did not</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1928194,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "09/06/2022 10:11:40",
      "content": "<p>Happy to see a spanish kaggler in the gold zone, congratulations and thanks for sharing! </p>\n<p>In would like to ask some questions:</p>\n<p><strong>1.</strong> Over which features did you compute the PCA?<br>\n<strong>2.</strong> Regarding the globally ranked features, \"globally\" means that you ranked them considering all customers or per customer?<br>\n<strong>3.</strong> How was the performance of the FCN compared with the GBDTs? Could you roughly share the architecture?<br>\n<strong>4.</strong> Did you experiment a boost after sequentially removing features by split importance? <br>\n<strong>5.</strong> The hybrid FCN + Transformer is very interesting, what was its CV/LB score? However, I think I didn't fully get the training procedure. For example, at first, you freeze Transformer weights and train the FCN: does this mean that during this step the <code>0.5FCN + 0.5Transformer</code> is still carried out but with Transformer initial (not trained) weights?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1928210,
          "author_name": "joseantonioalatorre",
          "author_url": "",
          "post_date": "09/06/2022 10:20:07",
          "content": "<p>Hi </p>\n<ol>\n<li>All of them.</li>\n<li>I considered all consumers.</li>\n<li>FCN and all my DL approaches perform ~.002 worst than other methods.</li>\n<li>Yes.</li>\n<li>(CV like in 3)  Correct as you described it. </li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1929711,
      "author_name": "ystrangex",
      "author_url": "",
      "post_date": "09/07/2022 09:57:48",
      "content": "<p>Congrats ! The ensembling method is amzaing. By the way, I wonder the software you use to plot the first figrue.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1913348": "Thank you Amex and Kaggle for hosting! Very fun competition with a lot of insights. \n\n## Solution\n\nThis is the first of two posts of our ensemble @gandagorn  .  My leg of the ensemble consisted of taking different preprocessing methods and passing them through different architectures (Thank you @jiweiliu @lucasmorin @raddar  ! for your excellent notebooks) \nsecond part  is [https://www.kaggle.com/competitions/amex-default-prediction/discussion/348058](url)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F7a02721f85aa7042adf4202cb9e7c624%2Fensemble.png?generation=1661418572999162&alt=media)\n\n\n## Preprocessing\n\n1. Used the @raddar  integer features  and raw features and then created  separate pre-processing pipelines including standard aggregations (min, max means, last, diff) \n2. Added top 50 PCA Features. \n3. Globally ranked all the features and built products of features that shared the same trend. This boosted all the models. \n4. Used embedded layers for all categorical features. \n\n## PostProcessing\n\nSequentially removed features accordingly to split importance. \n\n## Models\n\n### LGBM with Focal Loss + logloss\n\nStandard lgbm trained with dart with one-sided focal loss added to the log-loss. The addition of the one-side focal loss was with the objective of having a better classification of the positive classes.\n\n### XGBoost\n\nNothing special.\n\n### FCN\n\nFully connected architecture using the same features that were used in the boosted trees. 20 folds CV.\n\n### FCN + Transformer Encoder\n\nProbably the most interesting model.  A NN with double inputs: tabular features (from previous preprocessing) and full-time series of ranked raw features. The training process was as follows:\n\n1. Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]\n2. Freeze FCN weights and train the Transformer for ~30 epochs to produce 1d output [0,1]\n3. Freeze Transformer weights and train the FCN architecture for ~30 epochs to produce 1d output [0,1]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4900136%2F31f1f8ded2cba706ad095ed7595847fc%2Farchitecture.png?generation=1661419863567301&alt=media)\n\n\nFor the Transformer, I used the standard TransformerEncoderBlock with 8 heads and 2 layers",
    "1913358": "Muchas felicidades José Antonio (y a Daniel)",
    "1913360": "Gracias Santiago",
    "1913369": "Congratulations on your achievement @joseantonioalatorre, it would be nice to get these notebooks",
    "1913375": "This is really really impressive! \nDo you plan to release the code?",
    "1913437": "I am afraid I cant because I use custom wrappers for the models that are not public so will take me looooong to re-write the code :/",
    "1913507": "Thanks a lot for the mention. Did you use more than base aggregate ?",
    "1913532": "No only min,max,mean last, and diff last-first",
    "1913639": "Wonderful job! \nUnfortunately in the end we failed to merge.",
    "1913657": "Maybe next time!",
    "1913750": "Wishing you hearty congratulations @joseantonioalatorre for the innovative approach and the fantastic result it gave you!",
    "1913791": "Wow！ good job！Thanks for sharing！",
    "1914074": "Congrats ... THe hybrid approach is really interesting and seems the secret sauce we tried the two approaches separately and didn't bring fruits like the one you had merging the two as separate inputs. Hope if you could share the code for this Model architecture .",
    "1915083": "Hi Gaurav, its on private wrappers so I cant share them",
    "1915120": "np Thanks will try to make for my knowledge based on the notes you stated.. Think ensembles did at last achieve better scores ( but not selected I think this has been case for most :) ) ,but your model is great as it can be used in a more real world scenario .. 👍",
    "1915328": "Really impressive and innovative approach! Many congrats!",
    "1916950": "Hi @joseantonioalatorre I wanted to ask you a question, regarding the differences (diff) in feature engineering did you also use the differences in absolute value?",
    "1917002": "Hi Raimondo, no I did not",
    "1928194": "Happy to see a spanish kaggler in the gold zone, congratulations and thanks for sharing! \n\nIn would like to ask some questions:\n\n**1.** Over which features did you compute the PCA?\n**2.** Regarding the globally ranked features, \"globally\" means that you ranked them considering all customers or per customer?\n**3.** How was the performance of the FCN compared with the GBDTs? Could you roughly share the architecture?\n**4.** Did you experiment a boost after sequentially removing features by split importance? \n**5.** The hybrid FCN + Transformer is very interesting, what was its CV/LB score? However, I think I didn't fully get the training procedure. For example, at first, you freeze Transformer weights and train the FCN: does this mean that during this step the `0.5FCN + 0.5Transformer` is still carried out but with Transformer initial (not trained) weights?",
    "1928210": "Hi \n1. All of them.\n2.  I considered all consumers.\n3. FCN and all my DL approaches perform ~.002 worst than other methods.\n4. Yes.\n5. (CV like in 3)  Correct as you described it.",
    "1929711": "Congrats ! The ensembling method is amzaing. By the way, I wonder the software you use to plot the first figrue."
  },
  "source": "meta"
}