{
  "id": 347880,
  "title": "25th place solution",
  "url": "/competitions/amex-default-prediction/writeups/pavel-and-eran-25th-place-solution",
  "author_name": "",
  "post_date": "2022-08-27T17:58:52.857Z",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<h2>Intro</h2>\n<p>As a senior AI researcher in a start-up company called <strong>pecan.ai</strong> I mostly deal with similar types of transational data. I used my work experience to get a smooth start. Also a huge thanks to my bosses for providing me with computational power :) It will absolutely paid off by amount of usefull directions I learned from such a warm community :)</p>\n<h2>Feature Engineering:</h2>\n<p>We used a couple of versions of feature engineering (it was kind of chaotic) and not all techniques described here were applied for all models. It’s done for couple of reasons:</p>\n<ul>\n<li>Make datasets slightly more diverse</li>\n<li>Reengineering features after some time to be sure that there are no bugs.<br>\nas a basis, we used <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> 's dataset</li>\n</ul>\n<h3>1. Pre-flattening - engineering done on sequences:</h3>\n<p>We started with “after-pay” features and then advanced into feature interaction: We iterated over all possible combinations of the numerical features, calculated the difference between each pair of features and filled NAN with 0, if the linear correlation was significantly higher than either one of the standalone features then this difference, became a new standalone feature and was added to the model, e.g. <code>P_2–B_3</code>.<br>\nWe also used amex metric instead of linear correlation.<br>\nIn one of datasets categorical features was one-hot encoded into multiple binary sequences.</p>\n<h3>2. Flattening - reducing to a single value:</h3>\n<p>Basics - mean, min, max, std, first, last (indeed).<br>\nAlso - middles: <code>mean(1:12)</code>, <code>mean(2:13)</code>, <code>mean(2:12)</code>. The reason - post-flattening of <code>last-mean(1:13)</code> makes less sense because <code>mean</code> already has <code>last</code> inside, so <code>last-mean(1:12)</code> sounds as more correct solution. In practice, I do not know to measure how it is usefull.</p>\n<h3>3. Post-flattening - operating on created features:</h3>\n<ul>\n<li>Last - first, last - middle, middle - first, relative_pos: (last - min) / (max - min)</li>\n<li>Row features: count of NANs, sum of normalised values, difference between normalised values of P,D,S,B,R e.g. sum(normalisedP) – sum(normalisedD)</li>\n</ul>\n<h3>Specialized features</h3>\n<ul>\n<li>feature value momentum</li>\n<li>linear predictions of each feature 180 days into the future using the last 3 months only</li>\n<li>days since last NAN value</li>\n<li>(last value – prev value ) / (last_date – prev_date) </li>\n<li>and some other variation of the above features.<br>\nFeature Engineering was done in Python and C# .Net:</li>\n</ul>\n<h3>SequentialEncoder</h3>\n<p>We also created sequential features described in <a href=\"https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431\" target=\"_blank\">this</a> notebook (they was very significant).  </p>\n<h2>Feature Selection</h2>\n<p>We tested some simple methods, but they was reducing our score. We decided that loosing 4th point  of CV does not worth it, we are still able to run a model with 3.5k features dataset (our biggest one), so risk will not paid off.</p>\n<h2>Modeling:</h2>\n<h3>GBDT:</h3>\n<p>LightGBM Dart, CatBoost, XGBoost </p>\n<h3>Tabnet</h3>\n<p>We did not manage to get useful results (all ensembles almost nullify its contribution).<br>\n<strong>But!</strong> then we tried the trick I tested a couple of years ago. We took our best model (lightgbm dart) and calculated <strong>shap values</strong> in an out-of-fold manner. Back then I called it self-supervised pre training, but technically it's a feature transformation. As a result we got a dataset which is much easier to digest for NN models, because it was extracted with GBT. Then I trained tabnet using this dataset and achieved <strong>0.797</strong> on LB. That solution not so diverse, as straight-forward tabnet model. But Tabnet learned predictions in different manner than GBT, so it still was very useful for the ensemble. <br>\nYou can find more detailed explaination <a href=\"https://www.kaggle.com/code/pavelvod/gbm-supervised-pretraining\" target=\"_blank\">here</a> (from some previous competition)</p>\n<h3>Node</h3>\n<p>But my greatest excitement was trying the <strong>NODE</strong> model for this competition. It achieved good results - I do not think my results were optimal, I believe If I would finetune it more - it would achieve better results. But It was probably the heaviest tabular model I ever tried - my 12 gb GPU (thanks <strong>pecan.ai</strong>) was screaming with only 64 batch size.</p>\n<h3>Public models:</h3>\n<p>We used TF transformer by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> with 3 seeds.</p>\n<h3>Sequential models</h3>\n<p>We used <code>tsai</code> library which has plenty of sequential models, such as LSTM, 1D-CNN and many many more their advanced versions and implementations. Nothing was found usefull for us.</p>\n<h3>Training parameters</h3>\n<p>We used logloss for all our models both as loss and stopping metric. We tried FocalLoss and RankingLoss, but they not worked for us.<br>\nWe used 5-fold CV.<br>\nAlmost all models was trained and averaged with 4 seeds (3 seeds was stratified by target and 4th was stratified by target and P_2_last - idea by <a href=\"https://www.kaggle.com/bogorodvo\" target=\"_blank\">@bogorodvo</a> </p>\n<h3>Ensemble</h3>\n<p>As an EnsemblerClassifier, we developed an iterative process, known as a forward selection process, which was able to find the optimum weights to maximize AMEX score.<br>\nWe also tried other methods which optimised LogLoss, but they was not good enough.</p>\n<h2>Worth trying</h2>\n<p>Couple of key ideas that worked well for other people was also in our plans, but we somehow we gave up on this ideas. I will not mention them.<br>\nThe only thing I regret that I did not tried (and still not found someone tested it) is Tabnet self-supervised pretraining on the whole train and test data and then finetuning on the train data. I started to regret after seeing the knowledge distillation solution of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
  "messages": [
    {
      "id": "1914144",
      "postDate": "08/25/2022 19:08:19",
      "content": "<h2>Intro</h2>\n<p>As a senior AI researcher in a start-up company called <strong>pecan.ai</strong> I mostly deal with similar types of transational data. I used my work experience to get a smooth start. Also a huge thanks to my bosses for providing me with computational power :) It will absolutely paid off by amount of usefull directions I learned from such a warm community :)</p>\n<h2>Feature Engineering:</h2>\n<p>We used a couple of versions of feature engineering (it was kind of chaotic) and not all techniques described here were applied for all models. It’s done for couple of reasons:</p>\n<ul>\n<li>Make datasets slightly more diverse</li>\n<li>Reengineering features after some time to be sure that there are no bugs.<br>\nas a basis, we used <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> 's dataset</li>\n</ul>\n<h3>1. Pre-flattening - engineering done on sequences:</h3>\n<p>We started with “after-pay” features and then advanced into feature interaction: We iterated over all possible combinations of the numerical features, calculated the difference between each pair of features and filled NAN with 0, if the linear correlation was significantly higher than either one of the standalone features then this difference, became a new standalone feature and was added to the model, e.g. <code>P_2–B_3</code>.<br>\nWe also used amex metric instead of linear correlation.<br>\nIn one of datasets categorical features was one-hot encoded into multiple binary sequences.</p>\n<h3>2. Flattening - reducing to a single value:</h3>\n<p>Basics - mean, min, max, std, first, last (indeed).<br>\nAlso - middles: <code>mean(1:12)</code>, <code>mean(2:13)</code>, <code>mean(2:12)</code>. The reason - post-flattening of <code>last-mean(1:13)</code> makes less sense because <code>mean</code> already has <code>last</code> inside, so <code>last-mean(1:12)</code> sounds as more correct solution. In practice, I do not know to measure how it is usefull.</p>\n<h3>3. Post-flattening - operating on created features:</h3>\n<ul>\n<li>Last - first, last - middle, middle - first, relative_pos: (last - min) / (max - min)</li>\n<li>Row features: count of NANs, sum of normalised values, difference between normalised values of P,D,S,B,R e.g. sum(normalisedP) – sum(normalisedD)</li>\n</ul>\n<h3>Specialized features</h3>\n<ul>\n<li>feature value momentum</li>\n<li>linear predictions of each feature 180 days into the future using the last 3 months only</li>\n<li>days since last NAN value</li>\n<li>(last value – prev value ) / (last_date – prev_date) </li>\n<li>and some other variation of the above features.<br>\nFeature Engineering was done in Python and C# .Net:</li>\n</ul>\n<h3>SequentialEncoder</h3>\n<p>We also created sequential features described in <a href=\"https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431\" target=\"_blank\">this</a> notebook (they was very significant).  </p>\n<h2>Feature Selection</h2>\n<p>We tested some simple methods, but they was reducing our score. We decided that loosing 4th point  of CV does not worth it, we are still able to run a model with 3.5k features dataset (our biggest one), so risk will not paid off.</p>\n<h2>Modeling:</h2>\n<h3>GBDT:</h3>\n<p>LightGBM Dart, CatBoost, XGBoost </p>\n<h3>Tabnet</h3>\n<p>We did not manage to get useful results (all ensembles almost nullify its contribution).<br>\n<strong>But!</strong> then we tried the trick I tested a couple of years ago. We took our best model (lightgbm dart) and calculated <strong>shap values</strong> in an out-of-fold manner. Back then I called it self-supervised pre training, but technically it's a feature transformation. As a result we got a dataset which is much easier to digest for NN models, because it was extracted with GBT. Then I trained tabnet using this dataset and achieved <strong>0.797</strong> on LB. That solution not so diverse, as straight-forward tabnet model. But Tabnet learned predictions in different manner than GBT, so it still was very useful for the ensemble. <br>\nYou can find more detailed explaination <a href=\"https://www.kaggle.com/code/pavelvod/gbm-supervised-pretraining\" target=\"_blank\">here</a> (from some previous competition)</p>\n<h3>Node</h3>\n<p>But my greatest excitement was trying the <strong>NODE</strong> model for this competition. It achieved good results - I do not think my results were optimal, I believe If I would finetune it more - it would achieve better results. But It was probably the heaviest tabular model I ever tried - my 12 gb GPU (thanks <strong>pecan.ai</strong>) was screaming with only 64 batch size.</p>\n<h3>Public models:</h3>\n<p>We used TF transformer by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> with 3 seeds.</p>\n<h3>Sequential models</h3>\n<p>We used <code>tsai</code> library which has plenty of sequential models, such as LSTM, 1D-CNN and many many more their advanced versions and implementations. Nothing was found usefull for us.</p>\n<h3>Training parameters</h3>\n<p>We used logloss for all our models both as loss and stopping metric. We tried FocalLoss and RankingLoss, but they not worked for us.<br>\nWe used 5-fold CV.<br>\nAlmost all models was trained and averaged with 4 seeds (3 seeds was stratified by target and 4th was stratified by target and P_2_last - idea by <a href=\"https://www.kaggle.com/bogorodvo\" target=\"_blank\">@bogorodvo</a> </p>\n<h3>Ensemble</h3>\n<p>As an EnsemblerClassifier, we developed an iterative process, known as a forward selection process, which was able to find the optimum weights to maximize AMEX score.<br>\nWe also tried other methods which optimised LogLoss, but they was not good enough.</p>\n<h2>Worth trying</h2>\n<p>Couple of key ideas that worked well for other people was also in our plans, but we somehow we gave up on this ideas. I will not mention them.<br>\nThe only thing I regret that I did not tried (and still not found someone tested it) is Tabnet self-supervised pretraining on the whole train and test data and then finetuning on the train data. I started to regret after seeing the knowledge distillation solution of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "## Intro\nAs a senior AI researcher in a start-up company called **pecan.ai** I mostly deal with similar types of transational data. I used my work experience to get a smooth start. Also a huge thanks to my bosses for providing me with computational power :) It will absolutely paid off by amount of usefull directions I learned from such a warm community :)\n \n\n## Feature Engineering:\n\nWe used a couple of versions of feature engineering (it was kind of chaotic) and not all techniques described here were applied for all models. It’s done for couple of reasons:\n\n- Make datasets slightly more diverse\n- Reengineering features after some time to be sure that there are no bugs.\n\n  as a basis, we used @raddar 's dataset\n\n\n### 1. Pre-flattening - engineering done on sequences:\n\nWe started with “after-pay” features and then advanced into feature interaction: We iterated over all possible combinations of the numerical features, calculated the difference between each pair of features and filled NAN with 0, if the linear correlation was significantly higher than either one of the standalone features then this difference, became a new standalone feature and was added to the model, e.g. `P_2–B_3`.\nWe also used amex metric instead of linear correlation.\n\nIn one of datasets categorical features was one-hot encoded into multiple binary sequences.\n\n### 2. Flattening - reducing to a single value:\n\nBasics - mean, min, max, std, first, last (indeed).\n\nAlso - middles: `mean(1:12)`, `mean(2:13)`, `mean(2:12)`. The reason - post-flattening of `last-mean(1:13)` makes less sense because `mean` already has `last` inside, so `last-mean(1:12)` sounds as more correct solution. In practice, I do not know to measure how it is usefull.\n\n\n### 3. Post-flattening - operating on created features:\n\n* Last - first, last - middle, middle - first, relative_pos: (last - min) / (max - min)\n\n* Row features: count of NANs, sum of normalised values, difference between normalised values of P,D,S,B,R e.g. sum(normalisedP) – sum(normalisedD)\n\n### Specialized features\n* feature value momentum\n* linear predictions of each feature 180 days into the future using the last 3 months only\n*  days since last NAN value\n* (last value – prev value ) / (last_date – prev_date) \n* and some other variation of the above features.\n\nFeature Engineering was done in Python and C# .Net:\n\n### SequentialEncoder\nWe also created sequential features described in [this](https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431) notebook (they was very significant).  \n \n \n ## Feature Selection\n \n We tested some simple methods, but they was reducing our score. We decided that loosing 4th point  of CV does not worth it, we are still able to run a model with 3.5k features dataset (our biggest one), so risk will not paid off.\n  \n## Modeling:\n\n\n### GBDT:\nLightGBM Dart, CatBoost, XGBoost \n\n### Tabnet\nWe did not manage to get useful results (all ensembles almost nullify its contribution).\n\n**But!** then we tried the trick I tested a couple of years ago. We took our best model (lightgbm dart) and calculated **shap values** in an out-of-fold manner. Back then I called it self-supervised pre training, but technically it's a feature transformation. As a result we got a dataset which is much easier to digest for NN models, because it was extracted with GBT. Then I trained tabnet using this dataset and achieved **0.797** on LB. That solution not so diverse, as straight-forward tabnet model. But Tabnet learned predictions in different manner than GBT, so it still was very useful for the ensemble. \n\nYou can find more detailed explaination [here](https://www.kaggle.com/code/pavelvod/gbm-supervised-pretraining) (from some previous competition)\n\n\n ### Node\nBut my greatest excitement was trying the **NODE** model for this competition. It achieved good results - I do not think my results were optimal, I believe If I would finetune it more - it would achieve better results. But It was probably the heaviest tabular model I ever tried - my 12 gb GPU (thanks **pecan.ai**) was screaming with only 64 batch size.\n\n### Public models:\nWe used TF transformer by @cdeotte with 3 seeds.\n\n### Sequential models\nWe used `tsai` library which has plenty of sequential models, such as LSTM, 1D-CNN and many many more their advanced versions and implementations. Nothing was found usefull for us.\n\n### Training parameters\nWe used logloss for all our models both as loss and stopping metric. We tried FocalLoss and RankingLoss, but they not worked for us.\nWe used 5-fold CV.\nAlmost all models was trained and averaged with 4 seeds (3 seeds was stratified by target and 4th was stratified by target and P_2_last - idea by @bogorodvo \n\n### Ensemble\n\nAs an EnsemblerClassifier, we developed an iterative process, known as a forward selection process, which was able to find the optimum weights to maximize AMEX score.\nWe also tried other methods which optimised LogLoss, but they was not good enough.\n\n## Worth trying\n\nCouple of key ideas that worked well for other people was also in our plans, but we somehow we gave up on this ideas. I will not mention them.\n\n\nThe only thing I regret that I did not tried (and still not found someone tested it) is Tabnet self-supervised pretraining on the whole train and test data and then finetuning on the train data. I started to regret after seeing the knowledge distillation solution of @cdeotte",
      "votes": null
    },
    {
      "id": "1914159",
      "postDate": "08/25/2022 19:35:35",
      "content": "<p>What is NODE model? The name is common, so hard to search for online…</p>",
      "rawMarkdown": "What is NODE model? The name is common, so hard to search for online...",
      "votes": null
    },
    {
      "id": "1914173",
      "postDate": "08/25/2022 19:51:30",
      "content": "<p><a href=\"https://arxiv.org/abs/1909.06312v2\" target=\"_blank\">https://arxiv.org/abs/1909.06312v2</a></p>\n<p>I used implementation in <code>pytorch_tabular</code> library</p>",
      "rawMarkdown": "https://arxiv.org/abs/1909.06312v2\n\nI used implementation in `pytorch_tabular` library",
      "votes": null
    },
    {
      "id": "1914541",
      "postDate": "08/26/2022 07:18:04",
      "content": "<p>Congratulations on your placement. Just a question on your NODE model did you use Oblivious Decision Trees or Normal Decision Trees? How long did it take you to finish training with a GPU, i have heard takes pretty long.</p>",
      "rawMarkdown": "Congratulations on your placement. Just a question on your NODE model did you use Oblivious Decision Trees or Normal Decision Trees? How long did it take you to finish training with a GPU, i have heard takes pretty long.",
      "votes": null
    },
    {
      "id": "1915192",
      "postDate": "08/26/2022 18:22:09",
      "content": "<p>Thanks!<br>\nAFAIR It tooks about 8 hours for 5-fold training and predicting oof + test.<br>\nIncreasing number of trees or number of layers was going into GPU memory overflow (12 gb) with 512 batches or even less<br>\nWhile other models such as tabnet used about 2-3 Gb of GPU memory with much bigger batch_size (e.g. 2048) and radical changes in hyperparameters.</p>\n<p>Here is my config (I used <code>pytorch_tabular</code> package)</p>\n<pre><code>data_config = DataConfig(\n            target=['target'], \n            continuous_cols=feature_names,\n            categorical_cols=[],\n        )\ntrainer_config = TrainerConfig(\n    auto_lr_find=False, \n    batch_size=512,\n    gpus=1,\n)\noptimizer_config = OptimizerConfig()\n\nmodel_config = NodeConfig(\n    task=\"classification\",\n    num_layers=3,\n    num_trees=512,\n    learning_rate = 1e-3\n)\n\nmodel = TabularModel(\n    data_config=data_config,\n    model_config=model_config,\n    optimizer_config=optimizer_config,\n    trainer_config=trainer_config,\n)\n</code></pre>",
      "rawMarkdown": "Thanks!\nAFAIR It tooks about 8 hours for 5-fold training and predicting oof + test.\nIncreasing number of trees or number of layers was going into GPU memory overflow (12 gb) with 512 batches or even less\nWhile other models such as tabnet used about 2-3 Gb of GPU memory with much bigger batch_size (e.g. 2048) and radical changes in hyperparameters.\n\nHere is my config (I used `pytorch_tabular` package)\n\n```\ndata_config = DataConfig(\n            target=['target'], \n            continuous_cols=feature_names,\n            categorical_cols=[],\n        )\ntrainer_config = TrainerConfig(\n    auto_lr_find=False, \n    batch_size=512,\n    gpus=1,\n)\noptimizer_config = OptimizerConfig()\n\nmodel_config = NodeConfig(\n    task=\"classification\",\n    num_layers=3,\n    num_trees=512,\n    learning_rate = 1e-3\n)\n\nmodel = TabularModel(\n    data_config=data_config,\n    model_config=model_config,\n    optimizer_config=optimizer_config,\n    trainer_config=trainer_config,\n)\n```",
      "votes": null
    },
    {
      "id": "1922114",
      "postDate": "09/01/2022 09:22:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pavelvod\" target=\"_blank\">@pavelvod</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @pavelvod, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1914159,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/25/2022 19:35:35",
      "content": "<p>What is NODE model? The name is common, so hard to search for online…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1914173,
          "author_name": "pavelvod",
          "author_url": "",
          "post_date": "08/25/2022 19:51:30",
          "content": "<p><a href=\"https://arxiv.org/abs/1909.06312v2\" target=\"_blank\">https://arxiv.org/abs/1909.06312v2</a></p>\n<p>I used implementation in <code>pytorch_tabular</code> library</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1914541,
      "author_name": "tarrasque9",
      "author_url": "",
      "post_date": "08/26/2022 07:18:04",
      "content": "<p>Congratulations on your placement. Just a question on your NODE model did you use Oblivious Decision Trees or Normal Decision Trees? How long did it take you to finish training with a GPU, i have heard takes pretty long.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915192,
          "author_name": "pavelvod",
          "author_url": "",
          "post_date": "08/26/2022 18:22:09",
          "content": "<p>Thanks!<br>\nAFAIR It tooks about 8 hours for 5-fold training and predicting oof + test.<br>\nIncreasing number of trees or number of layers was going into GPU memory overflow (12 gb) with 512 batches or even less<br>\nWhile other models such as tabnet used about 2-3 Gb of GPU memory with much bigger batch_size (e.g. 2048) and radical changes in hyperparameters.</p>\n<p>Here is my config (I used <code>pytorch_tabular</code> package)</p>\n<pre><code>data_config = DataConfig(\n            target=['target'], \n            continuous_cols=feature_names,\n            categorical_cols=[],\n        )\ntrainer_config = TrainerConfig(\n    auto_lr_find=False, \n    batch_size=512,\n    gpus=1,\n)\noptimizer_config = OptimizerConfig()\n\nmodel_config = NodeConfig(\n    task=\"classification\",\n    num_layers=3,\n    num_trees=512,\n    learning_rate = 1e-3\n)\n\nmodel = TabularModel(\n    data_config=data_config,\n    model_config=model_config,\n    optimizer_config=optimizer_config,\n    trainer_config=trainer_config,\n)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1922114,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 09:22:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pavelvod\" target=\"_blank\">@pavelvod</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914144": "## Intro\nAs a senior AI researcher in a start-up company called **pecan.ai** I mostly deal with similar types of transational data. I used my work experience to get a smooth start. Also a huge thanks to my bosses for providing me with computational power :) It will absolutely paid off by amount of usefull directions I learned from such a warm community :)\n \n\n## Feature Engineering:\n\nWe used a couple of versions of feature engineering (it was kind of chaotic) and not all techniques described here were applied for all models. It’s done for couple of reasons:\n\n- Make datasets slightly more diverse\n- Reengineering features after some time to be sure that there are no bugs.\n\n  as a basis, we used @raddar 's dataset\n\n\n### 1. Pre-flattening - engineering done on sequences:\n\nWe started with “after-pay” features and then advanced into feature interaction: We iterated over all possible combinations of the numerical features, calculated the difference between each pair of features and filled NAN with 0, if the linear correlation was significantly higher than either one of the standalone features then this difference, became a new standalone feature and was added to the model, e.g. `P_2–B_3`.\nWe also used amex metric instead of linear correlation.\n\nIn one of datasets categorical features was one-hot encoded into multiple binary sequences.\n\n### 2. Flattening - reducing to a single value:\n\nBasics - mean, min, max, std, first, last (indeed).\n\nAlso - middles: `mean(1:12)`, `mean(2:13)`, `mean(2:12)`. The reason - post-flattening of `last-mean(1:13)` makes less sense because `mean` already has `last` inside, so `last-mean(1:12)` sounds as more correct solution. In practice, I do not know to measure how it is usefull.\n\n\n### 3. Post-flattening - operating on created features:\n\n* Last - first, last - middle, middle - first, relative_pos: (last - min) / (max - min)\n\n* Row features: count of NANs, sum of normalised values, difference between normalised values of P,D,S,B,R e.g. sum(normalisedP) – sum(normalisedD)\n\n### Specialized features\n* feature value momentum\n* linear predictions of each feature 180 days into the future using the last 3 months only\n*  days since last NAN value\n* (last value – prev value ) / (last_date – prev_date) \n* and some other variation of the above features.\n\nFeature Engineering was done in Python and C# .Net:\n\n### SequentialEncoder\nWe also created sequential features described in [this](https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431) notebook (they was very significant).  \n \n \n ## Feature Selection\n \n We tested some simple methods, but they was reducing our score. We decided that loosing 4th point  of CV does not worth it, we are still able to run a model with 3.5k features dataset (our biggest one), so risk will not paid off.\n  \n## Modeling:\n\n\n### GBDT:\nLightGBM Dart, CatBoost, XGBoost \n\n### Tabnet\nWe did not manage to get useful results (all ensembles almost nullify its contribution).\n\n**But!** then we tried the trick I tested a couple of years ago. We took our best model (lightgbm dart) and calculated **shap values** in an out-of-fold manner. Back then I called it self-supervised pre training, but technically it's a feature transformation. As a result we got a dataset which is much easier to digest for NN models, because it was extracted with GBT. Then I trained tabnet using this dataset and achieved **0.797** on LB. That solution not so diverse, as straight-forward tabnet model. But Tabnet learned predictions in different manner than GBT, so it still was very useful for the ensemble. \n\nYou can find more detailed explaination [here](https://www.kaggle.com/code/pavelvod/gbm-supervised-pretraining) (from some previous competition)\n\n\n ### Node\nBut my greatest excitement was trying the **NODE** model for this competition. It achieved good results - I do not think my results were optimal, I believe If I would finetune it more - it would achieve better results. But It was probably the heaviest tabular model I ever tried - my 12 gb GPU (thanks **pecan.ai**) was screaming with only 64 batch size.\n\n### Public models:\nWe used TF transformer by @cdeotte with 3 seeds.\n\n### Sequential models\nWe used `tsai` library which has plenty of sequential models, such as LSTM, 1D-CNN and many many more their advanced versions and implementations. Nothing was found usefull for us.\n\n### Training parameters\nWe used logloss for all our models both as loss and stopping metric. We tried FocalLoss and RankingLoss, but they not worked for us.\nWe used 5-fold CV.\nAlmost all models was trained and averaged with 4 seeds (3 seeds was stratified by target and 4th was stratified by target and P_2_last - idea by @bogorodvo \n\n### Ensemble\n\nAs an EnsemblerClassifier, we developed an iterative process, known as a forward selection process, which was able to find the optimum weights to maximize AMEX score.\nWe also tried other methods which optimised LogLoss, but they was not good enough.\n\n## Worth trying\n\nCouple of key ideas that worked well for other people was also in our plans, but we somehow we gave up on this ideas. I will not mention them.\n\n\nThe only thing I regret that I did not tried (and still not found someone tested it) is Tabnet self-supervised pretraining on the whole train and test data and then finetuning on the train data. I started to regret after seeing the knowledge distillation solution of @cdeotte",
    "1914159": "What is NODE model? The name is common, so hard to search for online...",
    "1914173": "https://arxiv.org/abs/1909.06312v2\n\nI used implementation in `pytorch_tabular` library",
    "1914541": "Congratulations on your placement. Just a question on your NODE model did you use Oblivious Decision Trees or Normal Decision Trees? How long did it take you to finish training with a GPU, i have heard takes pretty long.",
    "1915192": "Thanks!\nAFAIR It tooks about 8 hours for 5-fold training and predicting oof + test.\nIncreasing number of trees or number of layers was going into GPU memory overflow (12 gb) with 512 batches or even less\nWhile other models such as tabnet used about 2-3 Gb of GPU memory with much bigger batch_size (e.g. 2048) and radical changes in hyperparameters.\n\nHere is my config (I used `pytorch_tabular` package)\n\n```\ndata_config = DataConfig(\n            target=['target'], \n            continuous_cols=feature_names,\n            categorical_cols=[],\n        )\ntrainer_config = TrainerConfig(\n    auto_lr_find=False, \n    batch_size=512,\n    gpus=1,\n)\noptimizer_config = OptimizerConfig()\n\nmodel_config = NodeConfig(\n    task=\"classification\",\n    num_layers=3,\n    num_trees=512,\n    learning_rate = 1e-3\n)\n\nmodel = TabularModel(\n    data_config=data_config,\n    model_config=model_config,\n    optimizer_config=optimizer_config,\n    trainer_config=trainer_config,\n)\n```",
    "1922114": "Hi @pavelvod, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
  },
  "source": "meta"
}