{
  "id": 347745,
  "title": "2020 place overview",
  "url": "/competitions/amex-default-prediction/writeups/default-name-2061-place-overview",
  "author_name": "",
  "post_date": "2022-08-27T06:25:17.550Z",
  "votes": 28,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First and foremost we would like to thank Raddar for his cleaned and compressed dataset <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">[1]</a> (which we used) and Martin for his two magnificent DART notebooks <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963\" target=\"_blank\">[2]</a><a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">[3]</a>. We finally used a stacking of XGBoost, LightGBM and TabNet. When blending our solution with that of Martin we obtained a Private score of 0.80761, and playing the hypothetical <em>what if</em> game this would have landed around 138th on the LB. However, although acutely aware that our solution was not competitive, we still went ahead with selecting only our own work, and followed the sage advice of Raddar in submitting our best CV, and (our best CV + our best LB)/2 works. Doing this left our CV score on the wrong side of the CV of the work published by Martin, and consequently way down the LB, but with no regrets whatsoever. </p>\n<p>Things perhaps of mention:\nFeature engineering: We spent an inordinate  amount of time with the features in conjunction with the 13GB GPU memory limit on kaggle, and with each new feature sadly eventually having to sacrifice in the end around 30 other features (team-mate <a href=\"https://www.kaggle.com/danielhanchen\" target=\"_blank\">@danielhanchen</a> helped with that process) to stay within the notebook memory. In retrospect we should really have heavily sub-sampled the training data (if it wasn't for the noisy <em>D</em> component in the metric needing lots of data that would definitely be the way to go) in order to keep as many of these features as possible as almost all of them added just a little something (we did try the random under-sampling <a href=\"https://imbalanced-ensemble.readthedocs.io/en/latest/api/ensemble/_autosummary/imbalanced_ensemble.ensemble.RUSBoostClassifier.html\" target=\"_blank\">RUSBoostClassifier</a> from <a href=\"https://imbalanced-ensemble.readthedocs.io/en/latest/index.html\" target=\"_blank\">imbalanced-ensemble</a>, but it did not work well).\nAs well as a lot of futher cleaning to remove outliers <em>etc</em>, things of particular note were:</p>\n<ul>\n<li>The features <code>S_8</code> and <code>S_13</code> seemed reminiscent of some sort of FICO credit rating scores. With that in mind we created an additional two new binary features with a split (<code>np.where</code>) chosen by fitting just these features alone with a RandomForest 'stump' with the idea of 'good' or 'bad' creditor in mind.</li>\n<li>As well as generating the usual aggregated features for all of the (up to) 13 statements, we found that also using such aggregations just over the last 3 statements, and the last 6 statements, proved to produce some very informative features.</li>\n<li>We found that the feature <code>S_3/S_7</code> helped.</li>\n<li>We found that the average of the three features <code>D_115</code>, <code>D_118</code> and <code>D_119</code> also helped a little</li>\n<li>On the float features we used the <code>round(2)</code> trick suggested by Jiwei Liu <a href=\"https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick\" target=\"_blank\">[4]</a></li>\n<li>Our final number of features after aggregations was 1145</li>\n</ul>\n<p>Estimators:</p>\n<ul>\n<li>We started with also using CatBoost, but in the end we found the CV scores were not on par with those of XGBoost and LGBM</li>\n<li>We used DART with LGBM thanks to the work of Martin (but not with XGBoost; incredibly slow!)</li>\n<li>Our only novelty with XGBoost was using the <code>num_parallel_tree=7</code> parameter, which slowed down the training by around 7 times, but made for a smoother training curve and a CV with a lower standard deviation.</li>\n<li>As per the magnificent notbook by Chris Deotte <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-79391\" target=\"_blank\">[5]</a> we used surprisingly  shallow estimators (<code>max_depth=4</code>) as going deeper didn't really seem to help</li>\n<li>Combined multiple runs of each estimator, each using different CV splitting seeds</li>\n<li>Although TabNet alone did not have a spectacular CV score despite the best efforts of my team-mate <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> , it helped every ensemble it touched.</li>\n</ul>\n<p>The OOF predictions of these estimators were then calibrated and then fed into a stacking ensemble.</p>\n<p>We were also aware that the distribution of some of the features in the test data were significantly different from those in the training data (some wonderful notebooks on this were made on this by Pavel Vodolazov <a href=\"https://www.kaggle.com/code/pavelvod/amex-eda-revealing-time-patterns-of-features\" target=\"_blank\">[6]</a><a href=\"https://www.kaggle.com/code/pavelvod/amex-eda-even-more-insane-time-patterns-revealed\" target=\"_blank\">[7]</a>). In view of the fact that usually tree based estimators cannot extrapolate we tried out the <a href=\"https://github.com/cerlymarco/linear-tree\" target=\"_blank\"><code>linear-tree</code></a>  estimator, and also the <code>gblinear</code> booster in XGBoost, but neither were used in the end; they may well have eventually performed better on the test data for the features that had covariate shifts, but there was little way of telling beforehand. </p>\n<p>All in all a very enjoyable competition and many thanks to AmEx for providing us with tabular data to play with, and kaggle for the computational resources!</p>\n<p>All the best,\ncarl</p>",
  "messages": [
    {
      "id": "1913368",
      "postDate": "08/25/2022 09:50:23",
      "content": "<p>First and foremost we would like to thank Raddar for his cleaned and compressed dataset <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">[1]</a> (which we used) and Martin for his two magnificent DART notebooks <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963\" target=\"_blank\">[2]</a><a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">[3]</a>. We finally used a stacking of XGBoost, LightGBM and TabNet. When blending our solution with that of Martin we obtained a Private score of 0.80761, and playing the hypothetical <em>what if</em> game this would have landed around 138th on the LB. However, although acutely aware that our solution was not competitive, we still went ahead with selecting only our own work, and followed the sage advice of Raddar in submitting our best CV, and (our best CV + our best LB)/2 works. Doing this left our CV score on the wrong side of the CV of the work published by Martin, and consequently way down the LB, but with no regrets whatsoever. </p>\n<p>Things perhaps of mention:\nFeature engineering: We spent an inordinate  amount of time with the features in conjunction with the 13GB GPU memory limit on kaggle, and with each new feature sadly eventually having to sacrifice in the end around 30 other features (team-mate <a href=\"https://www.kaggle.com/danielhanchen\" target=\"_blank\">@danielhanchen</a> helped with that process) to stay within the notebook memory. In retrospect we should really have heavily sub-sampled the training data (if it wasn't for the noisy <em>D</em> component in the metric needing lots of data that would definitely be the way to go) in order to keep as many of these features as possible as almost all of them added just a little something (we did try the random under-sampling <a href=\"https://imbalanced-ensemble.readthedocs.io/en/latest/api/ensemble/_autosummary/imbalanced_ensemble.ensemble.RUSBoostClassifier.html\" target=\"_blank\">RUSBoostClassifier</a> from <a href=\"https://imbalanced-ensemble.readthedocs.io/en/latest/index.html\" target=\"_blank\">imbalanced-ensemble</a>, but it did not work well).\nAs well as a lot of futher cleaning to remove outliers <em>etc</em>, things of particular note were:</p>\n<ul>\n<li>The features <code>S_8</code> and <code>S_13</code> seemed reminiscent of some sort of FICO credit rating scores. With that in mind we created an additional two new binary features with a split (<code>np.where</code>) chosen by fitting just these features alone with a RandomForest 'stump' with the idea of 'good' or 'bad' creditor in mind.</li>\n<li>As well as generating the usual aggregated features for all of the (up to) 13 statements, we found that also using such aggregations just over the last 3 statements, and the last 6 statements, proved to produce some very informative features.</li>\n<li>We found that the feature <code>S_3/S_7</code> helped.</li>\n<li>We found that the average of the three features <code>D_115</code>, <code>D_118</code> and <code>D_119</code> also helped a little</li>\n<li>On the float features we used the <code>round(2)</code> trick suggested by Jiwei Liu <a href=\"https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick\" target=\"_blank\">[4]</a></li>\n<li>Our final number of features after aggregations was 1145</li>\n</ul>\n<p>Estimators:</p>\n<ul>\n<li>We started with also using CatBoost, but in the end we found the CV scores were not on par with those of XGBoost and LGBM</li>\n<li>We used DART with LGBM thanks to the work of Martin (but not with XGBoost; incredibly slow!)</li>\n<li>Our only novelty with XGBoost was using the <code>num_parallel_tree=7</code> parameter, which slowed down the training by around 7 times, but made for a smoother training curve and a CV with a lower standard deviation.</li>\n<li>As per the magnificent notbook by Chris Deotte <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-79391\" target=\"_blank\">[5]</a> we used surprisingly  shallow estimators (<code>max_depth=4</code>) as going deeper didn't really seem to help</li>\n<li>Combined multiple runs of each estimator, each using different CV splitting seeds</li>\n<li>Although TabNet alone did not have a spectacular CV score despite the best efforts of my team-mate <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> , it helped every ensemble it touched.</li>\n</ul>\n<p>The OOF predictions of these estimators were then calibrated and then fed into a stacking ensemble.</p>\n<p>We were also aware that the distribution of some of the features in the test data were significantly different from those in the training data (some wonderful notebooks on this were made on this by Pavel Vodolazov <a href=\"https://www.kaggle.com/code/pavelvod/amex-eda-revealing-time-patterns-of-features\" target=\"_blank\">[6]</a><a href=\"https://www.kaggle.com/code/pavelvod/amex-eda-even-more-insane-time-patterns-revealed\" target=\"_blank\">[7]</a>). In view of the fact that usually tree based estimators cannot extrapolate we tried out the <a href=\"https://github.com/cerlymarco/linear-tree\" target=\"_blank\"><code>linear-tree</code></a>  estimator, and also the <code>gblinear</code> booster in XGBoost, but neither were used in the end; they may well have eventually performed better on the test data for the features that had covariate shifts, but there was little way of telling beforehand. </p>\n<p>All in all a very enjoyable competition and many thanks to AmEx for providing us with tabular data to play with, and kaggle for the computational resources!</p>\n<p>All the best,\ncarl</p>",
      "rawMarkdown": "First and foremost we would like to thank Raddar for his cleaned and compressed dataset [[1]](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) (which we used) and Martin for his two magnificent DART notebooks [[2]](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963)[[3]](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977). We finally used a stacking of XGBoost, LightGBM and TabNet. When blending our solution with that of Martin we obtained a Private score of 0.80761, and playing the hypothetical *what if* game this would have landed around 138th on the LB. However, although acutely aware that our solution was not competitive, we still went ahead with selecting only our own work, and followed the sage advice of Raddar in submitting our best CV, and (our best CV + our best LB)/2 works. Doing this left our CV score on the wrong side of the CV of the work published by Martin, and consequently way down the LB, but with no regrets whatsoever. \n\nThings perhaps of mention:\nFeature engineering: We spent an inordinate  amount of time with the features in conjunction with the 13GB GPU memory limit on kaggle, and with each new feature sadly eventually having to sacrifice in the end around 30 other features (team-mate @danielhanchen helped with that process) to stay within the notebook memory. In retrospect we should really have heavily sub-sampled the training data (if it wasn't for the noisy *D* component in the metric needing lots of data that would definitely be the way to go) in order to keep as many of these features as possible as almost all of them added just a little something (we did try the random under-sampling [RUSBoostClassifier](https://imbalanced-ensemble.readthedocs.io/en/latest/api/ensemble/_autosummary/imbalanced_ensemble.ensemble.RUSBoostClassifier.html) from [imbalanced-ensemble](https://imbalanced-ensemble.readthedocs.io/en/latest/index.html), but it did not work well).\nAs well as a lot of futher cleaning to remove outliers *etc*, things of particular note were:\n* The features `S_8` and `S_13` seemed reminiscent of some sort of FICO credit rating scores. With that in mind we created an additional two new binary features with a split (`np.where`) chosen by fitting just these features alone with a RandomForest 'stump' with the idea of 'good' or 'bad' creditor in mind.\n* As well as generating the usual aggregated features for all of the (up to) 13 statements, we found that also using such aggregations just over the last 3 statements, and the last 6 statements, proved to produce some very informative features.\n* We found that the feature `S_3/S_7` helped.\n* We found that the average of the three features `D_115`, `D_118` and `D_119` also helped a little\n* On the float features we used the `round(2)` trick suggested by Jiwei Liu [[4]](https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick)\n* Our final number of features after aggregations was 1145\n\nEstimators:\n* We started with also using CatBoost, but in the end we found the CV scores were not on par with those of XGBoost and LGBM\n* We used DART with LGBM thanks to the work of Martin (but not with XGBoost; incredibly slow!)\n* Our only novelty with XGBoost was using the `num_parallel_tree=7` parameter, which slowed down the training by around 7 times, but made for a smoother training curve and a CV with a lower standard deviation.\n* As per the magnificent notbook by Chris Deotte [[5]](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-79391) we used surprisingly  shallow estimators (`max_depth=4`) as going deeper didn't really seem to help\n* Combined multiple runs of each estimator, each using different CV splitting seeds\n* Although TabNet alone did not have a spectacular CV score despite the best efforts of my team-mate @optimo , it helped every ensemble it touched.\n\nThe OOF predictions of these estimators were then calibrated and then fed into a stacking ensemble.\n\nWe were also aware that the distribution of some of the features in the test data were significantly different from those in the training data (some wonderful notebooks on this were made on this by Pavel Vodolazov [[6]](https://www.kaggle.com/code/pavelvod/amex-eda-revealing-time-patterns-of-features)[[7]](https://www.kaggle.com/code/pavelvod/amex-eda-even-more-insane-time-patterns-revealed)). In view of the fact that usually tree based estimators cannot extrapolate we tried out the [`linear-tree`](https://github.com/cerlymarco/linear-tree)  estimator, and also the `gblinear` booster in XGBoost, but neither were used in the end; they may well have eventually performed better on the test data for the features that had covariate shifts, but there was little way of telling beforehand. \n\n\nAll in all a very enjoyable competition and many thanks to AmEx for providing us with tabular data to play with, and kaggle for the computational resources!\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1913413",
      "postDate": "08/25/2022 10:12:31",
      "content": "<p>great Carl!, All respect for you for selecting only your own work, I didn't have the same courage but I really think that is what kaggle should be about.</p>",
      "rawMarkdown": "great Carl!, All respect for you for selecting only your own work, I didn't have the same courage but I really think that is what kaggle should be about.",
      "votes": null
    },
    {
      "id": "1913428",
      "postDate": "08/25/2022 10:18:09",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/pabuoro\" target=\"_blank\">@pabuoro</a> </p>\n<p>Thanks! Despite the very modest result I think that, just as with science, it is useful sometimes to read about what didn't make the cut as well as the wonderful work that did!</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @pabuoro \n\nThanks! Despite the very modest result I think that, just as with science, it is useful sometimes to read about what didn't make the cut as well as the wonderful work that did!\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1913443",
      "postDate": "08/25/2022 10:25:43",
      "content": "<p>Hi Carl. thanks for the feedback ! I also spent quite some time on time aggregation (first 6 v.s. last 6) and the usage of linear model for boosting. I also didn't get much out of it.</p>",
      "rawMarkdown": "Hi Carl. thanks for the feedback ! I also spent quite some time on time aggregation (first 6 v.s. last 6) and the usage of linear model for boosting. I also didn't get much out of it.",
      "votes": null
    },
    {
      "id": "1914987",
      "postDate": "08/26/2022 15:33:14",
      "content": "<p>Wonderful collaborating with you <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>! Apologies I wasn't that helpful in the end.</p>",
      "rawMarkdown": "Wonderful collaborating with you @carlmcbrideellis and @optimo! Apologies I wasn't that helpful in the end.",
      "votes": null
    },
    {
      "id": "1915089",
      "postDate": "08/26/2022 16:49:35",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/danielhanchen\" target=\"_blank\">@danielhanchen</a> </p>\n<p>It was wonderful having you onboard! Working together as a team is so much greater than the sum of all its parts; meeting new collaborators and learning together is one of the greatest rewards there is on kaggle, or anywhere else for that matter! </p>\n<p>Un gran abrazo,<br>\ncarl</p>",
      "rawMarkdown": "Dear @danielhanchen \n\nIt was wonderful having you onboard! Working together as a team is so much greater than the sum of all its parts; meeting new collaborators and learning together is one of the greatest rewards there is on kaggle, or anywhere else for that matter! \n\nUn gran abrazo,\ncarl",
      "votes": null
    },
    {
      "id": "1915100",
      "postDate": "08/26/2022 16:58:24",
      "content": "<p>During the competition I learned some really interesting tricks for memory management. </p>\n<p>XGBoost allows batched loading which really helps (store features in CPU, load into GPU). And float16 is viable, saved in pickle format. (though I probably should've kept the most important features float32). I managed 10k features in kaggle notebook! But then didn't spend the time on generating more features so stayed at \"only\" 3k. </p>\n<p>Memory management was tough, for sure!</p>",
      "rawMarkdown": "During the competition I learned some really interesting tricks for memory management. \n\nXGBoost allows batched loading which really helps (store features in CPU, load into GPU). And float16 is viable, saved in pickle format. (though I probably should've kept the most important features float32). I managed 10k features in kaggle notebook! But then didn't spend the time on generating more features so stayed at \"only\" 3k. \n\nMemory management was tough, for sure!",
      "votes": null
    },
    {
      "id": "1915123",
      "postDate": "08/26/2022 17:20:36",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>That is really interesting to know! </p>\n<p>I did vaguely think about paying for some cloud time; apart from big ones like <code>P_2</code> most of the features contributed very little but <em>something</em>, and all together they do make a difference, so I didn't enjoy dropping any of them (even though the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">permutation importance work of AmbrosM</a> indicated that some of the features were decidedly unhelpful. Chris Deotte found that <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641#1915259\" target=\"_blank\">\"<em>There are only 35 features [of his 1366 features] that are statistically significant hurt the model.</em>\"</a>) and would have liked to keep them all and let each estimator do its work. I cant remember exactly where but I remember one winner back in the golden age of kaggle tabular competitions saying that in general he never dropped any features.</p>\n<p>So resource availability and/or management was a big factor in this competition. There was an interesting topic posted a little back by James Trotman <a href=\"https://www.kaggle.com/discussions/general/322995\" target=\"_blank\">\"Incoming Feature: Paid User Plans?\"</a> which led to some thought provoking questions about the long term future regarding the size of kaggle competitions w.r.t. kaggle resources.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @roberthatch \n\nThat is really interesting to know! \n\nI did vaguely think about paying for some cloud time; apart from big ones like `P_2` most of the features contributed very little but *something*, and all together they do make a difference, so I didn't enjoy dropping any of them (even though the [permutation importance work of AmbrosM](https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131) indicated that some of the features were decidedly unhelpful. Chris Deotte found that [\"*There are only 35 features [of his 1366 features] that are statistically significant hurt the model.*\"](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641#1915259)) and would have liked to keep them all and let each estimator do its work. I cant remember exactly where but I remember one winner back in the golden age of kaggle tabular competitions saying that in general he never dropped any features.\n\nSo resource availability and/or management was a big factor in this competition. There was an interesting topic posted a little back by James Trotman [\"Incoming Feature: Paid User Plans?\"]( https://www.kaggle.com/discussions/general/322995) which led to some thought provoking questions about the long term future regarding the size of kaggle competitions w.r.t. kaggle resources.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1915136",
      "postDate": "08/26/2022 17:37:07",
      "content": "<p>The short version of tricks I used:</p>\n<ul>\n<li>Convert subset of features to float16 in CPU memory, save in pickle format</li>\n<li>Update iterLoad___ class in xgb starter notebook to handle list of dataframes instead of single. That way if I had 9gb features, I didn't have to concat them into single dataframes and run out of memory. (2 bytes times 10,000 times 458,000 is ~9GB pure data). I never found a way to concat three 3GB dataframes into single frame in 16gb mem. Probably something like dask could've worked better(?) but I had limited time to learn yet another tool. </li>\n</ul>\n<p>And lots of smaller batch jobs and multiple notebooks. Even for 3k features, I had three separate feature engg notebooks and combined them later. </p>",
      "rawMarkdown": "The short version of tricks I used:\n* Convert subset of features to float16 in CPU memory, save in pickle format\n* Update iterLoad___ class in xgb starter notebook to handle list of dataframes instead of single. That way if I had 9gb features, I didn't have to concat them into single dataframes and run out of memory. (2 bytes times 10,000 times 458,000 is ~9GB pure data). I never found a way to concat three 3GB dataframes into single frame in 16gb mem. Probably something like dask could've worked better(?) but I had limited time to learn yet another tool. \n\nAnd lots of smaller batch jobs and multiple notebooks. Even for 3k features, I had three separate feature engg notebooks and combined them later.",
      "votes": null
    },
    {
      "id": "1915146",
      "postDate": "08/26/2022 17:51:06",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>One thing I would really like to see in the future would be kaggle tabular competitions (!) that perhaps move away from monolithic <code>csv</code> files and perhaps provide a connection to a database via PySpark. For example, <a href=\"https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html\" target=\"_blank\">as of Spark 3.2 one can use pandas</a> via <code>import pyspark.pandas as ps</code> without the need for UDFs, and would also provide an environment where people could practice their SQL skills (something that all employers want) and, thanks to lazy evaluation, may not be so taxing on notebook CPU/GPU…</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @roberthatch \n\nOne thing I would really like to see in the future would be kaggle tabular competitions (!) that perhaps move away from monolithic `csv` files and perhaps provide a connection to a database via PySpark. For example, [as of Spark 3.2 one can use pandas](https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html) via `import pyspark.pandas as ps` without the need for UDFs, and would also provide an environment where people could practice their SQL skills (something that all employers want) and, thanks to lazy evaluation, may not be so taxing on notebook CPU/GPU...\n\nAll the best,\ncarl",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913413,
      "author_name": "pabuoro",
      "author_url": "",
      "post_date": "08/25/2022 10:12:31",
      "content": "<p>great Carl!, All respect for you for selecting only your own work, I didn't have the same courage but I really think that is what kaggle should be about.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913428,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "08/25/2022 10:18:09",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/pabuoro\" target=\"_blank\">@pabuoro</a> </p>\n<p>Thanks! Despite the very modest result I think that, just as with science, it is useful sometimes to read about what didn't make the cut as well as the wonderful work that did!</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913443,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "08/25/2022 10:25:43",
      "content": "<p>Hi Carl. thanks for the feedback ! I also spent quite some time on time aggregation (first 6 v.s. last 6) and the usage of linear model for boosting. I also didn't get much out of it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914987,
      "author_name": "danielhanchen",
      "author_url": "",
      "post_date": "08/26/2022 15:33:14",
      "content": "<p>Wonderful collaborating with you <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>! Apologies I wasn't that helpful in the end.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915089,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "08/26/2022 16:49:35",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/danielhanchen\" target=\"_blank\">@danielhanchen</a> </p>\n<p>It was wonderful having you onboard! Working together as a team is so much greater than the sum of all its parts; meeting new collaborators and learning together is one of the greatest rewards there is on kaggle, or anywhere else for that matter! </p>\n<p>Un gran abrazo,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915100,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/26/2022 16:58:24",
      "content": "<p>During the competition I learned some really interesting tricks for memory management. </p>\n<p>XGBoost allows batched loading which really helps (store features in CPU, load into GPU). And float16 is viable, saved in pickle format. (though I probably should've kept the most important features float32). I managed 10k features in kaggle notebook! But then didn't spend the time on generating more features so stayed at \"only\" 3k. </p>\n<p>Memory management was tough, for sure!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915123,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "08/26/2022 17:20:36",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>That is really interesting to know! </p>\n<p>I did vaguely think about paying for some cloud time; apart from big ones like <code>P_2</code> most of the features contributed very little but <em>something</em>, and all together they do make a difference, so I didn't enjoy dropping any of them (even though the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">permutation importance work of AmbrosM</a> indicated that some of the features were decidedly unhelpful. Chris Deotte found that <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641#1915259\" target=\"_blank\">\"<em>There are only 35 features [of his 1366 features] that are statistically significant hurt the model.</em>\"</a>) and would have liked to keep them all and let each estimator do its work. I cant remember exactly where but I remember one winner back in the golden age of kaggle tabular competitions saying that in general he never dropped any features.</p>\n<p>So resource availability and/or management was a big factor in this competition. There was an interesting topic posted a little back by James Trotman <a href=\"https://www.kaggle.com/discussions/general/322995\" target=\"_blank\">\"Incoming Feature: Paid User Plans?\"</a> which led to some thought provoking questions about the long term future regarding the size of kaggle competitions w.r.t. kaggle resources.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915136,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/26/2022 17:37:07",
          "content": "<p>The short version of tricks I used:</p>\n<ul>\n<li>Convert subset of features to float16 in CPU memory, save in pickle format</li>\n<li>Update iterLoad___ class in xgb starter notebook to handle list of dataframes instead of single. That way if I had 9gb features, I didn't have to concat them into single dataframes and run out of memory. (2 bytes times 10,000 times 458,000 is ~9GB pure data). I never found a way to concat three 3GB dataframes into single frame in 16gb mem. Probably something like dask could've worked better(?) but I had limited time to learn yet another tool. </li>\n</ul>\n<p>And lots of smaller batch jobs and multiple notebooks. Even for 3k features, I had three separate feature engg notebooks and combined them later. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1915146,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "08/26/2022 17:51:06",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>One thing I would really like to see in the future would be kaggle tabular competitions (!) that perhaps move away from monolithic <code>csv</code> files and perhaps provide a connection to a database via PySpark. For example, <a href=\"https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html\" target=\"_blank\">as of Spark 3.2 one can use pandas</a> via <code>import pyspark.pandas as ps</code> without the need for UDFs, and would also provide an environment where people could practice their SQL skills (something that all employers want) and, thanks to lazy evaluation, may not be so taxing on notebook CPU/GPU…</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1913368": "First and foremost we would like to thank Raddar for his cleaned and compressed dataset [[1]](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) (which we used) and Martin for his two magnificent DART notebooks [[2]](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963)[[3]](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977). We finally used a stacking of XGBoost, LightGBM and TabNet. When blending our solution with that of Martin we obtained a Private score of 0.80761, and playing the hypothetical *what if* game this would have landed around 138th on the LB. However, although acutely aware that our solution was not competitive, we still went ahead with selecting only our own work, and followed the sage advice of Raddar in submitting our best CV, and (our best CV + our best LB)/2 works. Doing this left our CV score on the wrong side of the CV of the work published by Martin, and consequently way down the LB, but with no regrets whatsoever. \n\nThings perhaps of mention:\nFeature engineering: We spent an inordinate  amount of time with the features in conjunction with the 13GB GPU memory limit on kaggle, and with each new feature sadly eventually having to sacrifice in the end around 30 other features (team-mate @danielhanchen helped with that process) to stay within the notebook memory. In retrospect we should really have heavily sub-sampled the training data (if it wasn't for the noisy *D* component in the metric needing lots of data that would definitely be the way to go) in order to keep as many of these features as possible as almost all of them added just a little something (we did try the random under-sampling [RUSBoostClassifier](https://imbalanced-ensemble.readthedocs.io/en/latest/api/ensemble/_autosummary/imbalanced_ensemble.ensemble.RUSBoostClassifier.html) from [imbalanced-ensemble](https://imbalanced-ensemble.readthedocs.io/en/latest/index.html), but it did not work well).\nAs well as a lot of futher cleaning to remove outliers *etc*, things of particular note were:\n* The features `S_8` and `S_13` seemed reminiscent of some sort of FICO credit rating scores. With that in mind we created an additional two new binary features with a split (`np.where`) chosen by fitting just these features alone with a RandomForest 'stump' with the idea of 'good' or 'bad' creditor in mind.\n* As well as generating the usual aggregated features for all of the (up to) 13 statements, we found that also using such aggregations just over the last 3 statements, and the last 6 statements, proved to produce some very informative features.\n* We found that the feature `S_3/S_7` helped.\n* We found that the average of the three features `D_115`, `D_118` and `D_119` also helped a little\n* On the float features we used the `round(2)` trick suggested by Jiwei Liu [[4]](https://www.kaggle.com/code/jiweiliu/amex-catboost-rounding-trick)\n* Our final number of features after aggregations was 1145\n\nEstimators:\n* We started with also using CatBoost, but in the end we found the CV scores were not on par with those of XGBoost and LGBM\n* We used DART with LGBM thanks to the work of Martin (but not with XGBoost; incredibly slow!)\n* Our only novelty with XGBoost was using the `num_parallel_tree=7` parameter, which slowed down the training by around 7 times, but made for a smoother training curve and a CV with a lower standard deviation.\n* As per the magnificent notbook by Chris Deotte [[5]](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-79391) we used surprisingly  shallow estimators (`max_depth=4`) as going deeper didn't really seem to help\n* Combined multiple runs of each estimator, each using different CV splitting seeds\n* Although TabNet alone did not have a spectacular CV score despite the best efforts of my team-mate @optimo , it helped every ensemble it touched.\n\nThe OOF predictions of these estimators were then calibrated and then fed into a stacking ensemble.\n\nWe were also aware that the distribution of some of the features in the test data were significantly different from those in the training data (some wonderful notebooks on this were made on this by Pavel Vodolazov [[6]](https://www.kaggle.com/code/pavelvod/amex-eda-revealing-time-patterns-of-features)[[7]](https://www.kaggle.com/code/pavelvod/amex-eda-even-more-insane-time-patterns-revealed)). In view of the fact that usually tree based estimators cannot extrapolate we tried out the [`linear-tree`](https://github.com/cerlymarco/linear-tree)  estimator, and also the `gblinear` booster in XGBoost, but neither were used in the end; they may well have eventually performed better on the test data for the features that had covariate shifts, but there was little way of telling beforehand. \n\n\nAll in all a very enjoyable competition and many thanks to AmEx for providing us with tabular data to play with, and kaggle for the computational resources!\n\nAll the best,\ncarl",
    "1913413": "great Carl!, All respect for you for selecting only your own work, I didn't have the same courage but I really think that is what kaggle should be about.",
    "1913428": "Dear @pabuoro \n\nThanks! Despite the very modest result I think that, just as with science, it is useful sometimes to read about what didn't make the cut as well as the wonderful work that did!\n\nAll the best,\ncarl",
    "1913443": "Hi Carl. thanks for the feedback ! I also spent quite some time on time aggregation (first 6 v.s. last 6) and the usage of linear model for boosting. I also didn't get much out of it.",
    "1914987": "Wonderful collaborating with you @carlmcbrideellis and @optimo! Apologies I wasn't that helpful in the end.",
    "1915089": "Dear @danielhanchen \n\nIt was wonderful having you onboard! Working together as a team is so much greater than the sum of all its parts; meeting new collaborators and learning together is one of the greatest rewards there is on kaggle, or anywhere else for that matter! \n\nUn gran abrazo,\ncarl",
    "1915100": "During the competition I learned some really interesting tricks for memory management. \n\nXGBoost allows batched loading which really helps (store features in CPU, load into GPU). And float16 is viable, saved in pickle format. (though I probably should've kept the most important features float32). I managed 10k features in kaggle notebook! But then didn't spend the time on generating more features so stayed at \"only\" 3k. \n\nMemory management was tough, for sure!",
    "1915123": "Dear @roberthatch \n\nThat is really interesting to know! \n\nI did vaguely think about paying for some cloud time; apart from big ones like `P_2` most of the features contributed very little but *something*, and all together they do make a difference, so I didn't enjoy dropping any of them (even though the [permutation importance work of AmbrosM](https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131) indicated that some of the features were decidedly unhelpful. Chris Deotte found that [\"*There are only 35 features [of his 1366 features] that are statistically significant hurt the model.*\"](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641#1915259)) and would have liked to keep them all and let each estimator do its work. I cant remember exactly where but I remember one winner back in the golden age of kaggle tabular competitions saying that in general he never dropped any features.\n\nSo resource availability and/or management was a big factor in this competition. There was an interesting topic posted a little back by James Trotman [\"Incoming Feature: Paid User Plans?\"]( https://www.kaggle.com/discussions/general/322995) which led to some thought provoking questions about the long term future regarding the size of kaggle competitions w.r.t. kaggle resources.\n\nAll the best,\ncarl",
    "1915136": "The short version of tricks I used:\n* Convert subset of features to float16 in CPU memory, save in pickle format\n* Update iterLoad___ class in xgb starter notebook to handle list of dataframes instead of single. That way if I had 9gb features, I didn't have to concat them into single dataframes and run out of memory. (2 bytes times 10,000 times 458,000 is ~9GB pure data). I never found a way to concat three 3GB dataframes into single frame in 16gb mem. Probably something like dask could've worked better(?) but I had limited time to learn yet another tool. \n\nAnd lots of smaller batch jobs and multiple notebooks. Even for 3k features, I had three separate feature engg notebooks and combined them later.",
    "1915146": "Dear @roberthatch \n\nOne thing I would really like to see in the future would be kaggle tabular competitions (!) that perhaps move away from monolithic `csv` files and perhaps provide a connection to a database via PySpark. For example, [as of Spark 3.2 one can use pandas](https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html) via `import pyspark.pandas as ps` without the need for UDFs, and would also provide an environment where people could practice their SQL skills (something that all employers want) and, thanks to lazy evaluation, may not be so taxing on notebook CPU/GPU...\n\nAll the best,\ncarl"
  },
  "source": "meta"
}