{
  "id": 336546,
  "title": "Sharing my ablations studies",
  "url": "/competitions/amex-default-prediction/discussion/336546",
  "author_name": "Andrea Cognolato",
  "post_date": "2022-07-11T19:01:43.079000",
  "votes": 63,
  "comment_count": 27,
  "views": 0,
  "content": "<p>2022-07-26 update: see the code in this public notebook <a href=\"https://www.kaggle.com/code/mrandri19/0-796-lb-sharing-my-ablations-dart-less-xgb\" target=\"_blank\">https://www.kaggle.com/code/mrandri19/0-796-lb-sharing-my-ablations-dart-less-xgb</a></p>\n<p>I am trying to achieve a single DART-less model with &gt;0.795 CV/LB score.<br>\nI am happy with the model and hyperparameters so I am working on feature engineering and feature selection.</p>\n<p>All the results have been obtained via sklearn's 5-fold StratifiedKFold cross-validation with seed=42.<br>\nKeep in mind that cdeotte says that the standard deviation of a model trained on different seeds is 0.0012.<br>\nWith a standard deviation of this magnitude, I don't feel confident in differences smaller than 0.0024, so take my results with a pinch of salt.</p>\n<p>Below is the description of the model, hyperparameters and the ablation studies I have ran so far:</p>\n<h2>Model</h2>\n<p>XGBoost trained with early stopping using the amex metric.<br>\nIn particular I use 9999 boosting rounds and 500 early stopping rounds.<br>\nEach model takes 10-20 minutes to train on a Kaggle GPU instance.</p>\n<h3>Hyperparameters</h3>\n<p>Taken from the excellent notebook <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>.</p>\n<pre><code>'max_depth': 7,\n'eta': 0.03,\n\n'subsample': 0.88,\n'colsample_bytree': 0.5,\n\n'objective': 'binary:logistic',\n\n'tree_method': 'gpu_hist',\n\n'random_state': 42,\n\n'gamma': 1.5,\n'min_child_weight': 8,\n'lambda': 70,\n</code></pre>\n<h2>Ablation studies</h2>\n<h3>(a) Baseline</h3>\n<p>The features, selecting the last month of data for each customer, minus customer_ID and S_2, fed directly to the model.</p>\n<ul>\n<li><strong>CV: 0.7885</strong></li>\n<li>9.5min training time</li>\n<li>188 features</li>\n</ul>\n<h3>(b) Round to 2 decimal digits</h3>\n<p>As raddar suggests, there is uniform noise injected to the data, we can ignore the noise by throwing signals smaller<br>\nthan 0.01 by rounding to 2 decimal digits.</p>\n<ul>\n<li><strong>CV: 0.7908, +0.0023</strong> compared with (a)</li>\n<li>9.5min training time</li>\n<li>188 features</li>\n</ul>\n<h3>(c) Drop correlated features</h3>\n<p>After filling the missing values with the mean, compute the pairwise pearson correlation between all 188 features.<br>\nThen drop features that are &gt;95% correlated with others.</p>\n<ul>\n<li><strong>CV: 0.7905, -0.0003</strong> compared with (b)</li>\n<li>10.5min training time</li>\n<li>178 features</li>\n</ul>\n<h3>(d) Add customer aggregated features</h3>\n<p>So fare we have only used the last value from each feature, let's use the first, last, mean, std, min, max as well.</p>\n<ul>\n<li><strong>CV: 0.7949, +0.0041</strong> compared with (b)</li>\n<li>21.5min training time</li>\n<li>1073 features</li>\n</ul>\n<h3>(e) Add customer aggregated features, then drop correlated features</h3>\n<p>Since we have a lot more features, we can again try to drop correlated ones.</p>\n<ul>\n<li><strong>CV: 0.7950, +0.0001</strong> compared with (d)</li>\n<li>22.5min training time</li>\n<li>875 features</li>\n</ul>\n<h3>(f) Add missing count features</h3>\n<p>We can count the number of missing values in each column, for every customer.<br>\nFor every column, this is an integer in [0, 13].</p>\n<ul>\n<li><strong>CV: 0.7946, -0.0003</strong> compared with (d)</li>\n<li>24.5min training time</li>\n<li>1261 features</li>\n</ul>\n<h3>(g) Add difference features</h3>\n<p>Replicate the feature enginering performed by Jiwei Liu in this notebook: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>.<br>\nIn practice, compute the subtraction of a bunch of B, D, S features minus two P (Payment) features: P_2, P_3</p>\n<ul>\n<li><strong>CV: 0.7953, +0.0004</strong> compared with (d)</li>\n<li>20.9min training time</li>\n<li>1157 features</li>\n</ul>\n<h3>(h) Add lag features</h3>\n<p>Replicate the feature engineering by thedevastator described in described in<br>\n<a href=\"https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need</a></p>\n<ul>\n<li><strong>CV: 0.7937, -0.0016</strong> compared with (g)</li>\n<li>24.2min training time</li>\n<li>1511 features</li>\n</ul>\n<p>This is very weird, as these features come from one of the best publicly available notebooks.<br>\nMy hypothesis is that too many (&gt;1500) features combined with an aggressive 0.5 colsample_bytree leads to<br>\nsome the model missing some interactions, but I am curious of the reader's opinion.</p>\n<h3>(i) Illidan7's Feature Set</h3>\n<p><a href=\"https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features</a></p>\n<ul>\n<li><strong>CV: 0.7950</strong></li>\n<li>41.2min training time</li>\n<li>1486 features</li>\n</ul>\n<h3>(j) Add ragnar123' first - mean lag feature</h3>\n<ul>\n<li><strong>CV: 0.7956</strong></li>\n<li>27.9min training time</li>\n<li>974 features</li>\n</ul>\n<h2>Conclusions and next steps</h2>\n<ul>\n<li>Rounding seems to be beneficial, should we round the values before performing the aggreagations, then?</li>\n<li>Adding first, mean, std, min, max aggregations has the biggest impact on performance, but they are almost 900 additional features. How to identify the ones that actually bring benefits?</li>\n<li>It is not clear if dropping correlated features has an effect. Why doesn't it reduce the training time? Is it because we need more trees?</li>\n<li>The difference features give a small (4 basis points) boost. I believe they are worth using as they are only 40.</li>\n<li>Next, I want to check 's <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></li>\n</ul>",
  "messages": [
    {
      "id": 1852093,
      "postDate": "2022-07-11T19:01:43.080Z",
      "content": "<p>2022-07-26 update: see the code in this public notebook <a href=\"https://www.kaggle.com/code/mrandri19/0-796-lb-sharing-my-ablations-dart-less-xgb\" target=\"_blank\">https://www.kaggle.com/code/mrandri19/0-796-lb-sharing-my-ablations-dart-less-xgb</a></p>\n<p>I am trying to achieve a single DART-less model with &gt;0.795 CV/LB score.<br>\nI am happy with the model and hyperparameters so I am working on feature engineering and feature selection.</p>\n<p>All the results have been obtained via sklearn's 5-fold StratifiedKFold cross-validation with seed=42.<br>\nKeep in mind that cdeotte says that the standard deviation of a model trained on different seeds is 0.0012.<br>\nWith a standard deviation of this magnitude, I don't feel confident in differences smaller than 0.0024, so take my results with a pinch of salt.</p>\n<p>Below is the description of the model, hyperparameters and the ablation studies I have ran so far:</p>\n<h2>Model</h2>\n<p>XGBoost trained with early stopping using the amex metric.<br>\nIn particular I use 9999 boosting rounds and 500 early stopping rounds.<br>\nEach model takes 10-20 minutes to train on a Kaggle GPU instance.</p>\n<h3>Hyperparameters</h3>\n<p>Taken from the excellent notebook <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>.</p>\n<pre><code>'max_depth': 7,\n'eta': 0.03,\n\n'subsample': 0.88,\n'colsample_bytree': 0.5,\n\n'objective': 'binary:logistic',\n\n'tree_method': 'gpu_hist',\n\n'random_state': 42,\n\n'gamma': 1.5,\n'min_child_weight': 8,\n'lambda': 70,\n</code></pre>\n<h2>Ablation studies</h2>\n<h3>(a) Baseline</h3>\n<p>The features, selecting the last month of data for each customer, minus customer_ID and S_2, fed directly to the model.</p>\n<ul>\n<li><strong>CV: 0.7885</strong></li>\n<li>9.5min training time</li>\n<li>188 features</li>\n</ul>\n<h3>(b) Round to 2 decimal digits</h3>\n<p>As raddar suggests, there is uniform noise injected to the data, we can ignore the noise by throwing signals smaller<br>\nthan 0.01 by rounding to 2 decimal digits.</p>\n<ul>\n<li><strong>CV: 0.7908, +0.0023</strong> compared with (a)</li>\n<li>9.5min training time</li>\n<li>188 features</li>\n</ul>\n<h3>(c) Drop correlated features</h3>\n<p>After filling the missing values with the mean, compute the pairwise pearson correlation between all 188 features.<br>\nThen drop features that are &gt;95% correlated with others.</p>\n<ul>\n<li><strong>CV: 0.7905, -0.0003</strong> compared with (b)</li>\n<li>10.5min training time</li>\n<li>178 features</li>\n</ul>\n<h3>(d) Add customer aggregated features</h3>\n<p>So fare we have only used the last value from each feature, let's use the first, last, mean, std, min, max as well.</p>\n<ul>\n<li><strong>CV: 0.7949, +0.0041</strong> compared with (b)</li>\n<li>21.5min training time</li>\n<li>1073 features</li>\n</ul>\n<h3>(e) Add customer aggregated features, then drop correlated features</h3>\n<p>Since we have a lot more features, we can again try to drop correlated ones.</p>\n<ul>\n<li><strong>CV: 0.7950, +0.0001</strong> compared with (d)</li>\n<li>22.5min training time</li>\n<li>875 features</li>\n</ul>\n<h3>(f) Add missing count features</h3>\n<p>We can count the number of missing values in each column, for every customer.<br>\nFor every column, this is an integer in [0, 13].</p>\n<ul>\n<li><strong>CV: 0.7946, -0.0003</strong> compared with (d)</li>\n<li>24.5min training time</li>\n<li>1261 features</li>\n</ul>\n<h3>(g) Add difference features</h3>\n<p>Replicate the feature enginering performed by Jiwei Liu in this notebook: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>.<br>\nIn practice, compute the subtraction of a bunch of B, D, S features minus two P (Payment) features: P_2, P_3</p>\n<ul>\n<li><strong>CV: 0.7953, +0.0004</strong> compared with (d)</li>\n<li>20.9min training time</li>\n<li>1157 features</li>\n</ul>\n<h3>(h) Add lag features</h3>\n<p>Replicate the feature engineering by thedevastator described in described in<br>\n<a href=\"https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need</a></p>\n<ul>\n<li><strong>CV: 0.7937, -0.0016</strong> compared with (g)</li>\n<li>24.2min training time</li>\n<li>1511 features</li>\n</ul>\n<p>This is very weird, as these features come from one of the best publicly available notebooks.<br>\nMy hypothesis is that too many (&gt;1500) features combined with an aggressive 0.5 colsample_bytree leads to<br>\nsome the model missing some interactions, but I am curious of the reader's opinion.</p>\n<h3>(i) Illidan7's Feature Set</h3>\n<p><a href=\"https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features</a></p>\n<ul>\n<li><strong>CV: 0.7950</strong></li>\n<li>41.2min training time</li>\n<li>1486 features</li>\n</ul>\n<h3>(j) Add ragnar123' first - mean lag feature</h3>\n<ul>\n<li><strong>CV: 0.7956</strong></li>\n<li>27.9min training time</li>\n<li>974 features</li>\n</ul>\n<h2>Conclusions and next steps</h2>\n<ul>\n<li>Rounding seems to be beneficial, should we round the values before performing the aggreagations, then?</li>\n<li>Adding first, mean, std, min, max aggregations has the biggest impact on performance, but they are almost 900 additional features. How to identify the ones that actually bring benefits?</li>\n<li>It is not clear if dropping correlated features has an effect. Why doesn't it reduce the training time? Is it because we need more trees?</li>\n<li>The difference features give a small (4 basis points) boost. I believe they are worth using as they are only 40.</li>\n<li>Next, I want to check 's <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></li>\n</ul>",
      "rawMarkdown": "2022-07-26 update: see the code in this public notebook https://www.kaggle.com/code/mrandri19/0-796-lb-sharing-my-ablations-dart-less-xgb\n\nI am trying to achieve a single DART-less model with >0.795 CV/LB score.\nI am happy with the model and hyperparameters so I am working on feature engineering and feature selection.\n\nAll the results have been obtained via sklearn's 5-fold StratifiedKFold cross-validation with seed=42.\nKeep in mind that cdeotte says that the standard deviation of a model trained on different seeds is 0.0012.\nWith a standard deviation of this magnitude, I don't feel confident in differences smaller than 0.0024, so take my results with a pinch of salt.\n\nBelow is the description of the model, hyperparameters and the ablation studies I have ran so far:\n\n## Model\n\nXGBoost trained with early stopping using the amex metric.\nIn particular I use 9999 boosting rounds and 500 early stopping rounds.\nEach model takes 10-20 minutes to train on a Kaggle GPU instance.\n\n### Hyperparameters\n\nTaken from the excellent notebook https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb.\n\n```\n'max_depth': 7,\n'eta': 0.03,\n\n'subsample': 0.88,\n'colsample_bytree': 0.5,\n        \n'objective': 'binary:logistic',\n        \n'tree_method': 'gpu_hist',\n        \n'random_state': 42,\n        \n'gamma': 1.5,\n'min_child_weight': 8,\n'lambda': 70,\n```\n\n## Ablation studies\n\n### (a) Baseline\n\nThe features, selecting the last month of data for each customer, minus customer_ID and S_2, fed directly to the model.\n- **CV: 0.7885**\n- 9.5min training time\n- 188 features\n\n### (b) Round to 2 decimal digits\n\nAs raddar suggests, there is uniform noise injected to the data, we can ignore the noise by throwing signals smaller\nthan 0.01 by rounding to 2 decimal digits.\n- **CV: 0.7908, +0.0023** compared with (a)\n- 9.5min training time\n- 188 features\n\n### (c) Drop correlated features\n\nAfter filling the missing values with the mean, compute the pairwise pearson correlation between all 188 features.\nThen drop features that are >95% correlated with others.\n\n- **CV: 0.7905, -0.0003** compared with (b)\n- 10.5min training time\n- 178 features\n\n### (d) Add customer aggregated features\n\nSo fare we have only used the last value from each feature, let's use the first, last, mean, std, min, max as well.\n\n- **CV: 0.7949, +0.0041** compared with (b)\n- 21.5min training time\n- 1073 features\n\n### (e) Add customer aggregated features, then drop correlated features\n\nSince we have a lot more features, we can again try to drop correlated ones.\n\n- **CV: 0.7950, +0.0001** compared with (d)\n- 22.5min training time\n- 875 features\n\n### (f) Add missing count features\n\nWe can count the number of missing values in each column, for every customer.\nFor every column, this is an integer in [0, 13].\n\n- **CV: 0.7946, -0.0003** compared with (d)\n- 24.5min training time\n- 1261 features\n\n### (g) Add difference features\n\nReplicate the feature enginering performed by Jiwei Liu in this notebook: https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb.\nIn practice, compute the subtraction of a bunch of B, D, S features minus two P (Payment) features: P_2, P_3\n\n- **CV: 0.7953, +0.0004** compared with (d)\n- 20.9min training time\n- 1157 features\n\n### (h) Add lag features\n\nReplicate the feature engineering by thedevastator described in described in\nhttps://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need\n\n- **CV: 0.7937, -0.0016** compared with (g)\n- 24.2min training time\n- 1511 features\n\nThis is very weird, as these features come from one of the best publicly available notebooks.\nMy hypothesis is that too many (>1500) features combined with an aggressive 0.5 colsample_bytree leads to\nsome the model missing some interactions, but I am curious of the reader's opinion.\n\n### (i) Illidan7's Feature Set\n\nhttps://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\n\n- **CV: 0.7950**\n- 41.2min training time\n- 1486 features\n\n### (j) Add ragnar123' first - mean lag feature\n\n- **CV: 0.7956**\n- 27.9min training time\n- 974 features\n\n\n## Conclusions and next steps\n\n- Rounding seems to be beneficial, should we round the values before performing the aggreagations, then?\n- Adding first, mean, std, min, max aggregations has the biggest impact on performance, but they are almost 900 additional features. How to identify the ones that actually bring benefits?\n- It is not clear if dropping correlated features has an effect. Why doesn't it reduce the training time? Is it because we need more trees?\n- The difference features give a small (4 basis points) boost. I believe they are worth using as they are only 40.\n- Next, I want to check 's https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977",
      "votes": 62
    },
    {
      "id": 1864319,
      "postDate": "2022-07-21T00:20:10.483Z",
      "content": "<p>Hi Andrea, update on my end: I went live with my XGB notebook today: <a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions</a> </p>\n<p>CV 0.7968 (~0.7965-7970); LB 0.797 and edged out other 0.797 public notebooks.</p>\n<p>The hyper-parameters aren't all that optimized, the feature engineering is unoptimized and pretty untested. But I'm <em>very</em> happy with the results of my \"Pyramid\" experiment.</p>\n<p>Essentially, I tried using other methods besides DART that I suspected might reduce \"over-specialization\", without the performance penalty that DART + XGB + GPU has. It was very successful, using two or three tricks, depending how you count it:</p>\n<ul>\n<li>Boosted forest early on, very directly reduces impact of initial tree, and a little more evenly spreads even at later points (with forest size of 100, then tree 101 has same impact as tree 200).</li>\n<li>Rescale the prediction results retroactively to give more 'room' for impact from later trees.</li>\n<li>These two innovations supported by the basic method of stopping a model early, and using it as the basis of the next model. There might be a more elegant way to do that, but the one I found that definitely works is to input the predictions from model A before starting training model B using \"set_base_margin\".</li>\n</ul>",
      "rawMarkdown": "Hi Andrea, update on my end: I went live with my XGB notebook today: https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions \n\nCV 0.7968 (~0.7965-7970); LB 0.797 and edged out other 0.797 public notebooks.\n\nThe hyper-parameters aren't all that optimized, the feature engineering is unoptimized and pretty untested. But I'm *very* happy with the results of my \"Pyramid\" experiment.\n\nEssentially, I tried using other methods besides DART that I suspected might reduce \"over-specialization\", without the performance penalty that DART + XGB + GPU has. It was very successful, using two or three tricks, depending how you count it:\n* Boosted forest early on, very directly reduces impact of initial tree, and a little more evenly spreads even at later points (with forest size of 100, then tree 101 has same impact as tree 200).\n* Rescale the prediction results retroactively to give more 'room' for impact from later trees.\n* These two innovations supported by the basic method of stopping a model early, and using it as the basis of the next model. There might be a more elegant way to do that, but the one I found that definitely works is to input the predictions from model A before starting training model B using \"set_base_margin\".",
      "votes": 1
    },
    {
      "id": 1853206,
      "postDate": "2022-07-12T17:24:21.607Z",
      "content": "<p>Great work, <a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.</p>",
      "rawMarkdown": "Great work, @mrandri19 , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.",
      "votes": 1,
      "replies": [
        {
          "id": 1853290,
          "postDate": "2022-07-12T19:24:19.427Z",
          "content": "<p>Humm, that's very interesting, thanks for sharing.<br>\nSo you tried ragnar123's features (last-first, last-avg) but also did not see any improvement.<br>\nI guess I should bite the bullet and try using DART (perhaps with LGBM as I hear it is faster for than XGB's DART).</p>",
          "rawMarkdown": "Humm, that's very interesting, thanks for sharing.\nSo you tried ragnar123's features (last-first, last-avg) but also did not see any improvement.\nI guess I should bite the bullet and try using DART (perhaps with LGBM as I hear it is faster for than XGB's DART).",
          "votes": 1
        },
        {
          "id": 1853982,
          "postDate": "2022-07-13T11:00:05.670Z",
          "content": "<p>It's true. XGB's DART is veeeery slow. I spent more than 2 hours learning 1 of 5 fold, but didn't wait for the end…</p>",
          "rawMarkdown": "It's true. XGB's DART is veeeery slow. I spent more than 2 hours learning 1 of 5 fold, but didn't wait for the end...",
          "votes": 1
        },
        {
          "id": 1854491,
          "postDate": "2022-07-13T18:28:52.467Z",
          "content": "<p>Dart vs non-Dart:</p>\n<p>With my own dataset, running everything in Kaggle, I got the following:</p>\n<ul>\n<li>XGB GPU non-Dart (varying seed and slight tweaks to hyper-params on different runs): CV 0.7951-0.7959, one hour training (~13 min per fold)</li>\n<li>LGBM CPU Dart: CV 0.7965, 5 hours training each fold in parallel, plus about an hour test inference using the 5 model outputs.</li>\n</ul>\n<p>I'd be careful coming to any conclusions on a single test, but, if not caring about CPU vs GPU, I wouldn't be too shocked if crazy low learning rate and row and col_sample could get non-Dart CV equal to Dart in similar time of 5 hours (60 min per fold) vs 5 CPUs doing Dart at once at 5 hours per fold. Shrug.</p>",
          "rawMarkdown": "Dart vs non-Dart:\n\nWith my own dataset, running everything in Kaggle, I got the following:\n* XGB GPU non-Dart (varying seed and slight tweaks to hyper-params on different runs): CV 0.7951-0.7959, one hour training (~13 min per fold)\n* LGBM CPU Dart: CV 0.7965, 5 hours training each fold in parallel, plus about an hour test inference using the 5 model outputs.\n\nI'd be careful coming to any conclusions on a single test, but, if not caring about CPU vs GPU, I wouldn't be too shocked if crazy low learning rate and row and col_sample could get non-Dart CV equal to Dart in similar time of 5 hours (60 min per fold) vs 5 CPUs doing Dart at once at 5 hours per fold. Shrug."
        },
        {
          "id": 1855640,
          "postDate": "2022-07-14T18:41:43.350Z",
          "content": "<p>I had a thought inspired by this sub thread:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337160\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/337160</a></p>\n<p>Basically, the issue with any two last-[thing], is that you are probably adding some highly correlated columns. This will hurt tree diversity. To deal with that, you need to drop sampling way down. But 'last' features are crucial… so let's double them(!). </p>\n<p>Oh wait, top notebook does both those things with 20% feature sampling and rounded last in addition to last. Hmm!</p>",
          "rawMarkdown": "I had a thought inspired by this sub thread:\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/337160\n\nBasically, the issue with any two last-[thing], is that you are probably adding some highly correlated columns. This will hurt tree diversity. To deal with that, you need to drop sampling way down. But 'last' features are crucial... so let's double them(!). \n\nOh wait, top notebook does both those things with 20% feature sampling and rounded last in addition to last. Hmm!",
          "votes": 1
        },
        {
          "id": 1858043,
          "postDate": "2022-07-16T15:50:20.847Z",
          "content": "<p><a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> I want to brag my new best results - XGB gbt CV(5 folds) 0.79662 (seed 42) LB 0.797. It was so difficult😅</p>",
          "rawMarkdown": "@mrandri19 I want to brag my new best results - XGB gbt CV(5 folds) 0.79662 (seed 42) LB 0.797. It was so difficult😅",
          "votes": 1
        },
        {
          "id": 1858162,
          "postDate": "2022-07-16T17:56:12.680Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1858755,
      "postDate": "2022-07-17T07:22:10.297Z",
      "content": "<p>The information provided is so clear and easy to understand, Awesome post :)<br>\nLooking forward to see more !!!<br>\nlet's grow together.</p>",
      "rawMarkdown": "The information provided is so clear and easy to understand, Awesome post :)\nLooking forward to see more !!!\nlet's grow together.",
      "votes": -3
    },
    {
      "id": 1857384,
      "postDate": "2022-07-16T05:13:45.507Z",
      "content": "<p>Great work, <a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.</p>",
      "rawMarkdown": "Great work, @mrandri19 , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.",
      "votes": -4
    },
    {
      "id": 1862506,
      "postDate": "2022-07-19T18:52:52.693Z",
      "content": "<p><a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> How are your research progressing? A week ago I was convinced that using DART is very early and it was true) At the moment I have 0.7973 CV(5) and 0.797 LB! And you know what? This is not the limit! I think the possible best CV result is around 0.7975 for XGB gbt.</p>",
      "rawMarkdown": "@mrandri19 How are your research progressing? A week ago I was convinced that using DART is very early and it was true) At the moment I have 0.7973 CV(5) and 0.797 LB! And you know what? This is not the limit! I think the possible best CV result is around 0.7975 for XGB gbt.",
      "replies": [
        {
          "id": 1872318,
          "postDate": "2022-07-26T20:17:51.517Z",
          "content": "<p><a href=\"https://www.kaggle.com/dmitryuarov\" target=\"_blank\">@dmitryuarov</a> I am slowly progressing on feature engineering. Right now I have:</p>\n<ul>\n<li>DART-less XGB 0.7951 CV 0.796 LB</li>\n<li>DART-less LGBM 0.7953 CV 0.796 LB</li>\n<li>(Work-In-Progress) NN 0.78XX CV (haven't submitted yet to LB and did not try with all folds as I still am iterating quickly)</li>\n</ul>\n<p>The biggest boosts have come from last, count, nunique for categoricals and last - mean.</p>",
          "rawMarkdown": "@dmitryuarov I am slowly progressing on feature engineering. Right now I have:\n\n- DART-less XGB 0.7951 CV 0.796 LB\n- DART-less LGBM 0.7953 CV 0.796 LB\n- (Work-In-Progress) NN 0.78XX CV (haven't submitted yet to LB and did not try with all folds as I still am iterating quickly)\n\nThe biggest boosts have come from last, count, nunique for categoricals and last - mean."
        }
      ]
    },
    {
      "id": 1861261,
      "postDate": "2022-07-18T22:42:10.480Z",
      "content": "<p>Thank you for sharing your ablations studies ! </p>",
      "rawMarkdown": "Thank you for sharing your ablations studies ! \n"
    },
    {
      "id": 1856280,
      "postDate": "2022-07-15T08:31:03.007Z",
      "content": "<p>thanks for sharing. on the std of the models, i wonder what std you get from 5 folds? mine is about 0.004. </p>",
      "rawMarkdown": "thanks for sharing. on the std of the models, i wonder what std you get from 5 folds? mine is about 0.004. ",
      "replies": [
        {
          "id": 1860812,
          "postDate": "2022-07-18T15:25:50.227Z",
          "content": "<p>around 0.002 or 0.003</p>",
          "rawMarkdown": "around 0.002 or 0.003"
        }
      ]
    },
    {
      "id": 1853608,
      "postDate": "2022-07-13T02:10:13.627Z",
      "content": "<p>In my experience (well, experiments), adding even more good features only really helps once you decrease colsample even lower AND lower Learning Rate. (0.01 eta and 35% col for me with ~1300 features)</p>",
      "rawMarkdown": "In my experience (well, experiments), adding even more good features only really helps once you decrease colsample even lower AND lower Learning Rate. (0.01 eta and 35% col for me with ~1300 features)",
      "replies": [
        {
          "id": 1853695,
          "postDate": "2022-07-13T04:41:02.183Z",
          "content": "<p>Re: my experiments and re your quote: \"I am trying to achieve a single DART-less model with &gt;0.795 CV/LB score.\"</p>\n<p>I've got: 0.79699 LB and 0.79557 CV. 74min train time for 5 folds.</p>\n<p>All XGB with feature engg and a only a little hyper param tinkering. Within the LB 0.796s, I'm listed first, so my score is about 0.79699.</p>\n<p>To get that score, I used the exact features in the current version of my public feature engg notebook: <a href=\"https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a></p>\n<p>And there's some interesting features you might want to try, but, as mentioned in above comment, I wouldn't be surprised if what you really need is some hyper-parameter tweaks and those alone get you your goal of &gt; 0.795 CV/LB</p>\n<p>Hyper-params:<br>\nExcept as below, should be identical to: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>.</p>\n<ol>\n<li>After getting good results with 0.01 learning rate, I tried this and got nearly the same train time and result, but marginally better in both, iirc. 'num_parallel_tree':3, 'learning_rate':0.03,<br>\n<em>EDIT</em>: if using num_parallel_tree, oddly enough you have to modify model.best_ntree_limit by adding integer divide by num_parallel_tree. Example:<br>\nmodel.predict(dvalid, iteration_range=(0,model.best_ntree_limit//TREE_MULTIPLE))</li>\n<li>'objective':'rank:pairwise', 'eval_metric':'auc',</li>\n<li>Sampling:    'subsample':0.75,    'colsample_bytree': 0.20,<br>\nNote that I think 35% cols gave better CV results, but I didn't create submission csv on that 0.7959 CV run, and haven't had GPU time to go back and fine-tune.</li>\n</ol>\n<p>Features engg besides the stuff in other notebooks and mentioned above.<br>\n(Again, from here: <a href=\"https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a> ) </p>\n<ol>\n<li>Simplified Hull Moving Average</li>\n<li>Magnitude = max-min. (And dropped min and max)</li>\n<li>CurrentLevel = (last-min) / (max-min)</li>\n</ol>\n<p>But all my experiments were about as unscientific as yours were scientific. Constantly changing 3 variables at once :)</p>\n<p>Let me know if you try any of the above, whether params or features!</p>\n<p>P.s.: planning to share my full notebook publicly, not just the feature engg, but working on an even higher score, with one additional innovation… and having some major pathfinding issues to sort out.</p>",
          "rawMarkdown": "Re: my experiments and re your quote: \"I am trying to achieve a single DART-less model with >0.795 CV/LB score.\"\n\nI've got: 0.79699 LB and 0.79557 CV. 74min train time for 5 folds.\n\nAll XGB with feature engg and a only a little hyper param tinkering. Within the LB 0.796s, I'm listed first, so my score is about 0.79699.\n\nTo get that score, I used the exact features in the current version of my public feature engg notebook: https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks\n\nAnd there's some interesting features you might want to try, but, as mentioned in above comment, I wouldn't be surprised if what you really need is some hyper-parameter tweaks and those alone get you your goal of > 0.795 CV/LB\n\nHyper-params:\nExcept as below, should be identical to: https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb.\n1. After getting good results with 0.01 learning rate, I tried this and got nearly the same train time and result, but marginally better in both, iirc. 'num_parallel_tree':3, 'learning_rate':0.03,\n*EDIT*: if using num_parallel_tree, oddly enough you have to modify model.best_ntree_limit by adding integer divide by num_parallel_tree. Example:\n  model.predict(dvalid, iteration_range=(0,model.best_ntree_limit//TREE_MULTIPLE))\n2. 'objective':'rank:pairwise', 'eval_metric':'auc',\n3. Sampling:    'subsample':0.75,    'colsample_bytree': 0.20,\nNote that I think 35% cols gave better CV results, but I didn't create submission csv on that 0.7959 CV run, and haven't had GPU time to go back and fine-tune.\n\nFeatures engg besides the stuff in other notebooks and mentioned above.\n(Again, from here: https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks ) \n1. Simplified Hull Moving Average\n2. Magnitude = max-min. (And dropped min and max)\n3. CurrentLevel = (last-min) / (max-min)\n\nBut all my experiments were about as unscientific as yours were scientific. Constantly changing 3 variables at once :)\n\nLet me know if you try any of the above, whether params or features!\n\nP.s.: planning to share my full notebook publicly, not just the feature engg, but working on an even higher score, with one additional innovation... and having some major pathfinding issues to sort out.",
          "votes": 2
        },
        {
          "id": 1854245,
          "postDate": "2022-07-13T14:35:09.800Z",
          "content": "<p>Wow, this is great thanks! I will try some of your suggestions and update you.</p>",
          "rawMarkdown": "Wow, this is great thanks! I will try some of your suggestions and update you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1853600,
      "postDate": "2022-07-13T01:54:28.247Z",
      "content": "<p>Thanks for sharing your ablation studies! <br>\nAre you using all of the features in the dataset? If not, what feature selection method do you use?</p>",
      "rawMarkdown": "Thanks for sharing your ablation studies! \nAre you using all of the features in the dataset? If not, what feature selection method do you use?\n\n",
      "replies": [
        {
          "id": 1854242,
          "postDate": "2022-07-13T14:33:34.133Z",
          "content": "<p>Yes, I am using all of the features.<br>\nI would like to try feature selection using feature importances but I did not have time yet/am not sure what to use among SHAP/Permutation Importance/Null Importance.</p>",
          "rawMarkdown": "Yes, I am using all of the features.\nI would like to try feature selection using feature importances but I did not have time yet/am not sure what to use among SHAP/Permutation Importance/Null Importance."
        }
      ]
    },
    {
      "id": 1853486,
      "postDate": "2022-07-12T23:20:23.360Z",
      "content": "<p>great work, thanks for sharing your ideas!</p>",
      "rawMarkdown": "great work, thanks for sharing your ideas!"
    },
    {
      "id": 1853294,
      "postDate": "2022-07-12T19:31:55.050Z",
      "content": "<p>Nice .. did lag features decrease ur CV/LB . I think one lag feature which ragner described did certainly help </p>",
      "rawMarkdown": "Nice .. did lag features decrease ur CV/LB . I think one lag feature which ragner described did certainly help ",
      "replies": [
        {
          "id": 1853302,
          "postDate": "2022-07-12T19:41:11.707Z",
          "content": "<p>Thanks :)<br>\nDo you mean the lag features (last - first, last - avg) described by ragnar or the ones (last / first, last - first) described by thedevastator?</p>",
          "rawMarkdown": "Thanks :)\nDo you mean the lag features (last - first, last - avg) described by ragnar or the ones (last / first, last - first) described by thedevastator?"
        },
        {
          "id": 1853372,
          "postDate": "2022-07-12T21:12:45.403Z",
          "content": "<p>I mean both .. I think Ragner last-Avg made sense to me the others I havent check the feature importance yet to assess .. but Ragner ones were up there in FE. </p>",
          "rawMarkdown": "I mean both .. I think Ragner last-Avg made sense to me the others I havent check the feature importance yet to assess .. but Ragner ones were up there in FE. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1852203,
      "postDate": "2022-07-11T21:38:26.727Z",
      "content": "<p>Always interesting to look at good feature engineering experiments, well done!</p>",
      "rawMarkdown": "Always interesting to look at good feature engineering experiments, well done!"
    },
    {
      "id": 1852120,
      "postDate": "2022-07-11T19:32:41.097Z",
      "content": "<p>Very well curated and summarized!</p>",
      "rawMarkdown": "Very well curated and summarized!"
    },
    {
      "id": 1857388,
      "postDate": "2022-07-16T05:14:05.723Z",
      "content": "<p>thanks for this help for us </p>",
      "rawMarkdown": "thanks for this help for us ",
      "votes": -3
    }
  ],
  "comments": [
    {
      "id": 1864319,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-07-21T00:20:10.483000",
      "content": "<p>Hi Andrea, update on my end: I went live with my XGB notebook today: <a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions</a> </p>\n<p>CV 0.7968 (~0.7965-7970); LB 0.797 and edged out other 0.797 public notebooks.</p>\n<p>The hyper-parameters aren't all that optimized, the feature engineering is unoptimized and pretty untested. But I'm <em>very</em> happy with the results of my \"Pyramid\" experiment.</p>\n<p>Essentially, I tried using other methods besides DART that I suspected might reduce \"over-specialization\", without the performance penalty that DART + XGB + GPU has. It was very successful, using two or three tricks, depending how you count it:</p>\n<ul>\n<li>Boosted forest early on, very directly reduces impact of initial tree, and a little more evenly spreads even at later points (with forest size of 100, then tree 101 has same impact as tree 200).</li>\n<li>Rescale the prediction results retroactively to give more 'room' for impact from later trees.</li>\n<li>These two innovations supported by the basic method of stopping a model early, and using it as the basis of the next model. There might be a more elegant way to do that, but the one I found that definitely works is to input the predictions from model A before starting training model B using \"set_base_margin\".</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1853206,
      "author_name": "Dmitry Uarov",
      "author_url": "",
      "post_date": "2022-07-12T17:24:21.607000",
      "content": "<p>Great work, <a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1853290,
          "author_name": "Andrea Cognolato",
          "author_url": "",
          "post_date": "2022-07-12T19:24:19.427000",
          "content": "<p>Humm, that's very interesting, thanks for sharing.<br>\nSo you tried ragnar123's features (last-first, last-avg) but also did not see any improvement.<br>\nI guess I should bite the bullet and try using DART (perhaps with LGBM as I hear it is faster for than XGB's DART).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1853982,
          "author_name": "Dmitry Uarov",
          "author_url": "",
          "post_date": "2022-07-13T11:00:05.670000",
          "content": "<p>It's true. XGB's DART is veeeery slow. I spent more than 2 hours learning 1 of 5 fold, but didn't wait for the end…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1854491,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-07-13T18:28:52.467000",
          "content": "<p>Dart vs non-Dart:</p>\n<p>With my own dataset, running everything in Kaggle, I got the following:</p>\n<ul>\n<li>XGB GPU non-Dart (varying seed and slight tweaks to hyper-params on different runs): CV 0.7951-0.7959, one hour training (~13 min per fold)</li>\n<li>LGBM CPU Dart: CV 0.7965, 5 hours training each fold in parallel, plus about an hour test inference using the 5 model outputs.</li>\n</ul>\n<p>I'd be careful coming to any conclusions on a single test, but, if not caring about CPU vs GPU, I wouldn't be too shocked if crazy low learning rate and row and col_sample could get non-Dart CV equal to Dart in similar time of 5 hours (60 min per fold) vs 5 CPUs doing Dart at once at 5 hours per fold. Shrug.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1855640,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-07-14T18:41:43.350000",
          "content": "<p>I had a thought inspired by this sub thread:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337160\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/337160</a></p>\n<p>Basically, the issue with any two last-[thing], is that you are probably adding some highly correlated columns. This will hurt tree diversity. To deal with that, you need to drop sampling way down. But 'last' features are crucial… so let's double them(!). </p>\n<p>Oh wait, top notebook does both those things with 20% feature sampling and rounded last in addition to last. Hmm!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1858043,
          "author_name": "Dmitry Uarov",
          "author_url": "",
          "post_date": "2022-07-16T15:50:20.847000",
          "content": "<p><a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> I want to brag my new best results - XGB gbt CV(5 folds) 0.79662 (seed 42) LB 0.797. It was so difficult😅</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1858162,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-16T17:56:12.680000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1858755,
      "author_name": "Bilal Suppal",
      "author_url": "",
      "post_date": "2022-07-17T07:22:10.297000",
      "content": "<p>The information provided is so clear and easy to understand, Awesome post :)<br>\nLooking forward to see more !!!<br>\nlet's grow together.</p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 1857384,
      "author_name": "Bilal Suppal",
      "author_url": "",
      "post_date": "2022-07-16T05:13:45.507000",
      "content": "<p>Great work, <a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.</p>",
      "votes": -4,
      "replies": []
    },
    {
      "id": 1862506,
      "author_name": "Dmitry Uarov",
      "author_url": "",
      "post_date": "2022-07-19T18:52:52.693000",
      "content": "<p><a href=\"https://www.kaggle.com/mrandri19\" target=\"_blank\">@mrandri19</a> How are your research progressing? A week ago I was convinced that using DART is very early and it was true) At the moment I have 0.7973 CV(5) and 0.797 LB! And you know what? This is not the limit! I think the possible best CV result is around 0.7975 for XGB gbt.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1872318,
          "author_name": "Andrea Cognolato",
          "author_url": "",
          "post_date": "2022-07-26T20:17:51.517000",
          "content": "<p><a href=\"https://www.kaggle.com/dmitryuarov\" target=\"_blank\">@dmitryuarov</a> I am slowly progressing on feature engineering. Right now I have:</p>\n<ul>\n<li>DART-less XGB 0.7951 CV 0.796 LB</li>\n<li>DART-less LGBM 0.7953 CV 0.796 LB</li>\n<li>(Work-In-Progress) NN 0.78XX CV (haven't submitted yet to LB and did not try with all folds as I still am iterating quickly)</li>\n</ul>\n<p>The biggest boosts have come from last, count, nunique for categoricals and last - mean.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1861261,
      "author_name": "Hou zhangyi ",
      "author_url": "",
      "post_date": "2022-07-18T22:42:10.480000",
      "content": "<p>Thank you for sharing your ablations studies ! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1856280,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "2022-07-15T08:31:03.007000",
      "content": "<p>thanks for sharing. on the std of the models, i wonder what std you get from 5 folds? mine is about 0.004. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1860812,
          "author_name": "Andrea Cognolato",
          "author_url": "",
          "post_date": "2022-07-18T15:25:50.227000",
          "content": "<p>around 0.002 or 0.003</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1853608,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-07-13T02:10:13.627000",
      "content": "<p>In my experience (well, experiments), adding even more good features only really helps once you decrease colsample even lower AND lower Learning Rate. (0.01 eta and 35% col for me with ~1300 features)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1853695,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-07-13T04:41:02.183000",
          "content": "<p>Re: my experiments and re your quote: \"I am trying to achieve a single DART-less model with &gt;0.795 CV/LB score.\"</p>\n<p>I've got: 0.79699 LB and 0.79557 CV. 74min train time for 5 folds.</p>\n<p>All XGB with feature engg and a only a little hyper param tinkering. Within the LB 0.796s, I'm listed first, so my score is about 0.79699.</p>\n<p>To get that score, I used the exact features in the current version of my public feature engg notebook: <a href=\"https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a></p>\n<p>And there's some interesting features you might want to try, but, as mentioned in above comment, I wouldn't be surprised if what you really need is some hyper-parameter tweaks and those alone get you your goal of &gt; 0.795 CV/LB</p>\n<p>Hyper-params:<br>\nExcept as below, should be identical to: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>.</p>\n<ol>\n<li>After getting good results with 0.01 learning rate, I tried this and got nearly the same train time and result, but marginally better in both, iirc. 'num_parallel_tree':3, 'learning_rate':0.03,<br>\n<em>EDIT</em>: if using num_parallel_tree, oddly enough you have to modify model.best_ntree_limit by adding integer divide by num_parallel_tree. Example:<br>\nmodel.predict(dvalid, iteration_range=(0,model.best_ntree_limit//TREE_MULTIPLE))</li>\n<li>'objective':'rank:pairwise', 'eval_metric':'auc',</li>\n<li>Sampling:    'subsample':0.75,    'colsample_bytree': 0.20,<br>\nNote that I think 35% cols gave better CV results, but I didn't create submission csv on that 0.7959 CV run, and haven't had GPU time to go back and fine-tune.</li>\n</ol>\n<p>Features engg besides the stuff in other notebooks and mentioned above.<br>\n(Again, from here: <a href=\"https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a> ) </p>\n<ol>\n<li>Simplified Hull Moving Average</li>\n<li>Magnitude = max-min. (And dropped min and max)</li>\n<li>CurrentLevel = (last-min) / (max-min)</li>\n</ol>\n<p>But all my experiments were about as unscientific as yours were scientific. Constantly changing 3 variables at once :)</p>\n<p>Let me know if you try any of the above, whether params or features!</p>\n<p>P.s.: planning to share my full notebook publicly, not just the feature engg, but working on an even higher score, with one additional innovation… and having some major pathfinding issues to sort out.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1854245,
          "author_name": "Andrea Cognolato",
          "author_url": "",
          "post_date": "2022-07-13T14:35:09.800000",
          "content": "<p>Wow, this is great thanks! I will try some of your suggestions and update you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1853600,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-13T01:54:28.247000",
      "content": "<p>Thanks for sharing your ablation studies! <br>\nAre you using all of the features in the dataset? If not, what feature selection method do you use?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1854242,
          "author_name": "Andrea Cognolato",
          "author_url": "",
          "post_date": "2022-07-13T14:33:34.133000",
          "content": "<p>Yes, I am using all of the features.<br>\nI would like to try feature selection using feature importances but I did not have time yet/am not sure what to use among SHAP/Permutation Importance/Null Importance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1853486,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-07-12T23:20:23.360000",
      "content": "<p>great work, thanks for sharing your ideas!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1853294,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-07-12T19:31:55.050000",
      "content": "<p>Nice .. did lag features decrease ur CV/LB . I think one lag feature which ragner described did certainly help </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1853302,
          "author_name": "Andrea Cognolato",
          "author_url": "",
          "post_date": "2022-07-12T19:41:11.707000",
          "content": "<p>Thanks :)<br>\nDo you mean the lag features (last - first, last - avg) described by ragnar or the ones (last / first, last - first) described by thedevastator?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1853372,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-07-12T21:12:45.403000",
          "content": "<p>I mean both .. I think Ragner last-Avg made sense to me the others I havent check the feature importance yet to assess .. but Ragner ones were up there in FE. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1852203,
      "author_name": "Oleg Sidorshin",
      "author_url": "",
      "post_date": "2022-07-11T21:38:26.727000",
      "content": "<p>Always interesting to look at good feature engineering experiments, well done!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1852120,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-07-11T19:32:41.097000",
      "content": "<p>Very well curated and summarized!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1857388,
      "author_name": "Bilal Suppal",
      "author_url": "",
      "post_date": "2022-07-16T05:14:05.723000",
      "content": "<p>thanks for this help for us </p>",
      "votes": -3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1852093": "2022-07-26 update: see the code in this public notebook https://www.kaggle.com/code/mrandri19/0-796-lb-sharing-my-ablations-dart-less-xgb\n\nI am trying to achieve a single DART-less model with >0.795 CV/LB score.\nI am happy with the model and hyperparameters so I am working on feature engineering and feature selection.\n\nAll the results have been obtained via sklearn's 5-fold StratifiedKFold cross-validation with seed=42.\nKeep in mind that cdeotte says that the standard deviation of a model trained on different seeds is 0.0012.\nWith a standard deviation of this magnitude, I don't feel confident in differences smaller than 0.0024, so take my results with a pinch of salt.\n\nBelow is the description of the model, hyperparameters and the ablation studies I have ran so far:\n\n## Model\n\nXGBoost trained with early stopping using the amex metric.\nIn particular I use 9999 boosting rounds and 500 early stopping rounds.\nEach model takes 10-20 minutes to train on a Kaggle GPU instance.\n\n### Hyperparameters\n\nTaken from the excellent notebook https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb.\n\n```\n'max_depth': 7,\n'eta': 0.03,\n\n'subsample': 0.88,\n'colsample_bytree': 0.5,\n        \n'objective': 'binary:logistic',\n        \n'tree_method': 'gpu_hist',\n        \n'random_state': 42,\n        \n'gamma': 1.5,\n'min_child_weight': 8,\n'lambda': 70,\n```\n\n## Ablation studies\n\n### (a) Baseline\n\nThe features, selecting the last month of data for each customer, minus customer_ID and S_2, fed directly to the model.\n- **CV: 0.7885**\n- 9.5min training time\n- 188 features\n\n### (b) Round to 2 decimal digits\n\nAs raddar suggests, there is uniform noise injected to the data, we can ignore the noise by throwing signals smaller\nthan 0.01 by rounding to 2 decimal digits.\n- **CV: 0.7908, +0.0023** compared with (a)\n- 9.5min training time\n- 188 features\n\n### (c) Drop correlated features\n\nAfter filling the missing values with the mean, compute the pairwise pearson correlation between all 188 features.\nThen drop features that are >95% correlated with others.\n\n- **CV: 0.7905, -0.0003** compared with (b)\n- 10.5min training time\n- 178 features\n\n### (d) Add customer aggregated features\n\nSo fare we have only used the last value from each feature, let's use the first, last, mean, std, min, max as well.\n\n- **CV: 0.7949, +0.0041** compared with (b)\n- 21.5min training time\n- 1073 features\n\n### (e) Add customer aggregated features, then drop correlated features\n\nSince we have a lot more features, we can again try to drop correlated ones.\n\n- **CV: 0.7950, +0.0001** compared with (d)\n- 22.5min training time\n- 875 features\n\n### (f) Add missing count features\n\nWe can count the number of missing values in each column, for every customer.\nFor every column, this is an integer in [0, 13].\n\n- **CV: 0.7946, -0.0003** compared with (d)\n- 24.5min training time\n- 1261 features\n\n### (g) Add difference features\n\nReplicate the feature enginering performed by Jiwei Liu in this notebook: https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb.\nIn practice, compute the subtraction of a bunch of B, D, S features minus two P (Payment) features: P_2, P_3\n\n- **CV: 0.7953, +0.0004** compared with (d)\n- 20.9min training time\n- 1157 features\n\n### (h) Add lag features\n\nReplicate the feature engineering by thedevastator described in described in\nhttps://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need\n\n- **CV: 0.7937, -0.0016** compared with (g)\n- 24.2min training time\n- 1511 features\n\nThis is very weird, as these features come from one of the best publicly available notebooks.\nMy hypothesis is that too many (>1500) features combined with an aggressive 0.5 colsample_bytree leads to\nsome the model missing some interactions, but I am curious of the reader's opinion.\n\n### (i) Illidan7's Feature Set\n\nhttps://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\n\n- **CV: 0.7950**\n- 41.2min training time\n- 1486 features\n\n### (j) Add ragnar123' first - mean lag feature\n\n- **CV: 0.7956**\n- 27.9min training time\n- 974 features\n\n\n## Conclusions and next steps\n\n- Rounding seems to be beneficial, should we round the values before performing the aggreagations, then?\n- Adding first, mean, std, min, max aggregations has the biggest impact on performance, but they are almost 900 additional features. How to identify the ones that actually bring benefits?\n- It is not clear if dropping correlated features has an effect. Why doesn't it reduce the training time? Is it because we need more trees?\n- The difference features give a small (4 basis points) boost. I believe they are worth using as they are only 40.\n- Next, I want to check 's https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977",
    "1864319": "Hi Andrea, update on my end: I went live with my XGB notebook today: https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions \n\nCV 0.7968 (~0.7965-7970); LB 0.797 and edged out other 0.797 public notebooks.\n\nThe hyper-parameters aren't all that optimized, the feature engineering is unoptimized and pretty untested. But I'm *very* happy with the results of my \"Pyramid\" experiment.\n\nEssentially, I tried using other methods besides DART that I suspected might reduce \"over-specialization\", without the performance penalty that DART + XGB + GPU has. It was very successful, using two or three tricks, depending how you count it:\n* Boosted forest early on, very directly reduces impact of initial tree, and a little more evenly spreads even at later points (with forest size of 100, then tree 101 has same impact as tree 200).\n* Rescale the prediction results retroactively to give more 'room' for impact from later trees.\n* These two innovations supported by the basic method of stopping a model early, and using it as the basis of the next model. There might be a more elegant way to do that, but the one I found that definitely works is to input the predictions from model A before starting training model B using \"set_base_margin\".",
    "1853206": "Great work, @mrandri19 , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.",
    "1858755": "The information provided is so clear and easy to understand, Awesome post :)\nLooking forward to see more !!!\nlet's grow together.",
    "1857384": "Great work, @mrandri19 , my best XGB gbt has CV 0.7958 and LB 0.796 (seed 42). Point (h) caused me the same question - I don't know why, but lags \"last-first\" and \"last-avg\" work worse together than if you leave one of them. Perhaps dart is more loyal to the number of features than gbt, but it's still strange. I also tried adding kurtosis and skew of most important features, but it didn't improve the result.",
    "1862506": "@mrandri19 How are your research progressing? A week ago I was convinced that using DART is very early and it was true) At the moment I have 0.7973 CV(5) and 0.797 LB! And you know what? This is not the limit! I think the possible best CV result is around 0.7975 for XGB gbt.",
    "1861261": "Thank you for sharing your ablations studies ! \n",
    "1856280": "thanks for sharing. on the std of the models, i wonder what std you get from 5 folds? mine is about 0.004. ",
    "1853608": "In my experience (well, experiments), adding even more good features only really helps once you decrease colsample even lower AND lower Learning Rate. (0.01 eta and 35% col for me with ~1300 features)",
    "1853600": "Thanks for sharing your ablation studies! \nAre you using all of the features in the dataset? If not, what feature selection method do you use?\n\n",
    "1853486": "great work, thanks for sharing your ideas!",
    "1853294": "Nice .. did lag features decrease ur CV/LB . I think one lag feature which ragner described did certainly help ",
    "1852203": "Always interesting to look at good feature engineering experiments, well done!",
    "1852120": "Very well curated and summarized!",
    "1857388": "thanks for this help for us "
  }
}