{
  "id": 338752,
  "title": "Boosting XGBoost model score - without slowness of DART",
  "url": "/competitions/amex-default-prediction/discussion/338752",
  "author_name": "",
  "post_date": "2022-07-21T22:03:55.687434900Z",
  "votes": 30,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I did pathfinding on an alternative to using DART to, ideally, see major score improvements through reducing the 'over-specialization' issue. For now, calling it the Pyramid.</p>\n<p>The Pyramid method was very successful for me, roughly ~0.001 score improvement with the <em>same</em> training time. If using the same base learning rate, it might see more like ~0.0010-0.0015 improvement with 25% longer training time.</p>\n<p>My initial example is here: <a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions</a> (links to the other notebooks used, it can all run on Kaggle as is :) )<br>\nCV 0.7968, LB 0.797.</p>\n<p>As a second proof of concept, I took an existing CV 0.7961; LB 0.796 notebook, and improved it with nothing more than the Pyramid method to CV 0.7968, LB 0.797.<br>\nOriginal: <a href=\"https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use\" target=\"_blank\">https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use</a> <br>\nMy straightforward update: <a href=\"https://www.kaggle.com/code/roberthatch/pyramid-on-statement-dates-notebook\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/pyramid-on-statement-dates-notebook</a> </p>\n<p>So what did I do?</p>\n<p>See the notebooks for full implementation details, but basically I separate a single model training into layers. Instead of a single 8000 tree model, add some 1000 tree layers, for example. If keeping all parameters the same, this would have no effect.</p>\n<p>The two functional changes built on to that:</p>\n<ul>\n<li>First (and second) layer trains a boosted forest instead of a normal booster, to reduce overspecialization by directly building batches of trees at once.</li>\n<li>Between layers, reduce the scale of the output predictions retroactively. This has a different impact than reducing the learning rate in advance. I can't be sure exactly the, ah, true underlying differences, but imagine two similar but different scenarios:<br>\nA: Train at 0.3 eta for 5 rounds. Logitraw predictions might be 30% to target value, then 50%, then 65%, 75%, end at 83%. Assuming the 5 trees happened to move it pretty smoothly towards the target value.<br>\nB: Train at 0.5 eta, then rescale retroactively to 60% of result: 50%, then 75%, 87%, 93%, 96%. Rescaled to 58%.</li>\n</ul>\n<p>So instead of starting with more impactful trees, then gradually adding less impactful trees to tweak the result at the end, you slightly intermix impactful trees, some less impactful tweaking, then jump back to more impactful trees, then again gradually less impactful trees doing more residual tweaking.</p>\n<p>This is all pathfinding and guesswork, I'm not a machine learning expert. Just on here to test fun hypotheses (like this one :) )</p>",
  "messages": [
    {
      "id": "1865506",
      "postDate": "07/21/2022 22:03:55",
      "content": "<p>I did pathfinding on an alternative to using DART to, ideally, see major score improvements through reducing the 'over-specialization' issue. For now, calling it the Pyramid.</p>\n<p>The Pyramid method was very successful for me, roughly ~0.001 score improvement with the <em>same</em> training time. If using the same base learning rate, it might see more like ~0.0010-0.0015 improvement with 25% longer training time.</p>\n<p>My initial example is here: <a href=\"https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions</a> (links to the other notebooks used, it can all run on Kaggle as is :) )<br>\nCV 0.7968, LB 0.797.</p>\n<p>As a second proof of concept, I took an existing CV 0.7961; LB 0.796 notebook, and improved it with nothing more than the Pyramid method to CV 0.7968, LB 0.797.<br>\nOriginal: <a href=\"https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use\" target=\"_blank\">https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use</a> <br>\nMy straightforward update: <a href=\"https://www.kaggle.com/code/roberthatch/pyramid-on-statement-dates-notebook\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/pyramid-on-statement-dates-notebook</a> </p>\n<p>So what did I do?</p>\n<p>See the notebooks for full implementation details, but basically I separate a single model training into layers. Instead of a single 8000 tree model, add some 1000 tree layers, for example. If keeping all parameters the same, this would have no effect.</p>\n<p>The two functional changes built on to that:</p>\n<ul>\n<li>First (and second) layer trains a boosted forest instead of a normal booster, to reduce overspecialization by directly building batches of trees at once.</li>\n<li>Between layers, reduce the scale of the output predictions retroactively. This has a different impact than reducing the learning rate in advance. I can't be sure exactly the, ah, true underlying differences, but imagine two similar but different scenarios:<br>\nA: Train at 0.3 eta for 5 rounds. Logitraw predictions might be 30% to target value, then 50%, then 65%, 75%, end at 83%. Assuming the 5 trees happened to move it pretty smoothly towards the target value.<br>\nB: Train at 0.5 eta, then rescale retroactively to 60% of result: 50%, then 75%, 87%, 93%, 96%. Rescaled to 58%.</li>\n</ul>\n<p>So instead of starting with more impactful trees, then gradually adding less impactful trees to tweak the result at the end, you slightly intermix impactful trees, some less impactful tweaking, then jump back to more impactful trees, then again gradually less impactful trees doing more residual tweaking.</p>\n<p>This is all pathfinding and guesswork, I'm not a machine learning expert. Just on here to test fun hypotheses (like this one :) )</p>",
      "rawMarkdown": "I did pathfinding on an alternative to using DART to, ideally, see major score improvements through reducing the 'over-specialization' issue. For now, calling it the Pyramid.\n\nThe Pyramid method was very successful for me, roughly ~0.001 score improvement with the *same* training time. If using the same base learning rate, it might see more like ~0.0010-0.0015 improvement with 25% longer training time.\n\nMy initial example is here: https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions (links to the other notebooks used, it can all run on Kaggle as is :) )\nCV 0.7968, LB 0.797.\n\nAs a second proof of concept, I took an existing CV 0.7961; LB 0.796 notebook, and improved it with nothing more than the Pyramid method to CV 0.7968, LB 0.797.\nOriginal: https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use \nMy straightforward update: https://www.kaggle.com/code/roberthatch/pyramid-on-statement-dates-notebook \n\nSo what did I do?\n\nSee the notebooks for full implementation details, but basically I separate a single model training into layers. Instead of a single 8000 tree model, add some 1000 tree layers, for example. If keeping all parameters the same, this would have no effect.\n\nThe two functional changes built on to that:\n* First (and second) layer trains a boosted forest instead of a normal booster, to reduce overspecialization by directly building batches of trees at once.\n* Between layers, reduce the scale of the output predictions retroactively. This has a different impact than reducing the learning rate in advance. I can't be sure exactly the, ah, true underlying differences, but imagine two similar but different scenarios:\nA: Train at 0.3 eta for 5 rounds. Logitraw predictions might be 30% to target value, then 50%, then 65%, 75%, end at 83%. Assuming the 5 trees happened to move it pretty smoothly towards the target value.\nB: Train at 0.5 eta, then rescale retroactively to 60% of result: 50%, then 75%, 87%, 93%, 96%. Rescaled to 58%.\n\nSo instead of starting with more impactful trees, then gradually adding less impactful trees to tweak the result at the end, you slightly intermix impactful trees, some less impactful tweaking, then jump back to more impactful trees, then again gradually less impactful trees doing more residual tweaking.\n\nThis is all pathfinding and guesswork, I'm not a machine learning expert. Just on here to test fun hypotheses (like this one :) )",
      "votes": null
    },
    {
      "id": "1866047",
      "postDate": "07/22/2022 08:35:33",
      "content": "<p>Good job <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>",
      "rawMarkdown": "Good job @roberthatch",
      "votes": null
    },
    {
      "id": "1866288",
      "postDate": "07/22/2022 11:20:41",
      "content": "<p>I didn't go through the code itself yet but the idea is impressive on its own.. </p>",
      "rawMarkdown": "I didn't go through the code itself yet but the idea is impressive on its own..",
      "votes": null
    },
    {
      "id": "1867211",
      "postDate": "07/23/2022 04:44:00",
      "content": "<p>I developed this on XGB, but it would be interesting to do it on LightGBM as well, so:</p>\n<p>LGBM vs XGB:</p>\n<ul>\n<li>LGBM's Dataset class has 'init_score' function which should be equivalent to 'set_base_margin' in XGB's DMatrix class.</li>\n<li>It appears that LGBM doesn't have any direct way to do a boosted forest like XGB does. You could simply do 100 to 1000 tree random forest as layer 1, reweight it as desired, and save the 'init_score' for the next layer to star the actual boosting. It might end up having about the same effect?</li>\n</ul>\n<p>Besides that, with either one, but maybe easier on LGBM, you can run Dart <em>and</em> try adding the pyramid and see if they are complementary? </p>",
      "rawMarkdown": "I developed this on XGB, but it would be interesting to do it on LightGBM as well, so:\n\nLGBM vs XGB:\n* LGBM's Dataset class has 'init_score' function which should be equivalent to 'set_base_margin' in XGB's DMatrix class.\n* It appears that LGBM doesn't have any direct way to do a boosted forest like XGB does. You could simply do 100 to 1000 tree random forest as layer 1, reweight it as desired, and save the 'init_score' for the next layer to star the actual boosting. It might end up having about the same effect?\n\nBesides that, with either one, but maybe easier on LGBM, you can run Dart *and* try adding the pyramid and see if they are complementary?",
      "votes": null
    },
    {
      "id": "1868593",
      "postDate": "07/24/2022 04:49:42",
      "content": "<p>Brilliant!</p>",
      "rawMarkdown": "Brilliant!",
      "votes": null
    },
    {
      "id": "1869372",
      "postDate": "07/24/2022 17:26:10",
      "content": "<p>Hi, I tried using the pyramid method on CPU-only machine and it raised an exception (that we cannot use np.array as an output type of the iterative loader), Do you know how to fix it? </p>",
      "rawMarkdown": "Hi, I tried using the pyramid method on CPU-only machine and it raised an exception (that we cannot use np.array as an output type of the iterative loader), Do you know how to fix it?",
      "votes": null
    },
    {
      "id": "1869426",
      "postDate": "07/24/2022 18:34:19",
      "content": "<p>Published an updated notebook that hopefully works well for fast drop-in use!</p>\n<p><a href=\"https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment\" target=\"_blank\">https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment</a></p>",
      "rawMarkdown": "Published an updated notebook that hopefully works well for fast drop-in use!\n\nhttps://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment",
      "votes": null
    },
    {
      "id": "1869431",
      "postDate": "07/24/2022 18:39:05",
      "content": "<p>First maybe try using the functions from this notebook and see if it works as is? (I just published it)</p>\n<p><a href=\"https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment\" target=\"_blank\">https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment</a></p>\n<p>If it persists, can you give a bit more detailed error information, pointing to which line, etc? I didn't try with CPU-only.</p>\n<p>I'm guessing failing line is one of these three?<br>\n// first time through the loop, fails at one of these?<br>\n            ptrain = ptrain * w<br>\n            dtrain.set_base_margin(ptrain)<br>\n// second time through the loop, fails at xgb.train?<br>\n        model = xgb.train(params, […])</p>",
      "rawMarkdown": "First maybe try using the functions from this notebook and see if it works as is? (I just published it)\n\nhttps://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment\n\nIf it persists, can you give a bit more detailed error information, pointing to which line, etc? I didn't try with CPU-only.\n\nI'm guessing failing line is one of these three?\n// first time through the loop, fails at one of these?\n            ptrain = ptrain * w\n            dtrain.set_base_margin(ptrain)\n// second time through the loop, fails at xgb.train?\n        model = xgb.train(params, [...])",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1866047,
      "author_name": "saberghaderi",
      "author_url": "",
      "post_date": "07/22/2022 08:35:33",
      "content": "<p>Good job <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1866288,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/22/2022 11:20:41",
      "content": "<p>I didn't go through the code itself yet but the idea is impressive on its own.. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1867211,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/23/2022 04:44:00",
      "content": "<p>I developed this on XGB, but it would be interesting to do it on LightGBM as well, so:</p>\n<p>LGBM vs XGB:</p>\n<ul>\n<li>LGBM's Dataset class has 'init_score' function which should be equivalent to 'set_base_margin' in XGB's DMatrix class.</li>\n<li>It appears that LGBM doesn't have any direct way to do a boosted forest like XGB does. You could simply do 100 to 1000 tree random forest as layer 1, reweight it as desired, and save the 'init_score' for the next layer to star the actual boosting. It might end up having about the same effect?</li>\n</ul>\n<p>Besides that, with either one, but maybe easier on LGBM, you can run Dart <em>and</em> try adding the pyramid and see if they are complementary? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1868593,
      "author_name": "lilgaussy",
      "author_url": "",
      "post_date": "07/24/2022 04:49:42",
      "content": "<p>Brilliant!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1869372,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/24/2022 17:26:10",
      "content": "<p>Hi, I tried using the pyramid method on CPU-only machine and it raised an exception (that we cannot use np.array as an output type of the iterative loader), Do you know how to fix it? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1869431,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/24/2022 18:39:05",
          "content": "<p>First maybe try using the functions from this notebook and see if it works as is? (I just published it)</p>\n<p><a href=\"https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment\" target=\"_blank\">https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment</a></p>\n<p>If it persists, can you give a bit more detailed error information, pointing to which line, etc? I didn't try with CPU-only.</p>\n<p>I'm guessing failing line is one of these three?<br>\n// first time through the loop, fails at one of these?<br>\n            ptrain = ptrain * w<br>\n            dtrain.set_base_margin(ptrain)<br>\n// second time through the loop, fails at xgb.train?<br>\n        model = xgb.train(params, […])</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1869426,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/24/2022 18:34:19",
      "content": "<p>Published an updated notebook that hopefully works well for fast drop-in use!</p>\n<p><a href=\"https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment\" target=\"_blank\">https://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1865506": "I did pathfinding on an alternative to using DART to, ideally, see major score improvements through reducing the 'over-specialization' issue. For now, calling it the Pyramid.\n\nThe Pyramid method was very successful for me, roughly ~0.001 score improvement with the *same* training time. If using the same base learning rate, it might see more like ~0.0010-0.0015 improvement with 25% longer training time.\n\nMy initial example is here: https://www.kaggle.com/code/roberthatch/xgboost-pyramid-test-predictions (links to the other notebooks used, it can all run on Kaggle as is :) )\nCV 0.7968, LB 0.797.\n\nAs a second proof of concept, I took an existing CV 0.7961; LB 0.796 notebook, and improved it with nothing more than the Pyramid method to CV 0.7968, LB 0.797.\nOriginal: https://www.kaggle.com/code/romaupgini/statement-dates-to-use-or-not-to-use \nMy straightforward update: https://www.kaggle.com/code/roberthatch/pyramid-on-statement-dates-notebook \n\nSo what did I do?\n\nSee the notebooks for full implementation details, but basically I separate a single model training into layers. Instead of a single 8000 tree model, add some 1000 tree layers, for example. If keeping all parameters the same, this would have no effect.\n\nThe two functional changes built on to that:\n* First (and second) layer trains a boosted forest instead of a normal booster, to reduce overspecialization by directly building batches of trees at once.\n* Between layers, reduce the scale of the output predictions retroactively. This has a different impact than reducing the learning rate in advance. I can't be sure exactly the, ah, true underlying differences, but imagine two similar but different scenarios:\nA: Train at 0.3 eta for 5 rounds. Logitraw predictions might be 30% to target value, then 50%, then 65%, 75%, end at 83%. Assuming the 5 trees happened to move it pretty smoothly towards the target value.\nB: Train at 0.5 eta, then rescale retroactively to 60% of result: 50%, then 75%, 87%, 93%, 96%. Rescaled to 58%.\n\nSo instead of starting with more impactful trees, then gradually adding less impactful trees to tweak the result at the end, you slightly intermix impactful trees, some less impactful tweaking, then jump back to more impactful trees, then again gradually less impactful trees doing more residual tweaking.\n\nThis is all pathfinding and guesswork, I'm not a machine learning expert. Just on here to test fun hypotheses (like this one :) )",
    "1866047": "Good job @roberthatch",
    "1866288": "I didn't go through the code itself yet but the idea is impressive on its own..",
    "1867211": "I developed this on XGB, but it would be interesting to do it on LightGBM as well, so:\n\nLGBM vs XGB:\n* LGBM's Dataset class has 'init_score' function which should be equivalent to 'set_base_margin' in XGB's DMatrix class.\n* It appears that LGBM doesn't have any direct way to do a boosted forest like XGB does. You could simply do 100 to 1000 tree random forest as layer 1, reweight it as desired, and save the 'init_score' for the next layer to star the actual boosting. It might end up having about the same effect?\n\nBesides that, with either one, but maybe easier on LGBM, you can run Dart *and* try adding the pyramid and see if they are complementary?",
    "1868593": "Brilliant!",
    "1869372": "Hi, I tried using the pyramid method on CPU-only machine and it raised an exception (that we cannot use np.array as an output type of the iterative loader), Do you know how to fix it?",
    "1869426": "Published an updated notebook that hopefully works well for fast drop-in use!\n\nhttps://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment",
    "1869431": "First maybe try using the functions from this notebook and see if it works as is? (I just published it)\n\nhttps://www.kaggle.com/roberthatch/pyramid-api-for-easy-deployment\n\nIf it persists, can you give a bit more detailed error information, pointing to which line, etc? I didn't try with CPU-only.\n\nI'm guessing failing line is one of these three?\n// first time through the loop, fails at one of these?\n            ptrain = ptrain * w\n            dtrain.set_base_margin(ptrain)\n// second time through the loop, fails at xgb.train?\n        model = xgb.train(params, [...])"
  },
  "source": "meta"
}