{
  "id": 338076,
  "title": "Does incremental training of LGBM negatively affect performance?",
  "url": "/competitions/amex-default-prediction/discussion/338076",
  "author_name": "",
  "post_date": "2022-07-19T04:10:13.748987200Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have found it quite challenging to run LGBM successfully on a free Colab, Paperspace or Kaggle instance without running out of memory</p>\n<p>I modified <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>'s notebook to run successfully within a Kaggle notebook, however in the process I sacrificed CV (just trusting their hyperparameters for now) and also CV ensembling as a result. This was done using incremental training to load the train data in 5 parts, and continuously train the model while dumping previous data through the use of parameters <code>keep_training_booster=True</code> and <code>init_model = model</code>.</p>\n<p>However, I noticed that my model performs significantly worse. at 0.785 with a single split and no lag features, and 0.793 by ensembling multiple seeds (for the incremental training splits) and adding lag features.</p>\n<p>I don't fully understand why, but I could guess that there would be a difference between training on the full train set for n rounds, versus training on one fifth of the train set each time, for n rounds, totaling 5n rounds.</p>\n<p>I am aware that there is a method to iteratively pass data to XGBoost using <code>DeviceQuantileDMatrix</code>, which seems to function identically to passing the full dataset initially.</p>\n<p>Would the corresponding method with LGBM be using the <code>Sequence</code> interface in LGBM instead of using <code>init_model</code> and training multiple times?</p>",
  "messages": [
    {
      "id": "1861481",
      "postDate": "07/19/2022 04:10:13",
      "content": "<p>I have found it quite challenging to run LGBM successfully on a free Colab, Paperspace or Kaggle instance without running out of memory</p>\n<p>I modified <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>'s notebook to run successfully within a Kaggle notebook, however in the process I sacrificed CV (just trusting their hyperparameters for now) and also CV ensembling as a result. This was done using incremental training to load the train data in 5 parts, and continuously train the model while dumping previous data through the use of parameters <code>keep_training_booster=True</code> and <code>init_model = model</code>.</p>\n<p>However, I noticed that my model performs significantly worse. at 0.785 with a single split and no lag features, and 0.793 by ensembling multiple seeds (for the incremental training splits) and adding lag features.</p>\n<p>I don't fully understand why, but I could guess that there would be a difference between training on the full train set for n rounds, versus training on one fifth of the train set each time, for n rounds, totaling 5n rounds.</p>\n<p>I am aware that there is a method to iteratively pass data to XGBoost using <code>DeviceQuantileDMatrix</code>, which seems to function identically to passing the full dataset initially.</p>\n<p>Would the corresponding method with LGBM be using the <code>Sequence</code> interface in LGBM instead of using <code>init_model</code> and training multiple times?</p>",
      "rawMarkdown": "I have found it quite challenging to run LGBM successfully on a free Colab, Paperspace or Kaggle instance without running out of memory\n\nI modified @ragnar123's notebook to run successfully within a Kaggle notebook, however in the process I sacrificed CV (just trusting their hyperparameters for now) and also CV ensembling as a result. This was done using incremental training to load the train data in 5 parts, and continuously train the model while dumping previous data through the use of parameters `keep_training_booster=True` and `init_model = model`.\n\nHowever, I noticed that my model performs significantly worse. at 0.785 with a single split and no lag features, and 0.793 by ensembling multiple seeds (for the incremental training splits) and adding lag features.\n\nI don't fully understand why, but I could guess that there would be a difference between training on the full train set for n rounds, versus training on one fifth of the train set each time, for n rounds, totaling 5n rounds.\n\nI am aware that there is a method to iteratively pass data to XGBoost using `DeviceQuantileDMatrix`, which seems to function identically to passing the full dataset initially.\n\nWould the corresponding method with LGBM be using the `Sequence` interface in LGBM instead of using `init_model` and training multiple times?",
      "votes": null
    },
    {
      "id": "1861745",
      "postDate": "07/19/2022 08:05:25",
      "content": "<p>To my understanding neither XGBoost nor LGBM support incremental training in the sense that you could train the model on data and than continue training on never seen data. The <em>continue training</em> feature is thought to continue training on the same data (more epochs).  For clarification: The same data means data that the model has already seen, not data from the same Kaggle competition. :-)</p>\n<p>The main reasons seem to be:</p>\n<ul>\n<li>tree construction algorithms currently depend on the availability of the whole data to choose optimal splits.</li>\n<li>a new tree should each time reduce the training loss over the <strong>whole</strong> training data.</li>\n</ul>\n<p>The above points are from a <a href=\"https://github.com/dmlc/xgboost/issues/3055#issuecomment-359505122\" target=\"_blank\">discussion on the XGBoost github about continued learing</a>.</p>\n<p>One of the participants adds:<br>\nHowever, there are applications when training continuation in new data makes good practical sense. E.g., in a situation when you get some new data that is related but has some sort of \"concept drift\", there are sometimes good chances that by taking an old model learned in old data as \"prior knowledge\", and adapting it to the new data by training continuation in that new data, you would get a better performing model for future use in data that would be like the new data than when training from scratch either with only this new data or with a combined sample of old + new data. </p>\n<p>So if you like you could try to split the data per month, train and predict for that month and then continue training the model on the next month and use the final model for inference. But I'm not sure if that helps with the current data.</p>\n<p>As you already mentioned you could pass the whole data incrementally to the model with <code>DeviceQuantileDMatrix</code> or a similar dataloader.<br>\nThe thing is that if you load the data from files, then random access to that data is hard to implement with csv or parquet, so <br>\n you probably don't want to use shuffling or splits over the whole dataset.<br>\nOne thing you could do is to split and shuffle the data into files beforehand and then yield them from the dataloader, meaning the index going into the dataloader to get the data is the number of the file and not the row of the dataset, thus the length returned by the dataloader is the number of files.</p>\n<p><a href=\"https://github.com/microsoft/LightGBM/issues/4672#issuecomment-941057024\" target=\"_blank\">Here are some suggestions</a> for \"train a model on a larger-than-memory dataset on a single node\" by an LGBM Maintainer:</p>\n<ul>\n<li>the Sequence Api you mentioned</li>\n<li>using lightgbm.dask</li>\n<li>Train a smaller model</li>\n<li>Use Training Continuation</li>\n</ul>\n<p>What? Use Training Continuation? How?<br>\nUnder the above link you'll find code.<br>\nThe trick seems to be:  \"reference=first_dataset\" in the loop over files.</p>\n<pre><code>dataset = lgb.Dataset(\n        next_file,\n        params={\"label_column\": \"y\"},\n        free_raw_data=True,\n        reference=first_dataset\n    )\n</code></pre>\n<p>This seems to <a href=\"https://github.com/microsoft/LightGBM/issues/4555#issuecomment-924533600\" target=\"_blank\">align the bin mapper for the dataset</a>.<br>\nI'm not sure about how this influences the training continuation:</p>\n<blockquote>\n  <p>If the distribution of  data is very different, the bin mapper could be very different.</p>\n</blockquote>\n<p>I don't know it they handle that completely in their dataloader or if one has to check that the distribution of values in the samples is similar beforehand.<br>\nIf you try this approch a followup would be nice. I'd like to know if that works for our data.</p>\n<p>Another, maybe easier idea is to take an ensemble approach.<br>\nSplit the data into multiple batches and train a model on each. <br>\nThen at inference time simply make a prediction with each model using all the new data and take the average to get the final prediction.</p>",
      "rawMarkdown": "To my understanding neither XGBoost nor LGBM support incremental training in the sense that you could train the model on data and than continue training on never seen data. The *continue training* feature is thought to continue training on the same data (more epochs).  For clarification: The same data means data that the model has already seen, not data from the same Kaggle competition. :-)\n\nThe main reasons seem to be:\n-  tree construction algorithms currently depend on the availability of the whole data to choose optimal splits.\n- a new tree should each time reduce the training loss over the **whole** training data.\n\nThe above points are from a [discussion on the XGBoost github about continued learing](https://github.com/dmlc/xgboost/issues/3055#issuecomment-359505122).\n\nOne of the participants adds:\nHowever, there are applications when training continuation in new data makes good practical sense. E.g., in a situation when you get some new data that is related but has some sort of \"concept drift\", there are sometimes good chances that by taking an old model learned in old data as \"prior knowledge\", and adapting it to the new data by training continuation in that new data, you would get a better performing model for future use in data that would be like the new data than when training from scratch either with only this new data or with a combined sample of old + new data. \n\nSo if you like you could try to split the data per month, train and predict for that month and then continue training the model on the next month and use the final model for inference. But I'm not sure if that helps with the current data.\n\nAs you already mentioned you could pass the whole data incrementally to the model with `DeviceQuantileDMatrix` or a similar dataloader.\nThe thing is that if you load the data from files, then random access to that data is hard to implement with csv or parquet, so \n you probably don't want to use shuffling or splits over the whole dataset.\nOne thing you could do is to split and shuffle the data into files beforehand and then yield them from the dataloader, meaning the index going into the dataloader to get the data is the number of the file and not the row of the dataset, thus the length returned by the dataloader is the number of files.\n\n[Here are some suggestions](https://github.com/microsoft/LightGBM/issues/4672#issuecomment-941057024) for \"train a model on a larger-than-memory dataset on a single node\" by an LGBM Maintainer:\n- the Sequence Api you mentioned\n- using lightgbm.dask\n- Train a smaller model\n- Use Training Continuation\n\nWhat? Use Training Continuation? How?\nUnder the above link you'll find code.\nThe trick seems to be:  \"reference=first_dataset\" in the loop over files.\n```\ndataset = lgb.Dataset(\n        next_file,\n        params={\"label_column\": \"y\"},\n        free_raw_data=True,\n        reference=first_dataset\n    )\n```\nThis seems to [align the bin mapper for the dataset](https://github.com/microsoft/LightGBM/issues/4555#issuecomment-924533600).\nI'm not sure about how this influences the training continuation:\n> If the distribution of ~~val and train~~ data is very different, the bin mapper could be very different.\n\nI don't know it they handle that completely in their dataloader or if one has to check that the distribution of values in the samples is similar beforehand.\nIf you try this approch a followup would be nice. I'd like to know if that works for our data.\n\nAnother, maybe easier idea is to take an ensemble approach.\nSplit the data into multiple batches and train a model on each. \nThen at inference time simply make a prediction with each model using all the new data and take the average to get the final prediction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1861745,
      "author_name": "arigion",
      "author_url": "",
      "post_date": "07/19/2022 08:05:25",
      "content": "<p>To my understanding neither XGBoost nor LGBM support incremental training in the sense that you could train the model on data and than continue training on never seen data. The <em>continue training</em> feature is thought to continue training on the same data (more epochs).  For clarification: The same data means data that the model has already seen, not data from the same Kaggle competition. :-)</p>\n<p>The main reasons seem to be:</p>\n<ul>\n<li>tree construction algorithms currently depend on the availability of the whole data to choose optimal splits.</li>\n<li>a new tree should each time reduce the training loss over the <strong>whole</strong> training data.</li>\n</ul>\n<p>The above points are from a <a href=\"https://github.com/dmlc/xgboost/issues/3055#issuecomment-359505122\" target=\"_blank\">discussion on the XGBoost github about continued learing</a>.</p>\n<p>One of the participants adds:<br>\nHowever, there are applications when training continuation in new data makes good practical sense. E.g., in a situation when you get some new data that is related but has some sort of \"concept drift\", there are sometimes good chances that by taking an old model learned in old data as \"prior knowledge\", and adapting it to the new data by training continuation in that new data, you would get a better performing model for future use in data that would be like the new data than when training from scratch either with only this new data or with a combined sample of old + new data. </p>\n<p>So if you like you could try to split the data per month, train and predict for that month and then continue training the model on the next month and use the final model for inference. But I'm not sure if that helps with the current data.</p>\n<p>As you already mentioned you could pass the whole data incrementally to the model with <code>DeviceQuantileDMatrix</code> or a similar dataloader.<br>\nThe thing is that if you load the data from files, then random access to that data is hard to implement with csv or parquet, so <br>\n you probably don't want to use shuffling or splits over the whole dataset.<br>\nOne thing you could do is to split and shuffle the data into files beforehand and then yield them from the dataloader, meaning the index going into the dataloader to get the data is the number of the file and not the row of the dataset, thus the length returned by the dataloader is the number of files.</p>\n<p><a href=\"https://github.com/microsoft/LightGBM/issues/4672#issuecomment-941057024\" target=\"_blank\">Here are some suggestions</a> for \"train a model on a larger-than-memory dataset on a single node\" by an LGBM Maintainer:</p>\n<ul>\n<li>the Sequence Api you mentioned</li>\n<li>using lightgbm.dask</li>\n<li>Train a smaller model</li>\n<li>Use Training Continuation</li>\n</ul>\n<p>What? Use Training Continuation? How?<br>\nUnder the above link you'll find code.<br>\nThe trick seems to be:  \"reference=first_dataset\" in the loop over files.</p>\n<pre><code>dataset = lgb.Dataset(\n        next_file,\n        params={\"label_column\": \"y\"},\n        free_raw_data=True,\n        reference=first_dataset\n    )\n</code></pre>\n<p>This seems to <a href=\"https://github.com/microsoft/LightGBM/issues/4555#issuecomment-924533600\" target=\"_blank\">align the bin mapper for the dataset</a>.<br>\nI'm not sure about how this influences the training continuation:</p>\n<blockquote>\n  <p>If the distribution of  data is very different, the bin mapper could be very different.</p>\n</blockquote>\n<p>I don't know it they handle that completely in their dataloader or if one has to check that the distribution of values in the samples is similar beforehand.<br>\nIf you try this approch a followup would be nice. I'd like to know if that works for our data.</p>\n<p>Another, maybe easier idea is to take an ensemble approach.<br>\nSplit the data into multiple batches and train a model on each. <br>\nThen at inference time simply make a prediction with each model using all the new data and take the average to get the final prediction.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1861481": "I have found it quite challenging to run LGBM successfully on a free Colab, Paperspace or Kaggle instance without running out of memory\n\nI modified @ragnar123's notebook to run successfully within a Kaggle notebook, however in the process I sacrificed CV (just trusting their hyperparameters for now) and also CV ensembling as a result. This was done using incremental training to load the train data in 5 parts, and continuously train the model while dumping previous data through the use of parameters `keep_training_booster=True` and `init_model = model`.\n\nHowever, I noticed that my model performs significantly worse. at 0.785 with a single split and no lag features, and 0.793 by ensembling multiple seeds (for the incremental training splits) and adding lag features.\n\nI don't fully understand why, but I could guess that there would be a difference between training on the full train set for n rounds, versus training on one fifth of the train set each time, for n rounds, totaling 5n rounds.\n\nI am aware that there is a method to iteratively pass data to XGBoost using `DeviceQuantileDMatrix`, which seems to function identically to passing the full dataset initially.\n\nWould the corresponding method with LGBM be using the `Sequence` interface in LGBM instead of using `init_model` and training multiple times?",
    "1861745": "To my understanding neither XGBoost nor LGBM support incremental training in the sense that you could train the model on data and than continue training on never seen data. The *continue training* feature is thought to continue training on the same data (more epochs).  For clarification: The same data means data that the model has already seen, not data from the same Kaggle competition. :-)\n\nThe main reasons seem to be:\n-  tree construction algorithms currently depend on the availability of the whole data to choose optimal splits.\n- a new tree should each time reduce the training loss over the **whole** training data.\n\nThe above points are from a [discussion on the XGBoost github about continued learing](https://github.com/dmlc/xgboost/issues/3055#issuecomment-359505122).\n\nOne of the participants adds:\nHowever, there are applications when training continuation in new data makes good practical sense. E.g., in a situation when you get some new data that is related but has some sort of \"concept drift\", there are sometimes good chances that by taking an old model learned in old data as \"prior knowledge\", and adapting it to the new data by training continuation in that new data, you would get a better performing model for future use in data that would be like the new data than when training from scratch either with only this new data or with a combined sample of old + new data. \n\nSo if you like you could try to split the data per month, train and predict for that month and then continue training the model on the next month and use the final model for inference. But I'm not sure if that helps with the current data.\n\nAs you already mentioned you could pass the whole data incrementally to the model with `DeviceQuantileDMatrix` or a similar dataloader.\nThe thing is that if you load the data from files, then random access to that data is hard to implement with csv or parquet, so \n you probably don't want to use shuffling or splits over the whole dataset.\nOne thing you could do is to split and shuffle the data into files beforehand and then yield them from the dataloader, meaning the index going into the dataloader to get the data is the number of the file and not the row of the dataset, thus the length returned by the dataloader is the number of files.\n\n[Here are some suggestions](https://github.com/microsoft/LightGBM/issues/4672#issuecomment-941057024) for \"train a model on a larger-than-memory dataset on a single node\" by an LGBM Maintainer:\n- the Sequence Api you mentioned\n- using lightgbm.dask\n- Train a smaller model\n- Use Training Continuation\n\nWhat? Use Training Continuation? How?\nUnder the above link you'll find code.\nThe trick seems to be:  \"reference=first_dataset\" in the loop over files.\n```\ndataset = lgb.Dataset(\n        next_file,\n        params={\"label_column\": \"y\"},\n        free_raw_data=True,\n        reference=first_dataset\n    )\n```\nThis seems to [align the bin mapper for the dataset](https://github.com/microsoft/LightGBM/issues/4555#issuecomment-924533600).\nI'm not sure about how this influences the training continuation:\n> If the distribution of ~~val and train~~ data is very different, the bin mapper could be very different.\n\nI don't know it they handle that completely in their dataloader or if one has to check that the distribution of values in the samples is similar beforehand.\nIf you try this approch a followup would be nice. I'd like to know if that works for our data.\n\nAnother, maybe easier idea is to take an ensemble approach.\nSplit the data into multiple batches and train a model on each. \nThen at inference time simply make a prediction with each model using all the new data and take the average to get the final prediction."
  },
  "source": "meta"
}