{
  "id": 356956,
  "title": "Ideas to use all train datasets for continues training",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356956",
  "author_name": "",
  "post_date": "2022-10-02T17:10:44.510463200Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was looking for how to use all training data provided in competition without having out of memory problem. Found out some features provided by LightGBM.</p>\n<p>LightGBM offers two options for continue training:</p>\n<ul>\n<li><p><strong>Refit:</strong> Uses existing tree structures to fit the new data. It doesn't add new trees and keeps model structure of a trained model. Only updates leaf counts and leaf values with new data. </p></li>\n<li><p><strong>Train:</strong> with init_model parameter continued training will add new trees to the existing trained model with new data. </p></li>\n</ul>\n<p><a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html\" target=\"_blank\">LightGBM Train Doc link</a></p>\n<p>My notebook which I tried 'train' method with LightGBM, using all 10 train datasets here: <a href=\"https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm</a></p>\n<p>Also XGboost offers a similar method to LightGBM train called XGBoost train with xgb_model parameter.</p>\n<ul>\n<li><strong>xgb_model</strong> (Optional[Union[str, PathLike, Booster, bytearray]]) – Xgb model to be loaded before training (allows training continuation).</li>\n</ul>\n<p><a href=\"https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.training\" target=\"_blank\">XGBoost Learning API Doc link</a></p>\n<p>Please comment any feedbacks, alternative methods and ideas.</p>",
  "messages": [
    {
      "id": "1967773",
      "postDate": "10/02/2022 17:10:44",
      "content": "<p>I was looking for how to use all training data provided in competition without having out of memory problem. Found out some features provided by LightGBM.</p>\n<p>LightGBM offers two options for continue training:</p>\n<ul>\n<li><p><strong>Refit:</strong> Uses existing tree structures to fit the new data. It doesn't add new trees and keeps model structure of a trained model. Only updates leaf counts and leaf values with new data. </p></li>\n<li><p><strong>Train:</strong> with init_model parameter continued training will add new trees to the existing trained model with new data. </p></li>\n</ul>\n<p><a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html\" target=\"_blank\">LightGBM Train Doc link</a></p>\n<p>My notebook which I tried 'train' method with LightGBM, using all 10 train datasets here: <a href=\"https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm</a></p>\n<p>Also XGboost offers a similar method to LightGBM train called XGBoost train with xgb_model parameter.</p>\n<ul>\n<li><strong>xgb_model</strong> (Optional[Union[str, PathLike, Booster, bytearray]]) – Xgb model to be loaded before training (allows training continuation).</li>\n</ul>\n<p><a href=\"https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.training\" target=\"_blank\">XGBoost Learning API Doc link</a></p>\n<p>Please comment any feedbacks, alternative methods and ideas.</p>",
      "rawMarkdown": "I was looking for how to use all training data provided in competition without having out of memory problem. Found out some features provided by LightGBM.\n\nLightGBM offers two options for continue training:\n\n- **Refit:** Uses existing tree structures to fit the new data. It doesn't add new trees and keeps model structure of a trained model. Only updates leaf counts and leaf values with new data. \n\n- **Train:** with init_model parameter continued training will add new trees to the existing trained model with new data. \n\n[LightGBM Train Doc link](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html)\n\nMy notebook which I tried 'train' method with LightGBM, using all 10 train datasets here: [https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm](https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm)\n\n\nAlso XGboost offers a similar method to LightGBM train called XGBoost train with xgb_model parameter.\n\n- **xgb_model** (Optional[Union[str, PathLike, Booster, bytearray]]) – Xgb model to be loaded before training (allows training continuation).\n\n[XGBoost Learning API Doc link](https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.training)\n\n\nPlease comment any feedbacks, alternative methods and ideas.",
      "votes": null
    },
    {
      "id": "1967802",
      "postDate": "10/02/2022 17:35:14",
      "content": "<p>Great observations <a href=\"https://www.kaggle.com/landfallmotto\" target=\"_blank\">@landfallmotto</a>, I think this will be extremely useful for people using batch ML models. <br>\nMany thanks!</p>",
      "rawMarkdown": "Great observations @landfallmotto, I think this will be extremely useful for people using batch ML models. \nMany thanks!",
      "votes": null
    },
    {
      "id": "1968116",
      "postDate": "10/02/2022 21:39:38",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/landfallmotto\" target=\"_blank\">@landfallmotto</a> Feel free to check out my notebooks <a href=\"https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-baseline\" target=\"_blank\">this one</a> and <a href=\"https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-catboost\" target=\"_blank\">this one</a> where we use boosting training for the gradient boosting. We manage to show that the results are as good as or better than using all the data at once. Thanks.</p>",
      "rawMarkdown": "Hi @landfallmotto Feel free to check out my notebooks [this one](https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-baseline) and [this one](https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-catboost) where we use boosting training for the gradient boosting. We manage to show that the results are as good as or better than using all the data at once. Thanks.",
      "votes": null
    },
    {
      "id": "1969041",
      "postDate": "10/03/2022 10:10:29",
      "content": "<p>This topic has covered my concept in machine learning keep it going bro and i will be glad if you will also check some of my notebooks.</p>",
      "rawMarkdown": "This topic has covered my concept in machine learning keep it going bro and i will be glad if you will also check some of my notebooks.",
      "votes": null
    },
    {
      "id": "1971263",
      "postDate": "10/04/2022 13:55:14",
      "content": "<p>Continues training, a nice topic for this month.</p>\n<p>Valid_0's binary_logloss is always the same 0.185145. No improvement with additional trainings?</p>",
      "rawMarkdown": "Continues training, a nice topic for this month.\n\nValid_0's binary_logloss is always the same 0.185145. No improvement with additional trainings?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1967802,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/02/2022 17:35:14",
      "content": "<p>Great observations <a href=\"https://www.kaggle.com/landfallmotto\" target=\"_blank\">@landfallmotto</a>, I think this will be extremely useful for people using batch ML models. <br>\nMany thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1968116,
      "author_name": "stautxie",
      "author_url": "",
      "post_date": "10/02/2022 21:39:38",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/landfallmotto\" target=\"_blank\">@landfallmotto</a> Feel free to check out my notebooks <a href=\"https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-baseline\" target=\"_blank\">this one</a> and <a href=\"https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-catboost\" target=\"_blank\">this one</a> where we use boosting training for the gradient boosting. We manage to show that the results are as good as or better than using all the data at once. Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1969041,
      "author_name": "faisaljanjua0555",
      "author_url": "",
      "post_date": "10/03/2022 10:10:29",
      "content": "<p>This topic has covered my concept in machine learning keep it going bro and i will be glad if you will also check some of my notebooks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1971263,
      "author_name": "stevesimons",
      "author_url": "",
      "post_date": "10/04/2022 13:55:14",
      "content": "<p>Continues training, a nice topic for this month.</p>\n<p>Valid_0's binary_logloss is always the same 0.185145. No improvement with additional trainings?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1967773": "I was looking for how to use all training data provided in competition without having out of memory problem. Found out some features provided by LightGBM.\n\nLightGBM offers two options for continue training:\n\n- **Refit:** Uses existing tree structures to fit the new data. It doesn't add new trees and keeps model structure of a trained model. Only updates leaf counts and leaf values with new data. \n\n- **Train:** with init_model parameter continued training will add new trees to the existing trained model with new data. \n\n[LightGBM Train Doc link](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html)\n\nMy notebook which I tried 'train' method with LightGBM, using all 10 train datasets here: [https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm](https://www.kaggle.com/code/landfallmotto/tps-oct-22-continue-training-method-lightgbm)\n\n\nAlso XGboost offers a similar method to LightGBM train called XGBoost train with xgb_model parameter.\n\n- **xgb_model** (Optional[Union[str, PathLike, Booster, bytearray]]) – Xgb model to be loaded before training (allows training continuation).\n\n[XGBoost Learning API Doc link](https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.training)\n\n\nPlease comment any feedbacks, alternative methods and ideas.",
    "1967802": "Great observations @landfallmotto, I think this will be extremely useful for people using batch ML models. \nMany thanks!",
    "1968116": "Hi @landfallmotto Feel free to check out my notebooks [this one](https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-baseline) and [this one](https://www.kaggle.com/code/stautxie/tabular-playground-oct-2022-catboost) where we use boosting training for the gradient boosting. We manage to show that the results are as good as or better than using all the data at once. Thanks.",
    "1969041": "This topic has covered my concept in machine learning keep it going bro and i will be glad if you will also check some of my notebooks.",
    "1971263": "Continues training, a nice topic for this month.\n\nValid_0's binary_logloss is always the same 0.185145. No improvement with additional trainings?"
  },
  "source": "meta"
}