{
  "id": 501170,
  "title": "How to train with much more data in lightgbm without running out of ram",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501170",
  "author_name": "",
  "post_date": "2024-05-08T10:34:25.516910800Z",
  "votes": 40,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Since this competition now has turned into \"hacking the metric as good as possible\" and I have lost interest, I will not keep this to myself.</p>\n<p>I've seen that lots of people having problem with out of ram error, and I've seen discussions where it is suggested that the memory consumption rises in the beginning and one should handle this by adjusting all kinds of model parameters. </p>\n<p>The reason this happens (at least in my case) is not due to model parameters, but rather that Lightgbm first transforms your data when dataset is constructed, see: <a href=\"https://github.com/microsoft/LightGBM/issues/1032\" target=\"_blank\">https://github.com/microsoft/LightGBM/issues/1032</a></p>\n<p>So what you can do is to utilize the fact that you can construct the lightgbm dataset before using train. This together with using the \"add_features_from()\" method makes it possible to break up the dataset into smaller chunks and construct each chunk independently and then add them back into a larger dataset. Like:</p>\n<p>`</p>\n<h1>break it up to not run out of memory</h1>\n<p>lgb_train_initial = lgb.Dataset(train[train_cols[:50]], train['target'],params=dataset_params1)<br>\nlgb_train_initial.construct()<br>\nlgb_train_initial2 = lgb.Dataset(train[train_cols[50:100]], train['target'],params=dataset_params1)<br>\nlgb_train_initial2.construct()<br>\nlgb_train_initial3 = lgb.Dataset(train[train_cols[100:150]], train['target'],params=dataset_params1)<br>\nlgb_train_initial3.construct()</p>\n<h1>combine them</h1>\n<p>lgb_train_initial.add_features_from(lgb_train_initial2)<br>\nlgb_train_initial.add_features_from(lgb_train_initial3)</p>\n<p>`</p>\n<p>Doing like this I can easily train with 800 features on 3 million rows etc</p>",
  "messages": [
    {
      "id": "2800741",
      "postDate": "05/08/2024 10:34:25",
      "content": "<p>Since this competition now has turned into \"hacking the metric as good as possible\" and I have lost interest, I will not keep this to myself.</p>\n<p>I've seen that lots of people having problem with out of ram error, and I've seen discussions where it is suggested that the memory consumption rises in the beginning and one should handle this by adjusting all kinds of model parameters. </p>\n<p>The reason this happens (at least in my case) is not due to model parameters, but rather that Lightgbm first transforms your data when dataset is constructed, see: <a href=\"https://github.com/microsoft/LightGBM/issues/1032\" target=\"_blank\">https://github.com/microsoft/LightGBM/issues/1032</a></p>\n<p>So what you can do is to utilize the fact that you can construct the lightgbm dataset before using train. This together with using the \"add_features_from()\" method makes it possible to break up the dataset into smaller chunks and construct each chunk independently and then add them back into a larger dataset. Like:</p>\n<p>`</p>\n<h1>break it up to not run out of memory</h1>\n<p>lgb_train_initial = lgb.Dataset(train[train_cols[:50]], train['target'],params=dataset_params1)<br>\nlgb_train_initial.construct()<br>\nlgb_train_initial2 = lgb.Dataset(train[train_cols[50:100]], train['target'],params=dataset_params1)<br>\nlgb_train_initial2.construct()<br>\nlgb_train_initial3 = lgb.Dataset(train[train_cols[100:150]], train['target'],params=dataset_params1)<br>\nlgb_train_initial3.construct()</p>\n<h1>combine them</h1>\n<p>lgb_train_initial.add_features_from(lgb_train_initial2)<br>\nlgb_train_initial.add_features_from(lgb_train_initial3)</p>\n<p>`</p>\n<p>Doing like this I can easily train with 800 features on 3 million rows etc</p>",
      "rawMarkdown": "Since this competition now has turned into \"hacking the metric as good as possible\" and I have lost interest, I will not keep this to myself.\n\nI've seen that lots of people having problem with out of ram error, and I've seen discussions where it is suggested that the memory consumption rises in the beginning and one should handle this by adjusting all kinds of model parameters. \n\nThe reason this happens (at least in my case) is not due to model parameters, but rather that Lightgbm first transforms your data when dataset is constructed, see: https://github.com/microsoft/LightGBM/issues/1032\n\nSo what you can do is to utilize the fact that you can construct the lightgbm dataset before using train. This together with using the \"add_features_from()\" method makes it possible to break up the dataset into smaller chunks and construct each chunk independently and then add them back into a larger dataset. Like:\n\n`\n# break it up to not run out of memory\nlgb_train_initial = lgb.Dataset(train[train_cols[:50]], train['target'],params=dataset_params1)\nlgb_train_initial.construct()\nlgb_train_initial2 = lgb.Dataset(train[train_cols[50:100]], train['target'],params=dataset_params1)\nlgb_train_initial2.construct()\nlgb_train_initial3 = lgb.Dataset(train[train_cols[100:150]], train['target'],params=dataset_params1)\nlgb_train_initial3.construct()\n\n# combine them\nlgb_train_initial.add_features_from(lgb_train_initial2)\nlgb_train_initial.add_features_from(lgb_train_initial3)\n\n`\n\nDoing like this I can easily train with 800 features on 3 million rows etc",
      "votes": null
    },
    {
      "id": "2807059",
      "postDate": "05/11/2024 13:25:50",
      "content": "<p>How do I start training, I try to start training but I get an error：</p>\n<p>model = lgb.LGBMClassifier(**params_gpu)<br>\nmodel.fit(lgbdataset,callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])<br>\nTraceback (most recent call last):<br>\n  File \"D:\\pycharm2024\\PyCharm Community Edition 2024.1\\plugins\\python-ce\\helpers\\pydev_pydevd_bundle\\pydevd_exec2.py\", line 3, in Exec<br>\n    exec(exp, global_vars, local_vars)<br>\n  File \"\", line 1, in <br>\nTypeError: fit() missing 1 required positional argument: 'y'</p>\n<p>But didn't I already enter the label when I built the dataset?</p>",
      "rawMarkdown": "How do I start training, I try to start training but I get an error：\n\nmodel = lgb.LGBMClassifier(**params_gpu)\nmodel.fit(lgbdataset,callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])\nTraceback (most recent call last):\n  File \"D:\\pycharm2024\\PyCharm Community Edition 2024.1\\plugins\\python-ce\\helpers\\pydev\\_pydevd_bundle\\pydevd_exec2.py\", line 3, in Exec\n    exec(exp, global_vars, local_vars)\n  File \"<input>\", line 1, in <module>\nTypeError: fit() missing 1 required positional argument: 'y'\n\nBut didn't I already enter the label when I built the dataset?",
      "votes": null
    },
    {
      "id": "2807077",
      "postDate": "05/11/2024 13:40:58",
      "content": "<p>It won't work with the lgb.LGBMClassifier, but rather with lgb.train(). Something like:</p>\n<p><code>lgb_model1 = lgb.train(params, lgb_train_initial)</code></p>\n<p>In params you can specify the objective etc. So in the end you will use predict (and not predict_proba) but if you specified the correct objective etc, it should be fine I believe. </p>\n<p>This one was quite useful for me: <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html</a></p>",
      "rawMarkdown": "It won't work with the lgb.LGBMClassifier, but rather with lgb.train(). Something like:\n\n`lgb_model1 = lgb.train(params, lgb_train_initial)`\n\nIn params you can specify the objective etc. So in the end you will use predict (and not predict_proba) but if you specified the correct objective etc, it should be fine I believe. \n\nThis one was quite useful for me: https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html",
      "votes": null
    },
    {
      "id": "2808402",
      "postDate": "05/12/2024 07:33:26",
      "content": "<p>Thank you so much it really helps a lot</p>",
      "rawMarkdown": "Thank you so much it really helps a lot",
      "votes": null
    },
    {
      "id": "2809253",
      "postDate": "05/12/2024 16:15:24",
      "content": "<p>Thank you so much it really helps a lot !</p>",
      "rawMarkdown": "Thank you so much it really helps a lot !",
      "votes": null
    },
    {
      "id": "2809483",
      "postDate": "05/12/2024 19:37:15",
      "content": "<p>Thanks a lot for the helpful suggestions! I found the out of memory issue more severe when I was using XGBoost. Using the same set of features, I was able to score using LGBM but failed with XGBoost <a href=\"https://www.kaggle.com/faithk7u/k7-cra-inference\" target=\"_blank\">(inference notebook link)</a>. Essentially everything is the same but I am using XGBoost, resulting an OOM error. Do you have any suggestions here? Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot for the helpful suggestions! I found the out of memory issue more severe when I was using XGBoost. Using the same set of features, I was able to score using LGBM but failed with XGBoost [(inference notebook link)](https://www.kaggle.com/faithk7u/k7-cra-inference). Essentially everything is the same but I am using XGBoost, resulting an OOM error. Do you have any suggestions here? Thanks a lot!",
      "votes": null
    },
    {
      "id": "2809593",
      "postDate": "05/12/2024 20:50:16",
      "content": "<p>No worries! No idea, I have not really tried XGBoost in this competition. Perhaps check the XGBoost documentation? (if you have not already done that) </p>",
      "rawMarkdown": "No worries! No idea, I have not really tried XGBoost in this competition. Perhaps check the XGBoost documentation? (if you have not already done that)",
      "votes": null
    },
    {
      "id": "2814655",
      "postDate": "05/15/2024 13:24:09",
      "content": "<p>Thank you for giving us your knowledge!!!</p>",
      "rawMarkdown": "Thank you for giving us your knowledge!!!",
      "votes": null
    },
    {
      "id": "2821603",
      "postDate": "05/18/2024 07:06:02",
      "content": "<p>thank you so much</p>",
      "rawMarkdown": "thank you so much",
      "votes": null
    },
    {
      "id": "2822025",
      "postDate": "05/18/2024 12:01:54",
      "content": "<p>For those that are also running models on their own computers, if you have windows, you can setup a higher paging file size(normal storage used as ram). This allowed me to test models with 1.5million rows and 4551 features, and then select the best ~900 features to use for the competition on the kaggle systems.</p>",
      "rawMarkdown": "For those that are also running models on their own computers, if you have windows, you can setup a higher paging file size(normal storage used as ram). This allowed me to test models with 1.5million rows and 4551 features, and then select the best ~900 features to use for the competition on the kaggle systems.",
      "votes": null
    },
    {
      "id": "2828993",
      "postDate": "05/22/2024 10:52:17",
      "content": "<p>XGBoost copies data into DMatrix internally, could not find a way to save memory with it</p>",
      "rawMarkdown": "XGBoost copies data into DMatrix internally, could not find a way to save memory with it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2807059,
      "author_name": "shuchangyu",
      "author_url": "",
      "post_date": "05/11/2024 13:25:50",
      "content": "<p>How do I start training, I try to start training but I get an error：</p>\n<p>model = lgb.LGBMClassifier(**params_gpu)<br>\nmodel.fit(lgbdataset,callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])<br>\nTraceback (most recent call last):<br>\n  File \"D:\\pycharm2024\\PyCharm Community Edition 2024.1\\plugins\\python-ce\\helpers\\pydev_pydevd_bundle\\pydevd_exec2.py\", line 3, in Exec<br>\n    exec(exp, global_vars, local_vars)<br>\n  File \"\", line 1, in <br>\nTypeError: fit() missing 1 required positional argument: 'y'</p>\n<p>But didn't I already enter the label when I built the dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2807077,
          "author_name": "ern711",
          "author_url": "",
          "post_date": "05/11/2024 13:40:58",
          "content": "<p>It won't work with the lgb.LGBMClassifier, but rather with lgb.train(). Something like:</p>\n<p><code>lgb_model1 = lgb.train(params, lgb_train_initial)</code></p>\n<p>In params you can specify the objective etc. So in the end you will use predict (and not predict_proba) but if you specified the correct objective etc, it should be fine I believe. </p>\n<p>This one was quite useful for me: <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2808402,
      "author_name": "jayeshkothavale",
      "author_url": "",
      "post_date": "05/12/2024 07:33:26",
      "content": "<p>Thank you so much it really helps a lot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2809253,
      "author_name": "sheemazain",
      "author_url": "",
      "post_date": "05/12/2024 16:15:24",
      "content": "<p>Thank you so much it really helps a lot !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2809483,
      "author_name": "faithk7u",
      "author_url": "",
      "post_date": "05/12/2024 19:37:15",
      "content": "<p>Thanks a lot for the helpful suggestions! I found the out of memory issue more severe when I was using XGBoost. Using the same set of features, I was able to score using LGBM but failed with XGBoost <a href=\"https://www.kaggle.com/faithk7u/k7-cra-inference\" target=\"_blank\">(inference notebook link)</a>. Essentially everything is the same but I am using XGBoost, resulting an OOM error. Do you have any suggestions here? Thanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2809593,
          "author_name": "ern711",
          "author_url": "",
          "post_date": "05/12/2024 20:50:16",
          "content": "<p>No worries! No idea, I have not really tried XGBoost in this competition. Perhaps check the XGBoost documentation? (if you have not already done that) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2828993,
          "author_name": "bluepill",
          "author_url": "",
          "post_date": "05/22/2024 10:52:17",
          "content": "<p>XGBoost copies data into DMatrix internally, could not find a way to save memory with it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2814655,
      "author_name": "yukinari298",
      "author_url": "",
      "post_date": "05/15/2024 13:24:09",
      "content": "<p>Thank you for giving us your knowledge!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2821603,
      "author_name": "alexanderegorovmephi",
      "author_url": "",
      "post_date": "05/18/2024 07:06:02",
      "content": "<p>thank you so much</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2822025,
      "author_name": "timotejfasiang",
      "author_url": "",
      "post_date": "05/18/2024 12:01:54",
      "content": "<p>For those that are also running models on their own computers, if you have windows, you can setup a higher paging file size(normal storage used as ram). This allowed me to test models with 1.5million rows and 4551 features, and then select the best ~900 features to use for the competition on the kaggle systems.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2800741": "Since this competition now has turned into \"hacking the metric as good as possible\" and I have lost interest, I will not keep this to myself.\n\nI've seen that lots of people having problem with out of ram error, and I've seen discussions where it is suggested that the memory consumption rises in the beginning and one should handle this by adjusting all kinds of model parameters. \n\nThe reason this happens (at least in my case) is not due to model parameters, but rather that Lightgbm first transforms your data when dataset is constructed, see: https://github.com/microsoft/LightGBM/issues/1032\n\nSo what you can do is to utilize the fact that you can construct the lightgbm dataset before using train. This together with using the \"add_features_from()\" method makes it possible to break up the dataset into smaller chunks and construct each chunk independently and then add them back into a larger dataset. Like:\n\n`\n# break it up to not run out of memory\nlgb_train_initial = lgb.Dataset(train[train_cols[:50]], train['target'],params=dataset_params1)\nlgb_train_initial.construct()\nlgb_train_initial2 = lgb.Dataset(train[train_cols[50:100]], train['target'],params=dataset_params1)\nlgb_train_initial2.construct()\nlgb_train_initial3 = lgb.Dataset(train[train_cols[100:150]], train['target'],params=dataset_params1)\nlgb_train_initial3.construct()\n\n# combine them\nlgb_train_initial.add_features_from(lgb_train_initial2)\nlgb_train_initial.add_features_from(lgb_train_initial3)\n\n`\n\nDoing like this I can easily train with 800 features on 3 million rows etc",
    "2807059": "How do I start training, I try to start training but I get an error：\n\nmodel = lgb.LGBMClassifier(**params_gpu)\nmodel.fit(lgbdataset,callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])\nTraceback (most recent call last):\n  File \"D:\\pycharm2024\\PyCharm Community Edition 2024.1\\plugins\\python-ce\\helpers\\pydev\\_pydevd_bundle\\pydevd_exec2.py\", line 3, in Exec\n    exec(exp, global_vars, local_vars)\n  File \"<input>\", line 1, in <module>\nTypeError: fit() missing 1 required positional argument: 'y'\n\nBut didn't I already enter the label when I built the dataset?",
    "2807077": "It won't work with the lgb.LGBMClassifier, but rather with lgb.train(). Something like:\n\n`lgb_model1 = lgb.train(params, lgb_train_initial)`\n\nIn params you can specify the objective etc. So in the end you will use predict (and not predict_proba) but if you specified the correct objective etc, it should be fine I believe. \n\nThis one was quite useful for me: https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html",
    "2808402": "Thank you so much it really helps a lot",
    "2809253": "Thank you so much it really helps a lot !",
    "2809483": "Thanks a lot for the helpful suggestions! I found the out of memory issue more severe when I was using XGBoost. Using the same set of features, I was able to score using LGBM but failed with XGBoost [(inference notebook link)](https://www.kaggle.com/faithk7u/k7-cra-inference). Essentially everything is the same but I am using XGBoost, resulting an OOM error. Do you have any suggestions here? Thanks a lot!",
    "2809593": "No worries! No idea, I have not really tried XGBoost in this competition. Perhaps check the XGBoost documentation? (if you have not already done that)",
    "2814655": "Thank you for giving us your knowledge!!!",
    "2821603": "thank you so much",
    "2822025": "For those that are also running models on their own computers, if you have windows, you can setup a higher paging file size(normal storage used as ram). This allowed me to test models with 1.5million rows and 4551 features, and then select the best ~900 features to use for the competition on the kaggle systems.",
    "2828993": "XGBoost copies data into DMatrix internally, could not find a way to save memory with it"
  },
  "source": "meta"
}