{
  "id": 502182,
  "title": "What is the maximal number of features that you are able to use without running into the Out of Memory error for LGBM?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/502182",
  "author_name": "",
  "post_date": "2024-05-12T12:43:59.968764300Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>For me it is around 520 without much optimization, but would like to know some upper bounds from other people. </p>",
  "messages": [
    {
      "id": "2808895",
      "postDate": "05/12/2024 12:43:59",
      "content": "<p>For me it is around 520 without much optimization, but would like to know some upper bounds from other people. </p>",
      "rawMarkdown": "For me it is around 520 without much optimization, but would like to know some upper bounds from other people.",
      "votes": null
    },
    {
      "id": "2808933",
      "postDate": "05/12/2024 13:08:39",
      "content": "<p>I am uncertain whether it's the limit, but I can run without OOM errors a notebook with:</p>\n<blockquote>\n  <p>The shape of the train data is (1526659, 610)<br>\n  Memory usage of train data is 2386.31 MB</p>\n</blockquote>\n<p>Note that I train and infer LGB seperately. </p>",
      "rawMarkdown": "I am uncertain whether it's the limit, but I can run without OOM errors a notebook with:\n\n> The shape of the train data is (1526659, 610)\n> Memory usage of train data is 2386.31 MB\n\nNote that I train and infer LGB seperately.",
      "votes": null
    },
    {
      "id": "2809013",
      "postDate": "05/12/2024 13:42:32",
      "content": "<p>about 700, <a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-training\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-training</a></p>",
      "rawMarkdown": "about 700, https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-training",
      "votes": null
    },
    {
      "id": "2809016",
      "postDate": "05/12/2024 13:43:28",
      "content": "<p>Hey guys, As a general guideline, if you're experiencing memory issues, you can try reducing the number of features by removing irrelevant or redundant features through feature selection techniques. Additionally, you can consider using LGBM's parameters such as max_bin and max_depth to control the memory usage during training. Technically, there isn't a hard maximum number of features that LightGBM can handle. </p>",
      "rawMarkdown": "Hey guys, As a general guideline, if you're experiencing memory issues, you can try reducing the number of features by removing irrelevant or redundant features through feature selection techniques. Additionally, you can consider using LGBM's parameters such as max_bin and max_depth to control the memory usage during training. Technically, there isn't a hard maximum number of features that LightGBM can handle.",
      "votes": null
    },
    {
      "id": "2809048",
      "postDate": "05/12/2024 13:53:03",
      "content": "<p>More than 2000</p>",
      "rawMarkdown": "More than 2000",
      "votes": null
    },
    {
      "id": "2809140",
      "postDate": "05/12/2024 14:49:48",
      "content": "<p>Thanks, will try calling gc.collect() proactively and see if I can include more features. </p>",
      "rawMarkdown": "Thanks, will try calling gc.collect() proactively and see if I can include more features.",
      "votes": null
    },
    {
      "id": "2809153",
      "postDate": "05/12/2024 14:55:43",
      "content": "<p>Thanks. I also did the training and inference separately, but during inference also included the training dataset because the columns used during inference have to be same as the ones used during training for the model. </p>\n<p>But I realized that one can just save the columns to a .csv file and load that file directly during inference, to avoid loading the training dataset. Will try that and see how it goes.  </p>",
      "rawMarkdown": "Thanks. I also did the training and inference separately, but during inference also included the training dataset because the columns used during inference have to be same as the ones used during training for the model. \n\nBut I realized that one can just save the columns to a .csv file and load that file directly during inference, to avoid loading the training dataset. Will try that and see how it goes.",
      "votes": null
    },
    {
      "id": "2809644",
      "postDate": "05/12/2024 21:43:01",
      "content": "<p>Or save the list of the columns, or even better - process in a pipeline and then apply the pipeline to the test</p>",
      "rawMarkdown": "Or save the list of the columns, or even better - process in a pipeline and then apply the pipeline to the test",
      "votes": null
    },
    {
      "id": "2810048",
      "postDate": "05/13/2024 05:46:03",
      "content": "<p>Additionally, the next approach can be used:</p>\n<pre><code>columns = model.feature_name_\npred = model.predict_proba(df[columns])[:, ]\n</code></pre>",
      "rawMarkdown": "Additionally, the next approach can be used:\n```python\ncolumns = model.feature_name_\npred = model.predict_proba(df[columns])[:, 1]\n```",
      "votes": null
    },
    {
      "id": "2817311",
      "postDate": "05/16/2024 20:13:36",
      "content": "<p><a href=\"https://www.kaggle.com/skrrydg\" target=\"_blank\">@skrrydg</a>, that's impressive. May I know how did you attempt to do the memory management during the inference stage in order to use this many features?</p>\n<p>I tried the below: </p>\n<ol>\n<li>Actively call <code>del var; gc.collect()</code> after being done with using that var</li>\n<li>Predict in batches </li>\n<li>Separate the training and inference into two notebooks</li>\n<li>Since the purpose of including the train dataset during the inference is just to select the features used during the training stage, instead of loading the whole train dataset, I just save the train feature columns to a .csv file, and load that .csv file during inference</li>\n</ol>\n<p>With the above approaches the model can predict with 770+ features in the private dataset without any problems, but when the number of features increases to 1100+, the notebook encounters OOM error during the submission. </p>\n<p>Do you have any other suggestions to reduce the memory usage? Thanks a lot!</p>",
      "rawMarkdown": "skrrydg, that's impressive. May I know how did you attempt to do the memory management during the inference stage in order to use this many features?\n\nI tried the below: \n1. Actively call `del var; gc.collect()` after being done with using that var\n2. Predict in batches \n3. Separate the training and inference into two notebooks\n4. Since the purpose of including the train dataset during the inference is just to select the features used during the training stage, instead of loading the whole train dataset, I just save the train feature columns to a .csv file, and load that .csv file during inference\n\nWith the above approaches the model can predict with 770+ features in the private dataset without any problems, but when the number of features increases to 1100+, the notebook encounters OOM error during the submission. \n\nDo you have any other suggestions to reduce the memory usage? Thanks a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2808933,
      "author_name": "andreasbis",
      "author_url": "",
      "post_date": "05/12/2024 13:08:39",
      "content": "<p>I am uncertain whether it's the limit, but I can run without OOM errors a notebook with:</p>\n<blockquote>\n  <p>The shape of the train data is (1526659, 610)<br>\n  Memory usage of train data is 2386.31 MB</p>\n</blockquote>\n<p>Note that I train and infer LGB seperately. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2809153,
          "author_name": "faithk7u",
          "author_url": "",
          "post_date": "05/12/2024 14:55:43",
          "content": "<p>Thanks. I also did the training and inference separately, but during inference also included the training dataset because the columns used during inference have to be same as the ones used during training for the model. </p>\n<p>But I realized that one can just save the columns to a .csv file and load that file directly during inference, to avoid loading the training dataset. Will try that and see how it goes.  </p>",
          "votes": null,
          "replies": [
            {
              "id": 2809644,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "05/12/2024 21:43:01",
              "content": "<p>Or save the list of the columns, or even better - process in a pipeline and then apply the pipeline to the test</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2810048,
              "author_name": "andreynesterov",
              "author_url": "",
              "post_date": "05/13/2024 05:46:03",
              "content": "<p>Additionally, the next approach can be used:</p>\n<pre><code>columns = model.feature_name_\npred = model.predict_proba(df[columns])[:, ]\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2809013,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "05/12/2024 13:42:32",
      "content": "<p>about 700, <a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-training\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-training</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2809140,
          "author_name": "faithk7u",
          "author_url": "",
          "post_date": "05/12/2024 14:49:48",
          "content": "<p>Thanks, will try calling gc.collect() proactively and see if I can include more features. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2809016,
      "author_name": "masroormohajerani",
      "author_url": "",
      "post_date": "05/12/2024 13:43:28",
      "content": "<p>Hey guys, As a general guideline, if you're experiencing memory issues, you can try reducing the number of features by removing irrelevant or redundant features through feature selection techniques. Additionally, you can consider using LGBM's parameters such as max_bin and max_depth to control the memory usage during training. Technically, there isn't a hard maximum number of features that LightGBM can handle. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2809048,
      "author_name": "skrrydg",
      "author_url": "",
      "post_date": "05/12/2024 13:53:03",
      "content": "<p>More than 2000</p>",
      "votes": null,
      "replies": [
        {
          "id": 2817311,
          "author_name": "faithk7u",
          "author_url": "",
          "post_date": "05/16/2024 20:13:36",
          "content": "<p><a href=\"https://www.kaggle.com/skrrydg\" target=\"_blank\">@skrrydg</a>, that's impressive. May I know how did you attempt to do the memory management during the inference stage in order to use this many features?</p>\n<p>I tried the below: </p>\n<ol>\n<li>Actively call <code>del var; gc.collect()</code> after being done with using that var</li>\n<li>Predict in batches </li>\n<li>Separate the training and inference into two notebooks</li>\n<li>Since the purpose of including the train dataset during the inference is just to select the features used during the training stage, instead of loading the whole train dataset, I just save the train feature columns to a .csv file, and load that .csv file during inference</li>\n</ol>\n<p>With the above approaches the model can predict with 770+ features in the private dataset without any problems, but when the number of features increases to 1100+, the notebook encounters OOM error during the submission. </p>\n<p>Do you have any other suggestions to reduce the memory usage? Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2808895": "For me it is around 520 without much optimization, but would like to know some upper bounds from other people.",
    "2808933": "I am uncertain whether it's the limit, but I can run without OOM errors a notebook with:\n\n> The shape of the train data is (1526659, 610)\n> Memory usage of train data is 2386.31 MB\n\nNote that I train and infer LGB seperately.",
    "2809013": "about 700, https://www.kaggle.com/code/yunsuxiaozi/home-credit-lgbm-with-677-features-training",
    "2809016": "Hey guys, As a general guideline, if you're experiencing memory issues, you can try reducing the number of features by removing irrelevant or redundant features through feature selection techniques. Additionally, you can consider using LGBM's parameters such as max_bin and max_depth to control the memory usage during training. Technically, there isn't a hard maximum number of features that LightGBM can handle.",
    "2809048": "More than 2000",
    "2809140": "Thanks, will try calling gc.collect() proactively and see if I can include more features.",
    "2809153": "Thanks. I also did the training and inference separately, but during inference also included the training dataset because the columns used during inference have to be same as the ones used during training for the model. \n\nBut I realized that one can just save the columns to a .csv file and load that file directly during inference, to avoid loading the training dataset. Will try that and see how it goes.",
    "2809644": "Or save the list of the columns, or even better - process in a pipeline and then apply the pipeline to the test",
    "2810048": "Additionally, the next approach can be used:\n```python\ncolumns = model.feature_name_\npred = model.predict_proba(df[columns])[:, 1]\n```",
    "2817311": "skrrydg, that's impressive. May I know how did you attempt to do the memory management during the inference stage in order to use this many features?\n\nI tried the below: \n1. Actively call `del var; gc.collect()` after being done with using that var\n2. Predict in batches \n3. Separate the training and inference into two notebooks\n4. Since the purpose of including the train dataset during the inference is just to select the features used during the training stage, instead of loading the whole train dataset, I just save the train feature columns to a .csv file, and load that .csv file during inference\n\nWith the above approaches the model can predict with 770+ features in the private dataset without any problems, but when the number of features increases to 1100+, the notebook encounters OOM error during the submission. \n\nDo you have any other suggestions to reduce the memory usage? Thanks a lot!"
  },
  "source": "meta"
}