{
  "id": 202606,
  "title": "How to reduce the memory cost when training LGBM with large amount of features?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202606",
  "author_name": "",
  "post_date": "2020-12-11T01:40:27.478489800Z",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I read several winner solutions using single LGBM in other competitions. They generated even more than 10K features. But in my code, memory cost easily reach 80G when training a model only with 40+ features. I could hardly imagine how large the memory will cost if I generate such amount of features. </p>\n<p>Is there a way to reduce the memory cost with large amount of features?</p>",
  "messages": [
    {
      "id": "1108770",
      "postDate": "12/11/2020 01:40:27",
      "content": "<p>I read several winner solutions using single LGBM in other competitions. They generated even more than 10K features. But in my code, memory cost easily reach 80G when training a model only with 40+ features. I could hardly imagine how large the memory will cost if I generate such amount of features. </p>\n<p>Is there a way to reduce the memory cost with large amount of features?</p>",
      "rawMarkdown": "I read several winner solutions using single LGBM in other competitions. They generated even more than 10K features. But in my code, memory cost easily reach 80G when training a model only with 40+ features. I could hardly imagine how large the memory will cost if I generate such amount of features. \n\nIs there a way to reduce the memory cost with large amount of features?",
      "votes": null
    },
    {
      "id": "1108960",
      "postDate": "12/11/2020 07:33:29",
      "content": "<p>Try casting feature data types to np.float16 before lgb.Dataset.. From my experiment, with 9kw training samples, around 100 features setup, you need about 50GB for training LGB model. </p>",
      "rawMarkdown": "Try casting feature data types to np.float16 before lgb.Dataset.. From my experiment, with 9kw training samples, around 100 features setup, you need about 50GB for training LGB model.",
      "votes": null
    },
    {
      "id": "1108992",
      "postDate": "12/11/2020 08:17:21",
      "content": "<p>Thanks, Rocky. I will try that. </p>",
      "rawMarkdown": "Thanks, Rocky. I will try that.",
      "votes": null
    },
    {
      "id": "1109005",
      "postDate": "12/11/2020 08:37:16",
      "content": "<p>Do you create your lgb.Dataset from a pandas dataframe or a numpy array?<br>\nIn case you are using a pandas dataframe, try converting the dataframe to a numpy array of data type float32.<br>\nI am currently using 50M training rows with 15 features and RAM usage is ~6GB<br>\nA discussion on creating a numpy array from a dataframe in an efficient way can be found <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Do you create your lgb.Dataset from a pandas dataframe or a numpy array?\nIn case you are using a pandas dataframe, try converting the dataframe to a numpy array of data type float32.\nI am currently using 50M training rows with 15 features and RAM usage is ~6GB\nA discussion on creating a numpy array from a dataframe in an efficient way can be found [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245)",
      "votes": null
    },
    {
      "id": "1109063",
      "postDate": "12/11/2020 09:48:05",
      "content": "<p>Thank you! I will try to convert df to np.array. </p>",
      "rawMarkdown": "Thank you! I will try to convert df to np.array.",
      "votes": null
    },
    {
      "id": "1110295",
      "postDate": "12/12/2020 15:48:47",
      "content": "<p>Check out these sources:</p>\n<ul>\n<li><a href=\"https://lightgbm.readthedocs.io/en/latest/FAQ.html\" target=\"_blank\">General LightGBM Questions</a></li>\n<li><a href=\"https://towardsdatascience.com/how-to-learn-from-bigdata-files-on-low-memory-incremental-learning-d377282d38ff\" target=\"_blank\">How to handle BigData Files on Low Memory?</a></li>\n</ul>",
      "rawMarkdown": "Check out these sources:\n- [General LightGBM Questions](https://lightgbm.readthedocs.io/en/latest/FAQ.html)\n- [How to handle BigData Files on Low Memory?](https://towardsdatascience.com/how-to-learn-from-bigdata-files-on-low-memory-incremental-learning-d377282d38ff)",
      "votes": null
    },
    {
      "id": "1111703",
      "postDate": "12/13/2020 23:36:09",
      "content": "<p><a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> <br>\nif you have some feature in pandas with dtype int8 ( or int16) and convert them to numpy float32 as you mentioned, wouldn't that increase memory usages ?</p>",
      "rawMarkdown": "markwijkhuizen \nif you have some feature in pandas with dtype int8 ( or int16) and convert them to numpy float32 as you mentioned, wouldn't that increase memory usages ?",
      "votes": null
    },
    {
      "id": "1112148",
      "postDate": "12/14/2020 10:28:31",
      "content": "<p>LighGBM converts all data to a numpy ndarray of type <strong>float32</strong> regardless the original data type, it will thus make a copy when using other data types. When creating the float32 ndarray yourself LightGBM will not make a copy but just use the given ndarray.<br>\nFor pandas DataFrames all data is copied to a new DataFrame of type <strong>float64</strong>, this is why you should avoid DataFrames when creating LightBM Datasets. If you would use int8 features in a Dataframe a new DataFrame will be created where this feature is converted to float64.</p>\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325\" target=\"_blank\">Here</a> is a discussion about LightGBM Datasets.</p>",
      "rawMarkdown": "LighGBM converts all data to a numpy ndarray of type **float32** regardless the original data type, it will thus make a copy when using other data types. When creating the float32 ndarray yourself LightGBM will not make a copy but just use the given ndarray.\nFor pandas DataFrames all data is copied to a new DataFrame of type **float64**, this is why you should avoid DataFrames when creating LightBM Datasets. If you would use int8 features in a Dataframe a new DataFrame will be created where this feature is converted to float64.\n\n[Here](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325) is a discussion about LightGBM Datasets.",
      "votes": null
    },
    {
      "id": "1112180",
      "postDate": "12/14/2020 10:55:10",
      "content": "<p>Aha, thats great explanation for the memory problems I have !</p>\n<p>That means it will be usesless to reduce the size of my pandas columns (to int8 or int16) because lgbm will be convert it to np.float any way …<br>\nAnd even it will consume more memory because I would have two copies of same data (original DF and lgbm np.float), right ?</p>",
      "rawMarkdown": "Aha, thats great explanation for the memory problems I have !\n\nThat means it will be usesless to reduce the size of my pandas columns (to int8 or int16) because lgbm will be convert it to np.float any way ...\nAnd even it will consume more memory because I would have two copies of same data (original DF and lgbm np.float), right ?",
      "votes": null
    },
    {
      "id": "1112516",
      "postDate": "12/14/2020 17:11:25",
      "content": "<p>Yes, if you have your features in a DataFrame, convert the DataFrame to a float32 ndarray and delete the DataFrame.</p>",
      "rawMarkdown": "Yes, if you have your features in a DataFrame, convert the DataFrame to a float32 ndarray and delete the DataFrame.",
      "votes": null
    },
    {
      "id": "1112540",
      "postDate": "12/14/2020 17:34:26",
      "content": "<p>As Mark says, it is really important to use np.float32 when constructing lgb.Dataset.<br>\nBut I think loading features from disk is more memory-efficient than converting pandas dataframe to numpy array. This is my code to save memory.</p>\n<pre><code>def load_arr(feature_names, data_type, n_rows = 99271300, path='../features/'):\n    X = np.zeros((n_rows, len(feature_names)), dtype = np.float32)\n    for i in tqdm(range(len(feature_names))):\n        X[:, i] = np.load(os.path.join(path, feature_names[i]) + '_' + data_type + '.npy')\n    return X\n</code></pre>",
      "rawMarkdown": "As Mark says, it is really important to use np.float32 when constructing lgb.Dataset.\nBut I think loading features from disk is more memory-efficient than converting pandas dataframe to numpy array. This is my code to save memory.\n```\ndef load_arr(feature_names, data_type, n_rows = 99271300, path='../features/'):\n    X = np.zeros((n_rows, len(feature_names)), dtype = np.float32)\n    for i in tqdm(range(len(feature_names))):\n        X[:, i] = np.load(os.path.join(path, feature_names[i]) + '_' + data_type + '.npy')\n    return X\n```",
      "votes": null
    },
    {
      "id": "1119019",
      "postDate": "12/19/2020 16:26:07",
      "content": "<p>Thanks for your kind share, <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> . It is really an awesome way to reduce memory cost. I will try to use it.</p>",
      "rawMarkdown": "Thanks for your kind share, @mamasinkgs . It is really an awesome way to reduce memory cost. I will try to use it.",
      "votes": null
    },
    {
      "id": "1119020",
      "postDate": "12/19/2020 16:26:46",
      "content": "<p>Thank you! I will read them. </p>",
      "rawMarkdown": "Thank you! I will read them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1108960,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "12/11/2020 07:33:29",
      "content": "<p>Try casting feature data types to np.float16 before lgb.Dataset.. From my experiment, with 9kw training samples, around 100 features setup, you need about 50GB for training LGB model. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1108992,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "12/11/2020 08:17:21",
          "content": "<p>Thanks, Rocky. I will try that. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109005,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "12/11/2020 08:37:16",
      "content": "<p>Do you create your lgb.Dataset from a pandas dataframe or a numpy array?<br>\nIn case you are using a pandas dataframe, try converting the dataframe to a numpy array of data type float32.<br>\nI am currently using 50M training rows with 15 features and RAM usage is ~6GB<br>\nA discussion on creating a numpy array from a dataframe in an efficient way can be found <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1109063,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "12/11/2020 09:48:05",
          "content": "<p>Thank you! I will try to convert df to np.array. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1111703,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "12/13/2020 23:36:09",
          "content": "<p><a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> <br>\nif you have some feature in pandas with dtype int8 ( or int16) and convert them to numpy float32 as you mentioned, wouldn't that increase memory usages ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1112148,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "12/14/2020 10:28:31",
          "content": "<p>LighGBM converts all data to a numpy ndarray of type <strong>float32</strong> regardless the original data type, it will thus make a copy when using other data types. When creating the float32 ndarray yourself LightGBM will not make a copy but just use the given ndarray.<br>\nFor pandas DataFrames all data is copied to a new DataFrame of type <strong>float64</strong>, this is why you should avoid DataFrames when creating LightBM Datasets. If you would use int8 features in a Dataframe a new DataFrame will be created where this feature is converted to float64.</p>\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325\" target=\"_blank\">Here</a> is a discussion about LightGBM Datasets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1112180,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "12/14/2020 10:55:10",
          "content": "<p>Aha, thats great explanation for the memory problems I have !</p>\n<p>That means it will be usesless to reduce the size of my pandas columns (to int8 or int16) because lgbm will be convert it to np.float any way …<br>\nAnd even it will consume more memory because I would have two copies of same data (original DF and lgbm np.float), right ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1112516,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "12/14/2020 17:11:25",
          "content": "<p>Yes, if you have your features in a DataFrame, convert the DataFrame to a float32 ndarray and delete the DataFrame.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1112540,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/14/2020 17:34:26",
          "content": "<p>As Mark says, it is really important to use np.float32 when constructing lgb.Dataset.<br>\nBut I think loading features from disk is more memory-efficient than converting pandas dataframe to numpy array. This is my code to save memory.</p>\n<pre><code>def load_arr(feature_names, data_type, n_rows = 99271300, path='../features/'):\n    X = np.zeros((n_rows, len(feature_names)), dtype = np.float32)\n    for i in tqdm(range(len(feature_names))):\n        X[:, i] = np.load(os.path.join(path, feature_names[i]) + '_' + data_type + '.npy')\n    return X\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119019,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "12/19/2020 16:26:07",
          "content": "<p>Thanks for your kind share, <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> . It is really an awesome way to reduce memory cost. I will try to use it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1110295,
      "author_name": "abdelghanibelgaid",
      "author_url": "",
      "post_date": "12/12/2020 15:48:47",
      "content": "<p>Check out these sources:</p>\n<ul>\n<li><a href=\"https://lightgbm.readthedocs.io/en/latest/FAQ.html\" target=\"_blank\">General LightGBM Questions</a></li>\n<li><a href=\"https://towardsdatascience.com/how-to-learn-from-bigdata-files-on-low-memory-incremental-learning-d377282d38ff\" target=\"_blank\">How to handle BigData Files on Low Memory?</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1119020,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "12/19/2020 16:26:46",
          "content": "<p>Thank you! I will read them. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1108770": "I read several winner solutions using single LGBM in other competitions. They generated even more than 10K features. But in my code, memory cost easily reach 80G when training a model only with 40+ features. I could hardly imagine how large the memory will cost if I generate such amount of features. \n\nIs there a way to reduce the memory cost with large amount of features?",
    "1108960": "Try casting feature data types to np.float16 before lgb.Dataset.. From my experiment, with 9kw training samples, around 100 features setup, you need about 50GB for training LGB model.",
    "1108992": "Thanks, Rocky. I will try that.",
    "1109005": "Do you create your lgb.Dataset from a pandas dataframe or a numpy array?\nIn case you are using a pandas dataframe, try converting the dataframe to a numpy array of data type float32.\nI am currently using 50M training rows with 15 features and RAM usage is ~6GB\nA discussion on creating a numpy array from a dataframe in an efficient way can be found [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245)",
    "1109063": "Thank you! I will try to convert df to np.array.",
    "1110295": "Check out these sources:\n- [General LightGBM Questions](https://lightgbm.readthedocs.io/en/latest/FAQ.html)\n- [How to handle BigData Files on Low Memory?](https://towardsdatascience.com/how-to-learn-from-bigdata-files-on-low-memory-incremental-learning-d377282d38ff)",
    "1111703": "markwijkhuizen \nif you have some feature in pandas with dtype int8 ( or int16) and convert them to numpy float32 as you mentioned, wouldn't that increase memory usages ?",
    "1112148": "LighGBM converts all data to a numpy ndarray of type **float32** regardless the original data type, it will thus make a copy when using other data types. When creating the float32 ndarray yourself LightGBM will not make a copy but just use the given ndarray.\nFor pandas DataFrames all data is copied to a new DataFrame of type **float64**, this is why you should avoid DataFrames when creating LightBM Datasets. If you would use int8 features in a Dataframe a new DataFrame will be created where this feature is converted to float64.\n\n[Here](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325) is a discussion about LightGBM Datasets.",
    "1112180": "Aha, thats great explanation for the memory problems I have !\n\nThat means it will be usesless to reduce the size of my pandas columns (to int8 or int16) because lgbm will be convert it to np.float any way ...\nAnd even it will consume more memory because I would have two copies of same data (original DF and lgbm np.float), right ?",
    "1112516": "Yes, if you have your features in a DataFrame, convert the DataFrame to a float32 ndarray and delete the DataFrame.",
    "1112540": "As Mark says, it is really important to use np.float32 when constructing lgb.Dataset.\nBut I think loading features from disk is more memory-efficient than converting pandas dataframe to numpy array. This is my code to save memory.\n```\ndef load_arr(feature_names, data_type, n_rows = 99271300, path='../features/'):\n    X = np.zeros((n_rows, len(feature_names)), dtype = np.float32)\n    for i in tqdm(range(len(feature_names))):\n        X[:, i] = np.load(os.path.join(path, feature_names[i]) + '_' + data_type + '.npy')\n    return X\n```",
    "1119019": "Thanks for your kind share, @mamasinkgs . It is really an awesome way to reduce memory cost. I will try to use it.",
    "1119020": "Thank you! I will read them."
  },
  "source": "meta"
}