{
  "id": 53773,
  "title": "Lightgbm: prevent RAM spike (explode) at the init training",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53773",
  "author_name": "",
  "post_date": "2018-04-04T23:31:01.911862100Z",
  "votes": 66,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I am currently using the whole training data, plus some feature engineering, giving me 25 gb training data. This is my first time using Lightgbm, and I am pretty impressed about the result. However, the RAM exploded so many times, and I have to scale machine RAM from 60 gb to 100 gb to 150gb. </p>\n\n<p>Today, when I was trying cross validation, the RAM exploded one more time, and I started to monitor the RAM.  For my example with 25gb training data, the memory usually spikes at the beginning to almost 150gb ram, and it usually goes down to 40gb and becomes stable afterward. Again, this is my first time using this method, and I followed most kernels' method,  putting the pandas object in lightgbm.Dataset and train. </p>\n\n<p>I did some study and found a solution - saving lightgbm into binary. With this method, you will get </p>\n\n<p><em><strong>1. much faster loading time \n2. much smaller data (1-2gb when first reading in memory)</strong></em></p>\n\n<p>References:\n1. \nAs <strong><em>guolinke</em></strong> pointed out in <a href=\"https://github.com/Microsoft/LightGBM/issues/1032\">https://github.com/Microsoft/LightGBM/issues/1032</a>:</p>\n\n<p>\"@beutlgb the memory overhead is caused before Dataset constructed.\nWhen constructing Dataset in python package, it will convert the whole Dataset to float32 type first (on python side), then pass the converted float32 dataset to LightGBM api.\nIf you pass filename to LightGBM.Dataset directly, the file will be read by LightGBM api, without python. As a result, it will be more memory efficient.\"</p>\n\n<p>2.\n<a href=\"http://lightgbm.readthedocs.io/en/latest/Python-Intro.html\">http://lightgbm.readthedocs.io/en/latest/Python-Intro.html</a></p>\n\n<p>So I did the following step:</p>\n\n<p>a. Convert the data into lightgbm format</p>\n\n<p><code># format train_data_v1 = lightgbm.Dataset(train[predictors],label=train['is_attributed'],feature_name=predictors, categorical_feature=categorical)\n</code></p>\n\n<p>b.  Store the data in binary format</p>\n\n<p><code>\ntrain_data_v1.save_binary('train_v1.bin')\n</code></p>\n\n<p>c.  Load back</p>\n\n<p><code>\ntrain = lightgbm.Dataset('train_v1.bin', feature_name=predictors,                      categorical_feature=categorical)\n</code></p>\n\n<p><strong>Observations</strong>:</p>\n\n<p>When I do the step 2, I observed a huge memory spike just like the one I saw when training a panda object. Once I started using the binary lightgbm dataset, the training became much smoother. I assume that lightgbm.Dataset(pandas object) does not convert data type until training starts, since the lightgbm binary data does not spike at all when training.</p>\n\n<p>Hopefully this is helpful for some new learners like me, and thanks ahead for any additional memory saving tips!</p>",
  "messages": [
    {
      "id": "309273",
      "postDate": "04/04/2018 23:31:01",
      "content": "<p>I am currently using the whole training data, plus some feature engineering, giving me 25 gb training data. This is my first time using Lightgbm, and I am pretty impressed about the result. However, the RAM exploded so many times, and I have to scale machine RAM from 60 gb to 100 gb to 150gb. </p>\n\n<p>Today, when I was trying cross validation, the RAM exploded one more time, and I started to monitor the RAM.  For my example with 25gb training data, the memory usually spikes at the beginning to almost 150gb ram, and it usually goes down to 40gb and becomes stable afterward. Again, this is my first time using this method, and I followed most kernels' method,  putting the pandas object in lightgbm.Dataset and train. </p>\n\n<p>I did some study and found a solution - saving lightgbm into binary. With this method, you will get </p>\n\n<p><em><strong>1. much faster loading time \n2. much smaller data (1-2gb when first reading in memory)</strong></em></p>\n\n<p>References:\n1. \nAs <strong><em>guolinke</em></strong> pointed out in <a href=\"https://github.com/Microsoft/LightGBM/issues/1032\">https://github.com/Microsoft/LightGBM/issues/1032</a>:</p>\n\n<p>\"@beutlgb the memory overhead is caused before Dataset constructed.\nWhen constructing Dataset in python package, it will convert the whole Dataset to float32 type first (on python side), then pass the converted float32 dataset to LightGBM api.\nIf you pass filename to LightGBM.Dataset directly, the file will be read by LightGBM api, without python. As a result, it will be more memory efficient.\"</p>\n\n<p>2.\n<a href=\"http://lightgbm.readthedocs.io/en/latest/Python-Intro.html\">http://lightgbm.readthedocs.io/en/latest/Python-Intro.html</a></p>\n\n<p>So I did the following step:</p>\n\n<p>a. Convert the data into lightgbm format</p>\n\n<p><code># format train_data_v1 = lightgbm.Dataset(train[predictors],label=train['is_attributed'],feature_name=predictors, categorical_feature=categorical)\n</code></p>\n\n<p>b.  Store the data in binary format</p>\n\n<p><code>\ntrain_data_v1.save_binary('train_v1.bin')\n</code></p>\n\n<p>c.  Load back</p>\n\n<p><code>\ntrain = lightgbm.Dataset('train_v1.bin', feature_name=predictors,                      categorical_feature=categorical)\n</code></p>\n\n<p><strong>Observations</strong>:</p>\n\n<p>When I do the step 2, I observed a huge memory spike just like the one I saw when training a panda object. Once I started using the binary lightgbm dataset, the training became much smoother. I assume that lightgbm.Dataset(pandas object) does not convert data type until training starts, since the lightgbm binary data does not spike at all when training.</p>\n\n<p>Hopefully this is helpful for some new learners like me, and thanks ahead for any additional memory saving tips!</p>",
      "rawMarkdown": "I am currently using the whole training data, plus some feature engineering, giving me 25 gb training data. This is my first time using Lightgbm, and I am pretty impressed about the result. However, the RAM exploded so many times, and I have to scale machine RAM from 60 gb to 100 gb to 150gb. \n\nToday, when I was trying cross validation, the RAM exploded one more time, and I started to monitor the RAM.  For my example with 25gb training data, the memory usually spikes at the beginning to almost 150gb ram, and it usually goes down to 40gb and becomes stable afterward. Again, this is my first time using this method, and I followed most kernels' method,  putting the pandas object in lightgbm.Dataset and train. \n\nI did some study and found a solution - saving lightgbm into binary. With this method, you will get \n\n***1. much faster loading time \n2. much smaller data (1-2gb when first reading in memory)***\n\nReferences:\n1. \nAs ***guolinke*** pointed out in https://github.com/Microsoft/LightGBM/issues/1032:\n\n\"@beutlgb the memory overhead is caused before Dataset constructed.\nWhen constructing Dataset in python package, it will convert the whole Dataset to float32 type first (on python side), then pass the converted float32 dataset to LightGBM api.\nIf you pass filename to LightGBM.Dataset directly, the file will be read by LightGBM api, without python. As a result, it will be more memory efficient.\"\n\n2.\nhttp://lightgbm.readthedocs.io/en/latest/Python-Intro.html\n\n\nSo I did the following step:\n\na. Convert the data into lightgbm format\n\n``` # format train_data_v1 = lightgbm.Dataset(train[predictors],label=train['is_attributed'],feature_name=predictors, categorical_feature=categorical)\n```\n\n b.  Store the data in binary format\n\n``` \ntrain_data_v1.save_binary('train_v1.bin')\n```\n\nc.  Load back\n\n```\ntrain = lightgbm.Dataset('train_v1.bin', feature_name=predictors,                      categorical_feature=categorical)\n```\n\n**Observations**:\n\nWhen I do the step 2, I observed a huge memory spike just like the one I saw when training a panda object. Once I started using the binary lightgbm dataset, the training became much smoother. I assume that lightgbm.Dataset(pandas object) does not convert data type until training starts, since the lightgbm binary data does not spike at all when training.\n\nHopefully this is helpful for some new learners like me, and thanks ahead for any additional memory saving tips!",
      "votes": null
    },
    {
      "id": "309828",
      "postDate": "04/06/2018 03:08:45",
      "content": "<p>That's very helpful, thx dude!</p>",
      "rawMarkdown": "That's very helpful, thx dude!",
      "votes": null
    },
    {
      "id": "309830",
      "postDate": "04/06/2018 03:21:41",
      "content": "<p>no prob, man</p>",
      "rawMarkdown": "no prob, man",
      "votes": null
    },
    {
      "id": "310027",
      "postDate": "04/06/2018 12:29:43",
      "content": "<p>Hey,</p>\n\n<p>Did you try this field by any chance ?</p>\n\n<blockquote>\n  <p>two_round, default=false, type=bool, alias=two_round_loading,\n  use_two_round_loading by default, LightGBM will map data file to\n  memory and load features from memory. This will provide faster data\n  loading speed. But it may run out of memory when the data file is very\n  big set this to true if data file is too big to fit in memory</p>\n</blockquote>",
      "rawMarkdown": "Hey,\n\nDid you try this field by any chance ?\n\n&gt; two_round, default=false, type=bool, alias=two_round_loading,\n&gt; use_two_round_loading by default, LightGBM will map data file to\n&gt; memory and load features from memory. This will provide faster data\n&gt; loading speed. But it may run out of memory when the data file is very\n&gt; big set this to true if data file is too big to fit in memory",
      "votes": null
    },
    {
      "id": "310171",
      "postDate": "04/06/2018 18:07:57",
      "content": "<p>No, I only tried the binary so far. My 25 gb data only takes 2 seconds to load right now. </p>",
      "rawMarkdown": "No, I only tried the binary so far. My 25 gb data only takes 2 seconds to load right now.",
      "votes": null
    },
    {
      "id": "310515",
      "postDate": "04/07/2018 20:13:03",
      "content": "<p>Do you have ram explore issues while doing feature engineering? If so, how do you solve this problem?</p>",
      "rawMarkdown": "Do you have ram explore issues while doing feature engineering? If so, how do you solve this problem?",
      "votes": null
    },
    {
      "id": "310523",
      "postDate": "04/07/2018 20:45:57",
      "content": "<p>Yes, I did. There are basically two ways that I hand this problem:\n1. Do some steps of your FE then save data as pickle format, pd.save_pickle(), shut down the python, load back, and do more FE (This method loads data really fast, I did this because I noticed there were some ram leak during the data merge process)</p>\n\n<ol>\n<li>I am using google cloud with the free tail (Google gives you $300 credits). When face Ram issue, usually ,I just spur a bigger linux box. Generally, I think the ram explodes because of the lightgbm train init.</li>\n</ol>\n\n<p>Also a quick question to you. This is my first doing kaggle. I have about 40 features at this point (25gb) , how do you determine which feature to keep and which to throw. Do you have any suggestion on that?</p>",
      "rawMarkdown": "Yes, I did. There are basically two ways that I hand this problem:\n1. Do some steps of your FE then save data as pickle format, pd.save_pickle(), shut down the python, load back, and do more FE (This method loads data really fast, I did this because I noticed there were some ram leak during the data merge process)\n\n2. I am using google cloud with the free tail (Google gives you $300 credits). When face Ram issue, usually ,I just spur a bigger linux box. Generally, I think the ram explodes because of the lightgbm train init.\n \nAlso a quick question to you. This is my first doing kaggle. I have about 40 features at this point (25gb) , how do you determine which feature to keep and which to throw. Do you have any suggestion on that?",
      "votes": null
    },
    {
      "id": "310530",
      "postDate": "04/07/2018 21:11:26",
      "content": "<p>I use a small subset (about 5000000 rows) for validation to test each new feature and simply utilize local cv to determine it whether is bad or good instead of lb.  If a new feature is bad, it will lead to overfit under 100 iterations with my lgbm model. </p>",
      "rawMarkdown": "I use a small subset (about 5000000 rows) for validation to test each new feature and simply utilize local cv to determine it whether is bad or good instead of lb.  If a new feature is bad, it will lead to overfit under 100 iterations with my lgbm model.",
      "votes": null
    },
    {
      "id": "310533",
      "postDate": "04/07/2018 21:14:08",
      "content": "<p>Cool, thanks a lot!</p>",
      "rawMarkdown": "Cool, thanks a lot!",
      "votes": null
    },
    {
      "id": "310723",
      "postDate": "04/08/2018 11:49:52",
      "content": "<p>That's gonna save me so much time, thanks a lot !\nYou were asking a question about which features to keep in the ones that you have created, for this there are different solutions ( and interpretations).\nFirst off, you can try to plot the feature importances at the end of the training of your lightgbm as you can see in <a href=\"https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code\">this kernel</a>. This way you can see the features that were the most relevant and select them accordingly.</p>\n\n<p>I have experimented also on running a chi2 test on the dataset, it gives you a metric of the \"quality\" ( pardon the shortcut ) of a certain feature in regards to its target ( available through sklearn ).</p>\n\n<p>Take a look a <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">this <strong>awesome</strong> notbook</a> there's a lot of very interesting tests there.</p>\n\n<p>Lastly you can look into the Information Value of each feature also in regards to the target and pick which one is relevant or not. This was done in a R notebook that I can't find anymore but I'll link it as soon as I find it.</p>\n\n<p>Good luck !</p>",
      "rawMarkdown": "That's gonna save me so much time, thanks a lot !\nYou were asking a question about which features to keep in the ones that you have created, for this there are different solutions ( and interpretations).\nFirst off, you can try to plot the feature importances at the end of the training of your lightgbm as you can see in [this kernel][1]. This way you can see the features that were the most relevant and select them accordingly.\n\nI have experimented also on running a chi2 test on the dataset, it gives you a metric of the \"quality\" ( pardon the shortcut ) of a certain feature in regards to its target ( available through sklearn ).\n\nTake a look a [this **awesome** notbook][2] there's a lot of very interesting tests there.\n\nLastly you can look into the Information Value of each feature also in regards to the target and pick which one is relevant or not. This was done in a R notebook that I can't find anymore but I'll link it as soon as I find it.\n\n\n\nGood luck !\n\n\n  [1]: https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code\n  [2]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing",
      "votes": null
    },
    {
      "id": "310787",
      "postDate": "04/08/2018 15:40:41",
      "content": "<p>cool man, thanks a lot for all useful advices! I have been stuck on this feature engineering for too long lol</p>",
      "rawMarkdown": "cool man, thanks a lot for all useful advices! I have been stuck on this feature engineering for too long lol",
      "votes": null
    },
    {
      "id": "321029",
      "postDate": "04/30/2018 12:40:02",
      "content": "<p>Thank you so much for this! I tried your 2nd method and got this error: \"lightgbm.basic.LightGBMError: b'std::bad_alloc' when trying to save the binary. Do you have any suggestions for that? Thanks!</p>",
      "rawMarkdown": "Thank you so much for this! I tried your 2nd method and got this error: \"lightgbm.basic.LightGBMError: b'std::bad_alloc' when trying to save the binary. Do you have any suggestions for that? Thanks!",
      "votes": null
    },
    {
      "id": "321197",
      "postDate": "04/30/2018 19:58:26",
      "content": "<p>I had that one before. I think it was either 1. memory program or 2. columns that contain missing value (if you have FE that calculate variance or next click) </p>",
      "rawMarkdown": "I had that one before. I think it was either 1. memory program or 2. columns that contain missing value (if you have FE that calculate variance or next click)",
      "votes": null
    },
    {
      "id": "321204",
      "postDate": "04/30/2018 20:22:48",
      "content": "<p>Thanks for your comment! Ah I see - I believe it's the 2nd case you mentioned. Were you able to find a solution for it?</p>",
      "rawMarkdown": "Thanks for your comment! Ah I see - I believe it's the 2nd case you mentioned. Were you able to find a solution for it?",
      "votes": null
    },
    {
      "id": "321205",
      "postDate": "04/30/2018 20:24:28",
      "content": "<p>I simply 1. drop the NaNs in my local machine or 2. use a bigger cloud machine (google cloud, aws, any kind) to run the ode</p>",
      "rawMarkdown": "I simply 1. drop the NaNs in my local machine or 2. use a bigger cloud machine (google cloud, aws, any kind) to run the ode",
      "votes": null
    },
    {
      "id": "321304",
      "postDate": "05/01/2018 01:19:23",
      "content": "<p>Can I use this approach in kaggle scripts/kernels? How can I save and load the binary files in kaggle server? Thanks!</p>",
      "rawMarkdown": "Can I use this approach in kaggle scripts/kernels? How can I save and load the binary files in kaggle server? Thanks!",
      "votes": null
    },
    {
      "id": "321334",
      "postDate": "05/01/2018 04:05:43",
      "content": "<p>Thanks a lot for your replies Dylan! For some reason, now I'm getting a different error: \"Length of feature_name(20) and num_feature(6) don't match\". </p>\n\n<p>xgtrain.save_binary('xgtrain_valid.bin')\nxgtest.save_binary('xgtest_valid.bin')</p>\n\n<p>xgtrain = lgb.Dataset('xgtrain_valid.bin', feature_name=predictors, categorical_feature=categorical)\nxgtest = lgb.Dataset('xgtest_valid.bin', feature_name=predictors, categorical_feature=categorical)</p>\n\n<p>The above are the only 4 lines of code I've added to try to load from binary but, for some reason, it's interpreting the data to only have 6 features where there were 20 columns when I saved it as binary? Am I missing something?</p>",
      "rawMarkdown": "Thanks a lot for your replies Dylan! For some reason, now I'm getting a different error: \"Length of feature_name(20) and num_feature(6) don't match\". \n\nxgtrain.save_binary('xgtrain_valid.bin')\nxgtest.save_binary('xgtest_valid.bin')\n\nxgtrain = lgb.Dataset('xgtrain_valid.bin', feature_name=predictors, categorical_feature=categorical)\nxgtest = lgb.Dataset('xgtest_valid.bin', feature_name=predictors, categorical_feature=categorical)\n\nThe above are the only 4 lines of code I've added to try to load from binary but, for some reason, it's interpreting the data to only have 6 features where there were 20 columns when I saved it as binary? Am I missing something?",
      "votes": null
    },
    {
      "id": "321382",
      "postDate": "05/01/2018 06:15:02",
      "content": "<p>I had that before as well. I think that you can only train on the variables that you save as binary. Let me know if you figure a way to by-pass this</p>",
      "rawMarkdown": "I had that before as well. I think that you can only train on the variables that you save as binary. Let me know if you figure a way to by-pass this",
      "votes": null
    },
    {
      "id": "321383",
      "postDate": "05/01/2018 06:17:06",
      "content": "<p>I only use this on my local computer and cloud server level. I don't know how kaggle server is configured, if a new kernel starts with a new virtual server every time, I think this probably not gonna work</p>",
      "rawMarkdown": "I only use this on my local computer and cloud server level. I don't know how kaggle server is configured, if a new kernel starts with a new virtual server every time, I think this probably not gonna work",
      "votes": null
    },
    {
      "id": "324983",
      "postDate": "05/08/2018 02:38:49",
      "content": "<p>Hey man.I can't understand your steps.In my opinion, what <em>guolinke</em> suggests is that don't use the python package to construct Dataset.But in your steps, aren't you still using the python package?And is that why you observed a huge memory spike on step 2?</p>",
      "rawMarkdown": "Hey man.I can't understand your steps.In my opinion, what *guolinke* suggests is that don't use the python package to construct Dataset.But in your steps, aren't you still using the python package?And is that why you observed a huge memory spike on step 2?",
      "votes": null
    },
    {
      "id": "324986",
      "postDate": "05/08/2018 02:41:31",
      "content": "<p>Could you send that person’s post? This is my first time using this package, and I was using python to construct the dataframe.</p>",
      "rawMarkdown": "Could you send that person’s post? This is my first time using this package, and I was using python to construct the dataframe.",
      "votes": null
    },
    {
      "id": "324988",
      "postDate": "05/08/2018 02:45:34",
      "content": "<p>That's in your reference 1  :D</p>\n\n<p>References: 1. As guolinke pointed out in <a href=\"https://github.com/Microsoft/LightGBM/issues/1032\">https://github.com/Microsoft/LightGBM/issues/1032</a></p>",
      "rawMarkdown": "That's in your reference 1  :D\n\nReferences: 1. As guolinke pointed out in https://github.com/Microsoft/LightGBM/issues/1032",
      "votes": null
    },
    {
      "id": "324993",
      "postDate": "05/08/2018 02:51:41",
      "content": "<p>Oh, this one. I did not try that method. I basically pre-creates the dataframe by saving it into binary format since lightgbm automatically creates the dataframe when saving it binary format. </p>",
      "rawMarkdown": "Oh, this one. I did not try that method. I basically pre-creates the dataframe by saving it into binary format since lightgbm automatically creates the dataframe when saving it binary format.",
      "votes": null
    },
    {
      "id": "324996",
      "postDate": "05/08/2018 02:59:08",
      "content": "<p>So you can prevent the RAM spike at init training but the memory may still explode when constructing Dataset?</p>",
      "rawMarkdown": "So you can prevent the RAM spike at init training but the memory may still explode when constructing Dataset?",
      "votes": null
    },
    {
      "id": "325008",
      "postDate": "05/08/2018 03:18:49",
      "content": "<p>That’s correct, there is a chance for that. Maybe the best way is from the link that you just sent </p>",
      "rawMarkdown": "That’s correct, there is a chance for that. Maybe the best way is from the link that you just sent",
      "votes": null
    },
    {
      "id": "325016",
      "postDate": "05/08/2018 03:22:13",
      "content": "<p>OK Thanks man</p>",
      "rawMarkdown": "OK Thanks man",
      "votes": null
    },
    {
      "id": "487766",
      "postDate": "03/11/2019 13:02:30",
      "content": "<p>Hi Dylan,</p>\n\n<p>Thanks a lot for the tips. I am new to LightGBM and was wondering why it kept crashing when the training (5k tfidf of 670k data) begins.</p>\n\n<p>Converting the data to binary and loading them back helps!</p>",
      "rawMarkdown": "Hi Dylan,\n\nThanks a lot for the tips. I am new to LightGBM and was wondering why it kept crashing when the training (5k tfidf of 670k data) begins.\n\nConverting the data to binary and loading them back helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 309828,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "04/06/2018 03:08:45",
      "content": "<p>That's very helpful, thx dude!</p>",
      "votes": null,
      "replies": [
        {
          "id": 309830,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "04/06/2018 03:21:41",
          "content": "<p>no prob, man</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310027,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "04/06/2018 12:29:43",
      "content": "<p>Hey,</p>\n\n<p>Did you try this field by any chance ?</p>\n\n<blockquote>\n  <p>two_round, default=false, type=bool, alias=two_round_loading,\n  use_two_round_loading by default, LightGBM will map data file to\n  memory and load features from memory. This will provide faster data\n  loading speed. But it may run out of memory when the data file is very\n  big set this to true if data file is too big to fit in memory</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 310171,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "04/06/2018 18:07:57",
          "content": "<p>No, I only tried the binary so far. My 25 gb data only takes 2 seconds to load right now. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310515,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "04/07/2018 20:13:03",
      "content": "<p>Do you have ram explore issues while doing feature engineering? If so, how do you solve this problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 310523,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "04/07/2018 20:45:57",
          "content": "<p>Yes, I did. There are basically two ways that I hand this problem:\n1. Do some steps of your FE then save data as pickle format, pd.save_pickle(), shut down the python, load back, and do more FE (This method loads data really fast, I did this because I noticed there were some ram leak during the data merge process)</p>\n\n<ol>\n<li>I am using google cloud with the free tail (Google gives you $300 credits). When face Ram issue, usually ,I just spur a bigger linux box. Generally, I think the ram explodes because of the lightgbm train init.</li>\n</ol>\n\n<p>Also a quick question to you. This is my first doing kaggle. I have about 40 features at this point (25gb) , how do you determine which feature to keep and which to throw. Do you have any suggestion on that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 310530,
          "author_name": "konohayui",
          "author_url": "",
          "post_date": "04/07/2018 21:11:26",
          "content": "<p>I use a small subset (about 5000000 rows) for validation to test each new feature and simply utilize local cv to determine it whether is bad or good instead of lb.  If a new feature is bad, it will lead to overfit under 100 iterations with my lgbm model. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 310533,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "04/07/2018 21:14:08",
          "content": "<p>Cool, thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310723,
      "author_name": "propanon",
      "author_url": "",
      "post_date": "04/08/2018 11:49:52",
      "content": "<p>That's gonna save me so much time, thanks a lot !\nYou were asking a question about which features to keep in the ones that you have created, for this there are different solutions ( and interpretations).\nFirst off, you can try to plot the feature importances at the end of the training of your lightgbm as you can see in <a href=\"https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code\">this kernel</a>. This way you can see the features that were the most relevant and select them accordingly.</p>\n\n<p>I have experimented also on running a chi2 test on the dataset, it gives you a metric of the \"quality\" ( pardon the shortcut ) of a certain feature in regards to its target ( available through sklearn ).</p>\n\n<p>Take a look a <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">this <strong>awesome</strong> notbook</a> there's a lot of very interesting tests there.</p>\n\n<p>Lastly you can look into the Information Value of each feature also in regards to the target and pick which one is relevant or not. This was done in a R notebook that I can't find anymore but I'll link it as soon as I find it.</p>\n\n<p>Good luck !</p>",
      "votes": null,
      "replies": [
        {
          "id": 310787,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "04/08/2018 15:40:41",
          "content": "<p>cool man, thanks a lot for all useful advices! I have been stuck on this feature engineering for too long lol</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321029,
      "author_name": "edchen8",
      "author_url": "",
      "post_date": "04/30/2018 12:40:02",
      "content": "<p>Thank you so much for this! I tried your 2nd method and got this error: \"lightgbm.basic.LightGBMError: b'std::bad_alloc' when trying to save the binary. Do you have any suggestions for that? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 321197,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "04/30/2018 19:58:26",
          "content": "<p>I had that one before. I think it was either 1. memory program or 2. columns that contain missing value (if you have FE that calculate variance or next click) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321204,
          "author_name": "edchen8",
          "author_url": "",
          "post_date": "04/30/2018 20:22:48",
          "content": "<p>Thanks for your comment! Ah I see - I believe it's the 2nd case you mentioned. Were you able to find a solution for it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321205,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "04/30/2018 20:24:28",
          "content": "<p>I simply 1. drop the NaNs in my local machine or 2. use a bigger cloud machine (google cloud, aws, any kind) to run the ode</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321334,
          "author_name": "edchen8",
          "author_url": "",
          "post_date": "05/01/2018 04:05:43",
          "content": "<p>Thanks a lot for your replies Dylan! For some reason, now I'm getting a different error: \"Length of feature_name(20) and num_feature(6) don't match\". </p>\n\n<p>xgtrain.save_binary('xgtrain_valid.bin')\nxgtest.save_binary('xgtest_valid.bin')</p>\n\n<p>xgtrain = lgb.Dataset('xgtrain_valid.bin', feature_name=predictors, categorical_feature=categorical)\nxgtest = lgb.Dataset('xgtest_valid.bin', feature_name=predictors, categorical_feature=categorical)</p>\n\n<p>The above are the only 4 lines of code I've added to try to load from binary but, for some reason, it's interpreting the data to only have 6 features where there were 20 columns when I saved it as binary? Am I missing something?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321382,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "05/01/2018 06:15:02",
          "content": "<p>I had that before as well. I think that you can only train on the variables that you save as binary. Let me know if you figure a way to by-pass this</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321304,
      "author_name": "jsaguiar",
      "author_url": "",
      "post_date": "05/01/2018 01:19:23",
      "content": "<p>Can I use this approach in kaggle scripts/kernels? How can I save and load the binary files in kaggle server? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 321383,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "05/01/2018 06:17:06",
          "content": "<p>I only use this on my local computer and cloud server level. I don't know how kaggle server is configured, if a new kernel starts with a new virtual server every time, I think this probably not gonna work</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324983,
      "author_name": "xiguapi",
      "author_url": "",
      "post_date": "05/08/2018 02:38:49",
      "content": "<p>Hey man.I can't understand your steps.In my opinion, what <em>guolinke</em> suggests is that don't use the python package to construct Dataset.But in your steps, aren't you still using the python package?And is that why you observed a huge memory spike on step 2?</p>",
      "votes": null,
      "replies": [
        {
          "id": 324986,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "05/08/2018 02:41:31",
          "content": "<p>Could you send that person’s post? This is my first time using this package, and I was using python to construct the dataframe.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324988,
          "author_name": "xiguapi",
          "author_url": "",
          "post_date": "05/08/2018 02:45:34",
          "content": "<p>That's in your reference 1  :D</p>\n\n<p>References: 1. As guolinke pointed out in <a href=\"https://github.com/Microsoft/LightGBM/issues/1032\">https://github.com/Microsoft/LightGBM/issues/1032</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324993,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "05/08/2018 02:51:41",
          "content": "<p>Oh, this one. I did not try that method. I basically pre-creates the dataframe by saving it into binary format since lightgbm automatically creates the dataframe when saving it binary format. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324996,
          "author_name": "xiguapi",
          "author_url": "",
          "post_date": "05/08/2018 02:59:08",
          "content": "<p>So you can prevent the RAM spike at init training but the memory may still explode when constructing Dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325008,
          "author_name": "wuyile516516",
          "author_url": "",
          "post_date": "05/08/2018 03:18:49",
          "content": "<p>That’s correct, there is a chance for that. Maybe the best way is from the link that you just sent </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325016,
          "author_name": "xiguapi",
          "author_url": "",
          "post_date": "05/08/2018 03:22:13",
          "content": "<p>OK Thanks man</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 487766,
      "author_name": "chewzy",
      "author_url": "",
      "post_date": "03/11/2019 13:02:30",
      "content": "<p>Hi Dylan,</p>\n\n<p>Thanks a lot for the tips. I am new to LightGBM and was wondering why it kept crashing when the training (5k tfidf of 670k data) begins.</p>\n\n<p>Converting the data to binary and loading them back helps!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "309273": "I am currently using the whole training data, plus some feature engineering, giving me 25 gb training data. This is my first time using Lightgbm, and I am pretty impressed about the result. However, the RAM exploded so many times, and I have to scale machine RAM from 60 gb to 100 gb to 150gb. \n\nToday, when I was trying cross validation, the RAM exploded one more time, and I started to monitor the RAM.  For my example with 25gb training data, the memory usually spikes at the beginning to almost 150gb ram, and it usually goes down to 40gb and becomes stable afterward. Again, this is my first time using this method, and I followed most kernels' method,  putting the pandas object in lightgbm.Dataset and train. \n\nI did some study and found a solution - saving lightgbm into binary. With this method, you will get \n\n***1. much faster loading time \n2. much smaller data (1-2gb when first reading in memory)***\n\nReferences:\n1. \nAs ***guolinke*** pointed out in https://github.com/Microsoft/LightGBM/issues/1032:\n\n\"@beutlgb the memory overhead is caused before Dataset constructed.\nWhen constructing Dataset in python package, it will convert the whole Dataset to float32 type first (on python side), then pass the converted float32 dataset to LightGBM api.\nIf you pass filename to LightGBM.Dataset directly, the file will be read by LightGBM api, without python. As a result, it will be more memory efficient.\"\n\n2.\nhttp://lightgbm.readthedocs.io/en/latest/Python-Intro.html\n\n\nSo I did the following step:\n\na. Convert the data into lightgbm format\n\n``` # format train_data_v1 = lightgbm.Dataset(train[predictors],label=train['is_attributed'],feature_name=predictors, categorical_feature=categorical)\n```\n\n b.  Store the data in binary format\n\n``` \ntrain_data_v1.save_binary('train_v1.bin')\n```\n\nc.  Load back\n\n```\ntrain = lightgbm.Dataset('train_v1.bin', feature_name=predictors,                      categorical_feature=categorical)\n```\n\n**Observations**:\n\nWhen I do the step 2, I observed a huge memory spike just like the one I saw when training a panda object. Once I started using the binary lightgbm dataset, the training became much smoother. I assume that lightgbm.Dataset(pandas object) does not convert data type until training starts, since the lightgbm binary data does not spike at all when training.\n\nHopefully this is helpful for some new learners like me, and thanks ahead for any additional memory saving tips!",
    "309828": "That's very helpful, thx dude!",
    "309830": "no prob, man",
    "310027": "Hey,\n\nDid you try this field by any chance ?\n\n&gt; two_round, default=false, type=bool, alias=two_round_loading,\n&gt; use_two_round_loading by default, LightGBM will map data file to\n&gt; memory and load features from memory. This will provide faster data\n&gt; loading speed. But it may run out of memory when the data file is very\n&gt; big set this to true if data file is too big to fit in memory",
    "310171": "No, I only tried the binary so far. My 25 gb data only takes 2 seconds to load right now.",
    "310515": "Do you have ram explore issues while doing feature engineering? If so, how do you solve this problem?",
    "310523": "Yes, I did. There are basically two ways that I hand this problem:\n1. Do some steps of your FE then save data as pickle format, pd.save_pickle(), shut down the python, load back, and do more FE (This method loads data really fast, I did this because I noticed there were some ram leak during the data merge process)\n\n2. I am using google cloud with the free tail (Google gives you $300 credits). When face Ram issue, usually ,I just spur a bigger linux box. Generally, I think the ram explodes because of the lightgbm train init.\n \nAlso a quick question to you. This is my first doing kaggle. I have about 40 features at this point (25gb) , how do you determine which feature to keep and which to throw. Do you have any suggestion on that?",
    "310530": "I use a small subset (about 5000000 rows) for validation to test each new feature and simply utilize local cv to determine it whether is bad or good instead of lb.  If a new feature is bad, it will lead to overfit under 100 iterations with my lgbm model.",
    "310533": "Cool, thanks a lot!",
    "310723": "That's gonna save me so much time, thanks a lot !\nYou were asking a question about which features to keep in the ones that you have created, for this there are different solutions ( and interpretations).\nFirst off, you can try to plot the feature importances at the end of the training of your lightgbm as you can see in [this kernel][1]. This way you can see the features that were the most relevant and select them accordingly.\n\nI have experimented also on running a chi2 test on the dataset, it gives you a metric of the \"quality\" ( pardon the shortcut ) of a certain feature in regards to its target ( available through sklearn ).\n\nTake a look a [this **awesome** notbook][2] there's a lot of very interesting tests there.\n\nLastly you can look into the Information Value of each feature also in regards to the target and pick which one is relevant or not. This was done in a R notebook that I can't find anymore but I'll link it as soon as I find it.\n\n\n\nGood luck !\n\n\n  [1]: https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code\n  [2]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing",
    "310787": "cool man, thanks a lot for all useful advices! I have been stuck on this feature engineering for too long lol",
    "321029": "Thank you so much for this! I tried your 2nd method and got this error: \"lightgbm.basic.LightGBMError: b'std::bad_alloc' when trying to save the binary. Do you have any suggestions for that? Thanks!",
    "321197": "I had that one before. I think it was either 1. memory program or 2. columns that contain missing value (if you have FE that calculate variance or next click)",
    "321204": "Thanks for your comment! Ah I see - I believe it's the 2nd case you mentioned. Were you able to find a solution for it?",
    "321205": "I simply 1. drop the NaNs in my local machine or 2. use a bigger cloud machine (google cloud, aws, any kind) to run the ode",
    "321304": "Can I use this approach in kaggle scripts/kernels? How can I save and load the binary files in kaggle server? Thanks!",
    "321334": "Thanks a lot for your replies Dylan! For some reason, now I'm getting a different error: \"Length of feature_name(20) and num_feature(6) don't match\". \n\nxgtrain.save_binary('xgtrain_valid.bin')\nxgtest.save_binary('xgtest_valid.bin')\n\nxgtrain = lgb.Dataset('xgtrain_valid.bin', feature_name=predictors, categorical_feature=categorical)\nxgtest = lgb.Dataset('xgtest_valid.bin', feature_name=predictors, categorical_feature=categorical)\n\nThe above are the only 4 lines of code I've added to try to load from binary but, for some reason, it's interpreting the data to only have 6 features where there were 20 columns when I saved it as binary? Am I missing something?",
    "321382": "I had that before as well. I think that you can only train on the variables that you save as binary. Let me know if you figure a way to by-pass this",
    "321383": "I only use this on my local computer and cloud server level. I don't know how kaggle server is configured, if a new kernel starts with a new virtual server every time, I think this probably not gonna work",
    "324983": "Hey man.I can't understand your steps.In my opinion, what *guolinke* suggests is that don't use the python package to construct Dataset.But in your steps, aren't you still using the python package?And is that why you observed a huge memory spike on step 2?",
    "324986": "Could you send that person’s post? This is my first time using this package, and I was using python to construct the dataframe.",
    "324988": "That's in your reference 1  :D\n\nReferences: 1. As guolinke pointed out in https://github.com/Microsoft/LightGBM/issues/1032",
    "324993": "Oh, this one. I did not try that method. I basically pre-creates the dataframe by saving it into binary format since lightgbm automatically creates the dataframe when saving it binary format.",
    "324996": "So you can prevent the RAM spike at init training but the memory may still explode when constructing Dataset?",
    "325008": "That’s correct, there is a chance for that. Maybe the best way is from the link that you just sent",
    "325016": "OK Thanks man",
    "487766": "Hi Dylan,\n\nThanks a lot for the tips. I am new to LightGBM and was wondering why it kept crashing when the training (5k tfidf of 670k data) begins.\n\nConverting the data to binary and loading them back helps!"
  },
  "source": "meta"
}