{
  "id": 55325,
  "title": "Reducing Lightgbm RAM spike using 1 difference",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55325",
  "author_name": "Yair Beer",
  "post_date": "2018-04-25T08:07:04.293000",
  "votes": 36,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi,\nAfter some digging I noticed that when using pandas DataFrame as a dataset it changes the type of the data to float which is a 64bit type and takes a lot of RAM because it has to save the new data into the memory.</p>\n\n<p>in lightgbm/basic.py. Line 258:</p>\n\n<pre><code>data = data.values.astype('float')\n</code></pre>\n\n<p>Using numpy array on the other hand uses np.float32 which is a 32bit type. \nAs in lightgbm/basic.py. Line 474:</p>\n\n<pre><code>data = np.array(mat.reshape(mat.size), dtype=np.float32)\n</code></pre>\n\n<p>Meaning reducing the memory spike by 50%!</p>\n\n<p>I think it is also possible to convert the data beforehand to np.float32 for no overhead at all because the data won't be copied. lightgbm/basic.py. Line 471/472:</p>\n\n<pre><code>if mat.dtype == np.float32 or mat.dtype == np.float64:\n        data = np.array(mat.reshape(mat.size), dtype=mat.dtype, copy=False)\n</code></pre>\n\n<p>If you used DataFrame as an input now you can use more data :).</p>",
  "messages": [
    {
      "id": 319095,
      "postDate": "2018-04-25T08:07:04.293Z",
      "content": "<p>Hi,\nAfter some digging I noticed that when using pandas DataFrame as a dataset it changes the type of the data to float which is a 64bit type and takes a lot of RAM because it has to save the new data into the memory.</p>\n\n<p>in lightgbm/basic.py. Line 258:</p>\n\n<pre><code>data = data.values.astype('float')\n</code></pre>\n\n<p>Using numpy array on the other hand uses np.float32 which is a 32bit type. \nAs in lightgbm/basic.py. Line 474:</p>\n\n<pre><code>data = np.array(mat.reshape(mat.size), dtype=np.float32)\n</code></pre>\n\n<p>Meaning reducing the memory spike by 50%!</p>\n\n<p>I think it is also possible to convert the data beforehand to np.float32 for no overhead at all because the data won't be copied. lightgbm/basic.py. Line 471/472:</p>\n\n<pre><code>if mat.dtype == np.float32 or mat.dtype == np.float64:\n        data = np.array(mat.reshape(mat.size), dtype=mat.dtype, copy=False)\n</code></pre>\n\n<p>If you used DataFrame as an input now you can use more data :).</p>",
      "rawMarkdown": "Hi,\nAfter some digging I noticed that when using pandas DataFrame as a dataset it changes the type of the data to float which is a 64bit type and takes a lot of RAM because it has to save the new data into the memory.\n\nin lightgbm/basic.py. Line 258:\n\n    data = data.values.astype('float')\n\nUsing numpy array on the other hand uses np.float32 which is a 32bit type. \nAs in lightgbm/basic.py. Line 474:\n\n    data = np.array(mat.reshape(mat.size), dtype=np.float32)\n\nMeaning reducing the memory spike by 50%!\n\nI think it is also possible to convert the data beforehand to np.float32 for no overhead at all because the data won't be copied. lightgbm/basic.py. Line 471/472:\n\n    if mat.dtype == np.float32 or mat.dtype == np.float64:\n            data = np.array(mat.reshape(mat.size), dtype=mat.dtype, copy=False)\n\nIf you used DataFrame as an input now you can use more data :).\n\n",
      "votes": 35
    },
    {
      "id": 321840,
      "postDate": "2018-05-02T02:11:07.200Z",
      "content": "<p>Be AWARE of 'uint32' to 'float32', especially for 'click_id', any integer number bigger than 16,777,216 in python will be rounded.</p>",
      "rawMarkdown": "Be AWARE of 'uint32' to 'float32', especially for 'click_id', any integer number bigger than 16,777,216 in python will be rounded.",
      "votes": 3
    },
    {
      "id": 319870,
      "postDate": "2018-04-27T02:12:10.087Z",
      "content": "<p>Nice. I made the change right in the lightgbm python binding <a href=\"https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/basic.py#L271\">https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/basic.py#L271</a> to <code>float32</code> since I've never used float64's with lightgbm before.</p>",
      "rawMarkdown": "Nice. I made the change right in the lightgbm python binding https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/basic.py#L271 to `float32` since I've never used float64's with lightgbm before.",
      "votes": 1,
      "replies": [
        {
          "id": 320046,
          "postDate": "2018-04-27T10:54:19.887Z",
          "content": "<p>Using \"float\" is basically np.float64.\nDo you mean you changed your local library file to np.float32?</p>\n\n<p>It is a good solution as well. </p>\n\n<p>I'm just afraid to forget to do it again when I update the library</p>",
          "rawMarkdown": "Using \"float\" is basically np.float64.\nDo you mean you changed your local library file to np.float32?\n\nIt is a good solution as well. \n\nI'm just afraid to forget to do it again when I update the library"
        },
        {
          "id": 320087,
          "postDate": "2018-04-27T12:49:54.763Z",
          "content": "<p>Yup, right in my local conda. I'm not afraid about forgetting because this competition has shredded the heck out of my SSD (swap space)... so both my computer as well as myself will remember to restore the original values once done.</p>",
          "rawMarkdown": "Yup, right in my local conda. I'm not afraid about forgetting because this competition has shredded the heck out of my SSD (swap space)... so both my computer as well as myself will remember to restore the original values once done.",
          "votes": 3
        },
        {
          "id": 320437,
          "postDate": "2018-04-28T17:00:50.667Z",
          "content": "<p>Hey guys. I just spent a day trying recovering from this since I've been lax in my source control the past week; but I believe I've been able to isolate an issue I started recently experiencing and would like to know if any of you have encountered similar.</p>\n\n<p>All the features I've attempted to add onto my highest scoring baseline the past day have underperformed to my great depression. So much so that I decided to run my highest scoring baseline again and noticed the auc dropped ~0.0003. Wtx? After much debugging, returning line 271 to 'float' from 'float32' restored the auc to what I had been expecting.</p>\n\n<p>Have any of you noticed this as well?</p>\n\n<p>My hypothesis is it might have something to do with floating point rounding of integer / categorical values . . . (?) Or alternatively, maybe it has to do something with accumulated errors in lgbm. After losing a day and 5 submissions I don't have time to re-attempt another run, but if any of you are currently on the float32, you might wanna try changing running it will full precision floats to see if you auc changes, or if I'm just tripping. Good luck!</p>",
          "rawMarkdown": "Hey guys. I just spent a day trying recovering from this since I've been lax in my source control the past week; but I believe I've been able to isolate an issue I started recently experiencing and would like to know if any of you have encountered similar.\n\nAll the features I've attempted to add onto my highest scoring baseline the past day have underperformed to my great depression. So much so that I decided to run my highest scoring baseline again and noticed the auc dropped ~0.0003. Wtx? After much debugging, returning line 271 to 'float' from 'float32' restored the auc to what I had been expecting.\n\nHave any of you noticed this as well?\n\nMy hypothesis is it might have something to do with floating point rounding of integer / categorical values . . . (?) Or alternatively, maybe it has to do something with accumulated errors in lgbm. After losing a day and 5 submissions I don't have time to re-attempt another run, but if any of you are currently on the float32, you might wanna try changing running it will full precision floats to see if you auc changes, or if I'm just tripping. Good luck!"
        },
        {
          "id": 320451,
          "postDate": "2018-04-28T18:04:52.750Z",
          "content": "<p>Seems weird because the predictions are always float64. Maybe you use very small numbers and the float32 round your numbers?</p>\n\n<pre><code>def __pred_for_np2d(self, mat, num_iteration, predict_type):\n    \"\"\"\n    Predict for a 2-D numpy matrix.\n    \"\"\"\n    if len(mat.shape) != 2:\n        raise ValueError('Input numpy.ndarray or list must be 2 dimensional')\n\n    if mat.dtype == np.float32 or mat.dtype == np.float64:\n        data = np.array(mat.reshape(mat.size), dtype=mat.dtype, copy=False)\n    else:\n        \"\"\"change non-float data to float data, need to copy\"\"\"\n        data = np.array(mat.reshape(mat.size), dtype=np.float32)\n    ptr_data, type_ptr_data = c_float_array(data)\n    n_preds = self.__get_num_preds(num_iteration, mat.shape[0],\n                                   predict_type)\n    preds = np.zeros(n_preds, dtype=np.float64)\n    out_num_preds = ctypes.c_int64(0)\n    _safe_call(_LIB.LGBM_BoosterPredictForMat(\n        self.handle,\n        ptr_data,\n        ctypes.c_int(type_ptr_data),\n        ctypes.c_int(mat.shape[0]),\n        ctypes.c_int(mat.shape[1]),\n        ctypes.c_int(C_API_IS_ROW_MAJOR),\n        ctypes.c_int(predict_type),\n        ctypes.c_int(num_iteration),\n        c_str(self.pred_parameter),\n        ctypes.byref(out_num_preds),\n        preds.ctypes.data_as(ctypes.POINTER(ctypes.c_double))))\n    if n_preds != out_num_preds.value:\n        raise ValueError(\"Wrong length for predict results\")\n    return preds, mat.shape[0]\n</code></pre>",
          "rawMarkdown": "Seems weird because the predictions are always float64. Maybe you use very small numbers and the float32 round your numbers?\n\n    def __pred_for_np2d(self, mat, num_iteration, predict_type):\n        \"\"\"\n        Predict for a 2-D numpy matrix.\n        \"\"\"\n        if len(mat.shape) != 2:\n            raise ValueError('Input numpy.ndarray or list must be 2 dimensional')\n    \n        if mat.dtype == np.float32 or mat.dtype == np.float64:\n            data = np.array(mat.reshape(mat.size), dtype=mat.dtype, copy=False)\n        else:\n            \"\"\"change non-float data to float data, need to copy\"\"\"\n            data = np.array(mat.reshape(mat.size), dtype=np.float32)\n        ptr_data, type_ptr_data = c_float_array(data)\n        n_preds = self.__get_num_preds(num_iteration, mat.shape[0],\n                                       predict_type)\n        preds = np.zeros(n_preds, dtype=np.float64)\n        out_num_preds = ctypes.c_int64(0)\n        _safe_call(_LIB.LGBM_BoosterPredictForMat(\n            self.handle,\n            ptr_data,\n            ctypes.c_int(type_ptr_data),\n            ctypes.c_int(mat.shape[0]),\n            ctypes.c_int(mat.shape[1]),\n            ctypes.c_int(C_API_IS_ROW_MAJOR),\n            ctypes.c_int(predict_type),\n            ctypes.c_int(num_iteration),\n            c_str(self.pred_parameter),\n            ctypes.byref(out_num_preds),\n            preds.ctypes.data_as(ctypes.POINTER(ctypes.c_double))))\n        if n_preds != out_num_preds.value:\n            raise ValueError(\"Wrong length for predict results\")\n        return preds, mat.shape[0]\n\n"
        }
      ]
    },
    {
      "id": 319189,
      "postDate": "2018-04-25T13:29:22.520Z",
      "content": "<p>Yes:</p>\n\n<pre><code>eval_data, hour_eval, y_eval = eval_data[predictors].values.astype(np.float32), eval_data[\"hour\"].values, eval_data[y_column].values\ntrain, hour_train, y_train = train[predictors].values.astype(np.float32), train[\"hour\"].values, train[y_column].values\n</code></pre>\n\n<p>This way the data won't be copied in addition to the df so it would take less total RAM.\nI think it should be done just before the lightgbm training because this way the data set is light in term of memory right to the end when you need it to the memory spikes of the groupby processing.</p>\n\n<p>I checked and if you send a np.float32 2D array there is no spike in the memory. I am using all the database with 5 variables and I am using less than 50% when trying with no memory spike when using the lgb dataset function :).</p>",
      "rawMarkdown": "Yes:\n\n    eval_data, hour_eval, y_eval = eval_data[predictors].values.astype(np.float32), eval_data[\"hour\"].values, eval_data[y_column].values\n    train, hour_train, y_train = train[predictors].values.astype(np.float32), train[\"hour\"].values, train[y_column].values\n\nThis way the data won't be copied in addition to the df so it would take less total RAM.\nI think it should be done just before the lightgbm training because this way the data set is light in term of memory right to the end when you need it to the memory spikes of the groupby processing.\n\nI checked and if you send a np.float32 2D array there is no spike in the memory. I am using all the database with 5 variables and I am using less than 50% when trying with no memory spike when using the lgb dataset function :).",
      "votes": 1,
      "replies": [
        {
          "id": 322565,
          "postDate": "2018-05-03T07:47:38.947Z",
          "content": "<p>Excuse me.When runing train = train[predictors].values.astype(np.float32),it comes out memoryerror.  How to deal with it?</p>",
          "rawMarkdown": "Excuse me.When runing train = train[predictors].values.astype(np.float32),it comes out memoryerror.  How to deal with it?"
        },
        {
          "id": 322571,
          "postDate": "2018-05-03T08:00:38.513Z",
          "content": "<p>It's because you are running out of your RAM(including swap), try to set more swap space or reduce the amount of entire data.</p>",
          "rawMarkdown": "It's because you are running out of your RAM(including swap), try to set more swap space or reduce the amount of entire data."
        },
        {
          "id": 322579,
          "postDate": "2018-05-03T08:23:02.730Z",
          "content": "<p>Hello Wenjie Bai.I cant reduce my data anymore.Is there any way to deal with this problem in python?</p>",
          "rawMarkdown": "Hello Wenjie Bai.I cant reduce my data anymore.Is there any way to deal with this problem in python?"
        },
        {
          "id": 322582,
          "postDate": "2018-05-03T08:31:31.060Z",
          "content": "<p>There are some ways to avoid memory spike but depending on what cases you are in, my QQ and Wechat are both 496852768, we can talk more in details.</p>",
          "rawMarkdown": "There are some ways to avoid memory spike but depending on what cases you are in, my QQ and Wechat are both 496852768, we can talk more in details."
        },
        {
          "id": 322673,
          "postDate": "2018-05-03T12:07:55.967Z",
          "content": "<p>Instead of converting you dfs yourself, add this additional parameter into your lgb_config: <a href=\"https://github.com/Microsoft/LightGBM/blob/84fef71528d4f84f38bd86f68150f9859276d721/examples/binary_classification/train.conf#L85\">https://github.com/Microsoft/LightGBM/blob/84fef71528d4f84f38bd86f68150f9859276d721/examples/binary_classification/train.conf#L85</a></p>",
          "rawMarkdown": "Instead of converting you dfs yourself, add this additional parameter into your lgb_config: https://github.com/Microsoft/LightGBM/blob/84fef71528d4f84f38bd86f68150f9859276d721/examples/binary_classification/train.conf#L85",
          "votes": 1
        }
      ]
    },
    {
      "id": 319866,
      "postDate": "2018-04-27T01:56:26.157Z",
      "content": "<p>I did try your trick, ram still spike, hmm </p>",
      "rawMarkdown": "I did try your trick, ram still spike, hmm ",
      "replies": [
        {
          "id": 320048,
          "postDate": "2018-04-27T10:55:14.067Z",
          "content": "<p>I tested it and it worked. maybe you save the values to a different variable and this way you have 2 copies of the dataset?</p>",
          "rawMarkdown": "I tested it and it worked. maybe you save the values to a different variable and this way you have 2 copies of the dataset?"
        },
        {
          "id": 320220,
          "postDate": "2018-04-27T22:26:39.910Z",
          "content": "<p>I did \n<code>dtrain = lgbm.Dataset(X_train[predictors].values.astype(np.float32), label = Y_train,\n                          feature_name = predictors, categorical_feature = categorical)\n    del X_train, Y_train; gc.collect()</code></p>",
          "rawMarkdown": "I did \n` dtrain = lgbm.Dataset(X_train[predictors].values.astype(np.float32), label = Y_train,\n                          feature_name = predictors, categorical_feature = categorical)\n    del X_train, Y_train; gc.collect()`"
        },
        {
          "id": 320229,
          "postDate": "2018-04-27T23:59:24.510Z",
          "content": "<p>I don't think it is enough. I might be wrong but intuitively you should do it before.\nIf you only send X_train[predictors].values.astype(np.float32) to the Dataset class, I'm sure the original DataFrame still exists. so you still store both the df and the numpy array you send in the memory.</p>\n\n<p>If you transform the df to numpy array beforehand. and you send the whole variable you would only send the pointer and not a copy (or transformed copy for that instance).</p>\n\n<p>I managed to work like that with 200M rows of 30 columns. Using 32GB.</p>",
          "rawMarkdown": "I don't think it is enough. I might be wrong but intuitively you should do it before.\nIf you only send X_train[predictors].values.astype(np.float32) to the Dataset class, I'm sure the original DataFrame still exists. so you still store both the df and the numpy array you send in the memory.\n\nIf you transform the df to numpy array beforehand. and you send the whole variable you would only send the pointer and not a copy (or transformed copy for that instance).\n\nI managed to work like that with 200M rows of 30 columns. Using 32GB.",
          "votes": 1
        },
        {
          "id": 320277,
          "postDate": "2018-04-28T05:37:05.223Z",
          "content": "<p>Thanks, now I don't see any huge ram spike with kernel</p>",
          "rawMarkdown": "Thanks, now I don't see any huge ram spike with kernel",
          "votes": 1
        }
      ]
    },
    {
      "id": 319175,
      "postDate": "2018-04-25T12:51:04.767Z",
      "content": "<p>Do you mean we should convert data from dataframe to np.array first?\nlabelv1 = train_df[target].values\ntrain_df = train_df[predictors].values\nxgtrain = lgb.Dataset(dtrain, label=labelv1,\n                          feature_name=predictors,\n                          categorical_feature=categorical_features\n                          )</p>",
      "rawMarkdown": "Do you mean we should convert data from dataframe to np.array first?\nlabelv1 = train_df[target].values\ntrain_df = train_df[predictors].values\nxgtrain = lgb.Dataset(dtrain, label=labelv1,\n                          feature_name=predictors,\n                          categorical_feature=categorical_features\n                          )\n",
      "replies": [
        {
          "id": 320205,
          "postDate": "2018-04-27T21:14:27.450Z",
          "content": "<p>you can simply do this when making dataset</p>\n\n<pre><code>lgb_train = lgb.Dataset(X_train[PREDICTORS].values.astype(np.float32),\\ label=y_train,feature_name=PREDICTORS, categorical_feature=CATEGORICAL)\n</code></pre>",
          "rawMarkdown": "you can simply do this when making dataset\n\n    lgb_train = lgb.Dataset(X_train[PREDICTORS].values.astype(np.float32),\\ label=y_train,feature_name=PREDICTORS, categorical_feature=CATEGORICAL)",
          "votes": 2
        },
        {
          "id": 320287,
          "postDate": "2018-04-28T06:44:55.240Z",
          "content": "<p>Great ! Thanks.</p>",
          "rawMarkdown": "Great ! Thanks."
        }
      ]
    },
    {
      "id": 319154,
      "postDate": "2018-04-25T11:31:15.657Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 321840,
      "author_name": "Wenjie Bai",
      "author_url": "",
      "post_date": "2018-05-02T02:11:07.200000",
      "content": "<p>Be AWARE of 'uint32' to 'float32', especially for 'click_id', any integer number bigger than 16,777,216 in python will be rounded.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 319870,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2018-04-27T02:12:10.087000",
      "content": "<p>Nice. I made the change right in the lightgbm python binding <a href=\"https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/basic.py#L271\">https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/basic.py#L271</a> to <code>float32</code> since I've never used float64's with lightgbm before.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 320046,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-04-27T10:54:19.887000",
          "content": "<p>Using \"float\" is basically np.float64.\nDo you mean you changed your local library file to np.float32?</p>\n\n<p>It is a good solution as well. </p>\n\n<p>I'm just afraid to forget to do it again when I update the library</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320087,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-27T12:49:54.763000",
          "content": "<p>Yup, right in my local conda. I'm not afraid about forgetting because this competition has shredded the heck out of my SSD (swap space)... so both my computer as well as myself will remember to restore the original values once done.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 320437,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-28T17:00:50.667000",
          "content": "<p>Hey guys. I just spent a day trying recovering from this since I've been lax in my source control the past week; but I believe I've been able to isolate an issue I started recently experiencing and would like to know if any of you have encountered similar.</p>\n\n<p>All the features I've attempted to add onto my highest scoring baseline the past day have underperformed to my great depression. So much so that I decided to run my highest scoring baseline again and noticed the auc dropped ~0.0003. Wtx? After much debugging, returning line 271 to 'float' from 'float32' restored the auc to what I had been expecting.</p>\n\n<p>Have any of you noticed this as well?</p>\n\n<p>My hypothesis is it might have something to do with floating point rounding of integer / categorical values . . . (?) Or alternatively, maybe it has to do something with accumulated errors in lgbm. After losing a day and 5 submissions I don't have time to re-attempt another run, but if any of you are currently on the float32, you might wanna try changing running it will full precision floats to see if you auc changes, or if I'm just tripping. Good luck!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320451,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-04-28T18:04:52.750000",
          "content": "<p>Seems weird because the predictions are always float64. Maybe you use very small numbers and the float32 round your numbers?</p>\n\n<pre><code>def __pred_for_np2d(self, mat, num_iteration, predict_type):\n    \"\"\"\n    Predict for a 2-D numpy matrix.\n    \"\"\"\n    if len(mat.shape) != 2:\n        raise ValueError('Input numpy.ndarray or list must be 2 dimensional')\n\n    if mat.dtype == np.float32 or mat.dtype == np.float64:\n        data = np.array(mat.reshape(mat.size), dtype=mat.dtype, copy=False)\n    else:\n        \"\"\"change non-float data to float data, need to copy\"\"\"\n        data = np.array(mat.reshape(mat.size), dtype=np.float32)\n    ptr_data, type_ptr_data = c_float_array(data)\n    n_preds = self.__get_num_preds(num_iteration, mat.shape[0],\n                                   predict_type)\n    preds = np.zeros(n_preds, dtype=np.float64)\n    out_num_preds = ctypes.c_int64(0)\n    _safe_call(_LIB.LGBM_BoosterPredictForMat(\n        self.handle,\n        ptr_data,\n        ctypes.c_int(type_ptr_data),\n        ctypes.c_int(mat.shape[0]),\n        ctypes.c_int(mat.shape[1]),\n        ctypes.c_int(C_API_IS_ROW_MAJOR),\n        ctypes.c_int(predict_type),\n        ctypes.c_int(num_iteration),\n        c_str(self.pred_parameter),\n        ctypes.byref(out_num_preds),\n        preds.ctypes.data_as(ctypes.POINTER(ctypes.c_double))))\n    if n_preds != out_num_preds.value:\n        raise ValueError(\"Wrong length for predict results\")\n    return preds, mat.shape[0]\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319189,
      "author_name": "Yair Beer",
      "author_url": "",
      "post_date": "2018-04-25T13:29:22.520000",
      "content": "<p>Yes:</p>\n\n<pre><code>eval_data, hour_eval, y_eval = eval_data[predictors].values.astype(np.float32), eval_data[\"hour\"].values, eval_data[y_column].values\ntrain, hour_train, y_train = train[predictors].values.astype(np.float32), train[\"hour\"].values, train[y_column].values\n</code></pre>\n\n<p>This way the data won't be copied in addition to the df so it would take less total RAM.\nI think it should be done just before the lightgbm training because this way the data set is light in term of memory right to the end when you need it to the memory spikes of the groupby processing.</p>\n\n<p>I checked and if you send a np.float32 2D array there is no spike in the memory. I am using all the database with 5 variables and I am using less than 50% when trying with no memory spike when using the lgb dataset function :).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 322565,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-05-03T07:47:38.947000",
          "content": "<p>Excuse me.When runing train = train[predictors].values.astype(np.float32),it comes out memoryerror.  How to deal with it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322571,
          "author_name": "Wenjie Bai",
          "author_url": "",
          "post_date": "2018-05-03T08:00:38.513000",
          "content": "<p>It's because you are running out of your RAM(including swap), try to set more swap space or reduce the amount of entire data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322579,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-05-03T08:23:02.730000",
          "content": "<p>Hello Wenjie Bai.I cant reduce my data anymore.Is there any way to deal with this problem in python?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322582,
          "author_name": "Wenjie Bai",
          "author_url": "",
          "post_date": "2018-05-03T08:31:31.060000",
          "content": "<p>There are some ways to avoid memory spike but depending on what cases you are in, my QQ and Wechat are both 496852768, we can talk more in details.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 322673,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-03T12:07:55.967000",
          "content": "<p>Instead of converting you dfs yourself, add this additional parameter into your lgb_config: <a href=\"https://github.com/Microsoft/LightGBM/blob/84fef71528d4f84f38bd86f68150f9859276d721/examples/binary_classification/train.conf#L85\">https://github.com/Microsoft/LightGBM/blob/84fef71528d4f84f38bd86f68150f9859276d721/examples/binary_classification/train.conf#L85</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 319866,
      "author_name": "MengYe",
      "author_url": "",
      "post_date": "2018-04-27T01:56:26.157000",
      "content": "<p>I did try your trick, ram still spike, hmm </p>",
      "votes": 0,
      "replies": [
        {
          "id": 320048,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-04-27T10:55:14.067000",
          "content": "<p>I tested it and it worked. maybe you save the values to a different variable and this way you have 2 copies of the dataset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320220,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-04-27T22:26:39.910000",
          "content": "<p>I did \n<code>dtrain = lgbm.Dataset(X_train[predictors].values.astype(np.float32), label = Y_train,\n                          feature_name = predictors, categorical_feature = categorical)\n    del X_train, Y_train; gc.collect()</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320229,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-04-27T23:59:24.510000",
          "content": "<p>I don't think it is enough. I might be wrong but intuitively you should do it before.\nIf you only send X_train[predictors].values.astype(np.float32) to the Dataset class, I'm sure the original DataFrame still exists. so you still store both the df and the numpy array you send in the memory.</p>\n\n<p>If you transform the df to numpy array beforehand. and you send the whole variable you would only send the pointer and not a copy (or transformed copy for that instance).</p>\n\n<p>I managed to work like that with 200M rows of 30 columns. Using 32GB.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 320277,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-04-28T05:37:05.223000",
          "content": "<p>Thanks, now I don't see any huge ram spike with kernel</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 319175,
      "author_name": "shark",
      "author_url": "",
      "post_date": "2018-04-25T12:51:04.767000",
      "content": "<p>Do you mean we should convert data from dataframe to np.array first?\nlabelv1 = train_df[target].values\ntrain_df = train_df[predictors].values\nxgtrain = lgb.Dataset(dtrain, label=labelv1,\n                          feature_name=predictors,\n                          categorical_feature=categorical_features\n                          )</p>",
      "votes": 0,
      "replies": [
        {
          "id": 320205,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-04-27T21:14:27.450000",
          "content": "<p>you can simply do this when making dataset</p>\n\n<pre><code>lgb_train = lgb.Dataset(X_train[PREDICTORS].values.astype(np.float32),\\ label=y_train,feature_name=PREDICTORS, categorical_feature=CATEGORICAL)\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 320287,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-28T06:44:55.240000",
          "content": "<p>Great ! Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319154,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-25T11:31:15.657000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "319095": "Hi,\nAfter some digging I noticed that when using pandas DataFrame as a dataset it changes the type of the data to float which is a 64bit type and takes a lot of RAM because it has to save the new data into the memory.\n\nin lightgbm/basic.py. Line 258:\n\n    data = data.values.astype('float')\n\nUsing numpy array on the other hand uses np.float32 which is a 32bit type. \nAs in lightgbm/basic.py. Line 474:\n\n    data = np.array(mat.reshape(mat.size), dtype=np.float32)\n\nMeaning reducing the memory spike by 50%!\n\nI think it is also possible to convert the data beforehand to np.float32 for no overhead at all because the data won't be copied. lightgbm/basic.py. Line 471/472:\n\n    if mat.dtype == np.float32 or mat.dtype == np.float64:\n            data = np.array(mat.reshape(mat.size), dtype=mat.dtype, copy=False)\n\nIf you used DataFrame as an input now you can use more data :).\n\n",
    "321840": "Be AWARE of 'uint32' to 'float32', especially for 'click_id', any integer number bigger than 16,777,216 in python will be rounded.",
    "319870": "Nice. I made the change right in the lightgbm python binding https://github.com/Microsoft/LightGBM/blob/master/python-package/lightgbm/basic.py#L271 to `float32` since I've never used float64's with lightgbm before.",
    "319189": "Yes:\n\n    eval_data, hour_eval, y_eval = eval_data[predictors].values.astype(np.float32), eval_data[\"hour\"].values, eval_data[y_column].values\n    train, hour_train, y_train = train[predictors].values.astype(np.float32), train[\"hour\"].values, train[y_column].values\n\nThis way the data won't be copied in addition to the df so it would take less total RAM.\nI think it should be done just before the lightgbm training because this way the data set is light in term of memory right to the end when you need it to the memory spikes of the groupby processing.\n\nI checked and if you send a np.float32 2D array there is no spike in the memory. I am using all the database with 5 variables and I am using less than 50% when trying with no memory spike when using the lgb dataset function :).",
    "319866": "I did try your trick, ram still spike, hmm ",
    "319175": "Do you mean we should convert data from dataframe to np.array first?\nlabelv1 = train_df[target].values\ntrain_df = train_df[predictors].values\nxgtrain = lgb.Dataset(dtrain, label=labelv1,\n                          feature_name=predictors,\n                          categorical_feature=categorical_features\n                          )\n",
    "319154": ""
  }
}