{
  "id": 198245,
  "title": "Reduce memory spike while creating lightgbm dataset",
  "url": "/competitions/riiid-test-answer-prediction/discussion/198245",
  "author_name": "Anurag Trivedi",
  "post_date": "2020-11-20T11:18:10.502000",
  "votes": 38,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Don't use DF to create lightgbm dataset, rather use np array:</p>\n<p>Sample code:</p>\n<pre><code>X_train_np = train.values.astype(np.float32)\nX_valid_np = valid.values.astype(np.float32)\nfeatures = train.columns\nlgb_train = lgb.Dataset(X_train_np, label=y_tr, feature_name=list(features))\nlgb_valid = lgb.Dataset(X_valid_np, label=y_va, feature_name=list(features))\ndel train, y_tr\n_=gc.collect()\n</code></pre>\n<p>Reference - <a href=\"https://www.kaggle.com/c/m5-forecasting-accuracy/discussion/149754\" target=\"_blank\">https://www.kaggle.com/c/m5-forecasting-accuracy/discussion/149754</a><br>\ncredit: <a href=\"https://www.kaggle.com/alankabisov\" target=\"_blank\">@alankabisov</a> </p>",
  "messages": [
    {
      "id": 1084772,
      "postDate": "2020-11-20T11:18:10.503Z",
      "content": "<p>Don't use DF to create lightgbm dataset, rather use np array:</p>\n<p>Sample code:</p>\n<pre><code>X_train_np = train.values.astype(np.float32)\nX_valid_np = valid.values.astype(np.float32)\nfeatures = train.columns\nlgb_train = lgb.Dataset(X_train_np, label=y_tr, feature_name=list(features))\nlgb_valid = lgb.Dataset(X_valid_np, label=y_va, feature_name=list(features))\ndel train, y_tr\n_=gc.collect()\n</code></pre>\n<p>Reference - <a href=\"https://www.kaggle.com/c/m5-forecasting-accuracy/discussion/149754\" target=\"_blank\">https://www.kaggle.com/c/m5-forecasting-accuracy/discussion/149754</a><br>\ncredit: <a href=\"https://www.kaggle.com/alankabisov\" target=\"_blank\">@alankabisov</a> </p>",
      "rawMarkdown": "Don't use DF to create lightgbm dataset, rather use np array:\n\nSample code:\n```\nX_train_np = train.values.astype(np.float32)\nX_valid_np = valid.values.astype(np.float32)\nfeatures = train.columns\nlgb_train = lgb.Dataset(X_train_np, label=y_tr, feature_name=list(features))\nlgb_valid = lgb.Dataset(X_valid_np, label=y_va, feature_name=list(features))\ndel train, y_tr\n_=gc.collect()\n```\n\nReference - https://www.kaggle.com/c/m5-forecasting-accuracy/discussion/149754\ncredit: @alankabisov ",
      "votes": 37
    },
    {
      "id": 1084857,
      "postDate": "2020-11-20T13:13:02.557Z",
      "content": "<p>Good tip!<br>\nYou can however get a memory spike when creating the numpy array from the dataframe. My solution was to convert the columns one by one.</p>\n<pre><code>X_train = np.ndarray(shape=(N_ROWS, N_COLS), dtype=np.float32)\nfor idx, feature in enumerate(features):\n        X_train[:,idx] = train[feature].values.astype(np.float32)\n</code></pre>",
      "rawMarkdown": "Good tip!\nYou can however get a memory spike when creating the numpy array from the dataframe. My solution was to convert the columns one by one.\n```\nX_train = np.ndarray(shape=(N_ROWS, N_COLS), dtype=np.float32)\nfor idx, feature in enumerate(features):\n        X_train[:,idx] = train[feature].values.astype(np.float32)\n```",
      "votes": 8,
      "replies": [
        {
          "id": 1087454,
          "postDate": "2020-11-22T18:06:04.687Z",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> , I will try your method for further improvement</p>",
          "rawMarkdown": "thanks @markwijkhuizen , I will try your method for further improvement"
        },
        {
          "id": 1090454,
          "postDate": "2020-11-25T10:40:27.570Z",
          "content": "<p>This is the solution that worked for me. Many thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>! However instead of casting it column wise and assigning them to an empty array (this would still exhaust memory if the array you are initializing is large enough). A safer solution would be to take chunks of the DF row wise, cast them to float and do a stack operation. Putting a <code>time.sleep()</code> in between helped me as well. </p>\n<p>Also, if you are using features such as <code>timestamp</code> it is <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325#321840\" target=\"_blank\">probably safer</a> to cast them to <code>np.float64</code>.</p>",
          "rawMarkdown": "This is the solution that worked for me. Many thanks @markwijkhuizen! However instead of casting it column wise and assigning them to an empty array (this would still exhaust memory if the array you are initializing is large enough). A safer solution would be to take chunks of the DF row wise, cast them to float and do a stack operation. Putting a `time.sleep()` in between helped me as well. \n\nAlso, if you are using features such as `timestamp` it is [probably safer](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325#321840) to cast them to `np.float64`.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1094140,
      "postDate": "2020-11-28T10:58:09.510Z",
      "content": "<p>Thanks, it works to avoid memory spike but converting to <code>np.float32</code> looks to modify AUC (from 0.7721 to 0.7742 on one for my test) even if not I'm not using any <code>np.int64</code> (for timestamp). Pandas dataframe only contains <code>np.int8</code>, <code>bool</code>, <code>np.int32</code>, <code>np.int16</code>, <code>np.float32</code>. That's weird as <code>lgb.Dataset</code> converts everything to <code>np.float32</code>:<br>\n<a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/basic.html#Dataset\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/basic.html#Dataset</a><br>\nAnyone got similar issue?</p>",
      "rawMarkdown": "Thanks, it works to avoid memory spike but converting to `np.float32` looks to modify AUC (from 0.7721 to 0.7742 on one for my test) even if not I'm not using any `np.int64` (for timestamp). Pandas dataframe only contains `np.int8`, `bool`, `np.int32`, `np.int16`, `np.float32`. That's weird as `lgb.Dataset` converts everything to `np.float32`:\nhttps://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/basic.html#Dataset\nAnyone got similar issue?",
      "votes": 3
    },
    {
      "id": 1119483,
      "postDate": "2020-12-20T05:27:39.003Z",
      "content": "<p>What about categorical features, like <code>user_id</code>? If I convert it to <code>float32</code> type, does the model recognize it as a categorical features?</p>",
      "rawMarkdown": "What about categorical features, like `user_id`? If I convert it to `float32` type, does the model recognize it as a categorical features?",
      "votes": 1,
      "replies": [
        {
          "id": 1119877,
          "postDate": "2020-12-20T12:24:56.343Z",
          "content": "<p>Yes, It would. We can specify this using the <code>categorical_feature</code> parameter during model training when we do a <code>lgb.train</code>. We also need to specify the <code>feature_name</code> parameter when we create the lgb.Dataset. </p>",
          "rawMarkdown": "Yes, It would. We can specify this using the `categorical_feature` parameter during model training when we do a `lgb.train`. We also need to specify the `feature_name` parameter when we create the lgb.Dataset. "
        },
        {
          "id": 1119898,
          "postDate": "2020-12-20T12:39:21.523Z",
          "content": "<p>I see, but why do we need to specify <code>feature_name</code>? If this is because of <code>categorical_feature</code>, we can specify <code>feature_names</code> by their index!</p>",
          "rawMarkdown": "I see, but why do we need to specify `feature_name`? If this is because of `categorical_feature`, we can specify `feature_names` by their index!"
        },
        {
          "id": 1120006,
          "postDate": "2020-12-20T14:44:50.477Z",
          "content": "<p>The model does not recognize categorical features automatically when using numpy arrays, you need to specify the indices of the columns containing categorical features. The code below is a method I use to obtain these indices and specify them for training.</p>\n<pre><code>categorical_feature_idxs = []\nfor v in ['part',  'bundle_id', 'etc']: # your cat features\n    # TRAIN_FEATURES is a list with your training features\n    categorical_feature_idxs.append(TRAIN_FEATURES.index(v))\n\n# Create dataset\ntrain_data = lgb.Dataset(data=X_train, label=y_train, categorical_feature=None)\n\n# train    \nlgb.train(\n    train_set = train_data,\n    # specify the column indices of the categorical features\n   categorical_feature = categorical_feature_idxs,\n)\n</code></pre>",
          "rawMarkdown": "The model does not recognize categorical features automatically when using numpy arrays, you need to specify the indices of the columns containing categorical features. The code below is a method I use to obtain these indices and specify them for training.\n\n```\ncategorical_feature_idxs = []\nfor v in ['part',  'bundle_id', 'etc']: # your cat features\n    # TRAIN_FEATURES is a list with your training features\n    categorical_feature_idxs.append(TRAIN_FEATURES.index(v))\n\n# Create dataset\ntrain_data = lgb.Dataset(data=X_train, label=y_train, categorical_feature=None)\n\n# train    \nlgb.train(\n    train_set = train_data,\n    # specify the column indices of the categorical features\n   categorical_feature = categorical_feature_idxs,\n)\n```",
          "votes": 1
        },
        {
          "id": 1120014,
          "postDate": "2020-12-20T14:49:26.070Z",
          "content": "<p>It's like you said, <code>feature_name</code> is needed only when we have specify <code>categorical_feature</code> by their column names. This is not required if we specify the columns using their index. It's just that I find it more convenient to use the column names so as to not keep track of the indices :D</p>",
          "rawMarkdown": "It's like you said, `feature_name` is needed only when we have specify `categorical_feature` by their column names. This is not required if we specify the columns using their index. It's just that I find it more convenient to use the column names so as to not keep track of the indices :D"
        },
        {
          "id": 1121271,
          "postDate": "2020-12-21T14:04:39.803Z",
          "content": "<p>I tried but this converting to np.array makes my training slower :(</p>",
          "rawMarkdown": "I tried but this converting to np.array makes my training slower :("
        }
      ]
    },
    {
      "id": 1085643,
      "postDate": "2020-11-21T03:48:39.750Z",
      "content": "<p>Does this hurt auc score? Although I had same condition (imputation of missing value, and keeping float32), auc score in validation of pandas improves (auc over 0.7), but not in numpy (auc 0.5).  I guess part of reason is I use xgboost, not lightgbm</p>",
      "rawMarkdown": "Does this hurt auc score? Although I had same condition (imputation of missing value, and keeping float32), auc score in validation of pandas improves (auc over 0.7), but not in numpy (auc 0.5).  I guess part of reason is I use xgboost, not lightgbm",
      "replies": [
        {
          "id": 1087456,
          "postDate": "2020-11-22T18:08:08.843Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1090439,
          "postDate": "2020-11-25T10:32:34.240Z",
          "content": "<p>The problem may arise if when you are using features such as <code>timestamp</code>. Which has a max value much greater than what <code>np.float32</code> could contain. These values would then be rounded off leading to drop in AUC. Check discussion <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325#321840\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "The problem may arise if when you are using features such as `timestamp`. Which has a max value much greater than what `np.float32` could contain. These values would then be rounded off leading to drop in AUC. Check discussion [here](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325#321840)."
        },
        {
          "id": 1090462,
          "postDate": "2020-11-25T10:51:23.943Z",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\nThanks. So is it about trade-off, getting rid of features like timestamp or setting all features as float64?</p>",
          "rawMarkdown": "@doctorkael \nThanks. So is it about trade-off, getting rid of features like timestamp or setting all features as float64?"
        },
        {
          "id": 1090475,
          "postDate": "2020-11-25T11:05:46.790Z",
          "content": "<p>You could still use <code>np.float64</code> without running into memory issues. <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> solution would help you do just that. If column wise casting still causes you memory issues, trying casting rows in chunks, say 10,000 rows at a time, stacking them to a numpy array going all the way to end of your dataset. The problem is caused <em>not</em> since Kaggle RAM can't contain all the data in float64 format (16 GB can actually hold quite a lot more). Rather it is caused due to a <em><strong>sudden</strong> memory spike</em>. This can be mitigated by doing the operation ourselves in a <em>slower fashion</em> (time.sleep comes to mind).</p>\n<p><code>&gt;&gt;&gt; lgb.Dataset(dataframe)</code></p>\n<p>would internally be translated to this -&gt; </p>\n<p><code>&gt;&gt;&gt; lgb.Dataset(dataframe.values.astype('float'))</code></p>\n<p>Well, not exactly but a rough equivalent. Hence the abrupt memory spike. Hope this helps :)</p>",
          "rawMarkdown": "You could still use `np.float64` without running into memory issues. @markwijkhuizen solution would help you do just that. If column wise casting still causes you memory issues, trying casting rows in chunks, say 10,000 rows at a time, stacking them to a numpy array going all the way to end of your dataset. The problem is caused *not* since Kaggle RAM can't contain all the data in float64 format (16 GB can actually hold quite a lot more). Rather it is caused due to a _**sudden** memory spike_. This can be mitigated by doing the operation ourselves in a *slower fashion* (time.sleep comes to mind).\n\n```>>> lgb.Dataset(dataframe)```\n\nwould internally be translated to this -> \n\n```>>> lgb.Dataset(dataframe.values.astype('float'))```\n\nWell, not exactly but a rough equivalent. Hence the abrupt memory spike. Hope this helps :)",
          "votes": 1
        },
        {
          "id": 1091726,
          "postDate": "2020-11-26T08:30:40.587Z",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\nThis \"sudden\" memory spike makes sense to me because success/error occurs arbitrary despite same cell trial… <br>\nI'm using sklearn api and not sure your method could work as same as python-api. But I'll try it anyway, thank you!</p>",
          "rawMarkdown": "@doctorkael \nThis \"sudden\" memory spike makes sense to me because success/error occurs arbitrary despite same cell trial... \nI'm using sklearn api and not sure your method could work as same as python-api. But I'll try it anyway, thank you!",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1084857,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2020-11-20T13:13:02.557000",
      "content": "<p>Good tip!<br>\nYou can however get a memory spike when creating the numpy array from the dataframe. My solution was to convert the columns one by one.</p>\n<pre><code>X_train = np.ndarray(shape=(N_ROWS, N_COLS), dtype=np.float32)\nfor idx, feature in enumerate(features):\n        X_train[:,idx] = train[feature].values.astype(np.float32)\n</code></pre>",
      "votes": 8,
      "replies": [
        {
          "id": 1087454,
          "author_name": "Anurag Trivedi",
          "author_url": "",
          "post_date": "2020-11-22T18:06:04.687000",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> , I will try your method for further improvement</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1090454,
          "author_name": "Naresh Jagadeesan",
          "author_url": "",
          "post_date": "2020-11-25T10:40:27.570000",
          "content": "<p>This is the solution that worked for me. Many thanks <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>! However instead of casting it column wise and assigning them to an empty array (this would still exhaust memory if the array you are initializing is large enough). A safer solution would be to take chunks of the DF row wise, cast them to float and do a stack operation. Putting a <code>time.sleep()</code> in between helped me as well. </p>\n<p>Also, if you are using features such as <code>timestamp</code> it is <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325#321840\" target=\"_blank\">probably safer</a> to cast them to <code>np.float64</code>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1094140,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-11-28T10:58:09.510000",
      "content": "<p>Thanks, it works to avoid memory spike but converting to <code>np.float32</code> looks to modify AUC (from 0.7721 to 0.7742 on one for my test) even if not I'm not using any <code>np.int64</code> (for timestamp). Pandas dataframe only contains <code>np.int8</code>, <code>bool</code>, <code>np.int32</code>, <code>np.int16</code>, <code>np.float32</code>. That's weird as <code>lgb.Dataset</code> converts everything to <code>np.float32</code>:<br>\n<a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/basic.html#Dataset\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/basic.html#Dataset</a><br>\nAnyone got similar issue?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1119483,
      "author_name": "jwc",
      "author_url": "",
      "post_date": "2020-12-20T05:27:39.003000",
      "content": "<p>What about categorical features, like <code>user_id</code>? If I convert it to <code>float32</code> type, does the model recognize it as a categorical features?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1119877,
          "author_name": "Naresh Jagadeesan",
          "author_url": "",
          "post_date": "2020-12-20T12:24:56.343000",
          "content": "<p>Yes, It would. We can specify this using the <code>categorical_feature</code> parameter during model training when we do a <code>lgb.train</code>. We also need to specify the <code>feature_name</code> parameter when we create the lgb.Dataset. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119898,
          "author_name": "jwc",
          "author_url": "",
          "post_date": "2020-12-20T12:39:21.523000",
          "content": "<p>I see, but why do we need to specify <code>feature_name</code>? If this is because of <code>categorical_feature</code>, we can specify <code>feature_names</code> by their index!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1120006,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2020-12-20T14:44:50.477000",
          "content": "<p>The model does not recognize categorical features automatically when using numpy arrays, you need to specify the indices of the columns containing categorical features. The code below is a method I use to obtain these indices and specify them for training.</p>\n<pre><code>categorical_feature_idxs = []\nfor v in ['part',  'bundle_id', 'etc']: # your cat features\n    # TRAIN_FEATURES is a list with your training features\n    categorical_feature_idxs.append(TRAIN_FEATURES.index(v))\n\n# Create dataset\ntrain_data = lgb.Dataset(data=X_train, label=y_train, categorical_feature=None)\n\n# train    \nlgb.train(\n    train_set = train_data,\n    # specify the column indices of the categorical features\n   categorical_feature = categorical_feature_idxs,\n)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120014,
          "author_name": "Naresh Jagadeesan",
          "author_url": "",
          "post_date": "2020-12-20T14:49:26.070000",
          "content": "<p>It's like you said, <code>feature_name</code> is needed only when we have specify <code>categorical_feature</code> by their column names. This is not required if we specify the columns using their index. It's just that I find it more convenient to use the column names so as to not keep track of the indices :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1121271,
          "author_name": "jwc",
          "author_url": "",
          "post_date": "2020-12-21T14:04:39.803000",
          "content": "<p>I tried but this converting to np.array makes my training slower :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1085643,
      "author_name": "tand",
      "author_url": "",
      "post_date": "2020-11-21T03:48:39.750000",
      "content": "<p>Does this hurt auc score? Although I had same condition (imputation of missing value, and keeping float32), auc score in validation of pandas improves (auc over 0.7), but not in numpy (auc 0.5).  I guess part of reason is I use xgboost, not lightgbm</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1087456,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-22T18:08:08.843000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1090439,
          "author_name": "Naresh Jagadeesan",
          "author_url": "",
          "post_date": "2020-11-25T10:32:34.240000",
          "content": "<p>The problem may arise if when you are using features such as <code>timestamp</code>. Which has a max value much greater than what <code>np.float32</code> could contain. These values would then be rounded off leading to drop in AUC. Check discussion <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55325#321840\" target=\"_blank\">here</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1090462,
          "author_name": "tand",
          "author_url": "",
          "post_date": "2020-11-25T10:51:23.943000",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\nThanks. So is it about trade-off, getting rid of features like timestamp or setting all features as float64?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1090475,
          "author_name": "Naresh Jagadeesan",
          "author_url": "",
          "post_date": "2020-11-25T11:05:46.790000",
          "content": "<p>You could still use <code>np.float64</code> without running into memory issues. <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> solution would help you do just that. If column wise casting still causes you memory issues, trying casting rows in chunks, say 10,000 rows at a time, stacking them to a numpy array going all the way to end of your dataset. The problem is caused <em>not</em> since Kaggle RAM can't contain all the data in float64 format (16 GB can actually hold quite a lot more). Rather it is caused due to a <em><strong>sudden</strong> memory spike</em>. This can be mitigated by doing the operation ourselves in a <em>slower fashion</em> (time.sleep comes to mind).</p>\n<p><code>&gt;&gt;&gt; lgb.Dataset(dataframe)</code></p>\n<p>would internally be translated to this -&gt; </p>\n<p><code>&gt;&gt;&gt; lgb.Dataset(dataframe.values.astype('float'))</code></p>\n<p>Well, not exactly but a rough equivalent. Hence the abrupt memory spike. Hope this helps :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091726,
          "author_name": "tand",
          "author_url": "",
          "post_date": "2020-11-26T08:30:40.587000",
          "content": "<p><a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a> <br>\nThis \"sudden\" memory spike makes sense to me because success/error occurs arbitrary despite same cell trial… <br>\nI'm using sklearn api and not sure your method could work as same as python-api. But I'll try it anyway, thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1084772": "Don't use DF to create lightgbm dataset, rather use np array:\n\nSample code:\n```\nX_train_np = train.values.astype(np.float32)\nX_valid_np = valid.values.astype(np.float32)\nfeatures = train.columns\nlgb_train = lgb.Dataset(X_train_np, label=y_tr, feature_name=list(features))\nlgb_valid = lgb.Dataset(X_valid_np, label=y_va, feature_name=list(features))\ndel train, y_tr\n_=gc.collect()\n```\n\nReference - https://www.kaggle.com/c/m5-forecasting-accuracy/discussion/149754\ncredit: @alankabisov ",
    "1084857": "Good tip!\nYou can however get a memory spike when creating the numpy array from the dataframe. My solution was to convert the columns one by one.\n```\nX_train = np.ndarray(shape=(N_ROWS, N_COLS), dtype=np.float32)\nfor idx, feature in enumerate(features):\n        X_train[:,idx] = train[feature].values.astype(np.float32)\n```",
    "1094140": "Thanks, it works to avoid memory spike but converting to `np.float32` looks to modify AUC (from 0.7721 to 0.7742 on one for my test) even if not I'm not using any `np.int64` (for timestamp). Pandas dataframe only contains `np.int8`, `bool`, `np.int32`, `np.int16`, `np.float32`. That's weird as `lgb.Dataset` converts everything to `np.float32`:\nhttps://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/basic.html#Dataset\nAnyone got similar issue?",
    "1119483": "What about categorical features, like `user_id`? If I convert it to `float32` type, does the model recognize it as a categorical features?",
    "1085643": "Does this hurt auc score? Although I had same condition (imputation of missing value, and keeping float32), auc score in validation of pandas improves (auc over 0.7), but not in numpy (auc 0.5).  I guess part of reason is I use xgboost, not lightgbm"
  }
}