{
  "id": 206538,
  "title": "Any clue about avoiding memory explosion?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206538",
  "author_name": "",
  "post_date": "2020-12-25T06:31:01.522372600Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>My dataset has 30M rows with 16 features. It uses 1.6 GB at the beginning.</p>\n<p>But when I use the following conversion for LGBM, the momery increases dramatically to over 16GB.<br>\n<code>train = train[FEATS].values.astype('float32')</code><br>\n<code>valid = valid[FEATS].values.astype('float32')</code></p>\n<p>Anyone knows how to optimize this operation?</p>",
  "messages": [
    {
      "id": "1125873",
      "postDate": "12/25/2020 06:31:01",
      "content": "<p>My dataset has 30M rows with 16 features. It uses 1.6 GB at the beginning.</p>\n<p>But when I use the following conversion for LGBM, the momery increases dramatically to over 16GB.<br>\n<code>train = train[FEATS].values.astype('float32')</code><br>\n<code>valid = valid[FEATS].values.astype('float32')</code></p>\n<p>Anyone knows how to optimize this operation?</p>",
      "rawMarkdown": "My dataset has 30M rows with 16 features. It uses 1.6 GB at the beginning.\n\nBut when I use the following conversion for LGBM, the momery increases dramatically to over 16GB.\n`train = train[FEATS].values.astype('float32')`\n`valid = valid[FEATS].values.astype('float32')`\n\nAnyone knows how to optimize this operation?",
      "votes": null
    },
    {
      "id": "1125907",
      "postDate": "12/25/2020 07:06:43",
      "content": "<p>I believe the problem is using <code>values</code>. That converts the Pandas dataframe into a Numpy array and then both are in memory. Try the following.</p>\n<pre><code>import gc\nfor f in FEATS:\n    train[f] = train[f].astype('float32')\n    valid[f] = valid[f].astype('float32')\n    _ = gc.collect()\n</code></pre>",
      "rawMarkdown": "I believe the problem is using `values`. That converts the Pandas dataframe into a Numpy array and then both are in memory. Try the following.\n\n    import gc\n    for f in FEATS:\n        train[f] = train[f].astype('float32')\n        valid[f] = valid[f].astype('float32')\n        _ = gc.collect()",
      "votes": null
    },
    {
      "id": "1125989",
      "postDate": "12/25/2020 08:13:57",
      "content": "<p>NB -&gt; This reduced performance for me; And yes, this avoid's the spike when creating gbm datasets as otherwise it will create a copy! </p>\n<p>Also, I was able to train whole 100M rows in &lt; 64GB. (peak mem was at 51GB)</p>",
      "rawMarkdown": "NB -> This reduced performance for me; And yes, this avoid's the spike when creating gbm datasets as otherwise it will create a copy! \n\nAlso, I was able to train whole 100M rows in < 64GB. (peak mem was at 51GB)",
      "votes": null
    },
    {
      "id": "1126006",
      "postDate": "12/25/2020 08:32:55",
      "content": "<p>still quite a lot of memory :)</p>",
      "rawMarkdown": "still quite a lot of memory :)",
      "votes": null
    },
    {
      "id": "1126011",
      "postDate": "12/25/2020 08:41:31",
      "content": "<blockquote>\n  <p>still quite a lot of memory :)</p>\n</blockquote>\n<p>:(; Ya, it's still quite a lot; Any tips to reduce it further would be appreciated highly Yih!</p>",
      "rawMarkdown": ">still quite a lot of memory :)\n\n:(; Ya, it's still quite a lot; Any tips to reduce it further would be appreciated highly Yih!",
      "votes": null
    },
    {
      "id": "1126013",
      "postDate": "12/25/2020 08:48:17",
      "content": "<p>Not from me - I use tensorflow record file to store my data, so I don't have the memory issue for training. You are using pytorch I think, and I am not familiar with it ….</p>",
      "rawMarkdown": "Not from me - I use tensorflow record file to store my data, so I don't have the memory issue for training. You are using pytorch I think, and I am not familiar with it ....",
      "votes": null
    },
    {
      "id": "1126028",
      "postDate": "12/25/2020 09:15:49",
      "content": "<p>I had the same problem.<br>\nThe method by <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">this discussion</a> worked for me.</p>",
      "rawMarkdown": "I had the same problem.\nThe method by @markwijkhuizen in [this discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245) worked for me.",
      "votes": null
    },
    {
      "id": "1126552",
      "postDate": "12/25/2020 17:32:37",
      "content": "<p>Somehow it does not work for me😂</p>",
      "rawMarkdown": "Somehow it does not work for me😂",
      "votes": null
    },
    {
      "id": "1126623",
      "postDate": "12/25/2020 18:26:05",
      "content": "<p>Updated:</p>\n<p>Finally I found a way that works for me. This will also cause memory spike but it is lighter than my original way. At least I can train my model now.</p>\n<p><code>train = train.astype('float32')</code><br>\n<code>valid = valid.astype('float32')</code></p>\n<p><code>lgb_train = lgb.Dataset(train[FEATS], train[TARGET])</code><br>\n<code>lgb_valid = lgb.Dataset(valid[FEATS], valid[TARGET])</code><br>\n<code>del train, valid</code><br>\n<code>_=gc.collect()</code></p>\n<p>Then, <br>\n<code>model = lgb.train(...)</code></p>",
      "rawMarkdown": "Updated:\n\nFinally I found a way that works for me. This will also cause memory spike but it is lighter than my original way. At least I can train my model now.\n\n`train = train.astype('float32')`\n`valid = valid.astype('float32')`\n\n`lgb_train = lgb.Dataset(train[FEATS], train[TARGET])`\n`lgb_valid = lgb.Dataset(valid[FEATS], valid[TARGET])`\n`del train, valid`\n`_=gc.collect()`\n\nThen, \n`model = lgb.train(...)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1125907,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/25/2020 07:06:43",
      "content": "<p>I believe the problem is using <code>values</code>. That converts the Pandas dataframe into a Numpy array and then both are in memory. Try the following.</p>\n<pre><code>import gc\nfor f in FEATS:\n    train[f] = train[f].astype('float32')\n    valid[f] = valid[f].astype('float32')\n    _ = gc.collect()\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1125989,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/25/2020 08:13:57",
          "content": "<p>NB -&gt; This reduced performance for me; And yes, this avoid's the spike when creating gbm datasets as otherwise it will create a copy! </p>\n<p>Also, I was able to train whole 100M rows in &lt; 64GB. (peak mem was at 51GB)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126006,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/25/2020 08:32:55",
          "content": "<p>still quite a lot of memory :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126011,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/25/2020 08:41:31",
          "content": "<blockquote>\n  <p>still quite a lot of memory :)</p>\n</blockquote>\n<p>:(; Ya, it's still quite a lot; Any tips to reduce it further would be appreciated highly Yih!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126013,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/25/2020 08:48:17",
          "content": "<p>Not from me - I use tensorflow record file to store my data, so I don't have the memory issue for training. You are using pytorch I think, and I am not familiar with it ….</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1126028,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "12/25/2020 09:15:49",
      "content": "<p>I had the same problem.<br>\nThe method by <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245\" target=\"_blank\">this discussion</a> worked for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1126552,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "12/25/2020 17:32:37",
          "content": "<p>Somehow it does not work for me😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1126623,
      "author_name": "woshifym",
      "author_url": "",
      "post_date": "12/25/2020 18:26:05",
      "content": "<p>Updated:</p>\n<p>Finally I found a way that works for me. This will also cause memory spike but it is lighter than my original way. At least I can train my model now.</p>\n<p><code>train = train.astype('float32')</code><br>\n<code>valid = valid.astype('float32')</code></p>\n<p><code>lgb_train = lgb.Dataset(train[FEATS], train[TARGET])</code><br>\n<code>lgb_valid = lgb.Dataset(valid[FEATS], valid[TARGET])</code><br>\n<code>del train, valid</code><br>\n<code>_=gc.collect()</code></p>\n<p>Then, <br>\n<code>model = lgb.train(...)</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125873": "My dataset has 30M rows with 16 features. It uses 1.6 GB at the beginning.\n\nBut when I use the following conversion for LGBM, the momery increases dramatically to over 16GB.\n`train = train[FEATS].values.astype('float32')`\n`valid = valid[FEATS].values.astype('float32')`\n\nAnyone knows how to optimize this operation?",
    "1125907": "I believe the problem is using `values`. That converts the Pandas dataframe into a Numpy array and then both are in memory. Try the following.\n\n    import gc\n    for f in FEATS:\n        train[f] = train[f].astype('float32')\n        valid[f] = valid[f].astype('float32')\n        _ = gc.collect()",
    "1125989": "NB -> This reduced performance for me; And yes, this avoid's the spike when creating gbm datasets as otherwise it will create a copy! \n\nAlso, I was able to train whole 100M rows in < 64GB. (peak mem was at 51GB)",
    "1126006": "still quite a lot of memory :)",
    "1126011": ">still quite a lot of memory :)\n\n:(; Ya, it's still quite a lot; Any tips to reduce it further would be appreciated highly Yih!",
    "1126013": "Not from me - I use tensorflow record file to store my data, so I don't have the memory issue for training. You are using pytorch I think, and I am not familiar with it ....",
    "1126028": "I had the same problem.\nThe method by @markwijkhuizen in [this discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198245) worked for me.",
    "1126552": "Somehow it does not work for me😂",
    "1126623": "Updated:\n\nFinally I found a way that works for me. This will also cause memory spike but it is lighter than my original way. At least I can train my model now.\n\n`train = train.astype('float32')`\n`valid = valid.astype('float32')`\n\n`lgb_train = lgb.Dataset(train[FEATS], train[TARGET])`\n`lgb_valid = lgb.Dataset(valid[FEATS], valid[TARGET])`\n`del train, valid`\n`_=gc.collect()`\n\nThen, \n`model = lgb.train(...)`"
  },
  "source": "meta"
}