{
  "id": 553925,
  "title": "GPU RAM usage in TF and Torch?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/553925",
  "author_name": "",
  "post_date": "2024-12-29T09:38:10.785448900Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Tensorflow taking a lot of GPU RAM while torch only take a little? </p>\n<p>I tried to load only 2 day of data by tensorflow:</p>\n<pre><code> ():\n    data = pl.scan_parquet(path).select(\n        pl.(),).(\n        pl.col().gt(start_dt),\n        pl.col().le(end_dt),\n    ).fill_null().fill_null()\n\n    data = data.collect().to_pandas()\n\n    data.replace([np.inf, -np.inf], , inplace=)\n     data\n</code></pre>\n<pre><code>X_train = load_data(training_resp_lag_path, start_dt=, end_dt=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7c2b9f0aaa0b042e6f04a6aae13ead38%2F_2024-12-29_171636_504.png?generation=1735463813123187&amp;alt=media\" alt=\"\"></p>\n<p>and after running tensorflow:</p>\n<pre><code>test = tf.data.Dataset.from_tensor_slices(X_train)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F3a93bc58373e85f56f2bba8b8c3c9220%2Fwechat_2024-12-29_171842_716.png?generation=1735463938077122&amp;alt=media\" alt=\"\"></p>\n<p>only two day of data already taking 14 gig data? how is that. and if I load more days, like 200</p>\n<pre><code>X_train = load_data(training_resp_lag_path, start_dt=, end_dt=)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fdb6e7436adefff1cf156a8ce1a1d7746%2Fwechat_2024-12-29_172010_446.png?generation=1735464033383761&amp;alt=media\" alt=\"\"></p>\n<p>the memory usage did not seem to change much.<br>\nbut if I use even more like from 600-1690, it will blow up the GPU RAM.</p>\n<pre><code>X_train = load_data(training_resp_lag_path, start_dt=, end_dt=)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n\nW external/local_xla/xla/tsl/framework/bfc_allocator.cc:] Allocator (GPU_0_bfc) ran out of memory trying to allocate GiB (rounded to )requested by op _EagerConst\nIf the cause  memory fragmentation maybe the environment variable  will improve the situation. ...\n</code></pre>\n<p>so this does really limit my ability to use tensorflow to train model. Batching the datasize did not really help with gpu memory, using batch size =1000 or batch size =8000 did not seemed to make much difference, but how much data I was loading was really matters. How to solve this?</p>\n<p>while using torch, refering to the public code that got 0.0077:<br>\n<a href=\"https://www.kaggle.com/code/voix97/jane-street-rmf-training-nn\" target=\"_blank\">https://www.kaggle.com/code/voix97/jane-street-rmf-training-nn</a></p>\n<p>while training with this code, the GPU usage was only 0.3/15 gig space.<br>\nhow to config tensorflow to not taking so much of space?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fa337c65052a14a3ea0ff0c16e90a37fd%2Fwechat_2024-12-29_173649_987.png?generation=1735465073974414&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3083313",
      "postDate": "12/29/2024 09:38:10",
      "content": "<p>Tensorflow taking a lot of GPU RAM while torch only take a little? </p>\n<p>I tried to load only 2 day of data by tensorflow:</p>\n<pre><code> ():\n    data = pl.scan_parquet(path).select(\n        pl.(),).(\n        pl.col().gt(start_dt),\n        pl.col().le(end_dt),\n    ).fill_null().fill_null()\n\n    data = data.collect().to_pandas()\n\n    data.replace([np.inf, -np.inf], , inplace=)\n     data\n</code></pre>\n<pre><code>X_train = load_data(training_resp_lag_path, start_dt=, end_dt=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7c2b9f0aaa0b042e6f04a6aae13ead38%2F_2024-12-29_171636_504.png?generation=1735463813123187&amp;alt=media\" alt=\"\"></p>\n<p>and after running tensorflow:</p>\n<pre><code>test = tf.data.Dataset.from_tensor_slices(X_train)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F3a93bc58373e85f56f2bba8b8c3c9220%2Fwechat_2024-12-29_171842_716.png?generation=1735463938077122&amp;alt=media\" alt=\"\"></p>\n<p>only two day of data already taking 14 gig data? how is that. and if I load more days, like 200</p>\n<pre><code>X_train = load_data(training_resp_lag_path, start_dt=, end_dt=)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fdb6e7436adefff1cf156a8ce1a1d7746%2Fwechat_2024-12-29_172010_446.png?generation=1735464033383761&amp;alt=media\" alt=\"\"></p>\n<p>the memory usage did not seem to change much.<br>\nbut if I use even more like from 600-1690, it will blow up the GPU RAM.</p>\n<pre><code>X_train = load_data(training_resp_lag_path, start_dt=, end_dt=)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n\nW external/local_xla/xla/tsl/framework/bfc_allocator.cc:] Allocator (GPU_0_bfc) ran out of memory trying to allocate GiB (rounded to )requested by op _EagerConst\nIf the cause  memory fragmentation maybe the environment variable  will improve the situation. ...\n</code></pre>\n<p>so this does really limit my ability to use tensorflow to train model. Batching the datasize did not really help with gpu memory, using batch size =1000 or batch size =8000 did not seemed to make much difference, but how much data I was loading was really matters. How to solve this?</p>\n<p>while using torch, refering to the public code that got 0.0077:<br>\n<a href=\"https://www.kaggle.com/code/voix97/jane-street-rmf-training-nn\" target=\"_blank\">https://www.kaggle.com/code/voix97/jane-street-rmf-training-nn</a></p>\n<p>while training with this code, the GPU usage was only 0.3/15 gig space.<br>\nhow to config tensorflow to not taking so much of space?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fa337c65052a14a3ea0ff0c16e90a37fd%2Fwechat_2024-12-29_173649_987.png?generation=1735465073974414&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Tensorflow taking a lot of GPU RAM while torch only take a little? \n\nI tried to load only 2 day of data by tensorflow:\n```python\ndef load_data(path, start_dt, end_dt):\n    data = pl.scan_parquet(path).select(\n        pl.all(),).filter(\n        pl.col(\"date_id\").gt(start_dt),\n        pl.col(\"date_id\").le(end_dt),\n    ).fill_null(0).fill_null(0)\n\n    data = data.collect().to_pandas()\n\n    data.replace([np.inf, -np.inf], 0, inplace=True)\n    return data\n```\n\n```python\nX_train = load_data(training_resp_lag_path, start_dt=1688, end_dt=1690)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7c2b9f0aaa0b042e6f04a6aae13ead38%2F_2024-12-29_171636_504.png?generation=1735463813123187&alt=media)\n\nand after running tensorflow:\n\n```python\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F3a93bc58373e85f56f2bba8b8c3c9220%2Fwechat_2024-12-29_171842_716.png?generation=1735463938077122&alt=media)\n\n\nonly two day of data already taking 14 gig data? how is that. and if I load more days, like 200\n```python\nX_train = load_data(training_resp_lag_path, start_dt=1500, end_dt=1690)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fdb6e7436adefff1cf156a8ce1a1d7746%2Fwechat_2024-12-29_172010_446.png?generation=1735464033383761&alt=media)\n\n\nthe memory usage did not seem to change much.\nbut if I use even more like from 600-1690, it will blow up the GPU RAM.\n\n```python\nX_train = load_data(training_resp_lag_path, start_dt=600, end_dt=1690)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n\nW external/local_xla/xla/tsl/framework/bfc_allocator.cc:497] Allocator (GPU_0_bfc) ran out of memory trying to allocate 28.57GiB (rounded to 30673070336)requested by op _EagerConst\nIf the cause is memory fragmentation maybe the environment variable 'TF_GPU_ALLOCATOR=cuda_malloc_async' will improve the situation. ...\n```\n\nso this does really limit my ability to use tensorflow to train model. Batching the datasize did not really help with gpu memory, using batch size =1000 or batch size =8000 did not seemed to make much difference, but how much data I was loading was really matters. How to solve this?\n\nwhile using torch, refering to the public code that got 0.0077:\nhttps://www.kaggle.com/code/voix97/jane-street-rmf-training-nn\n\nwhile training with this code, the GPU usage was only 0.3/15 gig space.\nhow to config tensorflow to not taking so much of space?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fa337c65052a14a3ea0ff0c16e90a37fd%2Fwechat_2024-12-29_173649_987.png?generation=1735465073974414&alt=media)",
      "votes": null
    },
    {
      "id": "3083373",
      "postDate": "12/29/2024 11:54:25",
      "content": "<p>Tensorflow always allocates the maximum available memory on the GPU. There are probably ways to force it to allocate only a part, but the total allocation is the usual behavior.  <br>\nWhat is strange here is that test = tf.data.Dataset.from_tensor_slices(X_train) tried to save the data on the GPU. The behavior I am familiar with is saving the dataset on the system RAM. Only the model itself is supposed to be saved on the GPU RAM. Maybe you forced it somehow to save the data on the GPU? </p>",
      "rawMarkdown": "Tensorflow always allocates the maximum available memory on the GPU. There are probably ways to force it to allocate only a part, but the total allocation is the usual behavior.  \nWhat is strange here is that test = tf.data.Dataset.from_tensor_slices(X_train) tried to save the data on the GPU. The behavior I am familiar with is saving the dataset on the system RAM. Only the model itself is supposed to be saved on the GPU RAM. Maybe you forced it somehow to save the data on the GPU?",
      "votes": null
    },
    {
      "id": "3083377",
      "postDate": "12/29/2024 12:09:34",
      "content": "<p>Specify the device when creating tensor slices and model fitting</p>\n<pre><code>with tf():\n          train_dataset = tf(\n                (features_batch, labels_batch, weights_batch)\n            )\n.....\n\nwith tf():\n            model(\n                train_dataset,\n                epochs=epochs,\n                validation_data=valid_dataset,\n                callbacks=,\n            )\n</code></pre>",
      "rawMarkdown": "Specify the device when creating tensor slices and model fitting\n\n```\nwith tf.device(\"/CPU:0\"):\n          train_dataset = tf.data.Dataset.from_tensor_slices(\n                (features_batch, labels_batch, weights_batch)\n            )\n.....\n\nwith tf.device(\"/GPU:0\"):\n            model.fit(\n                train_dataset,\n                epochs=epochs,\n                validation_data=valid_dataset,\n                callbacks=[callback],\n            )\n```",
      "votes": null
    },
    {
      "id": "3083447",
      "postDate": "12/29/2024 13:56:57",
      "content": "<p>my question is, using tensorflow cannot really control how many gpu is using and easy to blow the VRAM, how to use a large dataset while not blow up the gpu ram?</p>",
      "rawMarkdown": "my question is, using tensorflow cannot really control how many gpu is using and easy to blow the VRAM, how to use a large dataset while not blow up the gpu ram?",
      "votes": null
    },
    {
      "id": "3083465",
      "postDate": "12/29/2024 14:16:14",
      "content": "<p>Per Tensorflow documentation, would this help?</p>\n<blockquote>\n  <p>If memory growth is enabled for a PhysicalDevice, the runtime initialization will not allocate all memory on the device.<br>\n  <code>physical_devices = tf.config.list_physical_devices('GPU')\ntry:\n tf.config.experimental.set_memory_growth(physical_devices[0], True)\nexcept:\n # Invalid device or cannot modify virtual devices once initialized.\n pass</code></p>\n</blockquote>",
      "rawMarkdown": "Per Tensorflow documentation, would this help?\n> If memory growth is enabled for a PhysicalDevice, the runtime initialization will not allocate all memory on the device.\n\n`physical_devices = tf.config.list_physical_devices('GPU')\ntry:\n  tf.config.experimental.set_memory_growth(physical_devices[0], True)\nexcept:\n  # Invalid device or cannot modify virtual devices once initialized.\n  pass`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3083373,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "12/29/2024 11:54:25",
      "content": "<p>Tensorflow always allocates the maximum available memory on the GPU. There are probably ways to force it to allocate only a part, but the total allocation is the usual behavior.  <br>\nWhat is strange here is that test = tf.data.Dataset.from_tensor_slices(X_train) tried to save the data on the GPU. The behavior I am familiar with is saving the dataset on the system RAM. Only the model itself is supposed to be saved on the GPU RAM. Maybe you forced it somehow to save the data on the GPU? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3083377,
          "author_name": "edmund870",
          "author_url": "",
          "post_date": "12/29/2024 12:09:34",
          "content": "<p>Specify the device when creating tensor slices and model fitting</p>\n<pre><code>with tf():\n          train_dataset = tf(\n                (features_batch, labels_batch, weights_batch)\n            )\n.....\n\nwith tf():\n            model(\n                train_dataset,\n                epochs=epochs,\n                validation_data=valid_dataset,\n                callbacks=,\n            )\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3083447,
      "author_name": "zoutain",
      "author_url": "",
      "post_date": "12/29/2024 13:56:57",
      "content": "<p>my question is, using tensorflow cannot really control how many gpu is using and easy to blow the VRAM, how to use a large dataset while not blow up the gpu ram?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3083465,
          "author_name": "edmund870",
          "author_url": "",
          "post_date": "12/29/2024 14:16:14",
          "content": "<p>Per Tensorflow documentation, would this help?</p>\n<blockquote>\n  <p>If memory growth is enabled for a PhysicalDevice, the runtime initialization will not allocate all memory on the device.<br>\n  <code>physical_devices = tf.config.list_physical_devices('GPU')\ntry:\n tf.config.experimental.set_memory_growth(physical_devices[0], True)\nexcept:\n # Invalid device or cannot modify virtual devices once initialized.\n pass</code></p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3083313": "Tensorflow taking a lot of GPU RAM while torch only take a little? \n\nI tried to load only 2 day of data by tensorflow:\n```python\ndef load_data(path, start_dt, end_dt):\n    data = pl.scan_parquet(path).select(\n        pl.all(),).filter(\n        pl.col(\"date_id\").gt(start_dt),\n        pl.col(\"date_id\").le(end_dt),\n    ).fill_null(0).fill_null(0)\n\n    data = data.collect().to_pandas()\n\n    data.replace([np.inf, -np.inf], 0, inplace=True)\n    return data\n```\n\n```python\nX_train = load_data(training_resp_lag_path, start_dt=1688, end_dt=1690)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F7c2b9f0aaa0b042e6f04a6aae13ead38%2F_2024-12-29_171636_504.png?generation=1735463813123187&alt=media)\n\nand after running tensorflow:\n\n```python\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2F3a93bc58373e85f56f2bba8b8c3c9220%2Fwechat_2024-12-29_171842_716.png?generation=1735463938077122&alt=media)\n\n\nonly two day of data already taking 14 gig data? how is that. and if I load more days, like 200\n```python\nX_train = load_data(training_resp_lag_path, start_dt=1500, end_dt=1690)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fdb6e7436adefff1cf156a8ce1a1d7746%2Fwechat_2024-12-29_172010_446.png?generation=1735464033383761&alt=media)\n\n\nthe memory usage did not seem to change much.\nbut if I use even more like from 600-1690, it will blow up the GPU RAM.\n\n```python\nX_train = load_data(training_resp_lag_path, start_dt=600, end_dt=1690)\ntest = tf.data.Dataset.from_tensor_slices(X_train)\n\nW external/local_xla/xla/tsl/framework/bfc_allocator.cc:497] Allocator (GPU_0_bfc) ran out of memory trying to allocate 28.57GiB (rounded to 30673070336)requested by op _EagerConst\nIf the cause is memory fragmentation maybe the environment variable 'TF_GPU_ALLOCATOR=cuda_malloc_async' will improve the situation. ...\n```\n\nso this does really limit my ability to use tensorflow to train model. Batching the datasize did not really help with gpu memory, using batch size =1000 or batch size =8000 did not seemed to make much difference, but how much data I was loading was really matters. How to solve this?\n\nwhile using torch, refering to the public code that got 0.0077:\nhttps://www.kaggle.com/code/voix97/jane-street-rmf-training-nn\n\nwhile training with this code, the GPU usage was only 0.3/15 gig space.\nhow to config tensorflow to not taking so much of space?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4858569%2Fa337c65052a14a3ea0ff0c16e90a37fd%2Fwechat_2024-12-29_173649_987.png?generation=1735465073974414&alt=media)",
    "3083373": "Tensorflow always allocates the maximum available memory on the GPU. There are probably ways to force it to allocate only a part, but the total allocation is the usual behavior.  \nWhat is strange here is that test = tf.data.Dataset.from_tensor_slices(X_train) tried to save the data on the GPU. The behavior I am familiar with is saving the dataset on the system RAM. Only the model itself is supposed to be saved on the GPU RAM. Maybe you forced it somehow to save the data on the GPU?",
    "3083377": "Specify the device when creating tensor slices and model fitting\n\n```\nwith tf.device(\"/CPU:0\"):\n          train_dataset = tf.data.Dataset.from_tensor_slices(\n                (features_batch, labels_batch, weights_batch)\n            )\n.....\n\nwith tf.device(\"/GPU:0\"):\n            model.fit(\n                train_dataset,\n                epochs=epochs,\n                validation_data=valid_dataset,\n                callbacks=[callback],\n            )\n```",
    "3083447": "my question is, using tensorflow cannot really control how many gpu is using and easy to blow the VRAM, how to use a large dataset while not blow up the gpu ram?",
    "3083465": "Per Tensorflow documentation, would this help?\n> If memory growth is enabled for a PhysicalDevice, the runtime initialization will not allocate all memory on the device.\n\n`physical_devices = tf.config.list_physical_devices('GPU')\ntry:\n  tf.config.experimental.set_memory_growth(physical_devices[0], True)\nexcept:\n  # Invalid device or cannot modify virtual devices once initialized.\n  pass`"
  },
  "source": "meta"
}