{
  "id": 142043,
  "title": "How to freeup GPU memory",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/142043",
  "author_name": "yuvaramsingh",
  "post_date": "2020-04-08T17:13:38.823000",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi\ni am trying to run two models. one in TF and another in Pytorch . i am facing issue with loading pytorch model after finishing up with tf . is there a better way to cleanup the GPU RAM used by tf </p>",
  "messages": [
    {
      "id": 801664,
      "postDate": "2020-04-08T17:13:38.823Z",
      "content": "<p>Hi\ni am trying to run two models. one in TF and another in Pytorch . i am facing issue with loading pytorch model after finishing up with tf . is there a better way to cleanup the GPU RAM used by tf </p>",
      "rawMarkdown": "Hi\ni am trying to run two models. one in TF and another in Pytorch . i am facing issue with loading pytorch model after finishing up with tf . is there a better way to cleanup the GPU RAM used by tf ",
      "votes": 1
    },
    {
      "id": 874491,
      "postDate": "2020-06-05T03:41:15.423Z",
      "content": "<p>This is tricky business. TensorFlow won't release memory after it has run. Therefore the best way to run both TensorFlow and PyTorch in the same Kaggle notebook is to limit the amount of memory TF can use</p>\n\n<pre><code>import tensorflow as tf\nGIGS = 8\ngpus = tf.config.experimental.list_physical_devices('GPU')\nif gpus:\n  # Restrict TensorFlow memory\n  try:\n    tf.config.experimental.set_virtual_device_configuration(\n        gpus[0],\n        [tf.config.experimental.VirtualDeviceConfiguration(memory_limit=1024*GIGS)])\n    logical_gpus = tf.config.experimental.list_logical_devices('GPU')\n    print(len(gpus), \"Physical GPUs,\", len(logical_gpus), \"Logical GPUs\")\n  except RuntimeError as e:\n     # Virtual devices must be set before GPUs have been initialized\n     print(e)\n</code></pre>\n\n<p>Set the variable <code>GIGS</code> to the amount of memory you wish to restrict TF to. Note that Kaggle notebooks use GPU P100 and have a total of 16GB VRAM</p>",
      "rawMarkdown": "This is tricky business. TensorFlow won't release memory after it has run. Therefore the best way to run both TensorFlow and PyTorch in the same Kaggle notebook is to limit the amount of memory TF can use\n\n    import tensorflow as tf\n    GIGS = 8\n    gpus = tf.config.experimental.list_physical_devices('GPU')\n    if gpus:\n      # Restrict TensorFlow memory\n      try:\n        tf.config.experimental.set_virtual_device_configuration(\n            gpus[0],\n            [tf.config.experimental.VirtualDeviceConfiguration(memory_limit=1024*GIGS)])\n        logical_gpus = tf.config.experimental.list_logical_devices('GPU')\n        print(len(gpus), \"Physical GPUs,\", len(logical_gpus), \"Logical GPUs\")\n      except RuntimeError as e:\n         # Virtual devices must be set before GPUs have been initialized\n         print(e)\n\nSet the variable `GIGS` to the amount of memory you wish to restrict TF to. Note that Kaggle notebooks use GPU P100 and have a total of 16GB VRAM",
      "votes": 2,
      "replies": [
        {
          "id": 877454,
          "postDate": "2020-06-07T16:02:37.663Z",
          "content": "<p>Thanks for the info, just wonder, if memory for TF is set small, would it have out of memory error? </p>",
          "rawMarkdown": "Thanks for the info, just wonder, if memory for TF is set small, would it have out of memory error? ",
          "isDeleted": true
        },
        {
          "id": 877546,
          "postDate": "2020-06-07T17:28:46.650Z",
          "content": "<p>Depends what you're doing. Inference takes very little VRAM and training takes lots of VRAM.</p>\n\n<p>I suggest training TensorFlow and PyTorch in separate notebooks. Then make an inference notebook (that only loads models and predicts test) that uses both TensorFlow and PyTorch.</p>",
          "rawMarkdown": "Depends what you're doing. Inference takes very little VRAM and training takes lots of VRAM.\n\nI suggest training TensorFlow and PyTorch in separate notebooks. Then make an inference notebook (that only loads models and predicts test) that uses both TensorFlow and PyTorch."
        },
        {
          "id": 877548,
          "postDate": "2020-06-07T17:30:13.443Z",
          "content": "<p>Another way out is here (self-plug) <a href=\"https://www.kaggle.com/c/google-quest-challenge/discussion/129963\">link</a>!</p>\n\n<p>Use <code>os.system</code>!</p>",
          "rawMarkdown": "Another way out is here (self-plug) [link](https://www.kaggle.com/c/google-quest-challenge/discussion/129963)!\n\nUse `os.system`!"
        }
      ]
    },
    {
      "id": 877293,
      "postDate": "2020-06-07T13:22:29.987Z",
      "content": "<p>Thank you for the insight it really helped me out.</p>",
      "rawMarkdown": "Thank you for the insight it really helped me out.\n"
    },
    {
      "id": 873565,
      "postDate": "2020-06-04T09:16:32.303Z",
      "content": "<p>This was really helpful!</p>",
      "rawMarkdown": "This was really helpful!"
    },
    {
      "id": 873258,
      "postDate": "2020-06-04T01:13:26.963Z",
      "content": "<p>try <code>pip install git+https://github.com/nicolargo/nvidia-ml-py3\n</code> then <code>!pip install git+https://github.com/stas00/ipyexperiments.git</code></p>\n\n<p>then some like</p>\n\n<p>```</p>\n\n<h1>cell 1</h1>\n\n<p>exp1 = IPyExperimentsPytorch() # new experiment</p>\n\n<h1>cell 2</h1>\n\n<p>learn1 = language_model_learner(data_lm, bptt=60, drop_mult=0.25, pretrained_model=URLs.WT103)</p>\n\n<h1>cell 3</h1>\n\n<p>learn1.lr_find()</p>\n\n<h1>cell 4</h1>\n\n<p>del exp1\n```\nsee docs <a href=\"https://github.com/stas00/ipyexperiments/#documentation\">https://github.com/stas00/ipyexperiments/#documentation</a></p>",
      "rawMarkdown": "try `pip install git+https://github.com/nicolargo/nvidia-ml-py3\n` then `!pip install git+https://github.com/stas00/ipyexperiments.git`\n\n\nthen some like\n\n```\n# cell 1\nexp1 = IPyExperimentsPytorch() # new experiment\n# cell 2\nlearn1 = language_model_learner(data_lm, bptt=60, drop_mult=0.25, pretrained_model=URLs.WT103)\n# cell 3\nlearn1.lr_find()\n# cell 4\ndel exp1\n```\nsee docs https://github.com/stas00/ipyexperiments/#documentation"
    },
    {
      "id": 877460,
      "postDate": "2020-06-07T16:04:48.120Z",
      "content": "<p>I got same problem when using TF. I used <code>tf.one_hot</code> to process a numpy array but it exploded the memory from about 800M to 15G, no idea why. Changed to keras's <code>to_categorical</code>, it is total fine. </p>",
      "rawMarkdown": "I got same problem when using TF. I used `tf.one_hot` to process a numpy array but it exploded the memory from about 800M to 15G, no idea why. Changed to keras's `to_categorical`, it is total fine. ",
      "isDeleted": true
    },
    {
      "id": 801668,
      "postDate": "2020-04-08T17:16:11.437Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 874491,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-05T03:41:15.423000",
      "content": "<p>This is tricky business. TensorFlow won't release memory after it has run. Therefore the best way to run both TensorFlow and PyTorch in the same Kaggle notebook is to limit the amount of memory TF can use</p>\n\n<pre><code>import tensorflow as tf\nGIGS = 8\ngpus = tf.config.experimental.list_physical_devices('GPU')\nif gpus:\n  # Restrict TensorFlow memory\n  try:\n    tf.config.experimental.set_virtual_device_configuration(\n        gpus[0],\n        [tf.config.experimental.VirtualDeviceConfiguration(memory_limit=1024*GIGS)])\n    logical_gpus = tf.config.experimental.list_logical_devices('GPU')\n    print(len(gpus), \"Physical GPUs,\", len(logical_gpus), \"Logical GPUs\")\n  except RuntimeError as e:\n     # Virtual devices must be set before GPUs have been initialized\n     print(e)\n</code></pre>\n\n<p>Set the variable <code>GIGS</code> to the amount of memory you wish to restrict TF to. Note that Kaggle notebooks use GPU P100 and have a total of 16GB VRAM</p>",
      "votes": 2,
      "replies": [
        {
          "id": 877454,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-07T16:02:37.663000",
          "content": "<p>Thanks for the info, just wonder, if memory for TF is set small, would it have out of memory error? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 877546,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-07T17:28:46.650000",
          "content": "<p>Depends what you're doing. Inference takes very little VRAM and training takes lots of VRAM.</p>\n\n<p>I suggest training TensorFlow and PyTorch in separate notebooks. Then make an inference notebook (that only loads models and predicts test) that uses both TensorFlow and PyTorch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 877548,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-06-07T17:30:13.443000",
          "content": "<p>Another way out is here (self-plug) <a href=\"https://www.kaggle.com/c/google-quest-challenge/discussion/129963\">link</a>!</p>\n\n<p>Use <code>os.system</code>!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 877293,
      "author_name": "Tamizhiniyan S R",
      "author_url": "",
      "post_date": "2020-06-07T13:22:29.987000",
      "content": "<p>Thank you for the insight it really helped me out.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 873565,
      "author_name": "Adhiraj Bhagawati",
      "author_url": "",
      "post_date": "2020-06-04T09:16:32.303000",
      "content": "<p>This was really helpful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 873258,
      "author_name": "tyoc213",
      "author_url": "",
      "post_date": "2020-06-04T01:13:26.963000",
      "content": "<p>try <code>pip install git+https://github.com/nicolargo/nvidia-ml-py3\n</code> then <code>!pip install git+https://github.com/stas00/ipyexperiments.git</code></p>\n\n<p>then some like</p>\n\n<p>```</p>\n\n<h1>cell 1</h1>\n\n<p>exp1 = IPyExperimentsPytorch() # new experiment</p>\n\n<h1>cell 2</h1>\n\n<p>learn1 = language_model_learner(data_lm, bptt=60, drop_mult=0.25, pretrained_model=URLs.WT103)</p>\n\n<h1>cell 3</h1>\n\n<p>learn1.lr_find()</p>\n\n<h1>cell 4</h1>\n\n<p>del exp1\n```\nsee docs <a href=\"https://github.com/stas00/ipyexperiments/#documentation\">https://github.com/stas00/ipyexperiments/#documentation</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 877460,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-07T16:04:48.120000",
      "content": "<p>I got same problem when using TF. I used <code>tf.one_hot</code> to process a numpy array but it exploded the memory from about 800M to 15G, no idea why. Changed to keras's <code>to_categorical</code>, it is total fine. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 801668,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-08T17:16:11.437000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "801664": "Hi\ni am trying to run two models. one in TF and another in Pytorch . i am facing issue with loading pytorch model after finishing up with tf . is there a better way to cleanup the GPU RAM used by tf ",
    "874491": "This is tricky business. TensorFlow won't release memory after it has run. Therefore the best way to run both TensorFlow and PyTorch in the same Kaggle notebook is to limit the amount of memory TF can use\n\n    import tensorflow as tf\n    GIGS = 8\n    gpus = tf.config.experimental.list_physical_devices('GPU')\n    if gpus:\n      # Restrict TensorFlow memory\n      try:\n        tf.config.experimental.set_virtual_device_configuration(\n            gpus[0],\n            [tf.config.experimental.VirtualDeviceConfiguration(memory_limit=1024*GIGS)])\n        logical_gpus = tf.config.experimental.list_logical_devices('GPU')\n        print(len(gpus), \"Physical GPUs,\", len(logical_gpus), \"Logical GPUs\")\n      except RuntimeError as e:\n         # Virtual devices must be set before GPUs have been initialized\n         print(e)\n\nSet the variable `GIGS` to the amount of memory you wish to restrict TF to. Note that Kaggle notebooks use GPU P100 and have a total of 16GB VRAM",
    "877293": "Thank you for the insight it really helped me out.\n",
    "873565": "This was really helpful!",
    "873258": "try `pip install git+https://github.com/nicolargo/nvidia-ml-py3\n` then `!pip install git+https://github.com/stas00/ipyexperiments.git`\n\n\nthen some like\n\n```\n# cell 1\nexp1 = IPyExperimentsPytorch() # new experiment\n# cell 2\nlearn1 = language_model_learner(data_lm, bptt=60, drop_mult=0.25, pretrained_model=URLs.WT103)\n# cell 3\nlearn1.lr_find()\n# cell 4\ndel exp1\n```\nsee docs https://github.com/stas00/ipyexperiments/#documentation",
    "877460": "I got same problem when using TF. I used `tf.one_hot` to process a numpy array but it exploded the memory from about 800M to 15G, no idea why. Changed to keras's `to_categorical`, it is total fine. ",
    "801668": ""
  }
}