{
  "id": 131895,
  "title": "local VM (GCP) memory gets 100% usage during hyperparmeter searching",
  "url": "/competitions/flower-classification-with-tpus/discussion/131895",
  "author_name": "",
  "post_date": "2020-02-22T12:54:26.936457500Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have a code like:</p>\n\n<pre><code>import gc\ngc.enable()\n\nfor hyper_parametes in hyperparameter_spaces:\n\n    with strategy.scope:\n        model = Flower_Classifier(...)\n        model.compile(...)\n    model.fit(...)\n    del model\n    gc.collect()\n</code></pre>\n\n<p>I am running it on GCP, and after ~10 iterations, the local VM uses 100% of the memory, and the process is killed. Is any further step to free both the local / remote TPU memory??</p>",
  "messages": [
    {
      "id": "753601",
      "postDate": "02/22/2020 12:54:26",
      "content": "<p>I have a code like:</p>\n\n<pre><code>import gc\ngc.enable()\n\nfor hyper_parametes in hyperparameter_spaces:\n\n    with strategy.scope:\n        model = Flower_Classifier(...)\n        model.compile(...)\n    model.fit(...)\n    del model\n    gc.collect()\n</code></pre>\n\n<p>I am running it on GCP, and after ~10 iterations, the local VM uses 100% of the memory, and the process is killed. Is any further step to free both the local / remote TPU memory??</p>",
      "rawMarkdown": "I have a code like:\n\n    import gc\n    gc.enable()\n\n    for hyper_parametes in hyperparameter_spaces:\n \n        with strategy.scope:\n            model = Flower_Classifier(...)\n            model.compile(...)\n        model.fit(...)\n        del model\n        gc.collect()\n\nI am running it on GCP, and after ~10 iterations, the local VM uses 100% of the memory, and the process is killed. Is any further step to free both the local / remote TPU memory??",
      "votes": null
    },
    {
      "id": "755496",
      "postDate": "02/24/2020 20:23:49",
      "content": "<p>Probably a typing mistake but it's <code>strategy.scope()</code> with parentheses.\nIt looks like the model is not being released locally. Probably a stray pointer on the model somewhere.</p>\n\n<p>Also, if you are on GCP, I recommend you use GCP's HP Tuning service: <a href=\"https://cloud.google.com/ai-platform/training/docs/hyperparameter-tuning-overview\">https://cloud.google.com/ai-platform/training/docs/hyperparameter-tuning-overview</a>\nThe service creates a new machine for each run so you will not have any issues with memory. Also, it shuts down everything at the end so no overcharges either.</p>",
      "rawMarkdown": "Probably a typing mistake but it's `strategy.scope()` with parentheses.\nIt looks like the model is not being released locally. Probably a stray pointer on the model somewhere.\n\nAlso, if you are on GCP, I recommend you use GCP's HP Tuning service: https://cloud.google.com/ai-platform/training/docs/hyperparameter-tuning-overview\nThe service creates a new machine for each run so you will not have any issues with memory. Also, it shuts down everything at the end so no overcharges either.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 755496,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/24/2020 20:23:49",
      "content": "<p>Probably a typing mistake but it's <code>strategy.scope()</code> with parentheses.\nIt looks like the model is not being released locally. Probably a stray pointer on the model somewhere.</p>\n\n<p>Also, if you are on GCP, I recommend you use GCP's HP Tuning service: <a href=\"https://cloud.google.com/ai-platform/training/docs/hyperparameter-tuning-overview\">https://cloud.google.com/ai-platform/training/docs/hyperparameter-tuning-overview</a>\nThe service creates a new machine for each run so you will not have any issues with memory. Also, it shuts down everything at the end so no overcharges either.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "753601": "I have a code like:\n\n    import gc\n    gc.enable()\n\n    for hyper_parametes in hyperparameter_spaces:\n \n        with strategy.scope:\n            model = Flower_Classifier(...)\n            model.compile(...)\n        model.fit(...)\n        del model\n        gc.collect()\n\nI am running it on GCP, and after ~10 iterations, the local VM uses 100% of the memory, and the process is killed. Is any further step to free both the local / remote TPU memory??",
    "755496": "Probably a typing mistake but it's `strategy.scope()` with parentheses.\nIt looks like the model is not being released locally. Probably a stray pointer on the model somewhere.\n\nAlso, if you are on GCP, I recommend you use GCP's HP Tuning service: https://cloud.google.com/ai-platform/training/docs/hyperparameter-tuning-overview\nThe service creates a new machine for each run so you will not have any issues with memory. Also, it shuts down everything at the end so no overcharges either."
  },
  "source": "meta"
}