{
  "id": 173178,
  "title": "Memory Leak ??",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173178",
  "author_name": "",
  "post_date": "2020-08-08T07:31:41.182266300Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>During training on GPU I noticed that that my RAM keeps on growing. I am not saving the best weights(which I suspected might be the cause). What's happening? What's casuing the memory growth??</p>\n<p>GPU is operating at 90-100% .<br>\ntf.dataAPI is used for input pipeline using TFRecords.</p>",
  "messages": [
    {
      "id": "962519",
      "postDate": "08/08/2020 07:31:41",
      "content": "<p>During training on GPU I noticed that that my RAM keeps on growing. I am not saving the best weights(which I suspected might be the cause). What's happening? What's casuing the memory growth??</p>\n<p>GPU is operating at 90-100% .<br>\ntf.dataAPI is used for input pipeline using TFRecords.</p>",
      "rawMarkdown": "During training on GPU I noticed that that my RAM keeps on growing. I am not saving the best weights(which I suspected might be the cause). What's happening? What's casuing the memory growth??\n\nGPU is operating at 90-100% .\ntf.dataAPI is used for input pipeline using TFRecords.",
      "votes": null
    },
    {
      "id": "962703",
      "postDate": "08/08/2020 11:10:29",
      "content": "<p>if using tfrecord, why not using TPU?</p>",
      "rawMarkdown": "if using tfrecord, why not using TPU?",
      "votes": null
    },
    {
      "id": "962708",
      "postDate": "08/08/2020 11:18:00",
      "content": "<p>Quota exhausted for TPU.</p>",
      "rawMarkdown": "Quota exhausted for TPU.",
      "votes": null
    },
    {
      "id": "962849",
      "postDate": "08/08/2020 13:39:59",
      "content": "<p>I think TensorFlow has a memory leak. I have seen this problem for years. Is your code crashing? If so, you should describe your error message because it may be something besides the memory leak causing it.</p>",
      "rawMarkdown": "I think TensorFlow has a memory leak. I have seen this problem for years. Is your code crashing? If so, you should describe your error message because it may be something besides the memory leak causing it.",
      "votes": null
    },
    {
      "id": "962921",
      "postDate": "08/08/2020 14:40:57",
      "content": "<p>No no, the code isn't crashing. The RAM just fills up slowly and the kernel automatically stops</p>",
      "rawMarkdown": "No no, the code isn't crashing. The RAM just fills up slowly and the kernel automatically stops",
      "votes": null
    },
    {
      "id": "962991",
      "postDate": "08/08/2020 15:31:20",
      "content": "<p>I'm not sure what setting you're using but don't use tensorflow's mixed precision (float16) on the P100 GPU's, it doesn't work in my experience. That was a main issue I ran into on GPU kernels.</p>",
      "rawMarkdown": "I'm not sure what setting you're using but don't use tensorflow's mixed precision (float16) on the P100 GPU's, it doesn't work in my experience. That was a main issue I ran into on GPU kernels.",
      "votes": null
    },
    {
      "id": "962994",
      "postDate": "08/08/2020 15:31:57",
      "content": "<p>I suggest deleting variables (models, training data, etc) after you are done using them and then call <code>import gc; gc.collect()</code>. Also you can call <code>K.clear_session()</code> at the beginning of for-loops. This will help the memory leak some.</p>",
      "rawMarkdown": "I suggest deleting variables (models, training data, etc) after you are done using them and then call `import gc; gc.collect()`. Also you can call `K.clear_session()` at the beginning of for-loops. This will help the memory leak some.",
      "votes": null
    },
    {
      "id": "965097",
      "postDate": "08/10/2020 11:35:01",
      "content": "<p>Okay, sure!</p>",
      "rawMarkdown": "Okay, sure!",
      "votes": null
    },
    {
      "id": "2147826",
      "postDate": "02/16/2023 22:44:28",
      "content": "<p>Did this get resolved? I am experiencing the same problem</p>",
      "rawMarkdown": "Did this get resolved? I am experiencing the same problem",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 962703,
      "author_name": "yash612",
      "author_url": "",
      "post_date": "08/08/2020 11:10:29",
      "content": "<p>if using tfrecord, why not using TPU?</p>",
      "votes": null,
      "replies": [
        {
          "id": 962708,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "08/08/2020 11:18:00",
          "content": "<p>Quota exhausted for TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 962849,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/08/2020 13:39:59",
      "content": "<p>I think TensorFlow has a memory leak. I have seen this problem for years. Is your code crashing? If so, you should describe your error message because it may be something besides the memory leak causing it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 962921,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "08/08/2020 14:40:57",
          "content": "<p>No no, the code isn't crashing. The RAM just fills up slowly and the kernel automatically stops</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962994,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/08/2020 15:31:57",
          "content": "<p>I suggest deleting variables (models, training data, etc) after you are done using them and then call <code>import gc; gc.collect()</code>. Also you can call <code>K.clear_session()</code> at the beginning of for-loops. This will help the memory leak some.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965097,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "08/10/2020 11:35:01",
          "content": "<p>Okay, sure!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2147826,
      "author_name": "aghalsa",
      "author_url": "",
      "post_date": "02/16/2023 22:44:28",
      "content": "<p>Did this get resolved? I am experiencing the same problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962991,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "08/08/2020 15:31:20",
      "content": "<p>I'm not sure what setting you're using but don't use tensorflow's mixed precision (float16) on the P100 GPU's, it doesn't work in my experience. That was a main issue I ran into on GPU kernels.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "962519": "During training on GPU I noticed that that my RAM keeps on growing. I am not saving the best weights(which I suspected might be the cause). What's happening? What's casuing the memory growth??\n\nGPU is operating at 90-100% .\ntf.dataAPI is used for input pipeline using TFRecords.",
    "962703": "if using tfrecord, why not using TPU?",
    "962708": "Quota exhausted for TPU.",
    "962849": "I think TensorFlow has a memory leak. I have seen this problem for years. Is your code crashing? If so, you should describe your error message because it may be something besides the memory leak causing it.",
    "962921": "No no, the code isn't crashing. The RAM just fills up slowly and the kernel automatically stops",
    "962991": "I'm not sure what setting you're using but don't use tensorflow's mixed precision (float16) on the P100 GPU's, it doesn't work in my experience. That was a main issue I ran into on GPU kernels.",
    "962994": "I suggest deleting variables (models, training data, etc) after you are done using them and then call `import gc; gc.collect()`. Also you can call `K.clear_session()` at the beginning of for-loops. This will help the memory leak some.",
    "965097": "Okay, sure!",
    "2147826": "Did this get resolved? I am experiencing the same problem"
  },
  "source": "meta"
}