{
  "id": 217014,
  "title": "Help - Memory Leak related to tf data",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217014",
  "author_name": "",
  "post_date": "2021-02-04T22:23:37.563449100Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I don't know if this is the place for discussing about help or not. If i'm doing wrong, kindly notify me. The issue is related to my <a href=\"https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec\" target=\"_blank\">notebook</a>. In the last defined function, the CV models, at each iteration of the cv, the program seems to not be able to garbage collect the loaded data of the dataset and eventually crashes. The memory rises each time a new model is made and a new dataset is processed. Can anyone eyeball the problem? Maybe i'm doing a known mistake or something.</p>",
  "messages": [
    {
      "id": "1186557",
      "postDate": "02/04/2021 22:23:37",
      "content": "<p>I don't know if this is the place for discussing about help or not. If i'm doing wrong, kindly notify me. The issue is related to my <a href=\"https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec\" target=\"_blank\">notebook</a>. In the last defined function, the CV models, at each iteration of the cv, the program seems to not be able to garbage collect the loaded data of the dataset and eventually crashes. The memory rises each time a new model is made and a new dataset is processed. Can anyone eyeball the problem? Maybe i'm doing a known mistake or something.</p>",
      "rawMarkdown": "I don't know if this is the place for discussing about help or not. If i'm doing wrong, kindly notify me. The issue is related to my [notebook](https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec). In the last defined function, the CV models, at each iteration of the cv, the program seems to not be able to garbage collect the loaded data of the dataset and eventually crashes. The memory rises each time a new model is made and a new dataset is processed. Can anyone eyeball the problem? Maybe i'm doing a known mistake or something.",
      "votes": null
    },
    {
      "id": "1186567",
      "postDate": "02/04/2021 22:38:53",
      "content": "<p>Not sure if it solves your problem… but maybe removing the 'cache' (and/or the data_augmentation prior to this which I am not sure if it is called) solves it.</p>",
      "rawMarkdown": "Not sure if it solves your problem... but maybe removing the 'cache' (and/or the data_augmentation prior to this which I am not sure if it is called) solves it.",
      "votes": null
    },
    {
      "id": "1186591",
      "postDate": "02/04/2021 23:10:24",
      "content": "<p>I have not yet looked at your code but believe <a href=\"https://www.kaggle.com/morodertobias\" target=\"_blank\">tmoroder</a> is likely correct.   In a past competition I was getting a CPU RAM memory error - since I had 256GB it came as a bit of a surprise.  </p>\n<p>The memory was building up with each epoch - I was trying to use cache to speed up the model fit.</p>\n<p>It may be OK to cache the validation set in a tabular data model, but cache is likely a mistake for vision models.</p>\n<p>Update - after I looked :)</p>\n<p>You have this in at least two places - kill them both and try.</p>\n<p><code>ds = ds.cache()</code></p>",
      "rawMarkdown": "I have not yet looked at your code but believe [tmoroder](https://www.kaggle.com/morodertobias) is likely correct.   In a past competition I was getting a CPU RAM memory error - since I had 256GB it came as a bit of a surprise.  \n\nThe memory was building up with each epoch - I was trying to use cache to speed up the model fit.\n\nIt may be OK to cache the validation set in a tabular data model, but cache is likely a mistake for vision models.\n\nUpdate - after I looked :)\n\nYou have this in at least two places - kill them both and try.\n\n`ds = ds.cache()`",
      "votes": null
    },
    {
      "id": "1197829",
      "postDate": "02/12/2021 13:21:46",
      "content": "<p>Hello, sorry for my tardiness, was messing around with the code. I've deleted all the cache lines in the training, and its still leaking memory. Even though i'm deleting the models and the dataset, and using the garbage collection module to clean it up, the problem still persists. I honestly tried everything I could on the internet relevant to that tensorflow version and nothing worked. The link for the new version can be found <a href=\"https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Hello, sorry for my tardiness, was messing around with the code. I've deleted all the cache lines in the training, and its still leaking memory. Even though i'm deleting the models and the dataset, and using the garbage collection module to clean it up, the problem still persists. I honestly tried everything I could on the internet relevant to that tensorflow version and nothing worked. The link for the new version can be found [here](https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec).",
      "votes": null
    },
    {
      "id": "1198081",
      "postDate": "02/12/2021 17:34:33",
      "content": "<p>Not sure if these will solve your problems, but I can only give those comments:</p>\n<ul>\n<li>With EfficientNetB5, try batch size 128 instead of 256.</li>\n<li>Clear session before you create the next model, so call <code>tf.keras.backend.clear_session()</code> before your <code>with strategy.scope()</code>.</li>\n<li>Put the loss definition also into the <code>with strategy.scope()</code> block. Not sure, if you current code would mean that the loss function is placed on the CPU.</li>\n</ul>\n<p>If this all fails, my last suggestion would be to put all code that you execute in the fold loop into a separate function and call it in the loop… in this way the garbage collection might be more direct.</p>",
      "rawMarkdown": "Not sure if these will solve your problems, but I can only give those comments:\n\n- With EfficientNetB5, try batch size 128 instead of 256.\n- Clear session before you create the next model, so call ``tf.keras.backend.clear_session()`` before your ``with strategy.scope()``.\n- Put the loss definition also into the ``with strategy.scope()`` block. Not sure, if you current code would mean that the loss function is placed on the CPU.\n\nIf this all fails, my last suggestion would be to put all code that you execute in the fold loop into a separate function and call it in the loop... in this way the garbage collection might be more direct.",
      "votes": null
    },
    {
      "id": "1221320",
      "postDate": "02/28/2021 22:07:04",
      "content": "<p>No luck with that unfortunately.</p>",
      "rawMarkdown": "No luck with that unfortunately.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1186567,
      "author_name": "morodertobias",
      "author_url": "",
      "post_date": "02/04/2021 22:38:53",
      "content": "<p>Not sure if it solves your problem… but maybe removing the 'cache' (and/or the data_augmentation prior to this which I am not sure if it is called) solves it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186591,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "02/04/2021 23:10:24",
          "content": "<p>I have not yet looked at your code but believe <a href=\"https://www.kaggle.com/morodertobias\" target=\"_blank\">tmoroder</a> is likely correct.   In a past competition I was getting a CPU RAM memory error - since I had 256GB it came as a bit of a surprise.  </p>\n<p>The memory was building up with each epoch - I was trying to use cache to speed up the model fit.</p>\n<p>It may be OK to cache the validation set in a tabular data model, but cache is likely a mistake for vision models.</p>\n<p>Update - after I looked :)</p>\n<p>You have this in at least two places - kill them both and try.</p>\n<p><code>ds = ds.cache()</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1197829,
          "author_name": "elysium1436",
          "author_url": "",
          "post_date": "02/12/2021 13:21:46",
          "content": "<p>Hello, sorry for my tardiness, was messing around with the code. I've deleted all the cache lines in the training, and its still leaking memory. Even though i'm deleting the models and the dataset, and using the garbage collection module to clean it up, the problem still persists. I honestly tried everything I could on the internet relevant to that tensorflow version and nothing worked. The link for the new version can be found <a href=\"https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1198081,
          "author_name": "morodertobias",
          "author_url": "",
          "post_date": "02/12/2021 17:34:33",
          "content": "<p>Not sure if these will solve your problems, but I can only give those comments:</p>\n<ul>\n<li>With EfficientNetB5, try batch size 128 instead of 256.</li>\n<li>Clear session before you create the next model, so call <code>tf.keras.backend.clear_session()</code> before your <code>with strategy.scope()</code>.</li>\n<li>Put the loss definition also into the <code>with strategy.scope()</code> block. Not sure, if you current code would mean that the loss function is placed on the CPU.</li>\n</ul>\n<p>If this all fails, my last suggestion would be to put all code that you execute in the fold loop into a separate function and call it in the loop… in this way the garbage collection might be more direct.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221320,
          "author_name": "elysium1436",
          "author_url": "",
          "post_date": "02/28/2021 22:07:04",
          "content": "<p>No luck with that unfortunately.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1186557": "I don't know if this is the place for discussing about help or not. If i'm doing wrong, kindly notify me. The issue is related to my [notebook](https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec). In the last defined function, the CV models, at each iteration of the cv, the program seems to not be able to garbage collect the loaded data of the dataset and eventually crashes. The memory rises each time a new model is made and a new dataset is processed. Can anyone eyeball the problem? Maybe i'm doing a known mistake or something.",
    "1186567": "Not sure if it solves your problem... but maybe removing the 'cache' (and/or the data_augmentation prior to this which I am not sure if it is called) solves it.",
    "1186591": "I have not yet looked at your code but believe [tmoroder](https://www.kaggle.com/morodertobias) is likely correct.   In a past competition I was getting a CPU RAM memory error - since I had 256GB it came as a bit of a surprise.  \n\nThe memory was building up with each epoch - I was trying to use cache to speed up the model fit.\n\nIt may be OK to cache the validation set in a tabular data model, but cache is likely a mistake for vision models.\n\nUpdate - after I looked :)\n\nYou have this in at least two places - kill them both and try.\n\n`ds = ds.cache()`",
    "1197829": "Hello, sorry for my tardiness, was messing around with the code. I've deleted all the cache lines in the training, and its still leaking memory. Even though i'm deleting the models and the dataset, and using the garbage collection module to clean it up, the problem still persists. I honestly tried everything I could on the internet relevant to that tensorflow version and nothing worked. The link for the new version can be found [here](https://www.kaggle.com/elysium1436/tpu-bitempered-radam-cutmix-cv5-tf-data-tfrec).",
    "1198081": "Not sure if these will solve your problems, but I can only give those comments:\n\n- With EfficientNetB5, try batch size 128 instead of 256.\n- Clear session before you create the next model, so call ``tf.keras.backend.clear_session()`` before your ``with strategy.scope()``.\n- Put the loss definition also into the ``with strategy.scope()`` block. Not sure, if you current code would mean that the loss function is placed on the CPU.\n\nIf this all fails, my last suggestion would be to put all code that you execute in the fold loop into a separate function and call it in the loop... in this way the garbage collection might be more direct.",
    "1221320": "No luck with that unfortunately."
  },
  "source": "meta"
}