{
  "id": 205818,
  "title": "Resource exhausted: OOM resource allocation error",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/205818",
  "author_name": "",
  "post_date": "2020-12-22T02:18:27.967482500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This error happens when I am training a large network on GPU with a relatively large batch size and image size or for some other reasons that I am not aware of. How can I overcome this? Is switching to TPU the only possible solution? Please advise :)</p>",
  "messages": [
    {
      "id": "1121893",
      "postDate": "12/22/2020 02:18:27",
      "content": "<p>This error happens when I am training a large network on GPU with a relatively large batch size and image size or for some other reasons that I am not aware of. How can I overcome this? Is switching to TPU the only possible solution? Please advise :)</p>",
      "rawMarkdown": "This error happens when I am training a large network on GPU with a relatively large batch size and image size or for some other reasons that I am not aware of. How can I overcome this? Is switching to TPU the only possible solution? Please advise :)",
      "votes": null
    },
    {
      "id": "1123372",
      "postDate": "12/23/2020 07:03:05",
      "content": "<p>Have you tried mixed-precision? I can try <code>EfficientNetB2</code> with Image size <code>260</code> and batch size <code>128</code> with that. </p>\n<pre><code>policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n</code></pre>",
      "rawMarkdown": "Have you tried mixed-precision? I can try `EfficientNetB2` with Image size `260` and batch size `128` with that. \n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n```",
      "votes": null
    },
    {
      "id": "1123882",
      "postDate": "12/23/2020 14:59:30",
      "content": "<p>Reducing the batch size would be the standard way of resolving the issue. Obviously, at some point the batch size becomes pretty small, at which point you'd want to accumulate gradients for a few batches before backpropagating. Costlier options include going outside of Kaggle to train with multiple GPUs or with a single larger GPU (e.g. the <a href=\"https://www.nvidia.com/en-us/data-center/a100/#specifications\" target=\"_blank\">A100</a> has 40 GB of memory).</p>",
      "rawMarkdown": "Reducing the batch size would be the standard way of resolving the issue. Obviously, at some point the batch size becomes pretty small, at which point you'd want to accumulate gradients for a few batches before backpropagating. Costlier options include going outside of Kaggle to train with multiple GPUs or with a single larger GPU (e.g. the [A100](https://www.nvidia.com/en-us/data-center/a100/#specifications) has 40 GB of memory).",
      "votes": null
    },
    {
      "id": "1124362",
      "postDate": "12/23/2020 21:20:37",
      "content": "<p>Thank you for your reply! I did not try this but I will look into it! Thank you!</p>",
      "rawMarkdown": "Thank you for your reply! I did not try this but I will look into it! Thank you!",
      "votes": null
    },
    {
      "id": "1124365",
      "postDate": "12/23/2020 21:22:22",
      "content": "<p>Thank you for your reply and suggestions</p>",
      "rawMarkdown": "Thank you for your reply and suggestions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123372,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "12/23/2020 07:03:05",
      "content": "<p>Have you tried mixed-precision? I can try <code>EfficientNetB2</code> with Image size <code>260</code> and batch size <code>128</code> with that. </p>\n<pre><code>policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1124362,
          "author_name": "sinamhd9",
          "author_url": "",
          "post_date": "12/23/2020 21:20:37",
          "content": "<p>Thank you for your reply! I did not try this but I will look into it! Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123882,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/23/2020 14:59:30",
      "content": "<p>Reducing the batch size would be the standard way of resolving the issue. Obviously, at some point the batch size becomes pretty small, at which point you'd want to accumulate gradients for a few batches before backpropagating. Costlier options include going outside of Kaggle to train with multiple GPUs or with a single larger GPU (e.g. the <a href=\"https://www.nvidia.com/en-us/data-center/a100/#specifications\" target=\"_blank\">A100</a> has 40 GB of memory).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1124365,
          "author_name": "sinamhd9",
          "author_url": "",
          "post_date": "12/23/2020 21:22:22",
          "content": "<p>Thank you for your reply and suggestions</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1121893": "This error happens when I am training a large network on GPU with a relatively large batch size and image size or for some other reasons that I am not aware of. How can I overcome this? Is switching to TPU the only possible solution? Please advise :)",
    "1123372": "Have you tried mixed-precision? I can try `EfficientNetB2` with Image size `260` and batch size `128` with that. \n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\ntf.keras.mixed_precision.experimental.set_policy(policy)\n```",
    "1123882": "Reducing the batch size would be the standard way of resolving the issue. Obviously, at some point the batch size becomes pretty small, at which point you'd want to accumulate gradients for a few batches before backpropagating. Costlier options include going outside of Kaggle to train with multiple GPUs or with a single larger GPU (e.g. the [A100](https://www.nvidia.com/en-us/data-center/a100/#specifications) has 40 GB of memory).",
    "1124362": "Thank you for your reply! I did not try this but I will look into it! Thank you!",
    "1124365": "Thank you for your reply and suggestions"
  },
  "source": "meta"
}