{
  "id": 310870,
  "title": "ResourceExhaustedError:  OOM when allocating tensor with shape[2097152,512]",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/310870",
  "author_name": "paulreiners",
  "post_date": "2022-03-03T14:08:05.833000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have a {[simple notebook](<a href=\"https://www.kaggle.com/paulreiners/myherbariumnotebook]}\" target=\"_blank\">https://www.kaggle.com/paulreiners/myherbariumnotebook]}</a>.  I am getting this error</p>\n<blockquote>\n  <p>ResourceExhaustedError:  OOM when allocating tensor with shape[2097152,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc</p>\n</blockquote>\n<p>if I use this model:</p>\n<pre><code>from keras.models import Sequential\nfrom keras.layers import Conv2D\nfrom keras.layers import Dense, Flatten\nfrom keras import layers\n\nmodel = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.Flatten(),\n    layers.Dense(512, activation=\"relu\"),\n    layers.Dense(15501, activation='softmax')\n])\n\nmodel.compile(optimizer='rmsprop', \n              loss=\"categorical_crossentropy\", \n              metrics=[\"accuracy\"])\n</code></pre>\n<p>What am I doing wrong and how do I fix it?</p>",
  "messages": [
    {
      "id": 1713424,
      "postDate": "2022-03-06T02:14:44.097Z",
      "content": "<p>I guess the tip of the iceberg reason why this is happening is because you are flattening a Conv2D layer with a really large shape of (256, 256, 32). You will want to perform some pooling (e.g., MaxPooling, AvgPooling, etc.) to down sample/pool feature maps together.</p>\n<p>To give you an idea of why this model doesn't fit in memory, the above original code the output shape of the flatten layer is 256 * 256 * 32 = 2097152. Then, you'll have to multiply that with the number of Dense nodes (512) and you get an absurd amount of parameters (1073741824 or 1 billion+ params). Very large non-transformer based models like EfficientNet-L2 are \"only\" 480M parameters and that model already is pretty infeasible to train for most people. Note that the top 5 competitors last year (according to the publication about last year's competition) had param counts ranging from 28 million to 226 million after ensembling—no where near 1 billion parameters like the code above is trying to create.</p>\n<p>If you're just trying to create a very simple model, adding stacked pairs of Conv2D and pooling layers like down below, will decrease the amount of params you will have and allow it to fit within memory. </p>\n<pre><code>model = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Flatten(),\n...\n</code></pre>\n<p>Someone more knowledgeable should correct me if I'm wrong (especially if my explanation/terminology isn't correct).</p>",
      "rawMarkdown": "I guess the tip of the iceberg reason why this is happening is because you are flattening a Conv2D layer with a really large shape of (256, 256, 32). You will want to perform some pooling (e.g., MaxPooling, AvgPooling, etc.) to down sample/pool feature maps together.\n\nTo give you an idea of why this model doesn't fit in memory, the above original code the output shape of the flatten layer is 256 * 256 * 32 = 2097152. Then, you'll have to multiply that with the number of Dense nodes (512) and you get an absurd amount of parameters (1073741824 or 1 billion+ params). Very large non-transformer based models like EfficientNet-L2 are \"only\" 480M parameters and that model already is pretty infeasible to train for most people. Note that the top 5 competitors last year (according to the publication about last year's competition) had param counts ranging from 28 million to 226 million after ensembling—no where near 1 billion parameters like the code above is trying to create.\n\nIf you're just trying to create a very simple model, adding stacked pairs of Conv2D and pooling layers like down below, will decrease the amount of params you will have and allow it to fit within memory. \n\n```\nmodel = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Flatten(),\n...\n```\n\nSomeone more knowledgeable should correct me if I'm wrong (especially if my explanation/terminology isn't correct).",
      "votes": 1,
      "replies": [
        {
          "id": 1715071,
          "postDate": "2022-03-07T16:00:12.853Z",
          "content": "<p>Thanks for the great explanation, <a href=\"https://www.kaggle.com/dakilaledesma\" target=\"_blank\">@dakilaledesma</a>. Yes, I believe it is a standard approach to use pooling layers instead of flattening to connect feature extraction layer with the classification layer. One thing to add is that my impression is global average pooling works well for the pooling layer.  For example, <a href=\"https://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242233\" target=\"_blank\">second place solution of herbarium 2021</a> used global average pooling layer after the backbone, followed up fully connected layers: FC (2048x2048)+FC(2048x512)+FC(512x64500). You may also want to consider normalization layers and dropout layers to prevent overfitting. </p>",
          "rawMarkdown": "Thanks for the great explanation, @dakilaledesma. Yes, I believe it is a standard approach to use pooling layers instead of flattening to connect feature extraction layer with the classification layer. One thing to add is that my impression is global average pooling works well for the pooling layer.  For example, [second place solution of herbarium 2021](https://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242233) used global average pooling layer after the backbone, followed up fully connected layers: FC (2048x2048)+FC(2048x512)+FC(512x64500). You may also want to consider normalization layers and dropout layers to prevent overfitting. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1713067,
      "postDate": "2022-03-05T16:32:55.537Z",
      "content": "<p>I recently faced the same issue.<br>\nOOM (Out of memory) error means that you have a problem with memory so try :</p>\n<ol>\n<li>decrease image size </li>\n<li>decrease the batch size </li>\n</ol>",
      "rawMarkdown": "I recently faced the same issue.\nOOM (Out of memory) error means that you have a problem with memory so try :\n1. decrease image size \n2. decrease the batch size \n",
      "votes": 1
    },
    {
      "id": 1710967,
      "postDate": "2022-03-03T14:08:05.833Z",
      "content": "<p>I have a {[simple notebook](<a href=\"https://www.kaggle.com/paulreiners/myherbariumnotebook]}\" target=\"_blank\">https://www.kaggle.com/paulreiners/myherbariumnotebook]}</a>.  I am getting this error</p>\n<blockquote>\n  <p>ResourceExhaustedError:  OOM when allocating tensor with shape[2097152,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc</p>\n</blockquote>\n<p>if I use this model:</p>\n<pre><code>from keras.models import Sequential\nfrom keras.layers import Conv2D\nfrom keras.layers import Dense, Flatten\nfrom keras import layers\n\nmodel = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.Flatten(),\n    layers.Dense(512, activation=\"relu\"),\n    layers.Dense(15501, activation='softmax')\n])\n\nmodel.compile(optimizer='rmsprop', \n              loss=\"categorical_crossentropy\", \n              metrics=[\"accuracy\"])\n</code></pre>\n<p>What am I doing wrong and how do I fix it?</p>",
      "rawMarkdown": "I have a {[simple notebook](https://www.kaggle.com/paulreiners/myherbariumnotebook]}.  I am getting this error\n\n> ResourceExhaustedError:  OOM when allocating tensor with shape[2097152,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc\n\nif I use this model:\n\n```\nfrom keras.models import Sequential\nfrom keras.layers import Conv2D\nfrom keras.layers import Dense, Flatten\nfrom keras import layers\n\nmodel = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.Flatten(),\n    layers.Dense(512, activation=\"relu\"),\n    layers.Dense(15501, activation='softmax')\n])\n\nmodel.compile(optimizer='rmsprop', \n              loss=\"categorical_crossentropy\", \n              metrics=[\"accuracy\"])\n```\n\nWhat am I doing wrong and how do I fix it?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1713424,
      "author_name": "Dax Ledesma",
      "author_url": "",
      "post_date": "2022-03-06T02:14:44.097000",
      "content": "<p>I guess the tip of the iceberg reason why this is happening is because you are flattening a Conv2D layer with a really large shape of (256, 256, 32). You will want to perform some pooling (e.g., MaxPooling, AvgPooling, etc.) to down sample/pool feature maps together.</p>\n<p>To give you an idea of why this model doesn't fit in memory, the above original code the output shape of the flatten layer is 256 * 256 * 32 = 2097152. Then, you'll have to multiply that with the number of Dense nodes (512) and you get an absurd amount of parameters (1073741824 or 1 billion+ params). Very large non-transformer based models like EfficientNet-L2 are \"only\" 480M parameters and that model already is pretty infeasible to train for most people. Note that the top 5 competitors last year (according to the publication about last year's competition) had param counts ranging from 28 million to 226 million after ensembling—no where near 1 billion parameters like the code above is trying to create.</p>\n<p>If you're just trying to create a very simple model, adding stacked pairs of Conv2D and pooling layers like down below, will decrease the amount of params you will have and allow it to fit within memory. </p>\n<pre><code>model = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Flatten(),\n...\n</code></pre>\n<p>Someone more knowledgeable should correct me if I'm wrong (especially if my explanation/terminology isn't correct).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1715071,
          "author_name": "John Park",
          "author_url": "",
          "post_date": "2022-03-07T16:00:12.853000",
          "content": "<p>Thanks for the great explanation, <a href=\"https://www.kaggle.com/dakilaledesma\" target=\"_blank\">@dakilaledesma</a>. Yes, I believe it is a standard approach to use pooling layers instead of flattening to connect feature extraction layer with the classification layer. One thing to add is that my impression is global average pooling works well for the pooling layer.  For example, <a href=\"https://www.kaggle.com/c/herbarium-2021-fgvc8/discussion/242233\" target=\"_blank\">second place solution of herbarium 2021</a> used global average pooling layer after the backbone, followed up fully connected layers: FC (2048x2048)+FC(2048x512)+FC(512x64500). You may also want to consider normalization layers and dropout layers to prevent overfitting. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1713067,
      "author_name": "Hamza",
      "author_url": "",
      "post_date": "2022-03-05T16:32:55.537000",
      "content": "<p>I recently faced the same issue.<br>\nOOM (Out of memory) error means that you have a problem with memory so try :</p>\n<ol>\n<li>decrease image size </li>\n<li>decrease the batch size </li>\n</ol>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1713424": "I guess the tip of the iceberg reason why this is happening is because you are flattening a Conv2D layer with a really large shape of (256, 256, 32). You will want to perform some pooling (e.g., MaxPooling, AvgPooling, etc.) to down sample/pool feature maps together.\n\nTo give you an idea of why this model doesn't fit in memory, the above original code the output shape of the flatten layer is 256 * 256 * 32 = 2097152. Then, you'll have to multiply that with the number of Dense nodes (512) and you get an absurd amount of parameters (1073741824 or 1 billion+ params). Very large non-transformer based models like EfficientNet-L2 are \"only\" 480M parameters and that model already is pretty infeasible to train for most people. Note that the top 5 competitors last year (according to the publication about last year's competition) had param counts ranging from 28 million to 226 million after ensembling—no where near 1 billion parameters like the code above is trying to create.\n\nIf you're just trying to create a very simple model, adding stacked pairs of Conv2D and pooling layers like down below, will decrease the amount of params you will have and allow it to fit within memory. \n\n```\nmodel = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Conv2D(32, (3, 3), padding='same'),\n    layers.MaxPooling2D(),\n    layers.Flatten(),\n...\n```\n\nSomeone more knowledgeable should correct me if I'm wrong (especially if my explanation/terminology isn't correct).",
    "1713067": "I recently faced the same issue.\nOOM (Out of memory) error means that you have a problem with memory so try :\n1. decrease image size \n2. decrease the batch size \n",
    "1710967": "I have a {[simple notebook](https://www.kaggle.com/paulreiners/myherbariumnotebook]}.  I am getting this error\n\n> ResourceExhaustedError:  OOM when allocating tensor with shape[2097152,512] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc\n\nif I use this model:\n\n```\nfrom keras.models import Sequential\nfrom keras.layers import Conv2D\nfrom keras.layers import Dense, Flatten\nfrom keras import layers\n\nmodel = Sequential([\n    layers.Conv2D(32, (3, 3), padding='same',\n                 input_shape=(256, 256, 3)),\n    layers.Flatten(),\n    layers.Dense(512, activation=\"relu\"),\n    layers.Dense(15501, activation='softmax')\n])\n\nmodel.compile(optimizer='rmsprop', \n              loss=\"categorical_crossentropy\", \n              metrics=[\"accuracy\"])\n```\n\nWhat am I doing wrong and how do I fix it?"
  }
}