{
  "id": 242697,
  "title": "Unexpected CUDA Memory error",
  "url": "/competitions/seti-breakthrough-listen/discussion/242697",
  "author_name": "Bhuvan Sachdeva",
  "post_date": "2021-05-30T09:51:53.682000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>CUDA is running out of memory unexpectedly,</p>\n<p>In one case, I trained a custom model with a batch size of 1024.<br>\nThat model had 5 conv layers and 4 linear layers.</p>\n<p>Then, if I replace my conv layers with a pretrained resnet, even a batch size of 128 causes the program to run out of memory. I tried with 32 and it is training but this shouldn't happen as resnet model is not taking that much memory.<br>\nIs there any particular reason behind this? And any solutions?</p>",
  "messages": [
    {
      "id": 1328998,
      "postDate": "2021-05-30T18:26:56.487Z",
      "content": "<p>Not clear if your using kaggle or a local PC.  On local I often run into memory errors if I have not stopped older scripts running.  On occasion I can lock things up enough that a reboot needed before the error goes away.</p>\n<p>You custom model does not seem that big so 1024 batch makes sense.  For this competition I have found that accuracy is sensitive to batch size were too big and too small batch sizes are less accurate.</p>\n<p>Most of the models I have created in this completion are happiest at batch size of 16 for most of the pre-trained models (8GB or 11GB GPU size on my machines).  Unless I use a very small image size 128 as batch would be too big for me on this dataset.</p>\n<p>So - I don't think anything wrong - except for you expectations :)</p>",
      "rawMarkdown": "Not clear if your using kaggle or a local PC.  On local I often run into memory errors if I have not stopped older scripts running.  On occasion I can lock things up enough that a reboot needed before the error goes away.\n\nYou custom model does not seem that big so 1024 batch makes sense.  For this competition I have found that accuracy is sensitive to batch size were too big and too small batch sizes are less accurate.\n\nMost of the models I have created in this completion are happiest at batch size of 16 for most of the pre-trained models (8GB or 11GB GPU size on my machines).  Unless I use a very small image size 128 as batch would be too big for me on this dataset.\n\nSo - I don't think anything wrong - except for you expectations :)",
      "votes": 1,
      "replies": [
        {
          "id": 1329066,
          "postDate": "2021-05-30T19:49:00.977Z",
          "content": "<p>I get your point and that is fine but, I don't get why there is such a huge difference between the two scenarios. the possible batch size decreases 10 fold, Resnet50 isn't that big a model.</p>",
          "rawMarkdown": "I get your point and that is fine but, I don't get why there is such a huge difference between the two scenarios. the possible batch size decreases 10 fold, Resnet50 isn't that big a model."
        }
      ]
    },
    {
      "id": 1328974,
      "postDate": "2021-05-30T18:01:30.110Z",
      "content": "<p>Try by restarting the kernel. It may be happened because resnet have skip connections. Try to use small size images.</p>",
      "rawMarkdown": "Try by restarting the kernel. It may be happened because resnet have skip connections. Try to use small size images.",
      "votes": 1,
      "replies": [
        {
          "id": 1329063,
          "postDate": "2021-05-30T19:47:32.133Z",
          "content": "<p>Tried that, didn't help much.<br>\nThanks though.</p>",
          "rawMarkdown": "Tried that, didn't help much.\nThanks though."
        }
      ]
    },
    {
      "id": 1328506,
      "postDate": "2021-05-30T10:02:51.937Z",
      "content": "<p>can you share your code?<br>\nif related to model size : use small batch size + reduce wide of channel + mixed precision<br>\nbut it can be related to dataset or GC </p>",
      "rawMarkdown": "can you share your code?\nif related to model size : use small batch size + reduce wide of channel + mixed precision\nbut it can be related to dataset or GC ",
      "replies": [
        {
          "id": 1328525,
          "postDate": "2021-05-30T10:25:36.477Z",
          "content": "<p>These are the models, dataset is the same one in this competition.</p>\n<pre><code>class Extractor(nn.Module):\n    def __init__(self):\n        super(Extractor, self).__init__()\n\n        self.res_net = models.resnet50(pretrained = True)\n\n        for param in self.res_net.parameters():\n            param.require_grad = False\n\n\n    def forward(self, x):\n        x = self.res_net(x)\n        return x\n\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n\n        self.fc1 = nn.Linear(1000, 256)\n        self.fc2 = nn.Linear(256, 64)\n        self.fc3 = nn.Linear(64, 8)\n        self.fc4 = nn.Linear(8, 1)\n\n    def forward(self, x):   \n\n        x = F.leaky_relu(self.fc1(x))\n        x = F.leaky_relu(self.fc2(x))\n        x = F.leaky_relu(self.fc3(x))\n        x = self.fc4(x)\n\n        return x\n</code></pre>\n<p>This model works with the batch size 1024</p>\n<pre><code>class Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.conv1 = nn.Conv2d(6, 16, 5, 2)\n        self.conv2 = nn.Conv2d(16, 32, 5, 2)\n        self.conv3 = nn.Conv2d(32, 64, 5, 2)\n        self.conv4 = nn.Conv2d(64, 32, 3, 2)\n\n        self.fc1 = nn.Linear(32*15*14, 1024)\n        self.fc2 = nn.Linear(1024, 128)\n        self.fc3 = nn.Linear(128, 8)\n        self.fc4 = nn.Linear(8, 1)\n\n    def forward(self, x):   \n        x = F.leaky_relu(self.conv1(x))\n        x = F.leaky_relu(self.conv2(x))\n        x = F.leaky_relu(self.conv3(x))\n        x = F.leaky_relu(self.conv4(x))\n        x = x.view(x.shape[0], -1)\n        x = F.leaky_relu(self.fc1(x))\n        x = F.leaky_relu(self.fc2(x))\n        x = F.leaky_relu(self.fc3(x))\n        x = self.fc4(x)\n\n        return x\n</code></pre>",
          "rawMarkdown": "These are the models, dataset is the same one in this competition.\n\n```\nclass Extractor(nn.Module):\n    def __init__(self):\n        super(Extractor, self).__init__()\n        \n        self.res_net = models.resnet50(pretrained = True)\n        \n        for param in self.res_net.parameters():\n            param.require_grad = False\n            \n            \n    def forward(self, x):\n        x = self.res_net(x)\n        return x\n\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        \n        self.fc1 = nn.Linear(1000, 256)\n        self.fc2 = nn.Linear(256, 64)\n        self.fc3 = nn.Linear(64, 8)\n        self.fc4 = nn.Linear(8, 1)\n        \n    def forward(self, x):   \n        \n        x = F.leaky_relu(self.fc1(x))\n        x = F.leaky_relu(self.fc2(x))\n        x = F.leaky_relu(self.fc3(x))\n        x = self.fc4(x)\n        \n        return x\n```\n\nThis model works with the batch size 1024\n\n```\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.conv1 = nn.Conv2d(6, 16, 5, 2)\n        self.conv2 = nn.Conv2d(16, 32, 5, 2)\n        self.conv3 = nn.Conv2d(32, 64, 5, 2)\n        self.conv4 = nn.Conv2d(64, 32, 3, 2)\n        \n        self.fc1 = nn.Linear(32*15*14, 1024)\n        self.fc2 = nn.Linear(1024, 128)\n        self.fc3 = nn.Linear(128, 8)\n        self.fc4 = nn.Linear(8, 1)\n        \n    def forward(self, x):   \n        x = F.leaky_relu(self.conv1(x))\n        x = F.leaky_relu(self.conv2(x))\n        x = F.leaky_relu(self.conv3(x))\n        x = F.leaky_relu(self.conv4(x))\n        x = x.view(x.shape[0], -1)\n        x = F.leaky_relu(self.fc1(x))\n        x = F.leaky_relu(self.fc2(x))\n        x = F.leaky_relu(self.fc3(x))\n        x = self.fc4(x)\n        \n        return x\n```"
        }
      ]
    },
    {
      "id": 1328494,
      "postDate": "2021-05-30T09:51:53.683Z",
      "content": "<p>CUDA is running out of memory unexpectedly,</p>\n<p>In one case, I trained a custom model with a batch size of 1024.<br>\nThat model had 5 conv layers and 4 linear layers.</p>\n<p>Then, if I replace my conv layers with a pretrained resnet, even a batch size of 128 causes the program to run out of memory. I tried with 32 and it is training but this shouldn't happen as resnet model is not taking that much memory.<br>\nIs there any particular reason behind this? And any solutions?</p>",
      "rawMarkdown": "CUDA is running out of memory unexpectedly,\n\nIn one case, I trained a custom model with a batch size of 1024.\nThat model had 5 conv layers and 4 linear layers.\n\nThen, if I replace my conv layers with a pretrained resnet, even a batch size of 128 causes the program to run out of memory. I tried with 32 and it is training but this shouldn't happen as resnet model is not taking that much memory.\nIs there any particular reason behind this? And any solutions?"
    },
    {
      "id": 1328524,
      "postDate": "2021-05-30T10:24:55.430Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1328998,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-05-30T18:26:56.487000",
      "content": "<p>Not clear if your using kaggle or a local PC.  On local I often run into memory errors if I have not stopped older scripts running.  On occasion I can lock things up enough that a reboot needed before the error goes away.</p>\n<p>You custom model does not seem that big so 1024 batch makes sense.  For this competition I have found that accuracy is sensitive to batch size were too big and too small batch sizes are less accurate.</p>\n<p>Most of the models I have created in this completion are happiest at batch size of 16 for most of the pre-trained models (8GB or 11GB GPU size on my machines).  Unless I use a very small image size 128 as batch would be too big for me on this dataset.</p>\n<p>So - I don't think anything wrong - except for you expectations :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1329066,
          "author_name": "Bhuvan Sachdeva",
          "author_url": "",
          "post_date": "2021-05-30T19:49:00.977000",
          "content": "<p>I get your point and that is fine but, I don't get why there is such a huge difference between the two scenarios. the possible batch size decreases 10 fold, Resnet50 isn't that big a model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1328974,
      "author_name": "Aman Deep Gupta",
      "author_url": "",
      "post_date": "2021-05-30T18:01:30.110000",
      "content": "<p>Try by restarting the kernel. It may be happened because resnet have skip connections. Try to use small size images.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1329063,
          "author_name": "Bhuvan Sachdeva",
          "author_url": "",
          "post_date": "2021-05-30T19:47:32.133000",
          "content": "<p>Tried that, didn't help much.<br>\nThanks though.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1328506,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-05-30T10:02:51.937000",
      "content": "<p>can you share your code?<br>\nif related to model size : use small batch size + reduce wide of channel + mixed precision<br>\nbut it can be related to dataset or GC </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1328525,
          "author_name": "Bhuvan Sachdeva",
          "author_url": "",
          "post_date": "2021-05-30T10:25:36.477000",
          "content": "<p>These are the models, dataset is the same one in this competition.</p>\n<pre><code>class Extractor(nn.Module):\n    def __init__(self):\n        super(Extractor, self).__init__()\n\n        self.res_net = models.resnet50(pretrained = True)\n\n        for param in self.res_net.parameters():\n            param.require_grad = False\n\n\n    def forward(self, x):\n        x = self.res_net(x)\n        return x\n\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n\n        self.fc1 = nn.Linear(1000, 256)\n        self.fc2 = nn.Linear(256, 64)\n        self.fc3 = nn.Linear(64, 8)\n        self.fc4 = nn.Linear(8, 1)\n\n    def forward(self, x):   \n\n        x = F.leaky_relu(self.fc1(x))\n        x = F.leaky_relu(self.fc2(x))\n        x = F.leaky_relu(self.fc3(x))\n        x = self.fc4(x)\n\n        return x\n</code></pre>\n<p>This model works with the batch size 1024</p>\n<pre><code>class Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.conv1 = nn.Conv2d(6, 16, 5, 2)\n        self.conv2 = nn.Conv2d(16, 32, 5, 2)\n        self.conv3 = nn.Conv2d(32, 64, 5, 2)\n        self.conv4 = nn.Conv2d(64, 32, 3, 2)\n\n        self.fc1 = nn.Linear(32*15*14, 1024)\n        self.fc2 = nn.Linear(1024, 128)\n        self.fc3 = nn.Linear(128, 8)\n        self.fc4 = nn.Linear(8, 1)\n\n    def forward(self, x):   \n        x = F.leaky_relu(self.conv1(x))\n        x = F.leaky_relu(self.conv2(x))\n        x = F.leaky_relu(self.conv3(x))\n        x = F.leaky_relu(self.conv4(x))\n        x = x.view(x.shape[0], -1)\n        x = F.leaky_relu(self.fc1(x))\n        x = F.leaky_relu(self.fc2(x))\n        x = F.leaky_relu(self.fc3(x))\n        x = self.fc4(x)\n\n        return x\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1328524,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-30T10:24:55.430000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1328998": "Not clear if your using kaggle or a local PC.  On local I often run into memory errors if I have not stopped older scripts running.  On occasion I can lock things up enough that a reboot needed before the error goes away.\n\nYou custom model does not seem that big so 1024 batch makes sense.  For this competition I have found that accuracy is sensitive to batch size were too big and too small batch sizes are less accurate.\n\nMost of the models I have created in this completion are happiest at batch size of 16 for most of the pre-trained models (8GB or 11GB GPU size on my machines).  Unless I use a very small image size 128 as batch would be too big for me on this dataset.\n\nSo - I don't think anything wrong - except for you expectations :)",
    "1328974": "Try by restarting the kernel. It may be happened because resnet have skip connections. Try to use small size images.",
    "1328506": "can you share your code?\nif related to model size : use small batch size + reduce wide of channel + mixed precision\nbut it can be related to dataset or GC ",
    "1328494": "CUDA is running out of memory unexpectedly,\n\nIn one case, I trained a custom model with a batch size of 1024.\nThat model had 5 conv layers and 4 linear layers.\n\nThen, if I replace my conv layers with a pretrained resnet, even a batch size of 128 causes the program to run out of memory. I tried with 32 and it is training but this shouldn't happen as resnet model is not taking that much memory.\nIs there any particular reason behind this? And any solutions?",
    "1328524": ""
  }
}