{
  "id": 103749,
  "title": "Do we need multiple GPUs to train Densenets or Resnet50 and higher?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/103749",
  "author_name": "",
  "post_date": "2019-08-11T14:02:38.025098300Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'd like to try \"heavier models\" but if I try resnet50 or any densenets with a GPU w/ 12GB I'm running out of memory. I'm obviously new to the field, so are you guys using multiple GPUs to train such models?</p>",
  "messages": [
    {
      "id": "596917",
      "postDate": "08/11/2019 14:02:38",
      "content": "<p>I'd like to try \"heavier models\" but if I try resnet50 or any densenets with a GPU w/ 12GB I'm running out of memory. I'm obviously new to the field, so are you guys using multiple GPUs to train such models?</p>",
      "rawMarkdown": "I'd like to try \"heavier models\" but if I try resnet50 or any densenets with a GPU w/ 12GB I'm running out of memory. I'm obviously new to the field, so are you guys using multiple GPUs to train such models?",
      "votes": null
    },
    {
      "id": "596924",
      "postDate": "08/11/2019 14:08:43",
      "content": "<p>hey Michel, you can use a smaller batch size or run on GCP if you applied for the credits since there isn't a 9-hour time limit. My experience has been it's harder to find the optimal learning rate - hence harder to converge on multiple GPUs. But that's just me and I am also new. I definitely appreciate other people's insight or tips on training on multiple GPUs.</p>",
      "rawMarkdown": "hey Michel, you can use a smaller batch size or run on GCP if you applied for the credits since there isn't a 9-hour time limit. My experience has been it's harder to find the optimal learning rate - hence harder to converge on multiple GPUs. But that's just me and I am also new. I definitely appreciate other people's insight or tips on training on multiple GPUs.",
      "votes": null
    },
    {
      "id": "596993",
      "postDate": "08/11/2019 16:17:34",
      "content": "<p>Not really answering your question but the P100 have already at least 16GB (on kaggle or GCP) and is much faster than the K80 that you're probably using (T4 and V100 are again much faster, but have the same capacity)\nAlso, using mixed precision can basically double your batch size...</p>",
      "rawMarkdown": "Not really answering your question but the P100 have already at least 16GB (on kaggle or GCP) and is much faster than the K80 that you're probably using (T4 and V100 are again much faster, but have the same capacity)\nAlso, using mixed precision can basically double your batch size...",
      "votes": null
    },
    {
      "id": "597060",
      "postDate": "08/11/2019 18:32:04",
      "content": "<p>12GB should be enough if you batch size is small enough (16 maybe) you can always use gradient accumulation to emulate larger batchs. </p>",
      "rawMarkdown": "12GB should be enough if you batch size is small enough (16 maybe) you can always use gradient accumulation to emulate larger batchs.",
      "votes": null
    },
    {
      "id": "597156",
      "postDate": "08/11/2019 23:55:21",
      "content": "<p>Thanks guys, that helped!</p>",
      "rawMarkdown": "Thanks guys, that helped!",
      "votes": null
    },
    {
      "id": "599408",
      "postDate": "08/14/2019 23:27:54",
      "content": "<p>Hi yuval, may I ask whether is gradient accumulation every 2 epoch essentially the same as using not using accumulation but twice the batch size? As in, given the same learning rate, will both roughly have the same performance? Thanks</p>",
      "rawMarkdown": "Hi yuval, may I ask whether is gradient accumulation every 2 epoch essentially the same as using not using accumulation but twice the batch size? As in, given the same learning rate, will both roughly have the same performance? Thanks",
      "votes": null
    },
    {
      "id": "599574",
      "postDate": "08/15/2019 06:31:45",
      "content": "<p>Yes using gradient accumulation of 2 is like double the batch size (but slower). look <a href=\"https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/13\">here</a> for more details. </p>",
      "rawMarkdown": "Yes using gradient accumulation of 2 is like double the batch size (but slower). look [here](https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/13) for more details.",
      "votes": null
    },
    {
      "id": "600042",
      "postDate": "08/15/2019 15:11:29",
      "content": "<p>Thanks, that helps! :) </p>",
      "rawMarkdown": "Thanks, that helps! :)",
      "votes": null
    },
    {
      "id": "600203",
      "postDate": "08/15/2019 19:38:05",
      "content": "<p>mixed precision was next on my todo, thanks! seems pretty convincing based on the bag of tricks paper</p>",
      "rawMarkdown": "mixed precision was next on my todo, thanks! seems pretty convincing based on the bag of tricks paper",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 596924,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "08/11/2019 14:08:43",
      "content": "<p>hey Michel, you can use a smaller batch size or run on GCP if you applied for the credits since there isn't a 9-hour time limit. My experience has been it's harder to find the optimal learning rate - hence harder to converge on multiple GPUs. But that's just me and I am also new. I definitely appreciate other people's insight or tips on training on multiple GPUs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 596993,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "08/11/2019 16:17:34",
      "content": "<p>Not really answering your question but the P100 have already at least 16GB (on kaggle or GCP) and is much faster than the K80 that you're probably using (T4 and V100 are again much faster, but have the same capacity)\nAlso, using mixed precision can basically double your batch size...</p>",
      "votes": null,
      "replies": [
        {
          "id": 600203,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/15/2019 19:38:05",
          "content": "<p>mixed precision was next on my todo, thanks! seems pretty convincing based on the bag of tricks paper</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 597060,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "08/11/2019 18:32:04",
      "content": "<p>12GB should be enough if you batch size is small enough (16 maybe) you can always use gradient accumulation to emulate larger batchs. </p>",
      "votes": null,
      "replies": [
        {
          "id": 599408,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/14/2019 23:27:54",
          "content": "<p>Hi yuval, may I ask whether is gradient accumulation every 2 epoch essentially the same as using not using accumulation but twice the batch size? As in, given the same learning rate, will both roughly have the same performance? Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599574,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "08/15/2019 06:31:45",
          "content": "<p>Yes using gradient accumulation of 2 is like double the batch size (but slower). look <a href=\"https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/13\">here</a> for more details. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 600042,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/15/2019 15:11:29",
          "content": "<p>Thanks, that helps! :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 597156,
      "author_name": "michelml",
      "author_url": "",
      "post_date": "08/11/2019 23:55:21",
      "content": "<p>Thanks guys, that helped!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "596917": "I'd like to try \"heavier models\" but if I try resnet50 or any densenets with a GPU w/ 12GB I'm running out of memory. I'm obviously new to the field, so are you guys using multiple GPUs to train such models?",
    "596924": "hey Michel, you can use a smaller batch size or run on GCP if you applied for the credits since there isn't a 9-hour time limit. My experience has been it's harder to find the optimal learning rate - hence harder to converge on multiple GPUs. But that's just me and I am also new. I definitely appreciate other people's insight or tips on training on multiple GPUs.",
    "596993": "Not really answering your question but the P100 have already at least 16GB (on kaggle or GCP) and is much faster than the K80 that you're probably using (T4 and V100 are again much faster, but have the same capacity)\nAlso, using mixed precision can basically double your batch size...",
    "597060": "12GB should be enough if you batch size is small enough (16 maybe) you can always use gradient accumulation to emulate larger batchs.",
    "597156": "Thanks guys, that helped!",
    "599408": "Hi yuval, may I ask whether is gradient accumulation every 2 epoch essentially the same as using not using accumulation but twice the batch size? As in, given the same learning rate, will both roughly have the same performance? Thanks",
    "599574": "Yes using gradient accumulation of 2 is like double the batch size (but slower). look [here](https://discuss.pytorch.org/t/why-do-we-need-to-set-the-gradients-manually-to-zero-in-pytorch/4903/13) for more details.",
    "600042": "Thanks, that helps! :)",
    "600203": "mixed precision was next on my todo, thanks! seems pretty convincing based on the bag of tricks paper"
  },
  "source": "meta"
}