{
  "id": 22058,
  "title": "Memory issues while running VGG16",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22058",
  "author_name": "",
  "post_date": "2016-07-06T03:01:27.047Z",
  "votes": null,
  "comment_count": 3,
  "views": 805,
  "content": "<p>Looking for some pointers on how much memory is required to run VGG16 using Keras based on a previous thread in this forum.\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20747/simple-lb-0-23800-solution-keras-vgg-16-pretrained/124670\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20747/simple-lb-0-23800-solution-keras-vgg-16-pretrained/124670</a></p>\n\n<p>I'm using AWS GPUs but I get an out of memory error even for a batch size of 8. I'm using a 8 GPU cluster with 60GB RAM but 4GB on each GPU nodes. Could this be the reason for the memory error?</p>\n\n<p>Alternately, if I were to try Resnet say using Caffe, would the above GPU configuration be able to run it?</p>",
  "messages": [
    {
      "id": "126086",
      "postDate": "07/06/2016 03:01:27",
      "content": "<p>Looking for some pointers on how much memory is required to run VGG16 using Keras based on a previous thread in this forum.\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20747/simple-lb-0-23800-solution-keras-vgg-16-pretrained/124670\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20747/simple-lb-0-23800-solution-keras-vgg-16-pretrained/124670</a></p>\n\n<p>I'm using AWS GPUs but I get an out of memory error even for a batch size of 8. I'm using a 8 GPU cluster with 60GB RAM but 4GB on each GPU nodes. Could this be the reason for the memory error?</p>\n\n<p>Alternately, if I were to try Resnet say using Caffe, would the above GPU configuration be able to run it?</p>",
      "rawMarkdown": "Looking for some pointers on how much memory is required to run VGG16 using Keras based on a previous thread in this forum.\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20747/simple-lb-0-23800-solution-keras-vgg-16-pretrained/124670\r\n\r\nI'm using AWS GPUs but I get an out of memory error even for a batch size of 8. I'm using a 8 GPU cluster with 60GB RAM but 4GB on each GPU nodes. Could this be the reason for the memory error?\r\n\r\nAlternately, if I were to try Resnet say using Caffe, would the above GPU configuration be able to run it?",
      "votes": null
    },
    {
      "id": "126089",
      "postDate": "07/06/2016 03:39:30",
      "content": "<p>Are you sure that you use SGD? Is should fit in memory with batch size 16.</p>",
      "rawMarkdown": "Are you sure that you use SGD? Is should fit in memory with batch size 16.",
      "votes": null
    },
    {
      "id": "126116",
      "postDate": "07/06/2016 12:11:58",
      "content": "<p>Are you sure your script does not download all the image data at once in any time? </p>\n\n<p>In this Jiao Dong's script ( <a href=\"https://www.kaggle.com/jiaodong/state-farm-distracted-driver-detection/vgg-16-pretrained-loss-0-23800/comments\">https://www.kaggle.com/jiaodong/state-farm-distracted-driver-detection/vgg-16-pretrained-loss-0-23800/comments</a>), function called read_and_normalize_and_shuffle_train_data download all the images at once and later train models with defined batch size.</p>\n\n<p>In my case, script got work when I changed his script not to download all the images at once and also change batch size to 16. As in your case, my GPU also could not handle batch size of 64. </p>\n\n<p>I am glad if this could be a help</p>",
      "rawMarkdown": "Are you sure your script does not download all the image data at once in any time? \r\n\r\nIn this Jiao Dong's script ( https://www.kaggle.com/jiaodong/state-farm-distracted-driver-detection/vgg-16-pretrained-loss-0-23800/comments), function called read_and_normalize_and_shuffle_train_data download all the images at once and later train models with defined batch size.\r\n\r\nIn my case, script got work when I changed his script not to download all the images at once and also change batch size to 16. As in your case, my GPU also could not handle batch size of 64. \r\n\r\nI am glad if this could be a help",
      "votes": null
    },
    {
      "id": "126220",
      "postDate": "07/07/2016 02:56:30",
      "content": "<p>I'm using SGD. Will try not downloading everything. Good idea. Thanks for the idea</p>",
      "rawMarkdown": "I'm using SGD. Will try not downloading everything. Good idea. Thanks for the idea",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 126089,
      "author_name": "ehsanma",
      "author_url": "",
      "post_date": "07/06/2016 03:39:30",
      "content": "<p>Are you sure that you use SGD? Is should fit in memory with batch size 16.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126116,
      "author_name": "katotakuya",
      "author_url": "",
      "post_date": "07/06/2016 12:11:58",
      "content": "<p>Are you sure your script does not download all the image data at once in any time? </p>\n\n<p>In this Jiao Dong's script ( <a href=\"https://www.kaggle.com/jiaodong/state-farm-distracted-driver-detection/vgg-16-pretrained-loss-0-23800/comments\">https://www.kaggle.com/jiaodong/state-farm-distracted-driver-detection/vgg-16-pretrained-loss-0-23800/comments</a>), function called read_and_normalize_and_shuffle_train_data download all the images at once and later train models with defined batch size.</p>\n\n<p>In my case, script got work when I changed his script not to download all the images at once and also change batch size to 16. As in your case, my GPU also could not handle batch size of 64. </p>\n\n<p>I am glad if this could be a help</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126220,
      "author_name": "pradeepram80",
      "author_url": "",
      "post_date": "07/07/2016 02:56:30",
      "content": "<p>I'm using SGD. Will try not downloading everything. Good idea. Thanks for the idea</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "126086": "Looking for some pointers on how much memory is required to run VGG16 using Keras based on a previous thread in this forum.\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20747/simple-lb-0-23800-solution-keras-vgg-16-pretrained/124670\r\n\r\nI'm using AWS GPUs but I get an out of memory error even for a batch size of 8. I'm using a 8 GPU cluster with 60GB RAM but 4GB on each GPU nodes. Could this be the reason for the memory error?\r\n\r\nAlternately, if I were to try Resnet say using Caffe, would the above GPU configuration be able to run it?",
    "126089": "Are you sure that you use SGD? Is should fit in memory with batch size 16.",
    "126116": "Are you sure your script does not download all the image data at once in any time? \r\n\r\nIn this Jiao Dong's script ( https://www.kaggle.com/jiaodong/state-farm-distracted-driver-detection/vgg-16-pretrained-loss-0-23800/comments), function called read_and_normalize_and_shuffle_train_data download all the images at once and later train models with defined batch size.\r\n\r\nIn my case, script got work when I changed his script not to download all the images at once and also change batch size to 16. As in your case, my GPU also could not handle batch size of 64. \r\n\r\nI am glad if this could be a help",
    "126220": "I'm using SGD. Will try not downloading everything. Good idea. Thanks for the idea"
  },
  "source": "meta"
}