{
  "id": 38703,
  "title": "Reached the edge of my limitations",
  "url": "/competitions/carvana-image-masking-challenge/discussion/38703",
  "author_name": "Ramiro Debbe",
  "post_date": "2017-08-29T12:55:38.610000",
  "votes": 1,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I have trained the U-Net model I picked up from LUNA16 challenge (<a href=\"https://www.kaggle.com/rdebbe/training-a-u-net-model-in-keras-theano\">kernel</a> ) with 512X512 images in keras TensorFlow back end and reached LB=0.996.</p>\n\n<p>I worked with an AWS p2.xlarge instance but when I tried to train images of higher resolution I started to run out of GPU memory even at 640X640.\nThe GPU in the instance I used is a Tesla K80 with 11.17 GiB memory. After the model is defined and compiled the free memory has not changed by much: 11.05 GiB\nTraining in TensorFlow with 640X640 (batch_size=16) the job runs out of memory just after the first step of epoch 1; I get a message from bfc_allocator as it runs out of memory trying to allocate 23.53 MiB</p>\n\n<p>I tried reducing the batch size but all I get is more steps before the crash happens.\nI also switched to the keras  Theano back end but the training continues to be memory limited.\nI was under the impression that the use of batches was intended to avoid what seems to be happening in my case; no memory is released after each batch is processed.</p>\n\n<p>I think I'm suffering from my being unfamiliar with the use of GPU's.\nWhen I train on my laptop (on CPU), I can use 1280X1280 images but I takes forever and with lots of page swapping. I hope some of you can give me some advice. Thanks in advance.</p>",
  "messages": [
    {
      "id": 217652,
      "postDate": "2017-08-31T14:00:06.100Z",
      "content": "<p>The number of channels near the input &amp; output layer drastically affect memory size, because feature maps have high resolution there.\nMy U-Net has only 8 channels on next to the input &amp; output layer by 1024x1024 size. And I can train the net on batchsize=4(11GB RAM).</p>",
      "rawMarkdown": "The number of channels near the input &amp; output layer drastically affect memory size, because feature maps have high resolution there.\nMy U-Net has only 8 channels on next to the input &amp; output layer by 1024x1024 size. And I can train the net on batchsize=4(11GB RAM).",
      "votes": 4,
      "replies": [
        {
          "id": 217656,
          "postDate": "2017-08-31T14:29:04.340Z",
          "content": "<p>Thanks for your suggestion,\nI was trying to raise the number of filters in first and last layer after I trained 1280X1280 images with 5 channels at beginning and end which gave me poor LB. Maybe I was too ambitious running at high resolution. The job is now set to use your parameters and is running well.</p>\n\n<p>And Congratulations for your fantastic LB position, I'm tempted to address you as \"Your Topness\"</p>",
          "rawMarkdown": "Thanks for your suggestion,\nI was trying to raise the number of filters in first and last layer after I trained 1280X1280 images with 5 channels at beginning and end which gave me poor LB. Maybe I was too ambitious running at high resolution. The job is now set to use your parameters and is running well.\n\nAnd Congratulations for your fantastic LB position, I'm tempted to address you as \"Your Topness\""
        },
        {
          "id": 217660,
          "postDate": "2017-08-31T14:47:11.463Z",
          "content": "<p>I'm glad I could help. </p>\n\n<p>As a matter of a fact, my position is undoubtedly not top since LB is broken...</p>",
          "rawMarkdown": "I'm glad I could help. \n\nAs a matter of a fact, my position is undoubtedly not top since LB is broken..."
        }
      ]
    },
    {
      "id": 217223,
      "postDate": "2017-08-29T20:05:58.990Z",
      "content": "<p>I'd try reducing number of features (channels) and using batch size = 1. </p>\n\n<p>With a 1080 Ti (11 Gb) I train at full res 1920 x 1280 w/ a custom net with \"only\" 115k params. GPU memory is the limiting factor for me; others are training at lower res and upscaling or slicing the image. Do some back of the envelope calculations of the total activation size for each layer, and it's ~2X iirc for the backward pass.</p>",
      "rawMarkdown": "I'd try reducing number of features (channels) and using batch size = 1. \n\nWith a 1080 Ti (11 Gb) I train at full res 1920 x 1280 w/ a custom net with \"only\" 115k params. GPU memory is the limiting factor for me; others are training at lower res and upscaling or slicing the image. Do some back of the envelope calculations of the total activation size for each layer, and it's ~2X iirc for the backward pass.",
      "votes": 1,
      "replies": [
        {
          "id": 217331,
          "postDate": "2017-08-30T07:18:06.253Z",
          "content": "<blockquote>\n  <p>Do some back of the envelope calculations of the total activation size for each layer, and it's ~2X iirc for the backward pass.\n  That's the way to go. The by far biggest part of memory consumption is not the number of parameters but the sizes of feature maps. Imagine you perform a convolution with an 3x3 filter producing 32 feature maps of size 512x512. In terms of trainable parameters, the GPU has to hold only 3*3*32 = 288 parameters. But this convolution is also producing an output: 32 feature maps of size 512x512. The training mechanism has to hold this output in memory (for example for backpropagation). So, these are 32*512*512 = 8.126.464 units/activations/numbers. For wide and deep networks, these activations add up quickly and lead to out of memory problems.</p>\n</blockquote>",
          "rawMarkdown": "&gt; Do some back of the envelope calculations of the total activation size for each layer, and it's ~2X iirc for the backward pass.\nThat's the way to go. The by far biggest part of memory consumption is not the number of parameters but the sizes of feature maps. Imagine you perform a convolution with an 3x3 filter producing 32 feature maps of size 512x512. In terms of trainable parameters, the GPU has to hold only 3*3*32 = 288 parameters. But this convolution is also producing an output: 32 feature maps of size 512x512. The training mechanism has to hold this output in memory (for example for backpropagation). So, these are 32*512*512 = 8.126.464 units/activations/numbers. For wide and deep networks, these activations add up quickly and lead to out of memory problems."
        }
      ]
    },
    {
      "id": 217170,
      "postDate": "2017-08-29T16:37:36.510Z",
      "content": "<p>It also depends on how many channels your Unet starts (double the channels means double the tensor size output from the first convolution) with and the depth of your Unet. Naturally there is a tradeoff between the accuracy, resolution, depth and channels #.. I would try lowering the depth and the channels, till it starts working, then work from there.</p>",
      "rawMarkdown": "It also depends on how many channels your Unet starts (double the channels means double the tensor size output from the first convolution) with and the depth of your Unet. Naturally there is a tradeoff between the accuracy, resolution, depth and channels #.. I would try lowering the depth and the channels, till it starts working, then work from there.",
      "votes": 1,
      "replies": [
        {
          "id": 217203,
          "postDate": "2017-08-29T18:34:57.610Z",
          "content": "<p>Hi Francisco,\nLowering the number of channels would certainly be the easiest, just open the images with cv2 as monochrome. I understand that depth is the number of layers in the model, I will also try that and I will try not to mess the model up. Thanks for the suggestions.\np.s. Looking at the code I realized that the batch generator delivers arrays of numpy.float32. If I change that to int8 I could lower some of the memory use, but I imagine the weights of the model have to be floats and that is where most of the  memory is used. </p>\n\n<p>Edited:\nNow I understand that you were referring to the number of channels in the first convolution, not the input. </p>",
          "rawMarkdown": "Hi Francisco,\nLowering the number of channels would certainly be the easiest, just open the images with cv2 as monochrome. I understand that depth is the number of layers in the model, I will also try that and I will try not to mess the model up. Thanks for the suggestions.\np.s. Looking at the code I realized that the batch generator delivers arrays of numpy.float32. If I change that to int8 I could lower some of the memory use, but I imagine the weights of the model have to be floats and that is where most of the  memory is used. \n\nEdited:\nNow I understand that you were referring to the number of channels in the first convolution, not the input. "
        },
        {
          "id": 217208,
          "postDate": "2017-08-29T19:00:22.590Z",
          "content": "<p>Even if your generator returns int8 data, the images on the GPU will be in float because that is necessary to perform operations like convolutions (at least to my knowledge). So, changing the generator's output to int8 will only save you some internal memory. But I think that's not the problem here, but rather GPU memory.</p>",
          "rawMarkdown": "Even if your generator returns int8 data, the images on the GPU will be in float because that is necessary to perform operations like convolutions (at least to my knowledge). So, changing the generator's output to int8 will only save you some internal memory. But I think that's not the problem here, but rather GPU memory."
        },
        {
          "id": 217210,
          "postDate": "2017-08-29T19:07:53.653Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 217215,
          "postDate": "2017-08-29T19:41:46.170Z",
          "content": "<p>Thanks Ramtin,\nI have no experience looking for memory leaks in python. But I can change my model altogether, I have seen several versions in this competition and more in github.</p>\n\n<p>Edited, I switched to ecobill's model (which is the same as the LUNA16 one). It allows me to change the number of channels (filters) in the first convolution. I changed that number to 16 (I had 32) and the job is training happily with 640X640 images and batch size of 2, much better that what I was seeing yesterday. It appears that some hurdle has been cleared.</p>",
          "rawMarkdown": "Thanks Ramtin,\nI have no experience looking for memory leaks in python. But I can change my model altogether, I have seen several versions in this competition and more in github.\n\nEdited, I switched to ecobill's model (which is the same as the LUNA16 one). It allows me to change the number of channels (filters) in the first convolution. I changed that number to 16 (I had 32) and the job is training happily with 640X640 images and batch size of 2, much better that what I was seeing yesterday. It appears that some hurdle has been cleared."
        }
      ]
    },
    {
      "id": 217111,
      "postDate": "2017-08-29T12:55:38.610Z",
      "content": "<p>I have trained the U-Net model I picked up from LUNA16 challenge (<a href=\"https://www.kaggle.com/rdebbe/training-a-u-net-model-in-keras-theano\">kernel</a> ) with 512X512 images in keras TensorFlow back end and reached LB=0.996.</p>\n\n<p>I worked with an AWS p2.xlarge instance but when I tried to train images of higher resolution I started to run out of GPU memory even at 640X640.\nThe GPU in the instance I used is a Tesla K80 with 11.17 GiB memory. After the model is defined and compiled the free memory has not changed by much: 11.05 GiB\nTraining in TensorFlow with 640X640 (batch_size=16) the job runs out of memory just after the first step of epoch 1; I get a message from bfc_allocator as it runs out of memory trying to allocate 23.53 MiB</p>\n\n<p>I tried reducing the batch size but all I get is more steps before the crash happens.\nI also switched to the keras  Theano back end but the training continues to be memory limited.\nI was under the impression that the use of batches was intended to avoid what seems to be happening in my case; no memory is released after each batch is processed.</p>\n\n<p>I think I'm suffering from my being unfamiliar with the use of GPU's.\nWhen I train on my laptop (on CPU), I can use 1280X1280 images but I takes forever and with lots of page swapping. I hope some of you can give me some advice. Thanks in advance.</p>",
      "rawMarkdown": "I have trained the U-Net model I picked up from LUNA16 challenge ([kernel](https://www.kaggle.com/rdebbe/training-a-u-net-model-in-keras-theano) ) with 512X512 images in keras TensorFlow back end and reached LB=0.996.\n\nI worked with an AWS p2.xlarge instance but when I tried to train images of higher resolution I started to run out of GPU memory even at 640X640.\nThe GPU in the instance I used is a Tesla K80 with 11.17 GiB memory. After the model is defined and compiled the free memory has not changed by much: 11.05 GiB\nTraining in TensorFlow with 640X640 (batch_size=16) the job runs out of memory just after the first step of epoch 1; I get a message from bfc_allocator as it runs out of memory trying to allocate 23.53 MiB\n\nI tried reducing the batch size but all I get is more steps before the crash happens.\nI also switched to the keras  Theano back end but the training continues to be memory limited.\nI was under the impression that the use of batches was intended to avoid what seems to be happening in my case; no memory is released after each batch is processed.\n\nI think I'm suffering from my being unfamiliar with the use of GPU's.\nWhen I train on my laptop (on CPU), I can use 1280X1280 images but I takes forever and with lots of page swapping. I hope some of you can give me some advice. Thanks in advance.\n\n",
      "votes": 1
    },
    {
      "id": 217371,
      "postDate": "2017-08-30T10:59:26.817Z",
      "content": "<p>Many thanks to all. \nI lowered the number of filters (channels) in the model to 5. I'm training with 1280X1280 input images in batches with 10 samples. The number of parameters is low:  \"Trainable params: 190,026.0\" !\nThe job has been running overnight and reached epoch 19. It seems to be training well, the latest performance shows val_dice_loss = 0.9708. Using the top command I can see at most 12% of memory use. (I imagine that is GPU memory)\nI will then face the next hurdle to produce the predictions, but it looks like you helped me a lot.</p>",
      "rawMarkdown": "Many thanks to all. \nI lowered the number of filters (channels) in the model to 5. I'm training with 1280X1280 input images in batches with 10 samples. The number of parameters is low:  \"Trainable params: 190,026.0\" !\nThe job has been running overnight and reached epoch 19. It seems to be training well, the latest performance shows val_dice_loss = 0.9708. Using the top command I can see at most 12% of memory use. (I imagine that is GPU memory)\nI will then face the next hurdle to produce the predictions, but it looks like you helped me a lot.",
      "replies": [
        {
          "id": 217396,
          "postDate": "2017-08-30T12:39:43.873Z",
          "content": "<p>The top command usually shows you the internal memory. You can check GPU memory using the command <code>nvidia-smi</code>. If you want a \"live version\" of it, you can use <code>watch -n 1 nvidia-smi</code> (this updates the output of <code>nvidia-smi</code> every 1 second).</p>",
          "rawMarkdown": "The top command usually shows you the internal memory. You can check GPU memory using the command `nvidia-smi`. If you want a \"live version\" of it, you can use `watch -n 1 nvidia-smi` (this updates the output of `nvidia-smi` every 1 second).",
          "votes": 1
        },
        {
          "id": 217403,
          "postDate": "2017-08-30T12:56:12.417Z",
          "content": "<p>I'm learning so much! </p>\n\n<p>It looks like I'm using all of the 11 GiB of the single GPU in the instance!</p>\n\n<pre><code>\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 367.57                 Driver Version: 367.57                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  Tesla K80           On   | 0000:00:1E.0     Off |                    0 |\n| N/A   63C    P0   155W / 149W |  10938MiB / 11439MiB |    100%      Default |\n+-------------------------------+----------------------+----------------------+\n\n+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID  Type  Process name                               Usage      |\n|=============================================================================|\n|    0     18923    C   python                                       10934MiB |\n+-----------------------------------------------------------------------------+\n</code></pre>",
          "rawMarkdown": "I'm learning so much! \n\nIt looks like I'm using all of the 11 GiB of the single GPU in the instance!\n\n<pre><code>\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 367.57                 Driver Version: 367.57                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  Tesla K80           On   | 0000:00:1E.0     Off |                    0 |\n| N/A   63C    P0   155W / 149W |  10938MiB / 11439MiB |    100%      Default |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID  Type  Process name                               Usage      |\n|=============================================================================|\n|    0     18923    C   python                                       10934MiB |\n+-----------------------------------------------------------------------------+\n</code></pre>"
        },
        {
          "id": 217406,
          "postDate": "2017-08-30T13:30:53.897Z",
          "content": "<p>Given the input image size, you probably are using that much - However, unless specified otherwise, Tensorflow will occupy all the GPU memory - so '10934MiB' may not be the true about actually being used.</p>",
          "rawMarkdown": "Given the input image size, you probably are using that much - However, unless specified otherwise, Tensorflow will occupy all the GPU memory - so '10934MiB' may not be the true about actually being used."
        }
      ]
    },
    {
      "id": 217129,
      "postDate": "2017-08-29T13:52:14.857Z",
      "content": "<p>Are you saying even with a batch size of 1 you cannot train on a Telsa K80? I'm training a U-net on 1024x1024 images with a batch size of 4 on a 1080Ti 11Gb - So you should manage?</p>\n\n<p>How many parameters is your U-net? - Mine is approximately 33M.</p>",
      "rawMarkdown": "Are you saying even with a batch size of 1 you cannot train on a Telsa K80? I'm training a U-net on 1024x1024 images with a batch size of 4 on a 1080Ti 11Gb - So you should manage?\n\nHow many parameters is your U-net? - Mine is approximately 33M.",
      "replies": [
        {
          "id": 217153,
          "postDate": "2017-08-29T15:40:51.590Z",
          "content": "<p>Hi Craig,\nThe model summary gives me: Trainable params: 7,846,657.0\nMuch smaller than your number.</p>",
          "rawMarkdown": "Hi Craig,\nThe model summary gives me: Trainable params: 7,846,657.0\nMuch smaller than your number."
        },
        {
          "id": 217234,
          "postDate": "2017-08-29T21:19:35.547Z",
          "content": "<p>It's likely an issue with the way you're minibatching your data - Can you supply the code? Else it's very difficult to help you any further.</p>",
          "rawMarkdown": "It's likely an issue with the way you're minibatching your data - Can you supply the code? Else it's very difficult to help you any further."
        },
        {
          "id": 217245,
          "postDate": "2017-08-29T22:40:16.007Z",
          "content": "<p>Let me show the relevant parts of the code for the batch generators based on petrosgk \n(<a href=\"https://github.com/petrosgk/Kaggle-Carvana-Image-Masking-Challenge\">https://github.com/petrosgk/Kaggle-Carvana-Image-Masking-Challenge</a>)\nwith some small modifications</p>\n\n<pre><code>\ndef data_gen_small(data_dir, train_images, batch_size, input_size):\n        while True:\n            #\n            # use all data sequentially\n            #\n            for start in range(0, len(train_images), batch_size):\n                x_batch = []\n                y_batch = []\n                end = min(start + batch_size, len(train_images))\n                ix = train_images[start:end] \n                #print('len ix ', len(ix))\n                imgs = []\n                labels = []\n                for i in ix:\n                    img = cv2.imread(i)\n                    img = cv2.resize(img, (input_size, input_size))\n                    mask_filename = basename(i)\n                    no_extension = os.path.splitext(mask_filename)[0]\n                    correct_mask = INPUT_PATH + 'train_masks/sandbox/' + no_extension + '_mask.png' \n                    mask = cv2.imread(correct_mask, cv2.IMREAD_GRAYSCALE)\n                    mask = cv2.resize(mask, (input_size, input_size))\n                    img = randomHueSaturationValue(img,\n                                               hue_shift_limit=(-50, 50),\n                                               sat_shift_limit=(-5, 5),\n                                               val_shift_limit=(-15, 15))\n                    img, mask = randomShiftScaleRotate(img, mask,\n                                                   shift_limit=(-0.0625, 0.0625),\n                                                   scale_limit=(-0.1, 0.1),\n                                                   rotate_limit=(-0, 0))\n                    img, mask = randomHorizontalFlip(img, mask)\n                    mask = np.expand_dims(mask, axis=2)\n                    x_batch.append(img)\n                    y_batch.append(mask)\n                x_batch = np.array(x_batch, np.float32) / 255\n                y_batch = np.array(y_batch, np.float32) / 255\n                yield x_batch, y_batch\ndef valid_generator(validation_images, batch_size):\n    while True:\n        for start in range(0, len(validation_images), batch_size, input_size):\n            x_batch = []\n            y_batch = []\n            end = min(start + batch_size, len(validation_images))\n            ids_valid_batch = validation_images[start:end]\n            for id in ids_valid_batch:\n                img = cv2.imread(id)\n                img = cv2.resize(img, (input_size, input_size))\n                mask_filename = basename(id)\n                no_extension = os.path.splitext(mask_filename)[0]\n                correct_mask = INPUT_PATH + 'train_masks/sandbox/' + no_extension + '_mask.png'               \n                mask = cv2.imread(correct_mask, cv2.IMREAD_GRAYSCALE)\n                mask = cv2.resize(mask, (input_size, input_size))\n                mask = np.expand_dims(mask, axis=2)\n                x_batch.append(img)\n                y_batch.append(mask)\n            x_batch = np.array(x_batch, np.float32) / 255\n            y_batch = np.array(y_batch, np.float32) / 255\n            yield x_batch, y_batch\n#\ntrain_gen = data_gen_small(data_dir, masks, train_images, 10) \nimg, msk = next(train_gen) \n#\n# create an instance of a validation generator:\nvalidation_gen = valid_generator(validation_images, 10) \nimgv, mskv = next(validation_gen)\n</code></pre>\n\n<p>The actual call to do the training:</p>\n\n<pre><code>\nmodel.fit_generator(generator=train_gen,\n                    steps_per_epoch=np.ceil(float(len(train_images)) / float(batch_size)),\n                    epochs=max_epochs,\n                    verbose=1,\n                    callbacks=callbacks,\n                    validation_data=validation_gen,\n                    validation_steps=np.ceil(float(len(validation_images)) / float(batch_size)))\n</code></pre>",
          "rawMarkdown": "Let me show the relevant parts of the code for the batch generators based on petrosgk \n(https://github.com/petrosgk/Kaggle-Carvana-Image-Masking-Challenge)\nwith some small modifications\n\n<pre><code>\ndef data_gen_small(data_dir, train_images, batch_size, input_size):\n        while True:\n            #\n            # use all data sequentially\n            #\n            for start in range(0, len(train_images), batch_size):\n                x_batch = []\n                y_batch = []\n                end = min(start + batch_size, len(train_images))\n                ix = train_images[start:end] \n                #print('len ix ', len(ix))\n                imgs = []\n                labels = []\n                for i in ix:\n                    img = cv2.imread(i)\n                    img = cv2.resize(img, (input_size, input_size))\n                    mask_filename = basename(i)\n                    no_extension = os.path.splitext(mask_filename)[0]\n                    correct_mask = INPUT_PATH + 'train_masks/sandbox/' + no_extension + '_mask.png' \n                    mask = cv2.imread(correct_mask, cv2.IMREAD_GRAYSCALE)\n                    mask = cv2.resize(mask, (input_size, input_size))\n                    img = randomHueSaturationValue(img,\n                                               hue_shift_limit=(-50, 50),\n                                               sat_shift_limit=(-5, 5),\n                                               val_shift_limit=(-15, 15))\n                    img, mask = randomShiftScaleRotate(img, mask,\n                                                   shift_limit=(-0.0625, 0.0625),\n                                                   scale_limit=(-0.1, 0.1),\n                                                   rotate_limit=(-0, 0))\n                    img, mask = randomHorizontalFlip(img, mask)\n                    mask = np.expand_dims(mask, axis=2)\n                    x_batch.append(img)\n                    y_batch.append(mask)\n                x_batch = np.array(x_batch, np.float32) / 255\n                y_batch = np.array(y_batch, np.float32) / 255\n                yield x_batch, y_batch\ndef valid_generator(validation_images, batch_size):\n    while True:\n        for start in range(0, len(validation_images), batch_size, input_size):\n            x_batch = []\n            y_batch = []\n            end = min(start + batch_size, len(validation_images))\n            ids_valid_batch = validation_images[start:end]\n            for id in ids_valid_batch:\n                img = cv2.imread(id)\n                img = cv2.resize(img, (input_size, input_size))\n                mask_filename = basename(id)\n                no_extension = os.path.splitext(mask_filename)[0]\n                correct_mask = INPUT_PATH + 'train_masks/sandbox/' + no_extension + '_mask.png'               \n                mask = cv2.imread(correct_mask, cv2.IMREAD_GRAYSCALE)\n                mask = cv2.resize(mask, (input_size, input_size))\n                mask = np.expand_dims(mask, axis=2)\n                x_batch.append(img)\n                y_batch.append(mask)\n            x_batch = np.array(x_batch, np.float32) / 255\n            y_batch = np.array(y_batch, np.float32) / 255\n            yield x_batch, y_batch\n#\ntrain_gen = data_gen_small(data_dir, masks, train_images, 10) \nimg, msk = next(train_gen) \n#\n# create an instance of a validation generator:\nvalidation_gen = valid_generator(validation_images, 10) \nimgv, mskv = next(validation_gen)\n</code></pre>\n\nThe actual call to do the training:\n<pre><code>\nmodel.fit_generator(generator=train_gen,\n                    steps_per_epoch=np.ceil(float(len(train_images)) / float(batch_size)),\n                    epochs=max_epochs,\n                    verbose=1,\n                    callbacks=callbacks,\n                    validation_data=validation_gen,\n                    validation_steps=np.ceil(float(len(validation_images)) / float(batch_size)))\n</code></pre>\n\n\n    "
        },
        {
          "id": 217328,
          "postDate": "2017-08-30T07:06:48.737Z",
          "content": "<p>Totally unrelated to your problem: Your generators for training and validation do exactly the same except the augmentation part. In general you can save a lot of code (and possible bugs) by having just one generator with an additional boolean <code>augment</code> argument.</p>\n\n<p>Another thing I have noticed is that your batches are not shuffled during training. It is common practice to shuffle the data before each epoch in order to have different compositions of batches.</p>",
          "rawMarkdown": "Totally unrelated to your problem: Your generators for training and validation do exactly the same except the augmentation part. In general you can save a lot of code (and possible bugs) by having just one generator with an additional boolean `augment` argument.\n\nAnother thing I have noticed is that your batches are not shuffled during training. It is common practice to shuffle the data before each epoch in order to have different compositions of batches."
        },
        {
          "id": 217364,
          "postDate": "2017-08-30T10:32:34.837Z",
          "content": "<p>Another good idea is to check what you're actually outputting from the generator and feeding into your network,  e.g. using the .shape attribute:</p>\n\n<p>img, msk = next(train_gen) \nprint img.shape\nprint msk.shape</p>\n\n<p>In your case, this must be something like (batch_size, input_dim1, input_dim2, channels) or translated (1, 1024, 1024, 3). </p>",
          "rawMarkdown": "Another good idea is to check what you're actually outputting from the generator and feeding into your network,  e.g. using the .shape attribute:\n\nimg, msk = next(train_gen) \nprint img.shape\nprint msk.shape\n\nIn your case, this must be something like (batch_size, input_dim1, input_dim2, channels) or translated (1, 1024, 1024, 3). "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 217652,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "2017-08-31T14:00:06.100000",
      "content": "<p>The number of channels near the input &amp; output layer drastically affect memory size, because feature maps have high resolution there.\nMy U-Net has only 8 channels on next to the input &amp; output layer by 1024x1024 size. And I can train the net on batchsize=4(11GB RAM).</p>",
      "votes": 4,
      "replies": [
        {
          "id": 217656,
          "author_name": "Ramiro Debbe",
          "author_url": "",
          "post_date": "2017-08-31T14:29:04.340000",
          "content": "<p>Thanks for your suggestion,\nI was trying to raise the number of filters in first and last layer after I trained 1280X1280 images with 5 channels at beginning and end which gave me poor LB. Maybe I was too ambitious running at high resolution. The job is now set to use your parameters and is running well.</p>\n\n<p>And Congratulations for your fantastic LB position, I'm tempted to address you as \"Your Topness\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217660,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2017-08-31T14:47:11.463000",
          "content": "<p>I'm glad I could help. </p>\n\n<p>As a matter of a fact, my position is undoubtedly not top since LB is broken...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 217223,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2017-08-29T20:05:58.990000",
      "content": "<p>I'd try reducing number of features (channels) and using batch size = 1. </p>\n\n<p>With a 1080 Ti (11 Gb) I train at full res 1920 x 1280 w/ a custom net with \"only\" 115k params. GPU memory is the limiting factor for me; others are training at lower res and upscaling or slicing the image. Do some back of the envelope calculations of the total activation size for each layer, and it's ~2X iirc for the backward pass.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 217331,
          "author_name": "Markus",
          "author_url": "",
          "post_date": "2017-08-30T07:18:06.253000",
          "content": "<blockquote>\n  <p>Do some back of the envelope calculations of the total activation size for each layer, and it's ~2X iirc for the backward pass.\n  That's the way to go. The by far biggest part of memory consumption is not the number of parameters but the sizes of feature maps. Imagine you perform a convolution with an 3x3 filter producing 32 feature maps of size 512x512. In terms of trainable parameters, the GPU has to hold only 3*3*32 = 288 parameters. But this convolution is also producing an output: 32 feature maps of size 512x512. The training mechanism has to hold this output in memory (for example for backpropagation). So, these are 32*512*512 = 8.126.464 units/activations/numbers. For wide and deep networks, these activations add up quickly and lead to out of memory problems.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 217170,
      "author_name": "HariSeldon",
      "author_url": "",
      "post_date": "2017-08-29T16:37:36.510000",
      "content": "<p>It also depends on how many channels your Unet starts (double the channels means double the tensor size output from the first convolution) with and the depth of your Unet. Naturally there is a tradeoff between the accuracy, resolution, depth and channels #.. I would try lowering the depth and the channels, till it starts working, then work from there.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 217203,
          "author_name": "Ramiro Debbe",
          "author_url": "",
          "post_date": "2017-08-29T18:34:57.610000",
          "content": "<p>Hi Francisco,\nLowering the number of channels would certainly be the easiest, just open the images with cv2 as monochrome. I understand that depth is the number of layers in the model, I will also try that and I will try not to mess the model up. Thanks for the suggestions.\np.s. Looking at the code I realized that the batch generator delivers arrays of numpy.float32. If I change that to int8 I could lower some of the memory use, but I imagine the weights of the model have to be floats and that is where most of the  memory is used. </p>\n\n<p>Edited:\nNow I understand that you were referring to the number of channels in the first convolution, not the input. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217208,
          "author_name": "Markus",
          "author_url": "",
          "post_date": "2017-08-29T19:00:22.590000",
          "content": "<p>Even if your generator returns int8 data, the images on the GPU will be in float because that is necessary to perform operations like convolutions (at least to my knowledge). So, changing the generator's output to int8 will only save you some internal memory. But I think that's not the problem here, but rather GPU memory.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217210,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-08-29T19:07:53.653000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217215,
          "author_name": "Ramiro Debbe",
          "author_url": "",
          "post_date": "2017-08-29T19:41:46.170000",
          "content": "<p>Thanks Ramtin,\nI have no experience looking for memory leaks in python. But I can change my model altogether, I have seen several versions in this competition and more in github.</p>\n\n<p>Edited, I switched to ecobill's model (which is the same as the LUNA16 one). It allows me to change the number of channels (filters) in the first convolution. I changed that number to 16 (I had 32) and the job is training happily with 640X640 images and batch size of 2, much better that what I was seeing yesterday. It appears that some hurdle has been cleared.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 217371,
      "author_name": "Ramiro Debbe",
      "author_url": "",
      "post_date": "2017-08-30T10:59:26.817000",
      "content": "<p>Many thanks to all. \nI lowered the number of filters (channels) in the model to 5. I'm training with 1280X1280 input images in batches with 10 samples. The number of parameters is low:  \"Trainable params: 190,026.0\" !\nThe job has been running overnight and reached epoch 19. It seems to be training well, the latest performance shows val_dice_loss = 0.9708. Using the top command I can see at most 12% of memory use. (I imagine that is GPU memory)\nI will then face the next hurdle to produce the predictions, but it looks like you helped me a lot.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 217396,
          "author_name": "Markus",
          "author_url": "",
          "post_date": "2017-08-30T12:39:43.873000",
          "content": "<p>The top command usually shows you the internal memory. You can check GPU memory using the command <code>nvidia-smi</code>. If you want a \"live version\" of it, you can use <code>watch -n 1 nvidia-smi</code> (this updates the output of <code>nvidia-smi</code> every 1 second).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 217403,
          "author_name": "Ramiro Debbe",
          "author_url": "",
          "post_date": "2017-08-30T12:56:12.417000",
          "content": "<p>I'm learning so much! </p>\n\n<p>It looks like I'm using all of the 11 GiB of the single GPU in the instance!</p>\n\n<pre><code>\n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 367.57                 Driver Version: 367.57                    |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|===============================+======================+======================|\n|   0  Tesla K80           On   | 0000:00:1E.0     Off |                    0 |\n| N/A   63C    P0   155W / 149W |  10938MiB / 11439MiB |    100%      Default |\n+-------------------------------+----------------------+----------------------+\n\n+-----------------------------------------------------------------------------+\n| Processes:                                                       GPU Memory |\n|  GPU       PID  Type  Process name                               Usage      |\n|=============================================================================|\n|    0     18923    C   python                                       10934MiB |\n+-----------------------------------------------------------------------------+\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217406,
          "author_name": "Craig Glastonbury",
          "author_url": "",
          "post_date": "2017-08-30T13:30:53.897000",
          "content": "<p>Given the input image size, you probably are using that much - However, unless specified otherwise, Tensorflow will occupy all the GPU memory - so '10934MiB' may not be the true about actually being used.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 217129,
      "author_name": "Craig Glastonbury",
      "author_url": "",
      "post_date": "2017-08-29T13:52:14.857000",
      "content": "<p>Are you saying even with a batch size of 1 you cannot train on a Telsa K80? I'm training a U-net on 1024x1024 images with a batch size of 4 on a 1080Ti 11Gb - So you should manage?</p>\n\n<p>How many parameters is your U-net? - Mine is approximately 33M.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 217153,
          "author_name": "Ramiro Debbe",
          "author_url": "",
          "post_date": "2017-08-29T15:40:51.590000",
          "content": "<p>Hi Craig,\nThe model summary gives me: Trainable params: 7,846,657.0\nMuch smaller than your number.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217234,
          "author_name": "Craig Glastonbury",
          "author_url": "",
          "post_date": "2017-08-29T21:19:35.547000",
          "content": "<p>It's likely an issue with the way you're minibatching your data - Can you supply the code? Else it's very difficult to help you any further.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217245,
          "author_name": "Ramiro Debbe",
          "author_url": "",
          "post_date": "2017-08-29T22:40:16.007000",
          "content": "<p>Let me show the relevant parts of the code for the batch generators based on petrosgk \n(<a href=\"https://github.com/petrosgk/Kaggle-Carvana-Image-Masking-Challenge\">https://github.com/petrosgk/Kaggle-Carvana-Image-Masking-Challenge</a>)\nwith some small modifications</p>\n\n<pre><code>\ndef data_gen_small(data_dir, train_images, batch_size, input_size):\n        while True:\n            #\n            # use all data sequentially\n            #\n            for start in range(0, len(train_images), batch_size):\n                x_batch = []\n                y_batch = []\n                end = min(start + batch_size, len(train_images))\n                ix = train_images[start:end] \n                #print('len ix ', len(ix))\n                imgs = []\n                labels = []\n                for i in ix:\n                    img = cv2.imread(i)\n                    img = cv2.resize(img, (input_size, input_size))\n                    mask_filename = basename(i)\n                    no_extension = os.path.splitext(mask_filename)[0]\n                    correct_mask = INPUT_PATH + 'train_masks/sandbox/' + no_extension + '_mask.png' \n                    mask = cv2.imread(correct_mask, cv2.IMREAD_GRAYSCALE)\n                    mask = cv2.resize(mask, (input_size, input_size))\n                    img = randomHueSaturationValue(img,\n                                               hue_shift_limit=(-50, 50),\n                                               sat_shift_limit=(-5, 5),\n                                               val_shift_limit=(-15, 15))\n                    img, mask = randomShiftScaleRotate(img, mask,\n                                                   shift_limit=(-0.0625, 0.0625),\n                                                   scale_limit=(-0.1, 0.1),\n                                                   rotate_limit=(-0, 0))\n                    img, mask = randomHorizontalFlip(img, mask)\n                    mask = np.expand_dims(mask, axis=2)\n                    x_batch.append(img)\n                    y_batch.append(mask)\n                x_batch = np.array(x_batch, np.float32) / 255\n                y_batch = np.array(y_batch, np.float32) / 255\n                yield x_batch, y_batch\ndef valid_generator(validation_images, batch_size):\n    while True:\n        for start in range(0, len(validation_images), batch_size, input_size):\n            x_batch = []\n            y_batch = []\n            end = min(start + batch_size, len(validation_images))\n            ids_valid_batch = validation_images[start:end]\n            for id in ids_valid_batch:\n                img = cv2.imread(id)\n                img = cv2.resize(img, (input_size, input_size))\n                mask_filename = basename(id)\n                no_extension = os.path.splitext(mask_filename)[0]\n                correct_mask = INPUT_PATH + 'train_masks/sandbox/' + no_extension + '_mask.png'               \n                mask = cv2.imread(correct_mask, cv2.IMREAD_GRAYSCALE)\n                mask = cv2.resize(mask, (input_size, input_size))\n                mask = np.expand_dims(mask, axis=2)\n                x_batch.append(img)\n                y_batch.append(mask)\n            x_batch = np.array(x_batch, np.float32) / 255\n            y_batch = np.array(y_batch, np.float32) / 255\n            yield x_batch, y_batch\n#\ntrain_gen = data_gen_small(data_dir, masks, train_images, 10) \nimg, msk = next(train_gen) \n#\n# create an instance of a validation generator:\nvalidation_gen = valid_generator(validation_images, 10) \nimgv, mskv = next(validation_gen)\n</code></pre>\n\n<p>The actual call to do the training:</p>\n\n<pre><code>\nmodel.fit_generator(generator=train_gen,\n                    steps_per_epoch=np.ceil(float(len(train_images)) / float(batch_size)),\n                    epochs=max_epochs,\n                    verbose=1,\n                    callbacks=callbacks,\n                    validation_data=validation_gen,\n                    validation_steps=np.ceil(float(len(validation_images)) / float(batch_size)))\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217328,
          "author_name": "Markus",
          "author_url": "",
          "post_date": "2017-08-30T07:06:48.737000",
          "content": "<p>Totally unrelated to your problem: Your generators for training and validation do exactly the same except the augmentation part. In general you can save a lot of code (and possible bugs) by having just one generator with an additional boolean <code>augment</code> argument.</p>\n\n<p>Another thing I have noticed is that your batches are not shuffled during training. It is common practice to shuffle the data before each epoch in order to have different compositions of batches.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 217364,
          "author_name": "rubenhx",
          "author_url": "",
          "post_date": "2017-08-30T10:32:34.837000",
          "content": "<p>Another good idea is to check what you're actually outputting from the generator and feeding into your network,  e.g. using the .shape attribute:</p>\n\n<p>img, msk = next(train_gen) \nprint img.shape\nprint msk.shape</p>\n\n<p>In your case, this must be something like (batch_size, input_dim1, input_dim2, channels) or translated (1, 1024, 1024, 3). </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "217652": "The number of channels near the input &amp; output layer drastically affect memory size, because feature maps have high resolution there.\nMy U-Net has only 8 channels on next to the input &amp; output layer by 1024x1024 size. And I can train the net on batchsize=4(11GB RAM).",
    "217223": "I'd try reducing number of features (channels) and using batch size = 1. \n\nWith a 1080 Ti (11 Gb) I train at full res 1920 x 1280 w/ a custom net with \"only\" 115k params. GPU memory is the limiting factor for me; others are training at lower res and upscaling or slicing the image. Do some back of the envelope calculations of the total activation size for each layer, and it's ~2X iirc for the backward pass.",
    "217170": "It also depends on how many channels your Unet starts (double the channels means double the tensor size output from the first convolution) with and the depth of your Unet. Naturally there is a tradeoff between the accuracy, resolution, depth and channels #.. I would try lowering the depth and the channels, till it starts working, then work from there.",
    "217111": "I have trained the U-Net model I picked up from LUNA16 challenge ([kernel](https://www.kaggle.com/rdebbe/training-a-u-net-model-in-keras-theano) ) with 512X512 images in keras TensorFlow back end and reached LB=0.996.\n\nI worked with an AWS p2.xlarge instance but when I tried to train images of higher resolution I started to run out of GPU memory even at 640X640.\nThe GPU in the instance I used is a Tesla K80 with 11.17 GiB memory. After the model is defined and compiled the free memory has not changed by much: 11.05 GiB\nTraining in TensorFlow with 640X640 (batch_size=16) the job runs out of memory just after the first step of epoch 1; I get a message from bfc_allocator as it runs out of memory trying to allocate 23.53 MiB\n\nI tried reducing the batch size but all I get is more steps before the crash happens.\nI also switched to the keras  Theano back end but the training continues to be memory limited.\nI was under the impression that the use of batches was intended to avoid what seems to be happening in my case; no memory is released after each batch is processed.\n\nI think I'm suffering from my being unfamiliar with the use of GPU's.\nWhen I train on my laptop (on CPU), I can use 1280X1280 images but I takes forever and with lots of page swapping. I hope some of you can give me some advice. Thanks in advance.\n\n",
    "217371": "Many thanks to all. \nI lowered the number of filters (channels) in the model to 5. I'm training with 1280X1280 input images in batches with 10 samples. The number of parameters is low:  \"Trainable params: 190,026.0\" !\nThe job has been running overnight and reached epoch 19. It seems to be training well, the latest performance shows val_dice_loss = 0.9708. Using the top command I can see at most 12% of memory use. (I imagine that is GPU memory)\nI will then face the next hurdle to produce the predictions, but it looks like you helped me a lot.",
    "217129": "Are you saying even with a batch size of 1 you cannot train on a Telsa K80? I'm training a U-net on 1024x1024 images with a batch size of 4 on a 1080Ti 11Gb - So you should manage?\n\nHow many parameters is your U-net? - Mine is approximately 33M."
  }
}