{
  "id": 37681,
  "title": "Trained Unet @ 1918x1280, got LB 99.6 ~ 1 hour per epoch. Drawbacks?",
  "url": "/competitions/carvana-image-masking-challenge/discussion/37681",
  "author_name": "",
  "post_date": "2017-08-07T11:59:13.699985200Z",
  "votes": 22,
  "comment_count": 49,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I trained a Unet with 1920x1280 images, each epoch takes ~1 hour (on a 1080 Ti) but I can only do batch size = 1 since otherwise it runs out of memory. Im using Keras w/ tensorflow.</p>\n\n<p>Has anybody successfully trained full size images with batch size &gt; 1?</p>",
  "messages": [
    {
      "id": "210863",
      "postDate": "08/07/2017 11:59:13",
      "content": "<p>Hi,</p>\n\n<p>I trained a Unet with 1920x1280 images, each epoch takes ~1 hour (on a 1080 Ti) but I can only do batch size = 1 since otherwise it runs out of memory. Im using Keras w/ tensorflow.</p>\n\n<p>Has anybody successfully trained full size images with batch size &gt; 1?</p>",
      "rawMarkdown": "Hi,\n\nI trained a Unet with 1920x1280 images, each epoch takes ~1 hour (on a 1080 Ti) but I can only do batch size = 1 since otherwise it runs out of memory. Im using Keras w/ tensorflow.\n\nHas anybody successfully trained full size images with batch size &gt; 1?",
      "votes": null
    },
    {
      "id": "210866",
      "postDate": "08/07/2017 12:13:18",
      "content": "<p>One major draw back is test set prediction time.\n1 hours per epoch ~= 1 hour for 10k images -&gt; 10 hours for 100k (whole test set)\nNow if you want to do test time augmentation (for example 4 rotations) this time will increase 4-fold to 40 hours. Now with train + test time for each model it is approximately 40 hours + 24 hours for training = 64 hours -&gt; nearly 3 days!</p>",
      "rawMarkdown": "One major draw back is test set prediction time.\n1 hours per epoch ~= 1 hour for 10k images -&gt; 10 hours for 100k (whole test set)\nNow if you want to do test time augmentation (for example 4 rotations) this time will increase 4-fold to 40 hours. Now with train + test time for each model it is approximately 40 hours + 24 hours for training = 64 hours -&gt; nearly 3 days!",
      "votes": null
    },
    {
      "id": "210868",
      "postDate": "08/07/2017 12:22:45",
      "content": "<p>1 hour per epoch is training time. Inference is a bit faster and I can do batch size = 2 for inference. The whole CSV generation process takes ~5 hours.</p>\n\n<p>Why would you want to do augmentation for the test set?</p>",
      "rawMarkdown": "1 hour per epoch is training time. Inference is a bit faster and I can do batch size = 2 for inference. The whole CSV generation process takes ~5 hours.\n\nWhy would you want to do augmentation for the test set?",
      "votes": null
    },
    {
      "id": "210876",
      "postDate": "08/07/2017 12:56:12",
      "content": "<p>Yea, I already used 50% for inference time (better be conservative with estimates)\nIt is general practice for classification to get better scores by using test time augmentation (TTA).\nEssentially you are showing the classifier the same image but transformed and obviously if these transformations are lossless (like rotation for example) we should get the same result as non transformed for a perfect network.\nSince our networks are not perfect TTA may help to reduce prediction variance!</p>\n\n<p>For this challenge the transformations used for TTA may be very limited, since cars and segmentation are not invariant to a lot of transformation!</p>",
      "rawMarkdown": "Yea, I already used 50% for inference time (better be conservative with estimates)\nIt is general practice for classification to get better scores by using test time augmentation (TTA).\nEssentially you are showing the classifier the same image but transformed and obviously if these transformations are lossless (like rotation for example) we should get the same result as non transformed for a perfect network.\nSince our networks are not perfect TTA may help to reduce prediction variance!\n\nFor this challenge the transformations used for TTA may be very limited, since cars and segmentation are not invariant to a lot of transformation!",
      "votes": null
    },
    {
      "id": "210879",
      "postDate": "08/07/2017 13:09:13",
      "content": "<p>This was already mentioned in a different discussion topic, but still relevant:\n<a href=\"https://github.com/fchollet/keras/issues/3556\">https://github.com/fchollet/keras/issues/3556</a></p>\n\n<p>In the code segments, two different users formulated a way to stall parameter updates for a certain number of mini-batches. That way, you could artificially increase your batch size without exceeding memory limitations set by Keras. </p>",
      "rawMarkdown": "This was already mentioned in a different discussion topic, but still relevant:\nhttps://github.com/fchollet/keras/issues/3556\n\nIn the code segments, two different users formulated a way to stall parameter updates for a certain number of mini-batches. That way, you could artificially increase your batch size without exceeding memory limitations set by Keras.",
      "votes": null
    },
    {
      "id": "210915",
      "postDate": "08/07/2017 15:54:52",
      "content": "<p>Yeah, especially considering that any transformation (except multiple integer scaling and mirroring) requires interpolation and the mask will cease being 0 or 1. </p>\n\n<p>I am doing tiny train-time augmentation and I am rounding up the transformed mask to closest integer (0 or 1). I need to do a proper experiment to see if  it is better to round transformed masks or not.</p>",
      "rawMarkdown": "Yeah, especially considering that any transformation (except multiple integer scaling and mirroring) requires interpolation and the mask will cease being 0 or 1. \n\nI am doing tiny train-time augmentation and I am rounding up the transformed mask to closest integer (0 or 1). I need to do a proper experiment to see if  it is better to round transformed masks or not.",
      "votes": null
    },
    {
      "id": "210992",
      "postDate": "08/07/2017 19:10:56",
      "content": "<p>How many upsampling/downsampling parts do you have in your network or how large is the downsampled representation? </p>\n\n<p>I tried the same in pytorch and at least batch size 5 worked well for me on a 1080 Ti. But I did not just train the u-net. On the same GPU I also trained some other networks for post- and pre-processing at the same time. Usually it should also work for bigger batch sizes if I would only train the u-net </p>",
      "rawMarkdown": "How many upsampling/downsampling parts do you have in your network or how large is the downsampled representation? \n\nI tried the same in pytorch and at least batch size 5 worked well for me on a 1080 Ti. But I did not just train the u-net. On the same GPU I also trained some other networks for post- and pre-processing at the same time. Usually it should also work for bigger batch sizes if I would only train the u-net",
      "votes": null
    },
    {
      "id": "210996",
      "postDate": "08/07/2017 19:19:55",
      "content": "<p>Just curious how much memory you need to load all training images @ 1918x1280.</p>",
      "rawMarkdown": "Just curious how much memory you need to load all training images @ 1918x1280.",
      "votes": null
    },
    {
      "id": "211016",
      "postDate": "08/07/2017 19:57:51",
      "content": "<p>I do not load them all together. We've written an own dataloader to load them on the fly which may cause a longer train time/epoch but avoids memory consumption. We always preload the next batch at CPU-RAM  while processing weight-update at GPU but at the GPU we only have one batch at once loaded</p>",
      "rawMarkdown": "I do not load them all together. We've written an own dataloader to load them on the fly which may cause a longer train time/epoch but avoids memory consumption. We always preload the next batch at CPU-RAM  while processing weight-update at GPU but at the GPU we only have one batch at once loaded",
      "votes": null
    },
    {
      "id": "211025",
      "postDate": "08/07/2017 20:51:07",
      "content": "<p>I'm using this UNet, the downsampled representation is 80x120x512.</p>\n\n<pre><code>inputs = Input(input_shape)\nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\npool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n\nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\npool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n\nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\npool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n\nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\npool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n\nconv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\nconv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n</code></pre>\n\n<p>So: input 1280x1920 @ conv1 -&gt; 640x 960  @ conv2  -&gt; 320x480 @ conv3 -&gt; 160x240 @ conv4 -&gt; 80x120 @conv5</p>",
      "rawMarkdown": "I'm using this UNet, the downsampled representation is 80x120x512.\n\n    inputs = Input(input_shape)\n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\n    pool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n\n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\n    pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n\n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\n    pool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n\n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\n    pool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n\n    conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\n    conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n\nSo: input 1280x1920 @ conv1 -&gt; 640x 960  @ conv2  -&gt; 320x480 @ conv3 -&gt; 160x240 @ conv4 -&gt; 80x120 @conv5",
      "votes": null
    },
    {
      "id": "211029",
      "postDate": "08/07/2017 20:54:15",
      "content": "<p>Thanks. I tried this just now but in one epoch is still converges slowly... </p>",
      "rawMarkdown": "Thanks. I tried this just now but in one epoch is still converges slowly...",
      "votes": null
    },
    {
      "id": "211037",
      "postDate": "08/07/2017 21:42:30",
      "content": "<p>Glad to hear you were able to implement it! And would your setup allow for BatchNormalization after each Conv2D layer? It demands quite some memory, but I would expect faster convergence (even with batch_size set at 1).</p>",
      "rawMarkdown": "Glad to hear you were able to implement it! And would your setup allow for BatchNormalization after each Conv2D layer? It demands quite some memory, but I would expect faster convergence (even with batch_size set at 1).",
      "votes": null
    },
    {
      "id": "211084",
      "postDate": "08/08/2017 01:29:50",
      "content": "<p>@Justus Great idea! Not sure how to implement this. Any guide or reference? Thanks!</p>",
      "rawMarkdown": "Justus Great idea! Not sure how to implement this. Any guide or reference? Thanks!",
      "votes": null
    },
    {
      "id": "211126",
      "postDate": "08/08/2017 05:19:14",
      "content": "<p>If using Keras with a <code>generator</code> and then use <code>fit_generator</code>.</p>",
      "rawMarkdown": "If using Keras with a `generator` and then use `fit_generator`.",
      "votes": null
    },
    {
      "id": "211169",
      "postDate": "08/08/2017 07:46:48",
      "content": "<p>In pytorch there is a dataloader implemented which we used as baseclass for our own </p>",
      "rawMarkdown": "In pytorch there is a dataloader implemented which we used as baseclass for our own",
      "votes": null
    },
    {
      "id": "211172",
      "postDate": "08/08/2017 07:52:23",
      "content": "<p>And how many parameters do you have? \nI go down to 30x20 and have about 70million parameters. And training with batchsize 5 needs 7 GB GPU-RAM. You should have less parameter. Do you use Theano or tensorflow as backend for keras? Because tensorflow allocates the maximum amount of storage. Actually I don't know why it dies not work with batchsize on your case. My first guess was to remove BN layers but there are no BNs in your code. </p>",
      "rawMarkdown": "And how many parameters do you have? \nI go down to 30x20 and have about 70million parameters. And training with batchsize 5 needs 7 GB GPU-RAM. You should have less parameter. Do you use Theano or tensorflow as backend for keras? Because tensorflow allocates the maximum amount of storage. Actually I don't know why it dies not work with batchsize on your case. My first guess was to remove BN layers but there are no BNs in your code.",
      "votes": null
    },
    {
      "id": "211178",
      "postDate": "08/08/2017 08:06:13",
      "content": "<p>And as far as I remember tensorflow has a function to approximate the whole network with uint8 instead of float32. I don't know how much accuracy this will drop but it should decrease the necessary GPU memory by a factor of 4</p>",
      "rawMarkdown": "And as far as I remember tensorflow has a function to approximate the whole network with uint8 instead of float32. I don't know how much accuracy this will drop but it should decrease the necessary GPU memory by a factor of 4",
      "votes": null
    },
    {
      "id": "211185",
      "postDate": "08/08/2017 08:38:44",
      "content": "<p>It has ~7 million parameters, and I don't have any batch norm.\nStacking 3x3 convolutions great in terms of expressive power (more non-linearities) and computing efficiency but for big images it's a hog memory-wise,</p>\n\n<pre><code>inputs = Input(input_shape)\nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\n**-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\n**-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \npool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n**-&gt; 960*640*32 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n\nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\n**-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\n**-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \npool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n**-&gt; 480*320*64 * 4 bytes (float 32) =&gt; 40 Megs for activations** \n\nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\n**-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\n**-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \npool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n**-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n\nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\n**-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\n**-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \npool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n**-&gt; 120*80*128 * 4 bytes (float 32) =&gt; 5 Megs for activations** \n\n conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\n**-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \nconv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n **-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n</code></pre>\n\n<p>So activations alone, the Unet will be: (314+314+79+157+157+40+79+79+20+39+39+5)*2 + 20 + 20 ~= 2. 6 Gb... I understand during training memory consumption doubles (<a href=\"https://www.youtube.com/watch?v=83bMCcPmFvE&amp;feature=youtu.be&amp;t=1m38s\">https://www.youtube.com/watch?v=83bMCcPmFvE&amp;feature=youtu.be&amp;t=1m38s</a>) 5.2 Gbytes. Assuming the numbers are right batch size =2 is too tight to fit on GPU (11 Gb).</p>",
      "rawMarkdown": "It has ~7 million parameters, and I don't have any batch norm.\nStacking 3x3 convolutions great in terms of expressive power (more non-linearities) and computing efficiency but for big images it's a hog memory-wise,\n\n    inputs = Input(input_shape)\n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\n    **-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\n    **-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \n    pool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n    **-&gt; 960*640*32 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n   \n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\n    **-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\n    **-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \n    pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n    **-&gt; 480*320*64 * 4 bytes (float 32) =&gt; 40 Megs for activations** \n\n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\n    **-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\n    **-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n    pool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n    **-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n    \n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\n    **-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\n    **-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \n    pool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n    **-&gt; 120*80*128 * 4 bytes (float 32) =&gt; 5 Megs for activations** \n\n     conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\n    **-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n    conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n     **-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n\nSo activations alone, the Unet will be: (314+314+79+157+157+40+79+79+20+39+39+5)*2 + 20 + 20 ~= 2. 6 Gb... I understand during training memory consumption doubles (https://www.youtube.com/watch?v=83bMCcPmFvE&amp;feature=youtu.be&amp;t=1m38s) 5.2 Gbytes. Assuming the numbers are right batch size =2 is too tight to fit on GPU (11 Gb).",
      "votes": null
    },
    {
      "id": "211194",
      "postDate": "08/08/2017 09:26:01",
      "content": "<p>I don't know whether the consumption doubles for the complete batchsize or only for single batch and if this is a constant offset for every batchsize eg. for the optimizer etc. If this is the case you should be able to train with more than batchsize 2. I think this is an implementation detail either of keras but more probably of the backend you are using.</p>",
      "rawMarkdown": "I don't know whether the consumption doubles for the complete batchsize or only for single batch and if this is a constant offset for every batchsize eg. for the optimizer etc. If this is the case you should be able to train with more than batchsize 2. I think this is an implementation detail either of keras but more probably of the backend you are using.",
      "votes": null
    },
    {
      "id": "211210",
      "postDate": "08/08/2017 10:13:58",
      "content": "<p>In tensorflow i'm using two queues with coordinate random seed, somenthing like this:</p>\n\n<pre><code>image_queue=tf.train.string_input_producer(\n    tf.constant(train_images_paths),\n    num_epochs=NUM_EPOCHS,\n    shuffle=True,\n    capacity=QUEUE_CAPACITY,\n    seed=QUEUE_SEED\n)\n\nmask_queue=tf.train.string_input_producer(\n    tf.constant(train_masks_paths),\n    num_epochs=NUM_EPOCHS,\n    shuffle=True,\n    capacity=QUEUE_CAPACITY,\n    seed=QUEUE_SEED\n)\n</code></pre>\n\n<p>Anyway i'm try to preprocess my data and than save image and mask in TFRecords format to speed-up training, but i'm figuring some difficult (i'm new to this low-level, necessary for real problem, mechanisms).</p>",
      "rawMarkdown": "In tensorflow i'm using two queues with coordinate random seed, somenthing like this:\n\n\n    image_queue=tf.train.string_input_producer(\n        tf.constant(train_images_paths),\n        num_epochs=NUM_EPOCHS,\n        shuffle=True,\n        capacity=QUEUE_CAPACITY,\n        seed=QUEUE_SEED\n    )\n    \n    mask_queue=tf.train.string_input_producer(\n        tf.constant(train_masks_paths),\n        num_epochs=NUM_EPOCHS,\n        shuffle=True,\n        capacity=QUEUE_CAPACITY,\n        seed=QUEUE_SEED\n    )\n    \n \nAnyway i'm try to preprocess my data and than save image and mask in TFRecords format to speed-up training, but i'm figuring some difficult (i'm new to this low-level, necessary for real problem, mechanisms).",
      "votes": null
    },
    {
      "id": "211263",
      "postDate": "08/08/2017 13:27:01",
      "content": "<p>Thanks so much guys! I will try it in Keras.</p>",
      "rawMarkdown": "Thanks so much guys! I will try it in Keras.",
      "votes": null
    },
    {
      "id": "211304",
      "postDate": "08/08/2017 15:00:58",
      "content": "<p>@Little Monkey, you can write a Keras generator (it should yield X, y tuple per batch) in the pure python:</p>\n\n<pre><code>def generator():\n    while True:\n        images, masks = load_on_the_fly()\n        yield images, masks\n</code></pre>",
      "rawMarkdown": "Little Monkey, you can write a Keras generator (it should yield X, y tuple per batch) in the pure python:\n\n\n    def generator():\n        while True:\n            images, masks = load_on_the_fly()\n            yield images, masks",
      "votes": null
    },
    {
      "id": "211356",
      "postDate": "08/08/2017 17:52:03",
      "content": "<p>Thanks! @Dmitry Pranchuk </p>",
      "rawMarkdown": "Thanks! @Dmitry Pranchuk",
      "votes": null
    },
    {
      "id": "211368",
      "postDate": "08/08/2017 18:43:04",
      "content": "<p>Just found a concrete example to use generator with keras - <a href=\"https://github.com/fchollet/keras/issues/1627\">https://github.com/fchollet/keras/issues/1627</a> Just in case someone else needs it.</p>",
      "rawMarkdown": "Just found a concrete example to use generator with keras - https://github.com/fchollet/keras/issues/1627 Just in case someone else needs it.",
      "votes": null
    },
    {
      "id": "211701",
      "postDate": "08/09/2017 18:54:20",
      "content": "<p>Im training a modified UNet now, I thought I would be fun to post some intermediary results. </p>\n\n<p>While the simple Unet  has blurry/shadowy areas this one has more non-linearities and while distorted shows wavy contrast.</p>",
      "rawMarkdown": "Im training a modified UNet now, I thought I would be fun to post some intermediary results. \n\nWhile the simple Unet  has blurry/shadowy areas this one has more non-linearities and while distorted shows wavy contrast.",
      "votes": null
    },
    {
      "id": "211763",
      "postDate": "08/09/2017 22:02:45",
      "content": "<p>pretty cool picture</p>",
      "rawMarkdown": "pretty cool picture",
      "votes": null
    },
    {
      "id": "211765",
      "postDate": "08/09/2017 22:04:10",
      "content": "<p>I trained a simple full resolution u-net days ago and the predictions had no blurry areas. Which loss function did you use? </p>",
      "rawMarkdown": "I trained a simple full resolution u-net days ago and the predictions had no blurry areas. Which loss function did you use?",
      "votes": null
    },
    {
      "id": "211881",
      "postDate": "08/10/2017 06:16:54",
      "content": "<p>I doubt how many epochs you train？My UNet down to 15X10X2048 and have <strong>Total params: 124,497,457</strong>. I do batch normalization.</p>",
      "rawMarkdown": "I doubt how many epochs you train？My UNet down to 15X10X2048 and have **Total params: 124,497,457**. I do batch normalization.",
      "votes": null
    },
    {
      "id": "211900",
      "postDate": "08/10/2017 07:20:30",
      "content": "<p>I think this is because you forget to rescale the input? I got the same result when I forgot to rescale test image</p>",
      "rawMarkdown": "I think this is because you forget to rescale the input? I got the same result when I forgot to rescale test image",
      "votes": null
    },
    {
      "id": "211905",
      "postDate": "08/10/2017 07:37:29",
      "content": "<p>Usually at least 25 epochs and if the results look good 50 and more epochs</p>",
      "rawMarkdown": "Usually at least 25 epochs and if the results look good 50 and more epochs",
      "votes": null
    },
    {
      "id": "211906",
      "postDate": "08/10/2017 07:37:38",
      "content": "<p>This is full resolution training/inference.</p>",
      "rawMarkdown": "This is full resolution training/inference.",
      "votes": null
    },
    {
      "id": "211910",
      "postDate": "08/10/2017 07:48:30",
      "content": "<p>The model training is so slow!!! It cost about 3 h per epoch!   If train 25 epochs, then it takes couples of days.....</p>",
      "rawMarkdown": "The model training is so slow!!! It cost about 3 h per epoch!   If train 25 epochs, then it takes couples of days.....",
      "votes": null
    },
    {
      "id": "211916",
      "postDate": "08/10/2017 08:03:20",
      "content": "<p>Intuitively 100+ million parameters seems way too much. Alexnet is ~60 million parameters and had a bigger training set. </p>\n\n<p>My best performer so far is ~7m which to me seems high. I'm testing different approaches.</p>",
      "rawMarkdown": "Intuitively 100+ million parameters seems way too much. Alexnet is ~60 million parameters and had a bigger training set. \n\nMy best performer so far is ~7m which to me seems high. I'm testing different approaches.",
      "votes": null
    },
    {
      "id": "211929",
      "postDate": "08/10/2017 09:12:23",
      "content": "<p>i have made negative experiences using BatchNorm on tiny batches. BatchRenorm might be a way to go.</p>",
      "rawMarkdown": "i have made negative experiences using BatchNorm on tiny batches. BatchRenorm might be a way to go.",
      "votes": null
    },
    {
      "id": "211932",
      "postDate": "08/10/2017 09:20:29",
      "content": "<p>To the best of my knowledge the original U-Net has 31 million parameters. Starting with 32 channels should result in 7.8 million.</p>",
      "rawMarkdown": "To the best of my knowledge the original U-Net has 31 million parameters. Starting with 32 channels should result in 7.8 million.",
      "votes": null
    },
    {
      "id": "211992",
      "postDate": "08/10/2017 12:30:32",
      "content": "<p>I don't think so. I also have about 100 million parameters in my current version. Depends on the implementation details. For example I added some extra layers to increase robustness and these layers have lot's of parameters</p>",
      "rawMarkdown": "I don't think so. I also have about 100 million parameters in my current version. Depends on the implementation details. For example I added some extra layers to increase robustness and these layers have lot's of parameters",
      "votes": null
    },
    {
      "id": "211993",
      "postDate": "08/10/2017 12:31:44",
      "content": "<p>@heroxrq:\ndid you try to remove the BN layers? it may need a few more steps to converge but it should be suspiciously faster.</p>",
      "rawMarkdown": "heroxrq:\ndid you try to remove the BN layers? it may need a few more steps to converge but it should be suspiciously faster.",
      "votes": null
    },
    {
      "id": "212771",
      "postDate": "08/12/2017 16:32:36",
      "content": "<p>@Dmitry Pranchuk follow up question: you still load all images into GPU memory when using generator. Correct? </p>",
      "rawMarkdown": "Dmitry Pranchuk follow up question: you still load all images into GPU memory when using generator. Correct?",
      "votes": null
    },
    {
      "id": "213157",
      "postDate": "08/13/2017 22:15:22",
      "content": "<p>@Little Monkey: Not at once. It depends on the load_on_the_fly function. If this function only loads one batch at once and your code is not very strange this batch should be the only one in gpu memory at one point of time. </p>\n\n<p>See <a href=\"https://stackoverflow.com/questions/231767/what-does-the-yield-keyword-do\">here</a> for more detailed information about yield and generators. </p>",
      "rawMarkdown": "Little Monkey: Not at once. It depends on the load_on_the_fly function. If this function only loads one batch at once and your code is not very strange this batch should be the only one in gpu memory at one point of time. \n\nSee [here][1] for more detailed information about yield and generators. \n\n\n  [1]: https://stackoverflow.com/questions/231767/what-does-the-yield-keyword-do",
      "votes": null
    },
    {
      "id": "213195",
      "postDate": "08/14/2017 02:18:39",
      "content": "<p>@Justus, Thanks! So generator actually helps reduce the pressure for loading all data into CPU memory since it only holds a couple of batches in the memory, while you still need enough GPU memory to do the training. </p>",
      "rawMarkdown": "Justus, Thanks! So generator actually helps reduce the pressure for loading all data into CPU memory since it only holds a couple of batches in the memory, while you still need enough GPU memory to do the training.",
      "votes": null
    },
    {
      "id": "214771",
      "postDate": "08/18/2017 08:05:59",
      "content": "<p>Any chance you might remember how to do \"approximate the whole network with uint8 instead of float32\", Justus?</p>",
      "rawMarkdown": "Any chance you might remember how to do \"approximate the whole network with uint8 instead of float32\", Justus?",
      "votes": null
    },
    {
      "id": "215626",
      "postDate": "08/22/2017 12:23:53",
      "content": "<p>have a look at <a href=\"https://www.tensorflow.org/performance/quantization\">this part of the tensorflow API</a> as a starting point</p>",
      "rawMarkdown": "have a look at [this part of the tensorflow API][1] as a starting point\n\n\n  [1]: https://www.tensorflow.org/performance/quantization",
      "votes": null
    },
    {
      "id": "218652",
      "postDate": "09/05/2017 12:49:46",
      "content": "<p>I also use Keras. But when concatenating two layers, the (640,959,16) layer and the unsample layer whose parameter is (640,960,32) can not be concatenated. How do you solve it?</p>",
      "rawMarkdown": "I also use Keras. But when concatenating two layers, the (640,959,16) layer and the unsample layer whose parameter is (640,960,32) can not be concatenated. How do you solve it?",
      "votes": null
    },
    {
      "id": "218654",
      "postDate": "09/05/2017 12:59:11",
      "content": "<p>I think your (640,959,16) should be (640,960,16), did you use <code>padding='same'</code> ?</p>",
      "rawMarkdown": "I think your (640,959,16) should be (640,960,16), did you use `padding='same'` ?",
      "votes": null
    },
    {
      "id": "218823",
      "postDate": "09/06/2017 00:53:38",
      "content": "<p>My u-net code is </p>\n\n<pre><code>down0b = Conv2D(8, (3, 3), padding='same')(inputs)\ndown0b = BatchNormalization()(down0b)\ndown0b = Activation('relu')(down0b)\ndown0b = Conv2D(8, (3, 3), padding='same')(down0b)\ndown0b = BatchNormalization()(down0b)\ndown0b = Activation('relu')(down0b)\ndown0b_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0b)\n\ndown0a = Conv2D(16, (3, 3), padding='same')(down0b_pool)\ndown0a = BatchNormalization()(down0a)\ndown0a = Activation('relu')(down0a)\ndown0a = Conv2D(16, (3, 3), padding='same')(down0a)\ndown0a = BatchNormalization()(down0a)\ndown0a = Activation('relu')(down0a)\ndown0a_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0a)\n\ndown0 = Conv2D(32, (3, 3), padding='same')(down0a_pool)\ndown0 = BatchNormalization()(down0)\ndown0 = Activation('relu')(down0)\ndown0 = Conv2D(32, (3, 3), padding='same')(down0)\ndown0 = BatchNormalization()(down0)\ndown0 = Activation('relu')(down0)\ndown0_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0)\n\ndown1 = Conv2D(64, (3, 3), padding='same')(down0_pool)\ndown1 = BatchNormalization()(down1)\ndown1 = Activation('relu')(down1)\ndown1 = Conv2D(64, (3, 3), padding='same')(down1)\ndown1 = BatchNormalization()(down1)\ndown1 = Activation('relu')(down1)\ndown1_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down1)\n\ndown2 = Conv2D(128, (3, 3), padding='same')(down1_pool)\ndown2 = BatchNormalization()(down2)\ndown2 = Activation('relu')(down2)\ndown2 = Conv2D(128, (3, 3), padding='same')(down2)\ndown2 = BatchNormalization()(down2)\ndown2 = Activation('relu')(down2)\ndown2_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down2)\n\ndown3 = Conv2D(256, (3, 3), padding='same')(down2_pool)\ndown3 = BatchNormalization()(down3)\ndown3 = Activation('relu')(down3)\ndown3 = Conv2D(256, (3, 3), padding='same')(down3)\ndown3 = BatchNormalization()(down3)\ndown3 = Activation('relu')(down3)\ndown3_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down3)\n\ndown4 = Conv2D(512, (3, 3), padding='same')(down3_pool)\ndown4 = BatchNormalization()(down4)\ndown4 = Activation('relu')(down4)\ndown4 = Conv2D(512, (3, 3), padding='same')(down4)\ndown4 = BatchNormalization()(down4)\ndown4 = Activation('relu')(down4)\ndown4_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down4)\n\ncenter = Conv2D(1024, (3, 3), padding='same')(down4_pool)\ncenter = BatchNormalization()(center)\ncenter = Activation('relu')(center)\ncenter = Conv2D(1024, (3, 3), padding='same')(center)\ncenter = BatchNormalization()(center)\ncenter = Activation('relu')(center)\n\nup4 = UpSampling2D((2, 2))(center)\nup4 = concatenate([down4, up4], axis=3)\nup4 = Conv2D(512, (3, 3), padding='same')(up4)\nup4 = BatchNormalization()(up4)\nup4 = Activation('relu')(up4)\nup4 = Conv2D(512, (3, 3), padding='same')(up4)\nup4 = BatchNormalization()(up4)\nup4 = Activation('relu')(up4)\nup4 = Conv2D(512, (3, 3), padding='same')(up4)\nup4 = BatchNormalization()(up4)\nup4 = Activation('relu')(up4)\n\nup3 = UpSampling2D((2, 2))(up4)\nup3 = concatenate([down3, up3], axis=3)\nup3 = Conv2D(256, (3, 3), padding='same')(up3)\nup3 = BatchNormalization()(up3)\nup3 = Activation('relu')(up3)\nup3 = Conv2D(256, (3, 3), padding='same')(up3)\nup3 = BatchNormalization()(up3)\nup3 = Activation('relu')(up3)\nup3 = Conv2D(256, (3, 3), padding='same')(up3)\nup3 = BatchNormalization()(up3)\nup3 = Activation('relu')(up3)\n\nup2 = UpSampling2D((2, 2))(up3)\nup2 = concatenate([down2, up2], axis=3)\nup2 = Conv2D(128, (3, 3), padding='same')(up2)\nup2 = BatchNormalization()(up2)\nup2 = Activation('relu')(up2)\nup2 = Conv2D(128, (3, 3), padding='same')(up2)\nup2 = BatchNormalization()(up2)\nup2 = Activation('relu')(up2)\nup2 = Conv2D(128, (3, 3), padding='same')(up2)\nup2 = BatchNormalization()(up2)\nup2 = Activation('relu')(up2)\n\nup1 = UpSampling2D((2, 2))(up2)\nup1 = concatenate([down1, up1], axis=3)\nup1 = Conv2D(64, (3, 3), padding='same')(up1)\nup1 = BatchNormalization()(up1)\nup1 = Activation('relu')(up1)\nup1 = Conv2D(64, (3, 3), padding='same')(up1)\nup1 = BatchNormalization()(up1)\nup1 = Activation('relu')(up1)\nup1 = Conv2D(64, (3, 3), padding='same')(up1)\nup1 = BatchNormalization()(up1)\nup1 = Activation('relu')(up1)\n\nup0 = UpSampling2D((2, 2))(up1)\nup0 = concatenate([down0, up0], axis=3)\nup0 = Conv2D(32, (3, 3), padding='same')(up0)\nup0 = BatchNormalization()(up0)\nup0 = Activation('relu')(up0)\nup0 = Conv2D(32, (3, 3), padding='same')(up0)\nup0 = BatchNormalization()(up0)\nup0 = Activation('relu')(up0)\nup0 = Conv2D(32, (3, 3), padding='same')(up0)\nup0 = BatchNormalization()(up0)\nup0 = Activation('relu')(up0)\n\nup0a = UpSampling2D((2, 2))(up0)\nup0a = concatenate([down0a, up0a], axis=3)\nup0a = Conv2D(16, (3, 3), padding='same')(up0a)\nup0a = BatchNormalization()(up0a)\nup0a = Activation('relu')(up0a)\nup0a = Conv2D(16, (3, 3), padding='same')(up0a)\nup0a = BatchNormalization()(up0a)\nup0a = Activation('relu')(up0a)\nup0a = Conv2D(16, (3, 3), padding='same')(up0a)\nup0a = BatchNormalization()(up0a)\nup0a = Activation('relu')(up0a)\n\nup0b = UpSampling2D((2, 2))(up0a)\nup0b = concatenate([down0b, up0b], axis=3)\nup0b = Conv2D(8, (3, 3), padding='same')(up0b)\nup0b = BatchNormalization()(up0b)\nup0b = Activation('relu')(up0b)\nup0b = Conv2D(8, (3, 3), padding='same')(up0b)\nup0b = BatchNormalization()(up0b)\nup0b = Activation('relu')(up0b)\nup0b = Conv2D(8, (3, 3), padding='same')(up0b)\nup0b = BatchNormalization()(up0b)\nup0b = Activation('relu')(up0b)\n</code></pre>\n\n<p>If the original size is 1918x1280,  the output size will be 959x640 after the first maxplooing layer. But the output size is 960x640 after the first unsampling layer. The two layers cannot be concatenated. I can not find where the mistake happened.\nThank you for your advice.</p>",
      "rawMarkdown": "My u-net code is \n\n\n    down0b = Conv2D(8, (3, 3), padding='same')(inputs)\n    down0b = BatchNormalization()(down0b)\n    down0b = Activation('relu')(down0b)\n    down0b = Conv2D(8, (3, 3), padding='same')(down0b)\n    down0b = BatchNormalization()(down0b)\n    down0b = Activation('relu')(down0b)\n    down0b_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0b)\n\n    down0a = Conv2D(16, (3, 3), padding='same')(down0b_pool)\n    down0a = BatchNormalization()(down0a)\n    down0a = Activation('relu')(down0a)\n    down0a = Conv2D(16, (3, 3), padding='same')(down0a)\n    down0a = BatchNormalization()(down0a)\n    down0a = Activation('relu')(down0a)\n    down0a_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0a)\n\n    down0 = Conv2D(32, (3, 3), padding='same')(down0a_pool)\n    down0 = BatchNormalization()(down0)\n    down0 = Activation('relu')(down0)\n    down0 = Conv2D(32, (3, 3), padding='same')(down0)\n    down0 = BatchNormalization()(down0)\n    down0 = Activation('relu')(down0)\n    down0_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0)\n\n    down1 = Conv2D(64, (3, 3), padding='same')(down0_pool)\n    down1 = BatchNormalization()(down1)\n    down1 = Activation('relu')(down1)\n    down1 = Conv2D(64, (3, 3), padding='same')(down1)\n    down1 = BatchNormalization()(down1)\n    down1 = Activation('relu')(down1)\n    down1_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down1)\n\n    down2 = Conv2D(128, (3, 3), padding='same')(down1_pool)\n    down2 = BatchNormalization()(down2)\n    down2 = Activation('relu')(down2)\n    down2 = Conv2D(128, (3, 3), padding='same')(down2)\n    down2 = BatchNormalization()(down2)\n    down2 = Activation('relu')(down2)\n    down2_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down2)\n\n    down3 = Conv2D(256, (3, 3), padding='same')(down2_pool)\n    down3 = BatchNormalization()(down3)\n    down3 = Activation('relu')(down3)\n    down3 = Conv2D(256, (3, 3), padding='same')(down3)\n    down3 = BatchNormalization()(down3)\n    down3 = Activation('relu')(down3)\n    down3_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down3)\n\n    down4 = Conv2D(512, (3, 3), padding='same')(down3_pool)\n    down4 = BatchNormalization()(down4)\n    down4 = Activation('relu')(down4)\n    down4 = Conv2D(512, (3, 3), padding='same')(down4)\n    down4 = BatchNormalization()(down4)\n    down4 = Activation('relu')(down4)\n    down4_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down4)\n\n    center = Conv2D(1024, (3, 3), padding='same')(down4_pool)\n    center = BatchNormalization()(center)\n    center = Activation('relu')(center)\n    center = Conv2D(1024, (3, 3), padding='same')(center)\n    center = BatchNormalization()(center)\n    center = Activation('relu')(center)\n\n    up4 = UpSampling2D((2, 2))(center)\n    up4 = concatenate([down4, up4], axis=3)\n    up4 = Conv2D(512, (3, 3), padding='same')(up4)\n    up4 = BatchNormalization()(up4)\n    up4 = Activation('relu')(up4)\n    up4 = Conv2D(512, (3, 3), padding='same')(up4)\n    up4 = BatchNormalization()(up4)\n    up4 = Activation('relu')(up4)\n    up4 = Conv2D(512, (3, 3), padding='same')(up4)\n    up4 = BatchNormalization()(up4)\n    up4 = Activation('relu')(up4)\n\n    up3 = UpSampling2D((2, 2))(up4)\n    up3 = concatenate([down3, up3], axis=3)\n    up3 = Conv2D(256, (3, 3), padding='same')(up3)\n    up3 = BatchNormalization()(up3)\n    up3 = Activation('relu')(up3)\n    up3 = Conv2D(256, (3, 3), padding='same')(up3)\n    up3 = BatchNormalization()(up3)\n    up3 = Activation('relu')(up3)\n    up3 = Conv2D(256, (3, 3), padding='same')(up3)\n    up3 = BatchNormalization()(up3)\n    up3 = Activation('relu')(up3)\n\n    up2 = UpSampling2D((2, 2))(up3)\n    up2 = concatenate([down2, up2], axis=3)\n    up2 = Conv2D(128, (3, 3), padding='same')(up2)\n    up2 = BatchNormalization()(up2)\n    up2 = Activation('relu')(up2)\n    up2 = Conv2D(128, (3, 3), padding='same')(up2)\n    up2 = BatchNormalization()(up2)\n    up2 = Activation('relu')(up2)\n    up2 = Conv2D(128, (3, 3), padding='same')(up2)\n    up2 = BatchNormalization()(up2)\n    up2 = Activation('relu')(up2)\n\n    up1 = UpSampling2D((2, 2))(up2)\n    up1 = concatenate([down1, up1], axis=3)\n    up1 = Conv2D(64, (3, 3), padding='same')(up1)\n    up1 = BatchNormalization()(up1)\n    up1 = Activation('relu')(up1)\n    up1 = Conv2D(64, (3, 3), padding='same')(up1)\n    up1 = BatchNormalization()(up1)\n    up1 = Activation('relu')(up1)\n    up1 = Conv2D(64, (3, 3), padding='same')(up1)\n    up1 = BatchNormalization()(up1)\n    up1 = Activation('relu')(up1)\n\n    up0 = UpSampling2D((2, 2))(up1)\n    up0 = concatenate([down0, up0], axis=3)\n    up0 = Conv2D(32, (3, 3), padding='same')(up0)\n    up0 = BatchNormalization()(up0)\n    up0 = Activation('relu')(up0)\n    up0 = Conv2D(32, (3, 3), padding='same')(up0)\n    up0 = BatchNormalization()(up0)\n    up0 = Activation('relu')(up0)\n    up0 = Conv2D(32, (3, 3), padding='same')(up0)\n    up0 = BatchNormalization()(up0)\n    up0 = Activation('relu')(up0)\n\n    up0a = UpSampling2D((2, 2))(up0)\n    up0a = concatenate([down0a, up0a], axis=3)\n    up0a = Conv2D(16, (3, 3), padding='same')(up0a)\n    up0a = BatchNormalization()(up0a)\n    up0a = Activation('relu')(up0a)\n    up0a = Conv2D(16, (3, 3), padding='same')(up0a)\n    up0a = BatchNormalization()(up0a)\n    up0a = Activation('relu')(up0a)\n    up0a = Conv2D(16, (3, 3), padding='same')(up0a)\n    up0a = BatchNormalization()(up0a)\n    up0a = Activation('relu')(up0a)\n\n    up0b = UpSampling2D((2, 2))(up0a)\n    up0b = concatenate([down0b, up0b], axis=3)\n    up0b = Conv2D(8, (3, 3), padding='same')(up0b)\n    up0b = BatchNormalization()(up0b)\n    up0b = Activation('relu')(up0b)\n    up0b = Conv2D(8, (3, 3), padding='same')(up0b)\n    up0b = BatchNormalization()(up0b)\n    up0b = Activation('relu')(up0b)\n    up0b = Conv2D(8, (3, 3), padding='same')(up0b)\n    up0b = BatchNormalization()(up0b)\n    up0b = Activation('relu')(up0b)\n\nIf the original size is 1918x1280,  the output size will be 959x640 after the first maxplooing layer. But the output size is 960x640 after the first unsampling layer. The two layers cannot be concatenated. I can not find where the mistake happened.\nThank you for your advice.",
      "votes": null
    },
    {
      "id": "218826",
      "postDate": "09/06/2017 01:22:16",
      "content": "<p>Resize the image to 1920x1280 before feeding it to the network. This way the image will be divisible by 2 through the maxpooling layers. Hope this helps.</p>",
      "rawMarkdown": "Resize the image to 1920x1280 before feeding it to the network. This way the image will be divisible by 2 through the maxpooling layers. Hope this helps.",
      "votes": null
    },
    {
      "id": "218874",
      "postDate": "09/06/2017 06:32:39",
      "content": "<p>Thank you for your advice, I will try it later. By the way, did you do it in this way?</p>",
      "rawMarkdown": "Thank you for your advice, I will try it later. By the way, did you do it in this way?",
      "votes": null
    },
    {
      "id": "218876",
      "postDate": "09/06/2017 06:50:34",
      "content": "<p>I pad with 2px of zeros the 1918 so it becomes 1920. You can do it in the generator or in the net itself.</p>",
      "rawMarkdown": "I pad with 2px of zeros the 1918 so it becomes 1920. You can do it in the generator or in the net itself.",
      "votes": null
    },
    {
      "id": "219112",
      "postDate": "09/07/2017 02:13:15",
      "content": "<p>Then how do you solve it? Using BR layer rather BN layer? Or any other solution? I only have a signal GPU (1080Ti).</p>",
      "rawMarkdown": "Then how do you solve it? Using BR layer rather BN layer? Or any other solution? I only have a signal GPU (1080Ti).",
      "votes": null
    },
    {
      "id": "219125",
      "postDate": "09/07/2017 04:00:53",
      "content": "<p>I am training on half resolution, 960x640. I tried 959x640 with padding same. Didn't work.</p>",
      "rawMarkdown": "I am training on half resolution, 960x640. I tried 959x640 with padding same. Didn't work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 210866,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "08/07/2017 12:13:18",
      "content": "<p>One major draw back is test set prediction time.\n1 hours per epoch ~= 1 hour for 10k images -&gt; 10 hours for 100k (whole test set)\nNow if you want to do test time augmentation (for example 4 rotations) this time will increase 4-fold to 40 hours. Now with train + test time for each model it is approximately 40 hours + 24 hours for training = 64 hours -&gt; nearly 3 days!</p>",
      "votes": null,
      "replies": [
        {
          "id": 210868,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/07/2017 12:22:45",
          "content": "<p>1 hour per epoch is training time. Inference is a bit faster and I can do batch size = 2 for inference. The whole CSV generation process takes ~5 hours.</p>\n\n<p>Why would you want to do augmentation for the test set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210876,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/07/2017 12:56:12",
          "content": "<p>Yea, I already used 50% for inference time (better be conservative with estimates)\nIt is general practice for classification to get better scores by using test time augmentation (TTA).\nEssentially you are showing the classifier the same image but transformed and obviously if these transformations are lossless (like rotation for example) we should get the same result as non transformed for a perfect network.\nSince our networks are not perfect TTA may help to reduce prediction variance!</p>\n\n<p>For this challenge the transformations used for TTA may be very limited, since cars and segmentation are not invariant to a lot of transformation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210915,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/07/2017 15:54:52",
          "content": "<p>Yeah, especially considering that any transformation (except multiple integer scaling and mirroring) requires interpolation and the mask will cease being 0 or 1. </p>\n\n<p>I am doing tiny train-time augmentation and I am rounding up the transformed mask to closest integer (0 or 1). I need to do a proper experiment to see if  it is better to round transformed masks or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 210879,
      "author_name": "rubenhx",
      "author_url": "",
      "post_date": "08/07/2017 13:09:13",
      "content": "<p>This was already mentioned in a different discussion topic, but still relevant:\n<a href=\"https://github.com/fchollet/keras/issues/3556\">https://github.com/fchollet/keras/issues/3556</a></p>\n\n<p>In the code segments, two different users formulated a way to stall parameter updates for a certain number of mini-batches. That way, you could artificially increase your batch size without exceeding memory limitations set by Keras. </p>",
      "votes": null,
      "replies": [
        {
          "id": 211029,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/07/2017 20:54:15",
          "content": "<p>Thanks. I tried this just now but in one epoch is still converges slowly... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211037,
          "author_name": "rubenhx",
          "author_url": "",
          "post_date": "08/07/2017 21:42:30",
          "content": "<p>Glad to hear you were able to implement it! And would your setup allow for BatchNormalization after each Conv2D layer? It demands quite some memory, but I would expect faster convergence (even with batch_size set at 1).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211929,
          "author_name": "gopietz",
          "author_url": "",
          "post_date": "08/10/2017 09:12:23",
          "content": "<p>i have made negative experiences using BatchNorm on tiny batches. BatchRenorm might be a way to go.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 210992,
      "author_name": "pama328",
      "author_url": "",
      "post_date": "08/07/2017 19:10:56",
      "content": "<p>How many upsampling/downsampling parts do you have in your network or how large is the downsampled representation? </p>\n\n<p>I tried the same in pytorch and at least batch size 5 worked well for me on a 1080 Ti. But I did not just train the u-net. On the same GPU I also trained some other networks for post- and pre-processing at the same time. Usually it should also work for bigger batch sizes if I would only train the u-net </p>",
      "votes": null,
      "replies": [
        {
          "id": 211025,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/07/2017 20:51:07",
          "content": "<p>I'm using this UNet, the downsampled representation is 80x120x512.</p>\n\n<pre><code>inputs = Input(input_shape)\nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\npool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n\nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\npool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n\nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\npool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n\nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\npool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n\nconv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\nconv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n</code></pre>\n\n<p>So: input 1280x1920 @ conv1 -&gt; 640x 960  @ conv2  -&gt; 320x480 @ conv3 -&gt; 160x240 @ conv4 -&gt; 80x120 @conv5</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211172,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/08/2017 07:52:23",
          "content": "<p>And how many parameters do you have? \nI go down to 30x20 and have about 70million parameters. And training with batchsize 5 needs 7 GB GPU-RAM. You should have less parameter. Do you use Theano or tensorflow as backend for keras? Because tensorflow allocates the maximum amount of storage. Actually I don't know why it dies not work with batchsize on your case. My first guess was to remove BN layers but there are no BNs in your code. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211178,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/08/2017 08:06:13",
          "content": "<p>And as far as I remember tensorflow has a function to approximate the whole network with uint8 instead of float32. I don't know how much accuracy this will drop but it should decrease the necessary GPU memory by a factor of 4</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211185,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/08/2017 08:38:44",
          "content": "<p>It has ~7 million parameters, and I don't have any batch norm.\nStacking 3x3 convolutions great in terms of expressive power (more non-linearities) and computing efficiency but for big images it's a hog memory-wise,</p>\n\n<pre><code>inputs = Input(input_shape)\nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\n**-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \nconv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\n**-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \npool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n**-&gt; 960*640*32 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n\nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\n**-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \nconv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\n**-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \npool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n**-&gt; 480*320*64 * 4 bytes (float 32) =&gt; 40 Megs for activations** \n\nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\n**-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \nconv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\n**-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \npool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n**-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n\nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\n**-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \nconv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\n**-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \npool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n**-&gt; 120*80*128 * 4 bytes (float 32) =&gt; 5 Megs for activations** \n\n conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\n**-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \nconv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n **-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n</code></pre>\n\n<p>So activations alone, the Unet will be: (314+314+79+157+157+40+79+79+20+39+39+5)*2 + 20 + 20 ~= 2. 6 Gb... I understand during training memory consumption doubles (<a href=\"https://www.youtube.com/watch?v=83bMCcPmFvE&amp;feature=youtu.be&amp;t=1m38s\">https://www.youtube.com/watch?v=83bMCcPmFvE&amp;feature=youtu.be&amp;t=1m38s</a>) 5.2 Gbytes. Assuming the numbers are right batch size =2 is too tight to fit on GPU (11 Gb).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211194,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/08/2017 09:26:01",
          "content": "<p>I don't know whether the consumption doubles for the complete batchsize or only for single batch and if this is a constant offset for every batchsize eg. for the optimizer etc. If this is the case you should be able to train with more than batchsize 2. I think this is an implementation detail either of keras but more probably of the backend you are using.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211932,
          "author_name": "gopietz",
          "author_url": "",
          "post_date": "08/10/2017 09:20:29",
          "content": "<p>To the best of my knowledge the original U-Net has 31 million parameters. Starting with 32 channels should result in 7.8 million.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214771,
          "author_name": "dvdbos",
          "author_url": "",
          "post_date": "08/18/2017 08:05:59",
          "content": "<p>Any chance you might remember how to do \"approximate the whole network with uint8 instead of float32\", Justus?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215626,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/22/2017 12:23:53",
          "content": "<p>have a look at <a href=\"https://www.tensorflow.org/performance/quantization\">this part of the tensorflow API</a> as a starting point</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 210996,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "08/07/2017 19:19:55",
      "content": "<p>Just curious how much memory you need to load all training images @ 1918x1280.</p>",
      "votes": null,
      "replies": [
        {
          "id": 211016,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/07/2017 19:57:51",
          "content": "<p>I do not load them all together. We've written an own dataloader to load them on the fly which may cause a longer train time/epoch but avoids memory consumption. We always preload the next batch at CPU-RAM  while processing weight-update at GPU but at the GPU we only have one batch at once loaded</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211084,
          "author_name": "zhugds",
          "author_url": "",
          "post_date": "08/08/2017 01:29:50",
          "content": "<p>@Justus Great idea! Not sure how to implement this. Any guide or reference? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211126,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/08/2017 05:19:14",
          "content": "<p>If using Keras with a <code>generator</code> and then use <code>fit_generator</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211169,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/08/2017 07:46:48",
          "content": "<p>In pytorch there is a dataloader implemented which we used as baseclass for our own </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211210,
          "author_name": "brasnold",
          "author_url": "",
          "post_date": "08/08/2017 10:13:58",
          "content": "<p>In tensorflow i'm using two queues with coordinate random seed, somenthing like this:</p>\n\n<pre><code>image_queue=tf.train.string_input_producer(\n    tf.constant(train_images_paths),\n    num_epochs=NUM_EPOCHS,\n    shuffle=True,\n    capacity=QUEUE_CAPACITY,\n    seed=QUEUE_SEED\n)\n\nmask_queue=tf.train.string_input_producer(\n    tf.constant(train_masks_paths),\n    num_epochs=NUM_EPOCHS,\n    shuffle=True,\n    capacity=QUEUE_CAPACITY,\n    seed=QUEUE_SEED\n)\n</code></pre>\n\n<p>Anyway i'm try to preprocess my data and than save image and mask in TFRecords format to speed-up training, but i'm figuring some difficult (i'm new to this low-level, necessary for real problem, mechanisms).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211263,
          "author_name": "zhugds",
          "author_url": "",
          "post_date": "08/08/2017 13:27:01",
          "content": "<p>Thanks so much guys! I will try it in Keras.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211304,
          "author_name": "cortwave",
          "author_url": "",
          "post_date": "08/08/2017 15:00:58",
          "content": "<p>@Little Monkey, you can write a Keras generator (it should yield X, y tuple per batch) in the pure python:</p>\n\n<pre><code>def generator():\n    while True:\n        images, masks = load_on_the_fly()\n        yield images, masks\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211356,
          "author_name": "zhugds",
          "author_url": "",
          "post_date": "08/08/2017 17:52:03",
          "content": "<p>Thanks! @Dmitry Pranchuk </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211368,
          "author_name": "zhugds",
          "author_url": "",
          "post_date": "08/08/2017 18:43:04",
          "content": "<p>Just found a concrete example to use generator with keras - <a href=\"https://github.com/fchollet/keras/issues/1627\">https://github.com/fchollet/keras/issues/1627</a> Just in case someone else needs it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212771,
          "author_name": "zhugds",
          "author_url": "",
          "post_date": "08/12/2017 16:32:36",
          "content": "<p>@Dmitry Pranchuk follow up question: you still load all images into GPU memory when using generator. Correct? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213157,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/13/2017 22:15:22",
          "content": "<p>@Little Monkey: Not at once. It depends on the load_on_the_fly function. If this function only loads one batch at once and your code is not very strange this batch should be the only one in gpu memory at one point of time. </p>\n\n<p>See <a href=\"https://stackoverflow.com/questions/231767/what-does-the-yield-keyword-do\">here</a> for more detailed information about yield and generators. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213195,
          "author_name": "zhugds",
          "author_url": "",
          "post_date": "08/14/2017 02:18:39",
          "content": "<p>@Justus, Thanks! So generator actually helps reduce the pressure for loading all data into CPU memory since it only holds a couple of batches in the memory, while you still need enough GPU memory to do the training. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 211701,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "08/09/2017 18:54:20",
      "content": "<p>Im training a modified UNet now, I thought I would be fun to post some intermediary results. </p>\n\n<p>While the simple Unet  has blurry/shadowy areas this one has more non-linearities and while distorted shows wavy contrast.</p>",
      "votes": null,
      "replies": [
        {
          "id": 211763,
          "author_name": "cjansen",
          "author_url": "",
          "post_date": "08/09/2017 22:02:45",
          "content": "<p>pretty cool picture</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211765,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/09/2017 22:04:10",
          "content": "<p>I trained a simple full resolution u-net days ago and the predictions had no blurry areas. Which loss function did you use? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211900,
          "author_name": "viebboy",
          "author_url": "",
          "post_date": "08/10/2017 07:20:30",
          "content": "<p>I think this is because you forget to rescale the input? I got the same result when I forgot to rescale test image</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211906,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/10/2017 07:37:38",
          "content": "<p>This is full resolution training/inference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 211881,
      "author_name": "heroxrq",
      "author_url": "",
      "post_date": "08/10/2017 06:16:54",
      "content": "<p>I doubt how many epochs you train？My UNet down to 15X10X2048 and have <strong>Total params: 124,497,457</strong>. I do batch normalization.</p>",
      "votes": null,
      "replies": [
        {
          "id": 211905,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/10/2017 07:37:29",
          "content": "<p>Usually at least 25 epochs and if the results look good 50 and more epochs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211910,
          "author_name": "heroxrq",
          "author_url": "",
          "post_date": "08/10/2017 07:48:30",
          "content": "<p>The model training is so slow!!! It cost about 3 h per epoch!   If train 25 epochs, then it takes couples of days.....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211916,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/10/2017 08:03:20",
          "content": "<p>Intuitively 100+ million parameters seems way too much. Alexnet is ~60 million parameters and had a bigger training set. </p>\n\n<p>My best performer so far is ~7m which to me seems high. I'm testing different approaches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211992,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/10/2017 12:30:32",
          "content": "<p>I don't think so. I also have about 100 million parameters in my current version. Depends on the implementation details. For example I added some extra layers to increase robustness and these layers have lot's of parameters</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211993,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/10/2017 12:31:44",
          "content": "<p>@heroxrq:\ndid you try to remove the BN layers? it may need a few more steps to converge but it should be suspiciously faster.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 218652,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "09/05/2017 12:49:46",
      "content": "<p>I also use Keras. But when concatenating two layers, the (640,959,16) layer and the unsample layer whose parameter is (640,960,32) can not be concatenated. How do you solve it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 218654,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/05/2017 12:59:11",
          "content": "<p>I think your (640,959,16) should be (640,960,16), did you use <code>padding='same'</code> ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218823,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/06/2017 00:53:38",
          "content": "<p>My u-net code is </p>\n\n<pre><code>down0b = Conv2D(8, (3, 3), padding='same')(inputs)\ndown0b = BatchNormalization()(down0b)\ndown0b = Activation('relu')(down0b)\ndown0b = Conv2D(8, (3, 3), padding='same')(down0b)\ndown0b = BatchNormalization()(down0b)\ndown0b = Activation('relu')(down0b)\ndown0b_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0b)\n\ndown0a = Conv2D(16, (3, 3), padding='same')(down0b_pool)\ndown0a = BatchNormalization()(down0a)\ndown0a = Activation('relu')(down0a)\ndown0a = Conv2D(16, (3, 3), padding='same')(down0a)\ndown0a = BatchNormalization()(down0a)\ndown0a = Activation('relu')(down0a)\ndown0a_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0a)\n\ndown0 = Conv2D(32, (3, 3), padding='same')(down0a_pool)\ndown0 = BatchNormalization()(down0)\ndown0 = Activation('relu')(down0)\ndown0 = Conv2D(32, (3, 3), padding='same')(down0)\ndown0 = BatchNormalization()(down0)\ndown0 = Activation('relu')(down0)\ndown0_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0)\n\ndown1 = Conv2D(64, (3, 3), padding='same')(down0_pool)\ndown1 = BatchNormalization()(down1)\ndown1 = Activation('relu')(down1)\ndown1 = Conv2D(64, (3, 3), padding='same')(down1)\ndown1 = BatchNormalization()(down1)\ndown1 = Activation('relu')(down1)\ndown1_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down1)\n\ndown2 = Conv2D(128, (3, 3), padding='same')(down1_pool)\ndown2 = BatchNormalization()(down2)\ndown2 = Activation('relu')(down2)\ndown2 = Conv2D(128, (3, 3), padding='same')(down2)\ndown2 = BatchNormalization()(down2)\ndown2 = Activation('relu')(down2)\ndown2_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down2)\n\ndown3 = Conv2D(256, (3, 3), padding='same')(down2_pool)\ndown3 = BatchNormalization()(down3)\ndown3 = Activation('relu')(down3)\ndown3 = Conv2D(256, (3, 3), padding='same')(down3)\ndown3 = BatchNormalization()(down3)\ndown3 = Activation('relu')(down3)\ndown3_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down3)\n\ndown4 = Conv2D(512, (3, 3), padding='same')(down3_pool)\ndown4 = BatchNormalization()(down4)\ndown4 = Activation('relu')(down4)\ndown4 = Conv2D(512, (3, 3), padding='same')(down4)\ndown4 = BatchNormalization()(down4)\ndown4 = Activation('relu')(down4)\ndown4_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down4)\n\ncenter = Conv2D(1024, (3, 3), padding='same')(down4_pool)\ncenter = BatchNormalization()(center)\ncenter = Activation('relu')(center)\ncenter = Conv2D(1024, (3, 3), padding='same')(center)\ncenter = BatchNormalization()(center)\ncenter = Activation('relu')(center)\n\nup4 = UpSampling2D((2, 2))(center)\nup4 = concatenate([down4, up4], axis=3)\nup4 = Conv2D(512, (3, 3), padding='same')(up4)\nup4 = BatchNormalization()(up4)\nup4 = Activation('relu')(up4)\nup4 = Conv2D(512, (3, 3), padding='same')(up4)\nup4 = BatchNormalization()(up4)\nup4 = Activation('relu')(up4)\nup4 = Conv2D(512, (3, 3), padding='same')(up4)\nup4 = BatchNormalization()(up4)\nup4 = Activation('relu')(up4)\n\nup3 = UpSampling2D((2, 2))(up4)\nup3 = concatenate([down3, up3], axis=3)\nup3 = Conv2D(256, (3, 3), padding='same')(up3)\nup3 = BatchNormalization()(up3)\nup3 = Activation('relu')(up3)\nup3 = Conv2D(256, (3, 3), padding='same')(up3)\nup3 = BatchNormalization()(up3)\nup3 = Activation('relu')(up3)\nup3 = Conv2D(256, (3, 3), padding='same')(up3)\nup3 = BatchNormalization()(up3)\nup3 = Activation('relu')(up3)\n\nup2 = UpSampling2D((2, 2))(up3)\nup2 = concatenate([down2, up2], axis=3)\nup2 = Conv2D(128, (3, 3), padding='same')(up2)\nup2 = BatchNormalization()(up2)\nup2 = Activation('relu')(up2)\nup2 = Conv2D(128, (3, 3), padding='same')(up2)\nup2 = BatchNormalization()(up2)\nup2 = Activation('relu')(up2)\nup2 = Conv2D(128, (3, 3), padding='same')(up2)\nup2 = BatchNormalization()(up2)\nup2 = Activation('relu')(up2)\n\nup1 = UpSampling2D((2, 2))(up2)\nup1 = concatenate([down1, up1], axis=3)\nup1 = Conv2D(64, (3, 3), padding='same')(up1)\nup1 = BatchNormalization()(up1)\nup1 = Activation('relu')(up1)\nup1 = Conv2D(64, (3, 3), padding='same')(up1)\nup1 = BatchNormalization()(up1)\nup1 = Activation('relu')(up1)\nup1 = Conv2D(64, (3, 3), padding='same')(up1)\nup1 = BatchNormalization()(up1)\nup1 = Activation('relu')(up1)\n\nup0 = UpSampling2D((2, 2))(up1)\nup0 = concatenate([down0, up0], axis=3)\nup0 = Conv2D(32, (3, 3), padding='same')(up0)\nup0 = BatchNormalization()(up0)\nup0 = Activation('relu')(up0)\nup0 = Conv2D(32, (3, 3), padding='same')(up0)\nup0 = BatchNormalization()(up0)\nup0 = Activation('relu')(up0)\nup0 = Conv2D(32, (3, 3), padding='same')(up0)\nup0 = BatchNormalization()(up0)\nup0 = Activation('relu')(up0)\n\nup0a = UpSampling2D((2, 2))(up0)\nup0a = concatenate([down0a, up0a], axis=3)\nup0a = Conv2D(16, (3, 3), padding='same')(up0a)\nup0a = BatchNormalization()(up0a)\nup0a = Activation('relu')(up0a)\nup0a = Conv2D(16, (3, 3), padding='same')(up0a)\nup0a = BatchNormalization()(up0a)\nup0a = Activation('relu')(up0a)\nup0a = Conv2D(16, (3, 3), padding='same')(up0a)\nup0a = BatchNormalization()(up0a)\nup0a = Activation('relu')(up0a)\n\nup0b = UpSampling2D((2, 2))(up0a)\nup0b = concatenate([down0b, up0b], axis=3)\nup0b = Conv2D(8, (3, 3), padding='same')(up0b)\nup0b = BatchNormalization()(up0b)\nup0b = Activation('relu')(up0b)\nup0b = Conv2D(8, (3, 3), padding='same')(up0b)\nup0b = BatchNormalization()(up0b)\nup0b = Activation('relu')(up0b)\nup0b = Conv2D(8, (3, 3), padding='same')(up0b)\nup0b = BatchNormalization()(up0b)\nup0b = Activation('relu')(up0b)\n</code></pre>\n\n<p>If the original size is 1918x1280,  the output size will be 959x640 after the first maxplooing layer. But the output size is 960x640 after the first unsampling layer. The two layers cannot be concatenated. I can not find where the mistake happened.\nThank you for your advice.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218826,
          "author_name": "sgalib",
          "author_url": "",
          "post_date": "09/06/2017 01:22:16",
          "content": "<p>Resize the image to 1920x1280 before feeding it to the network. This way the image will be divisible by 2 through the maxpooling layers. Hope this helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218874,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/06/2017 06:32:39",
          "content": "<p>Thank you for your advice, I will try it later. By the way, did you do it in this way?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 218876,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/06/2017 06:50:34",
          "content": "<p>I pad with 2px of zeros the 1918 so it becomes 1920. You can do it in the generator or in the net itself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 219125,
          "author_name": "sgalib",
          "author_url": "",
          "post_date": "09/07/2017 04:00:53",
          "content": "<p>I am training on half resolution, 960x640. I tried 959x640 with padding same. Didn't work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 219112,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "09/07/2017 02:13:15",
      "content": "<p>Then how do you solve it? Using BR layer rather BN layer? Or any other solution? I only have a signal GPU (1080Ti).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "210863": "Hi,\n\nI trained a Unet with 1920x1280 images, each epoch takes ~1 hour (on a 1080 Ti) but I can only do batch size = 1 since otherwise it runs out of memory. Im using Keras w/ tensorflow.\n\nHas anybody successfully trained full size images with batch size &gt; 1?",
    "210866": "One major draw back is test set prediction time.\n1 hours per epoch ~= 1 hour for 10k images -&gt; 10 hours for 100k (whole test set)\nNow if you want to do test time augmentation (for example 4 rotations) this time will increase 4-fold to 40 hours. Now with train + test time for each model it is approximately 40 hours + 24 hours for training = 64 hours -&gt; nearly 3 days!",
    "210868": "1 hour per epoch is training time. Inference is a bit faster and I can do batch size = 2 for inference. The whole CSV generation process takes ~5 hours.\n\nWhy would you want to do augmentation for the test set?",
    "210876": "Yea, I already used 50% for inference time (better be conservative with estimates)\nIt is general practice for classification to get better scores by using test time augmentation (TTA).\nEssentially you are showing the classifier the same image but transformed and obviously if these transformations are lossless (like rotation for example) we should get the same result as non transformed for a perfect network.\nSince our networks are not perfect TTA may help to reduce prediction variance!\n\nFor this challenge the transformations used for TTA may be very limited, since cars and segmentation are not invariant to a lot of transformation!",
    "210879": "This was already mentioned in a different discussion topic, but still relevant:\nhttps://github.com/fchollet/keras/issues/3556\n\nIn the code segments, two different users formulated a way to stall parameter updates for a certain number of mini-batches. That way, you could artificially increase your batch size without exceeding memory limitations set by Keras.",
    "210915": "Yeah, especially considering that any transformation (except multiple integer scaling and mirroring) requires interpolation and the mask will cease being 0 or 1. \n\nI am doing tiny train-time augmentation and I am rounding up the transformed mask to closest integer (0 or 1). I need to do a proper experiment to see if  it is better to round transformed masks or not.",
    "210992": "How many upsampling/downsampling parts do you have in your network or how large is the downsampled representation? \n\nI tried the same in pytorch and at least batch size 5 worked well for me on a 1080 Ti. But I did not just train the u-net. On the same GPU I also trained some other networks for post- and pre-processing at the same time. Usually it should also work for bigger batch sizes if I would only train the u-net",
    "210996": "Just curious how much memory you need to load all training images @ 1918x1280.",
    "211016": "I do not load them all together. We've written an own dataloader to load them on the fly which may cause a longer train time/epoch but avoids memory consumption. We always preload the next batch at CPU-RAM  while processing weight-update at GPU but at the GPU we only have one batch at once loaded",
    "211025": "I'm using this UNet, the downsampled representation is 80x120x512.\n\n    inputs = Input(input_shape)\n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\n    pool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n\n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\n    pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n\n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\n    pool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n\n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\n    pool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n\n    conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\n    conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n\nSo: input 1280x1920 @ conv1 -&gt; 640x 960  @ conv2  -&gt; 320x480 @ conv3 -&gt; 160x240 @ conv4 -&gt; 80x120 @conv5",
    "211029": "Thanks. I tried this just now but in one epoch is still converges slowly...",
    "211037": "Glad to hear you were able to implement it! And would your setup allow for BatchNormalization after each Conv2D layer? It demands quite some memory, but I would expect faster convergence (even with batch_size set at 1).",
    "211084": "Justus Great idea! Not sure how to implement this. Any guide or reference? Thanks!",
    "211126": "If using Keras with a `generator` and then use `fit_generator`.",
    "211169": "In pytorch there is a dataloader implemented which we used as baseclass for our own",
    "211172": "And how many parameters do you have? \nI go down to 30x20 and have about 70million parameters. And training with batchsize 5 needs 7 GB GPU-RAM. You should have less parameter. Do you use Theano or tensorflow as backend for keras? Because tensorflow allocates the maximum amount of storage. Actually I don't know why it dies not work with batchsize on your case. My first guess was to remove BN layers but there are no BNs in your code.",
    "211178": "And as far as I remember tensorflow has a function to approximate the whole network with uint8 instead of float32. I don't know how much accuracy this will drop but it should decrease the necessary GPU memory by a factor of 4",
    "211185": "It has ~7 million parameters, and I don't have any batch norm.\nStacking 3x3 convolutions great in terms of expressive power (more non-linearities) and computing efficiency but for big images it's a hog memory-wise,\n\n    inputs = Input(input_shape)\n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(inputs)\n    **-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \n    conv1 = Conv2D(32, (3, 3), activation=int_activation, padding='same')(conv1)\n    **-&gt; 1920*1280*32 * 4 bytes (float 32) =&gt; 314 Megs for activations** \n    pool1 = MaxPooling2D(pool_size=(2, 2))(conv1)\n    **-&gt; 960*640*32 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n   \n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(pool1)\n    **-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \n    conv2 = Conv2D(64, (3, 3), activation=int_activation, padding='same')(conv2)\n    **-&gt; 960*640*64 * 4 bytes (float 32) =&gt; 157 Megs for activations** \n    pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)\n    **-&gt; 480*320*64 * 4 bytes (float 32) =&gt; 40 Megs for activations** \n\n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(pool2)\n    **-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n    conv3 = Conv2D(128, (3, 3), activation=int_activation, padding='same')(conv3)\n    **-&gt; 480*320*128 * 4 bytes (float 32) =&gt; 79 Megs for activations** \n    pool3 = MaxPooling2D(pool_size=(2, 2))(conv3)\n    **-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n    \n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(pool3)\n    **-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \n    conv4 = Conv2D(256, (3, 3), activation=int_activation, padding='same')(conv4)\n    **-&gt; 240*160*128 * 4 bytes (float 32) =&gt; 39 Megs for activations** \n    pool4 = MaxPooling2D(pool_size=(2, 2))(conv4)\n    **-&gt; 120*80*128 * 4 bytes (float 32) =&gt; 5 Megs for activations** \n\n     conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(pool4)\n    **-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n    conv5 = Conv2D(512, (3, 3), activation=int_activation, padding='same')(conv5)\n     **-&gt; 120*80*512 * 4 bytes (float 32) =&gt; 20 Megs for activations** \n\nSo activations alone, the Unet will be: (314+314+79+157+157+40+79+79+20+39+39+5)*2 + 20 + 20 ~= 2. 6 Gb... I understand during training memory consumption doubles (https://www.youtube.com/watch?v=83bMCcPmFvE&amp;feature=youtu.be&amp;t=1m38s) 5.2 Gbytes. Assuming the numbers are right batch size =2 is too tight to fit on GPU (11 Gb).",
    "211194": "I don't know whether the consumption doubles for the complete batchsize or only for single batch and if this is a constant offset for every batchsize eg. for the optimizer etc. If this is the case you should be able to train with more than batchsize 2. I think this is an implementation detail either of keras but more probably of the backend you are using.",
    "211210": "In tensorflow i'm using two queues with coordinate random seed, somenthing like this:\n\n\n    image_queue=tf.train.string_input_producer(\n        tf.constant(train_images_paths),\n        num_epochs=NUM_EPOCHS,\n        shuffle=True,\n        capacity=QUEUE_CAPACITY,\n        seed=QUEUE_SEED\n    )\n    \n    mask_queue=tf.train.string_input_producer(\n        tf.constant(train_masks_paths),\n        num_epochs=NUM_EPOCHS,\n        shuffle=True,\n        capacity=QUEUE_CAPACITY,\n        seed=QUEUE_SEED\n    )\n    \n \nAnyway i'm try to preprocess my data and than save image and mask in TFRecords format to speed-up training, but i'm figuring some difficult (i'm new to this low-level, necessary for real problem, mechanisms).",
    "211263": "Thanks so much guys! I will try it in Keras.",
    "211304": "Little Monkey, you can write a Keras generator (it should yield X, y tuple per batch) in the pure python:\n\n\n    def generator():\n        while True:\n            images, masks = load_on_the_fly()\n            yield images, masks",
    "211356": "Thanks! @Dmitry Pranchuk",
    "211368": "Just found a concrete example to use generator with keras - https://github.com/fchollet/keras/issues/1627 Just in case someone else needs it.",
    "211701": "Im training a modified UNet now, I thought I would be fun to post some intermediary results. \n\nWhile the simple Unet  has blurry/shadowy areas this one has more non-linearities and while distorted shows wavy contrast.",
    "211763": "pretty cool picture",
    "211765": "I trained a simple full resolution u-net days ago and the predictions had no blurry areas. Which loss function did you use?",
    "211881": "I doubt how many epochs you train？My UNet down to 15X10X2048 and have **Total params: 124,497,457**. I do batch normalization.",
    "211900": "I think this is because you forget to rescale the input? I got the same result when I forgot to rescale test image",
    "211905": "Usually at least 25 epochs and if the results look good 50 and more epochs",
    "211906": "This is full resolution training/inference.",
    "211910": "The model training is so slow!!! It cost about 3 h per epoch!   If train 25 epochs, then it takes couples of days.....",
    "211916": "Intuitively 100+ million parameters seems way too much. Alexnet is ~60 million parameters and had a bigger training set. \n\nMy best performer so far is ~7m which to me seems high. I'm testing different approaches.",
    "211929": "i have made negative experiences using BatchNorm on tiny batches. BatchRenorm might be a way to go.",
    "211932": "To the best of my knowledge the original U-Net has 31 million parameters. Starting with 32 channels should result in 7.8 million.",
    "211992": "I don't think so. I also have about 100 million parameters in my current version. Depends on the implementation details. For example I added some extra layers to increase robustness and these layers have lot's of parameters",
    "211993": "heroxrq:\ndid you try to remove the BN layers? it may need a few more steps to converge but it should be suspiciously faster.",
    "212771": "Dmitry Pranchuk follow up question: you still load all images into GPU memory when using generator. Correct?",
    "213157": "Little Monkey: Not at once. It depends on the load_on_the_fly function. If this function only loads one batch at once and your code is not very strange this batch should be the only one in gpu memory at one point of time. \n\nSee [here][1] for more detailed information about yield and generators. \n\n\n  [1]: https://stackoverflow.com/questions/231767/what-does-the-yield-keyword-do",
    "213195": "Justus, Thanks! So generator actually helps reduce the pressure for loading all data into CPU memory since it only holds a couple of batches in the memory, while you still need enough GPU memory to do the training.",
    "214771": "Any chance you might remember how to do \"approximate the whole network with uint8 instead of float32\", Justus?",
    "215626": "have a look at [this part of the tensorflow API][1] as a starting point\n\n\n  [1]: https://www.tensorflow.org/performance/quantization",
    "218652": "I also use Keras. But when concatenating two layers, the (640,959,16) layer and the unsample layer whose parameter is (640,960,32) can not be concatenated. How do you solve it?",
    "218654": "I think your (640,959,16) should be (640,960,16), did you use `padding='same'` ?",
    "218823": "My u-net code is \n\n\n    down0b = Conv2D(8, (3, 3), padding='same')(inputs)\n    down0b = BatchNormalization()(down0b)\n    down0b = Activation('relu')(down0b)\n    down0b = Conv2D(8, (3, 3), padding='same')(down0b)\n    down0b = BatchNormalization()(down0b)\n    down0b = Activation('relu')(down0b)\n    down0b_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0b)\n\n    down0a = Conv2D(16, (3, 3), padding='same')(down0b_pool)\n    down0a = BatchNormalization()(down0a)\n    down0a = Activation('relu')(down0a)\n    down0a = Conv2D(16, (3, 3), padding='same')(down0a)\n    down0a = BatchNormalization()(down0a)\n    down0a = Activation('relu')(down0a)\n    down0a_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0a)\n\n    down0 = Conv2D(32, (3, 3), padding='same')(down0a_pool)\n    down0 = BatchNormalization()(down0)\n    down0 = Activation('relu')(down0)\n    down0 = Conv2D(32, (3, 3), padding='same')(down0)\n    down0 = BatchNormalization()(down0)\n    down0 = Activation('relu')(down0)\n    down0_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down0)\n\n    down1 = Conv2D(64, (3, 3), padding='same')(down0_pool)\n    down1 = BatchNormalization()(down1)\n    down1 = Activation('relu')(down1)\n    down1 = Conv2D(64, (3, 3), padding='same')(down1)\n    down1 = BatchNormalization()(down1)\n    down1 = Activation('relu')(down1)\n    down1_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down1)\n\n    down2 = Conv2D(128, (3, 3), padding='same')(down1_pool)\n    down2 = BatchNormalization()(down2)\n    down2 = Activation('relu')(down2)\n    down2 = Conv2D(128, (3, 3), padding='same')(down2)\n    down2 = BatchNormalization()(down2)\n    down2 = Activation('relu')(down2)\n    down2_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down2)\n\n    down3 = Conv2D(256, (3, 3), padding='same')(down2_pool)\n    down3 = BatchNormalization()(down3)\n    down3 = Activation('relu')(down3)\n    down3 = Conv2D(256, (3, 3), padding='same')(down3)\n    down3 = BatchNormalization()(down3)\n    down3 = Activation('relu')(down3)\n    down3_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down3)\n\n    down4 = Conv2D(512, (3, 3), padding='same')(down3_pool)\n    down4 = BatchNormalization()(down4)\n    down4 = Activation('relu')(down4)\n    down4 = Conv2D(512, (3, 3), padding='same')(down4)\n    down4 = BatchNormalization()(down4)\n    down4 = Activation('relu')(down4)\n    down4_pool = MaxPooling2D((2, 2), strides=(2, 2),padding='same')(down4)\n\n    center = Conv2D(1024, (3, 3), padding='same')(down4_pool)\n    center = BatchNormalization()(center)\n    center = Activation('relu')(center)\n    center = Conv2D(1024, (3, 3), padding='same')(center)\n    center = BatchNormalization()(center)\n    center = Activation('relu')(center)\n\n    up4 = UpSampling2D((2, 2))(center)\n    up4 = concatenate([down4, up4], axis=3)\n    up4 = Conv2D(512, (3, 3), padding='same')(up4)\n    up4 = BatchNormalization()(up4)\n    up4 = Activation('relu')(up4)\n    up4 = Conv2D(512, (3, 3), padding='same')(up4)\n    up4 = BatchNormalization()(up4)\n    up4 = Activation('relu')(up4)\n    up4 = Conv2D(512, (3, 3), padding='same')(up4)\n    up4 = BatchNormalization()(up4)\n    up4 = Activation('relu')(up4)\n\n    up3 = UpSampling2D((2, 2))(up4)\n    up3 = concatenate([down3, up3], axis=3)\n    up3 = Conv2D(256, (3, 3), padding='same')(up3)\n    up3 = BatchNormalization()(up3)\n    up3 = Activation('relu')(up3)\n    up3 = Conv2D(256, (3, 3), padding='same')(up3)\n    up3 = BatchNormalization()(up3)\n    up3 = Activation('relu')(up3)\n    up3 = Conv2D(256, (3, 3), padding='same')(up3)\n    up3 = BatchNormalization()(up3)\n    up3 = Activation('relu')(up3)\n\n    up2 = UpSampling2D((2, 2))(up3)\n    up2 = concatenate([down2, up2], axis=3)\n    up2 = Conv2D(128, (3, 3), padding='same')(up2)\n    up2 = BatchNormalization()(up2)\n    up2 = Activation('relu')(up2)\n    up2 = Conv2D(128, (3, 3), padding='same')(up2)\n    up2 = BatchNormalization()(up2)\n    up2 = Activation('relu')(up2)\n    up2 = Conv2D(128, (3, 3), padding='same')(up2)\n    up2 = BatchNormalization()(up2)\n    up2 = Activation('relu')(up2)\n\n    up1 = UpSampling2D((2, 2))(up2)\n    up1 = concatenate([down1, up1], axis=3)\n    up1 = Conv2D(64, (3, 3), padding='same')(up1)\n    up1 = BatchNormalization()(up1)\n    up1 = Activation('relu')(up1)\n    up1 = Conv2D(64, (3, 3), padding='same')(up1)\n    up1 = BatchNormalization()(up1)\n    up1 = Activation('relu')(up1)\n    up1 = Conv2D(64, (3, 3), padding='same')(up1)\n    up1 = BatchNormalization()(up1)\n    up1 = Activation('relu')(up1)\n\n    up0 = UpSampling2D((2, 2))(up1)\n    up0 = concatenate([down0, up0], axis=3)\n    up0 = Conv2D(32, (3, 3), padding='same')(up0)\n    up0 = BatchNormalization()(up0)\n    up0 = Activation('relu')(up0)\n    up0 = Conv2D(32, (3, 3), padding='same')(up0)\n    up0 = BatchNormalization()(up0)\n    up0 = Activation('relu')(up0)\n    up0 = Conv2D(32, (3, 3), padding='same')(up0)\n    up0 = BatchNormalization()(up0)\n    up0 = Activation('relu')(up0)\n\n    up0a = UpSampling2D((2, 2))(up0)\n    up0a = concatenate([down0a, up0a], axis=3)\n    up0a = Conv2D(16, (3, 3), padding='same')(up0a)\n    up0a = BatchNormalization()(up0a)\n    up0a = Activation('relu')(up0a)\n    up0a = Conv2D(16, (3, 3), padding='same')(up0a)\n    up0a = BatchNormalization()(up0a)\n    up0a = Activation('relu')(up0a)\n    up0a = Conv2D(16, (3, 3), padding='same')(up0a)\n    up0a = BatchNormalization()(up0a)\n    up0a = Activation('relu')(up0a)\n\n    up0b = UpSampling2D((2, 2))(up0a)\n    up0b = concatenate([down0b, up0b], axis=3)\n    up0b = Conv2D(8, (3, 3), padding='same')(up0b)\n    up0b = BatchNormalization()(up0b)\n    up0b = Activation('relu')(up0b)\n    up0b = Conv2D(8, (3, 3), padding='same')(up0b)\n    up0b = BatchNormalization()(up0b)\n    up0b = Activation('relu')(up0b)\n    up0b = Conv2D(8, (3, 3), padding='same')(up0b)\n    up0b = BatchNormalization()(up0b)\n    up0b = Activation('relu')(up0b)\n\nIf the original size is 1918x1280,  the output size will be 959x640 after the first maxplooing layer. But the output size is 960x640 after the first unsampling layer. The two layers cannot be concatenated. I can not find where the mistake happened.\nThank you for your advice.",
    "218826": "Resize the image to 1920x1280 before feeding it to the network. This way the image will be divisible by 2 through the maxpooling layers. Hope this helps.",
    "218874": "Thank you for your advice, I will try it later. By the way, did you do it in this way?",
    "218876": "I pad with 2px of zeros the 1918 so it becomes 1920. You can do it in the generator or in the net itself.",
    "219112": "Then how do you solve it? Using BR layer rather BN layer? Or any other solution? I only have a signal GPU (1080Ti).",
    "219125": "I am training on half resolution, 960x640. I tried 959x640 with padding same. Didn't work."
  },
  "source": "meta"
}