{
  "id": 31505,
  "title": "4th place overview from deepsense.io",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/writeups/deepsense-io-4th-place-overview-from-deepsense-io",
  "author_name": "",
  "post_date": "2017-04-12T15:38:12.089880400Z",
  "votes": 11,
  "comment_count": 12,
  "views": 4,
  "content": "<p>We've just posted high level overview of our solution. It's written for people who haven't participated in the competition, so may lack a lot of details. But if you have any questions, I'll try to answer them here, in this thread.</p>\n\n<p><a href=\"https://deepsense.io/deep-learning-for-satellite-imagery-via-image-segmentation/\">https://deepsense.io/deep-learning-for-satellite-imagery-via-image-segmentation/</a></p>",
  "messages": [
    {
      "id": "174684",
      "postDate": "04/12/2017 15:38:12",
      "content": "<p>We've just posted high level overview of our solution. It's written for people who haven't participated in the competition, so may lack a lot of details. But if you have any questions, I'll try to answer them here, in this thread.</p>\n\n<p><a href=\"https://deepsense.io/deep-learning-for-satellite-imagery-via-image-segmentation/\">https://deepsense.io/deep-learning-for-satellite-imagery-via-image-segmentation/</a></p>",
      "rawMarkdown": "We've just posted high level overview of our solution. It's written for people who haven't participated in the competition, so may lack a lot of details. But if you have any questions, I'll try to answer them here, in this thread.\n\nhttps://deepsense.io/deep-learning-for-satellite-imagery-via-image-segmentation/",
      "votes": null
    },
    {
      "id": "174952",
      "postDate": "04/13/2017 09:25:55",
      "content": "<p>Great post and solution, guys! </p>\n\n<p>One thing I notice as very unusual -- do you really used 4 patches per batch? Why so little? </p>\n\n<p>And also, do you have breakdown to classes? If so, could you share? (it is very interesting to compare with our results!)</p>",
      "rawMarkdown": "Great post and solution, guys! \n\nOne thing I notice as very unusual -- do you really used 4 patches per batch? Why so little? \n\nAnd also, do you have breakdown to classes? If so, could you share? (it is very interesting to compare with our results!)",
      "votes": null
    },
    {
      "id": "175192",
      "postDate": "04/14/2017 07:52:03",
      "content": "<p>Yes, it was 4 256x256 patches per batch. \nWe used GPU with 8GB memory. It should be enough to train with larger batch size but because of some implementation details we quickly got to memory limit. When I think of it now, I wish I hadn't optimized the code for larger batch size.</p>\n\n<p>After competition we found simple way to train models with arbitrarily large batch size in Pytorch. Because in Pytorch you control when optimizer does updates and when gradients are zeroed out, the algorithm is straightforward:</p>\n\n<pre><code>1 zero out gradients values\n2 run a couple of times forward pass and backword pass (sample new batch at each time) - gradients are automatically accumulated\n3 call optimizer which does weights update\n4 Repeat 1-&gt;2-&gt;3 \n</code></pre>\n\n<p>It's definitely something I will try in further competitions. It may be slow, but good for convergence.</p>\n\n<p>Our results per class in final submission:</p>\n\n<pre><code>classid,private,public\n1,0.06295,0.07791\n2,0.02311,0.01732\n3,0.04618,0.07758\n4,0.04410,0.04219\n5,0.06445,0.05492\n6,0.08549,0.07358\n7,0.08887,0.09295\n8,0.01817,0.05559\n9,0.01414,0.03328\n10,0.01140,0.01590\ntotal,0.45886,0.54122\n</code></pre>",
      "rawMarkdown": "Yes, it was 4 256x256 patches per batch. \nWe used GPU with 8GB memory. It should be enough to train with larger batch size but because of some implementation details we quickly got to memory limit. When I think of it now, I wish I hadn't optimized the code for larger batch size.\n\nAfter competition we found simple way to train models with arbitrarily large batch size in Pytorch. Because in Pytorch you control when optimizer does updates and when gradients are zeroed out, the algorithm is straightforward:\n\n    1 zero out gradients values\n    2 run a couple of times forward pass and backword pass (sample new batch at each time) - gradients are automatically accumulated\n    3 call optimizer which does weights update\n    4 Repeat 1->2->3 \n\nIt's definitely something I will try in further competitions. It may be slow, but good for convergence.\n\nOur results per class in final submission:\n\n    classid,private,public\n    1,0.06295,0.07791\n    2,0.02311,0.01732\n    3,0.04618,0.07758\n    4,0.04410,0.04219\n    5,0.06445,0.05492\n    6,0.08549,0.07358\n    7,0.08887,0.09295\n    8,0.01817,0.05559\n    9,0.01414,0.03328\n    10,0.01140,0.01590\n    total,0.45886,0.54122",
      "votes": null
    },
    {
      "id": "175492",
      "postDate": "04/15/2017 19:01:04",
      "content": "<p>Thanks for sharing! Some of you classes are extremely good, like cars (given how hard it to detect) or crops (almost perfect detection)\nIt is strange that you hit memory limits with such small batch. I have 8Gb GPU too and experimented with 224*224*20 patches (it is a bit smaller than 256*256*20) and batch about 24 using Keras+Theano. \nHowever, great finding how to accumulate gradient in PyTorch, it is definitely worth knowing!</p>",
      "rawMarkdown": "Thanks for sharing! Some of you classes are extremely good, like cars (given how hard it to detect) or crops (almost perfect detection)\nIt is strange that you hit memory limits with such small batch. I have 8Gb GPU too and experimented with 224*224*20 patches (it is a bit smaller than 256*256*20) and batch about 24 using Keras+Theano. \nHowever, great finding how to accumulate gradient in PyTorch, it is definitely worth knowing!",
      "votes": null
    },
    {
      "id": "178081",
      "postDate": "04/26/2017 16:20:03",
      "content": "<p>Thanks so much for sharing. \nI would like to ask for your up-conv step, is it a repeating of nearby pixel or is a real 3x3 kernel learned for upsampling? Thanks a lot.</p>",
      "rawMarkdown": "Thanks so much for sharing. \nI would like to ask for your up-conv step, is it a repeating of nearby pixel or is a real 3x3 kernel learned for upsampling? Thanks a lot.",
      "votes": null
    },
    {
      "id": "178168",
      "postDate": "04/26/2017 20:47:05",
      "content": "<p>It's transposed convolution with 3x3 kernel which is learned to upsample image. Here you'll find details for UPCONV layer:\n<a href=\"https://deepsense.io/wp-content/uploads/2017/04/architecture_details.png\">https://deepsense.io/wp-content/uploads/2017/04/architecture_details.png</a></p>",
      "rawMarkdown": "It's transposed convolution with 3x3 kernel which is learned to upsample image. Here you'll find details for UPCONV layer:\nhttps://deepsense.io/wp-content/uploads/2017/04/architecture_details.png",
      "votes": null
    },
    {
      "id": "179450",
      "postDate": "05/01/2017 15:01:46",
      "content": "<p>Thx a lot for reply. I use Keras, and I found there was not 3x3 upsamle kernel, it only had an upsampling kernel to duplicate pixel value. So this is by PyTorch or you implemented this in Keras? For learning purpose, I am thinking to reimplement your idea and see how well can I score. Thanks.</p>\n\n<p>For \"Alignment was necessary to remove shifts between channels\", do you think OpenCV \"MOTION_TRANSLATION \" (<a href=\"http://www.learnopencv.com/image-alignment-ecc-in-opencv-c-python/\">http://www.learnopencv.com/image-alignment-ecc-in-opencv-c-python/</a>) is sufficient? Also, after alignment, there will be some missing contents for some channels, how do you handle this? Many thanks for sharing knowledge. </p>",
      "rawMarkdown": "Thx a lot for reply. I use Keras, and I found there was not 3x3 upsamle kernel, it only had an upsampling kernel to duplicate pixel value. So this is by PyTorch or you implemented this in Keras? For learning purpose, I am thinking to reimplement your idea and see how well can I score. Thanks.\n\nFor \"Alignment was necessary to remove shifts between channels\", do you think OpenCV \"MOTION_TRANSLATION \" (http://www.learnopencv.com/image-alignment-ecc-in-opencv-c-python/) is sufficient? Also, after alignment, there will be some missing contents for some channels, how do you handle this? Many thanks for sharing knowledge.",
      "votes": null
    },
    {
      "id": "180171",
      "postDate": "05/04/2017 10:08:21",
      "content": "<p>First, all the images were upsampled to the resolution of the RGB image. Then we did our registration with translation only, using OpenCV \"MOTION_TRANSLATION\".</p>\n\n<p>Also, we registered both of the multispectral images to the panspectral one, then panspectral to RGB and then we composed translations to have all images aligned with the RGB ones. This type of registration seemed to be less likely to diverge than registering directly to RGB, although we still had divergence for a few images (these were just a few fairly homogenous ones so we just kept misaligned images, but maybe some more intelligent backup would be better).</p>\n\n<p>Once registered, the images were aligned and we concatenated all the channels. The missing parts after alignment were just filled with the image average value for each channel.</p>",
      "rawMarkdown": "First, all the images were upsampled to the resolution of the RGB image. Then we did our registration with translation only, using OpenCV \"MOTION_TRANSLATION\".\n\nAlso, we registered both of the multispectral images to the panspectral one, then panspectral to RGB and then we composed translations to have all images aligned with the RGB ones. This type of registration seemed to be less likely to diverge than registering directly to RGB, although we still had divergence for a few images (these were just a few fairly homogenous ones so we just kept misaligned images, but maybe some more intelligent backup would be better).\n\nOnce registered, the images were aligned and we concatenated all the channels. The missing parts after alignment were just filled with the image average value for each channel.",
      "votes": null
    },
    {
      "id": "180203",
      "postDate": "05/04/2017 12:25:33",
      "content": "<p>Final architecture was implemented in PyTorch. For convolutional upsampling we used this: <a href=\"http://pytorch.org/docs/nn.html#convtranspose2d\">http://pytorch.org/docs/nn.html#convtranspose2d</a>\nbut Keras has something similar:\n<a href=\"https://keras.io/layers/convolutional/#conv2dtranspose\">https://keras.io/layers/convolutional/#conv2dtranspose</a></p>",
      "rawMarkdown": "Final architecture was implemented in PyTorch. For convolutional upsampling we used this: http://pytorch.org/docs/nn.html#convtranspose2d\nbut Keras has something similar:\nhttps://keras.io/layers/convolutional/#conv2dtranspose",
      "votes": null
    },
    {
      "id": "180789",
      "postDate": "05/07/2017 08:30:28",
      "content": "<p>Share the code please.</p>",
      "rawMarkdown": "Share the code please.",
      "votes": null
    },
    {
      "id": "192664",
      "postDate": "06/14/2017 10:30:57",
      "content": "<p>Can this code be used as it is for segmentation and classification of Landsat 8 image (30m resolution) ? If no, Will you let me know how it is to be done ?</p>",
      "rawMarkdown": "Can this code be used as it is for segmentation and classification of Landsat 8 image (30m resolution) ? If no, Will you let me know how it is to be done ?",
      "votes": null
    },
    {
      "id": "192778",
      "postDate": "06/14/2017 19:08:53",
      "content": "<p>Thanks for sharing! Did you guys do any preprocessing to the images (smoothing, blurring, boosting contrast, etc.) before feeding them in?</p>",
      "rawMarkdown": "Thanks for sharing! Did you guys do any preprocessing to the images (smoothing, blurring, boosting contrast, etc.) before feeding them in?",
      "votes": null
    },
    {
      "id": "193407",
      "postDate": "06/16/2017 10:35:48",
      "content": "<p>We normalized each input channel independently to have a zero mean and unit variance. This approach requires some pre-computed mean and std per channel. You can get this numbers in a couple of ways:</p>\n\n<p>1) for each input image compute its own mean and std and use them during normalization</p>\n\n<p>2) compute mean and std for entire dataset and use always the same number or any input image</p>\n\n<p>3) compute mean and std per each 5x5 region and normalize input images by region statistics to which image belongs to (I think it worked the best in our case)</p>",
      "rawMarkdown": "We normalized each input channel independently to have a zero mean and unit variance. This approach requires some pre-computed mean and std per channel. You can get this numbers in a couple of ways:\n\n1) for each input image compute its own mean and std and use them during normalization\n\n2) compute mean and std for entire dataset and use always the same number or any input image\n\n3) compute mean and std per each 5x5 region and normalize input images by region statistics to which image belongs to (I think it worked the best in our case)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 174952,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "04/13/2017 09:25:55",
      "content": "<p>Great post and solution, guys! </p>\n\n<p>One thing I notice as very unusual -- do you really used 4 patches per batch? Why so little? </p>\n\n<p>And also, do you have breakdown to classes? If so, could you share? (it is very interesting to compare with our results!)</p>",
      "votes": null,
      "replies": [
        {
          "id": 175192,
          "author_name": "arnowaczynski",
          "author_url": "",
          "post_date": "04/14/2017 07:52:03",
          "content": "<p>Yes, it was 4 256x256 patches per batch. \nWe used GPU with 8GB memory. It should be enough to train with larger batch size but because of some implementation details we quickly got to memory limit. When I think of it now, I wish I hadn't optimized the code for larger batch size.</p>\n\n<p>After competition we found simple way to train models with arbitrarily large batch size in Pytorch. Because in Pytorch you control when optimizer does updates and when gradients are zeroed out, the algorithm is straightforward:</p>\n\n<pre><code>1 zero out gradients values\n2 run a couple of times forward pass and backword pass (sample new batch at each time) - gradients are automatically accumulated\n3 call optimizer which does weights update\n4 Repeat 1-&gt;2-&gt;3 \n</code></pre>\n\n<p>It's definitely something I will try in further competitions. It may be slow, but good for convergence.</p>\n\n<p>Our results per class in final submission:</p>\n\n<pre><code>classid,private,public\n1,0.06295,0.07791\n2,0.02311,0.01732\n3,0.04618,0.07758\n4,0.04410,0.04219\n5,0.06445,0.05492\n6,0.08549,0.07358\n7,0.08887,0.09295\n8,0.01817,0.05559\n9,0.01414,0.03328\n10,0.01140,0.01590\ntotal,0.45886,0.54122\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 175492,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "04/15/2017 19:01:04",
          "content": "<p>Thanks for sharing! Some of you classes are extremely good, like cars (given how hard it to detect) or crops (almost perfect detection)\nIt is strange that you hit memory limits with such small batch. I have 8Gb GPU too and experimented with 224*224*20 patches (it is a bit smaller than 256*256*20) and batch about 24 using Keras+Theano. \nHowever, great finding how to accumulate gradient in PyTorch, it is definitely worth knowing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 178081,
      "author_name": "craymond",
      "author_url": "",
      "post_date": "04/26/2017 16:20:03",
      "content": "<p>Thanks so much for sharing. \nI would like to ask for your up-conv step, is it a repeating of nearby pixel or is a real 3x3 kernel learned for upsampling? Thanks a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 178168,
          "author_name": "arnowaczynski",
          "author_url": "",
          "post_date": "04/26/2017 20:47:05",
          "content": "<p>It's transposed convolution with 3x3 kernel which is learned to upsample image. Here you'll find details for UPCONV layer:\n<a href=\"https://deepsense.io/wp-content/uploads/2017/04/architecture_details.png\">https://deepsense.io/wp-content/uploads/2017/04/architecture_details.png</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 179450,
          "author_name": "craymond",
          "author_url": "",
          "post_date": "05/01/2017 15:01:46",
          "content": "<p>Thx a lot for reply. I use Keras, and I found there was not 3x3 upsamle kernel, it only had an upsampling kernel to duplicate pixel value. So this is by PyTorch or you implemented this in Keras? For learning purpose, I am thinking to reimplement your idea and see how well can I score. Thanks.</p>\n\n<p>For \"Alignment was necessary to remove shifts between channels\", do you think OpenCV \"MOTION_TRANSLATION \" (<a href=\"http://www.learnopencv.com/image-alignment-ecc-in-opencv-c-python/\">http://www.learnopencv.com/image-alignment-ecc-in-opencv-c-python/</a>) is sufficient? Also, after alignment, there will be some missing contents for some channels, how do you handle this? Many thanks for sharing knowledge. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 180171,
          "author_name": "romaniukm",
          "author_url": "",
          "post_date": "05/04/2017 10:08:21",
          "content": "<p>First, all the images were upsampled to the resolution of the RGB image. Then we did our registration with translation only, using OpenCV \"MOTION_TRANSLATION\".</p>\n\n<p>Also, we registered both of the multispectral images to the panspectral one, then panspectral to RGB and then we composed translations to have all images aligned with the RGB ones. This type of registration seemed to be less likely to diverge than registering directly to RGB, although we still had divergence for a few images (these were just a few fairly homogenous ones so we just kept misaligned images, but maybe some more intelligent backup would be better).</p>\n\n<p>Once registered, the images were aligned and we concatenated all the channels. The missing parts after alignment were just filled with the image average value for each channel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 180203,
          "author_name": "arnowaczynski",
          "author_url": "",
          "post_date": "05/04/2017 12:25:33",
          "content": "<p>Final architecture was implemented in PyTorch. For convolutional upsampling we used this: <a href=\"http://pytorch.org/docs/nn.html#convtranspose2d\">http://pytorch.org/docs/nn.html#convtranspose2d</a>\nbut Keras has something similar:\n<a href=\"https://keras.io/layers/convolutional/#conv2dtranspose\">https://keras.io/layers/convolutional/#conv2dtranspose</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 180789,
      "author_name": "quest13",
      "author_url": "",
      "post_date": "05/07/2017 08:30:28",
      "content": "<p>Share the code please.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192664,
      "author_name": "researcher1",
      "author_url": "",
      "post_date": "06/14/2017 10:30:57",
      "content": "<p>Can this code be used as it is for segmentation and classification of Landsat 8 image (30m resolution) ? If no, Will you let me know how it is to be done ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192778,
      "author_name": "adsnash",
      "author_url": "",
      "post_date": "06/14/2017 19:08:53",
      "content": "<p>Thanks for sharing! Did you guys do any preprocessing to the images (smoothing, blurring, boosting contrast, etc.) before feeding them in?</p>",
      "votes": null,
      "replies": [
        {
          "id": 193407,
          "author_name": "arnowaczynski",
          "author_url": "",
          "post_date": "06/16/2017 10:35:48",
          "content": "<p>We normalized each input channel independently to have a zero mean and unit variance. This approach requires some pre-computed mean and std per channel. You can get this numbers in a couple of ways:</p>\n\n<p>1) for each input image compute its own mean and std and use them during normalization</p>\n\n<p>2) compute mean and std for entire dataset and use always the same number or any input image</p>\n\n<p>3) compute mean and std per each 5x5 region and normalize input images by region statistics to which image belongs to (I think it worked the best in our case)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "174684": "We've just posted high level overview of our solution. It's written for people who haven't participated in the competition, so may lack a lot of details. But if you have any questions, I'll try to answer them here, in this thread.\n\nhttps://deepsense.io/deep-learning-for-satellite-imagery-via-image-segmentation/",
    "174952": "Great post and solution, guys! \n\nOne thing I notice as very unusual -- do you really used 4 patches per batch? Why so little? \n\nAnd also, do you have breakdown to classes? If so, could you share? (it is very interesting to compare with our results!)",
    "175192": "Yes, it was 4 256x256 patches per batch. \nWe used GPU with 8GB memory. It should be enough to train with larger batch size but because of some implementation details we quickly got to memory limit. When I think of it now, I wish I hadn't optimized the code for larger batch size.\n\nAfter competition we found simple way to train models with arbitrarily large batch size in Pytorch. Because in Pytorch you control when optimizer does updates and when gradients are zeroed out, the algorithm is straightforward:\n\n    1 zero out gradients values\n    2 run a couple of times forward pass and backword pass (sample new batch at each time) - gradients are automatically accumulated\n    3 call optimizer which does weights update\n    4 Repeat 1->2->3 \n\nIt's definitely something I will try in further competitions. It may be slow, but good for convergence.\n\nOur results per class in final submission:\n\n    classid,private,public\n    1,0.06295,0.07791\n    2,0.02311,0.01732\n    3,0.04618,0.07758\n    4,0.04410,0.04219\n    5,0.06445,0.05492\n    6,0.08549,0.07358\n    7,0.08887,0.09295\n    8,0.01817,0.05559\n    9,0.01414,0.03328\n    10,0.01140,0.01590\n    total,0.45886,0.54122",
    "175492": "Thanks for sharing! Some of you classes are extremely good, like cars (given how hard it to detect) or crops (almost perfect detection)\nIt is strange that you hit memory limits with such small batch. I have 8Gb GPU too and experimented with 224*224*20 patches (it is a bit smaller than 256*256*20) and batch about 24 using Keras+Theano. \nHowever, great finding how to accumulate gradient in PyTorch, it is definitely worth knowing!",
    "178081": "Thanks so much for sharing. \nI would like to ask for your up-conv step, is it a repeating of nearby pixel or is a real 3x3 kernel learned for upsampling? Thanks a lot.",
    "178168": "It's transposed convolution with 3x3 kernel which is learned to upsample image. Here you'll find details for UPCONV layer:\nhttps://deepsense.io/wp-content/uploads/2017/04/architecture_details.png",
    "179450": "Thx a lot for reply. I use Keras, and I found there was not 3x3 upsamle kernel, it only had an upsampling kernel to duplicate pixel value. So this is by PyTorch or you implemented this in Keras? For learning purpose, I am thinking to reimplement your idea and see how well can I score. Thanks.\n\nFor \"Alignment was necessary to remove shifts between channels\", do you think OpenCV \"MOTION_TRANSLATION \" (http://www.learnopencv.com/image-alignment-ecc-in-opencv-c-python/) is sufficient? Also, after alignment, there will be some missing contents for some channels, how do you handle this? Many thanks for sharing knowledge.",
    "180171": "First, all the images were upsampled to the resolution of the RGB image. Then we did our registration with translation only, using OpenCV \"MOTION_TRANSLATION\".\n\nAlso, we registered both of the multispectral images to the panspectral one, then panspectral to RGB and then we composed translations to have all images aligned with the RGB ones. This type of registration seemed to be less likely to diverge than registering directly to RGB, although we still had divergence for a few images (these were just a few fairly homogenous ones so we just kept misaligned images, but maybe some more intelligent backup would be better).\n\nOnce registered, the images were aligned and we concatenated all the channels. The missing parts after alignment were just filled with the image average value for each channel.",
    "180203": "Final architecture was implemented in PyTorch. For convolutional upsampling we used this: http://pytorch.org/docs/nn.html#convtranspose2d\nbut Keras has something similar:\nhttps://keras.io/layers/convolutional/#conv2dtranspose",
    "180789": "Share the code please.",
    "192664": "Can this code be used as it is for segmentation and classification of Landsat 8 image (30m resolution) ? If no, Will you let me know how it is to be done ?",
    "192778": "Thanks for sharing! Did you guys do any preprocessing to the images (smoothing, blurring, boosting contrast, etc.) before feeding them in?",
    "193407": "We normalized each input channel independently to have a zero mean and unit variance. This approach requires some pre-computed mean and std per channel. You can get this numbers in a couple of ways:\n\n1) for each input image compute its own mean and std and use them during normalization\n\n2) compute mean and std for entire dataset and use always the same number or any input image\n\n3) compute mean and std per each 5x5 region and normalize input images by region statistics to which image belongs to (I think it worked the best in our case)"
  },
  "source": "meta"
}