{
  "id": 27377,
  "title": "3 band vs 16 band",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/27377",
  "author_name": "",
  "post_date": "2017-01-06T03:57:24.880Z",
  "votes": 2,
  "comment_count": 18,
  "views": 1208,
  "content": "<p>On the one hand I would expect that data in different channels is extremely correlated, but on the other hand correlated does not mean exactly the same.</p>\n\n<p>Question is: Is there really a big difference between 3 band and 16 band images?</p>\n\n<p>Answers like: </p>\n\n<ol>\n<li>Based on the paper XXX there should be a big difference for trees / water / roads / xxx, because in the infrared they are very distinguishable.</li>\n<li>My current model is highly / not at all affected by a switch from 3 to 16 channels.</li>\n</ol>\n\n<p>are highly appreciated.</p>",
  "messages": [
    {
      "id": "154416",
      "postDate": "01/06/2017 03:57:24",
      "content": "<p>On the one hand I would expect that data in different channels is extremely correlated, but on the other hand correlated does not mean exactly the same.</p>\n\n<p>Question is: Is there really a big difference between 3 band and 16 band images?</p>\n\n<p>Answers like: </p>\n\n<ol>\n<li>Based on the paper XXX there should be a big difference for trees / water / roads / xxx, because in the infrared they are very distinguishable.</li>\n<li>My current model is highly / not at all affected by a switch from 3 to 16 channels.</li>\n</ol>\n\n<p>are highly appreciated.</p>",
      "rawMarkdown": "On the one hand I would expect that data in different channels is extremely correlated, but on the other hand correlated does not mean exactly the same.\r\n\r\nQuestion is: Is there really a big difference between 3 band and 16 band images?\r\n\r\nAnswers like: \r\n\r\n 1. Based on the paper XXX there should be a big difference for trees / water / roads / xxx, because in the infrared they are very distinguishable.\r\n 2. My current model is highly / not at all affected by a switch from 3 to 16 channels.\r\n\r\nare highly appreciated.",
      "votes": null
    },
    {
      "id": "154450",
      "postDate": "01/06/2017 08:45:29",
      "content": "<p>3) Based on 3 year remote sensing (ML based) experience, YES! Regular earth observation satellites have specific wavelengths to detect certain physical phenomena. Oxygen absorption band, chlorophyll fluorescence, temperature bands, etc. etc. etc. Now, I didn't do <em>any</em> reading on how the WorldView bands were designated. </p>\n\n<p>I noted that in some bands (from 16 bands), the trees/vegetation are brighter than the rest of the features. This must be true for water (probably another channel) and maybe some land types too (eg. roads are probably hotter (brighter in infrared) and less bright in UV. </p>\n\n<p>But this is based on my previous experience, I didn't do any statistics on current dataset. Anyway, you can't beat a system designated to directly observe the physical phenomena (eg. extract the chlorophyll levels when you don't have dedicated bands for that)</p>\n\n<p>Cristi</p>",
      "rawMarkdown": "3) Based on 3 year remote sensing (ML based) experience, YES! Regular earth observation satellites have specific wavelengths to detect certain physical phenomena. Oxygen absorption band, chlorophyll fluorescence, temperature bands, etc. etc. etc. Now, I didn't do *any* reading on how the WorldView bands were designated. \r\n\r\nI noted that in some bands (from 16 bands), the trees/vegetation are brighter than the rest of the features. This must be true for water (probably another channel) and maybe some land types too (eg. roads are probably hotter (brighter in infrared) and less bright in UV. \r\n\r\nBut this is based on my previous experience, I didn't do any statistics on current dataset. Anyway, you can't beat a system designated to directly observe the physical phenomena (eg. extract the chlorophyll levels when you don't have dedicated bands for that)\r\n\r\nCristi",
      "votes": null
    },
    {
      "id": "154486",
      "postDate": "01/06/2017 12:35:50",
      "content": "<p>Yes, there is a big difference. That's the reason for their existence. The first difference is due to the wavelength. As Cristi said different materials have different fingerprints (roads made of asphalt vs other \"roads\" look differently.</p>\n\n<p>Important to notice a wider spectrum doesn't come for free. Multispectral bands have a lower resolution (each pixel cover a greater land area), while P have a res of 0.3m which is great for fine details.\nAlso number of bits, each sensor havs a different number of bitsfot each pxel, we havd more than 8 bits of information.</p>\n\n<p>IMHO, regardless of the approach you choose, you can make use of all that information. People using DL does not need to restrain to RGB or 3 channels only.</p>",
      "rawMarkdown": "Yes, there is a big difference. That's the reason for their existence. The first difference is due to the wavelength. As Cristi said different materials have different fingerprints (roads made of asphalt vs other \"roads\" look differently.\r\n\r\nImportant to notice a wider spectrum doesn't come for free. Multispectral bands have a lower resolution (each pixel cover a greater land area), while P have a res of 0.3m which is great for fine details.\r\nAlso number of bits, each sensor havs a different number of bitsfot each pxel, we havd more than 8 bits of information.\r\n\r\nIMHO, regardless of the approach you choose, you can make use of all that information. People using DL does not need to restrain to RGB or 3 channels only.",
      "votes": null
    },
    {
      "id": "154500",
      "postDate": "01/06/2017 14:35:09",
      "content": "<p>I think if you are using CNN's there is probably not a big difference, as those rely a lot on edges which are best defined in the 3-band true color image. But there is a lot that can be done using pixel wise classification, which rely heavily on the spectral properties as @amaia and @visoft have pointed out. </p>",
      "rawMarkdown": "I think if you are using CNN's there is probably not a big difference, as those rely a lot on edges which are best defined in the 3-band true color image. But there is a lot that can be done using pixel wise classification, which rely heavily on the spectral properties as @amaia and @visoft have pointed out.",
      "votes": null
    },
    {
      "id": "154532",
      "postDate": "01/06/2017 17:56:44",
      "content": "<p>Thank you. </p>\n\n<p>I would like to concatenate all different bands and make an array that can be fed into the CNN. I am looking at image <strong>6010_0_0</strong></p>\n\n<ul>\n<li>3 band =&gt; shape (3, 3349, 3396)</li>\n<li>A =&gt; shape (8, 134, 136)</li>\n<li>M =&gt; shape (8, 837, 849)</li>\n<li>P =&gt; shape (3348, 3396)</li>\n</ul>\n\n<p>[1] Is it true that P - is just grayscale, meaning linear combination version of <em>3 band</em>?</p>\n\n<p>[2] I have tried to extract RGB image from <em>M</em> following <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26763/can-we-get-the-frequency-of-the-diffrent-sixteen-band\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26763/can-we-get-the-frequency-of-the-diffrent-sixteen-band</a>\nbut I did not have success with it. Are you able to do it?</p>\n\n<p>[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than</p>\n\n<p><code>np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))</code></p>\n\n<p>?</p>",
      "rawMarkdown": "Thank you. \r\n\r\nI would like to concatenate all different bands and make an array that can be fed into the CNN. I am looking at image **6010_0_0**\r\n\r\n - 3 band => shape (3, 3349, 3396)\r\n - A => shape (8, 134, 136)\r\n - M => shape (8, 837, 849)\r\n - P => shape (3348, 3396)\r\n\r\n[1] Is it true that P - is just grayscale, meaning linear combination version of *3 band*?\r\n\r\n[2] I have tried to extract RGB image from *M* following https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26763/can-we-get-the-frequency-of-the-diffrent-sixteen-band\r\nbut I did not have success with it. Are you able to do it?\r\n\r\n[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than\r\n\r\n`np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))`\r\n\r\n?",
      "votes": null
    },
    {
      "id": "154572",
      "postDate": "01/06/2017 21:50:00",
      "content": "<p>Vladimir, maybe this will help:</p>\n\n<p><a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2/code\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2/code</a></p>\n\n<p>In the example above there are only 2 images concatenated but of course it is easy extendable. And it takes care of A-P displacement. Feel free to register M to P (or _3) also.</p>\n\n<p>On <a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/export-pixel-wise-mask\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/export-pixel-wise-mask</a>  you will find how to build pixel-wise masks from polygons.</p>\n\n<p>Hope it helps!</p>\n\n<p>p.s. Please note that different scenes have different image sizes and form factors! (aka not all P images have the same size)</p>",
      "rawMarkdown": "Vladimir, maybe this will help:\r\n\r\nhttps://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2/code\r\n\r\nIn the example above there are only 2 images concatenated but of course it is easy extendable. And it takes care of A-P displacement. Feel free to register M to P (or _3) also.\r\n\r\nOn https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/export-pixel-wise-mask  you will find how to build pixel-wise masks from polygons.\r\n\r\nHope it helps!\r\n\r\np.s. Please note that different scenes have different image sizes and form factors! (aka not all P images have the same size)",
      "votes": null
    },
    {
      "id": "154672",
      "postDate": "01/07/2017 15:23:28",
      "content": "<p>Check out this video for a good intro on what the different bands mean and how they are different.   </p>\n\n<p><a href=\"https://www.youtube.com/watch?v=3iaFzafWJQE\">https://www.youtube.com/watch?v=3iaFzafWJQE</a></p>",
      "rawMarkdown": "Check out this video for a good intro on what the different bands mean and how they are different.   \r\n\r\nhttps://www.youtube.com/watch?v=3iaFzafWJQE",
      "votes": null
    },
    {
      "id": "156900",
      "postDate": "01/18/2017 13:48:31",
      "content": "<p>Hi All, </p>\n\n<p>This is the first competition that I am entering at Kaggle and so my question might be too naive for you experts.</p>\n\n<p>Above video shared by @shawn is good place to start. But I want to go to deeper level understanding. After going through multiple threads, I have understood bits and pieces but now I am confused on <strong>whether RGB bands and A,M,P are NULL in intersection or not.</strong></p>\n\n<p>I see that RGB overlaps with M at few wavelengths.\nPlease correct me if I am wrong</p>",
      "rawMarkdown": "Hi All, \r\n\r\nThis is the first competition that I am entering at Kaggle and so my question might be too naive for you experts.\r\n\r\nAbove video shared by @shawn is good place to start. But I want to go to deeper level understanding. After going through multiple threads, I have understood bits and pieces but now I am confused on **whether RGB bands and A,M,P are NULL in intersection or not.**\r\n \r\nI see that RGB overlaps with M at few wavelengths.\r\nPlease correct me if I am wrong",
      "votes": null
    },
    {
      "id": "156939",
      "postDate": "01/18/2017 17:30:13",
      "content": "<p>[quote=Puneet Jindal;156900]</p>\n\n<p>Please correct me if I am wrong</p>\n\n<p>[/quote]</p>\n\n<p>This satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.</p>",
      "rawMarkdown": "[quote=Puneet Jindal;156900]\r\n\r\nPlease correct me if I am wrong\r\n\r\n[/quote]\r\n\r\nThis satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.",
      "votes": null
    },
    {
      "id": "156983",
      "postDate": "01/18/2017 20:36:54",
      "content": "<p>[quote=amaia;156939]</p>\n\n<p>This satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.</p>\n\n<p>[/quote]</p>\n\n<p>That's correct in principle. The RGB bands of the 3-band data correspond to 3 of the Multispectral bands.  But in practice the data isn't the same. There is a small misalignment of a few pixels between the 3-band and M-bands. (Presumably the class polygons are draw with respect to the 3-band images.) And, although visually similar at large scales, when you zoom in the M-band images look fuzzier (i.e. lower resolution) than the 3-bands.</p>\n\n<p>(attached are views of the top left corder or 6100_1_3.  Red band and the corresponding band from Multispectral data.)</p>\n\n<p>So we've decided to keep all the data. We pad and resize all the images to the same size, realign the A, M and P bands to the mean of the 3-band data, and then stack them all into a 20-layer sandwich. And then add another 10 layers for the category masks on top of that.</p>",
      "rawMarkdown": "[quote=amaia;156939]\r\n\r\nThis satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.\r\n\r\n[/quote]\r\n\r\nThat's correct in principle. The RGB bands of the 3-band data correspond to 3 of the Multispectral bands.  But in practice the data isn't the same. There is a small misalignment of a few pixels between the 3-band and M-bands. (Presumably the class polygons are draw with respect to the 3-band images.) And, although visually similar at large scales, when you zoom in the M-band images look fuzzier (i.e. lower resolution) than the 3-bands.\r\n\r\n(attached are views of the top left corder or 6100_1_3.  Red band and the corresponding band from Multispectral data.)\r\n\r\nSo we've decided to keep all the data. We pad and resize all the images to the same size, realign the A, M and P bands to the mean of the 3-band data, and then stack them all into a 20-layer sandwich. And then add another 10 layers for the category masks on top of that.",
      "votes": null
    },
    {
      "id": "156989",
      "postDate": "01/18/2017 20:54:14",
      "content": "<p>I agree with you, I was responding to the \"wavelenghts\" part.</p>",
      "rawMarkdown": "I agree with you, I was responding to the \"wavelenghts\" part.",
      "votes": null
    },
    {
      "id": "157163",
      "postDate": "01/19/2017 14:14:07",
      "content": "<p>@amaia: Thanks for the explanation however i still have question: if it is 17 then why we say it as sixteen band data?</p>\n\n<p>According to me, CNN should work only on 3 band(RGB) images as these are true color images. </p>\n\n<p>Rest is converting all the images to spectral signatures and using train dataset polygon from the images-class mapping to be fed into the model which shall be able to predict for test as well.</p>",
      "rawMarkdown": "amaia: Thanks for the explanation however i still have question: if it is 17 then why we say it as sixteen band data?\r\n\r\nAccording to me, CNN should work only on 3 band(RGB) images as these are true color images. \r\n\r\nRest is converting all the images to spectral signatures and using train dataset polygon from the images-class mapping to be fed into the model which shall be able to predict for test as well.",
      "votes": null
    },
    {
      "id": "157170",
      "postDate": "01/19/2017 14:37:28",
      "content": "<p>Threeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.</p>\n\n<p>But i am just bouncing ideas here.</p>\n\n<p>Puneet, of course they work. Popular nets are for rgb but 20 channel input shouldn't be a problem.\nFrom what images I saw in forums i think you can predict several labels only with thresholds and some band arithmetic.</p>\n\n<p>Cristi</p>",
      "rawMarkdown": "Threeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.\r\n\r\nBut i am just bouncing ideas here.\r\n\r\nPuneet, of course they work. Popular nets are for rgb but 20 channel input shouldn't be a problem.\r\nFrom what images I saw in forums i think you can predict several labels only with thresholds and some band arithmetic.\r\n\r\nCristi",
      "votes": null
    },
    {
      "id": "157227",
      "postDate": "01/19/2017 19:15:42",
      "content": "<p>[quote=visoft;157170]</p>\n\n<p>Threeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.</p>\n\n<p>[/quote]</p>\n\n<p>That sounds about right to me, although we're using a sigmoid to give 10 outputs bits.</p>",
      "rawMarkdown": "[quote=visoft;157170]\r\n\r\nThreeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.\r\n\r\n[/quote]\r\n\r\nThat sounds about right to me, although we're using a sigmoid to give 10 outputs bits.",
      "votes": null
    },
    {
      "id": "157228",
      "postDate": "01/19/2017 19:19:22",
      "content": "<p>[quote=Vladimir Iglovikov;154532]</p>\n\n<p>[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than</p>\n\n<p><code>np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))</code></p>\n\n<p>[/quote]</p>\n\n<p>Not that I've found. Everything seems to assume greyscale or RGB. So I'm rescaling one band at a time. This is slow, but in principle only has to be done once.</p>",
      "rawMarkdown": "[quote=Vladimir Iglovikov;154532]\r\n\r\n[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than\r\n\r\n`np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))`\r\n\r\n[/quote]\r\n\r\nNot that I've found. Everything seems to assume greyscale or RGB. So I'm rescaling one band at a time. This is slow, but in principle only has to be done once.",
      "votes": null
    },
    {
      "id": "157245",
      "postDate": "01/19/2017 21:15:24",
      "content": "<p>Opencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \n<a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2</a></p>\n\n<p>Anaconda packs opencv 3.1so no compile hussle!</p>",
      "rawMarkdown": "Opencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \r\nhttps://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\r\n\r\nAnaconda packs opencv 3.1so no compile hussle!",
      "votes": null
    },
    {
      "id": "157273",
      "postDate": "01/19/2017 23:40:04",
      "content": "<p>[quote=visoft;157245]</p>\n\n<p>Opencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \n<a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2</a></p>\n\n<p>[/quote]</p>\n\n<p>Aha! I used your script as a template to write my own realignment code. But I overlooked that you're shifting multiple bands at once. Time for another rewrite...</p>",
      "rawMarkdown": "[quote=visoft;157245]\r\n\r\nOpencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \r\nhttps://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\r\n\r\n[/quote]\r\n\r\nAha! I used your script as a template to write my own realignment code. But I overlooked that you're shifting multiple bands at once. Time for another rewrite...",
      "votes": null
    },
    {
      "id": "158494",
      "postDate": "01/27/2017 17:39:52",
      "content": "<p>Now, when I finally survived through the submission errors, I would love to repeat my initial question.</p>\n\n<p>For those that are above <strong>0.4</strong>, do you guys see a significant difference between 3 vs 16 band images for the input.</p>\n\n<p>P.S. My last submission <strong>0.36</strong> is on 3 band.</p>",
      "rawMarkdown": "Now, when I finally survived through the submission errors, I would love to repeat my initial question.\r\n\r\nFor those that are above **0.4**, do you guys see a significant difference between 3 vs 16 band images for the input.\r\n\r\nP.S. My last submission **0.36** is on 3 band.",
      "votes": null
    },
    {
      "id": "158509",
      "postDate": "01/27/2017 19:01:57",
      "content": "<p>All I would say is that certain classes are difficult with only 3 bands, and misalignment is overrated if you are using deep learning.  The miss-alignment might be a bonus just saying.</p>",
      "rawMarkdown": "All I would say is that certain classes are difficult with only 3 bands, and misalignment is overrated if you are using deep learning.  The miss-alignment might be a bonus just saying.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 154450,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "01/06/2017 08:45:29",
      "content": "<p>3) Based on 3 year remote sensing (ML based) experience, YES! Regular earth observation satellites have specific wavelengths to detect certain physical phenomena. Oxygen absorption band, chlorophyll fluorescence, temperature bands, etc. etc. etc. Now, I didn't do <em>any</em> reading on how the WorldView bands were designated. </p>\n\n<p>I noted that in some bands (from 16 bands), the trees/vegetation are brighter than the rest of the features. This must be true for water (probably another channel) and maybe some land types too (eg. roads are probably hotter (brighter in infrared) and less bright in UV. </p>\n\n<p>But this is based on my previous experience, I didn't do any statistics on current dataset. Anyway, you can't beat a system designated to directly observe the physical phenomena (eg. extract the chlorophyll levels when you don't have dedicated bands for that)</p>\n\n<p>Cristi</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154486,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "01/06/2017 12:35:50",
      "content": "<p>Yes, there is a big difference. That's the reason for their existence. The first difference is due to the wavelength. As Cristi said different materials have different fingerprints (roads made of asphalt vs other \"roads\" look differently.</p>\n\n<p>Important to notice a wider spectrum doesn't come for free. Multispectral bands have a lower resolution (each pixel cover a greater land area), while P have a res of 0.3m which is great for fine details.\nAlso number of bits, each sensor havs a different number of bitsfot each pxel, we havd more than 8 bits of information.</p>\n\n<p>IMHO, regardless of the approach you choose, you can make use of all that information. People using DL does not need to restrain to RGB or 3 channels only.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154500,
      "author_name": "shawn775",
      "author_url": "",
      "post_date": "01/06/2017 14:35:09",
      "content": "<p>I think if you are using CNN's there is probably not a big difference, as those rely a lot on edges which are best defined in the 3-band true color image. But there is a lot that can be done using pixel wise classification, which rely heavily on the spectral properties as @amaia and @visoft have pointed out. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154532,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "01/06/2017 17:56:44",
      "content": "<p>Thank you. </p>\n\n<p>I would like to concatenate all different bands and make an array that can be fed into the CNN. I am looking at image <strong>6010_0_0</strong></p>\n\n<ul>\n<li>3 band =&gt; shape (3, 3349, 3396)</li>\n<li>A =&gt; shape (8, 134, 136)</li>\n<li>M =&gt; shape (8, 837, 849)</li>\n<li>P =&gt; shape (3348, 3396)</li>\n</ul>\n\n<p>[1] Is it true that P - is just grayscale, meaning linear combination version of <em>3 band</em>?</p>\n\n<p>[2] I have tried to extract RGB image from <em>M</em> following <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26763/can-we-get-the-frequency-of-the-diffrent-sixteen-band\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26763/can-we-get-the-frequency-of-the-diffrent-sixteen-band</a>\nbut I did not have success with it. Are you able to do it?</p>\n\n<p>[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than</p>\n\n<p><code>np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))</code></p>\n\n<p>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 157228,
          "author_name": "threeplusone",
          "author_url": "",
          "post_date": "01/19/2017 19:19:22",
          "content": "<p>[quote=Vladimir Iglovikov;154532]</p>\n\n<p>[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than</p>\n\n<p><code>np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))</code></p>\n\n<p>[/quote]</p>\n\n<p>Not that I've found. Everything seems to assume greyscale or RGB. So I'm rescaling one band at a time. This is slow, but in principle only has to be done once.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 154572,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "01/06/2017 21:50:00",
      "content": "<p>Vladimir, maybe this will help:</p>\n\n<p><a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2/code\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2/code</a></p>\n\n<p>In the example above there are only 2 images concatenated but of course it is easy extendable. And it takes care of A-P displacement. Feel free to register M to P (or _3) also.</p>\n\n<p>On <a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/export-pixel-wise-mask\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/export-pixel-wise-mask</a>  you will find how to build pixel-wise masks from polygons.</p>\n\n<p>Hope it helps!</p>\n\n<p>p.s. Please note that different scenes have different image sizes and form factors! (aka not all P images have the same size)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154672,
      "author_name": "shawn775",
      "author_url": "",
      "post_date": "01/07/2017 15:23:28",
      "content": "<p>Check out this video for a good intro on what the different bands mean and how they are different.   </p>\n\n<p><a href=\"https://www.youtube.com/watch?v=3iaFzafWJQE\">https://www.youtube.com/watch?v=3iaFzafWJQE</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156900,
      "author_name": "puneetjindal",
      "author_url": "",
      "post_date": "01/18/2017 13:48:31",
      "content": "<p>Hi All, </p>\n\n<p>This is the first competition that I am entering at Kaggle and so my question might be too naive for you experts.</p>\n\n<p>Above video shared by @shawn is good place to start. But I want to go to deeper level understanding. After going through multiple threads, I have understood bits and pieces but now I am confused on <strong>whether RGB bands and A,M,P are NULL in intersection or not.</strong></p>\n\n<p>I see that RGB overlaps with M at few wavelengths.\nPlease correct me if I am wrong</p>",
      "votes": null,
      "replies": [
        {
          "id": 156939,
          "author_name": "aamaia",
          "author_url": "",
          "post_date": "01/18/2017 17:30:13",
          "content": "<p>[quote=Puneet Jindal;156900]</p>\n\n<p>Please correct me if I am wrong</p>\n\n<p>[/quote]</p>\n\n<p>This satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.</p>",
          "votes": null,
          "replies": [
            {
              "id": 156983,
              "author_name": "threeplusone",
              "author_url": "",
              "post_date": "01/18/2017 20:36:54",
              "content": "<p>[quote=amaia;156939]</p>\n\n<p>This satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.</p>\n\n<p>[/quote]</p>\n\n<p>That's correct in principle. The RGB bands of the 3-band data correspond to 3 of the Multispectral bands.  But in practice the data isn't the same. There is a small misalignment of a few pixels between the 3-band and M-bands. (Presumably the class polygons are draw with respect to the 3-band images.) And, although visually similar at large scales, when you zoom in the M-band images look fuzzier (i.e. lower resolution) than the 3-bands.</p>\n\n<p>(attached are views of the top left corder or 6100_1_3.  Red band and the corresponding band from Multispectral data.)</p>\n\n<p>So we've decided to keep all the data. We pad and resize all the images to the same size, realign the A, M and P bands to the mean of the 3-band data, and then stack them all into a 20-layer sandwich. And then add another 10 layers for the category masks on top of that.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 156989,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "01/18/2017 20:54:14",
      "content": "<p>I agree with you, I was responding to the \"wavelenghts\" part.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157163,
      "author_name": "puneetjindal",
      "author_url": "",
      "post_date": "01/19/2017 14:14:07",
      "content": "<p>@amaia: Thanks for the explanation however i still have question: if it is 17 then why we say it as sixteen band data?</p>\n\n<p>According to me, CNN should work only on 3 band(RGB) images as these are true color images. </p>\n\n<p>Rest is converting all the images to spectral signatures and using train dataset polygon from the images-class mapping to be fed into the model which shall be able to predict for test as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157170,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "01/19/2017 14:37:28",
      "content": "<p>Threeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.</p>\n\n<p>But i am just bouncing ideas here.</p>\n\n<p>Puneet, of course they work. Popular nets are for rgb but 20 channel input shouldn't be a problem.\nFrom what images I saw in forums i think you can predict several labels only with thresholds and some band arithmetic.</p>\n\n<p>Cristi</p>",
      "votes": null,
      "replies": [
        {
          "id": 157227,
          "author_name": "threeplusone",
          "author_url": "",
          "post_date": "01/19/2017 19:15:42",
          "content": "<p>[quote=visoft;157170]</p>\n\n<p>Threeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.</p>\n\n<p>[/quote]</p>\n\n<p>That sounds about right to me, although we're using a sigmoid to give 10 outputs bits.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157245,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "01/19/2017 21:15:24",
      "content": "<p>Opencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \n<a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2</a></p>\n\n<p>Anaconda packs opencv 3.1so no compile hussle!</p>",
      "votes": null,
      "replies": [
        {
          "id": 157273,
          "author_name": "threeplusone",
          "author_url": "",
          "post_date": "01/19/2017 23:40:04",
          "content": "<p>[quote=visoft;157245]</p>\n\n<p>Opencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \n<a href=\"https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\">https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2</a></p>\n\n<p>[/quote]</p>\n\n<p>Aha! I used your script as a template to write my own realignment code. But I overlooked that you're shifting multiple bands at once. Time for another rewrite...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 158494,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "01/27/2017 17:39:52",
      "content": "<p>Now, when I finally survived through the submission errors, I would love to repeat my initial question.</p>\n\n<p>For those that are above <strong>0.4</strong>, do you guys see a significant difference between 3 vs 16 band images for the input.</p>\n\n<p>P.S. My last submission <strong>0.36</strong> is on 3 band.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 158509,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "01/27/2017 19:01:57",
      "content": "<p>All I would say is that certain classes are difficult with only 3 bands, and misalignment is overrated if you are using deep learning.  The miss-alignment might be a bonus just saying.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "154416": "On the one hand I would expect that data in different channels is extremely correlated, but on the other hand correlated does not mean exactly the same.\r\n\r\nQuestion is: Is there really a big difference between 3 band and 16 band images?\r\n\r\nAnswers like: \r\n\r\n 1. Based on the paper XXX there should be a big difference for trees / water / roads / xxx, because in the infrared they are very distinguishable.\r\n 2. My current model is highly / not at all affected by a switch from 3 to 16 channels.\r\n\r\nare highly appreciated.",
    "154450": "3) Based on 3 year remote sensing (ML based) experience, YES! Regular earth observation satellites have specific wavelengths to detect certain physical phenomena. Oxygen absorption band, chlorophyll fluorescence, temperature bands, etc. etc. etc. Now, I didn't do *any* reading on how the WorldView bands were designated. \r\n\r\nI noted that in some bands (from 16 bands), the trees/vegetation are brighter than the rest of the features. This must be true for water (probably another channel) and maybe some land types too (eg. roads are probably hotter (brighter in infrared) and less bright in UV. \r\n\r\nBut this is based on my previous experience, I didn't do any statistics on current dataset. Anyway, you can't beat a system designated to directly observe the physical phenomena (eg. extract the chlorophyll levels when you don't have dedicated bands for that)\r\n\r\nCristi",
    "154486": "Yes, there is a big difference. That's the reason for their existence. The first difference is due to the wavelength. As Cristi said different materials have different fingerprints (roads made of asphalt vs other \"roads\" look differently.\r\n\r\nImportant to notice a wider spectrum doesn't come for free. Multispectral bands have a lower resolution (each pixel cover a greater land area), while P have a res of 0.3m which is great for fine details.\r\nAlso number of bits, each sensor havs a different number of bitsfot each pxel, we havd more than 8 bits of information.\r\n\r\nIMHO, regardless of the approach you choose, you can make use of all that information. People using DL does not need to restrain to RGB or 3 channels only.",
    "154500": "I think if you are using CNN's there is probably not a big difference, as those rely a lot on edges which are best defined in the 3-band true color image. But there is a lot that can be done using pixel wise classification, which rely heavily on the spectral properties as @amaia and @visoft have pointed out.",
    "154532": "Thank you. \r\n\r\nI would like to concatenate all different bands and make an array that can be fed into the CNN. I am looking at image **6010_0_0**\r\n\r\n - 3 band => shape (3, 3349, 3396)\r\n - A => shape (8, 134, 136)\r\n - M => shape (8, 837, 849)\r\n - P => shape (3348, 3396)\r\n\r\n[1] Is it true that P - is just grayscale, meaning linear combination version of *3 band*?\r\n\r\n[2] I have tried to extract RGB image from *M* following https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/26763/can-we-get-the-frequency-of-the-diffrent-sixteen-band\r\nbut I did not have success with it. Are you able to do it?\r\n\r\n[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than\r\n\r\n`np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))`\r\n\r\n?",
    "154572": "Vladimir, maybe this will help:\r\n\r\nhttps://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2/code\r\n\r\nIn the example above there are only 2 images concatenated but of course it is easy extendable. And it takes care of A-P displacement. Feel free to register M to P (or _3) also.\r\n\r\nOn https://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/export-pixel-wise-mask  you will find how to build pixel-wise masks from polygons.\r\n\r\nHope it helps!\r\n\r\np.s. Please note that different scenes have different image sizes and form factors! (aka not all P images have the same size)",
    "154672": "Check out this video for a good intro on what the different bands mean and how they are different.   \r\n\r\nhttps://www.youtube.com/watch?v=3iaFzafWJQE",
    "156900": "Hi All, \r\n\r\nThis is the first competition that I am entering at Kaggle and so my question might be too naive for you experts.\r\n\r\nAbove video shared by @shawn is good place to start. But I want to go to deeper level understanding. After going through multiple threads, I have understood bits and pieces but now I am confused on **whether RGB bands and A,M,P are NULL in intersection or not.**\r\n \r\nI see that RGB overlaps with M at few wavelengths.\r\nPlease correct me if I am wrong",
    "156939": "[quote=Puneet Jindal;156900]\r\n\r\nPlease correct me if I am wrong\r\n\r\n[/quote]\r\n\r\nThis satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.",
    "156983": "[quote=amaia;156939]\r\n\r\nThis satellite works with 17 wavebands (8 + 8 + 1). The three-band natural color image is built using those.\r\n\r\n[/quote]\r\n\r\nThat's correct in principle. The RGB bands of the 3-band data correspond to 3 of the Multispectral bands.  But in practice the data isn't the same. There is a small misalignment of a few pixels between the 3-band and M-bands. (Presumably the class polygons are draw with respect to the 3-band images.) And, although visually similar at large scales, when you zoom in the M-band images look fuzzier (i.e. lower resolution) than the 3-bands.\r\n\r\n(attached are views of the top left corder or 6100_1_3.  Red band and the corresponding band from Multispectral data.)\r\n\r\nSo we've decided to keep all the data. We pad and resize all the images to the same size, realign the A, M and P bands to the mean of the 3-band data, and then stack them all into a 20-layer sandwich. And then add another 10 layers for the category masks on top of that.",
    "156989": "I agree with you, I was responding to the \"wavelenghts\" part.",
    "157163": "amaia: Thanks for the explanation however i still have question: if it is 17 then why we say it as sixteen band data?\r\n\r\nAccording to me, CNN should work only on 3 band(RGB) images as these are true color images. \r\n\r\nRest is converting all the images to spectral signatures and using train dataset polygon from the images-class mapping to be fed into the model which shall be able to predict for test as well.",
    "157170": "Threeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.\r\n\r\nBut i am just bouncing ideas here.\r\n\r\nPuneet, of course they work. Popular nets are for rgb but 20 channel input shouldn't be a problem.\r\nFrom what images I saw in forums i think you can predict several labels only with thresholds and some band arithmetic.\r\n\r\nCristi",
    "157227": "[quote=visoft;157170]\r\n\r\nThreeplusone what objective/loss do you use? I ask because you have only 10 output layers. While I am not there, I was thinking to have a softmax+crossentropy for each lanel. That should ammount to 2x10 output \"channels\". The architecture should be common up to the last layer.\r\n\r\n[/quote]\r\n\r\nThat sounds about right to me, although we're using a sigmoid to give 10 outputs bits.",
    "157228": "[quote=Vladimir Iglovikov;154532]\r\n\r\n[3] More technical question. I would love to stack all these images into 20 layer sandwich, but all images have different shape. Do you know a way to rescale images that have more than 3 channels in python that is more straightforward than\r\n\r\n`np.transpose(cv2.resize(np.transpose(A, (1, 2, 0)), (3396, 3349)), (2, 0, 1))`\r\n\r\n[/quote]\r\n\r\nNot that I've found. Everything seems to assume greyscale or RGB. So I'm rescaling one band at a time. This is slow, but in principle only has to be done once.",
    "157245": "Opencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \r\nhttps://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\r\n\r\nAnaconda packs opencv 3.1so no compile hussle!",
    "157273": "[quote=visoft;157245]\r\n\r\nOpencv works well up to (i think) 32 bands. I posted a script doing both rescaling and a rigid registration btween bands. \r\nhttps://www.kaggle.com/visoft/dstl-satellite-imagery-feature-detection/correct-image-missalignment-v2\r\n\r\n[/quote]\r\n\r\nAha! I used your script as a template to write my own realignment code. But I overlooked that you're shifting multiple bands at once. Time for another rewrite...",
    "158494": "Now, when I finally survived through the submission errors, I would love to repeat my initial question.\r\n\r\nFor those that are above **0.4**, do you guys see a significant difference between 3 vs 16 band images for the input.\r\n\r\nP.S. My last submission **0.36** is on 3 band.",
    "158509": "All I would say is that certain classes are difficult with only 3 bands, and misalignment is overrated if you are using deep learning.  The miss-alignment might be a bonus just saying."
  },
  "source": "meta"
}