{
  "id": 29606,
  "title": "how I got LB 0.2, how to reach LB >0.3?",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/29606",
  "author_name": "",
  "post_date": "2017-03-05T09:39:38.609810600Z",
  "votes": null,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Let have some discussions on how to do better in few days left:</p>\n\n<p>I current use Unet (by n01z3 code) with 3 brands dataset:\n3 brand dataset is down sampled to 0.37 of original, hence I got around 800 pixels.\nUnet uses input of 160x160 image cropped from 3 brands dataset.</p>\n\n<p>The most I can get is LB 0.22.\nIf i use 8 channel M brands, the most I can get is LB 0.20.</p>\n\n<p>I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.</p>\n\n<p>I run out of idea on how to get LB 0.3, given some claim we can use unet+3 brands to get 0.36.. Any hints you think is ok to share?\nThank a lot.</p>",
  "messages": [
    {
      "id": "165365",
      "postDate": "03/05/2017 09:39:38",
      "content": "<p>Let have some discussions on how to do better in few days left:</p>\n\n<p>I current use Unet (by n01z3 code) with 3 brands dataset:\n3 brand dataset is down sampled to 0.37 of original, hence I got around 800 pixels.\nUnet uses input of 160x160 image cropped from 3 brands dataset.</p>\n\n<p>The most I can get is LB 0.22.\nIf i use 8 channel M brands, the most I can get is LB 0.20.</p>\n\n<p>I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.</p>\n\n<p>I run out of idea on how to get LB 0.3, given some claim we can use unet+3 brands to get 0.36.. Any hints you think is ok to share?\nThank a lot.</p>",
      "rawMarkdown": "Let have some discussions on how to do better in few days left:\n\nI current use Unet (by n01z3 code) with 3 brands dataset:\n3 brand dataset is down sampled to 0.37 of original, hence I got around 800 pixels.\nUnet uses input of 160x160 image cropped from 3 brands dataset.\n\nThe most I can get is LB 0.22.\nIf i use 8 channel M brands, the most I can get is LB 0.20.\n\nI tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.\n\nI run out of idea on how to get LB 0.3, given some claim we can use unet+3 brands to get 0.36.. Any hints you think is ok to share?\nThank a lot.",
      "votes": null
    },
    {
      "id": "165374",
      "postDate": "03/05/2017 11:08:37",
      "content": "<p>To reach LB &gt; 0.3, you need to deal with the unbalance problem of different classes. </p>",
      "rawMarkdown": "To reach LB > 0.3, you need to deal with the unbalance problem of different classes.",
      "votes": null
    },
    {
      "id": "165375",
      "postDate": "03/05/2017 11:13:51",
      "content": "<p>\"I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image.\"</p>\n\n<p>For this add upscaled bands of M (I didn't repeat the RGB though). Try a patch size around 200, I am using 256.</p>",
      "rawMarkdown": "\"I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image.\"\n\nFor this add upscaled bands of M (I didn't repeat the RGB though). Try a patch size around 200, I am using 256.",
      "votes": null
    },
    {
      "id": "165386",
      "postDate": "03/05/2017 11:50:37",
      "content": "<p>First, thank a lot for reply.</p>\n\n<p>Do you mean keep the RGB as original resolution 3k x 3k. Then upscale M bands to 3k x 3k too. Then crop it into 11x320x320 image, and use patch size of 256 for unet training?</p>\n\n<p>I am thinking if M bands and RGB bands need to do alignment? I read post that without alignment can still get &gt;0.3.</p>\n\n<p>Patch size of 256 is quite large, I am using a K80 with 12G ram, I can use 64 patch size with 3x160x160, and and only 32 when the image is 17x160x160 (P.S. I tried with all 17 channels and got worse result..) What config/trick you are using such that patch size of 256 is possible?</p>\n\n<p>Thank a lot for sharing and advice.</p>",
      "rawMarkdown": "First, thank a lot for reply.\n\nDo you mean keep the RGB as original resolution 3k x 3k. Then upscale M bands to 3k x 3k too. Then crop it into 11x320x320 image, and use patch size of 256 for unet training?\n\nI am thinking if M bands and RGB bands need to do alignment? I read post that without alignment can still get >0.3.\n\nPatch size of 256 is quite large, I am using a K80 with 12G ram, I can use 64 patch size with 3x160x160, and and only 32 when the image is 17x160x160 (P.S. I tried with all 17 channels and got worse result..) What config/trick you are using such that patch size of 256 is possible?\n\nThank a lot for sharing and advice.",
      "votes": null
    },
    {
      "id": "165387",
      "postDate": "03/05/2017 11:53:43",
      "content": "<p>Thank for advice. </p>\n\n<p>Can you tell me more on conceptual idea to deal with unbalance problem issue.</p>\n\n<p>When I did nothing on unbalance problem, I still got reasonable good accuracy on most class. But indeed, my LB score is poor even I got 0.8954 px level overall accuracy.</p>\n\n<p>0 0.855698007325 0.3\n1 0.748657416246 0.3\n2 0.974797826147 0.6\n3 0.711063017797 0.3\n4 0.76722388704 0.6\n5 0.973309108224 0.7\n6 0.994745028157 0.7\n7 0.996887983821 0.2\n8 0.986518691589 0.4\n9 0.945433176315 0.3\nunet_3c_10_jk0.8954</p>\n\n<p>Thank a lot for sharing.</p>",
      "rawMarkdown": "Thank for advice. \n\nCan you tell me more on conceptual idea to deal with unbalance problem issue.\n\nWhen I did nothing on unbalance problem, I still got reasonable good accuracy on most class. But indeed, my LB score is poor even I got 0.8954 px level overall accuracy.\n\n0 0.855698007325 0.3\n1 0.748657416246 0.3\n2 0.974797826147 0.6\n3 0.711063017797 0.3\n4 0.76722388704 0.6\n5 0.973309108224 0.7\n6 0.994745028157 0.7\n7 0.996887983821 0.2\n8 0.986518691589 0.4\n9 0.945433176315 0.3\nunet_3c_10_jk0.8954\n\nThank a lot for sharing.",
      "votes": null
    },
    {
      "id": "165388",
      "postDate": "03/05/2017 12:06:06",
      "content": "<p>Hi! I (re)started working on the competition this weekend.</p>\n\n<p>My best entry is at 0.16 for 5 classes (easy ones, trees, waterway, houses, crops and roads).\nI used full pancromatic resolution with the other bands scaled and registered(!) to pancromatic. This includes RGB bands (redundant but doh, they help me visualize stuff) and 4 additional derived params (NVDI and some dirt/vegetation indexes, what I found discussed on the forums/stackoverflow) So 24 bands in total.</p>\n\n<p>I used randomly cut 256x256 patches on a custom unet. I didn't try the n01z3 net, I don't think I will have time. But I did one net per class. I got far better results. Probably I have architecture issues, the architecture is \"recycled\" from another project. </p>\n\n<p>My problem is that with that many rasters, it takes about 2hrs to train a model with ~100 epochs (gtx980ti) and another 3 hrs to generate a solution. Not much time to do architecture tuning or CV. I barely did one submission this weekend.</p>",
      "rawMarkdown": "Hi! I (re)started working on the competition this weekend.\n\nMy best entry is at 0.16 for 5 classes (easy ones, trees, waterway, houses, crops and roads).\nI used full pancromatic resolution with the other bands scaled and registered(!) to pancromatic. This includes RGB bands (redundant but doh, they help me visualize stuff) and 4 additional derived params (NVDI and some dirt/vegetation indexes, what I found discussed on the forums/stackoverflow) So 24 bands in total.\n\nI used randomly cut 256x256 patches on a custom unet. I didn't try the n01z3 net, I don't think I will have time. But I did one net per class. I got far better results. Probably I have architecture issues, the architecture is \"recycled\" from another project. \n\nMy problem is that with that many rasters, it takes about 2hrs to train a model with ~100 epochs (gtx980ti) and another 3 hrs to generate a solution. Not much time to do architecture tuning or CV. I barely did one submission this weekend.",
      "votes": null
    },
    {
      "id": "165390",
      "postDate": "03/05/2017 12:11:26",
      "content": "<p>I think you didn't understand the problem. The accuracy is not a good indicator in this competition. For example for small cars, when you predict all the pixel 0, which means all pixels are background, no cars, you can still get one accuracy like 0.98+, but obviously it's meaningless. </p>",
      "rawMarkdown": "I think you didn't understand the problem. The accuracy is not a good indicator in this competition. For example for small cars, when you predict all the pixel 0, which means all pixels are background, no cars, you can still get one accuracy like 0.98+, but obviously it's meaningless.",
      "votes": null
    },
    {
      "id": "165398",
      "postDate": "03/05/2017 12:53:08",
      "content": "<p>i get your point. that is definite a issue for higher accuracy.  But I just guess to &gt;0.3, the good balance issue is not yet the bottleneck. \nWith 0.8954 accuracy, but the LB is just 0.2, something very wrong, but no idea what to do to analyze/improve it..\nThx a lot.</p>",
      "rawMarkdown": "i get your point. that is definite a issue for higher accuracy.  But I just guess to >0.3, the good balance issue is not yet the bottleneck. \nWith 0.8954 accuracy, but the LB is just 0.2, something very wrong, but no idea what to do to analyze/improve it..\nThx a lot.",
      "votes": null
    },
    {
      "id": "165399",
      "postDate": "03/05/2017 12:56:50",
      "content": "<p>Great to see you again, I read your kernels and learned a lot!</p>\n\n<p>I am thinking that if I can read good on RGB bands, then I should be able to use deeplearning to solve it. I tried used 17 bands before, but I could only train at 16 patches which is not good..</p>\n\n<p>I am thinking and trying architecture tuning, but I got no good idea to do tuning..</p>",
      "rawMarkdown": "Great to see you again, I read your kernels and learned a lot!\n\nI am thinking that if I can read good on RGB bands, then I should be able to use deeplearning to solve it. I tried used 17 bands before, but I could only train at 16 patches which is not good..\n\nI am thinking and trying architecture tuning, but I got no good idea to do tuning..",
      "votes": null
    },
    {
      "id": "165404",
      "postDate": "03/05/2017 13:18:21",
      "content": "<p>when \"CNMeM is disabled\", i can trained with 128 patches!!</p>",
      "rawMarkdown": "when \"CNMeM is disabled\", i can trained with 128 patches!!",
      "votes": null
    },
    {
      "id": "165405",
      "postDate": "03/05/2017 13:22:37",
      "content": "<p>also, I think using 160x160 is not good.. but i have not got a good idea to deal with this \"block and discontinuous\" effect.. any good suggestion? Thank a lot.</p>\n\n<p><a href=\"http://imgur.com/a/QulYk\">http://imgur.com/a/QulYk</a></p>",
      "rawMarkdown": "also, I think using 160x160 is not good.. but i have not got a good idea to deal with this \"block and discontinuous\" effect.. any good suggestion? Thank a lot.\n\nhttp://imgur.com/a/QulYk",
      "votes": null
    },
    {
      "id": "165434",
      "postDate": "03/05/2017 16:43:17",
      "content": "<p>n01z3 kernel is a great baseline, however, it contains some very limiting defaults (unintended, it is just quick and dirty first try). First of all, stretch_n function is great for viewing but unnecessary for network learning and with parameters in kernel it throws away 10% (!) of most important information (because it cuts the most bright and dark pixels). If you didn't fixed it -- this change alone could give you 0.3+ \nSecond, default mask to polygon routine uses very high epsilon of 5 which has extremely degrading effect on small objects like individual trees or buildings. Use 1 and get another boost. \nThird, it may take a lot of luck to get good initialization at learning start. Two separate trainings of the same network with the same parameters could give vastly different results, unfortunately. </p>\n\n<p>And last, about full sized RGB: it requires a lot more training time, fine tuning and even more luck to get decent results from RGB, because there are a lot more spatial information but a lot less spectral information. And network gets overwhelmed by this (less relevant) info (and your GPU couldn't support batch size big enough to smooth out training and so on). It is useless to downsample RGB because it was originally upsampled from part of the M-band images. Also, some channels in M images are extremely useful, for example, NIR channel alone contains enough information to highlight vegetation (especially trees) and water. </p>",
      "rawMarkdown": "n01z3 kernel is a great baseline, however, it contains some very limiting defaults (unintended, it is just quick and dirty first try). First of all, stretch_n function is great for viewing but unnecessary for network learning and with parameters in kernel it throws away 10% (!) of most important information (because it cuts the most bright and dark pixels). If you didn't fixed it -- this change alone could give you 0.3+ \nSecond, default mask to polygon routine uses very high epsilon of 5 which has extremely degrading effect on small objects like individual trees or buildings. Use 1 and get another boost. \nThird, it may take a lot of luck to get good initialization at learning start. Two separate trainings of the same network with the same parameters could give vastly different results, unfortunately. \n\nAnd last, about full sized RGB: it requires a lot more training time, fine tuning and even more luck to get decent results from RGB, because there are a lot more spatial information but a lot less spectral information. And network gets overwhelmed by this (less relevant) info (and your GPU couldn't support batch size big enough to smooth out training and so on). It is useless to downsample RGB because it was originally upsampled from part of the M-band images. Also, some channels in M images are extremely useful, for example, NIR channel alone contains enough information to highlight vegetation (especially trees) and water.",
      "votes": null
    },
    {
      "id": "165451",
      "postDate": "03/05/2017 18:15:50",
      "content": "<p>In my experience inbalance is not an issue if you are using jaccard loss.</p>",
      "rawMarkdown": "In my experience inbalance is not an issue if you are using jaccard loss.",
      "votes": null
    },
    {
      "id": "165458",
      "postDate": "03/05/2017 18:51:20",
      "content": "<p>Hey sorry for the late reply. I got carried away with training the models :P\nWhat I meant is, resize M bands(i.e. non RGB). concatenate to RGB. Cut your model sized patches i.e. 10x256x256 as the data. (Since a class might not exist everywhere you can try to cut 50% of data patches locating around a class). 128 didn't help me much. May be you can try for 192, 224.</p>",
      "rawMarkdown": "Hey sorry for the late reply. I got carried away with training the models :P\nWhat I meant is, resize M bands(i.e. non RGB). concatenate to RGB. Cut your model sized patches i.e. 10x256x256 as the data. (Since a class might not exist everywhere you can try to cut 50% of data patches locating around a class). 128 didn't help me much. May be you can try for 192, 224.",
      "votes": null
    },
    {
      "id": "165459",
      "postDate": "03/05/2017 18:53:56",
      "content": "<p>for 256 patch size I am using 16 as batch_size.\nFor alignment you can simply stitch each 5x5 images, resize to that of 5x5 RGB and cut again as per RGB dimensions.</p>",
      "rawMarkdown": "for 256 patch size I am using 16 as batch_size.\nFor alignment you can simply stitch each 5x5 images, resize to that of 5x5 RGB and cut again as per RGB dimensions.",
      "votes": null
    },
    {
      "id": "165471",
      "postDate": "03/05/2017 20:02:43",
      "content": "<blockquote>\n  <p><strong>CRaymond wrote</strong></p>\n  \n  <p>I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.</p>\n</blockquote>\n\n<p>I just submitted predictions created from RGB 320X320 patches with UNET and got 0.37. I am using a deeper unet with custom loss in mxnet.</p>",
      "rawMarkdown": "> **CRaymond wrote**\n> \n\n> I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.\n> \n\nI just submitted predictions created from RGB 320X320 patches with UNET and got 0.37. I am using a deeper unet with custom loss in mxnet.",
      "votes": null
    },
    {
      "id": "165535",
      "postDate": "03/06/2017 04:15:33",
      "content": "<p>how do you know the 3-band RGB is upsampled from part of M-band images ?</p>",
      "rawMarkdown": "how do you know the 3-band RGB is upsampled from part of M-band images ?",
      "votes": null
    },
    {
      "id": "165567",
      "postDate": "03/06/2017 07:54:39",
      "content": "<p>Because Worldview-3 doesn't provide any \"RGB\" data besides M-band. RGB was made with pansharpening from P and part of M that contains visible light. </p>",
      "rawMarkdown": "Because Worldview-3 doesn't provide any \"RGB\" data besides M-band. RGB was made with pansharpening from P and part of M that contains visible light.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 165374,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "03/05/2017 11:08:37",
      "content": "<p>To reach LB &gt; 0.3, you need to deal with the unbalance problem of different classes. </p>",
      "votes": null,
      "replies": [
        {
          "id": 165387,
          "author_name": "craymond",
          "author_url": "",
          "post_date": "03/05/2017 11:53:43",
          "content": "<p>Thank for advice. </p>\n\n<p>Can you tell me more on conceptual idea to deal with unbalance problem issue.</p>\n\n<p>When I did nothing on unbalance problem, I still got reasonable good accuracy on most class. But indeed, my LB score is poor even I got 0.8954 px level overall accuracy.</p>\n\n<p>0 0.855698007325 0.3\n1 0.748657416246 0.3\n2 0.974797826147 0.6\n3 0.711063017797 0.3\n4 0.76722388704 0.6\n5 0.973309108224 0.7\n6 0.994745028157 0.7\n7 0.996887983821 0.2\n8 0.986518691589 0.4\n9 0.945433176315 0.3\nunet_3c_10_jk0.8954</p>\n\n<p>Thank a lot for sharing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 165390,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "03/05/2017 12:11:26",
          "content": "<p>I think you didn't understand the problem. The accuracy is not a good indicator in this competition. For example for small cars, when you predict all the pixel 0, which means all pixels are background, no cars, you can still get one accuracy like 0.98+, but obviously it's meaningless. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 165398,
          "author_name": "craymond",
          "author_url": "",
          "post_date": "03/05/2017 12:53:08",
          "content": "<p>i get your point. that is definite a issue for higher accuracy.  But I just guess to &gt;0.3, the good balance issue is not yet the bottleneck. \nWith 0.8954 accuracy, but the LB is just 0.2, something very wrong, but no idea what to do to analyze/improve it..\nThx a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 165451,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "03/05/2017 18:15:50",
          "content": "<p>In my experience inbalance is not an issue if you are using jaccard loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 165375,
      "author_name": "kghanareddy",
      "author_url": "",
      "post_date": "03/05/2017 11:13:51",
      "content": "<p>\"I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image.\"</p>\n\n<p>For this add upscaled bands of M (I didn't repeat the RGB though). Try a patch size around 200, I am using 256.</p>",
      "votes": null,
      "replies": [
        {
          "id": 165386,
          "author_name": "craymond",
          "author_url": "",
          "post_date": "03/05/2017 11:50:37",
          "content": "<p>First, thank a lot for reply.</p>\n\n<p>Do you mean keep the RGB as original resolution 3k x 3k. Then upscale M bands to 3k x 3k too. Then crop it into 11x320x320 image, and use patch size of 256 for unet training?</p>\n\n<p>I am thinking if M bands and RGB bands need to do alignment? I read post that without alignment can still get &gt;0.3.</p>\n\n<p>Patch size of 256 is quite large, I am using a K80 with 12G ram, I can use 64 patch size with 3x160x160, and and only 32 when the image is 17x160x160 (P.S. I tried with all 17 channels and got worse result..) What config/trick you are using such that patch size of 256 is possible?</p>\n\n<p>Thank a lot for sharing and advice.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 165404,
          "author_name": "craymond",
          "author_url": "",
          "post_date": "03/05/2017 13:18:21",
          "content": "<p>when \"CNMeM is disabled\", i can trained with 128 patches!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 165458,
          "author_name": "kghanareddy",
          "author_url": "",
          "post_date": "03/05/2017 18:51:20",
          "content": "<p>Hey sorry for the late reply. I got carried away with training the models :P\nWhat I meant is, resize M bands(i.e. non RGB). concatenate to RGB. Cut your model sized patches i.e. 10x256x256 as the data. (Since a class might not exist everywhere you can try to cut 50% of data patches locating around a class). 128 didn't help me much. May be you can try for 192, 224.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 165459,
          "author_name": "kghanareddy",
          "author_url": "",
          "post_date": "03/05/2017 18:53:56",
          "content": "<p>for 256 patch size I am using 16 as batch_size.\nFor alignment you can simply stitch each 5x5 images, resize to that of 5x5 RGB and cut again as per RGB dimensions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 165388,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "03/05/2017 12:06:06",
      "content": "<p>Hi! I (re)started working on the competition this weekend.</p>\n\n<p>My best entry is at 0.16 for 5 classes (easy ones, trees, waterway, houses, crops and roads).\nI used full pancromatic resolution with the other bands scaled and registered(!) to pancromatic. This includes RGB bands (redundant but doh, they help me visualize stuff) and 4 additional derived params (NVDI and some dirt/vegetation indexes, what I found discussed on the forums/stackoverflow) So 24 bands in total.</p>\n\n<p>I used randomly cut 256x256 patches on a custom unet. I didn't try the n01z3 net, I don't think I will have time. But I did one net per class. I got far better results. Probably I have architecture issues, the architecture is \"recycled\" from another project. </p>\n\n<p>My problem is that with that many rasters, it takes about 2hrs to train a model with ~100 epochs (gtx980ti) and another 3 hrs to generate a solution. Not much time to do architecture tuning or CV. I barely did one submission this weekend.</p>",
      "votes": null,
      "replies": [
        {
          "id": 165399,
          "author_name": "craymond",
          "author_url": "",
          "post_date": "03/05/2017 12:56:50",
          "content": "<p>Great to see you again, I read your kernels and learned a lot!</p>\n\n<p>I am thinking that if I can read good on RGB bands, then I should be able to use deeplearning to solve it. I tried used 17 bands before, but I could only train at 16 patches which is not good..</p>\n\n<p>I am thinking and trying architecture tuning, but I got no good idea to do tuning..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 165405,
      "author_name": "craymond",
      "author_url": "",
      "post_date": "03/05/2017 13:22:37",
      "content": "<p>also, I think using 160x160 is not good.. but i have not got a good idea to deal with this \"block and discontinuous\" effect.. any good suggestion? Thank a lot.</p>\n\n<p><a href=\"http://imgur.com/a/QulYk\">http://imgur.com/a/QulYk</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 165434,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "03/05/2017 16:43:17",
      "content": "<p>n01z3 kernel is a great baseline, however, it contains some very limiting defaults (unintended, it is just quick and dirty first try). First of all, stretch_n function is great for viewing but unnecessary for network learning and with parameters in kernel it throws away 10% (!) of most important information (because it cuts the most bright and dark pixels). If you didn't fixed it -- this change alone could give you 0.3+ \nSecond, default mask to polygon routine uses very high epsilon of 5 which has extremely degrading effect on small objects like individual trees or buildings. Use 1 and get another boost. \nThird, it may take a lot of luck to get good initialization at learning start. Two separate trainings of the same network with the same parameters could give vastly different results, unfortunately. </p>\n\n<p>And last, about full sized RGB: it requires a lot more training time, fine tuning and even more luck to get decent results from RGB, because there are a lot more spatial information but a lot less spectral information. And network gets overwhelmed by this (less relevant) info (and your GPU couldn't support batch size big enough to smooth out training and so on). It is useless to downsample RGB because it was originally upsampled from part of the M-band images. Also, some channels in M images are extremely useful, for example, NIR channel alone contains enough information to highlight vegetation (especially trees) and water. </p>",
      "votes": null,
      "replies": [
        {
          "id": 165535,
          "author_name": "zeliek",
          "author_url": "",
          "post_date": "03/06/2017 04:15:33",
          "content": "<p>how do you know the 3-band RGB is upsampled from part of M-band images ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 165567,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "03/06/2017 07:54:39",
          "content": "<p>Because Worldview-3 doesn't provide any \"RGB\" data besides M-band. RGB was made with pansharpening from P and part of M that contains visible light. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 165471,
      "author_name": "aakansh9",
      "author_url": "",
      "post_date": "03/05/2017 20:02:43",
      "content": "<blockquote>\n  <p><strong>CRaymond wrote</strong></p>\n  \n  <p>I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.</p>\n</blockquote>\n\n<p>I just submitted predictions created from RGB 320X320 patches with UNET and got 0.37. I am using a deeper unet with custom loss in mxnet.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "165365": "Let have some discussions on how to do better in few days left:\n\nI current use Unet (by n01z3 code) with 3 brands dataset:\n3 brand dataset is down sampled to 0.37 of original, hence I got around 800 pixels.\nUnet uses input of 160x160 image cropped from 3 brands dataset.\n\nThe most I can get is LB 0.22.\nIf i use 8 channel M brands, the most I can get is LB 0.20.\n\nI tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.\n\nI run out of idea on how to get LB 0.3, given some claim we can use unet+3 brands to get 0.36.. Any hints you think is ok to share?\nThank a lot.",
    "165374": "To reach LB > 0.3, you need to deal with the unbalance problem of different classes.",
    "165375": "\"I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image.\"\n\nFor this add upscaled bands of M (I didn't repeat the RGB though). Try a patch size around 200, I am using 256.",
    "165386": "First, thank a lot for reply.\n\nDo you mean keep the RGB as original resolution 3k x 3k. Then upscale M bands to 3k x 3k too. Then crop it into 11x320x320 image, and use patch size of 256 for unet training?\n\nI am thinking if M bands and RGB bands need to do alignment? I read post that without alignment can still get >0.3.\n\nPatch size of 256 is quite large, I am using a K80 with 12G ram, I can use 64 patch size with 3x160x160, and and only 32 when the image is 17x160x160 (P.S. I tried with all 17 channels and got worse result..) What config/trick you are using such that patch size of 256 is possible?\n\nThank a lot for sharing and advice.",
    "165387": "Thank for advice. \n\nCan you tell me more on conceptual idea to deal with unbalance problem issue.\n\nWhen I did nothing on unbalance problem, I still got reasonable good accuracy on most class. But indeed, my LB score is poor even I got 0.8954 px level overall accuracy.\n\n0 0.855698007325 0.3\n1 0.748657416246 0.3\n2 0.974797826147 0.6\n3 0.711063017797 0.3\n4 0.76722388704 0.6\n5 0.973309108224 0.7\n6 0.994745028157 0.7\n7 0.996887983821 0.2\n8 0.986518691589 0.4\n9 0.945433176315 0.3\nunet_3c_10_jk0.8954\n\nThank a lot for sharing.",
    "165388": "Hi! I (re)started working on the competition this weekend.\n\nMy best entry is at 0.16 for 5 classes (easy ones, trees, waterway, houses, crops and roads).\nI used full pancromatic resolution with the other bands scaled and registered(!) to pancromatic. This includes RGB bands (redundant but doh, they help me visualize stuff) and 4 additional derived params (NVDI and some dirt/vegetation indexes, what I found discussed on the forums/stackoverflow) So 24 bands in total.\n\nI used randomly cut 256x256 patches on a custom unet. I didn't try the n01z3 net, I don't think I will have time. But I did one net per class. I got far better results. Probably I have architecture issues, the architecture is \"recycled\" from another project. \n\nMy problem is that with that many rasters, it takes about 2hrs to train a model with ~100 epochs (gtx980ti) and another 3 hrs to generate a solution. Not much time to do architecture tuning or CV. I barely did one submission this weekend.",
    "165390": "I think you didn't understand the problem. The accuracy is not a good indicator in this competition. For example for small cars, when you predict all the pixel 0, which means all pixels are background, no cars, you can still get one accuracy like 0.98+, but obviously it's meaningless.",
    "165398": "i get your point. that is definite a issue for higher accuracy.  But I just guess to >0.3, the good balance issue is not yet the bottleneck. \nWith 0.8954 accuracy, but the LB is just 0.2, something very wrong, but no idea what to do to analyze/improve it..\nThx a lot.",
    "165399": "Great to see you again, I read your kernels and learned a lot!\n\nI am thinking that if I can read good on RGB bands, then I should be able to use deeplearning to solve it. I tried used 17 bands before, but I could only train at 16 patches which is not good..\n\nI am thinking and trying architecture tuning, but I got no good idea to do tuning..",
    "165404": "when \"CNMeM is disabled\", i can trained with 128 patches!!",
    "165405": "also, I think using 160x160 is not good.. but i have not got a good idea to deal with this \"block and discontinuous\" effect.. any good suggestion? Thank a lot.\n\nhttp://imgur.com/a/QulYk",
    "165434": "n01z3 kernel is a great baseline, however, it contains some very limiting defaults (unintended, it is just quick and dirty first try). First of all, stretch_n function is great for viewing but unnecessary for network learning and with parameters in kernel it throws away 10% (!) of most important information (because it cuts the most bright and dark pixels). If you didn't fixed it -- this change alone could give you 0.3+ \nSecond, default mask to polygon routine uses very high epsilon of 5 which has extremely degrading effect on small objects like individual trees or buildings. Use 1 and get another boost. \nThird, it may take a lot of luck to get good initialization at learning start. Two separate trainings of the same network with the same parameters could give vastly different results, unfortunately. \n\nAnd last, about full sized RGB: it requires a lot more training time, fine tuning and even more luck to get decent results from RGB, because there are a lot more spatial information but a lot less spectral information. And network gets overwhelmed by this (less relevant) info (and your GPU couldn't support batch size big enough to smooth out training and so on). It is useless to downsample RGB because it was originally upsampled from part of the M-band images. Also, some channels in M images are extremely useful, for example, NIR channel alone contains enough information to highlight vegetation (especially trees) and water.",
    "165451": "In my experience inbalance is not an issue if you are using jaccard loss.",
    "165458": "Hey sorry for the late reply. I got carried away with training the models :P\nWhat I meant is, resize M bands(i.e. non RGB). concatenate to RGB. Cut your model sized patches i.e. 10x256x256 as the data. (Since a class might not exist everywhere you can try to cut 50% of data patches locating around a class). 128 didn't help me much. May be you can try for 192, 224.",
    "165459": "for 256 patch size I am using 16 as batch_size.\nFor alignment you can simply stitch each 5x5 images, resize to that of 5x5 RGB and cut again as per RGB dimensions.",
    "165471": "> **CRaymond wrote**\n> \n\n> I tried to change not downsample RGB image (i.e. 3000x3000) with unet input to 320x320 image, but it led to worse result, LB 0.1 only.\n> \n\nI just submitted predictions created from RGB 320X320 patches with UNET and got 0.37. I am using a deeper unet with custom loss in mxnet.",
    "165535": "how do you know the 3-band RGB is upsampled from part of M-band images ?",
    "165567": "Because Worldview-3 doesn't provide any \"RGB\" data besides M-band. RGB was made with pansharpening from P and part of M that contains visible light."
  },
  "source": "meta"
}