{
  "id": 38189,
  "title": "from 0.997 to 0.998",
  "url": "/competitions/carvana-image-masking-challenge/discussion/38189",
  "author_name": "",
  "post_date": "2017-08-16T11:02:37.994705100Z",
  "votes": 28,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Let's discuss here! Here is how a 0.998 error would look like. It is about 1-pixel boundary error.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/214243/7133/0.998.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Here are my ideas:</p>\n\n<p>(1) will changing network structure (e.g. resnet, densenet for decoder/encoder) help? </p>\n\n<p>not useful unless you can increase the number of channels when you scale up. Actually, in my experiments, you still get improved results if you use more channels, even in simple UNet. But it becomes so slow and it is difficult to extend this low image resolution.</p>\n\n<p>.</p>\n\n<p>(2) How about increase resolution for input?</p>\n\n<p>There will be improvement. However, larger resolution also means deeper network. Difficult to scale up. One crazy idea is to train up 2 x actual size (i.e. larger than input resolution) </p>\n\n<p>.</p>\n\n<p>(3) weak supervised learning.</p>\n\n<p>We actually already have pretty good results on the test images  (0.997x) human error is about (0.999?).  CVPR/ICCV has some works to show that partial label can achieve good results. e.g. \"Simple does it: Weakly Supervised Instance and Semantic Segmentation\" -CVPR 2017.</p>\n\n<p>If you are considering \"512x512\", you have 0.998x accurate pseudo-label on the test set.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>In summary, this is my best bet:</p>\n\n<ul>\n<li><p>my current  best LB results 0.997, more precisely 0.99662(<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38301\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38301</a>. i need to revised the targets below later)</p></li>\n<li><p>best results attain by deep CNN (better image resolution, more parameters, more data, etc) 0.9972</p></li>\n<li><p>smart but simple post image/meta-data processing 0.9973 (Note you surely needs some method to complement deep CNN)</p></li>\n<li><p>ensemble to stabilize and make robust results? 0.99735</p></li>\n</ul>\n\n<p>... to be updated ...</p>",
  "messages": [
    {
      "id": "214243",
      "postDate": "08/16/2017 11:02:37",
      "content": "<p>Let's discuss here! Here is how a 0.998 error would look like. It is about 1-pixel boundary error.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/214243/7133/0.998.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>Here are my ideas:</p>\n\n<p>(1) will changing network structure (e.g. resnet, densenet for decoder/encoder) help? </p>\n\n<p>not useful unless you can increase the number of channels when you scale up. Actually, in my experiments, you still get improved results if you use more channels, even in simple UNet. But it becomes so slow and it is difficult to extend this low image resolution.</p>\n\n<p>.</p>\n\n<p>(2) How about increase resolution for input?</p>\n\n<p>There will be improvement. However, larger resolution also means deeper network. Difficult to scale up. One crazy idea is to train up 2 x actual size (i.e. larger than input resolution) </p>\n\n<p>.</p>\n\n<p>(3) weak supervised learning.</p>\n\n<p>We actually already have pretty good results on the test images  (0.997x) human error is about (0.999?).  CVPR/ICCV has some works to show that partial label can achieve good results. e.g. \"Simple does it: Weakly Supervised Instance and Semantic Segmentation\" -CVPR 2017.</p>\n\n<p>If you are considering \"512x512\", you have 0.998x accurate pseudo-label on the test set.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>In summary, this is my best bet:</p>\n\n<ul>\n<li><p>my current  best LB results 0.997, more precisely 0.99662(<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38301\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38301</a>. i need to revised the targets below later)</p></li>\n<li><p>best results attain by deep CNN (better image resolution, more parameters, more data, etc) 0.9972</p></li>\n<li><p>smart but simple post image/meta-data processing 0.9973 (Note you surely needs some method to complement deep CNN)</p></li>\n<li><p>ensemble to stabilize and make robust results? 0.99735</p></li>\n</ul>\n\n<p>... to be updated ...</p>",
      "rawMarkdown": "Let's discuss here! Here is how a 0.998 error would look like. It is about 1-pixel boundary error.\n\n ![enter image description here][1]\n\n\nHere are my ideas:\n\n(1) will changing network structure (e.g. resnet, densenet for decoder/encoder) help? \n\nnot useful unless you can increase the number of channels when you scale up. Actually, in my experiments, you still get improved results if you use more channels, even in simple UNet. But it becomes so slow and it is difficult to extend this low image resolution.\n\n.\n\n\n(2) How about increase resolution for input?\n\nThere will be improvement. However, larger resolution also means deeper network. Difficult to scale up. One crazy idea is to train up 2 x actual size (i.e. larger than input resolution) \n\n.\n\n(3) weak supervised learning.\n\nWe actually already have pretty good results on the test images  (0.997x) human error is about (0.999?).  CVPR/ICCV has some works to show that partial label can achieve good results. e.g. \"Simple does it: Weakly Supervised Instance and Semantic Segmentation\" -CVPR 2017.\n\nIf you are considering \"512x512\", you have 0.998x accurate pseudo-label on the test set.\n\n\n.\n\n.\n\n.\n\n\nIn summary, this is my best bet:\n\n- my current  best LB results 0.997, more precisely 0.99662(https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38301. i need to revised the targets below later)\n\n- best results attain by deep CNN (better image resolution, more parameters, more data, etc) 0.9972\n\n- smart but simple post image/meta-data processing 0.9973 (Note you surely needs some method to complement deep CNN)\n\n- ensemble to stabilize and make robust results? 0.99735\n \n\n\n\n... to be updated ...\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/214243/7133/0.998.png",
      "votes": null
    },
    {
      "id": "214250",
      "postDate": "08/16/2017 11:43:41",
      "content": "<p>wow。。from 0.99X to 1</p>",
      "rawMarkdown": "wow。。from 0.99X to 1",
      "votes": null
    },
    {
      "id": "214263",
      "postDate": "08/16/2017 12:28:06",
      "content": "<ol>\n<li>wheel spokes fix </li>\n<li>SUVs fix: roof and \"down sidebars\" </li>\n<li>light car color preprocess </li>\n<li>smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots</li>\n</ol>",
      "rawMarkdown": "1. wheel spokes fix \n 2. SUVs fix: roof and \"down sidebars\" \n 3. light car color preprocess \n 4.  smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots",
      "votes": null
    },
    {
      "id": "214264",
      "postDate": "08/16/2017 12:33:24",
      "content": "<p>1 is unreachable\neven 0.999 is unreachable.\nNow I'm near 0.998 and it seems to be limit, cause most of \"dice errors\" are due to ground truth errors</p>\n\n<p>Limiting factors:</p>\n\n<ol>\n<li>Darkness under car's bottom + most of cars have black bottom =&gt; sometimes its impossible to see border\n<ol><li>Ground truth errors due to human factors - I see lot of them nearly in every picture</li>\n<li>Low picture quality</li></ol></li>\n</ol>\n\n<p>That's why I dont believe in 0.999. 0.9985 seems to be adequate limit</p>",
      "rawMarkdown": "1 is unreachable\neven 0.999 is unreachable.\nNow I'm near 0.998 and it seems to be limit, cause most of \"dice errors\" are due to ground truth errors\n\nLimiting factors:\n\n 1. Darkness under car's bottom + most of cars have black bottom =&gt; sometimes its impossible to see border\n2. Ground truth errors due to human factors - I see lot of them nearly in every picture\n3. Low picture quality\n\nThat's why I dont believe in 0.999. 0.9985 seems to be adequate limit",
      "votes": null
    },
    {
      "id": "214272",
      "postDate": "08/16/2017 13:20:12",
      "content": "<blockquote>\n  <p>smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots</p>\n</blockquote>\n\n<p>Who did say that there is no holes inside the car body in private set? In this case you mask could be a perfect one as it gets - but gives an error due to private test set </p>",
      "rawMarkdown": "&gt; smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots\n\nWho did say that there is no holes inside the car body in private set? In this case you mask could be a perfect one as it gets - but gives an error due to private test set",
      "votes": null
    },
    {
      "id": "214364",
      "postDate": "08/16/2017 16:39:57",
      "content": "<p>Hi Heng!</p>\n\n<p>Why do you use zero-padded U-net?</p>\n\n<p>Even in their paper they used non-padded net</p>",
      "rawMarkdown": "Hi Heng!\n\nWhy do you use zero-padded U-net?\n\nEven in their paper they used non-padded net",
      "votes": null
    },
    {
      "id": "214443",
      "postDate": "08/16/2017 23:00:50",
      "content": "<p>I'm not in this competition but how about eroding prediction with 1 pixel using opencv or or something similar.  </p>",
      "rawMarkdown": "I'm not in this competition but how about eroding prediction with 1 pixel using opencv or or something similar.",
      "votes": null
    },
    {
      "id": "214538",
      "postDate": "08/17/2017 09:03:53",
      "content": "<p>In addition - wheel spokes</p>\n\n<p>On some \"ground truth\" images space between spokes marked as non_car, on some images - not. Suppose it depends on marker's judgment, and varies in test set as well</p>",
      "rawMarkdown": "In addition - wheel spokes\n\nOn some \"ground truth\" images space between spokes marked as non_car, on some images - not. Suppose it depends on marker's judgment, and varies in test set as well",
      "votes": null
    },
    {
      "id": "214683",
      "postDate": "08/17/2017 20:42:36",
      "content": "<p>The prediction isn't always larger.</p>",
      "rawMarkdown": "The prediction isn't always larger.",
      "votes": null
    },
    {
      "id": "214691",
      "postDate": "08/17/2017 21:42:35",
      "content": "<p>If the approach is useful one can also use dilation which is opposite of erode operation. How to choose the operation is a big question mark for me ... as I mentioned I'm not in this competition.</p>",
      "rawMarkdown": "If the approach is useful one can also use dilation which is opposite of erode operation. How to choose the operation is a big question mark for me ... as I mentioned I'm not in this competition.",
      "votes": null
    },
    {
      "id": "214806",
      "postDate": "08/18/2017 11:08:49",
      "content": "<p>I think a pixel error can not be eliminated, because of the reflection and shadow noise. But the counturing of the mirror on the side is coarse, at 0.998 we still make so big mistakes on the details?</p>",
      "rawMarkdown": "I think a pixel error can not be eliminated, because of the reflection and shadow noise. But the counturing of the mirror on the side is coarse, at 0.998 we still make so big mistakes on the details?",
      "votes": null
    },
    {
      "id": "214808",
      "postDate": "08/18/2017 11:21:40",
      "content": "<p>Isn't it possible to at least reduce the shadows and make car borders clearer with some light pre-processing? This is something I did quickly in PS, I'm sure something better (and automated) is possible. There are some publications dealing with shadow removal (eg. <a href=\"http://bit.ly/2uWNzZZ\">http://bit.ly/2uWNzZZ</a>)</p>\n\n<p><img src=\"https://www.dropbox.com/s/7bv8wrku3vi4pt8/0eeaf1ff136d_04.jpg?raw=1\" alt=\"original\" title=\"\">\n<img src=\"https://www.dropbox.com/s/accbb6e8vpavxog/0eeaf1ff136d_04_tweaked.jpg?raw=1\" alt=\"processed\" title=\"\"></p>",
      "rawMarkdown": "Isn't it possible to at least reduce the shadows and make car borders clearer with some light pre-processing? This is something I did quickly in PS, I'm sure something better (and automated) is possible. There are some publications dealing with shadow removal (eg. http://bit.ly/2uWNzZZ)\n\n![original][2]\n![processed][3]\n\n\n  [1]: http://www.inf.u-szeged.hu/projectdirs/ssip2011/teamF/ShadowRemoval/Documentation/Shadow%20detection%20and%20removal%20from%20a%20single%20image.pdf\n  [2]: https://www.dropbox.com/s/7bv8wrku3vi4pt8/0eeaf1ff136d_04.jpg?raw=1\n  [3]: https://www.dropbox.com/s/accbb6e8vpavxog/0eeaf1ff136d_04_tweaked.jpg?raw=1",
      "votes": null
    },
    {
      "id": "214817",
      "postDate": "08/18/2017 12:07:50",
      "content": "<p>thank you for the suggestion. I did something like local normalization:  x = (x=mean)/std. I concat this new channel to the original bgr input channels. There are some minor improvements  but i think the wrong segmentation problem at the bottom of cars is not only due to shadow, but also due to the human error in labeling, etc</p>\n\n<p>In tf, you can try the LRN layer too.</p>\n\n<pre><code>def make_norm_channel(x):\n     z = torch.unsqueeze(torch.sum(x,dim=1),1)\n     z2_mean = F.avg_pool2d(z*z,kernel_size=7,padding=3,stride=1)\n     z_mean  = F.avg_pool2d(z,  kernel_size=7,padding=3,stride=1)\n     z_std   = torch.sqrt(z2_mean-z_mean*z_mean+0.03)\n     z = (z-z_mean)/(z_std)\n     return z\n</code></pre>",
      "rawMarkdown": "thank you for the suggestion. I did something like local normalization:  x = (x=mean)/std. I concat this new channel to the original bgr input channels. There are some minor improvements  but i think the wrong segmentation problem at the bottom of cars is not only due to shadow, but also due to the human error in labeling, etc\n\nIn tf, you can try the LRN layer too.\n\n\n    def make_norm_channel(x):\n         z = torch.unsqueeze(torch.sum(x,dim=1),1)\n         z2_mean = F.avg_pool2d(z*z,kernel_size=7,padding=3,stride=1)\n         z_mean  = F.avg_pool2d(z,  kernel_size=7,padding=3,stride=1)\n         z_std   = torch.sqrt(z2_mean-z_mean*z_mean+0.03)\n         z = (z-z_mean)/(z_std)\n         return z",
      "votes": null
    },
    {
      "id": "214891",
      "postDate": "08/18/2017 17:31:00",
      "content": "<p>@Peter Giannakopoulos</p>\n\n<p>A fast way to do proof-of-concept is to use PS to process all test and train images first. Then get the validation results (and compare visual results for LB images). If it works, then think of way to replace PS function with python code.</p>",
      "rawMarkdown": "Peter Giannakopoulos\n\nA fast way to do proof-of-concept is to use PS to process all test and train images first. Then get the validation results (and compare visual results for LB images). If it works, then think of way to replace PS function with python code.",
      "votes": null
    },
    {
      "id": "214897",
      "postDate": "08/18/2017 17:47:22",
      "content": "<p>simple image processing method like the follows should improve results by 0.0001</p>\n\n<p>1) at a edge pixel, sample  inner pixels as foreground and outer pixels as background.</p>\n\n<p>2)for pixels next to the edge (e.g. 4-neighours), decide to erode or dilate based on color.</p>\n\n<p>3)if you are unsure (e.g is dark pixel shadow or car body), don't make changes.</p>\n\n<p>.</p>\n\n<p>there is a very big clue. the images of different view are <strong>stationary</strong>. at certain pixel locations, we know it is background or not by checking neighbouring images.</p>",
      "rawMarkdown": "simple image processing method like the follows should improve results by 0.0001\n\n1) at a edge pixel, sample  inner pixels as foreground and outer pixels as background.\n\n2)for pixels next to the edge (e.g. 4-neighours), decide to erode or dilate based on color.\n\n3)if you are unsure (e.g is dark pixel shadow or car body), don't make changes.\n\n.\n\nthere is a very big clue. the images of different view are **stationary**. at certain pixel locations, we know it is background or not by checking neighbouring images.",
      "votes": null
    },
    {
      "id": "216172",
      "postDate": "08/24/2017 16:22:42",
      "content": "<p>You also should to think about compression artifacts. All images are compressed with jpeg with middle quality factor, so all borders are distorted and it is hard to тел where actual mask borders is.\n<a href=\"https://yadi.sk/i/Z2CEYELj3MJJeN\">https://yadi.sk/i/Z2CEYELj3MJJeN</a></p>",
      "rawMarkdown": "You also should to think about compression artifacts. All images are compressed with jpeg with middle quality factor, so all borders are distorted and it is hard to тел where actual mask borders is.\n<img>https://yadi.sk/i/Z2CEYELj3MJJeN",
      "votes": null
    },
    {
      "id": "216735",
      "postDate": "08/27/2017 22:06:19",
      "content": "<p>For idea (2), I am wondering if you couldn't train a segmenter specifically for edges.  It could be given random high resolution patches of edges and be used to clean up the edges of the predicted segmentations.  </p>",
      "rawMarkdown": "For idea (2), I am wondering if you couldn't train a segmenter specifically for edges.  It could be given random high resolution patches of edges and be used to clean up the edges of the predicted segmentations.",
      "votes": null
    },
    {
      "id": "217619",
      "postDate": "08/31/2017 12:33:36",
      "content": "<p>Heng, in your summary you are talking about your local validation scores I guess? Are these your estimates or did you already reach these scores?</p>",
      "rawMarkdown": "Heng, in your summary you are talking about your local validation scores I guess? Are these your estimates or did you already reach these scores?",
      "votes": null
    },
    {
      "id": "217648",
      "postDate": "08/31/2017 13:48:52",
      "content": "<p>I tried something like that. Trained a CNN on 256x256 patches along the border. Didn't work for me, results were always worse then an standard unet trained on the whole image.</p>",
      "rawMarkdown": "I tried something like that. Trained a CNN on 256x256 patches along the border. Didn't work for me, results were always worse then an standard unet trained on the whole image.",
      "votes": null
    },
    {
      "id": "217658",
      "postDate": "08/31/2017 14:43:24",
      "content": "<p>these are estimations for LB score. i have not obtained these scores yet</p>",
      "rawMarkdown": "these are estimations for LB score. i have not obtained these scores yet",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 214250,
      "author_name": "zhihang",
      "author_url": "",
      "post_date": "08/16/2017 11:43:41",
      "content": "<p>wow。。from 0.99X to 1</p>",
      "votes": null,
      "replies": [
        {
          "id": 214264,
          "author_name": "gadgysaidoff",
          "author_url": "",
          "post_date": "08/16/2017 12:33:24",
          "content": "<p>1 is unreachable\neven 0.999 is unreachable.\nNow I'm near 0.998 and it seems to be limit, cause most of \"dice errors\" are due to ground truth errors</p>\n\n<p>Limiting factors:</p>\n\n<ol>\n<li>Darkness under car's bottom + most of cars have black bottom =&gt; sometimes its impossible to see border\n<ol><li>Ground truth errors due to human factors - I see lot of them nearly in every picture</li>\n<li>Low picture quality</li></ol></li>\n</ol>\n\n<p>That's why I dont believe in 0.999. 0.9985 seems to be adequate limit</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214538,
          "author_name": "gadgysaidoff",
          "author_url": "",
          "post_date": "08/17/2017 09:03:53",
          "content": "<p>In addition - wheel spokes</p>\n\n<p>On some \"ground truth\" images space between spokes marked as non_car, on some images - not. Suppose it depends on marker's judgment, and varies in test set as well</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214263,
      "author_name": "heyt0ny",
      "author_url": "",
      "post_date": "08/16/2017 12:28:06",
      "content": "<ol>\n<li>wheel spokes fix </li>\n<li>SUVs fix: roof and \"down sidebars\" </li>\n<li>light car color preprocess </li>\n<li>smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 214272,
          "author_name": "venheads",
          "author_url": "",
          "post_date": "08/16/2017 13:20:12",
          "content": "<blockquote>\n  <p>smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots</p>\n</blockquote>\n\n<p>Who did say that there is no holes inside the car body in private set? In this case you mask could be a perfect one as it gets - but gives an error due to private test set </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214364,
      "author_name": "gadgysaidoff",
      "author_url": "",
      "post_date": "08/16/2017 16:39:57",
      "content": "<p>Hi Heng!</p>\n\n<p>Why do you use zero-padded U-net?</p>\n\n<p>Even in their paper they used non-padded net</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 214443,
      "author_name": "sso2000",
      "author_url": "",
      "post_date": "08/16/2017 23:00:50",
      "content": "<p>I'm not in this competition but how about eroding prediction with 1 pixel using opencv or or something similar.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 214683,
          "author_name": "gopietz",
          "author_url": "",
          "post_date": "08/17/2017 20:42:36",
          "content": "<p>The prediction isn't always larger.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214691,
          "author_name": "sso2000",
          "author_url": "",
          "post_date": "08/17/2017 21:42:35",
          "content": "<p>If the approach is useful one can also use dilation which is opposite of erode operation. How to choose the operation is a big question mark for me ... as I mentioned I'm not in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214897,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/18/2017 17:47:22",
          "content": "<p>simple image processing method like the follows should improve results by 0.0001</p>\n\n<p>1) at a edge pixel, sample  inner pixels as foreground and outer pixels as background.</p>\n\n<p>2)for pixels next to the edge (e.g. 4-neighours), decide to erode or dilate based on color.</p>\n\n<p>3)if you are unsure (e.g is dark pixel shadow or car body), don't make changes.</p>\n\n<p>.</p>\n\n<p>there is a very big clue. the images of different view are <strong>stationary</strong>. at certain pixel locations, we know it is background or not by checking neighbouring images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214806,
      "author_name": "brasnold",
      "author_url": "",
      "post_date": "08/18/2017 11:08:49",
      "content": "<p>I think a pixel error can not be eliminated, because of the reflection and shadow noise. But the counturing of the mirror on the side is coarse, at 0.998 we still make so big mistakes on the details?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 214808,
      "author_name": "petrosgk",
      "author_url": "",
      "post_date": "08/18/2017 11:21:40",
      "content": "<p>Isn't it possible to at least reduce the shadows and make car borders clearer with some light pre-processing? This is something I did quickly in PS, I'm sure something better (and automated) is possible. There are some publications dealing with shadow removal (eg. <a href=\"http://bit.ly/2uWNzZZ\">http://bit.ly/2uWNzZZ</a>)</p>\n\n<p><img src=\"https://www.dropbox.com/s/7bv8wrku3vi4pt8/0eeaf1ff136d_04.jpg?raw=1\" alt=\"original\" title=\"\">\n<img src=\"https://www.dropbox.com/s/accbb6e8vpavxog/0eeaf1ff136d_04_tweaked.jpg?raw=1\" alt=\"processed\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 214817,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/18/2017 12:07:50",
          "content": "<p>thank you for the suggestion. I did something like local normalization:  x = (x=mean)/std. I concat this new channel to the original bgr input channels. There are some minor improvements  but i think the wrong segmentation problem at the bottom of cars is not only due to shadow, but also due to the human error in labeling, etc</p>\n\n<p>In tf, you can try the LRN layer too.</p>\n\n<pre><code>def make_norm_channel(x):\n     z = torch.unsqueeze(torch.sum(x,dim=1),1)\n     z2_mean = F.avg_pool2d(z*z,kernel_size=7,padding=3,stride=1)\n     z_mean  = F.avg_pool2d(z,  kernel_size=7,padding=3,stride=1)\n     z_std   = torch.sqrt(z2_mean-z_mean*z_mean+0.03)\n     z = (z-z_mean)/(z_std)\n     return z\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214891,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/18/2017 17:31:00",
          "content": "<p>@Peter Giannakopoulos</p>\n\n<p>A fast way to do proof-of-concept is to use PS to process all test and train images first. Then get the validation results (and compare visual results for LB images). If it works, then think of way to replace PS function with python code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 216172,
      "author_name": "demesgal",
      "author_url": "",
      "post_date": "08/24/2017 16:22:42",
      "content": "<p>You also should to think about compression artifacts. All images are compressed with jpeg with middle quality factor, so all borders are distorted and it is hard to тел where actual mask borders is.\n<a href=\"https://yadi.sk/i/Z2CEYELj3MJJeN\">https://yadi.sk/i/Z2CEYELj3MJJeN</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 216735,
      "author_name": "",
      "author_url": "",
      "post_date": "08/27/2017 22:06:19",
      "content": "<p>For idea (2), I am wondering if you couldn't train a segmenter specifically for edges.  It could be given random high resolution patches of edges and be used to clean up the edges of the predicted segmentations.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 217648,
          "author_name": "friedeks",
          "author_url": "",
          "post_date": "08/31/2017 13:48:52",
          "content": "<p>I tried something like that. Trained a CNN on 256x256 patches along the border. Didn't work for me, results were always worse then an standard unet trained on the whole image.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 217619,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "08/31/2017 12:33:36",
      "content": "<p>Heng, in your summary you are talking about your local validation scores I guess? Are these your estimates or did you already reach these scores?</p>",
      "votes": null,
      "replies": [
        {
          "id": 217658,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/31/2017 14:43:24",
          "content": "<p>these are estimations for LB score. i have not obtained these scores yet</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "214243": "Let's discuss here! Here is how a 0.998 error would look like. It is about 1-pixel boundary error.\n\n ![enter image description here][1]\n\n\nHere are my ideas:\n\n(1) will changing network structure (e.g. resnet, densenet for decoder/encoder) help? \n\nnot useful unless you can increase the number of channels when you scale up. Actually, in my experiments, you still get improved results if you use more channels, even in simple UNet. But it becomes so slow and it is difficult to extend this low image resolution.\n\n.\n\n\n(2) How about increase resolution for input?\n\nThere will be improvement. However, larger resolution also means deeper network. Difficult to scale up. One crazy idea is to train up 2 x actual size (i.e. larger than input resolution) \n\n.\n\n(3) weak supervised learning.\n\nWe actually already have pretty good results on the test images  (0.997x) human error is about (0.999?).  CVPR/ICCV has some works to show that partial label can achieve good results. e.g. \"Simple does it: Weakly Supervised Instance and Semantic Segmentation\" -CVPR 2017.\n\nIf you are considering \"512x512\", you have 0.998x accurate pseudo-label on the test set.\n\n\n.\n\n.\n\n.\n\n\nIn summary, this is my best bet:\n\n- my current  best LB results 0.997, more precisely 0.99662(https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38301. i need to revised the targets below later)\n\n- best results attain by deep CNN (better image resolution, more parameters, more data, etc) 0.9972\n\n- smart but simple post image/meta-data processing 0.9973 (Note you surely needs some method to complement deep CNN)\n\n- ensemble to stabilize and make robust results? 0.99735\n \n\n\n\n... to be updated ...\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/214243/7133/0.998.png",
    "214250": "wow。。from 0.99X to 1",
    "214263": "1. wheel spokes fix \n 2. SUVs fix: roof and \"down sidebars\" \n 3. light car color preprocess \n 4.  smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots",
    "214264": "1 is unreachable\neven 0.999 is unreachable.\nNow I'm near 0.998 and it seems to be limit, cause most of \"dice errors\" are due to ground truth errors\n\nLimiting factors:\n\n 1. Darkness under car's bottom + most of cars have black bottom =&gt; sometimes its impossible to see border\n2. Ground truth errors due to human factors - I see lot of them nearly in every picture\n3. Low picture quality\n\nThat's why I dont believe in 0.999. 0.9985 seems to be adequate limit",
    "214272": "&gt; smart postprocessing to fix dumb errors like holes inside the car body, unsymmetrical sides, redundant spots\n\nWho did say that there is no holes inside the car body in private set? In this case you mask could be a perfect one as it gets - but gives an error due to private test set",
    "214364": "Hi Heng!\n\nWhy do you use zero-padded U-net?\n\nEven in their paper they used non-padded net",
    "214443": "I'm not in this competition but how about eroding prediction with 1 pixel using opencv or or something similar.",
    "214538": "In addition - wheel spokes\n\nOn some \"ground truth\" images space between spokes marked as non_car, on some images - not. Suppose it depends on marker's judgment, and varies in test set as well",
    "214683": "The prediction isn't always larger.",
    "214691": "If the approach is useful one can also use dilation which is opposite of erode operation. How to choose the operation is a big question mark for me ... as I mentioned I'm not in this competition.",
    "214806": "I think a pixel error can not be eliminated, because of the reflection and shadow noise. But the counturing of the mirror on the side is coarse, at 0.998 we still make so big mistakes on the details?",
    "214808": "Isn't it possible to at least reduce the shadows and make car borders clearer with some light pre-processing? This is something I did quickly in PS, I'm sure something better (and automated) is possible. There are some publications dealing with shadow removal (eg. http://bit.ly/2uWNzZZ)\n\n![original][2]\n![processed][3]\n\n\n  [1]: http://www.inf.u-szeged.hu/projectdirs/ssip2011/teamF/ShadowRemoval/Documentation/Shadow%20detection%20and%20removal%20from%20a%20single%20image.pdf\n  [2]: https://www.dropbox.com/s/7bv8wrku3vi4pt8/0eeaf1ff136d_04.jpg?raw=1\n  [3]: https://www.dropbox.com/s/accbb6e8vpavxog/0eeaf1ff136d_04_tweaked.jpg?raw=1",
    "214817": "thank you for the suggestion. I did something like local normalization:  x = (x=mean)/std. I concat this new channel to the original bgr input channels. There are some minor improvements  but i think the wrong segmentation problem at the bottom of cars is not only due to shadow, but also due to the human error in labeling, etc\n\nIn tf, you can try the LRN layer too.\n\n\n    def make_norm_channel(x):\n         z = torch.unsqueeze(torch.sum(x,dim=1),1)\n         z2_mean = F.avg_pool2d(z*z,kernel_size=7,padding=3,stride=1)\n         z_mean  = F.avg_pool2d(z,  kernel_size=7,padding=3,stride=1)\n         z_std   = torch.sqrt(z2_mean-z_mean*z_mean+0.03)\n         z = (z-z_mean)/(z_std)\n         return z",
    "214891": "Peter Giannakopoulos\n\nA fast way to do proof-of-concept is to use PS to process all test and train images first. Then get the validation results (and compare visual results for LB images). If it works, then think of way to replace PS function with python code.",
    "214897": "simple image processing method like the follows should improve results by 0.0001\n\n1) at a edge pixel, sample  inner pixels as foreground and outer pixels as background.\n\n2)for pixels next to the edge (e.g. 4-neighours), decide to erode or dilate based on color.\n\n3)if you are unsure (e.g is dark pixel shadow or car body), don't make changes.\n\n.\n\nthere is a very big clue. the images of different view are **stationary**. at certain pixel locations, we know it is background or not by checking neighbouring images.",
    "216172": "You also should to think about compression artifacts. All images are compressed with jpeg with middle quality factor, so all borders are distorted and it is hard to тел where actual mask borders is.\n<img>https://yadi.sk/i/Z2CEYELj3MJJeN",
    "216735": "For idea (2), I am wondering if you couldn't train a segmenter specifically for edges.  It could be given random high resolution patches of edges and be used to clean up the edges of the predicted segmentations.",
    "217619": "Heng, in your summary you are talking about your local validation scores I guess? Are these your estimates or did you already reach these scores?",
    "217648": "I tried something like that. Trained a CNN on 256x256 patches along the border. Didn't work for me, results were always worse then an standard unet trained on the whole image.",
    "217658": "these are estimations for LB score. i have not obtained these scores yet"
  },
  "source": "meta"
}