{
  "id": 40126,
  "title": "LB 11th(0.9971) solution overview ",
  "url": "/competitions/carvana-image-masking-challenge/writeups/sukjae-cho-lb-11th-0-9971-solution-overview",
  "author_name": "",
  "post_date": "2017-09-28T02:33:30.227471Z",
  "votes": 33,
  "comment_count": 13,
  "views": 0,
  "content": "<p>My solution is in two stages.</p>\n\n<p><strong>Stage 1:  Predict rough outline of mask with 1/4 downscaled image</strong></p>\n\n<ol>\n<li>Scale down all images/masks to 320x480</li>\n<li>Train with default UNET.</li>\n<li>Predict mask for test set</li>\n</ol>\n\n<p><strong>Stage 2: Train/Predict mask only using images around edge areas.</strong></p>\n\n<p>This reduces the amount of data to process a lot, and I can use more complex network in reasonable time.</p>\n\n<p>Details are:</p>\n\n<ul>\n<li><p>Extract 256x256 tiles only along the edge of masks from stage 1 </p></li>\n<li><p>Used deeper UNET - 6 downsample/upsample. But for each depth, Conv\nlayers    with small number of layers than original UNET to avoid\nexplosion of\n   parameter numbers. Around 3million parameters are used.</p></li>\n<li><p>During training, tiles are overlapped so that each edge pixel    are\ncovered twice.</p></li>\n<li><p>During inferencing, tiles are overlapped so that       each edge\npixel are covered at least three times. So each pixel's    prediction\nis voted at least by 3 predictions(see attached image) UNET tends to    have    bad\nprediction around edge area, but this method reduced this problem.</p></li>\n<li><p>Used standard augmentation(scale+-10%, flip left-right, brightness, contrast, pixel shift, etc) and basic dice_coeff as loss.</p></li>\n<li>Used batch normalization, and to maximize the effectiveness of it, <br>\nthe batch consists of random tiles from multiple images. I guess this\nhelped network to converge much faster.</li>\n</ul>\n\n<p>I tried several other segmentation networks, but this one showed best result with large margin(0.0002) compared to other models. Single model LB was 0.9970 and with ensemble of several folds from same model, it got to 0.9971. </p>\n\n<p>I think this pipeline works really well considering its simplicity, but I made a big mistake.\n - I didn't noticed that some test outputs from stage 1 were horribly wrong until today and didn't have chance to correct it.  :(\nI should have found the issue early only if I had good visualization or simple sanity check logic.</p>\n\n<p>The best lesson from this competition is that good visualization/analyzing tool is really important.</p>\n\n<p>Oh, Heng CherKeng, thank you very much. I learned a lot from your posting!</p>",
  "messages": [
    {
      "id": "224981",
      "postDate": "09/28/2017 02:33:30",
      "content": "<p>My solution is in two stages.</p>\n\n<p><strong>Stage 1:  Predict rough outline of mask with 1/4 downscaled image</strong></p>\n\n<ol>\n<li>Scale down all images/masks to 320x480</li>\n<li>Train with default UNET.</li>\n<li>Predict mask for test set</li>\n</ol>\n\n<p><strong>Stage 2: Train/Predict mask only using images around edge areas.</strong></p>\n\n<p>This reduces the amount of data to process a lot, and I can use more complex network in reasonable time.</p>\n\n<p>Details are:</p>\n\n<ul>\n<li><p>Extract 256x256 tiles only along the edge of masks from stage 1 </p></li>\n<li><p>Used deeper UNET - 6 downsample/upsample. But for each depth, Conv\nlayers    with small number of layers than original UNET to avoid\nexplosion of\n   parameter numbers. Around 3million parameters are used.</p></li>\n<li><p>During training, tiles are overlapped so that each edge pixel    are\ncovered twice.</p></li>\n<li><p>During inferencing, tiles are overlapped so that       each edge\npixel are covered at least three times. So each pixel's    prediction\nis voted at least by 3 predictions(see attached image) UNET tends to    have    bad\nprediction around edge area, but this method reduced this problem.</p></li>\n<li><p>Used standard augmentation(scale+-10%, flip left-right, brightness, contrast, pixel shift, etc) and basic dice_coeff as loss.</p></li>\n<li>Used batch normalization, and to maximize the effectiveness of it, <br>\nthe batch consists of random tiles from multiple images. I guess this\nhelped network to converge much faster.</li>\n</ul>\n\n<p>I tried several other segmentation networks, but this one showed best result with large margin(0.0002) compared to other models. Single model LB was 0.9970 and with ensemble of several folds from same model, it got to 0.9971. </p>\n\n<p>I think this pipeline works really well considering its simplicity, but I made a big mistake.\n - I didn't noticed that some test outputs from stage 1 were horribly wrong until today and didn't have chance to correct it.  :(\nI should have found the issue early only if I had good visualization or simple sanity check logic.</p>\n\n<p>The best lesson from this competition is that good visualization/analyzing tool is really important.</p>\n\n<p>Oh, Heng CherKeng, thank you very much. I learned a lot from your posting!</p>",
      "rawMarkdown": "My solution is in two stages.\n\n**Stage 1:  Predict rough outline of mask with 1/4 downscaled image**\n\n 1. Scale down all images/masks to 320x480\n 2. Train with default UNET.\n 3. Predict mask for test set\n\n**Stage 2: Train/Predict mask only using images around edge areas.**\n\nThis reduces the amount of data to process a lot, and I can use more complex network in reasonable time.\n\nDetails are:\n\n - Extract 256x256 tiles only along the edge of masks from stage 1 \n\n - Used deeper UNET - 6 downsample/upsample. But for each depth, Conv\n   layers    with small number of layers than original UNET to avoid\n   explosion of\n       parameter numbers. Around 3million parameters are used.\n\n - During training, tiles are overlapped so that each edge pixel    are\n   covered twice.\n\n - During inferencing, tiles are overlapped so that       each edge\n   pixel are covered at least three times. So each pixel's    prediction\n   is voted at least by 3 predictions(see attached image) UNET tends to    have    bad\n   prediction around edge area, but this method reduced this problem.\n\n - Used standard augmentation(scale+-10%, flip left-right, brightness, contrast, pixel shift, etc) and basic dice_coeff as loss.\n - Used batch normalization, and to maximize the effectiveness of it,   \n   the batch consists of random tiles from multiple images. I guess this\n   helped network to converge much faster.\n\nI tried several other segmentation networks, but this one showed best result with large margin(0.0002) compared to other models. Single model LB was 0.9970 and with ensemble of several folds from same model, it got to 0.9971. \n\nI think this pipeline works really well considering its simplicity, but I made a big mistake.\n - I didn't noticed that some test outputs from stage 1 were horribly wrong until today and didn't have chance to correct it.  :(\nI should have found the issue early only if I had good visualization or simple sanity check logic.\n\n\nThe best lesson from this competition is that good visualization/analyzing tool is really important.\n\n\nOh, Heng CherKeng, thank you very much. I learned a lot from your posting!",
      "votes": null
    },
    {
      "id": "224983",
      "postDate": "09/28/2017 02:37:41",
      "content": "<p>Neat solution! Thanks!</p>",
      "rawMarkdown": "Neat solution! Thanks!",
      "votes": null
    },
    {
      "id": "224991",
      "postDate": "09/28/2017 03:01:23",
      "content": "<p>Amazing! This was our main idea, but I was worrying that uNet256 doesn't give a proper dice score on the edge. We even trained the edge detector using uNet, but stopped right there and decided to move another way. Nice to see you implemented this!</p>",
      "rawMarkdown": "Amazing! This was our main idea, but I was worrying that uNet256 doesn't give a proper dice score on the edge. We even trained the edge detector using uNet, but stopped right there and decided to move another way. Nice to see you implemented this!",
      "votes": null
    },
    {
      "id": "225002",
      "postDate": "09/28/2017 03:22:16",
      "content": "<p>This is cool !! <br>\nI have some questions~ <br>\n1. Did you train with images/masks of 320x480 for your best result ? <br>\n2. Do you think train higher resolution of images (e.g 1024x1024) will increase the score? <br>\n3. Do you think extracting higher resolution tiles along the edge (e.g. 512x512) will help? <br>\nAnd the most important, Can you share your code, please, hahah?  </p>",
      "rawMarkdown": "This is cool !!  \nI have some questions~  \n1. Did you train with images/masks of 320x480 for your best result ?  \n2. Do you think train higher resolution of images (e.g 1024x1024) will increase the score?  \n3. Do you think extracting higher resolution tiles along the edge (e.g. 512x512) will help?  \nAnd the most important, Can you share your code, please, hahah?",
      "votes": null
    },
    {
      "id": "225009",
      "postDate": "09/28/2017 03:44:27",
      "content": "<ol>\n<li><p>Stage 1's output mask is used for generating stage2's test tile and fill areas that are not covered by tiles.\nI thought the quality of stage 1 output doesn't matter much since details are corrected anyway in second stage anyway. But in reality, there were a few images that have big errors.</p></li>\n<li><p>I think higher resolution in stage 1 will increase score just little bit. With 320x480 image, details like antenna is lost, but higher resolution will help</p></li>\n<li><p>I tried several other tile sizes. I saw heavy over-fitting with smaller tile size than 256. With 64 tile size, it works extremely well in local validation(even reached 0.9981 - well partially caused by the way that I split train/val), but terrible in public leader board. I tried tile size of 384 also, but it took too long to converge so I gave up. With 256x256 tile, it got okay model in a day. </p></li>\n</ol>\n\n<p>For code, I guess I need to clean it up first (for myself also) and it can take some time, but I'll try to share it.</p>",
      "rawMarkdown": "1. Stage 1's output mask is used for generating stage2's test tile and fill areas that are not covered by tiles.\nI thought the quality of stage 1 output doesn't matter much since details are corrected anyway in second stage anyway. But in reality, there were a few images that have big errors.\n\n2. I think higher resolution in stage 1 will increase score just little bit. With 320x480 image, details like antenna is lost, but higher resolution will help\n\n3. I tried several other tile sizes. I saw heavy over-fitting with smaller tile size than 256. With 64 tile size, it works extremely well in local validation(even reached 0.9981 - well partially caused by the way that I split train/val), but terrible in public leader board. I tried tile size of 384 also, but it took too long to converge so I gave up. With 256x256 tile, it got okay model in a day. \n\nFor code, I guess I need to clean it up first (for myself also) and it can take some time, but I'll try to share it.",
      "votes": null
    },
    {
      "id": "225020",
      "postDate": "09/28/2017 03:58:56",
      "content": "<p>Thanks for your sharing! <br>\nI look forward to learn more from your code~</p>",
      "rawMarkdown": "Thanks for your sharing!  \nI look forward to learn more from your code~",
      "votes": null
    },
    {
      "id": "225131",
      "postDate": "09/28/2017 10:30:33",
      "content": "<p>Hey! We thought of the same... and also implemented it a patches-based one.</p>",
      "rawMarkdown": "Hey! We thought of the same... and also implemented it a patches-based one.",
      "votes": null
    },
    {
      "id": "225144",
      "postDate": "09/28/2017 11:17:21",
      "content": "<p>Hi JandJ\nMay I ask how to you extract tile along the edge of the mask?</p>\n\n<p>I image you have to slide the window over a image and capture ones that have roughly 50-50 share of true false pixels, for all the images.</p>\n\n<p>how long did this process take?</p>\n\n<p>many thanks!</p>",
      "rawMarkdown": "Hi JandJ\nMay I ask how to you extract tile along the edge of the mask?\n\nI image you have to slide the window over a image and capture ones that have roughly 50-50 share of true false pixels, for all the images.\n\nhow long did this process take?\n\nmany thanks!",
      "votes": null
    },
    {
      "id": "225272",
      "postDate": "09/28/2017 16:18:08",
      "content": "<p>Basically, generate image with edge, and eliminate pixels while getting tile location.</p>\n\n<p>For training set, it took 100 sec with 8 cores with overlap_ratio of 2.\nIn retrospect, I should have added some more code for checking sanity and correct big holes far inside the car edge, etc. </p>\n\n<pre><code>mask = mask &gt;= 127    \nedge = mask - morphology.binary_erosion(mask)\nedge *= overlap_ratio\nhs = tile_size//2\nqs = tile_size//4\nwhile (True):\n        nzs = np.transpose(np.nonzero(edge))\n        if len(nzs) == 0:\n            break\n        # Pick random pixels among non-zero values.\n        for i in range(len(nzs)//10):\n            nz = nzs[random.randint(0, len(nzs)-1)]\n            if edge[nz[0], nz[1]] &lt;= 0:\n                continue\n            # Make it fit inside valid area\n            nz[0] = np.clip(nz[0], hs, 1280 - hs)\n            nz[1] = np.clip(nz[1], hs, 1918 - hs)\n\n            tile_rcs.append((nz[0] - hs, nz[1] - hs))  # Record left-top coordinates\n\n            edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] -= edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] &gt; 0\n            # Mark center area as 'done' part to avoid sampling again from similar location\n            edge[nz[0] - qs:nz[0] + qs, nz[1] - qs:nz[1] + qs] = 0\n            if nz[1] == hs or nz[1] == 1918 - hs:  # Avoid oversampling on Left,right edge cases\n                edge[nz[0] - qs // 2:nz[0] + qs // 2, nz[1] - hs:nz[1] + hs] = 0\n</code></pre>",
      "rawMarkdown": "Basically, generate image with edge, and eliminate pixels while getting tile location.\n\nFor training set, it took 100 sec with 8 cores with overlap_ratio of 2.\nIn retrospect, I should have added some more code for checking sanity and correct big holes far inside the car edge, etc. \n\n    mask = mask &gt;= 127    \n    edge = mask - morphology.binary_erosion(mask)\n    edge *= overlap_ratio\n    hs = tile_size//2\n    qs = tile_size//4\n    while (True):\n            nzs = np.transpose(np.nonzero(edge))\n            if len(nzs) == 0:\n                break\n            # Pick random pixels among non-zero values.\n            for i in range(len(nzs)//10):\n                nz = nzs[random.randint(0, len(nzs)-1)]\n                if edge[nz[0], nz[1]] &lt;= 0:\n                    continue\n                # Make it fit inside valid area\n                nz[0] = np.clip(nz[0], hs, 1280 - hs)\n                nz[1] = np.clip(nz[1], hs, 1918 - hs)\n    \n                tile_rcs.append((nz[0] - hs, nz[1] - hs))  # Record left-top coordinates\n    \n                edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] -= edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] &gt; 0\n                # Mark center area as 'done' part to avoid sampling again from similar location\n                edge[nz[0] - qs:nz[0] + qs, nz[1] - qs:nz[1] + qs] = 0\n                if nz[1] == hs or nz[1] == 1918 - hs:  # Avoid oversampling on Left,right edge cases\n                    edge[nz[0] - qs // 2:nz[0] + qs // 2, nz[1] - hs:nz[1] + hs] = 0",
      "votes": null
    },
    {
      "id": "225354",
      "postDate": "09/28/2017 20:22:32",
      "content": "<p>Thanks @JandJ for sharing your solution and congratulations.</p>",
      "rawMarkdown": "Thanks @JandJ for sharing your solution and congratulations.",
      "votes": null
    },
    {
      "id": "225433",
      "postDate": "09/29/2017 03:50:20",
      "content": "<p>Thanks for sharing.  that's a great idea.</p>",
      "rawMarkdown": "Thanks for sharing.  that's a great idea.",
      "votes": null
    },
    {
      "id": "225517",
      "postDate": "09/29/2017 10:48:02",
      "content": "<p>Hi JandJ</p>\n\n<p>Thanks for sharing!</p>",
      "rawMarkdown": "Hi JandJ\n\nThanks for sharing!",
      "votes": null
    },
    {
      "id": "225759",
      "postDate": "09/30/2017 00:19:45",
      "content": "<p>Thanks for sharing!! It is very helpful.</p>",
      "rawMarkdown": "Thanks for sharing!! It is very helpful.",
      "votes": null
    },
    {
      "id": "225773",
      "postDate": "09/30/2017 01:46:13",
      "content": "<p>It is a very method with excellent tesults. You can consider a end to end solution with single network like faster rcnn. I think the features for stage one and two can be shared. You can write a new paper😀. I am think of focus loss for selecting the crops for stage two ...</p>\n\n<p>Thanks for sharing the work and it is a good job!</p>",
      "rawMarkdown": "It is a very method with excellent tesults. You can consider a end to end solution with single network like faster rcnn. I think the features for stage one and two can be shared. You can write a new paper😀. I am think of focus loss for selecting the crops for stage two ...\n\nThanks for sharing the work and it is a good job!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 224983,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "09/28/2017 02:37:41",
      "content": "<p>Neat solution! Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224991,
      "author_name": "truepk",
      "author_url": "",
      "post_date": "09/28/2017 03:01:23",
      "content": "<p>Amazing! This was our main idea, but I was worrying that uNet256 doesn't give a proper dice score on the edge. We even trained the edge detector using uNet, but stopped right there and decided to move another way. Nice to see you implemented this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225002,
      "author_name": "miaouhoho",
      "author_url": "",
      "post_date": "09/28/2017 03:22:16",
      "content": "<p>This is cool !! <br>\nI have some questions~ <br>\n1. Did you train with images/masks of 320x480 for your best result ? <br>\n2. Do you think train higher resolution of images (e.g 1024x1024) will increase the score? <br>\n3. Do you think extracting higher resolution tiles along the edge (e.g. 512x512) will help? <br>\nAnd the most important, Can you share your code, please, hahah?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 225009,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "09/28/2017 03:44:27",
          "content": "<ol>\n<li><p>Stage 1's output mask is used for generating stage2's test tile and fill areas that are not covered by tiles.\nI thought the quality of stage 1 output doesn't matter much since details are corrected anyway in second stage anyway. But in reality, there were a few images that have big errors.</p></li>\n<li><p>I think higher resolution in stage 1 will increase score just little bit. With 320x480 image, details like antenna is lost, but higher resolution will help</p></li>\n<li><p>I tried several other tile sizes. I saw heavy over-fitting with smaller tile size than 256. With 64 tile size, it works extremely well in local validation(even reached 0.9981 - well partially caused by the way that I split train/val), but terrible in public leader board. I tried tile size of 384 also, but it took too long to converge so I gave up. With 256x256 tile, it got okay model in a day. </p></li>\n</ol>\n\n<p>For code, I guess I need to clean it up first (for myself also) and it can take some time, but I'll try to share it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225020,
          "author_name": "miaouhoho",
          "author_url": "",
          "post_date": "09/28/2017 03:58:56",
          "content": "<p>Thanks for your sharing! <br>\nI look forward to learn more from your code~</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225131,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "09/28/2017 10:30:33",
      "content": "<p>Hey! We thought of the same... and also implemented it a patches-based one.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225144,
      "author_name": "yl1202",
      "author_url": "",
      "post_date": "09/28/2017 11:17:21",
      "content": "<p>Hi JandJ\nMay I ask how to you extract tile along the edge of the mask?</p>\n\n<p>I image you have to slide the window over a image and capture ones that have roughly 50-50 share of true false pixels, for all the images.</p>\n\n<p>how long did this process take?</p>\n\n<p>many thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 225272,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "09/28/2017 16:18:08",
          "content": "<p>Basically, generate image with edge, and eliminate pixels while getting tile location.</p>\n\n<p>For training set, it took 100 sec with 8 cores with overlap_ratio of 2.\nIn retrospect, I should have added some more code for checking sanity and correct big holes far inside the car edge, etc. </p>\n\n<pre><code>mask = mask &gt;= 127    \nedge = mask - morphology.binary_erosion(mask)\nedge *= overlap_ratio\nhs = tile_size//2\nqs = tile_size//4\nwhile (True):\n        nzs = np.transpose(np.nonzero(edge))\n        if len(nzs) == 0:\n            break\n        # Pick random pixels among non-zero values.\n        for i in range(len(nzs)//10):\n            nz = nzs[random.randint(0, len(nzs)-1)]\n            if edge[nz[0], nz[1]] &lt;= 0:\n                continue\n            # Make it fit inside valid area\n            nz[0] = np.clip(nz[0], hs, 1280 - hs)\n            nz[1] = np.clip(nz[1], hs, 1918 - hs)\n\n            tile_rcs.append((nz[0] - hs, nz[1] - hs))  # Record left-top coordinates\n\n            edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] -= edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] &gt; 0\n            # Mark center area as 'done' part to avoid sampling again from similar location\n            edge[nz[0] - qs:nz[0] + qs, nz[1] - qs:nz[1] + qs] = 0\n            if nz[1] == hs or nz[1] == 1918 - hs:  # Avoid oversampling on Left,right edge cases\n                edge[nz[0] - qs // 2:nz[0] + qs // 2, nz[1] - hs:nz[1] + hs] = 0\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225517,
          "author_name": "yl1202",
          "author_url": "",
          "post_date": "09/29/2017 10:48:02",
          "content": "<p>Hi JandJ</p>\n\n<p>Thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225354,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "09/28/2017 20:22:32",
      "content": "<p>Thanks @JandJ for sharing your solution and congratulations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225433,
      "author_name": "jackkwok",
      "author_url": "",
      "post_date": "09/29/2017 03:50:20",
      "content": "<p>Thanks for sharing.  that's a great idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225759,
      "author_name": "utkarshsingh",
      "author_url": "",
      "post_date": "09/30/2017 00:19:45",
      "content": "<p>Thanks for sharing!! It is very helpful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 225773,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/30/2017 01:46:13",
          "content": "<p>It is a very method with excellent tesults. You can consider a end to end solution with single network like faster rcnn. I think the features for stage one and two can be shared. You can write a new paper😀. I am think of focus loss for selecting the crops for stage two ...</p>\n\n<p>Thanks for sharing the work and it is a good job!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "224981": "My solution is in two stages.\n\n**Stage 1:  Predict rough outline of mask with 1/4 downscaled image**\n\n 1. Scale down all images/masks to 320x480\n 2. Train with default UNET.\n 3. Predict mask for test set\n\n**Stage 2: Train/Predict mask only using images around edge areas.**\n\nThis reduces the amount of data to process a lot, and I can use more complex network in reasonable time.\n\nDetails are:\n\n - Extract 256x256 tiles only along the edge of masks from stage 1 \n\n - Used deeper UNET - 6 downsample/upsample. But for each depth, Conv\n   layers    with small number of layers than original UNET to avoid\n   explosion of\n       parameter numbers. Around 3million parameters are used.\n\n - During training, tiles are overlapped so that each edge pixel    are\n   covered twice.\n\n - During inferencing, tiles are overlapped so that       each edge\n   pixel are covered at least three times. So each pixel's    prediction\n   is voted at least by 3 predictions(see attached image) UNET tends to    have    bad\n   prediction around edge area, but this method reduced this problem.\n\n - Used standard augmentation(scale+-10%, flip left-right, brightness, contrast, pixel shift, etc) and basic dice_coeff as loss.\n - Used batch normalization, and to maximize the effectiveness of it,   \n   the batch consists of random tiles from multiple images. I guess this\n   helped network to converge much faster.\n\nI tried several other segmentation networks, but this one showed best result with large margin(0.0002) compared to other models. Single model LB was 0.9970 and with ensemble of several folds from same model, it got to 0.9971. \n\nI think this pipeline works really well considering its simplicity, but I made a big mistake.\n - I didn't noticed that some test outputs from stage 1 were horribly wrong until today and didn't have chance to correct it.  :(\nI should have found the issue early only if I had good visualization or simple sanity check logic.\n\n\nThe best lesson from this competition is that good visualization/analyzing tool is really important.\n\n\nOh, Heng CherKeng, thank you very much. I learned a lot from your posting!",
    "224983": "Neat solution! Thanks!",
    "224991": "Amazing! This was our main idea, but I was worrying that uNet256 doesn't give a proper dice score on the edge. We even trained the edge detector using uNet, but stopped right there and decided to move another way. Nice to see you implemented this!",
    "225002": "This is cool !!  \nI have some questions~  \n1. Did you train with images/masks of 320x480 for your best result ?  \n2. Do you think train higher resolution of images (e.g 1024x1024) will increase the score?  \n3. Do you think extracting higher resolution tiles along the edge (e.g. 512x512) will help?  \nAnd the most important, Can you share your code, please, hahah?",
    "225009": "1. Stage 1's output mask is used for generating stage2's test tile and fill areas that are not covered by tiles.\nI thought the quality of stage 1 output doesn't matter much since details are corrected anyway in second stage anyway. But in reality, there were a few images that have big errors.\n\n2. I think higher resolution in stage 1 will increase score just little bit. With 320x480 image, details like antenna is lost, but higher resolution will help\n\n3. I tried several other tile sizes. I saw heavy over-fitting with smaller tile size than 256. With 64 tile size, it works extremely well in local validation(even reached 0.9981 - well partially caused by the way that I split train/val), but terrible in public leader board. I tried tile size of 384 also, but it took too long to converge so I gave up. With 256x256 tile, it got okay model in a day. \n\nFor code, I guess I need to clean it up first (for myself also) and it can take some time, but I'll try to share it.",
    "225020": "Thanks for your sharing!  \nI look forward to learn more from your code~",
    "225131": "Hey! We thought of the same... and also implemented it a patches-based one.",
    "225144": "Hi JandJ\nMay I ask how to you extract tile along the edge of the mask?\n\nI image you have to slide the window over a image and capture ones that have roughly 50-50 share of true false pixels, for all the images.\n\nhow long did this process take?\n\nmany thanks!",
    "225272": "Basically, generate image with edge, and eliminate pixels while getting tile location.\n\nFor training set, it took 100 sec with 8 cores with overlap_ratio of 2.\nIn retrospect, I should have added some more code for checking sanity and correct big holes far inside the car edge, etc. \n\n    mask = mask &gt;= 127    \n    edge = mask - morphology.binary_erosion(mask)\n    edge *= overlap_ratio\n    hs = tile_size//2\n    qs = tile_size//4\n    while (True):\n            nzs = np.transpose(np.nonzero(edge))\n            if len(nzs) == 0:\n                break\n            # Pick random pixels among non-zero values.\n            for i in range(len(nzs)//10):\n                nz = nzs[random.randint(0, len(nzs)-1)]\n                if edge[nz[0], nz[1]] &lt;= 0:\n                    continue\n                # Make it fit inside valid area\n                nz[0] = np.clip(nz[0], hs, 1280 - hs)\n                nz[1] = np.clip(nz[1], hs, 1918 - hs)\n    \n                tile_rcs.append((nz[0] - hs, nz[1] - hs))  # Record left-top coordinates\n    \n                edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] -= edge[nz[0] - hs:nz[0] + hs, nz[1] - hs:nz[1] + hs] &gt; 0\n                # Mark center area as 'done' part to avoid sampling again from similar location\n                edge[nz[0] - qs:nz[0] + qs, nz[1] - qs:nz[1] + qs] = 0\n                if nz[1] == hs or nz[1] == 1918 - hs:  # Avoid oversampling on Left,right edge cases\n                    edge[nz[0] - qs // 2:nz[0] + qs // 2, nz[1] - hs:nz[1] + hs] = 0",
    "225354": "Thanks @JandJ for sharing your solution and congratulations.",
    "225433": "Thanks for sharing.  that's a great idea.",
    "225517": "Hi JandJ\n\nThanks for sharing!",
    "225759": "Thanks for sharing!! It is very helpful.",
    "225773": "It is a very method with excellent tesults. You can consider a end to end solution with single network like faster rcnn. I think the features for stage one and two can be shared. You can write a new paper😀. I am think of focus loss for selecting the crops for stage two ...\n\nThanks for sharing the work and it is a good job!"
  },
  "source": "meta"
}