{
  "id": 40146,
  "title": "LB 15th (0.9970) Some new ideas (and a bug!)",
  "url": "/competitions/carvana-image-masking-challenge/writeups/team-73-lb-15th-0-9970-some-new-ideas-and-a-bug",
  "author_name": "",
  "post_date": "2017-09-28T10:25:08.075474200Z",
  "votes": 25,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Our solution has two stages that can be run recursively (but we only ran then once b/c of lack of time) =&gt;</p>\n\n<p><strong>1. Predict a coarse mask (CO)</strong>. For this we used a regular UNET @ full resolution.</p>\n\n<p><strong>2. Refine the mask using patches around the mask contour</strong> </p>\n\n<p><img src=\"http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/patches-net.png\" alt=\"Patchesnet\" title=\"\"></p>\n\n<p>This second network is a residual UNET (<a href=\"https://arxiv.org/pdf/1708.04747.pdf\">https://arxiv.org/pdf/1708.04747.pdf</a>) that works with patches, we used <code>384x384</code> patches.</p>\n\n<p>The input of the network is <code>384x384x6</code> =&gt; RGB (<code>3</code>) + X (<code>2</code>) + CO (<code>1</code>)</p>\n\n<p>Where X is the L2 difference of the patch RGB vs. a computed background. For each car, using its coarse mask we (offline) computed the background using the 16 views and basic pixel statistics:</p>\n\n<p><img src=\"http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/carvana-background.png\" alt=\"Computed background\" title=\"\"></p>\n\n<p>When feeding the L2 difference, pixels in magenta will be substituted by gaussian noise with same mean/variance that non-magenta pixels, that's the first channel of X, the second channel is just an 1,0 mask of whether the pixel is magenta or not.</p>\n\n<p>Finally CO is just the probabilities (saved as PNG files) of the coarse mask for the previous iteration.</p>\n\n<p>We computed the effective receptive field of our residual Unet and was ~224 px. Since patches are going to overlap on inference we weighted each pixel in the patch so center pixels contributed more than edge pixels (ratio 1:4).</p>\n\n<p>For augmentation we just used random horizontal flips, both for training and TTA.</p>\n\n<p><strong>BUG</strong>\nUnfortunately we had a bug and some predictions came out like this:\n<img src=\"http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/bad-prediction.png\" alt=\"Bad prediction\" title=\"\"></p>\n\n<p>Which is due (we think!) to a race condition in our Keras inference/generator pipeline... and meant we had to blindly reject predictions whose refined mask was different (dice &lt;0.999) than the coarse one... neglecting scenarios where the refined one made a lot of corrections.</p>",
  "messages": [
    {
      "id": "225129",
      "postDate": "09/28/2017 10:25:08",
      "content": "<p>Our solution has two stages that can be run recursively (but we only ran then once b/c of lack of time) =&gt;</p>\n\n<p><strong>1. Predict a coarse mask (CO)</strong>. For this we used a regular UNET @ full resolution.</p>\n\n<p><strong>2. Refine the mask using patches around the mask contour</strong> </p>\n\n<p><img src=\"http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/patches-net.png\" alt=\"Patchesnet\" title=\"\"></p>\n\n<p>This second network is a residual UNET (<a href=\"https://arxiv.org/pdf/1708.04747.pdf\">https://arxiv.org/pdf/1708.04747.pdf</a>) that works with patches, we used <code>384x384</code> patches.</p>\n\n<p>The input of the network is <code>384x384x6</code> =&gt; RGB (<code>3</code>) + X (<code>2</code>) + CO (<code>1</code>)</p>\n\n<p>Where X is the L2 difference of the patch RGB vs. a computed background. For each car, using its coarse mask we (offline) computed the background using the 16 views and basic pixel statistics:</p>\n\n<p><img src=\"http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/carvana-background.png\" alt=\"Computed background\" title=\"\"></p>\n\n<p>When feeding the L2 difference, pixels in magenta will be substituted by gaussian noise with same mean/variance that non-magenta pixels, that's the first channel of X, the second channel is just an 1,0 mask of whether the pixel is magenta or not.</p>\n\n<p>Finally CO is just the probabilities (saved as PNG files) of the coarse mask for the previous iteration.</p>\n\n<p>We computed the effective receptive field of our residual Unet and was ~224 px. Since patches are going to overlap on inference we weighted each pixel in the patch so center pixels contributed more than edge pixels (ratio 1:4).</p>\n\n<p>For augmentation we just used random horizontal flips, both for training and TTA.</p>\n\n<p><strong>BUG</strong>\nUnfortunately we had a bug and some predictions came out like this:\n<img src=\"http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/bad-prediction.png\" alt=\"Bad prediction\" title=\"\"></p>\n\n<p>Which is due (we think!) to a race condition in our Keras inference/generator pipeline... and meant we had to blindly reject predictions whose refined mask was different (dice &lt;0.999) than the coarse one... neglecting scenarios where the refined one made a lot of corrections.</p>",
      "rawMarkdown": "Our solution has two stages that can be run recursively (but we only ran then once b/c of lack of time) =&gt;\n\n**1. Predict a coarse mask (CO)**. For this we used a regular UNET @ full resolution.\n\n**2. Refine the mask using patches around the mask contour** \n\n![Patchesnet][1]\n\nThis second network is a residual UNET (https://arxiv.org/pdf/1708.04747.pdf) that works with patches, we used `384x384` patches.\n\nThe input of the network is `384x384x6` =&gt; RGB (`3`) + X (`2`) + CO (`1`)\n\nWhere X is the L2 difference of the patch RGB vs. a computed background. For each car, using its coarse mask we (offline) computed the background using the 16 views and basic pixel statistics:\n\n![Computed background][2]\n\nWhen feeding the L2 difference, pixels in magenta will be substituted by gaussian noise with same mean/variance that non-magenta pixels, that's the first channel of X, the second channel is just an 1,0 mask of whether the pixel is magenta or not.\n\nFinally CO is just the probabilities (saved as PNG files) of the coarse mask for the previous iteration.\n\nWe computed the effective receptive field of our residual Unet and was ~224 px. Since patches are going to overlap on inference we weighted each pixel in the patch so center pixels contributed more than edge pixels (ratio 1:4).\n\nFor augmentation we just used random horizontal flips, both for training and TTA.\n\n**BUG**\nUnfortunately we had a bug and some predictions came out like this:\n![Bad prediction][3]\n\nWhich is due (we think!) to a race condition in our Keras inference/generator pipeline... and meant we had to blindly reject predictions whose refined mask was different (dice &lt;0.999) than the coarse one... neglecting scenarios where the refined one made a lot of corrections.\n\n  [1]: http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/patches-net.png\n  [2]: http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/carvana-background.png\n  [3]: http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/bad-prediction.png",
      "votes": null
    },
    {
      "id": "225137",
      "postDate": "09/28/2017 10:42:50",
      "content": "<p>Very nice soultion. Try to submit fixed masks and tell us what the results could be!</p>",
      "rawMarkdown": "Very nice soultion. Try to submit fixed masks and tell us what the results could be!",
      "votes": null
    },
    {
      "id": "225139",
      "postDate": "09/28/2017 10:44:55",
      "content": "<p>our code is available here <a href=\"https://github.com/pavelgonchar/PatchesNet\">https://github.com/pavelgonchar/PatchesNet</a></p>",
      "rawMarkdown": "our code is available here https://github.com/pavelgonchar/PatchesNet",
      "votes": null
    },
    {
      "id": "225276",
      "postDate": "09/28/2017 16:30:36",
      "content": "<p>I used similar approach and got the same hole middle in the car bug!</p>\n\n<p>About giving information about the background, did it help? </p>\n\n<p>I tried somewhat similar approach and at first the results seem to improve - but there was a catch. The output can never be far better than given hint, and the improvement was marginal. I guess the network relies on the hint too much and doesn't learn much from the actual test image. </p>",
      "rawMarkdown": "I used similar approach and got the same hole middle in the car bug!\n\nAbout giving information about the background, did it help? \n\nI tried somewhat similar approach and at first the results seem to improve - but there was a catch. The output can never be far better than given hint, and the improvement was marginal. I guess the network relies on the hint too much and doesn't learn much from the actual test image.",
      "votes": null
    },
    {
      "id": "225341",
      "postDate": "09/28/2017 19:34:04",
      "content": "<p>Andres,\nCongrats on the solution.\nwould you have the paper in english for stage 2 by any chance ?</p>",
      "rawMarkdown": "Andres,\nCongrats on the solution.\nwould you have the paper in english for stage 2 by any chance ?",
      "votes": null
    },
    {
      "id": "225342",
      "postDate": "09/28/2017 19:38:57",
      "content": "<p>No, I just translated it using Google docs, but it's basically a Unet with some residual blocks.</p>",
      "rawMarkdown": "No, I just translated it using Google docs, but it's basically a Unet with some residual blocks.",
      "votes": null
    },
    {
      "id": "225356",
      "postDate": "09/28/2017 20:26:37",
      "content": "<p>Thanks @Andres and congratulations. Very good solution.</p>",
      "rawMarkdown": "Thanks @Andres and congratulations. Very good solution.",
      "votes": null
    },
    {
      "id": "225576",
      "postDate": "09/29/2017 14:49:08",
      "content": "<p>It did on cases where the background was too similar to the car so the net supposedly uses L2 distance to discern... but it's difficult to say for sure with more comparisons.</p>\n\n<p>We had better predictions with RGB + X + CO (we also tested the other variations)... but unfortunately the race condition bug and the fact that we only had time to refine ~80% of the coarse masks prevented us from getting better score.</p>\n\n<p>The other obvious thing we could have done in hindsight was to be a bit more aggressive with augmentations... we just did flips... and I think we should have done small rotations + zooming.</p>",
      "rawMarkdown": "It did on cases where the background was too similar to the car so the net supposedly uses L2 distance to discern... but it's difficult to say for sure with more comparisons.\n\nWe had better predictions with RGB + X + CO (we also tested the other variations)... but unfortunately the race condition bug and the fact that we only had time to refine ~80% of the coarse masks prevented us from getting better score.\n\nThe other obvious thing we could have done in hindsight was to be a bit more aggressive with augmentations... we just did flips... and I think we should have done small rotations + zooming.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 225137,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "09/28/2017 10:42:50",
      "content": "<p>Very nice soultion. Try to submit fixed masks and tell us what the results could be!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225139,
      "author_name": "pavelgonchar",
      "author_url": "",
      "post_date": "09/28/2017 10:44:55",
      "content": "<p>our code is available here <a href=\"https://github.com/pavelgonchar/PatchesNet\">https://github.com/pavelgonchar/PatchesNet</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 225276,
      "author_name": "jandjenter",
      "author_url": "",
      "post_date": "09/28/2017 16:30:36",
      "content": "<p>I used similar approach and got the same hole middle in the car bug!</p>\n\n<p>About giving information about the background, did it help? </p>\n\n<p>I tried somewhat similar approach and at first the results seem to improve - but there was a catch. The output can never be far better than given hint, and the improvement was marginal. I guess the network relies on the hint too much and doesn't learn much from the actual test image. </p>",
      "votes": null,
      "replies": [
        {
          "id": 225576,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/29/2017 14:49:08",
          "content": "<p>It did on cases where the background was too similar to the car so the net supposedly uses L2 distance to discern... but it's difficult to say for sure with more comparisons.</p>\n\n<p>We had better predictions with RGB + X + CO (we also tested the other variations)... but unfortunately the race condition bug and the fact that we only had time to refine ~80% of the coarse masks prevented us from getting better score.</p>\n\n<p>The other obvious thing we could have done in hindsight was to be a bit more aggressive with augmentations... we just did flips... and I think we should have done small rotations + zooming.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225341,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "09/28/2017 19:34:04",
      "content": "<p>Andres,\nCongrats on the solution.\nwould you have the paper in english for stage 2 by any chance ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 225342,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/28/2017 19:38:57",
          "content": "<p>No, I just translated it using Google docs, but it's basically a Unet with some residual blocks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225356,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "09/28/2017 20:26:37",
      "content": "<p>Thanks @Andres and congratulations. Very good solution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "225129": "Our solution has two stages that can be run recursively (but we only ran then once b/c of lack of time) =&gt;\n\n**1. Predict a coarse mask (CO)**. For this we used a regular UNET @ full resolution.\n\n**2. Refine the mask using patches around the mask contour** \n\n![Patchesnet][1]\n\nThis second network is a residual UNET (https://arxiv.org/pdf/1708.04747.pdf) that works with patches, we used `384x384` patches.\n\nThe input of the network is `384x384x6` =&gt; RGB (`3`) + X (`2`) + CO (`1`)\n\nWhere X is the L2 difference of the patch RGB vs. a computed background. For each car, using its coarse mask we (offline) computed the background using the 16 views and basic pixel statistics:\n\n![Computed background][2]\n\nWhen feeding the L2 difference, pixels in magenta will be substituted by gaussian noise with same mean/variance that non-magenta pixels, that's the first channel of X, the second channel is just an 1,0 mask of whether the pixel is magenta or not.\n\nFinally CO is just the probabilities (saved as PNG files) of the coarse mask for the previous iteration.\n\nWe computed the effective receptive field of our residual Unet and was ~224 px. Since patches are going to overlap on inference we weighted each pixel in the patch so center pixels contributed more than edge pixels (ratio 1:4).\n\nFor augmentation we just used random horizontal flips, both for training and TTA.\n\n**BUG**\nUnfortunately we had a bug and some predictions came out like this:\n![Bad prediction][3]\n\nWhich is due (we think!) to a race condition in our Keras inference/generator pipeline... and meant we had to blindly reject predictions whose refined mask was different (dice &lt;0.999) than the coarse one... neglecting scenarios where the refined one made a lot of corrections.\n\n  [1]: http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/patches-net.png\n  [2]: http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/carvana-background.png\n  [3]: http://s3.eu-central-1.amazonaws.com/kaggle-carvana-team73/bad-prediction.png",
    "225137": "Very nice soultion. Try to submit fixed masks and tell us what the results could be!",
    "225139": "our code is available here https://github.com/pavelgonchar/PatchesNet",
    "225276": "I used similar approach and got the same hole middle in the car bug!\n\nAbout giving information about the background, did it help? \n\nI tried somewhat similar approach and at first the results seem to improve - but there was a catch. The output can never be far better than given hint, and the improvement was marginal. I guess the network relies on the hint too much and doesn't learn much from the actual test image.",
    "225341": "Andres,\nCongrats on the solution.\nwould you have the paper in english for stage 2 by any chance ?",
    "225342": "No, I just translated it using Google docs, but it's basically a Unet with some residual blocks.",
    "225356": "Thanks @Andres and congratulations. Very good solution.",
    "225576": "It did on cases where the background was too similar to the car so the net supposedly uses L2 distance to discern... but it's difficult to say for sure with more comparisons.\n\nWe had better predictions with RGB + X + CO (we also tested the other variations)... but unfortunately the race condition bug and the fact that we only had time to refine ~80% of the coarse masks prevented us from getting better score.\n\nThe other obvious thing we could have done in hindsight was to be a bit more aggressive with augmentations... we just did flips... and I think we should have done small rotations + zooming."
  },
  "source": "meta"
}