{
  "id": 37096,
  "title": "Any advice/tips on how to solve this problem?",
  "url": "/competitions/carvana-image-masking-challenge/discussion/37096",
  "author_name": "",
  "post_date": "2017-07-26T23:58:44.539001600Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi Kaggler and welcome to one more challenge/competition. This is the first time I see a task like this, I would like to know from you guys what are reasonable things/algorithms, tips and tricks to try out  to automatically identify the boundaries of the car in an image? Is there good resources, previous competition or  good approaches used to handle  Image Masking or similar problems? </p>\n\n<p>I'm Just getting started in this competition and I would appreciate any help and some words of wisdom.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "207598",
      "postDate": "07/26/2017 23:58:44",
      "content": "<p>Hi Kaggler and welcome to one more challenge/competition. This is the first time I see a task like this, I would like to know from you guys what are reasonable things/algorithms, tips and tricks to try out  to automatically identify the boundaries of the car in an image? Is there good resources, previous competition or  good approaches used to handle  Image Masking or similar problems? </p>\n\n<p>I'm Just getting started in this competition and I would appreciate any help and some words of wisdom.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Kaggler and welcome to one more challenge/competition. This is the first time I see a task like this, I would like to know from you guys what are reasonable things/algorithms, tips and tricks to try out  to automatically identify the boundaries of the car in an image? Is there good resources, previous competition or  good approaches used to handle  Image Masking or similar problems? \n\nI'm Just getting started in this competition and I would appreciate any help and some words of wisdom.\n\nThanks!",
      "votes": null
    },
    {
      "id": "207602",
      "postDate": "07/27/2017 00:14:11",
      "content": "<p>A lot of knowledge from <a href=\"https://www.kaggle.com/c/ultrasound-nerve-segmentation\">Ultrasound Nerve Segmentation</a> can be transfered to this problem, although this problem looks significasntly easier.</p>",
      "rawMarkdown": "A lot of knowledge from [Ultrasound Nerve Segmentation][1] can be transfered to this problem, although this problem looks significasntly easier.\n\n\n  [1]: https://www.kaggle.com/c/ultrasound-nerve-segmentation",
      "votes": null
    },
    {
      "id": "207606",
      "postDate": "07/27/2017 01:11:26",
      "content": "<p>The biggest challenge with this type of problem is that edge detection doesn't work well with reflective objects, especially ones with irregular shapes. Fortunately, the vehicles are positioned and moved uniformly (one camera angle, 16 rotation steps at 22.5 degree intervals).</p>\n\n<p>I don't have any advice for ML frameworks or algorithms to use, but here are some tips for procedural methods you may be able to use as additional inputs:</p>\n\n<p>At a very basic level, I think first gathering edge information can be valuable. See <a href=\"https://en.wikipedia.org/wiki/Canny_edge_detector\">Canny edge detector</a>. This will result in a lot of noise. Some of the edges detected will be actual boundaries of the vehicle while others will be clues that an area is <em>not</em> part of the vehicle (e.g. the edges in \"CARVANA\" in the background). One useful method to deal with edge data for feature detection is the <a href=\"https://en.wikipedia.org/wiki/Hough_transform\">Hough transform</a>, which can help you detect straight lines. Perfectly straight lines are uncommon on vehicles, but if you're searching in localized areas, it can help you detect features like windshields and undercarriages.</p>\n\n<p>Some features are easier to detect using naive methods than others, and can be used as anchors. For example, tires are one of the most straightforward features to detect. </p>\n\n<p><img src=\"http://i.imgur.com/fcGMkEc.png\" alt=\"wheel detection\" title=\"\"></p>\n\n<p>Sharing detected features across images within a set (id_01.jpg..id_16.jpg) may also be helpful. For example, if you look at a vehicle's 90 or 270 rotation angles (labeled '05' and '13' respectively), the tires will be perpendicular to the camera and more easily detectable. The positions of the tires can be correlated to the size and shape of the vehicle. If you are using a 3D projection mapping approach, detecting the tires in one angle can tell you where to expect them to be in another angle even if they aren't visible (0 or 180 degrees)</p>\n\n<p>It's common for the top parts of the car to be missing edge pixels (black paint can reflect white pixels in front of a white background). The boundaries, then, would need to be interpolated based on anchors, such as features on the vehicle (the top of the windshield will be relative to the top of the car) or features outside of the vehicle (points where background features get blocked/altered).</p>",
      "rawMarkdown": "The biggest challenge with this type of problem is that edge detection doesn't work well with reflective objects, especially ones with irregular shapes. Fortunately, the vehicles are positioned and moved uniformly (one camera angle, 16 rotation steps at 22.5 degree intervals).\n\nI don't have any advice for ML frameworks or algorithms to use, but here are some tips for procedural methods you may be able to use as additional inputs:\n\nAt a very basic level, I think first gathering edge information can be valuable. See [Canny edge detector][1]. This will result in a lot of noise. Some of the edges detected will be actual boundaries of the vehicle while others will be clues that an area is *not* part of the vehicle (e.g. the edges in \"CARVANA\" in the background). One useful method to deal with edge data for feature detection is the [Hough transform][2], which can help you detect straight lines. Perfectly straight lines are uncommon on vehicles, but if you're searching in localized areas, it can help you detect features like windshields and undercarriages.\n\nSome features are easier to detect using naive methods than others, and can be used as anchors. For example, tires are one of the most straightforward features to detect. \n\n![wheel detection][3]\n\nSharing detected features across images within a set (id_01.jpg..id_16.jpg) may also be helpful. For example, if you look at a vehicle's 90 or 270 rotation angles (labeled '05' and '13' respectively), the tires will be perpendicular to the camera and more easily detectable. The positions of the tires can be correlated to the size and shape of the vehicle. If you are using a 3D projection mapping approach, detecting the tires in one angle can tell you where to expect them to be in another angle even if they aren't visible (0 or 180 degrees)\n\nIt's common for the top parts of the car to be missing edge pixels (black paint can reflect white pixels in front of a white background). The boundaries, then, would need to be interpolated based on anchors, such as features on the vehicle (the top of the windshield will be relative to the top of the car) or features outside of the vehicle (points where background features get blocked/altered).\n\n  [1]: https://en.wikipedia.org/wiki/Canny_edge_detector\n  [2]: https://en.wikipedia.org/wiki/Hough_transform\n  [3]: http://i.imgur.com/fcGMkEc.png",
      "votes": null
    },
    {
      "id": "207661",
      "postDate": "07/27/2017 05:26:00",
      "content": "<p>Hi @Pryor\nI suspect this is a potential candidate for a deep learning approach where you would be predicting a 0 or 1 for each pixel position. In other words, you could have a convolutional neural network (CNN) which has the image as the input and a manually segmented binary mask as the output: then train and predict:-)</p>\n\n<p>I guess a CNN approach would easily outperform other more naive approaches such as edge detection which the organisers could have tried very easily even in an image processing studio.\nCheers,\n@bk0000</p>",
      "rawMarkdown": "Hi @Pryor\nI suspect this is a potential candidate for a deep learning approach where you would be predicting a 0 or 1 for each pixel position. In other words, you could have a convolutional neural network (CNN) which has the image as the input and a manually segmented binary mask as the output: then train and predict:-)\n\nI guess a CNN approach would easily outperform other more naive approaches such as edge detection which the organisers could have tried very easily even in an image processing studio.\nCheers,\n@bk0000",
      "votes": null
    },
    {
      "id": "207743",
      "postDate": "07/27/2017 10:38:13",
      "content": "<p>Thanks Vladimir!</p>",
      "rawMarkdown": "Thanks Vladimir!",
      "votes": null
    },
    {
      "id": "207823",
      "postDate": "07/27/2017 15:40:59",
      "content": "<p><strong>Segmentation</strong></p>\n\n<ul>\n<li>Fully Convolutional Networks for Semantic Segmentation =&gt; <a href=\"https://arxiv.org/abs/1411.4038\">https://arxiv.org/abs/1411.4038</a></li>\n<li>U-Net: Convolutional Networks for Biomedical Image Segmentation =&gt; <a href=\"http://arxiv.org/abs/1505.04597\">http://arxiv.org/abs/1505.04597</a></li>\n<li>Multi-Scale Context Aggregation by Dilated Convolutions =&gt; <a href=\"https://arxiv.org/pdf/1511.07122v3.pdf\">https://arxiv.org/pdf/1511.07122v3.pdf</a></li>\n<li>Conditional Random Fields as Recurrent Neural Networks =&gt; <a href=\"https://arxiv.org/abs/1502.03240\">https://arxiv.org/abs/1502.03240</a></li>\n<li>ParseNet: Looking Wider to See Better =&gt; <a href=\"http://arxiv.org/abs/1506.04579\">http://arxiv.org/abs/1506.04579</a></li>\n<li>Instance-aware Semantic Segmentation via Multi-task Network Cascades =&gt; <a href=\"http://arxiv.org/abs/1512.04412\">http://arxiv.org/abs/1512.04412</a></li>\n</ul>\n\n<p><strong>Source</strong>: <a href=\"https://github.com/ternaus/NeuralNetworksSheet\">https://github.com/ternaus/NeuralNetworksSheet</a></p>",
      "rawMarkdown": "**Segmentation**\n\n* Fully Convolutional Networks for Semantic Segmentation =&gt; https://arxiv.org/abs/1411.4038\n* U-Net: Convolutional Networks for Biomedical Image Segmentation =&gt; http://arxiv.org/abs/1505.04597\n* Multi-Scale Context Aggregation by Dilated Convolutions =&gt; https://arxiv.org/pdf/1511.07122v3.pdf\n* Conditional Random Fields as Recurrent Neural Networks =&gt; https://arxiv.org/abs/1502.03240\n* ParseNet: Looking Wider to See Better =&gt; http://arxiv.org/abs/1506.04579\n* Instance-aware Semantic Segmentation via Multi-task Network Cascades =&gt; http://arxiv.org/abs/1512.04412\n\n**Source**: https://github.com/ternaus/NeuralNetworksSheet",
      "votes": null
    },
    {
      "id": "207824",
      "postDate": "07/27/2017 15:43:37",
      "content": "<p>Any info about winning solutions or interviews?</p>",
      "rawMarkdown": "Any info about winning solutions or interviews?",
      "votes": null
    },
    {
      "id": "208560",
      "postDate": "07/30/2017 07:32:14",
      "content": "<p>I believe we can use classical algorithms for image segmentation (like <a href=\"https://github.com/andrewssobral/bgslibrary\">https://github.com/andrewssobral/bgslibrary</a>) as a pre-processing step to generate masks with decent accuracy, then use them to train a CNN. This should achieve better results than training a CNN directly from the original images.</p>",
      "rawMarkdown": "I believe we can use classical algorithms for image segmentation (like https://github.com/andrewssobral/bgslibrary) as a pre-processing step to generate masks with decent accuracy, then use them to train a CNN. This should achieve better results than training a CNN directly from the original images.",
      "votes": null
    },
    {
      "id": "208741",
      "postDate": "07/30/2017 21:40:24",
      "content": "<p>Superb Shinto, thanks. this is quite helpful.</p>",
      "rawMarkdown": "Superb Shinto, thanks. this is quite helpful.",
      "votes": null
    },
    {
      "id": "210796",
      "postDate": "08/07/2017 04:40:42",
      "content": "<p>Adding to Shinto's list =&gt; \n<a href=\"https://arxiv.org/abs/1704.06857\">A Review on Deep Learning Techniques Applied to Semantic Segmentation</a></p>\n\n<p><a href=\"http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review\">Qure Blog on segmentation</a></p>",
      "rawMarkdown": "Adding to Shinto's list =&gt; \n[A Review on Deep Learning Techniques Applied to Semantic Segmentation][1]\n\n[Qure Blog on segmentation][2]\n\n\n  [1]: https://arxiv.org/abs/1704.06857\n  [2]: http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 207602,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "07/27/2017 00:14:11",
      "content": "<p>A lot of knowledge from <a href=\"https://www.kaggle.com/c/ultrasound-nerve-segmentation\">Ultrasound Nerve Segmentation</a> can be transfered to this problem, although this problem looks significasntly easier.</p>",
      "votes": null,
      "replies": [
        {
          "id": 207743,
          "author_name": "davidfumo",
          "author_url": "",
          "post_date": "07/27/2017 10:38:13",
          "content": "<p>Thanks Vladimir!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 207824,
          "author_name": "shinto",
          "author_url": "",
          "post_date": "07/27/2017 15:43:37",
          "content": "<p>Any info about winning solutions or interviews?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 207606,
      "author_name": "brianshaler",
      "author_url": "",
      "post_date": "07/27/2017 01:11:26",
      "content": "<p>The biggest challenge with this type of problem is that edge detection doesn't work well with reflective objects, especially ones with irregular shapes. Fortunately, the vehicles are positioned and moved uniformly (one camera angle, 16 rotation steps at 22.5 degree intervals).</p>\n\n<p>I don't have any advice for ML frameworks or algorithms to use, but here are some tips for procedural methods you may be able to use as additional inputs:</p>\n\n<p>At a very basic level, I think first gathering edge information can be valuable. See <a href=\"https://en.wikipedia.org/wiki/Canny_edge_detector\">Canny edge detector</a>. This will result in a lot of noise. Some of the edges detected will be actual boundaries of the vehicle while others will be clues that an area is <em>not</em> part of the vehicle (e.g. the edges in \"CARVANA\" in the background). One useful method to deal with edge data for feature detection is the <a href=\"https://en.wikipedia.org/wiki/Hough_transform\">Hough transform</a>, which can help you detect straight lines. Perfectly straight lines are uncommon on vehicles, but if you're searching in localized areas, it can help you detect features like windshields and undercarriages.</p>\n\n<p>Some features are easier to detect using naive methods than others, and can be used as anchors. For example, tires are one of the most straightforward features to detect. </p>\n\n<p><img src=\"http://i.imgur.com/fcGMkEc.png\" alt=\"wheel detection\" title=\"\"></p>\n\n<p>Sharing detected features across images within a set (id_01.jpg..id_16.jpg) may also be helpful. For example, if you look at a vehicle's 90 or 270 rotation angles (labeled '05' and '13' respectively), the tires will be perpendicular to the camera and more easily detectable. The positions of the tires can be correlated to the size and shape of the vehicle. If you are using a 3D projection mapping approach, detecting the tires in one angle can tell you where to expect them to be in another angle even if they aren't visible (0 or 180 degrees)</p>\n\n<p>It's common for the top parts of the car to be missing edge pixels (black paint can reflect white pixels in front of a white background). The boundaries, then, would need to be interpolated based on anchors, such as features on the vehicle (the top of the windshield will be relative to the top of the car) or features outside of the vehicle (points where background features get blocked/altered).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 207661,
      "author_name": "bk0000",
      "author_url": "",
      "post_date": "07/27/2017 05:26:00",
      "content": "<p>Hi @Pryor\nI suspect this is a potential candidate for a deep learning approach where you would be predicting a 0 or 1 for each pixel position. In other words, you could have a convolutional neural network (CNN) which has the image as the input and a manually segmented binary mask as the output: then train and predict:-)</p>\n\n<p>I guess a CNN approach would easily outperform other more naive approaches such as edge detection which the organisers could have tried very easily even in an image processing studio.\nCheers,\n@bk0000</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 207823,
      "author_name": "shinto",
      "author_url": "",
      "post_date": "07/27/2017 15:40:59",
      "content": "<p><strong>Segmentation</strong></p>\n\n<ul>\n<li>Fully Convolutional Networks for Semantic Segmentation =&gt; <a href=\"https://arxiv.org/abs/1411.4038\">https://arxiv.org/abs/1411.4038</a></li>\n<li>U-Net: Convolutional Networks for Biomedical Image Segmentation =&gt; <a href=\"http://arxiv.org/abs/1505.04597\">http://arxiv.org/abs/1505.04597</a></li>\n<li>Multi-Scale Context Aggregation by Dilated Convolutions =&gt; <a href=\"https://arxiv.org/pdf/1511.07122v3.pdf\">https://arxiv.org/pdf/1511.07122v3.pdf</a></li>\n<li>Conditional Random Fields as Recurrent Neural Networks =&gt; <a href=\"https://arxiv.org/abs/1502.03240\">https://arxiv.org/abs/1502.03240</a></li>\n<li>ParseNet: Looking Wider to See Better =&gt; <a href=\"http://arxiv.org/abs/1506.04579\">http://arxiv.org/abs/1506.04579</a></li>\n<li>Instance-aware Semantic Segmentation via Multi-task Network Cascades =&gt; <a href=\"http://arxiv.org/abs/1512.04412\">http://arxiv.org/abs/1512.04412</a></li>\n</ul>\n\n<p><strong>Source</strong>: <a href=\"https://github.com/ternaus/NeuralNetworksSheet\">https://github.com/ternaus/NeuralNetworksSheet</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 208741,
          "author_name": "davidfumo",
          "author_url": "",
          "post_date": "07/30/2017 21:40:24",
          "content": "<p>Superb Shinto, thanks. this is quite helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 210796,
          "author_name": "satya758",
          "author_url": "",
          "post_date": "08/07/2017 04:40:42",
          "content": "<p>Adding to Shinto's list =&gt; \n<a href=\"https://arxiv.org/abs/1704.06857\">A Review on Deep Learning Techniques Applied to Semantic Segmentation</a></p>\n\n<p><a href=\"http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review\">Qure Blog on segmentation</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208560,
      "author_name": "petrosgk",
      "author_url": "",
      "post_date": "07/30/2017 07:32:14",
      "content": "<p>I believe we can use classical algorithms for image segmentation (like <a href=\"https://github.com/andrewssobral/bgslibrary\">https://github.com/andrewssobral/bgslibrary</a>) as a pre-processing step to generate masks with decent accuracy, then use them to train a CNN. This should achieve better results than training a CNN directly from the original images.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "207598": "Hi Kaggler and welcome to one more challenge/competition. This is the first time I see a task like this, I would like to know from you guys what are reasonable things/algorithms, tips and tricks to try out  to automatically identify the boundaries of the car in an image? Is there good resources, previous competition or  good approaches used to handle  Image Masking or similar problems? \n\nI'm Just getting started in this competition and I would appreciate any help and some words of wisdom.\n\nThanks!",
    "207602": "A lot of knowledge from [Ultrasound Nerve Segmentation][1] can be transfered to this problem, although this problem looks significasntly easier.\n\n\n  [1]: https://www.kaggle.com/c/ultrasound-nerve-segmentation",
    "207606": "The biggest challenge with this type of problem is that edge detection doesn't work well with reflective objects, especially ones with irregular shapes. Fortunately, the vehicles are positioned and moved uniformly (one camera angle, 16 rotation steps at 22.5 degree intervals).\n\nI don't have any advice for ML frameworks or algorithms to use, but here are some tips for procedural methods you may be able to use as additional inputs:\n\nAt a very basic level, I think first gathering edge information can be valuable. See [Canny edge detector][1]. This will result in a lot of noise. Some of the edges detected will be actual boundaries of the vehicle while others will be clues that an area is *not* part of the vehicle (e.g. the edges in \"CARVANA\" in the background). One useful method to deal with edge data for feature detection is the [Hough transform][2], which can help you detect straight lines. Perfectly straight lines are uncommon on vehicles, but if you're searching in localized areas, it can help you detect features like windshields and undercarriages.\n\nSome features are easier to detect using naive methods than others, and can be used as anchors. For example, tires are one of the most straightforward features to detect. \n\n![wheel detection][3]\n\nSharing detected features across images within a set (id_01.jpg..id_16.jpg) may also be helpful. For example, if you look at a vehicle's 90 or 270 rotation angles (labeled '05' and '13' respectively), the tires will be perpendicular to the camera and more easily detectable. The positions of the tires can be correlated to the size and shape of the vehicle. If you are using a 3D projection mapping approach, detecting the tires in one angle can tell you where to expect them to be in another angle even if they aren't visible (0 or 180 degrees)\n\nIt's common for the top parts of the car to be missing edge pixels (black paint can reflect white pixels in front of a white background). The boundaries, then, would need to be interpolated based on anchors, such as features on the vehicle (the top of the windshield will be relative to the top of the car) or features outside of the vehicle (points where background features get blocked/altered).\n\n  [1]: https://en.wikipedia.org/wiki/Canny_edge_detector\n  [2]: https://en.wikipedia.org/wiki/Hough_transform\n  [3]: http://i.imgur.com/fcGMkEc.png",
    "207661": "Hi @Pryor\nI suspect this is a potential candidate for a deep learning approach where you would be predicting a 0 or 1 for each pixel position. In other words, you could have a convolutional neural network (CNN) which has the image as the input and a manually segmented binary mask as the output: then train and predict:-)\n\nI guess a CNN approach would easily outperform other more naive approaches such as edge detection which the organisers could have tried very easily even in an image processing studio.\nCheers,\n@bk0000",
    "207743": "Thanks Vladimir!",
    "207823": "**Segmentation**\n\n* Fully Convolutional Networks for Semantic Segmentation =&gt; https://arxiv.org/abs/1411.4038\n* U-Net: Convolutional Networks for Biomedical Image Segmentation =&gt; http://arxiv.org/abs/1505.04597\n* Multi-Scale Context Aggregation by Dilated Convolutions =&gt; https://arxiv.org/pdf/1511.07122v3.pdf\n* Conditional Random Fields as Recurrent Neural Networks =&gt; https://arxiv.org/abs/1502.03240\n* ParseNet: Looking Wider to See Better =&gt; http://arxiv.org/abs/1506.04579\n* Instance-aware Semantic Segmentation via Multi-task Network Cascades =&gt; http://arxiv.org/abs/1512.04412\n\n**Source**: https://github.com/ternaus/NeuralNetworksSheet",
    "207824": "Any info about winning solutions or interviews?",
    "208560": "I believe we can use classical algorithms for image segmentation (like https://github.com/andrewssobral/bgslibrary) as a pre-processing step to generate masks with decent accuracy, then use them to train a CNN. This should achieve better results than training a CNN directly from the original images.",
    "208741": "Superb Shinto, thanks. this is quite helpful.",
    "210796": "Adding to Shinto's list =&gt; \n[A Review on Deep Learning Techniques Applied to Semantic Segmentation][1]\n\n[Qure Blog on segmentation][2]\n\n\n  [1]: https://arxiv.org/abs/1704.06857\n  [2]: http://blog.qure.ai/notes/semantic-segmentation-deep-learning-review"
  },
  "source": "meta"
}