{
  "id": 114863,
  "title": "Bounding Boxes instead of Segmentation",
  "url": "/competitions/understanding_cloud_organization/discussion/114863",
  "author_name": "",
  "post_date": "2019-10-29T18:46:30.805791800Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>EDIT: The sponsors prefer non-rectangular solutions as explained <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a>. So consider this post for educational purposes only.</p>\n\n<p>As an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. You have probably noticed that when you upload photos to social media, vision recognition software locates all the faces in the photo in places boxes around the faces. Similarly, you can design a model to locale Fish, Flower, Gravel, and Sugar clouds in these images and place boxes around these cloud formations.  </p>\n\n<p>I'm just learning object localization myself, so I don't know what the best techniques are but I thought I'd mention that this competition is an opportunity to learn object localization. I posted a starter kernel <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a> that you can play around with. Currently it scores LB 0.611. I'm sure it can be improved.</p>\n\n<p>Here are some bounding boxes that my starter kernel made. The yellow outlines are the true masks and the blue outlines are the predicted bounding boxes:</p>\n\n<h1>Fish</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9607c619065d581f14822170fbc647a7%2Ffish.png?generation=1572374576590570&amp;alt=media\" alt=\"\"></p>\n\n<h1>Flower</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F338b9ea4ab999e4c265911d55f6e39b6%2Fflower.png?generation=1572374622258245&amp;alt=media\" alt=\"\"></p>\n\n<h1>Gravel</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fd77e03268eb17b4201cd30fe39b80e34%2Fgravel.png?generation=1572374636943367&amp;alt=media\" alt=\"\"></p>\n\n<h1>Sugar</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff2bafbbf4fccc3a0611f78bbee051927%2Fsugar.png?generation=1572374652327131&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "660894",
      "postDate": "10/29/2019 18:46:30",
      "content": "<p>EDIT: The sponsors prefer non-rectangular solutions as explained <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a>. So consider this post for educational purposes only.</p>\n\n<p>As an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. You have probably noticed that when you upload photos to social media, vision recognition software locates all the faces in the photo in places boxes around the faces. Similarly, you can design a model to locale Fish, Flower, Gravel, and Sugar clouds in these images and place boxes around these cloud formations.  </p>\n\n<p>I'm just learning object localization myself, so I don't know what the best techniques are but I thought I'd mention that this competition is an opportunity to learn object localization. I posted a starter kernel <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a> that you can play around with. Currently it scores LB 0.611. I'm sure it can be improved.</p>\n\n<p>Here are some bounding boxes that my starter kernel made. The yellow outlines are the true masks and the blue outlines are the predicted bounding boxes:</p>\n\n<h1>Fish</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9607c619065d581f14822170fbc647a7%2Ffish.png?generation=1572374576590570&amp;alt=media\" alt=\"\"></p>\n\n<h1>Flower</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F338b9ea4ab999e4c265911d55f6e39b6%2Fflower.png?generation=1572374622258245&amp;alt=media\" alt=\"\"></p>\n\n<h1>Gravel</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fd77e03268eb17b4201cd30fe39b80e34%2Fgravel.png?generation=1572374636943367&amp;alt=media\" alt=\"\"></p>\n\n<h1>Sugar</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff2bafbbf4fccc3a0611f78bbee051927%2Fsugar.png?generation=1572374652327131&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "EDIT: The sponsors prefer non-rectangular solutions as explained [here][1]. So consider this post for educational purposes only.\n\nAs an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. You have probably noticed that when you upload photos to social media, vision recognition software locates all the faces in the photo in places boxes around the faces. Similarly, you can design a model to locale Fish, Flower, Gravel, and Sugar clouds in these images and place boxes around these cloud formations.  \n  \nI'm just learning object localization myself, so I don't know what the best techniques are but I thought I'd mention that this competition is an opportunity to learn object localization. I posted a starter kernel [here][1] that you can play around with. Currently it scores LB 0.611. I'm sure it can be improved.\n\nHere are some bounding boxes that my starter kernel made. The yellow outlines are the true masks and the blue outlines are the predicted bounding boxes:\n# Fish\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9607c619065d581f14822170fbc647a7%2Ffish.png?generation=1572374576590570&amp;alt=media)\n# Flower\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F338b9ea4ab999e4c265911d55f6e39b6%2Fflower.png?generation=1572374622258245&amp;alt=media)\n# Gravel\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fd77e03268eb17b4201cd30fe39b80e34%2Fgravel.png?generation=1572374636943367&amp;alt=media)\n# Sugar\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff2bafbbf4fccc3a0611f78bbee051927%2Fsugar.png?generation=1572374652327131&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587\n\n\n\n\n[1]: https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58",
      "votes": null
    },
    {
      "id": "660905",
      "postDate": "10/29/2019 19:05:07",
      "content": "<blockquote>\n  <p>As an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. </p>\n</blockquote>\n\n<p>Why object detection is safer than segmentation in terms of overfitting?</p>",
      "rawMarkdown": "&gt; As an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. \n\nWhy object detection is safer than segmentation in terms of overfitting?",
      "votes": null
    },
    {
      "id": "660923",
      "postDate": "10/29/2019 19:19:19",
      "content": "<p>just skipped a beautiful localization tutorial earlier in the morning because i thought we need to focus on segmentation,thanks for this post,going back to that beautiful localization tutorial :)</p>",
      "rawMarkdown": "just skipped a beautiful localization tutorial earlier in the morning because i thought we need to focus on segmentation,thanks for this post,going back to that beautiful localization tutorial :)",
      "votes": null
    },
    {
      "id": "660932",
      "postDate": "10/29/2019 19:29:16",
      "content": "<p>In the same way that a linear model generalizes more than a non-linear model, I suspect that predicting rectangles generalizes more than predicting non-rectangle shapes. (It's the bias-variance trade off. Boxes have more bias less variance). </p>\n\n<p>Also note the the training masks are the <strong>union</strong> of annotators which is a non-rectangle shape. However the <strong>intersection</strong> of the annotators (the region of highest confidence) is a rectangle shape.</p>",
      "rawMarkdown": "In the same way that a linear model generalizes more than a non-linear model, I suspect that predicting rectangles generalizes more than predicting non-rectangle shapes. (It's the bias-variance trade off. Boxes have more bias less variance). \n\nAlso note the the training masks are the **union** of annotators which is a non-rectangle shape. However the **intersection** of the annotators (the region of highest confidence) is a rectangle shape.",
      "votes": null
    },
    {
      "id": "661015",
      "postDate": "10/29/2019 21:39:36",
      "content": "<p>I remember the organisers prefer segmentation masks than bounding boxes for their further research purposes.</p>",
      "rawMarkdown": "I remember the organisers prefer segmentation masks than bounding boxes for their further research purposes.",
      "votes": null
    },
    {
      "id": "661042",
      "postDate": "10/29/2019 22:47:05",
      "content": "<p>Thanks for the response. I found the original discussion about object detection <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587\">here</a>. It seems that the sponsors do prefer non-rectangular masks.</p>",
      "rawMarkdown": "Thanks for the response. I found the original discussion about object detection [here][1]. It seems that the sponsors do prefer non-rectangular masks.\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587",
      "votes": null
    },
    {
      "id": "661443",
      "postDate": "10/30/2019 09:59:02",
      "content": "<p>But they don't restrict us to not use bounding boxes. Perhaps a mixture would be better.</p>",
      "rawMarkdown": "But they don't restrict us to not use bounding boxes. Perhaps a mixture would be better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 660905,
      "author_name": "naivelamb",
      "author_url": "",
      "post_date": "10/29/2019 19:05:07",
      "content": "<blockquote>\n  <p>As an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. </p>\n</blockquote>\n\n<p>Why object detection is safer than segmentation in terms of overfitting?</p>",
      "votes": null,
      "replies": [
        {
          "id": 660932,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/29/2019 19:29:16",
          "content": "<p>In the same way that a linear model generalizes more than a non-linear model, I suspect that predicting rectangles generalizes more than predicting non-rectangle shapes. (It's the bias-variance trade off. Boxes have more bias less variance). </p>\n\n<p>Also note the the training masks are the <strong>union</strong> of annotators which is a non-rectangle shape. However the <strong>intersection</strong> of the annotators (the region of highest confidence) is a rectangle shape.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 660923,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "10/29/2019 19:19:19",
      "content": "<p>just skipped a beautiful localization tutorial earlier in the morning because i thought we need to focus on segmentation,thanks for this post,going back to that beautiful localization tutorial :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 661015,
      "author_name": "gogo827jz",
      "author_url": "",
      "post_date": "10/29/2019 21:39:36",
      "content": "<p>I remember the organisers prefer segmentation masks than bounding boxes for their further research purposes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 661042,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/29/2019 22:47:05",
          "content": "<p>Thanks for the response. I found the original discussion about object detection <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587\">here</a>. It seems that the sponsors do prefer non-rectangular masks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 661443,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "10/30/2019 09:59:02",
          "content": "<p>But they don't restrict us to not use bounding boxes. Perhaps a mixture would be better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "660894": "EDIT: The sponsors prefer non-rectangular solutions as explained [here][1]. So consider this post for educational purposes only.\n\nAs an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. You have probably noticed that when you upload photos to social media, vision recognition software locates all the faces in the photo in places boxes around the faces. Similarly, you can design a model to locale Fish, Flower, Gravel, and Sugar clouds in these images and place boxes around these cloud formations.  \n  \nI'm just learning object localization myself, so I don't know what the best techniques are but I thought I'd mention that this competition is an opportunity to learn object localization. I posted a starter kernel [here][1] that you can play around with. Currently it scores LB 0.611. I'm sure it can be improved.\n\nHere are some bounding boxes that my starter kernel made. The yellow outlines are the true masks and the blue outlines are the predicted bounding boxes:\n# Fish\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F9607c619065d581f14822170fbc647a7%2Ffish.png?generation=1572374576590570&amp;alt=media)\n# Flower\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F338b9ea4ab999e4c265911d55f6e39b6%2Fflower.png?generation=1572374622258245&amp;alt=media)\n# Gravel\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fd77e03268eb17b4201cd30fe39b80e34%2Fgravel.png?generation=1572374636943367&amp;alt=media)\n# Sugar\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff2bafbbf4fccc3a0611f78bbee051927%2Fsugar.png?generation=1572374652327131&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587\n\n\n\n\n[1]: https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58",
    "660905": "&gt; As an alternative to segmentation (and safety against overfitting), you can try to use object localization instead of segmentation. \n\nWhy object detection is safer than segmentation in terms of overfitting?",
    "660923": "just skipped a beautiful localization tutorial earlier in the morning because i thought we need to focus on segmentation,thanks for this post,going back to that beautiful localization tutorial :)",
    "660932": "In the same way that a linear model generalizes more than a non-linear model, I suspect that predicting rectangles generalizes more than predicting non-rectangle shapes. (It's the bias-variance trade off. Boxes have more bias less variance). \n\nAlso note the the training masks are the **union** of annotators which is a non-rectangle shape. However the **intersection** of the annotators (the region of highest confidence) is a rectangle shape.",
    "661015": "I remember the organisers prefer segmentation masks than bounding boxes for their further research purposes.",
    "661042": "Thanks for the response. I found the original discussion about object detection [here][1]. It seems that the sponsors do prefer non-rectangular masks.\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/104587",
    "661443": "But they don't restrict us to not use bounding boxes. Perhaps a mixture would be better."
  },
  "source": "meta"
}