{
  "id": 104587,
  "title": "Object Detection vs Instance Segmentation",
  "url": "/competitions/understanding_cloud_organization/discussion/104587",
  "author_name": "",
  "post_date": "2019-08-17T23:18:23.200096100Z",
  "votes": 35,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I'm sure that the organizers and Kaggle have a better idea of the subject than me, but one thing that confused me was: what are the advantages of building an instance segmentation model for this competition compared to an object detection (with bounding boxes) model? </p>\n\n<p>Looking at the <a href=\"https://www.kaggle.com/inversion/run-length-decoding-and-flower-submission\">official starter kernel</a>, I have the impression the masks can be represented as bounding boxes. I am also not sure I understand the advantages of segmentation masks for the end use case (help categorize a great number of unlabelled cloud images in order to alleviate climate scientists' work). Would anyone more familiar with the subject be willing to share some insights on that?</p>\n\n<p>Thank you.</p>",
  "messages": [
    {
      "id": "601633",
      "postDate": "08/17/2019 23:18:23",
      "content": "<p>I'm sure that the organizers and Kaggle have a better idea of the subject than me, but one thing that confused me was: what are the advantages of building an instance segmentation model for this competition compared to an object detection (with bounding boxes) model? </p>\n\n<p>Looking at the <a href=\"https://www.kaggle.com/inversion/run-length-decoding-and-flower-submission\">official starter kernel</a>, I have the impression the masks can be represented as bounding boxes. I am also not sure I understand the advantages of segmentation masks for the end use case (help categorize a great number of unlabelled cloud images in order to alleviate climate scientists' work). Would anyone more familiar with the subject be willing to share some insights on that?</p>\n\n<p>Thank you.</p>",
      "rawMarkdown": "I'm sure that the organizers and Kaggle have a better idea of the subject than me, but one thing that confused me was: what are the advantages of building an instance segmentation model for this competition compared to an object detection (with bounding boxes) model? \n\nLooking at the [official starter kernel](https://www.kaggle.com/inversion/run-length-decoding-and-flower-submission), I have the impression the masks can be represented as bounding boxes. I am also not sure I understand the advantages of segmentation masks for the end use case (help categorize a great number of unlabelled cloud images in order to alleviate climate scientists' work). Would anyone more familiar with the subject be willing to share some insights on that?\n\nThank you.",
      "votes": null
    },
    {
      "id": "601668",
      "postDate": "08/18/2019 02:20:33",
      "content": "<p>This is a really interesting question. Thanks for asking. I'll try to explain why we went for the segmentation masks instead of bounding boxes</p>\n\n<p>The reason we chose bounding boxes for the crowd-sourcing activity, where the labels were created, was to make it easy for the labelers. Drawing a rectangle is much easier than drawing a more detailed shape like a polygon. But, of course, the shapes we actually want to detect are not rectangular. In some first tests we saw that a segmentation algorithm was able to actually pick out the actual objects within the rectangles. Check out this picture from our paper, for example. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1178422%2F8cbd001a77e2c0c6157482f37e421a79%2FScreen%20Shot%202019-08-17%20at%2019.12.31.png?generation=1566094370827111&amp;alt=media\" alt=\"\"></p>\n\n<p>Then some practical reasons: Segmentation masks are also easier to score than bounding boxes. There you have to figure out how to deal with overlapping boxes, etc. This has been done in the literature, of course, but adds another layer of complexity. Further, segmentation masks are a little easier to use for any analysis we might do later. Bounding boxes require an additional processing step.</p>\n\n<p>Finally, in traditional bounding box problems like detecting people or cars it actually matters whether you draw two boxes for two adjacent objects (this would be correct) or just one box that contains both objects. For our scientific problem this is not the case. </p>\n\n<p>To summarize, the traditional advantages of bounding boxes are less important for our case, so the practical advantages and the ability to draw non-rectangular regions led us to the decision to use segmentation masks.</p>\n\n<p>I hope this makes sense. Let me know if you have any further questions.</p>",
      "rawMarkdown": "This is a really interesting question. Thanks for asking. I'll try to explain why we went for the segmentation masks instead of bounding boxes\n\nThe reason we chose bounding boxes for the crowd-sourcing activity, where the labels were created, was to make it easy for the labelers. Drawing a rectangle is much easier than drawing a more detailed shape like a polygon. But, of course, the shapes we actually want to detect are not rectangular. In some first tests we saw that a segmentation algorithm was able to actually pick out the actual objects within the rectangles. Check out this picture from our paper, for example. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1178422%2F8cbd001a77e2c0c6157482f37e421a79%2FScreen%20Shot%202019-08-17%20at%2019.12.31.png?generation=1566094370827111&amp;alt=media)\n\nThen some practical reasons: Segmentation masks are also easier to score than bounding boxes. There you have to figure out how to deal with overlapping boxes, etc. This has been done in the literature, of course, but adds another layer of complexity. Further, segmentation masks are a little easier to use for any analysis we might do later. Bounding boxes require an additional processing step.\n\nFinally, in traditional bounding box problems like detecting people or cars it actually matters whether you draw two boxes for two adjacent objects (this would be correct) or just one box that contains both objects. For our scientific problem this is not the case. \n\nTo summarize, the traditional advantages of bounding boxes are less important for our case, so the practical advantages and the ability to draw non-rectangular regions led us to the decision to use segmentation masks.\n\nI hope this makes sense. Let me know if you have any further questions.",
      "votes": null
    },
    {
      "id": "601672",
      "postDate": "08/18/2019 02:31:25",
      "content": "<p>Thank you for the quick response! Indeed, i think this answer a lot of my questions :) </p>",
      "rawMarkdown": "Thank you for the quick response! Indeed, i think this answer a lot of my questions :)",
      "votes": null
    },
    {
      "id": "601825",
      "postDate": "08/18/2019 07:41:33",
      "content": "<p>Wow, thank you so much for your feedback :) </p>",
      "rawMarkdown": "Wow, thank you so much for your feedback :)",
      "votes": null
    },
    {
      "id": "602347",
      "postDate": "08/19/2019 00:57:24",
      "content": "<p>Thanks for the detailed explanation.</p>",
      "rawMarkdown": "Thanks for the detailed explanation.",
      "votes": null
    },
    {
      "id": "602448",
      "postDate": "08/19/2019 04:30:34",
      "content": "<p>I also want to talk about this question. I don't have a lot of experience with instance segmentation, but it seems to me that even simple Unet gives masks which are finer than the original masks.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F727004%2F70f901092e595828620b604e97d71b55%2F2019-08-19_7-29-07.jpg?generation=1566188973244544&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I also want to talk about this question. I don't have a lot of experience with instance segmentation, but it seems to me that even simple Unet gives masks which are finer than the original masks.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F727004%2F70f901092e595828620b604e97d71b55%2F2019-08-19_7-29-07.jpg?generation=1566188973244544&amp;alt=media)",
      "votes": null
    },
    {
      "id": "602534",
      "postDate": "08/19/2019 06:54:53",
      "content": "<p>many thanks to both of you guys: <a href=\"/xhlulu\">@xhlulu</a> for raising these great questions and <a href=\"/raspstephan\">@raspstephan</a> for giving us a detailed explanation! :-)</p>",
      "rawMarkdown": "many thanks to both of you guys: @xhlulu for raising these great questions and @raspstephan for giving us a detailed explanation! :-)",
      "votes": null
    },
    {
      "id": "603607",
      "postDate": "08/20/2019 13:43:53",
      "content": "<p>Hey, cool visualizations. Yes, this was also our experience. It seems that despite the labels being rectangles, the U-Net seems to learn which features are actually interesting inside the rectangles by looking at a lot of labels. </p>\n\n<p>This is actually kind of cool because the U-Net labels often seem more accurate than the human labels!  As I mentioned below, this is one of the reasons why we went for segmentation rather than object detection. Let us know if you have any further questions.</p>",
      "rawMarkdown": "Hey, cool visualizations. Yes, this was also our experience. It seems that despite the labels being rectangles, the U-Net seems to learn which features are actually interesting inside the rectangles by looking at a lot of labels. \n\nThis is actually kind of cool because the U-Net labels often seem more accurate than the human labels!  As I mentioned below, this is one of the reasons why we went for segmentation rather than object detection. Let us know if you have any further questions.",
      "votes": null
    },
    {
      "id": "607429",
      "postDate": "08/25/2019 08:27:12",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.",
      "votes": null
    },
    {
      "id": "611967",
      "postDate": "08/29/2019 14:32:57",
      "content": "<p>Could you explain a little more on how these segmentations are going to help <code>\"to alleviate climate scientists' work\"</code>? \nReally curious about this part. </p>",
      "rawMarkdown": "Could you explain a little more on how these segmentations are going to help `\"to alleviate climate scientists' work\"`? \nReally curious about this part.",
      "votes": null
    },
    {
      "id": "620469",
      "postDate": "09/07/2019 14:37:01",
      "content": "<p>Thanks for the explanations! However, this leads to another question. How could non-rectangular masks give higher scores than the rectangular masks since all labels are rectangular? Should we resize the masks to rectangles like in this <a href=\"https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras\">kernel</a>?</p>\n\n<p>For example, assuming a model that generates extremely accurate segmentations that are exactly the contours of the clouds we want. When we calculating the Dice coefficient score, this model's results can probably be very low because it loses information from the rest parts of the rectangle label.</p>\n\n<p>Thank you very much.</p>",
      "rawMarkdown": "Thanks for the explanations! However, this leads to another question. How could non-rectangular masks give higher scores than the rectangular masks since all labels are rectangular? Should we resize the masks to rectangles like in this [kernel][1]?\n\nFor example, assuming a model that generates extremely accurate segmentations that are exactly the contours of the clouds we want. When we calculating the Dice coefficient score, this model's results can probably be very low because it loses information from the rest parts of the rectangle label.\n\nThank you very much.\n\n[1]: https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras",
      "votes": null
    },
    {
      "id": "622002",
      "postDate": "09/09/2019 06:38:48",
      "content": "<p>A bit confused on scoring for this challenge.   If I could create a model that did perfect bounding box (than make segments with the box) , like that shown for sugar in you b image I would get hopefully a perfect score for that test image.  </p>\n\n<p>If I create a model that does a terrific job of creating segments as shown in you figure c it seems to me that my score for the model that does a great job of segmentation will be much worse than a model that does a great bounding box.</p>\n\n<p>So it would seem that the winning model is one that best duplicates the bounding boxes that are used to score while the more useful and desired model that segments well can easily have a poorer score.</p>\n\n<p>My concerns go away if your scoring the test images on a \"true\" segmentation rather than just more bounding like boxes as seen in train.</p>",
      "rawMarkdown": "A bit confused on scoring for this challenge.   If I could create a model that did perfect bounding box (than make segments with the box) , like that shown for sugar in you b image I would get hopefully a perfect score for that test image.  \n\nIf I create a model that does a terrific job of creating segments as shown in you figure c it seems to me that my score for the model that does a great job of segmentation will be much worse than a model that does a great bounding box.\n\nSo it would seem that the winning model is one that best duplicates the bounding boxes that are used to score while the more useful and desired model that segments well can easily have a poorer score.\n\nMy concerns go away if your scoring the test images on a \"true\" segmentation rather than just more bounding like boxes as seen in train.",
      "votes": null
    },
    {
      "id": "622018",
      "postDate": "09/09/2019 07:09:05",
      "content": "<p>This is an interesting point and, as with any competition, understanding the score (which you do) is an important part of the competition. So I can't give you too much advice on this. I would like to point out two things though:</p>\n\n<ol>\n<li><p>Yes, the labels are made up of individual rectangles, but from several labelers. So often, the rectangles will overlap, so that the prediction is a more complex shape. Also, the black regions are omitted from the ground truth masks.</p></li>\n<li><p>In this competition, there is no objective ground truth. The cloud classes are subjective, i.e. every human thinks slightly differently about what is Sugar vs what is Gravel. This means that it is basically impossible to create a perfect algorithm. One consideration to make is whether using a detection algorithm that draws a box will actually lead to larger errors because of the inherent uncertainty of the data. I do not know the answer to this though. I guess figuring this out is part of the competition ;)</p></li>\n</ol>",
      "rawMarkdown": "This is an interesting point and, as with any competition, understanding the score (which you do) is an important part of the competition. So I can't give you too much advice on this. I would like to point out two things though:\n\n1. Yes, the labels are made up of individual rectangles, but from several labelers. So often, the rectangles will overlap, so that the prediction is a more complex shape. Also, the black regions are omitted from the ground truth masks.\n\n2. In this competition, there is no objective ground truth. The cloud classes are subjective, i.e. every human thinks slightly differently about what is Sugar vs what is Gravel. This means that it is basically impossible to create a perfect algorithm. One consideration to make is whether using a detection algorithm that draws a box will actually lead to larger errors because of the inherent uncertainty of the data. I do not know the answer to this though. I guess figuring this out is part of the competition ;)",
      "votes": null
    },
    {
      "id": "622019",
      "postDate": "09/09/2019 07:10:37",
      "content": "<p>Hi, good question! Someone else asked a related question (see above). My answer there should also apply to your question. </p>",
      "rawMarkdown": "Hi, good question! Someone else asked a related question (see above). My answer there should also apply to your question.",
      "votes": null
    },
    {
      "id": "624260",
      "postDate": "09/11/2019 21:44:57",
      "content": "<p>After having read this, I would definitely go for a rectangular boxing mask over a segmentation approach.\nI can not understand how a segmentation mask approach would perform better on the DICE calculation.</p>\n\n<p>Anyone who can convince elaborate on how a segmentation approachh would perform better on the evaluation metric if the subjective ground truths are boxes?</p>",
      "rawMarkdown": "After having read this, I would definitely go for a rectangular boxing mask over a segmentation approach.\nI can not understand how a segmentation mask approach would perform better on the DICE calculation.\n\nAnyone who can convince elaborate on how a segmentation approachh would perform better on the evaluation metric if the subjective ground truths are boxes?",
      "votes": null
    },
    {
      "id": "663757",
      "postDate": "11/02/2019 16:27:47",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Joining this late for practice purposes and I totally agree with you! I had my thoughts about this competition prior to joining and reading multiple posts they are getting confirmed.</p>\n\n<p>While the metric is how close your mask / segment is to \"ground truth\" (which we learned there isn't one), the \"ground truth masks\" themselves seem to be masks that are made from bounding boxes (because bounding boxes were easier to draw).</p>\n\n<p>So this preference of organizers to predicting segments rather than bounding boxes doesn't make sense to me when your training data and labels are not resembling it. Its like saying I would rather your models predict polar bears shapes, but all I have is data of polar bears with multiple bounding boxes from different labelers combined into one strange looking bounding box which was then converted into a mask because getting to this was easier… ?</p>\n\n<p>From there stems the so called post processing threads,</p>\n\n<p><a href=\"https://www.kaggle.com/ratthachat/cloud-convexhull-polygon-postprocessing-no-gpu/\">https://www.kaggle.com/ratthachat/cloud-convexhull-polygon-postprocessing-no-gpu/</a></p>\n\n<p>where people are building shapes from predicted segments to resemble squares, rectangles, etc. So trying to fix \"non square/rectangular mask\" into a \"square/rectangular hole\"…</p>\n\n<p>While organizers answer is well:</p>\n\n<p>1\nwe take multiple labels from different people and make more robust shapes than just one rectangle, by putting two rectangles/ squares together …\nand</p>\n\n<p>2\nblack regions are omitted from ground truths\nWell…</p>\n\n<p>1\nTo me this still seems like we are trying to predict rectangular / square shaped bounding boxes…\n2\n(This might actually help people).. once you have one square / rectangular shaped mask, take all pixel values with black values [0] - black pixels -- and break your square according to coordinates of these black value pixels…</p>\n\n<p>Then I saw people posting comments about trying bounding box models and actually not scoring too low with them -- -<a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/114863#latest-661443\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/114863#latest-661443</a></p>\n\n<p>So that's that. But hey, don't like it? Nobody is forcing you to participate, right?</p>\n\n<p>Perhaps a lot of time here will be spent fiddling around with outputs to make your results look better against this metric given \"no ground truth\" labels, which may not be the best use of time if you are trying to practice or learn. Just my observations, no ground truth to them either ;)</p>",
      "rawMarkdown": "pcjimmmy Joining this late for practice purposes and I totally agree with you! I had my thoughts about this competition prior to joining and reading multiple posts they are getting confirmed.\n\nWhile the metric is how close your mask / segment is to \"ground truth\" (which we learned there isn't one), the \"ground truth masks\" themselves seem to be masks that are made from bounding boxes (because bounding boxes were easier to draw).\n\nSo this preference of organizers to predicting segments rather than bounding boxes doesn't make sense to me when your training data and labels are not resembling it. Its like saying I would rather your models predict polar bears shapes, but all I have is data of polar bears with multiple bounding boxes from different labelers combined into one strange looking bounding box which was then converted into a mask because getting to this was easier… ?\n\nFrom there stems the so called post processing threads,\n\nhttps://www.kaggle.com/ratthachat/cloud-convexhull-polygon-postprocessing-no-gpu/\n\nwhere people are building shapes from predicted segments to resemble squares, rectangles, etc. So trying to fix \"non square/rectangular mask\" into a \"square/rectangular hole\"…\n\nWhile organizers answer is well:\n\n1\nwe take multiple labels from different people and make more robust shapes than just one rectangle, by putting two rectangles/ squares together …\nand\n\n2\nblack regions are omitted from ground truths\nWell…\n\n1\nTo me this still seems like we are trying to predict rectangular / square shaped bounding boxes…\n2\n(This might actually help people).. once you have one square / rectangular shaped mask, take all pixel values with black values [0] - black pixels -- and break your square according to coordinates of these black value pixels…\n\nThen I saw people posting comments about trying bounding box models and actually not scoring too low with them -- -https://www.kaggle.com/c/understanding_cloud_organization/discussion/114863#latest-661443\n\nSo that's that. But hey, don't like it? Nobody is forcing you to participate, right?\n\nPerhaps a lot of time here will be spent fiddling around with outputs to make your results look better against this metric given \"no ground truth\" labels, which may not be the best use of time if you are trying to practice or learn. Just my observations, no ground truth to them either ;)",
      "votes": null
    },
    {
      "id": "663768",
      "postDate": "11/02/2019 16:35:49",
      "content": "<p>However, I've never found the rectangular masks to be better than segmentation masks yet. I think the reason is the bounding boxes are over-confident about the area, which may result in too large or too small masks. Since the labels in this competition are very noisy and over-confidence may worsen the score, I suggest using hybrid-masks (e.g., some use bounding boxes or convex hull whilst others use orginal segmentation masks) based on some criterions.</p>",
      "rawMarkdown": "However, I've never found the rectangular masks to be better than segmentation masks yet. I think the reason is the bounding boxes are over-confident about the area, which may result in too large or too small masks. Since the labels in this competition are very noisy and over-confidence may worsen the score, I suggest using hybrid-masks (e.g., some use bounding boxes or convex hull whilst others use orginal segmentation masks) based on some criterions.",
      "votes": null
    },
    {
      "id": "663788",
      "postDate": "11/02/2019 17:04:37",
      "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a>  so go with segmentation then try post processing? are you including this logic while training your model or you still use only dice on predicted non-hybrid mask, trying to optimize that and then convert to hybrid?</p>",
      "rawMarkdown": "gogo827jz  so go with segmentation then try post processing? are you including this logic while training your model or you still use only dice on predicted non-hybrid mask, trying to optimize that and then convert to hybrid?",
      "votes": null
    },
    {
      "id": "663794",
      "postDate": "11/02/2019 17:11:02",
      "content": "<p>Train only segmentaton and do hybrid in post process</p>",
      "rawMarkdown": "Train only segmentaton and do hybrid in post process",
      "votes": null
    },
    {
      "id": "663799",
      "postDate": "11/02/2019 17:16:36",
      "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a>  cool are you using classifiers like Severstal competition. If class present probability &lt; X treshhold, predict 0 for Y mask class... ?</p>",
      "rawMarkdown": "gogo827jz  cool are you using classifiers like Severstal competition. If class present probability &lt; X treshhold, predict 0 for Y mask class... ?",
      "votes": null
    },
    {
      "id": "663809",
      "postDate": "11/02/2019 17:33:21",
      "content": "<p>Yes</p>",
      "rawMarkdown": "Yes",
      "votes": null
    },
    {
      "id": "663812",
      "postDate": "11/02/2019 17:37:37",
      "content": "<p>k, thanks! sooo much work =P</p>",
      "rawMarkdown": "k, thanks! sooo much work =P",
      "votes": null
    },
    {
      "id": "755031",
      "postDate": "02/24/2020 11:15:40",
      "content": "<p>nice pappy!!!</p>",
      "rawMarkdown": "nice pappy!!!",
      "votes": null
    },
    {
      "id": "992628",
      "postDate": "08/31/2020 09:39:49",
      "content": "<p>I like to Read your theory.</p>",
      "rawMarkdown": "I like to Read your theory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 992628,
      "author_name": "vatsalrakholiya",
      "author_url": "",
      "post_date": "08/31/2020 09:39:49",
      "content": "<p>I like to Read your theory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 601668,
      "author_name": "raspstephan",
      "author_url": "",
      "post_date": "08/18/2019 02:20:33",
      "content": "<p>This is a really interesting question. Thanks for asking. I'll try to explain why we went for the segmentation masks instead of bounding boxes</p>\n\n<p>The reason we chose bounding boxes for the crowd-sourcing activity, where the labels were created, was to make it easy for the labelers. Drawing a rectangle is much easier than drawing a more detailed shape like a polygon. But, of course, the shapes we actually want to detect are not rectangular. In some first tests we saw that a segmentation algorithm was able to actually pick out the actual objects within the rectangles. Check out this picture from our paper, for example. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1178422%2F8cbd001a77e2c0c6157482f37e421a79%2FScreen%20Shot%202019-08-17%20at%2019.12.31.png?generation=1566094370827111&amp;alt=media\" alt=\"\"></p>\n\n<p>Then some practical reasons: Segmentation masks are also easier to score than bounding boxes. There you have to figure out how to deal with overlapping boxes, etc. This has been done in the literature, of course, but adds another layer of complexity. Further, segmentation masks are a little easier to use for any analysis we might do later. Bounding boxes require an additional processing step.</p>\n\n<p>Finally, in traditional bounding box problems like detecting people or cars it actually matters whether you draw two boxes for two adjacent objects (this would be correct) or just one box that contains both objects. For our scientific problem this is not the case. </p>\n\n<p>To summarize, the traditional advantages of bounding boxes are less important for our case, so the practical advantages and the ability to draw non-rectangular regions led us to the decision to use segmentation masks.</p>\n\n<p>I hope this makes sense. Let me know if you have any further questions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 601672,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "08/18/2019 02:31:25",
          "content": "<p>Thank you for the quick response! Indeed, i think this answer a lot of my questions :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601825,
          "author_name": "jesucristo",
          "author_url": "",
          "post_date": "08/18/2019 07:41:33",
          "content": "<p>Wow, thank you so much for your feedback :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602347,
          "author_name": "grapestone5321",
          "author_url": "",
          "post_date": "08/19/2019 00:57:24",
          "content": "<p>Thanks for the detailed explanation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607429,
          "author_name": "chandraroy",
          "author_url": "",
          "post_date": "08/25/2019 08:27:12",
          "content": "<p>Thank you for sharing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 611967,
          "author_name": "rohitmidha23",
          "author_url": "",
          "post_date": "08/29/2019 14:32:57",
          "content": "<p>Could you explain a little more on how these segmentations are going to help <code>\"to alleviate climate scientists' work\"</code>? \nReally curious about this part. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622002,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "09/09/2019 06:38:48",
          "content": "<p>A bit confused on scoring for this challenge.   If I could create a model that did perfect bounding box (than make segments with the box) , like that shown for sugar in you b image I would get hopefully a perfect score for that test image.  </p>\n\n<p>If I create a model that does a terrific job of creating segments as shown in you figure c it seems to me that my score for the model that does a great job of segmentation will be much worse than a model that does a great bounding box.</p>\n\n<p>So it would seem that the winning model is one that best duplicates the bounding boxes that are used to score while the more useful and desired model that segments well can easily have a poorer score.</p>\n\n<p>My concerns go away if your scoring the test images on a \"true\" segmentation rather than just more bounding like boxes as seen in train.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622018,
          "author_name": "raspstephan",
          "author_url": "",
          "post_date": "09/09/2019 07:09:05",
          "content": "<p>This is an interesting point and, as with any competition, understanding the score (which you do) is an important part of the competition. So I can't give you too much advice on this. I would like to point out two things though:</p>\n\n<ol>\n<li><p>Yes, the labels are made up of individual rectangles, but from several labelers. So often, the rectangles will overlap, so that the prediction is a more complex shape. Also, the black regions are omitted from the ground truth masks.</p></li>\n<li><p>In this competition, there is no objective ground truth. The cloud classes are subjective, i.e. every human thinks slightly differently about what is Sugar vs what is Gravel. This means that it is basically impossible to create a perfect algorithm. One consideration to make is whether using a detection algorithm that draws a box will actually lead to larger errors because of the inherent uncertainty of the data. I do not know the answer to this though. I guess figuring this out is part of the competition ;)</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663757,
          "author_name": "yevguyduy",
          "author_url": "",
          "post_date": "11/02/2019 16:27:47",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Joining this late for practice purposes and I totally agree with you! I had my thoughts about this competition prior to joining and reading multiple posts they are getting confirmed.</p>\n\n<p>While the metric is how close your mask / segment is to \"ground truth\" (which we learned there isn't one), the \"ground truth masks\" themselves seem to be masks that are made from bounding boxes (because bounding boxes were easier to draw).</p>\n\n<p>So this preference of organizers to predicting segments rather than bounding boxes doesn't make sense to me when your training data and labels are not resembling it. Its like saying I would rather your models predict polar bears shapes, but all I have is data of polar bears with multiple bounding boxes from different labelers combined into one strange looking bounding box which was then converted into a mask because getting to this was easier… ?</p>\n\n<p>From there stems the so called post processing threads,</p>\n\n<p><a href=\"https://www.kaggle.com/ratthachat/cloud-convexhull-polygon-postprocessing-no-gpu/\">https://www.kaggle.com/ratthachat/cloud-convexhull-polygon-postprocessing-no-gpu/</a></p>\n\n<p>where people are building shapes from predicted segments to resemble squares, rectangles, etc. So trying to fix \"non square/rectangular mask\" into a \"square/rectangular hole\"…</p>\n\n<p>While organizers answer is well:</p>\n\n<p>1\nwe take multiple labels from different people and make more robust shapes than just one rectangle, by putting two rectangles/ squares together …\nand</p>\n\n<p>2\nblack regions are omitted from ground truths\nWell…</p>\n\n<p>1\nTo me this still seems like we are trying to predict rectangular / square shaped bounding boxes…\n2\n(This might actually help people).. once you have one square / rectangular shaped mask, take all pixel values with black values [0] - black pixels -- and break your square according to coordinates of these black value pixels…</p>\n\n<p>Then I saw people posting comments about trying bounding box models and actually not scoring too low with them -- -<a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/114863#latest-661443\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/114863#latest-661443</a></p>\n\n<p>So that's that. But hey, don't like it? Nobody is forcing you to participate, right?</p>\n\n<p>Perhaps a lot of time here will be spent fiddling around with outputs to make your results look better against this metric given \"no ground truth\" labels, which may not be the best use of time if you are trying to practice or learn. Just my observations, no ground truth to them either ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 602448,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "08/19/2019 04:30:34",
      "content": "<p>I also want to talk about this question. I don't have a lot of experience with instance segmentation, but it seems to me that even simple Unet gives masks which are finer than the original masks.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F727004%2F70f901092e595828620b604e97d71b55%2F2019-08-19_7-29-07.jpg?generation=1566188973244544&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 603607,
          "author_name": "raspstephan",
          "author_url": "",
          "post_date": "08/20/2019 13:43:53",
          "content": "<p>Hey, cool visualizations. Yes, this was also our experience. It seems that despite the labels being rectangles, the U-Net seems to learn which features are actually interesting inside the rectangles by looking at a lot of labels. </p>\n\n<p>This is actually kind of cool because the U-Net labels often seem more accurate than the human labels!  As I mentioned below, this is one of the reasons why we went for segmentation rather than object detection. Let us know if you have any further questions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620469,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/07/2019 14:37:01",
          "content": "<p>Thanks for the explanations! However, this leads to another question. How could non-rectangular masks give higher scores than the rectangular masks since all labels are rectangular? Should we resize the masks to rectangles like in this <a href=\"https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras\">kernel</a>?</p>\n\n<p>For example, assuming a model that generates extremely accurate segmentations that are exactly the contours of the clouds we want. When we calculating the Dice coefficient score, this model's results can probably be very low because it loses information from the rest parts of the rectangle label.</p>\n\n<p>Thank you very much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622019,
          "author_name": "raspstephan",
          "author_url": "",
          "post_date": "09/09/2019 07:10:37",
          "content": "<p>Hi, good question! Someone else asked a related question (see above). My answer there should also apply to your question. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 602534,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "08/19/2019 06:54:53",
      "content": "<p>many thanks to both of you guys: <a href=\"/xhlulu\">@xhlulu</a> for raising these great questions and <a href=\"/raspstephan\">@raspstephan</a> for giving us a detailed explanation! :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 624260,
      "author_name": "",
      "author_url": "",
      "post_date": "09/11/2019 21:44:57",
      "content": "<p>After having read this, I would definitely go for a rectangular boxing mask over a segmentation approach.\nI can not understand how a segmentation mask approach would perform better on the DICE calculation.</p>\n\n<p>Anyone who can convince elaborate on how a segmentation approachh would perform better on the evaluation metric if the subjective ground truths are boxes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 663768,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "11/02/2019 16:35:49",
          "content": "<p>However, I've never found the rectangular masks to be better than segmentation masks yet. I think the reason is the bounding boxes are over-confident about the area, which may result in too large or too small masks. Since the labels in this competition are very noisy and over-confidence may worsen the score, I suggest using hybrid-masks (e.g., some use bounding boxes or convex hull whilst others use orginal segmentation masks) based on some criterions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663788,
          "author_name": "yevguyduy",
          "author_url": "",
          "post_date": "11/02/2019 17:04:37",
          "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a>  so go with segmentation then try post processing? are you including this logic while training your model or you still use only dice on predicted non-hybrid mask, trying to optimize that and then convert to hybrid?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663794,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "11/02/2019 17:11:02",
          "content": "<p>Train only segmentaton and do hybrid in post process</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663799,
          "author_name": "yevguyduy",
          "author_url": "",
          "post_date": "11/02/2019 17:16:36",
          "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a>  cool are you using classifiers like Severstal competition. If class present probability &lt; X treshhold, predict 0 for Y mask class... ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663809,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "11/02/2019 17:33:21",
          "content": "<p>Yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663812,
          "author_name": "yevguyduy",
          "author_url": "",
          "post_date": "11/02/2019 17:37:37",
          "content": "<p>k, thanks! sooo much work =P</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755031,
      "author_name": "mariavoytyuk",
      "author_url": "",
      "post_date": "02/24/2020 11:15:40",
      "content": "<p>nice pappy!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "601633": "I'm sure that the organizers and Kaggle have a better idea of the subject than me, but one thing that confused me was: what are the advantages of building an instance segmentation model for this competition compared to an object detection (with bounding boxes) model? \n\nLooking at the [official starter kernel](https://www.kaggle.com/inversion/run-length-decoding-and-flower-submission), I have the impression the masks can be represented as bounding boxes. I am also not sure I understand the advantages of segmentation masks for the end use case (help categorize a great number of unlabelled cloud images in order to alleviate climate scientists' work). Would anyone more familiar with the subject be willing to share some insights on that?\n\nThank you.",
    "601668": "This is a really interesting question. Thanks for asking. I'll try to explain why we went for the segmentation masks instead of bounding boxes\n\nThe reason we chose bounding boxes for the crowd-sourcing activity, where the labels were created, was to make it easy for the labelers. Drawing a rectangle is much easier than drawing a more detailed shape like a polygon. But, of course, the shapes we actually want to detect are not rectangular. In some first tests we saw that a segmentation algorithm was able to actually pick out the actual objects within the rectangles. Check out this picture from our paper, for example. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1178422%2F8cbd001a77e2c0c6157482f37e421a79%2FScreen%20Shot%202019-08-17%20at%2019.12.31.png?generation=1566094370827111&amp;alt=media)\n\nThen some practical reasons: Segmentation masks are also easier to score than bounding boxes. There you have to figure out how to deal with overlapping boxes, etc. This has been done in the literature, of course, but adds another layer of complexity. Further, segmentation masks are a little easier to use for any analysis we might do later. Bounding boxes require an additional processing step.\n\nFinally, in traditional bounding box problems like detecting people or cars it actually matters whether you draw two boxes for two adjacent objects (this would be correct) or just one box that contains both objects. For our scientific problem this is not the case. \n\nTo summarize, the traditional advantages of bounding boxes are less important for our case, so the practical advantages and the ability to draw non-rectangular regions led us to the decision to use segmentation masks.\n\nI hope this makes sense. Let me know if you have any further questions.",
    "601672": "Thank you for the quick response! Indeed, i think this answer a lot of my questions :)",
    "601825": "Wow, thank you so much for your feedback :)",
    "602347": "Thanks for the detailed explanation.",
    "602448": "I also want to talk about this question. I don't have a lot of experience with instance segmentation, but it seems to me that even simple Unet gives masks which are finer than the original masks.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F727004%2F70f901092e595828620b604e97d71b55%2F2019-08-19_7-29-07.jpg?generation=1566188973244544&amp;alt=media)",
    "602534": "many thanks to both of you guys: @xhlulu for raising these great questions and @raspstephan for giving us a detailed explanation! :-)",
    "603607": "Hey, cool visualizations. Yes, this was also our experience. It seems that despite the labels being rectangles, the U-Net seems to learn which features are actually interesting inside the rectangles by looking at a lot of labels. \n\nThis is actually kind of cool because the U-Net labels often seem more accurate than the human labels!  As I mentioned below, this is one of the reasons why we went for segmentation rather than object detection. Let us know if you have any further questions.",
    "607429": "Thank you for sharing.",
    "611967": "Could you explain a little more on how these segmentations are going to help `\"to alleviate climate scientists' work\"`? \nReally curious about this part.",
    "620469": "Thanks for the explanations! However, this leads to another question. How could non-rectangular masks give higher scores than the rectangular masks since all labels are rectangular? Should we resize the masks to rectangles like in this [kernel][1]?\n\nFor example, assuming a model that generates extremely accurate segmentations that are exactly the contours of the clouds we want. When we calculating the Dice coefficient score, this model's results can probably be very low because it loses information from the rest parts of the rectangle label.\n\nThank you very much.\n\n[1]: https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras",
    "622002": "A bit confused on scoring for this challenge.   If I could create a model that did perfect bounding box (than make segments with the box) , like that shown for sugar in you b image I would get hopefully a perfect score for that test image.  \n\nIf I create a model that does a terrific job of creating segments as shown in you figure c it seems to me that my score for the model that does a great job of segmentation will be much worse than a model that does a great bounding box.\n\nSo it would seem that the winning model is one that best duplicates the bounding boxes that are used to score while the more useful and desired model that segments well can easily have a poorer score.\n\nMy concerns go away if your scoring the test images on a \"true\" segmentation rather than just more bounding like boxes as seen in train.",
    "622018": "This is an interesting point and, as with any competition, understanding the score (which you do) is an important part of the competition. So I can't give you too much advice on this. I would like to point out two things though:\n\n1. Yes, the labels are made up of individual rectangles, but from several labelers. So often, the rectangles will overlap, so that the prediction is a more complex shape. Also, the black regions are omitted from the ground truth masks.\n\n2. In this competition, there is no objective ground truth. The cloud classes are subjective, i.e. every human thinks slightly differently about what is Sugar vs what is Gravel. This means that it is basically impossible to create a perfect algorithm. One consideration to make is whether using a detection algorithm that draws a box will actually lead to larger errors because of the inherent uncertainty of the data. I do not know the answer to this though. I guess figuring this out is part of the competition ;)",
    "622019": "Hi, good question! Someone else asked a related question (see above). My answer there should also apply to your question.",
    "624260": "After having read this, I would definitely go for a rectangular boxing mask over a segmentation approach.\nI can not understand how a segmentation mask approach would perform better on the DICE calculation.\n\nAnyone who can convince elaborate on how a segmentation approachh would perform better on the evaluation metric if the subjective ground truths are boxes?",
    "663757": "pcjimmmy Joining this late for practice purposes and I totally agree with you! I had my thoughts about this competition prior to joining and reading multiple posts they are getting confirmed.\n\nWhile the metric is how close your mask / segment is to \"ground truth\" (which we learned there isn't one), the \"ground truth masks\" themselves seem to be masks that are made from bounding boxes (because bounding boxes were easier to draw).\n\nSo this preference of organizers to predicting segments rather than bounding boxes doesn't make sense to me when your training data and labels are not resembling it. Its like saying I would rather your models predict polar bears shapes, but all I have is data of polar bears with multiple bounding boxes from different labelers combined into one strange looking bounding box which was then converted into a mask because getting to this was easier… ?\n\nFrom there stems the so called post processing threads,\n\nhttps://www.kaggle.com/ratthachat/cloud-convexhull-polygon-postprocessing-no-gpu/\n\nwhere people are building shapes from predicted segments to resemble squares, rectangles, etc. So trying to fix \"non square/rectangular mask\" into a \"square/rectangular hole\"…\n\nWhile organizers answer is well:\n\n1\nwe take multiple labels from different people and make more robust shapes than just one rectangle, by putting two rectangles/ squares together …\nand\n\n2\nblack regions are omitted from ground truths\nWell…\n\n1\nTo me this still seems like we are trying to predict rectangular / square shaped bounding boxes…\n2\n(This might actually help people).. once you have one square / rectangular shaped mask, take all pixel values with black values [0] - black pixels -- and break your square according to coordinates of these black value pixels…\n\nThen I saw people posting comments about trying bounding box models and actually not scoring too low with them -- -https://www.kaggle.com/c/understanding_cloud_organization/discussion/114863#latest-661443\n\nSo that's that. But hey, don't like it? Nobody is forcing you to participate, right?\n\nPerhaps a lot of time here will be spent fiddling around with outputs to make your results look better against this metric given \"no ground truth\" labels, which may not be the best use of time if you are trying to practice or learn. Just my observations, no ground truth to them either ;)",
    "663768": "However, I've never found the rectangular masks to be better than segmentation masks yet. I think the reason is the bounding boxes are over-confident about the area, which may result in too large or too small masks. Since the labels in this competition are very noisy and over-confidence may worsen the score, I suggest using hybrid-masks (e.g., some use bounding boxes or convex hull whilst others use orginal segmentation masks) based on some criterions.",
    "663788": "gogo827jz  so go with segmentation then try post processing? are you including this logic while training your model or you still use only dice on predicted non-hybrid mask, trying to optimize that and then convert to hybrid?",
    "663794": "Train only segmentaton and do hybrid in post process",
    "663799": "gogo827jz  cool are you using classifiers like Severstal competition. If class present probability &lt; X treshhold, predict 0 for Y mask class... ?",
    "663809": "Yes",
    "663812": "k, thanks! sooo much work =P",
    "755031": "nice pappy!!!",
    "992628": "I like to Read your theory."
  },
  "source": "meta"
}