{
  "id": 91536,
  "title": "Which approach works best?",
  "url": "/competitions/imaterialist-fashion-2019-FGVC6/discussion/91536",
  "author_name": "Konstantin Lopukhin",
  "post_date": "2019-05-06T08:08:00.796000",
  "votes": 19,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I'm using Mask-RCNN so far (<a href=\"https://github.com/facebookresearch/maskrcnn-benchmark\">https://github.com/facebookresearch/maskrcnn-benchmark</a>) and it seems to work OK, and the code is quite easy to modify. Looks like current top-2 are also using Mask-RCNN?\nAnother possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here). And maybe also have a separate classification step to determine attributes.\nMain difference is that for Mask-RCNN, segmentation and classification are more explicitly separated, while for pure segmentation tasks, they are performed jointly. Although it seems that Mask-RCNN is likely to have worse quality masks, due to fixed per-object mask size and resampling.\nAlso attributes seem really hard to predict to me so far.</p>\n\n<p>What do you think?</p>",
  "messages": [
    {
      "id": 527746,
      "postDate": "2019-05-06T08:08:00.797Z",
      "content": "<p>I'm using Mask-RCNN so far (<a href=\"https://github.com/facebookresearch/maskrcnn-benchmark\">https://github.com/facebookresearch/maskrcnn-benchmark</a>) and it seems to work OK, and the code is quite easy to modify. Looks like current top-2 are also using Mask-RCNN?\nAnother possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here). And maybe also have a separate classification step to determine attributes.\nMain difference is that for Mask-RCNN, segmentation and classification are more explicitly separated, while for pure segmentation tasks, they are performed jointly. Although it seems that Mask-RCNN is likely to have worse quality masks, due to fixed per-object mask size and resampling.\nAlso attributes seem really hard to predict to me so far.</p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "I'm using Mask-RCNN so far (https://github.com/facebookresearch/maskrcnn-benchmark) and it seems to work OK, and the code is quite easy to modify. Looks like current top-2 are also using Mask-RCNN?\nAnother possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here). And maybe also have a separate classification step to determine attributes.\nMain difference is that for Mask-RCNN, segmentation and classification are more explicitly separated, while for pure segmentation tasks, they are performed jointly. Although it seems that Mask-RCNN is likely to have worse quality masks, due to fixed per-object mask size and resampling.\nAlso attributes seem really hard to predict to me so far.\n\nWhat do you think?",
      "votes": 19
    },
    {
      "id": 528027,
      "postDate": "2019-05-06T22:21:38.307Z",
      "content": "<p>Are you guys stacking the masks such that all masks associated with one image id are together? So the number of images and masks = unique image ids? Or take each row in train.csv as one image-mask pair so have number of samples = rows?</p>",
      "rawMarkdown": "Are you guys stacking the masks such that all masks associated with one image id are together? So the number of images and masks = unique image ids? Or take each row in train.csv as one image-mask pair so have number of samples = rows?",
      "votes": 1,
      "replies": [
        {
          "id": 528095,
          "postDate": "2019-05-07T04:28:37.837Z",
          "content": "<p>I think you need to take each row in train.csv as one image-mask pair so have number of samples = rows.</p>",
          "rawMarkdown": "I think you need to take each row in train.csv as one image-mask pair so have number of samples = rows.",
          "votes": 1
        },
        {
          "id": 528145,
          "postDate": "2019-05-07T06:39:06.120Z",
          "content": "<p>I don't think it's possible to stack the masks because some class ID's (not attributes) overlap - eg. segmentation of the sleeve overlaps completely with that of the jacket of the first training image.</p>\n\n<p>So training a model with different image-mask pairs is essentially training to classify a pixel as classID #n or not for each training example. How would inference look like then? You wouldn't be able to get all segmented classes through one forward pass right? </p>",
          "rawMarkdown": "I don't think it's possible to stack the masks because some class ID's (not attributes) overlap - eg. segmentation of the sleeve overlaps completely with that of the jacket of the first training image.\n\nSo training a model with different image-mask pairs is essentially training to classify a pixel as classID #n or not for each training example. How would inference look like then? You wouldn't be able to get all segmented classes through one forward pass right? "
        },
        {
          "id": 528867,
          "postDate": "2019-05-08T19:10:02.590Z",
          "content": "<blockquote>\n  <p>I don't think it's possible to stack the masks because some class ID's (not attributes) overlap - eg. segmentation of the sleeve overlaps completely with that of the jacket of the first training image.</p>\n</blockquote>\n\n<p>I agree that this makes things more complicated, but it's still possible to use multi-class segmentation here, non-linearity at the end would be sigmoid instead of softmax, and target mask would have number of channels equal to number of categories (as well as output mask). Predicted masks would be obtained by applying threshold to each output channel separately. These masks still would need to be divided into instances.</p>",
          "rawMarkdown": "&gt; I don't think it's possible to stack the masks because some class ID's (not attributes) overlap - eg. segmentation of the sleeve overlaps completely with that of the jacket of the first training image.\n\nI agree that this makes things more complicated, but it's still possible to use multi-class segmentation here, non-linearity at the end would be sigmoid instead of softmax, and target mask would have number of channels equal to number of categories (as well as output mask). Predicted masks would be obtained by applying threshold to each output channel separately. These masks still would need to be divided into instances."
        },
        {
          "id": 530244,
          "postDate": "2019-05-12T08:59:04.633Z",
          "content": "<p>That's interesting. Yeah, you're right. With this method, you'd be going from a (n x m x 3) input to  (n x m x num_classes) output</p>",
          "rawMarkdown": "That's interesting. Yeah, you're right. With this method, you'd be going from a (n x m x 3) input to  (n x m x num_classes) output",
          "votes": 1
        }
      ]
    },
    {
      "id": 534257,
      "postDate": "2019-05-21T01:57:21.353Z",
      "content": "<p>Thanks for your idea, I have tried mask-rcnn，then I have a baseline now. I'm confused about the direction of optimization, specifically about how to predict the attributes, or  classificate Finer-grained，can you have a share about it ?</p>",
      "rawMarkdown": "Thanks for your idea, I have tried mask-rcnn，then I have a baseline now. I'm confused about the direction of optimization, specifically about how to predict the attributes, or  classificate Finer-grained，can you have a share about it ?",
      "votes": 2
    },
    {
      "id": 528848,
      "postDate": "2019-05-08T17:59:40.633Z",
      "content": "<p>I'm also using Mask-RCNN, but in Keras. Still working on a trainable model adapted from <a href=\"https://github.com/matterport/Mask_RCNN\">Matterport's Mask_RCNN</a>.</p>",
      "rawMarkdown": "I'm also using Mask-RCNN, but in Keras. Still working on a trainable model adapted from [Matterport's Mask_RCNN](https://github.com/matterport/Mask_RCNN).",
      "votes": 2
    },
    {
      "id": 534909,
      "postDate": "2019-05-22T02:53:20.347Z",
      "content": "<p>How did you generate an instance?Just each row generate an instance ,or generate a mask for each class firstly, and then find  the envelope in the figure？(cv2.findContours), thank you for your sharing！</p>",
      "rawMarkdown": "How did you generate an instance?Just each row generate an instance ,or generate a mask for each class firstly, and then find  the envelope in the figure？(cv2.findContours), thank you for your sharing！",
      "replies": [
        {
          "id": 535062,
          "postDate": "2019-05-22T08:27:01.367Z",
          "content": "<p>It's one instance from one row.</p>",
          "rawMarkdown": "It's one instance from one row."
        }
      ]
    },
    {
      "id": 534899,
      "postDate": "2019-05-22T02:02:01.110Z",
      "content": "<p>Resize 512*512 this process loses much precision，I wonder if there is any need to pay attention to post-processing.</p>",
      "rawMarkdown": "Resize 512*512 this process loses much precision，I wonder if there is any need to pay attention to post-processing."
    },
    {
      "id": 534258,
      "postDate": "2019-05-21T01:58:15.980Z",
      "content": "<p>test</p>",
      "rawMarkdown": "test"
    },
    {
      "id": 534111,
      "postDate": "2019-05-20T17:53:44.450Z",
      "content": "<p>how much time does mask rcnn take for you? its taking nearly 30hrs for one epoch on my gpu. :/</p>",
      "rawMarkdown": "how much time does mask rcnn take for you? its taking nearly 30hrs for one epoch on my gpu. :/",
      "replies": [
        {
          "id": 534118,
          "postDate": "2019-05-20T18:01:09.913Z",
          "content": "<p>For me it's around 6 hours.</p>",
          "rawMarkdown": "For me it's around 6 hours."
        },
        {
          "id": 534152,
          "postDate": "2019-05-20T19:43:48.550Z",
          "content": "<p>It's similar for us, 30h per epoch sounds like a lot even for mask-rcnn</p>",
          "rawMarkdown": "It's similar for us, 30h per epoch sounds like a lot even for mask-rcnn"
        },
        {
          "id": 534163,
          "postDate": "2019-05-20T20:18:12.837Z",
          "content": "<p>what kind of gpu do you use?</p>",
          "rawMarkdown": "what kind of gpu do you use?"
        },
        {
          "id": 534164,
          "postDate": "2019-05-20T20:19:26.907Z",
          "content": "<p>8*1080Ti</p>",
          "rawMarkdown": "8*1080Ti",
          "votes": 5
        },
        {
          "id": 547636,
          "postDate": "2019-06-08T02:23:30.153Z",
          "content": "<p>How many images do you fit on one 1080TI? I only fit one, but I feel like I could fit 2 if I optimised some stuff.</p>",
          "rawMarkdown": "How many images do you fit on one 1080TI? I only fit one, but I feel like I could fit 2 if I optimised some stuff.",
          "votes": 1
        }
      ]
    },
    {
      "id": 527989,
      "postDate": "2019-05-06T18:52:01.947Z",
      "content": "<p>&gt; Another possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here)</p>\n\n<p>This part is not easy: e.g., how do you differentiate between two shoes and two patches of one shoe (where occlusion occurs in between)? cv2.connectedComponents cannot solve this problem directly.</p>",
      "rawMarkdown": "&gt; Another possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here)\n\nThis part is not easy: e.g., how do you differentiate between two shoes and two patches of one shoe (where occlusion occurs in between)? cv2.connectedComponents cannot solve this problem directly.",
      "replies": [
        {
          "id": 528013,
          "postDate": "2019-05-06T21:09:51.377Z",
          "content": "<p>I agree, \"easy\" is not the right word here, but it's still possible, for example in 2018 Data Science Bow the task was to separate cells which were really close, and 1st place used watershed to separate instances: <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741\">https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741</a></p>",
          "rawMarkdown": "I agree, \"easy\" is not the right word here, but it's still possible, for example in 2018 Data Science Bow the task was to separate cells which were really close, and 1st place used watershed to separate instances: https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741",
          "votes": 3
        }
      ]
    },
    {
      "id": 540227,
      "postDate": "2019-05-31T07:13:49.653Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 528027,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-05-06T22:21:38.307000",
      "content": "<p>Are you guys stacking the masks such that all masks associated with one image id are together? So the number of images and masks = unique image ids? Or take each row in train.csv as one image-mask pair so have number of samples = rows?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 528095,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-05-07T04:28:37.837000",
          "content": "<p>I think you need to take each row in train.csv as one image-mask pair so have number of samples = rows.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 528145,
          "author_name": "Sidhant",
          "author_url": "",
          "post_date": "2019-05-07T06:39:06.120000",
          "content": "<p>I don't think it's possible to stack the masks because some class ID's (not attributes) overlap - eg. segmentation of the sleeve overlaps completely with that of the jacket of the first training image.</p>\n\n<p>So training a model with different image-mask pairs is essentially training to classify a pixel as classID #n or not for each training example. How would inference look like then? You wouldn't be able to get all segmented classes through one forward pass right? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 528867,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-05-08T19:10:02.590000",
          "content": "<blockquote>\n  <p>I don't think it's possible to stack the masks because some class ID's (not attributes) overlap - eg. segmentation of the sleeve overlaps completely with that of the jacket of the first training image.</p>\n</blockquote>\n\n<p>I agree that this makes things more complicated, but it's still possible to use multi-class segmentation here, non-linearity at the end would be sigmoid instead of softmax, and target mask would have number of channels equal to number of categories (as well as output mask). Predicted masks would be obtained by applying threshold to each output channel separately. These masks still would need to be divided into instances.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 530244,
          "author_name": "Sidhant",
          "author_url": "",
          "post_date": "2019-05-12T08:59:04.633000",
          "content": "<p>That's interesting. Yeah, you're right. With this method, you'd be going from a (n x m x 3) input to  (n x m x num_classes) output</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 534257,
      "author_name": "sunxiaotao",
      "author_url": "",
      "post_date": "2019-05-21T01:57:21.353000",
      "content": "<p>Thanks for your idea, I have tried mask-rcnn，then I have a baseline now. I'm confused about the direction of optimization, specifically about how to predict the attributes, or  classificate Finer-grained，can you have a share about it ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 528848,
      "author_name": "JohnM",
      "author_url": "",
      "post_date": "2019-05-08T17:59:40.633000",
      "content": "<p>I'm also using Mask-RCNN, but in Keras. Still working on a trainable model adapted from <a href=\"https://github.com/matterport/Mask_RCNN\">Matterport's Mask_RCNN</a>.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 534909,
      "author_name": "donglee",
      "author_url": "",
      "post_date": "2019-05-22T02:53:20.347000",
      "content": "<p>How did you generate an instance?Just each row generate an instance ,or generate a mask for each class firstly, and then find  the envelope in the figure？(cv2.findContours), thank you for your sharing！</p>",
      "votes": 0,
      "replies": [
        {
          "id": 535062,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-05-22T08:27:01.367000",
          "content": "<p>It's one instance from one row.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 534899,
      "author_name": "donglee",
      "author_url": "",
      "post_date": "2019-05-22T02:02:01.110000",
      "content": "<p>Resize 512*512 this process loses much precision，I wonder if there is any need to pay attention to post-processing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 534258,
      "author_name": "sunxiaotao",
      "author_url": "",
      "post_date": "2019-05-21T01:58:15.980000",
      "content": "<p>test</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 534111,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2019-05-20T17:53:44.450000",
      "content": "<p>how much time does mask rcnn take for you? its taking nearly 30hrs for one epoch on my gpu. :/</p>",
      "votes": 0,
      "replies": [
        {
          "id": 534118,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2019-05-20T18:01:09.913000",
          "content": "<p>For me it's around 6 hours.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534152,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-05-20T19:43:48.550000",
          "content": "<p>It's similar for us, 30h per epoch sounds like a lot even for mask-rcnn</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534163,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-05-20T20:18:12.837000",
          "content": "<p>what kind of gpu do you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534164,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2019-05-20T20:19:26.907000",
          "content": "<p>8*1080Ti</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 547636,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-06-08T02:23:30.153000",
          "content": "<p>How many images do you fit on one 1080TI? I only fit one, but I feel like I could fit 2 if I optimised some stuff.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 527989,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2019-05-06T18:52:01.947000",
      "content": "<p>&gt; Another possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here)</p>\n\n<p>This part is not easy: e.g., how do you differentiate between two shoes and two patches of one shoe (where occlusion occurs in between)? cv2.connectedComponents cannot solve this problem directly.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 528013,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-05-06T21:09:51.377000",
          "content": "<p>I agree, \"easy\" is not the right word here, but it's still possible, for example in 2018 Data Science Bow the task was to separate cells which were really close, and 1st place used watershed to separate instances: <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741\">https://www.kaggle.com/c/data-science-bowl-2018/discussion/54741</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 540227,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-31T07:13:49.653000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "527746": "I'm using Mask-RCNN so far (https://github.com/facebookresearch/maskrcnn-benchmark) and it seems to work OK, and the code is quite easy to modify. Looks like current top-2 are also using Mask-RCNN?\nAnother possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here). And maybe also have a separate classification step to determine attributes.\nMain difference is that for Mask-RCNN, segmentation and classification are more explicitly separated, while for pure segmentation tasks, they are performed jointly. Although it seems that Mask-RCNN is likely to have worse quality masks, due to fixed per-object mask size and resampling.\nAlso attributes seem really hard to predict to me so far.\n\nWhat do you think?",
    "528027": "Are you guys stacking the masks such that all masks associated with one image id are together? So the number of images and masks = unique image ids? Or take each row in train.csv as one image-mask pair so have number of samples = rows?",
    "534257": "Thanks for your idea, I have tried mask-rcnn，then I have a baseline now. I'm confused about the direction of optimization, specifically about how to predict the attributes, or  classificate Finer-grained，can you have a share about it ?",
    "528848": "I'm also using Mask-RCNN, but in Keras. Still working on a trainable model adapted from [Matterport's Mask_RCNN](https://github.com/matterport/Mask_RCNN).",
    "534909": "How did you generate an instance?Just each row generate an instance ,or generate a mask for each class firstly, and then find  the envelope in the figure？(cv2.findContours), thank you for your sharing！",
    "534899": "Resize 512*512 this process loses much precision，I wonder if there is any need to pay attention to post-processing.",
    "534258": "test",
    "534111": "how much time does mask rcnn take for you? its taking nearly 30hrs for one epoch on my gpu. :/",
    "527989": "&gt; Another possible approach would be to do segmentation with something UNet-alike, and segment into instances (should be easy here)\n\nThis part is not easy: e.g., how do you differentiate between two shoes and two patches of one shoe (where occlusion occurs in between)? cv2.connectedComponents cannot solve this problem directly.",
    "540227": ""
  }
}