{
  "id": 95257,
  "title": "Did anyone use U-net in this competition?",
  "url": "/competitions/imaterialist-fashion-2019-FGVC6/discussion/95257",
  "author_name": "",
  "post_date": "2019-06-11T04:57:36.647495300Z",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>As the competition is over, I think it's the time to ask this question.\nI tried to use U-net in this competition, since I used it to do binary classification and segmentation before. However I faced a lot of difficulties to do a multiple classification and segmentation. \nMy solution is to transfer the input mask into a [-1, 224, 224, 47] np.array, code like this:  <code>sub_mask = cv2.resize(sub_mask, (SIZE, SIZE), interpolation=cv2.INTER_NEAREST)\n                    mask[:, :, int(label)] = sub_mask</code>, where 47 means 47 types(including the background), then when doing inference, I don't know how to choose the channel, and actually \nI think the channel is not matching the label.\nSo, as a beginner, I am really curious if anyone who used U-net in this competition and realized the multiple classification and segmentation. \nI will appreciate it if you can share the code or even just give some hints. Thank you.</p>\n\n<p>Additional: I'm using keras and tf.</p>",
  "messages": [
    {
      "id": "549851",
      "postDate": "06/11/2019 04:57:36",
      "content": "<p>As the competition is over, I think it's the time to ask this question.\nI tried to use U-net in this competition, since I used it to do binary classification and segmentation before. However I faced a lot of difficulties to do a multiple classification and segmentation. \nMy solution is to transfer the input mask into a [-1, 224, 224, 47] np.array, code like this:  <code>sub_mask = cv2.resize(sub_mask, (SIZE, SIZE), interpolation=cv2.INTER_NEAREST)\n                    mask[:, :, int(label)] = sub_mask</code>, where 47 means 47 types(including the background), then when doing inference, I don't know how to choose the channel, and actually \nI think the channel is not matching the label.\nSo, as a beginner, I am really curious if anyone who used U-net in this competition and realized the multiple classification and segmentation. \nI will appreciate it if you can share the code or even just give some hints. Thank you.</p>\n\n<p>Additional: I'm using keras and tf.</p>",
      "rawMarkdown": "As the competition is over, I think it's the time to ask this question.\nI tried to use U-net in this competition, since I used it to do binary classification and segmentation before. However I faced a lot of difficulties to do a multiple classification and segmentation. \nMy solution is to transfer the input mask into a [-1, 224, 224, 47] np.array, code like this:  `sub_mask = cv2.resize(sub_mask, (SIZE, SIZE), interpolation=cv2.INTER_NEAREST)\n                    mask[:, :, int(label)] = sub_mask`, where 47 means 47 types(including the background), then when doing inference, I don't know how to choose the channel, and actually \nI think the channel is not matching the label.\nSo, as a beginner, I am really curious if anyone who used U-net in this competition and realized the multiple classification and segmentation. \nI will appreciate it if you can share the code or even just give some hints. Thank you.\n\nAdditional: I'm using keras and tf.",
      "votes": null
    },
    {
      "id": "549865",
      "postDate": "06/11/2019 05:20:23",
      "content": "<p>I was using a lot of Unet with different encoders (actually more than 100,  2-3almost for each class) \nStarted with them to check submission, they had no good score. But they could be used to refine an outputted mask</p>",
      "rawMarkdown": "I was using a lot of Unet with different encoders (actually more than 100,  2-3almost for each class) \nStarted with them to check submission, they had no good score. But they could be used to refine an outputted mask",
      "votes": null
    },
    {
      "id": "549879",
      "postDate": "06/11/2019 05:32:03",
      "content": "<p>Sorry for that, Could you explain what do you mean by  &gt; actually more than 100, 2-3almost for each class? Does that mean you regard it as a 47 bi-classification &amp; segmentation?</p>",
      "rawMarkdown": "Sorry for that, Could you explain what do you mean by  &gt; actually more than 100, 2-3almost for each class? Does that mean you regard it as a 47 bi-classification &amp; segmentation?",
      "votes": null
    },
    {
      "id": "549884",
      "postDate": "06/11/2019 05:39:39",
      "content": "<p>You are correct\nModel for each class, different models as well</p>",
      "rawMarkdown": "You are correct\nModel for each class, different models as well",
      "votes": null
    },
    {
      "id": "549887",
      "postDate": "06/11/2019 05:44:29",
      "content": "<p>I used U-Net at the beginning of this competition. \ninput: 768 * 768 * 3, output: 768 * 768 * 46. \nOutput is activated by sigmoid and trained by BCELoss because one pixel may has two or more class.\nAfter some postprocessing, my U-Net best score is 0.10007(private)/0.10582(public), so I decided to use mask-rcnn. </p>",
      "rawMarkdown": "I used U-Net at the beginning of this competition. \ninput: 768 * 768 * 3, output: 768 * 768 * 46. \nOutput is activated by sigmoid and trained by BCELoss because one pixel may has two or more class.\nAfter some postprocessing, my U-Net best score is 0.10007(private)/0.10582(public), so I decided to use mask-rcnn.",
      "votes": null
    },
    {
      "id": "549917",
      "postDate": "06/11/2019 06:25:24",
      "content": "<p>Seems that demands excellent GPU...</p>",
      "rawMarkdown": "Seems that demands excellent GPU...",
      "votes": null
    },
    {
      "id": "549922",
      "postDate": "06/11/2019 06:28:16",
      "content": "<p>DGX-1 in my control</p>",
      "rawMarkdown": "DGX-1 in my control",
      "votes": null
    },
    {
      "id": "549928",
      "postDate": "06/11/2019 06:34:58",
      "content": "<p>If one pixel may has two or more class, why did you choose sigmoid &amp; bce instead of softmax &amp; cce? And, as you have 46 categories, where is the background?\nWell, and a little beg(perhaps offensive), since the competition is over, could you please share your code of the Unet? We have the similar idea but I got mad for I can not come out any output. I would appreciate it if you did. Thank you again.</p>",
      "rawMarkdown": "If one pixel may has two or more class, why did you choose sigmoid &amp; bce instead of softmax &amp; cce? And, as you have 46 categories, where is the background?\nWell, and a little beg(perhaps offensive), since the competition is over, could you please share your code of the Unet? We have the similar idea but I got mad for I can not come out any output. I would appreciate it if you did. Thank you again.",
      "votes": null
    },
    {
      "id": "550016",
      "postDate": "06/11/2019 08:06:02",
      "content": "<p>If you use sigmoid &amp; bce, each output value has 0 to 1 value. And you can thresholding each value(for example, threshold=0.5). So one pixel can have multi class. If values in all channels are under threshold,  you can consider this pixel as background.</p>\n\n<p>And I'm sorry but I have many work to do, and my code needs many refactoring, so I cannot premise that I can share my code.</p>\n\n<p>Instead, following repository may helps you:\n<a href=\"https://github.com/milesial/Pytorch-UNet\">https://github.com/milesial/Pytorch-UNet</a></p>\n\n<p>I didn't use this repository but this repository uses sigmoid &amp; bce.</p>",
      "rawMarkdown": "If you use sigmoid &amp; bce, each output value has 0 to 1 value. And you can thresholding each value(for example, threshold=0.5). So one pixel can have multi class. If values in all channels are under threshold,  you can consider this pixel as background.\n\nAnd I'm sorry but I have many work to do, and my code needs many refactoring, so I cannot premise that I can share my code.\n\nInstead, following repository may helps you:\nhttps://github.com/milesial/Pytorch-UNet\n\nI didn't use this repository but this repository uses sigmoid &amp; bce.",
      "votes": null
    },
    {
      "id": "550023",
      "postDate": "06/11/2019 08:20:22",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "550681",
      "postDate": "06/11/2019 23:09:45",
      "content": "<p>Not actually, If you are doing mask refinement. training 5 folds DenseNet121-&gt;FPN (Not exactly Unet, but still segmentation) takes 1h-24h on two 1080Ti(depending from class), with average of 2hours per 5 folds per class. A lot of classes have to few predicted classes to be trained. So if you have 1 1080 it will take ~100-200hours  to train all them. Not a zero time, but there is nothing that can not be done with one gpu or kaggle/collab kernels.</p>",
      "rawMarkdown": "Not actually, If you are doing mask refinement. training 5 folds DenseNet121-&gt;FPN (Not exactly Unet, but still segmentation) takes 1h-24h on two 1080Ti(depending from class), with average of 2hours per 5 folds per class. A lot of classes have to few predicted classes to be trained. So if you have 1 1080 it will take ~100-200hours  to train all them. Not a zero time, but there is nothing that can not be done with one gpu or kaggle/collab kernels.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 549865,
      "author_name": "venheads",
      "author_url": "",
      "post_date": "06/11/2019 05:20:23",
      "content": "<p>I was using a lot of Unet with different encoders (actually more than 100,  2-3almost for each class) \nStarted with them to check submission, they had no good score. But they could be used to refine an outputted mask</p>",
      "votes": null,
      "replies": [
        {
          "id": 549879,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "06/11/2019 05:32:03",
          "content": "<p>Sorry for that, Could you explain what do you mean by  &gt; actually more than 100, 2-3almost for each class? Does that mean you regard it as a 47 bi-classification &amp; segmentation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 549884,
          "author_name": "venheads",
          "author_url": "",
          "post_date": "06/11/2019 05:39:39",
          "content": "<p>You are correct\nModel for each class, different models as well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 549917,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "06/11/2019 06:25:24",
          "content": "<p>Seems that demands excellent GPU...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 549922,
          "author_name": "venheads",
          "author_url": "",
          "post_date": "06/11/2019 06:28:16",
          "content": "<p>DGX-1 in my control</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550681,
          "author_name": "realityfaker",
          "author_url": "",
          "post_date": "06/11/2019 23:09:45",
          "content": "<p>Not actually, If you are doing mask refinement. training 5 folds DenseNet121-&gt;FPN (Not exactly Unet, but still segmentation) takes 1h-24h on two 1080Ti(depending from class), with average of 2hours per 5 folds per class. A lot of classes have to few predicted classes to be trained. So if you have 1 1080 it will take ~100-200hours  to train all them. Not a zero time, but there is nothing that can not be done with one gpu or kaggle/collab kernels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 549887,
      "author_name": "rskmoi",
      "author_url": "",
      "post_date": "06/11/2019 05:44:29",
      "content": "<p>I used U-Net at the beginning of this competition. \ninput: 768 * 768 * 3, output: 768 * 768 * 46. \nOutput is activated by sigmoid and trained by BCELoss because one pixel may has two or more class.\nAfter some postprocessing, my U-Net best score is 0.10007(private)/0.10582(public), so I decided to use mask-rcnn. </p>",
      "votes": null,
      "replies": [
        {
          "id": 549928,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "06/11/2019 06:34:58",
          "content": "<p>If one pixel may has two or more class, why did you choose sigmoid &amp; bce instead of softmax &amp; cce? And, as you have 46 categories, where is the background?\nWell, and a little beg(perhaps offensive), since the competition is over, could you please share your code of the Unet? We have the similar idea but I got mad for I can not come out any output. I would appreciate it if you did. Thank you again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550016,
          "author_name": "rskmoi",
          "author_url": "",
          "post_date": "06/11/2019 08:06:02",
          "content": "<p>If you use sigmoid &amp; bce, each output value has 0 to 1 value. And you can thresholding each value(for example, threshold=0.5). So one pixel can have multi class. If values in all channels are under threshold,  you can consider this pixel as background.</p>\n\n<p>And I'm sorry but I have many work to do, and my code needs many refactoring, so I cannot premise that I can share my code.</p>\n\n<p>Instead, following repository may helps you:\n<a href=\"https://github.com/milesial/Pytorch-UNet\">https://github.com/milesial/Pytorch-UNet</a></p>\n\n<p>I didn't use this repository but this repository uses sigmoid &amp; bce.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550023,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "06/11/2019 08:20:22",
          "content": "<p>Thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "549851": "As the competition is over, I think it's the time to ask this question.\nI tried to use U-net in this competition, since I used it to do binary classification and segmentation before. However I faced a lot of difficulties to do a multiple classification and segmentation. \nMy solution is to transfer the input mask into a [-1, 224, 224, 47] np.array, code like this:  `sub_mask = cv2.resize(sub_mask, (SIZE, SIZE), interpolation=cv2.INTER_NEAREST)\n                    mask[:, :, int(label)] = sub_mask`, where 47 means 47 types(including the background), then when doing inference, I don't know how to choose the channel, and actually \nI think the channel is not matching the label.\nSo, as a beginner, I am really curious if anyone who used U-net in this competition and realized the multiple classification and segmentation. \nI will appreciate it if you can share the code or even just give some hints. Thank you.\n\nAdditional: I'm using keras and tf.",
    "549865": "I was using a lot of Unet with different encoders (actually more than 100,  2-3almost for each class) \nStarted with them to check submission, they had no good score. But they could be used to refine an outputted mask",
    "549879": "Sorry for that, Could you explain what do you mean by  &gt; actually more than 100, 2-3almost for each class? Does that mean you regard it as a 47 bi-classification &amp; segmentation?",
    "549884": "You are correct\nModel for each class, different models as well",
    "549887": "I used U-Net at the beginning of this competition. \ninput: 768 * 768 * 3, output: 768 * 768 * 46. \nOutput is activated by sigmoid and trained by BCELoss because one pixel may has two or more class.\nAfter some postprocessing, my U-Net best score is 0.10007(private)/0.10582(public), so I decided to use mask-rcnn.",
    "549917": "Seems that demands excellent GPU...",
    "549922": "DGX-1 in my control",
    "549928": "If one pixel may has two or more class, why did you choose sigmoid &amp; bce instead of softmax &amp; cce? And, as you have 46 categories, where is the background?\nWell, and a little beg(perhaps offensive), since the competition is over, could you please share your code of the Unet? We have the similar idea but I got mad for I can not come out any output. I would appreciate it if you did. Thank you again.",
    "550016": "If you use sigmoid &amp; bce, each output value has 0 to 1 value. And you can thresholding each value(for example, threshold=0.5). So one pixel can have multi class. If values in all channels are under threshold,  you can consider this pixel as background.\n\nAnd I'm sorry but I have many work to do, and my code needs many refactoring, so I cannot premise that I can share my code.\n\nInstead, following repository may helps you:\nhttps://github.com/milesial/Pytorch-UNet\n\nI didn't use this repository but this repository uses sigmoid &amp; bce.",
    "550023": "Thanks for sharing!",
    "550681": "Not actually, If you are doing mask refinement. training 5 folds DenseNet121-&gt;FPN (Not exactly Unet, but still segmentation) takes 1h-24h on two 1080Ti(depending from class), with average of 2hours per 5 folds per class. A lot of classes have to few predicted classes to be trained. So if you have 1 1080 it will take ~100-200hours  to train all them. Not a zero time, but there is nothing that can not be done with one gpu or kaggle/collab kernels."
  },
  "source": "meta"
}