{
  "id": 108306,
  "title": "Weakly supervised segmentation",
  "url": "/competitions/understanding_cloud_organization/discussion/108306",
  "author_name": "",
  "post_date": "2019-09-10T16:49:14.387964200Z",
  "votes": 27,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I've played with weak (image-level) labels, and surprisingly a straightforward DenseNet121 classifier + GradCAM led to reasonable baseline performance: 0.602 in local validation and 0.605 LB. All details on GradCAM-based masks can be found in <a href=\"https://www.kaggle.com/samusram/gradcam-extracting-masks-from-classifier\"><em>GradCAM: extracting masks from classifier</em> kernel</a>, and details on training DenseNet are in <a href=\"https://www.kaggle.com/samusram/cloud-classifier-for-post-processing\"><em>Cloud Classifier for Post-processing</em> kernel</a>.</p>\n\n<p>0.605 LB of the GradCAM is better than performance of several nice segmentation baselines:\n<a href=\"https://www.kaggle.com/dimitreoliveira/understanding-clouds-eda-and-keras-u-net\"><em>Understanding Clouds - EDA and Keras U-Net</em> kernel with 0.567 LB</a>,\n<a href=\"https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras\"><em>MaskRCNN for cloud classification (Keras)</em> kernel with 0.578 LB</a>,\n<a href=\"https://www.kaggle.com/xhlulu/satellite-clouds-u-net-with-resnet-encoder\"><em>Satellite Clouds: U-Net with ResNet Encoder</em> kernel with 0.594 LB</a>.</p>\n\n<p>This made me think about potential of weakly supervised segmentation methods for this competition. After quick googling, I came across some papers which seem relevant:\n1. <a href=\"https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y\">Learning from Weak and Noisy Labels for Semantic Segmentation</a>,\n    <em>I'd like to understand the paper better, after the first scanning it seems that the approach might be helpful for our task, with the obvious modification that we don't need to generate initial noisy pixel-level label, but rather we can use the provided subjective masks.</em>\n2. <a href=\"https://pdfs.semanticscholar.org/d566/73be998b3ed38ccbb53551e38758ae8cfc9d.pdf\">FeaBoost: Joint Feature and Label Refinement for Semantic Segmentation</a>,\n3. <a href=\"https://arxiv.org/pdf/1807.11719.pdf\">A Two-Stream Mutual Attention Network for Semi-supervised Biomedical Segmentation with Noisy Labels</a>,\n4. <a href=\"http://openaccess.thecvf.com/content_cvpr_2018/papers/Zhang_Deep_Unsupervised_Saliency_CVPR_2018_paper.pdf\">Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective</a>,\n5. <a href=\"http://openaccess.thecvf.com/content_ICCV_2017/papers/Zhang_Supervision_by_Fusion_ICCV_2017_paper.pdf\">Supervision by Fusion: Towards Unsupervised Learning of Deep Salient Object Detector</a>.\n<em>edit</em>\n6. <a href=\"https://arxiv.org/pdf/1802.10171.pdf\">Tell Me Where to Look: Guided Attention Inference Network</a> (<em>kindly shared by <a href=\"/ratthachat\">@ratthachat</a></em>)\n7. <a href=\"https://github.com/JackieZhangdx/WeakSupervisedSegmentationList\"><strong>list of state-or-art weakly supervised semantic segmentation works</strong></a></p>",
  "messages": [
    {
      "id": "623267",
      "postDate": "09/10/2019 16:49:14",
      "content": "<p>I've played with weak (image-level) labels, and surprisingly a straightforward DenseNet121 classifier + GradCAM led to reasonable baseline performance: 0.602 in local validation and 0.605 LB. All details on GradCAM-based masks can be found in <a href=\"https://www.kaggle.com/samusram/gradcam-extracting-masks-from-classifier\"><em>GradCAM: extracting masks from classifier</em> kernel</a>, and details on training DenseNet are in <a href=\"https://www.kaggle.com/samusram/cloud-classifier-for-post-processing\"><em>Cloud Classifier for Post-processing</em> kernel</a>.</p>\n\n<p>0.605 LB of the GradCAM is better than performance of several nice segmentation baselines:\n<a href=\"https://www.kaggle.com/dimitreoliveira/understanding-clouds-eda-and-keras-u-net\"><em>Understanding Clouds - EDA and Keras U-Net</em> kernel with 0.567 LB</a>,\n<a href=\"https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras\"><em>MaskRCNN for cloud classification (Keras)</em> kernel with 0.578 LB</a>,\n<a href=\"https://www.kaggle.com/xhlulu/satellite-clouds-u-net-with-resnet-encoder\"><em>Satellite Clouds: U-Net with ResNet Encoder</em> kernel with 0.594 LB</a>.</p>\n\n<p>This made me think about potential of weakly supervised segmentation methods for this competition. After quick googling, I came across some papers which seem relevant:\n1. <a href=\"https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y\">Learning from Weak and Noisy Labels for Semantic Segmentation</a>,\n    <em>I'd like to understand the paper better, after the first scanning it seems that the approach might be helpful for our task, with the obvious modification that we don't need to generate initial noisy pixel-level label, but rather we can use the provided subjective masks.</em>\n2. <a href=\"https://pdfs.semanticscholar.org/d566/73be998b3ed38ccbb53551e38758ae8cfc9d.pdf\">FeaBoost: Joint Feature and Label Refinement for Semantic Segmentation</a>,\n3. <a href=\"https://arxiv.org/pdf/1807.11719.pdf\">A Two-Stream Mutual Attention Network for Semi-supervised Biomedical Segmentation with Noisy Labels</a>,\n4. <a href=\"http://openaccess.thecvf.com/content_cvpr_2018/papers/Zhang_Deep_Unsupervised_Saliency_CVPR_2018_paper.pdf\">Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective</a>,\n5. <a href=\"http://openaccess.thecvf.com/content_ICCV_2017/papers/Zhang_Supervision_by_Fusion_ICCV_2017_paper.pdf\">Supervision by Fusion: Towards Unsupervised Learning of Deep Salient Object Detector</a>.\n<em>edit</em>\n6. <a href=\"https://arxiv.org/pdf/1802.10171.pdf\">Tell Me Where to Look: Guided Attention Inference Network</a> (<em>kindly shared by <a href=\"/ratthachat\">@ratthachat</a></em>)\n7. <a href=\"https://github.com/JackieZhangdx/WeakSupervisedSegmentationList\"><strong>list of state-or-art weakly supervised semantic segmentation works</strong></a></p>",
      "rawMarkdown": "I've played with weak (image-level) labels, and surprisingly a straightforward DenseNet121 classifier + GradCAM led to reasonable baseline performance: 0.602 in local validation and 0.605 LB. All details on GradCAM-based masks can be found in [*GradCAM: extracting masks from classifier* kernel](https://www.kaggle.com/samusram/gradcam-extracting-masks-from-classifier), and details on training DenseNet are in [*Cloud Classifier for Post-processing* kernel](https://www.kaggle.com/samusram/cloud-classifier-for-post-processing).\n\n0.605 LB of the GradCAM is better than performance of several nice segmentation baselines:\n[*Understanding Clouds - EDA and Keras U-Net* kernel with 0.567 LB](https://www.kaggle.com/dimitreoliveira/understanding-clouds-eda-and-keras-u-net),\n[*MaskRCNN for cloud classification (Keras)* kernel with 0.578 LB](https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras),\n[*Satellite Clouds: U-Net with ResNet Encoder* kernel with 0.594 LB](https://www.kaggle.com/xhlulu/satellite-clouds-u-net-with-resnet-encoder).\n\nThis made me think about potential of weakly supervised segmentation methods for this competition. After quick googling, I came across some papers which seem relevant:\n1. [Learning from Weak and Noisy Labels for Semantic Segmentation](https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y),\n    *I'd like to understand the paper better, after the first scanning it seems that the approach might be helpful for our task, with the obvious modification that we don't need to generate initial noisy pixel-level label, but rather we can use the provided subjective masks.*\n2. [FeaBoost: Joint Feature and Label Refinement for Semantic Segmentation](https://pdfs.semanticscholar.org/d566/73be998b3ed38ccbb53551e38758ae8cfc9d.pdf),\n3. [A Two-Stream Mutual Attention Network for Semi-supervised Biomedical Segmentation with Noisy Labels](https://arxiv.org/pdf/1807.11719.pdf),\n4. [Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective](http://openaccess.thecvf.com/content_cvpr_2018/papers/Zhang_Deep_Unsupervised_Saliency_CVPR_2018_paper.pdf),\n5. [Supervision by Fusion: Towards Unsupervised Learning of Deep Salient Object Detector](http://openaccess.thecvf.com/content_ICCV_2017/papers/Zhang_Supervision_by_Fusion_ICCV_2017_paper.pdf).\n*edit*\n6. [Tell Me Where to Look: Guided Attention Inference Network](https://arxiv.org/pdf/1802.10171.pdf) (*kindly shared by @ratthachat*)\n7. [**list of state-or-art weakly supervised semantic segmentation works**](https://github.com/JackieZhangdx/WeakSupervisedSegmentationList)",
      "votes": null
    },
    {
      "id": "623502",
      "postDate": "09/11/2019 03:11:30",
      "content": "<p>Great post and great kernels at the same time. I was thinking about the relationship between GradCAM and Masks, and suddenly found your post to study more.</p>\n\n<p>Thanks for sharing!</p>",
      "rawMarkdown": "Great post and great kernels at the same time. I was thinking about the relationship between GradCAM and Masks, and suddenly found your post to study more.\n\nThanks for sharing!",
      "votes": null
    },
    {
      "id": "623542",
      "postDate": "09/11/2019 04:22:09",
      "content": "<p>Thank you for the kind comment, it made my morning!</p>",
      "rawMarkdown": "Thank you for the kind comment, it made my morning!",
      "votes": null
    },
    {
      "id": "623913",
      "postDate": "09/11/2019 12:06:47",
      "content": "<p><a href=\"/samusram\">@samusram</a> Do you think this paper relevant ?\n<a href=\"https://arxiv.org/abs/1802.10171\">Tell Me Where to Look: Guided Attention Inference Network</a></p>",
      "rawMarkdown": "samusram Do you think this paper relevant ?\n[Tell Me Where to Look: Guided Attention Inference Network](https://arxiv.org/abs/1802.10171)",
      "votes": null
    },
    {
      "id": "623951",
      "postDate": "09/11/2019 12:58:10",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> , Wow! This is very much relevant!\nThanks for providing me with great tips, now I have to figure out what to do first😊 </p>\n\n<p>Btw, I came across <a href=\"https://github.com/JackieZhangdx/WeakSupervisedSegmentationList\">this list of state-or-art weakly supervised semantic segmentation works</a>. It seems like there are quite a few fascinating papers to study/play with😊 </p>",
      "rawMarkdown": "ratthachat , Wow! This is very much relevant!\nThanks for providing me with great tips, now I have to figure out what to do first😊 \n\nBtw, I came across [this list of state-or-art weakly supervised semantic segmentation works](https://github.com/JackieZhangdx/WeakSupervisedSegmentationList). It seems like there are quite a few fascinating papers to study/play with😊",
      "votes": null
    },
    {
      "id": "623952",
      "postDate": "09/11/2019 13:06:14",
      "content": "<p>I am very much interested in studying weakly supervised learning in general [related to this competition or not]. They are too much materials to read by one person, so I am happy to find you as a companion here. :D</p>\n\n<p>BTW, the above paper also was shared to me by our friend <a href=\"/backaggle\">@backaggle</a> . He just provided us an open source of the paper (nevertheless, in Pytorch, not in Keras :)</p>",
      "rawMarkdown": "I am very much interested in studying weakly supervised learning in general [related to this competition or not]. They are too much materials to read by one person, so I am happy to find you as a companion here. :D\n\nBTW, the above paper also was shared to me by our friend @backaggle . He just provided us an open source of the paper (nevertheless, in Pytorch, not in Keras :)",
      "votes": null
    },
    {
      "id": "624062",
      "postDate": "09/11/2019 16:05:13",
      "content": "<p>Thanks for sharing this wonderful kernel @Raman. </p>",
      "rawMarkdown": "Thanks for sharing this wonderful kernel @Raman.",
      "votes": null
    },
    {
      "id": "624067",
      "postDate": "09/11/2019 16:15:48",
      "content": "<p>I'm glad you liked it! Thanks for the comment</p>",
      "rawMarkdown": "I'm glad you liked it! Thanks for the comment",
      "votes": null
    },
    {
      "id": "624449",
      "postDate": "09/12/2019 05:02:56",
      "content": "<p>Yeah. Great Learning for me. Was always looking at Grad CAM more from Explainable AI perspective. Now, I am trying to correlate it in some different perspective. This is what I like more in this forum as well. Healthy discussions on different topics. </p>",
      "rawMarkdown": "Yeah. Great Learning for me. Was always looking at Grad CAM more from Explainable AI perspective. Now, I am trying to correlate it in some different perspective. This is what I like more in this forum as well. Healthy discussions on different topics.",
      "votes": null
    },
    {
      "id": "625321",
      "postDate": "09/12/2019 23:52:35",
      "content": "<p>Raman <a href=\"/samusram\">@samusram</a> , allow me to ask some questions on your kernel : </p>\n\n<ul>\n<li><p>Every image has  a <strong>\"black stripe\"</strong> of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this? </p></li>\n<li><p>since the shape of a mask is always consisting of straight lines. Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ? \n(after quick search I found this but not yet able to integrate : <a href=\"https://gist.github.com/bigsnarfdude/d811e31ee17495f82f10db12651ae82d\">https://gist.github.com/bigsnarfdude/d811e31ee17495f82f10db12651ae82d</a>) </p></li>\n<li><p>Just a small point1 : <code>np.empty</code> and <code>cv2.resize</code> indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)</p></li>\n<li><p>Small point2 : Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 ,  but changed the backbone to B2)</p></li>\n<li><p>Last interesting fact : I created some random noise pictures using <code>np.random.rand</code> and the model predicted all these pictures as <strong>sugar</strong> with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)</p></li>\n</ul>",
      "rawMarkdown": "Raman @samusram , allow me to ask some questions on your kernel : \n\n- Every image has  a **\"black stripe\"** of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this? \n\n- since the shape of a mask is always consisting of straight lines. Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ? \n(after quick search I found this but not yet able to integrate : https://gist.github.com/bigsnarfdude/d811e31ee17495f82f10db12651ae82d) \n\n- Just a small point1 : `np.empty` and `cv2.resize` indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)\n\n- Small point2 : Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 ,  but changed the backbone to B2)\n\n- Last interesting fact : I created some random noise pictures using `np.random.rand` and the model predicted all these pictures as **sugar** with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)",
      "votes": null
    },
    {
      "id": "625480",
      "postDate": "09/13/2019 05:58:17",
      "content": "<p>Jung <a href=\"/ratthachat\">@ratthachat</a> ,</p>\n\n<p>thank You so much for such detailed, extremely helpful post! I do appreciate your unique approach!</p>\n\n<blockquote>\n  <p>Every image has a \"black stripe\" of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this?</p>\n</blockquote>\n\n<p>Sure, in the function <code>generate_gradcam_masks</code> I simply do\n<code>binarized_gradcams_batch[imgs_batch[:,:,:,0]==0] = 0</code></p>\n\n<blockquote>\n  <p>Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ?</p>\n</blockquote>\n\n<p>I think it's definitely an interesting direction. I believe that if we had perfect bboxes and reasonable instance segmentation, then this type of post-processing would help a lot. However, as both the masks and bboxes seem to be rather noisy (at least based on some of visualizations in the kernels), I think at this stage such post-processing might hurt more than help. If there's enough time, I believe it's worth trying out anyway.</p>\n\n<blockquote>\n  <p>np.empty and cv2.resize indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)</p>\n</blockquote>\n\n<p>Thank You for the correction, it's very much appreciated! I've realized this type of bug after my mask-submission code was producing garbage, but has fixed it only in the submission-generation cell. Later I'll correct it elsewhere as well, thx! :)</p>\n\n<blockquote>\n  <p>Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 , but changed the backbone to B2)</p>\n</blockquote>\n\n<p>Thank you for the valuable point! Moreover, thanks for reporting the B2 performance! I'd like to dive into the suggested by you <em>Tell Me Where To Look</em> paper, but afterwards I'd definitely get back to the classifier kernel and will use the improved segmentation!</p>\n\n<blockquote>\n  <p>I created some random noise pictures using np.random.rand and the model predicted all these pictures as sugar with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)</p>\n</blockquote>\n\n<p>Interesting :) It's reassuring that even for the predicted class probability is around 50%, yet it's cool observation...sugar.. \nthx for sharing! :)</p>",
      "rawMarkdown": "Jung @ratthachat ,\n\nthank You so much for such detailed, extremely helpful post! I do appreciate your unique approach!\n\n&gt; Every image has a \"black stripe\" of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this?\n\nSure, in the function `generate_gradcam_masks` I simply do\n```  binarized_gradcams_batch[imgs_batch[:,:,:,0]==0] = 0 ```\n\n&gt; Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ?\n\nI think it's definitely an interesting direction. I believe that if we had perfect bboxes and reasonable instance segmentation, then this type of post-processing would help a lot. However, as both the masks and bboxes seem to be rather noisy (at least based on some of visualizations in the kernels), I think at this stage such post-processing might hurt more than help. If there's enough time, I believe it's worth trying out anyway.\n\n&gt; np.empty and cv2.resize indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)\n\nThank You for the correction, it's very much appreciated! I've realized this type of bug after my mask-submission code was producing garbage, but has fixed it only in the submission-generation cell. Later I'll correct it elsewhere as well, thx! :)\n\n&gt; Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 , but changed the backbone to B2)\n\nThank you for the valuable point! Moreover, thanks for reporting the B2 performance! I'd like to dive into the suggested by you *Tell Me Where To Look* paper, but afterwards I'd definitely get back to the classifier kernel and will use the improved segmentation!\n\n&gt;  I created some random noise pictures using np.random.rand and the model predicted all these pictures as sugar with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)\n\nInteresting :) It's reassuring that even for the predicted class probability is around 50%, yet it's cool observation...sugar.. \nthx for sharing! :)",
      "votes": null
    },
    {
      "id": "627746",
      "postDate": "09/16/2019 10:13:06",
      "content": "<p>Good information.. Thanks for sharing <a href=\"/samusram\">@samusram</a> </p>",
      "rawMarkdown": "Good information.. Thanks for sharing @samusram",
      "votes": null
    },
    {
      "id": "627777",
      "postDate": "09/16/2019 11:04:17",
      "content": "<p>I'm glad you found it interesting, <a href=\"/kranthi9\">@kranthi9</a> </p>",
      "rawMarkdown": "I'm glad you found it interesting, @kranthi9",
      "votes": null
    },
    {
      "id": "631556",
      "postDate": "09/22/2019 08:19:22",
      "content": "<p>I follow up the discussion here on using rectangle-shape post-processing on masks here : \n<a href=\"https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\">https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/</a>\n(I wrote this kernel when I run out of GPU quota :)</p>\n\n<p><strong>EDIT : Update to have more and better choices : convex-shape or approximate polygon shape</strong></p>\n\n<p><img src=\"https://i.ibb.co/w45jCdW/convex-mask.jpg\" alt=\"rectangle masks\"></p>",
      "rawMarkdown": "I follow up the discussion here on using rectangle-shape post-processing on masks here : \nhttps://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\n(I wrote this kernel when I run out of GPU quota :)\n\n**EDIT : Update to have more and better choices : convex-shape or approximate polygon shape**\n\n![rectangle masks](https://i.ibb.co/w45jCdW/convex-mask.jpg)",
      "votes": null
    },
    {
      "id": "667470",
      "postDate": "11/07/2019 09:14:22",
      "content": "<p>Thanks for sharing!! </p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": null
    },
    {
      "id": "667477",
      "postDate": "11/07/2019 09:25:54",
      "content": "<p>Sure, I hope the links would be helpful to us</p>",
      "rawMarkdown": "Sure, I hope the links would be helpful to us",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 623502,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/11/2019 03:11:30",
      "content": "<p>Great post and great kernels at the same time. I was thinking about the relationship between GradCAM and Masks, and suddenly found your post to study more.</p>\n\n<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 623542,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "09/11/2019 04:22:09",
          "content": "<p>Thank you for the kind comment, it made my morning!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 623913,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "09/11/2019 12:06:47",
          "content": "<p><a href=\"/samusram\">@samusram</a> Do you think this paper relevant ?\n<a href=\"https://arxiv.org/abs/1802.10171\">Tell Me Where to Look: Guided Attention Inference Network</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 623951,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "09/11/2019 12:58:10",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> , Wow! This is very much relevant!\nThanks for providing me with great tips, now I have to figure out what to do first😊 </p>\n\n<p>Btw, I came across <a href=\"https://github.com/JackieZhangdx/WeakSupervisedSegmentationList\">this list of state-or-art weakly supervised semantic segmentation works</a>. It seems like there are quite a few fascinating papers to study/play with😊 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 623952,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "09/11/2019 13:06:14",
          "content": "<p>I am very much interested in studying weakly supervised learning in general [related to this competition or not]. They are too much materials to read by one person, so I am happy to find you as a companion here. :D</p>\n\n<p>BTW, the above paper also was shared to me by our friend <a href=\"/backaggle\">@backaggle</a> . He just provided us an open source of the paper (nevertheless, in Pytorch, not in Keras :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 624062,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "09/11/2019 16:05:13",
      "content": "<p>Thanks for sharing this wonderful kernel @Raman. </p>",
      "votes": null,
      "replies": [
        {
          "id": 624067,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "09/11/2019 16:15:48",
          "content": "<p>I'm glad you liked it! Thanks for the comment</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 624449,
          "author_name": "manojprabhaakr",
          "author_url": "",
          "post_date": "09/12/2019 05:02:56",
          "content": "<p>Yeah. Great Learning for me. Was always looking at Grad CAM more from Explainable AI perspective. Now, I am trying to correlate it in some different perspective. This is what I like more in this forum as well. Healthy discussions on different topics. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 625321,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/12/2019 23:52:35",
      "content": "<p>Raman <a href=\"/samusram\">@samusram</a> , allow me to ask some questions on your kernel : </p>\n\n<ul>\n<li><p>Every image has  a <strong>\"black stripe\"</strong> of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this? </p></li>\n<li><p>since the shape of a mask is always consisting of straight lines. Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ? \n(after quick search I found this but not yet able to integrate : <a href=\"https://gist.github.com/bigsnarfdude/d811e31ee17495f82f10db12651ae82d\">https://gist.github.com/bigsnarfdude/d811e31ee17495f82f10db12651ae82d</a>) </p></li>\n<li><p>Just a small point1 : <code>np.empty</code> and <code>cv2.resize</code> indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)</p></li>\n<li><p>Small point2 : Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 ,  but changed the backbone to B2)</p></li>\n<li><p>Last interesting fact : I created some random noise pictures using <code>np.random.rand</code> and the model predicted all these pictures as <strong>sugar</strong> with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 625480,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "09/13/2019 05:58:17",
          "content": "<p>Jung <a href=\"/ratthachat\">@ratthachat</a> ,</p>\n\n<p>thank You so much for such detailed, extremely helpful post! I do appreciate your unique approach!</p>\n\n<blockquote>\n  <p>Every image has a \"black stripe\" of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this?</p>\n</blockquote>\n\n<p>Sure, in the function <code>generate_gradcam_masks</code> I simply do\n<code>binarized_gradcams_batch[imgs_batch[:,:,:,0]==0] = 0</code></p>\n\n<blockquote>\n  <p>Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ?</p>\n</blockquote>\n\n<p>I think it's definitely an interesting direction. I believe that if we had perfect bboxes and reasonable instance segmentation, then this type of post-processing would help a lot. However, as both the masks and bboxes seem to be rather noisy (at least based on some of visualizations in the kernels), I think at this stage such post-processing might hurt more than help. If there's enough time, I believe it's worth trying out anyway.</p>\n\n<blockquote>\n  <p>np.empty and cv2.resize indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)</p>\n</blockquote>\n\n<p>Thank You for the correction, it's very much appreciated! I've realized this type of bug after my mask-submission code was producing garbage, but has fixed it only in the submission-generation cell. Later I'll correct it elsewhere as well, thx! :)</p>\n\n<blockquote>\n  <p>Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 , but changed the backbone to B2)</p>\n</blockquote>\n\n<p>Thank you for the valuable point! Moreover, thanks for reporting the B2 performance! I'd like to dive into the suggested by you <em>Tell Me Where To Look</em> paper, but afterwards I'd definitely get back to the classifier kernel and will use the improved segmentation!</p>\n\n<blockquote>\n  <p>I created some random noise pictures using np.random.rand and the model predicted all these pictures as sugar with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)</p>\n</blockquote>\n\n<p>Interesting :) It's reassuring that even for the predicted class probability is around 50%, yet it's cool observation...sugar.. \nthx for sharing! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 627746,
      "author_name": "kranthi9",
      "author_url": "",
      "post_date": "09/16/2019 10:13:06",
      "content": "<p>Good information.. Thanks for sharing <a href=\"/samusram\">@samusram</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 627777,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "09/16/2019 11:04:17",
          "content": "<p>I'm glad you found it interesting, <a href=\"/kranthi9\">@kranthi9</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 631556,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/22/2019 08:19:22",
      "content": "<p>I follow up the discussion here on using rectangle-shape post-processing on masks here : \n<a href=\"https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\">https://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/</a>\n(I wrote this kernel when I run out of GPU quota :)</p>\n\n<p><strong>EDIT : Update to have more and better choices : convex-shape or approximate polygon shape</strong></p>\n\n<p><img src=\"https://i.ibb.co/w45jCdW/convex-mask.jpg\" alt=\"rectangle masks\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 667470,
      "author_name": "ashit10",
      "author_url": "",
      "post_date": "11/07/2019 09:14:22",
      "content": "<p>Thanks for sharing!! </p>",
      "votes": null,
      "replies": [
        {
          "id": 667477,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "11/07/2019 09:25:54",
          "content": "<p>Sure, I hope the links would be helpful to us</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "623267": "I've played with weak (image-level) labels, and surprisingly a straightforward DenseNet121 classifier + GradCAM led to reasonable baseline performance: 0.602 in local validation and 0.605 LB. All details on GradCAM-based masks can be found in [*GradCAM: extracting masks from classifier* kernel](https://www.kaggle.com/samusram/gradcam-extracting-masks-from-classifier), and details on training DenseNet are in [*Cloud Classifier for Post-processing* kernel](https://www.kaggle.com/samusram/cloud-classifier-for-post-processing).\n\n0.605 LB of the GradCAM is better than performance of several nice segmentation baselines:\n[*Understanding Clouds - EDA and Keras U-Net* kernel with 0.567 LB](https://www.kaggle.com/dimitreoliveira/understanding-clouds-eda-and-keras-u-net),\n[*MaskRCNN for cloud classification (Keras)* kernel with 0.578 LB](https://www.kaggle.com/frlemarchand/maskrcnn-for-cloud-classification-keras),\n[*Satellite Clouds: U-Net with ResNet Encoder* kernel with 0.594 LB](https://www.kaggle.com/xhlulu/satellite-clouds-u-net-with-resnet-encoder).\n\nThis made me think about potential of weakly supervised segmentation methods for this competition. After quick googling, I came across some papers which seem relevant:\n1. [Learning from Weak and Noisy Labels for Semantic Segmentation](https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y),\n    *I'd like to understand the paper better, after the first scanning it seems that the approach might be helpful for our task, with the obvious modification that we don't need to generate initial noisy pixel-level label, but rather we can use the provided subjective masks.*\n2. [FeaBoost: Joint Feature and Label Refinement for Semantic Segmentation](https://pdfs.semanticscholar.org/d566/73be998b3ed38ccbb53551e38758ae8cfc9d.pdf),\n3. [A Two-Stream Mutual Attention Network for Semi-supervised Biomedical Segmentation with Noisy Labels](https://arxiv.org/pdf/1807.11719.pdf),\n4. [Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective](http://openaccess.thecvf.com/content_cvpr_2018/papers/Zhang_Deep_Unsupervised_Saliency_CVPR_2018_paper.pdf),\n5. [Supervision by Fusion: Towards Unsupervised Learning of Deep Salient Object Detector](http://openaccess.thecvf.com/content_ICCV_2017/papers/Zhang_Supervision_by_Fusion_ICCV_2017_paper.pdf).\n*edit*\n6. [Tell Me Where to Look: Guided Attention Inference Network](https://arxiv.org/pdf/1802.10171.pdf) (*kindly shared by @ratthachat*)\n7. [**list of state-or-art weakly supervised semantic segmentation works**](https://github.com/JackieZhangdx/WeakSupervisedSegmentationList)",
    "623502": "Great post and great kernels at the same time. I was thinking about the relationship between GradCAM and Masks, and suddenly found your post to study more.\n\nThanks for sharing!",
    "623542": "Thank you for the kind comment, it made my morning!",
    "623913": "samusram Do you think this paper relevant ?\n[Tell Me Where to Look: Guided Attention Inference Network](https://arxiv.org/abs/1802.10171)",
    "623951": "ratthachat , Wow! This is very much relevant!\nThanks for providing me with great tips, now I have to figure out what to do first😊 \n\nBtw, I came across [this list of state-or-art weakly supervised semantic segmentation works](https://github.com/JackieZhangdx/WeakSupervisedSegmentationList). It seems like there are quite a few fascinating papers to study/play with😊",
    "623952": "I am very much interested in studying weakly supervised learning in general [related to this competition or not]. They are too much materials to read by one person, so I am happy to find you as a companion here. :D\n\nBTW, the above paper also was shared to me by our friend @backaggle . He just provided us an open source of the paper (nevertheless, in Pytorch, not in Keras :)",
    "624062": "Thanks for sharing this wonderful kernel @Raman.",
    "624067": "I'm glad you liked it! Thanks for the comment",
    "624449": "Yeah. Great Learning for me. Was always looking at Grad CAM more from Explainable AI perspective. Now, I am trying to correlate it in some different perspective. This is what I like more in this forum as well. Healthy discussions on different topics.",
    "625321": "Raman @samusram , allow me to ask some questions on your kernel : \n\n- Every image has  a **\"black stripe\"** of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this? \n\n- since the shape of a mask is always consisting of straight lines. Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ? \n(after quick search I found this but not yet able to integrate : https://gist.github.com/bigsnarfdude/d811e31ee17495f82f10db12651ae82d) \n\n- Just a small point1 : `np.empty` and `cv2.resize` indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)\n\n- Small point2 : Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 ,  but changed the backbone to B2)\n\n- Last interesting fact : I created some random noise pictures using `np.random.rand` and the model predicted all these pictures as **sugar** with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)",
    "625480": "Jung @ratthachat ,\n\nthank You so much for such detailed, extremely helpful post! I do appreciate your unique approach!\n\n&gt; Every image has a \"black stripe\" of satellite moving, so we should prevent our prediction mask not to predict on that region. But I see on your plot that GradCAM masks don't touch the black region, do you do some kind of post-processing (or a plot technique) to prevent this?\n\nSure, in the function `generate_gradcam_masks` I simply do\n```  binarized_gradcams_batch[imgs_batch[:,:,:,0]==0] = 0 ```\n\n&gt; Do you think creating a bounding box to cover a high-probability mass region in the GradCAM output make sense ?\n\nI think it's definitely an interesting direction. I believe that if we had perfect bboxes and reasonable instance segmentation, then this type of post-processing would help a lot. However, as both the masks and bboxes seem to be rather noisy (at least based on some of visualizations in the kernels), I think at this stage such post-processing might hurt more than help. If there's enough time, I believe it's worth trying out anyway.\n\n&gt; np.empty and cv2.resize indeed get height/width in an opposite way, and it seems your code on GradCAM still has some bugs on this (but the code was running fine since we used 224x224 :)\n\nThank You for the correction, it's very much appreciated! I've realized this type of bug after my mask-submission code was producing garbage, but has fixed it only in the submission-generation cell. Later I'll correct it elsewhere as well, thx! :)\n\n&gt; Since Andrew updated his kernel to get .645, it seems your multi-label post-processing kernel will have a bit more score. (I got around .646-.647 , but changed the backbone to B2)\n\nThank you for the valuable point! Moreover, thanks for reporting the B2 performance! I'd like to dive into the suggested by you *Tell Me Where To Look* paper, but afterwards I'd definitely get back to the classifier kernel and will use the improved segmentation!\n\n&gt;  I created some random noise pictures using np.random.rand and the model predicted all these pictures as sugar with around &gt; 50% probability. I guess this makes sense 😄 (or not -- so we can have a bit more adversarial training?)\n\nInteresting :) It's reassuring that even for the predicted class probability is around 50%, yet it's cool observation...sugar.. \nthx for sharing! :)",
    "627746": "Good information.. Thanks for sharing @samusram",
    "627777": "I'm glad you found it interesting, @kranthi9",
    "631556": "I follow up the discussion here on using rectangle-shape post-processing on masks here : \nhttps://www.kaggle.com/ratthachat/cloud-rectangle-mask-postprocessing-no-gpu/\n(I wrote this kernel when I run out of GPU quota :)\n\n**EDIT : Update to have more and better choices : convex-shape or approximate polygon shape**\n\n![rectangle masks](https://i.ibb.co/w45jCdW/convex-mask.jpg)",
    "667470": "Thanks for sharing!!",
    "667477": "Sure, I hope the links would be helpful to us"
  },
  "source": "meta"
}