{
  "id": 118019,
  "title": "Easy silver in last days [55th]",
  "url": "/competitions/understanding_cloud_organization/writeups/ods-ai-qubvel-easy-silver-in-last-days-55th",
  "author_name": "",
  "post_date": "2019-11-19T11:48:53.437Z",
  "votes": 49,
  "comment_count": 31,
  "views": 0,
  "content": "<h2>Easy silver in last days</h2>\n\n<p>I have adopted my pipeline from Severstal Defect Detection and was able to get silver medal in last two days with just 6 submissions, here is a short description of 55th place solution.</p>\n\n<p>2 step pileline\n 1) Multi-task network (classification + segmentation) as classifier to remove empty masks\n 2) Binary segmentation for each class</p>\n\n<h3>1st step.</h3>\n\n<p>I have trained 5-fold <code>FPN(resnet34) + aux classfication output</code> on  480x640 images using <code>Flip</code>, <code>RandomBrightness</code> as augmentations. Model trained just 6-7 epochs and than starts to overfit, I do nothing with that, just save top 5 checkpoints according to metric.</p>\n\n<p>Loss (segmentation head): bce+dice\nLoss (classification head): bce\nOptimizer: AdamW\nPostprocessing: remove masks less than 10000 pixels\nThresholds: [0.6, 0.6, 0.6, 0.6]\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F940556%2Faec538c6a536cc9dd070697b4e416870%2F2019-11-19%2010-58-58.png?generation=1574150366439784&amp;alt=media\" alt=\"\"></p>\n\n<h3>2nd step.</h3>\n\n<p>For each class trained <code>2 x Unet(se_resnext50_34x4d)</code> only on images with masks of that class!\nwith same optimizer, image size and augmentations.</p>\n\n<p>Loss: bce+dice\nThresholds: [0.4, 0.4, 0.4, 0.4]</p>\n\n<h3>Ensemble</h3>\n\n<p>For all models made checkpoints weights! averaging (+0.005-0.01 on validation).\nModels over each stage have been just averaged with Flip TTA.</p>\n\n<h3>Useful links</h3>\n\n<ul>\n<li>Segmentation Models: <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">https://github.com/qubvel/segmentation_models.pytorch</a></li>\n<li>Test Time Augmentation for PyTorch: <a href=\"https://github.com/qubvel/ttach\">https://github.com/qubvel/ttach</a></li>\n</ul>\n\n<p><strong>And congratulations to winners!</strong></p>",
  "messages": [
    {
      "id": "676474",
      "postDate": "11/19/2019 08:06:08",
      "content": "<h2>Easy silver in last days</h2>\n\n<p>I have adopted my pipeline from Severstal Defect Detection and was able to get silver medal in last two days with just 6 submissions, here is a short description of 55th place solution.</p>\n\n<p>2 step pileline\n 1) Multi-task network (classification + segmentation) as classifier to remove empty masks\n 2) Binary segmentation for each class</p>\n\n<h3>1st step.</h3>\n\n<p>I have trained 5-fold <code>FPN(resnet34) + aux classfication output</code> on  480x640 images using <code>Flip</code>, <code>RandomBrightness</code> as augmentations. Model trained just 6-7 epochs and than starts to overfit, I do nothing with that, just save top 5 checkpoints according to metric.</p>\n\n<p>Loss (segmentation head): bce+dice\nLoss (classification head): bce\nOptimizer: AdamW\nPostprocessing: remove masks less than 10000 pixels\nThresholds: [0.6, 0.6, 0.6, 0.6]\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F940556%2Faec538c6a536cc9dd070697b4e416870%2F2019-11-19%2010-58-58.png?generation=1574150366439784&amp;alt=media\" alt=\"\"></p>\n\n<h3>2nd step.</h3>\n\n<p>For each class trained <code>2 x Unet(se_resnext50_34x4d)</code> only on images with masks of that class!\nwith same optimizer, image size and augmentations.</p>\n\n<p>Loss: bce+dice\nThresholds: [0.4, 0.4, 0.4, 0.4]</p>\n\n<h3>Ensemble</h3>\n\n<p>For all models made checkpoints weights! averaging (+0.005-0.01 on validation).\nModels over each stage have been just averaged with Flip TTA.</p>\n\n<h3>Useful links</h3>\n\n<ul>\n<li>Segmentation Models: <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">https://github.com/qubvel/segmentation_models.pytorch</a></li>\n<li>Test Time Augmentation for PyTorch: <a href=\"https://github.com/qubvel/ttach\">https://github.com/qubvel/ttach</a></li>\n</ul>\n\n<p><strong>And congratulations to winners!</strong></p>",
      "rawMarkdown": "## Easy silver in last days\nI have adopted my pipeline from Severstal Defect Detection and was able to get silver medal in last two days with just 6 submissions, here is a short description of 55th place solution.\n\n2 step pileline\n 1) Multi-task network (classification + segmentation) as classifier to remove empty masks\n 2) Binary segmentation for each class\n \n### 1st step.\nI have trained 5-fold `FPN(resnet34) + aux classfication output` on  480x640 images using `Flip`, `RandomBrightness` as augmentations. Model trained just 6-7 epochs and than starts to overfit, I do nothing with that, just save top 5 checkpoints according to metric.\n\nLoss (segmentation head): bce+dice\nLoss (classification head): bce\nOptimizer: AdamW\nPostprocessing: remove masks less than 10000 pixels\nThresholds: [0.6, 0.6, 0.6, 0.6]\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F940556%2Faec538c6a536cc9dd070697b4e416870%2F2019-11-19%2010-58-58.png?generation=1574150366439784&amp;alt=media)\n\n### 2nd step.\nFor each class trained `2 x Unet(se_resnext50_34x4d)` only on images with masks of that class!\nwith same optimizer, image size and augmentations.\n\nLoss: bce+dice\nThresholds: [0.4, 0.4, 0.4, 0.4]\n\n### Ensemble\nFor all models made checkpoints weights! averaging (+0.005-0.01 on validation).\nModels over each stage have been just averaged with Flip TTA.\n\n### Useful links\n - Segmentation Models: https://github.com/qubvel/segmentation_models.pytorch\n - Test Time Augmentation for PyTorch: https://github.com/qubvel/ttach\n\n**And congratulations to winners!**",
      "votes": null
    },
    {
      "id": "676481",
      "postDate": "11/19/2019 08:13:57",
      "content": "<p>Sweet and Simple!! Thank you for sharing and also thank you for your <code>segmentation_models.pytorch</code> library. Extremely helpful and convenient </p>",
      "rawMarkdown": "Sweet and Simple!! Thank you for sharing and also thank you for your `segmentation_models.pytorch` library. Extremely helpful and convenient",
      "votes": null
    },
    {
      "id": "676482",
      "postDate": "11/19/2019 08:13:57",
      "content": "<p>most of the time i use your library,thanks a lot and congratulations</p>",
      "rawMarkdown": "most of the time i use your library,thanks a lot and congratulations",
      "votes": null
    },
    {
      "id": "676523",
      "postDate": "11/19/2019 09:03:50",
      "content": "<p>Congratulations\nThank you for Sharing your Insights &amp; Approach! <a href=\"/pavel92\">@pavel92</a> </p>",
      "rawMarkdown": "Congratulations\nThank you for Sharing your Insights &amp; Approach! @pavel92",
      "votes": null
    },
    {
      "id": "676551",
      "postDate": "11/19/2019 09:44:30",
      "content": "<p><a href=\"/pavel92\">@pavel92</a> thanks for your solution writeup. This was my first computer vision competition and I used only your library 'segmentation models' to implement my solution. I must say this library is really helpful for beginners like me. After reading the comments of top solutions, I came to know about the term 'weights averaging from different checkpoints'. Will you please share some resources to understand this concept and may be a sample boiler-plate implementation of this idea?</p>",
      "rawMarkdown": "pavel92 thanks for your solution writeup. This was my first computer vision competition and I used only your library 'segmentation models' to implement my solution. I must say this library is really helpful for beginners like me. After reading the comments of top solutions, I came to know about the term 'weights averaging from different checkpoints'. Will you please share some resources to understand this concept and may be a sample boiler-plate implementation of this idea?",
      "votes": null
    },
    {
      "id": "676573",
      "postDate": "11/19/2019 10:12:22",
      "content": "<p>```\nimport torch\nfrom collections import OrderedDict\nfrom typing import List</p>\n\n<p>checkpoints_weights_paths: List[str] = ...  # sorted in descending order by score\nmodel: torch.nn.Module = ...</p>\n\n<p>def average_weights(state_dicts: List[dict]):\n    everage_dict = OrderedDict()\n    for k in state_dicts[0].keys():\n        everage_dict[k] = sum([state_dict[k] for state_dict in state_dicts]) / len(state_dicts)\n    return everage_dict</p>\n\n<p>all_weights = [torch.load(path) for path in checkpoints_weights_paths]</p>\n\n<p>best_score = 0\nbest_weights = []</p>\n\n<p>for w in all_weights:\n    current_weights = best_weights + [w]\n    average_dict = average_weights(current_weights)\n    model.load_state_dict(average_dict)\n    score = evaluate_model(model, ...)\n    if score &gt; best_score:\n        best_score = score\n        best_weights.append(w)\n```</p>\n\n<p><a href=\"https://gist.github.com/qubvel/70c3d5e4cddcde731408f478e12ef87b\">https://gist.github.com/qubvel/70c3d5e4cddcde731408f478e12ef87b</a></p>",
      "rawMarkdown": "```\nimport torch\nfrom collections import OrderedDict\nfrom typing import List\n\ncheckpoints_weights_paths: List[str] = ...  # sorted in descending order by score\nmodel: torch.nn.Module = ...\n\n\ndef average_weights(state_dicts: List[dict]):\n    everage_dict = OrderedDict()\n    for k in state_dicts[0].keys():\n        everage_dict[k] = sum([state_dict[k] for state_dict in state_dicts]) / len(state_dicts)\n    return everage_dict\n\n\nall_weights = [torch.load(path) for path in checkpoints_weights_paths]\n\nbest_score = 0\nbest_weights = []\n\nfor w in all_weights:\n    current_weights = best_weights + [w]\n    average_dict = average_weights(current_weights)\n    model.load_state_dict(average_dict)\n    score = evaluate_model(model, ...)\n    if score &gt; best_score:\n        best_score = score\n        best_weights.append(w)\n```\n\nhttps://gist.github.com/qubvel/70c3d5e4cddcde731408f478e12ef87b",
      "votes": null
    },
    {
      "id": "676590",
      "postDate": "11/19/2019 10:30:49",
      "content": "<p>Very impressive!!! Congratulations 🎉  and thanks for sharing ❤️ </p>",
      "rawMarkdown": "Very impressive!!! Congratulations 🎉  and thanks for sharing ❤️",
      "votes": null
    },
    {
      "id": "676608",
      "postDate": "11/19/2019 10:50:36",
      "content": "<p>Thanks a lot :)</p>",
      "rawMarkdown": "Thanks a lot :)",
      "votes": null
    },
    {
      "id": "676609",
      "postDate": "11/19/2019 10:51:51",
      "content": "<p>Pavel, can you please leave here some info (maybe links to read) about such multi-task network and how to implement it? </p>",
      "rawMarkdown": "Pavel, can you please leave here some info (maybe links to read) about such multi-task network and how to implement it?",
      "votes": null
    },
    {
      "id": "676632",
      "postDate": "11/19/2019 11:27:40",
      "content": "<p>I have prepared new feature for SMP library that add aux output for models (look at <a href=\"https://github.com/qubvel/segmentation_models.pytorch/tree/models-refactoring\">https://github.com/qubvel/segmentation_models.pytorch/tree/models-refactoring</a>)</p>\n\n<p>According to this implementation my multi-task network defined as follows:\n```\nclass GatedFPN(smp.FPN):</p>\n\n<pre><code>def forward(self, x):\n    mask, label = super().forward(x)\n    return dict(\n        mask=mask*label.reshape(*label.size(), 1, 1),\n        label=label,\n    )\n</code></pre>\n\n<p>aux_params = dict(classes=4, activation='sigmoid', dropout=0.5, pooling='avg')\nmodel = GatedFPN('resnet34', encoder_weights='imagenet', classes=4, activation='sigmoid', aux_params=aux_params)\n```</p>",
      "rawMarkdown": "I have prepared new feature for SMP library that add aux output for models (look at https://github.com/qubvel/segmentation_models.pytorch/tree/models-refactoring)\n\nAccording to this implementation my multi-task network defined as follows:\n```\nclass GatedFPN(smp.FPN):\n\n    def forward(self, x):\n        mask, label = super().forward(x)\n        return dict(\n            mask=mask*label.reshape(*label.size(), 1, 1),\n            label=label,\n        )\n\naux_params = dict(classes=4, activation='sigmoid', dropout=0.5, pooling='avg')\nmodel = GatedFPN('resnet34', encoder_weights='imagenet', classes=4, activation='sigmoid', aux_params=aux_params)\n```",
      "votes": null
    },
    {
      "id": "676661",
      "postDate": "11/19/2019 11:56:26",
      "content": "<p>strong man, congratulationns qubvel</p>",
      "rawMarkdown": "strong man, congratulationns qubvel",
      "votes": null
    },
    {
      "id": "676663",
      "postDate": "11/19/2019 12:04:59",
      "content": "<p>Congratulations <a href=\"/pavel92\">@pavel92</a> , </p>\n\n<p>It's amazing how some people can get such great results with just a few days and a pipeline simples as that.</p>\n\n<p>I was planning to use AdamW as well, but did not have the time, if you used an implementation from a Git repository do you mind sharing?</p>\n\n<p>Also thanks for your amazing Git repositories.</p>",
      "rawMarkdown": "Congratulations @pavel92 , \n\nIt's amazing how some people can get such great results with just a few days and a pipeline simples as that.\n\nI was planning to use AdamW as well, but did not have the time, if you used an implementation from a Git repository do you mind sharing?\n\nAlso thanks for your amazing Git repositories.",
      "votes": null
    },
    {
      "id": "676670",
      "postDate": "11/19/2019 12:08:17",
      "content": "<p>If you use pytorch, then AdamW is already implemented <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.AdamW\">here</a></p>",
      "rawMarkdown": "If you use pytorch, then AdamW is already implemented [here](https://pytorch.org/docs/stable/optim.html#torch.optim.AdamW)",
      "votes": null
    },
    {
      "id": "676882",
      "postDate": "11/19/2019 15:11:56",
      "content": "<p>Thanks <a href=\"/bibek777\">@bibek777</a> , but I mainly use Keras 😄 </p>",
      "rawMarkdown": "Thanks @bibek777 , but I mainly use Keras 😄",
      "votes": null
    },
    {
      "id": "676885",
      "postDate": "11/19/2019 15:15:42",
      "content": "<p>Very! useful. Specialty TTAch ...</p>",
      "rawMarkdown": "Very! useful. Specialty TTAch ...",
      "votes": null
    },
    {
      "id": "676997",
      "postDate": "11/19/2019 17:35:35",
      "content": "<p>Congrats Qubvel. Simple and effective. </p>\n\n<p>In step 2, do you use multi-task networks (two headed models) again? In steps 1 and 2 does threshold refer to \"label threshold\" as in <code>if label &amp;lt; threshold: mask = np.zeros()</code> or \"pixel threshold\" as in <code>mask = (mask &amp;gt; threshold).astype(int)</code>. If label threshold, what is your pixel threshold?</p>\n\n<p>I also added a second head to your GitHub Unet model. I added it after the segmentation output, I should try to add it after the encoder and see if that works better. A picture of my design is <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118086\">here</a></p>",
      "rawMarkdown": "Congrats Qubvel. Simple and effective. \n  \nIn step 2, do you use multi-task networks (two headed models) again? In steps 1 and 2 does threshold refer to \"label threshold\" as in `if label &lt; threshold: mask = np.zeros()` or \"pixel threshold\" as in `mask = (mask &gt; threshold).astype(int)`. If label threshold, what is your pixel threshold?\n\nI also added a second head to your GitHub Unet model. I added it after the segmentation output, I should try to add it after the encoder and see if that works better. A picture of my design is [here][1]\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/118086",
      "votes": null
    },
    {
      "id": "677012",
      "postDate": "11/19/2019 17:54:59",
      "content": "<p>That's cool!\nFast track to medal zone is using <a href=\"https://github.com/qubvel/segmentation_models.pytorch/https://github.com/qubvel/segmentation_models.pytorch/\">segmentation_models</a> lib</p>",
      "rawMarkdown": "That's cool!\nFast track to medal zone is using [segmentation_models](https://github.com/qubvel/segmentation_models.pytorch/https://github.com/qubvel/segmentation_models.pytorch/) lib",
      "votes": null
    },
    {
      "id": "677413",
      "postDate": "11/20/2019 05:41:34",
      "content": "<p>Thanks!\nTo select images for the second stage I use max pixel value from <code>mask</code> output on first stage. Actually, at first, I was think that first stage will be enough and create a submission with only these models, but get not satisfied result. After that I decided to train stage 2 models and replace masks from that submission with new ones (so, may be using label output can give better score, but I did not have time and submissions to test it).\nIn step two I used just simple unet without classification head.</p>",
      "rawMarkdown": "Thanks!\nTo select images for the second stage I use max pixel value from `mask` output on first stage. Actually, at first, I was think that first stage will be enough and create a submission with only these models, but get not satisfied result. After that I decided to train stage 2 models and replace masks from that submission with new ones (so, may be using label output can give better score, but I did not have time and submissions to test it).\nIn step two I used just simple unet without classification head.",
      "votes": null
    },
    {
      "id": "677428",
      "postDate": "11/20/2019 06:18:04",
      "content": "<p>Congrats and thanks for sharing!! Great approach!!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!! Great approach!!",
      "votes": null
    },
    {
      "id": "677467",
      "postDate": "11/20/2019 07:52:54",
      "content": "<p>I use pytorch optimizers</p>",
      "rawMarkdown": "I use pytorch optimizers",
      "votes": null
    },
    {
      "id": "678097",
      "postDate": "11/21/2019 01:42:13",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> AdamW saved me in this competition. Few breakthrough I got in this competition are:\nMy network - 2 X FPN EfficientNetB1 + 2 X FPN EfficientNetB2 + 4 Classifiers(2 each of EfficientNetB2 and EfficientNetB3)\n1. Changed optimizer from Adam to AdamW and did not change anything - jumped from top 48% to top 15%.\n2. Increased min pixel area from 2500 to 10000 - jumped from top 15% to top 8%.\n3. Increased min pixel area from 10000 to 15000 - jumped from top 8% to top 4%.</p>\n\n<p>After that I tried many other methods but did not work.</p>",
      "rawMarkdown": "dimitreoliveira AdamW saved me in this competition. Few breakthrough I got in this competition are:\nMy network - 2 X FPN EfficientNetB1 + 2 X FPN EfficientNetB2 + 4 Classifiers(2 each of EfficientNetB2 and EfficientNetB3)\n1. Changed optimizer from Adam to AdamW and did not change anything - jumped from top 48% to top 15%.\n2. Increased min pixel area from 2500 to 10000 - jumped from top 15% to top 8%.\n3. Increased min pixel area from 10000 to 15000 - jumped from top 8% to top 4%.\n\nAfter that I tried many other methods but did not work.",
      "votes": null
    },
    {
      "id": "678433",
      "postDate": "11/21/2019 11:46:52",
      "content": "<p>Thanks <a href=\"/raghaw\">@raghaw</a> , your results were very impressive, I was planning to experiment with AdamW on the last 2 weeks, but with all the ensembling I ended not having the time.</p>\n\n<p>If you don't mind what LR Scheduler were you using? and if was Keras/Tensorflow, do you mind sharing your source/implementation of AdamW?</p>",
      "rawMarkdown": "Thanks @raghaw , your results were very impressive, I was planning to experiment with AdamW on the last 2 weeks, but with all the ensembling I ended not having the time.\n\nIf you don't mind what LR Scheduler were you using? and if was Keras/Tensorflow, do you mind sharing your source/implementation of AdamW?",
      "votes": null
    },
    {
      "id": "678550",
      "postDate": "11/21/2019 14:31:07",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>, I used ExponentialLR scheduler of pytorch with gamma=0.80 and initial learning rate 5e-4. AdamW is already in pytorch as mentioned by <a href=\"/bibek777\">@bibek777</a> so I did not face any problem. I found AdamW for keras version on github hope it will be helpfull to you.\n<a href=\"https://github.com/GLambard/AdamW_Keras\">https://github.com/GLambard/AdamW_Keras</a></p>",
      "rawMarkdown": "dimitreoliveira, I used ExponentialLR scheduler of pytorch with gamma=0.80 and initial learning rate 5e-4. AdamW is already in pytorch as mentioned by @bibek777 so I did not face any problem. I found AdamW for keras version on github hope it will be helpfull to you.\nhttps://github.com/GLambard/AdamW_Keras",
      "votes": null
    },
    {
      "id": "678573",
      "postDate": "11/21/2019 15:10:07",
      "content": "<p>Thanks again <a href=\"/raghaw\">@raghaw</a> , about the repository it seems that it has <a href=\"https://github.com/GLambard/AdamW_Keras/issues/2\">some issues</a>, but the topic author linked a more reliable source.</p>",
      "rawMarkdown": "Thanks again @raghaw , about the repository it seems that it has [some issues](https://github.com/GLambard/AdamW_Keras/issues/2), but the topic author linked a more reliable source.",
      "votes": null
    },
    {
      "id": "678626",
      "postDate": "11/21/2019 16:53:21",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> I was reading this blog:\n<a href=\"https://www.iprally.com/news/recent-improvements-to-the-adam-optimizer\">https://www.iprally.com/news/recent-improvements-to-the-adam-optimizer</a>\nIn AdamW section they have given hyperlink of its tensorflow implementation. It is already added in tensorflow. check this link.\n<a href=\"https://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/opt/AdamWOptimizer\">https://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/opt/AdamWOptimizer</a></p>",
      "rawMarkdown": "dimitreoliveira I was reading this blog:\nhttps://www.iprally.com/news/recent-improvements-to-the-adam-optimizer\nIn AdamW section they have given hyperlink of its tensorflow implementation. It is already added in tensorflow. check this link.\nhttps://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/opt/AdamWOptimizer",
      "votes": null
    },
    {
      "id": "678661",
      "postDate": "11/21/2019 17:53:45",
      "content": "<p>Nice, thanks a lot!</p>",
      "rawMarkdown": "Nice, thanks a lot!",
      "votes": null
    },
    {
      "id": "678707",
      "postDate": "11/21/2019 18:33:44",
      "content": "<p>Interesting. So you don't use the classification output from Stage 1. Then is the only purpose of the classification head to back propagate label error into the encoder? Or does it server some other purpose that I'm missing?</p>",
      "rawMarkdown": "Interesting. So you don't use the classification output from Stage 1. Then is the only purpose of the classification head to back propagate label error into the encoder? Or does it server some other purpose that I'm missing?",
      "votes": null
    },
    {
      "id": "679407",
      "postDate": "11/22/2019 18:45:20",
      "content": "<p>You are right, but also decoder output is multiplied by 'label' output to generate mask (look at the pic.).</p>",
      "rawMarkdown": "You are right, but also decoder output is multiplied by 'label' output to generate mask (look at the pic.).",
      "votes": null
    },
    {
      "id": "679699",
      "postDate": "11/23/2019 07:34:30",
      "content": "<p>Thank you , very useful</p>",
      "rawMarkdown": "Thank you , very useful",
      "votes": null
    },
    {
      "id": "679712",
      "postDate": "11/23/2019 08:36:07",
      "content": "<p>p.s. merged to master</p>",
      "rawMarkdown": "p.s. merged to master",
      "votes": null
    },
    {
      "id": "680925",
      "postDate": "11/25/2019 12:21:20",
      "content": "<p>Congrats！one question, In step 2 , for every class, you trained a binary-seg model or train a seg model for all class</p>",
      "rawMarkdown": "Congrats！one question, In step 2 , for every class, you trained a binary-seg model or train a seg model for all class",
      "votes": null
    },
    {
      "id": "681940",
      "postDate": "11/26/2019 18:02:55",
      "content": "<p>step 2 - separate binary segmentation model for each class</p>",
      "rawMarkdown": "step 2 - separate binary segmentation model for each class",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 676481,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "11/19/2019 08:13:57",
      "content": "<p>Sweet and Simple!! Thank you for sharing and also thank you for your <code>segmentation_models.pytorch</code> library. Extremely helpful and convenient </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 676482,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "11/19/2019 08:13:57",
      "content": "<p>most of the time i use your library,thanks a lot and congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 676523,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "11/19/2019 09:03:50",
      "content": "<p>Congratulations\nThank you for Sharing your Insights &amp; Approach! <a href=\"/pavel92\">@pavel92</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 676551,
      "author_name": "adish333",
      "author_url": "",
      "post_date": "11/19/2019 09:44:30",
      "content": "<p><a href=\"/pavel92\">@pavel92</a> thanks for your solution writeup. This was my first computer vision competition and I used only your library 'segmentation models' to implement my solution. I must say this library is really helpful for beginners like me. After reading the comments of top solutions, I came to know about the term 'weights averaging from different checkpoints'. Will you please share some resources to understand this concept and may be a sample boiler-plate implementation of this idea?</p>",
      "votes": null,
      "replies": [
        {
          "id": 676573,
          "author_name": "pavel92",
          "author_url": "",
          "post_date": "11/19/2019 10:12:22",
          "content": "<p>```\nimport torch\nfrom collections import OrderedDict\nfrom typing import List</p>\n\n<p>checkpoints_weights_paths: List[str] = ...  # sorted in descending order by score\nmodel: torch.nn.Module = ...</p>\n\n<p>def average_weights(state_dicts: List[dict]):\n    everage_dict = OrderedDict()\n    for k in state_dicts[0].keys():\n        everage_dict[k] = sum([state_dict[k] for state_dict in state_dicts]) / len(state_dicts)\n    return everage_dict</p>\n\n<p>all_weights = [torch.load(path) for path in checkpoints_weights_paths]</p>\n\n<p>best_score = 0\nbest_weights = []</p>\n\n<p>for w in all_weights:\n    current_weights = best_weights + [w]\n    average_dict = average_weights(current_weights)\n    model.load_state_dict(average_dict)\n    score = evaluate_model(model, ...)\n    if score &gt; best_score:\n        best_score = score\n        best_weights.append(w)\n```</p>\n\n<p><a href=\"https://gist.github.com/qubvel/70c3d5e4cddcde731408f478e12ef87b\">https://gist.github.com/qubvel/70c3d5e4cddcde731408f478e12ef87b</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 676608,
          "author_name": "adish333",
          "author_url": "",
          "post_date": "11/19/2019 10:50:36",
          "content": "<p>Thanks a lot :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 676590,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "11/19/2019 10:30:49",
      "content": "<p>Very impressive!!! Congratulations 🎉  and thanks for sharing ❤️ </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 676609,
      "author_name": "detkov",
      "author_url": "",
      "post_date": "11/19/2019 10:51:51",
      "content": "<p>Pavel, can you please leave here some info (maybe links to read) about such multi-task network and how to implement it? </p>",
      "votes": null,
      "replies": [
        {
          "id": 676632,
          "author_name": "pavel92",
          "author_url": "",
          "post_date": "11/19/2019 11:27:40",
          "content": "<p>I have prepared new feature for SMP library that add aux output for models (look at <a href=\"https://github.com/qubvel/segmentation_models.pytorch/tree/models-refactoring\">https://github.com/qubvel/segmentation_models.pytorch/tree/models-refactoring</a>)</p>\n\n<p>According to this implementation my multi-task network defined as follows:\n```\nclass GatedFPN(smp.FPN):</p>\n\n<pre><code>def forward(self, x):\n    mask, label = super().forward(x)\n    return dict(\n        mask=mask*label.reshape(*label.size(), 1, 1),\n        label=label,\n    )\n</code></pre>\n\n<p>aux_params = dict(classes=4, activation='sigmoid', dropout=0.5, pooling='avg')\nmodel = GatedFPN('resnet34', encoder_weights='imagenet', classes=4, activation='sigmoid', aux_params=aux_params)\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 679712,
          "author_name": "pavel92",
          "author_url": "",
          "post_date": "11/23/2019 08:36:07",
          "content": "<p>p.s. merged to master</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 676661,
      "author_name": "jt120lz",
      "author_url": "",
      "post_date": "11/19/2019 11:56:26",
      "content": "<p>strong man, congratulationns qubvel</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 676663,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "11/19/2019 12:04:59",
      "content": "<p>Congratulations <a href=\"/pavel92\">@pavel92</a> , </p>\n\n<p>It's amazing how some people can get such great results with just a few days and a pipeline simples as that.</p>\n\n<p>I was planning to use AdamW as well, but did not have the time, if you used an implementation from a Git repository do you mind sharing?</p>\n\n<p>Also thanks for your amazing Git repositories.</p>",
      "votes": null,
      "replies": [
        {
          "id": 676670,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "11/19/2019 12:08:17",
          "content": "<p>If you use pytorch, then AdamW is already implemented <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.AdamW\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 676882,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "11/19/2019 15:11:56",
          "content": "<p>Thanks <a href=\"/bibek777\">@bibek777</a> , but I mainly use Keras 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677467,
          "author_name": "pavel92",
          "author_url": "",
          "post_date": "11/20/2019 07:52:54",
          "content": "<p>I use pytorch optimizers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678097,
          "author_name": "raghaw",
          "author_url": "",
          "post_date": "11/21/2019 01:42:13",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> AdamW saved me in this competition. Few breakthrough I got in this competition are:\nMy network - 2 X FPN EfficientNetB1 + 2 X FPN EfficientNetB2 + 4 Classifiers(2 each of EfficientNetB2 and EfficientNetB3)\n1. Changed optimizer from Adam to AdamW and did not change anything - jumped from top 48% to top 15%.\n2. Increased min pixel area from 2500 to 10000 - jumped from top 15% to top 8%.\n3. Increased min pixel area from 10000 to 15000 - jumped from top 8% to top 4%.</p>\n\n<p>After that I tried many other methods but did not work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678433,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "11/21/2019 11:46:52",
          "content": "<p>Thanks <a href=\"/raghaw\">@raghaw</a> , your results were very impressive, I was planning to experiment with AdamW on the last 2 weeks, but with all the ensembling I ended not having the time.</p>\n\n<p>If you don't mind what LR Scheduler were you using? and if was Keras/Tensorflow, do you mind sharing your source/implementation of AdamW?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678550,
          "author_name": "raghaw",
          "author_url": "",
          "post_date": "11/21/2019 14:31:07",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>, I used ExponentialLR scheduler of pytorch with gamma=0.80 and initial learning rate 5e-4. AdamW is already in pytorch as mentioned by <a href=\"/bibek777\">@bibek777</a> so I did not face any problem. I found AdamW for keras version on github hope it will be helpfull to you.\n<a href=\"https://github.com/GLambard/AdamW_Keras\">https://github.com/GLambard/AdamW_Keras</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678573,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "11/21/2019 15:10:07",
          "content": "<p>Thanks again <a href=\"/raghaw\">@raghaw</a> , about the repository it seems that it has <a href=\"https://github.com/GLambard/AdamW_Keras/issues/2\">some issues</a>, but the topic author linked a more reliable source.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678626,
          "author_name": "raghaw",
          "author_url": "",
          "post_date": "11/21/2019 16:53:21",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> I was reading this blog:\n<a href=\"https://www.iprally.com/news/recent-improvements-to-the-adam-optimizer\">https://www.iprally.com/news/recent-improvements-to-the-adam-optimizer</a>\nIn AdamW section they have given hyperlink of its tensorflow implementation. It is already added in tensorflow. check this link.\n<a href=\"https://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/opt/AdamWOptimizer\">https://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/opt/AdamWOptimizer</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678661,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "11/21/2019 17:53:45",
          "content": "<p>Nice, thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 676885,
      "author_name": "vingou",
      "author_url": "",
      "post_date": "11/19/2019 15:15:42",
      "content": "<p>Very! useful. Specialty TTAch ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 676997,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/19/2019 17:35:35",
      "content": "<p>Congrats Qubvel. Simple and effective. </p>\n\n<p>In step 2, do you use multi-task networks (two headed models) again? In steps 1 and 2 does threshold refer to \"label threshold\" as in <code>if label &amp;lt; threshold: mask = np.zeros()</code> or \"pixel threshold\" as in <code>mask = (mask &amp;gt; threshold).astype(int)</code>. If label threshold, what is your pixel threshold?</p>\n\n<p>I also added a second head to your GitHub Unet model. I added it after the segmentation output, I should try to add it after the encoder and see if that works better. A picture of my design is <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118086\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 677413,
          "author_name": "pavel92",
          "author_url": "",
          "post_date": "11/20/2019 05:41:34",
          "content": "<p>Thanks!\nTo select images for the second stage I use max pixel value from <code>mask</code> output on first stage. Actually, at first, I was think that first stage will be enough and create a submission with only these models, but get not satisfied result. After that I decided to train stage 2 models and replace masks from that submission with new ones (so, may be using label output can give better score, but I did not have time and submissions to test it).\nIn step two I used just simple unet without classification head.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678707,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/21/2019 18:33:44",
          "content": "<p>Interesting. So you don't use the classification output from Stage 1. Then is the only purpose of the classification head to back propagate label error into the encoder? Or does it server some other purpose that I'm missing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 679407,
          "author_name": "pavel92",
          "author_url": "",
          "post_date": "11/22/2019 18:45:20",
          "content": "<p>You are right, but also decoder output is multiplied by 'label' output to generate mask (look at the pic.).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 680925,
          "author_name": "hustkevin1037",
          "author_url": "",
          "post_date": "11/25/2019 12:21:20",
          "content": "<p>Congrats！one question, In step 2 , for every class, you trained a binary-seg model or train a seg model for all class</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681940,
          "author_name": "pavel92",
          "author_url": "",
          "post_date": "11/26/2019 18:02:55",
          "content": "<p>step 2 - separate binary segmentation model for each class</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 677012,
      "author_name": "mystery",
      "author_url": "",
      "post_date": "11/19/2019 17:54:59",
      "content": "<p>That's cool!\nFast track to medal zone is using <a href=\"https://github.com/qubvel/segmentation_models.pytorch/https://github.com/qubvel/segmentation_models.pytorch/\">segmentation_models</a> lib</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 677428,
      "author_name": "sleepysleeping",
      "author_url": "",
      "post_date": "11/20/2019 06:18:04",
      "content": "<p>Congrats and thanks for sharing!! Great approach!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 679699,
      "author_name": "deividasmataciunas",
      "author_url": "",
      "post_date": "11/23/2019 07:34:30",
      "content": "<p>Thank you , very useful</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676474": "## Easy silver in last days\nI have adopted my pipeline from Severstal Defect Detection and was able to get silver medal in last two days with just 6 submissions, here is a short description of 55th place solution.\n\n2 step pileline\n 1) Multi-task network (classification + segmentation) as classifier to remove empty masks\n 2) Binary segmentation for each class\n \n### 1st step.\nI have trained 5-fold `FPN(resnet34) + aux classfication output` on  480x640 images using `Flip`, `RandomBrightness` as augmentations. Model trained just 6-7 epochs and than starts to overfit, I do nothing with that, just save top 5 checkpoints according to metric.\n\nLoss (segmentation head): bce+dice\nLoss (classification head): bce\nOptimizer: AdamW\nPostprocessing: remove masks less than 10000 pixels\nThresholds: [0.6, 0.6, 0.6, 0.6]\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F940556%2Faec538c6a536cc9dd070697b4e416870%2F2019-11-19%2010-58-58.png?generation=1574150366439784&amp;alt=media)\n\n### 2nd step.\nFor each class trained `2 x Unet(se_resnext50_34x4d)` only on images with masks of that class!\nwith same optimizer, image size and augmentations.\n\nLoss: bce+dice\nThresholds: [0.4, 0.4, 0.4, 0.4]\n\n### Ensemble\nFor all models made checkpoints weights! averaging (+0.005-0.01 on validation).\nModels over each stage have been just averaged with Flip TTA.\n\n### Useful links\n - Segmentation Models: https://github.com/qubvel/segmentation_models.pytorch\n - Test Time Augmentation for PyTorch: https://github.com/qubvel/ttach\n\n**And congratulations to winners!**",
    "676481": "Sweet and Simple!! Thank you for sharing and also thank you for your `segmentation_models.pytorch` library. Extremely helpful and convenient",
    "676482": "most of the time i use your library,thanks a lot and congratulations",
    "676523": "Congratulations\nThank you for Sharing your Insights &amp; Approach! @pavel92",
    "676551": "pavel92 thanks for your solution writeup. This was my first computer vision competition and I used only your library 'segmentation models' to implement my solution. I must say this library is really helpful for beginners like me. After reading the comments of top solutions, I came to know about the term 'weights averaging from different checkpoints'. Will you please share some resources to understand this concept and may be a sample boiler-plate implementation of this idea?",
    "676573": "```\nimport torch\nfrom collections import OrderedDict\nfrom typing import List\n\ncheckpoints_weights_paths: List[str] = ...  # sorted in descending order by score\nmodel: torch.nn.Module = ...\n\n\ndef average_weights(state_dicts: List[dict]):\n    everage_dict = OrderedDict()\n    for k in state_dicts[0].keys():\n        everage_dict[k] = sum([state_dict[k] for state_dict in state_dicts]) / len(state_dicts)\n    return everage_dict\n\n\nall_weights = [torch.load(path) for path in checkpoints_weights_paths]\n\nbest_score = 0\nbest_weights = []\n\nfor w in all_weights:\n    current_weights = best_weights + [w]\n    average_dict = average_weights(current_weights)\n    model.load_state_dict(average_dict)\n    score = evaluate_model(model, ...)\n    if score &gt; best_score:\n        best_score = score\n        best_weights.append(w)\n```\n\nhttps://gist.github.com/qubvel/70c3d5e4cddcde731408f478e12ef87b",
    "676590": "Very impressive!!! Congratulations 🎉  and thanks for sharing ❤️",
    "676608": "Thanks a lot :)",
    "676609": "Pavel, can you please leave here some info (maybe links to read) about such multi-task network and how to implement it?",
    "676632": "I have prepared new feature for SMP library that add aux output for models (look at https://github.com/qubvel/segmentation_models.pytorch/tree/models-refactoring)\n\nAccording to this implementation my multi-task network defined as follows:\n```\nclass GatedFPN(smp.FPN):\n\n    def forward(self, x):\n        mask, label = super().forward(x)\n        return dict(\n            mask=mask*label.reshape(*label.size(), 1, 1),\n            label=label,\n        )\n\naux_params = dict(classes=4, activation='sigmoid', dropout=0.5, pooling='avg')\nmodel = GatedFPN('resnet34', encoder_weights='imagenet', classes=4, activation='sigmoid', aux_params=aux_params)\n```",
    "676661": "strong man, congratulationns qubvel",
    "676663": "Congratulations @pavel92 , \n\nIt's amazing how some people can get such great results with just a few days and a pipeline simples as that.\n\nI was planning to use AdamW as well, but did not have the time, if you used an implementation from a Git repository do you mind sharing?\n\nAlso thanks for your amazing Git repositories.",
    "676670": "If you use pytorch, then AdamW is already implemented [here](https://pytorch.org/docs/stable/optim.html#torch.optim.AdamW)",
    "676882": "Thanks @bibek777 , but I mainly use Keras 😄",
    "676885": "Very! useful. Specialty TTAch ...",
    "676997": "Congrats Qubvel. Simple and effective. \n  \nIn step 2, do you use multi-task networks (two headed models) again? In steps 1 and 2 does threshold refer to \"label threshold\" as in `if label &lt; threshold: mask = np.zeros()` or \"pixel threshold\" as in `mask = (mask &gt; threshold).astype(int)`. If label threshold, what is your pixel threshold?\n\nI also added a second head to your GitHub Unet model. I added it after the segmentation output, I should try to add it after the encoder and see if that works better. A picture of my design is [here][1]\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/118086",
    "677012": "That's cool!\nFast track to medal zone is using [segmentation_models](https://github.com/qubvel/segmentation_models.pytorch/https://github.com/qubvel/segmentation_models.pytorch/) lib",
    "677413": "Thanks!\nTo select images for the second stage I use max pixel value from `mask` output on first stage. Actually, at first, I was think that first stage will be enough and create a submission with only these models, but get not satisfied result. After that I decided to train stage 2 models and replace masks from that submission with new ones (so, may be using label output can give better score, but I did not have time and submissions to test it).\nIn step two I used just simple unet without classification head.",
    "677428": "Congrats and thanks for sharing!! Great approach!!",
    "677467": "I use pytorch optimizers",
    "678097": "dimitreoliveira AdamW saved me in this competition. Few breakthrough I got in this competition are:\nMy network - 2 X FPN EfficientNetB1 + 2 X FPN EfficientNetB2 + 4 Classifiers(2 each of EfficientNetB2 and EfficientNetB3)\n1. Changed optimizer from Adam to AdamW and did not change anything - jumped from top 48% to top 15%.\n2. Increased min pixel area from 2500 to 10000 - jumped from top 15% to top 8%.\n3. Increased min pixel area from 10000 to 15000 - jumped from top 8% to top 4%.\n\nAfter that I tried many other methods but did not work.",
    "678433": "Thanks @raghaw , your results were very impressive, I was planning to experiment with AdamW on the last 2 weeks, but with all the ensembling I ended not having the time.\n\nIf you don't mind what LR Scheduler were you using? and if was Keras/Tensorflow, do you mind sharing your source/implementation of AdamW?",
    "678550": "dimitreoliveira, I used ExponentialLR scheduler of pytorch with gamma=0.80 and initial learning rate 5e-4. AdamW is already in pytorch as mentioned by @bibek777 so I did not face any problem. I found AdamW for keras version on github hope it will be helpfull to you.\nhttps://github.com/GLambard/AdamW_Keras",
    "678573": "Thanks again @raghaw , about the repository it seems that it has [some issues](https://github.com/GLambard/AdamW_Keras/issues/2), but the topic author linked a more reliable source.",
    "678626": "dimitreoliveira I was reading this blog:\nhttps://www.iprally.com/news/recent-improvements-to-the-adam-optimizer\nIn AdamW section they have given hyperlink of its tensorflow implementation. It is already added in tensorflow. check this link.\nhttps://www.tensorflow.org/versions/r1.15/api_docs/python/tf/contrib/opt/AdamWOptimizer",
    "678661": "Nice, thanks a lot!",
    "678707": "Interesting. So you don't use the classification output from Stage 1. Then is the only purpose of the classification head to back propagate label error into the encoder? Or does it server some other purpose that I'm missing?",
    "679407": "You are right, but also decoder output is multiplied by 'label' output to generate mask (look at the pic.).",
    "679699": "Thank you , very useful",
    "679712": "p.s. merged to master",
    "680925": "Congrats！one question, In step 2 , for every class, you trained a binary-seg model or train a seg model for all class",
    "681940": "step 2 - separate binary segmentation model for each class"
  },
  "source": "meta"
}