{
  "id": 118016,
  "title": "4th Place Solution: Stabilizing Convergence in Understanding Clouds",
  "url": "/competitions/understanding_cloud_organization/discussion/118016",
  "author_name": "Ching-Loong Seow",
  "post_date": "2019-11-19T07:38:49.314000",
  "votes": 29,
  "comment_count": 13,
  "views": 0,
  "content": "<p><img src=\"https://storage.googleapis.com/kaggle-media/competitions/MaxPlanck/Teaser_AnimationwLabels.gif\"></p>\n\n<p>First of all, I would like to express my gratitude and appreciation to the following parties for organizing such a great competition:\n - <a href=\"https://www.kaggle.com\">Kaggle</a>\n - [Max Planck Institute for Meteorology] (<a href=\"https://www.kaggle.com/MaxPlanckInstitute\">https://www.kaggle.com/MaxPlanckInstitute</a>)</p>\n\n<p>Besides, I would like to use this opportunity to thank my fellow kagglers for all the insightful posts in the discussion forum of various competitions. I have also learned a lot of stuffs and gained knowledge by reading from past solutions. There is a good thriving culture of idea sharing and contributions which I have found in every corner of Kaggle and I loved to be part of it.</p>\n\n<h2>The Main Challenge</h2>\n\n<p>The challenge that I have faced initially in this competition is that many models of different architectures tend to overfit easily in early training stage especially for the larger and deeper models, such as SE-ResNext-101 and EfficientNet B5-B7. I have suspected the culprit might be due to the labels given are too noisy and this increases the tendency of model to be overfitted to the noises of training data, as the labels were determined by the union of the areas marked by all annotators. Also, the shape of the label provided is rectangular instead of the exact shape fitted to the boundary of cloud patterns. I understand the reasons behind <a href=\"https://arxiv.org/pdf/1906.01906.pdf\">these decisions made by the competition host</a>, and here goes my whole journey of this competition, which is revolved around stabilizing the convergence in training models.</p>\n\n<h2>Solution Overview</h2>\n\n<p>My solution for this competition is mainly comprised of the followings:</p>\n\n<ul>\n<li><p><strong>Pure segmentation models without false positive classifier</strong>\nAfter reaching public LB 0.6752 with segmentation model, I've trained a few classifiers using Resnet34, SE-ResNext-50 and EfficientNet-B4 but the performance is pretty unstable (+/- 0.003 ~ 0.010) in local cross validations of 10 folds. Thus, I discarded the classifiers and decided to stick with segmentation models.</p></li>\n<li><p><strong>Network Architectures</strong>\nI've used the awesome implementations of various models from <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">segmentation_models.pytorch</a>, <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">pretrained-models.pytorch\n</a>, <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\">EfficientNet-PyTorch</a> and <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/115787#671393\">Resnet34-ASPP</a> from <a href=\"/hengck23\">@hengck23</a>. My final ensemble used 7 folds of EfficientNet-B4-FPN and 3 folds of Resnet34-ASPP as they have better performance and more stable in error convergence in my case after running rounds of experiments using various network architectures.</p></li>\n<li><strong>RAdam Optimizer</strong>\nRAdam helped to stabilize training error convergence as it is less sensitive to learning rate change in my case, thus minimizing the variance.</li>\n<li><strong>Flat threshold of 0.4 for all classes</strong>\nThreshold of 0.4 yielded the highest cross validation DICE score when compared in the range of [0.4, 0.5, 0.6], no further fine-tuning of threshold is done.</li>\n<li><strong>Minimum segmentation mask size of 5000 pixels for all classes</strong>\nThe mask size threshold is set to be just high enough to filter out noises, no any other post-processing methods is used.</li>\n<li><strong>Input Size</strong>\nDownsized from the raw size of 1400 x 2100 to 700 x 1050. After applying augmentations, it is downsized again from 700 x 1050 to 384 x 576.</li>\n<li><strong>Augmentations used in training</strong>\n<ul><li>horizontal flip</li>\n<li>vertical flip</li>\n<li>random shift, scale and rotate</li></ul></li>\n<li><p><strong>Test-time Augmentations (TTA)</strong>:</p>\n\n<ul><li>horizontal flip</li>\n<li>vertical flip</li>\n<li>180 degree flip (horizontal + vertical flip)</li></ul></li>\n<li><p><strong>Pseudo-labeling</strong>\nI've used two approach for pseudo-labeling, one in which only the confident pseudo-labels are selected and use in training, another in which pseudo-labels are generated from all the test data. In my case, the model training performance of using pseudo-labels from all test data is more robust and stable in terms of error convergence and achieve higher DICE score.</p></li>\n<li><strong>Ensemble with equal weight averaging</strong>  </li>\n<li><strong>Trained initially with BCE Loss, fine-tuned with Symmetric Lovasz Loss originated from this <a href=\"https://arxiv.org/abs/1705.08790\">paper</a> and modified by <a href=\"/tugstugi\">@tugstugi</a></strong>\nBelow is the PyTorch implementation code of Symmetric Lovasz Loss:\n<code>\ndef symmetric_lovasz_loss(outputs, targets):\nbatch_size, num_class, H, W = outputs.shape\noutputs = outputs.contiguous().view(-1, H, W)\ntargets = targets.contiguous().view(-1, H, W)\nreturn (lovasz_hinge(outputs, targets) \n&lt;ul&gt;&lt;li&gt;lovasz_hinge(-outputs, 1 - targets))/2\n</code></li></ul>\n<li><strong>GPU used</strong>\n<ul><li>2 x RTX2080Ti</li></ul></li>\n\n\n<h2>Conclusion</h2>\n\n<p>I think local cross validation is very important and we should always believe in it despite the score showed on Public LB might be lower or higher as it is only computed based on a minor subset of the test dataset. Besides, the <strong>combination of RAdam optimizer, Symmetric Lovasz Loss, Pseudo-labeling and ensembling</strong> has helped significantly in stabilizing the convergence and improving the score.\n<br><br>\nThanks for reading! See you again in upcoming competitions.</p>",
  "messages": [
    {
      "id": 676452,
      "postDate": "2019-11-19T07:38:49.313Z",
      "content": "<p><img src=\"https://storage.googleapis.com/kaggle-media/competitions/MaxPlanck/Teaser_AnimationwLabels.gif\"></p>\n\n<p>First of all, I would like to express my gratitude and appreciation to the following parties for organizing such a great competition:\n - <a href=\"https://www.kaggle.com\">Kaggle</a>\n - [Max Planck Institute for Meteorology] (<a href=\"https://www.kaggle.com/MaxPlanckInstitute\">https://www.kaggle.com/MaxPlanckInstitute</a>)</p>\n\n<p>Besides, I would like to use this opportunity to thank my fellow kagglers for all the insightful posts in the discussion forum of various competitions. I have also learned a lot of stuffs and gained knowledge by reading from past solutions. There is a good thriving culture of idea sharing and contributions which I have found in every corner of Kaggle and I loved to be part of it.</p>\n\n<h2>The Main Challenge</h2>\n\n<p>The challenge that I have faced initially in this competition is that many models of different architectures tend to overfit easily in early training stage especially for the larger and deeper models, such as SE-ResNext-101 and EfficientNet B5-B7. I have suspected the culprit might be due to the labels given are too noisy and this increases the tendency of model to be overfitted to the noises of training data, as the labels were determined by the union of the areas marked by all annotators. Also, the shape of the label provided is rectangular instead of the exact shape fitted to the boundary of cloud patterns. I understand the reasons behind <a href=\"https://arxiv.org/pdf/1906.01906.pdf\">these decisions made by the competition host</a>, and here goes my whole journey of this competition, which is revolved around stabilizing the convergence in training models.</p>\n\n<h2>Solution Overview</h2>\n\n<p>My solution for this competition is mainly comprised of the followings:</p>\n\n<ul>\n<li><p><strong>Pure segmentation models without false positive classifier</strong>\nAfter reaching public LB 0.6752 with segmentation model, I've trained a few classifiers using Resnet34, SE-ResNext-50 and EfficientNet-B4 but the performance is pretty unstable (+/- 0.003 ~ 0.010) in local cross validations of 10 folds. Thus, I discarded the classifiers and decided to stick with segmentation models.</p></li>\n<li><p><strong>Network Architectures</strong>\nI've used the awesome implementations of various models from <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">segmentation_models.pytorch</a>, <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">pretrained-models.pytorch\n</a>, <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\">EfficientNet-PyTorch</a> and <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/115787#671393\">Resnet34-ASPP</a> from <a href=\"/hengck23\">@hengck23</a>. My final ensemble used 7 folds of EfficientNet-B4-FPN and 3 folds of Resnet34-ASPP as they have better performance and more stable in error convergence in my case after running rounds of experiments using various network architectures.</p></li>\n<li><strong>RAdam Optimizer</strong>\nRAdam helped to stabilize training error convergence as it is less sensitive to learning rate change in my case, thus minimizing the variance.</li>\n<li><strong>Flat threshold of 0.4 for all classes</strong>\nThreshold of 0.4 yielded the highest cross validation DICE score when compared in the range of [0.4, 0.5, 0.6], no further fine-tuning of threshold is done.</li>\n<li><strong>Minimum segmentation mask size of 5000 pixels for all classes</strong>\nThe mask size threshold is set to be just high enough to filter out noises, no any other post-processing methods is used.</li>\n<li><strong>Input Size</strong>\nDownsized from the raw size of 1400 x 2100 to 700 x 1050. After applying augmentations, it is downsized again from 700 x 1050 to 384 x 576.</li>\n<li><strong>Augmentations used in training</strong>\n<ul><li>horizontal flip</li>\n<li>vertical flip</li>\n<li>random shift, scale and rotate</li></ul></li>\n<li><p><strong>Test-time Augmentations (TTA)</strong>:</p>\n\n<ul><li>horizontal flip</li>\n<li>vertical flip</li>\n<li>180 degree flip (horizontal + vertical flip)</li></ul></li>\n<li><p><strong>Pseudo-labeling</strong>\nI've used two approach for pseudo-labeling, one in which only the confident pseudo-labels are selected and use in training, another in which pseudo-labels are generated from all the test data. In my case, the model training performance of using pseudo-labels from all test data is more robust and stable in terms of error convergence and achieve higher DICE score.</p></li>\n<li><strong>Ensemble with equal weight averaging</strong>  </li>\n<li><strong>Trained initially with BCE Loss, fine-tuned with Symmetric Lovasz Loss originated from this <a href=\"https://arxiv.org/abs/1705.08790\">paper</a> and modified by <a href=\"/tugstugi\">@tugstugi</a></strong>\nBelow is the PyTorch implementation code of Symmetric Lovasz Loss:\n<code>\ndef symmetric_lovasz_loss(outputs, targets):\nbatch_size, num_class, H, W = outputs.shape\noutputs = outputs.contiguous().view(-1, H, W)\ntargets = targets.contiguous().view(-1, H, W)\nreturn (lovasz_hinge(outputs, targets) \n&lt;ul&gt;&lt;li&gt;lovasz_hinge(-outputs, 1 - targets))/2\n</code></li></ul>\n<li><strong>GPU used</strong>\n<ul><li>2 x RTX2080Ti</li></ul></li>\n\n\n<h2>Conclusion</h2>\n\n<p>I think local cross validation is very important and we should always believe in it despite the score showed on Public LB might be lower or higher as it is only computed based on a minor subset of the test dataset. Besides, the <strong>combination of RAdam optimizer, Symmetric Lovasz Loss, Pseudo-labeling and ensembling</strong> has helped significantly in stabilizing the convergence and improving the score.\n<br><br>\nThanks for reading! See you again in upcoming competitions.</p>",
      "rawMarkdown": "<img src=\"https://storage.googleapis.com/kaggle-media/competitions/MaxPlanck/Teaser_AnimationwLabels.gif\">\n\nFirst of all, I would like to express my gratitude and appreciation to the following parties for organizing such a great competition:\n - [Kaggle](https://www.kaggle.com)\n - [Max Planck Institute for Meteorology] (https://www.kaggle.com/MaxPlanckInstitute)\n\nBesides, I would like to use this opportunity to thank my fellow kagglers for all the insightful posts in the discussion forum of various competitions. I have also learned a lot of stuffs and gained knowledge by reading from past solutions. There is a good thriving culture of idea sharing and contributions which I have found in every corner of Kaggle and I loved to be part of it.\n\n## The Main Challenge\nThe challenge that I have faced initially in this competition is that many models of different architectures tend to overfit easily in early training stage especially for the larger and deeper models, such as SE-ResNext-101 and EfficientNet B5-B7. I have suspected the culprit might be due to the labels given are too noisy and this increases the tendency of model to be overfitted to the noises of training data, as the labels were determined by the union of the areas marked by all annotators. Also, the shape of the label provided is rectangular instead of the exact shape fitted to the boundary of cloud patterns. I understand the reasons behind [these decisions made by the competition host](https://arxiv.org/pdf/1906.01906.pdf), and here goes my whole journey of this competition, which is revolved around stabilizing the convergence in training models.\n \n## Solution Overview\nMy solution for this competition is mainly comprised of the followings:\n\n- **Pure segmentation models without false positive classifier**\nAfter reaching public LB 0.6752 with segmentation model, I've trained a few classifiers using Resnet34, SE-ResNext-50 and EfficientNet-B4 but the performance is pretty unstable (+/- 0.003 ~ 0.010) in local cross validations of 10 folds. Thus, I discarded the classifiers and decided to stick with segmentation models.\n\n- **Network Architectures**\nI've used the awesome implementations of various models from [segmentation_models.pytorch](https://github.com/qubvel/segmentation_models.pytorch), [pretrained-models.pytorch\n](https://github.com/Cadene/pretrained-models.pytorch), [EfficientNet-PyTorch](https://github.com/lukemelas/EfficientNet-PyTorch) and [Resnet34-ASPP](https://www.kaggle.com/c/understanding_cloud_organization/discussion/115787#671393) from @hengck23. My final ensemble used 7 folds of EfficientNet-B4-FPN and 3 folds of Resnet34-ASPP as they have better performance and more stable in error convergence in my case after running rounds of experiments using various network architectures.\n- **RAdam Optimizer**\nRAdam helped to stabilize training error convergence as it is less sensitive to learning rate change in my case, thus minimizing the variance.\n- **Flat threshold of 0.4 for all classes**\nThreshold of 0.4 yielded the highest cross validation DICE score when compared in the range of [0.4, 0.5, 0.6], no further fine-tuning of threshold is done.\n- **Minimum segmentation mask size of 5000 pixels for all classes**\nThe mask size threshold is set to be just high enough to filter out noises, no any other post-processing methods is used.\n- **Input Size**\nDownsized from the raw size of 1400 x 2100 to 700 x 1050. After applying augmentations, it is downsized again from 700 x 1050 to 384 x 576.\n- **Augmentations used in training**\n    - horizontal flip\n    - vertical flip\n    - random shift, scale and rotate\n- **Test-time Augmentations (TTA)**:\n  - horizontal flip\n  - vertical flip\n  - 180 degree flip (horizontal + vertical flip)\n\n- **Pseudo-labeling**\n I've used two approach for pseudo-labeling, one in which only the confident pseudo-labels are selected and use in training, another in which pseudo-labels are generated from all the test data. In my case, the model training performance of using pseudo-labels from all test data is more robust and stable in terms of error convergence and achieve higher DICE score.\n- **Ensemble with equal weight averaging**  \n- **Trained initially with BCE Loss, fine-tuned with Symmetric Lovasz Loss originated from this [paper](https://arxiv.org/abs/1705.08790) and modified by @tugstugi**\nBelow is the PyTorch implementation code of Symmetric Lovasz Loss:\n```\ndef symmetric_lovasz_loss(outputs, targets):\n    batch_size, num_class, H, W = outputs.shape\n    outputs = outputs.contiguous().view(-1, H, W)\n    targets = targets.contiguous().view(-1, H, W)\n    return (lovasz_hinge(outputs, targets) \n      + lovasz_hinge(-outputs, 1 - targets))/2\n```\n- **GPU used**\n  - 2 x RTX2080Ti\n\n## Conclusion\nI think local cross validation is very important and we should always believe in it despite the score showed on Public LB might be lower or higher as it is only computed based on a minor subset of the test dataset. Besides, the **combination of RAdam optimizer, Symmetric Lovasz Loss, Pseudo-labeling and ensembling** has helped significantly in stabilizing the convergence and improving the score.\n<br><br>\nThanks for reading! See you again in upcoming competitions.",
      "votes": 29
    },
    {
      "id": 677205,
      "postDate": "2019-11-19T23:19:43.817Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "replies": [
        {
          "id": 677333,
          "postDate": "2019-11-20T03:51:59.877Z",
          "content": "<p>Thanks! <a href=\"/corochann\">@corochann</a> </p>",
          "rawMarkdown": "Thanks! @corochann "
        }
      ]
    },
    {
      "id": 677046,
      "postDate": "2019-11-19T18:27:05.420Z",
      "content": "<p>Congratulations. Amazing solo performance and Gold finish. </p>\n\n<p>Your result is quite impressive. It seems that you built just segmentation model(s) and trained without modifying the training data and achieved one of the best solutions. Many others optimized two pipelines. One to predict empty mask and one to predict shape of mask.</p>",
      "rawMarkdown": "Congratulations. Amazing solo performance and Gold finish. \n\nYour result is quite impressive. It seems that you built just segmentation model(s) and trained without modifying the training data and achieved one of the best solutions. Many others optimized two pipelines. One to predict empty mask and one to predict shape of mask.",
      "replies": [
        {
          "id": 677332,
          "postDate": "2019-11-20T03:51:42.637Z",
          "content": "<p>Thanks Chris <a href=\"/cdeotte\">@cdeotte</a>! I have learned a lot from your posts too </p>",
          "rawMarkdown": "Thanks Chris @cdeotte! I have learned a lot from your posts too "
        }
      ]
    },
    {
      "id": 676683,
      "postDate": "2019-11-19T12:20:29.630Z",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing.",
      "replies": [
        {
          "id": 676763,
          "postDate": "2019-11-19T13:41:11.630Z",
          "content": "<p>Thanks! <a href=\"/titericz\">@titericz</a> </p>",
          "rawMarkdown": "Thanks! @titericz "
        }
      ]
    },
    {
      "id": 676595,
      "postDate": "2019-11-19T10:35:33.790Z",
      "content": "<p>Thanks for sharing! And congratulations 🎉 🎉 🎉 </p>",
      "rawMarkdown": "Thanks for sharing! And congratulations 🎉 🎉 🎉 ",
      "replies": [
        {
          "id": 676678,
          "postDate": "2019-11-19T12:16:56.780Z",
          "content": "<p>Thanks! Congrats for getting a silver medal too <a href=\"/phunghieu\">@phunghieu</a> </p>",
          "rawMarkdown": "Thanks! Congrats for getting a silver medal too @phunghieu ",
          "votes": 1
        }
      ]
    },
    {
      "id": 676577,
      "postDate": "2019-11-19T10:21:41.597Z",
      "content": "<p>Congratulations ;)</p>",
      "rawMarkdown": "Congratulations ;)",
      "replies": [
        {
          "id": 676679,
          "postDate": "2019-11-19T12:17:21.233Z",
          "content": "<p>Thanks! <a href=\"/hanjoonchoe\">@hanjoonchoe</a> </p>",
          "rawMarkdown": "Thanks! @hanjoonchoe "
        }
      ]
    },
    {
      "id": 676526,
      "postDate": "2019-11-19T09:08:26.633Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 676681,
          "postDate": "2019-11-19T12:18:07.347Z",
          "content": "<p>Thanks! <a href=\"/veeralakrishna\">@veeralakrishna</a> </p>",
          "rawMarkdown": "Thanks! @veeralakrishna "
        }
      ]
    },
    {
      "id": 676868,
      "postDate": "2019-11-19T15:02:27.220Z",
      "content": "<p>congratulations and thanks for sharing</p>",
      "rawMarkdown": "congratulations and thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 677205,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-11-19T23:19:43.817000",
      "content": "<p>Congrats!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677333,
          "author_name": "Ching-Loong Seow",
          "author_url": "",
          "post_date": "2019-11-20T03:51:59.877000",
          "content": "<p>Thanks! <a href=\"/corochann\">@corochann</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 677046,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-11-19T18:27:05.420000",
      "content": "<p>Congratulations. Amazing solo performance and Gold finish. </p>\n\n<p>Your result is quite impressive. It seems that you built just segmentation model(s) and trained without modifying the training data and achieved one of the best solutions. Many others optimized two pipelines. One to predict empty mask and one to predict shape of mask.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677332,
          "author_name": "Ching-Loong Seow",
          "author_url": "",
          "post_date": "2019-11-20T03:51:42.637000",
          "content": "<p>Thanks Chris <a href=\"/cdeotte\">@cdeotte</a>! I have learned a lot from your posts too </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676683,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-11-19T12:20:29.630000",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676763,
          "author_name": "Ching-Loong Seow",
          "author_url": "",
          "post_date": "2019-11-19T13:41:11.630000",
          "content": "<p>Thanks! <a href=\"/titericz\">@titericz</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676595,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2019-11-19T10:35:33.790000",
      "content": "<p>Thanks for sharing! And congratulations 🎉 🎉 🎉 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 676678,
          "author_name": "Ching-Loong Seow",
          "author_url": "",
          "post_date": "2019-11-19T12:16:56.780000",
          "content": "<p>Thanks! Congrats for getting a silver medal too <a href=\"/phunghieu\">@phunghieu</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676577,
      "author_name": "Hanjoon Choe",
      "author_url": "",
      "post_date": "2019-11-19T10:21:41.597000",
      "content": "<p>Congratulations ;)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676679,
          "author_name": "Ching-Loong Seow",
          "author_url": "",
          "post_date": "2019-11-19T12:17:21.233000",
          "content": "<p>Thanks! <a href=\"/hanjoonchoe\">@hanjoonchoe</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676526,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-19T09:08:26.633000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 676681,
          "author_name": "Ching-Loong Seow",
          "author_url": "",
          "post_date": "2019-11-19T12:18:07.347000",
          "content": "<p>Thanks! <a href=\"/veeralakrishna\">@veeralakrishna</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676868,
      "author_name": "liuze",
      "author_url": "",
      "post_date": "2019-11-19T15:02:27.220000",
      "content": "<p>congratulations and thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676452": "<img src=\"https://storage.googleapis.com/kaggle-media/competitions/MaxPlanck/Teaser_AnimationwLabels.gif\">\n\nFirst of all, I would like to express my gratitude and appreciation to the following parties for organizing such a great competition:\n - [Kaggle](https://www.kaggle.com)\n - [Max Planck Institute for Meteorology] (https://www.kaggle.com/MaxPlanckInstitute)\n\nBesides, I would like to use this opportunity to thank my fellow kagglers for all the insightful posts in the discussion forum of various competitions. I have also learned a lot of stuffs and gained knowledge by reading from past solutions. There is a good thriving culture of idea sharing and contributions which I have found in every corner of Kaggle and I loved to be part of it.\n\n## The Main Challenge\nThe challenge that I have faced initially in this competition is that many models of different architectures tend to overfit easily in early training stage especially for the larger and deeper models, such as SE-ResNext-101 and EfficientNet B5-B7. I have suspected the culprit might be due to the labels given are too noisy and this increases the tendency of model to be overfitted to the noises of training data, as the labels were determined by the union of the areas marked by all annotators. Also, the shape of the label provided is rectangular instead of the exact shape fitted to the boundary of cloud patterns. I understand the reasons behind [these decisions made by the competition host](https://arxiv.org/pdf/1906.01906.pdf), and here goes my whole journey of this competition, which is revolved around stabilizing the convergence in training models.\n \n## Solution Overview\nMy solution for this competition is mainly comprised of the followings:\n\n- **Pure segmentation models without false positive classifier**\nAfter reaching public LB 0.6752 with segmentation model, I've trained a few classifiers using Resnet34, SE-ResNext-50 and EfficientNet-B4 but the performance is pretty unstable (+/- 0.003 ~ 0.010) in local cross validations of 10 folds. Thus, I discarded the classifiers and decided to stick with segmentation models.\n\n- **Network Architectures**\nI've used the awesome implementations of various models from [segmentation_models.pytorch](https://github.com/qubvel/segmentation_models.pytorch), [pretrained-models.pytorch\n](https://github.com/Cadene/pretrained-models.pytorch), [EfficientNet-PyTorch](https://github.com/lukemelas/EfficientNet-PyTorch) and [Resnet34-ASPP](https://www.kaggle.com/c/understanding_cloud_organization/discussion/115787#671393) from @hengck23. My final ensemble used 7 folds of EfficientNet-B4-FPN and 3 folds of Resnet34-ASPP as they have better performance and more stable in error convergence in my case after running rounds of experiments using various network architectures.\n- **RAdam Optimizer**\nRAdam helped to stabilize training error convergence as it is less sensitive to learning rate change in my case, thus minimizing the variance.\n- **Flat threshold of 0.4 for all classes**\nThreshold of 0.4 yielded the highest cross validation DICE score when compared in the range of [0.4, 0.5, 0.6], no further fine-tuning of threshold is done.\n- **Minimum segmentation mask size of 5000 pixels for all classes**\nThe mask size threshold is set to be just high enough to filter out noises, no any other post-processing methods is used.\n- **Input Size**\nDownsized from the raw size of 1400 x 2100 to 700 x 1050. After applying augmentations, it is downsized again from 700 x 1050 to 384 x 576.\n- **Augmentations used in training**\n    - horizontal flip\n    - vertical flip\n    - random shift, scale and rotate\n- **Test-time Augmentations (TTA)**:\n  - horizontal flip\n  - vertical flip\n  - 180 degree flip (horizontal + vertical flip)\n\n- **Pseudo-labeling**\n I've used two approach for pseudo-labeling, one in which only the confident pseudo-labels are selected and use in training, another in which pseudo-labels are generated from all the test data. In my case, the model training performance of using pseudo-labels from all test data is more robust and stable in terms of error convergence and achieve higher DICE score.\n- **Ensemble with equal weight averaging**  \n- **Trained initially with BCE Loss, fine-tuned with Symmetric Lovasz Loss originated from this [paper](https://arxiv.org/abs/1705.08790) and modified by @tugstugi**\nBelow is the PyTorch implementation code of Symmetric Lovasz Loss:\n```\ndef symmetric_lovasz_loss(outputs, targets):\n    batch_size, num_class, H, W = outputs.shape\n    outputs = outputs.contiguous().view(-1, H, W)\n    targets = targets.contiguous().view(-1, H, W)\n    return (lovasz_hinge(outputs, targets) \n      + lovasz_hinge(-outputs, 1 - targets))/2\n```\n- **GPU used**\n  - 2 x RTX2080Ti\n\n## Conclusion\nI think local cross validation is very important and we should always believe in it despite the score showed on Public LB might be lower or higher as it is only computed based on a minor subset of the test dataset. Besides, the **combination of RAdam optimizer, Symmetric Lovasz Loss, Pseudo-labeling and ensembling** has helped significantly in stabilizing the convergence and improving the score.\n<br><br>\nThanks for reading! See you again in upcoming competitions.",
    "677205": "Congrats!",
    "677046": "Congratulations. Amazing solo performance and Gold finish. \n\nYour result is quite impressive. It seems that you built just segmentation model(s) and trained without modifying the training data and achieved one of the best solutions. Many others optimized two pipelines. One to predict empty mask and one to predict shape of mask.",
    "676683": "Congratulations and thanks for sharing.",
    "676595": "Thanks for sharing! And congratulations 🎉 🎉 🎉 ",
    "676577": "Congratulations ;)",
    "676526": "",
    "676868": "congratulations and thanks for sharing"
  }
}