{
  "id": 118072,
  "title": "13th place solution",
  "url": "/competitions/understanding_cloud_organization/writeups/tom-13th-place-solution",
  "author_name": "",
  "post_date": "2019-11-19T13:17:34.748936400Z",
  "votes": 15,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks Kaggle and Max Planck Institute for this interesting competition and congrats to all the winners!  Here is a brief summary of my solution (Public 0.67698, Private 0.66713).</p>\n\n<h2>Preprocessing</h2>\n\n<ul>\n<li>exclude bad images (removed 13 images)</li>\n<li>resize image size to (320, 512)</li>\n</ul>\n\n<h2>Augmentations (by Albumentations)</h2>\n\n<ul>\n<li>gamma (limit=(50,100), p=0.5)</li>\n<li>brightness (limit=0.2, p=0.5)</li>\n<li>shift (limit=0.2, border_mode=0, p=0.5)</li>\n<li>rotation (limit=30deg, border_mode=0, p=0.5)</li>\n<li>horizontal flip (p=0.5)</li>\n<li>vertical flip (p=0.5)</li>\n</ul>\n\n<h2>Validation</h2>\n\n<ul>\n<li>StratifiedKFold for the number of empty masks</li>\n</ul>\n\n<h2>Model (ensemble of 7 models x 5folds)</h2>\n\n<ol>\n<li>UNet-ResNet34 + CBAM + Hypercolumns</li>\n<li>same as 1. but with other seed</li>\n<li>UNet-ResNet18 + CBAM + Hypercolumns</li>\n<li>UNet-InceptionResNetV2 + CBAM+ Hypercolumns</li>\n<li>UNet-SeResNext50 + CBAM + Hypercolumns</li>\n<li>UNet-ResNet34 + CBAM + FPA</li>\n<li>UNet-ResNet18 + CBAM + FPA\nI used the weights of best validation score epochs.</li>\n</ol>\n\n<h2>Loss</h2>\n\n<ul>\n<li>BCE + LovaszHinge</li>\n<li>on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1</li>\n</ul>\n\n<h2>Optimizer &amp; Scheduler</h2>\n\n<ul>\n<li>Adam &amp; CosineAnnealingWarmRestart (20epoch cycle)</li>\n<li>learning rate : 1e-4 to 1e-6</li>\n</ul>\n\n<h2>Ensemble</h2>\n\n<ul>\n<li>simple average of the 7 models (x 5folds = total 35 models)</li>\n</ul>\n\n<h2>Postprocessing</h2>\n\n<ul>\n<li>TTA : None + h-flip + v-flip + h- and v- flip</li>\n<li>pixel threshold = 0.45</li>\n<li>small mask threshold = 18000\nBoth determined by the 5foldCV for model 1.</li>\n</ul>\n\n<h2>Final submission</h2>\n\n<ul>\n<li>I checked only Public LB score for ensembles. So I needed some criteria to choose the final submission. I decided to choose two submissions which were good in Public LB and stable against the small mask threshold, although these were not my best Public LB submission. Luckily I survived the shake up and got a gold medal.</li>\n</ul>",
  "messages": [
    {
      "id": "676728",
      "postDate": "11/19/2019 13:17:34",
      "content": "<p>Thanks Kaggle and Max Planck Institute for this interesting competition and congrats to all the winners!  Here is a brief summary of my solution (Public 0.67698, Private 0.66713).</p>\n\n<h2>Preprocessing</h2>\n\n<ul>\n<li>exclude bad images (removed 13 images)</li>\n<li>resize image size to (320, 512)</li>\n</ul>\n\n<h2>Augmentations (by Albumentations)</h2>\n\n<ul>\n<li>gamma (limit=(50,100), p=0.5)</li>\n<li>brightness (limit=0.2, p=0.5)</li>\n<li>shift (limit=0.2, border_mode=0, p=0.5)</li>\n<li>rotation (limit=30deg, border_mode=0, p=0.5)</li>\n<li>horizontal flip (p=0.5)</li>\n<li>vertical flip (p=0.5)</li>\n</ul>\n\n<h2>Validation</h2>\n\n<ul>\n<li>StratifiedKFold for the number of empty masks</li>\n</ul>\n\n<h2>Model (ensemble of 7 models x 5folds)</h2>\n\n<ol>\n<li>UNet-ResNet34 + CBAM + Hypercolumns</li>\n<li>same as 1. but with other seed</li>\n<li>UNet-ResNet18 + CBAM + Hypercolumns</li>\n<li>UNet-InceptionResNetV2 + CBAM+ Hypercolumns</li>\n<li>UNet-SeResNext50 + CBAM + Hypercolumns</li>\n<li>UNet-ResNet34 + CBAM + FPA</li>\n<li>UNet-ResNet18 + CBAM + FPA\nI used the weights of best validation score epochs.</li>\n</ol>\n\n<h2>Loss</h2>\n\n<ul>\n<li>BCE + LovaszHinge</li>\n<li>on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1</li>\n</ul>\n\n<h2>Optimizer &amp; Scheduler</h2>\n\n<ul>\n<li>Adam &amp; CosineAnnealingWarmRestart (20epoch cycle)</li>\n<li>learning rate : 1e-4 to 1e-6</li>\n</ul>\n\n<h2>Ensemble</h2>\n\n<ul>\n<li>simple average of the 7 models (x 5folds = total 35 models)</li>\n</ul>\n\n<h2>Postprocessing</h2>\n\n<ul>\n<li>TTA : None + h-flip + v-flip + h- and v- flip</li>\n<li>pixel threshold = 0.45</li>\n<li>small mask threshold = 18000\nBoth determined by the 5foldCV for model 1.</li>\n</ul>\n\n<h2>Final submission</h2>\n\n<ul>\n<li>I checked only Public LB score for ensembles. So I needed some criteria to choose the final submission. I decided to choose two submissions which were good in Public LB and stable against the small mask threshold, although these were not my best Public LB submission. Luckily I survived the shake up and got a gold medal.</li>\n</ul>",
      "rawMarkdown": "Thanks Kaggle and Max Planck Institute for this interesting competition and congrats to all the winners!  Here is a brief summary of my solution (Public 0.67698, Private 0.66713).\n\n\n##Preprocessing\n- exclude bad images (removed 13 images)\n- resize image size to (320, 512)\n\n\n##Augmentations (by Albumentations)\n- gamma (limit=(50,100), p=0.5)\n- brightness (limit=0.2, p=0.5)\n- shift (limit=0.2, border_mode=0, p=0.5)\n- rotation (limit=30deg, border_mode=0, p=0.5)\n- horizontal flip (p=0.5)\n- vertical flip (p=0.5)\n\n\n##Validation\n- StratifiedKFold for the number of empty masks\n\n\n##Model (ensemble of 7 models x 5folds)\n1. UNet-ResNet34 + CBAM + Hypercolumns\n2. same as 1. but with other seed\n3. UNet-ResNet18 + CBAM + Hypercolumns\n4. UNet-InceptionResNetV2 + CBAM+ Hypercolumns\n5. UNet-SeResNext50 + CBAM + Hypercolumns\n6. UNet-ResNet34 + CBAM + FPA\n7. UNet-ResNet18 + CBAM + FPA\nI used the weights of best validation score epochs.\n\n\n##Loss\n- BCE + LovaszHinge\n- on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1\n\n\n##Optimizer &amp; Scheduler\n- Adam &amp; CosineAnnealingWarmRestart (20epoch cycle)\n- learning rate : 1e-4 to 1e-6\n\n\n##Ensemble \n- simple average of the 7 models (x 5folds = total 35 models)\n\n\n##Postprocessing\n- TTA : None + h-flip + v-flip + h- and v- flip\n- pixel threshold = 0.45\n- small mask threshold = 18000\nBoth determined by the 5foldCV for model 1.\n\n\n##Final submission\n- I checked only Public LB score for ensembles. So I needed some criteria to choose the final submission. I decided to choose two submissions which were good in Public LB and stable against the small mask threshold, although these were not my best Public LB submission. Luckily I survived the shake up and got a gold medal.",
      "votes": null
    },
    {
      "id": "676848",
      "postDate": "11/19/2019 14:45:32",
      "content": "<p>Congratulations\nThank you for Sharing your Insights &amp; Approach! <a href=\"/tikutiku\">@tikutiku</a> </p>",
      "rawMarkdown": "Congratulations\nThank you for Sharing your Insights &amp; Approach! @tikutiku",
      "votes": null
    },
    {
      "id": "677022",
      "postDate": "11/19/2019 18:04:33",
      "content": "<p>Congratulations Tom on achieving another solo Gold. Your Dog model was great too.</p>\n\n<blockquote>\n  <p>on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1</p>\n</blockquote>\n\n<p>What is \"deep supervision\"? I notice you say \"for only non-emp.y masks\", that's smart. It appears all the top solutions did something like this. How did you accomplish this? i.e. did you modify the loss function, training samples, etc</p>",
      "rawMarkdown": "Congratulations Tom on achieving another solo Gold. Your Dog model was great too.\n\n&gt; on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1\n\nWhat is \"deep supervision\"? I notice you say \"for only non-emp.y masks\", that's smart. It appears all the top solutions did something like this. How did you accomplish this? i.e. did you modify the loss function, training samples, etc",
      "votes": null
    },
    {
      "id": "677185",
      "postDate": "11/19/2019 22:36:51",
      "content": "<p>Thanks Chris !\nIn deepsupervision, not only the final output mask but also intermediate feature maps are compared with label.  My loss function is like below:\n<code>\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(logits,label)\nloss += lovasz_hinge(logits.view(-1,h,w), label.view(-1,h,w))\nfor x in [x1,x2,x3,x4]: #x1,x2,x3,x4 are upsampled outputs of decoder layers (+ additional conv)\n    loss += 0.1 * criterion_lovasz_hinge_non_empty(criterion, x, label)\n</code></p>",
      "rawMarkdown": "Thanks Chris !\nIn deepsupervision, not only the final output mask but also intermediate feature maps are compared with label.  My loss function is like below:\n```\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(logits,label)\nloss += lovasz_hinge(logits.view(-1,h,w), label.view(-1,h,w))\nfor x in [x1,x2,x3,x4]: #x1,x2,x3,x4 are upsampled outputs of decoder layers (+ additional conv)\n    loss += 0.1 * criterion_lovasz_hinge_non_empty(criterion, x, label)\n```",
      "votes": null
    },
    {
      "id": "677190",
      "postDate": "11/19/2019 22:47:08",
      "content": "<p>Wow, that's cool, I never knew about that. I wonder what the Unet architecture looks like. If you start with an image of size 32x, then the encoder reduces it 1x (with 5 down samples). Then the decoder upsamples 5 times with intermediate sizes 2x, 4x, 8x, 16x, 32x. How do you compare the sizes 4x, 8x, 16x with the label mask which has size 32x?</p>\n\n<p>Is your variable x1 an 8 time upsample of 4x. And x2 is an 4 time upsample of 8x. And x3 is a 2 time upsample of 16x?</p>\n\n<p>(When I say 32x that is <code>320x512</code> and 16x is <code>160x256</code> and 8x is <code>80x128</code> and 4x is <code>40x64</code> and 2x is <code>20x32</code> and 1x is <code>10x16</code>).</p>",
      "rawMarkdown": "Wow, that's cool, I never knew about that. I wonder what the Unet architecture looks like. If you start with an image of size 32x, then the encoder reduces it 1x (with 5 down samples). Then the decoder upsamples 5 times with intermediate sizes 2x, 4x, 8x, 16x, 32x. How do you compare the sizes 4x, 8x, 16x with the label mask which has size 32x?\n\nIs your variable x1 an 8 time upsample of 4x. And x2 is an 4 time upsample of 8x. And x3 is a 2 time upsample of 16x?\n\n(When I say 32x that is `320x512` and 16x is `160x256` and 8x is `80x128` and 4x is `40x64` and 2x is `20x32` and 1x is `10x16`).",
      "votes": null
    },
    {
      "id": "677264",
      "postDate": "11/20/2019 01:24:12",
      "content": "<p>I compared them as you have mentioned. I also applied conv1x1 after upsample so that the number of channels matched with label's one (=3)</p>",
      "rawMarkdown": "I compared them as you have mentioned. I also applied conv1x1 after upsample so that the number of channels matched with label's one (=3)",
      "votes": null
    },
    {
      "id": "678716",
      "postDate": "11/21/2019 18:49:09",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  . I believe  i read about it in TGS Salt comp. Please check here . </p>\n\n<p><a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/68435\">https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/68435</a></p>\n\n<p>below is <a href=\"/shentao\">@shentao</a> 's  implementation of deep supervision .  I still could not fathom it , i tried to build a similar net called SeuTaoNet in this competition , but time ran out :) </p>\n\n<p><a href=\"https://github.com/SeuTao/TGS-Salt-Identification-Challenge-2018-_4th_place_solution\">https://github.com/SeuTao/TGS-Salt-Identification-Challenge-2018-_4th_place_solution</a></p>",
      "rawMarkdown": "cdeotte  . I believe  i read about it in TGS Salt comp. Please check here . \n\nhttps://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/68435\n\nbelow is @shentao 's  implementation of deep supervision .  I still could not fathom it , i tried to build a similar net called SeuTaoNet in this competition , but time ran out :) \n\nhttps://github.com/SeuTao/TGS-Salt-Identification-Challenge-2018-_4th_place_solution",
      "votes": null
    },
    {
      "id": "678718",
      "postDate": "11/21/2019 18:51:57",
      "content": "<p>Congratulations :) . Gold to remember .  I will try to implement this architecture , always fascinated by Deep Supervision and Hypercolumn.</p>",
      "rawMarkdown": "Congratulations :) . Gold to remember .  I will try to implement this architecture , always fascinated by Deep Supervision and Hypercolumn.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 676848,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "11/19/2019 14:45:32",
      "content": "<p>Congratulations\nThank you for Sharing your Insights &amp; Approach! <a href=\"/tikutiku\">@tikutiku</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 677022,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/19/2019 18:04:33",
      "content": "<p>Congratulations Tom on achieving another solo Gold. Your Dog model was great too.</p>\n\n<blockquote>\n  <p>on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1</p>\n</blockquote>\n\n<p>What is \"deep supervision\"? I notice you say \"for only non-emp.y masks\", that's smart. It appears all the top solutions did something like this. How did you accomplish this? i.e. did you modify the loss function, training samples, etc</p>",
      "votes": null,
      "replies": [
        {
          "id": 677185,
          "author_name": "tikutiku",
          "author_url": "",
          "post_date": "11/19/2019 22:36:51",
          "content": "<p>Thanks Chris !\nIn deepsupervision, not only the final output mask but also intermediate feature maps are compared with label.  My loss function is like below:\n<code>\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(logits,label)\nloss += lovasz_hinge(logits.view(-1,h,w), label.view(-1,h,w))\nfor x in [x1,x2,x3,x4]: #x1,x2,x3,x4 are upsampled outputs of decoder layers (+ additional conv)\n    loss += 0.1 * criterion_lovasz_hinge_non_empty(criterion, x, label)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677190,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/19/2019 22:47:08",
          "content": "<p>Wow, that's cool, I never knew about that. I wonder what the Unet architecture looks like. If you start with an image of size 32x, then the encoder reduces it 1x (with 5 down samples). Then the decoder upsamples 5 times with intermediate sizes 2x, 4x, 8x, 16x, 32x. How do you compare the sizes 4x, 8x, 16x with the label mask which has size 32x?</p>\n\n<p>Is your variable x1 an 8 time upsample of 4x. And x2 is an 4 time upsample of 8x. And x3 is a 2 time upsample of 16x?</p>\n\n<p>(When I say 32x that is <code>320x512</code> and 16x is <code>160x256</code> and 8x is <code>80x128</code> and 4x is <code>40x64</code> and 2x is <code>20x32</code> and 1x is <code>10x16</code>).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677264,
          "author_name": "tikutiku",
          "author_url": "",
          "post_date": "11/20/2019 01:24:12",
          "content": "<p>I compared them as you have mentioned. I also applied conv1x1 after upsample so that the number of channels matched with label's one (=3)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678716,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/21/2019 18:49:09",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  . I believe  i read about it in TGS Salt comp. Please check here . </p>\n\n<p><a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/68435\">https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/68435</a></p>\n\n<p>below is <a href=\"/shentao\">@shentao</a> 's  implementation of deep supervision .  I still could not fathom it , i tried to build a similar net called SeuTaoNet in this competition , but time ran out :) </p>\n\n<p><a href=\"https://github.com/SeuTao/TGS-Salt-Identification-Challenge-2018-_4th_place_solution\">https://github.com/SeuTao/TGS-Salt-Identification-Challenge-2018-_4th_place_solution</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 678718,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/21/2019 18:51:57",
      "content": "<p>Congratulations :) . Gold to remember .  I will try to implement this architecture , always fascinated by Deep Supervision and Hypercolumn.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676728": "Thanks Kaggle and Max Planck Institute for this interesting competition and congrats to all the winners!  Here is a brief summary of my solution (Public 0.67698, Private 0.66713).\n\n\n##Preprocessing\n- exclude bad images (removed 13 images)\n- resize image size to (320, 512)\n\n\n##Augmentations (by Albumentations)\n- gamma (limit=(50,100), p=0.5)\n- brightness (limit=0.2, p=0.5)\n- shift (limit=0.2, border_mode=0, p=0.5)\n- rotation (limit=30deg, border_mode=0, p=0.5)\n- horizontal flip (p=0.5)\n- vertical flip (p=0.5)\n\n\n##Validation\n- StratifiedKFold for the number of empty masks\n\n\n##Model (ensemble of 7 models x 5folds)\n1. UNet-ResNet34 + CBAM + Hypercolumns\n2. same as 1. but with other seed\n3. UNet-ResNet18 + CBAM + Hypercolumns\n4. UNet-InceptionResNetV2 + CBAM+ Hypercolumns\n5. UNet-SeResNext50 + CBAM + Hypercolumns\n6. UNet-ResNet34 + CBAM + FPA\n7. UNet-ResNet18 + CBAM + FPA\nI used the weights of best validation score epochs.\n\n\n##Loss\n- BCE + LovaszHinge\n- on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1\n\n\n##Optimizer &amp; Scheduler\n- Adam &amp; CosineAnnealingWarmRestart (20epoch cycle)\n- learning rate : 1e-4 to 1e-6\n\n\n##Ensemble \n- simple average of the 7 models (x 5folds = total 35 models)\n\n\n##Postprocessing\n- TTA : None + h-flip + v-flip + h- and v- flip\n- pixel threshold = 0.45\n- small mask threshold = 18000\nBoth determined by the 5foldCV for model 1.\n\n\n##Final submission\n- I checked only Public LB score for ensembles. So I needed some criteria to choose the final submission. I decided to choose two submissions which were good in Public LB and stable against the small mask threshold, although these were not my best Public LB submission. Luckily I survived the shake up and got a gold medal.",
    "676848": "Congratulations\nThank you for Sharing your Insights &amp; Approach! @tikutiku",
    "677022": "Congratulations Tom on achieving another solo Gold. Your Dog model was great too.\n\n&gt; on top oh that I used deep supervision with BCE+LovaszHinge loss (for only non-empty masks) multiplied by 0.1\n\nWhat is \"deep supervision\"? I notice you say \"for only non-emp.y masks\", that's smart. It appears all the top solutions did something like this. How did you accomplish this? i.e. did you modify the loss function, training samples, etc",
    "677185": "Thanks Chris !\nIn deepsupervision, not only the final output mask but also intermediate feature maps are compared with label.  My loss function is like below:\n```\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(logits,label)\nloss += lovasz_hinge(logits.view(-1,h,w), label.view(-1,h,w))\nfor x in [x1,x2,x3,x4]: #x1,x2,x3,x4 are upsampled outputs of decoder layers (+ additional conv)\n    loss += 0.1 * criterion_lovasz_hinge_non_empty(criterion, x, label)\n```",
    "677190": "Wow, that's cool, I never knew about that. I wonder what the Unet architecture looks like. If you start with an image of size 32x, then the encoder reduces it 1x (with 5 down samples). Then the decoder upsamples 5 times with intermediate sizes 2x, 4x, 8x, 16x, 32x. How do you compare the sizes 4x, 8x, 16x with the label mask which has size 32x?\n\nIs your variable x1 an 8 time upsample of 4x. And x2 is an 4 time upsample of 8x. And x3 is a 2 time upsample of 16x?\n\n(When I say 32x that is `320x512` and 16x is `160x256` and 8x is `80x128` and 4x is `40x64` and 2x is `20x32` and 1x is `10x16`).",
    "677264": "I compared them as you have mentioned. I also applied conv1x1 after upsample so that the number of channels matched with label's one (=3)",
    "678716": "cdeotte  . I believe  i read about it in TGS Salt comp. Please check here . \n\nhttps://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/68435\n\nbelow is @shentao 's  implementation of deep supervision .  I still could not fathom it , i tried to build a similar net called SeuTaoNet in this competition , but time ran out :) \n\nhttps://github.com/SeuTao/TGS-Salt-Identification-Challenge-2018-_4th_place_solution",
    "678718": "Congratulations :) . Gold to remember .  I will try to implement this architecture , always fascinated by Deep Supervision and Hypercolumn."
  },
  "source": "meta"
}