{
  "id": 117949,
  "title": "Finally GM & 1st time won prize money! And 3rd place solution. ",
  "url": "/competitions/understanding_cloud_organization/discussion/117949",
  "author_name": "Xuan Cao",
  "post_date": "2019-11-19T00:02:40.250000",
  "votes": 98,
  "comment_count": 45,
  "views": 0,
  "content": "<p>UPDATE: code is now available <a href=\"https://github.com/naivelamb/kaggle-cloud-organization\">here</a>.</p>\n\n<p>Thanks for Max Planck Institute for Meteorology and Kaggle for hosting such an interesting competition. Congrats to all the winners.</p>\n\n<p>The key in my solution is training two segmentation models: <strong>seg1</strong> trained on all data with BCE loss, and <strong>seg2</strong> trained on non-empty images only with soft DICE loss. I think it works because this competition basically has two tasks: 1) detect the empty images; 2) predict accurate masks for the non-empty images. The two segmentation models address these two tasks respectively. </p>\n\n<h2>How I come up with this.</h2>\n\n<p>I started the competition with resnet34-FPN using BCE loss (<strong>seg1</strong>). This model achieves ~0.608 on LB and the major contribution comes from capturing the empty mask: it captures ~80% empty masks. I tried a lot to improve the non-empty part, like using combo loss of BCE and DICE, but it is hard to improve the neg-dice (dice score for the empty masks) and pos-dice (dice score for the non-empty makes) simultaneously.</p>\n\n<p>To predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together. Since all the train images are non-empty, we can use soft DICE loss directly and the model would focus on predicting accurate masks. I used exactly the same network structure, resnet34-FPN (<strong>seg2</strong>). Then I simply replace all the non-empty predictions from <strong>seg1</strong> model using the predictions from ‘seg2’. Only 1 fold of this 2-stage segmentation pipeline, no TTA, no min-size remover, no classifier, no threshold adjustment (all 0.5) could achieve LB 0.652. After including a resnet34 classifier (0.5 threshold), I got LB 0.655. </p>\n\n<p>Later on, I managed to train all 4 classes in one model by implementing pos-only soft DICE loss. The code looks like:</p>\n\n<p><code>python\ndef dice_only_pos(logits, labels, labels_fc):\n    # logits -&amp;gt; pixel level predictions\n    # labels -&amp;gt; pixel level labels\n    # labels_fc -&amp;gt; image/channel level labels\n    pos_idx = (labels_fc &amp;gt; 0.5)\n    neg_idx = (labels_fc &amp;lt; 0.5)\n    loss = SoftDiceLoss()(logits[pos_idx], labels[pos_idx])\n    return loss\n</code>\nThis loss only counts the non-empty channels and ignores all the empty channels.</p>\n\n<p>In summary the pipeline looks like: \n&gt;1. <strong>seg1</strong>: a multi-label segmentation model trained with BCE loss\n&gt;2. <strong>seg2</strong>: a multi-label segmentation model trained with pos-only soft DICE loss\n&gt;3. <strong>cls</strong>: a multi-label classifier trained with BCE loss. </p>\n\n<p>The final submission is achieved by the following steps:\n&gt;1. Get predictions using <strong>seg1</strong>\n&gt;2. Replacing the non-empty masks from <strong>seg1</strong> by predictions from <strong>seg2</strong>\n&gt;3. Removing more empty masks using <strong>cls</strong></p>\n\n<p>Both pixel-level (segmentation) and image-level (classifier) thresholds are 0.5. </p>\n\n<h2>Baseline results for the 2-stage segmentation</h2>\n\n<p>Model summary:\n&gt;Network: Resnet34-FPN \n&gt;Image size: 384x576\n&gt;Batch size: 16\n&gt;Optimizer: Adam\n&gt;Scheduler: reduceLR for seg1, warmRestart for seg2.\n&gt;Augmentations: H/V flip,  ShiftScalerRotate and GridDistortion\n&gt;TTA: raw, Horizontal Flip, Vertical Flip</p>\n\n<p>Results:\n&gt;1-fold: 0.664 \n&gt;5-fold + TTA3: 0.669\n&gt;5-fold + TTA3 + classifier: 0.670. </p>\n\n<p><em>TTA1 means only raw images; TTA3 means raw + H/V flip.</em></p>\n\n<p>The rest of my work is just trying different backbones to find the best one. My final models are:</p>\n\n<p>&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b7-FPN, TTA3\n&gt;cls: b1, TTA1</p>\n\n<p>Results:\n&gt;1-fold LB: 0.673\n&gt;5-fold LB: 0.6788</p>\n\n<h2>Ensemble</h2>\n\n<p>I ensembled multiple seg2 models using major vote. By including 4 models (b5-Unet, InceptionResnetV2-FPN, b7-FPN and b7-Unet), I achieved 0.6792 on LB. </p>\n\n<h2>Pseudo Labeling</h2>\n\n<p>I selected the pseudo labels based a LB 0.6790 submission with the following rules:\n&gt;1. Empty channels with classifier prediction &lt; 0.3\n&gt;2. Non-empty channels with classifier prediction &gt; 0.7</p>\n\n<p>An image is selected when all the 4 channels satisfy one of the conditions. 835 images are selected. I retrained the b7-FPN and b1-classifier including the pseudo labeling samples, and the final models are:\n&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b5-Unet + InceptionResnetV2-FPN + b7-Unet + b7-FPN + b7-FPN-PL, TTA3\n&gt;cls: b1-PL, TTA3</p>\n\n<p><em>PL means the model is retrained with pseudo labels</em></p>\n\n<p>This model achieves 0.6794 LB. </p>\n\n<p>On the last day, I decide to optimize the classifier threshold channel wise to achieve the best local CV, which gives me 0.6805 LB. </p>\n\n<h2>Other things worth mentioning</h2>\n\n<ol>\n<li>My CV aligns pretty well with the LB. 1-fold CV = LB +- 0.005. 5-fold CV = LB - (0.010 ~ 0.012). This helps a lot during the model development.</li>\n<li>Resizing the image before training could significantly reduce the training time. My resnet34-FPN could finish 1 epoch of training and validation in around 1 mins on a 2080Ti. </li>\n<li>For <strong>seg1</strong> and <strong>cls</strong>, complicated networks do not work. This is probably due to the noisy labels. For <strong>seg2</strong>, I cannot make seresnext50 and seresnext101 work and I have no idea why. </li>\n</ol>",
  "messages": [
    {
      "id": 676052,
      "postDate": "2019-11-19T00:02:40.250Z",
      "content": "<p>UPDATE: code is now available <a href=\"https://github.com/naivelamb/kaggle-cloud-organization\">here</a>.</p>\n\n<p>Thanks for Max Planck Institute for Meteorology and Kaggle for hosting such an interesting competition. Congrats to all the winners.</p>\n\n<p>The key in my solution is training two segmentation models: <strong>seg1</strong> trained on all data with BCE loss, and <strong>seg2</strong> trained on non-empty images only with soft DICE loss. I think it works because this competition basically has two tasks: 1) detect the empty images; 2) predict accurate masks for the non-empty images. The two segmentation models address these two tasks respectively. </p>\n\n<h2>How I come up with this.</h2>\n\n<p>I started the competition with resnet34-FPN using BCE loss (<strong>seg1</strong>). This model achieves ~0.608 on LB and the major contribution comes from capturing the empty mask: it captures ~80% empty masks. I tried a lot to improve the non-empty part, like using combo loss of BCE and DICE, but it is hard to improve the neg-dice (dice score for the empty masks) and pos-dice (dice score for the non-empty makes) simultaneously.</p>\n\n<p>To predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together. Since all the train images are non-empty, we can use soft DICE loss directly and the model would focus on predicting accurate masks. I used exactly the same network structure, resnet34-FPN (<strong>seg2</strong>). Then I simply replace all the non-empty predictions from <strong>seg1</strong> model using the predictions from ‘seg2’. Only 1 fold of this 2-stage segmentation pipeline, no TTA, no min-size remover, no classifier, no threshold adjustment (all 0.5) could achieve LB 0.652. After including a resnet34 classifier (0.5 threshold), I got LB 0.655. </p>\n\n<p>Later on, I managed to train all 4 classes in one model by implementing pos-only soft DICE loss. The code looks like:</p>\n\n<p><code>python\ndef dice_only_pos(logits, labels, labels_fc):\n    # logits -&amp;gt; pixel level predictions\n    # labels -&amp;gt; pixel level labels\n    # labels_fc -&amp;gt; image/channel level labels\n    pos_idx = (labels_fc &amp;gt; 0.5)\n    neg_idx = (labels_fc &amp;lt; 0.5)\n    loss = SoftDiceLoss()(logits[pos_idx], labels[pos_idx])\n    return loss\n</code>\nThis loss only counts the non-empty channels and ignores all the empty channels.</p>\n\n<p>In summary the pipeline looks like: \n&gt;1. <strong>seg1</strong>: a multi-label segmentation model trained with BCE loss\n&gt;2. <strong>seg2</strong>: a multi-label segmentation model trained with pos-only soft DICE loss\n&gt;3. <strong>cls</strong>: a multi-label classifier trained with BCE loss. </p>\n\n<p>The final submission is achieved by the following steps:\n&gt;1. Get predictions using <strong>seg1</strong>\n&gt;2. Replacing the non-empty masks from <strong>seg1</strong> by predictions from <strong>seg2</strong>\n&gt;3. Removing more empty masks using <strong>cls</strong></p>\n\n<p>Both pixel-level (segmentation) and image-level (classifier) thresholds are 0.5. </p>\n\n<h2>Baseline results for the 2-stage segmentation</h2>\n\n<p>Model summary:\n&gt;Network: Resnet34-FPN \n&gt;Image size: 384x576\n&gt;Batch size: 16\n&gt;Optimizer: Adam\n&gt;Scheduler: reduceLR for seg1, warmRestart for seg2.\n&gt;Augmentations: H/V flip,  ShiftScalerRotate and GridDistortion\n&gt;TTA: raw, Horizontal Flip, Vertical Flip</p>\n\n<p>Results:\n&gt;1-fold: 0.664 \n&gt;5-fold + TTA3: 0.669\n&gt;5-fold + TTA3 + classifier: 0.670. </p>\n\n<p><em>TTA1 means only raw images; TTA3 means raw + H/V flip.</em></p>\n\n<p>The rest of my work is just trying different backbones to find the best one. My final models are:</p>\n\n<p>&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b7-FPN, TTA3\n&gt;cls: b1, TTA1</p>\n\n<p>Results:\n&gt;1-fold LB: 0.673\n&gt;5-fold LB: 0.6788</p>\n\n<h2>Ensemble</h2>\n\n<p>I ensembled multiple seg2 models using major vote. By including 4 models (b5-Unet, InceptionResnetV2-FPN, b7-FPN and b7-Unet), I achieved 0.6792 on LB. </p>\n\n<h2>Pseudo Labeling</h2>\n\n<p>I selected the pseudo labels based a LB 0.6790 submission with the following rules:\n&gt;1. Empty channels with classifier prediction &lt; 0.3\n&gt;2. Non-empty channels with classifier prediction &gt; 0.7</p>\n\n<p>An image is selected when all the 4 channels satisfy one of the conditions. 835 images are selected. I retrained the b7-FPN and b1-classifier including the pseudo labeling samples, and the final models are:\n&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b5-Unet + InceptionResnetV2-FPN + b7-Unet + b7-FPN + b7-FPN-PL, TTA3\n&gt;cls: b1-PL, TTA3</p>\n\n<p><em>PL means the model is retrained with pseudo labels</em></p>\n\n<p>This model achieves 0.6794 LB. </p>\n\n<p>On the last day, I decide to optimize the classifier threshold channel wise to achieve the best local CV, which gives me 0.6805 LB. </p>\n\n<h2>Other things worth mentioning</h2>\n\n<ol>\n<li>My CV aligns pretty well with the LB. 1-fold CV = LB +- 0.005. 5-fold CV = LB - (0.010 ~ 0.012). This helps a lot during the model development.</li>\n<li>Resizing the image before training could significantly reduce the training time. My resnet34-FPN could finish 1 epoch of training and validation in around 1 mins on a 2080Ti. </li>\n<li>For <strong>seg1</strong> and <strong>cls</strong>, complicated networks do not work. This is probably due to the noisy labels. For <strong>seg2</strong>, I cannot make seresnext50 and seresnext101 work and I have no idea why. </li>\n</ol>",
      "rawMarkdown": "UPDATE: code is now available [here](https://github.com/naivelamb/kaggle-cloud-organization).\n\nThanks for Max Planck Institute for Meteorology and Kaggle for hosting such an interesting competition. Congrats to all the winners.\n\nThe key in my solution is training two segmentation models: **seg1** trained on all data with BCE loss, and **seg2** trained on non-empty images only with soft DICE loss. I think it works because this competition basically has two tasks: 1) detect the empty images; 2) predict accurate masks for the non-empty images. The two segmentation models address these two tasks respectively. \n## How I come up with this. \nI started the competition with resnet34-FPN using BCE loss (**seg1**). This model achieves ~0.608 on LB and the major contribution comes from capturing the empty mask: it captures ~80% empty masks. I tried a lot to improve the non-empty part, like using combo loss of BCE and DICE, but it is hard to improve the neg-dice (dice score for the empty masks) and pos-dice (dice score for the non-empty makes) simultaneously.\n\nTo predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together. Since all the train images are non-empty, we can use soft DICE loss directly and the model would focus on predicting accurate masks. I used exactly the same network structure, resnet34-FPN (**seg2**). Then I simply replace all the non-empty predictions from **seg1** model using the predictions from ‘seg2’. Only 1 fold of this 2-stage segmentation pipeline, no TTA, no min-size remover, no classifier, no threshold adjustment (all 0.5) could achieve LB 0.652. After including a resnet34 classifier (0.5 threshold), I got LB 0.655. \n\nLater on, I managed to train all 4 classes in one model by implementing pos-only soft DICE loss. The code looks like:\n\n```python\ndef dice_only_pos(logits, labels, labels_fc):\n    # logits -&gt; pixel level predictions\n    # labels -&gt; pixel level labels\n    # labels_fc -&gt; image/channel level labels\n    pos_idx = (labels_fc &gt; 0.5)\n    neg_idx = (labels_fc &lt; 0.5)\n    loss = SoftDiceLoss()(logits[pos_idx], labels[pos_idx])\n    return loss\n```\nThis loss only counts the non-empty channels and ignores all the empty channels.\n\nIn summary the pipeline looks like: \n&gt;1. **seg1**: a multi-label segmentation model trained with BCE loss\n&gt;2. **seg2**: a multi-label segmentation model trained with pos-only soft DICE loss\n&gt;3. **cls**: a multi-label classifier trained with BCE loss. \n\nThe final submission is achieved by the following steps:\n&gt;1. Get predictions using **seg1**\n&gt;2. Replacing the non-empty masks from **seg1** by predictions from **seg2**\n&gt;3. Removing more empty masks using **cls**\n\nBoth pixel-level (segmentation) and image-level (classifier) thresholds are 0.5. \n\n## Baseline results for the 2-stage segmentation\nModel summary:\n&gt;Network: Resnet34-FPN \n&gt;Image size: 384x576\n&gt;Batch size: 16\n&gt;Optimizer: Adam\n&gt;Scheduler: reduceLR for seg1, warmRestart for seg2.\n&gt;Augmentations: H/V flip,  ShiftScalerRotate and GridDistortion\n&gt;TTA: raw, Horizontal Flip, Vertical Flip\n\nResults:\n&gt;1-fold: 0.664 \n&gt;5-fold + TTA3: 0.669\n&gt;5-fold + TTA3 + classifier: 0.670. \n\n*TTA1 means only raw images; TTA3 means raw + H/V flip.*\n\nThe rest of my work is just trying different backbones to find the best one. My final models are:\n\n&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b7-FPN, TTA3\n&gt;cls: b1, TTA1\n\nResults:\n&gt;1-fold LB: 0.673\n&gt;5-fold LB: 0.6788\n\n## Ensemble\n\nI ensembled multiple seg2 models using major vote. By including 4 models (b5-Unet, InceptionResnetV2-FPN, b7-FPN and b7-Unet), I achieved 0.6792 on LB. \n\n## Pseudo Labeling\nI selected the pseudo labels based a LB 0.6790 submission with the following rules:\n&gt;1. Empty channels with classifier prediction &lt; 0.3\n&gt;2. Non-empty channels with classifier prediction &gt; 0.7\n\nAn image is selected when all the 4 channels satisfy one of the conditions. 835 images are selected. I retrained the b7-FPN and b1-classifier including the pseudo labeling samples, and the final models are:\n&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b5-Unet + InceptionResnetV2-FPN + b7-Unet + b7-FPN + b7-FPN-PL, TTA3\n&gt;cls: b1-PL, TTA3\n\n*PL means the model is retrained with pseudo labels*\n\nThis model achieves 0.6794 LB. \n\nOn the last day, I decide to optimize the classifier threshold channel wise to achieve the best local CV, which gives me 0.6805 LB. \n\n## Other things worth mentioning\n1. My CV aligns pretty well with the LB. 1-fold CV = LB +- 0.005. 5-fold CV = LB - (0.010 ~ 0.012). This helps a lot during the model development.\n2. Resizing the image before training could significantly reduce the training time. My resnet34-FPN could finish 1 epoch of training and validation in around 1 mins on a 2080Ti. \n3. For **seg1** and **cls**, complicated networks do not work. This is probably due to the noisy labels. For **seg2**, I cannot make seresnext50 and seresnext101 work and I have no idea why. ",
      "votes": 98
    },
    {
      "id": 676104,
      "postDate": "2019-11-19T00:54:02.740Z",
      "content": "<blockquote>\n  <p>Finally GM&amp; 1st time won prize money! And 3rd place solution.</p>\n</blockquote>\n\n<p>I have 3x congratulations for you!!! Well done and well deserved.</p>",
      "rawMarkdown": "&gt; Finally GM&amp; 1st time won prize money! And 3rd place solution.\n\nI have 3x congratulations for you!!! Well done and well deserved.",
      "votes": 1
    },
    {
      "id": 676132,
      "postDate": "2019-11-19T01:32:28.413Z",
      "content": "<p>What you did is truly amazing! Congratulations, and thank you for sharing your solution.</p>\n\n<p>If I understood correctly, we can say that you used <strong>'seg model with pos-only dice loss'</strong> as a base network and used <strong>'seg model with bce loss'</strong> and <strong>'cls model'</strong> to zero-out some masks to get public score of <strong>0.670</strong>. Then you changed backbones of those three models and got <strong>0.6788</strong>.</p>\n\n<p>It is surprising that just changing backbones leaded to so much improvement of nearly <strong>0.09</strong>. Many of us didn't enjoy significant improvement from changing backbone architectures.</p>\n\n<p>What do you think made such an improvement possible?</p>\n\n<p>One more, you said you used 'soft dice-loss'. I googled it and found <a href=\"https://gist.github.com/jeremyjordan/9ea3032a32909f71dd2ab35fe3bacc08\">this</a>. (which references  <a href=\"https://mediatum.ub.tum.de/doc/1395260/1395260.pdf\">https://mediatum.ub.tum.de/doc/1395260/1395260.pdf</a> page 72) It squares masks element-wise.</p>\n\n<p><code>\ndef soft_dice_loss(y_true, y_pred, epsilon=1e-6): \n    axes = tuple(range(1, len(y_pred.shape)-1)) \n    numerator = 2. * np.sum(y_pred * y_true, axes)\n    denominator = np.sum(np.square(y_pred) + np.square(y_true), axes)\n    return 1 - np.mean(numerator / (denominator + epsilon))\n</code></p>\n\n<p>Did using it instead of  normal dice loss improve your score?</p>\n\n<p>Thanks again for your kind writeup.</p>",
      "rawMarkdown": "What you did is truly amazing! Congratulations, and thank you for sharing your solution.\n\nIf I understood correctly, we can say that you used **'seg model with pos-only dice loss'** as a base network and used **'seg model with bce loss'** and **'cls model'** to zero-out some masks to get public score of **0.670**. Then you changed backbones of those three models and got **0.6788**.\n\nIt is surprising that just changing backbones leaded to so much improvement of nearly **0.09**. Many of us didn't enjoy significant improvement from changing backbone architectures.\n\nWhat do you think made such an improvement possible?\n\nOne more, you said you used 'soft dice-loss'. I googled it and found [this](https://gist.github.com/jeremyjordan/9ea3032a32909f71dd2ab35fe3bacc08). (which references  https://mediatum.ub.tum.de/doc/1395260/1395260.pdf page 72) It squares masks element-wise.\n\n```\ndef soft_dice_loss(y_true, y_pred, epsilon=1e-6): \n    axes = tuple(range(1, len(y_pred.shape)-1)) \n    numerator = 2. * np.sum(y_pred * y_true, axes)\n    denominator = np.sum(np.square(y_pred) + np.square(y_true), axes)\n    return 1 - np.mean(numerator / (denominator + epsilon))\n```\n\nDid using it instead of  normal dice loss improve your score?\n\nThanks again for your kind writeup.",
      "votes": 2,
      "replies": [
        {
          "id": 676227,
          "postDate": "2019-11-19T03:19:58.580Z",
          "content": "<blockquote>\n  <p>What do you think made such an improvement possible?</p>\n</blockquote>\n\n<p>My 2-stage segmentation pipeline. It is much easier to improve empty and non-empty predictions respectively. </p>\n\n<p>The soft-dice loss is basically soft-f1 loss at pixel level. </p>",
          "rawMarkdown": "&gt;What do you think made such an improvement possible?\n\nMy 2-stage segmentation pipeline. It is much easier to improve empty and non-empty predictions respectively. \n\nThe soft-dice loss is basically soft-f1 loss at pixel level. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 676058,
      "postDate": "2019-11-19T00:09:15.323Z",
      "content": "<p>Congratulations. Well deserved. That solo gold is a doozy. </p>\n\n<p>Ahh I was so close. I have a variant of your non-empty loss as well. My rationale was that it would focus purely on improving the quality of the mask and less on reducing the false positives. Ended up not moving forward with it because I couldnt get anything out of it initially. Shouldve spent more time on it</p>",
      "rawMarkdown": "Congratulations. Well deserved. That solo gold is a doozy. \n\nAhh I was so close. I have a variant of your non-empty loss as well. My rationale was that it would focus purely on improving the quality of the mask and less on reducing the false positives. Ended up not moving forward with it because I couldnt get anything out of it initially. Shouldve spent more time on it",
      "votes": 2,
      "replies": [
        {
          "id": 676060,
          "postDate": "2019-11-19T00:12:31.323Z",
          "content": "<p>Yeah, you are so close! This technique makes improving models much easier, since you only need to focus on one thing at a time. </p>",
          "rawMarkdown": "Yeah, you are so close! This technique makes improving models much easier, since you only need to focus on one thing at a time. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 677203,
      "postDate": "2019-11-19T23:17:56.830Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "replies": [
        {
          "id": 678692,
          "postDate": "2019-11-21T18:23:21.023Z",
          "content": "<p>Thanks. </p>",
          "rawMarkdown": "Thanks. "
        }
      ]
    },
    {
      "id": 677149,
      "postDate": "2019-11-19T21:25:01.337Z",
      "content": "<p>Congratulations, great work and post !</p>",
      "rawMarkdown": "Congratulations, great work and post !",
      "replies": [
        {
          "id": 678693,
          "postDate": "2019-11-21T18:23:27.427Z",
          "content": "<p>Thanks. </p>",
          "rawMarkdown": "Thanks. "
        }
      ]
    },
    {
      "id": 677106,
      "postDate": "2019-11-19T19:59:45.377Z",
      "content": "<p>Congratz! </p>",
      "rawMarkdown": "Congratz! ",
      "replies": [
        {
          "id": 678695,
          "postDate": "2019-11-21T18:23:35.620Z",
          "content": "<p>Thanks. </p>",
          "rawMarkdown": "Thanks. "
        }
      ]
    },
    {
      "id": 677008,
      "postDate": "2019-11-19T17:53:03.477Z",
      "content": "<p>Congratulations Xuan on solo Gold and becoming Grand Master. </p>\n\n<p>I like your <code>def dice_only_pos(logits, labels, labels_fc)</code> function. I plan to use that in the future. That is key for Kaggle competitions with their discontinuous Dice metric. </p>\n\n<p>I also noticed how different segmentation models would achieve different empty mask prediction accuracies. I even added a metric to my models:</p>\n\n<pre><code>def kaggle_acc(y_true, y_pred0, pix=0.5, area=24000, dim=(384,576)):\n\n# PIXEL THRESHOLD\ny_pred = K.cast( K.greater(y_pred0,pix), K.floatx() )\n\n# MIN AREA THRESHOLD\ns = K.sum(y_pred, axis=(1,2))\ns = K.cast( K.greater(s, area), K.floatx() )\n\n# REMOVE MIN AREA\ns = K.reshape(s,(-1,1))\ns = K.repeat(s,dim[0]*dim[1])\ns = K.reshape(s,(-1,1))\ny_pred = K.permute_dimensions(y_pred,(0,3,1,2))\ny_pred = K.reshape(y_pred,shape=(-1,1))\ny_pred = s*y_pred\ny_pred = K.reshape(y_pred,(-1,y_pred0.shape[3],dim[0],dim[1]))\ny_pred = K.permute_dimensions(y_pred,(0,2,3,1))\n\n# COMPUTE KAGGLE ACC\ntotal_y_true = K.sum(y_true, axis=(1,2))\ntotal_y_true = K.cast( K.greater(total_y_true, 0), K.floatx() )\n\ntotal_y_pred = K.sum(y_pred, axis=(1,2))\ntotal_y_pred = K.cast( K.greater(total_y_pred, 0), K.floatx() )\n\nreturn 1 - K.mean( K.abs( total_y_pred - total_y_true ) )\n</code></pre>\n\n<p>And then similar to you, I noticed that some models had low LB score but very high accuracy predicting empty masks. While other models had higher LB score but lower accuracy predicting empty masks. I did something similar to you and converted the low LB high ACC models into classifiers to remove false positives from the high LB low ACC models.</p>",
      "rawMarkdown": "Congratulations Xuan on solo Gold and becoming Grand Master. \n\nI like your `def dice_only_pos(logits, labels, labels_fc)` function. I plan to use that in the future. That is key for Kaggle competitions with their discontinuous Dice metric. \n\nI also noticed how different segmentation models would achieve different empty mask prediction accuracies. I even added a metric to my models:\n\n    def kaggle_acc(y_true, y_pred0, pix=0.5, area=24000, dim=(384,576)):\n \n    # PIXEL THRESHOLD\n    y_pred = K.cast( K.greater(y_pred0,pix), K.floatx() )\n    \n    # MIN AREA THRESHOLD\n    s = K.sum(y_pred, axis=(1,2))\n    s = K.cast( K.greater(s, area), K.floatx() )\n\n    # REMOVE MIN AREA\n    s = K.reshape(s,(-1,1))\n    s = K.repeat(s,dim[0]*dim[1])\n    s = K.reshape(s,(-1,1))\n    y_pred = K.permute_dimensions(y_pred,(0,3,1,2))\n    y_pred = K.reshape(y_pred,shape=(-1,1))\n    y_pred = s*y_pred\n    y_pred = K.reshape(y_pred,(-1,y_pred0.shape[3],dim[0],dim[1]))\n    y_pred = K.permute_dimensions(y_pred,(0,2,3,1))\n\n    # COMPUTE KAGGLE ACC\n    total_y_true = K.sum(y_true, axis=(1,2))\n    total_y_true = K.cast( K.greater(total_y_true, 0), K.floatx() )\n\n    total_y_pred = K.sum(y_pred, axis=(1,2))\n    total_y_pred = K.cast( K.greater(total_y_pred, 0), K.floatx() )\n\n    return 1 - K.mean( K.abs( total_y_pred - total_y_true ) )\n\nAnd then similar to you, I noticed that some models had low LB score but very high accuracy predicting empty masks. While other models had higher LB score but lower accuracy predicting empty masks. I did something similar to you and converted the low LB high ACC models into classifiers to remove false positives from the high LB low ACC models.",
      "replies": [
        {
          "id": 677348,
          "postDate": "2019-11-20T04:05:08.257Z",
          "content": "<blockquote>\n  <p>That is key for Kaggle competitions with their discontinuous Dice metric.</p>\n</blockquote>\n\n<p>Yes, I realized it in steel, but was unable to overcome it there. There were too many empty masks in steel, and this approach would make each batch very unstable. I was thinking about training 4 segmentation models but gave up at last considering the inference time limit. </p>\n\n<p>Good luck to your image competition journal, wish you get your first CV gold medal soon! </p>",
          "rawMarkdown": "&gt;That is key for Kaggle competitions with their discontinuous Dice metric.\n\nYes, I realized it in steel, but was unable to overcome it there. There were too many empty masks in steel, and this approach would make each batch very unstable. I was thinking about training 4 segmentation models but gave up at last considering the inference time limit. \n\nGood luck to your image competition journal, wish you get your first CV gold medal soon! \n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 676988,
      "postDate": "2019-11-19T17:28:52.513Z",
      "content": "<p>Very nice and congratulations! I really like the write up. It is nice to be in the head of the winners sometimes ;-)</p>\n\n<p>The \"things worth mentioning\" I think are probably very important - especially point 1! Always trust your CV</p>",
      "rawMarkdown": "Very nice and congratulations! I really like the write up. It is nice to be in the head of the winners sometimes ;-)\n\nThe \"things worth mentioning\" I think are probably very important - especially point 1! Always trust your CV"
    },
    {
      "id": 676685,
      "postDate": "2019-11-19T12:23:24.150Z",
      "content": "<p>Double congrats <a href=\"/naivelamb\">@naivelamb</a> ! </p>",
      "rawMarkdown": "Double congrats @naivelamb ! ",
      "replies": [
        {
          "id": 677066,
          "postDate": "2019-11-19T18:49:35.027Z",
          "content": "<p>Thanks Giba! </p>",
          "rawMarkdown": "Thanks Giba! "
        }
      ]
    },
    {
      "id": 676660,
      "postDate": "2019-11-19T11:55:39.980Z",
      "content": "<p>Congratulations <a href=\"/naivelamb\">@naivelamb</a> , great work with your solution. Your GM title is well deserved :)</p>",
      "rawMarkdown": "Congratulations @naivelamb , great work with your solution. Your GM title is well deserved :)"
    },
    {
      "id": 676641,
      "postDate": "2019-11-19T11:39:23.020Z",
      "content": "<p>Сongratulations and thanks!</p>\n\n<p>seg1: a multi-label segmentation model trained with BCE loss\nseg2: a multi-label segmentation model trained with pos-only soft DICE loss\ncls: a multi-label classifier trained with BCE loss.</p>\n\n<p>Good idea!</p>",
      "rawMarkdown": "Сongratulations and thanks!\n\nseg1: a multi-label segmentation model trained with BCE loss\nseg2: a multi-label segmentation model trained with pos-only soft DICE loss\ncls: a multi-label classifier trained with BCE loss.\n\nGood idea!"
    },
    {
      "id": 676612,
      "postDate": "2019-11-19T10:54:07.650Z",
      "content": "<p>Great work, thanks for your write up, and congratulations!</p>\n\n<p>Just out of interest - were you using Kaggle Kernels, or do you have your own GPU?</p>",
      "rawMarkdown": "Great work, thanks for your write up, and congratulations!\n\nJust out of interest - were you using Kaggle Kernels, or do you have your own GPU?",
      "replies": [
        {
          "id": 677065,
          "postDate": "2019-11-19T18:49:25.317Z",
          "content": "<p>I used my own GPU. </p>",
          "rawMarkdown": "I used my own GPU. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 676401,
      "postDate": "2019-11-19T06:18:57.220Z",
      "content": "<p>Great write up! Well deserved GM!!</p>",
      "rawMarkdown": "Great write up! Well deserved GM!!"
    },
    {
      "id": 676288,
      "postDate": "2019-11-19T04:22:22.940Z",
      "content": "<p>Congratulations, good solution:)</p>",
      "rawMarkdown": "Congratulations, good solution:)"
    },
    {
      "id": 676213,
      "postDate": "2019-11-19T03:02:06.337Z",
      "content": "<p>Congratulations！\nI saw you trained several big models(with B5/B7 backbone), can you tell me how did you set the batchsize when you train your biggest model?  </p>",
      "rawMarkdown": "Congratulations！\nI saw you trained several big models(with B5/B7 backbone), can you tell me how did you set the batchsize when you train your biggest model?  ",
      "replies": [
        {
          "id": 676233,
          "postDate": "2019-11-19T03:23:41.593Z",
          "content": "<p>I have mentioned the batch size, 16 for all models. </p>",
          "rawMarkdown": "I have mentioned the batch size, 16 for all models. ",
          "votes": 1
        },
        {
          "id": 676242,
          "postDate": "2019-11-19T03:30:14.200Z",
          "content": "<p>Oh, sorry, I comment after I read your topic roughtly. Your topic includes so much information which is useful for me, thanks for your sharing!</p>",
          "rawMarkdown": "Oh, sorry, I comment after I read your topic roughtly. Your topic includes so much information which is useful for me, thanks for your sharing!"
        },
        {
          "id": 677124,
          "postDate": "2019-11-19T20:44:50.183Z",
          "content": "<p>Congratulations for your performance ! \nDid batch size of 16 fit in your GPU memory even for large architectures ? Did you use fp16 ?</p>",
          "rawMarkdown": "Congratulations for your performance ! \nDid batch size of 16 fit in your GPU memory even for large architectures ? Did you use fp16 ?"
        },
        {
          "id": 677323,
          "postDate": "2019-11-20T03:35:13.970Z",
          "content": "<p>apex is your good friend. </p>",
          "rawMarkdown": "apex is your good friend. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 676206,
      "postDate": "2019-11-19T02:53:15.597Z",
      "content": "<p>Congratulations😄 😄 </p>",
      "rawMarkdown": "Congratulations😄 😄 ",
      "replies": [
        {
          "id": 676235,
          "postDate": "2019-11-19T03:24:31.123Z",
          "content": "<p>Same to you. </p>",
          "rawMarkdown": "Same to you. "
        }
      ]
    },
    {
      "id": 676174,
      "postDate": "2019-11-19T02:18:59.670Z",
      "content": "<p>Congrats!!!!! Pseudo labeling part is really impressive!!</p>",
      "rawMarkdown": "Congrats!!!!! Pseudo labeling part is really impressive!!"
    },
    {
      "id": 676172,
      "postDate": "2019-11-19T02:15:46.563Z",
      "content": "<p>Hi, Congratulations for getting money and GM!!</p>\n\n<p>I have a question on some part of your solution.</p>\n\n<p>You mentioned the below sentences.</p>\n\n<p>&gt; To predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together.</p>\n\n<p>Please correct my understanding. </p>\n\n<p>Seg 1.</p>\n\n<p>Make a 4-class segmentation model. and get sub1.</p>\n\n<p>Seg 2.\nMake binary segmentation models for 4 each label. (4 segmentation models)\nImages of this competition can have multiple labels, so the same images can be used to train multiple labels. </p>\n\n<p>And, replace the predictions in sub1 with the predictions from 4 segmentation models.</p>\n\n<p>cls 1.\nMake a 4-class classifier(finding an empty mask). and remove the false-postive masks in sub1.</p>",
      "rawMarkdown": "Hi, Congratulations for getting money and GM!!\n\nI have a question on some part of your solution.\n\nYou mentioned the below sentences.\n\n&gt; To predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together.\n\nPlease correct my understanding. \n\nSeg 1.\n\nMake a 4-class segmentation model. and get sub1.\n\nSeg 2.\nMake binary segmentation models for 4 each label. (4 segmentation models)\nImages of this competition can have multiple labels, so the same images can be used to train multiple labels. \n\nAnd, replace the predictions in sub1 with the predictions from 4 segmentation models.\n\ncls 1.\nMake a 4-class classifier(finding an empty mask). and remove the false-postive masks in sub1.",
      "replies": [
        {
          "id": 676175,
          "postDate": "2019-11-19T02:19:19.450Z",
          "content": "<p>At first, you did that. Later on, you made one seg2 model with special loss (only pos), right..?</p>",
          "rawMarkdown": "At first, you did that. Later on, you made one seg2 model with special loss (only pos), right..?"
        },
        {
          "id": 676232,
          "postDate": "2019-11-19T03:23:15.030Z",
          "content": "<p>The very first version of seg2 is 4 binary segmentation models trained on non-empty images only (channel wise, so different channels will have different numbers of train images). </p>\n\n<p>The non-empty predictions in sub1 is replaced by those from seg2, since seg2 is better in giving accurate masks.  </p>",
          "rawMarkdown": "The very first version of seg2 is 4 binary segmentation models trained on non-empty images only (channel wise, so different channels will have different numbers of train images). \n\nThe non-empty predictions in sub1 is replaced by those from seg2, since seg2 is better in giving accurate masks.  "
        }
      ]
    },
    {
      "id": 676160,
      "postDate": "2019-11-19T02:04:17.497Z",
      "content": "<p>Congrats Cao.  </p>",
      "rawMarkdown": "Congrats Cao.  "
    },
    {
      "id": 676155,
      "postDate": "2019-11-19T01:57:26.630Z",
      "content": "<p>Congrats on winning and thanks for sharing the solution over to us.</p>",
      "rawMarkdown": "Congrats on winning and thanks for sharing the solution over to us."
    },
    {
      "id": 676148,
      "postDate": "2019-11-19T01:48:41.437Z",
      "content": "<p>Double congratulations <a href=\"/naivelamb\">@naivelamb</a> and well done. </p>\n\n<p>Your solution is interesting and thanks for sharing. Would you be rreleasing your code?</p>",
      "rawMarkdown": "Double congratulations @naivelamb and well done. \n\nYour solution is interesting and thanks for sharing. Would you be rreleasing your code?",
      "replies": [
        {
          "id": 676234,
          "postDate": "2019-11-19T03:24:20.337Z",
          "content": "<p>Thanks. I haven't decided for the code, it is messy. </p>",
          "rawMarkdown": "Thanks. I haven't decided for the code, it is messy. ",
          "votes": 1
        },
        {
          "id": 677181,
          "postDate": "2019-11-19T22:23:05.120Z",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a>, thanks for answering my question. I do  not mind a messy code as I would like to reproduce the top results of this comeptition with the aim of beating the 1st place scores. I want to do that with a single model that combines different top ideas in a couple of weeks when I have some free time.</p>",
          "rawMarkdown": "@naivelamb, thanks for answering my question. I do  not mind a messy code as I would like to reproduce the top results of this comeptition with the aim of beating the 1st place scores. I want to do that with a single model that combines different top ideas in a couple of weeks when I have some free time."
        }
      ]
    },
    {
      "id": 676121,
      "postDate": "2019-11-19T01:16:22.683Z",
      "content": "<p>Congratulations :) </p>",
      "rawMarkdown": "Congratulations :) "
    },
    {
      "id": 676061,
      "postDate": "2019-11-19T00:14:44.130Z",
      "content": "<p>Thank you for sharing and congratulations!</p>",
      "rawMarkdown": "Thank you for sharing and congratulations!"
    },
    {
      "id": 676056,
      "postDate": "2019-11-19T00:07:12.440Z",
      "content": "<p>Congrats! Thanks for sharing, very impressive!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing, very impressive!"
    },
    {
      "id": 676053,
      "postDate": "2019-11-19T00:05:16.813Z",
      "content": "<p>Congratulations! And thanks for sharing!!!</p>",
      "rawMarkdown": "Congratulations! And thanks for sharing!!!"
    },
    {
      "id": 676349,
      "postDate": "2019-11-19T05:22:39.070Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 676714,
      "postDate": "2019-11-19T12:59:36.957Z",
      "content": "<p>Congrats! Thanks for sharing!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing!"
    },
    {
      "id": 676397,
      "postDate": "2019-11-19T06:15:05.733Z",
      "content": "<p>Congrats Xuan! Thanks for sharing!</p>",
      "rawMarkdown": "Congrats Xuan! Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 676104,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2019-11-19T00:54:02.740000",
      "content": "<blockquote>\n  <p>Finally GM&amp; 1st time won prize money! And 3rd place solution.</p>\n</blockquote>\n\n<p>I have 3x congratulations for you!!! Well done and well deserved.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 676132,
      "author_name": "YoonSoo",
      "author_url": "",
      "post_date": "2019-11-19T01:32:28.413000",
      "content": "<p>What you did is truly amazing! Congratulations, and thank you for sharing your solution.</p>\n\n<p>If I understood correctly, we can say that you used <strong>'seg model with pos-only dice loss'</strong> as a base network and used <strong>'seg model with bce loss'</strong> and <strong>'cls model'</strong> to zero-out some masks to get public score of <strong>0.670</strong>. Then you changed backbones of those three models and got <strong>0.6788</strong>.</p>\n\n<p>It is surprising that just changing backbones leaded to so much improvement of nearly <strong>0.09</strong>. Many of us didn't enjoy significant improvement from changing backbone architectures.</p>\n\n<p>What do you think made such an improvement possible?</p>\n\n<p>One more, you said you used 'soft dice-loss'. I googled it and found <a href=\"https://gist.github.com/jeremyjordan/9ea3032a32909f71dd2ab35fe3bacc08\">this</a>. (which references  <a href=\"https://mediatum.ub.tum.de/doc/1395260/1395260.pdf\">https://mediatum.ub.tum.de/doc/1395260/1395260.pdf</a> page 72) It squares masks element-wise.</p>\n\n<p><code>\ndef soft_dice_loss(y_true, y_pred, epsilon=1e-6): \n    axes = tuple(range(1, len(y_pred.shape)-1)) \n    numerator = 2. * np.sum(y_pred * y_true, axes)\n    denominator = np.sum(np.square(y_pred) + np.square(y_true), axes)\n    return 1 - np.mean(numerator / (denominator + epsilon))\n</code></p>\n\n<p>Did using it instead of  normal dice loss improve your score?</p>\n\n<p>Thanks again for your kind writeup.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 676227,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T03:19:58.580000",
          "content": "<blockquote>\n  <p>What do you think made such an improvement possible?</p>\n</blockquote>\n\n<p>My 2-stage segmentation pipeline. It is much easier to improve empty and non-empty predictions respectively. </p>\n\n<p>The soft-dice loss is basically soft-f1 loss at pixel level. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676058,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-11-19T00:09:15.323000",
      "content": "<p>Congratulations. Well deserved. That solo gold is a doozy. </p>\n\n<p>Ahh I was so close. I have a variant of your non-empty loss as well. My rationale was that it would focus purely on improving the quality of the mask and less on reducing the false positives. Ended up not moving forward with it because I couldnt get anything out of it initially. Shouldve spent more time on it</p>",
      "votes": 2,
      "replies": [
        {
          "id": 676060,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T00:12:31.323000",
          "content": "<p>Yeah, you are so close! This technique makes improving models much easier, since you only need to focus on one thing at a time. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 677203,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-11-19T23:17:56.830000",
      "content": "<p>Congrats!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 678692,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-21T18:23:21.023000",
          "content": "<p>Thanks. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 677149,
      "author_name": "Rubnenkov Ivan",
      "author_url": "",
      "post_date": "2019-11-19T21:25:01.337000",
      "content": "<p>Congratulations, great work and post !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 678693,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-21T18:23:27.427000",
          "content": "<p>Thanks. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 677106,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-11-19T19:59:45.377000",
      "content": "<p>Congratz! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 678695,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-21T18:23:35.620000",
          "content": "<p>Thanks. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 677008,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-11-19T17:53:03.477000",
      "content": "<p>Congratulations Xuan on solo Gold and becoming Grand Master. </p>\n\n<p>I like your <code>def dice_only_pos(logits, labels, labels_fc)</code> function. I plan to use that in the future. That is key for Kaggle competitions with their discontinuous Dice metric. </p>\n\n<p>I also noticed how different segmentation models would achieve different empty mask prediction accuracies. I even added a metric to my models:</p>\n\n<pre><code>def kaggle_acc(y_true, y_pred0, pix=0.5, area=24000, dim=(384,576)):\n\n# PIXEL THRESHOLD\ny_pred = K.cast( K.greater(y_pred0,pix), K.floatx() )\n\n# MIN AREA THRESHOLD\ns = K.sum(y_pred, axis=(1,2))\ns = K.cast( K.greater(s, area), K.floatx() )\n\n# REMOVE MIN AREA\ns = K.reshape(s,(-1,1))\ns = K.repeat(s,dim[0]*dim[1])\ns = K.reshape(s,(-1,1))\ny_pred = K.permute_dimensions(y_pred,(0,3,1,2))\ny_pred = K.reshape(y_pred,shape=(-1,1))\ny_pred = s*y_pred\ny_pred = K.reshape(y_pred,(-1,y_pred0.shape[3],dim[0],dim[1]))\ny_pred = K.permute_dimensions(y_pred,(0,2,3,1))\n\n# COMPUTE KAGGLE ACC\ntotal_y_true = K.sum(y_true, axis=(1,2))\ntotal_y_true = K.cast( K.greater(total_y_true, 0), K.floatx() )\n\ntotal_y_pred = K.sum(y_pred, axis=(1,2))\ntotal_y_pred = K.cast( K.greater(total_y_pred, 0), K.floatx() )\n\nreturn 1 - K.mean( K.abs( total_y_pred - total_y_true ) )\n</code></pre>\n\n<p>And then similar to you, I noticed that some models had low LB score but very high accuracy predicting empty masks. While other models had higher LB score but lower accuracy predicting empty masks. I did something similar to you and converted the low LB high ACC models into classifiers to remove false positives from the high LB low ACC models.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677348,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-20T04:05:08.257000",
          "content": "<blockquote>\n  <p>That is key for Kaggle competitions with their discontinuous Dice metric.</p>\n</blockquote>\n\n<p>Yes, I realized it in steel, but was unable to overcome it there. There were too many empty masks in steel, and this approach would make each batch very unstable. I was thinking about training 4 segmentation models but gave up at last considering the inference time limit. </p>\n\n<p>Good luck to your image competition journal, wish you get your first CV gold medal soon! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676988,
      "author_name": "TheGruffalo",
      "author_url": "",
      "post_date": "2019-11-19T17:28:52.513000",
      "content": "<p>Very nice and congratulations! I really like the write up. It is nice to be in the head of the winners sometimes ;-)</p>\n\n<p>The \"things worth mentioning\" I think are probably very important - especially point 1! Always trust your CV</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676685,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-11-19T12:23:24.150000",
      "content": "<p>Double congrats <a href=\"/naivelamb\">@naivelamb</a> ! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 677066,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T18:49:35.027000",
          "content": "<p>Thanks Giba! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676660,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2019-11-19T11:55:39.980000",
      "content": "<p>Congratulations <a href=\"/naivelamb\">@naivelamb</a> , great work with your solution. Your GM title is well deserved :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676641,
      "author_name": "Ivan Bagmut",
      "author_url": "",
      "post_date": "2019-11-19T11:39:23.020000",
      "content": "<p>Сongratulations and thanks!</p>\n\n<p>seg1: a multi-label segmentation model trained with BCE loss\nseg2: a multi-label segmentation model trained with pos-only soft DICE loss\ncls: a multi-label classifier trained with BCE loss.</p>\n\n<p>Good idea!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676612,
      "author_name": "Marco Gorelli",
      "author_url": "",
      "post_date": "2019-11-19T10:54:07.650000",
      "content": "<p>Great work, thanks for your write up, and congratulations!</p>\n\n<p>Just out of interest - were you using Kaggle Kernels, or do you have your own GPU?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 677065,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T18:49:25.317000",
          "content": "<p>I used my own GPU. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676401,
      "author_name": "Hamish",
      "author_url": "",
      "post_date": "2019-11-19T06:18:57.220000",
      "content": "<p>Great write up! Well deserved GM!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676288,
      "author_name": "liuze",
      "author_url": "",
      "post_date": "2019-11-19T04:22:22.940000",
      "content": "<p>Congratulations, good solution:)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676213,
      "author_name": "lifengnan",
      "author_url": "",
      "post_date": "2019-11-19T03:02:06.337000",
      "content": "<p>Congratulations！\nI saw you trained several big models(with B5/B7 backbone), can you tell me how did you set the batchsize when you train your biggest model?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 676233,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T03:23:41.593000",
          "content": "<p>I have mentioned the batch size, 16 for all models. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 676242,
          "author_name": "lifengnan",
          "author_url": "",
          "post_date": "2019-11-19T03:30:14.200000",
          "content": "<p>Oh, sorry, I comment after I read your topic roughtly. Your topic includes so much information which is useful for me, thanks for your sharing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 677124,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2019-11-19T20:44:50.183000",
          "content": "<p>Congratulations for your performance ! \nDid batch size of 16 fit in your GPU memory even for large architectures ? Did you use fp16 ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 677323,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-20T03:35:13.970000",
          "content": "<p>apex is your good friend. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676206,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-11-19T02:53:15.597000",
      "content": "<p>Congratulations😄 😄 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 676235,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T03:24:31.123000",
          "content": "<p>Same to you. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676174,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-11-19T02:18:59.670000",
      "content": "<p>Congrats!!!!! Pseudo labeling part is really impressive!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676172,
      "author_name": "Youhan Lee",
      "author_url": "",
      "post_date": "2019-11-19T02:15:46.563000",
      "content": "<p>Hi, Congratulations for getting money and GM!!</p>\n\n<p>I have a question on some part of your solution.</p>\n\n<p>You mentioned the below sentences.</p>\n\n<p>&gt; To predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together.</p>\n\n<p>Please correct my understanding. </p>\n\n<p>Seg 1.</p>\n\n<p>Make a 4-class segmentation model. and get sub1.</p>\n\n<p>Seg 2.\nMake binary segmentation models for 4 each label. (4 segmentation models)\nImages of this competition can have multiple labels, so the same images can be used to train multiple labels. </p>\n\n<p>And, replace the predictions in sub1 with the predictions from 4 segmentation models.</p>\n\n<p>cls 1.\nMake a 4-class classifier(finding an empty mask). and remove the false-postive masks in sub1.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676175,
          "author_name": "Youhan Lee",
          "author_url": "",
          "post_date": "2019-11-19T02:19:19.450000",
          "content": "<p>At first, you did that. Later on, you made one seg2 model with special loss (only pos), right..?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676232,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T03:23:15.030000",
          "content": "<p>The very first version of seg2 is 4 binary segmentation models trained on non-empty images only (channel wise, so different channels will have different numbers of train images). </p>\n\n<p>The non-empty predictions in sub1 is replaced by those from seg2, since seg2 is better in giving accurate masks.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676160,
      "author_name": "llh1818",
      "author_url": "",
      "post_date": "2019-11-19T02:04:17.497000",
      "content": "<p>Congrats Cao.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676155,
      "author_name": "Vishy",
      "author_url": "",
      "post_date": "2019-11-19T01:57:26.630000",
      "content": "<p>Congrats on winning and thanks for sharing the solution over to us.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676148,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-11-19T01:48:41.437000",
      "content": "<p>Double congratulations <a href=\"/naivelamb\">@naivelamb</a> and well done. </p>\n\n<p>Your solution is interesting and thanks for sharing. Would you be rreleasing your code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676234,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-11-19T03:24:20.337000",
          "content": "<p>Thanks. I haven't decided for the code, it is messy. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 677181,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-11-19T22:23:05.120000",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a>, thanks for answering my question. I do  not mind a messy code as I would like to reproduce the top results of this comeptition with the aim of beating the 1st place scores. I want to do that with a single model that combines different top ideas in a couple of weeks when I have some free time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 676121,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2019-11-19T01:16:22.683000",
      "content": "<p>Congratulations :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676061,
      "author_name": "atfujita",
      "author_url": "",
      "post_date": "2019-11-19T00:14:44.130000",
      "content": "<p>Thank you for sharing and congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676056,
      "author_name": "Laevatein",
      "author_url": "",
      "post_date": "2019-11-19T00:07:12.440000",
      "content": "<p>Congrats! Thanks for sharing, very impressive!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676053,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2019-11-19T00:05:16.813000",
      "content": "<p>Congratulations! And thanks for sharing!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676349,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-19T05:22:39.070000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676714,
      "author_name": "bassbone(Junichi Ito)",
      "author_url": "",
      "post_date": "2019-11-19T12:59:36.957000",
      "content": "<p>Congrats! Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676397,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2019-11-19T06:15:05.733000",
      "content": "<p>Congrats Xuan! Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676052": "UPDATE: code is now available [here](https://github.com/naivelamb/kaggle-cloud-organization).\n\nThanks for Max Planck Institute for Meteorology and Kaggle for hosting such an interesting competition. Congrats to all the winners.\n\nThe key in my solution is training two segmentation models: **seg1** trained on all data with BCE loss, and **seg2** trained on non-empty images only with soft DICE loss. I think it works because this competition basically has two tasks: 1) detect the empty images; 2) predict accurate masks for the non-empty images. The two segmentation models address these two tasks respectively. \n## How I come up with this. \nI started the competition with resnet34-FPN using BCE loss (**seg1**). This model achieves ~0.608 on LB and the major contribution comes from capturing the empty mask: it captures ~80% empty masks. I tried a lot to improve the non-empty part, like using combo loss of BCE and DICE, but it is hard to improve the neg-dice (dice score for the empty masks) and pos-dice (dice score for the non-empty makes) simultaneously.\n\nTo predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together. Since all the train images are non-empty, we can use soft DICE loss directly and the model would focus on predicting accurate masks. I used exactly the same network structure, resnet34-FPN (**seg2**). Then I simply replace all the non-empty predictions from **seg1** model using the predictions from ‘seg2’. Only 1 fold of this 2-stage segmentation pipeline, no TTA, no min-size remover, no classifier, no threshold adjustment (all 0.5) could achieve LB 0.652. After including a resnet34 classifier (0.5 threshold), I got LB 0.655. \n\nLater on, I managed to train all 4 classes in one model by implementing pos-only soft DICE loss. The code looks like:\n\n```python\ndef dice_only_pos(logits, labels, labels_fc):\n    # logits -&gt; pixel level predictions\n    # labels -&gt; pixel level labels\n    # labels_fc -&gt; image/channel level labels\n    pos_idx = (labels_fc &gt; 0.5)\n    neg_idx = (labels_fc &lt; 0.5)\n    loss = SoftDiceLoss()(logits[pos_idx], labels[pos_idx])\n    return loss\n```\nThis loss only counts the non-empty channels and ignores all the empty channels.\n\nIn summary the pipeline looks like: \n&gt;1. **seg1**: a multi-label segmentation model trained with BCE loss\n&gt;2. **seg2**: a multi-label segmentation model trained with pos-only soft DICE loss\n&gt;3. **cls**: a multi-label classifier trained with BCE loss. \n\nThe final submission is achieved by the following steps:\n&gt;1. Get predictions using **seg1**\n&gt;2. Replacing the non-empty masks from **seg1** by predictions from **seg2**\n&gt;3. Removing more empty masks using **cls**\n\nBoth pixel-level (segmentation) and image-level (classifier) thresholds are 0.5. \n\n## Baseline results for the 2-stage segmentation\nModel summary:\n&gt;Network: Resnet34-FPN \n&gt;Image size: 384x576\n&gt;Batch size: 16\n&gt;Optimizer: Adam\n&gt;Scheduler: reduceLR for seg1, warmRestart for seg2.\n&gt;Augmentations: H/V flip,  ShiftScalerRotate and GridDistortion\n&gt;TTA: raw, Horizontal Flip, Vertical Flip\n\nResults:\n&gt;1-fold: 0.664 \n&gt;5-fold + TTA3: 0.669\n&gt;5-fold + TTA3 + classifier: 0.670. \n\n*TTA1 means only raw images; TTA3 means raw + H/V flip.*\n\nThe rest of my work is just trying different backbones to find the best one. My final models are:\n\n&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b7-FPN, TTA3\n&gt;cls: b1, TTA1\n\nResults:\n&gt;1-fold LB: 0.673\n&gt;5-fold LB: 0.6788\n\n## Ensemble\n\nI ensembled multiple seg2 models using major vote. By including 4 models (b5-Unet, InceptionResnetV2-FPN, b7-FPN and b7-Unet), I achieved 0.6792 on LB. \n\n## Pseudo Labeling\nI selected the pseudo labels based a LB 0.6790 submission with the following rules:\n&gt;1. Empty channels with classifier prediction &lt; 0.3\n&gt;2. Non-empty channels with classifier prediction &gt; 0.7\n\nAn image is selected when all the 4 channels satisfy one of the conditions. 835 images are selected. I retrained the b7-FPN and b1-classifier including the pseudo labeling samples, and the final models are:\n&gt;seg1: densenet121-FPN, TTA1\n&gt;seg2: b5-Unet + InceptionResnetV2-FPN + b7-Unet + b7-FPN + b7-FPN-PL, TTA3\n&gt;cls: b1-PL, TTA3\n\n*PL means the model is retrained with pseudo labels*\n\nThis model achieves 0.6794 LB. \n\nOn the last day, I decide to optimize the classifier threshold channel wise to achieve the best local CV, which gives me 0.6805 LB. \n\n## Other things worth mentioning\n1. My CV aligns pretty well with the LB. 1-fold CV = LB +- 0.005. 5-fold CV = LB - (0.010 ~ 0.012). This helps a lot during the model development.\n2. Resizing the image before training could significantly reduce the training time. My resnet34-FPN could finish 1 epoch of training and validation in around 1 mins on a 2080Ti. \n3. For **seg1** and **cls**, complicated networks do not work. This is probably due to the noisy labels. For **seg2**, I cannot make seresnext50 and seresnext101 work and I have no idea why. ",
    "676104": "&gt; Finally GM&amp; 1st time won prize money! And 3rd place solution.\n\nI have 3x congratulations for you!!! Well done and well deserved.",
    "676132": "What you did is truly amazing! Congratulations, and thank you for sharing your solution.\n\nIf I understood correctly, we can say that you used **'seg model with pos-only dice loss'** as a base network and used **'seg model with bce loss'** and **'cls model'** to zero-out some masks to get public score of **0.670**. Then you changed backbones of those three models and got **0.6788**.\n\nIt is surprising that just changing backbones leaded to so much improvement of nearly **0.09**. Many of us didn't enjoy significant improvement from changing backbone architectures.\n\nWhat do you think made such an improvement possible?\n\nOne more, you said you used 'soft dice-loss'. I googled it and found [this](https://gist.github.com/jeremyjordan/9ea3032a32909f71dd2ab35fe3bacc08). (which references  https://mediatum.ub.tum.de/doc/1395260/1395260.pdf page 72) It squares masks element-wise.\n\n```\ndef soft_dice_loss(y_true, y_pred, epsilon=1e-6): \n    axes = tuple(range(1, len(y_pred.shape)-1)) \n    numerator = 2. * np.sum(y_pred * y_true, axes)\n    denominator = np.sum(np.square(y_pred) + np.square(y_true), axes)\n    return 1 - np.mean(numerator / (denominator + epsilon))\n```\n\nDid using it instead of  normal dice loss improve your score?\n\nThanks again for your kind writeup.",
    "676058": "Congratulations. Well deserved. That solo gold is a doozy. \n\nAhh I was so close. I have a variant of your non-empty loss as well. My rationale was that it would focus purely on improving the quality of the mask and less on reducing the false positives. Ended up not moving forward with it because I couldnt get anything out of it initially. Shouldve spent more time on it",
    "677203": "Congrats!",
    "677149": "Congratulations, great work and post !",
    "677106": "Congratz! ",
    "677008": "Congratulations Xuan on solo Gold and becoming Grand Master. \n\nI like your `def dice_only_pos(logits, labels, labels_fc)` function. I plan to use that in the future. That is key for Kaggle competitions with their discontinuous Dice metric. \n\nI also noticed how different segmentation models would achieve different empty mask prediction accuracies. I even added a metric to my models:\n\n    def kaggle_acc(y_true, y_pred0, pix=0.5, area=24000, dim=(384,576)):\n \n    # PIXEL THRESHOLD\n    y_pred = K.cast( K.greater(y_pred0,pix), K.floatx() )\n    \n    # MIN AREA THRESHOLD\n    s = K.sum(y_pred, axis=(1,2))\n    s = K.cast( K.greater(s, area), K.floatx() )\n\n    # REMOVE MIN AREA\n    s = K.reshape(s,(-1,1))\n    s = K.repeat(s,dim[0]*dim[1])\n    s = K.reshape(s,(-1,1))\n    y_pred = K.permute_dimensions(y_pred,(0,3,1,2))\n    y_pred = K.reshape(y_pred,shape=(-1,1))\n    y_pred = s*y_pred\n    y_pred = K.reshape(y_pred,(-1,y_pred0.shape[3],dim[0],dim[1]))\n    y_pred = K.permute_dimensions(y_pred,(0,2,3,1))\n\n    # COMPUTE KAGGLE ACC\n    total_y_true = K.sum(y_true, axis=(1,2))\n    total_y_true = K.cast( K.greater(total_y_true, 0), K.floatx() )\n\n    total_y_pred = K.sum(y_pred, axis=(1,2))\n    total_y_pred = K.cast( K.greater(total_y_pred, 0), K.floatx() )\n\n    return 1 - K.mean( K.abs( total_y_pred - total_y_true ) )\n\nAnd then similar to you, I noticed that some models had low LB score but very high accuracy predicting empty masks. While other models had higher LB score but lower accuracy predicting empty masks. I did something similar to you and converted the low LB high ACC models into classifiers to remove false positives from the high LB low ACC models.",
    "676988": "Very nice and congratulations! I really like the write up. It is nice to be in the head of the winners sometimes ;-)\n\nThe \"things worth mentioning\" I think are probably very important - especially point 1! Always trust your CV",
    "676685": "Double congrats @naivelamb ! ",
    "676660": "Congratulations @naivelamb , great work with your solution. Your GM title is well deserved :)",
    "676641": "Сongratulations and thanks!\n\nseg1: a multi-label segmentation model trained with BCE loss\nseg2: a multi-label segmentation model trained with pos-only soft DICE loss\ncls: a multi-label classifier trained with BCE loss.\n\nGood idea!",
    "676612": "Great work, thanks for your write up, and congratulations!\n\nJust out of interest - were you using Kaggle Kernels, or do you have your own GPU?",
    "676401": "Great write up! Well deserved GM!!",
    "676288": "Congratulations, good solution:)",
    "676213": "Congratulations！\nI saw you trained several big models(with B5/B7 backbone), can you tell me how did you set the batchsize when you train your biggest model?  ",
    "676206": "Congratulations😄 😄 ",
    "676174": "Congrats!!!!! Pseudo labeling part is really impressive!!",
    "676172": "Hi, Congratulations for getting money and GM!!\n\nI have a question on some part of your solution.\n\nYou mentioned the below sentences.\n\n&gt; To predict the non-empty mask accurately, I decided to train 4 individual segmentation models for the non-empty images and then ensemble them together.\n\nPlease correct my understanding. \n\nSeg 1.\n\nMake a 4-class segmentation model. and get sub1.\n\nSeg 2.\nMake binary segmentation models for 4 each label. (4 segmentation models)\nImages of this competition can have multiple labels, so the same images can be used to train multiple labels. \n\nAnd, replace the predictions in sub1 with the predictions from 4 segmentation models.\n\ncls 1.\nMake a 4-class classifier(finding an empty mask). and remove the false-postive masks in sub1.",
    "676160": "Congrats Cao.  ",
    "676155": "Congrats on winning and thanks for sharing the solution over to us.",
    "676148": "Double congratulations @naivelamb and well done. \n\nYour solution is interesting and thanks for sharing. Would you be rreleasing your code?",
    "676121": "Congratulations :) ",
    "676061": "Thank you for sharing and congratulations!",
    "676056": "Congrats! Thanks for sharing, very impressive!",
    "676053": "Congratulations! And thanks for sharing!!!",
    "676349": "",
    "676714": "Congrats! Thanks for sharing!",
    "676397": "Congrats Xuan! Thanks for sharing!"
  }
}