{
  "id": 71782,
  "title": "6th Place Solution (41st in the Public LB)",
  "url": "/competitions/airbus-ship-detection/discussion/71782",
  "author_name": "Yauhen Babakhin",
  "post_date": "2018-11-16T13:00:41.170000",
  "votes": 50,
  "comment_count": 13,
  "views": 0,
  "content": "<p>First of all, I'd like to thank my teammates <a href=\"https://www.kaggle.com/zfturbo\">ZFTurbo</a> and <a href=\"https://www.kaggle.com/nicksergievskiy\">Nick Sergievskiy</a> for the nice teamwork and their great effort!</p>\n\n<p>We've entered the competition and merged into the team pretty late. Thus, we haven't had enough time to make lots of experiments. Another problem was the validation. While our local validation score was growing, Public LB score remained the same, at about 0.738, throughout the last week. The best reward for this struggle was a move from the 41st place in the Public LB to the 6th place in the Private LB.</p>\n\n<p>There have been already shared lots of great methods and ideas from the top teams. So, here is our brief solution outline.</p>\n\n<h2>Local Validation</h2>\n\n<p>We created 5 folds validation without a leak. As already mentioned, there were inconsistencies between Local and Public LB score movement. However, we tried to trust only the local validation.</p>\n\n<h2>Models</h2>\n\n<ol>\n<li>Classification (empty vs non-empty images). InceptionResNetV2, trained on 299x299 by <a href=\"https://www.kaggle.com/zfturbo\">ZFTurbo</a></li>\n<li>Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768</li>\n<li>Semantic segmentation: ResNet152 + U-Net by <a href=\"https://www.kaggle.com/zfturbo\">ZFTurbo</a>. Trained on 224x224 random crops, prediction with a sliding window</li>\n<li>Instance segmentation: ResNet18/ResNet50 + Mask R-CNN trained on 1000x1000 by <a href=\"https://www.kaggle.com/nicksergievskiy\">Nick Sergievskiy</a></li>\n</ol>\n\n<h2>Sampling</h2>\n\n<p>We used different percentages of non-empty / empty images in the batch. It was 50/50 for ResNet34 U-net, and even 90/10 for ResNet152 U-net. So, it generated lots of False Positive ships, and the role of the Classifier was pretty crucial.</p>\n\n<h2>Performance</h2>\n\n<ol>\n<li>Ensemble of 4 models + TTA for ResNet34 + U-Net. Private LB: 0.846 -&gt; 0.848 with Classification (deleting masks for confident empty images)</li>\n<li>Ensemble of 3 models + TTA for ResNet152 + U-Net. Private LB: 0.775 -&gt; 0.848 with Classification</li>\n<li>Ensemble of 3 models + TTA for Mask R-CNN. Private LB: 0.843 -&gt; 0.850 with Classification</li>\n</ol>\n\n<h2>Final Ensemble</h2>\n\n<p>As a final ensemble we used:</p>\n\n<ol>\n<li>Geometric mean of 7 U-Net models including one model with pseudolabels and one 2nd level model on OOF predictions. Denote predictions of this ensemble as <strong>unet_mask</strong></li>\n<li>Ensemble of 3 Mask R-CNN models</li>\n</ol>\n\n<p>For the Mask R-CNN ensemble <a href=\"https://www.kaggle.com/nicksergievskiy\">Nick Sergievskiy</a> chose two thresholds: thr_high and thr_mid. They gave the most confident predictions (<strong>rcnn_mask_high</strong>) and just confident predictions (<strong>rcnn_mask_mid</strong>). Further, <strong>rcnn_mask_high</strong> had the highest priority and replaced <strong>unet_mask</strong> objects; <strong>rcnn_mask_mid</strong> were added only if there was an intersection with <strong>unet_mask</strong> objects.</p>",
  "messages": [
    {
      "id": 422589,
      "postDate": "2018-11-16T13:00:41.170Z",
      "content": "<p>First of all, I'd like to thank my teammates <a href=\"https://www.kaggle.com/zfturbo\">ZFTurbo</a> and <a href=\"https://www.kaggle.com/nicksergievskiy\">Nick Sergievskiy</a> for the nice teamwork and their great effort!</p>\n\n<p>We've entered the competition and merged into the team pretty late. Thus, we haven't had enough time to make lots of experiments. Another problem was the validation. While our local validation score was growing, Public LB score remained the same, at about 0.738, throughout the last week. The best reward for this struggle was a move from the 41st place in the Public LB to the 6th place in the Private LB.</p>\n\n<p>There have been already shared lots of great methods and ideas from the top teams. So, here is our brief solution outline.</p>\n\n<h2>Local Validation</h2>\n\n<p>We created 5 folds validation without a leak. As already mentioned, there were inconsistencies between Local and Public LB score movement. However, we tried to trust only the local validation.</p>\n\n<h2>Models</h2>\n\n<ol>\n<li>Classification (empty vs non-empty images). InceptionResNetV2, trained on 299x299 by <a href=\"https://www.kaggle.com/zfturbo\">ZFTurbo</a></li>\n<li>Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768</li>\n<li>Semantic segmentation: ResNet152 + U-Net by <a href=\"https://www.kaggle.com/zfturbo\">ZFTurbo</a>. Trained on 224x224 random crops, prediction with a sliding window</li>\n<li>Instance segmentation: ResNet18/ResNet50 + Mask R-CNN trained on 1000x1000 by <a href=\"https://www.kaggle.com/nicksergievskiy\">Nick Sergievskiy</a></li>\n</ol>\n\n<h2>Sampling</h2>\n\n<p>We used different percentages of non-empty / empty images in the batch. It was 50/50 for ResNet34 U-net, and even 90/10 for ResNet152 U-net. So, it generated lots of False Positive ships, and the role of the Classifier was pretty crucial.</p>\n\n<h2>Performance</h2>\n\n<ol>\n<li>Ensemble of 4 models + TTA for ResNet34 + U-Net. Private LB: 0.846 -&gt; 0.848 with Classification (deleting masks for confident empty images)</li>\n<li>Ensemble of 3 models + TTA for ResNet152 + U-Net. Private LB: 0.775 -&gt; 0.848 with Classification</li>\n<li>Ensemble of 3 models + TTA for Mask R-CNN. Private LB: 0.843 -&gt; 0.850 with Classification</li>\n</ol>\n\n<h2>Final Ensemble</h2>\n\n<p>As a final ensemble we used:</p>\n\n<ol>\n<li>Geometric mean of 7 U-Net models including one model with pseudolabels and one 2nd level model on OOF predictions. Denote predictions of this ensemble as <strong>unet_mask</strong></li>\n<li>Ensemble of 3 Mask R-CNN models</li>\n</ol>\n\n<p>For the Mask R-CNN ensemble <a href=\"https://www.kaggle.com/nicksergievskiy\">Nick Sergievskiy</a> chose two thresholds: thr_high and thr_mid. They gave the most confident predictions (<strong>rcnn_mask_high</strong>) and just confident predictions (<strong>rcnn_mask_mid</strong>). Further, <strong>rcnn_mask_high</strong> had the highest priority and replaced <strong>unet_mask</strong> objects; <strong>rcnn_mask_mid</strong> were added only if there was an intersection with <strong>unet_mask</strong> objects.</p>",
      "rawMarkdown": "First of all, I'd like to thank my teammates [ZFTurbo][1] and [Nick Sergievskiy][2] for the nice teamwork and their great effort!\n\nWe've entered the competition and merged into the team pretty late. Thus, we haven't had enough time to make lots of experiments. Another problem was the validation. While our local validation score was growing, Public LB score remained the same, at about 0.738, throughout the last week. The best reward for this struggle was a move from the 41st place in the Public LB to the 6th place in the Private LB.\n\nThere have been already shared lots of great methods and ideas from the top teams. So, here is our brief solution outline.\n\n## Local Validation\nWe created 5 folds validation without a leak. As already mentioned, there were inconsistencies between Local and Public LB score movement. However, we tried to trust only the local validation.\n\n## Models\n 1. Classification (empty vs non-empty images). InceptionResNetV2, trained on 299x299 by [ZFTurbo][1]\n 2. Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768\n 3. Semantic segmentation: ResNet152 + U-Net by [ZFTurbo][1]. Trained on 224x224 random crops, prediction with a sliding window\n 4. Instance segmentation: ResNet18/ResNet50 + Mask R-CNN trained on 1000x1000 by [Nick Sergievskiy][2]\n\n## Sampling\nWe used different percentages of non-empty / empty images in the batch. It was 50/50 for ResNet34 U-net, and even 90/10 for ResNet152 U-net. So, it generated lots of False Positive ships, and the role of the Classifier was pretty crucial.\n\n## Performance\n 1. Ensemble of 4 models + TTA for ResNet34 + U-Net. Private LB: 0.846 -&gt; 0.848 with Classification (deleting masks for confident empty images)\n 2. Ensemble of 3 models + TTA for ResNet152 + U-Net. Private LB: 0.775 -&gt; 0.848 with Classification\n 3. Ensemble of 3 models + TTA for Mask R-CNN. Private LB: 0.843 -&gt; 0.850 with Classification\n\n## Final Ensemble\nAs a final ensemble we used:\n\n 1. Geometric mean of 7 U-Net models including one model with pseudolabels and one 2nd level model on OOF predictions. Denote predictions of this ensemble as **unet\\_mask**\n 2. Ensemble of 3 Mask R-CNN models\n\nFor the Mask R-CNN ensemble [Nick Sergievskiy][2] chose two thresholds: thr\\_high and thr\\_mid. They gave the most confident predictions (**rcnn\\_mask\\_high**) and just confident predictions (**rcnn\\_mask\\_mid**). Further, **rcnn\\_mask\\_high** had the highest priority and replaced **unet\\_mask** objects; **rcnn\\_mask\\_mid** were added only if there was an intersection with **unet\\_mask** objects.\n\n[1]: https://www.kaggle.com/zfturbo\n[2]: https://www.kaggle.com/nicksergievskiy",
      "votes": 50
    },
    {
      "id": 437585,
      "postDate": "2018-12-12T06:41:01.990Z",
      "content": "<p>Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768\nHow can you train on 256x256 and predict on 768x768.... which techniques u used? \nthanks.</p>",
      "rawMarkdown": "Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768\nHow can you train on 256x256 and predict on 768x768.... which techniques u used? \nthanks.",
      "replies": [
        {
          "id": 437667,
          "postDate": "2018-12-12T09:44:32.763Z",
          "content": "<p>The architecture used is fully convolutional. So, the network input could be of arbitrary size. You can read about fully convolutional networks here: <a href=\"https://people.eecs.berkeley.edu/~jonlong/long_shelhamer_fcn.pdf\">https://people.eecs.berkeley.edu/~jonlong/long_shelhamer_fcn.pdf</a></p>",
          "rawMarkdown": "The architecture used is fully convolutional. So, the network input could be of arbitrary size. You can read about fully convolutional networks here: https://people.eecs.berkeley.edu/~jonlong/long_shelhamer_fcn.pdf",
          "votes": 4
        },
        {
          "id": 558831,
          "postDate": "2019-06-23T02:44:03.423Z",
          "content": "<p>Isn't 256 the best input for a model trained with 256?</p>",
          "rawMarkdown": "Isn't 256 the best input for a model trained with 256?"
        },
        {
          "id": 627244,
          "postDate": "2019-09-15T16:03:59.400Z",
          "content": "<p>That's very interesting. Could you provide a link to a fastai/pytorch implementation that you found useful/used? Thanks in advance for your time and for the great solution. :)</p>",
          "rawMarkdown": "That's very interesting. Could you provide a link to a fastai/pytorch implementation that you found useful/used? Thanks in advance for your time and for the great solution. :)"
        }
      ]
    },
    {
      "id": 427684,
      "postDate": "2018-11-26T00:29:28.857Z",
      "content": "<p>Congrats <a href=\"/ybabakhin\">@ybabakhin</a>, and thanks for sharing.</p>",
      "rawMarkdown": "Congrats @ybabakhin, and thanks for sharing."
    },
    {
      "id": 425505,
      "postDate": "2018-11-21T17:31:59.770Z",
      "content": "<p>Congrats! Highly Informative... Keep up the great work...!</p>",
      "rawMarkdown": "Congrats! Highly Informative... Keep up the great work...!"
    },
    {
      "id": 425497,
      "postDate": "2018-11-21T17:12:39.027Z",
      "content": "<p>Congrats, why are you using 90/10 sampling ? and which sampling is better? 50:50 one ?</p>",
      "rawMarkdown": "Congrats, why are you using 90/10 sampling ? and which sampling is better? 50:50 one ?",
      "replies": [
        {
          "id": 426085,
          "postDate": "2018-11-22T15:23:13.507Z",
          "content": "<p>It's hard to tell which one is better. 50:50 mimics the distribution of empty images in the test set. And 90:10 finds almost all the ships, but has lots of False Positives. So, it needs a good emptiness Classifier on top.</p>",
          "rawMarkdown": "It's hard to tell which one is better. 50:50 mimics the distribution of empty images in the test set. And 90:10 finds almost all the ships, but has lots of False Positives. So, it needs a good emptiness Classifier on top."
        }
      ]
    },
    {
      "id": 422976,
      "postDate": "2018-11-17T06:55:14.603Z",
      "content": "<p>Thanks for sharing and congratulations to the team. Would you mind explaining and/or sharing resources about <strong>pseudolabels</strong>? Thanks in advance. :) </p>",
      "rawMarkdown": "Thanks for sharing and congratulations to the team. Would you mind explaining and/or sharing resources about **pseudolabels**? Thanks in advance. :) ",
      "replies": [
        {
          "id": 423016,
          "postDate": "2018-11-17T09:14:50.923Z",
          "content": "<p>It's just one of the semi-supervised learning methods. You could read more about it e.g. here: </p>\n\n<ol>\n<li><a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a></li>\n<li><a href=\"https://datawhatnow.com/pseudo-labeling-semi-supervised-learning/\">https://datawhatnow.com/pseudo-labeling-semi-supervised-learning/</a></li>\n</ol>",
          "rawMarkdown": "It's just one of the semi-supervised learning methods. You could read more about it e.g. here: \n\n1. http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\n2. https://datawhatnow.com/pseudo-labeling-semi-supervised-learning/",
          "votes": 2
        },
        {
          "id": 423021,
          "postDate": "2018-11-17T09:34:41.987Z",
          "content": "<p>Thanks. Much appreciated. :) </p>",
          "rawMarkdown": "Thanks. Much appreciated. :) "
        }
      ]
    },
    {
      "id": 422596,
      "postDate": "2018-11-16T13:18:54.843Z",
      "content": "<p>More proof to trust CV. Congrats!</p>",
      "rawMarkdown": "More proof to trust CV. Congrats!"
    },
    {
      "id": 423207,
      "postDate": "2018-11-17T17:54:36.643Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 437585,
      "author_name": "Mukesh",
      "author_url": "",
      "post_date": "2018-12-12T06:41:01.990000",
      "content": "<p>Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768\nHow can you train on 256x256 and predict on 768x768.... which techniques u used? \nthanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437667,
          "author_name": "Yauhen Babakhin",
          "author_url": "",
          "post_date": "2018-12-12T09:44:32.763000",
          "content": "<p>The architecture used is fully convolutional. So, the network input could be of arbitrary size. You can read about fully convolutional networks here: <a href=\"https://people.eecs.berkeley.edu/~jonlong/long_shelhamer_fcn.pdf\">https://people.eecs.berkeley.edu/~jonlong/long_shelhamer_fcn.pdf</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 558831,
          "author_name": "Jone",
          "author_url": "",
          "post_date": "2019-06-23T02:44:03.423000",
          "content": "<p>Isn't 256 the best input for a model trained with 256?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 627244,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2019-09-15T16:03:59.400000",
          "content": "<p>That's very interesting. Could you provide a link to a fastai/pytorch implementation that you found useful/used? Thanks in advance for your time and for the great solution. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 427684,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-11-26T00:29:28.857000",
      "content": "<p>Congrats <a href=\"/ybabakhin\">@ybabakhin</a>, and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425505,
      "author_name": "Karthik Chowdary Tsaliki",
      "author_url": "",
      "post_date": "2018-11-21T17:31:59.770000",
      "content": "<p>Congrats! Highly Informative... Keep up the great work...!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425497,
      "author_name": "soywu",
      "author_url": "",
      "post_date": "2018-11-21T17:12:39.027000",
      "content": "<p>Congrats, why are you using 90/10 sampling ? and which sampling is better? 50:50 one ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 426085,
          "author_name": "Yauhen Babakhin",
          "author_url": "",
          "post_date": "2018-11-22T15:23:13.507000",
          "content": "<p>It's hard to tell which one is better. 50:50 mimics the distribution of empty images in the test set. And 90:10 finds almost all the ships, but has lots of False Positives. So, it needs a good emptiness Classifier on top.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422976,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2018-11-17T06:55:14.603000",
      "content": "<p>Thanks for sharing and congratulations to the team. Would you mind explaining and/or sharing resources about <strong>pseudolabels</strong>? Thanks in advance. :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 423016,
          "author_name": "Yauhen Babakhin",
          "author_url": "",
          "post_date": "2018-11-17T09:14:50.923000",
          "content": "<p>It's just one of the semi-supervised learning methods. You could read more about it e.g. here: </p>\n\n<ol>\n<li><a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a></li>\n<li><a href=\"https://datawhatnow.com/pseudo-labeling-semi-supervised-learning/\">https://datawhatnow.com/pseudo-labeling-semi-supervised-learning/</a></li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423021,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2018-11-17T09:34:41.987000",
          "content": "<p>Thanks. Much appreciated. :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422596,
      "author_name": "Allen",
      "author_url": "",
      "post_date": "2018-11-16T13:18:54.843000",
      "content": "<p>More proof to trust CV. Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 423207,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-17T17:54:36.643000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422589": "First of all, I'd like to thank my teammates [ZFTurbo][1] and [Nick Sergievskiy][2] for the nice teamwork and their great effort!\n\nWe've entered the competition and merged into the team pretty late. Thus, we haven't had enough time to make lots of experiments. Another problem was the validation. While our local validation score was growing, Public LB score remained the same, at about 0.738, throughout the last week. The best reward for this struggle was a move from the 41st place in the Public LB to the 6th place in the Private LB.\n\nThere have been already shared lots of great methods and ideas from the top teams. So, here is our brief solution outline.\n\n## Local Validation\nWe created 5 folds validation without a leak. As already mentioned, there were inconsistencies between Local and Public LB score movement. However, we tried to trust only the local validation.\n\n## Models\n 1. Classification (empty vs non-empty images). InceptionResNetV2, trained on 299x299 by [ZFTurbo][1]\n 2. Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768\n 3. Semantic segmentation: ResNet152 + U-Net by [ZFTurbo][1]. Trained on 224x224 random crops, prediction with a sliding window\n 4. Instance segmentation: ResNet18/ResNet50 + Mask R-CNN trained on 1000x1000 by [Nick Sergievskiy][2]\n\n## Sampling\nWe used different percentages of non-empty / empty images in the batch. It was 50/50 for ResNet34 U-net, and even 90/10 for ResNet152 U-net. So, it generated lots of False Positive ships, and the role of the Classifier was pretty crucial.\n\n## Performance\n 1. Ensemble of 4 models + TTA for ResNet34 + U-Net. Private LB: 0.846 -&gt; 0.848 with Classification (deleting masks for confident empty images)\n 2. Ensemble of 3 models + TTA for ResNet152 + U-Net. Private LB: 0.775 -&gt; 0.848 with Classification\n 3. Ensemble of 3 models + TTA for Mask R-CNN. Private LB: 0.843 -&gt; 0.850 with Classification\n\n## Final Ensemble\nAs a final ensemble we used:\n\n 1. Geometric mean of 7 U-Net models including one model with pseudolabels and one 2nd level model on OOF predictions. Denote predictions of this ensemble as **unet\\_mask**\n 2. Ensemble of 3 Mask R-CNN models\n\nFor the Mask R-CNN ensemble [Nick Sergievskiy][2] chose two thresholds: thr\\_high and thr\\_mid. They gave the most confident predictions (**rcnn\\_mask\\_high**) and just confident predictions (**rcnn\\_mask\\_mid**). Further, **rcnn\\_mask\\_high** had the highest priority and replaced **unet\\_mask** objects; **rcnn\\_mask\\_mid** were added only if there was an intersection with **unet\\_mask** objects.\n\n[1]: https://www.kaggle.com/zfturbo\n[2]: https://www.kaggle.com/nicksergievskiy",
    "437585": "Semantic segmentation: ResNet34 + U-Net by me. Trained on 256x256 random crops, prediction on a full-size 768x768\nHow can you train on 256x256 and predict on 768x768.... which techniques u used? \nthanks.",
    "427684": "Congrats @ybabakhin, and thanks for sharing.",
    "425505": "Congrats! Highly Informative... Keep up the great work...!",
    "425497": "Congrats, why are you using 90/10 sampling ? and which sampling is better? 50:50 one ?",
    "422976": "Thanks for sharing and congratulations to the team. Would you mind explaining and/or sharing resources about **pseudolabels**? Thanks in advance. :) ",
    "422596": "More proof to trust CV. Congrats!",
    "423207": ""
  }
}