{
  "id": 111351,
  "title": "11th place solution [0.4796 private LB]",
  "url": "/competitions/open-images-2019-instance-segmentation/writeups/schwert-11th-place-solution-0-4796-private-lb",
  "author_name": "",
  "post_date": "2019-10-05T01:57:45.661282700Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I would like to thank the competition organizers and all the competitors! </p>\n\n<p>Here's my brief solution writeup:</p>\n\n<h2>1. Dataset</h2>\n\n<ul>\n<li>No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.</li>\n<li>Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal (1/300) across 300 classes. One instance is randomly picked from an image to train the segmentation network described below.</li>\n</ul>\n\n<h2>2. Pipiline and Models</h2>\n\n<p>A two-stage pipeline with detection and single-instance segmentation networks is employed.\n- Detection Model.\nThe detection baseline model is Feature Pyramid Network with ResNeXt152 backbone with modulated deformable convolution layers. (see <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110953\">my post at the detection track</a>). </p>\n\n<ul>\n<li>Segmentation Model.\nThe segmentation model is ResNet152-C4 with two upsampling layers and two U-net-like skip connections. </li>\n</ul>\n\n<p>Each instance is cropped from the image based on:\n1) At training time: the ground truth bounding boxes.\n2) At inference time: the bounding boxes detected by the (ensembled) detection model including the parent classes.\nThe cropped images are resized to (320, 320). The output mask resolution is (160, 160).</p>\n\n<p>The models and training pipeline are developed based on the maskrcnn-benchmark repo.</p>\n\n<h2>3. Training</h2>\n\n<p>The training conditions are optimized for single GPU (V100).</p>\n\n<ul>\n<li><p>Detection Model.\nThe detection model has been trained using 500-class box labels and eight models are ensembled (0.597 private LB at object detection track).</p></li>\n<li><p>Segmentation Model.\nThe segmentation model has been trained for 1.8 million iterations and cosine decay is scheduled for the last 0.2 million iterations. Batchsize is 8 and batchnorm layers are used.</p></li>\n</ul>\n\n<h2>4. Ensembling</h2>\n\n<ul>\n<li>Two models ensembling.\nTwo segmentation models with different image sampling seeds are ensembled with  and without horizontal flip. The output heatmaps are averaged.</li>\n<li>Results.\nModel Ensembling improved private LB score from 0.4740 (single segmentation model) to 0.4796.</li>\n</ul>",
  "messages": [
    {
      "id": "641704",
      "postDate": "10/05/2019 01:57:45",
      "content": "<p>I would like to thank the competition organizers and all the competitors! </p>\n\n<p>Here's my brief solution writeup:</p>\n\n<h2>1. Dataset</h2>\n\n<ul>\n<li>No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.</li>\n<li>Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal (1/300) across 300 classes. One instance is randomly picked from an image to train the segmentation network described below.</li>\n</ul>\n\n<h2>2. Pipiline and Models</h2>\n\n<p>A two-stage pipeline with detection and single-instance segmentation networks is employed.\n- Detection Model.\nThe detection baseline model is Feature Pyramid Network with ResNeXt152 backbone with modulated deformable convolution layers. (see <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110953\">my post at the detection track</a>). </p>\n\n<ul>\n<li>Segmentation Model.\nThe segmentation model is ResNet152-C4 with two upsampling layers and two U-net-like skip connections. </li>\n</ul>\n\n<p>Each instance is cropped from the image based on:\n1) At training time: the ground truth bounding boxes.\n2) At inference time: the bounding boxes detected by the (ensembled) detection model including the parent classes.\nThe cropped images are resized to (320, 320). The output mask resolution is (160, 160).</p>\n\n<p>The models and training pipeline are developed based on the maskrcnn-benchmark repo.</p>\n\n<h2>3. Training</h2>\n\n<p>The training conditions are optimized for single GPU (V100).</p>\n\n<ul>\n<li><p>Detection Model.\nThe detection model has been trained using 500-class box labels and eight models are ensembled (0.597 private LB at object detection track).</p></li>\n<li><p>Segmentation Model.\nThe segmentation model has been trained for 1.8 million iterations and cosine decay is scheduled for the last 0.2 million iterations. Batchsize is 8 and batchnorm layers are used.</p></li>\n</ul>\n\n<h2>4. Ensembling</h2>\n\n<ul>\n<li>Two models ensembling.\nTwo segmentation models with different image sampling seeds are ensembled with  and without horizontal flip. The output heatmaps are averaged.</li>\n<li>Results.\nModel Ensembling improved private LB score from 0.4740 (single segmentation model) to 0.4796.</li>\n</ul>",
      "rawMarkdown": "I would like to thank the competition organizers and all the competitors! \n\nHere's my brief solution writeup:\n\n## 1. Dataset\n- No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.\n- Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal (1/300) across 300 classes. One instance is randomly picked from an image to train the segmentation network described below.\n\n## 2. Pipiline and Models\nA two-stage pipeline with detection and single-instance segmentation networks is employed.\n- Detection Model.\nThe detection baseline model is Feature Pyramid Network with ResNeXt152 backbone with modulated deformable convolution layers. (see [my post at the detection track](https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110953)). \n\n- Segmentation Model.\nThe segmentation model is ResNet152-C4 with two upsampling layers and two U-net-like skip connections. \n\nEach instance is cropped from the image based on:\n1) At training time: the ground truth bounding boxes.\n2) At inference time: the bounding boxes detected by the (ensembled) detection model including the parent classes.\nThe cropped images are resized to (320, 320). The output mask resolution is (160, 160).\n\nThe models and training pipeline are developed based on the maskrcnn-benchmark repo.\n\n## 3. Training\nThe training conditions are optimized for single GPU (V100).\n\n- Detection Model.\nThe detection model has been trained using 500-class box labels and eight models are ensembled (0.597 private LB at object detection track).\n\n- Segmentation Model.\nThe segmentation model has been trained for 1.8 million iterations and cosine decay is scheduled for the last 0.2 million iterations. Batchsize is 8 and batchnorm layers are used.\n\n## 4. Ensembling\n- Two models ensembling.\nTwo segmentation models with different image sampling seeds are ensembled with  and without horizontal flip. The output heatmaps are averaged.\n- Results.\nModel Ensembling improved private LB score from 0.4740 (single segmentation model) to 0.4796.",
      "votes": null
    },
    {
      "id": "641778",
      "postDate": "10/05/2019 04:52:12",
      "content": "<p>Congratulations\nGreat Write-Up\nThank You for Sharing Your Approach &amp; Insights….!! <a href=\"/hirotoschwert\">@hirotoschwert</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThank You for Sharing Your Approach &amp; Insights….!! @hirotoschwert",
      "votes": null
    },
    {
      "id": "642211",
      "postDate": "10/05/2019 17:36:22",
      "content": "<p>Did you use only 1 V100 for training? \nIf so, how long did you train it for? </p>",
      "rawMarkdown": "Did you use only 1 V100 for training? \nIf so, how long did you train it for?",
      "votes": null
    },
    {
      "id": "642560",
      "postDate": "10/06/2019 09:37:52",
      "content": "<p>Congrats, thank you for sharing approach</p>",
      "rawMarkdown": "Congrats, thank you for sharing approach",
      "votes": null
    },
    {
      "id": "642585",
      "postDate": "10/06/2019 10:38:37",
      "content": "<p>Thanks for sharing. I’m not sure whether I understand correctly. How did you use detection and segmentation models togеther in one ensemble? </p>",
      "rawMarkdown": "Thanks for sharing. I’m not sure whether I understand correctly. How did you use detection and segmentation models togеther in one ensemble?",
      "votes": null
    },
    {
      "id": "642612",
      "postDate": "10/06/2019 11:34:54",
      "content": "<p>Thank you! It's not ensembling. I use detection results (bounding boxes) only to determine the regions to crop from an input image. The segmentation model infers mask heatmaps using the (resized) cropped images as input. </p>",
      "rawMarkdown": "Thank you! It's not ensembling. I use detection results (bounding boxes) only to determine the regions to crop from an input image. The segmentation model infers mask heatmaps using the (resized) cropped images as input.",
      "votes": null
    },
    {
      "id": "642613",
      "postDate": "10/06/2019 11:46:16",
      "content": "<p>It took about 6 days per segmentation model and around one month per detection model. I sometimes use multiple GPU instances to train models in parallel. </p>",
      "rawMarkdown": "It took about 6 days per segmentation model and around one month per detection model. I sometimes use multiple GPU instances to train models in parallel.",
      "votes": null
    },
    {
      "id": "643533",
      "postDate": "10/07/2019 15:36:12",
      "content": "<p>Interesting approach. Did you try to run your segmentations model on a raw images from initial dataset (maybe on validation split)? What was a score?</p>",
      "rawMarkdown": "Interesting approach. Did you try to run your segmentations model on a raw images from initial dataset (maybe on validation split)? What was a score?",
      "votes": null
    },
    {
      "id": "644140",
      "postDate": "10/08/2019 11:54:08",
      "content": "<p>Thank you, yes it would be a very important validation but I have not done that yet. I will report it when I do additional experiments!</p>",
      "rawMarkdown": "Thank you, yes it would be a very important validation but I have not done that yet. I will report it when I do additional experiments!",
      "votes": null
    },
    {
      "id": "644498",
      "postDate": "10/08/2019 22:39:22",
      "content": "<p>Congrats. Many thanks for sharing.\nA</p>",
      "rawMarkdown": "Congrats. Many thanks for sharing.\nA",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 641778,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/05/2019 04:52:12",
      "content": "<p>Congratulations\nGreat Write-Up\nThank You for Sharing Your Approach &amp; Insights….!! <a href=\"/hirotoschwert\">@hirotoschwert</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 642211,
      "author_name": "rohitmidha23",
      "author_url": "",
      "post_date": "10/05/2019 17:36:22",
      "content": "<p>Did you use only 1 V100 for training? \nIf so, how long did you train it for? </p>",
      "votes": null,
      "replies": [
        {
          "id": 642613,
          "author_name": "hirotoschwert",
          "author_url": "",
          "post_date": "10/06/2019 11:46:16",
          "content": "<p>It took about 6 days per segmentation model and around one month per detection model. I sometimes use multiple GPU instances to train models in parallel. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 642560,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "10/06/2019 09:37:52",
      "content": "<p>Congrats, thank you for sharing approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 642585,
      "author_name": "dsvolkov",
      "author_url": "",
      "post_date": "10/06/2019 10:38:37",
      "content": "<p>Thanks for sharing. I’m not sure whether I understand correctly. How did you use detection and segmentation models togеther in one ensemble? </p>",
      "votes": null,
      "replies": [
        {
          "id": 642612,
          "author_name": "hirotoschwert",
          "author_url": "",
          "post_date": "10/06/2019 11:34:54",
          "content": "<p>Thank you! It's not ensembling. I use detection results (bounding boxes) only to determine the regions to crop from an input image. The segmentation model infers mask heatmaps using the (resized) cropped images as input. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 643533,
          "author_name": "dsvolkov",
          "author_url": "",
          "post_date": "10/07/2019 15:36:12",
          "content": "<p>Interesting approach. Did you try to run your segmentations model on a raw images from initial dataset (maybe on validation split)? What was a score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 644140,
          "author_name": "hirotoschwert",
          "author_url": "",
          "post_date": "10/08/2019 11:54:08",
          "content": "<p>Thank you, yes it would be a very important validation but I have not done that yet. I will report it when I do additional experiments!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 644498,
      "author_name": "zinovadr",
      "author_url": "",
      "post_date": "10/08/2019 22:39:22",
      "content": "<p>Congrats. Many thanks for sharing.\nA</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "641704": "I would like to thank the competition organizers and all the competitors! \n\nHere's my brief solution writeup:\n\n## 1. Dataset\n- No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.\n- Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal (1/300) across 300 classes. One instance is randomly picked from an image to train the segmentation network described below.\n\n## 2. Pipiline and Models\nA two-stage pipeline with detection and single-instance segmentation networks is employed.\n- Detection Model.\nThe detection baseline model is Feature Pyramid Network with ResNeXt152 backbone with modulated deformable convolution layers. (see [my post at the detection track](https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110953)). \n\n- Segmentation Model.\nThe segmentation model is ResNet152-C4 with two upsampling layers and two U-net-like skip connections. \n\nEach instance is cropped from the image based on:\n1) At training time: the ground truth bounding boxes.\n2) At inference time: the bounding boxes detected by the (ensembled) detection model including the parent classes.\nThe cropped images are resized to (320, 320). The output mask resolution is (160, 160).\n\nThe models and training pipeline are developed based on the maskrcnn-benchmark repo.\n\n## 3. Training\nThe training conditions are optimized for single GPU (V100).\n\n- Detection Model.\nThe detection model has been trained using 500-class box labels and eight models are ensembled (0.597 private LB at object detection track).\n\n- Segmentation Model.\nThe segmentation model has been trained for 1.8 million iterations and cosine decay is scheduled for the last 0.2 million iterations. Batchsize is 8 and batchnorm layers are used.\n\n## 4. Ensembling\n- Two models ensembling.\nTwo segmentation models with different image sampling seeds are ensembled with  and without horizontal flip. The output heatmaps are averaged.\n- Results.\nModel Ensembling improved private LB score from 0.4740 (single segmentation model) to 0.4796.",
    "641778": "Congratulations\nGreat Write-Up\nThank You for Sharing Your Approach &amp; Insights….!! @hirotoschwert",
    "642211": "Did you use only 1 V100 for training? \nIf so, how long did you train it for?",
    "642560": "Congrats, thank you for sharing approach",
    "642585": "Thanks for sharing. I’m not sure whether I understand correctly. How did you use detection and segmentation models togеther in one ensemble?",
    "642612": "Thank you! It's not ensembling. I use detection results (bounding boxes) only to determine the regions to crop from an input image. The segmentation model infers mask heatmaps using the (resized) cropped images as input.",
    "642613": "It took about 6 days per segmentation model and around one month per detection model. I sometimes use multiple GPU instances to train models in parallel.",
    "643533": "Interesting approach. Did you try to run your segmentations model on a raw images from initial dataset (maybe on validation split)? What was a score?",
    "644140": "Thank you, yes it would be a very important validation but I have not done that yet. I will report it when I do additional experiments!",
    "644498": "Congrats. Many thanks for sharing.\nA"
  },
  "source": "meta"
}