{
  "id": 110953,
  "title": "6th place solution [0.6023 private LB]",
  "url": "/competitions/open-images-2019-object-detection/discussion/110953",
  "author_name": "Schwert",
  "post_date": "2019-10-02T13:47:01.482000",
  "votes": 56,
  "comment_count": 20,
  "views": 0,
  "content": "<p>First of all, I would like to thank the competition organizers and all the competitors!\nThis was my first Kaggle competition and I really had a great time here 😀 </p>\n\n<p>Here's my brief solution writeup:</p>\n\n<h2>1. Dataset</h2>\n\n<ul>\n<li>No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.</li>\n<li>Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal across 500 classes. For example, a model encounters very rare 'pressure cooker' images with probability of 1/500. For non-rare classes, the number of the images is limited.</li>\n</ul>\n\n<h2>2. Models</h2>\n\n<p>The baseline model is Feature Pyramid Network with ResNeXt152 backbone.\nModulated deformable convolution layers are introduced in the backbone network.\nThe model and training pipeline are developed based on the maskrcnn-benchmark repo.</p>\n\n<h2>3. Training</h2>\n\n<ul>\n<li>Single GPU training.\nThe training conditions are optimized for single GPU (V100) training.\nThe baseline model has been trained for 3 million iterations and cosine decay is scheduled for the last 1.2 million iterations. Batch size is 1 (!) and loss is accumulated for 4 batches.</li>\n<li>Parent class expansion.\nThe models are trained with the ground truth boxes without parent class expansion. Parent boxes are added after inference, which achieves empirically better AP than multi-class training.</li>\n<li>Mini-validation.\nA subset of validation dataset consisting of 5,700 images is used. Validation is performed every 0.2 million iterations using an instance with K80 GPU.</li>\n</ul>\n\n<h2>4. Ensembling</h2>\n\n<ul>\n<li>Ensembling eight models.\nEight models with different image sampling seeds and different model conditions (ResNeXt 152 / 101, with and without DCN) are chosen and ensembled (after NMS).</li>\n<li>Final NMS.\nNMS is performed again on the ensembled bounding boxes class by class. IoU threshold of NMS has been chosen carefully so that the resulting AP is maximized. Scores of box pairs with higher overlap than the threshold are added together.</li>\n<li>Results.\nModel Ensembling improved private LB score from 0.56369 (single model) to 0.60231.</li>\n</ul>",
  "messages": [
    {
      "id": 638855,
      "postDate": "2019-10-02T13:47:01.483Z",
      "content": "<p>First of all, I would like to thank the competition organizers and all the competitors!\nThis was my first Kaggle competition and I really had a great time here 😀 </p>\n\n<p>Here's my brief solution writeup:</p>\n\n<h2>1. Dataset</h2>\n\n<ul>\n<li>No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.</li>\n<li>Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal across 500 classes. For example, a model encounters very rare 'pressure cooker' images with probability of 1/500. For non-rare classes, the number of the images is limited.</li>\n</ul>\n\n<h2>2. Models</h2>\n\n<p>The baseline model is Feature Pyramid Network with ResNeXt152 backbone.\nModulated deformable convolution layers are introduced in the backbone network.\nThe model and training pipeline are developed based on the maskrcnn-benchmark repo.</p>\n\n<h2>3. Training</h2>\n\n<ul>\n<li>Single GPU training.\nThe training conditions are optimized for single GPU (V100) training.\nThe baseline model has been trained for 3 million iterations and cosine decay is scheduled for the last 1.2 million iterations. Batch size is 1 (!) and loss is accumulated for 4 batches.</li>\n<li>Parent class expansion.\nThe models are trained with the ground truth boxes without parent class expansion. Parent boxes are added after inference, which achieves empirically better AP than multi-class training.</li>\n<li>Mini-validation.\nA subset of validation dataset consisting of 5,700 images is used. Validation is performed every 0.2 million iterations using an instance with K80 GPU.</li>\n</ul>\n\n<h2>4. Ensembling</h2>\n\n<ul>\n<li>Ensembling eight models.\nEight models with different image sampling seeds and different model conditions (ResNeXt 152 / 101, with and without DCN) are chosen and ensembled (after NMS).</li>\n<li>Final NMS.\nNMS is performed again on the ensembled bounding boxes class by class. IoU threshold of NMS has been chosen carefully so that the resulting AP is maximized. Scores of box pairs with higher overlap than the threshold are added together.</li>\n<li>Results.\nModel Ensembling improved private LB score from 0.56369 (single model) to 0.60231.</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to thank the competition organizers and all the competitors!\nThis was my first Kaggle competition and I really had a great time here 😀 \n\nHere's my brief solution writeup:\n## 1. Dataset\n- No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.\n- Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal across 500 classes. For example, a model encounters very rare 'pressure cooker' images with probability of 1/500. For non-rare classes, the number of the images is limited.\n\n## 2. Models\nThe baseline model is Feature Pyramid Network with ResNeXt152 backbone.\nModulated deformable convolution layers are introduced in the backbone network.\nThe model and training pipeline are developed based on the maskrcnn-benchmark repo.\n\n## 3. Training\n- Single GPU training.\nThe training conditions are optimized for single GPU (V100) training.\nThe baseline model has been trained for 3 million iterations and cosine decay is scheduled for the last 1.2 million iterations. Batch size is 1 (!) and loss is accumulated for 4 batches.\n- Parent class expansion.\nThe models are trained with the ground truth boxes without parent class expansion. Parent boxes are added after inference, which achieves empirically better AP than multi-class training.\n- Mini-validation.\nA subset of validation dataset consisting of 5,700 images is used. Validation is performed every 0.2 million iterations using an instance with K80 GPU.\n\n## 4. Ensembling\n- Ensembling eight models.\nEight models with different image sampling seeds and different model conditions (ResNeXt 152 / 101, with and without DCN) are chosen and ensembled (after NMS).\n- Final NMS.\nNMS is performed again on the ensembled bounding boxes class by class. IoU threshold of NMS has been chosen carefully so that the resulting AP is maximized. Scores of box pairs with higher overlap than the threshold are added together.\n- Results.\nModel Ensembling improved private LB score from 0.56369 (single model) to 0.60231.",
      "votes": 54
    },
    {
      "id": 662113,
      "postDate": "2019-10-31T05:15:11.070Z",
      "content": "<p>may I ask, your solution is based on detectron or other framework?</p>",
      "rawMarkdown": "may I ask, your solution is based on detectron or other framework?",
      "votes": 1,
      "replies": [
        {
          "id": 668573,
          "postDate": "2019-11-08T15:19:03.070Z",
          "content": "<p>It's based on <a href=\"https://github.com/facebookresearch/maskrcnn-benchmark\">maskrcnn-benchmark</a>.</p>",
          "rawMarkdown": "It's based on [maskrcnn-benchmark](https://github.com/facebookresearch/maskrcnn-benchmark).",
          "votes": 1
        }
      ]
    },
    {
      "id": 648413,
      "postDate": "2019-10-14T06:26:27.550Z",
      "content": "<p>congrats on your medal .\n\"Batch size is 1\", what image size you used?</p>",
      "rawMarkdown": "congrats on your medal .\n\"Batch size is 1\", what image size you used?",
      "votes": 1,
      "replies": [
        {
          "id": 649253,
          "postDate": "2019-10-15T06:22:55.173Z",
          "content": "<p>Thank you!\nThe image size setting during training is \"maximum size of the side of the image (H or W) = 800\", which is the standard configuration of the R-CNN family.</p>",
          "rawMarkdown": "Thank you!\nThe image size setting during training is \"maximum size of the side of the image (H or W) = 800\", which is the standard configuration of the R-CNN family."
        }
      ]
    },
    {
      "id": 641103,
      "postDate": "2019-10-04T12:42:27.677Z",
      "content": "<p>Congratulations! Thanks for sharing. Quite Informative. </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing. Quite Informative. ",
      "votes": 1
    },
    {
      "id": 640189,
      "postDate": "2019-10-03T23:41:09.310Z",
      "content": "<p>Congrats! ensembling makes improvement a lot!</p>",
      "rawMarkdown": "Congrats! ensembling makes improvement a lot!",
      "votes": 1
    },
    {
      "id": 639030,
      "postDate": "2019-10-02T17:40:51.890Z",
      "content": "<p>Congratulations <a href=\"/hirotoschwert\">@hirotoschwert</a> !</p>",
      "rawMarkdown": "Congratulations @hirotoschwert !",
      "votes": 1
    },
    {
      "id": 638885,
      "postDate": "2019-10-02T14:25:00.077Z",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights… <a href=\"/hirotoschwert\">@hirotoschwert</a> </p>",
      "rawMarkdown": "Congrats\nThank you for Sharing your Approach &amp; Insights… @hirotoschwert ",
      "votes": 1
    },
    {
      "id": 638988,
      "postDate": "2019-10-02T16:46:19.130Z",
      "content": "<p>Congrats and thanks for sharing. Just curious to know how long it took to ensemble 8 models. My guess is you would have spent at-least 10 to 15 days to get everything in the current shape. </p>",
      "rawMarkdown": "Congrats and thanks for sharing. Just curious to know how long it took to ensemble 8 models. My guess is you would have spent at-least 10 to 15 days to get everything in the current shape. ",
      "votes": 2,
      "replies": [
        {
          "id": 639339,
          "postDate": "2019-10-03T05:09:25.613Z",
          "content": "<p>Thank you! Yes it took time. It takes 18 to 36 days to train one model. The models are trained in parallel using multiple GPU instances. As for ensembling, it takes almost one day per model for inference on test data and one more day for the final NMS.</p>",
          "rawMarkdown": "Thank you! Yes it took time. It takes 18 to 36 days to train one model. The models are trained in parallel using multiple GPU instances. As for ensembling, it takes almost one day per model for inference on test data and one more day for the final NMS.",
          "votes": 1
        }
      ]
    },
    {
      "id": 638961,
      "postDate": "2019-10-02T16:01:45.773Z",
      "content": "<p>Can you explain about ensemble methods for object detection? I had a hard time with the ensemble...</p>",
      "rawMarkdown": "Can you explain about ensemble methods for object detection? I had a hard time with the ensemble...",
      "votes": 2,
      "replies": [
        {
          "id": 639627,
          "postDate": "2019-10-03T12:16:17.867Z",
          "content": "<p>I simply concatenate the boxes which the models have predicted, and then perform NMS on the gathered boxes.</p>",
          "rawMarkdown": "I simply concatenate the boxes which the models have predicted, and then perform NMS on the gathered boxes.",
          "votes": 3
        }
      ]
    },
    {
      "id": 638907,
      "postDate": "2019-10-02T14:42:05.680Z",
      "content": "<p>Congratulations on winning your solo gold medal. Is it possible for you to open source your code solutions? :-)</p>",
      "rawMarkdown": "Congratulations on winning your solo gold medal. Is it possible for you to open source your code solutions? :-)",
      "votes": 2,
      "replies": [
        {
          "id": 639259,
          "postDate": "2019-10-03T02:05:20.237Z",
          "content": "<p>Thank you! My code (repo) is not ready right now but I will try to make it available soon.</p>",
          "rawMarkdown": "Thank you! My code (repo) is not ready right now but I will try to make it available soon.",
          "votes": 1
        }
      ]
    },
    {
      "id": 638891,
      "postDate": "2019-10-02T14:28:05.867Z",
      "content": "<p>Congrats and thank you for sharing. What is total training time for all models? Did you use some test time tricks like tta and so on? What was your validation strategy?</p>",
      "rawMarkdown": "Congrats and thank you for sharing. What is total training time for all models? Did you use some test time tricks like tta and so on? What was your validation strategy?",
      "votes": 2,
      "replies": [
        {
          "id": 639269,
          "postDate": "2019-10-03T02:26:06.477Z",
          "content": "<p>Thank you! It takes 18 to 36 days to train one model. The models are trained in parallel using multiple GPU instances. For TTA only horizontal flip is used. The validation strategy is added to the 3. Training section. </p>",
          "rawMarkdown": "Thank you! It takes 18 to 36 days to train one model. The models are trained in parallel using multiple GPU instances. For TTA only horizontal flip is used. The validation strategy is added to the 3. Training section. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 766750,
      "postDate": "2020-03-08T16:31:37.623Z",
      "content": "<p>Congratulations! 🤘 </p>",
      "rawMarkdown": "Congratulations! 🤘 \n"
    },
    {
      "id": 680767,
      "postDate": "2019-11-25T07:32:39.770Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 668180,
      "postDate": "2019-11-08T05:01:40.727Z",
      "content": "<p>Thank you for sharing and congrats!</p>",
      "rawMarkdown": "Thank you for sharing and congrats!",
      "votes": 1
    },
    {
      "id": 641728,
      "postDate": "2019-10-05T02:45:42.810Z",
      "content": "<p>Thank you for sharing~</p>",
      "rawMarkdown": "Thank you for sharing~",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 662113,
      "author_name": "Peterzhang",
      "author_url": "",
      "post_date": "2019-10-31T05:15:11.070000",
      "content": "<p>may I ask, your solution is based on detectron or other framework?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 668573,
          "author_name": "Schwert",
          "author_url": "",
          "post_date": "2019-11-08T15:19:03.070000",
          "content": "<p>It's based on <a href=\"https://github.com/facebookresearch/maskrcnn-benchmark\">maskrcnn-benchmark</a>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 648413,
      "author_name": "Bhaskar",
      "author_url": "",
      "post_date": "2019-10-14T06:26:27.550000",
      "content": "<p>congrats on your medal .\n\"Batch size is 1\", what image size you used?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 649253,
          "author_name": "Schwert",
          "author_url": "",
          "post_date": "2019-10-15T06:22:55.173000",
          "content": "<p>Thank you!\nThe image size setting during training is \"maximum size of the side of the image (H or W) = 800\", which is the standard configuration of the R-CNN family.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 641103,
      "author_name": "SanthoRiyu",
      "author_url": "",
      "post_date": "2019-10-04T12:42:27.677000",
      "content": "<p>Congratulations! Thanks for sharing. Quite Informative. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 640189,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-10-03T23:41:09.310000",
      "content": "<p>Congrats! ensembling makes improvement a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 639030,
      "author_name": "Nicolas H",
      "author_url": "",
      "post_date": "2019-10-02T17:40:51.890000",
      "content": "<p>Congratulations <a href=\"/hirotoschwert\">@hirotoschwert</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 638885,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-10-02T14:25:00.077000",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights… <a href=\"/hirotoschwert\">@hirotoschwert</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 638988,
      "author_name": "Manoj Prabhakar",
      "author_url": "",
      "post_date": "2019-10-02T16:46:19.130000",
      "content": "<p>Congrats and thanks for sharing. Just curious to know how long it took to ensemble 8 models. My guess is you would have spent at-least 10 to 15 days to get everything in the current shape. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 639339,
          "author_name": "Schwert",
          "author_url": "",
          "post_date": "2019-10-03T05:09:25.613000",
          "content": "<p>Thank you! Yes it took time. It takes 18 to 36 days to train one model. The models are trained in parallel using multiple GPU instances. As for ensembling, it takes almost one day per model for inference on test data and one more day for the final NMS.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 638961,
      "author_name": "Chanran Kim",
      "author_url": "",
      "post_date": "2019-10-02T16:01:45.773000",
      "content": "<p>Can you explain about ensemble methods for object detection? I had a hard time with the ensemble...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 639627,
          "author_name": "Schwert",
          "author_url": "",
          "post_date": "2019-10-03T12:16:17.867000",
          "content": "<p>I simply concatenate the boxes which the models have predicted, and then perform NMS on the gathered boxes.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 638907,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-10-02T14:42:05.680000",
      "content": "<p>Congratulations on winning your solo gold medal. Is it possible for you to open source your code solutions? :-)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 639259,
          "author_name": "Schwert",
          "author_url": "",
          "post_date": "2019-10-03T02:05:20.237000",
          "content": "<p>Thank you! My code (repo) is not ready right now but I will try to make it available soon.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 638891,
      "author_name": "Dmitry Volkov",
      "author_url": "",
      "post_date": "2019-10-02T14:28:05.867000",
      "content": "<p>Congrats and thank you for sharing. What is total training time for all models? Did you use some test time tricks like tta and so on? What was your validation strategy?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 639269,
          "author_name": "Schwert",
          "author_url": "",
          "post_date": "2019-10-03T02:26:06.477000",
          "content": "<p>Thank you! It takes 18 to 36 days to train one model. The models are trained in parallel using multiple GPU instances. For TTA only horizontal flip is used. The validation strategy is added to the 3. Training section. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 766750,
      "author_name": "Mahdi",
      "author_url": "",
      "post_date": "2020-03-08T16:31:37.623000",
      "content": "<p>Congratulations! 🤘 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 680767,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-25T07:32:39.770000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 668180,
      "author_name": "sjsjs",
      "author_url": "",
      "post_date": "2019-11-08T05:01:40.727000",
      "content": "<p>Thank you for sharing and congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 641728,
      "author_name": "Yilin Li",
      "author_url": "",
      "post_date": "2019-10-05T02:45:42.810000",
      "content": "<p>Thank you for sharing~</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "638855": "First of all, I would like to thank the competition organizers and all the competitors!\nThis was my first Kaggle competition and I really had a great time here 😀 \n\nHere's my brief solution writeup:\n## 1. Dataset\n- No external dataset.\nI only use FAIR's ImageNet pretrained weights for initialization, as I have described in the Official External Data Thread.\n- Class balancing.\nFor each class, images are sampled so that probability to have at least one instance of the class is equal across 500 classes. For example, a model encounters very rare 'pressure cooker' images with probability of 1/500. For non-rare classes, the number of the images is limited.\n\n## 2. Models\nThe baseline model is Feature Pyramid Network with ResNeXt152 backbone.\nModulated deformable convolution layers are introduced in the backbone network.\nThe model and training pipeline are developed based on the maskrcnn-benchmark repo.\n\n## 3. Training\n- Single GPU training.\nThe training conditions are optimized for single GPU (V100) training.\nThe baseline model has been trained for 3 million iterations and cosine decay is scheduled for the last 1.2 million iterations. Batch size is 1 (!) and loss is accumulated for 4 batches.\n- Parent class expansion.\nThe models are trained with the ground truth boxes without parent class expansion. Parent boxes are added after inference, which achieves empirically better AP than multi-class training.\n- Mini-validation.\nA subset of validation dataset consisting of 5,700 images is used. Validation is performed every 0.2 million iterations using an instance with K80 GPU.\n\n## 4. Ensembling\n- Ensembling eight models.\nEight models with different image sampling seeds and different model conditions (ResNeXt 152 / 101, with and without DCN) are chosen and ensembled (after NMS).\n- Final NMS.\nNMS is performed again on the ensembled bounding boxes class by class. IoU threshold of NMS has been chosen carefully so that the resulting AP is maximized. Scores of box pairs with higher overlap than the threshold are added together.\n- Results.\nModel Ensembling improved private LB score from 0.56369 (single model) to 0.60231.",
    "662113": "may I ask, your solution is based on detectron or other framework?",
    "648413": "congrats on your medal .\n\"Batch size is 1\", what image size you used?",
    "641103": "Congratulations! Thanks for sharing. Quite Informative. ",
    "640189": "Congrats! ensembling makes improvement a lot!",
    "639030": "Congratulations @hirotoschwert !",
    "638885": "Congrats\nThank you for Sharing your Approach &amp; Insights… @hirotoschwert ",
    "638988": "Congrats and thanks for sharing. Just curious to know how long it took to ensemble 8 models. My guess is you would have spent at-least 10 to 15 days to get everything in the current shape. ",
    "638961": "Can you explain about ensemble methods for object detection? I had a hard time with the ensemble...",
    "638907": "Congratulations on winning your solo gold medal. Is it possible for you to open source your code solutions? :-)",
    "638891": "Congrats and thank you for sharing. What is total training time for all models? Did you use some test time tricks like tta and so on? What was your validation strategy?",
    "766750": "Congratulations! 🤘 \n",
    "680767": "",
    "668180": "Thank you for sharing and congrats!",
    "641728": "Thank you for sharing~"
  }
}