{
  "id": 65120,
  "title": "8th place solution [0.49 on private LB]",
  "url": "/competitions/google-ai-open-images-object-detection-track/writeups/ringukraine-cloudresearch-8th-place-solution-0-49-",
  "author_name": "",
  "post_date": "2018-09-06T11:54:18.522Z",
  "votes": 30,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Using out-the-box solutions (pre-trained models)</h1>\n\n<p>To make initial submit we used pre-trained *faster_rcnn_inception_resnet_v2_atrous_oid* model from <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.mdhttps://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\">TF Model Zoo</a> . We achieved 0.216 on public LB. As this model was frozen with threshold 0.3 and we decided to refroze this model without any thresholds. After this model gave us 0.36 on public LB and we use it as base for our final submissions.</p>\n\n<h1>Data and Metric Analysis</h1>\n\n<ul>\n<li><strong>Unbalanced data.</strong> First thing we noticed that data is very unbalanced. For example Top-50 classes (by count of b-boxes) is 87.9% of training b-boxes or 84.3% of training images. On the other hand we have less than 1% of b-boxes and images for Bottom-50. It means that we have 200K samples per class for Top-50 and only 200 samples per class for Bottom-50. Due to that fact that evaluation metric is mAP all classes have the same impact on the final score.</li>\n<li><strong>Hierarchy structure.</strong> All classes in OID organized into one huge classes tree. Therefore there are a lot of classes that share common features and this is must help a lot for model, but such hierarchical structure also rise new problems: model confused by almost same classes (what class will have Girl age of 19?  Girl or Woman?) and meta-classes start to contain some trash (can not define is this Woman or Man? Let’s mark as Person) </li>\n<li><strong>Bounding Boxes.</strong> We analyzed b-boxes by its area size and aspect ratios but didn’t find strong correlation. By default aspect ratios in RPN Anchor Generator are 1, 2, 0.5. We added ⅓, 3, ¼, 4 ratios but it did not give notable improvement.</li>\n<li><strong>Missing annotations.</strong> We visualized classes b-boxes and looked threw some training images. We noticed that there are many images which contains partially annotated instances for large number of classes. Missing annotations would lead to false negatives during training. <img src=\"https://cdn1.savepice.ru/uploads/2018/9/6/e11465bd14a191d264444dc53911414f-full.png\" alt=\"Image\"></li>\n<li><strong>Validation dataset.</strong> We used validation set proposed by organizators. During submissions we noticed that results on Public LB are 0.13 lower than on validation. This gap left almost unchangeable from submission to submission.</li>\n</ul>\n\n<h1>Detection Architecture</h1>\n\n<p>As our experiments based on framework from tensorflow/models we choose Faster R-CNN + backbone. Also we tried Single Shot Multibox Detector approach but it gave much worse results for OID (even for small amount of classes).\nTraining single model on the whole dataset\nAt first stage we tried to train on all 500 classes and to figure out what accuracy we can achieve with single model. We experimented with different backbones for Faster R-CNN:</p>\n\n<ul>\n<li><strong>Inception v2.</strong> One of the fastest models for training/inference (except MobileNet) in TF Model Zoo. We knew that we won’t achieve comparable results to Inception ResNet v2 but we could check different solutions relatively fast. Best mAP: 0.298 on public LB</li>\n<li><strong>Inception ResNet v2.</strong> As it had pre-trained weights on OID we fine-tune it but stuck on mAP ~0.38 on public LB</li>\n<li><strong>NasNet Large.</strong> Consume large amount of resources during training/inference (2-3x in comparison to Inception ResNet v2). Best mAP ~ 0.42 on public LB</li>\n</ul>\n\n<h1>Dataset Split</h1>\n\n<p>Taking into account that data is very unbalanced we split it into 6 subsets according to its data distribution. Then we trained Faster R-CNN + ResNet-101 on each subset and gained improvement about 0.12 in comparison to single model.</p>\n\n<ul>\n<li><strong>Bottom 0-100:</strong> 263 imgs/class</li>\n<li><strong>Bottom 100-200:</strong> 691 imgs/class</li>\n<li><strong>Bottom 200-300:</strong> 1552 imgs/class</li>\n<li><strong>Bottom 300-400:</strong> 4206 imgs/class</li>\n<li><strong>Bottom 400-450:</strong> 14375 imgs/class</li>\n<li><strong>Bottom 450-500:</strong> 202171 imgs/class</li>\n</ul>\n\n<h1>Missed annotations:</h1>\n\n<p>Open Image Dataset contain a lot of unlabeled object. This is produce noisy and some time even wrong training signal for our models. To leverage this problem we tried following ideas:</p>\n\n<ul>\n<li><strong>Pseudo-labeling.</strong> This is common approach that used on competitions to get some additional data and to fix bad annotation. We applied this trick to only to bottom 0-100 classes subset with prediction threshold &gt; 0.5. This trick increase mAP for bottom 0-100 classes subset by 1%. </li>\n<li><strong>Overlap Soft Sampling (OSS, <a href=\"https://arxiv.org/abs/1806.06986\">paper</a>).</strong> Another idea is to reweight training samples depending on highest IOU with any ground of truth. Core idea here is when sample contain significant part of any annotated object then lower probability of wrong training signal. This approach in average increase performance on 0.3%, but if look closer, than we can sees that some classes decreased in performance up to 50% AP and some classes increased in performance up to 50% AP, so this is was really good news, in terms of ensembling.  </li>\n</ul>\n\n<h1>Ensembling</h1>\n\n<ul>\n<li><strong>Best only.</strong> In this approach for each class we simply select predictions from model that predicts this class the best, based on validation.  We used it at starting point. </li>\n<li><strong>Performance-based with restore procedure.</strong> This approach share same idea as “Best only”, but instead selecting just one model to predict some class, we select several models that perform best for this class and add they predictions. Before adding some predictions to final model we rescore all confidences of this model for this class by following formula: <img src=\"https://cdn1.savepice.ru/uploads/2018/9/6/0d7f56dae42f4d016d54ade327af674d-full.png\" alt=\"Image\">\nHere performance is class average precision and N - number of models that predict this class. We achieved ~1% improvement in comparison to base strategy (Best only) </li>\n</ul>\n\n<h1>Result Pipeline</h1>\n\n<ol>\n<li>Split training set into 6 subset based on training sample (see dataset split topic).</li>\n<li>Training models. Generally we train Faster R-CNN with ResNet-101 backbone for each subset.</li>\n<li>For each class we selected models (based on evaluation results) that will be used to predict this class. Note that we used pretrained Faster R-CNN with \nInception-ResNet v2 backbone as base model.</li>\n<li>Then we concatenate all selected prediction from each model, expand all classes with meta-classes predictions and apply Soft-NMS to suppress multiple predictions.</li>\n</ol>\n\n<h1>Hints and tricks:</h1>\n\n<ul>\n<li>We increase output size for RPN and Result predictor from 300 to 500. Improved mAP ~0.1%.</li>\n<li>We expanded predicted b-boxes from leaves on superclasses. Improved mAP ~0.5%</li>\n<li>Using <a href=\"https://arxiv.org/abs/1704.04503\">Soft-NMS</a> above default NMS improved mAP ~2.5% </li>\n</ul>\n\n<h1>Other approaches that did not improve performance:</h1>\n\n<ul>\n<li>Classify image classes and suppress false positives from detectors</li>\n<li>Split prediction head into classification and regression</li>\n<li>RMSProp/Adam instead SGD with momentum</li>\n<li>Dropout</li>\n<li>Using ImageLevel annotations during training</li>\n<li>Augmentations: Random Distort Color, Random Adjust Brightness,  Random Black Patches, Random Gaussian Noise</li>\n</ul>",
  "messages": [
    {
      "id": "382454",
      "postDate": "09/06/2018 11:54:18",
      "content": "<h1>Using out-the-box solutions (pre-trained models)</h1>\n\n<p>To make initial submit we used pre-trained *faster_rcnn_inception_resnet_v2_atrous_oid* model from <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.mdhttps://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\">TF Model Zoo</a> . We achieved 0.216 on public LB. As this model was frozen with threshold 0.3 and we decided to refroze this model without any thresholds. After this model gave us 0.36 on public LB and we use it as base for our final submissions.</p>\n\n<h1>Data and Metric Analysis</h1>\n\n<ul>\n<li><strong>Unbalanced data.</strong> First thing we noticed that data is very unbalanced. For example Top-50 classes (by count of b-boxes) is 87.9% of training b-boxes or 84.3% of training images. On the other hand we have less than 1% of b-boxes and images for Bottom-50. It means that we have 200K samples per class for Top-50 and only 200 samples per class for Bottom-50. Due to that fact that evaluation metric is mAP all classes have the same impact on the final score.</li>\n<li><strong>Hierarchy structure.</strong> All classes in OID organized into one huge classes tree. Therefore there are a lot of classes that share common features and this is must help a lot for model, but such hierarchical structure also rise new problems: model confused by almost same classes (what class will have Girl age of 19?  Girl or Woman?) and meta-classes start to contain some trash (can not define is this Woman or Man? Let’s mark as Person) </li>\n<li><strong>Bounding Boxes.</strong> We analyzed b-boxes by its area size and aspect ratios but didn’t find strong correlation. By default aspect ratios in RPN Anchor Generator are 1, 2, 0.5. We added ⅓, 3, ¼, 4 ratios but it did not give notable improvement.</li>\n<li><strong>Missing annotations.</strong> We visualized classes b-boxes and looked threw some training images. We noticed that there are many images which contains partially annotated instances for large number of classes. Missing annotations would lead to false negatives during training. <img src=\"https://cdn1.savepice.ru/uploads/2018/9/6/e11465bd14a191d264444dc53911414f-full.png\" alt=\"Image\"></li>\n<li><strong>Validation dataset.</strong> We used validation set proposed by organizators. During submissions we noticed that results on Public LB are 0.13 lower than on validation. This gap left almost unchangeable from submission to submission.</li>\n</ul>\n\n<h1>Detection Architecture</h1>\n\n<p>As our experiments based on framework from tensorflow/models we choose Faster R-CNN + backbone. Also we tried Single Shot Multibox Detector approach but it gave much worse results for OID (even for small amount of classes).\nTraining single model on the whole dataset\nAt first stage we tried to train on all 500 classes and to figure out what accuracy we can achieve with single model. We experimented with different backbones for Faster R-CNN:</p>\n\n<ul>\n<li><strong>Inception v2.</strong> One of the fastest models for training/inference (except MobileNet) in TF Model Zoo. We knew that we won’t achieve comparable results to Inception ResNet v2 but we could check different solutions relatively fast. Best mAP: 0.298 on public LB</li>\n<li><strong>Inception ResNet v2.</strong> As it had pre-trained weights on OID we fine-tune it but stuck on mAP ~0.38 on public LB</li>\n<li><strong>NasNet Large.</strong> Consume large amount of resources during training/inference (2-3x in comparison to Inception ResNet v2). Best mAP ~ 0.42 on public LB</li>\n</ul>\n\n<h1>Dataset Split</h1>\n\n<p>Taking into account that data is very unbalanced we split it into 6 subsets according to its data distribution. Then we trained Faster R-CNN + ResNet-101 on each subset and gained improvement about 0.12 in comparison to single model.</p>\n\n<ul>\n<li><strong>Bottom 0-100:</strong> 263 imgs/class</li>\n<li><strong>Bottom 100-200:</strong> 691 imgs/class</li>\n<li><strong>Bottom 200-300:</strong> 1552 imgs/class</li>\n<li><strong>Bottom 300-400:</strong> 4206 imgs/class</li>\n<li><strong>Bottom 400-450:</strong> 14375 imgs/class</li>\n<li><strong>Bottom 450-500:</strong> 202171 imgs/class</li>\n</ul>\n\n<h1>Missed annotations:</h1>\n\n<p>Open Image Dataset contain a lot of unlabeled object. This is produce noisy and some time even wrong training signal for our models. To leverage this problem we tried following ideas:</p>\n\n<ul>\n<li><strong>Pseudo-labeling.</strong> This is common approach that used on competitions to get some additional data and to fix bad annotation. We applied this trick to only to bottom 0-100 classes subset with prediction threshold &gt; 0.5. This trick increase mAP for bottom 0-100 classes subset by 1%. </li>\n<li><strong>Overlap Soft Sampling (OSS, <a href=\"https://arxiv.org/abs/1806.06986\">paper</a>).</strong> Another idea is to reweight training samples depending on highest IOU with any ground of truth. Core idea here is when sample contain significant part of any annotated object then lower probability of wrong training signal. This approach in average increase performance on 0.3%, but if look closer, than we can sees that some classes decreased in performance up to 50% AP and some classes increased in performance up to 50% AP, so this is was really good news, in terms of ensembling.  </li>\n</ul>\n\n<h1>Ensembling</h1>\n\n<ul>\n<li><strong>Best only.</strong> In this approach for each class we simply select predictions from model that predicts this class the best, based on validation.  We used it at starting point. </li>\n<li><strong>Performance-based with restore procedure.</strong> This approach share same idea as “Best only”, but instead selecting just one model to predict some class, we select several models that perform best for this class and add they predictions. Before adding some predictions to final model we rescore all confidences of this model for this class by following formula: <img src=\"https://cdn1.savepice.ru/uploads/2018/9/6/0d7f56dae42f4d016d54ade327af674d-full.png\" alt=\"Image\">\nHere performance is class average precision and N - number of models that predict this class. We achieved ~1% improvement in comparison to base strategy (Best only) </li>\n</ul>\n\n<h1>Result Pipeline</h1>\n\n<ol>\n<li>Split training set into 6 subset based on training sample (see dataset split topic).</li>\n<li>Training models. Generally we train Faster R-CNN with ResNet-101 backbone for each subset.</li>\n<li>For each class we selected models (based on evaluation results) that will be used to predict this class. Note that we used pretrained Faster R-CNN with \nInception-ResNet v2 backbone as base model.</li>\n<li>Then we concatenate all selected prediction from each model, expand all classes with meta-classes predictions and apply Soft-NMS to suppress multiple predictions.</li>\n</ol>\n\n<h1>Hints and tricks:</h1>\n\n<ul>\n<li>We increase output size for RPN and Result predictor from 300 to 500. Improved mAP ~0.1%.</li>\n<li>We expanded predicted b-boxes from leaves on superclasses. Improved mAP ~0.5%</li>\n<li>Using <a href=\"https://arxiv.org/abs/1704.04503\">Soft-NMS</a> above default NMS improved mAP ~2.5% </li>\n</ul>\n\n<h1>Other approaches that did not improve performance:</h1>\n\n<ul>\n<li>Classify image classes and suppress false positives from detectors</li>\n<li>Split prediction head into classification and regression</li>\n<li>RMSProp/Adam instead SGD with momentum</li>\n<li>Dropout</li>\n<li>Using ImageLevel annotations during training</li>\n<li>Augmentations: Random Distort Color, Random Adjust Brightness,  Random Black Patches, Random Gaussian Noise</li>\n</ul>",
      "rawMarkdown": "# Using out-the-box solutions (pre-trained models) \nTo make initial submit we used pre-trained *faster_rcnn_inception_resnet_v2_atrous_oid* model from [TF Model Zoo](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.mdhttps://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md) . We achieved 0.216 on public LB. As this model was frozen with threshold 0.3 and we decided to refroze this model without any thresholds. After this model gave us 0.36 on public LB and we use it as base for our final submissions.\n\n# Data and Metric Analysis \n* **Unbalanced data.** First thing we noticed that data is very unbalanced. For example Top-50 classes (by count of b-boxes) is 87.9% of training b-boxes or 84.3% of training images. On the other hand we have less than 1% of b-boxes and images for Bottom-50. It means that we have 200K samples per class for Top-50 and only 200 samples per class for Bottom-50. Due to that fact that evaluation metric is mAP all classes have the same impact on the final score.\n* **Hierarchy structure.** All classes in OID organized into one huge classes tree. Therefore there are a lot of classes that share common features and this is must help a lot for model, but such hierarchical structure also rise new problems: model confused by almost same classes (what class will have Girl age of 19?  Girl or Woman?) and meta-classes start to contain some trash (can not define is this Woman or Man? Let’s mark as Person) \n* **Bounding Boxes.** We analyzed b-boxes by its area size and aspect ratios but didn’t find strong correlation. By default aspect ratios in RPN Anchor Generator are 1, 2, 0.5. We added ⅓, 3, ¼, 4 ratios but it did not give notable improvement.\n* **Missing annotations.** We visualized classes b-boxes and looked threw some training images. We noticed that there are many images which contains partially annotated instances for large number of classes. Missing annotations would lead to false negatives during training. ![Image](https://cdn1.savepice.ru/uploads/2018/9/6/e11465bd14a191d264444dc53911414f-full.png)\n* **Validation dataset.** We used validation set proposed by organizators. During submissions we noticed that results on Public LB are 0.13 lower than on validation. This gap left almost unchangeable from submission to submission.\n\n# Detection Architecture\nAs our experiments based on framework from tensorflow/models we choose Faster R-CNN + backbone. Also we tried Single Shot Multibox Detector approach but it gave much worse results for OID (even for small amount of classes).\nTraining single model on the whole dataset\nAt first stage we tried to train on all 500 classes and to figure out what accuracy we can achieve with single model. We experimented with different backbones for Faster R-CNN:\n\n* **Inception v2.** One of the fastest models for training/inference (except MobileNet) in TF Model Zoo. We knew that we won’t achieve comparable results to Inception ResNet v2 but we could check different solutions relatively fast. Best mAP: 0.298 on public LB\n* **Inception ResNet v2.** As it had pre-trained weights on OID we fine-tune it but stuck on mAP ~0.38 on public LB\n* **NasNet Large.** Consume large amount of resources during training/inference (2-3x in comparison to Inception ResNet v2). Best mAP ~ 0.42 on public LB\n\n# Dataset Split\nTaking into account that data is very unbalanced we split it into 6 subsets according to its data distribution. Then we trained Faster R-CNN + ResNet-101 on each subset and gained improvement about 0.12 in comparison to single model.\n\n* **Bottom 0-100:** 263 imgs/class\n* **Bottom 100-200:** 691 imgs/class\n* **Bottom 200-300:** 1552 imgs/class\n* **Bottom 300-400:** 4206 imgs/class\n* **Bottom 400-450:** 14375 imgs/class\n* **Bottom 450-500:** 202171 imgs/class\n\n# Missed annotations:\nOpen Image Dataset contain a lot of unlabeled object. This is produce noisy and some time even wrong training signal for our models. To leverage this problem we tried following ideas:\n\n* **Pseudo-labeling.** This is common approach that used on competitions to get some additional data and to fix bad annotation. We applied this trick to only to bottom 0-100 classes subset with prediction threshold &gt; 0.5. This trick increase mAP for bottom 0-100 classes subset by 1%. \n* **Overlap Soft Sampling (OSS, [paper](https://arxiv.org/abs/1806.06986)).** Another idea is to reweight training samples depending on highest IOU with any ground of truth. Core idea here is when sample contain significant part of any annotated object then lower probability of wrong training signal. This approach in average increase performance on 0.3%, but if look closer, than we can sees that some classes decreased in performance up to 50% AP and some classes increased in performance up to 50% AP, so this is was really good news, in terms of ensembling.  \n\n# Ensembling\n* **Best only.** In this approach for each class we simply select predictions from model that predicts this class the best, based on validation.  We used it at starting point. \n* **Performance-based with restore procedure.** This approach share same idea as “Best only”, but instead selecting just one model to predict some class, we select several models that perform best for this class and add they predictions. Before adding some predictions to final model we rescore all confidences of this model for this class by following formula: ![Image](https://cdn1.savepice.ru/uploads/2018/9/6/0d7f56dae42f4d016d54ade327af674d-full.png)\nHere performance is class average precision and N - number of models that predict this class. We achieved ~1% improvement in comparison to base strategy (Best only) \n\n# Result Pipeline\n1. Split training set into 6 subset based on training sample (see dataset split topic).\n2. Training models. Generally we train Faster R-CNN with ResNet-101 backbone for each subset.\n3. For each class we selected models (based on evaluation results) that will be used to predict this class. Note that we used pretrained Faster R-CNN with \nInception-ResNet v2 backbone as base model.\n4. Then we concatenate all selected prediction from each model, expand all classes with meta-classes predictions and apply Soft-NMS to suppress multiple predictions.\n\n# Hints and tricks:\n* We increase output size for RPN and Result predictor from 300 to 500. Improved mAP ~0.1%.\n* We expanded predicted b-boxes from leaves on superclasses. Improved mAP ~0.5%\n* Using [Soft-NMS](https://arxiv.org/abs/1704.04503) above default NMS improved mAP ~2.5% \n\n# Other approaches that did not improve performance:\n* Classify image classes and suppress false positives from detectors\n* Split prediction head into classification and regression\n* RMSProp/Adam instead SGD with momentum\n* Dropout\n* Using ImageLevel annotations during training\n* Augmentations: Random Distort Color, Random Adjust Brightness,  Random Black Patches, Random Gaussian Noise",
      "votes": null
    },
    {
      "id": "399538",
      "postDate": "10/06/2018 03:49:41",
      "content": "<p>Thank you for your sharing. How many GPUs or CPUs you used for training?</p>",
      "rawMarkdown": "Thank you for your sharing. How many GPUs or CPUs you used for training?",
      "votes": null
    },
    {
      "id": "400734",
      "postDate": "10/08/2018 19:33:55",
      "content": "<p>We used 3 machines, each of it has Intel Core i7-7700k (4 cores/8 threads), 64Gb RAM, 2x Nvidia Geforce GTX 1080 Ti (11Gb) on a board. We trained only one model on each instance (3 experiments/models in parallel).</p>",
      "rawMarkdown": "We used 3 machines, each of it has Intel Core i7-7700k (4 cores/8 threads), 64Gb RAM, 2x Nvidia Geforce GTX 1080 Ti (11Gb) on a board. We trained only one model on each instance (3 experiments/models in parallel).",
      "votes": null
    },
    {
      "id": "459568",
      "postDate": "01/22/2019 02:18:04",
      "content": "<p>Hi, i want ask which part you implement Overlap-Soft-Sampling method on faster-rcnn, RPN and fast-rcnn or just fast-rcnn?</p>",
      "rawMarkdown": "Hi, i want ask which part you implement Overlap-Soft-Sampling method on faster-rcnn, RPN and fast-rcnn or just fast-rcnn?",
      "votes": null
    },
    {
      "id": "466716",
      "postDate": "02/05/2019 21:14:40",
      "content": "<p>We applied OSS only for RPN, because unannotated objects have higher impact for RPN`s training signal due to the high density of default (anchor) boxes.</p>",
      "rawMarkdown": "We applied OSS only for RPN, because unannotated objects have higher impact for RPN`s training signal due to the high density of default (anchor) boxes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 399538,
      "author_name": "dingyan",
      "author_url": "",
      "post_date": "10/06/2018 03:49:41",
      "content": "<p>Thank you for your sharing. How many GPUs or CPUs you used for training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 400734,
          "author_name": "alexkirnas",
          "author_url": "",
          "post_date": "10/08/2018 19:33:55",
          "content": "<p>We used 3 machines, each of it has Intel Core i7-7700k (4 cores/8 threads), 64Gb RAM, 2x Nvidia Geforce GTX 1080 Ti (11Gb) on a board. We trained only one model on each instance (3 experiments/models in parallel).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 459568,
      "author_name": "kdfans",
      "author_url": "",
      "post_date": "01/22/2019 02:18:04",
      "content": "<p>Hi, i want ask which part you implement Overlap-Soft-Sampling method on faster-rcnn, RPN and fast-rcnn or just fast-rcnn?</p>",
      "votes": null,
      "replies": [
        {
          "id": 466716,
          "author_name": "alexkirnas",
          "author_url": "",
          "post_date": "02/05/2019 21:14:40",
          "content": "<p>We applied OSS only for RPN, because unannotated objects have higher impact for RPN`s training signal due to the high density of default (anchor) boxes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "382454": "# Using out-the-box solutions (pre-trained models) \nTo make initial submit we used pre-trained *faster_rcnn_inception_resnet_v2_atrous_oid* model from [TF Model Zoo](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.mdhttps://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md) . We achieved 0.216 on public LB. As this model was frozen with threshold 0.3 and we decided to refroze this model without any thresholds. After this model gave us 0.36 on public LB and we use it as base for our final submissions.\n\n# Data and Metric Analysis \n* **Unbalanced data.** First thing we noticed that data is very unbalanced. For example Top-50 classes (by count of b-boxes) is 87.9% of training b-boxes or 84.3% of training images. On the other hand we have less than 1% of b-boxes and images for Bottom-50. It means that we have 200K samples per class for Top-50 and only 200 samples per class for Bottom-50. Due to that fact that evaluation metric is mAP all classes have the same impact on the final score.\n* **Hierarchy structure.** All classes in OID organized into one huge classes tree. Therefore there are a lot of classes that share common features and this is must help a lot for model, but such hierarchical structure also rise new problems: model confused by almost same classes (what class will have Girl age of 19?  Girl or Woman?) and meta-classes start to contain some trash (can not define is this Woman or Man? Let’s mark as Person) \n* **Bounding Boxes.** We analyzed b-boxes by its area size and aspect ratios but didn’t find strong correlation. By default aspect ratios in RPN Anchor Generator are 1, 2, 0.5. We added ⅓, 3, ¼, 4 ratios but it did not give notable improvement.\n* **Missing annotations.** We visualized classes b-boxes and looked threw some training images. We noticed that there are many images which contains partially annotated instances for large number of classes. Missing annotations would lead to false negatives during training. ![Image](https://cdn1.savepice.ru/uploads/2018/9/6/e11465bd14a191d264444dc53911414f-full.png)\n* **Validation dataset.** We used validation set proposed by organizators. During submissions we noticed that results on Public LB are 0.13 lower than on validation. This gap left almost unchangeable from submission to submission.\n\n# Detection Architecture\nAs our experiments based on framework from tensorflow/models we choose Faster R-CNN + backbone. Also we tried Single Shot Multibox Detector approach but it gave much worse results for OID (even for small amount of classes).\nTraining single model on the whole dataset\nAt first stage we tried to train on all 500 classes and to figure out what accuracy we can achieve with single model. We experimented with different backbones for Faster R-CNN:\n\n* **Inception v2.** One of the fastest models for training/inference (except MobileNet) in TF Model Zoo. We knew that we won’t achieve comparable results to Inception ResNet v2 but we could check different solutions relatively fast. Best mAP: 0.298 on public LB\n* **Inception ResNet v2.** As it had pre-trained weights on OID we fine-tune it but stuck on mAP ~0.38 on public LB\n* **NasNet Large.** Consume large amount of resources during training/inference (2-3x in comparison to Inception ResNet v2). Best mAP ~ 0.42 on public LB\n\n# Dataset Split\nTaking into account that data is very unbalanced we split it into 6 subsets according to its data distribution. Then we trained Faster R-CNN + ResNet-101 on each subset and gained improvement about 0.12 in comparison to single model.\n\n* **Bottom 0-100:** 263 imgs/class\n* **Bottom 100-200:** 691 imgs/class\n* **Bottom 200-300:** 1552 imgs/class\n* **Bottom 300-400:** 4206 imgs/class\n* **Bottom 400-450:** 14375 imgs/class\n* **Bottom 450-500:** 202171 imgs/class\n\n# Missed annotations:\nOpen Image Dataset contain a lot of unlabeled object. This is produce noisy and some time even wrong training signal for our models. To leverage this problem we tried following ideas:\n\n* **Pseudo-labeling.** This is common approach that used on competitions to get some additional data and to fix bad annotation. We applied this trick to only to bottom 0-100 classes subset with prediction threshold &gt; 0.5. This trick increase mAP for bottom 0-100 classes subset by 1%. \n* **Overlap Soft Sampling (OSS, [paper](https://arxiv.org/abs/1806.06986)).** Another idea is to reweight training samples depending on highest IOU with any ground of truth. Core idea here is when sample contain significant part of any annotated object then lower probability of wrong training signal. This approach in average increase performance on 0.3%, but if look closer, than we can sees that some classes decreased in performance up to 50% AP and some classes increased in performance up to 50% AP, so this is was really good news, in terms of ensembling.  \n\n# Ensembling\n* **Best only.** In this approach for each class we simply select predictions from model that predicts this class the best, based on validation.  We used it at starting point. \n* **Performance-based with restore procedure.** This approach share same idea as “Best only”, but instead selecting just one model to predict some class, we select several models that perform best for this class and add they predictions. Before adding some predictions to final model we rescore all confidences of this model for this class by following formula: ![Image](https://cdn1.savepice.ru/uploads/2018/9/6/0d7f56dae42f4d016d54ade327af674d-full.png)\nHere performance is class average precision and N - number of models that predict this class. We achieved ~1% improvement in comparison to base strategy (Best only) \n\n# Result Pipeline\n1. Split training set into 6 subset based on training sample (see dataset split topic).\n2. Training models. Generally we train Faster R-CNN with ResNet-101 backbone for each subset.\n3. For each class we selected models (based on evaluation results) that will be used to predict this class. Note that we used pretrained Faster R-CNN with \nInception-ResNet v2 backbone as base model.\n4. Then we concatenate all selected prediction from each model, expand all classes with meta-classes predictions and apply Soft-NMS to suppress multiple predictions.\n\n# Hints and tricks:\n* We increase output size for RPN and Result predictor from 300 to 500. Improved mAP ~0.1%.\n* We expanded predicted b-boxes from leaves on superclasses. Improved mAP ~0.5%\n* Using [Soft-NMS](https://arxiv.org/abs/1704.04503) above default NMS improved mAP ~2.5% \n\n# Other approaches that did not improve performance:\n* Classify image classes and suppress false positives from detectors\n* Split prediction head into classification and regression\n* RMSProp/Adam instead SGD with momentum\n* Dropout\n* Using ImageLevel annotations during training\n* Augmentations: Random Distort Color, Random Adjust Brightness,  Random Black Patches, Random Gaussian Noise",
    "399538": "Thank you for your sharing. How many GPUs or CPUs you used for training?",
    "400734": "We used 3 machines, each of it has Intel Core i7-7700k (4 cores/8 threads), 64Gb RAM, 2x Nvidia Geforce GTX 1080 Ti (11Gb) on a board. We trained only one model on each instance (3 experiments/models in parallel).",
    "459568": "Hi, i want ask which part you implement Overlap-Soft-Sampling method on faster-rcnn, RPN and fast-rcnn or just fast-rcnn?",
    "466716": "We applied OSS only for RPN, because unannotated objects have higher impact for RPN`s training signal due to the high density of default (anchor) boxes."
  },
  "source": "meta"
}