{
  "id": 64633,
  "title": "15th place solution [0.45 private LB]",
  "url": "/competitions/google-ai-open-images-object-detection-track/writeups/ods-ai-zfturbo-15th-place-solution-0-45-private-lb",
  "author_name": "",
  "post_date": "2018-08-31T15:30:39.013Z",
  "votes": 80,
  "comment_count": 18,
  "views": 0,
  "content": "<h1>Software</h1>\n\n<p>Windows 10 + Python 3.5 + Keras 2.2 + Keras-RetinaNet [0.4.1]: <a href=\"https://github.com/fizyr/keras-retinanet\">https://github.com/fizyr/keras-retinanet</a></p>\n\n<h1>Main approach</h1>\n\n<p>All classes were split on 5 levels depends on children level. As the basis I used official classes hierarchy: <a href=\"https://storage.googleapis.com/openimages/challenge_2018/bbox_labels_500_hierarchy_visualizer/circle.html\">https://storage.googleapis.com/openimages/challenge_2018/bbox_labels_500_hierarchy_visualizer/circle.html</a></p>\n\n<ul>\n<li><strong>Level 1</strong>: 443 classes - no children, have maximum impact on score</li>\n<li><strong>Level 2</strong>: 46 classes - have children only with level 1</li>\n<li><strong>Level 3</strong>: 4 classes: 'Seafood', 'Watercraft', 'Insect', 'Carnivore'</li>\n<li><strong>Level 4</strong>: 4 classes: 'Vegetable', 'Land vehicle', 'Reptile', 'Invertebrate'</li>\n<li><strong>Level 5</strong>: 3 classes: 'Furniture', 'Vehicle', 'Animal'</li>\n</ul>\n\n<p><strong>Note 1</strong>: It’s possible to split classes on 3rd, 4th and 5th levels differently. Only “Animal” class has real 5th level.</p>\n\n<p><strong>Note 2</strong>: 3rd, 4th and 5th don’t have large impact on final score so I didn’t really tune these models.</p>\n\n<p><strong>Note 3</strong>: Models of 2-5 levels optional, since we can generate predictions for them using Level 1 model. We just need to duplicate boxes for their children and use some NMS algo on them since there will appear some duplicates. Score for this approach will be slightly lower than using separate 2-5 level models (as I remember correctly ~0.01-0.02 lower).</p>\n\n<h1>Why we need the split? Why we don’t train using all 500 classes?</h1>\n\n<p>In process of dataset analysis I found out that there are some images which marked up only on higher level classes. The most explicit representatives are the Person class (2nd level) and Man, Woman, Boy, Girl (1st level). What's the problem? Images marked with the Person class do not have markup for Man, Woman, Boy and Girl. If we use these images for training at once for all classes - the model will be confused in the classes Man, Woman, Boy and Girl. You can throw out images for Person, but then it makes no sense to train this class as part of the overall model, but it's easier to generate markup for Person in the inference step (using boxes for Man, Woman, Boy and Girl).</p>\n\n<p>To train 2-5 levels models, we can use more data, including images marked for example by Person besides images including subclasses (Man, Woman, Boy and Girl) and this improves the result.</p>\n\n<h1>Preparing data for training</h1>\n\n<p><strong>Level 1</strong>: All images containing markup for one of the 443 classes were added + images were added that contained the markup of the first-level classes not included in Challenge 500, as negative samples. Excluded all images containing markup of the higher levels in order to reduce the number of potential False negative (that is, the markup should be, but it is not in the training set).</p>\n\n<p><strong>Levels 2 - 5</strong>: Images for each class contain both markups directly for this class, and for all children classes. Excluded images containing the higher level classes. Added images containing disjoint classes as negative samples.</p>\n\n<p>Validation is rather slow on RetinaNet. I did a separate validation for the training process, where for each class there were about 25 images containing this class. Thus, I balanced the classes a bit and reduced the time of the validation. My validation during training was about 6.5K images instead of 41K.</p>\n\n<h1>Best models</h1>\n\n<pre><code>1) Keras RetinaNet + ResNet152 Image size ranges 600-800 px\nmAP on small validation: 0.5028\nmAP on full validation: 0.384009\nLB score*: 0.47441\n\n2) Keras RetinaNet + ResNet101 Image size ranges 768-1024 px\nmAP on small validation: 0.4896\nmAP on full validation: 0.377631\nLB score*: 0.47549\n</code></pre>\n\n<ul>\n<li><ul><li>I have not tried individual submissions by levels, just sent a merged result for the best models for all levels.</li></ul></li>\n</ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10220/ResNet101-loss.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10218/ResNet101-valid-mAP.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10219/ResNet152-loss.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10217/ResNet152-valid-mAP.png\" alt=\"enter image description here\"></p>\n\n<p>Models continue improving when I stopped training them, so I think there is some room to increase score.</p>\n\n<h1>Process of training and inference</h1>\n\n<p>Keras-RetinaNet already has a implemented generator for training on Open Images Dataset (OID). But in the current form it is not very suitable. I made several changes to it:</p>\n\n<ul>\n<li>I added support for “empty” class - for images with negative samples that do not contain any boxes.</li>\n<li>In the current generator there are no augmentations associated with the color, I added a random change in the intensity of the channels - this gave a good increase in validation score. And it feels like augmentation set in Keras-Retinanet is not enough for training a strong model.</li>\n<li>Most important, selecting just random images for the batch will work poor at random. I replaced it with the following method:\n<ul><li>before the start of training for each class (including \"empty\") we create a list of images that contains boxes for the class.</li>\n<li>during the training, to add the next image to the batch, we first randomly select a class, then randomly select the image from the image list for the class. Thus, we achieve more or less uniform training by classes. Strictly speaking, the distribution is still not quite uniform because the images usually contain boxes for several classes at once.</li></ul></li>\n<li>From small things: I increased values ​​for augmentations transform-generator, especially scale. I changed the value of factor from 0.1 to 0.9 in ReduceLROnPlateau because the learning rate dropped too fast.</li>\n<li>At the stage of the convert model for inference, in the FilterDetections layer in RetinaNet - reduced the score_threshold from 0.05 to 0.01 and tried to increase the number of boxes from 300 to 500. This gives more flexibility in the ensemble stage. Also note that the default model RetinaNet for Inference already contains NMS in the FilterDetections layer with nms_threshold = 0.5 and there is a feeling that you can play with this parameter.</li>\n<li>I restarted training several times with a large LR = 1e-5, when LR fell too low. And each time the model became better. But the experiment is not very clean, because every time I changed the augmentation parameters.</li>\n</ul>\n\n<h1>Ensembles</h1>\n\n<ul>\n<li>For each model, I made a prediction on the image and its horizontal mirror.</li>\n<li>RetinaNet has internal layer with made NMS on single image prediction (it can be switched off to get the full set of boxes, but I didn’t try it).</li>\n<li>I merged boxes using two methods:\n<ul><li>standard NMS (<a href=\"https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py\"></a><a href=\"https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py\">https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py</a>), the optimal threshold I found was about 0.75. </li>\n<li>second method is my own heuristic. It works better on LB. Heuristics included a weighted addition of boxes, which changed their coordinates and Confidence score.</li></ul></li>\n<li>I also tried Soft NMS [https://github.com/bharatsingh430/soft-nms] last 3 days of competition, but it works almost the same as default NMS and worse than my ensemble approach. Probably I missed something.</li>\n</ul>\n\n<h1>My ensemble approach in short</h1>\n\n<ol>\n<li>On input we get set of boxes from different N models</li>\n<li>Set “Init_weight” = 1/N and initialize “result” set of boxes as empty list</li>\n<li>Add all boxes for best model with “init_weight” weight to “result” list. For all other boxes of other models try to find best matching box in result using IOU. If box with IOU &gt; THR (0.55) exists in “result” then merge best found box in result with it using weighted average for coordinates and confidence score. Increase weight of this box by Init_weight value. Otherwise add this new box to “result” with “init_weight” weight.</li>\n</ol>\n\n<h1>Other solutions</h1>\n\n<p>At the first stage I tried to solve the problem simply on pretrain models without any retraining:</p>\n\n<ul>\n<li>RetinanNet Pretrain Coco: <a href=\"https://github.com/fizyr/keras-retinanet/releases\">https://github.com/fizyr/keras-retinanet/releases</a> - gives an approximately 0.13 on LB, if the classes between OID and COCO are correctly matched.</li>\n<li>On the pretrain from here: <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/object_detection_tutorial.ipynb\">https://github.com/tensorflow/models/blob/master/research/object_detection/object_detection_tutorial.ipynb</a> using this model: faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28 you can get ~0.25 on LB.</li>\n</ul>\n\n<h1>Observations and small tricks</h1>\n\n<ol>\n<li>Because of the metric nature, it is better to output the maximum number of rectangles even with low confidence score.</li>\n<li>Validation works so-so. In most cases, it is worse than LB. Most likely this is due to two factors: the distribution in kaggle test is very different from the distribution for validation. Classes with a small number of elements affect the score as much as others.</li>\n<li>The limitation on the size of the CSV file on Kaggle in this task is ~2GB. This fact didn't allow to submit more boxes with low confidence score.</li>\n</ol>\n\n<h1>Proposed dataset improvement</h1>\n\n<p>In case we will have similar competitions next years:</p>\n\n<ul>\n<li>To simplify the task organizers should improve the training dataset, so that there are no situations when there is a markup for level 2 and there is no markup for level 1. </li>\n<li>It would be good to have a full markup for each image with all the boxes including the parent classes. To avoid discrepancies / errors, etc.</li>\n<li>It’s better to use actual class names like \"Ambulance\" instead of /m/012n7d</li>\n<li>A little confusing is the presence of the flags like “isGroupOf” - are there any images of such type in the Test set or not? I eventually excluded these images from training and still not sure if it was right thing to do.</li>\n<li>It seems to me that the current metric is not very good due to the fact that classes with a very small number of boxes influence just like classes with millions of boxes. Small mistakes can lead to significant score changes.</li>\n</ul>\n\n<h1>Code and PreTrained models</h1>\n\n<p>GitHub repository: <a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018</a></p>",
  "messages": [
    {
      "id": "379163",
      "postDate": "08/31/2018 00:37:35",
      "content": "<h1>Software</h1>\n\n<p>Windows 10 + Python 3.5 + Keras 2.2 + Keras-RetinaNet [0.4.1]: <a href=\"https://github.com/fizyr/keras-retinanet\">https://github.com/fizyr/keras-retinanet</a></p>\n\n<h1>Main approach</h1>\n\n<p>All classes were split on 5 levels depends on children level. As the basis I used official classes hierarchy: <a href=\"https://storage.googleapis.com/openimages/challenge_2018/bbox_labels_500_hierarchy_visualizer/circle.html\">https://storage.googleapis.com/openimages/challenge_2018/bbox_labels_500_hierarchy_visualizer/circle.html</a></p>\n\n<ul>\n<li><strong>Level 1</strong>: 443 classes - no children, have maximum impact on score</li>\n<li><strong>Level 2</strong>: 46 classes - have children only with level 1</li>\n<li><strong>Level 3</strong>: 4 classes: 'Seafood', 'Watercraft', 'Insect', 'Carnivore'</li>\n<li><strong>Level 4</strong>: 4 classes: 'Vegetable', 'Land vehicle', 'Reptile', 'Invertebrate'</li>\n<li><strong>Level 5</strong>: 3 classes: 'Furniture', 'Vehicle', 'Animal'</li>\n</ul>\n\n<p><strong>Note 1</strong>: It’s possible to split classes on 3rd, 4th and 5th levels differently. Only “Animal” class has real 5th level.</p>\n\n<p><strong>Note 2</strong>: 3rd, 4th and 5th don’t have large impact on final score so I didn’t really tune these models.</p>\n\n<p><strong>Note 3</strong>: Models of 2-5 levels optional, since we can generate predictions for them using Level 1 model. We just need to duplicate boxes for their children and use some NMS algo on them since there will appear some duplicates. Score for this approach will be slightly lower than using separate 2-5 level models (as I remember correctly ~0.01-0.02 lower).</p>\n\n<h1>Why we need the split? Why we don’t train using all 500 classes?</h1>\n\n<p>In process of dataset analysis I found out that there are some images which marked up only on higher level classes. The most explicit representatives are the Person class (2nd level) and Man, Woman, Boy, Girl (1st level). What's the problem? Images marked with the Person class do not have markup for Man, Woman, Boy and Girl. If we use these images for training at once for all classes - the model will be confused in the classes Man, Woman, Boy and Girl. You can throw out images for Person, but then it makes no sense to train this class as part of the overall model, but it's easier to generate markup for Person in the inference step (using boxes for Man, Woman, Boy and Girl).</p>\n\n<p>To train 2-5 levels models, we can use more data, including images marked for example by Person besides images including subclasses (Man, Woman, Boy and Girl) and this improves the result.</p>\n\n<h1>Preparing data for training</h1>\n\n<p><strong>Level 1</strong>: All images containing markup for one of the 443 classes were added + images were added that contained the markup of the first-level classes not included in Challenge 500, as negative samples. Excluded all images containing markup of the higher levels in order to reduce the number of potential False negative (that is, the markup should be, but it is not in the training set).</p>\n\n<p><strong>Levels 2 - 5</strong>: Images for each class contain both markups directly for this class, and for all children classes. Excluded images containing the higher level classes. Added images containing disjoint classes as negative samples.</p>\n\n<p>Validation is rather slow on RetinaNet. I did a separate validation for the training process, where for each class there were about 25 images containing this class. Thus, I balanced the classes a bit and reduced the time of the validation. My validation during training was about 6.5K images instead of 41K.</p>\n\n<h1>Best models</h1>\n\n<pre><code>1) Keras RetinaNet + ResNet152 Image size ranges 600-800 px\nmAP on small validation: 0.5028\nmAP on full validation: 0.384009\nLB score*: 0.47441\n\n2) Keras RetinaNet + ResNet101 Image size ranges 768-1024 px\nmAP on small validation: 0.4896\nmAP on full validation: 0.377631\nLB score*: 0.47549\n</code></pre>\n\n<ul>\n<li><ul><li>I have not tried individual submissions by levels, just sent a merged result for the best models for all levels.</li></ul></li>\n</ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10220/ResNet101-loss.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10218/ResNet101-valid-mAP.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10219/ResNet152-loss.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10217/ResNet152-valid-mAP.png\" alt=\"enter image description here\"></p>\n\n<p>Models continue improving when I stopped training them, so I think there is some room to increase score.</p>\n\n<h1>Process of training and inference</h1>\n\n<p>Keras-RetinaNet already has a implemented generator for training on Open Images Dataset (OID). But in the current form it is not very suitable. I made several changes to it:</p>\n\n<ul>\n<li>I added support for “empty” class - for images with negative samples that do not contain any boxes.</li>\n<li>In the current generator there are no augmentations associated with the color, I added a random change in the intensity of the channels - this gave a good increase in validation score. And it feels like augmentation set in Keras-Retinanet is not enough for training a strong model.</li>\n<li>Most important, selecting just random images for the batch will work poor at random. I replaced it with the following method:\n<ul><li>before the start of training for each class (including \"empty\") we create a list of images that contains boxes for the class.</li>\n<li>during the training, to add the next image to the batch, we first randomly select a class, then randomly select the image from the image list for the class. Thus, we achieve more or less uniform training by classes. Strictly speaking, the distribution is still not quite uniform because the images usually contain boxes for several classes at once.</li></ul></li>\n<li>From small things: I increased values ​​for augmentations transform-generator, especially scale. I changed the value of factor from 0.1 to 0.9 in ReduceLROnPlateau because the learning rate dropped too fast.</li>\n<li>At the stage of the convert model for inference, in the FilterDetections layer in RetinaNet - reduced the score_threshold from 0.05 to 0.01 and tried to increase the number of boxes from 300 to 500. This gives more flexibility in the ensemble stage. Also note that the default model RetinaNet for Inference already contains NMS in the FilterDetections layer with nms_threshold = 0.5 and there is a feeling that you can play with this parameter.</li>\n<li>I restarted training several times with a large LR = 1e-5, when LR fell too low. And each time the model became better. But the experiment is not very clean, because every time I changed the augmentation parameters.</li>\n</ul>\n\n<h1>Ensembles</h1>\n\n<ul>\n<li>For each model, I made a prediction on the image and its horizontal mirror.</li>\n<li>RetinaNet has internal layer with made NMS on single image prediction (it can be switched off to get the full set of boxes, but I didn’t try it).</li>\n<li>I merged boxes using two methods:\n<ul><li>standard NMS (<a href=\"https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py\"></a><a href=\"https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py\">https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py</a>), the optimal threshold I found was about 0.75. </li>\n<li>second method is my own heuristic. It works better on LB. Heuristics included a weighted addition of boxes, which changed their coordinates and Confidence score.</li></ul></li>\n<li>I also tried Soft NMS [https://github.com/bharatsingh430/soft-nms] last 3 days of competition, but it works almost the same as default NMS and worse than my ensemble approach. Probably I missed something.</li>\n</ul>\n\n<h1>My ensemble approach in short</h1>\n\n<ol>\n<li>On input we get set of boxes from different N models</li>\n<li>Set “Init_weight” = 1/N and initialize “result” set of boxes as empty list</li>\n<li>Add all boxes for best model with “init_weight” weight to “result” list. For all other boxes of other models try to find best matching box in result using IOU. If box with IOU &gt; THR (0.55) exists in “result” then merge best found box in result with it using weighted average for coordinates and confidence score. Increase weight of this box by Init_weight value. Otherwise add this new box to “result” with “init_weight” weight.</li>\n</ol>\n\n<h1>Other solutions</h1>\n\n<p>At the first stage I tried to solve the problem simply on pretrain models without any retraining:</p>\n\n<ul>\n<li>RetinanNet Pretrain Coco: <a href=\"https://github.com/fizyr/keras-retinanet/releases\">https://github.com/fizyr/keras-retinanet/releases</a> - gives an approximately 0.13 on LB, if the classes between OID and COCO are correctly matched.</li>\n<li>On the pretrain from here: <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/object_detection_tutorial.ipynb\">https://github.com/tensorflow/models/blob/master/research/object_detection/object_detection_tutorial.ipynb</a> using this model: faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28 you can get ~0.25 on LB.</li>\n</ul>\n\n<h1>Observations and small tricks</h1>\n\n<ol>\n<li>Because of the metric nature, it is better to output the maximum number of rectangles even with low confidence score.</li>\n<li>Validation works so-so. In most cases, it is worse than LB. Most likely this is due to two factors: the distribution in kaggle test is very different from the distribution for validation. Classes with a small number of elements affect the score as much as others.</li>\n<li>The limitation on the size of the CSV file on Kaggle in this task is ~2GB. This fact didn't allow to submit more boxes with low confidence score.</li>\n</ol>\n\n<h1>Proposed dataset improvement</h1>\n\n<p>In case we will have similar competitions next years:</p>\n\n<ul>\n<li>To simplify the task organizers should improve the training dataset, so that there are no situations when there is a markup for level 2 and there is no markup for level 1. </li>\n<li>It would be good to have a full markup for each image with all the boxes including the parent classes. To avoid discrepancies / errors, etc.</li>\n<li>It’s better to use actual class names like \"Ambulance\" instead of /m/012n7d</li>\n<li>A little confusing is the presence of the flags like “isGroupOf” - are there any images of such type in the Test set or not? I eventually excluded these images from training and still not sure if it was right thing to do.</li>\n<li>It seems to me that the current metric is not very good due to the fact that classes with a very small number of boxes influence just like classes with millions of boxes. Small mistakes can lead to significant score changes.</li>\n</ul>\n\n<h1>Code and PreTrained models</h1>\n\n<p>GitHub repository: <a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018</a></p>",
      "rawMarkdown": "Software\n========\n\nWindows 10 + Python 3.5 + Keras 2.2 + Keras-RetinaNet [0.4.1]: https://github.com/fizyr/keras-retinanet\n\nMain approach\n=============\n\nAll classes were split on 5 levels depends on children level. As the basis I used official classes hierarchy: https://storage.googleapis.com/openimages/challenge_2018/bbox_labels_500_hierarchy_visualizer/circle.html\n\n - **Level 1**: 443 classes - no children, have maximum impact on score\n - **Level 2**: 46 classes - have children only with level 1\n - **Level 3**: 4 classes: 'Seafood', 'Watercraft', 'Insect', 'Carnivore'\n - **Level 4**: 4 classes: 'Vegetable', 'Land vehicle', 'Reptile', 'Invertebrate'\n - **Level 5**: 3 classes: 'Furniture', 'Vehicle', 'Animal'\n\n**Note 1**: It’s possible to split classes on 3rd, 4th and 5th levels differently. Only “Animal” class has real 5th level.\n\n**Note 2**: 3rd, 4th and 5th don’t have large impact on final score so I didn’t really tune these models.\n\n**Note 3**: Models of 2-5 levels optional, since we can generate predictions for them using Level 1 model. We just need to duplicate boxes for their children and use some NMS algo on them since there will appear some duplicates. Score for this approach will be slightly lower than using separate 2-5 level models (as I remember correctly ~0.01-0.02 lower).\n\nWhy we need the split? Why we don’t train using all 500 classes?\n================================================================\n\nIn process of dataset analysis I found out that there are some images which marked up only on higher level classes. The most explicit representatives are the Person class (2nd level) and Man, Woman, Boy, Girl (1st level). What's the problem? Images marked with the Person class do not have markup for Man, Woman, Boy and Girl. If we use these images for training at once for all classes - the model will be confused in the classes Man, Woman, Boy and Girl. You can throw out images for Person, but then it makes no sense to train this class as part of the overall model, but it's easier to generate markup for Person in the inference step (using boxes for Man, Woman, Boy and Girl).\n\nTo train 2-5 levels models, we can use more data, including images marked for example by Person besides images including subclasses (Man, Woman, Boy and Girl) and this improves the result.\n\nPreparing data for training\n============================\n\n**Level 1**: All images containing markup for one of the 443 classes were added + images were added that contained the markup of the first-level classes not included in Challenge 500, as negative samples. Excluded all images containing markup of the higher levels in order to reduce the number of potential False negative (that is, the markup should be, but it is not in the training set).\n\n**Levels 2 - 5**: Images for each class contain both markups directly for this class, and for all children classes. Excluded images containing the higher level classes. Added images containing disjoint classes as negative samples.\n\nValidation is rather slow on RetinaNet. I did a separate validation for the training process, where for each class there were about 25 images containing this class. Thus, I balanced the classes a bit and reduced the time of the validation. My validation during training was about 6.5K images instead of 41K.\n\nBest models\n===========\n\n    1) Keras RetinaNet + ResNet152 Image size ranges 600-800 px\n    mAP on small validation: 0.5028\n    mAP on full validation: 0.384009\n    LB score*: 0.47441\n    \n    2) Keras RetinaNet + ResNet101 Image size ranges 768-1024 px\n    mAP on small validation: 0.4896\n    mAP on full validation: 0.377631\n    LB score*: 0.47549\n\n* - I have not tried individual submissions by levels, just sent a merged result for the best models for all levels.\n\n![enter image description here][1]\n\n![enter image description here][2]\n\n![enter image description here][3]\n\n![enter image description here][4]\n\nModels continue improving when I stopped training them, so I think there is some room to increase score.\n\nProcess of training and inference\n=================================\n\nKeras-RetinaNet already has a implemented generator for training on Open Images Dataset (OID). But in the current form it is not very suitable. I made several changes to it:\n\n - I added support for “empty” class - for images with negative samples that do not contain any boxes.\n - In the current generator there are no augmentations associated with the color, I added a random change in the intensity of the channels - this gave a good increase in validation score. And it feels like augmentation set in Keras-Retinanet is not enough for training a strong model.\n - Most important, selecting just random images for the batch will work poor at random. I replaced it with the following method:\n- before the start of training for each class (including \"empty\") we create a list of images that contains boxes for the class.\n- during the training, to add the next image to the batch, we first randomly select a class, then randomly select the image from the image list for the class. Thus, we achieve more or less uniform training by classes. Strictly speaking, the distribution is still not quite uniform because the images usually contain boxes for several classes at once.\n - From small things: I increased values ​​for augmentations transform-generator, especially scale. I changed the value of factor from 0.1 to 0.9 in ReduceLROnPlateau because the learning rate dropped too fast.\n - At the stage of the convert model for inference, in the FilterDetections layer in RetinaNet - reduced the score_threshold from 0.05 to 0.01 and tried to increase the number of boxes from 300 to 500. This gives more flexibility in the ensemble stage. Also note that the default model RetinaNet for Inference already contains NMS in the FilterDetections layer with nms_threshold = 0.5 and there is a feeling that you can play with this parameter.\n - I restarted training several times with a large LR = 1e-5, when LR fell too low. And each time the model became better. But the experiment is not very clean, because every time I changed the augmentation parameters.\n\nEnsembles\n=========\n\n - For each model, I made a prediction on the image and its horizontal mirror.\n - RetinaNet has internal layer with made NMS on single image prediction (it can be switched off to get the full set of boxes, but I didn’t try it).\n - I merged boxes using two methods:\n- standard NMS (https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py), the optimal threshold I found was about 0.75. \n- second method is my own heuristic. It works better on LB. Heuristics included a weighted addition of boxes, which changed their coordinates and Confidence score.\n - I also tried Soft NMS [https://github.com/bharatsingh430/soft-nms] last 3 days of competition, but it works almost the same as default NMS and worse than my ensemble approach. Probably I missed something.\n\nMy ensemble approach in short\n=============================\n\n 1. On input we get set of boxes from different N models\n 2. Set “Init_weight” = 1/N and initialize “result” set of boxes as empty list\n 3. Add all boxes for best model with “init_weight” weight to “result” list. For all other boxes of other models try to find best matching box in result using IOU. If box with IOU &gt; THR (0.55) exists in “result” then merge best found box in result with it using weighted average for coordinates and confidence score. Increase weight of this box by Init_weight value. Otherwise add this new box to “result” with “init_weight” weight.\n\nOther solutions\n===============\n\nAt the first stage I tried to solve the problem simply on pretrain models without any retraining:\n\n - RetinanNet Pretrain Coco: https://github.com/fizyr/keras-retinanet/releases - gives an approximately 0.13 on LB, if the classes between OID and COCO are correctly matched.\n - On the pretrain from here: https://github.com/tensorflow/models/blob/master/research/object_detection/object_detection_tutorial.ipynb using this model: faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28 you can get ~0.25 on LB.\n\nObservations and small tricks\n=============================\n\n 1. Because of the metric nature, it is better to output the maximum number of rectangles even with low confidence score.\n 2. Validation works so-so. In most cases, it is worse than LB. Most likely this is due to two factors: the distribution in kaggle test is very different from the distribution for validation. Classes with a small number of elements affect the score as much as others.\n 3. The limitation on the size of the CSV file on Kaggle in this task is ~2GB. This fact didn't allow to submit more boxes with low confidence score.\n\nProposed dataset improvement\n============================\n\nIn case we will have similar competitions next years:\n\n - To simplify the task organizers should improve the training dataset, so that there are no situations when there is a markup for level 2 and there is no markup for level 1. \n - It would be good to have a full markup for each image with all the boxes including the parent classes. To avoid discrepancies / errors, etc.\n - It’s better to use actual class names like \"Ambulance\" instead of /m/012n7d\n - A little confusing is the presence of the flags like “isGroupOf” - are there any images of such type in the Test set or not? I eventually excluded these images from training and still not sure if it was right thing to do.\n - It seems to me that the current metric is not very good due to the fact that classes with a very small number of boxes influence just like classes with millions of boxes. Small mistakes can lead to significant score changes.\n\nCode and PreTrained models\n==========================\n\nGitHub repository: https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10220/ResNet101-loss.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10218/ResNet101-valid-mAP.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10219/ResNet152-loss.png\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10217/ResNet152-valid-mAP.png",
      "votes": null
    },
    {
      "id": "379356",
      "postDate": "08/31/2018 08:39:51",
      "content": "<p>Congratulations! And nice to hear you've used keras-retinanet. Some others tried but reached nice, but lower scores than yours. As a contributor to this project, can i ask you (if it is possible, useful) to push the changes you've done to OID generator or others into a PR on github? Thanks and again, congratulations!</p>",
      "rawMarkdown": "Congratulations! And nice to hear you've used keras-retinanet. Some others tried but reached nice, but lower scores than yours. As a contributor to this project, can i ask you (if it is possible, useful) to push the changes you've done to OID generator or others into a PR on github? Thanks and again, congratulations!",
      "votes": null
    },
    {
      "id": "379390",
      "postDate": "08/31/2018 09:55:28",
      "content": "<p>That would be great since we also happened to use keras-retinanet for this competition.\nThanks again for sharing with us your solution.</p>",
      "rawMarkdown": "That would be great since we also happened to use keras-retinanet for this competition.\nThanks again for sharing with us your solution.",
      "votes": null
    },
    {
      "id": "379403",
      "postDate": "08/31/2018 10:10:10",
      "content": "<p>Interesting and educational... We tried using keras-retinanet but it stopped learning at about 0.50 mAP. Was res50 though.\nA question if you don't mind. How long did it take per step and per inference? I got really slow times (800ms) on a formidable machine. Might have been configuration problem. </p>",
      "rawMarkdown": "Interesting and educational... We tried using keras-retinanet but it stopped learning at about 0.50 mAP. Was res50 though.\nA question if you don't mind. How long did it take per step and per inference? I got really slow times (800ms) on a formidable machine. Might have been configuration problem.",
      "votes": null
    },
    {
      "id": "379470",
      "postDate": "08/31/2018 12:09:01",
      "content": "<p>I used NVIDIA GTX 1080 Ti</p>\n\n<ul>\n<li>Training with ResNet101 (img size 768-1024) one epoch 10000 images:  ~6400 sec </li>\n<li>Training with ResNet152 (img size 600-800) one epoch 10000 images: ~5200 sec</li>\n</ul>\n\n<p>Inference can't say exactly but something like 500 ms per image I guess.</p>",
      "rawMarkdown": "I used NVIDIA GTX 1080 Ti\n\n - Training with ResNet101 (img size 768-1024) one epoch 10000 images:  ~6400 sec \n - Training with ResNet152 (img size 600-800) one epoch 10000 images: ~5200 sec\n\nInference can't say exactly but something like 500 ms per image I guess.",
      "votes": null
    },
    {
      "id": "379476",
      "postDate": "08/31/2018 12:14:54",
      "content": "<p>Thanks. ) I'm currently preparing code with my solution for github. I'm not sure if I will be able to directly make PR to keras-retinanet since there are some dependencies on my code and directory structure. But code for updated generator will be available in my repository.</p>",
      "rawMarkdown": "Thanks. ) I'm currently preparing code with my solution for github. I'm not sure if I will be able to directly make PR to keras-retinanet since there are some dependencies on my code and directory structure. But code for updated generator will be available in my repository.",
      "votes": null
    },
    {
      "id": "379731",
      "postDate": "08/31/2018 20:49:33",
      "content": "<p>Mmmm ok, so there must have been something with my setup. Thank you for your quick reply. </p>",
      "rawMarkdown": "Mmmm ok, so there must have been something with my setup. Thank you for your quick reply.",
      "votes": null
    },
    {
      "id": "380684",
      "postDate": "09/03/2018 08:13:42",
      "content": "<p>Hi @ZFTurbo,\nCongratulations. Thanks a lot for sharing your nice approach. I tried pre-trained model 'faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28' using object_detection_tutorial.ipynb, however, ended with a low score one. I just wondering if you can share your modified version of 'object_detection_tutorial.ipynb'. I really appreciate your contribution to the community.</p>",
      "rawMarkdown": "Hi @ZFTurbo,\nCongratulations. Thanks a lot for sharing your nice approach. I tried pre-trained model 'faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28' using object_detection_tutorial.ipynb, however, ended with a low score one. I just wondering if you can share your modified version of 'object_detection_tutorial.ipynb'. I really appreciate your contribution to the community.",
      "votes": null
    },
    {
      "id": "380698",
      "postDate": "09/03/2018 08:56:02",
      "content": "<p>Congratulations! And also, very comprehensive explanation, well structured, important details highlighted and useful tips.</p>",
      "rawMarkdown": "Congratulations! And also, very comprehensive explanation, well structured, important details highlighted and useful tips.",
      "votes": null
    },
    {
      "id": "380699",
      "postDate": "09/03/2018 08:57:22",
      "content": "<p>Please add here a note when you upload your code. It will be very interesting to read the detailed solution.</p>",
      "rawMarkdown": "Please add here a note when you upload your code. It will be very interesting to read the detailed solution.",
      "votes": null
    },
    {
      "id": "380718",
      "postDate": "09/03/2018 10:11:17",
      "content": "<p>Code already available:\n<a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018</a></p>\n\n<p>There are inference and training examples.</p>",
      "rawMarkdown": "Code already available:\nhttps://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\n\nThere are inference and training examples.",
      "votes": null
    },
    {
      "id": "380751",
      "postDate": "09/03/2018 11:44:55",
      "content": "<p>Thank you very much.</p>",
      "rawMarkdown": "Thank you very much.",
      "votes": null
    },
    {
      "id": "380949",
      "postDate": "09/03/2018 19:16:23",
      "content": "<p>I put the TensorFlow code in the same repo. You can find details here:</p>\n\n<p><a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018/tree/master/tf_solution\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018/tree/master/tf_solution</a></p>\n\n<p>Some contestants tell me that there is possbility to lower threshold for this pretrained model (graph reexport?) and it will start to generate more boxes with lower confidence score. This leads to much higher mAP. One contestant reports 0.37872 public LB, 0.34736 private LB without any changes in model. But I didn't check it.</p>",
      "rawMarkdown": "I put the TensorFlow code in the same repo. You can find details here:\n\nhttps://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018/tree/master/tf_solution\n\nSome contestants tell me that there is possbility to lower threshold for this pretrained model (graph reexport?) and it will start to generate more boxes with lower confidence score. This leads to much higher mAP. One contestant reports 0.37872 public LB, 0.34736 private LB without any changes in model. But I didn't check it.",
      "votes": null
    },
    {
      "id": "381130",
      "postDate": "09/04/2018 05:48:11",
      "content": "<p>Congratulations! Thanks for this explanation!</p>",
      "rawMarkdown": "Congratulations! Thanks for this explanation!",
      "votes": null
    },
    {
      "id": "387100",
      "postDate": "09/14/2018 10:02:31",
      "content": "<h2>Congratulations on this amazing win : ) Thank you for sharing the code. Our team is trying to implement your solution, we were using 1 gpu Nvidia V100 , at training we ResourceExhaustedError so we added additional 7 gpu but then got the following error for the Adam optimizer for LR. We trace this to the create model and we suspect that model is not compiling correctly with multiple gpus. Any insights on either of these issues and how many gpus did you use would be highly appreciated. Thank you again :) </h2>\n\n<p>AttributeError                            Traceback (most recent call last)\n in ()\n     18         DATASET_PATH,\n     19     ]\n---&gt; 20     main(params)</p>\n\n<p> in main(args)\n     52     print (device_lib.list_local_devices())\n     53 \n---&gt; 54     print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n     55     # K.set_value(model.optimizer.lr, 1e-5)\n     56     print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))</p>\n\n<p>AttributeError: 'NoneType' object has no attribute 'lr'</p>",
      "rawMarkdown": "Congratulations on this amazing win : ) Thank you for sharing the code. Our team is trying to implement your solution, we were using 1 gpu Nvidia V100 , at training we ResourceExhaustedError so we added additional 7 gpu but then got the following error for the Adam optimizer for LR. We trace this to the create model and we suspect that model is not compiling correctly with multiple gpus. Any insights on either of these issues and how many gpus did you use would be highly appreciated. Thank you again :) \n---------------------------------------------------------------------------\nAttributeError                            Traceback (most recent call last)",
      "votes": null
    },
    {
      "id": "387141",
      "postDate": "09/14/2018 12:24:24",
      "content": "<p>I used single GPU for training. You can try multi GPU by uncommenting lines:</p>\n\n<pre><code># '--multi-gpu', '2',\n# '--multi-gpu-force',\n</code></pre>\n\n<p>You can remove these lines if they cause some trouble:</p>\n\n<pre><code>    print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n    # K.set_value(model.optimizer.lr, 1e-5)\n    print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n</code></pre>",
      "rawMarkdown": "I used single GPU for training. You can try multi GPU by uncommenting lines:\n\n    # '--multi-gpu', '2',\n    # '--multi-gpu-force',\n\nYou can remove these lines if they cause some trouble:\n\n        print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n        # K.set_value(model.optimizer.lr, 1e-5)\n        print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))",
      "votes": null
    },
    {
      "id": "387832",
      "postDate": "09/15/2018 18:24:19",
      "content": "<p>Thank you this is very helpful. We are interested in visual relationship solution as well. Would it be possible to share the code? Thank you </p>",
      "rawMarkdown": "Thank you this is very helpful. We are interested in visual relationship solution as well. Would it be possible to share the code? Thank you",
      "votes": null
    },
    {
      "id": "553116",
      "postDate": "06/15/2019 06:08:16",
      "content": "<p>Gracias.</p>",
      "rawMarkdown": "Gracias.",
      "votes": null
    },
    {
      "id": "590283",
      "postDate": "08/02/2019 02:26:05",
      "content": "<p>Nice explanation. very helpful. Thank you very much. </p>",
      "rawMarkdown": "Nice explanation. very helpful. Thank you very much.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 379356,
      "author_name": "lvaleriu",
      "author_url": "",
      "post_date": "08/31/2018 08:39:51",
      "content": "<p>Congratulations! And nice to hear you've used keras-retinanet. Some others tried but reached nice, but lower scores than yours. As a contributor to this project, can i ask you (if it is possible, useful) to push the changes you've done to OID generator or others into a PR on github? Thanks and again, congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 379390,
          "author_name": "andy2709",
          "author_url": "",
          "post_date": "08/31/2018 09:55:28",
          "content": "<p>That would be great since we also happened to use keras-retinanet for this competition.\nThanks again for sharing with us your solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 379476,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "08/31/2018 12:14:54",
          "content": "<p>Thanks. ) I'm currently preparing code with my solution for github. I'm not sure if I will be able to directly make PR to keras-retinanet since there are some dependencies on my code and directory structure. But code for updated generator will be available in my repository.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 380699,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "09/03/2018 08:57:22",
          "content": "<p>Please add here a note when you upload your code. It will be very interesting to read the detailed solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 380718,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "09/03/2018 10:11:17",
          "content": "<p>Code already available:\n<a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018</a></p>\n\n<p>There are inference and training examples.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 380751,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "09/03/2018 11:44:55",
          "content": "<p>Thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 379403,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "08/31/2018 10:10:10",
      "content": "<p>Interesting and educational... We tried using keras-retinanet but it stopped learning at about 0.50 mAP. Was res50 though.\nA question if you don't mind. How long did it take per step and per inference? I got really slow times (800ms) on a formidable machine. Might have been configuration problem. </p>",
      "votes": null,
      "replies": [
        {
          "id": 379470,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "08/31/2018 12:09:01",
          "content": "<p>I used NVIDIA GTX 1080 Ti</p>\n\n<ul>\n<li>Training with ResNet101 (img size 768-1024) one epoch 10000 images:  ~6400 sec </li>\n<li>Training with ResNet152 (img size 600-800) one epoch 10000 images: ~5200 sec</li>\n</ul>\n\n<p>Inference can't say exactly but something like 500 ms per image I guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 379731,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "08/31/2018 20:49:33",
          "content": "<p>Mmmm ok, so there must have been something with my setup. Thank you for your quick reply. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 380684,
      "author_name": "muhammedazamkhan",
      "author_url": "",
      "post_date": "09/03/2018 08:13:42",
      "content": "<p>Hi @ZFTurbo,\nCongratulations. Thanks a lot for sharing your nice approach. I tried pre-trained model 'faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28' using object_detection_tutorial.ipynb, however, ended with a low score one. I just wondering if you can share your modified version of 'object_detection_tutorial.ipynb'. I really appreciate your contribution to the community.</p>",
      "votes": null,
      "replies": [
        {
          "id": 380949,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "09/03/2018 19:16:23",
          "content": "<p>I put the TensorFlow code in the same repo. You can find details here:</p>\n\n<p><a href=\"https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018/tree/master/tf_solution\">https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018/tree/master/tf_solution</a></p>\n\n<p>Some contestants tell me that there is possbility to lower threshold for this pretrained model (graph reexport?) and it will start to generate more boxes with lower confidence score. This leads to much higher mAP. One contestant reports 0.37872 public LB, 0.34736 private LB without any changes in model. But I didn't check it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 380698,
      "author_name": "gpreda",
      "author_url": "",
      "post_date": "09/03/2018 08:56:02",
      "content": "<p>Congratulations! And also, very comprehensive explanation, well structured, important details highlighted and useful tips.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 381130,
      "author_name": "thomas647",
      "author_url": "",
      "post_date": "09/04/2018 05:48:11",
      "content": "<p>Congratulations! Thanks for this explanation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 387100,
      "author_name": "aeweida",
      "author_url": "",
      "post_date": "09/14/2018 10:02:31",
      "content": "<h2>Congratulations on this amazing win : ) Thank you for sharing the code. Our team is trying to implement your solution, we were using 1 gpu Nvidia V100 , at training we ResourceExhaustedError so we added additional 7 gpu but then got the following error for the Adam optimizer for LR. We trace this to the create model and we suspect that model is not compiling correctly with multiple gpus. Any insights on either of these issues and how many gpus did you use would be highly appreciated. Thank you again :) </h2>\n\n<p>AttributeError                            Traceback (most recent call last)\n in ()\n     18         DATASET_PATH,\n     19     ]\n---&gt; 20     main(params)</p>\n\n<p> in main(args)\n     52     print (device_lib.list_local_devices())\n     53 \n---&gt; 54     print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n     55     # K.set_value(model.optimizer.lr, 1e-5)\n     56     print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))</p>\n\n<p>AttributeError: 'NoneType' object has no attribute 'lr'</p>",
      "votes": null,
      "replies": [
        {
          "id": 387141,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "09/14/2018 12:24:24",
          "content": "<p>I used single GPU for training. You can try multi GPU by uncommenting lines:</p>\n\n<pre><code># '--multi-gpu', '2',\n# '--multi-gpu-force',\n</code></pre>\n\n<p>You can remove these lines if they cause some trouble:</p>\n\n<pre><code>    print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n    # K.set_value(model.optimizer.lr, 1e-5)\n    print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 387832,
          "author_name": "aeweida",
          "author_url": "",
          "post_date": "09/15/2018 18:24:19",
          "content": "<p>Thank you this is very helpful. We are interested in visual relationship solution as well. Would it be possible to share the code? Thank you </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 553116,
      "author_name": "jmorenog",
      "author_url": "",
      "post_date": "06/15/2019 06:08:16",
      "content": "<p>Gracias.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 590283,
      "author_name": "sangyul3",
      "author_url": "",
      "post_date": "08/02/2019 02:26:05",
      "content": "<p>Nice explanation. very helpful. Thank you very much. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "379163": "Software\n========\n\nWindows 10 + Python 3.5 + Keras 2.2 + Keras-RetinaNet [0.4.1]: https://github.com/fizyr/keras-retinanet\n\nMain approach\n=============\n\nAll classes were split on 5 levels depends on children level. As the basis I used official classes hierarchy: https://storage.googleapis.com/openimages/challenge_2018/bbox_labels_500_hierarchy_visualizer/circle.html\n\n - **Level 1**: 443 classes - no children, have maximum impact on score\n - **Level 2**: 46 classes - have children only with level 1\n - **Level 3**: 4 classes: 'Seafood', 'Watercraft', 'Insect', 'Carnivore'\n - **Level 4**: 4 classes: 'Vegetable', 'Land vehicle', 'Reptile', 'Invertebrate'\n - **Level 5**: 3 classes: 'Furniture', 'Vehicle', 'Animal'\n\n**Note 1**: It’s possible to split classes on 3rd, 4th and 5th levels differently. Only “Animal” class has real 5th level.\n\n**Note 2**: 3rd, 4th and 5th don’t have large impact on final score so I didn’t really tune these models.\n\n**Note 3**: Models of 2-5 levels optional, since we can generate predictions for them using Level 1 model. We just need to duplicate boxes for their children and use some NMS algo on them since there will appear some duplicates. Score for this approach will be slightly lower than using separate 2-5 level models (as I remember correctly ~0.01-0.02 lower).\n\nWhy we need the split? Why we don’t train using all 500 classes?\n================================================================\n\nIn process of dataset analysis I found out that there are some images which marked up only on higher level classes. The most explicit representatives are the Person class (2nd level) and Man, Woman, Boy, Girl (1st level). What's the problem? Images marked with the Person class do not have markup for Man, Woman, Boy and Girl. If we use these images for training at once for all classes - the model will be confused in the classes Man, Woman, Boy and Girl. You can throw out images for Person, but then it makes no sense to train this class as part of the overall model, but it's easier to generate markup for Person in the inference step (using boxes for Man, Woman, Boy and Girl).\n\nTo train 2-5 levels models, we can use more data, including images marked for example by Person besides images including subclasses (Man, Woman, Boy and Girl) and this improves the result.\n\nPreparing data for training\n============================\n\n**Level 1**: All images containing markup for one of the 443 classes were added + images were added that contained the markup of the first-level classes not included in Challenge 500, as negative samples. Excluded all images containing markup of the higher levels in order to reduce the number of potential False negative (that is, the markup should be, but it is not in the training set).\n\n**Levels 2 - 5**: Images for each class contain both markups directly for this class, and for all children classes. Excluded images containing the higher level classes. Added images containing disjoint classes as negative samples.\n\nValidation is rather slow on RetinaNet. I did a separate validation for the training process, where for each class there were about 25 images containing this class. Thus, I balanced the classes a bit and reduced the time of the validation. My validation during training was about 6.5K images instead of 41K.\n\nBest models\n===========\n\n    1) Keras RetinaNet + ResNet152 Image size ranges 600-800 px\n    mAP on small validation: 0.5028\n    mAP on full validation: 0.384009\n    LB score*: 0.47441\n    \n    2) Keras RetinaNet + ResNet101 Image size ranges 768-1024 px\n    mAP on small validation: 0.4896\n    mAP on full validation: 0.377631\n    LB score*: 0.47549\n\n* - I have not tried individual submissions by levels, just sent a merged result for the best models for all levels.\n\n![enter image description here][1]\n\n![enter image description here][2]\n\n![enter image description here][3]\n\n![enter image description here][4]\n\nModels continue improving when I stopped training them, so I think there is some room to increase score.\n\nProcess of training and inference\n=================================\n\nKeras-RetinaNet already has a implemented generator for training on Open Images Dataset (OID). But in the current form it is not very suitable. I made several changes to it:\n\n - I added support for “empty” class - for images with negative samples that do not contain any boxes.\n - In the current generator there are no augmentations associated with the color, I added a random change in the intensity of the channels - this gave a good increase in validation score. And it feels like augmentation set in Keras-Retinanet is not enough for training a strong model.\n - Most important, selecting just random images for the batch will work poor at random. I replaced it with the following method:\n- before the start of training for each class (including \"empty\") we create a list of images that contains boxes for the class.\n- during the training, to add the next image to the batch, we first randomly select a class, then randomly select the image from the image list for the class. Thus, we achieve more or less uniform training by classes. Strictly speaking, the distribution is still not quite uniform because the images usually contain boxes for several classes at once.\n - From small things: I increased values ​​for augmentations transform-generator, especially scale. I changed the value of factor from 0.1 to 0.9 in ReduceLROnPlateau because the learning rate dropped too fast.\n - At the stage of the convert model for inference, in the FilterDetections layer in RetinaNet - reduced the score_threshold from 0.05 to 0.01 and tried to increase the number of boxes from 300 to 500. This gives more flexibility in the ensemble stage. Also note that the default model RetinaNet for Inference already contains NMS in the FilterDetections layer with nms_threshold = 0.5 and there is a feeling that you can play with this parameter.\n - I restarted training several times with a large LR = 1e-5, when LR fell too low. And each time the model became better. But the experiment is not very clean, because every time I changed the augmentation parameters.\n\nEnsembles\n=========\n\n - For each model, I made a prediction on the image and its horizontal mirror.\n - RetinaNet has internal layer with made NMS on single image prediction (it can be switched off to get the full set of boxes, but I didn’t try it).\n - I merged boxes using two methods:\n- standard NMS (https://github.com/rbgirshick/fast-rcnn/blob/master/lib/utils/nms.py), the optimal threshold I found was about 0.75. \n- second method is my own heuristic. It works better on LB. Heuristics included a weighted addition of boxes, which changed their coordinates and Confidence score.\n - I also tried Soft NMS [https://github.com/bharatsingh430/soft-nms] last 3 days of competition, but it works almost the same as default NMS and worse than my ensemble approach. Probably I missed something.\n\nMy ensemble approach in short\n=============================\n\n 1. On input we get set of boxes from different N models\n 2. Set “Init_weight” = 1/N and initialize “result” set of boxes as empty list\n 3. Add all boxes for best model with “init_weight” weight to “result” list. For all other boxes of other models try to find best matching box in result using IOU. If box with IOU &gt; THR (0.55) exists in “result” then merge best found box in result with it using weighted average for coordinates and confidence score. Increase weight of this box by Init_weight value. Otherwise add this new box to “result” with “init_weight” weight.\n\nOther solutions\n===============\n\nAt the first stage I tried to solve the problem simply on pretrain models without any retraining:\n\n - RetinanNet Pretrain Coco: https://github.com/fizyr/keras-retinanet/releases - gives an approximately 0.13 on LB, if the classes between OID and COCO are correctly matched.\n - On the pretrain from here: https://github.com/tensorflow/models/blob/master/research/object_detection/object_detection_tutorial.ipynb using this model: faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28 you can get ~0.25 on LB.\n\nObservations and small tricks\n=============================\n\n 1. Because of the metric nature, it is better to output the maximum number of rectangles even with low confidence score.\n 2. Validation works so-so. In most cases, it is worse than LB. Most likely this is due to two factors: the distribution in kaggle test is very different from the distribution for validation. Classes with a small number of elements affect the score as much as others.\n 3. The limitation on the size of the CSV file on Kaggle in this task is ~2GB. This fact didn't allow to submit more boxes with low confidence score.\n\nProposed dataset improvement\n============================\n\nIn case we will have similar competitions next years:\n\n - To simplify the task organizers should improve the training dataset, so that there are no situations when there is a markup for level 2 and there is no markup for level 1. \n - It would be good to have a full markup for each image with all the boxes including the parent classes. To avoid discrepancies / errors, etc.\n - It’s better to use actual class names like \"Ambulance\" instead of /m/012n7d\n - A little confusing is the presence of the flags like “isGroupOf” - are there any images of such type in the Test set or not? I eventually excluded these images from training and still not sure if it was right thing to do.\n - It seems to me that the current metric is not very good due to the fact that classes with a very small number of boxes influence just like classes with millions of boxes. Small mistakes can lead to significant score changes.\n\nCode and PreTrained models\n==========================\n\nGitHub repository: https://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10220/ResNet101-loss.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10218/ResNet101-valid-mAP.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10219/ResNet152-loss.png\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/379163/10217/ResNet152-valid-mAP.png",
    "379356": "Congratulations! And nice to hear you've used keras-retinanet. Some others tried but reached nice, but lower scores than yours. As a contributor to this project, can i ask you (if it is possible, useful) to push the changes you've done to OID generator or others into a PR on github? Thanks and again, congratulations!",
    "379390": "That would be great since we also happened to use keras-retinanet for this competition.\nThanks again for sharing with us your solution.",
    "379403": "Interesting and educational... We tried using keras-retinanet but it stopped learning at about 0.50 mAP. Was res50 though.\nA question if you don't mind. How long did it take per step and per inference? I got really slow times (800ms) on a formidable machine. Might have been configuration problem.",
    "379470": "I used NVIDIA GTX 1080 Ti\n\n - Training with ResNet101 (img size 768-1024) one epoch 10000 images:  ~6400 sec \n - Training with ResNet152 (img size 600-800) one epoch 10000 images: ~5200 sec\n\nInference can't say exactly but something like 500 ms per image I guess.",
    "379476": "Thanks. ) I'm currently preparing code with my solution for github. I'm not sure if I will be able to directly make PR to keras-retinanet since there are some dependencies on my code and directory structure. But code for updated generator will be available in my repository.",
    "379731": "Mmmm ok, so there must have been something with my setup. Thank you for your quick reply.",
    "380684": "Hi @ZFTurbo,\nCongratulations. Thanks a lot for sharing your nice approach. I tried pre-trained model 'faster_rcnn_inception_resnet_v2_atrous_oid_2018_01_28' using object_detection_tutorial.ipynb, however, ended with a low score one. I just wondering if you can share your modified version of 'object_detection_tutorial.ipynb'. I really appreciate your contribution to the community.",
    "380698": "Congratulations! And also, very comprehensive explanation, well structured, important details highlighted and useful tips.",
    "380699": "Please add here a note when you upload your code. It will be very interesting to read the detailed solution.",
    "380718": "Code already available:\nhttps://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018\n\nThere are inference and training examples.",
    "380751": "Thank you very much.",
    "380949": "I put the TensorFlow code in the same repo. You can find details here:\n\nhttps://github.com/ZFTurbo/Keras-RetinaNet-for-Open-Images-Challenge-2018/tree/master/tf_solution\n\nSome contestants tell me that there is possbility to lower threshold for this pretrained model (graph reexport?) and it will start to generate more boxes with lower confidence score. This leads to much higher mAP. One contestant reports 0.37872 public LB, 0.34736 private LB without any changes in model. But I didn't check it.",
    "381130": "Congratulations! Thanks for this explanation!",
    "387100": "Congratulations on this amazing win : ) Thank you for sharing the code. Our team is trying to implement your solution, we were using 1 gpu Nvidia V100 , at training we ResourceExhaustedError so we added additional 7 gpu but then got the following error for the Adam optimizer for LR. We trace this to the create model and we suspect that model is not compiling correctly with multiple gpus. Any insights on either of these issues and how many gpus did you use would be highly appreciated. Thank you again :) \n---------------------------------------------------------------------------\nAttributeError                            Traceback (most recent call last)",
    "387141": "I used single GPU for training. You can try multi GPU by uncommenting lines:\n\n    # '--multi-gpu', '2',\n    # '--multi-gpu-force',\n\nYou can remove these lines if they cause some trouble:\n\n        print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))\n        # K.set_value(model.optimizer.lr, 1e-5)\n        print('Learning rate: {}'.format(K.get_value(model.optimizer.lr)))",
    "387832": "Thank you this is very helpful. We are interested in visual relationship solution as well. Would it be possible to share the code? Thank you",
    "553116": "Gracias.",
    "590283": "Nice explanation. very helpful. Thank you very much."
  },
  "source": "meta"
}