{
  "id": 111266,
  "title": "10th place solution",
  "url": "/competitions/open-images-2019-object-detection/discussion/111266",
  "author_name": "dingwoai",
  "post_date": "2019-10-04T11:14:07.516000",
  "votes": 18,
  "comment_count": 3,
  "views": 0,
  "content": "<p><strong>TL;DR</strong>: \n- I'm using mmdetection framework and it's really convenient;\n- Best single model (cascade rcnn with imagenet pretrained resnext101) + TTA (horizontal flip, multi-scale testing (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400,\n2100)) achieves 0.499 public;\n- Split datasets into 6 subsets by frequency and then finetune on them for 1-2 epochs. --&gt; appr.  0.05+ increase;\n- Parent class expansion gives appr. 0.01+ increase;\n- Weighted ensemble (from <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283\">ZFTurbo's solution last year</a>) of all my high score submissions --&gt; 0.608 final public score.</p>\n\n<p><strong>1. Single model</strong>\nAt the very beginning, I was going to attend the \"Visual Relationship Detection\" track. But then I realized that I didn't have a good object detection model for that one. So I started with faster rcnn+resnext101, it takes me about 20 days to train 24 epochs and results in 0.446 on public lb.\nSimilarly, I trained cascade rcnn+resnext101, cascade rcnn+senet154 for 12 and 8 epochs respectively.\nI just leave these models training for several weeks, do my daily work and give up \"visual relation detection\". \nThe best single model is cascade rcnn+resnext101, which was accidently trained for 19 epochs (6 epochs longer than planned). So maybe I should train longer for each model :).\n<strong>Conclusion</strong>: my single models are weak. They should be trained longer.</p>\n\n<p><strong>2. Finetune</strong>\nSince the classes are very unbalanced, I split the dataset classes into 6 subsets simply according to frequency and finetune on them using faster rcnn+resnext101 model:\n- Classes 0-50, appr. 1411368images, 2 epochs, lr 0.001\n- Classes 51-100, appr. 308352 images, 2 epochs, lr 0.001\n- Classes 101-200: appr. 208096 images, 2 epochs, lr 0.001\n- Classes 201-300, appr. 93140 images, 2epochs, lr 0.001\n- Classes 301-400, appr. 45840 images, 1epoch, lr 0.001\n- Classes 401-500, appr. 19316 images, 1epoch, lr 0.001\nLast two weeks before final deadline, I found one huge bug in my code.\nAfter solving this bug, ensembling the predictions of finetuned models gave appr. 0.05+ increase on public LB.\n<strong>Conclusion</strong>: solving class imbalance problem is the key to top silver or gold medal.</p>\n\n<p><strong>3. TTA and Final Ensemble</strong>\nSome teams merged and I may have the chance for solo gold.\nSo I did the followings:\n- TTA for each model: horizontal flip, multi-scale testing with (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400, 2100) image size;\n- Expand parent class for each prediction after inference.\n- Weighted ensemble (from <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283\">ZFTurbo's solution last year</a>) of all models.</p>\n\n<p><strong>Other tricks:</strong>\n- Increase the number of limited boxes for each image, even though they are of low confidence. I choose 600 as upper limit in my final submission.\n- I wasn't able to finish 12 epochs for my cascade rcnn+senet154 model. Nevertheless, ensembling it gives slight improvement.</p>\n\n<p><strong>Not work for me:</strong>\nSoft-NMS: tried to use it for ensemble, wasn't going well. </p>\n\n<p><strong>Planned but not implemented:</strong>\n- Use mask annotations from the segmentation track;\n- Multi-scale training and mixup augmentation.</p>\n\n<p><strong>Final words:</strong>\nLet's play fair. \nPeace&amp;love. 👍 </p>",
  "messages": [
    {
      "id": 640985,
      "postDate": "2019-10-04T11:14:07.517Z",
      "content": "<p><strong>TL;DR</strong>: \n- I'm using mmdetection framework and it's really convenient;\n- Best single model (cascade rcnn with imagenet pretrained resnext101) + TTA (horizontal flip, multi-scale testing (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400,\n2100)) achieves 0.499 public;\n- Split datasets into 6 subsets by frequency and then finetune on them for 1-2 epochs. --&gt; appr.  0.05+ increase;\n- Parent class expansion gives appr. 0.01+ increase;\n- Weighted ensemble (from <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283\">ZFTurbo's solution last year</a>) of all my high score submissions --&gt; 0.608 final public score.</p>\n\n<p><strong>1. Single model</strong>\nAt the very beginning, I was going to attend the \"Visual Relationship Detection\" track. But then I realized that I didn't have a good object detection model for that one. So I started with faster rcnn+resnext101, it takes me about 20 days to train 24 epochs and results in 0.446 on public lb.\nSimilarly, I trained cascade rcnn+resnext101, cascade rcnn+senet154 for 12 and 8 epochs respectively.\nI just leave these models training for several weeks, do my daily work and give up \"visual relation detection\". \nThe best single model is cascade rcnn+resnext101, which was accidently trained for 19 epochs (6 epochs longer than planned). So maybe I should train longer for each model :).\n<strong>Conclusion</strong>: my single models are weak. They should be trained longer.</p>\n\n<p><strong>2. Finetune</strong>\nSince the classes are very unbalanced, I split the dataset classes into 6 subsets simply according to frequency and finetune on them using faster rcnn+resnext101 model:\n- Classes 0-50, appr. 1411368images, 2 epochs, lr 0.001\n- Classes 51-100, appr. 308352 images, 2 epochs, lr 0.001\n- Classes 101-200: appr. 208096 images, 2 epochs, lr 0.001\n- Classes 201-300, appr. 93140 images, 2epochs, lr 0.001\n- Classes 301-400, appr. 45840 images, 1epoch, lr 0.001\n- Classes 401-500, appr. 19316 images, 1epoch, lr 0.001\nLast two weeks before final deadline, I found one huge bug in my code.\nAfter solving this bug, ensembling the predictions of finetuned models gave appr. 0.05+ increase on public LB.\n<strong>Conclusion</strong>: solving class imbalance problem is the key to top silver or gold medal.</p>\n\n<p><strong>3. TTA and Final Ensemble</strong>\nSome teams merged and I may have the chance for solo gold.\nSo I did the followings:\n- TTA for each model: horizontal flip, multi-scale testing with (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400, 2100) image size;\n- Expand parent class for each prediction after inference.\n- Weighted ensemble (from <a href=\"https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283\">ZFTurbo's solution last year</a>) of all models.</p>\n\n<p><strong>Other tricks:</strong>\n- Increase the number of limited boxes for each image, even though they are of low confidence. I choose 600 as upper limit in my final submission.\n- I wasn't able to finish 12 epochs for my cascade rcnn+senet154 model. Nevertheless, ensembling it gives slight improvement.</p>\n\n<p><strong>Not work for me:</strong>\nSoft-NMS: tried to use it for ensemble, wasn't going well. </p>\n\n<p><strong>Planned but not implemented:</strong>\n- Use mask annotations from the segmentation track;\n- Multi-scale training and mixup augmentation.</p>\n\n<p><strong>Final words:</strong>\nLet's play fair. \nPeace&amp;love. 👍 </p>",
      "rawMarkdown": "**TL;DR**: \n- I'm using mmdetection framework and it's really convenient;\n- Best single model (cascade rcnn with imagenet pretrained resnext101) + TTA (horizontal flip, multi-scale testing (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400,\n2100)) achieves 0.499 public;\n- Split datasets into 6 subsets by frequency and then finetune on them for 1-2 epochs. --&gt; appr.  0.05+ increase;\n- Parent class expansion gives appr. 0.01+ increase;\n- Weighted ensemble (from [ZFTurbo's solution last year](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283)) of all my high score submissions --&gt; 0.608 final public score.\n\n**1. Single model**\nAt the very beginning, I was going to attend the \"Visual Relationship Detection\" track. But then I realized that I didn't have a good object detection model for that one. So I started with faster rcnn+resnext101, it takes me about 20 days to train 24 epochs and results in 0.446 on public lb.\nSimilarly, I trained cascade rcnn+resnext101, cascade rcnn+senet154 for 12 and 8 epochs respectively.\nI just leave these models training for several weeks, do my daily work and give up \"visual relation detection\". \nThe best single model is cascade rcnn+resnext101, which was accidently trained for 19 epochs (6 epochs longer than planned). So maybe I should train longer for each model :).\n**Conclusion**: my single models are weak. They should be trained longer.\n\n**2. Finetune**\nSince the classes are very unbalanced, I split the dataset classes into 6 subsets simply according to frequency and finetune on them using faster rcnn+resnext101 model:\n- Classes 0-50, appr. 1411368images, 2 epochs, lr 0.001\n- Classes 51-100, appr. 308352 images, 2 epochs, lr 0.001\n- Classes 101-200: appr. 208096 images, 2 epochs, lr 0.001\n- Classes 201-300, appr. 93140 images, 2epochs, lr 0.001\n- Classes 301-400, appr. 45840 images, 1epoch, lr 0.001\n- Classes 401-500, appr. 19316 images, 1epoch, lr 0.001\nLast two weeks before final deadline, I found one huge bug in my code.\nAfter solving this bug, ensembling the predictions of finetuned models gave appr. 0.05+ increase on public LB.\n**Conclusion**: solving class imbalance problem is the key to top silver or gold medal.\n\n**3. TTA and Final Ensemble**\nSome teams merged and I may have the chance for solo gold.\nSo I did the followings:\n- TTA for each model: horizontal flip, multi-scale testing with (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400, 2100) image size;\n- Expand parent class for each prediction after inference.\n- Weighted ensemble (from [ZFTurbo's solution last year](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283)) of all models.\n\n**Other tricks:**\n- Increase the number of limited boxes for each image, even though they are of low confidence. I choose 600 as upper limit in my final submission.\n- I wasn't able to finish 12 epochs for my cascade rcnn+senet154 model. Nevertheless, ensembling it gives slight improvement.\n\n**Not work for me:**\nSoft-NMS: tried to use it for ensemble, wasn't going well. \n\n**Planned but not implemented:**\n- Use mask annotations from the segmentation track;\n- Multi-scale training and mixup augmentation.\n\n**Final words:**\nLet's play fair. \nPeace&amp;love. 👍 ",
      "votes": 16
    },
    {
      "id": 644368,
      "postDate": "2019-10-08T17:34:49.070Z",
      "content": "<p>Thanks for sharing your approach. What were the hardware specs you used and how much time did you train for? </p>",
      "rawMarkdown": "Thanks for sharing your approach. What were the hardware specs you used and how much time did you train for? "
    },
    {
      "id": 641896,
      "postDate": "2019-10-05T09:21:34.773Z",
      "content": "<p>Thanks and congrats! So where was the huge bug? :)</p>",
      "rawMarkdown": "Thanks and congrats! So where was the huge bug? :)"
    },
    {
      "id": 641027,
      "postDate": "2019-10-04T11:52:59.340Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 644368,
      "author_name": "Rohit Midha",
      "author_url": "",
      "post_date": "2019-10-08T17:34:49.070000",
      "content": "<p>Thanks for sharing your approach. What were the hardware specs you used and how much time did you train for? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 641896,
      "author_name": "Artyom Palvelev",
      "author_url": "",
      "post_date": "2019-10-05T09:21:34.773000",
      "content": "<p>Thanks and congrats! So where was the huge bug? :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 641027,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-04T11:52:59.340000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "640985": "**TL;DR**: \n- I'm using mmdetection framework and it's really convenient;\n- Best single model (cascade rcnn with imagenet pretrained resnext101) + TTA (horizontal flip, multi-scale testing (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400,\n2100)) achieves 0.499 public;\n- Split datasets into 6 subsets by frequency and then finetune on them for 1-2 epochs. --&gt; appr.  0.05+ increase;\n- Parent class expansion gives appr. 0.01+ increase;\n- Weighted ensemble (from [ZFTurbo's solution last year](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283)) of all my high score submissions --&gt; 0.608 final public score.\n\n**1. Single model**\nAt the very beginning, I was going to attend the \"Visual Relationship Detection\" track. But then I realized that I didn't have a good object detection model for that one. So I started with faster rcnn+resnext101, it takes me about 20 days to train 24 epochs and results in 0.446 on public lb.\nSimilarly, I trained cascade rcnn+resnext101, cascade rcnn+senet154 for 12 and 8 epochs respectively.\nI just leave these models training for several weeks, do my daily work and give up \"visual relation detection\". \nThe best single model is cascade rcnn+resnext101, which was accidently trained for 19 epochs (6 epochs longer than planned). So maybe I should train longer for each model :).\n**Conclusion**: my single models are weak. They should be trained longer.\n\n**2. Finetune**\nSince the classes are very unbalanced, I split the dataset classes into 6 subsets simply according to frequency and finetune on them using faster rcnn+resnext101 model:\n- Classes 0-50, appr. 1411368images, 2 epochs, lr 0.001\n- Classes 51-100, appr. 308352 images, 2 epochs, lr 0.001\n- Classes 101-200: appr. 208096 images, 2 epochs, lr 0.001\n- Classes 201-300, appr. 93140 images, 2epochs, lr 0.001\n- Classes 301-400, appr. 45840 images, 1epoch, lr 0.001\n- Classes 401-500, appr. 19316 images, 1epoch, lr 0.001\nLast two weeks before final deadline, I found one huge bug in my code.\nAfter solving this bug, ensembling the predictions of finetuned models gave appr. 0.05+ increase on public LB.\n**Conclusion**: solving class imbalance problem is the key to top silver or gold medal.\n\n**3. TTA and Final Ensemble**\nSome teams merged and I may have the chance for solo gold.\nSo I did the followings:\n- TTA for each model: horizontal flip, multi-scale testing with (600, 900), (800, 1200), (1000, 1500), (1200, 1800), (1400, 2100) image size;\n- Expand parent class for each prediction after inference.\n- Weighted ensemble (from [ZFTurbo's solution last year](https://www.kaggle.com/c/google-ai-open-images-object-detection-track/discussion/64633#latest-590283)) of all models.\n\n**Other tricks:**\n- Increase the number of limited boxes for each image, even though they are of low confidence. I choose 600 as upper limit in my final submission.\n- I wasn't able to finish 12 epochs for my cascade rcnn+senet154 model. Nevertheless, ensembling it gives slight improvement.\n\n**Not work for me:**\nSoft-NMS: tried to use it for ensemble, wasn't going well. \n\n**Planned but not implemented:**\n- Use mask annotations from the segmentation track;\n- Multi-scale training and mixup augmentation.\n\n**Final words:**\nLet's play fair. \nPeace&amp;love. 👍 ",
    "644368": "Thanks for sharing your approach. What were the hardware specs you used and how much time did you train for? ",
    "641896": "Thanks and congrats! So where was the huge bug? :)",
    "641027": ""
  }
}