{
  "id": 110983,
  "title": "7th place solution",
  "url": "/competitions/open-images-2019-instance-segmentation/writeups/s-p-7th-place-solution",
  "author_name": "",
  "post_date": "2019-10-07T17:30:56.453Z",
  "votes": 34,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Congratulations to all the winners. Here is our solution:</p>\n\n<h2>Codebase and hardware</h2>\n\n<p>We used <a href=\"https://github.com/open-mmlab/mmdetection\">mmdetection</a> with 4 Tesla T4 during training, 8 Tesla T4 during TTA inference.</p>\n\n<h2>Model setup</h2>\n\n<p>There are 300 classes, in 3 hierarchical levels: \n-  275 leaf classes\n-  23 parent classes \n-  2 grandparent classes (<code>Carnivore</code> and <code>Reptile</code>)</p>\n\n<p>For the 275 leaf classes, we train models with these 275 classes as labels and use inference results directly.</p>\n\n<p>For the 25 parent and grandparent classes, there are two methods to get the predictions:\n- (i). Use the prediction results from the leaf model, and “expand” to the parent and grandparent. For example, if the leaf model predicted a <code>Tortoise</code> mask, we add a <code>Turtle</code> prediction (its parent) and a <code>Reptile</code> prediction (its grandparent) with the same mask and score. <br>\n- (ii). We train models with only the 23 parent classes as labels, and expand to the 2 grandparent classes. Note that for these models, we need to create hierarchical expansion of the instance segmentation before training (as explained <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track\">here</a>).</p>\n\n<h2>Validation set</h2>\n\n<p>We sampled a small subset from the official validation set for faster inference, 2844 images for the leaf model and 3841 images for the parent model. We used the validation scores to decide when to adjust learning rate, and to choose best checkpoints. </p>\n\n<p>The validation scores are significantly higher than LB. For the 275 leaf classes, we have validation score over 0.8. But nevertheless, the delta between validation and LB are stable.</p>\n\n<h2>Training set rebalancing</h2>\n\n<p>The open images dataset is extremely imbalanced between the 300 classes. We rebalanced the training data for the leaf model as follows:\nSort the 275 classes by descending number of images. So class #0 is the largest class.</p>\n\n<p>| Group | Class      | original num of imgs per class | rebalancing                   |\n|-----|------------|--------------------------------|-------------------------------|\n| 1     | #241 to #274 | 150-13                         | oversample x10                |\n| 2     | #64 to #240  | 1500-150                       | oversample to 1500 imgs/class |\n| 3     | #24 to #63   | 6k-1500                        | no rebalancing                |\n| 4     | #0 to #23    | 89k-6k                         | downsample to 6k*            |</p>\n\n<p>Each epoch has about 450k images after rebalancing. </p>\n\n<p>*For the downsample, we use different random seeds for each epoch to feed the model as many images as possible. Also, this 6k number is a rough target. In actual sampling, the top few classes ended up having more images because many images have multiple labels, and top classes already have more than 6k images after Group 1-3 sampling are done.</p>\n\n<p>Similarly for the parent model, sort the 23 classes</p>\n\n<p>| Group | Class    | rebalancing                  |\n|-------|----------|------------------------------|\n| 1     | #10 to #22 | upsample to 10k imgs/class   |\n| 2     | #5 to #9   | no rebalancing               |\n| 3     | #0 to #4   | downsample to 30k imgs/class |</p>\n\n<p>Each epoch has about 330k images after rebalancing.</p>\n\n<h2>Single models</h2>\n\n<p>We trained 3 cascade mrcnn leaf models\n- L1: backbone x101\n- L2: backbone r101 + deformable module\n- L3: backbone x101 + deformable module</p>\n\n<p>and 1 cascade mrcnn parent model\n- P1: backbone x101</p>\n\n<p>hyper-paramters:\n<code>\nimgs_per_gpu=1\nnum_gpu=4\nimg_scale=[(1333, 640), (1333, 960)]\nmultiscale_mode='range'\n</code></p>\n\n<p>lr schedule:\nmodel L1: 0.005 for 760k iterations; 0.005/3 for 60k iterations; 0.005/15 for 152k iterations\nmodel L2: 0.005 for 675k iterations; 0.005/3 for 64k iterations; 0.005/15 for 48k iterations\nmodel L3: 0.005 for 226k iterations; 0.005/5 for 132k iterations; 0.005/50 for 26k iterations\nmodel P1: 0.005 for 148k iterations; 0.005/5 for 62k iterations; 0.005/50 for 8k iterations</p>\n\n<p>Training took 0.5-0.65 hour per 1000 iterations with 4 T4. So in total the training the 4 models took about 22, 18, 10 and 5 days respectively.</p>\n\n<p>Public/private scores of each (<code>max_per_img=120, thr=0</code>, without TTA):</p>\n\n<p>| model | public score | private score |                |\n|-------|--------------|---------------|----------------|\n| L1    | 0.4708       | 0.4354        | on 275 classes |\n| L2    | 0.4705       | 0.4275        | on 275 classes |\n| L3    | 0.4712       | 0.4324        | on 275 classes |\n| P1    | 0.0349       | 0.0344        | on 25 classes  |</p>\n\n<h2>TTA</h2>\n\n<p>We used the TTA implementation from <a href=\"https://github.com/amirassov/kaggle-imaterialist\">Miras Amir's winning solution</a> of iMaterialist (Fashion) 2019 </p>\n\n<p>For the leaf model, it’s the ensemble of 3 single models, 2 scale (1333,800) and (1600,960), and flip. At RPN and BB stages, it’s the NMS ensemble of 12 single models. In the end, it’s the mean of all 12 masks for each instance. </p>\n\n<p>TTA inference of the leaf model is very slow, which took about 525 T4-hours. We split the 99999 images into 25 chunks and ran it in parallel. </p>\n\n<p>For the parent model, it’s the ensemble of 2 scale and flip, but just one single model.</p>\n\n<p>Lastly, for the 25 parent and grandparent classes, we implemented NMS ensemble at mask level (i.e. calculating the mask iou instead of bbox iou when determining which ones to suppress) to ensemble (i) leaf class predications expanded to 25 classes, and (ii) parent class predications expand to 25 classes.</p>\n\n<p>Public/private scores of each (<code>max_per_img=120, thr=0</code>):</p>\n\n<p>| model                    | public score | private score |                |\n|--------------------------|--------------|---------------|----------------|\n| leaf model ensemble      | 0.5007       | 0.4548        | on 275 classes |\n| leaf ensemble expanded   | 0.0325       | 0.0330        | on 25 classes  |\n| ensemble of (i) and (ii) | 0.0370       | 0.0370        | on 25 classes  |</p>\n\n<p>At this point, the total score is public 0.5007+0.0370 = 0.5378, private 0.4548+0.0370 = 0.4918. </p>\n\n<p>There’s one last thing we did to boost scores to 0.5383/0.4922: set <code>max_per_img=200</code> for the leaf ensemble. But we only had time (also restricted by sub file size) to do this for 12/25 of the test images.</p>\n\n<h2>Code</h2>\n\n<p>Training, inference, pre- and post-process code are available at <a href=\"https://github.com/boliu61/open-images-2019-instance-segmentation\">https://github.com/boliu61/open-images-2019-instance-segmentation</a>\nTrained model weights are also linked in the readme there</p>",
  "messages": [
    {
      "id": "639047",
      "postDate": "10/02/2019 18:00:42",
      "content": "<p>Congratulations to all the winners. Here is our solution:</p>\n\n<h2>Codebase and hardware</h2>\n\n<p>We used <a href=\"https://github.com/open-mmlab/mmdetection\">mmdetection</a> with 4 Tesla T4 during training, 8 Tesla T4 during TTA inference.</p>\n\n<h2>Model setup</h2>\n\n<p>There are 300 classes, in 3 hierarchical levels: \n-  275 leaf classes\n-  23 parent classes \n-  2 grandparent classes (<code>Carnivore</code> and <code>Reptile</code>)</p>\n\n<p>For the 275 leaf classes, we train models with these 275 classes as labels and use inference results directly.</p>\n\n<p>For the 25 parent and grandparent classes, there are two methods to get the predictions:\n- (i). Use the prediction results from the leaf model, and “expand” to the parent and grandparent. For example, if the leaf model predicted a <code>Tortoise</code> mask, we add a <code>Turtle</code> prediction (its parent) and a <code>Reptile</code> prediction (its grandparent) with the same mask and score. <br>\n- (ii). We train models with only the 23 parent classes as labels, and expand to the 2 grandparent classes. Note that for these models, we need to create hierarchical expansion of the instance segmentation before training (as explained <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track\">here</a>).</p>\n\n<h2>Validation set</h2>\n\n<p>We sampled a small subset from the official validation set for faster inference, 2844 images for the leaf model and 3841 images for the parent model. We used the validation scores to decide when to adjust learning rate, and to choose best checkpoints. </p>\n\n<p>The validation scores are significantly higher than LB. For the 275 leaf classes, we have validation score over 0.8. But nevertheless, the delta between validation and LB are stable.</p>\n\n<h2>Training set rebalancing</h2>\n\n<p>The open images dataset is extremely imbalanced between the 300 classes. We rebalanced the training data for the leaf model as follows:\nSort the 275 classes by descending number of images. So class #0 is the largest class.</p>\n\n<p>| Group | Class      | original num of imgs per class | rebalancing                   |\n|-----|------------|--------------------------------|-------------------------------|\n| 1     | #241 to #274 | 150-13                         | oversample x10                |\n| 2     | #64 to #240  | 1500-150                       | oversample to 1500 imgs/class |\n| 3     | #24 to #63   | 6k-1500                        | no rebalancing                |\n| 4     | #0 to #23    | 89k-6k                         | downsample to 6k*            |</p>\n\n<p>Each epoch has about 450k images after rebalancing. </p>\n\n<p>*For the downsample, we use different random seeds for each epoch to feed the model as many images as possible. Also, this 6k number is a rough target. In actual sampling, the top few classes ended up having more images because many images have multiple labels, and top classes already have more than 6k images after Group 1-3 sampling are done.</p>\n\n<p>Similarly for the parent model, sort the 23 classes</p>\n\n<p>| Group | Class    | rebalancing                  |\n|-------|----------|------------------------------|\n| 1     | #10 to #22 | upsample to 10k imgs/class   |\n| 2     | #5 to #9   | no rebalancing               |\n| 3     | #0 to #4   | downsample to 30k imgs/class |</p>\n\n<p>Each epoch has about 330k images after rebalancing.</p>\n\n<h2>Single models</h2>\n\n<p>We trained 3 cascade mrcnn leaf models\n- L1: backbone x101\n- L2: backbone r101 + deformable module\n- L3: backbone x101 + deformable module</p>\n\n<p>and 1 cascade mrcnn parent model\n- P1: backbone x101</p>\n\n<p>hyper-paramters:\n<code>\nimgs_per_gpu=1\nnum_gpu=4\nimg_scale=[(1333, 640), (1333, 960)]\nmultiscale_mode='range'\n</code></p>\n\n<p>lr schedule:\nmodel L1: 0.005 for 760k iterations; 0.005/3 for 60k iterations; 0.005/15 for 152k iterations\nmodel L2: 0.005 for 675k iterations; 0.005/3 for 64k iterations; 0.005/15 for 48k iterations\nmodel L3: 0.005 for 226k iterations; 0.005/5 for 132k iterations; 0.005/50 for 26k iterations\nmodel P1: 0.005 for 148k iterations; 0.005/5 for 62k iterations; 0.005/50 for 8k iterations</p>\n\n<p>Training took 0.5-0.65 hour per 1000 iterations with 4 T4. So in total the training the 4 models took about 22, 18, 10 and 5 days respectively.</p>\n\n<p>Public/private scores of each (<code>max_per_img=120, thr=0</code>, without TTA):</p>\n\n<p>| model | public score | private score |                |\n|-------|--------------|---------------|----------------|\n| L1    | 0.4708       | 0.4354        | on 275 classes |\n| L2    | 0.4705       | 0.4275        | on 275 classes |\n| L3    | 0.4712       | 0.4324        | on 275 classes |\n| P1    | 0.0349       | 0.0344        | on 25 classes  |</p>\n\n<h2>TTA</h2>\n\n<p>We used the TTA implementation from <a href=\"https://github.com/amirassov/kaggle-imaterialist\">Miras Amir's winning solution</a> of iMaterialist (Fashion) 2019 </p>\n\n<p>For the leaf model, it’s the ensemble of 3 single models, 2 scale (1333,800) and (1600,960), and flip. At RPN and BB stages, it’s the NMS ensemble of 12 single models. In the end, it’s the mean of all 12 masks for each instance. </p>\n\n<p>TTA inference of the leaf model is very slow, which took about 525 T4-hours. We split the 99999 images into 25 chunks and ran it in parallel. </p>\n\n<p>For the parent model, it’s the ensemble of 2 scale and flip, but just one single model.</p>\n\n<p>Lastly, for the 25 parent and grandparent classes, we implemented NMS ensemble at mask level (i.e. calculating the mask iou instead of bbox iou when determining which ones to suppress) to ensemble (i) leaf class predications expanded to 25 classes, and (ii) parent class predications expand to 25 classes.</p>\n\n<p>Public/private scores of each (<code>max_per_img=120, thr=0</code>):</p>\n\n<p>| model                    | public score | private score |                |\n|--------------------------|--------------|---------------|----------------|\n| leaf model ensemble      | 0.5007       | 0.4548        | on 275 classes |\n| leaf ensemble expanded   | 0.0325       | 0.0330        | on 25 classes  |\n| ensemble of (i) and (ii) | 0.0370       | 0.0370        | on 25 classes  |</p>\n\n<p>At this point, the total score is public 0.5007+0.0370 = 0.5378, private 0.4548+0.0370 = 0.4918. </p>\n\n<p>There’s one last thing we did to boost scores to 0.5383/0.4922: set <code>max_per_img=200</code> for the leaf ensemble. But we only had time (also restricted by sub file size) to do this for 12/25 of the test images.</p>\n\n<h2>Code</h2>\n\n<p>Training, inference, pre- and post-process code are available at <a href=\"https://github.com/boliu61/open-images-2019-instance-segmentation\">https://github.com/boliu61/open-images-2019-instance-segmentation</a>\nTrained model weights are also linked in the readme there</p>",
      "rawMarkdown": "Congratulations to all the winners. Here is our solution:\n\n## Codebase and hardware\n\nWe used [mmdetection](https://github.com/open-mmlab/mmdetection) with 4 Tesla T4 during training, 8 Tesla T4 during TTA inference.\n\n## Model setup\n\nThere are 300 classes, in 3 hierarchical levels: \n-  275 leaf classes\n-  23 parent classes \n-  2 grandparent classes (`Carnivore` and `Reptile`)\n\nFor the 275 leaf classes, we train models with these 275 classes as labels and use inference results directly.\n\nFor the 25 parent and grandparent classes, there are two methods to get the predictions:\n- (i). Use the prediction results from the leaf model, and “expand” to the parent and grandparent. For example, if the leaf model predicted a `Tortoise` mask, we add a `Turtle` prediction (its parent) and a `Reptile` prediction (its grandparent) with the same mask and score.  \n- (ii). We train models with only the 23 parent classes as labels, and expand to the 2 grandparent classes. Note that for these models, we need to create hierarchical expansion of the instance segmentation before training (as explained [here](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track)).\n\n\n## Validation set\n\nWe sampled a small subset from the official validation set for faster inference, 2844 images for the leaf model and 3841 images for the parent model. We used the validation scores to decide when to adjust learning rate, and to choose best checkpoints. \n\nThe validation scores are significantly higher than LB. For the 275 leaf classes, we have validation score over 0.8. But nevertheless, the delta between validation and LB are stable.\n\n\n## Training set rebalancing\n\nThe open images dataset is extremely imbalanced between the 300 classes. We rebalanced the training data for the leaf model as follows:\nSort the 275 classes by descending number of images. So class #0 is the largest class.\n\n| Group | Class      | original num of imgs per class | rebalancing                   |\n|-----|------------|--------------------------------|-------------------------------|\n| 1     | #241 to #274 | 150-13                         | oversample x10                |\n| 2     | #64 to #240  | 1500-150                       | oversample to 1500 imgs/class |\n| 3     | #24 to #63   | 6k-1500                        | no rebalancing                |\n| 4     | #0 to #23    | 89k-6k                         | downsample to 6k*            |\n\nEach epoch has about 450k images after rebalancing. \n\n*For the downsample, we use different random seeds for each epoch to feed the model as many images as possible. Also, this 6k number is a rough target. In actual sampling, the top few classes ended up having more images because many images have multiple labels, and top classes already have more than 6k images after Group 1-3 sampling are done.\n\nSimilarly for the parent model, sort the 23 classes\n\n| Group | Class    | rebalancing                  |\n|-------|----------|------------------------------|\n| 1     | #10 to #22 | upsample to 10k imgs/class   |\n| 2     | #5 to #9   | no rebalancing               |\n| 3     | #0 to #4   | downsample to 30k imgs/class |\n\nEach epoch has about 330k images after rebalancing.\n\n## Single models\n\nWe trained 3 cascade mrcnn leaf models\n- L1: backbone x101\n- L2: backbone r101 + deformable module\n- L3: backbone x101 + deformable module\n\nand 1 cascade mrcnn parent model\n- P1: backbone x101\n\nhyper-paramters:\n```\nimgs_per_gpu=1\nnum_gpu=4\nimg_scale=[(1333, 640), (1333, 960)]\nmultiscale_mode='range'\n```\n\nlr schedule:\nmodel L1: 0.005 for 760k iterations; 0.005/3 for 60k iterations; 0.005/15 for 152k iterations\nmodel L2: 0.005 for 675k iterations; 0.005/3 for 64k iterations; 0.005/15 for 48k iterations\nmodel L3: 0.005 for 226k iterations; 0.005/5 for 132k iterations; 0.005/50 for 26k iterations\nmodel P1: 0.005 for 148k iterations; 0.005/5 for 62k iterations; 0.005/50 for 8k iterations\n\nTraining took 0.5-0.65 hour per 1000 iterations with 4 T4. So in total the training the 4 models took about 22, 18, 10 and 5 days respectively.\n\nPublic/private scores of each (`max_per_img=120, thr=0`, without TTA):\n\n| model | public score | private score |                |\n|-------|--------------|---------------|----------------|\n| L1    | 0.4708       | 0.4354        | on 275 classes |\n| L2    | 0.4705       | 0.4275        | on 275 classes |\n| L3    | 0.4712       | 0.4324        | on 275 classes |\n| P1    | 0.0349       | 0.0344        | on 25 classes  |\n\n\n## TTA\n\nWe used the TTA implementation from [Miras Amir's winning solution](https://github.com/amirassov/kaggle-imaterialist) of iMaterialist (Fashion) 2019 \n\nFor the leaf model, it’s the ensemble of 3 single models, 2 scale (1333,800) and (1600,960), and flip. At RPN and BB stages, it’s the NMS ensemble of 12 single models. In the end, it’s the mean of all 12 masks for each instance. \n\nTTA inference of the leaf model is very slow, which took about 525 T4-hours. We split the 99999 images into 25 chunks and ran it in parallel. \n\nFor the parent model, it’s the ensemble of 2 scale and flip, but just one single model.\n\nLastly, for the 25 parent and grandparent classes, we implemented NMS ensemble at mask level (i.e. calculating the mask iou instead of bbox iou when determining which ones to suppress) to ensemble (i) leaf class predications expanded to 25 classes, and (ii) parent class predications expand to 25 classes.\n\nPublic/private scores of each (`max_per_img=120, thr=0`):\n\n| model                    | public score | private score |                |\n|--------------------------|--------------|---------------|----------------|\n| leaf model ensemble      | 0.5007       | 0.4548        | on 275 classes |\n| leaf ensemble expanded   | 0.0325       | 0.0330        | on 25 classes  |\n| ensemble of (i) and (ii) | 0.0370       | 0.0370        | on 25 classes  |\n\nAt this point, the total score is public 0.5007+0.0370 = 0.5378, private 0.4548+0.0370 = 0.4918. \n\nThere’s one last thing we did to boost scores to 0.5383/0.4922: set `max_per_img=200` for the leaf ensemble. But we only had time (also restricted by sub file size) to do this for 12/25 of the test images.\n\n\n## Code\n\nTraining, inference, pre- and post-process code are available at https://github.com/boliu61/open-images-2019-instance-segmentation\nTrained model weights are also linked in the readme there",
      "votes": null
    },
    {
      "id": "639324",
      "postDate": "10/03/2019 04:30:35",
      "content": "<p>great!</p>",
      "rawMarkdown": "great!",
      "votes": null
    },
    {
      "id": "639343",
      "postDate": "10/03/2019 05:12:35",
      "content": "<p>Congrats.....\nThank you for Sharing your Approach &amp; Insights.....!! <a href=\"/boliu0\">@boliu0</a> </p>",
      "rawMarkdown": "Congrats.....\nThank you for Sharing your Approach &amp; Insights.....!! @boliu0",
      "votes": null
    },
    {
      "id": "639592",
      "postDate": "10/03/2019 11:32:18",
      "content": "<p>Congratulations on winning a gold medal! Very nice solution approaches! Looking forward to your code solutions. :-)</p>",
      "rawMarkdown": "Congratulations on winning a gold medal! Very nice solution approaches! Looking forward to your code solutions. :-)",
      "votes": null
    },
    {
      "id": "642385",
      "postDate": "10/06/2019 01:36:22",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "642561",
      "postDate": "10/06/2019 09:38:14",
      "content": "<p>Congrats, thank you for sharing approach</p>",
      "rawMarkdown": "Congrats, thank you for sharing approach",
      "votes": null
    },
    {
      "id": "643617",
      "postDate": "10/07/2019 17:32:23",
      "content": "<p><strong>Update</strong>: We have pushed code (training, inference, pre-process, post-process) to <a href=\"https://github.com/boliu61/open-images-2019-instance-segmentation\">https://github.com/boliu61/open-images-2019-instance-segmentation</a>\nTrained weights are also linked in the repo readme</p>",
      "rawMarkdown": "**Update**: We have pushed code (training, inference, pre-process, post-process) to https://github.com/boliu61/open-images-2019-instance-segmentation\nTrained weights are also linked in the repo readme",
      "votes": null
    },
    {
      "id": "644584",
      "postDate": "10/09/2019 03:19:28",
      "content": "<p>Congrats!\nI wonder how do you do the operation upsample when rebalancing. Just repeat the images?</p>",
      "rawMarkdown": "Congrats!\nI wonder how do you do the operation upsample when rebalancing. Just repeat the images?",
      "votes": null
    },
    {
      "id": "644897",
      "postDate": "10/09/2019 13:48:07",
      "content": "<p>Thanks. Yes, unsample and oversample mean the same thing in my writeup above. </p>\n\n<p>\"upsample to 10k imgs/class\" was done like this:\nIf a class has 3000 images, then repeat each image 3 times, then randomly choose 1000 (so that they appear 4 times);\nif a class has 7000 images, then randomly choose 3000 (so that these 3000 appear twice, the other 4000 appear once);\n...</p>\n\n<p>This was done using different seeds for each epoch, so that a different 3000 in above example was chosen in each epoch.</p>\n\n<p>Implementation at:\n<a href=\"https://github.com/boliu61/open-images-2019-instance-segmentation/blob/master/util/make_rebalanced_train_ann_parent.py#L116\">https://github.com/boliu61/open-images-2019-instance-segmentation/blob/master/util/make_rebalanced_train_ann_parent.py#L116</a></p>",
      "rawMarkdown": "Thanks. Yes, unsample and oversample mean the same thing in my writeup above. \n\n\"upsample to 10k imgs/class\" was done like this:\nIf a class has 3000 images, then repeat each image 3 times, then randomly choose 1000 (so that they appear 4 times);\nif a class has 7000 images, then randomly choose 3000 (so that these 3000 appear twice, the other 4000 appear once);\n...\n\nThis was done using different seeds for each epoch, so that a different 3000 in above example was chosen in each epoch.\n\nImplementation at:\nhttps://github.com/boliu61/open-images-2019-instance-segmentation/blob/master/util/make_rebalanced_train_ann_parent.py#L116",
      "votes": null
    },
    {
      "id": "645271",
      "postDate": "10/10/2019 00:26:30",
      "content": "<p>thx👍 </p>",
      "rawMarkdown": "thx👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 639324,
      "author_name": "ryyghfjk89",
      "author_url": "",
      "post_date": "10/03/2019 04:30:35",
      "content": "<p>great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 639343,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/03/2019 05:12:35",
      "content": "<p>Congrats.....\nThank you for Sharing your Approach &amp; Insights.....!! <a href=\"/boliu0\">@boliu0</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 639592,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "10/03/2019 11:32:18",
      "content": "<p>Congratulations on winning a gold medal! Very nice solution approaches! Looking forward to your code solutions. :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 642385,
      "author_name": "terenceliu4444",
      "author_url": "",
      "post_date": "10/06/2019 01:36:22",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 642561,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "10/06/2019 09:38:14",
      "content": "<p>Congrats, thank you for sharing approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 643617,
      "author_name": "boliu0",
      "author_url": "",
      "post_date": "10/07/2019 17:32:23",
      "content": "<p><strong>Update</strong>: We have pushed code (training, inference, pre-process, post-process) to <a href=\"https://github.com/boliu61/open-images-2019-instance-segmentation\">https://github.com/boliu61/open-images-2019-instance-segmentation</a>\nTrained weights are also linked in the repo readme</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 644584,
      "author_name": "wangchuanyuan",
      "author_url": "",
      "post_date": "10/09/2019 03:19:28",
      "content": "<p>Congrats!\nI wonder how do you do the operation upsample when rebalancing. Just repeat the images?</p>",
      "votes": null,
      "replies": [
        {
          "id": 644897,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "10/09/2019 13:48:07",
          "content": "<p>Thanks. Yes, unsample and oversample mean the same thing in my writeup above. </p>\n\n<p>\"upsample to 10k imgs/class\" was done like this:\nIf a class has 3000 images, then repeat each image 3 times, then randomly choose 1000 (so that they appear 4 times);\nif a class has 7000 images, then randomly choose 3000 (so that these 3000 appear twice, the other 4000 appear once);\n...</p>\n\n<p>This was done using different seeds for each epoch, so that a different 3000 in above example was chosen in each epoch.</p>\n\n<p>Implementation at:\n<a href=\"https://github.com/boliu61/open-images-2019-instance-segmentation/blob/master/util/make_rebalanced_train_ann_parent.py#L116\">https://github.com/boliu61/open-images-2019-instance-segmentation/blob/master/util/make_rebalanced_train_ann_parent.py#L116</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 645271,
          "author_name": "wangchuanyuan",
          "author_url": "",
          "post_date": "10/10/2019 00:26:30",
          "content": "<p>thx👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "639047": "Congratulations to all the winners. Here is our solution:\n\n## Codebase and hardware\n\nWe used [mmdetection](https://github.com/open-mmlab/mmdetection) with 4 Tesla T4 during training, 8 Tesla T4 during TTA inference.\n\n## Model setup\n\nThere are 300 classes, in 3 hierarchical levels: \n-  275 leaf classes\n-  23 parent classes \n-  2 grandparent classes (`Carnivore` and `Reptile`)\n\nFor the 275 leaf classes, we train models with these 275 classes as labels and use inference results directly.\n\nFor the 25 parent and grandparent classes, there are two methods to get the predictions:\n- (i). Use the prediction results from the leaf model, and “expand” to the parent and grandparent. For example, if the leaf model predicted a `Tortoise` mask, we add a `Turtle` prediction (its parent) and a `Reptile` prediction (its grandparent) with the same mask and score.  \n- (ii). We train models with only the 23 parent classes as labels, and expand to the 2 grandparent classes. Note that for these models, we need to create hierarchical expansion of the instance segmentation before training (as explained [here](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track)).\n\n\n## Validation set\n\nWe sampled a small subset from the official validation set for faster inference, 2844 images for the leaf model and 3841 images for the parent model. We used the validation scores to decide when to adjust learning rate, and to choose best checkpoints. \n\nThe validation scores are significantly higher than LB. For the 275 leaf classes, we have validation score over 0.8. But nevertheless, the delta between validation and LB are stable.\n\n\n## Training set rebalancing\n\nThe open images dataset is extremely imbalanced between the 300 classes. We rebalanced the training data for the leaf model as follows:\nSort the 275 classes by descending number of images. So class #0 is the largest class.\n\n| Group | Class      | original num of imgs per class | rebalancing                   |\n|-----|------------|--------------------------------|-------------------------------|\n| 1     | #241 to #274 | 150-13                         | oversample x10                |\n| 2     | #64 to #240  | 1500-150                       | oversample to 1500 imgs/class |\n| 3     | #24 to #63   | 6k-1500                        | no rebalancing                |\n| 4     | #0 to #23    | 89k-6k                         | downsample to 6k*            |\n\nEach epoch has about 450k images after rebalancing. \n\n*For the downsample, we use different random seeds for each epoch to feed the model as many images as possible. Also, this 6k number is a rough target. In actual sampling, the top few classes ended up having more images because many images have multiple labels, and top classes already have more than 6k images after Group 1-3 sampling are done.\n\nSimilarly for the parent model, sort the 23 classes\n\n| Group | Class    | rebalancing                  |\n|-------|----------|------------------------------|\n| 1     | #10 to #22 | upsample to 10k imgs/class   |\n| 2     | #5 to #9   | no rebalancing               |\n| 3     | #0 to #4   | downsample to 30k imgs/class |\n\nEach epoch has about 330k images after rebalancing.\n\n## Single models\n\nWe trained 3 cascade mrcnn leaf models\n- L1: backbone x101\n- L2: backbone r101 + deformable module\n- L3: backbone x101 + deformable module\n\nand 1 cascade mrcnn parent model\n- P1: backbone x101\n\nhyper-paramters:\n```\nimgs_per_gpu=1\nnum_gpu=4\nimg_scale=[(1333, 640), (1333, 960)]\nmultiscale_mode='range'\n```\n\nlr schedule:\nmodel L1: 0.005 for 760k iterations; 0.005/3 for 60k iterations; 0.005/15 for 152k iterations\nmodel L2: 0.005 for 675k iterations; 0.005/3 for 64k iterations; 0.005/15 for 48k iterations\nmodel L3: 0.005 for 226k iterations; 0.005/5 for 132k iterations; 0.005/50 for 26k iterations\nmodel P1: 0.005 for 148k iterations; 0.005/5 for 62k iterations; 0.005/50 for 8k iterations\n\nTraining took 0.5-0.65 hour per 1000 iterations with 4 T4. So in total the training the 4 models took about 22, 18, 10 and 5 days respectively.\n\nPublic/private scores of each (`max_per_img=120, thr=0`, without TTA):\n\n| model | public score | private score |                |\n|-------|--------------|---------------|----------------|\n| L1    | 0.4708       | 0.4354        | on 275 classes |\n| L2    | 0.4705       | 0.4275        | on 275 classes |\n| L3    | 0.4712       | 0.4324        | on 275 classes |\n| P1    | 0.0349       | 0.0344        | on 25 classes  |\n\n\n## TTA\n\nWe used the TTA implementation from [Miras Amir's winning solution](https://github.com/amirassov/kaggle-imaterialist) of iMaterialist (Fashion) 2019 \n\nFor the leaf model, it’s the ensemble of 3 single models, 2 scale (1333,800) and (1600,960), and flip. At RPN and BB stages, it’s the NMS ensemble of 12 single models. In the end, it’s the mean of all 12 masks for each instance. \n\nTTA inference of the leaf model is very slow, which took about 525 T4-hours. We split the 99999 images into 25 chunks and ran it in parallel. \n\nFor the parent model, it’s the ensemble of 2 scale and flip, but just one single model.\n\nLastly, for the 25 parent and grandparent classes, we implemented NMS ensemble at mask level (i.e. calculating the mask iou instead of bbox iou when determining which ones to suppress) to ensemble (i) leaf class predications expanded to 25 classes, and (ii) parent class predications expand to 25 classes.\n\nPublic/private scores of each (`max_per_img=120, thr=0`):\n\n| model                    | public score | private score |                |\n|--------------------------|--------------|---------------|----------------|\n| leaf model ensemble      | 0.5007       | 0.4548        | on 275 classes |\n| leaf ensemble expanded   | 0.0325       | 0.0330        | on 25 classes  |\n| ensemble of (i) and (ii) | 0.0370       | 0.0370        | on 25 classes  |\n\nAt this point, the total score is public 0.5007+0.0370 = 0.5378, private 0.4548+0.0370 = 0.4918. \n\nThere’s one last thing we did to boost scores to 0.5383/0.4922: set `max_per_img=200` for the leaf ensemble. But we only had time (also restricted by sub file size) to do this for 12/25 of the test images.\n\n\n## Code\n\nTraining, inference, pre- and post-process code are available at https://github.com/boliu61/open-images-2019-instance-segmentation\nTrained model weights are also linked in the readme there",
    "639324": "great!",
    "639343": "Congrats.....\nThank you for Sharing your Approach &amp; Insights.....!! @boliu0",
    "639592": "Congratulations on winning a gold medal! Very nice solution approaches! Looking forward to your code solutions. :-)",
    "642385": "Congratulations!",
    "642561": "Congrats, thank you for sharing approach",
    "643617": "**Update**: We have pushed code (training, inference, pre-process, post-process) to https://github.com/boliu61/open-images-2019-instance-segmentation\nTrained weights are also linked in the repo readme",
    "644584": "Congrats!\nI wonder how do you do the operation upsample when rebalancing. Just repeat the images?",
    "644897": "Thanks. Yes, unsample and oversample mean the same thing in my writeup above. \n\n\"upsample to 10k imgs/class\" was done like this:\nIf a class has 3000 images, then repeat each image 3 times, then randomly choose 1000 (so that they appear 4 times);\nif a class has 7000 images, then randomly choose 3000 (so that these 3000 appear twice, the other 4000 appear once);\n...\n\nThis was done using different seeds for each epoch, so that a different 3000 in above example was chosen in each epoch.\n\nImplementation at:\nhttps://github.com/boliu61/open-images-2019-instance-segmentation/blob/master/util/make_rebalanced_train_ann_parent.py#L116",
    "645271": "thx👍"
  },
  "source": "meta"
}