{
  "id": 110908,
  "title": " 20th place solution - maskrcnn-benchmark baseline",
  "url": "/competitions/open-images-2019-instance-segmentation/writeups/yu4u-20th-place-solution-maskrcnn-benchmark-baseli",
  "author_name": "",
  "post_date": "2023-01-10T15:21:13.337Z",
  "votes": 24,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to Open Images and Kaggle team for this great competition(s) and congrats to all (tentative) prize and medal winners!</p>\n<p>My result is not outstanding but the solution might be valuable to be shared because I used the famous maskrcnn-benchmark library 'as it is' and also used its outputs as it is without TTA or any post processing. Training two models requires only 14 hours (x2) using V100 8GPUs.</p>\n<p>All codes are available at: <a href=\"https://github.com/yu4u/kaggle-open-images-2019-instance-segmentation\" target=\"_blank\">https://github.com/yu4u/kaggle-open-images-2019-instance-segmentation</a></p>\n<p>There are mainly two issues to be solved in this competition and the Object Detection track: (1) class imbalance and (2) class hierarchy. I tackled these issues only on a dataset creation side. The former is easy to handle: use fixed number of training images for each class. In this post, I mainly describe how to handle class hierarchy.</p>\n<p>Firstly, I divided all classes into two groups: layer0 and layer1. From challenge-2019-label300-segmentable-hierarchy.json we can see that:</p>\n<ol>\n<li>Maximum depth of hierarchy is 2 (starting from 0)</li>\n<li>The number of depth 2 classes is only 5.</li>\n</ol>\n<pre><code>Carnivore\n└── Bear\n    ├── Brown bear\n    ├── Polar bear\n    └── Teddy bear &lt;---　Are you serious?\n\nReptile\n└── Turtle\n    ├── Tortoise\n    └── Sea turtle\n</code></pre>\n<p>Thus, I decided to group depth 0 classes as layer0 group and depth 1 and 2 classes together as layer1 group. The idea is to make different model for each of two groups.<br>\nIn training each model, a dedicated dataset is used, which includes only the target group class instances. By doing so, there is no need to care about class hierarchy.<br>\nHowever, practically, it is impossible to make dataset from only training images that includes only target classes and does not include non-target classes. Therefore, I removed non-target class instances from training images.</p>\n<p>For layer0 group dataset:</p>\n<ol>\n<li>Remove non-target class annotations that occlude target class object 25% or more</li>\n<li>Convert non-target class to its parent class (Thus it becomes target class. Some classes need to be processed twice. 'Teddy bear' is converted only to 'Toy', not 'Carnivore')</li>\n</ol>\n<p>For layer1 group dataset:</p>\n<ol>\n<li>Remove non-target class annotations that occlude target class object 25% or more</li>\n<li>Remove non-target class annotations that do not have any child class (no impact to layer1 group classes because there is no relationship between them)</li>\n<li>Remove non-target class annotations that have some child classes, and fill their bbox with gray in the training image (removing only annotations is not good idea because these cause 'false false positive' signal (loss) to the model)</li>\n</ol>\n<p>That's all, and let's train!</p>",
  "messages": [
    {
      "id": "638554",
      "postDate": "10/02/2019 03:22:52",
      "content": "<p>Thanks to Open Images and Kaggle team for this great competition(s) and congrats to all (tentative) prize and medal winners!</p>\n<p>My result is not outstanding but the solution might be valuable to be shared because I used the famous maskrcnn-benchmark library 'as it is' and also used its outputs as it is without TTA or any post processing. Training two models requires only 14 hours (x2) using V100 8GPUs.</p>\n<p>All codes are available at: <a href=\"https://github.com/yu4u/kaggle-open-images-2019-instance-segmentation\" target=\"_blank\">https://github.com/yu4u/kaggle-open-images-2019-instance-segmentation</a></p>\n<p>There are mainly two issues to be solved in this competition and the Object Detection track: (1) class imbalance and (2) class hierarchy. I tackled these issues only on a dataset creation side. The former is easy to handle: use fixed number of training images for each class. In this post, I mainly describe how to handle class hierarchy.</p>\n<p>Firstly, I divided all classes into two groups: layer0 and layer1. From challenge-2019-label300-segmentable-hierarchy.json we can see that:</p>\n<ol>\n<li>Maximum depth of hierarchy is 2 (starting from 0)</li>\n<li>The number of depth 2 classes is only 5.</li>\n</ol>\n<pre><code>Carnivore\n└── Bear\n    ├── Brown bear\n    ├── Polar bear\n    └── Teddy bear &lt;---　Are you serious?\n\nReptile\n└── Turtle\n    ├── Tortoise\n    └── Sea turtle\n</code></pre>\n<p>Thus, I decided to group depth 0 classes as layer0 group and depth 1 and 2 classes together as layer1 group. The idea is to make different model for each of two groups.<br>\nIn training each model, a dedicated dataset is used, which includes only the target group class instances. By doing so, there is no need to care about class hierarchy.<br>\nHowever, practically, it is impossible to make dataset from only training images that includes only target classes and does not include non-target classes. Therefore, I removed non-target class instances from training images.</p>\n<p>For layer0 group dataset:</p>\n<ol>\n<li>Remove non-target class annotations that occlude target class object 25% or more</li>\n<li>Convert non-target class to its parent class (Thus it becomes target class. Some classes need to be processed twice. 'Teddy bear' is converted only to 'Toy', not 'Carnivore')</li>\n</ol>\n<p>For layer1 group dataset:</p>\n<ol>\n<li>Remove non-target class annotations that occlude target class object 25% or more</li>\n<li>Remove non-target class annotations that do not have any child class (no impact to layer1 group classes because there is no relationship between them)</li>\n<li>Remove non-target class annotations that have some child classes, and fill their bbox with gray in the training image (removing only annotations is not good idea because these cause 'false false positive' signal (loss) to the model)</li>\n</ol>\n<p>That's all, and let's train!</p>",
      "rawMarkdown": "Thanks to Open Images and Kaggle team for this great competition(s) and congrats to all (tentative) prize and medal winners!\n\nMy result is not outstanding but the solution might be valuable to be shared because I used the famous maskrcnn-benchmark library 'as it is' and also used its outputs as it is without TTA or any post processing. Training two models requires only 14 hours (x2) using V100 8GPUs.\n\nAll codes are available at: https://github.com/yu4u/kaggle-open-images-2019-instance-segmentation\n\nThere are mainly two issues to be solved in this competition and the Object Detection track: (1) class imbalance and (2) class hierarchy. I tackled these issues only on a dataset creation side. The former is easy to handle: use fixed number of training images for each class. In this post, I mainly describe how to handle class hierarchy.\n\nFirstly, I divided all classes into two groups: layer0 and layer1. From challenge-2019-label300-segmentable-hierarchy.json we can see that:\n1. Maximum depth of hierarchy is 2 (starting from 0)\n2. The number of depth 2 classes is only 5.\n\n```\nCarnivore\n└── Bear\n    ├── Brown bear\n    ├── Polar bear\n    └── Teddy bear &lt;---　Are you serious?\n\nReptile\n└── Turtle\n    ├── Tortoise\n    └── Sea turtle\n```\n\nThus, I decided to group depth 0 classes as layer0 group and depth 1 and 2 classes together as layer1 group. The idea is to make different model for each of two groups.\nIn training each model, a dedicated dataset is used, which includes only the target group class instances. By doing so, there is no need to care about class hierarchy.\nHowever, practically, it is impossible to make dataset from only training images that includes only target classes and does not include non-target classes. Therefore, I removed non-target class instances from training images.\n\nFor layer0 group dataset:\n\n1. Remove non-target class annotations that occlude target class object 25% or more\n2. Convert non-target class to its parent class (Thus it becomes target class. Some classes need to be processed twice. 'Teddy bear' is converted only to 'Toy', not 'Carnivore')\n\nFor layer1 group dataset:\n\n1. Remove non-target class annotations that occlude target class object 25% or more\n2. Remove non-target class annotations that do not have any child class (no impact to layer1 group classes because there is no relationship between them)\n3. Remove non-target class annotations that have some child classes, and fill their bbox with gray in the training image (removing only annotations is not good idea because these cause 'false false positive' signal (loss) to the model)\n\nThat's all, and let's train!",
      "votes": null
    },
    {
      "id": "638585",
      "postDate": "10/02/2019 04:56:41",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights... <a href=\"/ren4yu\">@ren4yu</a> </p>",
      "rawMarkdown": "Congrats\nThank you for Sharing your Approach &amp; Insights... @ren4yu",
      "votes": null
    },
    {
      "id": "638852",
      "postDate": "10/02/2019 13:39:50",
      "content": "<p>Yeah, teddy bear class is misplaced. I'm wondering why couldn't they make the class hierarchy a DAG?</p>",
      "rawMarkdown": "Yeah, teddy bear class is misplaced. I'm wondering why couldn't they make the class hierarchy a DAG?",
      "votes": null
    },
    {
      "id": "638897",
      "postDate": "10/02/2019 14:33:53",
      "content": "<p>Congratulations! :-)</p>",
      "rawMarkdown": "Congratulations! :-)",
      "votes": null
    },
    {
      "id": "642556",
      "postDate": "10/06/2019 09:29:24",
      "content": "<p>Congrats, thank you for sharing code</p>",
      "rawMarkdown": "Congrats, thank you for sharing code",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 638585,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/02/2019 04:56:41",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights... <a href=\"/ren4yu\">@ren4yu</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638852,
      "author_name": "artyomp",
      "author_url": "",
      "post_date": "10/02/2019 13:39:50",
      "content": "<p>Yeah, teddy bear class is misplaced. I'm wondering why couldn't they make the class hierarchy a DAG?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638897,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "10/02/2019 14:33:53",
      "content": "<p>Congratulations! :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 642556,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "10/06/2019 09:29:24",
      "content": "<p>Congrats, thank you for sharing code</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "638554": "Thanks to Open Images and Kaggle team for this great competition(s) and congrats to all (tentative) prize and medal winners!\n\nMy result is not outstanding but the solution might be valuable to be shared because I used the famous maskrcnn-benchmark library 'as it is' and also used its outputs as it is without TTA or any post processing. Training two models requires only 14 hours (x2) using V100 8GPUs.\n\nAll codes are available at: https://github.com/yu4u/kaggle-open-images-2019-instance-segmentation\n\nThere are mainly two issues to be solved in this competition and the Object Detection track: (1) class imbalance and (2) class hierarchy. I tackled these issues only on a dataset creation side. The former is easy to handle: use fixed number of training images for each class. In this post, I mainly describe how to handle class hierarchy.\n\nFirstly, I divided all classes into two groups: layer0 and layer1. From challenge-2019-label300-segmentable-hierarchy.json we can see that:\n1. Maximum depth of hierarchy is 2 (starting from 0)\n2. The number of depth 2 classes is only 5.\n\n```\nCarnivore\n└── Bear\n    ├── Brown bear\n    ├── Polar bear\n    └── Teddy bear &lt;---　Are you serious?\n\nReptile\n└── Turtle\n    ├── Tortoise\n    └── Sea turtle\n```\n\nThus, I decided to group depth 0 classes as layer0 group and depth 1 and 2 classes together as layer1 group. The idea is to make different model for each of two groups.\nIn training each model, a dedicated dataset is used, which includes only the target group class instances. By doing so, there is no need to care about class hierarchy.\nHowever, practically, it is impossible to make dataset from only training images that includes only target classes and does not include non-target classes. Therefore, I removed non-target class instances from training images.\n\nFor layer0 group dataset:\n\n1. Remove non-target class annotations that occlude target class object 25% or more\n2. Convert non-target class to its parent class (Thus it becomes target class. Some classes need to be processed twice. 'Teddy bear' is converted only to 'Toy', not 'Carnivore')\n\nFor layer1 group dataset:\n\n1. Remove non-target class annotations that occlude target class object 25% or more\n2. Remove non-target class annotations that do not have any child class (no impact to layer1 group classes because there is no relationship between them)\n3. Remove non-target class annotations that have some child classes, and fill their bbox with gray in the training image (removing only annotations is not good idea because these cause 'false false positive' signal (loss) to the model)\n\nThat's all, and let's train!",
    "638585": "Congrats\nThank you for Sharing your Approach &amp; Insights... @ren4yu",
    "638852": "Yeah, teddy bear class is misplaced. I'm wondering why couldn't they make the class hierarchy a DAG?",
    "638897": "Congratulations! :-)",
    "642556": "Congrats, thank you for sharing code"
  },
  "source": "meta"
}