{
  "id": 95249,
  "title": "16-th Solution: Mask-RCNN on Steroids",
  "url": "/competitions/imaterialist-fashion-2019-FGVC6/writeups/kaggle-cn-alex-doesn-t-know-how-to-code-16-th-solu",
  "author_name": "",
  "post_date": "2019-06-11T03:54:41.765531100Z",
  "votes": 28,
  "comment_count": 9,
  "views": 0,
  "content": "<p><em>(Also known as Alex doesn't know how to code)</em></p>\n\n<p>After seeing first place's solution, I was shocked at how close I was to reaching the gold area... This competition has taught me a lot on hyper-parameter fine-tuning and architecture selection for instance segmentation, and the greatest (and most painful) of them all is the threshold for non-maximum-suppression- and I fell for it. This is my first serious attempt at aiming for grandmaster status, and though the outcome is not so satisfactory, for the most part I enjoyed the competition.</p>\n\n<p>For my solution, it is based on mmdetection, an awesome codebase for detection related tasks. </p>\n\n<h3>I started with Hybrid Task Cascade with ResNeXt-101-64x4d-FPN backbone, with the following modifications:</h3>\n\n<ol>\n<li><p>DCN-v2 for part of the backbone</p></li>\n<li><p>Removing mask fusion branch (since it takes too much memory and semantic information in this task is tricky to define)</p></li>\n<li><p>Multi-scale training with flipping and resizing range (512,512) -&gt; (1333,1333)</p></li>\n<li><p>Guided-anchoring RPN</p></li>\n<li><p>Increasing initial mask branch output from 14 pixels to 28 pixels</p></li>\n</ol>\n\n<h3>Then I modified the training progress a bit.</h3>\n\n<ol>\n<li><p>Default training schedule, 1 image per GPU for 4/8 GPUs</p></li>\n<li><p>Changing mask loss from BCE to symmetric lovasz:</p></li>\n</ol>\n\n<p><code>def symmetric_lovasz(outputs, targets):\nreturn (lovasz_hinge(outputs, targets) + lovasz_hinge(-outputs, 1 - targets)) / 2</code></p>\n\n<ol>\n<li><p>Changing smooth L1 loss to GHM-R loss</p></li>\n<li><p>Changing sigmoid-based losses to Focal Loss and GHM-C loss</p></li>\n</ol>\n\n<h3>Then, for inference:</h3>\n\n<ol>\n<li><p>Used soft-nms ( </p>\n\n<h1>but I didn't fine-tune the threshold :(</h1></li>\n<li><p>Multi-scale testing in mmdetection wasn't implemented, so I wrote one myself that merge predictions in multiple scales in proposals, bboxes and masks. The following scales are used (with flipping): [(1333,1333), (1333,800), (800,1333), (900,600)]</p></li>\n<li><p>Deleting masks with low confidence &amp; low pixels (with fine-tuned threshold)</p></li>\n<li><p>Heuristic that make sure that the same pixel are not assigned to two instances w. the same ClassId</p></li>\n<li><p>(I couldn't submit for the final minutes so I didn't have a chance to test this out) Binary classifier that predict if an object is fine-grained or not. Then, in final submission, if an instance is determined to be fine-grained, it is excluded. The rationale is as below:</p>\n\n<p>a) An included object that is fine-grained but the mask generated is not: 1 FP + 1 FN\nb) Simply not predicting that object: 1 FN</p></li>\n</ol>\n\n<p>And for this network, it's a se-resnext50 based image classifier with 4-channel input (image + binary instance mask) and one-hot category fusion in the fc-layer. Honestly I could be doing a DAE before the classifier (from dirty testing mask to imagined clean train mask) but I ran out of GPUs and time.</p>\n\n<h3>What didn't work:</h3>\n\n<ol>\n<li><p>Un-freezing bn / gn / synchronized bn</p></li>\n<li><p>Parallelized testing</p></li>\n<li><p>After-nms ensemble (like the one used in Top3)</p></li>\n<li><p>Attribute classification (my local F1 was &lt; 0.1)</p></li>\n</ol>\n\n<h3>Machines</h3>\n\n<p>I used a variety of rented cloud machines, from 1 * 2080 Ti, 1 * P6000 to 8 * 1080 Ti to 4 * P100. I've basically spent all of the prizes I earned in the Whales competition on this one, plus a few hundred bucks :(</p>\n\n<p>For your reference:</p>\n\n<ol>\n<li><p>For 1-image per GPU, you can fit an (1333,1333) image into an RTX 2080 Ti w.o. the enlargement of mask branch or GA-RPN</p></li>\n<li><p>But if you want to add GA-RPN, only GTX 1080 Ti would suffice (the pitfall: <strong>1080 Ti has a little bit more memory than it's RTX counterpart</strong>)</p></li>\n<li><p>My full setting fits barely in an P100.</p></li>\n</ol>\n\n<h3>Final words</h3>\n\n<p>I have to say goodbye to kaggle for the following months due to internships and school-related stuffs, so it's a little discouraging to know that I missed the last chance this year to getting a solo gold by such a small margin. But I guess the journey and the things I've learnt is more important than the recognition itself. So I will come back with full strength after a small break. Also congratulations to the winners, you guys did a fantastic job!</p>",
  "messages": [
    {
      "id": "549807",
      "postDate": "06/11/2019 03:54:41",
      "content": "<p><em>(Also known as Alex doesn't know how to code)</em></p>\n\n<p>After seeing first place's solution, I was shocked at how close I was to reaching the gold area... This competition has taught me a lot on hyper-parameter fine-tuning and architecture selection for instance segmentation, and the greatest (and most painful) of them all is the threshold for non-maximum-suppression- and I fell for it. This is my first serious attempt at aiming for grandmaster status, and though the outcome is not so satisfactory, for the most part I enjoyed the competition.</p>\n\n<p>For my solution, it is based on mmdetection, an awesome codebase for detection related tasks. </p>\n\n<h3>I started with Hybrid Task Cascade with ResNeXt-101-64x4d-FPN backbone, with the following modifications:</h3>\n\n<ol>\n<li><p>DCN-v2 for part of the backbone</p></li>\n<li><p>Removing mask fusion branch (since it takes too much memory and semantic information in this task is tricky to define)</p></li>\n<li><p>Multi-scale training with flipping and resizing range (512,512) -&gt; (1333,1333)</p></li>\n<li><p>Guided-anchoring RPN</p></li>\n<li><p>Increasing initial mask branch output from 14 pixels to 28 pixels</p></li>\n</ol>\n\n<h3>Then I modified the training progress a bit.</h3>\n\n<ol>\n<li><p>Default training schedule, 1 image per GPU for 4/8 GPUs</p></li>\n<li><p>Changing mask loss from BCE to symmetric lovasz:</p></li>\n</ol>\n\n<p><code>def symmetric_lovasz(outputs, targets):\nreturn (lovasz_hinge(outputs, targets) + lovasz_hinge(-outputs, 1 - targets)) / 2</code></p>\n\n<ol>\n<li><p>Changing smooth L1 loss to GHM-R loss</p></li>\n<li><p>Changing sigmoid-based losses to Focal Loss and GHM-C loss</p></li>\n</ol>\n\n<h3>Then, for inference:</h3>\n\n<ol>\n<li><p>Used soft-nms ( </p>\n\n<h1>but I didn't fine-tune the threshold :(</h1></li>\n<li><p>Multi-scale testing in mmdetection wasn't implemented, so I wrote one myself that merge predictions in multiple scales in proposals, bboxes and masks. The following scales are used (with flipping): [(1333,1333), (1333,800), (800,1333), (900,600)]</p></li>\n<li><p>Deleting masks with low confidence &amp; low pixels (with fine-tuned threshold)</p></li>\n<li><p>Heuristic that make sure that the same pixel are not assigned to two instances w. the same ClassId</p></li>\n<li><p>(I couldn't submit for the final minutes so I didn't have a chance to test this out) Binary classifier that predict if an object is fine-grained or not. Then, in final submission, if an instance is determined to be fine-grained, it is excluded. The rationale is as below:</p>\n\n<p>a) An included object that is fine-grained but the mask generated is not: 1 FP + 1 FN\nb) Simply not predicting that object: 1 FN</p></li>\n</ol>\n\n<p>And for this network, it's a se-resnext50 based image classifier with 4-channel input (image + binary instance mask) and one-hot category fusion in the fc-layer. Honestly I could be doing a DAE before the classifier (from dirty testing mask to imagined clean train mask) but I ran out of GPUs and time.</p>\n\n<h3>What didn't work:</h3>\n\n<ol>\n<li><p>Un-freezing bn / gn / synchronized bn</p></li>\n<li><p>Parallelized testing</p></li>\n<li><p>After-nms ensemble (like the one used in Top3)</p></li>\n<li><p>Attribute classification (my local F1 was &lt; 0.1)</p></li>\n</ol>\n\n<h3>Machines</h3>\n\n<p>I used a variety of rented cloud machines, from 1 * 2080 Ti, 1 * P6000 to 8 * 1080 Ti to 4 * P100. I've basically spent all of the prizes I earned in the Whales competition on this one, plus a few hundred bucks :(</p>\n\n<p>For your reference:</p>\n\n<ol>\n<li><p>For 1-image per GPU, you can fit an (1333,1333) image into an RTX 2080 Ti w.o. the enlargement of mask branch or GA-RPN</p></li>\n<li><p>But if you want to add GA-RPN, only GTX 1080 Ti would suffice (the pitfall: <strong>1080 Ti has a little bit more memory than it's RTX counterpart</strong>)</p></li>\n<li><p>My full setting fits barely in an P100.</p></li>\n</ol>\n\n<h3>Final words</h3>\n\n<p>I have to say goodbye to kaggle for the following months due to internships and school-related stuffs, so it's a little discouraging to know that I missed the last chance this year to getting a solo gold by such a small margin. But I guess the journey and the things I've learnt is more important than the recognition itself. So I will come back with full strength after a small break. Also congratulations to the winners, you guys did a fantastic job!</p>",
      "rawMarkdown": "*(Also known as Alex doesn't know how to code)*\n\nAfter seeing first place's solution, I was shocked at how close I was to reaching the gold area... This competition has taught me a lot on hyper-parameter fine-tuning and architecture selection for instance segmentation, and the greatest (and most painful) of them all is the threshold for non-maximum-suppression- and I fell for it. This is my first serious attempt at aiming for grandmaster status, and though the outcome is not so satisfactory, for the most part I enjoyed the competition.\n\nFor my solution, it is based on mmdetection, an awesome codebase for detection related tasks. \n\n### I started with Hybrid Task Cascade with ResNeXt-101-64x4d-FPN backbone, with the following modifications:\n\n1. DCN-v2 for part of the backbone\n\n2. Removing mask fusion branch (since it takes too much memory and semantic information in this task is tricky to define)\n\n3. Multi-scale training with flipping and resizing range (512,512) -&gt; (1333,1333)\n\n4. Guided-anchoring RPN\n\n5. Increasing initial mask branch output from 14 pixels to 28 pixels\n\n### Then I modified the training progress a bit.\n\n1. Default training schedule, 1 image per GPU for 4/8 GPUs\n\n2. Changing mask loss from BCE to symmetric lovasz:\n\n`def symmetric_lovasz(outputs, targets):\nreturn (lovasz_hinge(outputs, targets) + lovasz_hinge(-outputs, 1 - targets)) / 2`\n\n3. Changing smooth L1 loss to GHM-R loss\n\n4. Changing sigmoid-based losses to Focal Loss and GHM-C loss\n\n### Then, for inference:\n\n1. Used soft-nms ( \n#but I didn't fine-tune the threshold :(\n\n\n2. Multi-scale testing in mmdetection wasn't implemented, so I wrote one myself that merge predictions in multiple scales in proposals, bboxes and masks. The following scales are used (with flipping): [(1333,1333), (1333,800), (800,1333), (900,600)]\n\n3. Deleting masks with low confidence &amp; low pixels (with fine-tuned threshold)\n\n4. Heuristic that make sure that the same pixel are not assigned to two instances w. the same ClassId\n\n5. (I couldn't submit for the final minutes so I didn't have a chance to test this out) Binary classifier that predict if an object is fine-grained or not. Then, in final submission, if an instance is determined to be fine-grained, it is excluded. The rationale is as below:\n\n  a) An included object that is fine-grained but the mask generated is not: 1 FP + 1 FN\n  b) Simply not predicting that object: 1 FN\n\nAnd for this network, it's a se-resnext50 based image classifier with 4-channel input (image + binary instance mask) and one-hot category fusion in the fc-layer. Honestly I could be doing a DAE before the classifier (from dirty testing mask to imagined clean train mask) but I ran out of GPUs and time.\n\n### What didn't work:\n\n1. Un-freezing bn / gn / synchronized bn\n\n2. Parallelized testing\n\n3. After-nms ensemble (like the one used in Top3)\n\n4. Attribute classification (my local F1 was &lt; 0.1)\n\n### Machines\n\nI used a variety of rented cloud machines, from 1 * 2080 Ti, 1 * P6000 to 8 * 1080 Ti to 4 * P100. I've basically spent all of the prizes I earned in the Whales competition on this one, plus a few hundred bucks :(\n\nFor your reference:\n\n1. For 1-image per GPU, you can fit an (1333,1333) image into an RTX 2080 Ti w.o. the enlargement of mask branch or GA-RPN\n\n2. But if you want to add GA-RPN, only GTX 1080 Ti would suffice (the pitfall: **1080 Ti has a little bit more memory than it's RTX counterpart**)\n\n3. My full setting fits barely in an P100.\n\n### Final words\n\nI have to say goodbye to kaggle for the following months due to internships and school-related stuffs, so it's a little discouraging to know that I missed the last chance this year to getting a solo gold by such a small margin. But I guess the journey and the things I've learnt is more important than the recognition itself. So I will come back with full strength after a small break. Also congratulations to the winners, you guys did a fantastic job!",
      "votes": null
    },
    {
      "id": "549820",
      "postDate": "06/11/2019 04:05:47",
      "content": "<p>Nice work, Alex!\nThanks for sharing.</p>",
      "rawMarkdown": "Nice work, Alex!\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "549821",
      "postDate": "06/11/2019 04:06:08",
      "content": "<p>I'll probably create a pull-request in mmdetection to add multi-scaled testing for HTC, but currently my code is way too dirty.....</p>",
      "rawMarkdown": "I'll probably create a pull-request in mmdetection to add multi-scaled testing for HTC, but currently my code is way too dirty.....",
      "votes": null
    },
    {
      "id": "549822",
      "postDate": "06/11/2019 04:07:10",
      "content": "<p>Thank you for sharing!!! </p>",
      "rawMarkdown": "Thank you for sharing!!!",
      "votes": null
    },
    {
      "id": "549926",
      "postDate": "06/11/2019 06:33:35",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "549990",
      "postDate": "06/11/2019 07:47:52",
      "content": "<p>congratulations to all participants! Thanks for sharing.</p>",
      "rawMarkdown": "congratulations to all participants! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "550063",
      "postDate": "06/11/2019 09:09:07",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    },
    {
      "id": "550265",
      "postDate": "06/11/2019 12:56:26",
      "content": "<p>computing power!</p>",
      "rawMarkdown": "computing power!",
      "votes": null
    },
    {
      "id": "550758",
      "postDate": "06/12/2019 02:21:28",
      "content": "<p>Hi <a href=\"/alexanderliao\">@alexanderliao</a>   Thank you for sharing your solution! Will you attend CVPR2019 this year? <a href=\"https://sites.google.com/view/fgvc6/program?authuser=0\">Schedule of our upcoming FGVC workshop can be found here</a>.</p>",
      "rawMarkdown": "Hi @alexanderliao   Thank you for sharing your solution! Will you attend CVPR2019 this year? [Schedule of our upcoming FGVC workshop can be found here](https://sites.google.com/view/fgvc6/program?authuser=0).",
      "votes": null
    },
    {
      "id": "550784",
      "postDate": "06/12/2019 03:06:50",
      "content": "<p>Sorry, but I'm pretty busy with my internship......</p>",
      "rawMarkdown": "Sorry, but I'm pretty busy with my internship......",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 549820,
      "author_name": "lzhbrian",
      "author_url": "",
      "post_date": "06/11/2019 04:05:47",
      "content": "<p>Nice work, Alex!\nThanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549821,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "06/11/2019 04:06:08",
      "content": "<p>I'll probably create a pull-request in mmdetection to add multi-scaled testing for HTC, but currently my code is way too dirty.....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549822,
      "author_name": "hesene",
      "author_url": "",
      "post_date": "06/11/2019 04:07:10",
      "content": "<p>Thank you for sharing!!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549926,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "06/11/2019 06:33:35",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549990,
      "author_name": "brunhs",
      "author_url": "",
      "post_date": "06/11/2019 07:47:52",
      "content": "<p>congratulations to all participants! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550063,
      "author_name": "sanikamal",
      "author_url": "",
      "post_date": "06/11/2019 09:09:07",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550265,
      "author_name": "rooshroosh",
      "author_url": "",
      "post_date": "06/11/2019 12:56:26",
      "content": "<p>computing power!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550758,
      "author_name": "makeitworkjml",
      "author_url": "",
      "post_date": "06/12/2019 02:21:28",
      "content": "<p>Hi <a href=\"/alexanderliao\">@alexanderliao</a>   Thank you for sharing your solution! Will you attend CVPR2019 this year? <a href=\"https://sites.google.com/view/fgvc6/program?authuser=0\">Schedule of our upcoming FGVC workshop can be found here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 550784,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "06/12/2019 03:06:50",
          "content": "<p>Sorry, but I'm pretty busy with my internship......</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "549807": "*(Also known as Alex doesn't know how to code)*\n\nAfter seeing first place's solution, I was shocked at how close I was to reaching the gold area... This competition has taught me a lot on hyper-parameter fine-tuning and architecture selection for instance segmentation, and the greatest (and most painful) of them all is the threshold for non-maximum-suppression- and I fell for it. This is my first serious attempt at aiming for grandmaster status, and though the outcome is not so satisfactory, for the most part I enjoyed the competition.\n\nFor my solution, it is based on mmdetection, an awesome codebase for detection related tasks. \n\n### I started with Hybrid Task Cascade with ResNeXt-101-64x4d-FPN backbone, with the following modifications:\n\n1. DCN-v2 for part of the backbone\n\n2. Removing mask fusion branch (since it takes too much memory and semantic information in this task is tricky to define)\n\n3. Multi-scale training with flipping and resizing range (512,512) -&gt; (1333,1333)\n\n4. Guided-anchoring RPN\n\n5. Increasing initial mask branch output from 14 pixels to 28 pixels\n\n### Then I modified the training progress a bit.\n\n1. Default training schedule, 1 image per GPU for 4/8 GPUs\n\n2. Changing mask loss from BCE to symmetric lovasz:\n\n`def symmetric_lovasz(outputs, targets):\nreturn (lovasz_hinge(outputs, targets) + lovasz_hinge(-outputs, 1 - targets)) / 2`\n\n3. Changing smooth L1 loss to GHM-R loss\n\n4. Changing sigmoid-based losses to Focal Loss and GHM-C loss\n\n### Then, for inference:\n\n1. Used soft-nms ( \n#but I didn't fine-tune the threshold :(\n\n\n2. Multi-scale testing in mmdetection wasn't implemented, so I wrote one myself that merge predictions in multiple scales in proposals, bboxes and masks. The following scales are used (with flipping): [(1333,1333), (1333,800), (800,1333), (900,600)]\n\n3. Deleting masks with low confidence &amp; low pixels (with fine-tuned threshold)\n\n4. Heuristic that make sure that the same pixel are not assigned to two instances w. the same ClassId\n\n5. (I couldn't submit for the final minutes so I didn't have a chance to test this out) Binary classifier that predict if an object is fine-grained or not. Then, in final submission, if an instance is determined to be fine-grained, it is excluded. The rationale is as below:\n\n  a) An included object that is fine-grained but the mask generated is not: 1 FP + 1 FN\n  b) Simply not predicting that object: 1 FN\n\nAnd for this network, it's a se-resnext50 based image classifier with 4-channel input (image + binary instance mask) and one-hot category fusion in the fc-layer. Honestly I could be doing a DAE before the classifier (from dirty testing mask to imagined clean train mask) but I ran out of GPUs and time.\n\n### What didn't work:\n\n1. Un-freezing bn / gn / synchronized bn\n\n2. Parallelized testing\n\n3. After-nms ensemble (like the one used in Top3)\n\n4. Attribute classification (my local F1 was &lt; 0.1)\n\n### Machines\n\nI used a variety of rented cloud machines, from 1 * 2080 Ti, 1 * P6000 to 8 * 1080 Ti to 4 * P100. I've basically spent all of the prizes I earned in the Whales competition on this one, plus a few hundred bucks :(\n\nFor your reference:\n\n1. For 1-image per GPU, you can fit an (1333,1333) image into an RTX 2080 Ti w.o. the enlargement of mask branch or GA-RPN\n\n2. But if you want to add GA-RPN, only GTX 1080 Ti would suffice (the pitfall: **1080 Ti has a little bit more memory than it's RTX counterpart**)\n\n3. My full setting fits barely in an P100.\n\n### Final words\n\nI have to say goodbye to kaggle for the following months due to internships and school-related stuffs, so it's a little discouraging to know that I missed the last chance this year to getting a solo gold by such a small margin. But I guess the journey and the things I've learnt is more important than the recognition itself. So I will come back with full strength after a small break. Also congratulations to the winners, you guys did a fantastic job!",
    "549820": "Nice work, Alex!\nThanks for sharing.",
    "549821": "I'll probably create a pull-request in mmdetection to add multi-scaled testing for HTC, but currently my code is way too dirty.....",
    "549822": "Thank you for sharing!!!",
    "549926": "Congrats!",
    "549990": "congratulations to all participants! Thanks for sharing.",
    "550063": "Thank you for sharing",
    "550265": "computing power!",
    "550758": "Hi @alexanderliao   Thank you for sharing your solution! Will you attend CVPR2019 this year? [Schedule of our upcoming FGVC workshop can be found here](https://sites.google.com/view/fgvc6/program?authuser=0).",
    "550784": "Sorry, but I'm pretty busy with my internship......"
  },
  "source": "meta"
}