{
  "id": 95234,
  "title": "Solution | Private Top 3",
  "url": "/competitions/imaterialist-fashion-2019-FGVC6/discussion/95234",
  "author_name": "Konstantin Gavrilchik",
  "post_date": "2019-06-10T23:47:26.560000",
  "votes": 39,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello everybody and congrats to all who finished this competitions in the gold zone. We want to share our solution and describe main ideas. </p>\n\n<p><strong>Validation metric:</strong>\nWe implemented it according to leaderboard evaluation metric (absolutely the same).</p>\n\n<p><strong>Data preprocessing:</strong>\nWe used different sizes for trainings:\nmin_size: (800, … 960); max_size &lt;=1600</p>\n\n<p><strong>Models:</strong>\n- facebook repo - Mask-RCNN x-101\n- mmdetection repo - Hybrid Task Cascade with X-101-64x4d-FPN backbone and c3-c5  DCN</p>\n\n<p><strong>Blend:</strong>\nWe decided to find a way to blend multi-class instant segmentation.\nFirst of all, we just should match all bounding boxes from all models. Then we can average respective masks. So, we iterate over all boxes and build IOU matrix (where mat[i, j] = iou between i and j boxes). We binarized matrix with IOU &gt; 0.7 (according to validation). Then we build a graph (networks library) and select connected components. Then we consider one connected component as a respective matched bboxes and averaged their masks. Confidence of boxes was calculated as a SUM(confidences all matches boxes) divided by number of models. </p>\n\n<p><strong>TTA-4:</strong>\nAs we have merger (blend algorithm) we can perform TTA:\n- hflip\n- different sizes (with min_size=800 and min_size=960)</p>\n\n<p>It gives ~ +0.01 - 0.015</p>\n\n<p><strong>Second level model:</strong>\nWe want to predict IOU between mask and real mask (according to metric if IOU &lt; 0.5 we have FP, we want to drop all masks with IOU &lt; 0.5 to avoid some FP)</p>\n\n<p>*We extracted the following features from masks:\n*area of predicted mask\n- max, mean, min, median, std of predicted values (for mask)\n- confidence of bboxes\n- left / right side of predicted bbox\n- area of predicted bbox\n- divide one side size to another\n- category_id (categorical feature)\n- max, mean, min, median, std aggregation of all above features by category_id</p>\n\n<p>We trained LGBM + XGB + CatBoost over this dataset with Labels = IOU (predicted and real-mask on out-of-fold predictions)</p>\n\n<p>And obtain ~0.92 roc-auc score (It gives ~ +0.01 according to final metric)</p>\n\n<p><strong>Attributes</strong></p>\n\n<p>We dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.\nRegarding attributes, we tried hard to predict them both inside Mask-RCNN and by a separate model, but didn’t get any improvement. </p>",
  "messages": [
    {
      "id": 549679,
      "postDate": "2019-06-10T23:47:26.560Z",
      "content": "<p>Hello everybody and congrats to all who finished this competitions in the gold zone. We want to share our solution and describe main ideas. </p>\n\n<p><strong>Validation metric:</strong>\nWe implemented it according to leaderboard evaluation metric (absolutely the same).</p>\n\n<p><strong>Data preprocessing:</strong>\nWe used different sizes for trainings:\nmin_size: (800, … 960); max_size &lt;=1600</p>\n\n<p><strong>Models:</strong>\n- facebook repo - Mask-RCNN x-101\n- mmdetection repo - Hybrid Task Cascade with X-101-64x4d-FPN backbone and c3-c5  DCN</p>\n\n<p><strong>Blend:</strong>\nWe decided to find a way to blend multi-class instant segmentation.\nFirst of all, we just should match all bounding boxes from all models. Then we can average respective masks. So, we iterate over all boxes and build IOU matrix (where mat[i, j] = iou between i and j boxes). We binarized matrix with IOU &gt; 0.7 (according to validation). Then we build a graph (networks library) and select connected components. Then we consider one connected component as a respective matched bboxes and averaged their masks. Confidence of boxes was calculated as a SUM(confidences all matches boxes) divided by number of models. </p>\n\n<p><strong>TTA-4:</strong>\nAs we have merger (blend algorithm) we can perform TTA:\n- hflip\n- different sizes (with min_size=800 and min_size=960)</p>\n\n<p>It gives ~ +0.01 - 0.015</p>\n\n<p><strong>Second level model:</strong>\nWe want to predict IOU between mask and real mask (according to metric if IOU &lt; 0.5 we have FP, we want to drop all masks with IOU &lt; 0.5 to avoid some FP)</p>\n\n<p>*We extracted the following features from masks:\n*area of predicted mask\n- max, mean, min, median, std of predicted values (for mask)\n- confidence of bboxes\n- left / right side of predicted bbox\n- area of predicted bbox\n- divide one side size to another\n- category_id (categorical feature)\n- max, mean, min, median, std aggregation of all above features by category_id</p>\n\n<p>We trained LGBM + XGB + CatBoost over this dataset with Labels = IOU (predicted and real-mask on out-of-fold predictions)</p>\n\n<p>And obtain ~0.92 roc-auc score (It gives ~ +0.01 according to final metric)</p>\n\n<p><strong>Attributes</strong></p>\n\n<p>We dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.\nRegarding attributes, we tried hard to predict them both inside Mask-RCNN and by a separate model, but didn’t get any improvement. </p>",
      "rawMarkdown": "Hello everybody and congrats to all who finished this competitions in the gold zone. We want to share our solution and describe main ideas. \n\n**Validation metric:**\nWe implemented it according to leaderboard evaluation metric (absolutely the same).\n\n**Data preprocessing:**\nWe used different sizes for trainings:\nmin_size: (800, … 960); max_size &lt;=1600\n\n**Models:**\n- facebook repo - Mask-RCNN x-101\n- mmdetection repo - Hybrid Task Cascade with X-101-64x4d-FPN backbone and c3-c5  DCN\n\n**Blend:**\nWe decided to find a way to blend multi-class instant segmentation.\nFirst of all, we just should match all bounding boxes from all models. Then we can average respective masks. So, we iterate over all boxes and build IOU matrix (where mat[i, j] = iou between i and j boxes). We binarized matrix with IOU &gt; 0.7 (according to validation). Then we build a graph (networks library) and select connected components. Then we consider one connected component as a respective matched bboxes and averaged their masks. Confidence of boxes was calculated as a SUM(confidences all matches boxes) divided by number of models. \n\n**TTA-4:**\nAs we have merger (blend algorithm) we can perform TTA:\n- hflip\n- different sizes (with min_size=800 and min_size=960)\n\nIt gives ~ +0.01 - 0.015\n\n**Second level model:**\nWe want to predict IOU between mask and real mask (according to metric if IOU &lt; 0.5 we have FP, we want to drop all masks with IOU &lt; 0.5 to avoid some FP)\n\n*We extracted the following features from masks:\n*area of predicted mask\n- max, mean, min, median, std of predicted values (for mask)\n- confidence of bboxes\n- left / right side of predicted bbox\n- area of predicted bbox\n- divide one side size to another\n- category_id (categorical feature)\n- max, mean, min, median, std aggregation of all above features by category_id\n\nWe trained LGBM + XGB + CatBoost over this dataset with Labels = IOU (predicted and real-mask on out-of-fold predictions)\n\nAnd obtain ~0.92 roc-auc score (It gives ~ +0.01 according to final metric)\n\n**Attributes**\n\nWe dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.\nRegarding attributes, we tried hard to predict them both inside Mask-RCNN and by a separate model, but didn’t get any improvement. ",
      "votes": 38
    },
    {
      "id": 555537,
      "postDate": "2019-06-19T03:27:38.517Z",
      "content": "<p>Congrats! Looking forward to your source code!</p>",
      "rawMarkdown": "Congrats! Looking forward to your source code!"
    },
    {
      "id": 550756,
      "postDate": "2019-06-12T02:19:50.233Z",
      "content": "<p>Hi <a href=\"/dempton\">@dempton</a>  Thank you for sharing your solution! Will you attend CVPR2019 this year? <a href=\"https://sites.google.com/view/fgvc6/program?authuser=0\">Schedule of our upcoming FGVC workshop can be found here</a>.</p>\n\n<p>As one of top 3 teams, you are invited to present your solution in our upcoming FGVC workshop at CVPR.</p>\n\n<ol>\n<li>Could you be able to send me 1-2 pages of google slides (or pdf) describing your method? So I can include it in my presentation of this challenge.</li>\n<li>You will also have access to a 4 foot x 4 foot poster board in our FGVC workshop if you want to present your method at the workshop. (If you couldn't make it to the workshop, another option is that you can send your poster to me before June 14. I can help you print it and hang in the board that day)</li>\n</ol>\n\n<p>Let me know if you have any questions!\nThank you!</p>",
      "rawMarkdown": "Hi @dempton  Thank you for sharing your solution! Will you attend CVPR2019 this year? [Schedule of our upcoming FGVC workshop can be found here](https://sites.google.com/view/fgvc6/program?authuser=0).\n\nAs one of top 3 teams, you are invited to present your solution in our upcoming FGVC workshop at CVPR.\n\n1. Could you be able to send me 1-2 pages of google slides (or pdf) describing your method? So I can include it in my presentation of this challenge.\n2. You will also have access to a 4 foot x 4 foot poster board in our FGVC workshop if you want to present your method at the workshop. (If you couldn't make it to the workshop, another option is that you can send your poster to me before June 14. I can help you print it and hang in the board that day)\n\nLet me know if you have any questions!\nThank you!"
    },
    {
      "id": 549801,
      "postDate": "2019-06-11T03:50:02.377Z",
      "content": "<p>Many thanks for sharing your solutions. Looking forward to your source code! Congratulations! :-)</p>",
      "rawMarkdown": "Many thanks for sharing your solutions. Looking forward to your source code! Congratulations! :-)"
    },
    {
      "id": 549773,
      "postDate": "2019-06-11T02:50:17.647Z",
      "content": "<p>Congrats! Pubic &amp; Private 3rd! Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats! Pubic &amp; Private 3rd! Thanks for sharing your solution."
    },
    {
      "id": 549760,
      "postDate": "2019-06-11T02:26:34.850Z",
      "content": "<p>Congrats! Thanks for sharing! </p>\n\n<p>By the way, how much boost can the blend gives？</p>",
      "rawMarkdown": "Congrats! Thanks for sharing! \n\nBy the way, how much boost can the blend gives？\n\n"
    },
    {
      "id": 549722,
      "postDate": "2019-06-11T01:02:26.337Z",
      "content": "<blockquote>\n  <p>We dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.</p>\n</blockquote>\n\n<p>Some images have only 0-12 detected objects so when dropping them you would have an error when submitting </p>\n\n<blockquote>\n  <p>Evaluation Exception: The submitted ids must contain all solution ids.</p>\n</blockquote>",
      "rawMarkdown": "&gt;We dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.\n\nSome images have only 0-12 detected objects so when dropping them you would have an error when submitting \n&gt;Evaluation Exception: The submitted ids must contain all solution ids.",
      "replies": [
        {
          "id": 549745,
          "postDate": "2019-06-11T01:57:40.953Z",
          "content": "<p>This error can be avoided by adding a virtual prediction to these images.</p>\n\n<p>It seems that dropping the predictions for categories 0 - 12 can significantly improve the score, although I think this deviates from the original intention of the competition.</p>",
          "rawMarkdown": "This error can be avoided by adding a virtual prediction to these images.\n\nIt seems that dropping the predictions for categories 0 - 12 can significantly improve the score, although I think this deviates from the original intention of the competition.",
          "votes": 1
        },
        {
          "id": 549747,
          "postDate": "2019-06-11T02:00:29.130Z",
          "content": "<p><a href=\"/raykoo\">@raykoo</a>  what do you mean by virtual?  </p>",
          "rawMarkdown": "@raykoo  what do you mean by virtual?  "
        },
        {
          "id": 549755,
          "postDate": "2019-06-11T02:09:34.867Z",
          "content": "<p>For the image <code>xxxx.jpg</code> that has only 0-12 objects, we can add a prediction like\n<code>{'ImageId': 'xxxx.jpg' , 'EncodedPixels': '1 1', 'ClassId': 1 }</code></p>",
          "rawMarkdown": "For the image `xxxx.jpg` that has only 0-12 objects, we can add a prediction like\n`{'ImageId': 'xxxx.jpg' , 'EncodedPixels': '1 1', 'ClassId': 1 }`",
          "votes": 1
        }
      ]
    },
    {
      "id": 549793,
      "postDate": "2019-06-11T03:31:57.640Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 629439,
      "postDate": "2019-09-18T19:19:09.020Z",
      "content": "<p>Thanks for the explanation.</p>",
      "rawMarkdown": "Thanks for the explanation."
    },
    {
      "id": 549688,
      "postDate": "2019-06-10T23:56:30.600Z",
      "content": "<p>Great job thanks for sharing!</p>",
      "rawMarkdown": "Great job thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 555537,
      "author_name": "kongcraft",
      "author_url": "",
      "post_date": "2019-06-19T03:27:38.517000",
      "content": "<p>Congrats! Looking forward to your source code!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 550756,
      "author_name": "Menglin Jia",
      "author_url": "",
      "post_date": "2019-06-12T02:19:50.233000",
      "content": "<p>Hi <a href=\"/dempton\">@dempton</a>  Thank you for sharing your solution! Will you attend CVPR2019 this year? <a href=\"https://sites.google.com/view/fgvc6/program?authuser=0\">Schedule of our upcoming FGVC workshop can be found here</a>.</p>\n\n<p>As one of top 3 teams, you are invited to present your solution in our upcoming FGVC workshop at CVPR.</p>\n\n<ol>\n<li>Could you be able to send me 1-2 pages of google slides (or pdf) describing your method? So I can include it in my presentation of this challenge.</li>\n<li>You will also have access to a 4 foot x 4 foot poster board in our FGVC workshop if you want to present your method at the workshop. (If you couldn't make it to the workshop, another option is that you can send your poster to me before June 14. I can help you print it and hang in the board that day)</li>\n</ol>\n\n<p>Let me know if you have any questions!\nThank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 549801,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-06-11T03:50:02.377000",
      "content": "<p>Many thanks for sharing your solutions. Looking forward to your source code! Congratulations! :-)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 549773,
      "author_name": "dhaqui the kaggler",
      "author_url": "",
      "post_date": "2019-06-11T02:50:17.647000",
      "content": "<p>Congrats! Pubic &amp; Private 3rd! Thanks for sharing your solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 549760,
      "author_name": "Ke",
      "author_url": "",
      "post_date": "2019-06-11T02:26:34.850000",
      "content": "<p>Congrats! Thanks for sharing! </p>\n\n<p>By the way, how much boost can the blend gives？</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 549722,
      "author_name": "Firas Baba",
      "author_url": "",
      "post_date": "2019-06-11T01:02:26.337000",
      "content": "<blockquote>\n  <p>We dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.</p>\n</blockquote>\n\n<p>Some images have only 0-12 detected objects so when dropping them you would have an error when submitting </p>\n\n<blockquote>\n  <p>Evaluation Exception: The submitted ids must contain all solution ids.</p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 549745,
          "author_name": "Ke",
          "author_url": "",
          "post_date": "2019-06-11T01:57:40.953000",
          "content": "<p>This error can be avoided by adding a virtual prediction to these images.</p>\n\n<p>It seems that dropping the predictions for categories 0 - 12 can significantly improve the score, although I think this deviates from the original intention of the competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 549747,
          "author_name": "Firas Baba",
          "author_url": "",
          "post_date": "2019-06-11T02:00:29.130000",
          "content": "<p><a href=\"/raykoo\">@raykoo</a>  what do you mean by virtual?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 549755,
          "author_name": "Ke",
          "author_url": "",
          "post_date": "2019-06-11T02:09:34.867000",
          "content": "<p>For the image <code>xxxx.jpg</code> that has only 0-12 objects, we can add a prediction like\n<code>{'ImageId': 'xxxx.jpg' , 'EncodedPixels': '1 1', 'ClassId': 1 }</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 549793,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-11T03:31:57.640000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 629439,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2019-09-18T19:19:09.020000",
      "content": "<p>Thanks for the explanation.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 549688,
      "author_name": "Justin Faler",
      "author_url": "",
      "post_date": "2019-06-10T23:56:30.600000",
      "content": "<p>Great job thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "549679": "Hello everybody and congrats to all who finished this competitions in the gold zone. We want to share our solution and describe main ideas. \n\n**Validation metric:**\nWe implemented it according to leaderboard evaluation metric (absolutely the same).\n\n**Data preprocessing:**\nWe used different sizes for trainings:\nmin_size: (800, … 960); max_size &lt;=1600\n\n**Models:**\n- facebook repo - Mask-RCNN x-101\n- mmdetection repo - Hybrid Task Cascade with X-101-64x4d-FPN backbone and c3-c5  DCN\n\n**Blend:**\nWe decided to find a way to blend multi-class instant segmentation.\nFirst of all, we just should match all bounding boxes from all models. Then we can average respective masks. So, we iterate over all boxes and build IOU matrix (where mat[i, j] = iou between i and j boxes). We binarized matrix with IOU &gt; 0.7 (according to validation). Then we build a graph (networks library) and select connected components. Then we consider one connected component as a respective matched bboxes and averaged their masks. Confidence of boxes was calculated as a SUM(confidences all matches boxes) divided by number of models. \n\n**TTA-4:**\nAs we have merger (blend algorithm) we can perform TTA:\n- hflip\n- different sizes (with min_size=800 and min_size=960)\n\nIt gives ~ +0.01 - 0.015\n\n**Second level model:**\nWe want to predict IOU between mask and real mask (according to metric if IOU &lt; 0.5 we have FP, we want to drop all masks with IOU &lt; 0.5 to avoid some FP)\n\n*We extracted the following features from masks:\n*area of predicted mask\n- max, mean, min, median, std of predicted values (for mask)\n- confidence of bboxes\n- left / right side of predicted bbox\n- area of predicted bbox\n- divide one side size to another\n- category_id (categorical feature)\n- max, mean, min, median, std aggregation of all above features by category_id\n\nWe trained LGBM + XGB + CatBoost over this dataset with Labels = IOU (predicted and real-mask on out-of-fold predictions)\n\nAnd obtain ~0.92 roc-auc score (It gives ~ +0.01 according to final metric)\n\n**Attributes**\n\nWe dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.\nRegarding attributes, we tried hard to predict them both inside Mask-RCNN and by a separate model, but didn’t get any improvement. ",
    "555537": "Congrats! Looking forward to your source code!",
    "550756": "Hi @dempton  Thank you for sharing your solution! Will you attend CVPR2019 this year? [Schedule of our upcoming FGVC workshop can be found here](https://sites.google.com/view/fgvc6/program?authuser=0).\n\nAs one of top 3 teams, you are invited to present your solution in our upcoming FGVC workshop at CVPR.\n\n1. Could you be able to send me 1-2 pages of google slides (or pdf) describing your method? So I can include it in my presentation of this challenge.\n2. You will also have access to a 4 foot x 4 foot poster board in our FGVC workshop if you want to present your method at the workshop. (If you couldn't make it to the workshop, another option is that you can send your poster to me before June 14. I can help you print it and hang in the board that day)\n\nLet me know if you have any questions!\nThank you!",
    "549801": "Many thanks for sharing your solutions. Looking forward to your source code! Congratulations! :-)",
    "549773": "Congrats! Pubic &amp; Private 3rd! Thanks for sharing your solution.",
    "549760": "Congrats! Thanks for sharing! \n\nBy the way, how much boost can the blend gives？\n\n",
    "549722": "&gt;We dropped all predictions for categories 0 -- 12, as we didn’t predict attributes and all such predictions would be both FP and FN. Once we drop them, we are penalized for FN only, which gives a big boost.\n\nSome images have only 0-12 detected objects so when dropping them you would have an error when submitting \n&gt;Evaluation Exception: The submitted ids must contain all solution ids.",
    "549793": "",
    "629439": "Thanks for the explanation.",
    "549688": "Great job thanks for sharing!"
  }
}