{
  "id": 154306,
  "title": "1st place solution",
  "url": "/competitions/imaterialist-fashion-2020-fgvc7/writeups/oleg-polosin-1st-place-solution",
  "author_name": "",
  "post_date": "2021-04-27T02:56:17.573Z",
  "votes": 37,
  "comment_count": 21,
  "views": 0,
  "content": "<h1>Model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F853978%2F378129d2e9afb90bbd1f320858b9b73b%2Fimat2020_2.jpg?generation=1591441185889910&amp;alt=media\" alt=\"\"></p>\n<p>I used a single Mask R-CNN model with the <a href=\"https://arxiv.org/abs/1912.05027\" target=\"_blank\">SpineNet-143</a> + FPN backbone and added an extra head to classify attributes. The attributes head was trained with the focal loss. For the augmentations, I used one of the <a href=\"https://arxiv.org/abs/1906.11172\" target=\"_blank\">AutoAugment</a> policies (see details below). No TTAs during the inference.</p>\n<p>All the changes were made on top of the <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/detection\" target=\"_blank\">TPU Object Detection and Segmentation Framework</a>. You can find the <strong>code and weights</strong> for the best model in <strong><a href=\"https://github.com/apls777/kaggle-imaterialist2020-model\" target=\"_blank\">this repo</a></strong>.</p>\n<p>Switching from the ResNet-50 to the SpineNet-96 backbone improved my score by on the LB by ~+0.07 (private), ~+0.05 (public). Not sure though how better the SpineNet-143 backbone was as I trained it with a slightly different configuration.</p>\n<h1>Data</h1>\n<p>I split the training data to training and validation datasets in a way that the validation dataset contains at least 10% of images for each class and each attribute. I ended up with 39932 images in the training dataset and 5691 images in the validation dataset.</p>\n<h1>Training</h1>\n<ul>\n<li>The model was trained on top of pre-trained on the COCO dataset weights.</li>\n<li>It was trained on resolution 1280x1280.</li>\n<li>The attributes head was trained with the focal loss. Switching to the focal loss improved the score by ~+0.012 (private), ~+0.018 (public)</li>\n<li>For the augmentations I used random scaling (0.5 - 2.0) and <a href=\"https://github.com/tensorflow/tpu/blob/2d9507360e3712715c584e2c21c639b39efd6ad1/models/official/detection/utils/autoaugment_utils.py#L126\" target=\"_blank\">v3 policy</a> from the Google's <a href=\"https://arxiv.org/abs/1906.11172\" target=\"_blank\">AutoAugment</a> implementation. I modified the code to make it working with masks as it supports only object detection case at the moment.</li>\n</ul>\n<p>The model was trained with a batch size 64 for 91.6k steps on a v3-8 TPU for ~69 hours. Big thank you to <a href=\"https://www.tensorflow.org/tfrc\" target=\"_blank\">TensorFlow Research Cloud</a> for giving me free access to TPUs, it helped a lot!</p>\n<h1>Predictions</h1>\n<h2>Attributes</h2>\n<p>For the attribute predictions, I used thresholds that maximize F1-score for each individual attribute within a category (so it's 294*46=13524 thresholds, but most of them actually &gt;1 as there are no training examples).</p>\n<h2>Best Predictions</h2>\n<p>I don't think the metric used in this competition was good. Unfortunately, it's not taking into account false-negative predictions at all. That basically means that with just 1 prediction per image the score of 1.0 is still achievable (I first asked about FNs in <a href=\"https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/discussion/141891\" target=\"_blank\">this</a> discussion and later I also contacted the organizer by email to make sure it's not a mistake, but they assured me that the metric is okay.).</p>\n<p>So at first I just used one most confident class prediction per image. But high class confidence does not necessarily mean that the segmentation or attributes are good. So next I tried to compute a score for each prediction as an average of class confidence, category mask AP and category attributes F1-score, then I used it to select 1 best prediction per image. It improved my score on the LB by ~+0.041 (private), ~+0.036 (public).</p>\n<p>In the end, I implemented <a href=\"https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/overview/evaluation\" target=\"_blank\">the metric</a> used in the competition, computed AP for each model prediction (based on the validation dataset) and trained a regression model to predict APs for the Mask R-CNN predictions. Then, as usual, I just used 1 best prediction per image based on the predicted APs. It improved the score on the LB by ~+0.065 (private), ~+0.058 (public). For the regression model, I ended up using a random forest regressor from the scikit-learn package (100 estimators, max depth 8). It was trained on features like category ID, class confidence, mask area, number of predicted attributes, category mask AP, category attributes F1-score, and so on - 13 features in total.</p>",
  "messages": [
    {
      "id": "864343",
      "postDate": "05/28/2020 00:39:16",
      "content": "<h1>Model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F853978%2F378129d2e9afb90bbd1f320858b9b73b%2Fimat2020_2.jpg?generation=1591441185889910&amp;alt=media\" alt=\"\"></p>\n<p>I used a single Mask R-CNN model with the <a href=\"https://arxiv.org/abs/1912.05027\" target=\"_blank\">SpineNet-143</a> + FPN backbone and added an extra head to classify attributes. The attributes head was trained with the focal loss. For the augmentations, I used one of the <a href=\"https://arxiv.org/abs/1906.11172\" target=\"_blank\">AutoAugment</a> policies (see details below). No TTAs during the inference.</p>\n<p>All the changes were made on top of the <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/detection\" target=\"_blank\">TPU Object Detection and Segmentation Framework</a>. You can find the <strong>code and weights</strong> for the best model in <strong><a href=\"https://github.com/apls777/kaggle-imaterialist2020-model\" target=\"_blank\">this repo</a></strong>.</p>\n<p>Switching from the ResNet-50 to the SpineNet-96 backbone improved my score by on the LB by ~+0.07 (private), ~+0.05 (public). Not sure though how better the SpineNet-143 backbone was as I trained it with a slightly different configuration.</p>\n<h1>Data</h1>\n<p>I split the training data to training and validation datasets in a way that the validation dataset contains at least 10% of images for each class and each attribute. I ended up with 39932 images in the training dataset and 5691 images in the validation dataset.</p>\n<h1>Training</h1>\n<ul>\n<li>The model was trained on top of pre-trained on the COCO dataset weights.</li>\n<li>It was trained on resolution 1280x1280.</li>\n<li>The attributes head was trained with the focal loss. Switching to the focal loss improved the score by ~+0.012 (private), ~+0.018 (public)</li>\n<li>For the augmentations I used random scaling (0.5 - 2.0) and <a href=\"https://github.com/tensorflow/tpu/blob/2d9507360e3712715c584e2c21c639b39efd6ad1/models/official/detection/utils/autoaugment_utils.py#L126\" target=\"_blank\">v3 policy</a> from the Google's <a href=\"https://arxiv.org/abs/1906.11172\" target=\"_blank\">AutoAugment</a> implementation. I modified the code to make it working with masks as it supports only object detection case at the moment.</li>\n</ul>\n<p>The model was trained with a batch size 64 for 91.6k steps on a v3-8 TPU for ~69 hours. Big thank you to <a href=\"https://www.tensorflow.org/tfrc\" target=\"_blank\">TensorFlow Research Cloud</a> for giving me free access to TPUs, it helped a lot!</p>\n<h1>Predictions</h1>\n<h2>Attributes</h2>\n<p>For the attribute predictions, I used thresholds that maximize F1-score for each individual attribute within a category (so it's 294*46=13524 thresholds, but most of them actually &gt;1 as there are no training examples).</p>\n<h2>Best Predictions</h2>\n<p>I don't think the metric used in this competition was good. Unfortunately, it's not taking into account false-negative predictions at all. That basically means that with just 1 prediction per image the score of 1.0 is still achievable (I first asked about FNs in <a href=\"https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/discussion/141891\" target=\"_blank\">this</a> discussion and later I also contacted the organizer by email to make sure it's not a mistake, but they assured me that the metric is okay.).</p>\n<p>So at first I just used one most confident class prediction per image. But high class confidence does not necessarily mean that the segmentation or attributes are good. So next I tried to compute a score for each prediction as an average of class confidence, category mask AP and category attributes F1-score, then I used it to select 1 best prediction per image. It improved my score on the LB by ~+0.041 (private), ~+0.036 (public).</p>\n<p>In the end, I implemented <a href=\"https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/overview/evaluation\" target=\"_blank\">the metric</a> used in the competition, computed AP for each model prediction (based on the validation dataset) and trained a regression model to predict APs for the Mask R-CNN predictions. Then, as usual, I just used 1 best prediction per image based on the predicted APs. It improved the score on the LB by ~+0.065 (private), ~+0.058 (public). For the regression model, I ended up using a random forest regressor from the scikit-learn package (100 estimators, max depth 8). It was trained on features like category ID, class confidence, mask area, number of predicted attributes, category mask AP, category attributes F1-score, and so on - 13 features in total.</p>",
      "rawMarkdown": "# Model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F853978%2F378129d2e9afb90bbd1f320858b9b73b%2Fimat2020_2.jpg?generation=1591441185889910&amp;alt=media)\n\nI used a single Mask R-CNN model with the [SpineNet-143](https://arxiv.org/abs/1912.05027) + FPN backbone and added an extra head to classify attributes. The attributes head was trained with the focal loss. For the augmentations, I used one of the [AutoAugment](https://arxiv.org/abs/1906.11172) policies (see details below). No TTAs during the inference.\n\nAll the changes were made on top of the [TPU Object Detection and Segmentation Framework](https://github.com/tensorflow/tpu/tree/master/models/official/detection). You can find the **code and weights** for the best model in **[this repo](https://github.com/apls777/kaggle-imaterialist2020-model)**.\n\nSwitching from the ResNet-50 to the SpineNet-96 backbone improved my score by on the LB by ~+0.07 (private), ~+0.05 (public). Not sure though how better the SpineNet-143 backbone was as I trained it with a slightly different configuration.\n\n# Data\n\nI split the training data to training and validation datasets in a way that the validation dataset contains at least 10% of images for each class and each attribute. I ended up with 39932 images in the training dataset and 5691 images in the validation dataset.\n\n# Training\n\n* The model was trained on top of pre-trained on the COCO dataset weights.\n* It was trained on resolution 1280x1280.\n* The attributes head was trained with the focal loss. Switching to the focal loss improved the score by ~+0.012 (private), ~+0.018 (public)\n* For the augmentations I used random scaling (0.5 - 2.0) and [v3 policy](https://github.com/tensorflow/tpu/blob/2d9507360e3712715c584e2c21c639b39efd6ad1/models/official/detection/utils/autoaugment_utils.py#L126) from the Google's [AutoAugment](https://arxiv.org/abs/1906.11172) implementation. I modified the code to make it working with masks as it supports only object detection case at the moment.\n\nThe model was trained with a batch size 64 for 91.6k steps on a v3-8 TPU for ~69 hours. Big thank you to [TensorFlow Research Cloud](https://www.tensorflow.org/tfrc) for giving me free access to TPUs, it helped a lot!\n\n# Predictions\n\n## Attributes\n\nFor the attribute predictions, I used thresholds that maximize F1-score for each individual attribute within a category (so it's 294*46=13524 thresholds, but most of them actually &gt;1 as there are no training examples).\n\n## Best Predictions\n\nI don't think the metric used in this competition was good. Unfortunately, it's not taking into account false-negative predictions at all. That basically means that with just 1 prediction per image the score of 1.0 is still achievable (I first asked about FNs in [this](https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/discussion/141891) discussion and later I also contacted the organizer by email to make sure it's not a mistake, but they assured me that the metric is okay.).\n\nSo at first I just used one most confident class prediction per image. But high class confidence does not necessarily mean that the segmentation or attributes are good. So next I tried to compute a score for each prediction as an average of class confidence, category mask AP and category attributes F1-score, then I used it to select 1 best prediction per image. It improved my score on the LB by ~+0.041 (private), ~+0.036 (public).\n\nIn the end, I implemented [the metric](https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/overview/evaluation) used in the competition, computed AP for each model prediction (based on the validation dataset) and trained a regression model to predict APs for the Mask R-CNN predictions. Then, as usual, I just used 1 best prediction per image based on the predicted APs. It improved the score on the LB by ~+0.065 (private), ~+0.058 (public). For the regression model, I ended up using a random forest regressor from the scikit-learn package (100 estimators, max depth 8). It was trained on features like category ID, class confidence, mask area, number of predicted attributes, category mask AP, category attributes F1-score, and so on - 13 features in total.",
      "votes": null
    },
    {
      "id": "864573",
      "postDate": "05/28/2020 04:31:43",
      "content": "<p>Congrats on the great finish and nice write up. Did you try anything with the fashionpedia attribute mask rcnn? </p>",
      "rawMarkdown": "Congrats on the great finish and nice write up. Did you try anything with the fashionpedia attribute mask rcnn?",
      "votes": null
    },
    {
      "id": "864903",
      "postDate": "05/28/2020 09:00:21",
      "content": "<p>Congratulations!\nMay I ask you,  Can you share your code on the github?  :)</p>",
      "rawMarkdown": "Congratulations!\nMay I ask you,  Can you share your code on the github?  :)",
      "votes": null
    },
    {
      "id": "865103",
      "postDate": "05/28/2020 11:42:43",
      "content": "<p>Hi，Could I ask some questions? : <br>\n1.How do you set 13524 thresholds？\n2.What is the effect of the  random forest regressor ? <br>\nThanks in advance!   :)</p>",
      "rawMarkdown": "Hi，Could I ask some questions? :  \n1.How do you set 13524 thresholds？\n2.What is the effect of the  random forest regressor ?  \nThanks in advance!   :)",
      "votes": null
    },
    {
      "id": "865768",
      "postDate": "05/28/2020 21:38:36",
      "content": "<p>Thank you! I'm going to release it, but a little bit later, need to groom it a little.</p>",
      "rawMarkdown": "Thank you! I'm going to release it, but a little bit later, need to groom it a little.",
      "votes": null
    },
    {
      "id": "865771",
      "postDate": "05/28/2020 21:40:30",
      "content": "<p>Thank you! No, I actually didn't know they released the model. But anyway it's just an exported model without the actual code, it would be hard to do anything with it.\nDid you try to use it in the competition?</p>",
      "rawMarkdown": "Thank you! No, I actually didn't know they released the model. But anyway it's just an exported model without the actual code, it would be hard to do anything with it.\nDid you try to use it in the competition?",
      "votes": null
    },
    {
      "id": "865782",
      "postDate": "05/28/2020 21:56:00",
      "content": "<p>&gt; 1.How do you set 13524 thresholds？</p>\n\n<p>I take prediction confidences and ground truth labels for each category and each attribute, compute precision-recall values at different thresholds and take one that corresponds to the best F1-score (see <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_recall_curve.html\">sklearn.metrics.precision_recall_curve</a>).</p>\n\n<p>&gt; 2.What is the effect of the random forest regressor ?</p>\n\n<p>I use it to estimate Average Precision for each of the model predictions. Then I use the best predictions for the submission.</p>\n\n<p>You can compute AP for each prediction on the validation dataset using the following function:\n```\ndef get_precision_value(mask_iou: float, f1_score: float) -&gt; float:\n    thresholds = [0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95]</p>\n\n<pre><code>tps = 0\nfor ti in thresholds:\n    for tf in thresholds:\n        if (mask_iou &gt; ti) and (f1_score &gt; tf):\n            tps += 1\n\nreturn tps / (len(thresholds) ** 2)\n</code></pre>\n\n<p>```\nThen you can use computed APs together with predictions' features (i.e. category ID, mask area, number of predicted attributes) to train a random forest regressor model (see <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html\">sklearn.ensemble.RandomForestRegressor</a>).</p>",
      "rawMarkdown": "&gt; 1.How do you set 13524 thresholds？\n\nI take prediction confidences and ground truth labels for each category and each attribute, compute precision-recall values at different thresholds and take one that corresponds to the best F1-score (see [sklearn.metrics.precision\\_recall\\_curve](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_recall_curve.html)).\n\n&gt; 2.What is the effect of the random forest regressor ?\n\nI use it to estimate Average Precision for each of the model predictions. Then I use the best predictions for the submission.\n\nYou can compute AP for each prediction on the validation dataset using the following function:\n```\ndef get_precision_value(mask_iou: float, f1_score: float) -&gt; float:\n    thresholds = [0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95]\n\n    tps = 0\n    for ti in thresholds:\n        for tf in thresholds:\n            if (mask_iou &gt; ti) and (f1_score &gt; tf):\n                tps += 1\n\n    return tps / (len(thresholds) ** 2)\n```\nThen you can use computed APs together with predictions' features (i.e. category ID, mask area, number of predicted attributes) to train a random forest regressor model (see [sklearn.ensemble.RandomForestRegressor](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html)).",
      "votes": null
    },
    {
      "id": "865816",
      "postDate": "05/28/2020 22:54:58",
      "content": "<p>Congratulations Oleg! And thanks for sharing you methodology details. Looking forward to see your code too!</p>",
      "rawMarkdown": "Congratulations Oleg! And thanks for sharing you methodology details. Looking forward to see your code too!",
      "votes": null
    },
    {
      "id": "865838",
      "postDate": "05/28/2020 23:36:58",
      "content": "<p>I only found it last week by chance while looking for better ways to add classifier heads onto mask rcnn. </p>\n\n<p>I was a little concerned to make a submission with it but as far as I could tell it was not trained on anything against the rules here so I gave it a try. I made one submission using only their demo inference script and it scored about 0.18 with no post-processing at all and a submission file around 60k rows. Not bad actually, but I ran out of time to do much with it and the masks from my mmdet and detectron2 models were looking much better anyway. </p>\n\n<p>Also, thanks for spotty, it’s really good. </p>",
      "rawMarkdown": "I only found it last week by chance while looking for better ways to add classifier heads onto mask rcnn. \n\nI was a little concerned to make a submission with it but as far as I could tell it was not trained on anything against the rules here so I gave it a try. I made one submission using only their demo inference script and it scored about 0.18 with no post-processing at all and a submission file around 60k rows. Not bad actually, but I ran out of time to do much with it and the masks from my mmdet and detectron2 models were looking much better anyway. \n\nAlso, thanks for spotty, it’s really good.",
      "votes": null
    },
    {
      "id": "865892",
      "postDate": "05/29/2020 00:53:29",
      "content": "<p>You're welcome! Thank you for using it! :)</p>",
      "rawMarkdown": "You're welcome! Thank you for using it! :)",
      "votes": null
    },
    {
      "id": "865893",
      "postDate": "05/29/2020 00:54:02",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "865982",
      "postDate": "05/29/2020 02:47:33",
      "content": "<p>Please tell us, after you push it to the github~  :)</p>",
      "rawMarkdown": "Please tell us, after you push it to the github~  :)",
      "votes": null
    },
    {
      "id": "865984",
      "postDate": "05/29/2020 02:48:44",
      "content": "<p>Thanks for your patient explanation， ：)</p>",
      "rawMarkdown": "Thanks for your patient explanation， ：)",
      "votes": null
    },
    {
      "id": "878118",
      "postDate": "06/08/2020 09:02:37",
      "content": "<p>Hi~~\nCould you share your code?</p>",
      "rawMarkdown": "Hi~~\nCould you share your code?",
      "votes": null
    },
    {
      "id": "902148",
      "postDate": "06/26/2020 00:44:13",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!",
      "votes": null
    },
    {
      "id": "911271",
      "postDate": "07/01/2020 16:29:43",
      "content": "<p>Congratulations! Could you open-source your code for us enthusiasts to learn from it?</p>",
      "rawMarkdown": "Congratulations! Could you open-source your code for us enthusiasts to learn from it?",
      "votes": null
    },
    {
      "id": "911472",
      "postDate": "07/01/2020 18:50:22",
      "content": "<p>Thank you, Krish! I'll come back to the code and will open-source it right after I finish with the next Spotty release.</p>",
      "rawMarkdown": "Thank you, Krish! I'll come back to the code and will open-source it right after I finish with the next Spotty release.",
      "votes": null
    },
    {
      "id": "914245",
      "postDate": "07/03/2020 17:39:17",
      "content": "<p>Hi Oleg, your method is brilliant. The insight about false negative is amazing. It would be great if beginners like me can try to understand your code. Hope you open-source it soon.</p>",
      "rawMarkdown": "Hi Oleg, your method is brilliant. The insight about false negative is amazing. It would be great if beginners like me can try to understand your code. Hope you open-source it soon.",
      "votes": null
    },
    {
      "id": "914655",
      "postDate": "07/04/2020 06:02:49",
      "content": "<p>Please tell us, after you push it to the github</p>",
      "rawMarkdown": "Please tell us, after you push it to the github",
      "votes": null
    },
    {
      "id": "918747",
      "postDate": "07/07/2020 13:07:42",
      "content": "<p>Thank you so much! Looking forward to it!</p>",
      "rawMarkdown": "Thank you so much! Looking forward to it!",
      "votes": null
    },
    {
      "id": "919000",
      "postDate": "07/07/2020 16:33:53",
      "content": "<p>Nice. Congratulations</p>",
      "rawMarkdown": "Nice. Congratulations",
      "votes": null
    },
    {
      "id": "1574388",
      "postDate": "11/07/2021 13:58:29",
      "content": "<p>Congrats Oleg!<br>\nyour way of adding attribute heads to Mask-RCNN is really helpful. <br>\nI am facing one issue while training \"mask_rcnn\" model(it's working fine for \"retinanet\", but unfortunately that does not have an attribute head). I have provided a snippet of my traceback below, please let me know if you ever encounter that issue. </p>\n<pre><code>Parsing Inputs...\nTraceback (most recent call last):\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1365, in _do_call\n    return fn(*args)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1350, in _run_fn\n    target_list, run_metadata)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1443, in _call_tf_sessionrun\n    run_metadata)\ntensorflow.python.framework.errors_impl.InvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument: {{function_node __inference_Dataset_map_&lt;class 'dataloader.maskrcnn_parser.Parser'&gt;_2124}} indices[7] = 7 is not in [0, 0)\n     [[{{node parser/GatherV2_2}}]]\n     [[MultiDeviceIteratorGetNextFromShard]]\n     [[RemoteCall]]\n     [[IteratorGetNext]]\n     [[ArithmeticOptimizer/AddOpsRewrite_add_5/_9827]]\n  (1) Invalid argument: {{function_node __inference_Dataset_map_&lt;class 'dataloader.maskrcnn_parser.Parser'&gt;_2124}} indices[7] = 7 is not in [0, 0)\n     [[{{node parser/GatherV2_2}}]]\n     [[MultiDeviceIteratorGetNextFromShard]]\n     [[RemoteCall]]\n     [[IteratorGetNext]]\n</code></pre>\n<p>just an additional note, I am using fashionpedia data and converted it to tfrecord using <a href=\"https://github.com/tensorflow/tpu/blob/master/tools/datasets/create_coco_tf_record.py\" target=\"_blank\">this script</a>. I am not sure it's the tfrecord causing this issue or the \"maskrcnn_parser\" (because the same tfrecord file working fine for \"retinanet_parser\").</p>",
      "rawMarkdown": "Congrats Oleg!\nyour way of adding attribute heads to Mask-RCNN is really helpful. \nI am facing one issue while training \"mask_rcnn\" model(it's working fine for \"retinanet\", but unfortunately that does not have an attribute head). I have provided a snippet of my traceback below, please let me know if you ever encounter that issue. \n```\nParsing Inputs...\nTraceback (most recent call last):\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1365, in _do_call\n    return fn(*args)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1350, in _run_fn\n    target_list, run_metadata)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1443, in _call_tf_sessionrun\n    run_metadata)\ntensorflow.python.framework.errors_impl.InvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument: {{function_node __inference_Dataset_map_<class 'dataloader.maskrcnn_parser.Parser'>_2124}} indices[7] = 7 is not in [0, 0)\n\t [[{{node parser/GatherV2_2}}]]\n\t [[MultiDeviceIteratorGetNextFromShard]]\n\t [[RemoteCall]]\n\t [[IteratorGetNext]]\n\t [[ArithmeticOptimizer/AddOpsRewrite_add_5/_9827]]\n  (1) Invalid argument: {{function_node __inference_Dataset_map_<class 'dataloader.maskrcnn_parser.Parser'>_2124}} indices[7] = 7 is not in [0, 0)\n\t [[{{node parser/GatherV2_2}}]]\n\t [[MultiDeviceIteratorGetNextFromShard]]\n\t [[RemoteCall]]\n\t [[IteratorGetNext]]\n```\njust an additional note, I am using fashionpedia data and converted it to tfrecord using [this script](https://github.com/tensorflow/tpu/blob/master/tools/datasets/create_coco_tf_record.py). I am not sure it's the tfrecord causing this issue or the \"maskrcnn_parser\" (because the same tfrecord file working fine for \"retinanet_parser\").",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1574388,
      "author_name": "janmejaya",
      "author_url": "",
      "post_date": "11/07/2021 13:58:29",
      "content": "<p>Congrats Oleg!<br>\nyour way of adding attribute heads to Mask-RCNN is really helpful. <br>\nI am facing one issue while training \"mask_rcnn\" model(it's working fine for \"retinanet\", but unfortunately that does not have an attribute head). I have provided a snippet of my traceback below, please let me know if you ever encounter that issue. </p>\n<pre><code>Parsing Inputs...\nTraceback (most recent call last):\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1365, in _do_call\n    return fn(*args)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1350, in _run_fn\n    target_list, run_metadata)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1443, in _call_tf_sessionrun\n    run_metadata)\ntensorflow.python.framework.errors_impl.InvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument: {{function_node __inference_Dataset_map_&lt;class 'dataloader.maskrcnn_parser.Parser'&gt;_2124}} indices[7] = 7 is not in [0, 0)\n     [[{{node parser/GatherV2_2}}]]\n     [[MultiDeviceIteratorGetNextFromShard]]\n     [[RemoteCall]]\n     [[IteratorGetNext]]\n     [[ArithmeticOptimizer/AddOpsRewrite_add_5/_9827]]\n  (1) Invalid argument: {{function_node __inference_Dataset_map_&lt;class 'dataloader.maskrcnn_parser.Parser'&gt;_2124}} indices[7] = 7 is not in [0, 0)\n     [[{{node parser/GatherV2_2}}]]\n     [[MultiDeviceIteratorGetNextFromShard]]\n     [[RemoteCall]]\n     [[IteratorGetNext]]\n</code></pre>\n<p>just an additional note, I am using fashionpedia data and converted it to tfrecord using <a href=\"https://github.com/tensorflow/tpu/blob/master/tools/datasets/create_coco_tf_record.py\" target=\"_blank\">this script</a>. I am not sure it's the tfrecord causing this issue or the \"maskrcnn_parser\" (because the same tfrecord file working fine for \"retinanet_parser\").</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 864573,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "05/28/2020 04:31:43",
      "content": "<p>Congrats on the great finish and nice write up. Did you try anything with the fashionpedia attribute mask rcnn? </p>",
      "votes": null,
      "replies": [
        {
          "id": 865771,
          "author_name": "polosin",
          "author_url": "",
          "post_date": "05/28/2020 21:40:30",
          "content": "<p>Thank you! No, I actually didn't know they released the model. But anyway it's just an exported model without the actual code, it would be hard to do anything with it.\nDid you try to use it in the competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865838,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "05/28/2020 23:36:58",
          "content": "<p>I only found it last week by chance while looking for better ways to add classifier heads onto mask rcnn. </p>\n\n<p>I was a little concerned to make a submission with it but as far as I could tell it was not trained on anything against the rules here so I gave it a try. I made one submission using only their demo inference script and it scored about 0.18 with no post-processing at all and a submission file around 60k rows. Not bad actually, but I ran out of time to do much with it and the masks from my mmdet and detectron2 models were looking much better anyway. </p>\n\n<p>Also, thanks for spotty, it’s really good. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865892,
          "author_name": "polosin",
          "author_url": "",
          "post_date": "05/29/2020 00:53:29",
          "content": "<p>You're welcome! Thank you for using it! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 864903,
      "author_name": "fanhaobei",
      "author_url": "",
      "post_date": "05/28/2020 09:00:21",
      "content": "<p>Congratulations!\nMay I ask you,  Can you share your code on the github?  :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 865768,
          "author_name": "polosin",
          "author_url": "",
          "post_date": "05/28/2020 21:38:36",
          "content": "<p>Thank you! I'm going to release it, but a little bit later, need to groom it a little.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865982,
          "author_name": "fanhaobei",
          "author_url": "",
          "post_date": "05/29/2020 02:47:33",
          "content": "<p>Please tell us, after you push it to the github~  :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914655,
          "author_name": "ajax0564",
          "author_url": "",
          "post_date": "07/04/2020 06:02:49",
          "content": "<p>Please tell us, after you push it to the github</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 865103,
      "author_name": "fanhaobei",
      "author_url": "",
      "post_date": "05/28/2020 11:42:43",
      "content": "<p>Hi，Could I ask some questions? : <br>\n1.How do you set 13524 thresholds？\n2.What is the effect of the  random forest regressor ? <br>\nThanks in advance!   :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 865782,
          "author_name": "polosin",
          "author_url": "",
          "post_date": "05/28/2020 21:56:00",
          "content": "<p>&gt; 1.How do you set 13524 thresholds？</p>\n\n<p>I take prediction confidences and ground truth labels for each category and each attribute, compute precision-recall values at different thresholds and take one that corresponds to the best F1-score (see <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_recall_curve.html\">sklearn.metrics.precision_recall_curve</a>).</p>\n\n<p>&gt; 2.What is the effect of the random forest regressor ?</p>\n\n<p>I use it to estimate Average Precision for each of the model predictions. Then I use the best predictions for the submission.</p>\n\n<p>You can compute AP for each prediction on the validation dataset using the following function:\n```\ndef get_precision_value(mask_iou: float, f1_score: float) -&gt; float:\n    thresholds = [0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95]</p>\n\n<pre><code>tps = 0\nfor ti in thresholds:\n    for tf in thresholds:\n        if (mask_iou &gt; ti) and (f1_score &gt; tf):\n            tps += 1\n\nreturn tps / (len(thresholds) ** 2)\n</code></pre>\n\n<p>```\nThen you can use computed APs together with predictions' features (i.e. category ID, mask area, number of predicted attributes) to train a random forest regressor model (see <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html\">sklearn.ensemble.RandomForestRegressor</a>).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865984,
          "author_name": "fanhaobei",
          "author_url": "",
          "post_date": "05/29/2020 02:48:44",
          "content": "<p>Thanks for your patient explanation， ：)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 878118,
          "author_name": "fanhaobei",
          "author_url": "",
          "post_date": "06/08/2020 09:02:37",
          "content": "<p>Hi~~\nCould you share your code?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 865816,
      "author_name": "mihirmavalankar",
      "author_url": "",
      "post_date": "05/28/2020 22:54:58",
      "content": "<p>Congratulations Oleg! And thanks for sharing you methodology details. Looking forward to see your code too!</p>",
      "votes": null,
      "replies": [
        {
          "id": 865893,
          "author_name": "polosin",
          "author_url": "",
          "post_date": "05/29/2020 00:54:02",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 902148,
      "author_name": "rashidulhasanhridoy",
      "author_url": "",
      "post_date": "06/26/2020 00:44:13",
      "content": "<p>Great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 911271,
      "author_name": "krisg04",
      "author_url": "",
      "post_date": "07/01/2020 16:29:43",
      "content": "<p>Congratulations! Could you open-source your code for us enthusiasts to learn from it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 911472,
          "author_name": "polosin",
          "author_url": "",
          "post_date": "07/01/2020 18:50:22",
          "content": "<p>Thank you, Krish! I'll come back to the code and will open-source it right after I finish with the next Spotty release.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918747,
          "author_name": "krisg04",
          "author_url": "",
          "post_date": "07/07/2020 13:07:42",
          "content": "<p>Thank you so much! Looking forward to it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 914245,
      "author_name": "deeeeeeeep",
      "author_url": "",
      "post_date": "07/03/2020 17:39:17",
      "content": "<p>Hi Oleg, your method is brilliant. The insight about false negative is amazing. It would be great if beginners like me can try to understand your code. Hope you open-source it soon.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 919000,
      "author_name": "jqsfire125",
      "author_url": "",
      "post_date": "07/07/2020 16:33:53",
      "content": "<p>Nice. Congratulations</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864343": "# Model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F853978%2F378129d2e9afb90bbd1f320858b9b73b%2Fimat2020_2.jpg?generation=1591441185889910&amp;alt=media)\n\nI used a single Mask R-CNN model with the [SpineNet-143](https://arxiv.org/abs/1912.05027) + FPN backbone and added an extra head to classify attributes. The attributes head was trained with the focal loss. For the augmentations, I used one of the [AutoAugment](https://arxiv.org/abs/1906.11172) policies (see details below). No TTAs during the inference.\n\nAll the changes were made on top of the [TPU Object Detection and Segmentation Framework](https://github.com/tensorflow/tpu/tree/master/models/official/detection). You can find the **code and weights** for the best model in **[this repo](https://github.com/apls777/kaggle-imaterialist2020-model)**.\n\nSwitching from the ResNet-50 to the SpineNet-96 backbone improved my score by on the LB by ~+0.07 (private), ~+0.05 (public). Not sure though how better the SpineNet-143 backbone was as I trained it with a slightly different configuration.\n\n# Data\n\nI split the training data to training and validation datasets in a way that the validation dataset contains at least 10% of images for each class and each attribute. I ended up with 39932 images in the training dataset and 5691 images in the validation dataset.\n\n# Training\n\n* The model was trained on top of pre-trained on the COCO dataset weights.\n* It was trained on resolution 1280x1280.\n* The attributes head was trained with the focal loss. Switching to the focal loss improved the score by ~+0.012 (private), ~+0.018 (public)\n* For the augmentations I used random scaling (0.5 - 2.0) and [v3 policy](https://github.com/tensorflow/tpu/blob/2d9507360e3712715c584e2c21c639b39efd6ad1/models/official/detection/utils/autoaugment_utils.py#L126) from the Google's [AutoAugment](https://arxiv.org/abs/1906.11172) implementation. I modified the code to make it working with masks as it supports only object detection case at the moment.\n\nThe model was trained with a batch size 64 for 91.6k steps on a v3-8 TPU for ~69 hours. Big thank you to [TensorFlow Research Cloud](https://www.tensorflow.org/tfrc) for giving me free access to TPUs, it helped a lot!\n\n# Predictions\n\n## Attributes\n\nFor the attribute predictions, I used thresholds that maximize F1-score for each individual attribute within a category (so it's 294*46=13524 thresholds, but most of them actually &gt;1 as there are no training examples).\n\n## Best Predictions\n\nI don't think the metric used in this competition was good. Unfortunately, it's not taking into account false-negative predictions at all. That basically means that with just 1 prediction per image the score of 1.0 is still achievable (I first asked about FNs in [this](https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/discussion/141891) discussion and later I also contacted the organizer by email to make sure it's not a mistake, but they assured me that the metric is okay.).\n\nSo at first I just used one most confident class prediction per image. But high class confidence does not necessarily mean that the segmentation or attributes are good. So next I tried to compute a score for each prediction as an average of class confidence, category mask AP and category attributes F1-score, then I used it to select 1 best prediction per image. It improved my score on the LB by ~+0.041 (private), ~+0.036 (public).\n\nIn the end, I implemented [the metric](https://www.kaggle.com/c/imaterialist-fashion-2020-fgvc7/overview/evaluation) used in the competition, computed AP for each model prediction (based on the validation dataset) and trained a regression model to predict APs for the Mask R-CNN predictions. Then, as usual, I just used 1 best prediction per image based on the predicted APs. It improved the score on the LB by ~+0.065 (private), ~+0.058 (public). For the regression model, I ended up using a random forest regressor from the scikit-learn package (100 estimators, max depth 8). It was trained on features like category ID, class confidence, mask area, number of predicted attributes, category mask AP, category attributes F1-score, and so on - 13 features in total.",
    "864573": "Congrats on the great finish and nice write up. Did you try anything with the fashionpedia attribute mask rcnn?",
    "864903": "Congratulations!\nMay I ask you,  Can you share your code on the github?  :)",
    "865103": "Hi，Could I ask some questions? :  \n1.How do you set 13524 thresholds？\n2.What is the effect of the  random forest regressor ?  \nThanks in advance!   :)",
    "865768": "Thank you! I'm going to release it, but a little bit later, need to groom it a little.",
    "865771": "Thank you! No, I actually didn't know they released the model. But anyway it's just an exported model without the actual code, it would be hard to do anything with it.\nDid you try to use it in the competition?",
    "865782": "&gt; 1.How do you set 13524 thresholds？\n\nI take prediction confidences and ground truth labels for each category and each attribute, compute precision-recall values at different thresholds and take one that corresponds to the best F1-score (see [sklearn.metrics.precision\\_recall\\_curve](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_recall_curve.html)).\n\n&gt; 2.What is the effect of the random forest regressor ?\n\nI use it to estimate Average Precision for each of the model predictions. Then I use the best predictions for the submission.\n\nYou can compute AP for each prediction on the validation dataset using the following function:\n```\ndef get_precision_value(mask_iou: float, f1_score: float) -&gt; float:\n    thresholds = [0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95]\n\n    tps = 0\n    for ti in thresholds:\n        for tf in thresholds:\n            if (mask_iou &gt; ti) and (f1_score &gt; tf):\n                tps += 1\n\n    return tps / (len(thresholds) ** 2)\n```\nThen you can use computed APs together with predictions' features (i.e. category ID, mask area, number of predicted attributes) to train a random forest regressor model (see [sklearn.ensemble.RandomForestRegressor](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html)).",
    "865816": "Congratulations Oleg! And thanks for sharing you methodology details. Looking forward to see your code too!",
    "865838": "I only found it last week by chance while looking for better ways to add classifier heads onto mask rcnn. \n\nI was a little concerned to make a submission with it but as far as I could tell it was not trained on anything against the rules here so I gave it a try. I made one submission using only their demo inference script and it scored about 0.18 with no post-processing at all and a submission file around 60k rows. Not bad actually, but I ran out of time to do much with it and the masks from my mmdet and detectron2 models were looking much better anyway. \n\nAlso, thanks for spotty, it’s really good.",
    "865892": "You're welcome! Thank you for using it! :)",
    "865893": "Thank you!",
    "865982": "Please tell us, after you push it to the github~  :)",
    "865984": "Thanks for your patient explanation， ：)",
    "878118": "Hi~~\nCould you share your code?",
    "902148": "Great!",
    "911271": "Congratulations! Could you open-source your code for us enthusiasts to learn from it?",
    "911472": "Thank you, Krish! I'll come back to the code and will open-source it right after I finish with the next Spotty release.",
    "914245": "Hi Oleg, your method is brilliant. The insight about false negative is amazing. It would be great if beginners like me can try to understand your code. Hope you open-source it soon.",
    "914655": "Please tell us, after you push it to the github",
    "918747": "Thank you so much! Looking forward to it!",
    "919000": "Nice. Congratulations",
    "1574388": "Congrats Oleg!\nyour way of adding attribute heads to Mask-RCNN is really helpful. \nI am facing one issue while training \"mask_rcnn\" model(it's working fine for \"retinanet\", but unfortunately that does not have an attribute head). I have provided a snippet of my traceback below, please let me know if you ever encounter that issue. \n```\nParsing Inputs...\nTraceback (most recent call last):\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1365, in _do_call\n    return fn(*args)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1350, in _run_fn\n    target_list, run_metadata)\n  File \"/home/impact/.virtualenv/attribute-smart-tpu/lib/python3.6/site-packages/tensorflow_core/python/client/session.py\", line 1443, in _call_tf_sessionrun\n    run_metadata)\ntensorflow.python.framework.errors_impl.InvalidArgumentError: 2 root error(s) found.\n  (0) Invalid argument: {{function_node __inference_Dataset_map_<class 'dataloader.maskrcnn_parser.Parser'>_2124}} indices[7] = 7 is not in [0, 0)\n\t [[{{node parser/GatherV2_2}}]]\n\t [[MultiDeviceIteratorGetNextFromShard]]\n\t [[RemoteCall]]\n\t [[IteratorGetNext]]\n\t [[ArithmeticOptimizer/AddOpsRewrite_add_5/_9827]]\n  (1) Invalid argument: {{function_node __inference_Dataset_map_<class 'dataloader.maskrcnn_parser.Parser'>_2124}} indices[7] = 7 is not in [0, 0)\n\t [[{{node parser/GatherV2_2}}]]\n\t [[MultiDeviceIteratorGetNextFromShard]]\n\t [[RemoteCall]]\n\t [[IteratorGetNext]]\n```\njust an additional note, I am using fashionpedia data and converted it to tfrecord using [this script](https://github.com/tensorflow/tpu/blob/master/tools/datasets/create_coco_tf_record.py). I am not sure it's the tfrecord causing this issue or the \"maskrcnn_parser\" (because the same tfrecord file working fine for \"retinanet_parser\")."
  },
  "source": "meta"
}