{
  "id": 287774,
  "title": "Let's talk about TTA and Ensembling...",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/287774",
  "author_name": "",
  "post_date": "2021-11-15T14:05:16.488447Z",
  "votes": 37,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi everyone. I was experimenting with different tta/ensembling schemes on the weekends. However, nothing works reliably for me. My general TTA/Ensembling scheme looks like this:</p>\n<ol>\n<li>Inference, get T*N sets of predictions from T augmentations and N models (currently T=4 and N=1, just have a simple baseline Mask-RCNN model right now lol).</li>\n<li>Filter out small masks from those predictions.</li>\n<li>Combine predictions (using methods below).</li>\n<li>Filter out overlapping masks, giving priority to predictions with higher score.</li>\n</ol>\n<p>Here are the methods I've tried:</p>\n<ol>\n<li>Take unions of masks with large overlap and discard masks with small overlap (as suggested by winners of DS Bowl 2018, e.g. <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326\" target=\"_blank\">here</a>. I tried different thresholds for that. No visible difference in local CV and on LB.</li>\n<li>I tried the following scheme as well (from <a href=\"https://www.kaggle.com/c/imaterialist-fashion-2019-FGVC6/discussion/95247\" target=\"_blank\">1st place in iMaterialist competition</a>):<br>\n<img src=\"https://raw.githubusercontent.com/amirassov/kaggle-imaterialist/master/figures/tta.png\" alt=\"TTA scheme\"><br>\nbut the metrics in local CV only got worse (maybe because of hard NMS, perhaps I should try WBF?)</li>\n</ol>\n<p>I'm wondering if anyone has managed to reliably get improvements from ensembling and/or TTA?</p>",
  "messages": [
    {
      "id": "1583062",
      "postDate": "11/15/2021 14:05:16",
      "content": "<p>Hi everyone. I was experimenting with different tta/ensembling schemes on the weekends. However, nothing works reliably for me. My general TTA/Ensembling scheme looks like this:</p>\n<ol>\n<li>Inference, get T*N sets of predictions from T augmentations and N models (currently T=4 and N=1, just have a simple baseline Mask-RCNN model right now lol).</li>\n<li>Filter out small masks from those predictions.</li>\n<li>Combine predictions (using methods below).</li>\n<li>Filter out overlapping masks, giving priority to predictions with higher score.</li>\n</ol>\n<p>Here are the methods I've tried:</p>\n<ol>\n<li>Take unions of masks with large overlap and discard masks with small overlap (as suggested by winners of DS Bowl 2018, e.g. <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326\" target=\"_blank\">here</a>. I tried different thresholds for that. No visible difference in local CV and on LB.</li>\n<li>I tried the following scheme as well (from <a href=\"https://www.kaggle.com/c/imaterialist-fashion-2019-FGVC6/discussion/95247\" target=\"_blank\">1st place in iMaterialist competition</a>):<br>\n<img src=\"https://raw.githubusercontent.com/amirassov/kaggle-imaterialist/master/figures/tta.png\" alt=\"TTA scheme\"><br>\nbut the metrics in local CV only got worse (maybe because of hard NMS, perhaps I should try WBF?)</li>\n</ol>\n<p>I'm wondering if anyone has managed to reliably get improvements from ensembling and/or TTA?</p>",
      "rawMarkdown": "Hi everyone. I was experimenting with different tta/ensembling schemes on the weekends. However, nothing works reliably for me. My general TTA/Ensembling scheme looks like this:\n\n1. Inference, get T*N sets of predictions from T augmentations and N models (currently T=4 and N=1, just have a simple baseline Mask-RCNN model right now lol).\n2. Filter out small masks from those predictions.\n3. Combine predictions (using methods below).\n4. Filter out overlapping masks, giving priority to predictions with higher score.\n\nHere are the methods I've tried:\n\n1. Take unions of masks with large overlap and discard masks with small overlap (as suggested by winners of DS Bowl 2018, e.g. [here](https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326). I tried different thresholds for that. No visible difference in local CV and on LB.\n2. I tried the following scheme as well (from [1st place in iMaterialist competition](https://www.kaggle.com/c/imaterialist-fashion-2019-FGVC6/discussion/95247)):\n    ![TTA scheme](https://raw.githubusercontent.com/amirassov/kaggle-imaterialist/master/figures/tta.png)\n   but the metrics in local CV only got worse (maybe because of hard NMS, perhaps I should try WBF?)\n\nI'm wondering if anyone has managed to reliably get improvements from ensembling and/or TTA?",
      "votes": null
    },
    {
      "id": "1583204",
      "postDate": "11/15/2021 16:12:36",
      "content": "<p>I have faced the same problem. The method I've used is the union of the masks (with horizontal flip as you mentioned) and have not seen any improvement in my scores. </p>",
      "rawMarkdown": "I have faced the same problem. The method I've used is the union of the masks (with horizontal flip as you mentioned) and have not seen any improvement in my scores.",
      "votes": null
    },
    {
      "id": "1583985",
      "postDate": "11/16/2021 07:28:04",
      "content": "<p>Can I know, What Data-Augmentation transformations have you performed during test time.</p>",
      "rawMarkdown": "Can I know, What Data-Augmentation transformations have you performed during test time.",
      "votes": null
    },
    {
      "id": "1584010",
      "postDate": "11/16/2021 07:55:21",
      "content": "<p>As mentioned in <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326:\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326:</a><br>\nCombined predictions on actual image and horizontally flipped image: took unions of masks with maximum overlap and removed false positive masks with small overlap.</p>",
      "rawMarkdown": "As mentioned in https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326:\nCombined predictions on actual image and horizontally flipped image: took unions of masks with maximum overlap and removed false positive masks with small overlap.",
      "votes": null
    },
    {
      "id": "1584142",
      "postDate": "11/16/2021 10:02:31",
      "content": "<p>Excuse me, what is TTA?</p>",
      "rawMarkdown": "Excuse me, what is TTA?",
      "votes": null
    },
    {
      "id": "1584144",
      "postDate": "11/16/2021 10:05:09",
      "content": "<p>I learn what TTA is.<br>\nTTA is test time augmentation.<br>\nI found the good source for TTA.<br>\n<a href=\"https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it\" target=\"_blank\">https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it</a></p>\n<p>According to this article, TTA for mask rcnn has good effect on the accuracy.</p>\n<p><a href=\"https://www.researchgate.net/publication/340031950_Test-time_augmentation_for_deep_learning-based_cell_segmentation_on_microscopy_images\" target=\"_blank\">https://www.researchgate.net/publication/340031950_Test-time_augmentation_for_deep_learning-based_cell_segmentation_on_microscopy_images</a></p>",
      "rawMarkdown": "I learn what TTA is.\nTTA is test time augmentation.\nI found the good source for TTA.\nhttps://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it\n\nAccording to this article, TTA for mask rcnn has good effect on the accuracy.\n\nhttps://www.researchgate.net/publication/340031950_Test-time_augmentation_for_deep_learning-based_cell_segmentation_on_microscopy_images",
      "votes": null
    },
    {
      "id": "1584250",
      "postDate": "11/16/2021 11:43:01",
      "content": "<p>I could improve it from 0.299 to 0.300 with ensemble. I will explore more.</p>",
      "rawMarkdown": "I could improve it from 0.299 to 0.300 with ensemble. I will explore more.",
      "votes": null
    },
    {
      "id": "1584260",
      "postDate": "11/16/2021 11:52:46",
      "content": "<p>Detectron's implementation is averaging the augmented mask.</p>\n<pre><code>def _reduce_pred_masks(self, outputs, tfms):\n        # Should apply inverse transforms on masks.\n        # We assume only resize &amp; flip are used. pred_masks is a scale-invariant\n        # representation, so we handle flip specially\n        for output, tfm in zip(outputs, tfms):\n            if any(isinstance(t, HFlipTransform) for t in tfm.transforms):\n                output.pred_masks = output.pred_masks.flip(dims=[3])\n        all_pred_masks = torch.stack([o.pred_masks for o in outputs], dim=0)\n        avg_pred_masks = torch.mean(all_pred_masks, dim=0)\n        return avg_pred_masks\n</code></pre>\n<p><a href=\"https://github.com/facebookresearch/detectron2/blob/cbbc1ce26473cb2a5cc8f58e8ada9ae14cb41052/detectron2/modeling/test_time_augmentation.py#L262\" target=\"_blank\">https://github.com/facebookresearch/detectron2/blob/cbbc1ce26473cb2a5cc8f58e8ada9ae14cb41052/detectron2/modeling/test_time_augmentation.py#L262</a></p>",
      "rawMarkdown": "Detectron's implementation is averaging the augmented mask.\n\n```\ndef _reduce_pred_masks(self, outputs, tfms):\n        # Should apply inverse transforms on masks.\n        # We assume only resize & flip are used. pred_masks is a scale-invariant\n        # representation, so we handle flip specially\n        for output, tfm in zip(outputs, tfms):\n            if any(isinstance(t, HFlipTransform) for t in tfm.transforms):\n                output.pred_masks = output.pred_masks.flip(dims=[3])\n        all_pred_masks = torch.stack([o.pred_masks for o in outputs], dim=0)\n        avg_pred_masks = torch.mean(all_pred_masks, dim=0)\n        return avg_pred_masks\n```\nhttps://github.com/facebookresearch/detectron2/blob/cbbc1ce26473cb2a5cc8f58e8ada9ae14cb41052/detectron2/modeling/test_time_augmentation.py#L262",
      "votes": null
    },
    {
      "id": "1584944",
      "postDate": "11/16/2021 22:23:59",
      "content": "<p>Oh, yeah, that's the implementation I used :)</p>",
      "rawMarkdown": "Oh, yeah, that's the implementation I used :)",
      "votes": null
    },
    {
      "id": "1585802",
      "postDate": "11/17/2021 14:56:00",
      "content": "<p>Alright guys, we had fun and all.. buttt it's over. <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> started ensembling. <br>\nThe competition is over! we can all go home now! </p>\n<p>Seriously, I just watched a video of you on youtube and you spoke about some past winning solution.. <br>\nMan.. you pull out some insaaaaaaane stacking pipelines. <br>\nI couldn't figure how you manage to never overfit with such a huge ensemble. I usually overfit already with 10% of the models. </p>\n<p>Respect! </p>",
      "rawMarkdown": "Alright guys, we had fun and all.. buttt it's over. @onodera started ensembling. \nThe competition is over! we can all go home now! \n\nSeriously, I just watched a video of you on youtube and you spoke about some past winning solution.. \nMan.. you pull out some insaaaaaaane stacking pipelines. \nI couldn't figure how you manage to never overfit with such a huge ensemble. I usually overfit already with 10% of the models. \n\nRespect!",
      "votes": null
    },
    {
      "id": "1586225",
      "postDate": "11/17/2021 22:51:36",
      "content": "<p>I also actually found some thresholds/parameters to make my ensemble improve by 0.001-0.002, but it's not stable (3x tta is fine, but 4x or more will produce worse results). However, it seems a bit meaningless since the metric itself is quite noisy, so I doubt such small difference may propagate to private LB. I might be terribly wrong though.</p>",
      "rawMarkdown": "I also actually found some thresholds/parameters to make my ensemble improve by 0.001-0.002, but it's not stable (3x tta is fine, but 4x or more will produce worse results). However, it seems a bit meaningless since the metric itself is quite noisy, so I doubt such small difference may propagate to private LB. I might be terribly wrong though.",
      "votes": null
    },
    {
      "id": "1589523",
      "postDate": "11/20/2021 11:10:13",
      "content": "<p>I also tried two simple ensembles, both of which improved my score from 0.305 to 0.306. </p>\n<p>I'm thinking of trying TTA as well. When I tried to use only vertical flip for prediction, the score became much worse.<br>\nI think this has something to do with the fact that only horizontal flip is used for training. However, that was surprising since I thought there was no vertical difference in optical micrographs. There was almost no deterioration in the score with horizontal flip. </p>",
      "rawMarkdown": "I also tried two simple ensembles, both of which improved my score from 0.305 to 0.306. \n\nI'm thinking of trying TTA as well. When I tried to use only vertical flip for prediction, the score became much worse.\nI think this has something to do with the fact that only horizontal flip is used for training. However, that was surprising since I thought there was no vertical difference in optical micrographs. There was almost no deterioration in the score with horizontal flip.",
      "votes": null
    },
    {
      "id": "1589632",
      "postDate": "11/20/2021 12:56:47",
      "content": "<p>Hi, How to add it?</p>",
      "rawMarkdown": "Hi, How to add it?",
      "votes": null
    },
    {
      "id": "1591382",
      "postDate": "11/22/2021 09:14:59",
      "content": "<p><a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a> I have tried NMS at stage 2, WBF at stage 2, NMS at each stage like mentioned above, and also average the masks, however none of them gave a significant gain in CV and LB unfortunately.</p>",
      "rawMarkdown": "chankhavu I have tried NMS at stage 2, WBF at stage 2, NMS at each stage like mentioned above, and also average the masks, however none of them gave a significant gain in CV and LB unfortunately.",
      "votes": null
    },
    {
      "id": "1594444",
      "postDate": "11/24/2021 20:42:41",
      "content": "<p>I also experimented with TTA and ended up writing a <a href=\"https://www.kaggle.com/mistag/sartorius-tta-with-weighted-segments-fusion/notebook\" target=\"_blank\">Weighted Segments Fusion</a> algorithm - works well on a mediocre performing model. Haven’t succeeded in training a good performing model yet, so can’t say if it also improves score on that…</p>",
      "rawMarkdown": "I also experimented with TTA and ended up writing a [Weighted Segments Fusion](https://www.kaggle.com/mistag/sartorius-tta-with-weighted-segments-fusion/notebook) algorithm - works well on a mediocre performing model. Haven’t succeeded in training a good performing model yet, so can’t say if it also improves score on that…",
      "votes": null
    },
    {
      "id": "1613570",
      "postDate": "12/10/2021 05:29:45",
      "content": "<p>I found the paper for applying TTA to cell instance segmentation. It should have good effect.<br>\n<a href=\"https://www.nature.com/articles/s41598-020-61808-3\" target=\"_blank\">https://www.nature.com/articles/s41598-020-61808-3</a></p>",
      "rawMarkdown": "I found the paper for applying TTA to cell instance segmentation. It should have good effect.\nhttps://www.nature.com/articles/s41598-020-61808-3",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1583204,
      "author_name": "hsadeghian",
      "author_url": "",
      "post_date": "11/15/2021 16:12:36",
      "content": "<p>I have faced the same problem. The method I've used is the union of the masks (with horizontal flip as you mentioned) and have not seen any improvement in my scores. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1583985,
      "author_name": "gowrishankarp",
      "author_url": "",
      "post_date": "11/16/2021 07:28:04",
      "content": "<p>Can I know, What Data-Augmentation transformations have you performed during test time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1584010,
          "author_name": "hsadeghian",
          "author_url": "",
          "post_date": "11/16/2021 07:55:21",
          "content": "<p>As mentioned in <a href=\"https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326:\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326:</a><br>\nCombined predictions on actual image and horizontally flipped image: took unions of masks with maximum overlap and removed false positive masks with small overlap.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1584142,
      "author_name": "osamurai",
      "author_url": "",
      "post_date": "11/16/2021 10:02:31",
      "content": "<p>Excuse me, what is TTA?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1584144,
          "author_name": "osamurai",
          "author_url": "",
          "post_date": "11/16/2021 10:05:09",
          "content": "<p>I learn what TTA is.<br>\nTTA is test time augmentation.<br>\nI found the good source for TTA.<br>\n<a href=\"https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it\" target=\"_blank\">https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it</a></p>\n<p>According to this article, TTA for mask rcnn has good effect on the accuracy.</p>\n<p><a href=\"https://www.researchgate.net/publication/340031950_Test-time_augmentation_for_deep_learning-based_cell_segmentation_on_microscopy_images\" target=\"_blank\">https://www.researchgate.net/publication/340031950_Test-time_augmentation_for_deep_learning-based_cell_segmentation_on_microscopy_images</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1584250,
      "author_name": "onodera",
      "author_url": "",
      "post_date": "11/16/2021 11:43:01",
      "content": "<p>I could improve it from 0.299 to 0.300 with ensemble. I will explore more.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1585802,
          "author_name": "",
          "author_url": "",
          "post_date": "11/17/2021 14:56:00",
          "content": "<p>Alright guys, we had fun and all.. buttt it's over. <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> started ensembling. <br>\nThe competition is over! we can all go home now! </p>\n<p>Seriously, I just watched a video of you on youtube and you spoke about some past winning solution.. <br>\nMan.. you pull out some insaaaaaaane stacking pipelines. <br>\nI couldn't figure how you manage to never overfit with such a huge ensemble. I usually overfit already with 10% of the models. </p>\n<p>Respect! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1586225,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "11/17/2021 22:51:36",
          "content": "<p>I also actually found some thresholds/parameters to make my ensemble improve by 0.001-0.002, but it's not stable (3x tta is fine, but 4x or more will produce worse results). However, it seems a bit meaningless since the metric itself is quite noisy, so I doubt such small difference may propagate to private LB. I might be terribly wrong though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1589523,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "11/20/2021 11:10:13",
          "content": "<p>I also tried two simple ensembles, both of which improved my score from 0.305 to 0.306. </p>\n<p>I'm thinking of trying TTA as well. When I tried to use only vertical flip for prediction, the score became much worse.<br>\nI think this has something to do with the fact that only horizontal flip is used for training. However, that was surprising since I thought there was no vertical difference in optical micrographs. There was almost no deterioration in the score with horizontal flip. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1584260,
      "author_name": "ctawong",
      "author_url": "",
      "post_date": "11/16/2021 11:52:46",
      "content": "<p>Detectron's implementation is averaging the augmented mask.</p>\n<pre><code>def _reduce_pred_masks(self, outputs, tfms):\n        # Should apply inverse transforms on masks.\n        # We assume only resize &amp; flip are used. pred_masks is a scale-invariant\n        # representation, so we handle flip specially\n        for output, tfm in zip(outputs, tfms):\n            if any(isinstance(t, HFlipTransform) for t in tfm.transforms):\n                output.pred_masks = output.pred_masks.flip(dims=[3])\n        all_pred_masks = torch.stack([o.pred_masks for o in outputs], dim=0)\n        avg_pred_masks = torch.mean(all_pred_masks, dim=0)\n        return avg_pred_masks\n</code></pre>\n<p><a href=\"https://github.com/facebookresearch/detectron2/blob/cbbc1ce26473cb2a5cc8f58e8ada9ae14cb41052/detectron2/modeling/test_time_augmentation.py#L262\" target=\"_blank\">https://github.com/facebookresearch/detectron2/blob/cbbc1ce26473cb2a5cc8f58e8ada9ae14cb41052/detectron2/modeling/test_time_augmentation.py#L262</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1584944,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "11/16/2021 22:23:59",
          "content": "<p>Oh, yeah, that's the implementation I used :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1589632,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "11/20/2021 12:56:47",
          "content": "<p>Hi, How to add it?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1591382,
      "author_name": "namgalielei",
      "author_url": "",
      "post_date": "11/22/2021 09:14:59",
      "content": "<p><a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a> I have tried NMS at stage 2, WBF at stage 2, NMS at each stage like mentioned above, and also average the masks, however none of them gave a significant gain in CV and LB unfortunately.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1594444,
      "author_name": "mistag",
      "author_url": "",
      "post_date": "11/24/2021 20:42:41",
      "content": "<p>I also experimented with TTA and ended up writing a <a href=\"https://www.kaggle.com/mistag/sartorius-tta-with-weighted-segments-fusion/notebook\" target=\"_blank\">Weighted Segments Fusion</a> algorithm - works well on a mediocre performing model. Haven’t succeeded in training a good performing model yet, so can’t say if it also improves score on that…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1613570,
      "author_name": "osamurai",
      "author_url": "",
      "post_date": "12/10/2021 05:29:45",
      "content": "<p>I found the paper for applying TTA to cell instance segmentation. It should have good effect.<br>\n<a href=\"https://www.nature.com/articles/s41598-020-61808-3\" target=\"_blank\">https://www.nature.com/articles/s41598-020-61808-3</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1583062": "Hi everyone. I was experimenting with different tta/ensembling schemes on the weekends. However, nothing works reliably for me. My general TTA/Ensembling scheme looks like this:\n\n1. Inference, get T*N sets of predictions from T augmentations and N models (currently T=4 and N=1, just have a simple baseline Mask-RCNN model right now lol).\n2. Filter out small masks from those predictions.\n3. Combine predictions (using methods below).\n4. Filter out overlapping masks, giving priority to predictions with higher score.\n\nHere are the methods I've tried:\n\n1. Take unions of masks with large overlap and discard masks with small overlap (as suggested by winners of DS Bowl 2018, e.g. [here](https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326). I tried different thresholds for that. No visible difference in local CV and on LB.\n2. I tried the following scheme as well (from [1st place in iMaterialist competition](https://www.kaggle.com/c/imaterialist-fashion-2019-FGVC6/discussion/95247)):\n    ![TTA scheme](https://raw.githubusercontent.com/amirassov/kaggle-imaterialist/master/figures/tta.png)\n   but the metrics in local CV only got worse (maybe because of hard NMS, perhaps I should try WBF?)\n\nI'm wondering if anyone has managed to reliably get improvements from ensembling and/or TTA?",
    "1583204": "I have faced the same problem. The method I've used is the union of the masks (with horizontal flip as you mentioned) and have not seen any improvement in my scores.",
    "1583985": "Can I know, What Data-Augmentation transformations have you performed during test time.",
    "1584010": "As mentioned in https://www.kaggle.com/c/data-science-bowl-2018/discussion/56326:\nCombined predictions on actual image and horizontally flipped image: took unions of masks with maximum overlap and removed false positive masks with small overlap.",
    "1584142": "Excuse me, what is TTA?",
    "1584144": "I learn what TTA is.\nTTA is test time augmentation.\nI found the good source for TTA.\nhttps://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it\n\nAccording to this article, TTA for mask rcnn has good effect on the accuracy.\n\nhttps://www.researchgate.net/publication/340031950_Test-time_augmentation_for_deep_learning-based_cell_segmentation_on_microscopy_images",
    "1584250": "I could improve it from 0.299 to 0.300 with ensemble. I will explore more.",
    "1584260": "Detectron's implementation is averaging the augmented mask.\n\n```\ndef _reduce_pred_masks(self, outputs, tfms):\n        # Should apply inverse transforms on masks.\n        # We assume only resize & flip are used. pred_masks is a scale-invariant\n        # representation, so we handle flip specially\n        for output, tfm in zip(outputs, tfms):\n            if any(isinstance(t, HFlipTransform) for t in tfm.transforms):\n                output.pred_masks = output.pred_masks.flip(dims=[3])\n        all_pred_masks = torch.stack([o.pred_masks for o in outputs], dim=0)\n        avg_pred_masks = torch.mean(all_pred_masks, dim=0)\n        return avg_pred_masks\n```\nhttps://github.com/facebookresearch/detectron2/blob/cbbc1ce26473cb2a5cc8f58e8ada9ae14cb41052/detectron2/modeling/test_time_augmentation.py#L262",
    "1584944": "Oh, yeah, that's the implementation I used :)",
    "1585802": "Alright guys, we had fun and all.. buttt it's over. @onodera started ensembling. \nThe competition is over! we can all go home now! \n\nSeriously, I just watched a video of you on youtube and you spoke about some past winning solution.. \nMan.. you pull out some insaaaaaaane stacking pipelines. \nI couldn't figure how you manage to never overfit with such a huge ensemble. I usually overfit already with 10% of the models. \n\nRespect!",
    "1586225": "I also actually found some thresholds/parameters to make my ensemble improve by 0.001-0.002, but it's not stable (3x tta is fine, but 4x or more will produce worse results). However, it seems a bit meaningless since the metric itself is quite noisy, so I doubt such small difference may propagate to private LB. I might be terribly wrong though.",
    "1589523": "I also tried two simple ensembles, both of which improved my score from 0.305 to 0.306. \n\nI'm thinking of trying TTA as well. When I tried to use only vertical flip for prediction, the score became much worse.\nI think this has something to do with the fact that only horizontal flip is used for training. However, that was surprising since I thought there was no vertical difference in optical micrographs. There was almost no deterioration in the score with horizontal flip.",
    "1589632": "Hi, How to add it?",
    "1591382": "chankhavu I have tried NMS at stage 2, WBF at stage 2, NMS at each stage like mentioned above, and also average the masks, however none of them gave a significant gain in CV and LB unfortunately.",
    "1594444": "I also experimented with TTA and ended up writing a [Weighted Segments Fusion](https://www.kaggle.com/mistag/sartorius-tta-with-weighted-segments-fusion/notebook) algorithm - works well on a mediocre performing model. Haven’t succeeded in training a good performing model yet, so can’t say if it also improves score on that…",
    "1613570": "I found the paper for applying TTA to cell instance segmentation. It should have good effect.\nhttps://www.nature.com/articles/s41598-020-61808-3"
  },
  "source": "meta"
}