{
  "id": 207472,
  "title": "What might be the cause of score drop when using TTA?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/207472",
  "author_name": "",
  "post_date": "2020-12-29T20:14:28.451911200Z",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>I hope you all doing well. I wonder that what might be the root cause of score drop when using TTA,  especially in the case of segmentation. Like too much distortion, information loss etc. Which augmentation methods can cause them. If you have previous experiences and enlighten me about the potential problems, it would be too much appreciated. Thanks!</p>",
  "messages": [
    {
      "id": "1131569",
      "postDate": "12/29/2020 20:14:28",
      "content": "<p>Hi everyone, </p>\n<p>I hope you all doing well. I wonder that what might be the root cause of score drop when using TTA,  especially in the case of segmentation. Like too much distortion, information loss etc. Which augmentation methods can cause them. If you have previous experiences and enlighten me about the potential problems, it would be too much appreciated. Thanks!</p>",
      "rawMarkdown": "Hi everyone, \n\nI hope you all doing well. I wonder that what might be the root cause of score drop when using TTA,  especially in the case of segmentation. Like too much distortion, information loss etc. Which augmentation methods can cause them. If you have previous experiences and enlighten me about the potential problems, it would be too much appreciated. Thanks!",
      "votes": null
    },
    {
      "id": "1133382",
      "postDate": "12/31/2020 07:07:19",
      "content": "<p>To me, this is a hint, maybe you have not used TTA correctly, right? But I haven't used TTA, and I am going to use it this time.</p>",
      "rawMarkdown": "To me, this is a hint, maybe you have not used TTA correctly, right? But I haven't used TTA, and I am going to use it this time.",
      "votes": null
    },
    {
      "id": "1133464",
      "postDate": "12/31/2020 08:30:57",
      "content": "<p>Now I have the same problem.<br>\nI think this is because of the original incorrect shift and my technical problem in TTA in the first place.</p>",
      "rawMarkdown": "Now I have the same problem.\nI think this is because of the original incorrect shift and my technical problem in TTA in the first place.",
      "votes": null
    },
    {
      "id": "1133849",
      "postDate": "12/31/2020 15:29:17",
      "content": "<p>I actually tried the flips and some hue. It's important to deaugment the flips btw. I plotted the masks and it seemed like there was no mistake. I am still confused unfortunately. No progress so far.</p>",
      "rawMarkdown": "I actually tried the flips and some hue. It's important to deaugment the flips btw. I plotted the masks and it seemed like there was no mistake. I am still confused unfortunately. No progress so far.",
      "votes": null
    },
    {
      "id": "1133854",
      "postDate": "12/31/2020 15:31:26",
      "content": "<p>Yes, the labelling errors might be a problem. But, in some high scoring public notebooks, TTA is applied. I don't currently know how to interpret this one.</p>",
      "rawMarkdown": "Yes, the labelling errors might be a problem. But, in some high scoring public notebooks, TTA is applied. I don't currently know how to interpret this one.",
      "votes": null
    },
    {
      "id": "1135404",
      "postDate": "01/02/2021 08:18:41",
      "content": "<p>TTA can affect predictions both ways, the best approach is to test different augmentations systematically. For segmentation models, flipping, rotating and cropping usually works well in TTA, while hue and brightness adjustments might hurt performance. Also consider that predicted masks must be de-augmented before being combined, so e.g. distortion effects (shearing etc.) are good for training but not so much for TTA.</p>",
      "rawMarkdown": "TTA can affect predictions both ways, the best approach is to test different augmentations systematically. For segmentation models, flipping, rotating and cropping usually works well in TTA, while hue and brightness adjustments might hurt performance. Also consider that predicted masks must be de-augmented before being combined, so e.g. distortion effects (shearing etc.) are good for training but not so much for TTA.",
      "votes": null
    },
    {
      "id": "1136182",
      "postDate": "01/02/2021 19:51:04",
      "content": "<p>for me it'a all trial-and-error. Add them one by one and identify the ones that cause problems</p>",
      "rawMarkdown": "for me it'a all trial-and-error. Add them one by one and identify the ones that cause problems",
      "votes": null
    },
    {
      "id": "1136358",
      "postDate": "01/03/2021 02:06:34",
      "content": "<p>You may want to test your augmentation code and confirm that it does not produce identically-augmented images for each inference fold.</p>",
      "rawMarkdown": "You may want to test your augmentation code and confirm that it does not produce identically-augmented images for each inference fold.",
      "votes": null
    },
    {
      "id": "1139816",
      "postDate": "01/05/2021 16:30:09",
      "content": "<p>BTW, can confirm that all combinations of image flipping 90, horizontally and vertically resulted in a 0.002 drop LB score. Freakishly, the exact value for all combinations!!</p>",
      "rawMarkdown": "BTW, can confirm that all combinations of image flipping 90, horizontally and vertically resulted in a 0.002 drop LB score. Freakishly, the exact value for all combinations!!",
      "votes": null
    },
    {
      "id": "1139910",
      "postDate": "01/05/2021 17:35:31",
      "content": "<p>That's a nice observation. Thanks</p>",
      "rawMarkdown": "That's a nice observation. Thanks",
      "votes": null
    },
    {
      "id": "1139912",
      "postDate": "01/05/2021 17:37:37",
      "content": "<p>I actually plotted the masks and they seem to be ok for me comparing to the ground truth masks. I will go one by one after this time. Thanks for your answer as well.</p>",
      "rawMarkdown": "I actually plotted the masks and they seem to be ok for me comparing to the ground truth masks. I will go one by one after this time. Thanks for your answer as well.",
      "votes": null
    },
    {
      "id": "1141474",
      "postDate": "01/06/2021 18:18:59",
      "content": "<p>TTA in segmentation tasks usually improve predictions at an object's border where predictions are less confident.  Because the labels in this case are rather noisy especially around the border region, TTA is unlikely to help much.  The other factor here is many of the labels appear to have been machine generated with polygon type characteristics (sharp lines at borders).  Look at the delta maps of TTA vs non-TTA predictions and you'll see the extra smoothing with TTA which loses some of the polygon type features contained in the actual labels.</p>",
      "rawMarkdown": "TTA in segmentation tasks usually improve predictions at an object's border where predictions are less confident.  Because the labels in this case are rather noisy especially around the border region, TTA is unlikely to help much.  The other factor here is many of the labels appear to have been machine generated with polygon type characteristics (sharp lines at borders).  Look at the delta maps of TTA vs non-TTA predictions and you'll see the extra smoothing with TTA which loses some of the polygon type features contained in the actual labels.",
      "votes": null
    },
    {
      "id": "1142773",
      "postDate": "01/07/2021 15:29:04",
      "content": "<p>Thanks a lot for your insightful comment. Do you have any recommendation about the border part to fix (?) that problem?</p>",
      "rawMarkdown": "Thanks a lot for your insightful comment. Do you have any recommendation about the border part to fix (?) that problem?",
      "votes": null
    },
    {
      "id": "1163099",
      "postDate": "01/21/2021 14:27:53",
      "content": "<p>hue won't help, in my opinion, as TTA mainly improves border prediction. All the flips, rotations, etc. seemed to add 0.001-ish points on the LB for me. </p>",
      "rawMarkdown": "hue won't help, in my opinion, as TTA mainly improves border prediction. All the flips, rotations, etc. seemed to add 0.001-ish points on the LB for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1133382,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/31/2020 07:07:19",
      "content": "<p>To me, this is a hint, maybe you have not used TTA correctly, right? But I haven't used TTA, and I am going to use it this time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1133849,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "12/31/2020 15:29:17",
          "content": "<p>I actually tried the flips and some hue. It's important to deaugment the flips btw. I plotted the masks and it seemed like there was no mistake. I am still confused unfortunately. No progress so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1163099,
          "author_name": "andrasferenczi",
          "author_url": "",
          "post_date": "01/21/2021 14:27:53",
          "content": "<p>hue won't help, in my opinion, as TTA mainly improves border prediction. All the flips, rotations, etc. seemed to add 0.001-ish points on the LB for me. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1133464,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "12/31/2020 08:30:57",
      "content": "<p>Now I have the same problem.<br>\nI think this is because of the original incorrect shift and my technical problem in TTA in the first place.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1133854,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "12/31/2020 15:31:26",
          "content": "<p>Yes, the labelling errors might be a problem. But, in some high scoring public notebooks, TTA is applied. I don't currently know how to interpret this one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1135404,
      "author_name": "mistag",
      "author_url": "",
      "post_date": "01/02/2021 08:18:41",
      "content": "<p>TTA can affect predictions both ways, the best approach is to test different augmentations systematically. For segmentation models, flipping, rotating and cropping usually works well in TTA, while hue and brightness adjustments might hurt performance. Also consider that predicted masks must be de-augmented before being combined, so e.g. distortion effects (shearing etc.) are good for training but not so much for TTA.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1139912,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/05/2021 17:37:37",
          "content": "<p>I actually plotted the masks and they seem to be ok for me comparing to the ground truth masks. I will go one by one after this time. Thanks for your answer as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136182,
      "author_name": "andrasferenczi",
      "author_url": "",
      "post_date": "01/02/2021 19:51:04",
      "content": "<p>for me it'a all trial-and-error. Add them one by one and identify the ones that cause problems</p>",
      "votes": null,
      "replies": [
        {
          "id": 1139816,
          "author_name": "andrasferenczi",
          "author_url": "",
          "post_date": "01/05/2021 16:30:09",
          "content": "<p>BTW, can confirm that all combinations of image flipping 90, horizontally and vertically resulted in a 0.002 drop LB score. Freakishly, the exact value for all combinations!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139910,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/05/2021 17:35:31",
          "content": "<p>That's a nice observation. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136358,
      "author_name": "eschibli",
      "author_url": "",
      "post_date": "01/03/2021 02:06:34",
      "content": "<p>You may want to test your augmentation code and confirm that it does not produce identically-augmented images for each inference fold.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1141474,
      "author_name": "tivfrvqhs5",
      "author_url": "",
      "post_date": "01/06/2021 18:18:59",
      "content": "<p>TTA in segmentation tasks usually improve predictions at an object's border where predictions are less confident.  Because the labels in this case are rather noisy especially around the border region, TTA is unlikely to help much.  The other factor here is many of the labels appear to have been machine generated with polygon type characteristics (sharp lines at borders).  Look at the delta maps of TTA vs non-TTA predictions and you'll see the extra smoothing with TTA which loses some of the polygon type features contained in the actual labels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1142773,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/07/2021 15:29:04",
          "content": "<p>Thanks a lot for your insightful comment. Do you have any recommendation about the border part to fix (?) that problem?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1131569": "Hi everyone, \n\nI hope you all doing well. I wonder that what might be the root cause of score drop when using TTA,  especially in the case of segmentation. Like too much distortion, information loss etc. Which augmentation methods can cause them. If you have previous experiences and enlighten me about the potential problems, it would be too much appreciated. Thanks!",
    "1133382": "To me, this is a hint, maybe you have not used TTA correctly, right? But I haven't used TTA, and I am going to use it this time.",
    "1133464": "Now I have the same problem.\nI think this is because of the original incorrect shift and my technical problem in TTA in the first place.",
    "1133849": "I actually tried the flips and some hue. It's important to deaugment the flips btw. I plotted the masks and it seemed like there was no mistake. I am still confused unfortunately. No progress so far.",
    "1133854": "Yes, the labelling errors might be a problem. But, in some high scoring public notebooks, TTA is applied. I don't currently know how to interpret this one.",
    "1135404": "TTA can affect predictions both ways, the best approach is to test different augmentations systematically. For segmentation models, flipping, rotating and cropping usually works well in TTA, while hue and brightness adjustments might hurt performance. Also consider that predicted masks must be de-augmented before being combined, so e.g. distortion effects (shearing etc.) are good for training but not so much for TTA.",
    "1136182": "for me it'a all trial-and-error. Add them one by one and identify the ones that cause problems",
    "1136358": "You may want to test your augmentation code and confirm that it does not produce identically-augmented images for each inference fold.",
    "1139816": "BTW, can confirm that all combinations of image flipping 90, horizontally and vertically resulted in a 0.002 drop LB score. Freakishly, the exact value for all combinations!!",
    "1139910": "That's a nice observation. Thanks",
    "1139912": "I actually plotted the masks and they seem to be ok for me comparing to the ground truth masks. I will go one by one after this time. Thanks for your answer as well.",
    "1141474": "TTA in segmentation tasks usually improve predictions at an object's border where predictions are less confident.  Because the labels in this case are rather noisy especially around the border region, TTA is unlikely to help much.  The other factor here is many of the labels appear to have been machine generated with polygon type characteristics (sharp lines at borders).  Look at the delta maps of TTA vs non-TTA predictions and you'll see the extra smoothing with TTA which loses some of the polygon type features contained in the actual labels.",
    "1142773": "Thanks a lot for your insightful comment. Do you have any recommendation about the border part to fix (?) that problem?",
    "1163099": "hue won't help, in my opinion, as TTA mainly improves border prediction. All the flips, rotations, etc. seemed to add 0.001-ish points on the LB for me."
  },
  "source": "meta"
}