{
  "id": 210921,
  "title": "People say no TTA for this competition but are people doing VTA?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210921",
  "author_name": "",
  "post_date": "2021-01-13T01:35:21.987729100Z",
  "votes": 4,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I have seen some discussion posts regarding Test Time Augmentation, but what are peoples thoughts of Validation Time Augmentation.</p>\n<p>Do the same rules apply? <br>\nAre people augmenting their validation dataset when training models?</p>",
  "messages": [
    {
      "id": "1150940",
      "postDate": "01/13/2021 01:35:21",
      "content": "<p>I have seen some discussion posts regarding Test Time Augmentation, but what are peoples thoughts of Validation Time Augmentation.</p>\n<p>Do the same rules apply? <br>\nAre people augmenting their validation dataset when training models?</p>",
      "rawMarkdown": "I have seen some discussion posts regarding Test Time Augmentation, but what are peoples thoughts of Validation Time Augmentation.\n\nDo the same rules apply? \nAre people augmenting their validation dataset when training models?",
      "votes": null
    },
    {
      "id": "1150950",
      "postDate": "01/13/2021 01:52:44",
      "content": "<p>When calculating the validation score, I do the same augmentation as in TTA. Otherwise, the validation score will not be calculated properly and the correlation between CV and LB may worsen. I usually apply <code>Resize</code> or <code>CenterCrop</code> and <code>Normalize</code> when validation and final prediction.</p>",
      "rawMarkdown": "When calculating the validation score, I do the same augmentation as in TTA. Otherwise, the validation score will not be calculated properly and the correlation between CV and LB may worsen. I usually apply `Resize` or `CenterCrop` and `Normalize` when validation and final prediction.",
      "votes": null
    },
    {
      "id": "1150978",
      "postDate": "01/13/2021 03:11:58",
      "content": "<p>Interesting. Could you just get lucky with validation augmentation and end up with a validation_accuracy that isn't representative of the true model accuracy?</p>",
      "rawMarkdown": "Interesting. Could you just get lucky with validation augmentation and end up with a validation_accuracy that isn't representative of the true model accuracy?",
      "votes": null
    },
    {
      "id": "1151106",
      "postDate": "01/13/2021 06:29:31",
      "content": "<p>VTA? Means verification set TTA? Then the verification set and the test set use the same TTA to ensure that the scores of the verification set and the test set are similar? So CV is closer to LB? Excuse me: is my understanding right? Use <strong>crop</strong> with caution! Maybe you will drop the important features of <strong>crop</strong>!</p>",
      "rawMarkdown": "VTA? Means verification set TTA? Then the verification set and the test set use the same TTA to ensure that the scores of the verification set and the test set are similar? So CV is closer to LB? Excuse me: is my understanding right? Use **crop** with caution! Maybe you will drop the important features of **crop**!",
      "votes": null
    },
    {
      "id": "1151115",
      "postDate": "01/13/2021 06:32:10",
      "content": "<p>VTA will increase the training &amp; verification time. Of course, if you use TPU and there are a lot of TPU quotas, you can ignore my idea, but if you train on the local computer, it will increase the time a lot</p>",
      "rawMarkdown": "VTA will increase the training & verification time. Of course, if you use TPU and there are a lot of TPU quotas, you can ignore my idea, but if you train on the local computer, it will increase the time a lot",
      "votes": null
    },
    {
      "id": "1151453",
      "postDate": "01/13/2021 10:36:29",
      "content": "<p>Resizing may also drop the important features. If a small spot in an image is important, resizing will make it harder to predict the correct label. In my experiment, <code>CenterCrop</code> gives me better result than <code>Resize</code>. Of course, I don't use <code>CenterCrop</code> when input size is less than 512.</p>",
      "rawMarkdown": "Resizing may also drop the important features. If a small spot in an image is important, resizing will make it harder to predict the correct label. In my experiment, `CenterCrop` gives me better result than `Resize`. Of course, I don't use `CenterCrop` when input size is less than 512.",
      "votes": null
    },
    {
      "id": "1151713",
      "postDate": "01/13/2021 14:09:57",
      "content": "<blockquote>\n  <p>People say no TTA for this competition</p>\n</blockquote>\n<p>Which people ?  Did you try it yourself ? </p>\n<p>You should take what is said in discussions with a bit grain of salt and do your own experiments.  Because some tricks may be helpful for some and not for others, depending on the general setup</p>\n<p>And don't forget their is a private LB ^^</p>",
      "rawMarkdown": "> People say no TTA for this competition\n\nWhich people ?  Did you try it yourself ? \n\nYou should take what is said in discussions with a bit grain of salt and do your own experiments.  Because some tricks may be helpful for some and not for others, depending on the general setup\n\nAnd don't forget their is a private LB ^^",
      "votes": null
    },
    {
      "id": "1151736",
      "postDate": "01/13/2021 14:22:21",
      "content": "<p>Yes my tests for my submissions suggest that 3xTTA gives good results. 4 or more is worse as is no TTA. I’m yet to try one or two TTA.</p>",
      "rawMarkdown": "Yes my tests for my submissions suggest that 3xTTA gives good results. 4 or more is worse as is no TTA. I’m yet to try one or two TTA.",
      "votes": null
    },
    {
      "id": "1151796",
      "postDate": "01/13/2021 15:03:29",
      "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/mutantspore\" target=\"_blank\">@mutantspore</a> and <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> .</p>\n<p>Yes I have tried submits with identical models with TTA and without TTA and was just wondering what peoples thoughts were..</p>\n<p>In regards to Visible LB score vs Final LB score on all data I personally think TTA might overfit on the visible data. </p>",
      "rawMarkdown": "Thanks for the reply @mutantspore and @serigne .\n\nYes I have tried submits with identical models with TTA and without TTA and was just wondering what peoples thoughts were..\n\nIn regards to Visible LB score vs Final LB score on all data I personally think TTA might overfit on the visible data.",
      "votes": null
    },
    {
      "id": "1151801",
      "postDate": "01/13/2021 15:07:11",
      "content": "<p><a href=\"https://www.kaggle.com/zhengeng\" target=\"_blank\">@zhengeng</a> Yep exactly right. </p>\n<p>Do you use the same Augmentation techniques when submitting a model to the test data, as when you are evaluating the model accuracy during training?</p>\n<p>Thanks for the reply!</p>",
      "rawMarkdown": "zhengeng Yep exactly right. \n\nDo you use the same Augmentation techniques when submitting a model to the test data, as when you are evaluating the model accuracy during training?\n\nThanks for the reply!",
      "votes": null
    },
    {
      "id": "1153187",
      "postDate": "01/14/2021 16:59:40",
      "content": "<p>My understanding is that TTA helps because it averages out one model which overfitted a certain class or an augmentation on an image which made classification more difficult. If you do that on the validation set you will have an increased overall accuracy, but there is a limit to the improvement. In my experimentation i find useful to do VTA after training, just to test if the model learned how to interpret augmentations properly with the best trained model from a single augmentation.<br>\nYou can see here: <a href=\"https://www.kaggle.com/capiru/effnetb4-autoaugment-lossfns-ranger-easy-to-start\" target=\"_blank\">https://www.kaggle.com/capiru/effnetb4-autoaugment-lossfns-ranger-easy-to-start</a><br>\nThe model by training just one epoch didn't learn enough generalizations about transforms, but if you train for ~10 epochs, you will see that the model can go from 88.x to ~89ish.</p>",
      "rawMarkdown": "My understanding is that TTA helps because it averages out one model which overfitted a certain class or an augmentation on an image which made classification more difficult. If you do that on the validation set you will have an increased overall accuracy, but there is a limit to the improvement. In my experimentation i find useful to do VTA after training, just to test if the model learned how to interpret augmentations properly with the best trained model from a single augmentation.\nYou can see here: https://www.kaggle.com/capiru/effnetb4-autoaugment-lossfns-ranger-easy-to-start\nThe model by training just one epoch didn't learn enough generalizations about transforms, but if you train for ~10 epochs, you will see that the model can go from 88.x to ~89ish.",
      "votes": null
    },
    {
      "id": "1153396",
      "postDate": "01/14/2021 21:09:46",
      "content": "<p>Regarding doing TTA, don't forget to add the original image to your TTA pipeline :) We were missing it while back then. It improved 0.2-0.3 in the public LB. Simple but it is so easy to overlook 😃</p>",
      "rawMarkdown": "Regarding doing TTA, don't forget to add the original image to your TTA pipeline :) We were missing it while back then. It improved 0.2-0.3 in the public LB. Simple but it is so easy to overlook 😃",
      "votes": null
    },
    {
      "id": "1164474",
      "postDate": "01/22/2021 11:48:33",
      "content": "<p>Do you mind providing an example of the standard tta?</p>",
      "rawMarkdown": "Do you mind providing an example of the standard tta?",
      "votes": null
    },
    {
      "id": "1165320",
      "postDate": "01/22/2021 21:46:57",
      "content": "<p>Hi sure, I can provide you a pseudocode but it might seem a little bit easy. </p>\n<pre><code>tta_steps = 3\n# Pytorch datasets &amp; dataloaders\n# The important part is the transforms that you pass to the dataset class. \nraw_dataset = CassavaDataset(..., transforms=None) \nraw_loader = DataLoader(raw_dataset)\n\ntta_dataset = CassavaDataset(..., transforms=augment_transforms())\ntta_loader = DataLoader(tta_dataset)\n\npreds = None\nfor i in range(tta_steps):\n    predictions = inference(model, tta_loader)\n    if preds is None:\n        preds = predictions\n    else:\n        preds += predictions\n\nraw_preds = inference(model, raw_loader)\npreds += raw_preds\n\n# 1 is for the raw_loader\npreds /= (tta_steps + 1)\n</code></pre>\n<p>This is how we do it.</p>",
      "rawMarkdown": "Hi sure, I can provide you a pseudocode but it might seem a little bit easy. \n\n```python\n\ntta_steps = 3\n# Pytorch datasets & dataloaders\n# The important part is the transforms that you pass to the dataset class. \nraw_dataset = CassavaDataset(..., transforms=None) \nraw_loader = DataLoader(raw_dataset)\n\ntta_dataset = CassavaDataset(..., transforms=augment_transforms())\ntta_loader = DataLoader(tta_dataset)\n\npreds = None\nfor i in range(tta_steps):\n    predictions = inference(model, tta_loader)\n    if preds is None:\n        preds = predictions\n    else:\n        preds += predictions\n\nraw_preds = inference(model, raw_loader)\npreds += raw_preds\n\n# 1 is for the raw_loader\npreds /= (tta_steps + 1)\n\n```\n\nThis is how we do it.",
      "votes": null
    },
    {
      "id": "1169916",
      "postDate": "01/25/2021 20:27:35",
      "content": "<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> Many Many thanks for this sudo code. </p>",
      "rawMarkdown": "snnclsr Many Many thanks for this sudo code.",
      "votes": null
    },
    {
      "id": "1169938",
      "postDate": "01/25/2021 20:48:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/durbin164\" target=\"_blank\">@durbin164</a> </p>\n<p>You are welcome. Hope it is useful 🙏</p>",
      "rawMarkdown": "Hi @durbin164 \n\nYou are welcome. Hope it is useful 🙏",
      "votes": null
    },
    {
      "id": "1170074",
      "postDate": "01/26/2021 00:21:28",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> all of them, such as flipping and clipping, can be used in TTA, but cutout, mixup, cutmix, etc. can't be used in TTA. At least I haven't seen or tried a similar enhancement scheme for TTA</p>",
      "rawMarkdown": "brendanartley all of them, such as flipping and clipping, can be used in TTA, but cutout, mixup, cutmix, etc. can't be used in TTA. At least I haven't seen or tried a similar enhancement scheme for TTA",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1150950,
      "author_name": "tmhrkt",
      "author_url": "",
      "post_date": "01/13/2021 01:52:44",
      "content": "<p>When calculating the validation score, I do the same augmentation as in TTA. Otherwise, the validation score will not be calculated properly and the correlation between CV and LB may worsen. I usually apply <code>Resize</code> or <code>CenterCrop</code> and <code>Normalize</code> when validation and final prediction.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1150978,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "01/13/2021 03:11:58",
          "content": "<p>Interesting. Could you just get lucky with validation augmentation and end up with a validation_accuracy that isn't representative of the true model accuracy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1153187,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "01/14/2021 16:59:40",
          "content": "<p>My understanding is that TTA helps because it averages out one model which overfitted a certain class or an augmentation on an image which made classification more difficult. If you do that on the validation set you will have an increased overall accuracy, but there is a limit to the improvement. In my experimentation i find useful to do VTA after training, just to test if the model learned how to interpret augmentations properly with the best trained model from a single augmentation.<br>\nYou can see here: <a href=\"https://www.kaggle.com/capiru/effnetb4-autoaugment-lossfns-ranger-easy-to-start\" target=\"_blank\">https://www.kaggle.com/capiru/effnetb4-autoaugment-lossfns-ranger-easy-to-start</a><br>\nThe model by training just one epoch didn't learn enough generalizations about transforms, but if you train for ~10 epochs, you will see that the model can go from 88.x to ~89ish.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1151106,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "01/13/2021 06:29:31",
      "content": "<p>VTA? Means verification set TTA? Then the verification set and the test set use the same TTA to ensure that the scores of the verification set and the test set are similar? So CV is closer to LB? Excuse me: is my understanding right? Use <strong>crop</strong> with caution! Maybe you will drop the important features of <strong>crop</strong>!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1151453,
          "author_name": "tmhrkt",
          "author_url": "",
          "post_date": "01/13/2021 10:36:29",
          "content": "<p>Resizing may also drop the important features. If a small spot in an image is important, resizing will make it harder to predict the correct label. In my experiment, <code>CenterCrop</code> gives me better result than <code>Resize</code>. Of course, I don't use <code>CenterCrop</code> when input size is less than 512.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151801,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "01/13/2021 15:07:11",
          "content": "<p><a href=\"https://www.kaggle.com/zhengeng\" target=\"_blank\">@zhengeng</a> Yep exactly right. </p>\n<p>Do you use the same Augmentation techniques when submitting a model to the test data, as when you are evaluating the model accuracy during training?</p>\n<p>Thanks for the reply!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1170074,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "01/26/2021 00:21:28",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> all of them, such as flipping and clipping, can be used in TTA, but cutout, mixup, cutmix, etc. can't be used in TTA. At least I haven't seen or tried a similar enhancement scheme for TTA</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1151115,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "01/13/2021 06:32:10",
      "content": "<p>VTA will increase the training &amp; verification time. Of course, if you use TPU and there are a lot of TPU quotas, you can ignore my idea, but if you train on the local computer, it will increase the time a lot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1151713,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/13/2021 14:09:57",
      "content": "<blockquote>\n  <p>People say no TTA for this competition</p>\n</blockquote>\n<p>Which people ?  Did you try it yourself ? </p>\n<p>You should take what is said in discussions with a bit grain of salt and do your own experiments.  Because some tricks may be helpful for some and not for others, depending on the general setup</p>\n<p>And don't forget their is a private LB ^^</p>",
      "votes": null,
      "replies": [
        {
          "id": 1151736,
          "author_name": "mutantspore",
          "author_url": "",
          "post_date": "01/13/2021 14:22:21",
          "content": "<p>Yes my tests for my submissions suggest that 3xTTA gives good results. 4 or more is worse as is no TTA. I’m yet to try one or two TTA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151796,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "01/13/2021 15:03:29",
          "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/mutantspore\" target=\"_blank\">@mutantspore</a> and <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> .</p>\n<p>Yes I have tried submits with identical models with TTA and without TTA and was just wondering what peoples thoughts were..</p>\n<p>In regards to Visible LB score vs Final LB score on all data I personally think TTA might overfit on the visible data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1153396,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "01/14/2021 21:09:46",
      "content": "<p>Regarding doing TTA, don't forget to add the original image to your TTA pipeline :) We were missing it while back then. It improved 0.2-0.3 in the public LB. Simple but it is so easy to overlook 😃</p>",
      "votes": null,
      "replies": [
        {
          "id": 1164474,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "01/22/2021 11:48:33",
          "content": "<p>Do you mind providing an example of the standard tta?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1165320,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/22/2021 21:46:57",
          "content": "<p>Hi sure, I can provide you a pseudocode but it might seem a little bit easy. </p>\n<pre><code>tta_steps = 3\n# Pytorch datasets &amp; dataloaders\n# The important part is the transforms that you pass to the dataset class. \nraw_dataset = CassavaDataset(..., transforms=None) \nraw_loader = DataLoader(raw_dataset)\n\ntta_dataset = CassavaDataset(..., transforms=augment_transforms())\ntta_loader = DataLoader(tta_dataset)\n\npreds = None\nfor i in range(tta_steps):\n    predictions = inference(model, tta_loader)\n    if preds is None:\n        preds = predictions\n    else:\n        preds += predictions\n\nraw_preds = inference(model, raw_loader)\npreds += raw_preds\n\n# 1 is for the raw_loader\npreds /= (tta_steps + 1)\n</code></pre>\n<p>This is how we do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1169916,
          "author_name": "durbin164",
          "author_url": "",
          "post_date": "01/25/2021 20:27:35",
          "content": "<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> Many Many thanks for this sudo code. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1169938,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "01/25/2021 20:48:34",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/durbin164\" target=\"_blank\">@durbin164</a> </p>\n<p>You are welcome. Hope it is useful 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1150940": "I have seen some discussion posts regarding Test Time Augmentation, but what are peoples thoughts of Validation Time Augmentation.\n\nDo the same rules apply? \nAre people augmenting their validation dataset when training models?",
    "1150950": "When calculating the validation score, I do the same augmentation as in TTA. Otherwise, the validation score will not be calculated properly and the correlation between CV and LB may worsen. I usually apply `Resize` or `CenterCrop` and `Normalize` when validation and final prediction.",
    "1150978": "Interesting. Could you just get lucky with validation augmentation and end up with a validation_accuracy that isn't representative of the true model accuracy?",
    "1151106": "VTA? Means verification set TTA? Then the verification set and the test set use the same TTA to ensure that the scores of the verification set and the test set are similar? So CV is closer to LB? Excuse me: is my understanding right? Use **crop** with caution! Maybe you will drop the important features of **crop**!",
    "1151115": "VTA will increase the training & verification time. Of course, if you use TPU and there are a lot of TPU quotas, you can ignore my idea, but if you train on the local computer, it will increase the time a lot",
    "1151453": "Resizing may also drop the important features. If a small spot in an image is important, resizing will make it harder to predict the correct label. In my experiment, `CenterCrop` gives me better result than `Resize`. Of course, I don't use `CenterCrop` when input size is less than 512.",
    "1151713": "> People say no TTA for this competition\n\nWhich people ?  Did you try it yourself ? \n\nYou should take what is said in discussions with a bit grain of salt and do your own experiments.  Because some tricks may be helpful for some and not for others, depending on the general setup\n\nAnd don't forget their is a private LB ^^",
    "1151736": "Yes my tests for my submissions suggest that 3xTTA gives good results. 4 or more is worse as is no TTA. I’m yet to try one or two TTA.",
    "1151796": "Thanks for the reply @mutantspore and @serigne .\n\nYes I have tried submits with identical models with TTA and without TTA and was just wondering what peoples thoughts were..\n\nIn regards to Visible LB score vs Final LB score on all data I personally think TTA might overfit on the visible data.",
    "1151801": "zhengeng Yep exactly right. \n\nDo you use the same Augmentation techniques when submitting a model to the test data, as when you are evaluating the model accuracy during training?\n\nThanks for the reply!",
    "1153187": "My understanding is that TTA helps because it averages out one model which overfitted a certain class or an augmentation on an image which made classification more difficult. If you do that on the validation set you will have an increased overall accuracy, but there is a limit to the improvement. In my experimentation i find useful to do VTA after training, just to test if the model learned how to interpret augmentations properly with the best trained model from a single augmentation.\nYou can see here: https://www.kaggle.com/capiru/effnetb4-autoaugment-lossfns-ranger-easy-to-start\nThe model by training just one epoch didn't learn enough generalizations about transforms, but if you train for ~10 epochs, you will see that the model can go from 88.x to ~89ish.",
    "1153396": "Regarding doing TTA, don't forget to add the original image to your TTA pipeline :) We were missing it while back then. It improved 0.2-0.3 in the public LB. Simple but it is so easy to overlook 😃",
    "1164474": "Do you mind providing an example of the standard tta?",
    "1165320": "Hi sure, I can provide you a pseudocode but it might seem a little bit easy. \n\n```python\n\ntta_steps = 3\n# Pytorch datasets & dataloaders\n# The important part is the transforms that you pass to the dataset class. \nraw_dataset = CassavaDataset(..., transforms=None) \nraw_loader = DataLoader(raw_dataset)\n\ntta_dataset = CassavaDataset(..., transforms=augment_transforms())\ntta_loader = DataLoader(tta_dataset)\n\npreds = None\nfor i in range(tta_steps):\n    predictions = inference(model, tta_loader)\n    if preds is None:\n        preds = predictions\n    else:\n        preds += predictions\n\nraw_preds = inference(model, raw_loader)\npreds += raw_preds\n\n# 1 is for the raw_loader\npreds /= (tta_steps + 1)\n\n```\n\nThis is how we do it.",
    "1169916": "snnclsr Many Many thanks for this sudo code.",
    "1169938": "Hi @durbin164 \n\nYou are welcome. Hope it is useful 🙏",
    "1170074": "brendanartley all of them, such as flipping and clipping, can be used in TTA, but cutout, mixup, cutmix, etc. can't be used in TTA. At least I haven't seen or tried a similar enhancement scheme for TTA"
  },
  "source": "meta"
}