{
  "id": 172019,
  "title": "Test time augmentations increases CV",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172019",
  "author_name": "",
  "post_date": "2020-08-03T11:01:06.998170100Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Many competitors have reported that test time augmentation is effective at increasing CV/LB. Chris Deotte even reported that people were using 25+ augmentation steps, but there isn't a formal topic dedicated to test time augmentations yet. </p>\n\n<p>Firstly, I wanted to share an experiment I did : looking at how the number of TTA changed my mean CV score. The model uses images of dimensions (128, 128), EfnB2 and is evaluated using 5 (triple) stratified folds. I only do basic affine transformations and brightness/saturation changes (e.g., no dropout or hair augmentation is used).\n![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F403021ac2e7cb5866e2f8380673fa49a%2Ftta_steps_investigation.png?generation=1596451769524367&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F403021ac2e7cb5866e2f8380673fa49a%2Ftta_steps_investigation.png?generation=1596451769524367&amp;alt=media</a> =350x150)</p>\n\n<p>As you can see, TTA improves the CV significantly and more iterations are better - my 'simple' model is still benefiting from increasing the number of augmentations at 24 augmentations per image. I expect (but haven't tested) that the more aggressive your augmentations (i.e., the greater the difference between augmented and un-augmented image) the more you will the benefit from increasing the number of TTA iterations. It is likely TTA will always lower the variance of your classifier, which could be a bonus.</p>\n\n<p>If you haven't used TTA yet, it looks like a low-cost way of improving your models performance since you don't need to re-train your models to use it. I recommend looking at <a href=\"/cdeotte\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">notebook</a> if you are interested in implementing it and look forward to hearing your experience of using TTA.</p>",
  "messages": [
    {
      "id": "956247",
      "postDate": "08/03/2020 11:01:06",
      "content": "<p>Many competitors have reported that test time augmentation is effective at increasing CV/LB. Chris Deotte even reported that people were using 25+ augmentation steps, but there isn't a formal topic dedicated to test time augmentations yet. </p>\n\n<p>Firstly, I wanted to share an experiment I did : looking at how the number of TTA changed my mean CV score. The model uses images of dimensions (128, 128), EfnB2 and is evaluated using 5 (triple) stratified folds. I only do basic affine transformations and brightness/saturation changes (e.g., no dropout or hair augmentation is used).\n![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F403021ac2e7cb5866e2f8380673fa49a%2Ftta_steps_investigation.png?generation=1596451769524367&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F403021ac2e7cb5866e2f8380673fa49a%2Ftta_steps_investigation.png?generation=1596451769524367&amp;alt=media</a> =350x150)</p>\n\n<p>As you can see, TTA improves the CV significantly and more iterations are better - my 'simple' model is still benefiting from increasing the number of augmentations at 24 augmentations per image. I expect (but haven't tested) that the more aggressive your augmentations (i.e., the greater the difference between augmented and un-augmented image) the more you will the benefit from increasing the number of TTA iterations. It is likely TTA will always lower the variance of your classifier, which could be a bonus.</p>\n\n<p>If you haven't used TTA yet, it looks like a low-cost way of improving your models performance since you don't need to re-train your models to use it. I recommend looking at <a href=\"/cdeotte\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">notebook</a> if you are interested in implementing it and look forward to hearing your experience of using TTA.</p>",
      "rawMarkdown": "Many competitors have reported that test time augmentation is effective at increasing CV/LB. Chris Deotte even reported that people were using 25+ augmentation steps, but there isn't a formal topic dedicated to test time augmentations yet. \n\nFirstly, I wanted to share an experiment I did : looking at how the number of TTA changed my mean CV score. The model uses images of dimensions (128, 128), EfnB2 and is evaluated using 5 (triple) stratified folds. I only do basic affine transformations and brightness/saturation changes (e.g., no dropout or hair augmentation is used).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F403021ac2e7cb5866e2f8380673fa49a%2Ftta_steps_investigation.png?generation=1596451769524367&amp;alt=media =350x150)\n\nAs you can see, TTA improves the CV significantly and more iterations are better - my 'simple' model is still benefiting from increasing the number of augmentations at 24 augmentations per image. I expect (but haven't tested) that the more aggressive your augmentations (i.e., the greater the difference between augmented and un-augmented image) the more you will the benefit from increasing the number of TTA iterations. It is likely TTA will always lower the variance of your classifier, which could be a bonus.\n\nIf you haven't used TTA yet, it looks like a low-cost way of improving your models performance since you don't need to re-train your models to use it. I recommend looking at @cdeotte [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) if you are interested in implementing it and look forward to hearing your experience of using TTA.",
      "votes": null
    },
    {
      "id": "956434",
      "postDate": "08/03/2020 13:49:07",
      "content": "<p>Great analysis FChmiel. I believe TTA is particularly helpful in this comp.</p>\n\n<p>I assume you compared TTA with the following method. After training your model using my triple stratified notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>, you will have the saved model for each fold. You can then run the notebook a second time and instead of training new models, you can read the saved models from disk. Using the exact same saved models, run the notebook 5 more times and try TTA 5x, 10x, 15x, 20x, 25x etc and compare the difference using the exact same saved models. (The notebook will run fast because you are not training new fold models). Each new notebook commit will report an overall CV score. Is this what you did?</p>",
      "rawMarkdown": "Great analysis FChmiel. I believe TTA is particularly helpful in this comp.\n\nI assume you compared TTA with the following method. After training your model using my triple stratified notebook [here][1], you will have the saved model for each fold. You can then run the notebook a second time and instead of training new models, you can read the saved models from disk. Using the exact same saved models, run the notebook 5 more times and try TTA 5x, 10x, 15x, 20x, 25x etc and compare the difference using the exact same saved models. (The notebook will run fast because you are not training new fold models). Each new notebook commit will report an overall CV score. Is this what you did?\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": null
    },
    {
      "id": "956442",
      "postDate": "08/03/2020 14:00:36",
      "content": "<p>Yes, I think identical to what you describe. I've adapted your script and made an inference only script which I iterated over the number of TTA iterations using the same model (loading in previously stored weights for each fold, which I believe is identical to 'save_model').</p>\n\n<p>As an aside, I can't stress how excellent <a href=\"https://neptune.ai/\">neptune.ai</a> is for experiment management - particularly when running from a script rather than a python notebook.  Makes it so easy to go back, load the model and run more TTA.</p>",
      "rawMarkdown": "Yes, I think identical to what you describe. I've adapted your script and made an inference only script which I iterated over the number of TTA iterations using the same model (loading in previously stored weights for each fold, which I believe is identical to 'save_model').\n\nAs an aside, I can't stress how excellent [neptune.ai](https://neptune.ai/) is for experiment management - particularly when running from a script rather than a python notebook.  Makes it so easy to go back, load the model and run more TTA.",
      "votes": null
    },
    {
      "id": "957095",
      "postDate": "08/04/2020 03:48:13",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    },
    {
      "id": "959788",
      "postDate": "08/05/2020 22:21:26",
      "content": "<p>I have seen the same.  Just adds so much time to the building, I am debating just continuing to test at TTA 12 and then goto TTA 25 during final submissions.</p>",
      "rawMarkdown": "I have seen the same.  Just adds so much time to the building, I am debating just continuing to test at TTA 12 and then goto TTA 25 during final submissions.",
      "votes": null
    },
    {
      "id": "960119",
      "postDate": "08/06/2020 06:45:59",
      "content": "<p>I use TTA = 4 when training the model. Once I have a final model I simply do the inference step with increased TTA steps. You may only gain in time with the larger images.</p>",
      "rawMarkdown": "I use TTA = 4 when training the model. Once I have a final model I simply do the inference step with increased TTA steps. You may only gain in time with the larger images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 956434,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/03/2020 13:49:07",
      "content": "<p>Great analysis FChmiel. I believe TTA is particularly helpful in this comp.</p>\n\n<p>I assume you compared TTA with the following method. After training your model using my triple stratified notebook <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>, you will have the saved model for each fold. You can then run the notebook a second time and instead of training new models, you can read the saved models from disk. Using the exact same saved models, run the notebook 5 more times and try TTA 5x, 10x, 15x, 20x, 25x etc and compare the difference using the exact same saved models. (The notebook will run fast because you are not training new fold models). Each new notebook commit will report an overall CV score. Is this what you did?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 956442,
      "author_name": "fchmiel",
      "author_url": "",
      "post_date": "08/03/2020 14:00:36",
      "content": "<p>Yes, I think identical to what you describe. I've adapted your script and made an inference only script which I iterated over the number of TTA iterations using the same model (loading in previously stored weights for each fold, which I believe is identical to 'save_model').</p>\n\n<p>As an aside, I can't stress how excellent <a href=\"https://neptune.ai/\">neptune.ai</a> is for experiment management - particularly when running from a script rather than a python notebook.  Makes it so easy to go back, load the model and run more TTA.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 957095,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "08/04/2020 03:48:13",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 959788,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "08/05/2020 22:21:26",
      "content": "<p>I have seen the same.  Just adds so much time to the building, I am debating just continuing to test at TTA 12 and then goto TTA 25 during final submissions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 960119,
          "author_name": "fchmiel",
          "author_url": "",
          "post_date": "08/06/2020 06:45:59",
          "content": "<p>I use TTA = 4 when training the model. Once I have a final model I simply do the inference step with increased TTA steps. You may only gain in time with the larger images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "956247": "Many competitors have reported that test time augmentation is effective at increasing CV/LB. Chris Deotte even reported that people were using 25+ augmentation steps, but there isn't a formal topic dedicated to test time augmentations yet. \n\nFirstly, I wanted to share an experiment I did : looking at how the number of TTA changed my mean CV score. The model uses images of dimensions (128, 128), EfnB2 and is evaluated using 5 (triple) stratified folds. I only do basic affine transformations and brightness/saturation changes (e.g., no dropout or hair augmentation is used).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2370491%2F403021ac2e7cb5866e2f8380673fa49a%2Ftta_steps_investigation.png?generation=1596451769524367&amp;alt=media =350x150)\n\nAs you can see, TTA improves the CV significantly and more iterations are better - my 'simple' model is still benefiting from increasing the number of augmentations at 24 augmentations per image. I expect (but haven't tested) that the more aggressive your augmentations (i.e., the greater the difference between augmented and un-augmented image) the more you will the benefit from increasing the number of TTA iterations. It is likely TTA will always lower the variance of your classifier, which could be a bonus.\n\nIf you haven't used TTA yet, it looks like a low-cost way of improving your models performance since you don't need to re-train your models to use it. I recommend looking at @cdeotte [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) if you are interested in implementing it and look forward to hearing your experience of using TTA.",
    "956434": "Great analysis FChmiel. I believe TTA is particularly helpful in this comp.\n\nI assume you compared TTA with the following method. After training your model using my triple stratified notebook [here][1], you will have the saved model for each fold. You can then run the notebook a second time and instead of training new models, you can read the saved models from disk. Using the exact same saved models, run the notebook 5 more times and try TTA 5x, 10x, 15x, 20x, 25x etc and compare the difference using the exact same saved models. (The notebook will run fast because you are not training new fold models). Each new notebook commit will report an overall CV score. Is this what you did?\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "956442": "Yes, I think identical to what you describe. I've adapted your script and made an inference only script which I iterated over the number of TTA iterations using the same model (loading in previously stored weights for each fold, which I believe is identical to 'save_model').\n\nAs an aside, I can't stress how excellent [neptune.ai](https://neptune.ai/) is for experiment management - particularly when running from a script rather than a python notebook.  Makes it so easy to go back, load the model and run more TTA.",
    "957095": "Thank you for sharing",
    "959788": "I have seen the same.  Just adds so much time to the building, I am debating just continuing to test at TTA 12 and then goto TTA 25 during final submissions.",
    "960119": "I use TTA = 4 when training the model. Once I have a final model I simply do the inference step with increased TTA steps. You may only gain in time with the larger images."
  },
  "source": "meta"
}