{
  "id": 216275,
  "title": "Measuring the destructiveness of augmentations",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/216275",
  "author_name": "",
  "post_date": "2021-02-02T08:42:48.096872300Z",
  "votes": 59,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Like with most image problems augmentations are quite important to aid in generalization of the model. Especially given that we only have 30k images it is beneficial to help the model learn general characteristics rather than memorizing the inputs. </p>\n<p>Some reasonable options have been laid out in the starter kernels, but I thought it would be beneficial to measure the effect of applying them by benchmarking against the validation set. What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time. </p>\n<p>This can accomplish 2 things. It can tell us which augmentations are destroying important information and it can also possibly give us candidates for test time augmentation. We want predictions that are still of high quality but have lower correlation with the non-TTA predictions. </p>\n<p>Here is the list of the augmentations explored. Taken from public kernels. </p>\n<p><img src=\"https://i.imgur.com/hdof2bl.png\" alt=\"\"></p>\n<p>And here is their respective performance in terms of AUC</p>\n<p><img src=\"https://i.imgur.com/PDTlf2z.png\" alt=\"\"></p>\n<p>And loss</p>\n<p><img src=\"https://i.imgur.com/0NArcgy.png\" alt=\"\"></p>\n<p>A few of them seem to be particularly destructive. Grid distortion, gaussian noise, median blur, gaussian blur and to a lesser extent several others seem to really hurt validation performance. This may have to do with these augmentations not being seen often enough during training to make the model durable to these kinds of alterations or it could be that these augmentations are too aggressive and simply remove the required information to make this judgement. </p>",
  "messages": [
    {
      "id": "1181995",
      "postDate": "02/02/2021 08:42:48",
      "content": "<p>Like with most image problems augmentations are quite important to aid in generalization of the model. Especially given that we only have 30k images it is beneficial to help the model learn general characteristics rather than memorizing the inputs. </p>\n<p>Some reasonable options have been laid out in the starter kernels, but I thought it would be beneficial to measure the effect of applying them by benchmarking against the validation set. What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time. </p>\n<p>This can accomplish 2 things. It can tell us which augmentations are destroying important information and it can also possibly give us candidates for test time augmentation. We want predictions that are still of high quality but have lower correlation with the non-TTA predictions. </p>\n<p>Here is the list of the augmentations explored. Taken from public kernels. </p>\n<p><img src=\"https://i.imgur.com/hdof2bl.png\" alt=\"\"></p>\n<p>And here is their respective performance in terms of AUC</p>\n<p><img src=\"https://i.imgur.com/PDTlf2z.png\" alt=\"\"></p>\n<p>And loss</p>\n<p><img src=\"https://i.imgur.com/0NArcgy.png\" alt=\"\"></p>\n<p>A few of them seem to be particularly destructive. Grid distortion, gaussian noise, median blur, gaussian blur and to a lesser extent several others seem to really hurt validation performance. This may have to do with these augmentations not being seen often enough during training to make the model durable to these kinds of alterations or it could be that these augmentations are too aggressive and simply remove the required information to make this judgement. </p>",
      "rawMarkdown": "Like with most image problems augmentations are quite important to aid in generalization of the model. Especially given that we only have 30k images it is beneficial to help the model learn general characteristics rather than memorizing the inputs. \n\nSome reasonable options have been laid out in the starter kernels, but I thought it would be beneficial to measure the effect of applying them by benchmarking against the validation set. What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time. \n\nThis can accomplish 2 things. It can tell us which augmentations are destroying important information and it can also possibly give us candidates for test time augmentation. We want predictions that are still of high quality but have lower correlation with the non-TTA predictions. \n\nHere is the list of the augmentations explored. Taken from public kernels. \n\n![](https://i.imgur.com/hdof2bl.png)\n\nAnd here is their respective performance in terms of AUC\n\n![](https://i.imgur.com/PDTlf2z.png)\n\nAnd loss\n\n![](https://i.imgur.com/0NArcgy.png)\n\nA few of them seem to be particularly destructive. Grid distortion, gaussian noise, median blur, gaussian blur and to a lesser extent several others seem to really hurt validation performance. This may have to do with these augmentations not being seen often enough during training to make the model durable to these kinds of alterations or it could be that these augmentations are too aggressive and simply remove the required information to make this judgement.",
      "votes": null
    },
    {
      "id": "1182021",
      "postDate": "02/02/2021 09:00:29",
      "content": "<p>Great work!Thanks!</p>",
      "rawMarkdown": "Great work!Thanks!",
      "votes": null
    },
    {
      "id": "1182098",
      "postDate": "02/02/2021 09:51:42",
      "content": "<p>Great Work.<br>\nWhy have you ignored RandomResizeCrop? Which is a popular augmentation in many of the public notebooks i have seen.</p>",
      "rawMarkdown": "Great Work.\nWhy have you ignored RandomResizeCrop? Which is a popular augmentation in many of the public notebooks i have seen.",
      "votes": null
    },
    {
      "id": "1182116",
      "postDate": "02/02/2021 10:03:49",
      "content": "<p>That was just an oversight when copying over the augmentations. I ran that one and it performed very well 0.939 auc. I wouldn't expect it to be particularly problematic because it is one of the operations that is applied 100% of the time during training</p>",
      "rawMarkdown": "That was just an oversight when copying over the augmentations. I ran that one and it performed very well 0.939 auc. I wouldn't expect it to be particularly problematic because it is one of the operations that is applied 100% of the time during training",
      "votes": null
    },
    {
      "id": "1182306",
      "postDate": "02/02/2021 12:04:55",
      "content": "<p>Cool. Thanks for the reply</p>",
      "rawMarkdown": "Cool. Thanks for the reply",
      "votes": null
    },
    {
      "id": "1184971",
      "postDate": "02/03/2021 23:18:03",
      "content": "<blockquote>\n  <p>What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time.</p>\n</blockquote>\n<p>When you say 'take a single model from a single fold', you mean to take a model trained on 4/5th of the data, and use the remaining 1/5th of data as the validation set?</p>\n<p>Was the original model trained on augmented data? I would assume that if the model is trained on the same type of augmentations it will see during validation, that will help avoid 'destroying important information'.</p>",
      "rawMarkdown": "> What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time.\n\nWhen you say 'take a single model from a single fold', you mean to take a model trained on 4/5th of the data, and use the remaining 1/5th of data as the validation set?\n\nWas the original model trained on augmented data? I would assume that if the model is trained on the same type of augmentations it will see during validation, that will help avoid 'destroying important information'.",
      "votes": null
    },
    {
      "id": "1185067",
      "postDate": "02/04/2021 00:29:17",
      "content": "<p>Yes, that is what I did. And yes the model was trained on these augmentations.</p>\n<p>That is what is so telling that information has been destroyed or at least the augmented version is teaching something at odds with the reality of the images. </p>",
      "rawMarkdown": "Yes, that is what I did. And yes the model was trained on these augmentations.\n\nThat is what is so telling that information has been destroyed or at least the augmented version is teaching something at odds with the reality of the images.",
      "votes": null
    },
    {
      "id": "1194893",
      "postDate": "02/10/2021 12:51:21",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Cool viz. It can be clearly seen that GridDistortion and MedianBlur have the lowest AUC score relative to the other augmentations. And that effect is magnified in the loss graph. What loss function did you use here?</p>",
      "rawMarkdown": "ryches Cool viz. It can be clearly seen that GridDistortion and MedianBlur have the lowest AUC score relative to the other augmentations. And that effect is magnified in the loss graph. What loss function did you use here?",
      "votes": null
    },
    {
      "id": "1195237",
      "postDate": "02/10/2021 16:19:43",
      "content": "<p>Binary cross entropy was the loss I used</p>",
      "rawMarkdown": "Binary cross entropy was the loss I used",
      "votes": null
    },
    {
      "id": "1197079",
      "postDate": "02/11/2021 23:21:23",
      "content": "<p>Thanks for sharing. Really neat idea for quickly testing augmentations. Will definitely be adding this to my workflows :)</p>",
      "rawMarkdown": "Thanks for sharing. Really neat idea for quickly testing augmentations. Will definitely be adding this to my workflows :)",
      "votes": null
    },
    {
      "id": "1210884",
      "postDate": "02/19/2021 19:39:30",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> !! I am new in Kaggle. Can you provide me code how you calculated AUC for each augmentations.<br>\nThanks.</p>",
      "rawMarkdown": "Thanks @ryches !! I am new in Kaggle. Can you provide me code how you calculated AUC for each augmentations.\nThanks.",
      "votes": null
    },
    {
      "id": "1210885",
      "postDate": "02/19/2021 19:41:07",
      "content": "<p>Haaa😄 Cool</p>",
      "rawMarkdown": "Haaa😄 Cool",
      "votes": null
    },
    {
      "id": "1211640",
      "postDate": "02/20/2021 12:00:37",
      "content": "<p>i add to add that sometimes destructive augmentation is useful too.<br>\ni read a paper (but forget the title for the time being) that you can treat extreme (and bad) augmentation as unlabelled or noisy labelled. Then you can use pseudo labels on them. </p>",
      "rawMarkdown": "i add to add that sometimes destructive augmentation is useful too.\ni read a paper (but forget the title for the time being) that you can treat extreme (and bad) augmentation as unlabelled or noisy labelled. Then you can use pseudo labels on them.",
      "votes": null
    },
    {
      "id": "1235746",
      "postDate": "03/12/2021 12:46:26",
      "content": "<p>follow up on this:<br>\n<a href=\"https://www.youtube.com/watch?v=K-1mN2mz66k\" target=\"_blank\">https://www.youtube.com/watch?v=K-1mN2mz66k</a><br>\nNegative Data Augmentation</p>\n<p>\" Negative Data Augmentation, a strategy for using label-corrupting, rather than label-preserving transformations in Deep Learning.\"</p>\n<p>this uses  destructive augmentation</p>",
      "rawMarkdown": "follow up on this:\nhttps://www.youtube.com/watch?v=K-1mN2mz66k\nNegative Data Augmentation\n\n\" Negative Data Augmentation, a strategy for using label-corrupting, rather than label-preserving transformations in Deep Learning.\"\n\nthis uses  destructive augmentation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1182021,
      "author_name": "bcwang",
      "author_url": "",
      "post_date": "02/02/2021 09:00:29",
      "content": "<p>Great work!Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1182098,
      "author_name": "aravindpadman",
      "author_url": "",
      "post_date": "02/02/2021 09:51:42",
      "content": "<p>Great Work.<br>\nWhy have you ignored RandomResizeCrop? Which is a popular augmentation in many of the public notebooks i have seen.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1182116,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/02/2021 10:03:49",
          "content": "<p>That was just an oversight when copying over the augmentations. I ran that one and it performed very well 0.939 auc. I wouldn't expect it to be particularly problematic because it is one of the operations that is applied 100% of the time during training</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1182306,
          "author_name": "aravindpadman",
          "author_url": "",
          "post_date": "02/02/2021 12:04:55",
          "content": "<p>Cool. Thanks for the reply</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210885,
          "author_name": "adg1822",
          "author_url": "",
          "post_date": "02/19/2021 19:41:07",
          "content": "<p>Haaa😄 Cool</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1184971,
      "author_name": "gianlucarossi",
      "author_url": "",
      "post_date": "02/03/2021 23:18:03",
      "content": "<blockquote>\n  <p>What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time.</p>\n</blockquote>\n<p>When you say 'take a single model from a single fold', you mean to take a model trained on 4/5th of the data, and use the remaining 1/5th of data as the validation set?</p>\n<p>Was the original model trained on augmented data? I would assume that if the model is trained on the same type of augmentations it will see during validation, that will help avoid 'destroying important information'.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1185067,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/04/2021 00:29:17",
          "content": "<p>Yes, that is what I did. And yes the model was trained on these augmentations.</p>\n<p>That is what is so telling that information has been destroyed or at least the augmented version is teaching something at odds with the reality of the images. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1194893,
      "author_name": "prvnkmr",
      "author_url": "",
      "post_date": "02/10/2021 12:51:21",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Cool viz. It can be clearly seen that GridDistortion and MedianBlur have the lowest AUC score relative to the other augmentations. And that effect is magnified in the loss graph. What loss function did you use here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1195237,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/10/2021 16:19:43",
          "content": "<p>Binary cross entropy was the loss I used</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1197079,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "02/11/2021 23:21:23",
      "content": "<p>Thanks for sharing. Really neat idea for quickly testing augmentations. Will definitely be adding this to my workflows :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210884,
      "author_name": "adg1822",
      "author_url": "",
      "post_date": "02/19/2021 19:39:30",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> !! I am new in Kaggle. Can you provide me code how you calculated AUC for each augmentations.<br>\nThanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211640,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/20/2021 12:00:37",
      "content": "<p>i add to add that sometimes destructive augmentation is useful too.<br>\ni read a paper (but forget the title for the time being) that you can treat extreme (and bad) augmentation as unlabelled or noisy labelled. Then you can use pseudo labels on them. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1235746,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/12/2021 12:46:26",
          "content": "<p>follow up on this:<br>\n<a href=\"https://www.youtube.com/watch?v=K-1mN2mz66k\" target=\"_blank\">https://www.youtube.com/watch?v=K-1mN2mz66k</a><br>\nNegative Data Augmentation</p>\n<p>\" Negative Data Augmentation, a strategy for using label-corrupting, rather than label-preserving transformations in Deep Learning.\"</p>\n<p>this uses  destructive augmentation</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1181995": "Like with most image problems augmentations are quite important to aid in generalization of the model. Especially given that we only have 30k images it is beneficial to help the model learn general characteristics rather than memorizing the inputs. \n\nSome reasonable options have been laid out in the starter kernels, but I thought it would be beneficial to measure the effect of applying them by benchmarking against the validation set. What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time. \n\nThis can accomplish 2 things. It can tell us which augmentations are destroying important information and it can also possibly give us candidates for test time augmentation. We want predictions that are still of high quality but have lower correlation with the non-TTA predictions. \n\nHere is the list of the augmentations explored. Taken from public kernels. \n\n![](https://i.imgur.com/hdof2bl.png)\n\nAnd here is their respective performance in terms of AUC\n\n![](https://i.imgur.com/PDTlf2z.png)\n\nAnd loss\n\n![](https://i.imgur.com/0NArcgy.png)\n\nA few of them seem to be particularly destructive. Grid distortion, gaussian noise, median blur, gaussian blur and to a lesser extent several others seem to really hurt validation performance. This may have to do with these augmentations not being seen often enough during training to make the model durable to these kinds of alterations or it could be that these augmentations are too aggressive and simply remove the required information to make this judgement.",
    "1182021": "Great work!Thanks!",
    "1182098": "Great Work.\nWhy have you ignored RandomResizeCrop? Which is a popular augmentation in many of the public notebooks i have seen.",
    "1182116": "That was just an oversight when copying over the augmentations. I ran that one and it performed very well 0.939 auc. I wouldn't expect it to be particularly problematic because it is one of the operations that is applied 100% of the time during training",
    "1182306": "Cool. Thanks for the reply",
    "1184971": "> What I did was take a single model from a single fold and do inference on its validation set with each of the augmentations individually turned on and applied 100% of the time.\n\nWhen you say 'take a single model from a single fold', you mean to take a model trained on 4/5th of the data, and use the remaining 1/5th of data as the validation set?\n\nWas the original model trained on augmented data? I would assume that if the model is trained on the same type of augmentations it will see during validation, that will help avoid 'destroying important information'.",
    "1185067": "Yes, that is what I did. And yes the model was trained on these augmentations.\n\nThat is what is so telling that information has been destroyed or at least the augmented version is teaching something at odds with the reality of the images.",
    "1194893": "ryches Cool viz. It can be clearly seen that GridDistortion and MedianBlur have the lowest AUC score relative to the other augmentations. And that effect is magnified in the loss graph. What loss function did you use here?",
    "1195237": "Binary cross entropy was the loss I used",
    "1197079": "Thanks for sharing. Really neat idea for quickly testing augmentations. Will definitely be adding this to my workflows :)",
    "1210884": "Thanks @ryches !! I am new in Kaggle. Can you provide me code how you calculated AUC for each augmentations.\nThanks.",
    "1210885": "Haaa😄 Cool",
    "1211640": "i add to add that sometimes destructive augmentation is useful too.\ni read a paper (but forget the title for the time being) that you can treat extreme (and bad) augmentation as unlabelled or noisy labelled. Then you can use pseudo labels on them.",
    "1235746": "follow up on this:\nhttps://www.youtube.com/watch?v=K-1mN2mz66k\nNegative Data Augmentation\n\n\" Negative Data Augmentation, a strategy for using label-corrupting, rather than label-preserving transformations in Deep Learning.\"\n\nthis uses  destructive augmentation"
  },
  "source": "meta"
}