{
  "id": 132642,
  "title": "Useful Baseline Data Augmentations?",
  "url": "/competitions/bengaliai-cv19/discussion/132642",
  "author_name": "Nicholas Lyu",
  "post_date": "2020-02-27T03:43:56.141000",
  "votes": 18,
  "comment_count": 24,
  "views": 0,
  "content": "<p>It's no news now that this competition might be very appropriately called Bengali.AI augmentation challenge. I wonder what augmentations are you guys using? Just baseline, no cutmix/mixup. Maybe better baseline augmentations is the key to reaching .996+CV</p>\n\n<p>Off the top of my head, I am using <code>ShiftScaleRotate(rotate_limit=10, scale_limit=.1)</code> right now; really haven't tried no-aug or tougher augmentations. Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods. </p>\n\n<p>Any comments welcome, thx :D</p>",
  "messages": [
    {
      "id": 759415,
      "postDate": "2020-02-29T02:23:11.743Z",
      "content": "<p>I would like to disclose that my current LB score is reached using very minimal baseline augmentations (aside cutmixup, augmix, etc). Just want to save some time for you guys...spent weeks experimenting with baseline augmentations nearly getting nowhere.</p>\n\n<p>Here is my augmentation:\n<code>\ntrain_transform = Compose([\n                OneOf([\n                    ShiftScaleRotate(scale_limit=.15, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT),\n                    IAAAffine(shear=20, mode='constant'),\n                    IAAPerspective(),\n                ])\n            ])\n</code></p>\n\n<p>I am worried that rotate=20 is a little too aggressive, but sticking with it. Good luck!</p>",
      "rawMarkdown": "I would like to disclose that my current LB score is reached using very minimal baseline augmentations (aside cutmixup, augmix, etc). Just want to save some time for you guys...spent weeks experimenting with baseline augmentations nearly getting nowhere.\n\nHere is my augmentation:\n```\ntrain_transform = Compose([\n                OneOf([\n                    ShiftScaleRotate(scale_limit=.15, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT),\n                    IAAAffine(shear=20, mode='constant'),\n                    IAAPerspective(),\n                ])\n            ])\n```\n\nI am worried that rotate=20 is a little too aggressive, but sticking with it. Good luck!",
      "votes": 22,
      "replies": [
        {
          "id": 761471,
          "postDate": "2020-03-02T14:57:02.410Z",
          "content": "<p>Thanks for sharing such insight. I was also thinking the same about hard augmentation.</p>",
          "rawMarkdown": "Thanks for sharing such insight. I was also thinking the same about hard augmentation.",
          "votes": 1
        },
        {
          "id": 762779,
          "postDate": "2020-03-03T19:52:55.903Z",
          "content": "<p>Thanks for sharing! It's better than my previous augmentation! I put rotation and shearing together, I guess that's the reason. :)</p>",
          "rawMarkdown": "Thanks for sharing! It's better than my previous augmentation! I put rotation and shearing together, I guess that's the reason. :)"
        },
        {
          "id": 762806,
          "postDate": "2020-03-03T20:24:13.010Z",
          "content": "<p>THANK YOU! you saved me so much time! Just a question tho, did you set shift=0 in ShiftScaleRotate or let it at the default setting? </p>",
          "rawMarkdown": "THANK YOU! you saved me so much time! Just a question tho, did you set shift=0 in ShiftScaleRotate or let it at the default setting? "
        },
        {
          "id": 762963,
          "postDate": "2020-03-04T01:29:22.117Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Default is .0625 I think. It's default</p>",
          "rawMarkdown": "@yannmajewski Default is .0625 I think. It's default",
          "votes": 1
        },
        {
          "id": 766174,
          "postDate": "2020-03-07T19:49:26.453Z",
          "content": "<p>Thank you for an example.\nBut should be careful with this augmentation strategy since sometimes it comes up with a really useless output.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1696976%2F4abbd95d37c19b6ca4efe3a16cf09cc2%2F.png?generation=1583610471068342&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you for an example.\nBut should be careful with this augmentation strategy since sometimes it comes up with a really useless output.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1696976%2F4abbd95d37c19b6ca4efe3a16cf09cc2%2F.png?generation=1583610471068342&amp;alt=media)",
          "votes": 3
        },
        {
          "id": 766312,
          "postDate": "2020-03-08T01:29:29.990Z",
          "content": "<p><a href=\"/nroman\">@nroman</a> Wow. What the hell. I have never even encountered something like this. Are you able to pin down exactly which transformation is responsible for this??</p>",
          "rawMarkdown": "@nroman Wow. What the hell. I have never even encountered something like this. Are you able to pin down exactly which transformation is responsible for this??"
        },
        {
          "id": 766515,
          "postDate": "2020-03-08T09:34:16.877Z",
          "content": "<p>I am sorry, this was my mistake. The problem is that train's shape is (c, h, w) as pytorch expects it. Meanwhile albumentations expects an image to be in a shape of (h, w, c).</p>",
          "rawMarkdown": "I am sorry, this was my mistake. The problem is that train's shape is (c, h, w) as pytorch expects it. Meanwhile albumentations expects an image to be in a shape of (h, w, c)."
        }
      ]
    },
    {
      "id": 757723,
      "postDate": "2020-02-27T03:43:56.143Z",
      "content": "<p>It's no news now that this competition might be very appropriately called Bengali.AI augmentation challenge. I wonder what augmentations are you guys using? Just baseline, no cutmix/mixup. Maybe better baseline augmentations is the key to reaching .996+CV</p>\n\n<p>Off the top of my head, I am using <code>ShiftScaleRotate(rotate_limit=10, scale_limit=.1)</code> right now; really haven't tried no-aug or tougher augmentations. Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods. </p>\n\n<p>Any comments welcome, thx :D</p>",
      "rawMarkdown": "It's no news now that this competition might be very appropriately called Bengali.AI augmentation challenge. I wonder what augmentations are you guys using? Just baseline, no cutmix/mixup. Maybe better baseline augmentations is the key to reaching .996+CV\n\nOff the top of my head, I am using `ShiftScaleRotate(rotate_limit=10, scale_limit=.1)` right now; really haven't tried no-aug or tougher augmentations. Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods. \n\nAny comments welcome, thx :D",
      "votes": 17
    },
    {
      "id": 757772,
      "postDate": "2020-02-27T04:59:30.513Z",
      "content": "<p>you can conduct fast experiments or refer to the reference papers, here are the cases:</p>\n\n<p>[1] baseline result (without any augmentation at all)\n[2] basic augmentation (scale, rotate, shift, other affine warp)\n[3] complex augmentation (include intensity, noise augment, complex distort, etc) without hyperparameter search (i.e random). complex means that that is a risk of over-augmentation, i.e. the over-augmented images may become rubbish image by human eye\n[4] complex augmentation with hyperparameter search  (i.e. autoaugment like RL, PPO, population based training, bayesian hyperopt, ...)\n[5] [2]+ cutout/mixup and other variants\n[6] adversial augmentation</p>\n\n<hr>\n\n<p>this is what i think (based on cifar10/100 paper results and my experiments on cifar10/100 bengali.ai dataset):</p>\n\n<p>assume baseline[1] has accuracy say 95%, [2] gives  +1.5% to 2 %.  [5] will further add about 1% to 1.5%</p>\n\n<hr>\n\n<p>if you compare [2] and [3],  maybe [3] is about 0.5% more if applied correctly. [4] will further improve this \"0.5%\" to about \"1.0%\" ?</p>\n\n<hr>\n\n<p>[6] probably gives 0.5% to 1.5% on top of [2] ? but adversial network is not easily to train</p>",
      "rawMarkdown": "you can conduct fast experiments or refer to the reference papers, here are the cases:\n\n[1] baseline result (without any augmentation at all)\n[2] basic augmentation (scale, rotate, shift, other affine warp)\n[3] complex augmentation (include intensity, noise augment, complex distort, etc) without hyperparameter search (i.e random). complex means that that is a risk of over-augmentation, i.e. the over-augmented images may become rubbish image by human eye\n[4] complex augmentation with hyperparameter search  (i.e. autoaugment like RL, PPO, population based training, bayesian hyperopt, ...)\n[5] [2]+ cutout/mixup and other variants\n[6] adversial augmentation\n\n----\nthis is what i think (based on cifar10/100 paper results and my experiments on cifar10/100 bengali.ai dataset):\n\nassume baseline[1] has accuracy say 95%, [2] gives  +1.5% to 2 %.  [5] will further add about 1% to 1.5%\n\n---\nif you compare [2] and [3],  maybe [3] is about 0.5% more if applied correctly. [4] will further improve this \"0.5%\" to about \"1.0%\" ?\n\n---\n\n[6] probably gives 0.5% to 1.5% on top of [2] ? but adversial network is not easily to train\n\n\n ",
      "votes": 12,
      "replies": [
        {
          "id": 757821,
          "postDate": "2020-02-27T06:11:40.130Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks very much Heng! Helpful as ever! I am learning from your starter kit right now </p>",
          "rawMarkdown": "@hengck23 Thanks very much Heng! Helpful as ever! I am learning from your starter kit right now "
        },
        {
          "id": 760268,
          "postDate": "2020-03-01T03:49:22.940Z",
          "content": "<p>there is yet another option:</p>\n\n<p>[7] extra extreme augmentation, i.e. augmented images may be come distorted and change label (i.e. become other label or \"none of the class\"). augmented samples  are treated as noisy labels. apply weak-supervised or self-supervised techniques for training.</p>",
          "rawMarkdown": "there is yet another option:\n\n[7] extra extreme augmentation, i.e. augmented images may be come distorted and change label (i.e. become other label or \"none of the class\"). augmented samples  are treated as noisy labels. apply weak-supervised or self-supervised techniques for training.",
          "votes": 1
        },
        {
          "id": 761536,
          "postDate": "2020-03-02T16:30:23.213Z",
          "content": "<p>there are yet few other options:</p>\n\n<p>[a] augment in feature space (not at input):  dropout, dropchannel, dropblock, dropconnect, manifold mixup</p>\n\n<p>[b] augmented sample weighing, e.g. give a train sample x0, perform augmentation to get x1,x2,x3 ...\nuse only sample xi with max loss, or median loss, etc ...</p>",
          "rawMarkdown": "there are yet few other options:\n\n[a] augment in feature space (not at input):  dropout, dropchannel, dropblock, dropconnect, manifold mixup\n\n[b] augmented sample weighing, e.g. give a train sample x0, perform augmentation to get x1,x2,x3 ...\nuse only sample xi with max loss, or median loss, etc ..."
        }
      ]
    },
    {
      "id": 757785,
      "postDate": "2020-02-27T05:14:50.070Z",
      "content": "<p>\"Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods.\"</p>\n\n<hr>\n\n<p>this is the \"explore and exploit problem\" over \"short or long term gain\". </p>\n\n<p>focusing on engineering tricks/efforts on one's already familiar methods will save time and efforts and more likely to produce better results in short run.</p>\n\n<p>on the other hands, if we try newer and adventurous methods, it is less likely to give improvement over the short runs because some time is required for the long learning curve. but it may give breakthrough results and maybe useful for future competitions since you are learning a new tool/weapon.</p>\n\n<p>what is the best solution? i would recommend Epsilon Greedy Bandit Solution. Spend most of the time to exploit with epsilon probability for explore</p>",
      "rawMarkdown": "\"Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods.\"\n\n---\n\nthis is the \"explore and exploit problem\" over \"short or long term gain\". \n\nfocusing on engineering tricks/efforts on one's already familiar methods will save time and efforts and more likely to produce better results in short run.\n\non the other hands, if we try newer and adventurous methods, it is less likely to give improvement over the short runs because some time is required for the long learning curve. but it may give breakthrough results and maybe useful for future competitions since you are learning a new tool/weapon.\n\nwhat is the best solution? i would recommend Epsilon Greedy Bandit Solution. Spend most of the time to exploit with epsilon probability for explore\n\n",
      "votes": 5
    },
    {
      "id": 757888,
      "postDate": "2020-02-27T08:03:13.557Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc7d62bdb9761aaf4954bae7d03344c26%2FSelection_061.png?generation=1582790590608413&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc7d62bdb9761aaf4954bae7d03344c26%2FSelection_061.png?generation=1582790590608413&amp;alt=media)\n",
      "votes": 3,
      "replies": [
        {
          "id": 758752,
          "postDate": "2020-02-28T05:26:19.357Z",
          "content": "<p>thanks <a href=\"/hengck23\">@hengck23</a> </p>",
          "rawMarkdown": "thanks @hengck23 "
        }
      ]
    },
    {
      "id": 759210,
      "postDate": "2020-02-28T18:18:24.093Z",
      "content": "<p>\"It's no news now that this competition might be very appropriately called Bengali.AI augmentation\"</p>\n\n<p>it seems that to win the cloud autoML prize, one needs to create more augmented data and upload to google cloud.</p>",
      "rawMarkdown": "\"It's no news now that this competition might be very appropriately called Bengali.AI augmentation\"\n\nit seems that to win the cloud autoML prize, one needs to create more augmented data and upload to google cloud.",
      "votes": 1
    },
    {
      "id": 761469,
      "postDate": "2020-03-02T14:53:33.497Z",
      "content": "<p>为啥我的总是提交失败，您可以解答一下吗<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2F027c480ed20ee5734d5748fabf844555%2FSharedScreenshot.jpg?generation=1583160786636595&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "为啥我的总是提交失败，您可以解答一下吗![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2F027c480ed20ee5734d5748fabf844555%2FSharedScreenshot.jpg?generation=1583160786636595&amp;alt=media)\n",
      "replies": [
        {
          "id": 762786,
          "postDate": "2020-03-03T20:00:21.037Z",
          "content": "<p>Use the whole training set to generate the submission CSV file.  I had the same problem yesterday. It turned out to be the  submission file contained much less result than it should be (around 200k)</p>",
          "rawMarkdown": "Use the whole training set to generate the submission CSV file.  I had the same problem yesterday. It turned out to be the  submission file contained much less result than it should be (around 200k)"
        }
      ]
    },
    {
      "id": 757767,
      "postDate": "2020-02-27T04:51:18.150Z",
      "content": "<p>You once told me you where using cutmix, so you took a step back and you are now using other augmentations than cutmix?</p>",
      "rawMarkdown": "You once told me you where using cutmix, so you took a step back and you are now using other augmentations than cutmix?"
    },
    {
      "id": 757751,
      "postDate": "2020-02-27T04:20:50.163Z",
      "content": "<p>Is your LB score without any Cutout, Cutmix, Mixup, Augmix? You're only using shift scale rotate? That's impressive!</p>",
      "rawMarkdown": "Is your LB score without any Cutout, Cutmix, Mixup, Augmix? You're only using shift scale rotate? That's impressive!",
      "replies": [
        {
          "id": 757819,
          "postDate": "2020-02-27T06:10:16.433Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Unfortunately, no. If only I were able to do this with shiftscalerotate only :(. By the way I have not been able to get augmix working. Any positive results?</p>",
          "rawMarkdown": "@cdeotte Unfortunately, no. If only I were able to do this with shiftscalerotate only :(. By the way I have not been able to get augmix working. Any positive results?"
        },
        {
          "id": 757837,
          "postDate": "2020-02-27T06:46:59.680Z",
          "content": "<p>The concept of AuxMix works on Bengali which is do lots of augmentation randomly. If you read the paper, <a href=\"https://openreview.net/pdf?id=S1gmrxHFvB\">here</a> they use all sorts of augmentation like posterize, equalize, posterize. Bengali is grayscale therefore the color augmentations have not helped me but random combinations of the others is useful.</p>",
          "rawMarkdown": "The concept of AuxMix works on Bengali which is do lots of augmentation randomly. If you read the paper, [here][1] they use all sorts of augmentation like posterize, equalize, posterize. Bengali is grayscale therefore the color augmentations have not helped me but random combinations of the others is useful.\n\n[1]:https://openreview.net/pdf?id=S1gmrxHFvB",
          "votes": 3
        },
        {
          "id": 757923,
          "postDate": "2020-02-27T08:46:56.583Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you very much for the information! I have just started an experiment on augmix. Really not sure if it is going to work. Just wondering, when you visualize the images, do the transformations look aggressive to the eye? It's really up to intuition (and luck) to determine correct augmentation configurations in this comp, I guess</p>",
          "rawMarkdown": "@cdeotte Thank you very much for the information! I have just started an experiment on augmix. Really not sure if it is going to work. Just wondering, when you visualize the images, do the transformations look aggressive to the eye? It's really up to intuition (and luck) to determine correct augmentation configurations in this comp, I guess"
        }
      ]
    },
    {
      "id": 757780,
      "postDate": "2020-02-27T05:09:11.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 759415,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-02-29T02:23:11.743000",
      "content": "<p>I would like to disclose that my current LB score is reached using very minimal baseline augmentations (aside cutmixup, augmix, etc). Just want to save some time for you guys...spent weeks experimenting with baseline augmentations nearly getting nowhere.</p>\n\n<p>Here is my augmentation:\n<code>\ntrain_transform = Compose([\n                OneOf([\n                    ShiftScaleRotate(scale_limit=.15, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT),\n                    IAAAffine(shear=20, mode='constant'),\n                    IAAPerspective(),\n                ])\n            ])\n</code></p>\n\n<p>I am worried that rotate=20 is a little too aggressive, but sticking with it. Good luck!</p>",
      "votes": 22,
      "replies": [
        {
          "id": 761471,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-02T14:57:02.410000",
          "content": "<p>Thanks for sharing such insight. I was also thinking the same about hard augmentation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762779,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-03-03T19:52:55.903000",
          "content": "<p>Thanks for sharing! It's better than my previous augmentation! I put rotation and shearing together, I guess that's the reason. :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762806,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-03-03T20:24:13.010000",
          "content": "<p>THANK YOU! you saved me so much time! Just a question tho, did you set shift=0 in ShiftScaleRotate or let it at the default setting? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762963,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-04T01:29:22.117000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Default is .0625 I think. It's default</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766174,
          "author_name": "Roman",
          "author_url": "",
          "post_date": "2020-03-07T19:49:26.453000",
          "content": "<p>Thank you for an example.\nBut should be careful with this augmentation strategy since sometimes it comes up with a really useless output.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1696976%2F4abbd95d37c19b6ca4efe3a16cf09cc2%2F.png?generation=1583610471068342&amp;alt=media\" alt=\"\"></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 766312,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-08T01:29:29.990000",
          "content": "<p><a href=\"/nroman\">@nroman</a> Wow. What the hell. I have never even encountered something like this. Are you able to pin down exactly which transformation is responsible for this??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766515,
          "author_name": "Roman",
          "author_url": "",
          "post_date": "2020-03-08T09:34:16.877000",
          "content": "<p>I am sorry, this was my mistake. The problem is that train's shape is (c, h, w) as pytorch expects it. Meanwhile albumentations expects an image to be in a shape of (h, w, c).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757772,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-27T04:59:30.513000",
      "content": "<p>you can conduct fast experiments or refer to the reference papers, here are the cases:</p>\n\n<p>[1] baseline result (without any augmentation at all)\n[2] basic augmentation (scale, rotate, shift, other affine warp)\n[3] complex augmentation (include intensity, noise augment, complex distort, etc) without hyperparameter search (i.e random). complex means that that is a risk of over-augmentation, i.e. the over-augmented images may become rubbish image by human eye\n[4] complex augmentation with hyperparameter search  (i.e. autoaugment like RL, PPO, population based training, bayesian hyperopt, ...)\n[5] [2]+ cutout/mixup and other variants\n[6] adversial augmentation</p>\n\n<hr>\n\n<p>this is what i think (based on cifar10/100 paper results and my experiments on cifar10/100 bengali.ai dataset):</p>\n\n<p>assume baseline[1] has accuracy say 95%, [2] gives  +1.5% to 2 %.  [5] will further add about 1% to 1.5%</p>\n\n<hr>\n\n<p>if you compare [2] and [3],  maybe [3] is about 0.5% more if applied correctly. [4] will further improve this \"0.5%\" to about \"1.0%\" ?</p>\n\n<hr>\n\n<p>[6] probably gives 0.5% to 1.5% on top of [2] ? but adversial network is not easily to train</p>",
      "votes": 12,
      "replies": [
        {
          "id": 757821,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-27T06:11:40.130000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks very much Heng! Helpful as ever! I am learning from your starter kit right now </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 760268,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-01T03:49:22.940000",
          "content": "<p>there is yet another option:</p>\n\n<p>[7] extra extreme augmentation, i.e. augmented images may be come distorted and change label (i.e. become other label or \"none of the class\"). augmented samples  are treated as noisy labels. apply weak-supervised or self-supervised techniques for training.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 761536,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-02T16:30:23.213000",
          "content": "<p>there are yet few other options:</p>\n\n<p>[a] augment in feature space (not at input):  dropout, dropchannel, dropblock, dropconnect, manifold mixup</p>\n\n<p>[b] augmented sample weighing, e.g. give a train sample x0, perform augmentation to get x1,x2,x3 ...\nuse only sample xi with max loss, or median loss, etc ...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757785,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-27T05:14:50.070000",
      "content": "<p>\"Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods.\"</p>\n\n<hr>\n\n<p>this is the \"explore and exploit problem\" over \"short or long term gain\". </p>\n\n<p>focusing on engineering tricks/efforts on one's already familiar methods will save time and efforts and more likely to produce better results in short run.</p>\n\n<p>on the other hands, if we try newer and adventurous methods, it is less likely to give improvement over the short runs because some time is required for the long learning curve. but it may give breakthrough results and maybe useful for future competitions since you are learning a new tool/weapon.</p>\n\n<p>what is the best solution? i would recommend Epsilon Greedy Bandit Solution. Spend most of the time to exploit with epsilon probability for explore</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 757888,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-27T08:03:13.557000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc7d62bdb9761aaf4954bae7d03344c26%2FSelection_061.png?generation=1582790590608413&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 758752,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-02-28T05:26:19.357000",
          "content": "<p>thanks <a href=\"/hengck23\">@hengck23</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 759210,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-28T18:18:24.093000",
      "content": "<p>\"It's no news now that this competition might be very appropriately called Bengali.AI augmentation\"</p>\n\n<p>it seems that to win the cloud autoML prize, one needs to create more augmented data and upload to google cloud.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 761469,
      "author_name": "Email:986285044@qq.com",
      "author_url": "",
      "post_date": "2020-03-02T14:53:33.497000",
      "content": "<p>为啥我的总是提交失败，您可以解答一下吗<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2F027c480ed20ee5734d5748fabf844555%2FSharedScreenshot.jpg?generation=1583160786636595&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 762786,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-03-03T20:00:21.037000",
          "content": "<p>Use the whole training set to generate the submission CSV file.  I had the same problem yesterday. It turned out to be the  submission file contained much less result than it should be (around 200k)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757767,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-02-27T04:51:18.150000",
      "content": "<p>You once told me you where using cutmix, so you took a step back and you are now using other augmentations than cutmix?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 757751,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-02-27T04:20:50.163000",
      "content": "<p>Is your LB score without any Cutout, Cutmix, Mixup, Augmix? You're only using shift scale rotate? That's impressive!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 757819,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-27T06:10:16.433000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Unfortunately, no. If only I were able to do this with shiftscalerotate only :(. By the way I have not been able to get augmix working. Any positive results?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 757837,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-02-27T06:46:59.680000",
          "content": "<p>The concept of AuxMix works on Bengali which is do lots of augmentation randomly. If you read the paper, <a href=\"https://openreview.net/pdf?id=S1gmrxHFvB\">here</a> they use all sorts of augmentation like posterize, equalize, posterize. Bengali is grayscale therefore the color augmentations have not helped me but random combinations of the others is useful.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 757923,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-27T08:46:56.583000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you very much for the information! I have just started an experiment on augmix. Really not sure if it is going to work. Just wondering, when you visualize the images, do the transformations look aggressive to the eye? It's really up to intuition (and luck) to determine correct augmentation configurations in this comp, I guess</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757780,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-27T05:09:11.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "759415": "I would like to disclose that my current LB score is reached using very minimal baseline augmentations (aside cutmixup, augmix, etc). Just want to save some time for you guys...spent weeks experimenting with baseline augmentations nearly getting nowhere.\n\nHere is my augmentation:\n```\ntrain_transform = Compose([\n                OneOf([\n                    ShiftScaleRotate(scale_limit=.15, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT),\n                    IAAAffine(shear=20, mode='constant'),\n                    IAAPerspective(),\n                ])\n            ])\n```\n\nI am worried that rotate=20 is a little too aggressive, but sticking with it. Good luck!",
    "757723": "It's no news now that this competition might be very appropriately called Bengali.AI augmentation challenge. I wonder what augmentations are you guys using? Just baseline, no cutmix/mixup. Maybe better baseline augmentations is the key to reaching .996+CV\n\nOff the top of my head, I am using `ShiftScaleRotate(rotate_limit=10, scale_limit=.1)` right now; really haven't tried no-aug or tougher augmentations. Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods. \n\nAny comments welcome, thx :D",
    "757772": "you can conduct fast experiments or refer to the reference papers, here are the cases:\n\n[1] baseline result (without any augmentation at all)\n[2] basic augmentation (scale, rotate, shift, other affine warp)\n[3] complex augmentation (include intensity, noise augment, complex distort, etc) without hyperparameter search (i.e random). complex means that that is a risk of over-augmentation, i.e. the over-augmented images may become rubbish image by human eye\n[4] complex augmentation with hyperparameter search  (i.e. autoaugment like RL, PPO, population based training, bayesian hyperopt, ...)\n[5] [2]+ cutout/mixup and other variants\n[6] adversial augmentation\n\n----\nthis is what i think (based on cifar10/100 paper results and my experiments on cifar10/100 bengali.ai dataset):\n\nassume baseline[1] has accuracy say 95%, [2] gives  +1.5% to 2 %.  [5] will further add about 1% to 1.5%\n\n---\nif you compare [2] and [3],  maybe [3] is about 0.5% more if applied correctly. [4] will further improve this \"0.5%\" to about \"1.0%\" ?\n\n---\n\n[6] probably gives 0.5% to 1.5% on top of [2] ? but adversial network is not easily to train\n\n\n ",
    "757785": "\"Doing experiments right now, but time and money is limited so I will not try automated augmentation searching methods.\"\n\n---\n\nthis is the \"explore and exploit problem\" over \"short or long term gain\". \n\nfocusing on engineering tricks/efforts on one's already familiar methods will save time and efforts and more likely to produce better results in short run.\n\non the other hands, if we try newer and adventurous methods, it is less likely to give improvement over the short runs because some time is required for the long learning curve. but it may give breakthrough results and maybe useful for future competitions since you are learning a new tool/weapon.\n\nwhat is the best solution? i would recommend Epsilon Greedy Bandit Solution. Spend most of the time to exploit with epsilon probability for explore\n\n",
    "757888": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc7d62bdb9761aaf4954bae7d03344c26%2FSelection_061.png?generation=1582790590608413&amp;alt=media)\n",
    "759210": "\"It's no news now that this competition might be very appropriately called Bengali.AI augmentation\"\n\nit seems that to win the cloud autoML prize, one needs to create more augmented data and upload to google cloud.",
    "761469": "为啥我的总是提交失败，您可以解答一下吗![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2F027c480ed20ee5734d5748fabf844555%2FSharedScreenshot.jpg?generation=1583160786636595&amp;alt=media)\n",
    "757767": "You once told me you where using cutmix, so you took a step back and you are now using other augmentations than cutmix?",
    "757751": "Is your LB score without any Cutout, Cutmix, Mixup, Augmix? You're only using shift scale rotate? That's impressive!",
    "757780": ""
  }
}