{
  "id": 132228,
  "title": "Which kind of Augmentation is useful in this competition？",
  "url": "/competitions/bengaliai-cv19/discussion/132228",
  "author_name": "",
  "post_date": "2020-02-25T01:51:45.437196200Z",
  "votes": 19,
  "comment_count": 45,
  "views": 0,
  "content": "<p>Update 2020/2/27\nHey guys,recently I do some experiment to check the best performence about data augmentations.\nIn my experiment，the setting is:</p>\n\n<p>*dataset:top fifth part\nmodel:resnet34\nopt:Over9000\nlr_schedule:fastai's OneCycle\nbasic augmentation:rotation+light+wrap*</p>\n\n<p>And the result looks like these:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2793161%2F543e7a3121eef98f37db96bf5e3bd153%2F20200227101845.png?generation=1582770008940180&amp;alt=media\" alt=\"\"></p>\n\n<p>So,I think the different combinations of** Cutmix,Mixup and Cutout** is much better for this dataset.</p>\n\n<p>Origin 2020/2/25\nHey guys，I am wondering which kind of data augmentation is really helpful in this dataset.\nI have already tired the basic Image transform(flip,rotation...), Gridmask, Augmix , Mixup and Cutmix on a small model(resnet34).And the score is:\nbaseline 0.947\nrotation+light+wrap 0.957\nMixup 0.961\nGridmask 0.944\nAugmix 0.946\nCutmix 0.95\nAs you can see only basic transform with out flip and Mixup help me improve the CV score.\nThe combination of basic transform and Mixup 0.969\nAnd I can't reach any more.\nCan any one share your experiment score on your CV?</p>",
  "messages": [
    {
      "id": "755661",
      "postDate": "02/25/2020 01:51:45",
      "content": "<p>Update 2020/2/27\nHey guys,recently I do some experiment to check the best performence about data augmentations.\nIn my experiment，the setting is:</p>\n\n<p>*dataset:top fifth part\nmodel:resnet34\nopt:Over9000\nlr_schedule:fastai's OneCycle\nbasic augmentation:rotation+light+wrap*</p>\n\n<p>And the result looks like these:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2793161%2F543e7a3121eef98f37db96bf5e3bd153%2F20200227101845.png?generation=1582770008940180&amp;alt=media\" alt=\"\"></p>\n\n<p>So,I think the different combinations of** Cutmix,Mixup and Cutout** is much better for this dataset.</p>\n\n<p>Origin 2020/2/25\nHey guys，I am wondering which kind of data augmentation is really helpful in this dataset.\nI have already tired the basic Image transform(flip,rotation...), Gridmask, Augmix , Mixup and Cutmix on a small model(resnet34).And the score is:\nbaseline 0.947\nrotation+light+wrap 0.957\nMixup 0.961\nGridmask 0.944\nAugmix 0.946\nCutmix 0.95\nAs you can see only basic transform with out flip and Mixup help me improve the CV score.\nThe combination of basic transform and Mixup 0.969\nAnd I can't reach any more.\nCan any one share your experiment score on your CV?</p>",
      "rawMarkdown": "Update 2020/2/27\nHey guys,recently I do some experiment to check the best performence about data augmentations.\nIn my experiment，the setting is:\n\n*dataset:top fifth part\nmodel:resnet34\nopt:Over9000\nlr_schedule:fastai's OneCycle\nbasic augmentation:rotation+light+wrap*\n\nAnd the result looks like these:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2793161%2F543e7a3121eef98f37db96bf5e3bd153%2F20200227101845.png?generation=1582770008940180&amp;alt=media)\n\nSo,I think the different combinations of** Cutmix,Mixup and Cutout** is much better for this dataset.\n\n\nOrigin 2020/2/25\nHey guys，I am wondering which kind of data augmentation is really helpful in this dataset.\nI have already tired the basic Image transform(flip,rotation...), Gridmask, Augmix , Mixup and Cutmix on a small model(resnet34).And the score is:\nbaseline 0.947\nrotation+light+wrap 0.957\nMixup 0.961\nGridmask 0.944\nAugmix 0.946\nCutmix 0.95\nAs you can see only basic transform with out flip and Mixup help me improve the CV score.\nThe combination of basic transform and Mixup 0.969\nAnd I can't reach any more.\nCan any one share your experiment score on your CV?",
      "votes": null
    },
    {
      "id": "755673",
      "postDate": "02/25/2020 02:31:56",
      "content": "<p>I have a similar problem</p>",
      "rawMarkdown": "I have a similar problem",
      "votes": null
    },
    {
      "id": "755675",
      "postDate": "02/25/2020 02:34:54",
      "content": "<p>You can give a try on different combination.</p>",
      "rawMarkdown": "You can give a try on different combination.",
      "votes": null
    },
    {
      "id": "755676",
      "postDate": "02/25/2020 02:35:29",
      "content": "<p>For me Cutout works best!</p>",
      "rawMarkdown": "For me Cutout works best!",
      "votes": null
    },
    {
      "id": "755706",
      "postDate": "02/25/2020 03:34:32",
      "content": "<p>I  also have a similar problem</p>",
      "rawMarkdown": "I  also have a similar problem",
      "votes": null
    },
    {
      "id": "755709",
      "postDate": "02/25/2020 03:41:50",
      "content": "<p>can we talk in private？</p>",
      "rawMarkdown": "can we talk in private？",
      "votes": null
    },
    {
      "id": "756114",
      "postDate": "02/25/2020 12:30:24",
      "content": "<p>Try other models like ResNext50, DenseNet, etc.</p>",
      "rawMarkdown": "Try other models like ResNext50, DenseNet, etc.",
      "votes": null
    },
    {
      "id": "756226",
      "postDate": "02/25/2020 14:27:20",
      "content": "<p>Do you include mixup and cutmix along with cutout?</p>",
      "rawMarkdown": "Do you include mixup and cutmix along with cutout?",
      "votes": null
    },
    {
      "id": "756799",
      "postDate": "02/26/2020 05:01:10",
      "content": "<p>Try gridmask + cutout in combination</p>",
      "rawMarkdown": "Try gridmask + cutout in combination",
      "votes": null
    },
    {
      "id": "756816",
      "postDate": "02/26/2020 05:37:07",
      "content": "<p>o_O</p>",
      "rawMarkdown": "o_O",
      "votes": null
    },
    {
      "id": "756965",
      "postDate": "02/26/2020 09:22:58",
      "content": "<p>Hey guys,Thanks for your advice.I do find cutout is really a good augmentation in this dataset.\nI am making a Comparative Test.Later I will upload a chart about my result</p>",
      "rawMarkdown": "Hey guys,Thanks for your advice.I do find cutout is really a good augmentation in this dataset.\nI am making a Comparative Test.Later I will upload a chart about my result",
      "votes": null
    },
    {
      "id": "757525",
      "postDate": "02/26/2020 21:39:01",
      "content": "<p>Just by using a solid baseline model I have 0.9724. I have not added any augmentation yet.\nWhen I'am unable to improve it anymore I will likely try either Mixup or GridMask.</p>\n\n<p>Mixup I used in 1 or 2 earlier competition and works Ok. GridMask I tried but it decreased my score...so likely not the right configuration..I think it will be interresting enough to give it a few more tries.</p>",
      "rawMarkdown": "Just by using a solid baseline model I have 0.9724. I have not added any augmentation yet.\nWhen I'am unable to improve it anymore I will likely try either Mixup or GridMask.\n\nMixup I used in 1 or 2 earlier competition and works Ok. GridMask I tried but it decreased my score...so likely not the right configuration..I think it will be interresting enough to give it a few more tries.",
      "votes": null
    },
    {
      "id": "757794",
      "postDate": "02/27/2020 05:30:20",
      "content": "<ol>\n<li>change backbone model to <strong><em>seresnext50</em></strong> will improve about 0.5%.</li>\n<li>use <strong><em>ImageNetPolicy</em></strong> of AutoAugment will improve 0.8-1.0%</li>\n<li>use [2] + Cutout: + 0.2%</li>\n<li>use [2] + Mixup: + 0.4%</li>\n<li>change Mixup in [4] to Mixup(50%)/Cutmix(50%): [4]‘s score + ~0.2%</li>\n</ol>\n\n<p>so, [1] + [2] + [5] will imporove about 1.5-2.0%. </p>\n\n<p>plus: My score is based on these + some modifications of model structure, I think these augments can give LB about 0.97+</p>",
      "rawMarkdown": "1. change backbone model to ***seresnext50*** will improve about 0.5%.\n2. use ***ImageNetPolicy*** of AutoAugment will improve 0.8-1.0%\n3. use [2] + Cutout: + 0.2%\n4. use [2] + Mixup: + 0.4%\n5. change Mixup in [4] to Mixup(50%)/Cutmix(50%): [4]‘s score + ~0.2%\n\nso, [1] + [2] + [5] will imporove about 1.5-2.0%. \n\nplus: My score is based on these + some modifications of model structure, I think these augments can give LB about 0.97+",
      "votes": null
    },
    {
      "id": "757806",
      "postDate": "02/27/2020 05:54:29",
      "content": "<p>How about ensemble of several methods? I would like to know if the different methods are correlated, i.e same or different kind of error?</p>",
      "rawMarkdown": "How about ensemble of several methods? I would like to know if the different methods are correlated, i.e same or different kind of error?",
      "votes": null
    },
    {
      "id": "757829",
      "postDate": "02/27/2020 06:36:24",
      "content": "<p>@ Robin Smits</p>\n\n<p>\" I have not added any augmentation yet\"</p>\n\n<p>do you mean \"no augmentation at all\" or you have used \"some basic augmentation but no mixup/cutup\"</p>",
      "rawMarkdown": "Robin Smits\n\n\" I have not added any augmentation yet\"\n\ndo you mean \"no augmentation at all\" or you have used \"some basic augmentation but no mixup/cutup\"",
      "votes": null
    },
    {
      "id": "757862",
      "postDate": "02/27/2020 07:30:39",
      "content": "<p>Thanks for your sharing.\n I update the result of my experiment.It seems that gridmask is worse than use nothing.\nAnd I suggest Cutmix+Mixup.It has the best performance in my experiment</p>",
      "rawMarkdown": "Thanks for your sharing.\n I update the result of my experiment.It seems that gridmask is worse than use nothing.\nAnd I suggest Cutmix+Mixup.It has the best performance in my experiment",
      "votes": null
    },
    {
      "id": "757867",
      "postDate": "02/27/2020 07:36:50",
      "content": "<p>Thank you for your suggest.\nI tried some ensemble augmentations.\nAfter a lot of try,the combinations didn't perform better.\nI think maybe the augmentations (like gridmask,mixup,cutout,cutmix) they are trying to create occlusion.It means that these kinds of augmentations won't add any new featuremaps but increase the robustness of the model by blocking features.</p>",
      "rawMarkdown": "Thank you for your suggest.\nI tried some ensemble augmentations.\nAfter a lot of try,the combinations didn't perform better.\nI think maybe the augmentations (like gridmask,mixup,cutout,cutmix) they are trying to create occlusion.It means that these kinds of augmentations won't add any new featuremaps but increase the robustness of the model by blocking features.",
      "votes": null
    },
    {
      "id": "757868",
      "postDate": "02/27/2020 07:37:55",
      "content": "<p>Thank you a lot.I will try them one by one</p>",
      "rawMarkdown": "Thank you a lot.I will try them one by one",
      "votes": null
    },
    {
      "id": "757870",
      "postDate": "02/27/2020 07:38:50",
      "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> The first part. At this moment I have completely no augmentation implemented. And I keep being able to further improve my model...although slowly.</p>\n\n<p>As long as I can improve without augmentation I will keep trying. When I'am stuck I will start adding augmentation. There are a lot of good discussions (like this one) showing the results of different augmentations...so it should be no problem to add it last minute.</p>",
      "rawMarkdown": "Hi @hengck23 The first part. At this moment I have completely no augmentation implemented. And I keep being able to further improve my model...although slowly.\n\nAs long as I can improve without augmentation I will keep trying. When I'am stuck I will start adding augmentation. There are a lot of good discussions (like this one) showing the results of different augmentations...so it should be no problem to add it last minute.",
      "votes": null
    },
    {
      "id": "757872",
      "postDate": "02/27/2020 07:48:20",
      "content": "<p>if ensemble, doesn't work, there there is no correlation. you can further confirm it with a table</p>\n\n<p>```\n                cutout   mixup   etc\ncutout      %...        %...      %...\nmixup      %...        %...      %...\netc           %...        %...      %...</p>\n\n<p>compute table of disagreement</p>\n\n<p>agree means cutout_trained_predictor(sample x) == mixup_trained_predictor(sample x)\n```</p>\n\n<p>if that is the case, some other augmentation is needed or new parameters of cutout, mixup is needed</p>",
      "rawMarkdown": "if ensemble, doesn't work, there there is no correlation. you can further confirm it with a table\n\n```\n                cutout   mixup   etc\ncutout      %...        %...      %...\nmixup      %...        %...      %...\netc           %...        %...      %...\n\ncompute table of disagreement\n\nagree means cutout_trained_predictor(sample x) == mixup_trained_predictor(sample x)\n```\n\nif that is the case, some other augmentation is needed or new parameters of cutout, mixup is needed",
      "votes": null
    },
    {
      "id": "758124",
      "postDate": "02/27/2020 13:04:08",
      "content": "<p>Just for FYI.\nI had 2 experiments without any augmentation, one got CV 0.985799 by 128x128 input, and another got CV 0.990776 by 224x224 input. But I didn't submit them.\nGood luck!</p>",
      "rawMarkdown": "Just for FYI.\nI had 2 experiments without any augmentation, one got CV 0.985799 by 128x128 input, and another got CV 0.990776 by 224x224 input. But I didn't submit them.\nGood luck!",
      "votes": null
    },
    {
      "id": "758127",
      "postDate": "02/27/2020 13:09:19",
      "content": "<p>@Robin Smits</p>\n\n<p>\"Hi <a href=\"/hengck23\">@hengck23</a> The first part. At this moment ... \"</p>\n\n<p>if you check my post at <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757</a>, augmentation improve my results by 0.02 LB.</p>\n\n<p>so it seems that your results will be in the 0.99+ range if you add augmentation</p>",
      "rawMarkdown": "Robin Smits\n \n\"Hi @hengck23 The first part. At this moment ... \"\n\nif you check my post at https://www.kaggle.com/c/bengaliai-cv19/discussion/123757, augmentation improve my results by 0.02 LB.\n\nso it seems that your results will be in the 0.99+ range if you add augmentation",
      "votes": null
    },
    {
      "id": "758179",
      "postDate": "02/27/2020 14:20:42",
      "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> with augmentation in the 0.99+ range ... Wow that would be something :-)\nI took a quick look at your post. I will go through it in detail tonight.</p>\n\n<p>Let me know what your thoughts are about teaming up...your augmentation...my model...\nSounds like we could help each other a lot :-)\nIn case you do then send me a PM. In case you don't also not a problem... I will continue as is.</p>",
      "rawMarkdown": "Hi @hengck23 with augmentation in the 0.99+ range ... Wow that would be something :-)\nI took a quick look at your post. I will go through it in detail tonight.\n\nLet me know what your thoughts are about teaming up...your augmentation...my model...\nSounds like we could help each other a lot :-)\nIn case you do then send me a PM. In case you don't also not a problem... I will continue as is.",
      "votes": null
    },
    {
      "id": "758187",
      "postDate": "02/27/2020 14:30:14",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> Real magic! Can't wait to hear out your solution</p>",
      "rawMarkdown": "haqishen Real magic! Can't wait to hear out your solution",
      "votes": null
    },
    {
      "id": "758722",
      "postDate": "02/28/2020 04:40:22",
      "content": "<p>@Robin Smits</p>\n\n<p>\"Just by using a solid baseline model I have 0.9724.\"  i check your github code and did my own experiments in efficientnetb3.</p>\n\n<p>can i confirm again your reported 0.9724 here is local CV or LB score?</p>",
      "rawMarkdown": "Robin Smits\n \n\"Just by using a solid baseline model I have 0.9724.\"  i check your github code and did my own experiments in efficientnetb3.\n\ncan i confirm again your reported 0.9724 here is local CV or LB score?",
      "votes": null
    },
    {
      "id": "758865",
      "postDate": "02/28/2020 09:13:29",
      "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> The 0.9724 score is LB  Local CV (Val Root Recall) was around 0.977 .. see <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530</a></p>\n\n<p>The model which gave a score of 0.9724 also gave already some scores of 0.9727 and 0.9728. It however fluctuates a little bit more compared to a normal fixed train/validation set as I vary the training distribution with each epoch.\nThats why in my kernel I do an ensemble of multiple weight files from different epochs.. it removes any fluctuations and combines into a higher score which follows quitte closely the 'val_root_recall'.</p>\n\n<p>Please do note that the 0.9724 and 0.9728 scores come from single model and from a model that's based on my github however it already contains a few changed and added lines.</p>",
      "rawMarkdown": "Hi @hengck23 The 0.9724 score is LB  Local CV (Val Root Recall) was around 0.977 .. see [https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530](https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530)\n\nThe model which gave a score of 0.9724 also gave already some scores of 0.9727 and 0.9728. It however fluctuates a little bit more compared to a normal fixed train/validation set as I vary the training distribution with each epoch.\nThats why in my kernel I do an ensemble of multiple weight files from different epochs.. it removes any fluctuations and combines into a higher score which follows quitte closely the 'val_root_recall'.\n\nPlease do note that the 0.9724 and 0.9728 scores come from single model and from a model that's based on my github however it already contains a few changed and added lines.",
      "votes": null
    },
    {
      "id": "758960",
      "postDate": "02/28/2020 11:53:42",
      "content": "<p><a href=\"/rsmits\">@rsmits</a> </p>\n\n<p>Thanks for answering my questions. I was puzzled on how to get LB0.97+ without augmentation. And i did a series of experiments to confirm this (i will publish my results soon).</p>\n\n<p>(1)  You mentioned at <a href=\"https://www.kaggle.com/rsmits/keras-efficientnet-b3-training-inference\">https://www.kaggle.com/rsmits/keras-efficientnet-b3-training-inference</a>:\n\"With a fixed training set and this model I would get about 0.970 - 972. LB varied around 0.961 - 0.963.\"</p>\n\n<p>i setup pytorch efficientb3 under same setting as yours. i get local CV of 0.973 under fixed training set (random split). I think this is close to yours</p>\n\n<p>(2) I deduced that your single model LB 0.97+ comes from:\n     - using changing dataset for each epoch (in fact you are using all train dataset in the long run)\n     - ensemble of the models at the final epoches. \n    (i further deduce that 15% additional data improves LB 0.96+ to LB 0.97+) </p>",
      "rawMarkdown": "rsmits \n\nThanks for answering my questions. I was puzzled on how to get LB0.97+ without augmentation. And i did a series of experiments to confirm this (i will publish my results soon).\n\n(1)  You mentioned at https://www.kaggle.com/rsmits/keras-efficientnet-b3-training-inference:\n\"With a fixed training set and this model I would get about 0.970 - 972. LB varied around 0.961 - 0.963.\"\n\ni setup pytorch efficientb3 under same setting as yours. i get local CV of 0.973 under fixed training set (random split). I think this is close to yours\n\n(2) I deduced that your single model LB 0.97+ comes from:\n     - using changing dataset for each epoch (in fact you are using all train dataset in the long run)\n     - ensemble of the models at the final epoches. \n    (i further deduce that 15% additional data improves LB 0.96+ to LB 0.97+)",
      "votes": null
    },
    {
      "id": "758974",
      "postDate": "02/28/2020 12:23:04",
      "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> No problem. We all learn from these discussions :-)\nThanks for such a nice reproduction in Pytorch. </p>\n\n<p>And yes indeed the performance comes for a large part from using all the train data. With the hardware I have available a full 5 or 6 fold CV usually takes to much time..thats why I came up with this solution. Used it in multiple competitions already...\nI occasionally compared this method with a 6 fold CV and the way I train and setup the data distribution for each epoch comes pretty close to what a CV ensemble with a model from each fold would give you. I just have the results within about 1,5 to 2 days..instead of close to 10 days.</p>",
      "rawMarkdown": "Hi @hengck23 No problem. We all learn from these discussions :-)\nThanks for such a nice reproduction in Pytorch. \n\nAnd yes indeed the performance comes for a large part from using all the train data. With the hardware I have available a full 5 or 6 fold CV usually takes to much time..thats why I came up with this solution. Used it in multiple competitions already...\nI occasionally compared this method with a 6 fold CV and the way I train and setup the data distribution for each epoch comes pretty close to what a CV ensemble with a model from each fold would give you. I just have the results within about 1,5 to 2 days..instead of close to 10 days.",
      "votes": null
    },
    {
      "id": "759079",
      "postDate": "02/28/2020 14:45:58",
      "content": "<p>one way to test augmentation is:\n1. split into 10% validation, 70% train1, 20% train2\n2. get baseline results-a for train1+train2 (without augmentation or just use basic augmentation)\n3. get baseline results-b for train1 (without augmentation or just use basic augmentation)\n4. now try train1+aug(train) and see if aug() can get same results (or even better) as  train1+train2 ?\n5. you can even compare results of \" train1+aug(train) \" against  \" train1+0.2 of train2\"  ,\" train1+0.4 of train2\" , \" train1+0.6 of train2\"   ....  This shows aug() is effectively creating how many new samples</p>",
      "rawMarkdown": "one way to test augmentation is:\n1. split into 10% validation, 70% train1, 20% train2\n2. get baseline results-a for train1+train2 (without augmentation or just use basic augmentation)\n3. get baseline results-b for train1 (without augmentation or just use basic augmentation)\n4. now try train1+aug(train) and see if aug() can get same results (or even better) as  train1+train2 ?\n5. you can even compare results of \" train1+aug(train) \" against  \" train1+0.2 of train2\"  ,\" train1+0.4 of train2\" , \" train1+0.6 of train2\"   ....  This shows aug() is effectively creating how many new samples",
      "votes": null
    },
    {
      "id": "759202",
      "postDate": "02/28/2020 18:05:45",
      "content": "<p>I'm really curious about this rolling training set per epoch. I don't think I've heard about this before, but it sounds like a really interesting idea when it comes to using all the data without having to train 4, 5 or 6 models to do a multiple Fold weighted model ensembling. It does mean though that you don't really have a CV and are purely basing yourself on the LB, if I understand correctly?</p>\n\n<p>I might give it a try, without augmentation and with as well, to see how things go.</p>\n\n<p>Have you also tried with resized images <a href=\"/rsmits\">@rsmits</a>? Did that improve any of your scores? The notebook you published is only resizing the original images, but did you try any other pre-processing?</p>",
      "rawMarkdown": "I'm really curious about this rolling training set per epoch. I don't think I've heard about this before, but it sounds like a really interesting idea when it comes to using all the data without having to train 4, 5 or 6 models to do a multiple Fold weighted model ensembling. It does mean though that you don't really have a CV and are purely basing yourself on the LB, if I understand correctly?\n\nI might give it a try, without augmentation and with as well, to see how things go.\n\nHave you also tried with resized images @rsmits? Did that improve any of your scores? The notebook you published is only resizing the original images, but did you try any other pre-processing?",
      "votes": null
    },
    {
      "id": "759267",
      "postDate": "02/28/2020 20:13:38",
      "content": "<p><a href=\"/rsmits\">@rsmits</a>, I took a serious look at your code and trained a model for 15 epochs that gave CV (Val Root Recall) of 0.928382  and LB = 0.9521 Thanks for sharing this very clean and easy to read code. </p>\n\n<p>From the way you are splitting the data, how do you continue training from a previously trained model without leakage?</p>",
      "rawMarkdown": "rsmits, I took a serious look at your code and trained a model for 15 epochs that gave CV (Val Root Recall) of 0.928382  and LB = 0.9521 Thanks for sharing this very clean and easy to read code. \n\nFrom the way you are splitting the data, how do you continue training from a previously trained model without leakage?",
      "votes": null
    },
    {
      "id": "759282",
      "postDate": "02/28/2020 20:54:33",
      "content": "<p>Echoing <a href=\"/sheriytm\">@sheriytm</a>'s sentiment—the repo is ridiculously straightforward and easy to go through. Also, I dunno if <a href=\"/rsmits\">@rsmits</a> continues training from a previous model? Looks like he just snapshot ensembles a bunch of the previous saved epochs and uses the LB for validation? That stated, since his <code>msss</code> is seeded, I'd assume we could replicate the stratification for a particular epoch by instantiating the <code>msss</code> object again with the same seed, then looping over the <code>for _ in msss.split(X_train, Y_train): pass</code> call <code>num_trained_epoch</code> times before we resume training.</p>",
      "rawMarkdown": "Echoing @sheriytm's sentiment—the repo is ridiculously straightforward and easy to go through. Also, I dunno if @rsmits continues training from a previous model? Looks like he just snapshot ensembles a bunch of the previous saved epochs and uses the LB for validation? That stated, since his `msss` is seeded, I'd assume we could replicate the stratification for a particular epoch by instantiating the `msss` object again with the same seed, then looping over the `for _ in msss.split(X_train, Y_train): pass` call `num_trained_epoch` times before we resume training.",
      "votes": null
    },
    {
      "id": "759338",
      "postDate": "02/28/2020 22:51:18",
      "content": "<p><a href=\"/sheriytm\">@sheriytm</a> <a href=\"/authman\">@authman</a> \nThanks and your welcome. so basically with the way the data splits are setup (as below)</p>\n\n<p><code># Multi Label Stratified Split stuff...\nmsss = MultilabelStratifiedShuffleSplit(n_splits = EPOCHS, test_size = TEST_SIZE, random_state = SEED)\n</code>\nI create a number of different splits equal to the amount of EPOCHS...so regardless of whether you restart training or not .. There is data leakage. because I use all of the training data for training and validation.\nWith the custom loop around model.fit you could basically see that as restarting training after fitting each model. To use a different train distribution this was necessary however.</p>\n\n<p>Even with the data leakage  the local score for 'val_root_recall' is quitte close to either single models or ensembles of single models and there score on the leaderboard.</p>\n\n<p>So I'am making my assumptions based on local CV and the LB. Up to various degrees I used this approach in previous competitions and I'am personally happy with the results.</p>\n\n<p>for example my current best score of 0.9742 was matched with a local CV score of about 0.977 - 0.978. And that already for multiple models.. I've seen larger gaps in the various discussions ;-)</p>",
      "rawMarkdown": "sheriytm @authman \nThanks and your welcome. so basically with the way the data splits are setup (as below)\n\n`# Multi Label Stratified Split stuff...\nmsss = MultilabelStratifiedShuffleSplit(n_splits = EPOCHS, test_size = TEST_SIZE, random_state = SEED)\n`\nI create a number of different splits equal to the amount of EPOCHS...so regardless of whether you restart training or not .. There is data leakage. because I use all of the training data for training and validation.\nWith the custom loop around model.fit you could basically see that as restarting training after fitting each model. To use a different train distribution this was necessary however.\n\nEven with the data leakage  the local score for 'val_root_recall' is quitte close to either single models or ensembles of single models and there score on the leaderboard.\n\nSo I'am making my assumptions based on local CV and the LB. Up to various degrees I used this approach in previous competitions and I'am personally happy with the results.\n\nfor example my current best score of 0.9742 was matched with a local CV score of about 0.977 - 0.978. And that already for multiple models.. I've seen larger gaps in the various discussions ;-)",
      "votes": null
    },
    {
      "id": "759345",
      "postDate": "02/28/2020 22:56:24",
      "content": "<p>Yup, I was more talking about from a reproduceability standpoint ;-)</p>",
      "rawMarkdown": "Yup, I was more talking about from a reproduceability standpoint ;-)",
      "votes": null
    },
    {
      "id": "760049",
      "postDate": "02/29/2020 19:20:03",
      "content": "<p>Just out of curiosity, what are your GPU ?</p>",
      "rawMarkdown": "Just out of curiosity, what are your GPU ?",
      "votes": null
    },
    {
      "id": "760187",
      "postDate": "03/01/2020 00:17:22",
      "content": "<p><a href=\"/rsmits\">@rsmits</a> Bro I think you should start adding augmentations😃 . Experimenting with combinations of them are extremely time-consuming, especially when you also have to tune the number of epochs to train. If you want a sound ensemble by the deadline, I think you should head that way now:)</p>",
      "rawMarkdown": "rsmits Bro I think you should start adding augmentations😃 . Experimenting with combinations of them are extremely time-consuming, especially when you also have to tune the number of epochs to train. If you want a sound ensemble by the deadline, I think you should head that way now:)",
      "votes": null
    },
    {
      "id": "760438",
      "postDate": "03/01/2020 09:43:14",
      "content": "<p>Cutout working pretty well for me as well but I'm not able to push my CV beyond 0.99.\nI'm using DenseNet121 only though, can't get any good results with seresnext50 :|</p>",
      "rawMarkdown": "Cutout working pretty well for me as well but I'm not able to push my CV beyond 0.99.\nI'm using DenseNet121 only though, can't get any good results with seresnext50 :|",
      "votes": null
    },
    {
      "id": "760670",
      "postDate": "03/01/2020 15:46:21",
      "content": "<p>Hi <a href=\"/roguekk007\">@roguekk007</a> Thanks for the feedback. Yes I've teamed up with fellow kagglers and we are now further experimenting with that. A lot easier to try out new things in that way and indeed also finetune augmentation.</p>",
      "rawMarkdown": "Hi @roguekk007 Thanks for the feedback. Yes I've teamed up with fellow kagglers and we are now further experimenting with that. A lot easier to try out new things in that way and indeed also finetune augmentation.",
      "votes": null
    },
    {
      "id": "760740",
      "postDate": "03/01/2020 17:23:24",
      "content": "<p><a href=\"/rsmits\">@rsmits</a> How does your procedure compare with training on all data? Do you find it to yield better results?</p>",
      "rawMarkdown": "rsmits How does your procedure compare with training on all data? Do you find it to yield better results?",
      "votes": null
    },
    {
      "id": "760819",
      "postDate": "03/01/2020 18:54:45",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> When training on all data I would have no validation at all..with this method there is at least some validation (altough with some data leakage). As it turns out that even with the leakage the 'val_root_recall' metric gives a good enough estimation of what the leaderboard will bring.</p>\n\n<p>I can imagine that training with all data could even bring some better results .. but then finding the right epoch(s) to use would really come down to LB probing.</p>\n\n<p>Or are there ways to do that? Any tips or hints you might have would be appreciated :-)</p>",
      "rawMarkdown": "cpmpml When training on all data I would have no validation at all..with this method there is at least some validation (altough with some data leakage). As it turns out that even with the leakage the 'val_root_recall' metric gives a good enough estimation of what the leaderboard will bring.\n\nI can imagine that training with all data could even bring some better results .. but then finding the right epoch(s) to use would really come down to LB probing.\n\nOr are there ways to do that? Any tips or hints you might have would be appreciated :-)",
      "votes": null
    },
    {
      "id": "761392",
      "postDate": "03/02/2020 13:17:32",
      "content": "<p>I don't have a real answer to your question, which is why I asked about the effectiveness of your method.  I saw it used by a team mate a year ago or so, and always wanted to try it since then, as he was getting better LB results than our regular k fold cross validation).  I cannot point you to a writeup about it given our team got removed (because of the same team mate).  We did not disclose what we did.</p>",
      "rawMarkdown": "I don't have a real answer to your question, which is why I asked about the effectiveness of your method.  I saw it used by a team mate a year ago or so, and always wanted to try it since then, as he was getting better LB results than our regular k fold cross validation).  I cannot point you to a writeup about it given our team got removed (because of the same team mate).  We did not disclose what we did.",
      "votes": null
    },
    {
      "id": "761468",
      "postDate": "03/02/2020 14:53:18",
      "content": "<p>Ok <a href=\"/cpmpml\">@cpmpml</a> When I started using it this way I did a couple of comparisons with smaller datasets (MNIST etc.) So train a full 5 fold CV and train through using all data and varying datastributions.\nThen do a 5 ensemble...take 1 model from each CV fold and take for example the last 5 models from the varying datadistributions model. \nI haven't written a report about it but I remember that performance was mostly almost equal..but it required 5 times less time.</p>\n\n<p>I've used it equally in the Google Landmark Detection competition last year and the RSNA competition late 2019. For me it works on par with Cross Validation...with the current sizes of some models and images I don't even attempt a CV anymore.\nAnother thing which I think works as an advantage is that every epoch has its own different datadistribution and that may'be also makes the model more generalized.</p>\n\n<p>After the competition ends I will see if I can make a general kernel based on MNIST that compares the 2 methods.</p>",
      "rawMarkdown": "Ok @cpmpml When I started using it this way I did a couple of comparisons with smaller datasets (MNIST etc.) So train a full 5 fold CV and train through using all data and varying datastributions.\nThen do a 5 ensemble...take 1 model from each CV fold and take for example the last 5 models from the varying datadistributions model. \nI haven't written a report about it but I remember that performance was mostly almost equal..but it required 5 times less time.\n\nI've used it equally in the Google Landmark Detection competition last year and the RSNA competition late 2019. For me it works on par with Cross Validation...with the current sizes of some models and images I don't even attempt a CV anymore.\nAnother thing which I think works as an advantage is that every epoch has its own different datadistribution and that may'be also makes the model more generalized.\n\nAfter the competition ends I will see if I can make a general kernel based on MNIST that compares the 2 methods.",
      "votes": null
    },
    {
      "id": "761475",
      "postDate": "03/02/2020 15:01:27",
      "content": "<p>兄弟帮个忙。看看这是啥问题，还在天津的话请你吃饭\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2Fd433e0cdea11211b55df8aaac6ca12c7%2FSharedScreenshot.jpg?generation=1583161208465617&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "兄弟帮个忙。看看这是啥问题，还在天津的话请你吃饭\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2Fd433e0cdea11211b55df8aaac6ca12c7%2FSharedScreenshot.jpg?generation=1583161208465617&amp;alt=media)",
      "votes": null
    },
    {
      "id": "763126",
      "postDate": "03/04/2020 06:53:50",
      "content": "<p>I had the same error when i was submitting a empty <code>submission.csv</code>. Check if its empty or not.</p>",
      "rawMarkdown": "I had the same error when i was submitting a empty `submission.csv`. Check if its empty or not.",
      "votes": null
    },
    {
      "id": "763642",
      "postDate": "03/04/2020 17:28:03",
      "content": "<blockquote>\n  <p>can we talk in private？</p>\n</blockquote>\n\n<p>Only if you team, see the competition rules.</p>",
      "rawMarkdown": "&gt; can we talk in private？\n\nOnly if you team, see the competition rules.",
      "votes": null
    },
    {
      "id": "1321842",
      "postDate": "05/25/2021 03:17:27",
      "content": "<p>hi，did u upload of your result?</p>",
      "rawMarkdown": "hi，did u upload of your result?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 755673,
      "author_name": "fengguninghun",
      "author_url": "",
      "post_date": "02/25/2020 02:31:56",
      "content": "<p>I have a similar problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755675,
      "author_name": "markson14",
      "author_url": "",
      "post_date": "02/25/2020 02:34:54",
      "content": "<p>You can give a try on different combination.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755676,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "02/25/2020 02:35:29",
      "content": "<p>For me Cutout works best!</p>",
      "votes": null,
      "replies": [
        {
          "id": 756226,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/25/2020 14:27:20",
          "content": "<p>Do you include mixup and cutmix along with cutout?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760438,
          "author_name": "karolzak",
          "author_url": "",
          "post_date": "03/01/2020 09:43:14",
          "content": "<p>Cutout working pretty well for me as well but I'm not able to push my CV beyond 0.99.\nI'm using DenseNet121 only though, can't get any good results with seresnext50 :|</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755706,
      "author_name": "lishuoshi1996",
      "author_url": "",
      "post_date": "02/25/2020 03:34:32",
      "content": "<p>I  also have a similar problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755709,
      "author_name": "wuzexu",
      "author_url": "",
      "post_date": "02/25/2020 03:41:50",
      "content": "<p>can we talk in private？</p>",
      "votes": null,
      "replies": [
        {
          "id": 756816,
          "author_name": "authman",
          "author_url": "",
          "post_date": "02/26/2020 05:37:07",
          "content": "<p>o_O</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763642,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/04/2020 17:28:03",
          "content": "<blockquote>\n  <p>can we talk in private？</p>\n</blockquote>\n\n<p>Only if you team, see the competition rules.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 756114,
      "author_name": "santosh16k",
      "author_url": "",
      "post_date": "02/25/2020 12:30:24",
      "content": "<p>Try other models like ResNext50, DenseNet, etc.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 756799,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "02/26/2020 05:01:10",
      "content": "<p>Try gridmask + cutout in combination</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 756965,
      "author_name": "erichen188",
      "author_url": "",
      "post_date": "02/26/2020 09:22:58",
      "content": "<p>Hey guys,Thanks for your advice.I do find cutout is really a good augmentation in this dataset.\nI am making a Comparative Test.Later I will upload a chart about my result</p>",
      "votes": null,
      "replies": [
        {
          "id": 1321842,
          "author_name": "martinlee07",
          "author_url": "",
          "post_date": "05/25/2021 03:17:27",
          "content": "<p>hi，did u upload of your result?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 757525,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "02/26/2020 21:39:01",
      "content": "<p>Just by using a solid baseline model I have 0.9724. I have not added any augmentation yet.\nWhen I'am unable to improve it anymore I will likely try either Mixup or GridMask.</p>\n\n<p>Mixup I used in 1 or 2 earlier competition and works Ok. GridMask I tried but it decreased my score...so likely not the right configuration..I think it will be interresting enough to give it a few more tries.</p>",
      "votes": null,
      "replies": [
        {
          "id": 757829,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/27/2020 06:36:24",
          "content": "<p>@ Robin Smits</p>\n\n<p>\" I have not added any augmentation yet\"</p>\n\n<p>do you mean \"no augmentation at all\" or you have used \"some basic augmentation but no mixup/cutup\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757862,
          "author_name": "erichen188",
          "author_url": "",
          "post_date": "02/27/2020 07:30:39",
          "content": "<p>Thanks for your sharing.\n I update the result of my experiment.It seems that gridmask is worse than use nothing.\nAnd I suggest Cutmix+Mixup.It has the best performance in my experiment</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757870,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "02/27/2020 07:38:50",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> The first part. At this moment I have completely no augmentation implemented. And I keep being able to further improve my model...although slowly.</p>\n\n<p>As long as I can improve without augmentation I will keep trying. When I'am stuck I will start adding augmentation. There are a lot of good discussions (like this one) showing the results of different augmentations...so it should be no problem to add it last minute.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758124,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/27/2020 13:04:08",
          "content": "<p>Just for FYI.\nI had 2 experiments without any augmentation, one got CV 0.985799 by 128x128 input, and another got CV 0.990776 by 224x224 input. But I didn't submit them.\nGood luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758127,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/27/2020 13:09:19",
          "content": "<p>@Robin Smits</p>\n\n<p>\"Hi <a href=\"/hengck23\">@hengck23</a> The first part. At this moment ... \"</p>\n\n<p>if you check my post at <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757</a>, augmentation improve my results by 0.02 LB.</p>\n\n<p>so it seems that your results will be in the 0.99+ range if you add augmentation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758179,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "02/27/2020 14:20:42",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> with augmentation in the 0.99+ range ... Wow that would be something :-)\nI took a quick look at your post. I will go through it in detail tonight.</p>\n\n<p>Let me know what your thoughts are about teaming up...your augmentation...my model...\nSounds like we could help each other a lot :-)\nIn case you do then send me a PM. In case you don't also not a problem... I will continue as is.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758187,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/27/2020 14:30:14",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> Real magic! Can't wait to hear out your solution</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758722,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/28/2020 04:40:22",
          "content": "<p>@Robin Smits</p>\n\n<p>\"Just by using a solid baseline model I have 0.9724.\"  i check your github code and did my own experiments in efficientnetb3.</p>\n\n<p>can i confirm again your reported 0.9724 here is local CV or LB score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758865,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "02/28/2020 09:13:29",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> The 0.9724 score is LB  Local CV (Val Root Recall) was around 0.977 .. see <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530</a></p>\n\n<p>The model which gave a score of 0.9724 also gave already some scores of 0.9727 and 0.9728. It however fluctuates a little bit more compared to a normal fixed train/validation set as I vary the training distribution with each epoch.\nThats why in my kernel I do an ensemble of multiple weight files from different epochs.. it removes any fluctuations and combines into a higher score which follows quitte closely the 'val_root_recall'.</p>\n\n<p>Please do note that the 0.9724 and 0.9728 scores come from single model and from a model that's based on my github however it already contains a few changed and added lines.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758960,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/28/2020 11:53:42",
          "content": "<p><a href=\"/rsmits\">@rsmits</a> </p>\n\n<p>Thanks for answering my questions. I was puzzled on how to get LB0.97+ without augmentation. And i did a series of experiments to confirm this (i will publish my results soon).</p>\n\n<p>(1)  You mentioned at <a href=\"https://www.kaggle.com/rsmits/keras-efficientnet-b3-training-inference\">https://www.kaggle.com/rsmits/keras-efficientnet-b3-training-inference</a>:\n\"With a fixed training set and this model I would get about 0.970 - 972. LB varied around 0.961 - 0.963.\"</p>\n\n<p>i setup pytorch efficientb3 under same setting as yours. i get local CV of 0.973 under fixed training set (random split). I think this is close to yours</p>\n\n<p>(2) I deduced that your single model LB 0.97+ comes from:\n     - using changing dataset for each epoch (in fact you are using all train dataset in the long run)\n     - ensemble of the models at the final epoches. \n    (i further deduce that 15% additional data improves LB 0.96+ to LB 0.97+) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758974,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "02/28/2020 12:23:04",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> No problem. We all learn from these discussions :-)\nThanks for such a nice reproduction in Pytorch. </p>\n\n<p>And yes indeed the performance comes for a large part from using all the train data. With the hardware I have available a full 5 or 6 fold CV usually takes to much time..thats why I came up with this solution. Used it in multiple competitions already...\nI occasionally compared this method with a 6 fold CV and the way I train and setup the data distribution for each epoch comes pretty close to what a CV ensemble with a model from each fold would give you. I just have the results within about 1,5 to 2 days..instead of close to 10 days.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759202,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/28/2020 18:05:45",
          "content": "<p>I'm really curious about this rolling training set per epoch. I don't think I've heard about this before, but it sounds like a really interesting idea when it comes to using all the data without having to train 4, 5 or 6 models to do a multiple Fold weighted model ensembling. It does mean though that you don't really have a CV and are purely basing yourself on the LB, if I understand correctly?</p>\n\n<p>I might give it a try, without augmentation and with as well, to see how things go.</p>\n\n<p>Have you also tried with resized images <a href=\"/rsmits\">@rsmits</a>? Did that improve any of your scores? The notebook you published is only resizing the original images, but did you try any other pre-processing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759267,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/28/2020 20:13:38",
          "content": "<p><a href=\"/rsmits\">@rsmits</a>, I took a serious look at your code and trained a model for 15 epochs that gave CV (Val Root Recall) of 0.928382  and LB = 0.9521 Thanks for sharing this very clean and easy to read code. </p>\n\n<p>From the way you are splitting the data, how do you continue training from a previously trained model without leakage?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759282,
          "author_name": "authman",
          "author_url": "",
          "post_date": "02/28/2020 20:54:33",
          "content": "<p>Echoing <a href=\"/sheriytm\">@sheriytm</a>'s sentiment—the repo is ridiculously straightforward and easy to go through. Also, I dunno if <a href=\"/rsmits\">@rsmits</a> continues training from a previous model? Looks like he just snapshot ensembles a bunch of the previous saved epochs and uses the LB for validation? That stated, since his <code>msss</code> is seeded, I'd assume we could replicate the stratification for a particular epoch by instantiating the <code>msss</code> object again with the same seed, then looping over the <code>for _ in msss.split(X_train, Y_train): pass</code> call <code>num_trained_epoch</code> times before we resume training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759338,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "02/28/2020 22:51:18",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> <a href=\"/authman\">@authman</a> \nThanks and your welcome. so basically with the way the data splits are setup (as below)</p>\n\n<p><code># Multi Label Stratified Split stuff...\nmsss = MultilabelStratifiedShuffleSplit(n_splits = EPOCHS, test_size = TEST_SIZE, random_state = SEED)\n</code>\nI create a number of different splits equal to the amount of EPOCHS...so regardless of whether you restart training or not .. There is data leakage. because I use all of the training data for training and validation.\nWith the custom loop around model.fit you could basically see that as restarting training after fitting each model. To use a different train distribution this was necessary however.</p>\n\n<p>Even with the data leakage  the local score for 'val_root_recall' is quitte close to either single models or ensembles of single models and there score on the leaderboard.</p>\n\n<p>So I'am making my assumptions based on local CV and the LB. Up to various degrees I used this approach in previous competitions and I'am personally happy with the results.</p>\n\n<p>for example my current best score of 0.9742 was matched with a local CV score of about 0.977 - 0.978. And that already for multiple models.. I've seen larger gaps in the various discussions ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759345,
          "author_name": "authman",
          "author_url": "",
          "post_date": "02/28/2020 22:56:24",
          "content": "<p>Yup, I was more talking about from a reproduceability standpoint ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760187,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/01/2020 00:17:22",
          "content": "<p><a href=\"/rsmits\">@rsmits</a> Bro I think you should start adding augmentations😃 . Experimenting with combinations of them are extremely time-consuming, especially when you also have to tune the number of epochs to train. If you want a sound ensemble by the deadline, I think you should head that way now:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760670,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "03/01/2020 15:46:21",
          "content": "<p>Hi <a href=\"/roguekk007\">@roguekk007</a> Thanks for the feedback. Yes I've teamed up with fellow kagglers and we are now further experimenting with that. A lot easier to try out new things in that way and indeed also finetune augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760740,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/01/2020 17:23:24",
          "content": "<p><a href=\"/rsmits\">@rsmits</a> How does your procedure compare with training on all data? Do you find it to yield better results?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760819,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "03/01/2020 18:54:45",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> When training on all data I would have no validation at all..with this method there is at least some validation (altough with some data leakage). As it turns out that even with the leakage the 'val_root_recall' metric gives a good enough estimation of what the leaderboard will bring.</p>\n\n<p>I can imagine that training with all data could even bring some better results .. but then finding the right epoch(s) to use would really come down to LB probing.</p>\n\n<p>Or are there ways to do that? Any tips or hints you might have would be appreciated :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761392,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/02/2020 13:17:32",
          "content": "<p>I don't have a real answer to your question, which is why I asked about the effectiveness of your method.  I saw it used by a team mate a year ago or so, and always wanted to try it since then, as he was getting better LB results than our regular k fold cross validation).  I cannot point you to a writeup about it given our team got removed (because of the same team mate).  We did not disclose what we did.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761468,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "03/02/2020 14:53:18",
          "content": "<p>Ok <a href=\"/cpmpml\">@cpmpml</a> When I started using it this way I did a couple of comparisons with smaller datasets (MNIST etc.) So train a full 5 fold CV and train through using all data and varying datastributions.\nThen do a 5 ensemble...take 1 model from each CV fold and take for example the last 5 models from the varying datadistributions model. \nI haven't written a report about it but I remember that performance was mostly almost equal..but it required 5 times less time.</p>\n\n<p>I've used it equally in the Google Landmark Detection competition last year and the RSNA competition late 2019. For me it works on par with Cross Validation...with the current sizes of some models and images I don't even attempt a CV anymore.\nAnother thing which I think works as an advantage is that every epoch has its own different datadistribution and that may'be also makes the model more generalized.</p>\n\n<p>After the competition ends I will see if I can make a general kernel based on MNIST that compares the 2 methods.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 757794,
      "author_name": "welkinfeng",
      "author_url": "",
      "post_date": "02/27/2020 05:30:20",
      "content": "<ol>\n<li>change backbone model to <strong><em>seresnext50</em></strong> will improve about 0.5%.</li>\n<li>use <strong><em>ImageNetPolicy</em></strong> of AutoAugment will improve 0.8-1.0%</li>\n<li>use [2] + Cutout: + 0.2%</li>\n<li>use [2] + Mixup: + 0.4%</li>\n<li>change Mixup in [4] to Mixup(50%)/Cutmix(50%): [4]‘s score + ~0.2%</li>\n</ol>\n\n<p>so, [1] + [2] + [5] will imporove about 1.5-2.0%. </p>\n\n<p>plus: My score is based on these + some modifications of model structure, I think these augments can give LB about 0.97+</p>",
      "votes": null,
      "replies": [
        {
          "id": 757868,
          "author_name": "erichen188",
          "author_url": "",
          "post_date": "02/27/2020 07:37:55",
          "content": "<p>Thank you a lot.I will try them one by one</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 757806,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/27/2020 05:54:29",
      "content": "<p>How about ensemble of several methods? I would like to know if the different methods are correlated, i.e same or different kind of error?</p>",
      "votes": null,
      "replies": [
        {
          "id": 757867,
          "author_name": "erichen188",
          "author_url": "",
          "post_date": "02/27/2020 07:36:50",
          "content": "<p>Thank you for your suggest.\nI tried some ensemble augmentations.\nAfter a lot of try,the combinations didn't perform better.\nI think maybe the augmentations (like gridmask,mixup,cutout,cutmix) they are trying to create occlusion.It means that these kinds of augmentations won't add any new featuremaps but increase the robustness of the model by blocking features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757872,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/27/2020 07:48:20",
          "content": "<p>if ensemble, doesn't work, there there is no correlation. you can further confirm it with a table</p>\n\n<p>```\n                cutout   mixup   etc\ncutout      %...        %...      %...\nmixup      %...        %...      %...\netc           %...        %...      %...</p>\n\n<p>compute table of disagreement</p>\n\n<p>agree means cutout_trained_predictor(sample x) == mixup_trained_predictor(sample x)\n```</p>\n\n<p>if that is the case, some other augmentation is needed or new parameters of cutout, mixup is needed</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759079,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/28/2020 14:45:58",
      "content": "<p>one way to test augmentation is:\n1. split into 10% validation, 70% train1, 20% train2\n2. get baseline results-a for train1+train2 (without augmentation or just use basic augmentation)\n3. get baseline results-b for train1 (without augmentation or just use basic augmentation)\n4. now try train1+aug(train) and see if aug() can get same results (or even better) as  train1+train2 ?\n5. you can even compare results of \" train1+aug(train) \" against  \" train1+0.2 of train2\"  ,\" train1+0.4 of train2\" , \" train1+0.6 of train2\"   ....  This shows aug() is effectively creating how many new samples</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 760049,
      "author_name": "johnnyhalliday",
      "author_url": "",
      "post_date": "02/29/2020 19:20:03",
      "content": "<p>Just out of curiosity, what are your GPU ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 761475,
      "author_name": "littlechicken",
      "author_url": "",
      "post_date": "03/02/2020 15:01:27",
      "content": "<p>兄弟帮个忙。看看这是啥问题，还在天津的话请你吃饭\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2Fd433e0cdea11211b55df8aaac6ca12c7%2FSharedScreenshot.jpg?generation=1583161208465617&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 763126,
          "author_name": "prajwalprashanth",
          "author_url": "",
          "post_date": "03/04/2020 06:53:50",
          "content": "<p>I had the same error when i was submitting a empty <code>submission.csv</code>. Check if its empty or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "755661": "Update 2020/2/27\nHey guys,recently I do some experiment to check the best performence about data augmentations.\nIn my experiment，the setting is:\n\n*dataset:top fifth part\nmodel:resnet34\nopt:Over9000\nlr_schedule:fastai's OneCycle\nbasic augmentation:rotation+light+wrap*\n\nAnd the result looks like these:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2793161%2F543e7a3121eef98f37db96bf5e3bd153%2F20200227101845.png?generation=1582770008940180&amp;alt=media)\n\nSo,I think the different combinations of** Cutmix,Mixup and Cutout** is much better for this dataset.\n\n\nOrigin 2020/2/25\nHey guys，I am wondering which kind of data augmentation is really helpful in this dataset.\nI have already tired the basic Image transform(flip,rotation...), Gridmask, Augmix , Mixup and Cutmix on a small model(resnet34).And the score is:\nbaseline 0.947\nrotation+light+wrap 0.957\nMixup 0.961\nGridmask 0.944\nAugmix 0.946\nCutmix 0.95\nAs you can see only basic transform with out flip and Mixup help me improve the CV score.\nThe combination of basic transform and Mixup 0.969\nAnd I can't reach any more.\nCan any one share your experiment score on your CV?",
    "755673": "I have a similar problem",
    "755675": "You can give a try on different combination.",
    "755676": "For me Cutout works best!",
    "755706": "I  also have a similar problem",
    "755709": "can we talk in private？",
    "756114": "Try other models like ResNext50, DenseNet, etc.",
    "756226": "Do you include mixup and cutmix along with cutout?",
    "756799": "Try gridmask + cutout in combination",
    "756816": "o_O",
    "756965": "Hey guys,Thanks for your advice.I do find cutout is really a good augmentation in this dataset.\nI am making a Comparative Test.Later I will upload a chart about my result",
    "757525": "Just by using a solid baseline model I have 0.9724. I have not added any augmentation yet.\nWhen I'am unable to improve it anymore I will likely try either Mixup or GridMask.\n\nMixup I used in 1 or 2 earlier competition and works Ok. GridMask I tried but it decreased my score...so likely not the right configuration..I think it will be interresting enough to give it a few more tries.",
    "757794": "1. change backbone model to ***seresnext50*** will improve about 0.5%.\n2. use ***ImageNetPolicy*** of AutoAugment will improve 0.8-1.0%\n3. use [2] + Cutout: + 0.2%\n4. use [2] + Mixup: + 0.4%\n5. change Mixup in [4] to Mixup(50%)/Cutmix(50%): [4]‘s score + ~0.2%\n\nso, [1] + [2] + [5] will imporove about 1.5-2.0%. \n\nplus: My score is based on these + some modifications of model structure, I think these augments can give LB about 0.97+",
    "757806": "How about ensemble of several methods? I would like to know if the different methods are correlated, i.e same or different kind of error?",
    "757829": "Robin Smits\n\n\" I have not added any augmentation yet\"\n\ndo you mean \"no augmentation at all\" or you have used \"some basic augmentation but no mixup/cutup\"",
    "757862": "Thanks for your sharing.\n I update the result of my experiment.It seems that gridmask is worse than use nothing.\nAnd I suggest Cutmix+Mixup.It has the best performance in my experiment",
    "757867": "Thank you for your suggest.\nI tried some ensemble augmentations.\nAfter a lot of try,the combinations didn't perform better.\nI think maybe the augmentations (like gridmask,mixup,cutout,cutmix) they are trying to create occlusion.It means that these kinds of augmentations won't add any new featuremaps but increase the robustness of the model by blocking features.",
    "757868": "Thank you a lot.I will try them one by one",
    "757870": "Hi @hengck23 The first part. At this moment I have completely no augmentation implemented. And I keep being able to further improve my model...although slowly.\n\nAs long as I can improve without augmentation I will keep trying. When I'am stuck I will start adding augmentation. There are a lot of good discussions (like this one) showing the results of different augmentations...so it should be no problem to add it last minute.",
    "757872": "if ensemble, doesn't work, there there is no correlation. you can further confirm it with a table\n\n```\n                cutout   mixup   etc\ncutout      %...        %...      %...\nmixup      %...        %...      %...\netc           %...        %...      %...\n\ncompute table of disagreement\n\nagree means cutout_trained_predictor(sample x) == mixup_trained_predictor(sample x)\n```\n\nif that is the case, some other augmentation is needed or new parameters of cutout, mixup is needed",
    "758124": "Just for FYI.\nI had 2 experiments without any augmentation, one got CV 0.985799 by 128x128 input, and another got CV 0.990776 by 224x224 input. But I didn't submit them.\nGood luck!",
    "758127": "Robin Smits\n \n\"Hi @hengck23 The first part. At this moment ... \"\n\nif you check my post at https://www.kaggle.com/c/bengaliai-cv19/discussion/123757, augmentation improve my results by 0.02 LB.\n\nso it seems that your results will be in the 0.99+ range if you add augmentation",
    "758179": "Hi @hengck23 with augmentation in the 0.99+ range ... Wow that would be something :-)\nI took a quick look at your post. I will go through it in detail tonight.\n\nLet me know what your thoughts are about teaming up...your augmentation...my model...\nSounds like we could help each other a lot :-)\nIn case you do then send me a PM. In case you don't also not a problem... I will continue as is.",
    "758187": "haqishen Real magic! Can't wait to hear out your solution",
    "758722": "Robin Smits\n \n\"Just by using a solid baseline model I have 0.9724.\"  i check your github code and did my own experiments in efficientnetb3.\n\ncan i confirm again your reported 0.9724 here is local CV or LB score?",
    "758865": "Hi @hengck23 The 0.9724 score is LB  Local CV (Val Root Recall) was around 0.977 .. see [https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530](https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#757530)\n\nThe model which gave a score of 0.9724 also gave already some scores of 0.9727 and 0.9728. It however fluctuates a little bit more compared to a normal fixed train/validation set as I vary the training distribution with each epoch.\nThats why in my kernel I do an ensemble of multiple weight files from different epochs.. it removes any fluctuations and combines into a higher score which follows quitte closely the 'val_root_recall'.\n\nPlease do note that the 0.9724 and 0.9728 scores come from single model and from a model that's based on my github however it already contains a few changed and added lines.",
    "758960": "rsmits \n\nThanks for answering my questions. I was puzzled on how to get LB0.97+ without augmentation. And i did a series of experiments to confirm this (i will publish my results soon).\n\n(1)  You mentioned at https://www.kaggle.com/rsmits/keras-efficientnet-b3-training-inference:\n\"With a fixed training set and this model I would get about 0.970 - 972. LB varied around 0.961 - 0.963.\"\n\ni setup pytorch efficientb3 under same setting as yours. i get local CV of 0.973 under fixed training set (random split). I think this is close to yours\n\n(2) I deduced that your single model LB 0.97+ comes from:\n     - using changing dataset for each epoch (in fact you are using all train dataset in the long run)\n     - ensemble of the models at the final epoches. \n    (i further deduce that 15% additional data improves LB 0.96+ to LB 0.97+)",
    "758974": "Hi @hengck23 No problem. We all learn from these discussions :-)\nThanks for such a nice reproduction in Pytorch. \n\nAnd yes indeed the performance comes for a large part from using all the train data. With the hardware I have available a full 5 or 6 fold CV usually takes to much time..thats why I came up with this solution. Used it in multiple competitions already...\nI occasionally compared this method with a 6 fold CV and the way I train and setup the data distribution for each epoch comes pretty close to what a CV ensemble with a model from each fold would give you. I just have the results within about 1,5 to 2 days..instead of close to 10 days.",
    "759079": "one way to test augmentation is:\n1. split into 10% validation, 70% train1, 20% train2\n2. get baseline results-a for train1+train2 (without augmentation or just use basic augmentation)\n3. get baseline results-b for train1 (without augmentation or just use basic augmentation)\n4. now try train1+aug(train) and see if aug() can get same results (or even better) as  train1+train2 ?\n5. you can even compare results of \" train1+aug(train) \" against  \" train1+0.2 of train2\"  ,\" train1+0.4 of train2\" , \" train1+0.6 of train2\"   ....  This shows aug() is effectively creating how many new samples",
    "759202": "I'm really curious about this rolling training set per epoch. I don't think I've heard about this before, but it sounds like a really interesting idea when it comes to using all the data without having to train 4, 5 or 6 models to do a multiple Fold weighted model ensembling. It does mean though that you don't really have a CV and are purely basing yourself on the LB, if I understand correctly?\n\nI might give it a try, without augmentation and with as well, to see how things go.\n\nHave you also tried with resized images @rsmits? Did that improve any of your scores? The notebook you published is only resizing the original images, but did you try any other pre-processing?",
    "759267": "rsmits, I took a serious look at your code and trained a model for 15 epochs that gave CV (Val Root Recall) of 0.928382  and LB = 0.9521 Thanks for sharing this very clean and easy to read code. \n\nFrom the way you are splitting the data, how do you continue training from a previously trained model without leakage?",
    "759282": "Echoing @sheriytm's sentiment—the repo is ridiculously straightforward and easy to go through. Also, I dunno if @rsmits continues training from a previous model? Looks like he just snapshot ensembles a bunch of the previous saved epochs and uses the LB for validation? That stated, since his `msss` is seeded, I'd assume we could replicate the stratification for a particular epoch by instantiating the `msss` object again with the same seed, then looping over the `for _ in msss.split(X_train, Y_train): pass` call `num_trained_epoch` times before we resume training.",
    "759338": "sheriytm @authman \nThanks and your welcome. so basically with the way the data splits are setup (as below)\n\n`# Multi Label Stratified Split stuff...\nmsss = MultilabelStratifiedShuffleSplit(n_splits = EPOCHS, test_size = TEST_SIZE, random_state = SEED)\n`\nI create a number of different splits equal to the amount of EPOCHS...so regardless of whether you restart training or not .. There is data leakage. because I use all of the training data for training and validation.\nWith the custom loop around model.fit you could basically see that as restarting training after fitting each model. To use a different train distribution this was necessary however.\n\nEven with the data leakage  the local score for 'val_root_recall' is quitte close to either single models or ensembles of single models and there score on the leaderboard.\n\nSo I'am making my assumptions based on local CV and the LB. Up to various degrees I used this approach in previous competitions and I'am personally happy with the results.\n\nfor example my current best score of 0.9742 was matched with a local CV score of about 0.977 - 0.978. And that already for multiple models.. I've seen larger gaps in the various discussions ;-)",
    "759345": "Yup, I was more talking about from a reproduceability standpoint ;-)",
    "760049": "Just out of curiosity, what are your GPU ?",
    "760187": "rsmits Bro I think you should start adding augmentations😃 . Experimenting with combinations of them are extremely time-consuming, especially when you also have to tune the number of epochs to train. If you want a sound ensemble by the deadline, I think you should head that way now:)",
    "760438": "Cutout working pretty well for me as well but I'm not able to push my CV beyond 0.99.\nI'm using DenseNet121 only though, can't get any good results with seresnext50 :|",
    "760670": "Hi @roguekk007 Thanks for the feedback. Yes I've teamed up with fellow kagglers and we are now further experimenting with that. A lot easier to try out new things in that way and indeed also finetune augmentation.",
    "760740": "rsmits How does your procedure compare with training on all data? Do you find it to yield better results?",
    "760819": "cpmpml When training on all data I would have no validation at all..with this method there is at least some validation (altough with some data leakage). As it turns out that even with the leakage the 'val_root_recall' metric gives a good enough estimation of what the leaderboard will bring.\n\nI can imagine that training with all data could even bring some better results .. but then finding the right epoch(s) to use would really come down to LB probing.\n\nOr are there ways to do that? Any tips or hints you might have would be appreciated :-)",
    "761392": "I don't have a real answer to your question, which is why I asked about the effectiveness of your method.  I saw it used by a team mate a year ago or so, and always wanted to try it since then, as he was getting better LB results than our regular k fold cross validation).  I cannot point you to a writeup about it given our team got removed (because of the same team mate).  We did not disclose what we did.",
    "761468": "Ok @cpmpml When I started using it this way I did a couple of comparisons with smaller datasets (MNIST etc.) So train a full 5 fold CV and train through using all data and varying datastributions.\nThen do a 5 ensemble...take 1 model from each CV fold and take for example the last 5 models from the varying datadistributions model. \nI haven't written a report about it but I remember that performance was mostly almost equal..but it required 5 times less time.\n\nI've used it equally in the Google Landmark Detection competition last year and the RSNA competition late 2019. For me it works on par with Cross Validation...with the current sizes of some models and images I don't even attempt a CV anymore.\nAnother thing which I think works as an advantage is that every epoch has its own different datadistribution and that may'be also makes the model more generalized.\n\nAfter the competition ends I will see if I can make a general kernel based on MNIST that compares the 2 methods.",
    "761475": "兄弟帮个忙。看看这是啥问题，还在天津的话请你吃饭\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1523520%2Fd433e0cdea11211b55df8aaac6ca12c7%2FSharedScreenshot.jpg?generation=1583161208465617&amp;alt=media)",
    "763126": "I had the same error when i was submitting a empty `submission.csv`. Check if its empty or not.",
    "763642": "&gt; can we talk in private？\n\nOnly if you team, see the competition rules.",
    "1321842": "hi，did u upload of your result?"
  },
  "source": "meta"
}