{
  "id": 132898,
  "title": "report on experiments on  \"No Augmentation\"",
  "url": "/competitions/bengaliai-cv19/discussion/132898",
  "author_name": "",
  "post_date": "2020-02-28T13:57:03.776724700Z",
  "votes": 43,
  "comment_count": 51,
  "views": 0,
  "content": "<p>part.0,1,2,3 of experiment as attached. more results later?</p>",
  "messages": [
    {
      "id": "759050",
      "postDate": "02/28/2020 13:57:03",
      "content": "<p>part.0,1,2,3 of experiment as attached. more results later?</p>",
      "rawMarkdown": "part.0,1,2,3 of experiment as attached. more results later?",
      "votes": null
    },
    {
      "id": "759051",
      "postDate": "02/28/2020 14:05:37",
      "content": "<p>coming up:\n1. modification of Robin Smits method:\n  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974\">https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974</a> </p>\n\n<ul>\n<li><p>keep a fixed validation set</p></li>\n<li><p>use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)</p></li>\n</ul>",
      "rawMarkdown": "coming up:\n1. modification of Robin Smits method:\n  https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974 \n\n- keep a fixed validation set\n\n- use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)",
      "votes": null
    },
    {
      "id": "759072",
      "postDate": "02/28/2020 14:29:48",
      "content": "<p>Some day I will be able to understand your pptx reports))</p>",
      "rawMarkdown": "Some day I will be able to understand your pptx reports))",
      "votes": null
    },
    {
      "id": "759120",
      "postDate": "02/28/2020 16:02:34",
      "content": "<p>Thanks for your inspiring slides! I'm considering using multi-scale now. Would you mind telling me the LB score of your multi-scale experiment (experiment 3)? </p>",
      "rawMarkdown": "Thanks for your inspiring slides! I'm considering using multi-scale now. Would you mind telling me the LB score of your multi-scale experiment (experiment 3)?",
      "votes": null
    },
    {
      "id": "759138",
      "postDate": "02/28/2020 16:24:25",
      "content": "<p><a href=\"/ivanwang\">@ivanwang</a></p>\n\n<p>I did not submit. but i can guess it would be around cv=0.977, lb=0.967</p>\n\n<p>the experiment shows that scale is an important data variation.</p>\n\n<p>there are many ways to use this information, e.g add scale augmentation in train (and test). </p>\n\n<p>max pooling over scale is another solution but may be inefficient.\nyou can google for more efficient multi-scale network or  multi-scale input</p>\n\n<p>other possibility includes adding consisentcy loss:</p>\n\n<p>net(image) = feature1\nnet(resized  image) = feature2</p>\n\n<p>feature 1 should be same as feature2</p>",
      "rawMarkdown": "ivanwang\n\nI did not submit. but i can guess it would be around cv=0.977, lb=0.967\n\nthe experiment shows that scale is an important data variation.\n\nthere are many ways to use this information, e.g add scale augmentation in train (and test). \n\nmax pooling over scale is another solution but may be inefficient.\nyou can google for more efficient multi-scale network or  multi-scale input\n\nother possibility includes adding consisentcy loss:\n\nnet(image) = feature1\nnet(resized  image) = feature2\n\nfeature 1 should be same as feature2",
      "votes": null
    },
    {
      "id": "759164",
      "postDate": "02/28/2020 16:49:39",
      "content": "<blockquote>\n  <p><strong>Heng CherKeng wrote:</strong></p>\n  \n  <p>coming up:\n  1. modification of Robin Smits method:\n    <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974\">https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974</a> </p>\n  \n  <ul>\n  <li><p>keep a fixed validation set</p></li>\n  <li><p>use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)</p></li>\n  </ul>\n</blockquote>\n\n<p>... modification of the Robin Smits method....that sounds cool ;-)</p>\n\n<p>Very nice overview <a href=\"/hengck23\">@hengck23</a> what hardware do you have being able to produce such an amount of results so quickly? Likely not an 1070 Ti...</p>",
      "rawMarkdown": "&gt; **Heng CherKeng wrote:**\n&gt; \n&gt; coming up:\n&gt; 1. modification of Robin Smits method:\n&gt;   https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974 \n&gt; \n&gt; - keep a fixed validation set\n&gt; \n&gt; - use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)\n&gt; \n\n... modification of the Robin Smits method....that sounds cool ;-)\n\nVery nice overview @hengck23 what hardware do you have being able to produce such an amount of results so quickly? Likely not an 1070 Ti...",
      "votes": null
    },
    {
      "id": "759198",
      "postDate": "02/28/2020 17:56:24",
      "content": "<p>I keep being impressed by these results without any augmentation!</p>\n\n<p>And same question as <a href=\"/rsmits\">@rsmits</a>, what are you using to get so many results? Both on the hardware level as well as some tricks you might have?</p>",
      "rawMarkdown": "I keep being impressed by these results without any augmentation!\n\nAnd same question as @rsmits, what are you using to get so many results? Both on the hardware level as well as some tricks you might have?",
      "votes": null
    },
    {
      "id": "759208",
      "postDate": "02/28/2020 18:14:41",
      "content": "<p><a href=\"/rsmits\">@rsmits</a> </p>\n\n<p>\" Likely not an 1070 Ti…\"</p>\n\n<p>I have 4x1080Ti and 1xPascal(TitianX).</p>\n\n<p>To make a fair comparison to see the effectiveness of the changing dataset per epoch, we should \n(1) train using all train\n(2) train using a fixed 80% of (1)\n(3) train using changing 80% of (1) </p>\n\n<p>further, we should test with fixed weight + moving average of weight for (1),(2),(3)</p>",
      "rawMarkdown": "rsmits \n\n\" Likely not an 1070 Ti…\"\n\nI have 4x1080Ti and 1xPascal(TitianX).\n\nTo make a fair comparison to see the effectiveness of the changing dataset per epoch, we should \n(1) train using all train\n(2) train using a fixed 80% of (1)\n(3) train using changing 80% of (1) \n\nfurther, we should test with fixed weight + moving average of weight for (1),(2),(3)",
      "votes": null
    },
    {
      "id": "759211",
      "postDate": "02/28/2020 18:21:39",
      "content": "<p>If we want to train 3 times with 100% / 80% / 80% data, isn't it worth doing 3 or even 4-Fold validation for similar training times?</p>\n\n<p>And 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P</p>",
      "rawMarkdown": "If we want to train 3 times with 100% / 80% / 80% data, isn't it worth doing 3 or even 4-Fold validation for similar training times?\n\nAnd 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P",
      "votes": null
    },
    {
      "id": "759225",
      "postDate": "02/28/2020 18:50:27",
      "content": "<p>Cool! 5 GPU's....that explains the speed perfectly ;-)\nLooking forward to your full investigation <a href=\"/hengck23\">@hengck23</a>  This is extremely usefull for all Kagglers. \nThank You! </p>",
      "rawMarkdown": "Cool! 5 GPU's....that explains the speed perfectly ;-)\nLooking forward to your full investigation @hengck23  This is extremely usefull for all Kagglers. \nThank You!",
      "votes": null
    },
    {
      "id": "759313",
      "postDate": "02/28/2020 21:55:13",
      "content": "<blockquote>\n  <p>And 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P</p>\n</blockquote>\n\n<p>Trust me, even after you stop being a student, it won't be that easy to rack up 4x2080TIs</p>",
      "rawMarkdown": "&gt; And 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P\n\nTrust me, even after you stop being a student, it won't be that easy to rack up 4x2080TIs",
      "votes": null
    },
    {
      "id": "759414",
      "postDate": "02/29/2020 02:22:37",
      "content": "<p>P106-100 6G headless mining card on ebay, $75-$100.  I got mine for $50.</p>",
      "rawMarkdown": "P106-100 6G headless mining card on ebay, $75-$100.  I got mine for $50.",
      "votes": null
    },
    {
      "id": "759465",
      "postDate": "02/29/2020 04:15:50",
      "content": "<p>Great work</p>",
      "rawMarkdown": "Great work",
      "votes": null
    },
    {
      "id": "759487",
      "postDate": "02/29/2020 05:12:46",
      "content": "<p>In your experiment.3 you use dual scale streams with 64x112 &amp; 96x168 to get a better score, did you compare it with using 96x168 only? 🤔 </p>",
      "rawMarkdown": "In your experiment.3 you use dual scale streams with 64x112 &amp; 96x168 to get a better score, did you compare it with using 96x168 only? 🤔",
      "votes": null
    },
    {
      "id": "759493",
      "postDate": "02/29/2020 05:32:03",
      "content": "<p>yes, the training is in progress. i also have scale1+scale2 instead of max(scale,1,scale2)</p>",
      "rawMarkdown": "yes, the training is in progress. i also have scale1+scale2 instead of max(scale,1,scale2)",
      "votes": null
    },
    {
      "id": "759495",
      "postDate": "02/29/2020 05:33:29",
      "content": "<p>new! experiment.5\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F267ed76905fb347c99cc4094980e55c8%2FSelection_087.png?generation=1582954406367873&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "new! experiment.5\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F267ed76905fb347c99cc4094980e55c8%2FSelection_087.png?generation=1582954406367873&amp;alt=media)",
      "votes": null
    },
    {
      "id": "759573",
      "postDate": "02/29/2020 07:56:05",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> </p>\n\n<p>please see new results at report_bengali_1.pptx</p>\n\n<p>96x168 obtain the best results. Hence it is not required to use dual scale</p>",
      "rawMarkdown": "haqishen \n\nplease see new results at report_bengali_1.pptx\n\n96x168 obtain the best results. Hence it is not required to use dual scale",
      "votes": null
    },
    {
      "id": "759574",
      "postDate": "02/29/2020 07:56:28",
      "content": "<p><a href=\"/ivanwang2016\">@ivanwang2016</a> </p>\n\n<p>please see new results at report_bengali_1.pptx</p>\n\n<p>96x168 obtain the best results. Hence it is not required to use dual scale</p>",
      "rawMarkdown": "ivanwang2016 \n\nplease see new results at report_bengali_1.pptx\n\n96x168 obtain the best results. Hence it is not required to use dual scale",
      "votes": null
    },
    {
      "id": "759583",
      "postDate": "02/29/2020 08:11:04",
      "content": "<p>Oh I very well know the prices of these puppies, and I'm not even planning on getting one 2080Ti, probably going to start a little bit more mid-range than that!</p>",
      "rawMarkdown": "Oh I very well know the prices of these puppies, and I'm not even planning on getting one 2080Ti, probably going to start a little bit more mid-range than that!",
      "votes": null
    },
    {
      "id": "759587",
      "postDate": "02/29/2020 08:21:56",
      "content": "<p>Thanks for the result!</p>",
      "rawMarkdown": "Thanks for the result!",
      "votes": null
    },
    {
      "id": "759608",
      "postDate": "02/29/2020 09:00:00",
      "content": "<p>maybe you can start with vast.ai 😃 </p>",
      "rawMarkdown": "maybe you can start with vast.ai 😃",
      "votes": null
    },
    {
      "id": "760072",
      "postDate": "02/29/2020 20:08:40",
      "content": "<p>Thanks! That's quite an impressive work!</p>",
      "rawMarkdown": "Thanks! That's quite an impressive work!",
      "votes": null
    },
    {
      "id": "760272",
      "postDate": "03/01/2020 03:54:33",
      "content": "<p>3rd set of experiments are as in \"report_bengali_2.pptx (468.56 KB)\"</p>\n\n<p>without augmentation:\n(1) 128x128 on modified senext50+drop block : cv 0.985\n(2) 96x168 on modified senext50+drop block : cv 0.979 \n(3) 96x168 on modified senext50+drop block +2x label : cv 0.979 </p>\n\n<p>new experiments that i am considering : \n-  imiplicit mixup : manifold mixup (mixing feature maps) or shakedrop?\n- 224x224</p>",
      "rawMarkdown": "3rd set of experiments are as in \"report\\_bengali\\_2.pptx (468.56 KB)\"\n\nwithout augmentation:\n(1) 128x128 on modified senext50+drop block : cv 0.985\n(2) 96x168 on modified senext50+drop block : cv 0.979 \n(3) 96x168 on modified senext50+drop block +2x label : cv 0.979 \n\nnew experiments that i am considering : \n-  imiplicit mixup : manifold mixup (mixing feature maps) or shakedrop?\n- 224x224",
      "votes": null
    },
    {
      "id": "760279",
      "postDate": "03/01/2020 04:10:23",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks a lot for sharing your wonderful experiments. I am just wondering whether your current LB score is based on a single fold or ensemble of multiple folds.</p>",
      "rawMarkdown": "hengck23 Thanks a lot for sharing your wonderful experiments. I am just wondering whether your current LB score is based on a single fold or ensemble of multiple folds.",
      "votes": null
    },
    {
      "id": "760622",
      "postDate": "03/01/2020 14:39:58",
      "content": "<p>preview of report_bengali_3.pptx (experiments in progress)</p>",
      "rawMarkdown": "preview of report\\_bengali\\_3.pptx (experiments in progress)",
      "votes": null
    },
    {
      "id": "760834",
      "postDate": "03/01/2020 19:23:08",
      "content": "<p>I included your dropblock but with cutmix and ended with 0.9879 at about 110.0 epochs. it seems to align closely with yours. Looking at your current experiment, it seems manifold progress is about where i was at that point of training. also saw that a square image was better. Thanks for sharing your experiments. I’ll try manifold now. I’m finding these experiments to be more interesting than the competition. </p>",
      "rawMarkdown": "I included your dropblock but with cutmix and ended with 0.9879 at about 110.0 epochs. it seems to align closely with yours. Looking at your current experiment, it seems manifold progress is about where i was at that point of training. also saw that a square image was better. Thanks for sharing your experiments. I’ll try manifold now. I’m finding these experiments to be more interesting than the competition.",
      "votes": null
    },
    {
      "id": "761062",
      "postDate": "03/02/2020 04:54:21",
      "content": "<p><a href=\"/learnmower\">@learnmower</a> </p>\n\n<p>if you are using dropblock, try different block size. i find that bigger block size may further improve in some cases. different size may be required for different layers as well?</p>",
      "rawMarkdown": "learnmower \n\nif you are using dropblock, try different block size. i find that bigger block size may further improve in some cases. different size may be required for different layers as well?",
      "votes": null
    },
    {
      "id": "761150",
      "postDate": "03/02/2020 07:36:18",
      "content": "<p>good suggestion... it makes sense to set different block sizes at different layers - perhaps some relative sizing to the convolution. also, gamma appears to determine location/occurrence of drop block, and I guess it would be a useful parameter to tune... </p>\n\n<p>i noticed that while drop block was developed at google brain in 2018, it seems that image augmentation based approaches to regularize neural nets have had more momentum in the last year... i wonder why is that the case?</p>",
      "rawMarkdown": "good suggestion... it makes sense to set different block sizes at different layers - perhaps some relative sizing to the convolution. also, gamma appears to determine location/occurrence of drop block, and I guess it would be a useful parameter to tune... \n\ni noticed that while drop block was developed at google brain in 2018, it seems that image augmentation based approaches to regularize neural nets have had more momentum in the last year... i wonder why is that the case?",
      "votes": null
    },
    {
      "id": "761277",
      "postDate": "03/02/2020 10:55:13",
      "content": "<p>Thanks a lot to your exps, I always wonder whether 100+ epochs will impove the grapheme_root's score,  as I tried resnet34 for 50+  epochs and the score keeps around 0.956. In your slide I see after 100+ epochs it can reach around 0.98, so will score keep increase by large epochs?, and what loss did you use?</p>",
      "rawMarkdown": "Thanks a lot to your exps, I always wonder whether 100+ epochs will impove the grapheme_root's score,  as I tried resnet34 for 50+  epochs and the score keeps around 0.956. In your slide I see after 100+ epochs it can reach around 0.98, so will score keep increase by large epochs?, and what loss did you use?",
      "votes": null
    },
    {
      "id": "761350",
      "postDate": "03/02/2020 12:16:25",
      "content": "<p>standard cross entropy loss</p>",
      "rawMarkdown": "standard cross entropy loss",
      "votes": null
    },
    {
      "id": "761351",
      "postDate": "03/02/2020 12:17:04",
      "content": "<p>finalised report_bengali_3.pptx is in the top message above</p>",
      "rawMarkdown": "finalised report_bengali_3.pptx is in the top message above",
      "votes": null
    },
    {
      "id": "761525",
      "postDate": "03/02/2020 16:13:02",
      "content": "<p>\"will score keep increase by large epochs?,\"</p>\n\n<p>not necessary. you may ignore the epoches in my logfile. they may not represent the \"actual number of epoches trained\" . sometimes i just learn my machine run over-night. sometime there are bugs in the code, or i may reset some hyper parameters, etc ... but i just let keep the epoch number running</p>",
      "rawMarkdown": "\"will score keep increase by large epochs?,\"\n\nnot necessary. you may ignore the epoches in my logfile. they may not represent the \"actual number of epoches trained\" . sometimes i just learn my machine run over-night. sometime there are bugs in the code, or i may reset some hyper parameters, etc ... but i just let keep the epoch number running",
      "votes": null
    },
    {
      "id": "761615",
      "postDate": "03/02/2020 18:45:36",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": null
    },
    {
      "id": "761664",
      "postDate": "03/02/2020 20:37:57",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> , Thanks for the nice report, really informative. I'm curious about your experiments with dropblock applied to different resolutions: did you change the size of dropblock depending on the image size you use or considered a specific constant value?</p>",
      "rawMarkdown": "hengck23 , Thanks for the nice report, really informative. I'm curious about your experiments with dropblock applied to different resolutions: did you change the size of dropblock depending on the image size you use or considered a specific constant value?",
      "votes": null
    },
    {
      "id": "761676",
      "postDate": "03/02/2020 20:50:36",
      "content": "<p>In his most recent, he uses <code>block_size=10</code> after the first layer, and <code>block_size=5</code> after the second.</p>",
      "rawMarkdown": "In his most recent, he uses `block_size=10` after the first layer, and `block_size=5` after the second.",
      "votes": null
    },
    {
      "id": "761706",
      "postDate": "03/02/2020 21:18:52",
      "content": "<p>Thanks for clarifications. I just thought that the relative box size with respect to the image size may play a role (rather than the absolute value of blocks), and the improvement for larger images may be not only attributed to the larger image size but also to slightly different regularization. </p>",
      "rawMarkdown": "Thanks for clarifications. I just thought that the relative box size with respect to the image size may play a role (rather than the absolute value of blocks), and the improvement for larger images may be not only attributed to the larger image size but also to slightly different regularization.",
      "votes": null
    },
    {
      "id": "761787",
      "postDate": "03/02/2020 23:41:28",
      "content": "<p>use larger size for larger input. i tried up to 33% of the input size. 50% or more may work better, but i haven tried yet.</p>\n\n<p>it should be same as cutout?</p>",
      "rawMarkdown": "use larger size for larger input. i tried up to 33% of the input size. 50% or more may work better, but i haven tried yet.\n\nit should be same as cutout?",
      "votes": null
    },
    {
      "id": "762155",
      "postDate": "03/03/2020 09:00:36",
      "content": "<p>May I ask why do you replace residual sum with max? Is there any reference paper?</p>",
      "rawMarkdown": "May I ask why do you replace residual sum with max? Is there any reference paper?",
      "votes": null
    },
    {
      "id": "762193",
      "postDate": "03/03/2020 10:01:30",
      "content": "<p>it is by experiment. it is slightly better than then addition in some cases.</p>\n\n<p>i suggest you should stick to the normal  addition first, then conduct experiment to see if the change is suitable for your case.</p>",
      "rawMarkdown": "it is by experiment. it is slightly better than then addition in some cases.\n\ni suggest you should stick to the normal  addition first, then conduct experiment to see if the change is suitable for your case.",
      "votes": null
    },
    {
      "id": "762303",
      "postDate": "03/03/2020 12:16:53",
      "content": "<p>I see. Thanks for the info. I would have a try.</p>",
      "rawMarkdown": "I see. Thanks for the info. I would have a try.",
      "votes": null
    },
    {
      "id": "763633",
      "postDate": "03/04/2020 17:18:58",
      "content": "<p>update:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe4777a0e0bae763b2d7f01b8dcb29b31%2FSelection_138.png?generation=1583342336635585&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "update:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe4777a0e0bae763b2d7f01b8dcb29b31%2FSelection_138.png?generation=1583342336635585&amp;alt=media)",
      "votes": null
    },
    {
      "id": "763972",
      "postDate": "03/05/2020 03:30:05",
      "content": "<p>Thank you for your sharing.\nDid you use surgeryed version of resnext? </p>",
      "rawMarkdown": "Thank you for your sharing.\nDid you use surgeryed version of resnext?",
      "votes": null
    },
    {
      "id": "763978",
      "postDate": "03/05/2020 03:38:28",
      "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a>, really thanks for your kindly sharing. According to your report, i try seresnext50 with dropoutblock which the same as your ppt says, and i achieve CV 0.986 without any augments, but when i merge the cutmix + rotate with dropout, the result cv 0.9833 is worse than seresnext50 + cutmix + rotate(for me best cv 0.9913), and it seems these combination not benefit the fitting, and still need tuning, and it seems not just a question which can be solve by simply adding or reducing</p>",
      "rawMarkdown": "Hi @hengck23, really thanks for your kindly sharing. According to your report, i try seresnext50 with dropoutblock which the same as your ppt says, and i achieve CV 0.986 without any augments, but when i merge the cutmix + rotate with dropout, the result cv 0.9833 is worse than seresnext50 + cutmix + rotate(for me best cv 0.9913), and it seems these combination not benefit the fitting, and still need tuning, and it seems not just a question which can be solve by simply adding or reducing",
      "votes": null
    },
    {
      "id": "763995",
      "postDate": "03/05/2020 04:04:36",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> I'm curious about why train with 2x class. </p>",
      "rawMarkdown": "hengck23 I'm curious about why train with 2x class.",
      "votes": null
    },
    {
      "id": "764001",
      "postDate": "03/05/2020 04:09:32",
      "content": "<p>Same for me, dropblock helped me a lot but when adding too hard augmentations it cant converge as much</p>",
      "rawMarkdown": "Same for me, dropblock helped me a lot but when adding too hard augmentations it cant converge as much",
      "votes": null
    },
    {
      "id": "764013",
      "postDate": "03/05/2020 04:33:42",
      "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> </p>\n\n<p>i suppose cutmix  is not suitable to use with dropblock.\nmaybe you can try cutmix at feature map? i will call it mixblock</p>",
      "rawMarkdown": "cswwp347724 \n\ni suppose cutmix  is not suitable to use with dropblock.\nmaybe you can try cutmix at feature map? i will call it mixblock",
      "votes": null
    },
    {
      "id": "764077",
      "postDate": "03/05/2020 05:49:19",
      "content": "<p>Dropblock doesn't work much for me. I'm not sure whether I need more epochs to train it.</p>",
      "rawMarkdown": "Dropblock doesn't work much for me. I'm not sure whether I need more epochs to train it.",
      "votes": null
    },
    {
      "id": "764095",
      "postDate": "03/05/2020 06:15:59",
      "content": "<p>I guess dropblock and mixup are both strong regularization methods and their combinations caused underfitting. Maybe a deeper model would benefit more from that.</p>",
      "rawMarkdown": "I guess dropblock and mixup are both strong regularization methods and their combinations caused underfitting. Maybe a deeper model would benefit more from that.",
      "votes": null
    },
    {
      "id": "764599",
      "postDate": "03/05/2020 16:28:48",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> I might be a bit late on that subject but i can see that you are classifying with 1295 classes but how did you figure out the 1295 classes if there are only 1292 combinations in the train set?</p>",
      "rawMarkdown": "hengck23 I might be a bit late on that subject but i can see that you are classifying with 1295 classes but how did you figure out the 1295 classes if there are only 1292 combinations in the train set?",
      "votes": null
    },
    {
      "id": "764628",
      "postDate": "03/05/2020 17:00:26",
      "content": "<p>yet another update\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F972371de7eefd559d01c5ca331cd4f26%2FSelection_156.png?generation=1583427623253253&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "yet another update\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F972371de7eefd559d01c5ca331cd4f26%2FSelection_156.png?generation=1583427623253253&amp;alt=media)",
      "votes": null
    },
    {
      "id": "764813",
      "postDate": "03/05/2020 23:41:39",
      "content": "<p>an alternative is that if you cannot fuse the different augmentation into a single network training, ensemble two classifiers trained separately on each?</p>",
      "rawMarkdown": "an alternative is that if you cannot fuse the different augmentation into a single network training, ensemble two classifiers trained separately on each?",
      "votes": null
    },
    {
      "id": "764815",
      "postDate": "03/05/2020 23:46:54",
      "content": "<p>How about the gap? I got a CV 0.982 but LB 0.969💔 </p>",
      "rawMarkdown": "How about the gap? I got a CV 0.982 but LB 0.969💔",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 759051,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/28/2020 14:05:37",
      "content": "<p>coming up:\n1. modification of Robin Smits method:\n  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974\">https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974</a> </p>\n\n<ul>\n<li><p>keep a fixed validation set</p></li>\n<li><p>use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 759164,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "02/28/2020 16:49:39",
          "content": "<blockquote>\n  <p><strong>Heng CherKeng wrote:</strong></p>\n  \n  <p>coming up:\n  1. modification of Robin Smits method:\n    <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974\">https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974</a> </p>\n  \n  <ul>\n  <li><p>keep a fixed validation set</p></li>\n  <li><p>use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)</p></li>\n  </ul>\n</blockquote>\n\n<p>... modification of the Robin Smits method....that sounds cool ;-)</p>\n\n<p>Very nice overview <a href=\"/hengck23\">@hengck23</a> what hardware do you have being able to produce such an amount of results so quickly? Likely not an 1070 Ti...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759198,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/28/2020 17:56:24",
          "content": "<p>I keep being impressed by these results without any augmentation!</p>\n\n<p>And same question as <a href=\"/rsmits\">@rsmits</a>, what are you using to get so many results? Both on the hardware level as well as some tricks you might have?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759208,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/28/2020 18:14:41",
          "content": "<p><a href=\"/rsmits\">@rsmits</a> </p>\n\n<p>\" Likely not an 1070 Ti…\"</p>\n\n<p>I have 4x1080Ti and 1xPascal(TitianX).</p>\n\n<p>To make a fair comparison to see the effectiveness of the changing dataset per epoch, we should \n(1) train using all train\n(2) train using a fixed 80% of (1)\n(3) train using changing 80% of (1) </p>\n\n<p>further, we should test with fixed weight + moving average of weight for (1),(2),(3)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759211,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/28/2020 18:21:39",
          "content": "<p>If we want to train 3 times with 100% / 80% / 80% data, isn't it worth doing 3 or even 4-Fold validation for similar training times?</p>\n\n<p>And 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759225,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "02/28/2020 18:50:27",
          "content": "<p>Cool! 5 GPU's....that explains the speed perfectly ;-)\nLooking forward to your full investigation <a href=\"/hengck23\">@hengck23</a>  This is extremely usefull for all Kagglers. \nThank You! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759313,
          "author_name": "authman",
          "author_url": "",
          "post_date": "02/28/2020 21:55:13",
          "content": "<blockquote>\n  <p>And 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P</p>\n</blockquote>\n\n<p>Trust me, even after you stop being a student, it won't be that easy to rack up 4x2080TIs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759414,
          "author_name": "ljschuster",
          "author_url": "",
          "post_date": "02/29/2020 02:22:37",
          "content": "<p>P106-100 6G headless mining card on ebay, $75-$100.  I got mine for $50.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759583,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/29/2020 08:11:04",
          "content": "<p>Oh I very well know the prices of these puppies, and I'm not even planning on getting one 2080Ti, probably going to start a little bit more mid-range than that!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759608,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/29/2020 09:00:00",
          "content": "<p>maybe you can start with vast.ai 😃 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759072,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "02/28/2020 14:29:48",
      "content": "<p>Some day I will be able to understand your pptx reports))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 759120,
      "author_name": "ivanwang2016",
      "author_url": "",
      "post_date": "02/28/2020 16:02:34",
      "content": "<p>Thanks for your inspiring slides! I'm considering using multi-scale now. Would you mind telling me the LB score of your multi-scale experiment (experiment 3)? </p>",
      "votes": null,
      "replies": [
        {
          "id": 759138,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/28/2020 16:24:25",
          "content": "<p><a href=\"/ivanwang\">@ivanwang</a></p>\n\n<p>I did not submit. but i can guess it would be around cv=0.977, lb=0.967</p>\n\n<p>the experiment shows that scale is an important data variation.</p>\n\n<p>there are many ways to use this information, e.g add scale augmentation in train (and test). </p>\n\n<p>max pooling over scale is another solution but may be inefficient.\nyou can google for more efficient multi-scale network or  multi-scale input</p>\n\n<p>other possibility includes adding consisentcy loss:</p>\n\n<p>net(image) = feature1\nnet(resized  image) = feature2</p>\n\n<p>feature 1 should be same as feature2</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759574,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/29/2020 07:56:28",
          "content": "<p><a href=\"/ivanwang2016\">@ivanwang2016</a> </p>\n\n<p>please see new results at report_bengali_1.pptx</p>\n\n<p>96x168 obtain the best results. Hence it is not required to use dual scale</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760072,
          "author_name": "ivanwang2016",
          "author_url": "",
          "post_date": "02/29/2020 20:08:40",
          "content": "<p>Thanks! That's quite an impressive work!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759465,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "02/29/2020 04:15:50",
      "content": "<p>Great work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 759487,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "02/29/2020 05:12:46",
      "content": "<p>In your experiment.3 you use dual scale streams with 64x112 &amp; 96x168 to get a better score, did you compare it with using 96x168 only? 🤔 </p>",
      "votes": null,
      "replies": [
        {
          "id": 759493,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/29/2020 05:32:03",
          "content": "<p>yes, the training is in progress. i also have scale1+scale2 instead of max(scale,1,scale2)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759573,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/29/2020 07:56:05",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> </p>\n\n<p>please see new results at report_bengali_1.pptx</p>\n\n<p>96x168 obtain the best results. Hence it is not required to use dual scale</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759587,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/29/2020 08:21:56",
          "content": "<p>Thanks for the result!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759495,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/29/2020 05:33:29",
      "content": "<p>new! experiment.5\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F267ed76905fb347c99cc4094980e55c8%2FSelection_087.png?generation=1582954406367873&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 763995,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "03/05/2020 04:04:36",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I'm curious about why train with 2x class. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760272,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/01/2020 03:54:33",
      "content": "<p>3rd set of experiments are as in \"report_bengali_2.pptx (468.56 KB)\"</p>\n\n<p>without augmentation:\n(1) 128x128 on modified senext50+drop block : cv 0.985\n(2) 96x168 on modified senext50+drop block : cv 0.979 \n(3) 96x168 on modified senext50+drop block +2x label : cv 0.979 </p>\n\n<p>new experiments that i am considering : \n-  imiplicit mixup : manifold mixup (mixing feature maps) or shakedrop?\n- 224x224</p>",
      "votes": null,
      "replies": [
        {
          "id": 760279,
          "author_name": "muhammedazamkhan",
          "author_url": "",
          "post_date": "03/01/2020 04:10:23",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks a lot for sharing your wonderful experiments. I am just wondering whether your current LB score is based on a single fold or ensemble of multiple folds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760622,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/01/2020 14:39:58",
      "content": "<p>preview of report_bengali_3.pptx (experiments in progress)</p>",
      "votes": null,
      "replies": [
        {
          "id": 760834,
          "author_name": "learnmower",
          "author_url": "",
          "post_date": "03/01/2020 19:23:08",
          "content": "<p>I included your dropblock but with cutmix and ended with 0.9879 at about 110.0 epochs. it seems to align closely with yours. Looking at your current experiment, it seems manifold progress is about where i was at that point of training. also saw that a square image was better. Thanks for sharing your experiments. I’ll try manifold now. I’m finding these experiments to be more interesting than the competition. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761062,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/02/2020 04:54:21",
          "content": "<p><a href=\"/learnmower\">@learnmower</a> </p>\n\n<p>if you are using dropblock, try different block size. i find that bigger block size may further improve in some cases. different size may be required for different layers as well?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761150,
          "author_name": "learnmower",
          "author_url": "",
          "post_date": "03/02/2020 07:36:18",
          "content": "<p>good suggestion... it makes sense to set different block sizes at different layers - perhaps some relative sizing to the convolution. also, gamma appears to determine location/occurrence of drop block, and I guess it would be a useful parameter to tune... </p>\n\n<p>i noticed that while drop block was developed at google brain in 2018, it seems that image augmentation based approaches to regularize neural nets have had more momentum in the last year... i wonder why is that the case?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761351,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/02/2020 12:17:04",
          "content": "<p>finalised report_bengali_3.pptx is in the top message above</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761277,
      "author_name": "jiangjiangzuijiang",
      "author_url": "",
      "post_date": "03/02/2020 10:55:13",
      "content": "<p>Thanks a lot to your exps, I always wonder whether 100+ epochs will impove the grapheme_root's score,  as I tried resnet34 for 50+  epochs and the score keeps around 0.956. In your slide I see after 100+ epochs it can reach around 0.98, so will score keep increase by large epochs?, and what loss did you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 761350,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/02/2020 12:16:25",
          "content": "<p>standard cross entropy loss</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761525,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/02/2020 16:13:02",
          "content": "<p>\"will score keep increase by large epochs?,\"</p>\n\n<p>not necessary. you may ignore the epoches in my logfile. they may not represent the \"actual number of epoches trained\" . sometimes i just learn my machine run over-night. sometime there are bugs in the code, or i may reset some hyper parameters, etc ... but i just let keep the epoch number running</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761615,
      "author_name": "souhai",
      "author_url": "",
      "post_date": "03/02/2020 18:45:36",
      "content": "<p>Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 761664,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "03/02/2020 20:37:57",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> , Thanks for the nice report, really informative. I'm curious about your experiments with dropblock applied to different resolutions: did you change the size of dropblock depending on the image size you use or considered a specific constant value?</p>",
      "votes": null,
      "replies": [
        {
          "id": 761676,
          "author_name": "authman",
          "author_url": "",
          "post_date": "03/02/2020 20:50:36",
          "content": "<p>In his most recent, he uses <code>block_size=10</code> after the first layer, and <code>block_size=5</code> after the second.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761706,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "03/02/2020 21:18:52",
          "content": "<p>Thanks for clarifications. I just thought that the relative box size with respect to the image size may play a role (rather than the absolute value of blocks), and the improvement for larger images may be not only attributed to the larger image size but also to slightly different regularization. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761787,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/02/2020 23:41:28",
          "content": "<p>use larger size for larger input. i tried up to 33% of the input size. 50% or more may work better, but i haven tried yet.</p>\n\n<p>it should be same as cutout?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 762155,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "03/03/2020 09:00:36",
      "content": "<p>May I ask why do you replace residual sum with max? Is there any reference paper?</p>",
      "votes": null,
      "replies": [
        {
          "id": 762193,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/03/2020 10:01:30",
          "content": "<p>it is by experiment. it is slightly better than then addition in some cases.</p>\n\n<p>i suggest you should stick to the normal  addition first, then conduct experiment to see if the change is suitable for your case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 762303,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "03/03/2020 12:16:53",
          "content": "<p>I see. Thanks for the info. I would have a try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763633,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/04/2020 17:18:58",
      "content": "<p>update:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe4777a0e0bae763b2d7f01b8dcb29b31%2FSelection_138.png?generation=1583342336635585&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 763972,
          "author_name": "lisosia",
          "author_url": "",
          "post_date": "03/05/2020 03:30:05",
          "content": "<p>Thank you for your sharing.\nDid you use surgeryed version of resnext? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764077,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "03/05/2020 05:49:19",
          "content": "<p>Dropblock doesn't work much for me. I'm not sure whether I need more epochs to train it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764599,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "03/05/2020 16:28:48",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I might be a bit late on that subject but i can see that you are classifying with 1295 classes but how did you figure out the 1295 classes if there are only 1292 combinations in the train set?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763978,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "03/05/2020 03:38:28",
      "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a>, really thanks for your kindly sharing. According to your report, i try seresnext50 with dropoutblock which the same as your ppt says, and i achieve CV 0.986 without any augments, but when i merge the cutmix + rotate with dropout, the result cv 0.9833 is worse than seresnext50 + cutmix + rotate(for me best cv 0.9913), and it seems these combination not benefit the fitting, and still need tuning, and it seems not just a question which can be solve by simply adding or reducing</p>",
      "votes": null,
      "replies": [
        {
          "id": 764001,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "03/05/2020 04:09:32",
          "content": "<p>Same for me, dropblock helped me a lot but when adding too hard augmentations it cant converge as much</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764013,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/05/2020 04:33:42",
          "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> </p>\n\n<p>i suppose cutmix  is not suitable to use with dropblock.\nmaybe you can try cutmix at feature map? i will call it mixblock</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764095,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "03/05/2020 06:15:59",
          "content": "<p>I guess dropblock and mixup are both strong regularization methods and their combinations caused underfitting. Maybe a deeper model would benefit more from that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764813,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/05/2020 23:41:39",
          "content": "<p>an alternative is that if you cannot fuse the different augmentation into a single network training, ensemble two classifiers trained separately on each?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 764628,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/05/2020 17:00:26",
      "content": "<p>yet another update\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F972371de7eefd559d01c5ca331cd4f26%2FSelection_156.png?generation=1583427623253253&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 764815,
          "author_name": "jiangjiangzuijiang",
          "author_url": "",
          "post_date": "03/05/2020 23:46:54",
          "content": "<p>How about the gap? I got a CV 0.982 but LB 0.969💔 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "759050": "part.0,1,2,3 of experiment as attached. more results later?",
    "759051": "coming up:\n1. modification of Robin Smits method:\n  https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974 \n\n- keep a fixed validation set\n\n- use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)",
    "759072": "Some day I will be able to understand your pptx reports))",
    "759120": "Thanks for your inspiring slides! I'm considering using multi-scale now. Would you mind telling me the LB score of your multi-scale experiment (experiment 3)?",
    "759138": "ivanwang\n\nI did not submit. but i can guess it would be around cv=0.977, lb=0.967\n\nthe experiment shows that scale is an important data variation.\n\nthere are many ways to use this information, e.g add scale augmentation in train (and test). \n\nmax pooling over scale is another solution but may be inefficient.\nyou can google for more efficient multi-scale network or  multi-scale input\n\nother possibility includes adding consisentcy loss:\n\nnet(image) = feature1\nnet(resized  image) = feature2\n\nfeature 1 should be same as feature2",
    "759164": "&gt; **Heng CherKeng wrote:**\n&gt; \n&gt; coming up:\n&gt; 1. modification of Robin Smits method:\n&gt;   https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974 \n&gt; \n&gt; - keep a fixed validation set\n&gt; \n&gt; - use 80%? of train set at each epoch. keep a moving average of learned weights (this reminds me of ema weights of style-gan)\n&gt; \n\n... modification of the Robin Smits method....that sounds cool ;-)\n\nVery nice overview @hengck23 what hardware do you have being able to produce such an amount of results so quickly? Likely not an 1070 Ti...",
    "759198": "I keep being impressed by these results without any augmentation!\n\nAnd same question as @rsmits, what are you using to get so many results? Both on the hardware level as well as some tricks you might have?",
    "759208": "rsmits \n\n\" Likely not an 1070 Ti…\"\n\nI have 4x1080Ti and 1xPascal(TitianX).\n\nTo make a fair comparison to see the effectiveness of the changing dataset per epoch, we should \n(1) train using all train\n(2) train using a fixed 80% of (1)\n(3) train using changing 80% of (1) \n\nfurther, we should test with fixed weight + moving average of weight for (1),(2),(3)",
    "759211": "If we want to train 3 times with 100% / 80% / 80% data, isn't it worth doing 3 or even 4-Fold validation for similar training times?\n\nAnd 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P",
    "759225": "Cool! 5 GPU's....that explains the speed perfectly ;-)\nLooking forward to your full investigation @hengck23  This is extremely usefull for all Kagglers. \nThank You!",
    "759313": "&gt; And 4x1080Tis, can't wait to stop being a student, get myself a decent salary and give it all to Nvidia :P\n\nTrust me, even after you stop being a student, it won't be that easy to rack up 4x2080TIs",
    "759414": "P106-100 6G headless mining card on ebay, $75-$100.  I got mine for $50.",
    "759465": "Great work",
    "759487": "In your experiment.3 you use dual scale streams with 64x112 &amp; 96x168 to get a better score, did you compare it with using 96x168 only? 🤔",
    "759493": "yes, the training is in progress. i also have scale1+scale2 instead of max(scale,1,scale2)",
    "759495": "new! experiment.5\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F267ed76905fb347c99cc4094980e55c8%2FSelection_087.png?generation=1582954406367873&amp;alt=media)",
    "759573": "haqishen \n\nplease see new results at report_bengali_1.pptx\n\n96x168 obtain the best results. Hence it is not required to use dual scale",
    "759574": "ivanwang2016 \n\nplease see new results at report_bengali_1.pptx\n\n96x168 obtain the best results. Hence it is not required to use dual scale",
    "759583": "Oh I very well know the prices of these puppies, and I'm not even planning on getting one 2080Ti, probably going to start a little bit more mid-range than that!",
    "759587": "Thanks for the result!",
    "759608": "maybe you can start with vast.ai 😃",
    "760072": "Thanks! That's quite an impressive work!",
    "760272": "3rd set of experiments are as in \"report\\_bengali\\_2.pptx (468.56 KB)\"\n\nwithout augmentation:\n(1) 128x128 on modified senext50+drop block : cv 0.985\n(2) 96x168 on modified senext50+drop block : cv 0.979 \n(3) 96x168 on modified senext50+drop block +2x label : cv 0.979 \n\nnew experiments that i am considering : \n-  imiplicit mixup : manifold mixup (mixing feature maps) or shakedrop?\n- 224x224",
    "760279": "hengck23 Thanks a lot for sharing your wonderful experiments. I am just wondering whether your current LB score is based on a single fold or ensemble of multiple folds.",
    "760622": "preview of report\\_bengali\\_3.pptx (experiments in progress)",
    "760834": "I included your dropblock but with cutmix and ended with 0.9879 at about 110.0 epochs. it seems to align closely with yours. Looking at your current experiment, it seems manifold progress is about where i was at that point of training. also saw that a square image was better. Thanks for sharing your experiments. I’ll try manifold now. I’m finding these experiments to be more interesting than the competition.",
    "761062": "learnmower \n\nif you are using dropblock, try different block size. i find that bigger block size may further improve in some cases. different size may be required for different layers as well?",
    "761150": "good suggestion... it makes sense to set different block sizes at different layers - perhaps some relative sizing to the convolution. also, gamma appears to determine location/occurrence of drop block, and I guess it would be a useful parameter to tune... \n\ni noticed that while drop block was developed at google brain in 2018, it seems that image augmentation based approaches to regularize neural nets have had more momentum in the last year... i wonder why is that the case?",
    "761277": "Thanks a lot to your exps, I always wonder whether 100+ epochs will impove the grapheme_root's score,  as I tried resnet34 for 50+  epochs and the score keeps around 0.956. In your slide I see after 100+ epochs it can reach around 0.98, so will score keep increase by large epochs?, and what loss did you use?",
    "761350": "standard cross entropy loss",
    "761351": "finalised report_bengali_3.pptx is in the top message above",
    "761525": "\"will score keep increase by large epochs?,\"\n\nnot necessary. you may ignore the epoches in my logfile. they may not represent the \"actual number of epoches trained\" . sometimes i just learn my machine run over-night. sometime there are bugs in the code, or i may reset some hyper parameters, etc ... but i just let keep the epoch number running",
    "761615": "Great work!",
    "761664": "hengck23 , Thanks for the nice report, really informative. I'm curious about your experiments with dropblock applied to different resolutions: did you change the size of dropblock depending on the image size you use or considered a specific constant value?",
    "761676": "In his most recent, he uses `block_size=10` after the first layer, and `block_size=5` after the second.",
    "761706": "Thanks for clarifications. I just thought that the relative box size with respect to the image size may play a role (rather than the absolute value of blocks), and the improvement for larger images may be not only attributed to the larger image size but also to slightly different regularization.",
    "761787": "use larger size for larger input. i tried up to 33% of the input size. 50% or more may work better, but i haven tried yet.\n\nit should be same as cutout?",
    "762155": "May I ask why do you replace residual sum with max? Is there any reference paper?",
    "762193": "it is by experiment. it is slightly better than then addition in some cases.\n\ni suggest you should stick to the normal  addition first, then conduct experiment to see if the change is suitable for your case.",
    "762303": "I see. Thanks for the info. I would have a try.",
    "763633": "update:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe4777a0e0bae763b2d7f01b8dcb29b31%2FSelection_138.png?generation=1583342336635585&amp;alt=media)",
    "763972": "Thank you for your sharing.\nDid you use surgeryed version of resnext?",
    "763978": "Hi @hengck23, really thanks for your kindly sharing. According to your report, i try seresnext50 with dropoutblock which the same as your ppt says, and i achieve CV 0.986 without any augments, but when i merge the cutmix + rotate with dropout, the result cv 0.9833 is worse than seresnext50 + cutmix + rotate(for me best cv 0.9913), and it seems these combination not benefit the fitting, and still need tuning, and it seems not just a question which can be solve by simply adding or reducing",
    "763995": "hengck23 I'm curious about why train with 2x class.",
    "764001": "Same for me, dropblock helped me a lot but when adding too hard augmentations it cant converge as much",
    "764013": "cswwp347724 \n\ni suppose cutmix  is not suitable to use with dropblock.\nmaybe you can try cutmix at feature map? i will call it mixblock",
    "764077": "Dropblock doesn't work much for me. I'm not sure whether I need more epochs to train it.",
    "764095": "I guess dropblock and mixup are both strong regularization methods and their combinations caused underfitting. Maybe a deeper model would benefit more from that.",
    "764599": "hengck23 I might be a bit late on that subject but i can see that you are classifying with 1295 classes but how did you figure out the 1295 classes if there are only 1292 combinations in the train set?",
    "764628": "yet another update\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F972371de7eefd559d01c5ca331cd4f26%2FSelection_156.png?generation=1583427623253253&amp;alt=media)",
    "764813": "an alternative is that if you cannot fuse the different augmentation into a single network training, ensemble two classifiers trained separately on each?",
    "764815": "How about the gap? I got a CV 0.982 but LB 0.969💔"
  },
  "source": "meta"
}