{
  "id": 124904,
  "title": "Solving grapheme_root is the key to 0.98?",
  "url": "/competitions/bengaliai-cv19/discussion/124904",
  "author_name": "Bibek",
  "post_date": "2020-01-07T12:17:21.914000",
  "votes": 55,
  "comment_count": 38,
  "views": 0,
  "content": "<p>In my experiments, the model does pretty well(<code>0.98+</code>on local CV) for <code>vowel_diacritic</code>  and <code>consonant_diacritic</code> but only <code>0.96</code> for <code>grapheme_root</code>. I guess this is the case for most of us. So why not discuss to improve our model's performance on <code>grapheme_root</code>?</p>\n\n<p>The only thing that I have tried so far is adding <code>weights</code> in loss function to penalize more on wrong predictions of <code>grapheme_root</code>. What other tricks could be applied? </p>",
  "messages": [
    {
      "id": 712591,
      "postDate": "2020-01-07T12:17:21.913Z",
      "content": "<p>In my experiments, the model does pretty well(<code>0.98+</code>on local CV) for <code>vowel_diacritic</code>  and <code>consonant_diacritic</code> but only <code>0.96</code> for <code>grapheme_root</code>. I guess this is the case for most of us. So why not discuss to improve our model's performance on <code>grapheme_root</code>?</p>\n\n<p>The only thing that I have tried so far is adding <code>weights</code> in loss function to penalize more on wrong predictions of <code>grapheme_root</code>. What other tricks could be applied? </p>",
      "rawMarkdown": "In my experiments, the model does pretty well(`0.98+`on local CV) for `vowel_diacritic`  and `consonant_diacritic` but only `0.96` for `grapheme_root`. I guess this is the case for most of us. So why not discuss to improve our model's performance on `grapheme_root`?\n\nThe only thing that I have tried so far is adding `weights` in loss function to penalize more on wrong predictions of `grapheme_root`. What other tricks could be applied? ",
      "votes": 54
    },
    {
      "id": 712733,
      "postDate": "2020-01-07T14:42:55.537Z",
      "content": "<p>Good Topic! Thanks for creating!</p>\n\n<p>You are absolutely right regarding <code>grapheme roots</code>. My current scoring metric score for this class is <code>0.967307</code> (rest two are around <code>0,98</code>). What helped me is to increase score from <code>0.962</code> to <code>0.967'  for</code>grapheme roots` is MixUp training and tweaking a bit architecture. </p>\n\n<p>This is my list I and am going one by one to see if improvement happens. \n<code>\n-random erasing \n-adding attention layer\n-cutmix\n-larger networks ?\n</code></p>\n\n<p>Will post updates =) </p>\n\n<p>Edit:\nMaybe people will find it useful regarding mixUp training. I use some ideas from this paper <a href=\"https://arxiv.org/abs/1912.11370\">https://arxiv.org/abs/1912.11370</a>. They are currently have lowest Imagenet error on the leaderboard.</p>\n\n<p>When training with MixUp.\n<code>\n-no dropouts\n-no weight decays\n-only horizontal flips\n</code></p>\n\n<p>They also suggest to replace <code>BN</code> with <code>Group Normalization</code>and add <code>Weights Standardization</code> for <code>Conv2d</code> and train with <code>SGD</code>.</p>",
      "rawMarkdown": "Good Topic! Thanks for creating!\n\nYou are absolutely right regarding `grapheme roots`. My current scoring metric score for this class is `0.967307` (rest two are around `0,98`). What helped me is to increase score from `0.962` to `0.967'  for   `grapheme roots` is MixUp training and tweaking a bit architecture. \n\nThis is my list I and am going one by one to see if improvement happens. \n```\n-random erasing \n-adding attention layer\n-cutmix\n-larger networks ?\n```\n\nWill post updates =) \n\nEdit:\nMaybe people will find it useful regarding mixUp training. I use some ideas from this paper https://arxiv.org/abs/1912.11370. They are currently have lowest Imagenet error on the leaderboard.\n\nWhen training with MixUp.\n```\n-no dropouts\n-no weight decays\n-only horizontal flips\n```\n\nThey also suggest to replace `BN` with `Group Normalization `and add `Weights Standardization` for `Conv2d` and train with `SGD`.",
      "votes": 35,
      "replies": [
        {
          "id": 712884,
          "postDate": "2020-01-07T17:11:53.050Z",
          "content": "<p>Here is a nice repo with implementations of all the techniques mentioned above.</p>\n\n<blockquote>\n  <p><a href=\"https://github.com/hysts/pytorch_image_classification/tree/master/augmentations\">https://github.com/hysts/pytorch_image_classification/tree/master/augmentations</a></p>\n</blockquote>",
          "rawMarkdown": "Here is a nice repo with implementations of all the techniques mentioned above.\n&gt; https://github.com/hysts/pytorch_image_classification/tree/master/augmentations",
          "votes": 7
        },
        {
          "id": 714082,
          "postDate": "2020-01-09T02:53:19.630Z",
          "content": "<p>Hey! how would you implement that in our case? I'm new to this kind of data augments.. just trying to learn</p>",
          "rawMarkdown": "Hey! how would you implement that in our case? I'm new to this kind of data augments.. just trying to learn",
          "votes": 1
        },
        {
          "id": 714170,
          "postDate": "2020-01-09T06:22:04.660Z",
          "content": "<p>I'm also new to this. I am learning now and hopefully will post a kernel about it once I fully understand it</p>",
          "rawMarkdown": "I'm also new to this. I am learning now and hopefully will post a kernel about it once I fully understand it",
          "votes": 2
        },
        {
          "id": 714747,
          "postDate": "2020-01-09T17:55:41.717Z",
          "content": "<p>Using group norm requires us to specify the number of groups. What do u think is a reasonable number of groups?</p>",
          "rawMarkdown": "Using group norm requires us to specify the number of groups. What do u think is a reasonable number of groups?",
          "votes": 1
        },
        {
          "id": 714753,
          "postDate": "2020-01-09T18:03:22.917Z",
          "content": "<p>The code of this paper hasn't released. But In the original paper they were using <code>32</code> you can read more here: <a href=\"https://arxiv.org/pdf/1803.08494.pdf\">https://arxiv.org/pdf/1803.08494.pdf</a></p>",
          "rawMarkdown": "The code of this paper hasn't released. But In the original paper they were using `32` you can read more here: https://arxiv.org/pdf/1803.08494.pdf",
          "votes": 1
        },
        {
          "id": 714908,
          "postDate": "2020-01-09T23:03:35.493Z",
          "content": "<p>Dear DrHB, I managed to figure out how to do mix up training however i think my model is under fitting, i'm having trouble reaching good accuracy, do you have any tips for me? It would be so appreciated! :)</p>",
          "rawMarkdown": " Dear DrHB, I managed to figure out how to do mix up training however i think my model is under fitting, i'm having trouble reaching good accuracy, do you have any tips for me? It would be so appreciated! :)",
          "votes": 1
        },
        {
          "id": 714930,
          "postDate": "2020-01-10T00:02:17.123Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> train it longer. mixup needs lots of epochs (~150-200)</p>",
          "rawMarkdown": "@yannmajewski train it longer. mixup needs lots of epochs (~150-200)",
          "votes": 5
        },
        {
          "id": 714936,
          "postDate": "2020-01-10T00:21:47.457Z",
          "content": "<p>Peter gave good advice =) \nGenerally Mixup requires longer training.</p>\n\n<p>Here are two thing that might improve training:</p>\n\n<p>1) Try to use minimal <code>augmentations</code>. If you read article that I posted above they have a nice table and they recommended only <code>horizontal flips</code> (only if necessary) and that pretty much it. Also <code>mixup</code> is shown to generalize better (without a lot of  augmentations) on unseen data or data which is out of distribution. This is relevant for this competition (you can read here <a href=\"https://openreview.net/pdf?id=rJgxnSHg8r\">https://openreview.net/pdf?id=rJgxnSHg8r</a>) </p>\n\n<p>2)You can try to use modern optimizers to speed up training.  I haven't done experiments and also I cant find any articles where modern optimizers are used e.g Like <code>Lookahead</code>, or <code>Ranger</code> with <code>Mixup</code>. Since new optimizers tends to find global minima faster maybe they can speed up mixup training. Again I have zero experience in this field =)</p>\n\n<p>Good luck =) </p>",
          "rawMarkdown": "Peter gave good advice =) \nGenerally Mixup requires longer training.\n\nHere are two thing that might improve training:\n\n1) Try to use minimal `augmentations`. If you read article that I posted above they have a nice table and they recommended only `horizontal flips` (only if necessary) and that pretty much it. Also `mixup` is shown to generalize better (without a lot of  augmentations) on unseen data or data which is out of distribution. This is relevant for this competition (you can read here https://openreview.net/pdf?id=rJgxnSHg8r) \n\n2)You can try to use modern optimizers to speed up training.  I haven't done experiments and also I cant find any articles where modern optimizers are used e.g Like `Lookahead`, or `Ranger` with `Mixup`. Since new optimizers tends to find global minima faster maybe they can speed up mixup training. Again I have zero experience in this field =)\n\nGood luck =) ",
          "votes": 5
        },
        {
          "id": 714949,
          "postDate": "2020-01-10T01:04:33.360Z",
          "content": "<p>Thanks for both of your answers! I think you guys are right, i was only doing 40 epochs with Over9000 optim and got 95% and had a random rotation. I also reduced the alpha parameter which seems to help.</p>\n\n<p>I will test with more epochs and maybe SGD with Lookahead will give a better result!</p>",
          "rawMarkdown": "Thanks for both of your answers! I think you guys are right, i was only doing 40 epochs with Over9000 optim and got 95% and had a random rotation. I also reduced the alpha parameter which seems to help.\n \nI will test with more epochs and maybe SGD with Lookahead will give a better result!",
          "votes": 2
        },
        {
          "id": 714996,
          "postDate": "2020-01-10T02:32:13.003Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 715002,
          "postDate": "2020-01-10T02:37:02.453Z",
          "content": "<p>You also try this <a href=\"https://arxiv.org/abs/1911.09737\">Filter Response Normalization</a> instead of <code>BN</code>. Looks promising </p>",
          "rawMarkdown": "You also try this [Filter Response Normalization](https://arxiv.org/abs/1911.09737) instead of `BN`. Looks promising "
        },
        {
          "id": 719073,
          "postDate": "2020-01-15T05:23:16.817Z",
          "content": "<p>This is just a general question, is mix up only good for big data sets or can it work on small datasets?</p>",
          "rawMarkdown": "This is just a general question, is mix up only good for big data sets or can it work on small datasets?"
        },
        {
          "id": 719231,
          "postDate": "2020-01-15T09:07:58.690Z",
          "content": "<p>the best way to know would be to try it yourself and see if it works.</p>",
          "rawMarkdown": "the best way to know would be to try it yourself and see if it works.",
          "votes": 2
        },
        {
          "id": 726742,
          "postDate": "2020-01-23T07:59:29.057Z",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a>! Thanks for sharing this information!</p>\n\n<p>I have one question.</p>\n\n<p>If I want to turn off dropout, say in EfficientNet, is it enough to set the value of <code>efnet_model._dropout.p=0.0</code>?  (for Resnet's this would be iterate over all dropout layers and set <code>p=0.0</code>)</p>\n\n<p>Or should I copy all the layers and its weights except Dropout layers and make a new model instead?</p>\n\n<p>Another option I've seen is to iterate over all layers and call <code>eval()</code> method on dropout layers, though there are enough issues on Github stating that <code>eval()</code> method doesn't work properly.</p>\n\n<p>Not sure what is the right option here. </p>",
          "rawMarkdown": "Hi @drhabib! Thanks for sharing this information!\n\nI have one question.\n\nIf I want to turn off dropout, say in EfficientNet, is it enough to set the value of `efnet_model._dropout.p=0.0`?  (for Resnet's this would be iterate over all dropout layers and set `p=0.0`)\n\nOr should I copy all the layers and its weights except Dropout layers and make a new model instead?\n\nAnother option I've seen is to iterate over all layers and call `eval()` method on dropout layers, though there are enough issues on Github stating that `eval()` method doesn't work properly.\n\nNot sure what is the right option here. "
        },
        {
          "id": 727193,
          "postDate": "2020-01-23T14:20:05.347Z",
          "content": "<p>Hi <a href=\"/lightnezzofbeing\">@lightnezzofbeing</a> </p>\n\n<p>@Iafoss has excellent high scoring kernel, if you carefully analyze, it has function called <code>to_Mish</code> this function finds all <code>nn.ReLU</code> and substitutes it to <code>Mish</code>. We can rewrite this function to find and replace dropout probability. </p>\n\n<p>```\ndef dropout_replace(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.Dropout):\n            setattr(model, child_name, nn.Dropout(p=0.))\n        else:\n            dropout_replace(child)</p>\n\n<p>```\nI think this should work. </p>",
          "rawMarkdown": "Hi @lightnezzofbeing \n\n\n@Iafoss has excellent high scoring kernel, if you carefully analyze, it has function called `to_Mish` this function finds all `nn.ReLU` and substitutes it to `Mish`. We can rewrite this function to find and replace dropout probability. \n\n```\ndef dropout_replace(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.Dropout):\n            setattr(model, child_name, nn.Dropout(p=0.))\n        else:\n            dropout_replace(child)\n\n```\nI think this should work. \n",
          "votes": 9
        },
        {
          "id": 728097,
          "postDate": "2020-01-24T11:56:09.883Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> \nThanks for your answer!</p>\n\n<p>By the way, I learned a lot from your posts on forum. Keep it up! </p>",
          "rawMarkdown": "@drhabib \nThanks for your answer!\n\nBy the way, I learned a lot from your posts on forum. Keep it up! \n\n",
          "votes": 2
        },
        {
          "id": 728216,
          "postDate": "2020-01-24T14:07:58.680Z",
          "content": "<p>Thanks for the kind words =) </p>",
          "rawMarkdown": "Thanks for the kind words =) ",
          "votes": 3
        }
      ]
    },
    {
      "id": 738651,
      "postDate": "2020-02-06T20:32:26.613Z",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> can you share some tips you've used to get high score? I just wonder have you trained model at your own pc then just used kaggle kernel for submission on loaded weights or something else? This question may be strange for guy with experience in computer vision but I don't so much practice in that direction of data science yet.</p>",
      "rawMarkdown": "@drhabib can you share some tips you've used to get high score? I just wonder have you trained model at your own pc then just used kaggle kernel for submission on loaded weights or something else? This question may be strange for guy with experience in computer vision but I don't so much practice in that direction of data science yet.",
      "votes": 3,
      "replies": [
        {
          "id": 738653,
          "postDate": "2020-02-06T20:42:05.010Z",
          "content": "<p>The best way I think is to train your model in offline and save the optimized weights and upload that weights file to Kernel. I don't prefer to train any complex model in Kernel because of its resource limitation. </p>",
          "rawMarkdown": "The best way I think is to train your model in offline and save the optimized weights and upload that weights file to Kernel. I don't prefer to train any complex model in Kernel because of its resource limitation. ",
          "votes": 3
        },
        {
          "id": 738661,
          "postDate": "2020-02-06T20:46:11.373Z",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> thx for advice, it will save some time for me!</p>",
          "rawMarkdown": "@ipythonx thx for advice, it will save some time for me!",
          "votes": 1
        },
        {
          "id": 738667,
          "postDate": "2020-02-06T21:00:16.820Z",
          "content": "<p>Hi Dmitry, \nThanks for asking questions. I use my local machine which has 2 old gpus =) My model is se_resnext50.</p>\n\n<p>Bonus info because I got today my green card approved =): </p>\n\n<p>It seems like in this competition there is no magic. Based on discussion and personal experince so far you can get high score using some combination of augmentations mixup, cutmix, cutout and etc (i might be wrong). Its hard to give specific advice on tips but I can tell you how I am approching this problem. Maybe this will be useful for you. So far I am using my main model which <code>se_resnext_50</code> and trying one by one different ideas. You can get list of ideas from this post: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127976#731262\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127976#731262</a> . If you are new user or have little experince I would suggest take top scoring kernal and try to modify it slowly. </p>\n\n<p>Every time you make a small modification or test idea, make sure to keep \"LAB BOOK\" on what you have done. Sometimes when you are out of ideas just looking on well documented \"LAB BOOK\" could results in some new insparation. </p>\n\n<p>Good luck! </p>",
          "rawMarkdown": "Hi Dmitry, \nThanks for asking questions. I use my local machine which has 2 old gpus =) My model is se_resnext50.\n\nBonus info because I got today my green card approved =): \n\nIt seems like in this competition there is no magic. Based on discussion and personal experince so far you can get high score using some combination of augmentations mixup, cutmix, cutout and etc (i might be wrong). Its hard to give specific advice on tips but I can tell you how I am approching this problem. Maybe this will be useful for you. So far I am using my main model which `se_resnext_50` and trying one by one different ideas. You can get list of ideas from this post: https://www.kaggle.com/c/bengaliai-cv19/discussion/127976#731262 . If you are new user or have little experince I would suggest take top scoring kernal and try to modify it slowly. \n\nEvery time you make a small modification or test idea, make sure to keep \"LAB BOOK\" on what you have done. Sometimes when you are out of ideas just looking on well documented \"LAB BOOK\" could results in some new insparation. \n\nGood luck! \n",
          "votes": 27
        },
        {
          "id": 738902,
          "postDate": "2020-02-07T06:42:25.743Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks for your advices. I have big experience working with tensorflow but not in computer vision (mostly in regression and classification based on 2-dimension datasets), so it's a good chance to learn something new))</p>",
          "rawMarkdown": "@drhabib thanks for your advices. I have big experience working with tensorflow but not in computer vision (mostly in regression and classification based on 2-dimension datasets), so it's a good chance to learn something new))"
        },
        {
          "id": 740902,
          "postDate": "2020-02-09T23:50:09.773Z",
          "content": "<p>Great infos, thanks <a href=\"/drhabib\">@drhabib</a>.\nAnyways congrats on your green!</p>",
          "rawMarkdown": "Great infos, thanks @drhabib.\nAnyways congrats on your green!",
          "votes": 1
        },
        {
          "id": 740923,
          "postDate": "2020-02-10T00:31:49.627Z",
          "content": "<p>Seems like the only trick for this comp is just being able to train a model for 100-150 epochs.</p>",
          "rawMarkdown": "Seems like the only trick for this comp is just being able to train a model for 100-150 epochs.",
          "votes": 2
        },
        {
          "id": 743785,
          "postDate": "2020-02-12T10:01:59.853Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fdff401277f4b488c1a717073980ed6a3%2Ftrainhard.jpg?generation=1581501706068894&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fdff401277f4b488c1a717073980ed6a3%2Ftrainhard.jpg?generation=1581501706068894&amp;alt=media)\n",
          "votes": 5
        }
      ]
    },
    {
      "id": 712603,
      "postDate": "2020-01-07T12:47:23.753Z",
      "content": "<p>yea that's obvious since it has highest error and doubled in kaggle metric</p>",
      "rawMarkdown": "yea that's obvious since it has highest error and doubled in kaggle metric",
      "votes": 2
    },
    {
      "id": 766643,
      "postDate": "2020-03-08T14:06:16.210Z",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> hello, I am also facing such a scenario about <code>grapheme_root</code> class; what I am thinking now is to account the following things:</p>\n\n<ul>\n<li>tweaking model arch. </li>\n<li>regularization (dropout, LR)</li>\n</ul>\n\n<p>do you have any suggestions on this? What are the possible things that can be taken to tackle this issue? Thanks. </p>",
      "rawMarkdown": "@bibek777 hello, I am also facing such a scenario about `grapheme_root` class; what I am thinking now is to account the following things:\n\n- tweaking model arch. \n- regularization (dropout, LR)\n\ndo you have any suggestions on this? What are the possible things that can be taken to tackle this issue? Thanks. ",
      "votes": 1
    },
    {
      "id": 745858,
      "postDate": "2020-02-14T09:42:00.237Z",
      "content": "<p>For me, adding weights in loss function doesn't make a difference.</p>",
      "rawMarkdown": "For me, adding weights in loss function doesn't make a difference.",
      "votes": 1
    },
    {
      "id": 730014,
      "postDate": "2020-01-27T01:22:50.923Z",
      "content": "<p>Do you guys do data-preprocessing? I got to 0.97 without preprocessing (just cv2.resize to 64x64), but since there is usually a lot of empty space on the images, I wonder if people doing better actually preprocess their images</p>",
      "rawMarkdown": "Do you guys do data-preprocessing? I got to 0.97 without preprocessing (just cv2.resize to 64x64), but since there is usually a lot of empty space on the images, I wonder if people doing better actually preprocess their images",
      "votes": 1
    },
    {
      "id": 712716,
      "postDate": "2020-01-07T14:27:58.143Z",
      "content": "<p>Some grapheme roots are lower in number. Should we not equalize the number of example images per grapheme root? Again, some are visually similar, so the network in getting confused.</p>",
      "rawMarkdown": "Some grapheme roots are lower in number. Should we not equalize the number of example images per grapheme root? Again, some are visually similar, so the network in getting confused.",
      "votes": 1
    },
    {
      "id": 714062,
      "postDate": "2020-01-09T02:10:00.103Z",
      "content": "<p>Good question, Train a separate model for grapheme_root only maybe improve</p>",
      "rawMarkdown": "Good question, Train a separate model for grapheme_root only maybe improve",
      "votes": 2,
      "replies": [
        {
          "id": 714168,
          "postDate": "2020-01-09T06:20:47.660Z",
          "content": "<p>Nice Idea!! I haven't tried that yet. Did you have any success with that?</p>",
          "rawMarkdown": "Nice Idea!! I haven't tried that yet. Did you have any success with that?"
        },
        {
          "id": 714950,
          "postDate": "2020-01-10T01:05:10.073Z",
          "content": "<p>Not yet, I will try</p>",
          "rawMarkdown": "Not yet, I will try"
        },
        {
          "id": 716717,
          "postDate": "2020-01-12T06:29:06.333Z",
          "content": "<p>I was thinking about that too. Can you tell us about your progress once you tried that? Much thanks!</p>",
          "rawMarkdown": "I was thinking about that too. Can you tell us about your progress once you tried that? Much thanks!",
          "votes": 1
        },
        {
          "id": 717018,
          "postDate": "2020-01-12T16:07:07.463Z",
          "content": "<p>I'm getting 0.98 for vowel and cons diacritics but stuck at 0.93~0.94 for grapheme_root. Any advice?</p>",
          "rawMarkdown": "I'm getting 0.98 for vowel and cons diacritics but stuck at 0.93~0.94 for grapheme_root. Any advice?",
          "votes": 1
        }
      ]
    },
    {
      "id": 762726,
      "postDate": "2020-03-03T19:06:53.237Z",
      "content": "<p>Did anyone figure out a way?</p>",
      "rawMarkdown": "Did anyone figure out a way?"
    },
    {
      "id": 719914,
      "postDate": "2020-01-16T01:26:00.777Z",
      "content": "<p>i dont think so</p>",
      "rawMarkdown": "i dont think so"
    }
  ],
  "comments": [
    {
      "id": 712733,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-01-07T14:42:55.537000",
      "content": "<p>Good Topic! Thanks for creating!</p>\n\n<p>You are absolutely right regarding <code>grapheme roots</code>. My current scoring metric score for this class is <code>0.967307</code> (rest two are around <code>0,98</code>). What helped me is to increase score from <code>0.962</code> to <code>0.967'  for</code>grapheme roots` is MixUp training and tweaking a bit architecture. </p>\n\n<p>This is my list I and am going one by one to see if improvement happens. \n<code>\n-random erasing \n-adding attention layer\n-cutmix\n-larger networks ?\n</code></p>\n\n<p>Will post updates =) </p>\n\n<p>Edit:\nMaybe people will find it useful regarding mixUp training. I use some ideas from this paper <a href=\"https://arxiv.org/abs/1912.11370\">https://arxiv.org/abs/1912.11370</a>. They are currently have lowest Imagenet error on the leaderboard.</p>\n\n<p>When training with MixUp.\n<code>\n-no dropouts\n-no weight decays\n-only horizontal flips\n</code></p>\n\n<p>They also suggest to replace <code>BN</code> with <code>Group Normalization</code>and add <code>Weights Standardization</code> for <code>Conv2d</code> and train with <code>SGD</code>.</p>",
      "votes": 35,
      "replies": [
        {
          "id": 712884,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-07T17:11:53.050000",
          "content": "<p>Here is a nice repo with implementations of all the techniques mentioned above.</p>\n\n<blockquote>\n  <p><a href=\"https://github.com/hysts/pytorch_image_classification/tree/master/augmentations\">https://github.com/hysts/pytorch_image_classification/tree/master/augmentations</a></p>\n</blockquote>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 714082,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-09T02:53:19.630000",
          "content": "<p>Hey! how would you implement that in our case? I'm new to this kind of data augments.. just trying to learn</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 714170,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-09T06:22:04.660000",
          "content": "<p>I'm also new to this. I am learning now and hopefully will post a kernel about it once I fully understand it</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 714747,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2020-01-09T17:55:41.717000",
          "content": "<p>Using group norm requires us to specify the number of groups. What do u think is a reasonable number of groups?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 714753,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-09T18:03:22.917000",
          "content": "<p>The code of this paper hasn't released. But In the original paper they were using <code>32</code> you can read more here: <a href=\"https://arxiv.org/pdf/1803.08494.pdf\">https://arxiv.org/pdf/1803.08494.pdf</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 714908,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-09T23:03:35.493000",
          "content": "<p>Dear DrHB, I managed to figure out how to do mix up training however i think my model is under fitting, i'm having trouble reaching good accuracy, do you have any tips for me? It would be so appreciated! :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 714930,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-01-10T00:02:17.123000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> train it longer. mixup needs lots of epochs (~150-200)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 714936,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-10T00:21:47.457000",
          "content": "<p>Peter gave good advice =) \nGenerally Mixup requires longer training.</p>\n\n<p>Here are two thing that might improve training:</p>\n\n<p>1) Try to use minimal <code>augmentations</code>. If you read article that I posted above they have a nice table and they recommended only <code>horizontal flips</code> (only if necessary) and that pretty much it. Also <code>mixup</code> is shown to generalize better (without a lot of  augmentations) on unseen data or data which is out of distribution. This is relevant for this competition (you can read here <a href=\"https://openreview.net/pdf?id=rJgxnSHg8r\">https://openreview.net/pdf?id=rJgxnSHg8r</a>) </p>\n\n<p>2)You can try to use modern optimizers to speed up training.  I haven't done experiments and also I cant find any articles where modern optimizers are used e.g Like <code>Lookahead</code>, or <code>Ranger</code> with <code>Mixup</code>. Since new optimizers tends to find global minima faster maybe they can speed up mixup training. Again I have zero experience in this field =)</p>\n\n<p>Good luck =) </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 714949,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-10T01:04:33.360000",
          "content": "<p>Thanks for both of your answers! I think you guys are right, i was only doing 40 epochs with Over9000 optim and got 95% and had a random rotation. I also reduced the alpha parameter which seems to help.</p>\n\n<p>I will test with more epochs and maybe SGD with Lookahead will give a better result!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 714996,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-10T02:32:13.003000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715002,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-10T02:37:02.453000",
          "content": "<p>You also try this <a href=\"https://arxiv.org/abs/1911.09737\">Filter Response Normalization</a> instead of <code>BN</code>. Looks promising </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 719073,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-15T05:23:16.817000",
          "content": "<p>This is just a general question, is mix up only good for big data sets or can it work on small datasets?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 719231,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-15T09:07:58.690000",
          "content": "<p>the best way to know would be to try it yourself and see if it works.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 726742,
          "author_name": "Cyr1ll",
          "author_url": "",
          "post_date": "2020-01-23T07:59:29.057000",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a>! Thanks for sharing this information!</p>\n\n<p>I have one question.</p>\n\n<p>If I want to turn off dropout, say in EfficientNet, is it enough to set the value of <code>efnet_model._dropout.p=0.0</code>?  (for Resnet's this would be iterate over all dropout layers and set <code>p=0.0</code>)</p>\n\n<p>Or should I copy all the layers and its weights except Dropout layers and make a new model instead?</p>\n\n<p>Another option I've seen is to iterate over all layers and call <code>eval()</code> method on dropout layers, though there are enough issues on Github stating that <code>eval()</code> method doesn't work properly.</p>\n\n<p>Not sure what is the right option here. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 727193,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-23T14:20:05.347000",
          "content": "<p>Hi <a href=\"/lightnezzofbeing\">@lightnezzofbeing</a> </p>\n\n<p>@Iafoss has excellent high scoring kernel, if you carefully analyze, it has function called <code>to_Mish</code> this function finds all <code>nn.ReLU</code> and substitutes it to <code>Mish</code>. We can rewrite this function to find and replace dropout probability. </p>\n\n<p>```\ndef dropout_replace(model):\n    for child_name, child in model.named_children():\n        if isinstance(child, nn.Dropout):\n            setattr(model, child_name, nn.Dropout(p=0.))\n        else:\n            dropout_replace(child)</p>\n\n<p>```\nI think this should work. </p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 728097,
          "author_name": "Cyr1ll",
          "author_url": "",
          "post_date": "2020-01-24T11:56:09.883000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> \nThanks for your answer!</p>\n\n<p>By the way, I learned a lot from your posts on forum. Keep it up! </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 728216,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-24T14:07:58.680000",
          "content": "<p>Thanks for the kind words =) </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 738651,
      "author_name": "Dmitry Kravchuk",
      "author_url": "",
      "post_date": "2020-02-06T20:32:26.613000",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> can you share some tips you've used to get high score? I just wonder have you trained model at your own pc then just used kaggle kernel for submission on loaded weights or something else? This question may be strange for guy with experience in computer vision but I don't so much practice in that direction of data science yet.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 738653,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-06T20:42:05.010000",
          "content": "<p>The best way I think is to train your model in offline and save the optimized weights and upload that weights file to Kernel. I don't prefer to train any complex model in Kernel because of its resource limitation. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 738661,
          "author_name": "Dmitry Kravchuk",
          "author_url": "",
          "post_date": "2020-02-06T20:46:11.373000",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> thx for advice, it will save some time for me!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 738667,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-02-06T21:00:16.820000",
          "content": "<p>Hi Dmitry, \nThanks for asking questions. I use my local machine which has 2 old gpus =) My model is se_resnext50.</p>\n\n<p>Bonus info because I got today my green card approved =): </p>\n\n<p>It seems like in this competition there is no magic. Based on discussion and personal experince so far you can get high score using some combination of augmentations mixup, cutmix, cutout and etc (i might be wrong). Its hard to give specific advice on tips but I can tell you how I am approching this problem. Maybe this will be useful for you. So far I am using my main model which <code>se_resnext_50</code> and trying one by one different ideas. You can get list of ideas from this post: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127976#731262\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127976#731262</a> . If you are new user or have little experince I would suggest take top scoring kernal and try to modify it slowly. </p>\n\n<p>Every time you make a small modification or test idea, make sure to keep \"LAB BOOK\" on what you have done. Sometimes when you are out of ideas just looking on well documented \"LAB BOOK\" could results in some new insparation. </p>\n\n<p>Good luck! </p>",
          "votes": 27,
          "replies": []
        },
        {
          "id": 738902,
          "author_name": "Dmitry Kravchuk",
          "author_url": "",
          "post_date": "2020-02-07T06:42:25.743000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks for your advices. I have big experience working with tensorflow but not in computer vision (mostly in regression and classification based on 2-dimension datasets), so it's a good chance to learn something new))</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 740902,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2020-02-09T23:50:09.773000",
          "content": "<p>Great infos, thanks <a href=\"/drhabib\">@drhabib</a>.\nAnyways congrats on your green!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 740923,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-10T00:31:49.627000",
          "content": "<p>Seems like the only trick for this comp is just being able to train a model for 100-150 epochs.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 743785,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-02-12T10:01:59.853000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fdff401277f4b488c1a717073980ed6a3%2Ftrainhard.jpg?generation=1581501706068894&amp;alt=media\" alt=\"\"></p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 712603,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-01-07T12:47:23.753000",
      "content": "<p>yea that's obvious since it has highest error and doubled in kaggle metric</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 766643,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-03-08T14:06:16.210000",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> hello, I am also facing such a scenario about <code>grapheme_root</code> class; what I am thinking now is to account the following things:</p>\n\n<ul>\n<li>tweaking model arch. </li>\n<li>regularization (dropout, LR)</li>\n</ul>\n\n<p>do you have any suggestions on this? What are the possible things that can be taken to tackle this issue? Thanks. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 745858,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2020-02-14T09:42:00.237000",
      "content": "<p>For me, adding weights in loss function doesn't make a difference.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 730014,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-01-27T01:22:50.923000",
      "content": "<p>Do you guys do data-preprocessing? I got to 0.97 without preprocessing (just cv2.resize to 64x64), but since there is usually a lot of empty space on the images, I wonder if people doing better actually preprocess their images</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 712716,
      "author_name": "Shayekh Islam",
      "author_url": "",
      "post_date": "2020-01-07T14:27:58.143000",
      "content": "<p>Some grapheme roots are lower in number. Should we not equalize the number of example images per grapheme root? Again, some are visually similar, so the network in getting confused.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 714062,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2020-01-09T02:10:00.103000",
      "content": "<p>Good question, Train a separate model for grapheme_root only maybe improve</p>",
      "votes": 2,
      "replies": [
        {
          "id": 714168,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-09T06:20:47.660000",
          "content": "<p>Nice Idea!! I haven't tried that yet. Did you have any success with that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714950,
          "author_name": "liuzhangzhen",
          "author_url": "",
          "post_date": "2020-01-10T01:05:10.073000",
          "content": "<p>Not yet, I will try</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716717,
          "author_name": "Stephen Lau",
          "author_url": "",
          "post_date": "2020-01-12T06:29:06.333000",
          "content": "<p>I was thinking about that too. Can you tell us about your progress once you tried that? Much thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717018,
          "author_name": "Rafid Abyaad",
          "author_url": "",
          "post_date": "2020-01-12T16:07:07.463000",
          "content": "<p>I'm getting 0.98 for vowel and cons diacritics but stuck at 0.93~0.94 for grapheme_root. Any advice?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 762726,
      "author_name": "Son of Anton v3.0",
      "author_url": "",
      "post_date": "2020-03-03T19:06:53.237000",
      "content": "<p>Did anyone figure out a way?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 719914,
      "author_name": "Chen Jia Jeng",
      "author_url": "",
      "post_date": "2020-01-16T01:26:00.777000",
      "content": "<p>i dont think so</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "712591": "In my experiments, the model does pretty well(`0.98+`on local CV) for `vowel_diacritic`  and `consonant_diacritic` but only `0.96` for `grapheme_root`. I guess this is the case for most of us. So why not discuss to improve our model's performance on `grapheme_root`?\n\nThe only thing that I have tried so far is adding `weights` in loss function to penalize more on wrong predictions of `grapheme_root`. What other tricks could be applied? ",
    "712733": "Good Topic! Thanks for creating!\n\nYou are absolutely right regarding `grapheme roots`. My current scoring metric score for this class is `0.967307` (rest two are around `0,98`). What helped me is to increase score from `0.962` to `0.967'  for   `grapheme roots` is MixUp training and tweaking a bit architecture. \n\nThis is my list I and am going one by one to see if improvement happens. \n```\n-random erasing \n-adding attention layer\n-cutmix\n-larger networks ?\n```\n\nWill post updates =) \n\nEdit:\nMaybe people will find it useful regarding mixUp training. I use some ideas from this paper https://arxiv.org/abs/1912.11370. They are currently have lowest Imagenet error on the leaderboard.\n\nWhen training with MixUp.\n```\n-no dropouts\n-no weight decays\n-only horizontal flips\n```\n\nThey also suggest to replace `BN` with `Group Normalization `and add `Weights Standardization` for `Conv2d` and train with `SGD`.",
    "738651": "@drhabib can you share some tips you've used to get high score? I just wonder have you trained model at your own pc then just used kaggle kernel for submission on loaded weights or something else? This question may be strange for guy with experience in computer vision but I don't so much practice in that direction of data science yet.",
    "712603": "yea that's obvious since it has highest error and doubled in kaggle metric",
    "766643": "@bibek777 hello, I am also facing such a scenario about `grapheme_root` class; what I am thinking now is to account the following things:\n\n- tweaking model arch. \n- regularization (dropout, LR)\n\ndo you have any suggestions on this? What are the possible things that can be taken to tackle this issue? Thanks. ",
    "745858": "For me, adding weights in loss function doesn't make a difference.",
    "730014": "Do you guys do data-preprocessing? I got to 0.97 without preprocessing (just cv2.resize to 64x64), but since there is usually a lot of empty space on the images, I wonder if people doing better actually preprocess their images",
    "712716": "Some grapheme roots are lower in number. Should we not equalize the number of example images per grapheme root? Again, some are visually similar, so the network in getting confused.",
    "714062": "Good question, Train a separate model for grapheme_root only maybe improve",
    "762726": "Did anyone figure out a way?",
    "719914": "i dont think so"
  }
}