{
  "id": 131734,
  "title": "Other 7 things that did not worked (part 2)",
  "url": "/competitions/bengaliai-cv19/discussion/131734",
  "author_name": "Vlad Vaduva",
  "post_date": "2020-02-21T11:38:18.799000",
  "votes": 15,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Along the things I have shared in this topic <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/130311\">https://www.kaggle.com/c/bengaliai-cv19/discussion/130311</a> I will share more of some methods, that for me did not result in a positive outcome.</p>\n\n<ol>\n<li>Switch from Relu to LeakyRelu for avoiding dying gradient problem (no improvement and training time increased)</li>\n<li>Upscale image to big resolution (over 200x200)</li>\n<li>Too many dropout layers (sometimes enough is enough)</li>\n<li>Small batch sizes (my best results were in batch_size&gt;48)</li>\n<li>AugMix did not improve my results (I spend some time tune the params and try to combine it with Cutmix/Mixup or Cutout)</li>\n<li>OHEM loss (no improvement)</li>\n<li>AdamW  optimizer (similar results as Adam, next on the list is the old SGD)</li>\n</ol>",
  "messages": [
    {
      "id": 752764,
      "postDate": "2020-02-21T11:38:18.800Z",
      "content": "<p>Along the things I have shared in this topic <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/130311\">https://www.kaggle.com/c/bengaliai-cv19/discussion/130311</a> I will share more of some methods, that for me did not result in a positive outcome.</p>\n\n<ol>\n<li>Switch from Relu to LeakyRelu for avoiding dying gradient problem (no improvement and training time increased)</li>\n<li>Upscale image to big resolution (over 200x200)</li>\n<li>Too many dropout layers (sometimes enough is enough)</li>\n<li>Small batch sizes (my best results were in batch_size&gt;48)</li>\n<li>AugMix did not improve my results (I spend some time tune the params and try to combine it with Cutmix/Mixup or Cutout)</li>\n<li>OHEM loss (no improvement)</li>\n<li>AdamW  optimizer (similar results as Adam, next on the list is the old SGD)</li>\n</ol>",
      "rawMarkdown": "Along the things I have shared in this topic https://www.kaggle.com/c/bengaliai-cv19/discussion/130311 I will share more of some methods, that for me did not result in a positive outcome.\n\n1. Switch from Relu to LeakyRelu for avoiding dying gradient problem (no improvement and training time increased)\n2. Upscale image to big resolution (over 200x200)\n3. Too many dropout layers (sometimes enough is enough)\n4. Small batch sizes (my best results were in batch_size&gt;48)\n5. AugMix did not improve my results (I spend some time tune the params and try to combine it with Cutmix/Mixup or Cutout)\n6. OHEM loss (no improvement)\n7. AdamW  optimizer (similar results as Adam, next on the list is the old SGD)",
      "votes": 15
    },
    {
      "id": 753030,
      "postDate": "2020-02-21T16:25:48.157Z",
      "content": "<p>Thanks again for sharing! Very helpful cutting down experiment times. I've had similar results on points 2,3 and 5-7.\n- I'd like to add for OHEM that I get worse results overall.\n- UPDATE: RAdam converges faster than Adam and performs better (less gap between CV,LB), next will be Lookahead SGD.\n- Seresnext101 performs better than Seresnext50 but takes longer to train.\n- Having .4 mixup, .4 cutmix and .2 aug only performs worse than just .5 mixup,cutmix.\n- UPDATE: Cutmix alone seems to perform better (less gap between CV,LB) will continue to experiment...</p>\n\n<p>My current score is using ensemble of models so I still haven't figured out the single model tricks, best single model score is .9683.</p>",
      "rawMarkdown": "Thanks again for sharing! Very helpful cutting down experiment times. I've had similar results on points 2,3 and 5-7.\n- I'd like to add for OHEM that I get worse results overall.\n- UPDATE: RAdam converges faster than Adam and performs better (less gap between CV,LB), next will be Lookahead SGD.\n- Seresnext101 performs better than Seresnext50 but takes longer to train.\n- Having .4 mixup, .4 cutmix and .2 aug only performs worse than just .5 mixup,cutmix.\n- UPDATE: Cutmix alone seems to perform better (less gap between CV,LB) will continue to experiment...\n\nMy current score is using ensemble of models so I still haven't figured out the single model tricks, best single model score is .9683.",
      "votes": 2,
      "replies": [
        {
          "id": 753054,
          "postDate": "2020-02-21T17:01:43.903Z",
          "content": "<p>Good feedback <a href=\"/greatgamedota\">@greatgamedota</a>  .\nPlease let me know how it goes with Lookahead SGD, I am really  curious</p>\n\n<p>Good luck !!!</p>",
          "rawMarkdown": "Good feedback @greatgamedota  .\nPlease let me know how it goes with Lookahead SGD, I am really  curious\n\nGood luck !!!"
        },
        {
          "id": 753065,
          "postDate": "2020-02-21T17:15:36.870Z",
          "content": "<p>I'm also testing OneCycle again for 300 epochs, it'll take a long time but I'll update with results</p>",
          "rawMarkdown": "I'm also testing OneCycle again for 300 epochs, it'll take a long time but I'll update with results"
        },
        {
          "id": 753071,
          "postDate": "2020-02-21T17:19:49.290Z",
          "content": "<p>300 epochs... what hardware are you using for training ?</p>",
          "rawMarkdown": "300 epochs... what hardware are you using for training ?"
        },
        {
          "id": 753073,
          "postDate": "2020-02-21T17:21:03.680Z",
          "content": "<p>RTX 2080 for long experiments, google colab for everything else.</p>",
          "rawMarkdown": "RTX 2080 for long experiments, google colab for everything else."
        },
        {
          "id": 753749,
          "postDate": "2020-02-22T16:11:22.440Z",
          "content": "<p>About .4 mixup, .4 cutmix and .2 aug, could I roughly know what kind of data augmentation are you using?\nDid you use cutout on that .2 data augmentation?\nBecause .4 mixup, .4 cutmix and .2 aug is what I am doing now, I cannot get better results...</p>",
          "rawMarkdown": "About .4 mixup, .4 cutmix and .2 aug, could I roughly know what kind of data augmentation are you using?\nDid you use cutout on that .2 data augmentation?\nBecause .4 mixup, .4 cutmix and .2 aug is what I am doing now, I cannot get better results..."
        },
        {
          "id": 753764,
          "postDate": "2020-02-22T16:25:31.003Z",
          "content": "<p>I'm using the same aug as this kernal: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a> (the transform part not the cropping)</p>",
          "rawMarkdown": "I'm using the same aug as this kernal: https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch (the transform part not the cropping)"
        },
        {
          "id": 753802,
          "postDate": "2020-02-22T16:54:28.960Z",
          "content": "<p>Got it. Thank you for your sharing!</p>",
          "rawMarkdown": "Got it. Thank you for your sharing!"
        },
        {
          "id": 753823,
          "postDate": "2020-02-22T17:19:08.847Z",
          "content": "<p>Lookahead(SGD) gives me a 0.3-0.4% accuracy boost, however I did not spend too much time testing adam because I just have more experience tweaking parameters with SGD and always seem to get better results with SGD</p>",
          "rawMarkdown": "Lookahead(SGD) gives me a 0.3-0.4% accuracy boost, however I did not spend too much time testing adam because I just have more experience tweaking parameters with SGD and always seem to get better results with SGD"
        },
        {
          "id": 753837,
          "postDate": "2020-02-22T17:52:05.217Z",
          "content": "<p>How long did you have to train for? SGD usually converged slower than Adam when I tested it.</p>",
          "rawMarkdown": "How long did you have to train for? SGD usually converged slower than Adam when I tested it."
        },
        {
          "id": 753844,
          "postDate": "2020-02-22T17:58:24.183Z",
          "content": "<p>I'm doing 100 epochs and using the ideas of: <a href=\"https://arxiv.org/pdf/1708.07120.pdf\">https://arxiv.org/pdf/1708.07120.pdf</a>\nSo high learning rate, large batchsize and OneCycleLR</p>\n\n<p>Lookahead also helps to converge in a fewer epochs</p>",
          "rawMarkdown": "I'm doing 100 epochs and using the ideas of: https://arxiv.org/pdf/1708.07120.pdf\nSo high learning rate, large batchsize and OneCycleLR\n\nLookahead also helps to converge in a fewer epochs",
          "votes": 1
        }
      ]
    },
    {
      "id": 752996,
      "postDate": "2020-02-21T15:54:32.717Z",
      "content": "<p>tried arcface loss against 1295 grapheme class. Seemed to help vowel and consonant classification but grapheme root still suffer. Have not totally given up on arcface loss yet.</p>",
      "rawMarkdown": "tried arcface loss against 1295 grapheme class. Seemed to help vowel and consonant classification but grapheme root still suffer. Have not totally given up on arcface loss yet.\n",
      "votes": 1,
      "replies": [
        {
          "id": 753001,
          "postDate": "2020-02-21T15:57:22.770Z",
          "content": "<p>Interesting, I have never use arcface loss. Please keep up to date with the progress</p>",
          "rawMarkdown": "Interesting, I have never use arcface loss. Please keep up to date with the progress"
        },
        {
          "id": 753076,
          "postDate": "2020-02-21T17:23:18.127Z",
          "content": "<p>What kind of architecture you are using for getting features ? </p>",
          "rawMarkdown": "What kind of architecture you are using for getting features ? "
        },
        {
          "id": 753195,
          "postDate": "2020-02-21T20:43:42.287Z",
          "content": "<p>one head for root, vowel, and consonant, and a seperate head from the feature layer trained with arcface</p>",
          "rawMarkdown": "one head for root, vowel, and consonant, and a seperate head from the feature layer trained with arcface"
        },
        {
          "id": 753196,
          "postDate": "2020-02-21T20:43:53.903Z",
          "content": "<p>well it's not working...</p>",
          "rawMarkdown": "well it's not working..."
        },
        {
          "id": 753579,
          "postDate": "2020-02-22T12:14:12.950Z",
          "content": "<p>unfortunately i still having figure out how to apply arcface loss properly ...</p>",
          "rawMarkdown": "unfortunately i still having figure out how to apply arcface loss properly ..."
        },
        {
          "id": 753584,
          "postDate": "2020-02-22T12:17:17.687Z",
          "content": "<p>it sounds like you @Nirjhar Roy\n may have some valuable tip(s) for us on this topic ? please do share as I am scratching my head...😉 </p>\n\n<p>I tried 2 things:\n1) use arcloss to traing feature extractor and then traing feature extractor with classification head\n2) use the model has two heads, one for classification (root, vowel, and consonant), the other for arcface loss.</p>\n\n<p>But none seems to work......</p>",
          "rawMarkdown": "it sounds like you @Nirjhar Roy\n may have some valuable tip(s) for us on this topic ? please do share as I am scratching my head...😉 \n\nI tried 2 things:\n1) use arcloss to traing feature extractor and then traing feature extractor with classification head\n2) use the model has two heads, one for classification (root, vowel, and consonant), the other for arcface loss.\n\nBut none seems to work......"
        }
      ]
    },
    {
      "id": 756243,
      "postDate": "2020-02-25T14:45:36.077Z",
      "content": "<p>Sharing everything that does not work is also science! Thanks for sharing Vlad Vaduva</p>",
      "rawMarkdown": "Sharing everything that does not work is also science! Thanks for sharing Vlad Vaduva"
    },
    {
      "id": 756222,
      "postDate": "2020-02-25T14:26:00.833Z",
      "content": "<p>OHEM FOCAL LOSS no improvement</p>",
      "rawMarkdown": "OHEM FOCAL LOSS no improvement"
    },
    {
      "id": 752993,
      "postDate": "2020-02-21T15:52:16.397Z",
      "content": "<p>Thanks for sharing those!</p>\n\n<p>As I'm still in the 2nd half of the leaderboard, I don't really know if this might be any useful, but I can definitely relate to your point about Dropouts. Removing some gave me a nice boost and allowed me to get passed the 0.96 mark!</p>",
      "rawMarkdown": "Thanks for sharing those!\n\nAs I'm still in the 2nd half of the leaderboard, I don't really know if this might be any useful, but I can definitely relate to your point about Dropouts. Removing some gave me a nice boost and allowed me to get passed the 0.96 mark!"
    },
    {
      "id": 752965,
      "postDate": "2020-02-21T15:27:19.793Z",
      "content": "<p>How you deal with grapheme_root: wrong labeling, imbalance ... etc. How's your validation score of it? I can't get more than .975 score on the only grapheme_root while others two achieve above .995+! I wonder how some people reach on it above .99!</p>",
      "rawMarkdown": "How you deal with grapheme_root: wrong labeling, imbalance ... etc. How's your validation score of it? I can't get more than .975 score on the only grapheme_root while others two achieve above .995+! I wonder how some people reach on it above .99!",
      "replies": [
        {
          "id": 752979,
          "postDate": "2020-02-21T15:37:45.283Z",
          "content": "<p>If you are training with different folds be sure that the folds are stratified. My validation score is very similar with the leaderboard score, you seem to have a big gap between validation score and leaderboard score, my advice is: \n- try to generalize better (use dropout is you are not)\n- use more augmentation (or smarter)\n- use a smaller arhitecture of the model</p>",
          "rawMarkdown": "If you are training with different folds be sure that the folds are stratified. My validation score is very similar with the leaderboard score, you seem to have a big gap between validation score and leaderboard score, my advice is: \n- try to generalize better (use dropout is you are not)\n- use more augmentation (or smarter)\n- use a smaller arhitecture of the model",
          "votes": 1
        },
        {
          "id": 752988,
          "postDate": "2020-02-21T15:46:30.633Z",
          "content": "<p>Thanks. I am not doing different folds, just plain splitting 8:2. The validation score on grapheme_root approx 0.975 while others two are approx 0.994. I have used drop-out but as you suggest I will look carefully now. Augmentation is a big issue here. I just want to ensure if we really need to give attention to grapheme_root specifically.  </p>",
          "rawMarkdown": "Thanks. I am not doing different folds, just plain splitting 8:2. The validation score on grapheme_root approx 0.975 while others two are approx 0.994. I have used drop-out but as you suggest I will look carefully now. Augmentation is a big issue here. I just want to ensure if we really need to give attention to grapheme_root specifically.  "
        },
        {
          "id": 752999,
          "postDate": "2020-02-21T15:56:15.037Z",
          "content": "<p>Try in the loss function to multiply by 2 the error from grapheme ! </p>",
          "rawMarkdown": "Try in the loss function to multiply by 2 the error from grapheme ! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 752767,
      "postDate": "2020-02-21T11:41:37.770Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 752776,
          "postDate": "2020-02-21T11:44:37.390Z",
          "content": "<p>Yes, I did but got better results with Cutmix/Mixup with cross entropy on a rexnext101 arhitecture</p>",
          "rawMarkdown": "Yes, I did but got better results with Cutmix/Mixup with cross entropy on a rexnext101 arhitecture"
        },
        {
          "id": 752851,
          "postDate": "2020-02-21T13:24:28.337Z",
          "content": "<p>That function returns a list of [loss1, loss2, loss3], so you need to add them first and call backward() then.</p>",
          "rawMarkdown": "That function returns a list of [loss1, loss2, loss3], so you need to add them first and call backward() then.",
          "votes": 1
        },
        {
          "id": 752894,
          "postDate": "2020-02-21T14:10:53.220Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 753030,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-02-21T16:25:48.157000",
      "content": "<p>Thanks again for sharing! Very helpful cutting down experiment times. I've had similar results on points 2,3 and 5-7.\n- I'd like to add for OHEM that I get worse results overall.\n- UPDATE: RAdam converges faster than Adam and performs better (less gap between CV,LB), next will be Lookahead SGD.\n- Seresnext101 performs better than Seresnext50 but takes longer to train.\n- Having .4 mixup, .4 cutmix and .2 aug only performs worse than just .5 mixup,cutmix.\n- UPDATE: Cutmix alone seems to perform better (less gap between CV,LB) will continue to experiment...</p>\n\n<p>My current score is using ensemble of models so I still haven't figured out the single model tricks, best single model score is .9683.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 753054,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-21T17:01:43.903000",
          "content": "<p>Good feedback <a href=\"/greatgamedota\">@greatgamedota</a>  .\nPlease let me know how it goes with Lookahead SGD, I am really  curious</p>\n\n<p>Good luck !!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753065,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-21T17:15:36.870000",
          "content": "<p>I'm also testing OneCycle again for 300 epochs, it'll take a long time but I'll update with results</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753071,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-21T17:19:49.290000",
          "content": "<p>300 epochs... what hardware are you using for training ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753073,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-21T17:21:03.680000",
          "content": "<p>RTX 2080 for long experiments, google colab for everything else.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753749,
          "author_name": "YS",
          "author_url": "",
          "post_date": "2020-02-22T16:11:22.440000",
          "content": "<p>About .4 mixup, .4 cutmix and .2 aug, could I roughly know what kind of data augmentation are you using?\nDid you use cutout on that .2 data augmentation?\nBecause .4 mixup, .4 cutmix and .2 aug is what I am doing now, I cannot get better results...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753764,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-22T16:25:31.003000",
          "content": "<p>I'm using the same aug as this kernal: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a> (the transform part not the cropping)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753802,
          "author_name": "YS",
          "author_url": "",
          "post_date": "2020-02-22T16:54:28.960000",
          "content": "<p>Got it. Thank you for your sharing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753823,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-22T17:19:08.847000",
          "content": "<p>Lookahead(SGD) gives me a 0.3-0.4% accuracy boost, however I did not spend too much time testing adam because I just have more experience tweaking parameters with SGD and always seem to get better results with SGD</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753837,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-22T17:52:05.217000",
          "content": "<p>How long did you have to train for? SGD usually converged slower than Adam when I tested it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753844,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-22T17:58:24.183000",
          "content": "<p>I'm doing 100 epochs and using the ideas of: <a href=\"https://arxiv.org/pdf/1708.07120.pdf\">https://arxiv.org/pdf/1708.07120.pdf</a>\nSo high learning rate, large batchsize and OneCycleLR</p>\n\n<p>Lookahead also helps to converge in a fewer epochs</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 752996,
      "author_name": "YL",
      "author_url": "",
      "post_date": "2020-02-21T15:54:32.717000",
      "content": "<p>tried arcface loss against 1295 grapheme class. Seemed to help vowel and consonant classification but grapheme root still suffer. Have not totally given up on arcface loss yet.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 753001,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-21T15:57:22.770000",
          "content": "<p>Interesting, I have never use arcface loss. Please keep up to date with the progress</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753076,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-02-21T17:23:18.127000",
          "content": "<p>What kind of architecture you are using for getting features ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753195,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-21T20:43:42.287000",
          "content": "<p>one head for root, vowel, and consonant, and a seperate head from the feature layer trained with arcface</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753196,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-21T20:43:53.903000",
          "content": "<p>well it's not working...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753579,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-22T12:14:12.950000",
          "content": "<p>unfortunately i still having figure out how to apply arcface loss properly ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753584,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-22T12:17:17.687000",
          "content": "<p>it sounds like you @Nirjhar Roy\n may have some valuable tip(s) for us on this topic ? please do share as I am scratching my head...😉 </p>\n\n<p>I tried 2 things:\n1) use arcloss to traing feature extractor and then traing feature extractor with classification head\n2) use the model has two heads, one for classification (root, vowel, and consonant), the other for arcface loss.</p>\n\n<p>But none seems to work......</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 756243,
      "author_name": "G-Dant",
      "author_url": "",
      "post_date": "2020-02-25T14:45:36.077000",
      "content": "<p>Sharing everything that does not work is also science! Thanks for sharing Vlad Vaduva</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 756222,
      "author_name": "cswwp",
      "author_url": "",
      "post_date": "2020-02-25T14:26:00.833000",
      "content": "<p>OHEM FOCAL LOSS no improvement</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 752993,
      "author_name": "Maxime Lenormand",
      "author_url": "",
      "post_date": "2020-02-21T15:52:16.397000",
      "content": "<p>Thanks for sharing those!</p>\n\n<p>As I'm still in the 2nd half of the leaderboard, I don't really know if this might be any useful, but I can definitely relate to your point about Dropouts. Removing some gave me a nice boost and allowed me to get passed the 0.96 mark!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 752965,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-02-21T15:27:19.793000",
      "content": "<p>How you deal with grapheme_root: wrong labeling, imbalance ... etc. How's your validation score of it? I can't get more than .975 score on the only grapheme_root while others two achieve above .995+! I wonder how some people reach on it above .99!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 752979,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-21T15:37:45.283000",
          "content": "<p>If you are training with different folds be sure that the folds are stratified. My validation score is very similar with the leaderboard score, you seem to have a big gap between validation score and leaderboard score, my advice is: \n- try to generalize better (use dropout is you are not)\n- use more augmentation (or smarter)\n- use a smaller arhitecture of the model</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 752988,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-21T15:46:30.633000",
          "content": "<p>Thanks. I am not doing different folds, just plain splitting 8:2. The validation score on grapheme_root approx 0.975 while others two are approx 0.994. I have used drop-out but as you suggest I will look carefully now. Augmentation is a big issue here. I just want to ensure if we really need to give attention to grapheme_root specifically.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752999,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-21T15:56:15.037000",
          "content": "<p>Try in the loss function to multiply by 2 the error from grapheme ! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 752767,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-21T11:41:37.770000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 752776,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-21T11:44:37.390000",
          "content": "<p>Yes, I did but got better results with Cutmix/Mixup with cross entropy on a rexnext101 arhitecture</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752851,
          "author_name": "hawkey",
          "author_url": "",
          "post_date": "2020-02-21T13:24:28.337000",
          "content": "<p>That function returns a list of [loss1, loss2, loss3], so you need to add them first and call backward() then.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 752894,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-21T14:10:53.220000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "752764": "Along the things I have shared in this topic https://www.kaggle.com/c/bengaliai-cv19/discussion/130311 I will share more of some methods, that for me did not result in a positive outcome.\n\n1. Switch from Relu to LeakyRelu for avoiding dying gradient problem (no improvement and training time increased)\n2. Upscale image to big resolution (over 200x200)\n3. Too many dropout layers (sometimes enough is enough)\n4. Small batch sizes (my best results were in batch_size&gt;48)\n5. AugMix did not improve my results (I spend some time tune the params and try to combine it with Cutmix/Mixup or Cutout)\n6. OHEM loss (no improvement)\n7. AdamW  optimizer (similar results as Adam, next on the list is the old SGD)",
    "753030": "Thanks again for sharing! Very helpful cutting down experiment times. I've had similar results on points 2,3 and 5-7.\n- I'd like to add for OHEM that I get worse results overall.\n- UPDATE: RAdam converges faster than Adam and performs better (less gap between CV,LB), next will be Lookahead SGD.\n- Seresnext101 performs better than Seresnext50 but takes longer to train.\n- Having .4 mixup, .4 cutmix and .2 aug only performs worse than just .5 mixup,cutmix.\n- UPDATE: Cutmix alone seems to perform better (less gap between CV,LB) will continue to experiment...\n\nMy current score is using ensemble of models so I still haven't figured out the single model tricks, best single model score is .9683.",
    "752996": "tried arcface loss against 1295 grapheme class. Seemed to help vowel and consonant classification but grapheme root still suffer. Have not totally given up on arcface loss yet.\n",
    "756243": "Sharing everything that does not work is also science! Thanks for sharing Vlad Vaduva",
    "756222": "OHEM FOCAL LOSS no improvement",
    "752993": "Thanks for sharing those!\n\nAs I'm still in the 2nd half of the leaderboard, I don't really know if this might be any useful, but I can definitely relate to your point about Dropouts. Removing some gave me a nice boost and allowed me to get passed the 0.96 mark!",
    "752965": "How you deal with grapheme_root: wrong labeling, imbalance ... etc. How's your validation score of it? I can't get more than .975 score on the only grapheme_root while others two achieve above .995+! I wonder how some people reach on it above .99!",
    "752767": ""
  }
}