{
  "id": 205491,
  "title": "7 things that did not worked so far",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/205491",
  "author_name": "",
  "post_date": "2020-12-20T11:42:36.411189500Z",
  "votes": 80,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Among other things that worked that I will partly share soon, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower</p>\n<p>Even if everybody wants to know what works, I think it's always helpful for the community to know what failed </p>\n<ol>\n<li><p>Optimizers: AdamW and RangerLars did not provide better results than good old ADAM.</p></li>\n<li><p>Very large architectures (Effnet B8, Effnet B7, Effnet B6). For these types of architectures, my limited hardware capabilities are forcing me to enter with a very low batch size (even using apex) which leads to noisy gradient backpropagation on a noisy set, which is a very bad combination. If you are having an advanced GPU which supports more than 11GB you may have better results.</p></li>\n<li><p>So far the old data from the similar competition did not bring any improvements in my models. Here you can get very easily tricked on the OOF score if you are accumulating the data and then stratify them into folds. These kind of auxiliary data even if is coming from an almost identical competition is mandatory to not enter the validation set. Otherwise, it will not be representative and correlated to the leaderboard.<br>\nI know that this worked for some of you, so it is not completely out of my todo list</p></li>\n<li><p>Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.</p></li>\n<li><p>Too small arhitectures (Resnet18, Resnet 50, even Effnet B0) are underfiting comparative with B2's, B3's and B4's</p></li>\n<li><p>Fmix did not give better results than Cutmix (On Bengali competition where this technique gave good results, it helped create new symbols by spiting core components(roots+vocals+consonant) that were not provided by the training data but found on the testing set.</p></li>\n<li><p>OHEM loss, gamblers loss and also focal loss did not performed better than CE for me. Did not tried yet cosine focal loss</p></li>\n</ol>",
  "messages": [
    {
      "id": "1119828",
      "postDate": "12/20/2020 11:42:36",
      "content": "<p>Among other things that worked that I will partly share soon, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower</p>\n<p>Even if everybody wants to know what works, I think it's always helpful for the community to know what failed </p>\n<ol>\n<li><p>Optimizers: AdamW and RangerLars did not provide better results than good old ADAM.</p></li>\n<li><p>Very large architectures (Effnet B8, Effnet B7, Effnet B6). For these types of architectures, my limited hardware capabilities are forcing me to enter with a very low batch size (even using apex) which leads to noisy gradient backpropagation on a noisy set, which is a very bad combination. If you are having an advanced GPU which supports more than 11GB you may have better results.</p></li>\n<li><p>So far the old data from the similar competition did not bring any improvements in my models. Here you can get very easily tricked on the OOF score if you are accumulating the data and then stratify them into folds. These kind of auxiliary data even if is coming from an almost identical competition is mandatory to not enter the validation set. Otherwise, it will not be representative and correlated to the leaderboard.<br>\nI know that this worked for some of you, so it is not completely out of my todo list</p></li>\n<li><p>Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.</p></li>\n<li><p>Too small arhitectures (Resnet18, Resnet 50, even Effnet B0) are underfiting comparative with B2's, B3's and B4's</p></li>\n<li><p>Fmix did not give better results than Cutmix (On Bengali competition where this technique gave good results, it helped create new symbols by spiting core components(roots+vocals+consonant) that were not provided by the training data but found on the testing set.</p></li>\n<li><p>OHEM loss, gamblers loss and also focal loss did not performed better than CE for me. Did not tried yet cosine focal loss</p></li>\n</ol>",
      "rawMarkdown": "Among other things that worked that I will partly share soon, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower\n\nEven if everybody wants to know what works, I think it's always helpful for the community to know what failed \n\n1. Optimizers: AdamW and RangerLars did not provide better results than good old ADAM.\n\n2. Very large architectures (Effnet B8, Effnet B7, Effnet B6). For these types of architectures, my limited hardware capabilities are forcing me to enter with a very low batch size (even using apex) which leads to noisy gradient backpropagation on a noisy set, which is a very bad combination. If you are having an advanced GPU which supports more than 11GB you may have better results.\n\n3. So far the old data from the similar competition did not bring any improvements in my models. Here you can get very easily tricked on the OOF score if you are accumulating the data and then stratify them into folds. These kind of auxiliary data even if is coming from an almost identical competition is mandatory to not enter the validation set. Otherwise, it will not be representative and correlated to the leaderboard.\nI know that this worked for some of you, so it is not completely out of my todo list\n\n4. Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.\n\n5. Too small arhitectures (Resnet18, Resnet 50, even Effnet B0) are underfiting comparative with B2's, B3's and B4's\n\n6. Fmix did not give better results than Cutmix (On Bengali competition where this technique gave good results, it helped create new symbols by spiting core components(roots+vocals+consonant) that were not provided by the training data but found on the testing set.\n\n7. OHEM loss, gamblers loss and also focal loss did not performed better than CE for me. Did not tried yet cosine focal loss",
      "votes": null
    },
    {
      "id": "1119923",
      "postDate": "12/20/2020 13:01:25",
      "content": "<p>Thanks for sharing. <br>\n1) I agree Adam is working better than AdamW<br>\n2) Yes smaller architectures are underfitting atleast for me.<br>\n3) I haven't tried OHEM loss. Can you please share your implementation ?</p>",
      "rawMarkdown": "Thanks for sharing. \n1) I agree Adam is working better than AdamW\n2) Yes smaller architectures are underfitting atleast for me.\n3) I haven't tried OHEM loss. Can you please share your implementation ?",
      "votes": null
    },
    {
      "id": "1119949",
      "postDate": "12/20/2020 13:31:52",
      "content": "<p><strong>What does OHEM loss do:</strong><br>\nGiven a batch sized K, it performs regular forward propagation and computes per instance losses. Then, it finds M&lt;K hard examples in the batch with high loss values and it only back-propagates the loss computed over the  selected instances.</p>\n<p><strong>This is the implementation:</strong><br>\ndef ohem_loss( rate, cls_pred, cls_target ):<br>\n   batch_size = cls_pred.size(0) <br>\n   ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)<br>\n   sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)<br>\n   keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )<br>\n   if keep_num &lt; sorted_ohem_loss.size()[0]:<br>\n       keep_idx_cuda = idx[:keep_num]<br>\n       ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]<br>\n   cls_loss = ohem_cls_loss.sum() / keep_num<br>\n   return cls_loss</p>",
      "rawMarkdown": "**What does OHEM loss do:**\nGiven a batch sized K, it performs regular forward propagation and computes per instance losses. Then, it finds M<K hard examples in the batch with high loss values and it only back-propagates the loss computed over the  selected instances.\n  \n\n**This is the implementation:**\ndef ohem_loss( rate, cls_pred, cls_target ):\n    batch_size = cls_pred.size(0) \n    ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n\n    sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\n    keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\n    if keep_num < sorted_ohem_loss.size()[0]:\n        keep_idx_cuda = idx[:keep_num]\n        ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\n    cls_loss = ohem_cls_loss.sum() / keep_num\n    return cls_loss",
      "votes": null
    },
    {
      "id": "1120007",
      "postDate": "12/20/2020 14:45:16",
      "content": "<p>What's the \"similar competition\" that you spoke about? Mind tagging a link here?</p>",
      "rawMarkdown": "What's the \"similar competition\" that you spoke about? Mind tagging a link here?",
      "votes": null
    },
    {
      "id": "1120013",
      "postDate": "12/20/2020 14:49:19",
      "content": "<p><a href=\"https://www.kaggle.com/tyqiangz\" target=\"_blank\">@tyqiangz</a> <br>\nThis is the one:<br>\n<a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease</a></p>",
      "rawMarkdown": "tyqiangz \nThis is the one:\nhttps://www.kaggle.com/c/cassava-disease",
      "votes": null
    },
    {
      "id": "1120201",
      "postDate": "12/20/2020 16:57:13",
      "content": "<p>Also for me, Cutmix works better than FMix.</p>",
      "rawMarkdown": "Also for me, Cutmix works better than FMix.",
      "votes": null
    },
    {
      "id": "1120340",
      "postDate": "12/20/2020 18:48:29",
      "content": "<p>Hey i have implemented <strong>gamblers loss</strong>  </p>\n<p>I trained effnet using CE loss and got 80% acc and then i use gamblers loss and for 10 epochs its accuracy was 20-30%.</p>\n<p>As you said gamblers loss did not worked out, By this you mean same or your accuracy is not this bad ? </p>\n<p>Can you please have look to my loss function <br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424</a></p>\n<p>Just to verify implementation was correct. </p>",
      "rawMarkdown": "Hey i have implemented **gamblers loss**  \n\nI trained effnet using CE loss and got 80% acc and then i use gamblers loss and for 10 epochs its accuracy was 20-30%.\n\nAs you said gamblers loss did not worked out, By this you mean same or your accuracy is not this bad ? \n\nCan you please have look to my loss function \nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424\n\nJust to verify implementation was correct.",
      "votes": null
    },
    {
      "id": "1120353",
      "postDate": "12/20/2020 19:00:03",
      "content": "<p>Something ive noticed is that i get very different CV scores on the same folds depending if i use original aspect ratio, or using resize or using centercrop. So ensembling these models could help!</p>",
      "rawMarkdown": "Something ive noticed is that i get very different CV scores on the same folds depending if i use original aspect ratio, or using resize or using centercrop. So ensembling these models could help!",
      "votes": null
    },
    {
      "id": "1120618",
      "postDate": "12/21/2020 01:37:22",
      "content": "<p>Compared to ADAM, does SGD with momentum perform better, except for training speed?</p>",
      "rawMarkdown": "Compared to ADAM, does SGD with momentum perform better, except for training speed?",
      "votes": null
    },
    {
      "id": "1120941",
      "postDate": "12/21/2020 08:15:29",
      "content": "<p>When you tried to train deeper nets (B7, B6) did you use Adam or SGD? <br>\nAdam is more memory consuming, so SGD could save you a bit of VRAM.</p>",
      "rawMarkdown": "When you tried to train deeper nets (B7, B6) did you use Adam or SGD? \nAdam is more memory consuming, so SGD could save you a bit of VRAM.",
      "votes": null
    },
    {
      "id": "1121016",
      "postDate": "12/21/2020 09:37:57",
      "content": "<p><a href=\"https://www.kaggle.com/altprof\" target=\"_blank\">@altprof</a> , <a href=\"https://www.kaggle.com/sjtuyxc\" target=\"_blank\">@sjtuyxc</a> Did not tried SGD yet. It is on my todo list.</p>",
      "rawMarkdown": "altprof , @sjtuyxc Did not tried SGD yet. It is on my todo list.",
      "votes": null
    },
    {
      "id": "1121220",
      "postDate": "12/21/2020 13:05:49",
      "content": "<p>Thanks~ I got it!</p>",
      "rawMarkdown": "Thanks~ I got it!",
      "votes": null
    },
    {
      "id": "1121307",
      "postDate": "12/21/2020 14:34:45",
      "content": "<p>Good information ! Great </p>",
      "rawMarkdown": "Good information ! Great",
      "votes": null
    },
    {
      "id": "1122480",
      "postDate": "12/22/2020 13:24:11",
      "content": "<p>I had a similar experience with effnets and image sizes. </p>\n<p>I was using older data and saw some small accuracy increases. I'm curious for those who might have seen increases using these as validation of what their methods may have been. </p>",
      "rawMarkdown": "I had a similar experience with effnets and image sizes. \n\nI was using older data and saw some small accuracy increases. I'm curious for those who might have seen increases using these as validation of what their methods may have been.",
      "votes": null
    },
    {
      "id": "1122532",
      "postDate": "12/22/2020 13:57:30",
      "content": "<p>I have similar results for my experiments:</p>\n<ul>\n<li>Other resolutions than <code>512x512</code> give me worse performance</li>\n<li><code>EfficientNet B3-B6</code> are the best architectures i've found so far</li>\n<li>Adam optimizer (I was planning to try ranger or similar optimizers, maybe a better idea could be to use <a href=\"https://arxiv.org/pdf/2010.01412.pdf\" target=\"_blank\">SAM</a>)</li>\n<li>External data (from 2019 competition) does not improve performance</li>\n<li>Label smoothing improves the LB score of <code>~0.01</code></li>\n</ul>\n<p>I use as data augmentation techniques CutMix and MixUp + random flips, crops and rotations, but maybe it's worth a try to implement snapmix and see if it gives any improvements.</p>\n<p>I'm planning to use bi-tempered logistic loss, but up to now my results are the same as those obtained with standard cross-entropy (single fold model LB score <code>0.890</code>)</p>",
      "rawMarkdown": "I have similar results for my experiments:\n\n- Other resolutions than `512x512` give me worse performance\n- `EfficientNet B3-B6` are the best architectures i've found so far\n- Adam optimizer (I was planning to try ranger or similar optimizers, maybe a better idea could be to use [SAM](https://arxiv.org/pdf/2010.01412.pdf))\n- External data (from 2019 competition) does not improve performance\n- Label smoothing improves the LB score of `~0.01`\n\n\nI use as data augmentation techniques CutMix and MixUp + random flips, crops and rotations, but maybe it's worth a try to implement snapmix and see if it gives any improvements.\n\nI'm planning to use bi-tempered logistic loss, but up to now my results are the same as those obtained with standard cross-entropy (single fold model LB score `0.890`)",
      "votes": null
    },
    {
      "id": "1123536",
      "postDate": "12/23/2020 10:05:28",
      "content": "<p><a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> you using CutMix and MixUp in the same time ?</p>",
      "rawMarkdown": "lazcoder you using CutMix and MixUp in the same time ?",
      "votes": null
    },
    {
      "id": "1123540",
      "postDate": "12/23/2020 10:07:45",
      "content": "<p><a href=\"https://www.kaggle.com/rajan\" target=\"_blank\">@rajan</a> The accuracy was lower pretty significant with gambler loss. Personally, I would go with other losses for this competition</p>",
      "rawMarkdown": "rajan The accuracy was lower pretty significant with gambler loss. Personally, I would go with other losses for this competition",
      "votes": null
    },
    {
      "id": "1123543",
      "postDate": "12/23/2020 10:09:07",
      "content": "<p>Training resize or validation resize?</p>",
      "rawMarkdown": "Training resize or validation resize?",
      "votes": null
    },
    {
      "id": "1123546",
      "postDate": "12/23/2020 10:13:43",
      "content": "<p>No, I use either CutMix or MixUp randomly on each batch.</p>\n<p>I use something like</p>\n<pre><code>if p &gt; p_mix:\n    do_mixup(batch)\nelse:\n    do_cutmix(batch)\n</code></pre>",
      "rawMarkdown": "No, I use either CutMix or MixUp randomly on each batch.\n\nI use something like\n\n```\nif p > p_mix:\n    do_mixup(batch)\nelse:\n    do_cutmix(batch)\n```",
      "votes": null
    },
    {
      "id": "1123550",
      "postDate": "12/23/2020 10:16:24",
      "content": "<p>Yes, I agree, It's the best way to do it ! </p>",
      "rawMarkdown": "Yes, I agree, It's the best way to do it !",
      "votes": null
    },
    {
      "id": "1123556",
      "postDate": "12/23/2020 10:21:28",
      "content": "<p>Actually I've read somewhere that mixup is harmful for this competition (even if in my experiments it seems not). <br>\nMaybe an idea that is worth to try is to use snapmix instead of mixup </p>",
      "rawMarkdown": "Actually I've read somewhere that mixup is harmful for this competition (even if in my experiments it seems not). \nMaybe an idea that is worth to try is to use snapmix instead of mixup",
      "votes": null
    },
    {
      "id": "1123559",
      "postDate": "12/23/2020 10:25:31",
      "content": "<p>If you use them both in the same time, use them separately like in your code piece<br>\nThe opportunity of using one or both of them in this competition is another discussion.<br>\nFor example, I am using just one of them</p>",
      "rawMarkdown": "If you use them both in the same time, use them separately like in your code piece\nThe opportunity of using one or both of them in this competition is another discussion.\nFor example, I am using just one of them",
      "votes": null
    },
    {
      "id": "1125896",
      "postDate": "12/25/2020 06:50:54",
      "content": "<p>Normal SGD worked better than SGD with momentum of 0.9 for me (learning rate=0.001 on both cases). </p>",
      "rawMarkdown": "Normal SGD worked better than SGD with momentum of 0.9 for me (learning rate=0.001 on both cases).",
      "votes": null
    },
    {
      "id": "1137023",
      "postDate": "01/03/2021 15:53:53",
      "content": "<blockquote>\n  <ol>\n  <li>Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.</li>\n  </ol>\n</blockquote>\n<p>(1) What do you mean by 'Image size smaller than 512' ? The image size in this competition is all 800x600?<br>\n(2) Maybe the cropped  lower limit 0.08 is too small, we need to set it to 0.6 or 0.8 ? Thanks!</p>",
      "rawMarkdown": "> 4. Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.\n\n(1) What do you mean by 'Image size smaller than 512' ? The image size in this competition is all 800x600?\n(2) Maybe the cropped  lower limit 0.08 is too small, we need to set it to 0.6 or 0.8 ? Thanks!",
      "votes": null
    },
    {
      "id": "1137161",
      "postDate": "01/03/2021 18:16:24",
      "content": "<p>i have used </p>\n<ul>\n<li>LabelSmooth</li>\n<li>cutmix</li>\n<li>res50 </li>\n<li>eff B3-b5</li>\n<li>SCELoss</li>\n<li>TTA</li>\n<li>Ranger<br>\nWhat else can i do ? I have seen many people with high accuracy suggesting some one of them. What it is that i am missing </li>\n</ul>\n<p>Getting 87.7 max accuracy </p>\n<p>my notebook <br>\n<a href=\"https://www.kaggle.com/rajanlagah/noob-programmer-s-code\" target=\"_blank\">https://www.kaggle.com/rajanlagah/noob-programmer-s-code</a></p>\n<p>Will appreciate your feedback </p>",
      "rawMarkdown": "i have used \n- LabelSmooth\n- cutmix\n- res50 \n- eff B3-b5\n- SCELoss\n- TTA\n- Ranger\nWhat else can i do ? I have seen many people with high accuracy suggesting some one of them. What it is that i am missing \n \nGetting 87.7 max accuracy \n\nmy notebook \nhttps://www.kaggle.com/rajanlagah/noob-programmer-s-code\n\nWill appreciate your feedback",
      "votes": null
    },
    {
      "id": "1137677",
      "postDate": "01/04/2021 06:33:23",
      "content": "<p>how to calculate training acc when using mixup? ( it becomes two labels, such as class_A=0.4 and class_B = 0.6)</p>",
      "rawMarkdown": "how to calculate training acc when using mixup? ( it becomes two labels, such as class_A=0.4 and class_B = 0.6)",
      "votes": null
    },
    {
      "id": "1144909",
      "postDate": "01/08/2021 18:36:57",
      "content": "<p>Why didn't you try it on Google Colab,  it would have given you better compute.</p>",
      "rawMarkdown": "Why didn't you try it on Google Colab,  it would have given you better compute.",
      "votes": null
    },
    {
      "id": "1226651",
      "postDate": "03/04/2021 18:03:17",
      "content": "<p>what about </p>\n<ol>\n<li>custom mean and std deviation ?</li>\n<li>SuperConvergence or OneCycleLR?</li>\n</ol>",
      "rawMarkdown": "what about \n1. custom mean and std deviation ?\n2. SuperConvergence or OneCycleLR?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1119923,
      "author_name": "atharvap329",
      "author_url": "",
      "post_date": "12/20/2020 13:01:25",
      "content": "<p>Thanks for sharing. <br>\n1) I agree Adam is working better than AdamW<br>\n2) Yes smaller architectures are underfitting atleast for me.<br>\n3) I haven't tried OHEM loss. Can you please share your implementation ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119949,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/20/2020 13:31:52",
          "content": "<p><strong>What does OHEM loss do:</strong><br>\nGiven a batch sized K, it performs regular forward propagation and computes per instance losses. Then, it finds M&lt;K hard examples in the batch with high loss values and it only back-propagates the loss computed over the  selected instances.</p>\n<p><strong>This is the implementation:</strong><br>\ndef ohem_loss( rate, cls_pred, cls_target ):<br>\n   batch_size = cls_pred.size(0) <br>\n   ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)<br>\n   sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)<br>\n   keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )<br>\n   if keep_num &lt; sorted_ohem_loss.size()[0]:<br>\n       keep_idx_cuda = idx[:keep_num]<br>\n       ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]<br>\n   cls_loss = ohem_cls_loss.sum() / keep_num<br>\n   return cls_loss</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120007,
      "author_name": "tyqiangz",
      "author_url": "",
      "post_date": "12/20/2020 14:45:16",
      "content": "<p>What's the \"similar competition\" that you spoke about? Mind tagging a link here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1120013,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/20/2020 14:49:19",
          "content": "<p><a href=\"https://www.kaggle.com/tyqiangz\" target=\"_blank\">@tyqiangz</a> <br>\nThis is the one:<br>\n<a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120201,
      "author_name": "kaushal2896",
      "author_url": "",
      "post_date": "12/20/2020 16:57:13",
      "content": "<p>Also for me, Cutmix works better than FMix.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1120340,
      "author_name": "rajanlagah",
      "author_url": "",
      "post_date": "12/20/2020 18:48:29",
      "content": "<p>Hey i have implemented <strong>gamblers loss</strong>  </p>\n<p>I trained effnet using CE loss and got 80% acc and then i use gamblers loss and for 10 epochs its accuracy was 20-30%.</p>\n<p>As you said gamblers loss did not worked out, By this you mean same or your accuracy is not this bad ? </p>\n<p>Can you please have look to my loss function <br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424</a></p>\n<p>Just to verify implementation was correct. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1123540,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/23/2020 10:07:45",
          "content": "<p><a href=\"https://www.kaggle.com/rajan\" target=\"_blank\">@rajan</a> The accuracy was lower pretty significant with gambler loss. Personally, I would go with other losses for this competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137161,
          "author_name": "rajanlagah",
          "author_url": "",
          "post_date": "01/03/2021 18:16:24",
          "content": "<p>i have used </p>\n<ul>\n<li>LabelSmooth</li>\n<li>cutmix</li>\n<li>res50 </li>\n<li>eff B3-b5</li>\n<li>SCELoss</li>\n<li>TTA</li>\n<li>Ranger<br>\nWhat else can i do ? I have seen many people with high accuracy suggesting some one of them. What it is that i am missing </li>\n</ul>\n<p>Getting 87.7 max accuracy </p>\n<p>my notebook <br>\n<a href=\"https://www.kaggle.com/rajanlagah/noob-programmer-s-code\" target=\"_blank\">https://www.kaggle.com/rajanlagah/noob-programmer-s-code</a></p>\n<p>Will appreciate your feedback </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120353,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "12/20/2020 19:00:03",
      "content": "<p>Something ive noticed is that i get very different CV scores on the same folds depending if i use original aspect ratio, or using resize or using centercrop. So ensembling these models could help!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123543,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/23/2020 10:09:07",
          "content": "<p>Training resize or validation resize?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120618,
      "author_name": "sjtuyxc",
      "author_url": "",
      "post_date": "12/21/2020 01:37:22",
      "content": "<p>Compared to ADAM, does SGD with momentum perform better, except for training speed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1125896,
          "author_name": "suryaaseran",
          "author_url": "",
          "post_date": "12/25/2020 06:50:54",
          "content": "<p>Normal SGD worked better than SGD with momentum of 0.9 for me (learning rate=0.001 on both cases). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120941,
      "author_name": "altprof",
      "author_url": "",
      "post_date": "12/21/2020 08:15:29",
      "content": "<p>When you tried to train deeper nets (B7, B6) did you use Adam or SGD? <br>\nAdam is more memory consuming, so SGD could save you a bit of VRAM.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1121016,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "12/21/2020 09:37:57",
      "content": "<p><a href=\"https://www.kaggle.com/altprof\" target=\"_blank\">@altprof</a> , <a href=\"https://www.kaggle.com/sjtuyxc\" target=\"_blank\">@sjtuyxc</a> Did not tried SGD yet. It is on my todo list.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1121220,
          "author_name": "sjtuyxc",
          "author_url": "",
          "post_date": "12/21/2020 13:05:49",
          "content": "<p>Thanks~ I got it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121307,
      "author_name": "",
      "author_url": "",
      "post_date": "12/21/2020 14:34:45",
      "content": "<p>Good information ! Great </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122480,
      "author_name": "crained",
      "author_url": "",
      "post_date": "12/22/2020 13:24:11",
      "content": "<p>I had a similar experience with effnets and image sizes. </p>\n<p>I was using older data and saw some small accuracy increases. I'm curious for those who might have seen increases using these as validation of what their methods may have been. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122532,
      "author_name": "lazcoder",
      "author_url": "",
      "post_date": "12/22/2020 13:57:30",
      "content": "<p>I have similar results for my experiments:</p>\n<ul>\n<li>Other resolutions than <code>512x512</code> give me worse performance</li>\n<li><code>EfficientNet B3-B6</code> are the best architectures i've found so far</li>\n<li>Adam optimizer (I was planning to try ranger or similar optimizers, maybe a better idea could be to use <a href=\"https://arxiv.org/pdf/2010.01412.pdf\" target=\"_blank\">SAM</a>)</li>\n<li>External data (from 2019 competition) does not improve performance</li>\n<li>Label smoothing improves the LB score of <code>~0.01</code></li>\n</ul>\n<p>I use as data augmentation techniques CutMix and MixUp + random flips, crops and rotations, but maybe it's worth a try to implement snapmix and see if it gives any improvements.</p>\n<p>I'm planning to use bi-tempered logistic loss, but up to now my results are the same as those obtained with standard cross-entropy (single fold model LB score <code>0.890</code>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123536,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/23/2020 10:05:28",
          "content": "<p><a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> you using CutMix and MixUp in the same time ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123546,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/23/2020 10:13:43",
          "content": "<p>No, I use either CutMix or MixUp randomly on each batch.</p>\n<p>I use something like</p>\n<pre><code>if p &gt; p_mix:\n    do_mixup(batch)\nelse:\n    do_cutmix(batch)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123550,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/23/2020 10:16:24",
          "content": "<p>Yes, I agree, It's the best way to do it ! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123556,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "12/23/2020 10:21:28",
          "content": "<p>Actually I've read somewhere that mixup is harmful for this competition (even if in my experiments it seems not). <br>\nMaybe an idea that is worth to try is to use snapmix instead of mixup </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123559,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "12/23/2020 10:25:31",
          "content": "<p>If you use them both in the same time, use them separately like in your code piece<br>\nThe opportunity of using one or both of them in this competition is another discussion.<br>\nFor example, I am using just one of them</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137677,
          "author_name": "clwwlc",
          "author_url": "",
          "post_date": "01/04/2021 06:33:23",
          "content": "<p>how to calculate training acc when using mixup? ( it becomes two labels, such as class_A=0.4 and class_B = 0.6)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1137023,
      "author_name": "clwwlc",
      "author_url": "",
      "post_date": "01/03/2021 15:53:53",
      "content": "<blockquote>\n  <ol>\n  <li>Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.</li>\n  </ol>\n</blockquote>\n<p>(1) What do you mean by 'Image size smaller than 512' ? The image size in this competition is all 800x600?<br>\n(2) Maybe the cropped  lower limit 0.08 is too small, we need to set it to 0.6 or 0.8 ? Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144909,
      "author_name": "deeprakesh",
      "author_url": "",
      "post_date": "01/08/2021 18:36:57",
      "content": "<p>Why didn't you try it on Google Colab,  it would have given you better compute.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1226651,
      "author_name": "abhiagwl",
      "author_url": "",
      "post_date": "03/04/2021 18:03:17",
      "content": "<p>what about </p>\n<ol>\n<li>custom mean and std deviation ?</li>\n<li>SuperConvergence or OneCycleLR?</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1119828": "Among other things that worked that I will partly share soon, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower\n\nEven if everybody wants to know what works, I think it's always helpful for the community to know what failed \n\n1. Optimizers: AdamW and RangerLars did not provide better results than good old ADAM.\n\n2. Very large architectures (Effnet B8, Effnet B7, Effnet B6). For these types of architectures, my limited hardware capabilities are forcing me to enter with a very low batch size (even using apex) which leads to noisy gradient backpropagation on a noisy set, which is a very bad combination. If you are having an advanced GPU which supports more than 11GB you may have better results.\n\n3. So far the old data from the similar competition did not bring any improvements in my models. Here you can get very easily tricked on the OOF score if you are accumulating the data and then stratify them into folds. These kind of auxiliary data even if is coming from an almost identical competition is mandatory to not enter the validation set. Otherwise, it will not be representative and correlated to the leaderboard.\nI know that this worked for some of you, so it is not completely out of my todo list\n\n4. Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.\n\n5. Too small arhitectures (Resnet18, Resnet 50, even Effnet B0) are underfiting comparative with B2's, B3's and B4's\n\n6. Fmix did not give better results than Cutmix (On Bengali competition where this technique gave good results, it helped create new symbols by spiting core components(roots+vocals+consonant) that were not provided by the training data but found on the testing set.\n\n7. OHEM loss, gamblers loss and also focal loss did not performed better than CE for me. Did not tried yet cosine focal loss",
    "1119923": "Thanks for sharing. \n1) I agree Adam is working better than AdamW\n2) Yes smaller architectures are underfitting atleast for me.\n3) I haven't tried OHEM loss. Can you please share your implementation ?",
    "1119949": "**What does OHEM loss do:**\nGiven a batch sized K, it performs regular forward propagation and computes per instance losses. Then, it finds M<K hard examples in the batch with high loss values and it only back-propagates the loss computed over the  selected instances.\n  \n\n**This is the implementation:**\ndef ohem_loss( rate, cls_pred, cls_target ):\n    batch_size = cls_pred.size(0) \n    ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n\n    sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\n    keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\n    if keep_num < sorted_ohem_loss.size()[0]:\n        keep_idx_cuda = idx[:keep_num]\n        ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\n    cls_loss = ohem_cls_loss.sum() / keep_num\n    return cls_loss",
    "1120007": "What's the \"similar competition\" that you spoke about? Mind tagging a link here?",
    "1120013": "tyqiangz \nThis is the one:\nhttps://www.kaggle.com/c/cassava-disease",
    "1120201": "Also for me, Cutmix works better than FMix.",
    "1120340": "Hey i have implemented **gamblers loss**  \n\nI trained effnet using CE loss and got 80% acc and then i use gamblers loss and for 10 epochs its accuracy was 20-30%.\n\nAs you said gamblers loss did not worked out, By this you mean same or your accuracy is not this bad ? \n\nCan you please have look to my loss function \nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424\n\nJust to verify implementation was correct.",
    "1120353": "Something ive noticed is that i get very different CV scores on the same folds depending if i use original aspect ratio, or using resize or using centercrop. So ensembling these models could help!",
    "1120618": "Compared to ADAM, does SGD with momentum perform better, except for training speed?",
    "1120941": "When you tried to train deeper nets (B7, B6) did you use Adam or SGD? \nAdam is more memory consuming, so SGD could save you a bit of VRAM.",
    "1121016": "altprof , @sjtuyxc Did not tried SGD yet. It is on my todo list.",
    "1121220": "Thanks~ I got it!",
    "1121307": "Good information ! Great",
    "1122480": "I had a similar experience with effnets and image sizes. \n\nI was using older data and saw some small accuracy increases. I'm curious for those who might have seen increases using these as validation of what their methods may have been.",
    "1122532": "I have similar results for my experiments:\n\n- Other resolutions than `512x512` give me worse performance\n- `EfficientNet B3-B6` are the best architectures i've found so far\n- Adam optimizer (I was planning to try ranger or similar optimizers, maybe a better idea could be to use [SAM](https://arxiv.org/pdf/2010.01412.pdf))\n- External data (from 2019 competition) does not improve performance\n- Label smoothing improves the LB score of `~0.01`\n\n\nI use as data augmentation techniques CutMix and MixUp + random flips, crops and rotations, but maybe it's worth a try to implement snapmix and see if it gives any improvements.\n\nI'm planning to use bi-tempered logistic loss, but up to now my results are the same as those obtained with standard cross-entropy (single fold model LB score `0.890`)",
    "1123536": "lazcoder you using CutMix and MixUp in the same time ?",
    "1123540": "rajan The accuracy was lower pretty significant with gambler loss. Personally, I would go with other losses for this competition",
    "1123543": "Training resize or validation resize?",
    "1123546": "No, I use either CutMix or MixUp randomly on each batch.\n\nI use something like\n\n```\nif p > p_mix:\n    do_mixup(batch)\nelse:\n    do_cutmix(batch)\n```",
    "1123550": "Yes, I agree, It's the best way to do it !",
    "1123556": "Actually I've read somewhere that mixup is harmful for this competition (even if in my experiments it seems not). \nMaybe an idea that is worth to try is to use snapmix instead of mixup",
    "1123559": "If you use them both in the same time, use them separately like in your code piece\nThe opportunity of using one or both of them in this competition is another discussion.\nFor example, I am using just one of them",
    "1125896": "Normal SGD worked better than SGD with momentum of 0.9 for me (learning rate=0.001 on both cases).",
    "1137023": "> 4. Image size smaller than 512 using random resize crop. Although, using random resized crop with default parameters (range of size of the origin size cropped: 0.08, 1.0 and  range of aspect ratio of the origin aspect ratio cropped: 0.75, 1.33 ) decreasing the image size seems to not perform. One thing to try here is the find the optimum parameters of random resize crop methodology that will allow smaller images which will further allow higher batch sizes.\n\n(1) What do you mean by 'Image size smaller than 512' ? The image size in this competition is all 800x600?\n(2) Maybe the cropped  lower limit 0.08 is too small, we need to set it to 0.6 or 0.8 ? Thanks!",
    "1137161": "i have used \n- LabelSmooth\n- cutmix\n- res50 \n- eff B3-b5\n- SCELoss\n- TTA\n- Ranger\nWhat else can i do ? I have seen many people with high accuracy suggesting some one of them. What it is that i am missing \n \nGetting 87.7 max accuracy \n\nmy notebook \nhttps://www.kaggle.com/rajanlagah/noob-programmer-s-code\n\nWill appreciate your feedback",
    "1137677": "how to calculate training acc when using mixup? ( it becomes two labels, such as class_A=0.4 and class_B = 0.6)",
    "1144909": "Why didn't you try it on Google Colab,  it would have given you better compute.",
    "1226651": "what about \n1. custom mean and std deviation ?\n2. SuperConvergence or OneCycleLR?"
  },
  "source": "meta"
}