{
  "id": 200726,
  "title": "What did/did not work for me so far?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200726",
  "author_name": "",
  "post_date": "2020-12-01T15:54:30.232986500Z",
  "votes": 55,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I know this is little early but it might help to someone!</p>\n<p><strong>Things that did work:</strong></p>\n<ul>\n<li>ResNet18 over SeResNext50_32x4d and EfficientNetb4</li>\n<li>CutMix, CutOut, Mixup, FMix and other basic image aug techniques</li>\n<li>Basic TTA</li>\n<li>Ensemble of different model architectures</li>\n</ul>\n<p><strong>Things that didn't work:</strong></p>\n<ul>\n<li>GridMask Aug</li>\n<li>Bigger models like SeResNext50_32x4d and EfficientNetb4</li>\n<li>Pseudo Labeling</li>\n<li>Label Smoothing</li>\n<li>Training with class weights to balance the data</li>\n<li>2019 competition dataset</li>\n</ul>\n<p><strong>Things I have not tried yet:</strong></p>\n<ul>\n<li>Image sizes other than 512x512</li>\n<li>Multiple fold crossvalidation</li>\n<li>Pseudo Labeling previous competition's unlabeled data</li>\n</ul>\n<p>Using all of these, the highest I could achieve is 0.898 on public LB.</p>\n<p>What about you?</p>",
  "messages": [
    {
      "id": "1098369",
      "postDate": "12/01/2020 15:54:30",
      "content": "<p>I know this is little early but it might help to someone!</p>\n<p><strong>Things that did work:</strong></p>\n<ul>\n<li>ResNet18 over SeResNext50_32x4d and EfficientNetb4</li>\n<li>CutMix, CutOut, Mixup, FMix and other basic image aug techniques</li>\n<li>Basic TTA</li>\n<li>Ensemble of different model architectures</li>\n</ul>\n<p><strong>Things that didn't work:</strong></p>\n<ul>\n<li>GridMask Aug</li>\n<li>Bigger models like SeResNext50_32x4d and EfficientNetb4</li>\n<li>Pseudo Labeling</li>\n<li>Label Smoothing</li>\n<li>Training with class weights to balance the data</li>\n<li>2019 competition dataset</li>\n</ul>\n<p><strong>Things I have not tried yet:</strong></p>\n<ul>\n<li>Image sizes other than 512x512</li>\n<li>Multiple fold crossvalidation</li>\n<li>Pseudo Labeling previous competition's unlabeled data</li>\n</ul>\n<p>Using all of these, the highest I could achieve is 0.898 on public LB.</p>\n<p>What about you?</p>",
      "rawMarkdown": "I know this is little early but it might help to someone!\n\n**Things that did work:**\n- ResNet18 over SeResNext50_32x4d and EfficientNetb4\n- CutMix, CutOut, Mixup, FMix and other basic image aug techniques\n- Basic TTA\n- Ensemble of different model architectures\n\n**Things that didn't work:**\n- GridMask Aug\n- Bigger models like SeResNext50_32x4d and EfficientNetb4\n- Pseudo Labeling\n- Label Smoothing\n- Training with class weights to balance the data\n- 2019 competition dataset\n\n**Things I have not tried yet:**\n- Image sizes other than 512x512\n- Multiple fold crossvalidation\n- Pseudo Labeling previous competition's unlabeled data\n\nUsing all of these, the highest I could achieve is 0.898 on public LB.\n\nWhat about you?",
      "votes": null
    },
    {
      "id": "1098378",
      "postDate": "12/01/2020 15:57:57",
      "content": "<p>yes mine too; simple models like Resnet working better than the new and better ones; I don't understand  why</p>",
      "rawMarkdown": "yes mine too; simple models like Resnet working better than the new and better ones; I don't understand  why",
      "votes": null
    },
    {
      "id": "1098449",
      "postDate": "12/01/2020 16:45:17",
      "content": "<p>Same for me, I cant seem to make efficientnets work better than ResNest50, ill give a try to resnet18!</p>",
      "rawMarkdown": "Same for me, I cant seem to make efficientnets work better than ResNest50, ill give a try to resnet18!",
      "votes": null
    },
    {
      "id": "1098765",
      "postDate": "12/01/2020 20:40:07",
      "content": "<p>I got 0.9 with efficientNetB7 but I train for a long time (5 folds). I need to try with some others architectures</p>",
      "rawMarkdown": "I got 0.9 with efficientNetB7 but I train for a long time (5 folds). I need to try with some others architectures",
      "votes": null
    },
    {
      "id": "1098766",
      "postDate": "12/01/2020 20:40:57",
      "content": "<p>how many epochs is a long time?</p>",
      "rawMarkdown": "how many epochs is a long time?",
      "votes": null
    },
    {
      "id": "1098824",
      "postDate": "12/01/2020 21:48:34",
      "content": "<p>I think we don't need some deep architectures like B7, it's enough to use B4 or use Resnets 18,34 and 50. Thank you</p>",
      "rawMarkdown": "I think we don't need some deep architectures like B7, it's enough to use B4 or use Resnets 18,34 and 50. Thank you",
      "votes": null
    },
    {
      "id": "1098849",
      "postDate": "12/01/2020 22:29:20",
      "content": "<p>For me;</p>\n<p><code>CutMix,TTA,seresnext50,labelsmoothing worked</code></p>\n<p><code>512x512, pseudolabeling test set and finetuning, 2019 dataset, self supervised learning didn't work</code></p>\n<p>my best 0.895 on a single fold model </p>\n<p>That being said beware of the variance and trust your CV!</p>\n<p>I consider something to be worked if it improves my CV and LB at the same time, and by a good margin.</p>",
      "rawMarkdown": "For me;\n\n`CutMix,TTA,seresnext50,labelsmoothing worked`\n\n`512x512, pseudolabeling test set and finetuning, 2019 dataset, self supervised learning didn't work`\n\nmy best 0.895 on a single fold model \n\nThat being said beware of the variance and trust your CV!\n\nI consider something to be worked if it improves my CV and LB at the same time, and by a good margin.",
      "votes": null
    },
    {
      "id": "1098891",
      "postDate": "12/01/2020 23:38:54",
      "content": "<p>The dataset might just be too limited in \"complexity\" for the more advanced models (which can learn even more advanced features) to add anything (and can even lead to overfitting). Unlike the Imagenet dataset, or the flower classification playground competition here on Kaggle, which can have hundreds of classes with substantial variation in appearance, the fact that we here are dealing with the same basic structure for all the classes (Cassava plants), makes it a less \"complex\" task.</p>",
      "rawMarkdown": "The dataset might just be too limited in \"complexity\" for the more advanced models (which can learn even more advanced features) to add anything (and can even lead to overfitting). Unlike the Imagenet dataset, or the flower classification playground competition here on Kaggle, which can have hundreds of classes with substantial variation in appearance, the fact that we here are dealing with the same basic structure for all the classes (Cassava plants), makes it a less \"complex\" task.",
      "votes": null
    },
    {
      "id": "1099420",
      "postDate": "12/02/2020 10:48:39",
      "content": "<p>Yes I think you are correct</p>",
      "rawMarkdown": "Yes I think you are correct",
      "votes": null
    },
    {
      "id": "1099617",
      "postDate": "12/02/2020 13:35:02",
      "content": "<p>100 epochs.</p>\n<p>Yes, I am trying some others architectures or smaller one, I think effb7 is not needed for this competition (at least from the results I got )</p>",
      "rawMarkdown": "100 epochs.\n\nYes, I am trying some others architectures or smaller one, I think effb7 is not needed for this competition (at least from the results I got )",
      "votes": null
    },
    {
      "id": "1099716",
      "postDate": "12/02/2020 14:59:20",
      "content": "<p>2 out of 6 positions from your list of <code>Things that didn't work</code> did work for me.</p>",
      "rawMarkdown": "2 out of 6 positions from your list of `Things that didn't work` did work for me.",
      "votes": null
    },
    {
      "id": "1099789",
      "postDate": "12/02/2020 15:42:35",
      "content": "<p>I was using efficientnet b8-ap. It took almost 35 minutes per epoch, with a batch size of 4, in FastAI and using Cutmix.</p>",
      "rawMarkdown": "I was using efficientnet b8-ap. It took almost 35 minutes per epoch, with a batch size of 4, in FastAI and using Cutmix.",
      "votes": null
    },
    {
      "id": "1099794",
      "postDate": "12/02/2020 15:44:33",
      "content": "<p>If simple architectures like resnet50 are giving good performance, why many people are using GANs? Are they giving better performance than simple imagenet classification architectures ?</p>",
      "rawMarkdown": "If simple architectures like resnet50 are giving good performance, why many people are using GANs? Are they giving better performance than simple imagenet classification architectures ?",
      "votes": null
    },
    {
      "id": "1099808",
      "postDate": "12/02/2020 15:56:20",
      "content": "<p>GANs are not for classification. People are generating more training data using GANs and then they use that data in classification models like resnet, effnet etc. So GAN (Generative Adversarial Network) is just data generator.</p>",
      "rawMarkdown": "GANs are not for classification. People are generating more training data using GANs and then they use that data in classification models like resnet, effnet etc. So GAN (Generative Adversarial Network) is just data generator.",
      "votes": null
    },
    {
      "id": "1099809",
      "postDate": "12/02/2020 15:57:49",
      "content": "<p>Let me guess: It's More complex model and Pseudo Labeling.</p>",
      "rawMarkdown": "Let me guess: It's More complex model and Pseudo Labeling.",
      "votes": null
    },
    {
      "id": "1099812",
      "postDate": "12/02/2020 15:58:27",
      "content": "<blockquote>\n  <p>ResNet18 over SeResNext50_32x4d and EfficientNetb4</p>\n</blockquote>\n<p>If tiny model like ResNet18 can learn from such high resolutions then EFFB4 can definitely learn even more</p>\n<p>May be you need to change your training schedule. </p>",
      "rawMarkdown": "> ResNet18 over SeResNext50_32x4d and EfficientNetb4\n\nIf tiny model like ResNet18 can learn from such high resolutions then EFFB4 can definitely learn even more\n\nMay be you need to change your training schedule.",
      "votes": null
    },
    {
      "id": "1099820",
      "postDate": "12/02/2020 16:02:55",
      "content": "<blockquote>\n  <p>why many people are using GANs</p>\n</blockquote>\n<p>Are you sure of that ?   </p>\n<p>You may use GAN just for Data Augmentation, but you can get good results with just traditional Data Augmentation </p>",
      "rawMarkdown": ">  why many people are using GANs\n\nAre you sure of that ?   \n\nYou may use GAN just for Data Augmentation, but you can get good results with just traditional Data Augmentation",
      "votes": null
    },
    {
      "id": "1099828",
      "postDate": "12/02/2020 16:09:40",
      "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Yes that makes complete sense…In deep learning, as model complexity increases, performance should always improve/stay as it is.</p>\n<p>By training schedule you mean optimizer/LR schedule?</p>",
      "rawMarkdown": "serigne Yes that makes complete sense...In deep learning, as model complexity increases, performance should always improve/stay as it is.\n\nBy training schedule you mean optimizer/LR schedule?",
      "votes": null
    },
    {
      "id": "1099848",
      "postDate": "12/02/2020 16:30:22",
      "content": "<blockquote>\n  <p>Yes that makes complete sense…In deep learning, as model complexity increases, performance should always improve/stay as it is.</p>\n</blockquote>\n<p>It's not always the case.  But from my experiments , with such noise and high resolutions, these tiny models can quickly be trapped in the noise and overfit on them. </p>\n<blockquote>\n  <p>By training schedule you mean optimizer/LR schedule?</p>\n</blockquote>\n<p>Optimizer/loss/LR/ number of epochs. </p>",
      "rawMarkdown": "> Yes that makes complete sense…In deep learning, as model complexity increases, performance should always improve/stay as it is.\n\n\nIt's not always the case.  But from my experiments , with such noise and high resolutions, these tiny models can quickly be trapped in the noise and overfit on them. \n\n\n>  By training schedule you mean optimizer/LR schedule?\n\nOptimizer/loss/LR/ number of epochs.",
      "votes": null
    },
    {
      "id": "1100421",
      "postDate": "12/03/2020 04:42:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a>  and <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> , I trained Se Resnext 50 but it started overfitting to training data after 10 epochs only . From My experience I have seen that CV models take time to converge and hence I was training for 50 epochs . Can you tell me if this is too much because the competition lacks that much complexity </p>\n<p>Also is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?</p>",
      "rawMarkdown": "Hi @kaushal2896  and @serigne , I trained Se Resnext 50 but it started overfitting to training data after 10 epochs only . From My experience I have seen that CV models take time to converge and hence I was training for 50 epochs . Can you tell me if this is too much because the competition lacks that much complexity \n\nAlso is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?",
      "votes": null
    },
    {
      "id": "1100437",
      "postDate": "12/03/2020 04:58:15",
      "content": "<p>vit can be almost the same as the general CNN model, so you can try vit</p>",
      "rawMarkdown": "vit can be almost the same as the general CNN model, so you can try vit",
      "votes": null
    },
    {
      "id": "1100467",
      "postDate": "12/03/2020 05:19:14",
      "content": "<p><a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a> Yea, it work for me also </p>",
      "rawMarkdown": "nroman Yea, it work for me also",
      "votes": null
    },
    {
      "id": "1100479",
      "postDate": "12/03/2020 05:33:01",
      "content": "<p>Not sure about your questions but here's what I experienced with this competition: In any model (smaller/larger), when I train for more epochs loss don't change much after ~25 epochs. Also it doesn't seem to overfit as I've strong regularizers in place. Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.</p>",
      "rawMarkdown": "Not sure about your questions but here's what I experienced with this competition: In any model (smaller/larger), when I train for more epochs loss don't change much after ~25 epochs. Also it doesn't seem to overfit as I've strong regularizers in place. Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.",
      "votes": null
    },
    {
      "id": "1100511",
      "postDate": "12/03/2020 06:11:05",
      "content": "<p><code>as I've strong regularizers in place.</code></p>\n<p>Can you elaborate on this ?</p>\n<p><code>Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.</code></p>\n<p>I think this might be due to the noise being present in the test data as well </p>",
      "rawMarkdown": "`as I've strong regularizers in place.`\n\nCan you elaborate on this ?\n\n`Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.`\n\nI think this might be due to the noise being present in the test data as well",
      "votes": null
    },
    {
      "id": "1100521",
      "postDate": "12/03/2020 06:21:02",
      "content": "<p>By strong regularizers, I mean I have implemented many image augmentations and because of that my validation accuracy remains higher than training accuracy while training. I've noticed that both the values are being stabilized after ~25 epochs.</p>\n<blockquote>\n  <p>I think this is might be due to the noise being present in the test data as well</p>\n</blockquote>\n<p>Shouldn't noise affect equally to smaller models as well?</p>",
      "rawMarkdown": "By strong regularizers, I mean I have implemented many image augmentations and because of that my validation accuracy remains higher than training accuracy while training. I've noticed that both the values are being stabilized after ~25 epochs.\n\n> I think this is might be due to the noise being present in the test data as well\n\nShouldn't noise affect equally to smaller models as well?",
      "votes": null
    },
    {
      "id": "1100592",
      "postDate": "12/03/2020 07:11:38",
      "content": "<p>Is anybody using Progressive Resizing Technique; it's really increasing my cv scores</p>",
      "rawMarkdown": "Is anybody using Progressive Resizing Technique; it's really increasing my cv scores",
      "votes": null
    },
    {
      "id": "1100652",
      "postDate": "12/03/2020 08:10:34",
      "content": "<p>The Bigger models might be doing a better job at classification ie they might be getting the right labels for the right images whereas the smaller models which are prone to overfit the noise at times might be doing a bad job at classification and as their are mislabels present this might be helping in getting better lb</p>\n<p>This is just for the sake of argument though , we don't know  for sure </p>",
      "rawMarkdown": "The Bigger models might be doing a better job at classification ie they might be getting the right labels for the right images whereas the smaller models which are prone to overfit the noise at times might be doing a bad job at classification and as their are mislabels present this might be helping in getting better lb\n\nThis is just for the sake of argument though , we don't know  for sure",
      "votes": null
    },
    {
      "id": "1100737",
      "postDate": "12/03/2020 09:41:53",
      "content": "<blockquote>\n  <p>Also is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?</p>\n</blockquote>\n<p>I use large model for large  epochs. But my training is semi-supervised. </p>",
      "rawMarkdown": "> Also is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?\n\nI use large model for large  epochs. But my training is semi-supervised.",
      "votes": null
    },
    {
      "id": "1100808",
      "postDate": "12/03/2020 11:08:47",
      "content": "<blockquote>\n  <p>Let me guess: It's More complex model and Pseudo Labeling.</p>\n</blockquote>\n<p>Warm :)</p>",
      "rawMarkdown": "> Let me guess: It's More complex model and Pseudo Labeling.\n\nWarm :)",
      "votes": null
    },
    {
      "id": "1100961",
      "postDate": "12/03/2020 14:10:44",
      "content": "<p>i haven't tried any augmentation or finetuning, my baseline for vit trained on imagenet-1k is around 83.3. </p>",
      "rawMarkdown": "i haven't tried any augmentation or finetuning, my baseline for vit trained on imagenet-1k is around 83.3.",
      "votes": null
    },
    {
      "id": "1102577",
      "postDate": "12/05/2020 04:32:09",
      "content": "<p>Which is better, CutMix or Mixup ?</p>",
      "rawMarkdown": "Which is better, CutMix or Mixup ?",
      "votes": null
    },
    {
      "id": "1102692",
      "postDate": "12/05/2020 07:43:35",
      "content": "<p>Cutmix, it's compared in cutmix paper <a href=\"https://arxiv.org/pdf/1905.04899.pdf\" target=\"_blank\">https://arxiv.org/pdf/1905.04899.pdf</a></p>",
      "rawMarkdown": "Cutmix, it's compared in cutmix paper https://arxiv.org/pdf/1905.04899.pdf",
      "votes": null
    },
    {
      "id": "1105685",
      "postDate": "12/08/2020 04:57:50",
      "content": "<p>I am not sure if I am qualified to speak out my thoughts, but still give it a shot. this one is fine-grained classification, with just 5 classes, and differences among classes are not too much, an ordinary people find difficult to make the correct choice. Thus complex models bring more biases.</p>",
      "rawMarkdown": "I am not sure if I am qualified to speak out my thoughts, but still give it a shot. this one is fine-grained classification, with just 5 classes, and differences among classes are not too much, an ordinary people find difficult to make the correct choice. Thus complex models bring more biases.",
      "votes": null
    },
    {
      "id": "1105719",
      "postDate": "12/08/2020 06:12:38",
      "content": "<p>how to do Pseudo Label since test image is invisible, saw many using Pseudo Label on 2019 dataset</p>",
      "rawMarkdown": "how to do Pseudo Label since test image is invisible, saw many using Pseudo Label on 2019 dataset",
      "votes": null
    },
    {
      "id": "1105770",
      "postDate": "12/08/2020 07:24:56",
      "content": "<p>Did you try ensemble of different models trained with different image sizes?</p>",
      "rawMarkdown": "Did you try ensemble of different models trained with different image sizes?",
      "votes": null
    },
    {
      "id": "1106552",
      "postDate": "12/08/2020 23:16:06",
      "content": "<p><a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> \"as model complexity increases, performance should always improve/stay as it is\" Not necessarily. It really depends on the data. In my practice I had situations when the heavy model overfitted to data, while the lighter model did not.</p>",
      "rawMarkdown": "kaushal2896 \"as model complexity increases, performance should always improve/stay as it is\" Not necessarily. It really depends on the data. In my practice I had situations when the heavy model overfitted to data, while the lighter model did not.",
      "votes": null
    },
    {
      "id": "1139221",
      "postDate": "01/05/2021 08:57:19",
      "content": "<p>512x512 not work? Your input size is bigger or smaller</p>",
      "rawMarkdown": "512x512 not work? Your input size is bigger or smaller",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1098378,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/01/2020 15:57:57",
      "content": "<p>yes mine too; simple models like Resnet working better than the new and better ones; I don't understand  why</p>",
      "votes": null,
      "replies": [
        {
          "id": 1098891,
          "author_name": "thomasbrekkunnvik",
          "author_url": "",
          "post_date": "12/01/2020 23:38:54",
          "content": "<p>The dataset might just be too limited in \"complexity\" for the more advanced models (which can learn even more advanced features) to add anything (and can even lead to overfitting). Unlike the Imagenet dataset, or the flower classification playground competition here on Kaggle, which can have hundreds of classes with substantial variation in appearance, the fact that we here are dealing with the same basic structure for all the classes (Cassava plants), makes it a less \"complex\" task.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099420,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "12/02/2020 10:48:39",
          "content": "<p>Yes I think you are correct</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105685,
          "author_name": "plugin1689",
          "author_url": "",
          "post_date": "12/08/2020 04:57:50",
          "content": "<p>I am not sure if I am qualified to speak out my thoughts, but still give it a shot. this one is fine-grained classification, with just 5 classes, and differences among classes are not too much, an ordinary people find difficult to make the correct choice. Thus complex models bring more biases.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1098449,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "12/01/2020 16:45:17",
      "content": "<p>Same for me, I cant seem to make efficientnets work better than ResNest50, ill give a try to resnet18!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1098765,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "12/01/2020 20:40:07",
      "content": "<p>I got 0.9 with efficientNetB7 but I train for a long time (5 folds). I need to try with some others architectures</p>",
      "votes": null,
      "replies": [
        {
          "id": 1098766,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "12/01/2020 20:40:57",
          "content": "<p>how many epochs is a long time?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1098824,
          "author_name": "",
          "author_url": "",
          "post_date": "12/01/2020 21:48:34",
          "content": "<p>I think we don't need some deep architectures like B7, it's enough to use B4 or use Resnets 18,34 and 50. Thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099617,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "12/02/2020 13:35:02",
          "content": "<p>100 epochs.</p>\n<p>Yes, I am trying some others architectures or smaller one, I think effb7 is not needed for this competition (at least from the results I got )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099789,
          "author_name": "adamdavis99",
          "author_url": "",
          "post_date": "12/02/2020 15:42:35",
          "content": "<p>I was using efficientnet b8-ap. It took almost 35 minutes per epoch, with a batch size of 4, in FastAI and using Cutmix.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1098849,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "12/01/2020 22:29:20",
      "content": "<p>For me;</p>\n<p><code>CutMix,TTA,seresnext50,labelsmoothing worked</code></p>\n<p><code>512x512, pseudolabeling test set and finetuning, 2019 dataset, self supervised learning didn't work</code></p>\n<p>my best 0.895 on a single fold model </p>\n<p>That being said beware of the variance and trust your CV!</p>\n<p>I consider something to be worked if it improves my CV and LB at the same time, and by a good margin.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1139221,
          "author_name": "clwwlc",
          "author_url": "",
          "post_date": "01/05/2021 08:57:19",
          "content": "<p>512x512 not work? Your input size is bigger or smaller</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1099716,
      "author_name": "nroman",
      "author_url": "",
      "post_date": "12/02/2020 14:59:20",
      "content": "<p>2 out of 6 positions from your list of <code>Things that didn't work</code> did work for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1099809,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "12/02/2020 15:57:49",
          "content": "<p>Let me guess: It's More complex model and Pseudo Labeling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100467,
          "author_name": "khursani8",
          "author_url": "",
          "post_date": "12/03/2020 05:19:14",
          "content": "<p><a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a> Yea, it work for me also </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100808,
          "author_name": "nroman",
          "author_url": "",
          "post_date": "12/03/2020 11:08:47",
          "content": "<blockquote>\n  <p>Let me guess: It's More complex model and Pseudo Labeling.</p>\n</blockquote>\n<p>Warm :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105719,
          "author_name": "plugin1689",
          "author_url": "",
          "post_date": "12/08/2020 06:12:38",
          "content": "<p>how to do Pseudo Label since test image is invisible, saw many using Pseudo Label on 2019 dataset</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1099794,
      "author_name": "adamdavis99",
      "author_url": "",
      "post_date": "12/02/2020 15:44:33",
      "content": "<p>If simple architectures like resnet50 are giving good performance, why many people are using GANs? Are they giving better performance than simple imagenet classification architectures ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1099808,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "12/02/2020 15:56:20",
          "content": "<p>GANs are not for classification. People are generating more training data using GANs and then they use that data in classification models like resnet, effnet etc. So GAN (Generative Adversarial Network) is just data generator.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099820,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "12/02/2020 16:02:55",
          "content": "<blockquote>\n  <p>why many people are using GANs</p>\n</blockquote>\n<p>Are you sure of that ?   </p>\n<p>You may use GAN just for Data Augmentation, but you can get good results with just traditional Data Augmentation </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1099812,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "12/02/2020 15:58:27",
      "content": "<blockquote>\n  <p>ResNet18 over SeResNext50_32x4d and EfficientNetb4</p>\n</blockquote>\n<p>If tiny model like ResNet18 can learn from such high resolutions then EFFB4 can definitely learn even more</p>\n<p>May be you need to change your training schedule. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1099828,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "12/02/2020 16:09:40",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Yes that makes complete sense…In deep learning, as model complexity increases, performance should always improve/stay as it is.</p>\n<p>By training schedule you mean optimizer/LR schedule?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099848,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "12/02/2020 16:30:22",
          "content": "<blockquote>\n  <p>Yes that makes complete sense…In deep learning, as model complexity increases, performance should always improve/stay as it is.</p>\n</blockquote>\n<p>It's not always the case.  But from my experiments , with such noise and high resolutions, these tiny models can quickly be trapped in the noise and overfit on them. </p>\n<blockquote>\n  <p>By training schedule you mean optimizer/LR schedule?</p>\n</blockquote>\n<p>Optimizer/loss/LR/ number of epochs. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100421,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "12/03/2020 04:42:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a>  and <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> , I trained Se Resnext 50 but it started overfitting to training data after 10 epochs only . From My experience I have seen that CV models take time to converge and hence I was training for 50 epochs . Can you tell me if this is too much because the competition lacks that much complexity </p>\n<p>Also is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100479,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "12/03/2020 05:33:01",
          "content": "<p>Not sure about your questions but here's what I experienced with this competition: In any model (smaller/larger), when I train for more epochs loss don't change much after ~25 epochs. Also it doesn't seem to overfit as I've strong regularizers in place. Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100511,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "12/03/2020 06:11:05",
          "content": "<p><code>as I've strong regularizers in place.</code></p>\n<p>Can you elaborate on this ?</p>\n<p><code>Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.</code></p>\n<p>I think this might be due to the noise being present in the test data as well </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100521,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "12/03/2020 06:21:02",
          "content": "<p>By strong regularizers, I mean I have implemented many image augmentations and because of that my validation accuracy remains higher than training accuracy while training. I've noticed that both the values are being stabilized after ~25 epochs.</p>\n<blockquote>\n  <p>I think this is might be due to the noise being present in the test data as well</p>\n</blockquote>\n<p>Shouldn't noise affect equally to smaller models as well?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100652,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "12/03/2020 08:10:34",
          "content": "<p>The Bigger models might be doing a better job at classification ie they might be getting the right labels for the right images whereas the smaller models which are prone to overfit the noise at times might be doing a bad job at classification and as their are mislabels present this might be helping in getting better lb</p>\n<p>This is just for the sake of argument though , we don't know  for sure </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100737,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "12/03/2020 09:41:53",
          "content": "<blockquote>\n  <p>Also is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?</p>\n</blockquote>\n<p>I use large model for large  epochs. But my training is semi-supervised. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100437,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/03/2020 04:58:15",
      "content": "<p>vit can be almost the same as the general CNN model, so you can try vit</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100961,
          "author_name": "nachiket273",
          "author_url": "",
          "post_date": "12/03/2020 14:10:44",
          "content": "<p>i haven't tried any augmentation or finetuning, my baseline for vit trained on imagenet-1k is around 83.3. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100592,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/03/2020 07:11:38",
      "content": "<p>Is anybody using Progressive Resizing Technique; it's really increasing my cv scores</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105770,
          "author_name": "ajaykumar7778",
          "author_url": "",
          "post_date": "12/08/2020 07:24:56",
          "content": "<p>Did you try ensemble of different models trained with different image sizes?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1102577,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "12/05/2020 04:32:09",
      "content": "<p>Which is better, CutMix or Mixup ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1102692,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "12/05/2020 07:43:35",
          "content": "<p>Cutmix, it's compared in cutmix paper <a href=\"https://arxiv.org/pdf/1905.04899.pdf\" target=\"_blank\">https://arxiv.org/pdf/1905.04899.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1106552,
      "author_name": "etagiev",
      "author_url": "",
      "post_date": "12/08/2020 23:16:06",
      "content": "<p><a href=\"https://www.kaggle.com/kaushal2896\" target=\"_blank\">@kaushal2896</a> \"as model complexity increases, performance should always improve/stay as it is\" Not necessarily. It really depends on the data. In my practice I had situations when the heavy model overfitted to data, while the lighter model did not.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1098369": "I know this is little early but it might help to someone!\n\n**Things that did work:**\n- ResNet18 over SeResNext50_32x4d and EfficientNetb4\n- CutMix, CutOut, Mixup, FMix and other basic image aug techniques\n- Basic TTA\n- Ensemble of different model architectures\n\n**Things that didn't work:**\n- GridMask Aug\n- Bigger models like SeResNext50_32x4d and EfficientNetb4\n- Pseudo Labeling\n- Label Smoothing\n- Training with class weights to balance the data\n- 2019 competition dataset\n\n**Things I have not tried yet:**\n- Image sizes other than 512x512\n- Multiple fold crossvalidation\n- Pseudo Labeling previous competition's unlabeled data\n\nUsing all of these, the highest I could achieve is 0.898 on public LB.\n\nWhat about you?",
    "1098378": "yes mine too; simple models like Resnet working better than the new and better ones; I don't understand  why",
    "1098449": "Same for me, I cant seem to make efficientnets work better than ResNest50, ill give a try to resnet18!",
    "1098765": "I got 0.9 with efficientNetB7 but I train for a long time (5 folds). I need to try with some others architectures",
    "1098766": "how many epochs is a long time?",
    "1098824": "I think we don't need some deep architectures like B7, it's enough to use B4 or use Resnets 18,34 and 50. Thank you",
    "1098849": "For me;\n\n`CutMix,TTA,seresnext50,labelsmoothing worked`\n\n`512x512, pseudolabeling test set and finetuning, 2019 dataset, self supervised learning didn't work`\n\nmy best 0.895 on a single fold model \n\nThat being said beware of the variance and trust your CV!\n\nI consider something to be worked if it improves my CV and LB at the same time, and by a good margin.",
    "1098891": "The dataset might just be too limited in \"complexity\" for the more advanced models (which can learn even more advanced features) to add anything (and can even lead to overfitting). Unlike the Imagenet dataset, or the flower classification playground competition here on Kaggle, which can have hundreds of classes with substantial variation in appearance, the fact that we here are dealing with the same basic structure for all the classes (Cassava plants), makes it a less \"complex\" task.",
    "1099420": "Yes I think you are correct",
    "1099617": "100 epochs.\n\nYes, I am trying some others architectures or smaller one, I think effb7 is not needed for this competition (at least from the results I got )",
    "1099716": "2 out of 6 positions from your list of `Things that didn't work` did work for me.",
    "1099789": "I was using efficientnet b8-ap. It took almost 35 minutes per epoch, with a batch size of 4, in FastAI and using Cutmix.",
    "1099794": "If simple architectures like resnet50 are giving good performance, why many people are using GANs? Are they giving better performance than simple imagenet classification architectures ?",
    "1099808": "GANs are not for classification. People are generating more training data using GANs and then they use that data in classification models like resnet, effnet etc. So GAN (Generative Adversarial Network) is just data generator.",
    "1099809": "Let me guess: It's More complex model and Pseudo Labeling.",
    "1099812": "> ResNet18 over SeResNext50_32x4d and EfficientNetb4\n\nIf tiny model like ResNet18 can learn from such high resolutions then EFFB4 can definitely learn even more\n\nMay be you need to change your training schedule.",
    "1099820": ">  why many people are using GANs\n\nAre you sure of that ?   \n\nYou may use GAN just for Data Augmentation, but you can get good results with just traditional Data Augmentation",
    "1099828": "serigne Yes that makes complete sense...In deep learning, as model complexity increases, performance should always improve/stay as it is.\n\nBy training schedule you mean optimizer/LR schedule?",
    "1099848": "> Yes that makes complete sense…In deep learning, as model complexity increases, performance should always improve/stay as it is.\n\n\nIt's not always the case.  But from my experiments , with such noise and high resolutions, these tiny models can quickly be trapped in the noise and overfit on them. \n\n\n>  By training schedule you mean optimizer/LR schedule?\n\nOptimizer/loss/LR/ number of epochs.",
    "1100421": "Hi @kaushal2896  and @serigne , I trained Se Resnext 50 but it started overfitting to training data after 10 epochs only . From My experience I have seen that CV models take time to converge and hence I was training for 50 epochs . Can you tell me if this is too much because the competition lacks that much complexity \n\nAlso is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?",
    "1100437": "vit can be almost the same as the general CNN model, so you can try vit",
    "1100467": "nroman Yea, it work for me also",
    "1100479": "Not sure about your questions but here's what I experienced with this competition: In any model (smaller/larger), when I train for more epochs loss don't change much after ~25 epochs. Also it doesn't seem to overfit as I've strong regularizers in place. Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.",
    "1100511": "`as I've strong regularizers in place.`\n\nCan you elaborate on this ?\n\n`Larger models tend to have slightly smaller validation loss and higher validation accuracy but still they perform worse on public LB as compared to smaller models. This behaviour confuses me.`\n\nI think this might be due to the noise being present in the test data as well",
    "1100521": "By strong regularizers, I mean I have implemented many image augmentations and because of that my validation accuracy remains higher than training accuracy while training. I've noticed that both the values are being stabilized after ~25 epochs.\n\n> I think this is might be due to the noise being present in the test data as well\n\nShouldn't noise affect equally to smaller models as well?",
    "1100592": "Is anybody using Progressive Resizing Technique; it's really increasing my cv scores",
    "1100652": "The Bigger models might be doing a better job at classification ie they might be getting the right labels for the right images whereas the smaller models which are prone to overfit the noise at times might be doing a bad job at classification and as their are mislabels present this might be helping in getting better lb\n\nThis is just for the sake of argument though , we don't know  for sure",
    "1100737": "> Also is it something like that we need to train for smaller epochs the larger models and more epochs for smaller ones?\n\nI use large model for large  epochs. But my training is semi-supervised.",
    "1100808": "> Let me guess: It's More complex model and Pseudo Labeling.\n\nWarm :)",
    "1100961": "i haven't tried any augmentation or finetuning, my baseline for vit trained on imagenet-1k is around 83.3.",
    "1102577": "Which is better, CutMix or Mixup ?",
    "1102692": "Cutmix, it's compared in cutmix paper https://arxiv.org/pdf/1905.04899.pdf",
    "1105685": "I am not sure if I am qualified to speak out my thoughts, but still give it a shot. this one is fine-grained classification, with just 5 classes, and differences among classes are not too much, an ordinary people find difficult to make the correct choice. Thus complex models bring more biases.",
    "1105719": "how to do Pseudo Label since test image is invisible, saw many using Pseudo Label on 2019 dataset",
    "1105770": "Did you try ensemble of different models trained with different image sizes?",
    "1106552": "kaushal2896 \"as model complexity increases, performance should always improve/stay as it is\" Not necessarily. It really depends on the data. In my practice I had situations when the heavy model overfitted to data, while the lighter model did not.",
    "1139221": "512x512 not work? Your input size is bigger or smaller"
  },
  "source": "meta"
}