{
  "id": 202351,
  "title": "8 techniques to experiment for improvement the model performance ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202351",
  "author_name": "Vlad Vaduva",
  "post_date": "2020-12-09T15:01:48.894000",
  "votes": 43,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This is a draft of my todo shortlist techniques for improvement the models performance .<br>\nIf you have experiment some of them join the conversation with useful feedback related to your experience and also feel free to add new things and ideas to the list.</p>\n<ol>\n<li><p>I see a lot of people using noisy students weights. For me so far in my previous competitions and real application they did not bring any improvements. It would be nice to see a comparison head to head \"imagenet weights\" vs \"noisy student weights\" in different scenarios and see if they really do bring a improvement and if yes, in what circumstances.</p></li>\n<li><p>Fmix really leads to better results than cutmix in this competition ? For me, it hasn't so far</p></li>\n<li><p>Beside cutmix and Fmix, mixup can be a useful technique that can improve the model training. Definitely will be on the todo list </p></li>\n<li><p>So far all my attempts were with Adam optimizer, a comparison will be interesting with ADAMW, RANGER LARS or any other SOTA optimizers.</p></li>\n<li><p>Other losses than crossentropy loss can lead to better results ? (focal loss, etc)</p></li>\n<li><p>Parameters of cutout are also very important. Find out the optimum cover areas of image so it would not cover too much useful information but also help with generalizing better</p></li>\n<li><p>Extending the head of the model with additional layers including dropout for increasing the generalizing power.</p></li>\n<li><p>What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.<br>\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.<br>\nI did not try it yet to this competition but it's on my todo list.</p></li>\n</ol>",
  "messages": [
    {
      "id": 1107282,
      "postDate": "2020-12-09T15:01:48.893Z",
      "content": "<p>This is a draft of my todo shortlist techniques for improvement the models performance .<br>\nIf you have experiment some of them join the conversation with useful feedback related to your experience and also feel free to add new things and ideas to the list.</p>\n<ol>\n<li><p>I see a lot of people using noisy students weights. For me so far in my previous competitions and real application they did not bring any improvements. It would be nice to see a comparison head to head \"imagenet weights\" vs \"noisy student weights\" in different scenarios and see if they really do bring a improvement and if yes, in what circumstances.</p></li>\n<li><p>Fmix really leads to better results than cutmix in this competition ? For me, it hasn't so far</p></li>\n<li><p>Beside cutmix and Fmix, mixup can be a useful technique that can improve the model training. Definitely will be on the todo list </p></li>\n<li><p>So far all my attempts were with Adam optimizer, a comparison will be interesting with ADAMW, RANGER LARS or any other SOTA optimizers.</p></li>\n<li><p>Other losses than crossentropy loss can lead to better results ? (focal loss, etc)</p></li>\n<li><p>Parameters of cutout are also very important. Find out the optimum cover areas of image so it would not cover too much useful information but also help with generalizing better</p></li>\n<li><p>Extending the head of the model with additional layers including dropout for increasing the generalizing power.</p></li>\n<li><p>What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.<br>\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.<br>\nI did not try it yet to this competition but it's on my todo list.</p></li>\n</ol>",
      "rawMarkdown": "This is a draft of my todo shortlist techniques for improvement the models performance .\nIf you have experiment some of them join the conversation with useful feedback related to your experience and also feel free to add new things and ideas to the list.\n\n1. I see a lot of people using noisy students weights. For me so far in my previous competitions and real application they did not bring any improvements. It would be nice to see a comparison head to head \"imagenet weights\" vs \"noisy student weights\" in different scenarios and see if they really do bring a improvement and if yes, in what circumstances.\n\n2. Fmix really leads to better results than cutmix in this competition ? For me, it hasn't so far\n\n3. Beside cutmix and Fmix, mixup can be a useful technique that can improve the model training. Definitely will be on the todo list \n\n4. So far all my attempts were with Adam optimizer, a comparison will be interesting with ADAMW, RANGER LARS or any other SOTA optimizers.\n\n5. Other losses than crossentropy loss can lead to better results ? (focal loss, etc)\n\n6. Parameters of cutout are also very important. Find out the optimum cover areas of image so it would not cover too much useful information but also help with generalizing better\n\n7. Extending the head of the model with additional layers including dropout for increasing the generalizing power.\n\n8. What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.\nI did not try it yet to this competition but it's on my todo list.",
      "votes": 43
    },
    {
      "id": 1108090,
      "postDate": "2020-12-10T08:55:31.570Z",
      "content": "<p>My thoughts are kinda close.</p>\n<ol>\n<li>Mixup should be great here, but only with proper hyperparameters (applies to Cutmix (something like half one image half another));</li>\n<li>Definitely \"high-res\" (512 and higher) images should be used. I would say 600x600 random crop with EfficientNetB7 (or 526x526 with EfficientNetB6) with drop connect=0.2…0.3 as a baseline. Mixed precision could help get everything into the memory;</li>\n<li>Drop duplicates, because they will interfere with class weights. I don’t expect models without class weights get good results on private. Something like 'balanced' or 'sqrt_balanced' schemes should be ok;</li>\n<li>Sigmoid in last layer + BinaryCrossentropy + label smoothing should be used to avoid softmax bottleneck;</li>\n<li>Focal Loss, Hinge losses could be ok, but could also lead to overfit on outliers/noisy data. Robust losses could be huge (bootstrap losses, bce+vat loss, lq loss, generalized cce etc);</li>\n<li>GlobalGEMPooling is a standard option here, but could be messy. Maybe try to use GlobalMean + GlobalMax pooling to save the details (max) and save the gradients (mean);</li>\n<li>Geometric transforms (flip, rotate) as augmentations, maybe dropout. None or very light noise/blur/color augs;</li>\n<li>Adam with gradient clipping and weight decay as a start. SGD with decay and clipping could be better;</li>\n<li>I have a feeling that scraped data and similar datasets added to the training data could be a deciding factor;</li>\n<li>Classic TTA with flips - worth trying;</li>\n<li>Don’t think that dropping non-leaves images is a good idea. They could be present in test set for all we know.</li>\n</ol>",
      "rawMarkdown": "My thoughts are kinda close.\n1. Mixup should be great here, but only with proper hyperparameters (applies to Cutmix (something like half one image half another));\n2. Definitely \"high-res\" (512 and higher) images should be used. I would say 600x600 random crop with EfficientNetB7 (or 526x526 with EfficientNetB6) with drop connect=0.2...0.3 as a baseline. Mixed precision could help get everything into the memory;\n3. Drop duplicates, because they will interfere with class weights. I don’t expect models without class weights get good results on private. Something like 'balanced' or 'sqrt_balanced' schemes should be ok;\n4. Sigmoid in last layer + BinaryCrossentropy + label smoothing should be used to avoid softmax bottleneck;\n5. Focal Loss, Hinge losses could be ok, but could also lead to overfit on outliers/noisy data. Robust losses could be huge (bootstrap losses, bce+vat loss, lq loss, generalized cce etc);\n6. GlobalGEMPooling is a standard option here, but could be messy. Maybe try to use GlobalMean + GlobalMax pooling to save the details (max) and save the gradients (mean);\n7. Geometric transforms (flip, rotate) as augmentations, maybe dropout. None or very light noise/blur/color augs;\n8. Adam with gradient clipping and weight decay as a start. SGD with decay and clipping could be better;\n9. I have a feeling that scraped data and similar datasets added to the training data could be a deciding factor;\n10. Classic TTA with flips - worth trying;\n11. Don’t think that dropping non-leaves images is a good idea. They could be present in test set for all we know.",
      "votes": 6,
      "replies": [
        {
          "id": 1108319,
          "postDate": "2020-12-10T14:40:30.093Z",
          "content": "<p>Good input <a href=\"https://www.kaggle.com/altprof\" target=\"_blank\">@altprof</a> . Label smoothing indeed seems to improve a little the results and GlobalGEMPooling  is a thing to try</p>",
          "rawMarkdown": "Good input @altprof . Label smoothing indeed seems to improve a little the results and GlobalGEMPooling  is a thing to try",
          "votes": 1
        }
      ]
    },
    {
      "id": 1107336,
      "postDate": "2020-12-09T15:49:45.383Z",
      "content": "<p>For me AdamW did not perforn well(Adam did best with cosineLR), label smoothing gave 0.1~ up, Cutmix + Mixup gave good results overall. EFFnetB4 perform very well in many tests.</p>",
      "rawMarkdown": "For me AdamW did not perforn well(Adam did best with cosineLR), label smoothing gave 0.1~ up, Cutmix + Mixup gave good results overall. EFFnetB4 perform very well in many tests.",
      "votes": 1,
      "replies": [
        {
          "id": 1107347,
          "postDate": "2020-12-09T16:00:11.020Z",
          "content": "<p>So far I only used Adam + OneCycle. Tried different maximum learning rate for onecycle and for me, the best was 0.0003. <br>\n<a href=\"https://www.kaggle.com/anku5hk\" target=\"_blank\">@anku5hk</a> Try doing warmup for a couple of epochs and then cosine</p>",
          "rawMarkdown": "So far I only used Adam + OneCycle. Tried different maximum learning rate for onecycle and for me, the best was 0.0003. \n@anku5hk Try doing warmup for a couple of epochs and then cosine",
          "votes": 1
        },
        {
          "id": 1107367,
          "postDate": "2020-12-09T16:19:19.780Z",
          "content": "<p>Okay, I will.👍</p>",
          "rawMarkdown": "Okay, I will.👍"
        },
        {
          "id": 1108171,
          "postDate": "2020-12-10T11:08:25.180Z",
          "content": "<p>In cosineLR when should I restart? I am currently restarting after every epoch.</p>",
          "rawMarkdown": "In cosineLR when should I restart? I am currently restarting after every epoch."
        },
        {
          "id": 1108197,
          "postDate": "2020-12-10T11:46:08.280Z",
          "content": "<p><a href=\"https://www.kaggle.com/koushiksahu\" target=\"_blank\">@koushiksahu</a> restart is not a must, you can get good results with a few warmup epochs and then simple cosine annealing scheduler</p>",
          "rawMarkdown": "@koushiksahu restart is not a must, you can get good results with a few warmup epochs and then simple cosine annealing scheduler"
        },
        {
          "id": 1108331,
          "postDate": "2020-12-10T14:52:02.537Z",
          "content": "<p>Cool, will try this out and see what works</p>",
          "rawMarkdown": "Cool, will try this out and see what works"
        }
      ]
    },
    {
      "id": 1110266,
      "postDate": "2020-12-12T15:09:07.937Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1108090,
      "author_name": "Hleb Levitski",
      "author_url": "",
      "post_date": "2020-12-10T08:55:31.570000",
      "content": "<p>My thoughts are kinda close.</p>\n<ol>\n<li>Mixup should be great here, but only with proper hyperparameters (applies to Cutmix (something like half one image half another));</li>\n<li>Definitely \"high-res\" (512 and higher) images should be used. I would say 600x600 random crop with EfficientNetB7 (or 526x526 with EfficientNetB6) with drop connect=0.2…0.3 as a baseline. Mixed precision could help get everything into the memory;</li>\n<li>Drop duplicates, because they will interfere with class weights. I don’t expect models without class weights get good results on private. Something like 'balanced' or 'sqrt_balanced' schemes should be ok;</li>\n<li>Sigmoid in last layer + BinaryCrossentropy + label smoothing should be used to avoid softmax bottleneck;</li>\n<li>Focal Loss, Hinge losses could be ok, but could also lead to overfit on outliers/noisy data. Robust losses could be huge (bootstrap losses, bce+vat loss, lq loss, generalized cce etc);</li>\n<li>GlobalGEMPooling is a standard option here, but could be messy. Maybe try to use GlobalMean + GlobalMax pooling to save the details (max) and save the gradients (mean);</li>\n<li>Geometric transforms (flip, rotate) as augmentations, maybe dropout. None or very light noise/blur/color augs;</li>\n<li>Adam with gradient clipping and weight decay as a start. SGD with decay and clipping could be better;</li>\n<li>I have a feeling that scraped data and similar datasets added to the training data could be a deciding factor;</li>\n<li>Classic TTA with flips - worth trying;</li>\n<li>Don’t think that dropping non-leaves images is a good idea. They could be present in test set for all we know.</li>\n</ol>",
      "votes": 6,
      "replies": [
        {
          "id": 1108319,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-12-10T14:40:30.093000",
          "content": "<p>Good input <a href=\"https://www.kaggle.com/altprof\" target=\"_blank\">@altprof</a> . Label smoothing indeed seems to improve a little the results and GlobalGEMPooling  is a thing to try</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1107336,
      "author_name": "Ankush kuwar",
      "author_url": "",
      "post_date": "2020-12-09T15:49:45.383000",
      "content": "<p>For me AdamW did not perforn well(Adam did best with cosineLR), label smoothing gave 0.1~ up, Cutmix + Mixup gave good results overall. EFFnetB4 perform very well in many tests.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1107347,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-12-09T16:00:11.020000",
          "content": "<p>So far I only used Adam + OneCycle. Tried different maximum learning rate for onecycle and for me, the best was 0.0003. <br>\n<a href=\"https://www.kaggle.com/anku5hk\" target=\"_blank\">@anku5hk</a> Try doing warmup for a couple of epochs and then cosine</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107367,
          "author_name": "Ankush kuwar",
          "author_url": "",
          "post_date": "2020-12-09T16:19:19.780000",
          "content": "<p>Okay, I will.👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1108171,
          "author_name": "Koushik Sahu",
          "author_url": "",
          "post_date": "2020-12-10T11:08:25.180000",
          "content": "<p>In cosineLR when should I restart? I am currently restarting after every epoch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1108197,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-12-10T11:46:08.280000",
          "content": "<p><a href=\"https://www.kaggle.com/koushiksahu\" target=\"_blank\">@koushiksahu</a> restart is not a must, you can get good results with a few warmup epochs and then simple cosine annealing scheduler</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1108331,
          "author_name": "Koushik Sahu",
          "author_url": "",
          "post_date": "2020-12-10T14:52:02.537000",
          "content": "<p>Cool, will try this out and see what works</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1110266,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-12T15:09:07.937000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1107282": "This is a draft of my todo shortlist techniques for improvement the models performance .\nIf you have experiment some of them join the conversation with useful feedback related to your experience and also feel free to add new things and ideas to the list.\n\n1. I see a lot of people using noisy students weights. For me so far in my previous competitions and real application they did not bring any improvements. It would be nice to see a comparison head to head \"imagenet weights\" vs \"noisy student weights\" in different scenarios and see if they really do bring a improvement and if yes, in what circumstances.\n\n2. Fmix really leads to better results than cutmix in this competition ? For me, it hasn't so far\n\n3. Beside cutmix and Fmix, mixup can be a useful technique that can improve the model training. Definitely will be on the todo list \n\n4. So far all my attempts were with Adam optimizer, a comparison will be interesting with ADAMW, RANGER LARS or any other SOTA optimizers.\n\n5. Other losses than crossentropy loss can lead to better results ? (focal loss, etc)\n\n6. Parameters of cutout are also very important. Find out the optimum cover areas of image so it would not cover too much useful information but also help with generalizing better\n\n7. Extending the head of the model with additional layers including dropout for increasing the generalizing power.\n\n8. What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.\nI did not try it yet to this competition but it's on my todo list.",
    "1108090": "My thoughts are kinda close.\n1. Mixup should be great here, but only with proper hyperparameters (applies to Cutmix (something like half one image half another));\n2. Definitely \"high-res\" (512 and higher) images should be used. I would say 600x600 random crop with EfficientNetB7 (or 526x526 with EfficientNetB6) with drop connect=0.2...0.3 as a baseline. Mixed precision could help get everything into the memory;\n3. Drop duplicates, because they will interfere with class weights. I don’t expect models without class weights get good results on private. Something like 'balanced' or 'sqrt_balanced' schemes should be ok;\n4. Sigmoid in last layer + BinaryCrossentropy + label smoothing should be used to avoid softmax bottleneck;\n5. Focal Loss, Hinge losses could be ok, but could also lead to overfit on outliers/noisy data. Robust losses could be huge (bootstrap losses, bce+vat loss, lq loss, generalized cce etc);\n6. GlobalGEMPooling is a standard option here, but could be messy. Maybe try to use GlobalMean + GlobalMax pooling to save the details (max) and save the gradients (mean);\n7. Geometric transforms (flip, rotate) as augmentations, maybe dropout. None or very light noise/blur/color augs;\n8. Adam with gradient clipping and weight decay as a start. SGD with decay and clipping could be better;\n9. I have a feeling that scraped data and similar datasets added to the training data could be a deciding factor;\n10. Classic TTA with flips - worth trying;\n11. Don’t think that dropping non-leaves images is a good idea. They could be present in test set for all we know.",
    "1107336": "For me AdamW did not perforn well(Adam did best with cosineLR), label smoothing gave 0.1~ up, Cutmix + Mixup gave good results overall. EFFnetB4 perform very well in many tests.",
    "1110266": ""
  }
}