{
  "id": 220612,
  "title": "Solution Arhitecture Pipeline for Bronze Medal ",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/vlad-vaduva-solution-arhitecture-pipeline-for-bron",
  "author_name": "",
  "post_date": "2021-02-19T01:04:11.898440700Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It had been a very interesting competition filled with of a lot of challenges, starting with the noise from the dataset, the possibility of using images from the previous competition dataset, and the diversity of possible solutions. <br>\nI would like to share to the community one of the designed architectures that lead me and my team to the bronze medal<br>\n<img src=\"https://i.ibb.co/M8QLDP8/general-Schematic.jpg\" alt=\"\"></p>\n<p>Presented above is the general schema of the architecture which was designed from 4 steps.</p>\n<p><strong>First step: Finding the best 3 models for the original provided data</strong></p>\n<p>There were tuned:</p>\n<ul>\n<li>architecture type: EfficientNet B2, EfficientNet B3, EfficientNet B4</li>\n<li>input image size: 480x480, 512x512, 520x520, 600x600</li>\n<li>training epochs number: 30, 40</li>\n<li>learning rate: 0.0001, 0.00015, 0.0002. 0.0003</li>\n<li>learning rate scheduler: Cosine Anealing w/wo warmup, One Cycle, Reduce on Plateau</li>\n<li>aditional layers on top of the network</li>\n<li>random resize crop paramers</li>\n<li>cutout probability, cutout size, cutout number, random brightness, random contrast parameters,     - hue&amp;saturation</li>\n<li>label smoothing w/wo</li>\n<li>cutmix w/wo and the proportions of cutting ([0.3, 0.4, 0.5, 0.6, 0.7])</li>\n<li>imagenet vs noisy student initial weights</li>\n<li>optimizer: Adam, RangerLars, AdamW</li>\n</ul>\n<p>The best 3 models configurations were:</p>\n<p>Model 1:</p>\n<ul>\n<li>Model architecture: EfficientNet B2 (imagenet weights)</li>\n<li>Image size: 600x600</li>\n<li>Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix</li>\n<li>Adam + One cycle (max value: 0.0003)</li>\n</ul>\n<p>Model 2:</p>\n<ul>\n<li>Model architecture: EfficientNet B4 (noisy student weights)</li>\n<li>Image size: 512x512</li>\n<li>Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, </li>\n<li>Adam + One cycle (max value: 0.0002) + Label Smoothing</li>\n</ul>\n<p>Model 3: </p>\n<ul>\n<li>Model architecture: EfficientNet B3 (imagenet weights)</li>\n<li>Image size: 512x512</li>\n<li>Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, </li>\n<li>Adam + One cycle (max value: 0.0002) + Label Smoothing</li>\n</ul>\n<p><strong>Second step: Denoising as much as possible the original data</strong><br>\n<img src=\"https://i.ibb.co/nCqjXT4/cleanImg.jpg\" alt=\"\"><br>\nThe OOF prediction of the presented 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to eliminate the samples that most likely were wrongly label. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We eliminated data where the predicted probabily for ground truth is lower than 0.20 (random chance). <br>\nUsing the new clean dataset, we tune the parameters and find the best possible model on this data</p>\n<p><strong>Third step: Adding useful data from the last year similar competition</strong><br>\n<img src=\"https://i.ibb.co/gtWBnxv/addData.jpg\" alt=\"\"><br>\nThe OOF prediction of the original 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to determine what data are most likely to be correct from the last year dataset. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We added data where the predicted probabily for ground truth is higher than 0.90. <br>\nUsing the new clean dataset, we tune the parameters and find the best possible model on this data</p>\n<p><strong>Forth step: Designing a meta classifier on top of the 3 models (original, cleaned and extended dataset)</strong><br>\n<img src=\"https://i.ibb.co/R7f093v/stacking.jpg\" alt=\"\"></p>\n<p>The input for the final step of the pipeline is the predictions of 3 models(original, cleaned and extended dataset) each with 4 TTA. So in total 12 predictions.<br>\nThere had been tested bagging, boosting and classical stacking and the best result is the one presented below<br>\n<img src=\"https://i.ibb.co/n7NJF2w/meta.jpg\" alt=\"\"></p>\n<p>So, we have a bagging of 100 meta classifiers, where each of the classifier is a two layer system with Random Forest, SVC, LGBM and Ridge at the bottom and SVC on the second layer. The main parameters of the bagging system is max_features = 0.3 and max_sample = 0.1 . <br>\nThe important parameters were tuned with a Bayesian approach. <br>\nAn interesting comparison was made between bagging and boosting methodology. As it you know, boosting is an algorithm which lead to more good results but prone to overfitting and bagging is more robust, generalize better but can give lower accuracy. And our results confirm this hypothesis. In the validation set, boosting provide much better accuracy but on the leaderboard due to the fact that we don't have a very good alignment between validation and learboard, bagging wins due to extra robustness.</p>\n<p>In the end I want to congratulate the winners and to all participants !</p>",
  "messages": [
    {
      "id": "1209608",
      "postDate": "02/19/2021 01:04:11",
      "content": "<p>It had been a very interesting competition filled with of a lot of challenges, starting with the noise from the dataset, the possibility of using images from the previous competition dataset, and the diversity of possible solutions. <br>\nI would like to share to the community one of the designed architectures that lead me and my team to the bronze medal<br>\n<img src=\"https://i.ibb.co/M8QLDP8/general-Schematic.jpg\" alt=\"\"></p>\n<p>Presented above is the general schema of the architecture which was designed from 4 steps.</p>\n<p><strong>First step: Finding the best 3 models for the original provided data</strong></p>\n<p>There were tuned:</p>\n<ul>\n<li>architecture type: EfficientNet B2, EfficientNet B3, EfficientNet B4</li>\n<li>input image size: 480x480, 512x512, 520x520, 600x600</li>\n<li>training epochs number: 30, 40</li>\n<li>learning rate: 0.0001, 0.00015, 0.0002. 0.0003</li>\n<li>learning rate scheduler: Cosine Anealing w/wo warmup, One Cycle, Reduce on Plateau</li>\n<li>aditional layers on top of the network</li>\n<li>random resize crop paramers</li>\n<li>cutout probability, cutout size, cutout number, random brightness, random contrast parameters,     - hue&amp;saturation</li>\n<li>label smoothing w/wo</li>\n<li>cutmix w/wo and the proportions of cutting ([0.3, 0.4, 0.5, 0.6, 0.7])</li>\n<li>imagenet vs noisy student initial weights</li>\n<li>optimizer: Adam, RangerLars, AdamW</li>\n</ul>\n<p>The best 3 models configurations were:</p>\n<p>Model 1:</p>\n<ul>\n<li>Model architecture: EfficientNet B2 (imagenet weights)</li>\n<li>Image size: 600x600</li>\n<li>Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix</li>\n<li>Adam + One cycle (max value: 0.0003)</li>\n</ul>\n<p>Model 2:</p>\n<ul>\n<li>Model architecture: EfficientNet B4 (noisy student weights)</li>\n<li>Image size: 512x512</li>\n<li>Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, </li>\n<li>Adam + One cycle (max value: 0.0002) + Label Smoothing</li>\n</ul>\n<p>Model 3: </p>\n<ul>\n<li>Model architecture: EfficientNet B3 (imagenet weights)</li>\n<li>Image size: 512x512</li>\n<li>Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, </li>\n<li>Adam + One cycle (max value: 0.0002) + Label Smoothing</li>\n</ul>\n<p><strong>Second step: Denoising as much as possible the original data</strong><br>\n<img src=\"https://i.ibb.co/nCqjXT4/cleanImg.jpg\" alt=\"\"><br>\nThe OOF prediction of the presented 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to eliminate the samples that most likely were wrongly label. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We eliminated data where the predicted probabily for ground truth is lower than 0.20 (random chance). <br>\nUsing the new clean dataset, we tune the parameters and find the best possible model on this data</p>\n<p><strong>Third step: Adding useful data from the last year similar competition</strong><br>\n<img src=\"https://i.ibb.co/gtWBnxv/addData.jpg\" alt=\"\"><br>\nThe OOF prediction of the original 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to determine what data are most likely to be correct from the last year dataset. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We added data where the predicted probabily for ground truth is higher than 0.90. <br>\nUsing the new clean dataset, we tune the parameters and find the best possible model on this data</p>\n<p><strong>Forth step: Designing a meta classifier on top of the 3 models (original, cleaned and extended dataset)</strong><br>\n<img src=\"https://i.ibb.co/R7f093v/stacking.jpg\" alt=\"\"></p>\n<p>The input for the final step of the pipeline is the predictions of 3 models(original, cleaned and extended dataset) each with 4 TTA. So in total 12 predictions.<br>\nThere had been tested bagging, boosting and classical stacking and the best result is the one presented below<br>\n<img src=\"https://i.ibb.co/n7NJF2w/meta.jpg\" alt=\"\"></p>\n<p>So, we have a bagging of 100 meta classifiers, where each of the classifier is a two layer system with Random Forest, SVC, LGBM and Ridge at the bottom and SVC on the second layer. The main parameters of the bagging system is max_features = 0.3 and max_sample = 0.1 . <br>\nThe important parameters were tuned with a Bayesian approach. <br>\nAn interesting comparison was made between bagging and boosting methodology. As it you know, boosting is an algorithm which lead to more good results but prone to overfitting and bagging is more robust, generalize better but can give lower accuracy. And our results confirm this hypothesis. In the validation set, boosting provide much better accuracy but on the leaderboard due to the fact that we don't have a very good alignment between validation and learboard, bagging wins due to extra robustness.</p>\n<p>In the end I want to congratulate the winners and to all participants !</p>",
      "rawMarkdown": "It had been a very interesting competition filled with of a lot of challenges, starting with the noise from the dataset, the possibility of using images from the previous competition dataset, and the diversity of possible solutions. \nI would like to share to the community one of the designed architectures that lead me and my team to the bronze medal\n![](https://i.ibb.co/M8QLDP8/general-Schematic.jpg)\n\nPresented above is the general schema of the architecture which was designed from 4 steps.\n\n**First step: Finding the best 3 models for the original provided data**\n\nThere were tuned:\n- architecture type: EfficientNet B2, EfficientNet B3, EfficientNet B4\n- input image size: 480x480, 512x512, 520x520, 600x600\n- training epochs number: 30, 40\n- learning rate: 0.0001, 0.00015, 0.0002. 0.0003\n- learning rate scheduler: Cosine Anealing w/wo warmup, One Cycle, Reduce on Plateau\n- aditional layers on top of the network\n- random resize crop paramers\n- cutout probability, cutout size, cutout number, random brightness, random contrast parameters,     - hue&saturation\n- label smoothing w/wo\n- cutmix w/wo and the proportions of cutting ([0.3, 0.4, 0.5, 0.6, 0.7])\n- imagenet vs noisy student initial weights\n- optimizer: Adam, RangerLars, AdamW\n\nThe best 3 models configurations were:\n\nModel 1:\n- Model architecture: EfficientNet B2 (imagenet weights)\n- Image size: 600x600\n- Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix\n- Adam + One cycle (max value: 0.0003)\n\nModel 2:\n- Model architecture: EfficientNet B4 (noisy student weights)\n- Image size: 512x512\n- Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, \n- Adam + One cycle (max value: 0.0002) + Label Smoothing\n\nModel 3: \n- Model architecture: EfficientNet B3 (imagenet weights)\n- Image size: 512x512\n- Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, \n- Adam + One cycle (max value: 0.0002) + Label Smoothing\n\n\n**Second step: Denoising as much as possible the original data**\n![](https://i.ibb.co/nCqjXT4/cleanImg.jpg)\nThe OOF prediction of the presented 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to eliminate the samples that most likely were wrongly label. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We eliminated data where the predicted probabily for ground truth is lower than 0.20 (random chance). \nUsing the new clean dataset, we tune the parameters and find the best possible model on this data\n\n\n**Third step: Adding useful data from the last year similar competition**\n![](https://i.ibb.co/gtWBnxv/addData.jpg)\nThe OOF prediction of the original 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to determine what data are most likely to be correct from the last year dataset. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We added data where the predicted probabily for ground truth is higher than 0.90. \nUsing the new clean dataset, we tune the parameters and find the best possible model on this data\n\n\n**Forth step: Designing a meta classifier on top of the 3 models (original, cleaned and extended dataset)**\n![](https://i.ibb.co/R7f093v/stacking.jpg)\n\nThe input for the final step of the pipeline is the predictions of 3 models(original, cleaned and extended dataset) each with 4 TTA. So in total 12 predictions.\nThere had been tested bagging, boosting and classical stacking and the best result is the one presented below\n![](https://i.ibb.co/n7NJF2w/meta.jpg)\n\nSo, we have a bagging of 100 meta classifiers, where each of the classifier is a two layer system with Random Forest, SVC, LGBM and Ridge at the bottom and SVC on the second layer. The main parameters of the bagging system is max_features = 0.3 and max_sample = 0.1 . \nThe important parameters were tuned with a Bayesian approach. \nAn interesting comparison was made between bagging and boosting methodology. As it you know, boosting is an algorithm which lead to more good results but prone to overfitting and bagging is more robust, generalize better but can give lower accuracy. And our results confirm this hypothesis. In the validation set, boosting provide much better accuracy but on the leaderboard due to the fact that we don't have a very good alignment between validation and learboard, bagging wins due to extra robustness.\n\n\n\nIn the end I want to congratulate the winners and to all participants !",
      "votes": null
    },
    {
      "id": "1209683",
      "postDate": "02/19/2021 02:28:26",
      "content": "<p>Thank you Vlad for sharing! I have 2 questions: </p>\n<ul>\n<li>Did denosing the data (prob for ground truth &gt; 0.9) give you a boost compared to retraining with original data? Because in my case, all my denoising attempts failed.</li>\n<li>Did you use logits (prediction probabilities of all the classes) or the hard labels (0,1,2,3,4) in stacking? I was thinking to do stacking but thought it wouldn't work. Glad to see it worked for you :)</li>\n</ul>",
      "rawMarkdown": "Thank you Vlad for sharing! I have 2 questions: \n* Did denosing the data (prob for ground truth > 0.9) give you a boost compared to retraining with original data? Because in my case, all my denoising attempts failed.\n* Did you use logits (prediction probabilities of all the classes) or the hard labels (0,1,2,3,4) in stacking? I was thinking to do stacking but thought it wouldn't work. Glad to see it worked for you :)",
      "votes": null
    },
    {
      "id": "1210252",
      "postDate": "02/19/2021 09:42:43",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> . On denoising I eliminated cases where prob for  ground truth &lt; 0.2 (random chance). For that model run solo, the result were indeed lower than on the original data.<br>\nMy team have tried both soft labels and hard labels, for the selected method we used hard labels</p>\n<p>Unfortunate, the submission selection was a huge drawback for me, having cv with leaderboard not aligned made me choose wrong final submissions </p>",
      "rawMarkdown": "Hi @amiiiney . On denoising I eliminated cases where prob for  ground truth < 0.2 (random chance). For that model run solo, the result were indeed lower than on the original data.\nMy team have tried both soft labels and hard labels, for the selected method we used hard labels\n\nUnfortunate, the submission selection was a huge drawback for me, having cv with leaderboard not aligned made me choose wrong final submissions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1209683,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "02/19/2021 02:28:26",
      "content": "<p>Thank you Vlad for sharing! I have 2 questions: </p>\n<ul>\n<li>Did denosing the data (prob for ground truth &gt; 0.9) give you a boost compared to retraining with original data? Because in my case, all my denoising attempts failed.</li>\n<li>Did you use logits (prediction probabilities of all the classes) or the hard labels (0,1,2,3,4) in stacking? I was thinking to do stacking but thought it wouldn't work. Glad to see it worked for you :)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1210252,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/19/2021 09:42:43",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> . On denoising I eliminated cases where prob for  ground truth &lt; 0.2 (random chance). For that model run solo, the result were indeed lower than on the original data.<br>\nMy team have tried both soft labels and hard labels, for the selected method we used hard labels</p>\n<p>Unfortunate, the submission selection was a huge drawback for me, having cv with leaderboard not aligned made me choose wrong final submissions </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1209608": "It had been a very interesting competition filled with of a lot of challenges, starting with the noise from the dataset, the possibility of using images from the previous competition dataset, and the diversity of possible solutions. \nI would like to share to the community one of the designed architectures that lead me and my team to the bronze medal\n![](https://i.ibb.co/M8QLDP8/general-Schematic.jpg)\n\nPresented above is the general schema of the architecture which was designed from 4 steps.\n\n**First step: Finding the best 3 models for the original provided data**\n\nThere were tuned:\n- architecture type: EfficientNet B2, EfficientNet B3, EfficientNet B4\n- input image size: 480x480, 512x512, 520x520, 600x600\n- training epochs number: 30, 40\n- learning rate: 0.0001, 0.00015, 0.0002. 0.0003\n- learning rate scheduler: Cosine Anealing w/wo warmup, One Cycle, Reduce on Plateau\n- aditional layers on top of the network\n- random resize crop paramers\n- cutout probability, cutout size, cutout number, random brightness, random contrast parameters,     - hue&saturation\n- label smoothing w/wo\n- cutmix w/wo and the proportions of cutting ([0.3, 0.4, 0.5, 0.6, 0.7])\n- imagenet vs noisy student initial weights\n- optimizer: Adam, RangerLars, AdamW\n\nThe best 3 models configurations were:\n\nModel 1:\n- Model architecture: EfficientNet B2 (imagenet weights)\n- Image size: 600x600\n- Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix\n- Adam + One cycle (max value: 0.0003)\n\nModel 2:\n- Model architecture: EfficientNet B4 (noisy student weights)\n- Image size: 512x512\n- Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, \n- Adam + One cycle (max value: 0.0002) + Label Smoothing\n\nModel 3: \n- Model architecture: EfficientNet B3 (imagenet weights)\n- Image size: 512x512\n- Augmentations: Cutout(nr:32, size:32, prob:0.5), Random Brightness, Random Contrast, Transpose + Flip, Cutmix, \n- Adam + One cycle (max value: 0.0002) + Label Smoothing\n\n\n**Second step: Denoising as much as possible the original data**\n![](https://i.ibb.co/nCqjXT4/cleanImg.jpg)\nThe OOF prediction of the presented 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to eliminate the samples that most likely were wrongly label. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We eliminated data where the predicted probabily for ground truth is lower than 0.20 (random chance). \nUsing the new clean dataset, we tune the parameters and find the best possible model on this data\n\n\n**Third step: Adding useful data from the last year similar competition**\n![](https://i.ibb.co/gtWBnxv/addData.jpg)\nThe OOF prediction of the original 3 models using 8 TTA combined using a weighted mean(weighted using the OOF score) was used to determine what data are most likely to be correct from the last year dataset. So we use the best OOF accuracy 3 models, each with 8 TTA added for extra robustness. We added data where the predicted probabily for ground truth is higher than 0.90. \nUsing the new clean dataset, we tune the parameters and find the best possible model on this data\n\n\n**Forth step: Designing a meta classifier on top of the 3 models (original, cleaned and extended dataset)**\n![](https://i.ibb.co/R7f093v/stacking.jpg)\n\nThe input for the final step of the pipeline is the predictions of 3 models(original, cleaned and extended dataset) each with 4 TTA. So in total 12 predictions.\nThere had been tested bagging, boosting and classical stacking and the best result is the one presented below\n![](https://i.ibb.co/n7NJF2w/meta.jpg)\n\nSo, we have a bagging of 100 meta classifiers, where each of the classifier is a two layer system with Random Forest, SVC, LGBM and Ridge at the bottom and SVC on the second layer. The main parameters of the bagging system is max_features = 0.3 and max_sample = 0.1 . \nThe important parameters were tuned with a Bayesian approach. \nAn interesting comparison was made between bagging and boosting methodology. As it you know, boosting is an algorithm which lead to more good results but prone to overfitting and bagging is more robust, generalize better but can give lower accuracy. And our results confirm this hypothesis. In the validation set, boosting provide much better accuracy but on the leaderboard due to the fact that we don't have a very good alignment between validation and learboard, bagging wins due to extra robustness.\n\n\n\nIn the end I want to congratulate the winners and to all participants !",
    "1209683": "Thank you Vlad for sharing! I have 2 questions: \n* Did denosing the data (prob for ground truth > 0.9) give you a boost compared to retraining with original data? Because in my case, all my denoising attempts failed.\n* Did you use logits (prediction probabilities of all the classes) or the hard labels (0,1,2,3,4) in stacking? I was thinking to do stacking but thought it wouldn't work. Glad to see it worked for you :)",
    "1210252": "Hi @amiiiney . On denoising I eliminated cases where prob for  ground truth < 0.2 (random chance). For that model run solo, the result were indeed lower than on the original data.\nMy team have tried both soft labels and hard labels, for the selected method we used hard labels\n\nUnfortunate, the submission selection was a huge drawback for me, having cv with leaderboard not aligned made me choose wrong final submissions"
  },
  "source": "meta"
}