{
  "id": 220934,
  "title": "Diversity is all we need",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/overfit-diversity-is-all-we-need",
  "author_name": "",
  "post_date": "2021-02-20T06:43:58.774757500Z",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Congrats to all. Learned a lot from this community.</p>\n<p>I am really happy to team up with my teammates <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> and <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> who have strong engineering capabilities and are thus able to put any ideas into practice. Our final solutions consist of 4 models. Unfortunately, same as most teams, we missed selecting a sub that can beat 3rd place in Private LB. We three have totally different training processes, which I think helps us to stay gold zone on both public LB and private LB. The diversity makes us win.</p>\n<h1>Models</h1>\n<ul>\n<li>Shirok’s effb5_ns with 8xTTA</li>\n<li>toxu’s vit_base_patch16_384 with 8xTTA</li>\n<li>Yimin’s effb5_ns with 8xTTA</li>\n<li>Yimin’s seresnext50 with no TTA</li>\n<li>Same weight for ensemble.</li>\n<li>Without any one of these models, we cannot win.</li>\n<li>The oof corrcoef is as below:<br>\n[[1.        , 0.79848116, 0.81904038, 0.81931724],<br>\n   [0.79848116, 1.        , 0.80287945, 0.79185962],<br>\n   [0.81904038, 0.80287945, 1.        , 0.82176191],<br>\n   [0.81931724, 0.79185962, 0.82176191, 1.        ]]</li>\n</ul>\n<h1>Diversity</h1>\n<h2>Teammates’ part</h2>\n<p>As I know, Shirok used tf TPU for training b5, and toxu used skd softlabel for vit. Welcome them to add more details at this part if would like.</p>\n<h2>In my part</h2>\n<h3>Dataset</h3>\n<ul>\n<li>2019+2020 dataset</li>\n<li>Removed some duplicates searched by <a href=\"https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions\" target=\"_blank\">https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions</a> and <a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan</a>. Thanks again for sharing useful kernels that can be used anywhere.</li>\n<li>Tried Pseudo Label on 2019 test data and extraimage data (removed duplicates). Got better CV but worse LB. Probably because of data leak. I used the avg. probs of 5 folds, which may bring in leak info. between folds. Due to time limit, never chose it after.</li>\n<li>Tried self knowledge distillation softlabel. Improved 0.003 on CV but decreased 0.001-0.002 on LB. Did not choose it as the final sub. -0.0015 around compared with hard label on private LB.</li>\n<li>Hard Label</li>\n<li>Stratified 5 folds</li>\n<li>Batch size 32</li>\n<li>Image size 512 x 512</li>\n</ul>\n<h3>Training AUG</h3>\n<p>Random Rotate<br>\nShiftScaleRotate<br>\nHueSaturationValue<br>\nRandomBrightnessContrast<br>\nFlip, transpose<br>\nGuaussNoise<br>\nBlur<br>\nDistortion<br>\nCutout</p>\n<h3>Loss</h3>\n<p>Bi-tempered</p>\n<h3>TTA</h3>\n<p>You must be curious about why I did not use TTA for se50 but used TTA for b5. I was stuck at 0.900 (se50), 0.902 (b5), and 0.902 (ensembles using both 8xTTA) before I checked out this <a href=\"https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903\" target=\"_blank\">notebook</a>. I was also curious about why Kaito did not use TTA for a model but used it for another model. Then I tried to inference my oof with and without various TTAs. And I found that TTA always decreased my cv score of se50 and increased the cv score of b5. The same situation happened on LB. Here is my previous <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/214559\" target=\"_blank\">discussion</a> regarding this. After using the similar strategy, I climbed up to 0.905 on LB. So, I just trusted the cv and LB at that moment, although the private LB shows the contrary result regarding se50. But our ensembles took almost 8 hrs to run submit, so it actually did not allow another model with TTA. It was the fate. Anyway, I was probably wrong with insights about TTA at my that discussion.</p>\n<h2>ensemble</h2>\n<p>My se50+b5 scored 0.905 on public LB, and Shirok’s b5 + toxu’s vit scored 0.906. After merge, we got 0.908. After finetuning each single model, we finally got 0.910 which stand 2nd place on public LB. Our ensemble is somehow robust. It keeps scoring 0.901-0.902 on private LB where the gold zone is. I attribute it to the diversity of our models (to be proved).</p>\n<h1>Trust CV or LB?</h1>\n<p>Unlike the Riiid, I have totally no idea about this comp. The dataset (both train and test) is noisy, and I do not think that test data split is balanced. Our CV is range from 0.905-0.906 on 2020 dataset. It is not aligned with LB for us. Anyone can share your scores may help to determine. Anyway, trusting cv is always the better way. Make sure there is no leak in your cv.</p>\n<h1>Additional</h1>\n<p>The Class Activation Map (CAM) can somewhat help to track your model performance. Introduced by toxu, inspired by the winner in cvpr2020-plant-pathology. I used it for error analysis to determine if the model cares about the right position of the image. I made a simple <a href=\"https://www.kaggle.com/woshifym/cld-class-activation-mapping-cam-wip\" target=\"_blank\">notebook</a> to show how it works. Hope you would enjoy it.</p>\n<h1>Others</h1>\n<p>Forgive my messy formatting.<br>\nWill update if I think of something missing…</p>",
  "messages": [
    {
      "id": "1211358",
      "postDate": "02/20/2021 06:43:58",
      "content": "<p>Congrats to all. Learned a lot from this community.</p>\n<p>I am really happy to team up with my teammates <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> and <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> who have strong engineering capabilities and are thus able to put any ideas into practice. Our final solutions consist of 4 models. Unfortunately, same as most teams, we missed selecting a sub that can beat 3rd place in Private LB. We three have totally different training processes, which I think helps us to stay gold zone on both public LB and private LB. The diversity makes us win.</p>\n<h1>Models</h1>\n<ul>\n<li>Shirok’s effb5_ns with 8xTTA</li>\n<li>toxu’s vit_base_patch16_384 with 8xTTA</li>\n<li>Yimin’s effb5_ns with 8xTTA</li>\n<li>Yimin’s seresnext50 with no TTA</li>\n<li>Same weight for ensemble.</li>\n<li>Without any one of these models, we cannot win.</li>\n<li>The oof corrcoef is as below:<br>\n[[1.        , 0.79848116, 0.81904038, 0.81931724],<br>\n   [0.79848116, 1.        , 0.80287945, 0.79185962],<br>\n   [0.81904038, 0.80287945, 1.        , 0.82176191],<br>\n   [0.81931724, 0.79185962, 0.82176191, 1.        ]]</li>\n</ul>\n<h1>Diversity</h1>\n<h2>Teammates’ part</h2>\n<p>As I know, Shirok used tf TPU for training b5, and toxu used skd softlabel for vit. Welcome them to add more details at this part if would like.</p>\n<h2>In my part</h2>\n<h3>Dataset</h3>\n<ul>\n<li>2019+2020 dataset</li>\n<li>Removed some duplicates searched by <a href=\"https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions\" target=\"_blank\">https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions</a> and <a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan</a>. Thanks again for sharing useful kernels that can be used anywhere.</li>\n<li>Tried Pseudo Label on 2019 test data and extraimage data (removed duplicates). Got better CV but worse LB. Probably because of data leak. I used the avg. probs of 5 folds, which may bring in leak info. between folds. Due to time limit, never chose it after.</li>\n<li>Tried self knowledge distillation softlabel. Improved 0.003 on CV but decreased 0.001-0.002 on LB. Did not choose it as the final sub. -0.0015 around compared with hard label on private LB.</li>\n<li>Hard Label</li>\n<li>Stratified 5 folds</li>\n<li>Batch size 32</li>\n<li>Image size 512 x 512</li>\n</ul>\n<h3>Training AUG</h3>\n<p>Random Rotate<br>\nShiftScaleRotate<br>\nHueSaturationValue<br>\nRandomBrightnessContrast<br>\nFlip, transpose<br>\nGuaussNoise<br>\nBlur<br>\nDistortion<br>\nCutout</p>\n<h3>Loss</h3>\n<p>Bi-tempered</p>\n<h3>TTA</h3>\n<p>You must be curious about why I did not use TTA for se50 but used TTA for b5. I was stuck at 0.900 (se50), 0.902 (b5), and 0.902 (ensembles using both 8xTTA) before I checked out this <a href=\"https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903\" target=\"_blank\">notebook</a>. I was also curious about why Kaito did not use TTA for a model but used it for another model. Then I tried to inference my oof with and without various TTAs. And I found that TTA always decreased my cv score of se50 and increased the cv score of b5. The same situation happened on LB. Here is my previous <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/214559\" target=\"_blank\">discussion</a> regarding this. After using the similar strategy, I climbed up to 0.905 on LB. So, I just trusted the cv and LB at that moment, although the private LB shows the contrary result regarding se50. But our ensembles took almost 8 hrs to run submit, so it actually did not allow another model with TTA. It was the fate. Anyway, I was probably wrong with insights about TTA at my that discussion.</p>\n<h2>ensemble</h2>\n<p>My se50+b5 scored 0.905 on public LB, and Shirok’s b5 + toxu’s vit scored 0.906. After merge, we got 0.908. After finetuning each single model, we finally got 0.910 which stand 2nd place on public LB. Our ensemble is somehow robust. It keeps scoring 0.901-0.902 on private LB where the gold zone is. I attribute it to the diversity of our models (to be proved).</p>\n<h1>Trust CV or LB?</h1>\n<p>Unlike the Riiid, I have totally no idea about this comp. The dataset (both train and test) is noisy, and I do not think that test data split is balanced. Our CV is range from 0.905-0.906 on 2020 dataset. It is not aligned with LB for us. Anyone can share your scores may help to determine. Anyway, trusting cv is always the better way. Make sure there is no leak in your cv.</p>\n<h1>Additional</h1>\n<p>The Class Activation Map (CAM) can somewhat help to track your model performance. Introduced by toxu, inspired by the winner in cvpr2020-plant-pathology. I used it for error analysis to determine if the model cares about the right position of the image. I made a simple <a href=\"https://www.kaggle.com/woshifym/cld-class-activation-mapping-cam-wip\" target=\"_blank\">notebook</a> to show how it works. Hope you would enjoy it.</p>\n<h1>Others</h1>\n<p>Forgive my messy formatting.<br>\nWill update if I think of something missing…</p>",
      "rawMarkdown": "Congrats to all. Learned a lot from this community.\n\nI am really happy to team up with my teammates @ludovick and @tonyxu who have strong engineering capabilities and are thus able to put any ideas into practice. Our final solutions consist of 4 models. Unfortunately, same as most teams, we missed selecting a sub that can beat 3rd place in Private LB. We three have totally different training processes, which I think helps us to stay gold zone on both public LB and private LB. The diversity makes us win.\n\n# Models\n- Shirok’s effb5_ns with 8xTTA\n- toxu’s vit_base_patch16_384 with 8xTTA\n- Yimin’s effb5_ns with 8xTTA\n- Yimin’s seresnext50 with no TTA\n- Same weight for ensemble.\n- Without any one of these models, we cannot win.\n- The oof corrcoef is as below:\n[[1.        , 0.79848116, 0.81904038, 0.81931724],\n       [0.79848116, 1.        , 0.80287945, 0.79185962],\n       [0.81904038, 0.80287945, 1.        , 0.82176191],\n       [0.81931724, 0.79185962, 0.82176191, 1.        ]]\n\n# Diversity\n## Teammates’ part\nAs I know, Shirok used tf TPU for training b5, and toxu used skd softlabel for vit. Welcome them to add more details at this part if would like.\n## In my part\n### Dataset\n- 2019+2020 dataset\n- Removed some duplicates searched by https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions and https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan. Thanks again for sharing useful kernels that can be used anywhere.\n- Tried Pseudo Label on 2019 test data and extraimage data (removed duplicates). Got better CV but worse LB. Probably because of data leak. I used the avg. probs of 5 folds, which may bring in leak info. between folds. Due to time limit, never chose it after.\n- Tried self knowledge distillation softlabel. Improved 0.003 on CV but decreased 0.001-0.002 on LB. Did not choose it as the final sub. -0.0015 around compared with hard label on private LB.\n- Hard Label\n- Stratified 5 folds\n- Batch size 32\n- Image size 512 x 512\n### Training AUG\nRandom Rotate\nShiftScaleRotate\nHueSaturationValue\nRandomBrightnessContrast\nFlip, transpose\nGuaussNoise\nBlur\nDistortion\nCutout\n### Loss\nBi-tempered\n### TTA\nYou must be curious about why I did not use TTA for se50 but used TTA for b5. I was stuck at 0.900 (se50), 0.902 (b5), and 0.902 (ensembles using both 8xTTA) before I checked out this [notebook](https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903). I was also curious about why Kaito did not use TTA for a model but used it for another model. Then I tried to inference my oof with and without various TTAs. And I found that TTA always decreased my cv score of se50 and increased the cv score of b5. The same situation happened on LB. Here is my previous [discussion]( https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/214559) regarding this. After using the similar strategy, I climbed up to 0.905 on LB. So, I just trusted the cv and LB at that moment, although the private LB shows the contrary result regarding se50. But our ensembles took almost 8 hrs to run submit, so it actually did not allow another model with TTA. It was the fate. Anyway, I was probably wrong with insights about TTA at my that discussion.\n\n## ensemble\nMy se50+b5 scored 0.905 on public LB, and Shirok’s b5 + toxu’s vit scored 0.906. After merge, we got 0.908. After finetuning each single model, we finally got 0.910 which stand 2nd place on public LB. Our ensemble is somehow robust. It keeps scoring 0.901-0.902 on private LB where the gold zone is. I attribute it to the diversity of our models (to be proved).\n\n# Trust CV or LB?\nUnlike the Riiid, I have totally no idea about this comp. The dataset (both train and test) is noisy, and I do not think that test data split is balanced. Our CV is range from 0.905-0.906 on 2020 dataset. It is not aligned with LB for us. Anyone can share your scores may help to determine. Anyway, trusting cv is always the better way. Make sure there is no leak in your cv.\n\n# Additional\nThe Class Activation Map (CAM) can somewhat help to track your model performance. Introduced by toxu, inspired by the winner in cvpr2020-plant-pathology. I used it for error analysis to determine if the model cares about the right position of the image. I made a simple [notebook](https://www.kaggle.com/woshifym/cld-class-activation-mapping-cam-wip) to show how it works. Hope you would enjoy it.\n\n# Others\nForgive my messy formatting.\nWill update if I think of something missing...",
      "votes": null
    },
    {
      "id": "1211564",
      "postDate": "02/20/2021 10:33:08",
      "content": "<p>Congratulations! Perhaps you could change the title to indicate your place :)</p>",
      "rawMarkdown": "Congratulations! Perhaps you could change the title to indicate your place :)",
      "votes": null
    },
    {
      "id": "1211685",
      "postDate": "02/20/2021 12:53:51",
      "content": "<p>Congratulations on your first gold medal and learned a lot from you.🎉🎉🎉</p>",
      "rawMarkdown": "Congratulations on your first gold medal and learned a lot from you.🎉🎉🎉",
      "votes": null
    },
    {
      "id": "1211786",
      "postDate": "02/20/2021 15:01:00",
      "content": "<p>Same for me, i removed all duplicates to make sure of no data leak and had a ensemble CV of 0.905. I trusted CV and my best submission was my best CV however i got 0.898 in private. Congrats on your gold medal!!!</p>",
      "rawMarkdown": "Same for me, i removed all duplicates to make sure of no data leak and had a ensemble CV of 0.905. I trusted CV and my best submission was my best CV however i got 0.898 in private. Congrats on your gold medal!!!",
      "votes": null
    },
    {
      "id": "1211794",
      "postDate": "02/20/2021 15:09:33",
      "content": "<p>Congratulations on your 11th place! </p>\n<p>By the way, how did you calculate the correlations?<br>\nIn my oofs, the correlations between each models are higher than yours. <br>\nNever dropped below 0.92 :)</p>",
      "rawMarkdown": "Congratulations on your 11th place! \n\nBy the way, how did you calculate the correlations?\nIn my oofs, the correlations between each models are higher than yours. \nNever dropped below 0.92 :)",
      "votes": null
    },
    {
      "id": "1211875",
      "postDate": "02/20/2021 16:18:01",
      "content": "<p>Congratulations on your medal mate, rly learned a lot from you. I am new to this, i want to ask a question, since almost everybody is talking about that in this competition. I want to ask how do you make the out of fold predictions? Does that mean we make 5 fold cv or whatever number of folds, and we test the model on the validation data where the model is not being trained on? Which means we can also add TTAs and average the predictions. With that, we generate four more predictions (if k=5), and based on that we repeat the same procedure 4 more times, and we obtain the matrix that you showed above? One more question, how did you calculate the correlation to produce the matrix as shown above? <br>\nThank you, cheers 🎉</p>",
      "rawMarkdown": "Congratulations on your medal mate, rly learned a lot from you. I am new to this, i want to ask a question, since almost everybody is talking about that in this competition. I want to ask how do you make the out of fold predictions? Does that mean we make 5 fold cv or whatever number of folds, and we test the model on the validation data where the model is not being trained on? Which means we can also add TTAs and average the predictions. With that, we generate four more predictions (if k=5), and based on that we repeat the same procedure 4 more times, and we obtain the matrix that you showed above? One more question, how did you calculate the correlation to produce the matrix as shown above? \nThank you, cheers 🎉",
      "votes": null
    },
    {
      "id": "1212177",
      "postDate": "02/21/2021 00:58:02",
      "content": "<p>Thanks mane. OOF means your predictions for validation data. If you split your data into 5 folds, you will predict 5 validation fold where the model is not being trained on. The OOF is the combination of this 5 folds. You can use any strategy on the OOF, like TTA or postprocess, to evaluate your model.</p>\n<p>You can use np.corrcoef([predict1, predict2, predict3]) to calculate the correlation coefs for your different predictions.</p>",
      "rawMarkdown": "Thanks mane. OOF means your predictions for validation data. If you split your data into 5 folds, you will predict 5 validation fold where the model is not being trained on. The OOF is the combination of this 5 folds. You can use any strategy on the OOF, like TTA or postprocess, to evaluate your model.\n\nYou can use np.corrcoef([predict1, predict2, predict3]) to calculate the correlation coefs for your different predictions.",
      "votes": null
    },
    {
      "id": "1212471",
      "postDate": "02/21/2021 08:45:00",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> for the wonderful writeup and explanation. Congrats on your gold too. I got to learn a lot from you.  Just a noob question. How to use the corelation matrix that you got for ensembling. Do we weight each model by the corresponding row values rather than averaging the model predictions?</p>",
      "rawMarkdown": "Thanks @woshifym for the wonderful writeup and explanation. Congrats on your gold too. I got to learn a lot from you.  Just a noob question. How to use the corelation matrix that you got for ensembling. Do we weight each model by the corresponding row values rather than averaging the model predictions?",
      "votes": null
    },
    {
      "id": "1213448",
      "postDate": "02/22/2021 05:55:01",
      "content": "<p>Congratulations💥💥💥</p>",
      "rawMarkdown": "Congratulations💥💥💥",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1211564,
      "author_name": "xfffrank",
      "author_url": "",
      "post_date": "02/20/2021 10:33:08",
      "content": "<p>Congratulations! Perhaps you could change the title to indicate your place :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211685,
      "author_name": "tonyxu",
      "author_url": "",
      "post_date": "02/20/2021 12:53:51",
      "content": "<p>Congratulations on your first gold medal and learned a lot from you.🎉🎉🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211786,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "02/20/2021 15:01:00",
      "content": "<p>Same for me, i removed all duplicates to make sure of no data leak and had a ensemble CV of 0.905. I trusted CV and my best submission was my best CV however i got 0.898 in private. Congrats on your gold medal!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211794,
      "author_name": "tmhrkt",
      "author_url": "",
      "post_date": "02/20/2021 15:09:33",
      "content": "<p>Congratulations on your 11th place! </p>\n<p>By the way, how did you calculate the correlations?<br>\nIn my oofs, the correlations between each models are higher than yours. <br>\nNever dropped below 0.92 :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211875,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "02/20/2021 16:18:01",
      "content": "<p>Congratulations on your medal mate, rly learned a lot from you. I am new to this, i want to ask a question, since almost everybody is talking about that in this competition. I want to ask how do you make the out of fold predictions? Does that mean we make 5 fold cv or whatever number of folds, and we test the model on the validation data where the model is not being trained on? Which means we can also add TTAs and average the predictions. With that, we generate four more predictions (if k=5), and based on that we repeat the same procedure 4 more times, and we obtain the matrix that you showed above? One more question, how did you calculate the correlation to produce the matrix as shown above? <br>\nThank you, cheers 🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212177,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "02/21/2021 00:58:02",
          "content": "<p>Thanks mane. OOF means your predictions for validation data. If you split your data into 5 folds, you will predict 5 validation fold where the model is not being trained on. The OOF is the combination of this 5 folds. You can use any strategy on the OOF, like TTA or postprocess, to evaluate your model.</p>\n<p>You can use np.corrcoef([predict1, predict2, predict3]) to calculate the correlation coefs for your different predictions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1212471,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "02/21/2021 08:45:00",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> for the wonderful writeup and explanation. Congrats on your gold too. I got to learn a lot from you.  Just a noob question. How to use the corelation matrix that you got for ensembling. Do we weight each model by the corresponding row values rather than averaging the model predictions?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1213448,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "02/22/2021 05:55:01",
      "content": "<p>Congratulations💥💥💥</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1211358": "Congrats to all. Learned a lot from this community.\n\nI am really happy to team up with my teammates @ludovick and @tonyxu who have strong engineering capabilities and are thus able to put any ideas into practice. Our final solutions consist of 4 models. Unfortunately, same as most teams, we missed selecting a sub that can beat 3rd place in Private LB. We three have totally different training processes, which I think helps us to stay gold zone on both public LB and private LB. The diversity makes us win.\n\n# Models\n- Shirok’s effb5_ns with 8xTTA\n- toxu’s vit_base_patch16_384 with 8xTTA\n- Yimin’s effb5_ns with 8xTTA\n- Yimin’s seresnext50 with no TTA\n- Same weight for ensemble.\n- Without any one of these models, we cannot win.\n- The oof corrcoef is as below:\n[[1.        , 0.79848116, 0.81904038, 0.81931724],\n       [0.79848116, 1.        , 0.80287945, 0.79185962],\n       [0.81904038, 0.80287945, 1.        , 0.82176191],\n       [0.81931724, 0.79185962, 0.82176191, 1.        ]]\n\n# Diversity\n## Teammates’ part\nAs I know, Shirok used tf TPU for training b5, and toxu used skd softlabel for vit. Welcome them to add more details at this part if would like.\n## In my part\n### Dataset\n- 2019+2020 dataset\n- Removed some duplicates searched by https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions and https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan. Thanks again for sharing useful kernels that can be used anywhere.\n- Tried Pseudo Label on 2019 test data and extraimage data (removed duplicates). Got better CV but worse LB. Probably because of data leak. I used the avg. probs of 5 folds, which may bring in leak info. between folds. Due to time limit, never chose it after.\n- Tried self knowledge distillation softlabel. Improved 0.003 on CV but decreased 0.001-0.002 on LB. Did not choose it as the final sub. -0.0015 around compared with hard label on private LB.\n- Hard Label\n- Stratified 5 folds\n- Batch size 32\n- Image size 512 x 512\n### Training AUG\nRandom Rotate\nShiftScaleRotate\nHueSaturationValue\nRandomBrightnessContrast\nFlip, transpose\nGuaussNoise\nBlur\nDistortion\nCutout\n### Loss\nBi-tempered\n### TTA\nYou must be curious about why I did not use TTA for se50 but used TTA for b5. I was stuck at 0.900 (se50), 0.902 (b5), and 0.902 (ensembles using both 8xTTA) before I checked out this [notebook](https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903). I was also curious about why Kaito did not use TTA for a model but used it for another model. Then I tried to inference my oof with and without various TTAs. And I found that TTA always decreased my cv score of se50 and increased the cv score of b5. The same situation happened on LB. Here is my previous [discussion]( https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/214559) regarding this. After using the similar strategy, I climbed up to 0.905 on LB. So, I just trusted the cv and LB at that moment, although the private LB shows the contrary result regarding se50. But our ensembles took almost 8 hrs to run submit, so it actually did not allow another model with TTA. It was the fate. Anyway, I was probably wrong with insights about TTA at my that discussion.\n\n## ensemble\nMy se50+b5 scored 0.905 on public LB, and Shirok’s b5 + toxu’s vit scored 0.906. After merge, we got 0.908. After finetuning each single model, we finally got 0.910 which stand 2nd place on public LB. Our ensemble is somehow robust. It keeps scoring 0.901-0.902 on private LB where the gold zone is. I attribute it to the diversity of our models (to be proved).\n\n# Trust CV or LB?\nUnlike the Riiid, I have totally no idea about this comp. The dataset (both train and test) is noisy, and I do not think that test data split is balanced. Our CV is range from 0.905-0.906 on 2020 dataset. It is not aligned with LB for us. Anyone can share your scores may help to determine. Anyway, trusting cv is always the better way. Make sure there is no leak in your cv.\n\n# Additional\nThe Class Activation Map (CAM) can somewhat help to track your model performance. Introduced by toxu, inspired by the winner in cvpr2020-plant-pathology. I used it for error analysis to determine if the model cares about the right position of the image. I made a simple [notebook](https://www.kaggle.com/woshifym/cld-class-activation-mapping-cam-wip) to show how it works. Hope you would enjoy it.\n\n# Others\nForgive my messy formatting.\nWill update if I think of something missing...",
    "1211564": "Congratulations! Perhaps you could change the title to indicate your place :)",
    "1211685": "Congratulations on your first gold medal and learned a lot from you.🎉🎉🎉",
    "1211786": "Same for me, i removed all duplicates to make sure of no data leak and had a ensemble CV of 0.905. I trusted CV and my best submission was my best CV however i got 0.898 in private. Congrats on your gold medal!!!",
    "1211794": "Congratulations on your 11th place! \n\nBy the way, how did you calculate the correlations?\nIn my oofs, the correlations between each models are higher than yours. \nNever dropped below 0.92 :)",
    "1211875": "Congratulations on your medal mate, rly learned a lot from you. I am new to this, i want to ask a question, since almost everybody is talking about that in this competition. I want to ask how do you make the out of fold predictions? Does that mean we make 5 fold cv or whatever number of folds, and we test the model on the validation data where the model is not being trained on? Which means we can also add TTAs and average the predictions. With that, we generate four more predictions (if k=5), and based on that we repeat the same procedure 4 more times, and we obtain the matrix that you showed above? One more question, how did you calculate the correlation to produce the matrix as shown above? \nThank you, cheers 🎉",
    "1212177": "Thanks mane. OOF means your predictions for validation data. If you split your data into 5 folds, you will predict 5 validation fold where the model is not being trained on. The OOF is the combination of this 5 folds. You can use any strategy on the OOF, like TTA or postprocess, to evaluate your model.\n\nYou can use np.corrcoef([predict1, predict2, predict3]) to calculate the correlation coefs for your different predictions.",
    "1212471": "Thanks @woshifym for the wonderful writeup and explanation. Congrats on your gold too. I got to learn a lot from you.  Just a noob question. How to use the corelation matrix that you got for ensembling. Do we weight each model by the corresponding row values rather than averaging the model predictions?",
    "1213448": "Congratulations💥💥💥"
  },
  "source": "meta"
}