{
  "id": 215607,
  "title": "Winner's Tips + 0.9 model targets",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215607",
  "author_name": "",
  "post_date": "2021-01-30T14:48:29.761037500Z",
  "votes": 42,
  "comment_count": 38,
  "views": 0,
  "content": "<p>Hello! </p>\n<p>As you know, the winner of <strong><a href=\"https://www.kaggle.com/c/cassava-disease/discussion/94114\" target=\"_blank\">2019 competition</a></strong> filtered out the training dataset part producing less than 0.95 prediction confidence. The winner of <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2020-fgvc7/discussion/154056\" target=\"_blank\">Plant Pathology 2020 Competition</a></strong> used knowledge distillation by combining soft outputs of ensemble and the ground truth. The data from both competitions is also reported to be severely imbalanced and suffers from label noise, so these techniques are definitely worth trying out.</p>\n<p>So I decided to save my 5 EfficientNetB4 ensemble (currently scoring 0.900 on LB) soft outputs to this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model\" target=\"_blank\">dataset</a></strong> to make it easier to boost new models. </p>\n<p>Hope this helps someone. Happy coding!</p>\n<h3>Updates</h3>\n<ul>\n<li>added soft targets for 2019 competition data as well (except for the duplicates)</li>\n<li>released new version with only out-of-fold predictions (a good catch by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a>)</li>\n<li>released a <strong><a href=\"https://www.kaggle.com/nickuzmenkov/tfrecords-preparation-for-6x-faster-training/\" target=\"_blank\">starter notebook</a></strong> for quick TFRecord serialization from soft targets for faster training</li>\n</ul>",
  "messages": [
    {
      "id": "1177850",
      "postDate": "01/30/2021 14:48:29",
      "content": "<p>Hello! </p>\n<p>As you know, the winner of <strong><a href=\"https://www.kaggle.com/c/cassava-disease/discussion/94114\" target=\"_blank\">2019 competition</a></strong> filtered out the training dataset part producing less than 0.95 prediction confidence. The winner of <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2020-fgvc7/discussion/154056\" target=\"_blank\">Plant Pathology 2020 Competition</a></strong> used knowledge distillation by combining soft outputs of ensemble and the ground truth. The data from both competitions is also reported to be severely imbalanced and suffers from label noise, so these techniques are definitely worth trying out.</p>\n<p>So I decided to save my 5 EfficientNetB4 ensemble (currently scoring 0.900 on LB) soft outputs to this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model\" target=\"_blank\">dataset</a></strong> to make it easier to boost new models. </p>\n<p>Hope this helps someone. Happy coding!</p>\n<h3>Updates</h3>\n<ul>\n<li>added soft targets for 2019 competition data as well (except for the duplicates)</li>\n<li>released new version with only out-of-fold predictions (a good catch by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a>)</li>\n<li>released a <strong><a href=\"https://www.kaggle.com/nickuzmenkov/tfrecords-preparation-for-6x-faster-training/\" target=\"_blank\">starter notebook</a></strong> for quick TFRecord serialization from soft targets for faster training</li>\n</ul>",
      "rawMarkdown": "Hello! \n\nAs you know, the winner of **[2019 competition](https://www.kaggle.com/c/cassava-disease/discussion/94114)** filtered out the training dataset part producing less than 0.95 prediction confidence. The winner of **[Plant Pathology 2020 Competition](https://www.kaggle.com/c/plant-pathology-2020-fgvc7/discussion/154056)** used knowledge distillation by combining soft outputs of ensemble and the ground truth. The data from both competitions is also reported to be severely imbalanced and suffers from label noise, so these techniques are definitely worth trying out.\n\nSo I decided to save my 5 EfficientNetB4 ensemble (currently scoring 0.900 on LB) soft outputs to this **[dataset](https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model)** to make it easier to boost new models. \n\nHope this helps someone. Happy coding!\n\n### Updates\n\n* added soft targets for 2019 competition data as well (except for the duplicates)\n* released new version with only out-of-fold predictions (a good catch by @ks2019)\n* released a **[starter notebook](https://www.kaggle.com/nickuzmenkov/tfrecords-preparation-for-6x-faster-training/)** for quick TFRecord serialization from soft targets for faster training",
      "votes": null
    },
    {
      "id": "1177884",
      "postDate": "01/30/2021 15:05:50",
      "content": "<p>thanks for you sharing for your knowledge to get 0.9 model targets and this is very useful for me. good luck to you🙏</p>",
      "rawMarkdown": "thanks for you sharing for your knowledge to get 0.9 model targets and this is very useful for me. good luck to you🙏",
      "votes": null
    },
    {
      "id": "1178054",
      "postDate": "01/30/2021 16:20:17",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> Thanks for sharing this excellent idea. Best of luck.</p>",
      "rawMarkdown": "nickuzmenkov Thanks for sharing this excellent idea. Best of luck.",
      "votes": null
    },
    {
      "id": "1178081",
      "postDate": "01/30/2021 16:35:33",
      "content": "<p>I took the time to visually look at all images from the current competition and from the training part of 2019. (very fast look searching for roots)</p>\n<p>They both may be noisy, but the 2020 images are much worse.  For example, there are almost 100 pictures of roots in 2020 and I think only one in 2019.  2020 contains lots of images with humans in the photo - don't think I saw a single one in 2019.  2020 contains a few images were the leaves are themselves a printed photo.  2020 contains lots of close-ups of stalks of the plant and don't think I saw many in 2019.  The 2019 images look to be mostly at around the same distance from the plant while 2020 ranges from closeups to scenic distances.</p>",
      "rawMarkdown": "I took the time to visually look at all images from the current competition and from the training part of 2019. (very fast look searching for roots)\n\nThey both may be noisy, but the 2020 images are much worse.  For example, there are almost 100 pictures of roots in 2020 and I think only one in 2019.  2020 contains lots of images with humans in the photo - don't think I saw a single one in 2019.  2020 contains a few images were the leaves are themselves a printed photo.  2020 contains lots of close-ups of stalks of the plant and don't think I saw many in 2019.  The 2019 images look to be mostly at around the same distance from the plant while 2020 ranges from closeups to scenic distances.",
      "votes": null
    },
    {
      "id": "1179065",
      "postDate": "01/31/2021 09:38:40",
      "content": "<p>So this competition is spicier ^^</p>",
      "rawMarkdown": "So this competition is spicier ^^",
      "votes": null
    },
    {
      "id": "1179126",
      "postDate": "01/31/2021 10:45:11",
      "content": "<p>Yes - just finished developing a very accurate model for roots - now that I got it not sure how to use it :)</p>",
      "rawMarkdown": "Yes - just finished developing a very accurate model for roots - now that I got it not sure how to use it :)",
      "votes": null
    },
    {
      "id": "1181527",
      "postDate": "02/02/2021 00:40:10",
      "content": "<p>Thank you for sharing !!!!!</p>",
      "rawMarkdown": "Thank you for sharing !!!!!",
      "votes": null
    },
    {
      "id": "1182811",
      "postDate": "02/02/2021 15:54:02",
      "content": "<p>If there are enough root images you can try, 1. train a model to classify roots and leafs, 2. train model to classify root images to respective classes, 3. train another model that classifies leaf images. All this only if root images are enough for each class. </p>",
      "rawMarkdown": "If there are enough root images you can try, 1. train a model to classify roots and leafs, 2. train model to classify root images to respective classes, 3. train another model that classifies leaf images. All this only if root images are enough for each class.",
      "votes": null
    },
    {
      "id": "1183892",
      "postDate": "02/03/2021 09:39:35",
      "content": "<p>Thanks for sharing mate!<br>\nwould you mind revealing whether you tried it and noticed any boost?<br>\nThanks</p>",
      "rawMarkdown": "Thanks for sharing mate!\nwould you mind revealing whether you tried it and noticed any boost?\nThanks",
      "votes": null
    },
    {
      "id": "1183993",
      "postDate": "02/03/2021 10:42:32",
      "content": "<p>Hello!</p>\n<p>I tried knowledge distillation by combining these soft outputs and the ground truth, which gave me a +0.002 LB boost, and suppose there's still plenty of room to improve.</p>",
      "rawMarkdown": "Hello!\n\nI tried knowledge distillation by combining these soft outputs and the ground truth, which gave me a +0.002 LB boost, and suppose there's still plenty of room to improve.",
      "votes": null
    },
    {
      "id": "1184020",
      "postDate": "02/03/2021 11:02:24",
      "content": "<p>Wow thanks for sharing it!</p>",
      "rawMarkdown": "Wow thanks for sharing it!",
      "votes": null
    },
    {
      "id": "1184029",
      "postDate": "02/03/2021 11:07:26",
      "content": "<p>what are the root images classified as in the dataset?</p>",
      "rawMarkdown": "what are the root images classified as in the dataset?",
      "votes": null
    },
    {
      "id": "1184669",
      "postDate": "02/03/2021 17:00:17",
      "content": "<p>Great! Thanks for sharing. May I ask if using the soft label, how you build your cv? i.e., using soft label or orginal label?</p>",
      "rawMarkdown": "Great! Thanks for sharing. May I ask if using the soft label, how you build your cv? i.e., using soft label or orginal label?",
      "votes": null
    },
    {
      "id": "1184792",
      "postDate": "02/03/2021 18:45:38",
      "content": "<p>Hello!</p>\n<p>Still using original labels. </p>",
      "rawMarkdown": "Hello!\n\nStill using original labels.",
      "votes": null
    },
    {
      "id": "1185163",
      "postDate": "02/04/2021 02:13:45",
      "content": "<p>Sure, thanks for your kind reply.</p>",
      "rawMarkdown": "Sure, thanks for your kind reply.",
      "votes": null
    },
    {
      "id": "1185281",
      "postDate": "02/04/2021 05:12:12",
      "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> most of them belong to class 1. Though I saw some in class 3 as well. I dont think I saw any root image in other classes.</p>",
      "rawMarkdown": "alexanderriedel most of them belong to class 1. Though I saw some in class 3 as well. I dont think I saw any root image in other classes.",
      "votes": null
    },
    {
      "id": "1185405",
      "postDate": "02/04/2021 06:34:32",
      "content": "<p>I added this dataset to the competition under Datasets --&gt; Add Data so it might be easy for people to find there in case they miss this discussion.  Still your link and same as under main Datasets.  Thanks for making it available! </p>",
      "rawMarkdown": "I added this dataset to the competition under Datasets --> Add Data so it might be easy for people to find there in case they miss this discussion.  Still your link and same as under main Datasets.  Thanks for making it available!",
      "votes": null
    },
    {
      "id": "1185500",
      "postDate": "02/04/2021 07:19:05",
      "content": "<p>Thank you</p>\n<p>I'm relatively new to Kaggle, so I didn't know about that option.</p>",
      "rawMarkdown": "Thank you\n\nI'm relatively new to Kaggle, so I didn't know about that option.",
      "votes": null
    },
    {
      "id": "1185544",
      "postDate": "02/04/2021 07:51:37",
      "content": "<p>I only realised it existed today!!  Not sure if it is something new with the rollout of new look datasets, but it is a nice feature to find competition related work easily.  </p>",
      "rawMarkdown": "I only realised it existed today!!  Not sure if it is something new with the rollout of new look datasets, but it is a nice feature to find competition related work easily.",
      "votes": null
    },
    {
      "id": "1186724",
      "postDate": "02/05/2021 02:32:47",
      "content": "<p>Thanks for great strategy<br>\nI'm quiet newbie for Knowledge distillation, can i ask you a question?</p>\n<p>in case we have soft label and model ouputs<br>\nex) soft label : [0.1, 0.1, 0.2, 0.5, 0.1] <br>\nex) model outputs : [0.01, 0.01, 0.01, 0.9, 0.07]<br>\nhow can I calculate loss function between them </p>\n<p>I've read paper for knowledge distillation, and it says<br>\nloss function for knowledge distillation =  <br>\nlambda * ( loss(student model_outputs, label)) + (1-lambda) * (loss(student model_outputs, teacher_model_outputs))</p>\n<p>but i don't know how to calculate loss between soft label and model outputs,<br>\nsoft label should be logits ? or softmaxed logits?</p>",
      "rawMarkdown": "Thanks for great strategy\nI'm quiet newbie for Knowledge distillation, can i ask you a question?\n\nin case we have soft label and model ouputs\nex) soft label : [0.1, 0.1, 0.2, 0.5, 0.1] \nex) model outputs : [0.01, 0.01, 0.01, 0.9, 0.07]\nhow can I calculate loss function between them \n\nI've read paper for knowledge distillation, and it says\nloss function for knowledge distillation =  \nlambda * ( loss(student model_outputs, label)) + (1-lambda) * (loss(student model_outputs, teacher_model_outputs))\n\nbut i don't know how to calculate loss between soft label and model outputs,\nsoft label should be logits ? or softmaxed logits?",
      "votes": null
    },
    {
      "id": "1186955",
      "postDate": "02/05/2021 06:06:20",
      "content": "<p>Hello! </p>\n<p>I've seen this paper, but Keras's implementation may be tricky. However, if you feel enthusiastic, there's an even more detailed description on distillation in <strong><a href=\"https://arxiv.org/abs/1503.02531\" target=\"_blank\">this paper</a></strong>. </p>\n<p>One simpler way to implement this is to calculate the weighted sum between the ground truth and soft labels:<br>\n<code>new_labels = w1 * ground_truth + w2 * soft_labels</code><br>\nwhere w1 / w2 &gt;= 1 (so that <code>accuracy</code> metric is still calculated based on the ground truth, not the teacher's predictions). The loss is then calculated just as usual, i.e. between the softmax logits and the new labels.</p>\n<p>I'm also not so familiar with knowledge distillation, so if someone experienced finds the above description inaccurate, please, correct me.</p>",
      "rawMarkdown": "Hello! \n\nI've seen this paper, but Keras's implementation may be tricky. However, if you feel enthusiastic, there's an even more detailed description on distillation in **[this paper](https://arxiv.org/abs/1503.02531)**. \n\nOne simpler way to implement this is to calculate the weighted sum between the ground truth and soft labels:\n`new_labels = w1 * ground_truth + w2 * soft_labels`\nwhere w1 / w2 >= 1 (so that `accuracy` metric is still calculated based on the ground truth, not the teacher's predictions). The loss is then calculated just as usual, i.e. between the softmax logits and the new labels.\n\nI'm also not so familiar with knowledge distillation, so if someone experienced finds the above description inaccurate, please, correct me.",
      "votes": null
    },
    {
      "id": "1196633",
      "postDate": "02/11/2021 14:17:48",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> for sharing! I tried your soft labels. Indeed, they gave me a +0.003 boost with small models scoring 0.88x (from 0.887 to 0.890). However, they don't seem to work with stronger models scoring +0.9!</p>\n<ul>\n<li>Did you get the +0.002 boost with models scoring around 0.9 (more or less)?</li>\n<li>Did you also use 0.7:0.3 for distillation?</li>\n</ul>",
      "rawMarkdown": "Thank you @nickuzmenkov for sharing! I tried your soft labels. Indeed, they gave me a +0.003 boost with small models scoring 0.88x (from 0.887 to 0.890). However, they don't seem to work with stronger models scoring +0.9!\n* Did you get the +0.002 boost with models scoring around 0.9 (more or less)?\n* Did you also use 0.7:0.3 for distillation?",
      "votes": null
    },
    {
      "id": "1197519",
      "postDate": "02/12/2021 07:55:35",
      "content": "<p>Hello! </p>\n<ul>\n<li>When creating this topic, my highest LB score was already 0.900 (with CV a bit lower). Retraining with this labels gave +0.002 LB boost (and some little tweaks which are better to stay secret for now gave +0.004 LB boost). However, secondary distillation on soft labels of these distilled models gave the same results even on larger models. </li>\n<li>I was focused on other things, keeping distillation weights simple 1:1. Unfortunately, I cannot experiment with them for now, as I'm still awaiting my TPU quota to renew.</li>\n</ul>",
      "rawMarkdown": "Hello! \n\n* When creating this topic, my highest LB score was already 0.900 (with CV a bit lower). Retraining with this labels gave +0.002 LB boost (and some little tweaks which are better to stay secret for now gave +0.004 LB boost). However, secondary distillation on soft labels of these distilled models gave the same results even on larger models. \n* I was focused on other things, keeping distillation weights simple 1:1. Unfortunately, I cannot experiment with them for now, as I'm still awaiting my TPU quota to renew.",
      "votes": null
    },
    {
      "id": "1197964",
      "postDate": "02/12/2021 15:09:50",
      "content": "<p>I see your accuracy for 2020 dataset is 0.916 !! Did you predict on both train and val set for each of the model?</p>",
      "rawMarkdown": "I see your accuracy for 2020 dataset is 0.916 !! Did you predict on both train and val set for each of the model?",
      "votes": null
    },
    {
      "id": "1197985",
      "postDate": "02/12/2021 15:35:05",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a>  hi, Did you use TF or Pytorch? Can you share an example kernel to use?</p>",
      "rawMarkdown": "nickuzmenkov  hi, Did you use TF or Pytorch? Can you share an example kernel to use?",
      "votes": null
    },
    {
      "id": "1198648",
      "postDate": "02/13/2021 07:22:12",
      "content": "<p>Hello!</p>\n<p>Yes, with 8 light TTA.</p>",
      "rawMarkdown": "Hello!\n\nYes, with 8 light TTA.",
      "votes": null
    },
    {
      "id": "1198656",
      "postDate": "02/13/2021 07:41:35",
      "content": "<p>Winning solution of Plant Pathology used only out of fold predictions for knowledge distillation. Using predictions on train set can cause your model to overfit. </p>",
      "rawMarkdown": "Winning solution of Plant Pathology used only out of fold predictions for knowledge distillation. Using predictions on train set can cause your model to overfit.",
      "votes": null
    },
    {
      "id": "1198669",
      "postDate": "02/13/2021 07:50:47",
      "content": "<p>Thank you for the warning! It seems like I've completely mixed up knowledge distillation and noisy student training concepts. Fortunately, there's still time to fix it up.</p>",
      "rawMarkdown": "Thank you for the warning! It seems like I've completely mixed up knowledge distillation and noisy student training concepts. Fortunately, there's still time to fix it up.",
      "votes": null
    },
    {
      "id": "1198690",
      "postDate": "02/13/2021 08:02:31",
      "content": "<p>Yes, best of luck !! I have a question though (if it's not a secret)!! How did you remove duplicates? I am using image dedup for duplicate detection and with that, I am getting 620 duplicates in 2019 dataset!! But I see your methods revealed 100 more duplicates. </p>",
      "rawMarkdown": "Yes, best of luck !! I have a question though (if it's not a secret)!! How did you remove duplicates? I am using image dedup for duplicate detection and with that, I am getting 620 duplicates in 2019 dataset!! But I see your methods revealed 100 more duplicates.",
      "votes": null
    },
    {
      "id": "1198718",
      "postDate": "02/13/2021 08:11:39",
      "content": "<p>Actually, this work is done by <a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a>, not me. I took images from <strong><a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">this dataset</a></strong>.</p>",
      "rawMarkdown": "Actually, this work is done by @tahsin, not me. I took images from **[this dataset](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged)**.",
      "votes": null
    },
    {
      "id": "1201778",
      "postDate": "02/15/2021 16:44:54",
      "content": "<p>Hey Nick, I want to ask, did the merged dataset gave a little boost, is it worth to try? If so, did you evaluate your model on the validation data from the merged dataset, or evaluated on the validation data from this competition, while training on combined dataset?<br>\nCheers.</p>",
      "rawMarkdown": "Hey Nick, I want to ask, did the merged dataset gave a little boost, is it worth to try? If so, did you evaluate your model on the validation data from the merged dataset, or evaluated on the validation data from this competition, while training on combined dataset?\nCheers.",
      "votes": null
    },
    {
      "id": "1201795",
      "postDate": "02/15/2021 17:01:47",
      "content": "<p>Hello!</p>\n<p>Sure, using this merged data gave me a little boost as well. Nevertheless, the 2019 data is from literally the same competition, I validated only on the data of the present competition to produce a trustworthy score. </p>\n<p>Another reason for that is the relative cleanliness of the 2019 data, which is reported by many users, and my results also prove that: the ensemble I've used for making predictions is scoring 0.916 CV on 2019 data while only 0.899 CV on the present data.</p>",
      "rawMarkdown": "Hello!\n\nSure, using this merged data gave me a little boost as well. Nevertheless, the 2019 data is from literally the same competition, I validated only on the data of the present competition to produce a trustworthy score. \n\nAnother reason for that is the relative cleanliness of the 2019 data, which is reported by many users, and my results also prove that: the ensemble I've used for making predictions is scoring 0.916 CV on 2019 data while only 0.899 CV on the present data.",
      "votes": null
    },
    {
      "id": "1201811",
      "postDate": "02/15/2021 17:10:39",
      "content": "<p>Hello!</p>\n<p>The thing is, I'm just a beginner in ML, having only a little practical and theoretical background. So I suppose sharing my kernel won't uncover any new ideas, but the ideas of more experienced users.</p>\n<p>The most inspiring work I took a lot from was <strong><a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">this notebook</a></strong> by <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>.</p>",
      "rawMarkdown": "Hello!\n\nThe thing is, I'm just a beginner in ML, having only a little practical and theoretical background. So I suppose sharing my kernel won't uncover any new ideas, but the ideas of more experienced users.\n\nThe most inspiring work I took a lot from was **[this notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods)** by @dimitreoliveira.",
      "votes": null
    },
    {
      "id": "1201858",
      "postDate": "02/15/2021 17:39:48",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a>  should we use max of soft labels ?</p>",
      "rawMarkdown": "nickuzmenkov  should we use max of soft labels ?",
      "votes": null
    },
    {
      "id": "1201877",
      "postDate": "02/15/2021 17:52:58",
      "content": "<p>Hello! </p>\n<p>The idea behind this is to use a combination of soft labels the ground truth, rather than soft labels standalone, thus paying attention to those heavily misclassed samples (which are probably the noisy ones) while still training on the original labels.</p>",
      "rawMarkdown": "Hello! \n\nThe idea behind this is to use a combination of soft labels the ground truth, rather than soft labels standalone, thus paying attention to those heavily misclassed samples (which are probably the noisy ones) while still training on the original labels.",
      "votes": null
    },
    {
      "id": "1201880",
      "postDate": "02/15/2021 17:57:32",
      "content": "<p>Any eg of label  That you can give  pls</p>",
      "rawMarkdown": "Any eg of label  That you can give  pls",
      "votes": null
    },
    {
      "id": "1201882",
      "postDate": "02/15/2021 18:02:38",
      "content": "<p>Sorry, I don't understand. Could you please paraphrase that?</p>",
      "rawMarkdown": "Sorry, I don't understand. Could you please paraphrase that?",
      "votes": null
    },
    {
      "id": "1201905",
      "postDate": "02/15/2021 18:25:24",
      "content": "<p>Nice to hear, my four models which I ensemble are scoring 0.908, 0.906, 0.904 and 0.902 (CV), while the average of the produce 0.904 (LB). <br>\nI am satisfied for first competition, I came here to learn in the first place, if i get medal that would be bonus. <br>\nThanks and good luck :) !</p>",
      "rawMarkdown": "Nice to hear, my four models which I ensemble are scoring 0.908, 0.906, 0.904 and 0.902 (CV), while the average of the produce 0.904 (LB). \nI am satisfied for first competition, I came here to learn in the first place, if i get medal that would be bonus. \nThanks and good luck :) !",
      "votes": null
    },
    {
      "id": "1201910",
      "postDate": "02/15/2021 18:29:35",
      "content": "<p>Check the plant pathology's 1st place solution linked above, you should one hot encode your labels and blend them with the soft labels</p>\n<pre><code>if self.soft_labels is not None:\n            label = torch.FloatTensor(\n                (self.data.iloc[index, 1:].values * 0.7).astype(np.float16)\n                + (self.soft_labels.iloc[index, 1:].values * 0.3).astype(np.float16)\n</code></pre>\n<p>And make sure to tweak your loss function to accept one hot encoded labels.</p>",
      "rawMarkdown": "Check the plant pathology's 1st place solution linked above, you should one hot encode your labels and blend them with the soft labels\n```\nif self.soft_labels is not None:\n            label = torch.FloatTensor(\n                (self.data.iloc[index, 1:].values * 0.7).astype(np.float16)\n                + (self.soft_labels.iloc[index, 1:].values * 0.3).astype(np.float16)\n```\nAnd make sure to tweak your loss function to accept one hot encoded labels.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1177884,
      "author_name": "fauzanalfariz",
      "author_url": "",
      "post_date": "01/30/2021 15:05:50",
      "content": "<p>thanks for you sharing for your knowledge to get 0.9 model targets and this is very useful for me. good luck to you🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1178054,
      "author_name": "durbin164",
      "author_url": "",
      "post_date": "01/30/2021 16:20:17",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> Thanks for sharing this excellent idea. Best of luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1178081,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/30/2021 16:35:33",
      "content": "<p>I took the time to visually look at all images from the current competition and from the training part of 2019. (very fast look searching for roots)</p>\n<p>They both may be noisy, but the 2020 images are much worse.  For example, there are almost 100 pictures of roots in 2020 and I think only one in 2019.  2020 contains lots of images with humans in the photo - don't think I saw a single one in 2019.  2020 contains a few images were the leaves are themselves a printed photo.  2020 contains lots of close-ups of stalks of the plant and don't think I saw many in 2019.  The 2019 images look to be mostly at around the same distance from the plant while 2020 ranges from closeups to scenic distances.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1179065,
          "author_name": "killimi",
          "author_url": "",
          "post_date": "01/31/2021 09:38:40",
          "content": "<p>So this competition is spicier ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1179126,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/31/2021 10:45:11",
          "content": "<p>Yes - just finished developing a very accurate model for roots - now that I got it not sure how to use it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1182811,
          "author_name": "anku5hk",
          "author_url": "",
          "post_date": "02/02/2021 15:54:02",
          "content": "<p>If there are enough root images you can try, 1. train a model to classify roots and leafs, 2. train model to classify root images to respective classes, 3. train another model that classifies leaf images. All this only if root images are enough for each class. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1184029,
          "author_name": "alexanderriedel",
          "author_url": "",
          "post_date": "02/03/2021 11:07:26",
          "content": "<p>what are the root images classified as in the dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1185281,
          "author_name": "hammaadali",
          "author_url": "",
          "post_date": "02/04/2021 05:12:12",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> most of them belong to class 1. Though I saw some in class 3 as well. I dont think I saw any root image in other classes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1181527,
      "author_name": "nistik03",
      "author_url": "",
      "post_date": "02/02/2021 00:40:10",
      "content": "<p>Thank you for sharing !!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1183892,
      "author_name": "saimanojakondi",
      "author_url": "",
      "post_date": "02/03/2021 09:39:35",
      "content": "<p>Thanks for sharing mate!<br>\nwould you mind revealing whether you tried it and noticed any boost?<br>\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1183993,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/03/2021 10:42:32",
          "content": "<p>Hello!</p>\n<p>I tried knowledge distillation by combining these soft outputs and the ground truth, which gave me a +0.002 LB boost, and suppose there's still plenty of room to improve.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1184020,
          "author_name": "saimanojakondi",
          "author_url": "",
          "post_date": "02/03/2021 11:02:24",
          "content": "<p>Wow thanks for sharing it!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1184669,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "02/03/2021 17:00:17",
          "content": "<p>Great! Thanks for sharing. May I ask if using the soft label, how you build your cv? i.e., using soft label or orginal label?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1184792,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/03/2021 18:45:38",
          "content": "<p>Hello!</p>\n<p>Still using original labels. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1185163,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "02/04/2021 02:13:45",
          "content": "<p>Sure, thanks for your kind reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1185405,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "02/04/2021 06:34:32",
      "content": "<p>I added this dataset to the competition under Datasets --&gt; Add Data so it might be easy for people to find there in case they miss this discussion.  Still your link and same as under main Datasets.  Thanks for making it available! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1185500,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/04/2021 07:19:05",
          "content": "<p>Thank you</p>\n<p>I'm relatively new to Kaggle, so I didn't know about that option.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1185544,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "02/04/2021 07:51:37",
          "content": "<p>I only realised it existed today!!  Not sure if it is something new with the rollout of new look datasets, but it is a nice feature to find competition related work easily.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1186724,
      "author_name": "deepbluebird",
      "author_url": "",
      "post_date": "02/05/2021 02:32:47",
      "content": "<p>Thanks for great strategy<br>\nI'm quiet newbie for Knowledge distillation, can i ask you a question?</p>\n<p>in case we have soft label and model ouputs<br>\nex) soft label : [0.1, 0.1, 0.2, 0.5, 0.1] <br>\nex) model outputs : [0.01, 0.01, 0.01, 0.9, 0.07]<br>\nhow can I calculate loss function between them </p>\n<p>I've read paper for knowledge distillation, and it says<br>\nloss function for knowledge distillation =  <br>\nlambda * ( loss(student model_outputs, label)) + (1-lambda) * (loss(student model_outputs, teacher_model_outputs))</p>\n<p>but i don't know how to calculate loss between soft label and model outputs,<br>\nsoft label should be logits ? or softmaxed logits?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186955,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/05/2021 06:06:20",
          "content": "<p>Hello! </p>\n<p>I've seen this paper, but Keras's implementation may be tricky. However, if you feel enthusiastic, there's an even more detailed description on distillation in <strong><a href=\"https://arxiv.org/abs/1503.02531\" target=\"_blank\">this paper</a></strong>. </p>\n<p>One simpler way to implement this is to calculate the weighted sum between the ground truth and soft labels:<br>\n<code>new_labels = w1 * ground_truth + w2 * soft_labels</code><br>\nwhere w1 / w2 &gt;= 1 (so that <code>accuracy</code> metric is still calculated based on the ground truth, not the teacher's predictions). The loss is then calculated just as usual, i.e. between the softmax logits and the new labels.</p>\n<p>I'm also not so familiar with knowledge distillation, so if someone experienced finds the above description inaccurate, please, correct me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1196633,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "02/11/2021 14:17:48",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> for sharing! I tried your soft labels. Indeed, they gave me a +0.003 boost with small models scoring 0.88x (from 0.887 to 0.890). However, they don't seem to work with stronger models scoring +0.9!</p>\n<ul>\n<li>Did you get the +0.002 boost with models scoring around 0.9 (more or less)?</li>\n<li>Did you also use 0.7:0.3 for distillation?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1197519,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/12/2021 07:55:35",
          "content": "<p>Hello! </p>\n<ul>\n<li>When creating this topic, my highest LB score was already 0.900 (with CV a bit lower). Retraining with this labels gave +0.002 LB boost (and some little tweaks which are better to stay secret for now gave +0.004 LB boost). However, secondary distillation on soft labels of these distilled models gave the same results even on larger models. </li>\n<li>I was focused on other things, keeping distillation weights simple 1:1. Unfortunately, I cannot experiment with them for now, as I'm still awaiting my TPU quota to renew.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1197964,
      "author_name": "ks2019",
      "author_url": "",
      "post_date": "02/12/2021 15:09:50",
      "content": "<p>I see your accuracy for 2020 dataset is 0.916 !! Did you predict on both train and val set for each of the model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1198648,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/13/2021 07:22:12",
          "content": "<p>Hello!</p>\n<p>Yes, with 8 light TTA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1198656,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "02/13/2021 07:41:35",
          "content": "<p>Winning solution of Plant Pathology used only out of fold predictions for knowledge distillation. Using predictions on train set can cause your model to overfit. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1198669,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/13/2021 07:50:47",
          "content": "<p>Thank you for the warning! It seems like I've completely mixed up knowledge distillation and noisy student training concepts. Fortunately, there's still time to fix it up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1198690,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "02/13/2021 08:02:31",
          "content": "<p>Yes, best of luck !! I have a question though (if it's not a secret)!! How did you remove duplicates? I am using image dedup for duplicate detection and with that, I am getting 620 duplicates in 2019 dataset!! But I see your methods revealed 100 more duplicates. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1198718,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/13/2021 08:11:39",
          "content": "<p>Actually, this work is done by <a href=\"https://www.kaggle.com/tahsin\" target=\"_blank\">@tahsin</a>, not me. I took images from <strong><a href=\"https://www.kaggle.com/tahsin/cassava-leaf-disease-merged\" target=\"_blank\">this dataset</a></strong>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201778,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/15/2021 16:44:54",
          "content": "<p>Hey Nick, I want to ask, did the merged dataset gave a little boost, is it worth to try? If so, did you evaluate your model on the validation data from the merged dataset, or evaluated on the validation data from this competition, while training on combined dataset?<br>\nCheers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201795,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/15/2021 17:01:47",
          "content": "<p>Hello!</p>\n<p>Sure, using this merged data gave me a little boost as well. Nevertheless, the 2019 data is from literally the same competition, I validated only on the data of the present competition to produce a trustworthy score. </p>\n<p>Another reason for that is the relative cleanliness of the 2019 data, which is reported by many users, and my results also prove that: the ensemble I've used for making predictions is scoring 0.916 CV on 2019 data while only 0.899 CV on the present data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201905,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/15/2021 18:25:24",
          "content": "<p>Nice to hear, my four models which I ensemble are scoring 0.908, 0.906, 0.904 and 0.902 (CV), while the average of the produce 0.904 (LB). <br>\nI am satisfied for first competition, I came here to learn in the first place, if i get medal that would be bonus. <br>\nThanks and good luck :) !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1197985,
      "author_name": "durbin164",
      "author_url": "",
      "post_date": "02/12/2021 15:35:05",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a>  hi, Did you use TF or Pytorch? Can you share an example kernel to use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1201811,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/15/2021 17:10:39",
          "content": "<p>Hello!</p>\n<p>The thing is, I'm just a beginner in ML, having only a little practical and theoretical background. So I suppose sharing my kernel won't uncover any new ideas, but the ideas of more experienced users.</p>\n<p>The most inspiring work I took a lot from was <strong><a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">this notebook</a></strong> by <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1201858,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "02/15/2021 17:39:48",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a>  should we use max of soft labels ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1201877,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/15/2021 17:52:58",
          "content": "<p>Hello! </p>\n<p>The idea behind this is to use a combination of soft labels the ground truth, rather than soft labels standalone, thus paying attention to those heavily misclassed samples (which are probably the noisy ones) while still training on the original labels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201880,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/15/2021 17:57:32",
          "content": "<p>Any eg of label  That you can give  pls</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201882,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/15/2021 18:02:38",
          "content": "<p>Sorry, I don't understand. Could you please paraphrase that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1201910,
          "author_name": "amiiiney",
          "author_url": "",
          "post_date": "02/15/2021 18:29:35",
          "content": "<p>Check the plant pathology's 1st place solution linked above, you should one hot encode your labels and blend them with the soft labels</p>\n<pre><code>if self.soft_labels is not None:\n            label = torch.FloatTensor(\n                (self.data.iloc[index, 1:].values * 0.7).astype(np.float16)\n                + (self.soft_labels.iloc[index, 1:].values * 0.3).astype(np.float16)\n</code></pre>\n<p>And make sure to tweak your loss function to accept one hot encoded labels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1177850": "Hello! \n\nAs you know, the winner of **[2019 competition](https://www.kaggle.com/c/cassava-disease/discussion/94114)** filtered out the training dataset part producing less than 0.95 prediction confidence. The winner of **[Plant Pathology 2020 Competition](https://www.kaggle.com/c/plant-pathology-2020-fgvc7/discussion/154056)** used knowledge distillation by combining soft outputs of ensemble and the ground truth. The data from both competitions is also reported to be severely imbalanced and suffers from label noise, so these techniques are definitely worth trying out.\n\nSo I decided to save my 5 EfficientNetB4 ensemble (currently scoring 0.900 on LB) soft outputs to this **[dataset](https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-soft-targets-09-model)** to make it easier to boost new models. \n\nHope this helps someone. Happy coding!\n\n### Updates\n\n* added soft targets for 2019 competition data as well (except for the duplicates)\n* released new version with only out-of-fold predictions (a good catch by @ks2019)\n* released a **[starter notebook](https://www.kaggle.com/nickuzmenkov/tfrecords-preparation-for-6x-faster-training/)** for quick TFRecord serialization from soft targets for faster training",
    "1177884": "thanks for you sharing for your knowledge to get 0.9 model targets and this is very useful for me. good luck to you🙏",
    "1178054": "nickuzmenkov Thanks for sharing this excellent idea. Best of luck.",
    "1178081": "I took the time to visually look at all images from the current competition and from the training part of 2019. (very fast look searching for roots)\n\nThey both may be noisy, but the 2020 images are much worse.  For example, there are almost 100 pictures of roots in 2020 and I think only one in 2019.  2020 contains lots of images with humans in the photo - don't think I saw a single one in 2019.  2020 contains a few images were the leaves are themselves a printed photo.  2020 contains lots of close-ups of stalks of the plant and don't think I saw many in 2019.  The 2019 images look to be mostly at around the same distance from the plant while 2020 ranges from closeups to scenic distances.",
    "1179065": "So this competition is spicier ^^",
    "1179126": "Yes - just finished developing a very accurate model for roots - now that I got it not sure how to use it :)",
    "1181527": "Thank you for sharing !!!!!",
    "1182811": "If there are enough root images you can try, 1. train a model to classify roots and leafs, 2. train model to classify root images to respective classes, 3. train another model that classifies leaf images. All this only if root images are enough for each class.",
    "1183892": "Thanks for sharing mate!\nwould you mind revealing whether you tried it and noticed any boost?\nThanks",
    "1183993": "Hello!\n\nI tried knowledge distillation by combining these soft outputs and the ground truth, which gave me a +0.002 LB boost, and suppose there's still plenty of room to improve.",
    "1184020": "Wow thanks for sharing it!",
    "1184029": "what are the root images classified as in the dataset?",
    "1184669": "Great! Thanks for sharing. May I ask if using the soft label, how you build your cv? i.e., using soft label or orginal label?",
    "1184792": "Hello!\n\nStill using original labels.",
    "1185163": "Sure, thanks for your kind reply.",
    "1185281": "alexanderriedel most of them belong to class 1. Though I saw some in class 3 as well. I dont think I saw any root image in other classes.",
    "1185405": "I added this dataset to the competition under Datasets --> Add Data so it might be easy for people to find there in case they miss this discussion.  Still your link and same as under main Datasets.  Thanks for making it available!",
    "1185500": "Thank you\n\nI'm relatively new to Kaggle, so I didn't know about that option.",
    "1185544": "I only realised it existed today!!  Not sure if it is something new with the rollout of new look datasets, but it is a nice feature to find competition related work easily.",
    "1186724": "Thanks for great strategy\nI'm quiet newbie for Knowledge distillation, can i ask you a question?\n\nin case we have soft label and model ouputs\nex) soft label : [0.1, 0.1, 0.2, 0.5, 0.1] \nex) model outputs : [0.01, 0.01, 0.01, 0.9, 0.07]\nhow can I calculate loss function between them \n\nI've read paper for knowledge distillation, and it says\nloss function for knowledge distillation =  \nlambda * ( loss(student model_outputs, label)) + (1-lambda) * (loss(student model_outputs, teacher_model_outputs))\n\nbut i don't know how to calculate loss between soft label and model outputs,\nsoft label should be logits ? or softmaxed logits?",
    "1186955": "Hello! \n\nI've seen this paper, but Keras's implementation may be tricky. However, if you feel enthusiastic, there's an even more detailed description on distillation in **[this paper](https://arxiv.org/abs/1503.02531)**. \n\nOne simpler way to implement this is to calculate the weighted sum between the ground truth and soft labels:\n`new_labels = w1 * ground_truth + w2 * soft_labels`\nwhere w1 / w2 >= 1 (so that `accuracy` metric is still calculated based on the ground truth, not the teacher's predictions). The loss is then calculated just as usual, i.e. between the softmax logits and the new labels.\n\nI'm also not so familiar with knowledge distillation, so if someone experienced finds the above description inaccurate, please, correct me.",
    "1196633": "Thank you @nickuzmenkov for sharing! I tried your soft labels. Indeed, they gave me a +0.003 boost with small models scoring 0.88x (from 0.887 to 0.890). However, they don't seem to work with stronger models scoring +0.9!\n* Did you get the +0.002 boost with models scoring around 0.9 (more or less)?\n* Did you also use 0.7:0.3 for distillation?",
    "1197519": "Hello! \n\n* When creating this topic, my highest LB score was already 0.900 (with CV a bit lower). Retraining with this labels gave +0.002 LB boost (and some little tweaks which are better to stay secret for now gave +0.004 LB boost). However, secondary distillation on soft labels of these distilled models gave the same results even on larger models. \n* I was focused on other things, keeping distillation weights simple 1:1. Unfortunately, I cannot experiment with them for now, as I'm still awaiting my TPU quota to renew.",
    "1197964": "I see your accuracy for 2020 dataset is 0.916 !! Did you predict on both train and val set for each of the model?",
    "1197985": "nickuzmenkov  hi, Did you use TF or Pytorch? Can you share an example kernel to use?",
    "1198648": "Hello!\n\nYes, with 8 light TTA.",
    "1198656": "Winning solution of Plant Pathology used only out of fold predictions for knowledge distillation. Using predictions on train set can cause your model to overfit.",
    "1198669": "Thank you for the warning! It seems like I've completely mixed up knowledge distillation and noisy student training concepts. Fortunately, there's still time to fix it up.",
    "1198690": "Yes, best of luck !! I have a question though (if it's not a secret)!! How did you remove duplicates? I am using image dedup for duplicate detection and with that, I am getting 620 duplicates in 2019 dataset!! But I see your methods revealed 100 more duplicates.",
    "1198718": "Actually, this work is done by @tahsin, not me. I took images from **[this dataset](https://www.kaggle.com/tahsin/cassava-leaf-disease-merged)**.",
    "1201778": "Hey Nick, I want to ask, did the merged dataset gave a little boost, is it worth to try? If so, did you evaluate your model on the validation data from the merged dataset, or evaluated on the validation data from this competition, while training on combined dataset?\nCheers.",
    "1201795": "Hello!\n\nSure, using this merged data gave me a little boost as well. Nevertheless, the 2019 data is from literally the same competition, I validated only on the data of the present competition to produce a trustworthy score. \n\nAnother reason for that is the relative cleanliness of the 2019 data, which is reported by many users, and my results also prove that: the ensemble I've used for making predictions is scoring 0.916 CV on 2019 data while only 0.899 CV on the present data.",
    "1201811": "Hello!\n\nThe thing is, I'm just a beginner in ML, having only a little practical and theoretical background. So I suppose sharing my kernel won't uncover any new ideas, but the ideas of more experienced users.\n\nThe most inspiring work I took a lot from was **[this notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods)** by @dimitreoliveira.",
    "1201858": "nickuzmenkov  should we use max of soft labels ?",
    "1201877": "Hello! \n\nThe idea behind this is to use a combination of soft labels the ground truth, rather than soft labels standalone, thus paying attention to those heavily misclassed samples (which are probably the noisy ones) while still training on the original labels.",
    "1201880": "Any eg of label  That you can give  pls",
    "1201882": "Sorry, I don't understand. Could you please paraphrase that?",
    "1201905": "Nice to hear, my four models which I ensemble are scoring 0.908, 0.906, 0.904 and 0.902 (CV), while the average of the produce 0.904 (LB). \nI am satisfied for first competition, I came here to learn in the first place, if i get medal that would be bonus. \nThanks and good luck :) !",
    "1201910": "Check the plant pathology's 1st place solution linked above, you should one hot encode your labels and blend them with the soft labels\n```\nif self.soft_labels is not None:\n            label = torch.FloatTensor(\n                (self.data.iloc[index, 1:].values * 0.7).astype(np.float16)\n                + (self.soft_labels.iloc[index, 1:].values * 0.3).astype(np.float16)\n```\nAnd make sure to tweak your loss function to accept one hot encoded labels."
  },
  "source": "meta"
}