{
  "id": 211853,
  "title": "Bit worried about the usefulness of this competition",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/211853",
  "author_name": "",
  "post_date": "2021-01-16T14:42:45.113023400Z",
  "votes": 70,
  "comment_count": 26,
  "views": 0,
  "content": "<p>TLDR: We aren't just dealing with label \"noise\" here. Pretty much all healthy labels are sick. This is label cacophony.</p>\n<p>I tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.</p>\n<p>I decided to look at random false positives (sick incorrectly classified as healthy)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2F72ce9acb4a4dd904450c514b67bd08f5%2F2021-01-01-12-41-59.png?generation=1610807784880412&amp;alt=media\" alt=\"\"></p>\n<p>and I thought hmm… seems my model sucks.</p>\n<p>Then I looked at random false negatives (healthy incorrectly classified as sick)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2Fd15d03beccf85192955800f1d9457855%2F2021-01-01-12-41-42.png?generation=1610807689114473&amp;alt=media\" alt=\"\"></p>\n<p>and realised what was really wrong. The resolution is not right here, but ALL (apart from the first one - don't know what that is) of these are sick.</p>\n<p>Don't get me wrong here, I know that people have been mentioning noisy labels for quite some time on this discussion forum. But this is probably two steps up from what I would call just \"noisy\". I then took it a step further and decided to copy all false negative images to a folder, then go through each of them one by one. I'm pretty sure it's at least 90% mislabelled, where the other 10% I'm not even sure of. Maybe 50% plus are dead obvious. The rest are more subtle. But seeing as many healthy looking images with one sick leaf are labelled sick, I think my reasoning was fair.</p>\n<p>There are about 10% healthy labels in the dataset. The leaderboard is like a traffic pileup at 90%. Unfortunately I have serious doubts about the real world usefulness of this competition. Not that it's not useful to be able to diagnose plant diseases from images, but more that we aren't going to be able to squeeze the juice out of Kaggle's top minds. Sorry to be a downer … :(</p>\n<hr>\n<p>For anyone who would like to check themselves. Here are the \"predicted sick, ground truth healthy\" image ids from 1 of 5 folds <a href=\"https://drive.google.com/file/d/1-1w7xXaBwP-gs-J8WRXCNDZnm8pRD5Vr/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1-1w7xXaBwP-gs-J8WRXCNDZnm8pRD5Vr/view?usp=sharing</a></p>\n<hr>\n<p><strong>NEVERTHELESS</strong>, great work to those who are persevering despite the challenges and leading the way at the top of the leader board!</p>",
  "messages": [
    {
      "id": "1155653",
      "postDate": "01/16/2021 14:42:45",
      "content": "<p>TLDR: We aren't just dealing with label \"noise\" here. Pretty much all healthy labels are sick. This is label cacophony.</p>\n<p>I tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.</p>\n<p>I decided to look at random false positives (sick incorrectly classified as healthy)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2F72ce9acb4a4dd904450c514b67bd08f5%2F2021-01-01-12-41-59.png?generation=1610807784880412&amp;alt=media\" alt=\"\"></p>\n<p>and I thought hmm… seems my model sucks.</p>\n<p>Then I looked at random false negatives (healthy incorrectly classified as sick)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2Fd15d03beccf85192955800f1d9457855%2F2021-01-01-12-41-42.png?generation=1610807689114473&amp;alt=media\" alt=\"\"></p>\n<p>and realised what was really wrong. The resolution is not right here, but ALL (apart from the first one - don't know what that is) of these are sick.</p>\n<p>Don't get me wrong here, I know that people have been mentioning noisy labels for quite some time on this discussion forum. But this is probably two steps up from what I would call just \"noisy\". I then took it a step further and decided to copy all false negative images to a folder, then go through each of them one by one. I'm pretty sure it's at least 90% mislabelled, where the other 10% I'm not even sure of. Maybe 50% plus are dead obvious. The rest are more subtle. But seeing as many healthy looking images with one sick leaf are labelled sick, I think my reasoning was fair.</p>\n<p>There are about 10% healthy labels in the dataset. The leaderboard is like a traffic pileup at 90%. Unfortunately I have serious doubts about the real world usefulness of this competition. Not that it's not useful to be able to diagnose plant diseases from images, but more that we aren't going to be able to squeeze the juice out of Kaggle's top minds. Sorry to be a downer … :(</p>\n<hr>\n<p>For anyone who would like to check themselves. Here are the \"predicted sick, ground truth healthy\" image ids from 1 of 5 folds <a href=\"https://drive.google.com/file/d/1-1w7xXaBwP-gs-J8WRXCNDZnm8pRD5Vr/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1-1w7xXaBwP-gs-J8WRXCNDZnm8pRD5Vr/view?usp=sharing</a></p>\n<hr>\n<p><strong>NEVERTHELESS</strong>, great work to those who are persevering despite the challenges and leading the way at the top of the leader board!</p>",
      "rawMarkdown": "TLDR: We aren't just dealing with label \"noise\" here. Pretty much all healthy labels are sick. This is label cacophony.\n\nI tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.\n\nI decided to look at random false positives (sick incorrectly classified as healthy)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2F72ce9acb4a4dd904450c514b67bd08f5%2F2021-01-01-12-41-59.png?generation=1610807784880412&alt=media)\n\nand I thought hmm... seems my model sucks.\n\nThen I looked at random false negatives (healthy incorrectly classified as sick)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2Fd15d03beccf85192955800f1d9457855%2F2021-01-01-12-41-42.png?generation=1610807689114473&alt=media)\n\nand realised what was really wrong. The resolution is not right here, but ALL (apart from the first one - don't know what that is) of these are sick.\n\nDon't get me wrong here, I know that people have been mentioning noisy labels for quite some time on this discussion forum. But this is probably two steps up from what I would call just \"noisy\". I then took it a step further and decided to copy all false negative images to a folder, then go through each of them one by one. I'm pretty sure it's at least 90% mislabelled, where the other 10% I'm not even sure of. Maybe 50% plus are dead obvious. The rest are more subtle. But seeing as many healthy looking images with one sick leaf are labelled sick, I think my reasoning was fair.\n\nThere are about 10% healthy labels in the dataset. The leaderboard is like a traffic pileup at 90%. Unfortunately I have serious doubts about the real world usefulness of this competition. Not that it's not useful to be able to diagnose plant diseases from images, but more that we aren't going to be able to squeeze the juice out of Kaggle's top minds. Sorry to be a downer ... :(\n\n---\n\nFor anyone who would like to check themselves. Here are the \"predicted sick, ground truth healthy\" image ids from 1 of 5 folds https://drive.google.com/file/d/1-1w7xXaBwP-gs-J8WRXCNDZnm8pRD5Vr/view?usp=sharing\n\n---\n\n**NEVERTHELESS**, great work to those who are persevering despite the challenges and leading the way at the top of the leader board!",
      "votes": null
    },
    {
      "id": "1155983",
      "postDate": "01/16/2021 20:24:40",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> for mining so deep.<br>\nWow, it is really unbelievable that there are at least 90% mislabelled images. If it is true, I am really doubt what our models are learning and predicting. If there are at most 20% mislabelled images, the competition would be still useful because it can mimic the real situation, and it just probably narrows the gap between each category. But, 90%!, I think we are learning a random distribution.</p>",
      "rawMarkdown": "Thanks @alexandersoare for mining so deep.\nWow, it is really unbelievable that there are at least 90% mislabelled images. If it is true, I am really doubt what our models are learning and predicting. If there are at most 20% mislabelled images, the competition would be still useful because it can mimic the real situation, and it just probably narrows the gap between each category. But, 90%!, I think we are learning a random distribution.",
      "votes": null
    },
    {
      "id": "1156005",
      "postDate": "01/16/2021 20:44:46",
      "content": "<p>I'd really doubt that it's true for the disease classes as it wouldn't be possible to achieve 90%. I think there's a serious problem with the healthy labels.</p>",
      "rawMarkdown": "I'd really doubt that it's true for the disease classes as it wouldn't be possible to achieve 90%. I think there's a serious problem with the healthy labels.",
      "votes": null
    },
    {
      "id": "1156193",
      "postDate": "01/17/2021 02:17:58",
      "content": "<p>I think the problem here is really how we define \"sick\" or \"healthy\". According to the organizers, the training set was labelled by one expert, while the test set is labelled by a consensus of three experts. However, from the private LB leak, we know that the distribution of the private test set is similar to the public test set. </p>\n<p>So the question now is, do we really think ourselves as being able to label the leaves better than the experts? After all, whether the leaves are healthy or sick are subjective. For example, if an image has 10 leaves, but only 2 are sick, they might still label it as healthy. Or even if we disregard the question, maybe what the organizers are looking for is a model able to label the leaves according to their standards, not our own. </p>\n<p>Though, I have to agree with you, the healthy labels are the ones with the worst noise. Denoising the healthy labels increased my CV tremendously, though public LB not so much :)</p>",
      "rawMarkdown": "I think the problem here is really how we define \"sick\" or \"healthy\". According to the organizers, the training set was labelled by one expert, while the test set is labelled by a consensus of three experts. However, from the private LB leak, we know that the distribution of the private test set is similar to the public test set. \n\nSo the question now is, do we really think ourselves as being able to label the leaves better than the experts? After all, whether the leaves are healthy or sick are subjective. For example, if an image has 10 leaves, but only 2 are sick, they might still label it as healthy. Or even if we disregard the question, maybe what the organizers are looking for is a model able to label the leaves according to their standards, not our own. \n\nThough, I have to agree with you, the healthy labels are the ones with the worst noise. Denoising the healthy labels increased my CV tremendously, though public LB not so much :)",
      "votes": null
    },
    {
      "id": "1156195",
      "postDate": "01/17/2021 02:21:31",
      "content": "<p>90% mislabelled…. I thought maybe 10% mislabelled but it was opposite. If I can't trust my own train label, what can I trust in this competition? Thanks for sharing your analyze. </p>",
      "rawMarkdown": "90% mislabelled.... I thought maybe 10% mislabelled but it was opposite. If I can't trust my own train label, what can I trust in this competition? Thanks for sharing your analyze.",
      "votes": null
    },
    {
      "id": "1156223",
      "postDate": "01/17/2021 03:04:54",
      "content": "<p>How would you go about denoising the data?</p>",
      "rawMarkdown": "How would you go about denoising the data?",
      "votes": null
    },
    {
      "id": "1156271",
      "postDate": "01/17/2021 04:03:03",
      "content": "<p>Using OOF accuracy. Then select a confidence level and remove data based on that. Or you can personally relabel, though I didn't take the time to do that.</p>",
      "rawMarkdown": "Using OOF accuracy. Then select a confidence level and remove data based on that. Or you can personally relabel, though I didn't take the time to do that.",
      "votes": null
    },
    {
      "id": "1156324",
      "postDate": "01/17/2021 04:50:48",
      "content": "<p>From the pictures you shown, you might be doing too much augmentation on the color. There are some mislabeled data but definitely not at 90%. And, also noisy label is a pretty common problem in data science and machine learning. Find a good model to predict with noisy label is definitely very helpful learning.</p>",
      "rawMarkdown": "From the pictures you shown, you might be doing too much augmentation on the color. There are some mislabeled data but definitely not at 90%. And, also noisy label is a pretty common problem in data science and machine learning. Find a good model to predict with noisy label is definitely very helpful learning.",
      "votes": null
    },
    {
      "id": "1156593",
      "postDate": "01/17/2021 09:18:02",
      "content": "<p><a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a> these images are not augmented in Python. They are a screen snip of a screen snip though, so some things may have changed in translation.</p>\n<p>As for \"definitely not 90%\", sure if you think so. I'm just wondering if you've explicitly done the exercise though. I've provided a link in an edit to a text file containing the file names of the culprits.</p>",
      "rawMarkdown": "louis925 these images are not augmented in Python. They are a screen snip of a screen snip though, so some things may have changed in translation.\n\nAs for \"definitely not 90%\", sure if you think so. I'm just wondering if you've explicitly done the exercise though. I've provided a link in an edit to a text file containing the file names of the culprits.",
      "votes": null
    },
    {
      "id": "1156640",
      "postDate": "01/17/2021 09:49:32",
      "content": "<p>I take my IPhone into the fields.   </p>\n<p>Do I take one picture and assume from the model prediction that it's telling me the status of my field.  Pretty sure the answer is NO - I take lots of pictures and process them all on my phone - I than act on what appears to be the correct signal.</p>\n<p>IF I am a framer - I probably know what a healthy leave looks like. I don;t know crap about plants but pretty sure I can tell the healthy ones.  </p>\n<p>What I might have a need to know is what type of disease is affecting the plants that do not look good so that I can apply the correct corrective action.</p>\n<p>So - first assumption for useful nature of the model - will it end up on a phone that can be used by NON experts.<br>\nNext assumption - can they afford the phone?</p>",
      "rawMarkdown": "I take my IPhone into the fields.   \n\nDo I take one picture and assume from the model prediction that it's telling me the status of my field.  Pretty sure the answer is NO - I take lots of pictures and process them all on my phone - I than act on what appears to be the correct signal.\n\nIF I am a framer - I probably know what a healthy leave looks like. I don;t know crap about plants but pretty sure I can tell the healthy ones.  \n\nWhat I might have a need to know is what type of disease is affecting the plants that do not look good so that I can apply the correct corrective action.\n\nSo - first assumption for useful nature of the model - will it end up on a phone that can be used by NON experts.\nNext assumption - can they afford the phone?",
      "votes": null
    },
    {
      "id": "1157182",
      "postDate": "01/17/2021 18:06:08",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/junyingsg\" target=\"_blank\">@junyingsg</a> can you tell me a bit more about OOF accuracy. I have seen it many times but am not sure what it actually does? Any links will also be helpful. TIA</p>",
      "rawMarkdown": "Hey @junyingsg can you tell me a bit more about OOF accuracy. I have seen it many times but am not sure what it actually does? Any links will also be helpful. TIA",
      "votes": null
    },
    {
      "id": "1157215",
      "postDate": "01/17/2021 18:26:09",
      "content": "<p>I made a notebook detailing my methodology for denoising the labels, do check it out: <a href=\"https://www.kaggle.com/junyingsg/step-by-step-guide-to-denoising-your-labels/\" target=\"_blank\">https://www.kaggle.com/junyingsg/step-by-step-guide-to-denoising-your-labels/</a></p>",
      "rawMarkdown": "I made a notebook detailing my methodology for denoising the labels, do check it out: https://www.kaggle.com/junyingsg/step-by-step-guide-to-denoising-your-labels/",
      "votes": null
    },
    {
      "id": "1157385",
      "postDate": "01/17/2021 21:09:54",
      "content": "<p>Thanks for the very useful analysis! <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> </p>\n<p>Suppose it contains 90% of mislabelled. We are now able to build a model that calculates accuracy of about 90%(on LB score). This seems a bit odd.</p>\n<p>I believe that there is a way of labeling that only an expert can determine.</p>",
      "rawMarkdown": "Thanks for the very useful analysis! @alexandersoare \n\nSuppose it contains 90% of mislabelled. We are now able to build a model that calculates accuracy of about 90%(on LB score). This seems a bit odd.\n\nI believe that there is a way of labeling that only an expert can determine.",
      "votes": null
    },
    {
      "id": "1157413",
      "postDate": "01/17/2021 21:42:49",
      "content": "<p>Let me clarify. <strong>Of the healthy labels that my binary classifier predicted as sick, 90%+ are actually sick</strong>.</p>\n<p>And to give even more detail. I went ahead and relabelled these. Then I trained another binary classifier on the fixed labels. Then I got the same result as above. Then I did it one more time, and got the same result yet again. So I concluded 90% of ALL healthy labels are actually sick. Although I did not actually audit this by hand labelling ALL healthy labels 1 by 1.</p>",
      "rawMarkdown": "Let me clarify. **Of the healthy labels that my binary classifier predicted as sick, 90%+ are actually sick**.\n\nAnd to give even more detail. I went ahead and relabelled these. Then I trained another binary classifier on the fixed labels. Then I got the same result as above. Then I did it one more time, and got the same result yet again. So I concluded 90% of ALL healthy labels are actually sick. Although I did not actually audit this by hand labelling ALL healthy labels 1 by 1.",
      "votes": null
    },
    {
      "id": "1157429",
      "postDate": "01/17/2021 22:22:31",
      "content": "<p>You discuss only healthy labels. I understand it better now. Thank you very much.</p>",
      "rawMarkdown": "You discuss only healthy labels. I understand it better now. Thank you very much.",
      "votes": null
    },
    {
      "id": "1157430",
      "postDate": "01/17/2021 22:23:15",
      "content": "<p>When you relabeled the leaves, did it improve test accuracy? Do you think the test data contains the same type of noise?</p>",
      "rawMarkdown": "When you relabeled the leaves, did it improve test accuracy? Do you think the test data contains the same type of noise?",
      "votes": null
    },
    {
      "id": "1157510",
      "postDate": "01/18/2021 00:57:37",
      "content": "<blockquote>\n  <p>I tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.</p>\n</blockquote>\n<p>The CMD class is the most widely distributed class of this data, not the healthy class. Therefore, the way I use to classify the most widely distributed class (CMD class) and non CMD class (including health class) is not the same as you. In this way, I use the double verification method to first confirm whether it is a CMD class (class = ‘3’), and then do multi classification</p>",
      "rawMarkdown": "> I tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.\n\n\nThe CMD class is the most widely distributed class of this data, not the healthy class. Therefore, the way I use to classify the most widely distributed class (CMD class) and non CMD class (including health class) is not the same as you. In this way, I use the double verification method to first confirm whether it is a CMD class (class = ‘3’), and then do multi classification",
      "votes": null
    },
    {
      "id": "1157523",
      "postDate": "01/18/2021 01:21:52",
      "content": "<p>My model gets 96% accuracy for class 3. Looks like the problem is in the rest of the classes, especially 0 and 4.<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212197\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212197</a></p>",
      "rawMarkdown": "My model gets 96% accuracy for class 3. Looks like the problem is in the rest of the classes, especially 0 and 4.\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212197",
      "votes": null
    },
    {
      "id": "1157524",
      "postDate": "01/18/2021 01:25:47",
      "content": "<p>I'm interested why the LB is so different from the previous cassava competition. (LB 0.92, PB 0.93)<br>\nThe hosts should have labeled it in the same manner, why the big differences?<br>\n<a href=\"https://www.kaggle.com/c/cassava-disease/leaderboard\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease/leaderboard</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3508221%2F9ea11cfd591b2c1d999e103f116b9f70%2F2021-01-18%20102534.png?generation=1610933145733453&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'm interested why the LB is so different from the previous cassava competition. (LB 0.92, PB 0.93)\nThe hosts should have labeled it in the same manner, why the big differences?\nhttps://www.kaggle.com/c/cassava-disease/leaderboard\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3508221%2F9ea11cfd591b2c1d999e103f116b9f70%2F2021-01-18%20102534.png?generation=1610933145733453&alt=media)",
      "votes": null
    },
    {
      "id": "1157773",
      "postDate": "01/18/2021 06:13:33",
      "content": "<p>I agree, class 3 is well classified, the problem is the class 4.</p>",
      "rawMarkdown": "I agree, class 3 is well classified, the problem is the class 4.",
      "votes": null
    },
    {
      "id": "1157948",
      "postDate": "01/18/2021 09:20:53",
      "content": "<p>It improved the accuracy but not because anything interesting was happening. Simply because there were less healthy labels to predict.</p>",
      "rawMarkdown": "It improved the accuracy but not because anything interesting was happening. Simply because there were less healthy labels to predict.",
      "votes": null
    },
    {
      "id": "1159055",
      "postDate": "01/19/2021 02:40:51",
      "content": "<p>I was wondering the same! Why are models performing (significantly) worse this time around?</p>",
      "rawMarkdown": "I was wondering the same! Why are models performing (significantly) worse this time around?",
      "votes": null
    },
    {
      "id": "1159854",
      "postDate": "01/19/2021 14:12:52",
      "content": "<p>This really makes sense. Thanks for sharing ur insights <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> </p>",
      "rawMarkdown": "This really makes sense. Thanks for sharing ur insights @alexandersoare",
      "votes": null
    },
    {
      "id": "1160578",
      "postDate": "01/20/2021 02:55:26",
      "content": "<p>I haven't looked extensively into the previous competition's dataset but I think the data was less noisy relative to this competition.</p>",
      "rawMarkdown": "I haven't looked extensively into the previous competition's dataset but I think the data was less noisy relative to this competition.",
      "votes": null
    },
    {
      "id": "1169499",
      "postDate": "01/25/2021 14:26:38",
      "content": "<p>Good set of articles to work with Noisy Labels</p>\n<p><a href=\"https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise\" target=\"_blank\">https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise</a></p>",
      "rawMarkdown": "Good set of articles to work with Noisy Labels\n\nhttps://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise",
      "votes": null
    },
    {
      "id": "1170291",
      "postDate": "01/26/2021 05:44:56",
      "content": "<p>This looks pretty worrying. I haven't looked closely at the labels but if it is systemic mislabeling maybe it is worth it for the hosts to evaluate where things may have gone wrong in their labeling pipeline. It is curious why the performance is so different from the previous similar competitions</p>",
      "rawMarkdown": "This looks pretty worrying. I haven't looked closely at the labels but if it is systemic mislabeling maybe it is worth it for the hosts to evaluate where things may have gone wrong in their labeling pipeline. It is curious why the performance is so different from the previous similar competitions",
      "votes": null
    },
    {
      "id": "1177570",
      "postDate": "01/30/2021 11:36:07",
      "content": "<blockquote>\n  <p>the training set was labelled by one expert, while the test set is labelled by a consensus of three experts.</p>\n</blockquote>\n<p>Where can I find this information?<br>\nThis is important for estimating the reliability of LB.</p>",
      "rawMarkdown": "> the training set was labelled by one expert, while the test set is labelled by a consensus of three experts.\n\nWhere can I find this information?\nThis is important for estimating the reliability of LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1155983,
      "author_name": "woshifym",
      "author_url": "",
      "post_date": "01/16/2021 20:24:40",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> for mining so deep.<br>\nWow, it is really unbelievable that there are at least 90% mislabelled images. If it is true, I am really doubt what our models are learning and predicting. If there are at most 20% mislabelled images, the competition would be still useful because it can mimic the real situation, and it just probably narrows the gap between each category. But, 90%!, I think we are learning a random distribution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1156005,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "01/16/2021 20:44:46",
          "content": "<p>I'd really doubt that it's true for the disease classes as it wouldn't be possible to achieve 90%. I think there's a serious problem with the healthy labels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1156193,
      "author_name": "junyingsg",
      "author_url": "",
      "post_date": "01/17/2021 02:17:58",
      "content": "<p>I think the problem here is really how we define \"sick\" or \"healthy\". According to the organizers, the training set was labelled by one expert, while the test set is labelled by a consensus of three experts. However, from the private LB leak, we know that the distribution of the private test set is similar to the public test set. </p>\n<p>So the question now is, do we really think ourselves as being able to label the leaves better than the experts? After all, whether the leaves are healthy or sick are subjective. For example, if an image has 10 leaves, but only 2 are sick, they might still label it as healthy. Or even if we disregard the question, maybe what the organizers are looking for is a model able to label the leaves according to their standards, not our own. </p>\n<p>Though, I have to agree with you, the healthy labels are the ones with the worst noise. Denoising the healthy labels increased my CV tremendously, though public LB not so much :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1156223,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/17/2021 03:04:54",
          "content": "<p>How would you go about denoising the data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1156271,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/17/2021 04:03:03",
          "content": "<p>Using OOF accuracy. Then select a confidence level and remove data based on that. Or you can personally relabel, though I didn't take the time to do that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157182,
          "author_name": "sadmanaraf",
          "author_url": "",
          "post_date": "01/17/2021 18:06:08",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/junyingsg\" target=\"_blank\">@junyingsg</a> can you tell me a bit more about OOF accuracy. I have seen it many times but am not sure what it actually does? Any links will also be helpful. TIA</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157215,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/17/2021 18:26:09",
          "content": "<p>I made a notebook detailing my methodology for denoising the labels, do check it out: <a href=\"https://www.kaggle.com/junyingsg/step-by-step-guide-to-denoising-your-labels/\" target=\"_blank\">https://www.kaggle.com/junyingsg/step-by-step-guide-to-denoising-your-labels/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1177570,
          "author_name": "sunakuzira",
          "author_url": "",
          "post_date": "01/30/2021 11:36:07",
          "content": "<blockquote>\n  <p>the training set was labelled by one expert, while the test set is labelled by a consensus of three experts.</p>\n</blockquote>\n<p>Where can I find this information?<br>\nThis is important for estimating the reliability of LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1156195,
      "author_name": "vkehfdl1",
      "author_url": "",
      "post_date": "01/17/2021 02:21:31",
      "content": "<p>90% mislabelled…. I thought maybe 10% mislabelled but it was opposite. If I can't trust my own train label, what can I trust in this competition? Thanks for sharing your analyze. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156324,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "01/17/2021 04:50:48",
      "content": "<p>From the pictures you shown, you might be doing too much augmentation on the color. There are some mislabeled data but definitely not at 90%. And, also noisy label is a pretty common problem in data science and machine learning. Find a good model to predict with noisy label is definitely very helpful learning.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1156593,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "01/17/2021 09:18:02",
          "content": "<p><a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a> these images are not augmented in Python. They are a screen snip of a screen snip though, so some things may have changed in translation.</p>\n<p>As for \"definitely not 90%\", sure if you think so. I'm just wondering if you've explicitly done the exercise though. I've provided a link in an edit to a text file containing the file names of the culprits.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1156640,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/17/2021 09:49:32",
      "content": "<p>I take my IPhone into the fields.   </p>\n<p>Do I take one picture and assume from the model prediction that it's telling me the status of my field.  Pretty sure the answer is NO - I take lots of pictures and process them all on my phone - I than act on what appears to be the correct signal.</p>\n<p>IF I am a framer - I probably know what a healthy leave looks like. I don;t know crap about plants but pretty sure I can tell the healthy ones.  </p>\n<p>What I might have a need to know is what type of disease is affecting the plants that do not look good so that I can apply the correct corrective action.</p>\n<p>So - first assumption for useful nature of the model - will it end up on a phone that can be used by NON experts.<br>\nNext assumption - can they afford the phone?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1157385,
      "author_name": "sinchir0",
      "author_url": "",
      "post_date": "01/17/2021 21:09:54",
      "content": "<p>Thanks for the very useful analysis! <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> </p>\n<p>Suppose it contains 90% of mislabelled. We are now able to build a model that calculates accuracy of about 90%(on LB score). This seems a bit odd.</p>\n<p>I believe that there is a way of labeling that only an expert can determine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1157413,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "01/17/2021 21:42:49",
          "content": "<p>Let me clarify. <strong>Of the healthy labels that my binary classifier predicted as sick, 90%+ are actually sick</strong>.</p>\n<p>And to give even more detail. I went ahead and relabelled these. Then I trained another binary classifier on the fixed labels. Then I got the same result as above. Then I did it one more time, and got the same result yet again. So I concluded 90% of ALL healthy labels are actually sick. Although I did not actually audit this by hand labelling ALL healthy labels 1 by 1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157429,
          "author_name": "sinchir0",
          "author_url": "",
          "post_date": "01/17/2021 22:22:31",
          "content": "<p>You discuss only healthy labels. I understand it better now. Thank you very much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157430,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/17/2021 22:23:15",
          "content": "<p>When you relabeled the leaves, did it improve test accuracy? Do you think the test data contains the same type of noise?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157948,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "01/18/2021 09:20:53",
          "content": "<p>It improved the accuracy but not because anything interesting was happening. Simply because there were less healthy labels to predict.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1157510,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "01/18/2021 00:57:37",
      "content": "<blockquote>\n  <p>I tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.</p>\n</blockquote>\n<p>The CMD class is the most widely distributed class of this data, not the healthy class. Therefore, the way I use to classify the most widely distributed class (CMD class) and non CMD class (including health class) is not the same as you. In this way, I use the double verification method to first confirm whether it is a CMD class (class = ‘3’), and then do multi classification</p>",
      "votes": null,
      "replies": [
        {
          "id": 1157523,
          "author_name": "kyoshioka47",
          "author_url": "",
          "post_date": "01/18/2021 01:21:52",
          "content": "<p>My model gets 96% accuracy for class 3. Looks like the problem is in the rest of the classes, especially 0 and 4.<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212197\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212197</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157773,
          "author_name": "wantsu",
          "author_url": "",
          "post_date": "01/18/2021 06:13:33",
          "content": "<p>I agree, class 3 is well classified, the problem is the class 4.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1157524,
      "author_name": "kyoshioka47",
      "author_url": "",
      "post_date": "01/18/2021 01:25:47",
      "content": "<p>I'm interested why the LB is so different from the previous cassava competition. (LB 0.92, PB 0.93)<br>\nThe hosts should have labeled it in the same manner, why the big differences?<br>\n<a href=\"https://www.kaggle.com/c/cassava-disease/leaderboard\" target=\"_blank\">https://www.kaggle.com/c/cassava-disease/leaderboard</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3508221%2F9ea11cfd591b2c1d999e103f116b9f70%2F2021-01-18%20102534.png?generation=1610933145733453&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1159055,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "01/19/2021 02:40:51",
          "content": "<p>I was wondering the same! Why are models performing (significantly) worse this time around?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1160578,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/20/2021 02:55:26",
          "content": "<p>I haven't looked extensively into the previous competition's dataset but I think the data was less noisy relative to this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1159854,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "01/19/2021 14:12:52",
      "content": "<p>This really makes sense. Thanks for sharing ur insights <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1169499,
      "author_name": "idrabenia",
      "author_url": "",
      "post_date": "01/25/2021 14:26:38",
      "content": "<p>Good set of articles to work with Noisy Labels</p>\n<p><a href=\"https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise\" target=\"_blank\">https://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1170291,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "01/26/2021 05:44:56",
      "content": "<p>This looks pretty worrying. I haven't looked closely at the labels but if it is systemic mislabeling maybe it is worth it for the hosts to evaluate where things may have gone wrong in their labeling pipeline. It is curious why the performance is so different from the previous similar competitions</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1155653": "TLDR: We aren't just dealing with label \"noise\" here. Pretty much all healthy labels are sick. This is label cacophony.\n\nI tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.\n\nI decided to look at random false positives (sick incorrectly classified as healthy)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2F72ce9acb4a4dd904450c514b67bd08f5%2F2021-01-01-12-41-59.png?generation=1610807784880412&alt=media)\n\nand I thought hmm... seems my model sucks.\n\nThen I looked at random false negatives (healthy incorrectly classified as sick)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4256010%2Fd15d03beccf85192955800f1d9457855%2F2021-01-01-12-41-42.png?generation=1610807689114473&alt=media)\n\nand realised what was really wrong. The resolution is not right here, but ALL (apart from the first one - don't know what that is) of these are sick.\n\nDon't get me wrong here, I know that people have been mentioning noisy labels for quite some time on this discussion forum. But this is probably two steps up from what I would call just \"noisy\". I then took it a step further and decided to copy all false negative images to a folder, then go through each of them one by one. I'm pretty sure it's at least 90% mislabelled, where the other 10% I'm not even sure of. Maybe 50% plus are dead obvious. The rest are more subtle. But seeing as many healthy looking images with one sick leaf are labelled sick, I think my reasoning was fair.\n\nThere are about 10% healthy labels in the dataset. The leaderboard is like a traffic pileup at 90%. Unfortunately I have serious doubts about the real world usefulness of this competition. Not that it's not useful to be able to diagnose plant diseases from images, but more that we aren't going to be able to squeeze the juice out of Kaggle's top minds. Sorry to be a downer ... :(\n\n---\n\nFor anyone who would like to check themselves. Here are the \"predicted sick, ground truth healthy\" image ids from 1 of 5 folds https://drive.google.com/file/d/1-1w7xXaBwP-gs-J8WRXCNDZnm8pRD5Vr/view?usp=sharing\n\n---\n\n**NEVERTHELESS**, great work to those who are persevering despite the challenges and leading the way at the top of the leader board!",
    "1155983": "Thanks @alexandersoare for mining so deep.\nWow, it is really unbelievable that there are at least 90% mislabelled images. If it is true, I am really doubt what our models are learning and predicting. If there are at most 20% mislabelled images, the competition would be still useful because it can mimic the real situation, and it just probably narrows the gap between each category. But, 90%!, I think we are learning a random distribution.",
    "1156005": "I'd really doubt that it's true for the disease classes as it wouldn't be possible to achieve 90%. I think there's a serious problem with the healthy labels.",
    "1156193": "I think the problem here is really how we define \"sick\" or \"healthy\". According to the organizers, the training set was labelled by one expert, while the test set is labelled by a consensus of three experts. However, from the private LB leak, we know that the distribution of the private test set is similar to the public test set. \n\nSo the question now is, do we really think ourselves as being able to label the leaves better than the experts? After all, whether the leaves are healthy or sick are subjective. For example, if an image has 10 leaves, but only 2 are sick, they might still label it as healthy. Or even if we disregard the question, maybe what the organizers are looking for is a model able to label the leaves according to their standards, not our own. \n\nThough, I have to agree with you, the healthy labels are the ones with the worst noise. Denoising the healthy labels increased my CV tremendously, though public LB not so much :)",
    "1156195": "90% mislabelled.... I thought maybe 10% mislabelled but it was opposite. If I can't trust my own train label, what can I trust in this competition? Thanks for sharing your analyze.",
    "1156223": "How would you go about denoising the data?",
    "1156271": "Using OOF accuracy. Then select a confidence level and remove data based on that. Or you can personally relabel, though I didn't take the time to do that.",
    "1156324": "From the pictures you shown, you might be doing too much augmentation on the color. There are some mislabeled data but definitely not at 90%. And, also noisy label is a pretty common problem in data science and machine learning. Find a good model to predict with noisy label is definitely very helpful learning.",
    "1156593": "louis925 these images are not augmented in Python. They are a screen snip of a screen snip though, so some things may have changed in translation.\n\nAs for \"definitely not 90%\", sure if you think so. I'm just wondering if you've explicitly done the exercise though. I've provided a link in an edit to a text file containing the file names of the culprits.",
    "1156640": "I take my IPhone into the fields.   \n\nDo I take one picture and assume from the model prediction that it's telling me the status of my field.  Pretty sure the answer is NO - I take lots of pictures and process them all on my phone - I than act on what appears to be the correct signal.\n\nIF I am a framer - I probably know what a healthy leave looks like. I don;t know crap about plants but pretty sure I can tell the healthy ones.  \n\nWhat I might have a need to know is what type of disease is affecting the plants that do not look good so that I can apply the correct corrective action.\n\nSo - first assumption for useful nature of the model - will it end up on a phone that can be used by NON experts.\nNext assumption - can they afford the phone?",
    "1157182": "Hey @junyingsg can you tell me a bit more about OOF accuracy. I have seen it many times but am not sure what it actually does? Any links will also be helpful. TIA",
    "1157215": "I made a notebook detailing my methodology for denoising the labels, do check it out: https://www.kaggle.com/junyingsg/step-by-step-guide-to-denoising-your-labels/",
    "1157385": "Thanks for the very useful analysis! @alexandersoare \n\nSuppose it contains 90% of mislabelled. We are now able to build a model that calculates accuracy of about 90%(on LB score). This seems a bit odd.\n\nI believe that there is a way of labeling that only an expert can determine.",
    "1157413": "Let me clarify. **Of the healthy labels that my binary classifier predicted as sick, 90%+ are actually sick**.\n\nAnd to give even more detail. I went ahead and relabelled these. Then I trained another binary classifier on the fixed labels. Then I got the same result as above. Then I did it one more time, and got the same result yet again. So I concluded 90% of ALL healthy labels are actually sick. Although I did not actually audit this by hand labelling ALL healthy labels 1 by 1.",
    "1157429": "You discuss only healthy labels. I understand it better now. Thank you very much.",
    "1157430": "When you relabeled the leaves, did it improve test accuracy? Do you think the test data contains the same type of noise?",
    "1157510": "> I tried making a binary classifier for healthy vs sick, and the results lead me to pretty much abandoning this comp.\n\n\nThe CMD class is the most widely distributed class of this data, not the healthy class. Therefore, the way I use to classify the most widely distributed class (CMD class) and non CMD class (including health class) is not the same as you. In this way, I use the double verification method to first confirm whether it is a CMD class (class = ‘3’), and then do multi classification",
    "1157523": "My model gets 96% accuracy for class 3. Looks like the problem is in the rest of the classes, especially 0 and 4.\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/212197",
    "1157524": "I'm interested why the LB is so different from the previous cassava competition. (LB 0.92, PB 0.93)\nThe hosts should have labeled it in the same manner, why the big differences?\nhttps://www.kaggle.com/c/cassava-disease/leaderboard\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3508221%2F9ea11cfd591b2c1d999e103f116b9f70%2F2021-01-18%20102534.png?generation=1610933145733453&alt=media)",
    "1157773": "I agree, class 3 is well classified, the problem is the class 4.",
    "1157948": "It improved the accuracy but not because anything interesting was happening. Simply because there were less healthy labels to predict.",
    "1159055": "I was wondering the same! Why are models performing (significantly) worse this time around?",
    "1159854": "This really makes sense. Thanks for sharing ur insights @alexandersoare",
    "1160578": "I haven't looked extensively into the previous competition's dataset but I think the data was less noisy relative to this competition.",
    "1169499": "Good set of articles to work with Noisy Labels\n\nhttps://github.com/subeeshvasu/Awesome-Learning-with-Label-Noise",
    "1170291": "This looks pretty worrying. I haven't looked closely at the labels but if it is systemic mislabeling maybe it is worth it for the hosts to evaluate where things may have gone wrong in their labeling pipeline. It is curious why the performance is so different from the previous similar competitions",
    "1177570": "> the training set was labelled by one expert, while the test set is labelled by a consensus of three experts.\n\nWhere can I find this information?\nThis is important for estimating the reliability of LB."
  },
  "source": "meta"
}