{
  "id": 198584,
  "title": "Incidence of disease in Cassava plants in Africa (and impact on metric choice)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198584",
  "author_name": "Thomas Brekk Unnvik",
  "post_date": "2020-11-21T23:33:47.671000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Having just started checking out this competition, I was alerted to the high percentage of images showing plants with Cassava Mosaic Disease (CMD) (about 61%) by a post by <a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a> . <br>\nWondering if infected plants perhaps were highly overrepresented in the dataset compared to healthy ones, I decided to research the topic a bit. </p>\n<p>It turns out that CMD is indeed a very large problem for Cassava growers in Africa. In this recent article from the Journal of Plant Pathology about the impact of CMD in Zambia, there were regions were in 2009 where 67 and 71 percent prevalence rates were recorded, for example, with infection even reaching 100% in isolated areas.<br>\n<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6951474/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6951474/</a></p>\n<p>So, the high occurence of CMD in the dataset seems to correlate with what is the real situation on the ground in many places in Africa. An interesting question, though, is how this will impact the learning algorithms and the ability to distinguish between the different diseases (and healthy plants). The above mentioned article mentions that Cassava Brown Streak Disease is on the rise, and may be more devastating when it occurs than CMD. So, since the results from this competition will hopefully be able to be used in practice, it's important to have a dicussion on the metric used (as <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> and others have brought up), as in practical cases it might be important not to misclassify one or more of the less prevalent cases in this dataset. Even though the penalty might not be that big in terms of score in this particular competition, it could potentially be a significant problem in practice (like the well-known, but more extreme, case of detecting rare forms of cancers, where predicting \"negative\" every time will give you an accuracy of, say, 99.5%, but is obviously not very useful). Would be interesting to hear what others of you think about this question. In any case, in a practical application it will be important to study the confusion matrix of the algorithm to make sure that no \"problematic\" misclassifications are systemically made.</p>",
  "messages": [
    {
      "id": 1086681,
      "postDate": "2020-11-21T23:33:47.670Z",
      "content": "<p>Having just started checking out this competition, I was alerted to the high percentage of images showing plants with Cassava Mosaic Disease (CMD) (about 61%) by a post by <a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a> . <br>\nWondering if infected plants perhaps were highly overrepresented in the dataset compared to healthy ones, I decided to research the topic a bit. </p>\n<p>It turns out that CMD is indeed a very large problem for Cassava growers in Africa. In this recent article from the Journal of Plant Pathology about the impact of CMD in Zambia, there were regions were in 2009 where 67 and 71 percent prevalence rates were recorded, for example, with infection even reaching 100% in isolated areas.<br>\n<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6951474/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6951474/</a></p>\n<p>So, the high occurence of CMD in the dataset seems to correlate with what is the real situation on the ground in many places in Africa. An interesting question, though, is how this will impact the learning algorithms and the ability to distinguish between the different diseases (and healthy plants). The above mentioned article mentions that Cassava Brown Streak Disease is on the rise, and may be more devastating when it occurs than CMD. So, since the results from this competition will hopefully be able to be used in practice, it's important to have a dicussion on the metric used (as <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> and others have brought up), as in practical cases it might be important not to misclassify one or more of the less prevalent cases in this dataset. Even though the penalty might not be that big in terms of score in this particular competition, it could potentially be a significant problem in practice (like the well-known, but more extreme, case of detecting rare forms of cancers, where predicting \"negative\" every time will give you an accuracy of, say, 99.5%, but is obviously not very useful). Would be interesting to hear what others of you think about this question. In any case, in a practical application it will be important to study the confusion matrix of the algorithm to make sure that no \"problematic\" misclassifications are systemically made.</p>",
      "rawMarkdown": "Having just started checking out this competition, I was alerted to the high percentage of images showing plants with Cassava Mosaic Disease (CMD) (about 61%) by a post by @kyoshioka47 . \nWondering if infected plants perhaps were highly overrepresented in the dataset compared to healthy ones, I decided to research the topic a bit. \n\nIt turns out that CMD is indeed a very large problem for Cassava growers in Africa. In this recent article from the Journal of Plant Pathology about the impact of CMD in Zambia, there were regions were in 2009 where 67 and 71 percent prevalence rates were recorded, for example, with infection even reaching 100% in isolated areas.\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6951474/\n\nSo, the high occurence of CMD in the dataset seems to correlate with what is the real situation on the ground in many places in Africa. An interesting question, though, is how this will impact the learning algorithms and the ability to distinguish between the different diseases (and healthy plants). The above mentioned article mentions that Cassava Brown Streak Disease is on the rise, and may be more devastating when it occurs than CMD. So, since the results from this competition will hopefully be able to be used in practice, it's important to have a dicussion on the metric used (as @ihelon and others have brought up), as in practical cases it might be important not to misclassify one or more of the less prevalent cases in this dataset. Even though the penalty might not be that big in terms of score in this particular competition, it could potentially be a significant problem in practice (like the well-known, but more extreme, case of detecting rare forms of cancers, where predicting \"negative\" every time will give you an accuracy of, say, 99.5%, but is obviously not very useful). Would be interesting to hear what others of you think about this question. In any case, in a practical application it will be important to study the confusion matrix of the algorithm to make sure that no \"problematic\" misclassifications are systemically made.\n\n",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1086681": "Having just started checking out this competition, I was alerted to the high percentage of images showing plants with Cassava Mosaic Disease (CMD) (about 61%) by a post by @kyoshioka47 . \nWondering if infected plants perhaps were highly overrepresented in the dataset compared to healthy ones, I decided to research the topic a bit. \n\nIt turns out that CMD is indeed a very large problem for Cassava growers in Africa. In this recent article from the Journal of Plant Pathology about the impact of CMD in Zambia, there were regions were in 2009 where 67 and 71 percent prevalence rates were recorded, for example, with infection even reaching 100% in isolated areas.\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6951474/\n\nSo, the high occurence of CMD in the dataset seems to correlate with what is the real situation on the ground in many places in Africa. An interesting question, though, is how this will impact the learning algorithms and the ability to distinguish between the different diseases (and healthy plants). The above mentioned article mentions that Cassava Brown Streak Disease is on the rise, and may be more devastating when it occurs than CMD. So, since the results from this competition will hopefully be able to be used in practice, it's important to have a dicussion on the metric used (as @ihelon and others have brought up), as in practical cases it might be important not to misclassify one or more of the less prevalent cases in this dataset. Even though the penalty might not be that big in terms of score in this particular competition, it could potentially be a significant problem in practice (like the well-known, but more extreme, case of detecting rare forms of cancers, where predicting \"negative\" every time will give you an accuracy of, say, 99.5%, but is obviously not very useful). Would be interesting to hear what others of you think about this question. In any case, in a practical application it will be important to study the confusion matrix of the algorithm to make sure that no \"problematic\" misclassifications are systemically made.\n\n"
  }
}