{
  "id": 212197,
  "title": "Analyzing the model's errors by confusion matrix",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212197",
  "author_name": "arutema47",
  "post_date": "2021-01-18T01:06:45.155000",
  "votes": 24,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Notebook:<br>\n<a href=\"https://www.kaggle.com/kyoshioka47/analyze-your-model-performance-by-confusion-matrix\" target=\"_blank\">https://www.kaggle.com/kyoshioka47/analyze-your-model-performance-by-confusion-matrix</a></p>\n<p>A typical way to analyze your model performance is using the confusion matrix.<br>\nAfter getting the 5-fold prediction results of my CV 0.89, LB 0.9 model, you can get the confusion matrix by</p>\n<pre><code>from sklearn.metrics import confusion_matrix\ncm = confusion_matrix(df_train[\"label\"], logits.argmax(1))\n\nprint(cm)\n\n[[  719    59    22    49   238]\n [  108  1751    53    84   193]\n [   33    50  1924   251   128]\n [   36    71   199 12748   104]\n [  207   131   135   192  1912]]\n</code></pre>\n<p>Interestingly, the accuracy is very high for class-3 but very low for others.</p>\n<pre><code>for i, val in enumerate(cm):\n    print(\"for class {}: accuracy: {}\".format(i, val[i]/sum(val)*100))\n\nfor class 0: accuracy: 66.14535418583257\nfor class 1: accuracy: 79.99086340794884\nfor class 2: accuracy: 80.63704945515508\nfor class 3: accuracy: 96.88402492780058\nfor class 4: accuracy: 74.19480015521924\n</code></pre>\n<p>Especially for class 0, it is <strong>very</strong> likely to mistake for healthy!</p>\n<pre><code>for i, val in enumerate(cm[:-1]):\n    print(\"for class {}: possibility to mistake for healthy: {}\".format(i, val[4]/val[i]*100))\n\nfor class 0: possibility to mistake for healthy: 33.101529902642554\nfor class 1: possibility to mistake for healthy: 11.02227298686465\nfor class 2: possibility to mistake for healthy: 6.652806652806653\nfor class 3: possibility to mistake for healthy: 0.815814245371823\n</code></pre>\n<pre><code>from sklearn.metrics import classification_report\n\nprint(classification_report(df_train[\"label\"], logits.argmax(1)))\n\nprecision    recall  f1-score   support\n\n           0       0.65      0.66      0.66      1087\n           1       0.85      0.80      0.82      2189\n           2       0.82      0.81      0.82      2386\n           3       0.96      0.97      0.96     13158\n           4       0.74      0.74      0.74      2577\n\n    accuracy                           0.89     21397\n</code></pre>",
  "messages": [
    {
      "id": 1157513,
      "postDate": "2021-01-18T01:06:45.157Z",
      "content": "<p>Notebook:<br>\n<a href=\"https://www.kaggle.com/kyoshioka47/analyze-your-model-performance-by-confusion-matrix\" target=\"_blank\">https://www.kaggle.com/kyoshioka47/analyze-your-model-performance-by-confusion-matrix</a></p>\n<p>A typical way to analyze your model performance is using the confusion matrix.<br>\nAfter getting the 5-fold prediction results of my CV 0.89, LB 0.9 model, you can get the confusion matrix by</p>\n<pre><code>from sklearn.metrics import confusion_matrix\ncm = confusion_matrix(df_train[\"label\"], logits.argmax(1))\n\nprint(cm)\n\n[[  719    59    22    49   238]\n [  108  1751    53    84   193]\n [   33    50  1924   251   128]\n [   36    71   199 12748   104]\n [  207   131   135   192  1912]]\n</code></pre>\n<p>Interestingly, the accuracy is very high for class-3 but very low for others.</p>\n<pre><code>for i, val in enumerate(cm):\n    print(\"for class {}: accuracy: {}\".format(i, val[i]/sum(val)*100))\n\nfor class 0: accuracy: 66.14535418583257\nfor class 1: accuracy: 79.99086340794884\nfor class 2: accuracy: 80.63704945515508\nfor class 3: accuracy: 96.88402492780058\nfor class 4: accuracy: 74.19480015521924\n</code></pre>\n<p>Especially for class 0, it is <strong>very</strong> likely to mistake for healthy!</p>\n<pre><code>for i, val in enumerate(cm[:-1]):\n    print(\"for class {}: possibility to mistake for healthy: {}\".format(i, val[4]/val[i]*100))\n\nfor class 0: possibility to mistake for healthy: 33.101529902642554\nfor class 1: possibility to mistake for healthy: 11.02227298686465\nfor class 2: possibility to mistake for healthy: 6.652806652806653\nfor class 3: possibility to mistake for healthy: 0.815814245371823\n</code></pre>\n<pre><code>from sklearn.metrics import classification_report\n\nprint(classification_report(df_train[\"label\"], logits.argmax(1)))\n\nprecision    recall  f1-score   support\n\n           0       0.65      0.66      0.66      1087\n           1       0.85      0.80      0.82      2189\n           2       0.82      0.81      0.82      2386\n           3       0.96      0.97      0.96     13158\n           4       0.74      0.74      0.74      2577\n\n    accuracy                           0.89     21397\n</code></pre>",
      "rawMarkdown": "Notebook:\nhttps://www.kaggle.com/kyoshioka47/analyze-your-model-performance-by-confusion-matrix\n\nA typical way to analyze your model performance is using the confusion matrix.\nAfter getting the 5-fold prediction results of my CV 0.89, LB 0.9 model, you can get the confusion matrix by\n\n```python\nfrom sklearn.metrics import confusion_matrix\ncm = confusion_matrix(df_train[\"label\"], logits.argmax(1))\n\nprint(cm)\n\n[[  719    59    22    49   238]\n [  108  1751    53    84   193]\n [   33    50  1924   251   128]\n [   36    71   199 12748   104]\n [  207   131   135   192  1912]]\n\n```\n\nInterestingly, the accuracy is very high for class-3 but very low for others.\n\n```\nfor i, val in enumerate(cm):\n    print(\"for class {}: accuracy: {}\".format(i, val[i]/sum(val)*100))\n\nfor class 0: accuracy: 66.14535418583257\nfor class 1: accuracy: 79.99086340794884\nfor class 2: accuracy: 80.63704945515508\nfor class 3: accuracy: 96.88402492780058\nfor class 4: accuracy: 74.19480015521924\n\n```\n\nEspecially for class 0, it is **very** likely to mistake for healthy!\n\n```\nfor i, val in enumerate(cm[:-1]):\n    print(\"for class {}: possibility to mistake for healthy: {}\".format(i, val[4]/val[i]*100))\n\nfor class 0: possibility to mistake for healthy: 33.101529902642554\nfor class 1: possibility to mistake for healthy: 11.02227298686465\nfor class 2: possibility to mistake for healthy: 6.652806652806653\nfor class 3: possibility to mistake for healthy: 0.815814245371823\n```\n\n```\nfrom sklearn.metrics import classification_report\n\nprint(classification_report(df_train[\"label\"], logits.argmax(1)))\n\nprecision    recall  f1-score   support\n\n           0       0.65      0.66      0.66      1087\n           1       0.85      0.80      0.82      2189\n           2       0.82      0.81      0.82      2386\n           3       0.96      0.97      0.96     13158\n           4       0.74      0.74      0.74      2577\n\n    accuracy                           0.89     21397\n```",
      "votes": 24
    },
    {
      "id": 1184175,
      "postDate": "2021-02-03T12:23:44.437Z",
      "content": "<p>One more confusion matrix for B6 and images resized to 384x384:<br>\n<img src=\"https://i.ibb.co/LxPxB3F/image.png\" alt=\"\"></p>\n<p>In my case, the network mainly has a hard time recognizing healthy images, with the largest confusion to class 0</p>",
      "rawMarkdown": "One more confusion matrix for B6 and images resized to 384x384:\n![](https://i.ibb.co/LxPxB3F/image.png)\n\nIn my case, the network mainly has a hard time recognizing healthy images, with the largest confusion to class 0",
      "votes": 1
    },
    {
      "id": 1157525,
      "postDate": "2021-01-18T01:31:30.467Z",
      "content": "<p>my classification report:<br>\nprecision    recall  f1-score   support<br>\n           0       0.73      0.68      0.70       108<br>\n           1       0.84      0.83      0.83       219<br>\n           2       0.81      0.79      0.80       239<br>\n           3       0.96      0.98      0.97      1315<br>\n           4       0.78      0.74      0.76       258</p>",
      "rawMarkdown": "my classification report:\nprecision    recall  f1-score   support\n           0       0.73      0.68      0.70       108\n           1       0.84      0.83      0.83       219\n           2       0.81      0.79      0.80       239\n           3       0.96      0.98      0.97      1315\n           4       0.78      0.74      0.76       258",
      "votes": 1,
      "replies": [
        {
          "id": 1157643,
          "postDate": "2021-01-18T04:12:16.943Z",
          "content": "<p>Thanks, I added the P,R f1-score as well</p>",
          "rawMarkdown": "Thanks, I added the P,R f1-score as well",
          "votes": 1
        }
      ]
    },
    {
      "id": 1157522,
      "postDate": "2021-01-18T01:21:35.833Z",
      "content": "<p>This analysis confirms what was shared in this discussion topic:  <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211853\" target=\"_blank\">Bit worried about the usefulness of this competition</a>. The model is very likely to mistake for healthy because a lot of the healthy images are mislabeled. </p>",
      "rawMarkdown": "This analysis confirms what was shared in this discussion topic:  [Bit worried about the usefulness of this competition](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211853). The model is very likely to mistake for healthy because a lot of the healthy images are mislabeled. ",
      "votes": 1,
      "replies": [
        {
          "id": 1157621,
          "postDate": "2021-01-18T03:56:12.353Z",
          "content": "<p>Are we sure that these images are mislabeled? I know that many of the label 4 images seem to be yellow in color and have a disease but perhaps the yellow color is due to a natural lifecycle of the leaf or environmental factors?</p>",
          "rawMarkdown": "Are we sure that these images are mislabeled? I know that many of the label 4 images seem to be yellow in color and have a disease but perhaps the yellow color is due to a natural lifecycle of the leaf or environmental factors?",
          "votes": 3
        },
        {
          "id": 1159558,
          "postDate": "2021-01-19T11:04:04.503Z",
          "content": "<p>I agree with <a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> and I would like to give my point of view on what we consider \"noise\".</p>\n<p>I think that if an image shows a plant/leaf with symptoms in the early stage of disease, it could have been labeled (or mis-labeled) as healthy for 2 reasons: the more obvious is by mistake, but could it be due to the fact that the disease is too weak to harm the plant and thus the plant is actually healthy?</p>",
          "rawMarkdown": "I agree with @ayu055 and I would like to give my point of view on what we consider \"noise\".\n\nI think that if an image shows a plant/leaf with symptoms in the early stage of disease, it could have been labeled (or mis-labeled) as healthy for 2 reasons: the more obvious is by mistake, but could it be due to the fact that the disease is too weak to harm the plant and thus the plant is actually healthy?",
          "votes": 1
        },
        {
          "id": 1159765,
          "postDate": "2021-01-19T13:00:31.723Z",
          "content": "<p>You are right! It can be that the disease is too small to be qualified as \"a disease\", in this case, the best solution would be to add a sixth label to differentiate the 100% healthy leaves and the leaves with irrelevant small spots. Because when the model finds the 'small disease' in healthy leaves, it gets confused and might assign it to the closest class, which in this case is class 0.</p>",
          "rawMarkdown": "You are right! It can be that the disease is too small to be qualified as \"a disease\", in this case, the best solution would be to add a sixth label to differentiate the 100% healthy leaves and the leaves with irrelevant small spots. Because when the model finds the 'small disease' in healthy leaves, it gets confused and might assign it to the closest class, which in this case is class 0.",
          "votes": 1
        },
        {
          "id": 1159768,
          "postDate": "2021-01-19T13:03:06.103Z",
          "content": "<p>Maybe another approach could be to use only disease classes (i.e. remove healthy class) for training and use this 4-classes model at inference time, filtering out predictions with low confidence for all diseases and mark them as healthy.</p>",
          "rawMarkdown": "Maybe another approach could be to use only disease classes (i.e. remove healthy class) for training and use this 4-classes model at inference time, filtering out predictions with low confidence for all diseases and mark them as healthy.",
          "votes": 2
        },
        {
          "id": 1159776,
          "postDate": "2021-01-19T13:10:43.183Z",
          "content": "<p>Great idea :) I will try this and see if it helps to improve the accuracy of class 0, if it does, then there is a potential for post-processing to correct the class 0 'predictions' predicted as healthy.</p>",
          "rawMarkdown": "Great idea :) I will try this and see if it helps to improve the accuracy of class 0, if it does, then there is a potential for post-processing to correct the class 0 'predictions' predicted as healthy."
        },
        {
          "id": 1160584,
          "postDate": "2021-01-20T02:59:41.437Z",
          "content": "<p><a href=\"https://www.kaggle.com/LAZCoder\" target=\"_blank\">@LAZCoder</a> Looking at the dataset, label 4 seems to be filled with many yellow leaves that look diseased. Some look very similar to the diseased labeled leaves. Based off this, I think that the model may give a high confidence for a disease but it would still be interesting to see what happens.</p>",
          "rawMarkdown": "@LAZCoder Looking at the dataset, label 4 seems to be filled with many yellow leaves that look diseased. Some look very similar to the diseased labeled leaves. Based off this, I think that the model may give a high confidence for a disease but it would still be interesting to see what happens."
        },
        {
          "id": 1160593,
          "postDate": "2021-01-20T03:08:50.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> For your idea, would we then have to relabel the training set, putting the noisy healthy leaves as label 5?</p>",
          "rawMarkdown": "@amiiiney For your idea, would we then have to relabel the training set, putting the noisy healthy leaves as label 5?"
        },
        {
          "id": 1161274,
          "postDate": "2021-01-20T13:11:01.953Z",
          "content": "<p><a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> Exactly! I am thinking to relabel the noisy healthy leaves as number 5, hoping this new label will improve the accuracy of class 0 (66%) since the noisy healthy images will be predicted as 5 instead of 0.<br>\nAfter inference, you can post-process and replace 5 by 4</p>\n<pre><code>test['label'].replace(5, 4)\n</code></pre>\n<p>The challenge is to manually relabel more than 2000 images, this will take some time :)</p>",
          "rawMarkdown": "@ayu055 Exactly! I am thinking to relabel the noisy healthy leaves as number 5, hoping this new label will improve the accuracy of class 0 (66%) since the noisy healthy images will be predicted as 5 instead of 0.\nAfter inference, you can post-process and replace 5 by 4\n```\ntest['label'].replace(5, 4)\n```\nThe challenge is to manually relabel more than 2000 images, this will take some time :)"
        },
        {
          "id": 1161384,
          "postDate": "2021-01-20T14:37:49.627Z",
          "content": "<p>Let us know if it works! I've been stuck at 0.901 - 0.902 for a long time :(</p>",
          "rawMarkdown": "Let us know if it works! I've been stuck at 0.901 - 0.902 for a long time :(",
          "votes": 1
        }
      ]
    },
    {
      "id": 1157855,
      "postDate": "2021-01-18T07:56:33.283Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1184175,
      "author_name": "Gena",
      "author_url": "",
      "post_date": "2021-02-03T12:23:44.437000",
      "content": "<p>One more confusion matrix for B6 and images resized to 384x384:<br>\n<img src=\"https://i.ibb.co/LxPxB3F/image.png\" alt=\"\"></p>\n<p>In my case, the network mainly has a hard time recognizing healthy images, with the largest confusion to class 0</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1157525,
      "author_name": "mknzfr",
      "author_url": "",
      "post_date": "2021-01-18T01:31:30.467000",
      "content": "<p>my classification report:<br>\nprecision    recall  f1-score   support<br>\n           0       0.73      0.68      0.70       108<br>\n           1       0.84      0.83      0.83       219<br>\n           2       0.81      0.79      0.80       239<br>\n           3       0.96      0.98      0.97      1315<br>\n           4       0.78      0.74      0.76       258</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1157643,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-18T04:12:16.943000",
          "content": "<p>Thanks, I added the P,R f1-score as well</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1157522,
      "author_name": "Amin",
      "author_url": "",
      "post_date": "2021-01-18T01:21:35.833000",
      "content": "<p>This analysis confirms what was shared in this discussion topic:  <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211853\" target=\"_blank\">Bit worried about the usefulness of this competition</a>. The model is very likely to mistake for healthy because a lot of the healthy images are mislabeled. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1157621,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "2021-01-18T03:56:12.353000",
          "content": "<p>Are we sure that these images are mislabeled? I know that many of the label 4 images seem to be yellow in color and have a disease but perhaps the yellow color is due to a natural lifecycle of the leaf or environmental factors?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1159558,
          "author_name": "LAZCoder",
          "author_url": "",
          "post_date": "2021-01-19T11:04:04.503000",
          "content": "<p>I agree with <a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> and I would like to give my point of view on what we consider \"noise\".</p>\n<p>I think that if an image shows a plant/leaf with symptoms in the early stage of disease, it could have been labeled (or mis-labeled) as healthy for 2 reasons: the more obvious is by mistake, but could it be due to the fact that the disease is too weak to harm the plant and thus the plant is actually healthy?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1159765,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2021-01-19T13:00:31.723000",
          "content": "<p>You are right! It can be that the disease is too small to be qualified as \"a disease\", in this case, the best solution would be to add a sixth label to differentiate the 100% healthy leaves and the leaves with irrelevant small spots. Because when the model finds the 'small disease' in healthy leaves, it gets confused and might assign it to the closest class, which in this case is class 0.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1159768,
          "author_name": "LAZCoder",
          "author_url": "",
          "post_date": "2021-01-19T13:03:06.103000",
          "content": "<p>Maybe another approach could be to use only disease classes (i.e. remove healthy class) for training and use this 4-classes model at inference time, filtering out predictions with low confidence for all diseases and mark them as healthy.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1159776,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2021-01-19T13:10:43.183000",
          "content": "<p>Great idea :) I will try this and see if it helps to improve the accuracy of class 0, if it does, then there is a potential for post-processing to correct the class 0 'predictions' predicted as healthy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1160584,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "2021-01-20T02:59:41.437000",
          "content": "<p><a href=\"https://www.kaggle.com/LAZCoder\" target=\"_blank\">@LAZCoder</a> Looking at the dataset, label 4 seems to be filled with many yellow leaves that look diseased. Some look very similar to the diseased labeled leaves. Based off this, I think that the model may give a high confidence for a disease but it would still be interesting to see what happens.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1160593,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "2021-01-20T03:08:50.927000",
          "content": "<p><a href=\"https://www.kaggle.com/amiiiney\" target=\"_blank\">@amiiiney</a> For your idea, would we then have to relabel the training set, putting the noisy healthy leaves as label 5?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161274,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2021-01-20T13:11:01.953000",
          "content": "<p><a href=\"https://www.kaggle.com/ayu055\" target=\"_blank\">@ayu055</a> Exactly! I am thinking to relabel the noisy healthy leaves as number 5, hoping this new label will improve the accuracy of class 0 (66%) since the noisy healthy images will be predicted as 5 instead of 0.<br>\nAfter inference, you can post-process and replace 5 by 4</p>\n<pre><code>test['label'].replace(5, 4)\n</code></pre>\n<p>The challenge is to manually relabel more than 2000 images, this will take some time :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161384,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "2021-01-20T14:37:49.627000",
          "content": "<p>Let us know if it works! I've been stuck at 0.901 - 0.902 for a long time :(</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1157855,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-18T07:56:33.283000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1157513": "Notebook:\nhttps://www.kaggle.com/kyoshioka47/analyze-your-model-performance-by-confusion-matrix\n\nA typical way to analyze your model performance is using the confusion matrix.\nAfter getting the 5-fold prediction results of my CV 0.89, LB 0.9 model, you can get the confusion matrix by\n\n```python\nfrom sklearn.metrics import confusion_matrix\ncm = confusion_matrix(df_train[\"label\"], logits.argmax(1))\n\nprint(cm)\n\n[[  719    59    22    49   238]\n [  108  1751    53    84   193]\n [   33    50  1924   251   128]\n [   36    71   199 12748   104]\n [  207   131   135   192  1912]]\n\n```\n\nInterestingly, the accuracy is very high for class-3 but very low for others.\n\n```\nfor i, val in enumerate(cm):\n    print(\"for class {}: accuracy: {}\".format(i, val[i]/sum(val)*100))\n\nfor class 0: accuracy: 66.14535418583257\nfor class 1: accuracy: 79.99086340794884\nfor class 2: accuracy: 80.63704945515508\nfor class 3: accuracy: 96.88402492780058\nfor class 4: accuracy: 74.19480015521924\n\n```\n\nEspecially for class 0, it is **very** likely to mistake for healthy!\n\n```\nfor i, val in enumerate(cm[:-1]):\n    print(\"for class {}: possibility to mistake for healthy: {}\".format(i, val[4]/val[i]*100))\n\nfor class 0: possibility to mistake for healthy: 33.101529902642554\nfor class 1: possibility to mistake for healthy: 11.02227298686465\nfor class 2: possibility to mistake for healthy: 6.652806652806653\nfor class 3: possibility to mistake for healthy: 0.815814245371823\n```\n\n```\nfrom sklearn.metrics import classification_report\n\nprint(classification_report(df_train[\"label\"], logits.argmax(1)))\n\nprecision    recall  f1-score   support\n\n           0       0.65      0.66      0.66      1087\n           1       0.85      0.80      0.82      2189\n           2       0.82      0.81      0.82      2386\n           3       0.96      0.97      0.96     13158\n           4       0.74      0.74      0.74      2577\n\n    accuracy                           0.89     21397\n```",
    "1184175": "One more confusion matrix for B6 and images resized to 384x384:\n![](https://i.ibb.co/LxPxB3F/image.png)\n\nIn my case, the network mainly has a hard time recognizing healthy images, with the largest confusion to class 0",
    "1157525": "my classification report:\nprecision    recall  f1-score   support\n           0       0.73      0.68      0.70       108\n           1       0.84      0.83      0.83       219\n           2       0.81      0.79      0.80       239\n           3       0.96      0.98      0.97      1315\n           4       0.78      0.74      0.76       258",
    "1157522": "This analysis confirms what was shared in this discussion topic:  [Bit worried about the usefulness of this competition](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211853). The model is very likely to mistake for healthy because a lot of the healthy images are mislabeled. ",
    "1157855": ""
  }
}