{
  "id": 202247,
  "title": "Key Classes in this Competition",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202247",
  "author_name": "",
  "post_date": "2020-12-09T04:30:37.772364100Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Here I am sharing confusion matrix from my best model (LB 0.901). We know there are many noisy/mislabeled images, and it's probably why we have lot of false negatives and false positives for \"Healthy\" class. Also, you can notice, 0.21 error rate for CBB which is confused by healthy. This again might be due to mislabel or bad performance for detecting CBB. Another high confusion is between CGM - CMD, around 0.1 error rate. Remaining parts are almost all below 0.05. </p>\n<p>Somehow targeting these issues will be key factors for performing well in this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F59ae5ca05fc1781d8904b56f2e3d6296%2FScreen%20Shot%202020-12-08%20at%2010.47.18%20PM.png?generation=1607489261668986&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Fca4efa2ad3403957bcddf3ea6084a6bf%2FScreen%20Shot%202020-12-08%20at%2010.23.07%20PM.png?generation=1607487823245242&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1106733",
      "postDate": "12/09/2020 04:30:37",
      "content": "<p>Here I am sharing confusion matrix from my best model (LB 0.901). We know there are many noisy/mislabeled images, and it's probably why we have lot of false negatives and false positives for \"Healthy\" class. Also, you can notice, 0.21 error rate for CBB which is confused by healthy. This again might be due to mislabel or bad performance for detecting CBB. Another high confusion is between CGM - CMD, around 0.1 error rate. Remaining parts are almost all below 0.05. </p>\n<p>Somehow targeting these issues will be key factors for performing well in this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F59ae5ca05fc1781d8904b56f2e3d6296%2FScreen%20Shot%202020-12-08%20at%2010.47.18%20PM.png?generation=1607489261668986&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Fca4efa2ad3403957bcddf3ea6084a6bf%2FScreen%20Shot%202020-12-08%20at%2010.23.07%20PM.png?generation=1607487823245242&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here I am sharing confusion matrix from my best model (LB 0.901). We know there are many noisy/mislabeled images, and it's probably why we have lot of false negatives and false positives for \"Healthy\" class. Also, you can notice, 0.21 error rate for CBB which is confused by healthy. This again might be due to mislabel or bad performance for detecting CBB. Another high confusion is between CGM - CMD, around 0.1 error rate. Remaining parts are almost all below 0.05. \n\nSomehow targeting these issues will be key factors for performing well in this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F59ae5ca05fc1781d8904b56f2e3d6296%2FScreen%20Shot%202020-12-08%20at%2010.47.18%20PM.png?generation=1607489261668986&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Fca4efa2ad3403957bcddf3ea6084a6bf%2FScreen%20Shot%202020-12-08%20at%2010.23.07%20PM.png?generation=1607487823245242&alt=media)",
      "votes": null
    },
    {
      "id": "1106739",
      "postDate": "12/09/2020 04:39:26",
      "content": "<p>thanks for the confusion matrix. The fact the confusion is not symmetric is good news.<br>\nyou can make confusion matrix by different score band, e.g. one matrix for p&gt;0.9, another for 0.8&gt;p&gt;0.9, etc …<br>\nif you can make confusion matrix by leaf size, etc</p>\n<p>if you can probe the distrubtution of each class in the hidden set (at least for public), you can post adjust your prediction score to gain a few 0.001 scores.</p>",
      "rawMarkdown": "thanks for the confusion matrix. The fact the confusion is not symmetric is good news.\nyou can make confusion matrix by different score band, e.g. one matrix for p>0.9, another for 0.8>p>0.9, etc ...\nif you can make confusion matrix by leaf size, etc\n\nif you can probe the distrubtution of each class in the hidden set (at least for public), you can post adjust your prediction score to gain a few 0.001 scores.",
      "votes": null
    },
    {
      "id": "1106907",
      "postDate": "12/09/2020 08:01:37",
      "content": "<p>Thanks for sharing.<br>\nBut I have very basic question. <br>\nWith this confusion matrix, what can we do? I understand what you wrote but I can't catch idea.<br>\nWe need to do something  that makes predict Healthy class well?</p>",
      "rawMarkdown": "Thanks for sharing.\nBut I have very basic question. \nWith this confusion matrix, what can we do? I understand what you wrote but I can't catch idea.\nWe need to do something  that makes predict Healthy class well?",
      "votes": null
    },
    {
      "id": "1106933",
      "postDate": "12/09/2020 08:27:13",
      "content": "<p>Because of the problems mentioned above, <br>\npublic lb score can be significantly different, even if the CV is the same or a little lower.</p>\n<p>This points also gives the possibility of improving performance using ensembles.<br>\n<a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a></p>",
      "rawMarkdown": "Because of the problems mentioned above, \npublic lb score can be significantly different, even if the CV is the same or a little lower.\n\nThis points also gives the possibility of improving performance using ensembles.\n@keremt",
      "votes": null
    },
    {
      "id": "1107611",
      "postDate": "12/09/2020 20:10:33",
      "content": "<p>There are multiple things we can \"try\". Others also mentioned a few here; </p>\n<ul>\n<li>You can try to do post processing based on model predictions and the actual LB distribution</li>\n<li>You can try to fix mislabeled samples in the most confused class pairs</li>\n<li>You can gather more data for the classes your model perform bad</li>\n</ul>",
      "rawMarkdown": "There are multiple things we can \"try\". Others also mentioned a few here; \n\n- You can try to do post processing based on model predictions and the actual LB distribution\n- You can try to fix mislabeled samples in the most confused class pairs\n- You can gather more data for the classes your model perform bad",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1106739,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/09/2020 04:39:26",
      "content": "<p>thanks for the confusion matrix. The fact the confusion is not symmetric is good news.<br>\nyou can make confusion matrix by different score band, e.g. one matrix for p&gt;0.9, another for 0.8&gt;p&gt;0.9, etc …<br>\nif you can make confusion matrix by leaf size, etc</p>\n<p>if you can probe the distrubtution of each class in the hidden set (at least for public), you can post adjust your prediction score to gain a few 0.001 scores.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1106907,
      "author_name": "jinkyh",
      "author_url": "",
      "post_date": "12/09/2020 08:01:37",
      "content": "<p>Thanks for sharing.<br>\nBut I have very basic question. <br>\nWith this confusion matrix, what can we do? I understand what you wrote but I can't catch idea.<br>\nWe need to do something  that makes predict Healthy class well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107611,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "12/09/2020 20:10:33",
          "content": "<p>There are multiple things we can \"try\". Others also mentioned a few here; </p>\n<ul>\n<li>You can try to do post processing based on model predictions and the actual LB distribution</li>\n<li>You can try to fix mislabeled samples in the most confused class pairs</li>\n<li>You can gather more data for the classes your model perform bad</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1106933,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "12/09/2020 08:27:13",
      "content": "<p>Because of the problems mentioned above, <br>\npublic lb score can be significantly different, even if the CV is the same or a little lower.</p>\n<p>This points also gives the possibility of improving performance using ensembles.<br>\n<a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1106733": "Here I am sharing confusion matrix from my best model (LB 0.901). We know there are many noisy/mislabeled images, and it's probably why we have lot of false negatives and false positives for \"Healthy\" class. Also, you can notice, 0.21 error rate for CBB which is confused by healthy. This again might be due to mislabel or bad performance for detecting CBB. Another high confusion is between CGM - CMD, around 0.1 error rate. Remaining parts are almost all below 0.05. \n\nSomehow targeting these issues will be key factors for performing well in this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F59ae5ca05fc1781d8904b56f2e3d6296%2FScreen%20Shot%202020-12-08%20at%2010.47.18%20PM.png?generation=1607489261668986&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Fca4efa2ad3403957bcddf3ea6084a6bf%2FScreen%20Shot%202020-12-08%20at%2010.23.07%20PM.png?generation=1607487823245242&alt=media)",
    "1106739": "thanks for the confusion matrix. The fact the confusion is not symmetric is good news.\nyou can make confusion matrix by different score band, e.g. one matrix for p>0.9, another for 0.8>p>0.9, etc ...\nif you can make confusion matrix by leaf size, etc\n\nif you can probe the distrubtution of each class in the hidden set (at least for public), you can post adjust your prediction score to gain a few 0.001 scores.",
    "1106907": "Thanks for sharing.\nBut I have very basic question. \nWith this confusion matrix, what can we do? I understand what you wrote but I can't catch idea.\nWe need to do something  that makes predict Healthy class well?",
    "1106933": "Because of the problems mentioned above, \npublic lb score can be significantly different, even if the CV is the same or a little lower.\n\nThis points also gives the possibility of improving performance using ensembles.\n@keremt",
    "1107611": "There are multiple things we can \"try\". Others also mentioned a few here; \n\n- You can try to do post processing based on model predictions and the actual LB distribution\n- You can try to fix mislabeled samples in the most confused class pairs\n- You can gather more data for the classes your model perform bad"
  },
  "source": "meta"
}