{
  "id": 207780,
  "title": "New approach for ... Cassava Leaf Disease Classification?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207780",
  "author_name": "",
  "post_date": "2020-12-31T08:22:11.191309400Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>So far for this problem, almost every notebook I have seen follows a very generic template approach.<br>\nThat is:  </p>\n<ol>\n<li>load the image data </li>\n<li>apply some image augmentation techniques</li>\n<li>Load a pre-trained model. Add or subtract few layers on top. </li>\n<li>Train the model and get the output.</li>\n</ol>\n<p>Note that by no means I am belittling this approach. I am just interested in learning newer approaches to the problem.</p>\n<p>So, is it possible to do some image segmentation first maybe in RGB, HSV, or some other color space, and then proceed with the model training and classification? Or maybe some entirely different approach that would prove helpful in the problem. Or maybe some small image processing technique that can prove useful here?<br>\nPlease share your valuable thoughts and suggestions! </p>",
  "messages": [
    {
      "id": "1133453",
      "postDate": "12/31/2020 08:22:11",
      "content": "<p>So far for this problem, almost every notebook I have seen follows a very generic template approach.<br>\nThat is:  </p>\n<ol>\n<li>load the image data </li>\n<li>apply some image augmentation techniques</li>\n<li>Load a pre-trained model. Add or subtract few layers on top. </li>\n<li>Train the model and get the output.</li>\n</ol>\n<p>Note that by no means I am belittling this approach. I am just interested in learning newer approaches to the problem.</p>\n<p>So, is it possible to do some image segmentation first maybe in RGB, HSV, or some other color space, and then proceed with the model training and classification? Or maybe some entirely different approach that would prove helpful in the problem. Or maybe some small image processing technique that can prove useful here?<br>\nPlease share your valuable thoughts and suggestions! </p>",
      "rawMarkdown": "So far for this problem, almost every notebook I have seen follows a very generic template approach.\nThat is:  \n1. load the image data \n2. apply some image augmentation techniques\n3. Load a pre-trained model. Add or subtract few layers on top. \n4. Train the model and get the output.\n\nNote that by no means I am belittling this approach. I am just interested in learning newer approaches to the problem.\n\nSo, is it possible to do some image segmentation first maybe in RGB, HSV, or some other color space, and then proceed with the model training and classification? Or maybe some entirely different approach that would prove helpful in the problem. Or maybe some small image processing technique that can prove useful here?\nPlease share your valuable thoughts and suggestions!",
      "votes": null
    },
    {
      "id": "1133482",
      "postDate": "12/31/2020 08:55:11",
      "content": "<p>There are 5 categories in total. Let's count the number of images in each category of these 5 categories, and sort them from high to low. The category with the highest number corresponds to other categories. This is a two-class classification problem. First eliminate the most frequently occurring ones, and then perform multi-classification of the rest. Since the remaining four categories are more evenly distributed, multi-classification can be used. In short, the problem of a small category and uneven distribution of data Converting to a two-category + a multi-category, whether this can balance the relationship between categories, I only temporarily think about so much.</p>",
      "rawMarkdown": "There are 5 categories in total. Let's count the number of images in each category of these 5 categories, and sort them from high to low. The category with the highest number corresponds to other categories. This is a two-class classification problem. First eliminate the most frequently occurring ones, and then perform multi-classification of the rest. Since the remaining four categories are more evenly distributed, multi-classification can be used. In short, the problem of a small category and uneven distribution of data Converting to a two-category + a multi-category, whether this can balance the relationship between categories, I only temporarily think about so much.",
      "votes": null
    },
    {
      "id": "1133484",
      "postDate": "12/31/2020 08:57:04",
      "content": "<p>it can be transformed into a two-category (select the class with the most distribution and not that class) + a multi-class problem (select the remaining 4 categories)</p>",
      "rawMarkdown": "it can be transformed into a two-category (select the class with the most distribution and not that class) + a multi-class problem (select the remaining 4 categories)",
      "votes": null
    },
    {
      "id": "1133522",
      "postDate": "12/31/2020 09:30:55",
      "content": "<p>Only problem I see with this approach is that it will result in 2 cascaded classifiers. One for the two-category and second one for multi-category. This will result in the combined accuracy being the multiplication of the cascaded classifiers. <br>\nIn my experience it does not outperform the original classifier :).</p>",
      "rawMarkdown": "Only problem I see with this approach is that it will result in 2 cascaded classifiers. One for the two-category and second one for multi-category. This will result in the combined accuracy being the multiplication of the cascaded classifiers. \nIn my experience it does not outperform the original classifier :).",
      "votes": null
    },
    {
      "id": "1133528",
      "postDate": "12/31/2020 09:35:50",
      "content": "<p>I think this can achieve a category balance. You see that the number of Cassava Mosaic Disease (CMD) categories is more than half of the other four categories. Then you can perform two categories to exclude this category first, and then you can look at the other four categories. Tend to balance, so long to solve the problem of uneven categories. Anyway, I will experiment with this method tomorrow. My goal is to make the data more balanced.</p>",
      "rawMarkdown": "I think this can achieve a category balance. You see that the number of Cassava Mosaic Disease (CMD) categories is more than half of the other four categories. Then you can perform two categories to exclude this category first, and then you can look at the other four categories. Tend to balance, so long to solve the problem of uneven categories. Anyway, I will experiment with this method tomorrow. My goal is to make the data more balanced.",
      "votes": null
    },
    {
      "id": "1133602",
      "postDate": "12/31/2020 11:22:05",
      "content": "<p>Do let me know if you achieve see interesting results!</p>",
      "rawMarkdown": "Do let me know if you achieve see interesting results!",
      "votes": null
    },
    {
      "id": "1133919",
      "postDate": "12/31/2020 16:36:19",
      "content": "<p>All ideas are not made public before the end. </p>\n<p>Just because you don't see them on public notebooks dosn't mean people are not trying different things.</p>\n<p>You may try to do some searchs, read papers, read solutions from past competitions etc.  All these may help to learn new ideas and try them out.  </p>",
      "rawMarkdown": "All ideas are not made public before the end. \n\nJust because you don't see them on public notebooks dosn't mean people are not trying different things.\n\n\nYou may try to do some searchs, read papers, read solutions from past competitions etc.  All these may help to learn new ideas and try them out.",
      "votes": null
    },
    {
      "id": "1134911",
      "postDate": "01/01/2021 17:40:15",
      "content": "<p>I am new to the platform, so don't know some things. I guess you are right, it's good to wait for the competition to end. <br>\nThanks a lot for your feedback. </p>",
      "rawMarkdown": "I am new to the platform, so don't know some things. I guess you are right, it's good to wait for the competition to end. \nThanks a lot for your feedback.",
      "votes": null
    },
    {
      "id": "1186720",
      "postDate": "02/05/2021 02:27:45",
      "content": "<p>i started down a path of the 2-stage category split being \"Healthy\" and \"NotHealthy\" as the 1st stage, and then the 2nd stage was just the 4 NotHealthy categories.<br>\ni don't have the chops to do it all in one model, so i do them separately, save the weights, and then do a 3rd inference run.  takes like all day for 1 cycle.<br>\nthe results were actually promising but then i got discouraged by the label noise factor, and started working on some denoising approaches.<br>\nbut i must agree with <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> - i don't think it was outperforming a single stage model. </p>",
      "rawMarkdown": "i started down a path of the 2-stage category split being \"Healthy\" and \"NotHealthy\" as the 1st stage, and then the 2nd stage was just the 4 NotHealthy categories.\ni don't have the chops to do it all in one model, so i do them separately, save the weights, and then do a 3rd inference run.  takes like all day for 1 cycle.\nthe results were actually promising but then i got discouraged by the label noise factor, and started working on some denoising approaches.\nbut i must agree with @prvnkmr - i don't think it was outperforming a single stage model.",
      "votes": null
    },
    {
      "id": "1186730",
      "postDate": "02/05/2021 02:38:13",
      "content": "<p>i have not seen any discussion or notebooks that employ an 'active' custom transform approach, such as smartly choosing a specific center-point for cropping, for example.<br>\ni use randomcentercrop(512), and it actually improves things over just centercrop(512), but i have some ideas around ways to choose the center-points that might maximize the amount of 'leaf info' that is in a given cropped image.<br>\nbut i don't know how to code a custom transform in the transform chain.<br>\ni assume it's just a matter of digging into the module, and i'm really impressed with the albumentation module.<br>\n<a href=\"https://albumentations.ai/\" target=\"_blank\">https://albumentations.ai/</a> </p>\n<p>i think it may be possible to get a little better at ideal classification, but then the noisy label issues swamp any real progress beyond the noise rate, it seems.</p>",
      "rawMarkdown": "i have not seen any discussion or notebooks that employ an 'active' custom transform approach, such as smartly choosing a specific center-point for cropping, for example.\ni use randomcentercrop(512), and it actually improves things over just centercrop(512), but i have some ideas around ways to choose the center-points that might maximize the amount of 'leaf info' that is in a given cropped image.\nbut i don't know how to code a custom transform in the transform chain.\ni assume it's just a matter of digging into the module, and i'm really impressed with the albumentation module.\nhttps://albumentations.ai/ \n\ni think it may be possible to get a little better at ideal classification, but then the noisy label issues swamp any real progress beyond the noise rate, it seems.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1133482,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/31/2020 08:55:11",
      "content": "<p>There are 5 categories in total. Let's count the number of images in each category of these 5 categories, and sort them from high to low. The category with the highest number corresponds to other categories. This is a two-class classification problem. First eliminate the most frequently occurring ones, and then perform multi-classification of the rest. Since the remaining four categories are more evenly distributed, multi-classification can be used. In short, the problem of a small category and uneven distribution of data Converting to a two-category + a multi-category, whether this can balance the relationship between categories, I only temporarily think about so much.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1133522,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "12/31/2020 09:30:55",
          "content": "<p>Only problem I see with this approach is that it will result in 2 cascaded classifiers. One for the two-category and second one for multi-category. This will result in the combined accuracy being the multiplication of the cascaded classifiers. <br>\nIn my experience it does not outperform the original classifier :).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133528,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "12/31/2020 09:35:50",
          "content": "<p>I think this can achieve a category balance. You see that the number of Cassava Mosaic Disease (CMD) categories is more than half of the other four categories. Then you can perform two categories to exclude this category first, and then you can look at the other four categories. Tend to balance, so long to solve the problem of uneven categories. Anyway, I will experiment with this method tomorrow. My goal is to make the data more balanced.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133602,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "12/31/2020 11:22:05",
          "content": "<p>Do let me know if you achieve see interesting results!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1186720,
          "author_name": "daveccampbell",
          "author_url": "",
          "post_date": "02/05/2021 02:27:45",
          "content": "<p>i started down a path of the 2-stage category split being \"Healthy\" and \"NotHealthy\" as the 1st stage, and then the 2nd stage was just the 4 NotHealthy categories.<br>\ni don't have the chops to do it all in one model, so i do them separately, save the weights, and then do a 3rd inference run.  takes like all day for 1 cycle.<br>\nthe results were actually promising but then i got discouraged by the label noise factor, and started working on some denoising approaches.<br>\nbut i must agree with <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> - i don't think it was outperforming a single stage model. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1133484,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "12/31/2020 08:57:04",
      "content": "<p>it can be transformed into a two-category (select the class with the most distribution and not that class) + a multi-class problem (select the remaining 4 categories)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1133919,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "12/31/2020 16:36:19",
      "content": "<p>All ideas are not made public before the end. </p>\n<p>Just because you don't see them on public notebooks dosn't mean people are not trying different things.</p>\n<p>You may try to do some searchs, read papers, read solutions from past competitions etc.  All these may help to learn new ideas and try them out.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1134911,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "01/01/2021 17:40:15",
          "content": "<p>I am new to the platform, so don't know some things. I guess you are right, it's good to wait for the competition to end. <br>\nThanks a lot for your feedback. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1186730,
      "author_name": "daveccampbell",
      "author_url": "",
      "post_date": "02/05/2021 02:38:13",
      "content": "<p>i have not seen any discussion or notebooks that employ an 'active' custom transform approach, such as smartly choosing a specific center-point for cropping, for example.<br>\ni use randomcentercrop(512), and it actually improves things over just centercrop(512), but i have some ideas around ways to choose the center-points that might maximize the amount of 'leaf info' that is in a given cropped image.<br>\nbut i don't know how to code a custom transform in the transform chain.<br>\ni assume it's just a matter of digging into the module, and i'm really impressed with the albumentation module.<br>\n<a href=\"https://albumentations.ai/\" target=\"_blank\">https://albumentations.ai/</a> </p>\n<p>i think it may be possible to get a little better at ideal classification, but then the noisy label issues swamp any real progress beyond the noise rate, it seems.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1133453": "So far for this problem, almost every notebook I have seen follows a very generic template approach.\nThat is:  \n1. load the image data \n2. apply some image augmentation techniques\n3. Load a pre-trained model. Add or subtract few layers on top. \n4. Train the model and get the output.\n\nNote that by no means I am belittling this approach. I am just interested in learning newer approaches to the problem.\n\nSo, is it possible to do some image segmentation first maybe in RGB, HSV, or some other color space, and then proceed with the model training and classification? Or maybe some entirely different approach that would prove helpful in the problem. Or maybe some small image processing technique that can prove useful here?\nPlease share your valuable thoughts and suggestions!",
    "1133482": "There are 5 categories in total. Let's count the number of images in each category of these 5 categories, and sort them from high to low. The category with the highest number corresponds to other categories. This is a two-class classification problem. First eliminate the most frequently occurring ones, and then perform multi-classification of the rest. Since the remaining four categories are more evenly distributed, multi-classification can be used. In short, the problem of a small category and uneven distribution of data Converting to a two-category + a multi-category, whether this can balance the relationship between categories, I only temporarily think about so much.",
    "1133484": "it can be transformed into a two-category (select the class with the most distribution and not that class) + a multi-class problem (select the remaining 4 categories)",
    "1133522": "Only problem I see with this approach is that it will result in 2 cascaded classifiers. One for the two-category and second one for multi-category. This will result in the combined accuracy being the multiplication of the cascaded classifiers. \nIn my experience it does not outperform the original classifier :).",
    "1133528": "I think this can achieve a category balance. You see that the number of Cassava Mosaic Disease (CMD) categories is more than half of the other four categories. Then you can perform two categories to exclude this category first, and then you can look at the other four categories. Tend to balance, so long to solve the problem of uneven categories. Anyway, I will experiment with this method tomorrow. My goal is to make the data more balanced.",
    "1133602": "Do let me know if you achieve see interesting results!",
    "1133919": "All ideas are not made public before the end. \n\nJust because you don't see them on public notebooks dosn't mean people are not trying different things.\n\n\nYou may try to do some searchs, read papers, read solutions from past competitions etc.  All these may help to learn new ideas and try them out.",
    "1134911": "I am new to the platform, so don't know some things. I guess you are right, it's good to wait for the competition to end. \nThanks a lot for your feedback.",
    "1186720": "i started down a path of the 2-stage category split being \"Healthy\" and \"NotHealthy\" as the 1st stage, and then the 2nd stage was just the 4 NotHealthy categories.\ni don't have the chops to do it all in one model, so i do them separately, save the weights, and then do a 3rd inference run.  takes like all day for 1 cycle.\nthe results were actually promising but then i got discouraged by the label noise factor, and started working on some denoising approaches.\nbut i must agree with @prvnkmr - i don't think it was outperforming a single stage model.",
    "1186730": "i have not seen any discussion or notebooks that employ an 'active' custom transform approach, such as smartly choosing a specific center-point for cropping, for example.\ni use randomcentercrop(512), and it actually improves things over just centercrop(512), but i have some ideas around ways to choose the center-points that might maximize the amount of 'leaf info' that is in a given cropped image.\nbut i don't know how to code a custom transform in the transform chain.\ni assume it's just a matter of digging into the module, and i'm really impressed with the albumentation module.\nhttps://albumentations.ai/ \n\ni think it may be possible to get a little better at ideal classification, but then the noisy label issues swamp any real progress beyond the noise rate, it seems."
  },
  "source": "meta"
}