{
  "id": 199902,
  "title": "External data: Mendeley Leaves",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199902",
  "author_name": "",
  "post_date": "2020-11-27T21:53:13.162232300Z",
  "votes": 39,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Since <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/overview/code-requirements\" target=\"_blank\">Code Requirements</a> says the following:</p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>\n<p>I've decided to search for a similar datasets wich can be used to encrease performance in this competition.<br>\nAnd I've found this very good <a href=\"https://data.mendeley.com/datasets/hb74ynkjcn/4\" target=\"_blank\">Database of Leaf Images</a>.</p>\n<p>I've downloaded all of the images, resized them to 1200x800 and created a csv file with images filenames and labels.</p>\n<p>You can find this dataset here: <a href=\"https://www.kaggle.com/nroman/mendeley-leaves\" target=\"_blank\">https://www.kaggle.com/nroman/mendeley-leaves</a><br>\nIt contains 4335 images.</p>\n<p>As you can notice this dataset has a binary label: leaves are either healthy or diseased. So how can one use this to benefit in this multilabel competition? Well, I see 2 ways:</p>\n<ol>\n<li>First traing a net for binary classification using <a href=\"https://www.kaggle.com/nroman/mendeley-leaves\" target=\"_blank\">Mendeley Leaves Dataset</a>. Then replace your network's head and use it to train on this competitions data. I've already tried this and got some improvements.</li>\n<li>Extend existing dataset with images of healthy leaves as class 4. </li>\n</ol>",
  "messages": [
    {
      "id": "1093644",
      "postDate": "11/27/2020 21:53:13",
      "content": "<p>Since <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/overview/code-requirements\" target=\"_blank\">Code Requirements</a> says the following:</p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>\n<p>I've decided to search for a similar datasets wich can be used to encrease performance in this competition.<br>\nAnd I've found this very good <a href=\"https://data.mendeley.com/datasets/hb74ynkjcn/4\" target=\"_blank\">Database of Leaf Images</a>.</p>\n<p>I've downloaded all of the images, resized them to 1200x800 and created a csv file with images filenames and labels.</p>\n<p>You can find this dataset here: <a href=\"https://www.kaggle.com/nroman/mendeley-leaves\" target=\"_blank\">https://www.kaggle.com/nroman/mendeley-leaves</a><br>\nIt contains 4335 images.</p>\n<p>As you can notice this dataset has a binary label: leaves are either healthy or diseased. So how can one use this to benefit in this multilabel competition? Well, I see 2 ways:</p>\n<ol>\n<li>First traing a net for binary classification using <a href=\"https://www.kaggle.com/nroman/mendeley-leaves\" target=\"_blank\">Mendeley Leaves Dataset</a>. Then replace your network's head and use it to train on this competitions data. I've already tried this and got some improvements.</li>\n<li>Extend existing dataset with images of healthy leaves as class 4. </li>\n</ol>",
      "rawMarkdown": "Since [Code Requirements](https://www.kaggle.com/c/cassava-leaf-disease-classification/overview/code-requirements) says the following:\n\n> Freely & publicly available external data is allowed, including pre-trained models\n\nI've decided to search for a similar datasets wich can be used to encrease performance in this competition.\nAnd I've found this very good [Database of Leaf Images](https://data.mendeley.com/datasets/hb74ynkjcn/4).\n\nI've downloaded all of the images, resized them to 1200x800 and created a csv file with images filenames and labels.\n\nYou can find this dataset here: https://www.kaggle.com/nroman/mendeley-leaves\nIt contains 4335 images.\n\nAs you can notice this dataset has a binary label: leaves are either healthy or diseased. So how can one use this to benefit in this multilabel competition? Well, I see 2 ways:\n1. First traing a net for binary classification using [Mendeley Leaves Dataset](https://www.kaggle.com/nroman/mendeley-leaves). Then replace your network's head and use it to train on this competitions data. I've already tried this and got some improvements.\n2. Extend existing dataset with images of healthy leaves as class 4.",
      "votes": null
    },
    {
      "id": "1096180",
      "postDate": "11/30/2020 09:36:07",
      "content": "<p>Thanks for sharing! I read the description of external data and didn't find Cassava among plants here. So, did you use all of the available plants (Mango, Arjun, Alstonia Scholaris, Guava, Bael, Jamun, Jatropha, Pongamia Pinnata, Basil, Pomegranate, Lemon, and Chinar) or you chose a subset? And does it mean that plant type doesn't matter for disease detection?</p>\n<p>And can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?</p>",
      "rawMarkdown": "Thanks for sharing! I read the description of external data and didn't find Cassava among plants here. So, did you use all of the available plants (Mango, Arjun, Alstonia Scholaris, Guava, Bael, Jamun, Jatropha, Pongamia Pinnata, Basil, Pomegranate, Lemon, and Chinar) or you chose a subset? And does it mean that plant type doesn't matter for disease detection?\n\nAnd can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?",
      "votes": null
    },
    {
      "id": "1096270",
      "postDate": "11/30/2020 11:14:42",
      "content": "<blockquote>\n  <p>didn't find Cassava among plants here</p>\n</blockquote>\n<p>Yes, those are the pictures of leaves of a different plants but it still makes sense to use by several reasons. And one of them is that we are focusion on distinguishing between healthy leaves and deseased ones. So adding a portion of healthy leaves will help a model to better undestand what kind of leaf can be called 'healthy'.</p>\n<blockquote>\n  <p>So, did you use all of the available plants</p>\n</blockquote>\n<p>Yes</p>\n<blockquote>\n  <p>And does it mean that plant type doesn't matter for disease detection?</p>\n</blockquote>\n<p>If you use pictures for pretraining a model which you are planning to train on a concrete dataset of this competition then it is safe to say \"yes, plant type doesn't matter\".</p>\n<blockquote>\n  <p>And can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?</p>\n</blockquote>\n<p>This is not a trick at all. When you are using any pretrained model, say resnet from torchvision, you are basically using a model that was trained on imagenet pictures, which includes a lot of different classes and plants are among them. In other words - those models already know how to distinguish between plants and other objects. And using a pretrained model to train it on this competition data allows you to get a better performance.<br>\nSo why don't make another step forward and not pretrain a model that specializes on difference between healthy and not healthy plants? And then use those weights to train it on cassava leaves?</p>\n<blockquote>\n  <p>Did it significantly contribute to the accuracy score?</p>\n</blockquote>\n<p>First point gave me +0.001 accuracy. Second point alone +0.004. Combination of both of these points improve the score even more.</p>",
      "rawMarkdown": "> didn't find Cassava among plants here\n\nYes, those are the pictures of leaves of a different plants but it still makes sense to use by several reasons. And one of them is that we are focusion on distinguishing between healthy leaves and deseased ones. So adding a portion of healthy leaves will help a model to better undestand what kind of leaf can be called 'healthy'.\n\n> So, did you use all of the available plants\n\nYes\n\n> And does it mean that plant type doesn't matter for disease detection?\n\nIf you use pictures for pretraining a model which you are planning to train on a concrete dataset of this competition then it is safe to say \"yes, plant type doesn't matter\".\n\n> And can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?\n\nThis is not a trick at all. When you are using any pretrained model, say resnet from torchvision, you are basically using a model that was trained on imagenet pictures, which includes a lot of different classes and plants are among them. In other words - those models already know how to distinguish between plants and other objects. And using a pretrained model to train it on this competition data allows you to get a better performance.\nSo why don't make another step forward and not pretrain a model that specializes on difference between healthy and not healthy plants? And then use those weights to train it on cassava leaves?\n\n> Did it significantly contribute to the accuracy score?\n\nFirst point gave me +0.001 accuracy. Second point alone +0.004. Combination of both of these points improve the score even more.",
      "votes": null
    },
    {
      "id": "1096909",
      "postDate": "11/30/2020 21:41:10",
      "content": "<p>Thanks a lot for detailed explanation!</p>",
      "rawMarkdown": "Thanks a lot for detailed explanation!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1096180,
      "author_name": "lashby",
      "author_url": "",
      "post_date": "11/30/2020 09:36:07",
      "content": "<p>Thanks for sharing! I read the description of external data and didn't find Cassava among plants here. So, did you use all of the available plants (Mango, Arjun, Alstonia Scholaris, Guava, Bael, Jamun, Jatropha, Pongamia Pinnata, Basil, Pomegranate, Lemon, and Chinar) or you chose a subset? And does it mean that plant type doesn't matter for disease detection?</p>\n<p>And can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1096270,
          "author_name": "nroman",
          "author_url": "",
          "post_date": "11/30/2020 11:14:42",
          "content": "<blockquote>\n  <p>didn't find Cassava among plants here</p>\n</blockquote>\n<p>Yes, those are the pictures of leaves of a different plants but it still makes sense to use by several reasons. And one of them is that we are focusion on distinguishing between healthy leaves and deseased ones. So adding a portion of healthy leaves will help a model to better undestand what kind of leaf can be called 'healthy'.</p>\n<blockquote>\n  <p>So, did you use all of the available plants</p>\n</blockquote>\n<p>Yes</p>\n<blockquote>\n  <p>And does it mean that plant type doesn't matter for disease detection?</p>\n</blockquote>\n<p>If you use pictures for pretraining a model which you are planning to train on a concrete dataset of this competition then it is safe to say \"yes, plant type doesn't matter\".</p>\n<blockquote>\n  <p>And can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?</p>\n</blockquote>\n<p>This is not a trick at all. When you are using any pretrained model, say resnet from torchvision, you are basically using a model that was trained on imagenet pictures, which includes a lot of different classes and plants are among them. In other words - those models already know how to distinguish between plants and other objects. And using a pretrained model to train it on this competition data allows you to get a better performance.<br>\nSo why don't make another step forward and not pretrain a model that specializes on difference between healthy and not healthy plants? And then use those weights to train it on cassava leaves?</p>\n<blockquote>\n  <p>Did it significantly contribute to the accuracy score?</p>\n</blockquote>\n<p>First point gave me +0.001 accuracy. Second point alone +0.004. Combination of both of these points improve the score even more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096909,
          "author_name": "lashby",
          "author_url": "",
          "post_date": "11/30/2020 21:41:10",
          "content": "<p>Thanks a lot for detailed explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1093644": "Since [Code Requirements](https://www.kaggle.com/c/cassava-leaf-disease-classification/overview/code-requirements) says the following:\n\n> Freely & publicly available external data is allowed, including pre-trained models\n\nI've decided to search for a similar datasets wich can be used to encrease performance in this competition.\nAnd I've found this very good [Database of Leaf Images](https://data.mendeley.com/datasets/hb74ynkjcn/4).\n\nI've downloaded all of the images, resized them to 1200x800 and created a csv file with images filenames and labels.\n\nYou can find this dataset here: https://www.kaggle.com/nroman/mendeley-leaves\nIt contains 4335 images.\n\nAs you can notice this dataset has a binary label: leaves are either healthy or diseased. So how can one use this to benefit in this multilabel competition? Well, I see 2 ways:\n1. First traing a net for binary classification using [Mendeley Leaves Dataset](https://www.kaggle.com/nroman/mendeley-leaves). Then replace your network's head and use it to train on this competitions data. I've already tried this and got some improvements.\n2. Extend existing dataset with images of healthy leaves as class 4.",
    "1096180": "Thanks for sharing! I read the description of external data and didn't find Cassava among plants here. So, did you use all of the available plants (Mango, Arjun, Alstonia Scholaris, Guava, Bael, Jamun, Jatropha, Pongamia Pinnata, Basil, Pomegranate, Lemon, and Chinar) or you chose a subset? And does it mean that plant type doesn't matter for disease detection?\n\nAnd can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?",
    "1096270": "> didn't find Cassava among plants here\n\nYes, those are the pictures of leaves of a different plants but it still makes sense to use by several reasons. And one of them is that we are focusion on distinguishing between healthy leaves and deseased ones. So adding a portion of healthy leaves will help a model to better undestand what kind of leaf can be called 'healthy'.\n\n> So, did you use all of the available plants\n\nYes\n\n> And does it mean that plant type doesn't matter for disease detection?\n\nIf you use pictures for pretraining a model which you are planning to train on a concrete dataset of this competition then it is safe to say \"yes, plant type doesn't matter\".\n\n> And can you please tell us more about the finetuning trick, which you've described at 1st point? Did it significantly contribute to the accuracy score?\n\nThis is not a trick at all. When you are using any pretrained model, say resnet from torchvision, you are basically using a model that was trained on imagenet pictures, which includes a lot of different classes and plants are among them. In other words - those models already know how to distinguish between plants and other objects. And using a pretrained model to train it on this competition data allows you to get a better performance.\nSo why don't make another step forward and not pretrain a model that specializes on difference between healthy and not healthy plants? And then use those weights to train it on cassava leaves?\n\n> Did it significantly contribute to the accuracy score?\n\nFirst point gave me +0.001 accuracy. Second point alone +0.004. Combination of both of these points improve the score even more.",
    "1096909": "Thanks a lot for detailed explanation!"
  },
  "source": "meta"
}