{
  "id": 308713,
  "title": "Summary of plant specimen classification outcomes from the previous literature ",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/308713",
  "author_name": "John Park",
  "post_date": "2022-02-19T20:57:43.973000",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Thank you for all the valuable information about summarizing the previous threads and listing potential relevant literature to this competition.</p>\n<p>Building upon <a href=\"https://www.kaggle.com/sytuannguyen\" target=\"_blank\">@sytuannguyen</a>'s discussion thread about <a href=\"https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307810\" target=\"_blank\">some relevant literature</a>, I would like to start this discussion thread to summarize all the previous literature on plant specimen classification. I hope this could serve as a reference point for you guys when building models for our competition. </p>\n<p>Focusing on the following questions:</p>\n<ul>\n<li>What dataset do they use: How many taxa to classify? How many training images and testing images?</li>\n<li>What are the models and the settings (backbone, loss, scheduler, etc.)?</li>\n<li>What is the accuracy of their model?</li>\n</ul>\n<p>To my knowledge, state-of-the-art methods are most likely from our previous competition (Please correct me if I am wrong), e.g., <a href=\"https://www.kaggle.com/c/herbarium-2020-fgvc7\" target=\"_blank\">herbarium 2020</a> and <a href=\"https://www.kaggle.com/c/herbarium-2021-fgvc8\" target=\"_blank\">herbarium 2021</a>. Therefore, for those of you who are running out of time, please refer to the paper about <a href=\"https://www.frontiersin.org/articles/10.3389/fpls.2021.787127/full\" target=\"_blank\">herbarium 2021</a>. It contains a detailed description of models, losses, and schedulers to achieve state-of-the-art performance for herbarium specimen classification. Unfortunately, we don't have a publication for herbarium 2020, but many winning teams posted their solutions on the <a href=\"https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307624\" target=\"_blank\">previous winning solutions</a>. Many thanks to <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> for the summary.</p>\n<h3>Summary of the previous plant specimen classification results.</h3>\n<p><a href=\"https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-016-0827-5\" target=\"_blank\">Computer vision applied to herbarium specimens of German trees: testing the future utility of the millions of herbarium specimen images for automated identification (Unger et al., 2016)</a></p>\n<ul>\n<li>26 most common taxa in Germany,  totaling 260 images for training. 0.72- 0.85 Top-1 accuracy<ul>\n<li>SVM with Fourier descriptors, multiple leaf shape parameters, weighted orientation histogram.</li></ul></li>\n</ul>\n<p><a href=\"https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-017-1014-z\" target=\"_blank\">Going deeper in the automated identification of herbarium specimens (Carranza-Rojas et al., 2017)</a></p>\n<ul>\n<li><p>255 taxa totaling 11K images. Top-1 accruacy: 0.745 Top-5 accuracy: 0.872</p>\n<ul>\n<li>input resolution: 256 by 256. </li>\n<li>best model pre-trained on imageNet and Herbarium1K dataset</li>\n<li>GoogleNet model that stacks inception layers</li>\n<li>Batch Normalization</li>\n<li>PReLU</li>\n<li>simple crop and resize data augmentation, default with Caffe.</li></ul></li>\n<li><p>1204 taxa totaling 292K images for training and 51K images for testing. Top-1 accuracy: 0.796, Top-5 accuracy:0.903</p>\n<ul>\n<li>Same setting</li></ul></li>\n</ul>\n<p><a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005993\" target=\"_blank\">Automated plant species identification - Trends and future directions (Waldchen et al., 2018)</a> <br></p>\n<ul>\n<li>No results found for plant specimen identification</li>\n</ul>\n<p><a href=\"https://bsapubs.onlinelibrary.wiley.com/doi/full/10.1002/aps3.11365\" target=\"_blank\">An algorithm competition for automatic species identification from herbarium specimens (Little et al., 2020)</a></p>\n<ul>\n<li>683 taxa totaling 46K images,34K for training, 2679 for validating, and 9565 for testing. Top -1 accuracy: 0.898. <ul>\n<li>Top 4 competitors utilized SeResNeXt-101, ResNet-50, SENet-154. </li></ul></li>\n</ul>\n<p><a href=\"https://www.frontiersin.org/articles/10.3389/fpls.2021.806407/full\" target=\"_blank\">Harnessing Large-Scale Herbarium Image Datasets Through Representation Learning (Walker et al., 2021)</a></p>\n<ul>\n<li>6.4K taxa totaling 424K images for training, 106K for validating, 139K for testing. <ul>\n<li>Top -1 accuracy: 0.473</li>\n<li>Compared autoencoder, CE loss, and triplet loss (triplet loss did the best)</li>\n<li>ResNet-18</li>\n<li>trained 25 epochs on a Tesla V100 GPU</li></ul></li>\n</ul>\n<p>Please feel free to extend this thread as you read other papers about plant specimen identification!</p>",
  "messages": [
    {
      "id": 1697792,
      "postDate": "2022-02-19T20:57:43.973Z",
      "content": "<p>Hi everyone,</p>\n<p>Thank you for all the valuable information about summarizing the previous threads and listing potential relevant literature to this competition.</p>\n<p>Building upon <a href=\"https://www.kaggle.com/sytuannguyen\" target=\"_blank\">@sytuannguyen</a>'s discussion thread about <a href=\"https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307810\" target=\"_blank\">some relevant literature</a>, I would like to start this discussion thread to summarize all the previous literature on plant specimen classification. I hope this could serve as a reference point for you guys when building models for our competition. </p>\n<p>Focusing on the following questions:</p>\n<ul>\n<li>What dataset do they use: How many taxa to classify? How many training images and testing images?</li>\n<li>What are the models and the settings (backbone, loss, scheduler, etc.)?</li>\n<li>What is the accuracy of their model?</li>\n</ul>\n<p>To my knowledge, state-of-the-art methods are most likely from our previous competition (Please correct me if I am wrong), e.g., <a href=\"https://www.kaggle.com/c/herbarium-2020-fgvc7\" target=\"_blank\">herbarium 2020</a> and <a href=\"https://www.kaggle.com/c/herbarium-2021-fgvc8\" target=\"_blank\">herbarium 2021</a>. Therefore, for those of you who are running out of time, please refer to the paper about <a href=\"https://www.frontiersin.org/articles/10.3389/fpls.2021.787127/full\" target=\"_blank\">herbarium 2021</a>. It contains a detailed description of models, losses, and schedulers to achieve state-of-the-art performance for herbarium specimen classification. Unfortunately, we don't have a publication for herbarium 2020, but many winning teams posted their solutions on the <a href=\"https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307624\" target=\"_blank\">previous winning solutions</a>. Many thanks to <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> for the summary.</p>\n<h3>Summary of the previous plant specimen classification results.</h3>\n<p><a href=\"https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-016-0827-5\" target=\"_blank\">Computer vision applied to herbarium specimens of German trees: testing the future utility of the millions of herbarium specimen images for automated identification (Unger et al., 2016)</a></p>\n<ul>\n<li>26 most common taxa in Germany,  totaling 260 images for training. 0.72- 0.85 Top-1 accuracy<ul>\n<li>SVM with Fourier descriptors, multiple leaf shape parameters, weighted orientation histogram.</li></ul></li>\n</ul>\n<p><a href=\"https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-017-1014-z\" target=\"_blank\">Going deeper in the automated identification of herbarium specimens (Carranza-Rojas et al., 2017)</a></p>\n<ul>\n<li><p>255 taxa totaling 11K images. Top-1 accruacy: 0.745 Top-5 accuracy: 0.872</p>\n<ul>\n<li>input resolution: 256 by 256. </li>\n<li>best model pre-trained on imageNet and Herbarium1K dataset</li>\n<li>GoogleNet model that stacks inception layers</li>\n<li>Batch Normalization</li>\n<li>PReLU</li>\n<li>simple crop and resize data augmentation, default with Caffe.</li></ul></li>\n<li><p>1204 taxa totaling 292K images for training and 51K images for testing. Top-1 accuracy: 0.796, Top-5 accuracy:0.903</p>\n<ul>\n<li>Same setting</li></ul></li>\n</ul>\n<p><a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005993\" target=\"_blank\">Automated plant species identification - Trends and future directions (Waldchen et al., 2018)</a> <br></p>\n<ul>\n<li>No results found for plant specimen identification</li>\n</ul>\n<p><a href=\"https://bsapubs.onlinelibrary.wiley.com/doi/full/10.1002/aps3.11365\" target=\"_blank\">An algorithm competition for automatic species identification from herbarium specimens (Little et al., 2020)</a></p>\n<ul>\n<li>683 taxa totaling 46K images,34K for training, 2679 for validating, and 9565 for testing. Top -1 accuracy: 0.898. <ul>\n<li>Top 4 competitors utilized SeResNeXt-101, ResNet-50, SENet-154. </li></ul></li>\n</ul>\n<p><a href=\"https://www.frontiersin.org/articles/10.3389/fpls.2021.806407/full\" target=\"_blank\">Harnessing Large-Scale Herbarium Image Datasets Through Representation Learning (Walker et al., 2021)</a></p>\n<ul>\n<li>6.4K taxa totaling 424K images for training, 106K for validating, 139K for testing. <ul>\n<li>Top -1 accuracy: 0.473</li>\n<li>Compared autoencoder, CE loss, and triplet loss (triplet loss did the best)</li>\n<li>ResNet-18</li>\n<li>trained 25 epochs on a Tesla V100 GPU</li></ul></li>\n</ul>\n<p>Please feel free to extend this thread as you read other papers about plant specimen identification!</p>",
      "rawMarkdown": "Hi everyone,\n\nThank you for all the valuable information about summarizing the previous threads and listing potential relevant literature to this competition.\n\nBuilding upon @sytuannguyen's discussion thread about [some relevant literature](https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307810), I would like to start this discussion thread to summarize all the previous literature on plant specimen classification. I hope this could serve as a reference point for you guys when building models for our competition. \n\nFocusing on the following questions:\n- What dataset do they use: How many taxa to classify? How many training images and testing images?\n- What are the models and the settings (backbone, loss, scheduler, etc.)?\n- What is the accuracy of their model?\n\nTo my knowledge, state-of-the-art methods are most likely from our previous competition (Please correct me if I am wrong), e.g., [herbarium 2020](https://www.kaggle.com/c/herbarium-2020-fgvc7) and [herbarium 2021](https://www.kaggle.com/c/herbarium-2021-fgvc8). Therefore, for those of you who are running out of time, please refer to the paper about [herbarium 2021](https://www.frontiersin.org/articles/10.3389/fpls.2021.787127/full). It contains a detailed description of models, losses, and schedulers to achieve state-of-the-art performance for herbarium specimen classification. Unfortunately, we don't have a publication for herbarium 2020, but many winning teams posted their solutions on the [previous winning solutions](https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307624). Many thanks to @init27 for the summary.\n\n\n### Summary of the previous plant specimen classification results. \n\n[Computer vision applied to herbarium specimens of German trees: testing the future utility of the millions of herbarium specimen images for automated identification (Unger et al., 2016)](https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-016-0827-5)\n- 26 most common taxa in Germany,  totaling 260 images for training. 0.72- 0.85 Top-1 accuracy\n    - SVM with Fourier descriptors, multiple leaf shape parameters, weighted orientation histogram.\n\n[Going deeper in the automated identification of herbarium specimens (Carranza-Rojas et al., 2017)](https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-017-1014-z)\n- 255 taxa totaling 11K images. Top-1 accruacy: 0.745 Top-5 accuracy: 0.872\n    - input resolution: 256 by 256. \n    - best model pre-trained on imageNet and Herbarium1K dataset\n    - GoogleNet model that stacks inception layers\n    - Batch Normalization\n    - PReLU\n    - simple crop and resize data augmentation, default with Caffe.\n    \n- 1204 taxa totaling 292K images for training and 51K images for testing. Top-1 accuracy: 0.796, Top-5 accuracy:0.903\n    - Same setting\n\n[Automated plant species identification - Trends and future directions (Waldchen et al., 2018)](https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005993) <br>\n- No results found for plant specimen identification\n\n[An algorithm competition for automatic species identification from herbarium specimens (Little et al., 2020)](https://bsapubs.onlinelibrary.wiley.com/doi/full/10.1002/aps3.11365)\n- 683 taxa totaling 46K images,34K for training, 2679 for validating, and 9565 for testing. Top -1 accuracy: 0.898. \n    - Top 4 competitors utilized SeResNeXt-101, ResNet-50, SENet-154. \n\n[Harnessing Large-Scale Herbarium Image Datasets Through Representation Learning (Walker et al., 2021)](https://www.frontiersin.org/articles/10.3389/fpls.2021.806407/full)\n- 6.4K taxa totaling 424K images for training, 106K for validating, 139K for testing. \n    - Top -1 accuracy: 0.473\n    - Compared autoencoder, CE loss, and triplet loss (triplet loss did the best)\n    - ResNet-18\n    - trained 25 epochs on a Tesla V100 GPU\n    \nPlease feel free to extend this thread as you read other papers about plant specimen identification!\n",
      "votes": 6
    },
    {
      "id": 1698003,
      "postDate": "2022-02-20T04:04:51.340Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1698003,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-20T04:04:51.340000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1697792": "Hi everyone,\n\nThank you for all the valuable information about summarizing the previous threads and listing potential relevant literature to this competition.\n\nBuilding upon @sytuannguyen's discussion thread about [some relevant literature](https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307810), I would like to start this discussion thread to summarize all the previous literature on plant specimen classification. I hope this could serve as a reference point for you guys when building models for our competition. \n\nFocusing on the following questions:\n- What dataset do they use: How many taxa to classify? How many training images and testing images?\n- What are the models and the settings (backbone, loss, scheduler, etc.)?\n- What is the accuracy of their model?\n\nTo my knowledge, state-of-the-art methods are most likely from our previous competition (Please correct me if I am wrong), e.g., [herbarium 2020](https://www.kaggle.com/c/herbarium-2020-fgvc7) and [herbarium 2021](https://www.kaggle.com/c/herbarium-2021-fgvc8). Therefore, for those of you who are running out of time, please refer to the paper about [herbarium 2021](https://www.frontiersin.org/articles/10.3389/fpls.2021.787127/full). It contains a detailed description of models, losses, and schedulers to achieve state-of-the-art performance for herbarium specimen classification. Unfortunately, we don't have a publication for herbarium 2020, but many winning teams posted their solutions on the [previous winning solutions](https://www.kaggle.com/c/herbarium-2022-fgvc9/discussion/307624). Many thanks to @init27 for the summary.\n\n\n### Summary of the previous plant specimen classification results. \n\n[Computer vision applied to herbarium specimens of German trees: testing the future utility of the millions of herbarium specimen images for automated identification (Unger et al., 2016)](https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-016-0827-5)\n- 26 most common taxa in Germany,  totaling 260 images for training. 0.72- 0.85 Top-1 accuracy\n    - SVM with Fourier descriptors, multiple leaf shape parameters, weighted orientation histogram.\n\n[Going deeper in the automated identification of herbarium specimens (Carranza-Rojas et al., 2017)](https://bmcecolevol.biomedcentral.com/articles/10.1186/s12862-017-1014-z)\n- 255 taxa totaling 11K images. Top-1 accruacy: 0.745 Top-5 accuracy: 0.872\n    - input resolution: 256 by 256. \n    - best model pre-trained on imageNet and Herbarium1K dataset\n    - GoogleNet model that stacks inception layers\n    - Batch Normalization\n    - PReLU\n    - simple crop and resize data augmentation, default with Caffe.\n    \n- 1204 taxa totaling 292K images for training and 51K images for testing. Top-1 accuracy: 0.796, Top-5 accuracy:0.903\n    - Same setting\n\n[Automated plant species identification - Trends and future directions (Waldchen et al., 2018)](https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005993) <br>\n- No results found for plant specimen identification\n\n[An algorithm competition for automatic species identification from herbarium specimens (Little et al., 2020)](https://bsapubs.onlinelibrary.wiley.com/doi/full/10.1002/aps3.11365)\n- 683 taxa totaling 46K images,34K for training, 2679 for validating, and 9565 for testing. Top -1 accuracy: 0.898. \n    - Top 4 competitors utilized SeResNeXt-101, ResNet-50, SENet-154. \n\n[Harnessing Large-Scale Herbarium Image Datasets Through Representation Learning (Walker et al., 2021)](https://www.frontiersin.org/articles/10.3389/fpls.2021.806407/full)\n- 6.4K taxa totaling 424K images for training, 106K for validating, 139K for testing. \n    - Top -1 accuracy: 0.473\n    - Compared autoencoder, CE loss, and triplet loss (triplet loss did the best)\n    - ResNet-18\n    - trained 25 epochs on a Tesla V100 GPU\n    \nPlease feel free to extend this thread as you read other papers about plant specimen identification!\n",
    "1698003": ""
  }
}