{
  "id": 308836,
  "title": "Taxonomy Prediction",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/308836",
  "author_name": "Marília Prata",
  "post_date": "2022-02-20T14:09:57.430000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>\"Deep Learning for Taxonomy Prediction\"</h1>\n<p>\"The last decade has seen great advances in Next-Generation Sequencing technologies, and, as a result, there has been a rise in the number of genomes sequenced each year. In 2017, there were as many as 10,000 new organisms sequenced and added into the RefSeq Database.\"<br>\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. In this research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"<br>\n\"Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction. In the Plinko strategy, each network takes advantage of different word usage patterns corresponding to different levels of evolutionary divergence. Plinko has the advantages of relatively low storage, GPGPU parallel training and inference, making the solution portable, and scalable with anticipated genome database growth.\"<br>\n\"To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional,alignment-free approach to taxonomy prediction.\"<br>\n<a href=\"https://vtechworks.lib.vt.edu/handle/10919/89752\" target=\"_blank\">https://vtechworks.lib.vt.edu/handle/10919/89752</a></p>\n<h1>Taxonomy relationships</h1>\n<p>\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"<br>\n\"In that research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"<br>\n\"Three major challenges in taxonomy prediction are (i) large dataset sizes (order of 109 sequences) (ii) large label spaces (order of 103 labels) and (iii) low resolution inputs (100 base pairs or less). Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction for hard to classify sequences under the three conditions stated.\"<br>\n\"Plinko has the advantage of relatively low storage footprint, making the solution portable, and scalable with anticipated genome database growth. To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional, alignment-free approach to taxonomy prediction.\"<br>\n<a href=\"https://vtechworks.lib.vt.edu/handle/10919/89752\" target=\"_blank\">https://vtechworks.lib.vt.edu/handle/10919/89752</a></p>\n<h1>Taxonomic Prediction with Tree-Structured Covariances</h1>\n<p>Authors: Matthew B. Blaschko, Wojciech Zaremba, Arthur Gretton DOI <a href=\"https://doi.org/10.1007/978-3-642-40991-2_20\" target=\"_blank\">https://doi.org/10.1007/978-3-642-40991-2_20</a><br>\n\"Taxonomies have been proposed numerous times in the literature in order to encode semantic relationships between classes. Such taxonomies have been used to improve classification results by increasing the statistical efficiency of learning, as similarities between classes can be used to increase the amount of relevant data during training.\"<br>\n\"In that paper, the authors show how data-derived taxonomies may be used in a structured prediction framework, and compare the performance of learned and semantically constructed taxonomies. Structured prediction in this case is multi-class categorization with the assumption that categories are taxonomically related.\"<br>\n\"They made three main contributions: (i) They proved the equivalence between tree-structured covariance matrices and taxonomies; (ii) They used this covariance representation to develop a highly computationally efficient optimization algorithm for structured prediction with taxonomies; (iii) They showed that the taxonomies learned from data using the Hilbert- Schmidt Independence Criterion (HSIC) often perform better than imputed semantic taxonomies.\"<br>\n\"Source code of this implementation, as well as machine readable learned taxonomies are available for download from https://github.com/blaschko/tree-structured-covariance.\"<br>\n<a href=\"https://link.springer.com/chapter/10.1007/978-3-642-40991-2_20\" target=\"_blank\">https://link.springer.com/chapter/10.1007/978-3-642-40991-2_20</a></p>\n<h1>Accuracy of Taxonomy Prediction</h1>\n<p>\"Accuracy of taxonomy prediction for 16S rRNA and fungal ITS sequences\"<br>\nAuthor: Robert C. Edgar<br>\n\"Prediction of taxonomy for marker gene sequences such as 16S ribosomal RNA (rRNA) is a fundamental task in microbiology. Most experimentally observed sequences are diverged from reference sequences of authoritatively named organisms, creating a challenge for prediction methods.\"<br>\n\"The author assessed the accuracy of several algorithms using cross-validation by identity, a new benchmark strategy which explicitly models the variation in distances between query sequences and the closest entry in a reference database. When the accuracy of genus predictions was averaged over a representative range of identities with the reference database (100%, 99%, 97%, 95% and 90%), all tested methods had ≤50% accuracy on the currently-popular V4 region of 16S rRNA.\"<br>\n\"Accuracy was found to fall rapidly with identity; for example, better methods were found to have V4 genus prediction accuracy of ∼100% at 100% identity but ∼50% at 97% identity. The relationship between identity and taxonomy was quantified as the probability that a rank is the lowest shared by a pair of sequences with a given pair-wise identity. With the V4 region, 95% identity was found to be a twilight zone where taxonomy is highly ambiguous because the probabilities that the lowest shared rank between pairs of sequences is genus, family, order or class are approximately equal.\"<br>\n<a href=\"https://peerj.com/articles/4652/\" target=\"_blank\">https://peerj.com/articles/4652/</a></p>\n<h1>Plants Conservation Relevant Predictions</h1>\n<p>\"The conservation status of most plant species is currently unknown, despite the fundamental role of plants in ecosystem health. To facilitate the costly process of conservation assessment, the authors developed a predictive protocol using a ML approach to predict conservation status of over 150,000 land plant species. Our study uses open-source geographic, environmental, and morphological trait data, making this the largest assessment of conservation risk to date and the only global assessment for plants.\"<br>\n\"Their results indicated that a large number of unassessed species are likely at risk and identify several geographic regions with the highest need of conservation efforts, many of which are not currently recognized as regions of global concern.\"<br>\n\"By providing conservation-relevant predictions at multiple spatial and taxonomic scales, predictive frameworks such as the one developed here fill a pressing need for biodiversity science.\"<br>\n<a href=\"https://www.pnas.org/content/115/51/13027\" target=\"_blank\">https://www.pnas.org/content/115/51/13027</a></p>\n<h1>Taxonomy prediction Concept</h1>\n<p>\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"<br>\n<a href=\"https://vtechworks.lib.vt.edu/handle/10919/89752\" target=\"_blank\">https://vtechworks.lib.vt.edu/handle/10919/89752</a></p>\n<h1>Plants Biodiversity</h1>\n<p>\"Biodiversity is essential for ecosystem function yet is being lost at an unprecedented rate. This threat to ecosystem function has downstream economic and cultural consequences that affect human health and well-being.\"<br>\n\"Plants are the foundation of ecosystem architecture and agriculture, and as such, changes in plant species diversity strongly influence processes such as biomass production, decomposition, and nutrient cycling. Plant diversity is therefore critical for diversity on other trophic levels.\"</p>",
  "messages": [
    {
      "id": 1698593,
      "postDate": "2022-02-20T14:09:57.430Z",
      "content": "<h1>\"Deep Learning for Taxonomy Prediction\"</h1>\n<p>\"The last decade has seen great advances in Next-Generation Sequencing technologies, and, as a result, there has been a rise in the number of genomes sequenced each year. In 2017, there were as many as 10,000 new organisms sequenced and added into the RefSeq Database.\"<br>\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. In this research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"<br>\n\"Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction. In the Plinko strategy, each network takes advantage of different word usage patterns corresponding to different levels of evolutionary divergence. Plinko has the advantages of relatively low storage, GPGPU parallel training and inference, making the solution portable, and scalable with anticipated genome database growth.\"<br>\n\"To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional,alignment-free approach to taxonomy prediction.\"<br>\n<a href=\"https://vtechworks.lib.vt.edu/handle/10919/89752\" target=\"_blank\">https://vtechworks.lib.vt.edu/handle/10919/89752</a></p>\n<h1>Taxonomy relationships</h1>\n<p>\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"<br>\n\"In that research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"<br>\n\"Three major challenges in taxonomy prediction are (i) large dataset sizes (order of 109 sequences) (ii) large label spaces (order of 103 labels) and (iii) low resolution inputs (100 base pairs or less). Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction for hard to classify sequences under the three conditions stated.\"<br>\n\"Plinko has the advantage of relatively low storage footprint, making the solution portable, and scalable with anticipated genome database growth. To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional, alignment-free approach to taxonomy prediction.\"<br>\n<a href=\"https://vtechworks.lib.vt.edu/handle/10919/89752\" target=\"_blank\">https://vtechworks.lib.vt.edu/handle/10919/89752</a></p>\n<h1>Taxonomic Prediction with Tree-Structured Covariances</h1>\n<p>Authors: Matthew B. Blaschko, Wojciech Zaremba, Arthur Gretton DOI <a href=\"https://doi.org/10.1007/978-3-642-40991-2_20\" target=\"_blank\">https://doi.org/10.1007/978-3-642-40991-2_20</a><br>\n\"Taxonomies have been proposed numerous times in the literature in order to encode semantic relationships between classes. Such taxonomies have been used to improve classification results by increasing the statistical efficiency of learning, as similarities between classes can be used to increase the amount of relevant data during training.\"<br>\n\"In that paper, the authors show how data-derived taxonomies may be used in a structured prediction framework, and compare the performance of learned and semantically constructed taxonomies. Structured prediction in this case is multi-class categorization with the assumption that categories are taxonomically related.\"<br>\n\"They made three main contributions: (i) They proved the equivalence between tree-structured covariance matrices and taxonomies; (ii) They used this covariance representation to develop a highly computationally efficient optimization algorithm for structured prediction with taxonomies; (iii) They showed that the taxonomies learned from data using the Hilbert- Schmidt Independence Criterion (HSIC) often perform better than imputed semantic taxonomies.\"<br>\n\"Source code of this implementation, as well as machine readable learned taxonomies are available for download from https://github.com/blaschko/tree-structured-covariance.\"<br>\n<a href=\"https://link.springer.com/chapter/10.1007/978-3-642-40991-2_20\" target=\"_blank\">https://link.springer.com/chapter/10.1007/978-3-642-40991-2_20</a></p>\n<h1>Accuracy of Taxonomy Prediction</h1>\n<p>\"Accuracy of taxonomy prediction for 16S rRNA and fungal ITS sequences\"<br>\nAuthor: Robert C. Edgar<br>\n\"Prediction of taxonomy for marker gene sequences such as 16S ribosomal RNA (rRNA) is a fundamental task in microbiology. Most experimentally observed sequences are diverged from reference sequences of authoritatively named organisms, creating a challenge for prediction methods.\"<br>\n\"The author assessed the accuracy of several algorithms using cross-validation by identity, a new benchmark strategy which explicitly models the variation in distances between query sequences and the closest entry in a reference database. When the accuracy of genus predictions was averaged over a representative range of identities with the reference database (100%, 99%, 97%, 95% and 90%), all tested methods had ≤50% accuracy on the currently-popular V4 region of 16S rRNA.\"<br>\n\"Accuracy was found to fall rapidly with identity; for example, better methods were found to have V4 genus prediction accuracy of ∼100% at 100% identity but ∼50% at 97% identity. The relationship between identity and taxonomy was quantified as the probability that a rank is the lowest shared by a pair of sequences with a given pair-wise identity. With the V4 region, 95% identity was found to be a twilight zone where taxonomy is highly ambiguous because the probabilities that the lowest shared rank between pairs of sequences is genus, family, order or class are approximately equal.\"<br>\n<a href=\"https://peerj.com/articles/4652/\" target=\"_blank\">https://peerj.com/articles/4652/</a></p>\n<h1>Plants Conservation Relevant Predictions</h1>\n<p>\"The conservation status of most plant species is currently unknown, despite the fundamental role of plants in ecosystem health. To facilitate the costly process of conservation assessment, the authors developed a predictive protocol using a ML approach to predict conservation status of over 150,000 land plant species. Our study uses open-source geographic, environmental, and morphological trait data, making this the largest assessment of conservation risk to date and the only global assessment for plants.\"<br>\n\"Their results indicated that a large number of unassessed species are likely at risk and identify several geographic regions with the highest need of conservation efforts, many of which are not currently recognized as regions of global concern.\"<br>\n\"By providing conservation-relevant predictions at multiple spatial and taxonomic scales, predictive frameworks such as the one developed here fill a pressing need for biodiversity science.\"<br>\n<a href=\"https://www.pnas.org/content/115/51/13027\" target=\"_blank\">https://www.pnas.org/content/115/51/13027</a></p>\n<h1>Taxonomy prediction Concept</h1>\n<p>\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"<br>\n<a href=\"https://vtechworks.lib.vt.edu/handle/10919/89752\" target=\"_blank\">https://vtechworks.lib.vt.edu/handle/10919/89752</a></p>\n<h1>Plants Biodiversity</h1>\n<p>\"Biodiversity is essential for ecosystem function yet is being lost at an unprecedented rate. This threat to ecosystem function has downstream economic and cultural consequences that affect human health and well-being.\"<br>\n\"Plants are the foundation of ecosystem architecture and agriculture, and as such, changes in plant species diversity strongly influence processes such as biomass production, decomposition, and nutrient cycling. Plant diversity is therefore critical for diversity on other trophic levels.\"</p>",
      "rawMarkdown": "\n#\"Deep Learning for Taxonomy Prediction\"\n\n\"The last decade has seen great advances in Next-Generation Sequencing technologies, and, as a result, there has been a rise in the number of genomes sequenced each year. In 2017, there were as many as 10,000 new organisms sequenced and added into the RefSeq Database.\"\n\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. In this research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"\n\n\"Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction. In the Plinko strategy, each network takes advantage of different word usage patterns corresponding to different levels of evolutionary divergence. Plinko has the advantages of relatively low storage, GPGPU parallel training and inference, making the solution portable, and scalable with anticipated genome database growth.\"\n\n\"To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional,alignment-free approach to taxonomy prediction.\"\n\nhttps://vtechworks.lib.vt.edu/handle/10919/89752\n\n\n#Taxonomy relationships\n\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"\n\n\"In that research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"\n\n\"Three major challenges in taxonomy prediction are (i) large dataset sizes (order of 109 sequences) (ii) large label spaces (order of 103 labels) and (iii) low resolution inputs (100 base pairs or less). Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction for hard to classify sequences under the three conditions stated.\"\n\n\"Plinko has the advantage of relatively low storage footprint, making the solution portable, and scalable with anticipated genome database growth. To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional, alignment-free approach to taxonomy prediction.\"\n\nhttps://vtechworks.lib.vt.edu/handle/10919/89752\n\n\n#Taxonomic Prediction with Tree-Structured Covariances\n\n\nAuthors: Matthew B. Blaschko, Wojciech Zaremba, Arthur Gretton DOI https://doi.org/10.1007/978-3-642-40991-2_20\n\n\n\"Taxonomies have been proposed numerous times in the literature in order to encode semantic relationships between classes. Such taxonomies have been used to improve classification results by increasing the statistical efficiency of learning, as similarities between classes can be used to increase the amount of relevant data during training.\"\n\n\"In that paper, the authors show how data-derived taxonomies may be used in a structured prediction framework, and compare the performance of learned and semantically constructed taxonomies. Structured prediction in this case is multi-class categorization with the assumption that categories are taxonomically related.\"\n\n\"They made three main contributions: (i) They proved the equivalence between tree-structured covariance matrices and taxonomies; (ii) They used this covariance representation to develop a highly computationally efficient optimization algorithm for structured prediction with taxonomies; (iii) They showed that the taxonomies learned from data using the Hilbert- Schmidt Independence Criterion (HSIC) often perform better than imputed semantic taxonomies.\"\n\n\"Source code of this implementation, as well as machine readable learned taxonomies are available for download from https://github.com/blaschko/tree-structured-covariance.\"\n\nhttps://link.springer.com/chapter/10.1007/978-3-642-40991-2_20\n\n#Accuracy of Taxonomy Prediction\n\n\"Accuracy of taxonomy prediction for 16S rRNA and fungal ITS sequences\"\n\n\nAuthor: Robert C. Edgar\n\n\"Prediction of taxonomy for marker gene sequences such as 16S ribosomal RNA (rRNA) is a fundamental task in microbiology. Most experimentally observed sequences are diverged from reference sequences of authoritatively named organisms, creating a challenge for prediction methods.\"\n\n\"The author assessed the accuracy of several algorithms using cross-validation by identity, a new benchmark strategy which explicitly models the variation in distances between query sequences and the closest entry in a reference database. When the accuracy of genus predictions was averaged over a representative range of identities with the reference database (100%, 99%, 97%, 95% and 90%), all tested methods had ≤50% accuracy on the currently-popular V4 region of 16S rRNA.\"\n\n\"Accuracy was found to fall rapidly with identity; for example, better methods were found to have V4 genus prediction accuracy of ∼100% at 100% identity but ∼50% at 97% identity. The relationship between identity and taxonomy was quantified as the probability that a rank is the lowest shared by a pair of sequences with a given pair-wise identity. With the V4 region, 95% identity was found to be a twilight zone where taxonomy is highly ambiguous because the probabilities that the lowest shared rank between pairs of sequences is genus, family, order or class are approximately equal.\"\n\nhttps://peerj.com/articles/4652/\n\n#Plants Conservation Relevant Predictions\n\n\"The conservation status of most plant species is currently unknown, despite the fundamental role of plants in ecosystem health. To facilitate the costly process of conservation assessment, the authors developed a predictive protocol using a ML approach to predict conservation status of over 150,000 land plant species. Our study uses open-source geographic, environmental, and morphological trait data, making this the largest assessment of conservation risk to date and the only global assessment for plants.\"\n\n\"Their results indicated that a large number of unassessed species are likely at risk and identify several geographic regions with the highest need of conservation efforts, many of which are not currently recognized as regions of global concern.\"\n\n\"By providing conservation-relevant predictions at multiple spatial and taxonomic scales, predictive frameworks such as the one developed here fill a pressing need for biodiversity science.\"\n\nhttps://www.pnas.org/content/115/51/13027\n\n#Taxonomy prediction Concept \n\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"\n\nhttps://vtechworks.lib.vt.edu/handle/10919/89752\n\n#Plants Biodiversity\n\n\"Biodiversity is essential for ecosystem function yet is being lost at an unprecedented rate. This threat to ecosystem function has downstream economic and cultural consequences that affect human health and well-being.\"\n\n\"Plants are the foundation of ecosystem architecture and agriculture, and as such, changes in plant species diversity strongly influence processes such as biomass production, decomposition, and nutrient cycling. Plant diversity is therefore critical for diversity on other trophic levels.\"",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1698593": "\n#\"Deep Learning for Taxonomy Prediction\"\n\n\"The last decade has seen great advances in Next-Generation Sequencing technologies, and, as a result, there has been a rise in the number of genomes sequenced each year. In 2017, there were as many as 10,000 new organisms sequenced and added into the RefSeq Database.\"\n\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. In this research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"\n\n\"Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction. In the Plinko strategy, each network takes advantage of different word usage patterns corresponding to different levels of evolutionary divergence. Plinko has the advantages of relatively low storage, GPGPU parallel training and inference, making the solution portable, and scalable with anticipated genome database growth.\"\n\n\"To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional,alignment-free approach to taxonomy prediction.\"\n\nhttps://vtechworks.lib.vt.edu/handle/10919/89752\n\n\n#Taxonomy relationships\n\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"\n\n\"In that research, the authors introduced Predicting Linked Organisms, Plinko, for short. Plinko is a fully-functioning, state-of-the-art predictive system that accurately captures DNA - Taxonomy relationships where other state-of-the-art algorithms falter.\"\n\n\"Three major challenges in taxonomy prediction are (i) large dataset sizes (order of 109 sequences) (ii) large label spaces (order of 103 labels) and (iii) low resolution inputs (100 base pairs or less). Plinko leverages multi-view CNNs and the pre-defined taxonomy tree structure to improve multi-level taxonomy prediction for hard to classify sequences under the three conditions stated.\"\n\n\"Plinko has the advantage of relatively low storage footprint, making the solution portable, and scalable with anticipated genome database growth. To the best of our knowledge, Plinko is the first to use multi-view CNN as the core algorithm in a compositional, alignment-free approach to taxonomy prediction.\"\n\nhttps://vtechworks.lib.vt.edu/handle/10919/89752\n\n\n#Taxonomic Prediction with Tree-Structured Covariances\n\n\nAuthors: Matthew B. Blaschko, Wojciech Zaremba, Arthur Gretton DOI https://doi.org/10.1007/978-3-642-40991-2_20\n\n\n\"Taxonomies have been proposed numerous times in the literature in order to encode semantic relationships between classes. Such taxonomies have been used to improve classification results by increasing the statistical efficiency of learning, as similarities between classes can be used to increase the amount of relevant data during training.\"\n\n\"In that paper, the authors show how data-derived taxonomies may be used in a structured prediction framework, and compare the performance of learned and semantically constructed taxonomies. Structured prediction in this case is multi-class categorization with the assumption that categories are taxonomically related.\"\n\n\"They made three main contributions: (i) They proved the equivalence between tree-structured covariance matrices and taxonomies; (ii) They used this covariance representation to develop a highly computationally efficient optimization algorithm for structured prediction with taxonomies; (iii) They showed that the taxonomies learned from data using the Hilbert- Schmidt Independence Criterion (HSIC) often perform better than imputed semantic taxonomies.\"\n\n\"Source code of this implementation, as well as machine readable learned taxonomies are available for download from https://github.com/blaschko/tree-structured-covariance.\"\n\nhttps://link.springer.com/chapter/10.1007/978-3-642-40991-2_20\n\n#Accuracy of Taxonomy Prediction\n\n\"Accuracy of taxonomy prediction for 16S rRNA and fungal ITS sequences\"\n\n\nAuthor: Robert C. Edgar\n\n\"Prediction of taxonomy for marker gene sequences such as 16S ribosomal RNA (rRNA) is a fundamental task in microbiology. Most experimentally observed sequences are diverged from reference sequences of authoritatively named organisms, creating a challenge for prediction methods.\"\n\n\"The author assessed the accuracy of several algorithms using cross-validation by identity, a new benchmark strategy which explicitly models the variation in distances between query sequences and the closest entry in a reference database. When the accuracy of genus predictions was averaged over a representative range of identities with the reference database (100%, 99%, 97%, 95% and 90%), all tested methods had ≤50% accuracy on the currently-popular V4 region of 16S rRNA.\"\n\n\"Accuracy was found to fall rapidly with identity; for example, better methods were found to have V4 genus prediction accuracy of ∼100% at 100% identity but ∼50% at 97% identity. The relationship between identity and taxonomy was quantified as the probability that a rank is the lowest shared by a pair of sequences with a given pair-wise identity. With the V4 region, 95% identity was found to be a twilight zone where taxonomy is highly ambiguous because the probabilities that the lowest shared rank between pairs of sequences is genus, family, order or class are approximately equal.\"\n\nhttps://peerj.com/articles/4652/\n\n#Plants Conservation Relevant Predictions\n\n\"The conservation status of most plant species is currently unknown, despite the fundamental role of plants in ecosystem health. To facilitate the costly process of conservation assessment, the authors developed a predictive protocol using a ML approach to predict conservation status of over 150,000 land plant species. Our study uses open-source geographic, environmental, and morphological trait data, making this the largest assessment of conservation risk to date and the only global assessment for plants.\"\n\n\"Their results indicated that a large number of unassessed species are likely at risk and identify several geographic regions with the highest need of conservation efforts, many of which are not currently recognized as regions of global concern.\"\n\n\"By providing conservation-relevant predictions at multiple spatial and taxonomic scales, predictive frameworks such as the one developed here fill a pressing need for biodiversity science.\"\n\nhttps://www.pnas.org/content/115/51/13027\n\n#Taxonomy prediction Concept \n\n\"Taxonomy prediction is a science involving the hierarchical classification of DNA fragments up to the rank species. Given species diversity on Earth, taxonomy prediction gets challenging with (i) increasing number of species (labels) to classify and (ii) decreasing input (DNA) size.\"\n\nhttps://vtechworks.lib.vt.edu/handle/10919/89752\n\n#Plants Biodiversity\n\n\"Biodiversity is essential for ecosystem function yet is being lost at an unprecedented rate. This threat to ecosystem function has downstream economic and cultural consequences that affect human health and well-being.\"\n\n\"Plants are the foundation of ecosystem architecture and agriculture, and as such, changes in plant species diversity strongly influence processes such as biomass production, decomposition, and nutrient cycling. Plant diversity is therefore critical for diversity on other trophic levels.\""
  }
}