{
  "id": 214891,
  "title": "Papers on multi-label classification protein and predicting protein patterns",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/214891",
  "author_name": "Charlie Craine",
  "post_date": "2021-01-28T02:47:38.433000",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey everyone!</p>\n<p>I wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!</p>\n<p>Research Papers:</p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1703.10663\" target=\"_blank\">Near Perfect Protein Multi-Label Classification with Deep Neural Networks</a> - Artificial neural networks (ANNs) have gained a well-deserved popularity among machine learning tools upon their recent successful applications in image- and sound processing and classification problems. ANNs have also been applied for predicting the family or function of a protein, knowing its residue sequence. Here we present two new ANNs with multi-label classification ability, showing impressive accuracy when classifying protein sequences into 698 UniProt families (AUC=99.99%) and 983 Gene Ontology classes (AUC=99.45%).</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1704.05204\" target=\"_blank\">HPSLPred: An Ensemble Multi-label Classifier for Human Protein Subcellular Location Prediction with Imbalanced Source</a> - Predicting the subcellular localization of proteins is an important and challenging problem. Traditional experimental approaches are often expensive and time-consuming. Consequently, a growing number of research efforts employ a series of machine learning approaches to predict the subcellular location of proteins. There are two main challenges among the state-of-the-art prediction methods. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/2011.11259\" target=\"_blank\">Sparse generative modeling of protein-sequence families</a> -  In this work, we introduce a parameter-reduction procedure via iterative decimation of the less statistically significant couplings. We propose an information-based criterion that identifies couplings that are either weak, or statistically unsupported. We show that our procedure allows one to remove more than 90% of the PM couplings, while preserving the predictive and generative properties of the original dense PM. The resulting model is far away from criticality, meaning that it is more robust to noise, and its couplings are more easily interpretable.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2012.03084\" target=\"_blank\">Pre-training Protein Language Models with Label-Agnostic Binding Pairs Enhances Performance in Downstream Tasks</a> - Less than 1% of protein sequences are structurally and functionally annotated. Natural Language Processing (NLP) community has recently embraced self-supervised learning as a powerful approach to learn representations from unlabeled text, in large part due to the attention-based context-aware Transformer models. In this work we present a modification to the RoBERTa model by inputting during pre-training a mixture of binding and non-binding protein sequences (from STRING database).</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2008.12473\" target=\"_blank\">Pre-training of Graph Neural Network for Modeling Effects of Mutations on Protein-Protein Binding Affinity</a> - Modeling the effects of mutations on the binding affinity plays a crucial role in protein engineering and drug design. In this study, we develop a novel deep learning based framework, named GraphPPI, to predict the binding affinity changes upon mutations based on the features provided by a graph neural network (GNN). </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1905.12126\" target=\"_blank\">Using Ontologies To Improve Performance In Massively Multi-label Prediction Models</a> - We apply this method to the two massively multi-label tasks of disease prediction (ICD-9 codes) and protein function prediction (Gene Ontology terms) and obtain significant improvements in per-label AUROC and average precision for less common labels.</p></li>\n</ul>",
  "messages": [
    {
      "id": 1173627,
      "postDate": "2021-01-28T02:47:38.433Z",
      "content": "<p>Hey everyone!</p>\n<p>I wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!</p>\n<p>Research Papers:</p>\n<ul>\n<li><p><a href=\"https://arxiv.org/abs/1703.10663\" target=\"_blank\">Near Perfect Protein Multi-Label Classification with Deep Neural Networks</a> - Artificial neural networks (ANNs) have gained a well-deserved popularity among machine learning tools upon their recent successful applications in image- and sound processing and classification problems. ANNs have also been applied for predicting the family or function of a protein, knowing its residue sequence. Here we present two new ANNs with multi-label classification ability, showing impressive accuracy when classifying protein sequences into 698 UniProt families (AUC=99.99%) and 983 Gene Ontology classes (AUC=99.45%).</p></li>\n<li><p><a href=\"https://arxiv.org/abs/1704.05204\" target=\"_blank\">HPSLPred: An Ensemble Multi-label Classifier for Human Protein Subcellular Location Prediction with Imbalanced Source</a> - Predicting the subcellular localization of proteins is an important and challenging problem. Traditional experimental approaches are often expensive and time-consuming. Consequently, a growing number of research efforts employ a series of machine learning approaches to predict the subcellular location of proteins. There are two main challenges among the state-of-the-art prediction methods. </p></li>\n<li><p><a href=\"https://arxiv.org/abs/2011.11259\" target=\"_blank\">Sparse generative modeling of protein-sequence families</a> -  In this work, we introduce a parameter-reduction procedure via iterative decimation of the less statistically significant couplings. We propose an information-based criterion that identifies couplings that are either weak, or statistically unsupported. We show that our procedure allows one to remove more than 90% of the PM couplings, while preserving the predictive and generative properties of the original dense PM. The resulting model is far away from criticality, meaning that it is more robust to noise, and its couplings are more easily interpretable.</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2012.03084\" target=\"_blank\">Pre-training Protein Language Models with Label-Agnostic Binding Pairs Enhances Performance in Downstream Tasks</a> - Less than 1% of protein sequences are structurally and functionally annotated. Natural Language Processing (NLP) community has recently embraced self-supervised learning as a powerful approach to learn representations from unlabeled text, in large part due to the attention-based context-aware Transformer models. In this work we present a modification to the RoBERTa model by inputting during pre-training a mixture of binding and non-binding protein sequences (from STRING database).</p></li>\n<li><p><a href=\"https://arxiv.org/abs/2008.12473\" target=\"_blank\">Pre-training of Graph Neural Network for Modeling Effects of Mutations on Protein-Protein Binding Affinity</a> - Modeling the effects of mutations on the binding affinity plays a crucial role in protein engineering and drug design. In this study, we develop a novel deep learning based framework, named GraphPPI, to predict the binding affinity changes upon mutations based on the features provided by a graph neural network (GNN). </p></li>\n<li><p><a href=\"https://arxiv.org/abs/1905.12126\" target=\"_blank\">Using Ontologies To Improve Performance In Massively Multi-label Prediction Models</a> - We apply this method to the two massively multi-label tasks of disease prediction (ICD-9 codes) and protein function prediction (Gene Ontology terms) and obtain significant improvements in per-label AUROC and average precision for less common labels.</p></li>\n</ul>",
      "rawMarkdown": "Hey everyone!\n\nI wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!\n\nResearch Papers:\n* [Near Perfect Protein Multi-Label Classification with Deep Neural Networks](https://arxiv.org/abs/1703.10663) - Artificial neural networks (ANNs) have gained a well-deserved popularity among machine learning tools upon their recent successful applications in image- and sound processing and classification problems. ANNs have also been applied for predicting the family or function of a protein, knowing its residue sequence. Here we present two new ANNs with multi-label classification ability, showing impressive accuracy when classifying protein sequences into 698 UniProt families (AUC=99.99%) and 983 Gene Ontology classes (AUC=99.45%).\n\n* [HPSLPred: An Ensemble Multi-label Classifier for Human Protein Subcellular Location Prediction with Imbalanced Source](https://arxiv.org/abs/1704.05204) - Predicting the subcellular localization of proteins is an important and challenging problem. Traditional experimental approaches are often expensive and time-consuming. Consequently, a growing number of research efforts employ a series of machine learning approaches to predict the subcellular location of proteins. There are two main challenges among the state-of-the-art prediction methods. \n\n* [Sparse generative modeling of protein-sequence families](https://arxiv.org/abs/2011.11259) -  In this work, we introduce a parameter-reduction procedure via iterative decimation of the less statistically significant couplings. We propose an information-based criterion that identifies couplings that are either weak, or statistically unsupported. We show that our procedure allows one to remove more than 90% of the PM couplings, while preserving the predictive and generative properties of the original dense PM. The resulting model is far away from criticality, meaning that it is more robust to noise, and its couplings are more easily interpretable.\n\n* [Pre-training Protein Language Models with Label-Agnostic Binding Pairs Enhances Performance in Downstream Tasks](https://arxiv.org/abs/2012.03084) - Less than 1% of protein sequences are structurally and functionally annotated. Natural Language Processing (NLP) community has recently embraced self-supervised learning as a powerful approach to learn representations from unlabeled text, in large part due to the attention-based context-aware Transformer models. In this work we present a modification to the RoBERTa model by inputting during pre-training a mixture of binding and non-binding protein sequences (from STRING database).\n\n* [Pre-training of Graph Neural Network for Modeling Effects of Mutations on Protein-Protein Binding Affinity](https://arxiv.org/abs/2008.12473) - Modeling the effects of mutations on the binding affinity plays a crucial role in protein engineering and drug design. In this study, we develop a novel deep learning based framework, named GraphPPI, to predict the binding affinity changes upon mutations based on the features provided by a graph neural network (GNN). \n\n* [Using Ontologies To Improve Performance In Massively Multi-label Prediction Models](https://arxiv.org/abs/1905.12126) - We apply this method to the two massively multi-label tasks of disease prediction (ICD-9 codes) and protein function prediction (Gene Ontology terms) and obtain significant improvements in per-label AUROC and average precision for less common labels.\n",
      "votes": 4
    },
    {
      "id": 1174263,
      "postDate": "2021-01-28T11:17:32.313Z",
      "content": "<p>Segmentation Loss Odyssey, Jun Ma: <a href=\"https://arxiv.org/pdf/2005.13449v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2005.13449v1.pdf</a><br>\nCode: <a href=\"https://github.com/JunMa11/SegLoss\" target=\"_blank\">https://github.com/JunMa11/SegLoss</a><br>\nI found this helpful as a starting point for various options in penalizing segmentation output</p>",
      "rawMarkdown": "Segmentation Loss Odyssey, Jun Ma: https://arxiv.org/pdf/2005.13449v1.pdf\nCode: https://github.com/JunMa11/SegLoss\nI found this helpful as a starting point for various options in penalizing segmentation output"
    }
  ],
  "comments": [
    {
      "id": 1174263,
      "author_name": "Geet Sharma",
      "author_url": "",
      "post_date": "2021-01-28T11:17:32.313000",
      "content": "<p>Segmentation Loss Odyssey, Jun Ma: <a href=\"https://arxiv.org/pdf/2005.13449v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2005.13449v1.pdf</a><br>\nCode: <a href=\"https://github.com/JunMa11/SegLoss\" target=\"_blank\">https://github.com/JunMa11/SegLoss</a><br>\nI found this helpful as a starting point for various options in penalizing segmentation output</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1173627": "Hey everyone!\n\nI wanted to start a papers thread and build on it, and hope others share as well, papers to gain domain knowledge. I have no domain knowledge in this area so I have downloaded papers that appeared to be relevant after reading their abstracts. I'll be reading these over the coming days and commenting more as I go through all of them. Hope this helps!\n\nResearch Papers:\n* [Near Perfect Protein Multi-Label Classification with Deep Neural Networks](https://arxiv.org/abs/1703.10663) - Artificial neural networks (ANNs) have gained a well-deserved popularity among machine learning tools upon their recent successful applications in image- and sound processing and classification problems. ANNs have also been applied for predicting the family or function of a protein, knowing its residue sequence. Here we present two new ANNs with multi-label classification ability, showing impressive accuracy when classifying protein sequences into 698 UniProt families (AUC=99.99%) and 983 Gene Ontology classes (AUC=99.45%).\n\n* [HPSLPred: An Ensemble Multi-label Classifier for Human Protein Subcellular Location Prediction with Imbalanced Source](https://arxiv.org/abs/1704.05204) - Predicting the subcellular localization of proteins is an important and challenging problem. Traditional experimental approaches are often expensive and time-consuming. Consequently, a growing number of research efforts employ a series of machine learning approaches to predict the subcellular location of proteins. There are two main challenges among the state-of-the-art prediction methods. \n\n* [Sparse generative modeling of protein-sequence families](https://arxiv.org/abs/2011.11259) -  In this work, we introduce a parameter-reduction procedure via iterative decimation of the less statistically significant couplings. We propose an information-based criterion that identifies couplings that are either weak, or statistically unsupported. We show that our procedure allows one to remove more than 90% of the PM couplings, while preserving the predictive and generative properties of the original dense PM. The resulting model is far away from criticality, meaning that it is more robust to noise, and its couplings are more easily interpretable.\n\n* [Pre-training Protein Language Models with Label-Agnostic Binding Pairs Enhances Performance in Downstream Tasks](https://arxiv.org/abs/2012.03084) - Less than 1% of protein sequences are structurally and functionally annotated. Natural Language Processing (NLP) community has recently embraced self-supervised learning as a powerful approach to learn representations from unlabeled text, in large part due to the attention-based context-aware Transformer models. In this work we present a modification to the RoBERTa model by inputting during pre-training a mixture of binding and non-binding protein sequences (from STRING database).\n\n* [Pre-training of Graph Neural Network for Modeling Effects of Mutations on Protein-Protein Binding Affinity](https://arxiv.org/abs/2008.12473) - Modeling the effects of mutations on the binding affinity plays a crucial role in protein engineering and drug design. In this study, we develop a novel deep learning based framework, named GraphPPI, to predict the binding affinity changes upon mutations based on the features provided by a graph neural network (GNN). \n\n* [Using Ontologies To Improve Performance In Massively Multi-label Prediction Models](https://arxiv.org/abs/1905.12126) - We apply this method to the two massively multi-label tasks of disease prediction (ICD-9 codes) and protein function prediction (Gene Ontology terms) and obtain significant improvements in per-label AUROC and average precision for less common labels.\n",
    "1174263": "Segmentation Loss Odyssey, Jun Ma: https://arxiv.org/pdf/2005.13449v1.pdf\nCode: https://github.com/JunMa11/SegLoss\nI found this helpful as a starting point for various options in penalizing segmentation output"
  }
}