{
  "id": 491401,
  "title": "🧬 Papers: DNA-encoded chemical library (DEL) 🧬",
  "url": "/competitions/leash-BELKA/discussion/491401",
  "author_name": "",
  "post_date": "2024-04-05T18:25:32.346887400Z",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F00ebc78a6f43f18e99c9cbac102c0e5c%2Fdel-en-02.jpg?generation=1712602026140587&amp;alt=media\"></p>\n<p><strong><a href=\"https://arxiv.org/pdf/2205.08020.pdf\" target=\"_blank\">Paper: PARTIAL PRODUCT AWARE MACHINE LEARNING ON DNA-ENCODED LIBRARIES</a></strong> - is this paper more relevant to this competition?</p>\n<blockquote>\n  <p>DNA encoded libraries (DELs) are used for rapid large-scale screening of small molecules against a protein target. These combinatorial libraries are built through several cycles of chemistry and DNA ligation, producing large sets of DNA-tagged molecules. Training machine learning models on DEL data has been shown to be effective at predicting molecules of interest dissimilar from those in the original DEL. Machine learning chemical property prediction approaches rely on the assumption that the property of interest is linked to a single chemical structure. In the context of DNA-encoded libraries, this is equivalent to assuming that every chemical reaction fully yields the desired product. However, in practice, multistep chemical synthesis sometimes generates partial molecules. Each unique DNA tag in a DEL therefore corresponds to a set of possible molecules. Here, we leverage reaction yield data to enumerate the set of possible molecules corresponding to a given DNA tag. This paper demonstrates that training a custom GNN on this richer dataset improves accuracy and generalization performance.<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6606e59f14c456b6c6c75f2b9d0b8867%2FScreenshot%202024-04-05%20at%2011.50.50PM.png?generation=1712341274372067&amp;alt=media\"></p>\n</blockquote>\n<hr>\n<h1>Architecture</h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F783eb8657361087fc67aef878d3f248f%2FScreenshot%202024-04-05%20at%2011.52.41PM.png?generation=1712341383691106&amp;alt=media\"></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>DNA encoded libraries are powerful tools for screening large small molecule libraries against a target protein. In this work, we present a modeling approach that leverages information about full and partial products in DELs and their proportions. We demonstrate that our approach that models full and partial products as well as their proportions outperforms models with only some of the products or a model that does not include proportions of the products. Using only full product (trisynthon) data can produce datasets that are too noisy to support effective training, while also inadequately describing the underlying data. Using only partial (disynthon) data is an effective approach to aggregating and de-noising DEL data, but altogether disregards potentially useful data. Moreover, our approach can be used to identify molecules on an external data set that are strong binders to a target protein of interest.</p>\n</blockquote>\n<hr>\n<p><strong><a href=\"https://pubs.acs.org/doi/epdf/10.1021/acsomega.3c02152\" target=\"_blank\">Paper: Machine-Learning-Based Data Analysis Method for Cell-Based Selection of DNA-Encoded Libraries</a></strong></p>\n<blockquote>\n  <p>DNA-encoded library (DEL) is a powerful ligand discovery technology that has been widely adopted in the pharmaceutical industry. DEL selections are typically performed with a purified protein target immobilized on a matrix or in solution phase. Recently, DELs have also been used to interrogate the targets in the complex biological environment, such as membrane proteins on live cells. However, due to the complex landscape of the cell surface, the selection inevitably involves significant nonspecific interactions, and the selection data are much noisier than the ones with purified proteins, making reliable hit identification highly challenging. Researchers have developed several approaches to denoise DEL datasets, but it remains unclear whether they are suitable for cell-based DEL selections. Here, we report the proof-of-principle of a new machine-learning (ML)-based approach to process cell-based DEL selection datasets by using a Maximum A Posteriori (MAP) estimation loss function, a probabilistic framework that can account for and quantify uncertainties of noisy data. We applied the approach to a DEL selection dataset, where a library of 7,721,415 compounds was selected against a purified carbonic anhydrase 2 (CA-2) and a cell line expressing the membrane protein carbonic anhydrase 12 (CA-12). The extended-connectivity fingerprint (ECFP)-based regression model using the MAP loss function was able to identify true binders and also reliable structure–activity relationship (SAR) from the noisy cell-based selection datasets. In addition, the regularized enrichment metric (known as MAP enrichment) could also be calculated directly without involving the specific machine-learning model, effectively suppressing low-confidence outliers and enhancing the signal-to-noise ratio. Future applications of this method will focus on de novo ligand discovery from cell-based DEL selections.<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F990d42fcb5af8be0353e9cbce4e2072c%2Fimages_medium_ao3c02152_0006.gif?generation=1712342150609848&amp;alt=media\"></p>\n</blockquote>",
  "messages": [
    {
      "id": "2737368",
      "postDate": "04/05/2024 18:25:32",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F00ebc78a6f43f18e99c9cbac102c0e5c%2Fdel-en-02.jpg?generation=1712602026140587&amp;alt=media\"></p>\n<p><strong><a href=\"https://arxiv.org/pdf/2205.08020.pdf\" target=\"_blank\">Paper: PARTIAL PRODUCT AWARE MACHINE LEARNING ON DNA-ENCODED LIBRARIES</a></strong> - is this paper more relevant to this competition?</p>\n<blockquote>\n  <p>DNA encoded libraries (DELs) are used for rapid large-scale screening of small molecules against a protein target. These combinatorial libraries are built through several cycles of chemistry and DNA ligation, producing large sets of DNA-tagged molecules. Training machine learning models on DEL data has been shown to be effective at predicting molecules of interest dissimilar from those in the original DEL. Machine learning chemical property prediction approaches rely on the assumption that the property of interest is linked to a single chemical structure. In the context of DNA-encoded libraries, this is equivalent to assuming that every chemical reaction fully yields the desired product. However, in practice, multistep chemical synthesis sometimes generates partial molecules. Each unique DNA tag in a DEL therefore corresponds to a set of possible molecules. Here, we leverage reaction yield data to enumerate the set of possible molecules corresponding to a given DNA tag. This paper demonstrates that training a custom GNN on this richer dataset improves accuracy and generalization performance.<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6606e59f14c456b6c6c75f2b9d0b8867%2FScreenshot%202024-04-05%20at%2011.50.50PM.png?generation=1712341274372067&amp;alt=media\"></p>\n</blockquote>\n<hr>\n<h1>Architecture</h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F783eb8657361087fc67aef878d3f248f%2FScreenshot%202024-04-05%20at%2011.52.41PM.png?generation=1712341383691106&amp;alt=media\"></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>DNA encoded libraries are powerful tools for screening large small molecule libraries against a target protein. In this work, we present a modeling approach that leverages information about full and partial products in DELs and their proportions. We demonstrate that our approach that models full and partial products as well as their proportions outperforms models with only some of the products or a model that does not include proportions of the products. Using only full product (trisynthon) data can produce datasets that are too noisy to support effective training, while also inadequately describing the underlying data. Using only partial (disynthon) data is an effective approach to aggregating and de-noising DEL data, but altogether disregards potentially useful data. Moreover, our approach can be used to identify molecules on an external data set that are strong binders to a target protein of interest.</p>\n</blockquote>\n<hr>\n<p><strong><a href=\"https://pubs.acs.org/doi/epdf/10.1021/acsomega.3c02152\" target=\"_blank\">Paper: Machine-Learning-Based Data Analysis Method for Cell-Based Selection of DNA-Encoded Libraries</a></strong></p>\n<blockquote>\n  <p>DNA-encoded library (DEL) is a powerful ligand discovery technology that has been widely adopted in the pharmaceutical industry. DEL selections are typically performed with a purified protein target immobilized on a matrix or in solution phase. Recently, DELs have also been used to interrogate the targets in the complex biological environment, such as membrane proteins on live cells. However, due to the complex landscape of the cell surface, the selection inevitably involves significant nonspecific interactions, and the selection data are much noisier than the ones with purified proteins, making reliable hit identification highly challenging. Researchers have developed several approaches to denoise DEL datasets, but it remains unclear whether they are suitable for cell-based DEL selections. Here, we report the proof-of-principle of a new machine-learning (ML)-based approach to process cell-based DEL selection datasets by using a Maximum A Posteriori (MAP) estimation loss function, a probabilistic framework that can account for and quantify uncertainties of noisy data. We applied the approach to a DEL selection dataset, where a library of 7,721,415 compounds was selected against a purified carbonic anhydrase 2 (CA-2) and a cell line expressing the membrane protein carbonic anhydrase 12 (CA-12). The extended-connectivity fingerprint (ECFP)-based regression model using the MAP loss function was able to identify true binders and also reliable structure–activity relationship (SAR) from the noisy cell-based selection datasets. In addition, the regularized enrichment metric (known as MAP enrichment) could also be calculated directly without involving the specific machine-learning model, effectively suppressing low-confidence outliers and enhancing the signal-to-noise ratio. Future applications of this method will focus on de novo ligand discovery from cell-based DEL selections.<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F990d42fcb5af8be0353e9cbce4e2072c%2Fimages_medium_ao3c02152_0006.gif?generation=1712342150609848&amp;alt=media\"></p>\n</blockquote>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F00ebc78a6f43f18e99c9cbac102c0e5c%2Fdel-en-02.jpg?generation=1712602026140587&alt=media)\n\n**[Paper: PARTIAL PRODUCT AWARE MACHINE LEARNING ON DNA-ENCODED LIBRARIES](https://arxiv.org/pdf/2205.08020.pdf)** - is this paper more relevant to this competition?\n\n> DNA encoded libraries (DELs) are used for rapid large-scale screening of small molecules against a protein target. These combinatorial libraries are built through several cycles of chemistry and DNA ligation, producing large sets of DNA-tagged molecules. Training machine learning models on DEL data has been shown to be effective at predicting molecules of interest dissimilar from those in the original DEL. Machine learning chemical property prediction approaches rely on the assumption that the property of interest is linked to a single chemical structure. In the context of DNA-encoded libraries, this is equivalent to assuming that every chemical reaction fully yields the desired product. However, in practice, multistep chemical synthesis sometimes generates partial molecules. Each unique DNA tag in a DEL therefore corresponds to a set of possible molecules. Here, we leverage reaction yield data to enumerate the set of possible molecules corresponding to a given DNA tag. This paper demonstrates that training a custom GNN on this richer dataset improves accuracy and generalization performance.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6606e59f14c456b6c6c75f2b9d0b8867%2FScreenshot%202024-04-05%20at%2011.50.50PM.png?generation=1712341274372067&alt=media)\n\n---\n\n# Architecture \n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F783eb8657361087fc67aef878d3f248f%2FScreenshot%202024-04-05%20at%2011.52.41PM.png?generation=1712341383691106&alt=media)\n\n---\n\n> DNA encoded libraries are powerful tools for screening large small molecule libraries against a target protein. In this work, we present a modeling approach that leverages information about full and partial products in DELs and their proportions. We demonstrate that our approach that models full and partial products as well as their proportions outperforms models with only some of the products or a model that does not include proportions of the products. Using only full product (trisynthon) data can produce datasets that are too noisy to support effective training, while also inadequately describing the underlying data. Using only partial (disynthon) data is an effective approach to aggregating and de-noising DEL data, but altogether disregards potentially useful data. Moreover, our approach can be used to identify molecules on an external data set that are strong binders to a target protein of interest.\n\n---\n**[Paper: Machine-Learning-Based Data Analysis Method for Cell-Based Selection of DNA-Encoded Libraries](https://pubs.acs.org/doi/epdf/10.1021/acsomega.3c02152)**\n\n> DNA-encoded library (DEL) is a powerful ligand discovery technology that has been widely adopted in the pharmaceutical industry. DEL selections are typically performed with a purified protein target immobilized on a matrix or in solution phase. Recently, DELs have also been used to interrogate the targets in the complex biological environment, such as membrane proteins on live cells. However, due to the complex landscape of the cell surface, the selection inevitably involves significant nonspecific interactions, and the selection data are much noisier than the ones with purified proteins, making reliable hit identification highly challenging. Researchers have developed several approaches to denoise DEL datasets, but it remains unclear whether they are suitable for cell-based DEL selections. Here, we report the proof-of-principle of a new machine-learning (ML)-based approach to process cell-based DEL selection datasets by using a Maximum A Posteriori (MAP) estimation loss function, a probabilistic framework that can account for and quantify uncertainties of noisy data. We applied the approach to a DEL selection dataset, where a library of 7,721,415 compounds was selected against a purified carbonic anhydrase 2 (CA-2) and a cell line expressing the membrane protein carbonic anhydrase 12 (CA-12). The extended-connectivity fingerprint (ECFP)-based regression model using the MAP loss function was able to identify true binders and also reliable structure–activity relationship (SAR) from the noisy cell-based selection datasets. In addition, the regularized enrichment metric (known as MAP enrichment) could also be calculated directly without involving the specific machine-learning model, effectively suppressing low-confidence outliers and enhancing the signal-to-noise ratio. Future applications of this method will focus on de novo ligand discovery from cell-based DEL selections.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F990d42fcb5af8be0353e9cbce4e2072c%2Fimages_medium_ao3c02152_0006.gif?generation=1712342150609848&alt=media)",
      "votes": null
    },
    {
      "id": "2786037",
      "postDate": "05/01/2024 05:59:53",
      "content": "<p>Thank you so much for sharing. I am a cs student trying to become a ds in biochem. Thanks a lot.</p>",
      "rawMarkdown": "Thank you so much for sharing. I am a cs student trying to become a ds in biochem. Thanks a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2786037,
      "author_name": "corneliaz",
      "author_url": "",
      "post_date": "05/01/2024 05:59:53",
      "content": "<p>Thank you so much for sharing. I am a cs student trying to become a ds in biochem. Thanks a lot.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2737368": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F00ebc78a6f43f18e99c9cbac102c0e5c%2Fdel-en-02.jpg?generation=1712602026140587&alt=media)\n\n**[Paper: PARTIAL PRODUCT AWARE MACHINE LEARNING ON DNA-ENCODED LIBRARIES](https://arxiv.org/pdf/2205.08020.pdf)** - is this paper more relevant to this competition?\n\n> DNA encoded libraries (DELs) are used for rapid large-scale screening of small molecules against a protein target. These combinatorial libraries are built through several cycles of chemistry and DNA ligation, producing large sets of DNA-tagged molecules. Training machine learning models on DEL data has been shown to be effective at predicting molecules of interest dissimilar from those in the original DEL. Machine learning chemical property prediction approaches rely on the assumption that the property of interest is linked to a single chemical structure. In the context of DNA-encoded libraries, this is equivalent to assuming that every chemical reaction fully yields the desired product. However, in practice, multistep chemical synthesis sometimes generates partial molecules. Each unique DNA tag in a DEL therefore corresponds to a set of possible molecules. Here, we leverage reaction yield data to enumerate the set of possible molecules corresponding to a given DNA tag. This paper demonstrates that training a custom GNN on this richer dataset improves accuracy and generalization performance.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F6606e59f14c456b6c6c75f2b9d0b8867%2FScreenshot%202024-04-05%20at%2011.50.50PM.png?generation=1712341274372067&alt=media)\n\n---\n\n# Architecture \n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F783eb8657361087fc67aef878d3f248f%2FScreenshot%202024-04-05%20at%2011.52.41PM.png?generation=1712341383691106&alt=media)\n\n---\n\n> DNA encoded libraries are powerful tools for screening large small molecule libraries against a target protein. In this work, we present a modeling approach that leverages information about full and partial products in DELs and their proportions. We demonstrate that our approach that models full and partial products as well as their proportions outperforms models with only some of the products or a model that does not include proportions of the products. Using only full product (trisynthon) data can produce datasets that are too noisy to support effective training, while also inadequately describing the underlying data. Using only partial (disynthon) data is an effective approach to aggregating and de-noising DEL data, but altogether disregards potentially useful data. Moreover, our approach can be used to identify molecules on an external data set that are strong binders to a target protein of interest.\n\n---\n**[Paper: Machine-Learning-Based Data Analysis Method for Cell-Based Selection of DNA-Encoded Libraries](https://pubs.acs.org/doi/epdf/10.1021/acsomega.3c02152)**\n\n> DNA-encoded library (DEL) is a powerful ligand discovery technology that has been widely adopted in the pharmaceutical industry. DEL selections are typically performed with a purified protein target immobilized on a matrix or in solution phase. Recently, DELs have also been used to interrogate the targets in the complex biological environment, such as membrane proteins on live cells. However, due to the complex landscape of the cell surface, the selection inevitably involves significant nonspecific interactions, and the selection data are much noisier than the ones with purified proteins, making reliable hit identification highly challenging. Researchers have developed several approaches to denoise DEL datasets, but it remains unclear whether they are suitable for cell-based DEL selections. Here, we report the proof-of-principle of a new machine-learning (ML)-based approach to process cell-based DEL selection datasets by using a Maximum A Posteriori (MAP) estimation loss function, a probabilistic framework that can account for and quantify uncertainties of noisy data. We applied the approach to a DEL selection dataset, where a library of 7,721,415 compounds was selected against a purified carbonic anhydrase 2 (CA-2) and a cell line expressing the membrane protein carbonic anhydrase 12 (CA-12). The extended-connectivity fingerprint (ECFP)-based regression model using the MAP loss function was able to identify true binders and also reliable structure–activity relationship (SAR) from the noisy cell-based selection datasets. In addition, the regularized enrichment metric (known as MAP enrichment) could also be calculated directly without involving the specific machine-learning model, effectively suppressing low-confidence outliers and enhancing the signal-to-noise ratio. Future applications of this method will focus on de novo ligand discovery from cell-based DEL selections.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F990d42fcb5af8be0353e9cbce4e2072c%2Fimages_medium_ao3c02152_0006.gif?generation=1712342150609848&alt=media)",
    "2786037": "Thank you so much for sharing. I am a cs student trying to become a ds in biochem. Thanks a lot."
  },
  "source": "meta"
}