{
  "id": 500852,
  "title": "any reason for the poor performance of HSA (ap only about 0.40)?",
  "url": "/competitions/leash-BELKA/discussion/500852",
  "author_name": "",
  "post_date": "2024-05-07T06:36:21.851207800Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>any reason for the poor performance of HSA? Its ap only about 0.40.<br>\nEven for training, the model cannot fit train data.</p>\n<p>i roughly google for binding sit of HSA. I thought they are quite distinctive. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6bcf4328added73c9ced5350cd399aaf%2FThe-crystal-structure-of-human-albumin-The-illustration-shows-the-crystal-structure-of.jpeg?generation=1715064140240607&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2798204",
      "postDate": "05/07/2024 06:36:21",
      "content": "<p>any reason for the poor performance of HSA? Its ap only about 0.40.<br>\nEven for training, the model cannot fit train data.</p>\n<p>i roughly google for binding sit of HSA. I thought they are quite distinctive. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6bcf4328added73c9ced5350cd399aaf%2FThe-crystal-structure-of-human-albumin-The-illustration-shows-the-crystal-structure-of.jpeg?generation=1715064140240607&amp;alt=media\"></p>",
      "rawMarkdown": "any reason for the poor performance of HSA? Its ap only about 0.40.\nEven for training, the model cannot fit train data.\n\ni roughly google for binding sit of HSA. I thought they are quite distinctive. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6bcf4328added73c9ced5350cd399aaf%2FThe-crystal-structure-of-human-albumin-The-illustration-shows-the-crystal-structure-of.jpeg?generation=1715064140240607&alt=media)",
      "votes": null
    },
    {
      "id": "2798410",
      "postDate": "05/07/2024 08:11:46",
      "content": "<p>I've noticed the same and thought it may be because of the length/trimming of the proteins.<br>\nHSA has the shortest (amino acid sequence) of the 3 proteins but not the most trimmed.<br>\n(I'm not sure if trim is the right word, see the data section for details.)<br>\nPerhaps a domain expert can better interpret this.</p>\n<p>BRD: cv~0.6 1362 -&gt; 585 <br>\nSEH: cv~0.9 555 -&gt; 554 <br>\nHSA: cv~0.4 609 -&gt; 477</p>",
      "rawMarkdown": "I've noticed the same and thought it may be because of the length/trimming of the proteins.\nHSA has the shortest (amino acid sequence) of the 3 proteins but not the most trimmed.\n(I'm not sure if trim is the right word, see the data section for details.)\nPerhaps a domain expert can better interpret this.\n\nBRD: cv~0.6 1362 -> 585 \nSEH: cv~0.9 555 -> 554 \nHSA: cv~0.4 609 -> 477",
      "votes": null
    },
    {
      "id": "2799686",
      "postDate": "05/07/2024 22:15:49",
      "content": "<p>Great question - I just checked myself trying to memorize targets (valid as subset of train) and it works well for the other proteins but not for HSA. Could it be that we have some label noise for this protein? Or maybe some other factors play a role here? </p>",
      "rawMarkdown": "Great question - I just checked myself trying to memorize targets (valid as subset of train) and it works well for the other proteins but not for HSA. Could it be that we have some label noise for this protein? Or maybe some other factors play a role here?",
      "votes": null
    },
    {
      "id": "2799837",
      "postDate": "05/08/2024 01:50:09",
      "content": "<p>I intend to look a bit closer into this with the data we have, but here are my thoughts on the proteins we're working with in this competition. </p>\n<p>HSA is notoriously promiscuous and it binds with many compounds. Binding tends to loosely trend with LogP (octanol/water partition of a compound, i.e. \"grease\"). From what I understand it is also very flexible so the binding pockets can be a bit dynamic. Its promiscuity is unsurprising given that it has many functions and it tends to bind to things like fatty acids.</p>\n<p>BRD4 binds to other proteins when their lysines are acetylated. This means that it's built to recognize features on macromolecular targets and as far as I'm aware, it hasn't been evolved to interact strongly with any endogenous small molecules, so the binding sites don't have an inherent set of recognition elements for a small molecule.</p>\n<p>sEH on the other hand is a bifunctional enzyme that catalyzes the opening of an epoxide or cleaves phosphate bonds. This means it has built in binding sites that are evolved to selectively bind to specific substructures and catalyze a reaction containing those substructures.</p>\n<p>My hypothesis: Because sEH is an enzyme with evolved \"active sites\" to catalyze different reactions with small molecules, the things that bind to it are likely to be the most specific, whereas the promiscuous HSA will have a noisier set of binders because it binds many things with lower selectivity than sEH. This is somewhat confirmed with a 3D PCA or UMAP plot of a random subset of compounds that bind to at least one of the proteins we're interested in. If you pull a random sample of 30000 compounds that bind to at least one of the 3 proteins of interest, calculate the Morgan fingerprints, do PCA on the fingerprints, and plot the first 3 components (only 14% of the total variance explained), it becomes clear that many of the sEH binders already separate from the pack pretty clearly.</p>",
      "rawMarkdown": "I intend to look a bit closer into this with the data we have, but here are my thoughts on the proteins we're working with in this competition. \n\nHSA is notoriously promiscuous and it binds with many compounds. Binding tends to loosely trend with LogP (octanol/water partition of a compound, i.e. \"grease\"). From what I understand it is also very flexible so the binding pockets can be a bit dynamic. Its promiscuity is unsurprising given that it has many functions and it tends to bind to things like fatty acids.\n\nBRD4 binds to other proteins when their lysines are acetylated. This means that it's built to recognize features on macromolecular targets and as far as I'm aware, it hasn't been evolved to interact strongly with any endogenous small molecules, so the binding sites don't have an inherent set of recognition elements for a small molecule.\n\nsEH on the other hand is a bifunctional enzyme that catalyzes the opening of an epoxide or cleaves phosphate bonds. This means it has built in binding sites that are evolved to selectively bind to specific substructures and catalyze a reaction containing those substructures.\n\nMy hypothesis: Because sEH is an enzyme with evolved \"active sites\" to catalyze different reactions with small molecules, the things that bind to it are likely to be the most specific, whereas the promiscuous HSA will have a noisier set of binders because it binds many things with lower selectivity than sEH. This is somewhat confirmed with a 3D PCA or UMAP plot of a random subset of compounds that bind to at least one of the proteins we're interested in. If you pull a random sample of 30000 compounds that bind to at least one of the 3 proteins of interest, calculate the Morgan fingerprints, do PCA on the fingerprints, and plot the first 3 components (only 14% of the total variance explained), it becomes clear that many of the sEH binders already separate from the pack pretty clearly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2798410,
      "author_name": "sroger",
      "author_url": "",
      "post_date": "05/07/2024 08:11:46",
      "content": "<p>I've noticed the same and thought it may be because of the length/trimming of the proteins.<br>\nHSA has the shortest (amino acid sequence) of the 3 proteins but not the most trimmed.<br>\n(I'm not sure if trim is the right word, see the data section for details.)<br>\nPerhaps a domain expert can better interpret this.</p>\n<p>BRD: cv~0.6 1362 -&gt; 585 <br>\nSEH: cv~0.9 555 -&gt; 554 <br>\nHSA: cv~0.4 609 -&gt; 477</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2799686,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "05/07/2024 22:15:49",
      "content": "<p>Great question - I just checked myself trying to memorize targets (valid as subset of train) and it works well for the other proteins but not for HSA. Could it be that we have some label noise for this protein? Or maybe some other factors play a role here? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2799837,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "05/08/2024 01:50:09",
      "content": "<p>I intend to look a bit closer into this with the data we have, but here are my thoughts on the proteins we're working with in this competition. </p>\n<p>HSA is notoriously promiscuous and it binds with many compounds. Binding tends to loosely trend with LogP (octanol/water partition of a compound, i.e. \"grease\"). From what I understand it is also very flexible so the binding pockets can be a bit dynamic. Its promiscuity is unsurprising given that it has many functions and it tends to bind to things like fatty acids.</p>\n<p>BRD4 binds to other proteins when their lysines are acetylated. This means that it's built to recognize features on macromolecular targets and as far as I'm aware, it hasn't been evolved to interact strongly with any endogenous small molecules, so the binding sites don't have an inherent set of recognition elements for a small molecule.</p>\n<p>sEH on the other hand is a bifunctional enzyme that catalyzes the opening of an epoxide or cleaves phosphate bonds. This means it has built in binding sites that are evolved to selectively bind to specific substructures and catalyze a reaction containing those substructures.</p>\n<p>My hypothesis: Because sEH is an enzyme with evolved \"active sites\" to catalyze different reactions with small molecules, the things that bind to it are likely to be the most specific, whereas the promiscuous HSA will have a noisier set of binders because it binds many things with lower selectivity than sEH. This is somewhat confirmed with a 3D PCA or UMAP plot of a random subset of compounds that bind to at least one of the proteins we're interested in. If you pull a random sample of 30000 compounds that bind to at least one of the 3 proteins of interest, calculate the Morgan fingerprints, do PCA on the fingerprints, and plot the first 3 components (only 14% of the total variance explained), it becomes clear that many of the sEH binders already separate from the pack pretty clearly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2798204": "any reason for the poor performance of HSA? Its ap only about 0.40.\nEven for training, the model cannot fit train data.\n\ni roughly google for binding sit of HSA. I thought they are quite distinctive. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F6bcf4328added73c9ced5350cd399aaf%2FThe-crystal-structure-of-human-albumin-The-illustration-shows-the-crystal-structure-of.jpeg?generation=1715064140240607&alt=media)",
    "2798410": "I've noticed the same and thought it may be because of the length/trimming of the proteins.\nHSA has the shortest (amino acid sequence) of the 3 proteins but not the most trimmed.\n(I'm not sure if trim is the right word, see the data section for details.)\nPerhaps a domain expert can better interpret this.\n\nBRD: cv~0.6 1362 -> 585 \nSEH: cv~0.9 555 -> 554 \nHSA: cv~0.4 609 -> 477",
    "2799686": "Great question - I just checked myself trying to memorize targets (valid as subset of train) and it works well for the other proteins but not for HSA. Could it be that we have some label noise for this protein? Or maybe some other factors play a role here?",
    "2799837": "I intend to look a bit closer into this with the data we have, but here are my thoughts on the proteins we're working with in this competition. \n\nHSA is notoriously promiscuous and it binds with many compounds. Binding tends to loosely trend with LogP (octanol/water partition of a compound, i.e. \"grease\"). From what I understand it is also very flexible so the binding pockets can be a bit dynamic. Its promiscuity is unsurprising given that it has many functions and it tends to bind to things like fatty acids.\n\nBRD4 binds to other proteins when their lysines are acetylated. This means that it's built to recognize features on macromolecular targets and as far as I'm aware, it hasn't been evolved to interact strongly with any endogenous small molecules, so the binding sites don't have an inherent set of recognition elements for a small molecule.\n\nsEH on the other hand is a bifunctional enzyme that catalyzes the opening of an epoxide or cleaves phosphate bonds. This means it has built in binding sites that are evolved to selectively bind to specific substructures and catalyze a reaction containing those substructures.\n\nMy hypothesis: Because sEH is an enzyme with evolved \"active sites\" to catalyze different reactions with small molecules, the things that bind to it are likely to be the most specific, whereas the promiscuous HSA will have a noisier set of binders because it binds many things with lower selectivity than sEH. This is somewhat confirmed with a 3D PCA or UMAP plot of a random subset of compounds that bind to at least one of the proteins we're interested in. If you pull a random sample of 30000 compounds that bind to at least one of the 3 proteins of interest, calculate the Morgan fingerprints, do PCA on the fingerprints, and plot the first 3 components (only 14% of the total variance explained), it becomes clear that many of the sEH binders already separate from the pack pretty clearly."
  },
  "source": "meta"
}