{
  "id": 508211,
  "title": "Can we Usefully Add High-throughput Docking to the Portfolio of Methods?",
  "url": "/competitions/leash-BELKA/discussion/508211",
  "author_name": "",
  "post_date": "2024-05-28T16:31:15.711112400Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In this competition, we've seen a variety of methods used, such as <a href=\"https://www.kaggle.com/code/horikitasaku/0-589-tokenize-in-terms-of-chemical-principles\" target=\"_blank\">tokenizing SMILES</a>, <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook\" target=\"_blank\">Convolutional Neural Networks</a> and more traditional feature-based <a href=\"https://www.kaggle.com/code/yujansaya/chemensemble-molecular-binding-with-ensemble\" target=\"_blank\">cheminformatics</a>. None of these use 3D protein-ligand structure information, which suggests that <a href=\"https://www.kaggle.com/code/hideakiogasawara/docking-simulation-using-dockstring\" target=\"_blank\">docking</a> ought to play a complementary role to the other techniques, especially with the likely poor performance of cheminformatics on the non-shared building blocks.</p>\n<p>Nonetheless, <a href=\"https://www.kaggle.com/hideakiogasawara\" target=\"_blank\">@hideakiogasawara</a>'s prelimanary results suggest that the raw docking scores don't obviously discriminate binders from non-binders and therefore some more thought is needed on how to incorporate docking results into our predictions. I'd be interested to discuss ideas on how this might be achieved.</p>",
  "messages": [
    {
      "id": "2841642",
      "postDate": "05/28/2024 16:31:15",
      "content": "<p>In this competition, we've seen a variety of methods used, such as <a href=\"https://www.kaggle.com/code/horikitasaku/0-589-tokenize-in-terms-of-chemical-principles\" target=\"_blank\">tokenizing SMILES</a>, <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook\" target=\"_blank\">Convolutional Neural Networks</a> and more traditional feature-based <a href=\"https://www.kaggle.com/code/yujansaya/chemensemble-molecular-binding-with-ensemble\" target=\"_blank\">cheminformatics</a>. None of these use 3D protein-ligand structure information, which suggests that <a href=\"https://www.kaggle.com/code/hideakiogasawara/docking-simulation-using-dockstring\" target=\"_blank\">docking</a> ought to play a complementary role to the other techniques, especially with the likely poor performance of cheminformatics on the non-shared building blocks.</p>\n<p>Nonetheless, <a href=\"https://www.kaggle.com/hideakiogasawara\" target=\"_blank\">@hideakiogasawara</a>'s prelimanary results suggest that the raw docking scores don't obviously discriminate binders from non-binders and therefore some more thought is needed on how to incorporate docking results into our predictions. I'd be interested to discuss ideas on how this might be achieved.</p>",
      "rawMarkdown": "In this competition, we've seen a variety of methods used, such as [tokenizing SMILES](https://www.kaggle.com/code/horikitasaku/0-589-tokenize-in-terms-of-chemical-principles), [Convolutional Neural Networks](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook) and more traditional feature-based [cheminformatics](https://www.kaggle.com/code/yujansaya/chemensemble-molecular-binding-with-ensemble). None of these use 3D protein-ligand structure information, which suggests that [docking](https://www.kaggle.com/code/hideakiogasawara/docking-simulation-using-dockstring) ought to play a complementary role to the other techniques, especially with the likely poor performance of cheminformatics on the non-shared building blocks.\n\nNonetheless, @hideakiogasawara's prelimanary results suggest that the raw docking scores don't obviously discriminate binders from non-binders and therefore some more thought is needed on how to incorporate docking results into our predictions. I'd be interested to discuss ideas on how this might be achieved.",
      "votes": null
    },
    {
      "id": "2842196",
      "postDate": "05/29/2024 00:41:16",
      "content": "<p>The problem with docking is most docking scores are based around \"enrichment\", which is the true positive rate within the top x% of molecules. The idea with docking is you dock a massive library, throw away almost everything and take the top 100 or so molecules for lab testing.</p>\n<p>This approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.</p>\n<p>There are some ML approaches that try to use protein + ligand docked structures to predict affinity or classify bind/no-bind, but these don't perform very well. You can see this paper for more:</p>\n<p><a href=\"https://pubs.acs.org/doi/10.1021/acs.jmedchem.2c00487\" target=\"_blank\">https://pubs.acs.org/doi/10.1021/acs.jmedchem.2c00487</a></p>",
      "rawMarkdown": "The problem with docking is most docking scores are based around \"enrichment\", which is the true positive rate within the top x% of molecules. The idea with docking is you dock a massive library, throw away almost everything and take the top 100 or so molecules for lab testing.\n\nThis approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.\n\nThere are some ML approaches that try to use protein + ligand docked structures to predict affinity or classify bind/no-bind, but these don't perform very well. You can see this paper for more:\n\nhttps://pubs.acs.org/doi/10.1021/acs.jmedchem.2c00487",
      "votes": null
    },
    {
      "id": "2842212",
      "postDate": "05/29/2024 01:06:47",
      "content": "<p>This is an excellent paper. Here is a preprint for anyone outside academia: <a href=\"https://hal.science/hal-03747976/document\" target=\"_blank\">https://hal.science/hal-03747976/document</a></p>",
      "rawMarkdown": "This is an excellent paper. Here is a preprint for anyone outside academia: https://hal.science/hal-03747976/document",
      "votes": null
    },
    {
      "id": "2842488",
      "postDate": "05/29/2024 05:35:26",
      "content": "<p>i think like machine learning, you need to have good experience to produce good docking results.<br>\nI believe that docking should discover some hits that machine learning cannot.</p>\n<p>\"This approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.\"</p>\n<p>if docking can identify some hits, than ML can learn to \"expand these\". a possible solution is to use docking for labeling the unknown for some top x%, then use ML to search for the rest of (100-x)%</p>\n<p>ML only works if we have have data. the molecular chemical space is vary large and i don't think current data is large enough  for generalized model.</p>",
      "rawMarkdown": "i think like machine learning, you need to have good experience to produce good docking results.\nI believe that docking should discover some hits that machine learning cannot.\n\n\"This approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.\"\n\nif docking can identify some hits, than ML can learn to \"expand these\". a possible solution is to use docking for labeling the unknown for some top x%, then use ML to search for the rest of (100-x)%\n\nML only works if we have have data. the molecular chemical space is vary large and i don't think current data is large enough  for generalized model.",
      "votes": null
    },
    {
      "id": "2842899",
      "postDate": "05/29/2024 10:18:12",
      "content": "<p>Thanks for your insightful comment <a href=\"https://www.kaggle.com/towardsentropy\" target=\"_blank\">@towardsentropy</a>. I believe that optimising enrichment (or similarly AUC) isn't really too different from the goal of this competition as per the scoring metric, which is to elevate the ~0.1% of binders up towards the top of our rankings. That is, if we have a way of using docking to identify the top 0.1% of compounds that gives decent overlap of predictions with ground truth, then we could boost the predicted probability of those compounds and hence significantly improve the LB score. </p>\n<p>The challenge I see is to get that kind of accuracy from the quick-and-dirty high-throughput docking methods, or alternatively to perform higher spec computational techniques on a well-chosen subset of compounds and ML-predict the remainder. I'm sure that right now my results are being confounded by molecular size effects, where the non-specific van der Waals or hydrophobic contributions are masking the signal I'm looking for. Some kind of correction for molecular size is possible, though it's possibly conformation-dependent?</p>",
      "rawMarkdown": "Thanks for your insightful comment @towardsentropy. I believe that optimising enrichment (or similarly AUC) isn't really too different from the goal of this competition as per the scoring metric, which is to elevate the ~0.1% of binders up towards the top of our rankings. That is, if we have a way of using docking to identify the top 0.1% of compounds that gives decent overlap of predictions with ground truth, then we could boost the predicted probability of those compounds and hence significantly improve the LB score. \n\nThe challenge I see is to get that kind of accuracy from the quick-and-dirty high-throughput docking methods, or alternatively to perform higher spec computational techniques on a well-chosen subset of compounds and ML-predict the remainder. I'm sure that right now my results are being confounded by molecular size effects, where the non-specific van der Waals or hydrophobic contributions are masking the signal I'm looking for. Some kind of correction for molecular size is possible, though it's possibly conformation-dependent?",
      "votes": null
    },
    {
      "id": "2844140",
      "postDate": "05/30/2024 01:11:42",
      "content": "<p>here are the numbers:</p>\n<pre><code>len(test_df)\n molecules\n\n(( test_df.triazine) &amp;(test_df.share)).()    \n((~.triazine) &amp;(test_df.share)).()    \n(( test_df.triazine) &amp;(test_df.nonshare)).() \n((~.triazine) &amp;(test_df.nonshare)).() \n\n\nnull submission %\ntrain hit fraction = %\ntest (triazine &amp; nonshare) hit fraction = %\n\nwhat we need is to dock   molecules to identify the % hit\n</code></pre>",
      "rawMarkdown": "here are the numbers:\n\n```\nlen(test_df)\n878022 molecules\n\n(( test_df.triazine) &(test_df.share)).mean()    # 0.4203072360373658\n((~test_df.triazine) &(test_df.share)).mean()    # 0\n(( test_df.triazine) &(test_df.nonshare)).mean() # 0.025731701483561915 #22593 molecules\n((~test_df.triazine) &(test_df.nonshare)).mean() # 0.5539610624790723\n\n\nnull submission 2.3%\ntrain hit fraction = 0.5%\ntest (triazine & nonshare) hit fraction = 1.8%\n\nwhat we need is to dock  22593 molecules to identify the 1.8% hit\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2842196,
      "author_name": "towardsentropy",
      "author_url": "",
      "post_date": "05/29/2024 00:41:16",
      "content": "<p>The problem with docking is most docking scores are based around \"enrichment\", which is the true positive rate within the top x% of molecules. The idea with docking is you dock a massive library, throw away almost everything and take the top 100 or so molecules for lab testing.</p>\n<p>This approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.</p>\n<p>There are some ML approaches that try to use protein + ligand docked structures to predict affinity or classify bind/no-bind, but these don't perform very well. You can see this paper for more:</p>\n<p><a href=\"https://pubs.acs.org/doi/10.1021/acs.jmedchem.2c00487\" target=\"_blank\">https://pubs.acs.org/doi/10.1021/acs.jmedchem.2c00487</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2842212,
          "author_name": "andrewdblevins",
          "author_url": "",
          "post_date": "05/29/2024 01:06:47",
          "content": "<p>This is an excellent paper. Here is a preprint for anyone outside academia: <a href=\"https://hal.science/hal-03747976/document\" target=\"_blank\">https://hal.science/hal-03747976/document</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2842899,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "05/29/2024 10:18:12",
          "content": "<p>Thanks for your insightful comment <a href=\"https://www.kaggle.com/towardsentropy\" target=\"_blank\">@towardsentropy</a>. I believe that optimising enrichment (or similarly AUC) isn't really too different from the goal of this competition as per the scoring metric, which is to elevate the ~0.1% of binders up towards the top of our rankings. That is, if we have a way of using docking to identify the top 0.1% of compounds that gives decent overlap of predictions with ground truth, then we could boost the predicted probability of those compounds and hence significantly improve the LB score. </p>\n<p>The challenge I see is to get that kind of accuracy from the quick-and-dirty high-throughput docking methods, or alternatively to perform higher spec computational techniques on a well-chosen subset of compounds and ML-predict the remainder. I'm sure that right now my results are being confounded by molecular size effects, where the non-specific van der Waals or hydrophobic contributions are masking the signal I'm looking for. Some kind of correction for molecular size is possible, though it's possibly conformation-dependent?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2842488,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/29/2024 05:35:26",
      "content": "<p>i think like machine learning, you need to have good experience to produce good docking results.<br>\nI believe that docking should discover some hits that machine learning cannot.</p>\n<p>\"This approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.\"</p>\n<p>if docking can identify some hits, than ML can learn to \"expand these\". a possible solution is to use docking for labeling the unknown for some top x%, then use ML to search for the rest of (100-x)%</p>\n<p>ML only works if we have have data. the molecular chemical space is vary large and i don't think current data is large enough  for generalized model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844140,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/30/2024 01:11:42",
      "content": "<p>here are the numbers:</p>\n<pre><code>len(test_df)\n molecules\n\n(( test_df.triazine) &amp;(test_df.share)).()    \n((~.triazine) &amp;(test_df.share)).()    \n(( test_df.triazine) &amp;(test_df.nonshare)).() \n((~.triazine) &amp;(test_df.nonshare)).() \n\n\nnull submission %\ntrain hit fraction = %\ntest (triazine &amp; nonshare) hit fraction = %\n\nwhat we need is to dock   molecules to identify the % hit\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2841642": "In this competition, we've seen a variety of methods used, such as [tokenizing SMILES](https://www.kaggle.com/code/horikitasaku/0-589-tokenize-in-terms-of-chemical-principles), [Convolutional Neural Networks](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook) and more traditional feature-based [cheminformatics](https://www.kaggle.com/code/yujansaya/chemensemble-molecular-binding-with-ensemble). None of these use 3D protein-ligand structure information, which suggests that [docking](https://www.kaggle.com/code/hideakiogasawara/docking-simulation-using-dockstring) ought to play a complementary role to the other techniques, especially with the likely poor performance of cheminformatics on the non-shared building blocks.\n\nNonetheless, @hideakiogasawara's prelimanary results suggest that the raw docking scores don't obviously discriminate binders from non-binders and therefore some more thought is needed on how to incorporate docking results into our predictions. I'd be interested to discuss ideas on how this might be achieved.",
    "2842196": "The problem with docking is most docking scores are based around \"enrichment\", which is the true positive rate within the top x% of molecules. The idea with docking is you dock a massive library, throw away almost everything and take the top 100 or so molecules for lab testing.\n\nThis approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.\n\nThere are some ML approaches that try to use protein + ligand docked structures to predict affinity or classify bind/no-bind, but these don't perform very well. You can see this paper for more:\n\nhttps://pubs.acs.org/doi/10.1021/acs.jmedchem.2c00487",
    "2842212": "This is an excellent paper. Here is a preprint for anyone outside academia: https://hal.science/hal-03747976/document",
    "2842488": "i think like machine learning, you need to have good experience to produce good docking results.\nI believe that docking should discover some hits that machine learning cannot.\n\n\"This approach makes sense when you're running a pharma lab with limited bandwidth, but doesn't help with a problem like this where we want to separate all molecules into distinct classes. We need to create a model that reliably classifies all molecules, not just the top x%.\"\n\nif docking can identify some hits, than ML can learn to \"expand these\". a possible solution is to use docking for labeling the unknown for some top x%, then use ML to search for the rest of (100-x)%\n\nML only works if we have have data. the molecular chemical space is vary large and i don't think current data is large enough  for generalized model.",
    "2842899": "Thanks for your insightful comment @towardsentropy. I believe that optimising enrichment (or similarly AUC) isn't really too different from the goal of this competition as per the scoring metric, which is to elevate the ~0.1% of binders up towards the top of our rankings. That is, if we have a way of using docking to identify the top 0.1% of compounds that gives decent overlap of predictions with ground truth, then we could boost the predicted probability of those compounds and hence significantly improve the LB score. \n\nThe challenge I see is to get that kind of accuracy from the quick-and-dirty high-throughput docking methods, or alternatively to perform higher spec computational techniques on a well-chosen subset of compounds and ML-predict the remainder. I'm sure that right now my results are being confounded by molecular size effects, where the non-specific van der Waals or hydrophobic contributions are masking the signal I'm looking for. Some kind of correction for molecular size is possible, though it's possibly conformation-dependent?",
    "2844140": "here are the numbers:\n\n```\nlen(test_df)\n878022 molecules\n\n(( test_df.triazine) &(test_df.share)).mean()    # 0.4203072360373658\n((~test_df.triazine) &(test_df.share)).mean()    # 0\n(( test_df.triazine) &(test_df.nonshare)).mean() # 0.025731701483561915 #22593 molecules\n((~test_df.triazine) &(test_df.nonshare)).mean() # 0.5539610624790723\n\n\nnull submission 2.3%\ntrain hit fraction = 0.5%\ntest (triazine & nonshare) hit fraction = 1.8%\n\nwhat we need is to dock  22593 molecules to identify the 1.8% hit\n```"
  },
  "source": "meta"
}