{
  "id": 562096,
  "title": "Docking Performance",
  "url": "/competitions/leash-BELKA/discussion/562096",
  "author_name": "",
  "post_date": "2025-02-09T21:00:58.865695700Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Slowly but surely I'm working through more ways to evaluate this dataset, and this time around, I've evaluated docking.  I only recall a post or two during the competition discussing docking (and I haven't exhaustively looked back though the discussion forum), but the computational burden was generally too large to make it a useful technique at the time.</p>\n<h2>Methods</h2>\n<p>For docking, I used <a href=\"https://github.com/ccsb-scripps/AutoDock-GPU\" target=\"_blank\">AutoDock-GPU v1.6</a>.  AutoDock performs very well in my experience and the GPU enabled version was able to dock about 7500 compounds per hour on my system.</p>\n<p>For each protein target, I selected a reasonable crystal structure to be used for docking.  To identify the binding site, I used the collection of structures from my dataset, <a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins\" target=\"_blank\">LeashBio_BELKA_Proteins</a>.  All crystal structures were aligned and the common binding location of the co-crystallized ligands used to define a binding box for docking.  HSA is more challenging, as there are multiple possible binding sites, 3 of which are particularly well represented with bound ligands.  For HSA, I defined 3 possible binding boxes and performed docking on all 3 independently.  All protein structures were prepared using the <a href=\"https://meeko.readthedocs.io/en/release-doc/#\" target=\"_blank\">Meeko interface for AutoDock</a></p>\n<p>For the training set, I docked a randomly selected 50,000 binders and 50,000 non-binders and used their docking scores to develop a logistic regression model to predict binders vs non-binders.  I tried different numbers of compounds for each group from 10,000 to 50,000 and found no difference in performance with increasing numbers and settled on the largest number tested just to be complete.</p>\n<p>For the test set, I docked all compounds that had binding data for each protein target.  Evaluations were done on docking scores as well as using the Logistic Regression model developed from the test set.</p>\n<h2>Results</h2>\n<p>In the following figures are shown results for the kin0 set of compounds in the test set.  Remember that this was the set of compounds that did not have the common triazine core and had R-groups not represented in the training set.   In other words, the hardest set of compounds.  Generally, though, results for the kin0 set were consistent with those seen for the other sets, public and private, shared and non-shared R-groups.</p>\n<h3>BRD4</h3>\n<p>The figure below illustrates various results.</p>\n<ul>\n<li>The upper left graph shows a boxplot of the actual docking scores (estimated free energy of binding) for all the test compounds.  The median of the binders is slighly lower than that of the non-binders, but not much.  Also, there are plenty of non-binders with docking scores lower than any observed binder.  </li>\n<li>The upper right graph shows a cumulative distribution function (CDF) plot of the predicted probabilities for binders and non-binders.  (The X-axis is the predicted probability of binding, and the y-axis is the probability of seeing that value or lower in the collection of scores.)  As with the docking scores, there is a very small increase in the population of binders probabilities compared to the non-binders.</li>\n<li>The lower left figure illustrates the observed frequency of binding in compounds with probability of binding scores (from the Logistic Regression model) in bins of size 0.05.  For example the tallest bar, centered on 0.6 with a height of 0.0024, indicates that for all the structures with scores from 0.55 to 0.60, the fraction of actives in that set is 0.0024.  The red line at 0.001 is the probability of a randomly selected structure being active.   The non-existent bars with labels of 0 across the bottom are due to no scores in that range occurring - for this target, scores were from around 0.4 to 0.65.  The limited range of predicted probabilities prevents any major conclusions, but there is a relationship between increasing predicted probability of binding (x-axis) and empirical observation of binding (y-axis).</li>\n<li>A more useful way to use docking scores (in my experience) is to selected the top N scoring structures for physical screening.  The bottom-right figure illlustrates what fraction of compounds are binders (y-axis) when N structures are selected (x-axis).  For example, selecting the top 1000 scoring structures gives 0.003 * 1000 = 3 binders.  The red line is again the probability of a randomly selected compound being a binder.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F00d36e22fa2b9996c9e530966fe24d52%2FBRD4_kin0_Private.jpeg?generation=1739136079512993&amp;alt=media\" alt=\"\"></p>\n<p>The interpretation for BRD4 from these results is that while docking scores improves the probability of finding a binder, it is only 3X better than random.  The results were similar for the other subgroups of compounds.</p>\n<h3>sEH</h3>\n<p>The results for sEH were somewhat better.  The separation in docking scores and predicted probability of binding scores were much better than BRD4.  Likewise the probability of finding a binder in set of higher predicted probabilities or by selection of the top 1000 scoring structures was around 9X that of random.  Not a huge boost, but certainly different from purely random selection.  It is also encouraging to see a positive relationship between predicted probability of binding and empirical probability of binding in the lower-left plot.  For other groups of structures in the dataset, the results were generally poorer.  The Public - Shared set (triazines with R-groups shared with the training set) only showed a very small and very probably not significant increase in probabiilty of 0.0176 compared to 0.012 expected at random.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff48d33515d19b425c631754ccf9d5b45%2FsEH_kin0_Private.jpeg?generation=1739136313040310&amp;alt=media\" alt=\"\"></p>\n<h3>HSA</h3>\n<p>HSA was a little trickier due to multiple binding sites.  In the figure below, the upper left image shows the minimum docking score for the 3 binding sites.  For the Logistic Regression model, all three docking scores were used in the model.  Differences in the median of the minimum docking scores is extremely small, and predicted probabilities are effectively no different between binders and non-binders. Actually, the binder probabilities are slightly lower than non-binders, as a population.  Finally, there is only about a 2X increase in probability of binding based on bins or top scoring structures.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Fa44eae551ffb9c399cceea4140bc2aa8%2FHSA_kin0_Private.jpeg?generation=1739137212301548&amp;alt=media\" alt=\"\"></p>\n<h2>Conclusions</h2>\n<p>I wasn't sure what to expect from docking for this set and, honestly, I was a bit underwhelmed.  sEH clearly performed the best, which is likely due to having a well defined, well enclosed binding site.  BRD4 has a well-defined binding site, but it is not as deeply formed.  This is the same situation for HSA.  Despite being the best performing, sEH increased the probability over random selection by at best 9X from these data.  </p>\n<p>Comparing back to my <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/548052\" target=\"_blank\">Fingerprint Exploration</a>, the three proteins performed in a simlar ranking with sEH, BRD4, and HSA showing at best 10.2X, 1.4X, and 7X over random.  Also, the Shared and Nonshared sets performed much, much better with about 60-80% accuracy on the private set.  The caveat being, however, that different fingerprints performed differently across the targets and there was no obvious way to identify <em>a priori</em> which fingerprint would perform best.  </p>\n<p>I appreciate any comments and discussion.</p>",
  "messages": [
    {
      "id": "3119901",
      "postDate": "02/09/2025 21:00:58",
      "content": "<p>Slowly but surely I'm working through more ways to evaluate this dataset, and this time around, I've evaluated docking.  I only recall a post or two during the competition discussing docking (and I haven't exhaustively looked back though the discussion forum), but the computational burden was generally too large to make it a useful technique at the time.</p>\n<h2>Methods</h2>\n<p>For docking, I used <a href=\"https://github.com/ccsb-scripps/AutoDock-GPU\" target=\"_blank\">AutoDock-GPU v1.6</a>.  AutoDock performs very well in my experience and the GPU enabled version was able to dock about 7500 compounds per hour on my system.</p>\n<p>For each protein target, I selected a reasonable crystal structure to be used for docking.  To identify the binding site, I used the collection of structures from my dataset, <a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins\" target=\"_blank\">LeashBio_BELKA_Proteins</a>.  All crystal structures were aligned and the common binding location of the co-crystallized ligands used to define a binding box for docking.  HSA is more challenging, as there are multiple possible binding sites, 3 of which are particularly well represented with bound ligands.  For HSA, I defined 3 possible binding boxes and performed docking on all 3 independently.  All protein structures were prepared using the <a href=\"https://meeko.readthedocs.io/en/release-doc/#\" target=\"_blank\">Meeko interface for AutoDock</a></p>\n<p>For the training set, I docked a randomly selected 50,000 binders and 50,000 non-binders and used their docking scores to develop a logistic regression model to predict binders vs non-binders.  I tried different numbers of compounds for each group from 10,000 to 50,000 and found no difference in performance with increasing numbers and settled on the largest number tested just to be complete.</p>\n<p>For the test set, I docked all compounds that had binding data for each protein target.  Evaluations were done on docking scores as well as using the Logistic Regression model developed from the test set.</p>\n<h2>Results</h2>\n<p>In the following figures are shown results for the kin0 set of compounds in the test set.  Remember that this was the set of compounds that did not have the common triazine core and had R-groups not represented in the training set.   In other words, the hardest set of compounds.  Generally, though, results for the kin0 set were consistent with those seen for the other sets, public and private, shared and non-shared R-groups.</p>\n<h3>BRD4</h3>\n<p>The figure below illustrates various results.</p>\n<ul>\n<li>The upper left graph shows a boxplot of the actual docking scores (estimated free energy of binding) for all the test compounds.  The median of the binders is slighly lower than that of the non-binders, but not much.  Also, there are plenty of non-binders with docking scores lower than any observed binder.  </li>\n<li>The upper right graph shows a cumulative distribution function (CDF) plot of the predicted probabilities for binders and non-binders.  (The X-axis is the predicted probability of binding, and the y-axis is the probability of seeing that value or lower in the collection of scores.)  As with the docking scores, there is a very small increase in the population of binders probabilities compared to the non-binders.</li>\n<li>The lower left figure illustrates the observed frequency of binding in compounds with probability of binding scores (from the Logistic Regression model) in bins of size 0.05.  For example the tallest bar, centered on 0.6 with a height of 0.0024, indicates that for all the structures with scores from 0.55 to 0.60, the fraction of actives in that set is 0.0024.  The red line at 0.001 is the probability of a randomly selected structure being active.   The non-existent bars with labels of 0 across the bottom are due to no scores in that range occurring - for this target, scores were from around 0.4 to 0.65.  The limited range of predicted probabilities prevents any major conclusions, but there is a relationship between increasing predicted probability of binding (x-axis) and empirical observation of binding (y-axis).</li>\n<li>A more useful way to use docking scores (in my experience) is to selected the top N scoring structures for physical screening.  The bottom-right figure illlustrates what fraction of compounds are binders (y-axis) when N structures are selected (x-axis).  For example, selecting the top 1000 scoring structures gives 0.003 * 1000 = 3 binders.  The red line is again the probability of a randomly selected compound being a binder.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F00d36e22fa2b9996c9e530966fe24d52%2FBRD4_kin0_Private.jpeg?generation=1739136079512993&amp;alt=media\" alt=\"\"></p>\n<p>The interpretation for BRD4 from these results is that while docking scores improves the probability of finding a binder, it is only 3X better than random.  The results were similar for the other subgroups of compounds.</p>\n<h3>sEH</h3>\n<p>The results for sEH were somewhat better.  The separation in docking scores and predicted probability of binding scores were much better than BRD4.  Likewise the probability of finding a binder in set of higher predicted probabilities or by selection of the top 1000 scoring structures was around 9X that of random.  Not a huge boost, but certainly different from purely random selection.  It is also encouraging to see a positive relationship between predicted probability of binding and empirical probability of binding in the lower-left plot.  For other groups of structures in the dataset, the results were generally poorer.  The Public - Shared set (triazines with R-groups shared with the training set) only showed a very small and very probably not significant increase in probabiilty of 0.0176 compared to 0.012 expected at random.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff48d33515d19b425c631754ccf9d5b45%2FsEH_kin0_Private.jpeg?generation=1739136313040310&amp;alt=media\" alt=\"\"></p>\n<h3>HSA</h3>\n<p>HSA was a little trickier due to multiple binding sites.  In the figure below, the upper left image shows the minimum docking score for the 3 binding sites.  For the Logistic Regression model, all three docking scores were used in the model.  Differences in the median of the minimum docking scores is extremely small, and predicted probabilities are effectively no different between binders and non-binders. Actually, the binder probabilities are slightly lower than non-binders, as a population.  Finally, there is only about a 2X increase in probability of binding based on bins or top scoring structures.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Fa44eae551ffb9c399cceea4140bc2aa8%2FHSA_kin0_Private.jpeg?generation=1739137212301548&amp;alt=media\" alt=\"\"></p>\n<h2>Conclusions</h2>\n<p>I wasn't sure what to expect from docking for this set and, honestly, I was a bit underwhelmed.  sEH clearly performed the best, which is likely due to having a well defined, well enclosed binding site.  BRD4 has a well-defined binding site, but it is not as deeply formed.  This is the same situation for HSA.  Despite being the best performing, sEH increased the probability over random selection by at best 9X from these data.  </p>\n<p>Comparing back to my <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/548052\" target=\"_blank\">Fingerprint Exploration</a>, the three proteins performed in a simlar ranking with sEH, BRD4, and HSA showing at best 10.2X, 1.4X, and 7X over random.  Also, the Shared and Nonshared sets performed much, much better with about 60-80% accuracy on the private set.  The caveat being, however, that different fingerprints performed differently across the targets and there was no obvious way to identify <em>a priori</em> which fingerprint would perform best.  </p>\n<p>I appreciate any comments and discussion.</p>",
      "rawMarkdown": "Slowly but surely I'm working through more ways to evaluate this dataset, and this time around, I've evaluated docking.  I only recall a post or two during the competition discussing docking (and I haven't exhaustively looked back though the discussion forum), but the computational burden was generally too large to make it a useful technique at the time.\n\n## Methods\n\nFor docking, I used [AutoDock-GPU v1.6](https://github.com/ccsb-scripps/AutoDock-GPU).  AutoDock performs very well in my experience and the GPU enabled version was able to dock about 7500 compounds per hour on my system.\n\nFor each protein target, I selected a reasonable crystal structure to be used for docking.  To identify the binding site, I used the collection of structures from my dataset, [LeashBio_BELKA_Proteins](https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins).  All crystal structures were aligned and the common binding location of the co-crystallized ligands used to define a binding box for docking.  HSA is more challenging, as there are multiple possible binding sites, 3 of which are particularly well represented with bound ligands.  For HSA, I defined 3 possible binding boxes and performed docking on all 3 independently.  All protein structures were prepared using the [Meeko interface for AutoDock](https://meeko.readthedocs.io/en/release-doc/#)\n\nFor the training set, I docked a randomly selected 50,000 binders and 50,000 non-binders and used their docking scores to develop a logistic regression model to predict binders vs non-binders.  I tried different numbers of compounds for each group from 10,000 to 50,000 and found no difference in performance with increasing numbers and settled on the largest number tested just to be complete.\n\nFor the test set, I docked all compounds that had binding data for each protein target.  Evaluations were done on docking scores as well as using the Logistic Regression model developed from the test set.\n\n## Results\n\nIn the following figures are shown results for the kin0 set of compounds in the test set.  Remember that this was the set of compounds that did not have the common triazine core and had R-groups not represented in the training set.   In other words, the hardest set of compounds.  Generally, though, results for the kin0 set were consistent with those seen for the other sets, public and private, shared and non-shared R-groups.\n\n### BRD4\n\nThe figure below illustrates various results.\n\n* The upper left graph shows a boxplot of the actual docking scores (estimated free energy of binding) for all the test compounds.  The median of the binders is slighly lower than that of the non-binders, but not much.  Also, there are plenty of non-binders with docking scores lower than any observed binder.  \n* The upper right graph shows a cumulative distribution function (CDF) plot of the predicted probabilities for binders and non-binders.  (The X-axis is the predicted probability of binding, and the y-axis is the probability of seeing that value or lower in the collection of scores.)  As with the docking scores, there is a very small increase in the population of binders probabilities compared to the non-binders.\n* The lower left figure illustrates the observed frequency of binding in compounds with probability of binding scores (from the Logistic Regression model) in bins of size 0.05.  For example the tallest bar, centered on 0.6 with a height of 0.0024, indicates that for all the structures with scores from 0.55 to 0.60, the fraction of actives in that set is 0.0024.  The red line at 0.001 is the probability of a randomly selected structure being active.   The non-existent bars with labels of 0 across the bottom are due to no scores in that range occurring - for this target, scores were from around 0.4 to 0.65.  The limited range of predicted probabilities prevents any major conclusions, but there is a relationship between increasing predicted probability of binding (x-axis) and empirical observation of binding (y-axis).\n* A more useful way to use docking scores (in my experience) is to selected the top N scoring structures for physical screening.  The bottom-right figure illlustrates what fraction of compounds are binders (y-axis) when N structures are selected (x-axis).  For example, selecting the top 1000 scoring structures gives 0.003 * 1000 = 3 binders.  The red line is again the probability of a randomly selected compound being a binder.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F00d36e22fa2b9996c9e530966fe24d52%2FBRD4_kin0_Private.jpeg?generation=1739136079512993&alt=media)\n\nThe interpretation for BRD4 from these results is that while docking scores improves the probability of finding a binder, it is only 3X better than random.  The results were similar for the other subgroups of compounds.\n\n### sEH\n \nThe results for sEH were somewhat better.  The separation in docking scores and predicted probability of binding scores were much better than BRD4.  Likewise the probability of finding a binder in set of higher predicted probabilities or by selection of the top 1000 scoring structures was around 9X that of random.  Not a huge boost, but certainly different from purely random selection.  It is also encouraging to see a positive relationship between predicted probability of binding and empirical probability of binding in the lower-left plot.  For other groups of structures in the dataset, the results were generally poorer.  The Public - Shared set (triazines with R-groups shared with the training set) only showed a very small and very probably not significant increase in probabiilty of 0.0176 compared to 0.012 expected at random.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff48d33515d19b425c631754ccf9d5b45%2FsEH_kin0_Private.jpeg?generation=1739136313040310&alt=media)\n\n### HSA\n\nHSA was a little trickier due to multiple binding sites.  In the figure below, the upper left image shows the minimum docking score for the 3 binding sites.  For the Logistic Regression model, all three docking scores were used in the model.  Differences in the median of the minimum docking scores is extremely small, and predicted probabilities are effectively no different between binders and non-binders. Actually, the binder probabilities are slightly lower than non-binders, as a population.  Finally, there is only about a 2X increase in probability of binding based on bins or top scoring structures.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Fa44eae551ffb9c399cceea4140bc2aa8%2FHSA_kin0_Private.jpeg?generation=1739137212301548&alt=media)\n\n## Conclusions\n\nI wasn't sure what to expect from docking for this set and, honestly, I was a bit underwhelmed.  sEH clearly performed the best, which is likely due to having a well defined, well enclosed binding site.  BRD4 has a well-defined binding site, but it is not as deeply formed.  This is the same situation for HSA.  Despite being the best performing, sEH increased the probability over random selection by at best 9X from these data.  \n\nComparing back to my [Fingerprint Exploration](https://www.kaggle.com/competitions/leash-BELKA/discussion/548052), the three proteins performed in a simlar ranking with sEH, BRD4, and HSA showing at best 10.2X, 1.4X, and 7X over random.  Also, the Shared and Nonshared sets performed much, much better with about 60-80% accuracy on the private set.  The caveat being, however, that different fingerprints performed differently across the targets and there was no obvious way to identify _a priori_ which fingerprint would perform best.  \n\nI appreciate any comments and discussion.",
      "votes": null
    },
    {
      "id": "3161639",
      "postDate": "03/28/2025 07:48:22",
      "content": "<p>Sorry I just saw the post after two months. Thank you so much for sharing! This is really a lot of work! </p>\n<p>If I understand correctly, it looks from the results that docking doesn’t perform well on this task (e.g., among top 1000 docking scored compounds, less than 10 are binders), no matter non-triazine core or shared BBs. This made me wonder how Shoichet lab can successfully run their docking and find new binders (<a href=\"https://bkslab.org/publications\" target=\"_blank\">https://bkslab.org/publications</a>, they published a lot of docking papers). It’s possible that different docking softwares could perform differently (but shouldn’t too much differences).</p>\n<p>Another possibility could be due to protein structure. The binding pocket changes after a compound binds to it. Here you used ligand free proteins for docking and the pocket here may only fit to a portion of ligands. It’s possible that the pocket structure doesn’t fit to most binders. One solution I can think of is to use diffusion models like diffdock or alphafold3 that doesn’t need to indicate pocket location and can adjust pocket conformations after fitting the compound. I’ll probably give a try when I have time and show the results. </p>",
      "rawMarkdown": "Sorry I just saw the post after two months. Thank you so much for sharing! This is really a lot of work! \n\nIf I understand correctly, it looks from the results that docking doesn’t perform well on this task (e.g., among top 1000 docking scored compounds, less than 10 are binders), no matter non-triazine core or shared BBs. This made me wonder how Shoichet lab can successfully run their docking and find new binders (https://bkslab.org/publications, they published a lot of docking papers). It’s possible that different docking softwares could perform differently (but shouldn’t too much differences).\n\nAnother possibility could be due to protein structure. The binding pocket changes after a compound binds to it. Here you used ligand free proteins for docking and the pocket here may only fit to a portion of ligands. It’s possible that the pocket structure doesn’t fit to most binders. One solution I can think of is to use diffusion models like diffdock or alphafold3 that doesn’t need to indicate pocket location and can adjust pocket conformations after fitting the compound. I’ll probably give a try when I have time and show the results.",
      "votes": null
    },
    {
      "id": "3162808",
      "postDate": "03/29/2025 19:44:31",
      "content": "<p>Thanks for the kind words.</p>\n<p>I was quite surprised docking didn't do very well.  I've tried to rationalize why the performance was so poor.  Here are few ideas.</p>\n<ul>\n<li>Bad setup and unoptimized parameters.  I used default settings and only one target structure for each docking process.  A more thorough optimization of the docking parameters might help.  As you mentioned, I used apo-structures and rigid targets so there could be some degree of induced fit that isn't considered.  Looking at the crystal structures (<a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins\" target=\"_blank\">dataset here</a>), there aren't any major structural changes between holo- and apo-structures, nor are there any major differences for different bound ligands.</li>\n<li>Only one method used - AutoDock.  AutoDock is quite good and I don't think it is necessarily to blame, but as you mentioned, there are other methods out there - GLIDE, GOLD, DiffDock, AlphaFold, Chai-1, Boltz-1, etc.  It would be super interesting to see a head to head comparison of docking methods.</li>\n<li>Incorrect ligand preparation.  I did not include any structural components for the DNA label for the docked compounds.  Based on other threads, the data were generated in multiples with different DNA tags for each run in order to avoid effects of specific DNA elements.  I thought about adding a tri-nucleotide tail at the [Dy] site, but I opted not to, just for simplicity.  I would expect the addition of such a tail would prevent the DNA-labelled R-group from interacting with the binding site and reducing the search space, but I wouldn't have expected it to have that much impact.  </li>\n<li>Nature of the targets.  The three targets, sEH, BRD4, and HSA, are a nice mix of well-defined binding pocket (sEH), moderately defined pocket/region (BRD4), and poorly defined binding site (HSA).  Performance on HSA wasn't particularly surprising to me as there are multiple binding sites in the protein (10 or more - I used 3), and most of them are more hydrophobic surfaces rather than clear binding sites.  I expected sEH to do well, though, given its very clear binding pocket where many ligands are shown to bind in crystal structures.</li>\n<li>Other binding sites.  This always feels like a last option type of claim, but there could certainly be other binding sites that I didn't consider.</li>\n</ul>\n<p>I can't speak for the full compendium of Shoichet papers, but in many cases they are docking very large libraries of extremely diverse compounds.  In my experience, docking can identify molecules that are simply incompatible with the binding site based on size and shape, but cannot discriminate among subtle changes within a congeneric series.  I've often thought that this type of docking approach is more of a size and shape fit exercise, rather than a real binding evaluation.  This dataset consists of very consistent shapes - a large majority with triazine cores and 3 R-groups.  Granted, there is good diversity among the R-groups used which I would have thought would be discriminating for docking.  </p>\n<p>I have not looked at the docking poses for the various binders that were found and those that were missed.  There could be a clue to why the performance was so poor in that data.  I may revisit it at some point when I have time.  Something else I should look at is the diversity of chemical matter that was found to bind.  The 1 in 1000 hit rate seemed consistent across subsets of the data (triazine shared, triazine non-shared, and non-triazine non-shared) but I'm curious how diverse the hits are between those groups.</p>\n<p>I'm also interested in other methods using 3D structures (pharmacophores, etc), in hopes there is a representation that crosses over from the triazines with known R-groups, to the triazines with unshared R-groups and the non-triazines.  While the fingerprint methods did quite well for triazines with known R-groups, they're not telling us more than what we already know.  That's not surprising, as we are really trying to expand our Domain of Applicability with the right representation, but ultimately any model doesn't do well when confronted with data that are out of distribution.  A representation having a distribution that is more generalizable is the Holy Grail, I suppose.</p>\n<p>If you post any further results, please tag me.  I'm still very interested in this dataset.</p>",
      "rawMarkdown": "Thanks for the kind words.\n\nI was quite surprised docking didn't do very well.  I've tried to rationalize why the performance was so poor.  Here are few ideas.\n\n* Bad setup and unoptimized parameters.  I used default settings and only one target structure for each docking process.  A more thorough optimization of the docking parameters might help.  As you mentioned, I used apo-structures and rigid targets so there could be some degree of induced fit that isn't considered.  Looking at the crystal structures ([dataset here](https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins)), there aren't any major structural changes between holo- and apo-structures, nor are there any major differences for different bound ligands.\n* Only one method used - AutoDock.  AutoDock is quite good and I don't think it is necessarily to blame, but as you mentioned, there are other methods out there - GLIDE, GOLD, DiffDock, AlphaFold, Chai-1, Boltz-1, etc.  It would be super interesting to see a head to head comparison of docking methods.\n* Incorrect ligand preparation.  I did not include any structural components for the DNA label for the docked compounds.  Based on other threads, the data were generated in multiples with different DNA tags for each run in order to avoid effects of specific DNA elements.  I thought about adding a tri-nucleotide tail at the [Dy] site, but I opted not to, just for simplicity.  I would expect the addition of such a tail would prevent the DNA-labelled R-group from interacting with the binding site and reducing the search space, but I wouldn't have expected it to have that much impact.  \n* Nature of the targets.  The three targets, sEH, BRD4, and HSA, are a nice mix of well-defined binding pocket (sEH), moderately defined pocket/region (BRD4), and poorly defined binding site (HSA).  Performance on HSA wasn't particularly surprising to me as there are multiple binding sites in the protein (10 or more - I used 3), and most of them are more hydrophobic surfaces rather than clear binding sites.  I expected sEH to do well, though, given its very clear binding pocket where many ligands are shown to bind in crystal structures.\n* Other binding sites.  This always feels like a last option type of claim, but there could certainly be other binding sites that I didn't consider.\n\nI can't speak for the full compendium of Shoichet papers, but in many cases they are docking very large libraries of extremely diverse compounds.  In my experience, docking can identify molecules that are simply incompatible with the binding site based on size and shape, but cannot discriminate among subtle changes within a congeneric series.  I've often thought that this type of docking approach is more of a size and shape fit exercise, rather than a real binding evaluation.  This dataset consists of very consistent shapes - a large majority with triazine cores and 3 R-groups.  Granted, there is good diversity among the R-groups used which I would have thought would be discriminating for docking.  \n\nI have not looked at the docking poses for the various binders that were found and those that were missed.  There could be a clue to why the performance was so poor in that data.  I may revisit it at some point when I have time.  Something else I should look at is the diversity of chemical matter that was found to bind.  The 1 in 1000 hit rate seemed consistent across subsets of the data (triazine shared, triazine non-shared, and non-triazine non-shared) but I'm curious how diverse the hits are between those groups.\n\nI'm also interested in other methods using 3D structures (pharmacophores, etc), in hopes there is a representation that crosses over from the triazines with known R-groups, to the triazines with unshared R-groups and the non-triazines.  While the fingerprint methods did quite well for triazines with known R-groups, they're not telling us more than what we already know.  That's not surprising, as we are really trying to expand our Domain of Applicability with the right representation, but ultimately any model doesn't do well when confronted with data that are out of distribution.  A representation having a distribution that is more generalizable is the Holy Grail, I suppose.\n\nIf you post any further results, please tag me.  I'm still very interested in this dataset.",
      "votes": null
    },
    {
      "id": "3163628",
      "postDate": "03/31/2025 03:41:01",
      "content": "<ul>\n<li>The preparation of protein structures look solid to me. </li>\n<li>Benchmark multiple docking methods would be super interesting, and publishable I think, if you can find anything that works. </li>\n<li>Incorporating tail at the [Dy] site will certainly help and reduce the search space.</li>\n<li>sEH, the better the binding pocket defined, the better docking tends to perform. </li>\n<li>I also agree with your take on docking - docking seems to discriminate the general shape of the compounds, but not the subtle change of the R-groups. If there's any algorithms that can distinguish the subtle change of the R-groups, that will be a game changer.</li>\n<li>For the docking poses, it's possible that the hits (1 in 1000) across the subsets of the data may share common interactions with the proteins, I'm thinking. It worth to look at any way.</li>\n<li>Thanks for reminding me the pharmocophores. One approach could be to convert all ligands into 3D conformers, then generate pharmacophore hypotheses from the actives using softwares like Schrödinger. Then screen the rest of the compounds in this library (some shared BB, unshared BB, and non-triazines) and see the hit rate. This way may also identify the important compound-protein interactions.</li>\n</ul>\n<p>I'll definitely take a look and run a few tests on my end. I'll tag you if updates. </p>",
      "rawMarkdown": "The preparation of protein structures look solid to me. \n- Benchmark multiple docking methods would be super interesting, and publishable I think, if you can find anything that works. \n- Incorporating tail at the [Dy] site will certainly help and reduce the search space.\n- sEH, the better the binding pocket defined, the better docking tends to perform. \n- I also agree with your take on docking - docking seems to discriminate the general shape of the compounds, but not the subtle change of the R-groups. If there's any algorithms that can distinguish the subtle change of the R-groups, that will be a game changer.\n- For the docking poses, it's possible that the hits (1 in 1000) across the subsets of the data may share common interactions with the proteins, I'm thinking. It worth to look at any way.\n- Thanks for reminding me the pharmocophores. One approach could be to convert all ligands into 3D conformers, then generate pharmacophore hypotheses from the actives using softwares like Schrödinger. Then screen the rest of the compounds in this library (some shared BB, unshared BB, and non-triazines) and see the hit rate. This way may also identify the important compound-protein interactions.\n\nI'll definitely take a look and run a few tests on my end. I'll tag you if updates.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3161639,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "03/28/2025 07:48:22",
      "content": "<p>Sorry I just saw the post after two months. Thank you so much for sharing! This is really a lot of work! </p>\n<p>If I understand correctly, it looks from the results that docking doesn’t perform well on this task (e.g., among top 1000 docking scored compounds, less than 10 are binders), no matter non-triazine core or shared BBs. This made me wonder how Shoichet lab can successfully run their docking and find new binders (<a href=\"https://bkslab.org/publications\" target=\"_blank\">https://bkslab.org/publications</a>, they published a lot of docking papers). It’s possible that different docking softwares could perform differently (but shouldn’t too much differences).</p>\n<p>Another possibility could be due to protein structure. The binding pocket changes after a compound binds to it. Here you used ligand free proteins for docking and the pocket here may only fit to a portion of ligands. It’s possible that the pocket structure doesn’t fit to most binders. One solution I can think of is to use diffusion models like diffdock or alphafold3 that doesn’t need to indicate pocket location and can adjust pocket conformations after fitting the compound. I’ll probably give a try when I have time and show the results. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3162808,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "03/29/2025 19:44:31",
          "content": "<p>Thanks for the kind words.</p>\n<p>I was quite surprised docking didn't do very well.  I've tried to rationalize why the performance was so poor.  Here are few ideas.</p>\n<ul>\n<li>Bad setup and unoptimized parameters.  I used default settings and only one target structure for each docking process.  A more thorough optimization of the docking parameters might help.  As you mentioned, I used apo-structures and rigid targets so there could be some degree of induced fit that isn't considered.  Looking at the crystal structures (<a href=\"https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins\" target=\"_blank\">dataset here</a>), there aren't any major structural changes between holo- and apo-structures, nor are there any major differences for different bound ligands.</li>\n<li>Only one method used - AutoDock.  AutoDock is quite good and I don't think it is necessarily to blame, but as you mentioned, there are other methods out there - GLIDE, GOLD, DiffDock, AlphaFold, Chai-1, Boltz-1, etc.  It would be super interesting to see a head to head comparison of docking methods.</li>\n<li>Incorrect ligand preparation.  I did not include any structural components for the DNA label for the docked compounds.  Based on other threads, the data were generated in multiples with different DNA tags for each run in order to avoid effects of specific DNA elements.  I thought about adding a tri-nucleotide tail at the [Dy] site, but I opted not to, just for simplicity.  I would expect the addition of such a tail would prevent the DNA-labelled R-group from interacting with the binding site and reducing the search space, but I wouldn't have expected it to have that much impact.  </li>\n<li>Nature of the targets.  The three targets, sEH, BRD4, and HSA, are a nice mix of well-defined binding pocket (sEH), moderately defined pocket/region (BRD4), and poorly defined binding site (HSA).  Performance on HSA wasn't particularly surprising to me as there are multiple binding sites in the protein (10 or more - I used 3), and most of them are more hydrophobic surfaces rather than clear binding sites.  I expected sEH to do well, though, given its very clear binding pocket where many ligands are shown to bind in crystal structures.</li>\n<li>Other binding sites.  This always feels like a last option type of claim, but there could certainly be other binding sites that I didn't consider.</li>\n</ul>\n<p>I can't speak for the full compendium of Shoichet papers, but in many cases they are docking very large libraries of extremely diverse compounds.  In my experience, docking can identify molecules that are simply incompatible with the binding site based on size and shape, but cannot discriminate among subtle changes within a congeneric series.  I've often thought that this type of docking approach is more of a size and shape fit exercise, rather than a real binding evaluation.  This dataset consists of very consistent shapes - a large majority with triazine cores and 3 R-groups.  Granted, there is good diversity among the R-groups used which I would have thought would be discriminating for docking.  </p>\n<p>I have not looked at the docking poses for the various binders that were found and those that were missed.  There could be a clue to why the performance was so poor in that data.  I may revisit it at some point when I have time.  Something else I should look at is the diversity of chemical matter that was found to bind.  The 1 in 1000 hit rate seemed consistent across subsets of the data (triazine shared, triazine non-shared, and non-triazine non-shared) but I'm curious how diverse the hits are between those groups.</p>\n<p>I'm also interested in other methods using 3D structures (pharmacophores, etc), in hopes there is a representation that crosses over from the triazines with known R-groups, to the triazines with unshared R-groups and the non-triazines.  While the fingerprint methods did quite well for triazines with known R-groups, they're not telling us more than what we already know.  That's not surprising, as we are really trying to expand our Domain of Applicability with the right representation, but ultimately any model doesn't do well when confronted with data that are out of distribution.  A representation having a distribution that is more generalizable is the Holy Grail, I suppose.</p>\n<p>If you post any further results, please tag me.  I'm still very interested in this dataset.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3163628,
              "author_name": "lililycai",
              "author_url": "",
              "post_date": "03/31/2025 03:41:01",
              "content": "<ul>\n<li>The preparation of protein structures look solid to me. </li>\n<li>Benchmark multiple docking methods would be super interesting, and publishable I think, if you can find anything that works. </li>\n<li>Incorporating tail at the [Dy] site will certainly help and reduce the search space.</li>\n<li>sEH, the better the binding pocket defined, the better docking tends to perform. </li>\n<li>I also agree with your take on docking - docking seems to discriminate the general shape of the compounds, but not the subtle change of the R-groups. If there's any algorithms that can distinguish the subtle change of the R-groups, that will be a game changer.</li>\n<li>For the docking poses, it's possible that the hits (1 in 1000) across the subsets of the data may share common interactions with the proteins, I'm thinking. It worth to look at any way.</li>\n<li>Thanks for reminding me the pharmocophores. One approach could be to convert all ligands into 3D conformers, then generate pharmacophore hypotheses from the actives using softwares like Schrödinger. Then screen the rest of the compounds in this library (some shared BB, unshared BB, and non-triazines) and see the hit rate. This way may also identify the important compound-protein interactions.</li>\n</ul>\n<p>I'll definitely take a look and run a few tests on my end. I'll tag you if updates. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3119901": "Slowly but surely I'm working through more ways to evaluate this dataset, and this time around, I've evaluated docking.  I only recall a post or two during the competition discussing docking (and I haven't exhaustively looked back though the discussion forum), but the computational burden was generally too large to make it a useful technique at the time.\n\n## Methods\n\nFor docking, I used [AutoDock-GPU v1.6](https://github.com/ccsb-scripps/AutoDock-GPU).  AutoDock performs very well in my experience and the GPU enabled version was able to dock about 7500 compounds per hour on my system.\n\nFor each protein target, I selected a reasonable crystal structure to be used for docking.  To identify the binding site, I used the collection of structures from my dataset, [LeashBio_BELKA_Proteins](https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins).  All crystal structures were aligned and the common binding location of the co-crystallized ligands used to define a binding box for docking.  HSA is more challenging, as there are multiple possible binding sites, 3 of which are particularly well represented with bound ligands.  For HSA, I defined 3 possible binding boxes and performed docking on all 3 independently.  All protein structures were prepared using the [Meeko interface for AutoDock](https://meeko.readthedocs.io/en/release-doc/#)\n\nFor the training set, I docked a randomly selected 50,000 binders and 50,000 non-binders and used their docking scores to develop a logistic regression model to predict binders vs non-binders.  I tried different numbers of compounds for each group from 10,000 to 50,000 and found no difference in performance with increasing numbers and settled on the largest number tested just to be complete.\n\nFor the test set, I docked all compounds that had binding data for each protein target.  Evaluations were done on docking scores as well as using the Logistic Regression model developed from the test set.\n\n## Results\n\nIn the following figures are shown results for the kin0 set of compounds in the test set.  Remember that this was the set of compounds that did not have the common triazine core and had R-groups not represented in the training set.   In other words, the hardest set of compounds.  Generally, though, results for the kin0 set were consistent with those seen for the other sets, public and private, shared and non-shared R-groups.\n\n### BRD4\n\nThe figure below illustrates various results.\n\n* The upper left graph shows a boxplot of the actual docking scores (estimated free energy of binding) for all the test compounds.  The median of the binders is slighly lower than that of the non-binders, but not much.  Also, there are plenty of non-binders with docking scores lower than any observed binder.  \n* The upper right graph shows a cumulative distribution function (CDF) plot of the predicted probabilities for binders and non-binders.  (The X-axis is the predicted probability of binding, and the y-axis is the probability of seeing that value or lower in the collection of scores.)  As with the docking scores, there is a very small increase in the population of binders probabilities compared to the non-binders.\n* The lower left figure illustrates the observed frequency of binding in compounds with probability of binding scores (from the Logistic Regression model) in bins of size 0.05.  For example the tallest bar, centered on 0.6 with a height of 0.0024, indicates that for all the structures with scores from 0.55 to 0.60, the fraction of actives in that set is 0.0024.  The red line at 0.001 is the probability of a randomly selected structure being active.   The non-existent bars with labels of 0 across the bottom are due to no scores in that range occurring - for this target, scores were from around 0.4 to 0.65.  The limited range of predicted probabilities prevents any major conclusions, but there is a relationship between increasing predicted probability of binding (x-axis) and empirical observation of binding (y-axis).\n* A more useful way to use docking scores (in my experience) is to selected the top N scoring structures for physical screening.  The bottom-right figure illlustrates what fraction of compounds are binders (y-axis) when N structures are selected (x-axis).  For example, selecting the top 1000 scoring structures gives 0.003 * 1000 = 3 binders.  The red line is again the probability of a randomly selected compound being a binder.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2F00d36e22fa2b9996c9e530966fe24d52%2FBRD4_kin0_Private.jpeg?generation=1739136079512993&alt=media)\n\nThe interpretation for BRD4 from these results is that while docking scores improves the probability of finding a binder, it is only 3X better than random.  The results were similar for the other subgroups of compounds.\n\n### sEH\n \nThe results for sEH were somewhat better.  The separation in docking scores and predicted probability of binding scores were much better than BRD4.  Likewise the probability of finding a binder in set of higher predicted probabilities or by selection of the top 1000 scoring structures was around 9X that of random.  Not a huge boost, but certainly different from purely random selection.  It is also encouraging to see a positive relationship between predicted probability of binding and empirical probability of binding in the lower-left plot.  For other groups of structures in the dataset, the results were generally poorer.  The Public - Shared set (triazines with R-groups shared with the training set) only showed a very small and very probably not significant increase in probabiilty of 0.0176 compared to 0.012 expected at random.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Ff48d33515d19b425c631754ccf9d5b45%2FsEH_kin0_Private.jpeg?generation=1739136313040310&alt=media)\n\n### HSA\n\nHSA was a little trickier due to multiple binding sites.  In the figure below, the upper left image shows the minimum docking score for the 3 binding sites.  For the Logistic Regression model, all three docking scores were used in the model.  Differences in the median of the minimum docking scores is extremely small, and predicted probabilities are effectively no different between binders and non-binders. Actually, the binder probabilities are slightly lower than non-binders, as a population.  Finally, there is only about a 2X increase in probability of binding based on bins or top scoring structures.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F779570%2Fa44eae551ffb9c399cceea4140bc2aa8%2FHSA_kin0_Private.jpeg?generation=1739137212301548&alt=media)\n\n## Conclusions\n\nI wasn't sure what to expect from docking for this set and, honestly, I was a bit underwhelmed.  sEH clearly performed the best, which is likely due to having a well defined, well enclosed binding site.  BRD4 has a well-defined binding site, but it is not as deeply formed.  This is the same situation for HSA.  Despite being the best performing, sEH increased the probability over random selection by at best 9X from these data.  \n\nComparing back to my [Fingerprint Exploration](https://www.kaggle.com/competitions/leash-BELKA/discussion/548052), the three proteins performed in a simlar ranking with sEH, BRD4, and HSA showing at best 10.2X, 1.4X, and 7X over random.  Also, the Shared and Nonshared sets performed much, much better with about 60-80% accuracy on the private set.  The caveat being, however, that different fingerprints performed differently across the targets and there was no obvious way to identify _a priori_ which fingerprint would perform best.  \n\nI appreciate any comments and discussion.",
    "3161639": "Sorry I just saw the post after two months. Thank you so much for sharing! This is really a lot of work! \n\nIf I understand correctly, it looks from the results that docking doesn’t perform well on this task (e.g., among top 1000 docking scored compounds, less than 10 are binders), no matter non-triazine core or shared BBs. This made me wonder how Shoichet lab can successfully run their docking and find new binders (https://bkslab.org/publications, they published a lot of docking papers). It’s possible that different docking softwares could perform differently (but shouldn’t too much differences).\n\nAnother possibility could be due to protein structure. The binding pocket changes after a compound binds to it. Here you used ligand free proteins for docking and the pocket here may only fit to a portion of ligands. It’s possible that the pocket structure doesn’t fit to most binders. One solution I can think of is to use diffusion models like diffdock or alphafold3 that doesn’t need to indicate pocket location and can adjust pocket conformations after fitting the compound. I’ll probably give a try when I have time and show the results.",
    "3162808": "Thanks for the kind words.\n\nI was quite surprised docking didn't do very well.  I've tried to rationalize why the performance was so poor.  Here are few ideas.\n\n* Bad setup and unoptimized parameters.  I used default settings and only one target structure for each docking process.  A more thorough optimization of the docking parameters might help.  As you mentioned, I used apo-structures and rigid targets so there could be some degree of induced fit that isn't considered.  Looking at the crystal structures ([dataset here](https://www.kaggle.com/datasets/kirkdco/leashbio-belka-proteins)), there aren't any major structural changes between holo- and apo-structures, nor are there any major differences for different bound ligands.\n* Only one method used - AutoDock.  AutoDock is quite good and I don't think it is necessarily to blame, but as you mentioned, there are other methods out there - GLIDE, GOLD, DiffDock, AlphaFold, Chai-1, Boltz-1, etc.  It would be super interesting to see a head to head comparison of docking methods.\n* Incorrect ligand preparation.  I did not include any structural components for the DNA label for the docked compounds.  Based on other threads, the data were generated in multiples with different DNA tags for each run in order to avoid effects of specific DNA elements.  I thought about adding a tri-nucleotide tail at the [Dy] site, but I opted not to, just for simplicity.  I would expect the addition of such a tail would prevent the DNA-labelled R-group from interacting with the binding site and reducing the search space, but I wouldn't have expected it to have that much impact.  \n* Nature of the targets.  The three targets, sEH, BRD4, and HSA, are a nice mix of well-defined binding pocket (sEH), moderately defined pocket/region (BRD4), and poorly defined binding site (HSA).  Performance on HSA wasn't particularly surprising to me as there are multiple binding sites in the protein (10 or more - I used 3), and most of them are more hydrophobic surfaces rather than clear binding sites.  I expected sEH to do well, though, given its very clear binding pocket where many ligands are shown to bind in crystal structures.\n* Other binding sites.  This always feels like a last option type of claim, but there could certainly be other binding sites that I didn't consider.\n\nI can't speak for the full compendium of Shoichet papers, but in many cases they are docking very large libraries of extremely diverse compounds.  In my experience, docking can identify molecules that are simply incompatible with the binding site based on size and shape, but cannot discriminate among subtle changes within a congeneric series.  I've often thought that this type of docking approach is more of a size and shape fit exercise, rather than a real binding evaluation.  This dataset consists of very consistent shapes - a large majority with triazine cores and 3 R-groups.  Granted, there is good diversity among the R-groups used which I would have thought would be discriminating for docking.  \n\nI have not looked at the docking poses for the various binders that were found and those that were missed.  There could be a clue to why the performance was so poor in that data.  I may revisit it at some point when I have time.  Something else I should look at is the diversity of chemical matter that was found to bind.  The 1 in 1000 hit rate seemed consistent across subsets of the data (triazine shared, triazine non-shared, and non-triazine non-shared) but I'm curious how diverse the hits are between those groups.\n\nI'm also interested in other methods using 3D structures (pharmacophores, etc), in hopes there is a representation that crosses over from the triazines with known R-groups, to the triazines with unshared R-groups and the non-triazines.  While the fingerprint methods did quite well for triazines with known R-groups, they're not telling us more than what we already know.  That's not surprising, as we are really trying to expand our Domain of Applicability with the right representation, but ultimately any model doesn't do well when confronted with data that are out of distribution.  A representation having a distribution that is more generalizable is the Holy Grail, I suppose.\n\nIf you post any further results, please tag me.  I'm still very interested in this dataset.",
    "3163628": "The preparation of protein structures look solid to me. \n- Benchmark multiple docking methods would be super interesting, and publishable I think, if you can find anything that works. \n- Incorporating tail at the [Dy] site will certainly help and reduce the search space.\n- sEH, the better the binding pocket defined, the better docking tends to perform. \n- I also agree with your take on docking - docking seems to discriminate the general shape of the compounds, but not the subtle change of the R-groups. If there's any algorithms that can distinguish the subtle change of the R-groups, that will be a game changer.\n- For the docking poses, it's possible that the hits (1 in 1000) across the subsets of the data may share common interactions with the proteins, I'm thinking. It worth to look at any way.\n- Thanks for reminding me the pharmocophores. One approach could be to convert all ligands into 3D conformers, then generate pharmacophore hypotheses from the actives using softwares like Schrödinger. Then screen the rest of the compounds in this library (some shared BB, unshared BB, and non-triazines) and see the hit rate. This way may also identify the important compound-protein interactions.\n\nI'll definitely take a look and run a few tests on my end. I'll tag you if updates."
  },
  "source": "meta"
}