{
  "id": 493297,
  "title": "Open Source Docking and Conformer Generating Libraries",
  "url": "/competitions/leash-BELKA/discussion/493297",
  "author_name": "",
  "post_date": "2024-04-12T20:40:22.300574500Z",
  "votes": 14,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>Given that our task is to build a model capable of successfully achieving scaffold hopping (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/493294)\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/493294)</a>, let's use this thread as a place to keep open-source docking and conformer generation implementations.</p>\n<p>I'm personally only really aware of commercial or proprietary software for this task, so I'm pretty excited to learn about open source implementations!</p>\n<p>The organizers got us started with DiffDock: <a href=\"https://arxiv.org/abs/2210.01776\" target=\"_blank\">https://arxiv.org/abs/2210.01776</a> (paper) GitHub (<a href=\"https://github.com/gcorso/DiffDock)\" target=\"_blank\">https://github.com/gcorso/DiffDock)</a>.</p>\n<p>For generating 3D conformers of the molecules we have 2D representations for (SMILES) I know that RDKit has some functionality for this, though I've not personally used it yet. I believe the conformers generated in RDKit are not physics based (i.e. no force-field applied) and are instead based on heuristics, though I could be wrong.</p>\n<p>I'll update as I find/learn more, but please feel free to share your findings as well!</p>",
  "messages": [
    {
      "id": "2749100",
      "postDate": "04/12/2024 20:40:22",
      "content": "<p>Hi Everyone,</p>\n<p>Given that our task is to build a model capable of successfully achieving scaffold hopping (<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/493294)\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/493294)</a>, let's use this thread as a place to keep open-source docking and conformer generation implementations.</p>\n<p>I'm personally only really aware of commercial or proprietary software for this task, so I'm pretty excited to learn about open source implementations!</p>\n<p>The organizers got us started with DiffDock: <a href=\"https://arxiv.org/abs/2210.01776\" target=\"_blank\">https://arxiv.org/abs/2210.01776</a> (paper) GitHub (<a href=\"https://github.com/gcorso/DiffDock)\" target=\"_blank\">https://github.com/gcorso/DiffDock)</a>.</p>\n<p>For generating 3D conformers of the molecules we have 2D representations for (SMILES) I know that RDKit has some functionality for this, though I've not personally used it yet. I believe the conformers generated in RDKit are not physics based (i.e. no force-field applied) and are instead based on heuristics, though I could be wrong.</p>\n<p>I'll update as I find/learn more, but please feel free to share your findings as well!</p>",
      "rawMarkdown": "Hi Everyone,\n\nGiven that our task is to build a model capable of successfully achieving scaffold hopping (https://www.kaggle.com/competitions/leash-BELKA/discussion/493294), let's use this thread as a place to keep open-source docking and conformer generation implementations.\n\nI'm personally only really aware of commercial or proprietary software for this task, so I'm pretty excited to learn about open source implementations!\n\nThe organizers got us started with DiffDock: https://arxiv.org/abs/2210.01776 (paper) GitHub (https://github.com/gcorso/DiffDock).\n\nFor generating 3D conformers of the molecules we have 2D representations for (SMILES) I know that RDKit has some functionality for this, though I've not personally used it yet. I believe the conformers generated in RDKit are not physics based (i.e. no force-field applied) and are instead based on heuristics, though I could be wrong.\n\nI'll update as I find/learn more, but please feel free to share your findings as well!",
      "votes": null
    },
    {
      "id": "2749105",
      "postDate": "04/12/2024 20:46:25",
      "content": "<p>This blog is pretty useful for working with conformer objects with RDKit: <a href=\"https://greglandrum.github.io/rdkit-blog/posts/2023-02-04-working-with-conformers.html\" target=\"_blank\">https://greglandrum.github.io/rdkit-blog/posts/2023-02-04-working-with-conformers.html</a></p>\n<p>Note: you can work with conformer objects in RDKit even if the 3D structure wasn't generated with RDKit.</p>",
      "rawMarkdown": "This blog is pretty useful for working with conformer objects with RDKit: https://greglandrum.github.io/rdkit-blog/posts/2023-02-04-working-with-conformers.html\n\nNote: you can work with conformer objects in RDKit even if the 3D structure wasn't generated with RDKit.",
      "votes": null
    },
    {
      "id": "2749123",
      "postDate": "04/12/2024 21:00:10",
      "content": "<p>Rdkit generates conformers based on distance geometry (e.g. ideal bond lengths &amp; angles) but it is simple to apply force-field minimization afterwards</p>\n<pre><code>AllChem.  # generates conformer\nAllChem.  # minimizes conformer\n</code></pre>\n<p>For docking, there are a number of open source tools: AutoDock Vina, Smina, SwissDock, as well as the deep-learning docking tools: EquiBind, LigPose, CarsiDock, DynamicBind, KarmaDock, RFAA, EDM-Dock, and so on. It is prohibitive to dock every ligand so one has to be creative when applying these tools ;)</p>",
      "rawMarkdown": "Rdkit generates conformers based on distance geometry (e.g. ideal bond lengths & angles) but it is simple to apply force-field minimization afterwards\n```\nAllChem.EmbedMolecule(mol)  # generates conformer\nAllChem.MMFFOptimizeMolecule(mol)  # minimizes conformer\n```\nFor docking, there are a number of open source tools: AutoDock Vina, Smina, SwissDock, as well as the deep-learning docking tools: EquiBind, LigPose, CarsiDock, DynamicBind, KarmaDock, RFAA, EDM-Dock, and so on. It is prohibitive to dock every ligand so one has to be creative when applying these tools ;)",
      "votes": null
    },
    {
      "id": "2749128",
      "postDate": "04/12/2024 21:08:32",
      "content": "<p>Thanks for the clarification on RDKit and the addition of free to use tools <a href=\"https://www.kaggle.com/matthewmasters\" target=\"_blank\">@matthewmasters</a> ! Looks like there's a bunch to get into here.</p>",
      "rawMarkdown": "Thanks for the clarification on RDKit and the addition of free to use tools @matthewmasters ! Looks like there's a bunch to get into here.",
      "votes": null
    },
    {
      "id": "2758706",
      "postDate": "04/18/2024 09:56:09",
      "content": "<p>What do you mean it is prohibitive? It would take to much computational resources?</p>",
      "rawMarkdown": "What do you mean it is prohibitive? It would take to much computational resources?",
      "votes": null
    },
    {
      "id": "2769231",
      "postDate": "04/23/2024 08:20:33",
      "content": "<p>anyone what to use opensource docking tools to create a list of binding sites/pockets for the 3 target protein and share with the public?</p>",
      "rawMarkdown": "anyone what to use opensource docking tools to create a list of binding sites/pockets for the 3 target protein and share with the public?",
      "votes": null
    },
    {
      "id": "2771753",
      "postDate": "04/24/2024 11:51:06",
      "content": "<p>i am getting 3d coordinates for input to 3d graph GNN (SMP spherical message passing).<br>\nHere is how i use rdkit ETKDG algo<br>\n(i later plan to search github for fast DFT approcimation with neural nets)</p>\n<pre><code>smiles = [\n    ,\n]\n\ncoordinates=[]\ns = smiles[0]\nmol = Chem.MolFromSmiles(s)\n\nmol = Chem.AddHs(mol)\n\n\n    ps = AllChem.ETKDGv2()\n    ps.useRandomCoords = True\n    AllChem.EmbedMolecule(mol, ps)\n    AllChem.MMFFOptimizeMolecule(mol, confId=0)\n    conf = mol.GetConformer()\n    c = conf.GetPositions()\n    coordinates.append(c)\n\n    print('error',s )\n\n\n\n</code></pre>\n<p>not sure if the fake Dy will affect results ….</p>",
      "rawMarkdown": "i am getting 3d coordinates for input to 3d graph GNN (SMP spherical message passing).\nHere is how i use rdkit ETKDG algo\n(i later plan to search github for fast DFT approcimation with neural nets)\n\n```\n\nsmiles = [\n\t\"C#CCOc1ccc(CNc2nc(NCc3cccc(Br)n3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1\",\n]\n\ncoordinates=[]\ns = smiles[0]\nmol = Chem.MolFromSmiles(s)\n# add hydrogen bonds to molecule because they are not in the smiles representation\nmol = Chem.AddHs(mol)\n\ntry:\n\tps = AllChem.ETKDGv2()\n\tps.useRandomCoords = True\n\tAllChem.EmbedMolecule(mol, ps)\n\tAllChem.MMFFOptimizeMolecule(mol, confId=0)\n\tconf = mol.GetConformer()\n\tc = conf.GetPositions()\n\tcoordinates.append(c)\nexcept:\n\tprint('error',s )\n\n[19:47:30] UFFTYPER: Unrecognized charge state for atom: 33\n[19:47:30] UFFTYPER: Unrecognized atom type: Dy5+3 (33)\n```\n\nnot sure if the fake Dy will affect results ....",
      "votes": null
    },
    {
      "id": "2771782",
      "postDate": "04/24/2024 12:03:50",
      "content": "<p>Thanks for sharing this here <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! </p>\n<p>A couple notes that might be helpful.</p>\n<p>1) you may want to generate multiple conformers and check for the lowest energy one/ones.</p>\n<p>2) You are correct, the Dy is not parameterized by this force field, so it's unlikely the MMFF will be able to give you accurate energies with this group attached.</p>",
      "rawMarkdown": "Thanks for sharing this here @hengck23 ! \n\nA couple notes that might be helpful.\n\n1) you may want to generate multiple conformers and check for the lowest energy one/ones.\n\n2) You are correct, the Dy is not parameterized by this force field, so it's unlikely the MMFF will be able to give you accurate energies with this group attached.",
      "votes": null
    },
    {
      "id": "2772336",
      "postDate": "04/24/2024 16:57:30",
      "content": "<p><a href=\"https://www.kaggle.com/gyulamaloveczky4\" target=\"_blank\">@gyulamaloveczky4</a> Sorry, didn't see this before. Yes, too many resources for the average kaggler. DiffDock for example takes around 10 seconds per molecule. Just to dock the test set would take something like 200 days of GPU time. And then you just have poses and scores. You still have to develop a classifier on top of it, meaning you have to dock training molecules too. Nevertheless, I still think docking can be useful. For example, to find the bioactive conformation of a known binder or to confirm a predicted binder. There are also some clever ideas out there in the docking/virtual screening literature</p>",
      "rawMarkdown": "gyulamaloveczky4 Sorry, didn't see this before. Yes, too many resources for the average kaggler. DiffDock for example takes around 10 seconds per molecule. Just to dock the test set would take something like 200 days of GPU time. And then you just have poses and scores. You still have to develop a classifier on top of it, meaning you have to dock training molecules too. Nevertheless, I still think docking can be useful. For example, to find the bioactive conformation of a known binder or to confirm a predicted binder. There are also some clever ideas out there in the docking/virtual screening literature",
      "votes": null
    },
    {
      "id": "2773473",
      "postDate": "04/24/2024 18:54:08",
      "content": "<p><a href=\"https://www.eyesopen.com/omega\" target=\"_blank\">https://www.eyesopen.com/omega</a><br>\nOMEGA samples the conformational space of drug-like molecules at speeds of hundreds of thousands of compounds per day on a CPU, and over a million molecules per day on a GPU.<br>\njust 1M per day!!!! (we have 100M)</p>",
      "rawMarkdown": "https://www.eyesopen.com/omega\nOMEGA samples the conformational space of drug-like molecules at speeds of hundreds of thousands of compounds per day on a CPU, and over a million molecules per day on a GPU.\njust 1M per day!!!! (we have 100M)",
      "votes": null
    },
    {
      "id": "2773544",
      "postDate": "04/24/2024 19:24:18",
      "content": "<p>For the task of binding site prediction, I would use P2Rank or DiffDock. Although DiffDock was designed to dock molecules, it's performance mostly comes from it's pocket finding ability. Paper: <a href=\"https://arxiv.org/pdf/2302.07134\" target=\"_blank\">Do Deep Learning Models Really Outperform\nTraditional Approaches in Molecular Docking?</a></p>",
      "rawMarkdown": "For the task of binding site prediction, I would use P2Rank or DiffDock. Although DiffDock was designed to dock molecules, it's performance mostly comes from it's pocket finding ability. Paper: [Do Deep Learning Models Really Outperform\nTraditional Approaches in Molecular Docking?](https://arxiv.org/pdf/2302.07134)",
      "votes": null
    },
    {
      "id": "2773937",
      "postDate": "04/25/2024 01:21:10",
      "content": "<p>High-Quality Conformer Generation with CONFORGE:Algorithm and Performance Assessment<br>\n<a href=\"https://pubs.acs.org/doi/epdf/10.1021/acs.jcim.3c00563\" target=\"_blank\">https://pubs.acs.org/doi/epdf/10.1021/acs.jcim.3c00563</a></p>\n<p>compare open-source CONFORGE with commercial OMEGA and CORINA in the paper. seems good.</p>",
      "rawMarkdown": "High-Quality Conformer Generation with CONFORGE:Algorithm and Performance Assessment\nhttps://pubs.acs.org/doi/epdf/10.1021/acs.jcim.3c00563\n\ncompare open-source CONFORGE with commercial OMEGA and CORINA in the paper. seems good.",
      "votes": null
    },
    {
      "id": "2776071",
      "postDate": "04/26/2024 01:12:21",
      "content": "<p>Hey everyone, if you're looking to explore different DNA attachment points (perhaps for building 3D models or whatever else you think it might be useful for), I wrote a notebook to help you get started with your experimentation: <a href=\"https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations</a>.</p>\n<p>Best of luck!</p>",
      "rawMarkdown": "Hey everyone, if you're looking to explore different DNA attachment points (perhaps for building 3D models or whatever else you think it might be useful for), I wrote a notebook to help you get started with your experimentation: https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations.\n\nBest of luck!",
      "votes": null
    },
    {
      "id": "2776097",
      "postDate": "04/26/2024 01:54:45",
      "content": "<p>thanks a lot for the code! ‎‎‎‎              </p>",
      "rawMarkdown": "thanks a lot for the code! ‎‎‎‎",
      "votes": null
    },
    {
      "id": "2776098",
      "postDate": "04/26/2024 01:56:19",
      "content": "<p>You're very welcome!</p>",
      "rawMarkdown": "You're very welcome!",
      "votes": null
    },
    {
      "id": "2776122",
      "postDate": "04/26/2024 02:26:23",
      "content": "<p>how to find pockets<br>\n<a href=\"https://github.com/yazdanimehdi/AttentionSiteDTI/blob/main/human_data.py\" target=\"_blank\">https://github.com/yazdanimehdi/AttentionSiteDTI/blob/main/human_data.py</a></p>\n<pre><code>pk = deepchem.dock.\n\ndef process:\n    m = Chem.\n    am = \n    pockets = pk.find\n    n2 = m.\n</code></pre>",
      "rawMarkdown": "how to find pockets\nhttps://github.com/yazdanimehdi/AttentionSiteDTI/blob/main/human_data.py\n\n```\npk = deepchem.dock.ConvexHullPocketFinder()\n\ndef process_protein(pdb_file):\n    m = Chem.MolFromPDBFile(pdb_file)\n    am = GetAdjacencyMatrix(m)\n    pockets = pk.find_pockets(pdb_file)\n    n2 = m.GetNumAtoms()\n\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2749105,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "04/12/2024 20:46:25",
      "content": "<p>This blog is pretty useful for working with conformer objects with RDKit: <a href=\"https://greglandrum.github.io/rdkit-blog/posts/2023-02-04-working-with-conformers.html\" target=\"_blank\">https://greglandrum.github.io/rdkit-blog/posts/2023-02-04-working-with-conformers.html</a></p>\n<p>Note: you can work with conformer objects in RDKit even if the 3D structure wasn't generated with RDKit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2749123,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "04/12/2024 21:00:10",
      "content": "<p>Rdkit generates conformers based on distance geometry (e.g. ideal bond lengths &amp; angles) but it is simple to apply force-field minimization afterwards</p>\n<pre><code>AllChem.  # generates conformer\nAllChem.  # minimizes conformer\n</code></pre>\n<p>For docking, there are a number of open source tools: AutoDock Vina, Smina, SwissDock, as well as the deep-learning docking tools: EquiBind, LigPose, CarsiDock, DynamicBind, KarmaDock, RFAA, EDM-Dock, and so on. It is prohibitive to dock every ligand so one has to be creative when applying these tools ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2749128,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/12/2024 21:08:32",
          "content": "<p>Thanks for the clarification on RDKit and the addition of free to use tools <a href=\"https://www.kaggle.com/matthewmasters\" target=\"_blank\">@matthewmasters</a> ! Looks like there's a bunch to get into here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2758706,
          "author_name": "gyulamaloveczky4",
          "author_url": "",
          "post_date": "04/18/2024 09:56:09",
          "content": "<p>What do you mean it is prohibitive? It would take to much computational resources?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2772336,
              "author_name": "matthewmasters",
              "author_url": "",
              "post_date": "04/24/2024 16:57:30",
              "content": "<p><a href=\"https://www.kaggle.com/gyulamaloveczky4\" target=\"_blank\">@gyulamaloveczky4</a> Sorry, didn't see this before. Yes, too many resources for the average kaggler. DiffDock for example takes around 10 seconds per molecule. Just to dock the test set would take something like 200 days of GPU time. And then you just have poses and scores. You still have to develop a classifier on top of it, meaning you have to dock training molecules too. Nevertheless, I still think docking can be useful. For example, to find the bioactive conformation of a known binder or to confirm a predicted binder. There are also some clever ideas out there in the docking/virtual screening literature</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2773473,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "04/24/2024 18:54:08",
                  "content": "<p><a href=\"https://www.eyesopen.com/omega\" target=\"_blank\">https://www.eyesopen.com/omega</a><br>\nOMEGA samples the conformational space of drug-like molecules at speeds of hundreds of thousands of compounds per day on a CPU, and over a million molecules per day on a GPU.<br>\njust 1M per day!!!! (we have 100M)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2769231,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/23/2024 08:20:33",
      "content": "<p>anyone what to use opensource docking tools to create a list of binding sites/pockets for the 3 target protein and share with the public?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2773544,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "04/24/2024 19:24:18",
          "content": "<p>For the task of binding site prediction, I would use P2Rank or DiffDock. Although DiffDock was designed to dock molecules, it's performance mostly comes from it's pocket finding ability. Paper: <a href=\"https://arxiv.org/pdf/2302.07134\" target=\"_blank\">Do Deep Learning Models Really Outperform\nTraditional Approaches in Molecular Docking?</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2776122,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "04/26/2024 02:26:23",
              "content": "<p>how to find pockets<br>\n<a href=\"https://github.com/yazdanimehdi/AttentionSiteDTI/blob/main/human_data.py\" target=\"_blank\">https://github.com/yazdanimehdi/AttentionSiteDTI/blob/main/human_data.py</a></p>\n<pre><code>pk = deepchem.dock.\n\ndef process:\n    m = Chem.\n    am = \n    pockets = pk.find\n    n2 = m.\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2771753,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/24/2024 11:51:06",
      "content": "<p>i am getting 3d coordinates for input to 3d graph GNN (SMP spherical message passing).<br>\nHere is how i use rdkit ETKDG algo<br>\n(i later plan to search github for fast DFT approcimation with neural nets)</p>\n<pre><code>smiles = [\n    ,\n]\n\ncoordinates=[]\ns = smiles[0]\nmol = Chem.MolFromSmiles(s)\n\nmol = Chem.AddHs(mol)\n\n\n    ps = AllChem.ETKDGv2()\n    ps.useRandomCoords = True\n    AllChem.EmbedMolecule(mol, ps)\n    AllChem.MMFFOptimizeMolecule(mol, confId=0)\n    conf = mol.GetConformer()\n    c = conf.GetPositions()\n    coordinates.append(c)\n\n    print('error',s )\n\n\n\n</code></pre>\n<p>not sure if the fake Dy will affect results ….</p>",
      "votes": null,
      "replies": [
        {
          "id": 2771782,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/24/2024 12:03:50",
          "content": "<p>Thanks for sharing this here <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! </p>\n<p>A couple notes that might be helpful.</p>\n<p>1) you may want to generate multiple conformers and check for the lowest energy one/ones.</p>\n<p>2) You are correct, the Dy is not parameterized by this force field, so it's unlikely the MMFF will be able to give you accurate energies with this group attached.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2773937,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/25/2024 01:21:10",
      "content": "<p>High-Quality Conformer Generation with CONFORGE:Algorithm and Performance Assessment<br>\n<a href=\"https://pubs.acs.org/doi/epdf/10.1021/acs.jcim.3c00563\" target=\"_blank\">https://pubs.acs.org/doi/epdf/10.1021/acs.jcim.3c00563</a></p>\n<p>compare open-source CONFORGE with commercial OMEGA and CORINA in the paper. seems good.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2776071,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "04/26/2024 01:12:21",
      "content": "<p>Hey everyone, if you're looking to explore different DNA attachment points (perhaps for building 3D models or whatever else you think it might be useful for), I wrote a notebook to help you get started with your experimentation: <a href=\"https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations\" target=\"_blank\">https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations</a>.</p>\n<p>Best of luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2776097,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/26/2024 01:54:45",
          "content": "<p>thanks a lot for the code! ‎‎‎‎              </p>",
          "votes": null,
          "replies": [
            {
              "id": 2776098,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "04/26/2024 01:56:19",
              "content": "<p>You're very welcome!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2749100": "Hi Everyone,\n\nGiven that our task is to build a model capable of successfully achieving scaffold hopping (https://www.kaggle.com/competitions/leash-BELKA/discussion/493294), let's use this thread as a place to keep open-source docking and conformer generation implementations.\n\nI'm personally only really aware of commercial or proprietary software for this task, so I'm pretty excited to learn about open source implementations!\n\nThe organizers got us started with DiffDock: https://arxiv.org/abs/2210.01776 (paper) GitHub (https://github.com/gcorso/DiffDock).\n\nFor generating 3D conformers of the molecules we have 2D representations for (SMILES) I know that RDKit has some functionality for this, though I've not personally used it yet. I believe the conformers generated in RDKit are not physics based (i.e. no force-field applied) and are instead based on heuristics, though I could be wrong.\n\nI'll update as I find/learn more, but please feel free to share your findings as well!",
    "2749105": "This blog is pretty useful for working with conformer objects with RDKit: https://greglandrum.github.io/rdkit-blog/posts/2023-02-04-working-with-conformers.html\n\nNote: you can work with conformer objects in RDKit even if the 3D structure wasn't generated with RDKit.",
    "2749123": "Rdkit generates conformers based on distance geometry (e.g. ideal bond lengths & angles) but it is simple to apply force-field minimization afterwards\n```\nAllChem.EmbedMolecule(mol)  # generates conformer\nAllChem.MMFFOptimizeMolecule(mol)  # minimizes conformer\n```\nFor docking, there are a number of open source tools: AutoDock Vina, Smina, SwissDock, as well as the deep-learning docking tools: EquiBind, LigPose, CarsiDock, DynamicBind, KarmaDock, RFAA, EDM-Dock, and so on. It is prohibitive to dock every ligand so one has to be creative when applying these tools ;)",
    "2749128": "Thanks for the clarification on RDKit and the addition of free to use tools @matthewmasters ! Looks like there's a bunch to get into here.",
    "2758706": "What do you mean it is prohibitive? It would take to much computational resources?",
    "2769231": "anyone what to use opensource docking tools to create a list of binding sites/pockets for the 3 target protein and share with the public?",
    "2771753": "i am getting 3d coordinates for input to 3d graph GNN (SMP spherical message passing).\nHere is how i use rdkit ETKDG algo\n(i later plan to search github for fast DFT approcimation with neural nets)\n\n```\n\nsmiles = [\n\t\"C#CCOc1ccc(CNc2nc(NCc3cccc(Br)n3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1\",\n]\n\ncoordinates=[]\ns = smiles[0]\nmol = Chem.MolFromSmiles(s)\n# add hydrogen bonds to molecule because they are not in the smiles representation\nmol = Chem.AddHs(mol)\n\ntry:\n\tps = AllChem.ETKDGv2()\n\tps.useRandomCoords = True\n\tAllChem.EmbedMolecule(mol, ps)\n\tAllChem.MMFFOptimizeMolecule(mol, confId=0)\n\tconf = mol.GetConformer()\n\tc = conf.GetPositions()\n\tcoordinates.append(c)\nexcept:\n\tprint('error',s )\n\n[19:47:30] UFFTYPER: Unrecognized charge state for atom: 33\n[19:47:30] UFFTYPER: Unrecognized atom type: Dy5+3 (33)\n```\n\nnot sure if the fake Dy will affect results ....",
    "2771782": "Thanks for sharing this here @hengck23 ! \n\nA couple notes that might be helpful.\n\n1) you may want to generate multiple conformers and check for the lowest energy one/ones.\n\n2) You are correct, the Dy is not parameterized by this force field, so it's unlikely the MMFF will be able to give you accurate energies with this group attached.",
    "2772336": "gyulamaloveczky4 Sorry, didn't see this before. Yes, too many resources for the average kaggler. DiffDock for example takes around 10 seconds per molecule. Just to dock the test set would take something like 200 days of GPU time. And then you just have poses and scores. You still have to develop a classifier on top of it, meaning you have to dock training molecules too. Nevertheless, I still think docking can be useful. For example, to find the bioactive conformation of a known binder or to confirm a predicted binder. There are also some clever ideas out there in the docking/virtual screening literature",
    "2773473": "https://www.eyesopen.com/omega\nOMEGA samples the conformational space of drug-like molecules at speeds of hundreds of thousands of compounds per day on a CPU, and over a million molecules per day on a GPU.\njust 1M per day!!!! (we have 100M)",
    "2773544": "For the task of binding site prediction, I would use P2Rank or DiffDock. Although DiffDock was designed to dock molecules, it's performance mostly comes from it's pocket finding ability. Paper: [Do Deep Learning Models Really Outperform\nTraditional Approaches in Molecular Docking?](https://arxiv.org/pdf/2302.07134)",
    "2773937": "High-Quality Conformer Generation with CONFORGE:Algorithm and Performance Assessment\nhttps://pubs.acs.org/doi/epdf/10.1021/acs.jcim.3c00563\n\ncompare open-source CONFORGE with commercial OMEGA and CORINA in the paper. seems good.",
    "2776071": "Hey everyone, if you're looking to explore different DNA attachment points (perhaps for building 3D models or whatever else you think it might be useful for), I wrote a notebook to help you get started with your experimentation: https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations.\n\nBest of luck!",
    "2776097": "thanks a lot for the code! ‎‎‎‎",
    "2776098": "You're very welcome!",
    "2776122": "how to find pockets\nhttps://github.com/yazdanimehdi/AttentionSiteDTI/blob/main/human_data.py\n\n```\npk = deepchem.dock.ConvexHullPocketFinder()\n\ndef process_protein(pdb_file):\n    m = Chem.MolFromPDBFile(pdb_file)\n    am = GetAdjacencyMatrix(m)\n    pockets = pk.find_pockets(pdb_file)\n    n2 = m.GetNumAtoms()\n\n```"
  },
  "source": "meta"
}