{
  "id": 496601,
  "title": "SMILES description of proteins",
  "url": "/competitions/leash-BELKA/discussion/496601",
  "author_name": "",
  "post_date": "2024-04-21T19:40:35.183030Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Are there API or packages available to provide SMILEs representation of proteins from their names as provided in this dataset (e.g. BRD4)?</p>",
  "messages": [
    {
      "id": "2766609",
      "postDate": "04/21/2024 19:40:35",
      "content": "<p>Are there API or packages available to provide SMILEs representation of proteins from their names as provided in this dataset (e.g. BRD4)?</p>",
      "rawMarkdown": "Are there API or packages available to provide SMILEs representation of proteins from their names as provided in this dataset (e.g. BRD4)?",
      "votes": null
    },
    {
      "id": "2766724",
      "postDate": "04/21/2024 21:04:14",
      "content": "<p>I haven't done it, but a partially manual process, if you:</p>\n<ul>\n<li>Take the Alphafold predicted structure linked from the Data page</li>\n<li>Manually edit the pdb text file to only ATOM entries within the \"positions [X] to [Y]\" from the same Data page description. Example below.</li>\n<li>Load the pdb in rdkit. Not sure the command.</li>\n<li>Use MolToSmiles() on the resulting molecule.</li>\n</ul>\n<p>I think that will work… it will be very long, proteins are generally represented more compactly by their amino acid (residue) text string.</p>\n<p>Example of ATOM entries in PDB file:</p>\n<pre><code>ATOM      N   ASN A           - -              N  \nATOM      CA  ASN A           - -              C  \nATOM      C   ASN A           - -              C  \nATOM      CB  ASN A           - -              C  \n</code></pre>\n<p>So to read the above, the atom number is the first number (304-307), but the number we care about is the 44 after (in this case) ASN A. That tells us this is the 44th position, and (for BRD4) we want positions 44-460, so simply delete the atom entries before and after (and change the TER line at end of file) and I think? it will load with a pdb reader without any issues.</p>",
      "rawMarkdown": "I haven't done it, but a partially manual process, if you:\n* Take the Alphafold predicted structure linked from the Data page\n* Manually edit the pdb text file to only ATOM entries within the \"positions [X] to [Y]\" from the same Data page description. Example below.\n* Load the pdb in rdkit. Not sure the command.\n* Use MolToSmiles() on the resulting molecule.\n\nI think that will work... it will be very long, proteins are generally represented more compactly by their amino acid (residue) text string.\n\nExample of ATOM entries in PDB file:\n```python\nATOM    304  N   ASN A  44       3.584  -4.695 -28.738  1.00 82.83           N  \nATOM    305  CA  ASN A  44       2.914  -5.294 -27.583  1.00 82.83           C  \nATOM    306  C   ASN A  44       1.557  -5.899 -28.021  1.00 82.83           C  \nATOM    307  CB  ASN A  44       2.721  -4.229 -26.485  1.00 82.83           C  \n```\n\nSo to read the above, the atom number is the first number (304-307), but the number we care about is the 44 after (in this case) ASN A. That tells us this is the 44th position, and (for BRD4) we want positions 44-460, so simply delete the atom entries before and after (and change the TER line at end of file) and I think? it will load with a pdb reader without any issues.",
      "votes": null
    },
    {
      "id": "2766793",
      "postDate": "04/21/2024 22:06:04",
      "content": "<p>check this paper and their code:<br>\n<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6477977/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6477977/</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F34ed66baa76ca2426012f488c78000d6%2Fbty757f1(1).jpg?generation=1713737161225823&amp;alt=media\"></p>\n<p><a href=\"https://github.com/oddt/oddt/blob/master/oddt/fingerprints.py\" target=\"_blank\">https://github.com/oddt/oddt/blob/master/oddt/fingerprints.py</a></p>\n<pre><code>        ligand_ecfp = _ECFP_atom_hash(ligand,\n                                      ligand_atom,\n                                      =depth_ligand,\n                                      =lig_atom_repr)\n\n\n        protein_ecfp = _ECFP_atom_hash(protein,\n                                       protein_atom,\n                                       =depth_protein,\n                                       =prot_atom_repr)\n</code></pre>",
      "rawMarkdown": "check this paper and their code:\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6477977/\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F34ed66baa76ca2426012f488c78000d6%2Fbty757f1(1).jpg?generation=1713737161225823&alt=media)\n\nhttps://github.com/oddt/oddt/blob/master/oddt/fingerprints.py\n```\n        ligand_ecfp = _ECFP_atom_hash(ligand,\n                                      ligand_atom,\n                                      depth=depth_ligand,\n                                      atom_repr_dict=lig_atom_repr)\n\n\n        protein_ecfp = _ECFP_atom_hash(protein,\n                                       protein_atom,\n                                       depth=depth_protein,\n                                       atom_repr_dict=prot_atom_repr)\n\n```",
      "votes": null
    },
    {
      "id": "2766797",
      "postDate": "04/21/2024 22:07:32",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> - Thank you kindly for the this information. This process appears a bit difficult to incorporate into our model, at least for me. I am looking to get the chemical structure for the protein, similar to what you get from <code>rdkit.Chem.Descriptors</code>, however this API works with SMILES description, which we do not have for the protein.</p>\n<p>Do you know of any other API or package that can provide the chemical description for the protein?</p>",
      "rawMarkdown": "roberthatch - Thank you kindly for the this information. This process appears a bit difficult to incorporate into our model, at least for me. I am looking to get the chemical structure for the protein, similar to what you get from `rdkit.Chem.Descriptors`, however this API works with SMILES description, which we do not have for the protein.\n\nDo you know of any other API or package that can provide the chemical description for the protein?",
      "votes": null
    },
    {
      "id": "2766799",
      "postDate": "04/21/2024 22:09:20",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - Thank you. This is a very informative article with lots of great suggestions and resources. Excellent find !!</p>",
      "rawMarkdown": "hengck23 - Thank you. This is a very informative article with lots of great suggestions and resources. Excellent find !!",
      "votes": null
    },
    {
      "id": "2766863",
      "postDate": "04/22/2024 00:50:06",
      "content": "<p>There should be multiple options for finding something that can read a PDB file. Like I thought I had seen, rdkit has something:</p>\n<blockquote>\n  <p><a href=\"https://www.rdkit.org/docs/source/rdkit.Chem.rdmolfiles.html\" target=\"_blank\">rdkit.Chem.rdmolfiles.MolFromPDBFile</a></p>\n</blockquote>\n<p>Then you can do all the rdkit things, Chem.Descriptors or anything else.</p>\n<p>Getting the PDB is easy, it's linked right from the Data page. You don't technically have to do anything manual with it, but no public PDB and no program will give you BRD4 that's <em>exactly</em> \"the amino acid sequence is positions 44-460\". If you want just 44-460, you will need to take one of the above inputs and strip the first 43 and all after 460. But either of the linked PDBs is <em>a</em> version of the protein, so you'll have something even without the manual edit step.</p>",
      "rawMarkdown": "There should be multiple options for finding something that can read a PDB file. Like I thought I had seen, rdkit has something:\n\n> [rdkit.Chem.rdmolfiles.MolFromPDBFile](https://www.rdkit.org/docs/source/rdkit.Chem.rdmolfiles.html)\n\nThen you can do all the rdkit things, Chem.Descriptors or anything else.\n\nGetting the PDB is easy, it's linked right from the Data page. You don't technically have to do anything manual with it, but no public PDB and no program will give you BRD4 that's *exactly* \"the amino acid sequence is positions 44-460\". If you want just 44-460, you will need to take one of the above inputs and strip the first 43 and all after 460. But either of the linked PDBs is *a* version of the protein, so you'll have something even without the manual edit step.",
      "votes": null
    },
    {
      "id": "2767638",
      "postDate": "04/22/2024 13:02:14",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> - Thank you again for the clarification and the guidance.</p>",
      "rawMarkdown": "roberthatch - Thank you again for the clarification and the guidance.",
      "votes": null
    },
    {
      "id": "2767795",
      "postDate": "04/22/2024 14:24:04",
      "content": "<p>not sure if these are helpful at all<br>\n<a href=\"https://github.com/yazdanimehdi/DeepDrugDomain/blob/b9930adb5a1e1bb2397134c8d40df60636a23638/deepdrugdomain/data/preprocessing/protein/protein_fingerprint.py#L36\" target=\"_blank\">https://github.com/yazdanimehdi/DeepDrugDomain/blob/b9930adb5a1e1bb2397134c8d40df60636a23638/deepdrugdomain/data/preprocessing/protein/protein_fingerprint.py#L36</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F39a8ce902a6ea0f76973d501dc52480b%2FSelection_034.png?generation=1713795842409606&amp;alt=media\"></p>\n<p>see also:<a href=\"https://pybiomed.readthedocs.io/en/latest/application.html#application-3-prediction-of-drugtarget-interaction-from-the-integration-of-chemical-and-protein-spaces\" target=\"_blank\">https://pybiomed.readthedocs.io/en/latest/application.html#application-3-prediction-of-drugtarget-interaction-from-the-integration-of-chemical-and-protein-spaces</a></p>",
      "rawMarkdown": "not sure if these are helpful at all\nhttps://github.com/yazdanimehdi/DeepDrugDomain/blob/b9930adb5a1e1bb2397134c8d40df60636a23638/deepdrugdomain/data/preprocessing/protein/protein_fingerprint.py#L36\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F39a8ce902a6ea0f76973d501dc52480b%2FSelection_034.png?generation=1713795842409606&alt=media)\n\nsee also:https://pybiomed.readthedocs.io/en/latest/application.html#application-3-prediction-of-drugtarget-interaction-from-the-integration-of-chemical-and-protein-spaces",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2766724,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/21/2024 21:04:14",
      "content": "<p>I haven't done it, but a partially manual process, if you:</p>\n<ul>\n<li>Take the Alphafold predicted structure linked from the Data page</li>\n<li>Manually edit the pdb text file to only ATOM entries within the \"positions [X] to [Y]\" from the same Data page description. Example below.</li>\n<li>Load the pdb in rdkit. Not sure the command.</li>\n<li>Use MolToSmiles() on the resulting molecule.</li>\n</ul>\n<p>I think that will work… it will be very long, proteins are generally represented more compactly by their amino acid (residue) text string.</p>\n<p>Example of ATOM entries in PDB file:</p>\n<pre><code>ATOM      N   ASN A           - -              N  \nATOM      CA  ASN A           - -              C  \nATOM      C   ASN A           - -              C  \nATOM      CB  ASN A           - -              C  \n</code></pre>\n<p>So to read the above, the atom number is the first number (304-307), but the number we care about is the 44 after (in this case) ASN A. That tells us this is the 44th position, and (for BRD4) we want positions 44-460, so simply delete the atom entries before and after (and change the TER line at end of file) and I think? it will load with a pdb reader without any issues.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2766793,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/21/2024 22:06:04",
          "content": "<p>check this paper and their code:<br>\n<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6477977/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6477977/</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F34ed66baa76ca2426012f488c78000d6%2Fbty757f1(1).jpg?generation=1713737161225823&amp;alt=media\"></p>\n<p><a href=\"https://github.com/oddt/oddt/blob/master/oddt/fingerprints.py\" target=\"_blank\">https://github.com/oddt/oddt/blob/master/oddt/fingerprints.py</a></p>\n<pre><code>        ligand_ecfp = _ECFP_atom_hash(ligand,\n                                      ligand_atom,\n                                      =depth_ligand,\n                                      =lig_atom_repr)\n\n\n        protein_ecfp = _ECFP_atom_hash(protein,\n                                       protein_atom,\n                                       =depth_protein,\n                                       =prot_atom_repr)\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2766799,
              "author_name": "nadereafshar",
              "author_url": "",
              "post_date": "04/21/2024 22:09:20",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - Thank you. This is a very informative article with lots of great suggestions and resources. Excellent find !!</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2766797,
          "author_name": "nadereafshar",
          "author_url": "",
          "post_date": "04/21/2024 22:07:32",
          "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> - Thank you kindly for the this information. This process appears a bit difficult to incorporate into our model, at least for me. I am looking to get the chemical structure for the protein, similar to what you get from <code>rdkit.Chem.Descriptors</code>, however this API works with SMILES description, which we do not have for the protein.</p>\n<p>Do you know of any other API or package that can provide the chemical description for the protein?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2766863,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "04/22/2024 00:50:06",
              "content": "<p>There should be multiple options for finding something that can read a PDB file. Like I thought I had seen, rdkit has something:</p>\n<blockquote>\n  <p><a href=\"https://www.rdkit.org/docs/source/rdkit.Chem.rdmolfiles.html\" target=\"_blank\">rdkit.Chem.rdmolfiles.MolFromPDBFile</a></p>\n</blockquote>\n<p>Then you can do all the rdkit things, Chem.Descriptors or anything else.</p>\n<p>Getting the PDB is easy, it's linked right from the Data page. You don't technically have to do anything manual with it, but no public PDB and no program will give you BRD4 that's <em>exactly</em> \"the amino acid sequence is positions 44-460\". If you want just 44-460, you will need to take one of the above inputs and strip the first 43 and all after 460. But either of the linked PDBs is <em>a</em> version of the protein, so you'll have something even without the manual edit step.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2767638,
                  "author_name": "nadereafshar",
                  "author_url": "",
                  "post_date": "04/22/2024 13:02:14",
                  "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> - Thank you again for the clarification and the guidance.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2767795,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/22/2024 14:24:04",
      "content": "<p>not sure if these are helpful at all<br>\n<a href=\"https://github.com/yazdanimehdi/DeepDrugDomain/blob/b9930adb5a1e1bb2397134c8d40df60636a23638/deepdrugdomain/data/preprocessing/protein/protein_fingerprint.py#L36\" target=\"_blank\">https://github.com/yazdanimehdi/DeepDrugDomain/blob/b9930adb5a1e1bb2397134c8d40df60636a23638/deepdrugdomain/data/preprocessing/protein/protein_fingerprint.py#L36</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F39a8ce902a6ea0f76973d501dc52480b%2FSelection_034.png?generation=1713795842409606&amp;alt=media\"></p>\n<p>see also:<a href=\"https://pybiomed.readthedocs.io/en/latest/application.html#application-3-prediction-of-drugtarget-interaction-from-the-integration-of-chemical-and-protein-spaces\" target=\"_blank\">https://pybiomed.readthedocs.io/en/latest/application.html#application-3-prediction-of-drugtarget-interaction-from-the-integration-of-chemical-and-protein-spaces</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2766609": "Are there API or packages available to provide SMILEs representation of proteins from their names as provided in this dataset (e.g. BRD4)?",
    "2766724": "I haven't done it, but a partially manual process, if you:\n* Take the Alphafold predicted structure linked from the Data page\n* Manually edit the pdb text file to only ATOM entries within the \"positions [X] to [Y]\" from the same Data page description. Example below.\n* Load the pdb in rdkit. Not sure the command.\n* Use MolToSmiles() on the resulting molecule.\n\nI think that will work... it will be very long, proteins are generally represented more compactly by their amino acid (residue) text string.\n\nExample of ATOM entries in PDB file:\n```python\nATOM    304  N   ASN A  44       3.584  -4.695 -28.738  1.00 82.83           N  \nATOM    305  CA  ASN A  44       2.914  -5.294 -27.583  1.00 82.83           C  \nATOM    306  C   ASN A  44       1.557  -5.899 -28.021  1.00 82.83           C  \nATOM    307  CB  ASN A  44       2.721  -4.229 -26.485  1.00 82.83           C  \n```\n\nSo to read the above, the atom number is the first number (304-307), but the number we care about is the 44 after (in this case) ASN A. That tells us this is the 44th position, and (for BRD4) we want positions 44-460, so simply delete the atom entries before and after (and change the TER line at end of file) and I think? it will load with a pdb reader without any issues.",
    "2766793": "check this paper and their code:\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6477977/\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F34ed66baa76ca2426012f488c78000d6%2Fbty757f1(1).jpg?generation=1713737161225823&alt=media)\n\nhttps://github.com/oddt/oddt/blob/master/oddt/fingerprints.py\n```\n        ligand_ecfp = _ECFP_atom_hash(ligand,\n                                      ligand_atom,\n                                      depth=depth_ligand,\n                                      atom_repr_dict=lig_atom_repr)\n\n\n        protein_ecfp = _ECFP_atom_hash(protein,\n                                       protein_atom,\n                                       depth=depth_protein,\n                                       atom_repr_dict=prot_atom_repr)\n\n```",
    "2766797": "roberthatch - Thank you kindly for the this information. This process appears a bit difficult to incorporate into our model, at least for me. I am looking to get the chemical structure for the protein, similar to what you get from `rdkit.Chem.Descriptors`, however this API works with SMILES description, which we do not have for the protein.\n\nDo you know of any other API or package that can provide the chemical description for the protein?",
    "2766799": "hengck23 - Thank you. This is a very informative article with lots of great suggestions and resources. Excellent find !!",
    "2766863": "There should be multiple options for finding something that can read a PDB file. Like I thought I had seen, rdkit has something:\n\n> [rdkit.Chem.rdmolfiles.MolFromPDBFile](https://www.rdkit.org/docs/source/rdkit.Chem.rdmolfiles.html)\n\nThen you can do all the rdkit things, Chem.Descriptors or anything else.\n\nGetting the PDB is easy, it's linked right from the Data page. You don't technically have to do anything manual with it, but no public PDB and no program will give you BRD4 that's *exactly* \"the amino acid sequence is positions 44-460\". If you want just 44-460, you will need to take one of the above inputs and strip the first 43 and all after 460. But either of the linked PDBs is *a* version of the protein, so you'll have something even without the manual edit step.",
    "2767638": "roberthatch - Thank you again for the clarification and the guidance.",
    "2767795": "not sure if these are helpful at all\nhttps://github.com/yazdanimehdi/DeepDrugDomain/blob/b9930adb5a1e1bb2397134c8d40df60636a23638/deepdrugdomain/data/preprocessing/protein/protein_fingerprint.py#L36\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F39a8ce902a6ea0f76973d501dc52480b%2FSelection_034.png?generation=1713795842409606&alt=media)\n\nsee also:https://pybiomed.readthedocs.io/en/latest/application.html#application-3-prediction-of-drugtarget-interaction-from-the-integration-of-chemical-and-protein-spaces"
  },
  "source": "meta"
}