{
  "id": 494914,
  "title": "Common Ways to Represent Molecules in ML",
  "url": "/competitions/leash-BELKA/discussion/494914",
  "author_name": "",
  "post_date": "2024-04-18T22:02:47.156697300Z",
  "votes": 24,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Aiming to consolidate landmark methods of representing molecules for machine learning. Check out this <a href=\"https://www.kaggle.com/code/sdlee94/belka-molecule-representations-for-ml-tutorial/notebook?scriptVersionId=172758977l\" target=\"_blank\">Notebook</a> for a tutorial for how to generate each type of embedding featured below.</p>\n<h2><strong>Molecular Descriptors</strong></h2>\n<ul>\n<li>quantitative properties that capture information about the chemical structure, such as its constitution, topology, geometry, electron distribution, and hydrophobicity. Descriptors can be generated using the RDKit python package.</li>\n<li>References:<ul>\n<li><a href=\"https://www.rdkit.org/docs/GettingStartedInPython.html#descriptor-calculation\" target=\"_blank\">RDKit Descriptors Documentation</a></li>\n<li><a href=\"https://github.com/rdkit/rdkit\" target=\"_blank\">RDKit GitHub</a></li></ul></li>\n</ul>\n<h2><strong>Molecular Fingerprints</strong></h2>\n<ul>\n<li>fixed-length numerical representations that encode molecular structures, usually as a binary vector. Each element (called a bit) of a fingerprint vector represents the presence of specific atomic or structural features within the chemical compound.</li>\n<li>examples of molecular fingerprints include<ul>\n<li>MACCS Keys: a 166-bit fingerprint based on a predefined list of molecular substructures or patterns</li>\n<li>Morgan Fingerprints: a fixed-length bit fingerprint based on hashed topological features of atoms and their bond connectivities within a specified radius. Also referred to as Extended-Connectivity Fingerprints (ECFPs)</li></ul></li>\n<li>References:<ul>\n<li>Durant, J. L., Leland, B. A., Henry, D. R., &amp; Nourse, J. G. (2002). Reoptimization of MDL keys for use in drug discovery. Journal of Chemical Information and Computer Sciences, 42(6), 1273-1280. <a href=\"https://pubs.acs.org/doi/10.1021/ci010132r\" target=\"_blank\">Link to Article</a></li>\n<li>Rogers, D., &amp; Hahn, M. (2010). Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5), 742-754. <a href=\"https://pubs.acs.org/doi/10.1021/ci100050t\" target=\"_blank\">Link to Article</a></li></ul></li>\n</ul>\n<h2><strong>Molecular Graphs</strong></h2>\n<ul>\n<li>a data structure in which molecules are represented as a graph with interconnected nodes (typically nodes are atoms and edges are bonds). Molecular graphs can be generated using the Deep Graph Library python package.</li>\n<li>References::<ul>\n<li>Wang, M., Zheng, D., Ye, Z., Gan, Q., Li, M., Song, X., Zhou, J., Ma, C., Yu, L., Gai, Y., Xiao, T., He, T., Karypis, G., Li, J., &amp; Zhang, Z. (2019). Deep graph library: A graph-centric, highly-performant package for graph neural networks. <a href=\"https://arxiv.org/abs/1909.01315\" target=\"_blank\">Link to Article</a></li>\n<li><a href=\"https://docs.dgl.ai/\" target=\"_blank\">Deep Graph Libarary Docs</a></li></ul></li>\n</ul>\n<h2><strong>Mol2Vec Embeddings</strong></h2>\n<ul>\n<li>continuous vector representations of molecules generated by an algorithm based on Word2Vec. Mol2Vec embeddings captures chemical similarity in a manner similar to how Word2Vec captures semantic similarity between words</li>\n<li>References:<ul>\n<li>Jaeger, S., Fulle, S., &amp; Turk, S. (2018). Mol2vec: unsupervised machine learning approach with chemical intuition. Journal of Chemical Information and Modeling, 58(1), 27-35. <a href=\"https://pubs.acs.org/doi/10.1021/acs.jcim.7b00616\" target=\"_blank\">Link to Article</a></li>\n<li><a href=\"https://github.com/samoturk/mol2vec\" target=\"_blank\">Mol2Vec GitHub Repo</a></li></ul></li>\n</ul>\n<h2><strong>Chemical Language Model Embeddings</strong></h2>\n<ul>\n<li>learned representations of molecules generated by transformer-based architectures - the same kind of models that forms the basis for traditional language models (e.g. BERT). Self-attention mechanisms compute the representation each chemical element (e.g. atom) to every other element in a given molecule.</li>\n<li>examples of Chemical Language Models include ChemBERTa and MolFormer</li>\n<li>References:<ul>\n<li>Chithrananda, S., Grand, G., &amp; Ramsundar, B. (2020). ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction. <a href=\"https://arxiv.org/abs/2010.09885\" target=\"_blank\">Link to Article</a></li>\n<li>Chithrananda, S., Grand, G., &amp; Ramsundar, B. (2022). ChemBERTa-2: Towards Chemical Foundation Models. <a href=\"https://arxiv.org/abs/2209.01712\" target=\"_blank\">Link to Article</a></li>\n<li>Ross, J., Belgodere, B., Chenthamarakshan, V., et al. (2022). Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4, 1256-1264. <a href=\"https://www.nature.com/articles/s42256-022-00580-7\" target=\"_blank\">Link to Article</a></li>\n<li><a href=\"https://huggingface.co/seyonec/ChemBERTa-zinc-base-v1\" target=\"_blank\">ChemBERTa on HuggingFace Model Repo</a></li>\n<li><a href=\"https://huggingface.co/ibm/MoLFormer-XL-both-10pct\" target=\"_blank\">MolFormer on HuggingFace Model Repo</a></li>\n<li><a href=\"https://github.com/IBM/molformer\" target=\"_blank\">MolFormer GitHub repo</a></li></ul></li>\n</ul>\n<p>Feel like anything is missing? Please provide your feedback!</p>",
  "messages": [
    {
      "id": "2759774",
      "postDate": "04/18/2024 22:02:47",
      "content": "<p>Aiming to consolidate landmark methods of representing molecules for machine learning. Check out this <a href=\"https://www.kaggle.com/code/sdlee94/belka-molecule-representations-for-ml-tutorial/notebook?scriptVersionId=172758977l\" target=\"_blank\">Notebook</a> for a tutorial for how to generate each type of embedding featured below.</p>\n<h2><strong>Molecular Descriptors</strong></h2>\n<ul>\n<li>quantitative properties that capture information about the chemical structure, such as its constitution, topology, geometry, electron distribution, and hydrophobicity. Descriptors can be generated using the RDKit python package.</li>\n<li>References:<ul>\n<li><a href=\"https://www.rdkit.org/docs/GettingStartedInPython.html#descriptor-calculation\" target=\"_blank\">RDKit Descriptors Documentation</a></li>\n<li><a href=\"https://github.com/rdkit/rdkit\" target=\"_blank\">RDKit GitHub</a></li></ul></li>\n</ul>\n<h2><strong>Molecular Fingerprints</strong></h2>\n<ul>\n<li>fixed-length numerical representations that encode molecular structures, usually as a binary vector. Each element (called a bit) of a fingerprint vector represents the presence of specific atomic or structural features within the chemical compound.</li>\n<li>examples of molecular fingerprints include<ul>\n<li>MACCS Keys: a 166-bit fingerprint based on a predefined list of molecular substructures or patterns</li>\n<li>Morgan Fingerprints: a fixed-length bit fingerprint based on hashed topological features of atoms and their bond connectivities within a specified radius. Also referred to as Extended-Connectivity Fingerprints (ECFPs)</li></ul></li>\n<li>References:<ul>\n<li>Durant, J. L., Leland, B. A., Henry, D. R., &amp; Nourse, J. G. (2002). Reoptimization of MDL keys for use in drug discovery. Journal of Chemical Information and Computer Sciences, 42(6), 1273-1280. <a href=\"https://pubs.acs.org/doi/10.1021/ci010132r\" target=\"_blank\">Link to Article</a></li>\n<li>Rogers, D., &amp; Hahn, M. (2010). Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5), 742-754. <a href=\"https://pubs.acs.org/doi/10.1021/ci100050t\" target=\"_blank\">Link to Article</a></li></ul></li>\n</ul>\n<h2><strong>Molecular Graphs</strong></h2>\n<ul>\n<li>a data structure in which molecules are represented as a graph with interconnected nodes (typically nodes are atoms and edges are bonds). Molecular graphs can be generated using the Deep Graph Library python package.</li>\n<li>References::<ul>\n<li>Wang, M., Zheng, D., Ye, Z., Gan, Q., Li, M., Song, X., Zhou, J., Ma, C., Yu, L., Gai, Y., Xiao, T., He, T., Karypis, G., Li, J., &amp; Zhang, Z. (2019). Deep graph library: A graph-centric, highly-performant package for graph neural networks. <a href=\"https://arxiv.org/abs/1909.01315\" target=\"_blank\">Link to Article</a></li>\n<li><a href=\"https://docs.dgl.ai/\" target=\"_blank\">Deep Graph Libarary Docs</a></li></ul></li>\n</ul>\n<h2><strong>Mol2Vec Embeddings</strong></h2>\n<ul>\n<li>continuous vector representations of molecules generated by an algorithm based on Word2Vec. Mol2Vec embeddings captures chemical similarity in a manner similar to how Word2Vec captures semantic similarity between words</li>\n<li>References:<ul>\n<li>Jaeger, S., Fulle, S., &amp; Turk, S. (2018). Mol2vec: unsupervised machine learning approach with chemical intuition. Journal of Chemical Information and Modeling, 58(1), 27-35. <a href=\"https://pubs.acs.org/doi/10.1021/acs.jcim.7b00616\" target=\"_blank\">Link to Article</a></li>\n<li><a href=\"https://github.com/samoturk/mol2vec\" target=\"_blank\">Mol2Vec GitHub Repo</a></li></ul></li>\n</ul>\n<h2><strong>Chemical Language Model Embeddings</strong></h2>\n<ul>\n<li>learned representations of molecules generated by transformer-based architectures - the same kind of models that forms the basis for traditional language models (e.g. BERT). Self-attention mechanisms compute the representation each chemical element (e.g. atom) to every other element in a given molecule.</li>\n<li>examples of Chemical Language Models include ChemBERTa and MolFormer</li>\n<li>References:<ul>\n<li>Chithrananda, S., Grand, G., &amp; Ramsundar, B. (2020). ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction. <a href=\"https://arxiv.org/abs/2010.09885\" target=\"_blank\">Link to Article</a></li>\n<li>Chithrananda, S., Grand, G., &amp; Ramsundar, B. (2022). ChemBERTa-2: Towards Chemical Foundation Models. <a href=\"https://arxiv.org/abs/2209.01712\" target=\"_blank\">Link to Article</a></li>\n<li>Ross, J., Belgodere, B., Chenthamarakshan, V., et al. (2022). Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4, 1256-1264. <a href=\"https://www.nature.com/articles/s42256-022-00580-7\" target=\"_blank\">Link to Article</a></li>\n<li><a href=\"https://huggingface.co/seyonec/ChemBERTa-zinc-base-v1\" target=\"_blank\">ChemBERTa on HuggingFace Model Repo</a></li>\n<li><a href=\"https://huggingface.co/ibm/MoLFormer-XL-both-10pct\" target=\"_blank\">MolFormer on HuggingFace Model Repo</a></li>\n<li><a href=\"https://github.com/IBM/molformer\" target=\"_blank\">MolFormer GitHub repo</a></li></ul></li>\n</ul>\n<p>Feel like anything is missing? Please provide your feedback!</p>",
      "rawMarkdown": "Aiming to consolidate landmark methods of representing molecules for machine learning. Check out this [Notebook](https://www.kaggle.com/code/sdlee94/belka-molecule-representations-for-ml-tutorial/notebook?scriptVersionId=172758977l) for a tutorial for how to generate each type of embedding featured below.\n\n##**Molecular Descriptors**\n- quantitative properties that capture information about the chemical structure, such as its constitution, topology, geometry, electron distribution, and hydrophobicity. Descriptors can be generated using the RDKit python package.\n- References:\n   - [RDKit Descriptors Documentation](https://www.rdkit.org/docs/GettingStartedInPython.html#descriptor-calculation)\n   - [RDKit GitHub](https://github.com/rdkit/rdkit)\n\n##**Molecular Fingerprints**\n- fixed-length numerical representations that encode molecular structures, usually as a binary vector. Each element (called a bit) of a fingerprint vector represents the presence of specific atomic or structural features within the chemical compound.\n- examples of molecular fingerprints include\n   - MACCS Keys: a 166-bit fingerprint based on a predefined list of molecular substructures or patterns\n   - Morgan Fingerprints: a fixed-length bit fingerprint based on hashed topological features of atoms and their bond connectivities within a specified radius. Also referred to as Extended-Connectivity Fingerprints (ECFPs)\n- References:\n   - Durant, J. L., Leland, B. A., Henry, D. R., & Nourse, J. G. (2002). Reoptimization of MDL keys for use in drug discovery. Journal of Chemical Information and Computer Sciences, 42(6), 1273-1280. [Link to Article](https://pubs.acs.org/doi/10.1021/ci010132r)\n   - Rogers, D., & Hahn, M. (2010). Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5), 742-754. [Link to Article](https://pubs.acs.org/doi/10.1021/ci100050t)\n\n##**Molecular Graphs**\n- a data structure in which molecules are represented as a graph with interconnected nodes (typically nodes are atoms and edges are bonds). Molecular graphs can be generated using the Deep Graph Library python package.\n- References::\n   - Wang, M., Zheng, D., Ye, Z., Gan, Q., Li, M., Song, X., Zhou, J., Ma, C., Yu, L., Gai, Y., Xiao, T., He, T., Karypis, G., Li, J., & Zhang, Z. (2019). Deep graph library: A graph-centric, highly-performant package for graph neural networks. [Link to Article](https://arxiv.org/abs/1909.01315)\n   - [Deep Graph Libarary Docs](https://docs.dgl.ai/)\n\n##**Mol2Vec Embeddings**\n- continuous vector representations of molecules generated by an algorithm based on Word2Vec. Mol2Vec embeddings captures chemical similarity in a manner similar to how Word2Vec captures semantic similarity between words\n- References:\n   - Jaeger, S., Fulle, S., & Turk, S. (2018). Mol2vec: unsupervised machine learning approach with chemical intuition. Journal of Chemical Information and Modeling, 58(1), 27-35. [Link to Article](https://pubs.acs.org/doi/10.1021/acs.jcim.7b00616)\n   - [Mol2Vec GitHub Repo](https://github.com/samoturk/mol2vec)\n\n##**Chemical Language Model Embeddings**\n- learned representations of molecules generated by transformer-based architectures - the same kind of models that forms the basis for traditional language models (e.g. BERT). Self-attention mechanisms compute the representation each chemical element (e.g. atom) to every other element in a given molecule.\n- examples of Chemical Language Models include ChemBERTa and MolFormer\n- References:\n   - Chithrananda, S., Grand, G., & Ramsundar, B. (2020). ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction. [Link to Article](https://arxiv.org/abs/2010.09885)\n   - Chithrananda, S., Grand, G., & Ramsundar, B. (2022). ChemBERTa-2: Towards Chemical Foundation Models. [Link to Article](https://arxiv.org/abs/2209.01712)\n   - Ross, J., Belgodere, B., Chenthamarakshan, V., et al. (2022). Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4, 1256-1264. [Link to Article](https://www.nature.com/articles/s42256-022-00580-7)\n   - [ChemBERTa on HuggingFace Model Repo](https://huggingface.co/seyonec/ChemBERTa-zinc-base-v1)\n   - [MolFormer on HuggingFace Model Repo](https://huggingface.co/ibm/MoLFormer-XL-both-10pct)\n   - [MolFormer GitHub repo](https://github.com/IBM/molformer)\n\nFeel like anything is missing? Please provide your feedback!",
      "votes": null
    },
    {
      "id": "2760015",
      "postDate": "04/19/2024 04:10:18",
      "content": "<p>I'd be remiss to not at least mention Taco S. Cohen, Mario Geiger, Jonas Koehler, Max Welling (2018). Spherical CNNs. <a href=\"https://arxiv.org/abs/1801.10130\" target=\"_blank\">Link to article</a></p>",
      "rawMarkdown": "I'd be remiss to not at least mention Taco S. Cohen, Mario Geiger, Jonas Koehler, Max Welling (2018). Spherical CNNs. [Link to article](https://arxiv.org/abs/1801.10130)",
      "votes": null
    },
    {
      "id": "2760227",
      "postDate": "04/19/2024 06:57:52",
      "content": "<p>Thanks for the link, this was not on my radar at all.</p>",
      "rawMarkdown": "Thanks for the link, this was not on my radar at all.",
      "votes": null
    },
    {
      "id": "2761098",
      "postDate": "04/19/2024 16:58:18",
      "content": "<p>Interesting - I've not encountered this approach in chemical ML before, thanks for the information! Are there any example code available that demonstrates how to convert SMILES to spherical data?</p>",
      "rawMarkdown": "Interesting - I've not encountered this approach in chemical ML before, thanks for the information! Are there any example code available that demonstrates how to convert SMILES to spherical data?",
      "votes": null
    },
    {
      "id": "2761263",
      "postDate": "04/19/2024 18:40:08",
      "content": "<p>Maybe poke around <a href=\"https://github.com/jonkhler/s2cnn/tree/master/examples/molecules\" target=\"_blank\">https://github.com/jonkhler/s2cnn/tree/master/examples/molecules</a>?</p>",
      "rawMarkdown": "Maybe poke around https://github.com/jonkhler/s2cnn/tree/master/examples/molecules?",
      "votes": null
    },
    {
      "id": "2761508",
      "postDate": "04/19/2024 21:34:26",
      "content": "<p>Not really any example code related to that, it just directly parses data from <a href=\"http://quantum-machine.org/datasets/\" target=\"_blank\">QM7 matlab dataset</a></p>\n<p>You'd probably need to either see if you can read and understand the dataset and transform into similar data, or poke around the spherical GitHub enough to understand more generally how to represent 3d data for that model. </p>\n<p>Converting smiles to 3d data with RdKit is pretty trivial, that part shouldn't be hard (though doing it in a reasonable timeframe for millions of a rows might be)</p>",
      "rawMarkdown": "Not really any example code related to that, it just directly parses data from [QM7 matlab dataset](http://quantum-machine.org/datasets/)\n\nYou'd probably need to either see if you can read and understand the dataset and transform into similar data, or poke around the spherical GitHub enough to understand more generally how to represent 3d data for that model. \n\nConverting smiles to 3d data with RdKit is pretty trivial, that part shouldn't be hard (though doing it in a reasonable timeframe for millions of a rows might be)",
      "votes": null
    },
    {
      "id": "2769998",
      "postDate": "04/23/2024 16:21:19",
      "content": "<p>I just use pytorch geometrics from_smiles function, makes me a nice graph</p>",
      "rawMarkdown": "I just use pytorch geometrics from_smiles function, makes me a nice graph",
      "votes": null
    },
    {
      "id": "2782752",
      "postDate": "04/29/2024 12:46:36",
      "content": "<p>Thank you for  this information!</p>\n<p>How effective do you think GBDT and Neural Network prediction models that use physical property values such as RDkit's Molecular Descriptors and Mordred (Journal of Cheminformatics volume 10, Article number: 4 (2018)) are effective?<br>\nDo you think this could be a complement to ECFP or transformer-based architectures?</p>",
      "rawMarkdown": "Thank you for  this information!\n\nHow effective do you think GBDT and Neural Network prediction models that use physical property values such as RDkit's Molecular Descriptors and Mordred (Journal of Cheminformatics volume 10, Article number: 4 (2018)) are effective?\nDo you think this could be a complement to ECFP or transformer-based architectures?",
      "votes": null
    },
    {
      "id": "2782892",
      "postDate": "04/29/2024 13:46:08",
      "content": "<p>Which article in <a href=\"https://link.springer.com/journal/13321/volumes-and-issues/10-1\" target=\"_blank\">Journal of Cheminformatics v10</a> are you referring to?  I can find anything in that issued called \"Mordred\".</p>\n<p>Physical properties may be important for certain types of drug-protein interactions, specifically those that are less binding-pocket oriented.</p>",
      "rawMarkdown": "Which article in [Journal of Cheminformatics v10](https://link.springer.com/journal/13321/volumes-and-issues/10-1) are you referring to?  I can find anything in that issued called \"Mordred\".\n\nPhysical properties may be important for certain types of drug-protein interactions, specifically those that are less binding-pocket oriented.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2760015,
      "author_name": "ianquigley",
      "author_url": "",
      "post_date": "04/19/2024 04:10:18",
      "content": "<p>I'd be remiss to not at least mention Taco S. Cohen, Mario Geiger, Jonas Koehler, Max Welling (2018). Spherical CNNs. <a href=\"https://arxiv.org/abs/1801.10130\" target=\"_blank\">Link to article</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2760227,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "04/19/2024 06:57:52",
          "content": "<p>Thanks for the link, this was not on my radar at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2761098,
          "author_name": "sdlee94",
          "author_url": "",
          "post_date": "04/19/2024 16:58:18",
          "content": "<p>Interesting - I've not encountered this approach in chemical ML before, thanks for the information! Are there any example code available that demonstrates how to convert SMILES to spherical data?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2761263,
              "author_name": "ianquigley",
              "author_url": "",
              "post_date": "04/19/2024 18:40:08",
              "content": "<p>Maybe poke around <a href=\"https://github.com/jonkhler/s2cnn/tree/master/examples/molecules\" target=\"_blank\">https://github.com/jonkhler/s2cnn/tree/master/examples/molecules</a>?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2761508,
                  "author_name": "roberthatch",
                  "author_url": "",
                  "post_date": "04/19/2024 21:34:26",
                  "content": "<p>Not really any example code related to that, it just directly parses data from <a href=\"http://quantum-machine.org/datasets/\" target=\"_blank\">QM7 matlab dataset</a></p>\n<p>You'd probably need to either see if you can read and understand the dataset and transform into similar data, or poke around the spherical GitHub enough to understand more generally how to represent 3d data for that model. </p>\n<p>Converting smiles to 3d data with RdKit is pretty trivial, that part shouldn't be hard (though doing it in a reasonable timeframe for millions of a rows might be)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2769998,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "04/23/2024 16:21:19",
      "content": "<p>I just use pytorch geometrics from_smiles function, makes me a nice graph</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2782752,
      "author_name": "higayuki",
      "author_url": "",
      "post_date": "04/29/2024 12:46:36",
      "content": "<p>Thank you for  this information!</p>\n<p>How effective do you think GBDT and Neural Network prediction models that use physical property values such as RDkit's Molecular Descriptors and Mordred (Journal of Cheminformatics volume 10, Article number: 4 (2018)) are effective?<br>\nDo you think this could be a complement to ECFP or transformer-based architectures?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2782892,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "04/29/2024 13:46:08",
          "content": "<p>Which article in <a href=\"https://link.springer.com/journal/13321/volumes-and-issues/10-1\" target=\"_blank\">Journal of Cheminformatics v10</a> are you referring to?  I can find anything in that issued called \"Mordred\".</p>\n<p>Physical properties may be important for certain types of drug-protein interactions, specifically those that are less binding-pocket oriented.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2759774": "Aiming to consolidate landmark methods of representing molecules for machine learning. Check out this [Notebook](https://www.kaggle.com/code/sdlee94/belka-molecule-representations-for-ml-tutorial/notebook?scriptVersionId=172758977l) for a tutorial for how to generate each type of embedding featured below.\n\n##**Molecular Descriptors**\n- quantitative properties that capture information about the chemical structure, such as its constitution, topology, geometry, electron distribution, and hydrophobicity. Descriptors can be generated using the RDKit python package.\n- References:\n   - [RDKit Descriptors Documentation](https://www.rdkit.org/docs/GettingStartedInPython.html#descriptor-calculation)\n   - [RDKit GitHub](https://github.com/rdkit/rdkit)\n\n##**Molecular Fingerprints**\n- fixed-length numerical representations that encode molecular structures, usually as a binary vector. Each element (called a bit) of a fingerprint vector represents the presence of specific atomic or structural features within the chemical compound.\n- examples of molecular fingerprints include\n   - MACCS Keys: a 166-bit fingerprint based on a predefined list of molecular substructures or patterns\n   - Morgan Fingerprints: a fixed-length bit fingerprint based on hashed topological features of atoms and their bond connectivities within a specified radius. Also referred to as Extended-Connectivity Fingerprints (ECFPs)\n- References:\n   - Durant, J. L., Leland, B. A., Henry, D. R., & Nourse, J. G. (2002). Reoptimization of MDL keys for use in drug discovery. Journal of Chemical Information and Computer Sciences, 42(6), 1273-1280. [Link to Article](https://pubs.acs.org/doi/10.1021/ci010132r)\n   - Rogers, D., & Hahn, M. (2010). Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5), 742-754. [Link to Article](https://pubs.acs.org/doi/10.1021/ci100050t)\n\n##**Molecular Graphs**\n- a data structure in which molecules are represented as a graph with interconnected nodes (typically nodes are atoms and edges are bonds). Molecular graphs can be generated using the Deep Graph Library python package.\n- References::\n   - Wang, M., Zheng, D., Ye, Z., Gan, Q., Li, M., Song, X., Zhou, J., Ma, C., Yu, L., Gai, Y., Xiao, T., He, T., Karypis, G., Li, J., & Zhang, Z. (2019). Deep graph library: A graph-centric, highly-performant package for graph neural networks. [Link to Article](https://arxiv.org/abs/1909.01315)\n   - [Deep Graph Libarary Docs](https://docs.dgl.ai/)\n\n##**Mol2Vec Embeddings**\n- continuous vector representations of molecules generated by an algorithm based on Word2Vec. Mol2Vec embeddings captures chemical similarity in a manner similar to how Word2Vec captures semantic similarity between words\n- References:\n   - Jaeger, S., Fulle, S., & Turk, S. (2018). Mol2vec: unsupervised machine learning approach with chemical intuition. Journal of Chemical Information and Modeling, 58(1), 27-35. [Link to Article](https://pubs.acs.org/doi/10.1021/acs.jcim.7b00616)\n   - [Mol2Vec GitHub Repo](https://github.com/samoturk/mol2vec)\n\n##**Chemical Language Model Embeddings**\n- learned representations of molecules generated by transformer-based architectures - the same kind of models that forms the basis for traditional language models (e.g. BERT). Self-attention mechanisms compute the representation each chemical element (e.g. atom) to every other element in a given molecule.\n- examples of Chemical Language Models include ChemBERTa and MolFormer\n- References:\n   - Chithrananda, S., Grand, G., & Ramsundar, B. (2020). ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction. [Link to Article](https://arxiv.org/abs/2010.09885)\n   - Chithrananda, S., Grand, G., & Ramsundar, B. (2022). ChemBERTa-2: Towards Chemical Foundation Models. [Link to Article](https://arxiv.org/abs/2209.01712)\n   - Ross, J., Belgodere, B., Chenthamarakshan, V., et al. (2022). Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4, 1256-1264. [Link to Article](https://www.nature.com/articles/s42256-022-00580-7)\n   - [ChemBERTa on HuggingFace Model Repo](https://huggingface.co/seyonec/ChemBERTa-zinc-base-v1)\n   - [MolFormer on HuggingFace Model Repo](https://huggingface.co/ibm/MoLFormer-XL-both-10pct)\n   - [MolFormer GitHub repo](https://github.com/IBM/molformer)\n\nFeel like anything is missing? Please provide your feedback!",
    "2760015": "I'd be remiss to not at least mention Taco S. Cohen, Mario Geiger, Jonas Koehler, Max Welling (2018). Spherical CNNs. [Link to article](https://arxiv.org/abs/1801.10130)",
    "2760227": "Thanks for the link, this was not on my radar at all.",
    "2761098": "Interesting - I've not encountered this approach in chemical ML before, thanks for the information! Are there any example code available that demonstrates how to convert SMILES to spherical data?",
    "2761263": "Maybe poke around https://github.com/jonkhler/s2cnn/tree/master/examples/molecules?",
    "2761508": "Not really any example code related to that, it just directly parses data from [QM7 matlab dataset](http://quantum-machine.org/datasets/)\n\nYou'd probably need to either see if you can read and understand the dataset and transform into similar data, or poke around the spherical GitHub enough to understand more generally how to represent 3d data for that model. \n\nConverting smiles to 3d data with RdKit is pretty trivial, that part shouldn't be hard (though doing it in a reasonable timeframe for millions of a rows might be)",
    "2769998": "I just use pytorch geometrics from_smiles function, makes me a nice graph",
    "2782752": "Thank you for  this information!\n\nHow effective do you think GBDT and Neural Network prediction models that use physical property values such as RDkit's Molecular Descriptors and Mordred (Journal of Cheminformatics volume 10, Article number: 4 (2018)) are effective?\nDo you think this could be a complement to ECFP or transformer-based architectures?",
    "2782892": "Which article in [Journal of Cheminformatics v10](https://link.springer.com/journal/13321/volumes-and-issues/10-1) are you referring to?  I can find anything in that issued called \"Mordred\".\n\nPhysical properties may be important for certain types of drug-protein interactions, specifically those that are less binding-pocket oriented."
  },
  "source": "meta"
}