{
  "id": 498385,
  "title": "A brief note on using GNNs",
  "url": "/competitions/leash-BELKA/discussion/498385",
  "author_name": "Kyle Sørensen",
  "post_date": "2024-04-28T04:55:58.276000",
  "votes": 15,
  "comment_count": 11,
  "views": 0,
  "content": "<pre><code> networkx  nx\n torch\n rdkit  Chem\n torch_geometric.utils.convert  from_networkx\n\n :\n     ():\n        \n\n     ():\n        molecule = Chem.MolFromSmiles(smiles)\n         molecule  :\n             ValueError()\n\n        G = nx.Graph()\n         atom  molecule.GetAtoms():\n            G.add_node(atom.GetIdx(),\n                       x=torch.tensor([\n                           atom.GetAtomicNum(),\n                           atom.GetFormalCharge(),\n                           atom.GetChiralTag().real,\n                           atom.GetHybridization().real,\n                           atom.GetNumExplicitHs(),\n                           atom.GetNumImplicitHs()],\n                           dtype=torch.))\n\n        bond_type_to_int = {Chem.rdchem.BondType.SINGLE: ,\n                            Chem.rdchem.BondType.DOUBLE: ,\n                            Chem.rdchem.BondType.TRIPLE: ,\n                            Chem.rdchem.BondType.AROMATIC: }\n         bond  molecule.GetBonds():\n            G.add_edge(bond.GetBeginAtomIdx(), bond.GetEndAtomIdx(),\n                       edge_attr=torch.tensor([\n                           bond_type_to_int[bond.GetBondType()]], dtype=torch.))\n\n        data = from_networkx(G, group_node_attrs=[], group_edge_attrs=[])\n         data\n\n     ():\n         self.smiles_to_graph(smiles)\n\n     ():\n         smiles  smiles_list:\n            :\n                data = self.smiles_to_gnn_input(smiles)\n                total_size = \n                 _, tensor  data:\n                    total_size += tensor.element_size() * tensor.nelement()\n                ()\n             ValueError  e:\n                ()\n\nsmiles_examples = [, , ]\nconverter = SMILESToGraph()\nconverter.print_graph_sizes(smiles_examples)\n</code></pre>\n<p><code>Memory size for C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1: 2744 bytes</code></p>\n<pre><code> torch\n torch_geometric.nn  TransformerConv, global_mean_pool\n\n\n (torch.nn.Module):\n     ():\n        (BindingAffinityPredictor, self).__init__()\n        self.conv1 = TransformerConv(node_features, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.conv2 = TransformerConv(hidden_dim * num_heads, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.out = torch.nn.Linear(hidden_dim * num_heads, )  \n\n     ():\n        x, edge_index, edge_attr, batch = data.x, data.edge_index, data.edge_attr, data.batch\n        x = self.conv1(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = self.conv2(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = global_mean_pool(x, batch)  \n        x = self.out(x)\n        x = torch.sigmoid(x)  \n         x\n</code></pre>\n<p>This is a warning for anyone who might go down this path. I attempted to take the smiles representations and and use rdkit to convert them to graphs and feed them into a binary class gnn with some fancy layers. I was able to get the model to converge, and from what I can tell this method is a state of the art way of handling molecules with GNNs. However, I hope you have a lot of $, because this thing is a beast. As you can see, if you convert the smiles to even a simple graph representation, a single molecule in memory is 2.7kb. There are 98 million samples for one 1/3 proteins. You will find that even simply converting the smiles to graphs will take months of processing. If you attempt to process them all before hand, which is what I attempted as well, you will have terabytes of graph data, which of course will not fit in memory and which will become just as difficult and slow to read from disk as in memory. I have since given up on trying to implement a GNN for this, you have to down sample so much to even get a model that will finish an epoch. If anyone has any suggestions, please let me know. </p>",
  "messages": [
    {
      "id": 2780170,
      "postDate": "2024-04-28T04:55:58.277Z",
      "content": "<pre><code> networkx  nx\n torch\n rdkit  Chem\n torch_geometric.utils.convert  from_networkx\n\n :\n     ():\n        \n\n     ():\n        molecule = Chem.MolFromSmiles(smiles)\n         molecule  :\n             ValueError()\n\n        G = nx.Graph()\n         atom  molecule.GetAtoms():\n            G.add_node(atom.GetIdx(),\n                       x=torch.tensor([\n                           atom.GetAtomicNum(),\n                           atom.GetFormalCharge(),\n                           atom.GetChiralTag().real,\n                           atom.GetHybridization().real,\n                           atom.GetNumExplicitHs(),\n                           atom.GetNumImplicitHs()],\n                           dtype=torch.))\n\n        bond_type_to_int = {Chem.rdchem.BondType.SINGLE: ,\n                            Chem.rdchem.BondType.DOUBLE: ,\n                            Chem.rdchem.BondType.TRIPLE: ,\n                            Chem.rdchem.BondType.AROMATIC: }\n         bond  molecule.GetBonds():\n            G.add_edge(bond.GetBeginAtomIdx(), bond.GetEndAtomIdx(),\n                       edge_attr=torch.tensor([\n                           bond_type_to_int[bond.GetBondType()]], dtype=torch.))\n\n        data = from_networkx(G, group_node_attrs=[], group_edge_attrs=[])\n         data\n\n     ():\n         self.smiles_to_graph(smiles)\n\n     ():\n         smiles  smiles_list:\n            :\n                data = self.smiles_to_gnn_input(smiles)\n                total_size = \n                 _, tensor  data:\n                    total_size += tensor.element_size() * tensor.nelement()\n                ()\n             ValueError  e:\n                ()\n\nsmiles_examples = [, , ]\nconverter = SMILESToGraph()\nconverter.print_graph_sizes(smiles_examples)\n</code></pre>\n<p><code>Memory size for C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1: 2744 bytes</code></p>\n<pre><code> torch\n torch_geometric.nn  TransformerConv, global_mean_pool\n\n\n (torch.nn.Module):\n     ():\n        (BindingAffinityPredictor, self).__init__()\n        self.conv1 = TransformerConv(node_features, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.conv2 = TransformerConv(hidden_dim * num_heads, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.out = torch.nn.Linear(hidden_dim * num_heads, )  \n\n     ():\n        x, edge_index, edge_attr, batch = data.x, data.edge_index, data.edge_attr, data.batch\n        x = self.conv1(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = self.conv2(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = global_mean_pool(x, batch)  \n        x = self.out(x)\n        x = torch.sigmoid(x)  \n         x\n</code></pre>\n<p>This is a warning for anyone who might go down this path. I attempted to take the smiles representations and and use rdkit to convert them to graphs and feed them into a binary class gnn with some fancy layers. I was able to get the model to converge, and from what I can tell this method is a state of the art way of handling molecules with GNNs. However, I hope you have a lot of $, because this thing is a beast. As you can see, if you convert the smiles to even a simple graph representation, a single molecule in memory is 2.7kb. There are 98 million samples for one 1/3 proteins. You will find that even simply converting the smiles to graphs will take months of processing. If you attempt to process them all before hand, which is what I attempted as well, you will have terabytes of graph data, which of course will not fit in memory and which will become just as difficult and slow to read from disk as in memory. I have since given up on trying to implement a GNN for this, you have to down sample so much to even get a model that will finish an epoch. If anyone has any suggestions, please let me know. </p>",
      "rawMarkdown": "```python\nimport networkx as nx\nimport torch\nfrom rdkit import Chem\nfrom torch_geometric.utils.convert import from_networkx\n\nclass SMILESToGraph:\n    def __init__(self):\n        pass\n\n    def smiles_to_graph(self, smiles):\n        molecule = Chem.MolFromSmiles(smiles)\n        if molecule is None:\n            raise ValueError(\"Invalid SMILES string provided.\")\n\n        G = nx.Graph()\n        for atom in molecule.GetAtoms():\n            G.add_node(atom.GetIdx(),\n                       x=torch.tensor([\n                           atom.GetAtomicNum(),\n                           atom.GetFormalCharge(),\n                           atom.GetChiralTag().real,\n                           atom.GetHybridization().real,\n                           atom.GetNumExplicitHs(),\n                           atom.GetNumImplicitHs()],\n                           dtype=torch.float))\n\n        bond_type_to_int = {Chem.rdchem.BondType.SINGLE: 0,\n                            Chem.rdchem.BondType.DOUBLE: 1,\n                            Chem.rdchem.BondType.TRIPLE: 2,\n                            Chem.rdchem.BondType.AROMATIC: 3}\n        for bond in molecule.GetBonds():\n            G.add_edge(bond.GetBeginAtomIdx(), bond.GetEndAtomIdx(),\n                       edge_attr=torch.tensor([\n                           bond_type_to_int[bond.GetBondType()]], dtype=torch.float))\n\n        data = from_networkx(G, group_node_attrs=['x'], group_edge_attrs=['edge_attr'])\n        return data\n\n    def smiles_to_gnn_input(self, smiles):\n        return self.smiles_to_graph(smiles)\n\n    def print_graph_sizes(self, smiles_list):\n        for smiles in smiles_list:\n            try:\n                data = self.smiles_to_gnn_input(smiles)\n                total_size = 0\n                for _, tensor in data:\n                    total_size += tensor.element_size() * tensor.nelement()\n                print(f\"Memory size for {smiles}: {total_size} bytes\")\n            except ValueError as e:\n                print(f\"Failed for {smiles}: {str(e)}\")\n\nsmiles_examples = [\"CCO\", \"O=C=O\", \"C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1\"]\nconverter = SMILESToGraph()\nconverter.print_graph_sizes(smiles_examples)\n\n```\n`Memory size for C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1: 2744 bytes`\n\n```python\nimport torch\nfrom torch_geometric.nn import TransformerConv, global_mean_pool\n\n\nclass BindingAffinityPredictor(torch.nn.Module):\n    def __init__(self, node_features, edge_features, hidden_dim, num_heads):\n        super(BindingAffinityPredictor, self).__init__()\n        self.conv1 = TransformerConv(node_features, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.conv2 = TransformerConv(hidden_dim * num_heads, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.out = torch.nn.Linear(hidden_dim * num_heads, 1)  # Output dimension set to 1\n\n    def forward(self, data):\n        x, edge_index, edge_attr, batch = data.x, data.edge_index, data.edge_attr, data.batch\n        x = self.conv1(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = self.conv2(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = global_mean_pool(x, batch)  # Average pooling over all nodes in each graph in the batch\n        x = self.out(x)\n        x = torch.sigmoid(x)  # Apply sigmoid activation function\n        return x\n```\n\n\nThis is a warning for anyone who might go down this path. I attempted to take the smiles representations and and use rdkit to convert them to graphs and feed them into a binary class gnn with some fancy layers. I was able to get the model to converge, and from what I can tell this method is a state of the art way of handling molecules with GNNs. However, I hope you have a lot of $, because this thing is a beast. As you can see, if you convert the smiles to even a simple graph representation, a single molecule in memory is 2.7kb. There are 98 million samples for one 1/3 proteins. You will find that even simply converting the smiles to graphs will take months of processing. If you attempt to process them all before hand, which is what I attempted as well, you will have terabytes of graph data, which of course will not fit in memory and which will become just as difficult and slow to read from disk as in memory. I have since given up on trying to implement a GNN for this, you have to down sample so much to even get a model that will finish an epoch. If anyone has any suggestions, please let me know. \n\n",
      "votes": 15
    },
    {
      "id": 2780265,
      "postDate": "2024-04-28T05:40:49.700Z",
      "content": "<p>i want to suggest a a possible solution:</p>\n<ol>\n<li>get to the top ranking fast </li>\n<li>work on small dataset to prove that you have a top solution and lack of resources</li>\n<li>write to some cloud provider, especially those for biomedical/drug discovery and ask them to sponsor you</li>\n</ol>",
      "rawMarkdown": "i want to suggest a a possible solution:\n1. get to the top ranking fast \n2. work on small dataset to prove that you have a top solution and lack of resources\n3. write to some cloud provider, especially those for biomedical/drug discovery and ask them to sponsor you",
      "votes": 6,
      "replies": [
        {
          "id": 2780299,
          "postDate": "2024-04-28T06:07:11.957Z",
          "content": "<p>Its not me claiming that a gnn using transformer cnn layers are a top solution :)</p>\n<p>Other suggestions are good. </p>",
          "rawMarkdown": "Its not me claiming that a gnn using transformer cnn layers are a top solution :)\n\nOther suggestions are good. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2781726,
      "postDate": "2024-04-29T00:55:47.507Z",
      "content": "<p>You can save all graph data in less than twice the size of the smile data. This is very manageable, especially on TPU. Keep the nodes and edges in int8 and the edges as sparse matrices. Good luck.</p>",
      "rawMarkdown": "You can save all graph data in less than twice the size of the smile data. This is very manageable, especially on TPU. Keep the nodes and edges in int8 and the edges as sparse matrices. Good luck.",
      "votes": 3,
      "replies": [
        {
          "id": 2781731,
          "postDate": "2024-04-29T01:13:11.077Z",
          "content": "<p>Also, this is how to process fast:  </p>\n<p>!pip install pysmiles<br>\nfrom pysmiles import read_smiles<br>\nimport networkx as nx<br>\nmol = read_smiles(smile)<br>\nnodes = mol.nodes(data='element')<br>\nnodes = [x for x in nodes]<br>\nadj_matrix = nx.to_scipy_sparse_array(mol, weight='order')  </p>\n<p>For me I calculate processing and saving all the data at about 100h = 20h with 5 Kaggle sessions…Not a problem. <br>\nAlso, try first training on all positives and about say 2-3M negatives, this is less than 1/10 of the data, you can get preliminary results fast.</p>",
          "rawMarkdown": "Also, this is how to process fast:  \n\n\n\n!pip install pysmiles\nfrom pysmiles import read_smiles\nimport networkx as nx\nmol = read_smiles(smile)\nnodes = mol.nodes(data='element')\nnodes = [x for x in nodes]\nadj_matrix = nx.to_scipy_sparse_array(mol, weight='order')  \n\n\n\nFor me I calculate processing and saving all the data at about 100h = 20h with 5 Kaggle sessions...Not a problem. \nAlso, try first training on all positives and about say 2-3M negatives, this is less than 1/10 of the data, you can get preliminary results fast.",
          "votes": 8,
          "replies": [
            {
              "id": 2783699,
              "postDate": "2024-04-29T22:10:23.257Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2783718,
              "postDate": "2024-04-29T22:41:57.570Z",
              "content": "<p>Thank you for your comment, I am looking into this, but your comment also made me realize that I made a huge mistake. I should use mixed precision, additionally you should be using multiple workers to convert the smiles dataset on the fly.</p>",
              "rawMarkdown": "Thank you for your comment, I am looking into this, but your comment also made me realize that I made a huge mistake. I should use mixed precision, additionally you should be using multiple workers to convert the smiles dataset on the fly.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2781070,
      "postDate": "2024-04-28T14:41:58.067Z",
      "content": "<p>I agree that training on 98 million graphs 3X might be a bit tricky. However, I believe if you are clever with your sampling scheme, you may find that you don't need all 98 million rows of data. This is something I will personally be looking into later in the competition, but I haven't done yet.</p>\n<p>For ideas on a sampling scheme, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> has a nice discussion post that helps understand the distributions of building blocks used to make this library: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">Hatch Discussion Post</a></p>\n<p>Perhaps you can use the fact that there &lt;1k building blocks at each position to help you when deciding on how to subsample the data for training.</p>",
      "rawMarkdown": "I agree that training on 98 million graphs 3X might be a bit tricky. However, I believe if you are clever with your sampling scheme, you may find that you don't need all 98 million rows of data. This is something I will personally be looking into later in the competition, but I haven't done yet.\n\nFor ideas on a sampling scheme, @roberthatch has a nice discussion post that helps understand the distributions of building blocks used to make this library: [Hatch Discussion Post](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576)\n\nPerhaps you can use the fact that there <1k building blocks at each position to help you when deciding on how to subsample the data for training.",
      "votes": 2,
      "replies": [
        {
          "id": 2781419,
          "postDate": "2024-04-28T19:10:52.213Z",
          "content": "<p>Yeah, I agree you don't need 98 million, but even when I down sampled to like 2 million, when you combine generating the graph representation AND the training, even on a tiny gnn with 8 hidden layers, it takes way too long, even on an A100. </p>",
          "rawMarkdown": "Yeah, I agree you don't need 98 million, but even when I down sampled to like 2 million, when you combine generating the graph representation AND the training, even on a tiny gnn with 8 hidden layers, it takes way too long, even on an A100. ",
          "votes": 1,
          "replies": [
            {
              "id": 2788997,
              "postDate": "2024-05-02T14:00:07.623Z",
              "content": "<p>Depending on the strategy you are using, 8 layers are not necessarily needed. For what I have seen, going over 4 layers has been unproductive for me. Also I am using a 50:50 subsample of 0s (with the 0s being chosen at random) and 1s</p>",
              "rawMarkdown": "Depending on the strategy you are using, 8 layers are not necessarily needed. For what I have seen, going over 4 layers has been unproductive for me. Also I am using a 50:50 subsample of 0s (with the 0s being chosen at random) and 1s",
              "votes": 1
            },
            {
              "id": 2789153,
              "postDate": "2024-05-02T15:19:41.420Z",
              "content": "<p>Depending on the specific message passing scheme, going beyond even 2-3 layers can be detrimental as you can have graph over-smoothing.</p>",
              "rawMarkdown": "Depending on the specific message passing scheme, going beyond even 2-3 layers can be detrimental as you can have graph over-smoothing."
            }
          ]
        }
      ]
    },
    {
      "id": 2780674,
      "postDate": "2024-04-28T11:19:25.920Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2780265,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-28T05:40:49.700000",
      "content": "<p>i want to suggest a a possible solution:</p>\n<ol>\n<li>get to the top ranking fast </li>\n<li>work on small dataset to prove that you have a top solution and lack of resources</li>\n<li>write to some cloud provider, especially those for biomedical/drug discovery and ask them to sponsor you</li>\n</ol>",
      "votes": 6,
      "replies": [
        {
          "id": 2780299,
          "author_name": "Kyle Sørensen",
          "author_url": "",
          "post_date": "2024-04-28T06:07:11.957000",
          "content": "<p>Its not me claiming that a gnn using transformer cnn layers are a top solution :)</p>\n<p>Other suggestions are good. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2781726,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-04-29T00:55:47.507000",
      "content": "<p>You can save all graph data in less than twice the size of the smile data. This is very manageable, especially on TPU. Keep the nodes and edges in int8 and the edges as sparse matrices. Good luck.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2781731,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-04-29T01:13:11.077000",
          "content": "<p>Also, this is how to process fast:  </p>\n<p>!pip install pysmiles<br>\nfrom pysmiles import read_smiles<br>\nimport networkx as nx<br>\nmol = read_smiles(smile)<br>\nnodes = mol.nodes(data='element')<br>\nnodes = [x for x in nodes]<br>\nadj_matrix = nx.to_scipy_sparse_array(mol, weight='order')  </p>\n<p>For me I calculate processing and saving all the data at about 100h = 20h with 5 Kaggle sessions…Not a problem. <br>\nAlso, try first training on all positives and about say 2-3M negatives, this is less than 1/10 of the data, you can get preliminary results fast.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2783699,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-29T22:10:23.257000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2783718,
              "author_name": "Kyle Sørensen",
              "author_url": "",
              "post_date": "2024-04-29T22:41:57.570000",
              "content": "<p>Thank you for your comment, I am looking into this, but your comment also made me realize that I made a huge mistake. I should use mixed precision, additionally you should be using multiple workers to convert the smiles dataset on the fly.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2781070,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-04-28T14:41:58.067000",
      "content": "<p>I agree that training on 98 million graphs 3X might be a bit tricky. However, I believe if you are clever with your sampling scheme, you may find that you don't need all 98 million rows of data. This is something I will personally be looking into later in the competition, but I haven't done yet.</p>\n<p>For ideas on a sampling scheme, <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> has a nice discussion post that helps understand the distributions of building blocks used to make this library: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">Hatch Discussion Post</a></p>\n<p>Perhaps you can use the fact that there &lt;1k building blocks at each position to help you when deciding on how to subsample the data for training.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2781419,
          "author_name": "Kyle Sørensen",
          "author_url": "",
          "post_date": "2024-04-28T19:10:52.213000",
          "content": "<p>Yeah, I agree you don't need 98 million, but even when I down sampled to like 2 million, when you combine generating the graph representation AND the training, even on a tiny gnn with 8 hidden layers, it takes way too long, even on an A100. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2788997,
              "author_name": "Giorgio Micaletto",
              "author_url": "",
              "post_date": "2024-05-02T14:00:07.623000",
              "content": "<p>Depending on the strategy you are using, 8 layers are not necessarily needed. For what I have seen, going over 4 layers has been unproductive for me. Also I am using a 50:50 subsample of 0s (with the 0s being chosen at random) and 1s</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2789153,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-05-02T15:19:41.420000",
              "content": "<p>Depending on the specific message passing scheme, going beyond even 2-3 layers can be detrimental as you can have graph over-smoothing.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2780674,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-28T11:19:25.920000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2780170": "```python\nimport networkx as nx\nimport torch\nfrom rdkit import Chem\nfrom torch_geometric.utils.convert import from_networkx\n\nclass SMILESToGraph:\n    def __init__(self):\n        pass\n\n    def smiles_to_graph(self, smiles):\n        molecule = Chem.MolFromSmiles(smiles)\n        if molecule is None:\n            raise ValueError(\"Invalid SMILES string provided.\")\n\n        G = nx.Graph()\n        for atom in molecule.GetAtoms():\n            G.add_node(atom.GetIdx(),\n                       x=torch.tensor([\n                           atom.GetAtomicNum(),\n                           atom.GetFormalCharge(),\n                           atom.GetChiralTag().real,\n                           atom.GetHybridization().real,\n                           atom.GetNumExplicitHs(),\n                           atom.GetNumImplicitHs()],\n                           dtype=torch.float))\n\n        bond_type_to_int = {Chem.rdchem.BondType.SINGLE: 0,\n                            Chem.rdchem.BondType.DOUBLE: 1,\n                            Chem.rdchem.BondType.TRIPLE: 2,\n                            Chem.rdchem.BondType.AROMATIC: 3}\n        for bond in molecule.GetBonds():\n            G.add_edge(bond.GetBeginAtomIdx(), bond.GetEndAtomIdx(),\n                       edge_attr=torch.tensor([\n                           bond_type_to_int[bond.GetBondType()]], dtype=torch.float))\n\n        data = from_networkx(G, group_node_attrs=['x'], group_edge_attrs=['edge_attr'])\n        return data\n\n    def smiles_to_gnn_input(self, smiles):\n        return self.smiles_to_graph(smiles)\n\n    def print_graph_sizes(self, smiles_list):\n        for smiles in smiles_list:\n            try:\n                data = self.smiles_to_gnn_input(smiles)\n                total_size = 0\n                for _, tensor in data:\n                    total_size += tensor.element_size() * tensor.nelement()\n                print(f\"Memory size for {smiles}: {total_size} bytes\")\n            except ValueError as e:\n                print(f\"Failed for {smiles}: {str(e)}\")\n\nsmiles_examples = [\"CCO\", \"O=C=O\", \"C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1\"]\nconverter = SMILESToGraph()\nconverter.print_graph_sizes(smiles_examples)\n\n```\n`Memory size for C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1: 2744 bytes`\n\n```python\nimport torch\nfrom torch_geometric.nn import TransformerConv, global_mean_pool\n\n\nclass BindingAffinityPredictor(torch.nn.Module):\n    def __init__(self, node_features, edge_features, hidden_dim, num_heads):\n        super(BindingAffinityPredictor, self).__init__()\n        self.conv1 = TransformerConv(node_features, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.conv2 = TransformerConv(hidden_dim * num_heads, hidden_dim, heads=num_heads, edge_dim=edge_features)\n        self.out = torch.nn.Linear(hidden_dim * num_heads, 1)  # Output dimension set to 1\n\n    def forward(self, data):\n        x, edge_index, edge_attr, batch = data.x, data.edge_index, data.edge_attr, data.batch\n        x = self.conv1(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = self.conv2(x, edge_index, edge_attr)\n        x = torch.nn.functional.relu(x)\n        x = global_mean_pool(x, batch)  # Average pooling over all nodes in each graph in the batch\n        x = self.out(x)\n        x = torch.sigmoid(x)  # Apply sigmoid activation function\n        return x\n```\n\n\nThis is a warning for anyone who might go down this path. I attempted to take the smiles representations and and use rdkit to convert them to graphs and feed them into a binary class gnn with some fancy layers. I was able to get the model to converge, and from what I can tell this method is a state of the art way of handling molecules with GNNs. However, I hope you have a lot of $, because this thing is a beast. As you can see, if you convert the smiles to even a simple graph representation, a single molecule in memory is 2.7kb. There are 98 million samples for one 1/3 proteins. You will find that even simply converting the smiles to graphs will take months of processing. If you attempt to process them all before hand, which is what I attempted as well, you will have terabytes of graph data, which of course will not fit in memory and which will become just as difficult and slow to read from disk as in memory. I have since given up on trying to implement a GNN for this, you have to down sample so much to even get a model that will finish an epoch. If anyone has any suggestions, please let me know. \n\n",
    "2780265": "i want to suggest a a possible solution:\n1. get to the top ranking fast \n2. work on small dataset to prove that you have a top solution and lack of resources\n3. write to some cloud provider, especially those for biomedical/drug discovery and ask them to sponsor you",
    "2781726": "You can save all graph data in less than twice the size of the smile data. This is very manageable, especially on TPU. Keep the nodes and edges in int8 and the edges as sparse matrices. Good luck.",
    "2781070": "I agree that training on 98 million graphs 3X might be a bit tricky. However, I believe if you are clever with your sampling scheme, you may find that you don't need all 98 million rows of data. This is something I will personally be looking into later in the competition, but I haven't done yet.\n\nFor ideas on a sampling scheme, @roberthatch has a nice discussion post that helps understand the distributions of building blocks used to make this library: [Hatch Discussion Post](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576)\n\nPerhaps you can use the fact that there <1k building blocks at each position to help you when deciding on how to subsample the data for training.",
    "2780674": ""
  }
}