{
  "id": 499607,
  "title": "[Question] Useful Graph Features",
  "url": "/competitions/leash-BELKA/discussion/499607",
  "author_name": "",
  "post_date": "2024-05-02T13:08:41.109499600Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n<p>Thanks to my university's support, I was able to train a relatively deep Graph Attention Network with 3477857 parameters. I've been using a combination of 3 million randomly chosen 0s and all available 1s for training.</p>\n<p>I'm currently achieving a score of 0.499 LB, and I'm reaching out to see if there might be ways to improve the model's feature representation. I've experimented with adding more data and going deeper in the network architecture, but the results haven't changed significantly (gained 0.002).</p>\n<p>Here are the features I'm currently using:</p>\n<pre><code> ():\n    features = []\n     atom  mol.GetAtoms():\n        features.append([\n            atom.GetAtomicNum(),  \n            atom.GetDegree(),     \n            atom.GetTotalValence(),    \n            atom.GetFormalCharge(),  \n            atom.GetHybridization().real,  \n            atom.GetIsAromatic(),  \n            atom.GetTotalNumHs(),  \n            atom.IsInRing(),       \n            atom.GetMass(),        \n            atom.GetChiralTag().real, \n            atom.GetExplicitValence(), \n            atom.GetImplicitValence(), \n            atom.GetNumRadicalElectrons() \n        ])\n     features\n\n ():\n    bond_features = []\n     bond  mol.GetBonds():\n        bond_features.append([\n            bond.GetBondTypeAsDouble(),  \n            bond.GetIsConjugated(),      \n            bond.GetIsAromatic(),        \n            bond.GetStereo(),            \n            bond.IsInRing(),             \n            bond.GetBeginAtomIdx(),      \n            bond.GetEndAtomIdx()         \n        ])\n     bond_features\n</code></pre>\n<p>I'm training the model on both the individual building blocks and the entire molecule.</p>\n<p>I'd be grateful for any insights you might have on the following:</p>\n<ul>\n<li>What additional features could be beneficial to consider?</li>\n<li>Have you encountered any effective data splitting strategies? (I tried leaving one (or more) building blocks out but it didn't really have any effect)</li>\n</ul>\n<p>I'm also very open to collaborating with anyone who has more experience in this area. This is only my second competition, and I'm always eager to learn. Thanks in advance for your time and advice!</p>",
  "messages": [
    {
      "id": "2788857",
      "postDate": "05/02/2024 13:08:41",
      "content": "<p>Hi everyone!</p>\n<p>Thanks to my university's support, I was able to train a relatively deep Graph Attention Network with 3477857 parameters. I've been using a combination of 3 million randomly chosen 0s and all available 1s for training.</p>\n<p>I'm currently achieving a score of 0.499 LB, and I'm reaching out to see if there might be ways to improve the model's feature representation. I've experimented with adding more data and going deeper in the network architecture, but the results haven't changed significantly (gained 0.002).</p>\n<p>Here are the features I'm currently using:</p>\n<pre><code> ():\n    features = []\n     atom  mol.GetAtoms():\n        features.append([\n            atom.GetAtomicNum(),  \n            atom.GetDegree(),     \n            atom.GetTotalValence(),    \n            atom.GetFormalCharge(),  \n            atom.GetHybridization().real,  \n            atom.GetIsAromatic(),  \n            atom.GetTotalNumHs(),  \n            atom.IsInRing(),       \n            atom.GetMass(),        \n            atom.GetChiralTag().real, \n            atom.GetExplicitValence(), \n            atom.GetImplicitValence(), \n            atom.GetNumRadicalElectrons() \n        ])\n     features\n\n ():\n    bond_features = []\n     bond  mol.GetBonds():\n        bond_features.append([\n            bond.GetBondTypeAsDouble(),  \n            bond.GetIsConjugated(),      \n            bond.GetIsAromatic(),        \n            bond.GetStereo(),            \n            bond.IsInRing(),             \n            bond.GetBeginAtomIdx(),      \n            bond.GetEndAtomIdx()         \n        ])\n     bond_features\n</code></pre>\n<p>I'm training the model on both the individual building blocks and the entire molecule.</p>\n<p>I'd be grateful for any insights you might have on the following:</p>\n<ul>\n<li>What additional features could be beneficial to consider?</li>\n<li>Have you encountered any effective data splitting strategies? (I tried leaving one (or more) building blocks out but it didn't really have any effect)</li>\n</ul>\n<p>I'm also very open to collaborating with anyone who has more experience in this area. This is only my second competition, and I'm always eager to learn. Thanks in advance for your time and advice!</p>",
      "rawMarkdown": "Hi everyone!\n\nThanks to my university's support, I was able to train a relatively deep Graph Attention Network with 3477857 parameters. I've been using a combination of 3 million randomly chosen 0s and all available 1s for training.\n\nI'm currently achieving a score of 0.499 LB, and I'm reaching out to see if there might be ways to improve the model's feature representation. I've experimented with adding more data and going deeper in the network architecture, but the results haven't changed significantly (gained 0.002).\n\nHere are the features I'm currently using:\n```python\ndef get_atom_features(mol):\n    features = []\n    for atom in mol.GetAtoms():\n        features.append([\n            atom.GetAtomicNum(),  ## Atomic number\n            atom.GetDegree(),     ## Degree, number of bonded atoms\n            atom.GetTotalValence(),    ## Valence\n            atom.GetFormalCharge(),  ## Formal charge\n            atom.GetHybridization().real,  ## Hybridization state, converted to a real number\n            atom.GetIsAromatic(),  ## Aromaticity, boolean converted to int\n            atom.GetTotalNumHs(),  ## Total number of attached hydrogen atoms\n            atom.IsInRing(),       ## Is the atom in a ring? Boolean converted to int\n            atom.GetMass(),        ## Atomic mass\n            atom.GetChiralTag().real, ## Chirality, converted to a real number\n            atom.GetExplicitValence(), ## Explicit valence\n            atom.GetImplicitValence(), ## Implicit valence\n            atom.GetNumRadicalElectrons() ## Number of radical electrons\n        ])\n    return features\n\ndef get_bond_features(mol):\n    bond_features = []\n    for bond in mol.GetBonds():\n        bond_features.append([\n            bond.GetBondTypeAsDouble(),  ## Bond type (single, double, etc.) as double\n            bond.GetIsConjugated(),      ## Conjugation, boolean converted to int\n            bond.GetIsAromatic(),        ## Aromaticity, boolean converted to int\n            bond.GetStereo(),            ## Stereochemistry of the bond\n            bond.IsInRing(),             ## Is the bond in a ring? Boolean converted to int\n            bond.GetBeginAtomIdx(),      ## Index of the start atom of the bond\n            bond.GetEndAtomIdx()         ## Index of the end atom of the bond\n        ])\n    return bond_features\n```\nI'm training the model on both the individual building blocks and the entire molecule.\n\nI'd be grateful for any insights you might have on the following:\n- What additional features could be beneficial to consider?\n- Have you encountered any effective data splitting strategies? (I tried leaving one (or more) building blocks out but it didn't really have any effect)\n\nI'm also very open to collaborating with anyone who has more experience in this area. This is only my second competition, and I'm always eager to learn. Thanks in advance for your time and advice!",
      "votes": null
    },
    {
      "id": "2789908",
      "postDate": "05/02/2024 22:26:20",
      "content": "<p>I don't know for the features, but for the GNN have you considered using rotation-invariant GNNs ?</p>",
      "rawMarkdown": "I don't know for the features, but for the GNN have you considered using rotation-invariant GNNs ?",
      "votes": null
    },
    {
      "id": "2789910",
      "postDate": "05/02/2024 22:27:39",
      "content": "<p>maybe using rdkit you can compute some additional easy-to-compute features ?</p>",
      "rawMarkdown": "maybe using rdkit you can compute some additional easy-to-compute features ?",
      "votes": null
    },
    {
      "id": "2790073",
      "postDate": "05/03/2024 02:48:22",
      "content": "<p>my suggestion is first to design GNN to match existing approaches.<br>\nyou have three here:</p>\n<p>1) xgboost + EFCP<br>\nEFCP is like partial graph. the input information is only atom and bonds. if you use radius=2, it means your gnn needs only 2 to 4 message passing layers. nBits=2048, so your GNN should be very wide.</p>\n<p>2) smiles string and transformer<br>\nif you look at chemberta, it only has 3 roberta layers. now transformer have global context, think of how to add<br>\nglobal  context to GNN?see if you find anything about the attention map and link to adjacency map.</p>\n<p>3) 1dCNN<br>\nThis suggest global context is not needed.</p>\n<p>all the solutions mentioned (1),(2),(3) have lb score in range of 0.565 to 0.595 (single model using all train samples). they only only atoms and bond. </p>\n<hr>\n<p>study current approaches. what features are current approach using? how to represent in graph?<br>\nwhat are the deciding factors that make current approaches  work?num of train samples? num of parameters? input information?</p>\n<p>if you can get GNN to meet current performance, then you can think of improvement layer.</p>\n<hr>\n<p>also try to repeat opensource/paper GNN models. that is maybe easier.</p>",
      "rawMarkdown": "my suggestion is first to design GNN to match existing approaches.\nyou have three here:\n\n1) xgboost + EFCP\nEFCP is like partial graph. the input information is only atom and bonds. if you use radius=2, it means your gnn needs only 2 to 4 message passing layers. nBits=2048, so your GNN should be very wide.\n\n2) smiles string and transformer\nif you look at chemberta, it only has 3 roberta layers. now transformer have global context, think of how to add\nglobal  context to GNN?see if you find anything about the attention map and link to adjacency map.\n\n3) 1dCNN\nThis suggest global context is not needed.\n\nall the solutions mentioned (1),(2),(3) have lb score in range of 0.565 to 0.595 (single model using all train samples). they only only atoms and bond. \n\n----\n\nstudy current approaches. what features are current approach using? how to represent in graph?\nwhat are the deciding factors that make current approaches  work?num of train samples? num of parameters? input information?\n\nif you can get GNN to meet current performance, then you can think of improvement layer.\n\n\n----\n\nalso try to repeat opensource/paper GNN models. that is maybe easier.",
      "votes": null
    },
    {
      "id": "2790350",
      "postDate": "05/03/2024 06:42:32",
      "content": "<p>After having seeing that the 1dCNN is capable of achieving such a high score I have also added in tandem with the original Transformer structure also a Convolutional one, but I will be honest, 16M parameters is pushing what I am comfortable with. I will try to make it wider but the resources I have been allotted may not be enough for it. Thank you for your tip!</p>",
      "rawMarkdown": "After having seeing that the 1dCNN is capable of achieving such a high score I have also added in tandem with the original Transformer structure also a Convolutional one, but I will be honest, 16M parameters is pushing what I am comfortable with. I will try to make it wider but the resources I have been allotted may not be enough for it. Thank you for your tip!",
      "votes": null
    },
    {
      "id": "2790366",
      "postDate": "05/03/2024 06:51:28",
      "content": "<p>I have added a couple more features, mainly about the bond being rotable or not, but as suggested by hengck23, it may be worthwhile to also explore other types of models, </p>",
      "rawMarkdown": "I have added a couple more features, mainly about the bond being rotable or not, but as suggested by hengck23, it may be worthwhile to also explore other types of models,",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2789908,
      "author_name": "louisstefanuto",
      "author_url": "",
      "post_date": "05/02/2024 22:26:20",
      "content": "<p>I don't know for the features, but for the GNN have you considered using rotation-invariant GNNs ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2789910,
      "author_name": "louisstefanuto",
      "author_url": "",
      "post_date": "05/02/2024 22:27:39",
      "content": "<p>maybe using rdkit you can compute some additional easy-to-compute features ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2790366,
          "author_name": "giorgiomicaletto",
          "author_url": "",
          "post_date": "05/03/2024 06:51:28",
          "content": "<p>I have added a couple more features, mainly about the bond being rotable or not, but as suggested by hengck23, it may be worthwhile to also explore other types of models, </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2790073,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/03/2024 02:48:22",
      "content": "<p>my suggestion is first to design GNN to match existing approaches.<br>\nyou have three here:</p>\n<p>1) xgboost + EFCP<br>\nEFCP is like partial graph. the input information is only atom and bonds. if you use radius=2, it means your gnn needs only 2 to 4 message passing layers. nBits=2048, so your GNN should be very wide.</p>\n<p>2) smiles string and transformer<br>\nif you look at chemberta, it only has 3 roberta layers. now transformer have global context, think of how to add<br>\nglobal  context to GNN?see if you find anything about the attention map and link to adjacency map.</p>\n<p>3) 1dCNN<br>\nThis suggest global context is not needed.</p>\n<p>all the solutions mentioned (1),(2),(3) have lb score in range of 0.565 to 0.595 (single model using all train samples). they only only atoms and bond. </p>\n<hr>\n<p>study current approaches. what features are current approach using? how to represent in graph?<br>\nwhat are the deciding factors that make current approaches  work?num of train samples? num of parameters? input information?</p>\n<p>if you can get GNN to meet current performance, then you can think of improvement layer.</p>\n<hr>\n<p>also try to repeat opensource/paper GNN models. that is maybe easier.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2790350,
          "author_name": "giorgiomicaletto",
          "author_url": "",
          "post_date": "05/03/2024 06:42:32",
          "content": "<p>After having seeing that the 1dCNN is capable of achieving such a high score I have also added in tandem with the original Transformer structure also a Convolutional one, but I will be honest, 16M parameters is pushing what I am comfortable with. I will try to make it wider but the resources I have been allotted may not be enough for it. Thank you for your tip!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2788857": "Hi everyone!\n\nThanks to my university's support, I was able to train a relatively deep Graph Attention Network with 3477857 parameters. I've been using a combination of 3 million randomly chosen 0s and all available 1s for training.\n\nI'm currently achieving a score of 0.499 LB, and I'm reaching out to see if there might be ways to improve the model's feature representation. I've experimented with adding more data and going deeper in the network architecture, but the results haven't changed significantly (gained 0.002).\n\nHere are the features I'm currently using:\n```python\ndef get_atom_features(mol):\n    features = []\n    for atom in mol.GetAtoms():\n        features.append([\n            atom.GetAtomicNum(),  ## Atomic number\n            atom.GetDegree(),     ## Degree, number of bonded atoms\n            atom.GetTotalValence(),    ## Valence\n            atom.GetFormalCharge(),  ## Formal charge\n            atom.GetHybridization().real,  ## Hybridization state, converted to a real number\n            atom.GetIsAromatic(),  ## Aromaticity, boolean converted to int\n            atom.GetTotalNumHs(),  ## Total number of attached hydrogen atoms\n            atom.IsInRing(),       ## Is the atom in a ring? Boolean converted to int\n            atom.GetMass(),        ## Atomic mass\n            atom.GetChiralTag().real, ## Chirality, converted to a real number\n            atom.GetExplicitValence(), ## Explicit valence\n            atom.GetImplicitValence(), ## Implicit valence\n            atom.GetNumRadicalElectrons() ## Number of radical electrons\n        ])\n    return features\n\ndef get_bond_features(mol):\n    bond_features = []\n    for bond in mol.GetBonds():\n        bond_features.append([\n            bond.GetBondTypeAsDouble(),  ## Bond type (single, double, etc.) as double\n            bond.GetIsConjugated(),      ## Conjugation, boolean converted to int\n            bond.GetIsAromatic(),        ## Aromaticity, boolean converted to int\n            bond.GetStereo(),            ## Stereochemistry of the bond\n            bond.IsInRing(),             ## Is the bond in a ring? Boolean converted to int\n            bond.GetBeginAtomIdx(),      ## Index of the start atom of the bond\n            bond.GetEndAtomIdx()         ## Index of the end atom of the bond\n        ])\n    return bond_features\n```\nI'm training the model on both the individual building blocks and the entire molecule.\n\nI'd be grateful for any insights you might have on the following:\n- What additional features could be beneficial to consider?\n- Have you encountered any effective data splitting strategies? (I tried leaving one (or more) building blocks out but it didn't really have any effect)\n\nI'm also very open to collaborating with anyone who has more experience in this area. This is only my second competition, and I'm always eager to learn. Thanks in advance for your time and advice!",
    "2789908": "I don't know for the features, but for the GNN have you considered using rotation-invariant GNNs ?",
    "2789910": "maybe using rdkit you can compute some additional easy-to-compute features ?",
    "2790073": "my suggestion is first to design GNN to match existing approaches.\nyou have three here:\n\n1) xgboost + EFCP\nEFCP is like partial graph. the input information is only atom and bonds. if you use radius=2, it means your gnn needs only 2 to 4 message passing layers. nBits=2048, so your GNN should be very wide.\n\n2) smiles string and transformer\nif you look at chemberta, it only has 3 roberta layers. now transformer have global context, think of how to add\nglobal  context to GNN?see if you find anything about the attention map and link to adjacency map.\n\n3) 1dCNN\nThis suggest global context is not needed.\n\nall the solutions mentioned (1),(2),(3) have lb score in range of 0.565 to 0.595 (single model using all train samples). they only only atoms and bond. \n\n----\n\nstudy current approaches. what features are current approach using? how to represent in graph?\nwhat are the deciding factors that make current approaches  work?num of train samples? num of parameters? input information?\n\nif you can get GNN to meet current performance, then you can think of improvement layer.\n\n\n----\n\nalso try to repeat opensource/paper GNN models. that is maybe easier.",
    "2790350": "After having seeing that the 1dCNN is capable of achieving such a high score I have also added in tandem with the original Transformer structure also a Convolutional one, but I will be honest, 16M parameters is pushing what I am comfortable with. I will try to make it wider but the resources I have been allotted may not be enough for it. Thank you for your tip!",
    "2790366": "I have added a couple more features, mainly about the bond being rotable or not, but as suggested by hengck23, it may be worthwhile to also explore other types of models,"
  },
  "source": "meta"
}