{
  "id": 500584,
  "title": "RDKit handling of [Dy] in smiles",
  "url": "/competitions/leash-BELKA/discussion/500584",
  "author_name": "",
  "post_date": "2024-05-06T07:49:07.022435800Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Question to folks with chemistry knowledge - I'm using RDKit to process the molecule smiles and extract atoms and features. Given that the DNA linker is encoded as [Dy], I believe RDKit assumes it's a dysprosium atom, and the feature vector for it will be wrong. Is my understanding correct? Would it make sense to try any pre/post-processing to handle that? </p>",
  "messages": [
    {
      "id": "2796389",
      "postDate": "05/06/2024 07:49:07",
      "content": "<p>Question to folks with chemistry knowledge - I'm using RDKit to process the molecule smiles and extract atoms and features. Given that the DNA linker is encoded as [Dy], I believe RDKit assumes it's a dysprosium atom, and the feature vector for it will be wrong. Is my understanding correct? Would it make sense to try any pre/post-processing to handle that? </p>",
      "rawMarkdown": "Question to folks with chemistry knowledge - I'm using RDKit to process the molecule smiles and extract atoms and features. Given that the DNA linker is encoded as [Dy], I believe RDKit assumes it's a dysprosium atom, and the feature vector for it will be wrong. Is my understanding correct? Would it make sense to try any pre/post-processing to handle that?",
      "votes": null
    },
    {
      "id": "2796432",
      "postDate": "05/06/2024 08:21:24",
      "content": "<p>Yes, RDKit considers it as a metal atom. This would pose problems if you use physchem descriptors (molecular weight, lipophilicity, etc).  One common approach to handle this is to replace Dy with a methyl group.</p>\n<pre><code> rdkit  Chem\n rdkit.Chem  AllChem\n ():\n\n    test_molecule = Chem.MolFromSmiles(smi)\n\n    Me = Chem.MolFromSmiles()\n\n    \n    Dy = Chem.MolFromSmiles()\n\n    new_mol = AllChem.ReplaceSubstructs(test_molecule, Dy, Me)\n\n     Chem.MolToSmiles(new_mol[])\n</code></pre>",
      "rawMarkdown": "Yes, RDKit considers it as a metal atom. This would pose problems if you use physchem descriptors (molecular weight, lipophilicity, etc).  One common approach to handle this is to replace Dy with a methyl group.\n\n\n```python\nfrom rdkit import Chem\nfrom rdkit.Chem import AllChem\ndef transform_Dy(smi):\n\n    test_molecule = Chem.MolFromSmiles(smi)\n\n    Me = Chem.MolFromSmiles('C')\n\n    #create Dy mol object\n    Dy = Chem.MolFromSmiles('[Dy]')\n\n    new_mol = AllChem.ReplaceSubstructs(test_molecule, Dy, Me)\n\n    return Chem.MolToSmiles(new_mol[0])\n```",
      "votes": null
    },
    {
      "id": "2796450",
      "postDate": "05/06/2024 08:32:29",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "2796693",
      "postDate": "05/06/2024 11:28:47",
      "content": "<p>Like <a href=\"https://www.kaggle.com/ruelcedeno\" target=\"_blank\">@ruelcedeno</a> said, if you want to use 3D structures or descriptors it definitely matters. I made a notebook about how to handle this: <a href=\"https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations\" target=\"_blank\">cheminformatics transformations</a></p>\n<p>For 2D descriptors like SMILES or fingerprints I think it's ok to leave it if you assume the attachment point doesn't contribute to binding much, though I haven't experimented much with this yet. I played around with ethylene glycol derived descriptors for 2D and I achieved slightly higher LB scores when I replaced it, but I haven't dug deeply into this yet in a systematic way.</p>",
      "rawMarkdown": "Like @ruelcedeno said, if you want to use 3D structures or descriptors it definitely matters. I made a notebook about how to handle this: [cheminformatics transformations](https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations)\n\nFor 2D descriptors like SMILES or fingerprints I think it's ok to leave it if you assume the attachment point doesn't contribute to binding much, though I haven't experimented much with this yet. I played around with ethylene glycol derived descriptors for 2D and I achieved slightly higher LB scores when I replaced it, but I haven't dug deeply into this yet in a systematic way.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2796432,
      "author_name": "ruelcedeno",
      "author_url": "",
      "post_date": "05/06/2024 08:21:24",
      "content": "<p>Yes, RDKit considers it as a metal atom. This would pose problems if you use physchem descriptors (molecular weight, lipophilicity, etc).  One common approach to handle this is to replace Dy with a methyl group.</p>\n<pre><code> rdkit  Chem\n rdkit.Chem  AllChem\n ():\n\n    test_molecule = Chem.MolFromSmiles(smi)\n\n    Me = Chem.MolFromSmiles()\n\n    \n    Dy = Chem.MolFromSmiles()\n\n    new_mol = AllChem.ReplaceSubstructs(test_molecule, Dy, Me)\n\n     Chem.MolToSmiles(new_mol[])\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2796450,
          "author_name": "thedrcat",
          "author_url": "",
          "post_date": "05/06/2024 08:32:29",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2796693,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "05/06/2024 11:28:47",
      "content": "<p>Like <a href=\"https://www.kaggle.com/ruelcedeno\" target=\"_blank\">@ruelcedeno</a> said, if you want to use 3D structures or descriptors it definitely matters. I made a notebook about how to handle this: <a href=\"https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations\" target=\"_blank\">cheminformatics transformations</a></p>\n<p>For 2D descriptors like SMILES or fingerprints I think it's ok to leave it if you assume the attachment point doesn't contribute to binding much, though I haven't experimented much with this yet. I played around with ethylene glycol derived descriptors for 2D and I achieved slightly higher LB scores when I replaced it, but I haven't dug deeply into this yet in a systematic way.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2796389": "Question to folks with chemistry knowledge - I'm using RDKit to process the molecule smiles and extract atoms and features. Given that the DNA linker is encoded as [Dy], I believe RDKit assumes it's a dysprosium atom, and the feature vector for it will be wrong. Is my understanding correct? Would it make sense to try any pre/post-processing to handle that?",
    "2796432": "Yes, RDKit considers it as a metal atom. This would pose problems if you use physchem descriptors (molecular weight, lipophilicity, etc).  One common approach to handle this is to replace Dy with a methyl group.\n\n\n```python\nfrom rdkit import Chem\nfrom rdkit.Chem import AllChem\ndef transform_Dy(smi):\n\n    test_molecule = Chem.MolFromSmiles(smi)\n\n    Me = Chem.MolFromSmiles('C')\n\n    #create Dy mol object\n    Dy = Chem.MolFromSmiles('[Dy]')\n\n    new_mol = AllChem.ReplaceSubstructs(test_molecule, Dy, Me)\n\n    return Chem.MolToSmiles(new_mol[0])\n```",
    "2796450": "Thanks a lot!",
    "2796693": "Like @ruelcedeno said, if you want to use 3D structures or descriptors it definitely matters. I made a notebook about how to handle this: [cheminformatics transformations](https://www.kaggle.com/code/chemdatafarmer/cheminformatics-transformations)\n\nFor 2D descriptors like SMILES or fingerprints I think it's ok to leave it if you assume the attachment point doesn't contribute to binding much, though I haven't experimented much with this yet. I played around with ethylene glycol derived descriptors for 2D and I achieved slightly higher LB scores when I replaced it, but I haven't dug deeply into this yet in a systematic way."
  },
  "source": "meta"
}