{
  "id": 491431,
  "title": "[Dy] in test compounds",
  "url": "/competitions/leash-BELKA/discussion/491431",
  "author_name": "",
  "post_date": "2024-04-05T21:24:44.155403600Z",
  "votes": 10,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Can the organisers explain if this is actually intended? Every SMILES in the test dataset has a metal atom [Dy] in the SMILES representation. Is this actually intended?</p>",
  "messages": [
    {
      "id": "2737655",
      "postDate": "04/05/2024 21:24:44",
      "content": "<p>Can the organisers explain if this is actually intended? Every SMILES in the test dataset has a metal atom [Dy] in the SMILES representation. Is this actually intended?</p>",
      "rawMarkdown": "Can the organisers explain if this is actually intended? Every SMILES in the test dataset has a metal atom [Dy] in the SMILES representation. Is this actually intended?",
      "votes": null
    },
    {
      "id": "2737658",
      "postDate": "04/05/2024 21:25:46",
      "content": "<p>We use [Dy] as a stand-in for where the molecule is attached to the DNA tag</p>",
      "rawMarkdown": "We use [Dy] as a stand-in for where the molecule is attached to the DNA tag",
      "votes": null
    },
    {
      "id": "2737665",
      "postDate": "04/05/2024 21:31:48",
      "content": "<p>That makes sense! Thank you :D</p>",
      "rawMarkdown": "That makes sense! Thank you :D",
      "votes": null
    },
    {
      "id": "2737722",
      "postDate": "04/05/2024 22:59:13",
      "content": "<p>Could you please elaborate on this? For someone with not much biology knowledge (I use smiles and so for polymers though). Thank you.</p>",
      "rawMarkdown": "Could you please elaborate on this? For someone with not much biology knowledge (I use smiles and so for polymers though). Thank you.",
      "votes": null
    },
    {
      "id": "2738750",
      "postDate": "04/06/2024 15:06:30",
      "content": "<p>We chose Dy (dysprosium) as a stand-in for the DNA because it's unlikely to be in druglike chemicals and started with the letter \"D\". Nothing more complicated than that!</p>",
      "rawMarkdown": "We chose Dy (dysprosium) as a stand-in for the DNA because it's unlikely to be in druglike chemicals and started with the letter \"D\". Nothing more complicated than that!",
      "votes": null
    },
    {
      "id": "2740776",
      "postDate": "04/07/2024 23:26:44",
      "content": "<p>According to the competition's data tab, it means the following:</p>\n<blockquote>\n  <p>molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.</p>\n</blockquote>\n<p>Hope that helps!</p>",
      "rawMarkdown": "According to the competition's data tab, it means the following:\n\n>molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.\n\nHope that helps!",
      "votes": null
    },
    {
      "id": "2741323",
      "postDate": "04/08/2024 08:57:17",
      "content": "<p>if we remove DNA linker [Dy] from smiles in kaggle dataset, it should not affect results. Is that correct?  </p>\n<p>i ask this because public dataset (external dataset) are unlikely to include DNA linker.<br>\nI am thinking of how to use external dataset.</p>",
      "rawMarkdown": "if we remove DNA linker [Dy] from smiles in kaggle dataset, it should not affect results. Is that correct?  \n\ni ask this because public dataset (external dataset) are unlikely to include DNA linker.\nI am thinking of how to use external dataset.",
      "votes": null
    },
    {
      "id": "2741810",
      "postDate": "04/08/2024 15:52:10",
      "content": "<p>Hi  <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>Removal of the [Dy] is a bit tricky. When the molecule binds to the protein, the DNA is still attached. Since DNA is a large molecule and there is a chain linking the molecule of interest to DNA, this can prohibit (or at least modify) the binding of the molecule to the protein.</p>\n<p>There are many ways you can experiment with handling this (convert the [Dy] to an alkyl group like Me or butyl, for example), but I'd suspect that keeping some information about where the DNA is bound will likely be valuable to the model.</p>\n<p>For context, medicinal chemists that use DEL hits as starting points often use the DNA binding location as a vector to add other chemical probes or as ways to modify the physical properties of the molecule (improve solubility, for example) without modifying the binding. That said, the DNA attachment point can still provide a contribution to the binding (I've seen this first hand).</p>",
      "rawMarkdown": "Hi  @hengck23 \n\nRemoval of the [Dy] is a bit tricky. When the molecule binds to the protein, the DNA is still attached. Since DNA is a large molecule and there is a chain linking the molecule of interest to DNA, this can prohibit (or at least modify) the binding of the molecule to the protein.\n\nThere are many ways you can experiment with handling this (convert the [Dy] to an alkyl group like Me or butyl, for example), but I'd suspect that keeping some information about where the DNA is bound will likely be valuable to the model.\n\nFor context, medicinal chemists that use DEL hits as starting points often use the DNA binding location as a vector to add other chemical probes or as ways to modify the physical properties of the molecule (improve solubility, for example) without modifying the binding. That said, the DNA attachment point can still provide a contribution to the binding (I've seen this first hand).",
      "votes": null
    },
    {
      "id": "2741969",
      "postDate": "04/08/2024 17:13:14",
      "content": "<p>Great explanation, thanks! This also got me thinking, can the molecule bind to the protein through a building block to which the DNA is attached?  Because if not, then the location of Dy could indicate at least, which of the three building blocks is not responsible for the binding. I thought so, but it seems that building blocks can indirectly affect a molecule's ability to bind, and there is no way to determine (using domain knowledge, I mean, not ML) which of them is more important, is it correct?</p>",
      "rawMarkdown": "Great explanation, thanks! This also got me thinking, can the molecule bind to the protein through a building block to which the DNA is attached?  Because if not, then the location of Dy could indicate at least, which of the three building blocks is not responsible for the binding. I thought so, but it seems that building blocks can indirectly affect a molecule's ability to bind, and there is no way to determine (using domain knowledge, I mean, not ML) which of them is more important, is it correct?",
      "votes": null
    },
    {
      "id": "2742147",
      "postDate": "04/08/2024 19:03:24",
      "content": "<p>Yes, a building block attached to the DNA can contribute to the binding and it often does. However, the closer you get to the linker attachment point, the more likely you are to be headed towards solvent (water) and not deeper into the protein.</p>\n<p>I hope that helps!</p>",
      "rawMarkdown": "Yes, a building block attached to the DNA can contribute to the binding and it often does. However, the closer you get to the linker attachment point, the more likely you are to be headed towards solvent (water) and not deeper into the protein.\n\nI hope that helps!",
      "votes": null
    },
    {
      "id": "2744877",
      "postDate": "04/10/2024 07:13:49",
      "content": "<p>Is is there a simple and fast way (just bbs SMILES strings, and no subgraph matching) to tell which bb is going to be attached to the DNA molecule?</p>",
      "rawMarkdown": "Is is there a simple and fast way (just bbs SMILES strings, and no subgraph matching) to tell which bb is going to be attached to the DNA molecule?",
      "votes": null
    },
    {
      "id": "2745095",
      "postDate": "04/10/2024 11:22:59",
      "content": "<p>I believe it will be the BB1 that is attached to the DNA. This is how I interpret the figure in the overview - but would be nice with a confirmation from the organizers.</p>",
      "rawMarkdown": "I believe it will be the BB1 that is attached to the DNA. This is how I interpret the figure in the overview - but would be nice with a confirmation from the organizers.",
      "votes": null
    },
    {
      "id": "2745384",
      "postDate": "04/10/2024 15:30:32",
      "content": "<p>BB1 is always attached to the DNA</p>",
      "rawMarkdown": "BB1 is always attached to the DNA",
      "votes": null
    },
    {
      "id": "2747839",
      "postDate": "04/12/2024 05:18:42",
      "content": "<p>Excuse my ignorance but what do you mean with the last part? (headed towards solvent not protein)</p>",
      "rawMarkdown": "Excuse my ignorance but what do you mean with the last part? (headed towards solvent not protein)",
      "votes": null
    },
    {
      "id": "2750040",
      "postDate": "04/13/2024 12:26:56",
      "content": "<p>Molecules bind in 3-dimensions. Imagine the protein is a small bowl and the air around it is water. Now, imagine a tennis ball you have on the end of a stick. You can put the tennis ball in the bowl (imagine a snug fit) and the stick is pointing up. In this analogy, the stick is the DNA attachment linker and the molecule is the ball. The parts of the ball that fill the bowl and the parts that touch the bowl would contribute to binding and the parts that don't have little to no effect on binding.</p>\n<p>This is an oversimplification, but gets the idea across </p>",
      "rawMarkdown": "Molecules bind in 3-dimensions. Imagine the protein is a small bowl and the air around it is water. Now, imagine a tennis ball you have on the end of a stick. You can put the tennis ball in the bowl (imagine a snug fit) and the stick is pointing up. In this analogy, the stick is the DNA attachment linker and the molecule is the ball. The parts of the ball that fill the bowl and the parts that touch the bowl would contribute to binding and the parts that don't have little to no effect on binding.\n\nThis is an oversimplification, but gets the idea across",
      "votes": null
    },
    {
      "id": "2750488",
      "postDate": "04/13/2024 17:27:46",
      "content": "<p>Thank you for the intuitive explanation!</p>",
      "rawMarkdown": "Thank you for the intuitive explanation!",
      "votes": null
    },
    {
      "id": "2753995",
      "postDate": "04/15/2024 18:41:06",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a>, if trying to generate 3d molecules from smiles, and just wanting to simplify for now, any suggestion for one \"mostly good enough\" method?</p>\n<ul>\n<li>Remove [Dy]?</li>\n<li>Replace the text string \"[Dy]\" with the text string <em>__</em>?</li>\n</ul>",
      "rawMarkdown": "Hi @chemdatafarmer, if trying to generate 3d molecules from smiles, and just wanting to simplify for now, any suggestion for one \"mostly good enough\" method?\n\n- Remove [Dy]?\n- Replace the text string \"[Dy]\" with the text string ____?",
      "votes": null
    },
    {
      "id": "2754223",
      "postDate": "04/15/2024 22:28:50",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> there are a couple ways to do this, but the simplest way is probably to replace the [Dy] with a methyl (CH3) group. I'll probably make a notebook on this once I get deeper into my own 3D workflows, but this should do for now.</p>\n<p>Probably the most robust way to do that is as follows.</p>\n<pre><code> rdkit  Chem\n rdkit.Chem  AllChem\n\n\nmol = Chem.MolFromSmiles(SMILES)\n\n\nnew_attachment = Chem.MolFromSmiles()\n\n\ndy_pattern = Chem.MolFromSmiles()\n\n\nnew_mol = AllChem.ReplaceSubstructs(mol, dy_pattern, new_attachment)[]\n\n\nChem.SanitizeMol(new_mol)\n\n\nChem.AddHs(new_mol)\n</code></pre>\n<p>Then convert this to whatever 3D representation you want.</p>",
      "rawMarkdown": "Hey @roberthatch there are a couple ways to do this, but the simplest way is probably to replace the [Dy] with a methyl (CH3) group. I'll probably make a notebook on this once I get deeper into my own 3D workflows, but this should do for now.\n\nProbably the most robust way to do that is as follows.\n\n\n```python\nfrom rdkit import Chem\nfrom rdkit.Chem import AllChem\n\n#Convert your SMILES to a mol object.\nmol = Chem.MolFromSmiles(SMILES)\n\n#Create a mol object to replace the Dy atom with.\nnew_attachment = Chem.MolFromSmiles('C')\n\n#Get the pattern for the Dy atom\ndy_pattern = Chem.MolFromSmiles('[Dy]')\n\n#This returns a tuple of all possible replacements, but we know there will only be one.\nnew_mol = AllChem.ReplaceSubstructs(mol, dy_pattern, new_attachment)[0]\n\n#Good idea to clean it up\nChem.SanitizeMol(new_mol)\n\n#Since you want 3D mols later, I'd suggest adding hydrogens. Note: this takes up a lot more memory for the obj.\nChem.AddHs(new_mol)\n```\n\n\nThen convert this to whatever 3D representation you want.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2737658,
      "author_name": "andrewdblevins",
      "author_url": "",
      "post_date": "04/05/2024 21:25:46",
      "content": "<p>We use [Dy] as a stand-in for where the molecule is attached to the DNA tag</p>",
      "votes": null,
      "replies": [
        {
          "id": 2737665,
          "author_name": "srijitseal",
          "author_url": "",
          "post_date": "04/05/2024 21:31:48",
          "content": "<p>That makes sense! Thank you :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2737722,
          "author_name": "luispintoc",
          "author_url": "",
          "post_date": "04/05/2024 22:59:13",
          "content": "<p>Could you please elaborate on this? For someone with not much biology knowledge (I use smiles and so for polymers though). Thank you.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2738750,
              "author_name": "ianquigley",
              "author_url": "",
              "post_date": "04/06/2024 15:06:30",
              "content": "<p>We chose Dy (dysprosium) as a stand-in for the DNA because it's unlikely to be in druglike chemicals and started with the letter \"D\". Nothing more complicated than that!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2744877,
                  "author_name": "beardypolonium",
                  "author_url": "",
                  "post_date": "04/10/2024 07:13:49",
                  "content": "<p>Is is there a simple and fast way (just bbs SMILES strings, and no subgraph matching) to tell which bb is going to be attached to the DNA molecule?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2745095,
                      "author_name": "danielvik",
                      "author_url": "",
                      "post_date": "04/10/2024 11:22:59",
                      "content": "<p>I believe it will be the BB1 that is attached to the DNA. This is how I interpret the figure in the overview - but would be nice with a confirmation from the organizers.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2745384,
                          "author_name": "andrewdblevins",
                          "author_url": "",
                          "post_date": "04/10/2024 15:30:32",
                          "content": "<p>BB1 is always attached to the DNA</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2740776,
      "author_name": "arnaldoneto",
      "author_url": "",
      "post_date": "04/07/2024 23:26:44",
      "content": "<p>According to the competition's data tab, it means the following:</p>\n<blockquote>\n  <p>molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.</p>\n</blockquote>\n<p>Hope that helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2741323,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/08/2024 08:57:17",
      "content": "<p>if we remove DNA linker [Dy] from smiles in kaggle dataset, it should not affect results. Is that correct?  </p>\n<p>i ask this because public dataset (external dataset) are unlikely to include DNA linker.<br>\nI am thinking of how to use external dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2741810,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/08/2024 15:52:10",
          "content": "<p>Hi  <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>Removal of the [Dy] is a bit tricky. When the molecule binds to the protein, the DNA is still attached. Since DNA is a large molecule and there is a chain linking the molecule of interest to DNA, this can prohibit (or at least modify) the binding of the molecule to the protein.</p>\n<p>There are many ways you can experiment with handling this (convert the [Dy] to an alkyl group like Me or butyl, for example), but I'd suspect that keeping some information about where the DNA is bound will likely be valuable to the model.</p>\n<p>For context, medicinal chemists that use DEL hits as starting points often use the DNA binding location as a vector to add other chemical probes or as ways to modify the physical properties of the molecule (improve solubility, for example) without modifying the binding. That said, the DNA attachment point can still provide a contribution to the binding (I've seen this first hand).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2741969,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "04/08/2024 17:13:14",
              "content": "<p>Great explanation, thanks! This also got me thinking, can the molecule bind to the protein through a building block to which the DNA is attached?  Because if not, then the location of Dy could indicate at least, which of the three building blocks is not responsible for the binding. I thought so, but it seems that building blocks can indirectly affect a molecule's ability to bind, and there is no way to determine (using domain knowledge, I mean, not ML) which of them is more important, is it correct?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2742147,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "04/08/2024 19:03:24",
                  "content": "<p>Yes, a building block attached to the DNA can contribute to the binding and it often does. However, the closer you get to the linker attachment point, the more likely you are to be headed towards solvent (water) and not deeper into the protein.</p>\n<p>I hope that helps!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2747839,
                      "author_name": "sroger",
                      "author_url": "",
                      "post_date": "04/12/2024 05:18:42",
                      "content": "<p>Excuse my ignorance but what do you mean with the last part? (headed towards solvent not protein)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2750040,
                          "author_name": "chemdatafarmer",
                          "author_url": "",
                          "post_date": "04/13/2024 12:26:56",
                          "content": "<p>Molecules bind in 3-dimensions. Imagine the protein is a small bowl and the air around it is water. Now, imagine a tennis ball you have on the end of a stick. You can put the tennis ball in the bowl (imagine a snug fit) and the stick is pointing up. In this analogy, the stick is the DNA attachment linker and the molecule is the ball. The parts of the ball that fill the bowl and the parts that touch the bowl would contribute to binding and the parts that don't have little to no effect on binding.</p>\n<p>This is an oversimplification, but gets the idea across </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2750488,
                              "author_name": "sroger",
                              "author_url": "",
                              "post_date": "04/13/2024 17:27:46",
                              "content": "<p>Thank you for the intuitive explanation!</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2753995,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "04/15/2024 18:41:06",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a>, if trying to generate 3d molecules from smiles, and just wanting to simplify for now, any suggestion for one \"mostly good enough\" method?</p>\n<ul>\n<li>Remove [Dy]?</li>\n<li>Replace the text string \"[Dy]\" with the text string <em>__</em>?</li>\n</ul>",
              "votes": null,
              "replies": [
                {
                  "id": 2754223,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "04/15/2024 22:28:50",
                  "content": "<p>Hey <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> there are a couple ways to do this, but the simplest way is probably to replace the [Dy] with a methyl (CH3) group. I'll probably make a notebook on this once I get deeper into my own 3D workflows, but this should do for now.</p>\n<p>Probably the most robust way to do that is as follows.</p>\n<pre><code> rdkit  Chem\n rdkit.Chem  AllChem\n\n\nmol = Chem.MolFromSmiles(SMILES)\n\n\nnew_attachment = Chem.MolFromSmiles()\n\n\ndy_pattern = Chem.MolFromSmiles()\n\n\nnew_mol = AllChem.ReplaceSubstructs(mol, dy_pattern, new_attachment)[]\n\n\nChem.SanitizeMol(new_mol)\n\n\nChem.AddHs(new_mol)\n</code></pre>\n<p>Then convert this to whatever 3D representation you want.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2737655": "Can the organisers explain if this is actually intended? Every SMILES in the test dataset has a metal atom [Dy] in the SMILES representation. Is this actually intended?",
    "2737658": "We use [Dy] as a stand-in for where the molecule is attached to the DNA tag",
    "2737665": "That makes sense! Thank you :D",
    "2737722": "Could you please elaborate on this? For someone with not much biology knowledge (I use smiles and so for polymers though). Thank you.",
    "2738750": "We chose Dy (dysprosium) as a stand-in for the DNA because it's unlikely to be in druglike chemicals and started with the letter \"D\". Nothing more complicated than that!",
    "2740776": "According to the competition's data tab, it means the following:\n\n>molecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.\n\nHope that helps!",
    "2741323": "if we remove DNA linker [Dy] from smiles in kaggle dataset, it should not affect results. Is that correct?  \n\ni ask this because public dataset (external dataset) are unlikely to include DNA linker.\nI am thinking of how to use external dataset.",
    "2741810": "Hi  @hengck23 \n\nRemoval of the [Dy] is a bit tricky. When the molecule binds to the protein, the DNA is still attached. Since DNA is a large molecule and there is a chain linking the molecule of interest to DNA, this can prohibit (or at least modify) the binding of the molecule to the protein.\n\nThere are many ways you can experiment with handling this (convert the [Dy] to an alkyl group like Me or butyl, for example), but I'd suspect that keeping some information about where the DNA is bound will likely be valuable to the model.\n\nFor context, medicinal chemists that use DEL hits as starting points often use the DNA binding location as a vector to add other chemical probes or as ways to modify the physical properties of the molecule (improve solubility, for example) without modifying the binding. That said, the DNA attachment point can still provide a contribution to the binding (I've seen this first hand).",
    "2741969": "Great explanation, thanks! This also got me thinking, can the molecule bind to the protein through a building block to which the DNA is attached?  Because if not, then the location of Dy could indicate at least, which of the three building blocks is not responsible for the binding. I thought so, but it seems that building blocks can indirectly affect a molecule's ability to bind, and there is no way to determine (using domain knowledge, I mean, not ML) which of them is more important, is it correct?",
    "2742147": "Yes, a building block attached to the DNA can contribute to the binding and it often does. However, the closer you get to the linker attachment point, the more likely you are to be headed towards solvent (water) and not deeper into the protein.\n\nI hope that helps!",
    "2744877": "Is is there a simple and fast way (just bbs SMILES strings, and no subgraph matching) to tell which bb is going to be attached to the DNA molecule?",
    "2745095": "I believe it will be the BB1 that is attached to the DNA. This is how I interpret the figure in the overview - but would be nice with a confirmation from the organizers.",
    "2745384": "BB1 is always attached to the DNA",
    "2747839": "Excuse my ignorance but what do you mean with the last part? (headed towards solvent not protein)",
    "2750040": "Molecules bind in 3-dimensions. Imagine the protein is a small bowl and the air around it is water. Now, imagine a tennis ball you have on the end of a stick. You can put the tennis ball in the bowl (imagine a snug fit) and the stick is pointing up. In this analogy, the stick is the DNA attachment linker and the molecule is the ball. The parts of the ball that fill the bowl and the parts that touch the bowl would contribute to binding and the parts that don't have little to no effect on binding.\n\nThis is an oversimplification, but gets the idea across",
    "2750488": "Thank you for the intuitive explanation!",
    "2753995": "Hi @chemdatafarmer, if trying to generate 3d molecules from smiles, and just wanting to simplify for now, any suggestion for one \"mostly good enough\" method?\n\n- Remove [Dy]?\n- Replace the text string \"[Dy]\" with the text string ____?",
    "2754223": "Hey @roberthatch there are a couple ways to do this, but the simplest way is probably to replace the [Dy] with a methyl (CH3) group. I'll probably make a notebook on this once I get deeper into my own 3D workflows, but this should do for now.\n\nProbably the most robust way to do that is as follows.\n\n\n```python\nfrom rdkit import Chem\nfrom rdkit.Chem import AllChem\n\n#Convert your SMILES to a mol object.\nmol = Chem.MolFromSmiles(SMILES)\n\n#Create a mol object to replace the Dy atom with.\nnew_attachment = Chem.MolFromSmiles('C')\n\n#Get the pattern for the Dy atom\ndy_pattern = Chem.MolFromSmiles('[Dy]')\n\n#This returns a tuple of all possible replacements, but we know there will only be one.\nnew_mol = AllChem.ReplaceSubstructs(mol, dy_pattern, new_attachment)[0]\n\n#Good idea to clean it up\nChem.SanitizeMol(new_mol)\n\n#Since you want 3D mols later, I'd suggest adding hydrogens. Note: this takes up a lot more memory for the obj.\nChem.AddHs(new_mol)\n```\n\n\nThen convert this to whatever 3D representation you want."
  },
  "source": "meta"
}