{
  "id": 243917,
  "title": "24th place (using molecular graph and atom coordinates)",
  "url": "/competitions/bms-molecular-translation/writeups/stas-sl-24th-place-using-molecular-graph-and-atom-",
  "author_name": "",
  "post_date": "2021-06-04T14:03:07.543Z",
  "votes": 35,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Since reading DACON solutions, especially one that used object detection to detect atom and bonds coordinates and then rebuilding molecule, I was thinking what would be the best representation for molecule that can be learned by NN more naturally. </p>\n<p>While directly predicting inchi is the simplest solution as it requires minimal post-processing and it can result in lower LD even if predicted string is not correct InChI. But it has some disadvantages as well:</p>\n<ul>\n<li>it is long</li>\n<li>it has pretty complicated atom numbering (and is very sensitive to it)</li>\n</ul>\n<p>So I was quite surprised after training, that NN can actually learn very well to directly predict correct InChI with high accuracy. Even LD &lt; 2 - seemed to me very good result, considering not perfect image quality.</p>\n<p>While being able to predict <em>/c</em> layer very well, I observed that it struggles with identifying stereochemestry (/t layer), which seemed to be a simpler task. If it can predict whole sequence of atom connections, why it can't just predict + or - ? This was not the biggest source of error as if other layers predicted correctly, it contributed only 1-4 LD, but still it was unclear to me.</p>\n<p>So I was looking for others molecule representations to overcome mentioned disadvantages. One of them is SMILES that produces shorter strings, another one is molecular graph. I have chosen the latter. It was not exactly what was used in DACON 1st place solution, which was fully object detection approach, as it predicted both atom and bonds coord, and then identified which atoms are connected based on geometrical coordinates of atoms and bonds. Instead, in my representation I had to keep atom numbering (though I changed it from InChI numbering to more simpler based on 2d coordinated of atoms - from left to right, from top to bottom) and for each atom, besides its properties (like element type, chirality, isotope, xy-coords) I predicted also connections from this atom to other atoms - you can think of it as an adjacency matrix.</p>\n<p><img src=\"https://user-images.githubusercontent.com/4602302/120774226-c0dded00-c52a-11eb-9a62-0a997467792d.png\" alt=\"image\"></p>\n<p>To recover molecule from atom properties and bonds, I used something like this. It is not full code, but just simple example how you can manually build molecule from scratch. Important methods here to correctly detect stereochemestry are <code>DetectBondStereochemistry</code> and <code>AssignChiralTypesFromBondDirs</code> - they are actually responsible for reversing the procedure that is happening during rendering, when coordinates of molecules are calculated, instead these methods assign E/Z and cis/trans stereo based on 2d coordinated (x, y) of atoms predicted by model.</p>\n<pre><code>m = Chem.RWMol()\n\n# nodes\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('F'))\nm.AddAtom(Chem.Atom('O'))\nm.AddAtom(Chem.Atom('Cl'))\n\n# edges\nm.AddBond(0, 1, Chem.BondType.SINGLE)\nm.AddBond(1, 2, Chem.BondType.DOUBLE)\nm.AddBond(2, 3, Chem.BondType.SINGLE)\nm.AddBond(3, 4, Chem.BondType.SINGLE)\nm.AddBond(3, 5, Chem.BondType.SINGLE)\nm.AddBond(3, 6, Chem.BondType.SINGLE)\nm.GetBondBetweenAtoms(3, 5).SetBondDir(Chem.BondDir.BEGINWEDGE)\n\n# coords\ncoords = ((-2, -1), (-1, 0), (1, 0), (2, 1), (2, 0), (3.5, 1), (2, 2))\nconf = Chem.Conformer(m.GetNumAtoms())\nconf.Set3D(False)\nfor i, (x, y) in enumerate(coords):\n    conf.SetAtomPosition(i, (x, y, 0))\nm.AddConformer(conf)\n\n# magic\nChem.SanitizeMol(m)\nChem.DetectBondStereochemistry(m)\nChem.AssignChiralTypesFromBondDirs(m)\nChem.AssignStereochemistry(m)\n\nm.Debug()\nopts = Draw.MolDrawOptions()\nopts.addAtomIndices = True\nopts.addStereoAnnotation = True\nDraw.MolToImage(m, options=opts)\n</code></pre>\n<p><img src=\"https://user-images.githubusercontent.com/4602302/115084791-06f6d700-9f12-11eb-8e07-0aaa5c739b61.png\" alt=\"image\"></p>\n<p>In parallel I was looking for ways to extend training data with other images that are as close to the original training images as possible. While rdkit was very popular and useful in this competition, its rendering of molecules was not exactly the same. So I was exploring other tools, and finally I managed to find the tool that produced the exact images (except noise). It was <a href=\"https://lifescience.opensource.epam.com/indigo/api/index.html\" target=\"_blank\">epam.indigo</a>. I was very happy, when I found it. Firstly, it made possible to use 10M molecules from the extra dataset, secondly I could use it calculate coordinates of atoms in all existing molecules.</p>\n<p>While I was able to simplify atom numbering scheme in my representation, that (hopefully) is easier for NN to learn, I still was not very satisfied with it. Ideally I'd like to get rid of any order of atoms. But I was not able to achieve that.</p>\n<p>Original image:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/120806284-58553700-c54f-11eb-83b5-51103c5c0edc.png\" alt=\"image\"></p>\n<p>Rendered by epam.indigo with atom numbers:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/120806360-6f942480-c54f-11eb-853e-d47f87e44dcd.png\" alt=\"image\"></p>\n<p>Rendered by epam.indigo with 2D (x, y) atom numbering:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/120806428-820e5e00-c54f-11eb-922b-c24eab3c3865.png\" alt=\"image\"></p>\n<p>What I was able to achieve with this representation is high accuracy of detecting stereo information. Going from almost 50% accuracy (for each stereocenter) when predicting inchi string directly to almost 100% accuracy when recovering it by rdkit from 2d coordinates. So I was able to achieve 95% exact inchi (LD=0) predictions on validation set, while predicting inchi as string I had somewhere below 90%, with ~5% mismatches only in 1-4 characters in stereolayer.</p>\n<p>One major disadvantage of this approach is that sometimes it can predict broken molecules with some bonds missing in the middle. It happened for &lt;1% of molecules, but then it produced very large LD, so in order to overcome this, I had to still use other model that predicted inchi as string. </p>\n<p>So my final submission (public LB = 0.88) was an ensemble of 2 models:</p>\n<ul>\n<li>model1: 384x384-&gt;16x16x256, effnet_b3 encoder, 6 layer transformer decoder (dim 256), inchi as string</li>\n<li>model2: 512x512-&gt;16x16x256, effnet_b0 encoder, 6 layer transformer decoder (dim 256), molecular graph representation</li>\n</ul>\n<p>In my submission only 4% of predicted inchis were from model1, others 96% were from pretty slim model2 (4M parameters in encoder, 5M in decoder). For all my training I was using only single Google Colab notebook. I didn't use beam search while inference.</p>",
  "messages": [
    {
      "id": "1335835",
      "postDate": "06/04/2021 13:14:49",
      "content": "<p>Since reading DACON solutions, especially one that used object detection to detect atom and bonds coordinates and then rebuilding molecule, I was thinking what would be the best representation for molecule that can be learned by NN more naturally. </p>\n<p>While directly predicting inchi is the simplest solution as it requires minimal post-processing and it can result in lower LD even if predicted string is not correct InChI. But it has some disadvantages as well:</p>\n<ul>\n<li>it is long</li>\n<li>it has pretty complicated atom numbering (and is very sensitive to it)</li>\n</ul>\n<p>So I was quite surprised after training, that NN can actually learn very well to directly predict correct InChI with high accuracy. Even LD &lt; 2 - seemed to me very good result, considering not perfect image quality.</p>\n<p>While being able to predict <em>/c</em> layer very well, I observed that it struggles with identifying stereochemestry (/t layer), which seemed to be a simpler task. If it can predict whole sequence of atom connections, why it can't just predict + or - ? This was not the biggest source of error as if other layers predicted correctly, it contributed only 1-4 LD, but still it was unclear to me.</p>\n<p>So I was looking for others molecule representations to overcome mentioned disadvantages. One of them is SMILES that produces shorter strings, another one is molecular graph. I have chosen the latter. It was not exactly what was used in DACON 1st place solution, which was fully object detection approach, as it predicted both atom and bonds coord, and then identified which atoms are connected based on geometrical coordinates of atoms and bonds. Instead, in my representation I had to keep atom numbering (though I changed it from InChI numbering to more simpler based on 2d coordinated of atoms - from left to right, from top to bottom) and for each atom, besides its properties (like element type, chirality, isotope, xy-coords) I predicted also connections from this atom to other atoms - you can think of it as an adjacency matrix.</p>\n<p><img src=\"https://user-images.githubusercontent.com/4602302/120774226-c0dded00-c52a-11eb-9a62-0a997467792d.png\" alt=\"image\"></p>\n<p>To recover molecule from atom properties and bonds, I used something like this. It is not full code, but just simple example how you can manually build molecule from scratch. Important methods here to correctly detect stereochemestry are <code>DetectBondStereochemistry</code> and <code>AssignChiralTypesFromBondDirs</code> - they are actually responsible for reversing the procedure that is happening during rendering, when coordinates of molecules are calculated, instead these methods assign E/Z and cis/trans stereo based on 2d coordinated (x, y) of atoms predicted by model.</p>\n<pre><code>m = Chem.RWMol()\n\n# nodes\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('F'))\nm.AddAtom(Chem.Atom('O'))\nm.AddAtom(Chem.Atom('Cl'))\n\n# edges\nm.AddBond(0, 1, Chem.BondType.SINGLE)\nm.AddBond(1, 2, Chem.BondType.DOUBLE)\nm.AddBond(2, 3, Chem.BondType.SINGLE)\nm.AddBond(3, 4, Chem.BondType.SINGLE)\nm.AddBond(3, 5, Chem.BondType.SINGLE)\nm.AddBond(3, 6, Chem.BondType.SINGLE)\nm.GetBondBetweenAtoms(3, 5).SetBondDir(Chem.BondDir.BEGINWEDGE)\n\n# coords\ncoords = ((-2, -1), (-1, 0), (1, 0), (2, 1), (2, 0), (3.5, 1), (2, 2))\nconf = Chem.Conformer(m.GetNumAtoms())\nconf.Set3D(False)\nfor i, (x, y) in enumerate(coords):\n    conf.SetAtomPosition(i, (x, y, 0))\nm.AddConformer(conf)\n\n# magic\nChem.SanitizeMol(m)\nChem.DetectBondStereochemistry(m)\nChem.AssignChiralTypesFromBondDirs(m)\nChem.AssignStereochemistry(m)\n\nm.Debug()\nopts = Draw.MolDrawOptions()\nopts.addAtomIndices = True\nopts.addStereoAnnotation = True\nDraw.MolToImage(m, options=opts)\n</code></pre>\n<p><img src=\"https://user-images.githubusercontent.com/4602302/115084791-06f6d700-9f12-11eb-8e07-0aaa5c739b61.png\" alt=\"image\"></p>\n<p>In parallel I was looking for ways to extend training data with other images that are as close to the original training images as possible. While rdkit was very popular and useful in this competition, its rendering of molecules was not exactly the same. So I was exploring other tools, and finally I managed to find the tool that produced the exact images (except noise). It was <a href=\"https://lifescience.opensource.epam.com/indigo/api/index.html\" target=\"_blank\">epam.indigo</a>. I was very happy, when I found it. Firstly, it made possible to use 10M molecules from the extra dataset, secondly I could use it calculate coordinates of atoms in all existing molecules.</p>\n<p>While I was able to simplify atom numbering scheme in my representation, that (hopefully) is easier for NN to learn, I still was not very satisfied with it. Ideally I'd like to get rid of any order of atoms. But I was not able to achieve that.</p>\n<p>Original image:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/120806284-58553700-c54f-11eb-83b5-51103c5c0edc.png\" alt=\"image\"></p>\n<p>Rendered by epam.indigo with atom numbers:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/120806360-6f942480-c54f-11eb-853e-d47f87e44dcd.png\" alt=\"image\"></p>\n<p>Rendered by epam.indigo with 2D (x, y) atom numbering:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/120806428-820e5e00-c54f-11eb-922b-c24eab3c3865.png\" alt=\"image\"></p>\n<p>What I was able to achieve with this representation is high accuracy of detecting stereo information. Going from almost 50% accuracy (for each stereocenter) when predicting inchi string directly to almost 100% accuracy when recovering it by rdkit from 2d coordinates. So I was able to achieve 95% exact inchi (LD=0) predictions on validation set, while predicting inchi as string I had somewhere below 90%, with ~5% mismatches only in 1-4 characters in stereolayer.</p>\n<p>One major disadvantage of this approach is that sometimes it can predict broken molecules with some bonds missing in the middle. It happened for &lt;1% of molecules, but then it produced very large LD, so in order to overcome this, I had to still use other model that predicted inchi as string. </p>\n<p>So my final submission (public LB = 0.88) was an ensemble of 2 models:</p>\n<ul>\n<li>model1: 384x384-&gt;16x16x256, effnet_b3 encoder, 6 layer transformer decoder (dim 256), inchi as string</li>\n<li>model2: 512x512-&gt;16x16x256, effnet_b0 encoder, 6 layer transformer decoder (dim 256), molecular graph representation</li>\n</ul>\n<p>In my submission only 4% of predicted inchis were from model1, others 96% were from pretty slim model2 (4M parameters in encoder, 5M in decoder). For all my training I was using only single Google Colab notebook. I didn't use beam search while inference.</p>",
      "rawMarkdown": "Since reading DACON solutions, especially one that used object detection to detect atom and bonds coordinates and then rebuilding molecule, I was thinking what would be the best representation for molecule that can be learned by NN more naturally. \n\nWhile directly predicting inchi is the simplest solution as it requires minimal post-processing and it can result in lower LD even if predicted string is not correct InChI. But it has some disadvantages as well:\n- it is long\n- it has pretty complicated atom numbering (and is very sensitive to it)\n\nSo I was quite surprised after training, that NN can actually learn very well to directly predict correct InChI with high accuracy. Even LD < 2 - seemed to me very good result, considering not perfect image quality.\n\nWhile being able to predict */c* layer very well, I observed that it struggles with identifying stereochemestry (/t layer), which seemed to be a simpler task. If it can predict whole sequence of atom connections, why it can't just predict + or - ? This was not the biggest source of error as if other layers predicted correctly, it contributed only 1-4 LD, but still it was unclear to me.\n\nSo I was looking for others molecule representations to overcome mentioned disadvantages. One of them is SMILES that produces shorter strings, another one is molecular graph. I have chosen the latter. It was not exactly what was used in DACON 1st place solution, which was fully object detection approach, as it predicted both atom and bonds coord, and then identified which atoms are connected based on geometrical coordinates of atoms and bonds. Instead, in my representation I had to keep atom numbering (though I changed it from InChI numbering to more simpler based on 2d coordinated of atoms - from left to right, from top to bottom) and for each atom, besides its properties (like element type, chirality, isotope, xy-coords) I predicted also connections from this atom to other atoms - you can think of it as an adjacency matrix.\n\n![image](https://user-images.githubusercontent.com/4602302/120774226-c0dded00-c52a-11eb-9a62-0a997467792d.png)\n\nTo recover molecule from atom properties and bonds, I used something like this. It is not full code, but just simple example how you can manually build molecule from scratch. Important methods here to correctly detect stereochemestry are `DetectBondStereochemistry` and `AssignChiralTypesFromBondDirs` - they are actually responsible for reversing the procedure that is happening during rendering, when coordinates of molecules are calculated, instead these methods assign E/Z and cis/trans stereo based on 2d coordinated (x, y) of atoms predicted by model.\n\n```\nm = Chem.RWMol()\n\n# nodes\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('F'))\nm.AddAtom(Chem.Atom('O'))\nm.AddAtom(Chem.Atom('Cl'))\n\n# edges\nm.AddBond(0, 1, Chem.BondType.SINGLE)\nm.AddBond(1, 2, Chem.BondType.DOUBLE)\nm.AddBond(2, 3, Chem.BondType.SINGLE)\nm.AddBond(3, 4, Chem.BondType.SINGLE)\nm.AddBond(3, 5, Chem.BondType.SINGLE)\nm.AddBond(3, 6, Chem.BondType.SINGLE)\nm.GetBondBetweenAtoms(3, 5).SetBondDir(Chem.BondDir.BEGINWEDGE)\n\n# coords\ncoords = ((-2, -1), (-1, 0), (1, 0), (2, 1), (2, 0), (3.5, 1), (2, 2))\nconf = Chem.Conformer(m.GetNumAtoms())\nconf.Set3D(False)\nfor i, (x, y) in enumerate(coords):\n    conf.SetAtomPosition(i, (x, y, 0))\nm.AddConformer(conf)\n\n# magic\nChem.SanitizeMol(m)\nChem.DetectBondStereochemistry(m)\nChem.AssignChiralTypesFromBondDirs(m)\nChem.AssignStereochemistry(m)\n\nm.Debug()\nopts = Draw.MolDrawOptions()\nopts.addAtomIndices = True\nopts.addStereoAnnotation = True\nDraw.MolToImage(m, options=opts)\n```\n\n![image](https://user-images.githubusercontent.com/4602302/115084791-06f6d700-9f12-11eb-8e07-0aaa5c739b61.png)\n\nIn parallel I was looking for ways to extend training data with other images that are as close to the original training images as possible. While rdkit was very popular and useful in this competition, its rendering of molecules was not exactly the same. So I was exploring other tools, and finally I managed to find the tool that produced the exact images (except noise). It was [epam.indigo](https://lifescience.opensource.epam.com/indigo/api/index.html). I was very happy, when I found it. Firstly, it made possible to use 10M molecules from the extra dataset, secondly I could use it calculate coordinates of atoms in all existing molecules.\n\nWhile I was able to simplify atom numbering scheme in my representation, that (hopefully) is easier for NN to learn, I still was not very satisfied with it. Ideally I'd like to get rid of any order of atoms. But I was not able to achieve that.\n\nOriginal image:\n![image](https://user-images.githubusercontent.com/4602302/120806284-58553700-c54f-11eb-83b5-51103c5c0edc.png)\n\nRendered by epam.indigo with atom numbers:\n![image](https://user-images.githubusercontent.com/4602302/120806360-6f942480-c54f-11eb-853e-d47f87e44dcd.png)\n\nRendered by epam.indigo with 2D (x, y) atom numbering:\n![image](https://user-images.githubusercontent.com/4602302/120806428-820e5e00-c54f-11eb-922b-c24eab3c3865.png)\n\nWhat I was able to achieve with this representation is high accuracy of detecting stereo information. Going from almost 50% accuracy (for each stereocenter) when predicting inchi string directly to almost 100% accuracy when recovering it by rdkit from 2d coordinates. So I was able to achieve 95% exact inchi (LD=0) predictions on validation set, while predicting inchi as string I had somewhere below 90%, with ~5% mismatches only in 1-4 characters in stereolayer.\n\nOne major disadvantage of this approach is that sometimes it can predict broken molecules with some bonds missing in the middle. It happened for <1% of molecules, but then it produced very large LD, so in order to overcome this, I had to still use other model that predicted inchi as string. \n\nSo my final submission (public LB = 0.88) was an ensemble of 2 models:\n- model1: 384x384->16x16x256, effnet_b3 encoder, 6 layer transformer decoder (dim 256), inchi as string\n- model2: 512x512->16x16x256, effnet_b0 encoder, 6 layer transformer decoder (dim 256), molecular graph representation\n\nIn my submission only 4% of predicted inchis were from model1, others 96% were from pretty slim model2 (4M parameters in encoder, 5M in decoder). For all my training I was using only single Google Colab notebook. I didn't use beam search while inference.",
      "votes": null
    },
    {
      "id": "1335857",
      "postDate": "06/04/2021 13:32:17",
      "content": "<p>Amazing work! Congrats!</p>",
      "rawMarkdown": "Amazing work! Congrats!",
      "votes": null
    },
    {
      "id": "1336388",
      "postDate": "06/04/2021 21:02:53",
      "content": "<p>Wonderfully inventive approach. Grreat job!</p>",
      "rawMarkdown": "Wonderfully inventive approach. Grreat job!",
      "votes": null
    },
    {
      "id": "1336465",
      "postDate": "06/05/2021 00:14:33",
      "content": "<p>Great work, I'm surprised we didn't see more people following this path! </p>",
      "rawMarkdown": "Great work, I'm surprised we didn't see more people following this path!",
      "votes": null
    },
    {
      "id": "1337700",
      "postDate": "06/05/2021 19:51:55",
      "content": "<p>Great work, the InChI modeling looks quite creative.</p>",
      "rawMarkdown": "Great work, the InChI modeling looks quite creative.",
      "votes": null
    },
    {
      "id": "1337766",
      "postDate": "06/05/2021 21:38:16",
      "content": "<p>This is beautiful!</p>",
      "rawMarkdown": "This is beautiful!",
      "votes": null
    },
    {
      "id": "1340292",
      "postDate": "06/07/2021 18:38:52",
      "content": "<p>After I wrote this post, I started to think, whether it actually helps to renumber atoms based on 2D coordinates. Intuitively, from a human perspective, it seems reasonable - you just look at image and assign numbers from left to right, from top to bottom, even for rather large molecules it shouldn't be very difficult. While algorithm described in InChI specification requires quite complicated calculations. But intuition might be misleading sometimes - some tasks that are easy for humans might be hard for NN and vice versa.</p>\n<p>To validate this empirically, I've conducted a quick experiment training for one epoch the same model with same parameters - the only difference was in a few lines responsible for atom reordering.</p>\n<p><img src=\"https://user-images.githubusercontent.com/4602302/121061194-374d4a00-c7cc-11eb-87db-47dc6367b7de.png\" alt=\"image\"></p>\n<p>First plot is percentage of exact InChI matches, second - levenshtein distance. You can see that, indeed, simplified atom numbering (🔴 red curve) makes the task much easier than when atoms are numbered according to the InChI spec (🔵 blue curve), and training goes much faster.</p>\n<p>There is no intrinsic order in a molecule (nor in image) - it is just a graph with unordered sets of nodes and edges. So when we are forcing NN to learn some non-trivial order of atoms, we make the task harder than it is required. RDKit is able to calculate atom numbers according to InChI spec perfectly, so why not to outsource this task to it and make thing simpler for NN.</p>\n<p>Another experiment is comparing performance of models one that outputs InChI as string directly (🟢 green curve) and another outputs mol graph (🔴 red curve). The result is ambiguous. From one side, if comparing percentage of exact matches, then model based on mol graph beats model based on InChI string with large margin (75% vs 35% after one epoch), even without simplified atom numbering (🔵 blue curve) - (55% vs 35%). From another side, model that outputs InChI string directly has lower LD, even if doesn't predict the result exactly. So in order to achieve better score, you should probably combine both.</p>",
      "rawMarkdown": "After I wrote this post, I started to think, whether it actually helps to renumber atoms based on 2D coordinates. Intuitively, from a human perspective, it seems reasonable - you just look at image and assign numbers from left to right, from top to bottom, even for rather large molecules it shouldn't be very difficult. While algorithm described in InChI specification requires quite complicated calculations. But intuition might be misleading sometimes - some tasks that are easy for humans might be hard for NN and vice versa.\n\nTo validate this empirically, I've conducted a quick experiment training for one epoch the same model with same parameters - the only difference was in a few lines responsible for atom reordering.\n\n![image](https://user-images.githubusercontent.com/4602302/121061194-374d4a00-c7cc-11eb-87db-47dc6367b7de.png)\n\nFirst plot is percentage of exact InChI matches, second - levenshtein distance. You can see that, indeed, simplified atom numbering (🔴 red curve) makes the task much easier than when atoms are numbered according to the InChI spec (🔵 blue curve), and training goes much faster.\n\nThere is no intrinsic order in a molecule (nor in image) - it is just a graph with unordered sets of nodes and edges. So when we are forcing NN to learn some non-trivial order of atoms, we make the task harder than it is required. RDKit is able to calculate atom numbers according to InChI spec perfectly, so why not to outsource this task to it and make thing simpler for NN.\n\nAnother experiment is comparing performance of models one that outputs InChI as string directly (🟢 green curve) and another outputs mol graph (🔴 red curve). The result is ambiguous. From one side, if comparing percentage of exact matches, then model based on mol graph beats model based on InChI string with large margin (75% vs 35% after one epoch), even without simplified atom numbering (🔵 blue curve) - (55% vs 35%). From another side, model that outputs InChI string directly has lower LD, even if doesn't predict the result exactly. So in order to achieve better score, you should probably combine both.",
      "votes": null
    },
    {
      "id": "1517498",
      "postDate": "09/19/2021 17:26:16",
      "content": "<p>wow, this is an awesome solution! Congratulations on the silver medal.<br>\nCould you please share the epam.indigo code you used to create the 2D (x, y) atom numbering image?</p>",
      "rawMarkdown": "wow, this is an awesome solution! Congratulations on the silver medal.\nCould you please share the epam.indigo code you used to create the 2D (x, y) atom numbering image?",
      "votes": null
    },
    {
      "id": "1517528",
      "postDate": "09/19/2021 18:14:57",
      "content": "<p>Thanks )</p>\n<p>It is not easy to navigate my code after several month 😜, but I believe this should be correct version:</p>\n<pre><code>mol_indigo = indigo.loadMolecule(inchi)\n\n# ask indigo to calculate 2d coords\nmol_indigo.layout()\n\n# serialize mol to molfile with 2d coords assigned and then load it with rdkit,\n# it is simplest way to pass coords from indigo to rdkit\nmol = Chem.MolFromMolBlock(mol_indigo.molfile(), removeHs=False, sanitize=False)\nChem.Kekulize(mol)\n\n# next few lines are responsible for correct stereo-bonds assignment based on 2d coords\n# so if you are interested in renumbering only, you can ignore them\nfor b in mol.GetBonds():\n    if b.GetBondDir() in (Chem.BondDir.BEGINDASH, Chem.BondDir.BEGINWEDGE):\n        b.SetBondDir(Chem.BondDir.NONE)\n\nconf = mol.GetConformer()\nChem.WedgeMolBonds(mol, mol.GetConformer())\n\ncoord_x = []\ncoord_y = []\nfor i in range(mol.GetNumAtoms()):\n    xyz = conf.GetAtomPosition(i)\n    coord_x.append(xyz[0])\n    coord_y.append(xyz[1])\ncoord_x = np.array(coord_x)\ncoord_y = np.array(coord_y)\ncoords = np.array([coord_x, coord_y]).T\n\nCOORDS_SCALE = 10\ncoords = (2 * (coords - coords.min(0)) / (\n        coords.max(0) - coords.min(0)) * COORDS_SCALE - COORDS_SCALE)\ncoords[np.isnan(coords)] = 0\norder = np.lexsort((-coord_y, coord_x))\nmol = Chem.RenumberAtoms(mol, order.tolist())\n\n# now atoms will be iterated in (x, y) order\nfor a in mol.GetAtoms():\n    # ...\n</code></pre>\n<p>But maybe if you don't use rdkit, or you don't need anything except  atom 2d coordinates, you can use epam.indigo directly to get coords and then compute atom order based on them.</p>",
      "rawMarkdown": "Thanks )\n\nIt is not easy to navigate my code after several month 😜, but I believe this should be correct version:\n```\nmol_indigo = indigo.loadMolecule(inchi)\n\n# ask indigo to calculate 2d coords\nmol_indigo.layout()\n\n# serialize mol to molfile with 2d coords assigned and then load it with rdkit,\n# it is simplest way to pass coords from indigo to rdkit\nmol = Chem.MolFromMolBlock(mol_indigo.molfile(), removeHs=False, sanitize=False)\nChem.Kekulize(mol)\n\n# next few lines are responsible for correct stereo-bonds assignment based on 2d coords\n# so if you are interested in renumbering only, you can ignore them\nfor b in mol.GetBonds():\n    if b.GetBondDir() in (Chem.BondDir.BEGINDASH, Chem.BondDir.BEGINWEDGE):\n        b.SetBondDir(Chem.BondDir.NONE)\n\nconf = mol.GetConformer()\nChem.WedgeMolBonds(mol, mol.GetConformer())\n\ncoord_x = []\ncoord_y = []\nfor i in range(mol.GetNumAtoms()):\n    xyz = conf.GetAtomPosition(i)\n    coord_x.append(xyz[0])\n    coord_y.append(xyz[1])\ncoord_x = np.array(coord_x)\ncoord_y = np.array(coord_y)\ncoords = np.array([coord_x, coord_y]).T\n\nCOORDS_SCALE = 10\ncoords = (2 * (coords - coords.min(0)) / (\n        coords.max(0) - coords.min(0)) * COORDS_SCALE - COORDS_SCALE)\ncoords[np.isnan(coords)] = 0\norder = np.lexsort((-coord_y, coord_x))\nmol = Chem.RenumberAtoms(mol, order.tolist())\n\n# now atoms will be iterated in (x, y) order\nfor a in mol.GetAtoms():\n    # ...\n```\n\nBut maybe if you don't use rdkit, or you don't need anything except  atom 2d coordinates, you can use epam.indigo directly to get coords and then compute atom order based on them.",
      "votes": null
    },
    {
      "id": "1518059",
      "postDate": "09/20/2021 11:13:36",
      "content": "<p>Thanks for the very kind reply. When I followed the code above, I confirmed that it was sorted. However, am I correct that the red arrow isn't explicitly declared in the image?</p>",
      "rawMarkdown": "Thanks for the very kind reply. When I followed the code above, I confirmed that it was sorted. However, am I correct that the red arrow isn't explicitly declared in the image?",
      "votes": null
    },
    {
      "id": "1518070",
      "postDate": "09/20/2021 11:22:29",
      "content": "<p>Yes, you are correct, I just draw it manually for demonstration, images are fed into network without any modification, the order is just used to sort network outputs.</p>",
      "rawMarkdown": "Yes, you are correct, I just draw it manually for demonstration, images are fed into network without any modification, the order is just used to sort network outputs.",
      "votes": null
    },
    {
      "id": "1518072",
      "postDate": "09/20/2021 11:28:22",
      "content": "<p>Thanks to the quick answer, I understood. Thank you so much. Good luck!</p>",
      "rawMarkdown": "Thanks to the quick answer, I understood. Thank you so much. Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335857,
      "author_name": "houndcl",
      "author_url": "",
      "post_date": "06/04/2021 13:32:17",
      "content": "<p>Amazing work! Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336388,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "06/04/2021 21:02:53",
      "content": "<p>Wonderfully inventive approach. Grreat job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336465,
      "author_name": "talktocharles",
      "author_url": "",
      "post_date": "06/05/2021 00:14:33",
      "content": "<p>Great work, I'm surprised we didn't see more people following this path! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1337700,
      "author_name": "rguo97",
      "author_url": "",
      "post_date": "06/05/2021 19:51:55",
      "content": "<p>Great work, the InChI modeling looks quite creative.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1337766,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "06/05/2021 21:38:16",
      "content": "<p>This is beautiful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1340292,
      "author_name": "stassl",
      "author_url": "",
      "post_date": "06/07/2021 18:38:52",
      "content": "<p>After I wrote this post, I started to think, whether it actually helps to renumber atoms based on 2D coordinates. Intuitively, from a human perspective, it seems reasonable - you just look at image and assign numbers from left to right, from top to bottom, even for rather large molecules it shouldn't be very difficult. While algorithm described in InChI specification requires quite complicated calculations. But intuition might be misleading sometimes - some tasks that are easy for humans might be hard for NN and vice versa.</p>\n<p>To validate this empirically, I've conducted a quick experiment training for one epoch the same model with same parameters - the only difference was in a few lines responsible for atom reordering.</p>\n<p><img src=\"https://user-images.githubusercontent.com/4602302/121061194-374d4a00-c7cc-11eb-87db-47dc6367b7de.png\" alt=\"image\"></p>\n<p>First plot is percentage of exact InChI matches, second - levenshtein distance. You can see that, indeed, simplified atom numbering (🔴 red curve) makes the task much easier than when atoms are numbered according to the InChI spec (🔵 blue curve), and training goes much faster.</p>\n<p>There is no intrinsic order in a molecule (nor in image) - it is just a graph with unordered sets of nodes and edges. So when we are forcing NN to learn some non-trivial order of atoms, we make the task harder than it is required. RDKit is able to calculate atom numbers according to InChI spec perfectly, so why not to outsource this task to it and make thing simpler for NN.</p>\n<p>Another experiment is comparing performance of models one that outputs InChI as string directly (🟢 green curve) and another outputs mol graph (🔴 red curve). The result is ambiguous. From one side, if comparing percentage of exact matches, then model based on mol graph beats model based on InChI string with large margin (75% vs 35% after one epoch), even without simplified atom numbering (🔵 blue curve) - (55% vs 35%). From another side, model that outputs InChI string directly has lower LD, even if doesn't predict the result exactly. So in order to achieve better score, you should probably combine both.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1517498,
      "author_name": "wooseokshin",
      "author_url": "",
      "post_date": "09/19/2021 17:26:16",
      "content": "<p>wow, this is an awesome solution! Congratulations on the silver medal.<br>\nCould you please share the epam.indigo code you used to create the 2D (x, y) atom numbering image?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1517528,
          "author_name": "stassl",
          "author_url": "",
          "post_date": "09/19/2021 18:14:57",
          "content": "<p>Thanks )</p>\n<p>It is not easy to navigate my code after several month 😜, but I believe this should be correct version:</p>\n<pre><code>mol_indigo = indigo.loadMolecule(inchi)\n\n# ask indigo to calculate 2d coords\nmol_indigo.layout()\n\n# serialize mol to molfile with 2d coords assigned and then load it with rdkit,\n# it is simplest way to pass coords from indigo to rdkit\nmol = Chem.MolFromMolBlock(mol_indigo.molfile(), removeHs=False, sanitize=False)\nChem.Kekulize(mol)\n\n# next few lines are responsible for correct stereo-bonds assignment based on 2d coords\n# so if you are interested in renumbering only, you can ignore them\nfor b in mol.GetBonds():\n    if b.GetBondDir() in (Chem.BondDir.BEGINDASH, Chem.BondDir.BEGINWEDGE):\n        b.SetBondDir(Chem.BondDir.NONE)\n\nconf = mol.GetConformer()\nChem.WedgeMolBonds(mol, mol.GetConformer())\n\ncoord_x = []\ncoord_y = []\nfor i in range(mol.GetNumAtoms()):\n    xyz = conf.GetAtomPosition(i)\n    coord_x.append(xyz[0])\n    coord_y.append(xyz[1])\ncoord_x = np.array(coord_x)\ncoord_y = np.array(coord_y)\ncoords = np.array([coord_x, coord_y]).T\n\nCOORDS_SCALE = 10\ncoords = (2 * (coords - coords.min(0)) / (\n        coords.max(0) - coords.min(0)) * COORDS_SCALE - COORDS_SCALE)\ncoords[np.isnan(coords)] = 0\norder = np.lexsort((-coord_y, coord_x))\nmol = Chem.RenumberAtoms(mol, order.tolist())\n\n# now atoms will be iterated in (x, y) order\nfor a in mol.GetAtoms():\n    # ...\n</code></pre>\n<p>But maybe if you don't use rdkit, or you don't need anything except  atom 2d coordinates, you can use epam.indigo directly to get coords and then compute atom order based on them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1518059,
          "author_name": "wooseokshin",
          "author_url": "",
          "post_date": "09/20/2021 11:13:36",
          "content": "<p>Thanks for the very kind reply. When I followed the code above, I confirmed that it was sorted. However, am I correct that the red arrow isn't explicitly declared in the image?</p>",
          "votes": null,
          "replies": [
            {
              "id": 1518070,
              "author_name": "stassl",
              "author_url": "",
              "post_date": "09/20/2021 11:22:29",
              "content": "<p>Yes, you are correct, I just draw it manually for demonstration, images are fed into network without any modification, the order is just used to sort network outputs.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1518072,
          "author_name": "wooseokshin",
          "author_url": "",
          "post_date": "09/20/2021 11:28:22",
          "content": "<p>Thanks to the quick answer, I understood. Thank you so much. Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1335835": "Since reading DACON solutions, especially one that used object detection to detect atom and bonds coordinates and then rebuilding molecule, I was thinking what would be the best representation for molecule that can be learned by NN more naturally. \n\nWhile directly predicting inchi is the simplest solution as it requires minimal post-processing and it can result in lower LD even if predicted string is not correct InChI. But it has some disadvantages as well:\n- it is long\n- it has pretty complicated atom numbering (and is very sensitive to it)\n\nSo I was quite surprised after training, that NN can actually learn very well to directly predict correct InChI with high accuracy. Even LD < 2 - seemed to me very good result, considering not perfect image quality.\n\nWhile being able to predict */c* layer very well, I observed that it struggles with identifying stereochemestry (/t layer), which seemed to be a simpler task. If it can predict whole sequence of atom connections, why it can't just predict + or - ? This was not the biggest source of error as if other layers predicted correctly, it contributed only 1-4 LD, but still it was unclear to me.\n\nSo I was looking for others molecule representations to overcome mentioned disadvantages. One of them is SMILES that produces shorter strings, another one is molecular graph. I have chosen the latter. It was not exactly what was used in DACON 1st place solution, which was fully object detection approach, as it predicted both atom and bonds coord, and then identified which atoms are connected based on geometrical coordinates of atoms and bonds. Instead, in my representation I had to keep atom numbering (though I changed it from InChI numbering to more simpler based on 2d coordinated of atoms - from left to right, from top to bottom) and for each atom, besides its properties (like element type, chirality, isotope, xy-coords) I predicted also connections from this atom to other atoms - you can think of it as an adjacency matrix.\n\n![image](https://user-images.githubusercontent.com/4602302/120774226-c0dded00-c52a-11eb-9a62-0a997467792d.png)\n\nTo recover molecule from atom properties and bonds, I used something like this. It is not full code, but just simple example how you can manually build molecule from scratch. Important methods here to correctly detect stereochemestry are `DetectBondStereochemistry` and `AssignChiralTypesFromBondDirs` - they are actually responsible for reversing the procedure that is happening during rendering, when coordinates of molecules are calculated, instead these methods assign E/Z and cis/trans stereo based on 2d coordinated (x, y) of atoms predicted by model.\n\n```\nm = Chem.RWMol()\n\n# nodes\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('C'))\nm.AddAtom(Chem.Atom('F'))\nm.AddAtom(Chem.Atom('O'))\nm.AddAtom(Chem.Atom('Cl'))\n\n# edges\nm.AddBond(0, 1, Chem.BondType.SINGLE)\nm.AddBond(1, 2, Chem.BondType.DOUBLE)\nm.AddBond(2, 3, Chem.BondType.SINGLE)\nm.AddBond(3, 4, Chem.BondType.SINGLE)\nm.AddBond(3, 5, Chem.BondType.SINGLE)\nm.AddBond(3, 6, Chem.BondType.SINGLE)\nm.GetBondBetweenAtoms(3, 5).SetBondDir(Chem.BondDir.BEGINWEDGE)\n\n# coords\ncoords = ((-2, -1), (-1, 0), (1, 0), (2, 1), (2, 0), (3.5, 1), (2, 2))\nconf = Chem.Conformer(m.GetNumAtoms())\nconf.Set3D(False)\nfor i, (x, y) in enumerate(coords):\n    conf.SetAtomPosition(i, (x, y, 0))\nm.AddConformer(conf)\n\n# magic\nChem.SanitizeMol(m)\nChem.DetectBondStereochemistry(m)\nChem.AssignChiralTypesFromBondDirs(m)\nChem.AssignStereochemistry(m)\n\nm.Debug()\nopts = Draw.MolDrawOptions()\nopts.addAtomIndices = True\nopts.addStereoAnnotation = True\nDraw.MolToImage(m, options=opts)\n```\n\n![image](https://user-images.githubusercontent.com/4602302/115084791-06f6d700-9f12-11eb-8e07-0aaa5c739b61.png)\n\nIn parallel I was looking for ways to extend training data with other images that are as close to the original training images as possible. While rdkit was very popular and useful in this competition, its rendering of molecules was not exactly the same. So I was exploring other tools, and finally I managed to find the tool that produced the exact images (except noise). It was [epam.indigo](https://lifescience.opensource.epam.com/indigo/api/index.html). I was very happy, when I found it. Firstly, it made possible to use 10M molecules from the extra dataset, secondly I could use it calculate coordinates of atoms in all existing molecules.\n\nWhile I was able to simplify atom numbering scheme in my representation, that (hopefully) is easier for NN to learn, I still was not very satisfied with it. Ideally I'd like to get rid of any order of atoms. But I was not able to achieve that.\n\nOriginal image:\n![image](https://user-images.githubusercontent.com/4602302/120806284-58553700-c54f-11eb-83b5-51103c5c0edc.png)\n\nRendered by epam.indigo with atom numbers:\n![image](https://user-images.githubusercontent.com/4602302/120806360-6f942480-c54f-11eb-853e-d47f87e44dcd.png)\n\nRendered by epam.indigo with 2D (x, y) atom numbering:\n![image](https://user-images.githubusercontent.com/4602302/120806428-820e5e00-c54f-11eb-922b-c24eab3c3865.png)\n\nWhat I was able to achieve with this representation is high accuracy of detecting stereo information. Going from almost 50% accuracy (for each stereocenter) when predicting inchi string directly to almost 100% accuracy when recovering it by rdkit from 2d coordinates. So I was able to achieve 95% exact inchi (LD=0) predictions on validation set, while predicting inchi as string I had somewhere below 90%, with ~5% mismatches only in 1-4 characters in stereolayer.\n\nOne major disadvantage of this approach is that sometimes it can predict broken molecules with some bonds missing in the middle. It happened for <1% of molecules, but then it produced very large LD, so in order to overcome this, I had to still use other model that predicted inchi as string. \n\nSo my final submission (public LB = 0.88) was an ensemble of 2 models:\n- model1: 384x384->16x16x256, effnet_b3 encoder, 6 layer transformer decoder (dim 256), inchi as string\n- model2: 512x512->16x16x256, effnet_b0 encoder, 6 layer transformer decoder (dim 256), molecular graph representation\n\nIn my submission only 4% of predicted inchis were from model1, others 96% were from pretty slim model2 (4M parameters in encoder, 5M in decoder). For all my training I was using only single Google Colab notebook. I didn't use beam search while inference.",
    "1335857": "Amazing work! Congrats!",
    "1336388": "Wonderfully inventive approach. Grreat job!",
    "1336465": "Great work, I'm surprised we didn't see more people following this path!",
    "1337700": "Great work, the InChI modeling looks quite creative.",
    "1337766": "This is beautiful!",
    "1340292": "After I wrote this post, I started to think, whether it actually helps to renumber atoms based on 2D coordinates. Intuitively, from a human perspective, it seems reasonable - you just look at image and assign numbers from left to right, from top to bottom, even for rather large molecules it shouldn't be very difficult. While algorithm described in InChI specification requires quite complicated calculations. But intuition might be misleading sometimes - some tasks that are easy for humans might be hard for NN and vice versa.\n\nTo validate this empirically, I've conducted a quick experiment training for one epoch the same model with same parameters - the only difference was in a few lines responsible for atom reordering.\n\n![image](https://user-images.githubusercontent.com/4602302/121061194-374d4a00-c7cc-11eb-87db-47dc6367b7de.png)\n\nFirst plot is percentage of exact InChI matches, second - levenshtein distance. You can see that, indeed, simplified atom numbering (🔴 red curve) makes the task much easier than when atoms are numbered according to the InChI spec (🔵 blue curve), and training goes much faster.\n\nThere is no intrinsic order in a molecule (nor in image) - it is just a graph with unordered sets of nodes and edges. So when we are forcing NN to learn some non-trivial order of atoms, we make the task harder than it is required. RDKit is able to calculate atom numbers according to InChI spec perfectly, so why not to outsource this task to it and make thing simpler for NN.\n\nAnother experiment is comparing performance of models one that outputs InChI as string directly (🟢 green curve) and another outputs mol graph (🔴 red curve). The result is ambiguous. From one side, if comparing percentage of exact matches, then model based on mol graph beats model based on InChI string with large margin (75% vs 35% after one epoch), even without simplified atom numbering (🔵 blue curve) - (55% vs 35%). From another side, model that outputs InChI string directly has lower LD, even if doesn't predict the result exactly. So in order to achieve better score, you should probably combine both.",
    "1517498": "wow, this is an awesome solution! Congratulations on the silver medal.\nCould you please share the epam.indigo code you used to create the 2D (x, y) atom numbering image?",
    "1517528": "Thanks )\n\nIt is not easy to navigate my code after several month 😜, but I believe this should be correct version:\n```\nmol_indigo = indigo.loadMolecule(inchi)\n\n# ask indigo to calculate 2d coords\nmol_indigo.layout()\n\n# serialize mol to molfile with 2d coords assigned and then load it with rdkit,\n# it is simplest way to pass coords from indigo to rdkit\nmol = Chem.MolFromMolBlock(mol_indigo.molfile(), removeHs=False, sanitize=False)\nChem.Kekulize(mol)\n\n# next few lines are responsible for correct stereo-bonds assignment based on 2d coords\n# so if you are interested in renumbering only, you can ignore them\nfor b in mol.GetBonds():\n    if b.GetBondDir() in (Chem.BondDir.BEGINDASH, Chem.BondDir.BEGINWEDGE):\n        b.SetBondDir(Chem.BondDir.NONE)\n\nconf = mol.GetConformer()\nChem.WedgeMolBonds(mol, mol.GetConformer())\n\ncoord_x = []\ncoord_y = []\nfor i in range(mol.GetNumAtoms()):\n    xyz = conf.GetAtomPosition(i)\n    coord_x.append(xyz[0])\n    coord_y.append(xyz[1])\ncoord_x = np.array(coord_x)\ncoord_y = np.array(coord_y)\ncoords = np.array([coord_x, coord_y]).T\n\nCOORDS_SCALE = 10\ncoords = (2 * (coords - coords.min(0)) / (\n        coords.max(0) - coords.min(0)) * COORDS_SCALE - COORDS_SCALE)\ncoords[np.isnan(coords)] = 0\norder = np.lexsort((-coord_y, coord_x))\nmol = Chem.RenumberAtoms(mol, order.tolist())\n\n# now atoms will be iterated in (x, y) order\nfor a in mol.GetAtoms():\n    # ...\n```\n\nBut maybe if you don't use rdkit, or you don't need anything except  atom 2d coordinates, you can use epam.indigo directly to get coords and then compute atom order based on them.",
    "1518059": "Thanks for the very kind reply. When I followed the code above, I confirmed that it was sorted. However, am I correct that the red arrow isn't explicitly declared in the image?",
    "1518070": "Yes, you are correct, I just draw it manually for demonstration, images are fed into network without any modification, the order is just used to sort network outputs.",
    "1518072": "Thanks to the quick answer, I understood. Thank you so much. Good luck!"
  },
  "source": "meta"
}