{
  "id": 224375,
  "title": "Understanding InChI (sublayer \"/c\")",
  "url": "/competitions/bms-molecular-translation/discussion/224375",
  "author_name": "",
  "post_date": "2021-03-08T07:51:37.924725600Z",
  "votes": 50,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The first question that I had after getting acquainted with the InChI notation – how exactly is the numbering of atoms performed? (and this question has already been asked on the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223503\" target=\"_blank\">forum</a>)</p>\n<p>After studying some resources, there was a certain understanding, which I want to describe in a short summary.<br>\nA couple of comments: 1) this is an <strong>incomplete</strong> description of the algorithm; 2) I am not a specialist in chemistry. Therefore, if there are any amendments and additions, please write in the comments</p>\n<blockquote>\n  <p>The main layer of the InChI ID consists of 3 sublayers that are separated \"/\":<br>\n  1) The chemical formula is represented according to Hill convention, that is, beginning with carbon atoms, then hydrogens, then all other elements in alphabetical order<br>\n  2) Sublayer of the structure (prefix \"/c\") - enumeration of <strong>canonical numbers</strong> in the chain of connected atoms (branches are indicated in parentheses)<br>\n  3) This layer prefixed with ‘/h’ lists the bonds between the atoms in the structure, partitioned into as many as three sublayers. The first sublayer represents all bonds other than those to non-bridging H-atoms, the second sublayer represents bonds of all immobile H-atoms, and the third sublayer provides locations of any mobile H-atoms. (<em>whatever that means</em>)</p>\n</blockquote>\n<p>In this topic, I will focus on <strong>point 2</strong> and the atom numbering algorithm.<br>\nIn the notebook <a href=\"https://www.kaggle.com/sapr3s/inchi-samples\" target=\"_blank\">1</a>, I showed several examples with images of simple molecules<br>\nEach molecule is displayed in 3 different forms: standard form; with atom numbers and with the number of hydrogen atoms for each atom of the basic structure.<br>\nFor the drawings, I use the techniques for the rdkit found in <a href=\"http://www.rdkit.org/docs/Cookbook.html\" target=\"_blank\">3</a><br>\nAnd the numbering algorithm itself is taken from <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">2</a> - <strong>Canonicalization</strong>.</p>\n<p>The list of possible numbers for the atoms depends on the chemical formula.<br>\nFor example, we have the formula \"C4H9Cl\". So the atoms \"C\" have numbers 1-4, and \"Cl\" – 5.</p>\n<p><strong>Description of the algorithm</strong></p>\n<p><img src=\"https://media.springernature.com/full/springer-static/image/art%3A10.1186%2Fs13321-015-0068-4/MediaObjects/13321_2015_68_Fig11_HTML.gif\" alt=\"Figure 11\"><br>\nFigure 11 from <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">2</a>  - <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4/figures/11\" target=\"_blank\">full size</a></p>\n<p>An iterative process is described (the steps are shown in the figure) using the example of the C4H9Cl molecule (2-chlorobutane) </p>\n<p><strong>Step A  - hydrogenless constitution</strong></p>\n<p>1) The skeletal atoms are labelled with numerical \"colors\" in the following order of precedence.</p>\n<ul>\n<li><p>Ordering number of chemical element in the sequence: carbon, other atoms in alphabetic order, bridging hydrogen.<br>\nIn case of C4H9Cl all C will be given color 1, Cl will be given 2.<br>\n(See 'a' on Figure)</p></li>\n<li><p>Determine the number of bonds for each atom. <br>\nIn 2-chlorobutane CH3CH2CH(Cl)CH3 these are (in curly brackets): C{1}C{2}C{3}(Cl{1})C{1};<br>\nThe resultant \"ordered lists of colors\" presented in order of the atoms in the semi-structural formula CH3CH2CH(Cl)CH3 are: C: (1, 1); C: (1, 2); C: (1, 3); Cl: (2, 1); C (1, 1)<br>\n(See 'b' on Figure)</p></li>\n</ul>\n<p>2) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order.<br>\nFor example, (1,1) &lt; (1,2) &lt; (2,1); (1, 2) &lt; (1, 2, 1)</p>\n<p>C: 1, 1 = &gt; 2;<br>\nC: 1, 2 = &gt; 3<br>\nC: 1, 3 = &gt; 4<br>\nCl: 2, 1 = &gt; 5<br>\nC 1, 1 = &gt; 2</p>\n<p>(See 'c' on Figure)</p>\n<p>Notice that each color is equal to the number of atoms that have this or smaller color</p>\n<p>3) Atoms are assigned new \"ordered lists of colors\": the first in the list is the color of the atom, the rest are sorted in ascending order colors of other atoms, connected to this atom</p>\n<p>C: 2, 3<br>\nC: 3, 2, 3 (<em>It seems to me that there is an error in the description and it should be '3, 2, 4'?</em>)<br>\nC: 4, 2, 3, 5<br>\nCl: 5, 4<br>\nC 2, 4</p>\n<p>(See 'd' on Figure)</p>\n<p>4) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order<br>\nC: 2, 3 = &gt; 1<br>\nC: 3, 2, 3 = &gt; 3<br>\nC: 4, 2, 3, 5 = &gt; 4<br>\nCl: 5, 4 = &gt; 5<br>\nC 2, 4 = &gt; 2</p>\n<p>(See 'e' on Figure)</p>\n<p>5) Steps 3–4 are repeated until all new colors are different or no more changes occur (for 2-chlorobutane the colors - canonical numbers - have already been found)<br>\n…</p>\n<p>there is a continuation of the algorithm: 6) … -9) and then Step B - Step D (<strong>I skip it</strong> … see <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">2</a>)</p>\n<p>=<br>\nAccordingly, then we just go in order, starting with the number 1 and write down the canonical numbers of the atoms, the bonds are denoted by the \"-\" sign.<br>\nAs a result, we get the \"/c \" sublayer for the InChI ID</p>\n<p><strong>Encoding of branches</strong></p>\n<p>It turns out that branches (and loops) are encoded using parentheses: \"(\"–the beginning of the branch; \")\" – end of the branch.<br>\nFor the cycle, the number of the atom that we start with is also repeated at the end.<br>\nFor example - 000afb9f46bb.png from train:<br>\nInChI=1S/C11H17ClN2/c1-8(9(6-13)7-14)10-4-2-3-5-11(10)12/h2-5,8-9H,6-7,13-14H2,1H3 (see <a href=\"https://www.kaggle.com/sapr3s/inchi-samples\" target=\"_blank\">1</a>)<br>\nHere, atom 10 is repeated twice - at the beginning of the cycle and at the end as a branch.</p>\n<p>So far, for me, the key question of the contest is how to get the 'canonical numbers' of atoms? <br>\nDo this algorithmically or using machine learning?</p>\n<p><strong>Links</strong><br>\n1 – Code with several examples of simple molecules<br>\n<a href=\"https://www.kaggle.com/sapr3s/inchi-samples\" target=\"_blank\">https://www.kaggle.com/sapr3s/inchi-samples</a></p>\n<p>2 – InChI, the IUPAC International Chemical Identifier <br>\n<a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4</a></p>\n<p>3 – RDKit Cookbook<br>\n<a href=\"http://www.rdkit.org/docs/Cookbook.html\" target=\"_blank\">http://www.rdkit.org/docs/Cookbook.html</a> </p>",
  "messages": [
    {
      "id": "1230519",
      "postDate": "03/08/2021 07:51:37",
      "content": "<p>The first question that I had after getting acquainted with the InChI notation – how exactly is the numbering of atoms performed? (and this question has already been asked on the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223503\" target=\"_blank\">forum</a>)</p>\n<p>After studying some resources, there was a certain understanding, which I want to describe in a short summary.<br>\nA couple of comments: 1) this is an <strong>incomplete</strong> description of the algorithm; 2) I am not a specialist in chemistry. Therefore, if there are any amendments and additions, please write in the comments</p>\n<blockquote>\n  <p>The main layer of the InChI ID consists of 3 sublayers that are separated \"/\":<br>\n  1) The chemical formula is represented according to Hill convention, that is, beginning with carbon atoms, then hydrogens, then all other elements in alphabetical order<br>\n  2) Sublayer of the structure (prefix \"/c\") - enumeration of <strong>canonical numbers</strong> in the chain of connected atoms (branches are indicated in parentheses)<br>\n  3) This layer prefixed with ‘/h’ lists the bonds between the atoms in the structure, partitioned into as many as three sublayers. The first sublayer represents all bonds other than those to non-bridging H-atoms, the second sublayer represents bonds of all immobile H-atoms, and the third sublayer provides locations of any mobile H-atoms. (<em>whatever that means</em>)</p>\n</blockquote>\n<p>In this topic, I will focus on <strong>point 2</strong> and the atom numbering algorithm.<br>\nIn the notebook <a href=\"https://www.kaggle.com/sapr3s/inchi-samples\" target=\"_blank\">1</a>, I showed several examples with images of simple molecules<br>\nEach molecule is displayed in 3 different forms: standard form; with atom numbers and with the number of hydrogen atoms for each atom of the basic structure.<br>\nFor the drawings, I use the techniques for the rdkit found in <a href=\"http://www.rdkit.org/docs/Cookbook.html\" target=\"_blank\">3</a><br>\nAnd the numbering algorithm itself is taken from <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">2</a> - <strong>Canonicalization</strong>.</p>\n<p>The list of possible numbers for the atoms depends on the chemical formula.<br>\nFor example, we have the formula \"C4H9Cl\". So the atoms \"C\" have numbers 1-4, and \"Cl\" – 5.</p>\n<p><strong>Description of the algorithm</strong></p>\n<p><img src=\"https://media.springernature.com/full/springer-static/image/art%3A10.1186%2Fs13321-015-0068-4/MediaObjects/13321_2015_68_Fig11_HTML.gif\" alt=\"Figure 11\"><br>\nFigure 11 from <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">2</a>  - <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4/figures/11\" target=\"_blank\">full size</a></p>\n<p>An iterative process is described (the steps are shown in the figure) using the example of the C4H9Cl molecule (2-chlorobutane) </p>\n<p><strong>Step A  - hydrogenless constitution</strong></p>\n<p>1) The skeletal atoms are labelled with numerical \"colors\" in the following order of precedence.</p>\n<ul>\n<li><p>Ordering number of chemical element in the sequence: carbon, other atoms in alphabetic order, bridging hydrogen.<br>\nIn case of C4H9Cl all C will be given color 1, Cl will be given 2.<br>\n(See 'a' on Figure)</p></li>\n<li><p>Determine the number of bonds for each atom. <br>\nIn 2-chlorobutane CH3CH2CH(Cl)CH3 these are (in curly brackets): C{1}C{2}C{3}(Cl{1})C{1};<br>\nThe resultant \"ordered lists of colors\" presented in order of the atoms in the semi-structural formula CH3CH2CH(Cl)CH3 are: C: (1, 1); C: (1, 2); C: (1, 3); Cl: (2, 1); C (1, 1)<br>\n(See 'b' on Figure)</p></li>\n</ul>\n<p>2) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order.<br>\nFor example, (1,1) &lt; (1,2) &lt; (2,1); (1, 2) &lt; (1, 2, 1)</p>\n<p>C: 1, 1 = &gt; 2;<br>\nC: 1, 2 = &gt; 3<br>\nC: 1, 3 = &gt; 4<br>\nCl: 2, 1 = &gt; 5<br>\nC 1, 1 = &gt; 2</p>\n<p>(See 'c' on Figure)</p>\n<p>Notice that each color is equal to the number of atoms that have this or smaller color</p>\n<p>3) Atoms are assigned new \"ordered lists of colors\": the first in the list is the color of the atom, the rest are sorted in ascending order colors of other atoms, connected to this atom</p>\n<p>C: 2, 3<br>\nC: 3, 2, 3 (<em>It seems to me that there is an error in the description and it should be '3, 2, 4'?</em>)<br>\nC: 4, 2, 3, 5<br>\nCl: 5, 4<br>\nC 2, 4</p>\n<p>(See 'd' on Figure)</p>\n<p>4) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order<br>\nC: 2, 3 = &gt; 1<br>\nC: 3, 2, 3 = &gt; 3<br>\nC: 4, 2, 3, 5 = &gt; 4<br>\nCl: 5, 4 = &gt; 5<br>\nC 2, 4 = &gt; 2</p>\n<p>(See 'e' on Figure)</p>\n<p>5) Steps 3–4 are repeated until all new colors are different or no more changes occur (for 2-chlorobutane the colors - canonical numbers - have already been found)<br>\n…</p>\n<p>there is a continuation of the algorithm: 6) … -9) and then Step B - Step D (<strong>I skip it</strong> … see <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">2</a>)</p>\n<p>=<br>\nAccordingly, then we just go in order, starting with the number 1 and write down the canonical numbers of the atoms, the bonds are denoted by the \"-\" sign.<br>\nAs a result, we get the \"/c \" sublayer for the InChI ID</p>\n<p><strong>Encoding of branches</strong></p>\n<p>It turns out that branches (and loops) are encoded using parentheses: \"(\"–the beginning of the branch; \")\" – end of the branch.<br>\nFor the cycle, the number of the atom that we start with is also repeated at the end.<br>\nFor example - 000afb9f46bb.png from train:<br>\nInChI=1S/C11H17ClN2/c1-8(9(6-13)7-14)10-4-2-3-5-11(10)12/h2-5,8-9H,6-7,13-14H2,1H3 (see <a href=\"https://www.kaggle.com/sapr3s/inchi-samples\" target=\"_blank\">1</a>)<br>\nHere, atom 10 is repeated twice - at the beginning of the cycle and at the end as a branch.</p>\n<p>So far, for me, the key question of the contest is how to get the 'canonical numbers' of atoms? <br>\nDo this algorithmically or using machine learning?</p>\n<p><strong>Links</strong><br>\n1 – Code with several examples of simple molecules<br>\n<a href=\"https://www.kaggle.com/sapr3s/inchi-samples\" target=\"_blank\">https://www.kaggle.com/sapr3s/inchi-samples</a></p>\n<p>2 – InChI, the IUPAC International Chemical Identifier <br>\n<a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\" target=\"_blank\">https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4</a></p>\n<p>3 – RDKit Cookbook<br>\n<a href=\"http://www.rdkit.org/docs/Cookbook.html\" target=\"_blank\">http://www.rdkit.org/docs/Cookbook.html</a> </p>",
      "rawMarkdown": "The first question that I had after getting acquainted with the InChI notation – how exactly is the numbering of atoms performed? (and this question has already been asked on the [forum](https://www.kaggle.com/c/bms-molecular-translation/discussion/223503))\n\nAfter studying some resources, there was a certain understanding, which I want to describe in a short summary.\nA couple of comments: 1) this is an **incomplete** description of the algorithm; 2) I am not a specialist in chemistry. Therefore, if there are any amendments and additions, please write in the comments\n\n\n> The main layer of the InChI ID consists of 3 sublayers that are separated \"/\":\n1) The chemical formula is represented according to Hill convention, that is, beginning with carbon atoms, then hydrogens, then all other elements in alphabetical order\n2) Sublayer of the structure (prefix \"/c\") - enumeration of **canonical numbers** in the chain of connected atoms (branches are indicated in parentheses)\n3) This layer prefixed with ‘/h’ lists the bonds between the atoms in the structure, partitioned into as many as three sublayers. The first sublayer represents all bonds other than those to non-bridging H-atoms, the second sublayer represents bonds of all immobile H-atoms, and the third sublayer provides locations of any mobile H-atoms. (*whatever that means*)\n\n\nIn this topic, I will focus on **point 2** and the atom numbering algorithm.\nIn the notebook [1](https://www.kaggle.com/sapr3s/inchi-samples), I showed several examples with images of simple molecules\nEach molecule is displayed in 3 different forms: standard form; with atom numbers and with the number of hydrogen atoms for each atom of the basic structure.\nFor the drawings, I use the techniques for the rdkit found in [3](http://www.rdkit.org/docs/Cookbook.html)\nAnd the numbering algorithm itself is taken from [2](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4) - **Canonicalization**.\n\nThe list of possible numbers for the atoms depends on the chemical formula.\nFor example, we have the formula \"C4H9Cl\". So the atoms \"C\" have numbers 1-4, and \"Cl\" – 5.\n\n**Description of the algorithm**\n\n![Figure 11](https://media.springernature.com/full/springer-static/image/art%3A10.1186%2Fs13321-015-0068-4/MediaObjects/13321_2015_68_Fig11_HTML.gif)\nFigure 11 from [2](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4)  - [full size](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4/figures/11)\n\nAn iterative process is described (the steps are shown in the figure) using the example of the C4H9Cl molecule (2-chlorobutane) \n\n**Step A  - hydrogenless constitution**\n\n1) The skeletal atoms are labelled with numerical \"colors\" in the following order of precedence.\n- Ordering number of chemical element in the sequence: carbon, other atoms in alphabetic order, bridging hydrogen.\nIn case of C4H9Cl all C will be given color 1, Cl will be given 2.\n(See 'a' on Figure)\n\n- Determine the number of bonds for each atom. \nIn 2-chlorobutane CH3CH2CH(Cl)CH3 these are (in curly brackets): C{1}C{2}C{3}(Cl{1})C{1};\nThe resultant \"ordered lists of colors\" presented in order of the atoms in the semi-structural formula CH3CH2CH(Cl)CH3 are: C: (1, 1); C: (1, 2); C: (1, 3); Cl: (2, 1); C (1, 1)\n(See 'b' on Figure)\n\n2) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order.\nFor example, (1,1) < (1,2) < (2,1); (1, 2) < (1, 2, 1)\n\nC: 1, 1 = > 2;\nC: 1, 2 = > 3\nC: 1, 3 = > 4\nCl: 2, 1 = > 5\nC 1, 1 = > 2\n\n(See 'c' on Figure)\n\nNotice that each color is equal to the number of atoms that have this or smaller color\n\n3) Atoms are assigned new \"ordered lists of colors\": the first in the list is the color of the atom, the rest are sorted in ascending order colors of other atoms, connected to this atom\n\nC: 2, 3\nC: 3, 2, 3 (*It seems to me that there is an error in the description and it should be '3, 2, 4'?*)\nC: 4, 2, 3, 5\nCl: 5, 4\nC 2, 4\n\n(See 'd' on Figure)\n\n4) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order\nC: 2, 3 = > 1\nC: 3, 2, 3 = > 3\nC: 4, 2, 3, 5 = > 4\nCl: 5, 4 = > 5\nC 2, 4 = > 2\n\n(See 'e' on Figure)\n\n5) Steps 3–4 are repeated until all new colors are different or no more changes occur (for 2-chlorobutane the colors - canonical numbers - have already been found)\n...\n\nthere is a continuation of the algorithm: 6) ... -9) and then Step B - Step D (**I skip it** ... see [2](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4))\n\n=\nAccordingly, then we just go in order, starting with the number 1 and write down the canonical numbers of the atoms, the bonds are denoted by the \"-\" sign.\nAs a result, we get the \"/c \" sublayer for the InChI ID\n\n**Encoding of branches**\n\nIt turns out that branches (and loops) are encoded using parentheses: \"(\"–the beginning of the branch; \")\" – end of the branch.\nFor the cycle, the number of the atom that we start with is also repeated at the end.\nFor example - 000afb9f46bb.png from train:\nInChI=1S/C11H17ClN2/c1-8(9(6-13)7-14)10-4-2-3-5-11(10)12/h2-5,8-9H,6-7,13-14H2,1H3 (see [1](https://www.kaggle.com/sapr3s/inchi-samples))\nHere, atom 10 is repeated twice - at the beginning of the cycle and at the end as a branch.\n\nSo far, for me, the key question of the contest is how to get the 'canonical numbers' of atoms? \nDo this algorithmically or using machine learning?\n\n**Links**\n1 – Code with several examples of simple molecules\nhttps://www.kaggle.com/sapr3s/inchi-samples\n\n2 – InChI, the IUPAC International Chemical Identifier \nhttps://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\n\n3 – RDKit Cookbook\nhttp://www.rdkit.org/docs/Cookbook.html",
      "votes": null
    },
    {
      "id": "1232481",
      "postDate": "03/09/2021 19:33:48",
      "content": "<p>Thanks a lot, I was wondering exactly about that and had just opened the JCheminf paper from Heller et al. (2015).<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">Here</a> the competion hosts give some insight into how it could be done. If I understood correctly they suggest to use an intermediate representation that can be converted algorithmically to a valid InChI.</p>",
      "rawMarkdown": "Thanks a lot, I was wondering exactly about that and had just opened the JCheminf paper from Heller et al. (2015).\n[Here](https://www.kaggle.com/c/bms-molecular-translation/discussion/224394) the competion hosts give some insight into how it could be done. If I understood correctly they suggest to use an intermediate representation that can be converted algorithmically to a valid InChI.",
      "votes": null
    },
    {
      "id": "1244273",
      "postDate": "03/18/2021 21:20:48",
      "content": "<p>\"that can be converted algorithmically to a valid InChI\"<br>\nThe question is what is this algorithm precisely and is it already implemented somewhere ?<br>\nHow do they came up with the InChI in the training set ?</p>",
      "rawMarkdown": "\"that can be converted algorithmically to a valid InChI\"\nThe question is what is this algorithm precisely and is it already implemented somewhere ?\nHow do they came up with the InChI in the training set ?",
      "votes": null
    },
    {
      "id": "1244551",
      "postDate": "03/19/2021 04:09:32",
      "content": "<p>Implemented, of course. For example, in RDKit :</p>\n<pre><code>from rdkit import Chem\nmol = Chem.MolFromInchi(inchi)\n# ...\nrestored_inchi = Chem.MolToInchi(mol )\n</code></pre>\n<p>'mol' is an object from which you can get atoms (<em>mol.GetAtoms()</em>) or bonds (<em>mol.GetBonds()</em>)</p>",
      "rawMarkdown": "Implemented, of course. For example, in RDKit :\n```\nfrom rdkit import Chem\nmol = Chem.MolFromInchi(inchi)\n# ...\nrestored_inchi = Chem.MolToInchi(mol )\n\n```\n'mol' is an object from which you can get atoms (*mol.GetAtoms()*) or bonds (*mol.GetBonds()*)",
      "votes": null
    },
    {
      "id": "1245420",
      "postDate": "03/19/2021 19:11:28",
      "content": "<p>Thanks Pavel for the code !<br>\nI checked training InChIs with rdkit and following code :</p>\n<pre><code>from rdkit import Chem\nfrom tqdm.notebook import tqdm\nimport pandas as pd\n\ndf = pd.read_csv(\"train_labels.csv\")\nnot_equal = []\n\nfor inchi in tqdm(df.InChI):\n    mol = Chem.MolFromInchi(inchi)\n    inchi2 = Chem.MolToInchi(mol)\n    if inchi != inchi2:\n        not_equal.append((inchi, inchi2))\n\nprint(len(not_equal)/len(df)*100)  # this prints 0.008173022470833305\n</code></pre>\n<p>I found that there is only <strong>0.008%</strong> of the training InChIs with conversion error (inchi != inchi2).<br>\nSo rdkit InChI implementation and training InChIs agree very well which is reassuring.</p>",
      "rawMarkdown": "Thanks Pavel for the code !\nI checked training InChIs with rdkit and following code :\n```python\nfrom rdkit import Chem\nfrom tqdm.notebook import tqdm\nimport pandas as pd\n\ndf = pd.read_csv(\"train_labels.csv\")\nnot_equal = []\n\nfor inchi in tqdm(df.InChI):\n    mol = Chem.MolFromInchi(inchi)\n    inchi2 = Chem.MolToInchi(mol)\n    if inchi != inchi2:\n        not_equal.append((inchi, inchi2))\n\nprint(len(not_equal)/len(df)*100)  # this prints 0.008173022470833305\n```\nI found that there is only **0.008%** of the training InChIs with conversion error (inchi != inchi2).\nSo rdkit InChI implementation and training InChIs agree very well which is reassuring.",
      "votes": null
    },
    {
      "id": "1247580",
      "postDate": "03/21/2021 20:49:51",
      "content": "<p>Thank for explaining this! I have cracked my head trying to understand this.</p>",
      "rawMarkdown": "Thank for explaining this! I have cracked my head trying to understand this.",
      "votes": null
    },
    {
      "id": "1255888",
      "postDate": "03/29/2021 10:13:41",
      "content": "<p>I found on 100 000 random InChIs that by predicting only the 3 first parts of InChI (formula, connections and hydrogens), the average levenstein is <strong>2.86</strong>.<br>\n<strong>It means that the best possible leaderboard score with a model predicting only the molecular graph (formula, connections and hydrogens) is 2.86.</strong><br>\nFor details, here is the code I use on the 100 000 InChIs :</p>\n<pre><code>from Levenshtein import distance\ninchi_trunc = \"/\".join(inchi.split(\"/\", maxsplit=4)[:4])\nlevenstein = distance(inchi_trunc, inchi)\n</code></pre>",
      "rawMarkdown": "I found on 100 000 random InChIs that by predicting only the 3 first parts of InChI (formula, connections and hydrogens), the average levenstein is **2.86**.\n**It means that the best possible leaderboard score with a model predicting only the molecular graph (formula, connections and hydrogens) is 2.86.**\nFor details, here is the code I use on the 100 000 InChIs :\n```python\nfrom Levenshtein import distance\ninchi_trunc = \"/\".join(inchi.split(\"/\", maxsplit=4)[:4])\nlevenstein = distance(inchi_trunc, inchi)\n```",
      "votes": null
    },
    {
      "id": "1256394",
      "postDate": "03/29/2021 20:34:40",
      "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> - thank you for the exhaustive explanation. Since no-one else mentioned it, I'll note that a mobile hydrogen is attached to different atoms in different tautomers.</p>",
      "rawMarkdown": "sapr3s - thank you for the exhaustive explanation. Since no-one else mentioned it, I'll note that a mobile hydrogen is attached to different atoms in different tautomers.",
      "votes": null
    },
    {
      "id": "1256630",
      "postDate": "03/30/2021 05:16:05",
      "content": "<p>Perhaps a model such as a <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229168\" target=\"_blank\">binary classifier</a> can be used for the remaining InChI string.</p>\n<p>Multi-step approach after /c and /h layers:</p>\n<ul>\n<li>Predict whether /b layer should be present <code>-&gt;</code> then predict the /b string and append</li>\n<li>Predict whether /t layer should be present <code>-&gt;</code> then predict the /t string and append</li>\n<li>Predict whether /m0 or /m1 or /s1 layer should be present and append these short strings at the end</li>\n</ul>\n<p>This should hopefully reduce Levenshtein distance below 2.0 or even below 1.0!</p>",
      "rawMarkdown": "Perhaps a model such as a [binary classifier](https://www.kaggle.com/c/bms-molecular-translation/discussion/229168) can be used for the remaining InChI string.\n\nMulti-step approach after /c and /h layers:\n- Predict whether /b layer should be present `->` then predict the /b string and append\n- Predict whether /t layer should be present `->` then predict the /t string and append\n- Predict whether /m0 or /m1 or /s1 layer should be present and append these short strings at the end\n\nThis should hopefully reduce Levenshtein distance below 2.0 or even below 1.0!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1232481,
      "author_name": "michaelwolff",
      "author_url": "",
      "post_date": "03/09/2021 19:33:48",
      "content": "<p>Thanks a lot, I was wondering exactly about that and had just opened the JCheminf paper from Heller et al. (2015).<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">Here</a> the competion hosts give some insight into how it could be done. If I understood correctly they suggest to use an intermediate representation that can be converted algorithmically to a valid InChI.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1244273,
          "author_name": "datanewb",
          "author_url": "",
          "post_date": "03/18/2021 21:20:48",
          "content": "<p>\"that can be converted algorithmically to a valid InChI\"<br>\nThe question is what is this algorithm precisely and is it already implemented somewhere ?<br>\nHow do they came up with the InChI in the training set ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1244551,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "03/19/2021 04:09:32",
          "content": "<p>Implemented, of course. For example, in RDKit :</p>\n<pre><code>from rdkit import Chem\nmol = Chem.MolFromInchi(inchi)\n# ...\nrestored_inchi = Chem.MolToInchi(mol )\n</code></pre>\n<p>'mol' is an object from which you can get atoms (<em>mol.GetAtoms()</em>) or bonds (<em>mol.GetBonds()</em>)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1245420,
          "author_name": "datanewb",
          "author_url": "",
          "post_date": "03/19/2021 19:11:28",
          "content": "<p>Thanks Pavel for the code !<br>\nI checked training InChIs with rdkit and following code :</p>\n<pre><code>from rdkit import Chem\nfrom tqdm.notebook import tqdm\nimport pandas as pd\n\ndf = pd.read_csv(\"train_labels.csv\")\nnot_equal = []\n\nfor inchi in tqdm(df.InChI):\n    mol = Chem.MolFromInchi(inchi)\n    inchi2 = Chem.MolToInchi(mol)\n    if inchi != inchi2:\n        not_equal.append((inchi, inchi2))\n\nprint(len(not_equal)/len(df)*100)  # this prints 0.008173022470833305\n</code></pre>\n<p>I found that there is only <strong>0.008%</strong> of the training InChIs with conversion error (inchi != inchi2).<br>\nSo rdkit InChI implementation and training InChIs agree very well which is reassuring.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1247580,
      "author_name": "dmitryyemelyanov",
      "author_url": "",
      "post_date": "03/21/2021 20:49:51",
      "content": "<p>Thank for explaining this! I have cracked my head trying to understand this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1255888,
      "author_name": "datanewb",
      "author_url": "",
      "post_date": "03/29/2021 10:13:41",
      "content": "<p>I found on 100 000 random InChIs that by predicting only the 3 first parts of InChI (formula, connections and hydrogens), the average levenstein is <strong>2.86</strong>.<br>\n<strong>It means that the best possible leaderboard score with a model predicting only the molecular graph (formula, connections and hydrogens) is 2.86.</strong><br>\nFor details, here is the code I use on the 100 000 InChIs :</p>\n<pre><code>from Levenshtein import distance\ninchi_trunc = \"/\".join(inchi.split(\"/\", maxsplit=4)[:4])\nlevenstein = distance(inchi_trunc, inchi)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1256630,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "03/30/2021 05:16:05",
          "content": "<p>Perhaps a model such as a <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229168\" target=\"_blank\">binary classifier</a> can be used for the remaining InChI string.</p>\n<p>Multi-step approach after /c and /h layers:</p>\n<ul>\n<li>Predict whether /b layer should be present <code>-&gt;</code> then predict the /b string and append</li>\n<li>Predict whether /t layer should be present <code>-&gt;</code> then predict the /t string and append</li>\n<li>Predict whether /m0 or /m1 or /s1 layer should be present and append these short strings at the end</li>\n</ul>\n<p>This should hopefully reduce Levenshtein distance below 2.0 or even below 1.0!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1256394,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "03/29/2021 20:34:40",
      "content": "<p><a href=\"https://www.kaggle.com/sapr3s\" target=\"_blank\">@sapr3s</a> - thank you for the exhaustive explanation. Since no-one else mentioned it, I'll note that a mobile hydrogen is attached to different atoms in different tautomers.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1230519": "The first question that I had after getting acquainted with the InChI notation – how exactly is the numbering of atoms performed? (and this question has already been asked on the [forum](https://www.kaggle.com/c/bms-molecular-translation/discussion/223503))\n\nAfter studying some resources, there was a certain understanding, which I want to describe in a short summary.\nA couple of comments: 1) this is an **incomplete** description of the algorithm; 2) I am not a specialist in chemistry. Therefore, if there are any amendments and additions, please write in the comments\n\n\n> The main layer of the InChI ID consists of 3 sublayers that are separated \"/\":\n1) The chemical formula is represented according to Hill convention, that is, beginning with carbon atoms, then hydrogens, then all other elements in alphabetical order\n2) Sublayer of the structure (prefix \"/c\") - enumeration of **canonical numbers** in the chain of connected atoms (branches are indicated in parentheses)\n3) This layer prefixed with ‘/h’ lists the bonds between the atoms in the structure, partitioned into as many as three sublayers. The first sublayer represents all bonds other than those to non-bridging H-atoms, the second sublayer represents bonds of all immobile H-atoms, and the third sublayer provides locations of any mobile H-atoms. (*whatever that means*)\n\n\nIn this topic, I will focus on **point 2** and the atom numbering algorithm.\nIn the notebook [1](https://www.kaggle.com/sapr3s/inchi-samples), I showed several examples with images of simple molecules\nEach molecule is displayed in 3 different forms: standard form; with atom numbers and with the number of hydrogen atoms for each atom of the basic structure.\nFor the drawings, I use the techniques for the rdkit found in [3](http://www.rdkit.org/docs/Cookbook.html)\nAnd the numbering algorithm itself is taken from [2](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4) - **Canonicalization**.\n\nThe list of possible numbers for the atoms depends on the chemical formula.\nFor example, we have the formula \"C4H9Cl\". So the atoms \"C\" have numbers 1-4, and \"Cl\" – 5.\n\n**Description of the algorithm**\n\n![Figure 11](https://media.springernature.com/full/springer-static/image/art%3A10.1186%2Fs13321-015-0068-4/MediaObjects/13321_2015_68_Fig11_HTML.gif)\nFigure 11 from [2](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4)  - [full size](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4/figures/11)\n\nAn iterative process is described (the steps are shown in the figure) using the example of the C4H9Cl molecule (2-chlorobutane) \n\n**Step A  - hydrogenless constitution**\n\n1) The skeletal atoms are labelled with numerical \"colors\" in the following order of precedence.\n- Ordering number of chemical element in the sequence: carbon, other atoms in alphabetic order, bridging hydrogen.\nIn case of C4H9Cl all C will be given color 1, Cl will be given 2.\n(See 'a' on Figure)\n\n- Determine the number of bonds for each atom. \nIn 2-chlorobutane CH3CH2CH(Cl)CH3 these are (in curly brackets): C{1}C{2}C{3}(Cl{1})C{1};\nThe resultant \"ordered lists of colors\" presented in order of the atoms in the semi-structural formula CH3CH2CH(Cl)CH3 are: C: (1, 1); C: (1, 2); C: (1, 3); Cl: (2, 1); C (1, 1)\n(See 'b' on Figure)\n\n2) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order.\nFor example, (1,1) < (1,2) < (2,1); (1, 2) < (1, 2, 1)\n\nC: 1, 1 = > 2;\nC: 1, 2 = > 3\nC: 1, 3 = > 4\nCl: 2, 1 = > 5\nC 1, 1 = > 2\n\n(See 'c' on Figure)\n\nNotice that each color is equal to the number of atoms that have this or smaller color\n\n3) Atoms are assigned new \"ordered lists of colors\": the first in the list is the color of the atom, the rest are sorted in ascending order colors of other atoms, connected to this atom\n\nC: 2, 3\nC: 3, 2, 3 (*It seems to me that there is an error in the description and it should be '3, 2, 4'?*)\nC: 4, 2, 3, 5\nCl: 5, 4\nC 2, 4\n\n(See 'd' on Figure)\n\n4) Atoms are assigned new colors according to lexicographical comparison of the \"color lists\", in ascending order\nC: 2, 3 = > 1\nC: 3, 2, 3 = > 3\nC: 4, 2, 3, 5 = > 4\nCl: 5, 4 = > 5\nC 2, 4 = > 2\n\n(See 'e' on Figure)\n\n5) Steps 3–4 are repeated until all new colors are different or no more changes occur (for 2-chlorobutane the colors - canonical numbers - have already been found)\n...\n\nthere is a continuation of the algorithm: 6) ... -9) and then Step B - Step D (**I skip it** ... see [2](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4))\n\n=\nAccordingly, then we just go in order, starting with the number 1 and write down the canonical numbers of the atoms, the bonds are denoted by the \"-\" sign.\nAs a result, we get the \"/c \" sublayer for the InChI ID\n\n**Encoding of branches**\n\nIt turns out that branches (and loops) are encoded using parentheses: \"(\"–the beginning of the branch; \")\" – end of the branch.\nFor the cycle, the number of the atom that we start with is also repeated at the end.\nFor example - 000afb9f46bb.png from train:\nInChI=1S/C11H17ClN2/c1-8(9(6-13)7-14)10-4-2-3-5-11(10)12/h2-5,8-9H,6-7,13-14H2,1H3 (see [1](https://www.kaggle.com/sapr3s/inchi-samples))\nHere, atom 10 is repeated twice - at the beginning of the cycle and at the end as a branch.\n\nSo far, for me, the key question of the contest is how to get the 'canonical numbers' of atoms? \nDo this algorithmically or using machine learning?\n\n**Links**\n1 – Code with several examples of simple molecules\nhttps://www.kaggle.com/sapr3s/inchi-samples\n\n2 – InChI, the IUPAC International Chemical Identifier \nhttps://jcheminf.biomedcentral.com/articles/10.1186/s13321-015-0068-4\n\n3 – RDKit Cookbook\nhttp://www.rdkit.org/docs/Cookbook.html",
    "1232481": "Thanks a lot, I was wondering exactly about that and had just opened the JCheminf paper from Heller et al. (2015).\n[Here](https://www.kaggle.com/c/bms-molecular-translation/discussion/224394) the competion hosts give some insight into how it could be done. If I understood correctly they suggest to use an intermediate representation that can be converted algorithmically to a valid InChI.",
    "1244273": "\"that can be converted algorithmically to a valid InChI\"\nThe question is what is this algorithm precisely and is it already implemented somewhere ?\nHow do they came up with the InChI in the training set ?",
    "1244551": "Implemented, of course. For example, in RDKit :\n```\nfrom rdkit import Chem\nmol = Chem.MolFromInchi(inchi)\n# ...\nrestored_inchi = Chem.MolToInchi(mol )\n\n```\n'mol' is an object from which you can get atoms (*mol.GetAtoms()*) or bonds (*mol.GetBonds()*)",
    "1245420": "Thanks Pavel for the code !\nI checked training InChIs with rdkit and following code :\n```python\nfrom rdkit import Chem\nfrom tqdm.notebook import tqdm\nimport pandas as pd\n\ndf = pd.read_csv(\"train_labels.csv\")\nnot_equal = []\n\nfor inchi in tqdm(df.InChI):\n    mol = Chem.MolFromInchi(inchi)\n    inchi2 = Chem.MolToInchi(mol)\n    if inchi != inchi2:\n        not_equal.append((inchi, inchi2))\n\nprint(len(not_equal)/len(df)*100)  # this prints 0.008173022470833305\n```\nI found that there is only **0.008%** of the training InChIs with conversion error (inchi != inchi2).\nSo rdkit InChI implementation and training InChIs agree very well which is reassuring.",
    "1247580": "Thank for explaining this! I have cracked my head trying to understand this.",
    "1255888": "I found on 100 000 random InChIs that by predicting only the 3 first parts of InChI (formula, connections and hydrogens), the average levenstein is **2.86**.\n**It means that the best possible leaderboard score with a model predicting only the molecular graph (formula, connections and hydrogens) is 2.86.**\nFor details, here is the code I use on the 100 000 InChIs :\n```python\nfrom Levenshtein import distance\ninchi_trunc = \"/\".join(inchi.split(\"/\", maxsplit=4)[:4])\nlevenstein = distance(inchi_trunc, inchi)\n```",
    "1256394": "sapr3s - thank you for the exhaustive explanation. Since no-one else mentioned it, I'll note that a mobile hydrogen is attached to different atoms in different tautomers.",
    "1256630": "Perhaps a model such as a [binary classifier](https://www.kaggle.com/c/bms-molecular-translation/discussion/229168) can be used for the remaining InChI string.\n\nMulti-step approach after /c and /h layers:\n- Predict whether /b layer should be present `->` then predict the /b string and append\n- Predict whether /t layer should be present `->` then predict the /t string and append\n- Predict whether /m0 or /m1 or /s1 layer should be present and append these short strings at the end\n\nThis should hopefully reduce Levenshtein distance below 2.0 or even below 1.0!"
  },
  "source": "meta"
}