{
  "id": 225943,
  "title": "Precise definition for InChI",
  "url": "/competitions/bms-molecular-translation/discussion/225943",
  "author_name": "",
  "post_date": "2021-03-14T14:45:10.228060800Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>If possible, I'd like to know Precise definition for InChI that how to make it from the chemical structure figure. According to the wiki page (<a href=\"https://en.wikipedia.org/wiki/International_Chemical_Identifier)\" target=\"_blank\">https://en.wikipedia.org/wiki/International_Chemical_Identifier)</a>, if we see \"Format and layers\" section of it, it might not express all rules to make InChI.  if we see \"ethanol\" example in it, it is understandable, but in case of \"L-ascorbic acid\", it is difficult to understand. \"-\" of \"h2,5,7-8,10-11H,1H2\" might means \"Hydrogen bonding\". but it doesn't show explicitly in the figure.  InChI  expression need some chemical knowledge which doesn't not necessary show in the figure. Also in case of \"ethanol\" example,  number of atoms are numbered from the left, but in case of \"L-ascorbic acid, the  rule is not necessary true. If  InChI expression include some variation in one molecular structure,  examples in the train dataset is not only one precise expression,  but also other expression might be true. In case  of it, whether we will tune up out DL model to the train dataset expression or we will take on strategy to make one of correct expressions from the figure. Please advice.</p>\n<p>I briefly see  InChI  page (<a href=\"https://iupac.org/who-we-are/divisions/division-details/inchi/\" target=\"_blank\">https://iupac.org/who-we-are/divisions/division-details/inchi/</a>) also</p>",
  "messages": [
    {
      "id": "1237976",
      "postDate": "03/14/2021 14:45:10",
      "content": "<p>If possible, I'd like to know Precise definition for InChI that how to make it from the chemical structure figure. According to the wiki page (<a href=\"https://en.wikipedia.org/wiki/International_Chemical_Identifier)\" target=\"_blank\">https://en.wikipedia.org/wiki/International_Chemical_Identifier)</a>, if we see \"Format and layers\" section of it, it might not express all rules to make InChI.  if we see \"ethanol\" example in it, it is understandable, but in case of \"L-ascorbic acid\", it is difficult to understand. \"-\" of \"h2,5,7-8,10-11H,1H2\" might means \"Hydrogen bonding\". but it doesn't show explicitly in the figure.  InChI  expression need some chemical knowledge which doesn't not necessary show in the figure. Also in case of \"ethanol\" example,  number of atoms are numbered from the left, but in case of \"L-ascorbic acid, the  rule is not necessary true. If  InChI expression include some variation in one molecular structure,  examples in the train dataset is not only one precise expression,  but also other expression might be true. In case  of it, whether we will tune up out DL model to the train dataset expression or we will take on strategy to make one of correct expressions from the figure. Please advice.</p>\n<p>I briefly see  InChI  page (<a href=\"https://iupac.org/who-we-are/divisions/division-details/inchi/\" target=\"_blank\">https://iupac.org/who-we-are/divisions/division-details/inchi/</a>) also</p>",
      "rawMarkdown": "If possible, I'd like to know Precise definition for InChI that how to make it from the chemical structure figure. According to the wiki page (https://en.wikipedia.org/wiki/International_Chemical_Identifier), if we see \"Format and layers\" section of it, it might not express all rules to make InChI.  if we see \"ethanol\" example in it, it is understandable, but in case of \"L-ascorbic acid\", it is difficult to understand. \"-\" of \"h2,5,7-8,10-11H,1H2\" might means \"Hydrogen bonding\". but it doesn't show explicitly in the figure.  InChI  expression need some chemical knowledge which doesn't not necessary show in the figure. Also in case of \"ethanol\" example,  number of atoms are numbered from the left, but in case of \"L-ascorbic acid, the  rule is not necessary true. If  InChI expression include some variation in one molecular structure,  examples in the train dataset is not only one precise expression,  but also other expression might be true. In case  of it, whether we will tune up out DL model to the train dataset expression or we will take on strategy to make one of correct expressions from the figure. Please advice.\n\nI briefly see  InChI  page (https://iupac.org/who-we-are/divisions/division-details/inchi/) also",
      "votes": null
    },
    {
      "id": "1238360",
      "postDate": "03/14/2021 23:33:17",
      "content": "<p>This is a very good question. In addition to your post, I will give an example where different InChI can be converted to the same molecule.</p>\n<pre><code>from rdkit.Chem import MolFromInchi, MolFromSmiles, MolToSmiles, MolToInchi\n\nfor smiles in [ 'c1cc(F)ccc1Cl', \n                'c1cc(Cl)ccc1F', \n                'Clc1ccc(F)cc1', \n                'Fc1ccc(Cl)cc1']:\n    mol = MolFromSmiles(smiles)\n    smiles = MolToSmiles(mol)\n    print(smiles)\n\nfor inchi in [  'InChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-1-3-6(8)4-2-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-4-2-5(7)1-3-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-4-2-5(7)1-3-6/h1-4H', \n                'InChI=1S/C6H4ClF/c7-5-2-4-6(8)3-1-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-2-4-6(8)3-1-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-3-1-5(7)2-4-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-3-1-5(7)2-4-6/h1-4H']:\n    mol = MolFromInchi(inchi)\n    inchi = MolToInchi(mol)\n    print(inchi)\n</code></pre>\n<p>Output:</p>\n<pre><code>Fc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\n</code></pre>\n<p>As you can see from the output, different InChI (and different SMILES) are actually the same thing 😐</p>",
      "rawMarkdown": "This is a very good question. In addition to your post, I will give an example where different InChI can be converted to the same molecule.\n\n```\nfrom rdkit.Chem import MolFromInchi, MolFromSmiles, MolToSmiles, MolToInchi\n\nfor smiles in [ 'c1cc(F)ccc1Cl', \n                'c1cc(Cl)ccc1F', \n                'Clc1ccc(F)cc1', \n                'Fc1ccc(Cl)cc1']:\n    mol = MolFromSmiles(smiles)\n    smiles = MolToSmiles(mol)\n    print(smiles)\n\nfor inchi in [  'InChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-1-3-6(8)4-2-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-4-2-5(7)1-3-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-4-2-5(7)1-3-6/h1-4H', \n                'InChI=1S/C6H4ClF/c7-5-2-4-6(8)3-1-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-2-4-6(8)3-1-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-3-1-5(7)2-4-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-3-1-5(7)2-4-6/h1-4H']:\n    mol = MolFromInchi(inchi)\n    inchi = MolToInchi(mol)\n    print(inchi)\n```\n\nOutput:\n```\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\n```\n\nAs you can see from the output, different InChI (and different SMILES) are actually the same thing 😐",
      "votes": null
    },
    {
      "id": "1239194",
      "postDate": "03/15/2021 14:24:39",
      "content": "<p>-Thank you, Igor !  This trial helps me with understand more. This means that different InChI expressions of one molecule(c5-1-3-6(8)4-2-5-7 etc) will go to one exact same expression (c7-5-1-3-6(8)4-2-5) by using the converter of rdkit. I think this is important thing.</p>\n<p>-My other question for InChI  definition was mentioned in  siero's discussion ([Paper]: About InChI chemical structure, <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223457)\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223457)</a>, depending on the paper referred in the page, InChI  will have one expression for one molecule (The same label always means the same substance, and the same substance always receives the same label (under the same labelling conditions). This is achieved through a well-defined procedure of obtaining canonical numbering of atoms.), Also, there is a procedure which make correct InChI expression shown in Figure 9 in the paper (Input structural data -&gt; Normalization -&gt; Canonicalization <br>\n-&gt;Serialization). this rdkit might follow similar procedure in convert between mol and InChI .</p>\n<p>-In this competition, difficulties in the problems are <br>\n(1)structure figures in training dataset have some noise. -&gt; to remove noise and get correct features<br>\n(2)one 2D structure figure might have some variations by (for example) molecule's rotation.<br>\n     even if 3D shape is same, 2D shadows will change depending on the light angle.<br>\n     all 2D shadow of the same 3D molecule have to  go to same InChI.<br>\n     (training dataset have only one of 2D shadows for one 3D molecule…how to make others..)<br>\n(3)  even if we can get correct features such as ..we can make correct 3D structure data from <br>\n     2D shadows structure, we have to make serial number for it to make InChl. I thought it is not <br>\n     so easy, but rdkit might make correct serialization. or   if some model can convert 3D model to  <br>\n     InChI which does not have perfect serial number of atom, rdkit will convert it to the perfect correct<br>\n     one.</p>\n<p>I'm not sure if 'end to end solution' such as 'figure captioning' can get these relation automatically,  <br>\nbut hope so.</p>",
      "rawMarkdown": "Thank you, Igor !  This trial helps me with understand more. This means that different InChI expressions of one molecule(c5-1-3-6(8)4-2-5-7 etc) will go to one exact same expression (c7-5-1-3-6(8)4-2-5) by using the converter of rdkit. I think this is important thing.\n\n-My other question for InChI  definition was mentioned in  siero's discussion ([Paper]: About InChI chemical structure, https://www.kaggle.com/c/bms-molecular-translation/discussion/223457), depending on the paper referred in the page, InChI  will have one expression for one molecule (The same label always means the same substance, and the same substance always receives the same label (under the same labelling conditions). This is achieved through a well-defined procedure of obtaining canonical numbering of atoms.), Also, there is a procedure which make correct InChI expression shown in Figure 9 in the paper (Input structural data -> Normalization -> Canonicalization \n->Serialization). this rdkit might follow similar procedure in convert between mol and InChI .\n\n-In this competition, difficulties in the problems are \n(1)structure figures in training dataset have some noise. -> to remove noise and get correct features\n(2)one 2D structure figure might have some variations by (for example) molecule's rotation.\n     even if 3D shape is same, 2D shadows will change depending on the light angle.\n     all 2D shadow of the same 3D molecule have to  go to same InChI.\n     (training dataset have only one of 2D shadows for one 3D molecule...how to make others..)\n(3)  even if we can get correct features such as ..we can make correct 3D structure data from \n     2D shadows structure, we have to make serial number for it to make InChl. I thought it is not \n     so easy, but rdkit might make correct serialization. or   if some model can convert 3D model to  \n     InChI which does not have perfect serial number of atom, rdkit will convert it to the perfect correct\n     one.\n\nI'm not sure if 'end to end solution' such as 'figure captioning' can get these relation automatically,  \nbut hope so.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1238360,
      "author_name": "igoryakovenko",
      "author_url": "",
      "post_date": "03/14/2021 23:33:17",
      "content": "<p>This is a very good question. In addition to your post, I will give an example where different InChI can be converted to the same molecule.</p>\n<pre><code>from rdkit.Chem import MolFromInchi, MolFromSmiles, MolToSmiles, MolToInchi\n\nfor smiles in [ 'c1cc(F)ccc1Cl', \n                'c1cc(Cl)ccc1F', \n                'Clc1ccc(F)cc1', \n                'Fc1ccc(Cl)cc1']:\n    mol = MolFromSmiles(smiles)\n    smiles = MolToSmiles(mol)\n    print(smiles)\n\nfor inchi in [  'InChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-1-3-6(8)4-2-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-4-2-5(7)1-3-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-4-2-5(7)1-3-6/h1-4H', \n                'InChI=1S/C6H4ClF/c7-5-2-4-6(8)3-1-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-2-4-6(8)3-1-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-3-1-5(7)2-4-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-3-1-5(7)2-4-6/h1-4H']:\n    mol = MolFromInchi(inchi)\n    inchi = MolToInchi(mol)\n    print(inchi)\n</code></pre>\n<p>Output:</p>\n<pre><code>Fc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\n</code></pre>\n<p>As you can see from the output, different InChI (and different SMILES) are actually the same thing 😐</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239194,
          "author_name": "kevin20181209",
          "author_url": "",
          "post_date": "03/15/2021 14:24:39",
          "content": "<p>-Thank you, Igor !  This trial helps me with understand more. This means that different InChI expressions of one molecule(c5-1-3-6(8)4-2-5-7 etc) will go to one exact same expression (c7-5-1-3-6(8)4-2-5) by using the converter of rdkit. I think this is important thing.</p>\n<p>-My other question for InChI  definition was mentioned in  siero's discussion ([Paper]: About InChI chemical structure, <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223457)\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223457)</a>, depending on the paper referred in the page, InChI  will have one expression for one molecule (The same label always means the same substance, and the same substance always receives the same label (under the same labelling conditions). This is achieved through a well-defined procedure of obtaining canonical numbering of atoms.), Also, there is a procedure which make correct InChI expression shown in Figure 9 in the paper (Input structural data -&gt; Normalization -&gt; Canonicalization <br>\n-&gt;Serialization). this rdkit might follow similar procedure in convert between mol and InChI .</p>\n<p>-In this competition, difficulties in the problems are <br>\n(1)structure figures in training dataset have some noise. -&gt; to remove noise and get correct features<br>\n(2)one 2D structure figure might have some variations by (for example) molecule's rotation.<br>\n     even if 3D shape is same, 2D shadows will change depending on the light angle.<br>\n     all 2D shadow of the same 3D molecule have to  go to same InChI.<br>\n     (training dataset have only one of 2D shadows for one 3D molecule…how to make others..)<br>\n(3)  even if we can get correct features such as ..we can make correct 3D structure data from <br>\n     2D shadows structure, we have to make serial number for it to make InChl. I thought it is not <br>\n     so easy, but rdkit might make correct serialization. or   if some model can convert 3D model to  <br>\n     InChI which does not have perfect serial number of atom, rdkit will convert it to the perfect correct<br>\n     one.</p>\n<p>I'm not sure if 'end to end solution' such as 'figure captioning' can get these relation automatically,  <br>\nbut hope so.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1237976": "If possible, I'd like to know Precise definition for InChI that how to make it from the chemical structure figure. According to the wiki page (https://en.wikipedia.org/wiki/International_Chemical_Identifier), if we see \"Format and layers\" section of it, it might not express all rules to make InChI.  if we see \"ethanol\" example in it, it is understandable, but in case of \"L-ascorbic acid\", it is difficult to understand. \"-\" of \"h2,5,7-8,10-11H,1H2\" might means \"Hydrogen bonding\". but it doesn't show explicitly in the figure.  InChI  expression need some chemical knowledge which doesn't not necessary show in the figure. Also in case of \"ethanol\" example,  number of atoms are numbered from the left, but in case of \"L-ascorbic acid, the  rule is not necessary true. If  InChI expression include some variation in one molecular structure,  examples in the train dataset is not only one precise expression,  but also other expression might be true. In case  of it, whether we will tune up out DL model to the train dataset expression or we will take on strategy to make one of correct expressions from the figure. Please advice.\n\nI briefly see  InChI  page (https://iupac.org/who-we-are/divisions/division-details/inchi/) also",
    "1238360": "This is a very good question. In addition to your post, I will give an example where different InChI can be converted to the same molecule.\n\n```\nfrom rdkit.Chem import MolFromInchi, MolFromSmiles, MolToSmiles, MolToInchi\n\nfor smiles in [ 'c1cc(F)ccc1Cl', \n                'c1cc(Cl)ccc1F', \n                'Clc1ccc(F)cc1', \n                'Fc1ccc(Cl)cc1']:\n    mol = MolFromSmiles(smiles)\n    smiles = MolToSmiles(mol)\n    print(smiles)\n\nfor inchi in [  'InChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-1-3-6(8)4-2-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-4-2-5(7)1-3-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-4-2-5(7)1-3-6/h1-4H', \n                'InChI=1S/C6H4ClF/c7-5-2-4-6(8)3-1-5/h1-4H', \n                'InChI=1S/C6H4ClF/c5-2-4-6(8)3-1-5-7/h1-4H', \n                'InChI=1S/C6H4ClF/c6-3-1-5(7)2-4-6-8/h1-4H', \n                'InChI=1S/C6H4ClF/c8-6-3-1-5(7)2-4-6/h1-4H']:\n    mol = MolFromInchi(inchi)\n    inchi = MolToInchi(mol)\n    print(inchi)\n```\n\nOutput:\n```\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nFc1ccc(Cl)cc1\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\nInChI=1S/C6H4ClF/c7-5-1-3-6(8)4-2-5/h1-4H\n```\n\nAs you can see from the output, different InChI (and different SMILES) are actually the same thing 😐",
    "1239194": "Thank you, Igor !  This trial helps me with understand more. This means that different InChI expressions of one molecule(c5-1-3-6(8)4-2-5-7 etc) will go to one exact same expression (c7-5-1-3-6(8)4-2-5) by using the converter of rdkit. I think this is important thing.\n\n-My other question for InChI  definition was mentioned in  siero's discussion ([Paper]: About InChI chemical structure, https://www.kaggle.com/c/bms-molecular-translation/discussion/223457), depending on the paper referred in the page, InChI  will have one expression for one molecule (The same label always means the same substance, and the same substance always receives the same label (under the same labelling conditions). This is achieved through a well-defined procedure of obtaining canonical numbering of atoms.), Also, there is a procedure which make correct InChI expression shown in Figure 9 in the paper (Input structural data -> Normalization -> Canonicalization \n->Serialization). this rdkit might follow similar procedure in convert between mol and InChI .\n\n-In this competition, difficulties in the problems are \n(1)structure figures in training dataset have some noise. -> to remove noise and get correct features\n(2)one 2D structure figure might have some variations by (for example) molecule's rotation.\n     even if 3D shape is same, 2D shadows will change depending on the light angle.\n     all 2D shadow of the same 3D molecule have to  go to same InChI.\n     (training dataset have only one of 2D shadows for one 3D molecule...how to make others..)\n(3)  even if we can get correct features such as ..we can make correct 3D structure data from \n     2D shadows structure, we have to make serial number for it to make InChl. I thought it is not \n     so easy, but rdkit might make correct serialization. or   if some model can convert 3D model to  \n     InChI which does not have perfect serial number of atom, rdkit will convert it to the perfect correct\n     one.\n\nI'm not sure if 'end to end solution' such as 'figure captioning' can get these relation automatically,  \nbut hope so."
  },
  "source": "meta"
}