{
  "id": 223706,
  "title": "InChI to SMILES? A good Idea?",
  "url": "/competitions/bms-molecular-translation/discussion/223706",
  "author_name": "",
  "post_date": "2021-03-05T07:21:40.762753200Z",
  "votes": 14,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm not very knowledgeable about chemistry, however, from what I've seen, SMILES format seems to be more succinct and interpretable, and thus it might be a better match for deep learning. So, my question is, do you recommend converting the labels to SMILES, training on those expressions, and just converting my predictions to InChI later?   </p>\n<p>EDIT: It looks like antifact made a SMILES dataset using RDKit<br>\nHis Post: <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223568\" target=\"_blank\">bms-molecular-translation</a><br>\nThe Dataset: <a href=\"https://www.kaggle.com/antifact/molecular-translation-smiles-csv\" target=\"_blank\">Molecular-Translation-smiles-csv</a></p>",
  "messages": [
    {
      "id": "1227086",
      "postDate": "03/05/2021 07:21:40",
      "content": "<p>I'm not very knowledgeable about chemistry, however, from what I've seen, SMILES format seems to be more succinct and interpretable, and thus it might be a better match for deep learning. So, my question is, do you recommend converting the labels to SMILES, training on those expressions, and just converting my predictions to InChI later?   </p>\n<p>EDIT: It looks like antifact made a SMILES dataset using RDKit<br>\nHis Post: <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223568\" target=\"_blank\">bms-molecular-translation</a><br>\nThe Dataset: <a href=\"https://www.kaggle.com/antifact/molecular-translation-smiles-csv\" target=\"_blank\">Molecular-Translation-smiles-csv</a></p>",
      "rawMarkdown": "I'm not very knowledgeable about chemistry, however, from what I've seen, SMILES format seems to be more succinct and interpretable, and thus it might be a better match for deep learning. So, my question is, do you recommend converting the labels to SMILES, training on those expressions, and just converting my predictions to InChI later?   \n\nEDIT: It looks like antifact made a SMILES dataset using RDKit\nHis Post: [bms-molecular-translation]( https://www.kaggle.com/c/bms-molecular-translation/discussion/223568)\nThe Dataset: [Molecular-Translation-smiles-csv](https://www.kaggle.com/antifact/molecular-translation-smiles-csv)",
      "votes": null
    },
    {
      "id": "1227339",
      "postDate": "03/05/2021 12:39:28",
      "content": "<p>I am not chemist but just reading from internet it seems like you can have multiple SMILES representation per structure e.g For example, CCO, OCC and C(O)C all specify the structure of ethanol. However InChI is quite unique and it’s guaranteed you will have 1 representation per structure. </p>\n<p>1) Also I am curious weather multiple SMILES are mapped to the same InChi ?</p>\n<p>2) If one would train using SMILE and let’s says your distance metric is 2. (predictions and valid) if you translate prediction and validation to InChi and measure the metric again. How much difference it will be ? </p>",
      "rawMarkdown": "I am not chemist but just reading from internet it seems like you can have multiple SMILES representation per structure e.g For example, CCO, OCC and C(O)C all specify the structure of ethanol. However InChI is quite unique and it’s guaranteed you will have 1 representation per structure. \n\n1) Also I am curious weather multiple SMILES are mapped to the same InChi ?\n\n2) If one would train using SMILE and let’s says your distance metric is 2. (predictions and valid) if you translate prediction and validation to InChi and measure the metric again. How much difference it will be ?",
      "votes": null
    },
    {
      "id": "1227429",
      "postDate": "03/05/2021 14:17:33",
      "content": "<ol>\n<li><p>I think the answer should be no because there maybe molecule M1 and M2 having common SMILES representation S1 and S2. In that case, the S1 and S2 can be mapped to InChI representations of both M1 and M2. This is hypothetical but doesn't look impossible to me.</p></li>\n<li><p>That's a tough question to answer without carrying out experiments I think</p></li>\n</ol>",
      "rawMarkdown": "1. I think the answer should be no because there maybe molecule M1 and M2 having common SMILES representation S1 and S2. In that case, the S1 and S2 can be mapped to InChI representations of both M1 and M2. This is hypothetical but doesn't look impossible to me.\n\n2. That's a tough question to answer without carrying out experiments I think",
      "votes": null
    },
    {
      "id": "1227581",
      "postDate": "03/05/2021 16:56:25",
      "content": "<p>Actually i just did quick test… its not a good idea… in some cases if your <code>SMILE</code> prediction is a bit wrong it wont always translate to InChI using <code>rdkit</code> … unless you can predict perfect <code>SMILE</code>=) </p>",
      "rawMarkdown": "Actually i just did quick test... its not a good idea... in some cases if your `SMILE` prediction is a bit wrong it wont always translate to InChI using `rdkit` ... unless you can predict perfect `SMILE`=)",
      "votes": null
    },
    {
      "id": "1227596",
      "postDate": "03/05/2021 17:03:15",
      "content": "<pre><code>from rdkit import Chem, DataStructs\ndef inch_to_smiles(name):\n    '''converts InChi name to smile string'''\n    mol = Chem.inchi.MolFromInchi(name)\n    smile_string = Chem.MolToSmiles(mol)\n    return smile_string\n\ndef smiles_to_inch(name):\n    '''converts smile string to InChi'''\n    chem_smile = Chem.MolFromSmiles(name)\n    return Chem.inchi.MolToInchi(chem_smile)\n</code></pre>\n<p>result</p>\n<pre><code>inch = 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'\nsm = inch_to_smiles(chn)\nsm\n&gt;'Cc1ccc(SCC(C)C)c(C(C)O)c1'\n\n#testing if sm to inch function is correct\nassert smiles_to_inch(sm)==inch\n\n#making small mistake (at the end 1 -&gt; 2)\nsm_corup = 'Cc1ccc(SCC(C)C)c(C(C)O)c2'\nsmiles_to_inch(sm_corup)\n&gt; ERROR\n</code></pre>",
      "rawMarkdown": "```\nfrom rdkit import Chem, DataStructs\ndef inch_to_smiles(name):\n    '''converts InChi name to smile string'''\n    mol = Chem.inchi.MolFromInchi(name)\n    smile_string = Chem.MolToSmiles(mol)\n    return smile_string\n\ndef smiles_to_inch(name):\n    '''converts smile string to InChi'''\n    chem_smile = Chem.MolFromSmiles(name)\n    return Chem.inchi.MolToInchi(chem_smile)\n\n```\nresult\n\n```\ninch = 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'\nsm = inch_to_smiles(chn)\nsm\n>'Cc1ccc(SCC(C)C)c(C(C)O)c1'\n\n#testing if sm to inch function is correct\nassert smiles_to_inch(sm)==inch\n\n#making small mistake (at the end 1 -> 2)\nsm_corup = 'Cc1ccc(SCC(C)C)c(C(C)O)c2'\nsmiles_to_inch(sm_corup)\n> ERROR\n```",
      "votes": null
    },
    {
      "id": "1227640",
      "postDate": "03/05/2021 17:47:37",
      "content": "<p>Yeah good bit of work confirming that, maybe not a good idea after all 😅</p>",
      "rawMarkdown": "Yeah good bit of work confirming that, maybe not a good idea after all 😅",
      "votes": null
    },
    {
      "id": "1229361",
      "postDate": "03/07/2021 09:34:54",
      "content": "<p>The SMILES strings are generated on the basis of the valence bond (VB) theory. So, for some exceptional chemicals, additional \"formal charges\" need to be introduced. (Sometimes it is very difficult.) If this can be done correctly, using SMILES is a good idea, I believe. If the InChI strings are directly generated,  there is no need to add the additional \"formal charges\". 😄</p>",
      "rawMarkdown": "The SMILES strings are generated on the basis of the valence bond (VB) theory. So, for some exceptional chemicals, additional \"formal charges\" need to be introduced. (Sometimes it is very difficult.) If this can be done correctly, using SMILES is a good idea, I believe. If the InChI strings are directly generated,  there is no need to add the additional \"formal charges\". 😄",
      "votes": null
    },
    {
      "id": "1238107",
      "postDate": "03/14/2021 17:15:08",
      "content": "<p>I think this is a really good question to ask! As I showed in my <a href=\"https://www.kaggle.com/rssrwn/inchi-tokenisation\" target=\"_blank\">notebook</a> SMILES are significantly shorter than InChIs for this dataset, which will make training faster and may even help the model to generate more accurate predictions. However, one downside is that, unlike InChIs, if the model were to get one token wrong when constructing the SMILES it could lead to a completely different InChI and so the LD would be very high. Another possible upside of using SMILES is that they can be augmented, which may help to improve a model's generalisation.</p>",
      "rawMarkdown": "I think this is a really good question to ask! As I showed in my [notebook](https://www.kaggle.com/rssrwn/inchi-tokenisation) SMILES are significantly shorter than InChIs for this dataset, which will make training faster and may even help the model to generate more accurate predictions. However, one downside is that, unlike InChIs, if the model were to get one token wrong when constructing the SMILES it could lead to a completely different InChI and so the LD would be very high. Another possible upside of using SMILES is that they can be augmented, which may help to improve a model's generalisation.",
      "votes": null
    },
    {
      "id": "1238302",
      "postDate": "03/14/2021 21:18:06",
      "content": "<p>The fact that it's relatively easy for a machine learning model to generate invalid SMILES is a known issue. To address this issue, one solution is to use an alternative representation such as SELFIES: <a href=\"https://github.com/aspuru-guzik-group/selfies\" target=\"_blank\">https://github.com/aspuru-guzik-group/selfies</a><br>\nSELFIES can be converted to/from SMILES…. so one could convert the training data from InChI-&gt;SMILES-&gt;SELFIES, train a model, then for the predictions go in the other direction SELFIES-&gt;SMILES-&gt;InChI</p>",
      "rawMarkdown": "The fact that it's relatively easy for a machine learning model to generate invalid SMILES is a known issue. To address this issue, one solution is to use an alternative representation such as SELFIES: https://github.com/aspuru-guzik-group/selfies\nSELFIES can be converted to/from SMILES.... so one could convert the training data from InChI->SMILES->SELFIES, train a model, then for the predictions go in the other direction SELFIES->SMILES->InChI",
      "votes": null
    },
    {
      "id": "1241302",
      "postDate": "03/17/2021 01:50:41",
      "content": "<p>why don't you do this:</p>\n<ol>\n<li><p>train a model</p></li>\n<li><p>use the model to predict the test, you end with prediction like<br>\ninchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>use rdkit to generate new train image samples for inchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>retrain and repeat</p></li>\n</ol>\n<hr>\n<p>note that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.<br>\nusing the generator should be part of your solution.</p>\n<p>it is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.</p>\n<p>data is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to memorize them</p>",
      "rawMarkdown": "why don't you do this:\n\n1. train a model\n\n2. use the model to predict the test, you end with prediction like\ninchi1,inchi2,inchi3 ...inchi1000\n\n3. use rdkit to generate new train image samples for inchi1,inchi2,inchi3 ...inchi1000\n\n4. retrain and repeat\n\n---\n\nnote that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.\nusing the generator should be part of your solution.\n\nit is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.\n\ndata is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to memorize them",
      "votes": null
    },
    {
      "id": "1241385",
      "postDate": "03/17/2021 02:45:14",
      "content": "<p>Are you suggesting retraining again on the data generated by rdkit? Isn't the distribution strikingly different from the given dataset?</p>",
      "rawMarkdown": "Are you suggesting retraining again on the data generated by rdkit? Isn't the distribution strikingly different from the given dataset?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1227339,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "03/05/2021 12:39:28",
      "content": "<p>I am not chemist but just reading from internet it seems like you can have multiple SMILES representation per structure e.g For example, CCO, OCC and C(O)C all specify the structure of ethanol. However InChI is quite unique and it’s guaranteed you will have 1 representation per structure. </p>\n<p>1) Also I am curious weather multiple SMILES are mapped to the same InChi ?</p>\n<p>2) If one would train using SMILE and let’s says your distance metric is 2. (predictions and valid) if you translate prediction and validation to InChi and measure the metric again. How much difference it will be ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1227429,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/05/2021 14:17:33",
          "content": "<ol>\n<li><p>I think the answer should be no because there maybe molecule M1 and M2 having common SMILES representation S1 and S2. In that case, the S1 and S2 can be mapped to InChI representations of both M1 and M2. This is hypothetical but doesn't look impossible to me.</p></li>\n<li><p>That's a tough question to answer without carrying out experiments I think</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1227581,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/05/2021 16:56:25",
          "content": "<p>Actually i just did quick test… its not a good idea… in some cases if your <code>SMILE</code> prediction is a bit wrong it wont always translate to InChI using <code>rdkit</code> … unless you can predict perfect <code>SMILE</code>=) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1227596,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/05/2021 17:03:15",
          "content": "<pre><code>from rdkit import Chem, DataStructs\ndef inch_to_smiles(name):\n    '''converts InChi name to smile string'''\n    mol = Chem.inchi.MolFromInchi(name)\n    smile_string = Chem.MolToSmiles(mol)\n    return smile_string\n\ndef smiles_to_inch(name):\n    '''converts smile string to InChi'''\n    chem_smile = Chem.MolFromSmiles(name)\n    return Chem.inchi.MolToInchi(chem_smile)\n</code></pre>\n<p>result</p>\n<pre><code>inch = 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'\nsm = inch_to_smiles(chn)\nsm\n&gt;'Cc1ccc(SCC(C)C)c(C(C)O)c1'\n\n#testing if sm to inch function is correct\nassert smiles_to_inch(sm)==inch\n\n#making small mistake (at the end 1 -&gt; 2)\nsm_corup = 'Cc1ccc(SCC(C)C)c(C(C)O)c2'\nsmiles_to_inch(sm_corup)\n&gt; ERROR\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1227640,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/05/2021 17:47:37",
          "content": "<p>Yeah good bit of work confirming that, maybe not a good idea after all 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238302,
          "author_name": "infy2097",
          "author_url": "",
          "post_date": "03/14/2021 21:18:06",
          "content": "<p>The fact that it's relatively easy for a machine learning model to generate invalid SMILES is a known issue. To address this issue, one solution is to use an alternative representation such as SELFIES: <a href=\"https://github.com/aspuru-guzik-group/selfies\" target=\"_blank\">https://github.com/aspuru-guzik-group/selfies</a><br>\nSELFIES can be converted to/from SMILES…. so one could convert the training data from InChI-&gt;SMILES-&gt;SELFIES, train a model, then for the predictions go in the other direction SELFIES-&gt;SMILES-&gt;InChI</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1229361,
      "author_name": "hiroshisakiyama",
      "author_url": "",
      "post_date": "03/07/2021 09:34:54",
      "content": "<p>The SMILES strings are generated on the basis of the valence bond (VB) theory. So, for some exceptional chemicals, additional \"formal charges\" need to be introduced. (Sometimes it is very difficult.) If this can be done correctly, using SMILES is a good idea, I believe. If the InChI strings are directly generated,  there is no need to add the additional \"formal charges\". 😄</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1238107,
      "author_name": "rssrwn",
      "author_url": "",
      "post_date": "03/14/2021 17:15:08",
      "content": "<p>I think this is a really good question to ask! As I showed in my <a href=\"https://www.kaggle.com/rssrwn/inchi-tokenisation\" target=\"_blank\">notebook</a> SMILES are significantly shorter than InChIs for this dataset, which will make training faster and may even help the model to generate more accurate predictions. However, one downside is that, unlike InChIs, if the model were to get one token wrong when constructing the SMILES it could lead to a completely different InChI and so the LD would be very high. Another possible upside of using SMILES is that they can be augmented, which may help to improve a model's generalisation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1241302,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/17/2021 01:50:41",
      "content": "<p>why don't you do this:</p>\n<ol>\n<li><p>train a model</p></li>\n<li><p>use the model to predict the test, you end with prediction like<br>\ninchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>use rdkit to generate new train image samples for inchi1,inchi2,inchi3 …inchi1000</p></li>\n<li><p>retrain and repeat</p></li>\n</ol>\n<hr>\n<p>note that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.<br>\nusing the generator should be part of your solution.</p>\n<p>it is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.</p>\n<p>data is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to memorize them</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241385,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/17/2021 02:45:14",
          "content": "<p>Are you suggesting retraining again on the data generated by rdkit? Isn't the distribution strikingly different from the given dataset?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1227086": "I'm not very knowledgeable about chemistry, however, from what I've seen, SMILES format seems to be more succinct and interpretable, and thus it might be a better match for deep learning. So, my question is, do you recommend converting the labels to SMILES, training on those expressions, and just converting my predictions to InChI later?   \n\nEDIT: It looks like antifact made a SMILES dataset using RDKit\nHis Post: [bms-molecular-translation]( https://www.kaggle.com/c/bms-molecular-translation/discussion/223568)\nThe Dataset: [Molecular-Translation-smiles-csv](https://www.kaggle.com/antifact/molecular-translation-smiles-csv)",
    "1227339": "I am not chemist but just reading from internet it seems like you can have multiple SMILES representation per structure e.g For example, CCO, OCC and C(O)C all specify the structure of ethanol. However InChI is quite unique and it’s guaranteed you will have 1 representation per structure. \n\n1) Also I am curious weather multiple SMILES are mapped to the same InChi ?\n\n2) If one would train using SMILE and let’s says your distance metric is 2. (predictions and valid) if you translate prediction and validation to InChi and measure the metric again. How much difference it will be ?",
    "1227429": "1. I think the answer should be no because there maybe molecule M1 and M2 having common SMILES representation S1 and S2. In that case, the S1 and S2 can be mapped to InChI representations of both M1 and M2. This is hypothetical but doesn't look impossible to me.\n\n2. That's a tough question to answer without carrying out experiments I think",
    "1227581": "Actually i just did quick test... its not a good idea... in some cases if your `SMILE` prediction is a bit wrong it wont always translate to InChI using `rdkit` ... unless you can predict perfect `SMILE`=)",
    "1227596": "```\nfrom rdkit import Chem, DataStructs\ndef inch_to_smiles(name):\n    '''converts InChi name to smile string'''\n    mol = Chem.inchi.MolFromInchi(name)\n    smile_string = Chem.MolToSmiles(mol)\n    return smile_string\n\ndef smiles_to_inch(name):\n    '''converts smile string to InChi'''\n    chem_smile = Chem.MolFromSmiles(name)\n    return Chem.inchi.MolToInchi(chem_smile)\n\n```\nresult\n\n```\ninch = 'InChI=1S/C13H20OS/c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14/h5-7,9,11,14H,8H2,1-4H3'\nsm = inch_to_smiles(chn)\nsm\n>'Cc1ccc(SCC(C)C)c(C(C)O)c1'\n\n#testing if sm to inch function is correct\nassert smiles_to_inch(sm)==inch\n\n#making small mistake (at the end 1 -> 2)\nsm_corup = 'Cc1ccc(SCC(C)C)c(C(C)O)c2'\nsmiles_to_inch(sm_corup)\n> ERROR\n```",
    "1227640": "Yeah good bit of work confirming that, maybe not a good idea after all 😅",
    "1229361": "The SMILES strings are generated on the basis of the valence bond (VB) theory. So, for some exceptional chemicals, additional \"formal charges\" need to be introduced. (Sometimes it is very difficult.) If this can be done correctly, using SMILES is a good idea, I believe. If the InChI strings are directly generated,  there is no need to add the additional \"formal charges\". 😄",
    "1238107": "I think this is a really good question to ask! As I showed in my [notebook](https://www.kaggle.com/rssrwn/inchi-tokenisation) SMILES are significantly shorter than InChIs for this dataset, which will make training faster and may even help the model to generate more accurate predictions. However, one downside is that, unlike InChIs, if the model were to get one token wrong when constructing the SMILES it could lead to a completely different InChI and so the LD would be very high. Another possible upside of using SMILES is that they can be augmented, which may help to improve a model's generalisation.",
    "1238302": "The fact that it's relatively easy for a machine learning model to generate invalid SMILES is a known issue. To address this issue, one solution is to use an alternative representation such as SELFIES: https://github.com/aspuru-guzik-group/selfies\nSELFIES can be converted to/from SMILES.... so one could convert the training data from InChI->SMILES->SELFIES, train a model, then for the predictions go in the other direction SELFIES->SMILES->InChI",
    "1241302": "why don't you do this:\n\n1. train a model\n\n2. use the model to predict the test, you end with prediction like\ninchi1,inchi2,inchi3 ...inchi1000\n\n3. use rdkit to generate new train image samples for inchi1,inchi2,inchi3 ...inchi1000\n\n4. retrain and repeat\n\n---\n\nnote that we have a perfect generator to go from text to image using rdkit. we are doing reverse engineering.\nusing the generator should be part of your solution.\n\nit is like computer graphics. we have a render to make the image. now we are going from image backward to 3d model.\n\ndata is the key to this competition. we have infinite data. (we have all test and train data). just think of a way to memorize them",
    "1241385": "Are you suggesting retraining again on the data generated by rdkit? Isn't the distribution strikingly different from the given dataset?"
  },
  "source": "meta"
}