{
  "id": 501663,
  "title": "SmilesEnumerator: Augmentation tool for SMILES",
  "url": "/competitions/leash-BELKA/discussion/501663",
  "author_name": "",
  "post_date": "2024-05-10T08:53:33.679750600Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>SmilesEnumerator: Augmentation for SMILES</h1>\n<blockquote>\n  <p>SMILES enumeration is the process of writing out all possible SMILES forms of a molecule. <br>\n  It's a useful technique for data augmentation before sequence based modeling of molecules. </p>\n</blockquote>\n<h3>You can see how to use this tool in <a href=\"https://www.kaggle.com/code/yusaku5739/smilesenumerator-augmentation-tool-for-smiles\" target=\"_blank\">my notebook</a></h3>\n<p><br> <br>\nOne of the way to address this competition is feeding SMILES into transformer or 1dcnn based models.  <br>\nSMILES can take many forms per molecule.  <br>\nIn this notebook, I introduce you SmiilesEnumerator, generating various smiles from a smiles.   <br>\nYou can use it for augmentation of smiles.    </p>\n<p>If you want more information, please visit (<a href=\"https://github.com/EBjerrum/SMILES-enumeration\" target=\"_blank\">https://github.com/EBjerrum/SMILES-enumeration</a>)</p>\n<h2>↓The example of the SmilesEnumerator</h2>\n<h3>example smiles</h3>\n<blockquote>\n  <p>C#CCCC<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@H</a>nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]</p>\n</blockquote>\n<h3>Augmented smiles</h3>\n<blockquote>\n  <p>c1cc(Nc2nc(N<a href=\"CCCC#C\" target=\"_blank\">C@@H</a>C(N[Dy])=O)nc(Nc3ccc(C=C)cc3)n2)ccc1C=C<br>\n  C(CC<a href=\"C(N[Dy])=O\" target=\"_blank\">C@@H</a>Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C#C<br>\n  c1(Nc2nc(Nc3ccc(C=C)cc3)nc(N<a href=\"C(=O)N[Dy]\" target=\"_blank\">C@H</a>CCCC#C)n2)ccc(C=C)cc1<br>\n  N(c1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)<a href=\"C(=O)N[Dy]\" target=\"_blank\">C@H</a>CCCC#C<br>\n  C(CCC#C)<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@H</a>nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]<br>\n  O=C(<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@@H</a>nc(Nc2ccc(C=C)cc2)n1)CCCC#C)N[Dy]<br>\n  c1c(Nc2nc(Nc3ccc(C=C)cc3)nc(N<a href=\"CCCC#C\" target=\"_blank\">C@@H</a>C(=O)N[Dy])n2)ccc(C=C)c1<br>\n  [Dy]NC(<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@@H</a>nc(Nc2ccc(C=C)cc2)n1)CCCC#C)=O<br>\n  C=Cc1ccc(Nc2nc(N<a href=\"C(N[Dy])=O\" target=\"_blank\">C@H</a>CCCC#C)nc(Nc3ccc(C=C)cc3)n2)cc1</p>\n</blockquote>\n<h3>As you can see from the following image, the raw and augmented smiles are the same.</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F6b48b565530d97d57a22730f19c433c4%2F2024-05-10%20084755.png?generation=1715330982097173&amp;alt=media\"></p>\n<h3>Please Upvote if you Find this Useful :)</h3>",
  "messages": [
    {
      "id": "2804890",
      "postDate": "05/10/2024 08:53:33",
      "content": "<h1>SmilesEnumerator: Augmentation for SMILES</h1>\n<blockquote>\n  <p>SMILES enumeration is the process of writing out all possible SMILES forms of a molecule. <br>\n  It's a useful technique for data augmentation before sequence based modeling of molecules. </p>\n</blockquote>\n<h3>You can see how to use this tool in <a href=\"https://www.kaggle.com/code/yusaku5739/smilesenumerator-augmentation-tool-for-smiles\" target=\"_blank\">my notebook</a></h3>\n<p><br> <br>\nOne of the way to address this competition is feeding SMILES into transformer or 1dcnn based models.  <br>\nSMILES can take many forms per molecule.  <br>\nIn this notebook, I introduce you SmiilesEnumerator, generating various smiles from a smiles.   <br>\nYou can use it for augmentation of smiles.    </p>\n<p>If you want more information, please visit (<a href=\"https://github.com/EBjerrum/SMILES-enumeration\" target=\"_blank\">https://github.com/EBjerrum/SMILES-enumeration</a>)</p>\n<h2>↓The example of the SmilesEnumerator</h2>\n<h3>example smiles</h3>\n<blockquote>\n  <p>C#CCCC<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@H</a>nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]</p>\n</blockquote>\n<h3>Augmented smiles</h3>\n<blockquote>\n  <p>c1cc(Nc2nc(N<a href=\"CCCC#C\" target=\"_blank\">C@@H</a>C(N[Dy])=O)nc(Nc3ccc(C=C)cc3)n2)ccc1C=C<br>\n  C(CC<a href=\"C(N[Dy])=O\" target=\"_blank\">C@@H</a>Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C#C<br>\n  c1(Nc2nc(Nc3ccc(C=C)cc3)nc(N<a href=\"C(=O)N[Dy]\" target=\"_blank\">C@H</a>CCCC#C)n2)ccc(C=C)cc1<br>\n  N(c1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)<a href=\"C(=O)N[Dy]\" target=\"_blank\">C@H</a>CCCC#C<br>\n  C(CCC#C)<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@H</a>nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]<br>\n  O=C(<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@@H</a>nc(Nc2ccc(C=C)cc2)n1)CCCC#C)N[Dy]<br>\n  c1c(Nc2nc(Nc3ccc(C=C)cc3)nc(N<a href=\"CCCC#C\" target=\"_blank\">C@@H</a>C(=O)N[Dy])n2)ccc(C=C)c1<br>\n  [Dy]NC(<a href=\"Nc1nc(Nc2ccc(C=C)cc2\" target=\"_blank\">C@@H</a>nc(Nc2ccc(C=C)cc2)n1)CCCC#C)=O<br>\n  C=Cc1ccc(Nc2nc(N<a href=\"C(N[Dy])=O\" target=\"_blank\">C@H</a>CCCC#C)nc(Nc3ccc(C=C)cc3)n2)cc1</p>\n</blockquote>\n<h3>As you can see from the following image, the raw and augmented smiles are the same.</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F6b48b565530d97d57a22730f19c433c4%2F2024-05-10%20084755.png?generation=1715330982097173&amp;alt=media\"></p>\n<h3>Please Upvote if you Find this Useful :)</h3>",
      "rawMarkdown": "# SmilesEnumerator: Augmentation for SMILES\n\n> SMILES enumeration is the process of writing out all possible SMILES forms of a molecule. \n> It's a useful technique for data augmentation before sequence based modeling of molecules. \n\n### You can see how to use this tool in [my notebook](https://www.kaggle.com/code/yusaku5739/smilesenumerator-augmentation-tool-for-smiles)  \n<br> \nOne of the way to address this competition is feeding SMILES into transformer or 1dcnn based models.  \nSMILES can take many forms per molecule.  \nIn this notebook, I introduce you SmiilesEnumerator, generating various smiles from a smiles.   \nYou can use it for augmentation of smiles.    \n  \nIf you want more information, please visit (https://github.com/EBjerrum/SMILES-enumeration)\n\n## ↓The example of the SmilesEnumerator \n\n### example smiles\n> C#CCCC[C@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]\n\n### Augmented smiles\n> c1cc(Nc2nc(N[C@@H](CCCC#C)C(N[Dy])=O)nc(Nc3ccc(C=C)cc3)n2)ccc1C=C\n> C(CC[C@@H](C(N[Dy])=O)Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C#C\n> c1(Nc2nc(Nc3ccc(C=C)cc3)nc(N[C@H](C(=O)N[Dy])CCCC#C)n2)ccc(C=C)cc1\n> N(c1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)[C@H](C(=O)N[Dy])CCCC#C\n> C(CCC#C)[C@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]\n> O=C([C@@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)CCCC#C)N[Dy]\n> c1c(Nc2nc(Nc3ccc(C=C)cc3)nc(N[C@@H](CCCC#C)C(=O)N[Dy])n2)ccc(C=C)c1\n> [Dy]NC([C@@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)CCCC#C)=O\n> C=Cc1ccc(Nc2nc(N[C@H](C(N[Dy])=O)CCCC#C)nc(Nc3ccc(C=C)cc3)n2)cc1\n\n### As you can see from the following image, the raw and augmented smiles are the same.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F6b48b565530d97d57a22730f19c433c4%2F2024-05-10%20084755.png?generation=1715330982097173&alt=media)\n\n### Please Upvote if you Find this Useful :)",
      "votes": null
    },
    {
      "id": "2804967",
      "postDate": "05/10/2024 09:32:21",
      "content": "<p>I tried smiles augmentation but haven't seen much improvement - I wonder if our dataset has all smiles in canonical format, do you know? Have you seen some improvement yourself? </p>",
      "rawMarkdown": "I tried smiles augmentation but haven't seen much improvement - I wonder if our dataset has all smiles in canonical format, do you know? Have you seen some improvement yourself?",
      "votes": null
    },
    {
      "id": "2806305",
      "postDate": "05/11/2024 03:11:31",
      "content": "<p>How is the speed of the transformation? </p>",
      "rawMarkdown": "How is the speed of the transformation?",
      "votes": null
    },
    {
      "id": "2806897",
      "postDate": "05/11/2024 11:32:28",
      "content": "<p>Yes it seems that all the smiles in the train and test set are in the canonical form,<br>\nhere is a better way to generate random smiles from the molecule with Rdkit : Chem.MolToSmiles(mol, doRandom=True)</p>",
      "rawMarkdown": "Yes it seems that all the smiles in the train and test set are in the canonical form,\nhere is a better way to generate random smiles from the molecule with Rdkit : Chem.MolToSmiles(mol, doRandom=True)",
      "votes": null
    },
    {
      "id": "2809004",
      "postDate": "05/12/2024 13:40:19",
      "content": "<p><a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> , have you tried? Because in my case doRandom with 1/8 augmentation probability led to worth results</p>",
      "rawMarkdown": "ahmedelfazouan , have you tried? Because in my case doRandom with 1/8 augmentation probability led to worth results",
      "votes": null
    },
    {
      "id": "2809100",
      "postDate": "05/12/2024 14:20:24",
      "content": "<p>I tried pretraining on random smiles then finetune on canonical ones, it improved slightly my cv but lb was the same, I didn't try augmentation</p>",
      "rawMarkdown": "I tried pretraining on random smiles then finetune on canonical ones, it improved slightly my cv but lb was the same, I didn't try augmentation",
      "votes": null
    },
    {
      "id": "2812635",
      "postDate": "05/14/2024 10:36:15",
      "content": "<p>If you are using SMILES in the way most of the cheminformatics community does, then the SMILES is translated into a set of features, descriptors, or fingerprints which are themselves deterministically dependent on the molecular graph (and stereochemistry, where relevant). So either your rearranged SMILES generates the same feature values as the original (putatively canonical) one, in which case it is a data duplication. Or it generates different ones, in which case you need to consider whether it is less representative of the real molecule (perhaps even \"wrong\", in simple terms).</p>\n<p>There may be some unusual use cases, perhaps if treating a SMILES like a word and tokenizing it, where a duplication strategy might be helpful.</p>",
      "rawMarkdown": "If you are using SMILES in the way most of the cheminformatics community does, then the SMILES is translated into a set of features, descriptors, or fingerprints which are themselves deterministically dependent on the molecular graph (and stereochemistry, where relevant). So either your rearranged SMILES generates the same feature values as the original (putatively canonical) one, in which case it is a data duplication. Or it generates different ones, in which case you need to consider whether it is less representative of the real molecule (perhaps even \"wrong\", in simple terms).\n\nThere may be some unusual use cases, perhaps if treating a SMILES like a word and tokenizing it, where a duplication strategy might be helpful.",
      "votes": null
    },
    {
      "id": "2813906",
      "postDate": "05/15/2024 04:53:47",
      "content": "<p>I would really appreciate something that could do reversals and other transformations automatically. Not as interested in random variations, though I may still try it. </p>",
      "rawMarkdown": "I would really appreciate something that could do reversals and other transformations automatically. Not as interested in random variations, though I may still try it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2804967,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "05/10/2024 09:32:21",
      "content": "<p>I tried smiles augmentation but haven't seen much improvement - I wonder if our dataset has all smiles in canonical format, do you know? Have you seen some improvement yourself? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2806897,
          "author_name": "ahmedelfazouan",
          "author_url": "",
          "post_date": "05/11/2024 11:32:28",
          "content": "<p>Yes it seems that all the smiles in the train and test set are in the canonical form,<br>\nhere is a better way to generate random smiles from the molecule with Rdkit : Chem.MolToSmiles(mol, doRandom=True)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2809004,
              "author_name": "ubique",
              "author_url": "",
              "post_date": "05/12/2024 13:40:19",
              "content": "<p><a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> , have you tried? Because in my case doRandom with 1/8 augmentation probability led to worth results</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2809100,
                  "author_name": "ahmedelfazouan",
                  "author_url": "",
                  "post_date": "05/12/2024 14:20:24",
                  "content": "<p>I tried pretraining on random smiles then finetune on canonical ones, it improved slightly my cv but lb was the same, I didn't try augmentation</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2806305,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "05/11/2024 03:11:31",
      "content": "<p>How is the speed of the transformation? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2812635,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "05/14/2024 10:36:15",
      "content": "<p>If you are using SMILES in the way most of the cheminformatics community does, then the SMILES is translated into a set of features, descriptors, or fingerprints which are themselves deterministically dependent on the molecular graph (and stereochemistry, where relevant). So either your rearranged SMILES generates the same feature values as the original (putatively canonical) one, in which case it is a data duplication. Or it generates different ones, in which case you need to consider whether it is less representative of the real molecule (perhaps even \"wrong\", in simple terms).</p>\n<p>There may be some unusual use cases, perhaps if treating a SMILES like a word and tokenizing it, where a duplication strategy might be helpful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2813906,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "05/15/2024 04:53:47",
      "content": "<p>I would really appreciate something that could do reversals and other transformations automatically. Not as interested in random variations, though I may still try it. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2804890": "# SmilesEnumerator: Augmentation for SMILES\n\n> SMILES enumeration is the process of writing out all possible SMILES forms of a molecule. \n> It's a useful technique for data augmentation before sequence based modeling of molecules. \n\n### You can see how to use this tool in [my notebook](https://www.kaggle.com/code/yusaku5739/smilesenumerator-augmentation-tool-for-smiles)  \n<br> \nOne of the way to address this competition is feeding SMILES into transformer or 1dcnn based models.  \nSMILES can take many forms per molecule.  \nIn this notebook, I introduce you SmiilesEnumerator, generating various smiles from a smiles.   \nYou can use it for augmentation of smiles.    \n  \nIf you want more information, please visit (https://github.com/EBjerrum/SMILES-enumeration)\n\n## ↓The example of the SmilesEnumerator \n\n### example smiles\n> C#CCCC[C@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]\n\n### Augmented smiles\n> c1cc(Nc2nc(N[C@@H](CCCC#C)C(N[Dy])=O)nc(Nc3ccc(C=C)cc3)n2)ccc1C=C\n> C(CC[C@@H](C(N[Dy])=O)Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C#C\n> c1(Nc2nc(Nc3ccc(C=C)cc3)nc(N[C@H](C(=O)N[Dy])CCCC#C)n2)ccc(C=C)cc1\n> N(c1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)[C@H](C(=O)N[Dy])CCCC#C\n> C(CCC#C)[C@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)C(=O)N[Dy]\n> O=C([C@@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)CCCC#C)N[Dy]\n> c1c(Nc2nc(Nc3ccc(C=C)cc3)nc(N[C@@H](CCCC#C)C(=O)N[Dy])n2)ccc(C=C)c1\n> [Dy]NC([C@@H](Nc1nc(Nc2ccc(C=C)cc2)nc(Nc2ccc(C=C)cc2)n1)CCCC#C)=O\n> C=Cc1ccc(Nc2nc(N[C@H](C(N[Dy])=O)CCCC#C)nc(Nc3ccc(C=C)cc3)n2)cc1\n\n### As you can see from the following image, the raw and augmented smiles are the same.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F6b48b565530d97d57a22730f19c433c4%2F2024-05-10%20084755.png?generation=1715330982097173&alt=media)\n\n### Please Upvote if you Find this Useful :)",
    "2804967": "I tried smiles augmentation but haven't seen much improvement - I wonder if our dataset has all smiles in canonical format, do you know? Have you seen some improvement yourself?",
    "2806305": "How is the speed of the transformation?",
    "2806897": "Yes it seems that all the smiles in the train and test set are in the canonical form,\nhere is a better way to generate random smiles from the molecule with Rdkit : Chem.MolToSmiles(mol, doRandom=True)",
    "2809004": "ahmedelfazouan , have you tried? Because in my case doRandom with 1/8 augmentation probability led to worth results",
    "2809100": "I tried pretraining on random smiles then finetune on canonical ones, it improved slightly my cv but lb was the same, I didn't try augmentation",
    "2812635": "If you are using SMILES in the way most of the cheminformatics community does, then the SMILES is translated into a set of features, descriptors, or fingerprints which are themselves deterministically dependent on the molecular graph (and stereochemistry, where relevant). So either your rearranged SMILES generates the same feature values as the original (putatively canonical) one, in which case it is a data duplication. Or it generates different ones, in which case you need to consider whether it is less representative of the real molecule (perhaps even \"wrong\", in simple terms).\n\nThere may be some unusual use cases, perhaps if treating a SMILES like a word and tokenizing it, where a duplication strategy might be helpful.",
    "2813906": "I would really appreciate something that could do reversals and other transformations automatically. Not as interested in random variations, though I may still try it."
  },
  "source": "meta"
}