{
  "id": 440177,
  "title": "SMILES notation: Rules. Aromaticity and Kekulization. PySmiles, Chem and RDKit.",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/440177",
  "author_name": "Marília Prata",
  "post_date": "2023-09-14T03:52:21.438000",
  "votes": 26,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>The SMILES chemical notation language</h1>\n<p>Citation: SMILES. 2. Algorithm for generation of unique SMILES notation<br>\nDavid Weininger, Arthur Weininger, and Joseph L. Weininger<br>\nJournal of Chemical Information and Computer Sciences 1989 29 (2), 97-101 - DOI: 10.1021/ci00062a008</p>\n<p>\"The SMILES chemical notation language was introduced in the first paper of this series.’ Processing chemical information with greater efficiency than conventional methods, it represents a new approach to computerized chemical nomenclature. SMILES is simple to write because rules and hierarchical procedures, which are inherently difficult for the chemist, are relegated to computer algorithms. For a given chemical structure, arbitrary SMILES notation can take many equally valid forms. One must emerge as “unique” to serve as the identifier of the structure for database and other computer applications.\"</p>\n<p><a href=\"http://organica1.org/seminario/smile_2_1988.pdf\" target=\"_blank\">http://organica1.org/seminario/smile_2_1988.pdf</a></p>\n<h1>SMILES Rules</h1>\n<p>RULE ONE: Atoms and Bonds</p>\n<p>\"SMILES supports all elements in the periodic table. An atom is represented using its respective atomic symbol. Upper case letters refer to non-aromatic atoms; lower case letters refer to aromatic atoms. If the atomic symbol has more than one letter the second letter must be lower case.\"</p>\n<p>\"Single bonds are the default and therefore need not be entered. For example, 'CC' would mean that there is a non-aromatic carbon attached to another non-aromatic carbon by a single bond, and the computer would identify the structure as the chemical ethane. It is also assumed that the bond between two lower case atom symbols is aromatic. A blank terminates the SMILES string.\"</p>\n<p>RULE TWO: Simple Chains</p>\n<p>\"By combining atomic symbols and bond symbols simple chain structures can be represented. The structures that are entered using SMILES are hydrogen-suppressed, that is to say that the molecules are represented without hydrogens. The SMILES software understands the number of possible connections that an atom can have. If enough bonds are not identified by the user through SMILES notation, the system will automatically assume that the other connections are satisfied by hydrogen bonds.\"</p>\n<p>RULE THREE: Branches</p>\n<p>\"A branch from a chain is specified by placing the SMILES symbol(s) for the branch between parenthesis. The string in parentheses is placed directly after the symbol for the atom to which it is connected. If it is connected by a double or triple bond, the bond symbol immediately follows the left parenthesis. Some examples:\"</p>\n<p>RULE FOUR: Rings</p>\n<p>\"SMILES allows a user to identify ring structures by using numbers to identify the opening and closing ring atom. For example, in C1CCCCC1, the first carbon has a number '1' which connects by a single bond with the last carbon which also has a number '1'. The resulting structure is cyclohexane. Chemicals that have multiple rings may be identified by using different numbers for each ring. If a double, single, or aromatic bond is used for the ring closure, the bond symbol is placed before the ring closure number.\"</p>\n<p>RULE FIVE: Charged Atoms</p>\n<p>\"Charges on an atom can be used to override the knowledge regarding valence that is built into SMILES software. The format for identifying a charged atom consists of the atom followed by brackets which enclose the charge on the atom. The number of charges may be explicitly stated ({-1}) or not ({-}). \" </p>\n<p><a href=\"https://archive.epa.gov/med/med_archive_03/web/html/smiles.html\" target=\"_blank\">https://archive.epa.gov/med/med_archive_03/web/html/smiles.html</a></p>\n<h1>Aromaticity and Kekulization</h1>\n<p>\"Aromaticity must be detected in a system that generates an unambiguous chemical nomenclature.  There can be no definition of \"aromaticity\" that is both rigorous and all encompassing; the word implies something about \"reactivity\" to a synthetic chemist, \"ring current\" to a NMR spectroscopist, \"symmetry\" to a crystallographer, and presumably \"odor\" to the original user of the word.  Although the SMILES algorithm produces results that most chemists find natural, nothing is implied by this definition about physical properties.\"</p>\n<p>\"The most important application of the DS for SMILES readers is kekulization. Kekulization is the process of assigning double bonds to a molecular graph using the DS as a guide. Kekulization occurs before the assignment of virtual hydrogens.\"</p>\n<p>\"The problem of assigning alternating double bonds is just one instance of a broader problem in graph theory known as matching. A matching is a subgraph in which each node has degree one. Three kinds of matching are of interest: maximal (no additional edges can be added); maximum (all possible edges have been added); and perfect (all nodes have been added).\"</p>\n<p><a href=\"https://depth-first.com/articles/2020/02/10/a-comprehensive-treatment-of-aromaticity-in-the-smiles-language/\" target=\"_blank\">https://depth-first.com/articles/2020/02/10/a-comprehensive-treatment-of-aromaticity-in-the-smiles-language/</a></p>\n<h1>Kekulé (a.k.a. Lewis Structures)</h1>\n<p>Kekulé structures are similar to Lewis Structures, but instead of covalent bonds being represented by electron dots, the two shared electrons are shown by a line.</p>\n<p><a href=\"https://chem.libretexts.org/Courses/Sacramento_City_College/SCC%3A_Chem_420_-_Organic_Chemistry_I/Text/01%3A_Introduction_and_Review/1.08%3A_Structural_Formulas_-_Lewis%2C_Kekule%2C_Bond-line%2C_Condensed%2C\" target=\"_blank\">https://chem.libretexts.org/Courses/Sacramento_City_College/SCC%3A_Chem_420_-_Organic_Chemistry_I/Text/01%3A_Introduction_and_Review/1.08%3A_Structural_Formulas_-_Lewis%2C_Kekule%2C_Bond-line%2C_Condensed%2C</a></p>\n<h1>Drawing the Structure of Organic Molecules</h1>\n<p>Pysmiles: The lightweight and pure-python SMILES reader and writer<br>\n<a href=\"https://pypi.org/project/pysmiles/\" target=\"_blank\">https://pypi.org/project/pysmiles/</a></p>\n<p>PySMILE: A similar named project, capable of encoding/decoding SMILE format objects. Doesn't deal with SMILES.</p>\n<p>RDKit: Open-Source Cheminformatics Software<br>\nA collection of cheminformatics and machine-learning software, capable of reading and writing SMILES, InChi, and others.<br>\n<a href=\"https://www.rdkit.org/\" target=\"_blank\">https://www.rdkit.org/</a></p>\n<p>OpenEye Chem toolkit: The OpenEye chemistry toolkit is a programming library for chemistry and cheminformatics. It is capable of dealing with (canonical) SMILES and InChi.<br>\n<a href=\"https://www.eyesopen.com/modeling-development-platform\" target=\"_blank\">https://www.eyesopen.com/modeling-development-platform</a></p>",
  "messages": [
    {
      "id": 2438096,
      "postDate": "2023-09-14T03:52:21.437Z",
      "content": "<h1>The SMILES chemical notation language</h1>\n<p>Citation: SMILES. 2. Algorithm for generation of unique SMILES notation<br>\nDavid Weininger, Arthur Weininger, and Joseph L. Weininger<br>\nJournal of Chemical Information and Computer Sciences 1989 29 (2), 97-101 - DOI: 10.1021/ci00062a008</p>\n<p>\"The SMILES chemical notation language was introduced in the first paper of this series.’ Processing chemical information with greater efficiency than conventional methods, it represents a new approach to computerized chemical nomenclature. SMILES is simple to write because rules and hierarchical procedures, which are inherently difficult for the chemist, are relegated to computer algorithms. For a given chemical structure, arbitrary SMILES notation can take many equally valid forms. One must emerge as “unique” to serve as the identifier of the structure for database and other computer applications.\"</p>\n<p><a href=\"http://organica1.org/seminario/smile_2_1988.pdf\" target=\"_blank\">http://organica1.org/seminario/smile_2_1988.pdf</a></p>\n<h1>SMILES Rules</h1>\n<p>RULE ONE: Atoms and Bonds</p>\n<p>\"SMILES supports all elements in the periodic table. An atom is represented using its respective atomic symbol. Upper case letters refer to non-aromatic atoms; lower case letters refer to aromatic atoms. If the atomic symbol has more than one letter the second letter must be lower case.\"</p>\n<p>\"Single bonds are the default and therefore need not be entered. For example, 'CC' would mean that there is a non-aromatic carbon attached to another non-aromatic carbon by a single bond, and the computer would identify the structure as the chemical ethane. It is also assumed that the bond between two lower case atom symbols is aromatic. A blank terminates the SMILES string.\"</p>\n<p>RULE TWO: Simple Chains</p>\n<p>\"By combining atomic symbols and bond symbols simple chain structures can be represented. The structures that are entered using SMILES are hydrogen-suppressed, that is to say that the molecules are represented without hydrogens. The SMILES software understands the number of possible connections that an atom can have. If enough bonds are not identified by the user through SMILES notation, the system will automatically assume that the other connections are satisfied by hydrogen bonds.\"</p>\n<p>RULE THREE: Branches</p>\n<p>\"A branch from a chain is specified by placing the SMILES symbol(s) for the branch between parenthesis. The string in parentheses is placed directly after the symbol for the atom to which it is connected. If it is connected by a double or triple bond, the bond symbol immediately follows the left parenthesis. Some examples:\"</p>\n<p>RULE FOUR: Rings</p>\n<p>\"SMILES allows a user to identify ring structures by using numbers to identify the opening and closing ring atom. For example, in C1CCCCC1, the first carbon has a number '1' which connects by a single bond with the last carbon which also has a number '1'. The resulting structure is cyclohexane. Chemicals that have multiple rings may be identified by using different numbers for each ring. If a double, single, or aromatic bond is used for the ring closure, the bond symbol is placed before the ring closure number.\"</p>\n<p>RULE FIVE: Charged Atoms</p>\n<p>\"Charges on an atom can be used to override the knowledge regarding valence that is built into SMILES software. The format for identifying a charged atom consists of the atom followed by brackets which enclose the charge on the atom. The number of charges may be explicitly stated ({-1}) or not ({-}). \" </p>\n<p><a href=\"https://archive.epa.gov/med/med_archive_03/web/html/smiles.html\" target=\"_blank\">https://archive.epa.gov/med/med_archive_03/web/html/smiles.html</a></p>\n<h1>Aromaticity and Kekulization</h1>\n<p>\"Aromaticity must be detected in a system that generates an unambiguous chemical nomenclature.  There can be no definition of \"aromaticity\" that is both rigorous and all encompassing; the word implies something about \"reactivity\" to a synthetic chemist, \"ring current\" to a NMR spectroscopist, \"symmetry\" to a crystallographer, and presumably \"odor\" to the original user of the word.  Although the SMILES algorithm produces results that most chemists find natural, nothing is implied by this definition about physical properties.\"</p>\n<p>\"The most important application of the DS for SMILES readers is kekulization. Kekulization is the process of assigning double bonds to a molecular graph using the DS as a guide. Kekulization occurs before the assignment of virtual hydrogens.\"</p>\n<p>\"The problem of assigning alternating double bonds is just one instance of a broader problem in graph theory known as matching. A matching is a subgraph in which each node has degree one. Three kinds of matching are of interest: maximal (no additional edges can be added); maximum (all possible edges have been added); and perfect (all nodes have been added).\"</p>\n<p><a href=\"https://depth-first.com/articles/2020/02/10/a-comprehensive-treatment-of-aromaticity-in-the-smiles-language/\" target=\"_blank\">https://depth-first.com/articles/2020/02/10/a-comprehensive-treatment-of-aromaticity-in-the-smiles-language/</a></p>\n<h1>Kekulé (a.k.a. Lewis Structures)</h1>\n<p>Kekulé structures are similar to Lewis Structures, but instead of covalent bonds being represented by electron dots, the two shared electrons are shown by a line.</p>\n<p><a href=\"https://chem.libretexts.org/Courses/Sacramento_City_College/SCC%3A_Chem_420_-_Organic_Chemistry_I/Text/01%3A_Introduction_and_Review/1.08%3A_Structural_Formulas_-_Lewis%2C_Kekule%2C_Bond-line%2C_Condensed%2C\" target=\"_blank\">https://chem.libretexts.org/Courses/Sacramento_City_College/SCC%3A_Chem_420_-_Organic_Chemistry_I/Text/01%3A_Introduction_and_Review/1.08%3A_Structural_Formulas_-_Lewis%2C_Kekule%2C_Bond-line%2C_Condensed%2C</a></p>\n<h1>Drawing the Structure of Organic Molecules</h1>\n<p>Pysmiles: The lightweight and pure-python SMILES reader and writer<br>\n<a href=\"https://pypi.org/project/pysmiles/\" target=\"_blank\">https://pypi.org/project/pysmiles/</a></p>\n<p>PySMILE: A similar named project, capable of encoding/decoding SMILE format objects. Doesn't deal with SMILES.</p>\n<p>RDKit: Open-Source Cheminformatics Software<br>\nA collection of cheminformatics and machine-learning software, capable of reading and writing SMILES, InChi, and others.<br>\n<a href=\"https://www.rdkit.org/\" target=\"_blank\">https://www.rdkit.org/</a></p>\n<p>OpenEye Chem toolkit: The OpenEye chemistry toolkit is a programming library for chemistry and cheminformatics. It is capable of dealing with (canonical) SMILES and InChi.<br>\n<a href=\"https://www.eyesopen.com/modeling-development-platform\" target=\"_blank\">https://www.eyesopen.com/modeling-development-platform</a></p>",
      "rawMarkdown": "#The SMILES chemical notation language\n\nCitation: SMILES. 2. Algorithm for generation of unique SMILES notation\nDavid Weininger, Arthur Weininger, and Joseph L. Weininger\nJournal of Chemical Information and Computer Sciences 1989 29 (2), 97-101 - DOI: 10.1021/ci00062a008\n\n\"The SMILES chemical notation language was introduced in the first paper of this series.’ Processing chemical information with greater efficiency than conventional methods, it represents a new approach to computerized chemical nomenclature. SMILES is simple to write because rules and hierarchical procedures, which are inherently difficult for the chemist, are relegated to computer algorithms. For a given chemical structure, arbitrary SMILES notation can take many equally valid forms. One must emerge as “unique” to serve as the identifier of the structure for database and other computer applications.\"\n\nhttp://organica1.org/seminario/smile_2_1988.pdf\n\n#SMILES Rules\n\nRULE ONE: Atoms and Bonds\n\n\"SMILES supports all elements in the periodic table. An atom is represented using its respective atomic symbol. Upper case letters refer to non-aromatic atoms; lower case letters refer to aromatic atoms. If the atomic symbol has more than one letter the second letter must be lower case.\"\n\n\"Single bonds are the default and therefore need not be entered. For example, 'CC' would mean that there is a non-aromatic carbon attached to another non-aromatic carbon by a single bond, and the computer would identify the structure as the chemical ethane. It is also assumed that the bond between two lower case atom symbols is aromatic. A blank terminates the SMILES string.\"\n\nRULE TWO: Simple Chains\n\n\"By combining atomic symbols and bond symbols simple chain structures can be represented. The structures that are entered using SMILES are hydrogen-suppressed, that is to say that the molecules are represented without hydrogens. The SMILES software understands the number of possible connections that an atom can have. If enough bonds are not identified by the user through SMILES notation, the system will automatically assume that the other connections are satisfied by hydrogen bonds.\"\n\nRULE THREE: Branches\n\n\"A branch from a chain is specified by placing the SMILES symbol(s) for the branch between parenthesis. The string in parentheses is placed directly after the symbol for the atom to which it is connected. If it is connected by a double or triple bond, the bond symbol immediately follows the left parenthesis. Some examples:\"\n\nRULE FOUR: Rings\n\n\"SMILES allows a user to identify ring structures by using numbers to identify the opening and closing ring atom. For example, in C1CCCCC1, the first carbon has a number '1' which connects by a single bond with the last carbon which also has a number '1'. The resulting structure is cyclohexane. Chemicals that have multiple rings may be identified by using different numbers for each ring. If a double, single, or aromatic bond is used for the ring closure, the bond symbol is placed before the ring closure number.\"\n\nRULE FIVE: Charged Atoms\n\n\"Charges on an atom can be used to override the knowledge regarding valence that is built into SMILES software. The format for identifying a charged atom consists of the atom followed by brackets which enclose the charge on the atom. The number of charges may be explicitly stated ({-1}) or not ({-}). \" \n\nhttps://archive.epa.gov/med/med_archive_03/web/html/smiles.html\n\n#Aromaticity and Kekulization\n\n\"Aromaticity must be detected in a system that generates an unambiguous chemical nomenclature.  There can be no definition of \"aromaticity\" that is both rigorous and all encompassing; the word implies something about \"reactivity\" to a synthetic chemist, \"ring current\" to a NMR spectroscopist, \"symmetry\" to a crystallographer, and presumably \"odor\" to the original user of the word.  Although the SMILES algorithm produces results that most chemists find natural, nothing is implied by this definition about physical properties.\"\n\n\"The most important application of the DS for SMILES readers is kekulization. Kekulization is the process of assigning double bonds to a molecular graph using the DS as a guide. Kekulization occurs before the assignment of virtual hydrogens.\"\n\n\"The problem of assigning alternating double bonds is just one instance of a broader problem in graph theory known as matching. A matching is a subgraph in which each node has degree one. Three kinds of matching are of interest: maximal (no additional edges can be added); maximum (all possible edges have been added); and perfect (all nodes have been added).\"\n\nhttps://depth-first.com/articles/2020/02/10/a-comprehensive-treatment-of-aromaticity-in-the-smiles-language/\n\n#Kekulé (a.k.a. Lewis Structures)\n\nKekulé structures are similar to Lewis Structures, but instead of covalent bonds being represented by electron dots, the two shared electrons are shown by a line.\n\nhttps://chem.libretexts.org/Courses/Sacramento_City_College/SCC%3A_Chem_420_-_Organic_Chemistry_I/Text/01%3A_Introduction_and_Review/1.08%3A_Structural_Formulas_-_Lewis%2C_Kekule%2C_Bond-line%2C_Condensed%2C\n\n#Drawing the Structure of Organic Molecules\n\nPysmiles: The lightweight and pure-python SMILES reader and writer\nhttps://pypi.org/project/pysmiles/\n\nPySMILE: A similar named project, capable of encoding/decoding SMILE format objects. Doesn't deal with SMILES.\n\nRDKit: Open-Source Cheminformatics Software\nA collection of cheminformatics and machine-learning software, capable of reading and writing SMILES, InChi, and others.\nhttps://www.rdkit.org/\n\nOpenEye Chem toolkit: The OpenEye chemistry toolkit is a programming library for chemistry and cheminformatics. It is capable of dealing with (canonical) SMILES and InChi.\nhttps://www.eyesopen.com/modeling-development-platform",
      "votes": 26
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2438096": "#The SMILES chemical notation language\n\nCitation: SMILES. 2. Algorithm for generation of unique SMILES notation\nDavid Weininger, Arthur Weininger, and Joseph L. Weininger\nJournal of Chemical Information and Computer Sciences 1989 29 (2), 97-101 - DOI: 10.1021/ci00062a008\n\n\"The SMILES chemical notation language was introduced in the first paper of this series.’ Processing chemical information with greater efficiency than conventional methods, it represents a new approach to computerized chemical nomenclature. SMILES is simple to write because rules and hierarchical procedures, which are inherently difficult for the chemist, are relegated to computer algorithms. For a given chemical structure, arbitrary SMILES notation can take many equally valid forms. One must emerge as “unique” to serve as the identifier of the structure for database and other computer applications.\"\n\nhttp://organica1.org/seminario/smile_2_1988.pdf\n\n#SMILES Rules\n\nRULE ONE: Atoms and Bonds\n\n\"SMILES supports all elements in the periodic table. An atom is represented using its respective atomic symbol. Upper case letters refer to non-aromatic atoms; lower case letters refer to aromatic atoms. If the atomic symbol has more than one letter the second letter must be lower case.\"\n\n\"Single bonds are the default and therefore need not be entered. For example, 'CC' would mean that there is a non-aromatic carbon attached to another non-aromatic carbon by a single bond, and the computer would identify the structure as the chemical ethane. It is also assumed that the bond between two lower case atom symbols is aromatic. A blank terminates the SMILES string.\"\n\nRULE TWO: Simple Chains\n\n\"By combining atomic symbols and bond symbols simple chain structures can be represented. The structures that are entered using SMILES are hydrogen-suppressed, that is to say that the molecules are represented without hydrogens. The SMILES software understands the number of possible connections that an atom can have. If enough bonds are not identified by the user through SMILES notation, the system will automatically assume that the other connections are satisfied by hydrogen bonds.\"\n\nRULE THREE: Branches\n\n\"A branch from a chain is specified by placing the SMILES symbol(s) for the branch between parenthesis. The string in parentheses is placed directly after the symbol for the atom to which it is connected. If it is connected by a double or triple bond, the bond symbol immediately follows the left parenthesis. Some examples:\"\n\nRULE FOUR: Rings\n\n\"SMILES allows a user to identify ring structures by using numbers to identify the opening and closing ring atom. For example, in C1CCCCC1, the first carbon has a number '1' which connects by a single bond with the last carbon which also has a number '1'. The resulting structure is cyclohexane. Chemicals that have multiple rings may be identified by using different numbers for each ring. If a double, single, or aromatic bond is used for the ring closure, the bond symbol is placed before the ring closure number.\"\n\nRULE FIVE: Charged Atoms\n\n\"Charges on an atom can be used to override the knowledge regarding valence that is built into SMILES software. The format for identifying a charged atom consists of the atom followed by brackets which enclose the charge on the atom. The number of charges may be explicitly stated ({-1}) or not ({-}). \" \n\nhttps://archive.epa.gov/med/med_archive_03/web/html/smiles.html\n\n#Aromaticity and Kekulization\n\n\"Aromaticity must be detected in a system that generates an unambiguous chemical nomenclature.  There can be no definition of \"aromaticity\" that is both rigorous and all encompassing; the word implies something about \"reactivity\" to a synthetic chemist, \"ring current\" to a NMR spectroscopist, \"symmetry\" to a crystallographer, and presumably \"odor\" to the original user of the word.  Although the SMILES algorithm produces results that most chemists find natural, nothing is implied by this definition about physical properties.\"\n\n\"The most important application of the DS for SMILES readers is kekulization. Kekulization is the process of assigning double bonds to a molecular graph using the DS as a guide. Kekulization occurs before the assignment of virtual hydrogens.\"\n\n\"The problem of assigning alternating double bonds is just one instance of a broader problem in graph theory known as matching. A matching is a subgraph in which each node has degree one. Three kinds of matching are of interest: maximal (no additional edges can be added); maximum (all possible edges have been added); and perfect (all nodes have been added).\"\n\nhttps://depth-first.com/articles/2020/02/10/a-comprehensive-treatment-of-aromaticity-in-the-smiles-language/\n\n#Kekulé (a.k.a. Lewis Structures)\n\nKekulé structures are similar to Lewis Structures, but instead of covalent bonds being represented by electron dots, the two shared electrons are shown by a line.\n\nhttps://chem.libretexts.org/Courses/Sacramento_City_College/SCC%3A_Chem_420_-_Organic_Chemistry_I/Text/01%3A_Introduction_and_Review/1.08%3A_Structural_Formulas_-_Lewis%2C_Kekule%2C_Bond-line%2C_Condensed%2C\n\n#Drawing the Structure of Organic Molecules\n\nPysmiles: The lightweight and pure-python SMILES reader and writer\nhttps://pypi.org/project/pysmiles/\n\nPySMILE: A similar named project, capable of encoding/decoding SMILE format objects. Doesn't deal with SMILES.\n\nRDKit: Open-Source Cheminformatics Software\nA collection of cheminformatics and machine-learning software, capable of reading and writing SMILES, InChi, and others.\nhttps://www.rdkit.org/\n\nOpenEye Chem toolkit: The OpenEye chemistry toolkit is a programming library for chemistry and cheminformatics. It is capable of dealing with (canonical) SMILES and InChi.\nhttps://www.eyesopen.com/modeling-development-platform"
  }
}