{
  "id": 242592,
  "title": "Some insights and observations",
  "url": "/competitions/bms-molecular-translation/discussion/242592",
  "author_name": "",
  "post_date": "2021-05-29T19:18:33.206106600Z",
  "votes": 19,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am very late to this competition but thought I should at least try to understand the problem and submit one submission. Here are some of the insights and observations I got from reading different forum discussions and searching on my own. </p>\n<h2>Keywords</h2>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Skeletal_formula\" target=\"_blank\">Skeletal formula</a>: a 2D representation of an organic molecule. </li>\n</ul>\n<p>Some interesting things to know: carbon and hydrogen atoms aren't represented but are implicit. Other than hydrogen and carbon, other additional atoms are drawn. Made by <a href=\"https://en.wikipedia.org/wiki/August_Kekul%C3%A9\" target=\"_blank\">August Kekulé</a>. As a fun fact: he discovered the <a href=\"https://en.wikipedia.org/wiki/Benzene\" target=\"_blank\">benzene</a> structure and apparently had a dream about it inspired by the <a href=\"https://en.wikipedia.org/wiki/Ouroboros\" target=\"_blank\">Ouroboros</a>.</p>\n<p>Here is an image of how bonds are represented: </p>\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/4/48/Skeletal_formula_samples_stereochemistry.svg/300px-Skeletal_formula_samples_stereochemistry.svg.png\" alt=\"chemical bonds\"></p>\n<ul>\n<li>InChI: short for <a href=\"https://en.wikipedia.org/wiki/International_Chemical_Identifier\" target=\"_blank\">International Chemical Identifier</a> is the target to predict. It is a machine-readable chemical description of a given molecule. Here is an example for Morphine (from Wikipedia):</li>\n</ul>\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/3/33/Morphin_-_Morphine.svg/440px-Morphin_-_Morphine.svg.png\" alt=\"morphine\"></p>\n<pre><code>InChI=1S/C17H19NO3/c1-18-7-6-17-10-3-5-13(20)16(17)21-15-12(19)4-2-9(14(15)17)8-11(10)18/h2-5,10-11,13,16,19-20H,6-8H2,1H3/t10-,11+,13-,16-,17-/m0/s1\n</code></pre>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Simplified_Molecular_Input_Line_Entry_Specification\" target=\"_blank\">SMILES</a>: short for Simplified Molecular Input Line Entry Specification, it is a <a href=\"https://en.wikipedia.org/wiki/Line_notation\" target=\"_blank\">line notation</a> for a molecular formula using ASCII characters. Continuing with the morphine example, here is its SMILES is:</li>\n</ul>\n<pre><code>CN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n</code></pre>\n<ul>\n<li>How to go from InChIkey to SMILES?</li>\n</ul>\n<p>Using the InChIkey, we need a lookup table to find the corresponding SMILES representation since all the information isn't contained in the key.</p>\n<p>Here is a code snippet from a SO <a href=\"https://bioinformatics.stackexchange.com/questions/10755/is-there-a-python-package-to-convert-inchi-to-molecular-structures\" target=\"_blank\">thread</a> that uses the PubChem API to do the lookup: </p>\n<pre><code>import requests\n\n\nmorphine_inchikey = \"BQJCRHHNABKAKU-KBQPJGBKSA-N\"\n\n\nr = requests.get(f'https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/inchikey/{morphine_inchikey}/property/CanonicalSMILES/JSON').json()\n\nmorphine_smiles = r['PropertyTable']['Properties'][0]['CanonicalSMILES']\n\nprint(morphine_smiles)\n\n# You should get: CN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n</code></pre>\n<p>To check that the computation is correct, check the following pubchem <a href=\"https://pubchem.ncbi.nlm.nih.gov/compound/Morphine\" target=\"_blank\">page</a></p>\n<p><img src=\"https://drive.google.com/uc?id=19H9rFf3Z_GeUQJjSpTsSCvW1Keec-89m\" alt=\"pubchem morphine\"></p>\n<ul>\n<li>Molecular formula: a concatenation of the atoms. Continuing with the morphine example, here is its molecular morphine: <code>C17H19NO3</code></li>\n</ul>\n<h2>Metric</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Levenshtein_distance\" target=\"_blank\">Levenshtein distance</a>: this metric is often used in NLP tasks. </p>\n<p>Indeed, it compares two strings (it is an <a href=\"https://en.wikipedia.org/wiki/Edit_distance\" target=\"_blank\">edit distance</a>) <code>a</code> and <code>b</code> using the following formula (from the Wikipedia page): </p>\n<p><img src=\"https://latex.codecogs.com/gif.latex?\\qquad\\operatorname{lev}(a,b)&amp;space;=&amp;space;\\begin{cases}&amp;space;|a|&amp;space;&amp;&amp;space;\\text{&amp;space;if&amp;space;}&amp;space;|b|&amp;space;=&amp;space;0,&amp;space;\\\\&amp;space;|b|&amp;space;&amp;&amp;space;\\text{&amp;space;if&amp;space;}&amp;space;|a|&amp;space;=&amp;space;0,&amp;space;\\\\&amp;space;\\operatorname{lev}(\\operatorname{tail}(a),\\operatorname{tail}(b))&amp;space;&amp;&amp;space;\\text{&amp;space;if&amp;space;}&amp;space;a[0]&amp;space;=&amp;space;b[0]&amp;space;\\\\&amp;space;1&amp;space;+&amp;space;\\min&amp;space;\\begin{cases}&amp;space;\\operatorname{lev}(\\operatorname{tail}(a),&amp;space;b)&amp;space;\\\\&amp;space;\\operatorname{lev}(a,&amp;space;\\operatorname{tail}(b))&amp;space;\\\\&amp;space;\\operatorname{lev}(\\operatorname{tail}(a),&amp;space;\\operatorname{tail}(b))&amp;space;\\\\&amp;space;\\end{cases}&amp;space;&amp;&amp;space;\\text{&amp;space;otherwise.}&amp;space;\\end{cases}\"></p>\n<p>where <code>tail(x)</code> is the string without the first character and <code>x[i]</code> is the character at position i.</p>\n<p>This is hard to parse so let's check with an example. Let's suppose that we have the following two strings (for example <code>a</code> is the true target and <code>b</code> is the predicted one): </p>\n<ul>\n<li><code>a = \"test\"</code></li>\n<li><code>b = \"tefl\"</code></li>\n</ul>\n<p>For the first two characters, we are in the third part of the definition. Starting from the third character, we need to use the fourth part of the definition since the characters are different</p>\n<p>Here is another example:</p>\n<ul>\n<li><code>a = \"test\"</code></li>\n<li><code>b = \"\"</code></li>\n</ul>\n<p>Since the <code>b</code> string is empty, the metric is equal to the length of <code>a</code>, i.e. <strong>4</strong>. </p>\n<p>There are also two bounds that can be obtained: </p>\n<ul>\n<li>lower: <code>abs(len(a) - len(b))</code></li>\n<li>upper: <code>max(len(a), len(b))</code></li>\n</ul>\n<p>Thus, if <code>len(a) == len(b)</code>, the lower bound is 0 and is attained when <code>a == b</code>.</p>\n<p>Now that the concept is better understood, let's apply this to a target (truncated to the first few characters): </p>\n<ul>\n<li><code>a = \"InChI=1S/C15H24O2\"</code></li>\n<li><code>b = \"InChI=1F/C14H23O1\"</code></li>\n</ul>\n<p>We can use a Python library to compute the distance: <code>pip install python-Levenshtein</code>.</p>\n<pre><code>import Levenshtein\n\na = \"InChI=1S/C15H24O2\"\nb = \"InChI=1F/C14H23O1\"\n\ndistance = Levenshtein.distance(a, b)\n\nprint(distance)\n\n# You should get: 4\n</code></pre>\n<h2>Some model ideas</h2>\n<ul>\n<li>extract the molecule's graph from the image then use an appropriate library to find the SMILES/InChI representation.</li>\n<li>train a vision-nlp mixed model that takes the images as an input then uses some recursive network (RNN, LSTM, transformer etc) to match the textual representation.</li>\n</ul>\n<h2>Useful packages</h2>\n<ul>\n<li><a href=\"https://www.rdkit.org/docs/index.html\" target=\"_blank\">RDKit</a>: everything related to chemistry and ML</li>\n<li><a href=\"https://pypi.org/project/pysmiles/\" target=\"_blank\">pysmiles</a>: read and write SMILES.</li>\n<li><a href=\"https://pytorch.org/\" target=\"_blank\">Pytorch</a>: as usual of course. ;)</li>\n<li><a href=\"https://github.com/huggingface/transformers\" target=\"_blank\">Transformers</a>: might be useful.</li>\n</ul>\n<p>and the usual pydata stack.</p>\n<h2>Similar competitions</h2>\n<ul>\n<li>The only one I have done so far: <a href=\"https://www.kaggle.com/c/champs-scalar-coupling\" target=\"_blank\">CHAMPS</a></li>\n<li>The closest to the current competition, from the DACON <a href=\"https://dacon.io/competitions/official/235640/overview/description\" target=\"_blank\">platform</a></li>\n</ul>",
  "messages": [
    {
      "id": "1327963",
      "postDate": "05/29/2021 19:18:33",
      "content": "<p>I am very late to this competition but thought I should at least try to understand the problem and submit one submission. Here are some of the insights and observations I got from reading different forum discussions and searching on my own. </p>\n<h2>Keywords</h2>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Skeletal_formula\" target=\"_blank\">Skeletal formula</a>: a 2D representation of an organic molecule. </li>\n</ul>\n<p>Some interesting things to know: carbon and hydrogen atoms aren't represented but are implicit. Other than hydrogen and carbon, other additional atoms are drawn. Made by <a href=\"https://en.wikipedia.org/wiki/August_Kekul%C3%A9\" target=\"_blank\">August Kekulé</a>. As a fun fact: he discovered the <a href=\"https://en.wikipedia.org/wiki/Benzene\" target=\"_blank\">benzene</a> structure and apparently had a dream about it inspired by the <a href=\"https://en.wikipedia.org/wiki/Ouroboros\" target=\"_blank\">Ouroboros</a>.</p>\n<p>Here is an image of how bonds are represented: </p>\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/4/48/Skeletal_formula_samples_stereochemistry.svg/300px-Skeletal_formula_samples_stereochemistry.svg.png\" alt=\"chemical bonds\"></p>\n<ul>\n<li>InChI: short for <a href=\"https://en.wikipedia.org/wiki/International_Chemical_Identifier\" target=\"_blank\">International Chemical Identifier</a> is the target to predict. It is a machine-readable chemical description of a given molecule. Here is an example for Morphine (from Wikipedia):</li>\n</ul>\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/3/33/Morphin_-_Morphine.svg/440px-Morphin_-_Morphine.svg.png\" alt=\"morphine\"></p>\n<pre><code>InChI=1S/C17H19NO3/c1-18-7-6-17-10-3-5-13(20)16(17)21-15-12(19)4-2-9(14(15)17)8-11(10)18/h2-5,10-11,13,16,19-20H,6-8H2,1H3/t10-,11+,13-,16-,17-/m0/s1\n</code></pre>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Simplified_Molecular_Input_Line_Entry_Specification\" target=\"_blank\">SMILES</a>: short for Simplified Molecular Input Line Entry Specification, it is a <a href=\"https://en.wikipedia.org/wiki/Line_notation\" target=\"_blank\">line notation</a> for a molecular formula using ASCII characters. Continuing with the morphine example, here is its SMILES is:</li>\n</ul>\n<pre><code>CN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n</code></pre>\n<ul>\n<li>How to go from InChIkey to SMILES?</li>\n</ul>\n<p>Using the InChIkey, we need a lookup table to find the corresponding SMILES representation since all the information isn't contained in the key.</p>\n<p>Here is a code snippet from a SO <a href=\"https://bioinformatics.stackexchange.com/questions/10755/is-there-a-python-package-to-convert-inchi-to-molecular-structures\" target=\"_blank\">thread</a> that uses the PubChem API to do the lookup: </p>\n<pre><code>import requests\n\n\nmorphine_inchikey = \"BQJCRHHNABKAKU-KBQPJGBKSA-N\"\n\n\nr = requests.get(f'https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/inchikey/{morphine_inchikey}/property/CanonicalSMILES/JSON').json()\n\nmorphine_smiles = r['PropertyTable']['Properties'][0]['CanonicalSMILES']\n\nprint(morphine_smiles)\n\n# You should get: CN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n</code></pre>\n<p>To check that the computation is correct, check the following pubchem <a href=\"https://pubchem.ncbi.nlm.nih.gov/compound/Morphine\" target=\"_blank\">page</a></p>\n<p><img src=\"https://drive.google.com/uc?id=19H9rFf3Z_GeUQJjSpTsSCvW1Keec-89m\" alt=\"pubchem morphine\"></p>\n<ul>\n<li>Molecular formula: a concatenation of the atoms. Continuing with the morphine example, here is its molecular morphine: <code>C17H19NO3</code></li>\n</ul>\n<h2>Metric</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Levenshtein_distance\" target=\"_blank\">Levenshtein distance</a>: this metric is often used in NLP tasks. </p>\n<p>Indeed, it compares two strings (it is an <a href=\"https://en.wikipedia.org/wiki/Edit_distance\" target=\"_blank\">edit distance</a>) <code>a</code> and <code>b</code> using the following formula (from the Wikipedia page): </p>\n<p><img src=\"https://latex.codecogs.com/gif.latex?\\qquad\\operatorname{lev}(a,b)&amp;space;=&amp;space;\\begin{cases}&amp;space;|a|&amp;space;&amp;&amp;space;\\text{&amp;space;if&amp;space;}&amp;space;|b|&amp;space;=&amp;space;0,&amp;space;\\\\&amp;space;|b|&amp;space;&amp;&amp;space;\\text{&amp;space;if&amp;space;}&amp;space;|a|&amp;space;=&amp;space;0,&amp;space;\\\\&amp;space;\\operatorname{lev}(\\operatorname{tail}(a),\\operatorname{tail}(b))&amp;space;&amp;&amp;space;\\text{&amp;space;if&amp;space;}&amp;space;a[0]&amp;space;=&amp;space;b[0]&amp;space;\\\\&amp;space;1&amp;space;+&amp;space;\\min&amp;space;\\begin{cases}&amp;space;\\operatorname{lev}(\\operatorname{tail}(a),&amp;space;b)&amp;space;\\\\&amp;space;\\operatorname{lev}(a,&amp;space;\\operatorname{tail}(b))&amp;space;\\\\&amp;space;\\operatorname{lev}(\\operatorname{tail}(a),&amp;space;\\operatorname{tail}(b))&amp;space;\\\\&amp;space;\\end{cases}&amp;space;&amp;&amp;space;\\text{&amp;space;otherwise.}&amp;space;\\end{cases}\"></p>\n<p>where <code>tail(x)</code> is the string without the first character and <code>x[i]</code> is the character at position i.</p>\n<p>This is hard to parse so let's check with an example. Let's suppose that we have the following two strings (for example <code>a</code> is the true target and <code>b</code> is the predicted one): </p>\n<ul>\n<li><code>a = \"test\"</code></li>\n<li><code>b = \"tefl\"</code></li>\n</ul>\n<p>For the first two characters, we are in the third part of the definition. Starting from the third character, we need to use the fourth part of the definition since the characters are different</p>\n<p>Here is another example:</p>\n<ul>\n<li><code>a = \"test\"</code></li>\n<li><code>b = \"\"</code></li>\n</ul>\n<p>Since the <code>b</code> string is empty, the metric is equal to the length of <code>a</code>, i.e. <strong>4</strong>. </p>\n<p>There are also two bounds that can be obtained: </p>\n<ul>\n<li>lower: <code>abs(len(a) - len(b))</code></li>\n<li>upper: <code>max(len(a), len(b))</code></li>\n</ul>\n<p>Thus, if <code>len(a) == len(b)</code>, the lower bound is 0 and is attained when <code>a == b</code>.</p>\n<p>Now that the concept is better understood, let's apply this to a target (truncated to the first few characters): </p>\n<ul>\n<li><code>a = \"InChI=1S/C15H24O2\"</code></li>\n<li><code>b = \"InChI=1F/C14H23O1\"</code></li>\n</ul>\n<p>We can use a Python library to compute the distance: <code>pip install python-Levenshtein</code>.</p>\n<pre><code>import Levenshtein\n\na = \"InChI=1S/C15H24O2\"\nb = \"InChI=1F/C14H23O1\"\n\ndistance = Levenshtein.distance(a, b)\n\nprint(distance)\n\n# You should get: 4\n</code></pre>\n<h2>Some model ideas</h2>\n<ul>\n<li>extract the molecule's graph from the image then use an appropriate library to find the SMILES/InChI representation.</li>\n<li>train a vision-nlp mixed model that takes the images as an input then uses some recursive network (RNN, LSTM, transformer etc) to match the textual representation.</li>\n</ul>\n<h2>Useful packages</h2>\n<ul>\n<li><a href=\"https://www.rdkit.org/docs/index.html\" target=\"_blank\">RDKit</a>: everything related to chemistry and ML</li>\n<li><a href=\"https://pypi.org/project/pysmiles/\" target=\"_blank\">pysmiles</a>: read and write SMILES.</li>\n<li><a href=\"https://pytorch.org/\" target=\"_blank\">Pytorch</a>: as usual of course. ;)</li>\n<li><a href=\"https://github.com/huggingface/transformers\" target=\"_blank\">Transformers</a>: might be useful.</li>\n</ul>\n<p>and the usual pydata stack.</p>\n<h2>Similar competitions</h2>\n<ul>\n<li>The only one I have done so far: <a href=\"https://www.kaggle.com/c/champs-scalar-coupling\" target=\"_blank\">CHAMPS</a></li>\n<li>The closest to the current competition, from the DACON <a href=\"https://dacon.io/competitions/official/235640/overview/description\" target=\"_blank\">platform</a></li>\n</ul>",
      "rawMarkdown": "I am very late to this competition but thought I should at least try to understand the problem and submit one submission. Here are some of the insights and observations I got from reading different forum discussions and searching on my own. \n\n\n## Keywords\n\n\n\n- [Skeletal formula](https://en.wikipedia.org/wiki/Skeletal_formula): a 2D representation of an organic molecule. \n\nSome interesting things to know: carbon and hydrogen atoms aren't represented but are implicit. Other than hydrogen and carbon, other additional atoms are drawn. Made by [August Kekulé](https://en.wikipedia.org/wiki/August_Kekul%C3%A9). As a fun fact: he discovered the [benzene](https://en.wikipedia.org/wiki/Benzene) structure and apparently had a dream about it inspired by the [Ouroboros](https://en.wikipedia.org/wiki/Ouroboros).\n\nHere is an image of how bonds are represented: \n\n\n![chemical bonds](https://upload.wikimedia.org/wikipedia/commons/thumb/4/48/Skeletal_formula_samples_stereochemistry.svg/300px-Skeletal_formula_samples_stereochemistry.svg.png)\n\n- InChI: short for [International Chemical Identifier](https://en.wikipedia.org/wiki/International_Chemical_Identifier) is the target to predict. It is a machine-readable chemical description of a given molecule. Here is an example for Morphine (from Wikipedia):\n\n\n![morphine](https://upload.wikimedia.org/wikipedia/commons/thumb/3/33/Morphin_-_Morphine.svg/440px-Morphin_-_Morphine.svg.png)\n\n```\nInChI=1S/C17H19NO3/c1-18-7-6-17-10-3-5-13(20)16(17)21-15-12(19)4-2-9(14(15)17)8-11(10)18/h2-5,10-11,13,16,19-20H,6-8H2,1H3/t10-,11+,13-,16-,17-/m0/s1\n```\n\n\n\n\n- [SMILES](https://en.wikipedia.org/wiki/Simplified_Molecular_Input_Line_Entry_Specification): short for Simplified Molecular Input Line Entry Specification, it is a [line notation](https://en.wikipedia.org/wiki/Line_notation) for a molecular formula using ASCII characters. Continuing with the morphine example, here is its SMILES is:\n\n```\nCN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n```\n\n\n\n- How to go from InChIkey to SMILES?\n\nUsing the InChIkey, we need a lookup table to find the corresponding SMILES representation since all the information isn't contained in the key.\n\nHere is a code snippet from a SO [thread](https://bioinformatics.stackexchange.com/questions/10755/is-there-a-python-package-to-convert-inchi-to-molecular-structures) that uses the PubChem API to do the lookup: \n\n```python\n\nimport requests\n\n\nmorphine_inchikey = \"BQJCRHHNABKAKU-KBQPJGBKSA-N\"\n\n\nr = requests.get(f'https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/inchikey/{morphine_inchikey}/property/CanonicalSMILES/JSON').json()\n\nmorphine_smiles = r['PropertyTable']['Properties'][0]['CanonicalSMILES']\n\nprint(morphine_smiles)\n\n# You should get: CN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n\n```\n\nTo check that the computation is correct, check the following pubchem [page](https://pubchem.ncbi.nlm.nih.gov/compound/Morphine)\n\n\n\n\n![pubchem morphine](https://drive.google.com/uc?id=19H9rFf3Z_GeUQJjSpTsSCvW1Keec-89m)\n\n- Molecular formula: a concatenation of the atoms. Continuing with the morphine example, here is its molecular morphine: `C17H19NO3`\n\n## Metric\n\n[Levenshtein distance](https://en.wikipedia.org/wiki/Levenshtein_distance): this metric is often used in NLP tasks. \n\nIndeed, it compares two strings (it is an [edit distance](https://en.wikipedia.org/wiki/Edit_distance)) `a` and `b` using the following formula (from the Wikipedia page): \n\n\n\n<img src=\"https://latex.codecogs.com/gif.latex?\\qquad\\operatorname{lev}(a,b)&space;=&space;\\begin{cases}&space;|a|&space;&&space;\\text{&space;if&space;}&space;|b|&space;=&space;0,&space;\\\\&space;|b|&space;&&space;\\text{&space;if&space;}&space;|a|&space;=&space;0,&space;\\\\&space;\\operatorname{lev}(\\operatorname{tail}(a),\\operatorname{tail}(b))&space;&&space;\\text{&space;if&space;}&space;a[0]&space;=&space;b[0]&space;\\\\&space;1&space;&plus;&space;\\min&space;\\begin{cases}&space;\\operatorname{lev}(\\operatorname{tail}(a),&space;b)&space;\\\\&space;\\operatorname{lev}(a,&space;\\operatorname{tail}(b))&space;\\\\&space;\\operatorname{lev}(\\operatorname{tail}(a),&space;\\operatorname{tail}(b))&space;\\\\&space;\\end{cases}&space;&&space;\\text{&space;otherwise.}&space;\\end{cases}\" title=\"\\qquad\\operatorname{lev}(a,b) = \\begin{cases} |a| & \\text{ if } |b| = 0, \\\\ |b| & \\text{ if } |a| = 0, \\\\ \\operatorname{lev}(\\operatorname{tail}(a),\\operatorname{tail}(b)) & \\text{ if } a[0] = b[0] \\\\ 1 + \\min \\begin{cases} \\operatorname{lev}(\\operatorname{tail}(a), b) \\\\ \\operatorname{lev}(a, \\operatorname{tail}(b)) \\\\ \\operatorname{lev}(\\operatorname{tail}(a), \\operatorname{tail}(b)) \\\\ \\end{cases} & \\text{ otherwise.} \\end{cases}\" />\n\nwhere `tail(x)` is the string without the first character and `x[i]` is the character at position i.\n\nThis is hard to parse so let's check with an example. Let's suppose that we have the following two strings (for example `a` is the true target and `b` is the predicted one): \n\n- `a = \"test\"`\n- `b = \"tefl\"`\n\nFor the first two characters, we are in the third part of the definition. Starting from the third character, we need to use the fourth part of the definition since the characters are different\n\n\nHere is another example:\n\n- `a = \"test\"`\n- `b = \"\"`\n\nSince the `b` string is empty, the metric is equal to the length of `a`, i.e. **4**. \n\n\nThere are also two bounds that can be obtained: \n\n- lower: `abs(len(a) - len(b))`\n- upper: `max(len(a), len(b))`\n\nThus, if `len(a) == len(b)`, the lower bound is 0 and is attained when `a == b`.\n\nNow that the concept is better understood, let's apply this to a target (truncated to the first few characters): \n\n- `a = \"InChI=1S/C15H24O2\"`\n- `b = \"InChI=1F/C14H23O1\"`\n\nWe can use a Python library to compute the distance: `pip install python-Levenshtein`.\n\n\n``` \n\nimport Levenshtein\n\na = \"InChI=1S/C15H24O2\"\nb = \"InChI=1F/C14H23O1\"\n\ndistance = Levenshtein.distance(a, b)\n\nprint(distance)\n\n# You should get: 4\n```\n\n## Some model ideas\n\n- extract the molecule's graph from the image then use an appropriate library to find the SMILES/InChI representation.\n- train a vision-nlp mixed model that takes the images as an input then uses some recursive network (RNN, LSTM, transformer etc) to match the textual representation.\n\n\n## Useful packages\n\n\n- [RDKit](https://www.rdkit.org/docs/index.html): everything related to chemistry and ML\n- [pysmiles](https://pypi.org/project/pysmiles/): read and write SMILES.\n- [Pytorch](https://pytorch.org/): as usual of course. ;)\n- [Transformers](https://github.com/huggingface/transformers): might be useful.\n\nand the usual pydata stack.\n\n## Similar competitions\n\n- The only one I have done so far: [CHAMPS](https://www.kaggle.com/c/champs-scalar-coupling)\n- The closest to the current competition, from the DACON [platform](https://dacon.io/competitions/official/235640/overview/description)",
      "votes": null
    },
    {
      "id": "1333860",
      "postDate": "06/03/2021 05:23:05",
      "content": "<p>Nice work 👍 keep it up.</p>",
      "rawMarkdown": "Nice work 👍 keep it up.",
      "votes": null
    },
    {
      "id": "1333868",
      "postDate": "06/03/2021 05:30:54",
      "content": "<p>Thanks! I  have started this competition very late so I thought that's the least I can do. :D</p>",
      "rawMarkdown": "Thanks! I  have started this competition very late so I thought that's the least I can do. :D",
      "votes": null
    },
    {
      "id": "1334842",
      "postDate": "06/03/2021 20:05:09",
      "content": "<p>good luck 👍</p>",
      "rawMarkdown": "good luck 👍",
      "votes": null
    },
    {
      "id": "1335535",
      "postDate": "06/04/2021 09:17:50",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> Very concise and relevant information, this is great! Upvoted. Keep up the great work!</p>",
      "rawMarkdown": "yassinealouini Very concise and relevant information, this is great! Upvoted. Keep up the great work!",
      "votes": null
    },
    {
      "id": "1335539",
      "postDate": "06/04/2021 09:19:58",
      "content": "<p>Thanks a lot, I am glad this is helpful. 😊</p>",
      "rawMarkdown": "Thanks a lot, I am glad this is helpful. 😊",
      "votes": null
    },
    {
      "id": "1336212",
      "postDate": "06/04/2021 18:04:26",
      "content": "<p>I am not sure for what (maybe for my single bad model?) but thanks. :D </p>",
      "rawMarkdown": "I am not sure for what (maybe for my single bad model?) but thanks. :D",
      "votes": null
    },
    {
      "id": "1337454",
      "postDate": "06/05/2021 16:07:40",
      "content": "<p>Thanks for the post! I always wanted to participate in the competition but wasn't sure, reading your observations will definitely help me understand it much better!</p>",
      "rawMarkdown": "Thanks for the post! I always wanted to participate in the competition but wasn't sure, reading your observations will definitely help me understand it much better!",
      "votes": null
    },
    {
      "id": "1337544",
      "postDate": "06/05/2021 17:21:28",
      "content": "<p>Awesome, I am glad this is useful to some extent! </p>",
      "rawMarkdown": "Awesome, I am glad this is useful to some extent!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1333860,
      "author_name": "yashdixit24",
      "author_url": "",
      "post_date": "06/03/2021 05:23:05",
      "content": "<p>Nice work 👍 keep it up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333868,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/03/2021 05:30:54",
          "content": "<p>Thanks! I  have started this competition very late so I thought that's the least I can do. :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1334842,
      "author_name": "satvicoder",
      "author_url": "",
      "post_date": "06/03/2021 20:05:09",
      "content": "<p>good luck 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336212,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/04/2021 18:04:26",
          "content": "<p>I am not sure for what (maybe for my single bad model?) but thanks. :D </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335535,
      "author_name": "yohmori02",
      "author_url": "",
      "post_date": "06/04/2021 09:17:50",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> Very concise and relevant information, this is great! Upvoted. Keep up the great work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335539,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/04/2021 09:19:58",
          "content": "<p>Thanks a lot, I am glad this is helpful. 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1337454,
      "author_name": "shubhamzorc9",
      "author_url": "",
      "post_date": "06/05/2021 16:07:40",
      "content": "<p>Thanks for the post! I always wanted to participate in the competition but wasn't sure, reading your observations will definitely help me understand it much better!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337544,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/05/2021 17:21:28",
          "content": "<p>Awesome, I am glad this is useful to some extent! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1327963": "I am very late to this competition but thought I should at least try to understand the problem and submit one submission. Here are some of the insights and observations I got from reading different forum discussions and searching on my own. \n\n\n## Keywords\n\n\n\n- [Skeletal formula](https://en.wikipedia.org/wiki/Skeletal_formula): a 2D representation of an organic molecule. \n\nSome interesting things to know: carbon and hydrogen atoms aren't represented but are implicit. Other than hydrogen and carbon, other additional atoms are drawn. Made by [August Kekulé](https://en.wikipedia.org/wiki/August_Kekul%C3%A9). As a fun fact: he discovered the [benzene](https://en.wikipedia.org/wiki/Benzene) structure and apparently had a dream about it inspired by the [Ouroboros](https://en.wikipedia.org/wiki/Ouroboros).\n\nHere is an image of how bonds are represented: \n\n\n![chemical bonds](https://upload.wikimedia.org/wikipedia/commons/thumb/4/48/Skeletal_formula_samples_stereochemistry.svg/300px-Skeletal_formula_samples_stereochemistry.svg.png)\n\n- InChI: short for [International Chemical Identifier](https://en.wikipedia.org/wiki/International_Chemical_Identifier) is the target to predict. It is a machine-readable chemical description of a given molecule. Here is an example for Morphine (from Wikipedia):\n\n\n![morphine](https://upload.wikimedia.org/wikipedia/commons/thumb/3/33/Morphin_-_Morphine.svg/440px-Morphin_-_Morphine.svg.png)\n\n```\nInChI=1S/C17H19NO3/c1-18-7-6-17-10-3-5-13(20)16(17)21-15-12(19)4-2-9(14(15)17)8-11(10)18/h2-5,10-11,13,16,19-20H,6-8H2,1H3/t10-,11+,13-,16-,17-/m0/s1\n```\n\n\n\n\n- [SMILES](https://en.wikipedia.org/wiki/Simplified_Molecular_Input_Line_Entry_Specification): short for Simplified Molecular Input Line Entry Specification, it is a [line notation](https://en.wikipedia.org/wiki/Line_notation) for a molecular formula using ASCII characters. Continuing with the morphine example, here is its SMILES is:\n\n```\nCN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n```\n\n\n\n- How to go from InChIkey to SMILES?\n\nUsing the InChIkey, we need a lookup table to find the corresponding SMILES representation since all the information isn't contained in the key.\n\nHere is a code snippet from a SO [thread](https://bioinformatics.stackexchange.com/questions/10755/is-there-a-python-package-to-convert-inchi-to-molecular-structures) that uses the PubChem API to do the lookup: \n\n```python\n\nimport requests\n\n\nmorphine_inchikey = \"BQJCRHHNABKAKU-KBQPJGBKSA-N\"\n\n\nr = requests.get(f'https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/inchikey/{morphine_inchikey}/property/CanonicalSMILES/JSON').json()\n\nmorphine_smiles = r['PropertyTable']['Properties'][0]['CanonicalSMILES']\n\nprint(morphine_smiles)\n\n# You should get: CN1CCC23C4C1CC5=C2C(=C(C=C5)O)OC3C(C=C4)O\n\n```\n\nTo check that the computation is correct, check the following pubchem [page](https://pubchem.ncbi.nlm.nih.gov/compound/Morphine)\n\n\n\n\n![pubchem morphine](https://drive.google.com/uc?id=19H9rFf3Z_GeUQJjSpTsSCvW1Keec-89m)\n\n- Molecular formula: a concatenation of the atoms. Continuing with the morphine example, here is its molecular morphine: `C17H19NO3`\n\n## Metric\n\n[Levenshtein distance](https://en.wikipedia.org/wiki/Levenshtein_distance): this metric is often used in NLP tasks. \n\nIndeed, it compares two strings (it is an [edit distance](https://en.wikipedia.org/wiki/Edit_distance)) `a` and `b` using the following formula (from the Wikipedia page): \n\n\n\n<img src=\"https://latex.codecogs.com/gif.latex?\\qquad\\operatorname{lev}(a,b)&space;=&space;\\begin{cases}&space;|a|&space;&&space;\\text{&space;if&space;}&space;|b|&space;=&space;0,&space;\\\\&space;|b|&space;&&space;\\text{&space;if&space;}&space;|a|&space;=&space;0,&space;\\\\&space;\\operatorname{lev}(\\operatorname{tail}(a),\\operatorname{tail}(b))&space;&&space;\\text{&space;if&space;}&space;a[0]&space;=&space;b[0]&space;\\\\&space;1&space;&plus;&space;\\min&space;\\begin{cases}&space;\\operatorname{lev}(\\operatorname{tail}(a),&space;b)&space;\\\\&space;\\operatorname{lev}(a,&space;\\operatorname{tail}(b))&space;\\\\&space;\\operatorname{lev}(\\operatorname{tail}(a),&space;\\operatorname{tail}(b))&space;\\\\&space;\\end{cases}&space;&&space;\\text{&space;otherwise.}&space;\\end{cases}\" title=\"\\qquad\\operatorname{lev}(a,b) = \\begin{cases} |a| & \\text{ if } |b| = 0, \\\\ |b| & \\text{ if } |a| = 0, \\\\ \\operatorname{lev}(\\operatorname{tail}(a),\\operatorname{tail}(b)) & \\text{ if } a[0] = b[0] \\\\ 1 + \\min \\begin{cases} \\operatorname{lev}(\\operatorname{tail}(a), b) \\\\ \\operatorname{lev}(a, \\operatorname{tail}(b)) \\\\ \\operatorname{lev}(\\operatorname{tail}(a), \\operatorname{tail}(b)) \\\\ \\end{cases} & \\text{ otherwise.} \\end{cases}\" />\n\nwhere `tail(x)` is the string without the first character and `x[i]` is the character at position i.\n\nThis is hard to parse so let's check with an example. Let's suppose that we have the following two strings (for example `a` is the true target and `b` is the predicted one): \n\n- `a = \"test\"`\n- `b = \"tefl\"`\n\nFor the first two characters, we are in the third part of the definition. Starting from the third character, we need to use the fourth part of the definition since the characters are different\n\n\nHere is another example:\n\n- `a = \"test\"`\n- `b = \"\"`\n\nSince the `b` string is empty, the metric is equal to the length of `a`, i.e. **4**. \n\n\nThere are also two bounds that can be obtained: \n\n- lower: `abs(len(a) - len(b))`\n- upper: `max(len(a), len(b))`\n\nThus, if `len(a) == len(b)`, the lower bound is 0 and is attained when `a == b`.\n\nNow that the concept is better understood, let's apply this to a target (truncated to the first few characters): \n\n- `a = \"InChI=1S/C15H24O2\"`\n- `b = \"InChI=1F/C14H23O1\"`\n\nWe can use a Python library to compute the distance: `pip install python-Levenshtein`.\n\n\n``` \n\nimport Levenshtein\n\na = \"InChI=1S/C15H24O2\"\nb = \"InChI=1F/C14H23O1\"\n\ndistance = Levenshtein.distance(a, b)\n\nprint(distance)\n\n# You should get: 4\n```\n\n## Some model ideas\n\n- extract the molecule's graph from the image then use an appropriate library to find the SMILES/InChI representation.\n- train a vision-nlp mixed model that takes the images as an input then uses some recursive network (RNN, LSTM, transformer etc) to match the textual representation.\n\n\n## Useful packages\n\n\n- [RDKit](https://www.rdkit.org/docs/index.html): everything related to chemistry and ML\n- [pysmiles](https://pypi.org/project/pysmiles/): read and write SMILES.\n- [Pytorch](https://pytorch.org/): as usual of course. ;)\n- [Transformers](https://github.com/huggingface/transformers): might be useful.\n\nand the usual pydata stack.\n\n## Similar competitions\n\n- The only one I have done so far: [CHAMPS](https://www.kaggle.com/c/champs-scalar-coupling)\n- The closest to the current competition, from the DACON [platform](https://dacon.io/competitions/official/235640/overview/description)",
    "1333860": "Nice work 👍 keep it up.",
    "1333868": "Thanks! I  have started this competition very late so I thought that's the least I can do. :D",
    "1334842": "good luck 👍",
    "1335535": "yassinealouini Very concise and relevant information, this is great! Upvoted. Keep up the great work!",
    "1335539": "Thanks a lot, I am glad this is helpful. 😊",
    "1336212": "I am not sure for what (maybe for my single bad model?) but thanks. :D",
    "1337454": "Thanks for the post! I always wanted to participate in the competition but wasn't sure, reading your observations will definitely help me understand it much better!",
    "1337544": "Awesome, I am glad this is useful to some extent!"
  },
  "source": "meta"
}