{
  "id": 224394,
  "title": "Problem with the evaluation metric",
  "url": "/competitions/bms-molecular-translation/discussion/224394",
  "author_name": "",
  "post_date": "2021-03-08T08:57:00.743616500Z",
  "votes": 35,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Scientifically, the <em>Levenshtein Distance</em> between targeted InChI string and predicted InChI string doesn't reflect the real chemical distance. One can find many \"one-pixel attack\" examples where small tweak on structure changes the entire InChI string. Another critical issue is the validity of output: not any InChI is syntactically meaningful. What's the point of predicting an <em>invalid</em> InChI even if it's highly similar to the real string?</p>\n<p>The proper metric is: 1) % exact InChI matches; or 2) mean chemical similarity or 3) Levenshtein distance for valid InChI (or whatever identifier)</p>\n<p>Below shows some interesting examples (aka, hacks): </p>\n<p>Example 1<br>\n<img src=\"https://pbs.twimg.com/media/Ev8bem6UcAQ2e45?format=png&amp;name=medium\" alt=\"\"></p>\n<p>Example 2<br>\n<img src=\"https://pbs.twimg.com/media/Ev2XsfMVEAMDdfs?format=jpg&amp;name=medium\" alt=\"\"></p>\n<blockquote>\n  <p>UPDATE: I'm less worried about the bias from direct InChI comparison, as the leaderboard is approaching LD ~ 10 with ~65% perfect predictions (aka, LD = 0).</p>\n</blockquote>",
  "messages": [
    {
      "id": "1230601",
      "postDate": "03/08/2021 08:57:00",
      "content": "<p>Scientifically, the <em>Levenshtein Distance</em> between targeted InChI string and predicted InChI string doesn't reflect the real chemical distance. One can find many \"one-pixel attack\" examples where small tweak on structure changes the entire InChI string. Another critical issue is the validity of output: not any InChI is syntactically meaningful. What's the point of predicting an <em>invalid</em> InChI even if it's highly similar to the real string?</p>\n<p>The proper metric is: 1) % exact InChI matches; or 2) mean chemical similarity or 3) Levenshtein distance for valid InChI (or whatever identifier)</p>\n<p>Below shows some interesting examples (aka, hacks): </p>\n<p>Example 1<br>\n<img src=\"https://pbs.twimg.com/media/Ev8bem6UcAQ2e45?format=png&amp;name=medium\" alt=\"\"></p>\n<p>Example 2<br>\n<img src=\"https://pbs.twimg.com/media/Ev2XsfMVEAMDdfs?format=jpg&amp;name=medium\" alt=\"\"></p>\n<blockquote>\n  <p>UPDATE: I'm less worried about the bias from direct InChI comparison, as the leaderboard is approaching LD ~ 10 with ~65% perfect predictions (aka, LD = 0).</p>\n</blockquote>",
      "rawMarkdown": "Scientifically, the *Levenshtein Distance* between targeted InChI string and predicted InChI string doesn't reflect the real chemical distance. One can find many \"one-pixel attack\" examples where small tweak on structure changes the entire InChI string. Another critical issue is the validity of output: not any InChI is syntactically meaningful. What's the point of predicting an *invalid* InChI even if it's highly similar to the real string?\n\nThe proper metric is: 1) % exact InChI matches; or 2) mean chemical similarity or 3) Levenshtein distance for valid InChI (or whatever identifier)\n\nBelow shows some interesting examples (aka, hacks): \n\nExample 1\n![](https://pbs.twimg.com/media/Ev8bem6UcAQ2e45?format=png&name=medium)\n\nExample 2\n![](https://pbs.twimg.com/media/Ev2XsfMVEAMDdfs?format=jpg&name=medium)\n\n> UPDATE: I'm less worried about the bias from direct InChI comparison, as the leaderboard is approaching LD ~ 10 with ~65% perfect predictions (aka, LD = 0).",
      "votes": null
    },
    {
      "id": "1230953",
      "postDate": "03/08/2021 15:05:42",
      "content": "<p>Great analysis! I’ll have to take a deeper look at this soon</p>",
      "rawMarkdown": "Great analysis! I’ll have to take a deeper look at this soon",
      "votes": null
    },
    {
      "id": "1231069",
      "postDate": "03/08/2021 16:39:00",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> kindly address the concern raised by <a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a> regarding the evaluation metric</p>",
      "rawMarkdown": "jakealbrecht1337 kindly address the concern raised by @houndcl regarding the evaluation metric",
      "votes": null
    },
    {
      "id": "1231178",
      "postDate": "03/08/2021 18:19:40",
      "content": "<p>I don't think that the outputs are going to be used for purposes where predictions must be valid InChIs.<br>\nI can think of use cases like searching the most similar substances to a query string amongst millions of predictions to images. Then you can look at the papers where the most similar ones came from.</p>",
      "rawMarkdown": "I don't think that the outputs are going to be used for purposes where predictions must be valid InChIs.\nI can think of use cases like searching the most similar substances to a query string amongst millions of predictions to images. Then you can look at the papers where the most similar ones came from.",
      "votes": null
    },
    {
      "id": "1231312",
      "postDate": "03/08/2021 21:36:08",
      "content": "<p>Thank you for your question!  Yes, the choice of a metric was a challenge.  The metric of fraction exactly matched was initially proposed, but we wanted to give more partial credit as in your examples.  Other metrics like Tanimoto (Jaccard) similarity or checks for valid InChIs were considered but would have been too demanding for automatic scoring.  Because the internal structure of InChIs makes them a poor target for direct prediction, ideally the winning submission would convert them to another format (e.g. <code>deepsmiles</code>) or molecular graphs prior to training, with the predictions converted back to InChI.  For these reasons, the risk of hacking the leaderboard was deemed low and a simple distance metric was used.</p>",
      "rawMarkdown": "Thank you for your question!  Yes, the choice of a metric was a challenge.  The metric of fraction exactly matched was initially proposed, but we wanted to give more partial credit as in your examples.  Other metrics like Tanimoto (Jaccard) similarity or checks for valid InChIs were considered but would have been too demanding for automatic scoring.  Because the internal structure of InChIs makes them a poor target for direct prediction, ideally the winning submission would convert them to another format (e.g. `deepsmiles`) or molecular graphs prior to training, with the predictions converted back to InChI.  For these reasons, the risk of hacking the leaderboard was deemed low and a simple distance metric was used.",
      "votes": null
    },
    {
      "id": "1231353",
      "postDate": "03/08/2021 22:17:40",
      "content": "<p>Thank you for the comprehensive response! :)</p>",
      "rawMarkdown": "Thank you for the comprehensive response! :)",
      "votes": null
    },
    {
      "id": "1232604",
      "postDate": "03/09/2021 21:16:33",
      "content": "<p>Thanks for the thorough response! With the best model pushing LD closer to 10, most are probably perfect predictions. <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> has confirmed this in his thread (65% LD = 0). This is amazing. I originally thought that the performance wouldn't be this good (e.g., LD ~ 40), which are more likely to suffer the mentioned bias from direct InChI comparison. </p>",
      "rawMarkdown": "Thanks for the thorough response! With the best model pushing LD closer to 10, most are probably perfect predictions. @drhabib has confirmed this in his thread (65% LD = 0). This is amazing. I originally thought that the performance wouldn't be this good (e.g., LD ~ 40), which are more likely to suffer the mentioned bias from direct InChI comparison.",
      "votes": null
    },
    {
      "id": "1235130",
      "postDate": "03/11/2021 21:19:35",
      "content": "<p>I've thought about this and there are some implications for predicting the InChI directly vs using a format like SMILES and converting. With the conversion approach, you can have your model predict a shorter, simpler sequence and then convert over, which feels easier. But this relies on having the ability to convert. If you use SMILES for example, you can have a single character syntax error that makes the entire string invalid, you lose the whole sample because you can't convert back to an InChI.</p>\n<p>Predicting InChI strings directly is more complicated (longer sequence to predict) but you're much more likely to get partially correct answers, which might perform better with the current metric.</p>",
      "rawMarkdown": "I've thought about this and there are some implications for predicting the InChI directly vs using a format like SMILES and converting. With the conversion approach, you can have your model predict a shorter, simpler sequence and then convert over, which feels easier. But this relies on having the ability to convert. If you use SMILES for example, you can have a single character syntax error that makes the entire string invalid, you lose the whole sample because you can't convert back to an InChI.\n\nPredicting InChI strings directly is more complicated (longer sequence to predict) but you're much more likely to get partially correct answers, which might perform better with the current metric.",
      "votes": null
    },
    {
      "id": "1235346",
      "postDate": "03/12/2021 04:32:43",
      "content": "<p>The layered structure of the InChI itself makes it amenable to some levels of fuzzy matching. For example a comparison could be done solely between the main layer (<a href=\"https://www.inchi-trust.org/technical-faq-2/#4.3\" target=\"_blank\">https://www.inchi-trust.org/technical-faq-2/#4.3</a>) by string splitting.<br>\nTechnically you could get even fuzzier and consider just the molecular formula (i.e. have the labels in the depiction been succesfully identified), albeit I expect you'd need to ignore hydrogen count, as if the connection/hydrogen layers are wrong this count is likely also wrong.</p>\n<p>From a cheminformatics viewpoint, ideally the dataset should have both SMILES and InChI, even if the InChI is the evaluation target. InChI, by design, considers difference resonance forms or tautomers to have the same InChI, but this distinction is made in chemical depictions (and also SMILES). <br>\ni.e. conversion from InChI to SMILES is a lossy conversion as there are multiple <strong>canonical</strong> SMILES for a single InChI.</p>\n<p>In practice the depictions appear to have been generated from the deterministically chosen structure you get if you convert an InChI to structure, so this is somewhat of a moot issue. One visible quirk of this is that the tautomer chosen for amides is not the typical one, hence you have C(O)=N instead of C(=O)N</p>",
      "rawMarkdown": "The layered structure of the InChI itself makes it amenable to some levels of fuzzy matching. For example a comparison could be done solely between the main layer (https://www.inchi-trust.org/technical-faq-2/#4.3) by string splitting.\nTechnically you could get even fuzzier and consider just the molecular formula (i.e. have the labels in the depiction been succesfully identified), albeit I expect you'd need to ignore hydrogen count, as if the connection/hydrogen layers are wrong this count is likely also wrong.\n\nFrom a cheminformatics viewpoint, ideally the dataset should have both SMILES and InChI, even if the InChI is the evaluation target. InChI, by design, considers difference resonance forms or tautomers to have the same InChI, but this distinction is made in chemical depictions (and also SMILES). \ni.e. conversion from InChI to SMILES is a lossy conversion as there are multiple **canonical** SMILES for a single InChI.\n\nIn practice the depictions appear to have been generated from the deterministically chosen structure you get if you convert an InChI to structure, so this is somewhat of a moot issue. One visible quirk of this is that the tautomer chosen for amides is not the typical one, hence you have C(O)=N instead of C(=O)N",
      "votes": null
    },
    {
      "id": "1235606",
      "postDate": "03/12/2021 09:45:11",
      "content": "<p>indeed InChi to canonical SMILES is a lossy compression, but in the train dataset you get &gt;99.99% valid InChi's when converting to SMILES and back. This discrepancies might become relevant when squeezing the last bits of performance towards the end of the competition</p>",
      "rawMarkdown": "indeed InChi to canonical SMILES is a lossy compression, but in the train dataset you get >99.99% valid InChi's when converting to SMILES and back. This discrepancies might become relevant when squeezing the last bits of performance towards the end of the competition",
      "votes": null
    },
    {
      "id": "1236214",
      "postDate": "03/12/2021 22:15:14",
      "content": "<p>Good discussion on the challenges of predicting invalid chemical strings.  I'm curious: have teams explored representations of the molecules based on generative models e.g. <a href=\"https://deepchem.readthedocs.io/en/latest/api_reference/featurizers.html#molganfeaturizer\" target=\"_blank\">MolGAN</a> ?</p>",
      "rawMarkdown": "Good discussion on the challenges of predicting invalid chemical strings.  I'm curious: have teams explored representations of the molecules based on generative models e.g. [MolGAN](https://deepchem.readthedocs.io/en/latest/api_reference/featurizers.html#molganfeaturizer) ?",
      "votes": null
    },
    {
      "id": "1238230",
      "postDate": "03/14/2021 19:04:22",
      "content": "<p>I published a baseline notebook to show that <strong>we can get LB 71.1 by always predicting the same InChI</strong> :<br>\n<a href=\"https://www.kaggle.com/datanewb/always-predict-same-inchi/\" target=\"_blank\">https://www.kaggle.com/datanewb/always-predict-same-inchi/</a><br>\nI found this \"central\" InChI with a method similar to <a href=\"https://rawgit.com/ztane/python-Levenshtein/master/docs/Levenshtein.html#Levenshtein-setmedian\" target=\"_blank\">https://rawgit.com/ztane/python-Levenshtein/master/docs/Levenshtein.html#Levenshtein-setmedian</a><br>\nThis set-median InChI is from the training InChIs but we could do even better than 71.1 with a median invalid InChI.<br>\nBut finding a median string among a set of strings is an NP-complete problem :<br>\n\"The generalized median is the more general concept and therefore provides a better representation of the set than the set median, but de la Higuera and Casacuberta [3] proved that the computation of the former is NP-complete, in other words very hard to compute.\" (c.f. <a href=\"http://www.ok.sc.e.titech.ac.jp/res/PCS/courses/comp644/\" target=\"_blank\">http://www.ok.sc.e.titech.ac.jp/res/PCS/courses/comp644/</a>)<br>\nThis is just to provide a constant prediction baseline and with top LB scores already below 3 it is impossible to \"hack\" the leaderboard with a median InChI.</p>",
      "rawMarkdown": "I published a baseline notebook to show that **we can get LB 71.1 by always predicting the same InChI** :\nhttps://www.kaggle.com/datanewb/always-predict-same-inchi/\nI found this \"central\" InChI with a method similar to https://rawgit.com/ztane/python-Levenshtein/master/docs/Levenshtein.html#Levenshtein-setmedian\nThis set-median InChI is from the training InChIs but we could do even better than 71.1 with a median invalid InChI.\nBut finding a median string among a set of strings is an NP-complete problem :\n\"The generalized median is the more general concept and therefore provides a better representation of the set than the set median, but de la Higuera and Casacuberta [3] proved that the computation of the former is NP-complete, in other words very hard to compute.\" (c.f. http://www.ok.sc.e.titech.ac.jp/res/PCS/courses/comp644/)\nThis is just to provide a constant prediction baseline and with top LB scores already below 3 it is impossible to \"hack\" the leaderboard with a median InChI.",
      "votes": null
    },
    {
      "id": "1238483",
      "postDate": "03/15/2021 03:03:15",
      "content": "<p>I do, i'm very instered on using <strong>latent spaces</strong> from <strong>graph autoencoders</strong>. I'm not sure if it would be the right approach but i'm going to learn a lot doing it.</p>",
      "rawMarkdown": "I do, i'm very instered on using **latent spaces** from **graph autoencoders**. I'm not sure if it would be the right approach but i'm going to learn a lot doing it.",
      "votes": null
    },
    {
      "id": "1244270",
      "postDate": "03/18/2021 21:13:16",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> \"the winning submission would convert them to another format (e.g. deepsmiles) or molecular graphs prior to training, with the predictions converted back to InChI\"<br>\nI would like to do as you say i.e. predict some intermediate representation for the molecular graph but the big issue is that <strong>we then miss a clear algorithm to produce a unique InChI from any molecular graph</strong>.</p>",
      "rawMarkdown": "jakealbrecht1337 \"the winning submission would convert them to another format (e.g. deepsmiles) or molecular graphs prior to training, with the predictions converted back to InChI\"\nI would like to do as you say i.e. predict some intermediate representation for the molecular graph but the big issue is that **we then miss a clear algorithm to produce a unique InChI from any molecular graph**.",
      "votes": null
    },
    {
      "id": "1244359",
      "postDate": "03/19/2021 00:12:46",
      "content": "<p><a href=\"https://www.kaggle.com/datanewb\" target=\"_blank\">@datanewb</a> Do you mean something like <a href=\"https://www.rdkit.org/docs/source/rdkit.Chem.inchi.html#rdkit.Chem.inchi.MolToInchi\" target=\"_blank\">https://www.rdkit.org/docs/source/rdkit.Chem.inchi.html#rdkit.Chem.inchi.MolToInchi</a> ?</p>\n<p>InChIs are produced using a specific algorithm, that for a given molecular graph (and variants of the graph that are chemically equivalent)  produces the same InChI. Partially due to the complexity of the algorithm, and partially to guarantee that the same algorithm is used, there is only one implementation. This library is then bundled with most of the major chemistry toolkits e.g. RDKit, Epam's Indigo, Chemistry Development Kit (CDK), and hence for a given molecular graph any of these will give the same InChI</p>\n<p>SELFIES (<a href=\"https://github.com/aspuru-guzik-group/selfies\" target=\"_blank\">https://github.com/aspuru-guzik-group/selfies</a>) are also worth considering as the intermediate representation.</p>",
      "rawMarkdown": "datanewb Do you mean something like https://www.rdkit.org/docs/source/rdkit.Chem.inchi.html#rdkit.Chem.inchi.MolToInchi ?\n\nInChIs are produced using a specific algorithm, that for a given molecular graph (and variants of the graph that are chemically equivalent)  produces the same InChI. Partially due to the complexity of the algorithm, and partially to guarantee that the same algorithm is used, there is only one implementation. This library is then bundled with most of the major chemistry toolkits e.g. RDKit, Epam's Indigo, Chemistry Development Kit (CDK), and hence for a given molecular graph any of these will give the same InChI\n\nSELFIES (https://github.com/aspuru-guzik-group/selfies) are also worth considering as the intermediate representation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1230953,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "03/08/2021 15:05:42",
      "content": "<p>Great analysis! I’ll have to take a deeper look at this soon</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1231069,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/08/2021 16:39:00",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> kindly address the concern raised by <a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a> regarding the evaluation metric</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1231178,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "03/08/2021 18:19:40",
      "content": "<p>I don't think that the outputs are going to be used for purposes where predictions must be valid InChIs.<br>\nI can think of use cases like searching the most similar substances to a query string amongst millions of predictions to images. Then you can look at the papers where the most similar ones came from.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1231312,
      "author_name": "jakealbrecht1337",
      "author_url": "",
      "post_date": "03/08/2021 21:36:08",
      "content": "<p>Thank you for your question!  Yes, the choice of a metric was a challenge.  The metric of fraction exactly matched was initially proposed, but we wanted to give more partial credit as in your examples.  Other metrics like Tanimoto (Jaccard) similarity or checks for valid InChIs were considered but would have been too demanding for automatic scoring.  Because the internal structure of InChIs makes them a poor target for direct prediction, ideally the winning submission would convert them to another format (e.g. <code>deepsmiles</code>) or molecular graphs prior to training, with the predictions converted back to InChI.  For these reasons, the risk of hacking the leaderboard was deemed low and a simple distance metric was used.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1231353,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/08/2021 22:17:40",
          "content": "<p>Thank you for the comprehensive response! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1232604,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "03/09/2021 21:16:33",
          "content": "<p>Thanks for the thorough response! With the best model pushing LD closer to 10, most are probably perfect predictions. <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> has confirmed this in his thread (65% LD = 0). This is amazing. I originally thought that the performance wouldn't be this good (e.g., LD ~ 40), which are more likely to suffer the mentioned bias from direct InChI comparison. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1235130,
          "author_name": "towardsentropy",
          "author_url": "",
          "post_date": "03/11/2021 21:19:35",
          "content": "<p>I've thought about this and there are some implications for predicting the InChI directly vs using a format like SMILES and converting. With the conversion approach, you can have your model predict a shorter, simpler sequence and then convert over, which feels easier. But this relies on having the ability to convert. If you use SMILES for example, you can have a single character syntax error that makes the entire string invalid, you lose the whole sample because you can't convert back to an InChI.</p>\n<p>Predicting InChI strings directly is more complicated (longer sequence to predict) but you're much more likely to get partially correct answers, which might perform better with the current metric.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1235346,
          "author_name": "infy2097",
          "author_url": "",
          "post_date": "03/12/2021 04:32:43",
          "content": "<p>The layered structure of the InChI itself makes it amenable to some levels of fuzzy matching. For example a comparison could be done solely between the main layer (<a href=\"https://www.inchi-trust.org/technical-faq-2/#4.3\" target=\"_blank\">https://www.inchi-trust.org/technical-faq-2/#4.3</a>) by string splitting.<br>\nTechnically you could get even fuzzier and consider just the molecular formula (i.e. have the labels in the depiction been succesfully identified), albeit I expect you'd need to ignore hydrogen count, as if the connection/hydrogen layers are wrong this count is likely also wrong.</p>\n<p>From a cheminformatics viewpoint, ideally the dataset should have both SMILES and InChI, even if the InChI is the evaluation target. InChI, by design, considers difference resonance forms or tautomers to have the same InChI, but this distinction is made in chemical depictions (and also SMILES). <br>\ni.e. conversion from InChI to SMILES is a lossy conversion as there are multiple <strong>canonical</strong> SMILES for a single InChI.</p>\n<p>In practice the depictions appear to have been generated from the deterministically chosen structure you get if you convert an InChI to structure, so this is somewhat of a moot issue. One visible quirk of this is that the tautomer chosen for amides is not the typical one, hence you have C(O)=N instead of C(=O)N</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1235606,
          "author_name": "bulatz",
          "author_url": "",
          "post_date": "03/12/2021 09:45:11",
          "content": "<p>indeed InChi to canonical SMILES is a lossy compression, but in the train dataset you get &gt;99.99% valid InChi's when converting to SMILES and back. This discrepancies might become relevant when squeezing the last bits of performance towards the end of the competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1236214,
          "author_name": "jakealbrecht1337",
          "author_url": "",
          "post_date": "03/12/2021 22:15:14",
          "content": "<p>Good discussion on the challenges of predicting invalid chemical strings.  I'm curious: have teams explored representations of the molecules based on generative models e.g. <a href=\"https://deepchem.readthedocs.io/en/latest/api_reference/featurizers.html#molganfeaturizer\" target=\"_blank\">MolGAN</a> ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238483,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "03/15/2021 03:03:15",
          "content": "<p>I do, i'm very instered on using <strong>latent spaces</strong> from <strong>graph autoencoders</strong>. I'm not sure if it would be the right approach but i'm going to learn a lot doing it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1244270,
          "author_name": "datanewb",
          "author_url": "",
          "post_date": "03/18/2021 21:13:16",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> \"the winning submission would convert them to another format (e.g. deepsmiles) or molecular graphs prior to training, with the predictions converted back to InChI\"<br>\nI would like to do as you say i.e. predict some intermediate representation for the molecular graph but the big issue is that <strong>we then miss a clear algorithm to produce a unique InChI from any molecular graph</strong>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1244359,
          "author_name": "infy2097",
          "author_url": "",
          "post_date": "03/19/2021 00:12:46",
          "content": "<p><a href=\"https://www.kaggle.com/datanewb\" target=\"_blank\">@datanewb</a> Do you mean something like <a href=\"https://www.rdkit.org/docs/source/rdkit.Chem.inchi.html#rdkit.Chem.inchi.MolToInchi\" target=\"_blank\">https://www.rdkit.org/docs/source/rdkit.Chem.inchi.html#rdkit.Chem.inchi.MolToInchi</a> ?</p>\n<p>InChIs are produced using a specific algorithm, that for a given molecular graph (and variants of the graph that are chemically equivalent)  produces the same InChI. Partially due to the complexity of the algorithm, and partially to guarantee that the same algorithm is used, there is only one implementation. This library is then bundled with most of the major chemistry toolkits e.g. RDKit, Epam's Indigo, Chemistry Development Kit (CDK), and hence for a given molecular graph any of these will give the same InChI</p>\n<p>SELFIES (<a href=\"https://github.com/aspuru-guzik-group/selfies\" target=\"_blank\">https://github.com/aspuru-guzik-group/selfies</a>) are also worth considering as the intermediate representation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1238230,
      "author_name": "datanewb",
      "author_url": "",
      "post_date": "03/14/2021 19:04:22",
      "content": "<p>I published a baseline notebook to show that <strong>we can get LB 71.1 by always predicting the same InChI</strong> :<br>\n<a href=\"https://www.kaggle.com/datanewb/always-predict-same-inchi/\" target=\"_blank\">https://www.kaggle.com/datanewb/always-predict-same-inchi/</a><br>\nI found this \"central\" InChI with a method similar to <a href=\"https://rawgit.com/ztane/python-Levenshtein/master/docs/Levenshtein.html#Levenshtein-setmedian\" target=\"_blank\">https://rawgit.com/ztane/python-Levenshtein/master/docs/Levenshtein.html#Levenshtein-setmedian</a><br>\nThis set-median InChI is from the training InChIs but we could do even better than 71.1 with a median invalid InChI.<br>\nBut finding a median string among a set of strings is an NP-complete problem :<br>\n\"The generalized median is the more general concept and therefore provides a better representation of the set than the set median, but de la Higuera and Casacuberta [3] proved that the computation of the former is NP-complete, in other words very hard to compute.\" (c.f. <a href=\"http://www.ok.sc.e.titech.ac.jp/res/PCS/courses/comp644/\" target=\"_blank\">http://www.ok.sc.e.titech.ac.jp/res/PCS/courses/comp644/</a>)<br>\nThis is just to provide a constant prediction baseline and with top LB scores already below 3 it is impossible to \"hack\" the leaderboard with a median InChI.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1230601": "Scientifically, the *Levenshtein Distance* between targeted InChI string and predicted InChI string doesn't reflect the real chemical distance. One can find many \"one-pixel attack\" examples where small tweak on structure changes the entire InChI string. Another critical issue is the validity of output: not any InChI is syntactically meaningful. What's the point of predicting an *invalid* InChI even if it's highly similar to the real string?\n\nThe proper metric is: 1) % exact InChI matches; or 2) mean chemical similarity or 3) Levenshtein distance for valid InChI (or whatever identifier)\n\nBelow shows some interesting examples (aka, hacks): \n\nExample 1\n![](https://pbs.twimg.com/media/Ev8bem6UcAQ2e45?format=png&name=medium)\n\nExample 2\n![](https://pbs.twimg.com/media/Ev2XsfMVEAMDdfs?format=jpg&name=medium)\n\n> UPDATE: I'm less worried about the bias from direct InChI comparison, as the leaderboard is approaching LD ~ 10 with ~65% perfect predictions (aka, LD = 0).",
    "1230953": "Great analysis! I’ll have to take a deeper look at this soon",
    "1231069": "jakealbrecht1337 kindly address the concern raised by @houndcl regarding the evaluation metric",
    "1231178": "I don't think that the outputs are going to be used for purposes where predictions must be valid InChIs.\nI can think of use cases like searching the most similar substances to a query string amongst millions of predictions to images. Then you can look at the papers where the most similar ones came from.",
    "1231312": "Thank you for your question!  Yes, the choice of a metric was a challenge.  The metric of fraction exactly matched was initially proposed, but we wanted to give more partial credit as in your examples.  Other metrics like Tanimoto (Jaccard) similarity or checks for valid InChIs were considered but would have been too demanding for automatic scoring.  Because the internal structure of InChIs makes them a poor target for direct prediction, ideally the winning submission would convert them to another format (e.g. `deepsmiles`) or molecular graphs prior to training, with the predictions converted back to InChI.  For these reasons, the risk of hacking the leaderboard was deemed low and a simple distance metric was used.",
    "1231353": "Thank you for the comprehensive response! :)",
    "1232604": "Thanks for the thorough response! With the best model pushing LD closer to 10, most are probably perfect predictions. @drhabib has confirmed this in his thread (65% LD = 0). This is amazing. I originally thought that the performance wouldn't be this good (e.g., LD ~ 40), which are more likely to suffer the mentioned bias from direct InChI comparison.",
    "1235130": "I've thought about this and there are some implications for predicting the InChI directly vs using a format like SMILES and converting. With the conversion approach, you can have your model predict a shorter, simpler sequence and then convert over, which feels easier. But this relies on having the ability to convert. If you use SMILES for example, you can have a single character syntax error that makes the entire string invalid, you lose the whole sample because you can't convert back to an InChI.\n\nPredicting InChI strings directly is more complicated (longer sequence to predict) but you're much more likely to get partially correct answers, which might perform better with the current metric.",
    "1235346": "The layered structure of the InChI itself makes it amenable to some levels of fuzzy matching. For example a comparison could be done solely between the main layer (https://www.inchi-trust.org/technical-faq-2/#4.3) by string splitting.\nTechnically you could get even fuzzier and consider just the molecular formula (i.e. have the labels in the depiction been succesfully identified), albeit I expect you'd need to ignore hydrogen count, as if the connection/hydrogen layers are wrong this count is likely also wrong.\n\nFrom a cheminformatics viewpoint, ideally the dataset should have both SMILES and InChI, even if the InChI is the evaluation target. InChI, by design, considers difference resonance forms or tautomers to have the same InChI, but this distinction is made in chemical depictions (and also SMILES). \ni.e. conversion from InChI to SMILES is a lossy conversion as there are multiple **canonical** SMILES for a single InChI.\n\nIn practice the depictions appear to have been generated from the deterministically chosen structure you get if you convert an InChI to structure, so this is somewhat of a moot issue. One visible quirk of this is that the tautomer chosen for amides is not the typical one, hence you have C(O)=N instead of C(=O)N",
    "1235606": "indeed InChi to canonical SMILES is a lossy compression, but in the train dataset you get >99.99% valid InChi's when converting to SMILES and back. This discrepancies might become relevant when squeezing the last bits of performance towards the end of the competition",
    "1236214": "Good discussion on the challenges of predicting invalid chemical strings.  I'm curious: have teams explored representations of the molecules based on generative models e.g. [MolGAN](https://deepchem.readthedocs.io/en/latest/api_reference/featurizers.html#molganfeaturizer) ?",
    "1238230": "I published a baseline notebook to show that **we can get LB 71.1 by always predicting the same InChI** :\nhttps://www.kaggle.com/datanewb/always-predict-same-inchi/\nI found this \"central\" InChI with a method similar to https://rawgit.com/ztane/python-Levenshtein/master/docs/Levenshtein.html#Levenshtein-setmedian\nThis set-median InChI is from the training InChIs but we could do even better than 71.1 with a median invalid InChI.\nBut finding a median string among a set of strings is an NP-complete problem :\n\"The generalized median is the more general concept and therefore provides a better representation of the set than the set median, but de la Higuera and Casacuberta [3] proved that the computation of the former is NP-complete, in other words very hard to compute.\" (c.f. http://www.ok.sc.e.titech.ac.jp/res/PCS/courses/comp644/)\nThis is just to provide a constant prediction baseline and with top LB scores already below 3 it is impossible to \"hack\" the leaderboard with a median InChI.",
    "1238483": "I do, i'm very instered on using **latent spaces** from **graph autoencoders**. I'm not sure if it would be the right approach but i'm going to learn a lot doing it.",
    "1244270": "jakealbrecht1337 \"the winning submission would convert them to another format (e.g. deepsmiles) or molecular graphs prior to training, with the predictions converted back to InChI\"\nI would like to do as you say i.e. predict some intermediate representation for the molecular graph but the big issue is that **we then miss a clear algorithm to produce a unique InChI from any molecular graph**.",
    "1244359": "datanewb Do you mean something like https://www.rdkit.org/docs/source/rdkit.Chem.inchi.html#rdkit.Chem.inchi.MolToInchi ?\n\nInChIs are produced using a specific algorithm, that for a given molecular graph (and variants of the graph that are chemically equivalent)  produces the same InChI. Partially due to the complexity of the algorithm, and partially to guarantee that the same algorithm is used, there is only one implementation. This library is then bundled with most of the major chemistry toolkits e.g. RDKit, Epam's Indigo, Chemistry Development Kit (CDK), and hence for a given molecular graph any of these will give the same InChI\n\nSELFIES (https://github.com/aspuru-guzik-group/selfies) are also worth considering as the intermediate representation."
  },
  "source": "meta"
}