{
  "id": 230185,
  "title": "Generated InCHI does not need to be valid",
  "url": "/competitions/bms-molecular-translation/discussion/230185",
  "author_name": "",
  "post_date": "2021-04-02T12:42:17.370600900Z",
  "votes": 6,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Title says it all.</p>\n<p>This maybe why it is possible to reach very low edit distance.  It can be better to output a wrong InCHI close to ground truth than a valid InCHI farther from it.  </p>",
  "messages": [
    {
      "id": "1260831",
      "postDate": "04/02/2021 12:42:17",
      "content": "<p>Title says it all.</p>\n<p>This maybe why it is possible to reach very low edit distance.  It can be better to output a wrong InCHI close to ground truth than a valid InCHI farther from it.  </p>",
      "rawMarkdown": "Title says it all.\n\nThis maybe why it is possible to reach very low edit distance.  It can be better to output a wrong InCHI close to ground truth than a valid InCHI farther from it.",
      "votes": null
    },
    {
      "id": "1261043",
      "postDate": "04/02/2021 16:39:58",
      "content": "<p>Do you mean normalizing with rdkit? There are other ways to normalize which helps a lot.</p>",
      "rawMarkdown": "Do you mean normalizing with rdkit? There are other ways to normalize which helps a lot.",
      "votes": null
    },
    {
      "id": "1261060",
      "postDate": "04/02/2021 16:58:42",
      "content": "<p>InCHI standard is unique for a given molecule.  My point is that we are not asked to predict valid InCHI standard string.</p>",
      "rawMarkdown": "InCHI standard is unique for a given molecule.  My point is that we are not asked to predict valid InCHI standard string.",
      "votes": null
    },
    {
      "id": "1261113",
      "postDate": "04/02/2021 17:49:48",
      "content": "<p>one may one to check this interactive inchi generator<br>\n<a href=\"http://www.cheminfo.org/Chemistry/Cheminformatics/Generate_InChI/index.html\" target=\"_blank\">http://www.cheminfo.org/Chemistry/Cheminformatics/Generate_InChI/index.html</a></p>\n<p>i think we are going to have great problems for generalisation?</p>\n<p><img src=\"https://i.ibb.co/znLtcGB/Selection-030.png\" alt=\"\"></p>\n<p>i am looking for API to convert from table of atoms connections to INCHI … this make more sense</p>",
      "rawMarkdown": "one may one to check this interactive inchi generator\nhttp://www.cheminfo.org/Chemistry/Cheminformatics/Generate_InChI/index.html\n\ni think we are going to have great problems for generalisation?\n\n![](https://i.ibb.co/znLtcGB/Selection-030.png)\n\ni am looking for API to convert from table of atoms connections to INCHI ... this make more sense",
      "votes": null
    },
    {
      "id": "1261121",
      "postDate": "04/02/2021 17:57:15",
      "content": "<p>I was using pubchem to normalize/validate the predictions which brings a big boost but it is now forbidden: <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313</a><br>\nSeems also the 1st place team used a similar approach.</p>",
      "rawMarkdown": "I was using pubchem to normalize/validate the predictions which brings a big boost but it is now forbidden: https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313\nSeems also the 1st place team used a similar approach.",
      "votes": null
    },
    {
      "id": "1261144",
      "postDate": "04/02/2021 18:15:58",
      "content": "<p>it is forbidden as a solution, but you can still use it for analysis of results and think of another accepted way that achieve the same purpose</p>",
      "rawMarkdown": "it is forbidden as a solution, but you can still use it for analysis of results and think of another accepted way that achieve the same purpose",
      "votes": null
    },
    {
      "id": "1261200",
      "postDate": "04/02/2021 19:12:05",
      "content": "<p>Can you please pin point where its says that we can't normalize our <code>prediction</code> (InChi to InChI using <code>rdkit</code>) ?. The only think I see is following:  <code>Fixing predictions to nearest published InChIs is not allowed.</code></p>",
      "rawMarkdown": "Can you please pin point where its says that we can't normalize our `prediction` (InChi to InChI using `rdkit`) ?. The only think I see is following:  `Fixing predictions to nearest published InChIs is not allowed. `",
      "votes": null
    },
    {
      "id": "1261201",
      "postDate": "04/02/2021 19:12:46",
      "content": "<p>This was also discussed here <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/224394</a></p>\n<p>With the top-scoring solutions now scoring so lowly the majority of results in these submissions need to be edit distance 0 and hence neccesarily have to be valid InChIs.</p>",
      "rawMarkdown": "This was also discussed here https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\n\nWith the top-scoring solutions now scoring so lowly the majority of results in these submissions need to be edit distance 0 and hence neccesarily have to be valid InChIs.",
      "votes": null
    },
    {
      "id": "1261239",
      "postDate": "04/02/2021 20:53:30",
      "content": "<p>Definitely the title speaks the truth. I understand (if it's not now forbidden knowledge) that the best scoring public kernels predict only small proportions of valid InChIs. I have given my opinions on the new rule interpretations <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229279\" target=\"_blank\">elsewhere</a>, but clearly predicting a valid InChI is in itself a significant challenge.</p>",
      "rawMarkdown": "Definitely the title speaks the truth. I understand (if it's not now forbidden knowledge) that the best scoring public kernels predict only small proportions of valid InChIs. I have given my opinions on the new rule interpretations [elsewhere](https://www.kaggle.com/c/bms-molecular-translation/discussion/229279), but clearly predicting a valid InChI is in itself a significant challenge.",
      "votes": null
    },
    {
      "id": "1261697",
      "postDate": "04/03/2021 11:03:07",
      "content": "<blockquote>\n  <p>This was also discussed here <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/224394</a></p>\n</blockquote>\n<p>I read that discussion, and it is mostly about the metric.  </p>\n<p>I think it was good to highlight what I discuss here separately.</p>",
      "rawMarkdown": "> This was also discussed here https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\n\nI read that discussion, and it is mostly about the metric.  \n\nI think it was good to highlight what I discuss here separately.",
      "votes": null
    },
    {
      "id": "1261724",
      "postDate": "04/03/2021 11:34:31",
      "content": "<p>Your discussion was very helpful to me.</p>",
      "rawMarkdown": "Your discussion was very helpful to me.",
      "votes": null
    },
    {
      "id": "1261739",
      "postDate": "04/03/2021 11:53:13",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  <a href=\"https://www.kaggle.com/infy2097\" target=\"_blank\">@infy2097</a> - can we agree that you are both right? Invalid InChIs are accepted by the scoring algorithm here and scored by edit distance (but rejected outright by every other piece of chemistry software I can find), but it is also true that to get to the top of this leaderboard you will need a lot of perfect InChIs.</p>",
      "rawMarkdown": "cpmpml  @infy2097 - can we agree that you are both right? Invalid InChIs are accepted by the scoring algorithm here and scored by edit distance (but rejected outright by every other piece of chemistry software I can find), but it is also true that to get to the top of this leaderboard you will need a lot of perfect InChIs.",
      "votes": null
    },
    {
      "id": "1261745",
      "postDate": "04/03/2021 11:58:10",
      "content": "<p>I edited my previous answer to make it clearer.  I don't see any disagreement,</p>",
      "rawMarkdown": "I edited my previous answer to make it clearer.  I don't see any disagreement,",
      "votes": null
    },
    {
      "id": "1261811",
      "postDate": "04/03/2021 13:27:16",
      "content": "<blockquote>\n  <p>Seems also the 1st place team used a similar approach.</p>\n</blockquote>\n<p>We haven't done this.</p>",
      "rawMarkdown": "> Seems also the 1st place team used a similar approach.\n\nWe haven't done this.",
      "votes": null
    },
    {
      "id": "1261822",
      "postDate": "04/03/2021 13:38:48",
      "content": "<p>I'm not using any external data, too. I think it is possible to reach 1.xx.</p>",
      "rawMarkdown": "I'm not using any external data, too. I think it is possible to reach 1.xx.",
      "votes": null
    },
    {
      "id": "1261826",
      "postDate": "04/03/2021 13:41:42",
      "content": "<p>Using pubchem brings up to +3 boost.</p>",
      "rawMarkdown": "Using pubchem brings up to +3 boost.",
      "votes": null
    },
    {
      "id": "1261837",
      "postDate": "04/03/2021 13:52:22",
      "content": "<p>we are also not using any external data. </p>",
      "rawMarkdown": "we are also not using any external data.",
      "votes": null
    },
    {
      "id": "1261867",
      "postDate": "04/03/2021 14:29:17",
      "content": "<p><a href=\"https://www.kaggle.com/DrHB\" target=\"_blank\">@DrHB</a> you keep listing what you don't use, when will you disclose what you actually use ?</p>\n<p>;)</p>",
      "rawMarkdown": "DrHB you keep listing what you don't use, when will you disclose what you actually use ?\n\n;)",
      "votes": null
    },
    {
      "id": "1261871",
      "postDate": "04/03/2021 14:31:26",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> just had quick chat with my teammates… we agreed that we will disclose on June 3rd 2021 =)  </p>",
      "rawMarkdown": "cpmpml just had quick chat with my teammates... we agreed that we will disclose on June 3rd 2021 =)",
      "votes": null
    },
    {
      "id": "1261878",
      "postDate": "04/03/2021 14:38:14",
      "content": "<p>I hope I'll have submitted by then…</p>",
      "rawMarkdown": "I hope I'll have submitted by then...",
      "votes": null
    },
    {
      "id": "1261887",
      "postDate": "04/03/2021 14:44:04",
      "content": "<p>I'll come back here once host settled on the external data and matching with known molecules.  </p>",
      "rawMarkdown": "I'll come back here once host settled on the external data and matching with known molecules.",
      "votes": null
    },
    {
      "id": "1262941",
      "postDate": "04/04/2021 22:43:46",
      "content": "<p><code>Fixing predictions to nearest published InChIs is not allowed</code> but <br>\n<code>Fixing predictions to the exact published InChI (aka whitelist search)</code> seems allowed. </p>\n<p>e.g. </p>\n<ul>\n<li>output: abcd1 -&gt; nearest PubChem entry is abcd2, then I submit <code>abcd2</code> (forbidden!)</li>\n<li>output: abcd1 (p=0.5), abcd2 (p=0.3), abcd3 (p=0.2) -&gt; I find that exact PubChem match is abcd2, then I submit <code>abcd2</code>. If I don't find a PubChem match, I submit abcd1 because it has the highest probability. (is it ok?) </li>\n</ul>\n<p></p>",
      "rawMarkdown": "```Fixing predictions to nearest published InChIs is not allowed``` but \n```Fixing predictions to the exact published InChI (aka whitelist search)``` seems allowed. \n\ne.g. \n* output: abcd1 -> nearest PubChem entry is abcd2, then I submit `abcd2` (forbidden!)\n* output: abcd1 (p=0.5), abcd2 (p=0.3), abcd3 (p=0.2) -> I find that exact PubChem match is abcd2, then I submit `abcd2`. If I don't find a PubChem match, I submit abcd1 because it has the highest probability. (is it ok?) \n\n~~Anyway, I don't care. ~~",
      "votes": null
    },
    {
      "id": "1262991",
      "postDate": "04/05/2021 01:47:07",
      "content": "<p>maybe you could use an [InChI  SMILES  DeepSMILES] mapping to get around this problem</p>",
      "rawMarkdown": "maybe you could use an [InChI <=> SMILES <=> DeepSMILES] mapping to get around this problem",
      "votes": null
    },
    {
      "id": "1262992",
      "postDate": "04/05/2021 01:49:10",
      "content": "<p>InChI to SMILES is not too difficult, although you need to be careful about the several different versions of SMILES (canonical, blah blah, thats why InChI is nicer than SMILES for many things). OpenBabel can do this.</p>",
      "rawMarkdown": "InChI to SMILES is not too difficult, although you need to be careful about the several different versions of SMILES (canonical, blah blah, thats why InChI is nicer than SMILES for many things). OpenBabel can do this.",
      "votes": null
    },
    {
      "id": "1263027",
      "postDate": "04/05/2021 03:13:10",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> Maybe I just couldn't follow discussion, can you please specify which rule says it's forbidden? It will affect public leaderboard a lot if some forbidden but possible technique to achieve good score on lb…</p>",
      "rawMarkdown": "tugstugi Maybe I just couldn't follow discussion, can you please specify which rule says it's forbidden? It will affect public leaderboard a lot if some forbidden but possible technique to achieve good score on lb...",
      "votes": null
    },
    {
      "id": "1263387",
      "postDate": "04/05/2021 11:22:48",
      "content": "<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> </p>\n<blockquote>\n  <p>Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified. </p>\n</blockquote>\n<p>It seems like training LM on pubchem and do beam search i.e. forbidden.</p>",
      "rawMarkdown": "bamps53 \n> Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified. \n\nIt seems like training LM on pubchem and do beam search i.e. forbidden.",
      "votes": null
    },
    {
      "id": "1263445",
      "postDate": "04/05/2021 12:28:02",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - I'd suggest that there are at least three kinds of solutions.</p>\n<ol>\n<li><p>(And this is one for the image recognition, computer vision, AI, and humungously massive data experts, so I know ~ 0 about it) - train the whatever out of your model à la AlphaZero so it it just knows the InChI rules instinctively.</p></li>\n<li><p>Code up the full InChI algorithm in true \"Blue Obelisk school of cheminformatics\" style.</p></li>\n<li><p>Leverage the fact the someone esle has coded up the InChI algorithm. Train your model to recognise SMILES or something similar, then use an external program like OpenBabel to convert to InChI.</p></li>\n<li><p>The clever things I haven't thought of …</p></li>\n</ol>",
      "rawMarkdown": "hengck23 - I'd suggest that there are at least three kinds of solutions.\n\n1. (And this is one for the image recognition, computer vision, AI, and humungously massive data experts, so I know ~ 0 about it) - train the whatever out of your model à la AlphaZero so it it just knows the InChI rules instinctively.\n\n2. Code up the full InChI algorithm in true \"Blue Obelisk school of cheminformatics\" style.\n\n3. Leverage the fact the someone esle has coded up the InChI algorithm. Train your model to recognise SMILES or something similar, then use an external program like OpenBabel to convert to InChI.\n\n4. The clever things I haven't thought of ...",
      "votes": null
    },
    {
      "id": "1263461",
      "postDate": "04/05/2021 12:41:44",
      "content": "<p>although the atomic numbering is not preserved, it still follows some fixed rules.<br>\nwith enough train samples, the neural net can actually learn the rule.</p>\n<p>the problem would be easier if each part has the same \"substring\"</p>\n<p>i think this is the reason why a strong and large image encoder is important. we need to look at the global context to determine the numbering</p>",
      "rawMarkdown": "although the atomic numbering is not preserved, it still follows some fixed rules.\nwith enough train samples, the neural net can actually learn the rule.\n\nthe problem would be easier if each part has the same \"substring\"\n\ni think this is the reason why a strong and large image encoder is important. we need to look at the global context to determine the numbering",
      "votes": null
    },
    {
      "id": "1263705",
      "postDate": "04/05/2021 16:02:20",
      "content": "<p>validating with pubchem should be allowed. the [space of compounds in pubchem] &gt;&gt;&gt; [the space of compounds that people have actually synthesized]</p>\n<p>theoretically its a bias, but in reality its not</p>",
      "rawMarkdown": "validating with pubchem should be allowed. the [space of compounds in pubchem] >>> [the space of compounds that people have actually synthesized]\n\ntheoretically its a bias, but in reality its not",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1261043,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "04/02/2021 16:39:58",
      "content": "<p>Do you mean normalizing with rdkit? There are other ways to normalize which helps a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1261060,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/02/2021 16:58:42",
          "content": "<p>InCHI standard is unique for a given molecule.  My point is that we are not asked to predict valid InCHI standard string.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261113,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/02/2021 17:49:48",
      "content": "<p>one may one to check this interactive inchi generator<br>\n<a href=\"http://www.cheminfo.org/Chemistry/Cheminformatics/Generate_InChI/index.html\" target=\"_blank\">http://www.cheminfo.org/Chemistry/Cheminformatics/Generate_InChI/index.html</a></p>\n<p>i think we are going to have great problems for generalisation?</p>\n<p><img src=\"https://i.ibb.co/znLtcGB/Selection-030.png\" alt=\"\"></p>\n<p>i am looking for API to convert from table of atoms connections to INCHI … this make more sense</p>",
      "votes": null,
      "replies": [
        {
          "id": 1263445,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "04/05/2021 12:28:02",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - I'd suggest that there are at least three kinds of solutions.</p>\n<ol>\n<li><p>(And this is one for the image recognition, computer vision, AI, and humungously massive data experts, so I know ~ 0 about it) - train the whatever out of your model à la AlphaZero so it it just knows the InChI rules instinctively.</p></li>\n<li><p>Code up the full InChI algorithm in true \"Blue Obelisk school of cheminformatics\" style.</p></li>\n<li><p>Leverage the fact the someone esle has coded up the InChI algorithm. Train your model to recognise SMILES or something similar, then use an external program like OpenBabel to convert to InChI.</p></li>\n<li><p>The clever things I haven't thought of …</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263461,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/05/2021 12:41:44",
          "content": "<p>although the atomic numbering is not preserved, it still follows some fixed rules.<br>\nwith enough train samples, the neural net can actually learn the rule.</p>\n<p>the problem would be easier if each part has the same \"substring\"</p>\n<p>i think this is the reason why a strong and large image encoder is important. we need to look at the global context to determine the numbering</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261121,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "04/02/2021 17:57:15",
      "content": "<p>I was using pubchem to normalize/validate the predictions which brings a big boost but it is now forbidden: <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313</a><br>\nSeems also the 1st place team used a similar approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1261144,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/02/2021 18:15:58",
          "content": "<p>it is forbidden as a solution, but you can still use it for analysis of results and think of another accepted way that achieve the same purpose</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261200,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/02/2021 19:12:05",
          "content": "<p>Can you please pin point where its says that we can't normalize our <code>prediction</code> (InChi to InChI using <code>rdkit</code>) ?. The only think I see is following:  <code>Fixing predictions to nearest published InChIs is not allowed.</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261811,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/03/2021 13:27:16",
          "content": "<blockquote>\n  <p>Seems also the 1st place team used a similar approach.</p>\n</blockquote>\n<p>We haven't done this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261822,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "04/03/2021 13:38:48",
          "content": "<p>I'm not using any external data, too. I think it is possible to reach 1.xx.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261826,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/03/2021 13:41:42",
          "content": "<p>Using pubchem brings up to +3 boost.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261837,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/03/2021 13:52:22",
          "content": "<p>we are also not using any external data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261867,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 14:29:17",
          "content": "<p><a href=\"https://www.kaggle.com/DrHB\" target=\"_blank\">@DrHB</a> you keep listing what you don't use, when will you disclose what you actually use ?</p>\n<p>;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261871,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/03/2021 14:31:26",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> just had quick chat with my teammates… we agreed that we will disclose on June 3rd 2021 =)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261878,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 14:38:14",
          "content": "<p>I hope I'll have submitted by then…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261887,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 14:44:04",
          "content": "<p>I'll come back here once host settled on the external data and matching with known molecules.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1262941,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/04/2021 22:43:46",
          "content": "<p><code>Fixing predictions to nearest published InChIs is not allowed</code> but <br>\n<code>Fixing predictions to the exact published InChI (aka whitelist search)</code> seems allowed. </p>\n<p>e.g. </p>\n<ul>\n<li>output: abcd1 -&gt; nearest PubChem entry is abcd2, then I submit <code>abcd2</code> (forbidden!)</li>\n<li>output: abcd1 (p=0.5), abcd2 (p=0.3), abcd3 (p=0.2) -&gt; I find that exact PubChem match is abcd2, then I submit <code>abcd2</code>. If I don't find a PubChem match, I submit abcd1 because it has the highest probability. (is it ok?) </li>\n</ul>\n<p></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263027,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/05/2021 03:13:10",
          "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> Maybe I just couldn't follow discussion, can you please specify which rule says it's forbidden? It will affect public leaderboard a lot if some forbidden but possible technique to achieve good score on lb…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263387,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/05/2021 11:22:48",
          "content": "<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> </p>\n<blockquote>\n  <p>Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified. </p>\n</blockquote>\n<p>It seems like training LM on pubchem and do beam search i.e. forbidden.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263705,
          "author_name": "plbremer",
          "author_url": "",
          "post_date": "04/05/2021 16:02:20",
          "content": "<p>validating with pubchem should be allowed. the [space of compounds in pubchem] &gt;&gt;&gt; [the space of compounds that people have actually synthesized]</p>\n<p>theoretically its a bias, but in reality its not</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261201,
      "author_name": "infy2097",
      "author_url": "",
      "post_date": "04/02/2021 19:12:46",
      "content": "<p>This was also discussed here <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/224394</a></p>\n<p>With the top-scoring solutions now scoring so lowly the majority of results in these submissions need to be edit distance 0 and hence neccesarily have to be valid InChIs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1261697,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 11:03:07",
          "content": "<blockquote>\n  <p>This was also discussed here <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/224394</a></p>\n</blockquote>\n<p>I read that discussion, and it is mostly about the metric.  </p>\n<p>I think it was good to highlight what I discuss here separately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261739,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "04/03/2021 11:53:13",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  <a href=\"https://www.kaggle.com/infy2097\" target=\"_blank\">@infy2097</a> - can we agree that you are both right? Invalid InChIs are accepted by the scoring algorithm here and scored by edit distance (but rejected outright by every other piece of chemistry software I can find), but it is also true that to get to the top of this leaderboard you will need a lot of perfect InChIs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261745,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 11:58:10",
          "content": "<p>I edited my previous answer to make it clearer.  I don't see any disagreement,</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261239,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "04/02/2021 20:53:30",
      "content": "<p>Definitely the title speaks the truth. I understand (if it's not now forbidden knowledge) that the best scoring public kernels predict only small proportions of valid InChIs. I have given my opinions on the new rule interpretations <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229279\" target=\"_blank\">elsewhere</a>, but clearly predicting a valid InChI is in itself a significant challenge.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1261724,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 11:34:31",
          "content": "<p>Your discussion was very helpful to me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1262991,
      "author_name": "plbremer",
      "author_url": "",
      "post_date": "04/05/2021 01:47:07",
      "content": "<p>maybe you could use an [InChI  SMILES  DeepSMILES] mapping to get around this problem</p>",
      "votes": null,
      "replies": [
        {
          "id": 1262992,
          "author_name": "plbremer",
          "author_url": "",
          "post_date": "04/05/2021 01:49:10",
          "content": "<p>InChI to SMILES is not too difficult, although you need to be careful about the several different versions of SMILES (canonical, blah blah, thats why InChI is nicer than SMILES for many things). OpenBabel can do this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1260831": "Title says it all.\n\nThis maybe why it is possible to reach very low edit distance.  It can be better to output a wrong InCHI close to ground truth than a valid InCHI farther from it.",
    "1261043": "Do you mean normalizing with rdkit? There are other ways to normalize which helps a lot.",
    "1261060": "InCHI standard is unique for a given molecule.  My point is that we are not asked to predict valid InCHI standard string.",
    "1261113": "one may one to check this interactive inchi generator\nhttp://www.cheminfo.org/Chemistry/Cheminformatics/Generate_InChI/index.html\n\ni think we are going to have great problems for generalisation?\n\n![](https://i.ibb.co/znLtcGB/Selection-030.png)\n\ni am looking for API to convert from table of atoms connections to INCHI ... this make more sense",
    "1261121": "I was using pubchem to normalize/validate the predictions which brings a big boost but it is now forbidden: https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313\nSeems also the 1st place team used a similar approach.",
    "1261144": "it is forbidden as a solution, but you can still use it for analysis of results and think of another accepted way that achieve the same purpose",
    "1261200": "Can you please pin point where its says that we can't normalize our `prediction` (InChi to InChI using `rdkit`) ?. The only think I see is following:  `Fixing predictions to nearest published InChIs is not allowed. `",
    "1261201": "This was also discussed here https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\n\nWith the top-scoring solutions now scoring so lowly the majority of results in these submissions need to be edit distance 0 and hence neccesarily have to be valid InChIs.",
    "1261239": "Definitely the title speaks the truth. I understand (if it's not now forbidden knowledge) that the best scoring public kernels predict only small proportions of valid InChIs. I have given my opinions on the new rule interpretations [elsewhere](https://www.kaggle.com/c/bms-molecular-translation/discussion/229279), but clearly predicting a valid InChI is in itself a significant challenge.",
    "1261697": "> This was also discussed here https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\n\nI read that discussion, and it is mostly about the metric.  \n\nI think it was good to highlight what I discuss here separately.",
    "1261724": "Your discussion was very helpful to me.",
    "1261739": "cpmpml  @infy2097 - can we agree that you are both right? Invalid InChIs are accepted by the scoring algorithm here and scored by edit distance (but rejected outright by every other piece of chemistry software I can find), but it is also true that to get to the top of this leaderboard you will need a lot of perfect InChIs.",
    "1261745": "I edited my previous answer to make it clearer.  I don't see any disagreement,",
    "1261811": "> Seems also the 1st place team used a similar approach.\n\nWe haven't done this.",
    "1261822": "I'm not using any external data, too. I think it is possible to reach 1.xx.",
    "1261826": "Using pubchem brings up to +3 boost.",
    "1261837": "we are also not using any external data.",
    "1261867": "DrHB you keep listing what you don't use, when will you disclose what you actually use ?\n\n;)",
    "1261871": "cpmpml just had quick chat with my teammates... we agreed that we will disclose on June 3rd 2021 =)",
    "1261878": "I hope I'll have submitted by then...",
    "1261887": "I'll come back here once host settled on the external data and matching with known molecules.",
    "1262941": "```Fixing predictions to nearest published InChIs is not allowed``` but \n```Fixing predictions to the exact published InChI (aka whitelist search)``` seems allowed. \n\ne.g. \n* output: abcd1 -> nearest PubChem entry is abcd2, then I submit `abcd2` (forbidden!)\n* output: abcd1 (p=0.5), abcd2 (p=0.3), abcd3 (p=0.2) -> I find that exact PubChem match is abcd2, then I submit `abcd2`. If I don't find a PubChem match, I submit abcd1 because it has the highest probability. (is it ok?) \n\n~~Anyway, I don't care. ~~",
    "1262991": "maybe you could use an [InChI <=> SMILES <=> DeepSMILES] mapping to get around this problem",
    "1262992": "InChI to SMILES is not too difficult, although you need to be careful about the several different versions of SMILES (canonical, blah blah, thats why InChI is nicer than SMILES for many things). OpenBabel can do this.",
    "1263027": "tugstugi Maybe I just couldn't follow discussion, can you please specify which rule says it's forbidden? It will affect public leaderboard a lot if some forbidden but possible technique to achieve good score on lb...",
    "1263387": "bamps53 \n> Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified. \n\nIt seems like training LM on pubchem and do beam search i.e. forbidden.",
    "1263445": "hengck23 - I'd suggest that there are at least three kinds of solutions.\n\n1. (And this is one for the image recognition, computer vision, AI, and humungously massive data experts, so I know ~ 0 about it) - train the whatever out of your model à la AlphaZero so it it just knows the InChI rules instinctively.\n\n2. Code up the full InChI algorithm in true \"Blue Obelisk school of cheminformatics\" style.\n\n3. Leverage the fact the someone esle has coded up the InChI algorithm. Train your model to recognise SMILES or something similar, then use an external program like OpenBabel to convert to InChI.\n\n4. The clever things I haven't thought of ...",
    "1263461": "although the atomic numbering is not preserved, it still follows some fixed rules.\nwith enough train samples, the neural net can actually learn the rule.\n\nthe problem would be easier if each part has the same \"substring\"\n\ni think this is the reason why a strong and large image encoder is important. we need to look at the global context to determine the numbering",
    "1263705": "validating with pubchem should be allowed. the [space of compounds in pubchem] >>> [the space of compounds that people have actually synthesized]\n\ntheoretically its a bias, but in reality its not"
  },
  "source": "meta"
}