{
  "id": 229279,
  "title": "Question about copyright issues",
  "url": "/competitions/bms-molecular-translation/discussion/229279",
  "author_name": "",
  "post_date": "2021-03-29T12:32:00.241763200Z",
  "votes": 26,
  "comment_count": 46,
  "views": 0,
  "content": "<p>A Kaggle discussion topic was <strong>deleted because of copyright issues</strong>, as per today's update in the topic <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223381\" target=\"_blank\">Winning solutions from similar competition: Molecular Translation into SMILES</a>.</p>\n<p>Deleted post - <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223299\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223299</a><br>\nDeleted kernel - <a href=\"https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection\" target=\"_blank\">https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection</a></p>\n<p>In order to avoid other such potential issues, can the Organizers please reply about usage of below public information:</p>\n<ol>\n<li>Are we allowed to use external data from PubChem for data augmentation and validation?<br>\nThe winners of the Dacon competition used additional data from pubchem to make their dataset more balanced, and also for model generalization on new and unseen types of compounds.<br>\nA related discussion post with <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/225876\" target=\"_blank\">REST API</a> for \"clean\" images<br>\nPython library <a href=\"https://pubchempy.readthedocs.io/en/latest/guide/introduction.html\" target=\"_blank\">PubChemPy</a> (MIT License)</li>\n<li>Similar to PubChem, can we use other open access chemistry databases such as ChemSpider, ChEMBL, and others?</li>\n<li>Can we use python libraries <a href=\"https://www.rdkit.org/\" target=\"_blank\">RDKit</a>, <a href=\"https://www.inchi-trust.org/downloads/\" target=\"_blank\">Inchi</a>, <a href=\"http://openbabel.org/wiki/Main_Page\" target=\"_blank\">OpenBabel</a> for back and forth conversions (inchi <code>&lt;-&gt;</code> smiles <code>&lt;-&gt;</code> mol), validations, image generation and new data generation?</li>\n<li>Other data augmentation techniques such as GANs using above external data</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/Inversion\" target=\"_blank\">@Inversion</a> - could you please confirm that above items meet the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/rules\" target=\"_blank\">competition Rules</a> wrt external data and licensing requirements.</p>",
  "messages": [
    {
      "id": "1255971",
      "postDate": "03/29/2021 12:32:00",
      "content": "<p>A Kaggle discussion topic was <strong>deleted because of copyright issues</strong>, as per today's update in the topic <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223381\" target=\"_blank\">Winning solutions from similar competition: Molecular Translation into SMILES</a>.</p>\n<p>Deleted post - <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223299\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223299</a><br>\nDeleted kernel - <a href=\"https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection\" target=\"_blank\">https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection</a></p>\n<p>In order to avoid other such potential issues, can the Organizers please reply about usage of below public information:</p>\n<ol>\n<li>Are we allowed to use external data from PubChem for data augmentation and validation?<br>\nThe winners of the Dacon competition used additional data from pubchem to make their dataset more balanced, and also for model generalization on new and unseen types of compounds.<br>\nA related discussion post with <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/225876\" target=\"_blank\">REST API</a> for \"clean\" images<br>\nPython library <a href=\"https://pubchempy.readthedocs.io/en/latest/guide/introduction.html\" target=\"_blank\">PubChemPy</a> (MIT License)</li>\n<li>Similar to PubChem, can we use other open access chemistry databases such as ChemSpider, ChEMBL, and others?</li>\n<li>Can we use python libraries <a href=\"https://www.rdkit.org/\" target=\"_blank\">RDKit</a>, <a href=\"https://www.inchi-trust.org/downloads/\" target=\"_blank\">Inchi</a>, <a href=\"http://openbabel.org/wiki/Main_Page\" target=\"_blank\">OpenBabel</a> for back and forth conversions (inchi <code>&lt;-&gt;</code> smiles <code>&lt;-&gt;</code> mol), validations, image generation and new data generation?</li>\n<li>Other data augmentation techniques such as GANs using above external data</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/Inversion\" target=\"_blank\">@Inversion</a> - could you please confirm that above items meet the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/rules\" target=\"_blank\">competition Rules</a> wrt external data and licensing requirements.</p>",
      "rawMarkdown": "A Kaggle discussion topic was **deleted because of copyright issues**, as per today's update in the topic [Winning solutions from similar competition: Molecular Translation into SMILES](https://www.kaggle.com/c/bms-molecular-translation/discussion/223381).\n\nDeleted post - [https://www.kaggle.com/c/bms-molecular-translation/discussion/223299](https://www.kaggle.com/c/bms-molecular-translation/discussion/223299)\nDeleted kernel - [https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection](https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection)\n\nIn order to avoid other such potential issues, can the Organizers please reply about usage of below public information:\n1. Are we allowed to use external data from PubChem for data augmentation and validation?\n   The winners of the Dacon competition used additional data from pubchem to make their dataset more balanced, and also for model generalization on new and unseen types of compounds.\n   A related discussion post with [REST API](https://www.kaggle.com/c/bms-molecular-translation/discussion/225876) for \"clean\" images\n   Python library [PubChemPy](https://pubchempy.readthedocs.io/en/latest/guide/introduction.html) (MIT License)\n2. Similar to PubChem, can we use other open access chemistry databases such as ChemSpider, ChEMBL, and others?\n3. Can we use python libraries [RDKit](https://www.rdkit.org/), [Inchi](https://www.inchi-trust.org/downloads/), [OpenBabel](http://openbabel.org/wiki/Main_Page) for back and forth conversions (inchi `<->` smiles `<->` mol), validations, image generation and new data generation?\n4. Other data augmentation techniques such as GANs using above external data\n\n@Inversion - could you please confirm that above items meet the [competition Rules](https://www.kaggle.com/c/bms-molecular-translation/rules) wrt external data and licensing requirements.",
      "votes": null
    },
    {
      "id": "1256553",
      "postDate": "03/30/2021 02:28:02",
      "content": "<p>Quoted from the competition rules:</p>\n<blockquote>\n  <p>External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>\n</blockquote>",
      "rawMarkdown": "Quoted from the competition rules:\n> External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).",
      "votes": null
    },
    {
      "id": "1260313",
      "postDate": "04/02/2021 02:31:55",
      "content": "<p>Thanks for your question!  Using open cheminformatics libraries is encouraged for the competition.  Data augmentation using GANs would also be an interesting approach to explore.  PubChem and ChemSpider are good resources to explore the diversity of chemistry, though they contain many compounds that are very different from the competition data.  The training and test set were selected to focus on molecules with properties relevant to pharmaceuticals with enough examples to reduce the value of external data, .  Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified.  Good Luck!</p>\n<p>Edit: After further discussion we've updated the rules on the use of external data, <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318\" target=\"_blank\">see this post</a></p>",
      "rawMarkdown": "Thanks for your question!  Using open cheminformatics libraries is encouraged for the competition.  Data augmentation using GANs would also be an interesting approach to explore.  PubChem and ChemSpider are good resources to explore the diversity of chemistry, though they contain many compounds that are very different from the competition data.  The training and test set were selected to focus on molecules with properties relevant to pharmaceuticals with enough examples to reduce the value of external data, ~~but there is no prohibition~~.  Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified.  Good Luck!\n\nEdit: After further discussion we've updated the rules on the use of external data, [see this post](https://www.kaggle.com/c/bms-molecular-translation/discussion/231318)",
      "votes": null
    },
    {
      "id": "1260641",
      "postDate": "04/02/2021 09:35:14",
      "content": "<blockquote>\n  <p>Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified</p>\n</blockquote>\n<p>Can you elaborate on this? What is meant by \"biasing the predictions\" - surely just training a model with external data would bias the model's predictions?</p>\n<p>I can't find anything in the competition rules that reflects your comment so it would be good to know what is and isn't allowed. Thanks!</p>",
      "rawMarkdown": "> Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified\n\nCan you elaborate on this? What is meant by \"biasing the predictions\" - surely just training a model with external data would bias the model's predictions?\n\nI can't find anything in the competition rules that reflects your comment so it would be good to know what is and isn't allowed. Thanks!",
      "votes": null
    },
    {
      "id": "1260677",
      "postDate": "04/02/2021 10:37:36",
      "content": "<p><a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> thanks for indicating that you ar e using external data  ;)</p>",
      "rawMarkdown": "anokas thanks for indicating that you ar e using external data  ;)",
      "votes": null
    },
    {
      "id": "1260891",
      "postDate": "04/02/2021 13:58:41",
      "content": "<p><a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> Some test molecules may be present in PubChem data. Most of the molecules used in this competition will be likely found in many public sources / publications.</p>\n<p>So, it is <strong>not allowed</strong> to use PubChem data for:</p>\n<ol>\n<li>Post-processing predictions to validate and \"fix\" the InChI output strings</li>\n<li>Training with data containing test set labels</li>\n<li>Finetuning on data containing test set labels</li>\n</ol>\n<p>Previous Kaggle <strong>disqualification</strong> references:</p>\n<ul>\n<li>1st prize winner of PetFinder contest was <strong>disqualified</strong> for <a href=\"https://www.kaggle.com/c/petfinder-adoption-prediction/discussion/125436\" target=\"_blank\">scraping a website to get test labels</a></li>\n<li>Scraping of test labels in <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80665\" target=\"_blank\">Quora Insincere Questions Classification</a></li>\n<li>Scraping in <a href=\"https://www.kaggle.com/c/ashrae-energy-prediction/discussion/116840\" target=\"_blank\">ASHRAE - Great Energy Predictor III</a></li>\n</ul>\n<p>So the Organizers (Jacob Albrecht) decided to disqualify such solutions because they do not generalize to new and unseen data, and so the <strong>solution is not very useful anyway</strong>.</p>\n<p>However, I believe that it is <strong>allowed</strong> to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.</p>\n<hr>\n<p>[Update] <a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> Jacob thank you for <strong>upvoting this comment</strong>. Could you please also reply to <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> below so that the competition rules and expectations are clear to everyone. Also please see my new comment about PubChem database below for more context.</p>",
      "rawMarkdown": "anokas Some test molecules may be present in PubChem data. Most of the molecules used in this competition will be likely found in many public sources / publications.\n\nSo, it is **not allowed** to use PubChem data for:\n1. Post-processing predictions to validate and \"fix\" the InChI output strings\n2. Training with data containing test set labels\n3. Finetuning on data containing test set labels\n\nPrevious Kaggle **disqualification** references:\n- 1st prize winner of PetFinder contest was **disqualified** for [scraping a website to get test labels](https://www.kaggle.com/c/petfinder-adoption-prediction/discussion/125436)\n- Scraping of test labels in [Quora Insincere Questions Classification](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80665)\n- Scraping in [ASHRAE - Great Energy Predictor III](https://www.kaggle.com/c/ashrae-energy-prediction/discussion/116840)\n\nSo the Organizers (Jacob Albrecht) decided to disqualify such solutions because they do not generalize to new and unseen data, and so the **solution is not very useful anyway**.\n\nHowever, I believe that it is **allowed** to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.\n__________________________________________________________________________________________\n[Update] @jakealbrecht1337 Jacob thank you for **upvoting this comment**. Could you please also reply to @anokas below so that the competition rules and expectations are clear to everyone. Also please see my new comment about PubChem database below for more context.",
      "votes": null
    },
    {
      "id": "1260904",
      "postDate": "04/02/2021 14:17:06",
      "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> Thank you, but I assume you are not speaking officially on behalf of the organisers. It would be good to get an official explanation (and addition to the ruleset) of what exactly is and isn't allowed.</p>\n<ol>\n<li><p>You say that \"it is allowed to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.\" As competitors do not have the test set labels, it would be impossible for a competitor to check whether any external data has 'unintentional' overlap. If the organisers are aware of any datasets that have overlap (for example, the datasets from which they generated their data), then they should make clear that these datasets are off-limits.</p></li>\n<li><p>Your examples of previous disqualification are a bit shaky I feel. Only your first link had someone disqualified, and in that case it was a very blatant attempt at concealed cheating. In the third link, the organisers <a href=\"https://www.kaggle.com/c/ashrae-energy-prediction/discussion/117357\" target=\"_blank\">specifically allowed</a> using this scraped data for the competition! (and it was also allowed in the second)</p></li>\n</ol>\n<p>I agree that this \"extreme\" scraping shouldn't be allowed (and I haven't made any attempt to scrape the test set labels), but what we need is an objective reasoning from the organisers on what they consider to be off-limits.</p>",
      "rawMarkdown": "sirishks Thank you, but I assume you are not speaking officially on behalf of the organisers. It would be good to get an official explanation (and addition to the ruleset) of what exactly is and isn't allowed.\n\n1. You say that \"it is allowed to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.\" As competitors do not have the test set labels, it would be impossible for a competitor to check whether any external data has 'unintentional' overlap. If the organisers are aware of any datasets that have overlap (for example, the datasets from which they generated their data), then they should make clear that these datasets are off-limits.\n\n2. Your examples of previous disqualification are a bit shaky I feel. Only your first link had someone disqualified, and in that case it was a very blatant attempt at concealed cheating. In the third link, the organisers [specifically allowed](https://www.kaggle.com/c/ashrae-energy-prediction/discussion/117357) using this scraped data for the competition! (and it was also allowed in the second)\n\nI agree that this \"extreme\" scraping shouldn't be allowed (and I haven't made any attempt to scrape the test set labels), but what we need is an objective reasoning from the organisers on what they consider to be off-limits.",
      "votes": null
    },
    {
      "id": "1261011",
      "postDate": "04/02/2021 15:57:41",
      "content": "<p>\"So, it is not allowed to use PubChem data for: Post-processing predictions to validate and \"fix\" the InChI output strings\"</p>\n<p>this is called white list filtering. In commercial car license plate applications, we are sometimes asked to identify the license plate registered cars coming in and out of a car park. Because cars are registered in the whitelist, we know the prediction must be one of those registered.</p>\n<p>This simple trick helps us to correct errors if we cannot get all numbers correct.</p>\n<p>verification and post-processing of InChI output strings is a smart and valid trick I think</p>",
      "rawMarkdown": "\"So, it is not allowed to use PubChem data for: Post-processing predictions to validate and \"fix\" the InChI output strings\"\n\n\nthis is called white list filtering. In commercial car license plate applications, we are sometimes asked to identify the license plate registered cars coming in and out of a car park. Because cars are registered in the whitelist, we know the prediction must be one of those registered.\n\nThis simple trick helps us to correct errors if we cannot get all numbers correct.\n\n\nverification and post-processing of InChI output strings is a smart and valid trick I think",
      "votes": null
    },
    {
      "id": "1261013",
      "postDate": "04/02/2021 16:04:13",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>, I have just checked and found some of the test molecules (<strong>private labels</strong>) to be present on PubChem.</p>\n<p>However PubChem is the largest (and still growing) public chemical database with <strong>103 Million</strong> molecules from 748 reputed sources, and it is not very easy to identify the competition test set (1.6M test molecules). Moreover, 103M is such a large number that it is reasonable to assume that most molecules are present in this dataset already i.e. a solution using this should be a \"general\" solution and should not be considered as scraping! Also, this approach \"may not compromise the utility of the model outside of the test set\" because other molecules encountered in publications in the future would <strong>also be added to PubChem anyway</strong>.</p>\n<p>Could you please make a decision to <strong>fully allow</strong> PubChem or <strong>disallow</strong> PubChem (and probably disqualify all other external data sources as well).</p>",
      "rawMarkdown": "jakealbrecht1337 @addisonhoward, I have just checked and found some of the test molecules (**private labels**) to be present on PubChem.\n\nHowever PubChem is the largest (and still growing) public chemical database with **103 Million** molecules from 748 reputed sources, and it is not very easy to identify the competition test set (1.6M test molecules). Moreover, 103M is such a large number that it is reasonable to assume that most molecules are present in this dataset already i.e. a solution using this should be a \"general\" solution and should not be considered as scraping! Also, this approach \"may not compromise the utility of the model outside of the test set\" because other molecules encountered in publications in the future would **also be added to PubChem anyway**.\n\nCould you please make a decision to **fully allow** PubChem or **disallow** PubChem (and probably disqualify all other external data sources as well).",
      "votes": null
    },
    {
      "id": "1261017",
      "postDate": "04/02/2021 16:09:54",
      "content": "<p>For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.  .  Fixing predictions to nearest published InChIs is not allowed.  A simple check that a predicted label was not included in the training data is good due diligence.</p>\n<p>Edit: We've updated the rules on the use of external data <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318\" target=\"_blank\">see this post</a></p>",
      "rawMarkdown": "For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.  ~~If you use external data in a prize winning submission, your model documentation should include the steps you took to ensure test labels were not included in training~~.  Fixing predictions to nearest published InChIs is not allowed.  A simple check that a predicted label was not included in the training data is good due diligence.\n\nEdit: We've updated the rules on the use of external data [see this post](https://www.kaggle.com/c/bms-molecular-translation/discussion/231318)",
      "votes": null
    },
    {
      "id": "1261023",
      "postDate": "04/02/2021 16:16:56",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> You replied that \"fixing predictions\" is not allowed, but what if someone \"unintentionally\" included valid InChI strings in the training set itself (and fixes predictions that way)?</p>\n<p>As pointed out by <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a>, it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.</p>\n<p>So, please see my comment below and decide whether to <strong>disallow all external data sources</strong> including PubChem, because any other approach will be ambiguous.</p>",
      "rawMarkdown": "jakealbrecht1337 You replied that \"fixing predictions\" is not allowed, but what if someone \"unintentionally\" included valid InChI strings in the training set itself (and fixes predictions that way)?\n\nAs pointed out by @anokas, it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.\n\nSo, please see my comment below and decide whether to **disallow all external data sources** including PubChem, because any other approach will be ambiguous.",
      "votes": null
    },
    {
      "id": "1261040",
      "postDate": "04/02/2021 16:36:57",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> what about unsupervised pretraining on test images or pseudolabeling them using model predictions - are they allowed?</p>",
      "rawMarkdown": "jakealbrecht1337 what about unsupervised pretraining on test images or pseudolabeling them using model predictions - are they allowed?",
      "votes": null
    },
    {
      "id": "1261041",
      "postDate": "04/02/2021 16:38:32",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> why limit this to prize winners?  It would be simpler to say that all entries must comply with what you wrote  (edited to remove reference to prize winners):</p>\n<blockquote>\n  <p>If you use external data in a prize winning submission, you should ensure that test labels were not included in training. Fixing predictions to nearest published InChIs is not allowed. A simple check that a predicted label was not included in the training data is good due diligence.</p>\n</blockquote>\n<p>Granted, you won't check this for each and every submissions, but same is true of any restriction set forth in competition rules.</p>",
      "rawMarkdown": "jakealbrecht1337 why limit this to prize winners?  It would be simpler to say that all entries must comply with what you wrote  (edited to remove reference to prize winners):\n\n>  If you use external data in a prize winning submission, you should ensure that test labels were not included in training. Fixing predictions to nearest published InChIs is not allowed. A simple check that a predicted label was not included in the training data is good due diligence.\n\nGranted, you won't check this for each and every submissions, but same is true of any restriction set forth in competition rules.",
      "votes": null
    },
    {
      "id": "1261045",
      "postDate": "04/02/2021 16:40:34",
      "content": "<blockquote>\n  <p>it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.</p>\n</blockquote>\n<p>Then don't include external data if you aren't sure.</p>",
      "rawMarkdown": ">  it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.\n\nThen don't include external data if you aren't sure.",
      "votes": null
    },
    {
      "id": "1261051",
      "postDate": "04/02/2021 16:51:16",
      "content": "<blockquote>\n  <p>For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.</p>\n</blockquote>\n<p>That's not what the rules says. (§12) says that you have the right to verify compliance with <strong>\"these rules\"</strong> (i.e. the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/rules\" target=\"_blank\">rules page</a>). I think any additional rules described in the forums should be formalised in the rules page for those not reading every post in the forums. One would expect the rules page to contain the full set of competition rules.</p>\n<p>As <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> asks, what about unsupervised learning/pseudo-learning using the test set (and without external data)? </p>\n<p>I think we need a better definition and explanation than \"we will decide what is allowed <em>after</em> the competition\". Competitors need to be saved working thousands of hours only to be told that the organisers didn't like their solution so they're being disqualified</p>",
      "rawMarkdown": "> For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.\n\nThat's not what the rules says. (§12) says that you have the right to verify compliance with **\"these rules\"** (i.e. the [rules page](https://www.kaggle.com/c/bms-molecular-translation/rules)). I think any additional rules described in the forums should be formalised in the rules page for those not reading every post in the forums. One would expect the rules page to contain the full set of competition rules.\n\nAs @stassl asks, what about unsupervised learning/pseudo-learning using the test set (and without external data)? \n\nI think we need a better definition and explanation than \"we will decide what is allowed _after_ the competition\". Competitors need to be saved working thousands of hours only to be told that the organisers didn't like their solution so they're being disqualified",
      "votes": null
    },
    {
      "id": "1261054",
      "postDate": "04/02/2021 16:53:09",
      "content": "<p>I agree. Having vague and post-hoc-decided rules has caused issues in past competitions. If it is not possible to have an objective measure of what external data is allowed, then it should be completely disallowed.</p>",
      "rawMarkdown": "I agree. Having vague and post-hoc-decided rules has caused issues in past competitions. If it is not possible to have an objective measure of what external data is allowed, then it should be completely disallowed.",
      "votes": null
    },
    {
      "id": "1261127",
      "postDate": "04/02/2021 18:01:29",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> A common approach in OCR/STT is to train a language model and do beam search. So it means training a LM on PubChem i.e. not allowed?</p>",
      "rawMarkdown": "jakealbrecht1337 A common approach in OCR/STT is to train a language model and do beam search. So it means training a LM on PubChem i.e. not allowed?",
      "votes": null
    },
    {
      "id": "1261171",
      "postDate": "04/02/2021 18:52:54",
      "content": "<p>a better solution for future competition is to allow the use of PubChem (including mapping to valid InChI), but a large percentage of the test data (e.g. 50%) has ground truth InChI not found in any public dataset like PubChem at all.</p>\n<p>in this case you really test the generalization of the algorithm and you do not restrict the creativity of kagglers to use external data.</p>",
      "rawMarkdown": "a better solution for future competition is to allow the use of PubChem (including mapping to valid InChI), but a large percentage of the test data (e.g. 50%) has ground truth InChI not found in any public dataset like PubChem at all.\n\nin this case you really test the generalization of the algorithm and you do not restrict the creativity of kagglers to use external data.",
      "votes": null
    },
    {
      "id": "1261224",
      "postDate": "04/02/2021 20:03:48",
      "content": "<p>At this point, I think the best solution is to disallow all external data. </p>",
      "rawMarkdown": "At this point, I think the best solution is to disallow all external data.",
      "votes": null
    },
    {
      "id": "1261227",
      "postDate": "04/02/2021 20:15:11",
      "content": "<p>From the train set only 326 moleculer formulas (haven't checked isomers) are not present in the PubChem. Test set should be similar.</p>",
      "rawMarkdown": "From the train set only 326 moleculer formulas (haven't checked isomers) are not present in the PubChem. Test set should be similar.",
      "votes": null
    },
    {
      "id": "1261238",
      "postDate": "04/02/2021 20:47:37",
      "content": "<p>First of all I will declare the vested interest that I have as a computational chemist. </p>\n<p>My view of Kaggle is that good competitions should and do allow people with quite distinct expertise to collaborate and compete on problems that defy the usual boundaries. Molecular Translation could and should be one such competition. Rather than rant or downvote comments I disagree with, I'll try to explain carefully why I fear that these interpretations of the rules are discriminating unreasonably against people with chemistry and chem(o)informatics skillsets, and unduly in favour of specialists in heavy duty machine learning and image recognition, thus reducing the scope and appeal of this contest.</p>\n<p>(1) On checking whether InChIs, either one-at-a-time or as a cohort of predictions, represent real molecules by cross-validating against external databases. This seems to me to be a fair and sensible way of assessing the proportion of fully correct InChIs coming out of a predictor. It doesn't tell us that the InChI represents the test molecule, just that it represents some molecule. It is manifestly not looking at the labels of the test data (images whose true labels don't exist anywhere on the internet), any more than feedback via a LB score is. The size of the dataset already precludes use of low-throughput methods such as browsing individual database entries to make a meaningful advance in LB score. Some of the test images also probably represent virtual compounds that won't be in PubChem or ChemSpider, so some correct InChIs won't be recognised by databases - and that is part of the challenge. </p>\n<p>(2) InChIs are extremely fiddly things to get right, and software &amp; databases seem to reject non-canonical InChIs outright even if they can handle non-canonical SMILES. So getting the InChI string right seems to be as big a challenge as identifying the compound (say by its graph or SMILES). The prohibition makes it harder to leverage chemoinformatics skills to check InChIs. I'd assume that this ruling must favour the image-to-SMILES-to-InChI route (the last step via external software) over direct image-to-InChI prediction. By making InChI checking harder, isn't it more likely that the winning solutions will be SMILES predictors?</p>\n<p>(3) Every Kaggle competition that I've been involved in has encouraged ensembling. An obvious ensembling route given predictions A &amp; B for molecule M is to choose whichever of A(M) or B(M) is in PubChem, whenever one and only one is present. If both or neither are present, stick with what you think is the better predictor.</p>\n<p>(4) Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction, better than predicting one 'average' InChI for everything. The leaderboard shows that the top teams are already doing better than this, but if someone wants to get a silver or bronze by the method I describe, why stop them?</p>\n<p>(5) There are alternative methods of checking InChI validity via external software, which I assume are still allowed. I expect that rejection of an InChI by chemoinformatics software usually means that the InChI fails to represent any feasible molecular graph, but I can't be absolutely sure that it doesn't in fact mean that a database lookup failed to generate a hit. So it seems hard to know what InChI checking procedures are compliant with the suggested interpretation and which aren't.</p>",
      "rawMarkdown": "First of all I will declare the vested interest that I have as a computational chemist. \n\nMy view of Kaggle is that good competitions should and do allow people with quite distinct expertise to collaborate and compete on problems that defy the usual boundaries. Molecular Translation could and should be one such competition. Rather than rant or downvote comments I disagree with, I'll try to explain carefully why I fear that these interpretations of the rules are discriminating unreasonably against people with chemistry and chem(o)informatics skillsets, and unduly in favour of specialists in heavy duty machine learning and image recognition, thus reducing the scope and appeal of this contest.\n\n(1) On checking whether InChIs, either one-at-a-time or as a cohort of predictions, represent real molecules by cross-validating against external databases. This seems to me to be a fair and sensible way of assessing the proportion of fully correct InChIs coming out of a predictor. It doesn't tell us that the InChI represents the test molecule, just that it represents some molecule. It is manifestly not looking at the labels of the test data (images whose true labels don't exist anywhere on the internet), any more than feedback via a LB score is. The size of the dataset already precludes use of low-throughput methods such as browsing individual database entries to make a meaningful advance in LB score. Some of the test images also probably represent virtual compounds that won't be in PubChem or ChemSpider, so some correct InChIs won't be recognised by databases - and that is part of the challenge. \n\n(2) InChIs are extremely fiddly things to get right, and software & databases seem to reject non-canonical InChIs outright even if they can handle non-canonical SMILES. So getting the InChI string right seems to be as big a challenge as identifying the compound (say by its graph or SMILES). The prohibition makes it harder to leverage chemoinformatics skills to check InChIs. I'd assume that this ruling must favour the image-to-SMILES-to-InChI route (the last step via external software) over direct image-to-InChI prediction. By making InChI checking harder, isn't it more likely that the winning solutions will be SMILES predictors?\n\n(3) Every Kaggle competition that I've been involved in has encouraged ensembling. An obvious ensembling route given predictions A & B for molecule M is to choose whichever of A(M) or B(M) is in PubChem, whenever one and only one is present. If both or neither are present, stick with what you think is the better predictor.\n\n(4) Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction, better than predicting one 'average' InChI for everything. The leaderboard shows that the top teams are already doing better than this, but if someone wants to get a silver or bronze by the method I describe, why stop them?\n\n(5) There are alternative methods of checking InChI validity via external software, which I assume are still allowed. I expect that rejection of an InChI by chemoinformatics software usually means that the InChI fails to represent any feasible molecular graph, but I can't be absolutely sure that it doesn't in fact mean that a database lookup failed to generate a hit. So it seems hard to know what InChI checking procedures are compliant with the suggested interpretation and which aren't.",
      "votes": null
    },
    {
      "id": "1261302",
      "postDate": "04/02/2021 22:45:46",
      "content": "<p>Hey all,</p>\n<p>We're working on a clarification but it will take a few days (and we recommend holding off on external data in the meantime, for those of you who want to be on the safe side).</p>\n<p>Stay tuned</p>",
      "rawMarkdown": "Hey all,\n\nWe're working on a clarification but it will take a few days (and we recommend holding off on external data in the meantime, for those of you who want to be on the safe side).\n\nStay tuned",
      "votes": null
    },
    {
      "id": "1261330",
      "postDate": "04/03/2021 00:37:49",
      "content": "<p><a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">@jbomitchell</a> I really like your points, and they reflect the same thoughts I have!</p>\n<p>Since the test set doesn't appear to have been generated from any public dataset, I can't see how a public dataset would unreasonably bias a solution towards <em>this</em> test set (as opposed to any other potential test set from the same distribution)</p>",
      "rawMarkdown": "jbomitchell I really like your points, and they reflect the same thoughts I have!\n\nSince the test set doesn't appear to have been generated from any public dataset, I can't see how a public dataset would unreasonably bias a solution towards _this_ test set (as opposed to any other potential test set from the same distribution)",
      "votes": null
    },
    {
      "id": "1261331",
      "postDate": "04/03/2021 00:38:11",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1261334",
      "postDate": "04/03/2021 00:58:39",
      "content": "<p>I suspect the reason for the organizer's concern is that in the real-world usage scenario the tool would be applied to internal documents and other documents that haven't had their chemistry indexed into Pubchem, so an assumption that the majority/all compounds will be present in PubChem will breakdown. Having said that I could imagine that even if none of the compounds were in PubChem, that a large dataset like PubChem would be useful in establishing what sort of motifs are common in real-world chemicals!</p>\n<p>Ideally the training/test set would contain mostly virtual compounds, that are plausible pharmaceutical compounds. Having some of them occur in PubChem would be fine, you don't want PubChem being used as a stop list either! Anyway the data set is what it is…</p>",
      "rawMarkdown": "I suspect the reason for the organizer's concern is that in the real-world usage scenario the tool would be applied to internal documents and other documents that haven't had their chemistry indexed into Pubchem, so an assumption that the majority/all compounds will be present in PubChem will breakdown. Having said that I could imagine that even if none of the compounds were in PubChem, that a large dataset like PubChem would be useful in establishing what sort of motifs are common in real-world chemicals!\n\nIdeally the training/test set would contain mostly virtual compounds, that are plausible pharmaceutical compounds. Having some of them occur in PubChem would be fine, you don't want PubChem being used as a stop list either! Anyway the data set is what it is...",
      "votes": null
    },
    {
      "id": "1261391",
      "postDate": "04/03/2021 03:24:43",
      "content": "<blockquote>\n  <p>to ensure test labels were not included in training…</p>\n</blockquote>\n<p>By the wording of this, pseudo labels shouldn't be allowed, since any decent model should get a great portion of test labels perfectly correct. But I don't see how you can enforce this on participants. </p>",
      "rawMarkdown": "> to ensure test labels were not included in training...\n\nBy the wording of this, pseudo labels shouldn't be allowed, since any decent model should get a great portion of test labels perfectly correct. But I don't see how you can enforce this on participants.",
      "votes": null
    },
    {
      "id": "1261497",
      "postDate": "04/03/2021 05:45:54",
      "content": "<p>Everyone knows using more training data and cross checking with external databases would be beneficial to improve the accuracy. I am sure the organizer knows it too. Try ask yourself why do they want to host a competition and pay others to do something they already know ? Why not they do it by themselves then ? If you are the organizer, would you prefer to see a competition of different new model ideas or a competition of getting a larger external training data set or a competition of tweaking predictions with a larger ensemble of external inchi databases ? It depends on what the organizer is looking for. </p>",
      "rawMarkdown": "Everyone knows using more training data and cross checking with external databases would be beneficial to improve the accuracy. I am sure the organizer knows it too. Try ask yourself why do they want to host a competition and pay others to do something they already know ? Why not they do it by themselves then ? If you are the organizer, would you prefer to see a competition of different new model ideas or a competition of getting a larger external training data set or a competition of tweaking predictions with a larger ensemble of external inchi databases ? It depends on what the organizer is looking for.",
      "votes": null
    },
    {
      "id": "1261650",
      "postDate": "04/03/2021 10:04:25",
      "content": "<p>And that's one of the reasons why I advocate for kernel competitions… no way to see private data upfront then.</p>",
      "rawMarkdown": "And that's one of the reasons why I advocate for kernel competitions... no way to see private data upfront then.",
      "votes": null
    },
    {
      "id": "1261653",
      "postDate": "04/03/2021 10:07:20",
      "content": "<p>My personal opinion is that it would be better to disallow all external data in most competition.</p>",
      "rawMarkdown": "My personal opinion is that it would be better to disallow all external data in most competition.",
      "votes": null
    },
    {
      "id": "1261681",
      "postDate": "04/03/2021 10:49:00",
      "content": "<p>i would opt for:</p>\n<ol>\n<li>public LB : csv submission and one can download data</li>\n<li>private LB : code submission and unseen data</li>\n</ol>",
      "rawMarkdown": "i would opt for:\n\n1. public LB : csv submission and one can download data\n2. private LB : code submission and unseen data",
      "votes": null
    },
    {
      "id": "1261705",
      "postDate": "04/03/2021 11:09:11",
      "content": "<p><a href=\"https://www.kaggle.com/Psi\" target=\"_blank\">@Psi</a> this is tempting indeed.  Issue with that is to still allow pretrained models.  And these models are trained on external data…</p>",
      "rawMarkdown": "Psi this is tempting indeed.  Issue with that is to still allow pretrained models.  And these models are trained on external data...",
      "votes": null
    },
    {
      "id": "1261708",
      "postDate": "04/03/2021 11:11:52",
      "content": "<blockquote>\n  <p>Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction</p>\n</blockquote>\n<p>This is useless for non published molecules.  It looks like host wants to improve retrieval from known molecules they have access to, and that may be different from pubchem.</p>",
      "rawMarkdown": "> Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction\n\nThis is useless for non published molecules.  It looks like host wants to improve retrieval from known molecules they have access to, and that may be different from pubchem.",
      "votes": null
    },
    {
      "id": "1261736",
      "postDate": "04/03/2021 11:47:36",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I agree, it is not an easy setup … Kaggle could approve certain pretrained repositories upfront (timm, tensorhub, huggingface, etc.) … but yeah there probably is no perfect setup unfortunately.</p>\n<p>I am just concerned about the general importance of external data in many competitions and I would rather prefer to focus on modeling than on finding the right data to use. And unfortunately there are often also leaks (not even on purpose) involved.</p>",
      "rawMarkdown": "cpmpml I agree, it is not an easy setup ... Kaggle could approve certain pretrained repositories upfront (timm, tensorhub, huggingface, etc.) ... but yeah there probably is no perfect setup unfortunately.\n\nI am just concerned about the general importance of external data in many competitions and I would rather prefer to focus on modeling than on finding the right data to use. And unfortunately there are often also leaks (not even on purpose) involved.",
      "votes": null
    },
    {
      "id": "1261741",
      "postDate": "04/03/2021 11:55:14",
      "content": "<p>I agree with you.  Very often what differentiate winner from other top finisher is the use of more external data.</p>",
      "rawMarkdown": "I agree with you.  Very often what differentiate winner from other top finisher is the use of more external data.",
      "votes": null
    },
    {
      "id": "1261747",
      "postDate": "04/03/2021 12:01:33",
      "content": "<p><a href=\"https://www.kaggle.com/infy2097\" target=\"_blank\">@infy2097</a> - a good point. I think there are at least two distinct real worlds here.</p>\n<p>One is the world of the Blue Obelisk gang, where the objective is to open up and add value to the chemical literature. They would be happy to see cooperation and interoperability between different databases, provided that everything is Open Access.</p>\n<p>The other is the world of Pharma, where the escape of your proprietary compound's InChI into the open internet might cost a billion dollars. Less keen on the use of external data without strong security!</p>",
      "rawMarkdown": "infy2097 - a good point. I think there are at least two distinct real worlds here.\n\nOne is the world of the Blue Obelisk gang, where the objective is to open up and add value to the chemical literature. They would be happy to see cooperation and interoperability between different databases, provided that everything is Open Access.\n\nThe other is the world of Pharma, where the escape of your proprietary compound's InChI into the open internet might cost a billion dollars. Less keen on the use of external data without strong security!",
      "votes": null
    },
    {
      "id": "1261849",
      "postDate": "04/03/2021 14:05:14",
      "content": "<p>I completely agree, Kaggle should disallow usage of any external data in all competitions, so everyone could focus in building better solutions for the competition dataset and not spend time searching for external data and leakages. It would save a lot of time for Kaggle and competitors and make rules easier to understand without any potentital obfuscated disqualification issue.</p>",
      "rawMarkdown": "I completely agree, Kaggle should disallow usage of any external data in all competitions, so everyone could focus in building better solutions for the competition dataset and not spend time searching for external data and leakages. It would save a lot of time for Kaggle and competitors and make rules easier to understand without any potentital obfuscated disqualification issue.",
      "votes": null
    },
    {
      "id": "1261855",
      "postDate": "04/03/2021 14:14:31",
      "content": "<p>I agree that external data should not be allowed, but the line can become blurry with pre-trained models (which I think are certainly a net positive for competitions). Where is the line between \" software library\" and \"external data\"?</p>",
      "rawMarkdown": "I agree that external data should not be allowed, but the line can become blurry with pre-trained models (which I think are certainly a net positive for competitions). Where is the line between \" software library\" and \"external data\"?",
      "votes": null
    },
    {
      "id": "1261873",
      "postDate": "04/03/2021 14:32:40",
      "content": "<p>A good way of handling external data was the last Lyft competition.  External data had to be explicitly approved by host.</p>",
      "rawMarkdown": "A good way of handling external data was the last Lyft competition.  External data had to be explicitly approved by host.",
      "votes": null
    },
    {
      "id": "1261947",
      "postDate": "04/03/2021 15:52:48",
      "content": "<p>I agree <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> - this was perfectly handled by <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> .</p>\n<p>I think it wouldn't be too hard to whitelist a few pretrained models and then for others make an approval process. It's not like competitors post 1000 different external types of pre-trained models, usually those are maybe 10-20.</p>",
      "rawMarkdown": "I agree @cpmpml - this was perfectly handled by @iglovikov .\n\nI think it wouldn't be too hard to whitelist a few pretrained models and then for others make an approval process. It's not like competitors post 1000 different external types of pre-trained models, usually those are maybe 10-20.",
      "votes": null
    },
    {
      "id": "1262953",
      "postDate": "04/04/2021 23:29:36",
      "content": "<p>I am guessing that more clarification is coming but I find this statement to be a questionable decision by the host.</p>\n<blockquote>\n  <p>Fixing predictions to nearest published InChIs is not allowed&gt; </p>\n</blockquote>\n<p>If I was handed this finished product to use, the very FIRST question I would have as a user - is the predicted molecule real/known.  If something like RDKIT returns None or crashes with the predicted InchI as a user I would want to KNOW.  ( I don't know enough about RDKIT to assume that a crash or a returned None is not a real molecule, but odds are not good.)</p>\n<p>I don't understand which would be the greatest issue - a model with a super low score where many of the molecules are not real or a model where the prediction has been post fixed to the nearest low scoring \"real\" InChi.  </p>\n<p>From my reading it seems that \"searching\" is done using InChi_key.  Can you generate a key with a non real molecule?</p>",
      "rawMarkdown": "I am guessing that more clarification is coming but I find this statement to be a questionable decision by the host.\n\n> Fixing predictions to nearest published InChIs is not allowed> \n\nIf I was handed this finished product to use, the very FIRST question I would have as a user - is the predicted molecule real/known.  If something like RDKIT returns None or crashes with the predicted InchI as a user I would want to KNOW.  ( I don't know enough about RDKIT to assume that a crash or a returned None is not a real molecule, but odds are not good.)\n\nI don't understand which would be the greatest issue - a model with a super low score where many of the molecules are not real or a model where the prediction has been post fixed to the nearest low scoring \"real\" InChi.  \n\nFrom my reading it seems that \"searching\" is done using InChi_key.  Can you generate a key with a non real molecule?",
      "votes": null
    },
    {
      "id": "1263395",
      "postDate": "04/05/2021 11:29:28",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> : Can you tell us if the usage of RDKit is allowed to check predictions ?</p>",
      "rawMarkdown": "jakealbrecht1337 : Can you tell us if the usage of RDKit is allowed to check predictions ?",
      "votes": null
    },
    {
      "id": "1263633",
      "postDate": "04/05/2021 14:50:47",
      "content": "<p>I completely disagree with the idea of banning external data and information. Leveraging suitable external sources brings a lot of richness and interest to Kaggle, opening up more avenues for creative problem solving. Without this you'd be left with just a (IMHO boring) coding competition, plus some advantage for those with access to local GPUs. </p>",
      "rawMarkdown": "I completely disagree with the idea of banning external data and information. Leveraging suitable external sources brings a lot of richness and interest to Kaggle, opening up more avenues for creative problem solving. Without this you'd be left with just a (IMHO boring) coding competition, plus some advantage for those with access to local GPUs.",
      "votes": null
    },
    {
      "id": "1263671",
      "postDate": "04/05/2021 15:30:31",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> - I very much agree with your main point, notwithstanding that some test molecules will be virtual compounds that don't exist, at least not yet, in databases.</p>\n<p>Yes, I think you can generate an InChIKey for any chemical structure you can make an InChI for, it is in essence a hashing operation on the InChI that yields the InChIKey.</p>",
      "rawMarkdown": "pcjimmmy - I very much agree with your main point, notwithstanding that some test molecules will be virtual compounds that don't exist, at least not yet, in databases.\n\nYes, I think you can generate an InChIKey for any chemical structure you can make an InChI for, it is in essence a hashing operation on the InChI that yields the InChIKey.",
      "votes": null
    },
    {
      "id": "1264292",
      "postDate": "04/06/2021 04:08:40",
      "content": "<p><a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">John</a> - thanks - I understand you to say that a hash of junk is possible :)</p>\n<p>As I have worked further on my post processing I found I had around 30 predicted InchI's that were so bad that they crash both RDKIT and openbabel fatally (model scores around 20 on LB).  </p>\n<p>This makes a \"not verified molecule\" prediction even more dangerous to future user when the prediction creates a fatal error that I can't seem to error trap in Jupyter or VS Code.  Both github's seem to have open issue for this type error.  </p>",
      "rawMarkdown": "[John](https://www.kaggle.com/jbomitchell) - thanks - I understand you to say that a hash of junk is possible :)\n\nAs I have worked further on my post processing I found I had around 30 predicted InchI's that were so bad that they crash both RDKIT and openbabel fatally (model scores around 20 on LB).  \n\nThis makes a \"not verified molecule\" prediction even more dangerous to future user when the prediction creates a fatal error that I can't seem to error trap in Jupyter or VS Code.  Both github's seem to have open issue for this type error.",
      "votes": null
    },
    {
      "id": "1264521",
      "postDate": "04/06/2021 08:10:48",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> - actually I mean to say something subtly different. A nonexistent molecule can still have a structure drawn for it, and that structure can be turned into a rule-compliant InChI, which in turn can be hashed into an InChI Key.</p>",
      "rawMarkdown": "pcjimmmy - actually I mean to say something subtly different. A nonexistent molecule can still have a structure drawn for it, and that structure can be turned into a rule-compliant InChI, which in turn can be hashed into an InChI Key.",
      "votes": null
    },
    {
      "id": "1266245",
      "postDate": "04/07/2021 15:37:49",
      "content": "<p><a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">John</a> - that is different.   Beyond RDKit and Openbabel is there a python library that can validate that structure is rule-compliant?</p>\n<p>My engineering degree was in extractive metallurgy - which meant I was classified as an inorganic chemistry major.  To validate a molecules structure we used a large set of balls and sticks to tinker toy together a molecule - that was a bit over 50 years ago.   I can recall it taking over an hour and a few broken sticks to build the structure of a single molecule on one of my chem final exams.  My post processing code is a bit slow - but I can now validate the structure of 1.6 million molecules in less than a day.</p>",
      "rawMarkdown": "[John](https://www.kaggle.com/jbomitchell) - that is different.   Beyond RDKit and Openbabel is there a python library that can validate that structure is rule-compliant?\n\nMy engineering degree was in extractive metallurgy - which meant I was classified as an inorganic chemistry major.  To validate a molecules structure we used a large set of balls and sticks to tinker toy together a molecule - that was a bit over 50 years ago.   I can recall it taking over an hour and a few broken sticks to build the structure of a single molecule on one of my chem final exams.  My post processing code is a bit slow - but I can now validate the structure of 1.6 million molecules in less than a day.",
      "votes": null
    },
    {
      "id": "1266727",
      "postDate": "04/08/2021 03:00:54",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> Thank you for the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318\" target=\"_blank\">clarification</a> specifying a list of <strong>approved external data</strong> <a href=\"https://www.kaggle.com/c/bms-molecular-translation/data\" target=\"_blank\">extra_approved_InChIs.csv</a></p>\n<ul>\n<li>2424186 Train rows</li>\n<li>9998711 Extra approved rows (which do not overlap with Test)</li>\n</ul>\n<p>i.e. 12422897 Total InChIs approved for use in training</p>",
      "rawMarkdown": "addisonhoward Thank you for the [clarification](https://www.kaggle.com/c/bms-molecular-translation/discussion/231318) specifying a list of **approved external data** [extra_approved_InChIs.csv](https://www.kaggle.com/c/bms-molecular-translation/data)\n - 2424186 Train rows\n - 9998711 Extra approved rows (which do not overlap with Test)\n\ni.e. 12422897 Total InChIs approved for use in training",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1256553,
      "author_name": "alansun17904",
      "author_url": "",
      "post_date": "03/30/2021 02:28:02",
      "content": "<p>Quoted from the competition rules:</p>\n<blockquote>\n  <p>External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1260313,
      "author_name": "jakealbrecht1337",
      "author_url": "",
      "post_date": "04/02/2021 02:31:55",
      "content": "<p>Thanks for your question!  Using open cheminformatics libraries is encouraged for the competition.  Data augmentation using GANs would also be an interesting approach to explore.  PubChem and ChemSpider are good resources to explore the diversity of chemistry, though they contain many compounds that are very different from the competition data.  The training and test set were selected to focus on molecules with properties relevant to pharmaceuticals with enough examples to reduce the value of external data, .  Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified.  Good Luck!</p>\n<p>Edit: After further discussion we've updated the rules on the use of external data, <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318\" target=\"_blank\">see this post</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1260641,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/02/2021 09:35:14",
          "content": "<blockquote>\n  <p>Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified</p>\n</blockquote>\n<p>Can you elaborate on this? What is meant by \"biasing the predictions\" - surely just training a model with external data would bias the model's predictions?</p>\n<p>I can't find anything in the competition rules that reflects your comment so it would be good to know what is and isn't allowed. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1260677,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/02/2021 10:37:36",
          "content": "<p><a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> thanks for indicating that you ar e using external data  ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1260891,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "04/02/2021 13:58:41",
          "content": "<p><a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> Some test molecules may be present in PubChem data. Most of the molecules used in this competition will be likely found in many public sources / publications.</p>\n<p>So, it is <strong>not allowed</strong> to use PubChem data for:</p>\n<ol>\n<li>Post-processing predictions to validate and \"fix\" the InChI output strings</li>\n<li>Training with data containing test set labels</li>\n<li>Finetuning on data containing test set labels</li>\n</ol>\n<p>Previous Kaggle <strong>disqualification</strong> references:</p>\n<ul>\n<li>1st prize winner of PetFinder contest was <strong>disqualified</strong> for <a href=\"https://www.kaggle.com/c/petfinder-adoption-prediction/discussion/125436\" target=\"_blank\">scraping a website to get test labels</a></li>\n<li>Scraping of test labels in <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80665\" target=\"_blank\">Quora Insincere Questions Classification</a></li>\n<li>Scraping in <a href=\"https://www.kaggle.com/c/ashrae-energy-prediction/discussion/116840\" target=\"_blank\">ASHRAE - Great Energy Predictor III</a></li>\n</ul>\n<p>So the Organizers (Jacob Albrecht) decided to disqualify such solutions because they do not generalize to new and unseen data, and so the <strong>solution is not very useful anyway</strong>.</p>\n<p>However, I believe that it is <strong>allowed</strong> to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.</p>\n<hr>\n<p>[Update] <a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> Jacob thank you for <strong>upvoting this comment</strong>. Could you please also reply to <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> below so that the competition rules and expectations are clear to everyone. Also please see my new comment about PubChem database below for more context.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1260904,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/02/2021 14:17:06",
          "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> Thank you, but I assume you are not speaking officially on behalf of the organisers. It would be good to get an official explanation (and addition to the ruleset) of what exactly is and isn't allowed.</p>\n<ol>\n<li><p>You say that \"it is allowed to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.\" As competitors do not have the test set labels, it would be impossible for a competitor to check whether any external data has 'unintentional' overlap. If the organisers are aware of any datasets that have overlap (for example, the datasets from which they generated their data), then they should make clear that these datasets are off-limits.</p></li>\n<li><p>Your examples of previous disqualification are a bit shaky I feel. Only your first link had someone disqualified, and in that case it was a very blatant attempt at concealed cheating. In the third link, the organisers <a href=\"https://www.kaggle.com/c/ashrae-energy-prediction/discussion/117357\" target=\"_blank\">specifically allowed</a> using this scraped data for the competition! (and it was also allowed in the second)</p></li>\n</ol>\n<p>I agree that this \"extreme\" scraping shouldn't be allowed (and I haven't made any attempt to scrape the test set labels), but what we need is an objective reasoning from the organisers on what they consider to be off-limits.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261011,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/02/2021 15:57:41",
          "content": "<p>\"So, it is not allowed to use PubChem data for: Post-processing predictions to validate and \"fix\" the InChI output strings\"</p>\n<p>this is called white list filtering. In commercial car license plate applications, we are sometimes asked to identify the license plate registered cars coming in and out of a car park. Because cars are registered in the whitelist, we know the prediction must be one of those registered.</p>\n<p>This simple trick helps us to correct errors if we cannot get all numbers correct.</p>\n<p>verification and post-processing of InChI output strings is a smart and valid trick I think</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261017,
          "author_name": "jakealbrecht1337",
          "author_url": "",
          "post_date": "04/02/2021 16:09:54",
          "content": "<p>For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.  .  Fixing predictions to nearest published InChIs is not allowed.  A simple check that a predicted label was not included in the training data is good due diligence.</p>\n<p>Edit: We've updated the rules on the use of external data <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318\" target=\"_blank\">see this post</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 1261040,
              "author_name": "stassl",
              "author_url": "",
              "post_date": "04/02/2021 16:36:57",
              "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> what about unsupervised pretraining on test images or pseudolabeling them using model predictions - are they allowed?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1261391,
              "author_name": "ryanzhang",
              "author_url": "",
              "post_date": "04/03/2021 03:24:43",
              "content": "<blockquote>\n  <p>to ensure test labels were not included in training…</p>\n</blockquote>\n<p>By the wording of this, pseudo labels shouldn't be allowed, since any decent model should get a great portion of test labels perfectly correct. But I don't see how you can enforce this on participants. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1261023,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "04/02/2021 16:16:56",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> You replied that \"fixing predictions\" is not allowed, but what if someone \"unintentionally\" included valid InChI strings in the training set itself (and fixes predictions that way)?</p>\n<p>As pointed out by <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a>, it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.</p>\n<p>So, please see my comment below and decide whether to <strong>disallow all external data sources</strong> including PubChem, because any other approach will be ambiguous.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261041,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/02/2021 16:38:32",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> why limit this to prize winners?  It would be simpler to say that all entries must comply with what you wrote  (edited to remove reference to prize winners):</p>\n<blockquote>\n  <p>If you use external data in a prize winning submission, you should ensure that test labels were not included in training. Fixing predictions to nearest published InChIs is not allowed. A simple check that a predicted label was not included in the training data is good due diligence.</p>\n</blockquote>\n<p>Granted, you won't check this for each and every submissions, but same is true of any restriction set forth in competition rules.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261045,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/02/2021 16:40:34",
          "content": "<blockquote>\n  <p>it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.</p>\n</blockquote>\n<p>Then don't include external data if you aren't sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261051,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/02/2021 16:51:16",
          "content": "<blockquote>\n  <p>For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.</p>\n</blockquote>\n<p>That's not what the rules says. (§12) says that you have the right to verify compliance with <strong>\"these rules\"</strong> (i.e. the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/rules\" target=\"_blank\">rules page</a>). I think any additional rules described in the forums should be formalised in the rules page for those not reading every post in the forums. One would expect the rules page to contain the full set of competition rules.</p>\n<p>As <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> asks, what about unsupervised learning/pseudo-learning using the test set (and without external data)? </p>\n<p>I think we need a better definition and explanation than \"we will decide what is allowed <em>after</em> the competition\". Competitors need to be saved working thousands of hours only to be told that the organisers didn't like their solution so they're being disqualified</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261127,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/02/2021 18:01:29",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> A common approach in OCR/STT is to train a language model and do beam search. So it means training a LM on PubChem i.e. not allowed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261171,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/02/2021 18:52:54",
          "content": "<p>a better solution for future competition is to allow the use of PubChem (including mapping to valid InChI), but a large percentage of the test data (e.g. 50%) has ground truth InChI not found in any public dataset like PubChem at all.</p>\n<p>in this case you really test the generalization of the algorithm and you do not restrict the creativity of kagglers to use external data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261238,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "04/02/2021 20:47:37",
          "content": "<p>First of all I will declare the vested interest that I have as a computational chemist. </p>\n<p>My view of Kaggle is that good competitions should and do allow people with quite distinct expertise to collaborate and compete on problems that defy the usual boundaries. Molecular Translation could and should be one such competition. Rather than rant or downvote comments I disagree with, I'll try to explain carefully why I fear that these interpretations of the rules are discriminating unreasonably against people with chemistry and chem(o)informatics skillsets, and unduly in favour of specialists in heavy duty machine learning and image recognition, thus reducing the scope and appeal of this contest.</p>\n<p>(1) On checking whether InChIs, either one-at-a-time or as a cohort of predictions, represent real molecules by cross-validating against external databases. This seems to me to be a fair and sensible way of assessing the proportion of fully correct InChIs coming out of a predictor. It doesn't tell us that the InChI represents the test molecule, just that it represents some molecule. It is manifestly not looking at the labels of the test data (images whose true labels don't exist anywhere on the internet), any more than feedback via a LB score is. The size of the dataset already precludes use of low-throughput methods such as browsing individual database entries to make a meaningful advance in LB score. Some of the test images also probably represent virtual compounds that won't be in PubChem or ChemSpider, so some correct InChIs won't be recognised by databases - and that is part of the challenge. </p>\n<p>(2) InChIs are extremely fiddly things to get right, and software &amp; databases seem to reject non-canonical InChIs outright even if they can handle non-canonical SMILES. So getting the InChI string right seems to be as big a challenge as identifying the compound (say by its graph or SMILES). The prohibition makes it harder to leverage chemoinformatics skills to check InChIs. I'd assume that this ruling must favour the image-to-SMILES-to-InChI route (the last step via external software) over direct image-to-InChI prediction. By making InChI checking harder, isn't it more likely that the winning solutions will be SMILES predictors?</p>\n<p>(3) Every Kaggle competition that I've been involved in has encouraged ensembling. An obvious ensembling route given predictions A &amp; B for molecule M is to choose whichever of A(M) or B(M) is in PubChem, whenever one and only one is present. If both or neither are present, stick with what you think is the better predictor.</p>\n<p>(4) Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction, better than predicting one 'average' InChI for everything. The leaderboard shows that the top teams are already doing better than this, but if someone wants to get a silver or bronze by the method I describe, why stop them?</p>\n<p>(5) There are alternative methods of checking InChI validity via external software, which I assume are still allowed. I expect that rejection of an InChI by chemoinformatics software usually means that the InChI fails to represent any feasible molecular graph, but I can't be absolutely sure that it doesn't in fact mean that a database lookup failed to generate a hit. So it seems hard to know what InChI checking procedures are compliant with the suggested interpretation and which aren't.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261330,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/03/2021 00:37:49",
          "content": "<p><a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">@jbomitchell</a> I really like your points, and they reflect the same thoughts I have!</p>\n<p>Since the test set doesn't appear to have been generated from any public dataset, I can't see how a public dataset would unreasonably bias a solution towards <em>this</em> test set (as opposed to any other potential test set from the same distribution)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261497,
          "author_name": "joblessphysicist",
          "author_url": "",
          "post_date": "04/03/2021 05:45:54",
          "content": "<p>Everyone knows using more training data and cross checking with external databases would be beneficial to improve the accuracy. I am sure the organizer knows it too. Try ask yourself why do they want to host a competition and pay others to do something they already know ? Why not they do it by themselves then ? If you are the organizer, would you prefer to see a competition of different new model ideas or a competition of getting a larger external training data set or a competition of tweaking predictions with a larger ensemble of external inchi databases ? It depends on what the organizer is looking for. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261650,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/03/2021 10:04:25",
          "content": "<p>And that's one of the reasons why I advocate for kernel competitions… no way to see private data upfront then.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261681,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/03/2021 10:49:00",
          "content": "<p>i would opt for:</p>\n<ol>\n<li>public LB : csv submission and one can download data</li>\n<li>private LB : code submission and unseen data</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261708,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 11:11:52",
          "content": "<blockquote>\n  <p>Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction</p>\n</blockquote>\n<p>This is useless for non published molecules.  It looks like host wants to improve retrieval from known molecules they have access to, and that may be different from pubchem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1262953,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "04/04/2021 23:29:36",
          "content": "<p>I am guessing that more clarification is coming but I find this statement to be a questionable decision by the host.</p>\n<blockquote>\n  <p>Fixing predictions to nearest published InChIs is not allowed&gt; </p>\n</blockquote>\n<p>If I was handed this finished product to use, the very FIRST question I would have as a user - is the predicted molecule real/known.  If something like RDKIT returns None or crashes with the predicted InchI as a user I would want to KNOW.  ( I don't know enough about RDKIT to assume that a crash or a returned None is not a real molecule, but odds are not good.)</p>\n<p>I don't understand which would be the greatest issue - a model with a super low score where many of the molecules are not real or a model where the prediction has been post fixed to the nearest low scoring \"real\" InChi.  </p>\n<p>From my reading it seems that \"searching\" is done using InChi_key.  Can you generate a key with a non real molecule?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263395,
          "author_name": "thomasseleck",
          "author_url": "",
          "post_date": "04/05/2021 11:29:28",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> : Can you tell us if the usage of RDKit is allowed to check predictions ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263671,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "04/05/2021 15:30:31",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> - I very much agree with your main point, notwithstanding that some test molecules will be virtual compounds that don't exist, at least not yet, in databases.</p>\n<p>Yes, I think you can generate an InChIKey for any chemical structure you can make an InChI for, it is in essence a hashing operation on the InChI that yields the InChIKey.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264292,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "04/06/2021 04:08:40",
          "content": "<p><a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">John</a> - thanks - I understand you to say that a hash of junk is possible :)</p>\n<p>As I have worked further on my post processing I found I had around 30 predicted InchI's that were so bad that they crash both RDKIT and openbabel fatally (model scores around 20 on LB).  </p>\n<p>This makes a \"not verified molecule\" prediction even more dangerous to future user when the prediction creates a fatal error that I can't seem to error trap in Jupyter or VS Code.  Both github's seem to have open issue for this type error.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264521,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "04/06/2021 08:10:48",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> - actually I mean to say something subtly different. A nonexistent molecule can still have a structure drawn for it, and that structure can be turned into a rule-compliant InChI, which in turn can be hashed into an InChI Key.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266245,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "04/07/2021 15:37:49",
          "content": "<p><a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">John</a> - that is different.   Beyond RDKit and Openbabel is there a python library that can validate that structure is rule-compliant?</p>\n<p>My engineering degree was in extractive metallurgy - which meant I was classified as an inorganic chemistry major.  To validate a molecules structure we used a large set of balls and sticks to tinker toy together a molecule - that was a bit over 50 years ago.   I can recall it taking over an hour and a few broken sticks to build the structure of a single molecule on one of my chem final exams.  My post processing code is a bit slow - but I can now validate the structure of 1.6 million molecules in less than a day.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261013,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "04/02/2021 16:04:13",
      "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>, I have just checked and found some of the test molecules (<strong>private labels</strong>) to be present on PubChem.</p>\n<p>However PubChem is the largest (and still growing) public chemical database with <strong>103 Million</strong> molecules from 748 reputed sources, and it is not very easy to identify the competition test set (1.6M test molecules). Moreover, 103M is such a large number that it is reasonable to assume that most molecules are present in this dataset already i.e. a solution using this should be a \"general\" solution and should not be considered as scraping! Also, this approach \"may not compromise the utility of the model outside of the test set\" because other molecules encountered in publications in the future would <strong>also be added to PubChem anyway</strong>.</p>\n<p>Could you please make a decision to <strong>fully allow</strong> PubChem or <strong>disallow</strong> PubChem (and probably disqualify all other external data sources as well).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1261054,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/02/2021 16:53:09",
          "content": "<p>I agree. Having vague and post-hoc-decided rules has caused issues in past competitions. If it is not possible to have an objective measure of what external data is allowed, then it should be completely disallowed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261227,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/02/2021 20:15:11",
          "content": "<p>From the train set only 326 moleculer formulas (haven't checked isomers) are not present in the PubChem. Test set should be similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261334,
          "author_name": "infy2097",
          "author_url": "",
          "post_date": "04/03/2021 00:58:39",
          "content": "<p>I suspect the reason for the organizer's concern is that in the real-world usage scenario the tool would be applied to internal documents and other documents that haven't had their chemistry indexed into Pubchem, so an assumption that the majority/all compounds will be present in PubChem will breakdown. Having said that I could imagine that even if none of the compounds were in PubChem, that a large dataset like PubChem would be useful in establishing what sort of motifs are common in real-world chemicals!</p>\n<p>Ideally the training/test set would contain mostly virtual compounds, that are plausible pharmaceutical compounds. Having some of them occur in PubChem would be fine, you don't want PubChem being used as a stop list either! Anyway the data set is what it is…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261747,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "04/03/2021 12:01:33",
          "content": "<p><a href=\"https://www.kaggle.com/infy2097\" target=\"_blank\">@infy2097</a> - a good point. I think there are at least two distinct real worlds here.</p>\n<p>One is the world of the Blue Obelisk gang, where the objective is to open up and add value to the chemical literature. They would be happy to see cooperation and interoperability between different databases, provided that everything is Open Access.</p>\n<p>The other is the world of Pharma, where the escape of your proprietary compound's InChI into the open internet might cost a billion dollars. Less keen on the use of external data without strong security!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261224,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "04/02/2021 20:03:48",
      "content": "<p>At this point, I think the best solution is to disallow all external data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1261653,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/03/2021 10:07:20",
          "content": "<p>My personal opinion is that it would be better to disallow all external data in most competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261705,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 11:09:11",
          "content": "<p><a href=\"https://www.kaggle.com/Psi\" target=\"_blank\">@Psi</a> this is tempting indeed.  Issue with that is to still allow pretrained models.  And these models are trained on external data…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261736,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/03/2021 11:47:36",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I agree, it is not an easy setup … Kaggle could approve certain pretrained repositories upfront (timm, tensorhub, huggingface, etc.) … but yeah there probably is no perfect setup unfortunately.</p>\n<p>I am just concerned about the general importance of external data in many competitions and I would rather prefer to focus on modeling than on finding the right data to use. And unfortunately there are often also leaks (not even on purpose) involved.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261741,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 11:55:14",
          "content": "<p>I agree with you.  Very often what differentiate winner from other top finisher is the use of more external data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261849,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "04/03/2021 14:05:14",
          "content": "<p>I completely agree, Kaggle should disallow usage of any external data in all competitions, so everyone could focus in building better solutions for the competition dataset and not spend time searching for external data and leakages. It would save a lot of time for Kaggle and competitors and make rules easier to understand without any potentital obfuscated disqualification issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261855,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/03/2021 14:14:31",
          "content": "<p>I agree that external data should not be allowed, but the line can become blurry with pre-trained models (which I think are certainly a net positive for competitions). Where is the line between \" software library\" and \"external data\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261873,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2021 14:32:40",
          "content": "<p>A good way of handling external data was the last Lyft competition.  External data had to be explicitly approved by host.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1261947,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/03/2021 15:52:48",
          "content": "<p>I agree <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> - this was perfectly handled by <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> .</p>\n<p>I think it wouldn't be too hard to whitelist a few pretrained models and then for others make an approval process. It's not like competitors post 1000 different external types of pre-trained models, usually those are maybe 10-20.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263633,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "04/05/2021 14:50:47",
          "content": "<p>I completely disagree with the idea of banning external data and information. Leveraging suitable external sources brings a lot of richness and interest to Kaggle, opening up more avenues for creative problem solving. Without this you'd be left with just a (IMHO boring) coding competition, plus some advantage for those with access to local GPUs. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1261302,
      "author_name": "addisonhoward",
      "author_url": "",
      "post_date": "04/02/2021 22:45:46",
      "content": "<p>Hey all,</p>\n<p>We're working on a clarification but it will take a few days (and we recommend holding off on external data in the meantime, for those of you who want to be on the safe side).</p>\n<p>Stay tuned</p>",
      "votes": null,
      "replies": [
        {
          "id": 1261331,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "04/03/2021 00:38:11",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266727,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "04/08/2021 03:00:54",
          "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> Thank you for the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318\" target=\"_blank\">clarification</a> specifying a list of <strong>approved external data</strong> <a href=\"https://www.kaggle.com/c/bms-molecular-translation/data\" target=\"_blank\">extra_approved_InChIs.csv</a></p>\n<ul>\n<li>2424186 Train rows</li>\n<li>9998711 Extra approved rows (which do not overlap with Test)</li>\n</ul>\n<p>i.e. 12422897 Total InChIs approved for use in training</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1255971": "A Kaggle discussion topic was **deleted because of copyright issues**, as per today's update in the topic [Winning solutions from similar competition: Molecular Translation into SMILES](https://www.kaggle.com/c/bms-molecular-translation/discussion/223381).\n\nDeleted post - [https://www.kaggle.com/c/bms-molecular-translation/discussion/223299](https://www.kaggle.com/c/bms-molecular-translation/discussion/223299)\nDeleted kernel - [https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection](https://www.kaggle.com/jinssaa/generate-molecular-bounding-box-for-detection)\n\nIn order to avoid other such potential issues, can the Organizers please reply about usage of below public information:\n1. Are we allowed to use external data from PubChem for data augmentation and validation?\n   The winners of the Dacon competition used additional data from pubchem to make their dataset more balanced, and also for model generalization on new and unseen types of compounds.\n   A related discussion post with [REST API](https://www.kaggle.com/c/bms-molecular-translation/discussion/225876) for \"clean\" images\n   Python library [PubChemPy](https://pubchempy.readthedocs.io/en/latest/guide/introduction.html) (MIT License)\n2. Similar to PubChem, can we use other open access chemistry databases such as ChemSpider, ChEMBL, and others?\n3. Can we use python libraries [RDKit](https://www.rdkit.org/), [Inchi](https://www.inchi-trust.org/downloads/), [OpenBabel](http://openbabel.org/wiki/Main_Page) for back and forth conversions (inchi `<->` smiles `<->` mol), validations, image generation and new data generation?\n4. Other data augmentation techniques such as GANs using above external data\n\n@Inversion - could you please confirm that above items meet the [competition Rules](https://www.kaggle.com/c/bms-molecular-translation/rules) wrt external data and licensing requirements.",
    "1256553": "Quoted from the competition rules:\n> External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).",
    "1260313": "Thanks for your question!  Using open cheminformatics libraries is encouraged for the competition.  Data augmentation using GANs would also be an interesting approach to explore.  PubChem and ChemSpider are good resources to explore the diversity of chemistry, though they contain many compounds that are very different from the competition data.  The training and test set were selected to focus on molecules with properties relevant to pharmaceuticals with enough examples to reduce the value of external data, ~~but there is no prohibition~~.  Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified.  Good Luck!\n\nEdit: After further discussion we've updated the rules on the use of external data, [see this post](https://www.kaggle.com/c/bms-molecular-translation/discussion/231318)",
    "1260641": "> Of course, given that the objective of the competition is an image to InChI translator, an approach that attempts to identify the test set compounds in external data sources or otherwise biases the predictions will be disqualified\n\nCan you elaborate on this? What is meant by \"biasing the predictions\" - surely just training a model with external data would bias the model's predictions?\n\nI can't find anything in the competition rules that reflects your comment so it would be good to know what is and isn't allowed. Thanks!",
    "1260677": "anokas thanks for indicating that you ar e using external data  ;)",
    "1260891": "anokas Some test molecules may be present in PubChem data. Most of the molecules used in this competition will be likely found in many public sources / publications.\n\nSo, it is **not allowed** to use PubChem data for:\n1. Post-processing predictions to validate and \"fix\" the InChI output strings\n2. Training with data containing test set labels\n3. Finetuning on data containing test set labels\n\nPrevious Kaggle **disqualification** references:\n- 1st prize winner of PetFinder contest was **disqualified** for [scraping a website to get test labels](https://www.kaggle.com/c/petfinder-adoption-prediction/discussion/125436)\n- Scraping of test labels in [Quora Insincere Questions Classification](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/80665)\n- Scraping in [ASHRAE - Great Energy Predictor III](https://www.kaggle.com/c/ashrae-energy-prediction/discussion/116840)\n\nSo the Organizers (Jacob Albrecht) decided to disqualify such solutions because they do not generalize to new and unseen data, and so the **solution is not very useful anyway**.\n\nHowever, I believe that it is **allowed** to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.\n__________________________________________________________________________________________\n[Update] @jakealbrecht1337 Jacob thank you for **upvoting this comment**. Could you please also reply to @anokas below so that the competition rules and expectations are clear to everyone. Also please see my new comment about PubChem database below for more context.",
    "1260904": "sirishks Thank you, but I assume you are not speaking officially on behalf of the organisers. It would be good to get an official explanation (and addition to the ruleset) of what exactly is and isn't allowed.\n\n1. You say that \"it is allowed to use external data such as PubChem to build robust solutions, as long as you ensure that your solution does not use any test molecules, even unintentionally.\" As competitors do not have the test set labels, it would be impossible for a competitor to check whether any external data has 'unintentional' overlap. If the organisers are aware of any datasets that have overlap (for example, the datasets from which they generated their data), then they should make clear that these datasets are off-limits.\n\n2. Your examples of previous disqualification are a bit shaky I feel. Only your first link had someone disqualified, and in that case it was a very blatant attempt at concealed cheating. In the third link, the organisers [specifically allowed](https://www.kaggle.com/c/ashrae-energy-prediction/discussion/117357) using this scraped data for the competition! (and it was also allowed in the second)\n\nI agree that this \"extreme\" scraping shouldn't be allowed (and I haven't made any attempt to scrape the test set labels), but what we need is an objective reasoning from the organisers on what they consider to be off-limits.",
    "1261011": "\"So, it is not allowed to use PubChem data for: Post-processing predictions to validate and \"fix\" the InChI output strings\"\n\n\nthis is called white list filtering. In commercial car license plate applications, we are sometimes asked to identify the license plate registered cars coming in and out of a car park. Because cars are registered in the whitelist, we know the prediction must be one of those registered.\n\nThis simple trick helps us to correct errors if we cannot get all numbers correct.\n\n\nverification and post-processing of InChI output strings is a smart and valid trick I think",
    "1261013": "jakealbrecht1337 @addisonhoward, I have just checked and found some of the test molecules (**private labels**) to be present on PubChem.\n\nHowever PubChem is the largest (and still growing) public chemical database with **103 Million** molecules from 748 reputed sources, and it is not very easy to identify the competition test set (1.6M test molecules). Moreover, 103M is such a large number that it is reasonable to assume that most molecules are present in this dataset already i.e. a solution using this should be a \"general\" solution and should not be considered as scraping! Also, this approach \"may not compromise the utility of the model outside of the test set\" because other molecules encountered in publications in the future would **also be added to PubChem anyway**.\n\nCould you please make a decision to **fully allow** PubChem or **disallow** PubChem (and probably disqualify all other external data sources as well).",
    "1261017": "For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.  ~~If you use external data in a prize winning submission, your model documentation should include the steps you took to ensure test labels were not included in training~~.  Fixing predictions to nearest published InChIs is not allowed.  A simple check that a predicted label was not included in the training data is good due diligence.\n\nEdit: We've updated the rules on the use of external data [see this post](https://www.kaggle.com/c/bms-molecular-translation/discussion/231318)",
    "1261023": "jakealbrecht1337 You replied that \"fixing predictions\" is not allowed, but what if someone \"unintentionally\" included valid InChI strings in the training set itself (and fixes predictions that way)?\n\nAs pointed out by @anokas, it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.\n\nSo, please see my comment below and decide whether to **disallow all external data sources** including PubChem, because any other approach will be ambiguous.",
    "1261040": "jakealbrecht1337 what about unsupervised pretraining on test images or pseudolabeling them using model predictions - are they allowed?",
    "1261041": "jakealbrecht1337 why limit this to prize winners?  It would be simpler to say that all entries must comply with what you wrote  (edited to remove reference to prize winners):\n\n>  If you use external data in a prize winning submission, you should ensure that test labels were not included in training. Fixing predictions to nearest published InChIs is not allowed. A simple check that a predicted label was not included in the training data is good due diligence.\n\nGranted, you won't check this for each and every submissions, but same is true of any restriction set forth in competition rules.",
    "1261045": ">  it is not reasonable to expect participants to know beforehand whether an external public data item is present in test set or not.\n\nThen don't include external data if you aren't sure.",
    "1261051": "> For prize winners we reserve the right (§12) to review the methodology and model to ensure that external augmentation (if used) hasn't compromised the utility of the model outside of the test set.\n\nThat's not what the rules says. (§12) says that you have the right to verify compliance with **\"these rules\"** (i.e. the [rules page](https://www.kaggle.com/c/bms-molecular-translation/rules)). I think any additional rules described in the forums should be formalised in the rules page for those not reading every post in the forums. One would expect the rules page to contain the full set of competition rules.\n\nAs @stassl asks, what about unsupervised learning/pseudo-learning using the test set (and without external data)? \n\nI think we need a better definition and explanation than \"we will decide what is allowed _after_ the competition\". Competitors need to be saved working thousands of hours only to be told that the organisers didn't like their solution so they're being disqualified",
    "1261054": "I agree. Having vague and post-hoc-decided rules has caused issues in past competitions. If it is not possible to have an objective measure of what external data is allowed, then it should be completely disallowed.",
    "1261127": "jakealbrecht1337 A common approach in OCR/STT is to train a language model and do beam search. So it means training a LM on PubChem i.e. not allowed?",
    "1261171": "a better solution for future competition is to allow the use of PubChem (including mapping to valid InChI), but a large percentage of the test data (e.g. 50%) has ground truth InChI not found in any public dataset like PubChem at all.\n\nin this case you really test the generalization of the algorithm and you do not restrict the creativity of kagglers to use external data.",
    "1261224": "At this point, I think the best solution is to disallow all external data.",
    "1261227": "From the train set only 326 moleculer formulas (haven't checked isomers) are not present in the PubChem. Test set should be similar.",
    "1261238": "First of all I will declare the vested interest that I have as a computational chemist. \n\nMy view of Kaggle is that good competitions should and do allow people with quite distinct expertise to collaborate and compete on problems that defy the usual boundaries. Molecular Translation could and should be one such competition. Rather than rant or downvote comments I disagree with, I'll try to explain carefully why I fear that these interpretations of the rules are discriminating unreasonably against people with chemistry and chem(o)informatics skillsets, and unduly in favour of specialists in heavy duty machine learning and image recognition, thus reducing the scope and appeal of this contest.\n\n(1) On checking whether InChIs, either one-at-a-time or as a cohort of predictions, represent real molecules by cross-validating against external databases. This seems to me to be a fair and sensible way of assessing the proportion of fully correct InChIs coming out of a predictor. It doesn't tell us that the InChI represents the test molecule, just that it represents some molecule. It is manifestly not looking at the labels of the test data (images whose true labels don't exist anywhere on the internet), any more than feedback via a LB score is. The size of the dataset already precludes use of low-throughput methods such as browsing individual database entries to make a meaningful advance in LB score. Some of the test images also probably represent virtual compounds that won't be in PubChem or ChemSpider, so some correct InChIs won't be recognised by databases - and that is part of the challenge. \n\n(2) InChIs are extremely fiddly things to get right, and software & databases seem to reject non-canonical InChIs outright even if they can handle non-canonical SMILES. So getting the InChI string right seems to be as big a challenge as identifying the compound (say by its graph or SMILES). The prohibition makes it harder to leverage chemoinformatics skills to check InChIs. I'd assume that this ruling must favour the image-to-SMILES-to-InChI route (the last step via external software) over direct image-to-InChI prediction. By making InChI checking harder, isn't it more likely that the winning solutions will be SMILES predictors?\n\n(3) Every Kaggle competition that I've been involved in has encouraged ensembling. An obvious ensembling route given predictions A & B for molecule M is to choose whichever of A(M) or B(M) is in PubChem, whenever one and only one is present. If both or neither are present, stick with what you think is the better predictor.\n\n(4) Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction, better than predicting one 'average' InChI for everything. The leaderboard shows that the top teams are already doing better than this, but if someone wants to get a silver or bronze by the method I describe, why stop them?\n\n(5) There are alternative methods of checking InChI validity via external software, which I assume are still allowed. I expect that rejection of an InChI by chemoinformatics software usually means that the InChI fails to represent any feasible molecular graph, but I can't be absolutely sure that it doesn't in fact mean that a database lookup failed to generate a hit. So it seems hard to know what InChI checking procedures are compliant with the suggested interpretation and which aren't.",
    "1261302": "Hey all,\n\nWe're working on a clarification but it will take a few days (and we recommend holding off on external data in the meantime, for those of you who want to be on the safe side).\n\nStay tuned",
    "1261330": "jbomitchell I really like your points, and they reflect the same thoughts I have!\n\nSince the test set doesn't appear to have been generated from any public dataset, I can't see how a public dataset would unreasonably bias a solution towards _this_ test set (as opposed to any other potential test set from the same distribution)",
    "1261331": "Thank you!",
    "1261334": "I suspect the reason for the organizer's concern is that in the real-world usage scenario the tool would be applied to internal documents and other documents that haven't had their chemistry indexed into Pubchem, so an assumption that the majority/all compounds will be present in PubChem will breakdown. Having said that I could imagine that even if none of the compounds were in PubChem, that a large dataset like PubChem would be useful in establishing what sort of motifs are common in real-world chemicals!\n\nIdeally the training/test set would contain mostly virtual compounds, that are plausible pharmaceutical compounds. Having some of them occur in PubChem would be fine, you don't want PubChem being used as a stop list either! Anyway the data set is what it is...",
    "1261391": "> to ensure test labels were not included in training...\n\nBy the wording of this, pseudo labels shouldn't be allowed, since any decent model should get a great portion of test labels perfectly correct. But I don't see how you can enforce this on participants.",
    "1261497": "Everyone knows using more training data and cross checking with external databases would be beneficial to improve the accuracy. I am sure the organizer knows it too. Try ask yourself why do they want to host a competition and pay others to do something they already know ? Why not they do it by themselves then ? If you are the organizer, would you prefer to see a competition of different new model ideas or a competition of getting a larger external training data set or a competition of tweaking predictions with a larger ensemble of external inchi databases ? It depends on what the organizer is looking for.",
    "1261650": "And that's one of the reasons why I advocate for kernel competitions... no way to see private data upfront then.",
    "1261653": "My personal opinion is that it would be better to disallow all external data in most competition.",
    "1261681": "i would opt for:\n\n1. public LB : csv submission and one can download data\n2. private LB : code submission and unseen data",
    "1261705": "Psi this is tempting indeed.  Issue with that is to still allow pretrained models.  And these models are trained on external data...",
    "1261708": "> Fixing to the nearest published InChI seems to me a very sensible way of generating a decent baseline prediction\n\nThis is useless for non published molecules.  It looks like host wants to improve retrieval from known molecules they have access to, and that may be different from pubchem.",
    "1261736": "cpmpml I agree, it is not an easy setup ... Kaggle could approve certain pretrained repositories upfront (timm, tensorhub, huggingface, etc.) ... but yeah there probably is no perfect setup unfortunately.\n\nI am just concerned about the general importance of external data in many competitions and I would rather prefer to focus on modeling than on finding the right data to use. And unfortunately there are often also leaks (not even on purpose) involved.",
    "1261741": "I agree with you.  Very often what differentiate winner from other top finisher is the use of more external data.",
    "1261747": "infy2097 - a good point. I think there are at least two distinct real worlds here.\n\nOne is the world of the Blue Obelisk gang, where the objective is to open up and add value to the chemical literature. They would be happy to see cooperation and interoperability between different databases, provided that everything is Open Access.\n\nThe other is the world of Pharma, where the escape of your proprietary compound's InChI into the open internet might cost a billion dollars. Less keen on the use of external data without strong security!",
    "1261849": "I completely agree, Kaggle should disallow usage of any external data in all competitions, so everyone could focus in building better solutions for the competition dataset and not spend time searching for external data and leakages. It would save a lot of time for Kaggle and competitors and make rules easier to understand without any potentital obfuscated disqualification issue.",
    "1261855": "I agree that external data should not be allowed, but the line can become blurry with pre-trained models (which I think are certainly a net positive for competitions). Where is the line between \" software library\" and \"external data\"?",
    "1261873": "A good way of handling external data was the last Lyft competition.  External data had to be explicitly approved by host.",
    "1261947": "I agree @cpmpml - this was perfectly handled by @iglovikov .\n\nI think it wouldn't be too hard to whitelist a few pretrained models and then for others make an approval process. It's not like competitors post 1000 different external types of pre-trained models, usually those are maybe 10-20.",
    "1262953": "I am guessing that more clarification is coming but I find this statement to be a questionable decision by the host.\n\n> Fixing predictions to nearest published InChIs is not allowed> \n\nIf I was handed this finished product to use, the very FIRST question I would have as a user - is the predicted molecule real/known.  If something like RDKIT returns None or crashes with the predicted InchI as a user I would want to KNOW.  ( I don't know enough about RDKIT to assume that a crash or a returned None is not a real molecule, but odds are not good.)\n\nI don't understand which would be the greatest issue - a model with a super low score where many of the molecules are not real or a model where the prediction has been post fixed to the nearest low scoring \"real\" InChi.  \n\nFrom my reading it seems that \"searching\" is done using InChi_key.  Can you generate a key with a non real molecule?",
    "1263395": "jakealbrecht1337 : Can you tell us if the usage of RDKit is allowed to check predictions ?",
    "1263633": "I completely disagree with the idea of banning external data and information. Leveraging suitable external sources brings a lot of richness and interest to Kaggle, opening up more avenues for creative problem solving. Without this you'd be left with just a (IMHO boring) coding competition, plus some advantage for those with access to local GPUs.",
    "1263671": "pcjimmmy - I very much agree with your main point, notwithstanding that some test molecules will be virtual compounds that don't exist, at least not yet, in databases.\n\nYes, I think you can generate an InChIKey for any chemical structure you can make an InChI for, it is in essence a hashing operation on the InChI that yields the InChIKey.",
    "1264292": "[John](https://www.kaggle.com/jbomitchell) - thanks - I understand you to say that a hash of junk is possible :)\n\nAs I have worked further on my post processing I found I had around 30 predicted InchI's that were so bad that they crash both RDKIT and openbabel fatally (model scores around 20 on LB).  \n\nThis makes a \"not verified molecule\" prediction even more dangerous to future user when the prediction creates a fatal error that I can't seem to error trap in Jupyter or VS Code.  Both github's seem to have open issue for this type error.",
    "1264521": "pcjimmmy - actually I mean to say something subtly different. A nonexistent molecule can still have a structure drawn for it, and that structure can be turned into a rule-compliant InChI, which in turn can be hashed into an InChI Key.",
    "1266245": "[John](https://www.kaggle.com/jbomitchell) - that is different.   Beyond RDKit and Openbabel is there a python library that can validate that structure is rule-compliant?\n\nMy engineering degree was in extractive metallurgy - which meant I was classified as an inorganic chemistry major.  To validate a molecules structure we used a large set of balls and sticks to tinker toy together a molecule - that was a bit over 50 years ago.   I can recall it taking over an hour and a few broken sticks to build the structure of a single molecule on one of my chem final exams.  My post processing code is a bit slow - but I can now validate the structure of 1.6 million molecules in less than a day.",
    "1266727": "addisonhoward Thank you for the [clarification](https://www.kaggle.com/c/bms-molecular-translation/discussion/231318) specifying a list of **approved external data** [extra_approved_InChIs.csv](https://www.kaggle.com/c/bms-molecular-translation/data)\n - 2424186 Train rows\n - 9998711 Extra approved rows (which do not overlap with Test)\n\ni.e. 12422897 Total InChIs approved for use in training"
  },
  "source": "meta"
}