{
  "id": 231318,
  "title": "Clarity about External Data use regarding PubChem, provided InChIs, etc.",
  "url": "/competitions/bms-molecular-translation/discussion/231318",
  "author_name": "Addison Howard",
  "post_date": "2021-04-08T01:52:06.407000",
  "votes": 33,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>In working with the BMS team, they have decided to limit any external chemical labels to only those provided in the competition dataset. A whitelist of those allowed (approx 10M) are now included on the Data tab as <code>Extra_Approved_InChIs.csv.gz</code> This limitation is specifically for chemical labels and also does not prohibit the use of any generated of data (i.e. GANs).</p>\n<p>The competition rules have been updated for avoidance of doubt with the following language:</p>\n<p>The following supersedes Section B.7.C below:</p>\n<p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. In addition, the only external chemical labels available for use are limited to records of molecules specified by the provided training and supplemental InChIs. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>",
  "messages": [
    {
      "id": 1266681,
      "postDate": "2021-04-08T01:52:06.407Z",
      "content": "<p>Hi All,</p>\n<p>In working with the BMS team, they have decided to limit any external chemical labels to only those provided in the competition dataset. A whitelist of those allowed (approx 10M) are now included on the Data tab as <code>Extra_Approved_InChIs.csv.gz</code> This limitation is specifically for chemical labels and also does not prohibit the use of any generated of data (i.e. GANs).</p>\n<p>The competition rules have been updated for avoidance of doubt with the following language:</p>\n<p>The following supersedes Section B.7.C below:</p>\n<p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. In addition, the only external chemical labels available for use are limited to records of molecules specified by the provided training and supplemental InChIs. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).</p>",
      "rawMarkdown": "Hi All,\n\nIn working with the BMS team, they have decided to limit any external chemical labels to only those provided in the competition dataset. A whitelist of those allowed (approx 10M) are now included on the Data tab as `Extra_Approved_InChIs.csv.gz` This limitation is specifically for chemical labels and also does not prohibit the use of any generated of data (i.e. GANs).\n\nThe competition rules have been updated for avoidance of doubt with the following language:\n\nThe following supersedes Section B.7.C below:\n\nC. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. In addition, the only external chemical labels available for use are limited to records of molecules specified by the provided training and supplemental InChIs. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).",
      "votes": 33
    },
    {
      "id": 1267556,
      "postDate": "2021-04-08T15:35:55.903Z",
      "content": "<p>Thanks for the clarification.</p>\n<p>What I am generally a little bit bothered with is that these restrictions practically only apply to competition winners. I wonder how Kaggle can ensure that all competitors adhere to these rules and similar ones in other competitions. I know this is not trivial, but kind-of important to the integrity of competitions.</p>",
      "rawMarkdown": "Thanks for the clarification.\n\nWhat I am generally a little bit bothered with is that these restrictions practically only apply to competition winners. I wonder how Kaggle can ensure that all competitors adhere to these rules and similar ones in other competitions. I know this is not trivial, but kind-of important to the integrity of competitions.",
      "votes": 8
    },
    {
      "id": 1267598,
      "postDate": "2021-04-08T16:08:48.737Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> this is crystal clear now.  I guess I have no more reason to procrastinate here ;)</p>",
      "rawMarkdown": "Thanks @addisonhoward @jakealbrecht1337 @inversion this is crystal clear now.  I guess I have no more reason to procrastinate here ;)",
      "votes": 5
    },
    {
      "id": 1266697,
      "postDate": "2021-04-08T02:13:34.747Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>!<br>\nWith the ultimate goal of this competition to develop a model capable of translating chemical images to a text representation, approaches that augment the training set with external data risk including test set compounds if they are in public databases.  Including these compounds would overfit the model, and reduce the ability to generalize to new chemical images.  Limiting the data with the compromise of supplemental InChIs will remove this risk, while also ensuring teams have freedom to develop innovative models and don't to waste the effort previously spent on developing code for synthetic data generation. </p>\n<p>Hopefully this makes the competition rules less ambiguous and ultimately increases the confidence in the capabilities of the winning solutions.  I want to thank the Kaggle team for their advice and effort over the past few days, and the community here for their work over the past month.  My expectations were exceeded within the first weeks of the competition, I am very impressed, and I look forward to the next two months!</p>",
      "rawMarkdown": "Thanks @addisonhoward!\nWith the ultimate goal of this competition to develop a model capable of translating chemical images to a text representation, approaches that augment the training set with external data risk including test set compounds if they are in public databases.  Including these compounds would overfit the model, and reduce the ability to generalize to new chemical images.  Limiting the data with the compromise of supplemental InChIs will remove this risk, while also ensuring teams have freedom to develop innovative models and don't to waste the effort previously spent on developing code for synthetic data generation. \n\nHopefully this makes the competition rules less ambiguous and ultimately increases the confidence in the capabilities of the winning solutions.  I want to thank the Kaggle team for their advice and effort over the past few days, and the community here for their work over the past month.  My expectations were exceeded within the first weeks of the competition, I am very impressed, and I look forward to the next two months!",
      "votes": 5,
      "replies": [
        {
          "id": 1266954,
          "postDate": "2021-04-08T07:40:00.900Z",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> thank you both for a clear update on competition rules. </p>\n<p>I have no clue about this, but it could be that top of LB includes submissions that do not comply with the new rules.  IN that case the top of LB could be over optimistic  I wonder if it is feasible and desirable to remove these submissions, if they exist.  Of course this would rely on declaration by submitter.</p>\n<p>I wonder what others think about this.  If no one else is bothered by too optimistic LB then so be it, no big deal.</p>",
          "rawMarkdown": "@jakealbrecht1337 @addisonhoward thank you both for a clear update on competition rules. \n\nI have no clue about this, but it could be that top of LB includes submissions that do not comply with the new rules.  IN that case the top of LB could be over optimistic  I wonder if it is feasible and desirable to remove these submissions, if they exist.  Of course this would rely on declaration by submitter.\n\nI wonder what others think about this.  If no one else is bothered by too optimistic LB then so be it, no big deal.",
          "votes": 2
        },
        {
          "id": 1267024,
          "postDate": "2021-04-08T08:52:10.100Z",
          "content": "<blockquote>\n  <p>don't waste the effort spent on developing code for synthetic data generation. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> Second thought.  Are you saying that generating new training data using training InCHI or supplemental InCHIs is not allowed?</p>\n<p>This would rule out common image augmentation at training time for instance.</p>",
          "rawMarkdown": ">  don't waste the effort spent on developing code for synthetic data generation. \n\n@jakealbrecht1337 Second thought.  Are you saying that generating new training data using training InCHI or supplemental InCHIs is not allowed?\n\nThis would rule out common image augmentation at training time for instance.\n\n"
        },
        {
          "id": 1267106,
          "postDate": "2021-04-08T10:23:12.813Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\n<code>it could be that top of LB includes submissions that do not comply with the new rules</code></p>\n<ul>\n<li>Organizers may not reset the LB, but <strong>Kaggle should definitely ensure</strong> that medals are not given to submissions made earlier, without following the new rule</li>\n</ul>\n<p><code>Are you saying that generating new training data using training InCHI or supplemental InCHIs is not allowed?</code></p>\n<ul>\n<li>I think he meant \"don't <strong>want to</strong> waste\" i.e. participants are encouraged to make use of synthetic data to build better models.</li>\n</ul>",
          "rawMarkdown": "@cpmpml \n` it could be that top of LB includes submissions that do not comply with the new rules`\n- Organizers may not reset the LB, but **Kaggle should definitely ensure** that medals are not given to submissions made earlier, without following the new rule\n\n`Are you saying that generating new training data using training InCHI or supplemental InCHIs is not allowed?`\n- I think he meant \"don't **want to** waste\" i.e. participants are encouraged to make use of synthetic data to build better models.",
          "votes": 1
        },
        {
          "id": 1267431,
          "postDate": "2021-04-08T14:06:39.963Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, I've updated the comment to clarify: generating synthetic images is ok, generating new molecules from a GAN trained on the supplied InChIs is ok.</p>",
          "rawMarkdown": "@cpmpml, I've updated the comment to clarify: generating synthetic images is ok, generating new molecules from a GAN trained on the supplied InChIs is ok.",
          "votes": 1
        },
        {
          "id": 1267546,
          "postDate": "2021-04-08T15:30:02.903Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> - if any top submissions are inadvertently breaking the rules, those users can be sure to unselect those submissions, and only select submissions that are in compliance.</p>",
          "rawMarkdown": "@cpmpml - if any top submissions are inadvertently breaking the rules, those users can be sure to unselect those submissions, and only select submissions that are in compliance.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1291427,
      "postDate": "2021-05-03T03:24:57.907Z",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>Thanks for clarification. But let me double check below:</p>\n<ol>\n<li>It is disallowed to validate my prediction using rdkit and fix it if it's invalid. Is it true?<br>\nI mean like <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229578\" target=\"_blank\">this method</a>.</li>\n<li>Above method is disallowed even if it's for pseudo labeling, not for final submission. Is it true?</li>\n<li>But it is OK to write my own code to validate whether the InChI is valid one or not, unless it doesn't lookup database which has possible chemicals in it. Is it true?</li>\n</ol>\n<p>P.S. Please pin this topic on top of discussions. It is the most important notice from organizer but currently hard to find in the bunch of community topics..</p>",
      "rawMarkdown": "@addisonhoward @inversion \n\nThanks for clarification. But let me double check below:\n1. It is disallowed to validate my prediction using rdkit and fix it if it's invalid. Is it true?\n    I mean like [this method](https://www.kaggle.com/c/bms-molecular-translation/discussion/229578).\n2. Above method is disallowed even if it's for pseudo labeling, not for final submission. Is it true?\n3. But it is OK to write my own code to validate whether the InChI is valid one or not, unless it doesn't lookup database which has possible chemicals in it. Is it true?\n\nP.S. Please pin this topic on top of discussions. It is the most important notice from organizer but currently hard to find in the bunch of community topics..\n",
      "votes": 3,
      "replies": [
        {
          "id": 1291687,
          "postDate": "2021-05-03T08:42:40.607Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313</a></p>\n<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> in his second sentence says:</p>\n<blockquote>\n  <p>Using open cheminformatics libraries is encouraged for the competition.</p>\n</blockquote>\n<p>But I also would like this to be verified to be sure!</p>\n<p>Edit: The link doesn't navigate me to the comment directly. ?</p>",
          "rawMarkdown": "https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313\n\n@jakealbrecht1337 in his second sentence says:\n>  Using open cheminformatics libraries is encouraged for the competition.\n\nBut I also would like this to be verified to be sure!\n\nEdit: The link doesn't navigate me to the comment directly. ?"
        },
        {
          "id": 1291706,
          "postDate": "2021-05-03T09:13:35.937Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Thanks, I can find the comment even though it doesn't directly navigate:) And also thank you for the original normalization script, it works for me too! <br>\nLet's wait for the host's comment.</p>",
          "rawMarkdown": "@nofreewill Thanks, I can find the comment even though it doesn't directly navigate:) And also thank you for the original normalization script, it works for me too! \nLet's wait for the host's comment.\n\n",
          "votes": 1
        },
        {
          "id": 1291715,
          "postDate": "2021-05-03T09:25:22.960Z",
          "content": "<p>Cheers! :)</p>",
          "rawMarkdown": "Cheers! :)"
        },
        {
          "id": 1292491,
          "postDate": "2021-05-04T02:47:51.613Z",
          "content": "<p>Yes, checking using <code>rdkit</code> is a smart way to canonicalize the InChI.  Other packages like <code>pubchempy</code> and <code>chemspipy</code> can query their respective databases; avoid that functionality please, thanks!</p>",
          "rawMarkdown": "Yes, checking using `rdkit` is a smart way to canonicalize the InChI.  Other packages like `pubchempy` and `chemspipy` can query their respective databases; avoid that functionality please, thanks!",
          "votes": 3
        },
        {
          "id": 1292521,
          "postDate": "2021-05-04T03:53:49.747Z",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> Thanks, it's crystal clear for me. Also thanks for hosting such a interesting competition, I'm really enjoying!</p>",
          "rawMarkdown": "@jakealbrecht1337 Thanks, it's crystal clear for me. Also thanks for hosting such a interesting competition, I'm really enjoying!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1267140,
      "postDate": "2021-04-08T10:44:36.913Z",
      "content": "<p>Hi organizers &amp; Kaggle,</p>\n<p>Thank you for the update. Can you confirm that using data <em>generated without external data</em> is not subject to this restriction (for example, for a team that does not use any external data at all)</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi organizers & Kaggle,\n\nThank you for the update. Can you confirm that using data _generated without external data_ is not subject to this restriction (for example, for a team that does not use any external data at all)\n\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 1267271,
          "postDate": "2021-04-08T12:34:18.410Z",
          "content": "<p>Do you mean, e.g., image augmentation? That is fine.</p>\n<p>The intent is to prevent, e.g., (to use a hypothetical extreme) generating or downloading images of all known chemical structures and training on those images.</p>",
          "rawMarkdown": "Do you mean, e.g., image augmentation? That is fine.\n\nThe intent is to prevent, e.g., (to use a hypothetical extreme) generating or downloading images of all known chemical structures and training on those images.",
          "votes": 2
        },
        {
          "id": 1267277,
          "postDate": "2021-04-08T12:38:53.880Z",
          "content": "<p>But is it fine to generate images from training InCHI or approved additional InCHI?</p>",
          "rawMarkdown": "But is it fine to generate images from training InCHI or approved additional InCHI?",
          "votes": 2
        },
        {
          "id": 1267288,
          "postDate": "2021-04-08T12:46:51.640Z",
          "content": "<p>I mean e.g. pseudo-labelling would be an example</p>",
          "rawMarkdown": "I mean e.g. pseudo-labelling would be an example",
          "votes": 3
        },
        {
          "id": 1267345,
          "postDate": "2021-04-08T13:31:00.070Z",
          "content": "<p>The issue is strictly the label. You can create (download, etc) whatever training data you want for the allowed labels. The model shouldn't be trained on data outside that list list of labels. </p>\n<p>Pseudo-labelling isn't an issue as long as (a) the model used to pseudo-labelling is only trained on the allowed labels, and (b) you're pseudo-labelling the test data.</p>\n<p>What is not allowed is pseudo-labelling that is effectively hand-labeling. For example, pseudo-labeling that predicts a test image, looks up that label from a database, adjusts it to a best match, etc. (Although there's no issue with adding a constraint in code that checks that the pseudo label is physically possible, e.g., has the correct number of bonds, etc.).</p>",
          "rawMarkdown": "The issue is strictly the label. You can create (download, etc) whatever training data you want for the allowed labels. The model shouldn't be trained on data outside that list list of labels. \n\nPseudo-labelling isn't an issue as long as (a) the model used to pseudo-labelling is only trained on the allowed labels, and (b) you're pseudo-labelling the test data.\n\nWhat is not allowed is pseudo-labelling that is effectively hand-labeling. For example, pseudo-labeling that predicts a test image, looks up that label from a database, adjusts it to a best match, etc. (Although there's no issue with adding a constraint in code that checks that the pseudo label is physically possible, e.g., has the correct number of bonds, etc.).",
          "votes": 5
        },
        {
          "id": 1267529,
          "postDate": "2021-04-08T15:18:35.423Z",
          "content": "<p>What if the model is NOT machine <em>trained</em>? It's a weird question in a data science competition, but it is the strategy we are using: generalized &amp; automated molecular recognition using simple image processing, similar to OSRA. Given a crappy image (e.g., missing bonds due to bad compression), multiple chemically reasonable candidates can be proposed. A database look up helps pick a most probable one. We only do <strong>exact</strong> match. <strong>No extra adjustment</strong> to the nearest. If there is no exact match, return the InChI that is most fidelitous to the input image. </p>\n<p>However, I guess this is not allowed. Current rule will only encourage more overfitting to the <code>molecular sketcher</code> and <code>noise generator</code> used in this competition. <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554</a> Broadly speaking, this is opposite to the general guidance: <code>method needs to generalize to new, unseen test data</code></p>",
          "rawMarkdown": "What if the model is NOT machine *trained*? It's a weird question in a data science competition, but it is the strategy we are using: generalized & automated molecular recognition using simple image processing, similar to OSRA. Given a crappy image (e.g., missing bonds due to bad compression), multiple chemically reasonable candidates can be proposed. A database look up helps pick a most probable one. We only do **exact** match. **No extra adjustment** to the nearest. If there is no exact match, return the InChI that is most fidelitous to the input image. \n\nHowever, I guess this is not allowed. Current rule will only encourage more overfitting to the `molecular sketcher` and `noise generator` used in this competition. https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554 Broadly speaking, this is opposite to the general guidance: `method needs to generalize to new, unseen test data`",
          "votes": 1
        },
        {
          "id": 1267548,
          "postDate": "2021-04-08T15:30:33.447Z",
          "content": "<p><a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a> As per above clarifications, participants can only use train+Extra_Approved_InChIs, but are <strong>not allowed to perform database lookups</strong> containing any other InChIs.</p>",
          "rawMarkdown": "@houndcl As per above clarifications, participants can only use train+Extra_Approved_InChIs, but are **not allowed to perform database lookups** containing any other InChIs."
        },
        {
          "id": 1267553,
          "postDate": "2021-04-08T15:33:01.217Z",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> What about randomly mixing parts of Inchis and then generating data. I am not sure if this makes sense, I am not in this competition.</p>",
          "rawMarkdown": "@inversion What about randomly mixing parts of Inchis and then generating data. I am not sure if this makes sense, I am not in this competition."
        },
        {
          "id": 1267559,
          "postDate": "2021-04-08T15:38:02.527Z",
          "content": "<p>To clarify - generated data (synthetic images, new molecules, etc) is allowed :)</p>",
          "rawMarkdown": "To clarify - generated data (synthetic images, new molecules, etc) is allowed :)",
          "votes": 4
        },
        {
          "id": 1267568,
          "postDate": "2021-04-08T15:44:33.200Z",
          "content": "<p>How about InChI Keys? :) I guess it's also forbidden, just want to make sure. </p>",
          "rawMarkdown": "How about InChI Keys? :) I guess it's also forbidden, just want to make sure. "
        }
      ]
    },
    {
      "id": 1267019,
      "postDate": "2021-04-08T08:50:30.347Z",
      "content": "<p>Sorry to be picky but can you explain this:</p>\n<blockquote>\n  <p>records of molecules specified by the provided training and supplemental InChIs</p>\n</blockquote>\n<p>I get what the training records are, but I don't get what are the records specified by supplemental InChis.</p>",
      "rawMarkdown": "Sorry to be picky but can you explain this:\n\n> records of molecules specified by the provided training and supplemental InChIs\n\nI get what the training records are, but I don't get what are the records specified by supplemental InChis.",
      "votes": 2,
      "replies": [
        {
          "id": 1267110,
          "postDate": "2021-04-08T10:24:09.347Z",
          "content": "<p>supplemental InChis = <a href=\"https://www.kaggle.com/c/bms-molecular-translation/data\" target=\"_blank\">extra_approved_InChIs.csv</a></p>\n<p>May be he meant that participants can use external information, say from PubChem, for these training+supplemental molecules. (Such as downloading \"clean\" images using <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/225876\" target=\"_blank\">REST API</a>)</p>",
          "rawMarkdown": "supplemental InChis = [extra_approved_InChIs.csv](https://www.kaggle.com/c/bms-molecular-translation/data)\n\nMay be he meant that participants can use external information, say from PubChem, for these training+supplemental molecules. (Such as downloading \"clean\" images using [REST API](https://www.kaggle.com/c/bms-molecular-translation/discussion/225876))"
        },
        {
          "id": 1267116,
          "postDate": "2021-04-08T10:26:27.823Z",
          "content": "<p>That's not what I am asking.  I ask what would be a molecule record for one of these extra InCHI.  For the training ones we have an image.</p>",
          "rawMarkdown": "That's not what I am asking.  I ask what would be a molecule record for one of these extra InCHI.  For the training ones we have an image.",
          "votes": 1
        },
        {
          "id": 1267279,
          "postDate": "2021-04-08T12:38:59.757Z",
          "content": "<p>Another way to say it - you can train you model on data corresponding to labels in in the training data and supplemental file (whether that data be an image, text, etc.). But you cannot train your models on data corresponding to labels that are not found in these lists.</p>\n<p>If that's still not clear, let us know!</p>",
          "rawMarkdown": "Another way to say it - you can train you model on data corresponding to labels in in the training data and supplemental file (whether that data be an image, text, etc.). But you cannot train your models on data corresponding to labels that are not found in these lists.\n\nIf that's still not clear, let us know!",
          "votes": 7
        },
        {
          "id": 1267283,
          "postDate": "2021-04-08T12:44:07.103Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1267284,
          "postDate": "2021-04-08T12:44:36.837Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1267349,
          "postDate": "2021-04-08T13:33:21.280Z",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> The general guidance is that any method needs to generalize to new, unseen test data. If the method won't work on new observations because of how it uses the test data, it's most likely not allowed.</p>",
          "rawMarkdown": "@drhabib The general guidance is that any method needs to generalize to new, unseen test data. If the method won't work on new observations because of how it uses the test data, it's most likely not allowed.",
          "votes": 3
        },
        {
          "id": 1267537,
          "postDate": "2021-04-08T15:22:20.323Z",
          "content": "<p>Thank you! <br>\nEDIT: for people who are interested my original question was (I accidentally deleted)<br>\n<code>weather its allowed to use test images in GAN or unsupervised training?</code></p>",
          "rawMarkdown": "Thank you! \nEDIT: for people who are interested my original question was (I accidentally deleted)\n`weather its allowed to use test images in GAN or unsupervised training?`",
          "votes": 1
        }
      ]
    },
    {
      "id": 1296448,
      "postDate": "2021-05-07T09:08:16.537Z",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>Hello! Can I use any label for training obtained by pseudo-labeling the test dataset, or only the ones specified in that file with external labels? If this is allowed, how many rounds can I do if model for each new round was trained with the PL labels from the previous round (without violating other rules and without going beyond the limits of the data available for use)?</p>",
      "rawMarkdown": "@addisonhoward @inversion \n\nHello! Can I use any label for training obtained by pseudo-labeling the test dataset, or only the ones specified in that file with external labels? If this is allowed, how many rounds can I do if model for each new round was trained with the PL labels from the previous round (without violating other rules and without going beyond the limits of the data available for use)?",
      "replies": [
        {
          "id": 1296624,
          "postDate": "2021-05-07T11:36:54.100Z",
          "content": "<p>the question was already answered .. </p>\n<pre><code>Pseudo-labelling isn't an issue as long as (a) the model used to pseudo-labelling is only trained on the allowed labels, and (b) you're pseudo-labelling the test data\nWhat is not allowed is pseudo-labelling that is effectively hand-labeling. For example, pseudo-labeling that predicts a test image, looks up that label from a database, adjusts it to a best match, etc. (Although there's no issue with adding a constraint in code that checks that the pseudo label is physically possible, e.g., has the correct number of bonds, etc.)\n</code></pre>\n<p>`<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267345\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267345</a></p>",
          "rawMarkdown": "the question was already answered .. \n\n```\nPseudo-labelling isn't an issue as long as (a) the model used to pseudo-labelling is only trained on the allowed labels, and (b) you're pseudo-labelling the test data\nWhat is not allowed is pseudo-labelling that is effectively hand-labeling. For example, pseudo-labeling that predicts a test image, looks up that label from a database, adjusts it to a best match, etc. (Although there's no issue with adding a constraint in code that checks that the pseudo label is physically possible, e.g., has the correct number of bonds, etc.)\n```\n\n`\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267345\n",
          "votes": 2
        },
        {
          "id": 1296745,
          "postDate": "2021-05-07T13:33:37.927Z",
          "content": "<p>Thank you for an answer, <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>. I read this message, however I still don't understand, is multi-round PL allowed on condition that model used for PL in each round was trained on PL labels from the previous ones?</p>",
          "rawMarkdown": "Thank you for an answer, @drhabib. I read this message, however I still don't understand, is multi-round PL allowed on condition that model used for PL in each round was trained on PL labels from the previous ones?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1267556,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-04-08T15:35:55.903000",
      "content": "<p>Thanks for the clarification.</p>\n<p>What I am generally a little bit bothered with is that these restrictions practically only apply to competition winners. I wonder how Kaggle can ensure that all competitors adhere to these rules and similar ones in other competitions. I know this is not trivial, but kind-of important to the integrity of competitions.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1267598,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-04-08T16:08:48.737000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> this is crystal clear now.  I guess I have no more reason to procrastinate here ;)</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1266697,
      "author_name": "Jacob Albrecht",
      "author_url": "",
      "post_date": "2021-04-08T02:13:34.747000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>!<br>\nWith the ultimate goal of this competition to develop a model capable of translating chemical images to a text representation, approaches that augment the training set with external data risk including test set compounds if they are in public databases.  Including these compounds would overfit the model, and reduce the ability to generalize to new chemical images.  Limiting the data with the compromise of supplemental InChIs will remove this risk, while also ensuring teams have freedom to develop innovative models and don't to waste the effort previously spent on developing code for synthetic data generation. </p>\n<p>Hopefully this makes the competition rules less ambiguous and ultimately increases the confidence in the capabilities of the winning solutions.  I want to thank the Kaggle team for their advice and effort over the past few days, and the community here for their work over the past month.  My expectations were exceeded within the first weeks of the competition, I am very impressed, and I look forward to the next two months!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1266954,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-04-08T07:40:00.900000",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> thank you both for a clear update on competition rules. </p>\n<p>I have no clue about this, but it could be that top of LB includes submissions that do not comply with the new rules.  IN that case the top of LB could be over optimistic  I wonder if it is feasible and desirable to remove these submissions, if they exist.  Of course this would rely on declaration by submitter.</p>\n<p>I wonder what others think about this.  If no one else is bothered by too optimistic LB then so be it, no big deal.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1267024,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-04-08T08:52:10.100000",
          "content": "<blockquote>\n  <p>don't waste the effort spent on developing code for synthetic data generation. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> Second thought.  Are you saying that generating new training data using training InCHI or supplemental InCHIs is not allowed?</p>\n<p>This would rule out common image augmentation at training time for instance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267106,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2021-04-08T10:23:12.813000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\n<code>it could be that top of LB includes submissions that do not comply with the new rules</code></p>\n<ul>\n<li>Organizers may not reset the LB, but <strong>Kaggle should definitely ensure</strong> that medals are not given to submissions made earlier, without following the new rule</li>\n</ul>\n<p><code>Are you saying that generating new training data using training InCHI or supplemental InCHIs is not allowed?</code></p>\n<ul>\n<li>I think he meant \"don't <strong>want to</strong> waste\" i.e. participants are encouraged to make use of synthetic data to build better models.</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1267431,
          "author_name": "Jacob Albrecht",
          "author_url": "",
          "post_date": "2021-04-08T14:06:39.963000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, I've updated the comment to clarify: generating synthetic images is ok, generating new molecules from a GAN trained on the supplied InChIs is ok.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1267546,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2021-04-08T15:30:02.903000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> - if any top submissions are inadvertently breaking the rules, those users can be sure to unselect those submissions, and only select submissions that are in compliance.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1291427,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2021-05-03T03:24:57.907000",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>Thanks for clarification. But let me double check below:</p>\n<ol>\n<li>It is disallowed to validate my prediction using rdkit and fix it if it's invalid. Is it true?<br>\nI mean like <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229578\" target=\"_blank\">this method</a>.</li>\n<li>Above method is disallowed even if it's for pseudo labeling, not for final submission. Is it true?</li>\n<li>But it is OK to write my own code to validate whether the InChI is valid one or not, unless it doesn't lookup database which has possible chemicals in it. Is it true?</li>\n</ol>\n<p>P.S. Please pin this topic on top of discussions. It is the most important notice from organizer but currently hard to find in the bunch of community topics..</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1291687,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-03T08:42:40.607000",
          "content": "<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/229279#1260313</a></p>\n<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> in his second sentence says:</p>\n<blockquote>\n  <p>Using open cheminformatics libraries is encouraged for the competition.</p>\n</blockquote>\n<p>But I also would like this to be verified to be sure!</p>\n<p>Edit: The link doesn't navigate me to the comment directly. ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1291706,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-05-03T09:13:35.937000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Thanks, I can find the comment even though it doesn't directly navigate:) And also thank you for the original normalization script, it works for me too! <br>\nLet's wait for the host's comment.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1291715,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-03T09:25:22.960000",
          "content": "<p>Cheers! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292491,
          "author_name": "Jacob Albrecht",
          "author_url": "",
          "post_date": "2021-05-04T02:47:51.613000",
          "content": "<p>Yes, checking using <code>rdkit</code> is a smart way to canonicalize the InChI.  Other packages like <code>pubchempy</code> and <code>chemspipy</code> can query their respective databases; avoid that functionality please, thanks!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1292521,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-05-04T03:53:49.747000",
          "content": "<p><a href=\"https://www.kaggle.com/jakealbrecht1337\" target=\"_blank\">@jakealbrecht1337</a> Thanks, it's crystal clear for me. Also thanks for hosting such a interesting competition, I'm really enjoying!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1267140,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "2021-04-08T10:44:36.913000",
      "content": "<p>Hi organizers &amp; Kaggle,</p>\n<p>Thank you for the update. Can you confirm that using data <em>generated without external data</em> is not subject to this restriction (for example, for a team that does not use any external data at all)</p>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1267271,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2021-04-08T12:34:18.410000",
          "content": "<p>Do you mean, e.g., image augmentation? That is fine.</p>\n<p>The intent is to prevent, e.g., (to use a hypothetical extreme) generating or downloading images of all known chemical structures and training on those images.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1267277,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-04-08T12:38:53.880000",
          "content": "<p>But is it fine to generate images from training InCHI or approved additional InCHI?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1267288,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2021-04-08T12:46:51.640000",
          "content": "<p>I mean e.g. pseudo-labelling would be an example</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1267345,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2021-04-08T13:31:00.070000",
          "content": "<p>The issue is strictly the label. You can create (download, etc) whatever training data you want for the allowed labels. The model shouldn't be trained on data outside that list list of labels. </p>\n<p>Pseudo-labelling isn't an issue as long as (a) the model used to pseudo-labelling is only trained on the allowed labels, and (b) you're pseudo-labelling the test data.</p>\n<p>What is not allowed is pseudo-labelling that is effectively hand-labeling. For example, pseudo-labeling that predicts a test image, looks up that label from a database, adjusts it to a best match, etc. (Although there's no issue with adding a constraint in code that checks that the pseudo label is physically possible, e.g., has the correct number of bonds, etc.).</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1267529,
          "author_name": "human intelligence",
          "author_url": "",
          "post_date": "2021-04-08T15:18:35.423000",
          "content": "<p>What if the model is NOT machine <em>trained</em>? It's a weird question in a data science competition, but it is the strategy we are using: generalized &amp; automated molecular recognition using simple image processing, similar to OSRA. Given a crappy image (e.g., missing bonds due to bad compression), multiple chemically reasonable candidates can be proposed. A database look up helps pick a most probable one. We only do <strong>exact</strong> match. <strong>No extra adjustment</strong> to the nearest. If there is no exact match, return the InChI that is most fidelitous to the input image. </p>\n<p>However, I guess this is not allowed. Current rule will only encourage more overfitting to the <code>molecular sketcher</code> and <code>noise generator</code> used in this competition. <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554</a> Broadly speaking, this is opposite to the general guidance: <code>method needs to generalize to new, unseen test data</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1267548,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2021-04-08T15:30:33.447000",
          "content": "<p><a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a> As per above clarifications, participants can only use train+Extra_Approved_InChIs, but are <strong>not allowed to perform database lookups</strong> containing any other InChIs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267553,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-04-08T15:33:01.217000",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> What about randomly mixing parts of Inchis and then generating data. I am not sure if this makes sense, I am not in this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267559,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2021-04-08T15:38:02.527000",
          "content": "<p>To clarify - generated data (synthetic images, new molecules, etc) is allowed :)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1267568,
          "author_name": "human intelligence",
          "author_url": "",
          "post_date": "2021-04-08T15:44:33.200000",
          "content": "<p>How about InChI Keys? :) I guess it's also forbidden, just want to make sure. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1267019,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-04-08T08:50:30.347000",
      "content": "<p>Sorry to be picky but can you explain this:</p>\n<blockquote>\n  <p>records of molecules specified by the provided training and supplemental InChIs</p>\n</blockquote>\n<p>I get what the training records are, but I don't get what are the records specified by supplemental InChis.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1267110,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2021-04-08T10:24:09.347000",
          "content": "<p>supplemental InChis = <a href=\"https://www.kaggle.com/c/bms-molecular-translation/data\" target=\"_blank\">extra_approved_InChIs.csv</a></p>\n<p>May be he meant that participants can use external information, say from PubChem, for these training+supplemental molecules. (Such as downloading \"clean\" images using <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/225876\" target=\"_blank\">REST API</a>)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267116,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-04-08T10:26:27.823000",
          "content": "<p>That's not what I am asking.  I ask what would be a molecule record for one of these extra InCHI.  For the training ones we have an image.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1267279,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2021-04-08T12:38:59.757000",
          "content": "<p>Another way to say it - you can train you model on data corresponding to labels in in the training data and supplemental file (whether that data be an image, text, etc.). But you cannot train your models on data corresponding to labels that are not found in these lists.</p>\n<p>If that's still not clear, let us know!</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1267283,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-08T12:44:07.103000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267284,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-08T12:44:36.837000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1267349,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2021-04-08T13:33:21.280000",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> The general guidance is that any method needs to generalize to new, unseen test data. If the method won't work on new observations because of how it uses the test data, it's most likely not allowed.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1267537,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-04-08T15:22:20.323000",
          "content": "<p>Thank you! <br>\nEDIT: for people who are interested my original question was (I accidentally deleted)<br>\n<code>weather its allowed to use test images in GAN or unsupervised training?</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1296448,
      "author_name": "GG",
      "author_url": "",
      "post_date": "2021-05-07T09:08:16.537000",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>Hello! Can I use any label for training obtained by pseudo-labeling the test dataset, or only the ones specified in that file with external labels? If this is allowed, how many rounds can I do if model for each new round was trained with the PL labels from the previous round (without violating other rules and without going beyond the limits of the data available for use)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1296624,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-05-07T11:36:54.100000",
          "content": "<p>the question was already answered .. </p>\n<pre><code>Pseudo-labelling isn't an issue as long as (a) the model used to pseudo-labelling is only trained on the allowed labels, and (b) you're pseudo-labelling the test data\nWhat is not allowed is pseudo-labelling that is effectively hand-labeling. For example, pseudo-labeling that predicts a test image, looks up that label from a database, adjusts it to a best match, etc. (Although there's no issue with adding a constraint in code that checks that the pseudo label is physically possible, e.g., has the correct number of bonds, etc.)\n</code></pre>\n<p>`<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267345\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267345</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1296745,
          "author_name": "GG",
          "author_url": "",
          "post_date": "2021-05-07T13:33:37.927000",
          "content": "<p>Thank you for an answer, <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>. I read this message, however I still don't understand, is multi-round PL allowed on condition that model used for PL in each round was trained on PL labels from the previous ones?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1266681": "Hi All,\n\nIn working with the BMS team, they have decided to limit any external chemical labels to only those provided in the competition dataset. A whitelist of those allowed (approx 10M) are now included on the Data tab as `Extra_Approved_InChIs.csv.gz` This limitation is specifically for chemical labels and also does not prohibit the use of any generated of data (i.e. GANs).\n\nThe competition rules have been updated for avoidance of doubt with the following language:\n\nThe following supersedes Section B.7.C below:\n\nC. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. In addition, the only external chemical labels available for use are limited to records of molecules specified by the provided training and supplemental InChIs. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).",
    "1267556": "Thanks for the clarification.\n\nWhat I am generally a little bit bothered with is that these restrictions practically only apply to competition winners. I wonder how Kaggle can ensure that all competitors adhere to these rules and similar ones in other competitions. I know this is not trivial, but kind-of important to the integrity of competitions.",
    "1267598": "Thanks @addisonhoward @jakealbrecht1337 @inversion this is crystal clear now.  I guess I have no more reason to procrastinate here ;)",
    "1266697": "Thanks @addisonhoward!\nWith the ultimate goal of this competition to develop a model capable of translating chemical images to a text representation, approaches that augment the training set with external data risk including test set compounds if they are in public databases.  Including these compounds would overfit the model, and reduce the ability to generalize to new chemical images.  Limiting the data with the compromise of supplemental InChIs will remove this risk, while also ensuring teams have freedom to develop innovative models and don't to waste the effort previously spent on developing code for synthetic data generation. \n\nHopefully this makes the competition rules less ambiguous and ultimately increases the confidence in the capabilities of the winning solutions.  I want to thank the Kaggle team for their advice and effort over the past few days, and the community here for their work over the past month.  My expectations were exceeded within the first weeks of the competition, I am very impressed, and I look forward to the next two months!",
    "1291427": "@addisonhoward @inversion \n\nThanks for clarification. But let me double check below:\n1. It is disallowed to validate my prediction using rdkit and fix it if it's invalid. Is it true?\n    I mean like [this method](https://www.kaggle.com/c/bms-molecular-translation/discussion/229578).\n2. Above method is disallowed even if it's for pseudo labeling, not for final submission. Is it true?\n3. But it is OK to write my own code to validate whether the InChI is valid one or not, unless it doesn't lookup database which has possible chemicals in it. Is it true?\n\nP.S. Please pin this topic on top of discussions. It is the most important notice from organizer but currently hard to find in the bunch of community topics..\n",
    "1267140": "Hi organizers & Kaggle,\n\nThank you for the update. Can you confirm that using data _generated without external data_ is not subject to this restriction (for example, for a team that does not use any external data at all)\n\nThanks!",
    "1267019": "Sorry to be picky but can you explain this:\n\n> records of molecules specified by the provided training and supplemental InChIs\n\nI get what the training records are, but I don't get what are the records specified by supplemental InChis.",
    "1296448": "@addisonhoward @inversion \n\nHello! Can I use any label for training obtained by pseudo-labeling the test dataset, or only the ones specified in that file with external labels? If this is allowed, how many rounds can I do if model for each new round was trained with the PL labels from the previous round (without violating other rules and without going beyond the limits of the data available for use)?"
  }
}