{
  "id": 225590,
  "title": "Concerns about the BMS dataset",
  "url": "/competitions/bms-molecular-translation/discussion/225590",
  "author_name": "",
  "post_date": "2021-03-13T02:58:15.755560900Z",
  "votes": 42,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Being a bit of an organic chemistry buff, this competition has really captured my imagination. However, it looks like a lot of work to do properly, and I have some doubts about how effective the dataset will be to reach the competition's desired objective.</p>\n<p>To see what I mean, do an image Google search for: [molecular structure diagram site:researchgate.net], and look at the image results. How well would the training set prepare our models for real-world diagrams?</p>\n<p>As far as I can see, the competition's dataset diagrams are all produced the same way—presumably by the same tool.</p>\n<ul>\n<li>The layouts and bond angles are entirely consistent across all diagrams.</li>\n<li>The structures are fully skeletonised: eg. every implied H is hidden, no \"CH₃\", \"COOH\",… units that I could see, etc.</li>\n<li>The same font is used.</li>\n<li>Other than the simulated background noise, the structures are clean: no annotations, lines, arrows, labels, icons, circled or bordered sub-structures, etc.</li>\n<li>Only 12 elements appear in the training set [refer to my notebook: <a href=\"https://www.kaggle.com/stainsby/bristol-myers-squibb-counting-elements\" target=\"_blank\">https://www.kaggle.com/stainsby/bristol-myers-squibb-counting-elements</a> ].</li>\n<li>It appears that the same simulated noise algorithms is used throughout, with varying intensity. There is also some scale variation.</li>\n</ul>\n<p>I suspect that any models designed just to <em>win</em> the competition will be too biased—not just in training but in the algorithms chosen by the designers—to work for the majority of real-world data.</p>\n<p>Clearly producing a real-world dataset of a similar size would be a huge task. I would think though that it would be possible to improve the synthetic dataset production process in several small ways to create a big difference in model robustness in real-world scenarios.</p>\n<p>I would also be interested in seeing the results of some of the models produced so far applied to a few real-world cases.</p>",
  "messages": [
    {
      "id": "1236336",
      "postDate": "03/13/2021 02:58:15",
      "content": "<p>Being a bit of an organic chemistry buff, this competition has really captured my imagination. However, it looks like a lot of work to do properly, and I have some doubts about how effective the dataset will be to reach the competition's desired objective.</p>\n<p>To see what I mean, do an image Google search for: [molecular structure diagram site:researchgate.net], and look at the image results. How well would the training set prepare our models for real-world diagrams?</p>\n<p>As far as I can see, the competition's dataset diagrams are all produced the same way—presumably by the same tool.</p>\n<ul>\n<li>The layouts and bond angles are entirely consistent across all diagrams.</li>\n<li>The structures are fully skeletonised: eg. every implied H is hidden, no \"CH₃\", \"COOH\",… units that I could see, etc.</li>\n<li>The same font is used.</li>\n<li>Other than the simulated background noise, the structures are clean: no annotations, lines, arrows, labels, icons, circled or bordered sub-structures, etc.</li>\n<li>Only 12 elements appear in the training set [refer to my notebook: <a href=\"https://www.kaggle.com/stainsby/bristol-myers-squibb-counting-elements\" target=\"_blank\">https://www.kaggle.com/stainsby/bristol-myers-squibb-counting-elements</a> ].</li>\n<li>It appears that the same simulated noise algorithms is used throughout, with varying intensity. There is also some scale variation.</li>\n</ul>\n<p>I suspect that any models designed just to <em>win</em> the competition will be too biased—not just in training but in the algorithms chosen by the designers—to work for the majority of real-world data.</p>\n<p>Clearly producing a real-world dataset of a similar size would be a huge task. I would think though that it would be possible to improve the synthetic dataset production process in several small ways to create a big difference in model robustness in real-world scenarios.</p>\n<p>I would also be interested in seeing the results of some of the models produced so far applied to a few real-world cases.</p>",
      "rawMarkdown": "Being a bit of an organic chemistry buff, this competition has really captured my imagination. However, it looks like a lot of work to do properly, and I have some doubts about how effective the dataset will be to reach the competition's desired objective.\n\nTo see what I mean, do an image Google search for: [molecular structure diagram site:researchgate.net], and look at the image results. How well would the training set prepare our models for real-world diagrams?\n\nAs far as I can see, the competition's dataset diagrams are all produced the same way—presumably by the same tool.\n  - The layouts and bond angles are entirely consistent across all diagrams.\n  - The structures are fully skeletonised: eg. every implied H is hidden, no \"CH₃\", \"COOH\",… units that I could see, etc.\n  - The same font is used.\n  - Other than the simulated background noise, the structures are clean: no annotations, lines, arrows, labels, icons, circled or bordered sub-structures, etc.\n  - Only 12 elements appear in the training set [refer to my notebook: https://www.kaggle.com/stainsby/bristol-myers-squibb-counting-elements ].\n  - It appears that the same simulated noise algorithms is used throughout, with varying intensity. There is also some scale variation.\n  \nI suspect that any models designed just to *win* the competition will be too biased—not just in training but in the algorithms chosen by the designers—to work for the majority of real-world data.\n\nClearly producing a real-world dataset of a similar size would be a huge task. I would think though that it would be possible to improve the synthetic dataset production process in several small ways to create a big difference in model robustness in real-world scenarios.\n\nI would also be interested in seeing the results of some of the models produced so far applied to a few real-world cases.",
      "votes": null
    },
    {
      "id": "1236357",
      "postDate": "03/13/2021 03:51:30",
      "content": "<p>Great discussion post! I agree with pretty much everything you said. The winning solutions aren't going to generalize to real-world examples and will be highly biased towards the test data.</p>\n<p>However, I do still think there are some valid challenges that make this competition worthwhile (at least for BMS)</p>",
      "rawMarkdown": "Great discussion post! I agree with pretty much everything you said. The winning solutions aren't going to generalize to real-world examples and will be highly biased towards the test data.\n\nHowever, I do still think there are some valid challenges that make this competition worthwhile (at least for BMS)",
      "votes": null
    },
    {
      "id": "1236477",
      "postDate": "03/13/2021 07:12:18",
      "content": "<p>From the competition's description, the second paragraph.</p>\n<blockquote>\n  <p>Unfortunately, most public data sets are too small to support modern machine learning models. Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. In these cases, time-consuming, manual work is required to reliably convert scanned chemical structure images into a machine-readable format.</p>\n</blockquote>\n<p>From this, I think that the algorithms developed here will be a step forward for them.</p>",
      "rawMarkdown": "From the competition's description, the second paragraph.\n\n> Unfortunately, most public data sets are too small to support modern machine learning models. Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. In these cases, time-consuming, manual work is required to reliably convert scanned chemical structure images into a machine-readable format.\n\nFrom this, I think that the algorithms developed here will be a step forward for them.",
      "votes": null
    },
    {
      "id": "1236483",
      "postDate": "03/13/2021 07:16:46",
      "content": "<p>Agree. One additional observation. I tried Imago, which an open source OCSR software, supposedly pretty good. It fell flat on its face, even when fed with the best quality samples from BMS dataset.</p>",
      "rawMarkdown": "Agree. One additional observation. I tried Imago, which an open source OCSR software, supposedly pretty good. It fell flat on its face, even when fed with the best quality samples from BMS dataset.",
      "votes": null
    },
    {
      "id": "1237264",
      "postDate": "03/14/2021 00:43:37",
      "content": "<p>One further quirk of the dataset is that all the compounds are single molecules; there are no compounds that are salts.</p>\n<p>A quirk of this is that if you're using a tool that is not explicitly trained for this task (e.g. Imago, MolVec, OSRA), when these predict salt-like structures you can know with certainty that these are not the expected result. Given that the simulated noise frequently results in bonds being missing, at least MolVec, frequently predicted such structures. If you want to game the metric you can then replace such cases with a more \"typical\" InChI. Disconnected components are indicated with a dot (.) in the InChI molecular formula layer… and the training data doesn't contain any InChIs that have a dot character.</p>",
      "rawMarkdown": "One further quirk of the dataset is that all the compounds are single molecules; there are no compounds that are salts.\n\nA quirk of this is that if you're using a tool that is not explicitly trained for this task (e.g. Imago, MolVec, OSRA), when these predict salt-like structures you can know with certainty that these are not the expected result. Given that the simulated noise frequently results in bonds being missing, at least MolVec, frequently predicted such structures. If you want to game the metric you can then replace such cases with a more \"typical\" InChI. Disconnected components are indicated with a dot (.) in the InChI molecular formula layer... and the training data doesn't contain any InChIs that have a dot character.",
      "votes": null
    },
    {
      "id": "1239461",
      "postDate": "03/15/2021 18:29:20",
      "content": "<p>Great points!  For a closer to reality application, there is a recent preprint describing the prediction of <a href=\"https://doi.org/10.26434/chemrxiv.14156957.v1\" target=\"_blank\">hand-drawn structures </a> that demonstrated good performance, though for a much less diverse set of molecules.  With this competition we are targeting a situation where structures were created and scanned within a single research organization, hence the standard fonts and style.  As more of these models are developed for their various use cases, my hope is that an ensemble of the best tools and approaches can achieve human-level performance for a wider range of applications.</p>",
      "rawMarkdown": "Great points!  For a closer to reality application, there is a recent preprint describing the prediction of [hand-drawn structures ](https://doi.org/10.26434/chemrxiv.14156957.v1) that demonstrated good performance, though for a much less diverse set of molecules.  With this competition we are targeting a situation where structures were created and scanned within a single research organization, hence the standard fonts and style.  As more of these models are developed for their various use cases, my hope is that an ensemble of the best tools and approaches can achieve human-level performance for a wider range of applications.",
      "votes": null
    },
    {
      "id": "1239472",
      "postDate": "03/15/2021 18:40:14",
      "content": "<ul>\n<li><p>Are we allowed to use additional data from PubChem? Like generating images with rdkit. Or did you get the compounds from there, so we should avoid using it?</p></li>\n<li><p>Are we allowed to pre-train on the test images with some kind of self-supervision, for example SimCLR?</p></li>\n</ul>",
      "rawMarkdown": "Are we allowed to use additional data from PubChem? Like generating images with rdkit. Or did you get the compounds from there, so we should avoid using it?\n\n- Are we allowed to pre-train on the test images with some kind of self-supervision, for example SimCLR?",
      "votes": null
    },
    {
      "id": "1239692",
      "postDate": "03/15/2021 23:36:35",
      "content": "<p>Thanks Jacob. That's good to know. I finalising a notebook today that generates less regular synthetic data using RDKIt, though it sounds like it's not needed for the current competition. I'll publish it a bit later on hopefully.</p>",
      "rawMarkdown": "Thanks Jacob. That's good to know. I finalising a notebook today that generates less regular synthetic data using RDKIt, though it sounds like it's not needed for the current competition. I'll publish it a bit later on hopefully.",
      "votes": null
    },
    {
      "id": "1240224",
      "postDate": "03/16/2021 09:52:50",
      "content": "<p>…and here it is: <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3</a></p>",
      "rawMarkdown": "…and here it is: https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3",
      "votes": null
    },
    {
      "id": "1247678",
      "postDate": "03/22/2021 00:32:37",
      "content": "<p>There's an updated/improved version of my synthetic data notebook for those still interested.</p>",
      "rawMarkdown": "There's an updated/improved version of my synthetic data notebook for those still interested.",
      "votes": null
    },
    {
      "id": "1247818",
      "postDate": "03/22/2021 04:53:49",
      "content": "<p>Great <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v2\" target=\"_blank\">notebook</a>. Your images are even worse than in the original dataset ) … in a good way. It is not clear what happens with \" Molecule # 3: 604a8299c132: InChI = 1S / C9H10O2 …\". With redkit, it is displayed differently. I wonder if this is a bug in train? Or is it just a second way to show this molecule (when the H atom is shown)?</p>",
      "rawMarkdown": "Great [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v2). Your images are even worse than in the original dataset ) ... in a good way. It is not clear what happens with \" Molecule # 3: 604a8299c132: InChI = 1S / C9H10O2 ...\". With redkit, it is displayed differently. I wonder if this is a bug in train? Or is it just a second way to show this molecule (when the H atom is shown)?",
      "votes": null
    },
    {
      "id": "1248005",
      "postDate": "03/22/2021 09:01:27",
      "content": "<p>If you look closer, you can hardly find any generated molecules that match the training ones. Yes, it is just another way of drawing the molecule.</p>",
      "rawMarkdown": "If you look closer, you can hardly find any generated molecules that match the training ones. Yes, it is just another way of drawing the molecule.",
      "votes": null
    },
    {
      "id": "1248013",
      "postDate": "03/22/2021 09:14:18",
      "content": "<p>Yes, but I did not mean differences in the arrangement or orientation of the atoms in space. I'm talking about the difference in the bonds and the displayed atoms. </p>\n<p>For example ,in # 3:  <br>\n   a) 'in train set' - shows an atom <strong>H</strong> and a triangular bond to it; <br>\n   b) 'in rdkit' - H is not shown and the triangular bond goes to the vertex (C) that is connected to O. </p>\n<p>There is a feeling that these are different molecules. Or is it still the same? (I don't have enough chemistry knowledge to understand this..)</p>\n<p>In all other examples, I do not see such differences.</p>",
      "rawMarkdown": "Yes, but I did not mean differences in the arrangement or orientation of the atoms in space. I'm talking about the difference in the bonds and the displayed atoms. \n\nFor example ,in # 3:  \n   a) 'in train set' - shows an atom **H** and a triangular bond to it; \n   b) 'in rdkit' - H is not shown and the triangular bond goes to the vertex (C) that is connected to O. \n\nThere is a feeling that these are different molecules. Or is it still the same? (I don't have enough chemistry knowledge to understand this..)\n\nIn all other examples, I do not see such differences.",
      "votes": null
    },
    {
      "id": "1248027",
      "postDate": "03/22/2021 09:23:56",
      "content": "<p>Yes, I've noticed that there are differences in how chirality is shown. I don't understand the concepts enough to add much more on that. Also some atom groups are shown differently. The choices of bond angles also varies in some cases.</p>",
      "rawMarkdown": "Yes, I've noticed that there are differences in how chirality is shown. I don't understand the concepts enough to add much more on that. Also some atom groups are shown differently. The choices of bond angles also varies in some cases.",
      "votes": null
    },
    {
      "id": "1248040",
      "postDate": "03/22/2021 09:35:50",
      "content": "<p>Oh, so I'm the one who should look closer :D<br>\n#3 is the same molecule, it is just that in the training image the extra H is not shown. I think that there are options in rdkit to draw them or not, but I may be wrong on that one (, too :D ).</p>",
      "rawMarkdown": "Oh, so I'm the one who should look closer :D\n\\#3 is the same molecule, it is just that in the training image the extra H is not shown. I think that there are options in rdkit to draw them or not, but I may be wrong on that one (, too :D ).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1236357,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "03/13/2021 03:51:30",
      "content": "<p>Great discussion post! I agree with pretty much everything you said. The winning solutions aren't going to generalize to real-world examples and will be highly biased towards the test data.</p>\n<p>However, I do still think there are some valid challenges that make this competition worthwhile (at least for BMS)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1236477,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "03/13/2021 07:12:18",
      "content": "<p>From the competition's description, the second paragraph.</p>\n<blockquote>\n  <p>Unfortunately, most public data sets are too small to support modern machine learning models. Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. In these cases, time-consuming, manual work is required to reliably convert scanned chemical structure images into a machine-readable format.</p>\n</blockquote>\n<p>From this, I think that the algorithms developed here will be a step forward for them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1236483,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "03/13/2021 07:16:46",
      "content": "<p>Agree. One additional observation. I tried Imago, which an open source OCSR software, supposedly pretty good. It fell flat on its face, even when fed with the best quality samples from BMS dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1237264,
      "author_name": "infy2097",
      "author_url": "",
      "post_date": "03/14/2021 00:43:37",
      "content": "<p>One further quirk of the dataset is that all the compounds are single molecules; there are no compounds that are salts.</p>\n<p>A quirk of this is that if you're using a tool that is not explicitly trained for this task (e.g. Imago, MolVec, OSRA), when these predict salt-like structures you can know with certainty that these are not the expected result. Given that the simulated noise frequently results in bonds being missing, at least MolVec, frequently predicted such structures. If you want to game the metric you can then replace such cases with a more \"typical\" InChI. Disconnected components are indicated with a dot (.) in the InChI molecular formula layer… and the training data doesn't contain any InChIs that have a dot character.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1239461,
      "author_name": "jakealbrecht1337",
      "author_url": "",
      "post_date": "03/15/2021 18:29:20",
      "content": "<p>Great points!  For a closer to reality application, there is a recent preprint describing the prediction of <a href=\"https://doi.org/10.26434/chemrxiv.14156957.v1\" target=\"_blank\">hand-drawn structures </a> that demonstrated good performance, though for a much less diverse set of molecules.  With this competition we are targeting a situation where structures were created and scanned within a single research organization, hence the standard fonts and style.  As more of these models are developed for their various use cases, my hope is that an ensemble of the best tools and approaches can achieve human-level performance for a wider range of applications.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239472,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "03/15/2021 18:40:14",
          "content": "<ul>\n<li><p>Are we allowed to use additional data from PubChem? Like generating images with rdkit. Or did you get the compounds from there, so we should avoid using it?</p></li>\n<li><p>Are we allowed to pre-train on the test images with some kind of self-supervision, for example SimCLR?</p></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239692,
          "author_name": "stainsby",
          "author_url": "",
          "post_date": "03/15/2021 23:36:35",
          "content": "<p>Thanks Jacob. That's good to know. I finalising a notebook today that generates less regular synthetic data using RDKIt, though it sounds like it's not needed for the current competition. I'll publish it a bit later on hopefully.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240224,
          "author_name": "stainsby",
          "author_url": "",
          "post_date": "03/16/2021 09:52:50",
          "content": "<p>…and here it is: <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1247678,
          "author_name": "stainsby",
          "author_url": "",
          "post_date": "03/22/2021 00:32:37",
          "content": "<p>There's an updated/improved version of my synthetic data notebook for those still interested.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1247818,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "03/22/2021 04:53:49",
      "content": "<p>Great <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v2\" target=\"_blank\">notebook</a>. Your images are even worse than in the original dataset ) … in a good way. It is not clear what happens with \" Molecule # 3: 604a8299c132: InChI = 1S / C9H10O2 …\". With redkit, it is displayed differently. I wonder if this is a bug in train? Or is it just a second way to show this molecule (when the H atom is shown)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1248005,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "03/22/2021 09:01:27",
          "content": "<p>If you look closer, you can hardly find any generated molecules that match the training ones. Yes, it is just another way of drawing the molecule.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1248013,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "03/22/2021 09:14:18",
          "content": "<p>Yes, but I did not mean differences in the arrangement or orientation of the atoms in space. I'm talking about the difference in the bonds and the displayed atoms. </p>\n<p>For example ,in # 3:  <br>\n   a) 'in train set' - shows an atom <strong>H</strong> and a triangular bond to it; <br>\n   b) 'in rdkit' - H is not shown and the triangular bond goes to the vertex (C) that is connected to O. </p>\n<p>There is a feeling that these are different molecules. Or is it still the same? (I don't have enough chemistry knowledge to understand this..)</p>\n<p>In all other examples, I do not see such differences.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1248027,
          "author_name": "stainsby",
          "author_url": "",
          "post_date": "03/22/2021 09:23:56",
          "content": "<p>Yes, I've noticed that there are differences in how chirality is shown. I don't understand the concepts enough to add much more on that. Also some atom groups are shown differently. The choices of bond angles also varies in some cases.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1248040,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "03/22/2021 09:35:50",
          "content": "<p>Oh, so I'm the one who should look closer :D<br>\n#3 is the same molecule, it is just that in the training image the extra H is not shown. I think that there are options in rdkit to draw them or not, but I may be wrong on that one (, too :D ).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1236336": "Being a bit of an organic chemistry buff, this competition has really captured my imagination. However, it looks like a lot of work to do properly, and I have some doubts about how effective the dataset will be to reach the competition's desired objective.\n\nTo see what I mean, do an image Google search for: [molecular structure diagram site:researchgate.net], and look at the image results. How well would the training set prepare our models for real-world diagrams?\n\nAs far as I can see, the competition's dataset diagrams are all produced the same way—presumably by the same tool.\n  - The layouts and bond angles are entirely consistent across all diagrams.\n  - The structures are fully skeletonised: eg. every implied H is hidden, no \"CH₃\", \"COOH\",… units that I could see, etc.\n  - The same font is used.\n  - Other than the simulated background noise, the structures are clean: no annotations, lines, arrows, labels, icons, circled or bordered sub-structures, etc.\n  - Only 12 elements appear in the training set [refer to my notebook: https://www.kaggle.com/stainsby/bristol-myers-squibb-counting-elements ].\n  - It appears that the same simulated noise algorithms is used throughout, with varying intensity. There is also some scale variation.\n  \nI suspect that any models designed just to *win* the competition will be too biased—not just in training but in the algorithms chosen by the designers—to work for the majority of real-world data.\n\nClearly producing a real-world dataset of a similar size would be a huge task. I would think though that it would be possible to improve the synthetic dataset production process in several small ways to create a big difference in model robustness in real-world scenarios.\n\nI would also be interested in seeing the results of some of the models produced so far applied to a few real-world cases.",
    "1236357": "Great discussion post! I agree with pretty much everything you said. The winning solutions aren't going to generalize to real-world examples and will be highly biased towards the test data.\n\nHowever, I do still think there are some valid challenges that make this competition worthwhile (at least for BMS)",
    "1236477": "From the competition's description, the second paragraph.\n\n> Unfortunately, most public data sets are too small to support modern machine learning models. Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. In these cases, time-consuming, manual work is required to reliably convert scanned chemical structure images into a machine-readable format.\n\nFrom this, I think that the algorithms developed here will be a step forward for them.",
    "1236483": "Agree. One additional observation. I tried Imago, which an open source OCSR software, supposedly pretty good. It fell flat on its face, even when fed with the best quality samples from BMS dataset.",
    "1237264": "One further quirk of the dataset is that all the compounds are single molecules; there are no compounds that are salts.\n\nA quirk of this is that if you're using a tool that is not explicitly trained for this task (e.g. Imago, MolVec, OSRA), when these predict salt-like structures you can know with certainty that these are not the expected result. Given that the simulated noise frequently results in bonds being missing, at least MolVec, frequently predicted such structures. If you want to game the metric you can then replace such cases with a more \"typical\" InChI. Disconnected components are indicated with a dot (.) in the InChI molecular formula layer... and the training data doesn't contain any InChIs that have a dot character.",
    "1239461": "Great points!  For a closer to reality application, there is a recent preprint describing the prediction of [hand-drawn structures ](https://doi.org/10.26434/chemrxiv.14156957.v1) that demonstrated good performance, though for a much less diverse set of molecules.  With this competition we are targeting a situation where structures were created and scanned within a single research organization, hence the standard fonts and style.  As more of these models are developed for their various use cases, my hope is that an ensemble of the best tools and approaches can achieve human-level performance for a wider range of applications.",
    "1239472": "Are we allowed to use additional data from PubChem? Like generating images with rdkit. Or did you get the compounds from there, so we should avoid using it?\n\n- Are we allowed to pre-train on the test images with some kind of self-supervision, for example SimCLR?",
    "1239692": "Thanks Jacob. That's good to know. I finalising a notebook today that generates less regular synthetic data using RDKIt, though it sounds like it's not needed for the current competition. I'll publish it a bit later on hopefully.",
    "1240224": "…and here it is: https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3",
    "1247678": "There's an updated/improved version of my synthetic data notebook for those still interested.",
    "1247818": "Great [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v2). Your images are even worse than in the original dataset ) ... in a good way. It is not clear what happens with \" Molecule # 3: 604a8299c132: InChI = 1S / C9H10O2 ...\". With redkit, it is displayed differently. I wonder if this is a bug in train? Or is it just a second way to show this molecule (when the H atom is shown)?",
    "1248005": "If you look closer, you can hardly find any generated molecules that match the training ones. Yes, it is just another way of drawing the molecule.",
    "1248013": "Yes, but I did not mean differences in the arrangement or orientation of the atoms in space. I'm talking about the difference in the bonds and the displayed atoms. \n\nFor example ,in # 3:  \n   a) 'in train set' - shows an atom **H** and a triangular bond to it; \n   b) 'in rdkit' - H is not shown and the triangular bond goes to the vertex (C) that is connected to O. \n\nThere is a feeling that these are different molecules. Or is it still the same? (I don't have enough chemistry knowledge to understand this..)\n\nIn all other examples, I do not see such differences.",
    "1248027": "Yes, I've noticed that there are differences in how chirality is shown. I don't understand the concepts enough to add much more on that. Also some atom groups are shown differently. The choices of bond angles also varies in some cases.",
    "1248040": "Oh, so I'm the one who should look closer :D\n\\#3 is the same molecule, it is just that in the training image the extra H is not shown. I think that there are options in rdkit to draw them or not, but I may be wrong on that one (, too :D )."
  },
  "source": "meta"
}