{
  "id": 224275,
  "title": "Attention: Errors in the dataset.",
  "url": "/competitions/bms-molecular-translation/discussion/224275",
  "author_name": "",
  "post_date": "2021-03-07T16:10:19.765728200Z",
  "votes": 8,
  "comment_count": 11,
  "views": 0,
  "content": "<p>The organizers may have used scaling while preparing the images.<br>\nAt the same time, some lines of chemical bonds disappeared. This is very bad.  😭😭😭 Some symbols of chemical elements were also distorted. An example is 00abb3b349a.png. I did not find such defects in the images at large scale.<br>\nPlease correct the dataset. </p>",
  "messages": [
    {
      "id": "1229837",
      "postDate": "03/07/2021 16:10:19",
      "content": "<p>The organizers may have used scaling while preparing the images.<br>\nAt the same time, some lines of chemical bonds disappeared. This is very bad.  😭😭😭 Some symbols of chemical elements were also distorted. An example is 00abb3b349a.png. I did not find such defects in the images at large scale.<br>\nPlease correct the dataset. </p>",
      "rawMarkdown": "The organizers may have used scaling while preparing the images.\nAt the same time, some lines of chemical bonds disappeared. This is very bad.  😭😭😭 Some symbols of chemical elements were also distorted. An example is 00abb3b349a.png. I did not find such defects in the images at large scale.\nPlease correct the dataset.",
      "votes": null
    },
    {
      "id": "1229861",
      "postDate": "03/07/2021 16:35:37",
      "content": "<p>I think this is the challenge…=)</p>\n<p>from the description…<br>\n<code>but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts.</code></p>\n<p>If you have tons of old paper or notes with bunch of chemical formulas its impossible to go back and correct all of them one by one. </p>",
      "rawMarkdown": "I think this is the challenge...=)\n\nfrom the description...\n`but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts.`\n\n If you have tons of old paper or notes with bunch of chemical formulas its impossible to go back and correct all of them one by one.",
      "votes": null
    },
    {
      "id": "1229885",
      "postDate": "03/07/2021 16:54:34",
      "content": "<p>This is not a challenge. This is an image processing error, not a paper or scan defect. This destroys chemical information that would have been preserved during scanning. </p>",
      "rawMarkdown": "This is not a challenge. This is an image processing error, not a paper or scan defect. This destroys chemical information that would have been preserved during scanning.",
      "votes": null
    },
    {
      "id": "1230002",
      "postDate": "03/07/2021 18:31:06",
      "content": "<p>This IS the challenge though. All those artifacts you talk about are intentionally added to the dataset, including missing bonds, occluded symbols, and the resizing. Even in the example you point out, a human expert could easily complete this molecule like so:<br>\n<img src=\"https://i.imgur.com/o84tDjB.png\" alt=\"\"></p>\n<p>Providing high definition images of molecules with no augmentation would make the challenge too easy and the solutions would not generalize to images of lower quality.</p>",
      "rawMarkdown": "This IS the challenge though. All those artifacts you talk about are intentionally added to the dataset, including missing bonds, occluded symbols, and the resizing. Even in the example you point out, a human expert could easily complete this molecule like so:\n![](https://i.imgur.com/o84tDjB.png)\n\nProviding high definition images of molecules with no augmentation would make the challenge too easy and the solutions would not generalize to images of lower quality.",
      "votes": null
    },
    {
      "id": "1230003",
      "postDate": "03/07/2021 18:31:11",
      "content": "<p>As far as I understand, the data is synthetic. Therefore, all these difficulties are artificial. And the task of the model (in my understanding) is to restore the data as completely as possible. If the 'line of chemical bond' is not visible, where it should be-the model should put it.</p>",
      "rawMarkdown": "As far as I understand, the data is synthetic. Therefore, all these difficulties are artificial. And the task of the model (in my understanding) is to restore the data as completely as possible. If the 'line of chemical bond' is not visible, where it should be-the model should put it.",
      "votes": null
    },
    {
      "id": "1230196",
      "postDate": "03/07/2021 21:33:47",
      "content": "<blockquote>\n  <p>the solutions would not generalize to images of lower quality</p>\n</blockquote>\n<p>Inverse is also true. When the training image quality is artificially lowered too far, the solution may not generalize to the real world image quality range. The MNIST data set is an example of this phenomenon. </p>",
      "rawMarkdown": "> the solutions would not generalize to images of lower quality\n\nInverse is also true. When the training image quality is artificially lowered too far, the solution may not generalize to the real world image quality range. The MNIST data set is an example of this phenomenon.",
      "votes": null
    },
    {
      "id": "1230797",
      "postDate": "03/08/2021 13:13:46",
      "content": "<p>I believe the hosts should clear the expectations now that people are starting to get geared up for the competitions by answering questions of this sort.</p>",
      "rawMarkdown": "I believe the hosts should clear the expectations now that people are starting to get geared up for the competitions by answering questions of this sort.",
      "votes": null
    },
    {
      "id": "1230947",
      "postDate": "03/08/2021 15:01:47",
      "content": "<p><a href=\"https://www.kaggle.com/arka47\" target=\"_blank\">@arka47</a> Whats the question though? Why do some images look so bad? It’s intentional.</p>\n<p>I’m not sure what answer you’re looking for beyond that.</p>",
      "rawMarkdown": "arka47 Whats the question though? Why do some images look so bad? It’s intentional.\n\nI’m not sure what answer you’re looking for beyond that.",
      "votes": null
    },
    {
      "id": "1230960",
      "postDate": "03/08/2021 15:09:13",
      "content": "<p>That might be a fair assumption for sure, but <a href=\"https://www.kaggle.com/pauljurczak\" target=\"_blank\">@pauljurczak</a> does have a point don't you think so? I think it's fair to ask them whether the non-trivial loss of chemical information in some images is intentional (which I think is rather probable, like you said) or an image processing error as <a href=\"https://www.kaggle.com/alien308\" target=\"_blank\">@alien308</a> points out</p>",
      "rawMarkdown": "That might be a fair assumption for sure, but @pauljurczak does have a point don't you think so? I think it's fair to ask them whether the non-trivial loss of chemical information in some images is intentional (which I think is rather probable, like you said) or an image processing error as @alien308 points out",
      "votes": null
    },
    {
      "id": "1230988",
      "postDate": "03/08/2021 15:39:17",
      "content": "<p>I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.</p>\n<p>And yes, I think Paul makes a fair point. The way I see it though, the point of the competition isn’t to build a production ready model, it’s to investigate many approaches and see which comes out on top. The main issue I see with providing HD, unaugmented images is that it allows for “dumb” methods that don’t actually learn the task to come out on top. One could just reverse engineer the image or use it to search a large library and get loss of 0 (especially considering there’s no hidden test set). Kaggle is good about discovering these exploits and doing what they can to prevent it before launching a competition.</p>",
      "rawMarkdown": "I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.\n\nAnd yes, I think Paul makes a fair point. The way I see it though, the point of the competition isn’t to build a production ready model, it’s to investigate many approaches and see which comes out on top. The main issue I see with providing HD, unaugmented images is that it allows for “dumb” methods that don’t actually learn the task to come out on top. One could just reverse engineer the image or use it to search a large library and get loss of 0 (especially considering there’s no hidden test set). Kaggle is good about discovering these exploits and doing what they can to prevent it before launching a competition.",
      "votes": null
    },
    {
      "id": "1231477",
      "postDate": "03/09/2021 02:43:55",
      "content": "<blockquote>\n  <p>I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.</p>\n</blockquote>\n<p>What about this test image (0111e04fe1c2)? It looks like a significant part of the formula is written in invisible ink. Can you be certain what the original formula is?</p>\n<p><a href=\"https://i.imgur.com/oeN62Ih.png\" target=\"_blank\">0111e04fe1c2.png</a></p>",
      "rawMarkdown": "> I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.\n\nWhat about this test image (0111e04fe1c2)? It looks like a significant part of the formula is written in invisible ink. Can you be certain what the original formula is?\n\n[0111e04fe1c2.png](https://i.imgur.com/oeN62Ih.png)",
      "votes": null
    },
    {
      "id": "1231496",
      "postDate": "03/09/2021 03:00:57",
      "content": "<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223230#1231323\" target=\"_blank\">This</a> is what the host says</p>",
      "rawMarkdown": "[This](https://www.kaggle.com/c/bms-molecular-translation/discussion/223230#1231323) is what the host says",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1229861,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "03/07/2021 16:35:37",
      "content": "<p>I think this is the challenge…=)</p>\n<p>from the description…<br>\n<code>but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts.</code></p>\n<p>If you have tons of old paper or notes with bunch of chemical formulas its impossible to go back and correct all of them one by one. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1229885,
      "author_name": "alien308",
      "author_url": "",
      "post_date": "03/07/2021 16:54:34",
      "content": "<p>This is not a challenge. This is an image processing error, not a paper or scan defect. This destroys chemical information that would have been preserved during scanning. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1230002,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/07/2021 18:31:06",
          "content": "<p>This IS the challenge though. All those artifacts you talk about are intentionally added to the dataset, including missing bonds, occluded symbols, and the resizing. Even in the example you point out, a human expert could easily complete this molecule like so:<br>\n<img src=\"https://i.imgur.com/o84tDjB.png\" alt=\"\"></p>\n<p>Providing high definition images of molecules with no augmentation would make the challenge too easy and the solutions would not generalize to images of lower quality.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230196,
          "author_name": "pauljurczak",
          "author_url": "",
          "post_date": "03/07/2021 21:33:47",
          "content": "<blockquote>\n  <p>the solutions would not generalize to images of lower quality</p>\n</blockquote>\n<p>Inverse is also true. When the training image quality is artificially lowered too far, the solution may not generalize to the real world image quality range. The MNIST data set is an example of this phenomenon. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230797,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/08/2021 13:13:46",
          "content": "<p>I believe the hosts should clear the expectations now that people are starting to get geared up for the competitions by answering questions of this sort.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230947,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/08/2021 15:01:47",
          "content": "<p><a href=\"https://www.kaggle.com/arka47\" target=\"_blank\">@arka47</a> Whats the question though? Why do some images look so bad? It’s intentional.</p>\n<p>I’m not sure what answer you’re looking for beyond that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230960,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/08/2021 15:09:13",
          "content": "<p>That might be a fair assumption for sure, but <a href=\"https://www.kaggle.com/pauljurczak\" target=\"_blank\">@pauljurczak</a> does have a point don't you think so? I think it's fair to ask them whether the non-trivial loss of chemical information in some images is intentional (which I think is rather probable, like you said) or an image processing error as <a href=\"https://www.kaggle.com/alien308\" target=\"_blank\">@alien308</a> points out</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230988,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/08/2021 15:39:17",
          "content": "<p>I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.</p>\n<p>And yes, I think Paul makes a fair point. The way I see it though, the point of the competition isn’t to build a production ready model, it’s to investigate many approaches and see which comes out on top. The main issue I see with providing HD, unaugmented images is that it allows for “dumb” methods that don’t actually learn the task to come out on top. One could just reverse engineer the image or use it to search a large library and get loss of 0 (especially considering there’s no hidden test set). Kaggle is good about discovering these exploits and doing what they can to prevent it before launching a competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231477,
          "author_name": "pauljurczak",
          "author_url": "",
          "post_date": "03/09/2021 02:43:55",
          "content": "<blockquote>\n  <p>I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.</p>\n</blockquote>\n<p>What about this test image (0111e04fe1c2)? It looks like a significant part of the formula is written in invisible ink. Can you be certain what the original formula is?</p>\n<p><a href=\"https://i.imgur.com/oeN62Ih.png\" target=\"_blank\">0111e04fe1c2.png</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231496,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/09/2021 03:00:57",
          "content": "<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223230#1231323\" target=\"_blank\">This</a> is what the host says</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1230003,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "03/07/2021 18:31:11",
      "content": "<p>As far as I understand, the data is synthetic. Therefore, all these difficulties are artificial. And the task of the model (in my understanding) is to restore the data as completely as possible. If the 'line of chemical bond' is not visible, where it should be-the model should put it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1229837": "The organizers may have used scaling while preparing the images.\nAt the same time, some lines of chemical bonds disappeared. This is very bad.  😭😭😭 Some symbols of chemical elements were also distorted. An example is 00abb3b349a.png. I did not find such defects in the images at large scale.\nPlease correct the dataset.",
    "1229861": "I think this is the challenge...=)\n\nfrom the description...\n`but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts.`\n\n If you have tons of old paper or notes with bunch of chemical formulas its impossible to go back and correct all of them one by one.",
    "1229885": "This is not a challenge. This is an image processing error, not a paper or scan defect. This destroys chemical information that would have been preserved during scanning.",
    "1230002": "This IS the challenge though. All those artifacts you talk about are intentionally added to the dataset, including missing bonds, occluded symbols, and the resizing. Even in the example you point out, a human expert could easily complete this molecule like so:\n![](https://i.imgur.com/o84tDjB.png)\n\nProviding high definition images of molecules with no augmentation would make the challenge too easy and the solutions would not generalize to images of lower quality.",
    "1230003": "As far as I understand, the data is synthetic. Therefore, all these difficulties are artificial. And the task of the model (in my understanding) is to restore the data as completely as possible. If the 'line of chemical bond' is not visible, where it should be-the model should put it.",
    "1230196": "> the solutions would not generalize to images of lower quality\n\nInverse is also true. When the training image quality is artificially lowered too far, the solution may not generalize to the real world image quality range. The MNIST data set is an example of this phenomenon.",
    "1230797": "I believe the hosts should clear the expectations now that people are starting to get geared up for the competitions by answering questions of this sort.",
    "1230947": "arka47 Whats the question though? Why do some images look so bad? It’s intentional.\n\nI’m not sure what answer you’re looking for beyond that.",
    "1230960": "That might be a fair assumption for sure, but @pauljurczak does have a point don't you think so? I think it's fair to ask them whether the non-trivial loss of chemical information in some images is intentional (which I think is rather probable, like you said) or an image processing error as @alien308 points out",
    "1230988": "I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.\n\nAnd yes, I think Paul makes a fair point. The way I see it though, the point of the competition isn’t to build a production ready model, it’s to investigate many approaches and see which comes out on top. The main issue I see with providing HD, unaugmented images is that it allows for “dumb” methods that don’t actually learn the task to come out on top. One could just reverse engineer the image or use it to search a large library and get loss of 0 (especially considering there’s no hidden test set). Kaggle is good about discovering these exploits and doing what they can to prevent it before launching a competition.",
    "1231477": "> I haven’t seen any examples of “non-trivial loss of chemical info”. Every example I see an expert could easily reconstruct the complete molecule.\n\nWhat about this test image (0111e04fe1c2)? It looks like a significant part of the formula is written in invisible ink. Can you be certain what the original formula is?\n\n[0111e04fe1c2.png](https://i.imgur.com/oeN62Ih.png)",
    "1231496": "[This](https://www.kaggle.com/c/bms-molecular-translation/discussion/223230#1231323) is what the host says"
  },
  "source": "meta"
}