{
  "id": 233672,
  "title": "Compliments To \"Eight V100s are Working Hard\" Getting Under 1.00!!",
  "url": "/competitions/bms-molecular-translation/discussion/233672",
  "author_name": "",
  "post_date": "2021-04-20T13:36:29.139723100Z",
  "votes": 11,
  "comment_count": 20,
  "views": 0,
  "content": "<p><strong>I just wanted to issue this team some compliments!</strong></p>\n<hr>\n<p>I know the competition is far from over, but personally, I find this level of performance to be incredibly inspiring, especially from a team of novices! Great job guys!!</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/zhongfeisheng6666\" target=\"_blank\">@zhongfeisheng6666</a> <br>\n<a href=\"https://www.kaggle.com/jiachengxiong\" target=\"_blank\">@jiachengxiong</a> <br>\n<a href=\"https://www.kaggle.com/xiaohongliua\" target=\"_blank\">@xiaohongliua</a> </p>",
  "messages": [
    {
      "id": "1278989",
      "postDate": "04/20/2021 13:36:29",
      "content": "<p><strong>I just wanted to issue this team some compliments!</strong></p>\n<hr>\n<p>I know the competition is far from over, but personally, I find this level of performance to be incredibly inspiring, especially from a team of novices! Great job guys!!</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/zhongfeisheng6666\" target=\"_blank\">@zhongfeisheng6666</a> <br>\n<a href=\"https://www.kaggle.com/jiachengxiong\" target=\"_blank\">@jiachengxiong</a> <br>\n<a href=\"https://www.kaggle.com/xiaohongliua\" target=\"_blank\">@xiaohongliua</a> </p>",
      "rawMarkdown": "**I just wanted to issue this team some compliments!**\n\n---\n\nI know the competition is far from over, but personally, I find this level of performance to be incredibly inspiring, especially from a team of novices! Great job guys!!\n\n---\n\n@zhongfeisheng6666 \n@jiachengxiong \n@xiaohongliua",
      "votes": null
    },
    {
      "id": "1279170",
      "postDate": "04/20/2021 16:55:28",
      "content": "<p>Congratulations too, if it is no overfitting.</p>\n<p>I wonder if they really have 8 V100s. Research grant by someone like Nvidia, University or temporary use of their company ones?</p>",
      "rawMarkdown": "Congratulations too, if it is no overfitting.\n\nI wonder if they really have 8 V100s. Research grant by someone like Nvidia, University or temporary use of their company ones?",
      "votes": null
    },
    {
      "id": "1279174",
      "postDate": "04/20/2021 17:00:42",
      "content": "<ul>\n<li>Pavel Orlov <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/226490\" target=\"_blank\">predicted scores less than 1.0</a> a month ago!</li>\n<li>Heng predicted that final top lb score could be in the range of <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/230887\" target=\"_blank\">0.5 to 0.8</a></li>\n</ul>",
      "rawMarkdown": "Pavel Orlov [predicted scores less than 1.0](https://www.kaggle.com/c/bms-molecular-translation/discussion/226490) a month ago!\n- Heng predicted that final top lb score could be in the range of [0.5 to 0.8](https://www.kaggle.com/c/bms-molecular-translation/discussion/230887)",
      "votes": null
    },
    {
      "id": "1279219",
      "postDate": "04/20/2021 17:56:28",
      "content": "<p>No if. It <strong>is</strong> overfitting to the <code>_____ layout</code> and <code>_____ noise generator</code>. Real-world examples are much more complex. Check out any folder to get a general idea about real world images: <a href=\"https://github.com/ncats/molvec/tree/master/src/test/resources/regressionTest\" target=\"_blank\">https://github.com/ncats/molvec/tree/master/src/test/resources/regressionTest</a>. Almost certain that <em>the</em> winning model will not be useful and susceptible to pixel attack and has to be re-trained from scratch. However, the training process (e.g., idea, pipeline) might be useful to build models for specific question.</p>",
      "rawMarkdown": "No if. It **is** overfitting to the `_____ layout` and `_____ noise generator`. Real-world examples are much more complex. Check out any folder to get a general idea about real world images: https://github.com/ncats/molvec/tree/master/src/test/resources/regressionTest. Almost certain that *the* winning model will not be useful and susceptible to pixel attack and has to be re-trained from scratch. However, the training process (e.g., idea, pipeline) might be useful to build models for specific question.",
      "votes": null
    },
    {
      "id": "1279286",
      "postDate": "04/20/2021 19:05:54",
      "content": "<p>The examples you linked look far cleaner and thereby easier than the examples we have. So I don't see a problem (after retraining at least). That they are zoomed in and cut out seems to be part of the task. But such mistakes could be prevented by simple image processing. Your linked examples are also obviously synthetic.<br>\nBTW, quite some test images have also a pretty clean look.</p>\n<p>I would real world pictures expect to be scans of relatively old prints. They haven't then such a white, even background and the print is at least partially faded (greyish). Their might also be rarely stains on it from such things as coffee. And they are aren't aligned, although that is easy to correct with classical methods.</p>\n<p>So I am indeed confused since the very beginning what the hosts intention is. But maybe their real world examples are close to the synthetic images. Only they know.</p>\n<p>TLDR: Linked examples are far cleaner, thereby easier to learn. Although I would expect real world examples to look different from training/test ones (dirtier/faded).</p>",
      "rawMarkdown": "The examples you linked look far cleaner and thereby easier than the examples we have. So I don't see a problem (after retraining at least). That they are zoomed in and cut out seems to be part of the task. But such mistakes could be prevented by simple image processing. Your linked examples are also obviously synthetic.\nBTW, quite some test images have also a pretty clean look.\n\nI would real world pictures expect to be scans of relatively old prints. They haven't then such a white, even background and the print is at least partially faded (greyish). Their might also be rarely stains on it from such things as coffee. And they are aren't aligned, although that is easy to correct with classical methods.\n\nSo I am indeed confused since the very beginning what the hosts intention is. But maybe their real world examples are close to the synthetic images. Only they know.\n\nTLDR: Linked examples are far cleaner, thereby easier to learn. Although I would expect real world examples to look different from training/test ones (dirtier/faded).",
      "votes": null
    },
    {
      "id": "1279309",
      "postDate": "04/20/2021 19:17:37",
      "content": "<p>Thumbs up for the TLDR :D</p>",
      "rawMarkdown": "Thumbs up for the TLDR :D",
      "votes": null
    },
    {
      "id": "1279316",
      "postDate": "04/20/2021 19:24:45",
      "content": "<p><code>Thumbs up for the TLDR :D</code><br>\nJust for you ;)</p>",
      "rawMarkdown": "`Thumbs up for the TLDR :D`\nJust for you ;)",
      "votes": null
    },
    {
      "id": "1279381",
      "postDate": "04/20/2021 20:32:53",
      "content": "<p>Yes. They are clean images. However, you can test your DL model to see if your model predicts well. What I have found is that the more we fine tune to fit the kaggle dataset, the worse performance on these cleaned image. This competition also have many images where pixel losses are so specific to the noise generator that organizer was used to generate the dataset, but they make no sense in real world application, such as <em>ghost</em> terminal methyl group <code>a08387ec61f4</code>.</p>\n<p>There are other important features than \"dirty/fade\" for a model to learn:</p>\n<ol>\n<li>chemistry</li>\n<li>depiction style (determined by molecular sketcher)</li>\n<li>pixel loss and non-chemical component (determined by noise generator)</li>\n</ol>\n<p><strong>Ideally, we want a model to learn (1) and do not overfit (2) and (3)</strong>. </p>\n<p>Since this dataset is artificially generated using <code>a fixed engine</code> and <code>a fixed parameter set</code>, some overfitting to (2) is acceptable, imo. However, we should understand the limitations. For example, the winning models may not be extended easily to predict InChI of chemical structures generated by other engines (e.g., CDK, RDKit, Marvin, ChemDraw, PubChem), with abbreviation (e.g., CF3, Me, Et, NO2), charge (COO-, N+) or salt (Mg2+) or wavy/dashed bond or colored image, because they have been so optimized to the images without these variations.  </p>\n<p>Overfitting to noise generator (3) is more problematic. The idea of molecular translation is to help human read. But there are &gt;3% images (in both training and test set) where even the experienced chemists will get 100% wrong. <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554</a> On the bright side, we can call it \"a DL model is able to recognize specific pixel patterns and helps recover the missing group due to image compression\". On the dark side, it is overfitting. If we stick to predicting <em>valid</em> InChI, one methyl group change corresponds to 30-50 edit distance. I'd loved to see how these cases are handled by the LB 0.5-0.8 models. </p>\n<p>Surely, every model has its applicability domain. While the top 1 model is pushing LB&lt;1.0, the generalizability is probably shrinking. This is purely from a computational chemist's perspective, and I hope I am wrong. Happy to create a dataset to test the robustness of model, if anyone is interested.</p>",
      "rawMarkdown": "Yes. They are clean images. However, you can test your DL model to see if your model predicts well. What I have found is that the more we fine tune to fit the kaggle dataset, the worse performance on these cleaned image. This competition also have many images where pixel losses are so specific to the noise generator that organizer was used to generate the dataset, but they make no sense in real world application, such as *ghost* terminal methyl group `a08387ec61f4`.\n\nThere are other important features than \"dirty/fade\" for a model to learn:\n1. chemistry\n2. depiction style (determined by molecular sketcher)\n3. pixel loss and non-chemical component (determined by noise generator)\n\n**Ideally, we want a model to learn (1) and do not overfit (2) and (3)**. \n\nSince this dataset is artificially generated using `a fixed engine` and `a fixed parameter set`, some overfitting to (2) is acceptable, imo. However, we should understand the limitations. For example, the winning models may not be extended easily to predict InChI of chemical structures generated by other engines (e.g., CDK, RDKit, Marvin, ChemDraw, PubChem), with abbreviation (e.g., CF3, Me, Et, NO2), charge (COO-, N+) or salt (Mg2+) or wavy/dashed bond or colored image, because they have been so optimized to the images without these variations.  \n\nOverfitting to noise generator (3) is more problematic. The idea of molecular translation is to help human read. But there are >3% images (in both training and test set) where even the experienced chemists will get 100% wrong. https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554 On the bright side, we can call it \"a DL model is able to recognize specific pixel patterns and helps recover the missing group due to image compression\". On the dark side, it is overfitting. If we stick to predicting *valid* InChI, one methyl group change corresponds to 30-50 edit distance. I'd loved to see how these cases are handled by the LB 0.5-0.8 models. \n\nSurely, every model has its applicability domain. While the top 1 model is pushing LB<1.0, the generalizability is probably shrinking. This is purely from a computational chemist's perspective, and I hope I am wrong. Happy to create a dataset to test the robustness of model, if anyone is interested.",
      "votes": null
    },
    {
      "id": "1279388",
      "postDate": "04/20/2021 20:40:45",
      "content": "<p>Sure, if you don't design your architecture well, do good augmentation and regularization, it wont work as well with clean images (from the same generator engine). <br>\nBUT if it can work very well with noise images, it will work at least as well with clean ones, if trained from scratch. That was my intended main point.</p>\n<blockquote>\n  <p>I'd loved to see how these cases are handled by the LB  ~1 models</p>\n</blockquote>\n<p>Me too.</p>\n<p><code>I am happy to create a dataset to test the robustness of model, if anyone is interested.</code><br>\nIf several people would test it with their model, that would be interesting insight. I am interested.</p>",
      "rawMarkdown": "Sure, if you don't design your architecture well, do good augmentation and regularization, it wont work as well with clean images (from the same generator engine). \nBUT if it can work very well with noise images, it will work at least as well with clean ones, if trained from scratch. That was my intended main point.\n\n> I'd loved to see how these cases are handled by the LB ~~0.5-0.8 ~~ ~1 models\n\nMe too.\n\n`I am happy to create a dataset to test the robustness of model, if anyone is interested.`\nIf several people would test it with their model, that would be interesting insight. I am interested.",
      "votes": null
    },
    {
      "id": "1279394",
      "postDate": "04/20/2021 20:56:59",
      "content": "<p>I'm glad you are interested. I will start with the structure in training set, like <code>a08387ec61f4</code></p>\n<p><em>label</em><br>\nInChI=1S/C18H26N2O5/c1-12(10-13(2)14-6-4-3-5-7-14)11-25-18(24)20-15(17(22)23)8-9-16(19)21/h3-7,12-13,15H,8-11H2,1-2H3,(H2,19,21)(H,20,24)(H,22,23)</p>\n<p><em>what image shows</em><br>\nInChI=1S/C17H24N2O5/c1-12(7-8-13-5-3-2-4-6-13)11-24-17(23)19-14(16(21)22)9-10-15(18)20/h2-6,12,14H,7-11H2,1H3,(H2,18,20)(H,19,23)(H,21,22)</p>\n<p>edit distance = 47 🙄 (There are &gt;3% images like this &amp; the \"Eight V100s are Working Hard\" model has LD &lt; 1. Just saying…)</p>",
      "rawMarkdown": "I'm glad you are interested. I will start with the structure in training set, like `a08387ec61f4`\n\n*label*\nInChI=1S/C18H26N2O5/c1-12(10-13(2)14-6-4-3-5-7-14)11-25-18(24)20-15(17(22)23)8-9-16(19)21/h3-7,12-13,15H,8-11H2,1-2H3,(H2,19,21)(H,20,24)(H,22,23)\n\n*what image shows*\nInChI=1S/C17H24N2O5/c1-12(7-8-13-5-3-2-4-6-13)11-24-17(23)19-14(16(21)22)9-10-15(18)20/h2-6,12,14H,7-11H2,1H3,(H2,18,20)(H,19,23)(H,21,22)\n\nedit distance = 47 🙄 (There are >3% images like this & the \"Eight V100s are Working Hard\" model has LD < 1. Just saying...)",
      "votes": null
    },
    {
      "id": "1279398",
      "postDate": "04/20/2021 21:07:38",
      "content": "<p>I will do myself qualitative analysis when I have implemented my inference generation code.<br>\nI think then it is finally time to read up on the InChI definition. Which means I can comprehend your examples better.</p>\n<p>Edit: could there be something missing in your example image which would make it correct again? I have looked at it, but for me it is impossible to catch non-obvious missing structures.</p>",
      "rawMarkdown": "I will do myself qualitative analysis when I have implemented my inference generation code.\nI think then it is finally time to read up on the InChI definition. Which means I can comprehend your examples better.\n\nEdit: could there be something missing in your example image which would make it correct again? I have looked at it, but for me it is impossible to catch non-obvious missing structures.",
      "votes": null
    },
    {
      "id": "1279405",
      "postDate": "04/20/2021 21:22:48",
      "content": "<p>THIS IS EXACTLY MY POINT <a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> it is impossible to catch non-obvious missing structures👍</p>\n<p>If a model is able to predict this InChI correct, I believe it's overfitting to the noise generator. I have manually scanned thousands of training set images, and found 3% are like this. There are also 2-3% cases where alternative structures are completely chemically sound e.g., <code>c11f56a93179</code>. Then think about the top1 with LD &lt; 1. </p>",
      "rawMarkdown": "THIS IS EXACTLY MY POINT @cepheidq it is impossible to catch non-obvious missing structures👍\n\nIf a model is able to predict this InChI correct, I believe it's overfitting to the noise generator. I have manually scanned thousands of training set images, and found 3% are like this. There are also 2-3% cases where alternative structures are completely chemically sound e.g., `c11f56a93179`. Then think about the top1 with LD < 1.",
      "votes": null
    },
    {
      "id": "1279411",
      "postDate": "04/20/2021 21:32:08",
      "content": "<p>Well, lets say we have 3% such cases and they lead at average to a difference of 50%. If everything else is predicted correctly we will get a expected distance of 1.5.</p>\n<p>As the distance of such cases is on average probably lower, I guess, we still have room for other miss-predictions. So a distance of ~1 isn't that unreasonable.<br>\nSomething like 0.5 though would be suspicious from that perspective.</p>\n<p>Still, you are correct of course.</p>",
      "rawMarkdown": "Well, lets say we have 3% such cases and they lead at average to a difference of 50%. If everything else is predicted correctly we will get a expected distance of 1.5.\n\nAs the distance of such cases is on average probably lower, I guess, we still have room for other miss-predictions. So a distance of ~1 isn't that unreasonable.\nSomething like 0.5 though would be suspicious from that perspective.\n\nStill, you are correct of course.",
      "votes": null
    },
    {
      "id": "1279412",
      "postDate": "04/20/2021 21:37:51",
      "content": "<blockquote>\n  <p>is impossible to catch non-obvious missing structures</p>\n</blockquote>\n<p>Is it always really though? Aren't there some molecules that are impossible to exist? So you can say when you see a e.g. OH at some position you know that is has to be a OH2.</p>\n<p>And isn't it a matter of probability in the end? Which is of course overfitting to the noise generator. But you could say from the positive perspective that the noise generator forces the model to better learn which kind of molecular high level structures are more likely.<br>\nI still think it is likely that it is an error on the hosts side. But it is also, although less likely that it was intentional</p>",
      "rawMarkdown": ">  is impossible to catch non-obvious missing structures\n\nIs it always really though? Aren't there some molecules that are impossible to exist? So you can say when you see a e.g. OH at some position you know that is has to be a OH2.\n\nAnd isn't it a matter of probability in the end? Which is of course overfitting to the noise generator. But you could say from the positive perspective that the noise generator forces the model to better learn which kind of molecular high level structures are more likely.\nI still think it is likely that it is an error on the hosts side. But it is also, although less likely that it was intentional",
      "votes": null
    },
    {
      "id": "1279414",
      "postDate": "04/20/2021 21:56:54",
      "content": "<p>I don't think it's labeling error. It is just image resizing that ruined the vertical/horizontal lines.</p>",
      "rawMarkdown": "I don't think it's labeling error. It is just image resizing that ruined the vertical/horizontal lines.",
      "votes": null
    },
    {
      "id": "1279417",
      "postDate": "04/20/2021 22:03:43",
      "content": "<p>In which way are the lines ruined that they lead to images which would represent different InChIs?<br>\nResizing usually works really well, btw.</p>",
      "rawMarkdown": "In which way are the lines ruined that they lead to images which would represent different InChIs?\nResizing usually works really well, btw.",
      "votes": null
    },
    {
      "id": "1279420",
      "postDate": "04/20/2021 22:15:07",
      "content": "<p>such image can be generated by nearest-neighbor interpolation or grid interpolation. Cubic interpolation then thresholding may also lead to entire h/v line loss.</p>",
      "rawMarkdown": "such image can be generated by nearest-neighbor interpolation or grid interpolation. Cubic interpolation then thresholding may also lead to entire h/v line loss.",
      "votes": null
    },
    {
      "id": "1279487",
      "postDate": "04/21/2021 01:04:28",
      "content": "<p>Thanks.Under your praise, 8 V100s will continue to work hard to reach poetry and beauty.</p>",
      "rawMarkdown": "Thanks.Under your praise, 8 V100s will continue to work hard to reach poetry and beauty.",
      "votes": null
    },
    {
      "id": "1280819",
      "postDate": "04/22/2021 11:36:15",
      "content": "<p>the race to the best results:<br>\n<a href=\"https://www.youtube.com/watch?time_continue=29&amp;v=fYVaHVHuARA&amp;feature=emb_logo\" target=\"_blank\">https://www.youtube.com/watch?time_continue=29&amp;v=fYVaHVHuARA&amp;feature=emb_logo</a></p>\n<p><img src=\"https://thumbs.gfycat.com/AdorableTestyHeterodontosaurus-size_restricted.gif\" alt=\"\"></p>",
      "rawMarkdown": "the race to the best results:\nhttps://www.youtube.com/watch?time_continue=29&v=fYVaHVHuARA&feature=emb_logo\n\n![](https://thumbs.gfycat.com/AdorableTestyHeterodontosaurus-size_restricted.gif)",
      "votes": null
    },
    {
      "id": "1280914",
      "postDate": "04/22/2021 13:29:17",
      "content": "<p>Great Job 1st place team =) It takes a lot of efforts to break 1 =) </p>",
      "rawMarkdown": "Great Job 1st place team =) It takes a lot of efforts to break 1 =)",
      "votes": null
    },
    {
      "id": "1283490",
      "postDate": "04/25/2021 01:25:54",
      "content": "<p>i shall given him a title:<br>\n刺客 = assassin = one shot learner</p>\n<p>Guanshuo Xu :  1.74 :  one submission</p>",
      "rawMarkdown": "i shall given him a title:\n刺客 = assassin = one shot learner\n\nGuanshuo Xu :  1.74 :  one submission",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1279170,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "04/20/2021 16:55:28",
      "content": "<p>Congratulations too, if it is no overfitting.</p>\n<p>I wonder if they really have 8 V100s. Research grant by someone like Nvidia, University or temporary use of their company ones?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1279219,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/20/2021 17:56:28",
          "content": "<p>No if. It <strong>is</strong> overfitting to the <code>_____ layout</code> and <code>_____ noise generator</code>. Real-world examples are much more complex. Check out any folder to get a general idea about real world images: <a href=\"https://github.com/ncats/molvec/tree/master/src/test/resources/regressionTest\" target=\"_blank\">https://github.com/ncats/molvec/tree/master/src/test/resources/regressionTest</a>. Almost certain that <em>the</em> winning model will not be useful and susceptible to pixel attack and has to be re-trained from scratch. However, the training process (e.g., idea, pipeline) might be useful to build models for specific question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279286,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 19:05:54",
          "content": "<p>The examples you linked look far cleaner and thereby easier than the examples we have. So I don't see a problem (after retraining at least). That they are zoomed in and cut out seems to be part of the task. But such mistakes could be prevented by simple image processing. Your linked examples are also obviously synthetic.<br>\nBTW, quite some test images have also a pretty clean look.</p>\n<p>I would real world pictures expect to be scans of relatively old prints. They haven't then such a white, even background and the print is at least partially faded (greyish). Their might also be rarely stains on it from such things as coffee. And they are aren't aligned, although that is easy to correct with classical methods.</p>\n<p>So I am indeed confused since the very beginning what the hosts intention is. But maybe their real world examples are close to the synthetic images. Only they know.</p>\n<p>TLDR: Linked examples are far cleaner, thereby easier to learn. Although I would expect real world examples to look different from training/test ones (dirtier/faded).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279309,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/20/2021 19:17:37",
          "content": "<p>Thumbs up for the TLDR :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279316,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 19:24:45",
          "content": "<p><code>Thumbs up for the TLDR :D</code><br>\nJust for you ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279381,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/20/2021 20:32:53",
          "content": "<p>Yes. They are clean images. However, you can test your DL model to see if your model predicts well. What I have found is that the more we fine tune to fit the kaggle dataset, the worse performance on these cleaned image. This competition also have many images where pixel losses are so specific to the noise generator that organizer was used to generate the dataset, but they make no sense in real world application, such as <em>ghost</em> terminal methyl group <code>a08387ec61f4</code>.</p>\n<p>There are other important features than \"dirty/fade\" for a model to learn:</p>\n<ol>\n<li>chemistry</li>\n<li>depiction style (determined by molecular sketcher)</li>\n<li>pixel loss and non-chemical component (determined by noise generator)</li>\n</ol>\n<p><strong>Ideally, we want a model to learn (1) and do not overfit (2) and (3)</strong>. </p>\n<p>Since this dataset is artificially generated using <code>a fixed engine</code> and <code>a fixed parameter set</code>, some overfitting to (2) is acceptable, imo. However, we should understand the limitations. For example, the winning models may not be extended easily to predict InChI of chemical structures generated by other engines (e.g., CDK, RDKit, Marvin, ChemDraw, PubChem), with abbreviation (e.g., CF3, Me, Et, NO2), charge (COO-, N+) or salt (Mg2+) or wavy/dashed bond or colored image, because they have been so optimized to the images without these variations.  </p>\n<p>Overfitting to noise generator (3) is more problematic. The idea of molecular translation is to help human read. But there are &gt;3% images (in both training and test set) where even the experienced chemists will get 100% wrong. <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554</a> On the bright side, we can call it \"a DL model is able to recognize specific pixel patterns and helps recover the missing group due to image compression\". On the dark side, it is overfitting. If we stick to predicting <em>valid</em> InChI, one methyl group change corresponds to 30-50 edit distance. I'd loved to see how these cases are handled by the LB 0.5-0.8 models. </p>\n<p>Surely, every model has its applicability domain. While the top 1 model is pushing LB&lt;1.0, the generalizability is probably shrinking. This is purely from a computational chemist's perspective, and I hope I am wrong. Happy to create a dataset to test the robustness of model, if anyone is interested.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279388,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 20:40:45",
          "content": "<p>Sure, if you don't design your architecture well, do good augmentation and regularization, it wont work as well with clean images (from the same generator engine). <br>\nBUT if it can work very well with noise images, it will work at least as well with clean ones, if trained from scratch. That was my intended main point.</p>\n<blockquote>\n  <p>I'd loved to see how these cases are handled by the LB  ~1 models</p>\n</blockquote>\n<p>Me too.</p>\n<p><code>I am happy to create a dataset to test the robustness of model, if anyone is interested.</code><br>\nIf several people would test it with their model, that would be interesting insight. I am interested.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279394,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/20/2021 20:56:59",
          "content": "<p>I'm glad you are interested. I will start with the structure in training set, like <code>a08387ec61f4</code></p>\n<p><em>label</em><br>\nInChI=1S/C18H26N2O5/c1-12(10-13(2)14-6-4-3-5-7-14)11-25-18(24)20-15(17(22)23)8-9-16(19)21/h3-7,12-13,15H,8-11H2,1-2H3,(H2,19,21)(H,20,24)(H,22,23)</p>\n<p><em>what image shows</em><br>\nInChI=1S/C17H24N2O5/c1-12(7-8-13-5-3-2-4-6-13)11-24-17(23)19-14(16(21)22)9-10-15(18)20/h2-6,12,14H,7-11H2,1H3,(H2,18,20)(H,19,23)(H,21,22)</p>\n<p>edit distance = 47 🙄 (There are &gt;3% images like this &amp; the \"Eight V100s are Working Hard\" model has LD &lt; 1. Just saying…)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279398,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 21:07:38",
          "content": "<p>I will do myself qualitative analysis when I have implemented my inference generation code.<br>\nI think then it is finally time to read up on the InChI definition. Which means I can comprehend your examples better.</p>\n<p>Edit: could there be something missing in your example image which would make it correct again? I have looked at it, but for me it is impossible to catch non-obvious missing structures.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279405,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/20/2021 21:22:48",
          "content": "<p>THIS IS EXACTLY MY POINT <a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> it is impossible to catch non-obvious missing structures👍</p>\n<p>If a model is able to predict this InChI correct, I believe it's overfitting to the noise generator. I have manually scanned thousands of training set images, and found 3% are like this. There are also 2-3% cases where alternative structures are completely chemically sound e.g., <code>c11f56a93179</code>. Then think about the top1 with LD &lt; 1. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279411,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 21:32:08",
          "content": "<p>Well, lets say we have 3% such cases and they lead at average to a difference of 50%. If everything else is predicted correctly we will get a expected distance of 1.5.</p>\n<p>As the distance of such cases is on average probably lower, I guess, we still have room for other miss-predictions. So a distance of ~1 isn't that unreasonable.<br>\nSomething like 0.5 though would be suspicious from that perspective.</p>\n<p>Still, you are correct of course.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279412,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 21:37:51",
          "content": "<blockquote>\n  <p>is impossible to catch non-obvious missing structures</p>\n</blockquote>\n<p>Is it always really though? Aren't there some molecules that are impossible to exist? So you can say when you see a e.g. OH at some position you know that is has to be a OH2.</p>\n<p>And isn't it a matter of probability in the end? Which is of course overfitting to the noise generator. But you could say from the positive perspective that the noise generator forces the model to better learn which kind of molecular high level structures are more likely.<br>\nI still think it is likely that it is an error on the hosts side. But it is also, although less likely that it was intentional</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279414,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/20/2021 21:56:54",
          "content": "<p>I don't think it's labeling error. It is just image resizing that ruined the vertical/horizontal lines.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279417,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/20/2021 22:03:43",
          "content": "<p>In which way are the lines ruined that they lead to images which would represent different InChIs?<br>\nResizing usually works really well, btw.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279420,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/20/2021 22:15:07",
          "content": "<p>such image can be generated by nearest-neighbor interpolation or grid interpolation. Cubic interpolation then thresholding may also lead to entire h/v line loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1279174,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "04/20/2021 17:00:42",
      "content": "<ul>\n<li>Pavel Orlov <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/226490\" target=\"_blank\">predicted scores less than 1.0</a> a month ago!</li>\n<li>Heng predicted that final top lb score could be in the range of <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/230887\" target=\"_blank\">0.5 to 0.8</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1279487,
      "author_name": "zhongfeisheng6666",
      "author_url": "",
      "post_date": "04/21/2021 01:04:28",
      "content": "<p>Thanks.Under your praise, 8 V100s will continue to work hard to reach poetry and beauty.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1280819,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/22/2021 11:36:15",
          "content": "<p>the race to the best results:<br>\n<a href=\"https://www.youtube.com/watch?time_continue=29&amp;v=fYVaHVHuARA&amp;feature=emb_logo\" target=\"_blank\">https://www.youtube.com/watch?time_continue=29&amp;v=fYVaHVHuARA&amp;feature=emb_logo</a></p>\n<p><img src=\"https://thumbs.gfycat.com/AdorableTestyHeterodontosaurus-size_restricted.gif\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1280914,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "04/22/2021 13:29:17",
      "content": "<p>Great Job 1st place team =) It takes a lot of efforts to break 1 =) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1283490,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/25/2021 01:25:54",
      "content": "<p>i shall given him a title:<br>\n刺客 = assassin = one shot learner</p>\n<p>Guanshuo Xu :  1.74 :  one submission</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1278989": "**I just wanted to issue this team some compliments!**\n\n---\n\nI know the competition is far from over, but personally, I find this level of performance to be incredibly inspiring, especially from a team of novices! Great job guys!!\n\n---\n\n@zhongfeisheng6666 \n@jiachengxiong \n@xiaohongliua",
    "1279170": "Congratulations too, if it is no overfitting.\n\nI wonder if they really have 8 V100s. Research grant by someone like Nvidia, University or temporary use of their company ones?",
    "1279174": "Pavel Orlov [predicted scores less than 1.0](https://www.kaggle.com/c/bms-molecular-translation/discussion/226490) a month ago!\n- Heng predicted that final top lb score could be in the range of [0.5 to 0.8](https://www.kaggle.com/c/bms-molecular-translation/discussion/230887)",
    "1279219": "No if. It **is** overfitting to the `_____ layout` and `_____ noise generator`. Real-world examples are much more complex. Check out any folder to get a general idea about real world images: https://github.com/ncats/molvec/tree/master/src/test/resources/regressionTest. Almost certain that *the* winning model will not be useful and susceptible to pixel attack and has to be re-trained from scratch. However, the training process (e.g., idea, pipeline) might be useful to build models for specific question.",
    "1279286": "The examples you linked look far cleaner and thereby easier than the examples we have. So I don't see a problem (after retraining at least). That they are zoomed in and cut out seems to be part of the task. But such mistakes could be prevented by simple image processing. Your linked examples are also obviously synthetic.\nBTW, quite some test images have also a pretty clean look.\n\nI would real world pictures expect to be scans of relatively old prints. They haven't then such a white, even background and the print is at least partially faded (greyish). Their might also be rarely stains on it from such things as coffee. And they are aren't aligned, although that is easy to correct with classical methods.\n\nSo I am indeed confused since the very beginning what the hosts intention is. But maybe their real world examples are close to the synthetic images. Only they know.\n\nTLDR: Linked examples are far cleaner, thereby easier to learn. Although I would expect real world examples to look different from training/test ones (dirtier/faded).",
    "1279309": "Thumbs up for the TLDR :D",
    "1279316": "`Thumbs up for the TLDR :D`\nJust for you ;)",
    "1279381": "Yes. They are clean images. However, you can test your DL model to see if your model predicts well. What I have found is that the more we fine tune to fit the kaggle dataset, the worse performance on these cleaned image. This competition also have many images where pixel losses are so specific to the noise generator that organizer was used to generate the dataset, but they make no sense in real world application, such as *ghost* terminal methyl group `a08387ec61f4`.\n\nThere are other important features than \"dirty/fade\" for a model to learn:\n1. chemistry\n2. depiction style (determined by molecular sketcher)\n3. pixel loss and non-chemical component (determined by noise generator)\n\n**Ideally, we want a model to learn (1) and do not overfit (2) and (3)**. \n\nSince this dataset is artificially generated using `a fixed engine` and `a fixed parameter set`, some overfitting to (2) is acceptable, imo. However, we should understand the limitations. For example, the winning models may not be extended easily to predict InChI of chemical structures generated by other engines (e.g., CDK, RDKit, Marvin, ChemDraw, PubChem), with abbreviation (e.g., CF3, Me, Et, NO2), charge (COO-, N+) or salt (Mg2+) or wavy/dashed bond or colored image, because they have been so optimized to the images without these variations.  \n\nOverfitting to noise generator (3) is more problematic. The idea of molecular translation is to help human read. But there are >3% images (in both training and test set) where even the experienced chemists will get 100% wrong. https://www.kaggle.com/c/bms-molecular-translation/discussion/230887#1265554 On the bright side, we can call it \"a DL model is able to recognize specific pixel patterns and helps recover the missing group due to image compression\". On the dark side, it is overfitting. If we stick to predicting *valid* InChI, one methyl group change corresponds to 30-50 edit distance. I'd loved to see how these cases are handled by the LB 0.5-0.8 models. \n\nSurely, every model has its applicability domain. While the top 1 model is pushing LB<1.0, the generalizability is probably shrinking. This is purely from a computational chemist's perspective, and I hope I am wrong. Happy to create a dataset to test the robustness of model, if anyone is interested.",
    "1279388": "Sure, if you don't design your architecture well, do good augmentation and regularization, it wont work as well with clean images (from the same generator engine). \nBUT if it can work very well with noise images, it will work at least as well with clean ones, if trained from scratch. That was my intended main point.\n\n> I'd loved to see how these cases are handled by the LB ~~0.5-0.8 ~~ ~1 models\n\nMe too.\n\n`I am happy to create a dataset to test the robustness of model, if anyone is interested.`\nIf several people would test it with their model, that would be interesting insight. I am interested.",
    "1279394": "I'm glad you are interested. I will start with the structure in training set, like `a08387ec61f4`\n\n*label*\nInChI=1S/C18H26N2O5/c1-12(10-13(2)14-6-4-3-5-7-14)11-25-18(24)20-15(17(22)23)8-9-16(19)21/h3-7,12-13,15H,8-11H2,1-2H3,(H2,19,21)(H,20,24)(H,22,23)\n\n*what image shows*\nInChI=1S/C17H24N2O5/c1-12(7-8-13-5-3-2-4-6-13)11-24-17(23)19-14(16(21)22)9-10-15(18)20/h2-6,12,14H,7-11H2,1H3,(H2,18,20)(H,19,23)(H,21,22)\n\nedit distance = 47 🙄 (There are >3% images like this & the \"Eight V100s are Working Hard\" model has LD < 1. Just saying...)",
    "1279398": "I will do myself qualitative analysis when I have implemented my inference generation code.\nI think then it is finally time to read up on the InChI definition. Which means I can comprehend your examples better.\n\nEdit: could there be something missing in your example image which would make it correct again? I have looked at it, but for me it is impossible to catch non-obvious missing structures.",
    "1279405": "THIS IS EXACTLY MY POINT @cepheidq it is impossible to catch non-obvious missing structures👍\n\nIf a model is able to predict this InChI correct, I believe it's overfitting to the noise generator. I have manually scanned thousands of training set images, and found 3% are like this. There are also 2-3% cases where alternative structures are completely chemically sound e.g., `c11f56a93179`. Then think about the top1 with LD < 1.",
    "1279411": "Well, lets say we have 3% such cases and they lead at average to a difference of 50%. If everything else is predicted correctly we will get a expected distance of 1.5.\n\nAs the distance of such cases is on average probably lower, I guess, we still have room for other miss-predictions. So a distance of ~1 isn't that unreasonable.\nSomething like 0.5 though would be suspicious from that perspective.\n\nStill, you are correct of course.",
    "1279412": ">  is impossible to catch non-obvious missing structures\n\nIs it always really though? Aren't there some molecules that are impossible to exist? So you can say when you see a e.g. OH at some position you know that is has to be a OH2.\n\nAnd isn't it a matter of probability in the end? Which is of course overfitting to the noise generator. But you could say from the positive perspective that the noise generator forces the model to better learn which kind of molecular high level structures are more likely.\nI still think it is likely that it is an error on the hosts side. But it is also, although less likely that it was intentional",
    "1279414": "I don't think it's labeling error. It is just image resizing that ruined the vertical/horizontal lines.",
    "1279417": "In which way are the lines ruined that they lead to images which would represent different InChIs?\nResizing usually works really well, btw.",
    "1279420": "such image can be generated by nearest-neighbor interpolation or grid interpolation. Cubic interpolation then thresholding may also lead to entire h/v line loss.",
    "1279487": "Thanks.Under your praise, 8 V100s will continue to work hard to reach poetry and beauty.",
    "1280819": "the race to the best results:\nhttps://www.youtube.com/watch?time_continue=29&v=fYVaHVHuARA&feature=emb_logo\n\n![](https://thumbs.gfycat.com/AdorableTestyHeterodontosaurus-size_restricted.gif)",
    "1280914": "Great Job 1st place team =) It takes a lot of efforts to break 1 =)",
    "1283490": "i shall given him a title:\n刺客 = assassin = one shot learner\n\nGuanshuo Xu :  1.74 :  one submission"
  },
  "source": "meta"
}