{
  "id": 233927,
  "title": "Information encoded in the image margins",
  "url": "/competitions/bms-molecular-translation/discussion/233927",
  "author_name": "",
  "post_date": "2021-04-21T20:27:07.686640100Z",
  "votes": 21,
  "comment_count": 9,
  "views": 0,
  "content": "<h2>[ Background ]</h2>\n<p>The chemical images in this competition were generated using a specific <code>engine</code> + a fixed <code>parameter set</code> and crappified by <code>artificial noise</code>. The noise involves image resizing which leads to pixel loss, rotation (much more in the test set) and random pixel addition. One of the challenges is to recover the real molecule from these pixel loss. Some loss are easy to recover, e.g., a line in aromatic ring. Some are more tricky.</p>\n<h2>[ What's encoded in the image margins? ]</h2>\n<p>Well, the correct question is <em>what is missing in the margins?</em> or <em>does the margin appear wider than it should have been?</em> As I mentioned in the background section, images were generated using a <code>fixed</code> parameter set without cropping &amp; padding, then right margin = left margin, top margin = bottom margin. Therefore, whenever you see uneven margins, let's say top margin &gt; bottom margin, it means some pixels (usually horizontal/vertical lines) was lost during the lossy resizing process. </p>\n<p>A little bit into chemistry. Some are correctable loss, e.g., C=N--CH3, if the methyl <code>CH3</code> is lost, since we know that nitrogen must have 3 bonds instead of 2, we know for sure that there must be something linking to the nitrogen. However, some are <em>incorrectable</em> loss, which means whether pixel is lost or not <em>does NOT</em> make the structure invalid. Here is an example, top margin = 27, bottom margin = 50. <img src=\"https://i.ibb.co/6rLsCY1/example.png\" alt=\"example\"></p>\n<p>I call the incorrectable pixel loss <strong>ghost methyl group</strong>. It sounds awful, but in fact, some ghost methyl group might be correctable, because pixel loss may leads to uneven margin! This shouldn't happen! There must be vertical/horizontal line there! On the other hand, some ghost methyl is completely incorrectable, because they didn't change the white margins. </p>\n<h2>[ But how many ghost methyl group in the dataset? ]</h2>\n<p>Here I <em>randomly</em> inspected 3,000 images from the training set, and found 85 cases (2.83%). Then I inspected another ~1,000 images, about same 3%. Note there is no cherry picking. Below listed some examples. </p>\n<h5>Correctable due to uneven margins</h5>\n<p><img src=\"https://i.ibb.co/8MMJhYH/correctable.png\" alt=\"correctable\"></p>\n<h5>Incorrectable</h5>\n<p><img src=\"https://i.ibb.co/mz1jX0V/m2.png\" alt=\"incorrectable\"></p>\n<h5>Complete list</h5>\n<p>01bb86b2e825,025b0a6c0ab0,09f23f281ac5,1091c2356821,115b3a2a6007,129f29ad58c5,16fa6e0a9e84,187183c350c6,1a6a7a357c48,1ab069aad819,1b3c1ade1202,1f4631762309,211bb8c440d1,211e9c4f22ca,229a5aae1ed0,2b9d54e524c5,312f53b2fb67,33a19834d838,365765eb72f2,3a49a5cd18bb,3a88219a277f,41129c013d4e,430eca312ee8,43244448c608,4414eec4a2d7,4815b5e8156c,4a4f6c36a791,553a2690f787,5c7425c79375,5d0e226a94c2,5ff1da202c41,681fd50b2a13,6c5ad8f1116c,72f418a14a23,76cb0208483f,776423bb98e4,7ca7f45ebf2f,7d9ee1264eb4,7f50dc5e336b,80ac0e096696,8774bc1cee3c,8c8340dde1be,8df25092b072,8f6e6cc4eb01,91af226dad0d,94000a9eb5e2,983a6ca4c528,995119009ce3,9bf73e56cdb1,9cae784aee4f,9cf55e180ee0,a08387ec61f4,a43af4631915,a4f0be16a961,a8d66e4a92cd,ab64a9f72959,ab6a6e53ea10,aba62e258602,af6723b1d516,b28d3306506b,b4a6ffb846be,b5d10215ca64,beb9dbd6b02e,c199408a807c,c2e9471aa44d,c8ae4451d4c9,caf5839872f2,d1504dc418d2,d65d17153e9b,d707d1bea71a,db83543d977f,e0e9fdf88cfd,e1531fa3080e,e1a3f26bcd4e,e5bb05bba291,e81a0a99741e,ea360f954322,eaaeb9d377c6,eab8f4fc6744,ed55215a20e4,f202b7cb5e5f,f206467a62b3,f262fd4a477a,faa7dd7b546d,fc94efaffe27</p>\n<h2>[ How does it affect Levenshtein distance? ]</h2>\n<p>If a model predicts <em>what the image literally shows</em> and always results in valid InChI, the Levenshtein distance between +- ghost methyl is about <code>41±13</code>, depending on the molecular size. Larger molecule -&gt; longer InChI -&gt; larger distance when removing a carbon. </p>\n<p>If a preprocessing is done by trimming out all white margins, the correctable cases may become incorrectable, because all the information encoded in the margins are lost.</p>\n<p>If uneven margins are found but there are <em>multiple</em> carbons where a bond can attach, as shown in the above example, good luck.</p>\n<h2>[ Science ]</h2>\n<p>As the great kagglers pushing the leaderboard to LD &lt; 1.0, what is our model learning then? I'm afraid the model is learning more and more on: </p>\n<ul>\n<li>how to read information in the image margins</li>\n<li>how to make invalid InChI when ghost methyl could happen (LD = 41 is too bad!)</li>\n<li>deciphering tiny pixel packs</li>\n<li>how to create bonds to fit the layout generated by software xyz</li>\n</ul>\n<p>It is a judgement call whether these things worth learning or not, but pretty sure as a scientist, I don't care about these things. Hopefully I am wrong. </p>",
  "messages": [
    {
      "id": "1280322",
      "postDate": "04/21/2021 20:27:07",
      "content": "<h2>[ Background ]</h2>\n<p>The chemical images in this competition were generated using a specific <code>engine</code> + a fixed <code>parameter set</code> and crappified by <code>artificial noise</code>. The noise involves image resizing which leads to pixel loss, rotation (much more in the test set) and random pixel addition. One of the challenges is to recover the real molecule from these pixel loss. Some loss are easy to recover, e.g., a line in aromatic ring. Some are more tricky.</p>\n<h2>[ What's encoded in the image margins? ]</h2>\n<p>Well, the correct question is <em>what is missing in the margins?</em> or <em>does the margin appear wider than it should have been?</em> As I mentioned in the background section, images were generated using a <code>fixed</code> parameter set without cropping &amp; padding, then right margin = left margin, top margin = bottom margin. Therefore, whenever you see uneven margins, let's say top margin &gt; bottom margin, it means some pixels (usually horizontal/vertical lines) was lost during the lossy resizing process. </p>\n<p>A little bit into chemistry. Some are correctable loss, e.g., C=N--CH3, if the methyl <code>CH3</code> is lost, since we know that nitrogen must have 3 bonds instead of 2, we know for sure that there must be something linking to the nitrogen. However, some are <em>incorrectable</em> loss, which means whether pixel is lost or not <em>does NOT</em> make the structure invalid. Here is an example, top margin = 27, bottom margin = 50. <img src=\"https://i.ibb.co/6rLsCY1/example.png\" alt=\"example\"></p>\n<p>I call the incorrectable pixel loss <strong>ghost methyl group</strong>. It sounds awful, but in fact, some ghost methyl group might be correctable, because pixel loss may leads to uneven margin! This shouldn't happen! There must be vertical/horizontal line there! On the other hand, some ghost methyl is completely incorrectable, because they didn't change the white margins. </p>\n<h2>[ But how many ghost methyl group in the dataset? ]</h2>\n<p>Here I <em>randomly</em> inspected 3,000 images from the training set, and found 85 cases (2.83%). Then I inspected another ~1,000 images, about same 3%. Note there is no cherry picking. Below listed some examples. </p>\n<h5>Correctable due to uneven margins</h5>\n<p><img src=\"https://i.ibb.co/8MMJhYH/correctable.png\" alt=\"correctable\"></p>\n<h5>Incorrectable</h5>\n<p><img src=\"https://i.ibb.co/mz1jX0V/m2.png\" alt=\"incorrectable\"></p>\n<h5>Complete list</h5>\n<p>01bb86b2e825,025b0a6c0ab0,09f23f281ac5,1091c2356821,115b3a2a6007,129f29ad58c5,16fa6e0a9e84,187183c350c6,1a6a7a357c48,1ab069aad819,1b3c1ade1202,1f4631762309,211bb8c440d1,211e9c4f22ca,229a5aae1ed0,2b9d54e524c5,312f53b2fb67,33a19834d838,365765eb72f2,3a49a5cd18bb,3a88219a277f,41129c013d4e,430eca312ee8,43244448c608,4414eec4a2d7,4815b5e8156c,4a4f6c36a791,553a2690f787,5c7425c79375,5d0e226a94c2,5ff1da202c41,681fd50b2a13,6c5ad8f1116c,72f418a14a23,76cb0208483f,776423bb98e4,7ca7f45ebf2f,7d9ee1264eb4,7f50dc5e336b,80ac0e096696,8774bc1cee3c,8c8340dde1be,8df25092b072,8f6e6cc4eb01,91af226dad0d,94000a9eb5e2,983a6ca4c528,995119009ce3,9bf73e56cdb1,9cae784aee4f,9cf55e180ee0,a08387ec61f4,a43af4631915,a4f0be16a961,a8d66e4a92cd,ab64a9f72959,ab6a6e53ea10,aba62e258602,af6723b1d516,b28d3306506b,b4a6ffb846be,b5d10215ca64,beb9dbd6b02e,c199408a807c,c2e9471aa44d,c8ae4451d4c9,caf5839872f2,d1504dc418d2,d65d17153e9b,d707d1bea71a,db83543d977f,e0e9fdf88cfd,e1531fa3080e,e1a3f26bcd4e,e5bb05bba291,e81a0a99741e,ea360f954322,eaaeb9d377c6,eab8f4fc6744,ed55215a20e4,f202b7cb5e5f,f206467a62b3,f262fd4a477a,faa7dd7b546d,fc94efaffe27</p>\n<h2>[ How does it affect Levenshtein distance? ]</h2>\n<p>If a model predicts <em>what the image literally shows</em> and always results in valid InChI, the Levenshtein distance between +- ghost methyl is about <code>41±13</code>, depending on the molecular size. Larger molecule -&gt; longer InChI -&gt; larger distance when removing a carbon. </p>\n<p>If a preprocessing is done by trimming out all white margins, the correctable cases may become incorrectable, because all the information encoded in the margins are lost.</p>\n<p>If uneven margins are found but there are <em>multiple</em> carbons where a bond can attach, as shown in the above example, good luck.</p>\n<h2>[ Science ]</h2>\n<p>As the great kagglers pushing the leaderboard to LD &lt; 1.0, what is our model learning then? I'm afraid the model is learning more and more on: </p>\n<ul>\n<li>how to read information in the image margins</li>\n<li>how to make invalid InChI when ghost methyl could happen (LD = 41 is too bad!)</li>\n<li>deciphering tiny pixel packs</li>\n<li>how to create bonds to fit the layout generated by software xyz</li>\n</ul>\n<p>It is a judgement call whether these things worth learning or not, but pretty sure as a scientist, I don't care about these things. Hopefully I am wrong. </p>",
      "rawMarkdown": "## [ Background ]\nThe chemical images in this competition were generated using a specific `engine` + a fixed `parameter set` and crappified by `artificial noise`. The noise involves image resizing which leads to pixel loss, rotation (much more in the test set) and random pixel addition. One of the challenges is to recover the real molecule from these pixel loss. Some loss are easy to recover, e.g., a line in aromatic ring. Some are more tricky.\n\n## [ What's encoded in the image margins? ]\nWell, the correct question is *what is missing in the margins?* or *does the margin appear wider than it should have been?* As I mentioned in the background section, images were generated using a `fixed` parameter set without cropping & padding, then right margin = left margin, top margin = bottom margin. Therefore, whenever you see uneven margins, let's say top margin > bottom margin, it means some pixels (usually horizontal/vertical lines) was lost during the lossy resizing process. \n\nA little bit into chemistry. Some are correctable loss, e.g., C=N--CH3, if the methyl `CH3` is lost, since we know that nitrogen must have 3 bonds instead of 2, we know for sure that there must be something linking to the nitrogen. However, some are *incorrectable* loss, which means whether pixel is lost or not *does NOT* make the structure invalid. Here is an example, top margin = 27, bottom margin = 50. ![example](https://i.ibb.co/6rLsCY1/example.png)\n\nI call the incorrectable pixel loss **ghost methyl group**. It sounds awful, but in fact, some ghost methyl group might be correctable, because pixel loss may leads to uneven margin! This shouldn't happen! There must be vertical/horizontal line there! On the other hand, some ghost methyl is completely incorrectable, because they didn't change the white margins. \n \n## [ But how many ghost methyl group in the dataset? ]\nHere I *randomly* inspected 3,000 images from the training set, and found 85 cases (2.83%). Then I inspected another ~1,000 images, about same 3%. Note there is no cherry picking. Below listed some examples. \n\n##### Correctable due to uneven margins\n![correctable](https://i.ibb.co/8MMJhYH/correctable.png)\n##### Incorrectable\n![incorrectable](https://i.ibb.co/mz1jX0V/m2.png)\n\n##### Complete list\n01bb86b2e825,025b0a6c0ab0,09f23f281ac5,1091c2356821,115b3a2a6007,129f29ad58c5,16fa6e0a9e84,187183c350c6,1a6a7a357c48,1ab069aad819,1b3c1ade1202,1f4631762309,211bb8c440d1,211e9c4f22ca,229a5aae1ed0,2b9d54e524c5,312f53b2fb67,33a19834d838,365765eb72f2,3a49a5cd18bb,3a88219a277f,41129c013d4e,430eca312ee8,43244448c608,4414eec4a2d7,4815b5e8156c,4a4f6c36a791,553a2690f787,5c7425c79375,5d0e226a94c2,5ff1da202c41,681fd50b2a13,6c5ad8f1116c,72f418a14a23,76cb0208483f,776423bb98e4,7ca7f45ebf2f,7d9ee1264eb4,7f50dc5e336b,80ac0e096696,8774bc1cee3c,8c8340dde1be,8df25092b072,8f6e6cc4eb01,91af226dad0d,94000a9eb5e2,983a6ca4c528,995119009ce3,9bf73e56cdb1,9cae784aee4f,9cf55e180ee0,a08387ec61f4,a43af4631915,a4f0be16a961,a8d66e4a92cd,ab64a9f72959,ab6a6e53ea10,aba62e258602,af6723b1d516,b28d3306506b,b4a6ffb846be,b5d10215ca64,beb9dbd6b02e,c199408a807c,c2e9471aa44d,c8ae4451d4c9,caf5839872f2,d1504dc418d2,d65d17153e9b,d707d1bea71a,db83543d977f,e0e9fdf88cfd,e1531fa3080e,e1a3f26bcd4e,e5bb05bba291,e81a0a99741e,ea360f954322,eaaeb9d377c6,eab8f4fc6744,ed55215a20e4,f202b7cb5e5f,f206467a62b3,f262fd4a477a,faa7dd7b546d,fc94efaffe27\n\n\n## [ How does it affect Levenshtein distance? ]\nIf a model predicts *what the image literally shows* and always results in valid InChI, the Levenshtein distance between +- ghost methyl is about `41±13`, depending on the molecular size. Larger molecule -> longer InChI -> larger distance when removing a carbon. \n\nIf a preprocessing is done by trimming out all white margins, the correctable cases may become incorrectable, because all the information encoded in the margins are lost.\n\nIf uneven margins are found but there are *multiple* carbons where a bond can attach, as shown in the above example, good luck.\n\n## [ Science ]\nAs the great kagglers pushing the leaderboard to LD < 1.0, what is our model learning then? I'm afraid the model is learning more and more on: \n* how to read information in the image margins\n* how to make invalid InChI when ghost methyl could happen (LD = 41 is too bad!)\n* deciphering tiny pixel packs\n* how to create bonds to fit the layout generated by software xyz\n\nIt is a judgement call whether these things worth learning or not, but pretty sure as a scientist, I don't care about these things. Hopefully I am wrong.",
      "votes": null
    },
    {
      "id": "1280368",
      "postDate": "04/21/2021 20:59:33",
      "content": "<p>So encode margin sizes if you crop the image. And maybe also, if you don't.</p>\n<p>Kaggle competitions got way better in terms of data leakage &amp; comparable things, why I finally decided to participate for real.<br>\nI think this case is still okay compared to past ones. Although best would be, if it weren't there at all.</p>\n<blockquote>\n  <p>crappified</p>\n</blockquote>\n<p>I like that neologism <em>Thumps up</em></p>",
      "rawMarkdown": "So encode margin sizes if you crop the image. And maybe also, if you don't.\n\nKaggle competitions got way better in terms of data leakage & comparable things, why I finally decided to participate for real.\nI think this case is still okay compared to past ones. Although best would be, if it weren't there at all.\n\n> crappified\n\nI like that neologism *Thumps up*",
      "votes": null
    },
    {
      "id": "1280457",
      "postDate": "04/22/2021 02:07:14",
      "content": "<p>i want to add that the \"size of the margin\" is related to scale used in the resizing.</p>\n<p>in theory, the size of the image also contains information (because of the way data is prepared)</p>",
      "rawMarkdown": "i want to add that the \"size of the margin\" is related to scale used in the resizing.\n\nin theory, the size of the image also contains information (because of the way data is prepared)",
      "votes": null
    },
    {
      "id": "1280756",
      "postDate": "04/22/2021 10:00:23",
      "content": "<p>It may not be leakage. If \"ghosting\" happens in real life scenarios, then just scan the whole image area to preserve the margins that may give us clues about these missing bonds.</p>",
      "rawMarkdown": "It may not be leakage. If \"ghosting\" happens in real life scenarios, then just scan the whole image area to preserve the margins that may give us clues about these missing bonds.",
      "votes": null
    },
    {
      "id": "1280872",
      "postDate": "04/22/2021 12:46:41",
      "content": "<p>Ghosting could happen in real world but reviewer/supervisor will reject such document. It is a \"my computer says no\" meme. </p>",
      "rawMarkdown": "Ghosting could happen in real world but reviewer/supervisor will reject such document. It is a \"my computer says no\" meme.",
      "votes": null
    },
    {
      "id": "1280883",
      "postDate": "04/22/2021 12:53:15",
      "content": "<p>Isn't ghosting happening because of scanning?</p>\n<p>Edit: I thought that the whole time, but now as I write it, sounds foolish :D</p>",
      "rawMarkdown": "Isn't ghosting happening because of scanning?\n\nEdit: I thought that the whole time, but now as I write it, sounds foolish :D",
      "votes": null
    },
    {
      "id": "1280891",
      "postDate": "04/22/2021 13:02:15",
      "content": "<p>Very rare. Usually a product of bad resizing and binarization. It brings problem e.g., I patent molecule A but literally it looks like B in the document. It doesn't make sense if I argue \"my AI model says it <em>is</em> A\".</p>",
      "rawMarkdown": "Very rare. Usually a product of bad resizing and binarization. It brings problem e.g., I patent molecule A but literally it looks like B in the document. It doesn't make sense if I argue \"my AI model says it *is* A\".",
      "votes": null
    },
    {
      "id": "1280937",
      "postDate": "04/22/2021 13:46:50",
      "content": "<p>Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points. What amaze me is that the leading teams can still get good results for normal data points. </p>",
      "rawMarkdown": "Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points. What amaze me is that the leading teams can still get good results for normal data points.",
      "votes": null
    },
    {
      "id": "1280947",
      "postDate": "04/22/2021 13:58:27",
      "content": "<blockquote>\n  <p>Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points.</p>\n</blockquote>\n<p>My experience is that this is not necessarily true. As I feed more images to my model (100k incrementing), the overall performance keeps increasing, for both clean images &amp; with ghost groups. I guess the model is getting better and better at recognizing <em>the layout algorithm</em> from specific software (I have figured out which =). </p>",
      "rawMarkdown": "> Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points.\n\nMy experience is that this is not necessarily true. As I feed more images to my model (100k incrementing), the overall performance keeps increasing, for both clean images & with ghost groups. I guess the model is getting better and better at recognizing *the layout algorithm* from specific software (I have figured out which =).",
      "votes": null
    },
    {
      "id": "1281093",
      "postDate": "04/22/2021 16:19:22",
      "content": "<p>Even modern image models are very depended on context. I found these papers: <a href=\"https://arxiv.org/pdf/1911.07349.pdf\" target=\"_blank\">Paper1</a> &amp; <a href=\"https://arxiv.org/pdf/2104.02215.pdf\" target=\"_blank\">Paper2</a>. So the <a href=\"https://www.gwern.net/Tanks\" target=\"_blank\">tank story</a> has its modern, real-life counter part.</p>\n<p>So it seems very likely, as stated in your <em>Science</em> section, that the image models will learn meta information like the margin, scale within the image or even not random enough noise.</p>\n<p>That means also that the models indeed wont generalize well, but as the images are synthetic anyway it is on the hosts to see too that. Sadly the way Kaggle competitions work, it would be stupid to not use the meta information if it gives you an advantage.</p>",
      "rawMarkdown": "Even modern image models are very depended on context. I found these papers: [Paper1](https://arxiv.org/pdf/1911.07349.pdf) & [Paper2](https://arxiv.org/pdf/2104.02215.pdf). So the [tank story](https://www.gwern.net/Tanks) has its modern, real-life counter part.\n\nSo it seems very likely, as stated in your *Science* section, that the image models will learn meta information like the margin, scale within the image or even not random enough noise.\n\nThat means also that the models indeed wont generalize well, but as the images are synthetic anyway it is on the hosts to see too that. Sadly the way Kaggle competitions work, it would be stupid to not use the meta information if it gives you an advantage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1280368,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "04/21/2021 20:59:33",
      "content": "<p>So encode margin sizes if you crop the image. And maybe also, if you don't.</p>\n<p>Kaggle competitions got way better in terms of data leakage &amp; comparable things, why I finally decided to participate for real.<br>\nI think this case is still okay compared to past ones. Although best would be, if it weren't there at all.</p>\n<blockquote>\n  <p>crappified</p>\n</blockquote>\n<p>I like that neologism <em>Thumps up</em></p>",
      "votes": null,
      "replies": [
        {
          "id": 1280756,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/22/2021 10:00:23",
          "content": "<p>It may not be leakage. If \"ghosting\" happens in real life scenarios, then just scan the whole image area to preserve the margins that may give us clues about these missing bonds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280872,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/22/2021 12:46:41",
          "content": "<p>Ghosting could happen in real world but reviewer/supervisor will reject such document. It is a \"my computer says no\" meme. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280883,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/22/2021 12:53:15",
          "content": "<p>Isn't ghosting happening because of scanning?</p>\n<p>Edit: I thought that the whole time, but now as I write it, sounds foolish :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280891,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/22/2021 13:02:15",
          "content": "<p>Very rare. Usually a product of bad resizing and binarization. It brings problem e.g., I patent molecule A but literally it looks like B in the document. It doesn't make sense if I argue \"my AI model says it <em>is</em> A\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280937,
          "author_name": "joblessphysicist",
          "author_url": "",
          "post_date": "04/22/2021 13:46:50",
          "content": "<p>Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points. What amaze me is that the leading teams can still get good results for normal data points. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280947,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "04/22/2021 13:58:27",
          "content": "<blockquote>\n  <p>Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points.</p>\n</blockquote>\n<p>My experience is that this is not necessarily true. As I feed more images to my model (100k incrementing), the overall performance keeps increasing, for both clean images &amp; with ghost groups. I guess the model is getting better and better at recognizing <em>the layout algorithm</em> from specific software (I have figured out which =). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1280457,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/22/2021 02:07:14",
      "content": "<p>i want to add that the \"size of the margin\" is related to scale used in the resizing.</p>\n<p>in theory, the size of the image also contains information (because of the way data is prepared)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1281093,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "04/22/2021 16:19:22",
      "content": "<p>Even modern image models are very depended on context. I found these papers: <a href=\"https://arxiv.org/pdf/1911.07349.pdf\" target=\"_blank\">Paper1</a> &amp; <a href=\"https://arxiv.org/pdf/2104.02215.pdf\" target=\"_blank\">Paper2</a>. So the <a href=\"https://www.gwern.net/Tanks\" target=\"_blank\">tank story</a> has its modern, real-life counter part.</p>\n<p>So it seems very likely, as stated in your <em>Science</em> section, that the image models will learn meta information like the margin, scale within the image or even not random enough noise.</p>\n<p>That means also that the models indeed wont generalize well, but as the images are synthetic anyway it is on the hosts to see too that. Sadly the way Kaggle competitions work, it would be stupid to not use the meta information if it gives you an advantage.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1280322": "## [ Background ]\nThe chemical images in this competition were generated using a specific `engine` + a fixed `parameter set` and crappified by `artificial noise`. The noise involves image resizing which leads to pixel loss, rotation (much more in the test set) and random pixel addition. One of the challenges is to recover the real molecule from these pixel loss. Some loss are easy to recover, e.g., a line in aromatic ring. Some are more tricky.\n\n## [ What's encoded in the image margins? ]\nWell, the correct question is *what is missing in the margins?* or *does the margin appear wider than it should have been?* As I mentioned in the background section, images were generated using a `fixed` parameter set without cropping & padding, then right margin = left margin, top margin = bottom margin. Therefore, whenever you see uneven margins, let's say top margin > bottom margin, it means some pixels (usually horizontal/vertical lines) was lost during the lossy resizing process. \n\nA little bit into chemistry. Some are correctable loss, e.g., C=N--CH3, if the methyl `CH3` is lost, since we know that nitrogen must have 3 bonds instead of 2, we know for sure that there must be something linking to the nitrogen. However, some are *incorrectable* loss, which means whether pixel is lost or not *does NOT* make the structure invalid. Here is an example, top margin = 27, bottom margin = 50. ![example](https://i.ibb.co/6rLsCY1/example.png)\n\nI call the incorrectable pixel loss **ghost methyl group**. It sounds awful, but in fact, some ghost methyl group might be correctable, because pixel loss may leads to uneven margin! This shouldn't happen! There must be vertical/horizontal line there! On the other hand, some ghost methyl is completely incorrectable, because they didn't change the white margins. \n \n## [ But how many ghost methyl group in the dataset? ]\nHere I *randomly* inspected 3,000 images from the training set, and found 85 cases (2.83%). Then I inspected another ~1,000 images, about same 3%. Note there is no cherry picking. Below listed some examples. \n\n##### Correctable due to uneven margins\n![correctable](https://i.ibb.co/8MMJhYH/correctable.png)\n##### Incorrectable\n![incorrectable](https://i.ibb.co/mz1jX0V/m2.png)\n\n##### Complete list\n01bb86b2e825,025b0a6c0ab0,09f23f281ac5,1091c2356821,115b3a2a6007,129f29ad58c5,16fa6e0a9e84,187183c350c6,1a6a7a357c48,1ab069aad819,1b3c1ade1202,1f4631762309,211bb8c440d1,211e9c4f22ca,229a5aae1ed0,2b9d54e524c5,312f53b2fb67,33a19834d838,365765eb72f2,3a49a5cd18bb,3a88219a277f,41129c013d4e,430eca312ee8,43244448c608,4414eec4a2d7,4815b5e8156c,4a4f6c36a791,553a2690f787,5c7425c79375,5d0e226a94c2,5ff1da202c41,681fd50b2a13,6c5ad8f1116c,72f418a14a23,76cb0208483f,776423bb98e4,7ca7f45ebf2f,7d9ee1264eb4,7f50dc5e336b,80ac0e096696,8774bc1cee3c,8c8340dde1be,8df25092b072,8f6e6cc4eb01,91af226dad0d,94000a9eb5e2,983a6ca4c528,995119009ce3,9bf73e56cdb1,9cae784aee4f,9cf55e180ee0,a08387ec61f4,a43af4631915,a4f0be16a961,a8d66e4a92cd,ab64a9f72959,ab6a6e53ea10,aba62e258602,af6723b1d516,b28d3306506b,b4a6ffb846be,b5d10215ca64,beb9dbd6b02e,c199408a807c,c2e9471aa44d,c8ae4451d4c9,caf5839872f2,d1504dc418d2,d65d17153e9b,d707d1bea71a,db83543d977f,e0e9fdf88cfd,e1531fa3080e,e1a3f26bcd4e,e5bb05bba291,e81a0a99741e,ea360f954322,eaaeb9d377c6,eab8f4fc6744,ed55215a20e4,f202b7cb5e5f,f206467a62b3,f262fd4a477a,faa7dd7b546d,fc94efaffe27\n\n\n## [ How does it affect Levenshtein distance? ]\nIf a model predicts *what the image literally shows* and always results in valid InChI, the Levenshtein distance between +- ghost methyl is about `41±13`, depending on the molecular size. Larger molecule -> longer InChI -> larger distance when removing a carbon. \n\nIf a preprocessing is done by trimming out all white margins, the correctable cases may become incorrectable, because all the information encoded in the margins are lost.\n\nIf uneven margins are found but there are *multiple* carbons where a bond can attach, as shown in the above example, good luck.\n\n## [ Science ]\nAs the great kagglers pushing the leaderboard to LD < 1.0, what is our model learning then? I'm afraid the model is learning more and more on: \n* how to read information in the image margins\n* how to make invalid InChI when ghost methyl could happen (LD = 41 is too bad!)\n* deciphering tiny pixel packs\n* how to create bonds to fit the layout generated by software xyz\n\nIt is a judgement call whether these things worth learning or not, but pretty sure as a scientist, I don't care about these things. Hopefully I am wrong.",
    "1280368": "So encode margin sizes if you crop the image. And maybe also, if you don't.\n\nKaggle competitions got way better in terms of data leakage & comparable things, why I finally decided to participate for real.\nI think this case is still okay compared to past ones. Although best would be, if it weren't there at all.\n\n> crappified\n\nI like that neologism *Thumps up*",
    "1280457": "i want to add that the \"size of the margin\" is related to scale used in the resizing.\n\nin theory, the size of the image also contains information (because of the way data is prepared)",
    "1280756": "It may not be leakage. If \"ghosting\" happens in real life scenarios, then just scan the whole image area to preserve the margins that may give us clues about these missing bonds.",
    "1280872": "Ghosting could happen in real world but reviewer/supervisor will reject such document. It is a \"my computer says no\" meme.",
    "1280883": "Isn't ghosting happening because of scanning?\n\nEdit: I thought that the whole time, but now as I write it, sounds foolish :D",
    "1280891": "Very rare. Usually a product of bad resizing and binarization. It brings problem e.g., I patent molecule A but literally it looks like B in the document. It doesn't make sense if I argue \"my AI model says it *is* A\".",
    "1280937": "Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points. What amaze me is that the leading teams can still get good results for normal data points.",
    "1280947": "> Usually if we over-fit outliners (e.g. molecules with invisible methyl group), we will see poorer predictions for normal data points.\n\nMy experience is that this is not necessarily true. As I feed more images to my model (100k incrementing), the overall performance keeps increasing, for both clean images & with ghost groups. I guess the model is getting better and better at recognizing *the layout algorithm* from specific software (I have figured out which =).",
    "1281093": "Even modern image models are very depended on context. I found these papers: [Paper1](https://arxiv.org/pdf/1911.07349.pdf) & [Paper2](https://arxiv.org/pdf/2104.02215.pdf). So the [tank story](https://www.gwern.net/Tanks) has its modern, real-life counter part.\n\nSo it seems very likely, as stated in your *Science* section, that the image models will learn meta information like the margin, scale within the image or even not random enough noise.\n\nThat means also that the models indeed wont generalize well, but as the images are synthetic anyway it is on the hosts to see too that. Sadly the way Kaggle competitions work, it would be stupid to not use the meta information if it gives you an advantage."
  },
  "source": "meta"
}