{
  "id": 228220,
  "title": "My model predictions are better than training labels",
  "url": "/competitions/bms-molecular-translation/discussion/228220",
  "author_name": "",
  "post_date": "2021-03-23T20:22:25.298813200Z",
  "votes": 44,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Looking through model prediction errors on validation set, I found that my model actually returns more reasonable labels than training labels themselves. If you look closer at pubchem images and compare them to the training image, you'll notice that some bonds are missing, and the model predicts formula that corresponds to the image without those bonds (which is also a valid molecule).   Although the molecules look very similar their formulas have levenshtein distance ~90-100.</p>\n<p>Some examples:</p>\n<p><strong>image_id:</strong> 0c21fc8b97b9<br>\n<strong>train label:</strong> InChI=1S/C32H34ClN5O2S/c1-3-40-28-12-8-7-11-27(28)37-17-19-38(20-18-37)31(39)26-15-13-25(14-16-26)23-41-32-34-29(33)21-30(35-32)36(2)22-24-9-5-4-6-10-24/h4-16,21H,3,17-20,22-23H2,1-2H3  <br>\n<strong>model prediction:</strong> InChI=1S/C31H32ClN5O2S/c1-35(21-23-8-4-3-5-9-23)29-20-28(32)33-31(34-29)40-22-24-12-14-25(15-13-24)30(38)37-18-16-36(17-19-37)26-10-6-7-11-27(26)39-2/h3-15,20H,16-19,21-22H2,1-2H3<br>\n<strong>levenshtein distance:</strong> 101</p>\n<table>\n<thead>\n<tr>\n<th>original from train</th>\n<th>generated from train label</th>\n<th>my prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112207876-bb118480-8c28-11eb-91d7-2bb3c8ef17c0.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112208112-088df180-8c29-11eb-8cc0-c2717adee033.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112208239-31ae8200-8c29-11eb-8981-ee6c683504f9.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>image_id:</strong> 859ee196bdf4<br>\n<strong>train label:</strong> InChI=1S/C31H37Cl2N3O4S/c1-6-22(4)34-31(38)23(5)35(19-25-14-15-26(32)18-28(25)33)30(37)20-36(29-11-9-8-10-24(29)7-2)41(39,40)27-16-12-21(3)13-17-27/h8-18,22-23H,6-7,19-20H2,1-5H3,(H,34,38)<br>\n<strong>model prediction:</strong> InChI=1S/C30H35Cl2N3O4S/c1-6-23-9-7-8-10-28(23)35(40(38,39)26-15-11-21(4)12-16-26)19-29(36)34(22(5)30(37)33-20(2)3)18-24-13-14-25(31)17-27(24)32/h7-17,20,22H,6,18-19H2,1-5H3,(H,33,37)<br>\n<strong>levenshtein distance:</strong> 99</p>\n<table>\n<thead>\n<tr>\n<th>original from train</th>\n<th>generated from train label</th>\n<th>my prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112209750-dc737000-8c2a-11eb-95a8-593e3bf39222.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112209837-fca32f00-8c2a-11eb-8496-4e0e6314a146.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210370-a84c7f00-8c2b-11eb-93a4-092cdf28bdc3.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>image_id:</strong> 5ff1da202c41<br>\n<strong>train label:</strong> InChI=1S/C27H27FN6O/c1-3-24-30-23-12-11-22(19-7-9-20(28)10-8-19)31-25(23)26(32-24)33-13-15-34(16-14-33)27(35)29-21-6-4-5-18(2)17-21/h4-12,17H,3,13-16H2,1-2H3,(H,29,35)<br>\n<strong>model prediction:</strong> InChI=1S/C26H25FN6O/c1-17-4-3-5-21(16-17)30-26(34)33-14-12-32(13-15-33)25-24-23(28-18(2)29-25)11-10-22(31-24)19-6-8-20(27)9-7-19/h3-11,16H,12-15H2,1-2H3,(H,30,34)<br>\n<strong>levenshtein distance:</strong> 88</p>\n<table>\n<thead>\n<tr>\n<th>original from train</th>\n<th>generated from train label</th>\n<th>my prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210815-3a548780-8c2c-11eb-96e6-650bf231f160.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210900-4fc9b180-8c2c-11eb-8c40-fba49261abb2.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210979-6839cc00-8c2c-11eb-93d7-1c1bacca8428.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<p>Other relevant discussions:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224275\" target=\"_blank\">Attention: Errors in the dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223230\" target=\"_blank\">Low quality of images</a></li>\n</ul>",
  "messages": [
    {
      "id": "1250181",
      "postDate": "03/23/2021 20:22:25",
      "content": "<p>Looking through model prediction errors on validation set, I found that my model actually returns more reasonable labels than training labels themselves. If you look closer at pubchem images and compare them to the training image, you'll notice that some bonds are missing, and the model predicts formula that corresponds to the image without those bonds (which is also a valid molecule).   Although the molecules look very similar their formulas have levenshtein distance ~90-100.</p>\n<p>Some examples:</p>\n<p><strong>image_id:</strong> 0c21fc8b97b9<br>\n<strong>train label:</strong> InChI=1S/C32H34ClN5O2S/c1-3-40-28-12-8-7-11-27(28)37-17-19-38(20-18-37)31(39)26-15-13-25(14-16-26)23-41-32-34-29(33)21-30(35-32)36(2)22-24-9-5-4-6-10-24/h4-16,21H,3,17-20,22-23H2,1-2H3  <br>\n<strong>model prediction:</strong> InChI=1S/C31H32ClN5O2S/c1-35(21-23-8-4-3-5-9-23)29-20-28(32)33-31(34-29)40-22-24-12-14-25(15-13-24)30(38)37-18-16-36(17-19-37)26-10-6-7-11-27(26)39-2/h3-15,20H,16-19,21-22H2,1-2H3<br>\n<strong>levenshtein distance:</strong> 101</p>\n<table>\n<thead>\n<tr>\n<th>original from train</th>\n<th>generated from train label</th>\n<th>my prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112207876-bb118480-8c28-11eb-91d7-2bb3c8ef17c0.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112208112-088df180-8c29-11eb-8cc0-c2717adee033.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112208239-31ae8200-8c29-11eb-8981-ee6c683504f9.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>image_id:</strong> 859ee196bdf4<br>\n<strong>train label:</strong> InChI=1S/C31H37Cl2N3O4S/c1-6-22(4)34-31(38)23(5)35(19-25-14-15-26(32)18-28(25)33)30(37)20-36(29-11-9-8-10-24(29)7-2)41(39,40)27-16-12-21(3)13-17-27/h8-18,22-23H,6-7,19-20H2,1-5H3,(H,34,38)<br>\n<strong>model prediction:</strong> InChI=1S/C30H35Cl2N3O4S/c1-6-23-9-7-8-10-28(23)35(40(38,39)26-15-11-21(4)12-16-26)19-29(36)34(22(5)30(37)33-20(2)3)18-24-13-14-25(31)17-27(24)32/h7-17,20,22H,6,18-19H2,1-5H3,(H,33,37)<br>\n<strong>levenshtein distance:</strong> 99</p>\n<table>\n<thead>\n<tr>\n<th>original from train</th>\n<th>generated from train label</th>\n<th>my prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112209750-dc737000-8c2a-11eb-95a8-593e3bf39222.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112209837-fca32f00-8c2a-11eb-8496-4e0e6314a146.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210370-a84c7f00-8c2b-11eb-93a4-092cdf28bdc3.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>image_id:</strong> 5ff1da202c41<br>\n<strong>train label:</strong> InChI=1S/C27H27FN6O/c1-3-24-30-23-12-11-22(19-7-9-20(28)10-8-19)31-25(23)26(32-24)33-13-15-34(16-14-33)27(35)29-21-6-4-5-18(2)17-21/h4-12,17H,3,13-16H2,1-2H3,(H,29,35)<br>\n<strong>model prediction:</strong> InChI=1S/C26H25FN6O/c1-17-4-3-5-21(16-17)30-26(34)33-14-12-32(13-15-33)25-24-23(28-18(2)29-25)11-10-22(31-24)19-6-8-20(27)9-7-19/h3-11,16H,12-15H2,1-2H3,(H,30,34)<br>\n<strong>levenshtein distance:</strong> 88</p>\n<table>\n<thead>\n<tr>\n<th>original from train</th>\n<th>generated from train label</th>\n<th>my prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210815-3a548780-8c2c-11eb-96e6-650bf231f160.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210900-4fc9b180-8c2c-11eb-8c40-fba49261abb2.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112210979-6839cc00-8c2c-11eb-93d7-1c1bacca8428.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<p>Other relevant discussions:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224275\" target=\"_blank\">Attention: Errors in the dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223230\" target=\"_blank\">Low quality of images</a></li>\n</ul>",
      "rawMarkdown": "Looking through model prediction errors on validation set, I found that my model actually returns more reasonable labels than training labels themselves. If you look closer at pubchem images and compare them to the training image, you'll notice that some bonds are missing, and the model predicts formula that corresponds to the image without those bonds (which is also a valid molecule).   Although the molecules look very similar their formulas have levenshtein distance ~90-100.\n\nSome examples:\n\n**image_id:** 0c21fc8b97b9\n**train label:** InChI=1S/C32H34ClN5O2S/c1-3-40-28-12-8-7-11-27(28)37-17-19-38(20-18-37)31(39)26-15-13-25(14-16-26)23-41-32-34-29(33)21-30(35-32)36(2)22-24-9-5-4-6-10-24/h4-16,21H,3,17-20,22-23H2,1-2H3  \n**model prediction:** InChI=1S/C31H32ClN5O2S/c1-35(21-23-8-4-3-5-9-23)29-20-28(32)33-31(34-29)40-22-24-12-14-25(15-13-24)30(38)37-18-16-36(17-19-37)26-10-6-7-11-27(26)39-2/h3-15,20H,16-19,21-22H2,1-2H3\n**levenshtein distance:** 101\n| original from train | generated from train label | my prediction |\n| --- | --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112207876-bb118480-8c28-11eb-91d7-2bb3c8ef17c0.png) | ![image](https://user-images.githubusercontent.com/4602302/112208112-088df180-8c29-11eb-8cc0-c2717adee033.png) | ![image](https://user-images.githubusercontent.com/4602302/112208239-31ae8200-8c29-11eb-8981-ee6c683504f9.png)|\n\n----\n\n**image_id:** 859ee196bdf4\n**train label:** InChI=1S/C31H37Cl2N3O4S/c1-6-22(4)34-31(38)23(5)35(19-25-14-15-26(32)18-28(25)33)30(37)20-36(29-11-9-8-10-24(29)7-2)41(39,40)27-16-12-21(3)13-17-27/h8-18,22-23H,6-7,19-20H2,1-5H3,(H,34,38)\n**model prediction:** InChI=1S/C30H35Cl2N3O4S/c1-6-23-9-7-8-10-28(23)35(40(38,39)26-15-11-21(4)12-16-26)19-29(36)34(22(5)30(37)33-20(2)3)18-24-13-14-25(31)17-27(24)32/h7-17,20,22H,6,18-19H2,1-5H3,(H,33,37)\n**levenshtein distance:** 99\n| original from train | generated from train label | my prediction |\n| --- | --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112209750-dc737000-8c2a-11eb-95a8-593e3bf39222.png) | ![image](https://user-images.githubusercontent.com/4602302/112209837-fca32f00-8c2a-11eb-8496-4e0e6314a146.png) | ![image](https://user-images.githubusercontent.com/4602302/112210370-a84c7f00-8c2b-11eb-93a4-092cdf28bdc3.png) |\n\n----\n\n**image_id:** 5ff1da202c41\n**train label:** InChI=1S/C27H27FN6O/c1-3-24-30-23-12-11-22(19-7-9-20(28)10-8-19)31-25(23)26(32-24)33-13-15-34(16-14-33)27(35)29-21-6-4-5-18(2)17-21/h4-12,17H,3,13-16H2,1-2H3,(H,29,35)\n**model prediction:** InChI=1S/C26H25FN6O/c1-17-4-3-5-21(16-17)30-26(34)33-14-12-32(13-15-33)25-24-23(28-18(2)29-25)11-10-22(31-24)19-6-8-20(27)9-7-19/h3-11,16H,12-15H2,1-2H3,(H,30,34)\n**levenshtein distance:** 88\n| original from train | generated from train label | my prediction |\n| --- | --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112210815-3a548780-8c2c-11eb-96e6-650bf231f160.png) | ![image](https://user-images.githubusercontent.com/4602302/112210900-4fc9b180-8c2c-11eb-8c40-fba49261abb2.png) | ![image](https://user-images.githubusercontent.com/4602302/112210979-6839cc00-8c2c-11eb-93d7-1c1bacca8428.png)|\n\nOther relevant discussions:\n- [Attention: Errors in the dataset](https://www.kaggle.com/c/bms-molecular-translation/discussion/224275)\n- [Low quality of images](https://www.kaggle.com/c/bms-molecular-translation/discussion/223230)",
      "votes": null
    },
    {
      "id": "1250448",
      "postDate": "03/24/2021 03:52:22",
      "content": "<p>As far as I know, it happens in most of the competitions. If there are not particularly many wrong labels, the host will not take action on them.</p>",
      "rawMarkdown": "As far as I know, it happens in most of the competitions. If there are not particularly many wrong labels, the host will not take action on them.",
      "votes": null
    },
    {
      "id": "1250531",
      "postDate": "03/24/2021 05:23:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a>, <br>\nDo you think by removing such images from the data we can improve the model?</p>",
      "rawMarkdown": "Hi @stassl, \nDo you think by removing such images from the data we can improve the model?",
      "votes": null
    },
    {
      "id": "1250551",
      "postDate": "03/24/2021 05:52:37",
      "content": "<p>For these examples, original image &amp; label generated image &amp; prediction generated image look same for me.</p>\n<ul>\n<li><a href=\"https://drive.google.com/file/d/1jKPJlBQ5gevcRnID4B6z3Y4jv-EQWOVi/view?usp=sharing\" target=\"_blank\">example1</a><ul>\n<li>file_path: ../input/bms-molecular-translation/train/a/a/f/aaf92f4b4be7.png</li>\n<li>label: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(13-24(33)27(25)26-18(4)11-17(3)12-19(26)5)20(6)15-36-23-9-7-22(8-10-23)29(30,31)32/h7-12,16,20-21H,13-15H2,1-6H3</li>\n<li>prediction: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(20(6)15-36-23-9-7-22(8-10-23)29(30,31)32)13-24(33)27(25)26-18(4)11-17(3)12-19(26)5/h7-12,16,20-21H,13-15H2,1-6H3</li>\n<li>levenshtein distance: 66</li></ul></li>\n<li><a href=\"https://drive.google.com/file/d/1gPvqSGlBQ2MXSkg2SOVVGOmRIrNklk0K/view?usp=sharing\" target=\"_blank\">example2</a><ul>\n<li>file_path: ../input/bms-molecular-translation/train/7/7/a/77a276f869b2.png</li>\n<li>label: InChI=1S/C22H27N3O5S2/c26-22(21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25)23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)</li>\n<li>prediction: InChI=1S/C22H27N3O5S2/c26-22(23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19)21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)</li>\n<li>levenshtein distance: 64</li></ul></li>\n</ul>\n<p>So even though model predicts same one, it could be error… 🤔?<br>\n(If they are not same, please kindly tell me)</p>",
      "rawMarkdown": "For these examples, original image & label generated image & prediction generated image look same for me.\n- [example1](https://drive.google.com/file/d/1jKPJlBQ5gevcRnID4B6z3Y4jv-EQWOVi/view?usp=sharing)\n  - file_path: ../input/bms-molecular-translation/train/a/a/f/aaf92f4b4be7.png\n  - label: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(13-24(33)27(25)26-18(4)11-17(3)12-19(26)5)20(6)15-36-23-9-7-22(8-10-23)29(30,31)32/h7-12,16,20-21H,13-15H2,1-6H3\n  - prediction: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(20(6)15-36-23-9-7-22(8-10-23)29(30,31)32)13-24(33)27(25)26-18(4)11-17(3)12-19(26)5/h7-12,16,20-21H,13-15H2,1-6H3\n  - levenshtein distance: 66\n- [example2](https://drive.google.com/file/d/1gPvqSGlBQ2MXSkg2SOVVGOmRIrNklk0K/view?usp=sharing)\n  - file_path: ../input/bms-molecular-translation/train/7/7/a/77a276f869b2.png\n  - label: InChI=1S/C22H27N3O5S2/c26-22(21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25)23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)\n  - prediction: InChI=1S/C22H27N3O5S2/c26-22(23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19)21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)\n  - levenshtein distance: 64\n\nSo even though model predicts same one, it could be error... 🤔?\n(If they are not same, please kindly tell me)",
      "votes": null
    },
    {
      "id": "1250604",
      "postDate": "03/24/2021 06:36:17",
      "content": "<p>Yeah, to me your examples also look 100% identical (visually). Can you paste text version of them here? Maybe worth looking at them at pubchem, maybe 3d structure? Did you try to convert them to mol and then back to inchi? I think there were a few labels that after  converting them to mol and then back to inchi produced different results.</p>",
      "rawMarkdown": "Yeah, to me your examples also look 100% identical (visually). Can you paste text version of them here? Maybe worth looking at them at pubchem, maybe 3d structure? Did you try to convert them to mol and then back to inchi? I think there were a few labels that after  converting them to mol and then back to inchi produced different results.",
      "votes": null
    },
    {
      "id": "1250632",
      "postDate": "03/24/2021 07:09:11",
      "content": "<blockquote>\n  <p>Can you paste text version of them here?</p>\n</blockquote>\n<p>I added text version!</p>\n<blockquote>\n  <p>Did you try to convert them to mol and then back to inchi?</p>\n</blockquote>\n<p>Yes, like below.</p>\n<pre><code>mol = rdkit.Chem.MolFromInchi(pred)\npred_image = rdkit.Chem.Draw.MolToImage(mol, size=(w, h))\n</code></pre>",
      "rawMarkdown": "> Can you paste text version of them here?\n\nI added text version!\n\n> Did you try to convert them to mol and then back to inchi?\n\nYes, like below.\n\n```\nmol = rdkit.Chem.MolFromInchi(pred)\npred_image = rdkit.Chem.Draw.MolToImage(mol, size=(w, h))\n```",
      "votes": null
    },
    {
      "id": "1250684",
      "postDate": "03/24/2021 07:49:30",
      "content": "<p>I meant like this:</p>\n<pre><code>Chem.MolToInchi(Chem.MolFromInchi(pred_label)) == train_label\n</code></pre>\n<p>I applied this transformation to your predictions and it produced same inchi as training labels. For some reason your model predicts not canonical inchis sometimes. Probably you can gain some points by normalizing them manually.</p>",
      "rawMarkdown": "I meant like this:\n```\nChem.MolToInchi(Chem.MolFromInchi(pred_label)) == train_label\n```\nI applied this transformation to your predictions and it produced same inchi as training labels. For some reason your model predicts not canonical inchis sometimes. Probably you can gain some points by normalizing them manually.",
      "votes": null
    },
    {
      "id": "1250699",
      "postDate": "03/24/2021 08:04:43",
      "content": "<p>That's a reasonable idea (or maybe re-label them), though for now I'm not sure how many of them exist. Hopefully, as training dataset is large the effect of those images is negligible.</p>",
      "rawMarkdown": "That's a reasonable idea (or maybe re-label them), though for now I'm not sure how many of them exist. Hopefully, as training dataset is large the effect of those images is negligible.",
      "votes": null
    },
    {
      "id": "1250700",
      "postDate": "03/24/2021 08:05:17",
      "content": "<p>Wow, as you say I confirmed this transformation gives me same inchi as training labels.<br>\nThanks for sharing! So then we should apply this transformation when monitoring validation and submission.</p>",
      "rawMarkdown": "Wow, as you say I confirmed this transformation gives me same inchi as training labels.\nThanks for sharing! So then we should apply this transformation when monitoring validation and submission.",
      "votes": null
    },
    {
      "id": "1251814",
      "postDate": "03/25/2021 06:53:46",
      "content": "<p><a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> , could you please explain what you mean by this…</p>\n<blockquote>\n  <p>Probably you can gain some points by normalizing them manually</p>\n</blockquote>",
      "rawMarkdown": "stassl , could you please explain what you mean by this...\n> Probably you can gain some points by normalizing them manually",
      "votes": null
    },
    {
      "id": "1251941",
      "postDate": "03/25/2021 09:22:04",
      "content": "<p>Holy.. this is gold for free.<br>\n<a href=\"https://www.kaggle.com/nitindatta\" target=\"_blank\">@nitindatta</a><br>\n<strong>predicted inchi</strong> describes the <strong>same molecule</strong> as its <strong>label</strong>, but the inchis <strong>differ</strong> -&gt; levenshtein can be huge<br>\nBut it seems that using rdkit: \"correctly\" predicted inchi -&gt; molfrominchi; moltoinchi -&gt; \"true\" inchi</p>",
      "rawMarkdown": "Holy.. this is gold for free.\n@nitindatta\n**predicted inchi** describes the **same molecule** as its **label**, but the inchis **differ** -> levenshtein can be huge\nBut it seems that using rdkit: \"correctly\" predicted inchi -> molfrominchi; moltoinchi -> \"true\" inchi",
      "votes": null
    },
    {
      "id": "1252192",
      "postDate": "03/25/2021 13:43:26",
      "content": "<p>They are not the wrong label. There are roughly 10% images in the training set (not sure about test set) belongs to this alternative-structure-possible category. It's due to</p>\n<ul>\n<li>either a methyl group attached to a carbon but lose the bond <em>due to compression</em> <code>1ab069aad819</code></li>\n<li>or extremely fuzzy atom depiction (e.g., F looks like I; I completely lose the line <code>0bb3c2bdcc2c</code>)</li>\n<li>There are other structures (roughly 2-3%) that even domain expert cannot recover the structure without context, due to crossing bonds and overlapped depiction. e.g., <code>9203b1e7df7b</code> <code>d88774e5cc9f</code></li>\n</ul>\n<p>If one sticks to making <strong>valid</strong> InChI prediction, the above 2 scenarios will significantly inflate the average LD, because the atom indexes will shift (10-&gt; 11, 11-&gt;12, so forth).</p>\n<p>The LB ~ 2.1 is amazing, but if any model manages to get below 2, I am more confident it's overfitting to the depiction style, given the aforementioned edge cases.</p>\n<p></p>",
      "rawMarkdown": "They are not the wrong label. There are roughly 10% images in the training set (not sure about test set) belongs to this alternative-structure-possible category. It's due to\n* either a methyl group attached to a carbon but lose the bond *due to compression* `1ab069aad819`\n* or extremely fuzzy atom depiction (e.g., F looks like I; I completely lose the line `0bb3c2bdcc2c`)\n* There are other structures (roughly 2-3%) that even domain expert cannot recover the structure without context, due to crossing bonds and overlapped depiction. e.g., `9203b1e7df7b` `d88774e5cc9f`\n\nIf one sticks to making **valid** InChI prediction, the above 2 scenarios will significantly inflate the average LD, because the atom indexes will shift (10-> 11, 11->12, so forth).\n\nThe LB ~ 2.1 is amazing, but if any model manages to get below 2, I am more confident it's overfitting to the depiction style, given the aforementioned edge cases.\n\n~~Of course, kaggle competition doesn't care...~~",
      "votes": null
    },
    {
      "id": "1252245",
      "postDate": "03/25/2021 14:28:09",
      "content": "<p>I don't think it's negligible. I checked 177 predictions in which I cannot find a PubChem match from a random 3k benchmark. I find</p>\n<ul>\n<li>35 images have missing methyl group attached to carbon</li>\n<li>20 images have extremely fuzzy atom depiction</li>\n</ul>\n<p>This already accounts for ~2% out of the 3,000. I am not even checking the rest 2833 which have PubChem matches. Each alternative structure may inflate the LD by ~40. Therefore, at least 40% LD are due to these alternative structures, as the average LD is approaching 2. </p>",
      "rawMarkdown": "I don't think it's negligible. I checked 177 predictions in which I cannot find a PubChem match from a random 3k benchmark. I find\n* 35 images have missing methyl group attached to carbon\n* 20 images have extremely fuzzy atom depiction\n\nThis already accounts for ~2% out of the 3,000. I am not even checking the rest 2833 which have PubChem matches. Each alternative structure may inflate the LD by ~40. Therefore, at least 40% LD are due to these alternative structures, as the average LD is approaching 2.",
      "votes": null
    },
    {
      "id": "1252353",
      "postDate": "03/25/2021 15:57:42",
      "content": "<p><a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a> thanks for chiming in. I've seen your topic <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">Problem with the evaluation metric</a> and this post is just another evidence of the problem you mentioned. </p>\n<p>Maybe the title of my post is a bit inaccurate in favor of catchiness - I didn't mean the labels are wrong. Or better to say I understand that labels were correct originally, but as result of some transformations (corruption/compression/resizing - intentional or not) they are no longer matching the image. </p>\n<p>As you say those corruptions may fall into several categories. Let’s maybe omit for now images that are convoluted, with multiple crossing bonds and complicated 3D-structure. If we consider rather simple molecules, and suppose it lost some bonds due to corruptions. I would say there could be following cases how it affects original label:</p>\n<ul>\n<li>Corruption leads to an invalid molecule, but it can be uniquely recovered by domain expert, like when some bond is lost in the middle of molecule and you know it should be there. Then I would say the original label is still correct and I believe NN can learn to recover them as they are trained for language modeling when some words are masked.</li>\n<li>Corruption leads to an invalid molecule, but it can be recovered to multiple valid ones. In that case model can predict distribution over possible labels and original label will be one of them.</li>\n<li>Corruption leads to a VALID molecule but this molecule is different from original, thus the label becomes irrelevant. Particularly in my post I wanted to pay attention to this case. </li>\n</ul>\n<blockquote>\n  <p>either a methyl group attached to a carbon but lose the bond&nbsp;due to compression&nbsp;1ab069aad819</p>\n</blockquote>\n<p>I’m attaching the image for convenience that your referenced:</p>\n<table>\n<thead>\n<tr>\n<th>train image</th>\n<th>generated from label</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112497893-13ff2b00-8d97-11eb-8a2f-e6040efd3d36.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112497943-1e212980-8d97-11eb-871f-26386dd67910.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<p>Its original label is <strong>InChI=1S/C10H10O4/c1-5-8(11)3-6-7(9(5)13-2)4-14-10(6)12/h3,11H,4H2,1-2H3</strong>, but when you remove a single bond you’ll get also a valid molecule with label <strong>InChI=1S/C9H8O4/c1-12-8-3-5(10)2-6-7(8)4-13-9(6)11/h2-3,10H,4H2,1H3</strong> which is <em>29</em> edit operation away from the original. And I believe that is what my model would predict and any domain expert that will see only the corrupted image without that bond, because why anyone would add any bonds if the molecule is already valid. </p>\n<p>Another interesting question that you brought up is what bonds can be lost and why. I was thinking that any bond can be dropped to methyl group, as you say, or somewhere in the middle or any other place. Do you observe that bonds can be missing only to methyl groups (is that equivalent to say at the border of molecules, but not in the middle, sorry my chemistry not so good)?</p>\n<p>UPD:<br>\nOk, considering my last question about where bonds can be dropped, I found confirmation of my initial thought that they can be dropped anywhere, including in the middle of molecule, and not only to methyl groups:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/112535690-a4e8fd00-8dbd-11eb-8984-11d5dfe7565d.png\" alt=\"image\"><br>\n<img src=\"https://user-images.githubusercontent.com/4602302/112536280-4f612000-8dbe-11eb-819a-53318cbf1bfe.png\" alt=\"image\"></p>",
      "rawMarkdown": "houndcl thanks for chiming in. I've seen your topic [Problem with the evaluation metric](https://www.kaggle.com/c/bms-molecular-translation/discussion/224394) and this post is just another evidence of the problem you mentioned. \n\nMaybe the title of my post is a bit inaccurate in favor of catchiness - I didn't mean the labels are wrong. Or better to say I understand that labels were correct originally, but as result of some transformations (corruption/compression/resizing - intentional or not) they are no longer matching the image. \n\nAs you say those corruptions may fall into several categories. Let’s maybe omit for now images that are convoluted, with multiple crossing bonds and complicated 3D-structure. If we consider rather simple molecules, and suppose it lost some bonds due to corruptions. I would say there could be following cases how it affects original label:\n\n- Corruption leads to an invalid molecule, but it can be uniquely recovered by domain expert, like when some bond is lost in the middle of molecule and you know it should be there. Then I would say the original label is still correct and I believe NN can learn to recover them as they are trained for language modeling when some words are masked.\n- Corruption leads to an invalid molecule, but it can be recovered to multiple valid ones. In that case model can predict distribution over possible labels and original label will be one of them.\n- Corruption leads to a VALID molecule but this molecule is different from original, thus the label becomes irrelevant. Particularly in my post I wanted to pay attention to this case. \n\n> either a methyl group attached to a carbon but lose the bond due to compression 1ab069aad819\n\nI’m attaching the image for convenience that your referenced:\n| train image | generated from label |\n| --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112497893-13ff2b00-8d97-11eb-8a2f-e6040efd3d36.png) | ![image](https://user-images.githubusercontent.com/4602302/112497943-1e212980-8d97-11eb-871f-26386dd67910.png) |\n\nIts original label is **InChI=1S/C10H10O4/c1-5-8(11)3-6-7(9(5)13-2)4-14-10(6)12/h3,11H,4H2,1-2H3**, but when you remove a single bond you’ll get also a valid molecule with label **InChI=1S/C9H8O4/c1-12-8-3-5(10)2-6-7(8)4-13-9(6)11/h2-3,10H,4H2,1H3** which is *29* edit operation away from the original. And I believe that is what my model would predict and any domain expert that will see only the corrupted image without that bond, because why anyone would add any bonds if the molecule is already valid. \n\nAnother interesting question that you brought up is what bonds can be lost and why. I was thinking that any bond can be dropped to methyl group, as you say, or somewhere in the middle or any other place. Do you observe that bonds can be missing only to methyl groups (is that equivalent to say at the border of molecules, but not in the middle, sorry my chemistry not so good)?\n\nUPD:\nOk, considering my last question about where bonds can be dropped, I found confirmation of my initial thought that they can be dropped anywhere, including in the middle of molecule, and not only to methyl groups:\n![image](https://user-images.githubusercontent.com/4602302/112535690-a4e8fd00-8dbd-11eb-8984-11d5dfe7565d.png)\n![image](https://user-images.githubusercontent.com/4602302/112536280-4f612000-8dbe-11eb-819a-53318cbf1bfe.png)",
      "votes": null
    },
    {
      "id": "1252409",
      "postDate": "03/25/2021 17:00:13",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> for sharing. I think maybe due to some representation missed on the training features.</p>",
      "rawMarkdown": "Thank you @stassl for sharing. I think maybe due to some representation missed on the training features.",
      "votes": null
    },
    {
      "id": "1259579",
      "postDate": "04/01/2021 14:15:29",
      "content": "<p>You did not submit anything with such a good model?  Also, what do yu used to create your images?  rdkit gives me different ones.</p>",
      "rawMarkdown": "You did not submit anything with such a good model?  Also, what do yu used to create your images?  rdkit gives me different ones.",
      "votes": null
    },
    {
      "id": "1259680",
      "postDate": "04/01/2021 15:22:29",
      "content": "<p>Not yet, as it takes too long to generate predictions for 1.6M images, for now working on improving my local CV score. </p>\n<p>Well, I just cherry-picked some examples where it was good, unfortunately there are lots of them where it is not so good, especially for convoluted or very large molecules. </p>\n<p>The images of molecules I took from pubchem site.</p>",
      "rawMarkdown": "Not yet, as it takes too long to generate predictions for 1.6M images, for now working on improving my local CV score. \n\nWell, I just cherry-picked some examples where it was good, unfortunately there are lots of them where it is not so good, especially for convoluted or very large molecules. \n\nThe images of molecules I took from pubchem site.",
      "votes": null
    },
    {
      "id": "1259733",
      "postDate": "04/01/2021 16:05:14",
      "content": "<p>Thanks for the reply and good luck.</p>",
      "rawMarkdown": "Thanks for the reply and good luck.",
      "votes": null
    },
    {
      "id": "1264187",
      "postDate": "04/06/2021 01:04:55",
      "content": "<p>Some InChIs can not be parsed by RDKit and cause segmentation faults 😲</p>\n<blockquote>\n  <p>Process finished with exit code 139 (interrupted by signal 11: SIGSEGV)</p>\n</blockquote>",
      "rawMarkdown": "Some InChIs can not be parsed by RDKit and cause segmentation faults 😲\n> Process finished with exit code 139 (interrupted by signal 11: SIGSEGV)",
      "votes": null
    },
    {
      "id": "1264250",
      "postDate": "04/06/2021 02:43:51",
      "content": "<p>There is a tool I made to do this:<br>\n<a href=\"https://www.kaggle.com/nofreewill/normalize-your-predictions\" target=\"_blank\">https://www.kaggle.com/nofreewill/normalize-your-predictions</a></p>",
      "rawMarkdown": "There is a tool I made to do this:\nhttps://www.kaggle.com/nofreewill/normalize-your-predictions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1250448,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "03/24/2021 03:52:22",
      "content": "<p>As far as I know, it happens in most of the competitions. If there are not particularly many wrong labels, the host will not take action on them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1250531,
      "author_name": "nitindatta",
      "author_url": "",
      "post_date": "03/24/2021 05:23:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a>, <br>\nDo you think by removing such images from the data we can improve the model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1250699,
          "author_name": "stassl",
          "author_url": "",
          "post_date": "03/24/2021 08:04:43",
          "content": "<p>That's a reasonable idea (or maybe re-label them), though for now I'm not sure how many of them exist. Hopefully, as training dataset is large the effect of those images is negligible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252245,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "03/25/2021 14:28:09",
          "content": "<p>I don't think it's negligible. I checked 177 predictions in which I cannot find a PubChem match from a random 3k benchmark. I find</p>\n<ul>\n<li>35 images have missing methyl group attached to carbon</li>\n<li>20 images have extremely fuzzy atom depiction</li>\n</ul>\n<p>This already accounts for ~2% out of the 3,000. I am not even checking the rest 2833 which have PubChem matches. Each alternative structure may inflate the LD by ~40. Therefore, at least 40% LD are due to these alternative structures, as the average LD is approaching 2. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1250551,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "03/24/2021 05:52:37",
      "content": "<p>For these examples, original image &amp; label generated image &amp; prediction generated image look same for me.</p>\n<ul>\n<li><a href=\"https://drive.google.com/file/d/1jKPJlBQ5gevcRnID4B6z3Y4jv-EQWOVi/view?usp=sharing\" target=\"_blank\">example1</a><ul>\n<li>file_path: ../input/bms-molecular-translation/train/a/a/f/aaf92f4b4be7.png</li>\n<li>label: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(13-24(33)27(25)26-18(4)11-17(3)12-19(26)5)20(6)15-36-23-9-7-22(8-10-23)29(30,31)32/h7-12,16,20-21H,13-15H2,1-6H3</li>\n<li>prediction: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(20(6)15-36-23-9-7-22(8-10-23)29(30,31)32)13-24(33)27(25)26-18(4)11-17(3)12-19(26)5/h7-12,16,20-21H,13-15H2,1-6H3</li>\n<li>levenshtein distance: 66</li></ul></li>\n<li><a href=\"https://drive.google.com/file/d/1gPvqSGlBQ2MXSkg2SOVVGOmRIrNklk0K/view?usp=sharing\" target=\"_blank\">example2</a><ul>\n<li>file_path: ../input/bms-molecular-translation/train/7/7/a/77a276f869b2.png</li>\n<li>label: InChI=1S/C22H27N3O5S2/c26-22(21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25)23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)</li>\n<li>prediction: InChI=1S/C22H27N3O5S2/c26-22(23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19)21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)</li>\n<li>levenshtein distance: 64</li></ul></li>\n</ul>\n<p>So even though model predicts same one, it could be error… 🤔?<br>\n(If they are not same, please kindly tell me)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1250604,
          "author_name": "stassl",
          "author_url": "",
          "post_date": "03/24/2021 06:36:17",
          "content": "<p>Yeah, to me your examples also look 100% identical (visually). Can you paste text version of them here? Maybe worth looking at them at pubchem, maybe 3d structure? Did you try to convert them to mol and then back to inchi? I think there were a few labels that after  converting them to mol and then back to inchi produced different results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250632,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/24/2021 07:09:11",
          "content": "<blockquote>\n  <p>Can you paste text version of them here?</p>\n</blockquote>\n<p>I added text version!</p>\n<blockquote>\n  <p>Did you try to convert them to mol and then back to inchi?</p>\n</blockquote>\n<p>Yes, like below.</p>\n<pre><code>mol = rdkit.Chem.MolFromInchi(pred)\npred_image = rdkit.Chem.Draw.MolToImage(mol, size=(w, h))\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 1250684,
              "author_name": "stassl",
              "author_url": "",
              "post_date": "03/24/2021 07:49:30",
              "content": "<p>I meant like this:</p>\n<pre><code>Chem.MolToInchi(Chem.MolFromInchi(pred_label)) == train_label\n</code></pre>\n<p>I applied this transformation to your predictions and it produced same inchi as training labels. For some reason your model predicts not canonical inchis sometimes. Probably you can gain some points by normalizing them manually.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1250700,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/24/2021 08:05:17",
          "content": "<p>Wow, as you say I confirmed this transformation gives me same inchi as training labels.<br>\nThanks for sharing! So then we should apply this transformation when monitoring validation and submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251814,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "03/25/2021 06:53:46",
          "content": "<p><a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> , could you please explain what you mean by this…</p>\n<blockquote>\n  <p>Probably you can gain some points by normalizing them manually</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251941,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "03/25/2021 09:22:04",
          "content": "<p>Holy.. this is gold for free.<br>\n<a href=\"https://www.kaggle.com/nitindatta\" target=\"_blank\">@nitindatta</a><br>\n<strong>predicted inchi</strong> describes the <strong>same molecule</strong> as its <strong>label</strong>, but the inchis <strong>differ</strong> -&gt; levenshtein can be huge<br>\nBut it seems that using rdkit: \"correctly\" predicted inchi -&gt; molfrominchi; moltoinchi -&gt; \"true\" inchi</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264187,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "04/06/2021 01:04:55",
          "content": "<p>Some InChIs can not be parsed by RDKit and cause segmentation faults 😲</p>\n<blockquote>\n  <p>Process finished with exit code 139 (interrupted by signal 11: SIGSEGV)</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264250,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/06/2021 02:43:51",
          "content": "<p>There is a tool I made to do this:<br>\n<a href=\"https://www.kaggle.com/nofreewill/normalize-your-predictions\" target=\"_blank\">https://www.kaggle.com/nofreewill/normalize-your-predictions</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252192,
      "author_name": "houndcl",
      "author_url": "",
      "post_date": "03/25/2021 13:43:26",
      "content": "<p>They are not the wrong label. There are roughly 10% images in the training set (not sure about test set) belongs to this alternative-structure-possible category. It's due to</p>\n<ul>\n<li>either a methyl group attached to a carbon but lose the bond <em>due to compression</em> <code>1ab069aad819</code></li>\n<li>or extremely fuzzy atom depiction (e.g., F looks like I; I completely lose the line <code>0bb3c2bdcc2c</code>)</li>\n<li>There are other structures (roughly 2-3%) that even domain expert cannot recover the structure without context, due to crossing bonds and overlapped depiction. e.g., <code>9203b1e7df7b</code> <code>d88774e5cc9f</code></li>\n</ul>\n<p>If one sticks to making <strong>valid</strong> InChI prediction, the above 2 scenarios will significantly inflate the average LD, because the atom indexes will shift (10-&gt; 11, 11-&gt;12, so forth).</p>\n<p>The LB ~ 2.1 is amazing, but if any model manages to get below 2, I am more confident it's overfitting to the depiction style, given the aforementioned edge cases.</p>\n<p></p>",
      "votes": null,
      "replies": [
        {
          "id": 1252353,
          "author_name": "stassl",
          "author_url": "",
          "post_date": "03/25/2021 15:57:42",
          "content": "<p><a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a> thanks for chiming in. I've seen your topic <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224394\" target=\"_blank\">Problem with the evaluation metric</a> and this post is just another evidence of the problem you mentioned. </p>\n<p>Maybe the title of my post is a bit inaccurate in favor of catchiness - I didn't mean the labels are wrong. Or better to say I understand that labels were correct originally, but as result of some transformations (corruption/compression/resizing - intentional or not) they are no longer matching the image. </p>\n<p>As you say those corruptions may fall into several categories. Let’s maybe omit for now images that are convoluted, with multiple crossing bonds and complicated 3D-structure. If we consider rather simple molecules, and suppose it lost some bonds due to corruptions. I would say there could be following cases how it affects original label:</p>\n<ul>\n<li>Corruption leads to an invalid molecule, but it can be uniquely recovered by domain expert, like when some bond is lost in the middle of molecule and you know it should be there. Then I would say the original label is still correct and I believe NN can learn to recover them as they are trained for language modeling when some words are masked.</li>\n<li>Corruption leads to an invalid molecule, but it can be recovered to multiple valid ones. In that case model can predict distribution over possible labels and original label will be one of them.</li>\n<li>Corruption leads to a VALID molecule but this molecule is different from original, thus the label becomes irrelevant. Particularly in my post I wanted to pay attention to this case. </li>\n</ul>\n<blockquote>\n  <p>either a methyl group attached to a carbon but lose the bond&nbsp;due to compression&nbsp;1ab069aad819</p>\n</blockquote>\n<p>I’m attaching the image for convenience that your referenced:</p>\n<table>\n<thead>\n<tr>\n<th>train image</th>\n<th>generated from label</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112497893-13ff2b00-8d97-11eb-8a2f-e6040efd3d36.png\" alt=\"image\"></td>\n<td><img src=\"https://user-images.githubusercontent.com/4602302/112497943-1e212980-8d97-11eb-871f-26386dd67910.png\" alt=\"image\"></td>\n</tr>\n</tbody>\n</table>\n<p>Its original label is <strong>InChI=1S/C10H10O4/c1-5-8(11)3-6-7(9(5)13-2)4-14-10(6)12/h3,11H,4H2,1-2H3</strong>, but when you remove a single bond you’ll get also a valid molecule with label <strong>InChI=1S/C9H8O4/c1-12-8-3-5(10)2-6-7(8)4-13-9(6)11/h2-3,10H,4H2,1H3</strong> which is <em>29</em> edit operation away from the original. And I believe that is what my model would predict and any domain expert that will see only the corrupted image without that bond, because why anyone would add any bonds if the molecule is already valid. </p>\n<p>Another interesting question that you brought up is what bonds can be lost and why. I was thinking that any bond can be dropped to methyl group, as you say, or somewhere in the middle or any other place. Do you observe that bonds can be missing only to methyl groups (is that equivalent to say at the border of molecules, but not in the middle, sorry my chemistry not so good)?</p>\n<p>UPD:<br>\nOk, considering my last question about where bonds can be dropped, I found confirmation of my initial thought that they can be dropped anywhere, including in the middle of molecule, and not only to methyl groups:<br>\n<img src=\"https://user-images.githubusercontent.com/4602302/112535690-a4e8fd00-8dbd-11eb-8984-11d5dfe7565d.png\" alt=\"image\"><br>\n<img src=\"https://user-images.githubusercontent.com/4602302/112536280-4f612000-8dbe-11eb-819a-53318cbf1bfe.png\" alt=\"image\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252409,
      "author_name": "olusesiadebisi",
      "author_url": "",
      "post_date": "03/25/2021 17:00:13",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> for sharing. I think maybe due to some representation missed on the training features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1259579,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/01/2021 14:15:29",
      "content": "<p>You did not submit anything with such a good model?  Also, what do yu used to create your images?  rdkit gives me different ones.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1259680,
          "author_name": "stassl",
          "author_url": "",
          "post_date": "04/01/2021 15:22:29",
          "content": "<p>Not yet, as it takes too long to generate predictions for 1.6M images, for now working on improving my local CV score. </p>\n<p>Well, I just cherry-picked some examples where it was good, unfortunately there are lots of them where it is not so good, especially for convoluted or very large molecules. </p>\n<p>The images of molecules I took from pubchem site.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1259733,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/01/2021 16:05:14",
          "content": "<p>Thanks for the reply and good luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1250181": "Looking through model prediction errors on validation set, I found that my model actually returns more reasonable labels than training labels themselves. If you look closer at pubchem images and compare them to the training image, you'll notice that some bonds are missing, and the model predicts formula that corresponds to the image without those bonds (which is also a valid molecule).   Although the molecules look very similar their formulas have levenshtein distance ~90-100.\n\nSome examples:\n\n**image_id:** 0c21fc8b97b9\n**train label:** InChI=1S/C32H34ClN5O2S/c1-3-40-28-12-8-7-11-27(28)37-17-19-38(20-18-37)31(39)26-15-13-25(14-16-26)23-41-32-34-29(33)21-30(35-32)36(2)22-24-9-5-4-6-10-24/h4-16,21H,3,17-20,22-23H2,1-2H3  \n**model prediction:** InChI=1S/C31H32ClN5O2S/c1-35(21-23-8-4-3-5-9-23)29-20-28(32)33-31(34-29)40-22-24-12-14-25(15-13-24)30(38)37-18-16-36(17-19-37)26-10-6-7-11-27(26)39-2/h3-15,20H,16-19,21-22H2,1-2H3\n**levenshtein distance:** 101\n| original from train | generated from train label | my prediction |\n| --- | --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112207876-bb118480-8c28-11eb-91d7-2bb3c8ef17c0.png) | ![image](https://user-images.githubusercontent.com/4602302/112208112-088df180-8c29-11eb-8cc0-c2717adee033.png) | ![image](https://user-images.githubusercontent.com/4602302/112208239-31ae8200-8c29-11eb-8981-ee6c683504f9.png)|\n\n----\n\n**image_id:** 859ee196bdf4\n**train label:** InChI=1S/C31H37Cl2N3O4S/c1-6-22(4)34-31(38)23(5)35(19-25-14-15-26(32)18-28(25)33)30(37)20-36(29-11-9-8-10-24(29)7-2)41(39,40)27-16-12-21(3)13-17-27/h8-18,22-23H,6-7,19-20H2,1-5H3,(H,34,38)\n**model prediction:** InChI=1S/C30H35Cl2N3O4S/c1-6-23-9-7-8-10-28(23)35(40(38,39)26-15-11-21(4)12-16-26)19-29(36)34(22(5)30(37)33-20(2)3)18-24-13-14-25(31)17-27(24)32/h7-17,20,22H,6,18-19H2,1-5H3,(H,33,37)\n**levenshtein distance:** 99\n| original from train | generated from train label | my prediction |\n| --- | --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112209750-dc737000-8c2a-11eb-95a8-593e3bf39222.png) | ![image](https://user-images.githubusercontent.com/4602302/112209837-fca32f00-8c2a-11eb-8496-4e0e6314a146.png) | ![image](https://user-images.githubusercontent.com/4602302/112210370-a84c7f00-8c2b-11eb-93a4-092cdf28bdc3.png) |\n\n----\n\n**image_id:** 5ff1da202c41\n**train label:** InChI=1S/C27H27FN6O/c1-3-24-30-23-12-11-22(19-7-9-20(28)10-8-19)31-25(23)26(32-24)33-13-15-34(16-14-33)27(35)29-21-6-4-5-18(2)17-21/h4-12,17H,3,13-16H2,1-2H3,(H,29,35)\n**model prediction:** InChI=1S/C26H25FN6O/c1-17-4-3-5-21(16-17)30-26(34)33-14-12-32(13-15-33)25-24-23(28-18(2)29-25)11-10-22(31-24)19-6-8-20(27)9-7-19/h3-11,16H,12-15H2,1-2H3,(H,30,34)\n**levenshtein distance:** 88\n| original from train | generated from train label | my prediction |\n| --- | --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112210815-3a548780-8c2c-11eb-96e6-650bf231f160.png) | ![image](https://user-images.githubusercontent.com/4602302/112210900-4fc9b180-8c2c-11eb-8c40-fba49261abb2.png) | ![image](https://user-images.githubusercontent.com/4602302/112210979-6839cc00-8c2c-11eb-93d7-1c1bacca8428.png)|\n\nOther relevant discussions:\n- [Attention: Errors in the dataset](https://www.kaggle.com/c/bms-molecular-translation/discussion/224275)\n- [Low quality of images](https://www.kaggle.com/c/bms-molecular-translation/discussion/223230)",
    "1250448": "As far as I know, it happens in most of the competitions. If there are not particularly many wrong labels, the host will not take action on them.",
    "1250531": "Hi @stassl, \nDo you think by removing such images from the data we can improve the model?",
    "1250551": "For these examples, original image & label generated image & prediction generated image look same for me.\n- [example1](https://drive.google.com/file/d/1jKPJlBQ5gevcRnID4B6z3Y4jv-EQWOVi/view?usp=sharing)\n  - file_path: ../input/bms-molecular-translation/train/a/a/f/aaf92f4b4be7.png\n  - label: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(13-24(33)27(25)26-18(4)11-17(3)12-19(26)5)20(6)15-36-23-9-7-22(8-10-23)29(30,31)32/h7-12,16,20-21H,13-15H2,1-6H3\n  - prediction: InChI=1S/C29H33F3O3S/c1-16(2)28(34)35-25-14-21(20(6)15-36-23-9-7-22(8-10-23)29(30,31)32)13-24(33)27(25)26-18(4)11-17(3)12-19(26)5/h7-12,16,20-21H,13-15H2,1-6H3\n  - levenshtein distance: 66\n- [example2](https://drive.google.com/file/d/1gPvqSGlBQ2MXSkg2SOVVGOmRIrNklk0K/view?usp=sharing)\n  - file_path: ../input/bms-molecular-translation/train/7/7/a/77a276f869b2.png\n  - label: InChI=1S/C22H27N3O5S2/c26-22(21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25)23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)\n  - prediction: InChI=1S/C22H27N3O5S2/c26-22(23-16-17-6-10-20(11-7-17)32(29,30)24-19-8-9-19)21(18-4-2-1-3-5-18)25-12-14-31(27,28)15-13-25/h1-7,10-11,19,21,24H,8-9,12-16H2,(H,23,26)\n  - levenshtein distance: 64\n\nSo even though model predicts same one, it could be error... 🤔?\n(If they are not same, please kindly tell me)",
    "1250604": "Yeah, to me your examples also look 100% identical (visually). Can you paste text version of them here? Maybe worth looking at them at pubchem, maybe 3d structure? Did you try to convert them to mol and then back to inchi? I think there were a few labels that after  converting them to mol and then back to inchi produced different results.",
    "1250632": "> Can you paste text version of them here?\n\nI added text version!\n\n> Did you try to convert them to mol and then back to inchi?\n\nYes, like below.\n\n```\nmol = rdkit.Chem.MolFromInchi(pred)\npred_image = rdkit.Chem.Draw.MolToImage(mol, size=(w, h))\n```",
    "1250684": "I meant like this:\n```\nChem.MolToInchi(Chem.MolFromInchi(pred_label)) == train_label\n```\nI applied this transformation to your predictions and it produced same inchi as training labels. For some reason your model predicts not canonical inchis sometimes. Probably you can gain some points by normalizing them manually.",
    "1250699": "That's a reasonable idea (or maybe re-label them), though for now I'm not sure how many of them exist. Hopefully, as training dataset is large the effect of those images is negligible.",
    "1250700": "Wow, as you say I confirmed this transformation gives me same inchi as training labels.\nThanks for sharing! So then we should apply this transformation when monitoring validation and submission.",
    "1251814": "stassl , could you please explain what you mean by this...\n> Probably you can gain some points by normalizing them manually",
    "1251941": "Holy.. this is gold for free.\n@nitindatta\n**predicted inchi** describes the **same molecule** as its **label**, but the inchis **differ** -> levenshtein can be huge\nBut it seems that using rdkit: \"correctly\" predicted inchi -> molfrominchi; moltoinchi -> \"true\" inchi",
    "1252192": "They are not the wrong label. There are roughly 10% images in the training set (not sure about test set) belongs to this alternative-structure-possible category. It's due to\n* either a methyl group attached to a carbon but lose the bond *due to compression* `1ab069aad819`\n* or extremely fuzzy atom depiction (e.g., F looks like I; I completely lose the line `0bb3c2bdcc2c`)\n* There are other structures (roughly 2-3%) that even domain expert cannot recover the structure without context, due to crossing bonds and overlapped depiction. e.g., `9203b1e7df7b` `d88774e5cc9f`\n\nIf one sticks to making **valid** InChI prediction, the above 2 scenarios will significantly inflate the average LD, because the atom indexes will shift (10-> 11, 11->12, so forth).\n\nThe LB ~ 2.1 is amazing, but if any model manages to get below 2, I am more confident it's overfitting to the depiction style, given the aforementioned edge cases.\n\n~~Of course, kaggle competition doesn't care...~~",
    "1252245": "I don't think it's negligible. I checked 177 predictions in which I cannot find a PubChem match from a random 3k benchmark. I find\n* 35 images have missing methyl group attached to carbon\n* 20 images have extremely fuzzy atom depiction\n\nThis already accounts for ~2% out of the 3,000. I am not even checking the rest 2833 which have PubChem matches. Each alternative structure may inflate the LD by ~40. Therefore, at least 40% LD are due to these alternative structures, as the average LD is approaching 2.",
    "1252353": "houndcl thanks for chiming in. I've seen your topic [Problem with the evaluation metric](https://www.kaggle.com/c/bms-molecular-translation/discussion/224394) and this post is just another evidence of the problem you mentioned. \n\nMaybe the title of my post is a bit inaccurate in favor of catchiness - I didn't mean the labels are wrong. Or better to say I understand that labels were correct originally, but as result of some transformations (corruption/compression/resizing - intentional or not) they are no longer matching the image. \n\nAs you say those corruptions may fall into several categories. Let’s maybe omit for now images that are convoluted, with multiple crossing bonds and complicated 3D-structure. If we consider rather simple molecules, and suppose it lost some bonds due to corruptions. I would say there could be following cases how it affects original label:\n\n- Corruption leads to an invalid molecule, but it can be uniquely recovered by domain expert, like when some bond is lost in the middle of molecule and you know it should be there. Then I would say the original label is still correct and I believe NN can learn to recover them as they are trained for language modeling when some words are masked.\n- Corruption leads to an invalid molecule, but it can be recovered to multiple valid ones. In that case model can predict distribution over possible labels and original label will be one of them.\n- Corruption leads to a VALID molecule but this molecule is different from original, thus the label becomes irrelevant. Particularly in my post I wanted to pay attention to this case. \n\n> either a methyl group attached to a carbon but lose the bond due to compression 1ab069aad819\n\nI’m attaching the image for convenience that your referenced:\n| train image | generated from label |\n| --- | --- |\n| ![image](https://user-images.githubusercontent.com/4602302/112497893-13ff2b00-8d97-11eb-8a2f-e6040efd3d36.png) | ![image](https://user-images.githubusercontent.com/4602302/112497943-1e212980-8d97-11eb-871f-26386dd67910.png) |\n\nIts original label is **InChI=1S/C10H10O4/c1-5-8(11)3-6-7(9(5)13-2)4-14-10(6)12/h3,11H,4H2,1-2H3**, but when you remove a single bond you’ll get also a valid molecule with label **InChI=1S/C9H8O4/c1-12-8-3-5(10)2-6-7(8)4-13-9(6)11/h2-3,10H,4H2,1H3** which is *29* edit operation away from the original. And I believe that is what my model would predict and any domain expert that will see only the corrupted image without that bond, because why anyone would add any bonds if the molecule is already valid. \n\nAnother interesting question that you brought up is what bonds can be lost and why. I was thinking that any bond can be dropped to methyl group, as you say, or somewhere in the middle or any other place. Do you observe that bonds can be missing only to methyl groups (is that equivalent to say at the border of molecules, but not in the middle, sorry my chemistry not so good)?\n\nUPD:\nOk, considering my last question about where bonds can be dropped, I found confirmation of my initial thought that they can be dropped anywhere, including in the middle of molecule, and not only to methyl groups:\n![image](https://user-images.githubusercontent.com/4602302/112535690-a4e8fd00-8dbd-11eb-8984-11d5dfe7565d.png)\n![image](https://user-images.githubusercontent.com/4602302/112536280-4f612000-8dbe-11eb-819a-53318cbf1bfe.png)",
    "1252409": "Thank you @stassl for sharing. I think maybe due to some representation missed on the training features.",
    "1259579": "You did not submit anything with such a good model?  Also, what do yu used to create your images?  rdkit gives me different ones.",
    "1259680": "Not yet, as it takes too long to generate predictions for 1.6M images, for now working on improving my local CV score. \n\nWell, I just cherry-picked some examples where it was good, unfortunately there are lots of them where it is not so good, especially for convoluted or very large molecules. \n\nThe images of molecules I took from pubchem site.",
    "1259733": "Thanks for the reply and good luck.",
    "1264187": "Some InChIs can not be parsed by RDKit and cause segmentation faults 😲\n> Process finished with exit code 139 (interrupted by signal 11: SIGSEGV)",
    "1264250": "There is a tool I made to do this:\nhttps://www.kaggle.com/nofreewill/normalize-your-predictions"
  },
  "source": "meta"
}