{
  "id": 243809,
  "title": "10th solution & thoughts - Part2. LB 0.61 [ Think out of the box version ]",
  "url": "/competitions/bms-molecular-translation/discussion/243809",
  "author_name": "",
  "post_date": "2021-06-04T03:54:01.855766500Z",
  "votes": 40,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Part 1 by the great leader <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>: <br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243766\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/243766</a> the TNT/VIT/CAIT ensemble model has provided a good baseline for my last-minute shot.</p>\n<h4>Before technical details</h4>\n<p>From the competition overview:</p>\n<blockquote>\n  <p>Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. </p>\n</blockquote>\n<p>I come from the so said <em>existing tools</em> world with zero deep learning experience. I joined the competition to prove this overview is not true (but also true). My colleagues are the developers of MolVec: <a href=\"https://molvec.ncats.io\" target=\"_blank\">https://molvec.ncats.io</a>. \"Traditional OCR method is able to achieve &gt;90% accuracy on near-perfect images\", but the first submission slapped my face hard: LD ~ 80, just slightly better than the H2O naïve baseline. Some heuristic fine tuning (link disconnected components, atomic/bond fixers) has made molvec perform much better on noised images, but we found that pure OCR method stuck on the LD = 5 line. Improvement is possible but it is limited by <em>\"what is presented by the image\"</em> rather than <em>\"what is encoded in the image\"</em>. So I wrote about <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/233927\" target=\"_blank\">ghost atom</a> problem in image resizing. In addition, OCR methods which heavily rely on heuristics (e.g., the appearance of atom/bond/connection) has a hard time dealing with complex <a href=\"https://molvec.ncats.io/c3bb46b5a4\" target=\"_blank\">cross-bond structure (example)</a>. I will discuss about OCR method in the thoughts section. </p>\n<p>So, while top teams are approaching LD ~ 1, what's the use of traditional OCR method which is stuck on LD = 5? We abandoned our arrogance and shifted overall strategy: patch deep learning predictions for what it doesn't work well.</p>\n<ol>\n<li>by feeding what OCR is good at (near-perfect image) </li>\n<li>on the areas where other method doesn't work (big image)</li>\n<li>borrow the idea (aka, build molecule atom by atom using chemistry rules)</li>\n</ol>\n<h4>1. Feed OCR with super-resolution image</h4>\n<p>\"Traditional OCR method is able to achieve &gt;90% accuracy on near-perfect images\". </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Training</td>\n<td>248k randomly selected kaggle image - Indigo rendered image pairs</td>\n</tr>\n<tr>\n<td>Model</td>\n<td>ResNet34 + U-Net with self-attention <a href=\"https://github.com/fastai/fastai/blob/master/dev_nbs/course/lesson7-superres.ipynb\" target=\"_blank\">fastai tutorial</a></td>\n</tr>\n<tr>\n<td>Image size</td>\n<td>832 x 832</td>\n</tr>\n<tr>\n<td>Inference</td>\n<td>- small image (any dimension &lt; 750): upscale to 832. big image (any dimension 750-1200): upscale to 1248. xbig image (any dimension &gt; 1200): upscale to 1504</td>\n</tr>\n<tr>\n<td>InChI generator</td>\n<td>super resolution images -&gt; molvec -&gt; .mol -&gt; InChI</td>\n</tr>\n<tr>\n<td>Performance</td>\n<td>LD ~ 3</td>\n</tr>\n<tr>\n<td>Fail</td>\n<td>complex molecule</td>\n</tr>\n</tbody>\n</table>\n<h4>2. Use OCR on extra large image (e.g., width &gt; 1000 pixels)</h4>\n<p>From the discussion board, there are lots of complaints like \"my model doesn't work well on large images\". But OCR works decently regardless of image size (10,000 pixels is fine)</p>\n<p><em>No training</em>. Just feed super resolution images to molvec -&gt; .mol -&gt; InChI<br>\nI didn't benchmark the performance for big images, but with 90% accuracy, it should outperforms deep learning model that downscaled to 384 or 224 in which most molecular features are lost.</p>\n<h4>3. Object detection route to build complex molecules from scratch</h4>\n<p>Idea identical to the DACON winning solution, but to build a valid molecule with acceptable layout from the detected objects on <em>noised</em> image requires lots of \"fixers\" as well as atom/bond imputation. We introduced the <strong>atom group</strong> idea to help the valency identification: e.g., CH2 and NH are different than C and N. We separate implicit carbon (shown as vertex) and explicit carbon (shown as Cxx). Meanwhile we enlarge the bbox a bit during the training to include some local context, e.g., a CH2 group usually has 2 bond tips in the bbox. There are totally 54 atom groups and bond types in the 1.6M training set, and we selected 49 for training.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Training</td>\n<td>691k training set images. Bounding boxes of atom/atom group/bond were obtained from Indigo rendered svg and custom parser. These 691k includes hard images with a. cross-bond structure or b. atom-atom close contact or c. rare atom/atom groups (e.g, isotope, explicit H, explicit CH2, etc)</td>\n</tr>\n<tr>\n<td>Model</td>\n<td>EfficientDet-D3</td>\n</tr>\n<tr>\n<td>Image size</td>\n<td>896 x 896 (bigger the better)</td>\n</tr>\n<tr>\n<td>Augmentation</td>\n<td>rotation / flip / salt &amp; pepper</td>\n</tr>\n<tr>\n<td>InChI generator</td>\n<td>bbox -&gt; custom molecular builder script -&gt; .mol -&gt; InChI</td>\n</tr>\n<tr>\n<td>Performance</td>\n<td>overall mAP=0.87. LD~0.65 for simple molecule. LD~3.5 for complex molecule</td>\n</tr>\n<tr>\n<td>Fail</td>\n<td>big images dim &gt; 896</td>\n</tr>\n</tbody>\n</table>\n<p>Manually check several hundreds hard molecules in the validation set where VIT failed to predict valid InChI, and found that on average MolBuild could reduce LD by ~25 even if MolBuild predicted the wrong molecule. Performance could be better if we have more time on training. It is almost a last-minute model.<br>\n<img src=\"https://i.ibb.co/BK6DWwX/builder.png\" alt=\"\"></p>\n<h4>4. Merge with TNT/VIT/CAIT ensemble model</h4>\n<p><strong>LB 0.71 -&gt; 0.61</strong> by updating only <strong>7,679</strong> invalid InChIs (totally 9,800) from VIT in the final submission. Criteria of substitution:</p>\n<ol>\n<li>Top tier: the formula component of InChI (the C3H4O2 part) matches the VIT prediction. </li>\n<li>Second tier: OCR/MolBuilder-to-VIT edit distance is lower than 10. </li>\n<li>Third tier: OCR matches MolBuilder prediction</li>\n<li>Finally, push in all OCR result for big and xbig images, regardless of distance</li>\n</ol>\n<p>Replacing 0.48% test set gains 0.1 overall LB score, indicating that <strong>on average each substitution reduces Levenshtein distance by 21</strong>. </p>\n<h4>Thanks</h4>\n<p>First I'd like to thank leader <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and members <a href=\"https://www.kaggle.com/mathurinache\" target=\"_blank\">@mathurinache</a> <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> who have worked very hard in the past few months. Special thank to <a href=\"https://www.kaggle.com/DrHB\" target=\"_blank\">@DrHB</a>'s team who crashed my day dream that OCR is undefeatable early in this competition. Also thank my colleagues who developed SOTA OCR method and fine tune the parameters. Last but not least, it's been a great &amp; tough competition. Thank you organizers &amp; staff behind the scene.</p>\n<h4>Thoughts</h4>\n<h6>End of molecular OCR?</h6>\n<p>With more training data augmented with different renderer parameters, I am convinced that deep learning will win with a big margin, at least for \"plain\" &amp; small molecule. On the other hand, traditional OCR method which is more generalized can be complementary to SOTA DL model when training data is unreachable (e.g., complex layout). <strong>There is always a sweet spot between domain knowledge which offers heuristics and data science which offers precision.</strong> Domain knowledge provides indispensable shortcut to the core of problem, and is cost-efficient in pre/post-processing and when big &amp; clean data / annotation are lacking. </p>\n<h6>Why InChI validation by RDKit works?</h6>\n<p>First we need to understand what is being learned: </p>\n<ol>\n<li>chemistry rules (e.g., a carbon cannot have 10 bonds. InChI string has strict rules)</li>\n<li>layout (for 99.9% cases, a benzene ring will be rendered as a hexagon)</li>\n<li>noise (vertical/horizontal lines are more likely to disappear in resizing)</li>\n<li>synthetic feasibility (wild structure could hardly be synthesized)</li>\n</ol>\n<p>From my understanding, most image-driven, encoder-decoder model is good at learning layout (#2) &amp; noise (#3). InChI validation introduces the chemistry rules (#1) elegantly. There is a great discussion <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243102\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>. On the other hand, OCR method which mainly looks at chemistry will be defeated when noise / layout is out of its domain. The message is that when we understand the problem better, we will solve it more efficiently.</p>\n<p>Happy think out of the box.</p>",
  "messages": [
    {
      "id": "1335176",
      "postDate": "06/04/2021 03:54:01",
      "content": "<p>Part 1 by the great leader <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>: <br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243766\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/243766</a> the TNT/VIT/CAIT ensemble model has provided a good baseline for my last-minute shot.</p>\n<h4>Before technical details</h4>\n<p>From the competition overview:</p>\n<blockquote>\n  <p>Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. </p>\n</blockquote>\n<p>I come from the so said <em>existing tools</em> world with zero deep learning experience. I joined the competition to prove this overview is not true (but also true). My colleagues are the developers of MolVec: <a href=\"https://molvec.ncats.io\" target=\"_blank\">https://molvec.ncats.io</a>. \"Traditional OCR method is able to achieve &gt;90% accuracy on near-perfect images\", but the first submission slapped my face hard: LD ~ 80, just slightly better than the H2O naïve baseline. Some heuristic fine tuning (link disconnected components, atomic/bond fixers) has made molvec perform much better on noised images, but we found that pure OCR method stuck on the LD = 5 line. Improvement is possible but it is limited by <em>\"what is presented by the image\"</em> rather than <em>\"what is encoded in the image\"</em>. So I wrote about <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/233927\" target=\"_blank\">ghost atom</a> problem in image resizing. In addition, OCR methods which heavily rely on heuristics (e.g., the appearance of atom/bond/connection) has a hard time dealing with complex <a href=\"https://molvec.ncats.io/c3bb46b5a4\" target=\"_blank\">cross-bond structure (example)</a>. I will discuss about OCR method in the thoughts section. </p>\n<p>So, while top teams are approaching LD ~ 1, what's the use of traditional OCR method which is stuck on LD = 5? We abandoned our arrogance and shifted overall strategy: patch deep learning predictions for what it doesn't work well.</p>\n<ol>\n<li>by feeding what OCR is good at (near-perfect image) </li>\n<li>on the areas where other method doesn't work (big image)</li>\n<li>borrow the idea (aka, build molecule atom by atom using chemistry rules)</li>\n</ol>\n<h4>1. Feed OCR with super-resolution image</h4>\n<p>\"Traditional OCR method is able to achieve &gt;90% accuracy on near-perfect images\". </p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Training</td>\n<td>248k randomly selected kaggle image - Indigo rendered image pairs</td>\n</tr>\n<tr>\n<td>Model</td>\n<td>ResNet34 + U-Net with self-attention <a href=\"https://github.com/fastai/fastai/blob/master/dev_nbs/course/lesson7-superres.ipynb\" target=\"_blank\">fastai tutorial</a></td>\n</tr>\n<tr>\n<td>Image size</td>\n<td>832 x 832</td>\n</tr>\n<tr>\n<td>Inference</td>\n<td>- small image (any dimension &lt; 750): upscale to 832. big image (any dimension 750-1200): upscale to 1248. xbig image (any dimension &gt; 1200): upscale to 1504</td>\n</tr>\n<tr>\n<td>InChI generator</td>\n<td>super resolution images -&gt; molvec -&gt; .mol -&gt; InChI</td>\n</tr>\n<tr>\n<td>Performance</td>\n<td>LD ~ 3</td>\n</tr>\n<tr>\n<td>Fail</td>\n<td>complex molecule</td>\n</tr>\n</tbody>\n</table>\n<h4>2. Use OCR on extra large image (e.g., width &gt; 1000 pixels)</h4>\n<p>From the discussion board, there are lots of complaints like \"my model doesn't work well on large images\". But OCR works decently regardless of image size (10,000 pixels is fine)</p>\n<p><em>No training</em>. Just feed super resolution images to molvec -&gt; .mol -&gt; InChI<br>\nI didn't benchmark the performance for big images, but with 90% accuracy, it should outperforms deep learning model that downscaled to 384 or 224 in which most molecular features are lost.</p>\n<h4>3. Object detection route to build complex molecules from scratch</h4>\n<p>Idea identical to the DACON winning solution, but to build a valid molecule with acceptable layout from the detected objects on <em>noised</em> image requires lots of \"fixers\" as well as atom/bond imputation. We introduced the <strong>atom group</strong> idea to help the valency identification: e.g., CH2 and NH are different than C and N. We separate implicit carbon (shown as vertex) and explicit carbon (shown as Cxx). Meanwhile we enlarge the bbox a bit during the training to include some local context, e.g., a CH2 group usually has 2 bond tips in the bbox. There are totally 54 atom groups and bond types in the 1.6M training set, and we selected 49 for training.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Training</td>\n<td>691k training set images. Bounding boxes of atom/atom group/bond were obtained from Indigo rendered svg and custom parser. These 691k includes hard images with a. cross-bond structure or b. atom-atom close contact or c. rare atom/atom groups (e.g, isotope, explicit H, explicit CH2, etc)</td>\n</tr>\n<tr>\n<td>Model</td>\n<td>EfficientDet-D3</td>\n</tr>\n<tr>\n<td>Image size</td>\n<td>896 x 896 (bigger the better)</td>\n</tr>\n<tr>\n<td>Augmentation</td>\n<td>rotation / flip / salt &amp; pepper</td>\n</tr>\n<tr>\n<td>InChI generator</td>\n<td>bbox -&gt; custom molecular builder script -&gt; .mol -&gt; InChI</td>\n</tr>\n<tr>\n<td>Performance</td>\n<td>overall mAP=0.87. LD~0.65 for simple molecule. LD~3.5 for complex molecule</td>\n</tr>\n<tr>\n<td>Fail</td>\n<td>big images dim &gt; 896</td>\n</tr>\n</tbody>\n</table>\n<p>Manually check several hundreds hard molecules in the validation set where VIT failed to predict valid InChI, and found that on average MolBuild could reduce LD by ~25 even if MolBuild predicted the wrong molecule. Performance could be better if we have more time on training. It is almost a last-minute model.<br>\n<img src=\"https://i.ibb.co/BK6DWwX/builder.png\" alt=\"\"></p>\n<h4>4. Merge with TNT/VIT/CAIT ensemble model</h4>\n<p><strong>LB 0.71 -&gt; 0.61</strong> by updating only <strong>7,679</strong> invalid InChIs (totally 9,800) from VIT in the final submission. Criteria of substitution:</p>\n<ol>\n<li>Top tier: the formula component of InChI (the C3H4O2 part) matches the VIT prediction. </li>\n<li>Second tier: OCR/MolBuilder-to-VIT edit distance is lower than 10. </li>\n<li>Third tier: OCR matches MolBuilder prediction</li>\n<li>Finally, push in all OCR result for big and xbig images, regardless of distance</li>\n</ol>\n<p>Replacing 0.48% test set gains 0.1 overall LB score, indicating that <strong>on average each substitution reduces Levenshtein distance by 21</strong>. </p>\n<h4>Thanks</h4>\n<p>First I'd like to thank leader <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and members <a href=\"https://www.kaggle.com/mathurinache\" target=\"_blank\">@mathurinache</a> <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> who have worked very hard in the past few months. Special thank to <a href=\"https://www.kaggle.com/DrHB\" target=\"_blank\">@DrHB</a>'s team who crashed my day dream that OCR is undefeatable early in this competition. Also thank my colleagues who developed SOTA OCR method and fine tune the parameters. Last but not least, it's been a great &amp; tough competition. Thank you organizers &amp; staff behind the scene.</p>\n<h4>Thoughts</h4>\n<h6>End of molecular OCR?</h6>\n<p>With more training data augmented with different renderer parameters, I am convinced that deep learning will win with a big margin, at least for \"plain\" &amp; small molecule. On the other hand, traditional OCR method which is more generalized can be complementary to SOTA DL model when training data is unreachable (e.g., complex layout). <strong>There is always a sweet spot between domain knowledge which offers heuristics and data science which offers precision.</strong> Domain knowledge provides indispensable shortcut to the core of problem, and is cost-efficient in pre/post-processing and when big &amp; clean data / annotation are lacking. </p>\n<h6>Why InChI validation by RDKit works?</h6>\n<p>First we need to understand what is being learned: </p>\n<ol>\n<li>chemistry rules (e.g., a carbon cannot have 10 bonds. InChI string has strict rules)</li>\n<li>layout (for 99.9% cases, a benzene ring will be rendered as a hexagon)</li>\n<li>noise (vertical/horizontal lines are more likely to disappear in resizing)</li>\n<li>synthetic feasibility (wild structure could hardly be synthesized)</li>\n</ol>\n<p>From my understanding, most image-driven, encoder-decoder model is good at learning layout (#2) &amp; noise (#3). InChI validation introduces the chemistry rules (#1) elegantly. There is a great discussion <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243102\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>. On the other hand, OCR method which mainly looks at chemistry will be defeated when noise / layout is out of its domain. The message is that when we understand the problem better, we will solve it more efficiently.</p>\n<p>Happy think out of the box.</p>",
      "rawMarkdown": "Part 1 by the great leader @hengck23: \nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/243766 the TNT/VIT/CAIT ensemble model has provided a good baseline for my last-minute shot.\n\n#### Before technical details\nFrom the competition overview:\n> Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. \n\nI come from the so said *existing tools* world with zero deep learning experience. I joined the competition to prove this overview is not true (but also true). My colleagues are the developers of MolVec: https://molvec.ncats.io. \"Traditional OCR method is able to achieve >90% accuracy on near-perfect images\", but the first submission slapped my face hard: LD ~ 80, just slightly better than the H2O naïve baseline. Some heuristic fine tuning (link disconnected components, atomic/bond fixers) has made molvec perform much better on noised images, but we found that pure OCR method stuck on the LD = 5 line. Improvement is possible but it is limited by *\"what is presented by the image\"* rather than *\"what is encoded in the image\"*. So I wrote about [ghost atom](https://www.kaggle.com/c/bms-molecular-translation/discussion/233927) problem in image resizing. In addition, OCR methods which heavily rely on heuristics (e.g., the appearance of atom/bond/connection) has a hard time dealing with complex [cross-bond structure (example)](https://molvec.ncats.io/c3bb46b5a4). I will discuss about OCR method in the thoughts section. \n\nSo, while top teams are approaching LD ~ 1, what's the use of traditional OCR method which is stuck on LD = 5? We abandoned our arrogance and shifted overall strategy: patch deep learning predictions for what it doesn't work well.\n1. by feeding what OCR is good at (near-perfect image) \n2. on the areas where other method doesn't work (big image)\n3. borrow the idea (aka, build molecule atom by atom using chemistry rules)\n\n#### 1. Feed OCR with super-resolution image\n\"Traditional OCR method is able to achieve >90% accuracy on near-perfect images\". \n|  |  |\n| --- | --- |\n| Training | 248k randomly selected kaggle image - Indigo rendered image pairs |\n| Model | ResNet34 + U-Net with self-attention [fastai tutorial](https://github.com/fastai/fastai/blob/master/dev_nbs/course/lesson7-superres.ipynb) |\n| Image size | 832 x 832 |\n| Inference | - small image (any dimension < 750): upscale to 832. big image (any dimension 750-1200): upscale to 1248. xbig image (any dimension > 1200): upscale to 1504\n|\n|InChI generator | super resolution images -> molvec -> .mol -> InChI |\n|Performance | LD ~ 3 |\n|Fail | complex molecule |\n\n#### 2. Use OCR on extra large image (e.g., width > 1000 pixels)\nFrom the discussion board, there are lots of complaints like \"my model doesn't work well on large images\". But OCR works decently regardless of image size (10,000 pixels is fine)\n\n*No training*. Just feed super resolution images to molvec -> .mol -> InChI\nI didn't benchmark the performance for big images, but with 90% accuracy, it should outperforms deep learning model that downscaled to 384 or 224 in which most molecular features are lost.\n\n#### 3. Object detection route to build complex molecules from scratch\nIdea identical to the DACON winning solution, but to build a valid molecule with acceptable layout from the detected objects on *noised* image requires lots of \"fixers\" as well as atom/bond imputation. We introduced the **atom group** idea to help the valency identification: e.g., CH2 and NH are different than C and N. We separate implicit carbon (shown as vertex) and explicit carbon (shown as Cxx). Meanwhile we enlarge the bbox a bit during the training to include some local context, e.g., a CH2 group usually has 2 bond tips in the bbox. There are totally 54 atom groups and bond types in the 1.6M training set, and we selected 49 for training.\n\n|  |  |\n| --- | --- |\n| Training | 691k training set images. Bounding boxes of atom/atom group/bond were obtained from Indigo rendered svg and custom parser. These 691k includes hard images with a. cross-bond structure or b. atom-atom close contact or c. rare atom/atom groups (e.g, isotope, explicit H, explicit CH2, etc) |\n| Model | EfficientDet-D3|\n| Image size| 896 x 896 (bigger the better) |\n| Augmentation | rotation / flip / salt & pepper|\n| InChI generator | bbox -> custom molecular builder script -> .mol -> InChI |\n| Performance | overall mAP=0.87. LD~0.65 for simple molecule. LD~3.5 for complex molecule |\n| Fail | big images dim > 896 |\n\nManually check several hundreds hard molecules in the validation set where VIT failed to predict valid InChI, and found that on average MolBuild could reduce LD by ~25 even if MolBuild predicted the wrong molecule. Performance could be better if we have more time on training. It is almost a last-minute model.\n![](https://i.ibb.co/BK6DWwX/builder.png)\n\n#### 4. Merge with TNT/VIT/CAIT ensemble model\n**LB 0.71 -> 0.61** by updating only **7,679** invalid InChIs (totally 9,800) from VIT in the final submission. Criteria of substitution:\n\n1. Top tier: the formula component of InChI (the C3H4O2 part) matches the VIT prediction. \n2. Second tier: OCR/MolBuilder-to-VIT edit distance is lower than 10. \n3. Third tier: OCR matches MolBuilder prediction\n4. Finally, push in all OCR result for big and xbig images, regardless of distance\n\nReplacing 0.48% test set gains 0.1 overall LB score, indicating that **on average each substitution reduces Levenshtein distance by 21**. \n\n#### Thanks\nFirst I'd like to thank leader @hengck23 and members @mathurinache @ruchi798 who have worked very hard in the past few months. Special thank to @DrHB's team who crashed my day dream that OCR is undefeatable early in this competition. Also thank my colleagues who developed SOTA OCR method and fine tune the parameters. Last but not least, it's been a great & tough competition. Thank you organizers & staff behind the scene.\n\n#### Thoughts\n###### End of molecular OCR?\nWith more training data augmented with different renderer parameters, I am convinced that deep learning will win with a big margin, at least for \"plain\" & small molecule. On the other hand, traditional OCR method which is more generalized can be complementary to SOTA DL model when training data is unreachable (e.g., complex layout). **There is always a sweet spot between domain knowledge which offers heuristics and data science which offers precision.** Domain knowledge provides indispensable shortcut to the core of problem, and is cost-efficient in pre/post-processing and when big & clean data / annotation are lacking. \n\n###### Why InChI validation by RDKit works?\nFirst we need to understand what is being learned: \n1. chemistry rules (e.g., a carbon cannot have 10 bonds. InChI string has strict rules)\n2. layout (for 99.9% cases, a benzene ring will be rendered as a hexagon)\n3. noise (vertical/horizontal lines are more likely to disappear in resizing)\n4. synthetic feasibility (wild structure could hardly be synthesized)\n\nFrom my understanding, most image-driven, encoder-decoder model is good at learning layout (#2) & noise (#3). InChI validation introduces the chemistry rules (#1) elegantly. There is a great discussion [here](https://www.kaggle.com/c/bms-molecular-translation/discussion/243102) by @nofreewill. On the other hand, OCR method which mainly looks at chemistry will be defeated when noise / layout is out of its domain. The message is that when we understand the problem better, we will solve it more efficiently.\n\nHappy think out of the box.",
      "votes": null
    },
    {
      "id": "1335262",
      "postDate": "06/04/2021 05:44:43",
      "content": "<p>Wow, this is a really cool trick) Congratulations on your gold medal!</p>",
      "rawMarkdown": "Wow, this is a really cool trick) Congratulations on your gold medal!",
      "votes": null
    },
    {
      "id": "1335713",
      "postDate": "06/04/2021 11:50:03",
      "content": "<p>It's truly amazing to see solutions like this ;)</p>",
      "rawMarkdown": "It's truly amazing to see solutions like this ;)",
      "votes": null
    },
    {
      "id": "1597039",
      "postDate": "11/27/2021 06:24:12",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a>, I really like your solution! Do you by any chance have ideas why the competition host used Indigo to visualize the structures? Would any other tools have worked similarly or does Indigo have some advantages compared to the other tools? Thanks!</p>",
      "rawMarkdown": "Dear @houndcl, I really like your solution! Do you by any chance have ideas why the competition host used Indigo to visualize the structures? Would any other tools have worked similarly or does Indigo have some advantages compared to the other tools? Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335262,
      "author_name": "timofeev25",
      "author_url": "",
      "post_date": "06/04/2021 05:44:43",
      "content": "<p>Wow, this is a really cool trick) Congratulations on your gold medal!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335713,
      "author_name": "artnotintelligence",
      "author_url": "",
      "post_date": "06/04/2021 11:50:03",
      "content": "<p>It's truly amazing to see solutions like this ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1597039,
      "author_name": "nurlybek17",
      "author_url": "",
      "post_date": "11/27/2021 06:24:12",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/houndcl\" target=\"_blank\">@houndcl</a>, I really like your solution! Do you by any chance have ideas why the competition host used Indigo to visualize the structures? Would any other tools have worked similarly or does Indigo have some advantages compared to the other tools? Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335176": "Part 1 by the great leader @hengck23: \nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/243766 the TNT/VIT/CAIT ensemble model has provided a good baseline for my last-minute shot.\n\n#### Before technical details\nFrom the competition overview:\n> Existing tools produce 90% accuracy but only under optimal conditions. Historical sources often have some level of image corruption, which reduces performance to near zero. \n\nI come from the so said *existing tools* world with zero deep learning experience. I joined the competition to prove this overview is not true (but also true). My colleagues are the developers of MolVec: https://molvec.ncats.io. \"Traditional OCR method is able to achieve >90% accuracy on near-perfect images\", but the first submission slapped my face hard: LD ~ 80, just slightly better than the H2O naïve baseline. Some heuristic fine tuning (link disconnected components, atomic/bond fixers) has made molvec perform much better on noised images, but we found that pure OCR method stuck on the LD = 5 line. Improvement is possible but it is limited by *\"what is presented by the image\"* rather than *\"what is encoded in the image\"*. So I wrote about [ghost atom](https://www.kaggle.com/c/bms-molecular-translation/discussion/233927) problem in image resizing. In addition, OCR methods which heavily rely on heuristics (e.g., the appearance of atom/bond/connection) has a hard time dealing with complex [cross-bond structure (example)](https://molvec.ncats.io/c3bb46b5a4). I will discuss about OCR method in the thoughts section. \n\nSo, while top teams are approaching LD ~ 1, what's the use of traditional OCR method which is stuck on LD = 5? We abandoned our arrogance and shifted overall strategy: patch deep learning predictions for what it doesn't work well.\n1. by feeding what OCR is good at (near-perfect image) \n2. on the areas where other method doesn't work (big image)\n3. borrow the idea (aka, build molecule atom by atom using chemistry rules)\n\n#### 1. Feed OCR with super-resolution image\n\"Traditional OCR method is able to achieve >90% accuracy on near-perfect images\". \n|  |  |\n| --- | --- |\n| Training | 248k randomly selected kaggle image - Indigo rendered image pairs |\n| Model | ResNet34 + U-Net with self-attention [fastai tutorial](https://github.com/fastai/fastai/blob/master/dev_nbs/course/lesson7-superres.ipynb) |\n| Image size | 832 x 832 |\n| Inference | - small image (any dimension < 750): upscale to 832. big image (any dimension 750-1200): upscale to 1248. xbig image (any dimension > 1200): upscale to 1504\n|\n|InChI generator | super resolution images -> molvec -> .mol -> InChI |\n|Performance | LD ~ 3 |\n|Fail | complex molecule |\n\n#### 2. Use OCR on extra large image (e.g., width > 1000 pixels)\nFrom the discussion board, there are lots of complaints like \"my model doesn't work well on large images\". But OCR works decently regardless of image size (10,000 pixels is fine)\n\n*No training*. Just feed super resolution images to molvec -> .mol -> InChI\nI didn't benchmark the performance for big images, but with 90% accuracy, it should outperforms deep learning model that downscaled to 384 or 224 in which most molecular features are lost.\n\n#### 3. Object detection route to build complex molecules from scratch\nIdea identical to the DACON winning solution, but to build a valid molecule with acceptable layout from the detected objects on *noised* image requires lots of \"fixers\" as well as atom/bond imputation. We introduced the **atom group** idea to help the valency identification: e.g., CH2 and NH are different than C and N. We separate implicit carbon (shown as vertex) and explicit carbon (shown as Cxx). Meanwhile we enlarge the bbox a bit during the training to include some local context, e.g., a CH2 group usually has 2 bond tips in the bbox. There are totally 54 atom groups and bond types in the 1.6M training set, and we selected 49 for training.\n\n|  |  |\n| --- | --- |\n| Training | 691k training set images. Bounding boxes of atom/atom group/bond were obtained from Indigo rendered svg and custom parser. These 691k includes hard images with a. cross-bond structure or b. atom-atom close contact or c. rare atom/atom groups (e.g, isotope, explicit H, explicit CH2, etc) |\n| Model | EfficientDet-D3|\n| Image size| 896 x 896 (bigger the better) |\n| Augmentation | rotation / flip / salt & pepper|\n| InChI generator | bbox -> custom molecular builder script -> .mol -> InChI |\n| Performance | overall mAP=0.87. LD~0.65 for simple molecule. LD~3.5 for complex molecule |\n| Fail | big images dim > 896 |\n\nManually check several hundreds hard molecules in the validation set where VIT failed to predict valid InChI, and found that on average MolBuild could reduce LD by ~25 even if MolBuild predicted the wrong molecule. Performance could be better if we have more time on training. It is almost a last-minute model.\n![](https://i.ibb.co/BK6DWwX/builder.png)\n\n#### 4. Merge with TNT/VIT/CAIT ensemble model\n**LB 0.71 -> 0.61** by updating only **7,679** invalid InChIs (totally 9,800) from VIT in the final submission. Criteria of substitution:\n\n1. Top tier: the formula component of InChI (the C3H4O2 part) matches the VIT prediction. \n2. Second tier: OCR/MolBuilder-to-VIT edit distance is lower than 10. \n3. Third tier: OCR matches MolBuilder prediction\n4. Finally, push in all OCR result for big and xbig images, regardless of distance\n\nReplacing 0.48% test set gains 0.1 overall LB score, indicating that **on average each substitution reduces Levenshtein distance by 21**. \n\n#### Thanks\nFirst I'd like to thank leader @hengck23 and members @mathurinache @ruchi798 who have worked very hard in the past few months. Special thank to @DrHB's team who crashed my day dream that OCR is undefeatable early in this competition. Also thank my colleagues who developed SOTA OCR method and fine tune the parameters. Last but not least, it's been a great & tough competition. Thank you organizers & staff behind the scene.\n\n#### Thoughts\n###### End of molecular OCR?\nWith more training data augmented with different renderer parameters, I am convinced that deep learning will win with a big margin, at least for \"plain\" & small molecule. On the other hand, traditional OCR method which is more generalized can be complementary to SOTA DL model when training data is unreachable (e.g., complex layout). **There is always a sweet spot between domain knowledge which offers heuristics and data science which offers precision.** Domain knowledge provides indispensable shortcut to the core of problem, and is cost-efficient in pre/post-processing and when big & clean data / annotation are lacking. \n\n###### Why InChI validation by RDKit works?\nFirst we need to understand what is being learned: \n1. chemistry rules (e.g., a carbon cannot have 10 bonds. InChI string has strict rules)\n2. layout (for 99.9% cases, a benzene ring will be rendered as a hexagon)\n3. noise (vertical/horizontal lines are more likely to disappear in resizing)\n4. synthetic feasibility (wild structure could hardly be synthesized)\n\nFrom my understanding, most image-driven, encoder-decoder model is good at learning layout (#2) & noise (#3). InChI validation introduces the chemistry rules (#1) elegantly. There is a great discussion [here](https://www.kaggle.com/c/bms-molecular-translation/discussion/243102) by @nofreewill. On the other hand, OCR method which mainly looks at chemistry will be defeated when noise / layout is out of its domain. The message is that when we understand the problem better, we will solve it more efficiently.\n\nHappy think out of the box.",
    "1335262": "Wow, this is a really cool trick) Congratulations on your gold medal!",
    "1335713": "It's truly amazing to see solutions like this ;)",
    "1597039": "Dear @houndcl, I really like your solution! Do you by any chance have ideas why the competition host used Indigo to visualize the structures? Would any other tools have worked similarly or does Indigo have some advantages compared to the other tools? Thanks!"
  },
  "source": "meta"
}