{
  "id": 244031,
  "title": "9th Place Solution",
  "url": "/competitions/bms-molecular-translation/writeups/translate-chemistry-together-9th-place-solution",
  "author_name": "",
  "post_date": "2021-06-05T00:26:00.932924Z",
  "votes": 36,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, I would like to thanks the organizers for organizing such an exciting competition.  Second I would like to thank my teammates. We had a great working environment with everybody working hard, and last one week, we had to stay up until 3 am local times. Below I will try to summarize the main idea of our best-performing model.</p>\n<h2>Image Preprocessing</h2>\n<p>One of the challenging things about this competition was that the chemical had very different structures, some were super long, and others were very short. This also meant that images were highly diverse in size and shapes (e.g., super long in 1 axis or super small). Therefore it was essential to preprocess these images correctly. After some initial discussion we decided to try to use two methods:<br><br>\n<code>1) simple rescaling</code><br>\nplus - we will get all the information (kind off) <br>\nminus - feature of the images were become distorted (extreme shrinking across one axis)<br><br>\n<code>2) rescaling with preserving the aspect ratio</code><br>\nplus - This method will preserve the feature<br>\nminus  - bigger images, when downscaled, will have extremely small features <br><br>\nBoth of these methods have some advantages and disadvantages, but we reasoned that it would be best that each of our team members pick one of these methods and stick to it. We also found this an excellent article showing that most modern computer vision libraries employ very buggy resizing methods (<a href=\"https://arxiv.org/pdf/2104.11222.pdf)\" target=\"_blank\">https://arxiv.org/pdf/2104.11222.pdf)</a>. Since we are dealing with a lot of lines when downscaling, it was essential to preserve most of the information (tl dr: use PIL)</p>\n<h2>Image Augmentation.</h2>\n<p>We decided to go with natural image augmentations that were found in the dataset. But to make things a little bit diverse, we decided that everyone should train on a slightly different set of augmentations.<br><br>\nSet 1:<br>\n<code>RandomRotate90</code><br><br>\nSet 2</p>\n<pre><code>RandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomRotate90\nRandomNoiseAugment\n</code></pre>\n<p>Set 3</p>\n<pre><code>RandomRotate90\nRandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomNoiseAugment\nRandomLineRemoval\nRandomPiexlThreshold (some of the chemical lines had different sheds of black pixel values so we perform)\n</code></pre>\n<pre><code>thr = random.randint(200, 250)\nimg[img &lt; thr] = 0\nimg[img &gt;= thr] = 255\n</code></pre>\n<h2>Tokenizers</h2>\n<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> tokenizers was great, but we were wondering if there is a better way to tokenize. We found a lot of research on different tokenization strategies using <code>SMILES</code> but almost no information about <code>InChI</code>. After some thinking, we decided to try <code>huggingface</code> <code>BPE</code> tokenization. </p>\n<pre><code>vocab = 128 # can be any\ntokenizer = Tokenizer(BPE(unk_token=\"[&lt;unk&gt;]\", fuse_unk=False))\npre_tokenizer = pre_tokenizers.Sequence([Whitespace(), Digits()])\ntokenizer.pre_tokenizer = pre_tokenizer\ntrainer = BpeTrainer(vocab = max_len, special_tokens=[\"[&lt;pad&gt;]\", \"[&lt;start&gt;]\", \"[&lt;end&gt;]\", \"[&lt;unk&gt;]\"])\n</code></pre>\n<p>We run a simple small-scale experiment using <code>resnet34</code> and several <code>BPE</code> <code>tokenizers</code> using different <code>vocab</code> sizes. Below are results</p>\n<pre><code>42 character tokenzation - 9.7\n64 BPE tokenization - 8.9\n128 BPE tokenization - 8.01\n246 BPE tokenization - 6.9  \n</code></pre>\n<p>It clearly shows that more vocab size leads to better-learned embeddings. As per usual, in the end, we decided to use three different tokenizers to bring more diversity. <br>\n1) Y.nakama tokenizer<br>\n2) <code>BPE - 512</code><br>\n3) <code>BPE - 1024</code></p>\n<h2>Models</h2>\n<p>We decided to train several different models with everything mentioned above, with progressive pseudo labeling on test data (take only valid inchis). In total, we had the following modules<br><br>\n1)<code>effnetv2 -&gt; encoder(6) -&gt; decoder(8, 10, 12)</code><br>\n2)<code>effnet3 -&gt; encoder(3) -&gt; decoder(3)</code><br>\n3)<code>vit/cait/swin</code><br><br>\nFor vision transformers, we used hengs decoders.<br><br>\nFor hybrid CNN transformers to increase the diversity of our models, we used both <code>Encoder/Decoder</code> with attention layer with following modifications:<br><br>\n1) gating with <code>GLUE</code> (<a href=\"https://arxiv.org/abs/2002.05202)\" target=\"_blank\">https://arxiv.org/abs/2002.05202)</a>. In feed-forward layer, use <code>GLUE</code> in place of the first linear transformation and the activation function.<br><br>\n2) like 4th place solutions, we use relative positional embedding as in <code>T5</code><br></p>\n<h2>Ensembling:</h2>\n<p>We trained our model, everything looks great, but we needed to come with an ensembling strategy because everyone was using different <code>tokenizers</code>, so combining probabilities was not an option. Luckily <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> wrote a custom ensemble script that operates on string levels. It deserves its own post, which he will share later. Using <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> script, we were able to get from weakly scoring models <code>(0.74 - 0.90) ~0.67</code></p>\n<h2>Classification for m0/m1</h2>\n<p>We noticed that most of our valid predictions were different only by <code>m0/m1</code>. To solve this problem, we build a simple <code>resnet50</code> model classifying images into <code>3</code> classes (<code>no m0/m1</code>, <code>m0</code>,<code>m1</code>). Using this classifier to post-process our <code>InChI</code>, we gain a boost from <code>0.67</code> to <code>0.65</code>.</p>\n<h2>Finetuning model only for long InChIs.</h2>\n<p>We observed that our models tend to make a more wrong prediction for long <code>InChIs</code>. As a result last <code>4</code> days we decided to finetune our model only on long <code>InChIs</code> (bigger then <code>180</code>). After finetuning, we took the longest 100k <code>InChI</code> in our test and made new predictions. We also applied <code>TTA</code> to images and got those predictions as well. Once we have those predictions, we used a modified version of the zfturbo script to make an ensemble only on the longest <code>InChI</code>. This eventually brought us to our final <code>LB</code> standing <code>0.65 -&gt; 0.61</code>.</p>\n<h2>Conclusion</h2>\n<p>I would like to thank my teammates for being a great teamates and working really hard <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>, <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a> and new GM <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> =) <br>\nP.S also a big congratulations to all new GMs =) </p>",
  "messages": [
    {
      "id": "1336470",
      "postDate": "06/05/2021 00:26:00",
      "content": "<p>First of all, I would like to thanks the organizers for organizing such an exciting competition.  Second I would like to thank my teammates. We had a great working environment with everybody working hard, and last one week, we had to stay up until 3 am local times. Below I will try to summarize the main idea of our best-performing model.</p>\n<h2>Image Preprocessing</h2>\n<p>One of the challenging things about this competition was that the chemical had very different structures, some were super long, and others were very short. This also meant that images were highly diverse in size and shapes (e.g., super long in 1 axis or super small). Therefore it was essential to preprocess these images correctly. After some initial discussion we decided to try to use two methods:<br><br>\n<code>1) simple rescaling</code><br>\nplus - we will get all the information (kind off) <br>\nminus - feature of the images were become distorted (extreme shrinking across one axis)<br><br>\n<code>2) rescaling with preserving the aspect ratio</code><br>\nplus - This method will preserve the feature<br>\nminus  - bigger images, when downscaled, will have extremely small features <br><br>\nBoth of these methods have some advantages and disadvantages, but we reasoned that it would be best that each of our team members pick one of these methods and stick to it. We also found this an excellent article showing that most modern computer vision libraries employ very buggy resizing methods (<a href=\"https://arxiv.org/pdf/2104.11222.pdf)\" target=\"_blank\">https://arxiv.org/pdf/2104.11222.pdf)</a>. Since we are dealing with a lot of lines when downscaling, it was essential to preserve most of the information (tl dr: use PIL)</p>\n<h2>Image Augmentation.</h2>\n<p>We decided to go with natural image augmentations that were found in the dataset. But to make things a little bit diverse, we decided that everyone should train on a slightly different set of augmentations.<br><br>\nSet 1:<br>\n<code>RandomRotate90</code><br><br>\nSet 2</p>\n<pre><code>RandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomRotate90\nRandomNoiseAugment\n</code></pre>\n<p>Set 3</p>\n<pre><code>RandomRotate90\nRandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomNoiseAugment\nRandomLineRemoval\nRandomPiexlThreshold (some of the chemical lines had different sheds of black pixel values so we perform)\n</code></pre>\n<pre><code>thr = random.randint(200, 250)\nimg[img &lt; thr] = 0\nimg[img &gt;= thr] = 255\n</code></pre>\n<h2>Tokenizers</h2>\n<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> tokenizers was great, but we were wondering if there is a better way to tokenize. We found a lot of research on different tokenization strategies using <code>SMILES</code> but almost no information about <code>InChI</code>. After some thinking, we decided to try <code>huggingface</code> <code>BPE</code> tokenization. </p>\n<pre><code>vocab = 128 # can be any\ntokenizer = Tokenizer(BPE(unk_token=\"[&lt;unk&gt;]\", fuse_unk=False))\npre_tokenizer = pre_tokenizers.Sequence([Whitespace(), Digits()])\ntokenizer.pre_tokenizer = pre_tokenizer\ntrainer = BpeTrainer(vocab = max_len, special_tokens=[\"[&lt;pad&gt;]\", \"[&lt;start&gt;]\", \"[&lt;end&gt;]\", \"[&lt;unk&gt;]\"])\n</code></pre>\n<p>We run a simple small-scale experiment using <code>resnet34</code> and several <code>BPE</code> <code>tokenizers</code> using different <code>vocab</code> sizes. Below are results</p>\n<pre><code>42 character tokenzation - 9.7\n64 BPE tokenization - 8.9\n128 BPE tokenization - 8.01\n246 BPE tokenization - 6.9  \n</code></pre>\n<p>It clearly shows that more vocab size leads to better-learned embeddings. As per usual, in the end, we decided to use three different tokenizers to bring more diversity. <br>\n1) Y.nakama tokenizer<br>\n2) <code>BPE - 512</code><br>\n3) <code>BPE - 1024</code></p>\n<h2>Models</h2>\n<p>We decided to train several different models with everything mentioned above, with progressive pseudo labeling on test data (take only valid inchis). In total, we had the following modules<br><br>\n1)<code>effnetv2 -&gt; encoder(6) -&gt; decoder(8, 10, 12)</code><br>\n2)<code>effnet3 -&gt; encoder(3) -&gt; decoder(3)</code><br>\n3)<code>vit/cait/swin</code><br><br>\nFor vision transformers, we used hengs decoders.<br><br>\nFor hybrid CNN transformers to increase the diversity of our models, we used both <code>Encoder/Decoder</code> with attention layer with following modifications:<br><br>\n1) gating with <code>GLUE</code> (<a href=\"https://arxiv.org/abs/2002.05202)\" target=\"_blank\">https://arxiv.org/abs/2002.05202)</a>. In feed-forward layer, use <code>GLUE</code> in place of the first linear transformation and the activation function.<br><br>\n2) like 4th place solutions, we use relative positional embedding as in <code>T5</code><br></p>\n<h2>Ensembling:</h2>\n<p>We trained our model, everything looks great, but we needed to come with an ensembling strategy because everyone was using different <code>tokenizers</code>, so combining probabilities was not an option. Luckily <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> wrote a custom ensemble script that operates on string levels. It deserves its own post, which he will share later. Using <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> script, we were able to get from weakly scoring models <code>(0.74 - 0.90) ~0.67</code></p>\n<h2>Classification for m0/m1</h2>\n<p>We noticed that most of our valid predictions were different only by <code>m0/m1</code>. To solve this problem, we build a simple <code>resnet50</code> model classifying images into <code>3</code> classes (<code>no m0/m1</code>, <code>m0</code>,<code>m1</code>). Using this classifier to post-process our <code>InChI</code>, we gain a boost from <code>0.67</code> to <code>0.65</code>.</p>\n<h2>Finetuning model only for long InChIs.</h2>\n<p>We observed that our models tend to make a more wrong prediction for long <code>InChIs</code>. As a result last <code>4</code> days we decided to finetune our model only on long <code>InChIs</code> (bigger then <code>180</code>). After finetuning, we took the longest 100k <code>InChI</code> in our test and made new predictions. We also applied <code>TTA</code> to images and got those predictions as well. Once we have those predictions, we used a modified version of the zfturbo script to make an ensemble only on the longest <code>InChI</code>. This eventually brought us to our final <code>LB</code> standing <code>0.65 -&gt; 0.61</code>.</p>\n<h2>Conclusion</h2>\n<p>I would like to thank my teammates for being a great teamates and working really hard <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>, <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a> and new GM <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> =) <br>\nP.S also a big congratulations to all new GMs =) </p>",
      "rawMarkdown": "First of all, I would like to thanks the organizers for organizing such an exciting competition.  Second I would like to thank my teammates. We had a great working environment with everybody working hard, and last one week, we had to stay up until 3 am local times. Below I will try to summarize the main idea of our best-performing model.\n\n## Image Preprocessing\nOne of the challenging things about this competition was that the chemical had very different structures, some were super long, and others were very short. This also meant that images were highly diverse in size and shapes (e.g., super long in 1 axis or super small). Therefore it was essential to preprocess these images correctly. After some initial discussion we decided to try to use two methods:<br/>\n\n`1) simple rescaling `\nplus - we will get all the information (kind off) \nminus - feature of the images were become distorted (extreme shrinking across one axis)<br/>\n\n\n`2) rescaling with preserving the aspect ratio  `\nplus - This method will preserve the feature\nminus  - bigger images, when downscaled, will have extremely small features <br/>\n\nBoth of these methods have some advantages and disadvantages, but we reasoned that it would be best that each of our team members pick one of these methods and stick to it. We also found this an excellent article showing that most modern computer vision libraries employ very buggy resizing methods (https://arxiv.org/pdf/2104.11222.pdf). Since we are dealing with a lot of lines when downscaling, it was essential to preserve most of the information (tl dr: use PIL)\n\n## Image Augmentation. \nWe decided to go with natural image augmentations that were found in the dataset. But to make things a little bit diverse, we decided that everyone should train on a slightly different set of augmentations.<br/>\n\nSet 1:\n`RandomRotate90`<br/>\n\nSet 2\n```\nRandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomRotate90\nRandomNoiseAugment\n```\nSet 3\n```\nRandomRotate90\nRandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomNoiseAugment\nRandomLineRemoval\nRandomPiexlThreshold (some of the chemical lines had different sheds of black pixel values so we perform)\n```\n```\nthr = random.randint(200, 250)\nimg[img < thr] = 0\nimg[img >= thr] = 255\n```\n## Tokenizers \n@yasufuminakama tokenizers was great, but we were wondering if there is a better way to tokenize. We found a lot of research on different tokenization strategies using `SMILES` but almost no information about `InChI`. After some thinking, we decided to try `huggingface` `BPE` tokenization. \n\n```\nvocab = 128 # can be any\ntokenizer = Tokenizer(BPE(unk_token=\"[<unk>]\", fuse_unk=False))\npre_tokenizer = pre_tokenizers.Sequence([Whitespace(), Digits()])\ntokenizer.pre_tokenizer = pre_tokenizer\ntrainer = BpeTrainer(vocab = max_len, special_tokens=[\"[<pad>]\", \"[<start>]\", \"[<end>]\", \"[<unk>]\"])\n```\n\nWe run a simple small-scale experiment using `resnet34` and several `BPE` `tokenizers` using different `vocab` sizes. Below are results\n\n```\n42 character tokenzation - 9.7\n64 BPE tokenization - 8.9\n128 BPE tokenization - 8.01\n246 BPE tokenization - 6.9  \n```\nIt clearly shows that more vocab size leads to better-learned embeddings. As per usual, in the end, we decided to use three different tokenizers to bring more diversity. \n1) Y.nakama tokenizer\n2) `BPE - 512`\n3) `BPE - 1024`\n\n## Models\nWe decided to train several different models with everything mentioned above, with progressive pseudo labeling on test data (take only valid inchis). In total, we had the following modules<br/>\n1)`effnetv2 -> encoder(6) -> decoder(8, 10, 12)`\n2)`effnet3 -> encoder(3) -> decoder(3)`\n3)`vit/cait/swin`<br/>\n\nFor vision transformers, we used hengs decoders.<br/>\n\nFor hybrid CNN transformers to increase the diversity of our models, we used both `Encoder/Decoder` with attention layer with following modifications:<br/>\n1) gating with `GLUE` (https://arxiv.org/abs/2002.05202). In feed-forward layer, use `GLUE` in place of the first linear transformation and the activation function.<br/>\n2) like 4th place solutions, we use relative positional embedding as in `T5`<br/>\n\n## Ensembling:\nWe trained our model, everything looks great, but we needed to come with an ensembling strategy because everyone was using different `tokenizers`, so combining probabilities was not an option. Luckily @zfturbo wrote a custom ensemble script that operates on string levels. It deserves its own post, which he will share later. Using @zfturbo script, we were able to get from weakly scoring models `(0.74 - 0.90) ~0.67`\n\n## Classification for m0/m1\nWe noticed that most of our valid predictions were different only by `m0/m1`. To solve this problem, we build a simple `resnet50` model classifying images into `3` classes (`no m0/m1`, `m0`,` m1`). Using this classifier to post-process our `InChI`, we gain a boost from `0.67` to `0.65`.\n\n## Finetuning model only for long InChIs.\nWe observed that our models tend to make a more wrong prediction for long `InChIs`. As a result last `4` days we decided to finetune our model only on long `InChIs` (bigger then `180`). After finetuning, we took the longest 100k `InChI` in our test and made new predictions. We also applied `TTA` to images and got those predictions as well. Once we have those predictions, we used a modified version of the zfturbo script to make an ensemble only on the longest `InChI`. This eventually brought us to our final `LB` standing `0.65 -> 0.61`.\n\n\n## Conclusion\nI would like to thank my teammates for being a great teamates and working really hard @zfturbo, @youhanlee and new GM @tugstugi =) \n\nP.S also a big congratulations to all new GMs =)",
      "votes": null
    },
    {
      "id": "1336481",
      "postDate": "06/05/2021 00:49:48",
      "content": "<p>Thx for sharing.  BPE and other tricks learned.</p>",
      "rawMarkdown": "Thx for sharing.  BPE and other tricks learned.",
      "votes": null
    },
    {
      "id": "1336482",
      "postDate": "06/05/2021 00:51:12",
      "content": "<p>Great work, thanks for sharing your insights on the tokenisers. Looking forward to seeing the ensemble script from <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>. </p>",
      "rawMarkdown": "Great work, thanks for sharing your insights on the tokenisers. Looking forward to seeing the ensemble script from @zfturbo.",
      "votes": null
    },
    {
      "id": "1336487",
      "postDate": "06/05/2021 01:04:12",
      "content": "<p>\"We found a lot of research on different tokenization strategies using SMILES but almost no information about InChI. \"</p>\n<p>you can check this:<br>\n<a href=\"https://www.aclweb.org/anthology/2020.aacl-main.19.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/2020.aacl-main.19.pdf</a><br>\nTransformer-based Approach for Predicting Chemical Compound Structures</p>\n<p>\"For tokenizing InChI strings, the model was learned on SentencePiece (Kudo and Richardson, 2018), a unigram-based unsupervised training method for word segmentation. \"</p>",
      "rawMarkdown": "\"We found a lot of research on different tokenization strategies using SMILES but almost no information about InChI. \"\n\nyou can check this:\nhttps://www.aclweb.org/anthology/2020.aacl-main.19.pdf\nTransformer-based Approach for Predicting Chemical Compound Structures\n\n\"For tokenizing InChI strings, the model was learned on SentencePiece (Kudo and Richardson, 2018), a unigram-based unsupervised training method for word segmentation. \"",
      "votes": null
    },
    {
      "id": "1336489",
      "postDate": "06/05/2021 01:08:14",
      "content": "<p>Thank you, I missed this article! </p>",
      "rawMarkdown": "Thank you, I missed this article!",
      "votes": null
    },
    {
      "id": "1336549",
      "postDate": "06/05/2021 03:14:28",
      "content": "<p>Wonderful! Thanks for sharing. I am wondering how the \"RandomLineRemoval\" and \"RandomPixelThreshold\" augmentations are implemented.</p>",
      "rawMarkdown": "Wonderful! Thanks for sharing. I am wondering how the \"RandomLineRemoval\" and \"RandomPixelThreshold\" augmentations are implemented.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1336481,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "06/05/2021 00:49:48",
      "content": "<p>Thx for sharing.  BPE and other tricks learned.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336482,
      "author_name": "talktocharles",
      "author_url": "",
      "post_date": "06/05/2021 00:51:12",
      "content": "<p>Great work, thanks for sharing your insights on the tokenisers. Looking forward to seeing the ensemble script from <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a>. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336487,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/05/2021 01:04:12",
      "content": "<p>\"We found a lot of research on different tokenization strategies using SMILES but almost no information about InChI. \"</p>\n<p>you can check this:<br>\n<a href=\"https://www.aclweb.org/anthology/2020.aacl-main.19.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/2020.aacl-main.19.pdf</a><br>\nTransformer-based Approach for Predicting Chemical Compound Structures</p>\n<p>\"For tokenizing InChI strings, the model was learned on SentencePiece (Kudo and Richardson, 2018), a unigram-based unsupervised training method for word segmentation. \"</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336489,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "06/05/2021 01:08:14",
          "content": "<p>Thank you, I missed this article! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336549,
      "author_name": "houndcl",
      "author_url": "",
      "post_date": "06/05/2021 03:14:28",
      "content": "<p>Wonderful! Thanks for sharing. I am wondering how the \"RandomLineRemoval\" and \"RandomPixelThreshold\" augmentations are implemented.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1336470": "First of all, I would like to thanks the organizers for organizing such an exciting competition.  Second I would like to thank my teammates. We had a great working environment with everybody working hard, and last one week, we had to stay up until 3 am local times. Below I will try to summarize the main idea of our best-performing model.\n\n## Image Preprocessing\nOne of the challenging things about this competition was that the chemical had very different structures, some were super long, and others were very short. This also meant that images were highly diverse in size and shapes (e.g., super long in 1 axis or super small). Therefore it was essential to preprocess these images correctly. After some initial discussion we decided to try to use two methods:<br/>\n\n`1) simple rescaling `\nplus - we will get all the information (kind off) \nminus - feature of the images were become distorted (extreme shrinking across one axis)<br/>\n\n\n`2) rescaling with preserving the aspect ratio  `\nplus - This method will preserve the feature\nminus  - bigger images, when downscaled, will have extremely small features <br/>\n\nBoth of these methods have some advantages and disadvantages, but we reasoned that it would be best that each of our team members pick one of these methods and stick to it. We also found this an excellent article showing that most modern computer vision libraries employ very buggy resizing methods (https://arxiv.org/pdf/2104.11222.pdf). Since we are dealing with a lot of lines when downscaling, it was essential to preserve most of the information (tl dr: use PIL)\n\n## Image Augmentation. \nWe decided to go with natural image augmentations that were found in the dataset. But to make things a little bit diverse, we decided that everyone should train on a slightly different set of augmentations.<br/>\n\nSet 1:\n`RandomRotate90`<br/>\n\nSet 2\n```\nRandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomRotate90\nRandomNoiseAugment\n```\nSet 3\n```\nRandomRotate90\nRandomBorderCropPixel (cropping 5-20 pixel at random from the border of the image)\nRandomNoiseAugment\nRandomLineRemoval\nRandomPiexlThreshold (some of the chemical lines had different sheds of black pixel values so we perform)\n```\n```\nthr = random.randint(200, 250)\nimg[img < thr] = 0\nimg[img >= thr] = 255\n```\n## Tokenizers \n@yasufuminakama tokenizers was great, but we were wondering if there is a better way to tokenize. We found a lot of research on different tokenization strategies using `SMILES` but almost no information about `InChI`. After some thinking, we decided to try `huggingface` `BPE` tokenization. \n\n```\nvocab = 128 # can be any\ntokenizer = Tokenizer(BPE(unk_token=\"[<unk>]\", fuse_unk=False))\npre_tokenizer = pre_tokenizers.Sequence([Whitespace(), Digits()])\ntokenizer.pre_tokenizer = pre_tokenizer\ntrainer = BpeTrainer(vocab = max_len, special_tokens=[\"[<pad>]\", \"[<start>]\", \"[<end>]\", \"[<unk>]\"])\n```\n\nWe run a simple small-scale experiment using `resnet34` and several `BPE` `tokenizers` using different `vocab` sizes. Below are results\n\n```\n42 character tokenzation - 9.7\n64 BPE tokenization - 8.9\n128 BPE tokenization - 8.01\n246 BPE tokenization - 6.9  \n```\nIt clearly shows that more vocab size leads to better-learned embeddings. As per usual, in the end, we decided to use three different tokenizers to bring more diversity. \n1) Y.nakama tokenizer\n2) `BPE - 512`\n3) `BPE - 1024`\n\n## Models\nWe decided to train several different models with everything mentioned above, with progressive pseudo labeling on test data (take only valid inchis). In total, we had the following modules<br/>\n1)`effnetv2 -> encoder(6) -> decoder(8, 10, 12)`\n2)`effnet3 -> encoder(3) -> decoder(3)`\n3)`vit/cait/swin`<br/>\n\nFor vision transformers, we used hengs decoders.<br/>\n\nFor hybrid CNN transformers to increase the diversity of our models, we used both `Encoder/Decoder` with attention layer with following modifications:<br/>\n1) gating with `GLUE` (https://arxiv.org/abs/2002.05202). In feed-forward layer, use `GLUE` in place of the first linear transformation and the activation function.<br/>\n2) like 4th place solutions, we use relative positional embedding as in `T5`<br/>\n\n## Ensembling:\nWe trained our model, everything looks great, but we needed to come with an ensembling strategy because everyone was using different `tokenizers`, so combining probabilities was not an option. Luckily @zfturbo wrote a custom ensemble script that operates on string levels. It deserves its own post, which he will share later. Using @zfturbo script, we were able to get from weakly scoring models `(0.74 - 0.90) ~0.67`\n\n## Classification for m0/m1\nWe noticed that most of our valid predictions were different only by `m0/m1`. To solve this problem, we build a simple `resnet50` model classifying images into `3` classes (`no m0/m1`, `m0`,` m1`). Using this classifier to post-process our `InChI`, we gain a boost from `0.67` to `0.65`.\n\n## Finetuning model only for long InChIs.\nWe observed that our models tend to make a more wrong prediction for long `InChIs`. As a result last `4` days we decided to finetune our model only on long `InChIs` (bigger then `180`). After finetuning, we took the longest 100k `InChI` in our test and made new predictions. We also applied `TTA` to images and got those predictions as well. Once we have those predictions, we used a modified version of the zfturbo script to make an ensemble only on the longest `InChI`. This eventually brought us to our final `LB` standing `0.65 -> 0.61`.\n\n\n## Conclusion\nI would like to thank my teammates for being a great teamates and working really hard @zfturbo, @youhanlee and new GM @tugstugi =) \n\nP.S also a big congratulations to all new GMs =)",
    "1336481": "Thx for sharing.  BPE and other tricks learned.",
    "1336482": "Great work, thanks for sharing your insights on the tokenisers. Looking forward to seeing the ensemble script from @zfturbo.",
    "1336487": "\"We found a lot of research on different tokenization strategies using SMILES but almost no information about InChI. \"\n\nyou can check this:\nhttps://www.aclweb.org/anthology/2020.aacl-main.19.pdf\nTransformer-based Approach for Predicting Chemical Compound Structures\n\n\"For tokenizing InChI strings, the model was learned on SentencePiece (Kudo and Richardson, 2018), a unigram-based unsupervised training method for word segmentation. \"",
    "1336489": "Thank you, I missed this article!",
    "1336549": "Wonderful! Thanks for sharing. I am wondering how the \"RandomLineRemoval\" and \"RandomPixelThreshold\" augmentations are implemented."
  },
  "source": "meta"
}