{
  "id": 243820,
  "title": "50th Place Solution",
  "url": "/competitions/bms-molecular-translation/writeups/jarvislabs-ai-50th-place-solution",
  "author_name": "",
  "post_date": "2021-06-04T06:01:37.552286500Z",
  "votes": 19,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The source code is available in <a href=\"https://github.com/affjljoo3581/BMS-Molecular-Translation\" target=\"_blank\">the github repository</a>.</p>\n<p>Hi all. Thanks to the kaggle and competition host. It was a greate experience. We tried hard and we barely hold the silver medal 🤣.</p>\n<p>Our approach is to handle this competition as a translation problem, not image-captioning or OCR. Because the <strong>ViT</strong> model showed that the images can be treated like sentences, we focused on the translation task (from <strong>image</strong> language to <strong>InChI</strong> language). Actually, the model structure and training procedures are not pretty different to other approaches, but the detailed configurations (e.g. hyperparameter tuning) are based on the translation papers.</p>\n<p>Our basic strategy is to pretrain the transformer model with large-scale images from <code>extra_approved_InChIs.csv</code>, and then finetune the model with the original images.</p>\n<h2>Model Structure</h2>\n<p>We use <strong>ViT</strong> as an encoder and <strong>the original transformer</strong> as a decoder. According to <a href=\"https://arxiv.org/abs/1906.01787\" target=\"_blank\">this paper</a>, it is more effective to increase the encoder depth rather than the decoder one. So we make a base version of our model to have 12-layer encoder, 6-layer decoder and a large version to have 24-layer encoder.</p>\n<h2>Dataset</h2>\n<p>As I mentioned above, we create the external dataset from <code>extra_approved_InChIs.csv</code> by synthesize the molecular images through <strong>rdkit</strong>. The number of original images is about <strong>2.4M</strong> and the number of entire pretraining images is about <strong>12M</strong>. It is comparable to the ImageNet-21k, mention in ViT paper. The authors of ViT paper showed that the performance of the image transformer is highly related to the dataset scale. And the <strong>12M</strong> images are indeed helpful.</p>\n<h2>Training</h2>\n<p>We train the base model for about 15 epochs with 224x224 <strong>12M</strong> images. And we finetune with 384x384 <strong>2.4M</strong> images. The CVs are 0.99 and 0.83 and LBs are 1.82 and 1.55 respectively.</p>\n<p>And we increase the model size by copying the 12 layers and stacking to the end of the model. So the large version of our model has 24 layers. We pretrain the model for 5 epochs and finetune for 15 epochs. It achieves 0.74 CV and 1.44 LB.</p>\n<p>Basically, we use single A100 GPU on GCP. To accelerate the training, we use <a href=\"https://github.com/NVIDIA/apex\" target=\"_blank\">apex</a>'s fused layers and O2 amp mode.<br>\nThanks to <a href=\"https://jarvislabs.ai/\" target=\"_blank\">jarvislabs.ai</a>, we can use 8xRTX5000s. And interestingly, we found that the multi-GPUs environment with DDP (DataDistributedParallel) and <strong>apex amp</strong> do not work correctly. I don't know why it happens, but it is well-known problem in <strong>PyTorch Lightning</strong>. So in 8xRTX5000 environment, we use native AMP and <a href=\"https://github.com/microsoft/DeepSpeed\" target=\"_blank\">ZeRO</a> to increase the computing performance.</p>\n<h2>Prediction</h2>\n<p>We use <strong>greedy search</strong>, instead of <strong>beam search</strong>. That's because the inference time is really long (large model with 384x384 spent about 7 hours).<br>\nThanks to <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>, we applied <a href=\"https://www.kaggle.com/nofreewill/normalize-your-predictions\" target=\"_blank\">rdkit normalization</a> and the performance is indeed enhanced to 1.37.</p>\n<p>Lastly, specially thanks to <a href=\"https://jarvislabs.ai/\" target=\"_blank\">jarvislabs.ai</a> to support some credits 👍👍</p>",
  "messages": [
    {
      "id": "1335275",
      "postDate": "06/04/2021 06:01:37",
      "content": "<p>The source code is available in <a href=\"https://github.com/affjljoo3581/BMS-Molecular-Translation\" target=\"_blank\">the github repository</a>.</p>\n<p>Hi all. Thanks to the kaggle and competition host. It was a greate experience. We tried hard and we barely hold the silver medal 🤣.</p>\n<p>Our approach is to handle this competition as a translation problem, not image-captioning or OCR. Because the <strong>ViT</strong> model showed that the images can be treated like sentences, we focused on the translation task (from <strong>image</strong> language to <strong>InChI</strong> language). Actually, the model structure and training procedures are not pretty different to other approaches, but the detailed configurations (e.g. hyperparameter tuning) are based on the translation papers.</p>\n<p>Our basic strategy is to pretrain the transformer model with large-scale images from <code>extra_approved_InChIs.csv</code>, and then finetune the model with the original images.</p>\n<h2>Model Structure</h2>\n<p>We use <strong>ViT</strong> as an encoder and <strong>the original transformer</strong> as a decoder. According to <a href=\"https://arxiv.org/abs/1906.01787\" target=\"_blank\">this paper</a>, it is more effective to increase the encoder depth rather than the decoder one. So we make a base version of our model to have 12-layer encoder, 6-layer decoder and a large version to have 24-layer encoder.</p>\n<h2>Dataset</h2>\n<p>As I mentioned above, we create the external dataset from <code>extra_approved_InChIs.csv</code> by synthesize the molecular images through <strong>rdkit</strong>. The number of original images is about <strong>2.4M</strong> and the number of entire pretraining images is about <strong>12M</strong>. It is comparable to the ImageNet-21k, mention in ViT paper. The authors of ViT paper showed that the performance of the image transformer is highly related to the dataset scale. And the <strong>12M</strong> images are indeed helpful.</p>\n<h2>Training</h2>\n<p>We train the base model for about 15 epochs with 224x224 <strong>12M</strong> images. And we finetune with 384x384 <strong>2.4M</strong> images. The CVs are 0.99 and 0.83 and LBs are 1.82 and 1.55 respectively.</p>\n<p>And we increase the model size by copying the 12 layers and stacking to the end of the model. So the large version of our model has 24 layers. We pretrain the model for 5 epochs and finetune for 15 epochs. It achieves 0.74 CV and 1.44 LB.</p>\n<p>Basically, we use single A100 GPU on GCP. To accelerate the training, we use <a href=\"https://github.com/NVIDIA/apex\" target=\"_blank\">apex</a>'s fused layers and O2 amp mode.<br>\nThanks to <a href=\"https://jarvislabs.ai/\" target=\"_blank\">jarvislabs.ai</a>, we can use 8xRTX5000s. And interestingly, we found that the multi-GPUs environment with DDP (DataDistributedParallel) and <strong>apex amp</strong> do not work correctly. I don't know why it happens, but it is well-known problem in <strong>PyTorch Lightning</strong>. So in 8xRTX5000 environment, we use native AMP and <a href=\"https://github.com/microsoft/DeepSpeed\" target=\"_blank\">ZeRO</a> to increase the computing performance.</p>\n<h2>Prediction</h2>\n<p>We use <strong>greedy search</strong>, instead of <strong>beam search</strong>. That's because the inference time is really long (large model with 384x384 spent about 7 hours).<br>\nThanks to <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>, we applied <a href=\"https://www.kaggle.com/nofreewill/normalize-your-predictions\" target=\"_blank\">rdkit normalization</a> and the performance is indeed enhanced to 1.37.</p>\n<p>Lastly, specially thanks to <a href=\"https://jarvislabs.ai/\" target=\"_blank\">jarvislabs.ai</a> to support some credits 👍👍</p>",
      "rawMarkdown": "The source code is available in [the github repository](https://github.com/affjljoo3581/BMS-Molecular-Translation).\n\nHi all. Thanks to the kaggle and competition host. It was a greate experience. We tried hard and we barely hold the silver medal 🤣.\n\nOur approach is to handle this competition as a translation problem, not image-captioning or OCR. Because the **ViT** model showed that the images can be treated like sentences, we focused on the translation task (from **image** language to **InChI** language). Actually, the model structure and training procedures are not pretty different to other approaches, but the detailed configurations (e.g. hyperparameter tuning) are based on the translation papers.\n\nOur basic strategy is to pretrain the transformer model with large-scale images from `extra_approved_InChIs.csv`, and then finetune the model with the original images.\n\n## Model Structure\nWe use **ViT** as an encoder and **the original transformer** as a decoder. According to [this paper](https://arxiv.org/abs/1906.01787), it is more effective to increase the encoder depth rather than the decoder one. So we make a base version of our model to have 12-layer encoder, 6-layer decoder and a large version to have 24-layer encoder.\n\n## Dataset\nAs I mentioned above, we create the external dataset from `extra_approved_InChIs.csv` by synthesize the molecular images through **rdkit**. The number of original images is about **2.4M** and the number of entire pretraining images is about **12M**. It is comparable to the ImageNet-21k, mention in ViT paper. The authors of ViT paper showed that the performance of the image transformer is highly related to the dataset scale. And the **12M** images are indeed helpful.\n\n## Training\nWe train the base model for about 15 epochs with 224x224 **12M** images. And we finetune with 384x384 **2.4M** images. The CVs are 0.99 and 0.83 and LBs are 1.82 and 1.55 respectively.\n\nAnd we increase the model size by copying the 12 layers and stacking to the end of the model. So the large version of our model has 24 layers. We pretrain the model for 5 epochs and finetune for 15 epochs. It achieves 0.74 CV and 1.44 LB.\n\nBasically, we use single A100 GPU on GCP. To accelerate the training, we use [apex](https://github.com/NVIDIA/apex)'s fused layers and O2 amp mode.\nThanks to [jarvislabs.ai](https://jarvislabs.ai/), we can use 8xRTX5000s. And interestingly, we found that the multi-GPUs environment with DDP (DataDistributedParallel) and **apex amp** do not work correctly. I don't know why it happens, but it is well-known problem in **PyTorch Lightning**. So in 8xRTX5000 environment, we use native AMP and [ZeRO](https://github.com/microsoft/DeepSpeed) to increase the computing performance.\n\n## Prediction\nWe use **greedy search**, instead of **beam search**. That's because the inference time is really long (large model with 384x384 spent about 7 hours).\nThanks to @nofreewill, we applied [rdkit normalization](https://www.kaggle.com/nofreewill/normalize-your-predictions) and the performance is indeed enhanced to 1.37.\n\nLastly, specially thanks to [jarvislabs.ai](https://jarvislabs.ai/) to support some credits 👍👍",
      "votes": null
    },
    {
      "id": "1335293",
      "postDate": "06/04/2021 06:17:29",
      "content": "<p>You're welcome, congratulations on the silver medal! ^^</p>",
      "rawMarkdown": "You're welcome, congratulations on the silver medal! ^^",
      "votes": null
    },
    {
      "id": "1335297",
      "postDate": "06/04/2021 06:21:54",
      "content": "<p>Congratulations!!!</p>",
      "rawMarkdown": "Congratulations!!!",
      "votes": null
    },
    {
      "id": "1335520",
      "postDate": "06/04/2021 09:11:24",
      "content": "<p>Congratulations🎉🥳</p>",
      "rawMarkdown": "Congratulations🎉🥳",
      "votes": null
    },
    {
      "id": "1335994",
      "postDate": "06/04/2021 15:15:05",
      "content": "<p>Congrats!!!!</p>",
      "rawMarkdown": "Congrats!!!!",
      "votes": null
    },
    {
      "id": "1338469",
      "postDate": "06/06/2021 13:36:45",
      "content": "<p>congrat and thanks for the great write-up!</p>",
      "rawMarkdown": "congrat and thanks for the great write-up!",
      "votes": null
    },
    {
      "id": "1339352",
      "postDate": "06/07/2021 07:40:50",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335293,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "06/04/2021 06:17:29",
      "content": "<p>You're welcome, congratulations on the silver medal! ^^</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335297,
      "author_name": "ripunjoygoswami",
      "author_url": "",
      "post_date": "06/04/2021 06:21:54",
      "content": "<p>Congratulations!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335520,
      "author_name": "saurabhjejurkar",
      "author_url": "",
      "post_date": "06/04/2021 09:11:24",
      "content": "<p>Congratulations🎉🥳</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335994,
      "author_name": "shubhamzorc9",
      "author_url": "",
      "post_date": "06/04/2021 15:15:05",
      "content": "<p>Congrats!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1338469,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/06/2021 13:36:45",
      "content": "<p>congrat and thanks for the great write-up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1339352,
      "author_name": "leewook",
      "author_url": "",
      "post_date": "06/07/2021 07:40:50",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335275": "The source code is available in [the github repository](https://github.com/affjljoo3581/BMS-Molecular-Translation).\n\nHi all. Thanks to the kaggle and competition host. It was a greate experience. We tried hard and we barely hold the silver medal 🤣.\n\nOur approach is to handle this competition as a translation problem, not image-captioning or OCR. Because the **ViT** model showed that the images can be treated like sentences, we focused on the translation task (from **image** language to **InChI** language). Actually, the model structure and training procedures are not pretty different to other approaches, but the detailed configurations (e.g. hyperparameter tuning) are based on the translation papers.\n\nOur basic strategy is to pretrain the transformer model with large-scale images from `extra_approved_InChIs.csv`, and then finetune the model with the original images.\n\n## Model Structure\nWe use **ViT** as an encoder and **the original transformer** as a decoder. According to [this paper](https://arxiv.org/abs/1906.01787), it is more effective to increase the encoder depth rather than the decoder one. So we make a base version of our model to have 12-layer encoder, 6-layer decoder and a large version to have 24-layer encoder.\n\n## Dataset\nAs I mentioned above, we create the external dataset from `extra_approved_InChIs.csv` by synthesize the molecular images through **rdkit**. The number of original images is about **2.4M** and the number of entire pretraining images is about **12M**. It is comparable to the ImageNet-21k, mention in ViT paper. The authors of ViT paper showed that the performance of the image transformer is highly related to the dataset scale. And the **12M** images are indeed helpful.\n\n## Training\nWe train the base model for about 15 epochs with 224x224 **12M** images. And we finetune with 384x384 **2.4M** images. The CVs are 0.99 and 0.83 and LBs are 1.82 and 1.55 respectively.\n\nAnd we increase the model size by copying the 12 layers and stacking to the end of the model. So the large version of our model has 24 layers. We pretrain the model for 5 epochs and finetune for 15 epochs. It achieves 0.74 CV and 1.44 LB.\n\nBasically, we use single A100 GPU on GCP. To accelerate the training, we use [apex](https://github.com/NVIDIA/apex)'s fused layers and O2 amp mode.\nThanks to [jarvislabs.ai](https://jarvislabs.ai/), we can use 8xRTX5000s. And interestingly, we found that the multi-GPUs environment with DDP (DataDistributedParallel) and **apex amp** do not work correctly. I don't know why it happens, but it is well-known problem in **PyTorch Lightning**. So in 8xRTX5000 environment, we use native AMP and [ZeRO](https://github.com/microsoft/DeepSpeed) to increase the computing performance.\n\n## Prediction\nWe use **greedy search**, instead of **beam search**. That's because the inference time is really long (large model with 384x384 spent about 7 hours).\nThanks to @nofreewill, we applied [rdkit normalization](https://www.kaggle.com/nofreewill/normalize-your-predictions) and the performance is indeed enhanced to 1.37.\n\nLastly, specially thanks to [jarvislabs.ai](https://jarvislabs.ai/) to support some credits 👍👍",
    "1335293": "You're welcome, congratulations on the silver medal! ^^",
    "1335297": "Congratulations!!!",
    "1335520": "Congratulations🎉🥳",
    "1335994": "Congrats!!!!",
    "1338469": "congrat and thanks for the great write-up!",
    "1339352": "Congratulations!"
  },
  "source": "meta"
}